Feature × Audience

How do MSPs use Model Fusion with Plugsky?

MSPs use Plugsky Model Fusion to keep AI service delivery profitable: most client traffic runs through cheap-first chains on plugsky-micro and plugsky-lite, hard tickets escalate automatically to stronger tiers, and each client gets its own workspace strategy and keys. Flat pricing means usage growth does not become variable cost, and fusion's logs show what actually ran.

Key facts

Router modelmodel="plugsky-fusion" applies a per-client chain (live)
Cheap-first chainDefault escalation runs micro to pro to max, keeping routine volume cheap
Per-client controlStrategies scoped per workspace or API key, one project per client
Cost modelFlat self-serve plans; no per-token charges, fair-use RPM per tier
White-labelDeliver under your own brand across all clients
ObservabilityModel, strategy and rule logged per request; tokens returned in every response
ResidencyPin each client workspace to the region their contract requires
RoadmapClassifier routing /v1/plugsky/route (model=auto) is coming soon

TL;DR

  • Default every client to a cheap-first chain with automatic escalation.
  • Give each client a project, strategy and scoped keys of their own.
  • Use flat pricing to keep margins stable as usage grows.
  • Show clients routing logs as proof of where their requests ran.
  • Escalate selectively — premium models only where tickets are hard.

How it works, step by step

  1. Standardise a single fusion chain and tune it once, instead of maintaining a model stack per client.
  2. Create one Plugsky project and key set per client, each with its own default strategy.
  3. Route triage and routine support to cheap tiers, and let the chain escalate complex tickets automatically.
  4. Set premium strategies only for clients or queues that pay for them.
  5. Pin each workspace to the contractually required region and configure retention accordingly.
  6. Report per-client routing and token telemetry through your portal to demonstrate value and control.
1Standardise asingle fusion chainand tune it once,2Create one Plugskyproject and key setper client, each3Route triage androutine support tocheap tiers, and4Set premiumstrategies only forclients or queues5Pin each workspaceto thecontractually6Report per-clientrouting and tokentelemetry through

Try it yourself

Open the LLM cost calculator →

Cheap-first economics, per client

An MSP's AI margin is decided by what runs on premium models. With model: "plugsky-fusion", every client workspace can run a cheap-first chain: routine classification, summarisation and answer drafting go to plugsky-micro or plugsky-lite, and only hard or ambiguous tickets escalate up the chain.

Because escalation is automatic, you do not need humans deciding model quality in real time. You tune the chain once, set a strategy per client, and the router handles the rest. Flat self-serve plans remove per-token charges, so more successful resolutions do not shrink the margin on that client.

Per-client isolation and control

A shared blob of AI usage is unmanageable at scale, and clients notice. Give each client a project with scoped keys, its own retrieval collection and its own default strategy. Onboarding becomes a checklist, and offboarding is a key revocation rather than an archaeology project.

  • Residency: pin each workspace to the region the contract requires — GCC, EU, US or APAC.
  • Overrides: if a client requires a fixed model for compliance, pin it on their key.
  • Audit: key and inference events are exportable to your SIEM for incident reviews.

Demonstrating value without per-token billing

Clients want to see value, not model names. Fusion logs give you the story: how many requests ran on fast, cheap models, how many escalated, and how often routing changed the outcome. Token usage returns in every response for internal reporting even though it does not drive billing.

Deliver the whole service white-label, expose only your portal, and keep Plugsky behind the API. When a client asks how costs are controlled, show the routing distribution rather than an invoice breakdown that would leak your unit economics.

Honest comparison

ConcernPlugsky Model FusionReselling per-token APIsSelf-hosted ensemble
Client unit costCheap-first chains, flat plansPer-token, variableGPU plus ops
EscalationAutomatic on hard promptsManual or noneCustom logic
Multi-tenancyProject and key per clientAccount per clientYou build isolation
White-labelFull brand controlProvider branding leaksFully yours
Ops burdenConfiguration onlyBilling reconciliationYou run everything

Frequently asked questions

How does fusion protect margin?

Most traffic stays on cheap tiers by default and escalates only when a prompt is hard. Flat self-serve plans mean no per-token charges, so volume growth does not become variable cost.

Can each client have a different chain?

Yes. Strategies can be set per workspace or key, so each client can run its own default, premium threshold and override rules.

Is client data kept separate?

Projects and scoped keys separate retrieval, logs and limits. Keys are rotatable and revocable without downtime.

Can we white-label the service?

Yes. Deliver chatbots, portals and reports under your own brand while Plugsky stays behind the API.

How do we prove value to clients?

Share routing summaries and resolution metrics from your portal. Logs record model, strategy and rule for every request, and tokens return in each response.

What about clients with residency requirements?

Pin each workspace to the required region, or use VPC, on-prem and air-gapped deployment for stricter contracts.

How should we start?

Run one client on a balanced or cost_saver strategy, measure quality and routing distribution for two weeks, then template that setup across your book.