Developer + API

What is model routing on Plugsky?

Model routing picks the right model for each request instead of sending everything to one endpoint. Plugsky classifies a request using signals such as token count, tool calls and prompt heuristics, then applies a strategy: Cost saver, Balanced, Max quality or custom rules. Strategies apply per workspace, per API key or per request, and every decision is logged.

Key facts

StrategiesCost saver, Balanced, Max quality and Custom
DefaultBalanced routes most calls to a mid-tier model
SignalsToken count, tool-call presence, prompt heuristics, user tier and time of day
Custom rulesOrdered rules where the first match wins, logged per request
OverrideCalling a specific model name bypasses routing
Latency impactThe router runs in the API gateway
ObservabilityEvery log records the chosen model, strategy and matched rule
Product statusLive

TL;DR

  • Route easy calls to cheap models and hard calls to strong ones automatically.
  • Balanced is a sane default; Cost saver and Max quality are one switch away.
  • Custom rules can key off prompt content, user plan or time of day.
  • Per-request model names always override routing.
  • Routing decisions are logged for audit and debugging.

How it works, step by step

  1. Instrument your application so requests carry user or workload context.
  2. Start with Balanced and compare quality against your current single model.
  3. Move simple workloads to Cost saver and measure the change.
  4. Add custom rules for known patterns such as refactors or long inputs.
  5. Keep an explicit fallback chain for outages and rate limits.
  6. Review routing logs weekly and tune rules that misfire.
1Instrument yourapplication sorequests carry user2Start with Balancedand compare qualityagainst your3Move simpleworkloads to Costsaver and measure4Add custom rulesfor known patternssuch as refactors5Keep an explicitfallback chain foroutages and rate6Review routing logsweekly and tunerules that misfire.

Try it yourself

Open the workload router simulator →

Why model routing

Most production AI systems route every request to a single model, which means over-paying for simple calls and under-serving hard ones. Model routing fixes the mismatch: cheap models answer the routine majority, strong models handle the demanding minority, and simple calls return faster because they are not queued behind a reasoning model.

As a resilience benefit, routing also gives you a natural place to define fallbacks. If one model is unavailable, traffic shifts to a peer profile instead of failing.

Routing strategies

Four strategies are available and can be mixed per workspace. Cost saver uses the cheapest model that can handle the call, classified by complexity. Balanced, the default, uses a mid-tier model for most calls with cheap and strong fallbacks. Max quality always uses the strongest model in your tier for work where quality dominates, such as code generation or financial analysis.

Custom lets you define your own rules when your traffic has patterns the presets do not capture.

Custom rules and logged decisions

Custom rules are evaluated in order and the first match wins. Examples that hold up in production: send free-plan users to plugsky-micro; send requests above a token threshold to a long-context model; send prompts containing code-review keywords to a stronger model; and vary the model by time of day to match load.

Every request log records the chosen model, the strategy and the rule that fired, so you can audit cost decisions and debug surprising quality changes. Use the router simulator to test rule sets before shipping them.

Routing also makes a 30+ model catalogue practical without hard-coding a model per feature.

Honest comparison

StrategyModel choiceBest forTrade-off
Cost saverCheapest capable modelHigh-volume simple callsSome hard prompts need retries
BalancedMid-tier with cheap and strong fallbacksMost production appsNot optimal at either extreme
Max qualityStrongest model in the tierCode, legal and financial analysisHigher compute per call
CustomYour ordered rulesKnown traffic patternsRequires maintenance
Per-request overrideA named modelTesting and special casesBypasses cost controls

Frequently asked questions

How does Plugsky know which model to use?

For preset strategies it classifies each request by token count, tool-call presence and prompt heuristics. For custom routing, you define the rules.

Can I override routing per request?

Yes. Passing a specific model name such as plugsky-pro bypasses routing for that request.

Does routing add latency?

The router runs in the API gateway, so the added latency is negligible compared with inference time.

Can I see routing decisions in logs?

Yes. Every request log includes the chosen model, the strategy and, for custom routing, the rule that fired.

Which strategy should I start with?

Start with Balanced, compare quality against your current single-model setup, then move specific workloads to Cost saver or Max quality.

Does routing work with fallbacks?

Yes. Combine routing with a fallback chain so an unavailable model shifts traffic to a same-profile peer.

Cite this page

Plugsky (2026). “Model Routing: Auto-Pick the Right Model”. Plugsky. Available at: https://plugsky.com/articles/model-routing (last updated 2026-09-25).