Key facts
| Strategies | Cost saver, Balanced, Max quality and Custom |
| Default | Balanced routes most calls to a mid-tier model |
| Signals | Token count, tool-call presence, prompt heuristics, user tier and time of day |
| Custom rules | Ordered rules where the first match wins, logged per request |
| Override | Calling a specific model name bypasses routing |
| Latency impact | The router runs in the API gateway |
| Observability | Every log records the chosen model, strategy and matched rule |
| Product status | Live |
TL;DR
- Route easy calls to cheap models and hard calls to strong ones automatically.
- Balanced is a sane default; Cost saver and Max quality are one switch away.
- Custom rules can key off prompt content, user plan or time of day.
- Per-request model names always override routing.
- Routing decisions are logged for audit and debugging.
How it works, step by step
- Instrument your application so requests carry user or workload context.
- Start with Balanced and compare quality against your current single model.
- Move simple workloads to Cost saver and measure the change.
- Add custom rules for known patterns such as refactors or long inputs.
- Keep an explicit fallback chain for outages and rate limits.
- Review routing logs weekly and tune rules that misfire.
Try it yourself
Open the workload router simulator →
Why model routing
Most production AI systems route every request to a single model, which means over-paying for simple calls and under-serving hard ones. Model routing fixes the mismatch: cheap models answer the routine majority, strong models handle the demanding minority, and simple calls return faster because they are not queued behind a reasoning model.
As a resilience benefit, routing also gives you a natural place to define fallbacks. If one model is unavailable, traffic shifts to a peer profile instead of failing.
Routing strategies
Four strategies are available and can be mixed per workspace. Cost saver uses the cheapest model that can handle the call, classified by complexity. Balanced, the default, uses a mid-tier model for most calls with cheap and strong fallbacks. Max quality always uses the strongest model in your tier for work where quality dominates, such as code generation or financial analysis.
Custom lets you define your own rules when your traffic has patterns the presets do not capture.
Custom rules and logged decisions
Custom rules are evaluated in order and the first match wins. Examples that hold up in production: send free-plan users to plugsky-micro; send requests above a token threshold to a long-context model; send prompts containing code-review keywords to a stronger model; and vary the model by time of day to match load.
Every request log records the chosen model, the strategy and the rule that fired, so you can audit cost decisions and debug surprising quality changes. Use the router simulator to test rule sets before shipping them.
Routing also makes a 30+ model catalogue practical without hard-coding a model per feature.Honest comparison
| Strategy | Model choice | Best for | Trade-off |
|---|---|---|---|
| Cost saver | Cheapest capable model | High-volume simple calls | Some hard prompts need retries |
| Balanced | Mid-tier with cheap and strong fallbacks | Most production apps | Not optimal at either extreme |
| Max quality | Strongest model in the tier | Code, legal and financial analysis | Higher compute per call |
| Custom | Your ordered rules | Known traffic patterns | Requires maintenance |
| Per-request override | A named model | Testing and special cases | Bypasses cost controls |
Frequently asked questions
How does Plugsky know which model to use?
For preset strategies it classifies each request by token count, tool-call presence and prompt heuristics. For custom routing, you define the rules.
Can I override routing per request?
Yes. Passing a specific model name such as plugsky-pro bypasses routing for that request.
Does routing add latency?
The router runs in the API gateway, so the added latency is negligible compared with inference time.
Can I see routing decisions in logs?
Yes. Every request log includes the chosen model, the strategy and, for custom routing, the rule that fired.
Which strategy should I start with?
Start with Balanced, compare quality against your current single-model setup, then move specific workloads to Cost saver or Max quality.
Does routing work with fallbacks?
Yes. Combine routing with a fallback chain so an unavailable model shifts traffic to a same-profile peer.
Plugsky (2026). “Model Routing: Auto-Pick the Right Model”. Plugsky. Available at: https://plugsky.com/articles/model-routing (last updated 2026-09-25).