Key facts
| Router model | plugsky-fusion escalates per request across tiers (live) |
| Strategies | cost_saver for deflection, balanced for mixed queues, custom rules for sensitive intents |
| Models | 30+ models; cheap tiers suit FAQs, triage and status updates |
| Function calling | Live for order lookups, subscription changes and escalation triggers |
| Pricing | Flat monthly self-serve plans with no per-token charges on self-serve |
| Free tier | plugsky-micro and plugsky-lite on the free plan, no card required |
| Governance | Scoped keys, RBAC and audit logs for support actions |
| Roadmap | The moderation endpoint is coming soon; apply your own filters today |
TL;DR
- Deflect volume on cheap tiers; escalate hard tickets early.
- Measure cost per resolved ticket, not cost per message.
- Reopened tickets are the hidden cost of over-eager deflection.
- Pin security, billing and safety intents to strong models or humans.
- Review intent-level economics monthly.
How it works, step by step
- Classify ticket intents and measure current volume, resolution and handle time per intent.
- Route FAQ, status and how-to intents to cheap tiers with knowledge retrieval.
- Escalate security, billing disputes, cancellations and negative sentiment to strong models or humans.
- Add tools for lookups and simple actions with server-side permission checks.
- Track cost per resolved ticket, reopen rate and escalation rate per intent.
- Move one intent at a time to a cheaper tier and watch quality metrics for a full cycle.
- Review thresholds monthly and document every routing change.
Try it yourself
Volume, repetition and escalation
Support traffic splits into two very different populations. The majority is repetitive: password resets, order status, how-to questions, policy confirmation. A minority is high-stakes: disputes, security concerns, cancellations, distressed customers. Serving both with the same model overpays for the majority and under-serves the minority.
Routing fixes the mismatch. Deflection and triage run on cheap tiers with retrieval over your help centre; sensitive intents escalate immediately to strong models with a human in the loop. The saving comes from volume, not from downgrading the cases that generate complaints.
Deflection without quality loss
Deflection is only a saving when the ticket stays closed. Optimising for containment pushes customers to repeat themselves, reopen tickets or churn — costs that never appear in token dashboards. Track reopen rate and repeat-contact rate per intent, and treat a rise as a signal to route that intent back up.
- Ground answers in retrieved knowledge and cite the source.
- Ask a clarifying question instead of guessing when confidence is low.
- Offer a human when policy or emotion demands it.
- Keep the transcript and tool calls for audit and for training reviewers.
Measuring support economics
Cost per resolved ticket is the honest metric: model usage plus retrieval plus human escalation, divided by tickets that stay closed. Compare it across intents to find where automation pays and where it only looks cheap. A ticket that takes three automated replies and a human will cost more than an immediate handoff.
Flat self-serve plans make the AI side of that equation stable — see the live pricing page — so the variable is human time and escalation policy, not token prices. Start on the free plan with plugsky-micro and plugsky-lite, then use the 14-day full-access trial to test stronger models on the intents that currently escalate.
Honest comparison
| Support intent | Routed support on Plugsky | One strong model for all | Rules-based deflection |
|---|---|---|---|
| FAQ and how-to | Cheap tier with retrieval | Strong tier for repetition | Fixed scripts |
| Billing disputes | Escalated quickly | Native strength | Often mishandled |
| Status and lookups | Cheap tier plus tools | Expensive per lookup | Works if integrated |
| Sentiment handling | Esclation thresholds | Native strength | Not detected |
| Economics | Cost per resolved ticket | Blended token spend | Containment rate only |
Frequently asked questions
What does support automation actually save?
The saving is real only when tickets stay resolved. Track cost per resolved ticket and reopen rate per intent instead of containment, which can hide frustrated customers repeating themselves.
Which tickets should escalate immediately?
Security concerns, billing disputes, cancellations, legal threats and distressed customers. These benefit most from strong models and human judgement, and are rare enough that their cost is marginal.
How do I use retrieval in support?
Index your help centre and resolved tickets, retrieve for the customer's question, and require answers to cite a source. Grounded answers reduce both wrong information and unnecessary escalations.
Can cheap models handle support tone?
For routine, factual replies yes, especially with a well-written system prompt and retrieval. Evaluate tone per intent on your own conversations before moving an intent down a tier.
What tools should a support agent have?
Read tools such as order and subscription lookup can run automatically; write tools such as refunds and plan changes need permission checks, limits and sometimes human approval.
Is there a moderation endpoint?
Not yet — it is coming soon. Apply your own input and output filters today, and log safety events for review.
How does flat pricing help support teams?
Self-serve plans are flat monthly with no per-token charges, so deflection experiments and prompt tuning do not create a variable bill. See the live pricing page for plan details.
Can I test routing for free?
Yes. plugsky-micro and plugsky-lite are on the free plan with no card, and the 14-day full-access trial lets you evaluate stronger models on escalated intents.