Key facts
| Core tools | Knowledge base retrieval, ticket read, tagging, routing, reply drafting |
| Grounding | Answers cite knowledge base articles with versions |
| Boundaries | Refunds, account changes and commitments require approval |
| Escalation | Sentiment, repeated contact and policy triggers hand off to a human |
| Quality metrics | Containment, citation accuracy, escalation quality, draft edit rate |
| Privacy | Redact personal data in logs and scope retrieval per customer |
| Models | Small models for triage, frontier for hard cases, on one API key |
| Status | Function calling, streaming, embeddings and RAG are live |
TL;DR
- Start with triage and drafted replies; let humans send the first hundred.
- Ground every answer in a cited article and admit gaps instead of improvising.
- Escalate on sentiment, repetition and policy — not only on confidence.
- Measure edit rate and escalation quality, not just deflection.
- Keep account changes and refunds behind explicit approval.
How it works, step by step
- Clean the knowledge base first; retrieval cannot fix contradictory articles.
- Connect ticket read and tagging tools so the agent sees real customer context.
- Expose retrieval as a tool with citations and version metadata.
- Draft replies for agent review before enabling any automatic sending.
- Add routing and escalation rules with clear triggers and an SLA.
- Track containment, citation accuracy, edit rate and resolution time weekly.
- Expand automation only in the categories where quality metrics hold.
Try it yourself
What a support agent should own
The safest valuable scope is read, understand, draft and route. The agent reads the ticket and account history, finds the right articles, classifies the issue and urgency, drafts a response with citations, and routes it to the right queue. That already removes most handling time while keeping the human in control of what customers receive.
As quality proves out, expand to low-risk actions: tagging, updating ticket fields, sending password-reset links, or answering in categories with short, verifiable answers. Anything with financial or contractual consequence stays human-approved.
The workflow: retrieve, decide, draft, escalate
- Retrieve: hybrid search over help centre articles, internal macros and past resolutions, filtered by product and region.
- Decide: classify intent, urgency and whether the issue is answerable from documentation.
- Draft: produce a reply that cites its sources and states any assumptions.
- Escalate: hand off on sentiment shifts, repeat contacts, legal or security keywords, or low retrieval confidence.
Escalation should carry context forward so customers do not repeat themselves. Summarise the thread, the attempted steps and what the human needs to check.
Measuring support quality honestly
Deflection rate is a vanity metric on its own: an agent that refuses to help inflates it. Track containment with resolution confirmed, citation accuracy on sampled replies, escalation appropriateness and the edit rate reviewers apply to drafts. A falling edit rate over time is the clearest signal that quality is improving.
Privacy is part of quality. Scope retrieval to the requesting customer and tenant, redact personal data in logs, and set retention for conversation content. On Plugsky, embeddings, RAG, function calling and streaming are live on the OpenAI-compatible API, with 30+ models on one key so triage runs on a small model and hard cases on a frontier one. Scoped keys, RBAC and audit logging cover review requirements. Plans are on the live pricing page; files and batch endpoints are coming soon.
Honest comparison
| Stage | Agent role | Automation level | Metric |
|---|---|---|---|
| Triage | Classify intent and urgency | Automatic | Routing accuracy |
| Answering | Draft cited replies | Review then send | Edit rate, citation accuracy |
| Low-risk actions | Tag, update fields, send links | Automatic with limits | Error rate, rework |
| Money or account changes | Prepare the action | Human approval | Approval latency, reversals |
| Escalation | Summarise and hand off | Automatic | Repeat contact rate |
Frequently asked questions
Will an AI support agent replace my team?
In most organisations it reduces handling time per ticket rather than headcount. Triage, drafting and routing are automated; judgement, escalations and sensitive conversations remain human.
How do I stop it hallucinating policy?
Ground every answer in retrieved knowledge base articles, require citations, and make the agent say when documentation does not cover the case. Contradictory articles must be cleaned up first.
Should the agent send replies automatically?
Not initially. Review drafts for the first weeks, measure the edit rate, and enable automatic sending only in categories where accuracy is consistently high.
How does escalation work?
Escalate on sentiment, repeated contact, policy keywords, low retrieval confidence or customer request. Pass a summary so the human starts with full context.
How do I handle personal data?
Scope retrieval per customer, redact personal data from logs, set retention limits and use region-locked deployment where regulation requires it.
Which metrics matter most?
Containment with confirmed resolution, citation accuracy, escalation appropriateness and reviewer edit rate. Track them together so improvement in one is not hiding regression in another.
Can it work in multiple languages?
Yes, with a multilingual model and a knowledge base in the customer's language. Validate quality per language rather than assuming parity.