Key facts
| Ingestion | Upload PDF, DOCX, TXT, MD and HTML; chunking and indexing are automatic |
| Retrieval | Keyword, vector and hybrid search with optional reranking |
| Citations | Every query returns ranked chunks with source attribution |
| Permissions | Scoped keys and per-collection isolation; RBAC and SSO on enterprise |
| Deployment | Managed, VPC, on-prem and air-gapped options |
| Models | 30+ models behind one OpenAI-compatible API |
| Free tier | plugsky-micro and plugsky-lite, no card required |
| Product status | Live |
TL;DR
- Ground every answer in published support content and show the source.
- Refresh the corpus on a schedule; stale policies produce wrong answers.
- Separate public help content from internal-only material by collection.
- Let the assistant refuse and hand off rather than guess.
- Measure deflection, escalation and answer accuracy together.
How it works, step by step
- Inventory the sources support agents actually use, and rate each for freshness.
- Separate public, internal and customer-specific content into different collections.
- Ingest with metadata such as product, version and last-reviewed date.
- Build a question set from real tickets with the correct article as the answer.
- Require citations and a visible handoff path when confidence is low.
- Monitor unanswered and escalated questions to find content gaps.
- Re-ingest on a schedule so updated articles replace outdated chunks.
Try it yourself
What support RAG is good at
Support work is repetitive and evidence-based: the same questions recur, and the correct answer usually exists in a help article, a macro or a policy document. That makes it a strong fit for retrieval-grounded answers. An assistant can surface the relevant passage, summarise it in the customer's context and link to the source, which reduces handling time and keeps answers consistent across agents.
It also scales onboarding. New agents can ask the assistant for the current process instead of searching across tools, and every answer points to the article it came from, so they learn the source of truth as they work.
Freshness and permissions are the hard parts
A support corpus changes constantly: prices, policies, product behaviour and workarounds. Stale chunks produce confidently wrong answers, which is worse than no answer. Give every document a last-reviewed date, re-ingest changed articles promptly, and retire content that no longer applies instead of leaving old versions in the index.
Permissions matter too. Draft policies, internal escalation notes and customer-specific history should not be retrievable by anyone. Keep public help content in a separate collection from internal material, use scoped keys for each surface, and verify with negative tests that a public query cannot reach internal chunks.
Designing the answer experience
The assistant should answer, cite and offer escalation. Show the source article so customers and agents can verify, keep the tone aligned with your support voice, and make the handoff to a human explicit when the corpus does not cover the question. A refusal with a link to contact support is a good outcome; an invented policy is not.
Keep a feedback signal on every answer, such as a helpful or not helpful action, and route negative feedback into a review queue. Those signals are the fastest way to find missing articles and poor chunking.
Implementing it on Plugsky
Upload help content to collections and let chunking, embedding and indexing run automatically. Query with keyword, vector or hybrid retrieval plus optional reranking, then pass the cited chunks to any of 30+ models on the same OpenAI-compatible API. Per-collection encryption and API-data-not-used-for-training commitments cover the privacy posture most support teams need.
Prototype with the RAG sandbox, then run it on the free plan with plugsky-micro and plugsky-lite or the 14-day full-access trial. Current plans are on the live pricing page.
Honest comparison
| Approach | RAG assistant with citations | Keyword help search | Ungrounded chatbot |
|---|---|---|---|
| Answer source | Retrieved help content | Matched articles | Model knowledge |
| Freshness | Depends on re-ingestion schedule | Depends on article updates | Unknown |
| Citations | Linked sources per answer | Article link | None |
| Escalation | Explicit when unsupported | User searches again | Model guesses |
| Internal content safety | Separate collections and scoped keys | Search permissions | No control |
Frequently asked questions
Can RAG answer from past support tickets?
Yes, if you ingest them as documents with metadata. Keep customer-specific tickets in restricted collections so they are not reachable from public surfaces.
How do I stop outdated answers?
Track a last-reviewed date on each document, re-ingest changes promptly, and remove retired articles. Stale chunks are the main cause of confidently wrong support answers.
Should the assistant answer without a source?
No. Require citations and an explicit handoff when the corpus does not support the answer. A clean refusal is better than an invented policy.
Can I keep internal material out of customer answers?
Yes. Separate public and internal content into different collections and scope keys so a customer-facing assistant cannot retrieve internal material.
Is there a free plan?
Yes. The free plan includes plugsky-micro and plugsky-lite with 2 API keys and no credit card, and a 14-day full-access trial is available.
How is pricing structured?
Self-serve plans are flat monthly with unlimited fair-use usage and no per-token charges or overage fees. See the live pricing page for current plans.