Key facts
| Grounding | No web-search grounding endpoint today; RAG over your own corpus is live |
| API compatibility | OpenAI-compatible /v1/chat/completions (drop-in base URL change) |
| Models | 30+ models in one catalogue, from free tiers to frontier reasoning |
| Embeddings | Embeddings API is live for retrieval over private corpora |
| Pricing | Flat monthly self-serve plans with unlimited fair-use usage |
| Free tier | plugsky-micro and plugsky-lite, no card required |
| Deployment | Plugsky cloud, your VPC, on-prem and air-gapped |
TL;DR
- Sonar grounds answers in live web results with citations; Plugsky does not.
- Plugsky grounds answers in your own corpus via embeddings and RAG.
- Use Sonar for fresh public facts and Plugsky for private knowledge.
- Both expose OpenAI-style chat, so routing between them is simple.
- Do not migrate web-grounded answers to a private model without a search layer.
How it works, step by step
- Split your questions into live-web and private-knowledge categories.
- Keep Sonar for questions that require current public information and citations.
- For private knowledge, build retrieval with Plugsky embeddings over your corpus.
- Return citations from your own retrieved chunks so answers stay verifiable.
- Route by intent in one service so callers do not choose providers.
- Evaluate answer correctness and citation quality per category.
Try it yourself
Open the Perplexity Sonar API cost calculator →
What Sonar does well
Sonar connects a language model to live search results and returns answers with citations. For news, market facts, recent releases and anything that changes daily, that grounding is the feature: the model is only as good as the freshness of its inputs.
It also sets expectations. Answers are tied to public web sources, not your private documents, and citation quality depends on retrieval quality. If your questions are mostly about internal policies, tickets or contracts, web grounding is the wrong tool.
Private RAG versus open-web grounding
Private RAG grounds answers in your data. You chunk and embed documents with the embeddings API, retrieve top passages per query, and instruct the model to answer only from that context. Plugsky supports this pattern with live chat, streaming and embeddings endpoints.
- Control what the model can see and cite.
- Keep sensitive documents inside your own data plane.
- Tune retrieval quality independently of generation.
- Audit answers against stored chunks.
- Log which sources were retrieved so audits can reproduce an answer.
- Refresh indexes on a schedule so private answers stay current.
Hybrid routing and honest limits
The practical architecture routes by intent: live-web questions to Sonar, private or general questions to Plugsky, with a shared response contract so callers see one interface. Both APIs are OpenAI-compatible, so the routing service stays small.
Be clear about the gap: web-search grounding is not available on Plugsky today. If fresh citations are essential, keep Sonar. Deployment options, flat monthly self-serve pricing and a free tier are where Plugsky adds value. See the live pricing page for current plans.
Honest comparison
| Capability | Plugsky | Perplexity Sonar API | Self-hosted search plus LLM |
|---|---|---|---|
| Web grounding | Not available today | Core feature with citations | You build the pipeline |
| Private RAG | Embeddings and RAG are live | Not the focus | You build the pipeline |
| API style | OpenAI-compatible | OpenAI-compatible | Varies |
| Pricing shape | Flat monthly self-serve | Usage-based | GPU plus ops cost |
| Deployment | Cloud, VPC, on-prem, air-gapped | Managed API | Your infrastructure |
Frequently asked questions
Can Plugsky replace Perplexity Sonar?
Not for live-web answers with citations. Plugsky has no web-search grounding endpoint today; it covers private RAG and general chat, so the two work best together.
How do I build private RAG with Plugsky?
Embed your corpus with the live embeddings API, retrieve top passages per query, and generate answers constrained to those passages with the chat completions endpoint.
Is there a free plan?
Yes. plugsky-micro and plugsky-lite are free with no credit card, and a 14-day full-access trial opens stronger models.
How is pricing structured?
Self-serve plans are flat monthly with unlimited fair-use usage and no per-token billing. See the live pricing page for current plans.
Can I route between Sonar and Plugsky?
Yes. Both expose OpenAI-style chat endpoints, so an intent router can send live-web questions one way and private questions the other.
What about citations from my own documents?
Return the retrieved chunks with their source metadata in your response layer. Grounding quality depends on retrieval, so evaluate it separately from generation: treat freshness, recall and citation accuracy as separate metrics so you know which layer to fix.