Key facts
| Request style | OpenAI-compatible chat completions |
| Built-in search | Sonar retrieves from the live web per request |
| Citations | Returned with the answer |
| Source control | Limited to filters and recency options |
| Pricing model | Usage-based with a per-search component |
| Embeddings | Not the focus of the Sonar API |
| Deployment | Vendor-hosted |
| Best at | Fresh web-grounded answers |
TL;DR
- Sonar is retrieval plus generation in one call; Plugsky is the model layer for your retrieval.
- Use Sonar when answers must reflect the live web with citations.
- Use Plugsky when answers must come from documents you control.
- Hybrid designs route open-web questions to Sonar and private questions to RAG.
- Judge both on citation accuracy, not answer fluency alone.
How it works, step by step
- Classify your questions: public web, private corpus, or both.
- Pilot Sonar on the public-web class and measure citation quality.
- Build a retrieval pipeline on Plugsky for the private-corpus class.
- Compare answers, citations and refusals on a shared golden question set.
- Define routing rules and keep both paths behind one interface.
- Monitor freshness, cost and accuracy as usage grows.
Try it yourself
Open the Perplexity Sonar API cost calculator →
Two architectures in one comparison
Sonar compresses a whole pipeline into a request: search, read, rank, synthesise and cite. That is valuable when the questions are about the public web and recency matters. The trade is that you inherit the vendor's index, ranking and coverage, and you pay a search component on top of token usage.
Plugsky sits one layer lower. It provides the chat model that writes the answer and the embedding model that finds the passages, while retrieval, storage and permissions stay in your application. That is more engineering, and it is also the only way to make citations map to documents you actually control.
Where each API is stronger
Choose Sonar for open-web research, news-style questions, competitive scans and anything where a missing result today is worse than imperfect control. Its cited answers are immediately useful, and the integration cost is one endpoint.
Choose Plugsky when the source of truth is internal: policies, tickets, contracts, product data. Embed your corpus with plugsky-embed or plugsky-embed-multilingual, filter retrieval by the requesting user's permissions, and instruct a chat model to answer only from retrieved context. Deployment can extend to VPC, on-prem or air-gapped, which matters when the corpus must not leave your environment; current plans are on the live pricing page.
A hybrid that stays testable
Most production systems end up with both paths. A router decides whether a question needs live web grounding or internal knowledge, and each path answers in a consistent format with citations. The routing decision is the risky part, so log it and review misroutes.
Keep one evaluation set that covers both classes and score citation accuracy, refusal behaviour and latency per path. That way the comparison never depends on intuition, and you can rebalance traffic as requirements or pricing change.
Honest comparison
| Dimension | Plugsky plus your retrieval | Sonar API | Decision signal |
|---|---|---|---|
| Source of truth | Your corpus | Live web | Internal or public knowledge |
| Freshness | As fresh as your index | Live per request | Recency requirement |
| Citation control | Full, with stored sources | Vendor-provided citations | Audit and traceability |
| Permissions | Enforced in retrieval | Not applicable | Multi-tenant access |
| Cost shape | Flat monthly plans | Usage-based with a search component | Volume and query mix |
| Deployment | Cloud, VPC, on-prem, air-gapped | Vendor-hosted | Residency requirements |
Frequently asked questions
Can Plugsky answer questions about current web content?
Not on its own. Plugsky has no built-in web index. Pair it with a search API, or build retrieval over a corpus that you refresh, when questions must reflect current information.
Which is more accurate for internal documents?
RAG on Plugsky, because retrieval is scoped to your corpus and permissions. Sonar is designed for the open web rather than private document sets.
Are both APIs OpenAI-compatible?
Both use the OpenAI chat completions shape in practice, so clients can call either with a base URL change. Verify response fields, especially citation metadata.
How do I evaluate the two fairly?
Use one golden set that includes answerable questions with known sources and unanswerable ones. Score whether citations contain the answer and whether the system refuses correctly.
Do I need embeddings if I use Sonar?
No. Sonar retrieves for you. If you also run internal RAG, you need an embedding model and a vector store for that path.
What about cost predictability?
Sonar usage includes a per-search component, so cost scales with queries. Plugsky self-serve plans are flat monthly with a free plan for plugsky-micro and plugsky-lite.
Can I route between Sonar and Plugsky automatically?
Yes. Classify each query by whether it needs live web grounding or internal knowledge, and keep the classification logic logged and testable.