Key facts
| API compatibility | OpenAI-compatible /v1/chat/completions (change the base URL) |
| Models | 30+ models from free to frontier tiers behind one API |
| Agent primitives | Function calling, JSON mode and streaming are live |
| Retrieval | Embeddings and RAG over your own corpus |
| Deployment | Plugsky cloud, VPC, on-prem or air-gapped |
| Pricing | Flat monthly self-serve plans with fair-use usage; see the live pricing page |
| Long context | Long-context models available for large excerpts |
| Sensitive data | On-prem and air-gapped deployment options |
TL;DR
- Keep your OpenAI SDK — change the base URL and model name.
- 30+ models behind one API, from free chat models to frontier reasoning.
- Deployment options from hosted cloud to VPC, on-prem and air-gapped.
- Ground synthesis in retrieval and verify every citation.
- Match the deployment tier to data sensitivity per project.
How it works, step by step
- Define the job, the permitted data sources and where a human must approve.
- Curate a clean corpus and pilot a literature review.
- Create a Plugsky account and generate an API key (free plan, no card required).
- Point your OpenAI SDK at the Plugsky base URL and map your model names.
- Index the approved corpus with embeddings and keep retrieval role-scoped.
- Verify citation accuracy before trusting synthesis.
- Measure quality on your own samples, then scale with usage monitoring.
Try it yourself
Where AI agents pay off in research teams
Research teams do not lack ideas for agents; they lack a safe path from demo to production. The pattern below targets repetitive, document-heavy work where a human can check the output, which is where agents earn their place first. Treat the agent as a new team member with a narrow brief, explicit permissions and a probation period, and rollout becomes an operations exercise rather than a leap of faith.
- Literature retrieval — find and summarise relevant work with citations
- Synthesis — draft structured reviews that link to source passages
- Analysis support — generate and explain code for data preparation
- Project memory — answer questions across notes, protocols and prior results
A reference architecture for research teams agents
A retrieval agent returns sourced passages, a synthesis agent organises them into an outline with references, and a coding agent drafts analysis scripts for a researcher to test. Publication and data interpretation remain with the research team.
- Corpus-scoped collections with citation metadata
- Tools into reference managers and storage
- Long-context models for whole-paper reading
- Versioned prompts and reproducible runs
Data governance and human oversight
Research data ranges from public to highly sensitive. Choose the deployment tier per project, keep unpublished results inside the boundary, and record provenance so findings can be reproduced.
- Project-scoped keys and collections
- On-prem or air-gapped options for sensitive data
- Audit logs and reproducible prompt versions
- Researcher review of every cited claim
From pilot to production
Pilot on a literature review with a well-defined corpus. Check citation accuracy before trusting synthesis, then extend to project memory.
Keep the rollout reversible: run the agent in shadow mode alongside the current process, compare outputs on your own samples, and move it into the workflow only when the evidence holds. Document what you measured so expanding to the next team is a decision, not a hope.
Honest comparison
| Capability | Plugsky | Typical cloud AI API | Building in-house |
|---|---|---|---|
| API compatibility | Drop-in base URL change | Usually compatible | Full rewrite |
| Model access | 30+ models behind one API | Vendor's own catalogue | You host each model |
| Pricing | Flat monthly self-serve plans; see live pricing | Often per-token | GPU + ops cost |
| Deployment | Cloud, VPC, on-prem or air-gapped | Usually vendor cloud regions | You own the stack |
| Citations | Retrieval-grounded with source links | Varies | You build it |
| Sensitive data | On-prem/air-gapped options | Often cloud-only | You own the stack |
Frequently asked questions
Do we have to rewrite our application?
No. The chat completions API is OpenAI-compatible, so you change the base URL and model name and keep your existing SDK.
Is there a free plan?
Yes — the free plan includes two free AI models, plugsky-micro and plugsky-lite, with no credit card required.
How is pricing structured?
Self-serve plans are flat monthly with fair-use usage and no per-token charges; see the live pricing page for current plans.
Which endpoints are live today?
Chat, streaming, JSON mode, function calling, embeddings, RAG and agents are live. Audio, images, moderation, files, batch, fine-tuning, assistants and responses endpoints are coming soon — check the docs before planning around them.
Can agents fabricate references?
Any generative model can produce incorrect citations; that is why retrieval-grounded answers with source links and human checking are the working pattern. Verify every reference.
Can I use it with sensitive datasets?
Yes — pick a deployment tier that keeps data inside your approved boundary, including on-prem or air-gapped options.
Does it support long papers?
Long-context models can process large excerpts; for large corpora, retrieval plus chunking is usually more reliable.