Key facts
| API compatibility | OpenAI-compatible /v1/chat/completions (drop-in base URL change) |
| Models | 30+ models in one catalogue, from free tiers to frontier reasoning |
| Pricing | Flat monthly self-serve plans with unlimited fair-use usage; no per-token billing on self-serve |
| Free tier | plugsky-micro and plugsky-lite on the free plan, no card required |
| Trial | 14-day full-access trial for stronger models |
| Deployment | Plugsky cloud, your VPC, on-prem and air-gapped |
| Media endpoints | Image, audio and moderation endpoints are coming soon; chat and embeddings are live |
TL;DR
- Aggregators simplify model access; flat pricing simplifies cost forecasting.
- Keep your OpenAI SDK and change the base URL plus model names.
- 30+ models behind one API, from free chat tiers to frontier reasoning.
- The free plan covers development, and the 14-day full-access trial tests stronger models.
- If image or audio endpoints are core, keep your current provider until those APIs ship.
How it works, step by step
- List the endpoints you actually call — chat, embeddings, tools, images, audio — and mark which are live on Plugsky today.
- Create a Plugsky API key on the free plan (plugsky-micro and plugsky-lite, no card).
- Point a staging environment at the Plugsky base URL and map your model names.
- Replay recorded requests and compare outputs, latency and error rates.
- Move text workloads first; leave media workloads on your current provider until those endpoints are live.
- Track usage in the dashboard and shift production traffic once your evals pass.
Original data
Try it yourself
Open the AI/ML API cost calculator →
Why teams look beyond AI/ML API
Aggregators solve discovery: one key, many models, including media models that a plain chat API does not cover. The trade-offs appear later. Multi-model usage-based billing makes forecasting hard, model availability can shift, and residency questions rarely have a simple answer. Teams that have settled on a stable text stack often decide they want fewer moving parts and one predictable invoice instead of a marketplace.
Plugsky takes the opposite approach: a curated catalogue of 30+ models behind one OpenAI-compatible endpoint, flat monthly self-serve pricing, and deployment choices that include your VPC, on-prem and air-gapped environments. See the live pricing page for current plans rather than comparing old numbers.
Compatibility and migration notes
Migration work is mostly mechanical. The chat completions schema, streaming, JSON mode and function calling follow OpenAI conventions, so existing client code keeps working after a base URL and model-name change. The real effort is in prompts and evaluations: replay recorded traffic, compare outputs and watch for model-specific quirks such as tool-call formatting or refusal behaviour.
- Keep a mapping table from your current model names to Plugsky models.
- Pin a default model so a missing mapping fails loudly instead of silently.
- Store prompts and evals in version control so switching back is cheap.
- Test streaming and tool calls, not just single completions.
What to verify before switching
Check the endpoint matrix first: chat, streaming, embeddings, function calling and agents are live; audio, images, moderation, files, batch, fine-tuning, assistants and responses are coming soon. If your product depends on a coming-soon endpoint, keep the current provider for that part of the stack and route the rest through Plugsky.
Then check operational fit: data residency and deployment model, rate limits under fair use, and how you will monitor usage. Because your code stays OpenAI-compatible, rollback is the same one-line change in reverse — which makes a staged migration low risk.
Honest comparison
| Capability | Plugsky | AI/ML API | Self-hosting models |
|---|---|---|---|
| Text chat and tools | OpenAI-compatible, live | Many models via one API | You run inference |
| Media endpoints | Images and audio coming soon | Broader media coverage | Build per model |
| Pricing shape | Flat monthly, unlimited fair use | Usage-based per model | GPU plus ops cost |
| Free tier | plugsky-micro and plugsky-lite, no card | Varies by model | No free tier |
| Deployment | Cloud, VPC, on-prem, air-gapped | Managed API | You own the stack |
Frequently asked questions
Can I keep my OpenAI SDK when switching from AI/ML API?
Yes. Plugsky exposes an OpenAI-compatible chat completions endpoint, so you change the base URL and model name and keep your SDK and most client code.
Is there a free plan?
Yes. The free plan includes plugsky-micro and plugsky-lite with no credit card, which is enough for development and small prototypes.
What does the 14-day full-access trial include?
It opens the full model catalogue for evaluation so you can test frontier models on your own prompts before choosing a paid plan.
How is pricing structured?
Self-serve plans are flat monthly with unlimited fair-use usage; there is no per-token billing on self-serve plans. Check the live pricing page for current plans.
Are image and audio endpoints available?
Not yet. Image, audio and moderation endpoints are coming soon, so keep your current provider for those workloads until they are live.
Does Plugsky support data residency?
Yes. You can choose regions and deploy in your VPC, on-prem or air-gapped environments for stricter residency requirements.
What if I want to move back?
Your code stays OpenAI-compatible, so switching back is the same one-line change in reverse.