Key facts
| Compatibility | OpenAI-compatible /v1/chat/completions with streaming, tools and JSON mode |
| Models | 30+ models from free chat to frontier reasoning and multimodal |
| Free tier | plugsky-micro and plugsky-lite on the free plan, no card required |
| Routing | Model routing and fallback chains across models |
| Context | Model-dependent windows, up to 1M tokens on long-context models |
| Pricing model | Flat monthly self-serve plans with unlimited fair-use usage |
| Deployment | Cloud, VPC, on-prem and air-gapped options |
| Product status | Live |
TL;DR
- One endpoint, 30+ models, no code rewrite.
- Start free with plugsky-micro and plugsky-lite, then scale to a paid tier.
- Route simple requests to small models and hard ones to frontier models.
- Fallback chains keep requests alive when a model is unavailable.
- Deployment options range from public cloud to air-gapped.
How it works, step by step
- Create a key on the free plan, with no card required.
- Point your OpenAI client at https://plugsky.com/v1.
- Choose a model per task: micro or lite for simple calls, pro for production, frontier for hard reasoning.
- Enable streaming for interactive applications.
- Add model routing or a fallback chain for resilience.
- Monitor usage and move to a higher tier when rate limits bind.
Original data
Try it yourself
The models on Plugsky
The catalogue spans several tiers. Free models such as plugsky-micro and plugsky-lite cover fast reasoning, classification and simple chat. Mid-tier models such as plugsky-plus and plugsky-pro handle general chat, content generation, code and function calling.
Above those sit long-context and multimodal models, hard-reasoning models such as plugsky-frontier and plugsky-reasoning, cost-optimised models such as plugsky-deepseek-flash, and embedding models with 1536 or 3072 dimensions. See the live catalogue for the current list and per-model context windows.
Choosing the right model for the task
The right model depends on three factors: quality required, context length and cost ceiling. Simple classification, intent detection and short replies belong on the free models. General chat and function calling belong on a mid-tier model. Book-length or codebase analysis needs a long-context model, and hard reasoning, maths or complex code belongs on the frontier tier.
Multimodal work moves to the vision models, and high-volume cost-sensitive work to the efficient models. Matching each request to the smallest model that can handle it is the single biggest cost lever.
Routing, failover and limits
Use model routing to pick the right model automatically, or configure a fallback chain so a 503 or rate limit on one model shifts traffic to a peer. The API returns 503 with a Retry-After header when a model is unavailable, which makes client-side fallbacks straightforward.
All models are unlimited under fair use on self-serve plans, with rate limits that scale by tier. Fine-tuning on open base models is available on Enterprise contracts and deploys into your tenant.
The same key covers the full 30+ model catalogue with no separate accounts.Honest comparison
| Need | Recommended Plugsky models | Why | Alternative |
|---|---|---|---|
| Free tier | plugsky-micro, plugsky-lite | No card, fast | Local small models |
| Production chat | plugsky-pro, plugsky-plus | Balanced quality and speed | Mid-tier models elsewhere |
| Long context | plugsky-max, plugsky-longctx | Windows up to 1M tokens | Chunking plus RAG |
| Hard reasoning | plugsky-frontier, plugsky-reasoning | Strongest reasoning tier | Multiple vendors |
| Multimodal | plugsky-vision, plugsky-ultra | Image and document understanding | Separate vision API |
| High volume | plugsky-deepseek-flash, plugsky-phi | Fast and efficient | Self-hosting at saturation |
Frequently asked questions
How do I switch between models?
Change the model parameter in your request. No other code changes are needed.
Can one application use multiple models?
Yes. Different requests can target different models, and model routing can choose automatically per request.
Is fine-tuning available?
On Enterprise contracts. Plugsky fine-tunes open base models on your data and deploys the result in your tenant.
What happens when a model is down?
The API returns 503 with a Retry-After header. Configure a fallback chain in your client to try another model automatically.
Is there a free plan?
Yes. The free plan includes plugsky-micro and plugsky-lite with no card required, and a 14-day full-access trial is available.
Are audio and image endpoints available?
Image generation and audio endpoints are coming soon. Chat, streaming, JSON mode, function calling, embeddings, RAG and agents are live today.
Plugsky (2026). “LLM API: 30+ Models, One Endpoint”. Plugsky. Available at: https://plugsky.com/articles/llm-api (last updated 2026-09-25).