Developer + API

What is the Plugsky LLM API?

The Plugsky LLM API is an HTTP service that runs 30+ large language models behind one OpenAI-compatible endpoint. Your existing OpenAI SDK, prompts and streaming code work after a base_url change. Models range from free plugsky-micro and plugsky-lite through production chat and code models to long-context, reasoning and multimodal options, with routing and fallback on the same key.

Key facts

CompatibilityOpenAI-compatible /v1/chat/completions with streaming, tools and JSON mode
Models30+ models from free chat to frontier reasoning and multimodal
Free tierplugsky-micro and plugsky-lite on the free plan, no card required
RoutingModel routing and fallback chains across models
ContextModel-dependent windows, up to 1M tokens on long-context models
Pricing modelFlat monthly self-serve plans with unlimited fair-use usage
DeploymentCloud, VPC, on-prem and air-gapped options
Product statusLive

TL;DR

  • One endpoint, 30+ models, no code rewrite.
  • Start free with plugsky-micro and plugsky-lite, then scale to a paid tier.
  • Route simple requests to small models and hard ones to frontier models.
  • Fallback chains keep requests alive when a model is unavailable.
  • Deployment options range from public cloud to air-gapped.

How it works, step by step

  1. Create a key on the free plan, with no card required.
  2. Point your OpenAI client at https://plugsky.com/v1.
  3. Choose a model per task: micro or lite for simple calls, pro for production, frontier for hard reasoning.
  4. Enable streaming for interactive applications.
  5. Add model routing or a fallback chain for resilience.
  6. Monitor usage and move to a higher tier when rate limits bind.
1Create a key on thefree plan, with nocard required.2Point your OpenAIclient athttps://plugsky.com/v1.3Choose a model pertask: micro or litefor simple calls,4Enable streamingfor interactiveapplications.5Add model routingor a fallback chainfor resilience.6Monitor usage andmove to a highertier when rate

Original data

OpenAI-compatiCompatibility30+ models froModelsModel-dependenContextSource: Plugsky facts table · updated 2026-09-25

Try it yourself

Open the model picker →

The models on Plugsky

The catalogue spans several tiers. Free models such as plugsky-micro and plugsky-lite cover fast reasoning, classification and simple chat. Mid-tier models such as plugsky-plus and plugsky-pro handle general chat, content generation, code and function calling.

Above those sit long-context and multimodal models, hard-reasoning models such as plugsky-frontier and plugsky-reasoning, cost-optimised models such as plugsky-deepseek-flash, and embedding models with 1536 or 3072 dimensions. See the live catalogue for the current list and per-model context windows.

Choosing the right model for the task

The right model depends on three factors: quality required, context length and cost ceiling. Simple classification, intent detection and short replies belong on the free models. General chat and function calling belong on a mid-tier model. Book-length or codebase analysis needs a long-context model, and hard reasoning, maths or complex code belongs on the frontier tier.

Multimodal work moves to the vision models, and high-volume cost-sensitive work to the efficient models. Matching each request to the smallest model that can handle it is the single biggest cost lever.

Routing, failover and limits

Use model routing to pick the right model automatically, or configure a fallback chain so a 503 or rate limit on one model shifts traffic to a peer. The API returns 503 with a Retry-After header when a model is unavailable, which makes client-side fallbacks straightforward.

All models are unlimited under fair use on self-serve plans, with rate limits that scale by tier. Fine-tuning on open base models is available on Enterprise contracts and deploys into your tenant.

The same key covers the full 30+ model catalogue with no separate accounts.

Honest comparison

NeedRecommended Plugsky modelsWhyAlternative
Free tierplugsky-micro, plugsky-liteNo card, fastLocal small models
Production chatplugsky-pro, plugsky-plusBalanced quality and speedMid-tier models elsewhere
Long contextplugsky-max, plugsky-longctxWindows up to 1M tokensChunking plus RAG
Hard reasoningplugsky-frontier, plugsky-reasoningStrongest reasoning tierMultiple vendors
Multimodalplugsky-vision, plugsky-ultraImage and document understandingSeparate vision API
High volumeplugsky-deepseek-flash, plugsky-phiFast and efficientSelf-hosting at saturation

Frequently asked questions

How do I switch between models?

Change the model parameter in your request. No other code changes are needed.

Can one application use multiple models?

Yes. Different requests can target different models, and model routing can choose automatically per request.

Is fine-tuning available?

On Enterprise contracts. Plugsky fine-tunes open base models on your data and deploys the result in your tenant.

What happens when a model is down?

The API returns 503 with a Retry-After header. Configure a fallback chain in your client to try another model automatically.

Is there a free plan?

Yes. The free plan includes plugsky-micro and plugsky-lite with no card required, and a 14-day full-access trial is available.

Are audio and image endpoints available?

Image generation and audio endpoints are coming soon. Chat, streaming, JSON mode, function calling, embeddings, RAG and agents are live today.

Cite this page

Plugsky (2026). “LLM API: 30+ Models, One Endpoint”. Plugsky. Available at: https://plugsky.com/articles/llm-api (last updated 2026-09-25).