Key facts
| Endpoint | OpenAI-compatible /v1/chat/completions |
| Streaming | Server-sent events, same client handling |
| Structured output | JSON mode is live |
| Tool use | Function calling is live |
| Models | 30+ models behind the same endpoint |
| Pricing | Flat monthly self-serve plans with unlimited fair-use usage |
| Rollback | Same one-line change in reverse |
TL;DR
- Drop-in means schema parity, not brand similarity.
- Verify streaming and tool calls, not just basic completions.
- Keep one SDK and change the base URL and model names.
- Test with recorded production traffic before cutover.
- Document rollback; it is the same one-line change in reverse.
How it works, step by step
- List every OpenAI API feature your code touches, including error handling.
- Generate a Plugsky key on the free plan and point a staging client at the endpoint.
- Run your existing test suite and compare response shapes, including streaming events.
- Replay recorded production requests and diff summaries of the outputs.
- Canary a small traffic share and watch latency, errors and quality signals.
- Cut over fully, keep rollback documented, and review after one week.
Try it yourself
Open the OpenAI compatibility checker →
What drop-in really means
Drop-in is a schema claim, not a marketing one. Your client expects specific request fields, a specific response envelope, streaming event shapes and error formats. If any of those differ, the change is not one line even if the endpoint looks OpenAI-shaped at a glance.
Plugsky implements the chat completions contract, including streaming, JSON mode and function calling, so OpenAI SDK code transfers with a base URL and model-name change. Embeddings and agents use the same key, which keeps configuration small.
Verify before cutover
A test matrix beats a demo. Exercise the primitives and the edges, because most migration incidents happen in error paths rather than happy paths.
- Single completions and multi-turn chat with a system prompt.
- Streaming, including cancellation and partial consumption.
- Function calling with the tool schemas you actually ship.
- JSON mode with strict parsers.
- Timeouts, retries and rate-limit responses.
Replay recorded traffic for realism, then verify that logs and tracing still capture what you need.
Rollback and dual-run
The strongest safety property of an OpenAI-compatible replacement is reversibility. Keep the old base URL and key in configuration so flipping back takes seconds, and use a feature flag or environment variable rather than code edits.
Dual-run a slice of traffic long enough to cover peak load and unusual prompts. Plugsky supports 30+ models, flat monthly self-serve pricing and a free tier for staging, plus a 14-day full-access trial for frontier models; see the live pricing page for current plans. Media and platform endpoints such as audio, images, moderation, files, batch, fine-tuning, assistants and responses are coming soon.
Honest comparison
| Concern | Plugsky | Typical drop-in replacement | Staying on OpenAI |
|---|---|---|---|
| Chat completions schema | OpenAI-compatible | Usually compatible | Native |
| Streaming and tools | Live | Varies | Native |
| Pricing shape | Flat monthly self-serve | Often usage-based | Usage-based |
| Deployment | Cloud, VPC, on-prem, air-gapped | Usually cloud-only | Managed cloud |
| Rollback | One line | Varies | None needed |
Frequently asked questions
What makes a replacement truly drop-in?
Schema parity across the features you use: chat completions, streaming, JSON mode, function calling, error shapes and timeouts. If those match, your SDK and client code keep working.
Will my OpenAI SDK work unchanged?
Yes, with a base URL and model-name change. Plugsky exposes an OpenAI-compatible endpoint and a single API key.
Is there a free plan?
Yes. plugsky-micro and plugsky-lite are available free with no credit card, which is enough for staging and evaluation.
How is pricing structured?
Self-serve plans are flat monthly with unlimited fair-use usage and no per-token billing. See the live pricing page for current plans.
What should I test before switching?
Streaming, tool calls, JSON mode, retries and rate-limit handling on recorded production traffic. Error paths cause most migration surprises.
How do I roll back?
Change the base URL back. Because your code stays OpenAI-compatible, rollback is the same one-line change in reverse.