Key facts
| API compatibility | OpenAI-compatible /v1/chat/completions (drop-in base URL change) |
| Models | 30+ models in one catalogue, from free tiers to frontier reasoning |
| Pricing | Flat monthly self-serve plans with unlimited fair-use usage |
| Free tier | plugsky-micro and plugsky-lite, no credit card |
| Trial | 14-day full-access trial for stronger models |
| Deployment | Plugsky cloud, your VPC, on-prem and air-gapped |
| Best fit | Backend services, agents, batch pipelines and RAG |
TL;DR
- Workers AI wins when inference must run inside a Worker at the edge.
- Backend, batch and agent workloads are usually better on a dedicated API.
- OpenAI-compatible calls keep your SDKs, agents and RAG pipelines portable.
- Flat monthly pricing and a free tier make cost and evaluation simple.
- Run both: edge inference on Cloudflare, central workloads on Plugsky.
How it works, step by step
- Separate request-path inference inside Workers from backend and batch workloads.
- For the backend set, list endpoints used: chat, tools, embeddings, RAG.
- Create a Plugsky key on the free plan and replay recorded prompts.
- Change the base URL and model names in the backend services; keep Worker bindings as they are.
- Compare quality, latency and error behaviour before shifting traffic.
- Add residency and deployment requirements to the evaluation if you operate in regulated markets.
Original data
Try it yourself
Open the Cloudflare Workers AI cost calculator →
Where Workers AI fits, and where it does not
Workers AI shines for lightweight inference on the request path: classify, summarise or answer inside the same Worker that serves the user, with no extra network hop. The trade-offs appear when workloads grow: long agent loops, large-context prompts, batch enrichment and pipelines that live outside Cloudflare are awkward to keep inside the edge runtime.
Those workloads want a standalone endpoint with predictable behaviour, model choice and limits you can reason about. That is the gap a dedicated API such as Plugsky fills, without asking you to restructure the edge code you already like.
Moving backend and agent workloads
Plugsky exposes an OpenAI-compatible chat completions API, so backend services written against the OpenAI SDK need only a base URL and model-name change. Agents, evaluation harnesses and RAG components that speak the same schema keep working, and one key covers chat, embeddings and agent calls.
- Keep Workers AI bindings for request-path inference.
- Route backend, batch and agent traffic to Plugsky.
- Replay production prompts to compare output quality and latency.
- Use embeddings from one provider so your vector index stays consistent.
Latency, residency and cost shape
Edge inference minimises distance to the user, but it cannot always satisfy residency requirements that demand a specific country or a private data plane. Plugsky supports region selection and deployment in your VPC, on-prem or air-gapped environments, which suits regulated teams better than a global edge network.
Cost shape is the other difference: Plugsky uses flat monthly self-serve plans with unlimited fair-use usage instead of per-request arithmetic. See the live pricing page for current plans. Audio, image, moderation, files, batch, fine-tuning, assistants and responses are coming soon; if you rely on those, keep them on Cloudflare or another provider for now.
Honest comparison
| Capability | Plugsky | Cloudflare Workers AI | Self-hosting inference |
|---|---|---|---|
| Run location | Cloud, customer VPC, on-prem, air-gapped | Cloudflare edge network | Your GPUs |
| API style | OpenAI-compatible | Workers bindings and REST API | Runtime-specific |
| Pricing shape | Flat monthly self-serve | Usage-based | GPU plus ops cost |
| Best fit | Backend, agents, batch, RAG | In-Worker request-path inference | Full control |
| Model choice | 30+ models, one endpoint | Hosted open models | Open-weight models only |
Frequently asked questions
Can Plugsky replace Workers AI inside a Cloudflare Worker?
Not for in-Worker bindings. Keep Workers AI for request-path inference and use Plugsky for backend, batch and agent workloads where an OpenAI-compatible endpoint is a better fit.
Is migration hard?
No. The chat completions schema is OpenAI-compatible, so most client code changes only its base URL and model name.
Is there a free plan?
Yes. plugsky-micro and plugsky-lite are free with no credit card, and a 14-day full-access trial opens stronger models.
How is pricing structured?
Self-serve plans are flat monthly with unlimited fair-use usage and no per-token billing. See the live pricing page for current plans.
Does Plugsky help with data residency?
Yes. Region selection and sovereign deployments including VPC, on-prem and air-gapped are available for teams that need them.
Can I run both providers at once?
Yes, and that is a common pattern: edge inference on Cloudflare, central inference and embeddings on Plugsky.