Key facts
| Provider | Cloudflare Workers AI — edge-hosted inference for a curated set of open models |
| API style | Workers binding for in-platform calls plus an OpenAI-compatible endpoint |
| Plugsky API | OpenAI-compatible /v1/chat/completions — change the base URL, keep your SDK |
| Models | 30+ models from free to frontier behind one API key |
| Pricing | Flat monthly plans with unlimited fair-use usage; no per-token billing on self-serve |
| Free tier | Free plan with plugsky-micro and plugsky-lite, no card; 14-day full-access trial |
| Deployment | Plugsky cloud, your VPC, on-prem or air-gapped; region choice for residency |
| Live vs roadmap | Chat, streaming, JSON mode, function calling, embeddings, RAG, agents live; audio, images, moderation, files, batch, fine-tuning, assistants, responses coming soon |
TL;DR
- Workers AI wins when latency is dominated by physical distance to users.
- Plugsky wins on catalogue breadth, flat billing and deployment options.
- Both speak OpenAI-format requests, so moving between them is mostly configuration.
- Plugsky free tier: plugsky-micro and plugsky-lite; 14-day full-access trial for the rest.
- Honest trade-off: Plugsky has no edge-local GPU story to match Cloudflare's.
How it works, step by step
- Profile where your latency budget goes: network distance, model time or application logic.
- Decide which inference can be edge-local and which benefits from a broader catalogue.
- Create a Plugsky account and test the models that cover your non-edge workloads.
- Replace Workers AI bindings with HTTP calls where portability is the goal.
- Compare measured latency and cost for both paths on the same prompts.
- Keep the edge for what it does best and consolidate the rest on one platform.
Original data
Try it yourself
Open the Cloudflare Workers AI cost calculator →
What the Workers AI API model gives you
Cloudflare's model is attractive because inference lives inside the same runtime as your application code. A binding call avoids API keys, extra network hops and separate deployment pipelines. Because the GPU pool sits on Cloudflare's global network, requests can be served near the user, which helps interactive features.
The trade is scope. The model list is curated, request limits are designed for edge workloads, and the service is inseparable from Cloudflare's platform. When your roadmap needs a wider catalogue or an in-region deployment you control, the edge stops being the whole answer.
What the Plugsky API model gives you
Plugsky decouples inference from any single cloud runtime. One key reaches 30+ models through OpenAI-format requests, with flat monthly self-serve plans on unlimited fair-use usage (live pricing). The free plan covers plugsky-micro and plugsky-lite, and a 14-day full-access trial covers paid models.
For teams with governance constraints, enterprise deployment runs in your VPC, on-prem or air-gapped, with region selection. Chat, streaming, JSON mode, function calling, embeddings, RAG and agents are live; audio, images, moderation, files, batch, fine-tuning, assistants and responses are coming soon.
How to split inference between edge and platform
Draw the line by workload, not by preference. Edge-local inference suits short, high-frequency calls where network distance dominates. Platform inference suits reasoning, long context, embeddings and anything that needs model choice.
- Keep classification, routing and caching at the edge.
- Send generation and multi-step agent work to a catalogue-rich platform.
- Standardize on OpenAI request shapes so either side can serve as fallback.
- Track per-workload cost and latency monthly; this trade-off shifts as both platforms evolve.
Honest comparison
| Capability | Plugsky | Workers AI | Building in-house |
|---|---|---|---|
| API style | OpenAI-compatible drop-in | Workers binding plus OpenAI-compatible endpoint | You define the schema |
| Placement | Platform endpoint with region choice | GPU inference inside the edge network | You choose the infrastructure |
| Model catalogue | 30+ models, free to frontier | Curated open models for edge workloads | You host each model |
| Billing | Flat monthly, unlimited fair use (see live pricing) | Usage-based, billed through Cloudflare | GPU + ops cost |
| Governance | Region choice, VPC, on-prem, air-gapped | Cloudflare locations, no self-host option | You control the infrastructure |
| Honest gap | No edge-local GPU proximity | Inference physically near users | You build it |
Frequently asked questions
Which is faster?
It depends on distance and model. Edge inference can win on network latency; platform inference can win on the models and context sizes available. Measure from your users' regions.
Can I call Plugsky from a Worker?
Yes. Plugsky exposes HTTP OpenAI-compatible endpoints, so Workers can call it with fetch and a stored API key.
Does Plugsky replace Workers AI entirely?
Not necessarily. Teams often keep edge logic on Cloudflare and move heavier inference to Plugsky.
Is there a free plan?
Yes — plugsky-micro and plugsky-lite are free with no card, and a 14-day full-access trial is available.
How is Plugsky billed?
Flat monthly self-serve plans with unlimited fair-use usage, no per-token billing. See the live pricing page.
Can I run inference privately?
Plugsky supports VPC, on-prem and air-gapped deployments; Workers AI runs only on Cloudflare's network.
How many models does each offer?
Plugsky 30+ models across tiers and tasks; Workers AI a curated set of open models for edge workloads.