Key facts
| API compatibility | OpenAI-compatible /v1/chat/completions; change the base URL |
| Models | 30+ models, from free chat tiers to frontier reasoning |
| Pricing | Flat monthly self-serve plans with unlimited fair-use usage |
| Free plan | plugsky-micro and plugsky-lite; no card required |
| Trial | 14-day full-access trial for higher tiers |
| Deployment | Plugsky cloud, your VPC, on-prem and air-gapped |
| Data residency | Region-locked data planes; see /data-residency for current regions |
| Local presence | No Budaiya office or data-centre claim; residency is region-based |
TL;DR
- OpenAI-compatible API: change the base URL and keep your SDK.
- 30+ models behind one endpoint, from free chat tiers to frontier reasoning.
- Flat monthly plans with no per-token billing on self-serve.
- Region-locked data planes plus VPC, on-prem and air-gapped options.
- Free plan with plugsky-micro and plugsky-lite; 14-day full-access trial.
How it works, step by step
- Create a Plugsky account and issue an API key on the free plan — no card required.
- Point your OpenAI SDK at api.plugsky.com and map your model names.
- Pick the deployment topology that satisfies your residency requirement: Plugsky cloud region, your VPC, on-prem or air-gapped.
- Replay your existing prompts and evaluations against the same test set.
- Measure latency from your Budaiya network against candidate regions before cut-over.
- Review the DPA, the SLA and the /data-residency documentation with security and legal.
- Cut production over behind a flag and keep the one-line rollback ready.
Original data
Try it yourself
Open the AI data residency checklist →
What enterprise AI cloud means for Budaiya
Budaiya is a residential town on Bahrain's northern coast, west of Manama. Local teams in retail, property, education and professional services are adopting AI at two speeds: experiments with frontier models, and production features that depend on small, inexpensive models for extraction, routing and chat. Plugsky matches that with 30+ models behind one OpenAI-compatible API and flat monthly self-serve plans, so a prototype does not need a rewrite to reach production.
Start on the free plan with plugsky-micro and plugsky-lite, then move to larger models when your evaluations justify it. Integration details are in the docs.
How data residency works for Budaiya teams
Residency at Plugsky is a deployment property, not a marketing claim: region-locked data planes keep requests, logs and artefacts inside the region you select. Plugsky does not announce a data centre in Budaiya, so the honest checks are the published region list and the contract — review the data-residency overview, the DPA and the SLA with your security team.
If policy requires a stricter boundary, run the same OpenAI-compatible API in your VPC, on-prem or air-gapped; your application code does not change.
Deployment options and latency trade-offs
For Budaiya teams the practical choice is between a managed region, your own cloud account, or a fully isolated deployment. Each keeps the API identical; what changes is who operates the hardware and where the processing boundary sits.
Latency is the second axis. A nearer data plane usually means shorter round trips, while strict residency can add distance. Test candidate regions from your own network with the same prompts, compare p50 and p95, and pick the configuration that meets policy without overshooting on latency.
Migrating and controlling cost from Budaiya
Switching is deliberately small: change the base URL, map model names, replay your test suite. Keep the old provider configured behind a flag so rollback is a one-line change.
On cost, self-serve plans are flat monthly with unlimited fair-use usage — no per-token billing — so budgeting for Budaiya projects stays a single line item. Evaluate on the free plan (plugsky-micro and plugsky-lite), then use the 14-day full-access trial for larger models. See the pricing page for current plans.
Honest comparison
| Capability | Plugsky | Typical global API provider | Building in-house |
|---|---|---|---|
| API compatibility | OpenAI-compatible — change the base URL | Usually compatible, varies by model | Full rewrite and integration work |
| Models | 30+ models behind one endpoint | Mainly the provider's own catalogue | You host and maintain each model |
| Pricing | Flat monthly self-serve plans, no per-token billing | Per-token billing, harder to forecast | GPU, operations and staffing costs |
| Data residency | Region-locked data planes; VPC, on-prem, air-gapped options | Limited region choices | You own the responsibility |
| Local presence in Budaiya | No office or data-centre claim; residency is region-based | Varies; few publish local commitments | Depends on your own facilities |
| Migration effort | Base URL plus model mapping | Depends on compatibility gaps | Long integration cycle |
Frequently asked questions
Does Plugsky have a data centre or office in Budaiya?
No. Plugsky does not claim a local office or data centre in Budaiya; residency is delivered through region-locked data planes, or through deployments in your own VPC, on-prem or air-gapped. See /data-residency for the current region list.
How does data residency work for Budaiya teams?
You select a region and Plugsky keeps the data plane locked to it, so requests, logs and artefacts are not routed globally by default. Validate the region list and the contractual terms in the DPA and the SLA before you commit.
Can we keep using the OpenAI SDK?
Yes. Plugsky exposes an OpenAI-compatible chat completions endpoint, so you change the base URL and model name and keep your existing SDK, prompts and evaluations.
Is there a free plan?
Yes — the free plan includes plugsky-micro and plugsky-lite with no credit card required. A 14-day full-access trial is available when you want to evaluate larger models.
How is pricing structured?
Self-serve plans are flat monthly with unlimited fair-use usage and no per-token billing. See the live pricing page for current plans and fair-use limits.
Which capabilities are live today?
Chat, streaming, JSON mode, function calling, embeddings, RAG and agents are live. Audio, images, moderation, files, batch, fine-tuning, assistants and responses are coming soon — check the docs before planning those workloads.
What about latency from Budaiya?
It depends on the region you select and how your network routes to it. Measure candidate regions from your own environment, compare p50 and p95, and treat residency and latency as an explicit trade-off.
How much migration work is involved?
For a standard chat application the code change is a base URL plus model mapping; the real work is replaying your evaluations and doing a staged cut-over.