Key facts
| API compatibility | OpenAI-compatible /v1/chat/completions; change the base URL |
| Models | 30+ models, from free chat tiers to frontier reasoning |
| Pricing | Flat monthly self-serve plans with unlimited fair-use usage; no per-token billing |
| Free plan | plugsky-micro and plugsky-lite; no card required |
| Trial | 14-day full-access trial for higher tiers |
| Deployment | Plugsky cloud, your VPC, on-prem and air-gapped |
| Data residency | Region-locked data planes; US (Virginia and Oregon) is the closest published plane for Toronto |
| Local presence | No Toronto office or data-centre claim; residency is region-based |
TL;DR
- OpenAI-compatible API: change the base URL and keep your SDK.
- 30+ models behind one endpoint, from free chat tiers to frontier reasoning.
- Flat monthly plans with no per-token billing on self-serve.
- Region-locked data planes plus VPC, on-prem and air-gapped options.
- Free plan with plugsky-micro and plugsky-lite; 14-day full-access trial.
How it works, step by step
- Create a Plugsky account and issue an API key on the free plan — no card required.
- Point your OpenAI SDK at api.plugsky.com and map your model names.
- Pick the deployment topology that satisfies your residency requirement: Plugsky cloud region, your VPC, on-prem or air-gapped.
- Replay your existing prompts and evaluations against the same test set.
- Measure latency from your Toronto network against candidate regions before cut-over.
- Review the DPA, the SLA and the /data-residency documentation with security and legal.
- Cut production over behind a flag and keep the one-line rollback ready.
Original data
Try it yourself
Open the AI data residency checklist →
What enterprise AI cloud means for Toronto
Canada's largest city and financial capital, Toronto has a deep software, healthcare and AI research ecosystem. Teams here run financial services, healthcare document workflows and SaaS products, and the pattern is familiar: the people building new products reach for frontier reasoning, while day-to-day traffic settles on smaller models for classification, extraction and chat. One OpenAI-compatible endpoint and 30+ models cover both, so you can prototype on plugsky-micro or plugsky-lite and scale up without changing SDKs.
Existing code, prompts and evaluations keep working; the docs list endpoints and model names.
How data residency works for Toronto teams
Plugsky supports region-locked data planes in the EU (Frankfurt), GCC (UAE and KSA), APAC (Singapore) and US (Virginia and Oregon). Requests, logs and stored artefacts stay in the plane you choose instead of being routed globally by default, and data does not leave it. For Toronto, the closest published plane is the US plane (Virginia and Oregon). There is no Plugsky office or data centre in Toronto; residency is a property of the region and deployment model, not of local premises.
The US planes are the closest published option for Canadian teams; if you require Canadian-only residency, confirm current availability at /data-residency before moving production. Validate the current list and the contractual wording in the data-residency overview, the DPA and the SLA before you sign.
Deployment options and latency trade-offs
Four paths cover most Toronto scenarios: a managed Plugsky cloud region, a private endpoint inside your own AWS, Azure or GCP account, on-prem for infrastructure you own, and air-gapped for classified or critical workloads with no internet egress, a local model registry and offline update channels. Residency and latency trade off: a closer plane shortens the network path, while a strict boundary can add round-trip time versus a globally distributed endpoint.
Measure before committing — run the same prompts from your Toronto network against candidate regions, compare p50 and p95, then decide. Region and deployment selection are configuration, not a rewrite.
Migrating and controlling cost from Toronto
Migration is one line: point the OpenAI SDK base URL at Plugsky and map model names, then replay your evaluations and cut over behind a flag so rollback stays trivial. Chat, streaming, JSON mode, function calling, embeddings, RAG and agents are live; audio, images, moderation, files, batch, fine-tuning, assistants and responses are coming soon.
Cost control is flat-rate on self-serve plans — unlimited fair-use usage with no per-token billing — so Toronto teams can forecast a monthly line item instead of token spend. Start free with plugsky-micro and plugsky-lite, and use the 14-day full-access trial to evaluate larger models. Current plans are listed on the pricing page.
Honest comparison
| Capability | Plugsky | Typical global API provider | Building in-house |
|---|---|---|---|
| API compatibility | OpenAI-compatible — change the base URL | Usually compatible, varies by model | Full rewrite and integration work |
| Models | 30+ models behind one endpoint | Mainly the provider's own catalogue | You host and maintain each model |
| Pricing | Flat monthly self-serve plans, no per-token billing | Per-token billing, harder to forecast | GPU, operations and staffing costs |
| Data residency | Region-locked data planes; VPC, on-prem, air-gapped options | Limited region choices | You own the responsibility |
| Local presence in Toronto | No office or data-centre claim; residency is region-based | Varies; few publish local commitments | Depends on your own facilities |
| Migration effort | Base URL plus model mapping | Depends on compatibility gaps | Long integration cycle |
Frequently asked questions
Does Plugsky have a data centre or office in Toronto?
No. Plugsky does not claim a local office or data centre in Toronto; residency is delivered through region-locked data planes, or through deployments in your own VPC, on-prem or air-gapped. See /data-residency for the current region list.
How does data residency work for Toronto teams?
You select a region and Plugsky keeps the data plane locked to it, so requests, logs and artefacts are not routed globally by default. Validate the region list and the contractual terms in the DPA and the SLA before you commit.
Which region should Toronto teams choose?
The closest published plane is the US plane (Virginia and Oregon), and the right choice depends on your jurisdiction and regulator rather than distance alone. Confirm the current region map at /data-residency.
Can we keep using the OpenAI SDK?
Yes. Plugsky exposes an OpenAI-compatible chat completions endpoint, so you change the base URL and model name and keep your existing SDK, prompts and evaluations.
Is there a free plan?
Yes — the free plan includes plugsky-micro and plugsky-lite with no credit card required. A 14-day full-access trial is available when you want to evaluate larger models.
How is pricing structured?
Self-serve plans are flat monthly with unlimited fair-use usage and no per-token billing. See the live pricing page for current plans and fair-use limits.
Which capabilities are live today?
Chat, streaming, JSON mode, function calling, embeddings, RAG and agents are live. Audio, images, moderation, files, batch, fine-tuning, assistants and responses are coming soon — check the docs before planning those workloads.
What about latency from Toronto?
It depends on the region you select and how your network routes to it. Measure candidate regions from your own environment, compare p50 and p95, and treat residency and latency as an explicit trade-off.