Key facts
| API surface | OpenAI-compatible /v1/chat/completions; change base_url and model name |
| Deployment | Plugsky cloud, your VPC, on-prem or air-gapped |
| Data residency | Region selection and sovereign deployment options |
| Models | 30+ models behind one API, including long-context and multilingual options |
| Access control | Scoped API keys and request logs; enterprise RBAC and SSO |
| Pricing | Flat monthly self-serve plans; see the live pricing page |
| Free tier | Free plan with 2 free AI models (plugsky-micro, plugsky-lite), no card |
| Contracting | DPA and SLA documents available for review |
TL;DR
- Start free without a card, then keep the same API as you scale: no re-platforming when enterprise deals arrive.
- Evaluate residency, cost predictability, auditability, model choice and exit before you commit.
- 30+ models behind one OpenAI-compatible API, so you can trade quality against cost per feature.
- Use the 14-day full-access trial to test frontier models on your real workload before you commit.
- Design for the compliance review you will face later: scoped keys, logs and an exit path from day one.
How it works, step by step
- Map where customer data, product metrics and investor material live today and which services are in scope.
- Write down residency, retention and access requirements for your target market and enterprise buyers.
- Shortlist deployment modes: managed cloud now, with VPC, on-prem or air-gapped as customer requirements grow.
- Run a pilot on the free plan or the 14-day full-access trial with a fixed set of real product questions.
- Score answers, citations, refusals, latency and cost on the same test.
- Keep prompts, evaluation sets and keys organized so due diligence and security reviews are quick.
- Re-run the evaluation as you change models or move workloads to a private deployment.
Original data
Try it yourself
Open the private LLM cost estimator →
What sovereign AI means for startups
Sovereign AI for startups means the models, prompts and retrieved data stay under your jurisdiction and operational control as you grow. In practice that is a deployment you can place in your own region or network, with identity, logging and retention governed by controls you can describe to buyers and investors.
The business case is not secrecy for its own sake: it is keeping customer records, product analytics and investor material inside a boundary you can audit, while still using modern models through a standard API.
The five things to evaluate
Score every candidate on these five areas, and require evidence rather than claims:
- Residency and deployment: region choice, and whether VPC, on-prem and air-gapped options actually exist.
- Data handling: whether prompts and outputs are used for training, how long they are retained, and which subprocessors are involved.
- Access and audit: scoped API keys, RBAC and SSO for enterprise, and request-level logs you can export.
- Model choice and portability: how many models you can route between, and how hard it is to move away.
- Commercials and exit: predictable pricing, SLA terms, and a documented data-return path.
Deployment modes and their trade-offs
Cloud is fastest: a managed, region-based deployment with residency selection. VPC puts inference inside your own cloud account and network boundary. On-prem runs in your data center; air-gapped runs with no outbound connectivity at all. The trade-off is operational: the further you move from managed cloud, the more you own upgrades, capacity and monitoring.
Startups usually start on managed cloud because it is fast and needs no infrastructure, then move the same OpenAI-compatible workload to a private deployment when an enterprise contract or a regulated market demands it. Because the API surface does not change, application code usually stays put.
A scorecard approach to procurement
Run a structured evaluation: define ten questions that represent real product work, test two or three models, and score answers, citations and refusals. Add a security review covering identity, key rotation, logging and data flow, then a commercial review covering pricing predictability, SLA and exit.
Begin on the free plan or the 14-day full-access trial, document the decision, and keep the evaluation set so you can re-run it after every model or deployment change.
Honest comparison
| Evaluation area | Plugsky | Vendor cloud-only | Building in-house |
|---|---|---|---|
| Residency | Region selection plus sovereign deployment options | Usually vendor regions only | Fully under your control |
| Deployment modes | Cloud, VPC, on-prem, air-gapped | Managed cloud only | Your infrastructure |
| Model choice | 30+ models behind one API | Single vendor catalogue | You host each model |
| Auditability | Scoped keys and request logs; enterprise RBAC and SSO | Vendor-defined logging | You build logging |
| Pricing | Flat monthly self-serve plans; see live pricing | Often per-token or per-seat | GPU plus operations cost |
| Time to production | Days on managed cloud; private options when enterprise deals require them | Fast but fixed residency | Months of build and ops |
Frequently asked questions
What makes an AI deployment sovereign?
Sovereign deployment keeps models, prompts and data under your jurisdiction and control: you choose the region, own the boundary, and keep identity, logging and retention under your existing governance.
Can we keep data inside our own country or network?
Yes. Plugsky supports region selection plus VPC, on-prem and air-gapped deployments, so inference and data can sit where your customers or regulators require.
Does Plugsky train on our data?
For strict requirements, choose a private or air-gapped deployment so content stays inside your environment. Review the current data-handling terms and DPA for cloud plans before rollout.
Can we start free and add a private deployment later?
Yes. Begin on the free plan with two free AI models and no card, or use the 14-day full-access trial, then move the same OpenAI-compatible workload when a customer or market requires it.
How do we avoid lock-in?
Plugsky is OpenAI-compatible: applications use a standard API, so you can change base_url and model name with minimal code changes, and keep your evaluation set to compare alternatives.
How do we keep AI costs predictable on a small budget?
Self-serve plans are flat monthly with fair-use usage rather than per-token billing, so you can forecast spend while you tune model choice per feature. See the live pricing page.
What should we prepare for enterprise security reviews?
Have scoped API keys, request logs, data-flow documentation and DPA and SLA documents ready, plus a documented path from managed cloud to VPC or on-prem.