Key facts
| API compatibility | OpenAI-compatible chat, embeddings and function calling |
| Deployment | VPC in AWS, Azure or GCP via Helm chart and Terraform module |
| Network | Private endpoint with no public ingress |
| Encryption and keys | Customer-managed KMS (AWS KMS, Azure Key Vault, GCP KMS) |
| Models | 30+ models; bring your own model registry and bucket |
| Observability | Logs, metrics and traces stay in your SIEM |
| Updates | Signed model and control-plane bundles applied on your cadence |
| Pricing | Flat monthly self-serve plans; Enterprise terms for VPC |
TL;DR
- The full control plane runs in your AWS, Azure or GCP account.
- No public ingress: the API is reachable only from your network.
- Your KMS keys and your model registry; weights stay in your bucket.
- AI features ship inside your product with per-tenant limits and cost you can forecast.
- Region-locked managed planes remain available when you do not need a VPC.
How it works, step by step
- Confirm which workloads move first and which must stay in a dedicated deployment.
- Prepare networking: private subnets, endpoints and no public ingress.
- Create KMS keys and a bucket for model weights and artifacts.
- Apply the Helm chart or Terraform module for AWS, Azure or GCP.
- Connect identity, logging and monitoring to your existing systems.
- Run acceptance tests and replay your prompts and evaluations.
- Enable signed update delivery and cut over with a rollback path ready.
Try it yourself
Open the self-hosting break-even calculator →
Why SaaS teams use a VPC deployment
SaaS teams ship AI features inside a product that already has tenants, billing and an uptime promise. The platform decision is really a multi-tenancy decision: per-tenant isolation, metering, quotas and predictable cost.
For SaaS teams, the deciding question is where the data sits: tenant data inside a multi-tenant product. A shared public endpoint pushes that decision into contracts and configuration; a VPC deployment puts the API inside the network boundary you already control, with your subnets, your identity provider and your logging. That is the difference between trusting a vendor's region and knowing the traffic never leaves your account.
What the Plugsky VPC architecture includes
Plugsky ships a Helm chart and Terraform module for AWS, Azure and GCP. The control plane installs inside your cloud account and the OpenAI-compatible API is exposed on a private endpoint with no public ingress. You supply a customer-managed KMS key for encryption at rest, and model weights live in your own bucket and registry. Logs, metrics and traces stay in your SIEM, and model and control-plane updates arrive as signed bundles you apply on your maintenance cadence with rollback.
When a full rollout is more than the workload needs, the same API runs in region-locked Plugsky planes with residency guarantees. Where customers demand residency, map tenant tiers to regions or dedicated deployments instead of running one global endpoint for everyone.
Controls to wire in from day one
Controls worth wiring in from the first deployment:
- Keep one private endpoint per environment and isolate tenant context in your application layer.
- Set per-tenant limits so one heavy tenant cannot exhaust shared capacity.
- Meter usage per tenant for entitlements and billing.
- Keep the OpenAI-compatible surface so feature code does not depend on one vendor SDK.
Deploy, measure, then cut over
AI features ship inside your product with per-tenant limits and cost you can forecast. The sequence is reversible: prepare the account and network, create keys and a weights bucket, apply the module, connect identity and observability, replay evaluations, then switch the base URL behind a flag. Latency is a trade-off: closer planes shorten the path, but strict boundaries can add round-trip time, so measure p50 and p95 from your own environment.
Chat, streaming, JSON mode, function calling, embeddings, RAG and agents are live; audio, images, moderation, files, batch, fine-tuning, assistants and responses are coming soon, so label those early. Self-serve plans are flat monthly with unlimited fair-use usage and no per-token billing; current plans are on the pricing page. Start with the free plan, then the 14-day full-access trial.
Honest comparison
| Factor | Plugsky VPC deployment | Region-locked Plugsky cloud | Building in-house |
|---|---|---|---|
| Data location | Your cloud account | Plugsky region you pin | Your data centre |
| Public ingress | None — private endpoint | Public API endpoint | Depends on your build |
| Key custody | Your KMS (AWS, Azure, GCP) | Vendor-managed; BYOK on Enterprise | Your HSM |
| Models | 30+ models, bring your own registry | 30+ managed models | You host and maintain each model |
| Updates | Signed bundles on your cadence | Managed by Plugsky | Your release engineering |
| Time to production | Days to weeks | Minutes | Weeks to months |
Frequently asked questions
Does a VPC deployment work for multi-tenant SaaS?
Yes. Run the platform in your account, keep tenant isolation in your application layer, and use per-tenant keys and limits for metering and blast-radius control.
How do we meter usage per tenant?
Use per-tenant API keys and limits, record usage at your application boundary, and consume usage.threshold events to trigger entitlements and billing.
Can we keep using the OpenAI SDK?
Yes. Plugsky exposes an OpenAI-compatible chat completions endpoint, so you change the base URL and model name and keep your existing SDK code, prompts and evaluations.
Is there a free plan?
Yes — the free plan includes plugsky-micro and plugsky-lite with no credit card. A 14-day full-access trial is available when you want to evaluate larger models.
How is pricing structured?
Self-serve plans are flat monthly with unlimited fair-use usage and no per-token billing. See the live pricing page for current plans.
Which capabilities are live today?
Chat, streaming, JSON mode, function calling, embeddings, RAG and agents are live; audio, images, moderation, files, batch, fine-tuning, assistants and responses are coming soon — check the docs before planning those workloads.
How is a VPC deployment priced?
VPC runs as an Enterprise engagement with terms agreed with the team; self-serve flat monthly plans cover the managed plane. See the live pricing page for current plans.
Who operates the deployment?
You run the infrastructure; Plugsky maintains the platform software and delivers signed updates for you to apply.