Key facts
| Install | Helm chart and Terraform module for AWS, Azure and GCP |
| Network | Private endpoint inside your VPC or VNet, no public ingress |
| Keys | Customer-managed KMS (AWS KMS, Azure Key Vault, GCP KMS) |
| Models | Bring your own model registry; weights stay in your bucket |
| Observability | Logs, metrics and traces stay in your SIEM |
| Updates | Frequent model updates and periodic control-plane releases as signed bundles |
| Operating model | You run the infrastructure; Plugsky maintains the software |
| Product status | Live |
TL;DR
- The full control plane runs in your own cloud account.
- No public ingress — the API is reachable only from your network.
- Your KMS keys and your model registry; weights stay in your bucket.
- Observability data lands in your SIEM, not a vendor dashboard.
- You operate the infrastructure; Plugsky maintains the platform software.
How it works, step by step
- Confirm which cloud account and region will host the deployment.
- Prepare networking: private subnets, endpoints and no public ingress.
- Create KMS keys and a bucket for model weights and artifacts.
- Install the Helm chart or apply the Terraform module.
- Connect identity, logging and monitoring to your existing systems.
- Run an acceptance test, then enable signed update delivery.
Try it yourself
Open the private LLM cost estimator →
How the deployment works
Plugsky ships infrastructure as code: a Helm chart for Kubernetes clusters and a Terraform module for AWS, Azure and GCP. The install places the control plane inside your account, so model serving, routing and the API gateway run on infrastructure you own and pay for directly. A typical configuration sets the region, points at a KMS key and keeps ingress internal.
Because the endpoint is private, clients reach it through your VPC network — VPN, private link or internal load balancer — and the API is not exposed to the public internet.
What stays in your account
- The OpenAI-compatible API on a private endpoint with no public ingress.
- Customer-managed KMS encryption for data at rest.
- Your own model registry; weights remain in your bucket.
- Logs, metrics and traces exported to your SIEM and monitoring stack.
This is the topology banks and regulated SaaS teams usually choose: the data plane and keys are theirs, while the platform's software maintenance stays with Plugsky. It also fits teams that already have strong cloud governance and want AI to inherit it rather than bypass it.
Updates and the operating model
You operate the infrastructure — capacity, networking, upgrades of the underlying cluster — while Plugsky maintains the platform software. Model updates and control-plane releases arrive as signed bundles on a regular cadence, so you gain new models without rebuilding the stack. Plan a maintenance window and rollback path for each release; even automated updates deserve a rehearsal in a staging account.
Document who holds KMS permissions, who approves upgrades and who owns on-call. VPC deployments fail at handover points more often than at install.
When VPC is the right choice
Choose VPC deployment when data must not leave your cloud account, when existing security controls must apply unchanged, or when procurement requires customer-managed keys and private networking. Choose region-locked Plugsky cloud when speed to production matters more and residency can be satisfied architecturally. Choose on-prem or air-gapped when even a cloud account is out of bounds.
Honest trade-off: VPC deployment moves operational responsibility to your team. If you do not want that responsibility, a managed private deployment or Managed AI Ops is the better fit.
Honest comparison
| Factor | VPC deployment | Region-locked cloud | On-prem or air-gapped |
|---|---|---|---|
| Data location | Your cloud account | Plugsky region | Your data center |
| Public ingress | None | Public API endpoint | None |
| Key custody | Your KMS | Vendor-managed or BYOK on Enterprise | Your HSM |
| Updates | Signed bundles you apply | Managed by Plugsky | Physical media |
| Ops owned by | You (infra), Plugsky (software) | Plugsky | You |
| Time to production | Days to weeks | Minutes | Weeks to months |
Frequently asked questions
Which clouds are supported?
AWS, Azure and GCP, using the provided Helm chart and Terraform module.
Is the API exposed to the internet?
No. The deployment uses a private endpoint with no public ingress; access is through your own network controls.
Who holds the encryption keys?
You do. Customer-managed KMS keys are used for encryption at rest, and model weights stay in your own storage bucket.
How do updates work?
Model and control-plane updates are delivered as signed bundles that you apply on your maintenance cadence, with rollback to the previous release.
Who operates the deployment?
Your team operates the infrastructure and networking. Plugsky maintains the platform software and provides support under an Enterprise agreement.
Can I keep my observability stack?
Yes. Logs, metrics and traces export to your SIEM, OpenTelemetry collector or monitoring platform.
How does billing work?
You pay your cloud provider for infrastructure directly, and Plugsky for the platform under an Enterprise contract; see the live pricing page for self-serve plan context.
Plugsky (2026). “VPC AI Deployment — Plugsky Endpoint in Your Cloud”. Plugsky. Available at: https://plugsky.com/solutions/vpc-ai-deployment (last updated 2026-09-25).