Key facts
| Delivery model | Repeatable reference install per client estate |
| API continuity | One OpenAI-compatible integration across all clients and tiers |
| Tenancy | Workspace and scoped keys per client, with per-client audit exports |
| Deployment | On-prem and air-gapped tiers alongside cloud and VPC options |
| Updates | Standard bundle pipeline with checksums, staging and rollback |
| Key management | Customer-managed keys in the client's KMS or HSM |
| Audit | Access, key and inference events exported to client SIEM |
| Status | Live platform; enterprise deployment terms per engagement |
TL;DR
- Industrialise one reference install instead of bespoke builds per client.
- Operate a standard update bundle pipeline with staging and rollback.
- Keep per-client workspaces, keys and audit streams for isolation.
- Support remote estates with bundle-based diagnostics, not outbound access.
- Charge for the managed service, not the inference.
How it works, step by step
- Define a reference architecture: model set, GPU sizing bands, gateway, key integration and log path that cover most client estates.
- Qualify clients against it — data sensitivity, network classification, hardware — and identify the exceptions.
- Build a provisioning kit: install runbook, model inventory, evaluation suite and client audit export.
- Operate a shared update pipeline producing verified bundles for every client, with staging validation and rollback.
- Keep a workspace and scoped key set per client, even in single-tenant estates, so reporting and offboarding stay clean.
- Package service tiers: deployment, quarterly model refresh, evaluation reporting and incident support.
- Rehearse recovery and failover with each client before go-live, and record the evidence in your service file.
Try it yourself
Open the self-hosting break-even calculator →
Productising air-gapped delivery
Air-gapped projects usually fail on bespoke engineering, not on models. MSPs win by standardising: one reference architecture with defined model sets and GPU sizing bands, one provisioning kit, one update process. Clients then receive a known quantity with documented behaviour rather than a research project.
Plugsky helps because the API contract is stable across tiers. Your integration, evaluation harness and runbooks are written once and reused in cloud, VPC, on-prem and disconnected estates, so each new client is a deployment exercise rather than a development one.
Update bundles as a managed service
The recurring value in an air-gapped estate is the refresh cycle: new models, runtime fixes and security patches delivered safely. Build a pipeline that produces verified bundles with checksums and release notes, validates them on your own staging estate or the client's mirror, and installs with a rollback path.
- Schedule: quarterly model refreshes and emergency security releases as separate tracks.
- Validate: run the client evaluation suite against the candidate bundle before promotion.
- Document: every installation leaves an updated manifest and change record.
Support and isolation without egress
Remote support in a disconnected estate is a process, not a connection. Ask clients to export diagnostics, metrics and audit summaries; analyse offline; and return an installable fix. Keep per-client workspaces and keys so access and reporting stay separated even when you operate multiple estates.
On the commercial side, air-gapped work justifies recurring fees because it is genuinely continuous: capacity reviews, refresh cycles, evaluation reporting, incident response and audit support. Price those as service lines and be explicit about what a client's own team owns, from hardware to key custody.
Honest comparison
| MSP concern | Plugsky air-gapped | Re-selling cloud AI | Bespoke builds |
|---|---|---|---|
| Repeatability | One reference install and kit | Single hosted tenancy | Each client differs |
| Update delivery | Standard verified bundles | Provider-managed | Ad hoc per client |
| Isolation | Workspace and keys per client | Shared platform | Varies |
| Support model | Export-based remote diagnostics | Online access | Whatever you improvise |
| Margin shape | Recurring service fees | Thin resale margin | Project-based |
| API portability | OpenAI-compatible offline | Vendor API | You define it |
Frequently asked questions
Do we need a separate codebase per client?
No. Keep one integration against the OpenAI-compatible endpoint and vary workspaces, keys, models and hardware per client. That is what makes the service profitable.
How do we update remote estates?
With a standard bundle pipeline: checksums, release notes, staging validation against the client evaluation suite, then installation with a tested rollback path.
How do we support a network with no egress?
Through exported diagnostics and metrics, offline analysis and installable fixes. Agree the process in the contract; expect it to be slower than online support.
Who owns the keys?
The client does. Customer-managed keys live in their KMS or HSM, and key administration should be separated from platform administration.
How do we price managed air-gapped AI?
As recurring service fees covering deployment, refresh cycles, evaluation reporting, incident support and audit assistance. Avoid pricing on inference volume, which is not the value.
Can we use the same team for cloud and disconnected clients?
Yes, if the integration, evaluation suite and runbooks are shared. The interface is identical; only the deployment environment differs.
How do we win the first air-gapped client?
Start with a deployment on their staging network using open-weight models and a bounded workload, and prove the update and recovery processes. That evidence wins the larger rollout.