Key facts
| Preparation | Download runtimes, models and dependencies on a connected staging machine |
| Model format | GGUF files are single-file and easy to transfer and verify |
| Transfer | Use approved media and scan before it enters the isolated network |
| Verification | Check checksums and record model revisions |
| Runtime | llama.cpp and Ollama run fully offline once models are present |
| Updates | Plan a recurring refresh for models, runtimes and security fixes |
| Cloud option | Plugsky supports air-gapped deployment for enterprise teams |
| Endpoint status | Chat, streaming, tools, JSON mode and embeddings are live |
TL;DR
- Build an offline bundle on a connected machine, then transfer it in.
- Single-file GGUF models are the simplest artefacts to move and verify.
- Checksum everything and record the exact model revision.
- Air-gapped stacks need a scheduled update process or they drift.
- For enterprise air-gapped deployment, Plugsky offers an on-prem option.
How it works, step by step
- List the runtimes, models and dependencies the isolated host needs.
- Download everything on a connected staging machine and record versions.
- Verify checksums and scan the bundle before transfer.
- Move the bundle by approved media and install on the isolated host.
- Test the model with representative prompts and confirm no network calls.
- Document the exact versions in use for reproducibility.
- Schedule a recurring offline refresh for models and security updates.
Try it yourself
Open the self-hosting requirements checker →
Building an offline bundle
An offline deployment starts on a connected machine. Install the runtime there, download the model files, and collect everything the isolated host needs: the runtime package, model weights, tokenizer files, configuration, and any container images or system libraries. Version-pin each item and write the list down.
Model format matters for portability. Single-file GGUF models are the easiest to transfer and verify, and llama.cpp-family runtimes consume them directly. If you plan to serve with a GPU stack, include its own model format and the exact engine build, because those stacks are less tolerant of version drift.
Verifying and installing in the isolated network
Treat the bundle as untrusted until verified. Compare hashes against the values you recorded at download time, scan media before it crosses the boundary, and keep a copy of the verified bundle in an internal repository for future hosts. Provenance records matter during audits: know which model revision is running and where it came from.
After installation, test with representative prompts rather than a single greeting. Confirm that generation works with networking disabled, and check that the runtime is not silently trying to fetch models or check for updates. Disable telemetry and any online features the deployment does not need.
Keeping an air-gapped stack alive
Connectivity is a one-time problem; currency is ongoing. Models improve, runtimes ship security fixes, and tokenizer or licence terms change. Without a refresh process, the isolated stack slowly becomes the least-maintained system you run.
Set a cadence: a recurring security review of the runtime, a model refresh when a validated improvement is worth the transfer, and a documented rollback path. For organisations that want managed inference inside their own boundary rather than self-run tooling, Plugsky offers VPC, on-prem and air-gapped deployment with published terms and SLA. Chat, streaming, tools, JSON mode and embeddings are live; see pricing for plan details.
Honest comparison
| Concern | Air-gapped local | Plugsky air-gapped deployment | Check before deciding |
|---|---|---|---|
| Connectivity | No internet by design | No internet inside your environment | Policy and physical controls |
| Model updates | You carry them in | Vendor-supported update path | Refresh cadence |
| Hardware | You own and maintain it | Deployed in your infrastructure | Capacity and uptime needs |
| Support | Internal team only | Enterprise support and SLA | Escalation expectations |
| Time to first model | Days to weeks | Project-based rollout | Deadline and staffing |
Frequently asked questions
Can Ollama or llama.cpp work fully offline?
Yes. After the runtime and model files are installed, generation runs locally. Only features that fetch models or check for updates need connectivity.
How do I transfer large model files?
Use approved removable media or an internal artefact repository, verify checksums on both sides, and keep a copy of the exact revision for reproducibility.
How often should air-gapped models be updated?
Set a cadence that matches your risk tolerance, such as quarterly security refreshes plus model updates when a validated improvement justifies the transfer.
Can I use a cloud API in an air-gapped environment?
Not directly. The usual alternative is a vendor-deployed private instance inside your network, which is what Plugsky offers for enterprise air-gapped deployments.
What about licensing offline?
Check that each model licence permits your use and internal redistribution, and keep licence records alongside the model artefacts.
Do embeddings and RAG work offline?
Yes, if the embedding model and vector store run locally. The full retrieval and generation loop can operate with no network access.
What is the hardest part of air-gapped AI?
Keeping the stack current. Transfer and installation are one-time problems; model, runtime and security updates are recurring and need an owner.