Key facts
| Model family | Llama is an open-weight family released by Meta |
| Licence | Llama Community Licence with acceptable-use terms |
| Hosting options | First-party preview, cloud marketplaces, serverless platforms, self-hosting |
| API style | Most hosted Llama endpoints are OpenAI-compatible chat completions |
| Version risk | Builds, quantisations and context limits vary by provider |
| Plugsky catalogue | Llama-family models served with other families under one API |
| Pricing model | Per-token on most hosts; flat monthly plans on Plugsky |
| Deployment | Plugsky cloud, VPC, on-prem or air-gapped |
TL;DR
- Llama is open-weight, so the real choice is where and how you serve it.
- First-party, hyperscaler, serverless and self-hosted options trade control against ops.
- Pin exact model versions: hosted builds and quantisations differ by provider.
- Plugsky serves Llama-family models under one OpenAI-compatible API.
- Start free with two models, then scale to a paid plan or private deployment.
How it works, step by step
- Write down the Llama variant, size and context length your workload needs.
- List candidate hosts and check the exact build, quantisation and context each serves.
- Benchmark your own prompts on two hosts rather than trusting generic scores.
- Check licence terms, region and data-handling policy for the chosen host.
- Decide between per-token hosting and a flat monthly plan based on your volume.
- Keep the model name in configuration so you can switch hosts without code changes.
Try it yourself
Open the LLM API cost calculator →
Why Llama is not a single API
Llama is a model family, not a service. Meta publishes the weights, and anyone can host them: cloud marketplaces package them as managed endpoints, serverless platforms rent them by the token, and companies with GPUs run them directly. The same model name can mean different builds, quantisations and context limits depending on the host.
That freedom is the appeal and the problem. You avoid single-vendor lock-in, but you inherit version drift, and output quality can differ between two endpoints that both claim to serve the same Llama model.
The ways to call Llama compared
Each route makes sense for a different team. A first-party preview endpoint is the easiest way to evaluate new releases. A hyperscaler marketplace wins when your infrastructure, identity and billing already live on that cloud. Serverless inference platforms are fast to adopt and bill per token. Self-hosting gives full control but makes you responsible for GPUs, scaling and upgrades.
A managed multi-model API sits between those extremes. Plugsky serves Llama-family models alongside 30+ other models behind one OpenAI-compatible endpoint, with flat monthly self-serve plans and a free plan covering plugsky-micro and plugsky-lite. Enterprise options add VPC, on-prem and air-gapped deployment, and current plan details are on the live pricing page.
Choosing for production
Pin an exact model version and context limit, then evaluate it on your own prompts. Treat published throughput and quality numbers as directional, because serving hardware and engine configuration change the result. Confirm the licence terms that apply to your scale, and check where inference happens if residency matters.
Finally, keep the model identifier in configuration so a future move is a setting rather than a rewrite. Where Plugsky does not replace another option yet: if you depend on a brand-new Llama release on day one, a first-party or hyperscaler endpoint may carry it sooner — verify the live catalogue before committing.
Honest comparison
| Option | Control | Ops burden | Typical fit |
|---|---|---|---|
| Meta first-party preview | Limited to offered endpoints | None | Early evaluation |
| Hyperscaler marketplace | Region and IAM controls | Managed but platform-bound | Enterprises already on that cloud |
| Serverless inference platform | Model and parameter choice | Low; per-token billing | Product teams shipping fast |
| Self-hosted on your GPUs | Full control of build and weights | High: GPUs, scaling, upgrades | Strict control or steady high load |
| Plugsky managed API | Model choice across families | None: one endpoint and plan | Teams wanting one API and flat pricing |
Frequently asked questions
Is there an official Llama API?
Meta has offered a first-party Llama API as a preview, but availability and model coverage change. Most teams use a cloud provider or inference platform instead, and the open weights can always be self-hosted.
Are Llama models free to use?
The weights are free to download under the Llama Community Licence, which includes acceptable-use terms and extra conditions for very large platforms. Hosting and serving costs still apply.
Which host serves Llama fastest?
Performance depends on hardware, quantisation and serving engine, so measure with your own prompts. Treat published throughput numbers as directional only.
Can I switch Llama hosts without changing code?
Often yes if both expose OpenAI-compatible endpoints: change the base URL and model ID. Keep the model name in configuration and pin an exact version.
Does Plugsky serve Llama models?
Yes. Llama-family models are part of the catalogue alongside other families, all behind one OpenAI-compatible API. Check the live catalogue for current versions and context limits.
Is pricing per token on Plugsky?
Self-serve Plugsky plans are flat monthly with unlimited fair use on paid tiers rather than per-token billing. See the live pricing page for current plans.
Can I run a fine-tuned Llama on Plugsky?
Fine-tuning endpoints are coming soon. Until then, run fine-tuned weights on your own infrastructure or a host that accepts custom weights.