Key facts
| Provider | Google — Vertex AI serves Gemini and partner models with MLOps tooling |
| API style | Google Cloud project-scoped APIs with IAM auth and regional endpoints |
| Plugsky API | OpenAI-compatible /v1/chat/completions — change the base URL, keep your SDK |
| Models | 30+ models from free to frontier behind one API key |
| Pricing | Flat monthly plans with unlimited fair-use usage; no per-token billing on self-serve |
| Free tier | Free plan with plugsky-micro and plugsky-lite, no card; 14-day full-access trial |
| Deployment | Plugsky cloud, your VPC, on-prem or air-gapped; region choice for residency |
| Live vs roadmap | Chat, streaming, JSON mode, function calling, embeddings, RAG, agents live; audio, images, moderation, files, batch, fine-tuning, assistants, responses coming soon |
TL;DR
- Vertex AI is a platform; Plugsky is a focused inference API.
- Vertex brings pipelines, registry and governance; Plugsky brings portability and flat pricing.
- Both can coexist — inference on Plugsky, pipelines on Vertex.
- Plugsky free tier: plugsky-micro and plugsky-lite; 14-day full-access trial.
- Honest trade-off: Plugsky does not replicate Vertex's data and MLOps ecosystem.
How it works, step by step
- List which Vertex capabilities are inference and which are platform (pipelines, registry, tuning).
- Create a Plugsky account and test the inference workloads on your eval set.
- Replace inference SDK calls with an OpenAI-compatible client.
- Re-test structured output, tools and long-context jobs for behavioural differences.
- Keep Vertex for pipeline and data work that depends on Google Cloud.
- Track combined spend so the split remains rational as usage grows.
Original data
Try it yourself
Open the Google Vertex AI cost calculator →
What the Vertex AI API provides
Vertex AI is designed for platform teams. Projects, IAM roles and service accounts govern access; endpoints are regional; and the surrounding services — pipelines, model registry, feature tooling — support the full model lifecycle. Model Garden adds Gemini and third-party options inside the same governance boundary.
That completeness costs simplicity. Getting a first chat call working takes more ceremony than an API key, model access varies by region, and billing is a blend of platform and usage charges that needs active management.
What the Plugsky API provides
Plugsky strips inference down to the essentials. One key, one OpenAI-compatible endpoint, 30+ models as parameters. The free plan covers plugsky-micro and plugsky-lite, and a 14-day full-access trial covers paid models.
Self-serve pricing is flat monthly with unlimited fair-use usage (live pricing), so cost does not spike with a successful launch. Enterprise deployment extends to your VPC, on-prem or air-gapped with region selection. Chat, streaming, JSON mode, function calling, embeddings, RAG and agents are live; audio, images, moderation, files, batch, fine-tuning, assistants and responses are coming soon.
Designing the split
The cleanest architectures separate the model lifecycle from the inference path.
- Train, tune and evaluate on the platform your data already lives in.
- Serve requests through an OpenAI-compatible API that is portable across models.
- Keep evaluation data and prompts versioned outside any single vendor.
- Document which workloads are allowed in which deployment (cloud, VPC, on-prem, air-gapped).
Document the boundary explicitly. A one-page architecture note that names which workloads run where prevents accidental coupling and keeps future migrations cheap. Do it before you migrate, not after.
Honest comparison
| Capability | Plugsky | Vertex AI | Building in-house |
|---|---|---|---|
| API style | OpenAI-compatible drop-in | Project-scoped Google Cloud APIs with IAM | You define the schema |
| Scope | Inference API plus deployment options | Full ML platform with MLOps tooling | You build everything |
| Model access | 30+ models, one key | Gemini and Model Garden partners | You host each model |
| Billing | Flat monthly, unlimited fair use (see live pricing) | Usage-based plus platform charges | GPU + ops cost |
| Residency | Region choice, VPC, on-prem, air-gapped | Google Cloud regions and controls | You control the infrastructure |
| Honest gap | Pipelines, registry and data tooling | Deep MLOps and Google data integration | You build it |
Frequently asked questions
Is Vertex AI just an API?
No. Vertex AI is a platform: model access, pipelines, a model registry and governance integrations all live under the same Google Cloud project model.
Why move inference off Vertex?
Simplicity and portability. Product teams often want an API key and an OpenAI-compatible client rather than project-scoped SDKs and regional model availability.
Can I keep Vertex for training and use Plugsky for serving?
Yes. That split is common: heavy lifecycle work stays on the cloud platform, while request-serving moves to a portable, flat-priced API.
Is there a free plan on Plugsky?
Yes — plugsky-micro and plugsky-lite are free with no credit card, and a 14-day full-access trial is available.
How is Plugsky billed?
Flat monthly self-serve plans with unlimited fair-use usage and no per-token billing. See the live pricing page.
Does Plugsky support private networking?
Enterprise deployments support VPC, on-prem and air-gapped options with region selection.
Which should handle regulated data?
Verify residency and controls for both paths. Plugsky offers region choice and self-managed deployment; Vertex provides Google Cloud's regional and governance controls.