Key facts
| API compatibility | OpenAI-compatible /v1/chat/completions (drop-in base URL change) |
| Models | 30+ models in one catalogue, from free tiers to frontier reasoning |
| Pricing | Flat monthly self-serve plans with unlimited fair-use usage |
| Free tier | plugsky-micro and plugsky-lite, no card required |
| Trial | 14-day full-access trial for stronger models |
| Deployment | Plugsky cloud, your VPC, on-prem and air-gapped |
| Scope | Inference API, not a full ML platform or training service |
TL;DR
- Vertex is a full ML platform; Plugsky is an inference API. Scope them separately.
- Keep Vertex for training, pipelines, IAM-heavy workflows and Google services.
- Move portable inference to an OpenAI-compatible endpoint with flat pricing.
- Region selection plus VPC, on-prem and air-gapped options cover residency needs.
- Splitting platform from inference is usually cheaper than replacing the platform.
How it works, step by step
- List Vertex usage by category: training, pipelines, inference, embeddings, Google integrations.
- Mark the inference workloads that do not depend on Google-specific features.
- Create a Plugsky key on the free plan and replay recorded prompts.
- Move OpenAI-style call sites to the Plugsky base URL and model names.
- Compare quality, latency and operational cost on a full traffic cycle.
- Keep Vertex for platform work and expand the inference split gradually.
Original data
Try it yourself
Open the Google Vertex AI cost calculator →
Vertex AI is a platform, not just a model endpoint
Vertex AI covers training, tuning, pipelines, model registry and deployment alongside Gemini and partner models. That breadth is why enterprises choose it, and it is also why the bill and the operational surface grow: every capability has to be learned, governed and paid for.
If a team only needs a reliable chat, tools and embeddings endpoint, most of that platform sits idle. Separating the inference layer from the training platform is a common architectural simplification, not a rejection of Google Cloud.
Splitting platform from inference
Plugsky speaks the OpenAI chat completions schema, which means portable inference workloads move with a base URL and model-name change. Anything genuinely tied to Google — BigQuery, Pub/Sub, Vertex Pipelines, custom training — stays where it is.
- One key for chat, embeddings, RAG and agents.
- Flat monthly self-serve pricing instead of per-token accounting.
- Free plan with plugsky-micro and plugsky-lite, no card.
- 14-day full-access trial for frontier models.
Governance, residency and gaps
Plugsky supports region selection and deployment in your VPC, on-prem or air-gapped environments, which helps when data must stay within a jurisdiction or a private network. It does not replace the Google Cloud control plane: IAM roles, audit logging and org policy remain Google's, and Plugsky is not a Google service.
Endpoint coverage is also narrower: audio, images, moderation, files, batch, fine-tuning, assistants and responses are coming soon. Chat, streaming, tools, embeddings and agents are live today. Keep platform and inference decisions independent so each can change without forcing the other. Start free, validate on real prompts, and see the live pricing page for current plans.
Honest comparison
| Capability | Plugsky | Google Vertex AI | Self-managed on GCP |
|---|---|---|---|
| Scope | Inference API | Full ML platform | Whatever you build |
| API style | OpenAI-compatible | Google SDKs and REST | Runtime-specific |
| Pricing shape | Flat monthly self-serve | Usage-based | VM and GPU cost plus ops |
| Deployment | Cloud, VPC, on-prem, air-gapped | Google Cloud regions | GKE or custom VMs |
| Platform features | Not a training platform | Training, pipelines, registry | You assemble tools |
Frequently asked questions
Is Plugsky a full replacement for Vertex AI?
No. Vertex is a platform covering training, pipelines and governance. Plugsky replaces the inference endpoint for portable text workloads; keep Vertex for platform work.
Can I move inference only?
Yes, and that is the recommended scope. Move chat, tool and embedding calls that do not depend on Google-specific features and leave the rest untouched.
Is there a free plan?
Yes. The free plan includes plugsky-micro and plugsky-lite with no credit card, and a 14-day full-access trial covers stronger models.
How is pricing structured?
Self-serve plans are flat monthly with unlimited fair-use usage and no per-token billing. See the live pricing page for current plans.
Does Plugsky meet residency requirements?
It supports region selection and VPC, on-prem or air-gapped deployment. Validate the exact data plane your policy requires during evaluation.
What endpoints are coming soon?
Audio, images, moderation, files, batch, fine-tuning, assistants and responses. Chat, streaming, JSON mode, function calling, embeddings, RAG and agents are live.
Can I keep model training on Vertex?
Yes. Training and pipelines remain on Google Cloud; only inference moves. The two layers do not conflict.