Key facts
| Text endpoints | Chat, streaming, JSON mode, function calling, embeddings, RAG and agents are live |
| Media endpoints | Image, audio and video endpoints are coming soon |
| Models | 30+ models in one catalogue, from free tiers to frontier reasoning |
| Pricing | Flat monthly self-serve plans with unlimited fair-use usage |
| Free tier | plugsky-micro and plugsky-lite, no card required |
| Deployment | Plugsky cloud, your VPC, on-prem and air-gapped |
| Positioning | Text and retrieval provider, not a media model marketplace |
TL;DR
- Replicate leads on media models; Plugsky leads on text and retrieval.
- Do not migrate image and video generation to a text API.
- Split by modality: Replicate for media, Plugsky for language.
- 30+ models and flat monthly pricing cover text workloads.
- Both expose simple HTTP APIs, so two providers stay manageable.
How it works, step by step
- Split your pipeline into media stages and language stages.
- Keep image, video and audio generation on Replicate.
- Move text stages: prompt rewriting, classification, summaries and embeddings, to Plugsky.
- Create a Plugsky key on the free plan and replay text-stage inputs.
- Verify that retrieval and generation share one embedding model consistently.
- Document the boundary so future features do not cross modalities by accident.
Try it yourself
Open the Replicate API cost calculator →
What Replicate does well
Replicate turned model hosting into a commodity: package a model, publish it, call it over HTTP. Its strength is breadth across modalities, especially image and video generation, and the community ecosystem around it. For creative tooling and media pipelines, that breadth is hard to replace.
Language workloads are a different shape. They need long-lived sessions, streaming, tool calling, embeddings and access to private corpora, and they are usually served better by a provider whose contract is designed around text and retrieval.
Where a text API fits
Plugsky exposes an OpenAI-compatible endpoint with 30+ models, live embeddings and agent primitives. A typical product pipeline mixes both providers: Replicate renders the image, Plugsky writes and checks the copy around it, and embeddings power search over generated assets and documentation.
- Keep media generation where the models live.
- Consolidate language stages behind one OpenAI-compatible key.
- Use embeddings consistently across text and metadata.
- Deploy privately where language data is sensitive.
- Keep API keys scoped per provider and per environment.
Endpoint status and hybrid architecture
Be precise about coverage. Plugsky does not ship image, audio or video endpoints yet; those are coming soon. Chat, streaming, JSON mode, function calling, embeddings, RAG and agents are live. That makes the boundary between providers a feature, not a compromise. Each side stays optimisable without rewriting the other.
The hybrid architecture is straightforward: two API clients, clear modality routing and separate spend lines. Flat monthly self-serve pricing on the Plugsky side keeps the language budget predictable. See the live pricing page for current plans and use the free tier to test.
Honest comparison
| Capability | Plugsky | Replicate | Self-hosting media models |
|---|---|---|---|
| Primary focus | Text, embeddings and agents | Media and community models | Whatever you deploy |
| Media endpoints | Coming soon | Broad image and video coverage | You run the GPUs |
| API style | OpenAI-compatible | HTTP model endpoints | Runtime-specific |
| Pricing shape | Flat monthly self-serve | Usage-based | GPU plus ops cost |
| Deployment | Cloud, VPC, on-prem, air-gapped | Managed platform | Your infrastructure |
Frequently asked questions
Can Plugsky replace Replicate?
Not for image, video or audio generation, which Plugsky does not ship yet. It replaces language workloads: chat, tools, embeddings, RAG and agents.
Should I move my whole pipeline to one provider?
No. Keep media generation on Replicate and consolidate text stages on Plugsky. Modality-specific providers remain the better choice for media today.
Is there a free plan?
Yes. plugsky-micro and plugsky-lite are free with no credit card, and a 14-day full-access trial opens stronger models.
How is pricing structured?
Self-serve plans are flat monthly with unlimited fair-use usage and no per-token billing. See the live pricing page for current plans.
Can Plugsky generate embeddings for my media metadata?
Yes. The embeddings API is live and works for text metadata, captions and descriptions, which is useful for search across generated assets.
Does Plugsky support private deployment?
Yes. VPC, on-prem and air-gapped deployments are available for language workloads with residency or security requirements.