Key facts
| Provider | Helicone — an LLM gateway and observability API that proxies provider traffic |
| Category | Logging, caching, budgets and analytics rather than model inference |
| Plugsky API | OpenAI-compatible /v1/chat/completions — change the base URL, keep your SDK |
| Models | 30+ models from free to frontier behind one API key |
| Pricing | Flat monthly plans with unlimited fair-use usage; no per-token billing on self-serve |
| Free tier | Free plan with plugsky-micro and plugsky-lite, no card; 14-day full-access trial |
| Deployment | Plugsky cloud, your VPC, on-prem or air-gapped; region choice for residency |
| Live vs roadmap | Chat, streaming, JSON mode, function calling, embeddings, RAG, agents live; audio, images, moderation, files, batch, fine-tuning, assistants, responses coming soon |
TL;DR
- Helicone instruments the path; Plugsky is the destination serving models.
- Plugsky offers 30+ models and dashboard analytics with flat monthly pricing.
- A gateway can sit in front of Plugsky because Plugsky is OpenAI-compatible.
- Free tier on Plugsky: plugsky-micro and plugsky-lite; 14-day full-access trial.
- Honest trade-off: prompt-level tracing across many vendors is Helicone's speciality.
How it works, step by step
- Clarify whether your immediate need is observability or model access.
- If both, design the request path: application, gateway, model platform.
- Create a Plugsky account and choose a model for the workload you are instrumenting.
- Configure the gateway with an OpenAI-compatible provider pointed at Plugsky.
- Verify that logs record what you need: latency, token counts, errors, model IDs.
- Review retention and residency for observability data before scaling traffic.
Original data
Try it yourself
Open the Helicone cost calculator →
What a gateway layer like Helicone provides
A gateway earns its place by being in the path. Helicone proxies OpenAI-compatible requests, so it can record prompts and completions, cache identical calls, enforce budgets and rate limits, and aggregate spend across providers. If your architecture spans several model vendors, that cross-cutting view is difficult to build well in-house.
The costs are the usual ones for an in-path component: an extra network hop, another system handling sensitive prompt data, and a dependency that must be monitored like any other production service.
What the model platform provides
Plugsky serves the inference itself. One OpenAI-compatible API reaches 30+ models, with a free plan (plugsky-micro and plugsky-lite) and a 14-day full-access trial. Self-serve plans are flat monthly with unlimited fair-use usage (live pricing), which reduces the cost-visibility problem gateways are often brought in to solve.
The dashboard provides usage analytics, and enterprise deployments can run in your VPC, on-prem or air-gapped with region selection. Chat, streaming, JSON mode, function calling, embeddings, RAG and agents are live; audio, images, moderation, files, batch, fine-tuning, assistants and responses are coming soon.
Composing the two correctly
There is no conflict between a gateway and a model platform; the order matters for latency and data handling.
- Application → gateway → Plugsky is the standard arrangement.
- Keep the gateway provider-agnostic so you can fail over between endpoints.
- Decide what must not be logged; prompt content may be regulated.
- Re-evaluate annually: more built-in analytics may make the extra hop unnecessary.
Also verify that the gateway preserves streaming and function-calling behaviour, since in-path components can subtly change those paths.
Honest comparison
| Capability | Plugsky | Helicone | Building in-house |
|---|---|---|---|
| Category | Model platform and API | Gateway and observability layer | You build both |
| Function | Serves chat, embeddings, agents | Logs, caches, limits and measures traffic | You instrument manually |
| Catalogue | 30+ models, one key | No models; proxies providers | You integrate each vendor |
| Pricing | Flat monthly, unlimited fair use (see live pricing) | Vendor plan for gateway usage | Engineering time |
| Deployment | Cloud, VPC, on-prem, air-gapped | Hosted with self-host options | You operate it |
| Honest gap | Multi-provider tracing depth | Specialist observability across vendors | You build everything |
Frequently asked questions
Is Helicone a model provider?
No. Helicone is a gateway and observability layer that proxies requests to model providers; it does not serve inference itself.
Can Helicone work with Plugsky?
Because Plugsky exposes OpenAI-compatible endpoints, a provider-agnostic gateway can proxy Plugsky traffic the same way it proxies other vendors.
Do I need a gateway at all?
If one platform covers your models and its analytics answer your questions, maybe not. If you route across vendors or need prompt-level tracing, keep one.
Is there a free plan on Plugsky?
Yes — plugsky-micro and plugsky-lite are free with no credit card, and a 14-day full-access trial is available.
How is Plugsky priced?
Flat monthly self-serve plans with unlimited fair-use usage and no per-token billing. See the live pricing page.
What about log data residency?
Treat observability logs as production data. Define retention and residency rules before routing prompts through any gateway.
Which order should they run in?
The gateway sits between your application and the model platform, so requests are instrumented before they reach Plugsky.