Key facts
| Router model | plugsky-fusion escalates per request across tiers (live) |
| Strategies | cost_saver for drafts and variants, max_quality for final copy |
| Models | 30+ models; cheap tiers handle outlines, variants and structured copy |
| JSON mode | Live for structured content plans and metadata |
| Pricing | Flat monthly self-serve plans with no per-token charges on self-serve |
| Free tier | plugsky-micro and plugsky-lite on the free plan, no card required |
| Usage controls | Scoped keys and usage analytics per content pipeline or channel |
| Roadmap | The batch endpoint is coming soon for bulk generation runs |
TL;DR
- Draft cheap, polish selectively — output volume is the bill.
- Separate generation from review so only final copy uses strong models.
- Cap length per format; unbounded outputs are unbounded cost.
- Reuse system prompts and style guides instead of rewriting them.
- Measure cost per published piece, not cost per generation call.
How it works, step by step
- Map the content pipeline: brief, outline, draft, variants, edit, final.
- Assign cheap tiers to briefs, outlines, variants and structured metadata.
- Reserve stronger models for final brand-sensitive copy and high-stakes channels.
- Cap output length per format and request outlines or bullets before prose.
- Store system prompts and style guides centrally and version them.
- Use JSON mode for plans, headlines and metadata so parsing is deterministic.
- Track cost per published piece and re-tune tiers as quality and volume change.
Try it yourself
Output tokens are the content bill
Content generation writes a lot: articles, product descriptions, ad variants, emails. Output tokens dominate the request, so the strongest lever is not which model writes but how much it writes and how often you regenerate. A plan-first workflow — outline, approve, then expand — prevents paying for prose that gets discarded.
Cheap tiers are genuinely good at structured work: outlines, headline sets, tone variants, metadata. Save stronger models for the final pass where voice, nuance and brand risk live. That division keeps quality where readers notice it and removes cost where they do not.
Routing a content pipeline
Treat the pipeline as stages with different quality bars:
- Brief and outline: cheap tier, JSON mode, short outputs.
- Variants: cheap tier, capped length, generated in groups.
- Draft: mid tier for long-form, cheap tier for short-form.
- Final polish: strong tier, one pass, human approval.
- Metadata and SEO: cheap tier, strict schemas.
Set strategies per channel key so a paid social pipeline and a landing-page pipeline do not share one policy. Pin regulated or brand-critical channels with custom rules.
Prompts, reuse and measurement
Most content prompts repeat the same instructions: voice, audience, forbidden claims, format. Keep these in a versioned system prompt and send them once, then let each request carry only the task specifics. Reusing structure reduces input size and makes style consistent across a team.
Measure cost per published piece alongside edit rate, time-to-publish and approval cycles. A cheap draft that a human rewrites entirely costs more than a stronger first draft. Run a small evaluation set per content type after any prompt or tier change. Flat self-serve plans keep the economics stable — see the live pricing page — and the free plan with plugsky-micro and plugsky-lite is enough to build the pipeline before scaling.
Honest comparison
| Content stage | Routed pipeline on Plugsky | Strongest model for all | Single cheap model |
|---|---|---|---|
| Briefs and outlines | Cheap tier with JSON mode | Strong tier for planning | Usually adequate |
| Variants | Cheap tier, capped length | Expensive experiments | Sometimes flat |
| Long-form draft | Mid or cheap tier | Strong tier, high cost | Weak structure |
| Final polish | Strong tier, one pass | Native strength | Brand risk |
| Cost visibility | Per-channel usage analytics | Blended, unclear | Blended, unclear |
Frequently asked questions
Why does content generation cost so much?
Output volume dominates. Long articles, variants and rewrites all consume output tokens, so length control and reducing discarded drafts matter more than input optimisation.
Which models should write first drafts?
Cheap tiers handle outlines, variants and structured copy well. Reserve stronger models for final, brand-sensitive or regulated content where nuance changes the outcome.
How do I avoid paying for discarded drafts?
Separate planning from writing: generate and approve outlines first, then expand. You only pay for prose that has a reason to exist.
Should metadata use the same model as articles?
No. Headlines, metadata and SEO fields are short structured outputs that cheap tiers handle well with JSON mode and strict schemas.
How do I keep brand voice consistent?
Version a shared system prompt with voice, audience and forbidden claims, and send only task specifics per request. Consistency is a prompt-management problem, not a model-size one.
Is batch generation available?
Not yet — the batch endpoint is coming soon. Run bulk generation with bounded-concurrency workers and a queue, and keep per-item status for retries.
How does flat pricing help content teams?
Self-serve plans are flat monthly with no per-token charges, so experimenting with variants and prompts does not create a variable bill. See the live pricing page for plan details.
Can I try the pipeline for free?
Yes. plugsky-micro and plugsky-lite are on the free plan with no card, and the 14-day full-access trial covers stronger models for final-copy evaluation.