Key facts
| RAG status | Collections, queries and citations are live on the OpenAI-compatible API |
| Knowledge freshness | Re-ingest a document and the next query sees it; no retraining |
| Citations | Answers return ranked chunks with source attribution |
| Models | 30+ models available for generation without custom training |
| Fine-tuning | Fine-tuning endpoints are a coming-soon capability on Plugsky |
| Deployment | Managed, VPC, on-prem and air-gapped options |
| Free tier | plugsky-micro and plugsky-lite, no card required |
| Product status | Live for RAG |
TL;DR
- RAG is for facts; fine-tuning is for form.
- RAG updates instantly by re-indexing; fine-tuning needs a training cycle.
- RAG can cite sources; a fine-tuned model cannot point to evidence.
- Fine-tuning suits stable tasks, style and structured output habits.
- Combine them when behaviour is stable but knowledge keeps changing.
How it works, step by step
- Write down the actual problem: missing knowledge, wrong format or wrong tone.
- If knowledge is missing or changes, start with retrieval over the documents.
- If output shape is wrong but knowledge is present, consider prompt and schema first.
- Only pursue fine-tuning when prompting and retrieval cannot change the behaviour.
- Assemble high-quality training examples if fine-tuning is genuinely needed.
- Evaluate retrieval and fine-tuned variants on the same task set.
- Keep knowledge in retrieval so updates do not require retraining.
Try it yourself
Open the best model for RAG selector →
What each approach changes
RAG changes what the model sees at inference time. Documents are chunked, embedded and retrieved so the answer is grounded in a corpus you control, and every answer can cite the chunks it used. Updating knowledge means re-ingesting documents, which is cheap and immediate.
Fine-tuning changes the model itself. It adjusts behaviour: response style, adherence to a format, domain vocabulary or a consistent classification pattern. It is poor at storing changing facts, because each update requires a new training run and the resulting model still cannot cite its sources.
The practical decision rule
Ask what is failing. If the model does not know something, that is a knowledge problem and retrieval is the fix. If the model knows the answer but writes it in the wrong shape, a better prompt or a JSON schema may solve it before training is considered. If the model consistently produces the wrong tone or structure across many examples despite good prompts, fine-tuning becomes a reasonable option.
Traceability usually settles the argument in regulated settings: a retrieval answer can show its evidence, while a fine-tuned model answers from weights with no citation to verify.
Cost, freshness and maintenance
RAG has ongoing costs for storage and retrieval but no training cycle, and it updates the same day a document changes. Fine-tuning has a one-time training cost plus the cost of maintaining datasets, retraining when data drifts, and versioning models in production. Both add evaluation work, but only RAG keeps knowledge and behaviour separable.
A common mistake is fine-tuning to inject facts. The model may memorise some of them, but it will drift, cannot cite, and becomes stale the moment the facts change.
How teams combine them
A durable pattern is fine-tune for form, retrieve for facts: a tuned or carefully prompted model that reliably follows your output contract, with retrieval supplying current, cited content. On Plugsky, retrieval is live today through collections and queries with citations, and 30+ models are available for generation, so behaviour can be tuned through prompt design while knowledge stays in the corpus.
Fine-tuning endpoints are listed as coming soon in the API docs, so check the capability matrix before planning a training workflow. Start retrieval now with the best model for RAG selector, then evaluate on the free plan with plugsky-micro and plugsky-lite or the 14-day full-access trial. Current plans are on the live pricing page.
Honest comparison
| Dimension | RAG | Fine-tuning | Prompting alone |
|---|---|---|---|
| Best for | Changing, citable knowledge | Stable behaviour and format | Simple format and tone fixes |
| Freshness | Update on re-ingestion | Requires retraining | Immediate but limited |
| Citations | Chunk-level source references | None | None |
| Cost shape | Storage and retrieval | Training plus maintenance | Prompt engineering only |
| Data needed | Documents | Curated input-output examples | A clear prompt |
| Risk | Retrieval quality gaps | Drift, overfitting, staleness | Inconsistent compliance |
Frequently asked questions
Can fine-tuning replace RAG?
For changing knowledge, no. Fine-tuning is better at behaviour and format, while retrieval handles facts that change and need citations.
Which is cheaper?
Depends on scale and task. Retrieval avoids training cycles but has ongoing storage and query work; fine-tuning has training and dataset maintenance costs.
Does Plugsky support fine-tuning?
Fine-tuning endpoints are marked as a coming-soon capability in the API docs. RAG, embeddings, chat, streaming, JSON mode and function calling are live today.
Can I do both?
Yes, and many teams do: fine-tune or carefully prompt for output behaviour, and use retrieval for current, citable content.
Is there a free plan?
Yes. The free plan includes plugsky-micro and plugsky-lite with 2 API keys and no credit card, and a 14-day full-access trial is available.
How is pricing structured?
Self-serve plans are flat monthly with unlimited fair-use usage and no per-token charges or overage fees. See the live pricing page for current plans.