RAG

Should you use RAG or fine-tuning?

Use RAG for knowledge that changes, must be cited, or cannot be baked into weights; use fine-tuning to change behaviour, tone or output format. RAG updates by re-indexing documents, while fine-tuning requires training data and a new model version. Many teams combine both, but knowledge grounding usually belongs in retrieval.

Key facts

RAG statusCollections, queries and citations are live on the OpenAI-compatible API
Knowledge freshnessRe-ingest a document and the next query sees it; no retraining
CitationsAnswers return ranked chunks with source attribution
Models30+ models available for generation without custom training
Fine-tuningFine-tuning endpoints are a coming-soon capability on Plugsky
DeploymentManaged, VPC, on-prem and air-gapped options
Free tierplugsky-micro and plugsky-lite, no card required
Product statusLive for RAG

TL;DR

  • RAG is for facts; fine-tuning is for form.
  • RAG updates instantly by re-indexing; fine-tuning needs a training cycle.
  • RAG can cite sources; a fine-tuned model cannot point to evidence.
  • Fine-tuning suits stable tasks, style and structured output habits.
  • Combine them when behaviour is stable but knowledge keeps changing.

How it works, step by step

  1. Write down the actual problem: missing knowledge, wrong format or wrong tone.
  2. If knowledge is missing or changes, start with retrieval over the documents.
  3. If output shape is wrong but knowledge is present, consider prompt and schema first.
  4. Only pursue fine-tuning when prompting and retrieval cannot change the behaviour.
  5. Assemble high-quality training examples if fine-tuning is genuinely needed.
  6. Evaluate retrieval and fine-tuned variants on the same task set.
  7. Keep knowledge in retrieval so updates do not require retraining.
1Write down theactual problem:missing knowledge,2If knowledge ismissing or changes,start with3If output shape iswrong but knowledgeis present,4Only pursuefine-tuning whenprompting and5Assemblehigh-qualitytraining examples6Evaluate retrievaland fine-tunedvariants on the

Try it yourself

Open the best model for RAG selector →

What each approach changes

RAG changes what the model sees at inference time. Documents are chunked, embedded and retrieved so the answer is grounded in a corpus you control, and every answer can cite the chunks it used. Updating knowledge means re-ingesting documents, which is cheap and immediate.

Fine-tuning changes the model itself. It adjusts behaviour: response style, adherence to a format, domain vocabulary or a consistent classification pattern. It is poor at storing changing facts, because each update requires a new training run and the resulting model still cannot cite its sources.

The practical decision rule

Ask what is failing. If the model does not know something, that is a knowledge problem and retrieval is the fix. If the model knows the answer but writes it in the wrong shape, a better prompt or a JSON schema may solve it before training is considered. If the model consistently produces the wrong tone or structure across many examples despite good prompts, fine-tuning becomes a reasonable option.

Traceability usually settles the argument in regulated settings: a retrieval answer can show its evidence, while a fine-tuned model answers from weights with no citation to verify.

Cost, freshness and maintenance

RAG has ongoing costs for storage and retrieval but no training cycle, and it updates the same day a document changes. Fine-tuning has a one-time training cost plus the cost of maintaining datasets, retraining when data drifts, and versioning models in production. Both add evaluation work, but only RAG keeps knowledge and behaviour separable.

A common mistake is fine-tuning to inject facts. The model may memorise some of them, but it will drift, cannot cite, and becomes stale the moment the facts change.

How teams combine them

A durable pattern is fine-tune for form, retrieve for facts: a tuned or carefully prompted model that reliably follows your output contract, with retrieval supplying current, cited content. On Plugsky, retrieval is live today through collections and queries with citations, and 30+ models are available for generation, so behaviour can be tuned through prompt design while knowledge stays in the corpus.

Fine-tuning endpoints are listed as coming soon in the API docs, so check the capability matrix before planning a training workflow. Start retrieval now with the best model for RAG selector, then evaluate on the free plan with plugsky-micro and plugsky-lite or the 14-day full-access trial. Current plans are on the live pricing page.

Honest comparison

DimensionRAGFine-tuningPrompting alone
Best forChanging, citable knowledgeStable behaviour and formatSimple format and tone fixes
FreshnessUpdate on re-ingestionRequires retrainingImmediate but limited
CitationsChunk-level source referencesNoneNone
Cost shapeStorage and retrievalTraining plus maintenancePrompt engineering only
Data neededDocumentsCurated input-output examplesA clear prompt
RiskRetrieval quality gapsDrift, overfitting, stalenessInconsistent compliance

Frequently asked questions

Can fine-tuning replace RAG?

For changing knowledge, no. Fine-tuning is better at behaviour and format, while retrieval handles facts that change and need citations.

Which is cheaper?

Depends on scale and task. Retrieval avoids training cycles but has ongoing storage and query work; fine-tuning has training and dataset maintenance costs.

Does Plugsky support fine-tuning?

Fine-tuning endpoints are marked as a coming-soon capability in the API docs. RAG, embeddings, chat, streaming, JSON mode and function calling are live today.

Can I do both?

Yes, and many teams do: fine-tune or carefully prompt for output behaviour, and use retrieval for current, citable content.

Is there a free plan?

Yes. The free plan includes plugsky-micro and plugsky-lite with 2 API keys and no credit card, and a 14-day full-access trial is available.

How is pricing structured?

Self-serve plans are flat monthly with unlimited fair-use usage and no per-token charges or overage fees. See the live pricing page for current plans.