The three compared
| RAG | Fine-tuning | Private LLM | |
|---|---|---|---|
| The problem it solves | Answers from your own documents | Consistent style, format or classification | Data that must stay inside your control |
| What changes | What the model can look up | How the model behaves | Where the model runs |
| Stays current as documents change? | Yes, re-index | No, retrain | Not by itself |
| Shows its sources? | Yes | No | Only if combined with RAG |
| Typical first step? | Yes | Rarely | When privacy or regulation requires it |
RAG: when the answer is in your documents
If staff or customers ask questions that your SOPs, contracts, quotes or support history can answer, RAG is the fit. It keeps working as documents change and it shows its sources, which is what makes the answers checkable.
Fine-tuning: when the model's behaviour needs to change
Fine-tuning trains a model on your examples so it writes in your format, follows your classification scheme or handles a narrow task more consistently. It does not teach the model facts reliably and it does not update when your documents do, so it is the wrong tool for 'answer from our policies'.
Private LLM: when the data cannot leave
A private LLM is an open-weight model that runs on hardware you own or a private cloud you control. Choose it when client files, patient records, deal data or trade secrets cannot go to a third-party AI provider. It answers a where-it-runs question, so it is usually combined with RAG.
How to choose
- Write down the task in one sentence. If it contains 'from our documents', start with RAG.
- If the task is about tone, format or sorting, test whether better prompts get you there before considering fine-tuning.
- Ask what data the AI will see. If any of it cannot be shared with a vendor, plan for a private deployment.
- Test on real questions before committing. A small test set decides more than any comparison table.