RAG vs fine-tuning: which one your use case actually needs
Both make a language model useful on material it was never trained on, and they are routinely presented as rivals. They are not. Retrieval looks your facts up at the moment of the question; fine-tuning bakes behaviour into the model in advance. Choosing between them is mostly one question: how often does the thing you want it to know change?
What each one actually does
Retrieval-augmented generation keeps your material outside the model. At question time it finds the passages that look relevant and hands them to the model along with the question, so the answer is assembled from text you can point at. Fine-tuning changes the model itself: you show it several hundred to several thousand example exchanges and it learns the shape of them - the tone, the format, the house style of an answer.
The deciding question is how often your facts change
Prices, opening hours, stock, policies and staff change weekly. Retrieval handles that by re-reading the source; fine-tuning would need retraining every time, which nobody does, so the model confidently states last quarter's price. If what you want to change is not the facts but the manner - always answer in two sentences, always open with the booking link, never use the word "unfortunately" - that is behaviour, and that is what fine-tuning is for.
Cost, time and what goes wrong
Retrieval is set up in minutes and costs a little on every question, because the retrieved passages are part of what the model reads. Its failure mode is visible and fixable: the wrong passage is retrieved, the answer is wrong, you look at the log and fix the source page. Fine-tuning costs a training run and the work of assembling the examples, and its failure mode is invisible: a model that has learned to sound right about something it has no current knowledge of. There is no citation to check, because there is no source.
When you want both, and when you want neither
A support assistant for a business is a retrieval problem with a thin behavioural layer - and in practice the behavioural part is usually solved with instructions rather than training. Fine-tuning earns its keep when the output format is strict and repetitive, or the domain language is unusual enough that general models mangle it. If your content fits in a single prompt and rarely changes, you need neither: paste it in.
Frequently asked questions
What is the difference between RAG and fine-tuning?
RAG keeps your knowledge outside the model and looks it up when a question arrives, so answers cite a source and update the moment the source does. Fine-tuning changes the model's weights using example exchanges, which teaches it style and format rather than facts. RAG is for what the model should know; fine-tuning is for how it should behave.
Is RAG better than fine-tuning?
Neither is better in general. For knowledge that changes - prices, availability, policies - retrieval is the right answer and fine-tuning is actively risky, because a retrained model states outdated facts with full confidence and no citation. For a strict output format or an unusual domain vocabulary, fine-tuning does something retrieval cannot.
Can you use RAG and fine-tuning together?
Yes, and for demanding cases that is the usual arrangement: fine-tune for the shape of the answer, retrieve for its content. Start with retrieval, because it is cheaper to build and its mistakes are visible, and add fine-tuning only when you can name the behaviour that instructions failed to produce.
Which is cheaper, RAG or fine-tuning?
Retrieval has almost no setup cost and a small recurring cost per question, because retrieved passages are part of what the model reads. Fine-tuning has a real up-front cost - assembling and cleaning the examples is most of it - and a lower per-question cost afterwards. The up-front cost repeats every time your material changes enough to need retraining.
Does a website chatbot need fine-tuning?
Usually not. A website assistant has to be current about prices, hours and services, which is exactly where fine-tuning is weakest, and it has to be able to say where an answer came from. SLAtech AI answers from your own pages by retrieval and cites them.