A general-purpose model is trained on broad data, then made available for many tasks. Fine-tuning is a second round of training that adapts a pre-trained model using examples from a narrower domain or task.
The goal is not to teach the model one fact for one request. The goal is to change how the model behaves more consistently across many future requests. In a common training path, fine-tuning comes after broad pre-training and post-training, but some teams also fine-tune a base model directly when that fits the use case.
What fine-tuning changes
During normal inference, the model’s parameters stay fixed. Prompt engineering can guide the model for one request, but it does not rewrite the model’s weights.
Fine-tuning is different. It runs additional training and updates some model parameters. After fine-tuning, the model itself has changed. It may follow a format more reliably, use domain language more naturally, or handle a repeated task with less instruction in every prompt.
Common fine-tuning goals include:
- Producing a consistent output format.
- Adapting tone or style to a specific product.
- Improving performance on a narrow classification or extraction task.
- Teaching domain-specific language patterns.
- Reducing the amount of repeated prompt instruction needed at runtime.
Fine-tuning does not guarantee factual accuracy. It changes behavior, not the need for evaluation.
A small example
Imagine a support team wants every issue summary to follow this shape:
Problem: ...
Likely cause: ...
Next step: ...
Priority: ...
Prompting can ask a general model to use that format. If the task is occasional, that may be enough. But if the product needs thousands of summaries every day and small formatting mistakes break downstream automation, fine-tuning may help.
The training set might include examples like:
Input:
Customer cannot reset password after receiving the reset email.
Output:
Problem: Password reset link does not complete successfully.
Likely cause: Expired link, blocked redirect, or account state issue.
Next step: Ask for timestamp and browser, then generate a fresh reset link.
Priority: Medium
After enough high-quality examples, the fine-tuned model can learn the desired structure and language pattern more reliably than a prompt alone. The important pattern is learned from many examples, not from one perfect demonstration.
Fine-tuning vs prompting vs retrieval
Fine-tuning is often confused with other ways to improve a model. They solve different problems.
| Approach | What changes | Best for |
|---|---|---|
| Prompting | The temporary input | Instructions, examples, formatting, task framing |
| Retrieval | The context supplied at runtime | Fresh facts, private documents, source-grounded answers |
| Fine-tuning | Model weights or added adapters | Repeated behavior, style, classification, specialized patterns |
If the model lacks current or private knowledge, fine-tuning is usually not the first answer. Retrieval and context engineering are often better because the source material can be updated without retraining the model.
If the model misunderstands the task format, start with better prompting and examples. Fine-tuning becomes more attractive when the prompt is large, brittle, expensive, or still unreliable after careful design.
Types of fine-tuning
There are several ways to adapt a model:
- Full fine-tuning updates many or all model parameters. It can be powerful, but expensive and harder to manage.
- Parameter-efficient fine-tuning (PEFT) usually trains a small set of added parameters while keeping most of the base model fixed.
- Instruction fine-tuning trains the model on input-output examples that teach it to follow task instructions.
- Preference tuning uses comparisons between outputs to encourage preferred responses. It includes RLHF and direct methods such as DPO.
Hosted model providers may hide these details behind a simpler API. Even then, it helps to understand that fine-tuning is training, not configuration.
LoRA and QLoRA adapters
LoRA, short for low-rank adaptation, is a common PEFT method. It freezes the base weights and trains small adapter matrices that are added to some layers. Instead of rewriting a huge weight matrix, LoRA learns a small update that nudges the layer in the right direction.
That adapter file can be much smaller than the base model: often megabytes instead of gigabytes. Teams can swap adapters for different tasks, keep them as separate artifacts, or merge the adapter update into the base weights for deployment. This is especially common with open-weight models where teams control the runtime.
QLoRA uses LoRA while the base model is loaded in a quantized format, often 4-bit, to save memory. The base stays mostly frozen and compressed; the small adapter remains trainable. See parameters and weights for the memory and quantization basics.
Continued pre-training, instruction tuning, and preference tuning
These names are easy to mix up:
- Continued pre-training uses more raw domain text and the same prediction-style objective. It helps the model absorb domain language patterns, not a specific response format.
- Instruction fine-tuning uses input-output demonstrations to teach a task, format, or assistant behavior.
- Preference tuning uses preferred and rejected outputs to move behavior toward what reviewers, rules, or reward signals prefer.
What makes good training data
Fine-tuning is only as good as the examples used to train it. A small, clean dataset often beats a large, noisy one.
Good fine-tuning data should be:
- Representative: examples match real production inputs.
- Consistent: similar inputs receive similar outputs.
- Specific: examples show the behavior you actually want.
- Reviewed: labels and answers are checked for errors.
- Deduplicated: repeated examples should not dominate the training signal.
- Separated from evaluation data: test examples should not be used for training.
Use clear splits: a training set for updates, a validation set for choosing checkpoints and settings, and a final test set for the last comparison. Keep a held-out evaluation set that reflects real inputs, including edge cases. Compare the tuned model against the original model plus a strong prompt baseline, not just against an old weak prompt. Avoid leakage: if test examples or near-duplicates appear in training, the measured gain may be fake.
If the examples contain contradictions, stale policy, private data that should not be learned, or inconsistent formatting, the fine-tuned model can learn those problems too.
Costs and risks
Like other post-training choices, fine-tuning adds operational work. Before choosing it, consider:
- Data collection: creating examples can take more time than writing prompts.
- Evaluation: you need tests to prove the tuned model is better.
- Versioning: each tuned model and adapter is a new artifact to track.
- Drift: the task, policy, or base model may change later.
- Overfitting: the model may memorize narrow examples and perform worse on new cases.
- Catastrophic forgetting: the model may lose general abilities while specializing. Mitigate with mixed general data, lower learning rates, fewer epochs, or adapters.
- Data privacy: examples can contain sensitive text, so collect and retain them deliberately.
- Deployment constraints: fine-tuned models may have different cost, latency, or provider support.
Operationally, tune the learning rate, number of epochs, and checkpoint choice on validation results. Keep versioned datasets, prompts, adapters, base model IDs, and evaluation reports so you can roll back when a new tuned model regresses.
Fine-tuning should improve a measured outcome. If you cannot define what better means, it is hard to know whether the tuning helped.
When fine-tuning is worth it
Fine-tuning is most useful when:
- The task repeats often enough to justify the setup cost.
- You have high-quality examples of the desired behavior.
- Prompting and retrieval have been tried and still fall short.
- The desired behavior is stable, not changing every week.
- You can evaluate the model before and after tuning.
It is less useful when the problem is missing knowledge, weak product requirements, or poor retrieval. In those cases, changing the model may hide the real issue.
The key idea
Fine-tuning adapts a pre-trained model by continuing training on task-specific examples. It can make repeated behavior more reliable, but it is not a shortcut around good data, prompting, retrieval, or evaluation. Use it when you need a model to consistently behave differently, not just when you need to give it more information.