Fine-Tuning

Fine-tuning takes a model that's already trained and trains it further on examples you supply. You give it pairs — an input, and the output you wanted for it — and training adjusts the model's weights, the numbers inside it that determine how it responds, so that kind of output becomes more likely. What comes out the other end is a new version of the model that behaves differently from the one you started with.

The mistake almost everyone makes

Fine-tuning a model on your documents does not reliably teach it those documents. Training on text teaches patterns — phrasing, structure, vocabulary, the shape of a good answer — it does not install a lookup table. A model fine-tuned on a company handbook starts sounding like that handbook while still getting details wrong, confidently, with no way to check where any answer actually came from. If the goal is "it should know our facts," that's a retrieval problem, not a training one — see RAG, and RAG vs Fine-Tuning for which one a given problem actually calls for.

Parameter-efficient training

Most fine-tuning today is parameter-efficient, meaning training adds a small set of new weights and leaves the original model frozen instead of adjusting everything inside it. LoRA — low-rank adaptation — is the common method. This matters for one practical reason: it's what makes fine-tuning affordable, since training a small add-on is far cheaper than retraining a model from scratch.

Real constraints

It needs a genuine dataset — typically hundreds to thousands of good example pairs — and assembling those well-labeled examples is usually more expensive than the training run itself. It can also make a model worse at anything outside the task it was tuned for, so a model tuned hard for one job is not a better general model afterward. And the model you want isn't always the one you're allowed to tune: providers typically offer fine-tuning on older or smaller models rather than their newest ones. As of 2026, Anthropic's own Claude API doesn't offer fine-tuning at all, and no current-generation Claude model supports it on Amazon Bedrock either — fine-tuning there covers Claude 3 Haiku, an older model, in one AWS region. Choosing to fine-tune can mean giving up the strongest model actually available to you.

What it's genuinely good for

Reshaping how a model responds when it already has the knowledge but delivers it wrong: always producing one structured format, matching a house writing style, using domain-specific vocabulary consistently, sorting things into fixed categories reliably. It's also a real lever for cost at scale — if a long set of instructions and examples would otherwise repeat on every one of a million calls, training that behavior into a smaller model can end up both cheaper and faster than repeating the same lengthy prompt every time.

When it's not the right tool

Before reaching for it, try prompting — a clear set of instructions with a few examples solves more problems than people expect, costs nothing to change, and takes an afternoon rather than a training run. Skip fine-tuning specifically when the actual goal is teaching the model new facts (that's retrieval), when different users need to see different information (a fine-tuned model can't filter per user — whatever it was trained on is inside every answer, for everyone), or when the task needs your provider's newest model, which fine-tuning access often doesn't cover.

In this guide
  1. The mistake almost everyone makes
  2. Parameter-efficient training
  3. Real constraints
  4. What it's genuinely good for
  5. When it's not the right tool
  6. FAQ

FAQ

Is fine-tuning the same as training a model from scratch?

No, and the gap is enormous. Training from scratch starts from random numbers and needs a vast dataset and a budget most companies don't have. Fine-tuning needs a comparatively tiny set of examples, because the hard and expensive part was already done by whoever trained the base model.

If fine-tuning doesn't reliably teach a model new facts, why does it seem to "learn" the content it was trained on?

It learns the style and structure of that content, which can look like knowledge on the surface — confident, correctly formatted answers that sound like they came from the source material. What it hasn't learned is a way to check any specific answer against that material, which is exactly why it can sound right while being factually wrong.

Practice interview questions on prompting vs fine-tuning →