Hallucination
A hallucination is when a model states something false as though it were fact — an invented citation, a function that doesn't exist, a confidently wrong date. Nothing goes wrong inside the model when this happens. It's doing exactly what it always does. There's no separate mode for "looking something up" versus "making something up," so a hallucination comes out of the same machinery as every correct answer you've ever had from it.
A fabrication up close
Ask for code using a library the model only half-knows and you might get client.batch_update(records) — right naming style for that library, plausible arguments, sitting in exactly the place such a method would sit. And no such method exists. The model doesn't hedge, because nothing in its process marks this call as different from the ones that are real.
The same shape turns up as academic citations with real-sounding authors and journals, court cases that were never filed, and product features that were never shipped. In every case the output is well-formed, which is the hard part: these models are very good at fluent writing, and fluent writing is what makes a wrong answer difficult to catch.
Why it happens
A language model generates text one token at a time — a token being a chunk of text, usually a word or part of one. At each step it works out a probability for every possible next token and picks from those. No database sits behind it, no fact-checking step runs, and nothing in the text it produces marks which statements it saw during training and which it assembled on the spot.
When the training data supported an answer strongly, the likely next tokens happen to be true ones. When it didn't, the model still produces the most plausible-sounding continuation, because producing plausible continuations is the only thing it does. Plausibility is what it optimizes for. Truth is not a separate target it can fall back on.
There's a second cause, and it's about how models are graded rather than how they run. Researchers at OpenAI and Georgia Tech argue that standard training and evaluation reward guessing over admitting uncertainty: a test that awards a point for a correct answer and nothing for a blank makes guessing the better strategy. In OpenAI's own summary of the work, a model asked for someone's birthday has a one-in-365 shot from a guess and a guaranteed zero from "I don't know." Over thousands of questions the guesser scores higher, and a model shaped by that scoring learns to produce an answer rather than admit it hasn't got one.
When you're most at risk
Hallucination isn't evenly distributed, and two things predict it well enough to plan around.
The first is how obscure the fact is. Anything documented thousands of times over — what a for loop does, the capital of France, how a popular framework's main API works — is well supported by training data and rarely fabricated. Specific numbers, dates, names, citations, pricing, the methods of a niche library, anything about a private or internal system: these are long-tail, and long-tail is where invention lives. A usable rule: if you wouldn't expect the answer to appear many times on the public internet, treat it as unverified.
The second is whether you supplied the source. A model answering from memory is guessing at recall. A model answering from a document you pasted in is reading. The risk drops sharply in the second case, which is what every real mitigation below is built on.
Why the obvious fixes don't work
Telling the model not to make things up does very little by itself. It can't compare its output against a source it was never given, so a general instruction to be truthful has nothing to act on. Specific instructions that name an acceptable alternative are a different matter, and they're covered below.
Lowering the temperature doesn't help either. Temperature controls how strongly a model favors its top candidate, so if the top candidate is the fabricated one, a lower setting just returns that same wrong answer more reliably. The temperature page covers why this is the wrong lever to reach for.
Reaching for a newer model seems the most obviously right move and is the least dependable. OpenAI's own system card for o3 and o4-mini measured both hallucinating more than the older o1 they followed — 33% and 48% against o1's 16% on a people-focused benchmark — and states that more research is needed to understand why. Newer models are often better here. They are not reliably better, and when a strong model is wrong, the fluency that makes it a better writer makes the error harder to notice.
What actually reduces it
One idea sits behind all of these: stop asking the model to remember, and give it something to read.
Ground the answer in retrieved text. Put the real source material into the request itself and the job changes from recall to reading comprehension, which models are far better at. That's retrieval-augmented generation, or RAG. It's the biggest single improvement available to most applications, and it doesn't eliminate the problem: a model can still answer from memory while ignoring the material you handed it.
Make it quote and cite. Requiring each claim to point at the passage it came from does two jobs at once. Every statement becomes checkable, and the model is pushed toward material that's actually in front of it.
Give it permission not to know. This is the practical counterweight to the scoring problem above. "If the documents don't answer the question, say so" works where "be accurate" doesn't, because it names an acceptable alternative output. Without one, producing something is the model's only option.
Shrink the space of possible answers. A classification task — sorting an input into one of a fixed set of categories — leaves far less room to invent than free text does. And if your provider supports constrained or schema-enforced output, where the response is forced to match a declared shape, a sixth label becomes impossible rather than merely unlikely. A prompt that merely lists five allowed labels does not give you that guarantee.
Check the result separately. A second pass — another model call, or plain code — that tests claims against the source will catch what the first pass invented.
Not every task needs all of this. Where a mistake is instantly obvious, the feedback loop is already doing your verification: code that won't compile and a query that errors announce their own failure, and building a checker around them is wasted effort. Where output is low-stakes and easily thrown away — brainstorming, naming, first drafts — the cost of verifying exceeds the cost of being wrong. And if the source material is small enough to paste straight into the request, do that instead of building a retrieval pipeline. Save the heavy machinery for answers someone will act on without checking.
Hallucination, or just out of date?
Not every wrong answer is a hallucination, and the difference decides which fix applies. A model that tells you the current version of a library is 3.2 when 4.0 shipped last month isn't fabricating anything — it's reporting what was true when its training data was collected. That's a knowledge cutoff problem, and the answer is to give it current information or a way to look it up.
A library method that never existed in any version is a hallucination. Identical symptom from the user's side, different cause, different fix.
In this guide
FAQ
Is "hallucination" even the right word?
Plenty of researchers think not. Hallucinating means perceiving something that isn't there, while the model is generating plausible filler — so "confabulation" gets proposed instead, and critics point out that both words describe a statistical process as though it had a mind. The term stuck anyway, largely because it's what everyone already searches for and writes.