LLM (Large Language Model)

A large language model, or LLM, is a program trained on an enormous amount of text to do one thing: predict what text comes next. Ask it a question and it doesn't look anything up. It works out, one small piece at a time, what text would most plausibly follow your question. That turns out to be enough to answer questions, write code, summarize documents, and translate — which is why nearly every AI product you've used has one underneath it.

Almost everything else on this site sits on top of that single mechanism. It's the reason a model can sound completely confident and still be wrong, the reason it doesn't remember your conversation from last week, and the reason so much engineering effort goes into deciding exactly what text to put in front of it.

What the "large" refers to

Two things: how much text it was trained on, and how many parameters it has. Parameters are the internal numbers the model adjusts during training to encode patterns it finds in that text — current models have billions of them. Neither number is a quality score on its own. A model with more parameters isn't automatically better at your task, and providers regularly ship smaller, cheaper models that outperform older larger ones.

How it actually produces an answer

The model doesn't compose a sentence all at once. It works in tokens — small chunks of text, often a whole word but frequently just part of one. Given everything it has seen so far, it calculates a probability for every token that could come next, one gets picked, that token is added to the text, and the loop runs again with the slightly longer text as its new input. A paragraph is that loop repeated a few hundred times.

How the pick gets made is adjustable. Temperature controls whether the model reliably takes the most likely token or sometimes takes a less likely one, which is what makes the same prompt produce different answers on different runs. Nearly every model in production use today generates text this way, one token at a time; other approaches exist and are an active research area, but they aren't what you'll be building on right now.

Training happens once — using the model doesn't change it

Training is a separate, finished event that happened before you ever sent a request. Once it ends, the model's parameters are frozen. Using the model doesn't teach it anything: your conversation doesn't update it, and nothing you tell it today is available to anyone else's request tomorrow.

Two consequences follow. First, the model's knowledge stops at whenever its training data ended — its knowledge cutoff — so it has no idea what happened afterward, and no idea what's in your company's documents at all. Second, when a model appears to remember something from earlier in a chat, that's not memory. The earlier messages are simply being sent again along with each new one, which is why the context window — the amount of text it can take in at once — is a hard constraint rather than a detail.

When you need a model to work with information it wasn't trained on, you supply that information at the moment you ask, which is what RAG does. When you need to change how it behaves rather than what it knows, that's fine-tuning. Those are different problems with different fixes, and conflating them is one of the most common early mistakes.

Why it makes things up

Producing a wrong answer and producing a right one are the same operation. The model generates whatever text is most plausible given what came before — and a fluent, confident, entirely invented citation is extremely plausible-looking text. Nothing in the mechanism checks a claim against a source, because there is no source; there's only a probability distribution over what usually follows. That's hallucination, and it's a property of how the thing works rather than a bug that a better model will eventually eliminate.

Where an LLM is and isn't the right tool

It's a strong fit for work involving language where some variability is acceptable: drafting, summarizing, rewriting, classifying text, answering questions from material you provide, and writing code. It's a poor fit when you need a guaranteed-correct answer that a database query or a few lines of ordinary code would produce exactly — arithmetic, looking up a record, applying a fixed rule. Reaching for a model there adds cost, latency — the delay before an answer comes back — and a failure mode you didn't previously have.

In this guide
  1. What the "large" refers to
  2. How it actually produces an answer
  3. Training happens once — using the model doesn't change it
  4. Why it makes things up
  5. Where an LLM is and isn't the right tool
  6. FAQ

FAQ

Is an LLM the same thing as AI?

No. AI is the broad field; an LLM is one specific kind of system within it. Plenty of AI has nothing to do with language models — image recognition, recommendation systems, the routing in a maps app. LLMs get most of the current attention, but they aren't the whole category.

Is ChatGPT an LLM?

ChatGPT is a product built around one. The distinction matters more than it sounds: the model does the text prediction, while the product wraps it in a chat interface, conversation history, safety filtering, and often tools like web search. When a chat assistant does something a raw model can't — cite a live web page, run code — that capability usually comes from the product, not the model.

Do LLMs learn from my conversations?

The model answering you does not change as you talk to it. Whether the provider separately stores your conversations and uses them to train some future model is a policy question that varies by provider and plan, and it's worth checking if you're sending anything sensitive — but that's a different thing from the model learning mid-conversation, which doesn't happen.

Does an LLM understand what it's saying?

It depends entirely on what you mean by "understand," which is why the question generates more argument than insight. What's observable and useful: it reliably produces text that behaves as if it grasped the meaning, and it also fails in ways that suggest it didn't. For building things, the practical stance is to judge it on output you can verify rather than on assumptions about what's happening internally.

Practice interview questions on LLMs →