In this guide
LLM Interview Questions and Answers
Forty questions on how large language models actually work, how they behave, and how they get used to build real applications — moving from fundamentals to practical engineering judgment. Dedicated depth on RAG and agents lives on their own pages; this one stays focused on the LLM itself.
LLM Fundamentals
1. What is a Large Language Model (LLM)?
A large language model is a neural network trained on enormous amounts of text to predict what word — or part of a word — comes next in a sequence. That one skill, applied at scale, turns out to be enough to answer questions, write code, summarize documents, and hold a conversation.
2. How does an LLM generate text?
An LLM generates text by repeatedly predicting what token — a small chunk of text, often a whole word or part of one — should come next, based on the tokens and context it has already received. Given "The capital of France is ___," it assigns a very high probability to a token representing "Paris," picks it, and repeats the process for the next token, and the next, until the response is complete. This is a useful mental model, not a complete description of everything happening inside the model.
3. What is a token?
An LLM doesn't usually read text as whole human words. Text gets broken into smaller units called tokens — a token might be a whole word, part of a word, punctuation, or another small piece of text. A long or unusual word can be represented by several tokens. Tokens matter because context limits, and often what you're charged, are measured in them. Learn more: Token.
4. What is tokenization?
The step that breaks raw text into tokens before a model can process it, and reassembles tokens back into readable text on the way out. It happens automatically, but the specific tokens a piece of text breaks into affects how much of the context window it uses.
5. What is a context window?
The maximum amount of text a model can read and act on for a single request — everything sent to it, and everything it writes back, shares that one fixed budget. Learn more: Context Window.
6. What is the difference between an LLM and a traditional language model?
Older language models were typically small, trained for one narrow task, and built on statistical or rule-based methods with a limited sense of longer context. An LLM is trained on vastly more data with far more parameters, which is what gives it broad, flexible language ability across many tasks instead of one narrow one.
7. What is a foundation model?
A large model trained on a broad range of data, meant to be adapted or built on for many different tasks rather than one narrow purpose. Most LLMs people actually use are foundation models, or built on top of one.
8. What is a reasoning model?
A model trained to work through a problem in explicit intermediate steps before giving a final answer, rather than producing an answer directly — closer to a person using scratch paper before writing down a result. Learn more: Reasoning Models.
9. What is a multimodal model?
A model that can take in, or produce, more than one kind of content — text and images, or text and audio, rather than text alone.
10. What is the difference between training and inference?
Training is when a model learns from data, adjusting its internal numbers over many passes. Inference is using that already-trained model to answer a real request, with no further learning happening. Training happens once, or occasionally; inference happens every single time someone uses the model.
Understanding LLM Behaviour
11. What is temperature in an LLM?
Temperature controls how much randomness goes into picking the next token. Low temperature makes the model stick to its most likely choice every time; high temperature lets it pick less-likely tokens more often. Learn more: Temperature.
12. Why can an LLM give different answers to the same prompt?
Because generating text means picking probabilistically among likely next tokens, not always taking the single most likely path, especially above temperature zero. Even at the lowest setting, small variation can still occur, since the underlying computation isn't always perfectly identical across runs on real hardware.
13. What is an LLM hallucination?
A confident, fluent answer that's factually wrong. The model isn't lying — it's producing plausible-sounding text, just without anything grounding this particular answer in fact. Learn more: Hallucination.
14. Why do LLMs hallucinate?
An LLM is always generating a plausible response based on learned patterns and whatever context it was given — generation doesn't inherently guarantee factual correctness. When there's no real grounding for an answer, it still produces something that sounds right, because sounding right is what it was trained to do.
15. Does an LLM understand information in the same way a human does?
Not in the way most people mean by "understand." It has no beliefs, awareness, or verified facts sitting behind an answer — it's producing the statistically plausible continuation of the text it was given, based on patterns learned from training. That process can look a great deal like understanding from the outside, especially when it's right, which is exactly why a wrong answer can be just as confident-sounding as a right one.
16. Why can an LLM confidently produce an incorrect answer?
Because fluency and correctness come from different places. The model's confidence in how it writes an answer reflects how plausible that text is as a continuation, not how factually accurate it is — nothing in the generation process checks the answer against reality before producing it.
17. What happens when the context window becomes full?
One of two things: the request fails outright with an error, or the application silently drops the oldest or least-recent part of the input to make room. The second case is the dangerous one — it keeps answering, just without whatever got cut, which can be an early instruction, a fact from a document, or a tool result the answer depended on. Learn more: Context Window.
18. What is context engineering?
Deciding what information actually goes into a model's context for a given task — which instructions, which retrieved documents, how much history — rather than treating everything as if it should just be included. Learn more: Context Engineering.
19. Why can adding more context sometimes make an LLM response worse?
Because more context is not automatically better context. Irrelevant, duplicated, conflicting, or poorly organized information can make it harder for the model to focus on what actually matters, and models are measurably less reliable at using information buried in the middle of a long input than information near the start or end. Adding more only helps when what's added is actually relevant.
Prompting and Outputs
20. What is a system prompt?
The instruction given to a model before the actual conversation starts, setting its role, tone, or rules for the whole session, separate from anything the user types.
21. What is zero-shot prompting?
Asking a model to do a task with no worked examples at all — just a plain description of what you want, relying entirely on what it already learned during training.
22. What is few-shot prompting?
Including a small number of worked examples in the prompt itself, showing the model the pattern you want before asking it to do the real task. It usually improves reliability over describing the task in words alone. Learn more: Few-Shot Prompting.
23. What is prompt chaining?
Breaking a task into a sequence of smaller prompts, where each step's output feeds into the next step's input, instead of asking for the whole thing in one request. It can make a complex task more reliable, at the cost of more requests and more latency.
24. What are structured outputs?
Forcing a model's response to match a schema you define — a fixed shape your code can parse directly. Instead of allowing "John is 32 and lives in Dublin," an application might require {"name": "John", "age": 32, "city": "Dublin"}, because predictable structure is what lets software process the response without guessing at its shape. Learn more: Structured Outputs.
25. How would you make an LLM return valid structured data?
Use real schema enforcement, not a prompt instruction. Asking a model to "only return JSON" is a request it can still drift from; enforced structured output blocks the model from producing anything outside the schema in the first place.
26. What is the difference between prompting and fine-tuning?
Prompting changes what's supplied to the model for one request — instructions, examples, context. Fine-tuning changes the model's behavior by training it further on examples beforehand. Prompting is cheap and instant to change; fine-tuning is heavier but can be worth it when the same instructions would otherwise repeat on every single call.
Giving LLMs External Capabilities
27. What is tool calling?
Tool calling lets a model request that a specific action run, with specific arguments, instead of only writing text back. The model never executes anything itself — it returns a structured request, and the application's code is what actually runs it. Learn more: Tool Calling.
28. What is function calling?
Function calling is tool calling applied to a custom function you wrote yourself, as distinct from a built-in capability a vendor supplies. Learn more: Function Calling.
29. Why would an LLM need tools?
Because on its own, a model can only produce text — it can't look up current information, run a calculation reliably, or take a real action like sending an email. Tools give it a way to request that something real happen, while the model itself still just decides what to ask for.
30. What is RAG?
A technique that retrieves relevant information and gives it to a model as context before it answers, instead of relying only on what the model learned during training. Learn more: RAG.
31. How does RAG help an LLM answer questions using external information?
It searches a store of documents for the passages most relevant to the question and hands them to the model alongside the question, so the model can answer using material it was never trained on — current documents, private company data — rather than only what it already knows.
32. What are embeddings and how are they used with LLM applications?
An embedding is a piece of text converted into numbers positioned so that text with similar meaning ends up near other text with similar meaning. LLM applications use them to power retrieval — finding the passages relevant to a question by meaning, not by matching exact words. Learn more: Embeddings.
33. What is AI memory, and how is it different from an LLM's context window?
The context window is the information the model can work with during one particular request — think of it as what's currently on the desk. AI memory is information stored elsewhere that an application deliberately saves and can bring to the desk later, across separate requests or sessions. That's an analogy, not a literal description of how either is actually built, but it captures the real distinction: the context window doesn't persist on its own, and memory is what makes something persist on purpose.
Practical LLM Engineering Questions
34. How would you choose between two LLMs for an application?
Start from the task's real constraints: cost per request at your expected volume, how fast it needs to respond, how much context the task needs, whether it needs to call tools reliably, and whether it needs step-by-step reasoning or just a fast, simple answer. The strongest model on a leaderboard is often the wrong choice if it's slower or pricier than the task actually requires.
35. When would you use a smaller LLM instead of a larger one?
When the task is straightforward, latency matters, request volume is high enough that cost adds up, the smaller model already meets the quality you need, or the output is narrow and predictable. Classifying support tickets into five categories likely doesn't need your most capable reasoning model to do it well.
36. How would you reduce LLM hallucinations?
Ground answers in retrieved source information, use RAG where the task calls for it, improve the instructions and context you give the model, require citations where that's appropriate, validate important outputs, use tools for calculations or current data instead of asking the model to produce them from memory, and add human review for high-impact decisions. None of these guarantee hallucinations never happen — treat it as a risk to manage, not a bug you fully close.
37. How would you reduce LLM latency?
Send less context, use a smaller or faster model where the task allows it, stream the response back instead of waiting for the whole thing, and avoid unnecessary sequential model calls where one call could do the job.
38. How would you reduce LLM API costs?
There's no single fix: use a smaller or cheaper model when the task doesn't need the strongest one, avoid sending context that isn't needed, cache repeated results, cut unnecessary model calls, route easy and hard requests to different models, and make sure retrieval only pulls in what's genuinely useful.
39. How would you evaluate an LLM application's output?
Build a fixed set of real examples with a defined way to score the output, and run it again after any change — a prompt, a model version — rather than judging quality from a handful of examples that happened to look right. Learn more: AI Evals.
40. What should you consider before using an LLM in production?
Whether the task can tolerate occasional wrong or inconsistent output, what happens when it does produce something bad, whether cost stays reasonable at real volume, and whether there's a way to measure quality over time rather than assuming a good demo stays good after launch.
Distinctions worth remembering
An LLM is not Generative AI as a whole. An LLM is a language-focused generative model. Generative AI also includes image, audio, video, and other generative systems built on different techniques.
Context is not long-term memory. Being available in the current context is different from information being stored and retrieved across separate interactions.
An LLM is not an AI agent. An LLM can generate a response without being an agent. An agentic system uses an LLM as part of a larger process that chooses actions, calls tools, observes results, and continues toward a goal.
RAG is not fine-tuning. RAG supplies retrieved information at the moment a request is made. Fine-tuning changes the model's behavior through additional training beforehand.
Tool calling is not the model performing the action itself. The model requests a tool; the surrounding application controls execution, permissions, and the result.
Going deeper on one area
For a broader revision pass: AI Interview Questions and Answers. For generative AI across images, audio, and video: Generative AI Interview Questions. For practical application-building judgment: AI Engineer Interview Questions. For retrieval specifically: RAG Interview Questions and Answers. For prompting specifically: Prompt Engineering Interview Questions.