In this guide
  1. AI Fundamentals
  2. Generative AI and LLMs
  3. Working With LLMs
  4. RAG, Agents and Modern AI Applications
  5. Reliability and Safety
  6. Practical Interview Questions
  7. Going deeper on one area

AI Interview Questions and Answers

Forty questions covering the concepts that come up most in AI interviews today — what they mean, how they fit together, and how to actually use them. Answers are written to be said out loud in an interview, not memorized as definitions. Where a concept has a full page on this site, it's linked so you can go deeper than this page goes.

AI Fundamentals

1. What is Artificial Intelligence?

Artificial intelligence is software built to perform tasks that normally need human judgment — understanding language, recognizing images, making a decision from incomplete information. It's a broad field, not one specific technique. Machine learning, deep learning, and generative AI are all approaches used to build it.

2. What is the difference between AI, Machine Learning, and Generative AI?

AI is the broad goal: software that performs tasks needing human-like judgment. Machine learning is the main approach used to get there today — instead of writing rules by hand, you train a system on examples and let it find the patterns itself. Generative AI is a specific application of machine learning, built to produce new content — text, images, audio — rather than just classify or predict something.

3. What is Machine Learning?

Machine learning is a way of building software that improves at a task by learning from examples, instead of following rules a person wrote by hand. You give it data where the right answer is already known, and training adjusts the system until it can produce the right answer for new examples it hasn't seen.

4. What is Deep Learning?

Deep learning is machine learning using neural networks with many layers stacked on top of each other. Each layer learns a slightly more complex pattern than the one before it — early layers in an image model might pick up edges, later ones combine those into shapes, then objects. It's the approach behind almost every major AI advance in the last decade, including large language models.

5. What is a neural network?

A neural network is a machine learning model loosely modeled on how neurons connect in a brain: layers of simple units, each doing a small calculation and passing its result to the next layer. No single unit understands anything on its own — the ability to recognize a pattern or generate an answer comes from how millions of these small calculations combine, with training adjusting the strength of each connection.

6. What is the difference between supervised and unsupervised learning?

Supervised learning trains on examples that already have the right answer labeled — show it a thousand photos marked "cat" or "dog" and it learns to tell them apart. Unsupervised learning gets data with no labels at all and has to find structure on its own, like grouping similar customers together without being told what the groups should be.

7. What is the difference between training and inference?

Training is when a model learns from data, adjusting its internal numbers over many passes until it gets better at the task. Inference is using that already-trained model to answer a real question or complete a real task, with no further learning happening. Training happens once, or occasionally; inference happens every single time someone actually uses the model.

Generative AI and LLMs

8. What is Generative AI?

Generative AI is a type of AI built to produce new content — text, images, audio, code — rather than just classify or predict something about existing content. A spam filter isn't generative; it labels an email. A model that writes the email is.

9. What is a Large Language Model (LLM)?

A large language model is a neural network trained on enormous amounts of text to predict what word — or part of a word — comes next in a sequence. That one skill, applied at scale, turns out to be enough to answer questions, write code, summarize documents, and hold a conversation.

10. What is a token?

A token is the small chunk of text a model actually reads and generates — often a whole word, sometimes just part of one. Text gets broken into tokens before a model processes it, and a model's context window is measured in tokens, not in words or characters. Learn more: Token.

11. What is a context window?

A context window is the maximum amount of text a model can read and act on for a single request — everything sent to it, and everything it writes back, shares that one fixed budget. Learn more: Context Window.

12. What is temperature in an LLM?

Temperature controls how much randomness goes into picking the next word. Low temperature makes the model stick to its most likely answer every time; high temperature lets it pick less-likely words more often, producing more varied and less predictable output. Learn more: Temperature.

13. What is an AI hallucination?

A hallucination is a confident, fluent answer that's factually wrong. The model isn't lying or guessing randomly — it's doing what it always does, predicting plausible-sounding text, just without anything grounding this particular answer in fact. Learn more: Hallucination.

14. What is a multimodal AI model?

A multimodal model can take in, or produce, more than one kind of content — text and images, or text and audio, rather than text alone. Asking it to describe a photo, or generate an image from a written description, uses its multimodal capability.

15. What is a reasoning model?

A reasoning model is trained to work through a problem in explicit intermediate steps before giving a final answer, rather than producing one directly. This tends to help on problems with a clear multi-step structure — math, multi-step planning — though it doesn't make the model immune to being confidently wrong, including about plain facts.

16. What are embeddings?

An embedding is a piece of text converted into a list of numbers, positioned so that text with similar meaning ends up near other text with similar meaning. That's what lets a system search by meaning instead of by matching exact words. Learn more: Embeddings.

Working With LLMs

17. What is prompt engineering?

Prompt engineering is writing and structuring the instructions you give a model to get a more reliable, more useful response — how you phrase a question, what examples you include, how you break down a task. It's the cheapest lever to pull before reaching for something heavier like retrieval or fine-tuning. Learn more: Prompt Engineering.

18. What is a system prompt?

A system prompt is the instruction given to a model before the actual conversation starts, setting its role, tone, or rules for the whole session, separate from anything the user types. It's how a product built on top of a model enforces consistent behavior across every conversation.

19. What is few-shot prompting?

Few-shot prompting means including a small number of worked examples in the prompt itself, showing the model the pattern you want before asking it to do the real task. It usually improves reliability over just describing the task in words, at the cost of using more of the context window. Learn more: Few-Shot Prompting.

20. What are structured outputs?

Structured output forces a model's response to match a schema you define — a fixed shape your code can parse directly — instead of just asking it nicely to format its answer a certain way. Learn more: Structured Outputs.

21. What is tool calling?

Tool calling lets a model request that a specific action run, with specific arguments, instead of only writing text back. The model never runs anything itself — it returns a structured request, and your code is what actually executes it. Learn more: Tool Calling.

22. What is function calling?

Function calling is tool calling applied to a custom function you wrote yourself, as opposed to a built-in capability a vendor supplies. The two terms are often used interchangeably in casual conversation. Learn more: Function Calling.

RAG, Agents and Modern AI Applications

23. What is Retrieval-Augmented Generation (RAG)?

RAG is a technique that retrieves relevant external information and gives it to a model as context before the model generates its answer. If an employee asks an AI assistant about the company's holiday policy, the application can retrieve the current policy document and hand the relevant section to the model before it answers. This helps the model answer using information that may not be in its training data. Learn more: RAG.

24. Why would you use RAG?

Use RAG when the model needs current or private information it wasn't trained on, when that information changes too often to retrain around, or when an answer needs to be traceable back to a real source. A support bot answering from this week's documentation is the typical case.

25. What is a vector database?

A vector database stores embeddings and is built to quickly find the ones most similar to a given query, out of millions of them. It's the piece of a RAG pipeline that makes fast retrieval by meaning possible at scale. Learn more: Vector Database.

26. What is an AI agent?

An AI agent is a system that decides its own sequence of actions to reach a goal — deciding, acting, checking the result, and deciding again — rather than following a fixed set of steps someone wrote in advance. Learn more: AI Agent.

27. What makes an AI system agentic?

How much of what happens next is decided by the model, rather than fixed in code. It's a scale, not a yes-or-no label — a system that makes one small decision on its own barely counts, while a full AI agent, deciding its whole sequence of actions, sits at the far end. Learn more: Agentic AI.

28. What is AI memory?

AI memory is what a system deliberately saves and reuses across separate requests or sessions — a user's preferences, a fact from an earlier conversation — since a model has no memory of its own beyond whatever's currently in its context window. Anything meant to persist has to be stored and re-supplied on purpose.

Reliability and Safety

29. How can you reduce AI hallucinations?

Ground answers in retrieved source material rather than the model's own memory, give it explicit permission to say it doesn't know instead of forcing an answer, and use a constrained output format for tasks with a fixed set of valid answers. None of these eliminate hallucination completely — treat it as a risk to manage, not a bug you fully close. Learn more: Hallucination.

30. What is prompt injection?

Prompt injection is when text your application treats as data — an email, a document, a web page — gets read by the model as an instruction instead, because the model has no reliable way to tell the two apart. Learn more: Prompt Injection.

31. What are AI guardrails?

Guardrails are checks placed around a model to catch or block bad behavior before it reaches a user or causes harm — filtering unsafe output, blocking a disallowed action, validating a response before it's used. They're a safety net around the model, not a change to the model itself. Learn more: Guardrails.

32. What are AI evals?

An eval is a repeatable test — a fixed set of real examples and a scoring method — used to measure how well an AI system performs on your actual task, instead of judging quality by trying a few prompts and going with a gut feeling. Learn more: AI Evals.

33. Why is human-in-the-loop important?

Because a model can be confidently wrong, and some actions are too costly to let it take unsupervised — approving a large refund, sending an external email, deleting data. Human-in-the-loop means a person reviews or approves specific actions before they happen, often the only reliable defense once an AI system has real tools and real permissions. Learn more: Human-in-the-Loop.

Practical Interview Questions

34. When would you use RAG instead of fine-tuning?

When the problem is that the model doesn't know your facts, not that it writes in the wrong style or format. RAG adds information at the moment a question is asked; fine-tuning changes the model itself beforehand and is better at reshaping how it responds than at reliably storing facts. If a model needs to answer from a document it's never seen, that's a retrieval problem, not a training one. Learn more: RAG vs Fine-Tuning.

35. When would you use an AI agent instead of a fixed workflow?

A fixed workflow is usually better when the steps are already known and predictable — it's cheaper, easier to test, and can't make a judgment call it wasn't supposed to make. An agent earns its cost when the right sequence of steps can't be known in advance, because it depends on what earlier steps find. A fixed invoice-processing pipeline doesn't need an agent, since the steps never change. Investigating why an application's performance suddenly dropped might, because the right path to the answer isn't known until you're partway through it. More agentic doesn't automatically mean better — it's a tradeoff, not an upgrade.

36. How would you choose an AI model for an application?

Start from the constraints that actually matter for the task: cost per request at your expected volume, how fast it needs to respond, how large a context window the task requires, whether it needs to call tools reliably, and whether the task needs step-by-step reasoning or just a fast, simple answer. Chasing the single "best" model on a leaderboard usually matters less than picking the one that fits these constraints for your specific use case.

37. How would you evaluate whether an AI application is working well?

Build an eval: a fixed set of real examples with a way to score the output, so quality can be measured and compared after any change, rather than judged by a handful of examples that happened to look right. Layer deterministic checks — exact-match rules with no judgment call involved — for anything with one correct answer, an LLM judge for subjective quality at scale, and periodic human review to keep the judge honest. Learn more: AI Evals.

38. What should you consider before giving an AI system access to tools?

Give it the narrowest set of tools that does the job, require human approval for anything irreversible or externally visible, and treat every tool result — and anything the tools read, like a fetched web page or document — as untrusted content the model could be manipulated by. The more real actions a system can take, the more a mistake or a manipulated input actually costs. Learn more: Prompt Injection.

39. How can you reduce the cost of an LLM application?

Send less: shorter prompts, retrieved context instead of a huge pasted document, and a smaller or cheaper model wherever the task doesn't need your strongest one. Caching repeated prompts, batching requests, and reserving a bigger, more expensive model for only the steps that genuinely need it are the other standard levers. Cost usually comes down to how many tokens you're sending and how often, more than which single model you picked.

40. What are common reasons an AI application fails in production?

A retrieval step that misses the right information, a prompt that worked in testing but breaks on real user phrasing, an agent given more tool access than the task actually needed, silently dropped context once a conversation grows past the model's context window, and nobody having built an eval to catch a regression before users did. Most production failures trace back to one of these, not to the model itself being incapable.

Going deeper on one area

This page stays broad on purpose. For harder, scenario-based questions on one specific topic: RAG Interview Questions covers retrieval, chunking, embeddings, and reranking; Agent & Subagent Interview Questions covers the agent loop, delegation, and multi-agent systems in more depth than the fundamentals above. For focused revision pages: LLM, Generative AI, Agentic AI, Prompt Engineering, and RAG interview questions. Asked to design a whole system instead of answer a definition: AI System Design Interview Questions.