In this guide
AI Engineer Interview Questions and Answers
Forty scenario-heavy questions on what it actually takes to build and operate a modern LLM application — model choice, context, retrieval, tools, agents, evaluation, and what breaks in production. This is a different test from the broad AI Interview Questions page: less "can you define this term," more "how would you actually build and run this system."
Understanding an AI Engineer's Work
1. What does an AI Engineer do?
An AI Engineer builds and operates applications powered by AI models — deciding which model to use, how to feed it the right context, how to give it tools when it needs to take real actions, and how to keep the whole system reliable, fast, and affordable once real users depend on it. The job is closer to building and running a system than to training a model.
2. How is an AI Engineer different from a Machine Learning Engineer?
A Machine Learning Engineer typically builds and trains models from data — feature engineering, training runs, model architecture. An AI Engineer typically builds applications on top of models that already exist, calling them through an API, and focuses on prompting, retrieval, tool access, and reliability rather than training. The two roles increasingly overlap, but most day-to-day AI Engineer work never touches a training run at all.
3. What are the main components of a modern LLM application?
A model to call, a way to give it the right context — a system prompt, retrieved documents, conversation history — tools it can call for real actions, and a layer that handles errors, logging, and evaluation around all of it. Most of the engineering work is in that surrounding layer, not in the model itself.
4. What happens from the moment a user sends a prompt until an LLM application returns an answer?
The application assembles the context — system instructions, any retrieved documents, the conversation so far, and the user's message — and sends it to the model. The model either answers directly or requests a tool call; if it's a tool call, the application executes it and sends the result back for the model to use in a follow-up response. Once the model produces a final answer, the application may validate or post-process it before it reaches the user.
5. How do you choose the right AI model for an application?
Start from the task's actual constraints, not a leaderboard: how fast a response needs to be, how much a request can cost at your real volume, how much context the task needs, whether it needs to call tools reliably, and whether it needs step-by-step reasoning or just a fast, simple answer. The strongest available model is often the wrong choice if it's slower or more expensive than the task requires.
Building Reliable LLM Applications
6. What is the difference between a system prompt and a user prompt?
A system prompt sets the model's role, tone, and rules for the whole session, written by the application before any conversation starts. A user prompt is whatever the person actually asks in the moment. The system prompt stays constant across a conversation; the user prompt changes with every message.
7. Why is context important in an LLM application?
Because a model only knows what's actually in front of it for that one request — it has no memory of anything else, no matter how obvious it seems. An application that doesn't hand over the right instructions, history, or supporting facts gets a plausible-sounding answer built on whatever the model happens to already know, which is often wrong for anything specific to your data or your users.
8. What is context engineering?
Deciding what information actually goes into a model's context for a given task — which instructions, which retrieved documents, how much conversation history — rather than treating everything as if it should just be included. It's a design decision with real tradeoffs: more context can mean better answers, but it also costs more and can bury the one fact that mattered among everything else that didn't. Learn more: Context Engineering.
9. What happens when an application exceeds a model's context window?
One of two things: the request fails outright with an error, or the application silently drops the oldest or least-recent part of the input to make room. The second case is the dangerous one — the application keeps answering, just without whatever got cut, which can be an early instruction, a fact from a document, or a tool result the answer depended on. Learn more: Context Window.
10. What are structured outputs and when would you use them?
Structured output forces a model's response to match a schema you define, so code downstream can parse it reliably instead of guessing at its shape. Reach for it whenever something other than a person — your code, a database, another system — has to read the response directly. Learn more: Structured Outputs.
11. How would you make an LLM return reliable JSON?
Use real schema enforcement, not a prompt instruction. Asking a model to "only return JSON" is a request it can still drift from, especially in a longer response; enforced structured output blocks the model from producing anything outside the schema in the first place, which is a mechanical guarantee a prompt instruction never gives you.
12. What causes hallucinations, and how would you reduce them?
A model is always predicting plausible next text — a hallucination is what that looks like when there's nothing grounding the specific answer in fact. Reducing it means grounding answers in retrieved material, giving the model permission to say it doesn't know, and constraining output for tasks with a fixed set of valid answers — not a single fix, and none of them eliminate it completely. Learn more: Hallucination.
13. When would you use a smaller model instead of the most capable model?
When the task doesn't need the extra capability — classifying an incoming request, extracting a few fields, a simple lookup — a smaller model is usually faster and cheaper with no real drop in quality for that job. A common pattern is routing: use a small model to classify how hard a request is, and only send the genuinely difficult ones to a larger, more expensive model.
RAG and Retrieval
14. When should you use RAG?
When the model needs current or private information it wasn't trained on, that information changes too often to retrain around, or the answer needs to be traceable back to a real source. Learn more: RAG.
15. How does a basic RAG pipeline work?
A question comes in, the system searches a store of pre-processed documents for the passages most relevant to it, and hands those passages to the model alongside the question. The model answers from what it was given, rather than from its own training data alone. Learn more: RAG.
16. Why are documents split into chunks?
Because embedding models and context windows both have size limits, and retrieval works better on a focused piece of text than on an entire document. Where you cut matters — split in the wrong place and a table or a clause gets separated from the sentence that explains it. Learn more: Chunking.
17. What are embeddings used for?
Embeddings convert text into numbers positioned so that text with similar meaning ends up near other text with similar meaning, which is what lets a system search by meaning instead of by matching exact words. Learn more: Embeddings.
18. What is a vector database?
A database built to quickly find the embeddings most similar to a given query out of millions of them — the piece of a RAG pipeline that makes fast retrieval by meaning possible at scale. Learn more: Vector Database.
19. What is semantic search?
Search that matches by meaning rather than by exact keywords — the query "cheap flights" can match a result that says "affordable airfare" even though no word overlaps, because their embeddings land near each other in meaning. It's the retrieval technique embeddings actually enable.
20. What is hybrid search?
Combining semantic search with traditional keyword search, so an exact term — a product code, an acronym — that meaning-based search alone might miss still gets found. It fixes a real gap, at the cost of now needing to tune how the two result sets get combined.
21. What is reranking?
A second, more careful scoring pass over a shortlist a first-stage search already narrowed down, to push the genuinely most relevant results higher. It's a fix for ranking within the shortlist, not a fix for something missing from the shortlist entirely. Learn more: Reranking.
22. Your RAG application retrieves irrelevant documents. How would you investigate the problem?
Start by reading the actual retrieved chunks against the actual question, not aggregate metrics. Check whether the embedding model captures the right kind of similarity for this domain, whether chunks are cut in a way that preserves meaning, and whether the query is phrased in a way that matches how the source material is worded. Irrelevant results usually trace back to embedding, chunking, or query phrasing — not to the model generating the final answer.
23. Your RAG system retrieves the correct document but still gives a poor answer. What would you check?
Retrieval succeeding and generation succeeding are separate problems. Check whether the specific relevant part of the document actually made it into the model's context, whether irrelevant retrieved content is crowding it out, whether the instructions are clear about how to use the retrieved material, whether multiple retrieved passages conflict with each other, and whether the model itself is capable enough for the reasoning the task needs. The right document being retrieved only guarantees the ingredients were available, not that they were used well.
Tools and AI Agents
24. What is tool calling?
Tool calling lets a model request that a specific action run, with specific arguments, instead of only writing text back. The model never executes anything itself — it returns a structured request, and the application's own code is what actually runs it. Learn more: Tool Calling.
25. What is the difference between tool calling and function calling?
Function calling is tool calling applied to a custom function you wrote yourself, as distinct from a built-in capability a vendor supplies, like web search. The two terms are used interchangeably in casual conversation, and the underlying mechanism — the model requests, your code executes — is identical either way. Learn more: Function Calling.
26. What is an AI agent?
A system that decides its own sequence of actions to reach a goal — deciding, acting, checking the result, and deciding again — rather than following a fixed set of steps someone wrote in advance. Learn more: AI Agent.
27. When would you use an AI agent instead of a fixed workflow?
A fixed workflow is better when the steps are already known and predictable — it's cheaper, easier to test, and can't make a wrong judgment call it wasn't supposed to make. An agent earns its cost when the right path can't be known in advance, because it depends on what earlier steps turn up. More agentic isn't automatically better; it's a tradeoff against predictability, not a straightforward upgrade.
28. How would you prevent an AI agent from running forever?
A hard limit on the number of steps it can take, detection for repeating the same action without making progress, and a clear stopping condition that tells it when the goal is actually met, not just attempted. Without at least a step limit, an agent that gets stuck in a bad loop has no reason to ever stop on its own.
29. What should happen when an AI tool fails?
The failure should come back to the model as a real, informative result — not silence, and not a crash — so the model can decide whether to retry, try a different approach, or tell the user it couldn't complete the task. An agent that gets no signal a tool failed will often act as if it succeeded, which is worse than the tool simply not existing.
30. When should an AI system require human approval?
Before anything irreversible or externally visible — sending money, sending an email, deleting data, posting publicly. The riskier and harder to undo an action is, the stronger the case for a person reviewing it before it happens, especially once the system is making its own decisions about when to take it. Learn more: Prompt Injection.
Evaluation and Production
31. How do you evaluate an LLM application?
Build a fixed set of real examples with a defined way to score the output, and run it again every time something changes — a prompt, a model version, a pipeline step — instead of judging quality from a handful of examples that happened to look right. Learn more: AI Evals.
32. What are AI evals?
A repeatable test — real examples plus a scoring method — used to measure how well an AI system performs on your actual task, rather than a generic benchmark. Learn more: AI Evals.
33. Why is testing an AI application different from testing traditional software?
Traditional software either produces the correct output or a bug — there's usually one right answer to check against. An LLM's output can be reasonable, unreasonable, or subtly wrong in ways that still look fluent, so testing often needs a scoring method, sometimes another model acting as a judge, rather than a simple pass/fail check, and the same input can legitimately produce slightly different outputs across runs.
34. How would you monitor an AI application in production?
Track the things that quietly go wrong without an obvious crash: response quality drifting on real traffic, latency — how long a request takes to get a response back — and cost per request, how often tool calls fail, and how often the model refuses or gives a low-confidence answer. Logging enough of each real request to reconstruct what happened is what makes any of this possible to investigate after the fact.
35. How would you reduce LLM application latency?
Send less context, use a smaller or faster model where the task allows it, stream the response back instead of waiting for the whole thing, and avoid unnecessary sequential model calls where a single call could do the job. Retrieval and tool calls both add real time on top of the model's own response time, so cutting an unnecessary one often helps more than optimizing the model call itself.
36. How would you reduce AI API costs?
There's no single fix. Use a smaller or cheaper model when the task doesn't need the strongest one, avoid sending context that isn't actually needed, cache repeated results where the same request comes up again, cut unnecessary model calls, route easy and hard requests to different models, control how long responses are allowed to run, and make sure retrieval only pulls in what's genuinely useful. A common pattern: use a small model to classify how hard an incoming request is, and only send the genuinely difficult ones to a more expensive reasoning model — model choice is an engineering decision, not just "always use the best model."
37. What would you log when debugging an LLM application?
The full input sent to the model, not just the user's message — the system prompt, any retrieved context, and conversation history, since the model saw all of it together. Also every tool call made, its arguments, and its result, plus the model version and any parameters like temperature that were in effect. Without the full input, a wrong answer is nearly impossible to reproduce or explain.
Security and Practical Scenarios
38. What is prompt injection, and why should an AI Engineer care about it?
Prompt injection is when text an application treats as data — an email, a document, a web page — gets read by the model as an instruction instead, because the model has no reliable way to tell the two apart. It matters specifically to an AI Engineer because an agent with real tools doesn't just produce a bad reply when this happens — it can take a real, damaging action. Learn more: Prompt Injection.
39. How would you safely give an AI agent access to sensitive tools or data?
Give it the narrowest set of tools that does the job, require human approval for anything irreversible or externally visible, and treat every tool result and every piece of fetched content as untrusted, even from a source you'd normally trust. Don't let the component reading untrusted content also hold the credentials that matter.
40. An AI application worked well during development but performs poorly for real users. How would you investigate it?
Real user input is messier than test cases. Check whether actual queries are phrased in ways your test set never covered, whether retrieval is missing documents that matter for real usage patterns, whether context is quietly getting dropped on longer real conversations than you tested with, and whether an eval was ever built to catch this kind of drift before users did. "It worked in my testing" and "it works for real traffic" are different claims, and the gap between them is usually where the actual bug lives.
Going deeper on one area
For a broader revision pass on definitions: AI Interview Questions and Answers. For harder, scenario-based questions on one specific topic: RAG Interview Questions or Agent & Subagent Interview Questions. For the full RAG pipeline end to end: RAG Interview Questions and Answers. For agent decision-making specifically: Agentic AI Interview Questions. For designing a whole system end to end: AI System Design Interview Questions.