AI System Design Interview Questions and Answers
Thirty questions on designing real AI applications and explaining the engineering decisions behind them. This isn't a definitions page — there's rarely one universally correct architecture, and a strong answer works through requirements, architecture, trade-offs, failure modes, and evaluation, in that spirit if not always that exact order.
In this guide
- A framework to use in the interview
- Basic Architecture Decisions
- RAG System Design
- Tool and Agent System Design
- Reliability and Failure Handling
- Cost, Latency and Scale
- Security and Evaluation
- Detailed scenario: internal company assistant
- Detailed scenario: customer support AI
- Detailed scenario: AI coding assistant
- Trade-offs, not universal answers
- Think about failure first
- Evaluation is part of the architecture
- Principles worth remembering
- Going deeper on one area
A framework to use in the interview
When asked "design an AI system that...", work through:
- What does the user actually need?
- What information does the AI need to do this?
- Which model is appropriate for the task?
- Does it need external knowledge it wasn't trained on?
- Does it need tools, or the ability to take real actions?
- Does it need memory across sessions?
- What output does the rest of the application require?
- What could go wrong?
- How will you evaluate it?
- What are the cost, latency, privacy, and security constraints?
This isn't a rigid sequence to follow in exactly this order — it's a checklist for making sure you haven't skipped something an interviewer will ask about anyway.
Basic Architecture Decisions
1. How would you design a basic LLM-powered application?
Start from what the user needs, then work outward: decide what information the model needs to do the task, pick a model that fits the task's real constraints, decide whether it needs external knowledge or the ability to take actions, decide what output shape the rest of the application requires, and only then think about memory, cost, and monitoring. The mistake to avoid is picking a model or an architecture pattern first and fitting the requirements to it afterward.
2. How would you choose the right model for an AI application?
Against the task's real constraints, not a leaderboard: how fast it needs to respond, how much a request can cost at real volume, how much context the task needs, whether it needs to call tools reliably, and whether it needs multi-step reasoning or just a fast, simple answer. The strongest available model is often the wrong choice if it's slower or pricier than the task actually requires.
3. When would you use one model versus multiple models?
Multiple models earn their place through routing: a smaller, cheaper model handles the easy or high-volume cases, and a more capable model only gets used for requests that genuinely need it. One model is simpler and usually the right starting point — add routing once you have real evidence that a meaningful share of requests don't need your most capable, most expensive option.
4. When would you use RAG?
When the model needs current or private information it wasn't trained on, that information changes too often to retrain around, or the answer needs to be traceable back to a source. If the relevant material is small and stable enough to paste directly into the prompt, RAG is probably unnecessary machinery.
5. When would you use tools or APIs?
When the task needs the application to actually do something — look up live data, take an action in another system, run a calculation — rather than just produce text. A model on its own can only write an answer; tools are what let a decision become a real effect.
6. When would you use an AI agent instead of a fixed workflow?
When the right sequence of steps can't be known in advance, because it depends on what earlier steps turn up. If the steps are already predictable, a fixed workflow is cheaper, easier to test, and easier to audit — reach for an agent only once that predictability genuinely isn't available.
7. When does an AI application need memory?
When something needs to persist across separate sessions or requests — a user's preferences, a fact from an earlier conversation — rather than being available only within the current context window. Not every application needs this; a lot of real usage is genuinely one-off.
8. When should an AI system require human approval?
Before anything irreversible or externally visible — sending money, sending an email, deleting data, deploying code. The more costly and harder to undo an action is, the stronger the case for a person reviewing it before it happens.
RAG System Design
9. Design an AI assistant that answers questions using a company's internal documents.
Company Documents
↓
Parse + Chunk
↓
Create Searchable Index
User Question
↓
Retrieve Relevant Content
↓
Build Context
↓
LLM
↓
Answer + Sources
Documents get parsed, cleaned, split into chunks, and indexed for search. At query time, the employee's question retrieves the relevant chunks, those get assembled into context, and the model answers from that context and cites its sources. Every real system varies the specifics — chunking strategy, whether reranking is used, how citations are formatted — but this is the shape underneath most of them.
10. How would you design RAG for hundreds of thousands of documents?
The pipeline is the same shape as a small system; what changes is what has to scale. Indexing needs to run incrementally as documents are added or updated, not as a full rebuild every time. Retrieval needs metadata filtering or hybrid search so a query doesn't compare against the entire collection unfiltered. And reranking becomes more valuable, since a first-stage search over that much content surfaces more noise to sort through.
11. How would you ensure users only retrieve documents they are authorized to access?
Apply permission checks as part of retrieval itself, not after the fact — filter by what the requesting user is allowed to see before or during the similarity search, not by generating an answer first and hoping nothing sensitive leaked in. Relevant does not automatically mean authorized.
12. How would you handle frequently changing documents?
Re-index incrementally whenever a source document changes, rather than only on a periodic full rebuild, and track the gap between when content changes and when the index reflects it. For genuinely fast-changing information, consider whether RAG is even the right mechanism versus calling a live tool for the current value.
13. What would you do if retrieval quality was poor?
Read actual retrieved chunks against actual questions rather than trusting an aggregate score. Check whether the embedding model fits the domain, whether chunking preserves meaning, and whether queries are phrased in a way that matches how the source material is worded — poor retrieval is usually one of those three, not something to fix by changing the model that generates the final answer.
14. How would you evaluate the RAG system?
Separately score retrieval quality — did it find the right information — and answer quality — did the model use it well — rather than one combined pass or fail on the final answer. Depending on the application, also track faithfulness to sources, citation usefulness, latency, cost, and whether permissions were actually enforced. Learn more: RAG Evaluation.
Tool and Agent System Design
15. Design an AI assistant that can use external tools.
Start from the narrowest set of tools the task actually needs, with each tool's inputs and outputs clearly defined. The model decides which tool to call and with what arguments; the application code actually executes the call, checks the result, and decides whether to hand that result back to the model for the next step. Treat every tool result as something to validate, not something to trust automatically.
16. How would you decide which tools an AI system can access?
By what the specific task requires, not by what's available. Giving a system every tool it might conceivably need "just in case" expands what can go wrong without expanding what it can actually accomplish for this task.
17. How would you handle a tool failure?
Return the failure to the model as a real, informative result rather than silence or a crash, so it can decide whether to retry, try something else, or report that it couldn't complete the task. An agent that gets no signal a tool failed will often act as if it succeeded.
18. How would you prevent an AI agent from repeatedly calling tools?
A hard limit on the number of steps or tool calls it can make, detection for calling the same tool with the same or similar arguments without the situation changing, and a clear condition for when the task is actually done. Without at least a step limit, a stuck agent has no built-in reason to stop.
19. How would you design approval for sensitive actions?
Identify which actions are irreversible or externally visible ahead of time, and require a person to review the specific action and its context — not just click an unlabeled "approve" button — before it executes.
AI travel assistant may automatically:
✓ Search flights
✓ Search hotels
✓ Check weather
✓ Build an itinerary
Requires explicit approval:
! Purchase flight
! Charge credit card
! Cancel booking
Tool access should follow the minimum permissions required for the task — freedom to search and plan, approval gates on anything that spends money or can't be undone.
20. When would you use multiple agents instead of one agent?
When a task genuinely benefits from separating context, tools, or responsibilities that would otherwise all crowd into one agent's growing context. It adds real coordination cost, so it's worth it only once a single agent juggling everything has become the actual bottleneck, not by default.
Reliability and Failure Handling
21. How would you reduce hallucinations in a production AI system?
Ground answers in retrieved source material where the task depends on specific facts, give the model explicit permission to say it doesn't know, use tools for anything requiring current or precise data instead of asking the model to produce it from memory, and validate high-impact outputs rather than trusting them by default. None of these eliminate hallucination completely.
22. What would you do when the model is unsure or lacks enough information?
Design the application so "I don't have enough information" is an acceptable output, not a failure state the model is implicitly being pushed to avoid. A system that only accepts a direct answer leaves the model no honest way to flag a real gap.
23. How would you design fallbacks if the primary model or service fails?
Have a secondary path — a backup model, a cached previous answer, or a clear message that the service is temporarily unavailable — rather than letting a single point of failure take down the whole application. What the fallback should do depends on how costly a wrong or missing answer is for that specific use case.
24. How would you prevent malformed AI output from breaking the application?
Use real structured output enforcement so the model's response is mechanically constrained to a schema, and validate it again on the application side before it's used. "Please always return valid JSON" is a request the model can still drift from; schema enforcement plus a validation check is what actually prevents a malformed response from reaching code that assumes it's well-formed.
25. How would you monitor an AI system after deployment?
Track the things that go wrong quietly: response quality drifting on real traffic, how often tool calls fail, latency and cost per request, and how often the model refuses or gives a low-confidence answer. Logging enough of each real request to reconstruct what happened is what makes any of this actionable.
Cost, Latency and Scale
26. How would you reduce the cost of a high-volume LLM application?
There's no single fix: use a smaller model where it's genuinely sufficient, route simple and difficult requests to different models, avoid sending unnecessary context, cache appropriate repeated work, cut unnecessary model calls, control output length, and make sure retrieval only pulls in what's actually useful. None of this should mean sacrificing quality the application actually requires — cost reduction that breaks the product isn't a saving.
Request
↓
Determine complexity
↓
┌─────────────┬──────────────┐
↓ ↓
Simple Complex
↓ ↓
Smaller model More capable model
Model routing only makes sense once the routing decision itself is reliable enough that misrouting a hard request to a weak model doesn't cause more problems than it solves.
27. How would you reduce AI application latency?
Send less context, use a smaller or faster model where the task allows it, stream responses instead of waiting for the whole thing, and cut unnecessary sequential model calls where one call could do the job.
28. How would you design an AI application that needs to handle a large number of users?
Design for the same principles as any high-traffic system — caching, request batching where technically appropriate, and horizontal scaling of the surrounding application — plus AI-specific levers like model routing and controlling how much context each request actually needs.
Security and Evaluation
29. How would you protect an AI application from prompt injection?
Treat every piece of fetched or tool-supplied content as untrusted, give the system the narrowest tool access the task needs, require human approval for anything irreversible or externally visible, and don't let the component that reads untrusted content also hold the credentials that matter. There's no complete fix — treat it as a standing risk to contain through architecture, not a prompt instruction to eliminate. Learn more: Prompt Injection.
30. How would you evaluate an AI system before and after deployment?
Before deployment, build a fixed set of representative test cases and a scoring method, and use it to compare architecture or prompt decisions rather than trusting a handful of examples that happened to look right. After deployment, monitor real traffic for the same measures, plus anything only visible at scale, and treat a new failure found in production as a new test case to add, not a one-off to patch and forget. Learn more: AI Evals.
Detailed scenario: internal company assistant
Design an AI assistant that lets employees ask questions about company policies, procedures, and internal documentation.
Start with requirements: what documents exist, how often do they change, who can access which ones, are citations required, how accurate must answers be, and what should happen when the answer is genuinely unknown.
Internal Documents
↓
Parsing / Cleaning
↓
Chunking + Metadata
↓
Search / Retrieval Index
Employee Question
↓
Authentication + Permissions
↓
Retrieval
↓
Reranking if useful
↓
Context
↓
LLM
↓
Answer + Source References
Failure cases worth naming up front: outdated documents, conflicting policies, incorrect retrieval, unauthorized documents surfacing, insufficient information for a real answer, and hallucinated answers presented with confidence.
Evaluation: retrieval quality, answer correctness, faithfulness to source material, permission enforcement, citation usefulness, latency, and cost. This scenario is meant to demonstrate system thinking, not just knowledge of RAG terminology.
Detailed scenario: customer support AI
Design an AI customer-support assistant that can answer questions and perform limited account actions.
Customer Message
↓
Identify request
↓
┌──────────────┬──────────────┐
↓ ↓
Knowledge question Account action
↓ ↓
RAG Approved tools
↓ ↓
└───────────→ LLM ←──────────┘
↓
Response
Discuss authentication, permissions, tool access, knowledge retrieval, handling of sensitive data, approval for risky actions, tool failures, escalation to a human, logging, and evaluation. Don't assume an autonomous agent is automatically the best design here — a controlled, more fixed workflow can be safer for many account actions than giving a model broad discretion over them.
Detailed scenario: AI coding assistant
Design an AI assistant that can inspect code, suggest changes, and run tests.
Read repository
↓
Understand task
↓
Search relevant files
↓
Suggest / make change
↓
Run tests
↓
Observe result
↓
Continue or stop
Capability boundaries: reading code, searching the repository, running safe tests, and proposing changes can happen automatically. Modifying sensitive configuration, accessing secrets, deploying production code, and deleting resources should require approval. Good AI system design includes these boundaries as a first-class part of the design, not an afterthought bolted onto model selection.
Trade-offs, not universal answers
Every architecture decision trades something for something else. A larger model may reason better and handle harder tasks, at higher cost and often higher latency. A smaller model is cheaper and faster, and fine for simpler tasks, but may perform worse on genuinely difficult ones. Neither is universally better — the right choice depends on the application's actual requirements and measured performance against them, not a general reputation for being "the best model."
Think about failure first
A strong candidate asks "how can this fail" before asking "how do I build this." For an agent: the wrong tool gets chosen, a tool fails, it gets stuck in a loop, it misreads a result, it costs more than expected, or it takes an unsafe action. For RAG: the wrong document gets retrieved, the right one never gets retrieved at all, content is outdated or conflicting, unauthorized content surfaces, or generation is poor despite correct retrieval. For structured outputs: fields go missing, values are wrong, the schema doesn't match, or the data is syntactically valid but semantically nonsense. Good system design anticipates these rather than assuming the model behaves perfectly.
Evaluation is part of the architecture
Requirements
↓
Build system
↓
Representative test cases
↓
Evaluate
↓
Find failures
↓
Improve
↓
Evaluate again
Evaluation isn't something bolted on after the application is built — it's part of the design from the start. Production monitoring then extends the same loop, surfacing new failure patterns that a pre-launch test set couldn't have anticipated.
Principles worth remembering
Use the simplest architecture that reliably solves the problem — don't reach for an agent when a predictable workflow is sufficient. More context is not automatically better; supply what's relevant. The largest model is not automatically the best model; measure quality, latency, and cost against the actual task. AI output should not be trusted automatically — how much validation it needs depends on the impact of being wrong. Autonomy should have boundaries, with permissions and approval where the action warrants it. Security is part of the architecture from the start, not something added at the end. Evaluation is part of development, not a handful of manual demonstrations that happened to look good.
Going deeper on one area
For a broader revision pass: AI Interview Questions and Answers. For the full RAG pipeline: RAG Interview Questions and Answers. For agent decision-making: Agentic AI Interview Questions. For practical application-building judgment: AI Engineer Interview Questions.