On this page
The seven layers in one table
Layer 1: Interface
Layer 2: Orchestrator
Layer 3: LLM layer
Layer 4: Tools
Layer 5: Memory
Layer 6: Guardrails
Layer 7: Observability
Why most candidates stop at layer three
How to spend a forty-five minute round
Frequently asked questions
Related reading
Design an AI Agent: The 7-Layer Blueprint


On This Page
The seven layers in one table
Layer 1: Interface
Layer 2: Orchestrator
Layer 3: LLM layer
Layer 4: Tools
Layer 5: Memory
Layer 6: Guardrails
Layer 7: Observability
Why most candidates stop at layer three
How to spend a forty-five minute round
Frequently asked questions
Related reading
Short answer: a complete design covers seven layers. Work arrives through an interface, an orchestrator runs the plan, act, observe loop, and an LLM layer routes each model call. Below those sit the tools the agent may call, the memory it carries between steps, the guardrails that limit what it may do, and the observability that records what it did and what it cost.
"Design an AI agent" has become the new "Design Twitter." It now appears in system design rounds at AI labs, at large technology companies, and at any product team adding an assistant to an existing application.
The seven layers are not graded equally. Most candidates describe the first three and stop there, because routing a model call feels like the interesting part of the problem. Senior candidates spend close to half the round on guardrails and observability, which are the two layers that decide whether the agent can run in production.
This article walks through each layer, names the decision an interviewer expects you to make inside it, and closes with a time budget for a forty-five minute round.

The seven layers in one table
| # | Layer | What it decides | The decision to state out loud |
|---|---|---|---|
| 1 | Interface | How work reaches the agent | Chat, API, or event, and the latency budget |
| 2 | Orchestrator | The plan, act, observe loop | Step limit, time limit, token budget |
| 3 | LLM layer | Which model runs each call | Routing policy and a fallback provider |
| 4 | Tools | What actions are possible | Tool count, schemas, and retry behavior |
| 5 | Memory | What the agent remembers | Context window policy and long-term store |
| 6 | Guardrails | What the agent may not do | Permission scope and approval points |
| 7 | Observability | What you learn after the run | Traces, evaluation set, cost per task |
Layer 1: Interface
The interface is how work reaches the agent, and three shapes cover almost every question you will be asked. A chat session where a person is waiting, a direct API call from another service, and an event trigger such as a new support ticket or an uploaded document.
The decision that matters is whether a person is waiting for the result. A waiting person limits the whole run to a few seconds, which in practice allows two or three model calls and little else. An event-triggered run can take minutes, so it needs a job queue, a run identifier the caller can poll, and a way to deliver the result later.
State the budget as a number rather than a quality. Something like "chat, so under two seconds to the first token, and the full run under twenty seconds" tells the interviewer you have already constrained every layer below.
Layer 2: Orchestrator
The orchestrator is the loop that makes this an agent instead of a pipeline. It plans a next step, calls a tool to act, observes what came back, and then decides again, repeating until the goal is met or a limit stops it.
Name the limits explicitly, because an unbounded loop is the most common correctness problem in agent systems. Three numbers are enough: a maximum of about twenty steps, a wall clock limit of a few minutes, and a token budget for the whole run. Without them, an agent that misreads a tool result will repeat the same call until someone notices the invoice.
The second decision is where run state lives. Keep the step history in a database row or a queue message rather than in process memory, so a worker that crashes halfway can resume the run instead of repeating work the user already paid for.

Layer 3: LLM layer
This layer decides which model handles each call, and it exists because one model for everything is both slow and expensive. Route cheap, high-volume calls such as classification and extraction to a small model, and reserve the large model for planning and for the final written answer.
Two reliability decisions belong here. The first is a fallback provider or a second model, because a provider outage otherwise takes your whole product offline. The second is caching, since agents repeat near-identical calls more often than most engineers expect.
Cost control sits here as well. Set a ceiling per run, measure tokens per completed task, and say plainly what you would do when a task exceeds the ceiling, which is usually to stop and hand the work to a person.
Layer 4: Tools
A tool is a function the model is allowed to call, described by a name, a schema of its inputs, and a short description the model reads to decide when the tool applies. Tools are what turn text generation into action, so this layer is where an agent stops being a chatbot.
There are two common ways to expose them. Function calling, where tool definitions are sent along with each model call, and MCP, short for Model Context Protocol, an open standard that lets a server publish a set of tools that any compatible model can use.
Every tool needs four properties: a timeout, a retry policy, an idempotency key so that a repeated call is safe, and a permission scope. Keep the number of tools small, roughly five to ten per agent, because selection accuracy drops as the list grows and the descriptions start to overlap.
Layer 5: Memory
Memory splits cleanly in two, and saying so immediately is worth real points. Short-term memory is the context window, which is the text the model can see inside a single call. Long-term memory is whatever survives between runs, normally a vector store for retrieval by meaning plus a short written summary of the user and the task.
The design problem is that step history grows faster than the context window allows. The standard answer is compaction: once the history passes a threshold, replace the older steps with a generated summary and keep the most recent steps in full.
Say what you deliberately do not store. Keeping raw tool outputs out of memory, and writing only the fields the agent needs later, is both a cost decision and a privacy decision.
If you want a structured path through all of this, Grokking the AI System Design Interview covers agent architecture and the rest of the AI design canon with worked numbers and timed case studies.
Layer 6: Guardrails
Guardrails limit what the agent may do, and the important point is that they are checks written in code around the model, not instructions written inside the question you send it. A model can be talked out of its instructions, and a permission check cannot.
Four of them cover most designs:
- Input filtering. Reject injected instructions arriving inside documents, tickets, or web pages the agent reads.
- Permission scope. Give each run a token with the narrowest access that completes the task, not the application credentials.
- Output filtering. Check generated text for leaked secrets and unsafe content before it reaches a user or another system.
- Human approval. Require a person to confirm any action that cannot be undone, such as a refund above a threshold, an outbound email, or a deletion.
The sentence to say out loud is this: every action the agent can take is either reversible, approved by a person, or capped by a number. That single rule answers most of the follow-up questions an interviewer has about damage.

Layer 7: Observability
Observability is how you learn what the agent did after it ran, and agents need more of it than ordinary services because the same input can produce different behavior twice. Three parts are expected.
A trace per run, recording every step, every tool call with its inputs and outputs, latency, and token count. An evaluation set, which is a fixed collection of tasks with known good outcomes that you run on every change to a model, a tool, or an instruction. And cost measured per completed task rather than per token, because that is the number a business can actually use.
Add one alert that most candidates miss. Watch the share of runs that finish successfully without human help, since a decline in quality appears there long before it appears in error rates.
Why most candidates stop at layer three
The first three layers are familiar territory for anyone who has used a model API, so they are comfortable to talk about and they consume time easily. The result is an answer that describes a demonstration rather than a system, and interviewers recognize the pattern quickly.
Layers four through seven are where the difficult questions live. What happens when a tool returns a wrong result confidently, how a run resumes after a crash, who approves a refund, and how you prove next month that quality has not declined. Reaching those layers is what separates a passing answer from a strong one.
| Answer that stops at layer 3 | Answer that covers all seven |
|---|---|
| Model routing and instruction design | Routing plus a bounded loop and resumable state |
| Tools mentioned by name | Tools with schemas, timeouts, and scopes |
| "We add safety checks" | Named checks, with approval for irreversible actions |
| No measurement | Traces, evaluation set, and cost per task |
How to spend a forty-five minute round
| Minutes | What to cover |
|---|---|
| 0 to 5 | Requirements, one concrete task, success criteria |
| 5 to 10 | Interface and the latency budget |
| 10 to 20 | Orchestrator loop, limits, and run state |
| 20 to 30 | Tools and memory, with one worked example each |
| 30 to 40 | Guardrails and observability |
| 40 to 45 | Failure modes, cost, and what you would build first |
Start by naming one concrete task the agent must complete, such as resolving a refund request from beginning to end. A single worked example keeps every later layer specific, and it stops the discussion from drifting into general statements about artificial intelligence.
Frequently asked questions
What is the difference between an AI agent and an AI workflow?
In a workflow, an engineer fixed the order of steps in advance and the model fills in the content. In an agent, the model chooses the next step at run time, based on what it has seen so far. The test is simply who decided the order, and when.
Do I need to name specific models or providers?
No, and naming them can date your answer. Describe the routing policy instead, such as a small model for classification and a large model for planning, with a second provider available for failover.
How many tools should an agent have?
Roughly five to ten. Beyond that, tool descriptions start to overlap and the model picks the wrong one more often. If you need more, split the work across several agents, each with its own narrow set.
What should I say about prompt injection?
Treat every document, ticket, and web page the agent reads as untrusted input. Defend with permission scope and approval steps rather than with instructions, because the model cannot reliably ignore text that tells it to misbehave.
Is a multi-agent design always better than a single agent?
No. One capable agent with good tools handles most problems, and every extra agent adds cost, latency, and places where coordination can fail. Add agents for context isolation, parallel work, or genuine specialization, and say which reason applies.
Which layer do interviewers probe hardest?
Guardrails and observability, because they reveal whether you have operated a system like this or only built one. Prepare a concrete example for each: one irreversible action you place behind approval, and one number you watch after release.
Related reading
What our users say
Tonya Sims
DesignGurus.io "Grokking the Coding Interview". One of the best resources I’ve found for learning the major patterns behind solving coding problems.
Roger Cruz
The world gets better inch by inch when you help someone else. If you haven't tried Grokking The Coding Interview, check it out, it's a great resource!
Eric
I've completed my first pass of "grokking the System Design Interview" and I can say this was an excellent use of money and time. I've grown as a developer and now know the secrets of how to build these really giant internet systems.
Access to 50+ courses
New content added monthly
Certificate of completion
$31.08
/month
Billed Annually
Recommended Course

Grokking the Object Oriented Design Interview
60,674+ students
4.2
Learn how to prepare for object oriented design interviews and practice common object oriented design interview questions. Master low level design interview.
View CourseRead More
Scalability in System Design: The Complete Guide to Techniques, Patterns & Trade-offs
Arslan Ahmad
How to Prepare for an AI/ML System Design Interview (Complete 2026 Roadmap)
Arslan Ahmad
Vibe Coding 101: A Beginner’s Guide to AI-Assisted Development
Arslan Ahmad
Cache Invalidation: Methods, Strategies, and Trade-offs
Arslan Ahmad