On this page

The seven layers in one table

Layer 1: Interface

Layer 2: Orchestrator

Layer 3: LLM layer

Layer 4: Tools

Layer 5: Memory

Layer 6: Guardrails

Layer 7: Observability

Why most candidates stop at layer three

How to spend a forty-five minute round

Frequently asked questions

Related reading

Design an AI Agent: The 7-Layer Blueprint

Image
Arslan Ahmad
Design an AI agent in seven layers: interface, orchestrator, LLM layer, tools, memory, guardrails, and observability. Most candidates stop at layer three, and the round is decided by six and seven.
Image

The seven layers in one table

Layer 1: Interface

Layer 2: Orchestrator

Layer 3: LLM layer

Layer 4: Tools

Layer 5: Memory

Layer 6: Guardrails

Layer 7: Observability

Why most candidates stop at layer three

How to spend a forty-five minute round

Frequently asked questions

Related reading

Short answer: a complete design covers seven layers. Work arrives through an interface, an orchestrator runs the plan, act, observe loop, and an LLM layer routes each model call. Below those sit the tools the agent may call, the memory it carries between steps, the guardrails that limit what it may do, and the observability that records what it did and what it cost.

"Design an AI agent" has become the new "Design Twitter." It now appears in system design rounds at AI labs, at large technology companies, and at any product team adding an assistant to an existing application.

The seven layers are not graded equally. Most candidates describe the first three and stop there, because routing a model call feels like the interesting part of the problem. Senior candidates spend close to half the round on guardrails and observability, which are the two layers that decide whether the agent can run in production.

This article walks through each layer, names the decision an interviewer expects you to make inside it, and closes with a time budget for a forty-five minute round.

The seven layers of an AI agent design: interface, orchestrator, LLM layer, tools, memory, guardrails, and observability

The seven layers in one table

#LayerWhat it decidesThe decision to state out loud
1InterfaceHow work reaches the agentChat, API, or event, and the latency budget
2OrchestratorThe plan, act, observe loopStep limit, time limit, token budget
3LLM layerWhich model runs each callRouting policy and a fallback provider
4ToolsWhat actions are possibleTool count, schemas, and retry behavior
5MemoryWhat the agent remembersContext window policy and long-term store
6GuardrailsWhat the agent may not doPermission scope and approval points
7ObservabilityWhat you learn after the runTraces, evaluation set, cost per task

Layer 1: Interface

The interface is how work reaches the agent, and three shapes cover almost every question you will be asked. A chat session where a person is waiting, a direct API call from another service, and an event trigger such as a new support ticket or an uploaded document.

The decision that matters is whether a person is waiting for the result. A waiting person limits the whole run to a few seconds, which in practice allows two or three model calls and little else. An event-triggered run can take minutes, so it needs a job queue, a run identifier the caller can poll, and a way to deliver the result later.

State the budget as a number rather than a quality. Something like "chat, so under two seconds to the first token, and the full run under twenty seconds" tells the interviewer you have already constrained every layer below.

Layer 2: Orchestrator

The orchestrator is the loop that makes this an agent instead of a pipeline. It plans a next step, calls a tool to act, observes what came back, and then decides again, repeating until the goal is met or a limit stops it.

Name the limits explicitly, because an unbounded loop is the most common correctness problem in agent systems. Three numbers are enough: a maximum of about twenty steps, a wall clock limit of a few minutes, and a token budget for the whole run. Without them, an agent that misreads a tool result will repeat the same call until someone notices the invoice.

The second decision is where run state lives. Keep the step history in a database row or a queue message rather than in process memory, so a worker that crashes halfway can resume the run instead of repeating work the user already paid for.

The agent loop: plan, act, observe, repeat, with a step limit, a time limit, and a token budget stopping the loop

Layer 3: LLM layer

This layer decides which model handles each call, and it exists because one model for everything is both slow and expensive. Route cheap, high-volume calls such as classification and extraction to a small model, and reserve the large model for planning and for the final written answer.

Two reliability decisions belong here. The first is a fallback provider or a second model, because a provider outage otherwise takes your whole product offline. The second is caching, since agents repeat near-identical calls more often than most engineers expect.

Cost control sits here as well. Set a ceiling per run, measure tokens per completed task, and say plainly what you would do when a task exceeds the ceiling, which is usually to stop and hand the work to a person.

Layer 4: Tools

A tool is a function the model is allowed to call, described by a name, a schema of its inputs, and a short description the model reads to decide when the tool applies. Tools are what turn text generation into action, so this layer is where an agent stops being a chatbot.

There are two common ways to expose them. Function calling, where tool definitions are sent along with each model call, and MCP, short for Model Context Protocol, an open standard that lets a server publish a set of tools that any compatible model can use.

Every tool needs four properties: a timeout, a retry policy, an idempotency key so that a repeated call is safe, and a permission scope. Keep the number of tools small, roughly five to ten per agent, because selection accuracy drops as the list grows and the descriptions start to overlap.

Layer 5: Memory

Memory splits cleanly in two, and saying so immediately is worth real points. Short-term memory is the context window, which is the text the model can see inside a single call. Long-term memory is whatever survives between runs, normally a vector store for retrieval by meaning plus a short written summary of the user and the task.

The design problem is that step history grows faster than the context window allows. The standard answer is compaction: once the history passes a threshold, replace the older steps with a generated summary and keep the most recent steps in full.

Say what you deliberately do not store. Keeping raw tool outputs out of memory, and writing only the fields the agent needs later, is both a cost decision and a privacy decision.

If you want a structured path through all of this, Grokking the AI System Design Interview covers agent architecture and the rest of the AI design canon with worked numbers and timed case studies.

Layer 6: Guardrails

Guardrails limit what the agent may do, and the important point is that they are checks written in code around the model, not instructions written inside the question you send it. A model can be talked out of its instructions, and a permission check cannot.

Four of them cover most designs:

  • Input filtering. Reject injected instructions arriving inside documents, tickets, or web pages the agent reads.
  • Permission scope. Give each run a token with the narrowest access that completes the task, not the application credentials.
  • Output filtering. Check generated text for leaked secrets and unsafe content before it reaches a user or another system.
  • Human approval. Require a person to confirm any action that cannot be undone, such as a refund above a threshold, an outbound email, or a deletion.

The sentence to say out loud is this: every action the agent can take is either reversible, approved by a person, or capped by a number. That single rule answers most of the follow-up questions an interviewer has about damage.

Guardrails placed around the agent loop: input filtering, permission scope, output filtering, and human approval before irreversible actions

Layer 7: Observability

Observability is how you learn what the agent did after it ran, and agents need more of it than ordinary services because the same input can produce different behavior twice. Three parts are expected.

A trace per run, recording every step, every tool call with its inputs and outputs, latency, and token count. An evaluation set, which is a fixed collection of tasks with known good outcomes that you run on every change to a model, a tool, or an instruction. And cost measured per completed task rather than per token, because that is the number a business can actually use.

Add one alert that most candidates miss. Watch the share of runs that finish successfully without human help, since a decline in quality appears there long before it appears in error rates.

Why most candidates stop at layer three

The first three layers are familiar territory for anyone who has used a model API, so they are comfortable to talk about and they consume time easily. The result is an answer that describes a demonstration rather than a system, and interviewers recognize the pattern quickly.

Layers four through seven are where the difficult questions live. What happens when a tool returns a wrong result confidently, how a run resumes after a crash, who approves a refund, and how you prove next month that quality has not declined. Reaching those layers is what separates a passing answer from a strong one.

Answer that stops at layer 3Answer that covers all seven
Model routing and instruction designRouting plus a bounded loop and resumable state
Tools mentioned by nameTools with schemas, timeouts, and scopes
"We add safety checks"Named checks, with approval for irreversible actions
No measurementTraces, evaluation set, and cost per task

How to spend a forty-five minute round

MinutesWhat to cover
0 to 5Requirements, one concrete task, success criteria
5 to 10Interface and the latency budget
10 to 20Orchestrator loop, limits, and run state
20 to 30Tools and memory, with one worked example each
30 to 40Guardrails and observability
40 to 45Failure modes, cost, and what you would build first

Start by naming one concrete task the agent must complete, such as resolving a refund request from beginning to end. A single worked example keeps every later layer specific, and it stops the discussion from drifting into general statements about artificial intelligence.

Frequently asked questions

What is the difference between an AI agent and an AI workflow?

In a workflow, an engineer fixed the order of steps in advance and the model fills in the content. In an agent, the model chooses the next step at run time, based on what it has seen so far. The test is simply who decided the order, and when.

Do I need to name specific models or providers?

No, and naming them can date your answer. Describe the routing policy instead, such as a small model for classification and a large model for planning, with a second provider available for failover.

How many tools should an agent have?

Roughly five to ten. Beyond that, tool descriptions start to overlap and the model picks the wrong one more often. If you need more, split the work across several agents, each with its own narrow set.

What should I say about prompt injection?

Treat every document, ticket, and web page the agent reads as untrusted input. Defend with permission scope and approval steps rather than with instructions, because the model cannot reliably ignore text that tells it to misbehave.

Is a multi-agent design always better than a single agent?

No. One capable agent with good tools handles most problems, and every extra agent adds cost, latency, and places where coordination can fail. Add agents for context isolation, parallel work, or genuine specialization, and say which reason applies.

Which layer do interviewers probe hardest?

Guardrails and observability, because they reveal whether you have operated a system like this or only built one. Prepare a concrete example for each: one irreversible action you place behind approval, and one number you watch after release.

AI
System Design Interview

What our users say

Tonya Sims

DesignGurus.io "Grokking the Coding Interview". One of the best resources I’ve found for learning the major patterns behind solving coding problems.

Roger Cruz

The world gets better inch by inch when you help someone else. If you haven't tried Grokking The Coding Interview, check it out, it's a great resource!

Eric

I've completed my first pass of "grokking the System Design Interview" and I can say this was an excellent use of money and time. I've grown as a developer and now know the secrets of how to build these really giant internet systems.

More From Designgurus
Annual Subscription
Get instant access to all current and upcoming courses for one year.

Access to 50+ courses

New content added monthly

Certificate of completion

$31.08

/month

Billed Annually

Recommended Course
Grokking the Object Oriented Design Interview

Grokking the Object Oriented Design Interview

60,674+ students

4.2

Learn how to prepare for object oriented design interviews and practice common object oriented design interview questions. Master low level design interview.

View Course
Join our Newsletter

Get the latest system design articles and interview tips delivered to your inbox.

Read More

Scalability in System Design: The Complete Guide to Techniques, Patterns & Trade-offs

Arslan Ahmad

Arslan Ahmad

How to Prepare for an AI/ML System Design Interview (Complete 2026 Roadmap)

Arslan Ahmad

Arslan Ahmad

Vibe Coding 101: A Beginner’s Guide to AI-Assisted Development

Arslan Ahmad

Arslan Ahmad

Cache Invalidation: Methods, Strategies, and Trade-offs

Arslan Ahmad

Arslan Ahmad

Design Gurus logo
One-Stop Portal For Tech Interviews.
Copyright © 2026 Design Gurus, LLC. All rights reserved.