Designing AI agent systems
The architecture behind LLM agents that use tools: the agent loop, tool definitions and execution, sandboxing, memory and context management, orchestration and multi-agent patterns, durable long-running tasks, guardrails and permissions, evaluation, cost control and observability.
Reading is half of it. See this used in a real interview: walk through Design an LLM Inference Service →
An AI agent is a language model in a loop: it reads a goal, decides on an action (call a tool, run code, search, ask the user), observes the result, and repeats until done. Coding assistants, research agents, customer support bots and automation tools all follow this pattern. Designing them raises classic system design questions (state, reliability, security, cost) in a new shape, and interviewers increasingly ask about it.
The agent loop
- Build the context: system instructions, the task, available tools, relevant memory and the conversation so far.
- Call the model; it returns either a final answer or one or more tool calls with arguments.
- Validate and execute the tool calls.
- Append results to the context.
- Repeat until the model produces an answer, a step limit is hit, or a human is needed.
Each iteration is a model call (seconds, and tokens that cost money), so a task may take minutes and dozens of calls.
Tools
Tools are functions the model can call: search, read a file, query a database, run code, send an email, call an internal API.
- Define each with a clear name, description and strict input schema; validate arguments before executing.
- Return concise, structured results; huge outputs waste context and money.
- Make tools idempotent where possible and pass idempotency keys for side effects, since the agent (or the system) may retry. See idempotency keys in practice.
- Standard protocols for exposing tools (such as MCP) let many agents reuse the same tool servers.
Sandboxing
Agents that run code or shell commands must do so in isolated environments: microVMs or hardened containers with CPU, memory and time limits, restricted network egress and no ambient credentials. Treat model output as untrusted code. See running untrusted code safely.
Permissions and guardrails
- Least privilege: the agent acts with the user's permissions (or less), never with a broad service account. See authorization and permissions.
- Human approval for risky or irreversible actions (payments, deleting data, sending messages to others).
- Prompt injection: content the agent reads (web pages, emails, documents) may contain instructions. Treat tool outputs as data, separate them clearly in the context, and restrict what actions can follow from untrusted input.
- Input and output filters for safety and data leakage. See LLM systems.
- Audit logs of every action taken. See security and auth.
Memory and context
Context windows are large but finite and expensive:
- Short-term: the current conversation and tool results, trimmed or summarised as it grows.
- Long-term: facts, preferences and past results stored externally and retrieved when relevant, often with vector search. See vector search and RAG.
- Working files: scratch files or notes the agent writes and reads, rather than holding everything in context.
- Prefix caching makes repeated system prompts and tool definitions cheap. See serving LLMs on GPUs.
Orchestration patterns
- Single agent with tools: simplest; good for most tasks.
- Planner and executors: one model plans steps, others carry them out.
- Multi-agent: specialised agents (researcher, coder, reviewer) coordinated by an orchestrator, often working in parallel on subtasks.
- Workflows with model steps: fixed control flow where some steps call a model. More predictable than open-ended agents; prefer it when the process is known.
More agents mean more cost, latency and failure modes; add them only when tasks genuinely parallelise or benefit from separation.
Long-running and durable tasks
Agent tasks can run for minutes or hours. Run them as durable workflows so a crash, deploy or rate limit does not lose progress: each model call and tool call is a recorded step, retried with backoff, resumable from history. Users get status updates and can cancel. See workflow orchestration and background jobs.
Reliability
- Step and budget limits: maximum iterations, tokens and wall time per task.
- Model API failures: retries with backoff, fallbacks to other models or providers, and handling rate limits.
- Loop detection: stop when the agent repeats the same failing action.
- Structured outputs validated against schemas, with a repair retry on invalid output.
Evaluation
Agents are hard to test because behaviour varies between runs:
- Task suites with known outcomes, run on every prompt, tool or model change, scoring success rate, cost and steps.
- Trajectory review: inspect the sequence of actions, not just the final answer.
- Production monitoring: success rate, user corrections, escalations, cost per task.
See testing distributed systems for the general mindset.
Observability and cost
Trace each task as a tree: model calls (with tokens, latency, model version) and tool calls (with arguments, results, errors). Attribute cost per task and per user, cap spending, and route simple steps to cheaper models. See distributed tracing and cost-aware system design.
In the interview
"Requests start a durable agent task; each iteration calls the model through a gateway with token-based limits; tool calls are validated against schemas and executed in sandboxes with the user's permissions; risky actions need approval; context is trimmed with summaries and long-term memory in a vector store; every step is traced with cost; tasks have iteration and budget caps; an evaluation suite gates changes." Pair it with Design LLM Inference for the serving layer and Design LeetCode for sandboxed execution.
Checklist
- Agent loop with clear termination and budgets.
- Strictly typed, idempotent tools with concise outputs.
- Sandboxed execution; least privilege; human approval for risky actions.
- Prompt injection treated as a real threat.
- Context management with summaries, retrieval and prefix caching.
- Durable, resumable tasks with retries and fallbacks.
- Evaluation suites, traces and per-task cost tracking.