← Back to concepts
8 min read

Agents

An AI agent is a system that uses a large language model to choose actions and make progress toward a goal. The key difference is not just chat versus multiple turns. It is goal-directed control: the model proposes the next action, and the host application controls which actions are allowed, executed, logged, and stopped.

For example, a travel assistant might search flights, compare prices, check calendar dates, ask for confirmation, and then book the selected option. The model is not just writing text. It is helping control a workflow.

What makes something an agent?

There is no single strict definition, but most agents have four parts:

  1. Goal: what the agent is trying to accomplish.
  2. Model: the reasoning engine that proposes what to do next during inference.
  3. Tools: functions or APIs the agent can call.
  4. State: task information about what has happened so far.

An agent loop is the repeatable cycle that lets the system decide, act, observe, and continue. The loop engineering article covers the policy details, but the basic loop often looks like this:

  1. Read the goal and current state.
  2. Decide the next action.
  3. Use a tool or ask the user a question.
  4. Observe the result.
  5. Repeat until the task is done or blocked.

This loop is powerful, but it also creates risk. If the agent chooses poor actions, it can waste time, spend money, or make unwanted changes.

The agent loop with application control A user goal enters the host application, which keeps task state and permissions. The model reads the goal and current state, then proposes a next step. The application either asks the user a question or validates and executes a tool call. The observation is written back to state. A stop check ends the task when it is done, blocked, waiting for approval, or over a limit; otherwise the loop repeats. User goal or request Application / agent host Model decides next step from goal + state Next action tool request or user question Tools / user act or clarify Observation tool result, user reply, error, or confirmation State completed steps, facts, errors, approvals repeat with updated state Stop when done, blocked, approved, or over limit stop check
An agent is not just a model; it is a host-controlled loop that updates state, uses tools, and stops at defined boundaries.

Tools

Tool calling lets an agent interact with the world. A tool might search documents, read a database, send an email, create a ticket, run a test, or update a file.

The model does not magically know how to use a private system. The application must expose safe tools with clear names, descriptions, inputs, and permissions.

Good tools are narrow and predictable. For example, search_docs(query) is safer than run_any_command(command). Narrow tools reduce the chance that the agent does something dangerous or confusing.

How safe tools sit between the model and the world The model proposes a tool call with a name and JSON arguments. The application boundary checks the tool name, argument schema, permissions, and whether human approval is needed. Safe narrow tools such as search docs, read order, and create draft ticket can run and return observations. A broad run any command tool is shown as blocked because it gives too much uncontrolled power. Model proposes call name + JSON args Application boundary schema validation permission check approval if risky Narrow tools that can run search_docs read_order create_draft_ticket observation returns to state run_any_command blocked: too broad
The model proposes actions, but the application decides which tools exist, validates every call, and blocks capabilities that are too broad.

Planning

Agents often need a plan. A plan breaks a goal into smaller steps. For example:

Goal: Prepare a release note.

Plan:
1. Read merged pull requests.
2. Group changes by feature, fix, and internal work.
3. Draft a release note.
4. Ask a human to review before publishing.

Planning helps users understand what the agent intends to do. It also gives the system a chance to add checkpoints before important actions.

Plans are not promises. They can go stale when a tool result, user reply, or error changes the situation. Good agents re-plan when observations change. Independent steps can run in parallel, but only when they do not depend on the same changing state.

Memory and state

Agents need task state so they can keep track of current work. State can include completed steps, tool results, discovered facts, errors, and pending confirmations. State helps the agent avoid repeat work, but it is not enough by itself. The harness still needs progress checks, step limits, and stop rules.

Memory is longer-lived information that may be reused across tasks, such as a user’s preferred coding style or team naming convention. Memory should be designed carefully. This is part of context engineering: deciding what the model should see now, what belongs only to the current task, and what should be stored longer. Some information may be sensitive and should not be stored unless the user clearly expects it.

The model may use a ReAct-style pattern internally or visibly: think about the next step, act through a tool, read the observation, and continue. Modern systems often implement that pattern with structured tool calls rather than plain text action labels.

When agents are useful

Agents are useful when a task is:

  • Multi-step.
  • Dependent on intermediate results.
  • Connected to tools or external systems.
  • Hard to express as one fixed workflow.
  • Valuable enough to justify extra complexity.

Examples include code maintenance, research, customer support investigation, data cleanup, report generation, and operations workflows.

When not to use an agent

Agents are not always the best answer. If the workflow is simple and predictable, a normal deterministic flow may be better.

For example, if a user clicks “reset password”, you do not need an agent to decide what to do. A normal backend flow is safer, faster, and easier to test.

Use agents when flexibility is needed. Use normal software when the steps are already known.

Choosing between a fixed workflow and an agent Two side-by-side panels compare a predictable password reset flow with a flexible support investigation. The fixed workflow follows known steps: user clicks reset, backend verifies identity, sends a reset link, and ends. The agent workflow starts with an open-ended issue, searches logs, reads docs, asks a clarifying question, updates state, and repeats until enough evidence is found. The contrast shows that agents are useful when the next step depends on intermediate results. Use normal software steps are known before the request starts click reset verify ID send link safer, faster, easier to test Use an agent next step depends on what was found investigate issue search logs read docs ask user repeat with evidence better for flexible, multi-step work
Choose an agent when flexibility matters; choose deterministic software when the correct path is already known.

Safety and control

A production agent should have boundaries:

  • Clear tool permissions.
  • Human approval for risky actions.
  • Logs of what it did and why.
  • Limits on retries, spending, and time.
  • Tests for common failure cases.
  • A way to stop or recover from mistakes.

The model should not be trusted with unlimited authority. Tool outputs, repo files, web pages, and emails can contain prompt injection, so the system should treat observations as untrusted input and follow prompt engineering defenses. Give the agent least-privilege tools, sandbox code execution, restrict file and network access where possible, and keep secrets out of the model context.

Evaluating agents

Agent quality is not just “did the final answer sound good?” Useful metrics include:

  • Task success rate: did the requested outcome actually happen?
  • Steps, cost, and latency per task: how much work did success require?
  • Tool error rate: how often calls fail, time out, or return unusable results.
  • Unsafe or unauthorized action attempts: how often the agent tries something outside policy.
  • Human intervention rate: how often people must correct, approve, or rescue the run.

Teams also replay saved traces when changing prompts, models, or tools. A trace records model calls, tool calls, observations, and decisions. Replaying those traces is part of harness engineering: it helps catch regressions before an updated agent reaches users.

Agents and agentic AI

An agent is a concrete system with a goal, loop, tools, and state. Agentic AI is the broader idea of giving AI systems more autonomy. A system can be highly agentic in one way and limited in another: it may plan several steps but only use read-only tools, or it may have write access but require approval every time.

The key idea

An agent is a model-driven workflow that can choose actions and use tools to complete a goal. Agents are useful for flexible, multi-step work, but they need strong boundaries. The best agents combine model reasoning with safe tools, clear state, and human control where it matters.