← Back to concepts
7 min read

Loop Engineering

Loop engineering is the practice of designing the repeated cycle an agent runs through — think, act, observe, repeat — so that it makes real progress, stops at the right time, and fails safely when it does not. The loop is the heartbeat of most agentic systems, though some use fixed workflows or graphs instead. Small design mistakes in a loop can cause big problems.

At its simplest, an agent loop looks like this:

while task is not done:
    decide next action
    take the action
    observe the result

That single while line hides most of the hard engineering work. Loop engineering is about answering: how does the agent decide it is “done”? What happens if it never is? What happens if it repeats the same failed step forever?

The pseudocode is simplified. Production loops also stop or pause on cancellation, tool errors, human approval waits, deadlines, and budget limits.

Why loops are risky by default

A loop with no limits will run until something forces it to stop. In an agentic system, that “something” needs to be designed deliberately, or a few things can go wrong:

  • Infinite loops: the model keeps trying the same failing action, or oscillates between two actions without progress.
  • Runaway cost: every iteration calls a model and possibly external tools, so an unbounded loop can become expensive quickly.
  • Silent drift: the agent keeps “making progress” in a direction that is no longer useful, without anyone noticing.
  • Compounding errors: a small mistake early in the loop feeds into every later step, growing worse each time.

Loop engineering exists to catch these problems before they cause damage.

Termination conditions

A well-engineered loop needs more than one way to stop:

  • Externally verified success: the goal is met — for example, all tests pass, a record exists, or the user confirms the answer is correct.
  • Model-reported completion: the model says it is done. This is useful, but it should be checked against external evidence when possible.
  • Step limit: a maximum number of iterations, so the loop cannot run forever even if it never technically succeeds.
  • Time or cost budget: a ceiling on how long or how much the loop is allowed to spend.
  • No-progress detection: if the last few steps produced no meaningful change, the loop should stop and report rather than keep repeating.
  • Stuck or failure signal: the model can report that it is stuck, and the harness should also detect stuck behavior independently.

Relying on just one of these (usually “the model says it’s done”) is fragile. Combining several makes the loop far more dependable, and those checks are usually enforced by the surrounding harness.

Multiple stop gates in an engineered agent loop After each think, act, observe iteration, the harness checks several stop gates: externally verified success, step limit, time or cost budget, no-progress detection, and a stuck signal. If any gate fires, the loop stops with success or a safe report. If none fire, the loop continues. The model saying done or stuck is only one signal, not the only stop condition. One iteration decide next action -> act -> observe STOP GATES CHECKED AFTER EACH ITERATION Verified goal met Step limit max turns Budget time or cost No progress stalled signal Stuck signal model or harness any gate fires Stop: return result or safe report none fire: continue
A dependable loop has several independent ways to stop, not just the model's own claim that it is finished.

Detecting progress, not just activity

A loop can look busy while going nowhere — calling tools, producing output, taking actions — without moving closer to the goal. Good loop engineering distinguishes activity from progress.

One approach is to track a measurable signal tied to the actual goal. For a coding agent fixing a bug, that might be “number of failing tests.” For a research agent, it might be “number of open questions answered.” If the signal is not improving after a few iterations, the loop should reconsider its approach instead of repeating it.

Progress should mean a measurable improvement: failing tests drop, a new fact is found, a blocker is removed, a file changes in the intended area, or a user-approved checkpoint is reached. False progress is activity without improvement: repeated identical tool calls, rewriting the same paragraph, searching the same logs, or producing longer plans without new evidence.

Activity versus progress in a loop Two panels compare loop histories. The real progress panel shows failing tests dropping from six to three to one to zero across four iterations, ending in success. The busy but stuck panel shows four iterations that all search the same logs and keep the open question unchanged, so the no-progress detector stops and reports. The important signal is movement toward the goal, not the number of actions. Real progress goal signal improves each iteration ITERATION SIGNAL STATUS 1 6 failing tests baseline 2 3 failing tests improving 3 1 failing test improving 4 0 failing tests stop: success Busy but stuck actions continue, but the goal signal does not move ITERATION ACTIVITY STATUS 1 search incident logs cause open 2 search same logs cause open 3 search same logs cause open 4 same logs stop: no progress A loop should measure whether the goal is closer, not whether the agent stayed busy.
Loop engineering turns progress into something the system can check instead of trusting activity as a proxy.

Feeding the right context back in

Each turn of the loop, the harness has to decide what the model sees next — the original goal, what has happened so far, and the latest result. Including too little makes the model repeat past mistakes. Including too much (like the full history of every step) can overwhelm the context window and slow the model down. This connects closely to context engineering: loop design and context design are solved together, not separately.

Parallel steps add another risk: stale state. If two branches read the same state and both act on it, one branch may make the other’s plan invalid. A loop that runs parallel work should merge observations carefully, re-check assumptions, and avoid write actions based on stale data.

Failure handling inside the loop

Real loops need more than “try again.” Common controls include:

  • Cancellation: stop promptly when the user cancels or the parent workflow ends.
  • Human waits: pause with saved state when approval or extra input is needed, then resume from the same checkpoint.
  • Tool timeouts: stop waiting on a slow tool and return a clear error to the loop.
  • Retries with backoff: retry transient failures after a delay, but only when the action is safe to repeat.
  • Circuit breakers: stop calling a failing tool after N failures and choose another path or report the blocker.

These controls keep a temporary service failure from becoming an expensive runaway loop.

An example: a research loop

Goal: Answer "What caused the outage last Tuesday?"

Iteration 1: Search incident logs. -> Found a timestamp, no root cause yet.
Iteration 2: Search deploys around that timestamp. -> Found a matching deploy.
Iteration 3: Read the deploy's change log. -> Found the likely change.
Iteration 4: Confirm by checking error rates before/after the change.
Iteration 5: Summarize findings and stop - goal met.

Each iteration should narrow the search or add new evidence. If iteration 6, 7, and 8 kept searching the same logs without new findings, a well-engineered loop would recognize the stall and stop rather than continue indefinitely, or hand the task to a human reviewer.

Loop metrics

Useful loop metrics include:

  • Success rate: how often the loop reaches an externally verified goal.
  • Premature-stop rate: how often it stops before enough work was done.
  • Runaway-loop rate: how often it hits a hard limit instead of a natural stop.
  • Average steps and cost per task: how much work successful runs take.
  • Tool failure rate: how often retries, timeouts, or circuit breakers trigger.

These metrics make loop changes testable. A new prompt, tool, or stop rule should improve at least one metric without making another important one worse.

The key idea

Loop engineering is about making the repeated think-act-observe cycle of an agent reliable: knowing when to stop, recognizing real progress versus busywork, and failing safely instead of running forever. A good loop is what keeps an agentic system from spinning out of control.