A single agent that drifts keeps working: loop drift is the failure mode where an agent stays busy — narrating progress, burning tokens, taking actions — without getting any closer to done. A multi-agent crew has a stranger version of the same failure, and it’s arguably worse, because nothing about it looks stuck. Every agent in the crew is making real, forward, individually defensible progress. The crew keeps shipping outputs. And the whole operation is still moving further from the goal you actually gave it with every step. Call it multi-agent goal drift: not a crew that stalls, but one that steadily walks off the original objective while every individual handoff looks like solid work.
This is the second companion post to failure modes in multi-agent teams, after split-brain. That post’s third failure mode, context fragmentation, sits close to this one and is worth separating cleanly. Context fragmentation is a single lossy event — the goal gets split across agents at one handoff, and a load-bearing detail falls into the gap between two context windows during that one split. Goal drift isn’t a single event, it’s compounding. It shows up in crews that run several sequential handoffs — a planner to a builder to a reviewer to a closer, or the same handoff shape looped across iterations — where each individual hop reshapes the goal by an amount too small to flag on its own, and only the sum, after enough hops, is unrecognizable. Split-brain, for comparison, is agents disagreeing on facts — two copies of a store diverging on what the current state is. Goal drift is agents agreeing perfectly on the facts and quietly disagreeing, without anyone noticing, on what they’re for.
It’s also adjacent to multi-agent context drift — the information a crew holds about a fixed objective going stale or inconsistent across agents — but that’s a claim about the crew’s beliefs. This is a claim about the crew’s target. A crew can have perfectly fresh, perfectly consistent context about the wrong goal, which is exactly what happens below.
A concrete failure
Three agents build a feature end to end: a planner turns a user request into a spec, a builder implements it, and a closer reviews the diff and merges. The request: “Add a cache in front of the search endpoint to cut p99 latency — do not change the public API contract.” The planner writes a reasonable spec: add a cache layer, wire it into the search handler, add invalidation on writes, expose a metric for hit rate. Each subtask inherits the tasks from that decomposition. It does not inherit the constraint. “Do not change the API contract” never makes it past the planner’s own reasoning, because it isn’t a task — it’s a boundary, and boundaries don’t fit cleanly into a task list.
The builder implements the cache. To make it debuggable, it adds an optional ?cache=bypass query parameter — a small, sensible, entirely well-reasoned addition, and a genuine change to the public contract. The builder’s own acceptance criteria — tests pass, cache works, hit rate is measurable — never mention the contract, because that criterion lived only in the original request, two hops upstream, and nothing carried it forward. The closer reviews the diff against the builder’s stated intent — “add a debug bypass parameter” — and approves it, because relative to that spec, the change is correct. Ship it. The contract broke. Every agent along the way did defensible work. The diff that actually broke the promise passed review, because review was checking the wrong reference point: the nearest upstream spec, not the original ask.
Why it isn’t just a game of telephone
The telephone-game framing undersells it, because telephone-game distortion is noise — random, undirected, roughly symmetric. Goal drift in an agent crew is closer to genetic drift than to noise: it’s directional, and it accumulates, because at every hop the crew has a structural reason to shed anything that isn’t a concrete, checkable task. A boundary condition like “don’t change the contract” has no natural home in a task list, a diff, or a pass/fail test — it’s exactly the kind of constraint each individual handoff is optimized to lose, not by accident but because it’s the informationally cheapest thing to drop. Multiply that by several sequential hops and the loss compounds in one direction: away from the constraints, toward whatever’s easiest to state as a concrete next action. No single hop looks unreasonable. The sum is a different problem than the one you asked for.
What causes it
Four things, and they compound:
- Local objective functions replace the global one. Every agent downstream of the planner grades its own output against its own subtask’s acceptance criteria, not the original request, because that’s the only spec it was ever handed.
- Constraints don’t have a task-shaped home. Concrete asks (“add a cache”) propagate cleanly through a task list. Boundaries (“don’t change the contract,” “keep it under 200ms,” “don’t touch the audit log”) don’t, and get silently paraphrased into something task-shaped, or dropped, because nothing in the pipeline’s structure requires them to survive verbatim.
- Nobody owns “does this still match the original ask.” Like diffused responsibility for actions, each agent assumes fidelity to the root goal is someone else’s job — usually whoever’s “closer” to the user, which in practice means nobody, because every agent in the middle is closer to its own predecessor than to the user.
- It compounds across hops, not within one. A single lossy handoff is context fragmentation; a crew that runs several handoffs — or loops the same handoff shape across iterations — accumulates many individually small, individually defensible substitutions in the same direction, and the result only looks wrong compared against hop zero, which is exactly the comparison nobody runs.
What to measure
You cannot catch this by checking whether each handoff’s own acceptance criteria passed — every hop above passes its own bar. You catch it by instrumenting the distance from the origin, not the local pass/fail:
| Signal | What it means | How to catch it |
|---|---|---|
| Constraint survival | The original ask’s non-negotiables are still stated, verbatim, in the current hop’s spec | Diff each downstream spec’s literal constraint list against the root request’s; a constraint present at hop zero and absent at hop N is drift, not a rounding error |
| Distance from origin | The working objective has moved away from the original ask, not just been refined | Embed the current spec and the original ask; track the distance across hops and flag a monotonic upward slope, not just one reading |
| Acceptance criteria without provenance | A subtask’s “definition of done” doesn’t cite which root constraint it protects | Require every subtask spec to tag which top-level constraints it’s responsible for; an untagged constraint is one nobody’s watching |
| Review against the wrong reference | The final check compares output to the nearest upstream spec instead of the original request | Run one check — cheap, separate from the crew — that diffs delivered output directly against the root ask, never against a downstream derivation of it |
What actually fixes it
- Carry constraints verbatim, not paraphrased, at every hop. A boundary condition should travel through a crew the way a retry budget travels through a loop: restated at each step, not re-derived from whatever the last agent said it was.
- Give one check the sole job of comparing final output to the original ask. Not the planner’s spec, not the builder’s diff description — the literal request the human or upstream system made, re-read fresh at the end, graded the same way any drift detector should be: on an external fact, never on an agent’s summary of its own fidelity.
- Make every subtask’s acceptance criteria trace to a named root constraint. An untraceable “definition of done” is exactly the gap goal drift lives in. Require the citation, and treat a subtask with no upstream constraint attached as one nobody’s actually checking against the ask.
- Bound the chain. The same shape as capping livelock or a retry budget: after some number of hops, force a re-grounding step that re-reads the literal original request before letting the crew continue, instead of letting the goal drift for as long as the crew keeps producing outputs.
The summary
Goal drift doesn’t look like failure while it’s happening, which is what makes it worse than loop drift: a stalled agent at least stops producing convincing output, while a drifting crew never stops — it just keeps shipping competent work that’s steadily less connected to what you asked for. It isn’t split-brain: the crew can agree completely on every fact and still drift, because the thing diverging isn’t the state, it’s the target. And it isn’t a single context-fragmentation event from failure modes in multi-agent teams — it’s what fragmentation looks like accumulated over many hops instead of one. If your crew runs more than a hop or two deep, assume the constraints from your original ask are decaying at every one of them, and go check — not by re-reading any single agent’s transcript, which will look reasonable, but by diffing what’s about to ship against the literal request you started with.