A cache that’s never invalidated doesn’t know it’s wrong. It read a value once, the value was correct at the moment it was read, and it will keep serving that exact value — confidently, correctly-looking, forever — even after the source of truth moved on without telling it. The context window is a cache: the same staleness that degrades one agent’s working memory over a long run degrades a whole crew’s shared understanding of what it’s even trying to do, except a crew has more places for a stale copy to sit unnoticed.

This closes out the cluster started by failure modes in multi-agent teams: after split-brain — the crew’s facts about current state disagreeing — and goal drift — the target itself moving, one small paraphrase per hop — there’s a third way a crew’s shared understanding breaks, and it isn’t either of those. Call it multi-agent context drift: the crew’s shared beliefs about a fixed objective go stale or become inconsistent across agents, even though the objective itself never moved. The goal-drift post names this distinction without developing it: a crew “can have perfectly fresh, perfectly consistent context about the wrong goal” — that’s goal drift. This is the mirror case: a crew with a perfectly correct, unmoving goal, and imperfectly, unevenly updated beliefs about what that goal currently requires.

A concrete failure

A support-escalation crew handles a billing dispute end to end: a triager reads the incoming ticket and routes it, a specialist pulls the account’s billing history and works out what’s owed, and a responder drafts the customer-facing resolution. The human coordinator hands all three the same policy brief at spawn: “Resolve the dispute. Do not offer refunds — refunds require legal sign-off. Offer account credit instead; credit doesn’t.” Triager, specialist, and responder each get their own copy of that sentence pasted into their own context window, and each starts working against it in parallel.

Twenty minutes in, legal actually clarifies the policy — to the triager, who’s the one still exchanging messages with the human coordinator: “Disputes under $50 can be refunded directly now, no sign-off needed. Leave everything else on the credit-only path.” The triager updates its own understanding and starts routing new tickets accordingly. The specialist and the responder, already spawned and already deep into this ticket with the original brief pinned in their own context, are never told. This particular dispute is $32. The responder, working from its own unrefreshed copy of the policy, drafts and sends: “We’ve applied $32 in account credit to your balance.” The customer was eligible for a cash refund. Nobody lied to them; the responder followed the policy it had, correctly, to the letter.

Nobody’s facts about the world were in conflict — this isn’t split-brain. There’s no shared mutable store with two writers racing on it; there’s one canonical clarification, issued once, that simply never reached two of the three agents holding a copy of the thing it clarified. And the objective didn’t move, hop by hop, into something unrecognizable — this isn’t goal drift either. “Resolve the dispute correctly under current policy” was true at spawn and stayed true throughout; what changed was a fact embedded in that objective, and the change reached one agent out of three. The crew ended the run holding three different, individually reasonable, collectively inconsistent beliefs about one target that itself never budged.

Why it isn’t split-brain, and it isn’t goal drift

It’s tempting to fold this into split-brain — both involve agents holding different values for something that’s supposed to be shared. But the shapes are different in a way that changes the fix. Split-brain needs a store: multiple writers, no lock, no arbiter, and the failure is a race between concurrent writes, resolved badly. Context drift needs no store and no race. The triager, specialist, and responder above never wrote to a shared document at all — each held its own private copy of the policy sentence, pasted in at spawn, and the failure was that a correction to the source of that sentence had no mechanism for reaching copies already handed out. You can fix split-brain with version numbers and optimistic locking on a store. None of that helps here, because there’s no store to version — there are three unsynced photocopies.

It’s also tempting to fold this into goal drift, because both are about a crew’s shared understanding failing to track something correctly over time. But goal drift is a compounding, directional failure that needs a chain: each sequential hop paraphrases the objective slightly, shedding whatever doesn’t fit neatly into the next agent’s task list, and the sum after several hops is a materially different goal than hop zero. Context drift needs no chain. It can happen with agents running fully in parallel and no handoffs between them at all, as above — three agents spawned once, off one brief, no agent ever re-deriving or re-stating what another said. The failure isn’t in how the objective gets reinterpreted along a path; it’s in whether a correction to a still-accurate objective reaches every agent already holding a copy of it. Goal drift is entropy accumulating across hops. Context drift is a broadcast that only reached one subscriber.

What causes it

Four things, and none of them require a bug:

  1. The objective is a value, copied, not a live reference. Each agent gets the policy brief pasted into its own prompt at spawn. From that point on, it’s just text sitting in that agent’s context — indistinguishable, to the agent, from any other fact it was told. Nothing marks it as “subject to change” or gives the agent a way to ask “is this still current.”
  2. Corrections propagate through whichever channel happens to be open, not to everyone holding a copy. Legal told the coordinator, the coordinator told the triager, because the triager was the one mid-conversation. Nothing in that exchange knows the specialist and responder are also operating on the same fact, so nothing routes the correction to them.
  3. Long-running agents don’t re-check their own inputs. An agent mid-task treats what it was told at spawn as settled input, the same way it treats the user’s original request — it has no built-in moment where it asks whether anything it was handed has since changed, because nothing about “keep working the task” prompts that question.
  4. No one compares beliefs across the crew until the output is already wrong. The mismatch between the triager’s updated understanding and the responder’s stale one existed the entire twenty minutes before the message went out, fully inspectable if anyone had asked all three agents to state the current policy and diffed the answers. Nobody did, because nothing in the architecture treats “do the agents currently agree on the objective” as a thing to check.

What to measure

You cannot catch this by checking any single agent’s transcript for a mistake — the responder in the example above made none. You catch it by instrumenting belief consistency, not correctness of any one agent’s local reasoning:

SignalWhat it meansHow to catch it
Belief divergence across agentsTwo agents would state the current objective’s constraints differently if asked right nowAt a checkpoint, prompt every active agent to restate its understanding of the current policy/objective in one sentence; diff the answers instead of assuming they match
Refresh age vs. correction ageAn agent is running on a copy of the objective older than the last correction issued to itTimestamp every agent’s last objective-refresh; timestamp every correction; flag any agent whose refresh predates a correction that should have reached it
Correction fan-outA clarification reached the one agent in the loop, not every agent holding a copy of what it clarifiedTrack, for every issued correction, what fraction of agents holding a copy of the relevant brief actually received it — a fan-out below 100% is the failure, not a rounding error
Output-vs-current-policy mismatchFinal output satisfies the original brief but violates something corrected mid-taskRe-check delivered output against the current canonical objective statement at send time, not the objective each agent believed when it started working

What actually fixes it

  • Don’t pin the objective as immutable — pin it as invalidatable. The cache-management discipline of admit, evict, summarize also needs a fourth verb here: invalidate on external change, so a pinned block can be told it’s gone stale instead of sitting there looking permanently authoritative.
  • Push corrections to every holder, not just whoever’s listening. If three agents were handed a copy of the policy, a change to that policy is a message to all three, not a reply to whichever one happened to be in the conversation when it changed.
  • Give long-running agents a refresh trigger, not just a start-of-task read. The same shape as a bounded retry budget: at defined checkpoints, an agent re-pulls its objective from the source instead of trusting the copy it’s been running on, and a stale copy past some age becomes a visible event, not a silent one.
  • Check final output against the current objective, not against what any agent believed. The same discipline goal drift argues for at the target level applies at the belief level: one check, external to the crew, that compares what’s about to ship against the live canonical statement of the objective — not the responder’s transcript, not the triager’s, either of which will look locally correct.

The summary

Context drift is the quiet member of this cluster because nothing about it looks like an error while it’s happening. Split-brain at least has two writers actively fighting over a store; goal drift at least has a chain of hops you can walk back and diff against hop zero. Context drift has neither — just several agents, each entirely correct relative to the copy of the objective it was handed, quietly diverging from each other because nobody ever came back to update the copies once handed out. Failure modes in multi-agent teams called the closest relative of this “context fragmentation” — a single lossy handoff that splits a goal so no one agent holds the whole of it. This is the opposite defect: every agent has the whole objective, complete and correct, at the moment it’s spawned. The failure happens after, when the world underneath that objective changes and the update stops at whichever agent happened to be listening. If more than one agent is running for more than a few minutes off the same brief, assume their copies are already diverging from the source, and go check — not by re-reading any one agent’s reasoning, which will look fine, but by asking every agent still running to state the objective, right now, and seeing whether they agree.