Your LLM-as-judge is lying to you

A model grading your agent’s output is the only thing that scales for subjective quality — and it’s a biased instrument you’re reading as a ruler. Here are the biases that actually move scores (position, verbosity, self-preference, leniency), why raw agreement with a human hides them, and how to validate and harden a judge with code — including why you should be reporting Cohen’s kappa, not accuracy.

July 8, 2026 · 10 min · 1958 words · Loop & Retry

What to actually measure when your agent "works"

“It works” is a demo result, not a measurement. An agent is a trajectory, not a function, and grading only the final answer throws away most of what decides whether it’s reliable. Here’s a five-layer scheme for what to measure — outcome, trajectory, cost, failure class, and stability under nondeterminism — with a small harness that computes it.

July 7, 2026 · 11 min · 2158 words · Loop & Retry

The context window is a cache, not a memory

Treating the context window as append-only memory is how agents get slow, expensive, and quietly wrong. The fix is to run it like a cache with a budget and an eviction policy: decide what earns its tokens every turn. Here’s the cost math, the accuracy failure mode, and a working context manager.

July 7, 2026 · 9 min · 1826 words · Loop & Retry

Loop drift: how agents convince themselves they're making progress

The worst agent failures don’t crash — they keep working. A postmortem on loop drift: an agent that stayed busy for 40 steps without getting closer to done, why the model’s own progress reports can’t catch it, and the external signals and evals that can.

July 7, 2026 · 9 min · 1737 words · Loop & Retry

Designing tools an LLM won't misuse

A tool schema is a contract with a caller that guesses. This is a concrete walkthrough of the four properties that separate a tool a model uses correctly from one it fumbles: legible schemas, validating boundaries, recoverable errors, and idempotency — with before-and-after code.

July 6, 2026 · 9 min · 1759 words · Loop & Retry

Retry budgets: why 20% per-step failure doubles your token bill

Retries feel cheap and local. In a multi-step agent they’re neither. A small cost model shows why 20% per-step failure can more than double your bill — and how your recovery architecture, not your failure rate, decides the multiplier.

July 5, 2026 · 8 min · 1651 words · Loop & Retry