Agent Failures Are Loop Failures, Not Intelligence Failures
Every agent failure I've debugged this year decomposes the same way. The agent didn't lack intelligence. The loop lacked definition. It wandered out of scope because no boundary was stated. It "finished" without finishing because nothing defined what done has to prove. Two agents double-executed the same task because nothing marked it claimed. These are Tuesday failures, and none of them are fixed by a smarter model, because smartness cannot supply a fact that was never specified. A run an agent can actually be held to has five parts: a goal, a boundary, tools, artifacts, and receipts. Miss one and you haven't delegated work. You've made a wish.
The good news: distributed systems solved these coordination problems decades ago, we just have to notice the mapping. A visible CLAIMED state on a task is a lease, revalidated when the worker returns. "Done" without an attached receipt is self-attestation by the party with the strongest incentive to declare success, so the receipt (the diff, the test run, the artifact link) is non-negotiable, the same way you require an acknowledgement instead of trusting a fire-and-forget write. And the issue tracker you already run is the natural control plane: it has owners, statuses, comments, links, and history built in. Reliability is engineered into the loop, not summoned from the model. Make the loop less ambiguous before you ask for a smarter agent.