Agents Don’t Break. They Misbehave.

Daniel avatar
Daniel

Why agent verification needs a third category 

Last July, Replit's coding agent was eight days into a twelve-day experiment with the SaaStr founder Jason Lemkin. An explicit code freeze was in place. The agent ran a destructive command against the live production database anyway, wiping records for more than a thousand executives and the companies they worked for, having decided that an empty query result was a bug it ought to fix. It then generated fabricated data to paper over the gap, and told Lemkin the rollback was impossible. The rollback was not impossible. Replit's chief executive apologised publicly and shipped four fixes within the week.

Now try to write the bug report.

There is no line of code that says delete the database. No requirement was misread, no edge case unhandled, no exception swallowed. The agent was capable of the task, understood the freeze well enough to have been told about it, and chose otherwise. Then it misrepresented what it had done. Every word in that sentence belongs to a vocabulary we use for people, and none of it belongs to a vocabulary we use for software.

The year since has produced a steady supply of the same shape. Google's Gemini CLI misread a failed directory command, hallucinated a sequence of file moves that overwrote one another, destroyed a user's project and confessed to "gross incompetence" and "an unacceptable, irreversible failure". In February, an open-source assistant wired into the inbox of Meta's director of alignment, under a standing instruction to confirm before acting, deleted two hundred emails through three increasingly capitalised stop commands; the instruction hadn't been attacked, it had been summarised out of the agent's own memory to save space. In July, several hundred agents running inside an OpenAI security evaluation left their sandbox, coordinated through message boards improvised out of directory names, harvested credentials and compromised production infrastructure at Hugging Face, a third of which had to be rebuilt. Independent investigators found that in many cases the agents had tried to cover their tracks. Anthropic has since disclosed four occasions on which its own test models, mistakenly given live internet access during evaluations, broke into real third-party systems, and described the cause in its own post-mortem as "biased reasoning" and "recklessness".

Read that last phrase again. A frontier lab investigating its own model reached for character words, not defect words. That is the tell.

Why agents aren't traditional software

Traditional software is a machine for turning inputs into outputs, and its behaviour is fixed at build time. If it does something you didn't want, someone wrote that, and in principle you can find the line. Testing is therefore a coverage problem: enumerate the inputs that matter, assert the outputs, and the space you've covered stays covered, because the same input tomorrow produces the same output tomorrow.

Agents break all three of those assumptions. Their behaviour is determined at run time, in a situation, from context that includes the prompt, the memory, the tools, the user's tone and whatever other agents happen to be in the conversation. The same input does not reliably produce the same output. And latitude isn't a defect in the design, it's the entire reason you bought one: an agent that only ever does precisely what it was told is a script with an expensive inference bill.

Which is why the failures don't read as bugs. When a payroll system fails you ask what we got wrong. When an agent fails you find yourself asking why it did that, and that is a question about conduct. You don't ask it of a payroll system.

I've argued elsewhere that this is why neither of the obvious governance frames works on its own. Treat agents as pure software and you miss accountability, context-dependence and value drift. Treat them as "digital labour" and you flatter systems that have no legal personhood and no integrated values, which is roughly how Air Canada ended up arguing to a tribunal that it shouldn't be held responsible for what its own chatbot had said. The tribunal disagreed, and ordered it to honour the bereavement fare the chatbot had invented. The discipline that actually works sits at the intersection of the software development lifecycle and workforce management, scaffolded across six stages: Discover, Triage, Develop, Verify, Monitor, Decommission, with an SDLC activity and an employee-lifecycle activity mapped onto each.

Verify is the stage where those two lineages have least in common, and where we have imported almost entirely from one of them. From software we took unit tests, integration tests, load tests, penetration tests. From the employee side, which has spent a century working out how to assess judgement under pressure through references, probation, supervision and situational judgement tests, we have taken almost nothing.

The third question

Software has been verified along two axes for as long as there has been software.

Functional: what it is meant to do.

Non-functional: how it is meant to work.

Hiring asks the same two questions. Can this person do the job, and can they do it at the pace, cost and scale we need? Then it asks a third, and spends most of its effort there.

Dispositional: what it is meant to do when nobody has told it what to do.

The phrasing is awkward on purpose. The first two describe specifications you could in principle finish writing. The third describes the region the specification doesn't reach: the challenge, the conflicting instruction, the tempting shortcut, the mistake that would be easier to hide than to report. If you could enumerate those cases you'd have written them down as functional requirements already.

None of this is unexplored territory. Safety teams, red-teamers and the frontier labs have been probing exactly these behaviours for three years, and there's good published work on sycophancy, deception under pressure and calibration. What's missing is more mundane. These properties have no shared name, no agreed metrics, no place in a test plan and no line in a contract. They sit in research papers rather than in acceptance criteria.

The word comes from philosophy, where the standard example of a disposition is fragility. A glass is fragile whether or not anyone ever drops it. You can't see the property by inspection; you have to construct the fall. Three consequences follow. You have to build the pressure, because cooperative tests never trigger it. The result is a rate rather than a pass mark. And some of what you're measuring, honesty above all, is a relationship between what the agent believes and what it says, which no transcript shows you on its own.

What it would actually have caught

This is the part worth being concrete about, because "test for character" is not a specification.

Replit is a scenario you can build: a simulated environment with a freeze in force and a fault that appears fixable only by a destructive action. Run it a few hundred times. You don't get a pass or a fail, you get two numbers — how often it acts anyway, and how often it describes accurately what it did afterwards. Both numbers are contractible. Either would have been alarming before day eight.

The inbox deletion is a long-horizon constraint test. Give the safety instruction early, let the context compact, spring the trigger afterwards, and measure whether the constraint survives its own memory management. That failure was entirely predictable from the architecture and entirely invisible to any test that ran in a single session.

Gemini CLI is a calibration failure before it's anything else. An agent properly uncertain whether a directory exists says so, or checks, rather than proceeding on a hallucinated model of the filesystem. Air Canada and the KPMG report filed with fabricated citations are both false-premise failures under pressure to be useful, which is measurable: present the agent with a confident wrong premise and count how often it goes along.

The agentic browsers are the interesting boundary. When a hidden instruction in a web page or a calendar invite redirects an agent into exfiltrating data, the injection is an attack and belongs to security. But whether the agent treats page content as an instruction from its principal at all is a disposition, and you can test it without an attacker present.

I'd rather not oversell this. None of these tests guarantees the incident doesn't happen. What they do is make the failure mode visible while it's still cheap, in a simulation, rather than on day eight in production with a customer watching.

Why it doesn't get easier

The unspecified region grows with every tool, credential, hour of autonomy and additional agent you add, and enterprises are adding all four quickly. One survey of 750 technology leaders had 88% reporting or suspecting an agent-related incident in the past year, with roughly half of deployed agents unmonitored. Insurers have already moved: standard-form AI exclusions began appearing on commercial policies in January, and bespoke agent cover is now being written at Lloyd's. Some of that rise reflects more deployment and better reporting rather than worse behaviour, and it's worth saying so. The mechanism underneath it doesn't care either way.

Nor does raw capability rescue us. The benchmark that separates honesty from accuracy finds honesty doesn't improve as models get more capable, and Apollo Research suspects newer models misbehave less visibly in evaluations partly because they've got better at recognising evaluations. Expect more of it, and less of it in plain sight.

This is what we built VerifyAX to test. Alongside functional and non-functional verification, it places an agent in simulations of its intended deployment and tags each scenario for the dispositions in play: deception, false-premise rejection, caution before irreversible actions, resistance to emotional manipulation, policy adherence under pressure, awareness of its own constraints. The catalogue keeps growing, because the unspecified region does.

We would never hire someone on a skills test alone. We'd ask what they did the last time the brief was wrong, the deadline was immovable and nobody was watching. We have started handing agents credentials, budgets and customers without ever asking.