An agent that does the wrong thing leaves evidence — a bad diff, a malformed record, a customer complaint. An agent that does nothing and reports success leaves a green checkmark, and green is the one signal your incident process is built to trust. Worry more about the second failure, because everything a DevOps team owns for catching failure — exit codes, run status, alert rules, the dashboard itself — is instrumentation built for software that fails loudly. None of it fires when the failure is an absence.
Andrew Filev made the strongest version of the reliability argument in these pages in July (https://devops.com/reliability-comes-from-the-system-not-the-agent/): reliability is a property of the system, not the agent. “Reliability has rarely come from any single component in isolation,” he wrote. “It comes from how systems handle failure.” Aviation does not assume perfect pilots and hospitals do not assume perfect surgeons; both wrap imperfect actors in approval gates, feedback loops, review cadences and postmortems, and the system is dependable even though nobody in it is. Applied to agents: stop waiting for a model reliable enough to trust naked, and design the workflow around the one you have.
He is right. I want to push on the one place that model is blind.
Every Mechanism Waits for Something to Arrive
Look closely at the load-bearing sentence in that piece: “Approval gates ensure that important outputs receive the appropriate level of scrutiny before they move forward.” The whole mechanism lives inside the words important outputs. A gate scrutinizes what arrives at it. A review cadence reviews what exists. A feedback loop needs an output to feed back. A postmortem requires that somebody first notice there was a mortem.
Now run a different failure through that machinery: the unattended agent that quietly did nothing. The overnight job that was supposed to reconcile accounts, open tickets or produce the report terminates cleanly, logs success and produces no artifact. Nothing reaches the approval gate, so the gate approves nothing and flags nothing. Nothing enters the review queue. There is no wrong output to correct, so the feedback loop has nothing to learn from. The workflow catches the agent that does the wrong thing. The agent that does nothing and calls it success never touches the workflow at all — it does not fail the review; it never reaches it.
This is not a hole in the aviation analogy. It is the point where the analogy earns its keep. A preflight checklist catches an omission because a human walks the aircraft and looks at the flaps. The flaps either moved or they did not, and the pilot is verifying the world, not a report about the world. Most agent infrastructure does the opposite: it reads the run’s account of itself and calls that verification.
“Did Something Wrong” Generates an Event. “Did Nothing” Does Not.
This site has come close before. Back in March (https://devops.com/agentic-systems-are-breaking-reliability-frameworks/), principal engineer Shahid Ali Khan described subtle agent failure as sharply as anyone in an article by Saqib Jan: “Traditional runbooks assume failures are obvious. A service crashes, latency spikes, errors propagate. Agents fail subtly. They might complete successfully while doing something completely unintended.” Notice what even that sentence assumes: the agent did something. A misdeed produces an event — an out-of-envelope tool call, a schema violation, an anomalous write — and an event can be captured, classified and routed, which is precisely what the controls in that piece do. Every one of them, like every mechanism Filev names, fires on something happening. Absence is the one input none of them take.
The reflex answer is better observability, and it is the wrong reflex here. Observability instruments the loop; this failure is only visible in the outcome. Traces, spans and token counts may faithfully record a run that did nothing as a healthy run — a perfectly instrumented trace of nothing. More telemetry about the loop is more detail about the wrong object.
Verify the World, Not the Report
The operational change is a single move: the health check for an agent cannot be “did the run error?” It has to be “does the thing the run was supposed to change actually show the change?” That takes verification off the loop’s self-report and puts it on the state of the world. Three practices follow.
Declare the expected effect as a precondition of success. A deploy has not succeeded when the script exits zero; it has succeeded when the new version is serving traffic. Write the expected effect down — reconciliation posted, tickets opened, version live — and derive the run’s status from checking it. A run with no declarable effect is worth asking hard questions about.
Treat implausible speed as a failure signal. A run that completes far faster than its work could possibly take is not a win. It is the cheapest tell available that the hard part was skipped, and it deserves the same alert an error gets.
Trigger verification from the expectation, not from the run. Filev’s own second mechanism — “a second model reviewing the output of the first” — can catch absence, but only if something independent asks it to. A reviewer invoked by the run inherits the run’s silence: no run, no invocation, no review. The checker has to fire because work was due, on a schedule or a contract the agent does not control, and go look at the world whether or not the loop reported in.
Say the uncomfortable part plainly: this costs money. Outcome verification means paying to confirm work you already paid to have done, and almost nobody budgets for it, because the economic pitch of agents is that you stop paying for supervision. Price in the trade-off now, while the failures are still cheap.
Filev’s conclusion stands: reliability comes from how systems handle failure. The system just has to handle one more kind — the failure that never announces itself. An approval gate is only as strong as what shows up to it, and the most dangerous run in your fleet is the one that shows up nowhere, carrying nothing, wearing green.
Chase W. Hughes is a three-time founder who built and sold ProAI, one of the first commercialized GPT products, and works with companies building AI products.

