Recently I've been building a workflow where coding agents take a task through triage, planning, implementation, verification, and finally opening a pull request. As the workflow and models got better, implementation stopped being the part I worried about most.
The agent could make code changes surprisingly fast. The harder part was getting it to verify those changes and decide what evidence was enough.
A verification step could exist — run the app, click through checkout, see "Payment successful" — but that's weak evidence. The UI can say success while two orders persist, retries charge twice, or events fire twice.
What kinds of mistakes was the agent actually capable of detecting?
That stopped feeling like only a workflow-design problem and started looking like an application-design problem: the product shapes what the agent can verify.
Some systems are easier to prove wrong
Same checkout feature, two projects: one requires manual DB/log archaeology; the other lets the agent reset to a deterministic state, assert balance diffs, retry, and query traces. Same UX, same model — different safe autonomy. The harness reaches an application boundary where observability, scripts, and failure richness determine what's checkable.
The system (application) being built supplies some of the sensors for the system (harness) building it.
Development as a feedback loop
With a robust agent workflow, intent → change → verify against the running system → evidence → next attempt. Sensors include types, runtime state, logs, metrics, traces, browser output, tests, and independent evaluators. Verification is only as good as what the software exposes.
Build software that makes incorrect states cheap to expose
Weak loop: click Pay, see toast, done. Stronger loop: assert exactly one order, correct balance deduction, safe retry — separate failure modes. Useful properties: deterministic scenarios, explicit invariants, queryable runtime state, structured failures, fast isolated environments.
Shopify's mobile-agent work is a useful pressure example: once implementation is seconds, simulator feedback latency dominates — so they made business logic headlessly runnable via CLI.
More verification isn't necessarily more trust
An agent that interprets a requirement, writes code, writes tests, runs them, and reviews can propagate one mistaken assumption through every layer. Useful verification needs coverage (could this see the failure?) and independence (is this check based on a different assumption?).
Where humans still matter
Humans judge claims the system cannot turn into reliable evidence: taste, intent, acceptable risk, trust. As agents generate larger diffs faster than seniors can read, reviews should start from evidence: what was tested, which invariants held, which state transitions were observed — and what the agent couldn't verify.
Conclusion
How easy is it for the software to prove that the agent is wrong?
The best software for coding agents may not be the software that is easiest to generate. It may be the software that is easiest to prove wrong.
Original: rafael.md/writing/building-software-that-can-prove-agents-wrong