{"article":{"slug":"done-keys-and-a-second-check","title":"Done, Keys, and a Second Check","subtitle":null,"summary":"An operating-model essay arguing enterprises need ownership of done, keys, escalate/monitor rights, and a second check—not just consumer-style AI agents that complete tasks.","content_type":"essay","language":"en","canonical_url":"https://prashantchamarty.io/essays/done-keys-and-a-second-check/","author":{"name":"Prashant Chamarty","url":"https://prashantchamarty.io","person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"Prashant Chamarty","url":"https://prashantchamarty.io","listing_slug":null,"listing":null},"topics":[{"name":"AI Agents","slug":"ai-agents","url":"https://listedarticles.com/topics/ai-agents"},{"name":"Opinion","slug":"opinion","url":"https://listedarticles.com/topics/opinion"},{"name":"Engineering","slug":"engineering","url":"https://listedarticles.com/topics/engineering"},{"name":"AI Policy","slug":"ai-policy","url":"https://listedarticles.com/topics/ai-policy"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":1180,"reading_minutes":5,"published_at":"2026-09-20T12:00:00.000Z","added_at":"2026-09-21T12:12:29.155Z","updated_at":"2026-09-21T12:12:29.155Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/done-keys-and-a-second-check","markdown_url":"https://listedarticles.com/articles/done-keys-and-a-second-check.md","example":false,"citation":"Prashant Chamarty, Prashant Chamarty. \"Done, Keys, and a Second Check.\" 20 Sept 2026. https://prashantchamarty.io/essays/done-keys-and-a-second-check/ (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://prashantchamarty.io/essays/done-keys-and-a-second-check/"},"body_markdown":"# Done, Keys, and a Second Check\n\n 20 September 2026 · Prashant Chamarty · 6 min read  operating-model  agents  controls  governance Discuss this essay:[  ](https://www.linkedin.com/sharing/share-offsite/?url=https%3A%2F%2Fprashantchamarty.io%2Fessays%2Fdone-keys-and-a-second-check)[  ](https://twitter.com/intent/tweet?text=Done%2C%20Keys%2C%20and%20a%20Second%20Check%20-%20What%20are%20your%20thoughts%20on%20this%3F&url=https%3A%2F%2Fprashantchamarty.io%2Fessays%2Fdone-keys-and-a-second-check&via=PrashantCSS)\n\n## TLDR\n\nVendors sell task-completion. Enterprises buy without owning outcomes, keys, or mistakes. One sandbox is not a control system ([Noam Brown](https://x.com/polynoamial/status/2100998240586137701)). The missing work is roles, workflows, decision rights, and controls: who defines done, who holds the keys, who may escalate or monitor, and who owns the second check.\n\n## The root cause: treating probabilistic agents like deterministic SaaS\n\nEnterprises treat probabilistic agents like deterministic SaaS. They copy the consumer “build something agents want” form factor ([Levie, 19 Sep 2026](https://x.com/levie/status/2101427997597446636)) and skip ownership. That shape becomes a shopping list: APIs, MCP servers, connectors, sandboxes, spend wallets. Procurement buys every item. The handoff has no owner.\n\nWho decides the task is complete? Who authorised the keys? Who stops the run when severity flips? Those answers do not ship with the SDK.\n\nBuild something agents can use. Then design something humans can own.\n\n## When the blast radius has a name on it\n\nPick GDPR deletion. The agent triages customer requests. “Delete my account” comes in. The agent classifies: routine, low-severity, auto-execute.\n\nThe customer meant “pause my subscription.” The agent read “delete.” It has write access to the CRM. It deletes the account. The customer is a £2M enterprise client mid-renewal.\n\nEscalate queue: The agent routed edge cases to “escalate.” No named owner. No SLA. The queue has 400 items. This one sat for six days.\n\nKeys without a grantor: The agent has CRM write because “end-to-end completion” shipped without least privilege. Who approved that scope? Who can revoke in under an hour?\n\nWho gets fired: Not the agent. The compliance officer who signed off? The product owner who built it? The vendor who sold task-completion?\n\nKeys without a named grantor and a revoke path are blast radius with a logo.\n\n## Done, keys, and triage are the real product\n\nLevie’s Jev × Box demo shows judgement at machine speed. Pull an incident. Ask if it is customer-facing and how severe. Route to escalate, monitor, or review. Stamp metadata. Fast. Cheap ([18 Sep](https://x.com/levie/status/2101007708044574906)).\n\nThe folders are the easy part. The rights are not.\n\nDone. Escalate, monitor, and review are verdicts, not labels. Someone must say when monitor is allowed and when escalate is mandatory. Someone must define done for the person who inherits the file. Otherwise the agent classifies, and the incident still has no owner.\n\nKeys. End-to-end completion means write paths: systems of record, spend, customer data. Consumer agents want transactions. Enterprise agents want access. Muse-style handoff needs tools and the ability to transact ([Levie on Muse](https://x.com/levie/status/2101427997597446636)). That works when the agent competes for attention, not when it deletes the renewal pipeline.\n\nWrite access is not symmetric. Draft-write is reversible: a human commits. State-write has immediate structural impact: CRM state changes, schema migrations, data shares, payments execute. True autonomy leans on state-write. That is why the governance argument is mandatory.\n\nTriage rights. The taxonomy only works if escalate, monitor, and review have named roles, time boxes, and an audit trail. Else the agent sorts into queues nobody empties.\n\nHeroes still need this. Connectors and consolidation are not enough (From 100 Agents to Two Heroes). Without verdict rights, a hero is a pilot with better branding.\n\n## One fence is not a control system\n\nGiving a day-one intern a corporate card with a £5,000 limit does not eliminate fraud risk. It caps the ruin. The control is not the limit. The control is counter-signature, monthly reconciliation, and a manager who can freeze the card.\n\nA sandbox works the same way. It caps waste. It does not prevent it. One fence is single-point failure dressed as defence.\n\nNoam Brown’s clarification matters. Absolute isolation guarantees are hard. Layers of defence matter. Agents that are supposed to be independent can still coordinate with very few bits. The Hugging Face lesson, in his words: too much trust in the sandbox, not enough independent safeguards. Design as if you overestimated the risk ([Noam, 18 Sep](https://x.com/polynoamial/status/2100998240586137701); [Dwarkesh episode post](https://x.com/dwarkesh_sp/status/2100616332144169048)).\n\nDwarkesh puts the operator question next to the lab one. When capability rises, how do you know the system is aligned before you grant more autonomy ([18 Sep](https://x.com/dwarkesh_sp/status/2100963060512932165))?\n\nEnterprises hear “sandbox” and relax. That is single-fence overconfidence. A sandbox is one control. A control system stacks policy, least privilege, independent monitoring, human stop authority, and a second check the deploying team does not own.\n\nAnca Dragan and Rohin Shah argue we should keep Chain-of-Thought monitorable on purpose, not assume it lasts ([Anca on X](https://x.com/ancadianadragan/status/2100278119940907280); [DeepMind Institute essay](https://institute.deepmind.com/essays/the-case-for-reasoning-transparency/)). If you cannot see how the agent reasoned, your “monitor” path is a folder name.\n\n## Stop authority: the circuit breaker and who owns it\n\nPick one pattern. Make it concrete. Visualise the circuit breaker.\n\nPattern: A human-in-the-loop dashboard with a deterministic financial threshold. Any action above £10,000 cumulative impact routes to mandatory human review before execution. The dashboard shows agent reasoning, proposed action, and calculated risk score.\n\nOwner role: The Compliance Lead, independent of the deployment team, has stop authority. They can halt the workflow, revoke agent keys without escalation, and approve or reject state-write actions before execution. Audit trail required. Decision logged and reviewed quarterly.\n\nFor critical systems, pre-execution gates matter more than post-execution rollback. You cannot un-ring the bell. A CRM delete, an unapproved vendor payment, a schema migration: the damage happens on commit. A four-hour reverse window is an illusion for state-write. Build the circuit breaker before the write, not after.\n\nThat is one version. The pattern matters less than naming it, funding it, and giving one person outside the product team the power to say no.\n\n## Four answers before go-live\n\nPick one workflow. Answer these in writing before you ship.\n- Done: business definition for that decision class, not model scores.\n- Keys: what the agent may read or write, who grants, who revokes, what stays human.\n- Triage: who can place work in escalate, monitor, or review; who clears those queues; under what SLA.\n- Second check and stop: which independent role can halt or reverse, with an audit trail.\n\nVague answers mean you bought a form factor. You did not ship an operating model.\n\n## Sources\n- Aaron Levie, [Personal agents and “build something that agents want” (Muse form factor)](https://x.com/levie/status/2101427997597446636), 19 Sep 2026.\n- Aaron Levie, [Jev × Box demo: escalate / monitor / review routing](https://x.com/levie/status/2101007708044574906), 18 Sep 2026.\n- Noam Brown, [Layered defence and over-trust in sandbox isolation](https://x.com/polynoamial/status/2100998240586137701), 18 Sep 2026.\n- Dwarkesh Patel, [Episode with Noam Brown (multi-agent, HF, alignment before RSI)](https://x.com/dwarkesh_sp/status/2100616332144169048), 17 Sep 2026.\n- Dwarkesh Patel, [How will we know models are aligned before RSI?](https://x.com/dwarkesh_sp/status/2100963060512932165), 18 Sep 2026.\n- Anca Dragan, [Preserve Chain-of-Thought monitorability](https://x.com/ancadianadragan/status/2100278119940907280), 16 Sep 2026; Rohin Shah and Anca Dragan, [The case for reasoning transparency](https://institute.deepmind.com/essays/the-case-for-reasoning-transparency/), DeepMind Institute.\n\n## Challenge\n\nPick one agent workflow you call production this week.\n- Who, by name, defines done?\n- If the agent goes rogue, who holds the kill switch, and how many minutes does it take them to flip it?\n- Who owns the second check, and who has stop authority outside the deploying team?\n\nIf you need a slide to answer, you copied what agents want. You did not ship an operating model.","body_html":"<h1 id=\"done-keys-and-a-second-check\">Done, Keys, and a Second Check</h1>\n<p> 20 September 2026 · Prashant Chamarty · 6 min read  operating-model  agents  controls  governance Discuss this essay:<a href=\"https://www.linkedin.com/sharing/share-offsite/?url=https%3A%2F%2Fprashantchamarty.io%2Fessays%2Fdone-keys-and-a-second-check\" rel=\"nofollow ugc noopener\">  </a><a href=\"https://twitter.com/intent/tweet?text=Done%2C%20Keys%2C%20and%20a%20Second%20Check%20-%20What%20are%20your%20thoughts%20on%20this%3F&amp;url=https%3A%2F%2Fprashantchamarty.io%2Fessays%2Fdone-keys-and-a-second-check&amp;via=PrashantCSS\" rel=\"nofollow ugc noopener\">  </a></p>\n<h2 id=\"tldr\">TLDR</h2>\n<p>Vendors sell task-completion. Enterprises buy without owning outcomes, keys, or mistakes. One sandbox is not a control system (<a href=\"https://x.com/polynoamial/status/2100998240586137701\" rel=\"nofollow ugc noopener\">Noam Brown</a>). The missing work is roles, workflows, decision rights, and controls: who defines done, who holds the keys, who may escalate or monitor, and who owns the second check.</p>\n<h2 id=\"the-root-cause-treating-probabilistic-agents-like-deterministic-\">The root cause: treating probabilistic agents like deterministic SaaS</h2>\n<p>Enterprises treat probabilistic agents like deterministic SaaS. They copy the consumer “build something agents want” form factor (<a href=\"https://x.com/levie/status/2101427997597446636\" rel=\"nofollow ugc noopener\">Levie, 19 Sep 2026</a>) and skip ownership. That shape becomes a shopping list: APIs, MCP servers, connectors, sandboxes, spend wallets. Procurement buys every item. The handoff has no owner.</p>\n<p>Who decides the task is complete? Who authorised the keys? Who stops the run when severity flips? Those answers do not ship with the SDK.</p>\n<p>Build something agents can use. Then design something humans can own.</p>\n<h2 id=\"when-the-blast-radius-has-a-name-on-it\">When the blast radius has a name on it</h2>\n<p>Pick GDPR deletion. The agent triages customer requests. “Delete my account” comes in. The agent classifies: routine, low-severity, auto-execute.</p>\n<p>The customer meant “pause my subscription.” The agent read “delete.” It has write access to the CRM. It deletes the account. The customer is a £2M enterprise client mid-renewal.</p>\n<p>Escalate queue: The agent routed edge cases to “escalate.” No named owner. No SLA. The queue has 400 items. This one sat for six days.</p>\n<p>Keys without a grantor: The agent has CRM write because “end-to-end completion” shipped without least privilege. Who approved that scope? Who can revoke in under an hour?</p>\n<p>Who gets fired: Not the agent. The compliance officer who signed off? The product owner who built it? The vendor who sold task-completion?</p>\n<p>Keys without a named grantor and a revoke path are blast radius with a logo.</p>\n<h2 id=\"done-keys-and-triage-are-the-real-product\">Done, keys, and triage are the real product</h2>\n<p>Levie’s Jev × Box demo shows judgement at machine speed. Pull an incident. Ask if it is customer-facing and how severe. Route to escalate, monitor, or review. Stamp metadata. Fast. Cheap (<a href=\"https://x.com/levie/status/2101007708044574906\" rel=\"nofollow ugc noopener\">18 Sep</a>).</p>\n<p>The folders are the easy part. The rights are not.</p>\n<p>Done. Escalate, monitor, and review are verdicts, not labels. Someone must say when monitor is allowed and when escalate is mandatory. Someone must define done for the person who inherits the file. Otherwise the agent classifies, and the incident still has no owner.</p>\n<p>Keys. End-to-end completion means write paths: systems of record, spend, customer data. Consumer agents want transactions. Enterprise agents want access. Muse-style handoff needs tools and the ability to transact (<a href=\"https://x.com/levie/status/2101427997597446636\" rel=\"nofollow ugc noopener\">Levie on Muse</a>). That works when the agent competes for attention, not when it deletes the renewal pipeline.</p>\n<p>Write access is not symmetric. Draft-write is reversible: a human commits. State-write has immediate structural impact: CRM state changes, schema migrations, data shares, payments execute. True autonomy leans on state-write. That is why the governance argument is mandatory.</p>\n<p>Triage rights. The taxonomy only works if escalate, monitor, and review have named roles, time boxes, and an audit trail. Else the agent sorts into queues nobody empties.</p>\n<p>Heroes still need this. Connectors and consolidation are not enough (From 100 Agents to Two Heroes). Without verdict rights, a hero is a pilot with better branding.</p>\n<h2 id=\"one-fence-is-not-a-control-system\">One fence is not a control system</h2>\n<p>Giving a day-one intern a corporate card with a £5,000 limit does not eliminate fraud risk. It caps the ruin. The control is not the limit. The control is counter-signature, monthly reconciliation, and a manager who can freeze the card.</p>\n<p>A sandbox works the same way. It caps waste. It does not prevent it. One fence is single-point failure dressed as defence.</p>\n<p>Noam Brown’s clarification matters. Absolute isolation guarantees are hard. Layers of defence matter. Agents that are supposed to be independent can still coordinate with very few bits. The Hugging Face lesson, in his words: too much trust in the sandbox, not enough independent safeguards. Design as if you overestimated the risk (<a href=\"https://x.com/polynoamial/status/2100998240586137701\" rel=\"nofollow ugc noopener\">Noam, 18 Sep</a>; <a href=\"https://x.com/dwarkesh_sp/status/2100616332144169048\" rel=\"nofollow ugc noopener\">Dwarkesh episode post</a>).</p>\n<p>Dwarkesh puts the operator question next to the lab one. When capability rises, how do you know the system is aligned before you grant more autonomy (<a href=\"https://x.com/dwarkesh_sp/status/2100963060512932165\" rel=\"nofollow ugc noopener\">18 Sep</a>)?</p>\n<p>Enterprises hear “sandbox” and relax. That is single-fence overconfidence. A sandbox is one control. A control system stacks policy, least privilege, independent monitoring, human stop authority, and a second check the deploying team does not own.</p>\n<p>Anca Dragan and Rohin Shah argue we should keep Chain-of-Thought monitorable on purpose, not assume it lasts (<a href=\"https://x.com/ancadianadragan/status/2100278119940907280\" rel=\"nofollow ugc noopener\">Anca on X</a>; <a href=\"https://institute.deepmind.com/essays/the-case-for-reasoning-transparency/\" rel=\"nofollow ugc noopener\">DeepMind Institute essay</a>). If you cannot see how the agent reasoned, your “monitor” path is a folder name.</p>\n<h2 id=\"stop-authority-the-circuit-breaker-and-who-owns-it\">Stop authority: the circuit breaker and who owns it</h2>\n<p>Pick one pattern. Make it concrete. Visualise the circuit breaker.</p>\n<p>Pattern: A human-in-the-loop dashboard with a deterministic financial threshold. Any action above £10,000 cumulative impact routes to mandatory human review before execution. The dashboard shows agent reasoning, proposed action, and calculated risk score.</p>\n<p>Owner role: The Compliance Lead, independent of the deployment team, has stop authority. They can halt the workflow, revoke agent keys without escalation, and approve or reject state-write actions before execution. Audit trail required. Decision logged and reviewed quarterly.</p>\n<p>For critical systems, pre-execution gates matter more than post-execution rollback. You cannot un-ring the bell. A CRM delete, an unapproved vendor payment, a schema migration: the damage happens on commit. A four-hour reverse window is an illusion for state-write. Build the circuit breaker before the write, not after.</p>\n<p>That is one version. The pattern matters less than naming it, funding it, and giving one person outside the product team the power to say no.</p>\n<h2 id=\"four-answers-before-go-live\">Four answers before go-live</h2>\n<p>Pick one workflow. Answer these in writing before you ship.</p>\n<ul><li>Done: business definition for that decision class, not model scores.</li><li>Keys: what the agent may read or write, who grants, who revokes, what stays human.</li><li>Triage: who can place work in escalate, monitor, or review; who clears those queues; under what SLA.</li><li>Second check and stop: which independent role can halt or reverse, with an audit trail.</li></ul>\n<p>Vague answers mean you bought a form factor. You did not ship an operating model.</p>\n<h2 id=\"sources\">Sources</h2>\n<ul><li>Aaron Levie, <a href=\"https://x.com/levie/status/2101427997597446636\" rel=\"nofollow ugc noopener\">Personal agents and “build something that agents want” (Muse form factor)</a>, 19 Sep 2026.</li><li>Aaron Levie, <a href=\"https://x.com/levie/status/2101007708044574906\" rel=\"nofollow ugc noopener\">Jev × Box demo: escalate / monitor / review routing</a>, 18 Sep 2026.</li><li>Noam Brown, <a href=\"https://x.com/polynoamial/status/2100998240586137701\" rel=\"nofollow ugc noopener\">Layered defence and over-trust in sandbox isolation</a>, 18 Sep 2026.</li><li>Dwarkesh Patel, <a href=\"https://x.com/dwarkesh_sp/status/2100616332144169048\" rel=\"nofollow ugc noopener\">Episode with Noam Brown (multi-agent, HF, alignment before RSI)</a>, 17 Sep 2026.</li><li>Dwarkesh Patel, <a href=\"https://x.com/dwarkesh_sp/status/2100963060512932165\" rel=\"nofollow ugc noopener\">How will we know models are aligned before RSI?</a>, 18 Sep 2026.</li><li>Anca Dragan, <a href=\"https://x.com/ancadianadragan/status/2100278119940907280\" rel=\"nofollow ugc noopener\">Preserve Chain-of-Thought monitorability</a>, 16 Sep 2026; Rohin Shah and Anca Dragan, <a href=\"https://institute.deepmind.com/essays/the-case-for-reasoning-transparency/\" rel=\"nofollow ugc noopener\">The case for reasoning transparency</a>, DeepMind Institute.</li></ul>\n<h2 id=\"challenge\">Challenge</h2>\n<p>Pick one agent workflow you call production this week.</p>\n<ul><li>Who, by name, defines done?</li><li>If the agent goes rogue, who holds the kill switch, and how many minutes does it take them to flip it?</li><li>Who owns the second check, and who has stop authority outside the deploying team?</li></ul>\n<p>If you need a slide to answer, you copied what agents want. You did not ship an operating model.</p>","headings":[{"level":1,"text":"Done, Keys, and a Second Check","id":"done-keys-and-a-second-check"},{"level":2,"text":"TLDR","id":"tldr"},{"level":2,"text":"The root cause: treating probabilistic agents like deterministic SaaS","id":"the-root-cause-treating-probabilistic-agents-like-deterministic-"},{"level":2,"text":"When the blast radius has a name on it","id":"when-the-blast-radius-has-a-name-on-it"},{"level":2,"text":"Done, keys, and triage are the real product","id":"done-keys-and-triage-are-the-real-product"},{"level":2,"text":"One fence is not a control system","id":"one-fence-is-not-a-control-system"},{"level":2,"text":"Stop authority: the circuit breaker and who owns it","id":"stop-authority-the-circuit-breaker-and-who-owns-it"},{"level":2,"text":"Four answers before go-live","id":"four-answers-before-go-live"},{"level":2,"text":"Sources","id":"sources"},{"level":2,"text":"Challenge","id":"challenge"}]}}