Human-in-the-loop risk gates for refunds, inventory, and spend: classify mutations, issue system approval ids, and verify with a second read before user-facing success copy.
HITL gates for agent mutations
Human-in-the-loop risk gates for refunds, inventory, and spend. Verify mutations after tool calls so agents cannot declare early victory.
Agents are optimistic. They call a tool, see a JSON blob that looks success-shaped, and tell the user "refund sent." Sometimes the refund is pending. Sometimes the tool returned a structured error the model ignored. Sometimes the write never happened.
Human-in-the-loop (HITL) gates and post-mutation verification are how you stop early victory. The model can propose. Policy decides. The system checks reality before anyone celebrates.
Which mutations need a gate
Not every tool needs a human. Read-only lookups should be fast. Mutations that move money, stock, or trust should slow down.
Risk class
Examples
Default gate
Money out
Refunds, payouts, store credit
Approve above threshold; auto only inside policy
Inventory
Force adjust, unpublish, bulk price
Approve or dual-control for large deltas
Spend / buy
Agent checkout complete, ad spend
Buyer confirm or spend ceiling
Credentials
Rotate keys, invite admins
Always human for production
Irreversible comms
Mass email, legal notices
Approve + preview
Thresholds are product decisions. Encode ceilings in policy config, not prompt text.
Gate shapes that work
Pre-tool approval: agent prepares a proposed_action payload; UI or Slack asks a human to approve; only then does execute_refund run with the approval id.
Two-step tools:draft_refund (safe) then commit_refund (requires approval_token).
Async review queue: mutation creates a pending record; worker waits for approve/deny; agent polls status.
Policy auto-approve: small, low-risk, fully validated cases skip the human but still write audit + verification.
Always bind approvals to actor, tenant, exact action hash / idempotency key, and expiry. Do not let the model invent an approval_id.
Anti early-victory: verify after tools
The failure mode is consistent across stacks:
Tool returns something the model interprets as success.
Model narrates completion to the user.
Downstream system never applied the change.
Verification is a second read (or event) that proves the world moved—after refund, inventory adjust, spend, or key rotate. Put verification in code the agent must call (or that your tool runner calls automatically).