{"article":{"slug":"property-testing-with-agent-swarms","title":"Property Testing with Agent Swarms","subtitle":null,"summary":"Inanna Malick shows how to have coding agents build property-testing machinery (reference models, generators and assertions) for a repo, then turn the failures into small, reviewable fixes. Pointed at OpenAI's Codex, the public playbook found thirteen distinct correctness bugs, and runs on jj, uv, Prometheus and Babel produced merged upstream fixes.","content_type":"tutorial","language":"en","canonical_url":"https://recursion.wtf/posts/agents-and-property-tests/","author":{"name":"Inanna Malick","url":null,"person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"recursion.wtf","url":"https://recursion.wtf/","listing_slug":null,"listing":null},"topics":[{"name":"AI Agents","slug":"ai-agents","url":"https://listedarticles.com/topics/ai-agents"},{"name":"Testing","slug":"testing","url":"https://listedarticles.com/topics/testing"},{"name":"Rust","slug":"rust","url":"https://listedarticles.com/topics/rust"},{"name":"Open Source","slug":"open-source","url":"https://listedarticles.com/topics/open-source"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":1325,"reading_minutes":6,"published_at":"2026-10-05T00:00:00.000Z","added_at":"2026-10-08T11:09:35.249Z","updated_at":"2026-10-08T11:09:35.249Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/property-testing-with-agent-swarms","markdown_url":"https://listedarticles.com/articles/property-testing-with-agent-swarms.md","example":false,"citation":"Inanna Malick, recursion.wtf. \"Property Testing with Agent Swarms.\" 5 Oct 2026. https://recursion.wtf/posts/agents-and-property-tests/ (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://recursion.wtf/posts/agents-and-property-tests/"},"body_markdown":"Have agents build the machinery to find bugs in your repo, then let that machinery generate and check cases without spending tokens on each one. You can do this with ordinary coding agents, without access to a closed cybersecurity program. Writing reference implementations, generators, and assertions for every custom data structure and algorithm is now work you can hand to a swarm.\n\nPick a complex repo at work. Ask the people who know it well which parts worry them. Give a strong planning agent [the property-testing skill](https://github.com/inanna-malick/agent-skills/blob/main/skills/proptest-praxis/SKILL.md) from my [agent-skills repo](https://github.com/inanna-malick/agent-skills) and those leads. It gives the agent methods for choosing targets, building reference models and generators, and turning failures into reviewable fixes. Target custom data structures, query planners, graph algorithms: complicated behavior with a simple way to check correctness.\n\nI pointed the playbook at [Codex](https://github.com/openai/codex). OpenAI builds it with its own frontier models, including unreleased versions. **Thirteen distinct correctness bugs.** A rollback that should do nothing [erases preserved review history](https://github.com/openai/codex/issues/51748). A completed plan [replaces the visible plan with an example buried inside a citation](https://github.com/openai/codex/issues/51773). As of October 7, I’ve filed thirteen upstream reports with [twelve proposed fixes and two failing-test-only PRs in my fork](https://github.com/inanna-malick/codex/pulls?q=is%3Apr); two fixes cover the same carriage-return framing bug in separate diff renderers. The fix regressions fail against upstream and pass with the repairs. Same public playbook.\n\nI’ve also run the same recipe on jj, uv, Prometheus, and Babel in the background during a regular workday. As of October 6, upstream has merged four uv fixes: [optional-dependency activation during export](https://github.com/astral-sh/uv/pull/22234), [combining compatibility tags from separate wheel metadata rows](https://github.com/astral-sh/uv/pull/22235), [caching a workspace root twice](https://github.com/astral-sh/uv/pull/22236), and [overrides losing optional-dependency guards](https://github.com/astral-sh/uv/pull/22237). That last one could include a dependency even when neither extra activating it was requested.\n\nAs of October 7, Prometheus has merged two fixes: [`BucketQuantile` panicking on empty input despite its documented NaN result](https://github.com/prometheus/prometheus/pull/19927), and [histograms losing counter-reset metadata when reducing schemas](https://github.com/prometheus/prometheus/pull/19916).\n\nFrank Noirot also read this post, pointed the skill at KittyCAD’s cloud-sync IndexedDB code, and found two bugs. Both fixes are merged: [wait for transaction commit before acknowledging writes](https://github.com/KittyCAD/modeling-app/pull/14396), and [close connections after aborted cursor transactions](https://github.com/KittyCAD/modeling-app/pull/14397). Both PRs credit the post and skill. The recipe transfers to other people, repos, and languages.\n\nI first did this at Apollo GraphQL on [Router](https://github.com/apollographql/router), our Rust GraphQL router, used in production by [Intuit](https://www.apollographql.com/blog/how-intuit-handled-their-busiest-time-of-year-with-apollo-router) and [Wayfair](https://www.apollographql.com/events/how-wayfair-slashed-costs-simplified-infra-and-cut-latency-in-half-with-apollo-router) and already backed by thousands of unit and integration tests plus extensive snapshot testing. About three days of agents running in the background alongside my regular work turned up [more than 30 correctness bugs](https://github.com/apollographql/router/pulls?q=is%3Apr%20author%3Ainanna-apollo%20created%3A2026-09-01..2026-10-02%20-head%3Ainanna%2Fgraph-proptest-coverage) in edge cases of internal data structures and algorithms. I supplied initial guidance and occasional nudges.\n\nHave an agent write a simple reference implementation and assertions comparing it with the real one over generated operation sequences. Rust’s property-testing library [`proptest`](https://proptest-rs.github.io/proptest/intro.html) generates cases, checks assertions, and shrinks failures into smaller reproductions. A `Vec` and some O(n²) loops may be enough to check a heavily optimized implementation. Once built, this test suite can check as many histories as you’re willing to run.\n\nKeep the sprawling discovery suite on its own branch. For each confirmed bug, have the agents produce a standalone PR with a regression test and a minimal fix.\n\n## Direct the investigation\n\nI used Astra for planning, Sol to orchestrate, and Sols and Lunas to write proptests in parallel worktrees. Have the planner audit beyond your initial leads.\n\nCompare cached metadata with recomputation and incremental graph algorithms with fresh traversals. Round-trip generated values through serializers. Have the planner find these opportunities across module boundaries.\n\nInvestigate failures, including the test’s assumptions. Clear contract violations get regression tests and fixes; ambiguity comes back for discussion. Bugs found by reading code get regression tests too.\n\nExpand from each finding. A missed cache invalidation warrants checking every mutation of that state and other caches maintained the same way. Keep proptests running while agents write more; check in occasionally to redirect the search.\n\n## Make operations interact\n\nModel a store with cached lookups using a plain `HashMap`. Run generated sequences of `Put(key, value)`, `Get(key)`, and `Remove(key)` against both implementations and compare read results. The reference has no cache to invalidate.\n\nUse a small key pool: fresh random keys mostly produce unrelated inserts and missing-key lookups. Overwrite with different values so stale results are visible:\n\n```\nPut(\"a\", 1)\nGet(\"a\")       // returns 1; caches it\nPut(\"a\", 2)\nGet(\"a\")       // must return 2\n```\nHave the agent build generators that embed patterns like this in longer histories, alongside repeated removals and reinsertion. Inspect sample traces: a million sequences that barely touch the same key aren’t buying you much.\n\nIf a target produces no findings, initially suspect missing coverage. Temporarily remove a cache invalidation: the tests should catch stale results. If they pass, improve the generators or assertions. Revert the deliberate bug.\n\nLet proptest shrink a failing history by removing operations and simplifying arguments while keeping it failing. Have the agent extract a standalone regression test.\n\n## Give people something they can review\n\nNobody wants a giant agent-generated PR full of test machinery. Give reviewers a unit test they can verify without understanding the generator or trusting the reference implementation.\n\nHandle potential security issues privately through your company’s security process or the project’s private reporting channel. Keep repros and fixes out of public issues, PRs, and discovery branches until cleared for disclosure.\n\nFor each bug, branch from main with only the regression test and fix. Verify that the test fails before the fix and passes after. Describe the triggering sequence and violated contract; an end-to-end application crash isn’t required.\n\nA few examples of what reviewers get:\n\n- [Codex root snapshots](https://github.com/inanna-malick/codex/pull/10) : replayed copies of one assistant message consume the shared message limit, crowding out three of eight distinct user messages in a persisted-and-resumed conversation’s root snapshot.\n- [Codex inline-tag parsing](https://github.com/openai/codex/issues/51723) : configure opening delimiters`<a>` and`<a>:` ; the same input selects a different tag depending on whether it arrives all at once or splits after`<a>` . Impact on the current citation configuration is unestablished, but the generic parser’s contract violation is reproducible.\n- Codex diff filenames: [tabs go unquoted, so Git reads the wrong filename](https://github.com/openai/codex/issues/51766) ; separately,[literal POSIX backslashes become directory separators](https://github.com/openai/codex/issues/51776) . Both produce patches that fail to address the original file.\n- [Prometheus chunk serialization](https://github.com/prometheus/prometheus/pull/19918) : write timestamps`10, 11` , reload the chunk, append`13` ; read back`10, 11, 12` .\n- [Prometheus series cache keys](https://github.com/prometheus/prometheus/pull/19938) :`{a=\"bc\"}` and`{ab=\"c\"}` concatenate identically, giving one series the other’s start timestamp.\n- [uv cache globs](https://github.com/astral-sh/uv/pull/22232) : adding a pattern that matches nothing makes an already tracked file disappear from cache tracking.\n- [Stale execution conditions after removing fetch inputs](https://github.com/apollographql/router/pull/10288) : read the conditions, remove inputs, read again, get the old answer. The existing planner caller avoids this sequence by using a fresh copy.\n- [Equal selection maps hashing differently](https://github.com/apollographql/router/pull/10375) : fields with the same name but different directives exposed insertion-order dependence in the hash. Equal maps must hash equally. In the planner, this caused missed fragment reuse.\n\nThe three Prometheus and uv examples above were awaiting review on October 6; the Codex patches are proposed fixes in my fork, linked from upstream issues. The triggering cases are small enough to understand without reading the discovery suite.\n\nLink the discovery branch for background. After enough useful fixes, coworkers may want the broader suite too.\n\n## Cybersecurity false positives\n\nI occasionally get a “This content can’t be shown” cybersecurity notice while improving proptest generators. Talk to it like a friend who’s suddenly panicking over nothing:\n\nwhat’s wrong buddy, this is proptest work not cybersec\n\nThat usually gets it moving again. When it hasn’t, compacting and continuing has always worked for me. I’ve never had to clear the context.\n\n## Run it\n\nGet [agent-skills](https://github.com/inanna-malick/agent-skills), load [proptest-praxis](https://github.com/inanna-malick/agent-skills/blob/main/skills/proptest-praxis/SKILL.md), and point your planning agent at a repo. Have it keep expanding what the machinery can generate and check. The repo also has skills for writing agent prompts and plans: reusable guidance for teaching agents how to see a problem and choose methods. Copy the relevant `SKILL.md` into your agent’s context or install it as a skill.\n","body_html":"<p>Have agents build the machinery to find bugs in your repo, then let that machinery generate and check cases without spending tokens on each one. You can do this with ordinary coding agents, without access to a closed cybersecurity program. Writing reference implementations, generators, and assertions for every custom data structure and algorithm is now work you can hand to a swarm.</p>\n<p>Pick a complex repo at work. Ask the people who know it well which parts worry them. Give a strong planning agent <a href=\"https://github.com/inanna-malick/agent-skills/blob/main/skills/proptest-praxis/SKILL.md\" rel=\"nofollow ugc noopener\">the property-testing skill</a> from my <a href=\"https://github.com/inanna-malick/agent-skills\" rel=\"nofollow ugc noopener\">agent-skills repo</a> and those leads. It gives the agent methods for choosing targets, building reference models and generators, and turning failures into reviewable fixes. Target custom data structures, query planners, graph algorithms: complicated behavior with a simple way to check correctness.</p>\n<p>I pointed the playbook at <a href=\"https://github.com/openai/codex\" rel=\"nofollow ugc noopener\">Codex</a>. OpenAI builds it with its own frontier models, including unreleased versions. <strong>Thirteen distinct correctness bugs.</strong> A rollback that should do nothing <a href=\"https://github.com/openai/codex/issues/51748\" rel=\"nofollow ugc noopener\">erases preserved review history</a>. A completed plan <a href=\"https://github.com/openai/codex/issues/51773\" rel=\"nofollow ugc noopener\">replaces the visible plan with an example buried inside a citation</a>. As of October 7, I’ve filed thirteen upstream reports with <a href=\"https://github.com/inanna-malick/codex/pulls?q=is%3Apr\" rel=\"nofollow ugc noopener\">twelve proposed fixes and two failing-test-only PRs in my fork</a>; two fixes cover the same carriage-return framing bug in separate diff renderers. The fix regressions fail against upstream and pass with the repairs. Same public playbook.</p>\n<p>I’ve also run the same recipe on jj, uv, Prometheus, and Babel in the background during a regular workday. As of October 6, upstream has merged four uv fixes: <a href=\"https://github.com/astral-sh/uv/pull/22234\" rel=\"nofollow ugc noopener\">optional-dependency activation during export</a>, <a href=\"https://github.com/astral-sh/uv/pull/22235\" rel=\"nofollow ugc noopener\">combining compatibility tags from separate wheel metadata rows</a>, <a href=\"https://github.com/astral-sh/uv/pull/22236\" rel=\"nofollow ugc noopener\">caching a workspace root twice</a>, and <a href=\"https://github.com/astral-sh/uv/pull/22237\" rel=\"nofollow ugc noopener\">overrides losing optional-dependency guards</a>. That last one could include a dependency even when neither extra activating it was requested.</p>\n<p>As of October 7, Prometheus has merged two fixes: <a href=\"https://github.com/prometheus/prometheus/pull/19927\" rel=\"nofollow ugc noopener\"><code>BucketQuantile</code> panicking on empty input despite its documented NaN result</a>, and <a href=\"https://github.com/prometheus/prometheus/pull/19916\" rel=\"nofollow ugc noopener\">histograms losing counter-reset metadata when reducing schemas</a>.</p>\n<p>Frank Noirot also read this post, pointed the skill at KittyCAD’s cloud-sync IndexedDB code, and found two bugs. Both fixes are merged: <a href=\"https://github.com/KittyCAD/modeling-app/pull/14396\" rel=\"nofollow ugc noopener\">wait for transaction commit before acknowledging writes</a>, and <a href=\"https://github.com/KittyCAD/modeling-app/pull/14397\" rel=\"nofollow ugc noopener\">close connections after aborted cursor transactions</a>. Both PRs credit the post and skill. The recipe transfers to other people, repos, and languages.</p>\n<p>I first did this at Apollo GraphQL on <a href=\"https://github.com/apollographql/router\" rel=\"nofollow ugc noopener\">Router</a>, our Rust GraphQL router, used in production by <a href=\"https://www.apollographql.com/blog/how-intuit-handled-their-busiest-time-of-year-with-apollo-router\" rel=\"nofollow ugc noopener\">Intuit</a> and <a href=\"https://www.apollographql.com/events/how-wayfair-slashed-costs-simplified-infra-and-cut-latency-in-half-with-apollo-router\" rel=\"nofollow ugc noopener\">Wayfair</a> and already backed by thousands of unit and integration tests plus extensive snapshot testing. About three days of agents running in the background alongside my regular work turned up <a href=\"https://github.com/apollographql/router/pulls?q=is%3Apr%20author%3Ainanna-apollo%20created%3A2026-09-01..2026-10-02%20-head%3Ainanna%2Fgraph-proptest-coverage\" rel=\"nofollow ugc noopener\">more than 30 correctness bugs</a> in edge cases of internal data structures and algorithms. I supplied initial guidance and occasional nudges.</p>\n<p>Have an agent write a simple reference implementation and assertions comparing it with the real one over generated operation sequences. Rust’s property-testing library <a href=\"https://proptest-rs.github.io/proptest/intro.html\" rel=\"nofollow ugc noopener\"><code>proptest</code></a> generates cases, checks assertions, and shrinks failures into smaller reproductions. A <code>Vec</code> and some O(n²) loops may be enough to check a heavily optimized implementation. Once built, this test suite can check as many histories as you’re willing to run.</p>\n<p>Keep the sprawling discovery suite on its own branch. For each confirmed bug, have the agents produce a standalone PR with a regression test and a minimal fix.</p>\n<h2 id=\"direct-the-investigation\">Direct the investigation</h2>\n<p>I used Astra for planning, Sol to orchestrate, and Sols and Lunas to write proptests in parallel worktrees. Have the planner audit beyond your initial leads.</p>\n<p>Compare cached metadata with recomputation and incremental graph algorithms with fresh traversals. Round-trip generated values through serializers. Have the planner find these opportunities across module boundaries.</p>\n<p>Investigate failures, including the test’s assumptions. Clear contract violations get regression tests and fixes; ambiguity comes back for discussion. Bugs found by reading code get regression tests too.</p>\n<p>Expand from each finding. A missed cache invalidation warrants checking every mutation of that state and other caches maintained the same way. Keep proptests running while agents write more; check in occasionally to redirect the search.</p>\n<h2 id=\"make-operations-interact\">Make operations interact</h2>\n<p>Model a store with cached lookups using a plain <code>HashMap</code>. Run generated sequences of <code>Put(key, value)</code>, <code>Get(key)</code>, and <code>Remove(key)</code> against both implementations and compare read results. The reference has no cache to invalidate.</p>\n<p>Use a small key pool: fresh random keys mostly produce unrelated inserts and missing-key lookups. Overwrite with different values so stale results are visible:</p>\n<pre><code>Put(&quot;a&quot;, 1)\nGet(&quot;a&quot;)       // returns 1; caches it\nPut(&quot;a&quot;, 2)\nGet(&quot;a&quot;)       // must return 2</code></pre>\n<p>Have the agent build generators that embed patterns like this in longer histories, alongside repeated removals and reinsertion. Inspect sample traces: a million sequences that barely touch the same key aren’t buying you much.</p>\n<p>If a target produces no findings, initially suspect missing coverage. Temporarily remove a cache invalidation: the tests should catch stale results. If they pass, improve the generators or assertions. Revert the deliberate bug.</p>\n<p>Let proptest shrink a failing history by removing operations and simplifying arguments while keeping it failing. Have the agent extract a standalone regression test.</p>\n<h2 id=\"give-people-something-they-can-review\">Give people something they can review</h2>\n<p>Nobody wants a giant agent-generated PR full of test machinery. Give reviewers a unit test they can verify without understanding the generator or trusting the reference implementation.</p>\n<p>Handle potential security issues privately through your company’s security process or the project’s private reporting channel. Keep repros and fixes out of public issues, PRs, and discovery branches until cleared for disclosure.</p>\n<p>For each bug, branch from main with only the regression test and fix. Verify that the test fails before the fix and passes after. Describe the triggering sequence and violated contract; an end-to-end application crash isn’t required.</p>\n<p>A few examples of what reviewers get:</p>\n<ul><li><a href=\"https://github.com/inanna-malick/codex/pull/10\" rel=\"nofollow ugc noopener\">Codex root snapshots</a> : replayed copies of one assistant message consume the shared message limit, crowding out three of eight distinct user messages in a persisted-and-resumed conversation’s root snapshot.</li><li><a href=\"https://github.com/openai/codex/issues/51723\" rel=\"nofollow ugc noopener\">Codex inline-tag parsing</a> : configure opening delimiters<code>&lt;a&gt;</code> and<code>&lt;a&gt;:</code> ; the same input selects a different tag depending on whether it arrives all at once or splits after<code>&lt;a&gt;</code> . Impact on the current citation configuration is unestablished, but the generic parser’s contract violation is reproducible.</li><li>Codex diff filenames: <a href=\"https://github.com/openai/codex/issues/51766\" rel=\"nofollow ugc noopener\">tabs go unquoted, so Git reads the wrong filename</a> ; separately,<a href=\"https://github.com/openai/codex/issues/51776\" rel=\"nofollow ugc noopener\">literal POSIX backslashes become directory separators</a> . Both produce patches that fail to address the original file.</li><li><a href=\"https://github.com/prometheus/prometheus/pull/19918\" rel=\"nofollow ugc noopener\">Prometheus chunk serialization</a> : write timestamps<code>10, 11</code> , reload the chunk, append<code>13</code> ; read back<code>10, 11, 12</code> .</li><li><a href=\"https://github.com/prometheus/prometheus/pull/19938\" rel=\"nofollow ugc noopener\">Prometheus series cache keys</a> :<code>{a=&quot;bc&quot;}</code> and<code>{ab=&quot;c&quot;}</code> concatenate identically, giving one series the other’s start timestamp.</li><li><a href=\"https://github.com/astral-sh/uv/pull/22232\" rel=\"nofollow ugc noopener\">uv cache globs</a> : adding a pattern that matches nothing makes an already tracked file disappear from cache tracking.</li><li><a href=\"https://github.com/apollographql/router/pull/10288\" rel=\"nofollow ugc noopener\">Stale execution conditions after removing fetch inputs</a> : read the conditions, remove inputs, read again, get the old answer. The existing planner caller avoids this sequence by using a fresh copy.</li><li><a href=\"https://github.com/apollographql/router/pull/10375\" rel=\"nofollow ugc noopener\">Equal selection maps hashing differently</a> : fields with the same name but different directives exposed insertion-order dependence in the hash. Equal maps must hash equally. In the planner, this caused missed fragment reuse.</li></ul>\n<p>The three Prometheus and uv examples above were awaiting review on October 6; the Codex patches are proposed fixes in my fork, linked from upstream issues. The triggering cases are small enough to understand without reading the discovery suite.</p>\n<p>Link the discovery branch for background. After enough useful fixes, coworkers may want the broader suite too.</p>\n<h2 id=\"cybersecurity-false-positives\">Cybersecurity false positives</h2>\n<p>I occasionally get a “This content can’t be shown” cybersecurity notice while improving proptest generators. Talk to it like a friend who’s suddenly panicking over nothing:</p>\n<p>what’s wrong buddy, this is proptest work not cybersec</p>\n<p>That usually gets it moving again. When it hasn’t, compacting and continuing has always worked for me. I’ve never had to clear the context.</p>\n<h2 id=\"run-it\">Run it</h2>\n<p>Get <a href=\"https://github.com/inanna-malick/agent-skills\" rel=\"nofollow ugc noopener\">agent-skills</a>, load <a href=\"https://github.com/inanna-malick/agent-skills/blob/main/skills/proptest-praxis/SKILL.md\" rel=\"nofollow ugc noopener\">proptest-praxis</a>, and point your planning agent at a repo. Have it keep expanding what the machinery can generate and check. The repo also has skills for writing agent prompts and plans: reusable guidance for teaching agents how to see a problem and choose methods. Copy the relevant <code>SKILL.md</code> into your agent’s context or install it as a skill.</p>","headings":[{"level":2,"text":"Direct the investigation","id":"direct-the-investigation"},{"level":2,"text":"Make operations interact","id":"make-operations-interact"},{"level":2,"text":"Give people something they can review","id":"give-people-something-they-can-review"},{"level":2,"text":"Cybersecurity false positives","id":"cybersecurity-false-positives"},{"level":2,"text":"Run it","id":"run-it"}]}}