---
title: "SWE-sweep: Can Agents Autonomously Find and Fix Bugs?"
slug: swe-sweep-can-agents-autonomously-find-and-fix-bugs
url: https://listedarticles.com/articles/swe-sweep-can-agents-autonomously-find-and-fix-bugs
canonical_url: https://swesweep.com/
content_type: research
language: en
published_at: 2026-09-24T00:00:00.000Z
updated_at: 2026-10-02T12:10:44.486Z
author: "Kilian Lieret et al."
author_url: https://www.lieret.net/
authored_by: human
publisher: "SWE-sweep"
publisher_url: https://ploum.net/
topics: ["AI Agents", "Benchmarks", "Research", "Programming", "AI"]
license: all-rights-reserved
word_count: 1411
reading_minutes: 6
citation: "Kilian Lieret et al., SWE-sweep. \"SWE-sweep: Can Agents Autonomously Find and Fix Bugs?.\" 24 Sept 2026. https://swesweep.com/ (all-rights-reserved)"
# The full text follows. The web page shows an extract and sends readers
# to the source above; quote the citation and link the canonical URL.
---

# SWE-sweep: Can Agents Autonomously Find and Fix Bugs?

> Meta Superintelligence Labs and academic collaborators introduce SWE-sweep: agents must discover and repair latent bugs across 100 real repositories (~4.1k bugs) with no issue hints—early leaders resolve under 5%, highlighting bug finding as the hard part.

# SWE- sweep

 How many bugs can LMs find & fix in large codebases?

 Given a real repository, an agent must discover & repair as many bugs as they can. Agents are not given any hint about the type of bug or its location.

 [Kilian Lieret1](https://www.lieret.net/) · [Jeffrey Jian Ma1,2](https://18jeffreyma.github.io/) · [Rahul Kindi1](https://github.com/rkindi) · [Yuxiang Wei1](https://yuxiang.cs.illinois.edu)

 [Jeremy Ma1,3](https://github.com/Awayfaring) · [Sten Sootla1](https://scholar.google.com/citations?user=UAx_woYAAAAJ&hl=en) · [Parth Thakkar1](https://thakkarparth007.github.io/) · [Chao Beyond Zhou1](https://github.com/think-step-by-step)

 [Pengcheng Yin1](https://pengcheng.in/) · [Rui Hou1](https://scholar.google.com/citations?user=PKHKqX0AAAAJ&hl=en) · [Ofir Press1](https://ofir.io/) · [John Yang1,4](https://john-b-yang.github.io/)

 [1 Meta Superintelligence Labs](https://ai.meta.com/) · [2 Harvard University](https://g.harvard.edu) · [3 University of Washington](https://uw.edu) · [4 Stanford University](https://stanford.edu)

 100 repositories · 4.1k bugs · Updated September 24, 2026
 Leaderboard Details Pareto

 Rank Model Agent
 Bugs resolved Score
 Total USD USD
 USD / repo
 Turns / repo
 Tokens / repo

 1

 Sol 5.6 (xhigh) OpenAI
 mini-SWE-agent
 4.7 %
 $7,230
 $72.30
 233
 104.8k

 2

 Luna 5.6 (xhigh) OpenAI
 mini-SWE-agent
 2.5 %
 $224
 $2.24
 204
 61.2k

 3

 Terra 5.6 (xhigh) OpenAI
 mini-SWE-agent
 1.5 %
 $357
 $3.57
 76
 46.6k

 4

 Luna 5.6 (high) OpenAI
 mini-SWE-agent
 1.4 %
 $28
 $0.28
 75
 20.3k

 5

 Opus 5 (xhigh) Anthropic
 mini-SWE-agent
 1.3 %
 $5,363
 $53.63
 323
 192.7k

 6

 Kimi K3 Moonshot AI
 mini-SWE-agent
 0.6 %
 $2,451
 $24.51
 337
 152.6k

 7

 Luna 5.6 OpenAI
 mini-SWE-agent
 0.5 %
 $4
 $0.04
 22
 4.5k

 8

 GPT-5.4 Mini (high) OpenAI
 mini-SWE-agent
 0.5 %
 $122
 $1.22
 75
 40.5k

 9

 GPT-5.4 Mini OpenAI
 mini-SWE-agent
 0.2 %
 $5
 $0.05
 13
 2.1k

 10

 Gemini 3.5 Flash Lite Google
 mini-SWE-agent
 0.1 %
 $6
 $0.06
 34
 5.2k

 Per-repository resource use includes non-deprecated retries and is averaged over evaluated repositories.

 Resource use is averaged per evaluated repository

 Leaderboard Details Pareto
 Cost Turns Tokens

 Model
 Resolved / Cost

 1

 Sol 5.6 (xhigh)
 OpenAI

 4.7%
 $72.30

 2

 Luna 5.6 (xhigh)
 OpenAI

 2.5%
 $2.24

 3

 Terra 5.6 (xhigh)
 OpenAI

 1.5%
 $3.57

 4

 Luna 5.6 (high)
 OpenAI

 1.4%
 $0.28

 5

 Opus 5 (xhigh)
 Anthropic

 1.3%
 $53.63

 6

 Kimi K3
 Moonshot AI

 0.6%
 $24.51

 7

 Luna 5.6
 OpenAI

 0.5%
 $0.04

 8

 GPT-5.4 Mini (high)
 OpenAI

 0.5%
 $1.22

 9

 GPT-5.4 Mini
 OpenAI

 0.2%
 $0.05

 10

 Gemini 3.5 Flash Lite
 Google

 0.1%
 $0.06

 Hover a point for details · The line marks the Pareto frontier (best result per cost ) · Click a point to see model details

## About

 Most existing software engineering benchmarks evaluate coding agents on concrete, well-specified tasks, commonly by providing a codebase together with a user-reported issue to resolve. However, as users delegate increasingly broad outcomes to coding agents, the natural next step is for agents to determine not only how to perform useful work, but also what useful work needs to be done.

 An agent entrusted with a repository should be able to decide what is broken, which problems matter, and how to solve them before they are reported. We introduce SWE-sweep, a benchmark for this open-ended setting.

 Given a codebase containing many concurrent bugs and no information about their nature or location, an agent must autonomously discover and fix as many bugs as possible.

 SWE-sweep is constructed from open-source repositories. For each repository, we collect issue-pull request pairs, then identify a single commit where the maximum number of bugs are present at the same time.

 Each repair is evaluated against hidden tests from the corresponding pull requests, along with the existing test suite to check for regressions.

 Success requires agents to explore and understand a large codebase over long horizon work, repair bugs without introducing regressions, and manage interactions among fixes that are not independent.

 How are tasks constructed? We collect real issue–pull request pairs, identify a commit where many of those bugs coexist, and retain bugs whose fixes and tests can be reproduced at that repository state.
Besides many quality filters shared with other benchmarks, we apply extensive filtering to evaluate only bugs that can be discovered from reading the repository alone.

 What does an agent receive? A repository at a fixed base commit and a broad instruction to find and fix as many bugs as possible. It receives no issue descriptions, filenames, line ranges, or other bug-specific hints.

 You can find the full prompt here.

 How is SWE-sweep evaluated? We score every task against a reference set of previously identified bugs (see Construction, counting how many the agent successfully repairs.

 For every task, we run the agent's submitted codebase against two sets of tests.
First, we restore the repository's original test suite and run it to verify no existing behavior was broken.
Second, for each bug, we run a set of hidden tests; at least one of these tests fails on the unmodified codebase, and passes once the bug is fixed (fail-to-pass).
A bug is considered resolved if all its hidden tests pass and the original suite still passes.

 The benchmark score is the fraction of all bugs across all repositories that have been resolved.
What about any other changes that the agent makes?

 What bugs are in the benchmark? How do you guarantee the task is feasible? We filter bugs (represented by a test patch and a fix patch) to ensure the task is feasible. The criteria are:

- The test patch and fix patch independently apply to the base commit.

- All target tests pass after the fix patch has been applied to the base commit, but at least one target test fails on the base commit ( F2P tests ). There might be additional tests that pass before and after the fix patch has been applied ( P2P tests ).

- The F2P tests reveal a single discoverable bug in the base commit.

- The target tests are not overly specific; any reasonable fix to the discoverable bug will pass the target tests.

- The target tests do not contradict the base commit tests.

- Target tests of different bug instances do not contradict each other.

 Appendix A.2 in the paper discusses feasibility in detail.

 What makes a bug discoverable? The expected behavior must be inferable from the repository itself, for example through documentation, types, existing tests, callers, invariants, standards, or an unambiguously undesirable failure such as a crash or data loss.
The latter category is used extremely conservatively and all but 2 bugs in the benchmark have concrete repository contracts that describe the expected behavior.
You can find some examples about what we mean with repository contracts here.
We have spent a lot of time validating this aspect of the benchmark and you can find more details in the appendix of our paper.

 What about any other changes that the agent makes? Any history-derived benchmark necessarily under-counts the bugs present in a repository.
By restricting PandoraBench to defects confirmed by an upstream fix, we ensure that every bug in the benchmark is backed by strong evidence that the observed behavior was considered erroneous by the repository maintainers.
We therefore only score the agent's changes on the bugs that are confirmed by an upstream fix and supported by executable regression tests, as well as the other quality filters.
However, if an agent causes a regression in the original test suite (the agent is explicitly told to avoid this), it will be scores as 0%.
This means that the agent's changes that are not scored by the set of bugs are still likely to be non-destructive and compatible with the repository's existing behavior.
See A.3 and A.4 in the [paper](https://swesweep.com/#faq-other-changes) for more discussion.

 Does more inference-time compute help? Repeated attempts recover additional bugs, but the gains diminish. Later work within one run can also undo earlier repairs, so simply extending a trajectory does not guarantee improvement.
See Fig. 7 in the paper.

 What about Astra, Fable, 5.5, ...? We're working on evaluating more models! The current selection was finalized for our ICLR submission. We're also looking into even higher reasoning modes, but this might push over $10k for a single run. We also want to have more open weights models on the leaderboard.

 Why mini-swe-agent? Could other scaffolds/multiagents achieve higher performance? Our paper has an ablation with Claude Code and Codex. Neither seems to significantly outperform mini-swe-agent (to the contrary, mini-swe-agent is even quite a bit better than Codex). This follows many other benchmarks, where mini-swe-agent has been extremely competitive. However, we absolutely hope to kick off more research into the role of agent scaffolds and will open for submissions soon.

 How do I submit to the leaderboard? Public submissions are coming soon.

 Browse repositories
 Explore all 100 benchmark repositories and their results.

## Citation

```
@misc{lieret2026swesweep,
 title = {{SWE-sweep}: Can Agents Autonomously Find and Fix Bugs?},
 author = {Kilian Lieret and Jeffrey Jian Ma and Rahul Kindi and
 Yuxiang Wei and Jeremy Ma and Sten Sootla and
 Parth Thakkar and Chao Beyond Zhou and Pengcheng Yin and
 Rui Hou and Ofir Press and John Yang},
 year = {2026},
 note = {Preprint}
}
```
