Anatomy of a bug-fixing agent
Every night, Serge, our CI agent, looks for real failures in Transformers, reproduces them on GPUs, investigates the cause, writes a patch, verifies it, and opens a PR for maintainers to review when everything checks out.
Over the last 80 days, 29 of those fixes have landed in Transformers.
This post explains how the system works, what happens before Serge is allowed to open a PR, and what we've learned so far.
Integration tests
In Transformers, integration tests run real model checkpoints end-to-end on real GPUs and verify their outputs.
They're expensive and slow, so running all of them on every commit would be wasteful. Instead, we run the relevant ones on model PRs and run a broader suite every night.
These tests are one of the main ways we verify that models still behave correctly against real checkpoints.
Why use an agent?
The Transformers codebase is huge, and we get a lot of contributions. New models are added frequently, frameworks change underneath us, and keeping every test green all the time is difficult.
Some failures are harmless numerical differences. Others expose real bugs after a framework or dependency changes. Either way, there is always a backlog of small failures that need someone to look at them.
The fixes themselves are often simple. The annoying part is the loop around them.
A maintainer might inspect a failure, make a tiny change, submit a GPU job, wait for it to start, inspect the result, adjust the patch, and run it again. The code change might take a few minutes, but the debugging process can repeatedly pull someone back into the same task.
That's exactly the kind of work we wanted to automate: repetitive, slow to orchestrate, and easy to verify.
The agent loop: 6 steps to produce a patch
To produce a patch, Serge goes through six steps and stops as early as possible when something doesn't look right.
The agent:
- filters out noise and selects failures worth investigating;
- reproduces the failure;
- checks that nobody is already working on it;
- generates a patch and runs quality checks;
- verifies that the patch actually fixes the failure;
- opens a PR for maintainers.
Each task runs in its own short-lived Kubernetes pod. The pod has no GitHub credentials and no cluster API token, and network access is restricted to GitHub, the Hugging Face inference router, and Serge's history index.
Serge can only push branches under serge/ and open pull requests. Nothing reaches main unless a maintainer reviews and merges it.
Step 1 - Pick the right test failures
We run almost 200k tests in CI, and some amount of flakiness is unavoidable.
Ideally, every source of flakiness would be fixed, whether that's a network call that needs better retry logic or a CI machine running out of disk space. In practice, though, we need to separate persistent failures from noise.
Serge looks back over several days and focuses on tests that have failed consistently throughout the past week.
If a test has been failing every night, it's much more likely to represent something worth investigating.
Step 2 - Reproduce the failure
Before trying to fix anything, Serge has to prove that the failure is reproducible.
It spins up a fresh GPU runner and executes the test several times. If the failure doesn't reproduce reliably, the task stops there.
There's no point asking an LLM to fix a failure we can't consistently observe.
Step 3 - Check whether someone is already fixing it
Transformers has a large contributor community, so it's common for several people to notice the same problem around the same time.
Before Serge starts writing code, it checks whether the failure is already being addressed.
We also use relore to give Serge context from previous issues, PRs, review comments, Git history, and maintainer discussions.
You can think of relore as repository memory. It lets Serge retrieve the reasoning behind earlier decisions instead of looking only at the current checkout.
For example, while investigating a GPTNeoXJapanese regression, Serge was about to implement a fix that someone else was already working on. relore found an open, approved PR addressing the same issue, so Serge stopped instead of opening a duplicate.
Step 4 - Generate a patch
This is where the LLM starts writing code.
The model gets a local checkout and iterates on the patch in a multi-turn loop. It can use the same repository tools a developer would use, including ruff, ty, and transformers-mlinter.
We also query relore again while generating the patch.
That matters because a fix that looks correct in isolation may conflict with an earlier repository-wide decision. Past review comments and maintainer discussions can provide context that isn't visible from the code alone.
This part is still experimental. You can read more about the project in the relore blog post.
If the patch survives those checks, Serge moves on to the important part: proving that it works.
Step 5 - Verify the patch
Serge spins up a GPU runner again and tests both versions of the code.
The targeted test runs five times on the unpatched tree and five times on the patched tree. This verifies both that the original failure is real and that the patch fixes it consistently.
This step also exposes an interesting failure mode.
An agent optimizes for the task we give it, which isn't always the same thing as solving the underlying problem.
Suppose a test expects the wrong output. In that case, changing the expected value may be exactly the right fix. But changing an assertion is also an easy way to make a test green without fixing anything.
So when Serge only changes expected values, we treat the patch with extra suspicion.
The test still has to pass repeatedly on GPU, but we don't yet have a perfect automatic way to distinguish a legitimate expectation update from reward hacking. We've rejected patches for this reason, and improving this check is one of the things we want to work on next.
There are also cases where Serge simply can't reach a verdict.
A runner may exit unexpectedly. A patch may fail to apply. GPU verification may time out.
Those are treated as infrastructure failures, not as evidence that the patch is good or bad. Serge opens no PR, marks the group as task failed on the tracking issue, and retries it the following night if the test is still failing.
For GPU timeouts, we keep the branch around without opening a PR so a human can inspect it. The patch may still be valid; we just don't have enough evidence to publish it automatically.
Step 6 - Open a PR
If a patch makes it through all five previous steps, Serge opens a PR and pings the maintainers.
Several failure groups run in parallel every night, so Serge also creates a tracking issue for each nightly run.
That issue gives maintainers one place to see which failures were dispatched, which ones produced a PR, which ended without a safe fix, and which failed because of infrastructure or verification problems.
For example, the September 23 run tracked all dispatched failure groups and linked successful ones to their resulting PRs. See issue #49034.
When Serge changes actual model code rather than a test, it often finds a real issue and produces a patch that's at least useful enough for maintainers to understand what went wrong.
Sometimes the patch is good enough to merge unchanged. Sometimes a maintainer tweaks it or writes a different fix. Either way, the investigation has already been done.
A real example
The PvtV2ModelIntegrationTest::test_inference_model test failed in all seven nightly runs in its window with the same assertion:
torch.Size([1, 256, 7, 7]) != torch.Size([1, 50, 512])
Serge reproduced the failure on GPU and started investigating.
The model spent 36 turns and 56 seconds of LLM time working out the cause. End-to-end wall-clock time was about 28 minutes.
The expected shape (1, 50, 512) had been copied from the PVT v1 test, but PVT v2 is hierarchical and returns 4D feature maps. For pvt_v2_b0, the correct shape is (1, 256, 7, 7).
Serge fixed the shape, updated the slice index, and added the correct expected values.
This is a useful example of why changes to expected values aren't automatically wrong. In this case, the test itself was incorrect.
Serge then ran the test five times on the unpatched tree and five times on the patched tree. The patch passed verification, a PR was opened, and it was approved and merged unchanged.
The fix itself wasn't especially difficult. What Serge removed was the need for a maintainer to keep coming back to an asynchronous GPU debugging loop.
See PR #49067.
What 80 days of Serge taught us
The first few weeks weren't great.
We had to rework prompts and eventually built relore because the agent needed more context about the history of the repository before it could make good decisions.
Over the last few weeks, things have improved noticeably.
One useful metric is simply how many of Serge's patches maintainers consider good enough to merge. So far, 29 fixes have landed in Transformers over roughly 80 days, or about 2.5 per week.
Doing the same work manually, even with AI assistance, still requires a developer to orchestrate the process: inspect results, launch runs, come back when they finish, retry failures, and decide what to do next.
We estimate that at roughly one to two hours of developer attention per fix, or around 43 hours across the 29 fixes that landed.
The important difference is that Serge doesn't behave like an assistant that a maintainer has to supervise continuously. It runs unattended and brings a human into the loop only when it has a verified patch worth reviewing. But only a fraction of dispatched failures ever reach the LLM, and only a fraction of those eventually become PRs.
That's intentional.
The pipeline is designed to throw away weak or unverifiable candidates early rather than maximize the number of patches it produces.
Over the past two weeks, the funnel looked like this:
We dispatched 117 failure groups.
Of those, 29 never reached the LLM because the GPU reproduction step found the test already passing at the base commit, sometimes because the problem had been fixed in the meantime.
Another two failed on their first request to the inference endpoint.
That left 86 groups that entered a real LLM session.
Of those 86:
- 24 produced a verified PR event, corresponding to 19 distinct PRs;
- 28 ended without a safe patch;
- 18 produced a patch that failed GPU verification;
- 5 timed out during verification;
- 11 failed mid-session because of patch-apply, runner, or output-parsing errors.
Inference Costs
So far, Serge has been using Kimi K-2.7-Code.
Most sessions cost a little under $2 in inference, although a few long runs pull the average closer to $3. Across the 86 LLM sessions from the last two weeks, we spent roughly $250.
A lot of that spend does not turn into a PR. Of those 86 sessions, 24 produced a verified PR event. Using the average session cost, that is around $72 of inference for the successful ones. The rest went into patches that failed verification, sessions that ended without a safe fix, timeouts, or other failures.
The reproduction step helps here. The 29 failure groups that turned out to be passing on the base commit never reached the LLM, so they cost us nothing in inference.
With the numbers we have today, inference comes out to roughly $14 per distinct PR and $43 per merged PR. That is still cheap compared with the developer time we estimate for doing the same debugging manually, but there is an obvious opportunity to improve it: we are still spending quite a bit on sessions whose result we throw away.
The model itself is one place to look. For instance, Qwen3.8-27B is smaller and cheaper, and has a higher DeepSWE 1.1 benchmark score than Kimi (42.2 vs 31). Our initial benchmarks confirm it's a better model for Serge use cases so far while being twice cheaper.
Conclusion
Serge is still an experiment, and there is plenty left to improve. But it has already shown us where this kind of agent is useful.
The sweet spot isn’t replacing maintainers. It’s automating repetitive, verifiable work and escalating only when human review is needed.
The resource we're really trying to save is maintainer attention.