NEWS

AI Agents Refactor 300,000 Lines in Three Weeks, But Experts Split on What It Proves

A CodeScene case study shows agents refactoring 300,000 lines of C for $4,000 in three weeks. Practitioners agree on the numbers, but disagree on what they actually prove.

A Fighting Game as a Proving Ground

CodeScene published a case study in which AI agents refactored a 300,000-line C codebase over three weeks, with a token cost of approximately $4,000, according to a report by InfoQ. The codebase chosen was the open source decompilation of Street Fighter III: 3rd Strike.

The measured numbers: 2,903 commits, 726 files changed, 252,055 lines modified. Code Health, CodeScene's proprietary metric, rose from 5.6 to 10.0. Adam Tornhill, the company's founder and author of "Your Code as a Crime Scene," described the result as the first time, in three decades working with large systems, that he saw what he called superhuman performance from agents at scale.

For those who refactor legacy systems day to day, the notable point isn't "agents refactor fast": that was already expected. It's the scale. 300,000 lines is on the order of magnitude of a real monolith, not a workshop example.

The Two Mechanisms Behind the Result

Two components underpinned the work, according to the report:

  • A deterministic quality signal: the CodeHealth MCP Server, which gives agents an objective score to optimize for and to judge whether a transformation helped or hurt the code.
  • A correctness harness: a replay-trace system that compares rollback state hashes frame by frame, checking whether the game's behavior changed after each alteration.

Without the second item, the first would be dangerous: an agent can raise Code Health and break behavior at the same time. It's the combination of the two that made it possible to trust the process without a line-by-line review of each of the 726 files.

The Recipes the AI Discovered on Its Own

The most unusual result wasn't the final score, but the discovery process. Instead of applying a fixed catalog of refactorings, the agents built up their own playbook over the course of the project, ending with 22 recipes and 82 supporting notes.

Some are textbook classics (Extract Function, Guard Clauses). Others emerged from patterns specific to the code of a 1990s fighting game:

  • Shared Index Range: captures repeated loops that differ only in the start and end of the range.
  • Action Parameter: handles duplicated control structures that differ mainly in which function they call.
  • Uniform Step Table: converts heterogeneous calls into table-driven dispatch.

Failed attempts were also recorded, including transformations that worsened Code Health. That detail matters: the playbook isn't a list of wins, it's a trial-and-error log the team kept.

The Model Chosen Made a Measurable Difference

The team chose Claude Opus for the bulk of the work. According to the report, Claude Code with Opus was significantly better than Codex with Sol at capturing and documenting the emerging patterns. Files often stalled when smaller models ran the task, as if they hit a local optimum they couldn't escape.

For those choosing a model for an agentic refactoring pipeline today, this suggests that the savings from using a cheaper model can come at a high cost in result quality, not just speed.

The Rift Among Practitioners

The reaction on LinkedIn, according to the InfoQ report, split not over whether the result happened, but over what it proves.

Who defendsArgument
Paolo PerroneTesting against a game replay is a higher bar than a green test suite
Mats Iremark, CTO of Omda ResponseReported a comparable experience using CodeScene MCP with current agents
Who questionsArgument
Konrad Otrębski, tech leadAsked for proof on a famous open source project, with a merge into the main master
Tracy Bannon, software architectQuestioned the "perfect result" framing used in the announcement
Denis BaltorPoints out that DRY is about knowledge duplication, not identical lines
Asko NõmmQuestions how much of the result measures the model versus the harness

I think the real true test of AI capability here would be to offer this courtesy of refactoring to some famous open source project, say Grafana. The definition of done is ofc merging it to master.

Konrad Otrębski, tech lead and consultant

Daniel Webb, CTO of NeoSee and one of the two engineers who did the work, replied to Otrębski confirming that the result was merged into the main branch of a fork, through 54 pull requests, rather than as a single merge.

most refactor claims i've read rest on a green test suite, which only tells you the tests survived. replaying traces against a fighting game sets a much higher bar.

Paolo Perrone

The Blind Spot: No Oracle, No Harness

The most uncomfortable argument for those who work with legacy systems didn't come from a critic, but from the case itself. The replay-trace worked because a decompiled game allows for deterministic frame-by-frame comparison. Most legacy systems have no equivalent: no oracle capable of saying, with the same precision, whether behavior changed after a refactoring.

That's the same reason refactoring these systems is risky. Tornhill acknowledges this when he writes that automated tests and equivalence checks are absolutely essential safeguards, which puts the weight right back on exactly what unhealthy codebases tend to lack.

The team's own selection process illustrates the point. Webb said they considered a UK government (Gov.UK) maritime licensing codebase he had worked on before, but it was too healthy to serve as research material, and they chose the game partly because they play it and are, therefore, users of it.

Questions the Authors Themselves Left Open

Marc Bouvier asked whether non-functional behavior had improved, since framerate, memory usage, and input latency matter in a game. Webb replied that a performance expert is being brought in to assess this, meaning there's no answer yet.

On whether the harness captured subtle frame-timing regressions, Webb was direct in saying there may be no definitive answer, because the harness ran as a pre-commit hook and some failures were fixed without the team ever observing them.

if you are not familiar with the code and no longer review every line, how big can a diff be?

Daniel Webb, CTO of NeoSee

Webb made no claim about the ideal diff size in this scenario; he simply left the question open, which is revealing: even the person who did the work acknowledges that line-by-line review stopped being practical at this volume.

Where This Research Goes Next

The three-week result produced two functionally equivalent versions of the same system: one with Code Health 5.6, the other with 10.0. It serves as base material for a study with Lund University, in which students will implement features in both versions using frontier models, comparing cost and quality.

Two numbers that circulated alongside the case belong to this future work, not to the experiment itself: CodeScene projects an approximately 70% reduction in AI-induced defects and roughly 45% less token waste from the healthier version, both extrapolated from the company's earlier research, not measured in this case. What was actually measured were the $4,000 and the three weeks.

What Changes for Those Who Refactor Real Code

Before replicating the idea in a production system, three practical takeaways from the case are worth noting:

  • The harness matters more than the model. Without a behavioral oracle (tests, replay, equivalence checking), turning agents loose on a large-scale refactoring is betting against the very Code Health they're optimizing for.
  • Small models stall. The report indicates that smaller models got stuck in local optima; if the goal is real scale, the token savings may not be worth it.
  • "Merged" needs context. The work was integrated into a fork through 54 pull requests, not into a third-party open source project in production. The bar Otrębski proposed, refactoring something like Grafana and merging it into the official master, hasn't been met by anyone publicly yet.

Translated from the Brazilian Portuguese original · Read the original