Compounding Intelligence

AI changes capability, not certainty

We Are All Iron Man Now

Nicholas Carlini, a researcher at Anthropic, built a 100,000-line C compiler that compiles the Linux kernel by directing 16 AI agents at once, alone, in about two weeks. The key idea is driven by the builder; the harness only gave one person the reach of a team.

By Christopher Hughes

4 August 2026

The 2008 film Iron Man put Tony Stark in a workshop, talking to his personal AI, Jarvis, to prototype the suit. We took it as fun sci-fi, unrealistic but the kind of thing you accept in a Marvel movie. It has turned out to be bizarrely close to how product and engineering teams now prototype. Add 3D printers and you can almost picture an engineer printing the suit itself, once you wave away the insane token budget, the ready supply of gold-titanium alloy, and a printer capable of printing it.

What's interesting is that this often repeated analogy highlights the real shift. Organisations still size ambition by headcount, treating how much one person can build as a roughly fixed ceiling, but that ceiling has moved. Whether a firm notices depends on how it reads what a single builder can accomplish.

Ambition Has Always Been Sized in Headcount

Traditionally, engineering leaders scope what is possible by how many people they can put on the problem. The output of one skilled person has been treated as a constant, so bigger ambitions meant bigger teams, and the size of the team set the size of the plan.

That constant has already started to move. Unfortunately, most organisations haven't yet adjusted to what one builder can now do and are still weighed down by process and procedure.

One Builder, 16 Agents, a Working Compiler

In about two weeks, Nicholas Carlini produced a 100,000-line C compiler in Rust without writing most of the code himself [1]. Carlini, a researcher on Anthropic's Safeguards team, directed 16 AI agents in parallel, an AI agent being a system that takes actions rather than only generating text.

The agents ran as 16 instances of Claude Code across nearly 2,000 sessions, for about $20,000 in API cost [1]. Carlini designed the orchestration and the checks personally and let the machine execute.

The compiler builds Linux 6.9 across x86, ARM, and RISC-V, and passes 99% of the GCC torture test suite, one of the least forgiving benchmarks in the field [1]. This impressive feat was futurist fantasy 10 years ago. Now imagine the same approach applied to the next version of Windows, or a new web browser.

The Idea Stays With the Builder, the Reach Does Not

A harness is not a chatbot. It is the orchestrated rig of agents and tools a builder directs at once, closer to Stark's whole workshop than to a single question put to Jarvis. The builder supplies the goal, the breakdown of the work, and the tests that decide whether each piece is correct. The agents supply the parallel execution.

Steve Faulkner, a director of engineering at Cloudflare, ran the same pattern independently. In under a week he rebuilt 94% of the Next.js 16 API surface by directing Claude through a harness called OpenCode, verified against more than 1,700 automated tests [2].

Faulkner is precise about where the line sits. "Architecture decisions, prioritisation, knowing when the AI was headed down a dead end: that was all me" [2].

This is not a senior reviewer signing off a junior's work, which is a division of labour inside a team. It is one person's reach stretched to the output of a team, with the same person still holding every judgement that matters.

Why This Matters Inside a Bank

A compiler either emits correct machine code or it does not. There is no partial credit and no marks for effort, which is exactly why 99% on the GCC torture suite is an externally verifiable bar rather than a soft claim about feeling more productive.

That unforgiving property is the same one that makes risk models, pricing and analytics tools, reconciliation engines, and controls tooling hard to build and hard to fake. Running agents against that class of work, in production, under a verification regime the builder designed, is a discipline in its own right [3].

The banks are already there. At BBVA, the retail banking legal team built its own assistant on ChatGPT Enterprise, a custom assistant being a configured version of the tool aimed at one task. That assistant now handles a share of the roughly 40,000 client legal enquiries branch managers raise each year, with responses in under 24 hours [4].

BBVA employees built nearly 3,000 such assistants in the tool's first six months, and more than 20,000 across 18 months [4][5]. These were built by the people who do the work, not handed down by a central AI team. That is the whole point. The reach moved to the builder, not to a new function you have to fund and staff.

A Harness Amplifies a Weak Idea as Fast as a Strong One

Reach without judgement scales mistakes at the same rate it scales output. A directed harness will build the wrong thing quickly, thoroughly, and in parallel, if the builder has specified the wrong thing.

Carlini is blunt that the verification was the real work. "The task verifier is nearly perfect, otherwise Claude will solve the wrong problem" [1]. He used the existing GCC compiler as a known-good reference to check the agents' output, because a human reading 100,000 lines was never the control.

Nobody ever credited Jarvis with the Iron Man suit. The idea was Stark's, and the machine only built it faster than a room of engineers could, which is the honest shape of what Carlini and Faulkner did.

The capability changed. The certainty did not. What one builder can now reach is far larger, and what decides whether that reach is worth anything is still the quality of the idea and the rigour of the test that proves it.

Key Takeaways

  • Stop sizing ambition by headcount. The output of one skilled builder directing a harness of agents is now closer to a small team's, so the question is what a person can direct, not how many people you can hire.
  • Treat the idea and the verification as the scarce inputs. Agents supply execution; the builder's goal, decomposition, and tests are what remain hard, and they are where to invest management attention.
  • Put verification at the centre before scaling any builder's reach. A harness amplifies a weak specification as fast as a strong one, so the test that decides correct from incorrect is the control, not an afterthought.

Sources

[1] Carlini, Nicholas. "Building a C compiler with a team of parallel Claudes." Anthropic Engineering Blog, 5 February 2026. https://www.anthropic.com/engineering/building-c-compiler (Vendor-reported figures.)

[2] Faulkner, Steve. "How we rebuilt Next.js with AI in one week." Cloudflare Blog, 24 February 2026. https://blog.cloudflare.com/vinext/ (Self-reported figures.)

[3] Hughes, Christopher. "Intelligent Ops." cgh.dev. https://cgh.dev/thinking/intelligent-ops/

[4] BBVA Newsroom. "BBVA sparks a wave of innovation among its employees with the deployment of ChatGPT Enterprise." BBVA, 2024. https://www.bbva.com/en/innovation/bbva-sparks-a-wave-of-innovation-among-its-employees-with-the-deployment-of-chatgpt-enterprise/ (Vendor-reported figures.)

[5] OpenAI. "BBVA puts AI at the core of banking with OpenAI." OpenAI, 2025. https://openai.com/index/bbva/ (Vendor-reported figures.)