
2,481 Cross-Reviews: Claude Code and Codex Check Each Other
For two months we made Claude Code and Codex review each other's work before anything shipped. 2,481 reviews later, here is what the numbers say, and the one rule we had to add.
Claude Code finishes the work, and Codex reads it before anything ships. When Codex writes, Claude reads.
A bank works the same way. Nobody approves their own loan. Somebody writes it up, somebody else signs, and nobody takes it as an insult. We call our version Dual Brain: two AI agents working as a pair on one job, one doing it, the other reviewing it. It's built into Sno Station, and we've been living inside it since the middle of July.
Every review leaves one line in a log file on our build machine. A while ago I got curious and counted them.
The count
From July 19 to September 21, the two agents completed 2,481 reviews of each other. They worked on 63 of those days, so about 39 reviews a day. They read 3.2 million lines across 9,452 files.
A review takes about two minutes. The median is 141 seconds, and nine out of ten finish in under six minutes. That number matters more than it looks. A human review takes a day, mostly spent waiting for the human. At two minutes you stop asking whether something deserves a review. Everything gets one.
I thought most of the work would pass.
Of the reviews that returned a verdict, 431 said approve. 1,953 said needs attention. So four times out of five, the second agent sent the work back.
Four out of five. These aren't sloppy first drafts from a tired intern. The model had already run its tests and said the work was done.
Code versus plans
Then I counted the code reviews and the plan reviews separately.
Code passed 27 percent of the time. 398 approvals out of 1,475 reviews. Not great, but you can live with it.
Plans passed 3.6 percent of the time. 33 out of 907.
I had to look at that twice. A plan is a document. It says what we're going to build and why. You'd think a document is easier to get right than code, because nothing has to compile. It's the other way around. Code gets corrected by reality every time you run it. A plan can say the database has a column it doesn't have, and the sentence will sit there looking perfectly reasonable until somebody builds on it.
So now we give the reviewer the facts along with the plan. Before a plan review, the writing agent runs the commands, counts the rows, dumps the schema, and attaches the raw output. Then the reviewer has something to check the claims against. Without that, it's a closed-book exam. The reviewer can find contradictions inside the document, but it can't see that the document and the world disagree.
Most of the worst findings come from exactly that gap. Over the same period the reviewers flagged 4,277 high-severity problems across 1,395 reviews. I haven't checked every one, and not all of them are real. But the ones that were real tended to be the same kind: something stated with confidence that nobody had actually looked at.
The part that went wrong
If I stopped here this would be a nice story. I'd like to think all those reviews made the work better.
It's not that simple. The thing is, a reviewer always finds something.
Look at the later rounds. After the writer fixed the findings and asked again, the work passed 21 percent of the time. In round one it was 17 percent. Fixing every finding barely moved the odds. One piece of work went fifteen rounds.
That's not the writer failing fifteen times. It's the reviewer doing what reviewers do. You ask a capable model to find problems, it finds problems. If the ordinary path is clean, it goes looking in the corners. Two requests arriving in the same millisecond. A cancel halfway through a write. Cases that are possible, sort of, in the way that being hit by a meteor is possible.
And the writer, being agreeable, fixes each one by adding something.
Earlier this month we had a small bug. The fix went through six review rounds. Each round the reviewer raised a new concern, and each round the writer added a defense. A ceiling on token counts. A lock per session. A separate table to track what had been served. An alias map. By the end the patch was a small fortress.
None of those mechanisms ever ran in production. And underneath all of them, the fix itself still had a one-line identity bug that nobody saw, because it was buried under the fortifications. It was like turning up the gain to hear a quiet instrument and getting mostly hiss.
The rule we added
So we wrote a rule, and it's the most useful thing to come out of the whole experiment.
A finding is not an order to add code.
The review now comes back in two piles. The first holds the problems a real user would hit on an ordinary day, where the smallest fix doesn't add anything new. Those get fixed. The second pile holds everything else: the exotic cases, the things that would need a new lock or a new retry or a new table. Those get written down and reported. They don't turn into code unless a person decides they should.
We also made the reviewer look for what's extra before it looks for what's missing. A lock nobody needed. A flag with one caller. A test that just restates the code. Deleting counts as a finding now.
And there's a cap. Three rounds for code. One for tests. After that the argument goes to a human, which is where it probably belonged.
What it's worth
I can't give you a clean number for the value. I'd like to. I don't have a control group, a second company that shipped the same work without review.
What I have is this. Four out of five times, work that one model called finished had something a different model wanted changed. Some of that is noise. But a single agent reviewing its own work almost never sends it back, because it already believes what it wrote. The second opinion has to come from somewhere else. Another model, trained by other people, that wasn't in the room when the first one decided everything was fine.
Two heads are better than one. It turns out that's true when neither head is human.
Sno Station is open source, and the cross-review is part of it. It's at sno.ai.
Written by Sno AI Team
Contributing writer at Sno.ai, sharing insights about AI, productivity, and knowledge management.
Related Articles
Comments
Comments coming soon. Configure Giscus at giscus.app


