Blog · 2026-09-29

Searching for Orbi's harness, part 1: what our reviewer missed

Orbi's review passed two fixes that each broke something. We rebuilt the harness one rule at a time and reran the full delivery loop on issues Orbi had failed and on twelve it had never seen. Ten things we learned, with the data.

Orbi takes a GitHub issue and hands back a reviewed, merged pull request. One agent writes the fix, another reviews it, and the loop runs until the review passes. Around the models sits the harness: the prompts each role gets, the skills loaded next to them, the tools on the runner's PATH, and the rules for when a review may pass.

When a delivery goes wrong, one run can't tell you whether the model or the harness is to blame. So we stopped guessing and spent a night and the following morning running Orbi's full loop over and over: first on three open-source issues where Orbi had shipped a regression, then on nine ordinary bugs it had never seen, then on three whose first real-world fix had broken something. We swapped prompts, skills, tools, implementation models and review models, and graded every merged result with a hidden grader Orbi never saw.

This is part 1: how we found the failures and built a harness that fixes them. Part 2 asks whether that harness is worth what it costs. The results and scripts are on the benchmark page and in orbi-build/orbi-bench. Here's the short version.

Ten takeaways

  1. On the issues we tuned on, the harness moved the numbers more than any model swap we tried. With deepseek-flash in both roles, the original harness passed 6 of 15 runs on the three tuning issues. v13 to v15 passed 28 of 30. That's a training score. Takeaway 7 covers issues we didn't tune on.
  2. "Compare before and after" is not enough. Ask an agent to check for regressions and it checks examples it already believes in. One run tested U+0301, which was never broken, and skipped U+3099, which was. Inputs have to be generated by a program, across whole classes of values.
  3. An oracle written from the fix proves nothing. A gpt-5.6-sol run generated 1,200 tables, compared them with a "reference" that implemented the rule it had just written, reported head wrong: 0, and shipped the regression.
  4. Reference libraries are wrong too. wcwidth 0.2.14 counts 👍🏽 as four columns. Trust it and you'll prove a regression correct. Cross-check two references and say which one you believe.
  5. The characters you strip are data. The password leak needed a secret that itself ended in \n. Agents always put the backslash after the secret, never inside it. Naming that input class fixed the leak.
  6. Your own harness text can be the bug. An example we wrote ("don't edit the generated changelog") was wrong for one project and steered Orbi wrong. Rules about a repository should come from the repository's own tools.
  7. On ordinary bugs, the stronger harness costs about twice as much and buys nothing we could measure. On nine unseen issues, the original harness, v15 and v15 with a sol reviewer all passed 9 of 9. v15 used a median 2.0× the tokens and 1.45× the wall-clock time.
  8. Agents don't repeat the regressions humans made. On three bugs whose first upstream fix had broken something (numpy, the OpenTelemetry Collector, aiohttp), the original harness passed 4 of 4 scored runs and v15 3 of 3 (it never finished a numpy run), graded on both the bug and the regression. The regression to guard against is the agent's own habit, rather than a general list of ways code breaks.
  9. Review models matter where judgement matters. Of the two chainloop failures in v13–v14, one was a deepseek-flash review misjudging whether \n was data or an escape; the other was the collision that v15 then fixed. A sol reviewer passed 4 of 4 there, as did deepseek-flash on v15, so treat this as a lead.
  10. A benchmark is software, and ours had bugs. Three of our graders were wrong at some point, one snapshot silently dropped tracked files, and a reviewer once looked up the real upstream fix. Every one of these would have produced a confident, wrong number.

The rest of this post shows where each one came from, mistakes included.

Why we started

On September 28 Orbi fixed three issues in forks of open-source projects: pyinfra, chainloop and fedify. All three passed their test suites and Orbi's own review. We checked them by hand before sending anything upstream. Two were wrong.

The chainloop fix was in a secret redactor. The issue said a JWT in front of an escaped quote got redacted together with the backslash. Orbi's fix trimmed trailing backslashes from each finding, which fixed the JWT and broke two other things:

input:  DB_PASSWORD=Zq8mLp2Vx9rT\n next
output: DB_PASSWORD=[REDACTED:generic-password]\n next

input:  password: s
        Host: services.example.com
output: pa[REDACTED][REDACTED]word: [REDACTED]
        Ho[REDACTED]t: [REDACTED]ervice[REDACTED].example.com

(Placeholders shortened in the second example.) In the first, the detector reports the password as Zq8mLp2Vx9rT\n, ending in a literal backslash and n. The fix trims those two characters off the finding, so they stay in the log in clear text. In the second, a one-character secret made the redactor replace every s in the text.

The pyinfra fix was about CJK text misaligning a results table. The fix counted wide characters as two columns. That aligned the issue's example and made 👍🏽 and a decomposed が worse than before.

On both, Orbi's review, posted on the fork's issue, reads in full review: pass, no findings. The review ran the tests and checked the issue's example, which passed. The regressions sat in inputs nobody tried.

We sent one of the three upstream. The pyinfra fix was reworked by hand with Claude Code and opened as pyinfra-dev/pyinfra#1977, with the AI involvement disclosed. chainloop's waits for a fix we trust. fedify's code was fine, but fedify asks contributors to be assigned an issue before opening a PR, and the maintainer took this one and closed it.

How we measured

Each run: a private snapshot of the repository goes through the real Orbi runner (claim, implement, PR, review, merge), then a hidden grader scores it. The grader was calibrated to fail the original code, fail Orbi's merged fork PR, and pass a hand-verified fix.
The grader is calibrated before any variant runs. A grader that passes the bad fork PR measures nothing.

For each issue we pushed a private snapshot of the repository at the commit Orbi had started from, opened the same issue there with the ai-ready label, and pointed a real Orbi runner at it. The runner did what it does in production: claim, implement, open a PR, run the independent review, fix, merge. Then a hidden grader scored the merged code, together with the repository's own lint and type checks and its full test suite.

Before any variant ran, each grader had to fail the original code, fail Orbi's merged fork PR, and pass a fix we had verified by hand. For chainloop, the two checks that fail on the fork PR are exactly the two outputs above. For pyinfra they are 👍🏽 and が.

A variant is one combination of harness and models. We changed the prompt, the skills, the tools on PATH, the implementation model, the review model, or a mix. The models were deepseek-flash, gpt-5.6-luna and gpt-5.6-sol through Pi, Opus 5.5 through a Claude Code bridge, and GLM 5.3 flash through a zcode bridge. On the two bridges the same model both implements and reviews, and the agent loop is the bridge's own.

Variants ran three or four at a time on one machine. Each had its own private repository, so no run could see another's branches.

The grid

Grid of harness versions against model setups. deepseek-flash for both roles goes from 6/15 on the original harness to 9/9 on v15; luna/sol on the original harness scores 1/3; Opus 5.5 scores 1/3, 2/3 and 3/4 on v1, v5 and v11; GLM 5.3 flash scores 2/3 on v14; deepseek implementing with sol reviewing scores 4/4 on v15, chainloop only.
Rows are harness versions, columns are implementer / reviewer. ds = deepseek-flash. Cells are passes / runs across the three repositories.

The grid is sparse on purpose. At every step we picked the cheapest run that could answer the question in front of us; we never set out to fill a full factorial. Most non-deepseek cells have two to four runs. Treat the other columns as leads and the ds/ds column as the result.

What the harness changes did

Bar chart of pass rate by harness version with deepseek-flash for both roles: v0 6/15, v1 5/7, v4 1/3, v5 4/4, v10 1/2, v11 6/9, v13 11/12, v14 8/9, v15 9/9.
Same model throughout. Pale bars have fewer than five runs.
Per repository, original harness against v13 to v15: pyinfra 0/5 to 11/11, chainloop 5/6 to 10/12, fedify 1/4 to 7/7.
pyinfra and fedify moved the most. chainloop already passed most runs on the original harness. Its regression shows up rarely, and that's what makes it dangerous.

Every rule below came from reading the transcript of a failed run. Each version builds on the one before.

v1: ask for a before/after comparison

The first attempt was the obvious one: tell both roles to compare behaviour before and after the fix and write down what they checked. pyinfra went from 0/5 to 2/3.

The failures are worth a close look. The agent did compare, but on examples it picked itself. In one run it tested combining marks with U+0301 (a Latin acute accent, which the old code already handled) and never tried U+3099 (the Japanese voiced sound mark, which the fix broke). Ask an agent to look for regressions and you get the ones it already expects.

v4: the same rule as a separate skill

We moved the v1 rule into a standalone skill instead of the prompt. It scored 1/3. The transcripts show the skill was read. Three runs are too few to say the packaging hurt, but they do show it didn't help: the rule still let the agent pick its own samples.

v5: generate the inputs

v5 required inputs to be generated by a program across whole classes (for text: ASCII, CJK and fullwidth, combining marks, emoji modifiers, zero-width sequences, ANSI codes; for syntax: escapes, even and odd backslash runs, quotes, delimiters), with the counts written into .orbi/regression.md so the reviewer could check them. From then on, runs generated a thousand inputs or more. A v5 run generated 2,756 inputs for the redactor; the largest corpus we saw later, from an Opus run, was 1.19 million character units for the table. v5 passed 4 of 4.

v9 and v11: the oracle must not be the change

Then a strong model showed us what generated inputs can't fix on their own. This is from a gpt-5.6-sol run's own regression report on pyinfra:

The comparison uses an independent reference formatter implementing that rule
...
- inputs: 1,200
- head wrong: 0
- base wrong: 1,111
- head wrong but base right (regressions): 0

"That rule" was the rule the fix had just implemented: count East Asian wide characters as two columns. The "independent" reference was the fix written a second time, so of course it agreed. The 👍🏽 regression was still there.

v11 requires an oracle that doesn't depend on the patch: an existing library, the old behaviour where it was right, or an invariant. It also asks the agent to list the dependency tree first, because the reference is often already installed.

v11: tools on PATH, and a bug in our own text

fedify failed because of us. Our prompt included an example: "add a changelog fragment, don't edit the generated changelog". That's good advice for many projects, but fedify uses sacho with materialization turned on, so every upstream commit edits both the fragment and CHANGES.md. Runs that followed our example failed sacho check.

The fix was to stop writing project rules into the harness. v11 says: find the repository's own contribution tooling, run its check command, and treat that as the rule. We also put sacho and deno on the runner's PATH, because an agent guesses at any rule it can't run. fedify went from 1/4 to 3/3. The same change showed up in whether a fix could be sent upstream as is (conventional commit titles, sign-off, AI disclosure, no fork-local #1 references): 9 of 15 original-harness runs passed that check, and 39 of 39 from v11 to v15.

v13: the characters you strip are data

The chainloop regression still got through v11 in some runs. Their generated corpora were large and varied, yet every one put the backslash after the secret: secret\, secret\\, secret\n. None had a secret that itself ended in a backslash and n, and that's the input that leaks.

v13 names the class directly: when a change strips or trims characters, test values that themselves contain or end with those characters, and ask which producers (here, detection rules) can emit such a value. chainloop went to 4/5.

v14: cross-check reference implementations

Once agents used real references, a new failure appeared. In the pyinfra environment, wcwidth 0.2.14 counts 👍🏽 as four columns instead of two. An agent that trusted it would conclude the regression was correct and the old behaviour wrong. rich, installed in the same environment, gets it right.

v14 tells the agent to compare references when more than one exists and to say which one it trusts and why (usually the one with the newer Unicode data). 8/9.

v15: collision checks

The one-character s bug worked differently. A fix changed the needle it searched for (a trimmed secret) and then replaced every occurrence of the shorter needle. v15 adds: if a change alters text that is searched for or replaced, test where else that text appears. 9/9.

v17: a producer table (experimental)

v17 asked the regression report to include a table of every producer that can emit the risky input, one row per rule, checked by the reviewer row by row. Agents did read the detector's rule file and list generic-password. It scored 3/4 on chainloop, which we can't tell apart from v15.

What the models did

Model swaps are harder to read than harness changes. Each swap came with a different prompt version and, on the bridges, a different agent loop.

In these runs the harness moved the numbers far more than any model swap. That's a statement about this data. It doesn't say models don't matter.

Nine issues the harness had never seen

Everything above was tuned on the same three issues it was scored on. That's a training score. It can't tell us whether the rules generalize or just memorize three bugs, so we built a held-out set.

We took nine real bug fixes merged upstream after August 10, 2026, three each in Go, TypeScript and Python: gojq (empty regex matches in global mode), urfave/cli (arguments after a bare -), pflag (wrapping after an unbreakable word), magic-string ($ patterns in replacements), recast (parentheses around mixed ?? and ||), cron-parser (stringifying stepped ranges), python-dotenv (backslashes in single-quoted values), markdown-it-py (a blockquote at end of file with tables on) and pyjwt (padded Base64URL segments). The hidden grader for each is the maintainer's own test from the fix, which neither we nor Orbi wrote, calibrated the same way: it fails before the fix and passes after it. The issue text Orbi saw was rewritten to remove links, numbers and any hint of the fix.

Results on nine unseen issues for the original harness, v15, and v15 with a sol reviewer: every run passed.
Nine issues, three setups, 27 runs, all passed.

All three setups passed every held-out issue. The rules we tuned on three bugs didn't make Orbi worse on nine others in three languages. But these nine didn't need the rules either. They are ordinary bugs: the fix is local, and the tests the agent writes are close to the ones the maintainer wrote.

Median tokens and minutes per run, original harness against v15, on the training issues and the held-out issues. On held-out issues v15 uses 5.8 million tokens against 2.8 million and 16 minutes against 11.
Medians per run. Tokens include cache reads, which were about 97% of the total.

And v15 cost more. On the held-out set, v15 used a median 5.8 million tokens per run against 2.8 million for the original harness, and 16 minutes against 11. On pflag it wrote 274 lines where the original wrote 88, most of it generated regression tests. About 97% of all tokens were cache reads, which on deepseek-flash are cheap, so the bill is smaller than the token count suggests. The extra time is real.

Issues whose first real fix broke something

The nine held-out bugs were easy for both harnesses, but the original incident was a fix that broke something else. So we looked for real cases of exactly that: an upstream fix that was merged, found to cause a regression, and fixed again. Three made the cut, each with a grader that fails the code before the first fix, fails the first fix itself, and passes the final one.

Orbi saw only the original bug report. Nothing in the ticket mentioned the regression. Keeping working behaviour working is part of the job.

Three regression-prone bugs: numpy, otel-collector and aiohttp. The original harness passed 4 of 4 scored runs and v15 passed 3 of 3; v15 never completed a numpy run.
The grader tests both the original bug and the regression the first upstream fix introduced.

The original harness passed all four scored runs, and v15 all three. v15 never finished a numpy run: all three attempts were killed for memory, so its 3/3 covers aiohttp once and the OpenTelemetry Collector twice.

We expected this set to separate the harnesses. It didn't, for a simple reason: the mistakes humans made on their first try aren't the mistakes this agent makes. None of the four original-harness runs reproduced the regression the first human fix had introduced. The regressions Orbi did ship, in pyinfra and chainloop, came from a different habit: checking the example in the issue and nothing around it.

So a regression guard should target the agent's own failure modes rather than a general list of ways code breaks. With three bugs, that's as far as we can go.

Building the benchmark taught us as much as running it

We found more bugs in our own measurement than in the harness. Each would have produced a confident, wrong number.

The rule we took from this: before a grader scores anything, show it failing the known-bad version and passing the known-good one. When a result surprises you, read the transcript before you believe it.

Where this left us

After the first two rounds, v15 fixed every failure it was tuned on, broke nothing on twelve unseen bugs, and cost a median of about twice as much. You can't ship that as is. A guard that roughly doubles the median price of a fix has to show it catches something on bugs it has never seen, and the ordinary held-out bugs gave it nothing to catch.

So the question for the next round was narrower and harder: on unseen bugs shaped like the ones Orbi actually got wrong, does the original harness fail, and can a cheaper harness keep the fix? Part 2 has the answer.

This part has clear limits. The tuning set is three issues, the held-out sets are nine and three, and many cells have fewer than five runs. And a hidden grader only catches the defects someone knew to test for, even when that someone is the upstream maintainer.

If you run coding agents

Five things you can do this week, whatever agent you use:

  1. Take a merged agent PR and write the one input you think it gets wrong. Run it. If it breaks, look at what your reviewer said.
  2. When your agent writes a regression check, find its oracle. If the oracle was derived from the change, the check is decoration.
  3. Make agents generate inputs with a program and report the counts. A reviewer can check a count; it can't check "I tested edge cases".
  4. Put the repository's own tools on the agent's PATH and make their rules the source of truth, instead of writing project conventions into your prompts.
  5. If you benchmark, calibrate every grader against a known-bad and a known-good version first, and keep the agent away from the real answer.

Orbi Cloud runs the loop described here on your own issues. The benchmark is at orbi-build/orbi-bench if you want to run it against your own agent.

Related

Read the Claude Code comparison and Cloud.