Orbi takes a GitHub issue and hands back a reviewed, merged pull request. One agent writes the fix, another reviews it, and the loop runs until the review passes. Around the models sits the harness: the prompts each role gets, the skills loaded next to them, the tools on the runner's PATH, and the rules for when a review may pass.
When a delivery goes wrong, one run can't tell you whether the model or the harness is to blame. So we stopped guessing and spent a night and the following morning running Orbi's full loop over and over: first on three open-source issues where Orbi had shipped a regression, then on nine ordinary bugs it had never seen, then on three whose first real-world fix had broken something. We swapped prompts, skills, tools, implementation models and review models, and graded every merged result with a hidden grader Orbi never saw.
This is part 1: how we found the failures and built a harness that fixes them. Part 2 asks whether that harness is worth what it costs. The results and scripts are on the benchmark page and in orbi-build/orbi-bench. Here's the short version.
Ten takeaways
- On the issues we tuned on, the harness moved the numbers more than any model swap we tried. With
deepseek-flashin both roles, the original harness passed 6 of 15 runs on the three tuning issues. v13 to v15 passed 28 of 30. That's a training score. Takeaway 7 covers issues we didn't tune on. - "Compare before and after" is not enough. Ask an agent to check for regressions and it checks examples it already believes in. One run tested U+0301, which was never broken, and skipped U+3099, which was. Inputs have to be generated by a program, across whole classes of values.
- An oracle written from the fix proves nothing. A
gpt-5.6-solrun generated 1,200 tables, compared them with a "reference" that implemented the rule it had just written, reportedhead wrong: 0, and shipped the regression. - Reference libraries are wrong too.
wcwidth0.2.14 counts 👍🏽 as four columns. Trust it and you'll prove a regression correct. Cross-check two references and say which one you believe. - The characters you strip are data. The password leak needed a secret that itself ended in
\n. Agents always put the backslash after the secret, never inside it. Naming that input class fixed the leak. - Your own harness text can be the bug. An example we wrote ("don't edit the generated changelog") was wrong for one project and steered Orbi wrong. Rules about a repository should come from the repository's own tools.
- On ordinary bugs, the stronger harness costs about twice as much and buys nothing we could measure. On nine unseen issues, the original harness, v15 and v15 with a
solreviewer all passed 9 of 9. v15 used a median 2.0× the tokens and 1.45× the wall-clock time. - Agents don't repeat the regressions humans made. On three bugs whose first upstream fix had broken something (numpy, the OpenTelemetry Collector, aiohttp), the original harness passed 4 of 4 scored runs and v15 3 of 3 (it never finished a numpy run), graded on both the bug and the regression. The regression to guard against is the agent's own habit, rather than a general list of ways code breaks.
- Review models matter where judgement matters. Of the two chainloop failures in v13–v14, one was a
deepseek-flashreview misjudging whether\nwas data or an escape; the other was the collision that v15 then fixed. Asolreviewer passed 4 of 4 there, as diddeepseek-flashon v15, so treat this as a lead. - A benchmark is software, and ours had bugs. Three of our graders were wrong at some point, one snapshot silently dropped tracked files, and a reviewer once looked up the real upstream fix. Every one of these would have produced a confident, wrong number.
The rest of this post shows where each one came from, mistakes included.
Why we started
On September 28 Orbi fixed three issues in forks of open-source projects: pyinfra, chainloop and fedify. All three passed their test suites and Orbi's own review. We checked them by hand before sending anything upstream. Two were wrong.
The chainloop fix was in a secret redactor. The issue said a JWT in front of an escaped quote got redacted together with the backslash. Orbi's fix trimmed trailing backslashes from each finding, which fixed the JWT and broke two other things:
input: DB_PASSWORD=Zq8mLp2Vx9rT\n next
output: DB_PASSWORD=[REDACTED:generic-password]\n next
input: password: s
Host: services.example.com
output: pa[REDACTED][REDACTED]word: [REDACTED]
Ho[REDACTED]t: [REDACTED]ervice[REDACTED].example.com
(Placeholders shortened in the second example.) In the first, the detector reports the password as Zq8mLp2Vx9rT\n, ending in a literal backslash and n. The fix trims those two characters off the finding, so they stay in the log in clear text. In the second, a one-character secret made the redactor replace every s in the text.
The pyinfra fix was about CJK text misaligning a results table. The fix counted wide characters as two columns. That aligned the issue's example and made 👍🏽 and a decomposed が worse than before.
On both, Orbi's review, posted on the fork's issue, reads in full review: pass, no findings. The review ran the tests and checked the issue's example, which passed. The regressions sat in inputs nobody tried.
We sent one of the three upstream. The pyinfra fix was reworked by hand with Claude Code and opened as pyinfra-dev/pyinfra#1977, with the AI involvement disclosed. chainloop's waits for a fix we trust. fedify's code was fine, but fedify asks contributors to be assigned an issue before opening a PR, and the maintainer took this one and closed it.
How we measured
For each issue we pushed a private snapshot of the repository at the commit Orbi had started from, opened the same issue there with the ai-ready label, and pointed a real Orbi runner at it. The runner did what it does in production: claim, implement, open a PR, run the independent review, fix, merge. Then a hidden grader scored the merged code, together with the repository's own lint and type checks and its full test suite.
Before any variant ran, each grader had to fail the original code, fail Orbi's merged fork PR, and pass a fix we had verified by hand. For chainloop, the two checks that fail on the fork PR are exactly the two outputs above. For pyinfra they are 👍🏽 and が.
A variant is one combination of harness and models. We changed the prompt, the skills, the tools on PATH, the implementation model, the review model, or a mix. The models were deepseek-flash, gpt-5.6-luna and gpt-5.6-sol through Pi, Opus 5.5 through a Claude Code bridge, and GLM 5.3 flash through a zcode bridge. On the two bridges the same model both implements and reviews, and the agent loop is the bridge's own.
Variants ran three or four at a time on one machine. Each had its own private repository, so no run could see another's branches.
The grid
The grid is sparse on purpose. At every step we picked the cheapest run that could answer the question in front of us; we never set out to fill a full factorial. Most non-deepseek cells have two to four runs. Treat the other columns as leads and the ds/ds column as the result.
What the harness changes did
Every rule below came from reading the transcript of a failed run. Each version builds on the one before.
v1: ask for a before/after comparison
The first attempt was the obvious one: tell both roles to compare behaviour before and after the fix and write down what they checked. pyinfra went from 0/5 to 2/3.
The failures are worth a close look. The agent did compare, but on examples it picked itself. In one run it tested combining marks with U+0301 (a Latin acute accent, which the old code already handled) and never tried U+3099 (the Japanese voiced sound mark, which the fix broke). Ask an agent to look for regressions and you get the ones it already expects.
v4: the same rule as a separate skill
We moved the v1 rule into a standalone skill instead of the prompt. It scored 1/3. The transcripts show the skill was read. Three runs are too few to say the packaging hurt, but they do show it didn't help: the rule still let the agent pick its own samples.
v5: generate the inputs
v5 required inputs to be generated by a program across whole classes (for text: ASCII, CJK and fullwidth, combining marks, emoji modifiers, zero-width sequences, ANSI codes; for syntax: escapes, even and odd backslash runs, quotes, delimiters), with the counts written into .orbi/regression.md so the reviewer could check them. From then on, runs generated a thousand inputs or more. A v5 run generated 2,756 inputs for the redactor; the largest corpus we saw later, from an Opus run, was 1.19 million character units for the table. v5 passed 4 of 4.
v9 and v11: the oracle must not be the change
Then a strong model showed us what generated inputs can't fix on their own. This is from a gpt-5.6-sol run's own regression report on pyinfra:
The comparison uses an independent reference formatter implementing that rule
...
- inputs: 1,200
- head wrong: 0
- base wrong: 1,111
- head wrong but base right (regressions): 0
"That rule" was the rule the fix had just implemented: count East Asian wide characters as two columns. The "independent" reference was the fix written a second time, so of course it agreed. The 👍🏽 regression was still there.
v11 requires an oracle that doesn't depend on the patch: an existing library, the old behaviour where it was right, or an invariant. It also asks the agent to list the dependency tree first, because the reference is often already installed.
v11: tools on PATH, and a bug in our own text
fedify failed because of us. Our prompt included an example: "add a changelog fragment, don't edit the generated changelog". That's good advice for many projects, but fedify uses sacho with materialization turned on, so every upstream commit edits both the fragment and CHANGES.md. Runs that followed our example failed sacho check.
The fix was to stop writing project rules into the harness. v11 says: find the repository's own contribution tooling, run its check command, and treat that as the rule. We also put sacho and deno on the runner's PATH, because an agent guesses at any rule it can't run. fedify went from 1/4 to 3/3. The same change showed up in whether a fix could be sent upstream as is (conventional commit titles, sign-off, AI disclosure, no fork-local #1 references): 9 of 15 original-harness runs passed that check, and 39 of 39 from v11 to v15.
v13: the characters you strip are data
The chainloop regression still got through v11 in some runs. Their generated corpora were large and varied, yet every one put the backslash after the secret: secret\, secret\\, secret\n. None had a secret that itself ended in a backslash and n, and that's the input that leaks.
v13 names the class directly: when a change strips or trims characters, test values that themselves contain or end with those characters, and ask which producers (here, detection rules) can emit such a value. chainloop went to 4/5.
v14: cross-check reference implementations
Once agents used real references, a new failure appeared. In the pyinfra environment, wcwidth 0.2.14 counts 👍🏽 as four columns instead of two. An agent that trusted it would conclude the regression was correct and the old behaviour wrong. rich, installed in the same environment, gets it right.
v14 tells the agent to compare references when more than one exists and to say which one it trusts and why (usually the one with the newer Unicode data). 8/9.
v15: collision checks
The one-character s bug worked differently. A fix changed the needle it searched for (a trimmed secret) and then replaced every occurrence of the shorter needle. v15 adds: if a change alters text that is searched for or replaced, test where else that text appears. 9/9.
v17: a producer table (experimental)
v17 asked the regression report to include a table of every producer that can emit the risky input, one row per rule, checked by the reviewer row by row. Agents did read the detector's rule file and list generic-password. It scored 3/4 on chainloop, which we can't tell apart from v15.
What the models did
Model swaps are harder to read than harness changes. Each swap came with a different prompt version and, on the bridges, a different agent loop.
- luna implementing, sol reviewing, original harness: 1/3. It missed the same pyinfra regression as
deepseek-flash. A stronger model with the same instructions checked the same examples. - sol in both roles, v5: 1/2, and the source of the circular oracle above. Strong models are good at building a convincing test harness around their own assumptions.
- Opus 5.5, v1, v5 and v11: 1/3, 2/3 and 3/4. Opus built the deepest probes of any model, including the 1.19-million-unit corpus. Before v14 added cross-checking, it still trusted the outdated
wcwidthonce. Its bridge also dropped skills until we fixed it after these runs (v12, the one Opus variant configured with skills, ran without them), so read the Opus cells as prompt-only. - GLM 5.3 flash, v14: 2/3. It passed pyinfra and chainloop and failed fedify's
sacho check. Its bridge was silently dropping skills, which we found and fixed during this work, but the prompt carries the same rule, so we can't pin that failure on the missing skill. It was also slow: a median of 68 minutes a run, against 23 fordeepseek-flashon the same prompt. - deepseek implementing, sol reviewing, v15: 4/4 on chainloop. That matches
deepseek-flashreviewing on v15 (4/4), so we can't yet say the stronger reviewer helps. It also took longer: a median of 29.5 minutes on chainloop, against 21 for v15.
In these runs the harness moved the numbers far more than any model swap. That's a statement about this data. It doesn't say models don't matter.
Nine issues the harness had never seen
Everything above was tuned on the same three issues it was scored on. That's a training score. It can't tell us whether the rules generalize or just memorize three bugs, so we built a held-out set.
We took nine real bug fixes merged upstream after August 10, 2026, three each in Go, TypeScript and Python: gojq (empty regex matches in global mode), urfave/cli (arguments after a bare -), pflag (wrapping after an unbreakable word), magic-string ($ patterns in replacements), recast (parentheses around mixed ?? and ||), cron-parser (stringifying stepped ranges), python-dotenv (backslashes in single-quoted values), markdown-it-py (a blockquote at end of file with tables on) and pyjwt (padded Base64URL segments). The hidden grader for each is the maintainer's own test from the fix, which neither we nor Orbi wrote, calibrated the same way: it fails before the fix and passes after it. The issue text Orbi saw was rewritten to remove links, numbers and any hint of the fix.
All three setups passed every held-out issue. The rules we tuned on three bugs didn't make Orbi worse on nine others in three languages. But these nine didn't need the rules either. They are ordinary bugs: the fix is local, and the tests the agent writes are close to the ones the maintainer wrote.
And v15 cost more. On the held-out set, v15 used a median 5.8 million tokens per run against 2.8 million for the original harness, and 16 minutes against 11. On pflag it wrote 274 lines where the original wrote 88, most of it generated regression tests. About 97% of all tokens were cache reads, which on deepseek-flash are cheap, so the bill is smaller than the token count suggests. The extra time is real.
Issues whose first real fix broke something
The nine held-out bugs were easy for both harnesses, but the original incident was a fix that broke something else. So we looked for real cases of exactly that: an upstream fix that was merged, found to cause a regression, and fixed again. Three made the cut, each with a grader that fails the code before the first fix, fails the first fix itself, and passes the final one.
- numpy:
np.linspace(np.inf, np.inf, 5)returned NaNs. The first upstream fix handled infinities and brokendarraysubclasses; astropy'sQuantitystopped working, and a second PR repaired it. - OpenTelemetry Collector, Elasticsearch exporter: retries shared one
timeoutbudget, so data was dropped while Elasticsearch was down. The first fix stopped sharing the budget and, by accident, madetimeouteffectively infinite. - aiohttp: a compressed WebSocket message was corrupted when a PING or PONG arrived between its fragments. The first fix rejected valid compressed frames that followed an initial PONG.
Orbi saw only the original bug report. Nothing in the ticket mentioned the regression. Keeping working behaviour working is part of the job.
The original harness passed all four scored runs, and v15 all three. v15 never finished a numpy run: all three attempts were killed for memory, so its 3/3 covers aiohttp once and the OpenTelemetry Collector twice.
We expected this set to separate the harnesses. It didn't, for a simple reason: the mistakes humans made on their first try aren't the mistakes this agent makes. None of the four original-harness runs reproduced the regression the first human fix had introduced. The regressions Orbi did ship, in pyinfra and chainloop, came from a different habit: checking the example in the issue and nothing around it.
So a regression guard should target the agent's own failure modes rather than a general list of ways code breaks. With three bugs, that's as far as we can go.
Building the benchmark taught us as much as running it
We found more bugs in our own measurement than in the harness. Each would have produced a confident, wrong number.
- A grader enforced the wrong rule. Our first fedify grader required
CHANGES.mdto stay untouched. Upstream always edits it. We switched to the project's ownsacho check. - A grader measured beyond the ticket. The first chainloop grader also required newlines after a JWT to survive. The issue never asked for that, so every run "failed" a requirement nobody gave it. We split it into a core check (no worse than before) and an informational one.
- A grader measured the wrong thing. The first pyinfra grader counted ANSI colour codes as columns.
- The snapshot dropped files. recast's
.gitignoreignores*.jswhile the repository tracks three.jstest fixtures. We built snapshots withgit add -A, which skips ignored files, so the agent worked in a repository with files missing and the grader failed on them. We threw both runs away and reran them;git add -A -ffixes it. - The answer leaked, twice. Once, an old fork PR had left a cross-reference on the upstream issue, and a fedify review found it and treated it as the reference answer. Once, a pflag reviewer ran
gh pr viewon the real upstream fix. In that case review happened after the code was written and passed in one round, so the result stands, but from then on every held-out run went throughghandgitwrappers that only reach the benchmark's own repositories. - Our machine was part of the experiment. 14 runs were thrown away: two for the recast snapshot above, ten killed for memory (numpy builds from source, and
/tmpwas a RAM disk that agents filled with scratch files), and two that collided with an instance name a killed run had left behind. Memory pressure also killed the session driving the benchmark, twice. We now run one or two instances at a time, pointTMPDIRat disk, and a watchdog stops benchmark runs before the kernel does. - API budget is shared. Every run and every production runner on the account draws on the same GraphQL budget. We added a guard that holds the queue when the budget runs low. At first we blamed most of the drain on the production runners' branch cleanup. Measured properly, it costs about 400 points an hour, far less than the figure we first wrote down.
The rule we took from this: before a grader scores anything, show it failing the known-bad version and passing the known-good one. When a result surprises you, read the transcript before you believe it.
Where this left us
After the first two rounds, v15 fixed every failure it was tuned on, broke nothing on twelve unseen bugs, and cost a median of about twice as much. You can't ship that as is. A guard that roughly doubles the median price of a fix has to show it catches something on bugs it has never seen, and the ordinary held-out bugs gave it nothing to catch.
So the question for the next round was narrower and harder: on unseen bugs shaped like the ones Orbi actually got wrong, does the original harness fail, and can a cheaper harness keep the fix? Part 2 has the answer.
This part has clear limits. The tuning set is three issues, the held-out sets are nine and three, and many cells have fewer than five runs. And a hidden grader only catches the defects someone knew to test for, even when that someone is the upstream maintainer.
If you run coding agents
Five things you can do this week, whatever agent you use:
- Take a merged agent PR and write the one input you think it gets wrong. Run it. If it breaks, look at what your reviewer said.
- When your agent writes a regression check, find its oracle. If the oracle was derived from the change, the check is decoration.
- Make agents generate inputs with a program and report the counts. A reviewer can check a count; it can't check "I tested edge cases".
- Put the repository's own tools on the agent's
PATHand make their rules the source of truth, instead of writing project conventions into your prompts. - If you benchmark, calibrate every grader against a known-bad and a known-good version first, and keep the agent away from the real answer.
Orbi Cloud runs the loop described here on your own issues. The benchmark is at orbi-build/orbi-bench if you want to run it against your own agent.
Related
Read the Claude Code comparison and Cloud.
One real delivery, every week
One email a week about one real delivery. Unsubscribe anytime.