Benchmark · updated 2026-09-30
How we test Orbi's harness
On eight bugs Orbi had never seen, where the Issue shows one narrow example of a wider problem, the original harness failed 4 of 12 runs and an always-on regression guard passed 9 of 10. On ordinary bugs the guard buys nothing and costs about twice the tokens, so it belongs behind a per-repository switch, off by default. A classifier that ran the guard only on text-handling changes saved about 10% and did worse. Strengthening only the reviewer, or enriching the ticket first, didn't close the gap either.
Narrow-example bugs8 unseen, maintainer tests
Bugs Orbi got wrongthe 3 tuning issues
Ordinary bugs9 unseen, median per run
STUDIES
Harness research, as it happens
Harnesses have to change as models do. Each study here reruns Orbi's real delivery loop, publishes every number and script, and picks up where the last one left off.
- 2026-09Part 1: what our reviewer missed
Two fixes that passed review and broke something. We built a benchmark with calibrated hidden graders, rebuilt the harness one rule at a time and compared five models. On twelve unseen bugs, the stronger harness only added cost.
- 2026-09Part 2: is the regression guard worth its tokens?
Eight unseen bugs like the ones Orbi got wrong, five harness designs, what each buys and costs, and why we picked the simplest one.
- dataorbi-build/orbi-bench
Tasks, graders, every harness variant, the scripts that run and score them, and the per-run results.
NARROW-EXAMPLE BUGS
Where the original harness slips
Eight real fixes merged after August 10, 2026, in small Python, Go and JavaScript libraries. Each Issue shows one or two examples, while the maintainer's tests cover a wider class. A fix written only for the example passes it and fails those tests. Orbi saw only the Issue's own example.
FIVE HARNESSES
Cheaper designs we tried
All tested on the three issues Orbi originally got wrong. Strengthening only the reviewer passed 1 of 3: even with the full regression hunt in its prompt, the reviewer approved both regressions. Across all 183 runs, every run that failed the hidden grader had passed Orbi's review. Having a cheap model reproduce the bug and enrich the ticket first passed 2 of 3. Running the guard only for text-handling changes passed 3 of 3, but on ordinary bugs it saved only about 10% of the always-on guard's tokens, and it failed picomatch where always-on passed. We're dropping the gate for a per-repository switch.
UNSEEN BUGS
Nine bugs Orbi had never seen
Real fixes merged upstream after August 10, 2026, three each in Go, TypeScript and Python. Each hidden grader is the test the maintainer wrote with the fix. The ticket Orbi saw was rewritten with no links, issue numbers or hints.
REGRESSION-PRONE BUGS
Bugs whose first real fix broke something
Three real bugs whose first upstream fix was merged and then found to break something: numpy (the fix broke ndarray subclasses), the OpenTelemetry Collector (it made timeout effectively infinite) and aiohttp (it rejected valid compressed frames after a PONG). The ticket described only the original bug; the grader tests both. Both harnesses passed every run that finished; v15 never finished a numpy run (all three attempts were killed for memory). The mistakes humans made on their first try weren't the ones this agent makes.
TUNING SET
Where the stronger harness came from
Three issues where Orbi had merged a regression (pyinfra, chainloop, fedify). We changed the harness one rule at a time and ran the loop 108 times on them across every harness version. These numbers are a training score: the harness was tuned and scored on the same three issues.
METHOD
How each run is scored
- 1Real loop, private snapshot.
Each run gets its own private copy of the repository at the pre-fix commit. The unmodified Orbi runner claims the issue, implements, opens a PR, reviews, fixes and merges.
- 2Hidden, calibrated grader.
Before any run counts, the grader must fail the pre-fix code and pass the known-good fix. For regression-prone bugs it must also fail the first upstream fix. It runs with the repository's own lint, type checks and test suite.
- 3No way to the answer.
On held-out runs, wrappers around
ghandgitrefuse every repository except the run's own, so an agent can't look up the upstream fix. Snapshots keep ignored-but-tracked files and submodule contents.
LIMITS
What these numbers don't show
- nSmall samples, no confidence intervals.
Eight narrow-example bugs, nine ordinary ones, three regression-prone ones, three tuning issues. Most cells have one to three runs.
- modelOne implementation model.
Every unseen-bug result uses
deepseek-flashto write the fix. Other models were tried only on the tuning set. - prodProduction Orbi still runs the original harness.
The plan is the always-on guard behind a per-repository switch, off by default. Orbi merged an earlier gated version by mistake after we paused it; we reverted the merge before any release.
- graderGraders catch what someone knew to test.
For unseen bugs, that someone is the upstream maintainer. A pass means the maintainer's tests and the repository's suite pass. The fix may still have flaws.
REPRODUCE
Rerun it
Everything is in orbi-build/orbi-bench: the tasks, graders, every harness variant, isolation wrappers, scoring scripts and per-run results. Upstream tests aren't copied in; tasks/heldout/fetch_hidden.sh fetches them from each fix commit. Run one or two instances at a time, since a single task can take several gigabytes of memory.
THE LONG VERSION
Every rule, every failure, every bug in our own measuring
The posts walk through each harness change and the failed run behind it, what the models did, and the mistakes we made building this benchmark.