Benchmark · updated 2026-09-30

How we test Orbi's harness

On eight bugs Orbi had never seen, where the Issue shows one narrow example of a wider problem, the original harness failed 4 of 12 runs and an always-on regression guard passed 9 of 10. On ordinary bugs the guard buys nothing and costs about twice the tokens, so it belongs behind a per-repository switch, off by default. A classifier that ran the guard only on text-handling changes saved about 10% and did worse. Strengthening only the reviewer, or enriching the ticket first, didn't close the gap either.

RESULTS · deepseek-flash implements · passes / runs

Narrow-example bugs8 unseen, maintainer tests

8/12original harness
9/10always-on guard (v15)

Bugs Orbi got wrongthe 3 tuning issues

6/15original harness
9/9always-on guard (v15)

Ordinary bugs9 unseen, median per run

9/9always-on guard (v15)
5.8Mtokens (original 2.8M)

Harness research, as it happens

Harnesses have to change as models do. Each study here reruns Orbi's real delivery loop, publishes every number and script, and picks up where the last one left off.

Where the original harness slips

Eight real fixes merged after August 10, 2026, in small Python, Go and JavaScript libraries. Each Issue shows one or two examples, while the maintainer's tests cover a wider class. A fix written only for the example passes it and fails those tests. Orbi saw only the Issue's own example.

Eight bugs whose Issue shows a narrow example. The original harness passed 8 of 12 runs, failing dateparser and picomatch twice each; the always-on guard passed 9 of 10; the gated guard passed 7 of 8.
dateparser and picomatch were run three times on the original harness and failed twice each.

Cheaper designs we tried

All tested on the three issues Orbi originally got wrong. Strengthening only the reviewer passed 1 of 3: even with the full regression hunt in its prompt, the reviewer approved both regressions. Across all 183 runs, every run that failed the hidden grader had passed Orbi's review. Having a cheap model reproduce the bug and enrich the ticket first passed 2 of 3. Running the guard only for text-handling changes passed 3 of 3, but on ordinary bugs it saved only about 10% of the always-on guard's tokens, and it failed picomatch where always-on passed. We're dropping the gate for a per-repository switch.

Five harnesses on the three issues Orbi got wrong: original 6 of 15, reviewer only 1 of 3, enrich ticket 2 of 3, gated guard 3 of 3, always-on guard 9 of 9.
Median million tokens per run under each bar. Pale bars have fewer than five runs.

Nine bugs Orbi had never seen

Real fixes merged upstream after August 10, 2026, three each in Go, TypeScript and Python. Each hidden grader is the test the maintainer wrote with the fix. The ticket Orbi saw was rewritten with no links, issue numbers or hints.

Results on nine unseen issues for the original harness, v15, and v15 with a sol reviewer: every run passed.
Nine issues, three setups, 27 runs, all passed.
Median tokens and minutes per run, original harness against v15. On unseen issues v15 uses 5.8 million tokens against 2.8 million and 16 minutes against 11.
What the stronger harness costs extra. Tokens include cache reads, about 97% of the total.

Bugs whose first real fix broke something

Three real bugs whose first upstream fix was merged and then found to break something: numpy (the fix broke ndarray subclasses), the OpenTelemetry Collector (it made timeout effectively infinite) and aiohttp (it rejected valid compressed frames after a PONG). The ticket described only the original bug; the grader tests both. Both harnesses passed every run that finished; v15 never finished a numpy run (all three attempts were killed for memory). The mistakes humans made on their first try weren't the ones this agent makes.

Three regression-prone bugs: the original harness passed 4 of 4 runs and v15 passed 3 of 3; v15 never completed a numpy run.

Where the stronger harness came from

Three issues where Orbi had merged a regression (pyinfra, chainloop, fedify). We changed the harness one rule at a time and ran the loop 108 times on them across every harness version. These numbers are a training score: the harness was tuned and scored on the same three issues.

Pass rate by harness version with deepseek-flash for both roles on the three tuning issues: v0 6/15 rising to v15 9/9.
Same model throughout, three tuning issues. Pale bars have fewer than five runs.
Grid of harness versions against model setups on the tuning issues; most non-deepseek cells have two to four runs.
Harness versions against implementer / reviewer models on the tuning issues. Read the other columns as leads.

How each run is scored

Each run: a private snapshot of the repository goes through the real Orbi runner, then a hidden grader scores it. The grader is calibrated first.

What these numbers don't show

Rerun it

Everything is in orbi-build/orbi-bench: the tasks, graders, every harness variant, isolation wrappers, scoring scripts and per-run results. Upstream tests aren't copied in; tasks/heldout/fetch_hidden.sh fetches them from each fix commit. Run one or two instances at a time, since a single task can take several gigabytes of memory.

Every rule, every failure, every bug in our own measuring

The posts walk through each harness change and the failed run behind it, what the models did, and the mistakes we made building this benchmark.