A horizontal "use any app" agent was the wrong first vertical.
No correctness oracle, head-on against frontier labs. Rejected by analysis — and the rejection shaped everything this lab became.
kept on the wall on purpose — failures are findings
IN PROGRESS
F.4 · OPEN PROGRAM
Does the delta grow as models get weaker?
Benchmark growing 9 → 25+ tasks, full behavioral judges, then the ablation matrix: Haiku · Opus · open-weight models. Watch it happen on the Lab Floor.
RETIRED
D.5 · 2026-05
A 96.7% math score, retired as a vanity metric.
Impressive number, wrong oracle, wrong domain. The lab deleted its own best-looking result. The benchmark framework survived; the score did not.
Gallery Index
CONFIRMED ......... 19
IN PROGRESS ....... 3
NEGATIVE RESULT ... 2
RETIRED ........... 1
a claim enters CONFIRMED only when a machine can re-check it — status is promoted by evals, not by people
"We publish what worked, what failed, and what we deleted — because the judge was never us."
FINDINGS PORTAL · PUBLIC · GENERATED FROM docs/research-log — THE LAB'S PAPER TRAIL, RENDERED