AI LABS
Findings Gallery · Public Wing
EVERY PLAQUE LINKS TO ITS REPRODUCTION COMMAND
CONFIRMED
EXP-025 · 2026-05
The verify loop beats the raw model by +18.5pp — with zero variance.
Same model, same 9 tasks, same kernel judge. Loop: 100% (3/3). Raw: 81.5% mean, 11pp spread. Cost ×2.6.
repro: python -m evals.f3_baseline_eval · report sha b7e2…
CONFIRMED
EXP-022 · 2026-05
"Verified" can mean behavior, not just acceptance.
Programs are attached to a live kernel, real syscalls fired, and map contents read back. A program that loads but miscounts now fails.
repro: python -m evals.f2_ebpf_benchmark_eval · 9/9
NEGATIVE RESULT
EXP-018 · 2026-05
A horizontal "use any app" agent was the wrong first vertical.
No correctness oracle, head-on against frontier labs. Rejected by analysis — and the rejection shaped everything this lab became.
kept on the wall on purpose — failures are findings
IN PROGRESS
F.4 · OPEN PROGRAM
Does the delta grow as models get weaker?
Benchmark growing 9 → 25+ tasks, full behavioral judges, then the ablation matrix: Haiku · Opus · open-weight models. Watch it happen on the Lab Floor.
RETIRED
D.5 · 2026-05
A 96.7% math score, retired as a vanity metric.
Impressive number, wrong oracle, wrong domain. The lab deleted its own best-looking result. The benchmark framework survived; the score did not.
Gallery Index
CONFIRMED ......... 19
IN PROGRESS ....... 3
NEGATIVE RESULT ... 2
RETIRED ........... 1
a claim enters CONFIRMED only when a machine can re-check it — status is promoted by evals, not by people
"We publish what worked, what failed, and what we deleted — because the judge was never us."
FINDINGS PORTAL · PUBLIC · GENERATED FROM docs/research-log — THE LAB'S PAPER TRAIL, RENDERED