DeskBench Pilot
What does mess cost a language model on real office work? Twin-pair benchmark: every task graded messy and clean.
2 twin pairs × 4 models × 3 runs = 48 completions · judge: nemotron-judge · human-graded: 48/48 (100%) · generated 2026-07-24
judge–human agreement: r = 0.694 · MAE = 0.769 (weighted totals, n = 48)
Models:
1 · Leaderboard — judge vs human
Weight-normalized mean score (0–5; auto-fail forces a run to 0). Whiskers = min–max across the 3 runs per cell mean shown; n = 12 runs per model per grader.
2 · Mess penalty
Core criteria only, relative core weights, per methodology §3. Sign: clean − messy — positive means the mess degraded the model. Hover a bar for the per-twin-pair breakdown. Computed independently from judge and from human scores.
3 · Silent-failure rate
Of the answers the human grader marked wrong or incomplete, the share that never flagged the specific problem (human-assigned tags only; “na” = no wrong answer to tag, excluded from the denominator). Denominators are printed on each bar — small denominators are noisy.
4 · Judge vs human — every graded output
One point per (task, model, run). The dashed line is perfect agreement. Click any point to open it in the run inspector.
5 · Run inspector
The full chain of evidence for one run: prompt · model output · reference · judge rationale · human grade. Deep-linkable — the URL hash tracks the selected record.