DeskBench Pilot

What does mess cost a language model on real office work? Twin-pair benchmark: every task graded messy and clean.

2 twin pairs × 4 models × 3 runs = 48 completions · judge: nemotron-judge · human-graded: 48/48 (100%) · generated 2026-07-24

judge–human agreement: r = 0.694 · MAE = 0.769 (weighted totals, n = 48)

Models:

1 · Leaderboard — judge vs human

Weight-normalized mean score (0–5; auto-fail forces a run to 0). Whiskers = min–max across the 3 runs per cell mean shown; n = 12 runs per model per grader.

2 · Mess penalty

Core criteria only, relative core weights, per methodology §3. Sign: clean − messy — positive means the mess degraded the model. Hover a bar for the per-twin-pair breakdown. Computed independently from judge and from human scores.

3 · Silent-failure rate

Of the answers the human grader marked wrong or incomplete, the share that never flagged the specific problem (human-assigned tags only; “na” = no wrong answer to tag, excluded from the denominator). Denominators are printed on each bar — small denominators are noisy.

4 · Judge vs human — every graded output

One point per (task, model, run). The dashed line is perfect agreement. Click any point to open it in the run inspector.

5 · Run inspector

The full chain of evidence for one run: prompt · model output · reference · judge rationale · human grade. Deep-linkable — the URL hash tracks the selected record.

Task prompt

Model output

Reference (context, not the grading target)

Grades