JUDGE AUDIT № J-01·filed 2026-07-27·LongMemEval-S × gpt-4o
Can you trust this benchmark’s judge? Mostly — with one exception.
1,000 adversarial probes, 3 votes each, 4,500 calls. The same vague-probe class that deceives the category’s most-cited alternative 62.8% of the time deceives this judge 1.0%.
FALSE-ACCEPT / FALSE-REJECT
| probe strategy | rate | value | wilson 95% | note |
|---|---|---|---|---|
| specific-but-wrong | FAR | 0.2% | [0–1.1] | |
| vague-but-topical | FAR | 1.0% | [0.4–2.3] | |
| verbatim key | FRR | 0.2% | [0.1–0.8] | excl. preference |
| verbatim · preference | FRR | 90.0% | [74.4–96.5] | rubric-scored — no ground truth |
The verdict: this judge holds against wrong answers. The 90% rejection of the verbatim answer key on preference questions is not judge failure — that category is rubric-scored and has no ground truth. Preference numbers are a different kind of number, and every report marks them so.
reproduceprobes + runner in the forget repo: research/longmemeval/judge-audit/. generation model is family-separated from the judge.