JUDGE AUDIT № J-01·filed 2026-07-27·LongMemEval-S × gpt-4o

Can you trust this benchmark’s judge? Mostly — with one exception.

1,000 adversarial probes, 3 votes each, 4,500 calls. The same vague-probe class that deceives the category’s most-cited alternative 62.8% of the time deceives this judge 1.0%.

FALSE-ACCEPT / FALSE-REJECT
probe strategyratevaluewilson 95%note
specific-but-wrongFAR0.2%[01.1]
vague-but-topicalFAR1.0%[0.42.3]
verbatim keyFRR0.2%[0.10.8]excl. preference
verbatim · preferenceFRR90.0%[74.496.5]rubric-scored — no ground truth

The verdict: this judge holds against wrong answers. The 90% rejection of the verbatim answer key on preference questions is not judge failure — that category is rubric-scored and has no ground truth. Preference numbers are a different kind of number, and every report marks them so.

reproduceprobes + runner in the forget repo: research/longmemeval/judge-audit/. generation model is family-separated from the judge.