Benchmarks
15 benchmarks documented from primary sources — with the known defects that make their numbers comparable or not. The defects column is the point.
LongMemEvalsource →
ICLR 2025 · UCLA + Tencent
500 q · S ≈115k tok / M ≈1.5M / Oracle · judge: GPT-4o, per-type prompts (>97% human agreement)
KNOWN ISSUES v1 haystacks deprecated by the authors (noisy sessions) — v1 / cleaned / v2 must be stated; v2 is a different generation, scores not comparable.
LoCoMosource →
ACL 2024 · Snap Research
public release: 10 of 50 conversations · 1,540 usable q · judge: paper: F1 — vendors swapped in LLM judges
KNOWN ISSUES 6.4% answer-key errors (audit); gpt-4o-mini judge accepts 62.8% of wrong answers; Cat-5 handling worth ~25 pts; solvable with grep at 74.0.
DMR (MemGPT)source →
UC Berkeley, 2023
MSC-derived consistency QA · judge: GPT-4 + ROUGE-L
KNOWN ISSUES Saturated: full-context alone scores 94.4–98.0. MemGPT's own number differs by arXiv revision (92.5 vs 93.4).
MemoryAgentBenchsource →
ICLR 2026 · UCSD
2,071 q · 103k–1.44M tok, incremental · judge: accuracy / R@5 / LLM judge by split
KNOWN ISSUES Places commercial memory products below naive BM25 RAG (Mem0 21.1 vs 41.5) — sharpest contradiction of vendor claims on record.
MemBenchsource →
ACL 2025 Findings · Renmin + Huawei
~53k q, fully synthetic · judge: multiple-choice — no LLM judge
KNOWN ISSUES At least two unrelated benchmarks share the name — cite by arXiv ID only.
GoodAI LTMsource →
NeurIPS 2024 D&B
11 tasks × 3 reps · spans to 500k tok · judge: scripted + GPT-4-turbo evaluator
KNOWN ISSUES "Living benchmark" — blog-config numbers differ from the NeurIPS tables; don't mix.
MSCsource →
ACL 2022 · FAIR
5 sessions, ~9k episodes · judge: perplexity + human eval
KNOWN ISSUES Not a QA benchmark — survives mainly as raw material for DMR.
PerLTQAsource →
SIGHAN 2024 · CUHK + Huawei
30 characters · 8,593 q · judge: P/R/F1 + model-scored synthesis
KNOWN ISSUES Entirely generated by gpt-3.5-turbo — contamination and artifact risk inherent.
PrefEvalsource →
ICLR 2025 Oral · UCLA + Amazon
3,000 preference-query pairs · to 100k tok · judge: Claude 3 Sonnet, 4 binary checks
KNOWN ISSUES Frontier models near 0.07 zero-shot at 10 turns — preference following collapses within 5 turns.
DialSimsource →
KAIST + SNU
1,300+ sessions · ~350k tok per show · judge: answer matching, no LLM judge
KNOWN ISSUES Title and setup changed across six arXiv versions; TV-script copyright latent; no agent exceeds 60%.
PersonaMemsource →
COLM 2025 · UPenn
~6k q · 32k/128k/1M contexts · judge: multiple-choice — no LLM judge
KNOWN ISSUES Frontier models ≈52% — the rare benchmark that is far from saturated.
ConvoMemsource →
Salesforce AI Research
75,336 q · 2–300 conversation histories · judge: multi-model validation, 2 consecutive correct
KNOWN ISSUES Long-context beats memory-RAG below 150 conversations (70-82 vs 30-45) — challenges the category's premise.
HaluMemsource →
MemTensor + IAAR
~15k memory points · 3.5k q · judge: operation-level scoring (extract/update/QA)
KNOWN ISSUES First benchmark scoring memory operations separately; run by its own vendor's ecosystem — provenance watch.
BEAMsource →
used by cognee, Hindsight, Mem0
100K–10M token corpora · judge: varies by runner
KNOWN ISSUES Dueling SOTA claims: cognee places itself above Hindsight on 100K; Hindsight claims #1 at 10M on its own self-operated leaderboard.
STaRKsource →
Stanford SNAP
MAG / Amazon / Prime variants · judge: hit@k / recall / MRR (retrieval)
KNOWN ISSUES Marketed as accuracy by at least one vendor; leaderboard entries can be on 10% test splits — variant and split must be stated.
Numbers reported on these benchmarks live in the Claims Index and on each system page.