Claims Index

Every published AI-memory benchmark claim we could trace to a primary source — 41 claims, indexed 2026-07-27. A claim here is a citation, not a verdict: nothing gets a verdict until we re-run it. Rows within a benchmark are generally not comparable — the config column is why.

Rules: primary sources only · judge and reader named or marked undisclosed · disputes carry both sides’ URLs · no score-sorted ranking, here or anywhere on this site.

LongMemEval-S

500 questions. The authors deprecated the v1 haystacks (noisy sessions) — every number must state v1 / cleaned / v2. Reader and judge choices swing results by 10+ points; rows here are NOT mutually comparable.

systemclaimedconfigstatussource
GPT-4o full-context (paper baseline)60.6full 115k context · GPT-4o judgeTHIRD-PARTYLongMemEval paper (ICLR 2025)2024-10
GPT-4o oracle retrieval (paper ceiling)87.0evidence-only sessions · GPT-4o judgeTHIRD-PARTYLongMemEval paper2024-10
forget · best config81.8gpt-4o observer · top-k 84 · reader v3 · GPT-4o judgeVERIFIEDour record №0001 — re-run 3×: 82.0 ± 0.352026-07
forget · fully-local pipeline76.2qwen2.5-14b observer · GPT-4o judgeVERIFIEDour record №0002 — re-run 3×: 76.1 ± 0.832026-07
Zep71.2gpt-4o reader · GPT-4o judgeSELF-REPORTEDZep paper2025-01
Zep90.2gpt-5.4 reader · gpt-5.4 CoT judgeSELF-REPORTEDgetzep.com/research (undated)2026
Zeptimed out under an independent harnessDNF (>2 days)unified harness · Qwen2.5-7B readerTHIRD-PARTYMemory in the LLM Era (survey)2026-04
Mem0README warns OSS SDK won't match94.4undisclosed judge/reader · managed platformSELF-REPORTEDmem0.ai blog2026-07
Mem066.4GPT-4o-mini · MemOS harnessTHIRD-PARTYMemOS paper v4, Table 42025-12
Mastradeclined to publish LoCoMo, citing unreliable scoring94.87gpt-5-mini reader · gpt-4o judge (LME prompts)SELF-REPORTEDmastra.ai research2026-02
MemMachine93.0gpt-5-mini reader · gpt-4o-mini judgeSELF-REPORTEDMemVerge paper2026-03
Hindsight (Vectorize)"#1 on every dataset" is a self-operated leaderboard91.4 / 94.6undisclosed reader/judgeSELF-REPORTEDlaunch PR / self-run leaderboard2025-12
Honcho (Plastic Labs)per-commit repro folders — best-in-class disclosure90.4gemini-2.5-flash-lite ingest · claude-haiku-4-5 chat · median 5% contextSELF-REPORTEDevals.honcho.dev (open harness)2026-05
Supermemoryheadline "95%" is recall@15, not QA accuracy85.4GPT-4o judge · competitor rows lifted from Zep paperSELF-REPORTEDsupermemory blog/research2026-05
EmergenceMemindependent reproduction fell 8–11 points short79.0–86.0hardcoded retrieval k=42DISPUTEDvendor claim vs third-party repro 71.22025-08
MemOS77.8GPT-4o-miniSELF-REPORTEDMemOS paper v4, Table 42025-12
EMem (UIUC)baseline rows copied from Nemori, not re-measured77.9gpt-4o-mini reader+judge · 3 runsTHIRD-PARTYEMem paper2025-11
LightMem (ICLR 2026)68.64GPT-4o-mini reader+judgeTHIRD-PARTYLightMem paper — also re-measured Mem0 at 53.61, LangMem at 37.202025-10
Nemori64.2 / 74.6gpt-4o-mini / gpt-4.1-mini reader · gpt-4o-mini judgeTHIRD-PARTYNemori paper (Mem0·Zep run via commercial APIs)2025-08
anchormind45.4 (recall@5 88.3)harness link 404UNTRACEDin-repo benchmark report2026

LoCoMo

Every vendor number is on the 10-conversation public subset with an LLM judge swapped in — not the paper's 50-conversation F1 protocol. A third-party audit found 6.4% answer-key errors and a 62.8% judge false-accept rate. Category-5 handling alone is worth ~25 points.

systemclaimedconfigstatussource
Human performance (paper)87.9F1, paper protocolTHIRD-PARTYLoCoMo paper (ACL 2024)2024-02
Mem0 full-context baseline72.90 ± 0.19GPT-4o-mini · unnamed judgeSELF-REPORTEDMem0's own paper — above Mem0's own score2025-04
Mem0 / Mem0-graph66.88 / 68.44GPT-4o-mini · judge unnamed in v1 · adversarial excludedSELF-REPORTEDMem0 paper2025-04
Mem092.5undisclosed judge/reader · 2026 algorithmSELF-REPORTEDmem0.ai blog2026-07
Zep — five numbers, one benchmarkcategory-5 in/out ≈ 25 points; role & timestamp integration contested58.44 → 65.99 → 75.14 → 79.09 → 94.7Mem0 re-run / Mem0 paper / Zep corrected / MIRIX run / Zep 2026 (gpt-5.4)DISPUTEDthe canonical dispute — both sides' URLs in the ledger2025–2026
Letta — filesystem agent (no memory system)the benchmark is solvable with grep74.0gpt-4o-mini + file tools + grepTHIRD-PARTYletta.com — beat Mem0-graph 68.52025-08
MemMachine91.69gpt-4.1-mini reader · gpt-4o-mini judgeSELF-REPORTEDMemVerge paper2026-03
ByteRover 2.092.2Gemini 3 Flash curation+judge (disclosed)SELF-REPORTEDbyterover blog2026-02
MemU92.09retrieval-based metric, judge undisclosedRETRACTEDwalked back: "historical… should not be read as a benchmark of the current architecture"2025
MIRIX85.38gpt-4.1-miniSELF-REPORTEDMIRIX paper — full-context row scores 87.52, above it2025-07
Memori (GibsonAI)81.95reader/judge/top-k undisclosed · GPT-4.1-mini cost basisSELF-REPORTEDmemorilabs.ai results page2026
Honcho89.9gemini-2.5-flash-lite ingest · claude-haiku-4-5 chatSELF-REPORTEDevals.honcho.dev2026-05
MemOSv1's OpenAI baseline (28.25 temporal) deleted from v4 entirely75.80 (paper) / 88.83 (README)GPT-4o-mini (paper) / "OmniMemEval" undisclosed (README)SELF-REPORTEDMemOS v4 + repo — the "+159% vs OpenAI" claim was silently removed in v42025-12
Memobase75.78gpt-4o-style judge · competitor rows copied from Mem0 paperSELF-REPORTEDin-repo benchmark README2025
LangMem — same system, two runnersa 20-point gap with no vendor claim in between58.10 vs 78.05Mem0's run vs Memori's run · configs undisclosedDISPUTEDLangChain itself has published no numbers2025–2026
Memvid"+35% SOTA"vs "the industry average" (undefined) · judge unnamedUNTRACEDREADME2026

DMR · MemoryAgentBench · STaRK · platform evals

Adjacent benchmarks where the same patterns repeat: saturated baselines, version drift, metric conflation, absent methodology.

systemclaimedconfigstatussource
MemGPT (DMR)cite by version; full-conversation baseline already 94.4–98.0 (Zep)92.5 or 93.4GPT-4/-Turbo — number differs by arXiv revisionSELF-REPORTEDMemGPT paper v1/v22023–2024
Zep (DMR)94.8gpt-4-turbo · GPT judgeSELF-REPORTEDZep paper — argues DMR is saturated2025-01
Mem0 (MemoryAgentBench)sharpest published contradiction of vendor claims21.1 — below BM25 RAG at 41.5independent academic harnessTHIRD-PARTYMemoryAgentBench (ICLR 2026)2025-07
Papr (STaRK)retrieval hit-rate marketed as accuracy; split disclosed only off-sitehit@5 92.04STaRK-MAG test-0.1 (10% split) · GPT-4o-mini rerankSELF-REPORTEDmarketed as "#1 on Stanford's STaRK, 91% accuracy"2026-01
OpenAI — ChatGPT "Dreaming"0 of 6 platforms published a reproducible memory benchmark at launch41.5 → 82.8"on its own evals" — no dataset, no harnessUNTRACEDthe only platform memory number ever attached (via secondary coverage)2026-06

Why so many DISPUTED rows?

Because the benchmarks don’t prescribe configs, the same product scores 58.44 or 94.7 depending on who runs it. That gap is the reason this index exists — and the reason claims graduate to the record only through a registered, signed re-run.

Read the method