Claims Index
Every published AI-memory benchmark claim we could trace to a primary source — 41 claims, indexed 2026-07-27. A claim here is a citation, not a verdict: nothing gets a verdict until we re-run it. Rows within a benchmark are generally not comparable — the config column is why.
Rules: primary sources only · judge and reader named or marked undisclosed · disputes carry both sides’ URLs · no score-sorted ranking, here or anywhere on this site.
LongMemEval-S
500 questions. The authors deprecated the v1 haystacks (noisy sessions) — every number must state v1 / cleaned / v2. Reader and judge choices swing results by 10+ points; rows here are NOT mutually comparable.
| system | claimed | config | status | source |
|---|---|---|---|---|
| GPT-4o full-context (paper baseline) | 60.6 | full 115k context · GPT-4o judge | THIRD-PARTY | LongMemEval paper (ICLR 2025)2024-10 |
| GPT-4o oracle retrieval (paper ceiling) | 87.0 | evidence-only sessions · GPT-4o judge | THIRD-PARTY | LongMemEval paper2024-10 |
| forget · best config | 81.8 | gpt-4o observer · top-k 84 · reader v3 · GPT-4o judge | VERIFIED | our record №0001 — re-run 3×: 82.0 ± 0.352026-07 |
| forget · fully-local pipeline | 76.2 | qwen2.5-14b observer · GPT-4o judge | VERIFIED | our record №0002 — re-run 3×: 76.1 ± 0.832026-07 |
| Zep | 71.2 | gpt-4o reader · GPT-4o judge | SELF-REPORTED | Zep paper2025-01 |
| Zep | 90.2 | gpt-5.4 reader · gpt-5.4 CoT judge | SELF-REPORTED | getzep.com/research (undated)2026 |
| Zeptimed out under an independent harness | DNF (>2 days) | unified harness · Qwen2.5-7B reader | THIRD-PARTY | Memory in the LLM Era (survey)2026-04 |
| Mem0README warns OSS SDK won't match | 94.4 | undisclosed judge/reader · managed platform | SELF-REPORTED | mem0.ai blog2026-07 |
| Mem0 | 66.4 | GPT-4o-mini · MemOS harness | THIRD-PARTY | MemOS paper v4, Table 42025-12 |
| Mastradeclined to publish LoCoMo, citing unreliable scoring | 94.87 | gpt-5-mini reader · gpt-4o judge (LME prompts) | SELF-REPORTED | mastra.ai research2026-02 |
| MemMachine | 93.0 | gpt-5-mini reader · gpt-4o-mini judge | SELF-REPORTED | MemVerge paper2026-03 |
| Hindsight (Vectorize)"#1 on every dataset" is a self-operated leaderboard | 91.4 / 94.6 | undisclosed reader/judge | SELF-REPORTED | launch PR / self-run leaderboard2025-12 |
| Honcho (Plastic Labs)per-commit repro folders — best-in-class disclosure | 90.4 | gemini-2.5-flash-lite ingest · claude-haiku-4-5 chat · median 5% context | SELF-REPORTED | evals.honcho.dev (open harness)2026-05 |
| Supermemoryheadline "95%" is recall@15, not QA accuracy | 85.4 | GPT-4o judge · competitor rows lifted from Zep paper | SELF-REPORTED | supermemory blog/research2026-05 |
| EmergenceMemindependent reproduction fell 8–11 points short | 79.0–86.0 | hardcoded retrieval k=42 | DISPUTED | vendor claim vs third-party repro 71.22025-08 |
| MemOS | 77.8 | GPT-4o-mini | SELF-REPORTED | MemOS paper v4, Table 42025-12 |
| EMem (UIUC)baseline rows copied from Nemori, not re-measured | 77.9 | gpt-4o-mini reader+judge · 3 runs | THIRD-PARTY | EMem paper2025-11 |
| LightMem (ICLR 2026) | 68.64 | GPT-4o-mini reader+judge | THIRD-PARTY | LightMem paper — also re-measured Mem0 at 53.61, LangMem at 37.202025-10 |
| Nemori | 64.2 / 74.6 | gpt-4o-mini / gpt-4.1-mini reader · gpt-4o-mini judge | THIRD-PARTY | Nemori paper (Mem0·Zep run via commercial APIs)2025-08 |
| anchormind | 45.4 (recall@5 88.3) | harness link 404 | UNTRACED | in-repo benchmark report2026 |
LoCoMo
Every vendor number is on the 10-conversation public subset with an LLM judge swapped in — not the paper's 50-conversation F1 protocol. A third-party audit found 6.4% answer-key errors and a 62.8% judge false-accept rate. Category-5 handling alone is worth ~25 points.
| system | claimed | config | status | source |
|---|---|---|---|---|
| Human performance (paper) | 87.9 | F1, paper protocol | THIRD-PARTY | LoCoMo paper (ACL 2024)2024-02 |
| Mem0 full-context baseline | 72.90 ± 0.19 | GPT-4o-mini · unnamed judge | SELF-REPORTED | Mem0's own paper — above Mem0's own score2025-04 |
| Mem0 / Mem0-graph | 66.88 / 68.44 | GPT-4o-mini · judge unnamed in v1 · adversarial excluded | SELF-REPORTED | Mem0 paper2025-04 |
| Mem0 | 92.5 | undisclosed judge/reader · 2026 algorithm | SELF-REPORTED | mem0.ai blog2026-07 |
| Zep — five numbers, one benchmarkcategory-5 in/out ≈ 25 points; role & timestamp integration contested | 58.44 → 65.99 → 75.14 → 79.09 → 94.7 | Mem0 re-run / Mem0 paper / Zep corrected / MIRIX run / Zep 2026 (gpt-5.4) | DISPUTED | the canonical dispute — both sides' URLs in the ledger2025–2026 |
| Letta — filesystem agent (no memory system)the benchmark is solvable with grep | 74.0 | gpt-4o-mini + file tools + grep | THIRD-PARTY | letta.com — beat Mem0-graph 68.52025-08 |
| MemMachine | 91.69 | gpt-4.1-mini reader · gpt-4o-mini judge | SELF-REPORTED | MemVerge paper2026-03 |
| ByteRover 2.0 | 92.2 | Gemini 3 Flash curation+judge (disclosed) | SELF-REPORTED | byterover blog2026-02 |
| MemU | 92.09 | retrieval-based metric, judge undisclosed | RETRACTED | walked back: "historical… should not be read as a benchmark of the current architecture"2025 |
| MIRIX | 85.38 | gpt-4.1-mini | SELF-REPORTED | MIRIX paper — full-context row scores 87.52, above it2025-07 |
| Memori (GibsonAI) | 81.95 | reader/judge/top-k undisclosed · GPT-4.1-mini cost basis | SELF-REPORTED | memorilabs.ai results page2026 |
| Honcho | 89.9 | gemini-2.5-flash-lite ingest · claude-haiku-4-5 chat | SELF-REPORTED | evals.honcho.dev2026-05 |
| MemOSv1's OpenAI baseline (28.25 temporal) deleted from v4 entirely | 75.80 (paper) / 88.83 (README) | GPT-4o-mini (paper) / "OmniMemEval" undisclosed (README) | SELF-REPORTED | MemOS v4 + repo — the "+159% vs OpenAI" claim was silently removed in v42025-12 |
| Memobase | 75.78 | gpt-4o-style judge · competitor rows copied from Mem0 paper | SELF-REPORTED | in-repo benchmark README2025 |
| LangMem — same system, two runnersa 20-point gap with no vendor claim in between | 58.10 vs 78.05 | Mem0's run vs Memori's run · configs undisclosed | DISPUTED | LangChain itself has published no numbers2025–2026 |
| Memvid | "+35% SOTA" | vs "the industry average" (undefined) · judge unnamed | UNTRACED | README2026 |
DMR · MemoryAgentBench · STaRK · platform evals
Adjacent benchmarks where the same patterns repeat: saturated baselines, version drift, metric conflation, absent methodology.
| system | claimed | config | status | source |
|---|---|---|---|---|
| MemGPT (DMR)cite by version; full-conversation baseline already 94.4–98.0 (Zep) | 92.5 or 93.4 | GPT-4/-Turbo — number differs by arXiv revision | SELF-REPORTED | MemGPT paper v1/v22023–2024 |
| Zep (DMR) | 94.8 | gpt-4-turbo · GPT judge | SELF-REPORTED | Zep paper — argues DMR is saturated2025-01 |
| Mem0 (MemoryAgentBench)sharpest published contradiction of vendor claims | 21.1 — below BM25 RAG at 41.5 | independent academic harness | THIRD-PARTY | MemoryAgentBench (ICLR 2026)2025-07 |
| Papr (STaRK)retrieval hit-rate marketed as accuracy; split disclosed only off-site | hit@5 92.04 | STaRK-MAG test-0.1 (10% split) · GPT-4o-mini rerank | SELF-REPORTED | marketed as "#1 on Stanford's STaRK, 91% accuracy"2026-01 |
| OpenAI — ChatGPT "Dreaming"0 of 6 platforms published a reproducible memory benchmark at launch | 41.5 → 82.8 | "on its own evals" — no dataset, no harness | UNTRACED | the only platform memory number ever attached (via secondary coverage)2026-06 |
Why so many DISPUTED rows?
Because the benchmarks don’t prescribe configs, the same product scores 58.44 or 94.7 depending on who runs it. That gap is the reason this index exists — and the reason claims graduate to the record only through a registered, signed re-run.