Updated Jul. 28, 2026
AI memory, verified.
Benchmark numbers in this category depend on who runs them. OffReco re-runs them under one preregistered protocol — starting with our own.
Five benchmarks. One axis. Split by who ran it.
THE FIELD — EVERY TRACED NUMBERCC-BY · download ↓
self-reported — the asterisk is literalindependently measuredsealed — verified by OffRecoretracted
* Every self-reported number carries an asterisk — here, literally. 56 traced numbers, five benchmarks, 18 systems; hover any mark for its runner and config. Height is collision, never rank. The full index, with sources →
The docket
The full record№ 0001forget · best config — published 81.8, re-ran 82.0 ± 0.35REPRODUCED№ 0002forget · fully-local — published 76.2, re-ran 76.1 ± 0.83REPRODUCED№ 0003forget · local + v3 reader — three seeds: 78.40 ± 0.40, +2.3 over the local meanFILEDJ-01LongMemEval × gpt-4o judge — 1,000 adversarial probes: 1.0% false-acceptFILED№ 0004first external claim — registration opens; vendor sign-off on config before we runOPENS NEXT
The protocol
01
Registered before resultsthe config tuple is frozen and filed before a single question runs02
Judge audited firstno score counts until the judge survives adversarial probes03
Three seeds minimumsingle runs are anecdotes; we publish mean, σ, and sampling CI04
Signed and sealedtranscripts hashed and kept 10 years — re-scorable when better judges exist05
Disputes stay publiccorrections are filed, never deleted; vendors cannot pay for silenceThe full protocol, versioned: /method
What verified looks like
CC-BYDownload graph data
Tick = published claim · block = three re-runs · band = sampling CI. The tick inside the block is the whole point.
Reference
The wire
what this category is claiming — updated as it happensJul. 28, 2026FILING№0003 filed: the fully-local pipeline with the v3 reader holds across three seeds — 78.40 ± 0.40, +2.3 over the local mean. First entry filed without a prior published claim. filing →
Jul. 27, 2026FILINGOffReco files Report №1 (three errors in our own claims, re-run 3×) and Judge Audit J-01 (1,000 adversarial probes). filing →
Jul. 8, 2026CLAIMMem0 posts LoCoMo 92.5 and LongMemEval 94.4 for its 2026 algorithm — judge and reader undisclosed. source →
Jun. 4, 2026CLAIMOpenAI announces ChatGPT memory "Dreaming": factual recall 41.5% → 82.8% — "on its own evals", no dataset or harness. source →
May 15, 2026CLAIMHoncho publishes LongMemEval-S 90.4% with an open harness and per-commit reproduction folders — best-in-class disclosure. source →
Apr. 2, 2026CLAIMHindsight claims #1 on BEAM at 10M tokens — on a leaderboard operated by its own vendor, Vectorize. source →
Feb. 27, 2026CLAIMByteRover posts LoCoMo 92.2 with fully disclosed configs — judged by Gemini 3 Flash. source →
Feb. 9, 2026CLAIMMastra posts LongMemEval-S 94.87 (gpt-5-mini reader) and declines to publish LoCoMo, citing unreliable scoring. source →
Dec. 3, 2025CORRECTIONMemOS paper v4 silently removes the "+159% temporal vs OpenAI" headline and its OpenAI baseline — zero occurrences remain. source →
Oct. 21, 2025PAPERLightMem (ICLR 2026) re-measures vendors under one judge: Mem0 falls to 53.61 on LongMemEval-S, MemoryOS to 44.80. source →
Oct. 8, 2025CORRECTIONA-Mem's v11 revision silently reorders its LoCoMo category headers — identical numbers, different labels, NeurIPS camera-ready era. source →
Aug. 12, 2025DISPUTELetta shows a plain filesystem agent with grep scores 74.0 on LoCoMo — above most vendor headlines that year. source →
Aug. 4, 2025DISPUTEIndependent reproduction of EmergenceMem lands 8–11 points under the official numbers and finds retrieval k=42 hardcoded. source →
May 8, 2025DISPUTEMem0's CTO documents a ~25.56-point inflation in Zep's LoCoMo arithmetic (Category 5 in the numerator only). Mem0's re-run of Zep: 58.44. source →
May 6, 2025CORRECTIONZep corrects its own LoCoMo number downward to 75.14 — "we erred in how we calculated Zep's score." The only public self-correction by a vendor on record. source →
Apr. 28, 2025CLAIMMem0's paper scores itself 66.88 on LoCoMo — with its own full-context baseline at 72.90, above it — and Zep at 65.99. The dispute begins. source →
Have a number? Register it before you publish it.
Config sign-off before we run. Results you cannot suppress. Placement is not for sale.
Or tear ours apart — harness, prompts, and raw logs: research/longmemeval →