Updated Jul. 28, 2026

AI memory, verified.

Benchmark numbers in this category depend on who runs them. OffReco re-runs them under one preregistered protocol — starting with our own.

Register a claim

Five benchmarks. One axis. Split by who ran it.

THE FIELD — EVERY TRACED NUMBERCC-BY · download ↓
self-reported — the asterisk is literalindependently measuredsealed — verified by OffRecoretracted
LongMemEval-S500 questions · configs differ per dot · n=22full-context 60.6oracle 87.037.2Mem0 53.6 under one judgeforget 82.0 — verified94.87LoCoMo10-conversation public subset · judge swapped per runner · n=24filesystem + grep 74.0Zep 58.4retractedZep 94.7MemoryAgentBenchone academic harness runs everyone — ICLR 2026 · n=6BM25 RAG 41.5Mem0 21.1DMR2023-era retrieval test — saturated by plain context · n=3full conversation 94.498.2BEAM · 10M tokens0–1 score shown ×100 · two vendors both claim #1 · n=6cognee 672030405060708090100

* Every self-reported number carries an asterisk — here, literally. 56 traced numbers, five benchmarks, 18 systems; hover any mark for its runner and config. Height is collision, never rank. The full index, with sources →

The protocol

01
Registered before resultsthe config tuple is frozen and filed before a single question runs
02
Judge audited firstno score counts until the judge survives adversarial probes
03
Three seeds minimumsingle runs are anecdotes; we publish mean, σ, and sampling CI
04
Signed and sealedtranscripts hashed and kept 10 years — re-scorable when better judges exist
05
Disputes stay publiccorrections are filed, never deleted; vendors cannot pay for silence

The full protocol, versioned: /method

What verified looks like

5560657075808590950001forget · best config0002forget · fully-local pipeline0003forget · local + v3 reader

Tick = published claim · block = three re-runs · band = sampling CI. The tick inside the block is the whole point.

The wire

what this category is claiming — updated as it happens
Jul. 28, 2026FILING№0003 filed: the fully-local pipeline with the v3 reader holds across three seeds — 78.40 ± 0.40, +2.3 over the local mean. First entry filed without a prior published claim. filing →
Jul. 27, 2026FILINGOffReco files Report №1 (three errors in our own claims, re-run 3×) and Judge Audit J-01 (1,000 adversarial probes). filing →
Jul. 8, 2026CLAIMMem0 posts LoCoMo 92.5 and LongMemEval 94.4 for its 2026 algorithm — judge and reader undisclosed. source →
Jun. 4, 2026CLAIMOpenAI announces ChatGPT memory "Dreaming": factual recall 41.5% → 82.8% — "on its own evals", no dataset or harness. source →
May 15, 2026CLAIMHoncho publishes LongMemEval-S 90.4% with an open harness and per-commit reproduction folders — best-in-class disclosure. source →
Apr. 2, 2026CLAIMHindsight claims #1 on BEAM at 10M tokens — on a leaderboard operated by its own vendor, Vectorize. source →
Feb. 27, 2026CLAIMByteRover posts LoCoMo 92.2 with fully disclosed configs — judged by Gemini 3 Flash. source →
Feb. 9, 2026CLAIMMastra posts LongMemEval-S 94.87 (gpt-5-mini reader) and declines to publish LoCoMo, citing unreliable scoring. source →
Feb. 1, 2026FUNDINGCognee raises a $7.5M seed led by Pebblebed. source →
Dec. 3, 2025CORRECTIONMemOS paper v4 silently removes the "+159% temporal vs OpenAI" headline and its OpenAI baseline — zero occurrences remain. source →
Oct. 28, 2025FUNDINGMem0 announces $24M (seed + Series A led by Basis Set). source →
Oct. 21, 2025PAPERLightMem (ICLR 2026) re-measures vendors under one judge: Mem0 falls to 53.61 on LongMemEval-S, MemoryOS to 44.80. source →
Oct. 8, 2025CORRECTIONA-Mem's v11 revision silently reorders its LoCoMo category headers — identical numbers, different labels, NeurIPS camera-ready era. source →
Aug. 12, 2025DISPUTELetta shows a plain filesystem agent with grep scores 74.0 on LoCoMo — above most vendor headlines that year. source →
Aug. 4, 2025DISPUTEIndependent reproduction of EmergenceMem lands 8–11 points under the official numbers and finds retrieval k=42 hardcoded. source →
May 8, 2025DISPUTEMem0's CTO documents a ~25.56-point inflation in Zep's LoCoMo arithmetic (Category 5 in the numerator only). Mem0's re-run of Zep: 58.44. source →
May 6, 2025CORRECTIONZep corrects its own LoCoMo number downward to 75.14 — "we erred in how we calculated Zep's score." The only public self-correction by a vendor on record. source →
Apr. 28, 2025CLAIMMem0's paper scores itself 66.88 on LoCoMo — with its own full-context baseline at 72.90, above it — and Zep at 65.99. The dispute begins. source →

Have a number? Register it before you publish it.

Config sign-off before we run. Results you cannot suppress. Placement is not for sale.

How registration works

Or tear ours apart — harness, prompts, and raw logs: research/longmemeval →

Numbers you can act on, in a category that grades its own homework.

Conflict, disclosed

OffReco is operated in the open by Junghun Kim, who also builds forget — the first system on the record. It appears unranked, its errors are in the corrections ledger, and its transcripts are public.

About →