We found three errors in our own benchmark claims. This is the report we filed about ourselves.
We build forget, an open-source memory system for AI agents. Like every vendor in this category, we had a benchmark number on our homepage: “81.8% on LongMemEval — above Mem0 (49%) and Zep (63.8%), 0.6 pp under the GPT-4o ceiling, 100% local.” Four claims in one sentence. Three of them were wrong.
Wrong claim 1: the comparison
The “Mem0 49% / Zep 63.8%” pair that floats around this category mixes a third-party measurement using one reader with Zep’s own paper number using a gpt-4o-mini reader (arXiv:2501.13956). Different readers, different measurers, different dates. Aggregator sites fused them into a single table, and we repeated it. With the same gpt-4o reader Zep reports 71.2%. We deleted the comparison from every surface — our Claims Index now exists so that numbers like these can never sit in one column without their configs again.
Wrong claim 2: the ceiling
“0.6 pp under the 82.4% ceiling” — the LongMemEval paper contains no such number. Full-context GPT-4o is 60.6%; the oracle ceiling is 87.0%. The 82.4 came from someone else’s oracle configuration in a different paper. We now cite the paper’s own baselines.
Wrong claim 3: the “100% local”
Our 81.8% run builds memory with a GPT-4o observer. The configuration where the memory pipeline never leaves your machine scores 76.2%. Both numbers are real; attaching “local” to the wrong one was not. Both are now published side by side, with our weakest category (43.3% on single-session-preference) next to them.
Then we re-ran ourselves
Three independent re-runs per configuration, full 500-question LongMemEval-S set, disclosed reader and judge (GPT-4o for both), per-question outputs public:
| configuration | published | re-runs | mean ± run σ | sampling 95% CI |
|---|---|---|---|---|
| best (GPT-4o observer) | 81.8% | 81.8 / 82.4 / 81.8 | 82.0 ± 0.35 | ±3.4 pp |
| fully-local pipeline | 76.2% | 76.8 / 76.4 / 75.2 | 76.1 ± 0.83 | ±3.7 pp |
Both published numbers sit inside their re-run bands. Read the two uncertainty columns separately, because they measure different things. Run σ (0.35 and 0.83 points) is decoding variance — what you get for re-running the same 500 questions. The sampling CI (±3.4 and ±3.7 points) is question-set variance — what you would get for drawing a different 500 questions. Three seeds cannot measure the second kind, which is why a single-run headline is not evidence, and why two systems 3 points apart on this benchmark are statistically indistinguishable. Most published rank claims in this category are unresolvable at n=500.
The judge was audited before any scoring
A score means nothing until the judge is measured, so we attacked ours first: 1,000 adversarial probes — specific-but-wrong answers, vague-but-topical answers, verbatim keys — three votes each, 4,500 calls. The LongMemEval GPT-4o judge held: 0.2% false-accept on specific-but-wrong, 1.0% on vague-but-topical, 0.2% false-reject on verbatim keys. The most-cited alternative benchmark’s judge accepts 62.8% of the same vague probes. One category failed ours too: preference questions are rubric-scored with no ground truth — 90% false-reject — so we exclude them from headline numbers. The full audit is filed separately.
Corrections ledger
- 2026-07-26
“above Mem0 (49%) and Zep (63.8%)”— retracted. Incompatible measurements, fused by aggregators; repeated by us. No competitor numbers in our tables since. - 2026-07-26
“0.6 pp under the 82.4% ceiling”— retracted. The paper's baselines: full-context 60.6, oracle 87.0. - 2026-07-26
“81.8% … 100% local”— re-attributed. 81.8 uses a GPT-4o observer; the fully-local number is 76.1 ± 0.8.
Why we’re publishing this
In the past year this category produced: two vendors re-scoring each other from 84% to 58% and back; a “+159%” headline silently deleted between paper revisions; a vendor whose official numbers reproduced 8–11 points lower in an independent run, with a hardcoded retrieval constant in the code; and the same product scoring anywhere from 11.62 to 94.4 depending on who ran it. We indexed 41 of these claims, each with its config and its dispute status. Every benchmark has an asterisk, and everyone knows it.
The fix isn’t a better leaderboard. It’s a norm: verified claims — multi-run, disclosed configs, public harness, adversarial judge review — starting with your own. This report is us applying that norm to ourselves. Next, someone else’s numbers, with their sign-off on the configuration before we run. Vendors cannot pay for placement; registration fees never buy silence; transcripts are sealed for ten years.
Harness, judge prompts, run configs, raw logs: research/longmemeval in the repo. If your re-run disagrees, open an issue — publicly. If you build a memory system and want your numbers verified under this protocol — or want to tear ours apart — write to us.
cosign verify-blob record.json --bundle ofr-2026-001.sigstore.json