Published claims
What every memory system claims.
The scores these projects publish about themselves, collected with a link to where each one came from.
These figures as charts
Every chart on this data lives on the front page, grouped by what it measures.
The projects
Licences come from GitHub. Closed products have no repository.
Mem0
Open-source SDK plus managed platform
Platform scores include proprietary optimizations
RepositoryMemPalace
Open source, local-first
Declines to compare itself with other projects
Repositorysupermemory
Open source plus hosted
Claims #1 without a figure for answer quality
RepositoryMemori
Memory infrastructure
LoCoMo paper arXiv 2603.19935
RepositoryByteRover
Memory layer for coding agents
LoCoMo run on its production codebase
RepositoryCORE
Personal memory layer
Benchmark repo published separately
RepositoryCognee
Open-source knowledge graph memory
BEAM runs scoped to a few questions
RepositoryZep
Graphiti open source; Zep Cloud managed
Disputes the figure in the Mem0 paper
RepositoryLightMem
Research memory framework
Its README compares several systems on one harness
RepositoryHoncho
Memory library for stateful agents
Points to an evals page; no figures in its README
RepositoryOpenAI memory
Closed, built into ChatGPT
Only appears via the Mem0 paper
Mnemosyne
Operator entry of this site
Held to the same rules as every system
RepositoryClaims with no usable figure:
Every claim
Type to filter. Each row links to where the number came from.
| System | Benchmark | Metric | Value | Family | Who ran it | Source |
|---|---|---|---|---|---|---|
| Hindsight | BEAM 100K | accuracyAgent Memory Benchmark comparison table. | 86.2% | Answer quality | Self-reported | Agent Memory Benchmark, run by Vectorize (maker of Hindsight) |
| Hindsight | BEAM 1M | accuracyAgent Memory Benchmark comparison table. | 79.1% | Answer quality | Self-reported | Agent Memory Benchmark, run by Vectorize (maker of Hindsight) |
| Cognee | BEAM 100K | score (0 to 1 scale)Reported 0.79: four rounds over 20 questions from one held-out conversation. | 79% | Answer quality | Self-reported | Cognee READMEcommit b32d8af line 291 |
| Hindsight · single query | BEAM 100K | accuracyHindsight results page, single-query mode. | 75% | Answer quality | Self-reported | Hindsight benchmarks page |
| Hindsight · single query | BEAM 1M | accuracyHindsight results page, single-query mode. | 73.9% | Answer quality | Self-reported | Hindsight benchmarks page |
| Cognee | BEAM 10M | score (0 to 1 scale)Reported 0.67: exploratory, question-type routing selected and scored on the same questions. | 67% | Answer quality | Self-reported | Cognee READMEcommit b32d8af line 292 |
| ReMe | BEAM 100K | agentic score20 cases, 400 questions. | 66.1% | Answer quality | Self-reported | ReMe READMEcommit 084c02e line 356 |
| ReMe | BEAM 1M | agentic score35 cases, 700 questions. | 65% | Answer quality | Self-reported | ReMe READMEcommit 084c02e line 357 |
| Hindsight | BEAM 10M | accuracyAgent Memory Benchmark and the Hindsight results page agree. | 64.1% | Answer quality | Self-reported | Agent Memory Benchmark, run by Vectorize (maker of Hindsight) |
| Mem0 | BEAM 1M | scoreSame setup. | 64.1% | Answer quality | Self-reported | Mem0 READMEcommit c93420c line 51 |
| MemOS | BEAM 10M | scoreEvaluated through OmniMemEval. | 56.75% | Answer quality | Self-reported | MemOS READMEcommit a7367d0 line 82 |
| Mem0 | BEAM 10M | scoreSame setup. | 48.6% | Answer quality | Self-reported | Mem0 READMEcommit c93420c line 52 |
| MemPalace | ConvoMem | average recallAll categories, 250 items, 50 per category. | 92.9% | Retrieval recall | Self-reported | MemPalace READMEcommit 35dc621 line 268 |
| supermemory | ConvoMem | ranking claim: #1States #1 with no figure in the README. | — | Metric not stated | Self-reported | supermemory READMEcommit 3535ff7 line 355 |
| MemOS | HaluMem | scoreEvaluated through OmniMemEval. | 80.91% | Answer quality | Self-reported | MemOS READMEcommit a7367d0 line 81 |
| HippoRAG 2 | HippoRAG: 2WikiMultiHopQA | Recall@5Published upstream result, Llama-3.3-70B. | 90.4% | Retrieval recall | Run by another party | Mnemetric Phase 11 evidence notecommit main line 35 |
| Mnemosyne | HippoRAG: 2WikiMultiHopQA | Recall@51,000 questions, July 2026, retrieval only. | 23.73% | Retrieval recall | Measured here | Mnemetric 2WikiMultiHopQA retrieval reportcommit main |
| HippoRAG 2 | HippoRAG: HotpotQA | Recall@5Published upstream result, Llama-3.3-70B. | 96.3% | Retrieval recall | Run by another party | Mnemetric Phase 11 evidence notecommit main line 36 |
| Mnemosyne | HippoRAG: HotpotQA | Recall@51,000 questions, July 2026, retrieval only. | 37.4% | Retrieval recall | Measured here | Mnemetric HotpotQA retrieval reportcommit main |
| HippoRAG 2 | HippoRAG: MuSiQue | Recall@5Published upstream result, Llama-3.3-70B. | 74.7% | Retrieval recall | Run by another party | Mnemetric Phase 11 evidence notecommit main line 34 |
| Mnemosyne | HippoRAG: MuSiQue | Recall@51,000 questions, July 2026, retrieval only. | 10.42% | Retrieval recall | Measured here | Mnemetric MuSiQue retrieval reportcommit main |
| Hindsight | LifeBench en | accuracyAgent Memory Benchmark. | 71.5% | Answer quality | Self-reported | Agent Memory Benchmark, run by Vectorize (maker of Hindsight) |
| hybrid-search | LifeBench en | accuracyAgent Memory Benchmark baseline: plain hybrid search. | 61% | Answer quality | Run by another party | Agent Memory Benchmark, run by Vectorize (maker of Hindsight) |
| ByteRover | LoCoMo | overall accuracy1,982 questions, 272 documents; production codebase, no separate prototype. | 96.1% | Answer quality | Self-reported | ByteRover READMEcommit 1052ac1 line 55 |
| Mem0 | LoCoMo | scoreApril 2026 algorithm on the managed platform, which includes proprietary optimizations; single pass, top_200; was 71.4. | 92.5% | Answer quality | Self-reported | Mem0 READMEcommit c93420c line 49 |
| Hindsight | LoCoMo | accuracyAgent Memory Benchmark; 1,986 questions. | 92% | Answer quality | Self-reported | Agent Memory Benchmark, run by Vectorize (maker of Hindsight) |
| MemOS | LoCoMo | scoreEvaluated through OmniMemEval, a framework run by the same organisation. | 88.83% | Answer quality | Self-reported | MemOS READMEcommit a7367d0 line 78 |
| CORE | LoCoMo | average accuracyAcross single-hop, multi-hop, open-domain and temporal questions. | 88.24% | Answer quality | Self-reported | CORE READMEcommit 4a5b18d line 244 |
| Memori | LoCoMo | overall accuracyAbout 721 tokens per query, 2.8% of the full-context footprint. | 87% | Answer quality | Self-reported | Memori READMEcommit 574b1ea line 130 |
| Nemori | LoCoMo | LLM scoreReported 0.8305 overall, version V5; 1,540 questions across the four categories. | 83.05% | Answer quality | Self-reported | Nemori READMEcommit d2a6dff line 171 |
| Cognee | LoCoMo | accuracyAgent Memory Benchmark, run by Hindsight's maker. | 80.3% | Answer quality | Run by another party | Agent Memory Benchmark, run by Vectorize (maker of Hindsight) |
| hybrid-search | LoCoMo | accuracyAgent Memory Benchmark baseline: plain hybrid search. | 79.1% | Answer quality | Run by another party | Agent Memory Benchmark, run by Vectorize (maker of Hindsight) |
| Zep | LoCoMo | J score (LLM judge)Zep corrected its own earlier post and reports 75.14 +/- 0.17, against 65.99 in the Mem0 paper. | 75.14% | Answer quality | Self-reported | Zep blog, corrected LoCoMo result |
| FullText | LoCoMo | ACCLightMem authors; backbone and judge gpt-4o-mini. | 73.83% | Answer quality | Run by another party | LightMem READMEcommit 8449d57 line 393 |
| Full context (no memory) | LoCoMo | J score (LLM judge)Whole conversation in the prompt, about 26,000 tokens. | 72.9% | Answer quality | Run by another party | Mem0 paper, arXiv 2504.19413 |
| Mem0 graph | LoCoMo | J score (LLM judge)Mem0 paper. | 68.44% | Answer quality | Self-reported | Mem0 paper, arXiv 2504.19413 |
| Mem0 | LoCoMo | J score (LLM judge)Mem0 paper. | 66.88% | Answer quality | Self-reported | Mem0 paper, arXiv 2504.19413 |
| Zep | LoCoMo | J score (LLM judge)Mem0 authors running Zep. | 65.99% | Answer quality | Run by another party | Mem0 paper, arXiv 2504.19413 |
| A-Mem | LoCoMo | ACCLightMem authors; backbone and judge gpt-4o-mini. | 64.16% | Answer quality | Run by another party | LightMem READMEcommit 8449d57 line 395 |
| NaiveRAG | LoCoMo | ACCLightMem authors; backbone and judge gpt-4o-mini. | 63.64% | Answer quality | Run by another party | LightMem READMEcommit 8449d57 line 394 |
| Mem0 · API | LoCoMo | ACCLightMem authors; the hosted API. | 61.69% | Answer quality | Run by another party | LightMem READMEcommit 8449d57 line 399 |
| Mem0 graph · API | LoCoMo | ACCLightMem authors; the hosted API with graph. | 60.32% | Answer quality | Run by another party | LightMem READMEcommit 8449d57 line 400 |
| MemoryOS · eval build | LoCoMo | ACCLightMem authors; the evaluation build. | 58.25% | Answer quality | Run by another party | LightMem READMEcommit 8449d57 line 396 |
| LangMem | LoCoMo | J score (LLM judge)Mem0 paper. | 58.1% | Answer quality | Run by another party | Mem0 paper, arXiv 2504.19413 |
| MemoryOS · PyPI | LoCoMo | ACCLightMem authors; the PyPI release. | 54.87% | Answer quality | Run by another party | LightMem READMEcommit 8449d57 line 397 |
| OpenAI memory | LoCoMo | J score (LLM judge)OpenAI memory as run by the Mem0 authors; closed source. | 52.9% | Answer quality | Run by another party | Mem0 paper, arXiv 2504.19413 |
| A-Mem | LoCoMo | J score (LLM judge)Mem0 paper. | 48.38% | Answer quality | Run by another party | Mem0 paper, arXiv 2504.19413 |
| Mem0 · open source | LoCoMo | ACCLightMem authors; the open-source package. | 36.49% | Answer quality | Run by another party | LightMem READMEcommit 8449d57 line 398 |
| OpenViking | LoCoMo | accuracyReports 80 to 83% across three agent integrations, against 24 to 57% on their native memory; reader Doubao 2.0 Pro. | 80 to 83% | Answer quality | Self-reported | OpenViking READMEcommit 10f3681 line 120 |
| LangMem | LoCoMo | total latency p95Search alone was 59.82 s. | 60.4 s | Latency | Run by another party | Mem0 paper, arXiv 2504.19413 |
| LangMem | LoCoMo | total latency p50Search alone was 17.99 s; the total is search plus answer. | 18.53 s | Latency | Run by another party | Mem0 paper, arXiv 2504.19413 |
| Full context (no memory) | LoCoMo | total latency p95Whole conversation in the prompt. | 17.117 s | Latency | Run by another party | Mem0 paper, arXiv 2504.19413 |
| Full context (no memory) | LoCoMo | total latency p50Whole conversation in the prompt. | 9.87 s | Latency | Run by another party | Mem0 paper, arXiv 2504.19413 |
| A-Mem | LoCoMo | total latency p95Search plus answer, seconds. | 4.374 s | Latency | Run by another party | Mem0 paper, arXiv 2504.19413 |
| Zep | LoCoMo | total latency p95Search plus answer, seconds. | 2.926 s | Latency | Run by another party | Mem0 paper, arXiv 2504.19413 |
| Mem0 graph | LoCoMo | total latency p95Search plus answer, seconds. | 2.59 s | Latency | Self-reported | Mem0 paper, arXiv 2504.19413 |
| Mem0 | LoCoMo | total latency p95Search plus answer, seconds. | 1.44 s | Latency | Self-reported | Mem0 paper, arXiv 2504.19413 |
| A-Mem | LoCoMo | total latency p50Search plus answer, seconds. | 1.41 s | Latency | Run by another party | Mem0 paper, arXiv 2504.19413 |
| Zep | LoCoMo | total latency p50Search plus answer, seconds. | 1.292 s | Latency | Run by another party | Mem0 paper, arXiv 2504.19413 |
| Mem0 graph | LoCoMo | total latency p50Search plus answer, seconds. | 1.091 s | Latency | Self-reported | Mem0 paper, arXiv 2504.19413 |
| OpenAI memory | LoCoMo | total latency p95No search step. | 0.889 s | Latency | Run by another party | Mem0 paper, arXiv 2504.19413 |
| Mem0 | LoCoMo | total latency p50Search plus answer, seconds. | 0.708 s | Latency | Self-reported | Mem0 paper, arXiv 2504.19413 |
| OpenAI memory | LoCoMo | total latency p50No search step; memories are extracted manually in the prompt. | 0.466 s | Latency | Run by another party | Mem0 paper, arXiv 2504.19413 |
| MemPalace · hybrid v5 | LoCoMo | R@10Hybrid v5, top-10, no rerank; same 1,986 questions. | 88.9% | Retrieval recall | Self-reported | MemPalace READMEcommit 35dc621 line 267 |
| MemPalace · session | LoCoMo | R@10Session level, top-10, no rerank; 1,986 questions. | 60.3% | Retrieval recall | Self-reported | MemPalace READMEcommit 35dc621 line 266 |
| Memanto | LoCoMo | scoreSame caveat. | 87.1% | Metric not stated | Self-reported | Memanto READMEcommit aac858c line 332 |
| supermemory | LoCoMo | ranking claim: #1States #1 with no figure in the README. | — | Metric not stated | Self-reported | supermemory READMEcommit 3535ff7 line 354 |
| Hindsight | LongMemEval S | accuracyAgent Memory Benchmark, operated by Vectorize, which makes Hindsight. | 94.6% | Answer quality | Self-reported | Agent Memory Benchmark, run by Vectorize (maker of Hindsight) |
| Mem0 | LongMemEval | scoreSame setup; was 67.8. | 94.4% | Answer quality | Self-reported | Mem0 READMEcommit c93420c line 50 |
| ByteRover | LongMemEval S | overall accuracy500 questions, 23,867 documents. | 92.8% | Answer quality | Self-reported | ByteRover READMEcommit 1052ac1 line 57 |
| gbrain | LongMemEval | accuracy453 of 500; house reader with reranker. | 90.6% | Answer quality | Self-reported | gbrain-evals READMEcommit 48dd47b line 50 |
| ReMe | LongMemEval cleaned-s | agentic score500 questions. | 89.4% | Answer quality | Self-reported | ReMe READMEcommit 084c02e line 355 |
| gbrain · gpt-5.4 reader | LongMemEval | accuracy447 of 500; gpt-5.4 reader on gbrain retrieval. | 89.4% | Answer quality | Self-reported | gbrain-evals READMEcommit 48dd47b line 51 |
| MemOS | LongMemEval | scoreEvaluated through OmniMemEval. | 89.2% | Answer quality | Self-reported | MemOS READMEcommit a7367d0 line 79 |
| hybrid-search | LongMemEval S | accuracyAgent Memory Benchmark baseline: plain hybrid search. | 74% | Answer quality | Run by another party | Agent Memory Benchmark, run by Vectorize (maker of Hindsight) |
| Zep · gpt-4o reader | LongMemEval S | accuracyZep paper; reader gpt-4o; full-context baseline 60.2; latency 2.58 s against 28.9 s. | 71.2% | Answer quality | Self-reported | Zep paper, arXiv 2501.13956, Table 2 |
| Zep · gpt-4o-mini reader | LongMemEval S | accuracyZep paper; reader gpt-4o-mini; full-context baseline 55.4; latency 3.20 s against 31.3 s. | 63.8% | Answer quality | Self-reported | Zep paper, arXiv 2501.13956, Table 2 |
| Full context (no memory) · gpt-4o reader | LongMemEval S | accuracyZep paper; reader gpt-4o; about 115k tokens of context. | 60.2% | Answer quality | Self-reported | Zep paper, arXiv 2501.13956, Table 2 |
| Full context (no memory) · gpt-4o-mini reader | LongMemEval S | accuracyZep paper; reader gpt-4o-mini; about 115k tokens of context. | 55.4% | Answer quality | Self-reported | Zep paper, arXiv 2501.13956, Table 2 |
| MemPalace · hybrid v4, held out | LongMemEval | R@5Hybrid v4, held-out 450 questions; tuned on 50 dev questions. | 98.4% | Retrieval recall | Self-reported | MemPalace READMEcommit 35dc621 line 246 |
| MemPalace · raw | LongMemEval | R@5500 questions; raw semantic search, no LLM, no heuristics. | 96.6% | Retrieval recall | Self-reported | MemPalace READMEcommit 35dc621 line 245 |
| gbrain | LongMemEval | strict recall_all@5451 of 470 questions; every required session must be in the top five; opaque session ids; with Voyage reranker. | 95.96% | Retrieval recall | Self-reported | gbrain-evals READMEcommit 48dd47b line 49 |
| agentmemory | LongMemEval S | R@5500 questions; all-MiniLM-L6-v2 embeddings, local. | 95.2% | Retrieval recall | Self-reported | agentmemory READMEcommit 007a1a7 line 356 |
| supermemory | LongMemEval | Recall@15Adds about 720 tokens of context, a 99.4% reduction. | 95% | Retrieval recall | Self-reported | supermemory READMEcommit 3535ff7 line 357 |
| gbrain · no reranker | LongMemEval | strict recall_all@5434 of 470; same without the reranker. | 92.34% | Retrieval recall | Self-reported | gbrain-evals READMEcommit 48dd47b line 88 |
| MemPalace · hybrid + LLM rerank, recounted | LongMemEval | strict recall_all@5gbrain authors recounting MemPalace's saved rankings strictly (423 of 470); hybrid search plus LLM rerank. | 90% | Retrieval recall | Run by another party | gbrain-evals READMEcommit 48dd47b line 89 |
| BM25-only fallback | LongMemEval S | R@5agentmemory's own keyword-only baseline on the same questions. | 86.2% | Retrieval recall | Self-reported | agentmemory READMEcommit 007a1a7 line 357 |
| MemPalace · raw, recounted | LongMemEval | strict recall_all@5Recount of the raw vector search (403 of 470). | 85.7% | Retrieval recall | Run by another party | gbrain-evals READMEcommit 48dd47b line 91 |
| Mnemosyne | LongMemEval | nDCG@5500 questions, interval 26.38 to 33.04. | 29.67% | Retrieval recall | Measured here | Mnemetric LongMemEval retrieval reportcommit main |
| Mnemosyne | LongMemEval | Recall@5500 questions, interval 24.88 to 31.37. | 28.06% | Retrieval recall | Measured here | Mnemetric LongMemEval retrieval reportcommit main |
| Memanto | LongMemEval | scoreREADME calls these public recall benchmarks but does not state the metric, and warns cross-project scores are not comparable. | 89.8% | Metric not stated | Self-reported | Memanto READMEcommit aac858c line 332 |
| supermemory | LongMemEval | ranking claim: #1States #1 with no figure in the README. | — | Metric not stated | Self-reported | supermemory READMEcommit 3535ff7 line 353 |
| MemPalace | MemBench | R@5ACL 2025, 8,500 items, all categories. | 80.3% | Retrieval recall | Self-reported | MemPalace READMEcommit 35dc621 line 269 |
| Hindsight | PersonaMem 32k | accuracyAgent Memory Benchmark. | 86.6% | Answer quality | Self-reported | Agent Memory Benchmark, run by Vectorize (maker of Hindsight) |
| hybrid-search | PersonaMem 32k | accuracyAgent Memory Benchmark baseline: plain hybrid search. | 84.4% | Answer quality | Run by another party | Agent Memory Benchmark, run by Vectorize (maker of Hindsight) |
| Cognee | PersonaMem 32k | accuracyAgent Memory Benchmark, run by Hindsight's maker. | 81.8% | Answer quality | Run by another party | Agent Memory Benchmark, run by Vectorize (maker of Hindsight) |
| MemOS | PersonaMem v2 | scoreEvaluated through OmniMemEval. | 40.58% | Answer quality | Self-reported | MemOS READMEcommit a7367d0 line 80 |
Corrections we have made
Numbers that were wrong or unsourced on this site, and what replaced them.
- Zep LongMemEval: an earlier version of the landscape page mixed the two readers in the Zep paper. The paper reports 63.8 with gpt-4o-mini (full context 55.4) and 71.2 with gpt-4o (full context 60.2).
- Zep latency: an earlier version used the Zep paper's latency in a chart of the Mem0 paper. The Mem0 paper measures Zep at 1.292 s p50 and 2.926 s p95 in total.
- LangMem latency: the figure shown was search time only (17.99 s and 59.82 s). The total is 18.53 s and 60.40 s.
- Removed: self-reported LoCoMo figures for supermemory (81.6) and MemOS (73.3) that could not be traced to a source. MemOS now reports 88.83; supermemory publishes a #1 claim without a figure.
- The capability table is an editorial summary and is labelled that way. Licences on this page come from GitHub's licence detection.
Collected 2026-10-06. Sources can change after that date. How the benchmarks differ · The charts