Research preview · operator-run
Which AI memory system actually works?
Memory systems let an AI remember past conversations. Every project claims to be the best. We collect what they publish, show who measured it, and run our own tests in the open.
How the memory systems compare
The highest figure each project publishes per benchmark. Not a ranking: every row was measured under different conditions, so treat small gaps as noise.
| System | Measured by | LongMemEval | LoCoMo | BEAM 1M | BEAM 10M | Recall |
|---|---|---|---|---|---|---|
| ByteRover | Self-reported | — | 96.1% | — | — | — |
| Mem0 | Self-reported | 94.4% | 92.5% | 64.1% | 48.6% | — |
| Hindsight | Self-reported | — | 92% | 79.1% | 64.1% | — |
| gbrain | Self-reported | 90.6% | — | — | — | 95.96% |
| MemOS | Self-reported | 89.2% | 88.83% | — | 56.75% | — |
| CORE | Self-reported | — | 88.24% | — | — | — |
| Memori | Self-reported | — | 87% | — | — | — |
| Nemori | Self-reported | — | 83.05% | — | — | — |
| Cognee | Self-reported | — | 80.3% | — | 67% | — |
| Zep | Self-reported | — | 75.14% | — | — | — |
| ReMe | Self-reported | — | — | 65% | — | — |
| A-Mem | Run by another party | — | 64.16% | — | — | — |
| MemoryOS | Run by another party | — | 58.25% | — | — | — |
| LangMem | Run by another party | — | 58.1% | — | — | — |
| OpenAI memory | Run by another party | — | 52.9% | — | — | — |
| MemPalace | Run by another party | — | — | — | — | 90% |
| supermemory | Self-reported | — | — | — | — | 95% |
| Mnemosyne Ours | Measured here | — | — | — | — | 28.06% |
Answering questions
The two benchmarks everyone quotes, plus how long a query takes.
LongMemEval
Answer accuracy · higher is better
| Hindsight | 94.6% |
|---|---|
| Mem0 | 94.4% |
| ByteRover | 92.8% |
| gbrain | 90.6% |
| gbrain | 89.4% |
| ReMe | 89.4% |
| MemOS | 89.2% |
| hybrid-search | 74% |
| Zep | 71.2% |
| Zep | 63.8% |
Can it answer questions about long chat histories? Each project used its own reader and judge, so these do not compare cleanly.
LoCoMo
Answer accuracy · higher is better
| ByteRover | 96.1% |
|---|---|
| Mem0 | 92.5% |
| Hindsight | 92% |
| MemOS | 88.83% |
| CORE | 88.24% |
| Memori | 87% |
| Nemori | 83.05% |
| Zep | 75.14% |
The most quoted memory benchmark, and the most disputed: its answer key has known errors.
Speed
Seconds per query · lower is better
| OpenAI memory | 0.47 s |
|---|---|
| Mem0 | 0.71 s |
| Mem0 graph | 1.09 s |
| Zep | 1.29 s |
| A-Mem | 1.41 s |
| Full context (no memory) | 9.87 s |
| LangMem | 18.53 s |
A typical query, search and answer together. One study, one machine.
Finding the right memory
Before a system can answer, it has to retrieve the right thing. This is measured separately and never mixed with accuracy.
Finding the evidence
Was the right session in the top five? · higher is better
LongMemEval, recall at five. Projects count a hit differently, so these are not directly comparable — including ours.
Strict recall
Every needed session found · higher is better
A harder count: every session a question needs must be in the top five. Run by gbrain’s authors.
Slow tail
Worst case, seconds · lower is better
| OpenAI memory | 0.89 s |
|---|---|
| Mem0 | 1.44 s |
| Mem0 graph | 2.59 s |
| Zep | 2.93 s |
| A-Mem | 4.37 s |
| Full context (no memory) | 17.12 s |
| LangMem | 60.40 s |
The slowest 5% of queries. LangMem hits 60 s, past the edge of the chart.
Very long conversations
BEAM runs the same memory over 100 thousand, 1 million and 10 million tokens of history.
BEAM 100K
Very long conversations · higher is better
| Hindsight | 86.2% |
|---|---|
| Cognee | 79% |
| ReMe | 66.1% |
Cognee’s figure covers 20 questions from one conversation.
BEAM 1M
Very long conversations · higher is better
| Hindsight | 79.1% |
|---|---|
| ReMe | 65% |
| Mem0 | 64.1% |
Hindsight’s other mode scores 73.9 on its own page.
BEAM 10M
Very long conversations · higher is better
| Cognee | 67% |
|---|---|
| Hindsight | 64.1% |
| MemOS | 56.75% |
| Mem0 | 48.6% |
Cognee’s figure was tuned on the questions it was scored on.
Who ran the test changes the answer
The same software scores differently depending on who measures it and how. This is the case for one neutral operator running everything.
Same system, two scorekeepers
Zep on LoCoMo · higher is better
| Run by Mem0 | 65.99% |
|---|---|
| Run by Zep | 75.14% |
Same software, two scorekeepers, nine points apart. This is why one neutral operator should run everything.
Reader changes the score
LongMemEval · higher is better
| Zep gpt-4o | 71.2% |
|---|---|
| Full context gpt-4o | 60.2% |
| Zep gpt-4o-mini | 63.8% |
| Full context gpt-4o-mini | 55.4% |
The same memory, two different answering models. Compare each bar only with the no-memory baseline beside it.
One harness, run by others
LoCoMo · higher is better
| FullText | 73.83% |
|---|---|
| A-Mem | 64.16% |
| NaiveRAG | 63.64% |
| Mem0 | 61.69% |
| Mem0 graph | 60.32% |
| MemoryOS | 58.25% |
| Mem0 | 36.49% |
Every row here ran on the same harness. Look at Mem0 free versus paid: the same product, 25 points apart.
What projects claim about themselves
LoCoMo · higher is better
| ByteRover | 96.1% |
|---|---|
| Mem0 | 92.5% |
| Hindsight | 92% |
| MemOS | 88.83% |
| CORE | 88.24% |
| Memori | 87% |
| Nemori | 83.05% |
| Zep | 75.14% |
From each project’s own README or blog, at the commit named on the claims page.
Same harness, vendor-operated
Higher is better
LoCoMo
LongMemEval S
PersonaMem 32k
LifeBench en
The Agent Memory Benchmark is run by Vectorize, who also make Hindsight. Same test for every row, but the scorekeeper is competing.
What we measured ourselves
Run by our own harness, signed, and reproducible from the evidence we publish. Our numbers are low and we publish them anyway.
Our LongMemEval run
Did the right evidence come back? · higher is better
| nDCG@5 | 29.67% (interval 26.38% to 33.04%) |
|---|---|
| Recall@5 | 28.06% (interval 24.88% to 31.37%) |
500 questions, signed and reproducible. Whiskers show the range. This asks only whether the right evidence came back, not whether the answer was right.
Multi-hop retrieval
Recall at five · higher is better
| Ours HotpotQA | 37.4% |
|---|---|
| HippoRAG 2 HotpotQA | 96.3% |
| Ours 2WikiMultiHopQA | 23.73% |
| HippoRAG 2 2WikiMultiHopQA | 90.4% |
| Ours MuSiQue | 10.42% |
| HippoRAG 2 MuSiQue | 74.7% |
1,000 questions each. Red is HippoRAG 2’s published result: the bar to beat, not a run here. Our graph search did not fire on a single query, which is why these are low.
How to read these
Finding evidence and answering well are different things, so they never share a chart.
Our results in full
Every metric we have measured, with its range and the evidence behind it.
| System | Benchmark | Metric | Result | Evidence |
|---|---|---|---|---|
| mnemosyne-localregistered-operator-retrieval | longmemeval-retrievallongmemeval-retrieval-v1 | recall_at_5retrieval | 0.2806 ratioInterval: 0.24880000000000002 to 0.31373333333333336 | View run 2026-10-04-longmemeval-retrieval-b0cdbd89operator-run; not publishableOperator: Mnemosyne projectRun disclosuresSystemmnemosyne-local Trackregistered-operator-retrieval Benchmarklongmemeval-retrieval longmemeval-retrieval-v1 Publicationoperator-run; not publishable OperatorMnemosyne project; disclosed: true Metrics
Immutable artifacts run_commit b0cdbd8986b94c91a3bf20833ed33f405612025abuild_fingerprint sha256:7ba4789e105be00461f40cdf5c3b84e29785b2234839976e3a20ef8d954b079cconfig_digest sha256:3cad9149f630c8c2d610ad87d9203e32713d071df6e96bd918c610291d3492a9bundle_digest sha256:c8449069cc96e7be7700283fec67d5ed53150c37d6ded0af634535c04e99f175trace_index_digest sha256:7146385bdc49de2c0353dd91689c08a4a60b74083de39f2bd5a7e802efabf20a |
| mnemosyne-localregistered-operator-retrieval | longmemeval-retrievallongmemeval-retrieval-v1 | ndcg_at_5retrieval | 0.2967188496001503 ratioInterval: 0.26378166565316336 to 0.33038741454372256 | View run 2026-10-04-longmemeval-retrieval-b0cdbd89operator-run; not publishableOperator: Mnemosyne projectRun disclosuresSystemmnemosyne-local Trackregistered-operator-retrieval Benchmarklongmemeval-retrieval longmemeval-retrieval-v1 Publicationoperator-run; not publishable OperatorMnemosyne project; disclosed: true Metrics
Immutable artifacts run_commit b0cdbd8986b94c91a3bf20833ed33f405612025abuild_fingerprint sha256:7ba4789e105be00461f40cdf5c3b84e29785b2234839976e3a20ef8d954b079cconfig_digest sha256:3cad9149f630c8c2d610ad87d9203e32713d071df6e96bd918c610291d3492a9bundle_digest sha256:c8449069cc96e7be7700283fec67d5ed53150c37d6ded0af634535c04e99f175trace_index_digest sha256:7146385bdc49de2c0353dd91689c08a4a60b74083de39f2bd5a7e802efabf20a |