Research preview · operator-run

Which AI memory system actually works?

Memory systems let an AI remember past conversations. Every project claims to be the best. We collect what they publish, show who measured it, and run our own tests in the open.

Systems tracked21Open source and closed
Published figures93Each linked to its source
Measured by us1Signed and reproducible
Benchmarks tracked14From LongMemEval to BEAM

How the memory systems compare

The highest figure each project publishes per benchmark. Not a ranking: every row was measured under different conditions, so treat small gaps as noise.

All claims and sources
SystemMeasured byLongMemEvalLoCoMoBEAM 1MBEAM 10MRecall
ByteRoverSelf-reported—96.1%———
Mem0Self-reported94.4%92.5%64.1%48.6%—
HindsightSelf-reported—92%79.1%64.1%—
gbrainSelf-reported90.6%———95.96%
MemOSSelf-reported89.2%88.83%—56.75%—
CORESelf-reported—88.24%———
MemoriSelf-reported—87%———
NemoriSelf-reported—83.05%———
CogneeSelf-reported—80.3%—67%—
ZepSelf-reported—75.14%———
ReMeSelf-reported——65%——
A-MemRun by another party—64.16%———
MemoryOSRun by another party—58.25%———
LangMemRun by another party—58.1%———
OpenAI memoryRun by another party—52.9%———
MemPalaceRun by another party————90%
supermemorySelf-reported————95%
Mnemosyne OursMeasured here————28.06%
Measured hereSelf-reportedRun by another partyNo-memory baselineNo figure published

Answering questions

The two benchmarks everyone quotes, plus how long a query takes.

LongMemEval

Answer accuracy · higher is better

Mixed sources
0%25%50%75%100%94.6%Hindsight94.4%Mem092.8%ByteRover90.6%gbrain89.4%gbrain89.4%ReMe89.2%MemOS74%hybrid-search71.2%Zep63.8%Zep
LongMemEval answer accuracy by system
Hindsight94.6%
Mem094.4%
ByteRover92.8%
gbrain90.6%
gbrain89.4%
ReMe89.4%
MemOS89.2%
hybrid-search74%
Zep71.2%
Zep63.8%

Can it answer questions about long chat histories? Each project used its own reader and judge, so these do not compare cleanly.

LoCoMo

Answer accuracy · higher is better

Mixed sources
0%25%50%75%100%96.1%ByteRover92.5%Mem092%Hindsight88.83%MemOS88.24%CORE87%Memori83.05%Nemori75.14%Zep
LoCoMo answer accuracy by system
ByteRover96.1%
Mem092.5%
Hindsight92%
MemOS88.83%
CORE88.24%
Memori87%
Nemori83.05%
Zep75.14%

The most quoted memory benchmark, and the most disputed: its answer key has known errors.

Speed

Seconds per query · lower is better

Run by Mem0
0s5s10s15s20s0.47 sOpenAI memory0.71 sMem01.09 sMem0 graph1.29 sZep1.41 sA-Mem9.87 sFull context (no memory)18.53 sLangMem
Seconds per query · lower is better
OpenAI memory0.47 s
Mem00.71 s
Mem0 graph1.09 s
Zep1.29 s
A-Mem1.41 s
Full context (no memory)9.87 s
LangMem18.53 s

A typical query, search and answer together. One study, one machine.

Finding the right memory

Before a system can answer, it has to retrieve the right thing. This is measured separately and never mixed with accuracy.

Finding the evidence

Was the right session in the top five? · higher is better

Mixed sources
  • MemPalace · hybrid v4, held outR@598.4%
  • MemPalace · rawR@596.6%
  • agentmemoryR@595.2%
  • BM25-only fallbackR@586.2%
  • MnemosyneRecall@528.06%

LongMemEval, recall at five. Projects count a hit differently, so these are not directly comparable — including ours.

Strict recall

Every needed session found · higher is better

Run by gbrain
  • gbrainstrict recall_all@595.96%
  • gbrain · no rerankerstrict recall_all@592.34%
  • MemPalace · hybrid + LLM rerank, recountedstrict recall_all@590%
  • MemPalace · raw, recountedstrict recall_all@585.7%

A harder count: every session a question needs must be in the top five. Run by gbrain’s authors.

Slow tail

Worst case, seconds · lower is better

Run by Mem0
0s5s10s15s20s0.89 sOpenAI memory1.44 sMem02.59 sMem0 graph2.93 sZep4.37 sA-Mem17.12 sFull context (no memory)60.40 s ↑LangMem
Worst case, seconds · lower is better
OpenAI memory0.89 s
Mem01.44 s
Mem0 graph2.59 s
Zep2.93 s
A-Mem4.37 s
Full context (no memory)17.12 s
LangMem60.40 s

The slowest 5% of queries. LangMem hits 60 s, past the edge of the chart.

Very long conversations

BEAM runs the same memory over 100 thousand, 1 million and 10 million tokens of history.

BEAM 100K

Very long conversations · higher is better

Mixed sources
0%25%50%75%100%86.2%Hindsight79%Cognee66.1%ReMe
BEAM 100K score by system
Hindsight86.2%
Cognee79%
ReMe66.1%

Cognee’s figure covers 20 questions from one conversation.

BEAM 1M

Very long conversations · higher is better

Mixed sources
0%25%50%75%100%79.1%Hindsight65%ReMe64.1%Mem0
BEAM 1M score by system
Hindsight79.1%
ReMe65%
Mem064.1%

Hindsight’s other mode scores 73.9 on its own page.

BEAM 10M

Very long conversations · higher is better

Mixed sources
0%25%50%75%100%67%Cognee64.1%Hindsight56.75%MemOS48.6%Mem0
BEAM 10M score by system
Cognee67%
Hindsight64.1%
MemOS56.75%
Mem048.6%

Cognee’s figure was tuned on the questions it was scored on.

Who ran the test changes the answer

The same software scores differently depending on who measures it and how. This is the case for one neutral operator running everything.

Same system, two scorekeepers

Zep on LoCoMo · higher is better

Disputed
0%25%50%75%100%65.99%Run by Mem075.14%Run by Zep
Zep LoCoMo accuracy measured by Mem0 and by Zep
Run by Mem065.99%
Run by Zep75.14%

Same software, two scorekeepers, nine points apart. This is why one neutral operator should run everything.

Reader changes the score

LongMemEval · higher is better

Self-reported
0%25%50%75%100%71.2%Zep gpt-4o60.2%Full context gpt-4o63.8%Zep gpt-4o-mini55.4%Full context gpt-4o-mini
LongMemEval accuracy from the Zep paper, two readers
Zep gpt-4o71.2%
Full context gpt-4o60.2%
Zep gpt-4o-mini63.8%
Full context gpt-4o-mini55.4%

The same memory, two different answering models. Compare each bar only with the no-memory baseline beside it.

One harness, run by others

LoCoMo · higher is better

Run by another party
0%25%50%75%100%73.83%FullText64.16%A-Mem63.64%NaiveRAG61.69%Mem060.32%Mem0 graph58.25%MemoryOS36.49%Mem0
LoCoMo accuracy measured by the LightMem authors
FullText73.83%
A-Mem64.16%
NaiveRAG63.64%
Mem061.69%
Mem0 graph60.32%
MemoryOS58.25%
Mem036.49%

Every row here ran on the same harness. Look at Mem0 free versus paid: the same product, 25 points apart.

What projects claim about themselves

LoCoMo · higher is better

Self-reported
0%25%50%75%100%96.1%ByteRover92.5%Mem092%Hindsight88.83%MemOS88.24%CORE87%Memori83.05%Nemori75.14%Zep
Self-reported LoCoMo accuracy by project
ByteRover96.1%
Mem092.5%
Hindsight92%
MemOS88.83%
CORE88.24%
Memori87%
Nemori83.05%
Zep75.14%

From each project’s own README or blog, at the commit named on the claims page.

Same harness, vendor-operated

Higher is better

Scorekeeper competes

LoCoMo

  • Hindsightaccuracy92%
  • Cogneeaccuracy80.3%
  • hybrid-searchaccuracy79.1%

LongMemEval S

  • Hindsightaccuracy94.6%
  • hybrid-searchaccuracy74%

PersonaMem 32k

  • Hindsightaccuracy86.6%
  • hybrid-searchaccuracy84.4%
  • Cogneeaccuracy81.8%

LifeBench en

  • Hindsightaccuracy71.5%
  • hybrid-searchaccuracy61%

The Agent Memory Benchmark is run by Vectorize, who also make Hindsight. Same test for every row, but the scorekeeper is competing.

What we measured ourselves

Run by our own harness, signed, and reproducible from the evidence we publish. Our numbers are low and we publish them anyway.

Compare runs

Our LongMemEval run

Did the right evidence come back? · higher is better

Measured here
0%25%50%75%100%29.67%nDCG@528.06%Recall@5
Mnemosyne LongMemEval retrieval with reported range
nDCG@529.67% (interval 26.38% to 33.04%)
Recall@528.06% (interval 24.88% to 31.37%)

500 questions, signed and reproducible. Whiskers show the range. This asks only whether the right evidence came back, not whether the answer was right.

Multi-hop retrieval

Recall at five · higher is better

Measured here
0%25%50%75%100%37.4%Ours HotpotQA96.3%HippoRAG 2 HotpotQA23.73%Ours 2WikiMultiHopQA90.4%HippoRAG 2 2WikiMultiHopQA10.42%Ours MuSiQue74.7%HippoRAG 2 MuSiQue
Mnemosyne versus published HippoRAG 2 recall on three datasets
Ours HotpotQA37.4%
HippoRAG 2 HotpotQA96.3%
Ours 2WikiMultiHopQA23.73%
HippoRAG 2 2WikiMultiHopQA90.4%
Ours MuSiQue10.42%
HippoRAG 2 MuSiQue74.7%

1,000 questions each. Red is HippoRAG 2’s published result: the bar to beat, not a run here. Our graph search did not fire on a single query, which is why these are low.

How to read these

Finding evidence and answering well are different things, so they never share a chart.

BarThe number the project reported
StripedThe project measured itself
VioletMnemosyne, our own system
GreyNo memory at all, for comparison
MissingNot measured, which is not the same as zero
Reading results

Our results in full

Every metric we have measured, with its range and the evidence behind it.

Download JSON
SystemBenchmarkMetricResultEvidence
mnemosyne-localregistered-operator-retrievallongmemeval-retrievallongmemeval-retrieval-v1recall_at_5retrieval0.2806 ratioInterval: 0.24880000000000002 to 0.31373333333333336View run 2026-10-04-longmemeval-retrieval-b0cdbd89operator-run; not publishableOperator: Mnemosyne project
Run disclosures
Systemmnemosyne-local
Trackregistered-operator-retrieval
Benchmarklongmemeval-retrieval longmemeval-retrieval-v1
Publicationoperator-run; not publishable
OperatorMnemosyne project; disclosed: true

Metrics

  • retrieval: ndcg_at_5 = 0.2967188496001503 ratioInterval: 0.26378166565316336 to 0.33038741454372256
  • retrieval: recall_at_5 = 0.2806 ratioInterval: 0.24880000000000002 to 0.31373333333333336

Immutable artifacts

run_commitb0cdbd8986b94c91a3bf20833ed33f405612025a
build_fingerprintsha256:7ba4789e105be00461f40cdf5c3b84e29785b2234839976e3a20ef8d954b079c
config_digestsha256:3cad9149f630c8c2d610ad87d9203e32713d071df6e96bd918c610291d3492a9
bundle_digestsha256:c8449069cc96e7be7700283fec67d5ed53150c37d6ded0af634535c04e99f175
trace_index_digestsha256:7146385bdc49de2c0353dd91689c08a4a60b74083de39f2bd5a7e802efabf20a
mnemosyne-localregistered-operator-retrievallongmemeval-retrievallongmemeval-retrieval-v1ndcg_at_5retrieval0.2967188496001503 ratioInterval: 0.26378166565316336 to 0.33038741454372256View run 2026-10-04-longmemeval-retrieval-b0cdbd89operator-run; not publishableOperator: Mnemosyne project
Run disclosures
Systemmnemosyne-local
Trackregistered-operator-retrieval
Benchmarklongmemeval-retrieval longmemeval-retrieval-v1
Publicationoperator-run; not publishable
OperatorMnemosyne project; disclosed: true

Metrics

  • retrieval: ndcg_at_5 = 0.2967188496001503 ratioInterval: 0.26378166565316336 to 0.33038741454372256
  • retrieval: recall_at_5 = 0.2806 ratioInterval: 0.24880000000000002 to 0.31373333333333336

Immutable artifacts

run_commitb0cdbd8986b94c91a3bf20833ed33f405612025a
build_fingerprintsha256:7ba4789e105be00461f40cdf5c3b84e29785b2234839976e3a20ef8d954b079c
config_digestsha256:3cad9149f630c8c2d610ad87d9203e32713d071df6e96bd918c610291d3492a9
bundle_digestsha256:c8449069cc96e7be7700283fec67d5ed53150c37d6ded0af634535c04e99f175
trace_index_digestsha256:7146385bdc49de2c0353dd91689c08a4a60b74083de39f2bd5a7e802efabf20a