Workspace Feature landscapeScope data
Official benchmark setup required. Current development checks and adapted results are not official benchmark runs. Official comparisons must use the upstream code, datasets, scoring and prescribed setup, with versions and deviations disclosed. Protocol status →

Memory is more than recall

What stands out in Mnemosyne, what other systems share, and what a benchmark score leaves unanswered.

Mnemosyne features · Benchmark coverage · What remains unproven

Primary-source review: 2026-10-04. This is a scoped comparison of documented designs and six evaluation projects, not an exhaustive survey or a measured ranking. “Not documented” does not mean “not supported”. No exclusive feature or best-system claim has been established.

What stands out in Mnemosyne

Evidence you can trace, rebuild and branch

Content-addressed evidence underlies rebuildable beliefs. Tenant memory can be branched, merged or discarded. Temporal queries preserve the distinction between current and historical beliefs.

Current evidence: Implemented interfaces and regression coverage; this is not a completed cross-system reliability benchmark. Mnemosyne source

Comparison: Graphiti also documents temporal history and provenance. Those are shared capabilities, not exclusive Mnemosyne features. Its documented graph design differs from Mnemosyne's evidence-ledger and branch/merge model. Graphiti

Memory for what to do next

TTL-bounded working memory can be promoted or expired. Signed, session-bound intentions support triggers and explicit recurrence policies.

Current evidence: Implemented and locally tested, with four-week recurrence regressions and an exact-time lateness diagnostic. M12 and M13 remain partial: registered multiweek benchmarks, calibrated controls, full trigger timing/cost and capacity evidence remain unfinished. Mnemosyne source

Comparison: Letta documents editable, shareable in-context memory blocks. MemOS documents asynchronous ingestion scheduling. Neither description by itself establishes the same intention-firing contract; this is not proof those products lack it. Letta memory blocks

Controlled learning and deletion

Consolidation uses a promotion gate and rollback; deletion has signed evidence and verification primitives. These controls make change auditable.

Current evidence: Implemented primitives, with deployment-specific evidence still required. Deletion does not establish unlearning in model weights or every external provider. Mnemosyne source

Comparison: Cognee documents session lessons, feedback and deletion; MemOS documents correction and reusable skills. Learning and forgetting are shared concerns. Their exact guarantees require contract-level tests, not a feature checklist. Cognee operations

Local memory with explicit access boundaries

Mnemosyne offers local SQLite and PostgreSQL paths, signed sessions, tenant boundaries and capability checks. BurnOS compatibility has dedicated tests.

Current evidence: Implemented paths and compatibility checks; production readiness still depends on the selected backend and deployment evidence. Mnemosyne source

Comparison: Mem0 offers open-source and managed variants; MemOS offers local and cloud deployments. Local operation is not unique. Supermemory's documented temporal graph API and HippoRAG's graph-assisted retrieval provide other design choices. MemOS deployment choices

Compare the combination and its guarantees, not a count of checkmarks. Read the source-linked profiles of all eight systems. Additional comparison sources: Mem0, MemOS, Supermemory, HippoRAG.

Where benchmark coverage stops

Existing benchmarks provide useful evidence, and broader suites are emerging. Our review does not establish that any one of the six projects below validates every requirement in Mnemosyne’s release plan. That is a coverage finding for this plan—not a claim that no solid memory benchmark exists.

BenchmarkDocumented focusAdditional evidence needed
LongMemEvalInformation extraction, multi-session reasoning, temporal reasoning, knowledge updates and abstention in conversational QA.A QA score does not by itself establish branch rollback, tenant isolation or scheduled-action correctness.
LoCoMoLong conversations with question answering, event summarization and multimodal dialogue-generation tasks; individual adapters may expose only QA.Report the task actually run. A QA-only adapter does not test the full release or establish erasure and operational guarantees.
MemoryAgentBenchIncremental interactions covering accurate retrieval, test-time learning, long-range understanding and conflict resolution (selective forgetting in the paper's terminology).Conflict-resolution scores do not establish physical erasure from every storage surface or model unlearning.
LoCoMo-PlusAdds a cognitive category: connecting a later trigger query to an earlier cue across conversations.Implicit recall in an answer is different from a persistent scheduler meeting deadlines, cancellation and recurrence contracts.
MemLensVisual and textual conversational memory across long contexts, including updates, temporal reasoning and answer refusal.Multimodal accuracy still needs separate access-control, provenance and resource envelope checks. Disclose full-dataset versus agent-subset evaluation.
OmniMemEvalA broad framework combining user-memory benchmarks and agent-task evaluation, including reasoning, information retrieval, knowledge work and coding.Broad task coverage is valuable. It does not automatically certify this project's specific invariants; map each release requirement to an actual test and artifact.

What a complete memory evaluation should establish

For this project: recall and grounded answers; change over time and contradictions; learning without regressions; working-memory limits and future intentions; provenance and verifiable deletion; isolation and poisoning resistance; confidence and abstention; multimodal fidelity; latency, cost and resource use; and reproducible results across supported deployments.

This checklist is derived from our traceability matrix, not a universal scientific definition of memory. Supplemental tests must preserve upstream benchmark scores, disclose limitations and apply the same rules to every entrant.

What we still have to prove

Mnemosyne has not completed this full evaluation. Real comparable runs, missing development coverage, deployment evidence and public reproduction remain open. Our benchmark framework is not proof of our own superiority.

Inspect the available results · Read the methods