How to read the evidence
A score is useful only when you can inspect how it was produced.
Why this project exists: the memory benchmark gap
Retrieval is not answer quality
Retrieval recall measures how much relevant evidence a system found within a stated result limit. It does not show that a generated answer was correct. Answer quality, security, calibration, latency and cost are separate metric families.
Compare like with like
Compare runs only when their dataset version, split, protocol, model and resource budgets support that comparison. An interval shows the uncertainty supplied by the run; its confidence level is not inferred. Missing intervals and missing measurements are not zero. This site does not calculate a universal winner across different tasks.
Inspect a run
Open a result to inspect its build and configuration digests, publication status and individual question traces. Trace pages show only stored evidence, retrieved evidence and answers actually disclosed in the source. Development results are not public benchmark rankings.
Publication requires more than rendering
A local preview may contain non-publishable development records. Rendering does not approve publication. Public release requires the signed ledger, registered experiment, reproducible bundle, permitted assets and governance evidence to pass the separate release gates.
About the name
Mnemetric (neh-MET-rik) combines memory with measurement. The platform hosts multiple benchmarks; our own suite is the Mnemetric Whole-Memory Benchmark. Each upstream benchmark retains its own name and methods.
Who operates this site
Mnemosyne is the operator entry. The project must run supported competitors under the same disclosed protocol, retain failed attempts, and explain missing systems. Operator-run does not mean independent or neutral evaluation.