Docs · Getting started

Glossary

The terms that appear on charts, badges and result pages.

Metrics

Recall@k
Share of questions whose correct evidence appears in the first k retrieved items.
nDCG@k
Like recall, but rewards placing the correct evidence near the top of the k items.
Judged accuracy
Share of answers a judge model grades as correct. Always reported with reader and judge.
p50 and p95
The typical query time and the time beaten by 95 of every 100 queries.
Reported interval
The uncertainty range a run supplies for its own metric, drawn as whiskers.
Full-context baseline
No memory system at all: the entire history is placed in the prompt.

Evidence

Bundle
The files that reproduce a number: traces, configs, build fingerprint, metrics and a script.
Trace
One question's record: what was stored, what was retrieved, and any answer.
Signed ledger
An append-only, hash-chained list of every attempt, including failures.
Held-out split
Data never used for tuning, kept apart so a score reflects unseen questions.

Status labels

Development
Code exists and runs on local fixtures. Not an admitted benchmark result.
Planned
A protocol is identified but adapters, pinned inputs or scoring are not built yet.
Operator entry
Mnemosyne, the system run by this site's operator, held to the same rules.
PBPP
The public-benchmark publication protocol that every public number must satisfy.