Docs · Getting started
Glossary
The terms that appear on charts, badges and result pages.
Metrics
- Recall@k
- Share of questions whose correct evidence appears in the first k retrieved items.
- nDCG@k
- Like recall, but rewards placing the correct evidence near the top of the k items.
- Judged accuracy
- Share of answers a judge model grades as correct. Always reported with reader and judge.
- p50 and p95
- The typical query time and the time beaten by 95 of every 100 queries.
- Reported interval
- The uncertainty range a run supplies for its own metric, drawn as whiskers.
- Full-context baseline
- No memory system at all: the entire history is placed in the prompt.
Evidence
- Bundle
- The files that reproduce a number: traces, configs, build fingerprint, metrics and a script.
- Trace
- One question's record: what was stored, what was retrieved, and any answer.
- Signed ledger
- An append-only, hash-chained list of every attempt, including failures.
- Held-out split
- Data never used for tuning, kept apart so a score reflects unseen questions.
Status labels
- Development
- Code exists and runs on local fixtures. Not an admitted benchmark result.
- Planned
- A protocol is identified but adapters, pinned inputs or scoring are not built yet.
- Operator entry
- Mnemosyne, the system run by this site's operator, held to the same rules.
- PBPP
- The public-benchmark publication protocol that every public number must satisfy.