Docs · Getting started
Reading results
A score is useful only when you can see how it was produced. This page explains the charts, intervals and labels used across the site.
Recall and accuracy are different
Was the right memory among the first k returned? Deterministic: no language model is involved, so the same bundle always gives the same number.
Did a language model answer correctly from what was retrieved, as graded by another model? It depends on the reader and the judge, which must be named.
The two are never placed on one axis. A low recall and a high accuracy can both be true for the same system, because they measure different stages.
What the bar styles mean
Intervals and ties
Whiskers show the range a run reported for itself. If two systems’ ranges overlap, call it a tie.
Compare like with like
Two numbers only compare if the dataset, split, models and budgets match. The Compare runs page groups runs that match and refuses to rank across groups.
Inspect a run
Open any result to see its digests, publication status and per-question traces.