Workspace Compare runsScope data
Official benchmark setup required. Current development checks and adapted results are not official benchmark runs. Official comparisons must use the upstream code, datasets, scoring and prescribed setup, with versions and deviations disclosed. Protocol status →

Compare memory evidence

Matching conditions. Visible limitations. Every run traceable.

These groups match declared benchmark, dataset, scoring, model and resource conditions. They are candidates for inspection, not certified rankings. Each row is one attempt; no overall winner or average is calculated. Reported intervals retain their original meaning and do not imply a confidence level.

No compatible result groups are available yet. Missing measurements are not zero scores. Explore planned benchmarks.

Outside comparison groups

Download comparison data and source IDs (JSON)

Dataset fingerprint: sha256:713fdcff46c8363d4ea5e39a904a6c3657db99835d15825082def55ebfbff718