Workspace Development evidenceScope data
Official benchmark setup required. Current development checks and adapted results are not official benchmark runs. Official comparisons must use the upstream code, datasets, scoring and prescribed setup, with versions and deviations disclosed. Protocol status →

DEVELOPMENT EVIDENCE

Experiments, including the failures.

Trace what ran, what failed, and what remains unknown. These operator-run development captures are not admitted leaderboard results.

No ranking or superiority claim. Format validity, stored-state accuracy and firing behavior are different checks. A completed diagnostic is not a completed benchmark. ZIPs include raw records, methodology and file hashes; hashes bind files but do not independently attest execution.

MNEMOSYNE · VERIFIED RETRIEVAL

A complete retrieval run is a starting point.

3,671 queries · 146 contexts · 6,795 captures. All 22 available task variants were included. Verification checked retained file hashes, ordered cases and each hit against its captured context.

566 queries returned no hits. These remain part of the result. The run used the recorded default policy; its resource budget is not matched to the official BM25 method. No generated-answer score follows from retrieval completion.

Download retrieval receipt — receipt only, not the full raw capture archive.

LOCAL GENERATION · FEASIBILITY ONLY

Real answers. Including incorrect ones.

A paired pilot sent one preselected case through a local Qwen3 1.7B model, once with Mnemosyne retrieval and once with BM25 retrieval. Both answers failed the official primary metric. The BM25 answer reached the unchanged 10-token output cap; the capped answer was retained and scored.

This establishes that the local path can execute that case. It does not establish a quality ranking, full-dataset feasibility or parity with the official GPT-4o-mini setup. Broader evaluation remains open.

Download pilot scoring receipt · How to interpret evidence

RETRIEVAL ONLY

Official BM25 method: all available task variants

22 task variants · 146 contexts · 3,671 queries completed.

The pinned official BM25 method retrieved context and prepared model requests. No answer model or judge ran. This is a baseline integration check, not a memory-system ranking or a completed official benchmark. The full upstream agent lifecycle remains unverified.

The index was rebuilt for each query, so elapsed time is not a fair latency comparison. Dependency versions, input and output hashes, source identity and deviations are in the receipt.

Download verification receipt. The receipt is not the full raw evidence: the 609 MB of prepared requests are retained locally. Their hashes alone do not independently prove execution or answer quality.

Completed paired development workload

Same workload, different transport

Persistent MCP stdio: 320/320 correct firings. CLI: 200/320, with 120 missed short windows. Both recovered every exact-time intention; no duplicate or cancelled firings. Programmed triggers only: this does not measure natural-language memory quality or competitor performance.

Completed development workload

Clocked trigger pressure

Five seeds, 320 live intentions: 207 fired correctly and 113 short-window triggers were missed. All 160 exact-time intentions recovered; no duplicate or cancelled firings. This measures the public CLI path, not intrinsic engine capacity.

Failed full-corpus attempt

Dependency formation

1 of 100 cases completed. The next case failed on task-ID reuse; 98 were not attempted. The completed case missed one expected firing. No full-corpus score.

Completed development diagnostic

JSON versus schema

22 first-turn responses across 11 inputs. All 11 plain-JSON outputs failed format checks; all 11 schema outputs passed them. Semantic failures remain. No tasks were executed.

Completed development diagnostic

Prompt semantics

22 first-turn responses across 11 inputs. Added instructions corrected one recurrence proposal, but event and condition errors remained and an unrelated task was introduced. No overall improvement established.