Ranked by what ran.
Task success on executable checkers, paired per task against the
claude_md baseline. Every number on this page traces to a published run
directory; nothing is typed in by hand.
This benchmark is built to break a memory system, not to demo one. Sixty-five percent of the cells scored here are conditions where the corpus is outdated, contradictory, inapplicable, or simply empty, and the best available outcome is to waste nothing. A memory layer is only handed a clean, governing fact in 35% of them.
A benchmark where memory can only help is a demo. The numbers below are what survives a corpus that is mostly trying to mislead.
Products and controls
ranked when a run is official| rank | arm | integration | task success | Δ vs baseline | 95% CI | discarded cells | total tokens | cost / task |
|---|
Deltas below the preregistered minimum effect are noise and are labeled as noise. Discarded cells are sessions that failed the admission gate: the treatment could not be proven applied, so the cell is thrown out rather than scored. Total tokens and cost are end-to-end, ingestion included; total tokens is shown for vendors so their measured spend is directly comparable.
Products, condition by condition
where a memory layer earns and where it costsOnly the products appear here. The controls exist to price the grid,
not to be studied condition by condition. present is the one condition holding a
clean governing fact; the other four are the traps, and a product's row across them is the
honest picture of what it does when memory is a liability rather than an asset.
Solved cells over admitted cells. A held product publishes nothing here either: withholding a headline while publishing its breakdown would defeat the hold.
Reference tracks
diagnostics, never rankedprotocol is the one to read first, because it prices the
question rather than any product. It carries the shared memory instruction with
no memory behind it. Asking an agent to consult a memory it does not have is
itself a treatment with its own effect, and it is an effect every memory product pays before
it retrieves a single document. Read its Δ as the entry fee, and read each
product's distance from it as what the product earned back.
| track | what it isolates | task success | Δ vs baseline |
|---|
A reference track answers where the memory path breaks, not which product wins, so it runs in the grid and is never ranked in the table above. The decomposition is on the method page.
What makes a result official
or it does not appear hereA run appears on this page only when all of the following hold, in this order:
| gate | requirement |
|---|---|
| 1 · preregistered | protocol committed under preregistration/ before the first session; the harness refuses to start while that directory is dirty |
| 2 · announced | the run is announced before it happens, not after it succeeds |
| 3 · gated | every scored cell passed the admission gate; discard counts published per arm |
| 4 · published in full | per-session logs, streams, admission verdicts and costs land in results/<run_id>/, wins and losses alike |
The results/ tree holds smoke tests, pilots and
diagnostics. They are bring-up, they are not previews of a ranking, and no number from them
is quoted here. Where the benchmark stands is in
docs/STATUS.md.