agent-memory-benchAMB / AMB
Leaderboard

Ranked by what ran.

Task success on executable checkers, paired per task against the claude_md baseline. Every number on this page traces to a published run directory; nothing is typed in by hand.

Adversarial by design

This benchmark is built to break a memory system, not to demo one. Sixty-five percent of the cells scored here are conditions where the corpus is outdated, contradictory, inapplicable, or simply empty, and the best available outcome is to waste nothing. A memory layer is only handed a clean, governing fact in 35% of them.

A benchmark where memory can only help is a demo. The numbers below are what survives a corpus that is mostly trying to mislead.

01

Products and controls

ranked when a run is official
rank arm integration task success Δ vs baseline 95% CI discarded cells total tokens cost / task

Deltas below the preregistered minimum effect are noise and are labeled as noise. Discarded cells are sessions that failed the admission gate: the treatment could not be proven applied, so the cell is thrown out rather than scored. Total tokens and cost are end-to-end, ingestion included; total tokens is shown for vendors so their measured spend is directly comparable.

02

Products, condition by condition

where a memory layer earns and where it costs

Only the products appear here. The controls exist to price the grid, not to be studied condition by condition. present is the one condition holding a clean governing fact; the other four are the traps, and a product's row across them is the honest picture of what it does when memory is a liability rather than an asset.

Solved cells over admitted cells. A held product publishes nothing here either: withholding a headline while publishing its breakdown would defeat the hold.

03

Reference tracks

diagnostics, never ranked

protocol is the one to read first, because it prices the question rather than any product. It carries the shared memory instruction with no memory behind it. Asking an agent to consult a memory it does not have is itself a treatment with its own effect, and it is an effect every memory product pays before it retrieves a single document. Read its Δ as the entry fee, and read each product's distance from it as what the product earned back.

track what it isolates task success Δ vs baseline

A reference track answers where the memory path breaks, not which product wins, so it runs in the grid and is never ranked in the table above. The decomposition is on the method page.

04

What makes a result official

or it does not appear here

A run appears on this page only when all of the following hold, in this order:

gaterequirement
1 · preregistered protocol committed under preregistration/ before the first session; the harness refuses to start while that directory is dirty
2 · announced the run is announced before it happens, not after it succeeds
3 · gated every scored cell passed the admission gate; discard counts published per arm
4 · published in full per-session logs, streams, admission verdicts and costs land in results/<run_id>/, wins and losses alike

The results/ tree holds smoke tests, pilots and diagnostics. They are bring-up, they are not previews of a ranking, and no number from them is quoted here. Where the benchmark stands is in docs/STATUS.md.