Broad coverage
Textual and coding memory tasks cover long conversations, changing user states, and reusable engineering experience.
Agent MemoryLeaderboard
Open · Fair · Reproducible
AML measures how well memory systems store, retrieve, and support the use of long-term information. Research methods and commercial products are compared under one unified, traceable evaluation contract.
The first official leaderboard is now live.
Leaderboard released
AMC / 2026
Agent Memory Challenge 2026
01 / Why AML
Memory systems are often reported with different datasets, answer models, and judges. AML brings public benchmarks, controlled evaluation data, a standard interface, unified answering, multi-agent judging, version review, and public governance into one long-running evaluation system.
Textual and coding memory tasks cover long conversations, changing user states, and reusable engineering experience.
Datasets, prompts, Answer models, judges, aggregation, and resource rules remain fixed within a release.
Capability-level profiles accompany the overall rank, making strengths and failure modes easier to understand.
02 / Unified evaluation
Candidates expose only two memory operations. The platform fixes generation and scoring, making it easier to attribute performance differences to memory storage and retrieval.
Read the evaluation documentation03 / Benchmark overview
Textual and coding memory use different tasks and primary metrics. Academic methods and commercial products are also published separately, so unlike systems are never mixed into one rank.
Measure what a system can reliably remember and use across long, evolving histories.
Measure whether historical engineering context helps an agent solve the software task in front of it.
04 / Capability profile
Textual results are mapped into seven shared dimensions. The profile reveals where a system succeeds, where it fails, and whether it respects changing state, evidence, and privacy boundaries.
05 / Open evaluation release
The public repository exposes per-benchmark evaluation contracts, shared runtime configuration, and documentation for methodological review. Protected evaluation data, participant traces, and production infrastructure remain private to preserve benchmark integrity.
Browse the public repositoryPublic Answer and scoring behavior can be reviewed against the reported release.
Comparable results bind to a benchmark bundle, pipeline revision, and model configuration.
Held-out questions, gold answers, participant traces, and private annotations are not distributed.
Every published row identifies the evaluated method, product, commit, image, or API version.
06 / Participant divisions
For papers, research projects, and open-source systems. Results bind to public source, configuration, and reproducibility evidence.
For stable hosted products and APIs. Internal implementation may remain closed, while the evaluated product and API version stay verifiable.
Agent Memory Challenge · Leaderboard released
Public rankings
Combined top-ten results from the two supplied workbooks, ranked together by total micro score. Division tags are shown inside each row for quick scanning.
Aggregate performance across the fixed textual evaluation suite.
All academic leaderboard results are evaluated using the same GPT-4o mini base model.
Textual and coding memory never share a ranking. Academic and commercial entries are also published separately.
Textual Memory ranks by Overall; Coding Memory ranks by Task Solve (%), with supporting detail retained.
Each public row identifies the evaluated source, product, image, commit, or API version and its review status.