Open · Fair · Reproducible

One rulebook for
memory.

AML measures how well memory systems store, retrieve, and support the use of long-term information. Research methods and commercial products are compared under one unified, traceable evaluation contract.

Evaluation portal

The first official leaderboard is now live.

Leaderboard released AMC / 2026

Agent Memory Challenge 2026

The first leaderboard
is now live

02Evaluation types 10+Text datasets ≈5KText questions 04Evaluation stages

01 / Why AML

A benchmark system,
not another isolated score.

Memory systems are often reported with different datasets, answer models, and judges. AML brings public benchmarks, controlled evaluation data, a standard interface, unified answering, multi-agent judging, version review, and public governance into one long-running evaluation system.

01

Broad coverage

Textual and coding memory tasks cover long conversations, changing user states, and reusable engineering experience.

02

Controlled comparison

Datasets, prompts, Answer models, judges, aggregation, and resource rules remain fixed within a release.

03

Actionable diagnosis

Capability-level profiles accompany the overall rank, making strengths and failure modes easier to understand.

02 / Unified evaluation

Your system owns memory.
The platform controls the rest.

Candidates expose only two memory operations. The platform fixes generation and scoring, making it easier to attribute performance differences to memory storage and retrieval.

Read the evaluation documentation
  1. 01 / StoreAddCandidate system
  2. 02 / RetrieveSearchCandidate system
  3. 03 / GenerateAnswerOfficial platform
  4. 04 / ScoreEvalOfficial platform

03 / Benchmark overview

Two memories,
two separate leaderboards.

Textual and coding memory use different tasks and primary metrics. Academic methods and commercial products are also published separately, so unlike systems are never mixed into one rank.

01 Textual

Textual Memory

Measure what a system can reliably remember and use across long, evolving histories.

≈150Mhistory characters 7capability dimensions
  • LoCoMo-Refined
  • ScriptMem
  • LongMemEval
  • CLBench
  • PersonaMem
  • BEAM
02 Coding

Coding Memory

Measure whether historical engineering context helps an agent solve the software task in front of it.

12source repositories 1,290annotated history tasks
  • 01Debug MemoryDiagnosis, fixes, and validation
  • 02Development MemoryArchitecture, patterns, and conventions

04 / Capability profile

Explain the rank,
not only the winner.

Textual results are mapped into seven shared dimensions. The profile reveals where a system succeeds, where it fails, and whether it respects changing state, evidence, and privacy boundaries.

01Recall evidenceRetrieve
AExplicit Fact RecallFacts, attributes, sources, entities
BCompositional InferenceRelations and multi-hop evidence
CTemporal & Event ReasoningDates, order, and state changes
02Maintain memoryAdapt
DMemory GovernanceUpdate, conflict, deletion, forgetting
EPersonalization & CarePreferences and sensitive context
03Safe executionAct
GContext Learning & ExecutionRules, procedures, and constraints
HSafety & PrivacyAbstention and evidence boundaries

05 / Open evaluation release

AML-memory / agent-memory-leaderboard

Inspect the contract behind every comparable score.

The public repository exposes per-benchmark evaluation contracts, shared runtime configuration, and documentation for methodological review. Protected evaluation data, participant traces, and production infrastructure remain private to preserve benchmark integrity.

Browse the public repository
01Inspectable

Public Answer and scoring behavior can be reviewed against the reported release.

02Versioned

Comparable results bind to a benchmark bundle, pipeline revision, and model configuration.

03Protected

Held-out questions, gold answers, participant traces, and private annotations are not distributed.

04Traceable

Every published row identifies the evaluated method, product, commit, image, or API version.

06 / Participant divisions

Open methods and commercial products,
evaluated without false equivalence.

A / Academic

Academic Methods

For papers, research projects, and open-source systems. Results bind to public source, configuration, and reproducibility evidence.

C / Commercial

Commercial Products

For stable hosted products and APIs. Internal implementation may remain closed, while the evaluated product and API version stay verifiable.

Agent Memory Challenge · Leaderboard released

Memory decides.

First verified leaderboard expected August 12, 2026.
Visit the challenge