README / README.md
AgentMemoryL's picture
Update README.md
690882d verified
|
Raw
History Blame Contribute Delete
2.41 kB
metadata
title: Agent Memory Leaderboard
emoji: 🧠
colorFrom: blue
colorTo: green
sdk: static
pinned: false
short_description: A unified evaluation platform for memory systems.

Agent Memory Leaderboard · 记忆之巅

A unified evaluation platform for long-term memory systems and memory-enabled agents.

Agent Memory Leaderboard (a.k.a. 记忆之巅) evaluates how effectively memory systems store, retrieve, and use information across long conversations, persistent user contexts, and memory-intensive agent tasks.

Official evaluation and submission are conducted exclusively through the Agent Memory Leaderboard website. This Hugging Face organization hosts public leaderboard releases, evaluation documentation, and community updates.

What We Evaluate

  • Long-term memory storage and retrieval
  • Long-context and multi-session understanding
  • Personalized and user-specific memory
  • Temporal and event-ordering reasoning

Evaluation Tracks

Track Intended for
Academic Methods Reproducible research systems, open-source methods, and academic implementations
Industry Systems Production APIs, hosted services, commercial systems, and closed-source products

All systems are evaluated under a versioned evaluation contract with fixed datasets, prompts, answer models, judge configurations, and pipeline hashes.

First Public Release

The inaugural public leaderboard is scheduled for mid-August 2026.

Results published on Hugging Face are official release snapshots. The live and authoritative leaderboard remains on the Agent Memory Leaderboard website.

Links

On Hugging Face

This organization will publish:

  • Versioned leaderboard result datasets
  • Public leaderboard Spaces
  • Benchmark and methodology documentation
  • Baselines and reproducibility materials
  • Release announcements and technical analyses

Follow this organization for the inaugural leaderboard release and future evaluation cycles.