UMA-LedgerQA-4B / README.md
dp66's picture
docs: add Specialist model card
d8ab9a7 verified
|
Raw
History Blame Contribute Delete
4.16 kB
metadata
license: apache-2.0
language:
  - en
library_name: transformers
pipeline_tag: text-generation
base_model:
  - Qwen/Qwen3-4B-Instruct-2507
base_model_relation: finetune
tags:
  - memory-agent
  - reinforcement-learning
  - long-context
  - tool-use
  - state-tracking
  - ledger-qa
  - qwen3
  - grpo
arxiv: 2602.18493

UMA-LedgerQA-4B (Specialist)

UMA-LedgerQA-4B is the Specialist checkpoint of the Unified Memory Agent (UMA) introduced in Learning to Remember: End-to-End Training of Memory Agents for Long-Context Reasoning.

UMA is a tool-using memory agent that incrementally maintains a compact core summary and a structured key-value Memory Bank. The same policy performs memory construction and downstream question answering through explicit memory and retrieval operations.

Checkpoint Variant

This repository contains the Specialist UMA checkpoint adapted to Ledger-QA, the paper's controlled benchmark for persistent state tracking under accumulated updates, overwrites, and aggregation.

Property Value
Base model Qwen/Qwen3-4B-Instruct-2507
Parameters 4B
Weight format BF16 Safetensors
Training method End-to-end reinforcement learning with Task-Stratified GRPO
Specialization data Ledger-QA
Reported default context budget 16K

The Generalist and Specialist checkpoints share the same UMA architecture and tool interface. Use this checkpoint to reproduce or extend the Ledger-QA Specialist setting; use ICTNLP/UMA-4B for the broader Generalist setting.

Intended Use

This checkpoint is intended for research on:

  • persistent state tracking over long input streams;
  • updates, overwrites, contradiction handling, and aggregation;
  • proactive structured memory construction;
  • memory maintenance with explicit tool calls;
  • evaluation and extension of the UMA framework and Ledger-QA.

The checkpoint is designed to run inside the UMA two-phase agent loop. A plain text-generation call loads the language model, but does not by itself instantiate the Memory Bank, retrieval tools, prompts, or memory-to-QA workflow.

Loading the Weights

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "ICTNLP/UMA-LedgerQA-4B"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16,
    device_map="auto",
)

The model can also be served through an OpenAI-compatible inference server:

vllm serve ICTNLP/UMA-LedgerQA-4B \
  --max-model-len 16384 \
  --gpu-memory-utilization 0.8

For full memory-agent inference, including the Memory Bank, memory tools, embedding retrieval, prompts, Ledger-QA data preparation, and evaluation, follow the official repository.

Limitations

  • This checkpoint is specialized for Ledger-QA and should not be presented as the paper's Generalist checkpoint.
  • Ledger-QA is a controlled diagnostic benchmark rather than comprehensive real-world validation.
  • The checkpoint is a research model and may generate incorrect answers or perform incorrect memory updates.
  • Agent behavior depends on the UMA prompt templates, tool implementations, retrieval backend, chunking policy, and inference configuration.
  • Persistent-memory applications can involve sensitive information. Deployments should provide appropriate privacy controls, retention policies, and user oversight.
  • This checkpoint should not be used as the sole basis for high-stakes decisions.

Citation

@article{zhang2026learning,
  title   = {Learning to Remember: End-to-End Training of Memory Agents for Long-Context Reasoning},
  author  = {Zhang, Kehao and Gui, Shangtong and Yang, Sheng and Chen, Wei and Feng, Yang},
  journal = {arXiv preprint arXiv:2602.18493},
  year    = {2026}
}