UMA-LedgerQA-4B / README.md
dp66's picture
docs: add Specialist model card
d8ab9a7 verified
|
Raw
History Blame Contribute Delete
4.16 kB
---
license: apache-2.0
language:
- en
library_name: transformers
pipeline_tag: text-generation
base_model:
- Qwen/Qwen3-4B-Instruct-2507
base_model_relation: finetune
tags:
- memory-agent
- reinforcement-learning
- long-context
- tool-use
- state-tracking
- ledger-qa
- qwen3
- grpo
arxiv: 2602.18493
---
# UMA-LedgerQA-4B (Specialist)
**UMA-LedgerQA-4B** is the Specialist checkpoint of the Unified Memory Agent (UMA) introduced in [Learning to Remember: End-to-End Training of Memory Agents for Long-Context Reasoning](https://arxiv.org/abs/2602.18493).
UMA is a tool-using memory agent that incrementally maintains a compact core summary and a structured key-value Memory Bank. The same policy performs memory construction and downstream question answering through explicit memory and retrieval operations.
- **Code:** [github.com/ictnlp/unified-memory-agent](https://github.com/ictnlp/unified-memory-agent)
- **Paper:** [arXiv:2602.18493](https://arxiv.org/abs/2602.18493)
- **Generalist checkpoint:** [ICTNLP/UMA-4B](https://huggingface.co/ICTNLP/UMA-4B)
## Checkpoint Variant
This repository contains the **Specialist** UMA checkpoint adapted to Ledger-QA, the paper's controlled benchmark for persistent state tracking under accumulated updates, overwrites, and aggregation.
| Property | Value |
| --- | --- |
| Base model | `Qwen/Qwen3-4B-Instruct-2507` |
| Parameters | 4B |
| Weight format | BF16 Safetensors |
| Training method | End-to-end reinforcement learning with Task-Stratified GRPO |
| Specialization data | Ledger-QA |
| Reported default context budget | 16K |
The Generalist and Specialist checkpoints share the same UMA architecture and tool interface. Use this checkpoint to reproduce or extend the Ledger-QA Specialist setting; use `ICTNLP/UMA-4B` for the broader Generalist setting.
## Intended Use
This checkpoint is intended for research on:
- persistent state tracking over long input streams;
- updates, overwrites, contradiction handling, and aggregation;
- proactive structured memory construction;
- memory maintenance with explicit tool calls;
- evaluation and extension of the UMA framework and Ledger-QA.
The checkpoint is designed to run inside the UMA two-phase agent loop. A plain text-generation call loads the language model, but does not by itself instantiate the Memory Bank, retrieval tools, prompts, or memory-to-QA workflow.
## Loading the Weights
```python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "ICTNLP/UMA-LedgerQA-4B"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
device_map="auto",
)
```
The model can also be served through an OpenAI-compatible inference server:
```bash
vllm serve ICTNLP/UMA-LedgerQA-4B \
--max-model-len 16384 \
--gpu-memory-utilization 0.8
```
For full memory-agent inference, including the Memory Bank, memory tools, embedding retrieval, prompts, Ledger-QA data preparation, and evaluation, follow the [official repository](https://github.com/ictnlp/unified-memory-agent).
## Limitations
- This checkpoint is specialized for Ledger-QA and should not be presented as the paper's Generalist checkpoint.
- Ledger-QA is a controlled diagnostic benchmark rather than comprehensive real-world validation.
- The checkpoint is a research model and may generate incorrect answers or perform incorrect memory updates.
- Agent behavior depends on the UMA prompt templates, tool implementations, retrieval backend, chunking policy, and inference configuration.
- Persistent-memory applications can involve sensitive information. Deployments should provide appropriate privacy controls, retention policies, and user oversight.
- This checkpoint should not be used as the sole basis for high-stakes decisions.
## Citation
```bibtex
@article{zhang2026learning,
title = {Learning to Remember: End-to-End Training of Memory Agents for Long-Context Reasoning},
author = {Zhang, Kehao and Gui, Shangtong and Yang, Sheng and Chen, Wei and Feng, Yang},
journal = {arXiv preprint arXiv:2602.18493},
year = {2026}
}
```