--- language: - en library_name: peft license: other license_name: memfold-and-upstream-model-licenses license_link: https://huggingface.co/Johnny221B/memfold/blob/main/LICENSE.md base_model: - Qwen/Qwen3-4B - Qwen/Qwen2.5-3B-Instruct - Qwen/Qwen2.5-7B-Instruct tags: - memfold - lora - long-context - memory inference: false --- # MemFold **Learning Compact Soft Memory for Long-Context Personalization via On-Policy Optimization** [![Code](https://img.shields.io/badge/GitHub-Code-181717?logo=github)](https://github.com/Johnny221B/memfold) [![Project](https://img.shields.io/badge/Project-Website-blue)](https://memfold.github.io/) [![Paper](https://img.shields.io/badge/arXiv-Paper-B31B1B?logo=arxiv)](https://drive.google.com/file/d/1WgRcUQ7mfzxNVd74B5kLYAUUQzVIea5v/view) This repository contains MemFold's selected **reader LoRA adapters and matching memory components** for Qwen3-4B, Qwen2.5-3B-Instruct, and Qwen2.5-7B-Instruct. The implementation and training instructions are in [GitHub](https://github.com/Johnny221B/memfold). ## Updated Qwen3-4B checkpoint · 2026-09-29 The default **PersonaMem-32K / Qwen3-4B** reader is now the validation-selected **more-grpo, epoch 5 (615 steps)** checkpoint. Its matching compressor is unchanged. Other model bundles are unchanged. | Split | Five-trial mean | Best trial | |---|---:|---:| | Validation | 75.2% | 86.0% | | Test | 76.0% | 82.0% | Both memory generation and answer decoding are greedy; trials use different fixed option permutations. **86% is a validation score, not the test score.** Checkpoint selection used validation mean. Configuration and all trial scores are in [OPTIMIZATION.json](OPTIMIZATION.json). Download `personamem-32k/qwen3-4b/` from the default branch for this updated reader. To reproduce the original paper release, set `revision="d39842eeca8132288c5e2814fdd66733fa92fc3d"` when downloading. The previous weights remain available at [that fixed revision](https://huggingface.co/Johnny221B/memfold/tree/d39842eeca8132288c5e2814fdd66733fa92fc3d). This post-release update does not replace historical paper results. ## Updated Qwen2.5-3B-Instruct checkpoint The default `personamem-32k/qwen2.5-3b/` reader is now **current-ratio-epoch-4**, selected by five-trial validation mean. Only the final OPD + GRPO phase was retrained; its compressor is unchanged. | Split | Mean accuracy | Best trial | |---|---:|---:| | Previous reader validation | 60.0% | 70.0% | | Updated reader validation | 67.6% | 76.0% | | Updated reader test | 61.2% | 66.0% | Memory generation and answers use greedy decoding; trials change answer-option order. Validation best scores are not test scores. The test split was evaluated after checkpoint selection. [Settings and all validation results](personamem-32k/qwen2.5-3b/optimization.json). Previous weights remain at revision `10e71a32171ccde6f3a53bc8ee5b407264321c3d`. This is a post-release update, not a replacement for historical paper results. ## Updated Qwen2.5-7B-Instruct checkpoint The default `personamem-32k/qwen2.5-7b/` reader is now **current-ratio-epoch-5**, selected by five-trial validation mean. Only the final OPD + GRPO phase was retrained; its compressor is unchanged. | Split | Mean accuracy | Best trial | |---|---:|---:| | Previous reader validation | 72.8% | 76.0% | | Updated reader validation | 74.4% | 78.0% | | Updated reader test | 80.0% | 86.0% | Memory generation and answers use greedy decoding; trials change answer-option order. Validation best scores are not test scores. The test split was evaluated after checkpoint selection. [Settings and all validation results](personamem-32k/qwen2.5-7b/optimization.json). Previous weights remain at revision `c86ee86e01369478b300007d7500a1a63ea582b2`. This is a post-release update, not a replacement for historical paper results. ## Updated PersonaMem-128K / Qwen3-4B Default reader: **highest-lr-step-1000**. Only the final OPD + GRPO phase was retrained; its compressor is unchanged. Selected by validation mean, with test evaluated afterward. | Split | Mean accuracy | Best trial | |---|---:|---:| | Previous reader validation | 87.03% | 89.74% | | Updated reader validation | 89.16% | 91.94% | | Updated reader test | 87.64% | 89.70% | 273 validation questions and 233 test questions; greedy writer and reader, five fixed option permutations. Use the [matching 128K code and protocol](https://github.com/Johnny221B/memfold/tree/c242a424bc98f2cb11e9d456f60ccd1e2158e74a) for YaRN4.5, full inputs without truncation, and the historical answer-format normalization. [Full settings and results](personamem-128k/qwen3-4b/optimization.json). Previous weights remain at revision `7de39f1557b44a3cdac0dec768f31e62e0c2180f`. These post-release results do not replace historical paper results. ## Updated PersonaMem-128K / Qwen2.5-3B-Instruct Default reader: **current-ratio-step-600**. Only the final OPD + GRPO phase was retrained; its compressor is unchanged. Selected by validation mean, with test evaluated afterward. | Split | Mean accuracy | Best trial | |---|---:|---:| | Previous reader validation | 89.01% | 91.58% | | Updated reader validation | 89.23% | 90.84% | | Updated reader test | 87.64% | 89.27% | 273 validation questions and 233 test questions; greedy writer and reader, five fixed option permutations. Use the [matching 128K code and protocol](https://github.com/Johnny221B/memfold/tree/c242a424bc98f2cb11e9d456f60ccd1e2158e74a) for YaRN4.5, full inputs without truncation, and the historical answer-format normalization. [Full settings and results](personamem-128k/qwen2.5-3b/optimization.json). Previous weights remain at revision `71a718d07b7c61b3fdf53509557e05571368ba8d`. These post-release results do not replace historical paper results. ## Components and names MemFold uses one workflow: ```text History + query → Writer → Text memory → Compressor → Soft tokens → Reader → Answer ``` - **Writer** generates textual memory. PersonaMem uses the selected reader adapter for this step; the LoCoMo transfer configuration uses the shared Qwen3-4B writer. - **Compressor** converts memory into soft tokens using a resampler and projector. PersonaMem stores this component in `bridge.pt`. - **Reader** is the Qwen backbone with a MemFold LoRA adapter, stored in `reader/`. Bundles are named by **training dataset / backbone**. Display names follow **MemFold · Backbone · Dataset**; they do not contain research-question or run numbers. Existing filenames and configuration keys remain compatible with the published loader. ## Available weights Each dataset directory contains `qwen3-4b`, `qwen2.5-3b`, and `qwen2.5-7b`. | Bundle path | Contents | Soft-memory budget | |---|---|---| | `personamem-32k//` | `reader/` + `bridge.pt` | 256 tokens | | `personamem-128k//` | `reader/` + `bridge.pt` | 256 tokens | | `locomo//` | `reader/` + `compressor.pt` + `mapper.pt` | 512 tokens per session | The LoCoMo-trained bundles are used for LongMemEval transfer with soft memory plus bounded text. Download `shared/locomo-writer-qwen3-4b/` for their shared memory writer. The budget is per session, not per full history. The session compressor and mapper must stay paired. `qwen2.5-3b` and `qwen2.5-7b` refer to the **Instruct** backbones. Each `bundle.json` records the exact base-model ID and revision. The release includes nine reader bundles and their inference dependencies; base-model weights and benchmark inputs are obtained separately. ## Quick start ### Install the code Use Python 3.11 and a CUDA-compatible PyTorch installation. This documentation targets code commit `d049fd293e3232358594cba7cdae0f7a138bb58d`. ```bash git clone https://github.com/Johnny221B/memfold.git cd memfold git checkout d049fd293e3232358594cba7cdae0f7a138bb58d pip install -e '.[inference]' ``` ### Download a bundle Run from the cloned repository: ```python from huggingface_hub import HfApi, snapshot_download repo_id = "Johnny221B/memfold" revision = HfApi().model_info(repo_id).sha snapshot_download( repo_id=repo_id, revision=revision, allow_patterns=[ "personamem-32k/qwen3-4b/*", "load_components.py", "manifest.json", "VALIDATION.json", "LICENSE.md", "NOTICE", "licenses/*", ], local_dir="checkpoints/memfold", ) print("HF revision:", revision) ``` Download the base model at the revision specified in `bundle.json`, using its matching tokenizer. ### Load memory components ```bash python checkpoints/memfold/load_components.py \ --bundle checkpoints/memfold/personamem-32k/qwen3-4b --code . ``` This checks component loading on CPU. The helper also exposes `load_reader(bundle, **model_kwargs)` for PEFT loading. Answer generation needs the complete memory pipeline below. ### Generate memory and evaluate Prepare the benchmark questions and writer inputs using the [data guide](https://github.com/Johnny221B/memfold/blob/d049fd293e3232358594cba7cdae0f7a138bb58d/docs/training.md). Then run the common inference example: ```bash export MODEL=/path/to/Qwen3-4B export BUNDLE="$PWD/checkpoints/memfold/personamem-32k/qwen3-4b" export QUESTIONS=/path/to/test/questions.jsonl export WRITER_INPUTS=/path/to/test/writer-inputs.jsonl export OUTPUT="$PWD/outputs/personamem-32k" bash examples/evaluate.sh ``` The example calls `memfold.py generate`, `memfold.py encode`, and `memfold.py evaluate`. It writes individual predictions and a summary containing each trial's accuracy, the mean, and the best result. PersonaMem uses greedy decoding; the five trials change answer-option order. PersonaMem-128K uses the same commands with its matching bundle and inputs; set `EXAMPLES=233` and a supported `MAX_MODEL_LEN` for the longer context. PrefEval and LoCoMo/LongMemEval use the [dataset-specific adapters](https://github.com/Johnny221B/memfold/blob/d049fd293e3232358594cba7cdae0f7a138bb58d/docs/training.md#dataset-specific-recipes). ## Training Use the same code entrypoint for both PersonaMem context lengths: ```text prepare → extract → train writer → train compressor → train warmup → train reasoning → train reader → train optimize ``` For example, `python memfold.py train reader --help` shows reader-initialization arguments. The optimization objective combines OPD and GRPO; reference KL defaults to zero. The [training guide](https://github.com/Johnny221B/memfold/blob/d049fd293e3232358594cba7cdae0f7a138bb58d/docs/training.md) defines each procedure and its trainable components. ## Verification [manifest.json](manifest.json) records artifact hashes and the matching code revision. [VALIDATION.json](VALIDATION.json) distinguishes the original CPU checks from the Qwen3-4B / PersonaMem-32K GPU reproduction and the subsequent code-refactor regression. These checks do not constitute a fresh benchmark of all released models and datasets. The documentation update preserves all adapter and memory-component weight files. Historical training settings can differ from the current code defaults. ## License MemFold-authored components and helper code use MIT, subject to the upstream model terms. Qwen3-4B and Qwen2.5-7B-Instruct use Apache-2.0. **Qwen2.5-3B-Instruct uses the Qwen Research License, including its non-commercial restriction.** See [LICENSE.md](LICENSE.md), [NOTICE](NOTICE), and the included license texts.