MemoryATHENA QA release

This repository contains the frozen MemoryATHENA QA deployment used for the five-task paper-aligned evaluation. It packages the shared source-memory artifact, the three-reader expert checkpoint, and the router/reader checkpoint; the model card reports the downstream results without redistributing raw examples or labels.

Training data

The source memory is the imported 20M-token XMemTransfer memory artifact. The MemoryATHENA reader/router training uses the English Wikipedia-2021 causal next-token-only stream, with seed 42 and a 20M-token budget. The source memory is a dependency artifact, not retrained in this release. Training data are referenced by name only; they are not included here.

Test data and protocol

The evaluation uses the paper-aligned test protocols for Natural Questions (NQ), WebQuestions/WebQA, TriviaQA rc.nocontext validation, TruthfulQA multiple_choice, and HotpotQA distractor. Open-domain QA reports EM/F1; TruthfulQA reports MC1/MC2/MC3 and their arithmetic mean. Labels are used for final scoring only and are not included in this repository.

Headline results (percent)

Condition NQ EM/F1 WebQA EM/F1 TriviaQA EM/F1 TruthfulQA MC1/MC2/MC3/mean HotpotQA EM/F1
Engram-only 20.28/28.18 14.86/33.28 63.92/69.18 27.42/44.32/22.69/31.47 17.95/25.92
E-only path 20.20/28.28 14.96/33.35 64.05/69.34 26.93/44.18/22.64/31.25 18.19/26.04
Three-source router 22.72/33.02 17.86/34.60 62.78/70.68 28.15/44.04/23.02/31.74 15.62/26.34
Mistral to Llama transfer 20.37/29.98 18.06/36.40 60.35/67.39 27.78/41.57/21.87/30.41 16.00/25.17

Results are single-seed (seed 42). The standalone Engram-only row and joint-checkpoint rows have different training histories; the transfer row is reported separately. Pair-restricted fusion diagnostics are in the dataset release.

Files and links

No ARTIFACTS.md, raw datasets, raw test examples, labels, prediction lists, optimizer states, credentials, or internal filesystem paths are part of this release.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for OLAResearchX/memoryathena-qa-20260922

Finetuned
(447)
this model

Datasets used to train OLAResearchX/memoryathena-qa-20260922

Collection including OLAResearchX/memoryathena-qa-20260922