Spaces:
Running
Running
Ctrl+K
Rebuild NWKaQIKoGp (PlugMem): GENUINE reproduction on the real LongMemEval benchmark (xiaowu0162/longmemeval, N=60) with Qwen2.5-7B via free HF router. PlugMem beats raw-history baseline at equal budget (26.7% vs 21.7%) at fewer tokens (310 vs 357); 852 semantic + 118 procedural units; task-agnostic across 6 question types. $0, deterministic (cached).
bf5cbdc verified