sourcedrecall notes models (English, int4 ONNX)

Two small models for one job: given one conversation, write the lasting facts the named person states about themselves, one short sentence per line, or NONE. sourcedrecall runs them inside its own process on onnxruntime-genai, on the CPU, with no server.

file base download per conversation (4 CPU threads) peak memory
sourcedrecall-notes-en-0.6b.tar.gz (small) Qwen3-0.6B 0.51 GB 2-3 s 1.9 GB
sourcedrecall-notes-en-1.7b.tar.gz (standard) Qwen3-1.7B 1.3 GB about 4.5 s 3.3 GB

They are optional. sourcedrecall stores and finds the user's own words without them; with one, short notes are added, each kept only if its words are in what the user wrote, and shown labelled as written by a model.

pip install 'sourcedrecall[notes]'
sourcedrecall-setup --notes            # small
sourcedrecall-setup --notes standard   # standard

How they were made

  • Full fine-tune of the base model with mlx-lm, then int4 ONNX with the onnxruntime-genai builder (k_quant_mixed).
  • Teacher: Qwen3-14B wrote notes for 3,000 LongMemEval-S conversations (MIT) and 1,500 SODA dialogues (CC BY 4.0), each SODA dialogue from both speakers' sides; a note was kept only if most of its words are in the person's own lines.
  • Every training conversation got a random owner name and date format, and its pronouns were made explicit (I/my -> the speaker, you/your -> the listener), as sourcedrecall does before the model reads a conversation.
  • Scripts: tools/distil in the sourcedrecall repository.

Measured (local Qwen3-14B reader and judge for every system)

LoCoMo, conversations 2-9 (1,307 questions):

overall multi-hop temporal single-hop
sourcedrecall, no notes 64.7% 53.6% 56.2% 74.8%
with the small model's notes 68.2% 55.2% 58.9% 79.0%
with the standard model's notes 70.3% 59.0% 57.4% 81.6%
with Qwen3-14B's notes (a server) 73.1% 62.3% 57.4% 85.4%
Mem0 2.2.1 64.6% 61.9% 37.2% 77.9%

Notes on the LoCoMo development conversations, judged against their conversation by Qwen3-14B: said by the person 80% (small), 87% (standard), 93% (Qwen3-14B); gave the person the other speaker's facts 13%, 9%, 3%. Held-out false-memory sets: v3 0 of 30 for both; v4, an old state given as current, 3 of 44 (small) and 4 of 44 (standard), against 3 with Qwen3-14B's notes and 2 without notes.

These answer scores come from local models and are not comparable with published LoCoMo results, which use GPT-4-class models.

Limits

English only. They sometimes give the user something the other person said ("has kids" from the other person's children). Neither closes sourcedrecall's multi-hop gap to Mem0.

Attribution

SODA (Kim et al., 2023), CC BY 4.0. LongMemEval (Wu et al., 2024), MIT. Qwen3 (Qwen team), Apache-2.0.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for synthiumjp/sourcedrecall-notes-en

Finetuned
Qwen/Qwen3-0.6B
Finetuned
(1342)
this model