oif-instok-rl / README.md
MagicLuke's picture
model card
e7512bd verified
|
Raw
History Blame Contribute Delete
12.3 kB
---
license: other
license_name: personaplex-derivative
base_model: nvidia/personaplex-7b-v1
library_name: moshi
tags: [full-duplex, speech, instruction-following, lora, moshi, personaplex, rl, gdpo]
extra_gated_prompt: >-
These adapters are GDPO post-training of PersonaPlex. `gdpo-v2f-r4` was trained on rollouts
seeded from Fisher (LDC); `gdpo-duplex` was not. Neither corpus is redistributed here. By
requesting access you agree not to attempt to reconstruct or redistribute the underlying
corpora.
extra_gated_fields:
Name: text
Affiliation: text
Intended use: text
---
# oif β€” online-instruction-following RL checkpoints
GDPO post-training of **[nvidia/personaplex-7b-v1](https://huggingface.co/nvidia/personaplex-7b-v1)**
for **instructions injected mid-conversation** by a separate text LLM. A directive is spliced
into the model's inner-monologue stream as extra timesteps (`ins_token` mode) with the audio
held silent, so the model **reads** the directive and acts on its next turn without speaking it.
RL optimises the free-run regime directly β€” it trains on what the model actually produces β€”
which is why these checkpoints, unlike the SFT line, are usable as conversational models.
Project: [MagicLuke/duplex-online-if](https://github.com/MagicLuke/duplex-online-if) Β·
lab notes: `docs/lab/rl-R1-gdpo-duplex.md`, `docs/lab/rl-R7-v2f-r4.md` Β·
SFT line: [`MagicLuke/oif-instok-sft`](https://huggingface.co/MagicLuke/oif-instok-sft)
## Contents
One subfolder per published checkpoint, `<run_tag>/v<step>/`.
| run | step | pool | reward | what it is |
|---|---|---|---|---|
| `gdpo-duplex` | 1200 | conv-synth goal arcs | v1 (4-term) | **Highest task-completion checkpoint we have.** Beats gpt-realtime-2.1 on IF-Bench completion |
| `gdpo-v2f-r4` | 600 | Fisher-only, goalless persona | v2 (8-dim, VAD) | **Best turn-taking checkpoint we have** on FDB v1.0 |
Of 60 launched recipes in this project, these two and one SFT checkpoint are the **only three**
with downstream benchmark numbers. Everything else was scored on in-house metrics only, and is
not published.
**These are not a curve.** Both were initialised from the *base* native checkpoint, never from
an SFT adapter, and they are alternative recipes rather than successive steps. Do not stack
them, and do not apply either on top of `oif-instok-sft`.
## Results
### IF-Bench v1.3 β€” 200 tasks, judge `Qwen/Qwen3.8-27B`, examinee `fdb-assistant-unaware`
"brain on" = steered live by a second LLM; "brain off" = the same weights unsteered. Completion
is verified stage completion (claimed% in parentheses); seat-flip and order are v1.1 metrics.
| examiner | topic 1–5 | completion | either-side | order | seat-flip ↓ | naturalness |
|---|---|---|---|---|---|---|
| **`gdpo-duplex` v1200** + brain | 4.80 | **85.8%** (86.4) | 91.4% | 47.5% | 42.5% | 3.81 |
| **`gdpo-duplex` v1200** brain off | 4.59 | **61.8%** (62.3) | 70.4% | 9.5% | 40.5% | 3.83 |
| **`gdpo-v2f-r4` v600** + brain | 4.33 | 57.9% (58.8) | 69.0% | 10.0% | 66.5% | 3.68 |
| **`gdpo-v2f-r4` v600** brain off | 4.02 | 36.8% (37.4) | 48.9% | 0.0% | 59.5% | 3.96 |
| *gpt-realtime-2.1* (reference) | 4.72 | 73.0% | 80.6% | 32.5% | 47.5% | 3.61 |
| *`oif-instok-sft` stage2-337* + brain | 3.79 | 49.1% | 58.4% | 9.5% | **4.6%** | 4.05 |
| *PersonaPlex-7B base* | 2.19 | 7.6% | 10.6% | 0.0% | 21.5% | 4.45 |
v600's either-side numbers are the **corrected** ones (2026-08-25). A dead judge shard had
scored 18 of 50 Safety dialogues as `[null, null, null]`, and the scorer counts an unparsed
verdict as "no" rather than excluding it, while `report.json` reported `missing: []`. That cost
the split 25 points and produced a published either-side figure *below* the strict one β€” which
is impossible, since either-side is a union. If you rerun the judge, check for null votes.
### FDB v1.0 β€” 727 stimuli, seed 42424242
**The two rows ran under different conditions and are not comparable to each other**: v1200
under its training card + voice, v600 under VARF0 + the benchmark's own prompts.
| checkpoint | pause TOR syn/candor ↓ | backchannel TOR/freq/JSD | smooth TOR/lat | interrupt TOR/rating/lat |
|---|---|---|---|---|
| `gdpo-v2f-r4` v600 (VARF0, FDB prompts) | **0.365 / 0.417** | 0.327 / 0.118 / 0.724 | 1.000 / 0.203 | 1.000 / **4.725** / 0.268 |
| `gdpo-duplex` v1200 (training card + voice) | 0.971 / 0.991 | 0.691 / 0.062 / 0.830 | 1.000 / 0.000 | 1.000 / 2.180 / 0.299 |
| *PersonaPlex (ours, VARF0)* | 0.642 / 0.718 | 0.418 / 0.176 / 0.700 | 0.983 / 0.020 | 0.935 / 4.684 / 0.214 |
Interruption ratings come from a local gemma-4-E4B judge served as `gpt-4-turbo`, comparable
only across our own rows.
**v1200's pause TOR of 0.971/0.991 is the worst row in that table** β€” it takes the floor during
a pause almost every time. That is the cost of the reward it was trained on: reward v1 has no
term paying the model to *stop* talking. If you need turn-taking discipline, v600 is the
checkpoint; if you need task completion, v1200 is.
### Training-data overlap with the benchmark β€” measured, and it does not move the score
**Disclosure:** conv-synth v1 was seeded from the 200 FDB-v2 staged tasks and split *by
scenario* into 180 train / 10 val / 10 test. Those 200 scenarios **are** the 200 IF-Bench task
ids. So `gdpo-duplex` (conv-synth pool) trained on scenarios matching 180 of the 200 benchmark
tasks. `gdpo-v2f-r4` did not β€” its pool is Fisher-only β€” which makes it the control.
We measured the gap rather than assuming it. Verified completion on the 20 genuinely held-out
scenarios minus the 180 trained ones, within each arm, bootstrap CI over tasks (n=20 vs 180):
| checkpoint | arm | held-out | trained | gap | 95% CI |
|---|---|---|---|---|---|
| `gdpo-duplex` v1200 | brain off | 66.2% | 61.3% | **+5.0** | [βˆ’7.2, +17.1] |
| `gdpo-duplex` v1200 | brain on | 88.8% | 85.4% | **+3.3** | [βˆ’3.9, +9.9] |
| `gdpo-v2f-r4` v600 *(control)* | brain off | 36.2% | 36.8% | βˆ’0.6 | [βˆ’11.8, +10.7] |
| `gdpo-v2f-r4` v600 *(control)* | brain on | 47.5% | 59.0% | βˆ’11.5 | [βˆ’25.3, +2.2] |
| *stage2-337 (SFT, conv-synth)* | brain on | 50.0% | 49.9% | +0.1 | [βˆ’13.0, +13.4] |
**No memorization advantage is detectable.** Every CI spans zero, and the direction is the
opposite of contamination: the two conv-synth-trained models score *higher* on scenarios they
never saw, while the Fisher-only control scores lower. The honest reading is that the overlap is
real in provenance but does not show up in the score.
**This test has limited power** β€” 20 held-out scenarios gives roughly Β±10 points, so it rules
out a large memorization effect, not a small one. The held-out ids are enumerated in
`fdb-v2-IF-bench/benchmark/oif_splits/convsynth_overlap.json`; score them separately with
`slurm/arm_compare.py --subsets` and report both halves.
## Usage
Fold the adapter into NVIDIA's base with the tested recipe from
[`duplex-brain`](https://github.com/MagicLuke/duplex-brain):
```bash
huggingface-cli download MagicLuke/oif-instok-rl --include "gdpo-duplex/v1200/*" --local-dir adapter
python -m duplex_brain.ckpt.build_checkpoint \
--hf-repo nvidia/personaplex-7b-v1 \
--adapter adapter/gdpo-duplex/v1200/lora.safetensors --scaling 2.0 \
--out ckpt/duplex-brain-gdpo1200.safetensors
```
`--scaling 2.0` is the run's `lora_scaling`, applied directly, **not** divided by rank. The
build writes `<stem>.depq.json` and `<stem>.vocab.json` beside the weights β€” **keep all three
files together.** The loader reads the architecture from those sidecars, and without them it
silently expands 8 β†’ 16 depth heads or aliases `<instr_start>` onto BOS: both load, run, and are
wrong. Ids β‰₯ 32000 are not SentencePiece pieces β€” never round-trip them through the tokenizer.
Checksums (the fold logs recorded these before the benchmark ran):
| file | bytes | sha256 |
|---|---|---|
| `gdpo-duplex/v1200/lora.safetensors` | 609,897,768 | `d71f286efc142150bfdac17ea15567f7e3ee68620aa8171bb544107be4bb2e7c` |
| `gdpo-v2f-r4/v600/lora.safetensors` | 609,897,768 | `1f1581a7f69a5e989b2886ee89f18270811d463a3bc41b78037e4c31c80305d2` |
## Recipe
**Model** β€” PersonaPlex-7B native, `dep_q` sliced 16 β†’ 8, text vocab extended 32000 β†’ 32002 for
`<instr_start>` (32000) / `<instr_end>` (32001), `context 3000` (240 s @ 12.5 Hz). The adapters
do **not** load onto stock PersonaPlex unchanged.
**LoRA** β€” rank **128** on the semantic path, rank **32** on the depth decoder, `scaling 2.0`,
embeddings frozen. 674 tensors.
**Trainer** β€” GDPO (per-dimension group z-scored advantages, arXiv 2601.05242), k3 KL against an
adapter-off reference. 24 GPUs: 3 trainer ranks + 16 rollout workers + 1 ASR/VAD score worker +
judge shards. G=8 samples per group, 6 groups per step, lr 1e-5. Scoring is
`nvidia/parakeet-tdt-0.6b-v2` ASR + silero VAD + a `Qwen/Qwen3.8-27B` judge.
| | `gdpo-duplex` v1200 (R1) | `gdpo-v2f-r4` v600 (R7) |
|---|---|---|
| pool | conv-synth goal-arc dialogues | Fisher-only, 9,417 train / 300 val, **goalless** persona |
| reward | **v1**: `C_goal βˆ’ 0.5Β·role_fail βˆ’ 0.5Β·meta_talk βˆ’ 0.25Β·jarring`, silent = βˆ’1 | **v2**, 8 dims: C_goal, fits, seat_ok, meta_talk + VAD overlap / false_takeover / latency / backchannel |
| KL | fixed Ξ² 0.01 | adaptive Ξ², `--beta 0.05 --kl-target 0.3`, clamp 5.0, floor Ξ²/2 |
| steps | ran to 1,479; **v1200** published | stopped at 1,181 by hand; **v600** best |
| in-domain | compliance@0.5 0.455 β†’ **0.779** by v600, flat after; jarring 0.375 β†’ 0.211 | compliance **0.579**, fits 0.906, latency 1.34 s; no-instruction control flat 0.277 β†’ 0.278, so the true conditioning delta is **+0.288** |
## Limitations β€” what is wrong with these checkpoints
- **Both flip seats far more than the SFT model.** 40–42% (v1200) and 59–66% (v600) of dialogues
contain a quote-verified seat flip, against 4.6% for `stage2-337`. RL bought task completion
and spent role stability to get it. If persona adherence is what you need, neither of these is
the right checkpoint.
- **v600's reward under-priced the failure it was meant to price.** A strict flip-only judge
finds 14 / 22 / 12% flipped turns at v600 / v900 / v1050 versus 4–8% by the training `seat_ok`
β€” recall 0.24–0.32, so `role_fail` at 0.5 was pricing about **one third** of real seat flips.
- **v600 drifts monotonically even while KL stays inside its band.** Per-3-step motion is 2.5Γ—
the per-150-step motion, and KL-to-base is 0.23 at v900 versus 0.22 at v1179 β€” both inside the
0.3 target. A base-only anchor admits the best and worst checkpoints alike, so there is no
restoring force toward v600. Later steps of that run are *worse*, not better.
- **Winning in-domain is not winning the benchmark.** v600 beats v1200 on the Fisher held-out
conditioning delta and loses to it on IF-Bench by 28 points of completion. They were scored on
different distributions β€” goalless Fisher windows versus scripted caller tasks β€” and neither
is simply "better".
- **The examiner sometimes adopts the wrong identity.** Free-running from a card that names a
business, the model has opened with *"Good morning, this is Morning Delight Cafe."* β€” i.e. it
became the callee named in its own card. Cards should state the speaking role and who opens,
in the imperative.
- **No as-launched command line survives for R7.** The reconstruction is
`slurm/runs/rl/env/R7-gdpo-v2f-r4.env` and is marked as reconstructed, not as-launched. R1's
topology is committed accurately (`rl/submit_rl.sh`).
## Provenance
Code: [`MagicLuke/duplex-online-if`](https://github.com/MagicLuke/duplex-online-if), trainer
`rl/`. Benchmark harnesses: `fdb-v2-IF-bench` (IF-Bench) and `fdb-v1-repro` (FDB v1.0).
Generated dialogues and per-dialogue judgments for the IF-Bench rows above are published as
[`MagicLuke/ifbench-conversations-v1`](https://huggingface.co/datasets/MagicLuke/ifbench-conversations-v1);
FDB v1.0 outputs as
[`MagicLuke/fdb-v1-outputs-v1`](https://huggingface.co/datasets/MagicLuke/fdb-v1-outputs-v1).
`gdpo-duplex` has **no Fisher lineage** β€” conv-synth pool, base init. `gdpo-v2f-r4` was trained
on rollouts seeded from Fisher (LDC) and is gated on that basis; both are gated for consistency.
Base weights are NVIDIA's PersonaPlex and their licence governs the base model.