otoearth 's Collections

Fullduplex Signals

Weekly signals in speech-to-speech and full-duplex voice AI. Latest: 2026-W35, Aug 17 - Aug 23, 2026. Archive: fullduplex.ai/signals


  • Note 2026-W35 · 480 persona-grounded scenarios hold the task fixed and vary whether the user's concern is stated in words or carried only in prosody, with objectively checkable outcomes. Giving the model the audio on top of the transcript moves the optimal-solution rate from 14.6% to 15.3%. Forcing it to first write the inferred concern into text takes the same models to 39.6%, against 40.7% for ground-truth state. The prosody is recoverable, and it still does not reach the action unless something ma


  • Note 2026-W35 · Synthesises from uncertain token prefixes instead of waiting for a sentence, using uncertainty-aware buffering and carrying decoder state across segment boundaries. Reported at 15.8ms median time to first token for a single request and 260.8ms at 128 concurrent. Five of its seven authors also wrote X2-Turn, the streaming turn-state model from last week, so one company is now assembling a real-time voice stack part by part without training an end-to-end duplex model at any point.


  • Note 2026-W35 · Instead of sweeping blindly through distortions, this uses cheap structural probes to locate the domain a watermark is embedded in, then applies a single attack matched to that domain, and reports a threshold-free fragility score per scheme. It needs no training and no access to the watermarking model. Anyone planning to satisfy a marking obligation with a neural audio watermark should read it before treating that watermark as the compliance artefact.


  • Note 2026-W35 · Adds real-time predicted backchannels and head nodding to a voice-cloned avatar and measures the effect in a within-subjects study of 35 people. Perceived attentiveness, the sense of talking with the real person, and co-presence all improve significantly. The argument is that a duplex agent feels present because of how it listens, not how well it speaks, which is a case for spending latency budget on the listening side.


  • Note 2026-W35 · Three behavioural probes, covering reference disagreement, masked-number recovery, and orthographic switching, show leading open ASR models reproducing verbatim benchmark reference spans even when the audio contradicts them, masks them, or leaves them ambiguous. The behaviour can be steered with a low-rank direction, which makes it a learned policy rather than an artefact. Anyone choosing a backbone off a WER leaderboard is reading a number that partly measures memorisation.


  • Note 2026-W35 · One backbone with two decoupled continuous paths, an audio encoder for understanding and a RedAE path for generation, covering ASR, audio understanding, zero-shot and instruct TTS, semantic and acoustic speech editing, and temporal grounding over recordings up to an hour. Weights are up under Apache-2.0. The benchmark claims on MMAU, MMSU, Seed-TTS-Eval and InstructTTSEval are self-reported and the linked paper is still a placeholder, so treat the scores as unverified.