Precompute cross-attention K/V in the text encoder

#20
ZeroWeight AI org

prefix_step took text_states as a per-frame input and re-projected every decoder layer's cross-attention K and V over the whole text on each frame - 45 us per text token per frame, for a value that cannot change within an utterance.

The text encoder now emits those K/V once as a new cross_kv output and prefix_step consumes them, replacing its text_states input.

Verified to produce byte-identical frame codes to the previous graphs for the same text, voice and random draws. Measured 1.07x on a 31-token sentence and 1.28x at 121 tokens (the ~120-token segments the browser demo's chunker produces); per-frame cost no longer grows with text length. Total download is unchanged - the weights moved between the two files.

Breaking for existing runtimes: prefix_step no longer accepts text_states. Callers must pass cross_kv from the text encoder.

zeroweightai changed pull request status to merged

Sign up or log in to comment