fix: RoPE inv_freq fp32 construction + clean-copy self-heal (NPU state-dict load corruption)

#4

Problem

On Ascend NPU (torch_npu), loading the state dict can leave the rotary-embedding inv_freq buffers corrupted — observed values include 9e-5, 3e-41, 0, 1e27 (non-persistent buffers can still be overwritten during weight loading on some transformers versions). Independent of corruption, an 8-bit-mantissa (bf16) frequency table amplifies RoPE phase error at large positions, compounding over long realtime sessions (visible response drift after ~30 turns on 910B2C).

What this PR changes

  • MossVLVisionRotaryEmbedding: build inv_freq in fp32 on CPU, register a non-persistent clean copy (_inv_freq_clean), and restore from it in forward when NaN/Inf is detected.
  • MossVLRotaryEmbedding (text side): same treatment — fp32 construction for the default rope type, clean copy (_inv_freq_computed), and a self-heal guard in forward that also restores original_inv_freq.
  • CUDA/GPU behavior is unchanged: the fp32 construction matches the stock computation, and the heal guard is a no-op when values are finite.

Evidence

  • With this change, 50+ turn realtime sessions on Ascend 910B2C show no positional drift; without it, drift appears after ~30 turns.
  • The adapter-side load-time validation counterpart is submitted separately in fnlp-vision/MOSS-VL-Realtime_Demo#8.
Publish this branch
This branch is in draft mode, publish it to be able to merge.

Sign up or log in to comment