TRIAGE: Direction-Aware Mismatch Stabilization of Native NVFP4 Reinforcement Learning Paper • 2610.07043 • Published 5 days ago • 39
TRIAGE: Direction-Aware Mismatch Stabilization of Native NVFP4 Reinforcement Learning Paper • 2610.07043 • Published 5 days ago • 39
A Practical Chinese Dependency Parser Based on A Large-scale Dataset Paper • 2009.00901 • Published Sep 3, 2020
z369217411/Qwen3-30B-A3B-GRPO-W4A4-NoOverlong-Step800 Reinforcement Learning • 31B • Updated Aug 9 • 25
z369217411/Qwen3-30B-A3B-GRPO-W4A4-NoOverlong-Step800 Reinforcement Learning • 31B • Updated Aug 9 • 25
z369217411/Qwen3-30B-A3B-GRPO-LowPrecision-Control-Step410 Text Generation • 31B • Updated Jul 27 • 21
z369217411/Qwen3-30B-A3B-GRPO-LowPrecision-Control-Step410 Text Generation • 31B • Updated Jul 27 • 21
Qwen3-30B-A3B W4A4-QAT vs BF16 Checkpoints Collection Seven step-matched Qwen3-30B-A3B DAPO pairs: FFN W4A4-QAT BF16 master weights (not packed W4A4) and a BF16 baseline. • 14 items • Updated Jul 12