--- license: mit tags: [tikz, latex, diagram-generation, reinforcement-learning, verl] --- # scale_base_blend_dr @ step 450 RL checkpoint. Base **TikZilla-3B (SFT)**, trained with **Dr.GRPO/DAPO** against the **DSV `blend_dr`** verifiable reward (0.7·DSV + 0.3·RenderGraph, both symbolic verifiers -- no learned reward model). * data: [`explcre/datikz_v4_rl_hybrid_428k`](https://huggingface.co/datasets/explcre/datikz_v4_rl_hybrid_428k) (427,553 rows, 46.6% Kimi-K3 descriptions / 53.4% DaTikZ-v4 `vlm_description`, tagged per row) * verifier: `explcre/DeTikZify` branch `dsv-verifier` @ `c348406` * config: batch 84, rollout n=8, max_response_length 2048, 1500 steps planned **Evaluate with the matching input distribution.** The same checkpoint scores differently on `vlmdesc984` vs `kimidesc984`, and differently again under different TeX/compile-gate versions -- always report the measurement config with the number.