explcre's picture
scale_base_blend_dr @ step 400
362bb4f verified
|
Raw
History Blame Contribute Delete
930 Bytes
---
license: mit
tags: [tikz, latex, diagram-generation, reinforcement-learning, verl]
---
# scale_base_blend_dr @ step 400
RL checkpoint. Base **TikZilla-3B (SFT)**, trained with **Dr.GRPO/DAPO**
against the **DSV `blend_dr`** verifiable reward (0.7路DSV + 0.3路RenderGraph, both symbolic verifiers --
no learned reward model).
* data: [`explcre/datikz_v4_rl_hybrid_428k`](https://huggingface.co/datasets/explcre/datikz_v4_rl_hybrid_428k)
(427,553 rows, 46.6% Kimi-K3 descriptions / 53.4% DaTikZ-v4 `vlm_description`, tagged per row)
* verifier: `explcre/DeTikZify` branch `dsv-verifier` @ `c348406`
* config: batch 84, rollout n=8, max_response_length 2048, 1500 steps planned
**Evaluate with the matching input distribution.** The same checkpoint scores differently on
`vlmdesc984` vs `kimidesc984`, and differently again under different TeX/compile-gate versions --
always report the measurement config with the number.