File size: 929 Bytes
a4d6b10
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
---
license: mit
tags: [tikz, latex, diagram-generation, reinforcement-learning, verl]
---

# scale_base_blend_dr @ step 50

RL checkpoint. Base **TikZilla-3B (SFT)**, trained with **Dr.GRPO/DAPO**
against the **DSV `blend_dr`** verifiable reward (0.7·DSV + 0.3·RenderGraph, both symbolic verifiers --
no learned reward model).

* data: [`explcre/datikz_v4_rl_hybrid_428k`](https://huggingface.co/datasets/explcre/datikz_v4_rl_hybrid_428k)
  (427,553 rows, 46.6% Kimi-K3 descriptions / 53.4% DaTikZ-v4 `vlm_description`, tagged per row)
* verifier: `explcre/DeTikZify` branch `dsv-verifier` @ `c348406`
* config: batch 84, rollout n=8, max_response_length 2048, 1500 steps planned

**Evaluate with the matching input distribution.** The same checkpoint scores differently on
`vlmdesc984` vs `kimidesc984`, and differently again under different TeX/compile-gate versions --
always report the measurement config with the number.