cross_tool_qwen3-32b

GRPO experiment from TinkerRL-Bench world-class experiment suite.

Training Details

  • Base model: Qwen/Qwen3-32B
  • Method: GRPO (Group Relative Policy Optimization)
  • Platform: Tinker API v0.18.1
  • Task: tool_use
  • Seed: 42
  • LoRA rank: 32
  • Learning rate: 3e-05
  • Group size: 8
  • Steps: 30

Results

  • First-5 avg reward: 0.0%
  • Last-10 avg reward: 0.0%
  • Peak reward: 0.0%
  • Zero-loss steps: 100%
  • Tinker Run ID: 73ff186a-d8e3-50c3-afb8-2b863cd09579:train:0

Reward Trace

[
  0.0,
  0.0,
  0.0,
  0.0,
  0.0,
  0.0,
  0.0,
  0.0,
  0.0,
  0.0,
  0.0,
  0.0,
  0.0,
  0.0,
  0.0,
  0.0,
  0.0,
  0.0,
  0.0,
  0.0,
  0.0,
  0.0,
  0.0,
  0.0,
  0.0,
  0.0,
  0.0,
  0.0,
  0.0,
  0.0
]

Citation

@misc{tinker-rl-bench-2026,
  title={TinkerRL-Bench: A Unified Benchmark for RL Post-Training},
  author={Arvind C R and Sandhya Jeyaraj and Madhu Kumara L and Mohammad Rafi and Dhruva N Murthy and Arumugam K},
  year={2026},
  url={https://github.com/arvindcr4/tinker-rl-lab}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Model tree for arvindcr4/tinker-rl-bench-cross_tool_qwen3-32b

Base model

Qwen/Qwen3-32B
Finetuned
(523)
this model

Dataset used to train arvindcr4/tinker-rl-bench-cross_tool_qwen3-32b