HypoDiverse: GRPO

This is the exact merged Hugging Face checkpoint evaluated on HypoDiverse.

Provenance

Field Value
method GRPO
Base model Qwen/Qwen3-4B
Dataset viciousa3gis/hypodiverse
Pinned dataset revision d16867cc49836f72ace9e3667164fa6e4ae76eda
Evaluation protocol standard

This checkpoint is the validity-reward GRPO baseline. Each completion is rewarded for producing a hypothesis that is consistent with the visible evidence; the reward has no explicit set-diversity term.

Exact training and evaluation configurations, per-file model hashes, and the pinned dataset revision are recorded in release_manifest.json and provenance/configs/.

Downloads last month
-
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for viciousa3gis/hypodiverse-grpo

Finetuned
Qwen/Qwen3-4B
Finetuned
(1070)
this model

Dataset used to train viciousa3gis/hypodiverse-grpo

Collection including viciousa3gis/hypodiverse-grpo