metadata
license: mit
base_model:
- CodeGoat24/UnifiedReward-Think-qwen3vl-32b
datasets:
- CodeGoat24/UnifiedReward-Flex-SFT-90K
Model Summary
UnifiedReward-Flex-qwen3vl-32b is a unified personalized reward model for vision generation that couples reward modeling with flexible and context-adaptive reasoning!!
๐ The inference code is available at Github.
For further details, please refer to the following resources:
- ๐ฐ Paper: https://arxiv.org/abs/2602.02380
- ๐ช Project Page: https://codegoat24.github.io/UnifiedReward/flex
- ๐ค Model Collections: https://huggingface.co/collections/CodeGoat24/unifiedreward-flex
- ๐ค Dataset: https://huggingface.co/datasets/CodeGoat24/UnifiedReward-Flex-SFT-90K
- ๐ Point of Contact: Yibin Wang
Citation
@article{unifiedreward-flex,
title={Unified Personalized Reward Model for Vision Generation},
author={Wang, Yibin and Zang, Yuhang and Han, Feng and Bu, Jiazi and Zhou, Yujie and Jin, Cheng and Wang, Jiaqi},
journal={arXiv preprint arXiv:2602.02380},
year={2026}
}