Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
Xiaoyang Cao
Sean13
6
Follow
0 followers
·
2 following
https://xiaoyangcao1113.github.io/
XiaoyangCao1113
xiaoyangcao
AI & ML interests
RLFH, Deep Reinfrocement Learning
Recent Activity
upvoted
a
paper
4 days ago
Do Modules Stay in Their Lane? Role Drift in Compound LLM Systems
updated
a model
3 months ago
Sean13/responsibility-decomposition
published
a model
3 months ago
Sean13/responsibility-decomposition
View all activity
Organizations
None yet
models
16
Sort: Recently updated
Sean13/responsibility-decomposition
Reinforcement Learning
•
Updated
May 27
Sean13/mistral-7b-instruct-v0.2-rdpo-full-alpha0.01
7B
•
Updated
Sep 22, 2025
•
4
Sean13/mistral-7b-instruct-v0.2-rdpo-full-alpha0.5
7B
•
Updated
Sep 22, 2025
•
3
Sean13/mistral-7b-instruct-v0.2-rdpo-full-alpha1.0
7B
•
Updated
Sep 22, 2025
•
4
Sean13/mistral-7b-instruct-v0.2-rdpo-full-alpha0.9
7B
•
Updated
Sep 19, 2025
•
2
Sean13/mistral-7b-instruct-v0.2-rdpo-full-alpha0.7
7B
•
Updated
Sep 19, 2025
•
4
Sean13/mistral-7b-instruct-v0.2-rdpo-full-alpha0.3
Updated
Sep 19, 2025
Sean13/mistral-7b-instruct-v0.2-rcpo-full
Text Generation
•
7B
•
Updated
Sep 15, 2025
•
5
Sean13/mistral-7b-instruct-v0.2-cpo-full
Text Generation
•
7B
•
Updated
Sep 11, 2025
•
4
Sean13/mistral-7b-instruct-v0.2-simpo-full
Text Generation
•
7B
•
Updated
Sep 6, 2025
•
8
View 16 models
datasets
0
None public yet