Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
SeanWang0027
/
ftb-sciworld-repro
like
0
Reinforcement Learning
English
on-policy-distillation
llm-agents
scienceworld
License:
apache-2.0
Model card
Files
Files and versions
xet
Community
Copy to bucket
new
main
ftb-sciworld-repro
/
tcod
/
trinity
/
common
/
rewards
52.5 kB
Ctrl+K
Ctrl+K
1 contributor
History:
1 commit
SeanWang0027
Upload folder using huggingface_hub
8c9ba62
verified
13 days ago
__init__.py
Safe
883 Bytes
Upload folder using huggingface_hub
13 days ago
accuracy_reward.py
Safe
2.53 kB
Upload folder using huggingface_hub
13 days ago
agents_reward.py
Safe
20 Bytes
Upload folder using huggingface_hub
13 days ago
countdown_reward.py
Safe
1.58 kB
Upload folder using huggingface_hub
13 days ago
dapo_reward.py
Safe
2.19 kB
Upload folder using huggingface_hub
13 days ago
eval_utils.py
Safe
8.17 kB
Upload folder using huggingface_hub
13 days ago
format_reward.py
Safe
847 Bytes
Upload folder using huggingface_hub
13 days ago
human_reward.py
Safe
20 Bytes
Upload folder using huggingface_hub
13 days ago
math_reward.py
Safe
1.92 kB
Upload folder using huggingface_hub
13 days ago
naive_dapo_score.py
Safe
14.4 kB
Upload folder using huggingface_hub
13 days ago
qwen25_eval.py
Safe
16.2 kB
Upload folder using huggingface_hub
13 days ago
reward_fn.py
Safe
3.15 kB
Upload folder using huggingface_hub
13 days ago
tool_reward.py
Safe
20 Bytes
Upload folder using huggingface_hub
13 days ago
utils.py
Safe
634 Bytes
Upload folder using huggingface_hub
13 days ago