File size: 1,418 Bytes
704d7b2
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
---
datasets: lilkm/stackblocks_recap_all_for_vf
library_name: lerobot
license: apache-2.0
model_name: distributional_value_function
pipeline_tag: robotics
tags:
- reward-model
- robotics
- distributional_value_function
- lerobot
---

# Reward Model Card for distributional_value_function

<!-- Provide a quick summary of what the reward model is/does. -->


_Reward model type not recognized — please update this template._


This reward model has been trained and pushed to the Hub using [LeRobot](https://github.com/huggingface/lerobot).
See the full documentation at [LeRobot Docs](https://huggingface.co/docs/lerobot/index).

---

## How to Get Started with the Reward Model

### Train from scratch

```bash
lerobot-train \
  --dataset.repo_id=${HF_USER}/<dataset> \
  --reward_model.type=distributional_value_function \
  --output_dir=outputs/train/<desired_reward_model_repo_id> \
  --job_name=lerobot_reward_training \
  --reward_model.device=cuda \
  --reward_model.repo_id=${HF_USER}/<desired_reward_model_repo_id> \
  --wandb.enable=true
```

_Writes checkpoints to `outputs/train/<desired_reward_model_repo_id>/checkpoints/`._

### Load the reward model in Python

```python
from lerobot.rewards import make_reward_model

reward_model = make_reward_model(pretrained_path="<hf_user>/<reward_model_repo_id>")
reward = reward_model.compute_reward(batch)
```

---

## Model Details

- **License:** apache-2.0