leonMW commited on
Commit
f74ab41
·
verified ·
1 Parent(s): d5794f3

Training in progress, epoch 1

Browse files
Files changed (4) hide show
  1. README.md +2 -4
  2. adapter_model.safetensors +1 -1
  3. tokenizer.json +2 -2
  4. training_args.bin +1 -1
README.md CHANGED
@@ -1,19 +1,17 @@
1
  ---
2
  base_model: deepseek-ai/DeepSeek-R1-Distill-Qwen-7B
3
- datasets: AIML-TUDA/SLR-Bench
4
  library_name: transformers
5
  model_name: DeepSeek-R1-Distill-Qwen-7B-S
6
  tags:
7
  - generated_from_trainer
8
  - grpo
9
- - open-r1
10
  - trl
11
  licence: license
12
  ---
13
 
14
  # Model Card for DeepSeek-R1-Distill-Qwen-7B-S
15
 
16
- This model is a fine-tuned version of [deepseek-ai/DeepSeek-R1-Distill-Qwen-7B](https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-7B) on the [AIML-TUDA/SLR-Bench](https://huggingface.co/datasets/AIML-TUDA/SLR-Bench) dataset.
17
  It has been trained using [TRL](https://github.com/huggingface/trl).
18
 
19
  ## Quick start
@@ -29,7 +27,7 @@ print(output["generated_text"])
29
 
30
  ## Training procedure
31
 
32
- [<img src="https://raw.githubusercontent.com/wandb/assets/main/wandb-github-badge-28.svg" alt="Visualize in Weights & Biases" width="150" height="24"/>](https://wandb.ai/leonwenderoth-tu-darmstadt/huggingface/runs/lm1y3lp0)
33
 
34
 
35
  This model was trained with GRPO, a method introduced in [DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models](https://huggingface.co/papers/2402.03300).
 
1
  ---
2
  base_model: deepseek-ai/DeepSeek-R1-Distill-Qwen-7B
 
3
  library_name: transformers
4
  model_name: DeepSeek-R1-Distill-Qwen-7B-S
5
  tags:
6
  - generated_from_trainer
7
  - grpo
 
8
  - trl
9
  licence: license
10
  ---
11
 
12
  # Model Card for DeepSeek-R1-Distill-Qwen-7B-S
13
 
14
+ This model is a fine-tuned version of [deepseek-ai/DeepSeek-R1-Distill-Qwen-7B](https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-7B).
15
  It has been trained using [TRL](https://github.com/huggingface/trl).
16
 
17
  ## Quick start
 
27
 
28
  ## Training procedure
29
 
30
+ [<img src="https://raw.githubusercontent.com/wandb/assets/main/wandb-github-badge-28.svg" alt="Visualize in Weights & Biases" width="150" height="24"/>](https://wandb.ai/leonwenderoth-tu-darmstadt/huggingface/runs/qwg16ch3)
31
 
32
 
33
  This model was trained with GRPO, a method introduced in [DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models](https://huggingface.co/papers/2402.03300).
adapter_model.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:52565e6aa76af43bc34670dbbfdbed35672b7b34ffd12aeb992113f8a00c3a00
3
  size 80792880
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:9e82218ce0f3e1be2d9ecbe7cb3f1323e66a4adf4bae86d16c980514dfd145db
3
  size 80792880
tokenizer.json CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:e20ddafc659ba90242154b55275402edeca0715e5dbb30f56815a4ce081f4893
3
- size 11422778
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:a4256422650d141f228fe954acee98679da412984c29a569877eefd3af69315a
3
+ size 11422959
training_args.bin CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:34d171fb43d579cf426e411af8692ec4c8f05b081506e460c53ab7aab072009e
3
  size 11729
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:8c03fb0a2519fe1221cc0794b826e3fa32b144f6afc892f441c388a80afa7bd0
3
  size 11729