KarthikRagunathAnandaKumar commited on
Commit
367f79a
·
verified ·
1 Parent(s): 9495534

Remove HTML from model description

Browse files
Files changed (1) hide show
  1. README.md +1 -1
README.md CHANGED
@@ -26,7 +26,7 @@ model-index:
26
 
27
  # LearningToPresent-RL-Qwen-2.5-Coder-7B-Instruct-GRPO-Finetuned
28
 
29
- A GRPO-finetuned LoRA adapter for [Qwen2.5-Coder-7B-Instruct](https://huggingface.co/Qwen/Qwen2.5-Coder-7B-Instruct), trained to generate professional HTML slide presentations through multi-turn tool use in a reinforcement learning environment.
30
 
31
  The model achieves **91.2% of Claude Opus 4.6's quality score** (0.724 vs. 0.794) while improving **33.1% over the untuned base model** (0.544), using only **0.53% trainable parameters** (40.4M out of 7.62B).
32
 
 
26
 
27
  # LearningToPresent-RL-Qwen-2.5-Coder-7B-Instruct-GRPO-Finetuned
28
 
29
+ A GRPO-finetuned LoRA adapter for [Qwen2.5-Coder-7B-Instruct](https://huggingface.co/Qwen/Qwen2.5-Coder-7B-Instruct), trained to generate professional slide presentations through multi-turn tool use in a reinforcement learning environment.
30
 
31
  The model achieves **91.2% of Claude Opus 4.6's quality score** (0.724 vs. 0.794) while improving **33.1% over the untuned base model** (0.544), using only **0.53% trainable parameters** (40.4M out of 7.62B).
32