YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
payelb/UltraFeedback_openbmb_Llama-3.2-1B_aligned_with_semantic_MARS_roberta_RM_KLsafe
Base model: meta-llama/Llama-3.2-1B-Instruct
Alignment dataset: openbmb/UltraFeedback
Reward model: payelb/UltraFeedback_openbmb_roberta-base_1k_fixed_MARS_semantic_refined
Reward model type: RoBERTa-base
RM method: semantic_mars
Method: PPO alignment with LoRA adapters.
Training details:
- NUM_TRAIN_SAMPLES: 1000
- TOTAL_PPO_STEPS: 250
- PPO_EPOCHS: 2
- LR: 5e-06
- Batch size: 16
- Mini-batch size: 4
- Gradient accumulation: 4
- MIN_NEW_TOKENS: 32
- MAX_NEW_TOKENS: 64
- USE_REWARD_NORMALIZATION: True
- REWARD_CLIP: 5.0
- KL control enabled: init_kl_coef=0.02, target=6.0, adap_kl_ctrl=True
- KL-safe rollout generation:
- do_sample=True
- temperature=0.7
- top_k=0.0
- top_p=1.0
- min_length=-1
- eos_token_id=None
- pad_token_id=policy_tokenizer.pad_token_id
- LoRA enabled:
- r=16
- alpha=32
- dropout=0.05
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support