YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

payelb/HHRLHF_TinyLlama-1.1B_aligned_with_semantic_MARS_RM_roberta_semantic_MARS_RM

Base model: TinyLlama/TinyLlama-1.1B-Chat-v1.0

Alignment dataset: Anthropic/hh-rlhf

Reward model: payelb/HHRLHF_roberta-base_1k_fixed_MARS_semantic_distance_synth

Method: PPO alignment with LoRA adapters.

This model aligns TinyLlama-1.1B-Chat-v1.0 using the semantic-MARS reward model.

Training setup matched to existing TinyLlama WoN/baseline aligned models:

  • NUM_TRAIN_SAMPLES: 1000
  • TOTAL_PPO_STEPS: 250
  • PPO_EPOCHS: 2
  • LR: 5e-06
  • Batch size: 16
  • Mini-batch size: 4
  • Gradient accumulation: 4
  • MIN_NEW_TOKENS: 32
  • MAX_NEW_TOKENS: 64
  • Reward normalization enabled: True
  • REWARD_CLIP: 5.0
  • KL control enabled: init_kl_coef=0.02, target=6.0, adap_kl_ctrl=True
  • Generation: do_sample=True, top_p=0.9, temperature=0.7
  • LoRA r=16, alpha=32, dropout=0.05
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support