DL HW3 Final LoRA Adapter

This repository contains only the final LoRA adapter files for DL HW3.

Base Model

  • unsloth/Qwen2.5-7B-Instruct

Training Pipeline

  1. Answer-focused SFT
  2. GRPO-style reward optimization
  3. Final answer-only warm-up

Output Format

答案: X

where X is one of A, B, C, or D.

Downloads last month
1
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Aaron314581048/dl_hw3

Base model

Qwen/Qwen2.5-7B
Adapter
(662)
this model