--- license: other license_name: llama3.2 license_link: https://github.com/meta-llama/llama-models/blob/main/models/llama3_2/LICENSE base_model: unsloth/Llama-3.2-1B tags: - math - fine-tuned - lora - unsloth - llama datasets: - MathLLMs/MathCodeInstruct language: - en library_name: transformers pipeline_tag: text-generation model-index: - name: Llama-3.2-1B-MathCodeInstruct-{{SIZE}} results: - task: type: text-generation name: GSM8K dataset: type: gsm8k name: GSM8K metrics: - type: exact_match value: {{GSM8K_ACC}} name: exact match (flexible-extract, 5-shot) - task: type: text-generation name: ARC-Challenge dataset: type: ai2_arc name: ARC-Challenge metrics: - type: acc_norm value: {{ARC_ACC}} name: acc_norm (25-shot) - task: type: text-generation name: HellaSwag dataset: type: hellaswag name: HellaSwag metrics: - type: acc_norm value: {{HELLASWAG_ACC}} name: acc_norm (10-shot) - task: type: text-generation name: WinoGrande dataset: type: winogrande name: WinoGrande metrics: - type: acc value: {{WINOGRANDE_ACC}} name: acc (5-shot) - task: type: text-generation name: MMLU dataset: type: mmlu name: MMLU metrics: - type: acc value: {{MMLU_ACC}} name: acc (5-shot) --- # Llama-3.2-1B-MathCodeInstruct-20k A [Llama-3.2-1B](https://huggingface.co/unsloth/Llama-3.2-1B) fine-tune on **20k examples** from [MathLLMs/MathCodeInstruct](https://huggingface.co/datasets/MathLLMs/MathCodeInstruct), trained to solve math word problems with step-by-step natural-language reasoning interleaved with executable Python. This is one of three sibling models trained on {5k, 10k, 20k}-example subsets of the same dataset, to study how fine-tuning data volume trades off against both math performance and general capability. See the [training write-up](https://github.com/OliverSundaram/finetuning-Llama3.2-1B) for the full comparison across all three. ## Training details | | | |---|-----------------------------------------------------------------------------------------| | Base model | `unsloth/Llama-3.2-1B` | | Method | LoRA (r=16, α=16, dropout=0) on all attention + MLP projections, merged to full weights | | Dataset | MathLLMs/MathCodeInstruct, 20k training examples | | Epochs | 1 | | Effective batch size | 16 (batch 1 × grad. accum. 16) | | Learning rate | 2e-4, cosine schedule, warmup ratio 0.03 | | Hardware | 1× RTX 4060 (8GB) | | Framework | Unsloth + TRL `SFTTrainer` | ## Benchmark results All benchmarks run with [lm-evaluation-harness](https://github.com/EleutherAI/lm-evaluation-harness), each at its standard published shot count, compared against the un-tuned base model. | Benchmark | Llama-3.2-1B (base) | This model | Change | |---|---|---|---| | GSM8K | 5.8% | 8.9% | 🟢 +3.1% | | ARC-Challenge | 36.9% | 35.8% | 🔴 -1.1% | | HellaSwag | 64.2% | 63.6% | 🔴 -0.6% | | WinoGrande | 60.8% | 61.4% | 🟢 +0.6% | **Speed** (single-request generation, greedy, RTX 4060): **37.84 tokens/sec** (base model: 12.74 tokens/sec) ### MMLU by category ![MMLU comparison](mmlu_20k.png) ## Usage ```python from transformers import AutoModelForCausalLM, AutoTokenizer model_id = "OliverSundaram/Llama-3.2-1B-MathCodeInstruct-20k}" tokenizer = AutoTokenizer.from_pretrained(model_id) model = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype="bfloat16", device_map="auto") messages = [ {"role": "system", "content": "Below is a math problem. Please solve it step by step."}, {"role": "user", "content": "If a train travels 60 miles in 45 minutes, what is its speed in miles per hour?"}, ] inputs = tokenizer.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device) output = model.generate(inputs, max_new_tokens=512, do_sample=False) print(tokenizer.decode(output[0], skip_special_tokens=True)) ``` ## Limitations - Trained on a single epoch of a 20k-example subset — not intended to be a general-purpose assistant. - MMLU/ARC/HellaSwag/WinoGrande scores reflect a small 1B-parameter base model and should be read relative to the base model's own scores, not against much larger models. - No safety alignment or RLHF was applied beyond what the base Llama-3.2-1B already has.