Instructions to use Raghul09/llama-code-gen-lora with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use Raghul09/llama-code-gen-lora with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("meta-llama/Llama-3.2-3B") model = PeftModel.from_pretrained(base_model, "Raghul09/llama-code-gen-lora") - Notebooks
- Google Colab
- Kaggle
File size: 5,479 Bytes
d706b93 acae1ff d706b93 acae1ff d706b93 acae1ff d706b93 acae1ff d706b93 acae1ff d706b93 acae1ff d706b93 acae1ff d706b93 acae1ff d706b93 acae1ff d706b93 acae1ff d706b93 acae1ff d706b93 acae1ff d706b93 acae1ff d706b93 acae1ff d706b93 acae1ff d706b93 acae1ff d706b93 acae1ff d706b93 acae1ff d706b93 acae1ff d706b93 acae1ff d706b93 acae1ff d706b93 acae1ff d706b93 acae1ff d706b93 acae1ff d706b93 acae1ff d706b93 acae1ff d706b93 acae1ff d706b93 acae1ff d706b93 acae1ff d706b93 acae1ff d706b93 acae1ff d706b93 acae1ff d706b93 acae1ff d706b93 acae1ff | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 | ---
base_model: meta-llama/Llama-3.2-3B
library_name: peft
tags: [lora, qlora, code-generation, python]
datasets: [sahil2801/CodeAlpaca-20k]
license: llama3.2
---
# Python code generation β LLaMA 3.2 3B + QLoRA
**HumanEval pass@1: 40.5% β 54.1%** on the uncontaminated subset β a 13.6-point gain
from a 9.2M-parameter adapter (0.285% of the model), trained in under two hours on a
single free-tier T4.
The base model already writes correct code. It just writes it in JavaScript 42% of the
time, wrapped in markdown fences. This adapter makes Python the default and the output
directly executable.
## Results
**HumanEval, first 50 problems** (greedy decoding, deterministic):
| | Base | Fine-tuned | Ξ |
|---|---|---|---|
| **Clean subset (37 problems)** | **40.5%** | **54.1%** | **+13.6 pts** |
| All 50 | 46.0% | 58.0% | +12.0 pts |
The clean subset excludes 13 problems whose function names appear in the training data
(see Contamination). It is the number to cite.
**Custom eval set** β 40 hand-written Python problems with executable unit tests, all 40
reference solutions verified to pass before use:
| | Base | Fine-tuned |
|---|---|---|
| Valid Python (free-form instruction) | 57.5% | **100%** |
| Language correct | 62.5% | **100%** |
| pass@1 | 82.5% | **90%** |
Without a specified function signature, the base model emitted non-Python for 15 of 40
problems and markdown-fenced (non-executable) output for 2 more. The adapter reaches
100% executable Python with no prompt engineering.
## Usage
The prompt format is load-bearing.
```python
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
from peft import PeftModel
import torch
bnb = BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_quant_type="nf4",
bnb_4bit_compute_dtype=torch.float16, bnb_4bit_use_double_quant=True)
base = AutoModelForCausalLM.from_pretrained(
"meta-llama/Llama-3.2-3B", quantization_config=bnb, device_map="auto")
model = PeftModel.from_pretrained(base, "Raghul09/llama-code-gen-lora")
tok = AutoTokenizer.from_pretrained("Raghul09/llama-code-gen-lora")
prompt = f"### Instruction:\n{instruction}\n\n### Response:\n"
```
## Training
| | |
|---|---|
| Base | meta-llama/Llama-3.2-3B |
| Quantization | 4-bit NF4 + double quant β 6.43 GB β 2.20 GB (66%) |
| LoRA | r=16, Ξ±=32, q/k/v/o across 28 layers |
| Trainable | 9,175,040 / 3,221,924,864 (0.285%) |
| Data | CodeAlpaca-20K, AST-filtered to Python β 6,418 / 802 / 803 |
| max_length | 256, set from measured distribution (p50=89, p95=211) |
| Hardware | single T4, 1.88 h, **4.17 GB peak VRAM** |
| Checkpoint | epoch 2 of 3 β val loss 0.475; epoch 3 rose to 0.485 |
Full fine-tuning of this model requires roughly 50 GB. QLoRA brought it to 4.17 GB,
a 12x reduction, which is what made it feasible on free-tier hardware.
## Rank sweep
| Rank | Trainable | Val loss | Valid Python | pass@1 |
|---|---|---|---|---|
| 8 | 4,587,520 | 0.4795 | 100% | 82.5% |
| **16** | **9,175,040** | **0.4753** | **100%** | **90.0%** |
| 32 | 18,350,080 | 0.4739 | 100% | 82.5% |
Quadrupling rank bought 1.2% lower validation loss and no consistent gain in pass@1.
The target behaviour is genuinely low-rank β consistent with the LoRA paper's central
hypothesis, tested here rather than assumed.
## Contamination
13 of the 50 HumanEval problems have their function names defined in CodeAlpaca. The
fine-tuned model scores 69.2% on those versus ~52% on clean problems, under both prompt
formats tested. Headline numbers use the clean subset only.
Function-name matching catches exact reuse but misses paraphrased problems, so 26% is a
lower bound on overlap, not an estimate.
## Limitations
- **Synthetic training data.** CodeAlpaca-20K is GPT-generated via self-instruct,
unverified, stylistically homogeneous. Quality is bounded by the teacher model.
- **Conformance over capability.** With a function signature specified in the prompt,
the base model already reaches 97.5% valid Python and 82.5% pass@1. This adapter's
main contribution is removing the need for that prompt engineering.
- **Python only.** Training data was AST-filtered to Python.
- **Short outputs.** Median training example was 89 tokens; long generations degrade.
- **Known failure modes:** repetition loops causing mid-generation truncation; calling
helper functions it never defines.
## Methodology notes
Three measurement errors found and corrected during evaluation:
1. **Naming confound.** Initial pass@1 read 35% base / 50% fine-tuned. 17 of 20 failures
were `NameError` β correct code under a different function name than the test called.
Specifying signatures corrected the baseline by 47 points, to 82.5%.
2. **Indentation destruction.** Raw-format HumanEval returned 0% for both models. The
fence-stripper called `.strip()`, removing leading indentation from function bodies.
All 50 failures were `IndentationError`.
3. **Training-data contamination.** A keyword-based Python filter kept 58.7% of
CodeAlpaca; AST-parsing a 300-example sample showed 38% weren't valid Python β mostly
Java and JavaScript matching on shared keywords like `for` and `class`. Replaced with
`ast.parse` plus a syntax-tree check: 0% contamination on re-check.
Sequence packing was disabled after batch inspection revealed it silently disabled
completion-only loss masking. Packing requires Flash Attention for block-diagonal
masking, which requires Ampere; the T4 is Turing.
|