PEFT
Safetensors
lora
qlora
code-generation
python
File size: 5,479 Bytes
d706b93
acae1ff
 
 
 
 
d706b93
 
acae1ff
d706b93
acae1ff
 
 
d706b93
acae1ff
 
 
d706b93
acae1ff
d706b93
acae1ff
d706b93
acae1ff
 
 
 
d706b93
acae1ff
 
d706b93
acae1ff
 
d706b93
acae1ff
 
 
 
 
d706b93
acae1ff
 
 
d706b93
acae1ff
d706b93
acae1ff
d706b93
acae1ff
 
 
 
d706b93
acae1ff
 
d706b93
acae1ff
 
 
 
d706b93
acae1ff
 
d706b93
acae1ff
d706b93
acae1ff
 
 
 
 
 
 
 
 
 
d706b93
acae1ff
 
d706b93
acae1ff
d706b93
acae1ff
 
 
 
 
d706b93
acae1ff
 
 
d706b93
acae1ff
d706b93
acae1ff
 
 
d706b93
acae1ff
 
d706b93
acae1ff
d706b93
acae1ff
 
 
 
 
 
 
 
 
d706b93
acae1ff
d706b93
acae1ff
d706b93
acae1ff
 
 
d706b93
acae1ff
 
 
d706b93
acae1ff
 
 
 
d706b93
acae1ff
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
---
base_model: meta-llama/Llama-3.2-3B
library_name: peft
tags: [lora, qlora, code-generation, python]
datasets: [sahil2801/CodeAlpaca-20k]
license: llama3.2
---

# Python code generation β€” LLaMA 3.2 3B + QLoRA

**HumanEval pass@1: 40.5% β†’ 54.1%** on the uncontaminated subset β€” a 13.6-point gain
from a 9.2M-parameter adapter (0.285% of the model), trained in under two hours on a
single free-tier T4.

The base model already writes correct code. It just writes it in JavaScript 42% of the
time, wrapped in markdown fences. This adapter makes Python the default and the output
directly executable.

## Results

**HumanEval, first 50 problems** (greedy decoding, deterministic):

| | Base | Fine-tuned | Ξ” |
|---|---|---|---|
| **Clean subset (37 problems)** | **40.5%** | **54.1%** | **+13.6 pts** |
| All 50 | 46.0% | 58.0% | +12.0 pts |

The clean subset excludes 13 problems whose function names appear in the training data
(see Contamination). It is the number to cite.

**Custom eval set** β€” 40 hand-written Python problems with executable unit tests, all 40
reference solutions verified to pass before use:

| | Base | Fine-tuned |
|---|---|---|
| Valid Python (free-form instruction) | 57.5% | **100%** |
| Language correct | 62.5% | **100%** |
| pass@1 | 82.5% | **90%** |

Without a specified function signature, the base model emitted non-Python for 15 of 40
problems and markdown-fenced (non-executable) output for 2 more. The adapter reaches
100% executable Python with no prompt engineering.

## Usage

The prompt format is load-bearing.

```python
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
from peft import PeftModel
import torch

bnb = BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_quant_type="nf4",
    bnb_4bit_compute_dtype=torch.float16, bnb_4bit_use_double_quant=True)

base = AutoModelForCausalLM.from_pretrained(
    "meta-llama/Llama-3.2-3B", quantization_config=bnb, device_map="auto")
model = PeftModel.from_pretrained(base, "Raghul09/llama-code-gen-lora")
tok = AutoTokenizer.from_pretrained("Raghul09/llama-code-gen-lora")

prompt = f"### Instruction:\n{instruction}\n\n### Response:\n"
```

## Training

| | |
|---|---|
| Base | meta-llama/Llama-3.2-3B |
| Quantization | 4-bit NF4 + double quant β€” 6.43 GB β†’ 2.20 GB (66%) |
| LoRA | r=16, Ξ±=32, q/k/v/o across 28 layers |
| Trainable | 9,175,040 / 3,221,924,864 (0.285%) |
| Data | CodeAlpaca-20K, AST-filtered to Python β€” 6,418 / 802 / 803 |
| max_length | 256, set from measured distribution (p50=89, p95=211) |
| Hardware | single T4, 1.88 h, **4.17 GB peak VRAM** |
| Checkpoint | epoch 2 of 3 β€” val loss 0.475; epoch 3 rose to 0.485 |

Full fine-tuning of this model requires roughly 50 GB. QLoRA brought it to 4.17 GB,
a 12x reduction, which is what made it feasible on free-tier hardware.

## Rank sweep

| Rank | Trainable | Val loss | Valid Python | pass@1 |
|---|---|---|---|---|
| 8 | 4,587,520 | 0.4795 | 100% | 82.5% |
| **16** | **9,175,040** | **0.4753** | **100%** | **90.0%** |
| 32 | 18,350,080 | 0.4739 | 100% | 82.5% |

Quadrupling rank bought 1.2% lower validation loss and no consistent gain in pass@1.
The target behaviour is genuinely low-rank β€” consistent with the LoRA paper's central
hypothesis, tested here rather than assumed.

## Contamination

13 of the 50 HumanEval problems have their function names defined in CodeAlpaca. The
fine-tuned model scores 69.2% on those versus ~52% on clean problems, under both prompt
formats tested. Headline numbers use the clean subset only.

Function-name matching catches exact reuse but misses paraphrased problems, so 26% is a
lower bound on overlap, not an estimate.

## Limitations

- **Synthetic training data.** CodeAlpaca-20K is GPT-generated via self-instruct,
  unverified, stylistically homogeneous. Quality is bounded by the teacher model.
- **Conformance over capability.** With a function signature specified in the prompt,
  the base model already reaches 97.5% valid Python and 82.5% pass@1. This adapter's
  main contribution is removing the need for that prompt engineering.
- **Python only.** Training data was AST-filtered to Python.
- **Short outputs.** Median training example was 89 tokens; long generations degrade.
- **Known failure modes:** repetition loops causing mid-generation truncation; calling
  helper functions it never defines.

## Methodology notes

Three measurement errors found and corrected during evaluation:

1. **Naming confound.** Initial pass@1 read 35% base / 50% fine-tuned. 17 of 20 failures
   were `NameError` β€” correct code under a different function name than the test called.
   Specifying signatures corrected the baseline by 47 points, to 82.5%.

2. **Indentation destruction.** Raw-format HumanEval returned 0% for both models. The
   fence-stripper called `.strip()`, removing leading indentation from function bodies.
   All 50 failures were `IndentationError`.

3. **Training-data contamination.** A keyword-based Python filter kept 58.7% of
   CodeAlpaca; AST-parsing a 300-example sample showed 38% weren't valid Python β€” mostly
   Java and JavaScript matching on shared keywords like `for` and `class`. Replaced with
   `ast.parse` plus a syntax-tree check: 0% contamination on re-check.

Sequence packing was disabled after batch inspection revealed it silently disabled
completion-only loss masking. Packing requires Flash Attention for block-diagonal
masking, which requires Ampere; the T4 is Turing.