Instructions to use Raghul09/llama-code-gen-lora with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use Raghul09/llama-code-gen-lora with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("meta-llama/Llama-3.2-3B") model = PeftModel.from_pretrained(base_model, "Raghul09/llama-code-gen-lora") - Notebooks
- Google Colab
- Kaggle
Python code generation β LLaMA 3.2 3B + QLoRA
HumanEval pass@1: 40.5% β 54.1% on the uncontaminated subset β a 13.6-point gain from a 9.2M-parameter adapter (0.285% of the model), trained in under two hours on a single free-tier T4.
The base model already writes correct code. It just writes it in JavaScript 42% of the time, wrapped in markdown fences. This adapter makes Python the default and the output directly executable.
Results
HumanEval, first 50 problems (greedy decoding, deterministic):
| Base | Fine-tuned | Ξ | |
|---|---|---|---|
| Clean subset (37 problems) | 40.5% | 54.1% | +13.6 pts |
| All 50 | 46.0% | 58.0% | +12.0 pts |
The clean subset excludes 13 problems whose function names appear in the training data (see Contamination). It is the number to cite.
Custom eval set β 40 hand-written Python problems with executable unit tests, all 40 reference solutions verified to pass before use:
| Base | Fine-tuned | |
|---|---|---|
| Valid Python (free-form instruction) | 57.5% | 100% |
| Language correct | 62.5% | 100% |
| pass@1 | 82.5% | 90% |
Without a specified function signature, the base model emitted non-Python for 15 of 40 problems and markdown-fenced (non-executable) output for 2 more. The adapter reaches 100% executable Python with no prompt engineering.
Usage
The prompt format is load-bearing.
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
from peft import PeftModel
import torch
bnb = BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_quant_type="nf4",
bnb_4bit_compute_dtype=torch.float16, bnb_4bit_use_double_quant=True)
base = AutoModelForCausalLM.from_pretrained(
"meta-llama/Llama-3.2-3B", quantization_config=bnb, device_map="auto")
model = PeftModel.from_pretrained(base, "Raghul09/llama-code-gen-lora")
tok = AutoTokenizer.from_pretrained("Raghul09/llama-code-gen-lora")
prompt = f"### Instruction:\n{instruction}\n\n### Response:\n"
Training
| Base | meta-llama/Llama-3.2-3B |
| Quantization | 4-bit NF4 + double quant β 6.43 GB β 2.20 GB (66%) |
| LoRA | r=16, Ξ±=32, q/k/v/o across 28 layers |
| Trainable | 9,175,040 / 3,221,924,864 (0.285%) |
| Data | CodeAlpaca-20K, AST-filtered to Python β 6,418 / 802 / 803 |
| max_length | 256, set from measured distribution (p50=89, p95=211) |
| Hardware | single T4, 1.88 h, 4.17 GB peak VRAM |
| Checkpoint | epoch 2 of 3 β val loss 0.475; epoch 3 rose to 0.485 |
Full fine-tuning of this model requires roughly 50 GB. QLoRA brought it to 4.17 GB, a 12x reduction, which is what made it feasible on free-tier hardware.
Rank sweep
| Rank | Trainable | Val loss | Valid Python | pass@1 |
|---|---|---|---|---|
| 8 | 4,587,520 | 0.4795 | 100% | 82.5% |
| 16 | 9,175,040 | 0.4753 | 100% | 90.0% |
| 32 | 18,350,080 | 0.4739 | 100% | 82.5% |
Quadrupling rank bought 1.2% lower validation loss and no consistent gain in pass@1. The target behaviour is genuinely low-rank β consistent with the LoRA paper's central hypothesis, tested here rather than assumed.
Contamination
13 of the 50 HumanEval problems have their function names defined in CodeAlpaca. The fine-tuned model scores 69.2% on those versus ~52% on clean problems, under both prompt formats tested. Headline numbers use the clean subset only.
Function-name matching catches exact reuse but misses paraphrased problems, so 26% is a lower bound on overlap, not an estimate.
Limitations
- Synthetic training data. CodeAlpaca-20K is GPT-generated via self-instruct, unverified, stylistically homogeneous. Quality is bounded by the teacher model.
- Conformance over capability. With a function signature specified in the prompt, the base model already reaches 97.5% valid Python and 82.5% pass@1. This adapter's main contribution is removing the need for that prompt engineering.
- Python only. Training data was AST-filtered to Python.
- Short outputs. Median training example was 89 tokens; long generations degrade.
- Known failure modes: repetition loops causing mid-generation truncation; calling helper functions it never defines.
Methodology notes
Three measurement errors found and corrected during evaluation:
Naming confound. Initial pass@1 read 35% base / 50% fine-tuned. 17 of 20 failures were
NameErrorβ correct code under a different function name than the test called. Specifying signatures corrected the baseline by 47 points, to 82.5%.Indentation destruction. Raw-format HumanEval returned 0% for both models. The fence-stripper called
.strip(), removing leading indentation from function bodies. All 50 failures wereIndentationError.Training-data contamination. A keyword-based Python filter kept 58.7% of CodeAlpaca; AST-parsing a 300-example sample showed 38% weren't valid Python β mostly Java and JavaScript matching on shared keywords like
forandclass. Replaced withast.parseplus a syntax-tree check: 0% contamination on re-check.
Sequence packing was disabled after batch inspection revealed it silently disabled completion-only loss masking. Packing requires Flash Attention for block-diagonal masking, which requires Ampere; the T4 is Turing.
- Downloads last month
- 23
Model tree for Raghul09/llama-code-gen-lora
Base model
meta-llama/Llama-3.2-3B