File size: 3,148 Bytes
60199c0
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
092272c
 
 
 
 
 
 
 
 
 
 
 
 
 
27c29da
092272c
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
---
language: en
license: mit
library_name: transformers
pipeline_tag: text-generation
tags:
- gpt2
- code-completion
- pytorch
- huggingface
- transformers
- codexglue
- nlp
- machine-learning
---


# GPT-2 Fine-Tuned for Python Code Completion

This repository contains a fine-tuned **GPT-2** model for **Python source code completion**. The model was trained on the **CodeXGLUE Python Code Completion** dataset using the Hugging Face Transformers library and PyTorch.

## Model Description

This model is designed to predict the next tokens in Python source code, enabling intelligent code completion for software development tasks.

- **Base Model:** GPT-2
- **Task:** Causal Language Modeling
- **Language:** Python
- **Framework:** PyTorch
- **Library:** Hugging Face Transformers

---

## Dataset

**Dataset:** CodeXGLUE – Python Code Completion

The dataset contains Python source code snippets used to train language models for next-token code prediction.

Note: A subset of approximately 13,000 training samples from the CodeXGLUE Python dataset was used for fine-tuning.

---

## Training Configuration

| Parameter | Value |
|-----------|--------|
| Model | GPT-2 |
| Epochs | 3 |
| Learning Rate | 2e-4 |
| Batch Size | 4 |
| Gradient Accumulation | 4 |
| Weight Decay | 0.01 |
| Max Sequence Length | 512 |
| Optimizer | AdamW |
| Framework | PyTorch |

Training was performed using the Hugging Face `Trainer` API.

---

## Evaluation Results

| Metric | Value |
|---------|---------|
| Validation Loss | **1.1869** |
| Perplexity | **3.28** |

The decreasing validation loss throughout training indicates successful adaptation of GPT-2 to the Python code completion task.

---

## Training Progress

| Step | Training Loss | Validation Loss |
|------|---------------|----------------|
|100|1.5613|1.3592|
|200|1.3962|1.2877|
|300|1.3317|1.2537|
|400|1.2437|1.2308|
|500|1.2253|1.2142|
|600|1.2014|1.2000|
|Final|—|**1.1869**|

---

## Usage

```python
from transformers import AutoTokenizer, AutoModelForCausalLM

tokenizer = AutoTokenizer.from_pretrained("YOUR_USERNAME/MODEL_NAME")
model = AutoModelForCausalLM.from_pretrained("YOUR_USERNAME/MODEL_NAME")

prompt = "def fibonacci(n):"

inputs = tokenizer(prompt, return_tensors="pt")

outputs = model.generate(
    **inputs,
    max_new_tokens=50,
    do_sample=True,
    temperature=0.7
)

print(tokenizer.decode(outputs[0], skip_special_tokens=True))
```

---

## Limitations

- Trained only on Python source code.
- Intended for research and educational purposes.
- May generate syntactically incorrect or incomplete code.
- Does not guarantee production-quality code suggestions.

---

## Technologies Used

- Python
- PyTorch
- Hugging Face Transformers
- Hugging Face Datasets
- CodeXGLUE Dataset

---

## Future Improvements

- Fine-tune larger transformer models.
- Train on larger subsets of CodeXGLUE.
- Evaluate using additional code generation metrics.
- Support multiple programming languages.
- Deploy as an inference API.

---

## Author

**Sai Nandu Vajhala**

GitHub: https://github.com/SaiNanduVajhala

LinkedIn: https://www.linkedin.com/in/sai-nandu-vajhala