GPT-2 Instruct

This project uses a GPT-2 model (124M parameters) fine-tuned on the SVAMP (Simple Variants of Arithmetic Math word Problems) model, available at FurkanNar/gpt-2_svamp.

Training Hyperparameters

  • Base Model: GPT-2 (124M parameters)
  • Dataset: Alpaca
  • Max Sequence Length: 256 tokens
  • Epochs: 3
  • Batch Size: 4
  • Learning Rate: 5e-5
  • Optimizer: AdamW
  • Loss Function: CrossEntropyLoss (tokens shifted by 1)
  • Gradient Clipping Max Norm: 1.0

Training Progress

Epoch 1/3

  • Average Train Loss: 0.7349
  • Average Validation Loss: 0.6526
  • Validation F1 Score (Macro): 0.2824

Epoch 2/3

  • Average Train Loss: 0.6583
  • Average Validation Loss: 0.6436
  • Validation F1 Score (Macro): 0.2839

Epoch 3/3

  • Average Train Loss: 0.6216
  • Average Validation Loss: 0.6411
  • Validation F1 Score (Macro): 0.2857

Training Metrics Visualization

Training Metrics

The model showed consistent improvement across epochs, with training loss decreasing from 0.7349 to 0.6216, indicating effective learning of the instruction-following task.

Model Files

  • config.json - Model configuration
  • generation_config.json - Generation parameters
  • model.safetensors - Fine-tuned model weights (475MB)

Usage

The application uses the official Alpaca instruction format for inference:

Below is an instruction that describes a task. Write a response that appropriately completes the request.

### Instruction:
{user_input}

### Response:

Proof of Concept

Example interaction with the fine-tuned model:

You: Hello
AI: Hi there! How can I help you today?

Configuration

The model generation can be configured with the following parameters:

  • model_name: Hugging Face model identifier or local path
  • system_prompt: System prompt for the assistant
  • max_length: Maximum response length
  • temperature: Sampling temperature (default: 0.5)
  • top_k: Top-k sampling parameter (default: 40)
  • top_p: Nucleus sampling parameter (default: 0.9)
  • repetition_penalty: Penalty for repeating tokens (default: 1.2)

Requirements

  • PyTorch
  • Transformers
  • CUDA-capable GPU

Notes

  • The model uses conversation history (last 10 messages) to maintain context
  • Generation parameters are tuned to reduce data leakage and improve response quality
  • CUDA GPU is required for inference
Downloads last month
267
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for FurkanNar/gpt-2_instruct

Finetuned
(2289)
this model

Datasets used to train FurkanNar/gpt-2_instruct