LogFix SmolLM2 360M

Release logfix-smollm2-360m-v1.0, packaged from experiment version V4 (Failure-Driven Dataset, run run_20260925_010).

Model Description

A LoRA fine-tuned version of HuggingFaceTB/SmolLM2-360M-Instruct specialized for explaining Python errors and debugging scenarios. Given an error message, traceback, or code plus error, it answers in four sections: ERROR TYPE, CAUSE, FIX, PREVENTION.

Base Model

HuggingFaceTB/SmolLM2-360M-Instruct (original weights, loaded unchanged; this repository holds only the LoRA adapter).

Training Method

LoRA (PEFT). Only the adapter weights were trained; the adapter was trained from the original base weights, not continued from an earlier version.

Dataset

  • LogFix Dataset V4 (logfix-v4): Failure-driven dataset: 250 de-duplicated Dataset V2 records plus new hard examples and new examples targeting the weaknesses measured on the frozen evaluation. Validation and test are the shared frozen sets.
  • Examples: 519 training; validation and test are shared frozen sets not trained on
  • Category distribution (largest): Virtual environments 24, pip and packaging 23, ModuleNotFoundError 22, pandas 21, File handling and paths 20, async/await 19, Classes and objects 18, Iteration errors 18, None handling 18, Flask 15, Environment variables 14, NumPy 14, datetime 14, Encoding 13, requests / HTTP APIs 13
  • Difficulty distribution: medium 205, easy 104, hard 210
  • Source: V2 records: synthetic-authored for V2. New records: synthetic-authored for V4 by authors who never saw the frozen set, checked by scripts (schema, format, length, duplicates, leakage); not human-reviewed.
  • The training examples themselves are not published with this model.

Training Configuration

Setting Value
epochs 3.0
learning_rate 0.0001
batch_size 4
gradient_accumulation 2
lora_rank 16
lora_alpha 32
lora_dropout 0.05
max_seq_length 512
seed 42

Evaluation

Two separate benchmarks, both never trained on, both sealed by checksum. Every version is scored with the same prompts and greedy decoding (temperature 0, 256 new tokens). Checks are deterministic: all four sections present, ERROR TYPE matches an accepted label, and which required fix concepts the FIX mentions. They measure format and mentions, not whether a fix is correct. The two benchmarks are reported separately and never averaged together.

Frozen Test (20 examples)

Metric Value
Format compliance (all four sections) 95.0%
Error type accuracy 80.0%
Required fix concept coverage 27.5%
Field completeness (sections present, of 4) 3.80
Test loss (reference answers) 1.9937
Mean / median generation time (CPU) 7.25 s / 6.73 s

Evaluation eval_20260928_050809, 20 examples.

Challenge Set (held-out generalization set)

Metric Value
Format compliance (all four sections) 100.0%
Error type accuracy 88.0%
Required fix concept coverage 28.0%
Field completeness (sections present, of 4) 4.00
Test loss (reference answers) 2.1137
Mean / median generation time (CPU) 14.13 s / 12.79 s

Evaluation eval_20260928_052201, 50 examples.

Known regressions against the previous experiment version on the frozen test:

  • generation seconds vs V3: 6.54 β†’ 7.25
  • category IndexError: fix concept coverage 100% β†’ 0% (1 case)
  • category OSError: fix concept coverage 50% β†’ 0% (1 case)
  • category StopIteration: fix concept coverage 50% β†’ 0% (1 case)

Intended Use

Educational and local Python debugging assistance: a first explanation of an error, checked by a person before any fix is applied.

Limitations

  • The model is built on a base of approximately 360M parameters.
  • It may fail on complex debugging tasks (long or multi-cause tracebacks, framework internals, version-specific behaviour).
  • Generated fixes should be verified.
  • It is not a replacement for testing or code review.
  • Training data is synthetic (written for this project and checked by scripts), and the benchmarks are small.

Version Lineage

Base (SmolLM2-360M-Instruct) β†’ V1 initial fine-tune β†’ V2 dataset experiment β†’ V3 hyperparameter experiment β†’ V4 failure-driven dataset experiment β†’ user-selected release candidate: V4.

Each version was trained independently from the base model. A later version is not necessarily better than an earlier one; the release was chosen by the user from the measured results.

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel

base_model = "HuggingFaceTB/SmolLM2-360M-Instruct"
adapter = "ManojP09/logfix-smollm2-360m"

tokenizer = AutoTokenizer.from_pretrained(base_model)
model = AutoModelForCausalLM.from_pretrained(base_model)
model = PeftModel.from_pretrained(model, adapter)
Downloads last month
25
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for ManojP09/logfix-smollm2-360m

Adapter
(59)
this model