File size: 7,133 Bytes
d2489d9
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
90645f4
d2489d9
90645f4
 
 
d2489d9
 
 
 
 
 
 
 
 
 
 
 
 
 
 
47b4dd8
d2489d9
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
---
license: other
license_name: lfm1.0
license_link: https://huggingface.co/LiquidAI/LFM2.5-230M/blob/main/LICENSE
base_model: LiquidAI/LFM2.5-230M
tags:
  - lfm2
  - lfm2.5
  - liquid
  - code
  - math
  - fine-tune
  - experimental
language:
  - en
pipeline_tag: text-generation
---

# LFM2.5-230M-Code-Math-Exp

An successful experimental fine-tune of [LiquidAI/LFM2.5-230M](https://huggingface.co/LiquidAI/LFM2.5-230M) (the instruct-tuned model, **not** the base checkpoint), focused on strengthening code generation across multiple languages, while retaining the general chat and instruction-following ability of the original instruct model.

Marked **successful and experimental** because training was stopped early (see below) and the model has only been evaluated on a fixed prompt suite, not a formal benchmark.

Shows major improvements over my previous model (hauser458b/lfm2.5-230m-code-math), delivering broader programming language support, better full-app generation, and reliable basic agentic behavior like sandbox file creation and terminal execution (tested locally via Q5_K_M quant).

## Why this exists

LiquidAI's own model card for LFM2.5-230M states it is **not recommended for reasoning-heavy workloads such as advanced math, code generation, or creative writing** β€” the model is tuned primarily for data extraction, structured outputs, and lightweight agentic/tool-use tasks. This fine-tune pushes the small instruct model further into multi-language code competence.

Fine-tuning started from the **instruct** checkpoint rather than the base pretrain checkpoint, to preserve chat and instruction-following behavior that the base model doesn't have.

## Training details

- **Base model**: `LiquidAI/LFM2.5-230M` (instruct)
- **Method**: Full fine-tune, bf16
- **Datasets**:
  - [`ise-uiuc/Magicoder-Evol-Instruct-110K`](https://huggingface.co/datasets/ise-uiuc/Magicoder-Evol-Instruct-110K) β€” random 60k-example subset
  - [`iamtarun/code_instructions_120k_alpaca`](https://huggingface.co/datasets/iamtarun/code_instructions_120k_alpaca) β€” full dataset
  - [`openai/gsm8k`](https://huggingface.co/datasets/openai/gsm8k) (main split) β€” full dataset
- **Epochs / batch size**: configured for 5 epochs, batch size 16, but training was stopped early at 9,800 steps β€” checkpoints beyond ~12k steps were tested and showed clearly worse output quality (less coherent, less accurate generations), so 9,800 was kept as the published checkpoint instead.
- **Sequence length**: 1024 tokens
- **Loss**: completion-only (loss computed only on assistant responses, not prompts)
- **Checkpoint**: saved directly from the training run at the stopping point (not selected via `load_best_model_at_end`, since the run was interrupted rather than completed) β€” see Known Limitations below

## Evaluation

Evaluated on a fixed suite of 90 hand-written prompts: 20 Python, 20 Java, 5 each of HTML/CSS/JavaScript, 5 combined HTML+CSS+JS mini-projects, 10 C++, 10 Rust, and 10 larger "build a complete program" prompts (calculator, to-do app, quiz game, etc.), compared against [`hauser458b/lfm2.5-230m-code-math`](https://huggingface.co/hauser458b/lfm2.5-230m-code-math).

This is a manual, single-run, single-temperature comparison (temperature 0.3) β€” not a formal benchmark, and results are anecdotal rather than statistically rigorous.

**Where this model did better:**
- **Java**: fewer broken/non-compiling outputs (e.g. no literal syntax errors, correct handling of tasks like title-casing and leap-year checks)
- **JavaScript**: more coherent DOM manipulation logic on the "add list item" style task
- **Combined HTML+CSS+JS mini-projects**: more functional end-to-end logic (e.g. quiz-answer checking actually evaluates a real answer rather than looping over hardcoded strings)
- **Full "build a project" prompts**: calculator, to-do app, and rock-paper-scissors implementations were more likely to be logically complete and runnable

**Where it was roughly even:** HTML, CSS, C++

**Where it did NOT clearly improve, or regressed:**
- **Rust remains weak in both models.** Neither dataset used here was Rust-heavy, and this shows: several Rust outputs from this model don't compile as-is (e.g. missing braces, incorrect struct/method logic). It's somewhat less broken than the comparison model's Rust output, but "less broken" is not "reliable" β€” treat all Rust output from this model as a rough draft, not working code.

## Known limitations

- **Early stopping**: training was manually interrupted before the configured 5 epochs completed, and the final checkpoint was saved directly rather than through automatic best-checkpoint selection. There is no guarantee this exact checkpoint is the single best point in the run β€” it's simply where training was stopped.
- **Rust output is unreliable** β€” expect syntax errors and logic bugs; always compile-check before trusting Rust output from this model.
- **Occasional logic bugs even in "solid" languages** β€” e.g. one Python anagram implementation in testing called a string method that doesn't exist. Review generated code before running it, especially anything involving user input, file I/O, or external calls.
- **Multi-file / multi-technology tasks** (e.g. combined HTML+CSS+JS builds) are the model's weakest spot even where it improved over the comparison model β€” interactivity is sometimes only partially wired up (e.g. a button that updates the wrong element).
- Not evaluated on data extraction, RAG, tool-calling, or general chat β€” this fine-tune was evaluated purely on code-generation prompts. If you need the original model's data-extraction / agentic strengths, this model has not been validated for that use case.
- Still a 230M-parameter model β€” do not expect deep multi-step reasoning or production-grade code without review.
- Not evaluated on safety-critical, medical, or legal use cases β€” do not use for those without additional safeguards.

## Usage

```python
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "hauser458b/lfm2.5-230m-code-math-exp"
model = AutoModelForCausalLM.from_pretrained(model_id, dtype="bfloat16", device_map="auto")
tokenizer = AutoTokenizer.from_pretrained(model_id)

messages = [{"role": "user", "content": "Write a Python function to check if a number is prime."}]
inputs = tokenizer.apply_chat_template(
    messages, add_generation_prompt=True, return_tensors="pt", return_dict=True
).to(model.device)

output = model.generate(**inputs, max_new_tokens=300, do_sample=True, temperature=0.3, top_p=0.9)
print(tokenizer.decode(output[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True))
```

## License

Inherits the [LFM Open License v1.0](https://huggingface.co/LiquidAI/LFM2.5-230M/blob/main/LICENSE) from the base model.

## Acknowledgements

Built on [LiquidAI/LFM2.5-230M](https://huggingface.co/LiquidAI/LFM2.5-230M). See the [LFM2 Technical Report](https://arxiv.org/abs/2511.23404) for details on the base architecture. Compared against [`hauser458b/lfm2.5-230m-code-math`](https://huggingface.co/hauser458b/lfm2.5-230m-code-math) during evaluation.