KoarAI LFM2.5-350M Thinking Banner

🐨 KoarAI / LFM2.5-350M-Thinking

License: Apache 2.0 Model Revision Fine-Tuning: 100% Full Weights Parameters GGUF Quantized

📌 Release Note: Model Code 0002 (Weight Architecture Update)

Model Code: 0002
This version underwent a comprehensive 100% Full Parameter Fine-Tuning across 9 epochs with a cosine learning rate scheduler. It integrates an expanded multi-teacher dataset (Reasoning CoT + DeepSeek-V4-Pro Agentic + MMLU-Pro + AIME 2026 Mathematics) and strict syntactic normalization for <think> ... </think> blocks.

🚀 KoarAI Release & Versioning Policy: Starting from the upcoming release (0003 and beyond), rather than overwriting existing models, each new iteration will be released into its own dedicated repository (e.g., KoarAI/LFM2.5-350M-Thinking-v3, KoarAI/LFM2.5-350M-Thinking-RU, etc.).


🌟 Overview

KoarAI/LFM2.5-350M-Thinking (Code: 0002) is an ultra-compact, high-efficiency language model featuring native Chain-of-Thought (CoT) reasoning capabilities.

Built upon the state-of-the-art Liquid Foundation Model architecture (LiquidAI/LFM2.5-350M), this model was trained using 100% Full Parameter Fine-Tuning on a balanced blend of distilled reasoning traces from frontier models:

  • Qwen 3.8 Max
  • GLM 5.2
  • Kimi K3
  • DeepSeek-V4-Pro 0813 Agentic
  • MMLU-Pro & AIME 2026 Mathematics

Despite having only 350 Million parameters, the model demonstrates strong multi-step logic, mathematical deduction, and structured problem-solving inside native <think> ... </think> blocks.


💡 Native Thinking Mode

The model natively reasons before outputting its final response:

<|im_start|>user
Solve: 32 + 32 - 42<|im_end|>
<|im_start|>assistant
<think>
1. Evaluate 32 + 32 = 64.
2. Subtract 42 from 64: 64 - 42 = 22.
</think>
\boxed{22}<|im_end|>

⚡ Quickstart

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "KoarAI/LFM2.5-350M-Thinking"

tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.float16,
    device_map="auto",
    trust_remote_code=True
)

messages = [
    {"role": "user", "content": "How many 'r' in strawberry?"}
]

prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)

outputs = model.generate(
    **inputs,
    max_new_tokens=512,
    temperature=0.6,
    top_p=0.9,
    do_sample=True,
    pad_token_id=tokenizer.eos_token_id
)

print(tokenizer.decode(outputs[0][inputs['input_ids'].shape[1]:], skip_special_tokens=False))

📦 GGUF & Quantization

Official quantized GGUF versions (FP16, Q8_0, Q5_K_M, Q4_K_M, Q4_0) for llama.cpp, Ollama, and LM Studio are available at:
👉 KoarAI/LFM2.5-350M-Thinking-GGUF


🐨 Maintained by KoarAI Lab

Released for the open-source AI community by KoarAI.

Downloads last month
-
Safetensors
Model size
0.4B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for KoarAI/LFM2.5-350M-Thinking

Finetuned
(62)
this model
Quantizations
2 models

Collection including KoarAI/LFM2.5-350M-Thinking