Instructions to use kurakurai/Luth-2-0.8B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use kurakurai/Luth-2-0.8B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="kurakurai/Luth-2-0.8B") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("kurakurai/Luth-2-0.8B") model = AutoModelForMultimodalLM.from_pretrained("kurakurai/Luth-2-0.8B", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use kurakurai/Luth-2-0.8B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "kurakurai/Luth-2-0.8B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "kurakurai/Luth-2-0.8B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/kurakurai/Luth-2-0.8B
- SGLang
How to use kurakurai/Luth-2-0.8B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "kurakurai/Luth-2-0.8B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "kurakurai/Luth-2-0.8B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "kurakurai/Luth-2-0.8B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "kurakurai/Luth-2-0.8B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use kurakurai/Luth-2-0.8B with Docker Model Runner:
docker model run hf.co/kurakurai/Luth-2-0.8B
Luth-2-0.8B
Luth-2-0.8B is a 750M-parameter (text only) non-reasoning model, setting a new state of the art in French for its size across math, code, instruction following, general knowledge and tool calling. It is trained on a 3B-token French SFT mixture followed by multi-domain on-policy distillation (MOPD). The model outperforms every other model in its size class on our selected French benchmarks and stays competitive with models 2 to 3 times larger. It is small enough for efficient local and on-device deployment.
- 📄 Blog: Luth-2: Pushing the French Capabilities of SLMs with MOPD
- 🤗 Models: Luth-2-0.8B · Luth-2-2B
- 📊 Datasets: SFT · RL
- 💻 Code: GitHub
- 🏆 Leaderboard: French LLM Leaderboard
Luth-2-0.8B inherits the VLM architecture of Qwen3.5-0.8B but was not trained on vision data. We do not recommend using it for vision tasks.
Model variants
| Model | Description |
|---|---|
| Luth-2-0.8B | Original checkpoint in native format. Best for fine-tuning or inference with Transformers, vLLM and SGLang. |
| Luth-2-0.8B-GGUF | Quantized format for llama.cpp and compatible tools. Optimized for CPU inference and reduced memory usage. |
Training
Luth-2-0.8B is post-trained from Qwen3.5-0.8B in two stages:
- Supervised fine-tuning on Luth-2-Post-Training-SFT, a 3B-token French mixture spanning math (37.2%), knowledge (27.9%), code (22.2%), instruction following (6.5%) and tool calling (6.3%). Prompts were translated from English SFT datasets and answers regenerated with strong open-source teachers.
- Multi-domain on-policy distillation (MOPD). Three specialists (math, code, instruction following) are trained separately with GRPO on Luth-2-Post-Training-RL, then distilled back into the SFT student.
Inference
Luth-2-0.8B is supported by Transformers, vLLM, SGLang and more.
Quick start with Transformers:
from transformers import AutoModelForCausalLM, AutoTokenizer, TextStreamer
model_id = "kurakurai/Luth-2-0.8B"
model = AutoModelForCausalLM.from_pretrained(
model_id,
device_map="auto",
dtype="bfloat16",
# attn_implementation="flash_attention_2" # uncomment on compatible GPU
)
tokenizer = AutoTokenizer.from_pretrained(model_id)
streamer = TextStreamer(tokenizer, skip_prompt=True, skip_special_tokens=True)
prompt = "Quelle est la capitale de la France?"
input_ids = tokenizer.apply_chat_template(
[{"role": "user", "content": prompt}],
add_generation_prompt=True,
return_tensors="pt",
tokenize=True,
)["input_ids"].to(model.device)
output = model.generate(
input_ids,
do_sample=True,
temperature=0.8,
top_p=0.95,
top_k=20,
max_new_tokens=512,
streamer=streamer,
)
Evaluation
Evaluations can be reproduced using our GitHub repository. The benchmarks are French subsets or verified translations, scored with temperature=0.6, top_p=0.95, top_k=20, thinking disabled, averaged over 10 runs.
| French Benchmarks | Luth-2-0.8B | Luth-0.6B-Instruct | Qwen3.5-0.8B |
|---|---|---|---|
| MGSM-rev2 | 72.92 | 58.52 | 35.20 |
| AIME 24 | 5.67 | 2.00 | 1.00 |
| AIME 25 | 8.67 | 1.33 | 0.33 |
| Math-500 | 57.60 | 44.74 | 27.46 |
| Global-MMLU-Lite | 53.30 | 40.20 | 44.00 |
| MMLU-ProX-Lite | 38.93 | 25.40 | 27.60 |
| GPQA-Diamond | 26.87 | 25.60 | 23.80 |
| IFEval | 71.23 | 51.23 | 44.47 |
| Multi-IF | 61.52 | 33.77 | 32.72 |
| HumanEval+ | 46.81 | 30.25 | 10.87 |
| MBPP+ | 42.33 | 34.74 | 18.20 |
| BFCL v2 | 64.02 | 61.72 | 51.49 |
See the French LLM Leaderboard for comparisons across models.
Contact
Questions or feedback? Reach us on LinkedIn: Maxence Lasbordes and Guillaume Pradel.
Citation
@misc{luth2,
title = {Luth-2: Pushing the French Capabilities of SLMs with MOPD},
author = {Maxence Lasbordes and Guillaume Pradel},
year = {2026},
url = {https://huggingface.co/blog/MaxLSB/luth-2}
}
- Downloads last month
- 134

