Instructions to use HuggingFaceTB/SmolLM-135M with libraries, inference providers, notebooks, and local apps. Follow these links to get started.

Libraries

How to use HuggingFaceTB/SmolLM-135M with Transformers:

# Use a pipeline as a high-level helper
from transformers import pipeline

pipe = pipeline("text-generation", model="HuggingFaceTB/SmolLM-135M")

# Load model directly
from transformers import AutoTokenizer, AutoModelForCausalLM

tokenizer = AutoTokenizer.from_pretrained("HuggingFaceTB/SmolLM-135M")
model = AutoModelForCausalLM.from_pretrained("HuggingFaceTB/SmolLM-135M")

Notebooks
Google Colab
Kaggle
Local Apps Settings

vLLM

How to use HuggingFaceTB/SmolLM-135M with vLLM:

Install from pip and serve model

# Install vLLM from pip:
pip install vllm
# Start the vLLM server:
vllm serve "HuggingFaceTB/SmolLM-135M"
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:8000/v1/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "HuggingFaceTB/SmolLM-135M",
		"prompt": "Once upon a time,",
		"max_tokens": 512,
		"temperature": 0.5
	}'

Use Docker

docker model run hf.co/HuggingFaceTB/SmolLM-135M

SGLang

How to use HuggingFaceTB/SmolLM-135M with SGLang:

Install from pip and serve model

# Install SGLang from pip:
pip install sglang
# Start the SGLang server:
python3 -m sglang.launch_server \
    --model-path "HuggingFaceTB/SmolLM-135M" \
    --host 0.0.0.0 \
    --port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "HuggingFaceTB/SmolLM-135M",
		"prompt": "Once upon a time,",
		"max_tokens": 512,
		"temperature": 0.5
	}'

Use Docker images

docker run --gpus all \
    --shm-size 32g \
    -p 30000:30000 \
    -v ~/.cache/huggingface:/root/.cache/huggingface \
    --env "HF_TOKEN=<secret>" \
    --ipc=host \
    lmsysorg/sglang:latest \
    python3 -m sglang.launch_server \
        --model-path "HuggingFaceTB/SmolLM-135M" \
        --host 0.0.0.0 \
        --port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "HuggingFaceTB/SmolLM-135M",
		"prompt": "Once upon a time,",
		"max_tokens": 512,
		"temperature": 0.5
	}'

Docker Model Runner
How to use HuggingFaceTB/SmolLM-135M with Docker Model Runner:
```
docker model run hf.co/HuggingFaceTB/SmolLM-135M
```

Model code for training from sractch

by Chrisneverdie - opened Jul 19, 2024

Discussion

Chrisneverdie

Jul 19, 2024

Hi there,

Thank you so much for the amazing work. I wonder if there is code and config available for us to train this model structure from scratch using our own dataset.

Thank you again!

loubnabnl

Hugging Face Smol Models Research org Aug 21, 2024

Hi, we will release the code soon

BasToTheMax

Sep 26, 2024

Hi, we will release the code soon

Hi. Any updates?

eliebak

Sep 26, 2024

Working on this, waiting for some PR to be merge and releasing it soon, for training we use nanotron https://github.com/huggingface/nanotron and here is a gist of the config we use in the meantime https://gist.github.com/eliebak/4263659706519536b7eebfe6d9815c60

BasToTheMax

Sep 26, 2024

Ah okay! Thanks for the quick response!

d3LLM-model

Oct 4, 2025

Hi, thanks for the updates so far! 🙏

Just wanted to kindly check back in — has the full training code been released yet, or is there an ETA on when it might be available? Would love to try training the model from scratch with our own dataset.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.

Tap or paste here to upload images

· Sign up or log in to comment