Instructions to use HuggingFaceTB/SmolLM2-135M with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use HuggingFaceTB/SmolLM2-135M with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="HuggingFaceTB/SmolLM2-135M")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("HuggingFaceTB/SmolLM2-135M") model = AutoModelForCausalLM.from_pretrained("HuggingFaceTB/SmolLM2-135M", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use HuggingFaceTB/SmolLM2-135M with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "HuggingFaceTB/SmolLM2-135M" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "HuggingFaceTB/SmolLM2-135M", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/HuggingFaceTB/SmolLM2-135M
- SGLang
How to use HuggingFaceTB/SmolLM2-135M with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "HuggingFaceTB/SmolLM2-135M" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "HuggingFaceTB/SmolLM2-135M", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "HuggingFaceTB/SmolLM2-135M" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "HuggingFaceTB/SmolLM2-135M", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use HuggingFaceTB/SmolLM2-135M with Docker Model Runner:
docker model run hf.co/HuggingFaceTB/SmolLM2-135M
OpenBookQA full-test result: 32.8% for SmolLM2-135M under a pinned zero-shot setup
Hi SmolLM team,
I ran the complete 500-item OpenBookQA test split for HuggingFaceTB/SmolLM2-135M to compare a pinned local configuration with the model card’s reported 34.6% result.
Configuration
- Model revision:
1e728f6bee47d2fd840b70f44e58d4ed97c0c32d - Weight SHA-256:
80521b40281d6ce74e35c9282c22539e75aa0ac8578892b2a59955ef78d55da1 - Dataset revision:
allenai/openbookqa@388097ea7776314e93a529163e0fea805b8a6454 - Dataset configuration/split:
main, all 500testrows in pinned source order - Task: zero-shot, four answer-text continuations prefixed with a space
- Metric: directly audited pinned LightEval
loglikelihood_acc_norm_nospacesemantics—summed continuation loglikelihood divided by the continuation’s character count excluding its leading space, followed by argmax - Runtime: LightEval
0.6.0.dev0, Transformers4.45.2, Torch2.4.0+cpu, Accelerate0.34.2 - Effective execution: Windows CPU, BF16 parameters, FP32 logits, batch size 1, maximum length 2,048, model seed 1234, Torch intra-op/inter-op threads both 1
Result
- 164/500 correct
acc_norm = 0.328(32.8%)- Difference from the model-card reference: −1.8 percentage points
- All 500 sample identities and all 2,000 planned likelihood-request identities matched
- All 2,000 likelihoods were finite
- Zero truncation, token-count, character-denominator, normalization, prediction, correctness, or independent-scorer discrepancies
The 95% Wilson interval is approximately 28.83%–37.03%, which contains 34.6%. That interval is only an inferential summary under a binomial/superpopulation interpretation; 32.8% is the exact observed score on this fixed test split.
This is therefore numerically compatible with 34.6% in that limited sense, but it is not an exact reproduction, confirmation, refutation, or equivalence result. The model card identifies LightEval but does not pin the evaluator version, split, backend, or complete configuration used for 34.6%.
Execution and evidence limitations
The run was offline. One expected requests→urllib3 IPv6 capability probe was intercepted before operating-system bind; prohibited socket operations and delegated network operations were zero, and 537 process-tree endpoint samples found no endpoints or audit errors.
A general loaded-module safety gate remained failed because four runtime DLLs coexisted: libiomp5md.dll, libiompstubs5md.dll, libomp140.x86_64.dll, and vcomp140.dll. The run proceeded only under a bounded exact-four-runtime, single-thread waiver. I make no general mixed-runtime safety or cross-machine reproducibility claim.
The sealed local evidence bundle contains 46,045 manifest entries:
- Manifest SHA-256:
4bd24d34a7de4e2b61712bec22c28339683aa70005b982ca26e9f2bf07a9541d - Completion record SHA-256:
9ce0c3e49c59f7fabf8113190a21788b67805ac34491c53c09eef6b826f9a678 - Controller result SHA-256:
8454ba68fbd71462683f8faf2ec2d192d3ac9631ddc643b06bb129a8739f2358 - Independent score-audit SHA-256:
ddac2ee76de0640d0154b36c16b88e3536debc6340560f03df94bf39f1dcba71
No raw artifacts are attached. The bundle includes dataset/model payloads, third-party runtime components, and machine-local paths; its redistribution, dependency-license, and privacy review is not complete.
If the evaluator revision, split, backend, or configuration used for the reported 34.6% is available, I would appreciate those details so the remaining configuration delta can be assessed.