Instructions to use roneneldan/TinyStories-1M with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use roneneldan/TinyStories-1M with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="roneneldan/TinyStories-1M")# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("roneneldan/TinyStories-1M") model = AutoModelForCausalLM.from_pretrained("roneneldan/TinyStories-1M", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use roneneldan/TinyStories-1M with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "roneneldan/TinyStories-1M" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "roneneldan/TinyStories-1M", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/roneneldan/TinyStories-1M
- SGLang
How to use roneneldan/TinyStories-1M with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "roneneldan/TinyStories-1M" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "roneneldan/TinyStories-1M", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "roneneldan/TinyStories-1M" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "roneneldan/TinyStories-1M", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use roneneldan/TinyStories-1M with Docker Model Runner:
docker model run hf.co/roneneldan/TinyStories-1M
Actual number of parameters?
Why is this model called -1M when it appears the actual number of parameters is 3745984?
model = AutoModelForCausalLM.from_pretrained("roneneldan/TinyStories-1M")
sum(p.numel() for p in model.parameters())
3745984
Or if you exclude the 3216448-parameter token embedding matrix (by far the bulk of the total parameters), the number of other parameters is 529536. But that's more like 500k than 1M. So why isn't this named either TinyStories-4M or TinyStories-500k? What does the -1M refer to?
Oh, I think I may have figured it out!
From the TinyStories paper:
"Our models are available on Huggingface named TinyStories-1M/3M/9M/28M/33M/1Layer/2Layer and TinyStories-Instruct-β. We
use GPT-Neo architecture with window size 256 and context length 512. We use GPT-Neo tokenizer but only keep the top 10K most
common tokens."
So, they only kept the top 10K most common tokens for the training. But the models here have the full vocabulary size 50257 for their embedding matrices. So I guess for distribution the trained models were sort of filled out (with what, zeros? garbage?) to plug-and-play into a much more common tokenizer?
The math works out, since instead of 3216448 (embedding matrix) + 529536 = 3745984 we would then have 640000 + 529536 = 1169536. This makes a lot more sense to me as a "-1M" model so I bet this is how it was trained.
Yes, but do you know to to limit the tokenizer to 10k tokens? Seems not to be trivial.