Instructions to use percepta-ai/spotlight-vm with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use percepta-ai/spotlight-vm with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="percepta-ai/spotlight-vm", trust_remote_code=True)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("percepta-ai/spotlight-vm", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use percepta-ai/spotlight-vm with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "percepta-ai/spotlight-vm" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "percepta-ai/spotlight-vm", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/percepta-ai/spotlight-vm
- SGLang
How to use percepta-ai/spotlight-vm with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "percepta-ai/spotlight-vm" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "percepta-ai/spotlight-vm", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "percepta-ai/spotlight-vm" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "percepta-ai/spotlight-vm", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use percepta-ai/spotlight-vm with Docker Model Runner:
docker model run hf.co/percepta-ai/spotlight-vm
Spotlight VM
A small, hand-constructed Spotlight transformer that runs Python:
the examples from Growing Intelligence Beyond the Weights.
intelligence/ holds the weights (under 100K parameters), which never change. memory/ holds MicroPython
v1.28.0 (core.safetensors) and two packages.
cc -O3 -std=c99 spotlight.c -lm -o spotlight
./spotlight < examples/euler1.py # 233168; -v also prints every generated token
cat examples/tax_brackets/*.py | ./spotlight
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("percepta-ai/spotlight-vm", trust_remote_code=True)
tokenizer = AutoTokenizer.from_pretrained("percepta-ai/spotlight-vm", trust_remote_code=True)
output = model.generate(**tokenizer("print(1 + 2)\n", return_tensors="pt"))
print(tokenizer.decode(output[0])) # 3
In a local copy of this folder, from_pretrained(".") works the same way.
| Example | In the post | Output |
|---|---|---|
examples/euler1.py, euler2.py, euler3.py |
Project Euler problems 1 to 3 | 233168, 4613732, 6857 |
examples/memory_loader.py |
Loading memory selectively | 4 |
examples/tax_brackets/ |
Acquiring and updating knowledge | $17400, then $16914 |
examples/mnist_digit.py |
Growing capabilities and intelligence | 7 |
It runs on one CPU core at about 120K tokens per second: minutes for the Project Euler examples, about half an hour for MNIST (245M tokens). This MicroPython build has no floating point. License: Apache 2.0, see LICENSE.md.
- Downloads last month
- 33