Instructions to use VantoraLabs/Vantora-Micro with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use VantoraLabs/Vantora-Micro with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="VantoraLabs/Vantora-Micro")# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("VantoraLabs/Vantora-Micro") model = AutoModelForCausalLM.from_pretrained("VantoraLabs/Vantora-Micro", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use VantoraLabs/Vantora-Micro with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "VantoraLabs/Vantora-Micro" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "VantoraLabs/Vantora-Micro", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/VantoraLabs/Vantora-Micro
- SGLang
How to use VantoraLabs/Vantora-Micro with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "VantoraLabs/Vantora-Micro" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "VantoraLabs/Vantora-Micro", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "VantoraLabs/Vantora-Micro" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "VantoraLabs/Vantora-Micro", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use VantoraLabs/Vantora-Micro with Docker Model Runner:
docker model run hf.co/VantoraLabs/Vantora-Micro
license: mit
language:
- en
tags:
- llama
- text-generation
- micro
- nano
- small
- base-model
- banana-mind
pipeline_tag: text-generation
library_name: transformers
datasets:
- HuggingFaceFW/fineweb-edu
Vantora Micro
A 9,800-parameter pure Llama-style causal language model, trained on a 100M-token slice of FineWeb-Edu. This is the "pure transformer" baseline in a head-to-head comparison against a hybrid Mamba-2 + attention model of the same size.
Model Details
| Property | Value |
|---|---|
| Architecture | LlamaForCausalLM (pure transformer) |
| Parameters | 9,800 |
| Vocab size | 1024 (ByteLevel BPE) |
| Hidden size (d_model) | 8 |
| Intermediate size | 22 (MLP ratio 2.77) |
| Hidden layers | 2 |
| Attention heads | 1 |
| Head dim | 8 |
| Context length | 512 |
| RoPE theta | 10000.0 |
| RMSNorm eps | 1e-6 |
| Tied embeddings | Yes |
| Dtype | float32 |
Training
- Data: first 100M tokens of
HuggingFaceFW/fineweb-edusample-10BT - Epochs: 1 (100M total tokens seen)
- Batch: 128 × seq 256 (3,051 steps)
- Optimizer: AdamW, lr 5e-3, cosine schedule + 15% warmup
- Grad clip: 1.0, seed 42
- Hardware: NVIDIA GTX 750 (Maxwell, 4 GB VRAM)
- Time: ~4.3 minutes
Benchmark: BananaMind Base Bench 1.1
Evaluated with the official BananaMind benchmark.py runner
(official_complete_run: true, exact 350-item split, SHA-256 verified).
| Metric | Value |
|---|---|
| Overall Elo | 810 |
| Accuracy | 26.00% (91/350) |
| Weighted accuracy | 25.48% |
| Category | Elo | Accuracy |
|---|---|---|
| Language Completion | 919 | 52.0% |
| Commonsense | 658 | 16.0% |
| World Knowledge | 702 | 20.0% |
| Context Tracking | 665 | 12.0% |
| Quantitative | 875 | 26.0% |
| Logical Reasoning | 839 | 22.0% |
| Code Completion | 982 | 34.0% |
⚠️ Length-bias caveat on Code Completion
The Code Completion score (Elo 982, 34%) is not evidence the model can code. It is a benchmark artifact:
- In the
code_completioncategory, the correct answer is the longest continuation 68% of the time (vs 12-34% in every other category). - This model has a length bias: it picks the longest continuation more often than random, because its token distribution is near uniform and longer sequences accumulate more probability.
- The two effects line up, so the length bias coincidentally matches the correct answer most of the time.
The BananaMind README itself warns: "Mean token log-probability reduces direct continuation-length bias but does not eliminate every tokenizer-dependent effect." Treat the Code Completion Elo as a length-bias artifact, not a real coding skill.
vs. Vantora-Micro-Hybrid (same size/data)
| Vantora-Micro | Vantora Micro Hybrid | |
|---|---|---|
| Params | 9,800 | 11,256 |
| Overall Elo | 810 | 863 |
| Accuracy | 26.00% | 30.29% |
| Val loss (edu) | 4.9097 | 4.8584 |
| Training time | ~4.3 min | ~49.5 min |
The hybrid edges out this model by +53 Elo and +4.3% accuracy, but most of that gap comes from the length-bias artifact on Code Completion, not real reasoning. On PIQA / HellaSwag / ARC-Easy the two are within noise (0.5-2%).
Why this model is the practical choice
For a 10K-param model on a 100M-token slice of web text, the pure transformer is the better tradeoff:
- 11.5× faster to train (4.3 min vs 49.5 min on the same GTX 750).
- Within noise of the hybrid on every benchmark that measures real ability.
- No custom architecture, no
trust_remote_code, loads with stockAutoModelForCausalLM.
The hybrid's SSM sequence memory was a clear win on TinyStories (where narrative memory mattered), but on this benchmark the extra training time buys almost nothing. This model gets the same result in a fraction of the time.
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("VantoraLabs/Vantora-Micro")
tokenizer = AutoTokenizer.from_pretrained("VantoraLabs/Vantora-Micro")
prompt = "Once upon a time"
inputs = tokenizer(prompt, return_tensors="pt")
outputs = model.generate(**inputs, max_new_tokens=50)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
Files
config.json # LlamaForCausalLM config
model.safetensors # 9,800-param weights
tokenizer.json # ByteLevel BPE (1024 vocab)
tokenizer_config.json # tokenizer settings
special_tokens_map.json # special token mapping
generation_config.json # generation defaults
Notes
This is an extremely small model — it is a research artifact for studying scaling laws and architecture comparisons at the sub-10K parameter scale, not a production language model. Its BananaMind score (Elo 810) is near the four-choice random baseline (25%), as expected for a model this size.