Instructions to use nesilabs/tosilos-24b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use nesilabs/tosilos-24b with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf nesilabs/tosilos-24b:Q4_K_M # Run inference directly in the terminal: llama cli -hf nesilabs/tosilos-24b:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf nesilabs/tosilos-24b:Q4_K_M # Run inference directly in the terminal: llama cli -hf nesilabs/tosilos-24b:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf nesilabs/tosilos-24b:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf nesilabs/tosilos-24b:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf nesilabs/tosilos-24b:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf nesilabs/tosilos-24b:Q4_K_M
Use Docker
docker model run hf.co/nesilabs/tosilos-24b:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use nesilabs/tosilos-24b with Ollama:
ollama run hf.co/nesilabs/tosilos-24b:Q4_K_M
- Unsloth Studio
How to use nesilabs/tosilos-24b with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for nesilabs/tosilos-24b to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for nesilabs/tosilos-24b to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for nesilabs/tosilos-24b to start chatting
- Pi
How to use nesilabs/tosilos-24b with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf nesilabs/tosilos-24b:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "nesilabs/tosilos-24b:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- OpenClaw new
How to use nesilabs/tosilos-24b with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf nesilabs/tosilos-24b:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "nesilabs/tosilos-24b:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Docker Model Runner
How to use nesilabs/tosilos-24b with Docker Model Runner:
docker model run hf.co/nesilabs/tosilos-24b:Q4_K_M
- Lemonade
How to use nesilabs/tosilos-24b with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull nesilabs/tosilos-24b:Q4_K_M
Run and chat with the model
lemonade run user.tosilos-24b-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use nesilabs/tosilos-24b with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf nesilabs/tosilos-24b:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default nesilabs/tosilos-24b:Q4_K_M
Run Hermes
hermes
- Atomic Chat
Tosilos-24b — home-runnable European cybersecurity model (Mistral Devstral 24B, fine-tuned)
Full 24B cybersecurity model (QLoRA r=64 alpha=128 fused into Devstral-Small-2505), trained on the Tosilos curated corpus (17k domain examples + 25% general replay). Designed to run locally on a single 32GB-class GPU (RTX 5090) in 4-bit.
Benchmarks (same harness, 500 questions each; judge: Opus 5, blind, randomized)
| Metric | Devstral 24B base | Tosilos-24b |
|---|---|---|
| CyberMetric | 91.6% | 92.8% (+1.2pp) |
| MMLU (general knowledge) | 77.2% | 77.4% (no forgetting) |
| Domain judge (Opus 5, 1-10) | 5.49 | 5.68 (+0.19, 26W/23L/8T) |
Tosilos-24b improves over its base on all three axes — modest but consistent. Recipe note: the 2-epoch run (834 steps over 6,666 examples) beat the 1-epoch XL run (17k examples), which had matched the base — epoch count mattered more than dataset size.
Usage (home server, 1× 32GB GPU)
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model = AutoModelForCausalLM.from_pretrained(
"nesilabs/tosilos-24b", load_in_4bit=True, device_map="auto")
tok = AutoTokenizer.from_pretrained("nesilabs/tosilos-24b")
Or vLLM tensor-parallel across 2 GPUs: vllm serve nesilabs/tosilos-24b -tp 2.
Files
model-*.safetensors— full merged model, bf16 (48GB). A GGUF Q4 build was evaluated but not published: current llama.cpp conversions of this architecture (tekken tokenizer, rope_theta=1e9) degrade output. Prefer the safetensors path until llama.cpp support matures.
Training recipe
QLoRA r=64 alpha=128, targets q/k/v/o + gate/up/down, 4-bit NF4, seq 4096, 2 epochs on 6,666 examples, lr 1e-4 cosine, 1×H200, 834 steps. Base: mistralai/Devstral-Small-2505.
Dual-use: intended for authorized security work only.
Disclaimer & responsible use
This model is released strictly for authorized security testing, research and education. Offensive security techniques are dual-use.
- You are solely responsible for how you use this model. Only use it against systems you own or have explicit, written authorization to test, and comply with all applicable laws and regulations.
- The authors and nesilabs accept no liability for any misuse, damage, or consequences arising from the use of this model. Use is entirely at your own risk.
- The model is provided "as is", without warranty of any kind, express or implied.
By downloading or using this model you accept these terms.
- Downloads last month
- 65