Instructions to use MrMoz33/tokioai-3b-cybersec with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use MrMoz33/tokioai-3b-cybersec with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="MrMoz33/tokioai-3b-cybersec")# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("MrMoz33/tokioai-3b-cybersec", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use MrMoz33/tokioai-3b-cybersec with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "MrMoz33/tokioai-3b-cybersec" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "MrMoz33/tokioai-3b-cybersec", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/MrMoz33/tokioai-3b-cybersec
- SGLang
How to use MrMoz33/tokioai-3b-cybersec with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "MrMoz33/tokioai-3b-cybersec" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "MrMoz33/tokioai-3b-cybersec", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "MrMoz33/tokioai-3b-cybersec" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "MrMoz33/tokioai-3b-cybersec", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use MrMoz33/tokioai-3b-cybersec with Docker Model Runner:
docker model run hf.co/MrMoz33/tokioai-3b-cybersec
TokioAI 3B Cybersec
A fine-tuned Qwen2.5-3B-Instruct model for autonomous cybersecurity and DevOps tool calling.
Built by TokioAI -- Intelligence for Evolution.
What is TokioAI?
TokioAI is an autonomous AI agent framework built on a radical philosophy: exploit the native capabilities of the model, don't reinvent them.
The core idea: the LLM already knows how to reason, plan, write code, and analyze problems. What it cannot do alone is act in the real world -- execute commands, connect to servers, read files, query APIs, respond to incidents. TokioAI provides the body for the AI brain:
- The model is the brain -- it thinks, plans, and decides.
- The engine is the nervous system -- ~1,000 lines of pure Python connecting the brain to reality.
- The tools are the hands -- 100+ autonomous tools that interact with the real world.
- The CLI/API is the skin -- the interface between the agent and the human operator.
Radical Minimalism
Every line of code must justify its existence. The entire TokioAI engine is approximately 1,000 lines controlling 100+ tools:
| Component | Lines | Role |
|---|---|---|
agent.py |
~400 | Agent loop: think, act, observe, learn |
registry.py |
~150 | Tool registry and discovery |
executor.py |
~150 | Tool execution and error handling |
loader.py |
~300 | Dynamic tool loading |
Zero external frameworks. No LangChain. No LlamaIndex. No CrewAI. Just pure Python, the model's native tool-calling capability, and carefully crafted tools.
The Body, Not the Mind
TokioAI is not the intelligence. TokioAI is the body. We don't try to make the model smarter -- we give it hands to act. The philosophy is:
"Don't build what the model already knows how to do. Build what the model cannot do alone."
This means: no prompt chains, no retrieval-augmented bloat, no framework overhead. Just a clean loop:
- Think -- the model receives the user request + available tools
- Act -- the model decides which tool to call with what arguments
- Observe -- the engine executes the tool and returns the result
- Learn -- the model incorporates the result and decides next action
This fine-tuned model is the natural evolution of that philosophy: take a small, capable open-source model (Qwen 2.5 3B) and teach it the specific tool-calling patterns that TokioAI uses in production. Instead of depending on closed frontier APIs for every inference, we exploit the model's native ability to learn patterns and specialize it for our exact use case. The same philosophy -- exploit native capabilities, don't add layers on top -- now applied to the model's weights themselves.
Why This Model Exists
TokioAI runs in production with frontier models (Claude, GPT, Gemini) as the brain. They work perfectly. But they are:
- Expensive -- API costs add up for high-volume operations
- Slow -- network latency on every inference call
- External -- dependent on third-party availability and policies
- Closed -- you can't modify the weights, can't specialize behavior
This fine-tuned 3B model aims to handle the most common tool-calling patterns locally, with zero latency and zero cost. The frontier models remain available for complex reasoning, but routine operations (check server status, run a command, read a file, connect via SSH) can run on a quantized local model.
We chose Qwen 2.5 3B because it's open-weight, commercially usable (Apache 2.0), runs on consumer hardware when quantized, and Qwen's architecture has excellent native tool-calling support that we can exploit rather than reinvent.
Training Details
| Parameter | Value |
|---|---|
| Base model | Qwen/Qwen2.5-3B-Instruct |
| Method | QLoRA (4-bit quantization + LoRA adapters) |
| LoRA rank | 16 |
| LoRA alpha | 32 |
| Target modules | q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj |
| Dataset | 120 curated cybersec/DevOps examples |
| Epochs | 3 |
| Batch size | 2 (effective 8 with gradient accumulation) |
| Learning rate | 2e-4 |
| Max sequence length | 2048 |
| Training hardware | NVIDIA T4 16GB |
| Training time | ~10 minutes |
| Training cost | ~$0.50 USD |
| Final loss | ~0.3 |
| Test accuracy | 10/10 (100% correct tool selection and argument generation) |
Tools in Training Data
The model was trained on 12 tools that cover the core TokioAI operational surface:
| Tool | Description | Example Use |
|---|---|---|
execute_local |
Run shell commands on local machine | nmap -sV target, docker ps, systemctl status |
execute_gcp |
Run commands on GCP VM via SSH | kubectl get pods, container deployment |
execute_raspi |
Run commands on Raspberry Pi | IoT monitoring, sensor data, GPIO control |
execute_router |
Run commands on network router | iptables, routing tables, firewall rules |
read_file |
Read files (local or remote) | Config files, logs, source code review |
write_file |
Write/create files | Configs, scripts, playbooks, Terraform, CI/CD |
edit_file |
Edit specific text in files | Patch configs, fix code, update values |
search_files |
Grep/search across files | Find patterns, secrets, vulnerabilities in code |
diagnose |
System health diagnostics | CPU, memory, disk, services, network status |
ssh_connect |
SSH to any server with credentials | Remote execution, multi-server management |
memory |
Persistent memory across sessions | Remember facts, preferences, ongoing context |
task |
Persistent task tracking | Track multi-step projects, resume interrupted work |
Training Data Composition
- 30% Cybersecurity (nmap, vulnerability scanning, log analysis, incident response, firewall rules, hardening)
- 25% DevOps/Infrastructure (Docker, Kubernetes, Terraform, Ansible, CI/CD pipelines, deployments)
- 20% System Administration (disk management, processes, services, networking, troubleshooting)
- 15% SSH Operations (multi-server management, credential handling, remote command execution)
- 10% File Operations (read/write configs, search patterns, edit source code, create scripts)
All examples are bilingual (English and Spanish) reflecting TokioAI's real production usage.
Model Files
Merged Model (Full Weights -- fp16)
The merged/ directory contains the complete model with adapter weights merged into the base. Ready for inference or GGUF conversion.
merged/
config.json
generation_config.json
model-00001-of-00002.safetensors (~3.1 GB)
model-00002-of-00002.safetensors (~3.1 GB)
model.safetensors.index.json
tokenizer.json, tokenizer_config.json, vocab.json, merges.txt
LoRA Adapter Only
The adapter/ directory contains just the QLoRA adapter (~67 MB). Apply this to your own Qwen2.5-3B-Instruct base model.
adapter/
adapter_config.json
adapter_model.safetensors (~67 MB)
tokenizer files...
Usage
With Transformers (Python)
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "MrMoz33/tokioai-3b-cybersec"
tokenizer = AutoTokenizer.from_pretrained(f"{model_id}", subfolder="merged")
model = AutoModelForCausalLM.from_pretrained(f"{model_id}", subfolder="merged", device_map="auto")
system = """You are TokioAI, a cybersecurity and DevOps AI assistant.
You have access to tools: execute_local, execute_gcp, read_file, write_file,
edit_file, search_files, diagnose, ssh_connect, memory, task.
Respond with a JSON tool call: {"tool": "tool_name", "args": {...}}"""
messages = [
{"role": "system", "content": system},
{"role": "user", "content": "Check disk space on my machine"}
]
inputs = tokenizer.apply_chat_template(messages, return_tensors="pt").to(model.device)
outputs = model.generate(inputs, max_new_tokens=256)
print(tokenizer.decode(outputs[0][inputs.shape[-1]:], skip_special_tokens=True))
# {"tool": "execute_local", "args": {"command": "df -h"}}
With Ollama (GGUF -- coming soon)
ollama create tokioai -f Modelfile
ollama run tokioai "scan 192.168.1.0/24 for open ports"
# {"tool": "execute_local", "args": {"command": "nmap -sV 192.168.1.0/24"}}
With LoRA Adapter
from peft import PeftModel
from transformers import AutoModelForCausalLM
base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-3B-Instruct")
model = PeftModel.from_pretrained(base, "MrMoz33/tokioai-3b-cybersec", subfolder="adapter")
Example Outputs
| Input | Model Output |
|---|---|
| "Check if nginx is running" | {"tool": "execute_local", "args": {"command": "systemctl status nginx"}} |
| "Scan 10.0.0.0/24 for open ports" | {"tool": "execute_local", "args": {"command": "nmap -sV 10.0.0.0/24"}} |
| "Read /etc/passwd" | {"tool": "read_file", "args": {"path": "/etc/passwd"}} |
| "SSH to 10.0.0.5 as admin pass Secret123" | {"tool": "ssh_connect", "args": {"host": "10.0.0.5", "username": "admin", "password": "Secret123"}} |
| "Deploy pods on GCP" | {"tool": "execute_gcp", "args": {"command": "kubectl get pods --all-namespaces"}} |
| "Remember that prod DB is on port 5433" | {"tool": "memory", "args": {"action": "append", "content": "prod DB is on port 5433"}} |
| "Diagnostico completo del sistema" | {"tool": "diagnose", "args": {"target": "all"}} |
| "Buscar passwords en los archivos .env" | {"tool": "search_files", "args": {"pattern": "password", "glob": "*.env"}} |
Limitations
- 3B parameters: Handles routine tool-calling well, but not complex multi-step reasoning chains. Use frontier models for those.
- Tool calling specialist: Optimized for selecting the right tool and generating correct JSON arguments. Not a general-purpose chatbot.
- 12 tools: Trained on 12 core TokioAI tools. May not generalize perfectly to arbitrary tool schemas without additional training.
- GGUF pending: Quantized GGUF files (Q4_K_M, Q8_0) for Ollama/llama.cpp are coming. The merged fp16 safetensors are available now for local conversion.
About TokioAI
TokioAI builds autonomous AI agents that act in the real world. Not chatbots -- operators.
From cybersecurity incident response to health monitoring, from autonomous navigation to exploring the frontiers of physics. The core philosophy: the model is the brain, we build the body.
Every agent is a minimal engine (~1,000 lines) that gives the LLM tools to interact with reality. No frameworks. No bloat. Just the model's native capabilities, amplified by precision-crafted tools.
Mission: Protect. Heal. Explore.
- Protect -- Autonomous cybersecurity: WAF, SOAR, red team, vulnerability scanning, incident response, threat intelligence.
- Heal -- AI-powered health monitoring: vital signs today, disease prevention and early detection tomorrow.
- Explore -- Push the boundaries of physics and knowledge with AI as a research partner.
Born in Buenos Aires. Open source at heart.
"No talking about AI. We put it in production."
License
Apache 2.0
Citation
@misc{tokioai-3b-cybersec-2026,
title = {TokioAI 3B Cybersec: Fine-tuned Qwen2.5-3B for Autonomous Tool Calling},
author = {TokioAI Security Research},
year = {2026},
url = {https://huggingface.co/MrMoz33/tokioai-3b-cybersec},
note = {QLoRA r=16 on 120 cybersecurity/DevOps examples, 12 tools, 10/10 test accuracy}
}