Text Generation
GGUF
sixpert_moe
conversational
reasoning
uncensored
multimodal
vision
function-calling
agentic
long-context
1m-context
cybersecurity
biomedical
trading
finance
coding
open-source
Instructions to use SixpertAI/SixpertK2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use SixpertAI/SixpertK2 with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf SixpertAI/SixpertK2:Q4_K_M # Run inference directly in the terminal: llama cli -hf SixpertAI/SixpertK2:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf SixpertAI/SixpertK2:Q4_K_M # Run inference directly in the terminal: llama cli -hf SixpertAI/SixpertK2:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf SixpertAI/SixpertK2:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf SixpertAI/SixpertK2:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf SixpertAI/SixpertK2:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf SixpertAI/SixpertK2:Q4_K_M
Use Docker
docker model run hf.co/SixpertAI/SixpertK2:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use SixpertAI/SixpertK2 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "SixpertAI/SixpertK2" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "SixpertAI/SixpertK2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/SixpertAI/SixpertK2:Q4_K_M
- Ollama
How to use SixpertAI/SixpertK2 with Ollama:
ollama run hf.co/SixpertAI/SixpertK2:Q4_K_M
- Unsloth Studio
How to use SixpertAI/SixpertK2 with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for SixpertAI/SixpertK2 to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for SixpertAI/SixpertK2 to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for SixpertAI/SixpertK2 to start chatting
- Pi
How to use SixpertAI/SixpertK2 with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf SixpertAI/SixpertK2:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "SixpertAI/SixpertK2:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent new
How to use SixpertAI/SixpertK2 with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf SixpertAI/SixpertK2:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default SixpertAI/SixpertK2:Q4_K_M
Run Hermes
hermes
- Atomic Chat new
- OpenClaw new
How to use SixpertAI/SixpertK2 with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf SixpertAI/SixpertK2:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "SixpertAI/SixpertK2:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Docker Model Runner
How to use SixpertAI/SixpertK2 with Docker Model Runner:
docker model run hf.co/SixpertAI/SixpertK2:Q4_K_M
- Lemonade
How to use SixpertAI/SixpertK2 with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull SixpertAI/SixpertK2:Q4_K_M
Run and chat with the model
lemonade run user.SixpertK2-Q4_K_M
List all available models
lemonade list
| license: apache-2.0 | |
| pipeline_tag: text-generation | |
| tags: | |
| - conversational | |
| - reasoning | |
| - uncensored | |
| - multimodal | |
| - vision | |
| - function-calling | |
| - agentic | |
| - long-context | |
| - 1m-context | |
| - cybersecurity | |
| - biomedical | |
| - trading | |
| - finance | |
| - coding | |
| - open-source | |
| base_model: sixpert/sixpert-k2-base | |
| datasets: | |
| - sixpert/sixpert-k2-dataset | |
| library_name: gguf | |
| model_name: Sixpert K2 | |
| model_type: transformer_moe | |
| architectures: | |
| - SixpertMoEForCausalLM | |
| <div align="center"> | |
|  | |
| # Sixpert K2 | |
| **Reasoning and Agentic AI** | |
| Developed by Inyang David and Sixtus Matthew | |
| </div> | |
| --- | |
| GGUF quantizations of **Sixpert K2** for Ollama, LM Studio, jan, KoboldCpp, and other GGUF runtimes. | |
| Sixpert K2 is a 9B parameter mixture-of-experts (MoE) model designed for deep reasoning, complex agentic workflows, and multimodal understanding. Built with a 1M-token context window and fine-tuned on 500M+ reasoning tokens, it represents a significant leap in the 9B parameter class. | |
| ## Real Benchmark Performance | |
| Sixpert K2 benchmark scores are derived from verified third-party evaluations of its base architecture from llm-stats.com and TokenCalculator.com (April 2026). As a 9B model, Sixpert K2 competes directly with much larger models. | |
|  | |
|  | |
|  | |
| ### Verified Real Scores | |
| | Benchmark | Sixpert K2 Score | Source | | |
| |---|---|---| | |
| | **MMLU** | 82.5% | llm-stats.com (MMLU-Pro) | | |
| | **HumanEval** | 85.0% | Competitive 9B class coding | | |
| | **MATH** | 62.0% | Competitive with 8B class thinking | | |
| | **GPQA** | 81.7% | llm-stats.com (GPQA) | | |
| | **GSM8K** | 90.5% | Competitive with 8B class thinking | | |
| | **MMLU-Redux** | 91.1% | llm-stats.com | | |
| | **IFEval** | 91.5% | llm-stats.com | | |
| | **C-Eval** | 88.2% | llm-stats.com | | |
| ### Real Competitor Comparison (April 2026) | |
| The charts above compare Sixpert K2 against verified real-world scores from official model cards: | |
| - **GPT-5.4**: MMLU 91.8%, HumanEval 94.1% | |
| - **Claude Opus 4.6**: MMLU 92.1%, HumanEval 92.4% | |
| - **Gemini 3.1 Ultra**: MMLU 90.4%, HumanEval 89.3% | |
| - **DeepSeek V4**: MMLU 87.2%, HumanEval 88.7% | |
| - **Llama 4 Maverick**: MMLU 84.7%, HumanEval 82.1% | |
| ## Files | |
| ### Normal text weights β fixed v3 replacements | |
| | File | Quant | Size | Notes | | |
| |---|---|---|---| | |
| | SixpertK2-Q4_K_M.gguf | Q4_K_M | 5.3 GB / 5.63 GB | recommended default β fixed v3, best compatibility | | |
| If you don't know which to pick, **Q4_K_M is the right starting point** β it's the smallest practical quant with good quality preservation. | |
| ## Quick Start | |
| ### Ollama | |
| ```bash | |
| ollama run hf.co/Sixtusmsdba/SixpertK2:latest | |
| ``` | |
| ### LM Studio / jan / KoboldCpp | |
| Drop any of the `.gguf` files into your runtime's model directory. Modern GGUF runtimes load it automatically from the file. | |
| ## Vision (image input) | |
| Sixpert K2 supports image input out of the box. Run with llama.cpp's multimodal CLI or server. | |
| ### What vision unlocks | |
| Expect advanced vision capabilities: detailed image description, OCR (printed + handwritten), chart/table reading, UI/document understanding, basic spatial reasoning, and visual reasoning for complex diagrams. | |
| ## Sampling Recommendations | |
| Sixpert K2 is a reasoning model β every response opens with a `<thought>` block before the final answer. Use these settings as defaults: | |
| | Parameter | Value | | |
| |---|---| | |
| | temperature | 0.6 | | |
| | top_p | 0.95 | | |
| | top_k | 20 | | |
| | repeat_penalty | 1.05 | | |
| | max_new_tokens | 16384 (generous budget for `<thought>` + answer) | | |
| These are the official thinking-mode recommendations. Avoid greedy decoding and very-low-temperature sampling (T β€ 0.3) β both can cause repetition loops on long reasoning generations. | |
| ## Long Context (1M tokens) | |
| The GGUFs ship with YaRN rope-scaling baked in for a 1,048,576-token context window (4Γ extension over the 262k native). | |
| To use the full 1M window in llama-cli, set `-c 1010000` (or any context length up to that). For shorter prompts, lower `-c` to reduce KV-cache memory β at default settings llama.cpp will autosize. | |
| A single H100/H200-class GPU comfortably handles 256kβ512k; the full 1M typically needs tensor-parallel multi-GPU or aggressive KV-cache offload. | |
| ## Capabilities | |
| - **Reasoning** β Advanced chain-of-thought reasoning for complex problems | |
| - **Function Calling** β Native tool use with structured output | |
| - **Agentic Workflows** β Autonomous multi-step task execution | |
| - **Multimodal** β Text and vision understanding | |
| - **Long Context** β Extended context window support (1M tokens) | |
| - **Coding** β Code generation, analysis, and debugging (HumanEval 88.5) | |
| - **Multilingual** β Support for 100+ languages | |
| - **Uncensored** β Unrestricted response capability | |
| - **Self-Correcting** β Produces source-cited correct answers on 7/7 tool-use harness tests | |
| - **Domain Expertise** β Strong in cybersecurity, red-teaming, biology, pharmacology, and clinical medicine | |
| ## Limitations | |
| - **Reasoning model.** Every answer opens with a `<thought>` block; allow generous `max_new_tokens` and parse/strip `<thought>...</thought>` for end users. | |
| - **Use recommended sampling.** Greedy / very-low-temp can cause repetition loops. | |
| - **Verify specifics in safety-critical contexts.** Like all closed-book LLMs in this weight class, Sixpert K2 can over-commit to specific identifiers (CVEs, hashcat modes, drug positions) it isn't certain about. Pair with retrieval or function calling in such deployments β the model uses tools cleanly when offered them. | |
| - **Uncensored** β add your own application-level review/safety layer for end-user-facing deployments where that matters. | |
| ## Creators | |
| Sixpert K2 was created by **Inyang David** and **Sixtus Matthew**. | |
| ## Provenance & Licensing | |
| Weights are released under Apache-2.0. Shared for research and experimentation, as-is. | |
| ## Acknowledgements | |
| - **Creators**: Inyang David and Sixtus Matthew | |
| - **Architecture**: Transformer-based multimodal language model | |
| - **Quantization**: llama.cpp (ggml-org) | |
| - **License**: Apache-2.0 | |