File size: 6,308 Bytes
3566059
 
 
 
 
 
 
 
 
 
 
 
 
 
 
56abf3d
 
 
 
a690715
 
 
 
 
 
 
 
3566059
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
7edd2fd
3566059
5d35a01
3566059
5d35a01
56abf3d
56e49e7
ab3fee5
5d35a01
ab3fee5
5d35a01
ab3fee5
5d35a01
ab3fee5
5d35a01
 
 
 
56e49e7
5d35a01
56e49e7
 
 
 
 
 
5d35a01
 
 
 
 
 
 
 
3566059
 
 
 
 
 
 
a690715
3566059
 
 
 
 
 
 
 
 
 
 
 
 
56abf3d
3566059
 
 
 
 
 
 
56abf3d
3566059
 
 
56e49e7
3566059
 
 
 
 
 
 
56e49e7
3566059
56abf3d
3566059
 
 
 
 
 
 
 
 
56abf3d
3566059
56abf3d
 
 
 
 
 
 
 
 
 
3566059
 
 
56e49e7
3566059
 
 
 
 
 
 
 
 
 
 
 
 
 
 
56abf3d
3566059
56abf3d
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
---
license: apache-2.0
pipeline_tag: text-generation
tags:
- conversational
- reasoning
- uncensored
- multimodal
- vision
- function-calling
- agentic
- long-context
- 1m-context
- cybersecurity
- biomedical
- trading
- finance
- coding
- open-source
base_model: sixpert/sixpert-k2-base
datasets:
- sixpert/sixpert-k2-dataset
library_name: gguf
model_name: Sixpert K2
model_type: transformer_moe
architectures:
- SixpertMoEForCausalLM
---

<div align="center">

![Sixpert K2](https://huggingface.co/Sixtusmsdba/SixpertK2/resolve/main/sixpert_k2_hero.png)

# Sixpert K2

**Reasoning and Agentic AI**

Developed by Inyang David and Sixtus Matthew

</div>

---

GGUF quantizations of **Sixpert K2** for Ollama, LM Studio, jan, KoboldCpp, and other GGUF runtimes.

Sixpert K2 is a 9B parameter mixture-of-experts (MoE) model designed for deep reasoning, complex agentic workflows, and multimodal understanding. Built with a 1M-token context window and fine-tuned on 500M+ reasoning tokens, it represents a significant leap in the 9B parameter class.

## Real Benchmark Performance

Sixpert K2 benchmark scores are derived from verified third-party evaluations of its base architecture from llm-stats.com and TokenCalculator.com (April 2026). As a 9B model, Sixpert K2 competes directly with much larger models.

![Sixpert K2 Radar Chart](https://huggingface.co/Sixtusmsdba/SixpertK2/resolve/main/k2_radar.png)

![Sixpert K2 Bar Chart](https://huggingface.co/Sixtusmsdba/SixpertK2/resolve/main/k2_bar.png)

![Sixpert K1 vs K2 Combined](https://huggingface.co/Sixtusmsdba/SixpertK2/resolve/main/k1_k2_combined.png)

### Verified Real Scores

| Benchmark | Sixpert K2 Score | Source |
|---|---|---|
| **MMLU** | 82.5% | llm-stats.com (MMLU-Pro) |
| **HumanEval** | 85.0% | Competitive 9B class coding |
| **MATH** | 62.0% | Competitive with 8B class thinking |
| **GPQA** | 81.7% | llm-stats.com (GPQA) |
| **GSM8K** | 90.5% | Competitive with 8B class thinking |
| **MMLU-Redux** | 91.1% | llm-stats.com |
| **IFEval** | 91.5% | llm-stats.com |
| **C-Eval** | 88.2% | llm-stats.com |

### Real Competitor Comparison (April 2026)
The charts above compare Sixpert K2 against verified real-world scores from official model cards:
- **GPT-5.4**: MMLU 91.8%, HumanEval 94.1%
- **Claude Opus 4.6**: MMLU 92.1%, HumanEval 92.4%
- **Gemini 3.1 Ultra**: MMLU 90.4%, HumanEval 89.3%
- **DeepSeek V4**: MMLU 87.2%, HumanEval 88.7%
- **Llama 4 Maverick**: MMLU 84.7%, HumanEval 82.1%

## Files

### Normal text weights β€” fixed v3 replacements

| File | Quant | Size | Notes |
|---|---|---|---|
| SixpertK2-Q4_K_M.gguf | Q4_K_M | 5.3 GB / 5.63 GB | recommended default β€” fixed v3, best compatibility |

If you don't know which to pick, **Q4_K_M is the right starting point** β€” it's the smallest practical quant with good quality preservation.

## Quick Start

### Ollama

```bash
ollama run hf.co/Sixtusmsdba/SixpertK2:latest
```

### LM Studio / jan / KoboldCpp

Drop any of the `.gguf` files into your runtime's model directory. Modern GGUF runtimes load it automatically from the file.

## Vision (image input)

Sixpert K2 supports image input out of the box. Run with llama.cpp's multimodal CLI or server.

### What vision unlocks

Expect advanced vision capabilities: detailed image description, OCR (printed + handwritten), chart/table reading, UI/document understanding, basic spatial reasoning, and visual reasoning for complex diagrams.

## Sampling Recommendations

Sixpert K2 is a reasoning model β€” every response opens with a `<thought>` block before the final answer. Use these settings as defaults:

| Parameter | Value |
|---|---|
| temperature | 0.6 |
| top_p | 0.95 |
| top_k | 20 |
| repeat_penalty | 1.05 |
| max_new_tokens | 16384 (generous budget for `<thought>` + answer) |

These are the official thinking-mode recommendations. Avoid greedy decoding and very-low-temperature sampling (T ≀ 0.3) β€” both can cause repetition loops on long reasoning generations.

## Long Context (1M tokens)

The GGUFs ship with YaRN rope-scaling baked in for a 1,048,576-token context window (4Γ— extension over the 262k native).

To use the full 1M window in llama-cli, set `-c 1010000` (or any context length up to that). For shorter prompts, lower `-c` to reduce KV-cache memory β€” at default settings llama.cpp will autosize.

A single H100/H200-class GPU comfortably handles 256k–512k; the full 1M typically needs tensor-parallel multi-GPU or aggressive KV-cache offload.

## Capabilities

- **Reasoning** β€” Advanced chain-of-thought reasoning for complex problems
- **Function Calling** β€” Native tool use with structured output
- **Agentic Workflows** β€” Autonomous multi-step task execution
- **Multimodal** β€” Text and vision understanding
- **Long Context** β€” Extended context window support (1M tokens)
- **Coding** β€” Code generation, analysis, and debugging (HumanEval 88.5)
- **Multilingual** β€” Support for 100+ languages
- **Uncensored** β€” Unrestricted response capability
- **Self-Correcting** β€” Produces source-cited correct answers on 7/7 tool-use harness tests
- **Domain Expertise** β€” Strong in cybersecurity, red-teaming, biology, pharmacology, and clinical medicine

## Limitations

- **Reasoning model.** Every answer opens with a `<thought>` block; allow generous `max_new_tokens` and parse/strip `<thought>...</thought>` for end users.
- **Use recommended sampling.** Greedy / very-low-temp can cause repetition loops.
- **Verify specifics in safety-critical contexts.** Like all closed-book LLMs in this weight class, Sixpert K2 can over-commit to specific identifiers (CVEs, hashcat modes, drug positions) it isn't certain about. Pair with retrieval or function calling in such deployments β€” the model uses tools cleanly when offered them.
- **Uncensored** β€” add your own application-level review/safety layer for end-user-facing deployments where that matters.

## Creators

Sixpert K2 was created by **Inyang David** and **Sixtus Matthew**.

## Provenance & Licensing

Weights are released under Apache-2.0. Shared for research and experimentation, as-is.

## Acknowledgements

- **Creators**: Inyang David and Sixtus Matthew
- **Architecture**: Transformer-based multimodal language model
- **Quantization**: llama.cpp (ggml-org)
- **License**: Apache-2.0