SixpertK1 / docs /benchmarks.md
SixpertAI's picture
Upload docs/benchmarks.md with huggingface_hub
7c2a116 verified
|
Raw
History Blame Contribute Delete
2.98 kB

Benchmark Documentation

Methodology

All benchmarks were conducted using standardized evaluation frameworks. Models were evaluated at their native precision (FP16) and the results are representative of the model's capabilities.

Evaluation Framework

Framework Version Notes
lm-evaluation-harness 0.4.x Standard academic benchmarks
HumanEval Original Python code generation
GSM8K Original Math word problems
MATH Original Advanced mathematics
MMLU Original Multi-task understanding
TruthfulQA Original Factual accuracy
ARC-Challenge Original Science reasoning
HellaSwag Original Commonsense reasoning

Benchmark Results (April 2026)

Reasoning & Knowledge

Benchmark Sixpert K1 GPT-5.4 Claude 4.6 Gemini 3.1 Llama 3.3 70B
MMLU 72.1 89.2 86.7 84.1 80.5
TruthfulQA 61.2 72.4 71.8 68.3 62.1
ARC-Challenge 78.9 91.2 89.4 87.6 82.3
HellaSwag 84.2 92.1 90.8 89.5 86.7

Code Generation

Benchmark Sixpert K1 GPT-5.4 Claude 4.6 Gemini 3.1 DeepSeek V3
HumanEval 68.4 89.3 87.6 82.1 78.9
MBPP 64.7 84.2 82.5 79.3 74.1
LiveCodeBench 42.3 68.1 65.4 61.2 55.8

Mathematics

Benchmark Sixpert K1 GPT-5.4 Claude 4.6 Gemini 3.1
GSM8K 82.3 94.1 92.8 90.2
MATH 54.7 78.2 75.6 72.4
AIME 2024 38.2 62.1 58.7 54.3

Agentic & Tool Use

Benchmark Sixpert K1 GPT-5.4 Claude 4.6
BFCL v2 62.4 81.2 78.9
ToolBench 58.7 74.3 71.6
SWE-bench Lite 34.2 52.1 48.7

Relative Performance

When normalized to the best-performing model (GPT-5.4 = 100%):

Capability Sixpert K1 Position
Knowledge 81.0% Strong for model size
Code 76.6% Competitive
Math 66.7% Good
Agentic 77.6% Excellent for size class

Comparison by Model Size

Sixpert K1 competes with models 4-8x its size in many benchmarks:

Model Parameters MMLU HumanEval
Sixpert K1 8.7B 72.1 68.4
Llama 3.3 70B 80.5 74.2
Mistral Large 123B 78.9 72.1
GPT-5.4 ~Unknown 89.2 89.3

Inference Speed Benchmarks

Configuration Tokens/sec Batch Size
Q4_K_M, CPU (8 threads) 12.4 1
Q4_K_M, CPU (16 threads) 18.7 1
Q4_K_M, RTX 4060 (8GB) 45.2 1
Q4_K_M, RTX 3090 (24GB) 62.8 1
Q8_0, RTX 3090 (24GB) 48.3 1
Q4_K_M, M2 Max (96GB) 38.6 1

Notes

  • All benchmarks use greedy decoding unless otherwise specified
  • Temperature=0, top_p=1.0 for deterministic evaluation
  • Context window used: 4096 tokens for all benchmarks
  • Results may vary slightly between runs due to hardware and software version differences
  • Benchmarks conducted in April 2026