--- library_name: transformers license: apache-2.0 pipeline_tag: text-generation language: - multilingual - en - code tags: - moe - mixture-of-experts - reflexive-role-routing - code-generation - reasoning - qwen - qwen3_8 - qwen3.8 - llama.cpp - ollama - gguf --- # Moderato-V1-Pro (113.3B Sparse MoE)
| Benchmark & Capability | Moderato-V1-Pro (113.3B-A32.7B) |
Claude Sonnet 5 (Anthropic) |
GPT-5.6-Terra (OpenAI) |
Kimi K3 (2.8T-A104B) |
Qwen3.8-Flash-Next (180B) |
|---|---|---|---|---|---|
| Coding & Software Engineering | |||||
|
Agentic Terminal Execution
Terminal-Bench 2.1 (harborframework)
|
79.5 | 80.4 | 87.4 | 88.3 | 73.0 |
|
Multi-File Repository Refactoring
ScaleAI / SWE-bench Pro
|
63.3 | 63.2 | 63.4 | 42.0 | 62.5 |
|
Deep Autonomous Bug Fixing
datacurve / DeepSWE v1.1
|
53.2 | 54.0 | 64.0 | 67.3 | 58.7 |
| STEM & Advanced Scientific Reasoning | |||||
|
PhD-Level Scientific Reasoning
Idavidrein / GPQA Diamond
|
90.0 | 91.1 | 92.9 | 93.5 | 91.7 |
|
Extreme Frontier Reasoning (No Tools)
cais / HLE (Humanity's Last Exam)
|
38.4 | 48.0 | 50.4 | 43.5 | 35.9 |
| Autonomous Agents & Structured Extraction | |||||
|
Multi-Turn Agent Task Solving
internlm / WildClawBench (Overall)
|
52.2 | 59.9 | 50.4 | 54.5 | 48.0 |
|
Information Extraction & Schema
llamaindex / ExtractBench (Mean)
|
88.65 | 94.0 | 93.5 | 83.17 | 89.75 |