1. Introduction
We're introducing GRM-3.2-Sky, our latest flagship model built for long-horizon agentic tasks and extremely difficult reasoning problems. GRM-3.2-Sky marks a substantial leap in long-horizon task capability over its predecessor, GRM-2.6-Plus, and is designed to serve as a dependable engine for complex, multi-step workflows.
The model is purpose-built for long-horizon agentic tasks and problems that are simply hard — difficult coding challenges, advanced mathematics, and rigorous logical reasoning. GRM-3.2-Sky aims to sustain coherent, goal-directed behavior over extended interactions, making it well suited for users who need a model that doesn't lose the thread across many steps of tool use, planning, and self-correction.
2. Key Capabilities
- Long-Horizon Agentic Mastery: GRM-3.2-Sky is specifically optimized to maintain coherence, planning quality, and task fidelity across long, multi-step agentic workflows, a significant step up from GRM-2.6-Plus.
- Elite Reasoning on Hard Problems: Strong performance on difficult coding, advanced mathematics, and logical reasoning tasks, with careful, structured step-by-step problem-solving.
- Robust Coding Ability: Handles complex, difficult codebases and multi-file coding tasks, including debugging, refactoring, and long-running terminal/agentic coding sessions.
- Consistent Logical Reasoning: Built to reason carefully through multi-constraint logic problems without losing track of intermediate steps.
- Flagship-Class Performance: Positioned as the top of the GRM lineup, intended to compete head-to-head with frontier-scale models on the hardest tasks.
3. Performance
GRM-3.2-Sky is designed as our most capable model to date for long-horizon agentic work and difficult reasoning. It builds directly on the strengths of GRM-2.6-Plus while specifically targeting the failure modes that emerge over long task horizons — drift, inconsistency, and loss of goal state — resulting in meaningfully improved reliability across extended sessions.
Its core strength is sustained intelligence over time: elite-level reasoning, resilient long-horizon planning, and the ability to stay on task through difficult, multi-step coding, math, and logic problems.
Detailed Benchmarks
| GRM-3.2-Sky | GRM-2.6-Plus-0628 | GPT-5.6-Luna | Sonnet 5 | Gemini 3 Pro | |
|---|---|---|---|---|---|
| Knowledge & STEM | |||||
| MMLU-Pro | 89.5 | 86.8 | — | — | 89.8 |
| MMLU-Redux | 96.9 | 94.2 | — | — | — |
| GPQA Diamond | 90.6 | 88.3 | 92.3 | — | 91.9 |
| Reasoning & Coding | |||||
| LiveCodeBench v6 | 87.7 | 84.8 | — | — | 82.9 |
| HMMT Feb 26 | 86.4 | 84.8 | — | — | — |
| AIME26 | 96.3 | 95.1 | — | — | — |
| General Agent | |||||
| SWE-bench Verified | 81.4 | 77.7 | — | 85.2 | 76.2 |
| SWE-bench Pro | 58.3 | 54.0 | 62.7 | 63.2 | — |
| Terminal-Bench 2.1 | 66.3 | — | 84.7 | 80.4 | — |
| NL2Repo | 35.6 | — | — | — | — |
Scores are taken from each provider's own published model card, blog post, or system card where available; "—" indicates a score was not publicly reported by that provider at the time of writing. Different labs may use different agent scaffolds when reporting SWE-bench and Terminal-Bench results, so cross-provider comparisons should be read with that caveat.
4. Family
The GRM-3.2 family is available in various sizes to suit every use case.
| Model | Size | Domain |
|---|---|---|
| GRM-3.2-Sky | 35B-A3B | Flagship model for long-horizon tasks |
| GRM-3.2-Cliff | 9B | Capable model for low GPU environments |
| GRM-3.2-Turf | 1.2B | Lightweight model for practical reasoning |
5. Architecture
GRM-3.2-Sky is built on the Ornith-1.0-35B architecture, a 35B-parameter Mixture-of-Experts model with ~3B active parameters (35B-A3B), optimized for long-horizon agentic workflows, difficult coding, advanced mathematics, and rigorous logical reasoning, while remaining efficient to deploy thanks to its sparse activation.
GRM-3.2-Sky is developed by OrionLLM and released under the Apache 2.0 License.
- Downloads last month
- 94
Model tree for OrionLLM/GRM-3.2-Sky
Space using OrionLLM/GRM-3.2-Sky 1
Collection including OrionLLM/GRM-3.2-Sky
Evaluation results
- SWE-bench/SWE-bench_Verified · Swe Bench Resolved View evaluation results leaderboard 81.4
- ScaleAI/SWE-bench_Pro · SWE Bench Pro View evaluation results leaderboard 58.3
- Idavidrein/gpqa · Diamond View evaluation results leaderboard 90.6
- TIGER-Lab/MMLU-Pro · Mmlu Pro View evaluation results leaderboard 89.5
- MathArena/hmmt_feb_2026 · MathArena Hmmt Feb 2026 View evaluation results leaderboard 86.4
