| --- |
| language: |
| - en |
| - zh |
| license: apache-2.0 |
| base_model: |
| - Qwen/Qwen3.5-397B-A17B |
| pipeline_tag: image-text-to-text |
| tags: |
| - agnes-ai |
| - reasoning |
| - multimodal |
| - long-context |
| - api |
| --- |
| |
| <p align="center"> |
| <img width="132" src="assets/agnes_logo.svg" alt="Agnes AI logo"> |
| </p> |
|
|
| <p align="center"> |
| <a href="https://agnes-ai.com/"><img src="https://img.shields.io/badge/Agnes_AI-Website-3248AF" alt="Agnes AI website"></a> |
| <a href="https://wiki.agnes-ai.com/en/docs/agnes-25-pro-alpha"><img src="https://img.shields.io/badge/Agnes_2.5_Pro_Alpha-API_Docs-3248AF" alt="Agnes API docs"></a> |
| <a href="https://artificialanalysis.ai/models/agnes-2-5-pro-alpha"><img src="https://img.shields.io/badge/Artificial_Analysis-Model_Page-111827" alt="Artificial Analysis model page"></a> |
| </p> |
|
|
| # Agnes 2.5 Pro Alpha |
|
|
| Hello! 👋 Today we are introducing **Agnes 2.5 Pro Alpha**, our most capable reasoning model for advanced coding, scientific problem solving, long-context analysis, multimodal understanding, and agentic workflows. |
|
|
| Highlights: |
|
|
| - **Competitive benchmark performance:** Agnes outperforms Qwen3.5-397B on six of the eight evaluations visualized below, with particularly clear gains on Terminal-Bench v2.1, CritPt, and AA-Omniscience Accuracy. |
| - **Built for demanding work:** a **1M-token context window**, up to **65,536 output tokens**, extended reasoning, tool calling, and text, image understanding. |
|
|
| <img style="width:100%;max-width:1100px" src="assets/agnes_benchmarks.svg" alt="LLM benchmark evaluation comparing Agnes 2.5 Pro Alpha with flagship-scale models" title="Agnes 2.5 Pro Alpha benchmark evaluation"> |
|
|
| ## Agnes 2.5 Pro Alpha |
|
|
| A multimodal reasoning model available through the Agnes AI API. The model combines long-context understanding with strong coding and scientific reasoning, while retaining the throughput and pricing needed for production workloads. |
|
|
| ### Benchmarks |
|
|
| Agnes 2.5 Pro Alpha is evaluated against the same comparison set selected for Ornith-1.0-397B: Qwen3.5-397B, Qwen3.7-Max, GLM-5.2-744B, MiniMax-M3-428B, DeepSeek-V4-Pro-1.6T, Claude Opus 4.7, and Claude Opus 4.8. Every result below is an independent Artificial Analysis benchmark measurement. |
|
|
| <div style="font-family:-apple-system,BlinkMacSystemFont,'Segoe UI',Roboto,sans-serif;width:100%;margin:0 auto;padding:16px 0;overflow-x:auto"> |
| <table style="width:100%;table-layout:fixed;border-collapse:collapse;font-size:11px;min-width:1100px"> |
| <thead><tr> |
| <th style="width:20%;padding:10px 6px;text-align:left;border-bottom:2px solid #3248AF;color:#3248AF">Benchmark</th> |
| <th style="width:10%;padding:10px 6px;text-align:center;font-weight:700;border-bottom:2px solid #3248AF;color:#3248AF;background:rgba(50,72,175,.09)">Agnes 2.5<br>Pro Alpha</th> |
| <th style="width:10%;padding:10px 6px;text-align:center;border-bottom:2px solid #3248AF">Qwen3.5<br>397B</th> |
| <th style="width:10%;padding:10px 6px;text-align:center;border-bottom:2px solid #3248AF">Qwen3.7<br>Max</th> |
| <th style="width:10%;padding:10px 6px;text-align:center;border-bottom:2px solid #3248AF">GLM-5.2<br>744B</th> |
| <th style="width:10%;padding:10px 6px;text-align:center;border-bottom:2px solid #3248AF">MiniMax-M3<br>428B</th> |
| <th style="width:10%;padding:10px 6px;text-align:center;border-bottom:2px solid #3248AF">DeepSeek-V4-Pro<br>1.6T</th> |
| <th style="width:10%;padding:10px 6px;text-align:center;border-bottom:2px solid #3248AF">Claude Opus<br>4.7</th> |
| <th style="width:10%;padding:10px 6px;text-align:center;border-bottom:2px solid #3248AF">Claude Opus<br>4.8</th> |
| </tr></thead> |
| <tbody> |
| <tr><td colspan="9" style="padding:8px 12px;font-weight:600;color:#3248AF;background:rgba(50,72,175,.07)">Agentic Work & Coding</td></tr> |
| <tr><td style="padding:7px">GDPval-AA v2 <sup>†</sup></td><td style="padding:7px;text-align:center;font-weight:700;color:#3248AF;background:rgba(50,72,175,.05)">33.8</td><td style="padding:7px;text-align:center">23.2</td><td style="padding:7px;text-align:center">38.6</td><td style="padding:7px;text-align:center">50.3</td><td style="padding:7px;text-align:center">44.3</td><td style="padding:7px;text-align:center">54.5</td><td style="padding:7px;text-align:center">49.5</td><td style="padding:7px;text-align:center">54.2</td></tr> |
| <tr><td style="padding:7px">τ³-Banking</td><td style="padding:7px;text-align:center;font-weight:700;color:#3248AF;background:rgba(50,72,175,.05)">12.4</td><td style="padding:7px;text-align:center">13.4</td><td style="padding:7px;text-align:center">11.8</td><td style="padding:7px;text-align:center">34.6</td><td style="padding:7px;text-align:center">15.3</td><td style="padding:7px;text-align:center">39.6</td><td style="padding:7px;text-align:center">34.6</td><td style="padding:7px;text-align:center">34.2</td></tr> |
| <tr><td style="padding:7px">Terminal-Bench v2.1</td><td style="padding:7px;text-align:center;font-weight:700;color:#3248AF;background:rgba(50,72,175,.05)">67.0</td><td style="padding:7px;text-align:center">51.3</td><td style="padding:7px;text-align:center">74.5</td><td style="padding:7px;text-align:center">77.9</td><td style="padding:7px;text-align:center">65.2</td><td style="padding:7px;text-align:center">78.7</td><td style="padding:7px;text-align:center">83.1</td><td style="padding:7px;text-align:center">84.6</td></tr> |
| <tr><td style="padding:7px">SciCode</td><td style="padding:7px;text-align:center;font-weight:700;color:#3248AF;background:rgba(50,72,175,.05)">42.2</td><td style="padding:7px;text-align:center">42.0</td><td style="padding:7px;text-align:center">48.8</td><td style="padding:7px;text-align:center">50.5</td><td style="padding:7px;text-align:center">45.4</td><td style="padding:7px;text-align:center">49.2</td><td style="padding:7px;text-align:center">54.5</td><td style="padding:7px;text-align:center">53.5</td></tr> |
| <tr><td colspan="9" style="padding:8px 12px;font-weight:600;color:#3248AF;background:rgba(50,72,175,.07)">Long Context & Scientific Reasoning</td></tr> |
| <tr><td style="padding:7px">AA-LCR</td><td style="padding:7px;text-align:center;font-weight:700;color:#3248AF;background:rgba(50,72,175,.05)">73.0</td><td style="padding:7px;text-align:center">72.7</td><td style="padding:7px;text-align:center">74.7</td><td style="padding:7px;text-align:center">76.7</td><td style="padding:7px;text-align:center">80.3</td><td style="padding:7px;text-align:center">75.3</td><td style="padding:7px;text-align:center">75.3</td><td style="padding:7px;text-align:center">73.0</td></tr> |
| <tr><td style="padding:7px">Humanity's Last Exam</td><td style="padding:7px;text-align:center;font-weight:700;color:#3248AF;background:rgba(50,72,175,.05)">33.6</td><td style="padding:7px;text-align:center">29.0</td><td style="padding:7px;text-align:center">40.5</td><td style="padding:7px;text-align:center">41.1</td><td style="padding:7px;text-align:center">39.0</td><td style="padding:7px;text-align:center">41.0</td><td style="padding:7px;text-align:center">42.3</td><td style="padding:7px;text-align:center">48.7</td></tr> |
| <tr><td style="padding:7px">GPQA Diamond</td><td style="padding:7px;text-align:center;font-weight:700;color:#3248AF;background:rgba(50,72,175,.05)">87.6</td><td style="padding:7px;text-align:center">89.3</td><td style="padding:7px;text-align:center">92.3</td><td style="padding:7px;text-align:center">89.5</td><td style="padding:7px;text-align:center">92.9</td><td style="padding:7px;text-align:center">92.8</td><td style="padding:7px;text-align:center">91.4</td><td style="padding:7px;text-align:center">92.0</td></tr> |
| <tr><td style="padding:7px">CritPt</td><td style="padding:7px;text-align:center;font-weight:700;color:#3248AF;background:rgba(50,72,175,.05)">10.9</td><td style="padding:7px;text-align:center">1.7</td><td style="padding:7px;text-align:center">13.4</td><td style="padding:7px;text-align:center">20.9</td><td style="padding:7px;text-align:center">3.7</td><td style="padding:7px;text-align:center">18.0</td><td style="padding:7px;text-align:center">12.0</td><td style="padding:7px;text-align:center">20.9</td></tr> |
| <tr><td colspan="9" style="padding:8px 12px;font-weight:600;color:#3248AF;background:rgba(50,72,175,.07)">Knowledge Reliability</td></tr> |
| <tr><td style="padding:7px">AA-Omniscience Accuracy</td><td style="padding:7px;text-align:center;font-weight:700;color:#3248AF;background:rgba(50,72,175,.05)">33.5</td><td style="padding:7px;text-align:center">30.8</td><td style="padding:7px;text-align:center">31.1</td><td style="padding:7px;text-align:center">24.3</td><td style="padding:7px;text-align:center">16.7</td><td style="padding:7px;text-align:center">49.1</td><td style="padding:7px;text-align:center">48.9</td><td style="padding:7px;text-align:center">48.8</td></tr> |
| <tr><td style="padding:7px">Non-Hallucination Rate</td><td style="padding:7px;text-align:center;font-weight:700;color:#3248AF;background:rgba(50,72,175,.05)">11.9</td><td style="padding:7px;text-align:center">11.1</td><td style="padding:7px;text-align:center">74.4</td><td style="padding:7px;text-align:center">73.7</td><td style="padding:7px;text-align:center">81.6</td><td style="padding:7px;text-align:center">5.2</td><td style="padding:7px;text-align:center">57.7</td><td style="padding:7px;text-align:center">60.7</td></tr> |
| </tbody> |
| </table> |
| </div> |
|
|
| <p style="font-size:11px;opacity:.72"> |
| † GDPval-AA v2 uses Artificial Analysis' normalized score, <code>(Elo − 500) / 2000</code>. Non-Hallucination Rate is <code>1 − hallucination rate</code>. Higher is better for every benchmark. Snapshot checked August 18, 2026; values may change as evaluations are updated. |
| </p> |
|
|
| ## Model Information |
|
|
| | Property | Value | |
| |---|---| |
| | Developed by | Agnes AI | |
| | Model name | Agnes 2.5 Pro Alpha | |
| | Model type | Multimodal reasoning model | |
| | License | Apache License 2.0 | |
| | Languages | English, Chinese | |
| | Context window | 1,048,576 tokens | |
| | Maximum output | 65,536 tokens | |
| | Input modalities | Text, image | |
| | Output modality | Text | |
| | Precision | BF16 | |
| | Tool calling | Yes | |
| | Streaming | Yes | |
| | Release date | July 2026 | |
|
|
| ## License |
|
|
| This repository is licensed under the [Apache License 2.0](LICENSE). |
|
|
| Agnes 2.5 Pro Alpha is a post-trained derivative of [Qwen/Qwen3.5-397B-A17B](https://huggingface.co/Qwen/Qwen3.5-397B-A17B), which is also licensed under Apache License 2.0. Original copyright notices are retained. Additional post-training was performed by Agnes AI. |
|
|
| See the `LICENSE` file in this repository for the full terms. |
|
|
| ## Hardware Requirements |
|
|
| The Quickstart launch command uses 8-GPU tensor parallelism. The checkpoint is a large multi-shard BF16 package; a single GPU is not sufficient. |
|
|
| | Resource | Recommendation | |
| |---|---| |
| | GPUs | 8× NVIDIA H200 (141 GB) or equivalent | |
| | Tensor parallel | `--tp 8` | |
| | Host memory / disk | Fast NVMe with about **1 TB** free for weights, tokenizer files, and download cache | |
| | Context length | The sample command sets `--context-length 1024000`. If you hit out-of-memory errors, lower this value | |
| | Network | Optional. The same model is also served at `https://apihub.agnes-ai.com/v1` without local GPUs | |
|
|
| ## Quickstart |
|
|
| <div style="border-left:4px solid #3248AF;background:rgba(50,72,175,.08);border-radius:6px;padding:12px 16px;font-family:-apple-system,BlinkMacSystemFont,'Segoe UI',Roboto,sans-serif;font-size:14px;line-height:1.6"> |
| <div style="font-weight:700;color:#3248AF;margin-bottom:6px">REASONING MODEL</div> |
| <p style="margin:0"><b>Agnes 2.5 Pro Alpha</b> uses extended reasoning for complex tasks and is available through OpenAI-compatible Chat Completions and Responses APIs. Keep API keys in environment variables and use publicly accessible URLs for image inputs.</p> |
| </div> |
|
|
| <h3>SGLang</h3> |
|
|
| ```bash |
| python -m sglang.launch_server \ |
| --model-path Agnes-AI/Agnes-2.5-Pro-Alpha \ |
| --served-model-name agnes-2.5-pro-alpha \ |
| --tp 8 \ |
| --host 0.0.0.0 --port 8000 \ |
| --context-length 1024000 \ |
| --mem-fraction-static 0.85 \ |
| --tool-call-parser qwen3_coder \ |
| --reasoning-parser qwen3 |
| ``` |
|
|
| ### Chat Completions |
|
|
| ```bash |
| export AGNES_API_KEY="your-api-key" |
| |
| curl https://apihub.agnes-ai.com/v1/chat/completions \ |
| -H "Authorization: Bearer ${AGNES_API_KEY}" \ |
| -H "Content-Type: application/json" \ |
| -d '{ |
| "model": "agnes-2.5-pro-alpha", |
| "messages": [ |
| { |
| "role": "user", |
| "content": "Review this API handler for security issues and provide a corrected version." |
| } |
| ], |
| "temperature": 1.0, |
| "max_tokens": 2000 |
| }' |
| ``` |
|
|
| ### Python |
|
|
| ```bash |
| pip install openai |
| ``` |
|
|
| ```python |
| import os |
| from openai import OpenAI |
| |
| client = OpenAI( |
| api_key=os.environ["AGNES_API_KEY"], |
| base_url="https://apihub.agnes-ai.com/v1", |
| ) |
| |
| response = client.chat.completions.create( |
| model="agnes-2.5-pro-alpha", |
| messages=[ |
| { |
| "role": "user", |
| "content": "Design a fault-tolerant event processing architecture.", |
| } |
| ], |
| temperature=1.0, |
| max_tokens=2000, |
| ) |
| |
| print(response.choices[0].message.content) |
| ``` |
|
|
| ### Image Understanding |
|
|
| ```python |
| response = client.chat.completions.create( |
| model="agnes-2.5-pro-alpha", |
| messages=[ |
| { |
| "role": "user", |
| "content": [ |
| {"type": "text", "text": "Explain this chart and call out anomalies."}, |
| { |
| "type": "image_url", |
| "image_url": {"url": "https://example.com/chart.png"}, |
| }, |
| ], |
| } |
| ], |
| ) |
| ``` |
|
|
| ### Responses API |
|
|
| ```bash |
| curl https://apihub.agnes-ai.com/v1/responses \ |
| -H "Authorization: Bearer ${AGNES_API_KEY}" \ |
| -H "Content-Type: application/json" \ |
| -d '{ |
| "model": "agnes-2.5-pro-alpha", |
| "input": "Create a step-by-step migration plan from a monolith to services.", |
| "max_output_tokens": 2000 |
| }' |
| ``` |
|
|
| ## Recommended Inference Settings |
|
|
| Use sampling rather than greedy decoding. Leave enough `max_tokens` / `max_output_tokens` for extended reasoning. |
|
|
| | Setting | Recommended | |
| |---|---| |
| | `temperature` | 1.0 | |
| | `top_p` | 0.95 | |
| | `top_k` | 20 | |
| | `repetition_penalty` | 1.05 | |
| | `max_tokens` | 2000 or higher | |
|
|
| Raise `max_tokens` if a response stops early. |
|
|
| ## Model Capabilities |
|
|
| | Capability | Support | |
| |---|---| |
| | Advanced reasoning | Yes | |
| | Coding and debugging | Yes | |
| | Long-context analysis | 1M tokens | |
| | Maximum output | 65,536 tokens | |
| | Image understanding | Yes, via public image URL | |
| | Tool calling | Yes | |
| | Streaming | Yes | |
| | OpenAI-compatible APIs | Chat Completions and Responses | |
|
|
| Agnes 2.5 Pro Alpha is especially well suited to repository-level coding, technical research, document synthesis, visual analysis, and tool-enabled agents that need to reason across long and complex contexts. |
|
|
| ## Responsible Use |
|
|
| Model outputs can contain errors. Validate high-impact decisions and tool actions in the application layer, and review the applicable Agnes AI service terms before sending sensitive or regulated data. |
|
|
| ## Citation |
|
|
| ```bibtex |
| @misc{agnes25proalpha2026, |
| title = {Agnes 2.5 Pro Alpha}, |
| author = {{Agnes AI}}, |
| year = {2026}, |
| month = jul, |
| howpublished = {API model}, |
| url = {https://wiki.agnes-ai.com/en/docs/agnes-25-pro-alpha} |
| } |
| ``` |
|
|