Fabric1.6 / README.md
tudor-iustin22's picture
Upload README.md
88502e1 verified
|
Raw
History Blame Contribute Delete
15.8 kB
---
license: apache-2.0
base_model: Qwen/Qwen3.5-35B-A3B-Base
library_name: transformers
pipeline_tag: image-text-to-text
datasets:
- open-thoughts/OpenThoughts3-1.2M
- open-r1/OpenR1-Math-220k
- HuggingFaceTB/smoltalk2
- NousResearch/hermes-function-calling-v1
- OpenAssistant/oasst2
- FabricAI/maple
tags:
- fabric1.6
- mixture-of-experts
- multimodal
- image-text-to-text
- video-text-to-text
- vision
- video
- reasoning
- agentic
- tool-calling
- long-context
- hybrid-attention
- transformers
- pytorch
- safetensors
- bf16
model-index:
- name: Fabric 1.6
results:
- task:
type: text-generation
name: Math and Reasoning
dataset:
type: aime-2025
name: AIME25
metrics:
- type: pass@1
value: 92.8
name: Pass@1
- task:
type: text-generation
name: Math and Reasoning
dataset:
type: aime-2026
name: AIME26
metrics:
- type: pass@1
value: 93.1
name: Pass@1
- task:
type: text-generation
name: Math and Reasoning
dataset:
type: hmmt-feb-2026
name: HMMT26
metrics:
- type: pass@1
value: 83.2
name: Pass@1
- task:
type: text-generation
name: Math and Reasoning
dataset:
type: imo-answerbench
name: IMOAB
metrics:
- type: pass@1
value: 79.2
name: Pass@1
- task:
type: text-generation
name: Math and Reasoning
dataset:
type: math-500
name: M500
metrics:
- type: accuracy
value: 84.8
name: Accuracy
- task:
type: text-generation
name: Science and Knowledge
dataset:
type: gpqa
name: GPQA
metrics:
- type: accuracy
value: 86.7
name: Accuracy
- task:
type: text-generation
name: Science and Knowledge
dataset:
type: gpqa-diamond
name: GPQA-D
metrics:
- type: accuracy
value: 84.9
name: Accuracy
- task:
type: text-generation
name: Science and Knowledge
dataset:
type: hle
name: HLE
metrics:
- type: accuracy
value: 21.4
name: Accuracy
- task:
type: text-generation
name: Science and Knowledge
dataset:
type: mmlu-pro
name: MMLU-P
metrics:
- type: accuracy
value: 85.6
name: Accuracy
- task:
type: text-generation
name: Science and Knowledge
dataset:
type: mmlu-redux
name: MMLU-R
metrics:
- type: accuracy
value: 93.5
name: Accuracy
- task:
type: text-generation
name: Science and Knowledge
dataset:
type: ceval
name: C-Eval
metrics:
- type: accuracy
value: 92.3
name: Accuracy
- task:
type: text-generation
name: Code Generation
dataset:
type: livecodebench-v6
name: LCB6
metrics:
- type: pass@1
value: 80.2
name: Pass@1
- task:
type: text-generation
name: Code Generation
dataset:
type: swe-bench-verified
name: SWEB-V
metrics:
- type: resolve-rate
value: 72.9
name: Resolve Rate
- task:
type: text-generation
name: Code Generation
dataset:
type: swe-bench-pro
name: SWEB-P
metrics:
- type: resolve-rate
value: 50.1
name: Resolve Rate
- task:
type: text-generation
name: Instruction Following
dataset:
type: ifeval
name: IFEval
metrics:
- type: instruction-level
value: 93.09
name: Instruction Level
- task:
type: text-generation
name: General Reasoning
dataset:
type: gsm8k-platinum
name: GSM8K-Pt
metrics:
- type: accuracy
value: 95.73
name: Accuracy
- task:
type: text-generation
name: Agentic Tool Use
dataset:
type: tau3-bench
name: TAU3
metrics:
- type: pass-rate
value: 67.2
name: Pass Rate
- task:
type: image-text-to-text
name: Visual Question Answering
dataset:
type: mmmu-pro
name: MMMU-P
metrics:
- type: accuracy
value: 74.10
name: Accuracy
- task:
type: image-text-to-text
name: Visual Question Answering
dataset:
type: realworldqa
name: RWQA
metrics:
- type: accuracy
value: 85.4
name: Accuracy
- task:
type: text-generation
name: Agentic Tool Use
dataset:
type: mcp-atlas
name: MCP-A
metrics:
- type: completion
value: 62.8
name: Completion
- task:
type: text-generation
name: Agentic Tool Use
dataset:
type: widesearch
name: WS
metrics:
- type: rubric-score
value: 60.3
name: Rubric Score
- task:
type: image-text-to-text
name: Visual Question Answering
dataset:
type: mathvista-mini
name: MV-mini
metrics:
- type: accuracy
value: 86.6
name: Accuracy
---
<div align="center">
<picture>
<img src="banner_fabric1.6.png" width="100%" alt="Fabric AI">
</picture>
</div>
<hr>
<div align="center" style="line-height:1">
<a href="https://huggingface.co/FabricAI" target="_blank"><img alt="Hugging Face" src="https://img.shields.io/badge/%F0%9F%A4%97%20Hugging%20Face-FabricAI-ffc107?color=e0a800&logoColor=white"/></a>
<a href="https://fabricai.co.uk" target="_blank"><img alt="Homepage" src="https://img.shields.io/badge/Homepage-Fabric%20AI-white?logo=globe&logoColor=white"/></a>
<a href="https://x.com/fabricai_uk" target="_blank"><img alt="X" src="https://img.shields.io/badge/X-%40fabricai_uk-white?logo=x&logoColor=black"/></a>
</div>
# Fabric 1.6
**Fabric 1.6** is a 35-billion-parameter Mixture-of-Experts (MoE) reasoning model developed by **Fabric AI**, with approximately **3 billion parameters activated per token**. It is a native multimodal, agentic model built on a hybrid **Gated DeltaNet + Gated Attention** architecture, with explicit chain-of-thought reasoning, a native 262,144-token context window, and built-in Multi-Token Prediction (MTP) for up to 50% faster generation.
Fabric 1.6 is designed for agentic use in harnesses such as **OpenCode**, **Pi Agent**, **Hermes Agent** and other OpenAI-compatible tool-calling environments, and offers the option to preserve thinking context from past messages across long multi-turn sessions.
## 1. Key Features
- **Hybrid Architecture**: Gated DeltaNet (linear attention) layers interleaved with Gated Attention layers inside a 256-expert MoE transformer — sub-quadratic scaling with full attention capacity where it matters.
- **Native Long Context**: 262,144 tokens natively, extensible up to **1,010,000 tokens**.
- **Multi-Token Prediction (MTP)**: predicts multiple future tokens per step for up to **50% faster generation**.
- **Native Multimodality**: accepts text, image and video inputs within the same model.
- **Explicit Reasoning**: produces an internal chain of thought before answering; reasoning is exposed in a structured format that can be streamed and stored.
- **Agentic by Design**: reliable structured tool-calling, long-horizon task execution, and preserved thinking context across turns.
## 2. Model Summary
<div align="center">
<table>
<tbody>
<tr>
<td align="center" style="vertical-align: middle; text-align: center"><strong>Architecture</strong></td>
<td align="center" style="vertical-align: middle; text-align: center">Hybrid Gated DeltaNet + Gated Attention, Mixture-of-Experts (MoE)</td>
</tr>
<tr>
<td align="center" style="vertical-align: middle; text-align: center"><strong>Total Parameters</strong></td>
<td align="center" style="vertical-align: middle; text-align: center">35B</td>
</tr>
<tr>
<td align="center" style="vertical-align: middle; text-align: center"><strong>Activated Parameters</strong></td>
<td align="center" style="vertical-align: middle; text-align: center">~3B</td>
</tr>
<tr>
<td align="center" style="vertical-align: middle; text-align: center"><strong>Number of Layers</strong></td>
<td align="center" style="vertical-align: middle; text-align: center">40</td>
</tr>
<tr>
<td align="center" style="vertical-align: middle; text-align: center"><strong>Layer Layout</strong></td>
<td align="center" style="vertical-align: middle; text-align: center">10 × (3 × (Gated DeltaNet → MoE) → 1 × (Gated Attention → MoE))</td>
</tr>
<tr>
<td align="center" style="vertical-align: middle; text-align: center"><strong>Hidden Dimension</strong></td>
<td align="center" style="vertical-align: middle; text-align: center">2048</td>
</tr>
<tr>
<td align="center" style="vertical-align: middle; text-align: center"><strong>Gated DeltaNet</strong></td>
<td align="center" style="vertical-align: middle; text-align: center">32 value heads, 16 QK heads, head dimension 128</td>
</tr>
<tr>
<td align="center" style="vertical-align: middle; text-align: center"><strong>Gated Attention</strong></td>
<td align="center" style="vertical-align: middle; text-align: center">16 Q heads, 2 KV heads, head dimension 256, RoPE dim 64</td>
</tr>
<tr>
<td align="center" style="vertical-align: middle; text-align: center"><strong>Number of Experts</strong></td>
<td align="center" style="vertical-align: middle; text-align: center">256</td>
</tr>
<tr>
<td align="center" style="vertical-align: middle; text-align: center"><strong>Selected Experts per Token</strong></td>
<td align="center" style="vertical-align: middle; text-align: center">8 routed + 1 shared</td>
</tr>
<tr>
<td align="center" style="vertical-align: middle; text-align: center"><strong>Expert Intermediate Dimension</strong></td>
<td align="center" style="vertical-align: middle; text-align: center">512</td>
</tr>
<tr>
<td align="center" style="vertical-align: middle; text-align: center"><strong>Vocabulary Size</strong></td>
<td align="center" style="vertical-align: middle; text-align: center">248,320</td>
</tr>
<tr>
<td align="center" style="vertical-align: middle; text-align: center"><strong>Context Length</strong></td>
<td align="center" style="vertical-align: middle; text-align: center">262,144 (extensible to 1,010,000)</td>
</tr>
<tr>
<td align="center" style="vertical-align: middle; text-align: center"><strong>Multi-Token Prediction</strong></td>
<td align="center" style="vertical-align: middle; text-align: center">1 MTP layer (up to 50% faster generation)</td>
</tr>
<tr>
<td align="center" style="vertical-align: middle; text-align: center"><strong>Vision Encoder</strong></td>
<td align="center" style="vertical-align: middle; text-align: center">27-layer ViT, hidden 1152, patch 16, temporal patch 2</td>
</tr>
<tr>
<td align="center" style="vertical-align: middle; text-align: center"><strong>Modality</strong></td>
<td align="center" style="vertical-align: middle; text-align: center">Text, Image, Video</td>
</tr>
<tr>
<td align="center" style="vertical-align: middle; text-align: center"><strong>Precision</strong></td>
<td align="center" style="vertical-align: middle; text-align: center">BF16 (this repo)</td>
</tr>
</tbody>
</table>
</div>
## 3. Datasets Used to Train
Fabric 1.6 was developed from the [Qwen3.5-35B-A3B-Base](https://huggingface.co/Qwen/Qwen3.5-35B-A3B-Base) foundation through **continuous pre-training** followed by **post-training** (supervised fine-tuning and reinforcement-learning-based alignment).
Pre-training was performed primarily on a large, in-house **proprietary synthetic dataset spanning code, mathematics and reasoning**, complemented by open reasoning corpora:
- [**OpenThoughts3-1.2M**](https://huggingface.co/datasets/open-thoughts/OpenThoughts3-1.2M) — 1.2M high-quality reasoning traces across mathematics, science, coding and general problem solving.
- [**OpenR1-Math-220k**](https://huggingface.co/datasets/open-r1/OpenR1-Math-220k) — 225k mathematical problems with think-style solutions.
Post-training instruction data combines permissively licensed open corpora with proprietary data:
- [**smoltalk2**](https://huggingface.co/datasets/HuggingFaceTB/smoltalk2) — an Apache-2.0 SFT subset (~340k examples) covering multilingual instruction following, multi-turn reasoning, tool-calling traces, system chats and table understanding.
- [**hermes-function-calling-v1**](https://huggingface.co/datasets/NousResearch/hermes-function-calling-v1) — structured tool-calling traces.
- [**oasst2**](https://huggingface.co/datasets/OpenAssistant/oasst2) — curated, reviewed conversational chains.
- [**maple**](https://huggingface.co/datasets/FabricAI/maple) — a proprietary instruction and reasoning corpus developed in-house by Fabric AI (CC-BY-4.0).
In total, approximately **12 billion tokens** were processed across the pre-training and post-training stages. Knowledge cutoff: July 2026.
## 4. Evaluation Results
Fabric 1.6 was evaluated on 22 benchmarks with greedy decoding (temperature 0).
| Category | Benchmark | Score |
|---|---|---|
| **Math & Reasoning** | AIME25 | 92.8 |
| | AIME26 | 93.1 |
| | HMMT26 | 83.2 |
| | IMOAB | 79.2 |
| | M500 | 84.8 |
| **Science & Knowledge** | GPQA | 86.7 |
| | GPQA-D | 84.9 |
| | HLE | 21.4 |
| | MMLU-P | 85.6 |
| | MMLU-R | 93.5 |
| | C-Eval | 92.3 |
| **Coding** | LCB6 | 80.2 |
| | SWEB-V | 72.9 |
| | SWEB-P | 50.1 |
| | IFEval | 93.09 |
| **General Reasoning** | GSM8K-Pt | 95.73 |
| **Agentic Tools** | TAU3 | 67.2 |
| | MMMU-P | 74.10 |
| | RWQA | 85.4 |
| | MCP-A | 62.8 |
| | WS | 60.3 |
| | MV-mini | 86.6 |
## 5. Deployment
> [!Note]
> Fabric 1.6 runs on the following inference engines with built-in MTP (multi-token prediction) support:
### vLLM
Install:
```bash
uv pip install vllm --torch-backend=auto
```
For deployment across an 8-GPU node (with built-in MTP support):
```bash
vllm serve FabricAI/Fabric1.6 --port 8000 --tensor-parallel-size 8 --max-model-len 262144 --reasoning-parser qwen3 --speculative-config '{"method":"qwen3_next_mtp","num_speculative_tokens":2}'
```
For tool use:
```bash
vllm serve FabricAI/Fabric1.6 --port 8000 --tensor-parallel-size 8 --max-model-len 262144 --reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder
```
### SGLang
Install:
```bash
uv pip install sglang[all]
```
For deployment across an 8-GPU node (with built-in MTP support):
```bash
python -m sglang.launch_server --model-path FabricAI/Fabric1.6 --port 8000 --tp-size 8 --mem-fraction-static 0.8 --context-length 262144 --reasoning-parser qwen3 --speculative-algo NEXTN --speculative-num-steps 3 --speculative-eagle-topk 1 --speculative-num-draft-tokens 4
```
For tool use:
```bash
python -m sglang.launch_server --model-path FabricAI/Fabric1.6 --port 8000 --tp-size 8 --mem-fraction-static 0.8 --context-length 262144 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
```
## 6. Model Usage
Fabric 1.6 always has thinking enabled and returns `reasoning_content` alongside the answer. The model was trained in **preserved thinking history** mode: for multi-turn conversations and tool calls, pass the complete assistant message returned by the API back to `messages` as-is — including `reasoning_content` and `tool_calls`, not just `content` — so that reasoning from earlier turns remains available to later ones.
## 7. License
The model weights are released under the [Apache License 2.0](LICENSE).
## 8. Citation
If you use Fabric 1.6 in your work, please cite it as:
```bibtex
@misc{fabric1.6,
title = {{Fabric1.6}: Agentic Open Model for Enterprises},
url = {https://huggingface.co/FabricAI/Fabric1.6},
author = {{Fabric AI}},
month = {August},
year = {2026}
}
```
## 9. Contact
For questions, collaborations or access requests, contact the Fabric AI research team at [research@fabricai.co.uk](mailto:research@fabricai.co.uk).