Text Generation
Transformers
Safetensors
English
fabric
efficient
0.7b
causal-lm
chunked-memory
conversational
custom_code
tudor-iustin22's picture
Transformers Serve Day-0 Support OpenAI API Compatible
27f97fc verified
|
Raw
History Blame Contribute Delete
9.11 kB
---
license: apache-2.0
language:
- en
library_name: transformers
pipeline_tag: text-generation
tags:
- fabric
- efficient
- 0.7b
- causal-lm
- chunked-memory
base_model: FabricAI/Fabric1.5-0.7B-Base
datasets:
- HuggingFaceTB/smoltalk2
- open-r1/Mixture-of-Thoughts
- nvidia/Nemotron-Post-Training-Dataset-v1
extra:
model_name: Fabric1.5-0.7B-Instruct
model_creator: Fabric AI
developer: Fabric AI
model_type: fabric
context_length: 32768
---
![Fabric 1.5 Banner](banner_fabric.png)
# Fabric 1.5 — 0.7B Instruct
## Overview
**Fabric 1.5** is a lightweight, high-capability language model developed by **Fabric AI**. It is built upon the **[Fabric1.5-0.7B-Base](https://huggingface.co/FabricAI/Fabric1.5-0.7B-Base)** pretrained checkpoint and fine-tuned for instruction following. It introduces a novel chunked memory architecture and demonstrates that a 0.7B parameter model can achieve meaningful performance on reasoning, knowledge, and instruction-following tasks.
## Model Details
- **Type:** Causal Language Model
- **Parameters:** 742M (0.7B)
- **Hidden Dimension:** 1,536
- **Vocabulary Size:** 65,536
- **Layers:** 24 (16 LocalBlock + 8 FabricMemoryBlock)
- **Attention:** GQA (24 query heads, 6 KV heads), RoPE (theta=1M)
- **Context Length:** 32,768 tokens native
- **Precision:** FP16
- **License:** Fabric AI Open License v1.0
## Architecture
Fabric 1.5 employs a **Chunked Fabric Memory** architecture — every 3rd block splits attention into two parallel branches:
- **Local causal window:** exact attention over the last 2,048 tokens
- **Chunked memory:** learned summaries (4 per 512-token chunk) of earlier completed chunks
A per-token scalar gate blends the outputs, giving the model 256 summary vectors of long-range memory at full 32K context without quadratic memory growth.
## Benchmark Results
Evaluated on an **NVIDIA A100-SXM4-40GB** using `lm-eval` harness with FP16 weights.
| Benchmark | Accuracy | Type |
|---|---|---|
| **ARC Easy** (0-shot) | **54.97%** | Loglikelihood |
| **ARC Challenge** (0-shot) | **26.88%** | Loglikelihood |
| **HellaSwag** (0-shot) | **35.00%** | Loglikelihood |
| **MMLU** (0-shot) | **28.96%** | Loglikelihood |
| **C-Eval** (0-shot) | **27.12%** | Loglikelihood |
### Comparison vs Qwen3.5-0.8B
| Benchmark | Fabric 1.5 | Qwen3.5-0.8B | Delta |
|---|---|---|---|
| **MMLU** (0-shot) | **28.96%** | 29.7% | -0.74% |
Fabric 1.5 performs within **0.74% of Qwen3.5-0.8B** on MMLU despite being 12% smaller (0.7B vs 0.8B) and having no vision encoder.
## Quickstart
### Serve with Hugging Face Transformers
Fabric 1.5 can be served through the Hugging Face `transformers serve` CLI using an OpenAI-compatible API.
**Install**
```bash
python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install --upgrade "transformers[serving]" openai
```
Fabric 1.5 uses custom Transformers code, so `--trust-remote-code` is required.
**Custom architecture compatibility patch**
Current versions of `transformers serve` may try to load `FabricForCausalLM` directly from the installed Transformers package instead of loading it from the model repository. Run this patch once inside the active virtual environment:
```bash
python - <<'PY'
from pathlib import Path
import shutil
import transformers.cli.serving.model_manager as model_manager
path = Path(model_manager.__file__)
backup = path.with_suffix(".py.bak")
old = """ architecture = getattr(transformers, config.architectures[0])
return architecture.from_pretrained(model_id, **model_kwargs)"""
new = """ from transformers import AutoModelForCausalLM
return AutoModelForCausalLM.from_pretrained(
model_id,
config=config,
**model_kwargs,
)"""
source = path.read_text()
if old not in source:
raise SystemExit(
f"Expected code was not found in:\n{path}\n"
"Your Transformers version may already contain a fix."
)
if not backup.exists():
shutil.copy2(path, backup)
path.write_text(source.replace(old, new))
print(f"Patched: {path}")
print(f"Backup: {backup}")
PY
```
**macOS with Apple Silicon**
```bash
transformers serve FabricAI/Fabric1.5-0.7B-Instruct \
--trust-remote-code \
--device mps \
--dtype float16 \
--attn-implementation eager \
--host 0.0.0.0 \
--port 8000
```
**NVIDIA GPU**
For NVIDIA GPUs with BF16 support:
```bash
transformers serve FabricAI/Fabric1.5-0.7B-Instruct \
--trust-remote-code \
--device cuda \
--dtype bfloat16 \
--attn-implementation eager \
--host 0.0.0.0 \
--port 8000
```
For NVIDIA GPUs that do not support BF16, use FP16:
```bash
transformers serve FabricAI/Fabric1.5-0.7B-Instruct \
--trust-remote-code \
--device cuda \
--dtype float16 \
--attn-implementation eager \
--host 0.0.0.0 \
--port 8000
```
**Stream a chat response**
Install `jq` if it is not already available. On macOS:
```bash
brew install jq
```
Then send an OpenAI-compatible streaming request:
```bash
curl -sN http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "FabricAI/Fabric1.5-0.7B-Instruct",
"messages": [
{
"role": "system",
"content": "You are Fabric, a helpful AI assistant."
},
{
"role": "user",
"content": "Explain neural networks in simple terms."
}
],
"max_tokens": 32768,
"stream": true,
"generation_config": "{\"do_sample\":true,\"temperature\":0.65,\"top_p\":0.9,\"top_k\":50,\"repetition_penalty\":1.05,\"no_repeat_ngram_size\":3}"
}' \
| sed -u -n 's/^data: //p' \
| jq --unbuffered -j 'select(.choices) | .choices[0].delta.content // empty'
printf '\n'
```
The server exposes an OpenAI-compatible API at `http://localhost:8000/v1`.
```python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "FabricAI/Fabric1.5-0.7B-Instruct"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_id,
trust_remote_code=True,
torch_dtype=torch.float16,
device_map="auto",
)
messages = [{"role": "user", "content": "What is gravity?"}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
with torch.no_grad():
output_ids = model.generate(
**inputs,
max_new_tokens=256,
do_sample=True,
temperature=0.65,
top_p=0.9,
repetition_penalty=1.05,
use_cache=True,
)
response = tokenizer.decode(output_ids[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True).strip()
print(response)
```
## Chat Template
```
<|bos|><|system|>
You are Fabric, a large language model developed by Fabric AI.<|end|>
<|user|>
Hello!<|end|>
<|assistant|>
Hi! How can I help you?<|end|>
```
Special tokens: `<bos>`, `<pad>`, `<|end|>`, `<|system|>`, `<|user|>`, `<|assistant|>`
## Sampling Parameters
| Mode | Temperature | Top-p | Top-k | Repetition Penalty |
|------|------------|-------|-------|-------------------|
| General chat | 0.65 | 0.90 | 50 | 1.05 |
| Creative writing | 0.80 | 0.95 | 50 | 1.00 |
| Factual/knowledge | 0.40 | 0.85 | 30 | 1.05 |
| Code generation | 0.30 | 0.90 | 40 | 1.02 |
## Training Data & Process
**Pre-training (18B tokens, 1× NVIDIA DGX H100 8-GPU):**
| Source | Weight | Description |
|---|---|---|
| FineWeb-Edu | 60% | High-quality educational web text |
| DCLM-Baseline | 25% | Deduplicated Common Crawl subset |
| OpenWebMath | 10% | Math-focused web text |
| Cosmopedia-v2 | 5% | Instruction-tuned educational data |
- **Optimizer:** AdamW (β1=0.9, β2=0.95, ε=1e-8)
- **Learning rate:** 3e-4 with cosine decay to 3e-5, 180M token warmup
- **Weight decay:** 0.1
- **Gradient clipping:** 1.0
- **Precision:** FP16
- **Distributed:** NCCL, 8 ranks, 4 sequences per GPU
**Post-training (SFT):**
1. Identity tuning — system prompt alignment and persona training
2. General SFT — diverse instruction-following datasets
## Hardware Notes
Fabric 1.5 was trained on a **single NVIDIA DGX H100** node with 8× H100 80GB GPUs. The model runs efficiently on:
| Hardware | Precision | Batch Size | Context |
|---|---|---|---|
| 1× NVIDIA H100 80GB | FP16 | 8 | 32K |
| 1× NVIDIA A100 40GB | FP16 | 4 | 32K |
| Apple Silicon (MPS) | FP16 | 1 | 32K |
| CPU | FP32 | 1 | 32K |
The chunked memory architecture ensures memory scales linearly with context, not quadratically.
## License
**Fabric AI Open License v1.0** — a permissive, attribution-required license. See [LICENSE](LICENSE) for full terms.
Key requirements:
- You may use, modify, and distribute the model freely
- You must **retain attribution** to Fabric AI when redistributing
- Modified files must carry notices of your changes
- No warranty or liability provisions
## Citation
```bibtex
@misc{fabric1.5,
title = {{Fabric 1.5}: A Causal Language Model with Chunked Fabric Memory},
author = {Fabric AI},
year = 2026,
}
```