Ornith-1.0-35B-INT4-AutoRound
This is a W4A16 (INT4 weight, BF16 activation) quantized version of deepreinforce-ai/Ornith-1.0-35B using spark-auto-round.
Base Model
Ornith-1.0-35B by DeepReinforce AI — a 33.4B parameter MoE model with vision capabilities, built on the Qwen3.5 architecture.
Quantization
| Detail | Value |
|---|---|
| Tool | spark-auto-round v0.14.3 by whpthomas |
| Format | W4A16 (INT4 weights, group_size=128, symmetric) |
| Dataset | opencode-instruct (128 samples, seqlen=2048) |
| Iterations | 200 per block |
| Batch size | 2 |
| Hardware | DGX Spark (GB10, 128 GB unified) |
| Peak RAM | 115.71 GB |
| Peak VRAM | 22.65 GB |
Quality
| Metric | Value |
|---|---|
| Avg cosine similarity | 0.9188 |
| Avg PSNR | 50.8 dB |
| Blocks | 39 transformer layers quantized |
Size
| Format | Size |
|---|---|
| Original (BF16) | ~62.2 GB |
| Quantized (INT4) | ~20 GB (5 safetensor shards) |
Usage
Compatible with vLLM (auto-detects quantization: auto-round, uses Marlin kernel):
vllm serve yourname/Ornith-1.0-35B-int4-AutoRound \
--trust-remote-code \
--enable-auto-tool-choice \
--tool-call-parser qwen3_xml \
--reasoning-parser qwen3 \
--load-format safetensors
Credits
- Original model: DeepReinforce AI — Ornith-1.0-35B
- Quantization tool: spark-auto-round by whpthomas (fork of Intel's AutoRound optimized for DGX Spark)
- Calibration data: opencode-instruct
- Downloads last month
- 151
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support
Model tree for rwillox/Ornith-1.0-35B-int4-AutoRound
Base model
ornith-ai/Ornith-1.0-35B