Ornith-1.0-9B-NVFP4

NVFP4 / FP8 mixed-precision PTQ checkpoint of deepreinforce-ai/Ornith-1.0-9B, produced with NVIDIA TensorRT Model Optimizer (hf_ptq).

This is a community quantized derivative. All credit for the base model belongs to DeepReinforce. Please cite / attribute the upstream model when using this checkpoint.

Lineage

deepreinforce-ai/Ornith-1.0-9B   (BF16, Qwen3.5 hybrid VLM/LLM)
        │
        │  ModelOpt PTQ
        │  recipe: huggingface/qwen3_5/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast
        ▼
dmnsh/Ornith-1.0-9B-NVFP4       (this repo)
Field Value
Base model deepreinforce-ai/Ornith-1.0-9B
Quantization tool NVIDIA ModelOpt (hf_ptq)
Recipe huggingface/qwen3_5/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast
Calibration data NVIDIA Nemotron SFT / science / math / coding / agentic mixes (ModelOpt default calib set)
License MIT (inherits from base)

Quantization recipe (what changed)

From hf_quant_config.json:

Component Format
MLP projections (gate / up / down) W4A16 NVFP4 (group size 16)
Attention projections (full + linear attn) FP8
KV cache FP8
Vision tower (model.visual.*) BF16 (not quantized)
lm_head BF16 (excluded)

Top-level algo in the export config is MIXED_PRECISION (W4A16_NVFP4 + FP8).

Intended use

Same as the base Ornith model (agentic coding / reasoning), with lower memory and higher decode throughput on Blackwell GPUs (e.g. DGX Spark GB10). NVFP4 weight kernels need Blackwell; calibration can be done on other GPUs.

Serve with vLLM

Validated on DGX Spark with vllm/vllm-openai:nightly-aarch64:

vllm serve dmnsh/Ornith-1.0-9B-NVFP4 \
  --quantization modelopt_mixed \
  --max-model-len 4096 \
  --gpu-memory-utilization 0.164 \
  --served-model-name Ornith-1.0-9B-NVFP4

Notes:

  • Use --quantization modelopt_mixed for this mixed NVFP4+FP8 checkpoint.
  • For longer contexts, raise --max-model-len and --gpu-memory-utilization as needed.
  • Ornith is a reasoning model; enable a reasoning / tool parser when serving for agents (see the base model card).

DGX Spark speedups (BF16 → NVFP4)

vLLM on GB10, max-model-len=4096, concurrency 1.

Metric BF16 NVFP4 Change
Memory (GiB) 30.25 25.66 −15%
Prefill TTFT (ms) 435.7 151.3 −65% (2.9× faster)
Throughput (tok/s) 12.2 28.8 +135% (2.35×)

Files to ignore in this local folder

If uploading from a local checkout that still has working artifacts, do not upload:

  • model.safetensors.pre_lmhead_fix
  • hf_quant_config.json.bak
  • .quant_summary.txt

Disclaimer

This checkpoint is provided as-is for experimentation. Quantization can change model quality; run your own evals before production use. Not affiliated with DeepReinforce or NVIDIA beyond use of open ModelOpt tooling.

Downloads last month
17
Safetensors
Model size
7B params
Tensor type
BF16
·
F8_E4M3
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for dmnsh/Ornith-1.0-9B-NVFP4

Quantized
(90)
this model