Instructions to use dmnsh/Ornith-1.0-9B-NVFP4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use dmnsh/Ornith-1.0-9B-NVFP4 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="dmnsh/Ornith-1.0-9B-NVFP4") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("dmnsh/Ornith-1.0-9B-NVFP4") model = AutoModelForMultimodalLM.from_pretrained("dmnsh/Ornith-1.0-9B-NVFP4", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use dmnsh/Ornith-1.0-9B-NVFP4 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "dmnsh/Ornith-1.0-9B-NVFP4" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "dmnsh/Ornith-1.0-9B-NVFP4", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/dmnsh/Ornith-1.0-9B-NVFP4
- SGLang
How to use dmnsh/Ornith-1.0-9B-NVFP4 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "dmnsh/Ornith-1.0-9B-NVFP4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "dmnsh/Ornith-1.0-9B-NVFP4", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "dmnsh/Ornith-1.0-9B-NVFP4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "dmnsh/Ornith-1.0-9B-NVFP4", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use dmnsh/Ornith-1.0-9B-NVFP4 with Docker Model Runner:
docker model run hf.co/dmnsh/Ornith-1.0-9B-NVFP4
Ornith-1.0-9B-NVFP4
NVFP4 / FP8 mixed-precision PTQ checkpoint of deepreinforce-ai/Ornith-1.0-9B, produced with NVIDIA TensorRT Model Optimizer (hf_ptq).
This is a community quantized derivative. All credit for the base model belongs to DeepReinforce. Please cite / attribute the upstream model when using this checkpoint.
Lineage
deepreinforce-ai/Ornith-1.0-9B (BF16, Qwen3.5 hybrid VLM/LLM)
│
│ ModelOpt PTQ
│ recipe: huggingface/qwen3_5/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast
▼
dmnsh/Ornith-1.0-9B-NVFP4 (this repo)
| Field | Value |
|---|---|
| Base model | deepreinforce-ai/Ornith-1.0-9B |
| Quantization tool | NVIDIA ModelOpt (hf_ptq) |
| Recipe | huggingface/qwen3_5/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast |
| Calibration data | NVIDIA Nemotron SFT / science / math / coding / agentic mixes (ModelOpt default calib set) |
| License | MIT (inherits from base) |
Quantization recipe (what changed)
From hf_quant_config.json:
| Component | Format |
|---|---|
MLP projections (gate / up / down) |
W4A16 NVFP4 (group size 16) |
| Attention projections (full + linear attn) | FP8 |
| KV cache | FP8 |
Vision tower (model.visual.*) |
BF16 (not quantized) |
lm_head |
BF16 (excluded) |
Top-level algo in the export config is MIXED_PRECISION (W4A16_NVFP4 + FP8).
Intended use
Same as the base Ornith model (agentic coding / reasoning), with lower memory and higher decode throughput on Blackwell GPUs (e.g. DGX Spark GB10). NVFP4 weight kernels need Blackwell; calibration can be done on other GPUs.
Serve with vLLM
Validated on DGX Spark with vllm/vllm-openai:nightly-aarch64:
vllm serve dmnsh/Ornith-1.0-9B-NVFP4 \
--quantization modelopt_mixed \
--max-model-len 4096 \
--gpu-memory-utilization 0.164 \
--served-model-name Ornith-1.0-9B-NVFP4
Notes:
- Use
--quantization modelopt_mixedfor this mixed NVFP4+FP8 checkpoint. - For longer contexts, raise
--max-model-lenand--gpu-memory-utilizationas needed. - Ornith is a reasoning model; enable a reasoning / tool parser when serving for agents (see the base model card).
DGX Spark speedups (BF16 → NVFP4)
vLLM on GB10, max-model-len=4096, concurrency 1.
| Metric | BF16 | NVFP4 | Change |
|---|---|---|---|
| Memory (GiB) | 30.25 | 25.66 | −15% |
| Prefill TTFT (ms) | 435.7 | 151.3 | −65% (2.9× faster) |
| Throughput (tok/s) | 12.2 | 28.8 | +135% (2.35×) |
Files to ignore in this local folder
If uploading from a local checkout that still has working artifacts, do not upload:
model.safetensors.pre_lmhead_fixhf_quant_config.json.bak.quant_summary.txt
Disclaimer
This checkpoint is provided as-is for experimentation. Quantization can change model quality; run your own evals before production use. Not affiliated with DeepReinforce or NVIDIA beyond use of open ModelOpt tooling.
- Downloads last month
- 17
Model tree for dmnsh/Ornith-1.0-9B-NVFP4
Base model
deepreinforce-ai/Ornith-1.0-9B