How to use from
llama.cpp
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf gsrunion/Ornith-1.0-35B-ROCmFPX-GGUF:
# Run inference directly in the terminal:
llama cli -hf gsrunion/Ornith-1.0-35B-ROCmFPX-GGUF:
Install from WinGet (Windows)
winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf gsrunion/Ornith-1.0-35B-ROCmFPX-GGUF:
# Run inference directly in the terminal:
llama cli -hf gsrunion/Ornith-1.0-35B-ROCmFPX-GGUF:
Use pre-built binary
# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases
# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf gsrunion/Ornith-1.0-35B-ROCmFPX-GGUF:
# Run inference directly in the terminal:
./llama-cli -hf gsrunion/Ornith-1.0-35B-ROCmFPX-GGUF:
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli
# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf gsrunion/Ornith-1.0-35B-ROCmFPX-GGUF:
# Run inference directly in the terminal:
./build/bin/llama-cli -hf gsrunion/Ornith-1.0-35B-ROCmFPX-GGUF:
Use Docker
docker model run hf.co/gsrunion/Ornith-1.0-35B-ROCmFPX-GGUF:
Quick Links

Ornith-1.0-35B — ROCmFPX quants for AMD Strix Halo

ROCmFPX-family GGUF quantization of deepreinforce-ai/Ornith-1.0-35B (Qwen3.5-MoE architecture, multimodal), produced for AMD Strix Halo (Ryzen AI Max, gfx1151) and the hal0 home inference platform.

⚠️ These files require the Hal0ai/Hal0_ROCmFPX llama.cpp fork (or the ghcr.io/hal0ai/hal0-rocmfpx container image that hal0 uses). Stock llama.cpp will reject the tensor types (invalid ggml type 101).

See also the companion repo: gsrunion/Ornith-1.0-9B-ROCmFPX-GGUF.

Files

File Quant BPW Size Notes
Ornith-1.0-35B-Q4_0_ROCMFP4_STRIX_LEAN.gguf Q4_0_ROCMFP4_STRIX_LEAN ~4.3 18.6 GB Size-biased Strix recipe — fits comfortably in a 64 GB GPU carve-out with long context
mmproj-BF16.gguf BF16 0.90 GB Vision projector (Ornith is multimodal) — load alongside the quant
imatrix_unsloth.gguf_file 184 MB Importance matrix used for calibration (from unsloth, included for reproducibility)

Measured performance

On AMD Ryzen AI Max+ 395 (Strix Halo, 128 GB unified LPDDR5X, ROCm backend, hal0-rocmfpx image, 64K ctx):

  • Q4_0_ROCMFP4_STRIX_LEAN: ~66 tok/s decode — the MoE's small active-parameter set decodes faster than the dense 9B (42 tok/s) despite the 3.7× file size.

How it was made

BF16 GGUF source and imatrix from unsloth/Ornith-1.0-35B-GGUF, quantized with the Hal0_ROCmFPX fork's llama-quantize:

llama-quantize --imatrix imatrix_unsloth.gguf_file \
  Ornith-1.0-35B-BF16.gguf Ornith-1.0-35B-Q4_0_ROCMFP4_STRIX_LEAN.gguf \
  Q4_0_ROCMFP4_STRIX_LEAN

Serving

Directly with the fork's llama-server (as done on hal0 boxes):

llama-server -m Ornith-1.0-35B-Q4_0_ROCMFP4_STRIX_LEAN.gguf \
  --mmproj mmproj-BF16.gguf -ngl 999 -fa on --jinja -c 65536

Or on hal0: download into the model store, register via add-from-path, then hal0 slot create <name> --type llm --hardware rocm --model <id> and load via a long-lived curl.

Credits

Downloads last month
378
GGUF
Model size
35B params
Architecture
qwen35moe
Hardware compatibility
Log In to add your hardware

4-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for gsrunion/Ornith-1.0-35B-ROCmFPX-GGUF

Quantized
(166)
this model