How to use from
llama.cpp
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf SpacemiT/Qwen3.5-0.8B
# Run inference directly in the terminal:
llama cli -hf SpacemiT/Qwen3.5-0.8B
Install from WinGet (Windows)
winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf SpacemiT/Qwen3.5-0.8B
# Run inference directly in the terminal:
llama cli -hf SpacemiT/Qwen3.5-0.8B
Use pre-built binary
# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases
# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf SpacemiT/Qwen3.5-0.8B
# Run inference directly in the terminal:
./llama-cli -hf SpacemiT/Qwen3.5-0.8B
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli
# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf SpacemiT/Qwen3.5-0.8B
# Run inference directly in the terminal:
./build/bin/llama-cli -hf SpacemiT/Qwen3.5-0.8B
Use Docker
docker model run hf.co/SpacemiT/Qwen3.5-0.8B
Quick Links

Qwen3.5-0.8B for SpacemiT K1/K3

This is the SpacemiT deployment package for Qwen/Qwen3.5-0.8B, an Apache-2.0 native multimodal vision-language model for image understanding, text, coding and agent tasks. Please see the Qwen Team's Qwen3.5 announcement and cite the original work:

@misc{qwen3.5, title={{Qwen3.5}: Towards Native Multimodal Agents}, author={{Qwen Team}}, month={February}, year={2026}, url={https://qwen.ai/blog?id=qwen3.5}}

The text decoder is Q4_1 GGUF and the vision encoder is ONNX. The default configuration uses the 384-pixel encoder; 224- and 768-pixel alternatives are included.

Files and platforms

qwen3_5vl_0.8b-text-q41.gguf, three qwen3_5vl_0.8b-vision-*-op23.f16.onnx files, and configs/K1/config.json / configs/K3/config.json are included. K1/X60 uses AI cores 0–3 and -t 4; K3/A100 uses AI cores 8–15 and -t 8. Use the matching configuration: its ep_config sets the SpaceMIT EP affinity.

Prerequisites

Install the SpacemiT ONNX Runtime release and an SMT-enabled SpacemiT llama.cpp. Prebuilt packages can be unpacked directly:

wget https://github.com/spacemit-com/onnxruntime/releases/download/2.0.6/spacemit-ort.riscv64.2.0.6.tar.gz
wget https://github.com/spacemit-com/llama.cpp/releases/download/v0.1.7/spacemit-llama.cpp.riscv64.0.1.7.tar.gz
tar -xf spacemit-ort.riscv64.2.0.6.tar.gz; tar -xf spacemit-llama.cpp.riscv64.0.1.7.tar.gz

To build llama.cpp, clone recursively, set RISCV_ROOT_PATH and SPACEMIT_ORT_DIR, then run bash build_spacemit.sh glibc.

Run

export MODEL_DIR=/path/to/Qwen3.5-0.8B-SpacemiT
export LLAMA_DIR=/path/to/llama.cpp-installed
export ORT_DIR=/path/to/spacemit-ort.riscv64.2.0.6
export LD_LIBRARY_PATH="$LLAMA_DIR/lib:$ORT_DIR/lib:${LD_LIBRARY_PATH:-}"

K1: "$LLAMA_DIR/bin/llama-server" -m "$MODEL_DIR/qwen3_5vl_0.8b-text-q41.gguf" --media-backend smt --smt-config-dir "$MODEL_DIR/configs/K1" -t 4 --host 0.0.0.0 --port 8080 --warmup

K3: use the same command with configs/K3 and -t 8.

Send an OpenAI-compatible request to /v1/chat/completions with an image_url data URL and text such as Describe the image content., max_tokens: 64, temperature: 0, and chat_template_kwargs: {"enable_thinking": false}.

Board verification

Using humanspeech.jpg, ORT 2.0.6 and SMT llama.cpp, both boards returned HTTP 200. K1 began The image captures a lively scene of a speech or presentation...; K3 began This image captures a lively scene, likely from a formal event or presentation.... These are functional smoke tests, not benchmarks.

The original Qwen3.5 model is Apache-2.0; dependency licenses remain with their respective projects.

Downloads last month
16
GGUF
Model size
0.8B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support