How to use from
Pi
Start the llama.cpp server
# Install llama.cpp:
brew install llama.cpp
# Start a local OpenAI-compatible server:
llama serve -hf SpacemiT/Qwen3.5-2B
Configure the model in Pi
# Install Pi:
npm install -g @earendil-works/pi-coding-agent
# Add to ~/.pi/agent/models.json:
{
  "providers": {
    "llama-cpp": {
      "baseUrl": "http://localhost:8080/v1",
      "api": "openai-completions",
      "apiKey": "none",
      "models": [
        {
          "id": "SpacemiT/Qwen3.5-2B"
        }
      ]
    }
  }
}
Run Pi
# Start Pi in your project directory:
pi
Quick Links

Qwen3.5-2B for SpacemiT K1/K3

This package deploys Qwen/Qwen3.5-2B, the Qwen Team's Apache-2.0 native multimodal vision-language model. It supports visual understanding, text, coding and agent workflows. See the Qwen3.5 blog and cite:

@misc{qwen3.5, title={{Qwen3.5}: Towards Native Multimodal Agents}, author={{Qwen Team}}, month={February}, year={2026}, url={https://qwen.ai/blog?id=qwen3.5}}

The Q4_1 GGUF text decoder is paired with ONNX vision encoders at 224, 384 (default), and 768 pixels. Files are qwen3_5_2b-text-q41.gguf, the three qwen3_5_2b-vision-*-op23.f16.onnx files, and configs/K1 / configs/K3.

Install SpacemiT ONNX Runtime and SMT-enabled llama.cpp, or unpack spacemit-ort.riscv64.2.0.6.tar.gz and spacemit-llama.cpp.riscv64.0.1.7.tar.gz. Source builds use git clone --recursive, RISCV_ROOT_PATH, SPACEMIT_ORT_DIR, and bash build_spacemit.sh glibc.

K1/X60 has AI cores 0–3: use configs/K1 and -t 4. K3/A100 has AI cores 8–15: use configs/K3 and -t 8. The configs' ep_config contains the corresponding affinity.

export MODEL_DIR=/path/to/Qwen3.5-2B-SpacemiT LLAMA_DIR=/path/to/llama.cpp-installed ORT_DIR=/path/to/spacemit-ort.riscv64.2.0.6
export LD_LIBRARY_PATH="$LLAMA_DIR/lib:$ORT_DIR/lib:${LD_LIBRARY_PATH:-}"
"$LLAMA_DIR/bin/llama-server" -m "$MODEL_DIR/qwen3_5_2b-text-q41.gguf" --media-backend smt --smt-config-dir "$MODEL_DIR/configs/K1" -t 4 --host 0.0.0.0 --port 8080 --warmup

For K3 change K1 to K3 and -t 4 to -t 8. POST an image data URL and Describe the image content. to /v1/chat/completions with max_tokens: 64, temperature: 0, and thinking disabled. Board smoke tests with humanspeech.jpg returned HTTP 200 on both K1 (0-3, output beginning Here's a detailed description of the image:) and K3 (8-15, output beginning The image displays a collection of books...).

The original model is Apache-2.0; llama.cpp and ONNX Runtime retain their own licenses.

Downloads last month
10
GGUF
Model size
2B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support