Hy-MT2-7B-NPU2
OpenFlowLM Q4NX conversion of tencent/Hy-MT2-7B-FP8 for AMD XDNA NPU inference.
This repository contains a quantized Q4NX port of the model, compiled for the OpenFlowLM (OFLM) runtime. It is not a GGUF file.
| Item | Value |
|---|---|
| Source model | tencent/Hy-MT2-7B-FP8 |
| Source GGUF | Hy-MT2-7B.Q8_0.gguf |
| Weights | model.q4nx (5.35 GB) |
| Modality | language |
| OFLM version | 0.1.0 |
| Converted | 2026-10-01 |
Install and run
This repository works with oflm-add, a small installer that copies the model
into the OpenFlowLM user directory and registers the tag. It never
modifies the system OpenFlowLM install.
pip install oflm-add or uv tool install oflm-add
uv tool install oflm-add
oflm-add Atomic-Germ/Hy-MT2-7B-NPU2 --family qwen3.5 --xclbin-from Hy-MT2-7B-NPU2
OFLM_CONFIG_PATH="$HOME/.config/oflm/model_list.json" OFLM_XCLBIN_PATH="$HOME/.config/oflm" oflm run Hy-MT2-7B-NPU2
Files
| File | Description |
|---|---|
model.q4nx |
Quantized weights (Q8_0 / Q4_1 / BF16) |
config.json |
OFLM runtime configuration |
tokenizer.json |
Tokenizer vocabulary |
tokenizer_config.json |
Tokenizer configuration |
chat_template.jinja |
Chat template |
Source model card
See the original model card: tencent/Hy-MT2-7B-FP8
- Downloads last month
- 11
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support