Onyx · Subject Video MLLM encoder (Qwen3-VL-2B-Thinking, bf16)

The multimodal feature encoder for Onyx's Subject Video (S2V) pipeline on OnyxMac: caption + reference images → [1, 1024, 2048] hidden-state features that guide the S2V DiT alongside the umT5 text conditioning. Unmodified Qwen3-VL-2B-Thinking weights in the HF layout; Onyx's own MLX extractor (S2VFeatureExtractor) loads this directory directly.

  • model.safetensors — 4.26 GB bf16 (model.visual.* + model.language_model.*).
  • sha256.json — per-file integrity manifest.

License: Apache-2.0 (inherited from the base model).

Downloads last month
51
Safetensors
Model size
2B params
Tensor type
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for wabibito/Onyx-Qwen3VL-2B-S2V-Encoder

Finetuned
(23)
this model