--- license: other license_name: mixed-apache-2.0-and-mit license_link: https://huggingface.co/tgo-app-dev/vpipe-supplement/blob/main/THIRD_PARTY_LICENSES.md tags: - coreml - vpipe - apple-silicon - object-detection - voice-activity-detection - vision-language-model --- # VPIPE supplementary CoreML models This repository distributes the supplementary, pre-converted **Core ML** model packages consumed by the VPIPE runtime — a concurrent pipeline framework for real-time video tasks on Apple silicon, by T-Go LLC. These are third-party models converted to Core ML; they are released here so the VPIPE runtime can fetch a pinned, ready-to-run artifact instead of re-running the conversion. > **Licensing note:** the bundled models are **not** all under a single > license. Four are Apache-2.0 and two (Silero VAD, BEATs) are MIT. See > [THIRD_PARTY_LICENSES.md](THIRD_PARTY_LICENSES.md) for the per-model > attribution and verbatim license texts. ## Distribution format Each archive is a `.tar` (tracked with Git LFS) containing a single top-level `*.mlpackage` directory — an unflattened Core ML model package with its `Manifest.json`, the `model.mlmodel` spec, and the `weights/weight.bin` blob. Unpack before use: ```sh tar xf yolox_l_1024x640_w8.tar # -> yolox_l_1024x640_w8.mlpackage/ ``` The VPIPE runtime loads the unpacked `.mlpackage` via its `model_path` configuration. In the filenames, `w8` denotes 8-bit weight quantization (activations run in FP16 on the Apple Neural Engine); resolutions are written `WxH`. ## Models included | Archive | Model | Task | License | | --- | --- | --- | --- | | `qwen3_5_mlx_4b_vision_vid_512x320_w8.tar` | Qwen3.5-4B vision tower @ 512×320 (Core ML) | Vision encoder for a multimodal LLM | Apache-2.0 | | `qwen3_5_mlx_4b_vision_vid_768x480_w8.tar` | Qwen3.5-4B vision tower @ 768×480 (Core ML) | Vision encoder for a multimodal LLM | Apache-2.0 | | `gemma4_mlx_e4b_vision_768x480_w8.tar` | Gemma 4 (e4b) vision tower @ 768×480 (Core ML) | Vision encoder for a multimodal LLM | Apache-2.0 | | `yolox_l_1024x640_w8.tar` | YOLOX-L (Core ML) | Object detection | Apache-2.0 | | `silero-vad-unified-v6.tar` | Silero VAD v6, unified (Core ML) | Voice activity detection | MIT | | `beats_as2m_10s.tar` | BEATs iter3+ (AS2M, cpt1) (Core ML) | Audio tagging / sound-event classification | MIT | ### `qwen3_5_mlx_4b_vision_vid_512x320_w8` Qwen3.5-4B vision tower (`variant=tensor_video_rgb`), the vision encoder that turns a pair of video frames into the embedding sequence consumed by the Qwen3.5-4B language model. - **Task:** video → vision feature embeddings (multimodal LLM front-end). - **Inputs:** `image0`, `image1` — the two consecutive video frames that form one temporal patch, each a planar RGB fp16 tensor `(1, 3, 320, 512)` (`1×3×H×W`, i.e. 512×320 W×H) with raw pixels in `0..255`. The `x/127.5 - 1` normalisation is baked into the graph (applied on the ANE). - **Output:** `image_features` — vision features of shape `(160, 2560)`, fp16. - **Precision:** 8-bit weight quantization, FP16 compute. - **Upstream:** [Qwen3.5](https://huggingface.co/Qwen) (Qwen team, Alibaba Cloud) — Apache-2.0. - **Conversion:** coremltools 9.0 from TorchScript (`torch==2.12.0`). ### `qwen3_5_mlx_4b_vision_vid_768x480_w8` The same Qwen3.5-4B vision tower (`variant=tensor_video_rgb`) at a higher 768×480 input — identical I/O contract to the 512×320 build, with a longer output token sequence. - **Task:** video → vision feature embeddings (multimodal LLM front-end). - **Inputs:** `image0`, `image1` — the two consecutive video frames that form one temporal patch, each a planar RGB fp16 tensor `(1, 3, 480, 768)` (`1×3×H×W`, i.e. 768×480 W×H) with raw pixels in `0..255`. The `x/127.5 - 1` normalisation is baked into the graph (applied on the ANE). - **Output:** `image_features` — vision features of shape `(360, 2560)`, fp16. - **Precision:** 8-bit weight quantization, FP16 compute. - **Upstream:** [Qwen3.5](https://huggingface.co/Qwen) (Qwen team, Alibaba Cloud) — Apache-2.0. - **Conversion:** coremltools 9.0 from TorchScript (`torch==2.12.0`). ### `gemma4_mlx_e4b_vision_768x480_w8` Gemma 4 (e4b) vision tower, converted to Core ML. It pools and projects the visual features into the language model's 2560-dim text embedding space, emitting "soft tokens" ready to splice into the Gemma 4 prompt. - **Task:** image → LLM soft tokens (multimodal LLM front-end). - **Input:** `image` — a single planar RGB fp16 MLMultiArray `[1, 3, 480, 768]` (`1×3×H×W`, i.e. 768×480 W×H), raw pixels in `0..255` (channel order R, G, B). Preprocessing is baked in: `x/255` then `2*(x - 0.5)`. - **Output:** `soft_tokens` — shape `[1, 160, 2560]`, fp16. - **Precision:** int8 weight-only quantization (per-block), FP16 compute. - **Upstream:** [Gemma 4 (e4b)](https://huggingface.co/google/gemma-4-E4B) (Google DeepMind) — Apache-2.0. - **Conversion:** coremltools 9.0 from TorchScript (`torch==2.12.1`). ### `yolox_l_1024x640_w8` YOLOX-L anchor-free object detector, converted to Core ML at a fixed **1024×640 (W×H)** input. - **Task:** object detection (bounding boxes + per-class scores). - **Input:** `image` — Core ML image input at 1024×640. - **Output:** `output` — detection tensor `[1, N, 4 + num_classes]` (`(cx, cy, w, h)` in model-input pixels plus sigmoid class scores). - **Precision:** 8-bit weight quantization, FP16 compute. - **Upstream:** [YOLOX](https://github.com/Megvii-BaseDetection/YOLOX) (Megvii Inc.) — Apache-2.0. - **Conversion:** coremltools 9.0 from TorchScript (`torch==2.12.1`); all operations placed on the Apple Neural Engine (no CPU fallback ops). ### `silero-vad-unified-v6` Silero VAD v6.0.0, the unified single-graph (STFT + encoder + decoder) Core ML conversion by FluidInference. - **Task:** voice activity detection (per-window speech probability). - **Inputs:** `audio_input` (mono PCM window) plus the recurrent `hidden_state` / `cell_state` carried between calls. - **Outputs:** `vad_output` (speech probability) plus `new_hidden_state` / `new_cell_state`. - **Precision:** FP16 compute. - **Upstream:** [Silero VAD](https://github.com/snakers4/silero-vad) (Silero Team) — MIT; Core ML conversion from [FluidInference/silero-vad-coreml](https://huggingface.co/FluidInference/silero-vad-coreml) — MIT. - **Conversion:** coremltools 9.0b1 from TorchScript (`torch==2.7.0`). ### `beats_as2m_10s` BEATs iter3+ audio representation model, fine-tuned on AudioSet-2M (the `cpt1` checkpoint), converted to Core ML over a fixed 10-second window. It classifies audio into the AudioSet sound-event ontology. - **Task:** audio tagging / sound-event classification. - **Input:** `waveform` — a mono 16 kHz waveform over a fixed 10-second window (160,000 samples). - **Output:** `probs` — per-class probabilities over the AudioSet ontology (527 sound-event classes). - **Precision:** FP16 compute (weights not quantized). - **Upstream:** [BEATs](https://github.com/microsoft/unilm/tree/master/beats) (Microsoft, `microsoft/unilm`) — MIT. - **Conversion:** coremltools 9.0 from TorchScript (`torch==2.7.0`). ## Verifying downloads SHA-256 checksums of the archives: ``` ed27924a97a873a019d1c8a1e1f9826bd30da2588683843d6c59cfb0b2289f30 qwen3_5_mlx_4b_vision_vid_512x320_w8.tar 3e928ea0c3e6848896278b68436d46f1aa306d95d8ca4a1faadfad354d705b87 qwen3_5_mlx_4b_vision_vid_768x480_w8.tar 17bebe9fdb416a760df4d00ed39c7edf296ba768d54139c8183c1323da646d3e gemma4_mlx_e4b_vision_768x480_w8.tar 7f7bf85947bb228f978470284e16a9d00c906a6cb07359f9a51c68806fa9c609 yolox_l_1024x640_w8.tar ee14d006d1cea87f15a4f0ac5f3e1e00b80a10c40eca855e8ff6741f389b4423 silero-vad-unified-v6.tar 47c63f83bfa47e248b199e820ae6a01a5ed40a5f6ef4b85f42f12d7ac5754638 beats_as2m_10s.tar ``` ## Licenses These are third-party models redistributed in converted form. Each retains its upstream license; see [THIRD_PARTY_LICENSES.md](THIRD_PARTY_LICENSES.md) for full attribution and verbatim license texts.