--- license: apache-2.0 base_model: robbyant/lingbot-video-dense-1.3b pipeline_tag: text-to-video library_name: coreai tags: - coreai - core-ai - coreai-fabric - aimodel - coreml - apple - apple-silicon - on-device - video - text-to-video - image-to-video - vae - diffusion --- > **Canonical:** [`kevinqz/LingBot-Video-Dense-1.3B-CoreAI`](https://huggingface.co/kevinqz/LingBot-Video-Dense-1.3B-CoreAI) — source of truth. # LingBot-Video Dense 1.3B (fabric) An Apple Core AI conversion of the **VAE decoder** from [robbyant/lingbot-video-dense-1.3b](https://huggingface.co/robbyant/lingbot-video-dense-1.3b) — the `AutoencoderKLWan` video autoencoder's decode path, mapping a **video latent** to **pixel frames**. Produced by [coreai-fabric](https://github.com/kevinqz/coreai-fabric/blob/main/recipes/lingbot-video-dense-1.3b.yaml) and indexed by [coreai-catalog](https://github.com/kevinqz/coreai-catalog). > **VAE decoder, not the full video pipeline.** A video diffusion model is four > separable blocks — text encoder, VAE encoder, denoising DiT, and this VAE > decoder. This asset is ONLY the decoder (latent → pixels). The **host owns** the > text encoder, the DiT few-step denoise loop, latent un-normalization > (`latents_mean`/`latents_std`), frame assembly, and — for multi-chunk streaming — > the causal `feat_cache`. It does not, by itself, generate video from a prompt. ## Model facts | Field | Value | |---|---| | Parameters (upstream) | 1.3B | | Architecture | diffusion | | Capabilities | text-to-video | | Deployable core | AutoencoderKLWan VAE decoder (latent → frames) | | Decoded frame shape | 1×3×1×480×832 | | Quantization / precision | none / float32 | | On-disk size | 273 MB | | Asset kind | single-graph VAE decoder (first-chunk decode) | | assetVersion | 2.0 | ## Use it — this needs host code you supply The bundle is a single static-size graph: `z [1,16,T,H/8,W/8]` in → `frames [1,3,Tp,H,W]` out (spatial 8× / temporal 4× upsampling, first-chunk decode). **You supply** the DiT denoise loop that produces the latent, the latent un-normalization, and frame assembly in your host code (Swift or Python). ```bash pip install coreai-catalog && coreai-catalog install lingbot-video-dense-1.3b ``` ## Requirements - **Deployment: macOS 27.0+ / iOS 27.0+, Xcode 27+.** The asset serializes with `minimum_os v27`, so the on-device Swift runtime requires macOS/iOS 27+. A Mac on macOS 26 can convert and inspect it but not run it on-device. - Apple Silicon. ## Verification (output parity) - **Gate A (structure): passed** — the bundle's layout + metadata were validated; the graph loads. - **Gate B — graph_output_cosine: 1.000000 min output cosine** (median 1.000000) vs the fp32 torch VAE decoder over 8 seeded latents, measured on apple_silicon. Certifies the export decodes the SAME pixels as the source VAE — a conversion-fidelity metric, not end video quality. - This certifies the export is **numerically faithful to the source VAE decoder** — it does **NOT** certify end video quality. Reproduce with `coreai-fabric verify`. ## Provenance | Field | Value | |---|---| | Base model | [robbyant/lingbot-video-dense-1.3b](https://huggingface.co/robbyant/lingbot-video-dense-1.3b) @ `f9789a7d9b4772a47aba62d4eb5282ddefd1da21` | | Converted by | `models/lingbotvideo/export.py` (version not reported) | | Recipe | [lingbot-video-dense-1.3b](https://github.com/kevinqz/coreai-fabric/blob/main/recipes/lingbot-video-dense-1.3b.yaml) (recipe_source: fabric) | | Precision / quantization | float32 / none | | Conversion date | 2026-07-09 | Machine-readable, in this repo: [`parity-report.json`](./parity-report.json) · [`reproduce-manifest.json`](./reproduce-manifest.json) · [`LICENSE`](./LICENSE). ## License and attribution Weights licensed **apache-2.0** — see the bundled `LICENSE`. This artifact is a **converted derivative** of the base VAE: its weights were converted to Apple Core AI format. The conversion itself is community work. ## Links - **Base model:** [robbyant/lingbot-video-dense-1.3b](https://huggingface.co/robbyant/lingbot-video-dense-1.3b) - **Reproduce:** [recipe `lingbot-video-dense-1.3b`](https://github.com/kevinqz/coreai-fabric/blob/main/recipes/lingbot-video-dense-1.3b.yaml) - **Index:** [coreai-catalog](https://github.com/kevinqz/coreai-catalog) - [HF Collection](https://huggingface.co/collections/kevinqz/coreai-apple-on-device-6a4879f21c7e1a87c99bcf5a) ## The on-device Core AI ecosystem - [coreai-fabric](https://github.com/kevinqz/coreai-fabric) — the reproducible recipe → `.aimodel` pipeline that produced this asset. - [coreai-catalog](https://github.com/kevinqz/coreai-catalog) — the index of Core AI models with provenance and integration snippets. - [apple/coreai-models](https://github.com/apple/coreai-models) — Apple's official exporters and runtimes. ## Not affiliated with Apple Community conversion. Not produced, hosted, or endorsed by Apple. Apple and Core AI are trademarks of Apple Inc., used here only to describe the target runtime/format.