--- license: apache-2.0 base_model: CSWRY/VOSR pipeline_tag: image-to-image tags: - image-super-resolution - image-restoration - quantized - mobile - qualcomm - qnn - snapdragon - npu --- # VOSR 1.4B Mobile `AllGPTORG/VOSR_1.4B_Mobile` is a mobile deployment package of the **VOSR-1.4B one-step** image super-resolution model, quantized and compiled as Qualcomm QNN DLC graphs for the Snapdragon 8 Gen 3 NPU. The package is designed for high-quality image restoration and super-resolution, including text-rich images. It preserves the original one-step VOSR pipeline while splitting the network into smaller graphs suitable for mobile integration. > This repository contains QNN deployment artifacts. It is not a Transformers or > Diffusers checkpoint and cannot be loaded with `from_pretrained()`. ## Model summary | Property | Value | |---|---| | Base model | [CSWRY/VOSR](https://huggingface.co/CSWRY/VOSR), VOSR-1.4B one-step | | Task | Generative image restoration and super-resolution | | Target SoC | Qualcomm Snapdragon 8 Gen 3 / SM8650 | | Compile target | Samsung Galaxy S24 family, Android 14 | | Runtime format | Qualcomm QNN DLC | | Static tile size | 512 x 512 pixels | | Batch size | 1 | | Activations | FP16 (A16) | | DiT block weights | INT4 (W4A16) | | Auxiliary/final graph weights | INT8 (W8A16), except the FP16 VAE decoder | | Total DLC size | 1,736,469,708 bytes / 1.617 GiB | ## Files The graphs must be executed in the order shown below. | Order | File | Precision | Size | |---:|---|---:|---:| | 1 | `vosr_dinov2l_layer17.dlc` | W8A16 | 229.50 MiB | | 2 | `vosr_qwen_vae_encoder.dlc` | W8A16 | 19.09 MiB | | 3 | `vosr_dit_prepare.dlc` | W8A16 | 42.61 MiB | | 4 | `vosr_dit_blocks_00_12.dlc` | W4A16 | 444.86 MiB | | 5 | `vosr_dit_blocks_12_18.dlc` | W4A16 | 222.62 MiB | | 6 | `vosr_dit_blocks_18_24.dlc` | W4A16 | 222.62 MiB | | 7 | `vosr_dit_blocks_24_36.dlc` | W4A16 | 444.86 MiB | | 8 | `vosr_dit_final.dlc` | W8A16 | 4.90 MiB | | 9 | `vosr_qwen_vae_decoder.dlc` | Legacy W8A16 DLC; EPContext uses FP16 | 24.97 MiB | `manifest.json` contains the same graph order, precision assignment, and exact byte size for programmatic use. ### Ready-to-run SM8650 EPContext package `sd8g3/qnn-context/` contains eleven pre-linked QNN context binaries for the Snapdragon 8 Gen 3 HTP. They were linked with QAIRT `2.45.0.260326154327`, target DSP v75 / SoC model 57, and use O1 graph finalization to stay inside the device's 8 MiB VTCM budget. The DiT is split into six contiguous stages: `00_06`, `06_12`, `12_18`, `18_24`, `24_30`, and `30_36`. RMSNorm pointwise multiplications are divided along the token axis before linking; a CPU reference comparison against the unmodified QDQ block measured 72.9 dB PSNR. The complete context package is 1,140,031,488 bytes (1.062 GiB). The small EPContext ONNX wrappers and their exact SHA-256, byte-size, tensor, graph-order, and immutable Hub-revision contracts are stored in `sd8g3/vosr_runtime_manifest.json`. The Qwen VAE decoder is compiled directly from the original FP16 ONNX graph. The earlier W8A16 QDQ decoder removed nearly all spatial detail. The FP16 SM8650 context matches the CPU ONNX decoder at 59.77 dB PSNR with a maximum per-channel error of one RGB level. ## Tensor interface All image and latent tensors use NCHW layout. | Graph | Inputs | Outputs | |---|---|---| | DINOv2-L layer 17 | `lq_image`: FP16 `[1,3,512,512]` | `dino_features`: FP16 `[1,1024,1024]` | | Qwen VAE encoder | `lq_image`: FP16 `[1,3,512,512]`; `posterior_noise`: FP16 `[1,16,64,64]` | `lq_latent`: FP16 `[1,16,64,64]` | | DiT prepare | `latent_pair`: FP16 `[1,32,64,64]`; `timestep`: FP32 `[1]`; `next_timestep`: FP32 `[1]`; `dino_features`: FP16 `[1,1024,1024]` | `hidden`: FP16 `[1,1024,1536]`; `conditioning`: FP16 `[1,1536]`; `block_conditioning`: FP16 `[1,9216]`; `projected_dino`: FP16 `[1,1024,1536]` | | DiT block stages | `hidden`, `block_conditioning`, `projected_dino` | `hidden_out`: FP16 `[1,1024,1536]` | | DiT final | `hidden`: FP16 `[1,1024,1536]`; `conditioning`: FP16 `[1,1536]` | `velocity`: FP16 `[1,16,64,64]` | | Qwen VAE decoder | `normalized_latent`: FP16 `[1,16,64,64]` | `sr_image`: FP16 `[1,3,512,512]` | ## Integration outline Use the Qualcomm AI Engine Direct SDK / QNN runtime to load and execute the DLCs. The host application is responsible for preprocessing, graph orchestration, random noise generation, the one-step latent update, tiling, and image postprocessing. 1. Resize the low-resolution image to the requested output resolution using bicubic interpolation, split it into 512 x 512 tiles if needed, convert RGB values to FP16 NCHW, and normalize them to `[-1, 1]`. 2. Run the DINO graph and Qwen VAE encoder on the same image tile. Supply seeded normal noise as `posterior_noise` if reproducible output is required. 3. Create an FP16 normal-noise latent `z` with shape `[1,16,64,64]`, concatenate `lq_latent` and `z` along the channel axis, and use the result as `latent_pair`. 4. For the one-step schedule, run DiT prepare with `timestep = [1.0]` and `next_timestep = [0.0]`. 5. Pass `hidden` sequentially through all four DiT block-stage DLCs, or through all six `sd8g3/qnn-context` DiT stages when using the EPContext package. Reuse `block_conditioning` and `projected_dino` for every stage. 6. Run DiT final and apply the one-step update `z = z - velocity`. 7. Decode the updated latent with the Qwen VAE decoder, clamp the output to `[-1, 1]`, convert it back to RGB, and blend overlapping tiles when tiling. For 4x super-resolution, bicubic-upscale the source to the final target resolution before creating the model tiles. The network then restores detail at that target resolution. ## Validation status - The DINO, VAE encoder, DiT prepare/final, and DiT blocks use their documented mobile quantization. The VAE decoder deliberately remains FP16. - All runtime graphs were successfully compiled and linked for SM8650. - The Qwen VAE encoder and DiT prepare graphs were profiled on the Snapdragon 8 Gen 3 NPU. - All eleven split EPContext graphs were linked and inspected as DSP v75 / SoC model 57 contexts. The runtime disables CPU execution-provider fallback. - The complete on-device chain preserves generated detail: on the S24 Ultra regression patch, edge energy rose from 1.54 for the bicubic input to 15.19 after VOSR. - The `00_06` six-block context was profiled successfully on the Galaxy S24 target: 382.9 ms estimated warm inference, 230.9 ms warm load, and 119.5 MB estimated peak memory. Every reported operator, including the VTCM chunks, executed on the NPU. - A monolithic context is intentionally not included. Splitting the DiT into six VTCM-safe contexts avoids the cloud linker's memory explosion and allows the application to release each session before opening the next one. Performance, memory use, and image quality depend on the QNN SDK version, device firmware, thermal state, tiling implementation, and host-side orchestration. Test on the exact target device before shipping a production application. ## Intended use This package is intended for research and mobile application development involving: - photo and screenshot restoration; - text-rich image enhancement; - single-image super-resolution; - Snapdragon NPU deployment experiments. It is not intended for forensic reconstruction or for recovering information that is not present in the source image. Generative restoration can introduce plausible but incorrect details. ## Attribution This is a quantized mobile derivative of VOSR. The original architecture, training, and checkpoints were created by the VOSR authors. See the [official project](https://github.com/cswry/VOSR) and [paper](https://arxiv.org/abs/2604.03225) for full details. ## Citation If you use this model, please cite the original VOSR work: ```bibtex @inproceedings{wu2026vosr, title = {VOSR: A Vision-Only Generative Model for Image Super-Resolution}, author = {Wu, Rongyuan and Sun, Lingchen and Zhang, Zhengqiang and Kong, Xiangtao and Zhao, Jixin and Wang, Shihao and Zhang, Lei}, booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition}, year = {2026} } ``` ## License Released under the Apache License 2.0, following the upstream VOSR repository. Users are responsible for reviewing and complying with the licenses and terms of all upstream components and the Qualcomm QNN SDK/runtime used for deployment.