--- license: cc-by-nc-4.0 library_name: libreyolo pipeline_tag: any-to-any base_model: sensenova/SenseNova-Vision-7B-MoT tags: - libreyolo - unified-multimodal - object-detection - keypoint-detection - depth-estimation - image-segmentation - ocr - arxiv:2607.06560 --- # SenseNovaVision7b > **Non-commercial weights.** These model weights are released by SenseTime > under **CC BY-NC 4.0** and may be used for non-commercial purposes only. > The LibreYOLO integration code is MIT; the vendored architecture code is > Apache-2.0. The license of the weights does not change by being mirrored > here. SenseNova-Vision-7B-MoT, mirrored for LibreYOLO's `LibreVLM` tier. A unified multimodal model (Bagel-MoT architecture: Qwen2.5-7B MoT decoder, SigLIP vision tower, FLUX autoencoder) that serves many vision tasks from one checkpoint: symbolic outputs (boxes, points, keypoints, OCR words) are generated as tagged text, dense outputs (depth maps, segmentation masks, panoptic maps) are generated as images decoded by the VAE. ## Source Byte-identical mirror of [sensenova/SenseNova-Vision-7B-MoT](https://huggingface.co/sensenova/SenseNova-Vision-7B-MoT) at revision `79548fcc5b954598799b9317f8d3ec5e347d5c0e`. Copyright (c) 2026 SenseTime Group Inc. and/or its affiliates. Reference implementation: [OpenSenseNova/SenseNova-Vision](https://github.com/OpenSenseNova/SenseNova-Vision) (commit `12ccd96e32b32967a11cacb6c5bd5fe3a555fc0c`, Apache-2.0). Paper: [Vision as Unified Multimodal Generation](https://arxiv.org/abs/2607.06560). Weight checksums (SHA-256): ``` ema.safetensors 96f29abd98791288c5a24087322e964bb9bcabfc2f185ece71543f827bc2b11e ae.safetensors afc8e28272cd15db3919bacdb6918ce9c1ed22e96cb12c4d5ed0fba823529e38 ``` ## Modifications None. The weights and configuration files are unmodified upstream bytes; only this card and the NOTICE file are added. One upstream quirk to be aware of: `tokenizer.json` assigns the chat/vision special tokens ids beyond the checkpoint's embedding table, while `tokenizer_config.json`'s `added_tokens_decoder` records the layout the model was trained with (structured tokens overriding ids 149632-151664). LibreYOLO reconstructs the trained layout at load time; other consumers should do the same or use a legacy slow tokenizer built from `vocab.json`/`merges.txt`. ## Usage ```python from libreyolo import LibreVLM model = LibreVLM("sensenova-vision", task="detect") model.set_classes(["person", "bicycle"]) results = model.predict("image.jpg") # Results.boxes model.set_task("depth") results = model.predict("image.jpg") # Results.depth_map model.set_task("segment").set_classes(["the person on the left"]) results = model.predict("image.jpg") # Results.masks ``` Supported LibreYOLO tasks: `detect`, `point`, `pose`, `ocr`, `depth`, `segment` (referring), `panoptic`. Free-form access via `model.chat(...)` and `model.generate(...)`. ## License Composition - Model weights: **CC BY-NC 4.0** (non-commercial), SenseTime - Reference implementation: Apache-2.0 (SenseTime, building on ByteDance's Bagel, Hugging Face transformers, and the Black Forest Labs FLUX autoencoder, all Apache-2.0) - LibreYOLO integration: MIT See [`LICENSE`](./LICENSE) (upstream license statement, verbatim) and [`NOTICE`](./NOTICE) for attribution details.