File size: 2,751 Bytes
41d0afb 70e1bae 41d0afb 70e1bae 41d0afb | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 | ---
license: other
tags:
- multimodal
- qwen3-vl
- covt
- segmentation
- depth
- training-bundle
---
# UMM Stage-2 portable training bundle
This public repository is a self-contained handoff for retraining UMM Stage 2: Qwen3-VL CoVT curriculum with SAM, DINOv2, VGGT, PiDiNet and SigLIP frozen-teacher supervision.
The deployed visual-token counts are `[8,4,4,4,4]`; VGGT/depth uses four tokens, aligned with CoVT.
## Download and deploy
```bash
unset http_proxy https_proxy HTTP_PROXY HTTPS_PROXY all_proxy ALL_PROXY
hf download Orangerl/umm --local-dir umm
cd umm
bash scripts/bootstrap.sh
DRY_RUN=1 bash code/umm/stage2_covt/run_train.sh
```
`bootstrap.sh` extracts the relocatable `unvideo` environment, reinstalls the exact bundled editable sources, materializes 756,894 images from tar shards, renders machine-local paths and runs the Stage-2 validation suite.
Reserve at least 650 GB of free disk space: the downloaded image shards and
their extracted images coexist after deployment. The target should be Linux
x86_64 with an NVIDIA driver compatible with CUDA 12.8.
For Code Agent handoff, use this prompt from the downloaded repository:
```text
Read skills/deploy-umm-stage2/SKILL.md completely and use it to deploy and
validate this repository. Do not start formal training without my explicit
instruction.
```
Provide `SWANLAB_API_KEY` or `secrets/swanlab_key.txt` before real training. Formal training is never started automatically.
The launcher defaults to GPUs `0,1,2,3`; override `CUDA_VISIBLE_DEVICES` for a
different topology.
After a completed run, merge the Stage-2 LoRA/non-LoRA output with:
```bash
MODEL_PATH="$PWD/outputs/stage2_covt_sam_vggt_covt4_v2" \
bash code/umm/stage2_covt/run_merge.sh
```
## Included artifacts
- portable UMM code and DeepSpeed config;
- exact customized Transformers and Diffusers sources;
- Qwen3-VL-4B-Instruct base model;
- Stage-1 `checkpoint-32000.bin` connector/embedding initialization;
- SAM ViT-H, VGGT-1B, DINOv2 ViT-L/14, PiDiNet and SigLIP2-Large teacher assets;
- exact local CoVT dataset: 889,489 conversations and 756,894 images;
- relocatable packed `unvideo` environment;
- `$deploy-umm-stage2` Code Agent skill;
- SHA256 integrity manifest and provenance metadata.
FLUX weights and the 560 GB Stage-1 DeepSpeed optimizer history are not included because the active Stage-2 code does not load them. Stage 2 consumes only the included 1.06 GB Stage-1 connector file.
## Licensing
No unified relicensing is asserted. Each bundled model, dataset subset and source tree remains subject to its upstream license. In particular, the original `facebook/VGGT-1B` checkpoint has non-commercial restrictions; replace it with the commercial checkpoint before commercial use.
|