| --- |
| license: other |
| tags: |
| - multimodal |
| - qwen3-vl |
| - covt |
| - segmentation |
| - depth |
| - training-bundle |
| --- |
| |
| # UMM Stage-2 portable training bundle |
|
|
| This public repository is a self-contained handoff for retraining UMM Stage 2: Qwen3-VL CoVT curriculum with SAM, DINOv2, VGGT, PiDiNet and SigLIP frozen-teacher supervision. |
|
|
| The deployed visual-token counts are `[8,4,4,4,4]`; VGGT/depth uses four tokens, aligned with CoVT. |
|
|
| ## Download and deploy |
|
|
| ```bash |
| unset http_proxy https_proxy HTTP_PROXY HTTPS_PROXY all_proxy ALL_PROXY |
| hf download Orangerl/umm --local-dir umm |
| cd umm |
| bash scripts/bootstrap.sh |
| DRY_RUN=1 bash code/umm/stage2_covt/run_train.sh |
| ``` |
|
|
| `bootstrap.sh` extracts the relocatable `unvideo` environment, reinstalls the exact bundled editable sources, materializes 756,894 images from tar shards, renders machine-local paths and runs the Stage-2 validation suite. |
|
|
| Reserve at least 650 GB of free disk space: the downloaded image shards and |
| their extracted images coexist after deployment. The target should be Linux |
| x86_64 with an NVIDIA driver compatible with CUDA 12.8. |
| |
| For Code Agent handoff, use this prompt from the downloaded repository: |
| |
| ```text |
| Read skills/deploy-umm-stage2/SKILL.md completely and use it to deploy and |
| validate this repository. Do not start formal training without my explicit |
| instruction. |
| ``` |
| |
| Provide `SWANLAB_API_KEY` or `secrets/swanlab_key.txt` before real training. Formal training is never started automatically. |
| The launcher defaults to GPUs `0,1,2,3`; override `CUDA_VISIBLE_DEVICES` for a |
| different topology. |
|
|
| After a completed run, merge the Stage-2 LoRA/non-LoRA output with: |
|
|
| ```bash |
| MODEL_PATH="$PWD/outputs/stage2_covt_sam_vggt_covt4_v2" \ |
| bash code/umm/stage2_covt/run_merge.sh |
| ``` |
|
|
| ## Included artifacts |
|
|
| - portable UMM code and DeepSpeed config; |
| - exact customized Transformers and Diffusers sources; |
| - Qwen3-VL-4B-Instruct base model; |
| - Stage-1 `checkpoint-32000.bin` connector/embedding initialization; |
| - SAM ViT-H, VGGT-1B, DINOv2 ViT-L/14, PiDiNet and SigLIP2-Large teacher assets; |
| - exact local CoVT dataset: 889,489 conversations and 756,894 images; |
| - relocatable packed `unvideo` environment; |
| - `$deploy-umm-stage2` Code Agent skill; |
| - SHA256 integrity manifest and provenance metadata. |
|
|
| FLUX weights and the 560 GB Stage-1 DeepSpeed optimizer history are not included because the active Stage-2 code does not load them. Stage 2 consumes only the included 1.06 GB Stage-1 connector file. |
|
|
| ## Licensing |
|
|
| No unified relicensing is asserted. Each bundled model, dataset subset and source tree remains subject to its upstream license. In particular, the original `facebook/VGGT-1B` checkpoint has non-commercial restrictions; replace it with the commercial checkpoint before commercial use. |
|
|