Companion Forge v7 β Distilled 3D Model Stack
Target: Hugging Face Jobs l4x1 / NVIDIA L4 (SM89, 24 GB VRAM).
v7 keeps the validated v6.4 ONNX/TensorRT runtime as the production fallback and adds a training/optimization layer inspired by the Kimi-K3 work: progressive few-step distillation, selective FP8, 2:4 structured sparsity, DINOv3 conditioning, multi-view fusion, geometry teachers, symmetry guidance, geometry-aware MoE, and timestep caching.
Pipeline
image/text
-> DINOv3 bridge + optional multi-view fusion
-> SS flow student (10 -> 4 -> 2 -> 1 steps)
-> SLat flow student (10 -> 4 -> 2 -> 1 steps)
-> full custom ONNX/TensorRT SLat DAE
-> MeshFlow/TRELLIS teacher losses during training only
-> AniGen skeleton/skin teacher
-> symmetry + rig guidance
-> rigged GLB
What is implemented
training/distill_core.py: progressive Euler macro-step KD + endpoint consistency + MeanFlow-style interval consistency.training/distill_flow.py: L4-friendly LoRA-then-merge training for SS/SLat flow students, stages10β4β2β1.training/fp8_sparse.py: selective ModelOpt FP8 hooks and exact magnitude-based 2:4 pruning for eligible Linear weights.training/dinov3_bridge.py+distill_dinov3.py: DINOv3 ViT-S bridge to AniGen's fixed1374Γ1024conditioning contract, distilled against DINOv2 ViT-L/14-reg.runtime/multiview.py: pose-Fourier multi-view token fusion for front/left/back/right conditioning.runtime/symmetry.py: flow-time reflection/C2/C4 velocity symmetrization plus rig symmetry loss.training/teacher_hybrid.py: cached MeshFlow/TRELLIS geometry teacher + AniGen rig teacher interface.training/geometry_moe.py: ModernMOE-inspired shared + top-k routed 3D experts.runtime/timestep_cache.py: residual timestep cache for >=4-step students; automatically unnecessary after 1β2-step distillation.bench/quality_eval.py: geometry/symmetry/joint/speed comparison harness.
L4 training order
- Distill DINOv3 bridge (keep current DINOv2 TRT runtime as fallback).
- Build teacher cache using v6.4 outputs plus optional MeshFlow/TRELLIS geometry targets.
- Distill SS
10β4, SLat10β4; validate. - Apply selective FP8 PTQ; if quality drops, run short QAT.
- Apply 2:4 pruning to MLP/projection weights, freeze masks, recovery distillation.
- Distill
4β2, then2β1only if geometry/rig metrics stay within thresholds. - Train geometry-inductive MoE as an optional higher-capacity student; distill it back to a dense/sparse deployment student if its routing overhead is not worthwhile on L4.
- Re-export ONNX external-data and compile SM89 TensorRT engines. Output heads/mesh remain FP32; sparse topology ops remain the validated custom plugins.
Safety/quality gates
A new student is not promoted merely because it is faster. Promotion requires finite tensors, exit code 0, GLB skin/skeleton validation, no lazy PyTorch fallback, and quality thresholds against the current v6.4 teacher. 1+1 is therefore an experimental final stage, not an automatic default.
External technology mapping
- NVIDIA FastGen concepts: progressive KD / MeanFlow / consistency-style few-step training.
- NVIDIA Model Optimizer: FP8 PTQ/QAT and export path.
- TensorRT: 2:4 structured sparse tactics where eligible.
- Meta DINOv3: stronger dense visual conditioning.
- Meta MeshFlow / Microsoft TRELLIS.2: geometry/topology teacher targets only; licensing/runtime isolation is preserved.
- SymTRELLIS idea: velocity symmetrization adapted to AniGen flow and rig constraints.
- ModernMOE idea: routed/shared lightweight experts specialized for geometry roles.
The current v6.4 runtime remains the rollback target until each v7 stage has its own L4 E2E validation artifact.