Reset torch default device to cpu after offload.profile(): fixes device mismatches in econdataset/pymaf/smplx code that assumes CPU-default tensor creation, without affecting mmgps explicit hook-based device management
Bump GPU duration to 600s: memory issue is fixed (attention slicing), task now just needs more time for the full multiview diffusion + mesh reconstruction pipeline
Implement attention slicing: chunk SDPA calls along the batch*heads dim since torch 2.8s fused kernels reject this shape on this GPU, forcing an O(seq_len^2 * batch) math fallback that OOMs even a 96GB card
Diagnose why SDPA auto-select avoids flash/efficient backends: try FLASH_ATTENTION explicitly and log the actual error, plus dtype/contiguity/mask info
Request ZeroGPU xlarge (full 96GB Blackwell card) instead of default large (48GB half-card): the multiview attention activation memory OOMs on the half-card even with mmgp weight offloading
Fix create_mean_pose() to build its numpy array from CPU tensors explicitly, instead of masking it with a global torch default-device reset that broke mmgp's own device handling
Bump diffusers to 0.29.0: mmgp (via optimum-quanto) needs PixArtTransformer2DModel, absent from the old 0.26.0 pin; verified all of PSHumans own diffusers imports (incl. LoRACompatibleConv/AdaLayerNorm) still resolve at 0.29.0
Drop xformers: its Hopper-specific flash-attention kernel crashes on Blackwell (sm_120) with 'CUDA error: invalid argument'. torch 2.8's native SDPA (used automatically by diffusers) replaces it.
Add nvidia-cuda-runtime + LD_LIBRARY_PATH shim so nvdiffrast's JIT-compiled CUDA extension can find libcudart.so.13 on Blackwell/CUDA-13 ZeroGPU workers
Pin torch/torchvision/torchaudio to 2.1.0+cu121: unpinned torch was drifting to latest PyPI (now 2.14.0/cu13x), breaking ABI with the hard-pinned kaolin/pytorch3d/torch_scatter cu121 wheels (libcudart.so.13 not found)