Spaces:
Running on Zero
Running on Zero
File size: 2,221 Bytes
d968973 83c5d6a d968973 83c5d6a d968973 83c5d6a | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 | ---
title: PatchAlign3D
emoji: 🧩
colorFrom: gray
colorTo: pink
sdk: gradio
sdk_version: 6.22.0
app_file: app.py
pinned: false
license: mit
short_description: Zero-shot 3D part segmentation from text queries
python_version: "3.12"
startup_duration_timeout: 1h
models:
- patchalign3d/patchalign3d-encoder
---
# PatchAlign3D · zero-shot 3D part segmentation
Name the parts you want **in words** and see them highlighted on a 3D shape. A single forward pass of a
point-cloud encoder whose *patch-level* features are aligned to CLIP text space — no test-time multi-view
rendering, no per-category training.
- Paper: https://huggingface.co/papers/2601.02457
- Code: https://github.com/souhail-hadgi/PatchAlign3D
- Weights: https://huggingface.co/patchalign3d/patchalign3d-encoder
## How it works
Points are centred and scaled to the unit sphere and the Y/Z axes are swapped to match the training
convention (exactly as in the official `infer.py`). Furthest-point sampling picks patch centres, a k-NN
neighbourhood around each becomes a patch token, and a 12-layer point transformer produces one feature per
patch. A learned linear head projects those into the CLIP `ViT-bigG-14 (laion2b_s39b_b160k)` text space,
where they are matched against a prompt ensemble. Each point takes the label of its nearest patch centre.
## Implementation notes
`patchalign3d.py` is a faithful port of the official inference path with three behaviour-preserving
deviations, so the Space runs without custom CUDA extensions:
1. `pointnet2_ops.furthest_point_sample` → pure-torch iterative FPS (same start index, squared distances,
argmax selection).
2. `knn_cuda.KNN` → `torch.cdist` + `topk`.
3. `open_clip ViT-bigG-14 / laion2b_s39b_b160k` text tower → the byte-identical (up to fp16 rounding)
`text_encoder_2` + `tokenizer_2` of `stabilityai/stable-diffusion-xl-base-1.0`, which avoids a 10 GB
download of the full image+text CLIP checkpoint.
Verified against ShapeNetPart ground truth on CPU before deployment: the checkpoint loads with 0 missing
and 0 unexpected keys, and per-shape zero-shot mIoU averages **0.59** with the default `part_only` prompts
(**0.64** with `ensemble`) across 8 held-out test shapes.
|