projseg-checkpoints / README.md
tianxia2's picture
Add card
0e39cc8 verified
|
Raw
History Blame Contribute Delete
1.87 kB
---
license: apache-2.0
tags:
- segmentation
- efficientsam
- aria-digital-twin
- reproducibility
---
# Checkpoints — Projection-Assisted Segmentation Reproduction
Model weights for reproducing projection-assisted gaze-prompted segmentation on
the Aria Digital Twin dataset.
Companion data: [`tianxia2/projseg-adt-seq144-subset`](https://huggingface.co/datasets/tianxia2/projseg-adt-seq144-subset).
## Files
| File | Size | What |
|---|---|---|
| `efficient_sam_vitt.pt` | 40 MB | **EfficientSAM ViT-Tiny.** The anchor segmenter, prompted with a single gaze point. Unmodified upstream weights from [yformer/EfficientSAM](https://github.com/yformer/EfficientSAM) (Apache-2.0), redistributed for convenience. |
| `refinenetgaze.pth` | 5.9 MB | **RefineNetGaze.** Optional U-Net that refines a reused mask from RGB + prior mask + gaze heatmap (5-channel input, base=16). Trained by us on ADT. |
## Usage
```bash
huggingface-cli download tianxia2/projseg-checkpoints --local-dir checkpoints
```
The reproduction artifact expects them at `checkpoints/` (override with
`PROJSEG_CKPT_ROOT`).
## Notes
- `refinenetgaze.pth` is stored as a **plain tensor state dict**, so it loads
under PyTorch ≥ 2.6's default `weights_only=True`. Construct the model as
`RefineNetGaze(base=16)`.
- Refinement is **optional and off by default**. On the reported EfficientSAM +
`dense_gt` configuration it *lowers* mIoU (0.3283 → 0.3111), because it was
trained against a different anchor-mask distribution than the one ESAM
anchors produce. It is shipped for completeness and ablation, not as part of
the headline result.
- EfficientSAM is class-agnostic. Prompted with one gaze point it segments
*some* coherent region around that point, which does not always coincide with
the annotated ADT instance — hence modest absolute mIoU on this benchmark.