bert_simpson / forgebench /code /baselines /AMODAL3R_FIXES.md
Ronaldo-GOAT's picture
Amodal3R wrapper fixes: occlusion-aware FORGE3DBench mask in batch_amodal3r.py
e546cc5 verified
|
Raw History Blame Contribute Delete
3.39 kB

Amodal3R wrapper fixes (2026-09-30)

1. FORGE3DBench occlusion mask (this directory, batch_amodal3r.py) - REPLACED IN PLACE

What was wrong. load_rgb_and_mask built the 3-value mask from the input crop's alpha only (alpha>0.5 -> 188 visible, else 255 background), so no pixel was ever marked occluded (0), even on FORGE3DBench, which ships real per-view amodal masks. That contradicts the documented Amodal3R mask contract (255 = background, 188 = visible, 0 = occluded part of the object).

What changed (mask construction only; crop, 518 resize, RGB, seed, views, sampler, to_glb unchanged):

  • new --mask-mode {auto (default), occ2, occ, prod}:
    • occ2: L = 188 on visible, 0 on (amodal & ~visible) inside the crop, 255 elsewhere. Amodal from <src_scene>/mask_amodal/cam<ci:02d>/<obj_id>.png, crop bbox from <exp>/renders/<obj>_<tag>.npz, src_scene/obj_id/views4 from selection.json. Occlusion is computed at native crop resolution and resampled with the same LANCZOS/>0.5 as the alpha.
    • auto: occ2 when the selection entry has src_scene+obj_id (FORGE3DBench), otherwise prod (clean toys4k/omni3d renders have no occlusion, so alpha-only is already correct there).
    • occ: v1 ablation (NEAREST-resized amodal, adds a spurious ~1 px ring). prod: old behaviour.
  • The mask functions are verbatim copies of the validated occmask.py used for the rerun; on 272 FORGE3DBench views (225 with occluded pixels) they give byte-identical masks to it.

Effect on scores: none measurable. 808-obj (FB150 + rest658) rerun: W-LPIPS .1776 -> .1780, CD-L2 3.50 -> 3.64 (x1e-3), F@0.01 .342 -> .337; all paired 95% CIs include 0 and the deltas are smaller than seed-42-vs-43 noise. Paper numbers are kept.

Full detail (build machine): /home/nvidia/jonghoon/mv-mesh/.debug/amodal3r_fb_maskfix/RESULTS.md, audit: /home/nvidia/jonghoon/mv-mesh/.debug/amodal3r_occlusion_audit/RESULTS.md.

2. HO3D / DexYCB pose-eval framing (pose-eval code lives in the private forge3d_release repo under pose_eval/any6d_mesh_eval/)

What was wrong. code/build_recon.py:151-159 and code/dexycb/build_recon.py:152-160 (baseline_glb_mv, amodal3r branch) pass the full uncropped 640x480 camera frame (+ its amodal_L mask) to Amodal3R. Amodal3R does no cropping itself, so a small off-centre object is squashed into 518^2. (The single-view baseline_glb branch, build_recon.py:107-109, has the same pattern; only MV was rerun.)

Fix. code/amodal3r_crop/rerun_amodal3r_crop.py reuses build_recon.py's own functions (same prep_inputs_mv views, first min(V,4), same masks, run_infer.sh, seed 42, sampler, simplify, anchor + query scripts) and only re-frames each (color, amodal_L) pair with the official HF-Space demo geometry (demo_crop from amodal3r_occlusion_audit/rerun_variants.py: visible-bbox, long side = 0.68*512, centred in 512^2). build_recon.py is kept unedited for provenance; its amodal3r branch is superseded.

Effect: large. Pose ADD/ADD-S/AR: HO3D 8.1/28.7/10.2 -> 58.7/97.3/56.4, DexYCB 41.2/62.0/41.1 -> 80.0/93.2/79.8. Mesh CD-L1 mm / F@1cm: HO3D 13.58/.593 -> 2.87/.985, DexYCB 14.58/.390 -> 3.00/.963.

Full detail (build machine): /home/nvidia/jonghoon/mv-mesh/.debug/amodal3r_fixed_rerun/RESULTS.md, audit: /home/nvidia/jonghoon/mv-mesh/.debug/amodal3r_occlusion_audit/RESULTS.md.