|
Download forgebench/code/baselines/AMODAL3R_FIXES.md from Ronaldo-GOAT/bert_simpson: direct link, hf CLI and curl.
- Browser
- Download file 3.39 kB
-
https://huggingface.co/Ronaldo-GOAT/bert_simpson/resolve/main/forgebench/code/baselines/AMODAL3R_FIXES.md
- Command line
-
hf download hf://Ronaldo-GOAT/bert_simpson/forgebench/code/baselines/AMODAL3R_FIXES.md
-
curl -L -o AMODAL3R_FIXES.md https://huggingface.co/Ronaldo-GOAT/bert_simpson/resolve/main/forgebench/code/baselines/AMODAL3R_FIXES.md
3.39 kB
| # Amodal3R wrapper fixes (2026-09-30) | |
| ## 1. FORGE3DBench occlusion mask (this directory, `batch_amodal3r.py`) - REPLACED IN PLACE | |
| **What was wrong.** `load_rgb_and_mask` built the 3-value mask from the input crop's alpha only | |
| (alpha>0.5 -> 188 visible, else 255 background), so **no pixel was ever marked occluded (0)**, even on | |
| FORGE3DBench, which ships real per-view amodal masks. That contradicts the documented Amodal3R | |
| mask contract (255 = background, 188 = visible, 0 = occluded part of the object). | |
| **What changed** (mask construction only; crop, 518 resize, RGB, seed, views, sampler, to_glb unchanged): | |
| - new `--mask-mode {auto (default), occ2, occ, prod}`: | |
| - `occ2`: L = 188 on visible, 0 on (amodal & ~visible) inside the crop, 255 elsewhere. Amodal from | |
| `<src_scene>/mask_amodal/cam<ci:02d>/<obj_id>.png`, crop bbox from `<exp>/renders/<obj>_<tag>.npz`, | |
| `src_scene`/`obj_id`/`views4` from `selection.json`. Occlusion is computed at native crop resolution | |
| and resampled with the same LANCZOS/>0.5 as the alpha. | |
| - `auto`: `occ2` when the selection entry has `src_scene`+`obj_id` (FORGE3DBench), otherwise `prod` | |
| (clean toys4k/omni3d renders have no occlusion, so alpha-only is already correct there). | |
| - `occ`: v1 ablation (NEAREST-resized amodal, adds a spurious ~1 px ring). `prod`: old behaviour. | |
| - The mask functions are verbatim copies of the validated `occmask.py` used for the rerun; on | |
| 272 FORGE3DBench views (225 with occluded pixels) they give byte-identical masks to it. | |
| **Effect on scores: none measurable.** 808-obj (FB150 + rest658) rerun: W-LPIPS .1776 -> .1780, | |
| CD-L2 3.50 -> 3.64 (x1e-3), F@0.01 .342 -> .337; all paired 95% CIs include 0 and the deltas are smaller | |
| than seed-42-vs-43 noise. Paper numbers are kept. | |
| Full detail (build machine): `/home/nvidia/jonghoon/mv-mesh/.debug/amodal3r_fb_maskfix/RESULTS.md`, | |
| audit: `/home/nvidia/jonghoon/mv-mesh/.debug/amodal3r_occlusion_audit/RESULTS.md`. | |
| ## 2. HO3D / DexYCB pose-eval framing (pose-eval code lives in the private `forge3d_release` repo under `pose_eval/any6d_mesh_eval/`) | |
| **What was wrong.** `code/build_recon.py:151-159` and `code/dexycb/build_recon.py:152-160` | |
| (`baseline_glb_mv`, `amodal3r` branch) pass the full uncropped 640x480 camera frame (+ its amodal_L mask) | |
| to Amodal3R. Amodal3R does no cropping itself, so a small off-centre object is squashed into 518^2. | |
| (The single-view `baseline_glb` branch, `build_recon.py:107-109`, has the same pattern; only MV was rerun.) | |
| **Fix.** `code/amodal3r_crop/rerun_amodal3r_crop.py` reuses build_recon.py's own functions (same | |
| `prep_inputs_mv` views, first min(V,4), same masks, run_infer.sh, seed 42, sampler, simplify, anchor + query | |
| scripts) and only re-frames each (color, amodal_L) pair with the official HF-Space demo geometry | |
| (`demo_crop` from `amodal3r_occlusion_audit/rerun_variants.py`: visible-bbox, long side = 0.68*512, | |
| centred in 512^2). `build_recon.py` is kept unedited for provenance; its amodal3r branch is superseded. | |
| **Effect: large.** Pose ADD/ADD-S/AR: HO3D 8.1/28.7/10.2 -> 58.7/97.3/56.4, DexYCB 41.2/62.0/41.1 -> | |
| 80.0/93.2/79.8. Mesh CD-L1 mm / F@1cm: HO3D 13.58/.593 -> 2.87/.985, DexYCB 14.58/.390 -> 3.00/.963. | |
| Full detail (build machine): `/home/nvidia/jonghoon/mv-mesh/.debug/amodal3r_fixed_rerun/RESULTS.md`, | |
| audit: `/home/nvidia/jonghoon/mv-mesh/.debug/amodal3r_occlusion_audit/RESULTS.md`. | |