DiVE β Learning Decomposed Visibility for Efficient Active Exploration of Cluttered Scenes
Checkpoints for DiVE, published at the Conference on Robot Learning (CoRL) 2026.
- Project page: https://lee-su-yun.github.io/DiVE/
- Code: https://github.com/lee-su-yun/DiVE (release in progress)
DiVE decomposes scene visibility into what a viewpoint change can reveal and what a push can reveal, and uses that decomposition to decide between looking and interacting before ranking candidate actions within the chosen mode.
Contents
| Folder | Paper notation | Role |
|---|---|---|
DiVE_decomposed_visibility_estimator/ |
f_DiVE |
Predicts the view-resolvable (phi_view) and push-resolvable (phi_push) visibility maps |
view_selection_policy/ |
pi_view |
Ranks the 98 candidate viewpoints |
push_selection_policy/ |
pi_push |
Ranks the candidate pushes |
belief_map_estimator/ |
f_scene |
Recursive semantic belief updater used as the central tracker |
The three DiVE folders each hold a single best.pth saved by the training script, with keys
epoch, model_state_dict, optimizer_state_dict, scheduler_state_dict,
scaler_state_dict and the validation metrics of that epoch. Load the weights with
torch.load(path, map_location="cpu")["model_state_dict"]. belief_map_estimator/estimator.pt
is a plain state_dict instead and loads directly.
The voxel grid is (D, H, W) = (60, 120, 80) throughout.
DiVE_decomposed_visibility_estimator β decomposed visibility estimator
- Model:
support_code/model_UAB_unified.pyβUNet3DPushUnified - Trainer:
2_belief_unified/train_UAB_unified.py - Architecture: 3D U-Net with a shared backbone and a 2-channel output head, 0.34M parameters
- Input:
(B, 3, 60, 120, 80)=[phi_view_t, phi_push_t, voxel_t] - Output:
(B, 2, 60, 120, 80)raw logits; applysigmoidfor visibility in[0, 1] - Trained for 200 epochs with teacher-forcing annealing; the released checkpoint is epoch 199 (self-feeding validation loss 3.92 β view head 2.48, push head 1.44)
- Original run name:
UAB_unified_ycb_v4_ep200_TFanneal
view_selection_policy β viewpoint selection policy
- Model:
support_code/model_reward.pyβRewardNet - Trainer:
4_policy/train_reward_UA.py - Architecture: 3D conv encoder collapsed onto the 7x14 camera grid, then a 2D conv head
(
base_ch=16,use_seen_mask=True), 1.02M parameters - Input: belief map
(B, 1, 60, 120, 80)plus a(B, 7, 14)binary mask of the viewpoints already visited - Output:
(B, 98)logits, row-major camera indexrow * 14 + col - Validation at the released epoch (47): top-1 0.830, top-3 0.957, P@10 0.904
- Original run name:
reward_UA_sigz_final_v3
python 4_policy/train_reward_UA.py \
--reward_roots /result/APOBU/reward_dataset/ua_reward_dataset_ycb_v3_unified \
--save_dir /result/DiVE/reward_UA_sigz_final_v3 \
--wandb_run_name reward_UA_sigz_final_v3 \
--loss_type bce_sigmoid_z \
--temperature 1.0 \
--use_seen_mask \
--device 1
push_selection_policy β push selection policy
- Model:
4_policy_test/model_reward_UB.pyβRewardNetUB(self-contained, nosupport_codedependency; not the 29-channel variant insupport_code/model_reward_UB.py) - Trainer:
4_policy_test/train_reward_UB.py - Architecture: siamese 3D CNN over
concat(belief, swept_map_i)with an overlap-residual score and a transformer block attending across candidate actions (base_ch=16,beta=0.5,use_attn=True, 4 heads), 1.43M parameters - Input: belief map
(B, 1, 60, 120, 80)and candidate swept maps(B, N, 60, 120, 80) - Output:
(B, N)scores over the candidate pushes - Validation at the released epoch (21): top-1 0.336, top-3 0.643, P@10 0.853
- This is the exact checkpoint used for the evaluation runs reported in the paper. A later epoch-28 checkpoint of the same run exists (marginally lower val loss, top-1 0.356 / top-3 0.629); the epoch-21 weights are released so the reported numbers reproduce
- Original run name:
reward_UB_full_siamese_sigz_b05_attn_final_v3
python 4_policy_test/train_reward_UB.py \
--reward_roots /result/APOBU/reward_dataset/ub_reward_dataset_ycb_v3_unified \
--data_roots /data/APOBU/beliefmap_high_occlusion_ycb_v3 \
/result/DiVE_data/beliefmap_low_occlusion_ycb_v3 \
--save_dir /result/APOBU/DiVE/reward_UB_full_siamese_sigz_b05_attn_final_v3 \
--wandb_run_name reward_UB_full_siamese_sigz_b05_attn_final_v3 \
--loss_type bce_sigmoid_z \
--temperature 1.0 \
--beta 0.5 \
--use_attn \
--attn_heads 4 \
--max_steps_per_epoch 3000 \
--max_val_steps 500 \
--batch_size 8 \
--epochs 50 \
--patience 6 \
--device 2,3
belief_map_estimator β recursive semantic belief updater
- Model:
BeliefUpdateNetRecursive(n_sem=15, base_ch=32)over a 3D U-Net backbone, 14.70M parameters - Input: 34 channels =
[alpha/50, beta/50, occ, swept_map, one_hot(sem_obs) x15, sem_belief x15] - Output: 17 channels =
[new_alpha, new_beta, sem_logits x15]; the softmaxed semantic logits and the Beta parameters are fed back at the next step (recursive) - YCB v3 uses classes 0-11 (12 active); slots 12-14 are reserved, hence
n_sem=12at evaluation withn_sem_model=15 - Stored as a plain
state_dict(58.8 MB):
import torch
from estimator import BeliefUpdateNetRecursive
net = BeliefUpdateNetRecursive(n_sem=15, base_ch=32)
net.load_state_dict(torch.load("belief_map_estimator/estimator.pt", map_location="cpu"))
- Trained from scratch with DDP on 4 GPUs (
--batch 32,--n_sem 15) onbeliefmap_{high,low}_occlusion_ycb_v3 - Reference evaluation, DiVE policy with this estimator as the central tracker, low-occlusion test set, 100 episodes x 25-step budget: mIoU 0.9279, occupancy IoU 0.8721, semantic mIoU 0.8116, 4.92 pushes per episode
Training data
Cluttered shelf scenes generated in NVIDIA Isaac Sim with YCB objects, at low and high occlusion. Extreme-occlusion scenes are held out for evaluation only. The dataset release is in preparation; see the project page.
Citation
@inproceedings{lee2026dive,
title = {Learning Decomposed Visibility for Efficient Active Exploration of Cluttered Scenes},
author = {Lee, Suyun and Choi, Minsoo and Gong, Jihwan and Nam, Unghui and Bae, Minji and Shim, Byonghyo},
booktitle = {Conference on Robot Learning (CoRL)},
year = {2026},
}