GRADE models and evaluation (MobiCom 2026)

This repository contains the complete GRADE model, paper ablations, and baselines: 19 model checkpoint files plus one auxiliary TAESD weight file (20 .safetensors files in total). It also includes the inference and metric code and the saved camera-ready quantitative results. The full GRADE model uses the radar, diffusion, and ControlNet checkpoints in checkpoints/grade/.

Two evaluation paths are supported:

  • E1 β€” saved-result reproduction: regenerate the quantitative tables and figures from the supplied camera-ready CSVs and auxiliary evaluation data. This requires no inference or dataset download.
  • E2 β€” local reproduction: run inference on the released evaluation data, compute 2D/3D metrics, then regenerate the quantitative tables and figures from those newly computed CSVs.

Figure 10 is qualitative-only and is intentionally outside this artifact's quantitative workflow.

See artifact review status for the checks completed and the archive and data-clearance items still awaiting author or account records.

Layout

grade-models/
β”œβ”€β”€ README.md                         # this guide
β”œβ”€β”€ environment.txt                   # Python 3.11 package specification
β”œβ”€β”€ checkpoints/                      # included model and auxiliary weights
β”œβ”€β”€ evaluation_dataset/               # download the separate Smoke-Eval data here
β”œβ”€β”€ src/                              # released model implementations and configs
β”‚   β”œβ”€β”€ Baselines/                    # DA3, GRT, CaFNet, RadarCam-Depth
β”‚   β”œβ”€β”€ Ablation/                     # ablation implementations
β”‚   β”œβ”€β”€ GRADE/                        # Stage 1 radar-depth and Stage 2 diffusion modules
β”‚   └── models/                       # per-model release-relative configurations
└── evaluation/
    β”œβ”€β”€ run_inference.py              # unified inference runner
    β”œβ”€β”€ run_metrics.py                # 2D, 3D, merged-CSV, and radar-robustness runner
    β”œβ”€β”€ reproduce_paper.py            # quantitative paper reproduction
    β”œβ”€β”€ utils/                         # internal metric-computation implementation
    β”œβ”€β”€ outputs/inference/            # newly generated predictions
    β”œβ”€β”€ metric_results/               # newly computed metric CSVs
    β”œβ”€β”€ reference_results/            # supplied camera-ready CSVs and paper assets
    └── reproduced_results/           # newly generated tables and figures

evaluation/reference_results/ is read-only golden material. The runners write only to evaluation/outputs/, evaluation/metric_results/, and evaluation/reproduced_results/. The package does not include saved inference .npy files.

Model-name mapping

Pass the following identifiers to --model. grade is the preferred name for the complete model; ours_full remains a compatibility alias and historical CSV name.

Code identifier Paper name
grade / ours_full GRADE
ours_radar Ours_radar
ours_diffusion Ours_diffusion
ours_radar_no_grad Ours_radar w/o gradient loss
ours_radar_no_doppler Ours_radar w/o Doppler
ours_full_no_3d GRADE w/o 3D losses
grt_refine_freeze GRT refinement (frozen)
grt_refine_retrain GRT refinement (retrained)
da3 DA3
grt GRT
grt_image GRT+Image
grt_no_doppler GRT w/o Doppler
cafnet CaFNet
cafnet_no_smoke CaFNet (No-Smoke)
radarcam-depth RadarCam-Depth

Download the models and evaluation data

Download this repository and the separate processed Smoke-Eval evaluation dataset:

hf download phi-lab-rice/GRADE --local-dir grade-models
cd grade-models
hf auth login  # recommended for the many small evaluation files; use the local secure prompt
hf download mypersonalsharingspot11/evaluation_dataset \
  --repo-type dataset --local-dir evaluation_dataset
python -m zipfile -e \
  evaluation_dataset/Smoke-Eval-RadarCam-Depth.zip evaluation_dataset

The model weights and saved E1 results are already in this repository. The evaluation data are public but contain about 27,000 small files, so anonymous downloads may be rate limited. You can also provide HF_TOKEN through your shell's secure credential mechanism. All models except radarcam-depth use evaluation_dataset/Smoke-Eval/; that baseline uses evaluation_dataset/Smoke-Eval-RadarCam-Depth/. If ZIP extraction creates an extra nested directory, move the inner directory up one level.

The raw GRADE dataset is a separate release. Its public version currently lacks the videos needed to create processed DJI RGB input; the processed Smoke-Eval repository linked above supplies the evaluation input used here.

Prepare the environment

Use Python 3.11 and install environment.txt. E1 can run on CPU. E2 requires an NVIDIA GPU and a CUDA-compatible PyTorch installation. For a CUDA 12.6 host, one setup is:

python3.11 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install torch==2.7.0 torchvision==0.22.0 \
  --index-url https://download.pytorch.org/whl/cu126
python -m pip install -r environment.txt

Choose the PyTorch wheel matching your installed CUDA driver when needed.

The first 2D metric run may download the approximately 233 MB TorchVision AlexNet weights used by LPIPS from download.pytorch.org. To run metrics offline, perform one LPIPS metric run online in the same environment first and keep the TorchVision cache. Set TORCH_HOME to choose a persistent cache location.

RadarCam-Depth also fetches its EfficientNet-Lite3 architecture from the upstream Torch Hub repository on first use. Run that baseline once online to populate the Torch Hub cache before using it offline.

E1: reproduce the saved camera-ready results

Run:

python evaluation/reproduce_paper.py --mode saved

The output is written to evaluation/reproduced_results/saved/. It contains Tables 2–7 and the panels for Figures 9 and 11–14. As quick checks, Table 2 reports GRADE clear/smoke MAE values of 0.303/0.313, Table 3 reports 0.295/0.304 for light/heavy smoke, Table 4 reports 0.308/0.320/0.360 for full/radar/diffusion, and Table 7 identifies eight DDIM steps as best.

E2: reproduce results from local inference

For each desired model, run the inference, metric, and paper-reproduction runners in that order:

python evaluation/run_inference.py --model grade --gpuid 0
python evaluation/run_metrics.py --model grade --workers 1
python evaluation/reproduce_paper.py --mode local

The inference runner launches through Accelerate by default and accepts direct accelerate launch invocation as well. Repeat the first two commands for every model needed by the desired table or figure before the final command. GRADE diffusion variants use the released fixed DDIM seed (42); each frame receives a reproducible, newly drawn noise sample from that seeded generator.

To reproduce the radar-robustness analysis after generating the prerequisite predictions and metric outputs, run:

python evaluation/run_metrics.py --radar-robustness

--mode local writes to evaluation/reproduced_results/local/. A complete GPU-based E2 run over all models is substantially more expensive than E1; provision roughly 400 GiB of free storage and expect the full workflow to take longer than a day on typical local hardware.

Local mode includes only available model rows in Tables 2–6 and reports skipped outputs when none of their required model CSVs exist. Table 7 labels its input source in the generated report. Unless evaluation/metric_results/sampling_step/sampling_step_ablation_pooled.csv has been separately regenerated, Table 7 is derived from the supplied reference results even in local mode. The ordinary E2 model commands do not recompute the sampling-step ablation.

Useful options and checks

  • All three runners provide --dry-run to validate argument dispatch without running a model.
  • run_inference.py refuses to overwrite existing predictions. Use a different output root when retaining an earlier run.
  • run_metrics.py supports --sequence, --target-resolution, and --workers for smaller functional checks. Use the default --workers 1: higher values have deadlocked on an evaluation host and are not validated for this release.
  • reproduce_paper.py supports --tables-only and --figures-only.

Changing the data subset, output resolution, sampling settings, or evaluation configuration is useful for a smoke test but will not reproduce the camera-ready numerical values exactly.

Run only the full GRADE model on your own frames

For the full GRADE model, config.yaml also supports direct inference from synchronized radar and DJI RGB arrays without ZED ground-truth depth. Put each sequence under inputs/<sequence>/ with radar.npy (complex64, shape (N, 64, 2, 8, 256)) and dji_rgb.npy (uint8 RGB, shape (N, 504, 896, 3)). The arrays must have the same frame count and order. The GRADE camera calibration is embedded in preprocessing.

accelerate launch --num_processes 1 --num_machines 1 \
  --mixed_precision fp16 --dynamo_backend no \
  src/GRADE/stage2_diffusion_refinement/inference_full.py \
  --config config.yaml

The output is outputs/ours_full/<sequence>_pred.npy, a float32 array of shape (N, 1, 288, 512) normalized to [0, 1]. Multiply by 11.2 for estimated depth in meters. Set data.smoke_eval_root and inference.output_root in config.yaml to use other paths. The script overwrites same-named outputs on rerun.

Checkpoint integrity and citation

SHA256SUMS lists the hashes of all 20 .safetensors files. Run sha256sum -c SHA256SUMS from the repository root after download. The checkpoints contain inference tensors, not optimizer or training state. TAESD is an auxiliary VAE from the TAESD project, whose upstream model is MIT licensed.

@inproceedings{zhao2026grade,
  author    = {Bin Zhao and Patrick Chiou and Nakul Garg},
  title     = {{GRADE}: Single-Frame Generative Radar Depth Estimation Under Visual Degradation},
  booktitle = {Proceedings of the 32nd Annual International Conference on Mobile Computing and Networking},
  series    = {MobiCom '26},
  year      = {2026},
  publisher = {ACM},
  doi       = {10.1145/3795866.3844478}
}
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Dataset used to train phi-lab-rice/GRADE