GRADE models and evaluation (MobiCom 2026)
This repository contains the complete GRADE model, paper ablations, and baselines: 19 model checkpoint files plus one auxiliary TAESD weight file (20 .safetensors files in total). It also includes the inference and metric code and the saved camera-ready quantitative results. The full GRADE model uses the radar, diffusion, and ControlNet checkpoints in checkpoints/grade/.
Two evaluation paths are supported:
- E1 β saved-result reproduction: regenerate the quantitative tables and figures from the supplied camera-ready CSVs and auxiliary evaluation data. This requires no inference or dataset download.
- E2 β local reproduction: run inference on the released evaluation data, compute 2D/3D metrics, then regenerate the quantitative tables and figures from those newly computed CSVs.
Figure 10 is qualitative-only and is intentionally outside this artifact's quantitative workflow.
See artifact review status for the checks completed and the archive and data-clearance items still awaiting author or account records.
Layout
grade-models/
βββ README.md # this guide
βββ environment.txt # Python 3.11 package specification
βββ checkpoints/ # included model and auxiliary weights
βββ evaluation_dataset/ # download the separate Smoke-Eval data here
βββ src/ # released model implementations and configs
β βββ Baselines/ # DA3, GRT, CaFNet, RadarCam-Depth
β βββ Ablation/ # ablation implementations
β βββ GRADE/ # Stage 1 radar-depth and Stage 2 diffusion modules
β βββ models/ # per-model release-relative configurations
βββ evaluation/
βββ run_inference.py # unified inference runner
βββ run_metrics.py # 2D, 3D, merged-CSV, and radar-robustness runner
βββ reproduce_paper.py # quantitative paper reproduction
βββ utils/ # internal metric-computation implementation
βββ outputs/inference/ # newly generated predictions
βββ metric_results/ # newly computed metric CSVs
βββ reference_results/ # supplied camera-ready CSVs and paper assets
βββ reproduced_results/ # newly generated tables and figures
evaluation/reference_results/ is read-only golden material. The runners write only to evaluation/outputs/, evaluation/metric_results/, and evaluation/reproduced_results/. The package does not include saved inference .npy files.
Model-name mapping
Pass the following identifiers to --model. grade is the preferred name for the complete model; ours_full remains a compatibility alias and historical CSV name.
| Code identifier | Paper name |
|---|---|
grade / ours_full |
GRADE |
ours_radar |
Ours_radar |
ours_diffusion |
Ours_diffusion |
ours_radar_no_grad |
Ours_radar w/o gradient loss |
ours_radar_no_doppler |
Ours_radar w/o Doppler |
ours_full_no_3d |
GRADE w/o 3D losses |
grt_refine_freeze |
GRT refinement (frozen) |
grt_refine_retrain |
GRT refinement (retrained) |
da3 |
DA3 |
grt |
GRT |
grt_image |
GRT+Image |
grt_no_doppler |
GRT w/o Doppler |
cafnet |
CaFNet |
cafnet_no_smoke |
CaFNet (No-Smoke) |
radarcam-depth |
RadarCam-Depth |
Download the models and evaluation data
Download this repository and the separate processed Smoke-Eval evaluation dataset:
hf download phi-lab-rice/GRADE --local-dir grade-models
cd grade-models
hf auth login # recommended for the many small evaluation files; use the local secure prompt
hf download mypersonalsharingspot11/evaluation_dataset \
--repo-type dataset --local-dir evaluation_dataset
python -m zipfile -e \
evaluation_dataset/Smoke-Eval-RadarCam-Depth.zip evaluation_dataset
The model weights and saved E1 results are already in this repository. The evaluation data are public but contain about 27,000 small files, so anonymous downloads may be rate limited. You can also provide HF_TOKEN through your shell's secure credential mechanism. All models except radarcam-depth use evaluation_dataset/Smoke-Eval/; that baseline uses evaluation_dataset/Smoke-Eval-RadarCam-Depth/. If ZIP extraction creates an extra nested directory, move the inner directory up one level.
The raw GRADE dataset is a separate release. Its public version currently lacks the videos needed to create processed DJI RGB input; the processed Smoke-Eval repository linked above supplies the evaluation input used here.
Prepare the environment
Use Python 3.11 and install environment.txt. E1 can run on CPU. E2 requires an NVIDIA GPU and a CUDA-compatible PyTorch installation. For a CUDA 12.6 host, one setup is:
python3.11 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install torch==2.7.0 torchvision==0.22.0 \
--index-url https://download.pytorch.org/whl/cu126
python -m pip install -r environment.txt
Choose the PyTorch wheel matching your installed CUDA driver when needed.
The first 2D metric run may download the approximately 233 MB TorchVision
AlexNet weights used by LPIPS from download.pytorch.org. To run metrics
offline, perform one LPIPS metric run online in the same environment first and
keep the TorchVision cache. Set TORCH_HOME to choose a persistent cache
location.
RadarCam-Depth also fetches its EfficientNet-Lite3 architecture from the upstream Torch Hub repository on first use. Run that baseline once online to populate the Torch Hub cache before using it offline.
E1: reproduce the saved camera-ready results
Run:
python evaluation/reproduce_paper.py --mode saved
The output is written to evaluation/reproduced_results/saved/. It contains Tables 2β7 and the panels for Figures 9 and 11β14. As quick checks, Table 2 reports GRADE clear/smoke MAE values of 0.303/0.313, Table 3 reports 0.295/0.304 for light/heavy smoke, Table 4 reports 0.308/0.320/0.360 for full/radar/diffusion, and Table 7 identifies eight DDIM steps as best.
E2: reproduce results from local inference
For each desired model, run the inference, metric, and paper-reproduction runners in that order:
python evaluation/run_inference.py --model grade --gpuid 0
python evaluation/run_metrics.py --model grade --workers 1
python evaluation/reproduce_paper.py --mode local
The inference runner launches through Accelerate by default and accepts direct accelerate launch invocation as well. Repeat the first two commands for every model needed by the desired table or figure before the final command. GRADE diffusion variants use the released fixed DDIM seed (42); each frame receives a reproducible, newly drawn noise sample from that seeded generator.
To reproduce the radar-robustness analysis after generating the prerequisite predictions and metric outputs, run:
python evaluation/run_metrics.py --radar-robustness
--mode local writes to evaluation/reproduced_results/local/. A complete GPU-based E2 run over all models is substantially more expensive than E1; provision roughly 400 GiB of free storage and expect the full workflow to take longer than a day on typical local hardware.
Local mode includes only available model rows in Tables 2β6 and reports
skipped outputs when none of their required model CSVs exist. Table 7 labels
its input source in the generated report. Unless
evaluation/metric_results/sampling_step/sampling_step_ablation_pooled.csv
has been separately regenerated, Table 7 is derived from the supplied
reference results even in local mode. The ordinary E2 model commands do not
recompute the sampling-step ablation.
Useful options and checks
- All three runners provide
--dry-runto validate argument dispatch without running a model. run_inference.pyrefuses to overwrite existing predictions. Use a different output root when retaining an earlier run.run_metrics.pysupports--sequence,--target-resolution, and--workersfor smaller functional checks. Use the default--workers 1: higher values have deadlocked on an evaluation host and are not validated for this release.reproduce_paper.pysupports--tables-onlyand--figures-only.
Changing the data subset, output resolution, sampling settings, or evaluation configuration is useful for a smoke test but will not reproduce the camera-ready numerical values exactly.
Run only the full GRADE model on your own frames
For the full GRADE model, config.yaml also supports direct inference from synchronized radar and DJI RGB arrays without ZED ground-truth depth. Put each sequence under inputs/<sequence>/ with radar.npy (complex64, shape (N, 64, 2, 8, 256)) and dji_rgb.npy (uint8 RGB, shape (N, 504, 896, 3)). The arrays must have the same frame count and order. The GRADE camera calibration is embedded in preprocessing.
accelerate launch --num_processes 1 --num_machines 1 \
--mixed_precision fp16 --dynamo_backend no \
src/GRADE/stage2_diffusion_refinement/inference_full.py \
--config config.yaml
The output is outputs/ours_full/<sequence>_pred.npy, a float32 array of shape (N, 1, 288, 512) normalized to [0, 1]. Multiply by 11.2 for estimated depth in meters. Set data.smoke_eval_root and inference.output_root in config.yaml to use other paths. The script overwrites same-named outputs on rerun.
Checkpoint integrity and citation
SHA256SUMS lists the hashes of all 20 .safetensors files. Run sha256sum -c SHA256SUMS from the repository root after download. The checkpoints contain inference tensors, not optimizer or training state. TAESD is an auxiliary VAE from the TAESD project, whose upstream model is MIT licensed.
@inproceedings{zhao2026grade,
author = {Bin Zhao and Patrick Chiou and Nakul Garg},
title = {{GRADE}: Single-Frame Generative Radar Depth Estimation Under Visual Degradation},
booktitle = {Proceedings of the 32nd Annual International Conference on Mobile Computing and Networking},
series = {MobiCom '26},
year = {2026},
publisher = {ACM},
doi = {10.1145/3795866.3844478}
}
- Downloads last month
- -