| --- |
| title: DermDepth |
| emoji: 🩺 |
| colorFrom: indigo |
| colorTo: pink |
| sdk: gradio |
| sdk_version: 6.20.0 |
| app_file: app.py |
| short_description: Metric-scale 3D from a single dermatology photo |
| python_version: "3.12" |
| startup_duration_timeout: 30m |
| license: cc-by-nc-4.0 |
| tags: |
| - depth-estimation |
| - 3d-reconstruction |
| - dermatology |
| - medical-imaging |
| - skin-lesion |
| - moge |
| models: |
| - hcarrion/DermDepth |
| - Ruicheng/moge-2-vitl-normal |
| datasets: |
| - hcarrion/D-Synth |
| --- |
| |
| # 🩺 DermDepth — Monocular Metric-Scale 3D for Dermatology |
|
|
| Official demo for **[DermDepth: Toward Monocular Metric Scale 3D Reconstruction Models for |
| Dermatology](https://arxiv.org/abs/2607.13010)** (Carrión & Norouzi, MICCAI 2026). |
|
|
| Dermatology is largely a **measurement** problem — clinicians screen and monitor lesions and wounds |
| by tracking size, border, elevation and texture over time. Those properties are inherently 3D, yet |
| point-of-care imaging is almost always a single 2D photograph. |
|
|
| DermDepth recovers **metric-scale** 3D from *one* ordinary photo — no depth sensor, no second view, |
| no ruler in frame. A 2.1M-parameter scale-and-normal head sits on a frozen |
| [MoGe-2](https://huggingface.co/Ruicheng/moge-2-vitl-normal) backbone and is trained progressively on |
| **[D-Synth](https://huggingface.co/datasets/hcarrion/D-Synth)** (synthetic renders with pixel-perfect |
| depth, normals and intrinsics) and then on real clinical data. |
|
|
| Benchmark results from the paper (held-out test sets) — accuracy on your own photograph may be worse: |
|
|
| | | MoGe-2 baseline | **DermDepth** | |
| |---|:---:|:---:| |
| | SKINL2 scale ratio (target 1.0×) | 16.10× | **0.87×** | |
| | WoundsDB scale ratio | 0.62× | **0.91×** | |
| | DDI scale ratio | 81.0× | **1.95×** | |
| | Fitzpatrick scale disparity | 10.90 | **1.02** | |
|
|
| ## Two checkpoints, on purpose |
|
|
| | Output | Checkpoint | Why | |
| |---|---|---| |
| | Metric depth + 3D mesh | `DermDepth_Synth_SKINL2_WoundsDB_DDI.pt` | The paper's best model (D-Synth → SKINL2 + WoundsDB → DDI pseudo-GT); best metric scale and lowest skin-tone disparity. | |
| | Surface normals | `DermDepth_Synth_Normals.pt` | The dedicated normal-head model. Real clinical normal ground truth is noisy (SKINL2 plenoptic depth has local planar noise; WoundsDB ToF is sparse and offset from RGB), so D-Synth's rendered normals are the only clean normal supervision. | |
|
|
| These are two complete, separately-trained models that do not share a trunk, so the demo loads both |
| rather than swapping a head. |
|
|
| ## Inference resolution |
|
|
| MoGe-2 converts `resolution_level` into a ViT token budget over `num_tokens_range = [1200, 3600]`. |
| This demo addresses `num_tokens` directly and keeps every option **inside** that range — "Ultra" |
| (3600) is the model's genuine maximum, and is the default. |
|
|
| ## Examples & credits |
|
|
| The **first three** examples are real clinical photographs from **WoundsDB** (Chronic Wounds |
| Multimodal Image Database, Silesian University of Technology), used under |
| [CC BY 4.0](https://creativecommons.org/licenses/by/4.0/) at their native 320×240 — the resolution |
| the paper evaluates WoundsDB at. They are **held-out** cases (the paper splits WoundsDB by case: |
| 1–30 train, 31+ test), chosen for anatomical variety: `case_45` (leg ulcer), `case_33` (hand), |
| `case_42` (foot). |
|
|
| > Kręcichwost, M., Czajkowska, J., Wijata, A., Juszczyk, J., Pyciński, B., Biesok, M., Rudzki, M., |
| > Majewski, J., Kostecki, J., & Pietka, E. (2021). Chronic wounds multimodal image database. |
| > *Computerized Medical Imaging and Graphics*, 88, 101844. |
| > [doi:10.1016/j.compmedimag.2020.101844](https://doi.org/10.1016/j.compmedimag.2020.101844) |
|
|
| The **last three** are **synthetic renders** from |
| [D-Synth](https://huggingface.co/datasets/hcarrion/D-Synth), released under |
| [CC BY-NC 4.0](https://creativecommons.org/licenses/by-nc/4.0/) — one per Fitzpatrick group |
| (I–II, III–IV, V–VI). They are renders, **not** patient photographs, and imply no clinical diagnosis. |
|
|
| No DDI imagery is bundled: Stanford's Research Use Agreement prohibits redistributing any portion |
| of that dataset. |
|
|
| > ⚠️ **Research demonstration only — not a medical device.** Outputs are not diagnostic and must not |
| > inform clinical decisions. Predictions on out-of-distribution images can fail silently. |
|
|
| ## Links |
|
|
| 📄 [Paper](https://arxiv.org/abs/2607.13010) · |
| 🤗 [Model](https://huggingface.co/hcarrion/DermDepth) · |
| 📊 [D-Synth](https://huggingface.co/datasets/hcarrion/D-Synth) · |
| 💻 [Code](https://github.com/hectorcarrion/dermdepth) |
|
|
| ```bibtex |
| @inproceedings{carrion2026dermdepth, |
| title = {DermDepth: Toward Monocular Metric Scale 3D Reconstruction Models for Dermatology}, |
| author = {Carri{\'o}n, H{\'e}ctor and Norouzi, Narges}, |
| booktitle = {Medical Image Computing and Computer-Assisted Intervention (MICCAI)}, |
| year = {2026}, |
| eprint = {2607.13010}, |
| archivePrefix = {arXiv} |
| } |
| ``` |
|
|
| ## License |
|
|
| Demo code: CC BY-NC 4.0, matching the [DermDepth weights](https://huggingface.co/hcarrion/DermDepth). |
| The MoGe-2 base weights remain under their original license. Example images are under their |
| respective licenses, credited above. |
|
|