mlboydaisuke's picture
Add README.md
0654387 verified
|
Raw
History Blame Contribute Delete
3.69 kB
---
license: apache-2.0
library_name: coreml
pipeline_tag: depth-estimation
base_model: depth-anything/DA3-SMALL
base_model_relation: quantized
tags:
- coreml
- core-ml
- ios
- macos
- apple
- on-device
- monocular-depth
- dinov2
- arxiv:2511.10647
---
# Depth Anything 3 Small (504×504) — Core ML
*ByteDance-Seed, ICLR 2026 oral*
Relative monocular depth from a single image. DA3 Main Series, Small (0.08B params, DINOv2 ViT-S/14 + DualDPT head). First public Core ML conversion of Depth Anything 3.
<p><img src="https://huggingface.co/mlboydaisuke/Depth-Anything-3-Small-CoreML/resolve/main/media/e3fad96333.jpg" alt="Depth Anything 3 Small (504×504) demo"> <img src="https://huggingface.co/mlboydaisuke/Depth-Anything-3-Small-CoreML/resolve/main/media/718482c065.jpg" alt="Depth Anything 3 Small (504×504) demo"> <img src="https://huggingface.co/mlboydaisuke/Depth-Anything-3-Small-CoreML/resolve/main/media/642cf23afa.jpg" alt="Depth Anything 3 Small (504×504) demo"> <img src="https://huggingface.co/mlboydaisuke/Depth-Anything-3-Small-CoreML/resolve/main/media/6c8d97048e.jpg" alt="Depth Anything 3 Small (504×504) demo"></p>
Core ML conversion of [ByteDance-Seed/Depth-Anything-3](https://github.com/ByteDance-Seed/Depth-Anything-3) for on-device inference on iPhone, iPad and Mac. Converted with `coremltools`; the packages are stateless, so all sequencing and buffering lives in your Swift code.
| | |
|---|---|
| Task | depth estimation |
| Upstream | [ByteDance-Seed/Depth-Anything-3](https://github.com/ByteDance-Seed/Depth-Anything-3) |
| Packages | 1 |
| Download size | 44 MB |
| Minimum iOS | 17.0 |
| Peak RAM | ~300 MB |
## Files
| File | Size | Compute units | SHA-256 |
|---|---:|---|---|
| `DepthAnythingV3_small_504.mlpackage.zip` | 44 MB | `all` | `c10f8afa01fdc1d2…` |
| **Total** | **44 MB** | | |
`compute_units` is not a suggestion -- it is the configuration the conversion was verified against. Moving a package to a different compute unit can silently change the numerics (FP16 attention overflow) or crash on the GPU.
## Download
```bash
hf download mlboydaisuke/coreml-zoo --include "depth_anything_v3/*" --local-dir ./depth_anything_v3_small_504
unzip './depth_anything_v3_small_504/depth_anything_v3/*.zip' -d ./depth_anything_v3_small_504
```
## Use in Swift
```swift
import CoreML
let config = MLModelConfiguration()
config.computeUnits = .all // as converted — see the table above
// Unzip the .mlpackage, drop it into your Xcode target and Xcode compiles it
// at build time:
let model = try DepthAnythingV3_small_504(configuration: config)
// ...or compile a downloaded .mlpackage at runtime:
let compiled = try await MLModel.compileModel(at: mlpackageURL)
let model = try MLModel(contentsOf: compiled, configuration: config)
```
## Demo
- **Models Zoo** — this model is downloadable and runnable inside the [Models Zoo app](https://apps.apple.com/app/id6762083207) on the App Store, no build required.
## Conversion
- Script: [`convert_depth_anything_v3.py`](https://github.com/john-rocky/CoreML-Models/blob/master/conversion_scripts/convert_depth_anything_v3.py)
- Pitfalls hit during conversion (FP16 overflow, ANE buffer limits, stride handling): [`docs/coreml_conversion_notes.md`](https://github.com/john-rocky/CoreML-Models/blob/master/docs/coreml_conversion_notes.md)
- Model index: [CoreML-Models](https://github.com/john-rocky/CoreML-Models)
## License
The conversion inherits the upstream license: **Apache-2.0**.
## Credits
- Upstream authors: [ByteDance-Seed/Depth-Anything-3](https://github.com/ByteDance-Seed/Depth-Anything-3), 2025
- Core ML conversion: john-rocky (Daisuke Majima)