|
Download README.md from rinpatch/dinov3-coreml: direct link, hf CLI and curl.
- Browser
- Download file 4.92 kB
-
https://huggingface.co/rinpatch/dinov3-coreml/resolve/main/README.md
- Command line
-
hf download hf://rinpatch/dinov3-coreml/README.md
-
curl -L -o README.md https://huggingface.co/rinpatch/dinov3-coreml/resolve/main/README.md
4.92 kB
| license: other | |
| license_name: dinov3 | |
| license_link: https://huggingface.co/rinpatch/dinov3-coreml/blob/main/LICENSE.dinov3.md | |
| base_model: facebook/dinov3-vitb16-pretrain-lvd1689m | |
| pipeline_tag: image-feature-extraction | |
| inference: false | |
| tags: | |
| - coreml | |
| - dinov3 | |
| - image-embeddings | |
| - macos | |
| # DINOv3 Core ML | |
| A Core ML conversion of Meta's DINOv3 ViT-B/16 image encoder. It returns an | |
| L2-normalized, 768-value CLS embedding for image similarity and retrieval, plus the | |
| 784 raw patch tokens for dense tasks such as segmentation, depth, and correspondence. | |
| This is a format conversion of pretrained weights; no additional training was performed. | |
| ## Model | |
| | Property | Value | | |
| |---|---| | |
| | Base model | `facebook/dinov3-vitb16-pretrain-lvd1689m` | | |
| | timm implementation | `vit_base_patch16_dinov3.lvd1689m` | | |
| | Artifact | `models/DINOv3ViTB16-FP32-448.mlpackage` | | |
| | Format | Core ML ML Program, FP32 | | |
| | Target | macOS 14 or newer | | |
| | Input | `image`: 448×448 RGB, pixel values 0–255 | | |
| | Output | `embedding`: float32, shape `[1, 768]`, unit L2 norm | | |
| | Output | `patch_embeddings`: float32, shape `[1, 784, 768]`, unnormalized | | |
| Correct EXIF orientation, convert to RGB, and resize to 448×448 before inference. | |
| The examples use a square resize, which can distort non-square images. Use the same | |
| preprocessing for all images being compared. Pixel scaling and ImageNet normalization | |
| are inside the model; do not apply them again. Similarity is the dot product of two | |
| `embedding` vectors. It does not provide text embeddings, captions, or coordinates. | |
| `patch_embeddings` is a 28x28 grid of 768-value tokens flattened in row-major order, | |
| one per 16x16 input patch. Unlike `embedding` these are not normalized, because dense | |
| heads generally want the magnitude; normalize per token yourself for cosine. The four | |
| register tokens are dropped: they exist to absorb high-norm artifacts that would | |
| otherwise pollute the patch tokens, and are not useful as features. | |
| ## Run | |
| Use an Apple silicon Mac and Python 3.12. Python inference requires macOS; | |
| Swift inference does not need Python or the conversion dependencies. | |
| ```sh | |
| uv sync --python 3.12 | |
| uv run python examples/embed.py | |
| uv run python examples/embed.py image.jpg --output build/embedding.npy | |
| ``` | |
| The first command after installation uses a synthetic image to check the model's | |
| output shape, finite values, and normalization. No photo library access is needed. | |
| For native Swift, provide an already resized, orientation-corrected 448×448 image: | |
| ```sh | |
| mkdir -p build | |
| swiftc -O examples/encode.swift -o build/encode | |
| build/encode models/DINOv3ViTB16-FP32-448.mlpackage image-448.png | |
| ``` | |
| The example compiles the package at runtime. An app can instead add the package to | |
| its Xcode target, which compiles and bundles it at build time, and load the `.mlmodelc` | |
| resource once for repeated predictions. | |
| ## Convert | |
| ```sh | |
| uv run python convert.py --output build/DINOv3ViTB16-FP32-448.mlpackage | |
| ``` | |
| The converter downloads pretrained weights through timm/Hugging Face. If the upstream | |
| checkpoint requires access, accept its terms and authenticate with Hugging Face first. | |
| It exports through `torch.export`, converts to Core ML, and compares both runtimes | |
| on a synthetic RGB gradient. A cosine below 0.999 on either output fails validation. | |
| `--size` changes the square input resolution; use a positive multiple of 16. | |
| `--precision fp16` is experimental and must pass its own validation. | |
| Existing output packages are protected unless `--force` is supplied. | |
| The copied FP32 artifact's original validation result is in | |
| `models/DINOv3ViTB16-FP32-448.validation.json`: cosine approximately 1.0 for both | |
| outputs, and CLS norm approximately 1.0. Floating-point rounding can put cosine slightly | |
| above 1. This single-image check establishes limited conversion parity, not retrieval | |
| accuracy across datasets. The supplied package retains its original conversion metadata. | |
| ## Attribution and license | |
| Base model and DINOv3 research by Meta: | |
| [upstream repository](https://github.com/facebookresearch/dinov3) and | |
| [model card](https://huggingface.co/facebook/dinov3-vitb16-pretrain-lvd1689m). | |
| The converted weights are subject to the [DINOv3 License](LICENSE.dinov3.md), | |
| included verbatim from upstream. This conversion is not an official Meta release. | |
| The conversion and example code in this repository — `convert.py` and `examples/` — | |
| is [MIT licensed](LICENSE). The two licenses are independent: the MIT terms never | |
| extend to the weights, and using the code to convert your own checkpoint does not | |
| put that output under the DINOv3 License. | |
| ## Repository files | |
| `convert.py` and the locked Python environment reproduce the conversion. | |
| `examples/` contains Python and Swift inference examples. `models/` contains the | |
| converted package and original validation report. `SHA256SUMS` records the artifact | |
| checksums. Large model files use Git LFS through `.gitattributes`. | |