Image classification on the Apple Neural Engine (via ANEForge)

ANEForge runs computation on the Apple Neural Engine (ANE) directly, without CoreML. load_vit loads a Hugging Face Vision Transformer image classifier (ViTForImageClassification) from the Hub by repo id and runs the whole forward pass on the engine.

This is a usage card, not a re-hosted model: it points at the upstream weights and shows how to run them on the ANE.

Install

pip install aneforge

Apple Silicon, macOS 14+.

Use

from aneforge.models import load_vit
from PIL import Image

vit = load_vit("google/vit-base-patch16-224")   # any HF ViT image classifier
image = Image.open("cat.jpg")
print(vit.classify(image, top_k=5))              # [(label, logit), ...]; forward on the ANE
# vit(image) -> raw logits [1, num_labels]

Measured

On an M5 Pro, google/vit-base-patch16-224 runs the full forward in ~27 ms/image, matching the Hugging Face reference (same top-1, relerr 4e-3). Preprocessing uses the model's own AutoImageProcessor.

Scope

ViT-family classifiers with a CLS token and a pre-norm encoder (ViTForImageClassification and compatible DeiT/BEiT-style models); both the modern and legacy HF weight namings are handled. ResNet / ConvNeXt loaders are tracked as follow-up issues in the repo.

Why the ANE

The ANE is the fixed-function accelerator on every recent Apple device. In production it is reachable only through CoreML, which can silently fall back to CPU/GPU; ANEForge compiles the classifier to a single ANE program and dispatches it through the same daemon and kernel-driver stack Apple's own frameworks use.

Links

Cite

Bryngelson, S. H. ANEForge: Python for direct computation on the Apple Neural Engine. arXiv:2606.17090 (2026).

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for aneforge/vit-image-classification