ViT-Base-P16 (ONNX) – Renesas X5H

Introduction

This repository hosts ViT-Base/16 (Vision Transformer, 16×16 patches), targeting the Renesas R-Car X5H platform for image classification inference on the NPX6 NPU.

  • Model Architecture: ViT-Base/16 — a Vision Transformer backbone, pretrained with MAE (Masked Autoencoder) self-supervision and then fine-tuned for ImageNet-1k classification.
  • Source Model: OpenMMLab config vit_base_p16_32xb128_mae_in1k (no HuggingFace mirror of these weights; see model.source in .metadata.yaml)
  • Task: Image Classification (ImageNet-1k, 1000 classes)
  • Parameters: 86M

Note: A vit_tiny variant also exists in the source benchmark data but produced zero passing compile/execute runs in the APM50 CI pipeline (both slices failed). It is intentionally not included in this repository — only the ViT-Base/16 checkpoint, which has passing benchmark numbers, is published here.

Deployment Flow

The FP32 ONNX model is auto-cast to INT8 by the Renesas MWMX toolchain at compile time — no separate quantization step is required.

vit_base_p16_..._optimized.onnx (FP32)
        │
        └─▶  MWMX Runtime  ──▶  INT8 auto-cast  ──▶  NPX6 NPU

Provided Artifacts

Artifact Status Notes
FP32 (ONNX) ✅ fp32/vit-base-p16_32xb128-mae_in1k.onnx — FP32 ONNX export

Performance

Measured on Renesas R-Car X5H via the MWMX runtime (APM50 ship-performance CI pipeline).

Benchmark configuration: Single NPU · Batch size: 1 · Input resolution: not available from source data (TBD)

Parameters Runtime Precision Device Latency (ms) Type
86M MWMX Runtime INT8 (auto) X5H · 1× NPU · 1 Core · 850 MHz 25.796629 Measured
86M MWMX Runtime INT8 (auto) X5H · 1× NPU · 1 Core · 850 MHz 25.923109 Measured (2026-09-16)
86M MWMX Runtime INT8 (auto) X5H · 1× NPU · 12 Cores · 850 MHz 6.873014 Measured

Accuracy

TBD — not yet measured/published for this repo.


Runtime Details

MWMX Runtime

  • Engine: Renesas MWMX (Middleware MX) native inference runtime
  • Input format: FP32 ONNX (compiled by the MWMX toolchain)
  • NPU execution precision: INT8 (auto-cast by MWMX toolchain)
  • Execution target: NPX6-48K NPU on R-Car X5H

Prerequisites

To run inference on Renesas R-Car X5H, you need:

  1. Renesas R-Car X5H board with NPX6 NPU
  2. Renesas MWMX Runtime
  3. Hugging Face CLI to download the model

Download

hf download Renesas/ViT-Base-P16-ONNX --repo-type=model --include "fp32/*"

Benchmark Methodology

  • HIL runs: Hardware-in-the-loop — measured on physical R-Car X5H silicon via the MWMX runtime (metawaremx_runtime CI pipeline, "APM50" ship-performance target)
  • Precision: FP32 ONNX input; INT8 execution (auto-cast by MWMX)
  • Slices: results reported for both 1 AI core and 12 AI cores per NPU instance
  • Excluded: the vit_tiny variant from the same source data batch had 0 passing runs across both slices and is not represented in this repository
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support