ONNX
onnxruntime
onnx-mlir
quantization
fp32
ONNX_Models / reports /conversion /pipeline_status.csv
purejomo's picture
Finalize public ONNX/ONNX-MLIR validation release
ed3aeeb
Raw
History Blame Contribute Delete
4.06 kB
model_id,task,fp32_source_format,public_quantized_source_format,public_quantization_scheme,s0_fp32_source,s1_public_quantized_source,pair,s2_fp32_onnx,s3_public_quantized_onnx,s6_fp32_mlir,s6_public_quantized_mlir,s6_combined
AD01,anomaly_detection,TFLite FlatBuffer FP32,TFLite FlatBuffer,FULL_INTEGER static PTQ performed upstream; published artifact consumed unchanged; TFLITE_BUILTINS_INT8 with int8 I/O,PASS,PASS,PASS,PASS,PASS,PASS_WITH_PATCH,PARTIAL,PARTIAL
LM04,language_modeling/text_classification,ONNX FP32,ONNX,dynamic ONNX quantization; MatMulInteger/DynamicQuantizeLinear with quantized embeddings,PASS,PASS,PASS,PASS,PASS,PASS,PARTIAL,PARTIAL
OD06,object_detection,TFLite FlatBuffer,TFLite FlatBuffer with metadata,FULL_INTEGER core with dequantized float raw-head outputs,PASS,PASS,PASS,PASS,PARTIAL,PARTIAL,PARTIAL,PARTIAL
OD07,object_detection,TFLite FlatBuffer,TFLite FlatBuffer with metadata,FULL_INTEGER core with dequantized float raw-head outputs,PASS,PASS,PASS,PASS,PARTIAL,PARTIAL,PARTIAL,PARTIAL
SG06,semantic_segmentation,TensorFlow GraphDef/SavedModel,TFLite FlatBuffer,FULL_INTEGER,PASS,PASS,PASS,PARTIAL,PARTIAL,PARTIAL,PARTIAL,PARTIAL
SG07,semantic_segmentation,TensorFlow GraphDef/SavedModel,TFLite FlatBuffer,FULL_INTEGER,PASS,PASS,PASS,PASS_WITH_PATCH,PASS,PARTIAL,PARTIAL,PARTIAL
SG08,semantic_segmentation,TensorFlow GraphDef/SavedModel,TFLite FlatBuffer,FULL_INTEGER,PASS,PASS,PASS,PASS_WITH_PATCH,PARTIAL,PARTIAL,PARTIAL,PARTIAL
SP01,keyword spotting,TFLite FlatBuffer; upstream float32-reference runtime with float32 I/O/activations and 5 per-channel INT8 constant tensors (HYBRID_OR_WEIGHT_ONLY observed),TFLite FlatBuffer,full-integer affine TFLite; upstream PTQ artifact used unchanged,PASS,PASS,PASS,PARTIAL,PASS,PASS,PARTIAL,PARTIAL
SP02,keyword spotting,Keras H5,TFLite FlatBuffer,full-integer affine TFLite public artifact; upstream training/conversion may use QAT and representative calibration,PASS,PASS,PASS,PASS_WITH_PATCH,PASS,PARTIAL,PARTIAL,PARTIAL
SP08,keyword spotting,TensorFlow GraphDef/SavedModel,TFLite FlatBuffer inside ZIP,full-integer affine TFLite; README says strictly int8 but measured legacy artifact tensors are uint8/int32,PASS,PASS,PASS,PASS_WITH_PATCH,PASS,PASS,PARTIAL,PARTIAL
SP09,keyword spotting,TFLite FlatBuffer,TFLite FlatBuffer,clustered (32 clusters; kmeans++) retrained FP32 then upstream post-training full-integer quantization,PASS,PASS,PASS,PASS,PASS,PASS,PARTIAL,PARTIAL
VC01,vision_classification,TFLite FlatBuffer,TFLite FlatBuffer,upstream static full-integer TFLite,PASS,PASS,PASS,PASS,PASS,PASS,PARTIAL,PARTIAL
VC02,vision_classification,TFLite FlatBuffer,TFLite FlatBuffer,upstream static full-integer TFLite,PASS,PASS,PASS,PASS,PASS,PASS,PARTIAL,PARTIAL
VC03,vision_classification,TFLite FlatBuffer,TFLite FlatBuffer inside TGZ,upstream quantization-aware training/FakeQuant then fully-quantized TFLite,PASS,PASS,PASS,PASS,PASS,PASS,PARTIAL,PARTIAL
VC04,vision_classification,TFLite FlatBuffer,TFLite FlatBuffer inside TGZ,upstream quantization-aware training/FakeQuant then fully-quantized TFLite,PASS,PASS,PASS,PASS,PASS,PASS,PARTIAL,PARTIAL
VC05,vision_classification,Keras v3,TFLite FlatBuffer,upstream static INT8 TFLite,PASS,PASS,PASS,PASS_WITH_PATCH,PARTIAL,PARTIAL,PARTIAL,PARTIAL
VC06,vision_classification,ONNX,ONNX QDQ,upstream static QDQ INT8 ONNX,PASS,PASS,PASS,PASS,PASS,PASS,PARTIAL,PARTIAL
VC09,vision_classification,ONNX,ONNX QOperator INT8,upstream static PTQ via Intel Neural Compressor/ONNX Runtime,PASS,PASS,PASS,PASS,PASS,PASS,PARTIAL,PARTIAL
VC11,vision_classification,TFLite FlatBuffer,TFLite FlatBuffer,static full-integer TFLite distributed by Google,PASS,PASS,PASS,PASS,PARTIAL,PASS,PARTIAL,PARTIAL
VC12,vision_classification,ONNX,ONNX QOperator INT8,upstream static PTQ via Intel Neural Compressor/ONNX Runtime,PASS,PASS,PASS,PASS,PASS,PASS,PARTIAL,PARTIAL
VC13,vision_classification,ONNX,ONNX QOperator INT8,upstream static PTQ via Intel Neural Compressor/ONNX Runtime,PASS,PASS,PASS,PASS,PASS,PASS,PARTIAL,PARTIAL