ai.onnx.DequantizeLinear

ai.onnx · standard ONNX operator · ONNX opset ≥ 25

Description

Dequantizes a quantized tensor back to full precision using the formula y = (x - x_zero_point) * x_scale. Scale and zero point determine quantization granularity: scalar for per-tensor, 1-D for per-axis, or same rank as input for blocked quantization. The output type matches x_scale unless overridden by output_dtype.

See the ONNX DequantizeLinear spec for the reference semantics.

Inputs

Name Bind key Logical dtype Rank Shape Description Presence
x x TQ N-D quantized input tensor to be dequantized. required
x_scale x_scale TF Scale for x; scalar for per-tensor, 1-D for per-axis, or same-rank tensor for blocked dequantization. required
x_zero_point x_zero_point TQ Zero point for x with shape matching x_scale; defaults to zero when absent. optional

Outputs

Name Bind key Logical dtype Rank Shape Description Presence
y y TF same as x same as x N-D full-precision output with the same shape as x, typed according to x_scale or output_dtype. required

Attributes

Default values (overridable per request):

Attribute Default Description
axis 1 The axis of the dequantizing dimension, used for per-axis and blocked quantization; negative values index from the end.
block_size 0 Number of elements along axis that share each scale value for blocked quantization; 0 means not blocked.
output_dtype 0 ONNX TensorProto element-type code for y; 0 uses the dtype of x_scale. The current float16/float32 routes require it to agree with x_scale.

Type constraints

Variable Allowed dtypes
TQ uint8, int8, int16, int32
TF float32, float16

Files

Use with @huggingface/kernels

The loader derives every required output's shape and logical dtype from the manifest contract and this call. It then allocates the result tensors automatically.

The version: 1 option selects the published kernel contract; it is independent of any operator opset, contrib since_version, or model version.

Replace each *Data placeholder with a typed array containing the corresponding input data.

import { getKernel } from "@huggingface/kernels";

const kernel = await getKernel("webgpu-kernels/ai.onnx.DequantizeLinear", { version: 1 });
const { y } = await kernel({
  x: { data: xData, shape: [4] },
  x_scale: { data: x_scaleData, shape: [] },
});
Downloads last month
-
kernel
webgpu
wgsl
apache-2.0
WebGPU

Requires WebGPU support. See the compatibility table.