ai.onnx.DequantizeLinear
ai.onnx · standard ONNX operator · ONNX opset ≥ 25
Description
Dequantizes a quantized tensor back to full precision using the formula y = (x - x_zero_point) * x_scale. Scale and zero point determine quantization granularity: scalar for per-tensor, 1-D for per-axis, or same rank as input for blocked quantization. The output type matches x_scale unless overridden by output_dtype.
See the ONNX DequantizeLinear spec for the reference semantics.
Inputs
| Name | Bind key | Logical dtype | Rank | Shape | Description | Presence |
|---|---|---|---|---|---|---|
x |
x |
TQ |
— | — | N-D quantized input tensor to be dequantized. | required |
x_scale |
x_scale |
TF |
— | — | Scale for x; scalar for per-tensor, 1-D for per-axis, or same-rank tensor for blocked dequantization. |
required |
x_zero_point |
x_zero_point |
TQ |
— | — | Zero point for x with shape matching x_scale; defaults to zero when absent. |
optional |
Outputs
| Name | Bind key | Logical dtype | Rank | Shape | Description | Presence |
|---|---|---|---|---|---|---|
y |
y |
TF |
same as x |
same as x |
N-D full-precision output with the same shape as x, typed according to x_scale or output_dtype. |
required |
Attributes
Default values (overridable per request):
| Attribute | Default | Description |
|---|---|---|
axis |
1 |
The axis of the dequantizing dimension, used for per-axis and blocked quantization; negative values index from the end. |
block_size |
0 |
Number of elements along axis that share each scale value for blocked quantization; 0 means not blocked. |
output_dtype |
0 |
ONNX TensorProto element-type code for y; 0 uses the dtype of x_scale. The current float16/float32 routes require it to agree with x_scale. |
Type constraints
| Variable | Allowed dtypes |
|---|---|
TQ |
uint8, int8, int16, int32 |
TF |
float32, float16 |
Files
metadata.json— kernel metadata (id, digests, provenance)manifest.json— the op contract (source of truth)test.json— correctness casesbench.json— benchmark + tuning casesquant-linear-blocked-axis.wgsl.jinjaquant-linear-scalar.wgsl.jinjaquant-linear-vec4.wgsl.jinja
Use with @huggingface/kernels
The loader derives every required output's shape and logical dtype from the manifest contract and this call. It then allocates the result tensors automatically.
The version: 1 option selects the published kernel contract; it is independent of any operator opset, contrib since_version, or model version.
Replace each *Data placeholder with a typed array containing the corresponding input data.
import { getKernel } from "@huggingface/kernels";
const kernel = await getKernel("webgpu-kernels/ai.onnx.DequantizeLinear", { version: 1 });
const { y } = await kernel({
x: { data: xData, shape: [4] },
x_scale: { data: x_scaleData, shape: [] },
});
- Downloads last month
- -
Requires WebGPU support. See the compatibility table.