File size: 4,799 Bytes
9e7cac1 31544ae 9e7cac1 31544ae 9e7cac1 31544ae 6d57be0 31544ae 6d57be0 31544ae 6d57be0 31544ae 6d57be0 31544ae 6d57be0 31544ae 6d57be0 31544ae 6d57be0 31544ae 6d57be0 31544ae 6d57be0 31544ae 6d57be0 31544ae 6d57be0 31544ae 6d57be0 31544ae 6d57be0 31544ae | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 | ---
library_name: kernels
license: apache-2.0
tags:
- kernel
- webgpu
- wgsl
---
# ai.onnx.ReduceL2
`ai.onnx` · standard ONNX operator · ONNX opset ≥ 18
## Description
Computes the L2 norm (Euclidean norm) of the input tensor along the specified axes: `sqrt(sum(x^2))`. The output rank matches the input when `keepdims` is 1; otherwise reduced dimensions are pruned. Reduction over an empty set of values yields 0.
See the [ONNX `ReduceL2` spec](https://onnx.ai/onnx/operators/onnx__ReduceL2.html) for the reference semantics.
## Inputs
| Name | Upstream name | Logical dtype | Rank | Shape | Description | Presence |
| --- | --- | --- | --- | --- | --- | --- |
| `x` | `data` | `T` | — | — | Input tensor to reduce. | required |
## Outputs
| Name | Upstream name | Logical dtype | Rank | Shape | Description | Presence |
| --- | --- | --- | --- | --- | --- | --- |
| `y` | `reduced` | `T` | derived | — | Reduced output tensor containing the L2 norm along the specified axes. | required |
## Attributes
Default values (overridable per request):
| Attribute | Default | Description |
| --- | --- | --- |
| `axes` | `[]` | Values of the optional ONNX `axes` tensor input, supplied through this request attribute; an empty list follows `noop_with_empty_axes`. |
| `keepdims` | `1` | If 1, retains reduced dimensions with size 1 in the output; if 0, removed dimensions are pruned. |
| `noop_with_empty_axes` | `0` | When 1 and axes is empty, skips reduction but still applies the elementwise square and square-root steps, yielding `abs(x)`; when 0 (default), reduces over all axes. |
## Type constraints
| Variable | Allowed dtypes |
| --- | --- |
| `T` | `float32`, `float16`, `int32` |
## Implementation variants
One implementation is selected per call from the device capabilities, the request shapes and the dtypes; these notes say what each one covers.
- `strided_axis_serial` — Flatten a single non-last reduction axis into outer/axis/inner geometry. Compile its strides and loop bound, keep one output per lane and float32 accumulation, and cap the workgroup by the device limits.
## Device requirements
Some implementation variants require `subgroups`. These are route-specific capabilities, not package-wide requirements; availability also depends on the request shape and dtype.
## Files
- [`metadata.json`](build/webgpu/metadata.json) — kernel metadata (id, digests, per-variant templates, provenance)
- [`manifest.json`](build/webgpu/manifest.json) — the op contract (source of truth)
- [`test.json`](build/webgpu/test.json) — correctness cases
- [`bench.json`](build/webgpu/bench.json) — benchmark + tuning cases
- [`reduce-axis-split-reduce.wgsl.jinja`](build/webgpu/reduce-axis-split-reduce.wgsl.jinja)
- [`reduce-axis0-splitk-combine.wgsl.jinja`](build/webgpu/reduce-axis0-splitk-combine.wgsl.jinja)
- [`reduce-axis0-splitk-reduce.wgsl.jinja`](build/webgpu/reduce-axis0-splitk-reduce.wgsl.jinja)
- [`reduce-axis0-tilecols.wgsl.jinja`](build/webgpu/reduce-axis0-tilecols.wgsl.jinja)
- [`reduce-flat-partial.wgsl.jinja`](build/webgpu/reduce-flat-partial.wgsl.jinja)
- [`reduce-multi-axis-coop.wgsl.jinja`](build/webgpu/reduce-multi-axis-coop.wgsl.jinja)
- [`reduce-noop-empty-axes.wgsl.jinja`](build/webgpu/reduce-noop-empty-axes.wgsl.jinja)
- [`reduce-row-subgroup-rows.wgsl.jinja`](build/webgpu/reduce-row-subgroup-rows.wgsl.jinja)
- [`reduce-row-subgroup.wgsl.jinja`](build/webgpu/reduce-row-subgroup.wgsl.jinja)
- [`reduce-row-tree.wgsl.jinja`](build/webgpu/reduce-row-tree.wgsl.jinja)
- [`reduce-serial-axis.wgsl.jinja`](build/webgpu/reduce-serial-axis.wgsl.jinja)
- [`reduce-strided-axis.wgsl.jinja`](build/webgpu/reduce-strided-axis.wgsl.jinja)
## Use with `@huggingface/kernels`
```sh
npm install --save-exact @huggingface/kernels@0.0.1-preview.2
```
Outputs with inferable metadata are allocated automatically. Explicit `outputs` entries request optional results or provide metadata that cannot be inferred from the supplied inputs and attributes.
This example supplies explicit metadata for:
- `y`
The `version: 1` option selects the published kernel contract; it is independent of any operator opset, contrib `since_version`, or model version.
It follows the `v1` branch as fixes land. To pin exact artifact bytes, pass a 40-character commit `revision` instead of `version`.
Replace each `*Data` placeholder with a typed array containing the corresponding input data.
```js
import { getKernel } from "@huggingface/kernels";
const kernel = await getKernel("webgpu-kernels/ai.onnx.ReduceL2", { version: 1 });
// Explicit destinations request optional results or supply metadata that cannot be inferred.
const { y } = await kernel({ x: { data: xData, shape: [] } }, {
outputs: { y: { shape: [], dtype: "float32" } },
});
```
|