File size: 5,060 Bytes
d298481 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 | ---
license: apache-2.0
library_name: kernels
tags:
- kernel
- audio
- vocoder
- bigvgan
- cpu
---
# vocoder-ops
Fused anti-aliased Snake activation for BigVGAN-class neural vocoders,
loadable through `kernels`. The resampling pair that keeps the nonlinearity
from aliasing is folded into the activation, so the oversampled signal never
reaches memory. The reference baseline is the framework formulation, matched
to 1.9e-7.
A neural vocoder's Snake activation must be anti-aliased, and the standard
way is to upsample, apply the nonlinearity, and downsample: two depthwise
convolutions with the oversampled tensor materialized twice, three times per
generator block. None of it is a GEMM, so the whole chain is memory-bound
and fuses. This kernel fuses it, and a whole BigVGAN block's activations run
in about a third less time with output identical to seven decimal places.

*BigVGAN generator stages at ratio 2 and 25 taps on 14 x86 threads: 1.10x,
1.79x, 2.07x, and 1.11x per stage, 1.60x for a whole block's three
activations (51.7 ms against 32.3), with output matching the framework to
1.9e-7.*
## Usage
```python
import kernels
vo = kernels.get_kernel("phanerozoic/vocoder-ops", version=1, trust_remote_code=True)
act = vo.AntiAliasedSnake(channels=512, ratio=2, taps=25, alpha=a, beta=b)
y = act(x) # [B, C, L] -> [B, C, L]
y = vo.snake(x, alpha, beta) # plain Snake, no resampling
```
`version` selects the release branch; `trust_remote_code` is required by
`kernels` for publishers without the trusted-publisher mark.
## API
| Symbol | Purpose |
|---|---|
| `AntiAliasedSnake(channels, ratio, taps, alpha, beta)` | fused upsample, Snake, downsample |
| `snake(x, alpha, beta)` | plain Snake activation, no resampling |
## Method
Three things carry the speedup.
Polyphase decomposition: both filters decompose by the residue of the
oversampled index modulo `ratio`, and within a phase the taps land on
consecutive samples. Walking the oversampled grid directly makes both
filters strided and forces the chain scalar; walking phases makes every
inner loop a unit-stride convolution. This alone was worth 6x.
One transcendental instead of two operations: Snake is `x + sin^2(ax)/b`,
and `sin^2(u) = (1 - cos 2u)/2`, so the activation is
`x + (1 - cos(2ax))/(2b)`. The cosine is evaluated with a Cody-Waite
reduction and a minimax polynomial, with quadrant selection done
branchlessly; a per-lane branch there spills the vector and runs scalar,
which costs more than the polynomial itself.
Multiversioned leaves: on x86-64 the hot loops carry both an AVX2 and a
baseline lowering behind an IFUNC resolver. The attribute has to sit on the
functions containing the loops rather than their callers, because
`at::parallel_for` takes a type-erased callable and a clone attribute never
reaches the lambda body.
## Measured
BigVGAN generator stages, ratio 2, 25 taps, against the framework
formulation; both produce identical output to 1.9e-7 relative:
| stage | vocoder-ops | framework | speedup |
|---|---:|---:|---:|
| **x86-64** (i9-12900H, 14 threads) | | | |
| C=256 L=800 | 0.93 ms | 1.02 ms | 1.10x |
| C=128 L=6400 | 3.17 ms | 5.66 ms | 1.79x |
| C=64 L=12800 | 3.30 ms | 6.83 ms | 2.07x |
| C=32 L=25600 | 3.37 ms | 3.72 ms | 1.11x |
| **whole block, 3 activations** | **32.30 ms** | **51.67 ms** | **1.60x** |
| **aarch64** (Cortex-A76, 4 threads) | | | |
| C=256 L=800 | 3.68 ms | 20.68 ms | 5.63x |
| C=128 L=6400 | 14.76 ms | 99.15 ms | 6.72x |
| C=64 L=12800 | 14.72 ms | 100.05 ms | 6.80x |
| C=32 L=25600 | 14.62 ms | 94.72 ms | 6.48x |
| **whole block, 3 activations** | **143.4 ms** | **943.8 ms** | **6.58x** |
Plain `snake` without resampling is 2.41x on aarch64 and 0.69x on x86-64,
where the framework's vectorized sine is already well optimized. The
anti-aliased path is the one BigVGAN actually uses.
## Correctness
Output matches the framework formulation to 1.9e-7 relative across every
stage measured. Expressed as two separate convolutions, the reference pads
the activation, so the oversampled signal is zero outside the span rather
than the activation of the filter tail; fusing removes that boundary, so
this kernel reimposes it and matches the framework result exactly. Without
it the interior agrees to floating-point rounding and only the edges differ.
## Requirements and limits
- CPU, `x86_64-linux` and `aarch64-linux`, for PyTorch 2.11, 2.12, and 2.13
(six prebuilt variants), each architecture built on its own host and
validated against the prebuilt binary.
- Plain `snake` is slower than the framework on x86-64; use the
anti-aliased path, which is what BigVGAN calls.
## References
Lee et al., BigVGAN (2023); Ziyin et al., "Neural Networks Fail to Learn
Periodic Functions and How to Fix It" (the Snake activation, 2020);
Cody-Waite range reduction.
## License
Apache-2.0.
|