WaveCut's picture
Add ordered and approximate INT8 attention decode layouts
99323da verified
|
Raw
History Blame Contribute Delete
319 Bytes

Ordered subwarp INT8 attention

Research-only eight-thread QK groups with four original lane partials per thread. Emulates original dot summation order and uses32-bit aligned key loads with scalar offset fallback. Exact-output hypothesis requires strict tests and model validation; no speed claim before measurement.