liangsu9988's picture
Promote latest kernel artifacts to main
b9ecbf8 verified
|
Raw
History Blame Contribute Delete
777 Bytes
---
tags: [kernel, cuda, attention, inference, cuda-graphs]
library_name: kernels
license: apache-2.0
---
# Masked MHA Runtime
Allocation-free FP16/BF16 attention that masks padded logits inside softmax,
removing the per-call `-inf` pre-fill. BF16 accepts fused-QKV token strides.
## API
- `forward(q, k, v, *, scale=None)`
- `forward_static(q, k, v, *, logits, out, scale=None)`
- `allocate_workspace(q, k)`
Inputs use `(sequence, heads, head_dim)`. `forward_static` is the CUDA Graph
hot-path API: allocate `logits` and `out` once and reuse their addresses.
Rows wider than 1024 keys use a deterministic multi-pass softmax.
This package contains the native masked-MHA execution path validated in
FlashRT's GROOT N1.7 Thor runtime. It is separate from FlashAttention-4.