metadata
tags:
- kernel
- cuda
- attention
- inference
- cuda-graphs
library_name: kernels
license: apache-2.0
Masked MHA Runtime
Allocation-free FP16/BF16 attention that masks padded logits inside softmax,
removing the per-call -inf pre-fill. BF16 accepts fused-QKV token strides.
API
forward(q, k, v, *, scale=None)forward_static(q, k, v, *, logits, out, scale=None)allocate_workspace(q, k)
Inputs use (sequence, heads, head_dim). forward_static is the CUDA Graph
hot-path API: allocate logits and out once and reuse their addresses.
Rows wider than 1024 keys use a deterministic multi-pass softmax.
This package contains the native masked-MHA execution path validated in FlashRT's GROOT N1.7 Thor runtime. It is separate from FlashAttention-4.