Ctrl+K
Update logbook: Reproduction: AsyncSpade: Efficient Test-Time Scaling with Asynchronous Sparse Decoding
c3f9073 verified - claim-1-asyncspade-achieves-over-20-reduction-in-time-per-output-token-tpot-compared-to-the-sota-sparse-attention-baseline-quest-evaluated-on-qwen3-8b-and-qwen3-32b-section-5
- claim-2-asyncspade-achieves-at-least-50-tpot-reduction-compared-to-full-attention-while-matching-or-surpassing-accuracy-on-aime-24-aime-25-gpqa-diamond-and-math-500-figure-9
- claim-3-asyncspade-fully-overlaps-kv-cache-management-operations-with-the-inference-pipeline-within-a-defined-workload-range-achieving-the-theoretical-optimal-tpot-section-4-figures-4-5
- claim-4-cache-management-latency-and-required-bandwidth-are-measured-at-3-92-14-43ms-across-configurations-on-a100-and-h100-8-gpu-nodes-compared-to-inference-latencies-of-5-47-14-26ms-table-1-table-2
- claim-5-the-methods-query-prediction-component-exploits-observed-temporal-locality-and-linear-correlation-in-attention-patterns-across-decoding-steps-to-select-sparse-kv-subsets-asynchronously-section-3-figures-4-5
- conclusion
- executive-summary
- 2.24 kB