Pretraining Transformers with Quantized Softmax in Attention Paper • 2609.33591 • Published 4 days ago • 4
Softmax Reparameterization for Output-Head Quantization Paper • 2609.31291 • Published 6 days ago • 5
Approximating Softmax in Pretrained LLMs: Model Sensitivity and Kernel Acceleration Paper • 2609.33586 • Published 4 days ago • 3
Approximating Softmax in Pretrained LLMs: Model Sensitivity and Kernel Acceleration Paper • 2609.33586 • Published 4 days ago • 3
Pretraining Transformers with Quantized Softmax in Attention Paper • 2609.33591 • Published 4 days ago • 4
Approximating Softmax in Pretrained LLMs: Model Sensitivity and Kernel Acceleration Paper • 2609.33586 • Published 4 days ago • 3
Pretraining Transformers with Quantized Softmax in Attention Paper • 2609.33591 • Published 4 days ago • 4