SparDA: Sparse Decoupled Attention for Efficient Long-Context LLM Inference Paper • 2606.04511 • Published Jun 3 • 3
somosnlp-hackathon-2022/readability-es-hackathon-pln-public Viewer • Updated Apr 13, 2023 • 1.02k • 79 • 3
nvidia/llama-nemotron-embed-vl-1b-v2 Sentence Similarity • 2B • Updated about 1 month ago • 75.5k • 96
Embarrassingly Simple Self-Distillation Improves Code Generation Paper • 2604.01193 • Published Apr 1 • 56
MSA: Memory Sparse Attention for Efficient End-to-End Memory Model Scaling to 100M Tokens Paper • 2603.23516 • Published Mar 6 • 53