Atom2.7m: Representation-Level Specialization for Arithmetic-Aware Small Language Models ucr-max • Jul 7 • 10
Two AI agents review each other's work. The night it went wrong was the useful part. Hug-zol • Jul 6 • 1
After the party comes the free lunch: regularizing ColBERT models to enhance pooling capabilities and reduce index footprint lightonai • Jul 6 • 15
Your local LLM forgets its cache after every restart. I measured how much time that wastes. vimalnakrani • Jul 6
Qwen3.6-27B-FP8 on One RTX 6000 Ada: Fast TTFT, 314 tok/s Decode Generation [Benchmark] hexgridcloud • Jul 6 • 1
Distilling OmniVoice into Aegis: Female Urdu TTS at 61 MB ONNX for CPU Inference mahwizzzz • Jul 5 • 3
Quantum Cryptanalysis on Real Hardware: Pushing Symmetric-Structure Key Recovery Beyond the Published Frontier FINAL-Bench • Jul 5 • 15