HARP Quantized Models Collection Quantized Llama 2 (7B/13B/70B) at 2-bit via HARP, a learnable orthogonal preprocessor for extreme LLM quantization. EMNLP 2026. • 3 items • Updated 30 days ago • 1
HARP Quantized Models Collection Quantized Llama 2 (7B/13B/70B) at 2-bit via HARP, a learnable orthogonal preprocessor for extreme LLM quantization. EMNLP 2026. • 3 items • Updated 30 days ago • 1
Gradient-Faithful Surrogates paper models Collection Final checkpoints of Llama3 and Qwen3 models from Gradient-Faithful Surrogates for KV-Cache Quantization Recovery paper • 4 items • Updated Aug 19