fix(rope): YaRN attention_factor is (0.1*ln(factor)+1)*attn_factor, not the bare multiplier d3cca23 verified joerowell commited on 29 days ago
Use chat_template.jinja as the single source: drop the {% include %} chat_template field from tokenizer_config.json ad2b974 verified joerowell commited on Jul 1
Note FP8 KV cache needs vLLM 0.22.0; drop scrambled-output workaround (vllm#42650) 571346d joerowell commited on Jun 5
Drop VLLM_USE_DEEP_GEMM=0 from vllm serve recipe (DeepGEMM is supported on Hopper and datacenter Blackwell) 514daf4 verified joerowell commited on May 20
Enable thinking by default in non-Hopper FP8-KV serve command 62a5860 verified joerowell commited on Apr 29
Update non-Hopper FP8-KV serve command and link to vLLM recipes page 92f8b44 verified joerowell commited on Apr 29