longchat-7b-starkv-proxima

STAR-KV checkpoint: lmsys/longchat-7b-v1.5-32k with attention K/V weights calibrated + fused into a low-rank form for KV-cache compression, produced by proxima_vllm (STAR-KV method).

What's different from the base model

Standard LongChat weights, plus a star_kv/ directory: weights.safetensors (fused low-rank K/V projections) and star_kv_config.json. config.json architecture set to ProximaStarKVLlamaForCausalLM. Directory loads directly as vLLM model= argument via additional_config.proxima_vllm.checkpoint.

Calibration recipe

  • Base model: lmsys/longchat-7b-v1.5-32k
  • Method: activation collection on K/V projections + distillation into low-rank basis, max_rank_k=64, max_rank_v=64, sharpness=10.0, mixed-precision fusion (top_fraction=0.25, top_bits=4, bottom_bits=3)
  • Calibration data: real text, HuggingFaceFW/fineweb-edu (CC-MAIN-2013-20 shard)
  • Active profile: balanced (gamma=1e-3); aggressive (gamma=5e-3) and conservative (gamma=2e-4) profiles also independently calibrated and included in checkpoint
  • Samples: 64 sequences x 256 tokens, batch size 2, 1 epoch

Deviation from paper (disclosed): paper calibrates on 3000 FineWeb-Edu samples x 2048 tokens (6 GPU-hours, 8B model). This checkpoint uses a scaled-down real-(not-synthetic)-text run (1/50th token volume) for single-session feasibility on one L4 GPU. Results should be read with that caveat — this is not a paper-scale calibration.

Benchmark results (measured, this checkpoint)

End-to-end vs plain vLLM (post-prefill-kernel-fix, 2026-08-06):

  • 4096 ctx: 1.65x slower / 177.3 tok/s, 3.11x more KV blocks used
  • 32768 ctx: 178.6 tok/s (STAR-KV-only run, requires gpu_memory_utilization=0.83, not 0.85)

Kernel-level:

  • Decode: 1.42-1.46x vs FlashAttention-2 at batch>=16
  • Prefill: 0.5-0.86x of FlashAttention-2, but 1.2-1.4x over this repo's own decode kernel

Usage

from vllm import LLM

llm = LLM(
    model="tenosra/longchat-7b-starkv-proxima",
    additional_config={"proxima_vllm": {"checkpoint": True}},
)

Intended use / limitations

Research checkpoint for benchmarking KV-cache compression at long context. Calibration sample count is far below the paper's protocol (see deviation note above) — treat compression/quality tradeoff numbers as directional, not final.

Downloads last month
13
Safetensors
Model size
7B params
Tensor type
F32
·
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for tenosra/longchat-7b-starkv-proxima

Finetuned
(1)
this model