Hugging Face's logo Hugging Face
  • Models
  • Datasets
  • Spaces
  • Buckets new
  • Docs
  • Enterprise
  • Pricing
    • Website
      • Tasks
      • HuggingChat
      • Collections
      • Languages
      • Organizations
    • Community
      • Blog
      • Posts
      • Daily Papers
      • Hardware
      • Learn
      • Discord
      • Forum
      • GitHub
    • Solutions
      • Team & Enterprise
      • Hugging Face PRO
      • Enterprise Support
      • Inference Providers
      • Inference Endpoints
      • Storage Buckets

  • Log In
  • Sign Up

pearsonkyle
/
gemma4-e4b-w4a16-mtp-vLLM

Text Generation
Safetensors
vllm
quantization
w4a16
gptq
compressed-tensors
tool-use
agentic
speculative-decoding
mtp
Model card Files Files and versions
xet
Community
gemma4-e4b-w4a16-mtp-vLLM / deploy
2.41 kB
Ctrl+K
Ctrl+K
  • 1 contributor
History: 1 commit
pearsonkyle's picture
pearsonkyle
gemma-4-E4B-it W4A16 (full-precision lm_head), calibrated on our own CLI/tool-call logs, + Gemma-4 MTP drafter for speculative decoding. ~130 tok/s single-stream on one RTX 4060 Ti, stock vLLM 0.26.0, correct tool calls, no vocab pruning.
c25f58d verified 1 day ago
  • launch.sh
    1.64 kB
    gemma-4-E4B-it W4A16 (full-precision lm_head), calibrated on our own CLI/tool-call logs, + Gemma-4 MTP drafter for speculative decoding. ~130 tok/s single-stream on one RTX 4060 Ti, stock vLLM 0.26.0, correct tool calls, no vocab pruning. 1 day ago
  • requirements.txt
    776 Bytes
    gemma-4-E4B-it W4A16 (full-precision lm_head), calibrated on our own CLI/tool-call logs, + Gemma-4 MTP drafter for speculative decoding. ~130 tok/s single-stream on one RTX 4060 Ti, stock vLLM 0.26.0, correct tool calls, no vocab pruning. 1 day ago