Hugging Face's logo Hugging Face
  • Models
  • Datasets
  • Spaces
  • Buckets new
  • Docs
  • Enterprise
  • Pricing
    • Website
      • Tasks
      • HuggingChat
      • Collections
      • Languages
      • Organizations
    • Community
      • Blog
      • Posts
      • Daily Papers
      • Hardware
      • Learn
      • Discord
      • Forum
      • GitHub
    • Solutions
      • Team & Enterprise
      • Hugging Face PRO
      • Enterprise Support
      • Inference Providers
      • Inference Endpoints
      • Storage Buckets

  • Log In
  • Sign Up

KaedeTai
/
mlx-mtp-graft

MLX
speculative-decoding
multi-token-prediction
mtp
apple-silicon
omlx
mtplx
qwen
qwen3.8
quantization
Model card Files Files and versions
xet
Community

Instructions to use KaedeTai/mlx-mtp-graft with libraries, inference providers, notebooks, and local apps. Follow these links to get started.

  • Libraries
  • MLX

    How to use KaedeTai/mlx-mtp-graft with MLX:

    # Download the model from the Hub
    pip install huggingface_hub[hf_xet]
    
    huggingface-cli download --local-dir mlx-mtp-graft KaedeTai/mlx-mtp-graft
  • Notebooks
  • Google Colab
  • Kaggle
  • Local Apps Settings
  • LM Studio
  • Atomic Chat
mlx-mtp-graft
314 MB
Ctrl+K
Ctrl+K
  • 1 contributor
History: 7 commits
KaedeTai's picture
KaedeTai
major revision: runtime dominates (MTPLX 3.79x vs oMLX 54.6 t/s on the same artifact); void the FP16-vs-4bit conclusion as runtime-confounded; add contract-parameter verification; correct the oMLX depth claim
7161354 verified 5 days ago
  • .gitattributes
    1.52 kB
    initial commit 5 days ago
  • README.md
    12 kB
    major revision: runtime dominates (MTPLX 3.79x vs oMLX 54.6 t/s on the same artifact); void the FP16-vs-4bit conclusion as runtime-confounded; add contract-parameter verification; correct the oMLX depth claim 5 days ago
  • graft_mtp.py
    6.14 kB
    MTP head graft for MLX Qwen3.8-27B quantizations that dropped theirs 5 days ago
  • qwen3.8-27b-mtp-4bit.safetensors
    314 MB
    xet
    MTP head graft for MLX Qwen3.8-27B quantizations that dropped theirs 5 days ago