Hugging Face's logo Hugging Face
  • Models
  • Datasets
  • Spaces
  • Buckets new
  • Docs
  • Enterprise
  • Pricing
    • Website
      • Tasks
      • HuggingChat
      • Collections
      • Languages
      • Organizations
    • Community
      • Blog
      • Posts
      • Daily Papers
      • Hardware
      • Learn
      • Discord
      • Forum
      • GitHub
    • Solutions
      • Team & Enterprise
      • Hugging Face PRO
      • Enterprise Support
      • Inference Providers
      • Inference Endpoints
      • Storage Buckets

  • Log In
  • Sign Up
hexgrid-cloud 's Collections
Best Open-Source Coding LLMs for Private Deployment
Production-Ready Quantized Chat LLMs — 4-bit & 8-bit
Open Source RAG Stack — Embed + Rerank + Generate
One-click LLM deployments on Private GPU

Production-Ready Quantized Chat LLMs — 4-bit & 8-bit

updated Jun 7

FP8, AWQ-4Bit and W8A8 quantized versions of popular models. Lower VRAM, same production quality. Deploy at hexgrid.cloud in one click.

Upvote
-

  • cyankiwi/Qwen3.5-9B-AWQ-4bit

    Image-Text-to-Text • 10B • Updated 11 days ago • 363k • 36

  • lovedheart/Qwen3.5-9B-FP8

    Image-Text-to-Text • 10B • Updated Mar 2 • 41.1k • 14

  • cyankiwi/Qwen3.5-27B-AWQ-4bit

    Image-Text-to-Text • 29B • Updated 11 days ago • 425k • 42

  • RedHatAI/gemma-4-31B-it-FP8-block

    Image-Text-to-Text • 31B • Updated 2 days ago • 4.95M • 44

  • QuantTrio/gemma-4-31B-it-AWQ

    Image-Text-to-Text • 31B • Updated 10 days ago • 252k • 14

  • nvidia/Llama-3.3-70B-Instruct-FP8

    71B • Updated Aug 22, 2025 • 126k • 27

  • ibnzterrell/Meta-Llama-3.3-70B-Instruct-AWQ-INT4

    Text Generation • 71B • Updated Dec 7, 2024 • 87.8k • 30

  • nvidia/Llama-3.1-8B-Instruct-FP8

    Text Generation • 8B • Updated Aug 22, 2025 • 75k • • 36

  • hugging-quants/Meta-Llama-3.1-8B-Instruct-AWQ-INT4

    Text Generation • 8B • Updated Aug 7, 2024 • 154k • 92
Upvote
-
  • Collection guide
  • Browse collections
Company
TOS Privacy About Careers
Website
Models Datasets Spaces Pricing Docs