Hy3-295B-A21B-R2fc / README.md
JasonW2025's picture
add model card (gated preview)
fac9230 verified
|
Raw
History Blame Contribute Delete
2.7 kB
metadata
base_model: tencent/Hy3-295B-A21B
library_name: vllm
pipeline_tag: text-generation
license: other
license_name: tencent-hunyuan-community
tags:
  - mixture-of-experts
  - quantized
  - gridbook
  - nvfp4
  - codebook
  - gb10
  - sm_121a
  - dgx-spark
extra_gated_prompt: >-
  This is a preview research build. Access is gated. By requesting access you
  acknowledge it is a GB10/sm_121a-specific serving build that requires a custom
  vLLM plugin, is not a stock-vLLM checkpoint, and inherits all terms of the
  base model tencent/Hy3-295B-A21B.
extra_gated_fields:
  Intended use: text
  I agree to the base model license terms: checkbox

Hy3-295B-A21B — GB10 / sm_121a single-node build (R2fc)

Status: preview, gated. This card is intentionally minimal. Quantization recipe and allocation details are withheld pending an official-release decision. Full benchmark numbers and the sm_121a build details will be added when the model is officially released. Not optimized for public download.

What this is

A single-node serving build of Hy3-295B-A21B (base: tencent/Hy3-295B-A21B) that fits and serves on one NVIDIA DGX Spark (GB10, sm_121a, 121 GB unified memory) under vLLM. Mixture-of-Experts, 80 transformer layers + 1 MTP head, 8-of-192 routed experts.

  • Resident weights: ~86.9 GiB
  • Context: up to ~137k tokens on a single Spark with nothing else resident (model natively supports 262,144)
  • Target hardware: NVIDIA GB10 (sm_121a). This build's serving kernels are compiled for compute_121a and are not portable to other GPUs as-is.

License

Inherits and preserves the license of the base model tencent/Hy3-295B-A21B. All base license terms and use restrictions apply unchanged. See the base model's license.

Serving

Serving requires a custom vLLM plugin (codebook-quantized MoE) plus a config overlay — this is not a stock-vLLM-loadable checkpoint. The corrected config.json is already baked into this repo. In brief:

  • Image: ghcr.io/spark-arena/dgx-vllm-eugr-nightly (sm_121a build)
  • Plugin: the gridbook codebook-MoE plugin + the load-chain patches (X3/X4)
  • Launch: --quantization gridbook, --max-num-batched-tokens 16384, fp8 KV cache

Intended use / limitations

Research and evaluation on GB10-class hardware. Chinese-trained MoE base; response-language behavior follows the base model. Not evaluated for production safety-critical use. Benchmark numbers are withheld from this preview card by design.

Provenance

Built 2026-07 for the DGX Spark single-node target. Quantization method, calibration set, and per-layer allocation are recorded internally and withheld from this preview.