ThinkingCap-Qwen3.8-27B — NInfer Blackwell V1

Production-oriented NInfer export of bottlecapai/ThinkingCap-Qwen3.8-27B, tuned and validated for NVIDIA Blackwell.

V1 status: frozen and production validated.

The final artifact combines:

  • compact 147,590-token public vocabulary
  • 147,712 physical vocabulary rows
  • 131,072-row indexed proposal head
  • native one-layer ThinkingCap MTP
  • NVFP4 target quantization
  • baked activation-divisor tuning
  • NVFP4 KV cache validation in the production benchmark path
  • NInfer upstream GDN/runtime changes
  • production prefill chunk 2048

No runtime DX override environment variables are required for this artifact.

Primary artifact

combined-147590/thinkingcap-v1-upstream-d48dx3p5-oattndx6-ogdndx1-igdndx1-iattndx1-g64dx1-mtpgateupdx15p40-combined-147590.ninfer

Field Value
File size 16308912900 bytes / 15.19 GiB
SHA256 d030303f73b658399810a1be097c3b682610c47df8ea1cd24eb0c1eb53255b27
Conversion sidecar combined-147590/thinkingcap-v1-upstream-d48dx3p5-oattndx6-ogdndx1-igdndx1-iattndx1-g64dx1-mtpgateupdx15p40-combined-147590.ninfer.conversion.json
Public vocabulary 147,590
Physical vocabulary 147,712
Proposal rows 131,072
Components text, mtp
Conversion device CPU
Converter rows/chunk 512

The final V1 artifact was rebuilt directly from the BF16 source checkpoint. Existing .ninfer artifacts were not requantized.

Frozen V1 quantization / activation profile

Family Activation input divisor
D48 — MLP down 3.5
O_GDN — GDN output 1.0
O_ATTN — attention output 6.0
I_GDN — GDN input parent 1.0
I_ATTN — attention input parent 1.0
G64 — MLP gate/up 1.0
MTP GateUp 15.40

The selected layer families remain the previously frozen:

  • D48
  • O44
  • I53
  • G64

Only the activation divisor calibration was re-optimized for the new upstream runtime where required.

Runtime recommendation

Validated production settings:

kv-dtype       = nvfp4
spec           = mtp
draft-tokens   = 5
lm-head-draft  = enabled
prefill-chunk  = 2048

The production binary was built for Blackwell CUDA architecture 120a, Release mode, using CUDA 13.1.2.

Validated local production image:

ninfer-upstream-combined-147k:v1

The corresponding public source is available in:

https://github.com/Xtravaganz/ninfer

Source tag:

thinkingcap-upstream-v1-2026-10-05

V1 source provenance

Item Commit
Published master ba8c4e726834c5e00cf5f68255ea6430c66e5319
Exact tested V1 source tag 883d0e434ecb54cb6d6a9274c6a1a680988fab9d
Compact-vocab/Q8 support b8685532
Runtime NVFP4 divisor hooks 883d0e43

The V1 production binary was proven byte-identical to the binary built from the validated worktree before publication.

Production binary hashes

ninfer
ee41f5be3a4f6b77148420b1bda8cf8fe12479ad8de9be5f9bcc63539557e7e3

ninfer-serve
2137f5c7d3d54343e5769735ca5103585b53d872ffb8422d3758b45bdf75058d

Baked V1 equivalence

The final CPU-built artifact was tested without any runtime DX override environment variables.

Its speculative acceptance matched the tuned runtime model exactly:

Workload Runtime-tuned Final baked V1 Delta
Mixed 8k 0.7509293680 0.7509293680 0
Mixed 32k 0.6134185304 0.6134185304 0
Mixed 48k 0.5483870968 0.5483870968 0
Mixed 110k 0.5864197531 0.5864197531 0

Baked equivalence: PASS

Old production vs V1

End-to-end comparison on the same physical GPU1.

Old production used the previous runtime with prefill chunk 1024. V1 uses the new upstream runtime, re-optimized D48 divisor, and prefill chunk 2048.

Context Acceptance old Acceptance V1 Delta Decode old Decode V1 Decode gain Prefill old Prefill V1 Prefill gain
Mixed 8k 54.02% 75.09% +21.07 pp 82.15 t/s 116.99 t/s +42.41% 3390.4 t/s 3664.2 t/s +8.07%
Mixed 32k 40.91% 61.34% +20.43 pp 64.00 t/s 93.16 t/s +45.56% 2866.7 t/s 3027.2 t/s +5.60%
Mixed 48k 55.33% 54.84% -0.49 pp 75.78 t/s 83.33 t/s +9.96% 2551.6 t/s 2707.2 t/s +6.10%
Mixed 110k 54.52% 58.64% +4.12 pp 67.55 t/s 76.57 t/s +13.35% 1798.4 t/s 1919.2 t/s +6.72%

This is an end-to-end production-stack comparison, not an isolated measurement of one kernel or one quantization parameter.

Why prefill chunk 2048

On the final V1 runtime, chunk 2048 also beat chunk 1024 directly:

Context Acc. 1024 Acc. 2048 Delta Decode gain Prefill gain
Mixed 32k 55.29% 61.34% +6.05 pp +6.28% +4.89%
Mixed 110k 43.53% 58.64% +15.11 pp +22.45% +4.79%

For this V1, prefill-chunk=2048 is therefore part of the validated production configuration.

Compact vocabulary

The source tokenizer exposes 248,077 public tokens with a 248,320-row aligned NInfer embedding/output head.

V1 uses the validated combined Script + CJK cut:

Original public vocabulary     248,077
V1 public vocabulary           147,590
V1 physical vocabulary         147,712
Removed                        100,487
Reduction                       40.51%
Padding rows                       122

The compact vocabulary was produced conservatively using:

  • Unicode/script classification
  • workload protection
  • BPE dependency closure
  • special/added-token protection
  • tokenizer parity checks
  • teacher-output audits
  • runtime A/B validation

The MTP proposal vocabulary remains 131,072 rows and is independent from the full target vocabulary cut.

Benchmark corpus

The V1 runtime tuning used mixed-context benchmark corpora intended to better match real coding/agent workloads than pure text-only or code-only streams.

Primary validation points:

  • 8,192-token mixed context
  • 32,768-token mixed context
  • 49,152-token mixed context
  • 112,640-token mixed context

The 112,640-token test is intended to approximate the long-context summary/compaction regime.

These benchmark-corpus acceptance values should not be interpreted as universal chat or coding acceptance rates.

Reproducibility

Final model:

thinkingcap-v1-upstream-d48dx3p5-oattndx6-ogdndx1-igdndx1-iattndx1-g64dx1-mtpgateupdx15p40-combined-147590.ninfer
SHA256: d030303f73b658399810a1be097c3b682610c47df8ea1cd24eb0c1eb53255b27

Conversion sidecar SHA256:

62eff698b11f1d393b518cf62937f918dc131e6893f65ec9cb57d5c5d8a9843b

Exact V1 source:

git clone https://github.com/Xtravaganz/ninfer.git
cd ninfer
git checkout thinkingcap-upstream-v1-2026-10-05

The V1 tag is intentionally kept on the exact tested source commit, while master also contains the repository's previous origin history.

Repository layout

combined-147590/
  thinkingcap-v1-upstream-d48dx3p5-oattndx6-ogdndx1-igdndx1-iattndx1-g64dx1-mtpgateupdx15p40-combined-147590.ninfer
  thinkingcap-v1-upstream-d48dx3p5-oattndx6-ogdndx1-igdndx1-iattndx1-g64dx1-mtpgateupdx15p40-combined-147590.ninfer.conversion.json
  thinkingcap-v1-upstream-d48dx3p5-oattndx6-ogdndx1-igdndx1-iattndx1-g64dx1-mtpgateupdx15p40-combined-147590.ninfer.sha256

provenance/
  V1-PRODUCTION.md

benchmarks/
  V1-RESULTS.md

MODEL_FILE.txt
FACTS.json
PROVENANCE.md
SHA256SUMS
README.md

Older RC/experimental artifacts are retained for historical comparison. The artifact listed in MODEL_FILE.txt is the current V1 production model.

License

See LICENSE.

Downloads last month
570
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Schestex/ThinkingCap-Qwen3.8-27B-NInfer

Base model

Qwen/Qwen3.8-27B
Quantized
(33)
this model