Instructions to use Schestex/ThinkingCap-Qwen3.8-27B-NInfer with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- NInfer
How to use Schestex/ThinkingCap-Qwen3.8-27B-NInfer with NInfer:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
ThinkingCap-Qwen3.8-27B — NInfer Blackwell V1
Production-oriented NInfer export of
bottlecapai/ThinkingCap-Qwen3.8-27B, tuned and validated for NVIDIA Blackwell.
V1 status: frozen and production validated.
The final artifact combines:
- compact 147,590-token public vocabulary
- 147,712 physical vocabulary rows
- 131,072-row indexed proposal head
- native one-layer ThinkingCap MTP
- NVFP4 target quantization
- baked activation-divisor tuning
- NVFP4 KV cache validation in the production benchmark path
- NInfer upstream GDN/runtime changes
- production prefill chunk 2048
No runtime DX override environment variables are required for this artifact.
Primary artifact
combined-147590/thinkingcap-v1-upstream-d48dx3p5-oattndx6-ogdndx1-igdndx1-iattndx1-g64dx1-mtpgateupdx15p40-combined-147590.ninfer
| Field | Value |
|---|---|
| File size | 16308912900 bytes / 15.19 GiB |
| SHA256 | d030303f73b658399810a1be097c3b682610c47df8ea1cd24eb0c1eb53255b27 |
| Conversion sidecar | combined-147590/thinkingcap-v1-upstream-d48dx3p5-oattndx6-ogdndx1-igdndx1-iattndx1-g64dx1-mtpgateupdx15p40-combined-147590.ninfer.conversion.json |
| Public vocabulary | 147,590 |
| Physical vocabulary | 147,712 |
| Proposal rows | 131,072 |
| Components | text, mtp |
| Conversion device | CPU |
| Converter rows/chunk | 512 |
The final V1 artifact was rebuilt directly from the BF16 source checkpoint.
Existing .ninfer artifacts were not requantized.
Frozen V1 quantization / activation profile
| Family | Activation input divisor |
|---|---|
| D48 — MLP down | 3.5 |
| O_GDN — GDN output | 1.0 |
| O_ATTN — attention output | 6.0 |
| I_GDN — GDN input parent | 1.0 |
| I_ATTN — attention input parent | 1.0 |
| G64 — MLP gate/up | 1.0 |
| MTP GateUp | 15.40 |
The selected layer families remain the previously frozen:
- D48
- O44
- I53
- G64
Only the activation divisor calibration was re-optimized for the new upstream runtime where required.
Runtime recommendation
Validated production settings:
kv-dtype = nvfp4
spec = mtp
draft-tokens = 5
lm-head-draft = enabled
prefill-chunk = 2048
The production binary was built for Blackwell CUDA architecture 120a,
Release mode, using CUDA 13.1.2.
Validated local production image:
ninfer-upstream-combined-147k:v1
The corresponding public source is available in:
https://github.com/Xtravaganz/ninfer
Source tag:
thinkingcap-upstream-v1-2026-10-05
V1 source provenance
| Item | Commit |
|---|---|
Published master |
ba8c4e726834c5e00cf5f68255ea6430c66e5319 |
| Exact tested V1 source tag | 883d0e434ecb54cb6d6a9274c6a1a680988fab9d |
| Compact-vocab/Q8 support | b8685532 |
| Runtime NVFP4 divisor hooks | 883d0e43 |
The V1 production binary was proven byte-identical to the binary built from the validated worktree before publication.
Production binary hashes
ninfer
ee41f5be3a4f6b77148420b1bda8cf8fe12479ad8de9be5f9bcc63539557e7e3
ninfer-serve
2137f5c7d3d54343e5769735ca5103585b53d872ffb8422d3758b45bdf75058d
Baked V1 equivalence
The final CPU-built artifact was tested without any runtime DX override environment variables.
Its speculative acceptance matched the tuned runtime model exactly:
| Workload | Runtime-tuned | Final baked V1 | Delta |
|---|---|---|---|
| Mixed 8k | 0.7509293680 | 0.7509293680 | 0 |
| Mixed 32k | 0.6134185304 | 0.6134185304 | 0 |
| Mixed 48k | 0.5483870968 | 0.5483870968 | 0 |
| Mixed 110k | 0.5864197531 | 0.5864197531 | 0 |
Baked equivalence: PASS
Old production vs V1
End-to-end comparison on the same physical GPU1.
Old production used the previous runtime with prefill chunk 1024. V1 uses the new upstream runtime, re-optimized D48 divisor, and prefill chunk 2048.
| Context | Acceptance old | Acceptance V1 | Delta | Decode old | Decode V1 | Decode gain | Prefill old | Prefill V1 | Prefill gain |
|---|---|---|---|---|---|---|---|---|---|
| Mixed 8k | 54.02% | 75.09% | +21.07 pp | 82.15 t/s | 116.99 t/s | +42.41% | 3390.4 t/s | 3664.2 t/s | +8.07% |
| Mixed 32k | 40.91% | 61.34% | +20.43 pp | 64.00 t/s | 93.16 t/s | +45.56% | 2866.7 t/s | 3027.2 t/s | +5.60% |
| Mixed 48k | 55.33% | 54.84% | -0.49 pp | 75.78 t/s | 83.33 t/s | +9.96% | 2551.6 t/s | 2707.2 t/s | +6.10% |
| Mixed 110k | 54.52% | 58.64% | +4.12 pp | 67.55 t/s | 76.57 t/s | +13.35% | 1798.4 t/s | 1919.2 t/s | +6.72% |
This is an end-to-end production-stack comparison, not an isolated measurement of one kernel or one quantization parameter.
Why prefill chunk 2048
On the final V1 runtime, chunk 2048 also beat chunk 1024 directly:
| Context | Acc. 1024 | Acc. 2048 | Delta | Decode gain | Prefill gain |
|---|---|---|---|---|---|
| Mixed 32k | 55.29% | 61.34% | +6.05 pp | +6.28% | +4.89% |
| Mixed 110k | 43.53% | 58.64% | +15.11 pp | +22.45% | +4.79% |
For this V1, prefill-chunk=2048 is therefore part of the validated
production configuration.
Compact vocabulary
The source tokenizer exposes 248,077 public tokens with a 248,320-row aligned NInfer embedding/output head.
V1 uses the validated combined Script + CJK cut:
Original public vocabulary 248,077
V1 public vocabulary 147,590
V1 physical vocabulary 147,712
Removed 100,487
Reduction 40.51%
Padding rows 122
The compact vocabulary was produced conservatively using:
- Unicode/script classification
- workload protection
- BPE dependency closure
- special/added-token protection
- tokenizer parity checks
- teacher-output audits
- runtime A/B validation
The MTP proposal vocabulary remains 131,072 rows and is independent from the full target vocabulary cut.
Benchmark corpus
The V1 runtime tuning used mixed-context benchmark corpora intended to better match real coding/agent workloads than pure text-only or code-only streams.
Primary validation points:
- 8,192-token mixed context
- 32,768-token mixed context
- 49,152-token mixed context
- 112,640-token mixed context
The 112,640-token test is intended to approximate the long-context summary/compaction regime.
These benchmark-corpus acceptance values should not be interpreted as universal chat or coding acceptance rates.
Reproducibility
Final model:
thinkingcap-v1-upstream-d48dx3p5-oattndx6-ogdndx1-igdndx1-iattndx1-g64dx1-mtpgateupdx15p40-combined-147590.ninfer
SHA256: d030303f73b658399810a1be097c3b682610c47df8ea1cd24eb0c1eb53255b27
Conversion sidecar SHA256:
62eff698b11f1d393b518cf62937f918dc131e6893f65ec9cb57d5c5d8a9843b
Exact V1 source:
git clone https://github.com/Xtravaganz/ninfer.git
cd ninfer
git checkout thinkingcap-upstream-v1-2026-10-05
The V1 tag is intentionally kept on the exact tested source commit, while
master also contains the repository's previous origin history.
Repository layout
combined-147590/
thinkingcap-v1-upstream-d48dx3p5-oattndx6-ogdndx1-igdndx1-iattndx1-g64dx1-mtpgateupdx15p40-combined-147590.ninfer
thinkingcap-v1-upstream-d48dx3p5-oattndx6-ogdndx1-igdndx1-iattndx1-g64dx1-mtpgateupdx15p40-combined-147590.ninfer.conversion.json
thinkingcap-v1-upstream-d48dx3p5-oattndx6-ogdndx1-igdndx1-iattndx1-g64dx1-mtpgateupdx15p40-combined-147590.ninfer.sha256
provenance/
V1-PRODUCTION.md
benchmarks/
V1-RESULTS.md
MODEL_FILE.txt
FACTS.json
PROVENANCE.md
SHA256SUMS
README.md
Older RC/experimental artifacts are retained for historical comparison.
The artifact listed in MODEL_FILE.txt is the current V1 production model.
License
See LICENSE.
- Downloads last month
- 570