Instructions to use pragmaticcs/PentaCoder-9B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use pragmaticcs/PentaCoder-9B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="pragmaticcs/PentaCoder-9B") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("pragmaticcs/PentaCoder-9B") model = AutoModelForCausalLM.from_pretrained("pragmaticcs/PentaCoder-9B", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use pragmaticcs/PentaCoder-9B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "pragmaticcs/PentaCoder-9B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "pragmaticcs/PentaCoder-9B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/pragmaticcs/PentaCoder-9B
- SGLang
How to use pragmaticcs/PentaCoder-9B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "pragmaticcs/PentaCoder-9B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "pragmaticcs/PentaCoder-9B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "pragmaticcs/PentaCoder-9B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "pragmaticcs/PentaCoder-9B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use pragmaticcs/PentaCoder-9B with Docker Model Runner:
docker model run hf.co/pragmaticcs/PentaCoder-9B
- Architectural Specifications
- Composition & Donor Weighting
- Merge Methodology & Mathematical Formulation
- 1. Task Vector Formulation
- 2. Hybrid-Aware Continuous Depth Modulation
- 3. Dynamic Gram-Matrix Task De-Biasing
- 4. Asymmetric GQA Disentangled QKV Slicing
- 5. Neuron-Coherent Row Gating
- 6. Heavy-Tailed DELLA Adaptive Rescaling
- 7. Coordinate-Wise Sign Consensus & Model Stock Scaling
- 8. Spectral Norm & Attention Entropy Anchoring
- 1. Task Vector Formulation
- Layer-Stratified Component Policies
- Agentic Chat Template & Operational Directives
- Recommended Generation Parameters
- How to Use
- Citation & References
Most sub-10B coding models fail in realistic agentic environments due to a common trade-off: aggressive fine-tuning on synthetic coding instructions improves immediate benchmark pass rates but introduces brittle syntactic degradation and catastrophic repetition loops when shell commands or compiler checks fail.
PentaCoder addresses these limitations by uniting five specialized post-trained checkpoints of Qwen 3.5 9B via GeoDELLA (Geometric Drop-and-Rescale with Task-Covariance De-Biasing and Spectral Norm Anchoring). The model synthesizes the distinct mathematical distributions of each donor:
- Algorithmic Correctness & Architectural Decomposition from Qwopus3.5-Coder (Claude 3.5 Opus distillation trajectories).
- Multi-Turn SWE-bench Planning & Tool Protocol Integrity from MiMo-V2.6-Distill (Large-scale agentic execution traces).
- Terminal Execution Discipline & Error-Recovery Heuristics from Ornith-1.5 (Reinforcement learning for anti-looping).
- Low-Level Systems Implementation & Runtime Robustness from OxCoder (Deep API, CLI, and operational coding specialization).
- Abstract Structural Reasoning & Syntax Grounding from Qwen3.8-Distill (Distilled frontier chain-of-thought representations).
The result is a lean, blisteringly fast 9B pure-text causal engine with a native 256k context window that runs comfortably on consumer GPUs.
Contents
- Architectural Specifications
- Composition & Donor Weighting
- Merge Methodology & Mathematical Formulation
- Layer-Stratified Component Policies
- Agentic Chat Template & Operational Directives
- Recommended Generation Parameters
- How to Use
- Citation & References
Architectural Specifications
| Parameter | Specification |
|---|---|
| Total Parameters | 8.8B (Pure Text Backbone) |
| Architecture Type | Hybrid Recurrent-Attention Causal LM (qwen3_5_text) |
| Hidden Dimension (dmodel) | 4096 |
| Intermediate Dimension (dmlp) | 12288 (SwiGLU) |
| Decoder Layers | 32 |
| Attention Layout | 8 Blocks × (3 Gated DeltaNet Linear Layers : 1 Gated Softmax Layer) |
| Full Attention Layers | Layers 3, 7, 11, 15, 19, 23, 27, 31 |
| Linear Attention Configuration | 16 Key Heads / 32 Value Heads (dk = dv = 128) |
| Full Attention Configuration | 16 Query Heads / 4 Key-Value Heads (GQA, dh = 256) |
| Rotary Position Embedding (RoPE) | 1D Partial RoPE (θ = 10⁷, Factor = 0.25 → 64 dimensions) |
| Context Window Length | 262,144 tokens (256k) |
| Native Precision | bfloat16 |
| Vocabulary Size | 248,320 (Padded for Fill-In-The-Middle and Tool Tokens) |
Composition & Donor Weighting
The foundational weights of Qwen/Qwen3.5-9B serve as the shared topological base (W₀). Five specialized donor checkpoints provide non-overlapping task vectors mapped across normalized layer depth u ∈ [0, 1]:
| Model | Primary Focus | Depth Target |
|---|---|---|
| Qwen/Qwen3.5-9B | Base pre-trained manifold and state-space anchors | Global (W₀) |
| empero-ai/Qwen3.8-9B-Distill | Abstract token synthesis, reasoning structure, syntax | Lower & Mid Decoders |
| Jackrong/Qwopus3.5-9B-Coder | Typing discipline, algorithm design, functional purity | Mid Decoders (Bell Curve) |
| ornith-ai/Ornith-1.5-9B | Circuit-breaker recovery, environment feedback integration | Upper-Mid Decoders |
| OrionLLM/OxCoder-9B | Systems engineering, runtime bug localization, CLI tooling | Deep Layers (Ascending Ramp) |
| XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B | SWE-bench multi-step planning, long-context tool invocation | Broad Central Decoders |
Merge Methodology & Mathematical Formulation
The merge was executed using the GeoDELLA-Coder pipeline, which solves weight interference through dynamic task-correlation de-biasing, row-wise sign consensus, sub-tensor GQA disentanglement, and power-iteration spectral norm stabilization.
1. Task Vector Formulation
For each donor checkpoint k ∈ {1, ..., 5}, the parameter update delta τk is computed relative to the base anchor W₀:
2. Hybrid-Aware Continuous Depth Modulation
Task vector mixing coefficients are continuously modulated over normalized depth u = l / (L - 1), where l ∈ [0, 31] and L = 32:
The initial coefficients are normalized to form a partition of unity across all layers:
3. Dynamic Gram-Matrix Task De-Biasing
Fine-tuned models frequently share underlying distillation datasets, leading to collinearity that can drown out specialized task vectors. To correct for this over-representation, the empirical Gram correlation matrix G ∈ ℝK × K is evaluated per tensor:
A uniqueness coefficient Uk is derived from the inverse column sum of task vector correlations:
The dynamic weights are then re-balanced and normalized:
This prevents dataset overlap from suppressing distinct algorithmic representations.
4. Asymmetric GQA Disentangled QKV Slicing
Qwen 3.5 9B features asymmetric head ratios in both its linear attention and full attention layers. Standard fused tensor merging causes cross-head pollution by treating routing projections identically to memory projections.
Fused projection tensors are sliced into their functional sub-matrices prior to merging:
- Gated DeltaNet Layers (8192 × dmodel): Sliced into Query (2048), Key (2048), and Value (4096).
- Gated Attention Layers (6144 × dmodel): Sliced into Query (4096), Key (1024), and Value (1024).
Query and Key slices are processed with a conservative retention density (ρ = 0.95) to preserve sharp context routing. Value matrices are processed with adaptive MLP density (ρ = 0.70) to maximize conceptual synthesis. The components are then re-concatenated along the head dimension.
5. Neuron-Coherent Row Gating
To eliminate destructive interference in Feed-Forward Networks (MLPs), task vectors are gated at the single-neuron (row) level:
For each row r of donor delta τk, the directional alignment with the consensus mean is evaluated:
Rows exhibiting severe directional opposition (cos θk, r < -0.10) are masked out:
This eliminates opposing gradient vectors that produce incoherent syntax generation.
6. Heavy-Tailed DELLA Adaptive Rescaling
Surviving parameters undergo non-linear magnitude-based sampling. Using parameter rank indices Rk, ij ∈ [0, 1] sorted by absolute magnitude, a Pareto-style retention probability pk, ij is established:
Parameters are sampled via a Bernoulli trial and rescaled by their inverse survival probability:
7. Coordinate-Wise Sign Consensus & Model Stock Scaling
Directional consensus is determined via weighted sign agreement:
The aggregated delta is projected onto the non-linear manifold using the Model Stock analytic scaling factor t*:
8. Spectral Norm & Attention Entropy Anchoring
Attention projection matrices (Wattn) are vulnerable to spectral explosion during merges, which contracts attention entropy and leads to repetitive generation loops.
The dominant singular value σ(W) is calculated via a three-iteration deterministic power iteration:
If the merged spectral radius grows more than 5% relative to the base model, it is scaled down:
Layer-Stratified Component Policies
| Parameter Class | Target Identifiers | Applied Policy | Density (ρ) | Mathematical Constraints |
|---|---|---|---|---|
| Embeddings & LM Head | embed_tokens, lm_head |
Low-Memory Streaming Blend | 1.0 | Convex iterative accumulation; vocab dimension aligned to 248,320. |
| Linear State-Space Projections | linear_attn.in_proj_qkv |
Asymmetric GQA DELLA | 0.95 (QK) / 0.70 (V) | Sub-tensor slicing; separate routing and associative memory passes. |
| Self-Attention Projections | self_attn.qkv_proj, o_proj |
Asymmetric GQA + Spectral Anchor | 0.95 (QK) / 0.70 (V) | Power-iteration clipping prevents σ > 1.05 σ₀. |
| Feed-Forward Blocks | mlp.gate_proj, up_proj, down_proj |
Neuron-Gated HG-DELLA | 0.50 – 0.70 | Row-wise cosine filtering (cos θ ≥ -0.10); sign consensus. |
| Recurrent Gates & Normalization | A_log, norm, conv1d |
Convex Parameter Blend | 1.0 | Preservation of Alog ≤ 0 to guarantee BIBO stability. |
Agentic Chat Template & Operational Directives
This model uses the Improved Chat Template for Qwen 3.x by Olivia Rossi to support multi-tier Chain-of-Thought (CoT) reasoning, dual-format agentic tool execution, automatic error-recovery heuristics, and strict token-waste elimination.
Recommended Generation Parameters
For deterministic software engineering and complex reasoning benchmarks:
| Parameter | Recommended Value | Description |
|---|---|---|
| Temperature | 0.6 |
Balances strict syntactic validity with algorithmic path exploration. |
| Top-P | 0.95 |
Eliminates low-probability token tails while preserving alternative logic paths. |
| Top-K | 20 |
Restricts token candidate pools to prevent architectural syntax drift. |
| Min-P | 0.0 (Off) |
Disabled in favor of explicit Top-K / Top-P governance. |
| Repetition Penalty | 1.0 (Off) |
Disabled to prevent syntax degradation in repetitive code patterns (indentation, braces). |
| Presence Penalty | 0.0 |
Prevents naming mutations across long-context symbol resolution. |
How to Use
Serving via vLLM
vllm serve pragmaticcs/PentaCoder-9B \
--dtype bfloat16 \
--max-model-len 65536 \
--gpu-memory-utilization 0.95 \
--enable-prefix-caching \
--enable-auto-tool-choice \
--tool-call-parser qwen3_coder \
--enable-reasoning \
--reasoning-parser qwen3
Inference via Transformers
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "pragmaticcs/PentaCoder-9B"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id, torch_dtype=torch.bfloat16, device_map="auto"
)
messages = [
{
"role": "system",
"content": "You are an expert systems engineer. Reason step by step and output clean, robust implementations.",
},
{
"role": "user",
"content": "Write an asynchronous connection pool in Python for TCP sockets with active health-checking, backpressure control, and graceful shutdown handling.",
},
]
inputs = tokenizer.apply_chat_template(
messages, add_generation_prompt=True, enable_thinking=True, return_tensors="pt"
).to(model.device)
outputs = model.generate(
inputs,
max_new_tokens=4096,
temperature=0.6,
top_p=0.95,
top_k=20,
do_sample=True,
)
response = tokenizer.decode(outputs[0][inputs.shape[-1] :], skip_special_tokens=True)
print(response)
Citation & References
- Qwen/Qwen3.5-9B
- XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B
- ornith-ai/Ornith-1.5-9B
- Jackrong/Qwopus3.5-9B-Coder
- OrionLLM/OxCoder-9B
- empero-ai/Qwen3.8-9B-Distill
- Improved Chat Template for Qwen 3.x
@inproceedings{yadav2023ties,
title={Resolving Interference When Merging Models},
author={Yadav, Prateek and Tam, Derek and Choshen, Leshem and Raffel, Colin and Bansal, Mohit},
booktitle={Advances in Neural Information Processing Systems (NeurIPS)},
volume={36},
pages={7093--7115},
year={2023}
}
@article{deep2024della,
title={DELLA-Merging: Reducing Interference in Model Merging through Magnitude-Based Sampling},
author={Deep, Pala Tej and Bhardwaj, Rishabh and Poria, Soujanya},
journal={arXiv preprint arXiv:2406.11617},
year={2024}
}
@article{jang2024modelstock,
title={Model Stock: All We Need Is just a Few Fine-Tuned Models},
author={Jang, Dong-Hwan and Yoon, Sang-Doo and Song, Gyeong-Moon},
journal={arXiv preprint arXiv:2403.19522},
year={2024}
}
- Downloads last month
- 293

