PentaCoder-9B

License Library Merge Method Architecture Attention Context

Most sub-10B coding models fail in realistic agentic environments due to a common trade-off: aggressive fine-tuning on synthetic coding instructions improves immediate benchmark pass rates but introduces brittle syntactic degradation and catastrophic repetition loops when shell commands or compiler checks fail.

PentaCoder addresses these limitations by uniting five specialized post-trained checkpoints of Qwen 3.5 9B via GeoDELLA (Geometric Drop-and-Rescale with Task-Covariance De-Biasing and Spectral Norm Anchoring). The model synthesizes the distinct mathematical distributions of each donor:

  • Algorithmic Correctness & Architectural Decomposition from Qwopus3.5-Coder (Claude 3.5 Opus distillation trajectories).
  • Multi-Turn SWE-bench Planning & Tool Protocol Integrity from MiMo-V2.6-Distill (Large-scale agentic execution traces).
  • Terminal Execution Discipline & Error-Recovery Heuristics from Ornith-1.5 (Reinforcement learning for anti-looping).
  • Low-Level Systems Implementation & Runtime Robustness from OxCoder (Deep API, CLI, and operational coding specialization).
  • Abstract Structural Reasoning & Syntax Grounding from Qwen3.8-Distill (Distilled frontier chain-of-thought representations).

The result is a lean, blisteringly fast 9B pure-text causal engine with a native 256k context window that runs comfortably on consumer GPUs.


Contents


Architectural Specifications

Parameter Specification
Total Parameters 8.8B (Pure Text Backbone)
Architecture Type Hybrid Recurrent-Attention Causal LM (qwen3_5_text)
Hidden Dimension (dmodel) 4096
Intermediate Dimension (dmlp) 12288 (SwiGLU)
Decoder Layers 32
Attention Layout 8 Blocks × (3 Gated DeltaNet Linear Layers : 1 Gated Softmax Layer)
Full Attention Layers Layers 3, 7, 11, 15, 19, 23, 27, 31
Linear Attention Configuration 16 Key Heads / 32 Value Heads (dk = dv = 128)
Full Attention Configuration 16 Query Heads / 4 Key-Value Heads (GQA, dh = 256)
Rotary Position Embedding (RoPE) 1D Partial RoPE (θ = 10⁷, Factor = 0.25 → 64 dimensions)
Context Window Length 262,144 tokens (256k)
Native Precision bfloat16
Vocabulary Size 248,320 (Padded for Fill-In-The-Middle and Tool Tokens)

Composition & Donor Weighting

The foundational weights of Qwen/Qwen3.5-9B serve as the shared topological base (W₀). Five specialized donor checkpoints provide non-overlapping task vectors mapped across normalized layer depth u ∈ [0, 1]:

Model Primary Focus Depth Target
Qwen/Qwen3.5-9B Base pre-trained manifold and state-space anchors Global (W₀)
empero-ai/Qwen3.8-9B-Distill Abstract token synthesis, reasoning structure, syntax Lower & Mid Decoders
Jackrong/Qwopus3.5-9B-Coder Typing discipline, algorithm design, functional purity Mid Decoders (Bell Curve)
ornith-ai/Ornith-1.5-9B Circuit-breaker recovery, environment feedback integration Upper-Mid Decoders
OrionLLM/OxCoder-9B Systems engineering, runtime bug localization, CLI tooling Deep Layers (Ascending Ramp)
XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B SWE-bench multi-step planning, long-context tool invocation Broad Central Decoders

Merge Methodology & Mathematical Formulation

The merge was executed using the GeoDELLA-Coder pipeline, which solves weight interference through dynamic task-correlation de-biasing, row-wise sign consensus, sub-tensor GQA disentanglement, and power-iteration spectral norm stabilization.

1. Task Vector Formulation

For each donor checkpoint k ∈ {1, ..., 5}, the parameter update delta τk is computed relative to the base anchor W₀:

τk=Dk−W0,k∈{Qwen3.8,Qwopus,Ornith,OxCoder,MiMo} \tau_k = D_k - W_0, \quad k \in \{\text{Qwen3.8}, \text{Qwopus}, \text{Ornith}, \text{OxCoder}, \text{MiMo}\}

2. Hybrid-Aware Continuous Depth Modulation

Task vector mixing coefficients are continuously modulated over normalized depth u = l / (L - 1), where l ∈ [0, 31] and L = 32:

uqwen38(u)=0.15+0.35cos⁡2(π2u)+0.15sin⁡2(πu) u_{\text{qwen38}}(u) = 0.15 + 0.35 \cos^2\left(\frac{\pi}{2} u\right) + 0.15 \sin^2(\pi u)

uqwopus(u)=0.05+0.35sin⁡2(πu) u_{\text{qwopus}}(u) = 0.05 + 0.35 \sin^2(\pi u)

uornith(u)=0.05+0.35sin⁡2(π2u) u_{\text{ornith}}(u) = 0.05 + 0.35 \sin^2\left(\frac{\pi}{2} u\right)

uoxcoder(u)=0.05+0.25sin⁡2(π2u) u_{\text{oxcoder}}(u) = 0.05 + 0.25 \sin^2\left(\frac{\pi}{2} u\right)

umimo(u)=0.05+0.15sin⁡(πu) u_{\text{mimo}}(u) = 0.05 + 0.15 \sin(\pi u)

The initial coefficients are normalized to form a partition of unity across all layers:

wk(l)=uk(u)∑j=15uj(u),∑k=15wk(l)=1.0 w_k(l) = \frac{u_k(u)}{\sum_{j=1}^5 u_j(u)}, \quad \sum_{k=1}^5 w_k(l) = 1.0

3. Dynamic Gram-Matrix Task De-Biasing

Fine-tuned models frequently share underlying distillation datasets, leading to collinearity that can drown out specialized task vectors. To correct for this over-representation, the empirical Gram correlation matrix G ∈ ℝK × K is evaluated per tensor:

Gij=∣⟨vec(τi),vec(τj)⟩∣∥τi∥2∥τj∥2 G_{ij} = \frac{|\langle \text{vec}(\tau_i), \text{vec}(\tau_j) \rangle|}{\|\tau_i\|_2 \|\tau_j\|_2}

A uniqueness coefficient Uk is derived from the inverse column sum of task vector correlations:

Uk=1∑j=1KGkj U_k = \frac{1}{\sum_{j=1}^K G_{kj}}

The dynamic weights are then re-balanced and normalized:

w~k=wkUk∑j=1KwjUj \tilde{w}_k = \frac{w_k U_k}{\sum_{j=1}^K w_j U_j}

This prevents dataset overlap from suppressing distinct algorithmic representations.

4. Asymmetric GQA Disentangled QKV Slicing

Qwen 3.5 9B features asymmetric head ratios in both its linear attention and full attention layers. Standard fused tensor merging causes cross-head pollution by treating routing projections identically to memory projections.

Fused projection tensors are sliced into their functional sub-matrices prior to merging:

  • Gated DeltaNet Layers (8192 × dmodel): Sliced into Query (2048), Key (2048), and Value (4096).
  • Gated Attention Layers (6144 × dmodel): Sliced into Query (4096), Key (1024), and Value (1024).

Query and Key slices are processed with a conservative retention density (ρ = 0.95) to preserve sharp context routing. Value matrices are processed with adaptive MLP density (ρ = 0.70) to maximize conceptual synthesis. The components are then re-concatenated along the head dimension.

5. Neuron-Coherent Row Gating

To eliminate destructive interference in Feed-Forward Networks (MLPs), task vectors are gated at the single-neuron (row) level:

τˉ=1K∑k=1Kτk \bar{\tau} = \frac{1}{K} \sum_{k=1}^K \tau_k

For each row r of donor delta τk, the directional alignment with the consensus mean is evaluated:

cos⁡θk,r=⟨τk,r,τˉr⟩∥τk,r∥2∥τˉr∥2+ϵ \cos \theta_{k, r} = \frac{\langle \tau_{k, r}, \bar{\tau}_r \rangle}{\|\tau_{k, r}\|_2 \|\bar{\tau}_r\|_2 + \epsilon}

Rows exhibiting severe directional opposition (cos θk, r < -0.10) are masked out:

τ^k,r=τk,r⋅I(cos⁡θk,r≥−0.10) \hat{\tau}_{k, r} = \tau_{k, r} \cdot \mathbb{I}\left(\cos \theta_{k, r} \ge -0.10\right)

This eliminates opposing gradient vectors that produce incoherent syntax generation.

6. Heavy-Tailed DELLA Adaptive Rescaling

Surviving parameters undergo non-linear magnitude-based sampling. Using parameter rank indices Rk, ij ∈ [0, 1] sorted by absolute magnitude, a Pareto-style retention probability pk, ij is established:

pk,ij=pmin⁡+(pmax⁡−pmin⁡)⋅Rk,ij p_{k, ij} = p_{\min} + (p_{\max} - p_{\min}) \cdot \sqrt{R_{k, ij}}

Parameters are sampled via a Bernoulli trial and rescaled by their inverse survival probability:

Mk,ij∼Bernoulli(pk,ij) M_{k, ij} \sim \text{Bernoulli}(p_{k, ij})

τ~k,ij=τ^k,ij⊙Mk,ijpk,ij \tilde{\tau}_{k, ij} = \frac{\hat{\tau}_{k, ij} \odot M_{k, ij}}{p_{k, ij}}

7. Coordinate-Wise Sign Consensus & Model Stock Scaling

Directional consensus is determined via weighted sign agreement:

Γ=sgn⁡(∑k=1Kw~kτ~k) \Gamma = \operatorname{sgn}\left(\sum_{k=1}^K \tilde{w}_k \tilde{\tau}_k\right)

Ak=I(sgn⁡(τ~k)=Γ)⊙I(τ~k≠0) A_k = \mathbb{I}\left(\operatorname{sgn}(\tilde{\tau}_k) = \Gamma\right) \odot \mathbb{I}\left(\tilde{\tau}_k \neq 0\right)

Δconsensus=∑k=1Kτ~k⊙Ak∑k=1KAk+ϵ \Delta_{\text{consensus}} = \frac{\sum_{k=1}^K \tilde{\tau}_k \odot A_k}{\sum_{k=1}^K A_k + \epsilon}

The aggregated delta is projected onto the non-linear manifold using the Model Stock analytic scaling factor t*:

ρˉ=2K(K−1)∑i<j⟨vec(τi),vec(τj)⟩∥τi∥2∥τj∥2 \bar{\rho} = \frac{2}{K(K-1)} \sum_{i < j} \frac{\langle \text{vec}(\tau_i), \text{vec}(\tau_j) \rangle}{\|\tau_i\|_2 \|\tau_j\|_2}

t∗=Kρˉ1+(K−1)ρˉ t^* = \frac{K \bar{\rho}}{1 + (K - 1)\bar{\rho}}

Δfinal=t∗⋅Δconsensus \Delta_{\text{final}} = t^* \cdot \Delta_{\text{consensus}}

8. Spectral Norm & Attention Entropy Anchoring

Attention projection matrices (Wattn) are vulnerable to spectral explosion during merges, which contracts attention entropy and leads to repetitive generation loops.

The dominant singular value σ(W) is calculated via a three-iteration deterministic power iteration:

v(t+1)=WTu(t)∥WTu(t)∥2,u(t+1)=Wv(t+1)∥Wv(t+1)∥2 v^{(t+1)} = \frac{W^T u^{(t)}}{\|W^T u^{(t)}\|_2}, \quad u^{(t+1)} = \frac{W v^{(t+1)}}{\|W v^{(t+1)}\|_2}

σ(W)≈u(3)TWv(3) \sigma(W) \approx {u^{(3)}}^T W v^{(3)}

If the merged spectral radius grows more than 5% relative to the base model, it is scaled down:

Wfinal={Wmerged⋅(1.05⋅σ(W0)σ(Wmerged))if σ(Wmerged)>1.05⋅σ(W0)Wmergedotherwise W_{\text{final}} = \begin{cases} W_{\text{merged}} \cdot \left(\frac{1.05 \cdot \sigma(W_0)}{\sigma(W_{\text{merged}})}\right) & \text{if } \sigma(W_{\text{merged}}) > 1.05 \cdot \sigma(W_0) \\ W_{\text{merged}} & \text{otherwise} \end{cases}


Layer-Stratified Component Policies

Parameter Class Target Identifiers Applied Policy Density (ρ) Mathematical Constraints
Embeddings & LM Head embed_tokens, lm_head Low-Memory Streaming Blend 1.0 Convex iterative accumulation; vocab dimension aligned to 248,320.
Linear State-Space Projections linear_attn.in_proj_qkv Asymmetric GQA DELLA 0.95 (QK) / 0.70 (V) Sub-tensor slicing; separate routing and associative memory passes.
Self-Attention Projections self_attn.qkv_proj, o_proj Asymmetric GQA + Spectral Anchor 0.95 (QK) / 0.70 (V) Power-iteration clipping prevents σ > 1.05 σ₀.
Feed-Forward Blocks mlp.gate_proj, up_proj, down_proj Neuron-Gated HG-DELLA 0.50 – 0.70 Row-wise cosine filtering (cos θ ≥ -0.10); sign consensus.
Recurrent Gates & Normalization A_log, norm, conv1d Convex Parameter Blend 1.0 Preservation of Alog ≤ 0 to guarantee BIBO stability.

Agentic Chat Template & Operational Directives

This model uses the Improved Chat Template for Qwen 3.x by Olivia Rossi to support multi-tier Chain-of-Thought (CoT) reasoning, dual-format agentic tool execution, automatic error-recovery heuristics, and strict token-waste elimination.


Recommended Generation Parameters

For deterministic software engineering and complex reasoning benchmarks:

Parameter Recommended Value Description
Temperature 0.6 Balances strict syntactic validity with algorithmic path exploration.
Top-P 0.95 Eliminates low-probability token tails while preserving alternative logic paths.
Top-K 20 Restricts token candidate pools to prevent architectural syntax drift.
Min-P 0.0 (Off) Disabled in favor of explicit Top-K / Top-P governance.
Repetition Penalty 1.0 (Off) Disabled to prevent syntax degradation in repetitive code patterns (indentation, braces).
Presence Penalty 0.0 Prevents naming mutations across long-context symbol resolution.

How to Use

Serving via vLLM

vllm serve pragmaticcs/PentaCoder-9B \
  --dtype bfloat16 \
  --max-model-len 65536 \
  --gpu-memory-utilization 0.95 \
  --enable-prefix-caching \
  --enable-auto-tool-choice \
  --tool-call-parser qwen3_coder \
  --enable-reasoning \
  --reasoning-parser qwen3

Inference via Transformers

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "pragmaticcs/PentaCoder-9B"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id, torch_dtype=torch.bfloat16, device_map="auto"
)

messages = [
    {
        "role": "system",
        "content": "You are an expert systems engineer. Reason step by step and output clean, robust implementations.",
    },
    {
        "role": "user",
        "content": "Write an asynchronous connection pool in Python for TCP sockets with active health-checking, backpressure control, and graceful shutdown handling.",
    },
]

inputs = tokenizer.apply_chat_template(
    messages, add_generation_prompt=True, enable_thinking=True, return_tensors="pt"
).to(model.device)

outputs = model.generate(
    inputs,
    max_new_tokens=4096,
    temperature=0.6,
    top_p=0.95,
    top_k=20,
    do_sample=True,
)

response = tokenizer.decode(outputs[0][inputs.shape[-1] :], skip_special_tokens=True)
print(response)

Citation & References

@inproceedings{yadav2023ties,
  title={Resolving Interference When Merging Models},
  author={Yadav, Prateek and Tam, Derek and Choshen, Leshem and Raffel, Colin and Bansal, Mohit},
  booktitle={Advances in Neural Information Processing Systems (NeurIPS)},
  volume={36},
  pages={7093--7115},
  year={2023}
}

@article{deep2024della,
  title={DELLA-Merging: Reducing Interference in Model Merging through Magnitude-Based Sampling},
  author={Deep, Pala Tej and Bhardwaj, Rishabh and Poria, Soujanya},
  journal={arXiv preprint arXiv:2406.11617},
  year={2024}
}

@article{jang2024modelstock,
  title={Model Stock: All We Need Is just a Few Fine-Tuned Models},
  author={Jang, Dong-Hwan and Yoon, Sang-Doo and Song, Gyeong-Moon},
  journal={arXiv preprint arXiv:2403.19522},
  year={2024}
}
Downloads last month
293
Safetensors
Model size
9B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for pragmaticcs/PentaCoder-9B

Collection including pragmaticcs/PentaCoder-9B

Papers for pragmaticcs/PentaCoder-9B