Ael-Coder-16B-MoE

15.71B Total / 2.36B Active Mixture-of-Experts Code Intelligence • ISOM-R2 Bounded-Memory Architecture

DOI Benchmark Suite Author


Overview

Ael-Coder-16B-MoE is a 15.71-billion-parameter Mixture-of-Experts code intelligence model (64 routed experts, 2 shared experts, 2.36B active parameters per token) equipped with Multi-Head Latent Attention (MLA, $d_k = 192$) and powered by the ISOM-R2 (Isometric State Operator Manifold) multi-million-token architecture.

Designed for repository-scale code synthesis and cross-module dependency resolution, Ael-Coder-16B-MoE streams 2,774,027 real codebase tokens (1,355 chunks across 82 source files from huggingface/transformers) in 12.18 seconds (227,686 tokens/sec) on a single NVIDIA A100-40GB GPU. Throughout the entire 2.77M-token prefill, active KV cache length remains locked at 2,112 tokens with only +0.06 GB streaming VRAM overhead (29.28 GB model weights $\rightarrow$ 29.34 GB streaming VRAM, 30.51 GB peak VRAM during synthesis).

Core Architectural Capabilities

  1. Hierarchical Paged Virtual SVD Cache (ISOMR2VirtualSVDCache)
    Streams multi-million-token repositories in 2,048-token chunks across MLA decoder layers while strictly bounding active GPU KV memory to 2,112 tokens (64 attention sinks + 2,048 rolling window).
  2. Entity-Balanced Multi-Hop Retrieval (ISOMR2Engine)
    Automatically intercepts long-context inputs inside model.generate(), balancing per-entity lexical and spectral recall to isolate exact target definitions separated by over 1.39 million tokens ([339, 1020], zero distractors).
  3. $SO(192)$ Multi-Head Latent Attention Lie-Manifold Transport
    Enforces skew-symmetric Cayley orthogonal transport ($R = (I - \frac{1}{2}A)^{-1}(I + \frac{1}{2}A) \in SO(192)$) across the 192-dimensional MLA head space, achieving an audited orthogonality error of $4.47 \times 10^{-5}$ and an $O(1)$ multi-hop state rollback error of $3.06 \times 10^{-6}$.

Ael Model Family — Audited Multi-Million-Token Benchmarks

All four models in the Ael benchmark suite are evaluated on unpadded, real-world open-source repositories (huggingface/transformers, sympy/sympy, and django/django) on NVIDIA A100-SXM4-40GB hardware. Interactive comparisons and telemetry logs are hosted on the Official Ael Benchmark Space.

Model Architecture & Parameters Benchmark Domain & Corpus Audited Context Inter-Hop Distance Total Time (Throughput) Weights / Peak VRAM Multi-Hop Recall & Verification
Ael-Coder-1.5B AelCoder15BForCausalLM
1.54B Dense (28L GQA)
Multi-File PyTorch Synthesis
transformers (82 files)
2,147,447
(1,049 chunks)
1,077,049 tokens
(526 chunks)
7.42 s
(289,500 tok/s)
2.98 GB / 3.14 GB
(+0.16 GB overhead)
100% ([388, 791, 265])
LoggedGELU (0.00e+00 err)
Ael-Reasoning-1.5B-Instruct AelReasoning15BForCausalLM
1.54B Dense (28L + $SO(64)$)
Symbolic Math & Combinatorics
sympy (42 modules)
2,311,513
(1,129 chunks)
932,802 tokens
(455 chunks)
9.97 s
(231,748 tok/s)
2.88 GB / 3.03 GB
(+0.15 GB overhead)
100% ([501, 956])
DerangedFibonacci (Exact)
Ael-Coder-16B-MoE AelCoder16BMoEForCausalLM
15.71B Total / 2.36B Active MoE
Repository-Scale MoE Synthesis
transformers (82 files)
2,774,027
(1,355 chunks)
1,394,629 tokens
(681 chunks)
12.18 s
(227,686 tok/s)
29.28 GB / 30.51 GB
(+1.23 GB overhead)
100% ([339, 1020])
LoggedGELU (0.00e+00 err)
Ael-Pro-40B AelPro40BForCausalLM
40.0B Dense (60L, 4-bit NF4)
Enterprise Security & Crypto Audit
django (122 modules)
2,399,330
(1,172 chunks)
1,171,652 tokens
(572 chunks)
13.94 s
(172,072 tok/s)
21.62 GB / 24.73 GB
(+3.11 GB overhead)
100% ([304, 876])
SignedPBKDF2Hasher (Verified)

Audited Benchmark 1: 2,774,027-Token Multi-File MoE Code Synthesis

Executed Notebook: Ael_Coder_16B.ipynb

Task Methodology

Evaluated on 82 real Python source files (9,779,186 characters, 2,774,027 real tokens across 1,355 chunks) from huggingface/transformers on an NVIDIA A100-SXM4-40GB GPU. The model must retrieve two target classes separated by 1,394,629 real tokens (681 chunks apart):

  • Hop 1 (Chunk 339, Token #695,058): class CaptureStdout(CaptureStd) in testing_utils.py
  • Hop 2 (Chunk 1020, Token #2,089,687): class GELUActivation(nn.Module) in activations.py

Using both retrieved definitions, the model synthesizes a LoggedGELU(GELUActivation) class that captures standard output during the forward pass and is verified live on GPU against PyTorch's native GELUActivation.

Hardware Execution Log (NVIDIA A100-SXM4-40GB)

Loaded AelCoder16BMoEForCausalLM (15.71B total / 2.36B active MoE params) | Weights VRAM: 29.28 GB
Tokenizing 2.25M+ real-world multi-file corpus (82 files | 9,779,186 chars)...
Hop 1 Target   : class CaptureStdout(CaptureStd) @ Token #695,058 -> Chunk 339
Hop 2 Target   : class GELUActivation(nn.Module) @ Token #2,089,687 -> Chunk 1020
Inter-File Gap : 1,394,629 real tokens (681 chunks apart)

  [ISOM-R2 Engine] Streaming 2,773,953 codebase context tokens across 1355 chunks (active GPU buffer < 400 MB)...
  [ISOM-R2] Prefill 552,960 / 2,773,953 tokens (19.9%) | Active buffer: 2112 tokens | VRAM: 29.34 GB
  [ISOM-R2] Prefill 1,105,920 / 2,773,953 tokens (39.9%) | Active buffer: 2112 tokens | VRAM: 29.34 GB
  [ISOM-R2] Prefill 1,658,880 / 2,773,953 tokens (59.8%) | Active buffer: 2112 tokens | VRAM: 29.34 GB
  [ISOM-R2] Prefill 2,211,840 / 2,773,953 tokens (79.7%) | Active buffer: 2112 tokens | VRAM: 29.34 GB
  [ISOM-R2] Prefill 2,764,800 / 2,773,953 tokens (99.7%) | Active buffer: 2112 tokens | VRAM: 29.34 GB
  [ISOM-R2] Prefill 2,773,953 / 2,773,953 tokens (100.0%) | Active buffer: 2112 tokens | VRAM: 29.34 GB
  [ISOM-R2] Retrieved salient context pages: [339, 1020] | Active KV: 2112 tokens

============================================================================================
2.77M-TOKEN MULTI-FILE MOE SYNTHESIS BENCHMARK: Prannesshkva/Ael-Coder-16B-MoE
============================================================================================
Total Real Context Tokens : 2,774,027 (82 files | 1355 chunks)
Inter-File Hop Distance   : 1,394,629 tokens (Chunk 339 <-> Chunk 1020)
Total Time (Prefill+Gen)  : 12.18 s (227,686 tok/s)
Model Weights VRAM        : 29.28 GB
Peak Total GPU VRAM       : 30.51 GB (KV + Activation Overhead: +1.23 GB)
Active KV Cache Length    : 2112 tokens
Retrieved Chunk Indices   : [339, 1020]
Multi-Hop Recall          : Hop 1 Hit: True | Hop 2 Hit: True
SO(192) MLA Orthogonality : 4.47e-05 ||R^T R - I||_F  |  O(1) Rollback Error: 3.06e-06
Live Execution Verification: PASSED (Output shape: [1, 4], Captured Stdout: 'torch.Size([1, 4])', max_err = 0.00e+00)
============================================================================================

Audited Benchmark 2: 1,089,849-Token Single-Hop Repository Retrieval

Evaluated on 93 real Python source files (3,867,627 characters, 1,089,849 tokens across 533 chunks) from huggingface/transformers and pytorch/pytorch:

Loaded AelCoder16BMoEForCausalLM | Model VRAM: 31.49 GB
Corpus: 93 real files | 3,867,627 chars | 1,089,849 tokens | Ground-Truth: Chunk 257
  [ISOM-R2 Engine] Streaming 1,089,795 codebase context tokens across 533 chunks (active GPU buffer < 400 MB)...
  [ISOM-R2] Prefill 1,089,795 / 1,089,795 tokens (100.0%) | Active buffer: 2112 tokens | VRAM: 31.49 GB
  [ISOM-R2] Retrieved salient context pages: [258, 257] | Active KV: 2112 tokens

================================================================================
BENCHMARK RESULTS: Prannesshkva/Ael-Coder-16B-MoE
================================================================================
Total Real Context Tokens : 1,089,849
Total Time (Prefill+Gen)  : 15.86 s (68,717 tok/s)
Model Weights VRAM        : 31.49 GB
Peak Total GPU VRAM       : 32.06 GB ( Overhead: +0.57 GB )
Active KV Cache Length    : 2112 tokens
Retrieved Chunk Indices   : [258, 257] (Exact Token Ground-Truth: Chunk 257)
Ground-Truth Base Class   : CaptureStd
Model Generated Output    : CaptureStd
================================================================================

Architectural Specifications

Parameter Specification
Model Architecture Class AelCoder16BMoEForCausalLM (AelCoder16BMoEConfig)
Total / Active Parameters 15.71 Billion Total / 2.36 Billion Active per Token
MoE & Attention Topology 27 Layers, 64 Routed Experts, 2 Shared Experts, Multi-Head Latent Attention ($d_k = 192$)
Long-Context Engine ISOM-R2 Paged Virtual SVD Cache (ISOMR2VirtualSVDCache) + $SO(192)$ Cayley Transport
Audited Context Length 2,774,027 real tokens (1,355 chunks of 2,048 tokens)
Active GPU KV Cache 2,112 tokens (64 attention sinks + 2,048 active sliding window)
VRAM Footprint (BF16) 29.28 GB Weights • 29.34 GB Streaming (+0.06 GB) • 30.51 GB Peak Synthesis

Quickstart Usage

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "Prannesshkva/Ael-Coder-16B-MoE"

tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16,
    device_map="auto",
    trust_remote_code=True,
).eval()

prompt = """<|im_start|>user
Write a high-performance concurrent queue in Python using lock-free atomic operations.<|im_end|>
<|im_start|>assistant
"""

inputs = tokenizer(prompt, return_tensors="pt").to(model.device)

with torch.no_grad():
    outputs = model.generate(
        **inputs,
        max_new_tokens=512,
        do_sample=False,
        tokenizer=tokenizer,
    )

print(tokenizer.decode(outputs[0][inputs.input_ids.shape[1]:], skip_special_tokens=True))

Citation, Dual-Layer Licensing & Upstream Attribution

@software{ael_coder_16b_moe_2026,
  author    = {Prannessh K. V. A.},
  title     = {Ael-Coder-16B-MoE: Multi-Million-Token Repository-Scale Bounded-Memory MoE Powered by ISOM-R2},
  year      = {2026},
  publisher = {CERN Zenodo},
  doi       = {10.5281/zenodo.22649142},
  url       = {https://huggingface.co/Prannesshkva/Ael-Coder-16B-MoE}
}

Dual-Layer License Structure (DeepSeek License Section 3 & MIT Compliance)

Pursuant to Section 3 of the DeepSeek License Agreement and the MIT License, this repository separates licensing between the unmodified upstream pretrained foundation weights and the author's original architectural modifications:

Component Copyright Holder Applicable License
ISOM-R2 Execution Engine & Architectural Modifications
(isom_r2_engine.py, ISOMR2VirtualSVDCache, $SO(192)$ MLA Cayley Lie-Group Transport, and custom AelCoder16BMoE* classes in modeling_isom_deepseek_coder_v2.py)
Copyright © 2026 Prannessh K. V. A. CC BY-NC-ND 4.0 (Non-Commercial Research) / Commercial Enterprise License via Author (LICENSE)
Base Pretrained Neural Weights & Unmodified Base Code
(Initialized from deepseek-ai/DeepSeek-Coder-V2-Lite-Instruct)
Copyright © 2024 DeepSeek DeepSeek License Agreement (Model Weights, including Attachment A Use-Based Restrictions) & MIT License (Base Code) (LICENSE & NOTICE)

Mandatory Upstream Notice & Trademark Disclaimer (DeepSeek License Sections 3 & 4)

"DeepSeek-Coder-V2 is licensed under the MODEL LICENSE. Copyright (c) 2024 DeepSeek. All Rights Reserved."

  1. Prominent Notice of Modification (Section 3(b)): Modified by Prannessh K. V. A. to integrate the ISOM-R2 (Isometric State Operator Manifold) Paged Virtual SVD Cache (ISOMR2VirtualSVDCache), entity-balanced multi-hop codebase retrieval (ISOMR2Engine), $SO(192)$ Multi-Head Latent Attention Lie-group transport, and AelCoder16BMoEForCausalLM execution bindings. Full modification logs and Attachment A Use-Based Restrictions are included in LICENSE and NOTICE.
  2. Distinct Naming & Non-Endorsement (Section 4 — Trademarks): In compliance with Section 4 of the DeepSeek License Agreement, this derivative architecture is published under the distinct Ael name (Ael-Coder-16B-MoE) so as not to imply endorsement by or affiliation with the original licensor. DeepSeek is a trademark of DeepSeek AI. This independent research work is not affiliated with, sponsored by, or endorsed by DeepSeek AI.
Downloads last month
1,349
Safetensors
Model size
16B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Prannesshkva/Ael-Coder-16B-MoE

Finetuned
(19)
this model

Space using Prannesshkva/Ael-Coder-16B-MoE 1

Collection including Prannesshkva/Ael-Coder-16B-MoE