Ael-Coder-1.5B

1.54B Dense Code Intelligence Model • ISOM-R2 Bounded-Memory Architecture

DOI Benchmark Suite Author


Overview

Ael-Coder-1.5B is the compact 1.54-billion-parameter code intelligence model in the Ael Model Family, built on the Qwen2.5-Coder-1.5B-Instruct foundation (28 decoder layers, 12 query heads, 2 key-value heads) and powered internally by the ISOM-R2 (Isometric State Operator Manifold) multi-million-token memory architecture.

In standard Grouped-Query Attention (GQA), a 2.15M-token context sequence requires 57.3 GiB of VRAM for the FP16 KV cache alone. Ael-Coder-1.5B eliminates this linear memory wall by bounding the active GPU attention window to 2,112 tokens (+0.16 GB peak overhead above the 2.98 GB model weights), streaming and synthesizing across 2,147,447 real codebase tokens in 7.42 seconds (289,500 tokens/sec) within 3.14 GB Peak VRAM.

Core Architectural Capabilities

  1. Hierarchical Paged Virtual SVD Cache (ISOMR2VirtualSVDCache)
    Streams multi-million-token repositories in 2,048-token pages while locking the active GPU KV cache at 2,112 tokens (64 attention sinks + 2,048 rolling window), holding streaming VRAM flat at 2.99 GB.
  2. Entity-Balanced Multi-Hop Retrieval (ISOMR2Engine)
    Automatically intercepts repository prompts $> 2,048$ tokens inside model.generate(), isolating target classes across million-token inter-file gaps and assembling a compact context window for code synthesis.
  3. Consumer & Edge GPU Deployment
    With a peak memory footprint of 3.14 GB VRAM at 2.15M tokens, repository-scale code intelligence runs natively on 4 GB, 6 GB, and 8 GB consumer GPUs.

Ael Model Family — Audited Multi-Million-Token Benchmarks

All four models in the Ael benchmark suite are evaluated on unpadded, real-world open-source repositories (huggingface/transformers, sympy/sympy, and django/django) on NVIDIA A100-SXM4-40GB hardware. Interactive comparisons and telemetry logs are hosted on the Official Ael Benchmark Space.

Model Architecture & Parameters Benchmark Domain & Corpus Audited Context Inter-Hop Distance Total Time (Throughput) Weights / Peak VRAM Multi-Hop Recall & Verification
Ael-Coder-1.5B AelCoder15BForCausalLM
1.54B Dense (28L GQA)
Multi-File PyTorch Synthesis
transformers (82 files)
2,147,447
(1,049 chunks)
1,077,049 tokens
(526 chunks)
7.42 s
(289,500 tok/s)
2.98 GB / 3.14 GB
(+0.16 GB overhead)
100% ([388, 791, 265])
LoggedGELU (0.00e+00 err)
Ael-Reasoning-1.5B-Instruct AelReasoning15BForCausalLM
1.54B Dense (28L + $SO(64)$)
Symbolic Math & Combinatorics
sympy (42 modules)
2,311,513
(1,129 chunks)
932,802 tokens
(455 chunks)
9.97 s
(231,748 tok/s)
2.88 GB / 3.03 GB
(+0.15 GB overhead)
100% ([501, 956])
DerangedFibonacci (Exact)
Ael-Coder-16B-MoE AelCoder16BMoEForCausalLM
15.71B Total / 2.36B Active MoE
Repository-Scale MoE Synthesis
transformers (82 files)
2,774,027
(1,355 chunks)
1,394,629 tokens
(681 chunks)
12.18 s
(227,686 tok/s)
29.28 GB / 30.51 GB
(+1.23 GB overhead)
100% ([339, 1020])
LoggedGELU (0.00e+00 err)
Ael-Pro-40B AelPro40BForCausalLM
40.0B Dense (60L, 4-bit NF4)
Enterprise Security & Crypto Audit
django (122 modules)
2,399,330
(1,172 chunks)
1,171,652 tokens
(572 chunks)
13.94 s
(172,072 tok/s)
21.62 GB / 24.73 GB
(+3.11 GB overhead)
100% ([304, 876])
SignedPBKDF2Hasher (Verified)

Memory Scaling: Standard Attention vs. Ael-Coder-1.5B (ISOM-R2)

In a standard 28-layer GQA architecture (2 KV heads, $d_{\text{head}} = 128$), KV cache memory grows linearly with sequence length $N$:

KV Cache Size=2×28×2×128×N×2 bytes\text{KV Cache Size} = 2 \times 28 \times 2 \times 128 \times N \times 2\text{ bytes}

Metric (at 2,147,447 Tokens) Standard Transformer GQA Ael-Coder-1.5B (ISOM-R2) Measured Improvement
KV Cache Footprint 57.34 GiB (Linear $O(N)$) +0.01 GB Streaming / +0.16 GB Peak 99.7% Reduction
Active KV Buffer Length 2,147,447 tokens 2,112 tokens Strict $O(1)$ State
Peak Execution VRAM $> 60.3\text{ GB}$ (CUDA OOM on 40GB) 3.14 GB Peak (Weights: 2.98 GB) 19.2× Lower Total VRAM
Prefill + Synthesis Latency Infeasible ($O(N^2)$ attention) 7.42 seconds (289,500 tok/s) Constant-Time Streaming

Audited Benchmark 1: 2,147,447-Token Multi-File Code Synthesis

Executed Notebook: Ael_Coder_benchmark.ipynb

Task Methodology

Streams 2,147,447 real codebase tokens (9,780,715 characters across 82 Python source files from huggingface/transformers), retrieves two distant target classes separated by 1,077,049 tokens (526 chunks apart) — class CaptureStdout(CaptureStd) (Chunk 265) and class GELUActivation(nn.Module) (Chunk 791) — synthesizes a unified LoggedGELU(GELUActivation) class combining both, and verifies the generated code via live GPU execution (max_err = 0.00e+00).

Hardware Execution Log (NVIDIA A100-SXM4-40GB)

Loaded AelCoder15BForCausalLM | Model Weights VRAM: 2.98 GB
Tokenizing 2.15M+ real-world multi-file corpus (82 files | 9,780,715 chars)...
Hop 1 Target   : class CaptureStdout(CaptureStd) @ Token #543,865 -> Chunk 265
Hop 2 Target   : class GELUActivation(nn.Module) @ Token #1,620,914 -> Chunk 791
Inter-File Gap : 1,077,049 real tokens (526 chunks apart)

  [ISOM-R2 Engine] Streaming 2,147,383 codebase context tokens across 1049 chunks (active GPU buffer < 400 MB)...
  [ISOM-R2] Prefill 428,032 / 2,147,383 tokens (19.9%) | Active buffer: 2112 tokens | VRAM: 2.99 GB
  [ISOM-R2] Prefill 856,064 / 2,147,383 tokens (39.9%) | Active buffer: 2112 tokens | VRAM: 2.99 GB
  [ISOM-R2] Prefill 1,284,096 / 2,147,383 tokens (59.8%) | Active buffer: 2112 tokens | VRAM: 2.99 GB
  [ISOM-R2] Prefill 1,712,128 / 2,147,383 tokens (79.7%) | Active buffer: 2112 tokens | VRAM: 2.99 GB
  [ISOM-R2] Prefill 2,140,160 / 2,147,383 tokens (99.7%) | Active buffer: 2112 tokens | VRAM: 2.99 GB
  [ISOM-R2] Prefill 2,147,383 / 2,147,383 tokens (100.0%) | Active buffer: 2112 tokens | VRAM: 2.99 GB
  [ISOM-R2] Retrieved salient context pages: [388, 791, 265] | Active KV: 2112 tokens

========================================================================================
2.15M-TOKEN MULTI-FILE SYNTHESIS BENCHMARK: Prannesshkva/Ael-Coder-1.5B
========================================================================================
Model Architecture Class  : AelCoder15BForCausalLM
Total Real Source Files   : 82 files (9,780,715 chars)
Total Real Context Tokens : 2,147,447
Distance Between Files    : 1,077,049 tokens (526 chunks apart)
Total Time (Prefill+Gen)  : 7.42 s (289,500 tok/s)
Model Weights VRAM        : 2.98 GB
Peak Total GPU VRAM       : 3.14 GB ( Overhead: +0.16 GB )
Active KV Cache Length    : 2112 tokens
Ground-Truth Chunks       : Hop 1 = Chunk 265 | Hop 2 = Chunk 791
Retrieved Chunk Indices   : [388, 791, 265] (Hop 1 Hit: True | Hop 2 Hit: True)
----------------------------------------------------------------------------------------
MODEL SYNTHESIZED MULTI-FILE CODE:
----------------------------------------------------------------------------------------
class LoggedGELU(GELUActivation):
    def forward(self, x):
        with CaptureStdout(replay=False) as cs:
            out = super().forward(x)
            print(x.shape)
        return (out, cs.out)
----------------------------------------------------------------------------------------
LIVE GPU EXECUTION VERIFICATION OF SYNTHESIZED CLASS:
  • Live Instantiation & Forward Pass : PASSED (Output shape: [1, 4])
  • Captured Stdout via CaptureStdout : 'torch.Size([1, 4])'
  • Numerical Parity vs GELUActivation: max_err = 0.00e+00 (PASSED)
========================================================================================

Audited Benchmark 2: 1,052,188-Token Single-Hop Repository Retrieval

Evaluated across 93 real Python source files (3,867,627 characters, 1,052,188 tokens across 514 chunks) from huggingface/transformers and pytorch/pytorch:

Loaded AelCoder15BForCausalLM | Model VRAM: 3.09 GB
Corpus: 93 real files | 3,867,627 chars | 1,052,188 tokens | Ground-Truth: Chunk 249
  [ISOM-R2 Engine] Streaming 1,052,137 codebase context tokens across 514 chunks (active GPU buffer < 400 MB)...
  [ISOM-R2] Prefill 1,052,137 / 1,052,137 tokens (100.0%) | Active buffer: 2112 tokens | VRAM: 3.09 GB
  [ISOM-R2] Retrieved salient context pages: [250, 249] | Active KV: 2112 tokens

================================================================================
BENCHMARK RESULTS: Prannesshkva/Ael-Coder-1.5B
================================================================================
Total Real Context Tokens : 1,052,188
Total Time (Prefill+Gen)  : 6.10 s (172,490 tok/s)
Model Weights VRAM        : 3.09 GB
Peak Total GPU VRAM       : 3.51 GB ( Overhead: +0.42 GB )
Active KV Cache Length    : 2112 tokens
Retrieved Chunk Indices   : [250, 249] (Exact Token Ground-Truth: Chunk 249)
Ground-Truth Base Class   : CaptureStd
Model Generated Output    : CaptureStd
================================================================================

Quickstart Usage

import torch
from transformers import AutoTokenizer, AutoModelForCausalLM

MODEL_ID = "Prannesshkva/Ael-Coder-1.5B"

tokenizer = AutoTokenizer.from_pretrained(MODEL_ID, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    MODEL_ID,
    trust_remote_code=True,
    torch_dtype=torch.bfloat16,
    device_map="auto",
).eval()

inputs = tokenizer(huge_codebase_prompt, return_tensors="pt").to(model.device)

with torch.no_grad():
    output = model.generate(
        inputs["input_ids"],
        max_new_tokens=256,
        do_sample=False,
        tokenizer=tokenizer,
    )

print(tokenizer.decode(output[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))

Citation, Dual-Layer Licensing & Upstream Attribution

@software{ael_coder_1_5b_2026,
  author    = {Prannessh K. V. A.},
  title     = {Ael-Coder-1.5B: Multi-Million-Token Bounded-Memory Code Intelligence Powered by ISOM-R2},
  year      = {2026},
  publisher = {Hugging Face / CERN Zenodo},
  doi       = {10.5281/zenodo.22649142},
  url       = {https://huggingface.co/Prannesshkva/Ael-Coder-1.5B}
}

Dual-Layer License Structure (Apache 2.0 Section 4 Compliance)

Pursuant to Section 4 of the Apache License, Version 2.0, this repository separates licensing between the unmodified upstream pretrained foundation weights and the author's original architectural modifications:

Component Copyright Holder Applicable License
ISOM-R2 Execution Engine & Architectural Modifications
(isom_r2_engine.py, ISOMR2VirtualSVDCache, $SO(d)$ Lie-Manifold Operators, and custom AelCoder15B* classes in modeling_isom_qwen25_coder.py)
Copyright © 2026 Prannessh K. V. A. CC BY-NC-ND 4.0 (Non-Commercial Research) / Commercial Enterprise License via Author (LICENSE)
Base Pretrained Neural Weights & Unmodified Base Qwen2 Code
(Initialized from Qwen/Qwen2.5-Coder-1.5B-Instruct)
Copyright © 2024 Alibaba Cloud (Qwen Team) Apache License, Version 2.0 (LICENSE & NOTICE)

Statement of Modifications & Trademark Notice (Apache 2.0 Sections 4 & 6)

  1. Prominent Notice of Modification (Section 4(b)): Modified by Prannessh K. V. A. to integrate the ISOM-R2 (Isometric State Operator Manifold) Paged Virtual SVD Cache (ISOMR2VirtualSVDCache), entity-balanced multi-hop codebase retrieval (ISOMR2Engine), dynamic symmetric INT8 KV quantization, and AelCoder15BForCausalLM execution bindings. Full modification logs are documented in NOTICE.
  2. Distinct Naming & Non-Endorsement (Section 6 — Trademarks): In compliance with Section 6 of the Apache License 2.0, this derivative architecture is published under the distinct Ael name (Ael-Coder-1.5B) so as not to imply endorsement by or affiliation with the original licensor. Qwen and Alibaba Cloud are trademarks of Alibaba Group. This independent research work is not affiliated with, sponsored by, or endorsed by Alibaba Cloud or the Qwen team.
Downloads last month
378
Safetensors
Model size
2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Prannesshkva/Ael-Coder-1.5B

Finetuned
(216)
this model

Space using Prannesshkva/Ael-Coder-1.5B 1

Collection including Prannesshkva/Ael-Coder-1.5B