- Ael-Coder-1.5B
- Overview
- Ael Model Family — Audited Multi-Million-Token Benchmarks
- Memory Scaling: Standard Attention vs. Ael-Coder-1.5B (ISOM-R2)
- Audited Benchmark 1: 2,147,447-Token Multi-File Code Synthesis
- Audited Benchmark 2: 1,052,188-Token Single-Hop Repository Retrieval
- Quickstart Usage
- Citation, Dual-Layer Licensing & Upstream Attribution
Ael-Coder-1.5B
1.54B Dense Code Intelligence Model • ISOM-R2 Bounded-Memory Architecture
Overview
Ael-Coder-1.5B is the compact 1.54-billion-parameter code intelligence model in the Ael Model Family, built on the Qwen2.5-Coder-1.5B-Instruct foundation (28 decoder layers, 12 query heads, 2 key-value heads) and powered internally by the ISOM-R2 (Isometric State Operator Manifold) multi-million-token memory architecture.
In standard Grouped-Query Attention (GQA), a 2.15M-token context sequence requires 57.3 GiB of VRAM for the FP16 KV cache alone. Ael-Coder-1.5B eliminates this linear memory wall by bounding the active GPU attention window to 2,112 tokens (+0.16 GB peak overhead above the 2.98 GB model weights), streaming and synthesizing across 2,147,447 real codebase tokens in 7.42 seconds (289,500 tokens/sec) within 3.14 GB Peak VRAM.
Core Architectural Capabilities
- Hierarchical Paged Virtual SVD Cache (
ISOMR2VirtualSVDCache)
Streams multi-million-token repositories in 2,048-token pages while locking the active GPU KV cache at 2,112 tokens (64 attention sinks + 2,048 rolling window), holding streaming VRAM flat at 2.99 GB. - Entity-Balanced Multi-Hop Retrieval (
ISOMR2Engine)
Automatically intercepts repository prompts $> 2,048$ tokens insidemodel.generate(), isolating target classes across million-token inter-file gaps and assembling a compact context window for code synthesis. - Consumer & Edge GPU Deployment
With a peak memory footprint of 3.14 GB VRAM at 2.15M tokens, repository-scale code intelligence runs natively on 4 GB, 6 GB, and 8 GB consumer GPUs.
Ael Model Family — Audited Multi-Million-Token Benchmarks
All four models in the Ael benchmark suite are evaluated on unpadded, real-world open-source repositories (huggingface/transformers, sympy/sympy, and django/django) on NVIDIA A100-SXM4-40GB hardware. Interactive comparisons and telemetry logs are hosted on the Official Ael Benchmark Space.
| Model | Architecture & Parameters | Benchmark Domain & Corpus | Audited Context | Inter-Hop Distance | Total Time (Throughput) | Weights / Peak VRAM | Multi-Hop Recall & Verification |
|---|---|---|---|---|---|---|---|
| Ael-Coder-1.5B | AelCoder15BForCausalLM1.54B Dense (28L GQA) |
Multi-File PyTorch Synthesistransformers (82 files) |
2,147,447 (1,049 chunks) |
1,077,049 tokens (526 chunks) |
7.42 s (289,500 tok/s) |
2.98 GB / 3.14 GB (+0.16 GB overhead) |
100% ([388, 791, 265])LoggedGELU (0.00e+00 err) |
| Ael-Reasoning-1.5B-Instruct | AelReasoning15BForCausalLM1.54B Dense (28L + $SO(64)$) |
Symbolic Math & Combinatoricssympy (42 modules) |
2,311,513 (1,129 chunks) |
932,802 tokens (455 chunks) |
9.97 s (231,748 tok/s) |
2.88 GB / 3.03 GB (+0.15 GB overhead) |
100% ([501, 956])DerangedFibonacci (Exact) |
| Ael-Coder-16B-MoE | AelCoder16BMoEForCausalLM15.71B Total / 2.36B Active MoE |
Repository-Scale MoE Synthesistransformers (82 files) |
2,774,027 (1,355 chunks) |
1,394,629 tokens (681 chunks) |
12.18 s (227,686 tok/s) |
29.28 GB / 30.51 GB (+1.23 GB overhead) |
100% ([339, 1020])LoggedGELU (0.00e+00 err) |
| Ael-Pro-40B | AelPro40BForCausalLM40.0B Dense (60L, 4-bit NF4) |
Enterprise Security & Crypto Auditdjango (122 modules) |
2,399,330 (1,172 chunks) |
1,171,652 tokens (572 chunks) |
13.94 s (172,072 tok/s) |
21.62 GB / 24.73 GB (+3.11 GB overhead) |
100% ([304, 876])SignedPBKDF2Hasher (Verified) |
Memory Scaling: Standard Attention vs. Ael-Coder-1.5B (ISOM-R2)
In a standard 28-layer GQA architecture (2 KV heads, $d_{\text{head}} = 128$), KV cache memory grows linearly with sequence length $N$:
| Metric (at 2,147,447 Tokens) | Standard Transformer GQA | Ael-Coder-1.5B (ISOM-R2) | Measured Improvement |
|---|---|---|---|
| KV Cache Footprint | 57.34 GiB (Linear $O(N)$) | +0.01 GB Streaming / +0.16 GB Peak | 99.7% Reduction |
| Active KV Buffer Length | 2,147,447 tokens | 2,112 tokens | Strict $O(1)$ State |
| Peak Execution VRAM | $> 60.3\text{ GB}$ (CUDA OOM on 40GB) | 3.14 GB Peak (Weights: 2.98 GB) | 19.2× Lower Total VRAM |
| Prefill + Synthesis Latency | Infeasible ($O(N^2)$ attention) | 7.42 seconds (289,500 tok/s) | Constant-Time Streaming |
Audited Benchmark 1: 2,147,447-Token Multi-File Code Synthesis
Executed Notebook: Ael_Coder_benchmark.ipynb
Task Methodology
Streams 2,147,447 real codebase tokens (9,780,715 characters across 82 Python source files from huggingface/transformers), retrieves two distant target classes separated by 1,077,049 tokens (526 chunks apart) — class CaptureStdout(CaptureStd) (Chunk 265) and class GELUActivation(nn.Module) (Chunk 791) — synthesizes a unified LoggedGELU(GELUActivation) class combining both, and verifies the generated code via live GPU execution (max_err = 0.00e+00).
Hardware Execution Log (NVIDIA A100-SXM4-40GB)
Loaded AelCoder15BForCausalLM | Model Weights VRAM: 2.98 GB
Tokenizing 2.15M+ real-world multi-file corpus (82 files | 9,780,715 chars)...
Hop 1 Target : class CaptureStdout(CaptureStd) @ Token #543,865 -> Chunk 265
Hop 2 Target : class GELUActivation(nn.Module) @ Token #1,620,914 -> Chunk 791
Inter-File Gap : 1,077,049 real tokens (526 chunks apart)
[ISOM-R2 Engine] Streaming 2,147,383 codebase context tokens across 1049 chunks (active GPU buffer < 400 MB)...
[ISOM-R2] Prefill 428,032 / 2,147,383 tokens (19.9%) | Active buffer: 2112 tokens | VRAM: 2.99 GB
[ISOM-R2] Prefill 856,064 / 2,147,383 tokens (39.9%) | Active buffer: 2112 tokens | VRAM: 2.99 GB
[ISOM-R2] Prefill 1,284,096 / 2,147,383 tokens (59.8%) | Active buffer: 2112 tokens | VRAM: 2.99 GB
[ISOM-R2] Prefill 1,712,128 / 2,147,383 tokens (79.7%) | Active buffer: 2112 tokens | VRAM: 2.99 GB
[ISOM-R2] Prefill 2,140,160 / 2,147,383 tokens (99.7%) | Active buffer: 2112 tokens | VRAM: 2.99 GB
[ISOM-R2] Prefill 2,147,383 / 2,147,383 tokens (100.0%) | Active buffer: 2112 tokens | VRAM: 2.99 GB
[ISOM-R2] Retrieved salient context pages: [388, 791, 265] | Active KV: 2112 tokens
========================================================================================
2.15M-TOKEN MULTI-FILE SYNTHESIS BENCHMARK: Prannesshkva/Ael-Coder-1.5B
========================================================================================
Model Architecture Class : AelCoder15BForCausalLM
Total Real Source Files : 82 files (9,780,715 chars)
Total Real Context Tokens : 2,147,447
Distance Between Files : 1,077,049 tokens (526 chunks apart)
Total Time (Prefill+Gen) : 7.42 s (289,500 tok/s)
Model Weights VRAM : 2.98 GB
Peak Total GPU VRAM : 3.14 GB ( Overhead: +0.16 GB )
Active KV Cache Length : 2112 tokens
Ground-Truth Chunks : Hop 1 = Chunk 265 | Hop 2 = Chunk 791
Retrieved Chunk Indices : [388, 791, 265] (Hop 1 Hit: True | Hop 2 Hit: True)
----------------------------------------------------------------------------------------
MODEL SYNTHESIZED MULTI-FILE CODE:
----------------------------------------------------------------------------------------
class LoggedGELU(GELUActivation):
def forward(self, x):
with CaptureStdout(replay=False) as cs:
out = super().forward(x)
print(x.shape)
return (out, cs.out)
----------------------------------------------------------------------------------------
LIVE GPU EXECUTION VERIFICATION OF SYNTHESIZED CLASS:
• Live Instantiation & Forward Pass : PASSED (Output shape: [1, 4])
• Captured Stdout via CaptureStdout : 'torch.Size([1, 4])'
• Numerical Parity vs GELUActivation: max_err = 0.00e+00 (PASSED)
========================================================================================
Audited Benchmark 2: 1,052,188-Token Single-Hop Repository Retrieval
Evaluated across 93 real Python source files (3,867,627 characters, 1,052,188 tokens across 514 chunks) from huggingface/transformers and pytorch/pytorch:
Loaded AelCoder15BForCausalLM | Model VRAM: 3.09 GB
Corpus: 93 real files | 3,867,627 chars | 1,052,188 tokens | Ground-Truth: Chunk 249
[ISOM-R2 Engine] Streaming 1,052,137 codebase context tokens across 514 chunks (active GPU buffer < 400 MB)...
[ISOM-R2] Prefill 1,052,137 / 1,052,137 tokens (100.0%) | Active buffer: 2112 tokens | VRAM: 3.09 GB
[ISOM-R2] Retrieved salient context pages: [250, 249] | Active KV: 2112 tokens
================================================================================
BENCHMARK RESULTS: Prannesshkva/Ael-Coder-1.5B
================================================================================
Total Real Context Tokens : 1,052,188
Total Time (Prefill+Gen) : 6.10 s (172,490 tok/s)
Model Weights VRAM : 3.09 GB
Peak Total GPU VRAM : 3.51 GB ( Overhead: +0.42 GB )
Active KV Cache Length : 2112 tokens
Retrieved Chunk Indices : [250, 249] (Exact Token Ground-Truth: Chunk 249)
Ground-Truth Base Class : CaptureStd
Model Generated Output : CaptureStd
================================================================================
Quickstart Usage
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
MODEL_ID = "Prannesshkva/Ael-Coder-1.5B"
tokenizer = AutoTokenizer.from_pretrained(MODEL_ID, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
MODEL_ID,
trust_remote_code=True,
torch_dtype=torch.bfloat16,
device_map="auto",
).eval()
inputs = tokenizer(huge_codebase_prompt, return_tensors="pt").to(model.device)
with torch.no_grad():
output = model.generate(
inputs["input_ids"],
max_new_tokens=256,
do_sample=False,
tokenizer=tokenizer,
)
print(tokenizer.decode(output[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))
Citation, Dual-Layer Licensing & Upstream Attribution
@software{ael_coder_1_5b_2026,
author = {Prannessh K. V. A.},
title = {Ael-Coder-1.5B: Multi-Million-Token Bounded-Memory Code Intelligence Powered by ISOM-R2},
year = {2026},
publisher = {Hugging Face / CERN Zenodo},
doi = {10.5281/zenodo.22649142},
url = {https://huggingface.co/Prannesshkva/Ael-Coder-1.5B}
}
Dual-Layer License Structure (Apache 2.0 Section 4 Compliance)
Pursuant to Section 4 of the Apache License, Version 2.0, this repository separates licensing between the unmodified upstream pretrained foundation weights and the author's original architectural modifications:
| Component | Copyright Holder | Applicable License |
|---|---|---|
| ISOM-R2 Execution Engine & Architectural Modifications ( isom_r2_engine.py, ISOMR2VirtualSVDCache, $SO(d)$ Lie-Manifold Operators, and custom AelCoder15B* classes in modeling_isom_qwen25_coder.py) |
Copyright © 2026 Prannessh K. V. A. | CC BY-NC-ND 4.0 (Non-Commercial Research) / Commercial Enterprise License via Author (LICENSE) |
| Base Pretrained Neural Weights & Unmodified Base Qwen2 Code (Initialized from Qwen/Qwen2.5-Coder-1.5B-Instruct) |
Copyright © 2024 Alibaba Cloud (Qwen Team) | Apache License, Version 2.0 (LICENSE & NOTICE) |
Statement of Modifications & Trademark Notice (Apache 2.0 Sections 4 & 6)
- Prominent Notice of Modification (Section 4(b)): Modified by Prannessh K. V. A. to integrate the ISOM-R2 (Isometric State Operator Manifold) Paged Virtual SVD Cache (
ISOMR2VirtualSVDCache), entity-balanced multi-hop codebase retrieval (ISOMR2Engine), dynamic symmetric INT8 KV quantization, andAelCoder15BForCausalLMexecution bindings. Full modification logs are documented inNOTICE. - Distinct Naming & Non-Endorsement (Section 6 — Trademarks): In compliance with Section 6 of the Apache License 2.0, this derivative architecture is published under the distinct Ael name (
Ael-Coder-1.5B) so as not to imply endorsement by or affiliation with the original licensor. Qwen and Alibaba Cloud are trademarks of Alibaba Group. This independent research work is not affiliated with, sponsored by, or endorsed by Alibaba Cloud or the Qwen team.
- Downloads last month
- 378
Model tree for Prannesshkva/Ael-Coder-1.5B
Base model
Qwen/Qwen2.5-1.5B