Text Generation
Transformers
Safetensors
English
forgelm
keystack
mla
Mixture of Experts
mrl
quarot
rotorquant
mtp
airllm
weight-transform
training-free
code
conversational
custom_code
Instructions to use GRKTheGreat/ForgeLM-v1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use GRKTheGreat/ForgeLM-v1 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="GRKTheGreat/ForgeLM-v1", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import ForgeLM model = ForgeLM.from_pretrained("GRKTheGreat/ForgeLM-v1", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use GRKTheGreat/ForgeLM-v1 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "GRKTheGreat/ForgeLM-v1" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "GRKTheGreat/ForgeLM-v1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/GRKTheGreat/ForgeLM-v1
- SGLang
How to use GRKTheGreat/ForgeLM-v1 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "GRKTheGreat/ForgeLM-v1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "GRKTheGreat/ForgeLM-v1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "GRKTheGreat/ForgeLM-v1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "GRKTheGreat/ForgeLM-v1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use GRKTheGreat/ForgeLM-v1 with Docker Model Runner:
docker model run hf.co/GRKTheGreat/ForgeLM-v1
| library_name: transformers | |
| license: apache-2.0 | |
| base_model: Qwen/Qwen2.5-Coder-1.5B-Instruct | |
| tags: | |
| - keystack | |
| - mla | |
| - moe | |
| - mrl | |
| - quarot | |
| - rotorquant | |
| - mtp | |
| - airllm | |
| - weight-transform | |
| - training-free | |
| - code | |
| language: | |
| - en | |
| pipeline_tag: text-generation | |
| # ForgeLM v1 | |
| > **A training-free architectural port of Qwen2.5-Coder-1.5B-Instruct with KeyStack transforms.** | |
| ForgeLM v1 is **not a trained model**. It is a **weight-transformed port** of [Qwen2.5-Coder-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-Coder-1.5B-Instruct), created by applying a series of closed-form mathematical transforms (called "Keys") that restructure the architecture without losing the original model's knowledge. | |
| --- | |
| ## ⚠️ Disclaimers | |
| ### Port Notice | |
| **This model is a port of Qwen2.5-Coder-1.5B-Instruct, not an independently trained model.** All of the model's knowledge, capabilities, and limitations come from the original Qwen model. The KeyStack transforms are mathematically lossless or near-lossless at initialization — they restructure the architecture but do not add new knowledge. The original Qwen2.5-Coder-1.5B-Instruct model is licensed under Apache 2.0 by Alibaba/Qwen Team. | |
| ### Vibe-Coded Notice | |
| This model and its entire codebase were **vibe-coded via [Devin](https://devin.ai) Desktop** — an AI coding agent by Cognition. No human wrote the transform code, inference engine, or model architecture. The project was directed through natural language prompts and the AI agent implemented everything: weight porting, KeyStack transforms, inference engine, fine-tuning scripts, and this documentation. | |
| ### System Support Disclaimer | |
| - **Tested on**: Windows 11, NVIDIA RTX 5070 (12GB VRAM), CUDA 13.1, Python 3.11, PyTorch 2.8 | |
| - **GPU**: Requires NVIDIA GPU with ≥8GB VRAM for fast path. CPU-only or <8GB VRAM uses AirLLM layer-streaming (slow but works). | |
| - **OS**: Windows 11 tested. Linux/macOS should work but are untested. | |
| - **torch.compile**: Works on GQA attention. MoE + torch.compile hits a Triton bug on Windows (use without compile for MoE). | |
| - **No warranty**: This is a research project. Do not use in production without thorough testing. | |
| --- | |
| ## Architecture | |
| | Component | Original (Qwen) | ForgeLM v1 | Transform | | |
| |-----------|----------------|------------|-----------| | |
| | **Attention** | GQA (12 Q heads, 2 KV heads) | MLA (d_c=512) | SVD-based GQA→MLA | | |
| | **FFN** | Dense SwiGLU (8960 intermediate) | MoE (4 routed + 1 shared, d_ff=1792) | Weight splitting | | |
| | **Norm** | RMSNorm | RMSNorm (unchanged) | Direct copy | | |
| | **Embedding** | 151936 vocab, 1536 dim | Same (unchanged) | Direct copy | | |
| | **LM Head** | Tied with embedding | Same (unchanged) | Direct copy | | |
| | **RoPE** | θ=1,000,000 | Same (unchanged) | Direct copy | | |
| ### KeyStack Transforms Applied | |
| All transforms are **lossless at initialization** (cosine similarity ≥ 0.9998 vs original Qwen): | |
| | Key | Type | Effect | Cosine Sim | | |
| |-----|------|--------|------------| | |
| | **MLA** | FULL | GQA→MLA via SVD, 4x KV cache compression | 1.0000 (100% energy) | | |
| | **MoE** | FULL | Dense FFN→5 experts, 60% active FLOPs | 1.0000 (exact split) | | |
| | **MRL** | FULL | Matryoshka dimension reordering | 1.0000 (permutation) | | |
| | **QuaRot** | FULL | Hadamard rotation on V/O | 1.0000 (rotation) | | |
| | **ValueResidual** | FULL | V0 stored + 28 gates (init=0) | 1.0000 (no-op at init) | | |
| | **RotorQuant** | FULL | Givens rotation matrices stored | 1.0000 (no-op at init) | | |
| | **MTP** | FULL | 4 prediction heads from LM head | 1.0000 (identity init) | | |
| | **AirLLM** | TRIVIAL | Streamable flag for low-VRAM | N/A (runtime) | | |
| ### What Was NOT Applied (and why) | |
| | Technique | Reason | | |
| |-----------|--------| | |
| | **GQA→MQA** | Catastrophic degradation without 5% pretraining compute | | |
| | **Wanda pruning** | Needs calibration data post-MRL | | |
| | **SliceGPT** | Needs calibration data | | |
| | **BitNet** | Needs native ternary training | | |
| | **ShortGPT** | Available as key, not applied to v1 | | |
| --- | |
| ## Performance | |
| ### Quality | |
| - **Cosine similarity vs Qwen2.5-Coder-1.5B**: 0.9998–1.0000 across all layers | |
| - **Generation**: Identical output to original Qwen on all test prompts | |
| - **No quality loss** from any KeyStack transform | |
| ### Speed (NVIDIA RTX 5070, 12GB VRAM) | |
| | Configuration | tok/s | KV Compression | Notes | | |
| |---------------|-------|----------------|-------| | |
| | Standard | 21 | 1x | Baseline | | |
| | Hadamard INT4 KV | 21 | 4x | Lossless | | |
| | Streaming KV | 12 | ∞ (sinks+window) | Infinite context | | |
| | RotorQuant KV | 13 | 3.88x | 0.94% error | | |
| | MTP self-spec | 11 | 1x | Correct output, needs batch verify | | |
| ### Memory | |
| - **Checkpoint**: 3.4 GB (bf16) | |
| - **GPTQ INT4**: 2.3 GB (63% of original, cos=0.99) | |
| - **VRAM (inference)**: ~6 GB (model + KV cache) | |
| - **Min VRAM**: 4 GB (AirLLM streaming, very slow) | |
| --- | |
| ## KeyStack System | |
| The KeyStack is a system of **Keys** — closed-form mathematical transforms that convert weights between architectures without training. Each Key implements: | |
| - `forward(data) → weights` — replace training with instant transform | |
| - `reverse(weights) → data` — extract what was learned | |
| - `classify() → KeyClass` — FULL (reversible+composable), PARTIAL (lossy), TRIVIAL (no weights) | |
| ### Key Classification | |
| | Class | Criteria | Examples | | |
| |-------|----------|----------| | |
| | **FULL** | Reversible + data→weight + composable | MLA, MoE, MRL, QuaRot, ValueResidual, RotorQuant, MTP | | |
| | **BI** | Both directions, round-trip identity | Embedding, RMSNorm, LM Head, RoPE, CausalMask | | |
| | **PARTIAL** | One direction only (lossy) | GPTQ, ShortGPT, Wanda, SliceGPT | | |
| | **TRIVIAL** | No weights (runtime/formula) | AirLLM, QK-Norm, LogitCap, StreamingLLM | | |
| --- | |
| ## Runtime Features | |
| ForgeLM v1 includes a unified inference engine (`ForgeEngine`) with pluggable strategies: | |
| ### KV Cache Strategies | |
| - `standard` — Basic tensor cache | |
| - `streaming` — StreamingLLM (attention sinks + sliding window, infinite context) | |
| - `snapkv` — Observation-window eviction (8x compression) | |
| - `hadamard_int4` — Block-diagonal Hadamard + INT4 (4x compression, lossless) | |
| - `rotorquant` — Givens rotation + Lloyd-Max quantization (3.88x compression) | |
| - `compressed` — H2O heavy-hitter eviction + KV quantization | |
| - `paged` — vLLM-style paged memory | |
| ### Decoding Strategies | |
| - `standard` — Autoregressive token-by-token | |
| - `mtp_selfspec` — MTP self-speculative decoding (draft + verify) | |
| - `speculative` — Draft model speculative decoding | |
| - `medusa` — Medusa multi-head decoding | |
| - `dspark` — DSpark semi-autoregressive decoding | |
| ### Acceleration | |
| - `airllm_streaming` — Smart layer-streaming (only when VRAM insufficient) | |
| - `cuda_graph` — CUDA graph capture for decode step | |
| - `torch.compile` — Inductor compilation (1.3-2x on GQA) | |
| - `prefix_cache` — KV cache reuse for repeated prompt prefixes | |
| --- | |
| ## Usage | |
| ### Quick Start | |
| ```python | |
| from research.inference.forge_engine import ForgeEngine | |
| # Load ForgeLM v1 | |
| engine = ForgeEngine.from_checkpoint( | |
| checkpoint="forgelm_v1.safetensors", | |
| config_name="forgelm_v1", | |
| device="cuda", | |
| ) | |
| # Activate with runtime strategies | |
| engine.activate( | |
| kv_cache="hadamard_int4", # 4x KV compression, lossless | |
| decoding="standard", | |
| use_prefix_cache=True, | |
| ) | |
| # Generate | |
| output = engine.generate("def fibonacci(n):", max_new_tokens=100) | |
| print(output) | |
| ``` | |
| ### With Streaming KV (infinite context) | |
| ```python | |
| engine.activate( | |
| kv_cache="streaming", # 4 sinks + 512 window | |
| decoding="standard", | |
| ) | |
| output = engine.generate(long_prompt, max_new_tokens=1000) | |
| ``` | |
| ### With MTP Self-Speculative Decoding | |
| ```python | |
| from research.mtp import MTPHead | |
| from safetensors.torch import load_file | |
| # Load fine-tuned MTP head | |
| mtp_state = load_file("forgelm_v1_mtp.safetensors") | |
| mtp_head = MTPHead(d_model=1536, vocab_size=151936, n_predict=4).cuda() | |
| mtp_head.load_state_dict(mtp_state) | |
| engine.model.mtp_head = mtp_head | |
| engine.activate( | |
| kv_cache="standard", | |
| decoding="mtp_selfspec", # Speculative decoding | |
| ) | |
| ``` | |
| ### AirLLM Streaming (low VRAM) | |
| ```python | |
| engine = ForgeEngine.from_checkpoint( | |
| checkpoint="forgelm_v1.safetensors", | |
| config_name="forgelm_v1", | |
| device="cuda", | |
| ) | |
| engine.activate( | |
| kv_cache="standard", | |
| decoding="standard", | |
| acceleration="airllm_streaming", # Auto-detects: streams only if VRAM < model size | |
| ) | |
| ``` | |
| --- | |
| ## Files | |
| | File | Description | | |
| |------|-------------| | |
| | `forgelm_v1.safetensors` | Model weights (MLA + MoE, 3.4 GB) | | |
| | `config.json` | HF-compatible model config | | |
| | `tokenizer.json` | Qwen2.5 tokenizer (unchanged) | | |
| | `tokenizer_config.json` | Tokenizer config | | |
| | `vocab.json` | Vocabulary | | |
| | `merges.txt` | BPE merges | | |
| | `README.md` | This file | | |
| --- | |
| ## Technical Details | |
| ### MLA Conversion | |
| GQA K/V projections are factored through a shared low-rank bottleneck via SVD: | |
| ``` | |
| W_KV = [W_K; W_V] → SVD → W_c (compress) + W_KC, W_VC (decompress) | |
| ``` | |
| With d_c=512 (rank of GQA KV), 100% energy is retained. | |
| ### MoE Conversion | |
| Dense SwiGLU FFN is split into 5 experts (4 routed + 1 shared): | |
| ``` | |
| Dense: Y = W_down(silu(W_gate(X)) * W_up(X)) [intermediate=8960] | |
| MoE: Y = sum_i gate_i * Expert_i(X) + Shared(X) | |
| Each expert: d_ff = 8960 / 5 = 1792 | |
| ``` | |
| With top-4 routing (all experts active), output is identical to dense. Top-2 routing (60% FLOPs) requires router fine-tuning. | |
| ### MTP Fine-tune | |
| MTP heads' shared trunk was fine-tuned for 200 steps (122 seconds): | |
| - **Trainable params**: 2.36M (trunk only, heads frozen from LM head) | |
| - **Loss**: 10.6 → 2.0 (81% reduction) | |
| - **Result**: Correct speculative decoding output | |
| --- | |
| ## Citation | |
| If you use ForgeLM v1, please cite both the original Qwen model and this work: | |
| ```bibtex | |
| @misc{forgelm_v1, | |
| title={ForgeLM v1: Training-Free Architectural Port of Qwen2.5-Coder via KeyStack Transforms}, | |
| author={ForgeAI Project}, | |
| year={2026}, | |
| note={Vibe-coded via Devin Desktop (Cognition)} | |
| } | |
| @misc{qwen25_coder, | |
| title={Qwen2.5-Coder Technical Report}, | |
| author={Qwen Team}, | |
| year={2024}, | |
| url={https://huggingface.co/Qwen/Qwen2.5-Coder-1.5B-Instruct} | |
| } | |
| ``` | |
| --- | |
| ## License | |
| Apache 2.0 (inherited from Qwen2.5-Coder-1.5B-Instruct) | |
| --- | |
| ## Acknowledgments | |
| - **Qwen Team** (Alibaba) for the original Qwen2.5-Coder-1.5B-Instruct model | |
| - **Devin** (Cognition) for the AI coding agent that implemented everything | |
| - **KeyStack research** — the concept of closed-form weight transforms replacing training | |