Text Generation
PyTorch
English
jl_code_python
code
python
python-code
code-generation
code-completion
causal-language-modeling
decoder-only
programming-language-model
small-language-model
from-scratch
random-initialization
custom-architecture
custom-pytorch
byte-level-bpe
rope
rmsnorm
swiglu
consumer-gpu
rtx-3060
research
jumplander
jl-code
| license: apache-2.0 | |
| language: | |
| - en | |
| pipeline_tag: text-generation | |
| datasets: | |
| - codeparrot/codeparrot-clean | |
| - google-research-datasets/mbpp | |
| tags: | |
| - code | |
| - python | |
| - python-code | |
| - code-generation | |
| - code-completion | |
| - causal-language-modeling | |
| - decoder-only | |
| - programming-language-model | |
| - small-language-model | |
| - from-scratch | |
| - random-initialization | |
| - custom-architecture | |
| - custom-pytorch | |
| - pytorch | |
| - byte-level-bpe | |
| - rope | |
| - rmsnorm | |
| - swiglu | |
| - consumer-gpu | |
| - rtx-3060 | |
| - research | |
| - jumplander | |
| - jl-code | |
| <p align="center"> | |
| <a href="https://jumplander.org/en/home"> | |
| <img | |
| src="https://cdn-avatars.huggingface.co/v1/production/uploads/69204763af796f2f22ad9f49/loC_Dutp1Rb4jHIlGsbkG.png" | |
| width="150" | |
| alt="JumpLander logo" | |
| /> | |
| </a> | |
| </p> | |
| <h1 align="center">JL-Code-Python-97M</h1> | |
| <p align="center"> | |
| <strong>A compact Python-focused causal language model trained from random initialization by JumpLander.</strong> | |
| </p> | |
| <p align="center"> | |
| <a href="https://huggingface.co/jumplander/JL-Code-Python-97M"> | |
| <img src="https://img.shields.io/badge/Hugging%20Face-JL--Code--Python--97M-FFD21E?logo=huggingface&logoColor=000000" alt="Hugging Face model"/> | |
| </a> | |
| <a href="https://jumplander.org/en/home"> | |
| <img src="https://img.shields.io/badge/JumpLander-Official%20Website-819e2e" alt="JumpLander website"/> | |
| </a> | |
| <a href="https://huggingface.co/jumplander"> | |
| <img src="https://img.shields.io/badge/Organization-jumplander-4b5d2a" alt="JumpLander Hugging Face"/> | |
| </a> | |
| </p> | |
| <p align="center"> | |
| <img src="https://img.shields.io/badge/Parameters-97.54M-28392b" alt="97.54M parameters"/> | |
| <img src="https://img.shields.io/badge/Language-Python-3776AB?logo=python&logoColor=white" alt="Python"/> | |
| <img src="https://img.shields.io/badge/Training-From%20Scratch-0c0c0e" alt="From scratch"/> | |
| <img src="https://img.shields.io/badge/GPU-RTX%203060%2012GB-76B900?logo=nvidia&logoColor=white" alt="RTX 3060"/> | |
| <img src="https://img.shields.io/badge/Release-v0.1--base-819e2e" alt="v0.1 base"/> | |
| </p> | |
| --- | |
| ## Overview | |
| **JL-Code-Python-97M** is an experimental, Python-specialized decoder-only language model developed by [JumpLander](https://jumplander.org/en/home). | |
| The checkpoint contains exactly **97,536,768 trainable parameters**—approximately **97.54 million parameters**. It was trained from **randomly initialized weights** and does not inherit pretrained weights from Qwen, Llama, DeepSeek, Code Llama, GPT, or another external model family. | |
| The release focuses on a narrow technical domain: | |
| - Python source-code continuation | |
| - Python function completion | |
| - technical English associated with Python code | |
| - docstring-to-code patterns | |
| - short English-to-Python programming tasks | |
| - fill-in-the-middle code reconstruction | |
| This is a **base research checkpoint**, not a general chat model and not a production coding assistant. | |
| > **Compatibility notice:** this release uses a custom PyTorch architecture and a custom `.pt` checkpoint. It is not currently loadable through `AutoModelForCausalLM.from_pretrained(...)`. Use the included [`jumplander_python_100m.py`](./jumplander_python_100m.py) file for loading and inference. | |
| --- | |
| ## JL-Code model family | |
| `JL-Code` is JumpLander's language-specific code-model line. Each branch is intended to focus on one programming language rather than mixing every language into a single small checkpoint. | |
| ```text | |
| JumpLander | |
| └── JL-Code | |
| ├── JL-Code-Python-97M ← current release | |
| ├── JL-Code-JavaScript-* ← planned | |
| ├── JL-Code-PHP-* ← planned | |
| ├── JL-Code-Rust-* ← planned | |
| ├── JL-Code-C-* ← planned | |
| ├── JL-Code-Cpp-* ← planned | |
| └── JL-Code-SQL-* ← planned | |
| ``` | |
| Planned names are directional and do not imply that those checkpoints have already been released. | |
| --- | |
| ## At a glance | |
| | Property | Value | | |
| |---|---| | |
| | Repository | [`jumplander/JL-Code-Python-97M`](https://huggingface.co/jumplander/JL-Code-Python-97M) | | |
| | Developer | [JumpLander](https://jumplander.org/en/home) | | |
| | Model family | JL-Code | | |
| | Release | `v0.1-base` | | |
| | Status | Experimental research release | | |
| | Primary domain | Python | | |
| | Natural-language context | Technical English | | |
| | Architecture | Custom decoder-only Transformer | | |
| | Parameters | **97,536,768** | | |
| | Initialization | Random initialization | | |
| | External pretrained checkpoint | None | | |
| | Context length | 1,024 tokens | | |
| | Tokenizer | Custom Byte-Level BPE | | |
| | Vocabulary | 16,384 tokens | | |
| | Training objective | Causal language modeling | | |
| | Framework | PyTorch | | |
| | Training GPU | NVIDIA RTX 3060 12GB | | |
| | Checkpoint format | Custom PyTorch training checkpoint (`.pt`) | | |
| --- | |
| ## Architecture | |
| The model architecture is implemented in [`jumplander_python_100m.py`](./jumplander_python_100m.py). | |
| | Component | Configuration | | |
| |---|---:| | |
| | Transformer blocks | 12 | | |
| | Hidden dimension | 768 | | |
| | Attention heads | 12 | | |
| | Head dimension | 64 | | |
| | Feed-forward dimension | 2,048 | | |
| | Maximum sequence length | 1,024 | | |
| | Vocabulary size | 16,384 | | |
| | Position encoding | Rotary Position Embeddings (RoPE) | | |
| | RoPE theta | 10,000 | | |
| | Normalization | RMSNorm | | |
| | RMSNorm epsilon | `1e-5` | | |
| | MLP | SwiGLU | | |
| | Attention | Causal self-attention | | |
| | Attention implementation | PyTorch scaled dot-product attention | | |
| | Attention/MLP bias | Disabled | | |
| | Dropout | `0.0` | | |
| | Input/output embedding tying | Enabled | | |
| | Initialization standard deviation | `0.02` | | |
| | Training gradient checkpointing | Enabled | | |
| The exact machine-readable architecture is available in [`config.json`](./config.json). | |
| ### Parameter count | |
| ```text | |
| 97,536,768 trainable parameters | |
| ≈ 97.54M parameters | |
| ≈ 100M-class model | |
| ``` | |
| The repository name uses `97M` to reflect the exact architecture more honestly than rounding it up to 100M. | |
| --- | |
| ## Tokenizer | |
| The tokenizer was trained specifically for this project from the prepared local training corpus. | |
| | Property | Value | | |
| |---|---| | |
| | Algorithm | Byte-Level BPE | | |
| | Vocabulary target | 16,384 | | |
| | Minimum token frequency | 2 | | |
| | Byte fallback | Enabled | | |
| | Pre-tokenizer | ByteLevel | | |
| | Decoder | ByteLevel | | |
| Special tokens: | |
| ```text | |
| <pad> | |
| <unk> | |
| <bos> | |
| <eos> | |
| <file_start> | |
| <file_end> | |
| <fim_prefix> | |
| <fim_suffix> | |
| <fim_middle> | |
| <instruction> | |
| <response> | |
| ``` | |
| The tokenizer file required for inference is [`tokenizer.json`](./tokenizer.json). | |
| --- | |
| ## Training data | |
| The local data-building pipeline uses two public Hugging Face datasets. | |
| ### 1. CodeParrot Clean | |
| - Dataset: [`codeparrot/codeparrot-clean`](https://huggingface.co/datasets/codeparrot/codeparrot-clean) | |
| - Purpose: Python source-code pretraining and code-structure learning | |
| - Source type: deduplicated Python files collected from public GitHub repositories | |
| The upstream dataset includes per-file license metadata and contains code under multiple software licenses. Users should review the dataset card and original repository licenses before downstream commercial use. | |
| ### 2. MBPP | |
| - Dataset: [`google-research-datasets/mbpp`](https://huggingface.co/datasets/google-research-datasets/mbpp) | |
| - Alternate dataset card: [`Muennighoff/mbpp`](https://huggingface.co/datasets/Muennighoff/mbpp) | |
| - Purpose: short technical-English-to-Python examples and held-out programming tasks | |
| - Content: natural-language tasks, Python reference solutions, and automated tests | |
| MBPP examples assigned to test or validation splits were kept outside the main training split by the project pipeline. | |
| --- | |
| ## Data preparation pipeline | |
| The dataset was not passed directly into training. The local pipeline applied additional filtering and transformation: | |
| 1. stream Python records from Hugging Face; | |
| 2. normalize line endings and remove null bytes; | |
| 3. reject very short or oversized files; | |
| 4. detect likely credentials, keys, and private-key blocks; | |
| 5. reject generated or minified files; | |
| 6. validate Python using `ast.parse`; | |
| 7. split large source files at Python function/class boundaries where possible; | |
| 8. remove exact duplicates using SHA-256; | |
| 9. create deterministic train, validation, and test splits; | |
| 10. extract selected function/docstring pairs as English-to-Python instruction rows; | |
| 11. add MBPP examples according to their upstream split; | |
| 12. generate deterministic fill-in-the-middle variants for base training rows; | |
| 13. tokenize and pack the corpus into contiguous `uint16` token streams. | |
| ### Prepared dataset statistics | |
| | Statistic | Count | | |
| |---|---:| | |
| | Accepted base samples | 20,000 | | |
| | Final training rows | 35,855 | | |
| | Final validation rows | 389 | | |
| | Final test rows | 793 | | |
| | Exact duplicates rejected | 391 | | |
| | Generated files rejected | 246 | | |
| | Secret-like samples rejected | 75 | | |
| | Syntax-invalid samples rejected | 1,523 | | |
| | Oversized samples rejected | 268 | | |
| | Source records read | 10,809 | | |
| A single accepted source file can produce more than one training row—for example, a base-code sample plus extracted function/docstring instruction samples. This is why the final training-row count is larger than the accepted-base-sample count. | |
| --- | |
| ## Training configuration | |
| The model was trained using next-token prediction over packed token sequences. | |
| | Setting | Value | | |
| |---|---:| | |
| | Optimizer | AdamW | | |
| | Optimizer betas | `(0.9, 0.95)` | | |
| | Optimizer epsilon | `1e-8` | | |
| | Weight decay | `0.1` | | |
| | Peak learning rate | `3e-4` | | |
| | Minimum learning rate | `3e-5` | | |
| | Scheduler | Cosine decay | | |
| | Warmup | 200 optimizer steps | | |
| | Total optimizer steps | 10,000 | | |
| | Micro-batch size | 1 sequence | | |
| | Gradient accumulation | 32 | | |
| | Effective batch | 32 sequences | | |
| | Sequence length | 1,024 tokens | | |
| | Tokens per optimizer step | 32,768 | | |
| | Approximate tokens processed | 327,680,000 | | |
| | Gradient clipping | `1.0` | | |
| | Evaluation interval | 250 steps | | |
| | Checkpoint interval | 500 steps | | |
| | Random seed | 1,337 | | |
| | Precision | BF16 mixed precision | | |
| | Hardware | NVIDIA RTX 3060 12GB | | |
| Observed during training: | |
| - approximately **8,400–9,000 tokens/second**; | |
| - approximately **1.86 GB peak allocated VRAM** in the reported run; | |
| - stable gradient norms during the observed training window. | |
| The complete machine-readable run settings are in [`training_config.json`](./training_config.json). | |
| --- | |
| ## Repository files | |
| Recommended repository layout: | |
| ```text | |
| JL-Code-Python-97M/ | |
| ├── README.md | |
| ├── config.json | |
| ├── training_config.json | |
| ├── jumplander_python_100m.pt | |
| ├── tokenizer.json | |
| ├── jumplander_python_100m.py | |
| ├── requirements.txt | |
| └── DATA_SOURCES.md | |
| ``` | |
| | File | Purpose | | |
| |---|---| | |
| | [`README.md`](./README.md) | Hugging Face model card | | |
| | [`config.json`](./config.json) | Machine-readable architecture description | | |
| | [`training_config.json`](./training_config.json) | Training and data-preparation settings | | |
| | [`jumplander_python_100m.pt`](./jumplander_python_100m.pt) | Trained model checkpoint | | |
| | [`tokenizer.json`](./tokenizer.json) | Byte-Level BPE tokenizer | | |
| | [`jumplander_python_100m.py`](./jumplander_python_100m.py) | Model architecture, loading, generation, and web UI | | |
| | [`requirements.txt`](./requirements.txt) | Python dependencies | | |
| | [`DATA_SOURCES.md`](./DATA_SOURCES.md) | Additional data provenance notes | | |
| Do not upload `.venv`, `__pycache__`, Hugging Face caches, tokenized `.bin` training files, or private training logs unless they are intentionally part of the release. | |
| --- | |
| ## Download | |
| ### Hugging Face CLI | |
| ```bash | |
| pip install -U huggingface_hub | |
| hf download jumplander/JL-Code-Python-97M --local-dir JL-Code-Python-97M | |
| cd JL-Code-Python-97M | |
| ``` | |
| ### Git and Git LFS | |
| ```bash | |
| git lfs install | |
| git clone https://huggingface.co/jumplander/JL-Code-Python-97M | |
| cd JL-Code-Python-97M | |
| ``` | |
| --- | |
| ## Installation | |
| ```bash | |
| python -m venv .venv | |
| ``` | |
| Windows PowerShell: | |
| ```powershell | |
| .\.venv\Scripts\Activate.ps1 | |
| ``` | |
| Linux/macOS: | |
| ```bash | |
| source .venv/bin/activate | |
| ``` | |
| Install dependencies: | |
| ```bash | |
| pip install -r requirements.txt | |
| ``` | |
| Inspect the environment and local files: | |
| ```bash | |
| python jumplander_python_100m.py check | |
| ``` | |
| Inspect the architecture and exact parameter count: | |
| ```bash | |
| python jumplander_python_100m.py info | |
| ``` | |
| --- | |
| ## Inference | |
| ### Terminal generation | |
| Instruction-style prompt: | |
| ```bash | |
| python jumplander_python_100m.py generate \ | |
| "Write a Python function that returns unique list items while preserving order." | |
| ``` | |
| Code completion: | |
| ```bash | |
| python jumplander_python_100m.py generate \ | |
| "def fibonacci(n):" \ | |
| --mode completion \ | |
| --max-new-tokens 160 \ | |
| --temperature 0.2 \ | |
| --top-p 0.95 | |
| ``` | |
| Windows PowerShell single-line example: | |
| ```powershell | |
| python jumplander_python_100m.py generate "Write a Python function that checks whether a number is prime." | |
| ``` | |
| ### Local browser UI | |
| ```bash | |
| python jumplander_python_100m.py chat | |
| ``` | |
| Default address: | |
| ```text | |
| http://127.0.0.1:7860 | |
| ``` | |
| The UI provides two modes: | |
| - instruction-to-Python; | |
| - Python code completion. | |
| --- | |
| ## Checkpoint format | |
| `jumplander_python_100m.pt` is a custom PyTorch training checkpoint containing: | |
| ```text | |
| model_name | |
| model_config | |
| train_config | |
| model_state | |
| optimizer_state | |
| step | |
| tokens_seen | |
| saved_at | |
| format_version | |
| ``` | |
| Because the checkpoint includes optimizer state, it can be larger than a weights-only release. | |
| A future release should also provide: | |
| - a weights-only checkpoint; | |
| - `model.safetensors`; | |
| - native Hugging Face Transformers integration; | |
| - `AutoModelForCausalLM` loading support. | |
| --- | |
| ## Evaluation status | |
| This release documents the completed training run, but it does **not yet claim a verified HumanEval, MBPP pass@1, or production code-generation score**. | |
| Training loss alone does not establish code correctness. Proper evaluation should include: | |
| - Python AST parse rate; | |
| - executable completion rate; | |
| - unit-test pass rate; | |
| - MBPP pass@1 and pass@k; | |
| - HumanEval evaluation after contamination review; | |
| - repetition and memorization checks; | |
| - security-oriented code review. | |
| Until those evaluations are published, treat this checkpoint as an experimental base model. | |
| --- | |
| ## Intended uses | |
| Appropriate uses: | |
| - research on small code language models; | |
| - experiments with Python code completion; | |
| - educational study of from-scratch Transformer training; | |
| - tokenizer and data-pipeline research; | |
| - local inference experiments; | |
| - continued pretraining and instruction tuning; | |
| - analysis of consumer-GPU model development. | |
| --- | |
| ## Out-of-scope uses | |
| This model is not intended for: | |
| - production-critical code generation; | |
| - security-sensitive implementation without review; | |
| - autonomous deployment of generated code; | |
| - legal, medical, financial, or safety-critical systems; | |
| - generating or executing untrusted code without sandboxing; | |
| - replacing human code review and automated testing. | |
| --- | |
| ## Limitations | |
| - The model is small and narrowly trained. | |
| - It has limited general-world knowledge. | |
| - It may generate invalid, incomplete, insecure, or fabricated Python code. | |
| - English understanding is primarily tied to technical programming patterns. | |
| - The context window is limited to 1,024 tokens. | |
| - The training corpus is much smaller than corpora used for leading code models. | |
| - GitHub-derived data can contain bugs, insecure patterns, biases, or licensing constraints. | |
| - The current architecture is custom and is not yet integrated with Transformers. | |
| - Generated code must be reviewed, sandboxed, and tested before use. | |
| --- | |
| ## Safety and responsible use | |
| Never execute generated code directly in a privileged environment. | |
| Recommended controls: | |
| 1. run generated code in a restricted sandbox; | |
| 2. disable network access where possible; | |
| 3. apply CPU, memory, and execution-time limits; | |
| 4. scan generated code for secrets and unsafe operations; | |
| 5. require unit tests and human review; | |
| 6. avoid exposing private repositories or credentials to untrusted systems. | |
| --- | |
| ## Licensing and data notice | |
| The repository declares the Apache License 2.0 for JumpLander-authored code and released model artifacts. | |
| That license declaration does not replace or override the licenses attached to upstream source files used by the training datasets. `codeparrot/codeparrot-clean` includes per-file license metadata from many public repositories. Users are responsible for reviewing the upstream dataset cards, source-code licenses, attribution requirements, and applicable law before commercial or regulated use. | |
| MBPP is commonly distributed under CC BY 4.0; consult the current dataset card for attribution and reuse requirements. | |
| This section is informational and is not legal advice. | |
| --- | |
| ## Roadmap | |
| Potential follow-up work: | |
| - publish a weights-only `safetensors` checkpoint; | |
| - add Transformers-compatible configuration and model classes; | |
| - publish reproducible evaluation scripts; | |
| - evaluate on MBPP and HumanEval; | |
| - expand high-quality Python training data; | |
| - create a stronger instruction-tuned Python variant; | |
| - train larger Python-specific models; | |
| - extend the JL-Code family to additional programming languages. | |
| --- | |
| ## Links | |
| - **Model:** [huggingface.co/jumplander/JL-Code-Python-97M](https://huggingface.co/jumplander/JL-Code-Python-97M) | |
| - **JumpLander on Hugging Face:** [huggingface.co/jumplander](https://huggingface.co/jumplander) | |
| - **JumpLander website:** [jumplander.org](https://jumplander.org/en/home) | |
| - **JumpLander documentation:** [jumplander.org/fa/docs](https://jumplander.org/fa/docs) | |
| - **CodeParrot Clean:** [huggingface.co/datasets/codeparrot/codeparrot-clean](https://huggingface.co/datasets/codeparrot/codeparrot-clean) | |
| - **MBPP:** [huggingface.co/datasets/google-research-datasets/mbpp](https://huggingface.co/datasets/google-research-datasets/mbpp) | |
| --- | |
| ## Citation | |
| ```bibtex | |
| @misc{jumplander_jl_code_python_97m_2026, | |
| author = {{JumpLander}}, | |
| title = {JL-Code-Python-97M: A From-Scratch Python Code Language Model}, | |
| year = {2026}, | |
| publisher = {Hugging Face}, | |
| howpublished = {\url{https://huggingface.co/jumplander/JL-Code-Python-97M}}, | |
| note = {Experimental base release, version 0.1} | |
| } | |
| ``` | |
| --- | |
| ## Acknowledgements | |
| This project uses: | |
| - [PyTorch](https://pytorch.org/) for model implementation and training; | |
| - [Hugging Face Datasets](https://huggingface.co/docs/datasets/) for dataset access; | |
| - [Hugging Face Tokenizers](https://huggingface.co/docs/tokenizers/) for Byte-Level BPE; | |
| - [CodeParrot Clean](https://huggingface.co/datasets/codeparrot/codeparrot-clean) for Python source data; | |
| - [MBPP](https://huggingface.co/datasets/google-research-datasets/mbpp) for English-to-Python programming tasks. | |
| --- | |
| <p align="center"> | |
| Built by <a href="https://jumplander.org/en/home"><strong>JumpLander</strong></a> | |
| · Programming intelligence, code models, datasets, and developer systems | |
| </p> | |