--- language: en pipeline_tag: text-generation library_name: pytorch tags: - causal-lm - small-language-model - research - text-generation license: other thumbnail: Aurora-5.png --- # Vortex Alpha ![Vortex Alpha banner](Aurora-5.png) [![Open In Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/#create=true&url=https://huggingface.co/North-ML1/vortex-alpha/resolve/main/Vortex_Alpha_Colab.ipynb) Vortex is the working name for a compact, experimental language model. The final public name has not been decided. This release is intended for research, local experimentation, and further fine-tuning—not as a finished general assistant. ## What is included - `model.safetensors`: the instruction/tool-format preview selected from the best small internal behavior pilot. - `base_model.safetensors`: the corresponding pretrained text-completion base. - `config.json`: the architecture configuration. - `tokenizer.model`: the 8,192-piece SentencePiece tokenizer used for both checkpoints. - `vortex_model.py` and `inference.py`: a minimal dependency-light PyTorch loader and sampler. - `configuration_vortex.py`, `modeling_vortex.py`, and `tokenization_vortex.py`: standard Transformers remote-code modules for `AutoModelForCausalLM` and `AutoTokenizer`. - `Vortex_Alpha_Colab.ipynb`: a one-click Google Colab quickstart. - `requirements.txt` and `chat_template.jinja`: convenience metadata for local and notebook use. - `Aurora-5.png`: the project thumbnail/banner. Optimizer state, private logs, local paths, credentials, and training-machine metadata are intentionally not included. ## Architecture Vortex is a dense decoder-only Transformer with 174,942,720 trainable parameters: | Component | Parameters | | --- | ---: | | Shared token embedding and tied output head | 8,388,608 | | Attention Q projections | 12,582,912 | | Attention K projections | 3,145,728 | | Attention V projections | 3,145,728 | | Attention output projections | 12,582,912 | | Per-head QK RMSNorm parameters | 1,536 | | SwiGLU feed-forward networks | 135,069,696 | | Transformer-block RMSNorm parameters | 24,576 | | Final RMSNorm | 1,024 | | **Total** | **174,942,720** | Configuration: 12 layers, hidden size 1,024, 16 query heads, 4 key/value heads, 64-dimensional heads, SwiGLU with intermediate size 3,664, pre-layer RMSNorm, per-head QK-Norm, RoPE with base 100,000, bias-free projections, 8,192-token vocabulary, and a 4,096-token training/inference limit. The published weights use the readable reference PyTorch layout. They do not require Transformer Engine to load. The input and output embeddings are tied. ## Quick start The minimal reference runner uses PyTorch, SentencePiece, and `safetensors`: ```bash python -m pip install torch sentencepiece safetensors python inference.py \ --weights model.safetensors \ --chat \ --prompt "Explain why the sky appears blue in two short paragraphs." ``` For the normal Hugging Face API, load the custom architecture through the repository's small remote-code modules: ```python from transformers import AutoModelForCausalLM, AutoTokenizer repo = "North-ML1/vortex-alpha" tokenizer = AutoTokenizer.from_pretrained(repo, trust_remote_code=True) model = AutoModelForCausalLM.from_pretrained( repo, trust_remote_code=True, torch_dtype="auto" ) inputs = tokenizer("Explain why the sky appears blue.", return_tensors="pt") outputs = model.generate(**inputs, max_new_tokens=80, do_sample=False) print(tokenizer.decode(outputs[0], skip_special_tokens=True)) ``` The tokenizer also exposes the chat template directly: ```python messages = [{"role": "user", "content": "What is photosynthesis?"}] chat_inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, return_tensors="pt" ) outputs = model.generate(**chat_inputs, max_new_tokens=80, do_sample=False) print(tokenizer.decode(outputs[0], skip_special_tokens=True)) ``` The `trust_remote_code=True` flag is required because Vortex's QK-Norm and GQA implementation is not one of the built-in Transformers model classes. For ordinary next-token completion, use the base checkpoint: ```bash python inference.py \ --weights base_model.safetensors \ --prompt "The sky appears blue because" ``` The reference runner recomputes the full prefix at each generated token and is deliberately simple. A production runner should add a KV cache and a fused attention implementation. The instruction preview was tuned with this compact serialization: ```text [SYSTEM] You are a helpful assistant. Follow instructions, answer clearly, and say when information is missing. [USER] Your question here [ASSISTANT] ``` The instruction preview may emit a `CALL {json}` calculator/search request when prompted for tool use. No tool server is included in this repository; without a tool runner, treat such output as ordinary text. The base checkpoint is the better starting point for continued pretraining. ## Training summary The base run used an approximate mixture of FineWeb, DCLM, educational/math material, and The Stack v3 code data. The recorded base checkpoint had seen about 8.48 billion pretraining tokens. Training used BF16 model computation, FP8-capable NVIDIA kernels where available, a WSD-style learning-rate tail, and token-budgeted batches designed for a 16 GB consumer GPU. The instruction preview is a lightweight supervised derivative of that base; its optimizer state and private training records are not part of this release. These data-mixture descriptions are a project-level summary, not a claim that every upstream document is suitable for every downstream use. Follow the licenses and terms of the upstream datasets. ## Evaluation snapshot These are exploratory measurements, not official leaderboard submissions. The base results used greedy decoding, no tools, and the stated sample sizes: | Test | Result | Notes | | --- | ---: | --- | | MMLU cloze sample | 511/2,000 = 25.55% | Wilson 95% interval: 23.69–27.51% | | GSM8K strict numeric sample | 0/256 = 0.00% | Wilson 95% upper bound: 1.48% | | GSM8K fallback numeric sample | 2/256 = 0.78% | Wilson 95% interval: 0.21–2.80% | The instruction checkpoint reached 9/12 arithmetic, 4/4 grounding, 4/4 abstention, 2/4 exact-format, and 1/2 JSON checks on a 26-prompt internal tool-format pilot when the calculator runner was available. That pilot is too small to support general capability claims, and the arithmetic result is not comparable to tool-free GSM8K. The results show why this is an alpha release: the model can produce useful local completions and structured tool calls, but it remains weak at reliable arithmetic, broad knowledge, long-form coherence, and hallucination control. ## Limitations and intended use Vortex is a small research model. It can be repetitive, overconfident, or factually wrong; architecture and training scale do not guarantee reliable answers. Do not use it as the sole basis for medical, legal, financial, safety-critical, or other high-stakes decisions. It has not been evaluated for privacy, bias, cybersecurity, or comprehensive safety. The repository is public, but no open-source license is asserted yet. The final name, licensing terms, and a production release decision are still open. ## Reproducibility note The conversion removed optimizer state and Transformer Engine-only auxiliary state, then wrote the model tensors in BF16 safetensors format. The exported reference tensors preserve the tied-embedding model weights and can be loaded with the included `vortex_model.py`.