File size: 7,609 Bytes
d3ab6f5 4e013bc bd25748 d3ab6f5 bd25748 d3ab6f5 35fd7fa bd25748 35fd7fa d3ab6f5 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 | ---
language: en
pipeline_tag: text-generation
library_name: pytorch
tags:
- causal-lm
- small-language-model
- research
- text-generation
license: other
thumbnail: Aurora-5.png
---
# Vortex Alpha

[](https://colab.research.google.com/#create=true&url=https://huggingface.co/North-ML1/vortex-alpha/resolve/main/Vortex_Alpha_Colab.ipynb)
Vortex is the working name for a compact, experimental language model. The
final public name has not been decided. This release is intended for research,
local experimentation, and further fine-tuning—not as a finished general
assistant.
## What is included
- `model.safetensors`: the instruction/tool-format preview selected from the
best small internal behavior pilot.
- `base_model.safetensors`: the corresponding pretrained text-completion
base.
- `config.json`: the architecture configuration.
- `tokenizer.model`: the 8,192-piece SentencePiece tokenizer used for both
checkpoints.
- `vortex_model.py` and `inference.py`: a minimal dependency-light PyTorch
loader and sampler.
- `configuration_vortex.py`, `modeling_vortex.py`, and
`tokenization_vortex.py`: standard Transformers remote-code modules for
`AutoModelForCausalLM` and `AutoTokenizer`.
- `Vortex_Alpha_Colab.ipynb`: a one-click Google Colab quickstart.
- `requirements.txt` and `chat_template.jinja`: convenience metadata for
local and notebook use.
- `Aurora-5.png`: the project thumbnail/banner.
Optimizer state, private logs, local paths, credentials, and training-machine
metadata are intentionally not included.
## Architecture
Vortex is a dense decoder-only Transformer with 174,942,720 trainable
parameters:
| Component | Parameters |
| --- | ---: |
| Shared token embedding and tied output head | 8,388,608 |
| Attention Q projections | 12,582,912 |
| Attention K projections | 3,145,728 |
| Attention V projections | 3,145,728 |
| Attention output projections | 12,582,912 |
| Per-head QK RMSNorm parameters | 1,536 |
| SwiGLU feed-forward networks | 135,069,696 |
| Transformer-block RMSNorm parameters | 24,576 |
| Final RMSNorm | 1,024 |
| **Total** | **174,942,720** |
Configuration: 12 layers, hidden size 1,024, 16 query heads, 4 key/value
heads, 64-dimensional heads, SwiGLU with intermediate size 3,664, pre-layer
RMSNorm, per-head QK-Norm, RoPE with base 100,000, bias-free projections,
8,192-token vocabulary, and a 4,096-token training/inference limit.
The published weights use the readable reference PyTorch layout. They do not
require Transformer Engine to load. The input and output embeddings are tied.
## Quick start
The minimal reference runner uses PyTorch, SentencePiece, and
`safetensors`:
```bash
python -m pip install torch sentencepiece safetensors
python inference.py \
--weights model.safetensors \
--chat \
--prompt "Explain why the sky appears blue in two short paragraphs."
```
For the normal Hugging Face API, load the custom architecture through the
repository's small remote-code modules:
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "North-ML1/vortex-alpha"
tokenizer = AutoTokenizer.from_pretrained(repo, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
repo, trust_remote_code=True, torch_dtype="auto"
)
inputs = tokenizer("Explain why the sky appears blue.", return_tensors="pt")
outputs = model.generate(**inputs, max_new_tokens=80, do_sample=False)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
```
The tokenizer also exposes the chat template directly:
```python
messages = [{"role": "user", "content": "What is photosynthesis?"}]
chat_inputs = tokenizer.apply_chat_template(
messages, add_generation_prompt=True, return_tensors="pt"
)
outputs = model.generate(**chat_inputs, max_new_tokens=80, do_sample=False)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
```
The `trust_remote_code=True` flag is required because Vortex's QK-Norm and
GQA implementation is not one of the built-in Transformers model classes.
For ordinary next-token completion, use the base checkpoint:
```bash
python inference.py \
--weights base_model.safetensors \
--prompt "The sky appears blue because"
```
The reference runner recomputes the full prefix at each generated token and
is deliberately simple. A production runner should add a KV cache and a
fused attention implementation.
The instruction preview was tuned with this compact serialization:
```text
[SYSTEM]
You are a helpful assistant. Follow instructions, answer clearly, and say
when information is missing.
</s>
[USER]
Your question here
</s>
[ASSISTANT]
```
The instruction preview may emit a `CALL {json}` calculator/search request
when prompted for tool use. No tool server is included in this repository;
without a tool runner, treat such output as ordinary text. The base checkpoint
is the better starting point for continued pretraining.
## Training summary
The base run used an approximate mixture of FineWeb, DCLM, educational/math
material, and The Stack v3 code data. The recorded base checkpoint had seen
about 8.48 billion pretraining tokens. Training used BF16 model computation,
FP8-capable NVIDIA kernels where available, a WSD-style learning-rate tail,
and token-budgeted batches designed for a 16 GB consumer GPU. The instruction
preview is a lightweight supervised derivative of that base; its optimizer
state and private training records are not part of this release.
These data-mixture descriptions are a project-level summary, not a claim that
every upstream document is suitable for every downstream use. Follow the
licenses and terms of the upstream datasets.
## Evaluation snapshot
These are exploratory measurements, not official leaderboard submissions.
The base results used greedy decoding, no tools, and the stated sample sizes:
| Test | Result | Notes |
| --- | ---: | --- |
| MMLU cloze sample | 511/2,000 = 25.55% | Wilson 95% interval: 23.69–27.51% |
| GSM8K strict numeric sample | 0/256 = 0.00% | Wilson 95% upper bound: 1.48% |
| GSM8K fallback numeric sample | 2/256 = 0.78% | Wilson 95% interval: 0.21–2.80% |
The instruction checkpoint reached 9/12 arithmetic, 4/4 grounding, 4/4
abstention, 2/4 exact-format, and 1/2 JSON checks on a 26-prompt internal
tool-format pilot when the calculator runner was available. That pilot is too
small to support general capability claims, and the arithmetic result is not
comparable to tool-free GSM8K.
The results show why this is an alpha release: the model can produce useful
local completions and structured tool calls, but it remains weak at reliable
arithmetic, broad knowledge, long-form coherence, and hallucination control.
## Limitations and intended use
Vortex is a small research model. It can be repetitive, overconfident, or
factually wrong; architecture and training scale do not guarantee reliable
answers. Do not use it as the sole basis for medical, legal, financial,
safety-critical, or other high-stakes decisions. It has not been evaluated for
privacy, bias, cybersecurity, or comprehensive safety.
The repository is public, but no open-source license is asserted yet. The
final name, licensing terms, and a production release decision are still open.
## Reproducibility note
The conversion removed optimizer state and Transformer Engine-only auxiliary
state, then wrote the model tensors in BF16 safetensors format. The exported
reference tensors preserve the tied-embedding model weights and can be loaded
with the included `vortex_model.py`.
|