File size: 7,609 Bytes
d3ab6f5
 
 
 
 
 
 
 
 
 
 
 
 
 
 
4e013bc
 
bd25748
 
d3ab6f5
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
bd25748
 
 
 
 
 
d3ab6f5
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
35fd7fa
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
bd25748
 
 
 
 
 
 
 
 
 
 
35fd7fa
 
 
d3ab6f5
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
---
language: en
pipeline_tag: text-generation
library_name: pytorch
tags:
- causal-lm
- small-language-model
- research
- text-generation
license: other
thumbnail: Aurora-5.png
---

# Vortex Alpha

![Vortex Alpha banner](Aurora-5.png)

[![Open In Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/#create=true&url=https://huggingface.co/North-ML1/vortex-alpha/resolve/main/Vortex_Alpha_Colab.ipynb)

Vortex is the working name for a compact, experimental language model. The
final public name has not been decided. This release is intended for research,
local experimentation, and further fine-tuning—not as a finished general
assistant.

## What is included

- `model.safetensors`: the instruction/tool-format preview selected from the
  best small internal behavior pilot.
- `base_model.safetensors`: the corresponding pretrained text-completion
  base.
- `config.json`: the architecture configuration.
- `tokenizer.model`: the 8,192-piece SentencePiece tokenizer used for both
  checkpoints.
- `vortex_model.py` and `inference.py`: a minimal dependency-light PyTorch
  loader and sampler.
- `configuration_vortex.py`, `modeling_vortex.py`, and
  `tokenization_vortex.py`: standard Transformers remote-code modules for
  `AutoModelForCausalLM` and `AutoTokenizer`.
- `Vortex_Alpha_Colab.ipynb`: a one-click Google Colab quickstart.
- `requirements.txt` and `chat_template.jinja`: convenience metadata for
  local and notebook use.
- `Aurora-5.png`: the project thumbnail/banner.

Optimizer state, private logs, local paths, credentials, and training-machine
metadata are intentionally not included.

## Architecture

Vortex is a dense decoder-only Transformer with 174,942,720 trainable
parameters:

| Component | Parameters |
| --- | ---: |
| Shared token embedding and tied output head | 8,388,608 |
| Attention Q projections | 12,582,912 |
| Attention K projections | 3,145,728 |
| Attention V projections | 3,145,728 |
| Attention output projections | 12,582,912 |
| Per-head QK RMSNorm parameters | 1,536 |
| SwiGLU feed-forward networks | 135,069,696 |
| Transformer-block RMSNorm parameters | 24,576 |
| Final RMSNorm | 1,024 |
| **Total** | **174,942,720** |

Configuration: 12 layers, hidden size 1,024, 16 query heads, 4 key/value
heads, 64-dimensional heads, SwiGLU with intermediate size 3,664, pre-layer
RMSNorm, per-head QK-Norm, RoPE with base 100,000, bias-free projections,
8,192-token vocabulary, and a 4,096-token training/inference limit.

The published weights use the readable reference PyTorch layout. They do not
require Transformer Engine to load. The input and output embeddings are tied.

## Quick start

The minimal reference runner uses PyTorch, SentencePiece, and
`safetensors`:

```bash
python -m pip install torch sentencepiece safetensors
python inference.py \
  --weights model.safetensors \
  --chat \
  --prompt "Explain why the sky appears blue in two short paragraphs."
```

For the normal Hugging Face API, load the custom architecture through the
repository's small remote-code modules:

```python
from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "North-ML1/vortex-alpha"
tokenizer = AutoTokenizer.from_pretrained(repo, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    repo, trust_remote_code=True, torch_dtype="auto"
)
inputs = tokenizer("Explain why the sky appears blue.", return_tensors="pt")
outputs = model.generate(**inputs, max_new_tokens=80, do_sample=False)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
```

The tokenizer also exposes the chat template directly:

```python
messages = [{"role": "user", "content": "What is photosynthesis?"}]
chat_inputs = tokenizer.apply_chat_template(
    messages, add_generation_prompt=True, return_tensors="pt"
)
outputs = model.generate(**chat_inputs, max_new_tokens=80, do_sample=False)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
```

The `trust_remote_code=True` flag is required because Vortex's QK-Norm and
GQA implementation is not one of the built-in Transformers model classes.

For ordinary next-token completion, use the base checkpoint:

```bash
python inference.py \
  --weights base_model.safetensors \
  --prompt "The sky appears blue because"
```

The reference runner recomputes the full prefix at each generated token and
is deliberately simple. A production runner should add a KV cache and a
fused attention implementation.

The instruction preview was tuned with this compact serialization:

```text
[SYSTEM]
You are a helpful assistant. Follow instructions, answer clearly, and say
when information is missing.
</s>
[USER]
Your question here
</s>
[ASSISTANT]
```

The instruction preview may emit a `CALL {json}` calculator/search request
when prompted for tool use. No tool server is included in this repository;
without a tool runner, treat such output as ordinary text. The base checkpoint
is the better starting point for continued pretraining.

## Training summary

The base run used an approximate mixture of FineWeb, DCLM, educational/math
material, and The Stack v3 code data. The recorded base checkpoint had seen
about 8.48 billion pretraining tokens. Training used BF16 model computation,
FP8-capable NVIDIA kernels where available, a WSD-style learning-rate tail,
and token-budgeted batches designed for a 16 GB consumer GPU. The instruction
preview is a lightweight supervised derivative of that base; its optimizer
state and private training records are not part of this release.

These data-mixture descriptions are a project-level summary, not a claim that
every upstream document is suitable for every downstream use. Follow the
licenses and terms of the upstream datasets.

## Evaluation snapshot

These are exploratory measurements, not official leaderboard submissions.
The base results used greedy decoding, no tools, and the stated sample sizes:

| Test | Result | Notes |
| --- | ---: | --- |
| MMLU cloze sample | 511/2,000 = 25.55% | Wilson 95% interval: 23.69–27.51% |
| GSM8K strict numeric sample | 0/256 = 0.00% | Wilson 95% upper bound: 1.48% |
| GSM8K fallback numeric sample | 2/256 = 0.78% | Wilson 95% interval: 0.21–2.80% |

The instruction checkpoint reached 9/12 arithmetic, 4/4 grounding, 4/4
abstention, 2/4 exact-format, and 1/2 JSON checks on a 26-prompt internal
tool-format pilot when the calculator runner was available. That pilot is too
small to support general capability claims, and the arithmetic result is not
comparable to tool-free GSM8K.

The results show why this is an alpha release: the model can produce useful
local completions and structured tool calls, but it remains weak at reliable
arithmetic, broad knowledge, long-form coherence, and hallucination control.

## Limitations and intended use

Vortex is a small research model. It can be repetitive, overconfident, or
factually wrong; architecture and training scale do not guarantee reliable
answers. Do not use it as the sole basis for medical, legal, financial,
safety-critical, or other high-stakes decisions. It has not been evaluated for
privacy, bias, cybersecurity, or comprehensive safety.

The repository is public, but no open-source license is asserted yet. The
final name, licensing terms, and a production release decision are still open.

## Reproducibility note

The conversion removed optimizer state and Transformer Engine-only auxiliary
state, then wrote the model tensors in BF16 safetensors format. The exported
reference tensors preserve the tied-embedding model weights and can be loaded
with the included `vortex_model.py`.