File size: 1,874 Bytes
c829dc9
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
---
license: apache-2.0
language:
- en
- ru
pipeline_tag: text-generation
tags:
- custom-code
- cubic-attention
- muon
- dpo
---

# Cubic Hier 150M

An experimental bilingual dialogue/reasoning language model with approximately
157M parameters. It keeps the faithful CubicV7 differentiable depth-memory path,
uses hierarchical cosine retrieval in the two penultimate attention layers, and
finishes with full global causal attention. Every sequence head uses QK RMSNorm.

## Training recipe

1. Base pretraining on a 72/13/10/5 stream of FineWeb-Edu, Russian FineWeb-2,
   deduplicated CodeParrot Python and FineMath-4+.
2. Assistant-only SFT on Smol-SmolTalk and English/Russian Aya examples.
3. Assistant-only reasoning SFT on the shortest verified-correct traces from
   OpenR1-Math-220k.
4. DPO with an SFT anchor on Anthropic HH-RLHF and corrected UltraFeedback.

The optimizer is the project's hybrid orthogonalized Muon (large matrices and
a boosted depth group) plus AdamW for embeddings, norms, scalars and gates.
Base training uses warmup-stable-decay and a one-token-ahead MTP auxiliary loss;
embeddings, RMSNorm scales and controls are excluded from AdamW decay.
SFT teaches explicit direct and `<think>...</think>` system-prompt modes; set
`CUBIC_REASONING=1` in chat mode to request the reasoning format.

## Important

This is a custom PyTorch architecture, not a drop-in Transformers model. Run
`train_and_chat.py` with `CUBIC_MODE=chat`; it downloads/loads all required
files and starts an interactive console. The model is small and experimental:
verify factual, safety-critical and mathematical answers independently.

Each source dataset retains its own license/terms. FineWeb corpora inherit the
Common Crawl terms described on their dataset cards; SmolTalk, Aya and OpenR1
are Apache-2.0; HH-RLHF and UltraFeedback use their published dataset licenses.