File size: 10,251 Bytes
58061e8
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
### **BLRN (Byte-level Linear Recurrent Network) Hierarchical Specification**

---

#### **1. Input Layer**
*   **1.1. Input Sequence ($x_{\text{raw}}$)**
    *   *Tensor Shape*: $[B, T]$ (Batch Size $B$ × Sequence Length $T$)
    *   *Data Type*: Raw unsigned 8-bit integers representing raw UTF-8 bytes ($\text{uint8}$, values $0 \le x_t \le 255$).
    *   *Context Length*: Up to $1,024\text{K}$ bytes, processed in linear space and time complexity $O(T)$.

---

#### **2. Embedding Layer**
*   **2.1. Byte Embedding**
    *   *Tensor Shape*: $[B, T] \rightarrow [B, T, d_{\text{model}}]$ (where $d_{\text{model}} = 256$)
    *   *Mathematical Operation*: Coordinate lookup table projection
        $$E_t = W_{\text{emb}}[x_t] \in \mathbb{R}^{256}$$
    *   *Weight Dimension*: $W_{\text{emb}} \in \mathbb{R}^{256 \times d_{\text{model}}}$
    *   *Parameter Calculation*: 
        $$\text{Params} = 256 \times d_{\text{model}} = 256 \times 256 = 65,536$$

---

#### **3. Dense Recurrent Block (BlrnDenseDecoderLayer) × 3 Layers**
*   Each block consists of a sequential stack of a Recurrent Sub-block and an MLP Sub-block, both enclosed by residual skip connections.
*   **3.1. Recurrent Sub-block (Temporal Processing)**
    *   **3.1.1. Pre-Normalization (RMSNorm)**
        *   *Tensor Shape*: $[B, T, 256] \rightarrow [B, T, 256]$
        *   *Mathematical Operation*: Root Mean Square Normalization with a learnable scaling vector $\gamma_{\text{rec}}$
            $$x_{\text{norm\_rec}} = \text{RMSNorm}(X_{\text{in}}) \odot \gamma_{\text{rec}} \quad (\gamma_{\text{rec}} \in \mathbb{R}^{256})$$
        *   *Parameter Calculation*: 
            $$\text{Params} = d_{\text{model}} = 256$$
    *   **3.1.2. Gated Linear Recurrent Unit (GLRU)**
        *   *Tensor Shape*: $[B, T, 256] \rightarrow [B, T, 256]$
        *   *Mathematical Operation*: 1st-order linear recurrence sequence update
            *   Forget Gate: $f_t = \sigma(W_f \cdot x_{\text{norm\_rec}} + b_f)$
            *   Input Gate: $i_t = \sigma(W_i \cdot x_{\text{norm\_rec}} + b_i)$
            *   Candidate State: $u_t = \tanh(W_u \cdot x_{\text{norm\_rec}} + b_u)$
            *   State Transition: $h_t = f_t \odot h_{t-1} + i_t \odot u_t$
        *   *Weight & Bias Dimensions*: $W_f, W_i, W_u \in \mathbb{R}^{256 \times 256}$, $b_f, b_i, b_u \in \mathbb{R}^{256}$
        *   *Parameter Calculation*: 
            $$\text{Params} = 3 \times (d_{\text{model}} \times d_{\text{model}} + d_{\text{model}}) = 3 \times (256 \times 256 + 256) = 197,376$$
    *   **3.1.3. Residual Addition**
        *   *Tensor Shape*: $[B, T, 256] + [B, T, 256] \rightarrow [B, T, 256]$
        *   *Mathematical Operation*: Element-wise addition of identity skip connection
            $$X_{\text{mid}} = X_{\text{in}} + h_t$$
        *   *Parameter Calculation*: $0$
*   **3.2. MLP Sub-block (Channel Projection)**
    *   **3.2.1. Mid-Normalization (RMSNorm)**
        *   *Tensor Shape*: $[B, T, 256] \rightarrow [B, T, 256]$
        *   *Mathematical Operation*: RMSNorm before non-linear projection
            $$x_{\text{norm\_mlp}} = \text{RMSNorm}(X_{\text{mid}}) \odot \gamma_{\text{mlp}} \quad (\gamma_{\text{mlp}} \in \mathbb{R}^{256})$$
        *   *Parameter Calculation*: 
            $$\text{Params} = d_{\text{model}} = 256$$
    *   **3.2.2. Gated MLP (SwiGLU)**
        *   *Tensor Shape*: $[B, T, 256] \rightarrow [B, T, 1024] \rightarrow [B, T, 256]$ (where $d_{\text{ffn}} = 1024$)
        *   *Mathematical Operation*: Swish-Gated Linear Unit projection
            *   Gate Projection: $x_{\text{gate}} = W_{\text{gate}} \cdot x_{\text{norm\_mlp}} + b_{\text{gate}}$
            *   Up-Projection: $x_{\text{up}} = W_{\text{up}} \cdot x_{\text{norm\_mlp}} + b_{\text{up}}$
            *   Non-linear Activation: $x_{\text{inter}} = \text{Swish}(x_{\text{gate}}) \odot x_{\text{up}}$
            *   Down-Projection: $x_{\text{out}} = W_{\text{down}} \cdot x_{\text{inter}} + b_{\text{down}}$
        *   *Weight & Bias Dimensions*: $W_{\text{gate}}, W_{\text{up}} \in \mathbb{R}^{1024 \times 256}$, $W_{\text{down}} \in \mathbb{R}^{256 \times 1024}$, $b_{\text{gate}}, b_{\text{up}} \in \mathbb{R}^{1024}$, $b_{\text{down}} \in \mathbb{R}^{256}$
        *   *Parameter Calculation*: 
            $$\text{Params} = (2 \times d_{\text{ffn}} \times d_{\text{model}}) + (d_{\text{model}} \times d_{\text{ffn}}) + (2 \times d_{\text{ffn}}) + d_{\text{model}}$$
            $$\text{Params} = (2 \times 1024 \times 256) + (256 \times 1024) + (2 \times 1024) + 256 = 788,736$$
    *   **3.2.3. Residual Addition**
        *   *Tensor Shape*: $[B, T, 256] + [B, T, 256] \rightarrow [B, T, 256]$
        *   *Mathematical Operation*: 
            $$X_{\text{dense\_out}} = X_{\text{mid}} + x_{\text{out}}$$
        *   *Parameter Calculation*: $0$

*   **Total Parameters for 3 Dense Blocks**: 
    $$\text{Params} = 3 \times (\text{Norm1} + \text{GLRU} + \text{Norm2} + \text{SwiGLU})$$
    $$\text{Params} = 3 \times (256 + 197,376 + 256 + 788,736) = 3 \times 986,624 = 2,959,872$$

---

#### **4. Bottleneck Recurrent Block (BlrnBottleneckDecoderLayer) × 12 Layers**
*   This module implements a narrower, parameter-efficient design featuring low-rank linear bottleneck operations in the channel-projection layer.
*   **4.1. Recurrent Sub-block (Temporal Processing)**
    *   **4.1.1. Pre-Normalization (RMSNorm)**: $256$ Parameters
    *   **4.1.2. Gated Linear Recurrent Unit (GLRU)**: $197,376$ Parameters
    *   **4.1.3. Residual Addition**: $0$ Parameters
*   **4.2. Low-Rank Bottleneck MLP Sub-block (Low-Rank Channel Projection)**
    *   **4.2.1. Mid-Normalization (RMSNorm)**: $256$ Parameters
    *   **4.2.2. Low-Rank MLP (Bottleneck Linear)**
        *   *Tensor Shape*: $[B, T, 256] \rightarrow [B, T, 64] \rightarrow [B, T, 256]$ (where $d_{\text{bottleneck}} = 64$)
        *   *Mathematical Operation*: Matrix projection into lower rank subspace followed by activation and up-projection
            *   Down-Rank Projection: $x_{\text{low}} = \text{GELU}(W_{\text{down\_rank}} \cdot x_{\text{norm\_mlp}} + b_{\text{down\_rank}})$
            *   Up-Rank Projection: $x_{\text{high}} = W_{\text{up\_rank}} \cdot x_{\text{low}} + b_{\text{up\_rank}}$
        *   *Weight & Bias Dimensions*: $W_{\text{down\_rank}} \in \mathbb{R}^{64 \times 256}$, $W_{\text{up\_rank}} \in \mathbb{R}^{256 \times 64}$, $b_{\text{down\_rank}} \in \mathbb{R}^{64}$, $b_{\text{up\_rank}} \in \mathbb{R}^{256}$
        *   *Parameter Calculation*: 
            $$\text{Params} = (d_{\text{bottleneck}} \times d_{\text{model}}) + (d_{\text{model}} \times d_{\text{bottleneck}}) + d_{\text{bottleneck}} + d_{\text{model}}$$
            $$\text{Params} = (64 \times 256) + (256 \times 64) + 64 + 256 = 33,088$$
    *   **4.2.3. Residual Addition**: $0$ Parameters

*   **Total Parameters for 12 Bottleneck Blocks**: 
    $$\text{Params} = 12 \times (\text{Norm1} + \text{GLRU} + \text{Norm2} + \text{Low-Rank MLP})$$
    $$\text{Params} = 12 \times (256 + 197,376 + 256 + 33,088) = 12 \times 230,976 = 2,771,712$$

---

#### **5. Final Normalization Layer**
*   **5.1. RMSNorm**
    *   *Tensor Shape*: $[B, T, 256] \rightarrow [B, T, 256]$
    *   *Mathematical Operation*: Amplitude standardization prior to linear projection
        $$x_{\text{final\_norm}} = \text{RMSNorm}(X_{\text{dec\_out}}) \odot \gamma_{\text{final}} \quad (\gamma_{\text{final}} \in \mathbb{R}^{256})$$
    *   *Parameter Calculation*: 
        $$\text{Params} = d_{\text{model}} = 256$$

---

#### **6. Output Prediction Branch (Bifurcated Output Path)**
*   The final normalized hidden state branches into a dual path to stabilize prediction distributions and model offset residuals.
*   **6.1. Linear Projection Branch (Base Estimator)**
    *   **6.1.1. Linear Head**
        *   *Tensor Shape*: $[B, T, 256] \rightarrow [B, T, 256]$ (Vocab Size $V = 256$ representing raw byte channels)
        *   *Mathematical Operation*: Linear output projection
            $$P_t = W_{\text{base}} \cdot x_{\text{final\_norm}} + b_{\text{base}}$$
        *   *Weight & Bias Dimensions*: $W_{\text{base}} \in \mathbb{R}^{256 \times 256}$, $b_{\text{base}} \in \mathbb{R}^{256}$
        *   *Parameter Calculation*: 
            $$\text{Params} = 256 \times d_{\text{model}} + 256 = 256 \times 256 + 256 = 65,792$$
*   **6.2. Gated Residual Offset Branch (Correction Estimator)**
    *   **6.2.1. Gated Offset Projection**
        *   *Tensor Shape*: $[B, T, 256] \rightarrow [B, T, 256]$
        *   *Mathematical Operation*: Element-wise multiplication of raw values modulated by an active SiLU gating factor
            $$O_t = \text{SiLU}(W_{\text{gate}} \cdot x_{\text{final\_norm}} + b_{\text{gate}}) \odot (W_{\text{val}} \cdot x_{\text{final\_norm}} + b_{\text{val}})$$
        *   *Weight & Bias Dimensions*: $W_{\text{gate}}, W_{\text{val}} \in \mathbb{R}^{256 \times 256}$, $b_{\text{gate}}, b_{\text{val}} \in \mathbb{R}^{256}$
        *   *Parameter Calculation*: 
            $$\text{Params} = 2 \times (256 \times d_{\text{model}} + 256) = 2 \times (256 \times 256 + 256) = 131,584$$

---

#### **7. Prediction Aggregation & Output**
*   **7.1. Accumulation Unit (Additive Connection)**
    *   *Tensor Shape*: $[B, T, 256] + [B, T, 256] \rightarrow [B, T, 256]$
    *   *Mathematical Operation*: Recombines baseline estimators with residual prediction coordinates
        $$Y_{\text{logits}} = P_t + O_t$$
    *   *Parameter Calculation*: $0$
*   **7.2. Probability Projection**
    *   *Tensor Shape*: $[B, T, 256] \rightarrow [B, T, 256]$
    *   *Mathematical Operation*: Softmax calculation for the categorical distribution over 256 UTF-8 byte outputs
        $$y_{\text{prob}} = \text{Softmax}(Y_{\text{logits}})$$
    *   *Parameter Calculation*: $0$

---

### **Summed Architecture Parameter Breakdown**

*   **Embedding Layer**: $65,536$
*   **Dense Recurrent Block (3 layers)**: $2,959,872$
*   **Bottleneck Recurrent Block (12 layers)**: $2,771,712$
*   **Final Normalization Layer**: $256$
*   **Linear Projection Branch**: $65,792$
*   **Gated Residual Offset Branch**: $131,584$

$$\text{Grand Total Parameter Count} = 5,994,752 \quad (\approx 5.99\text{ Million Parameters})$$