File size: 4,223 Bytes
de44a45 46c3d88 de44a45 46c3d88 de44a45 46c3d88 de44a45 46c3d88 de44a45 46c3d88 de44a45 46c3d88 de44a45 46c3d88 de44a45 46c3d88 de44a45 46c3d88 de44a45 46c3d88 de44a45 46c3d88 de44a45 46c3d88 de44a45 46c3d88 de44a45 46c3d88 de44a45 46c3d88 de44a45 46c3d88 de44a45 46c3d88 de44a45 46c3d88 de44a45 46c3d88 de44a45 46c3d88 de44a45 46c3d88 de44a45 46c3d88 de44a45 46c3d88 de44a45 46c3d88 de44a45 46c3d88 de44a45 46c3d88 de44a45 46c3d88 de44a45 46c3d88 de44a45 46c3d88 de44a45 46c3d88 de44a45 46c3d88 de44a45 46c3d88 de44a45 46c3d88 de44a45 46c3d88 de44a45 46c3d88 de44a45 46c3d88 de44a45 46c3d88 de44a45 46c3d88 de44a45 46c3d88 de44a45 46c3d88 de44a45 46c3d88 de44a45 46c3d88 de44a45 46c3d88 de44a45 46c3d88 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 | # TRM-text
TRM-text is an attention-free language model based on a Tiny Recursive Model (TRM) architecture.
Unlike Transformer-based language models, TRM-text removes self-attention entirely and replaces it with recursive computation built from causal dilated depthwise convolutions, hierarchical latent refinement, and parameter reuse.
The goal of TRM-text is to investigate whether recursive neural computation can provide competitive language modeling performance while dramatically reducing computational cost.
---
# Overview
TRM-text explores a different scaling path from Transformers.
Instead of increasing attention heads and context interactions, TRM-text repeatedly refines hidden representations using a hierarchy of recursive processing blocks.
Key properties:
* Attention-free
* Autoregressive language modeling
* Recursive computation
* Hierarchical latent refinement
* RoPE positional encoding
* Dilated depthwise convolution mixer
* Hugging Face compatible
* safetensors support
---
# Architecture
```text
Tokens
│
▼
Embedding
│
▼
Low-Level TRM
│
▼
Mid-Level TRM
│
▼
High-Level TRM
│
▼
Recursive Feedback
│
▼
LM Head
```
Each recursive block contains:
* RMSNorm
* RoPE
* Causal Dilated Depthwise Convolution
* SwiGLU Feed Forward Network
* Residual Recurrence
No self-attention layers are used.
---
# Model Configuration
Current release:
```text
Parameters: ~15M
dim = 256
hidden_dim = 512
low_steps = 6
mid_steps = 3
high_steps = 2
cycles = 3
kernel_size = 5
low_dilations = [1,2,4,8]
mid_dilations = [2,4,8,16]
high_dilations = [4,8,16,32]
```
---
# Compute Efficiency
Relative training cost:
| Architecture | Relative Cost |
| ------------ | ------------: |
| Transformer | 1200 |
| HRM | 100 |
| TRM-text | 1 |
These values represent relative compute requirements under the experimental scaling assumptions used during development.
The objective of TRM-text is to maximize efficiency through:
* parameter reuse
* recursive computation
* hierarchical refinement
* elimination of attention operations
---
# Training
## Base Pretraining
Dataset:
```text
FineWeb Sample-10BT
```
Tokenizer:
```text
GPT-2 BPE
```
Objective:
```text
Causal Language Modeling
```
---
# Instruction Tuning
Dataset:
```text
tatsu-lab/alpaca
```
Format:
```text
### Instruction:
...
### Response:
...
```
---
# Loading
```python
from transformers import AutoTokenizer
from transformers import AutoModelForCausalLM
tokenizer = AutoTokenizer.from_pretrained(
"summerMC/TRM-text",
trust_remote_code=True
)
model = AutoModelForCausalLM.from_pretrained(
"summerMC/TRM-text",
trust_remote_code=True
)
```
---
# Inference
```python
prompt = """
### Instruction:
Explain artificial intelligence in simple terms.
### Response:
"""
inputs = tokenizer(
prompt,
return_tensors="pt"
)
outputs = model.generate(
**inputs,
max_new_tokens=128,
do_sample=True,
temperature=0.7,
top_k=40
)
print(
tokenizer.decode(
outputs[0],
skip_special_tokens=True
)
)
```
---
# Research Motivation
TRM-text investigates whether recursive neural systems can replace attention mechanisms in language modeling.
Research directions:
* recursive reasoning
* hierarchical computation
* efficient language models
* attention-free architectures
* low-cost scaling laws
---
# Limitations
Current checkpoint is experimental.
Known limitations:
* small parameter count
* limited instruction tuning
* lower capability than modern frontier models
* research-focused implementation
* benchmark coverage still limited
---
# Intended Use
TRM-text is intended for:
* language model research
* efficient architecture experimentation
* recursive computation studies
* attention-free modeling research
Not intended for:
* safety-critical systems
* medical decision making
* legal advice
* financial advice
---
# License
Apache-2.0
---
# Citation
```bibtex
@software{trm_text_2026,
title={TRM-text: Attention-Free Recursive Language Modeling},
author={summerMC},
year={2026},
url={https://huggingface.co/summerMC/TRM-text}
}
```
|