TRM-text / README.md
summerMC's picture
Update README.md
de44a45 verified
|
Raw
History Blame Contribute Delete
4.22 kB
# TRM-text
TRM-text is an attention-free language model based on a Tiny Recursive Model (TRM) architecture.
Unlike Transformer-based language models, TRM-text removes self-attention entirely and replaces it with recursive computation built from causal dilated depthwise convolutions, hierarchical latent refinement, and parameter reuse.
The goal of TRM-text is to investigate whether recursive neural computation can provide competitive language modeling performance while dramatically reducing computational cost.
---
# Overview
TRM-text explores a different scaling path from Transformers.
Instead of increasing attention heads and context interactions, TRM-text repeatedly refines hidden representations using a hierarchy of recursive processing blocks.
Key properties:
* Attention-free
* Autoregressive language modeling
* Recursive computation
* Hierarchical latent refinement
* RoPE positional encoding
* Dilated depthwise convolution mixer
* Hugging Face compatible
* safetensors support
---
# Architecture
```text
Tokens
Embedding
Low-Level TRM
Mid-Level TRM
High-Level TRM
Recursive Feedback
LM Head
```
Each recursive block contains:
* RMSNorm
* RoPE
* Causal Dilated Depthwise Convolution
* SwiGLU Feed Forward Network
* Residual Recurrence
No self-attention layers are used.
---
# Model Configuration
Current release:
```text
Parameters: ~15M
dim = 256
hidden_dim = 512
low_steps = 6
mid_steps = 3
high_steps = 2
cycles = 3
kernel_size = 5
low_dilations = [1,2,4,8]
mid_dilations = [2,4,8,16]
high_dilations = [4,8,16,32]
```
---
# Compute Efficiency
Relative training cost:
| Architecture | Relative Cost |
| ------------ | ------------: |
| Transformer | 1200 |
| HRM | 100 |
| TRM-text | 1 |
These values represent relative compute requirements under the experimental scaling assumptions used during development.
The objective of TRM-text is to maximize efficiency through:
* parameter reuse
* recursive computation
* hierarchical refinement
* elimination of attention operations
---
# Training
## Base Pretraining
Dataset:
```text
FineWeb Sample-10BT
```
Tokenizer:
```text
GPT-2 BPE
```
Objective:
```text
Causal Language Modeling
```
---
# Instruction Tuning
Dataset:
```text
tatsu-lab/alpaca
```
Format:
```text
### Instruction:
...
### Response:
...
```
---
# Loading
```python
from transformers import AutoTokenizer
from transformers import AutoModelForCausalLM
tokenizer = AutoTokenizer.from_pretrained(
"summerMC/TRM-text",
trust_remote_code=True
)
model = AutoModelForCausalLM.from_pretrained(
"summerMC/TRM-text",
trust_remote_code=True
)
```
---
# Inference
```python
prompt = """
### Instruction:
Explain artificial intelligence in simple terms.
### Response:
"""
inputs = tokenizer(
prompt,
return_tensors="pt"
)
outputs = model.generate(
**inputs,
max_new_tokens=128,
do_sample=True,
temperature=0.7,
top_k=40
)
print(
tokenizer.decode(
outputs[0],
skip_special_tokens=True
)
)
```
---
# Research Motivation
TRM-text investigates whether recursive neural systems can replace attention mechanisms in language modeling.
Research directions:
* recursive reasoning
* hierarchical computation
* efficient language models
* attention-free architectures
* low-cost scaling laws
---
# Limitations
Current checkpoint is experimental.
Known limitations:
* small parameter count
* limited instruction tuning
* lower capability than modern frontier models
* research-focused implementation
* benchmark coverage still limited
---
# Intended Use
TRM-text is intended for:
* language model research
* efficient architecture experimentation
* recursive computation studies
* attention-free modeling research
Not intended for:
* safety-critical systems
* medical decision making
* legal advice
* financial advice
---
# License
Apache-2.0
---
# Citation
```bibtex
@software{trm_text_2026,
title={TRM-text: Attention-Free Recursive Language Modeling},
author={summerMC},
year={2026},
url={https://huggingface.co/summerMC/TRM-text}
}
```