File size: 10,409 Bytes
2510087
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
---

license: apache-2.0
datasets:
- HuggingFaceFW/fineweb-edu
- eliceai/korean-fineweb-edu-demo
language:
- en
- ko
base_model:
- RetentionLabs/TTT-Linear-1.3B-Base-Pile-8k
- Qwen/Qwen3-4B-Thinking-2507
base_model_relation: merge
pipeline_tag: text-generation
library_name: transformers
tags:
- Test-time Training
- Memory-augemted Transformer
---


# TTTPilot-Q-5B-Thinking-MAC

**A Pilot Implementation of Titans MAC Architecture**

Combining [TTT-Linear-1.3B-Base-Pile-8k](https://huggingface.co/test-time-training/TTT-Linear-1.3B-Base-Pile-8k) and [Qwen3-4B-Thinking-2507](https://huggingface.co/Qwen/Qwen3-4B-Thinking-2507) using the **Memory as Context (MAC)** architecture pattern.

---

## 🎯 Pilot Experiment Overview

This is an **experimental pilot** exploring how test-time training (TTT) memory layers can be combined with standard transformer cores in a modular architecture. The MAC pattern separates:

- **Memory Layers**: TTT-Linear's self-adaptation mechanism for dynamic context learning
- **Core Layers**: Qwen3's transformer decoder for reasoning and generation

### Key Idea: MAC Processing Flow

```

Input Sequence (Processed in Segments)

    ↓

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”

β”‚  1. Read-Only Retrieval (R-Mode)   β”‚  ← Memory layers (Q-only projection)

β”‚     Generate memory queries          β”‚    Enables parallel computation

β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

    ↓

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”

β”‚  2. Core Processing                 β”‚  ← Qwen3 transformer layers

β”‚     [Fixed Memory + Query + Input]  β”‚    Standard attention + MLP

β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

    ↓

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”

β”‚  3. Memory Update (W-Mode)          β”‚  ← Memory layers (full QKV)

β”‚     Update context representations  β”‚    Test-time adaptation

β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

    ↓

Final Output (Join updated context with core output)

```

### Architecture Benefits

1. **Parallel Memory Retrieval**: Q-only projection in R-mode enables efficient segment processing
2. **Weight Tying**: `retriever.q` shares weights with `memory.q` for efficiency
3. **Modular Design**: Memory and core can be independently scaled/fine-tuned
4. **Hybrid Capabilities**: Combines TTT's adaptive learning with transformer's proven performance

---

## πŸ“Š Model Statistics

| Component | Source Model | Layers | Parameters | Intermediate Size |
|-----------|-------------|--------|------------|-------------------|
| **Embedding & LM Head** | Qwen3-4B-Thinking | - | ~310M (vocab: 151,936) | - |
| **Memory Layers** | TTT-Linear-1.3B | 24 | ~1.3B | 5,504 |
| **Core Layers** | Qwen3-4B-Thinking | 36 | ~4B | 9,728 |
| **Total** | Combined MAC | **60 layers** | **~5.6B** | Mixed |

**Hidden Size**: 2048 (from TTT-Linear)  
**Attention Heads**: 32 (memory), 32/8 GQA (core)  
**Context Length**: Up to 262K tokens (Qwen3's max)  
**Precision**: BFloat16

---

## πŸ—οΈ Architecture Details


### Memory Module (TTT-Linear)

- **Purpose**: Dynamic context adaptation through test-time training
- **Key Components**:
  - Self-adaptation layers with learnable neural memory
  - Momentum-based learning rate gates
  - Q/K/V projections with RoPE
  - Mini-batch processing (chunk_size=16)

- **Special Features**:

  - Weight tying between retriever and full memory

  - Shared Q/K projections with separate conv layers

  - Learnable token-wise learning rates



### Core Module (Qwen3)



- **Purpose**: High-capacity reasoning and generation

- **Key Components**:

  - Multi-head attention with Grouped Query Attention (GQA)

  - SwiGLU MLP activations

  - RMSNorm for layer normalization

  - RoPE with theta=5,000,000 for long context

- **Special Features**:

  - 8 KV heads for efficient inference

  - No attention bias

  - Sliding window attention support



### Fixed Persistent Memory



- **Size**: 64 tokens Γ— 2048 dimensions

- **Purpose**: Store global context/knowledge across segments

- **Initialization**: Zeros (trainable parameter)



---



## πŸš€ Usage



### Installation



```bash

# Install dependencies

pip install torch transformers safetensors accelerate

```



### Quick Start



```python

from transformers import AutoModelForCausalLM, AutoTokenizer



# Load model

model = AutoModelForCausalLM.from_pretrained(
    "./TTTPilot-Q-5B-Thinking-MAC",

    trust_remote_code=True,

    torch_dtype="bfloat16",

    device_map="auto"

)


tokenizer = AutoTokenizer.from_pretrained("./TTTPilot-Q-5B-Thinking-MAC")



# Generate text

prompt = "The future of AI is"

inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=100, do_sample=True, temperature=0.7)



print(tokenizer.decode(outputs[0], skip_special_tokens=True))

```



### Advanced: Segment Processing



```python

# MAC processes inputs in segments for memory efficiency

# Segment size controlled by mini_batch_size (default: 16)



# For long sequences, the model automatically:

# 1. Chunks input into mini-batches

# 2. Processes each chunk through R-mode β†’ Core β†’ W-mode

# 3. Accumulates context updates across chunks



long_text = "..." * 1000  # Very long input

inputs = tokenizer(long_text, return_tensors="pt", truncation=False)

outputs = model.generate(**inputs, max_new_tokens=200)
```



---



## πŸ”§ Weight Conversion



To recreate this model from source checkpoints:



```bash

python convert_weights.py

```

The script:
1. Loads TTT-Linear-1.3B and Qwen3-4B-Thinking
2. Maps TTT weights β†’ memory layers (preserving exact key names)
3. Maps Qwen3 weights β†’ core layers (preserving exact key names)
4. Uses Qwen3's embedding & lm_head (for vocab compatibility)

5. Copies tokenizer files from Qwen3

6. Saves combined model in HuggingFace format



---



## βš™οΈ Configuration



Key hyperparameters in `config.json`:



```json

{

  "model_type": "tttpilot_mac",

  "vocab_size": 151936,          // Qwen3
  "hidden_size": 2048,            // TTT-Linear

  "num_memory_layers": 24,        // TTT-Linear

  "num_core_layers": 36,          // Qwen3

  "memory_intermediate_size": 5504,   // TTT MLP

  "core_intermediate_size": 9728,     // Qwen3 MLP

  "num_attention_heads": 32,

  "num_key_value_heads": 8,       // GQA in cores
  "mini_batch_size": 16,          // TTT chunk size
  "ttt_base_lr": 1.0,             // TTT learning rate
  "fixed_memory_size": 64,        // Persistent memory tokens
  "rope_theta": 5000000,          // Long context RoPE

  "max_position_embeddings": 262144   // Max sequence length

}

```



---



## πŸ§ͺ Pilot Experiment Status



### What Works

βœ… Model architecture defined  

βœ… Weight conversion pipeline  

βœ… Configuration files generated  

βœ… Tokenizer compatibility (Qwen3)  

βœ… Basic forward pass structure  



### What's Experimental

⚠️ **MAC segment processing**: Simplified in pilot, needs full TTT integration  

⚠️ **Retriever implementation**: Placeholder, requires Q-only inference mode  

⚠️ **Weight tying**: Defined but not verified in practice  

⚠️ **Memory update logic**: Full TTT adaptation step needs integration  



### Known Limitations

- Memory layers use placeholder identity functions (need full TTT code)

- No actual segment-based MAC flow yet (processes like standard transformer)

- TTT cache and Qwen3 KV cache not properly integrated

- No training/fine-tuning tested

- Generation quality not benchmarked



---



## πŸŽ“ Research Context



This pilot implements concepts from:



1. **Test-Time Training (TTT)**: Self-supervised adaptation during inference

   - Paper: [Learning to (Learn at Test Time)](https://arxiv.org/abs/2407.04620)

   - Code: [TTT-Linear](https://github.com/test-time-training/ttt-lm-pytorch)



2. **Titans Architecture**: Modular memory-augmented patterns

   - Inspiration: Memory-as-X design patterns (MAC, MAE, MAL, etc.)

   - Idea: Separate stateful memory from stateless reasoning



3. **Grouped Query Attention**: Efficient multi-head attention

   - From: Qwen3 and modern LLMs

   - Benefit: Faster inference with minimal quality loss



---



## πŸ“ Citation



If you use this pilot or build upon it:



```bibtex

@misc{tttpilot-mac-2026,

  title={TTTPilot-MAC: A Pilot Implementation of Memory-Augmented-Core Architecture},

  author={Your Name},

  year={2026},

  note={Pilot experiment combining TTT-Linear and Qwen3},

  howpublished={\url{https://github.com/...}}

}

```



**Source Models**:

- TTT-Linear: [test-time-training/TTT-Linear-1.3B-Base-Pile-8k](https://huggingface.co/test-time-training/TTT-Linear-1.3B-Base-Pile-8k)

- Qwen3: [Qwen/Qwen3-4B-Thinking-2507](https://huggingface.co/Qwen/Qwen3-4B-Thinking-2507)



---



## πŸ“„ License



This project combines:

- **TTT-Linear** (MIT License)

- **Qwen3** (Apache 2.0 License)



Final license: **Apache 2.0** (compatible with both)



See [LICENSE](./LICENSE) for details.



---



## 🀝 Contributing



This is a **pilot experiment** for research exploration. Contributions welcome:



1. Full MAC segment processing implementation

2. TTT-Linear integration (replace placeholders)

3. Retriever Q-only mode

4. Training scripts

5. Benchmark evaluations

6. Documentation improvements



---



## πŸ› Issues & Feedback



Found a bug or have suggestions? Open an issue!



**Important Notes**:

- This is NOT a production-ready model

- Use for research/experimentation only

- Performance not guaranteed

- May require significant compute resources (5.6B parameters)



---



## 🌟 Acknowledgments



- **TTT Team** for the test-time training paradigm

- **Qwen Team** for Qwen3-Thinking model

- **Titans Architecture** inspiration from modular design patterns



---



**Status**: 🚧 Experimental Pilot - Use with caution!