File size: 5,618 Bytes
67fddde
 
013f15a
 
 
 
 
 
 
 
 
 
67fddde
013f15a
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
af31e5b
 
013f15a
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
af31e5b
 
 
 
 
013f15a
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
---
license: mit
language:
- en
pipeline_tag: text-generation
tags:
- conversational
- dialogue
- latent-space
- thought-vectors
- research
- pytorch
---

# thoughtvec β€” conversation in thought-vector space

A ~48M-parameter dialogue system that never does token-level language
modeling in its reasoning loop. Trained from scratch in ~25 GPU-hours on one
consumer GPU (AMD RX 6700 XT, 12 GB) β€” a 12-hour codec run and a 12-hour
thinker run. (That's the cost of these weights, not of the research: the
ablation record behind them β€” 86 checkpointed runs β€” took on the order of
another hundred GPU-hours.)

It works in two parts:

- A **codec** (32.9M params) turns text into a sequence of "thought vectors"
  ordered by importance β€” any prefix of them decodes back to text, so
  compression ratio is a decode-time choice. Byte-perfect at 4:1, readable
  at 8:1.
- A **thinker** (15.1M params) converses directly in that latent space:
  history turns in, response thoughts out, and the frozen codec decoder
  renders the reply. No tokens anywhere in between.

```
user > i'm feeling really overwhelmed with work lately.
bot  > I can imagine. What's been going on?
user > my boss keeps piling on deadlines.
bot  > That sounds like a lot of pressure. Have you been able to talk to anyone else?
```

That transcript is greedy decoding (temperature 0), reproducible from
`FINAL_12H-best.pt`.

**Try it in the browser**: [thought-vectors-chat](https://huggingface.co/spaces/nochinator/thought-vectors-chat)
β€” a model picker lets you compare live against the paper's matched
token-LM baseline (Β§6.5) on the same conversation.

## What to expect β€” and not

This is a research model, and a small one. It holds coherent, grounded,
context-sensitive small talk: openings, follow-up questions, multi-turn
reference. That is the finding β€” conversational competence at this level
does not require a token-level LM in the loop.

It is *not* a general assistant. No knowledge tasks, English small talk
only. And it has one well-documented disease: it can reply cheerfully to
bad news, especially after an upbeat turn ("That's good to hear." to
insomnia). The paper traces that failure through three ablation rounds to
a training-data absence β€” no conversation in any corpus used ever reverses
mood mid-dialogue β€” and shows that even targeted data patching bought only
partial, topic-gated routing rather than the skill (the patched case-study
model commiserates with illness or job loss, and still cheers a burglary
or a hailstorm). If you probe the model, you will find it; it's a case
study, not a surprise.

## Files

| File | What it is |
|---|---|
| `FINAL_12H-best.pt` | Thinker flagship (use this one to chat) |
| `FINAL2_12H-best.pt` | Register case-study run (+EmpatheticDialogues +reversal splices) |
| `m5_frontier-best.pt` | Codec β€” needed alongside either thinker |
| `spm16k_bpe.model` | 16K SentencePiece tokenizer β€” all checkpoints depend on it |

SHA-256 (byte-identical to the [GitHub Release](https://github.com/nochinator/thought-vectors/releases/tag/v1.1.0) assets):

```
4685ab3dae35e7835ee0fdd469d23d39347b0cca0f71c0598bea30443f0d7672  FINAL_12H-best.pt
bc565782b51d74fe56de1c2d9dd4e44c52b1e6ff192bfcdb758291edeaf5240c  FINAL2_12H-best.pt
77bf080cc216f628fa80c260acc2cfe2ffb03c19bc3d44071047ad4377b0e4ad  m5_frontier-best.pt
```

The paper's matched token-LM baseline (`b3_lm_48m_24h-best.pt`, Β§6.5) is
the comparison model, not part of this system, so the raw checkpoint ships
as a GitHub Release asset only rather than in this repo. An ONNX export of
it is available for live comparison in the [chat demo](https://huggingface.co/spaces/nochinator/thought-vectors-chat)'s
model picker.

## Run it locally

CPU is enough β€” chat inference runs on CPU by design.

```bash
git clone https://github.com/nochinator/thought-vectors
cd thought-vectors
scripts/setup_env.sh --cpu

# place the files (thinker embeds codec weights but reads codec config
# from the codec file β€” you need both):
#   FINAL_12H-best.pt   -> checkpoints/FINAL_12H/best.pt
#   m5_frontier-best.pt -> checkpoints/m5_frontier/best.pt

.venv/bin/tv-chat --ckpt checkpoints/FINAL_12H/best.pt --device cpu
```

## Results at a glance

| Run | val_cos ↑ | ref_F1 ↑ | self_rep ↓ | ctx_sens ↓ |
|---|---|---|---|---|
| FINAL_12H (flagship) | **0.428** | **0.297** | 0.188 | **0.146** |
| FINAL2_12H (case study) | 0.415 | 0.278 | **0.154** | 0.206 |

Codec: byte-perfect reconstruction at 4:1 compression through 257 tokens;
graceful degradation at 8:1 and beyond. Every number maps to a named log
file in the repo β€” the full experimental record, including the negative
rounds and four incidents of models gaming lexical metrics, is in
[RESEARCH_LOG.md](https://github.com/nochinator/thought-vectors/blob/main/RESEARCH_LOG.md).

## Links

- **Paper**: [DOI 10.5281/zenodo.21262842](https://doi.org/10.5281/zenodo.21262842) β€” also as
  [PDF in the repo](https://github.com/nochinator/thought-vectors/blob/main/paper/main.pdf)
- **Code + logs**: https://github.com/nochinator/thought-vectors
- **Reproduce from scratch** (~25 GPU-hours, one RX 6700 XT):
  [docs/REPRODUCE.md](https://github.com/nochinator/thought-vectors/blob/main/docs/REPRODUCE.md)

## Citation

```bibtex
@misc{nochi2026thoughtvectors,
  author = {nochi},
  title  = {Codecs: Separating Meaning from Language --- Coherent
            Dialogue in Thought-Vector Space from 48M Parameters on One
            Consumer GPU},
  year   = {2026},
  doi    = {10.5281/zenodo.21262842},
  url    = {https://github.com/nochinator/thought-vectors}
}
```