File size: 6,227 Bytes
2fd8134
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9533e3b
2fd8134
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9533e3b
 
 
 
 
 
 
 
 
 
 
 
 
2fd8134
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
---
library_name: transformers
license: apache-2.0
language:
  - eu
  - en
  - es
tags:
  - mamba-2
  - basque
  - autocomplete
  - fim
  - fill-in-the-middle
  - ghost-text
  - continued-pretraining
  - low-resource
base_model: itzune/morpheus
pipeline_tag: text-generation
---

# Morpheus-FIM (Mamba-2) β€” Basque Fill-in-the-Middle Autocomplete

**Continued pre-training** of [itzune/morpheus](https://huggingface.co/itzune/morpheus) (the 91M AR-only Mamba-2 model, step 74K) for **Fill-in-the-Middle (FIM)** completion β€” the model can predict text at the cursor, not just at the end of a buffer.

This is the base model for the GGUF at [itzune/morpheus-gguf](https://huggingface.co/itzune/morpheus-gguf) (`v3_fim.Q5_K_M.gguf`).

## Model Details

- **Architecture:** Mamba-2 (State Space Model), 24 layers, d_model 768
- **Parameters:** 91M
- **Base model:** [itzune/morpheus](https://huggingface.co/itzune/morpheus) (step 74K, AR-only, PPL 7.13)
- **Checkpoint:** Phase 6 v3 `best.pt` (step 3500 of CPT)
- **AR perplexity:** 7.5 | **FIM perplexity:** 7.9
- **Vocabulary:** 4,016 (4,000 original Unigram pieces + `<PRE>` `<SUF>` `<MID>` `<EOT>` + 12 padding rows)
- **Trained without BOS token** (`add_bos_token=false`)

## Continued Pre-Training Recipe

| Parameter | Value |
|-----------|-------|
| Base checkpoint | step 74K AR (`itzune/morpheus`), embeddings resized 4000 β†’ 4016 |
| Token budget | 500M tokens |
| FIM/AR ratio | **70/30** |
| `<EOT>` loss weight | **5Γ—** (per-class cross-entropy weight on token id 4003) |
| Splitting | Token-level (BigCode/StarCoder), 20% at linguistic boundaries |
| Loss masking | None ("FIM-for-free" β€” loss on all tokens) |
| Packing | Greedy whole-example packing into 1025-token windows |
| Learning rate | 1.0e-3, cosine decay, ~3,815 steps |
| Tokenizer | `basque_unigram_fim.model` (original 4000 + 4 FIM tokens) |

The 70/30 ratio + 5Γ— `<EOT>` weight directly target the FIM stop-token reliability problem: the model must learn not just *what* to generate but *when to stop*, and the `<EOT>` signal is otherwise too sparse (one token per example) within a 500M-token budget for the model to reliably emit it.

## FIM Token Format

Code Llama-style FIM tokens (Bavarian et al., 2022; Roziere et al., 2023):

| Token | ID | Purpose |
|-------|----|---------|
| `<PRE>` | 4000 | Marks start of prefix |
| `<SUF>` | 4001 | Marks start of suffix |
| `<MID>` | 4002 | Marks start of generation (infill) |
| `<EOT>` | 4003 | End-of-infill (stop token) |
| 4004–4015 | β€” | Padding (kernel alignment, unused) |

To do a fill-in-the-middle completion, structure the prompt as:

```
<PRE>{prefix}<SUF>{suffix}<MID>
```

The model generates the infill and emits `<EOT>` when done.

## Evaluation Results

FIM eval on 147 held-out examples (token-level splits, 20% at linguistic boundaries):

| Metric | Result |
|--------|--------|
| `<EOT>` emission rate | 88.4% |
| Keystrokes saved | βˆ’5.9% |
| Exact-match rate | 6.8% |
| Avg char accuracy | 32.3% |
| Avg generation length (ref=45.0) | 40.3 |
| Prefix truncation (overall) | 1.4% |
| β”” long-bucket truncation | 2.3% (< 15% threshold) |
| AR valid PPL | 7.5 |
| FIM valid PPL | 7.9 |

The 5Γ— `<EOT>` loss weighting resolves the over-generation failure mode: `<EOT>` emission reaches 88.4% and generation length (40.3) sits near the 45.0-char reference, yielding near-break-even keystrokes saved (βˆ’5.9%). The feared premature-truncation failure mode β€” the dual risk of over-weighting the stop token β€” did not materialize (1.4% overall, 2.3% long-bucket, far below the 15% threshold). AR perplexity remained stable (7.5 vs. 7.13 AR-only base), confirming the 70/30 FIM ratio did not trade away AR capability.

## Usage

### With the Morpheus demo server (recommended)

The [Morpheus demo](https://github.com/itzune/morpheus/tree/main/demo) includes a FastAPI proxy that handles FIM templating, token-ID encoding, and an OpenAI-compatible API:

```bash
cd demo
MORPHEUS_MODEL=v3_fim.Q5_K_M.gguf docker compose -f docker-compose.yml -f docker-compose.local.yml up -d --build
# Open http://localhost:9090/editor.html
```

### Direct with transformers

```python
from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained("itzune/morpheus-fim")
tokenizer = AutoTokenizer.from_pretrained("itzune/morpheus-fim")

# FIM prompt
prompt = "<PRE>Kaixo, <SUF> moduz?<MID>"
inputs = tokenizer(prompt, return_tensors="pt")
output = model.generate(**inputs, max_new_tokens=20, temperature=0.2, top_k=5)
print(tokenizer.decode(output[0], skip_special_tokens=False))
```

### ⚠️ Critical: token-ID prompts for llama.cpp

When deploying via `llama.cpp`/`llama-server`, encode prompts with the `sentencepiece` library using `tokenizer.model` and send **token IDs** (not strings) to the `/completion` endpoint. This avoids the BOS auto-prepend and tokenizer-divergence issues documented on the [base model card](https://huggingface.co/itzune/morpheus). The demo proxy handles this automatically.

## Decoding Parameters (recommended)

| Parameter | Value | Rationale |
|-----------|-------|-----------|
| `temperature` | 0.2 | Low-but-nonzero: recovers rank-2 correct tokens greedy misses |
| `top_k` | 5 | Small nucleus; 5 correct answers sit at rank 2 in top-5 |
| `repeat_penalty` (FIM) | 1.0 | FIM legitimately reuses context words |
| `stop` (FIM) | `["<EOT>", "\n\n"]` | Model-emitted stop + paragraph-boundary fallback |

## Intended Use

Desktop text-editor **ghost-text autocompletion** for Basque prose. The Mamba-2 architecture's O(1) decode cost makes it well-suited to long editing sessions where per-token latency matters more than parallelism.

**Not intended for:** instruction following, chat, translation, or factual QA. This is a narrow autocomplete model.

## License

Apache-2.0.

## References

- Bavarian, M., et al. (2022). *Efficient Training of Language Models to Fill in the Middle*. arXiv:2207.14255.
- Roziere, B., et al. (2023). *Code Llama: Open Foundation Models for Code*. arXiv:2308.12950.
- Dao, T., & Gu, A. (2024). *Transformers are SSMs: Generalized Models and Efficient Algorithms through Structured State Space Duality*. arXiv:2405.21060.