File size: 15,947 Bytes
b14550d
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9c5f8cc
 
 
 
 
 
 
 
 
b14550d
 
 
 
 
 
 
 
 
 
 
 
9c5f8cc
b14550d
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
4eee74c
b14550d
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
---
license: other
library_name: transformers
tags:
- code-translation
- r
- python
- lora
- unsloth
- sas
- statistical
- programming
---

# Euclid-2.0

Bidirectional SAS ↔ R ↔ Python program translation. A 30B-A3B hybrid Mamba-2 / MoE / attention model, LoRA-tuned and merged to 16-bit weights.

## Model Summary

Euclid-2.0 translates complete statistical and data-processing programs across all six directed pairs over {SAS, Python, R}. The training objective is **behavioural equivalence**: given identical inputs, the translated program must compute identical values and place them in results carrying identical names. This criterion is encoded in the system prompt, in the dataset construction, and in the evaluation design — execution parity over CodeBLEU or exact match.

| | |
|---|---|
| Developer | ProCogia |
| Method | LoRA SFT (r=64), completion-only supervision, adapter merged to 16-bit |
| Parameters | ~30B total / ~3B active per token (A3B sparse activation) |
| Layers | 52 total — ~23 Mamba-2, ~23 MoE, ~6 attention (GQA) |
| Experts | 128 routed plus a shared expert; ~6 routed experts active per token |
| Expert layout | Ungated `up_proj` / `down_proj` (not the fused `gate_up_proj` layout) |
| Chat template | ChatML-style `<\|im_start\|>role\n … <\|im_end\|>` |
| Context | Up to 1M architectural; project cap and trained length 8,192 |
| Precision | BF16, ~63 GB merged |
| License | See `LICENSE` in this repository |

**Deployment mode is reasoning-OFF.** This is a deterministic translation engine; chain-of-thought is neither required nor wanted at inference. The training data contains no reasoning traces, and every row was rendered with `enable_thinking=False`, which injects an empty `<think></think>`. Because that empty block sits *inside* the masking marker, it is excluded from the loss — so the model learned "reasoning-off control token → answer directly" rather than "thinking is forbidden."

Reasoning-ON capability was explicitly sacrificed. It is presumed degraded and is unmeasured. Route reasoning-ON traffic to a general-purpose model.

## Intended Use

- Single-turn translation of complete programs across the six directed pairs over {SAS, Python, R}.
- Deterministic serving under the verbatim serving contract below: greedy decoding, `enable_thinking=False`, exact system prompt and user format.
- Output contract: the translated program only, no prose, no markdown fences, no placeholders, result names preserved exactly.

## How to Use

### Loader

```python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained(
    "ProCogia/Euclid-2.0",
    torch_dtype=torch.bfloat16,
    device_map={"": 0},
    trust_remote_code=True,
)
tokenizer = AutoTokenizer.from_pretrained("ProCogia/Euclid-2.0")
```

Requires `transformers >= 5.5, < 5.6` (5.5.4 proven); the MoE implementation is unavailable below 5.

`mamba_ssm` (2.3.2.post1) and `causal_conv1d` (1.6.2.post1) should be installed. The fused kernels give a 5–10× speedup on 23 of 52 layers. Build them once and cache the wheels: a rebuild takes roughly 25 minutes, a cached reinstall takes seconds. Record the torch version the wheels were built against — stale ABI wheels are poison, and a torch change must force a rebuild.

The torch `+cu128` build is load-bearing. The container toolkit is CUDA 12.8, and torch refuses to compile extensions across a CUDA major-version gap, so the default PyPI `+cu130` wheel cannot build the Mamba kernels here.

### vLLM

The merged 16-bit artifact is serving-ready for vLLM and SGLang.

```bash
vllm serve ProCogia/Euclid-2.0 --max-model-len 8192
```

8,192 is the trained regime. The architecture supports far more, but quality beyond the project cap is unvalidated.

### Serving contract

Mandatory. A fine-tune this narrow answers correctly **only** under its training conditioning:

| Setting | Value |
|---|---|
| System prompt | Byte-identical to the training prompt |
| User message | `Translate the following {SRC} program to {TGT}.` + the per-target constraint paragraph + the source fenced with the correct fence tag |
| Chat template | `enable_thinking=False` |
| Decoding | Greedy |
| Output | Bare program — no fences, no prose, no placeholders |

Extract the constraint paragraphs from the training data at runtime rather than transcribing them, so the serving layer reproduces them exactly.

### Prompt format

A single fixed template was used by design. For a deterministic translation engine, one conditioning format maximizes reliability; prompt diversity would trade that for robustness the product does not need. The consequence is the hard serving requirement above.

The system prompt is a compact SAS-semantics specification, byte-identical across all rows. It declares behavioural equivalence as the only criterion, requires structure, step-order and comment preservation with exact result naming, and legislates four SAS trap semantics:

```text
You are a code translation engine for statistical and data-processing programs written in SAS, Python, and R.

You are given one complete program in a source language and produce the equivalent program in the target language. Behavioural equivalence is the only criterion: given the same inputs, your program must compute the same values and place them in results carrying the same names.

- Preserve the source program's structure, step order, and intent. Carry its comments across as comments in the target language, and add a brief comment where the target expresses a source construct non-obviously.
- Name every result exactly as the source names it, so results can be compared name for name.
- SAS semantics decide the answer even when SAS is not the target language. A SAS date is whole days since 1960-01-01 and a datetime is seconds since 1960-01-01, both stored as plain numbers. A SAS FORMAT changes only how a value is displayed, never what is stored. Missing (.) sorts below every number, so `x < 5` is true when x is missing. Character values are blank-padded to a declared length and compare ignoring trailing blanks.
- A step that only prints or plots produces no data and has no translation outside SAS; leave it out rather than inventing an equivalent.
- Respond with the translated program and nothing else: no prose, no explanation, no markdown code fences, no placeholders.
```

The user turn is three deterministic blocks:

```
Translate the following {SRC} program to {TGT}.

{constraint paragraph — exactly one constant paragraph per TARGET language}

```{fence: python|r|sas}
{source program}
```


Exactly one constraint paragraph exists per target language across all 12,964 rows, identical for both source languages, with consistent fence tags and a header matching a fixed regex. The header doubles as a direction tag. The constraint paragraph — inputs preloaded, no file I/O, no printing, no randomness, named results — is what makes execution-parity evaluation possible.

The assistant turn is the translated program only, with no reasoning trace.

## Training Data

**Proprietary. Not released.** Composition is disclosed below for reproducibility of method, not of data.

| | |
|---|---|
| Source rows | 12,964 (JSON parse 12,964/12,964, zero errors) |
| After direction rebalance | ~15.4k |
| In-training eval subset | 120 (stratified, ~20 per direction) |
| Held-out eval set | 544 |
| Format | JSONL, single-turn `system` + `user` + `assistant` |
| Length | ~4.7k chars/row median; 0.1–0.4% of rows over the 8,192-token cap |

Source direction coverage is all six pairs: Python↔R 2,070 each, Python↔SAS 2,199 each, R↔SAS 2,213 each. Seed-program reuse is high — 2,070 `id`s appear in all six directions and 272 in two. Rebalancing upsampled every direction to the largest, then multiplied the two `*→SAS` directions by 1.5, giving four non-SAS-target directions at 2,213 and two to-SAS directions at ~3,320. SAS is the low-resource target and generation *into* SAS is the hardest case.

Filtering dropped 2 degenerate targets (under 20 characters) and rows over the 8,192-token cap. Overlong rows were **dropped, never truncated** — truncating a translation target teaches the model to emit incomplete programs.

A full-file audit confirmed: every row is exactly `[system, user, assistant]` with string content; the system prompt is byte-identical across all rows; `direction` is 100% consistent with `source_lang`/`target_lang`; targets carry no code fences, prose preambles, `<think>` blocks or special tokens; zero duplicate `(user, assistant)` pairs; and target-language sanity passes at 100%.

**Trap coverage** across the 8,824 SAS-involving rows:

| Trap construct | Rows | Share |
|---|---|---|
| FORMAT / display-only | 2,290 | 26.0% |
| Character padding / comparison | 1,154 | 13.1% |
| Missing-value semantics | 980 | 11.1% |
| **Date/datetime epoch** | **426** | **4.8%** |
| ≥2 traps co-occurring | 1,130 | 12.8% |
| `1960` epoch conversion visible on the Python/R side | 1,660 | 18.8% |

After SFT these rules migrate from prompt-instruction, which the model must apply in a single forward pass with no working space, to the training distribution as pattern recall. That migration only covers traps the data covers, so **date/epoch handling at 4.8% is the identified residual risk** and the thinnest slice.

## Training Procedure

### Adaptable surface

93% of this model was frozen by necessity rather than by choice.

**Routed experts are unreachable.** They hold 93% of the parameters. Unsloth's memory-efficient MoE path requires the `gate_up_proj` layout; the ungated up/down experts here fall back to PEFT's naive path, which materializes roughly a 2.3 GiB fp32 delta per MoE layer per forward — guaranteed OOM on 80 GB. The dotted target names are deliberate: Unsloth's MoE auto-enable fires only on bare `up_proj`/`down_proj`, and a `.` in the name excludes it, so PEFT suffix-matches the shared expert's `nn.Linear` and never wraps the fused routed experts.

**Quantization is unavailable.** bitsandbytes cannot quantize the fused MoE experts, so there are no 4-bit or 8-bit weights. BF16 LoRA only.

**Only 6 of 52 layers have attention.** Adapting attention alone would touch ~11% of layers. The Mamba `in_proj` (23 layers) and the shared expert (23 MoE layers, active on every token) are the load-bearing adaptable surfaces.

**Packing is forbidden.** Mamba-2 layers carry recurrent state across token positions with no attention mask able to sever it. Boundary-aware packing — the standard cross-contamination fix for transformers — does not work here, because packed documents leak state into one another.

Alternatives were considered and rejected: full fine-tuning on cost, catastrophic-forgetting risk, and its need for LR ≈ 5e-6; QLoRA 4-bit as impossible per above; and ESFT, best-in-class for retention in the literature but requiring routed-expert access this stack cannot provide.

### LoRA configuration

| Parameter | Value |
|---|---|
| Rank `r` | 64 |
| `lora_alpha` | 128 (α = 2r) |
| `lora_dropout` | 0 |
| Bias | none |
| Target modules | `q_proj`, `k_proj`, `v_proj`, `o_proj`, `in_proj`, `shared_experts.up_proj`, `shared_experts.down_proj` |
| Scope | Attention (6 of 52 layers), Mamba input projection (23 layers), shared expert (23 MoE layers) |
| Routed experts, router `.gate` | Never targeted |
| Gradient checkpointing | `"unsloth"` (owned by Unsloth; `gradient_checkpointing=False` in `SFTConfig`) |
| `random_state` | 3407 |
| Trainable params | ~220.7M measured at r=16; substantially more at r=64 |

`r=64` sits deliberately above the literature's 16–32: with the routed experts frozen, all adaptation must fit in a small trainable surface. Overfitting risk is handled by best-checkpoint selection rather than by low rank.

Three hard assertions enforce the invariant at build time — no routed-expert parameters requiring grad (would OOM), no gate parameters requiring grad (the router must stay frozen), and trainable > 0 (targets matched something). The router is frozen because on a narrow task a trainable router over-specializes to a few experts, degrading general ability and stability. The auxiliary load-balancing loss was left at its pretraining configuration, preventing expert collapse.

### Optimization

| Parameter | Value |
|---|---|
| Learning rate | 1e-4, cosine |
| Warmup | 3% of total steps |
| Optimizer | `adamw_8bit`, wd 0.01 |
| Gradient clipping | 0.5 |
| Micro-batch × accumulation | 1 × 32 (effective batch 32) |
| Sequence length | 8,192, packing off |
| Loss | Completion-only, prompt masked to −100 |
| Precision | BF16, `sdpa` attention, fp16 off |
| Epochs | 3 scheduled (~482 steps/epoch → max 1,446 steps) |
| Eval / save | Every 100 steps, `save_total_limit=10` |
| Best-model policy | `load_best_model_at_end=True`, `metric_for_best_model="eval_loss"` |
| Early stopping | `patience=3` |
| Seed | 3407 |
| Actual stop | Step 600 |

Loss masking was asserted rather than merely configured: over the first 100 rows, the minimum supervised fraction must exceed 0 (some rows with zero supervised tokens means masking failed) and the mean must stay below 0.95 (otherwise the model is training on the prompt too). The failure mode of broken masking — a permanent loss of 0.0, or training on prompts — is silent and expensive.

Gradient clipping fired on 0% of the last 200 steps, confirming LR 1e-4 as comfortably stable at r=64.

A memory canary ran one worst-case step at full sequence length before committing GPU-hours, with gradient checkpointing enabled exactly as the trainer uses it. A bare forward without checkpointing retains all 52 layers' activations, around 40 GB at 8,192 tokens, and measures a condition training never uses.

**Replay was implemented and deliberately not used.** The configuration already sits at the low end of forgetting risk — LoRA rather than full FT, router and 93% of parameters frozen, LR 1e-4, best-checkpoint selection — and at a fixed step budget 10% replay is a 10% dilution of task signal. The strategy was detect-then-treat: run without replay, measure with a HumanEval canary, add replay only on a >5-point drop. The canary returned 100%, so replay was never required.

### Infrastructure

| | |
|---|---|
| GPU | 1× H100 80GB (RunPod container, `nvcc` 12.8) |
| Parallelism | None — single GPU, `device_map={"": 0}`, no FSDP/DeepSpeed/expert parallelism |
| Framework | unsloth + unsloth_zoo (installed as a matched pair), torch `+cu128`, transformers `>=5.5,<5.6`, trl 1.10.0 |
| Kernels | `mamba_ssm` 2.3.2.post1, `causal_conv1d` 1.6.2.post1, built once and cached |
| Attention | `sdpa` |
| MoE backend | `grouped_mm` |
| Allocator | `expandable_segments:True` |
| Peak memory | 65.2 GiB / 80 GiB (~81%) |
| Throughput | ~5,519 tok/s; ~22 s per generated eval sample at batch-1 greedy |

Distributed-mode suppression is required: `accelerate` initializes a process group merely from the **presence** of distributed environment variables and then refuses to train a device-mapped model. Setting `WORLD_SIZE=1` is insufficient — the variables must be removed.

## Evaluation

**Execution-parity results are outstanding.** 
## Hardware Requirements

| | |
|---|---|
| Weights on disk | ~63 GB (merged 16-bit) |
| Practical single-GPU floor | 80 GB (H100 / H200 / A100 80GB) |
| Peak training memory | 65.2 GiB / 80 GiB (~81%) |
| Context | 8,192 trained; architecture supports up to 1M |
| Quantization | Unavailable — see below |

4-bit and 8-bit loading do not work on this architecture: bitsandbytes cannot quantize the fused MoE experts, which hold 93% of the parameters. The 80 GB floor is hard.

## Citation

```bibtex
@misc{euclid_2_0,
  title  = {Euclid-2.0: Bidirectional SAS/R/Python Program Translation},
  author = {ProCogia},
  year   = {2026},
  url    = {https://huggingface.co/ProCogia/Euclid-2.0}
}
```