File size: 7,930 Bytes
07c8ed1
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
0841ed7
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
07c8ed1
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
---

license: apache-2.0
library_name: onnx
tags:
  - onnx
  - onnxruntime
  - text-generation
  - automatic-speech-recognition
  - text-to-speech
  - image-feature-extraction
  - spike-whale
---


# ONNX-Library

ONNX exports of the **Quazim0t0 model family** β€” the SpikeWhale-DNA language models (Byrne / Escarda),
the Byrne ASR and TTS models, and the Byrne-VLM vision encoder. Every graph was exported from the
original PyTorch weights with the legacy TorchScript exporter (opset 17) and **parity-verified**
against the source model.

Each model lives in **its own folder**. The `.onnx` graphs contain the neural network only β€”
tokenizers, text frontends, and CTC/beam decoders stay in the original source repos (linked below),
exactly as they do at inference time.

## Contents

### Language models β€” `SpikeWhaleLM` (13)
`input_ids [B,T] (int64) β†’ logits [B,T,16512]` Β· dynamic batch & sequence Β· file `model.onnx` (~383 MB each)

| Folder | Family trait | Source |
|---|---|---|
| `Byrne-86M` | HRM | [Byrne-86M](https://huggingface.co/Quazim0t0/Byrne-86M) |
| `Byrne-86M-Base` | HRM (base) | [Byrne-86M-Base](https://huggingface.co/Quazim0t0/Byrne-86M-Base) |
| `Byrne-86M-Base-JL` | HRM (base, JL) | [Byrne-86M-Base-JL](https://huggingface.co/Quazim0t0/Byrne-86M-Base-JL) |
| `Byrne-86M-JL` | HRM (JL) | [Byrne-86M-JL](https://huggingface.co/Quazim0t0/Byrne-86M-JL) |
| `Byrne-TriAtn-86M` | HRM (tri-attention) | [Byrne-TriAtn-86M](https://huggingface.co/Quazim0t0/Byrne-TriAtn-86M) |
| `Byrne-TriAtn-86M-JL` | HRM (tri-attn, JL) | [Byrne-TriAtn-86M-JL](https://huggingface.co/Quazim0t0/Byrne-TriAtn-86M-JL) |
| `Escarda-86M` | HRM + JEPA | [Escarda-86M](https://huggingface.co/Quazim0t0/Escarda-86M) |
| `Escarda-86M-Base` | HRM+JEPA (base) | [Escarda-86M-Base](https://huggingface.co/Quazim0t0/Escarda-86M-Base) |
| `Escarda-86M-Base-JL` | HRM+JEPA (base, JL) | [Escarda-86M-Base-JL](https://huggingface.co/Quazim0t0/Escarda-86M-Base-JL) |
| `Escarda-86M-Identity` | HRM+JEPA (identity) | [Escarda-86M-Identity](https://huggingface.co/Quazim0t0/Escarda-86M-Identity) |
| `Escarda-86M-JL` | HRM+JEPA (JL) | [Escarda-86M-JL](https://huggingface.co/Quazim0t0/Escarda-86M-JL) |
| `Escarda-TriAtn-86M` | HRM+JEPA (tri-attn) | [Escarda-TriAtn-86M](https://huggingface.co/Quazim0t0/Escarda-TriAtn-86M) |
| `Escarda-TriAtn-86M-JL` | HRM+JEPA (tri-attn, JL) | [Escarda-TriAtn-86M-JL](https://huggingface.co/Quazim0t0/Escarda-TriAtn-86M-JL) |

Tokenizer (`tokenizer.json` + `spike_tokenizer.py`) is in each source repo. Verified **100 % argmax-token

agreement** with PyTorch (the deep custom ops add ~1e-1 fp noise to the wide logits β€” harmless; base
variants are near-exact).

### Speech recognition β€” `Byrne-ASR-English/model.onnx` (~50 MB)
`mel [B,80,T] (float32) β†’ logits [B,T',29]` (CTC, dynamic frames). Mel frontend params: sample_rate 24000,

n_fft 1024, hop 256, n_mels 80, log-mel. Vocab: `<blank>`, space, a–z, `'`. The lexicon / bigram / ARPA

beam-search decode lives in the source repo ([Byrne-ASR-English](https://huggingface.co/Quazim0t0/Byrne-ASR-English)).

Parity 7e-6.



### Vision encoder β€” `Byrne-VLM-131M/vision.onnx` (~167 MB)

`image [B,3,448,448] (float32, [-1,1]) β†’ pooled [B,512] + tokens [B,784,512]`. ViT-style, patch 16, native

**448Γ—448** (28Γ—28 patch grid), 2D axial RoPE. Source: [Byrne-VLM-131M](https://huggingface.co/Quazim0t0/Byrne-VLM-131M).

Parity 1e-6. *(The multimodal LM half is not included here.)*



### Text-to-speech β€” `Byrne-Speech/` (2-stage, ~49 MB)

1. `acoustic.onnx` β€” FastSpeech2: `ids [B,Tp] (int64) + plen [B] β†’ mel [B,80,Tm]` (variable length; length

   regulator generalizes across text lengths).

2. `vocoder.onnx` β€” HiFi-GAN: `mel [B,80,Tm] β†’ wav [B,1,Tm*256]` (24 kHz, hop 256). `weight_norm` folded.

Char text frontend (`text_to_char_sequence`) is in the source repo ([Byrne-Speech](https://huggingface.co/Quazim0t0/Byrne-Speech)).
Chain: `text β†’ acoustic.onnx β†’ mel β†’ vocoder.onnx β†’ wav`. Parity 7e-7.

### Tools (4)

| Folder | Task | I/O contract | Source |
|---|---|---|---|
| `Escarda-Rewrite/model.onnx` | text rewriting (causal LM) | `input_ids[B,T] -> logits[B,T,16512]` | [Escarda-Rewrite](https://huggingface.co/Quazim0t0/Escarda-Rewrite) |
| `Byrne-Embed/model.onnx` | text embeddings | `input_ids[B,T] -> embedding[B,768]` (pooled sentence vector) | [Byrne-Embed](https://huggingface.co/Quazim0t0/Byrne-Embed) |
| `Byrne-Anon/model.onnx` | PII tagging (BIOES) | `input_ids[B,T] -> pii_logits[B,T,33]` (labels in source pii_labels.json) | [Byrne-Anon](https://huggingface.co/Quazim0t0/Byrne-Anon) |

| `Byrne-Docling-131M/vision.onnx` | document VLM vision encoder | `image[B,3,448,448] -> pooled[B,512] + tokens[B,784,512]` | [Byrne-Docling-131M](https://huggingface.co/Quazim0t0/Byrne-Docling-131M) |



## Vision-language generation (full pipeline)



`Byrne-VLM-131M` (captioning) and `Byrne-Docling-131M` (document -> DocTags) are **generative** β€” each

ships THREE files for real image->text generation (the single `vision.onnx` is the encoder only):



- `vision_connector.onnx` : `image[1,3,448,448] -> image_embeds[1,784,640]` (vision encoder + projector)

- `lm_decode.onnx`        : `inputs_embeds[1,T,640] -> logits[1,T,V]` (LoRA-applied LM)

- `embed_tokens.npy`      : `[V,640]` token-embedding table (for generated text tokens)

**Generate:** encode the image once, feed `image_embeds` as the prefix, then autoregressively append
`embed_tokens[next_token]` and re-run `lm_decode`:
```python

import numpy as np, onnxruntime as ort

va=ort.InferenceSession("Byrne-VLM-131M/vision_connector.onnx")

lm=ort.InferenceSession("Byrne-VLM-131M/lm_decode.onnx")

emb=np.load("Byrne-VLM-131M/embed_tokens.npy")

ie=va.run(["image_embeds"],{"image":img})[0]          # img: [1,3,448,448] float32 in [-1,1]

out=[]

for _ in range(48):

    x = ie if not out else np.concatenate([ie, emb[out][None]], 1)

    nxt = int(lm.run(["logits"],{"inputs_embeds":x})[0][0,-1].argmax())

    if nxt==EOS: break

    out.append(nxt)                                    # decode with the model's tokenizer

```
Verified to produce PyTorch-identical greedy output (Byrne-VLM: *"A group of people standing on the

ground."*). Byrne-VLM decodes with the shared LM tokenizer; Byrne-Docling with its source
`tokenizer_doctags.json`.

## Usage (ONNX Runtime)

**Language model** (needs the tokenizer from the source repo):
```python

import onnxruntime as ort, numpy as np

sess = ort.InferenceSession("Byrne-86M/model.onnx", providers=["CPUExecutionProvider"])

input_ids = np.array([[1, 23, 45, 6]], dtype=np.int64)          # from the SpikeWhale tokenizer

logits = sess.run(["logits"], {"input_ids": input_ids})[0]      # [1, T, 16512]

next_id = logits[0, -1].argmax()                                # greedy next token

```

**TTS** (chain the two graphs):
```python

import onnxruntime as ort, numpy as np

ac = ort.InferenceSession("Byrne-Speech/acoustic.onnx", providers=["CPUExecutionProvider"])

vo = ort.InferenceSession("Byrne-Speech/vocoder.onnx",  providers=["CPUExecutionProvider"])

ids  = np.array([[...]], dtype=np.int64)                         # text_to_char_sequence(text)

plen = np.array([ids.shape[1]], dtype=np.int64)

mel = ac.run(["mel"], {"ids": ids, "plen": plen})[0]            # [1,80,Tm]

wav = vo.run(["wav"], {"mel": mel})[0]                          # [1,1,Tm*256] @ 24 kHz

```

**Vision encoder:**
```python

pooled, tokens = sess.run(["pooled","tokens"], {"images": img})  # img [B,3,448,448] float32 in [-1,1]

```

All graphs use dynamic batch (and dynamic sequence/frames where noted), so batching works out of the box.

## Provenance

Exported 1:1 from the PyTorch checkpoints in the linked source repos (opset 17, `dynamo=False`), each
verified against its original model. License: Apache-2.0.