File size: 7,930 Bytes
07c8ed1 0841ed7 07c8ed1 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 | ---
license: apache-2.0
library_name: onnx
tags:
- onnx
- onnxruntime
- text-generation
- automatic-speech-recognition
- text-to-speech
- image-feature-extraction
- spike-whale
---
# ONNX-Library
ONNX exports of the **Quazim0t0 model family** β the SpikeWhale-DNA language models (Byrne / Escarda),
the Byrne ASR and TTS models, and the Byrne-VLM vision encoder. Every graph was exported from the
original PyTorch weights with the legacy TorchScript exporter (opset 17) and **parity-verified**
against the source model.
Each model lives in **its own folder**. The `.onnx` graphs contain the neural network only β
tokenizers, text frontends, and CTC/beam decoders stay in the original source repos (linked below),
exactly as they do at inference time.
## Contents
### Language models β `SpikeWhaleLM` (13)
`input_ids [B,T] (int64) β logits [B,T,16512]` Β· dynamic batch & sequence Β· file `model.onnx` (~383 MB each)
| Folder | Family trait | Source |
|---|---|---|
| `Byrne-86M` | HRM | [Byrne-86M](https://huggingface.co/Quazim0t0/Byrne-86M) |
| `Byrne-86M-Base` | HRM (base) | [Byrne-86M-Base](https://huggingface.co/Quazim0t0/Byrne-86M-Base) |
| `Byrne-86M-Base-JL` | HRM (base, JL) | [Byrne-86M-Base-JL](https://huggingface.co/Quazim0t0/Byrne-86M-Base-JL) |
| `Byrne-86M-JL` | HRM (JL) | [Byrne-86M-JL](https://huggingface.co/Quazim0t0/Byrne-86M-JL) |
| `Byrne-TriAtn-86M` | HRM (tri-attention) | [Byrne-TriAtn-86M](https://huggingface.co/Quazim0t0/Byrne-TriAtn-86M) |
| `Byrne-TriAtn-86M-JL` | HRM (tri-attn, JL) | [Byrne-TriAtn-86M-JL](https://huggingface.co/Quazim0t0/Byrne-TriAtn-86M-JL) |
| `Escarda-86M` | HRM + JEPA | [Escarda-86M](https://huggingface.co/Quazim0t0/Escarda-86M) |
| `Escarda-86M-Base` | HRM+JEPA (base) | [Escarda-86M-Base](https://huggingface.co/Quazim0t0/Escarda-86M-Base) |
| `Escarda-86M-Base-JL` | HRM+JEPA (base, JL) | [Escarda-86M-Base-JL](https://huggingface.co/Quazim0t0/Escarda-86M-Base-JL) |
| `Escarda-86M-Identity` | HRM+JEPA (identity) | [Escarda-86M-Identity](https://huggingface.co/Quazim0t0/Escarda-86M-Identity) |
| `Escarda-86M-JL` | HRM+JEPA (JL) | [Escarda-86M-JL](https://huggingface.co/Quazim0t0/Escarda-86M-JL) |
| `Escarda-TriAtn-86M` | HRM+JEPA (tri-attn) | [Escarda-TriAtn-86M](https://huggingface.co/Quazim0t0/Escarda-TriAtn-86M) |
| `Escarda-TriAtn-86M-JL` | HRM+JEPA (tri-attn, JL) | [Escarda-TriAtn-86M-JL](https://huggingface.co/Quazim0t0/Escarda-TriAtn-86M-JL) |
Tokenizer (`tokenizer.json` + `spike_tokenizer.py`) is in each source repo. Verified **100 % argmax-token
agreement** with PyTorch (the deep custom ops add ~1e-1 fp noise to the wide logits β harmless; base
variants are near-exact).
### Speech recognition β `Byrne-ASR-English/model.onnx` (~50 MB)
`mel [B,80,T] (float32) β logits [B,T',29]` (CTC, dynamic frames). Mel frontend params: sample_rate 24000,
n_fft 1024, hop 256, n_mels 80, log-mel. Vocab: `<blank>`, space, aβz, `'`. The lexicon / bigram / ARPA
beam-search decode lives in the source repo ([Byrne-ASR-English](https://huggingface.co/Quazim0t0/Byrne-ASR-English)).
Parity 7e-6.
### Vision encoder β `Byrne-VLM-131M/vision.onnx` (~167 MB)
`image [B,3,448,448] (float32, [-1,1]) β pooled [B,512] + tokens [B,784,512]`. ViT-style, patch 16, native
**448Γ448** (28Γ28 patch grid), 2D axial RoPE. Source: [Byrne-VLM-131M](https://huggingface.co/Quazim0t0/Byrne-VLM-131M).
Parity 1e-6. *(The multimodal LM half is not included here.)*
### Text-to-speech β `Byrne-Speech/` (2-stage, ~49 MB)
1. `acoustic.onnx` β FastSpeech2: `ids [B,Tp] (int64) + plen [B] β mel [B,80,Tm]` (variable length; length
regulator generalizes across text lengths).
2. `vocoder.onnx` β HiFi-GAN: `mel [B,80,Tm] β wav [B,1,Tm*256]` (24 kHz, hop 256). `weight_norm` folded.
Char text frontend (`text_to_char_sequence`) is in the source repo ([Byrne-Speech](https://huggingface.co/Quazim0t0/Byrne-Speech)).
Chain: `text β acoustic.onnx β mel β vocoder.onnx β wav`. Parity 7e-7.
### Tools (4)
| Folder | Task | I/O contract | Source |
|---|---|---|---|
| `Escarda-Rewrite/model.onnx` | text rewriting (causal LM) | `input_ids[B,T] -> logits[B,T,16512]` | [Escarda-Rewrite](https://huggingface.co/Quazim0t0/Escarda-Rewrite) |
| `Byrne-Embed/model.onnx` | text embeddings | `input_ids[B,T] -> embedding[B,768]` (pooled sentence vector) | [Byrne-Embed](https://huggingface.co/Quazim0t0/Byrne-Embed) |
| `Byrne-Anon/model.onnx` | PII tagging (BIOES) | `input_ids[B,T] -> pii_logits[B,T,33]` (labels in source pii_labels.json) | [Byrne-Anon](https://huggingface.co/Quazim0t0/Byrne-Anon) |
| `Byrne-Docling-131M/vision.onnx` | document VLM vision encoder | `image[B,3,448,448] -> pooled[B,512] + tokens[B,784,512]` | [Byrne-Docling-131M](https://huggingface.co/Quazim0t0/Byrne-Docling-131M) |
## Vision-language generation (full pipeline)
`Byrne-VLM-131M` (captioning) and `Byrne-Docling-131M` (document -> DocTags) are **generative** β each
ships THREE files for real image->text generation (the single `vision.onnx` is the encoder only):
- `vision_connector.onnx` : `image[1,3,448,448] -> image_embeds[1,784,640]` (vision encoder + projector)
- `lm_decode.onnx` : `inputs_embeds[1,T,640] -> logits[1,T,V]` (LoRA-applied LM)
- `embed_tokens.npy` : `[V,640]` token-embedding table (for generated text tokens)
**Generate:** encode the image once, feed `image_embeds` as the prefix, then autoregressively append
`embed_tokens[next_token]` and re-run `lm_decode`:
```python
import numpy as np, onnxruntime as ort
va=ort.InferenceSession("Byrne-VLM-131M/vision_connector.onnx")
lm=ort.InferenceSession("Byrne-VLM-131M/lm_decode.onnx")
emb=np.load("Byrne-VLM-131M/embed_tokens.npy")
ie=va.run(["image_embeds"],{"image":img})[0] # img: [1,3,448,448] float32 in [-1,1]
out=[]
for _ in range(48):
x = ie if not out else np.concatenate([ie, emb[out][None]], 1)
nxt = int(lm.run(["logits"],{"inputs_embeds":x})[0][0,-1].argmax())
if nxt==EOS: break
out.append(nxt) # decode with the model's tokenizer
```
Verified to produce PyTorch-identical greedy output (Byrne-VLM: *"A group of people standing on the
ground."*). Byrne-VLM decodes with the shared LM tokenizer; Byrne-Docling with its source
`tokenizer_doctags.json`.
## Usage (ONNX Runtime)
**Language model** (needs the tokenizer from the source repo):
```python
import onnxruntime as ort, numpy as np
sess = ort.InferenceSession("Byrne-86M/model.onnx", providers=["CPUExecutionProvider"])
input_ids = np.array([[1, 23, 45, 6]], dtype=np.int64) # from the SpikeWhale tokenizer
logits = sess.run(["logits"], {"input_ids": input_ids})[0] # [1, T, 16512]
next_id = logits[0, -1].argmax() # greedy next token
```
**TTS** (chain the two graphs):
```python
import onnxruntime as ort, numpy as np
ac = ort.InferenceSession("Byrne-Speech/acoustic.onnx", providers=["CPUExecutionProvider"])
vo = ort.InferenceSession("Byrne-Speech/vocoder.onnx", providers=["CPUExecutionProvider"])
ids = np.array([[...]], dtype=np.int64) # text_to_char_sequence(text)
plen = np.array([ids.shape[1]], dtype=np.int64)
mel = ac.run(["mel"], {"ids": ids, "plen": plen})[0] # [1,80,Tm]
wav = vo.run(["wav"], {"mel": mel})[0] # [1,1,Tm*256] @ 24 kHz
```
**Vision encoder:**
```python
pooled, tokens = sess.run(["pooled","tokens"], {"images": img}) # img [B,3,448,448] float32 in [-1,1]
```
All graphs use dynamic batch (and dynamic sequence/frames where noted), so batching works out of the box.
## Provenance
Exported 1:1 from the PyTorch checkpoints in the linked source repos (opset 17, `dynamo=False`), each
verified against its original model. License: Apache-2.0.
|