SEPIA-0
SEPIA-0 is a 187,104-parameter character-level language model trained on the LUSCA open crypto web corpus. It is trained live by the LUSCA network: the server and volunteer GPUs (WebGPU in the browser) compute gradients for one shared optimizer, and the server checks every GPU gradient before applying it. This upload is the checkpoint at optimizer step 686,029, tagged step-686029.
Summary
| Parameters | 187,104 |
| Architecture | char-MLP · ctx 16 · emb 24 · hidden 384 · tanh |
| Vocabulary | 96 symbols: newline and printable ASCII |
| Context | 16 characters |
| Weights | model.safetensors, float32, 749,304 bytes |
| Step | 686,029 (saved 2026-10-06T05:49:43.175Z) |
| Train loss / validation loss | 1.5463 / 1.7042 nats per character |
| sha256 | 45ddcc4e9da790ded650c41f01f9e1dc204ee40758514d4a5871d233076bb2d1 |
What it is
SEPIA-0 is a proof of the LUSCA training system. It shows that gradients computed by volunteer GPUs can be checked by a server and applied to one shared model while the corpus keeps growing.
The model predicts the next character from the previous 16. It learns spelling, punctuation and short word sequences of crypto web text. It has no understanding of the text, follows no instructions, remembers nothing beyond 16 characters and does not read code. It is not an assistant.
The next model, SEPIA-1, is a transformer that reads blockchain protocol code. It is in design and has no released weights: docs/SEPIA-1.md.
How it was trained
Training runs continuously on the LUSCA server while the corpus grows. There is no fixed number of epochs; each upload here is a snapshot of the live run.
- One optimizer, two sources of steps. The server's own CPU loop (batch 64) and gradients returned by neurons step the same Adam state. Neurons are browser tabs computing with WebGPU (no install) or a desktop program. Every optimizer step increments the weights version.
- Every GPU gradient is checked. The server decodes it, checks that it is finite, checks its norm, compares it with a server gradient on random rows of the same batch (cosine and projection screens) and checks the reported loss. A weak score triggers a full audit.
- Full audits. The server recomputes the whole gradient from the exact weights the job was issued with (kept as float16 snapshots). Defaults: always for a neuron's first 3 results, then 20% at random; a result passes at cosine ≥ 0.99 and relative error ≤ 0.05. Results computed on weights more than 64 versions old are verified but not applied.
- Hyperparameters. Adam (β1 0.9, β2 0.99, ε 1e-8), global gradient-norm clip 1.0, learning rate warmed up linearly over 200 steps to 0.003, cosine decay to 0.0003 over 30,000 steps, then constant. Initialisation: embeddings N(0, 1), W1 N(0, 1/384), W2 N(0, (0.1/√384)²), biases 0.
- Precision. The weights here are the float32 master copy. Neurons receive them rounded to float16 for transfer and return float16 gradients with one power-of-two scale per tensor.
Numbers for this checkpoint, taken from the server's manifest when this folder was built:
| Optimizer steps (weights version) | 686,029 |
| Steps from the server CPU loop | 604,633 |
| Steps from GPU-neuron gradients | 81,396 |
| Training samples processed (context → next character) | 98,791,616 |
| Samples covered by GPU gradients | 60,095,104 |
| Full audits passed / failed | 16,913 / 0 |
| Results rejected by the cheap checks | 1 |
| Verified but too stale to apply | 901 |
| Distinct contributors in the 24 h before this checkpoint | 51 |
Of the 686,029 optimizer steps in this checkpoint, 81,396 were applied from gradients computed by volunteer GPUs and 604,633 by the server's own CPU loop.
Data
The LUSCA open crypto web corpus: public pages fetched by 24 data agents across eight sectors (Governance, Research, Docs, Standards, Markets, Codex, Security, Chronicle), covering DAO forums, protocol research, developer documentation, standards and security write-ups. Only pages the agents keep enter the dataset: each page is scored for relevance and near-duplicates are dropped.
robots.txtis honoured for LuscaBot and*.- AI-training opt-outs are respected: the
robots.txtgroups of AI-training bots,Content-Signal: ai-train=no, TDMRep andnoai. Links into opted-out pages are not followed. - Every host is paced (one request in flight, at least 2 s apart).
- E-mail addresses and phone numbers are redacted before text reaches the dataset or the model.
Text is NFKD-normalised and mapped to the 96-symbol vocabulary: accents are removed, typographic quotes and dashes are folded to ASCII, other characters become a space or are dropped. Every 20th kept document is held out for validation. The training corpus in memory is a window of the most recent pages (24M characters by default; older documents are evicted). The corpus is not included in this repository.
Architecture
ids[16] ─ emb ─▶ x[384] ─ W1, b1 ─▶ tanh ─▶ h[384] ─ W2, b2 ─▶ logits[96]
x = concat(emb[ids_1], …, emb[ids_16]) h = tanh(x·W1 + b1) logits = h·W2 + b2
| Tensor | Shape | Role |
|---|---|---|
emb |
[96, 24] | character embeddings |
W1 |
[384, 384] | hidden layer; input row t·24 + e is component e of context position t (oldest first) |
b1 |
[384] | hidden bias |
W2 |
[384, 96] | output layer |
b2 |
[96] | output bias |
Matrices are row-major [in, out] (x @ W); for torch.nn.Linear use the transpose. Contexts shorter than 16 characters are left-padded with id 0 (newline). The vocabulary in id order is in config.json (vocab) and vocab.json. Total: 96·24 + 384·384 + 384 + 384·96 + 96 = 187,104 parameters.
Usage
pip install numpy huggingface_hub
hf download LUSCAINK/SEPIA-0 --local-dir SEPIA-0
cd SEPIA-0
python inference.py "The validator" --n 240 --temperature 0.8 --seed 7
node sample.mjs "The validator" --n 240 --temperature 0.8 --seed 7
inference.py needs only numpy and reads model.safetensors with a built-in reader. sample.mjs needs only Node.js. Both sample exactly like POST https://lusca.ink/api/generate: the prompt is cut to 200 characters, --n is 1–600 (default 240), --temperature is 0.05–2 (default 0.8), and the output is the prompt followed by the continuation. With --seed, both use the same seeded generator and print the same text; without it, they use system randomness like the server.
Output of python inference.py "The validator" --n 240 --temperature 0.8 --seed 7 for this checkpoint:
The validator
Bysimplementation which confunes? What in the upgradenbase not jullers outabase investment GHO market volumeting and multiple the final $- applaneal opposed in any five holding and some minimon can und not variage minteractive. For exploit
From Python:
from inference import Sepia, encode, generate, mulberry32
model = Sepia("model.safetensors")
print(generate(model, "The validator", n=200, temperature=0.8, rand=mulberry32(7)))
ctx = ([0] * 16 + encode("The ", keep_trailing=True))[-16:] # 16 ids, oldest first
logits = model.logits(ctx) # float32[96], next-character scores
With the safetensors package: safetensors.numpy.load_file("model.safetensors") returns emb, W1, b1, W2, b2 as float32 arrays.
Evaluation
Cross-entropy in nats per character (lower is better), with perplexity and bits per character:
| Loss | Perplexity | Bits/char | |
|---|---|---|---|
| Uniform guess over 96 symbols | 4.5643 (ln 96) | 96.00 | 6.585 |
| Train (mean over the last 25 steps) | 1.5463 | 4.69 | 2.231 |
| Validation (held-out documents, 16 batches of 64) | 1.7042 | 5.50 | 2.459 |
The validation documents are every 20th kept document of the current corpus window, so the score follows the corpus as it changes. Character-level numbers are not comparable with token-level perplexities, and this is not a standard benchmark.
Limitations
- 16-character context: no coherence beyond a few words.
- Output is plausible-looking character sequences, not facts. Do not rely on anything it writes.
- 96-symbol vocabulary: non-ASCII text is folded to ASCII or dropped.
- No instruction tuning and no safety tuning. It can reproduce fragments of public web text it was trained on (contact data is redacted at ingest).
- Training never stops, so the numbers above describe this snapshot only.
- It does not read or write code.
Versions
Each upload is a snapshot of the live run, tagged step-<N> with N the optimizer step. This one is step-686029; main holds the newest. To pin one:
hf download LUSCAINK/SEPIA-0 --revision step-686029 --local-dir SEPIA-0-step-686029
License
inference.py and sample.mjs are MIT licensed, like the LUSCA repository. The weights in this snapshot are released under MIT as well. The project owner may choose a different license for the weights of later snapshots, so check the license of the revision you use.
Links
- https://lusca.ink: live training, network and corpus
- github.com/LUSCAINK/LUSCA: source code (model:
shared/sepia/model.mjs, trainer:server/trainer/, export:server/model/export.ts) - docs/SEPIA-1.md: what comes next
Citation
@misc{sepia0_2026,
title = {SEPIA-0: a character-level language model trained with server-checked volunteer GPU gradients},
author = {{LUSCA contributors}},
year = {2026},
howpublished = {\url{https://huggingface.co/LUSCAINK/SEPIA-0}},
note = {Snapshot step-686029}
}
- Downloads last month
- 19