SEPIA-0

SEPIA-0 is a 187,104-parameter character-level language model trained on the LUSCA open crypto web corpus. It is trained live by the LUSCA network: the server and volunteer GPUs (WebGPU in the browser) compute gradients for one shared optimizer, and the server checks every GPU gradient before applying it. This upload is the checkpoint at optimizer step 686,029, tagged step-686029.

Summary

Parameters 187,104
Architecture char-MLP · ctx 16 · emb 24 · hidden 384 · tanh
Vocabulary 96 symbols: newline and printable ASCII
Context 16 characters
Weights model.safetensors, float32, 749,304 bytes
Step 686,029 (saved 2026-10-06T05:49:43.175Z)
Train loss / validation loss 1.5463 / 1.7042 nats per character
sha256 45ddcc4e9da790ded650c41f01f9e1dc204ee40758514d4a5871d233076bb2d1

What it is

SEPIA-0 is a proof of the LUSCA training system. It shows that gradients computed by volunteer GPUs can be checked by a server and applied to one shared model while the corpus keeps growing.

The model predicts the next character from the previous 16. It learns spelling, punctuation and short word sequences of crypto web text. It has no understanding of the text, follows no instructions, remembers nothing beyond 16 characters and does not read code. It is not an assistant.

The next model, SEPIA-1, is a transformer that reads blockchain protocol code. It is in design and has no released weights: docs/SEPIA-1.md.

How it was trained

Training runs continuously on the LUSCA server while the corpus grows. There is no fixed number of epochs; each upload here is a snapshot of the live run.

  • One optimizer, two sources of steps. The server's own CPU loop (batch 64) and gradients returned by neurons step the same Adam state. Neurons are browser tabs computing with WebGPU (no install) or a desktop program. Every optimizer step increments the weights version.
  • Every GPU gradient is checked. The server decodes it, checks that it is finite, checks its norm, compares it with a server gradient on random rows of the same batch (cosine and projection screens) and checks the reported loss. A weak score triggers a full audit.
  • Full audits. The server recomputes the whole gradient from the exact weights the job was issued with (kept as float16 snapshots). Defaults: always for a neuron's first 3 results, then 20% at random; a result passes at cosine ≥ 0.99 and relative error ≤ 0.05. Results computed on weights more than 64 versions old are verified but not applied.
  • Hyperparameters. Adam (β1 0.9, β2 0.99, ε 1e-8), global gradient-norm clip 1.0, learning rate warmed up linearly over 200 steps to 0.003, cosine decay to 0.0003 over 30,000 steps, then constant. Initialisation: embeddings N(0, 1), W1 N(0, 1/384), W2 N(0, (0.1/√384)²), biases 0.
  • Precision. The weights here are the float32 master copy. Neurons receive them rounded to float16 for transfer and return float16 gradients with one power-of-two scale per tensor.

Numbers for this checkpoint, taken from the server's manifest when this folder was built:

Optimizer steps (weights version) 686,029
Steps from the server CPU loop 604,633
Steps from GPU-neuron gradients 81,396
Training samples processed (context → next character) 98,791,616
Samples covered by GPU gradients 60,095,104
Full audits passed / failed 16,913 / 0
Results rejected by the cheap checks 1
Verified but too stale to apply 901
Distinct contributors in the 24 h before this checkpoint 51

Of the 686,029 optimizer steps in this checkpoint, 81,396 were applied from gradients computed by volunteer GPUs and 604,633 by the server's own CPU loop.

Data

The LUSCA open crypto web corpus: public pages fetched by 24 data agents across eight sectors (Governance, Research, Docs, Standards, Markets, Codex, Security, Chronicle), covering DAO forums, protocol research, developer documentation, standards and security write-ups. Only pages the agents keep enter the dataset: each page is scored for relevance and near-duplicates are dropped.

  • robots.txt is honoured for LuscaBot and *.
  • AI-training opt-outs are respected: the robots.txt groups of AI-training bots, Content-Signal: ai-train=no, TDMRep and noai. Links into opted-out pages are not followed.
  • Every host is paced (one request in flight, at least 2 s apart).
  • E-mail addresses and phone numbers are redacted before text reaches the dataset or the model.

Text is NFKD-normalised and mapped to the 96-symbol vocabulary: accents are removed, typographic quotes and dashes are folded to ASCII, other characters become a space or are dropped. Every 20th kept document is held out for validation. The training corpus in memory is a window of the most recent pages (24M characters by default; older documents are evicted). The corpus is not included in this repository.

Architecture

ids[16] ─ emb ─▶ x[384] ─ W1, b1 ─▶ tanh ─▶ h[384] ─ W2, b2 ─▶ logits[96]
x = concat(emb[ids_1], …, emb[ids_16])     h = tanh(x·W1 + b1)     logits = h·W2 + b2
Tensor Shape Role
emb [96, 24] character embeddings
W1 [384, 384] hidden layer; input row t·24 + e is component e of context position t (oldest first)
b1 [384] hidden bias
W2 [384, 96] output layer
b2 [96] output bias

Matrices are row-major [in, out] (x @ W); for torch.nn.Linear use the transpose. Contexts shorter than 16 characters are left-padded with id 0 (newline). The vocabulary in id order is in config.json (vocab) and vocab.json. Total: 96·24 + 384·384 + 384 + 384·96 + 96 = 187,104 parameters.

Usage

pip install numpy huggingface_hub
hf download LUSCAINK/SEPIA-0 --local-dir SEPIA-0
cd SEPIA-0
python inference.py "The validator" --n 240 --temperature 0.8 --seed 7
node sample.mjs "The validator" --n 240 --temperature 0.8 --seed 7

inference.py needs only numpy and reads model.safetensors with a built-in reader. sample.mjs needs only Node.js. Both sample exactly like POST https://lusca.ink/api/generate: the prompt is cut to 200 characters, --n is 1–600 (default 240), --temperature is 0.05–2 (default 0.8), and the output is the prompt followed by the continuation. With --seed, both use the same seeded generator and print the same text; without it, they use system randomness like the server.

Output of python inference.py "The validator" --n 240 --temperature 0.8 --seed 7 for this checkpoint:

The validator
Bysimplementation which confunes? What in the upgradenbase not jullers outabase investment GHO market volumeting and multiple the final $- applaneal opposed in any five holding and some minimon can und not variage minteractive. For exploit

From Python:

from inference import Sepia, encode, generate, mulberry32
model = Sepia("model.safetensors")
print(generate(model, "The validator", n=200, temperature=0.8, rand=mulberry32(7)))
ctx = ([0] * 16 + encode("The ", keep_trailing=True))[-16:]  # 16 ids, oldest first
logits = model.logits(ctx)  # float32[96], next-character scores

With the safetensors package: safetensors.numpy.load_file("model.safetensors") returns emb, W1, b1, W2, b2 as float32 arrays.

Evaluation

Cross-entropy in nats per character (lower is better), with perplexity and bits per character:

Loss Perplexity Bits/char
Uniform guess over 96 symbols 4.5643 (ln 96) 96.00 6.585
Train (mean over the last 25 steps) 1.5463 4.69 2.231
Validation (held-out documents, 16 batches of 64) 1.7042 5.50 2.459

The validation documents are every 20th kept document of the current corpus window, so the score follows the corpus as it changes. Character-level numbers are not comparable with token-level perplexities, and this is not a standard benchmark.

Limitations

  • 16-character context: no coherence beyond a few words.
  • Output is plausible-looking character sequences, not facts. Do not rely on anything it writes.
  • 96-symbol vocabulary: non-ASCII text is folded to ASCII or dropped.
  • No instruction tuning and no safety tuning. It can reproduce fragments of public web text it was trained on (contact data is redacted at ingest).
  • Training never stops, so the numbers above describe this snapshot only.
  • It does not read or write code.

Versions

Each upload is a snapshot of the live run, tagged step-<N> with N the optimizer step. This one is step-686029; main holds the newest. To pin one:

hf download LUSCAINK/SEPIA-0 --revision step-686029 --local-dir SEPIA-0-step-686029

License

inference.py and sample.mjs are MIT licensed, like the LUSCA repository. The weights in this snapshot are released under MIT as well. The project owner may choose a different license for the weights of later snapshots, so check the license of the revision you use.

Links

Citation

@misc{sepia0_2026,
  title        = {SEPIA-0: a character-level language model trained with server-checked volunteer GPU gradients},
  author       = {{LUSCA contributors}},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/LUSCAINK/SEPIA-0}},
  note         = {Snapshot step-686029}
}
Downloads last month
19
Safetensors
Model size
187k params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support