docs: expand Z1T-0 model card from official Z1T publication
#1
by tljohnsilver - opened
README.md
CHANGED
|
@@ -1,10 +1,390 @@
|
|
| 1 |
---
|
| 2 |
library_name: z1t
|
| 3 |
pipeline_tag: text-generation
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 4 |
---
|
| 5 |
|
| 6 |
# Z1T-0
|
| 7 |
|
| 8 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 9 |
|
| 10 |
-
[GitHub](https://github.com/extropic-ai/sparse-transformers)
|
|
|
|
| 1 |
---
|
| 2 |
library_name: z1t
|
| 3 |
pipeline_tag: text-generation
|
| 4 |
+
language: en
|
| 5 |
+
license: apache-2.0
|
| 6 |
+
tags:
|
| 7 |
+
- z1t
|
| 8 |
+
- sparse
|
| 9 |
+
- probabilistic-computing
|
| 10 |
+
- thermodynamic-computing
|
| 11 |
+
- energy-efficient
|
| 12 |
+
- gated-convolutional-attention
|
| 13 |
+
- text-generation
|
| 14 |
+
datasets:
|
| 15 |
+
- openwebtext
|
| 16 |
---
|
| 17 |
|
| 18 |
# Z1T-0
|
| 19 |
|
| 20 |
+
## Summary
|
| 21 |
+
|
| 22 |
+
Z1T-0 is the first open weight release in the Z1T family from Extropic Corp. Z1T is a family of sparse transformer like language models designed for probabilistic hardware, pairing Extropic Z1 probabilistic chips with FPGA or XPU companions for disaggregated inference.
|
| 23 |
+
|
| 24 |
+
The design demonstrates substantial modeled energy efficiency gains over GPUs, reported as up to about 139x lower energy per token than an H100 at 10 percent model FLOPs utilization for the reference configuration, while preserving the familiar transformer outline with hardware compatible primitives.
|
| 25 |
+
|
| 26 |
+
This release provides the Z1T-0 checkpoint for research on sparse scaling, hardware algorithm codesign, and energy efficient decoding. Training and architecture code is available in the official sparse transformers repository.
|
| 27 |
+
|
| 28 |
+
Official resources:
|
| 29 |
+
|
| 30 |
+
* Publication: https://extropic.ai/writing/z1t/
|
| 31 |
+
* Training recipe: https://github.com/extropic-ai/sparse-transformers
|
| 32 |
+
* Organization: https://huggingface.co/Extropic-AI
|
| 33 |
+
* Z1 background: https://extropic.ai/writing/from-one-to-one-billion
|
| 34 |
+
* Z1 launch video: https://www.youtube.com/watch?v=ITkoPau3k98
|
| 35 |
+
|
| 36 |
+
## Model Details
|
| 37 |
+
|
| 38 |
+
* Developer: Extropic Corp, San Francisco
|
| 39 |
+
* Authors: Guillaume Verdon, Alexander Neagoe, Owen Lockwood, Seth Morton, and Extropic
|
| 40 |
+
* Release date: September 4, 2026
|
| 41 |
+
* Version: Z1T-0, first public Z1T checkpoint
|
| 42 |
+
* Model type: Sparse causal language model with gated convolutional attention
|
| 43 |
+
* Language: English
|
| 44 |
+
* Tokenizer: GPT-2 byte pair encoding, vocabulary size 50257
|
| 45 |
+
* Context length of released checkpoint: 256 tokens
|
| 46 |
+
* Framework: JAX with Equinox, custom `z1t` library
|
| 47 |
+
* Repository files: `config.json`, `load_model.py`, `model.eqx`
|
| 48 |
+
* Checkpoint size: approximately 4.97 GB
|
| 49 |
+
* License: Apache 2.0, matching the official sparse transformers repository. If Extropic publishes a distinct weight license, that notice will take precedence.
|
| 50 |
+
* Status: Research release. Energy and latency figures in the publication are modeled estimates, not full system silicon measurements.
|
| 51 |
+
|
| 52 |
+
### Released Checkpoint Configuration
|
| 53 |
+
|
| 54 |
+
Values below are taken directly from the released `config.json`:
|
| 55 |
+
|
| 56 |
+
* vocab: 50257
|
| 57 |
+
* sequence: 256
|
| 58 |
+
* n_layers: 4
|
| 59 |
+
* n_embed: 12288
|
| 60 |
+
* aft_kind: conv
|
| 61 |
+
* aft_heads: 4
|
| 62 |
+
* aft_ksize: 4
|
| 63 |
+
* dyt_alpha: 0.5
|
| 64 |
+
* linear_fan_in: 4
|
| 65 |
+
* attn_fan_in: null
|
| 66 |
+
* mlp_fan_in: null
|
| 67 |
+
* tanh_linear: true
|
| 68 |
+
* tanh_mlp: true
|
| 69 |
+
* remat: true
|
| 70 |
+
|
| 71 |
+
Note on reference versus released dimensions: the publication uses a compact operating point with L 4, D 512, T 1024, H 4, kernel 4, and sparse fan in 4 for concrete energy and latency modeling. The released Z1T-0 checkpoint is wider, with n_embed 12288 and sequence 256. Treat scaling plots and hardware estimates as family level evidence, and the configuration above as the exact artifact in this repository.
|
| 72 |
+
|
| 73 |
+
## Intended Use
|
| 74 |
+
|
| 75 |
+
### Intended research uses
|
| 76 |
+
|
| 77 |
+
* Study of sparse transformer scaling under fixed hardware connectivity
|
| 78 |
+
* Hardware algorithm codesign for probabilistic and thermodynamic substrates
|
| 79 |
+
* Energy efficient autoregressive decoding research
|
| 80 |
+
* Reproduction and extension of gated convolutional attention with tanh linear projections
|
| 81 |
+
* Analysis of sparsity, quantization, and sampling precision tradeoffs
|
| 82 |
+
|
| 83 |
+
### Out of scope uses
|
| 84 |
+
|
| 85 |
+
* Production deployment without additional evaluation for safety, bias, and robustness
|
| 86 |
+
* Long context prefill on Z1 style hardware, which the authors recommend running on GPUs for now
|
| 87 |
+
* High precision computation without accounting for sampling noise and quantization
|
| 88 |
+
* Use as a general instruction tuned assistant, as this is a base research checkpoint
|
| 89 |
+
* Any use implying measured full chip system guarantees, since published system figures are modeled
|
| 90 |
+
|
| 91 |
+
## Hardware Context: Z1 plus FPGA
|
| 92 |
+
|
| 93 |
+
Z1 was designed as a probabilistic graphical model sampler, not originally as a neural network accelerator. Z1T refactors dense transformer layers so compatible sparse operations compile onto the fixed Z1 topology, while remaining dense or irregular operations run on a companion FPGA or XPU.
|
| 94 |
+
|
| 95 |
+
Z1 characteristics reported by Extropic:
|
| 96 |
+
|
| 97 |
+
* 269,568 probabilistic bits across 8 cores, with 33,696 pbits per core
|
| 98 |
+
* 2,135,904 hardwired coupling edges
|
| 99 |
+
* Fixed node degree 16
|
| 100 |
+
* Chromatic Gibbs sampling in memory at 50 MHz internal update clock
|
| 101 |
+
* Subthreshold CMOS implementation
|
| 102 |
+
* Sampling energy estimate: 1.3e-14 J per sample
|
| 103 |
+
* Die level target: sampling rate above 50 MHz at below 1 W
|
| 104 |
+
* Data transfer assumption for modeling: 25.6 Gbit per s
|
| 105 |
+
|
| 106 |
+
The central constraint is connectivity. A program maps efficiently only if its interaction graph is a subgraph of the fixed coupling graph. Dense matrix multiplication therefore is not ported unchanged. It is rewritten as many parallel sparse tanh linear units that respect the degree 16 pattern.
|
| 107 |
+
|
| 108 |
+
Inference is disaggregated:
|
| 109 |
+
|
| 110 |
+
* Z1: sampled sparse projections, Dynamic Tanh style normalization fused with following layers, sparse MLP segments
|
| 111 |
+
* FPGA or XPU: embeddings, pooling, residuals, positional handling, orchestration, vocabulary readout and sampling
|
| 112 |
+
|
| 113 |
+
The publication notes that the companion processor could in principle be a GPU or another accelerator, but all concrete estimates assume the FPGA companion.
|
| 114 |
+
|
| 115 |
+
## Architecture
|
| 116 |
+
|
| 117 |
+
Z1T preserves the transformer block outline but replaces dense primitives with Z1 compatible forms.
|
| 118 |
+
|
| 119 |
+
### Dynamic Tanh normalization
|
| 120 |
+
|
| 121 |
+
The reference block replaces RMSNorm:
|
| 122 |
+
|
| 123 |
+
RMSNorm(x) = x / sqrt(mean(x^2) + eps) * gamma + beta
|
| 124 |
+
|
| 125 |
+
with a scaled Dynamic Tanh form:
|
| 126 |
+
|
| 127 |
+
DyT(x) = gamma * tanh(alpha * x) + beta
|
| 128 |
+
|
| 129 |
+
This substitution follows recent work on normalization free transformers and fuses naturally with subsequent sparse linear operations on Z1.
|
| 130 |
+
|
| 131 |
+
### Tanh linear unit
|
| 132 |
+
|
| 133 |
+
The core hardware primitive is a probabilistic cell whose visible spin expectation is a tanh of its local field. For visible spin v and 16 neighbors h:
|
| 134 |
+
|
| 135 |
+
E(v, h) = -(b_v * v + sum(J_j * v * h_j) + sum(b_j * h_j))
|
| 136 |
+
|
| 137 |
+
Conditioned on h, the expectation is:
|
| 138 |
+
|
| 139 |
+
E[v | h] = tanh(b_v + sum(J_j * h_j))
|
| 140 |
+
|
| 141 |
+
Averaged over parallel samples, this implements a sparse matrix vector product fused with tanh. Each output pbit uses its 16 couplings for 4 input values times 4 dyadic pbits.
|
| 142 |
+
|
| 143 |
+
### Continuous value encoding
|
| 144 |
+
|
| 145 |
+
Activations and weights are continuous, while Z1 operates on binary stochastic units represented as spins in (-1, 1). Z1T uses a 4 bit dyadic encoding called dy4p:
|
| 146 |
+
|
| 147 |
+
x approx sum(a_i * s_i), with s in (-1, 1)^4 and a_i = 2^(-i-1)
|
| 148 |
+
|
| 149 |
+
A weighted activation is therefore encoded as:
|
| 150 |
+
|
| 151 |
+
tanh(W x + b) approx tanh((W kron a) s + b)
|
| 152 |
+
|
| 153 |
+
For an input in R^D with 4 spins per value, the weight matrix is encoded directly into interactions over the flattened spin vector. Parallel samples are averaged to estimate the scalar tanh linear output. More samples reduce empirical variance and increase effective precision at the same energy per sample but higher latency if serialized.
|
| 154 |
+
|
| 155 |
+
### Gated convolutional attention
|
| 156 |
+
|
| 157 |
+
Instead of dense softmax self attention with Q K transpose scoring of order T squared D, Z1T adapts gated convolutional attention. Sparse 4 fan in projections compute gates Q, K, and V. For token position t and head i:
|
| 158 |
+
|
| 159 |
+
Y_t^i = tanh(Q_t^i) * (N_t^i / D_t^i)
|
| 160 |
+
|
| 161 |
+
N_t^i = conv1d(exp(K^i) * V^i, exp(w^i) - 1) + sum(exp(K_j^i) * V_j^i) over j 1 to t
|
| 162 |
+
|
| 163 |
+
D_t^i = conv1d(exp(K^i), exp(w^i) - 1) + sum(exp(K_j^i)) over j 1 to t
|
| 164 |
+
|
| 165 |
+
The convolution kernel provides local context while the running sum provides longer range accumulation. Transcendental functions, pooling, and residual adds are assigned to the digital companion.
|
| 166 |
+
|
| 167 |
+
### Feedforward block
|
| 168 |
+
|
| 169 |
+
An MLP layer has the form:
|
| 170 |
+
|
| 171 |
+
tanh(W x + b)
|
| 172 |
+
|
| 173 |
+
For sparse W, the layer is compiled as dout parallel tanh linear units placed across Z1 cores, with samples streamed between cores. Depth comes from composition of these sampled programs.
|
| 174 |
+
|
| 175 |
+
### Token path
|
| 176 |
+
|
| 177 |
+
For one decoding step across a four block stack, the publication traces:
|
| 178 |
+
|
| 179 |
+
FPGA embedding and input handling to Z1 sampled projections and normalization to FPGA pooling and residuals, repeated per block, followed by vocabulary readout and next token sampling. Z1 sampling and FPGA orchestration dominate modeled latency, while FPGA work dominates modeled energy.
|
| 180 |
+
|
| 181 |
+
## Training
|
| 182 |
+
|
| 183 |
+
### Data
|
| 184 |
+
|
| 185 |
+
* OpenWebText
|
| 186 |
+
* GPT-2 byte pair encoding for the Z1 matched sweep
|
| 187 |
+
* Byte tokenized OpenWebText for the general connectivity sweep with a standard decoder
|
| 188 |
+
* Final dense vocabulary to embedding matmul retained where applicable
|
| 189 |
+
|
| 190 |
+
### Setup
|
| 191 |
+
|
| 192 |
+
* JAX research implementation in `research/z1t` and related sparse transformer projects
|
| 193 |
+
* 4 bit weights and 4 incoming edges per output node for the Z1 matched architecture
|
| 194 |
+
* Fixed connectivity patterns rather than fixed percentage sparsity
|
| 195 |
+
* Sparsity therefore grows toward 100 percent as width grows, unlike conventional fixed percentage pruning
|
| 196 |
+
* Activation quantization is required for practical Z1 execution, but Figure 1 results in the publication do not include activation quantization
|
| 197 |
+
|
| 198 |
+
### Scaling studies
|
| 199 |
+
|
| 200 |
+
Two complementary directions are reported:
|
| 201 |
+
|
| 202 |
+
1. Z1 matched gated convolutional architecture sweep from about 2e14 to about 2e18 training FLOPs, with a log log frontier fit extended to about 9.5e19 FLOPs to match GPT-2 small loss near 3.4.
|
| 203 |
+
2. General sparsity study with a standard GPT-2 style decoder using RMS normalization and GELU feedforward blocks, varying connectivity c in 4, 16, 32, 64, 128 plus a dense baseline, sequence length 256, budgets from 3e14 to 1e18 FLOPs.
|
| 204 |
+
|
| 205 |
+
Key observations:
|
| 206 |
+
|
| 207 |
+
* Larger compute budgets reach lower validation loss at every connectivity.
|
| 208 |
+
* Iso FLOP curves show U shaped minima that move toward larger models with higher compute.
|
| 209 |
+
* Dense models appear more efficient per FLOP in these early results.
|
| 210 |
+
* Sparse Z1 operations are modeled as vastly more energy efficient per operation, so the full Z1 plus FPGA system can still be orders of magnitude more power efficient at matched performance.
|
| 211 |
+
|
| 212 |
+
## Evaluation
|
| 213 |
+
|
| 214 |
+
This is a research checkpoint without a conventional assistant benchmark suite in the release. The publication focuses on scaling behavior rather than task scores.
|
| 215 |
+
|
| 216 |
+
Reported reference points use the same body parameter convention, excluding the final embedding to vocabulary matmul:
|
| 217 |
+
|
| 218 |
+
* GPT-2 small: 85 million body parameters
|
| 219 |
+
* GPT-2 XL: 1476 million body parameters
|
| 220 |
+
* Z1T reference modeled block: 11.55 million body parameters in fp16 for H100 timing, with D 512, L 4
|
| 221 |
+
|
| 222 |
+
Users seeking task evaluation should report dataset, tokenizer handling, sampling settings, number of Z1 samples averaged, quantization, and whether the vocabulary readout is included, since readout dominates energy when included.
|
| 223 |
+
|
| 224 |
+
## Energy Cost Model
|
| 225 |
+
|
| 226 |
+
All figures below are modeled estimates based on theoretical Z1 chip energy anchored to experiments with similar pbits in X0. They exclude the dense final logit readout unless stated. They assume enough parallel Z1 chips to place all samples in a model parallel fashion. Required chip count is not modeled. Data movement between Z1 and FPGA is not included.
|
| 227 |
+
|
| 228 |
+
Reference modeled configuration: L 4, D 512, T 1024, GCA H 4, kernel 4, sparse fan in 4.
|
| 229 |
+
|
| 230 |
+
Assumptions:
|
| 231 |
+
|
| 232 |
+
* Z1 sampling: 1.3e-14 J per sample
|
| 233 |
+
* FPGA: 0.2 pJ per matrix multiply op, 3.0 pJ per scalar op, 1.5 W assumed static power
|
| 234 |
+
* H100 reference: 0.177 pJ per floating point op at 32 bit peak, same next token step run densely with no sparsity exploited
|
| 235 |
+
* H100 model FLOPs utilization varied because sparse models often achieve low utilization on GPUs
|
| 236 |
+
|
| 237 |
+
Modeled per token energy for the Z1T block without final readout:
|
| 238 |
+
|
| 239 |
+
* Total: 294.52 nJ per token
|
| 240 |
+
* Z1 sampling: 8.74 nJ
|
| 241 |
+
* Included FPGA work: 285.78 nJ
|
| 242 |
+
* Final logit readout on FPGA if included: approximately 136.4 uJ per token
|
| 243 |
+
|
| 244 |
+
Energy ratios, H100 relative to Z1T without final readout:
|
| 245 |
+
|
| 246 |
+
<table>
|
| 247 |
+
<thead>
|
| 248 |
+
<tr><th>H100 utilization</th><th>H100 energy per token</th><th>H100 to Z1T system</th><th>H100 to Z1 layers only</th></tr>
|
| 249 |
+
</thead>
|
| 250 |
+
<tbody>
|
| 251 |
+
<tr><td>10 percent</td><td>40.9 uJ</td><td>about 139x</td><td>about 4680x</td></tr>
|
| 252 |
+
<tr><td>50 percent</td><td>8.17 uJ</td><td>about 28x</td><td>about 935x</td></tr>
|
| 253 |
+
<tr><td>100 percent</td><td>4.09 uJ</td><td>about 14x</td><td>about 468x</td></tr>
|
| 254 |
+
</tbody>
|
| 255 |
+
</table>
|
| 256 |
+
|
| 257 |
+
Interpretation:
|
| 258 |
+
|
| 259 |
+
* At 10 percent utilization, representative of sparse workloads on GPUs, the modeled full system advantage is largest.
|
| 260 |
+
* At 100 percent utilization, the modeled full system advantage is smaller but still above one order of magnitude.
|
| 261 |
+
* Z1 only layers show much larger ratios because the FPGA accounts for more than 95 percent of modeled system energy.
|
| 262 |
+
* Including the vocabulary readout changes the absolute picture substantially and should always be disclosed.
|
| 263 |
+
|
| 264 |
+
## Latency Model
|
| 265 |
+
|
| 266 |
+
Throughput figures assume a single serial stream with no advanced pipelining. Parallel sampling would keep energy similar while reducing time to target precision.
|
| 267 |
+
|
| 268 |
+
Modeled serial path for one token, excluding final vocabulary logits:
|
| 269 |
+
|
| 270 |
+
* Total: 58.8 us per token, about 17000 tokens per s
|
| 271 |
+
* FPGA orchestration: 38.0 us, from 38 serial ops at 1 us
|
| 272 |
+
* Z1 sampling: 16.0 us, from 25 layers times 32 samples at 50 MHz
|
| 273 |
+
* Data reading: 3.2 us
|
| 274 |
+
* Data writing: 1.3 us
|
| 275 |
+
* FPGA sample averaging: 0.13 us
|
| 276 |
+
* FPGA pooling, positional, and residual work: 0.16 us
|
| 277 |
+
* Steps modeled: 25 sequential sampling layers, 8 FPGA to Z1 writes, 5 Z1 to FPGA reads
|
| 278 |
+
|
| 279 |
+
H100 baselines measured by Extropic on 2026-08-12, NVIDIA H100 80GB HBM3, torch 2.7.0, dense equivalent model, D 512, L 4, 11.55M body parameters in fp16, batch 1 sequential decoding, excluding final vocabulary logits:
|
| 280 |
+
|
| 281 |
+
<table>
|
| 282 |
+
<thead>
|
| 283 |
+
<tr><th>Operating point</th><th>Latency per token</th><th>Tokens per s</th></tr>
|
| 284 |
+
</thead>
|
| 285 |
+
<tbody>
|
| 286 |
+
<tr><td>Z1T conservative serial estimate</td><td>58.8 us</td><td>about 17000</td></tr>
|
| 287 |
+
<tr><td>H100 eager PyTorch</td><td>702 us</td><td>about 1425</td></tr>
|
| 288 |
+
<tr><td>H100 torch.compile</td><td>102 us</td><td>about 9764</td></tr>
|
| 289 |
+
</tbody>
|
| 290 |
+
</table>
|
| 291 |
+
|
| 292 |
+
Caveats:
|
| 293 |
+
|
| 294 |
+
* The test model is small relative to H100 scale, so kernel launch overhead is material. Reported H100 utilization for this case is 0.006 percent.
|
| 295 |
+
* Batching on GPU would improve GPU efficiency substantially.
|
| 296 |
+
* The authors recommend this Z1 plus XPU setup for decode rather than prefill.
|
| 297 |
+
|
| 298 |
+
## Usage
|
| 299 |
+
|
| 300 |
+
Install the research package from its own project directory. See `research/z1t` in the official repository for exact commands. Typical requirements include Python 3.11 or later, JAX, Equinox, `huggingface_hub`, and the local `z1t` package.
|
| 301 |
+
|
| 302 |
+
Minimal loading pattern, matching the released `load_model.py`:
|
| 303 |
+
|
| 304 |
+
```python
|
| 305 |
+
import json
|
| 306 |
+
import equinox as eqx
|
| 307 |
+
import jax
|
| 308 |
+
from huggingface_hub import hf_hub_download
|
| 309 |
+
from z1t.components import Config
|
| 310 |
+
from z1t.model import create_model
|
| 311 |
+
|
| 312 |
+
repo_id = "Extropic-AI/Z1T-0"
|
| 313 |
+
|
| 314 |
+
config_path = hf_hub_download(repo_id, "config.json")
|
| 315 |
+
weights_path = hf_hub_download(repo_id, "model.eqx")
|
| 316 |
+
|
| 317 |
+
with open(config_path) as f:
|
| 318 |
+
config = Config(**json.load(f))
|
| 319 |
+
|
| 320 |
+
model = create_model(config, jax.random.key(0))
|
| 321 |
+
model = eqx.tree_deserialise_leaves(weights_path, model)
|
| 322 |
+
```
|
| 323 |
+
|
| 324 |
+
Practical notes:
|
| 325 |
+
|
| 326 |
+
* The checkpoint is large. Ensure sufficient local storage and memory before download.
|
| 327 |
+
* Generation behavior depends on sampling configuration, tokenizer postprocessing, number of probabilistic samples averaged, and quantization settings.
|
| 328 |
+
* For fair comparison with the publication, state whether logits and sampling overhead are included.
|
| 329 |
+
* For training or fine tuning, use the official repository configs, experiments, and scripts rather than this inference only loading path.
|
| 330 |
+
|
| 331 |
+
## Limitations
|
| 332 |
+
|
| 333 |
+
* Preliminary codesign study. Z1 predates Z1T style ideas, so the mapping is unoptimized by design.
|
| 334 |
+
* System energy and latency are modeled, not end to end measured on deployed Z1 plus FPGA hardware for the full model.
|
| 335 |
+
* Parallel Z1 capacity is assumed. Physical chip count, yield, control overhead, and interconnect energy are not modeled.
|
| 336 |
+
* Activation quantization effects are not included in the main Z1 matched loss frontier.
|
| 337 |
+
* Sparse models require substantially deeper and wider structures to match dense parameter counts at fixed connectivity.
|
| 338 |
+
* OpenWebText has known quality, duplication, bias, and toxicity limitations. No additional safety tuning is documented for this checkpoint.
|
| 339 |
+
* English only validation is reported. Multilingual behavior is unknown.
|
| 340 |
+
* Future Z2 hardware may change the optimal sparsity, precision, and architecture choices.
|
| 341 |
+
|
| 342 |
+
## Safety, Bias, and Responsibility
|
| 343 |
+
|
| 344 |
+
No safety evaluation, red teaming, or bias mitigation is documented for Z1T-0. Language models trained on web text can reproduce harmful stereotypes, misinformation, sensitive content, and insecure code patterns.
|
| 345 |
+
|
| 346 |
+
Do not deploy in high stakes settings without domain evaluation and appropriate guardrails. Document data filtering, sampling temperature, refusal behavior if added, and any post training applied downstream.
|
| 347 |
+
|
| 348 |
+
## Citation
|
| 349 |
+
|
| 350 |
+
If you use Z1T-0, cite the official publication page and the training repository. The following BibTeX entry is suggested for this model card:
|
| 351 |
+
|
| 352 |
+
```bibtex
|
| 353 |
+
@misc{extropic2026z1t,
|
| 354 |
+
title = {Z1T: Sparse Transformer Like Models for Probabilistic Hardware},
|
| 355 |
+
author = {Verdon, Guillaume and Neagoe, Alexander and Lockwood, Owen and Morton, Seth and Extropic},
|
| 356 |
+
year = {2026},
|
| 357 |
+
month = {September},
|
| 358 |
+
publisher = {Extropic Corp},
|
| 359 |
+
howpublished = {Available at https://extropic.ai/writing/z1t/},
|
| 360 |
+
note = {Model weights: https://huggingface.co/Extropic-AI/Z1T-0; Code: https://github.com/extropic-ai/sparse-transformers}
|
| 361 |
+
}
|
| 362 |
+
```
|
| 363 |
+
|
| 364 |
+
## Selected References
|
| 365 |
+
|
| 366 |
+
* Vaswani et al. Attention Is All You Need. NeurIPS 30, 2017.
|
| 367 |
+
* Hoffmann et al. Training Compute Optimal Large Language Models. NeurIPS 35, 2022.
|
| 368 |
+
* Zhai et al. An Attention Free Transformer. arXiv 2105.14103, 2021.
|
| 369 |
+
* Zhang and Sennrich. Root Mean Square Layer Normalization. NeurIPS 32, 2019.
|
| 370 |
+
* Hendrycks and Gimpel. Gaussian Error Linear Units. arXiv 1606.08415, 2016.
|
| 371 |
+
* Zhu et al. Transformers Without Normalization. CVPR, 2025.
|
| 372 |
+
* Ackley, Hinton, and Sejnowski. A Learning Algorithm for Boltzmann Machines. Cognitive Science 9:147 to 169, 1985.
|
| 373 |
+
* Lockwood et al. A Blueprint for Equilibrium Based Differentiable Continuous Variable Thermodynamic Computing. arXiv 2607.16183, 2026.
|
| 374 |
+
* Jelincic et al. An Efficient Probabilistic Hardware Architecture for Diffusion Like Models. npj Unconventional Computing 3:30, 2026.
|
| 375 |
+
* Verdon et al. A Framework for Stochastic Differentiable Programming. arXiv 2608.01612, 2026.
|
| 376 |
+
* Amico et al. Thermalizing Stochastic Programs. arXiv 2608.01615, 2026.
|
| 377 |
+
* Extropic. From One to One Billion: Torx, Thermalizers, and Z1. 2026.
|
| 378 |
+
* Grattafiori et al. The Llama 3 Herd of Models. arXiv 2407.21783, 2024.
|
| 379 |
+
* Frantar et al. Scaling Laws for Sparsely Connected Foundation Models. arXiv 2309.08520, 2023.
|
| 380 |
+
* Gale et al. Sparse GPU Kernels for Deep Learning. SC20, 2020.
|
| 381 |
+
* Hooker. The Hardware Lottery. Communications of the ACM 64:58 to 65, 2021.
|
| 382 |
+
|
| 383 |
+
Full bibliography with 23 entries is available in the official publication.
|
| 384 |
+
|
| 385 |
+
## Contact
|
| 386 |
+
|
| 387 |
+
* Extropic: contact@extropic.ai
|
| 388 |
+
* Code issues: https://github.com/extropic-ai/sparse-transformers/issues
|
| 389 |
+
* Model discussion: use the Community tab on the Hugging Face repository
|
| 390 |
|
|
|