An integer execution method for reproducible inference from publicly available model weights, demonstrated on Qwen3-4B. Journaled bytes and all.
Keep an eye out for the gpt-oss-120B on the 24gb GPU- deterministically. We make AI models do the same things every time!βοΈπ i64systems/Qwen3-4B-openbob-i8
Weβre excited to release Pebble-25M and Pebble-25M-Chat!
Both models use our 3:1 Mamba2/Transformer hybrid architecture and were pretrained on 25B tokens. Pebble-25M-Chat was then further fine-tuned on an additional 250M tokens from smol-smoltalk, following the same approach used for the Pebble-10M models.
remat is no longer a one-model claim!!proved it on Qwen3-30B-A3B, K=32 of 128 experts resident, output task byte-identical to the full reference, zero bytes different *in bf16*π₯°π₯° GPU comes nextπ