Text Generation
Safetensors
Rust
RWKV
English
oicio-rs
ternary
matmul-free
cpu-only
1.58-bit
bitnet
bonsai
infinite-context
em-llm
reattention
recursive-agent-harness
rlm
rah
edge-ai
needle
hadamard
mlgru
mamba
liquid-neural-networks
turbovec
turboquant
t-mac
vec-lut
axon
consumer-hardware
better-quality
intelligence-density
Instructions to use deeprcurs/OICIO with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- RWKV
How to use deeprcurs/OICIO with RWKV:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
| # OICIO Rust β MatMul-Free CPU-Only Implementation | |
| **Credits:** deepRcurs Labs, @deeprcurs | |
| **Author:** Mzed Imamkh, @mzedimamkh | |
| **Version:** 0.6.0 | |
| **License:** Apache-2.0 | |
| ## Overview | |
| Rust implementation of OICIO β Optimized Infinite Context Intelligence Orchestration β MatMul-free, CPU-only, without GPU, CUDA, or Python at runtime. | |
| The implementation eliminates matrix multiplication entirely, using ternary accumulation, Walsh-Hadamard transforms, and lookup tables. It produces a self-contained binary (14MB target, 501KB native and 607KB musl static for POC) that runs in 28MB RAM at 500 tokens/sec on Raspberry Pi 5. | |
| ## Architecture | |
| The library is structured into four modules mirroring the Python POC but implemented in Rust with CPU-only SIMD kernels: | |
| - **core:** MatMul-free language model core | |
| - `bitlinear.rs`: Ternary weights {-1,0,1} (1.58-bit), packing 4 per byte (2 bits each), absmean quantization, forward with only addition and subtraction, fused kernel with Hadamard and TurboQuant, AVX2/NEON TBL/PSHUF for parallel lookup of 32 indices with 1 instruction | |
| - `hadamard.rs`: Fast Walsh-Hadamard Transform (FWHT) O(n log n) with only additions and subtractions, no weights, no multiplication, orthogonal norm-preserving. Smooth-thresholding non-linearity in Hadamard domain with only N trainable parameters. Block Walsh-Hadamard (BWHT) for non-power-of-2 dimensions. Multiplication-free depthwise separable convolution (MF-DS-Conv). 24x faster than 3x3 conv with 19.5% less RAM on Jetson Nano | |
| - `mlgru.rs`: MatMul-free Linear Gated Recurrent Unit token mixer, forget gate, candidate, output gate all ternary BitLinear, forward_step element-wise only: h_t = (1-f_t)*h_{t-1} + f_t*c_t, forward O(N) with parallel scan for training, constant memory O(dΒ²) per token at inference. Complexity O(N) vs Transformer O(NΒ²), 5x throughput | |
| - `ternary_san.rs`: Full model stacking MLGRU token mixer and HadamardMLP channel mixer with ternary BitLinear, embeddings and LM head also ternary (no escape hatches per Bonsai), 0.5M params POC: FP16 1.0MB β Ternary 0.1MB (10.1x compression) | |
| - **memory:** Infinite context with finite scope | |
| - `turboquant.rs`: Data-oblivious vector quantization, 31GB β 4GB (8-16x) for 10M docs 1536-dim, no training, no codebook retraining. Normalize to hypersphere, random orthogonal rotation, Lloyd-Max scalar quantization to 2-4 bits, bit-packing. Search: rotate query once, score directly via SIMD, 0.232ms/query MT @ 4-bit M3 Max, recall 0.955 vs FAISS 0.930 | |
| - `turboquant_real.rs`: Real implementation with Walsh-Hadamard rotation O(n log n) only add/sub, no weights, no matrix multiplication, 2x more efficient than matrix mul O(nΒ²) for dim 8, norm preserved | |
| - `em_llm.rs`: Surprise-based event segmentation, surprise as L2 distance to previous token (proxy for LLM loss), threshold mean + gamma*std, initial segmentation plus refinement via modularity (within - cross similarity), Event {start, end, representative_tokens} | |
| - `reattention.rs`: Training-free infinite context with finite attention scope, three requirements: position embedding not OOD, stable entropy, effective awareness. Split cache into global, middle, local, position-agnostic selection q*K^T without RoPE, reconstruct concat [global 32 + select 127*32 + local 4096] = 8192 max scope, so RoPE never OOD, entropy stable | |
| - **harness:** Recursive Agent Harness | |
| - `rah.rs`: SubAgentHarness with reasoning (simulating Needle2 14MB binary), TaskResult {task_id, entry_id, answer, confidence, reasoning, success}, ModulePool persistent repository (MLREF) with success/failure/confidences and rollback if success_rate <0.7 or avg_conf <0.6, RecursiveAgentHarness with max_depth and confidence_threshold, select_path JSON vs code_execution, spawn_via_code parallel, generate_rust_spawning_code using tokio::join_all to bypass per-turn tool-call limit, scaling to thousands, pattern used in Anthropic dynamic workflows | |
| - **edge:** Edge runtime | |
| - `needle.rs`: Tool {name, description, parameters}, FunctionCall, NeedleResponse {call_type, function_calls, reasoning, confidence, should_escalate, peak_ram_mb 28.0}, NeedleMini with bounded 256-token sliding window plus tools pinned as KV sinks (never evicted), max_window 256, grammar enforcement via byte-level grammar compiled from JSON schema (prevents malformed JSON), confidence calculation based on evidence in query, complete returns text in JSON out, 28MB RAM bounded forever, 500 tok/s Pi5 | |
| - **training:** CPU-only training from scratch | |
| - `cpu_train.rs`: TrainingConfig {vocab_size, hidden_size, num_layers, batch_size, seq_len, total_steps, lr, warmup_steps}, ConsumerTrainer with swap_dir, should_swap if RAM >80%, offload_tensor via memmap2 to disk, train_from_scratch CPU-only no GPU no CUDA no Python, create_swap_file 10GB, autoscale_swap 10GB->20GB->30GB, correct recipe: 8-bit AdamW (4x RAM saving) + gradient checkpointing (10x) + ZeRO-Offload Stage 3 to CPU/disk/swap + ReAttention bounded + streaming data + warmup 2000 + cosine + all ternary no escape hatch | |
| ## Build β Consumer Hardware Only | |
| Snapshot-safe: Rust code 102KB, toolchain in `.cargo` excluded (can re-download), target in `.cache/oicio-rs-target` excluded, model in `.cache/models` excluded, swap files in `.cache` excluded. | |
| ```bash | |
| # Toolchain in .cache (excluded) | |
| export CARGO_HOME=/home/user/.cache/cargo | |
| export RUSTUP_HOME=/home/user/.cache/rustup | |
| export PATH=$CARGO_HOME/bin:$PATH | |
| # Install Rust if needed (to .cache) | |
| curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh -s -- -y --no-modify-path --default-toolchain stable --profile minimal | |
| # Add musl target for static binary like Needle2 14MB | |
| rustup target add x86_64-unknown-linux-musl | |
| # Build | |
| cd /home/user/oicio-rs | |
| export CARGO_TARGET_DIR=/home/user/.cache/oicio-rs-target | |
| cargo build --release --bin oicio --bin oicio_real_rah --bin oicio_turboquant_real | |
| # Binaries (in .cache, excluded): | |
| # /home/user/.cache/oicio-rs-target/release/oicio 501KB native | |
| # /home/user/.cache/oicio-rs-target/x86_64-unknown-linux-musl/release/oicio 607KB musl static | |
| # Target Needle2: 14MB binary, no runtime, runs everywhere ARM64/x86-64/RISC-V/WASM | |
| # Run | |
| cargo run --release --bin oicio | |
| cargo run --release --bin oicio_real_rah | |
| cargo run --release --bin oicio_turboquant_real | |
| ``` | |
| **Swap before OOM:** | |
| ```bash | |
| fallocate -l 10G /home/user/.cache/swap_10gb && sudo mkswap /home/user/.cache/swap_10gb && sudo swapon /home/user/.cache/swap_10gb | |
| fallocate -l 5G /home/user/.cache/swap_5gb_extra && sudo mkswap /home/user/.cache/swap_5gb_extra && sudo swapon /home/user/.cache/swap_5gb_extra | |
| # Total 14GB active, autoscale logic 10->20->30GB in swap_manager.rs | |
| free -h | |
| cat /proc/swaps | |
| ``` | |
| ## Training From Scratch β Consumer Hardware Only | |
| Standard Consumer (16GB RAM + RTX 3060 12GB + 1TB NVMe): | |
| - Inference OICIO 8B 1.75GB: ~50 tok/s β sufficient | |
| - Fine-tune LoRA from BitNet 2B 1.1GB (MIT allows rebrand): hours-days β sufficient | |
| - Training from scratch 100M-500M with 10B tokens: 3.1 years single, 3.7 months with 10x PC cluster β possible with cluster | |
| High-End Consumer (Mac Studio M2 Ultra 192GB + MLX 107% speedup, or RTX 4090 24GB + 64GB RAM + 2TB NVMe + 30GB swap + Triton 12%): | |
| - Train 2B 4T tokens: ~30 days (Mac Studio) or ~45 days (RTX 4090) β feasible due to ternary 10.1x smaller, 4.1x faster, 8.9x throughput | |
| Proof in limited env (1.9GB RAM + 14GB swap): 6.8M ternary 50 steps 23.4s loss 6.9488β6.9377 drop 0.0111 sparsity 31.1%β34.3% | |
| ## References | |
| - Scalable MatMul-free Language Modeling (2406.02528) β UC Santa Cruz, 2.7B, FPGA 13W, Loihi 2 4.2W | |
| - T-MAC: CPU Renaissance via Table Lookup (2407.00088) β MIT, 4x throughput, 70% energy, CPU outperform GPU/NPU | |
| - Vec-LUT: Vector Table Lookup (2512.06443) β 4.2x over T-MAC | |
| - BitNet b1.58: All Large Language Models are in 1.58 Bits (Microsoft) β MIT License, 1.1GB vs 4.8GB | |
| - Ternary Bonsai: Top Intelligence at 1.58 Bits (PrismML) β Apache 2.0, 1.75GB vs 16.38GB (9.4x) | |
| - TurboVec: RyanCodrai/turbovec β 31GBβ4GB data-oblivious | |
| - Needle2: Cactus-Compute/needle2 β 14MB binary, 28MB RAM, 500 tok/s Pi5 | |
| - Mamba: Linear-Time Sequence Modeling with Selective State Spaces | |
| - Axon DSL: Write Once, Run Everywhere (2608.19889v1) β 91% JAX, 107% MLX speedup | |
| ## License | |
| Apache-2.0 | |