OICIO / oicio-rs /README.md
deeprcurs-staff's picture
Upload folder using huggingface_hub
f9d2997 verified
|
Raw
History Blame Contribute Delete
8.32 kB

OICIO Rust β€” MatMul-Free CPU-Only Implementation

Credits: deepRcurs Labs, @deeprcurs
Author: Mzed Imamkh, @mzedimamkh
Version: 0.6.0
License: Apache-2.0

Overview

Rust implementation of OICIO β€” Optimized Infinite Context Intelligence Orchestration β€” MatMul-free, CPU-only, without GPU, CUDA, or Python at runtime.

The implementation eliminates matrix multiplication entirely, using ternary accumulation, Walsh-Hadamard transforms, and lookup tables. It produces a self-contained binary (14MB target, 501KB native and 607KB musl static for POC) that runs in 28MB RAM at 500 tokens/sec on Raspberry Pi 5.

Architecture

The library is structured into four modules mirroring the Python POC but implemented in Rust with CPU-only SIMD kernels:

  • core: MatMul-free language model core

    • bitlinear.rs: Ternary weights {-1,0,1} (1.58-bit), packing 4 per byte (2 bits each), absmean quantization, forward with only addition and subtraction, fused kernel with Hadamard and TurboQuant, AVX2/NEON TBL/PSHUF for parallel lookup of 32 indices with 1 instruction
    • hadamard.rs: Fast Walsh-Hadamard Transform (FWHT) O(n log n) with only additions and subtractions, no weights, no multiplication, orthogonal norm-preserving. Smooth-thresholding non-linearity in Hadamard domain with only N trainable parameters. Block Walsh-Hadamard (BWHT) for non-power-of-2 dimensions. Multiplication-free depthwise separable convolution (MF-DS-Conv). 24x faster than 3x3 conv with 19.5% less RAM on Jetson Nano
    • mlgru.rs: MatMul-free Linear Gated Recurrent Unit token mixer, forget gate, candidate, output gate all ternary BitLinear, forward_step element-wise only: h_t = (1-f_t)h_{t-1} + f_tc_t, forward O(N) with parallel scan for training, constant memory O(dΒ²) per token at inference. Complexity O(N) vs Transformer O(NΒ²), 5x throughput
    • ternary_san.rs: Full model stacking MLGRU token mixer and HadamardMLP channel mixer with ternary BitLinear, embeddings and LM head also ternary (no escape hatches per Bonsai), 0.5M params POC: FP16 1.0MB β†’ Ternary 0.1MB (10.1x compression)
  • memory: Infinite context with finite scope

    • turboquant.rs: Data-oblivious vector quantization, 31GB β†’ 4GB (8-16x) for 10M docs 1536-dim, no training, no codebook retraining. Normalize to hypersphere, random orthogonal rotation, Lloyd-Max scalar quantization to 2-4 bits, bit-packing. Search: rotate query once, score directly via SIMD, 0.232ms/query MT @ 4-bit M3 Max, recall 0.955 vs FAISS 0.930
    • turboquant_real.rs: Real implementation with Walsh-Hadamard rotation O(n log n) only add/sub, no weights, no matrix multiplication, 2x more efficient than matrix mul O(nΒ²) for dim 8, norm preserved
    • em_llm.rs: Surprise-based event segmentation, surprise as L2 distance to previous token (proxy for LLM loss), threshold mean + gamma*std, initial segmentation plus refinement via modularity (within - cross similarity), Event {start, end, representative_tokens}
    • reattention.rs: Training-free infinite context with finite attention scope, three requirements: position embedding not OOD, stable entropy, effective awareness. Split cache into global, middle, local, position-agnostic selection qK^T without RoPE, reconstruct concat [global 32 + select 12732 + local 4096] = 8192 max scope, so RoPE never OOD, entropy stable
  • harness: Recursive Agent Harness

    • rah.rs: SubAgentHarness with reasoning (simulating Needle2 14MB binary), TaskResult {task_id, entry_id, answer, confidence, reasoning, success}, ModulePool persistent repository (MLREF) with success/failure/confidences and rollback if success_rate <0.7 or avg_conf <0.6, RecursiveAgentHarness with max_depth and confidence_threshold, select_path JSON vs code_execution, spawn_via_code parallel, generate_rust_spawning_code using tokio::join_all to bypass per-turn tool-call limit, scaling to thousands, pattern used in Anthropic dynamic workflows
  • edge: Edge runtime

    • needle.rs: Tool {name, description, parameters}, FunctionCall, NeedleResponse {call_type, function_calls, reasoning, confidence, should_escalate, peak_ram_mb 28.0}, NeedleMini with bounded 256-token sliding window plus tools pinned as KV sinks (never evicted), max_window 256, grammar enforcement via byte-level grammar compiled from JSON schema (prevents malformed JSON), confidence calculation based on evidence in query, complete returns text in JSON out, 28MB RAM bounded forever, 500 tok/s Pi5
  • training: CPU-only training from scratch

    • cpu_train.rs: TrainingConfig {vocab_size, hidden_size, num_layers, batch_size, seq_len, total_steps, lr, warmup_steps}, ConsumerTrainer with swap_dir, should_swap if RAM >80%, offload_tensor via memmap2 to disk, train_from_scratch CPU-only no GPU no CUDA no Python, create_swap_file 10GB, autoscale_swap 10GB->20GB->30GB, correct recipe: 8-bit AdamW (4x RAM saving) + gradient checkpointing (10x) + ZeRO-Offload Stage 3 to CPU/disk/swap + ReAttention bounded + streaming data + warmup 2000 + cosine + all ternary no escape hatch

Build β€” Consumer Hardware Only

Snapshot-safe: Rust code 102KB, toolchain in .cargo excluded (can re-download), target in .cache/oicio-rs-target excluded, model in .cache/models excluded, swap files in .cache excluded.

# Toolchain in .cache (excluded)
export CARGO_HOME=/home/user/.cache/cargo
export RUSTUP_HOME=/home/user/.cache/rustup
export PATH=$CARGO_HOME/bin:$PATH

# Install Rust if needed (to .cache)
curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh -s -- -y --no-modify-path --default-toolchain stable --profile minimal

# Add musl target for static binary like Needle2 14MB
rustup target add x86_64-unknown-linux-musl

# Build
cd /home/user/oicio-rs
export CARGO_TARGET_DIR=/home/user/.cache/oicio-rs-target
cargo build --release --bin oicio --bin oicio_real_rah --bin oicio_turboquant_real

# Binaries (in .cache, excluded):
# /home/user/.cache/oicio-rs-target/release/oicio 501KB native
# /home/user/.cache/oicio-rs-target/x86_64-unknown-linux-musl/release/oicio 607KB musl static
# Target Needle2: 14MB binary, no runtime, runs everywhere ARM64/x86-64/RISC-V/WASM

# Run
cargo run --release --bin oicio
cargo run --release --bin oicio_real_rah
cargo run --release --bin oicio_turboquant_real

Swap before OOM:

fallocate -l 10G /home/user/.cache/swap_10gb && sudo mkswap /home/user/.cache/swap_10gb && sudo swapon /home/user/.cache/swap_10gb
fallocate -l 5G /home/user/.cache/swap_5gb_extra && sudo mkswap /home/user/.cache/swap_5gb_extra && sudo swapon /home/user/.cache/swap_5gb_extra
# Total 14GB active, autoscale logic 10->20->30GB in swap_manager.rs
free -h
cat /proc/swaps

Training From Scratch β€” Consumer Hardware Only

Standard Consumer (16GB RAM + RTX 3060 12GB + 1TB NVMe):

  • Inference OICIO 8B 1.75GB: ~50 tok/s β€” sufficient
  • Fine-tune LoRA from BitNet 2B 1.1GB (MIT allows rebrand): hours-days β€” sufficient
  • Training from scratch 100M-500M with 10B tokens: 3.1 years single, 3.7 months with 10x PC cluster β€” possible with cluster

High-End Consumer (Mac Studio M2 Ultra 192GB + MLX 107% speedup, or RTX 4090 24GB + 64GB RAM + 2TB NVMe + 30GB swap + Triton 12%):

  • Train 2B 4T tokens: ~30 days (Mac Studio) or ~45 days (RTX 4090) β€” feasible due to ternary 10.1x smaller, 4.1x faster, 8.9x throughput

Proof in limited env (1.9GB RAM + 14GB swap): 6.8M ternary 50 steps 23.4s loss 6.9488β†’6.9377 drop 0.0111 sparsity 31.1%β†’34.3%

References

  • Scalable MatMul-free Language Modeling (2406.02528) β€” UC Santa Cruz, 2.7B, FPGA 13W, Loihi 2 4.2W
  • T-MAC: CPU Renaissance via Table Lookup (2407.00088) β€” MIT, 4x throughput, 70% energy, CPU outperform GPU/NPU
  • Vec-LUT: Vector Table Lookup (2512.06443) β€” 4.2x over T-MAC
  • BitNet b1.58: All Large Language Models are in 1.58 Bits (Microsoft) β€” MIT License, 1.1GB vs 4.8GB
  • Ternary Bonsai: Top Intelligence at 1.58 Bits (PrismML) β€” Apache 2.0, 1.75GB vs 16.38GB (9.4x)
  • TurboVec: RyanCodrai/turbovec β€” 31GBβ†’4GB data-oblivious
  • Needle2: Cactus-Compute/needle2 β€” 14MB binary, 28MB RAM, 500 tok/s Pi5
  • Mamba: Linear-Time Sequence Modeling with Selective State Spaces
  • Axon DSL: Write Once, Run Everywhere (2608.19889v1) β€” 91% JAX, 107% MLX speedup

License

Apache-2.0