Instructions to use deeprcurs/OICIO with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- RWKV
How to use deeprcurs/OICIO with RWKV:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
OICIO Rust β MatMul-Free CPU-Only Implementation
Credits: deepRcurs Labs, @deeprcurs
Author: Mzed Imamkh, @mzedimamkh
Version: 0.6.0
License: Apache-2.0
Overview
Rust implementation of OICIO β Optimized Infinite Context Intelligence Orchestration β MatMul-free, CPU-only, without GPU, CUDA, or Python at runtime.
The implementation eliminates matrix multiplication entirely, using ternary accumulation, Walsh-Hadamard transforms, and lookup tables. It produces a self-contained binary (14MB target, 501KB native and 607KB musl static for POC) that runs in 28MB RAM at 500 tokens/sec on Raspberry Pi 5.
Architecture
The library is structured into four modules mirroring the Python POC but implemented in Rust with CPU-only SIMD kernels:
core: MatMul-free language model core
bitlinear.rs: Ternary weights {-1,0,1} (1.58-bit), packing 4 per byte (2 bits each), absmean quantization, forward with only addition and subtraction, fused kernel with Hadamard and TurboQuant, AVX2/NEON TBL/PSHUF for parallel lookup of 32 indices with 1 instructionhadamard.rs: Fast Walsh-Hadamard Transform (FWHT) O(n log n) with only additions and subtractions, no weights, no multiplication, orthogonal norm-preserving. Smooth-thresholding non-linearity in Hadamard domain with only N trainable parameters. Block Walsh-Hadamard (BWHT) for non-power-of-2 dimensions. Multiplication-free depthwise separable convolution (MF-DS-Conv). 24x faster than 3x3 conv with 19.5% less RAM on Jetson Nanomlgru.rs: MatMul-free Linear Gated Recurrent Unit token mixer, forget gate, candidate, output gate all ternary BitLinear, forward_step element-wise only: h_t = (1-f_t)h_{t-1} + f_tc_t, forward O(N) with parallel scan for training, constant memory O(dΒ²) per token at inference. Complexity O(N) vs Transformer O(NΒ²), 5x throughputternary_san.rs: Full model stacking MLGRU token mixer and HadamardMLP channel mixer with ternary BitLinear, embeddings and LM head also ternary (no escape hatches per Bonsai), 0.5M params POC: FP16 1.0MB β Ternary 0.1MB (10.1x compression)
memory: Infinite context with finite scope
turboquant.rs: Data-oblivious vector quantization, 31GB β 4GB (8-16x) for 10M docs 1536-dim, no training, no codebook retraining. Normalize to hypersphere, random orthogonal rotation, Lloyd-Max scalar quantization to 2-4 bits, bit-packing. Search: rotate query once, score directly via SIMD, 0.232ms/query MT @ 4-bit M3 Max, recall 0.955 vs FAISS 0.930turboquant_real.rs: Real implementation with Walsh-Hadamard rotation O(n log n) only add/sub, no weights, no matrix multiplication, 2x more efficient than matrix mul O(nΒ²) for dim 8, norm preservedem_llm.rs: Surprise-based event segmentation, surprise as L2 distance to previous token (proxy for LLM loss), threshold mean + gamma*std, initial segmentation plus refinement via modularity (within - cross similarity), Event {start, end, representative_tokens}reattention.rs: Training-free infinite context with finite attention scope, three requirements: position embedding not OOD, stable entropy, effective awareness. Split cache into global, middle, local, position-agnostic selection qK^T without RoPE, reconstruct concat [global 32 + select 12732 + local 4096] = 8192 max scope, so RoPE never OOD, entropy stable
harness: Recursive Agent Harness
rah.rs: SubAgentHarness with reasoning (simulating Needle2 14MB binary), TaskResult {task_id, entry_id, answer, confidence, reasoning, success}, ModulePool persistent repository (MLREF) with success/failure/confidences and rollback if success_rate <0.7 or avg_conf <0.6, RecursiveAgentHarness with max_depth and confidence_threshold, select_path JSON vs code_execution, spawn_via_code parallel, generate_rust_spawning_code using tokio::join_all to bypass per-turn tool-call limit, scaling to thousands, pattern used in Anthropic dynamic workflows
edge: Edge runtime
needle.rs: Tool {name, description, parameters}, FunctionCall, NeedleResponse {call_type, function_calls, reasoning, confidence, should_escalate, peak_ram_mb 28.0}, NeedleMini with bounded 256-token sliding window plus tools pinned as KV sinks (never evicted), max_window 256, grammar enforcement via byte-level grammar compiled from JSON schema (prevents malformed JSON), confidence calculation based on evidence in query, complete returns text in JSON out, 28MB RAM bounded forever, 500 tok/s Pi5
training: CPU-only training from scratch
cpu_train.rs: TrainingConfig {vocab_size, hidden_size, num_layers, batch_size, seq_len, total_steps, lr, warmup_steps}, ConsumerTrainer with swap_dir, should_swap if RAM >80%, offload_tensor via memmap2 to disk, train_from_scratch CPU-only no GPU no CUDA no Python, create_swap_file 10GB, autoscale_swap 10GB->20GB->30GB, correct recipe: 8-bit AdamW (4x RAM saving) + gradient checkpointing (10x) + ZeRO-Offload Stage 3 to CPU/disk/swap + ReAttention bounded + streaming data + warmup 2000 + cosine + all ternary no escape hatch
Build β Consumer Hardware Only
Snapshot-safe: Rust code 102KB, toolchain in .cargo excluded (can re-download), target in .cache/oicio-rs-target excluded, model in .cache/models excluded, swap files in .cache excluded.
# Toolchain in .cache (excluded)
export CARGO_HOME=/home/user/.cache/cargo
export RUSTUP_HOME=/home/user/.cache/rustup
export PATH=$CARGO_HOME/bin:$PATH
# Install Rust if needed (to .cache)
curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh -s -- -y --no-modify-path --default-toolchain stable --profile minimal
# Add musl target for static binary like Needle2 14MB
rustup target add x86_64-unknown-linux-musl
# Build
cd /home/user/oicio-rs
export CARGO_TARGET_DIR=/home/user/.cache/oicio-rs-target
cargo build --release --bin oicio --bin oicio_real_rah --bin oicio_turboquant_real
# Binaries (in .cache, excluded):
# /home/user/.cache/oicio-rs-target/release/oicio 501KB native
# /home/user/.cache/oicio-rs-target/x86_64-unknown-linux-musl/release/oicio 607KB musl static
# Target Needle2: 14MB binary, no runtime, runs everywhere ARM64/x86-64/RISC-V/WASM
# Run
cargo run --release --bin oicio
cargo run --release --bin oicio_real_rah
cargo run --release --bin oicio_turboquant_real
Swap before OOM:
fallocate -l 10G /home/user/.cache/swap_10gb && sudo mkswap /home/user/.cache/swap_10gb && sudo swapon /home/user/.cache/swap_10gb
fallocate -l 5G /home/user/.cache/swap_5gb_extra && sudo mkswap /home/user/.cache/swap_5gb_extra && sudo swapon /home/user/.cache/swap_5gb_extra
# Total 14GB active, autoscale logic 10->20->30GB in swap_manager.rs
free -h
cat /proc/swaps
Training From Scratch β Consumer Hardware Only
Standard Consumer (16GB RAM + RTX 3060 12GB + 1TB NVMe):
- Inference OICIO 8B 1.75GB: ~50 tok/s β sufficient
- Fine-tune LoRA from BitNet 2B 1.1GB (MIT allows rebrand): hours-days β sufficient
- Training from scratch 100M-500M with 10B tokens: 3.1 years single, 3.7 months with 10x PC cluster β possible with cluster
High-End Consumer (Mac Studio M2 Ultra 192GB + MLX 107% speedup, or RTX 4090 24GB + 64GB RAM + 2TB NVMe + 30GB swap + Triton 12%):
- Train 2B 4T tokens: ~30 days (Mac Studio) or ~45 days (RTX 4090) β feasible due to ternary 10.1x smaller, 4.1x faster, 8.9x throughput
Proof in limited env (1.9GB RAM + 14GB swap): 6.8M ternary 50 steps 23.4s loss 6.9488β6.9377 drop 0.0111 sparsity 31.1%β34.3%
References
- Scalable MatMul-free Language Modeling (2406.02528) β UC Santa Cruz, 2.7B, FPGA 13W, Loihi 2 4.2W
- T-MAC: CPU Renaissance via Table Lookup (2407.00088) β MIT, 4x throughput, 70% energy, CPU outperform GPU/NPU
- Vec-LUT: Vector Table Lookup (2512.06443) β 4.2x over T-MAC
- BitNet b1.58: All Large Language Models are in 1.58 Bits (Microsoft) β MIT License, 1.1GB vs 4.8GB
- Ternary Bonsai: Top Intelligence at 1.58 Bits (PrismML) β Apache 2.0, 1.75GB vs 16.38GB (9.4x)
- TurboVec: RyanCodrai/turbovec β 31GBβ4GB data-oblivious
- Needle2: Cactus-Compute/needle2 β 14MB binary, 28MB RAM, 500 tok/s Pi5
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- Axon DSL: Write Once, Run Everywhere (2608.19889v1) β 91% JAX, 107% MLX speedup
License
Apache-2.0