Text Generation
Safetensors
Rust
RWKV
English
oicio-rs
ternary
matmul-free
cpu-only
1.58-bit
bitnet
bonsai
infinite-context
em-llm
reattention
recursive-agent-harness
rlm
rah
edge-ai
needle
hadamard
mlgru
mamba
liquid-neural-networks
turbovec
turboquant
t-mac
vec-lut
axon
consumer-hardware
better-quality
intelligence-density
Instructions to use deeprcurs/OICIO with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- RWKV
How to use deeprcurs/OICIO with RWKV:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
File size: 3,115 Bytes
ce20bc6 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 | # OICIO Data β Training Data and Checkpoints
**Credits:** deepRcurs Labs, @deeprcurs
**Author:** Mzed Imamkh, @mzedimamkh
## Overview
This directory contains training data and checkpoints for OICIO, following snapshot rules: code in snapshot-safe (<128MB), toolchain, dependencies, and large artifacts in `.cache` (excluded, can be re-downloaded).
As per requirements: dataset and trainer is LLM itself (LLM as teacher, source of knowledge, dataset, and auditor).
## Dataset Generation β LLM as Teacher
No large datasets are downloaded to snapshot (would exceed 128MB limit). Synthetic data is generated on-the-fly in RAM with swap offloading if needed.
**Synthetic datasets:**
- **OOLONG Synthetic:** Generates entries with user_id and entity classification, 3 topics with 90% coherence and 10% switch (surprise event boundary), mimicking Oolong-Synthetic benchmark (199 samples, 13 buckets 1K-4M tokens, average 629K tokens)
- **LongBench-like:** Generates QA, summarization, code tasks across 6 categories (SQA, MQA, Sum, FSL, Ret, Cod)
- **InfiniteBench-like:** Generates PassKey retrieval with hidden passkey at random position, tested up to 1M tokens (102400 chunks β 7144 events)
All generated on-the-fly in 1.9GB RAM + 14GB swap, not stored permanently (snapshot-safe).
## Checkpoints
- `training_log_here.json` β Training log from scratch HERE: 6.8M ternary, 50 steps, 23.4s, loss 6.9488β6.9377 drop 0.0111, sparsity 31.1%β34.3%, FP16 13MB β Ternary 1.3MB (10.1x), swap 14GB active, consumer hardware only
- Real checkpoints (BitNet 2B 1.1GB, Bonsai 8B 1.75GB) stored in `/home/user/.cache/models` (excluded from snapshot, can re-download via `hf download`)
- Large checkpoints (e.g., `oicio_from_scratch_here.pt` 27MB, `ternary_san_qat.pt` 5MB) moved to `/home/user/.cache/oicio_checkpoints` (excluded) to keep snapshot clean (316KB β 510KB after cleanup)
## Usage
```python
from oicio.training.qat_trainer import SyntheticOOLONGDataset
dataset = SyntheticOOLONGDataset(num_samples=1000, seq_len=128)
from oicio.training.train_from_scratch_here import LLMasTeacherDataset
dataset = LLMasTeacherDataset(vocab_size=1024, seq_len=128, num_samples=10000)
# Generates synthetic with 3 topics, LLM as teacher
```
LLM is teacher: generates data, trains, audits, repeats.
## Storage β Free Tier Without Credit Card/Phone
- **HuggingFace Hub:** Public best-effort up to 5TB, private 100GB free, no credit card, no phone verification, just email. Already proven push of BitNet 2B 1.1GB real weights + training logs via HF token.
- **Cloudflare R2:** 10GB free forever, 1M write, 10M read, unlimited egress, no credit card required per tutorial, S3-compatible.
- **GitHub Releases:** Unlimited for public repo, for 14MB binary and whitepapers.
- **MyBinder.org:** No account needed, just GitHub repo public, VM 2GB RAM, auto-build.
## Snapshot Compliance
Code in `oicio/data/` is snapshot-safe: README.md 1.2KB + training_log_here.json 570 bytes = ~2KB.
Large artifacts (*.pt, *.safetensors) excluded via `.gitignore` and stored in `.cache` (excluded from snapshot).
|