Spaces:
Running
Running
| title: Ares Static Lab | |
| emoji: 🛡️ | |
| colorFrom: indigo | |
| colorTo: red | |
| sdk: static | |
| app_file: index.html | |
| pinned: false | |
| short_description: Static UI and from-scratch LLM training scaffold | |
| # Ares Static Lab | |
| Ares is a from-scratch AI chatbot / general-intelligence research project scaffold. This repository is designed for a **Hugging Face Static Space**, so the visible app is pure HTML/CSS/JavaScript and does not call external AI APIs. | |
| The static browser UI is a control panel and local prototype. The heavy model-training code lives in `ares_core/` and is meant to run on Kaggle, Colab, a rented GPU, or any Python GPU environment. A Static Space cannot run server-side training or inference. | |
| ## What is included now | |
| - Static Ares chat/control UI in `index.html`, including official **Trainable Ares** and **Dataset** tabs. | |
| - Colab-first training notebook in `colab/Ares_Colab_Trainer.ipynb`. | |
| - Complete brain-training notebook in `colab/Ares_Wikipedia_Roleplay_Brain_Training.ipynb` for Wikipedia + ML-process + standalone roleplay JSONL + optional roleplay SFT. | |
| - Downloadable starter datasets in `starter_datasets/`: 1,000 roleplay JSONL scenarios and 500 ML-process records. | |
| - Public-repo safety rules in `.gitignore`, `PUBLICATION_CHECKLIST.md`, and `colab/README_COLAB.md`. | |
| - Bounded dataset builder, train/validation splitter, Wikipedia/ML/roleplay mixture builder, validation-aware trainer, evaluator, and offline HTML training report generator. | |
| - From-scratch decoder-only Transformer skeleton: | |
| - trained BPE tokenizer pipeline, | |
| - token embeddings, | |
| - RoPE positional embeddings, | |
| - RMSNorm, | |
| - multi-head / grouped-query attention, | |
| - KV cache support, | |
| - SwiGLU MLP, | |
| - tied/unembedding output matrix, | |
| - cross-entropy loss, | |
| - AdamW + cosine schedule training loop. | |
| - SQLite/FTS5 RAG bootstrap store with a no-external-model hashed-vector retriever. | |
| - SFT data formatting utilities. | |
| - Scaling configs from tiny smoke tests to a 1B-class architecture target. | |
| ## What is intentionally not claimed yet | |
| This initial repository does **not** contain a trained billion-parameter Ares checkpoint. Training a capable 1B model from scratch requires large datasets, careful evaluation, and significant GPU time. The project starts with a small, reproducible model and scales iteratively. | |
| ## Official Hugging Face repos | |
| - Static Space: https://huggingface.co/spaces/jacmor64/ares-static-lab | |
| - Starter dataset: https://huggingface.co/datasets/jacmor64/ares-starter-dataset | |
| - Tiny trained seed model 1M: https://huggingface.co/jacmor64/ares-seed-1m | |
| - Increased-parameter seed model 6M: https://huggingface.co/jacmor64/ares-seed-6m | |
| - 30M-class seed model: https://huggingface.co/jacmor64/ares-seed-30m | |
| ## Serious non-toy training notebook | |
| Use this when training Ares seriously: | |
| ```text | |
| colab/OFFICIAL_Ares_Serious_Dataset_Training_Run.ipynb | |
| ``` | |
| Guide: | |
| ```text | |
| SERIOUS_TRAINING.md | |
| ``` | |
| This notebook builds a real multi-source corpus: Wikipedia EN, FineWeb-Edu, OpenWebMath, Python Code Dataset, UltraChat, Ares ML-process curriculum, and 50,000 roleplay scenarios by default. The starter dataset is only a preview/smoke artifact. | |
| ## Official real-generation Colab notebook | |
| Use this first for the clearest Ares training loop: | |
| ```text | |
| colab/OFFICIAL_Ares_Colab_Training_Run.ipynb | |
| ``` | |
| Plain-English guide: | |
| ```text | |
| COLAB_OFFICIAL_RUN.md | |
| ``` | |
| This notebook uses `wikimedia/wikipedia`, optional `roneneldan/TinyStories`, Ares ML-process curriculum, thousands of roleplay scenarios, tokenizer training, base training, optional roleplay SFT, and real generation from the final checkpoint. | |
| ## Colab start | |
| For the original smoke pipeline: | |
| 1. Open `colab/Ares_Colab_Trainer.ipynb` in Google Colab. | |
| 2. Select **Runtime → Change runtime type → GPU**. | |
| 3. Run the notebook top-to-bottom. | |
| For the requested Ares brain training on Wikipedia, ML processes, and thousands of roleplay scenarios: | |
| 1. Open `colab/Ares_Wikipedia_Roleplay_Brain_Training.ipynb` in Google Colab. | |
| 2. Keep the first run on `20231101.simple` Wikipedia and `configs/ares_30m.json` unless your GPU is small, in which case use `configs/ares_8m.json`. | |
| 3. Run the notebook to build the mixture, train the tokenizer, train Ares weights, evaluate validation loss/perplexity, and generate responses from the trained checkpoint. | |
| 4. Keep artifacts in Google Drive unless you intentionally want to publish them. | |
| ## Quick local/Python start | |
| ```bash | |
| python -m venv .venv | |
| source .venv/bin/activate | |
| pip install -r requirements.txt | |
| # 1) Train a small tokenizer on the sample corpus | |
| python -m ares_core.tokenizer_train \ | |
| --input data/sample_corpus.txt \ | |
| --output artifacts/tokenizer.json \ | |
| --vocab-size 2048 | |
| # 2) Smoke-train a tiny model | |
| python -m ares_core.train \ | |
| --config configs/ares_8m.json \ | |
| --tokenizer artifacts/tokenizer.json \ | |
| --train data/sample_corpus.txt \ | |
| --out artifacts/ares_8m \ | |
| --steps 100 \ | |
| --batch-size 4 \ | |
| --device auto | |
| # 3) Generate from the tiny checkpoint | |
| python -m ares_core.generate \ | |
| --checkpoint artifacts/ares_8m/ckpt_last.pt \ | |
| --tokenizer artifacts/tokenizer.json \ | |
| --prompt "Ares is" \ | |
| --max-new-tokens 64 | |
| ``` | |
| ## Hugging Face Static Space deploy | |
| 1. Create a new Space on Hugging Face under `jacmor64`. | |
| 2. Choose **Static** SDK. | |
| 3. Upload/commit every file in this folder. | |
| 4. The README YAML block above must remain at the very top so the Space serves `index.html`. | |
| ## Roadmap | |
| See `ROADMAP.md` for the staged build plan and constraint analysis. | |