Spaces:
Running
Running
Ares Roadmap and Reality Checks
Constraint checks
Static Hugging Face Space
- Good for: UI, documentation, browser-only demos, downloadable checkpoints, localStorage/IndexedDB memory.
- Not good for: server-side Python, GPU training, persistent server databases, private compute, live web search.
No external AI APIs for core intelligence
- Good principle. Ares training and inference code here is from-scratch.
- Consequence: no OpenAI/Claude/Gemini/etc. behind the chatbot. Capabilities come only from Ares weights, RAG data, and deterministic tools.
Billion-parameter training from scratch
- Possible in principle.
- Not realistic on free static hosting. Requires many billions to trillions of tokens, multi-GPU training, checkpoint sharding, monitoring, and data governance.
Petabyte-scale database
- A single SQLite database is not a petabyte system.
- Practical path: SQLite/FTS5 now, then sharded SQLite or DuckDB/Parquet, then distributed object storage + vector index if compute becomes available.
Actual intelligence vs. prediction
- A Transformer learns by next-token prediction during pretraining.
- More agentic behavior comes from post-training, retrieval, planning loops, tool use, evaluation, self-critique, and environment feedback. Do not claim consciousness or full AGI.
Development phases
Phase 0 — Static Space and honest prototype
- Static UI.
- Local memory/RAG in browser.
- Architecture dashboard.
- Training instructions.
Phase 1 — Tokenizer and data pipeline
- Train BPE tokenizer.
- Dataset cleaning and deduplication.
- Packing into fixed-length token blocks.
- Data mixture manifests.
Phase 2 — Tiny Transformer smoke tests
- 1M–10M parameters.
- CPU/GPU smoke training.
- Loss decreases on sample data.
- Generate text from checkpoint.
Phase 3 — Small useful model
- 30M–150M parameters.
- More curated text/code/chat data.
- SFT formatting.
- Basic eval suite.
Phase 4 — RAG and long-term memory
- SQLite/FTS5 store.
- Hashed-vector bootstrap retrieval.
- Later replace with Ares-trained embedding head.
- Evidence-grounded answer templates.
Phase 5 — Planning and tools
- Planner proposes steps.
- Sandbox/code actuator runs controlled commands.
- Web search side-goal only after base model exists.
- Static Space can show UI; tools need a backend elsewhere.
Phase 6 — Scaling toward 1B
- 1B target config.
- Multi-GPU FSDP/DeepSpeed or similar.
- Checkpoint sharding.
- Mixed precision.
- Activation checkpointing.
- Larger context and RoPE scaling.
Phase 7 — SFT and preference alignment
- Instruction/chat/code data.
- Supervised fine-tuning.
- Preference data collection.
- DPO/RLHF-style optimization.
Immediate next success criterion
Ares v0.1 is successful when:
- The Static Space loads.
- The browser prototype can store and retrieve local memory.
- The tokenizer trains on sample text.
- The tiny model trains for 100 steps without crashing in a Python GPU/CPU environment.
- Loss decreases on the sample corpus.