NEXORA / docs /STATUS.md
devildasdf's picture
Release validated NEXORA research prototype, tiny weights and evidence
12496fc verified
|
Raw History Blame Contribute Delete
5.35 kB
# Implementation and validation status
The mission is much larger than one workstation can train or certify. This release is a completed **local prototype milestone**, not fulfillment of all production acceptance criteria. No unimplemented service is represented by a mock production endpoint.
Final validation: **36 passed, 1 skipped**. The skip is a native Windows symlink-creation privilege limitation. Real Qwen3.5 agent follow-up after capability filtering still failed by repeated reads; the loop guard terminated it. Streaming/nonstreaming output matched in the single-request experiment. These failures and limitations remain part of the release evidence.
| Area | Status | Evidence / limit |
|---|---|---|
| Hardware and feasibility | VALIDATED | `reports/environment.json`, `docs/ARCHITECTURE.md` |
| A–E cost and parameter comparison | IMPLEMENTED estimates | Calculator tested; scenarios, not cluster measurements |
| Tiny dense model and causal attention | VALIDATED | Model tests, actual 120-step training |
| Original weights and export | VALIDATED | `artifacts/tiny/model.safetensors`; safetensors loading |
| Byte tokenizer | VALIDATED | Unicode/whitespace round trips; multilingual sample report |
| Optimal vocabulary / industrial tokenizer | PLANNED | Small public sample cannot select vocab size |
| Provenance/quality/license/secret/dedup/shards | VALIDATED lab subset | Pipeline tests; heuristic PII/language and quadratic dedup limits |
| Industrial data ingestion and mixtures | PLANNED | No large rights-cleared corpus acquired |
| Checkpoint/RNG/optimizer recovery | VALIDATED CPU | Exact uninterrupted/resume equality and forced kill test |
| Distributed/async checkpoints | PLANNED | No cluster; no distributed recovery claim |
| LoRA and adapter merge | VALIDATED miniature | Frozen base/gradient/merge tests; toy experiment |
| SFT/DPO | VALIDATED miniature | Executed optimization, no generalization claim |
| GRPO/PPO production RL | PLANNED | Group-advantage primitive tested; no full rollout learner |
| Local pretrained inference | VALIDATED | Unmodified Qwen3 and Qwen3.5 CPU runs |
| Streaming local inference | VALIDATED single request | Greedy stream/nonstream parity; first-text measurement |
| Local HTTP endpoint | VALIDATED protocol tests | Auth/schema/SSE tests with test backend; not production load test |
| Structured tools and authority | VALIDATED bounded subset | Typed arguments, denied paths/permissions, hash-based writes, command timeout |
| Agent state machine | VALIDATED with test double | Independent verification and loop/failure budgets |
| Actual model agent reliability | LIMITED / experimental | Public smoke failures retained; policy-aware follow-up reported separately |
| Repository index | VALIDATED limited | Python AST/imports, lexical indexing for target language extensions |
| Repository-scale coding completion | UNKNOWN | No representative hidden multi-language repair suite run |
| Local memory and failure lifecycle | VALIDATED | Expiration/correction/deletion, staged consent/evidence tests |
| Embedding search, automatic model router | PLANNED | Designs only; no fabricated semantic retrieval or routing savings |
| Voice coordinator | VALIDATED with test doubles | Stale generation canceled on interruption |
| File ASR / local TTS adapters | IMPLEMENTED, unvalidated here | Optional dependencies/hardware not installed/tested |
| Full-duplex speech and WER | PLANNED / UNKNOWN | No actual microphone dataset, audio measurements or echo cancellation |
| CPU INT8 reference comparison | VALIDATED narrow experiment | One fixed batch, 20 timing samples; no compelling speedup |
| BF16/FP8/INT4 large-model serving | PLANNED | Engine/hardware and task-quality testing needed |
| Browser/MCP/database external tools | PLANNED | Not exposed by runtime |
| 32K/64K/128K/256K reliable context | UNKNOWN | Configured upstream capacity is not effective-context evidence |
| Injection-resistant sandbox | PLANNED | Host executor explicitly not a sandbox; OS symlink test may skip |
| Observability | IMPLEMENTED local subset | Receipts/latency/token reports/hash audit; distributed traces/GPU metrics absent |
| Telemetry/private uploads | Disabled | No telemetry endpoint; curated publication excludes private/cache data |
| 100M/1B/7B/120B scaling | PLANNED | Requires data, compute, budget and evidence gates |
## Mission traceability
Mission sections 1–5, 17–22, 31–32 and 39: architecture report, formulas and gates.
Sections 6–8, 27, 34 and 36: tokenizer/data/adapters/failure modules, configs and provenance.
Sections 9–11 and 24–26: post-training primitives, toy optimization, evaluation and quantization reports.
Sections 12–16, 23 and 40: bounded context, memory, tools/agent/coding and voice modules with the limits above.
Sections 28–30 and 33: receipts/privacy/security/checkpoint recovery.
Sections 35, 37–38 and 41–48: repository structure, tests, actual experiments, honesty labels, execution history and go/no-go gates.
Full-production acceptance remains unmet for base-model capability, multi-language repository coding, independent reasoning improvement, robust real-model agents, voice, load-tested serving and distributed reliability. The published artifacts allow continuing each item without confusing plans with a trained product.