NEXORA / docs /STATUS.md
devildasdf's picture
Release validated NEXORA research prototype, tiny weights and evidence
12496fc verified
|
Raw History Blame Contribute Delete
5.35 kB

Implementation and validation status

The mission is much larger than one workstation can train or certify. This release is a completed local prototype milestone, not fulfillment of all production acceptance criteria. No unimplemented service is represented by a mock production endpoint.

Final validation: 36 passed, 1 skipped. The skip is a native Windows symlink-creation privilege limitation. Real Qwen3.5 agent follow-up after capability filtering still failed by repeated reads; the loop guard terminated it. Streaming/nonstreaming output matched in the single-request experiment. These failures and limitations remain part of the release evidence.

Area Status Evidence / limit
Hardware and feasibility VALIDATED reports/environment.json, docs/ARCHITECTURE.md
A–E cost and parameter comparison IMPLEMENTED estimates Calculator tested; scenarios, not cluster measurements
Tiny dense model and causal attention VALIDATED Model tests, actual 120-step training
Original weights and export VALIDATED artifacts/tiny/model.safetensors; safetensors loading
Byte tokenizer VALIDATED Unicode/whitespace round trips; multilingual sample report
Optimal vocabulary / industrial tokenizer PLANNED Small public sample cannot select vocab size
Provenance/quality/license/secret/dedup/shards VALIDATED lab subset Pipeline tests; heuristic PII/language and quadratic dedup limits
Industrial data ingestion and mixtures PLANNED No large rights-cleared corpus acquired
Checkpoint/RNG/optimizer recovery VALIDATED CPU Exact uninterrupted/resume equality and forced kill test
Distributed/async checkpoints PLANNED No cluster; no distributed recovery claim
LoRA and adapter merge VALIDATED miniature Frozen base/gradient/merge tests; toy experiment
SFT/DPO VALIDATED miniature Executed optimization, no generalization claim
GRPO/PPO production RL PLANNED Group-advantage primitive tested; no full rollout learner
Local pretrained inference VALIDATED Unmodified Qwen3 and Qwen3.5 CPU runs
Streaming local inference VALIDATED single request Greedy stream/nonstream parity; first-text measurement
Local HTTP endpoint VALIDATED protocol tests Auth/schema/SSE tests with test backend; not production load test
Structured tools and authority VALIDATED bounded subset Typed arguments, denied paths/permissions, hash-based writes, command timeout
Agent state machine VALIDATED with test double Independent verification and loop/failure budgets
Actual model agent reliability LIMITED / experimental Public smoke failures retained; policy-aware follow-up reported separately
Repository index VALIDATED limited Python AST/imports, lexical indexing for target language extensions
Repository-scale coding completion UNKNOWN No representative hidden multi-language repair suite run
Local memory and failure lifecycle VALIDATED Expiration/correction/deletion, staged consent/evidence tests
Embedding search, automatic model router PLANNED Designs only; no fabricated semantic retrieval or routing savings
Voice coordinator VALIDATED with test doubles Stale generation canceled on interruption
File ASR / local TTS adapters IMPLEMENTED, unvalidated here Optional dependencies/hardware not installed/tested
Full-duplex speech and WER PLANNED / UNKNOWN No actual microphone dataset, audio measurements or echo cancellation
CPU INT8 reference comparison VALIDATED narrow experiment One fixed batch, 20 timing samples; no compelling speedup
BF16/FP8/INT4 large-model serving PLANNED Engine/hardware and task-quality testing needed
Browser/MCP/database external tools PLANNED Not exposed by runtime
32K/64K/128K/256K reliable context UNKNOWN Configured upstream capacity is not effective-context evidence
Injection-resistant sandbox PLANNED Host executor explicitly not a sandbox; OS symlink test may skip
Observability IMPLEMENTED local subset Receipts/latency/token reports/hash audit; distributed traces/GPU metrics absent
Telemetry/private uploads Disabled No telemetry endpoint; curated publication excludes private/cache data
100M/1B/7B/120B scaling PLANNED Requires data, compute, budget and evidence gates

Mission traceability

Mission sections 1–5, 17–22, 31–32 and 39: architecture report, formulas and gates. Sections 6–8, 27, 34 and 36: tokenizer/data/adapters/failure modules, configs and provenance. Sections 9–11 and 24–26: post-training primitives, toy optimization, evaluation and quantization reports. Sections 12–16, 23 and 40: bounded context, memory, tools/agent/coding and voice modules with the limits above. Sections 28–30 and 33: receipts/privacy/security/checkpoint recovery. Sections 35, 37–38 and 41–48: repository structure, tests, actual experiments, honesty labels, execution history and go/no-go gates.

Full-production acceptance remains unmet for base-model capability, multi-language repository coding, independent reasoning improvement, robust real-model agents, voice, load-tested serving and distributed reliability. The published artifacts allow continuing each item without confusing plans with a trained product.