NEXORA / docs /ARCHITECTURE.md
devildasdf's picture
Release validated NEXORA research prototype, tiny weights and evidence
12496fc verified
|
Raw History Blame Contribute Delete
31.4 kB

NEXORA-120B: feasibility and architecture decision

Assessment date: 2026-09-28. 120B is a product ambition, not the size of the released checkpoint. IMPLEMENTED means executable code exists; VALIDATED means a named test/experiment passed; PLANNED means design only; UNKNOWN means no supporting measurement. The release is a research prototype. It does not meet the mission's production acceptance criteria.

1. Executive Decision

Choose E + D: a small existing instruction model, retrieval, typed tools and independent verifiers, with adaptation only after a measured failure justifies it. Keep a tiny from-scratch model to validate training and recovery. Do not train dense 120B on this workstation.

Observed resources: Ryzen 7 7435HS, 8 cores, 16,989,736,960 bytes RAM, RTX 3050 Laptop 4096 MiB, approximately 322 GB free on workspace drive at discovery. Installed PyTorch 2.13.0 is CPU-only. Hugging Face authentication and network downloads work. No allocated cluster or spending budget was supplied; no paid compute was launched. See reports/environment.json.

The useful inference baseline and original training experiment are separate artifacts. Upstream Qwen weights retain upstream identity, licensing and attribution. They are never relabeled as a newly pretrained NEXORA foundation model.

2. Recommended Architecture

Built reference model

Dense decoder, 820,736 total and active parameters, 4 layers, width 128, SwiGLU intermediate 384, 4 query heads, 2 KV heads, head dimension 32, RMSNorm, RoPE theta 10,000, tied 259-token UTF-8 byte embeddings/output, maximum input 256 tokens. Training uses 128-token sequences. No MoE, learned memory, MTP or custom kernels. SDPA may choose a platform kernel; this CPU run is not a FlashAttention performance result. Generation recomputes context and is a correctness baseline, not an optimized KV-cached server.

Existing-model integration

Qwen3-0.6B is a compatibility baseline, not a frontier assistant: 28 layers, width 1024, FFN 3072, 16 query/8 KV heads, head dimension 128, vocabulary 151936, tied embeddings. The family name is not an exact count of all stored parameters; report the measured count when available. Its configured maximum is 40960 tokens; NEXORA caps requests at 4096 including output. Effective context is UNKNOWN.

The newer Qwen3.5-0.8B is also evaluated. Its text stack uses 24 layers, width 1024, FFN 3584, alternating Gated DeltaNet/full attention with one full-attention layer in each four-layer group. Full attention has 8 query/2 KV heads of dimension 256. It also includes a vision encoder; this release exercises text only. Original advertised long-context support is not validation on this machine. Model configs and immutable upstream revisions are in reports/sources.json.

Upgrade candidate: Qwen3.5-4B on a suitable quantized runtime or a stronger coding model on a separately provisioned server. Evaluate quality/latency before switching. The HTTP backend permits replacing the model without giving it execution authority.

Counterfactual MoE designs: PLANNED, not trained or recommended for immediate spend

Property B: roughly 120B total C: larger sparse design
Total parameters 120,960,230,400 247,047,262,208
Active/token convention 6,279,570,432 11,495,149,568
Layers / hidden 49 / 3072 80 / 4096
Expert FFN intermediate 2048 1536
Routed experts / selected 128 / 4 160 / 4
Shared experts per layer 1 1
Query heads / KV heads / dim 24 / 8 / 128 32 / 8 / 128
Attention parameters 1,233,125,376 3,355,443,200
All expert parameters 119,304,880,128 243,101,859,840
Router parameters 19,267,584 52,428,800
Tied embedding parameters 402,653,184 536,870,912
Normalization parameters 304,128 659,456
BF16 weights, decimal GB 241.920 494.095
FP8 / INT8 ideal weights, GB 120.960 247.047
INT4 ideal weights, GB 60.480 123.524
BF16 KV, batch 1, 32K, GB 6.577 10.737
BF16 KV, batch 1, 128K, GB 26.307 42.950

Counts assume every block is MoE, tied vocabulary 131072 and bias-free projections. Active count includes the full output matrix, router, all attention, selected and shared experts; it is a FLOPs-planning convention, not a count of unique embedding rows read. Quantized values exclude scales, unquantized tensors, padding, workspace, allocator and KV cache. Formula code is in scripts/research_reports.py; outputs are in reports/compute.json.

These high sparsity designs may fail to learn competitive routing or quality. Equal total parameters do not imply equal ability. A 49-layer model does not divide evenly across arbitrary pipeline layouts: PP=7 is possible by layer count, PP=8 requires explicit uneven partitioning. No claim is made that the example topology is an executable launch for these designs.

3. Why This Architecture

Keep intelligence and authority separate. Spend scarce memory on a working pretrained model and retrieved evidence. Keep expensive systems optional. Parameter count alone cannot predict coding or reasoning quality.

Technique decisions and measurable gates:

Technique Decision / expected measurable benefit Complexity and failure modes Simpler alternative
Decoder-only IMPLEMENTED; efficient causal loss/generation baseline Autoregressive latency, hallucinations Existing decoder weights
GQA IMPLEMENTED; KV size scales with KV heads, 2 vs 4 halves this model's nominal KV Capacity loss and head-shape bugs Full MHA
MQA PLANNED ablation; fewer KV bytes than GQA Reduced representational capacity GQA
MLA PLANNED only; evaluate latent-cache byte savings Different projections/kernels, conversion risk GQA
RoPE IMPLEMENTED; position-dependent attention without learned table Position extrapolation failure Trained-length RoPE
RoPE scaling PLANNED; requires long-context quality tests Aliasing and short-context regression Retrieval at trained length
RMSNorm / SwiGLU IMPLEMENTED, established baseline Numerical precision/FFN memory LayerNorm/GELU ablation
FlashAttention Use backend support; measure time/peak VRAM Unsupported GPU/dtype and kernel differences PyTorch SDPA/math
Sparse/sliding attention PLANNED; bound quadratic cost Lost distant dependencies, mask correctness Short context + retrieval
Hybrid attention Existing Qwen3.5 candidate; benchmark CPU fallback State-cache kernels and backend compatibility Qwen3 dense GQA
Routed/shared experts PLANNED B/C; lower active FLOPs than dense All-to-all traffic, expert collapse, load imbalance Dense pretrained model
Auxiliary-loss-free balancing PLANNED ablation Stateful bias updates, unstable routing Auxiliary balancing loss
MTP PLANNED; evaluate accepted draft tokens and latency Extra heads/objective and rejection cost One-token prediction
Speculative decoding PLANNED; require net wall-clock gain Draft overhead and low acceptance Single-model decode
FP8 training Not on this CPU/3050 environment Hardware support, overflow/scaling recipes BF16/FP32
INT8/INT4 inference INT8 reference experiment; larger quantization PLANNED Outliers and task-specific degradation Unquantized baseline
KV compression PLANNED; bytes/token and quality gate Retrieval/precision loss Short context
CP / SP / EP PLANNED distributed experiments Communication, topology, collective correctness Single process

Do not attach invented percentage quality gains. Architecture changes must improve held-out quality at matched tokens/FLOPs or reduce measured latency/memory at noninferior quality. DeepSeek-V3 provides evidence for sparse training at scale, not proof these custom shapes work. Technical report

4. From-Scratch vs Existing Base Model

Path Cost/complexity Quality and forgetting Rights and decision
A: dense 120B scratch Highest compute; ~1.92 TB mixed-precision Adam training state before activations Entire capability must be learned; quality UNKNOWN Full data rights/provenance required; NO-GO locally
B: 120.96B / 6.28B-active MoE scratch Less arithmetic, still all expert storage and heavy network Routing and data efficiency UNKNOWN Same corpus obligations; research-only
C: 247.05B / 11.50B-active MoE scratch More memory/checkpoints than B No established gain over a good existing model Reject without scaling-law evidence
D: existing 30.5B / ~3.3B-active model Training/serving state still reflects total parameters Strong starting point, domain shift can degrade general skills Apache-2.0 candidate; preserve notices; server tier
E: small existing model + tools Lowest local cost; executor engineering matters Limited base reasoning; tools do not cure all errors Best feasible starting point, measure failures
Continual pretraining 1–10B curated tokens as initial scenario Catastrophic forgetting; mix rehearsal data Verify every dataset; select only for missing domain coverage
LoRA / QLoRA Train adapter parameters; still forward/backward through base Limited capacity but reversible; quantization may degrade Base license survives; recommended first adaptation
Full fine-tuning Adam state/gradients for all weights Greater capacity and forgetting risk Use when adapter plateau is demonstrated
Distillation Teacher generation + filtering + student training Teacher errors and lower student ceiling Teacher terms and output/data rights must permit use
Dense-to-MoE conversion/composition Experimental architecture and recovery work Duplicating weights is not new knowledge No claimed free quality gain; defer

5. Compute Requirements

Reproducible calculator: python -m nexora.cli estimate. Decimal GB throughout.

training FLOPs ≈ 6 × active parameters × training tokens. hours = FLOPs / (GPU count × dense peak FLOPs/s × MFU × 3600). GPU-hours = hours × GPUs; compute charge = GPU-hours × assumed hourly price. KV bytes = 2 × layers × KV heads × head_dim × tokens × batch × bytes_per_element. mixed-precision Adam state ≈ 16 bytes/total parameter (BF16 weights+gradients, FP32 master+m+v). checkpoint weights+master+m+v ≈ 14 bytes/parameter; FP32-only and optimizer choices differ.

All scenarios below use an illustrative H100 SXM dense BF16 peak of 989 TFLOP/s. Low/expected/high cost assumptions are MFU .50/.35/.20 and $2/$3/$5 per GPU-hour; these are not current vendor quotes. MoE MFU can be substantially worse; 6NT omits attention quadratic work, routing, communication and recomputation. No forecast is a purchase recommendation.

Path Training tokens GPUs Expected hours Expected GPU-hours Compute USD low / expected / high
A dense 120B scratch 2.4T 1024 1354.2 ~1,386,682 1,941,355 / 4,160,046 / 12,133,468
B custom 120.96B sparse 2.4T 256 283.5 ~72,565 101,591 / 217,694 / 634,941
C custom 247.05B sparse 3T 512 324.3 ~166,043 232,460 / 498,129 / 1,452,875
D 30.5B continual 10B 8 19.9 ~159 222 / 477 / 1390
E 4B continual 1B 8 2.4 ~19 27 / 58 / 169

The lower D/E figures represent only a small incremental training pass, not pretraining the foundation. Small runs may never reach assumed MFU. Add data acquisition, evaluation, failed runs, storage, network, staff and opportunity cost. Hardware ownership/power is separate from rental pricing.

Weights for A require 240 GB BF16, 120 GB ideal FP8/INT8, 60 GB ideal INT4; D approximately 61/30.5/15.25 GB; E approximately 8/4/2 GB. Serving requires extra KV and runtime headroom. Throughput and latency for A–E at scale are UNKNOWN. Initial acceptance targets, not forecasts: interactive decode >=10 tokens/s and P95 TTFT <=2 seconds at batch 1 on the selected deployment profile; relax only through an explicit product decision. Measured local baseline numbers are published separately.

6. Training Data Requirements

Local experiment: 90 original synthetic training records, 11,600 byte tokens, three distinct validation records, 424 tokens. This tiny corpus is deliberately insufficient for general capability. It is public development data, not a private benchmark.

For a justified scratch experiment start with pilot datasets around 2B tokens at ~100M parameters, 20B at ~1B, and 140B at ~7B. The often-used ~20 tokens/parameter is a planning starting point for compute-optimal dense training, not a universal rule for MoE, post-training, or inference-optimal deployment. Re-estimate from learning curves. Scaling-law paper

Proposed pretraining mixture for experiments (not acquired): 35% licensed code/technical documentation, 25% quality web, 15% math/science, 10% legally usable books/reference, 10% multilingual including Hindi/Hinglish, 5% structured formats. Sample proportions by tokens and measure multilingual/code regressions. Separate post-training trajectories by repository and task lineage; never mix benchmark solutions into pretraining.

Implemented lab pipeline: explicit license/source/split → newline/NFC normalization preserving indentation → quality checks → credential-pattern rejection/email redaction → exact hash and shingle near-dedup → holdout-overlap rejection → declared domains and transparent heuristic score → byte tokens → checksummed NPY shards. The language hint is not industrial language ID. PII regex is incomplete. Near-dedup is quadratic and unsuitable for web scale. Corpus-wide scalable MinHash/LSH, learned quality/language models, provenance adjudication and mixture scheduler remain PLANNED.

Tokenizer: compare lossless bytes, byte-level BPE and unigram; preserve whitespace, arbitrary UTF-8 and structured syntax. Existing weights MUST retain their tokenizer unless embeddings are explicitly retrained. reports/tokenizer-*.json measures English/Hindi/Hinglish/Python/JSON/math/shell/XML samples. The 259 vocabulary is an exhaustive byte baseline, not an optimized natural-language vocabulary. A scratch BPE study should test 32K/64K/128K with equal corpus, embedding-adjusted FLOPs and held-out compression/loss. Eight examples cannot select an industrial vocabulary.

7. Training Pipeline

Implemented: seed-controlled AdamW training, cosine LR, gradient clipping/nonfinite checks, deterministic validation, atomic checkpoint pointer, SHA-256 validation, model/optimizer/CPU+CUDA/batch/Python/NumPy RNG recovery and safetensors export. Resume requires matching config and manifest. Existing state uses safe weights_only=True loading; load only trusted checkpoints.

R0: pretraining mechanics on tiny data. R1: assistant-token-masked SFT. R2: executable code/math/SQL verifiers in isolated environments. R3: rejection sampling from independently successful candidates. R4: outcome RL, group advantage primitive exists; full GRPO rollout/update pipeline PLANNED. R5: process supervision only when independently scored evidence beats outcome-only baseline. R6/R7: tool and long-horizon curricula require real successful trajectories. DPO and LoRA optimization are exercised on toy examples. PPO adds value/rollout complexity; defer until simpler methods plateau. No miniature result demonstrates improved general reasoning.

Coding examples must include repository snapshots, user request, investigation, failed tool receipts, minimal diff, compiler/tests and independently verified outcome. Split by repository/time, validate Python/JS/TS/C/C++/C#/Java/Go/Rust/SQL/Bash/PowerShell/HTML/CSS/Kotlin/Swift with native toolchains. Never give generated code a success label because its explanation sounds good. Keep verification code outside writable agent scope.

8. Coding Architecture

IMPLEMENTED: bounded source traversal, Python AST symbols/import edges and multilingual lexical indexing/retrieval; compare-and-swap file replacement via content hash; configured commands; final owner-selected verification; step/failure/repetition budgets. Unsupported languages are searchable text, not AST-aware. LSP integration, vector search, rich dependency graph, patch-hunk application and benchmark-grade repository repair remain PLANNED.

Target flow: task → repository map → retrieve relevant files → brief plan → minimal edits → build/test → inspect failures → repair → regression suite → diff and receipt-based report. Runtime tests use a scripted model solely to test state transitions. Real-model smoke results use an actual downloaded model and retain failures. Neither proves repository-scale coding capability.

9. Voice Architecture

Choose pipeline voice: local microphone/VAD → ASR → model/router → TTS → speaker. It isolates latency and allows replacing each model. Native audio models may improve prosody/full-duplex interaction but introduce a separate training/serving burden; no evidence here justifies building one.

IMPLEMENTED: async turn coordinator cancels stale generation on barge-in and invokes audio-stop callback; file ASR adapter using faster-whisper; OS TTS adapter via pyttsx3; WER and latency metrics. VALIDATED: cancellation with controlled async test doubles. Optional ASR/TTS hardware paths are not measured. Microphone capture, incremental ASR, VAD, acoustic echo cancellation, audio buffering/prosody and real full-duplex latency remain PLANNED. No fabricated WER or speech latency.

Track speech-end→transcript, transcript→first token, first token→first audio and total perceived latency separately. Collect consented English/Hindi/Hinglish recordings with interruptions and background noise. Pin ASR/TTS model versions and licenses. ASR implementation source

10. Agent Architecture

THINK is private model computation; exposed plans are concise action summaries. Runtime state is PLAN → ACT/OBSERVE → VERIFY → CONTINUE/FINISH. JSON Schema validates model actions and tool arguments. Unknown tools and extra arguments fail closed. READ/WRITE/EXECUTE permissions are enforced for implemented tools. NETWORK is an explicit inference opt-in; PRIVILEGED/IRREVERSIBLE are reserved classes with no implemented capabilities.

Filesystem tools resolve workspace containment, exclude private/control paths, bound reads and compare existing file hashes before overwrites. Commands are immutable owner-configured argument arrays, without shell expansion. Output capture is bounded; timeouts attempt process-tree termination. Receipts distinguish exit status, failure, elapsed time and truncation. Audit logs omit content and retain hashes. Per-process idempotency keys prevent duplicate execution within one executor; durable crash-safe transaction recovery is PLANNED. Do not automatically retry mutations. Read-only retries must remain bounded and observable.

The host command runner is not a sandbox. Enabling it runs code with the OS user's authority. A configured interpreter/test command can execute arbitrary repository code. Environment trimming does not isolate filesystem/network access. Use an external container/VM with no credentials, read-only hidden tests, restricted mounts and network policy before accepting hostile tasks. Concurrency/symlink race resistance requires OS-level primitives; path checks alone do not establish an adversarial security boundary. Browser/MCP/database integrations are PLANNED; do not expose arbitrary external actions through generic model strings.

11. Memory Architecture

L0 conversation in backend messages; L1 session context; L2 semantic records; L3 episodes; L4 projects; L5 preferences. SQLite stores layer/text/source/time/confidence/provenance/expiration, supports replacement by ID and deletion. Retrieval is lexical overlap, not embedding similarity; richer semantic retrieval is PLANNED. Avoid persisting unconfirmed inference as owner fact. TTL and confidence filters apply at retrieval. Audit deletion cannot promise removal from separately copied backups; define backup retention before production.

12. Inference Architecture

Local Transformers uses upstream KV caching and bounds input+output context. Small reference inference is intentionally simple. An HTTP chat adapter talks to explicit compatible endpoints, rejects non-HTTP URLs/embedded credentials, disables redirects/proxies, limits response size and requires remote-network opt-in. API credentials come from a named environment variable, never model messages.

Server-tier candidates: vLLM and SGLang for continuous batching, paged cache, prefix reuse, structured outputs and parallel inference; llama.cpp for CPU/GPU offload and GGUF on personal systems. Verify current engine/model/quantizer compatibility before deployment. vLLM's schema constraint support is evidence of an available backend feature, not proof NEXORA's local Transformers path constrains logits. Structured outputs

Router design: start with explicit fast/code/large endpoint routing, then calibrate an uncertainty policy on held-out failures. If 80% of requests use a cost-1 small model and 20% cost-10 large model, average model cost is 2.8 rather than 10 (72% arithmetic reduction), before verifier/retry overhead and quality loss. This is a scenario, not measured savings. Reject classifiers whose escalations conceal failure. No claim that a handcrafted prompt-length heuristic estimates difficulty.

Measure TTFT, TPOT, output tokens/s, requests/s, queue delay, P50/P95/P99, peak VRAM and failure rate under fixed prompts/concurrency. Report cold loading separately. FP32 baseline and CPU dynamic INT8 reference experiment are measured; BF16, FP8, INT4 and mixed precision task-quality suites remain PLANNED. Perplexity alone is insufficient.

13. Evaluation Framework

Published evidence: unit/integration tests, exact resume parity, checksum rejection, causal masking, tokenizer round trips, tool boundaries, timeouts, false-success prevention, memory expiry, toy post-training, local baseline inference, agent smoke, and reference quantization. All samples and limitations remain visible in reports/.

Required next suites: fresh repository repairs with hidden tests (SWE-bench-style isolation), execution-based multi-language tasks, verifiable math/science/SQL, multi-step agent target-state tests, long-context needles/multi-needles/cross-document synthesis/repo edits, voice WER/barge-in, injected file/web instructions and credential boundary tests. Keep hidden tests outside agent workspace and unavailable to the data pipeline. Pin test generation seed/date, dataset revision and harness source. Public samples have unknown contamination and are not a private holdout.

Report pass@1 and unbiased pass@k with sample counts, Wilson intervals for binomial rates, failure categories, environment and freshness. Verify compilation/execution rather than string similarity for coding. Vary reward attacks: modify tests, fabricate receipts, print expected output without state change, hide nonzero exit, trigger output truncation, exploit stale cache, manipulate a SQL oracle. Reject apparent success on these attacks. No benchmark-specific model branches or hardcoded solver answers.

14. Cost Model

Tier 1 (this workstation): tiny architecture training, small-model inference, runtime development, CPU experiments. Dense 120B scratch is economically and physically irrational. Tier 2 (8 GPUs, e.g. 80GB each): adapters and selected full/domain training of moderate models, 30B-class MoE serving. 640 GB aggregate is not 640 GB usable per rank; activations/KV/network matter. Tier 3 (8 nodes × 8): ~1B–7B scaling runs and research MoE, contingent on bandwidth and data. Tier 4 (32–128 nodes × 8): startup-scale pretraining only with product/data/learning-curve evidence and recovery drills. Tier 5 (128+ nodes × 8): 120B dense/large MoE feasible in principle, multi-million-dollar all-in program, not authorized here.

Distributed choice: PyTorch for prototype; FSDP2 for sharding supported dense models; Megatron-Core/NeMo for mature TP/PP/CP/EP at scale; DeepSpeed where its optimizer/offload integration wins measured benchmarks. Custom kernels only after profiles establish a bottleneck. SP often accompanies TP and is not another independent world-size factor.

Topology convention in calculator: world=TP×PP×CP×DP, EP divides the DP group. Examples: 1×8 GPUs=1×1×1×8; 8×8=2×2×1×16 with EP=8 and expert-DP=2; 128×8=8×8×2×8. Framework parallel folding/expert TP may use different decompositions; validate rank groups and divisibility against the actual stack. Megatron parallelism guide

Use high-bandwidth intra-node links for TP; scale-out all-to-all for MoE requires measured network bisection/latency. 200–400Gb/s per-GPU-class fabric is a planning target, not guaranteed sufficient. A 120B training checkpoint at 14 B/parameter is 1.68 TB; keeping three plus replicas consumes many TB. 2.4T uint32 tokens take 9.6 TB before raw text/intermediate copies. Add storage throughput sufficient to finish checkpoints within the tolerated pause; asynchronously saving without version consistency is unsafe. Activations depend on batch×sequence×layers×width plus attention strategy; profile representative steps instead of inventing a constant.

15. Prototype Roadmap

Current sprint produces executable training/runtime modules, an actual tiny checkpoint, actual inference tests and reproducibility artifacts. Fix failed tests before upload. Validate forced-kill recovery, record observed model limitations, and release without production claims. The smallest useful next experiment is a stronger model on the same fixed coding/tool harness, not immediately scaling the custom decoder.

16. Scaling Roadmap

Phase Entrance Exit
0 feasibility Hardware/rights inventory Scenarios, explicit budget and chosen path
1 tokenizer Provenance-cleared sample Round trips, domain compression and loss study
2 data Allowed sources and private holdout boundaries Reproducible shards, duplicate/PII/contamination audit
3 small prototype Validated model config Stable gradients, loss decrease and recovery
4 scaling laws >=3 sizes/token budgets with repeated seeds Fit+uncertainty predicts withheld run; no unstable routing
5 base pretraining Data/cluster/budget gate approved Held-out loss target, recovery drills, complete lineage
6 continual domain Evidence of domain gap Gain on private domain set without broad regression
7 coding Sandboxed multi-language trajectories Hidden-test repository success improvement
8 reasoning Reliable verifiable tasks Held-out accuracy/calibration improvement
9 tools Typed receipts and permission tests Multi-step target-state success without boundary escape
10 post-training/RL Audited rewards and baseline Generalization gain; reward hacks fail
11 long context Short-context gate 32K→64K→128K→256K only when retrieval/synthesis/repo gates pass
12 quantization Unquantized reference Noninferior task quality and measured resource win
13 serving Selected hardware/engine Load test P95/P99, queue/memory limits and failure recovery
14 voice Consented recordings and audio hardware WER, interruption and end-to-end latency targets
15 red team End-to-end deployment Injection, boundary, reward and privacy tests pass
16 release Reproducibility and honest card Immutable artifacts/checksums and clean install smoke
17 improvement Consented failure capture Reproduce→minimize→test→train→re-evaluate→controlled release

Do not extrapolate 820K results to 120B. Proposed 100M/1B/7B experiments remain PLANNED and require separate resources/data. No effective 32K+ context is claimed by this release.

17. Major Risks

Small pretrained models may produce invalid JSON and poor reasoning; rejection is preferable to fabricated success. Local CPU kernels are slow. Toy data overfits. MoE arithmetic savings can disappear into network inefficiency. QLoRA can still exceed 4GB after activations. Host execution is not isolated. License lists do not replace provenance due diligence. Short public tests have weak statistical power. Long context and native voice are unvalidated. Future dependencies may change; record tested versions and source hashes.

18. Repository Architecture

nexora/: model, tokenizer, data, training, adapters, posttraining, compute, evaluation, inference, agent, tools, coding, memory, voice and CLI. configs/: explicit experiment/permission files. examples/: original lab corpus. scripts/: reproducible experiments, research capture and release utilities. tests/: critical logic. docs/: decision, operations, scope and security. artifacts/: original tiny weights/checkpoints/data/experimental adapter. reports/: actual measurements, immutable source revisions and release manifest. .cache/: downloaded upstream weights, excluded from publication.

19. First Implementation Sprint

See docs/STATUS.md for the final evidence matrix. Changes are validated by tests rather than claiming every planned component exists. Publish only project-generated artifacts and source; exclude upstream weight caches, tokens, local private memory and machine-specific credentials. Provide commands for download, train, resume, inference, indexing, agent execution and re-running measurements.

20. Go / No-Go Gates

GO: publish the scoped research milestone when tests pass, checksums verify, actual experiment reports are included and the Hub upload is inspected. NO-GO: claim trained 120B, general coding competence, production sandboxing, reliable long context or full-duplex voice from this evidence. NO-GO: spend on scratch scaling before rights-cleared data, independent evaluation, a funded compute plan and kill/recovery tests exist. Only call a larger model a capability improvement after controlled evaluation.

Primary sources and evidence policy

Verified model metadata/configs and Apache-2.0 declarations are captured with revision hashes in reports/sources.json. Qwen cards report their own benchmark results; NEXORA does not treat those as local measurements. Additional sources: Qwen3.5-0.8B, Qwen3.5-4B, Qwen3-30B-A3B, PEFT LoRA, PyTorch distributed checkpoint.

Competitor lessons: Qwen provides reusable multilingual instruction weights; DeepSeek demonstrates sparsity/MLA/MTP at scale; Megatron offers tested parallel building blocks; vLLM offers serving/structured-output machinery. Transferring any of these into this constrained environment requires measurement. The custom reference model intentionally adds no unmeasured exotic architecture.