Download docs/ARCHITECTURE.md from devildasdf/NEXORA: direct link, hf CLI and curl.
- Browser
- Download file 31.4 kB
-
https://huggingface.co/devildasdf/NEXORA/resolve/main/docs/ARCHITECTURE.md
- Command line
-
hf download hf://devildasdf/NEXORA/docs/ARCHITECTURE.md
-
curl -L -o ARCHITECTURE.md https://huggingface.co/devildasdf/NEXORA/resolve/main/docs/ARCHITECTURE.md
NEXORA-120B: feasibility and architecture decision
Assessment date: 2026-09-28. 120B is a product ambition, not the size of the released checkpoint. IMPLEMENTED means executable code exists; VALIDATED means a named test/experiment passed; PLANNED means design only; UNKNOWN means no supporting measurement. The release is a research prototype. It does not meet the mission's production acceptance criteria.
1. Executive Decision
Choose E + D: a small existing instruction model, retrieval, typed tools and independent verifiers, with adaptation only after a measured failure justifies it. Keep a tiny from-scratch model to validate training and recovery. Do not train dense 120B on this workstation.
Observed resources: Ryzen 7 7435HS, 8 cores, 16,989,736,960 bytes RAM, RTX 3050 Laptop 4096 MiB, approximately 322 GB free on workspace drive at discovery. Installed PyTorch 2.13.0 is CPU-only. Hugging Face authentication and network downloads work. No allocated cluster or spending budget was supplied; no paid compute was launched. See reports/environment.json.
The useful inference baseline and original training experiment are separate artifacts. Upstream Qwen weights retain upstream identity, licensing and attribution. They are never relabeled as a newly pretrained NEXORA foundation model.
2. Recommended Architecture
Built reference model
Dense decoder, 820,736 total and active parameters, 4 layers, width 128, SwiGLU intermediate 384, 4 query heads, 2 KV heads, head dimension 32, RMSNorm, RoPE theta 10,000, tied 259-token UTF-8 byte embeddings/output, maximum input 256 tokens. Training uses 128-token sequences. No MoE, learned memory, MTP or custom kernels. SDPA may choose a platform kernel; this CPU run is not a FlashAttention performance result. Generation recomputes context and is a correctness baseline, not an optimized KV-cached server.
Existing-model integration
Qwen3-0.6B is a compatibility baseline, not a frontier assistant: 28 layers, width 1024, FFN 3072, 16 query/8 KV heads, head dimension 128, vocabulary 151936, tied embeddings. The family name is not an exact count of all stored parameters; report the measured count when available. Its configured maximum is 40960 tokens; NEXORA caps requests at 4096 including output. Effective context is UNKNOWN.
The newer Qwen3.5-0.8B is also evaluated. Its text stack uses 24 layers, width 1024, FFN 3584, alternating Gated DeltaNet/full attention with one full-attention layer in each four-layer group. Full attention has 8 query/2 KV heads of dimension 256. It also includes a vision encoder; this release exercises text only. Original advertised long-context support is not validation on this machine. Model configs and immutable upstream revisions are in reports/sources.json.
Upgrade candidate: Qwen3.5-4B on a suitable quantized runtime or a stronger coding model on a separately provisioned server. Evaluate quality/latency before switching. The HTTP backend permits replacing the model without giving it execution authority.
Counterfactual MoE designs: PLANNED, not trained or recommended for immediate spend
| Property | B: roughly 120B total | C: larger sparse design |
|---|---|---|
| Total parameters | 120,960,230,400 | 247,047,262,208 |
| Active/token convention | 6,279,570,432 | 11,495,149,568 |
| Layers / hidden | 49 / 3072 | 80 / 4096 |
| Expert FFN intermediate | 2048 | 1536 |
| Routed experts / selected | 128 / 4 | 160 / 4 |
| Shared experts per layer | 1 | 1 |
| Query heads / KV heads / dim | 24 / 8 / 128 | 32 / 8 / 128 |
| Attention parameters | 1,233,125,376 | 3,355,443,200 |
| All expert parameters | 119,304,880,128 | 243,101,859,840 |
| Router parameters | 19,267,584 | 52,428,800 |
| Tied embedding parameters | 402,653,184 | 536,870,912 |
| Normalization parameters | 304,128 | 659,456 |
| BF16 weights, decimal GB | 241.920 | 494.095 |
| FP8 / INT8 ideal weights, GB | 120.960 | 247.047 |
| INT4 ideal weights, GB | 60.480 | 123.524 |
| BF16 KV, batch 1, 32K, GB | 6.577 | 10.737 |
| BF16 KV, batch 1, 128K, GB | 26.307 | 42.950 |
Counts assume every block is MoE, tied vocabulary 131072 and bias-free projections. Active count includes the full output matrix, router, all attention, selected and shared experts; it is a FLOPs-planning convention, not a count of unique embedding rows read. Quantized values exclude scales, unquantized tensors, padding, workspace, allocator and KV cache. Formula code is in scripts/research_reports.py; outputs are in reports/compute.json.
These high sparsity designs may fail to learn competitive routing or quality. Equal total parameters do not imply equal ability. A 49-layer model does not divide evenly across arbitrary pipeline layouts: PP=7 is possible by layer count, PP=8 requires explicit uneven partitioning. No claim is made that the example topology is an executable launch for these designs.
3. Why This Architecture
Keep intelligence and authority separate. Spend scarce memory on a working pretrained model and retrieved evidence. Keep expensive systems optional. Parameter count alone cannot predict coding or reasoning quality.
Technique decisions and measurable gates:
| Technique | Decision / expected measurable benefit | Complexity and failure modes | Simpler alternative |
|---|---|---|---|
| Decoder-only | IMPLEMENTED; efficient causal loss/generation baseline | Autoregressive latency, hallucinations | Existing decoder weights |
| GQA | IMPLEMENTED; KV size scales with KV heads, 2 vs 4 halves this model's nominal KV | Capacity loss and head-shape bugs | Full MHA |
| MQA | PLANNED ablation; fewer KV bytes than GQA | Reduced representational capacity | GQA |
| MLA | PLANNED only; evaluate latent-cache byte savings | Different projections/kernels, conversion risk | GQA |
| RoPE | IMPLEMENTED; position-dependent attention without learned table | Position extrapolation failure | Trained-length RoPE |
| RoPE scaling | PLANNED; requires long-context quality tests | Aliasing and short-context regression | Retrieval at trained length |
| RMSNorm / SwiGLU | IMPLEMENTED, established baseline | Numerical precision/FFN memory | LayerNorm/GELU ablation |
| FlashAttention | Use backend support; measure time/peak VRAM | Unsupported GPU/dtype and kernel differences | PyTorch SDPA/math |
| Sparse/sliding attention | PLANNED; bound quadratic cost | Lost distant dependencies, mask correctness | Short context + retrieval |
| Hybrid attention | Existing Qwen3.5 candidate; benchmark CPU fallback | State-cache kernels and backend compatibility | Qwen3 dense GQA |
| Routed/shared experts | PLANNED B/C; lower active FLOPs than dense | All-to-all traffic, expert collapse, load imbalance | Dense pretrained model |
| Auxiliary-loss-free balancing | PLANNED ablation | Stateful bias updates, unstable routing | Auxiliary balancing loss |
| MTP | PLANNED; evaluate accepted draft tokens and latency | Extra heads/objective and rejection cost | One-token prediction |
| Speculative decoding | PLANNED; require net wall-clock gain | Draft overhead and low acceptance | Single-model decode |
| FP8 training | Not on this CPU/3050 environment | Hardware support, overflow/scaling recipes | BF16/FP32 |
| INT8/INT4 inference | INT8 reference experiment; larger quantization PLANNED | Outliers and task-specific degradation | Unquantized baseline |
| KV compression | PLANNED; bytes/token and quality gate | Retrieval/precision loss | Short context |
| CP / SP / EP | PLANNED distributed experiments | Communication, topology, collective correctness | Single process |
Do not attach invented percentage quality gains. Architecture changes must improve held-out quality at matched tokens/FLOPs or reduce measured latency/memory at noninferior quality. DeepSeek-V3 provides evidence for sparse training at scale, not proof these custom shapes work. Technical report
4. From-Scratch vs Existing Base Model
| Path | Cost/complexity | Quality and forgetting | Rights and decision |
|---|---|---|---|
| A: dense 120B scratch | Highest compute; ~1.92 TB mixed-precision Adam training state before activations | Entire capability must be learned; quality UNKNOWN | Full data rights/provenance required; NO-GO locally |
| B: 120.96B / 6.28B-active MoE scratch | Less arithmetic, still all expert storage and heavy network | Routing and data efficiency UNKNOWN | Same corpus obligations; research-only |
| C: 247.05B / 11.50B-active MoE scratch | More memory/checkpoints than B | No established gain over a good existing model | Reject without scaling-law evidence |
| D: existing 30.5B / ~3.3B-active model | Training/serving state still reflects total parameters | Strong starting point, domain shift can degrade general skills | Apache-2.0 candidate; preserve notices; server tier |
| E: small existing model + tools | Lowest local cost; executor engineering matters | Limited base reasoning; tools do not cure all errors | Best feasible starting point, measure failures |
| Continual pretraining | 1–10B curated tokens as initial scenario | Catastrophic forgetting; mix rehearsal data | Verify every dataset; select only for missing domain coverage |
| LoRA / QLoRA | Train adapter parameters; still forward/backward through base | Limited capacity but reversible; quantization may degrade | Base license survives; recommended first adaptation |
| Full fine-tuning | Adam state/gradients for all weights | Greater capacity and forgetting risk | Use when adapter plateau is demonstrated |
| Distillation | Teacher generation + filtering + student training | Teacher errors and lower student ceiling | Teacher terms and output/data rights must permit use |
| Dense-to-MoE conversion/composition | Experimental architecture and recovery work | Duplicating weights is not new knowledge | No claimed free quality gain; defer |
5. Compute Requirements
Reproducible calculator: python -m nexora.cli estimate. Decimal GB throughout.
training FLOPs ≈ 6 × active parameters × training tokens.
hours = FLOPs / (GPU count × dense peak FLOPs/s × MFU × 3600).
GPU-hours = hours × GPUs; compute charge = GPU-hours × assumed hourly price.
KV bytes = 2 × layers × KV heads × head_dim × tokens × batch × bytes_per_element.
mixed-precision Adam state ≈ 16 bytes/total parameter (BF16 weights+gradients, FP32 master+m+v).
checkpoint weights+master+m+v ≈ 14 bytes/parameter; FP32-only and optimizer choices differ.
All scenarios below use an illustrative H100 SXM dense BF16 peak of 989 TFLOP/s. Low/expected/high cost assumptions are MFU .50/.35/.20 and $2/$3/$5 per GPU-hour; these are not current vendor quotes. MoE MFU can be substantially worse; 6NT omits attention quadratic work, routing, communication and recomputation. No forecast is a purchase recommendation.
| Path | Training tokens | GPUs | Expected hours | Expected GPU-hours | Compute USD low / expected / high |
|---|---|---|---|---|---|
| A dense 120B scratch | 2.4T | 1024 | 1354.2 | ~1,386,682 | 1,941,355 / 4,160,046 / 12,133,468 |
| B custom 120.96B sparse | 2.4T | 256 | 283.5 | ~72,565 | 101,591 / 217,694 / 634,941 |
| C custom 247.05B sparse | 3T | 512 | 324.3 | ~166,043 | 232,460 / 498,129 / 1,452,875 |
| D 30.5B continual | 10B | 8 | 19.9 | ~159 | 222 / 477 / 1390 |
| E 4B continual | 1B | 8 | 2.4 | ~19 | 27 / 58 / 169 |
The lower D/E figures represent only a small incremental training pass, not pretraining the foundation. Small runs may never reach assumed MFU. Add data acquisition, evaluation, failed runs, storage, network, staff and opportunity cost. Hardware ownership/power is separate from rental pricing.
Weights for A require 240 GB BF16, 120 GB ideal FP8/INT8, 60 GB ideal INT4; D approximately 61/30.5/15.25 GB; E approximately 8/4/2 GB. Serving requires extra KV and runtime headroom. Throughput and latency for A–E at scale are UNKNOWN. Initial acceptance targets, not forecasts: interactive decode >=10 tokens/s and P95 TTFT <=2 seconds at batch 1 on the selected deployment profile; relax only through an explicit product decision. Measured local baseline numbers are published separately.
6. Training Data Requirements
Local experiment: 90 original synthetic training records, 11,600 byte tokens, three distinct validation records, 424 tokens. This tiny corpus is deliberately insufficient for general capability. It is public development data, not a private benchmark.
For a justified scratch experiment start with pilot datasets around 2B tokens at ~100M parameters, 20B at ~1B, and 140B at ~7B. The often-used ~20 tokens/parameter is a planning starting point for compute-optimal dense training, not a universal rule for MoE, post-training, or inference-optimal deployment. Re-estimate from learning curves. Scaling-law paper
Proposed pretraining mixture for experiments (not acquired): 35% licensed code/technical documentation, 25% quality web, 15% math/science, 10% legally usable books/reference, 10% multilingual including Hindi/Hinglish, 5% structured formats. Sample proportions by tokens and measure multilingual/code regressions. Separate post-training trajectories by repository and task lineage; never mix benchmark solutions into pretraining.
Implemented lab pipeline: explicit license/source/split → newline/NFC normalization preserving indentation → quality checks → credential-pattern rejection/email redaction → exact hash and shingle near-dedup → holdout-overlap rejection → declared domains and transparent heuristic score → byte tokens → checksummed NPY shards. The language hint is not industrial language ID. PII regex is incomplete. Near-dedup is quadratic and unsuitable for web scale. Corpus-wide scalable MinHash/LSH, learned quality/language models, provenance adjudication and mixture scheduler remain PLANNED.
Tokenizer: compare lossless bytes, byte-level BPE and unigram; preserve whitespace, arbitrary UTF-8 and structured syntax. Existing weights MUST retain their tokenizer unless embeddings are explicitly retrained. reports/tokenizer-*.json measures English/Hindi/Hinglish/Python/JSON/math/shell/XML samples. The 259 vocabulary is an exhaustive byte baseline, not an optimized natural-language vocabulary. A scratch BPE study should test 32K/64K/128K with equal corpus, embedding-adjusted FLOPs and held-out compression/loss. Eight examples cannot select an industrial vocabulary.
7. Training Pipeline
Implemented: seed-controlled AdamW training, cosine LR, gradient clipping/nonfinite checks, deterministic validation, atomic checkpoint pointer, SHA-256 validation, model/optimizer/CPU+CUDA/batch/Python/NumPy RNG recovery and safetensors export. Resume requires matching config and manifest. Existing state uses safe weights_only=True loading; load only trusted checkpoints.
R0: pretraining mechanics on tiny data. R1: assistant-token-masked SFT. R2: executable code/math/SQL verifiers in isolated environments. R3: rejection sampling from independently successful candidates. R4: outcome RL, group advantage primitive exists; full GRPO rollout/update pipeline PLANNED. R5: process supervision only when independently scored evidence beats outcome-only baseline. R6/R7: tool and long-horizon curricula require real successful trajectories. DPO and LoRA optimization are exercised on toy examples. PPO adds value/rollout complexity; defer until simpler methods plateau. No miniature result demonstrates improved general reasoning.
Coding examples must include repository snapshots, user request, investigation, failed tool receipts, minimal diff, compiler/tests and independently verified outcome. Split by repository/time, validate Python/JS/TS/C/C++/C#/Java/Go/Rust/SQL/Bash/PowerShell/HTML/CSS/Kotlin/Swift with native toolchains. Never give generated code a success label because its explanation sounds good. Keep verification code outside writable agent scope.
8. Coding Architecture
IMPLEMENTED: bounded source traversal, Python AST symbols/import edges and multilingual lexical indexing/retrieval; compare-and-swap file replacement via content hash; configured commands; final owner-selected verification; step/failure/repetition budgets. Unsupported languages are searchable text, not AST-aware. LSP integration, vector search, rich dependency graph, patch-hunk application and benchmark-grade repository repair remain PLANNED.
Target flow: task → repository map → retrieve relevant files → brief plan → minimal edits → build/test → inspect failures → repair → regression suite → diff and receipt-based report. Runtime tests use a scripted model solely to test state transitions. Real-model smoke results use an actual downloaded model and retain failures. Neither proves repository-scale coding capability.
9. Voice Architecture
Choose pipeline voice: local microphone/VAD → ASR → model/router → TTS → speaker. It isolates latency and allows replacing each model. Native audio models may improve prosody/full-duplex interaction but introduce a separate training/serving burden; no evidence here justifies building one.
IMPLEMENTED: async turn coordinator cancels stale generation on barge-in and invokes audio-stop callback; file ASR adapter using faster-whisper; OS TTS adapter via pyttsx3; WER and latency metrics. VALIDATED: cancellation with controlled async test doubles. Optional ASR/TTS hardware paths are not measured. Microphone capture, incremental ASR, VAD, acoustic echo cancellation, audio buffering/prosody and real full-duplex latency remain PLANNED. No fabricated WER or speech latency.
Track speech-end→transcript, transcript→first token, first token→first audio and total perceived latency separately. Collect consented English/Hindi/Hinglish recordings with interruptions and background noise. Pin ASR/TTS model versions and licenses. ASR implementation source
10. Agent Architecture
THINK is private model computation; exposed plans are concise action summaries. Runtime state is PLAN → ACT/OBSERVE → VERIFY → CONTINUE/FINISH. JSON Schema validates model actions and tool arguments. Unknown tools and extra arguments fail closed. READ/WRITE/EXECUTE permissions are enforced for implemented tools. NETWORK is an explicit inference opt-in; PRIVILEGED/IRREVERSIBLE are reserved classes with no implemented capabilities.
Filesystem tools resolve workspace containment, exclude private/control paths, bound reads and compare existing file hashes before overwrites. Commands are immutable owner-configured argument arrays, without shell expansion. Output capture is bounded; timeouts attempt process-tree termination. Receipts distinguish exit status, failure, elapsed time and truncation. Audit logs omit content and retain hashes. Per-process idempotency keys prevent duplicate execution within one executor; durable crash-safe transaction recovery is PLANNED. Do not automatically retry mutations. Read-only retries must remain bounded and observable.
The host command runner is not a sandbox. Enabling it runs code with the OS user's authority. A configured interpreter/test command can execute arbitrary repository code. Environment trimming does not isolate filesystem/network access. Use an external container/VM with no credentials, read-only hidden tests, restricted mounts and network policy before accepting hostile tasks. Concurrency/symlink race resistance requires OS-level primitives; path checks alone do not establish an adversarial security boundary. Browser/MCP/database integrations are PLANNED; do not expose arbitrary external actions through generic model strings.
11. Memory Architecture
L0 conversation in backend messages; L1 session context; L2 semantic records; L3 episodes; L4 projects; L5 preferences. SQLite stores layer/text/source/time/confidence/provenance/expiration, supports replacement by ID and deletion. Retrieval is lexical overlap, not embedding similarity; richer semantic retrieval is PLANNED. Avoid persisting unconfirmed inference as owner fact. TTL and confidence filters apply at retrieval. Audit deletion cannot promise removal from separately copied backups; define backup retention before production.
12. Inference Architecture
Local Transformers uses upstream KV caching and bounds input+output context. Small reference inference is intentionally simple. An HTTP chat adapter talks to explicit compatible endpoints, rejects non-HTTP URLs/embedded credentials, disables redirects/proxies, limits response size and requires remote-network opt-in. API credentials come from a named environment variable, never model messages.
Server-tier candidates: vLLM and SGLang for continuous batching, paged cache, prefix reuse, structured outputs and parallel inference; llama.cpp for CPU/GPU offload and GGUF on personal systems. Verify current engine/model/quantizer compatibility before deployment. vLLM's schema constraint support is evidence of an available backend feature, not proof NEXORA's local Transformers path constrains logits. Structured outputs
Router design: start with explicit fast/code/large endpoint routing, then calibrate an uncertainty policy on held-out failures. If 80% of requests use a cost-1 small model and 20% cost-10 large model, average model cost is 2.8 rather than 10 (72% arithmetic reduction), before verifier/retry overhead and quality loss. This is a scenario, not measured savings. Reject classifiers whose escalations conceal failure. No claim that a handcrafted prompt-length heuristic estimates difficulty.
Measure TTFT, TPOT, output tokens/s, requests/s, queue delay, P50/P95/P99, peak VRAM and failure rate under fixed prompts/concurrency. Report cold loading separately. FP32 baseline and CPU dynamic INT8 reference experiment are measured; BF16, FP8, INT4 and mixed precision task-quality suites remain PLANNED. Perplexity alone is insufficient.
13. Evaluation Framework
Published evidence: unit/integration tests, exact resume parity, checksum rejection, causal masking, tokenizer round trips, tool boundaries, timeouts, false-success prevention, memory expiry, toy post-training, local baseline inference, agent smoke, and reference quantization. All samples and limitations remain visible in reports/.
Required next suites: fresh repository repairs with hidden tests (SWE-bench-style isolation), execution-based multi-language tasks, verifiable math/science/SQL, multi-step agent target-state tests, long-context needles/multi-needles/cross-document synthesis/repo edits, voice WER/barge-in, injected file/web instructions and credential boundary tests. Keep hidden tests outside agent workspace and unavailable to the data pipeline. Pin test generation seed/date, dataset revision and harness source. Public samples have unknown contamination and are not a private holdout.
Report pass@1 and unbiased pass@k with sample counts, Wilson intervals for binomial rates, failure categories, environment and freshness. Verify compilation/execution rather than string similarity for coding. Vary reward attacks: modify tests, fabricate receipts, print expected output without state change, hide nonzero exit, trigger output truncation, exploit stale cache, manipulate a SQL oracle. Reject apparent success on these attacks. No benchmark-specific model branches or hardcoded solver answers.
14. Cost Model
Tier 1 (this workstation): tiny architecture training, small-model inference, runtime development, CPU experiments. Dense 120B scratch is economically and physically irrational. Tier 2 (8 GPUs, e.g. 80GB each): adapters and selected full/domain training of moderate models, 30B-class MoE serving. 640 GB aggregate is not 640 GB usable per rank; activations/KV/network matter. Tier 3 (8 nodes × 8): ~1B–7B scaling runs and research MoE, contingent on bandwidth and data. Tier 4 (32–128 nodes × 8): startup-scale pretraining only with product/data/learning-curve evidence and recovery drills. Tier 5 (128+ nodes × 8): 120B dense/large MoE feasible in principle, multi-million-dollar all-in program, not authorized here.
Distributed choice: PyTorch for prototype; FSDP2 for sharding supported dense models; Megatron-Core/NeMo for mature TP/PP/CP/EP at scale; DeepSpeed where its optimizer/offload integration wins measured benchmarks. Custom kernels only after profiles establish a bottleneck. SP often accompanies TP and is not another independent world-size factor.
Topology convention in calculator: world=TP×PP×CP×DP, EP divides the DP group. Examples: 1×8 GPUs=1×1×1×8; 8×8=2×2×1×16 with EP=8 and expert-DP=2; 128×8=8×8×2×8. Framework parallel folding/expert TP may use different decompositions; validate rank groups and divisibility against the actual stack. Megatron parallelism guide
Use high-bandwidth intra-node links for TP; scale-out all-to-all for MoE requires measured network bisection/latency. 200–400Gb/s per-GPU-class fabric is a planning target, not guaranteed sufficient. A 120B training checkpoint at 14 B/parameter is 1.68 TB; keeping three plus replicas consumes many TB. 2.4T uint32 tokens take 9.6 TB before raw text/intermediate copies. Add storage throughput sufficient to finish checkpoints within the tolerated pause; asynchronously saving without version consistency is unsafe. Activations depend on batch×sequence×layers×width plus attention strategy; profile representative steps instead of inventing a constant.
15. Prototype Roadmap
Current sprint produces executable training/runtime modules, an actual tiny checkpoint, actual inference tests and reproducibility artifacts. Fix failed tests before upload. Validate forced-kill recovery, record observed model limitations, and release without production claims. The smallest useful next experiment is a stronger model on the same fixed coding/tool harness, not immediately scaling the custom decoder.
16. Scaling Roadmap
| Phase | Entrance | Exit |
|---|---|---|
| 0 feasibility | Hardware/rights inventory | Scenarios, explicit budget and chosen path |
| 1 tokenizer | Provenance-cleared sample | Round trips, domain compression and loss study |
| 2 data | Allowed sources and private holdout boundaries | Reproducible shards, duplicate/PII/contamination audit |
| 3 small prototype | Validated model config | Stable gradients, loss decrease and recovery |
| 4 scaling laws | >=3 sizes/token budgets with repeated seeds | Fit+uncertainty predicts withheld run; no unstable routing |
| 5 base pretraining | Data/cluster/budget gate approved | Held-out loss target, recovery drills, complete lineage |
| 6 continual domain | Evidence of domain gap | Gain on private domain set without broad regression |
| 7 coding | Sandboxed multi-language trajectories | Hidden-test repository success improvement |
| 8 reasoning | Reliable verifiable tasks | Held-out accuracy/calibration improvement |
| 9 tools | Typed receipts and permission tests | Multi-step target-state success without boundary escape |
| 10 post-training/RL | Audited rewards and baseline | Generalization gain; reward hacks fail |
| 11 long context | Short-context gate | 32K→64K→128K→256K only when retrieval/synthesis/repo gates pass |
| 12 quantization | Unquantized reference | Noninferior task quality and measured resource win |
| 13 serving | Selected hardware/engine | Load test P95/P99, queue/memory limits and failure recovery |
| 14 voice | Consented recordings and audio hardware | WER, interruption and end-to-end latency targets |
| 15 red team | End-to-end deployment | Injection, boundary, reward and privacy tests pass |
| 16 release | Reproducibility and honest card | Immutable artifacts/checksums and clean install smoke |
| 17 improvement | Consented failure capture | Reproduce→minimize→test→train→re-evaluate→controlled release |
Do not extrapolate 820K results to 120B. Proposed 100M/1B/7B experiments remain PLANNED and require separate resources/data. No effective 32K+ context is claimed by this release.
17. Major Risks
Small pretrained models may produce invalid JSON and poor reasoning; rejection is preferable to fabricated success. Local CPU kernels are slow. Toy data overfits. MoE arithmetic savings can disappear into network inefficiency. QLoRA can still exceed 4GB after activations. Host execution is not isolated. License lists do not replace provenance due diligence. Short public tests have weak statistical power. Long context and native voice are unvalidated. Future dependencies may change; record tested versions and source hashes.
18. Repository Architecture
nexora/: model, tokenizer, data, training, adapters, posttraining, compute, evaluation, inference, agent, tools, coding, memory, voice and CLI. configs/: explicit experiment/permission files. examples/: original lab corpus. scripts/: reproducible experiments, research capture and release utilities. tests/: critical logic. docs/: decision, operations, scope and security. artifacts/: original tiny weights/checkpoints/data/experimental adapter. reports/: actual measurements, immutable source revisions and release manifest. .cache/: downloaded upstream weights, excluded from publication.
19. First Implementation Sprint
See docs/STATUS.md for the final evidence matrix. Changes are validated by tests rather than claiming every planned component exists. Publish only project-generated artifacts and source; exclude upstream weight caches, tokens, local private memory and machine-specific credentials. Provide commands for download, train, resume, inference, indexing, agent execution and re-running measurements.
20. Go / No-Go Gates
GO: publish the scoped research milestone when tests pass, checksums verify, actual experiment reports are included and the Hub upload is inspected. NO-GO: claim trained 120B, general coding competence, production sandboxing, reliable long context or full-duplex voice from this evidence. NO-GO: spend on scratch scaling before rights-cleared data, independent evaluation, a funded compute plan and kill/recovery tests exist. Only call a larger model a capability improvement after controlled evaluation.
Primary sources and evidence policy
Verified model metadata/configs and Apache-2.0 declarations are captured with revision hashes in reports/sources.json. Qwen cards report their own benchmark results; NEXORA does not treat those as local measurements. Additional sources: Qwen3.5-0.8B, Qwen3.5-4B, Qwen3-30B-A3B, PEFT LoRA, PyTorch distributed checkpoint.
Competitor lessons: Qwen provides reusable multilingual instruction weights; DeepSeek demonstrates sparsity/MLA/MTP at scale; Megatron offers tested parallel building blocks; vLLM offers serving/structured-output machinery. Transferring any of these into this constrained environment requires measurement. The custom reference model intentionally adds no unmeasured exotic architecture.