NEXORA-120B

A working local AI research prototype, with an original 820,736-parameter checkpoint. This is not a trained 120B model.

NEXORA means Neural EXecution, Optimization, Reasoning & Assistance. The long-term product ambition is a personal reasoning/coding/voice system. This release builds and tests the parts feasible on the available 16 GB RAM workstation; it clearly separates working code, measurements and future work.

What is included

  • A small PyTorch decoder with GQA, RoPE, RMSNorm and SwiGLU, trained for 120 steps on original synthetic text.
  • Checksummed data shards, actual weights, optimizer/RNG checkpoints and forced-kill recovery evidence.
  • Real miniature LoRA/SFT/DPO optimization and CPU INT8 comparison experiments.
  • Local upstream-model inference, an authenticated loopback streaming endpoint, and an optional external-model client.
  • Typed filesystem/command tools, explicit permissions, execution receipts, bounded agent loops and independent verification hooks.
  • Python AST/other-language lexical indexing, local memory lifecycle, failure triage and interruptible voice coordination.
  • Tests, reproducible experiment scripts and a researched architecture/cost/scaling assessment.

Evidence and limits

Measurement Result What it does not establish
Tiny training 820,736 parameters; validation loss 5.294 โ†’ 3.209 over 120 steps General language/coding/reasoning ability
Training throughput ~8,182 byte tokens/s on this CPU run, including eval/checkpoints GPU or large-model throughput
Recovery Forced worker kill after step 3; resumed step-9 weights bitwise equal to uninterrupted run Distributed/power-loss durability
Toy post-training SFT, DPO, adapter merge numerically exercised Capability improvement on independent tasks
Qwen3-0.6B smoke 0/4 strict tasks; file-agent task failed Broad model ranking
Qwen3.5-0.8B smoke 2/4 strict tasks; file-agent task failed Reliable autonomous coding
INT8 experiment Similar small-model CPU latency, validation loss 3.208 โ†’ 3.212 A demonstrated production optimization

The current local default is the unmodified upstream Qwen3.5-0.8B because it passed more of the same development smoke cases, with a small sample and no claim of statistically established superiority. Its weights are downloaded separately, attributed and pinned. It remains unreliable for agent tasks; failures are published in full. The original tiny checkpoint and its experimental adapter are not substitutes for a pretrained assistant.

Detailed reports: training, recovery, Qwen3.5 smoke, Qwen3 baseline, test results, post-training, quantization, environment.

The final test suite passed 36 tests, with 1 Windows symlink-privilege skip. After fixing capability advertisement, the real agent selected the permitted read tool but looped instead of completing the task; the loop guard stopped it. See the retained policy-aware follow-up. Streaming matched non-streaming output on one warm request, with first text at approximately 1.13 seconds; this is not a load-test percentile.

Measured training and CPU quantization results

Quick start

python -m pip install -e '.[inference,dev]'
python scripts/download_model.py
python -m nexora.cli chat "Explain how to verify a code repair."
python -m nexora.cli index . --query "checkpoint recovery"

The download uses the immutable upstream revision recorded in sources. The selected checkpoint includes vision components, but NEXORA only tests text inference. Local inference uses CPU by default and needs several GB of RAM. No paid cloud job is required or launched.

Train the tiny model:

python scripts/make_demo_data.py
python -m nexora.cli prepare examples/corpus.jsonl
python -m nexora.cli train --output artifacts/my-run
python -m nexora.cli generate "A reliable program" --checkpoint artifacts/my-run

PowerShell tests:

$env:PYTEST_DISABLE_PLUGIN_AUTOLOAD='1'
python -m pytest -q

See operations for resume, benchmarks, streaming HTTP, permissions and publishing. See architecture and economics for the requested 20-section decision, candidate comparisons, total versus active MoE parameters, lifecycle gates and primary sources. See status for implemented versus planned scope.

Model and data card

Original checkpoint: artifacts/tiny/model.safetensors, load with nexora.model.NexoraLM and artifacts/tiny/config.json. It is not an AutoModel checkpoint and does not claim Hugging Face inference-provider support. Architecture: 4 dense layers, width 128, FFN 384, 4 query/2 KV heads, head dimension 32, tied byte vocabulary 259. Maximum configured input 256 byte tokens; training sequences 128. Effective long context is untested.

Training data: 90 original synthetic records, 11,600 training byte tokens; three validation records, 424 tokens. Original examples are released under Apache-2.0 with the project. Provenance and split metadata are in the corpus/shard manifests. No scraped personal data, private chats, downloaded benchmark datasets or existing user's model caches were used for training. Validation is a development split, not a private capability benchmark.

Optimization: AdamW, seed 42, batch 8, 120 steps, cosine learning rate starting at 0.0005, clipping 1.0, CPU FP32. Training samples chunks with replacement and can pack across document boundaries. Checkpoints capture model, optimizer and RNG state. Lineage is recorded in training reports and the release hash manifest. The post-training adapter is an optional toy experiment on four SFT records/one preference pair and is not enabled by default.

Intended use: inspect, reproduce and extend training/runtime research. Out of scope: production autonomous execution, high-stakes decisions, claims of competitive 120B quality, reliable repository-scale coding or measured full-duplex speech. Native browser/MCP integration, distributed training, 32K+ effective context, production voice, large-model RL and large-scale quantization are planned.

Security and privacy

Model intelligence and executor authority are separate. Execution defaults off; remote inference requires opt-in. The host runner is not an OS sandbox. Only run trusted code there; hostile repositories require external isolation and immutable hidden tests. Local memory is plaintext and private by default. No telemetry endpoint is configured. Security boundaries

License and attribution

Original source, synthetic examples and tiny model artifacts: Apache-2.0. External models retain their original identity/license; no upstream weights are republished here. See NOTICE, LICENSE and pinned upstream sources. The model card does not claim endorsement by Qwen, Hugging Face, PyTorch or any other project.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support