Plans11's picture
RESERVED Session 6 | 200,000 examples | crash-safe reservation
a27bcf6 verified
|
Raw
History Blame Contribute Delete
1.74 kB
metadata
license: mit
library_name: transformers
tags:
  - causal-lm
  - decoder-only
  - reasoning
  - deep-reasoning
  - recurrent-transformer
  - adaptive-computation
  - chain-of-thought
  - adaptive-reasoning

Nano.Deep.Reasoner.11m-HyperMini

An approximately 11,094,003-parameter decoder-only adaptive recurrent reasoning language model.

Architecture

  • Parameters: 11,094,003
  • Context: 1096
  • Vocabulary: 16,000
  • Hidden size: 240
  • Base Transformer blocks: 6
  • Attention heads: 8
  • Head dimension: 30
  • Intermediate size: 1072
  • Shared recurrent reasoning block
  • Learned latent reasoning memory: 8 tokens
  • Adaptive reasoning depth: 2-16
  • Gated recurrent memory updates
  • Adaptive halting controller
  • Verification head
  • Revision head
  • RoPE
  • Tied input/output embeddings
  • Padding-aware causal attention
  • Explicit padded-state suppression

Reasoning tokens

Training

Dataset:

Plans11/Organized_PreTrain_1k_Context

Each session contains up to 200,000 NEW examples.

Examples are protected by SHA-256 hashes.

Session reservations are committed before training so a hard Kaggle interruption cannot cause the same reserved examples to be selected again.

Resume safety

The checkpoint contains:

  • model.safetensors
  • optimizer.pt
  • rng_state.pt
  • training_state.json
  • example_ledger.json
  • tokenizer.json
  • tokenizer_config.json
  • token_id_manifest.json
  • config.json

Dataset fingerprint and tokenizer artifact hashes are verified before resume.

Current state

Completed sessions: 5

Unique examples reserved/trained: 200,000

Unique completed examples: 460,000

Global optimizer steps: 7,189

Last session loss: 0.17293200694084168