Text Generation
Transformers
Safetensors
English
ivme_conversate_s
from-scratch
experimental
custom-architecture
causal-lm
small-language-model
custom_code
Eval Results (legacy)
Instructions to use IvmeLabs/Ivme-Conversate-S-v1-Base with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use IvmeLabs/Ivme-Conversate-S-v1-Base with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="IvmeLabs/Ivme-Conversate-S-v1-Base", trust_remote_code=True)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("IvmeLabs/Ivme-Conversate-S-v1-Base", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use IvmeLabs/Ivme-Conversate-S-v1-Base with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "IvmeLabs/Ivme-Conversate-S-v1-Base" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "IvmeLabs/Ivme-Conversate-S-v1-Base", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/IvmeLabs/Ivme-Conversate-S-v1-Base
- SGLang
How to use IvmeLabs/Ivme-Conversate-S-v1-Base with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "IvmeLabs/Ivme-Conversate-S-v1-Base" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "IvmeLabs/Ivme-Conversate-S-v1-Base", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "IvmeLabs/Ivme-Conversate-S-v1-Base" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "IvmeLabs/Ivme-Conversate-S-v1-Base", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use IvmeLabs/Ivme-Conversate-S-v1-Base with Docker Model Runner:
docker model run hf.co/IvmeLabs/Ivme-Conversate-S-v1-Base
| pipeline_tag: text-generation | |
| library_name: transformers | |
| language: | |
| - en | |
| tags: | |
| - from-scratch | |
| - experimental | |
| - custom-architecture | |
| - causal-lm | |
| - small-language-model | |
| datasets: | |
| - HuggingFaceFW/fineweb-edu | |
| - allenai/soda | |
| - roneneldan/TinyStories | |
| - nampdn-ai/tiny-codes | |
| - b-mc2/sql-create-context | |
| - rahular/simple-wikipedia | |
| - nampdn-ai/tiny-orca-textbooks | |
| - HuggingFaceH4/no_robots | |
| - databricks/databricks-dolly-15k | |
| - microsoft/orca-math-word-problems-200k | |
| - openbmb/UltraInteract_sft | |
| - allenai/sciq | |
| model-index: | |
| - name: Ivme-Conversate-S-v1-Base | |
| results: | |
| - task: | |
| type: text-generation | |
| dataset: | |
| name: arc_easy | |
| type: allenai/ai2_arc | |
| metrics: | |
| - type: acc | |
| value: 0.268 | |
| - type: acc_norm | |
| value: 0.276 | |
| - task: | |
| type: text-generation | |
| dataset: | |
| name: wikitext | |
| type: wikitext | |
| metrics: | |
| - type: word_perplexity | |
| value: 5174.4 | |
| - type: byte_perplexity | |
| value: 4.95 | |
| - type: bits_per_byte | |
| value: 2.31 | |
| - task: | |
| type: text-generation | |
| dataset: | |
| name: blimp | |
| type: blimp | |
| metrics: | |
| - type: acc | |
| value: 0.592 | |
| license: cc-by-nc-sa-4.0 | |
| # Ivme-Conversate-S-v1-Base | |
| **Codename: Small Apple 1** | |
|  | |
| A sub-10M parameter language model trained from scratch as an experiment in | |
| extreme data efficiency: how much can a genuinely tiny, architecturally | |
| unusual model learn from a single, unrepeated pass over a small, deliberately | |
| diverse token budget? | |
| This is not a production model. It is a research artifact from one long, | |
| mostly-nocturnal debugging session, documented here in full — including the | |
| mistakes, because the mistakes are half of what makes the result legible. | |
| --- | |
| ## TL;DR | |
| - **9,545,840 parameters.** Sub-10M, on purpose. | |
| - **~836M training tokens, single epoch, no repetition.** The entire point | |
| of the experiment was testing what a tiny model learns from one clean pass | |
| over diverse data, not what it memorizes from many passes over less. | |
| - **12 domain sources** — web/edu text, dialogue, code, math, instructions, | |
| reasoning traces, science QA — deliberately diverse rather than a few | |
| large homogeneous corpora. | |
| - **Architecturally exotic on purpose**: factorized + untied embeddings, | |
| GQA, real DIFF Transformer V2 attention, nGPT hypersphere normalization, | |
| immediate block-wise weight sharing, learnable meta tokens, RoPE. | |
| - **Benchmarked with EleutherAI's lm-evaluation-harness**, strictly, no | |
| custom scoring logic: ARC Easy, WikiText2, BLiMP (all 67 subtasks). | |
| - Trained on an NVIDIA B300, benchmarked on an L40S, served here via Modal + | |
| Hugging Face `trust_remote_code`. | |
| --- | |
| ## Why this exists | |
| The starting question was simple: **is diversity a substitute for scale, at | |
| the very bottom of the parameter range?** Conventional wisdom for tiny | |
| language models (TinyStories, SmolLM, MobileLLM) leans toward simple, clean, | |
| high-signal text over broad diversity — diversity is usually treated as a | |
| mid-model-size lever, something you can afford once you have enough capacity | |
| to actually exploit it. This project deliberately tested the opposite bet at | |
| the smallest end of the range: 12 genuinely different domains, one epoch, | |
| under 10M parameters, and see what happens. | |
| A companion question, layered on afterward: **how far can you push an | |
| architecture away from the field's converged defaults before the tooling | |
| itself starts fighting you?** The answer, documented below in the incident | |
| log, is "further than you'd expect works at all, but every fused kernel, | |
| every `torch.compile` mode, and every auto-batching heuristic in the modern | |
| training stack quietly assumes you're using something closer to a standard | |
| Transformer." Going exotic on purpose means paying for it in engineering | |
| time, not GPU-hours. | |
| --- | |
| ## Architecture | |
| | Component | Choice | Why | | |
| |---|---|---| | |
| | Embeddings | **Factorized, untied** — separate small-rank input/output projections, not shared | Untied embeddings were a deliberate constraint from the start (avoiding the safetensors weight-tying headaches of the prior model, V2). Factorization (via a bottleneck rank `r=48`) is what makes untying affordable at this parameter budget — a naive untied embedding at vocab=8000 would have eaten the entire parameter budget on its own. | | |
| | Attention | **GQA** (4:1 query:kv head ratio) + **DIFF Transformer V2** | GQA was a fixed requirement from the outset. DIFF V2 was chosen deliberately for its documented training-stability and loss benefits over a standard Transformer — see the DIFF V2 section below for what "V2" specifically means here and why it matters. | | |
| | Normalization | **nGPT** — every vector (embeddings, attention/FFN outputs, residual stream) constrained to a unit hypersphere; the residual update itself is a "move along the sphere" operation controlled by per-channel learnable "eigen learning rate" scale parameters, rather than standard LayerNorm/RMSNorm | Chosen for its reported 4–20x reduction in training steps needed to reach a given loss — directly relevant when the entire training budget is a single, unrepeated epoch. | | |
| | FFN | SwiGLU, fused single gate+up projection | Standard, efficient, one fewer matmul than the naive two-projection form. | | |
| | Positional encoding | RoPE, applied at full head_dim | DIFF V2 doesn't split head_dim (unlike V1), so RoPE is applied normally, no half-dimension bookkeeping required. | | |
| | Depth | 14 unique transformer blocks, each executed twice (**immediate block-wise weight sharing**) → 28 effective layers of depth at the parameter cost of 14 | Motivated by MobileLLM's finding that depth-over-width is the more parameter-efficient lever for small models, and that immediate block-wise sharing recovers accuracy with no parameter cost. | | |
| | Register tokens | 4 learnable "meta tokens" prepended to every sequence, dropped before the output head | A cheap (a few thousand parameters), Hymba-inspired addition: gives the model a place to accumulate global context without burdening ordinary attention to do all of that summarization from scratch. | | |
| **Parameter breakdown:** vocab (factorized embedding + head) 792,576 · body | |
| 8,752,240 · meta tokens 1,024 · **total 9,545,840**. | |
| ### On DIFF Transformer V2, specifically | |
| This deserves its own note because getting it right — and initially getting | |
| it *wrong* — was one of the more instructive parts of this project. | |
| DIFF attention computes attention as the *difference* of two softmax maps, | |
| which cancels noise and produces measurably lower language-modeling loss than | |
| a standard Transformer at equal parameter count. The original formulation | |
| (V1) splits each attention head's dimension in half to form the two | |
| subtraction branches, which works, but means the resulting Q/K tensors don't | |
| share a dimension with V — an awkward shape for modern fused attention | |
| kernels (FlashAttention, SDPA's fused backends), which expect Q, K, and V to | |
| share a last dimension. | |
| **V2** solves this architecturally rather than by working around the shape | |
| mismatch: it doubles the number of *query heads* instead of splitting | |
| head_dim, keeps K/V unchanged, and uses a single fused attention call whose | |
| `2h` output heads are then split — critically, by **interleaving** | |
| (`heads[0::2]`, `heads[1::2]`), not by halving the head list. The reference | |
| implementation is explicit that halving is a "Wrong Implementation": paired | |
| heads must share the same GQA group (the same K/V), and under standard | |
| head-to-group assignment, interleaved heads do and halved heads don't. Get | |
| this backwards and training is measurably less stable. | |
| This model was, for a while during development, running V1's shape-split | |
| math zero-padded into V2's shape contract — mathematically valid, verified | |
| exact, and genuinely functional, but not actually V2's real mechanism, and | |
| paying real overhead (two attention calls instead of one) for no benefit. | |
| Rebuilt correctly before the final training run: one fused call, no padding, | |
| natively FlashAttention-compatible, λ as a per-token per-head | |
| sigmoid-projected value rather than V1's global exponential form. | |
| --- | |
| ## Data | |
| **~836M tokens, one epoch, twelve sources**, deliberately mixed rather than | |
| dominated by one or two large corpora: | |
| | Domain | Source | Notes | | |
| |---|---|---| | |
| | General web/educational | FineWeb-Edu (`sample-10BT` config) | Largest single slice, ~30% of budget — the breadth anchor, kept a minority share on purpose. | | |
| | Dialogue | SODA | All dialogue routed through SODA alone after `daily_dialog` and `facebook/empathetic_dialogues` turned out to be unloadable — see incident log. | | |
| | Narrative | TinyStories | Small slice, kept modest to avoid duplicating what a different model in the same family already leaned on heavily. | | |
| | Code | `nampdn-ai/tiny-codes` (gated) + `b-mc2/sql-create-context` | | | |
| | Encyclopedic | `rahular/simple-wikipedia` | | | |
| | Explanatory/textbook | `nampdn-ai/tiny-orca-textbooks` (gated) | **Only the `textbook` field was used** — a probe of the raw data found the `question`/`response` fields in this dataset were frequently topically unrelated to the textbook content itself (e.g. a "problem-solving scenarios" textbook paired with an unrelated movie-plot question), so those fields were excluded rather than trained on as noise. | | |
| | Instructions | `HuggingFaceH4/no_robots` + `databricks-dolly-15k` | | | |
| | Math | `microsoft/orca-math-word-problems-200k` | | | |
| | Reasoning | `openbmb/UltraInteract_sft` | Rows treated independently; the dataset's `parent_id` tree structure wasn't resolved during packing. | | |
| | Science | `allenai/sciq` | Multiple-choice distractor fields explicitly excluded from the training text — only the question, correct answer, and supporting explanation were used, to avoid training on unlabeled wrong answers. | | |
| **Tokenizer:** custom 8,000-token byte-level BPE, trained on the collected | |
| mix itself (not reused from an existing model) — deliberately compact, since | |
| at this parameter scale the vocabulary/output-head cost is the single | |
| largest lever on how much of the parameter budget is left for the | |
| transformer body itself. | |
| **Packing:** flat, memory-mapped token stream, chunked into fixed-length | |
| windows, shuffled once into a single training permutation with no | |
| overlapping or repeated windows — one token, seen exactly once, with no | |
| mechanism by which the training loop could double back over data already | |
| covered. | |
| --- | |
| ## Training | |
| | | | | |
| |---|---| | |
| | Hardware | NVIDIA B300 (via Modal), single GPU | | |
| | Precision | bf16 autocast | | |
| | Attention kernel | Pinned FlashAttention-2 (via Hugging Face `kernels`), with an explicit compute-capability gate that falls back cleanly to unrestricted SDPA on hardware below Ampere | | |
| | Compile | `torch.compile`, regional (each of the 14 unique blocks compiled once, reused via the block-sharing execution order), `mode="reduce-overhead"` | | |
| | Optimizer | Muon (body matrices, 2D+) + AdamW (embeddings, norm scales, biases) — split via `torch.optim.Muon`, now native to recent PyTorch | | |
| | Data residency | Entire ~836M-token packed corpus loaded onto GPU VRAM once at startup (under 1% of a B300's 288GB), eliminating per-batch host transfer for the whole run | | |
| | Batch size | Auto-probed at startup via a doubling-then-binary-search OOM sweep, with a safety margin | | |
| | Epochs | **Exactly one.** Enforced structurally — the training loop walks a single fixed permutation and stops at the end rather than wrapping around, so it cannot silently repeat data even under a scheduling miscalculation. | | |
| | Throughput | Settled around ~230K tokens/sec at steady state after warmup/compile overhead | | |
| | Wall-clock | ~55–65 minutes for the full epoch once the pipeline was fully debugged | | |
| --- | |
| ## Results | |
| Evaluated with **EleutherAI's `lm-evaluation-harness`, strictly** — via the | |
| officially-supported path of wrapping this model in a minimal | |
| `transformers.PreTrainedModel` shim and passing it directly to `HFLM`, so | |
| every actual loglikelihood computation, batching decision, and metric | |
| aggregation is the harness's own tested code, not a reimplementation. 0-shot | |
| throughout. Run on an L40S with auto-batching. | |
| ### ARC Easy | |
| | Metric | Score | | |
| |---|---| | |
| | `acc` | 26.8% | | |
| | `acc_norm` | 27.6% | | |
| For context: random chance on a 4-option multiple-choice task is 25%. This | |
| model is barely above chance on ARC Easy — a genuine, honest result for a | |
| 9.5M-parameter, single-epoch model. ARC Easy requires a level of factual/ | |
| scientific-reasoning generalization this model's capacity and training | |
| budget simply weren't built to reach. | |
| ### WikiText2 | |
| | Metric | Score | | |
| |---|---| | |
| | Word perplexity | 5,174 | | |
| | Byte perplexity | 4.95 | | |
| | Bits per byte | 2.31 | | |
| High word-level perplexity is expected here for a structural reason, not | |
| just a capability one: the model's vocabulary was trained on this project's | |
| own 12-source mix, not on WikiText2's specific text distribution, and | |
| perplexity is highly sensitive to vocabulary/domain match. Byte-level | |
| perplexity (which is vocabulary-independent) is the more informative number | |
| of the two for a model with a from-scratch, non-standard tokenizer. | |
| ### BLiMP (all 67 subtasks, aggregate) | |
| | Metric | Score | | |
| |---|---| | |
| | `acc` (aggregate, `sample_count=67,000`) | **59.2%** | | |
| This is the most interesting result of the three, because the per-subtask | |
| breakdown is legible rather than uniform: | |
| - **Near-ceiling** on several subtasks: `principle_A_case_1` (99.8%), | |
| `sentential_negation_npi_licensor_present` (99.0%), | |
| `principle_A_domain_1` (98.3%), `wh_questions_subject_gap_long_distance` | |
| (98.2%). These cluster around **local binding/agreement and simple | |
| long-distance dependency patterns** — the kind of structure that shows up | |
| constantly, in a short window, across almost any register of English text. | |
| - **Near-floor** on a distinct cluster: `only_npi_licensor_present` (3.0%), | |
| `matrix_question_npi_licensor_present` (3.0%), | |
| `only_npi_scope` (16.7%), `wh_vs_that_with_gap_long_distance` (8.7%). | |
| These are almost entirely **negative polarity item (NPI) licensing** | |
| phenomena — a genuinely subtle syntactic dependency that requires tracking | |
| a licensing context across a clause, not just local agreement. | |
| Read together: the model learned real local grammatical structure — subject/ | |
| verb agreement, reflexive binding, some long-distance filler-gap patterns — | |
| robustly, from a single pass over diverse data. It essentially did not learn | |
| NPI licensing at all. That's a specific, falsifiable, and genuinely | |
| interesting finding about what a tiny, single-epoch, diversity-first model | |
| generalizes and what it doesn't, rather than a vague "it's a small model so | |
| of course it's bad at everything" shrug. | |
| ### Qualitative sample | |
| > **Prompt:** "Once upon a time" | |
| > | |
| > **Continuation:** "...you can be excited that her friends: he was a happy | |
| > total with an sorry most little a is and let from the number of fors. | |
| > Mivean: Just get your time about what you can was you in from, to be many | |
| > in the time, in which time at the number of: Rotes: That? Ining a likeing, | |
| > her person and" | |
| Locally plausible (consistent capitalization of name-like tokens, dialogue- | |
| style colon formatting clearly picked up from the SODA/no_robots portions of | |
| the training mix, roughly English sentence rhythm) and globally incoherent — | |
| exactly consistent with the BLiMP findings above and with the model's | |
| perplexity: it has learned surface statistics and some local syntax, not | |
| compositional semantics. | |
| --- | |
| ## What broke, and what that says about the field | |
| This section exists because the debugging process turned out to be as | |
| informative as the results — most of the field's tooling (FlashAttention, | |
| `torch.compile`, SDPA's backend selection, `lm-evaluation-harness`'s `HFLM`) | |
| is built around standard-architecture assumptions, and pushing outside them | |
| surfaces real, specific incompatibilities rather than vague friction. | |
| - **DIFF attention's shape contract genuinely fights fused kernels** unless | |
| implemented as V2 specifically intends (see architecture section above). | |
| The V1-style workaround was mathematically valid but paid a real, | |
| measurable performance tax for it. | |
| - **A GPU-resident dataset optimization introduced a genuine PyTorch | |
| kernel-coverage gap**: CUDA's advanced-indexing kernels don't support | |
| `uint16` tensors (`"index_cuda" not implemented for 'UInt16'"`), even | |
| though basic transfer operations do. Fixed by storing the token array as | |
| `int64` on GPU instead — a 4x memory cost that's still trivially small in | |
| absolute terms (under 3% of a B300's VRAM) for a corpus this size. | |
| - **`SDPBackend.FLASH_ATTENTION` is a generic label, not a pinned kernel | |
| version** — PyTorch's dispatcher silently resolves it to whatever | |
| FlashAttention generation the GPU's compute capability supports, which on | |
| Blackwell-class hardware meant FA4, empirically measured to *regress* | |
| throughput for this model's specific shape profile (small head_dim, small | |
| batches — the opposite of what FA4's warp-specialization and TMEM | |
| pipelining are tuned for). Fixed by pinning FlashAttention-2 specifically | |
| via Hugging Face's `kernels` library, with an explicit device | |
| compute-capability gate so the same code correctly falls back to | |
| unrestricted SDPA on hardware (like a T4) below FlashAttention's Ampere+ | |
| floor entirely. | |
| - **A silent, autocast-related dtype leak**: several raw `nn.Parameter` | |
| tensors in the model (the register/meta tokens, and the nGPT | |
| "eigen-learning-rate" scale parameters) never pass through an | |
| autocast-eligible operation, so multiplying them against bf16 activations | |
| silently promotes the *result* back to fp32 — with no error anywhere in | |
| the chain, only surfacing when that fp32 tensor eventually reached a | |
| kernel with a hard bf16-only assertion. The fix that actually held was an | |
| explicit, unconditional dtype cast immediately before that kernel call, | |
| rather than chasing every individual upstream leak point. | |
| - **The single most consequential bug**: an early version of the nGPT weight | |
| renormalization step applied the hypersphere constraint to *every* linear | |
| layer in the model, including the final output/vocabulary projection. | |
| Constraining that layer's weight norm directly caps the maximum logit | |
| magnitude the model can ever produce for any single token — which caps how | |
| confidently softmax can ever predict anything — which puts a hard, | |
| unmovable floor on achievable loss. This was diagnosed by the standard | |
| sanity check of confirming the model could trivially overfit a tiny fixed | |
| batch (it couldn't, capping at a stubborn ~2.1 loss no matter how long it | |
| trained); excluding the output head and token embedding from | |
| renormalization restored full learning capacity immediately. This is the | |
| reason the model that produced the results above needed a full retraining | |
| run partway through the project. | |
| - **`torch.compile(mode="max-autotune")` was tried and reverted** based on | |
| the run's own evidence: its Triton-kernel autotuning search consistently | |
| found nothing faster than plain cuBLAS `mm` at every matmul shape tested, | |
| while still paying minutes of upfront compile cost for the search — a net | |
| loss for a single-epoch run where that cost is never amortized. | |
| - **`from_pretrained()`'s fast/meta-device init path silently skips | |
| non-persistent buffer computation**, which left this model's RoPE cache as | |
| uninitialized memory (NaN logits) on every reload until the buffers were | |
| made persistent instead — a documented `transformers` behavior, not a bug | |
| in this codebase, but one that would have made the published model | |
| unloadable if it hadn't been caught before pushing. | |
| None of these were guesses that happened to work. Each had a measured | |
| number, a real error message, or a directly reproduced failure behind it — | |
| and several were tried, found wrong or insufficient, and revised at least | |
| once. That iteration is the actual cost of the architecture being genuinely | |
| exotic rather than a light reskin of a standard Transformer. | |
| --- | |
| ## Limitations | |
| - **No KV-cache.** `forward()` recomputes attention over the full sequence | |
| on every call. `.generate()` works but scales roughly quadratically with | |
| output length rather than linearly — fine for short samples, not built for | |
| long-form serving. | |
| - **Single epoch, no repetition, by design.** This model has seen each | |
| training token exactly once. It has not been given the opportunity to | |
| reinforce patterns through repeated exposure the way most small-model | |
| training recipes do. | |
| - **8,000-token custom vocabulary**, trained on this project's own data mix. | |
| Perplexity comparisons against models using larger, more standard | |
| vocabularies (GPT-2 BPE, etc.) are not directly comparable at the | |
| word-level; the byte-level numbers above are the fairer cross-model | |
| comparison point. | |
| - **NPI licensing and similar long-range/scope-sensitive syntactic | |
| phenomena are essentially unlearned**, per the BLiMP breakdown above. This | |
| is a specific, known gap, not a general disclaimer. | |
| - **This is an experimental architecture with no prior published | |
| implementation combining all of its pieces** (factorized+untied | |
| embeddings, GQA, DIFF V2, nGPT, block-sharing, meta tokens, together). | |
| Treat it as a research artifact, not a production-hardened model family. | |
| --- | |
| ## Usage | |
| ```python | |
| from transformers import AutoModelForCausalLM, AutoTokenizer | |
| model = AutoModelForCausalLM.from_pretrained( | |
| "ivmelabs/Ivme-Conversate-S-v1-Base", trust_remote_code=True | |
| ) | |
| tok = AutoTokenizer.from_pretrained("ivmelabs/Ivme-Conversate-S-v1-Base") | |
| ids = tok("Once upon a time", return_tensors="pt").input_ids | |
| out = model.generate(ids, max_new_tokens=80, do_sample=True, temperature=0.8, top_k=40) | |
| print(tok.decode(out[0])) | |
| ``` | |
| `trust_remote_code=True` is required — this is a genuinely custom | |
| architecture, not one of `transformers`' built-in model classes. | |
| --- | |
| ## Acknowledgments | |
| Architecture choices draw on: MobileLLM (depth-over-width, immediate | |
| block-wise weight sharing), the DIFF Transformer V2 work from | |
| Microsoft/UniLM, nGPT (NVIDIA), Hymba (meta/register tokens), and the general | |
| small-language-model efficiency literature (TinyStories, SmolLM, the | |
| BabyLM challenge, and the L20-Edu-135M single-GPU training study, which | |
| served as a useful real-world throughput benchmark during development even | |
| though it uses a substantially more conventional architecture than this | |
| model does). |