Aleph Differentiation, Parts 3 & 3-D: Two Laws, Five Days, One Framework

Community Article
Published July 20, 2026

The language line and the diffusion line ran in the same house style against four trunk classes — and the spine the program spent two years building held on every one of them: the same closed-form address, the same relay body, the same gates, the same optimizer, the same detach contract. Each arm also brought home a predictive law and buried comparative routing where it deserved to die. amoe-lora is the coalescence: the survivors as constructors, the falsifications as refusals, receipts in the error messages.

Part 1 established the addressing architecture on byte-level models. Part 2 moved it into a frozen pretrained trunk — Qwen2.5-0.5B — and certified the two-regime dispatch law with capacity controls. Then the program asked both remaining transfer questions inside the same five days. Part 3: does the law survive a different language substrate — a hybrid attention/SSM vision-language trunk whose recurrent state an adapter delta could in principle integrate into forever? Part 3-D: does any of the machinery survive leaving language entirely — a generative model with no token axis, a sigma axis instead, and three parameterizations of the same denoising problem?

Two four-day rented-GPU campaigns, thirty-five experiment packages (19 language, 16 diffusion), two full 2-seed programs. The joint answer, and the headline: the machinery transfers almost embarrassingly well. The same closed-form address, the same relay body, the same start-silent gates attach to a DeltaNet hybrid and to three denoiser classes with the math unchanged, and the detach contract verifies bit-exact on every one of them. Each campaign came home with a law that does something the program's earlier laws could not: predict, before spend, what used to be measurable only after. And each buried something honestly: the routing transfers only where the regime allows it — on diffusion, comparative routing not at all — and the record carries one formally published retraction, a corrected overclaim, and three falsifications with teeth, all shipped. What emerged from the two records is amoe-lora 0.2 — not a writeup of the campaigns but their point of coalescence: everything that survived, compiled; everything that died, encoded as a refusal.

Everything below links to a self-contained package in one of the two research records (language · diffusion) with raw ledgers and self-asserting results files. The production artifacts: the caption adapter, the curated aleph-diffusion-adapters, the amoe-lora framework, and the diffusion-pipe fork with the adapters baked in.

The continuities: what refused to change

The question both campaigns were built to answer — Part 2 asked it first — is which laws are substrate properties and which were artifacts of the beds that found them. The negatives below get their sections because they earned them; but the answer to the question is mostly a list of survivals, and it deserves its ledger before anything else.

  • The address. M̂ = Σₖ sinh(uₖ)·Aₖ / Σₖ cosh(uₖ) over 64 oriented half-axes — code-verified in the keystone, then carried unchanged through byte beds, GPT-2, Qwen2.5, the Qwen3.5 hybrid, two UNet classes, and a bf16 DiT. Zero revisions to the closed form, ever.
  • The relay body. Certified once in Part 2's architecture tournament (relay_pw, the first sub-14 relay), then attached essentially as-is to 24 hybrid blocks, 16 cross-attention sites, and 28 DiT blocks. The consumption law's placement default — spread the budget across every block, never concentrate — carried onto both new substrate families as the series default.
  • The gating discipline. Zero-init output weight and bias, gate at −3.0, adapters start silent and must earn amplitude — identical on every trunk in the program. Relay training has never diverged anywhere in the record (Part 2 logged zero divergences across every relay variant while the wide-MLP control lost a seed to NaN).
  • The toggle/detach contract. Bit-exact on every substrate in both campaigns — the single strongest joint certification the two records share, and the reason the product story is attach/toggle/detach rather than merge-and-hope.
  • Law-1. Architecture alone doesn't pay; differentiated pressure must — stated on the language line, reproduced on diffusion by exp008's monolith edge under uniform pressure.
  • The predictability principle. Established in Part 2, transferred whole to Part 3 as three independent data points — including the +91.6-perplexity cost of training against a capability the trunk already held at ceiling.
  • The probe-first method. Read the discrete surface before you build the mixture: the sign-code register probe predicted dispatch regime on language; its diffusion sibling found sigma as the only live register axis — and the banding architecture followed the probe, not the other way around.
  • The optimizer and precision discipline. Pure Adam, weight decay 0, fp32 judged gauges — the geometric-memory law, unbroken since it was named.
  • The no-selector doctrine. No selection event exists in any certified mechanism across the entire program: the dense signed dispatch on language and the cosine crossfade windows on diffusion are both the doctrine's constructive form. Where the diffusion arm probed routing anyway, the failure reproduced known program geometry — the raw high-dimensional address killing a dispatch key is the language line's disease geometry on a new organ.

That ledger is what "coalesced" means concretely: every default in amoe-lora is one of these rows, and every row has now held on at least two substrate families.

Four substrates, none of them bystanders

All four accepted the machinery; what differed is what each did with it. The language arm ran on Qwen/Qwen3.5-0.8B — a dual-tower VLM whose LLM tower is a hybrid: 18 Gated-DeltaNet linear-attention blocks and 6 full-attention blocks (indices 3,7,11,15,19,23), d=1024, a 248,320-token vocabulary with tied embeddings. Adapter deltas entering recurrent SSM state was a genuinely new regime — no competitive renormalization to bound a persistent offset — and it resolved cleanly: exp001 found no runaway drift within a 9k-token window (residual-norm ratio +2.3%→+4.0%, still rising at the window edge — flagged, not alarming), and all-24-block placement became the series default.

The diffusion arm ran on three substrate classes at once: SD15-Lune (an undertrained rectified-flow finetune — the cheap flow testbed), the stock SD1.5 core (epsilon objective, the primary target), and the Anima/Cosmos-Predict2 2B DiT (bf16, flow, 28 blocks, trained through the DeepSpeed pipeline fork rather than our own beds). Three parameterizations, two architectures, two dtypes, two trainers.

Put the two arms side by side and the fate of Part 2's gate-growth law is unmistakable. Relay gates grow on GPT-2 and on the stock SD1.5 core (0.047→0.070 — the trunk opts in); they shrink on the Qwen3.5 hybrid (while perplexity improves) and on the undertrained Lune. Four substrates, two votes each way: gate dynamics are a substrate diagnostic, not a constant — the mechanism survives, the number does not, and neither campaign smoothed that over.

Two more substrate lessons priced themselves in. The dtype law: attaching fp32 relays to a bf16 trunk spins fp32 noise into the environment, and the integration gate caught our own code doing exactly that — sniffing a first parameter that was still a meta tensor and reading fp32 off it. Adapters attach at the declared trunk dtype, never a sniffed one; judged gauges stay fp32. And the judging floors: the language trunk is a post-trained instruct model that captions natively and emits schema-valid JSON zero-shot, so every conditioning claim reads against that strong baseline; CLIP-L image-image cosine has a ~0.988 floor on pure-noise pairs, so every grounding number lives in a narrow band and the paired derangement-controlled design is load-bearing, not decorative. The diffusion zero-shot wall (exp000, exp000b) produced its own finding on the way: conditioning-format specialization is a double dissociation — the base model doubles its grounding on plain-English prompts exactly where the JSON-finetuned variant drops.

The deliverables first

A campaign earns its laws by shipping something that survives an audit. Both did.

Language (exp004): one always-on RelayPatchwork stack (6.26M params across 24 blocks) on ~4k fashion image–caption pairs. Held-out token-F1 0.408 → 0.706 (seed 0), 0.401 → 0.704 (seed 1) — a +0.30 distribution-match delta replicating across seeds within 0.005 — then the gauntlet: cross-judge |s0−s1| = 0.0019 (e004b); the checkpoint re-downloaded from the hub reproduces 0.706 exactly (e004c); the standalone relay_caption.py on a plain-transformers trunk — no project code — reproduces 0.706 / 0.7041 (e004d); a fresh machine on public files only captions correctly. And the honest caveat on the card: always-on, the adapter collapses the trunk's native depth-comparison JSON validity 1.0 → 0.0 (e004e). Use it for captioning; detach it for structured tasks. Detachment is free: the toggle law held bit-exact on the hybrid trunk (exp011: max|Δlogit| = 0.0 with all anchors off). Two smaller laws rode along — conditioning nearly doubles caption precision (0.356→0.70) while cutting the invented-attribute rate 0.20→0.14/0.10 across seeds (grounding is bought by tokens that are unpredictable without the image), and the predictability principle transferred whole.

Diffusion: the certified 16-slot relay (zero-init output weight and bias; gate −3.0; bit-exact code-path-skip toggle) went relay-favorable 3-for-3 across substrate classes, with matched controls every time. On Lune (exp001): paired val flow-MSE −0.85% vs frozen with round-trip grounding preserved (+0.2052 vs the wall's +0.2099), while the parameter-matched LoRA-r32 landed below frozen (+0.85% worse) and degraded grounding outright (→ +0.1694). On the stock SD1.5 core (exp006, 2-seed): relay −2.5%, LoRA behind it, and the relay gains grounding over the zero-shot core (+0.1809 → +0.2172) exactly where LoRA trades it away (+0.1596). On the Anima DiT (exp004, end-to-end through the diffusion-pipe fork, bf16): relay improves 0.123→0.115 while the fork's own native LoRA control degrades over training (0.138→0.142).

The cross-campaign contract underneath both deliverables: the toggle law held bit-exact on every substrate in both campaigns — a DeltaNet hybrid, two UNet classes, and a bf16 DiT — extending the contract the language line named. That contract, not any single benchmark, is why the production story is "attach, toggle, detach" and never "merge and hope."

The regime law (language)

The language campaign's pivotal mistake was running its image-task specialists (exp005: bbox grounding 0→0.69 valid / 0.89 IoU-score, depth-ordering 0→0.83, subject-fixation format 0.35→1.0) without measuring what each stack does to the other tasks. When the full sweep finally ran (exp018, retroactively, n=48), every specialist turned out to be a format monopolist: bbox kills fixation (0.0) and halves caption termination (0.56); fixation kills bbox and depth (0.0) and dents caption (F1 0.21 vs gauge 0.41); depth destroys captioning (token-F1 0.0014, zero terminations); caption kills depth JSON validity (1.0 → 0.0). Nothing in their training loss says so.

The dispatch is what rescues them — that is the point of it. exp017 put each specialist alone inside the trained dispatch (single-anchor masks on the exp007 collective: five anchors, 5000 steps, zero starvation alarms, training fingerprint replicated at a second seed — caption blend-dominance 0.779 vs 0.772): caption survives within 0.02 of gauge under every mask (0.389–0.406 vs 0.4076, termination 1.0), own-task validity is preserved for fixation and depth (bbox dips 0.69→0.44 valid at intact score). Against the exp018 destruction table, this is the strongest direct evidence in the program that the aleph denominator's damping genuinely contains specialize-regime anchors — the containment mechanism Part 2 flagged as a candidate law, now carrying weight on a second substrate. The reproduced cost: depth's content dilutes in-dispatch (0.826 solo → 0.123).

Math night supplied the other regime. It began with a weakness map (exp012: 10 domains × 3 tiers × 2 formats, internal control reproducing the earlier arithmetic ceiling at 1.0) — because exp006 had already trained a math anchor against a capability the trunk had at ceiling and bought nothing but a +91.6 perplexity tax. The experts (exp013) delivered the campaign's cleanest positive: derived-steps supervision beats direct-answer supervision — algebra with real worked derivations reaches 1.00/1.00 on provably-unseen held-out questions (direct: 0.79/0.57), replicating at a second seed to within 0.0001 of gain — plus the subtlest trap: two "striking" gains (sequences 0.21→1.00) dissolved under a held-out re-judge because the generators' question spaces (480 and 248 distinct questions) were smaller than the training draw (800/tier). An answer-space check does not protect against question-space exhaustion — now a law with a companion guard. And a replicated oddity: the always-on arith-steps anchor raises algebra off-domain by +0.375 at both seeds — the stepwise emission format itself transfers.

Then the basins collective (exp014: five frozen solo-trained math experts under a keys-only dispatch, preregistered P5a–d) measured the failure regime four ways: no surgical independence (own-domain removal deltas 0.04 / 0.00; usage purity 0.13–0.31; at the second key seed algebra is suppressed outright, 0.125 routed vs 0.958 solo); no damping (every expert fires at 0.86–1.6× on-domain amplitude on neutral prose against a ≥3× target); no composition, and worse — the bare trunk composes chain-of-thought at ceiling (1.0) and the full collective destroys it (0.21 / 0.0), with per-anchor attribution naming the tramplers. The perplexity side stays modest (+1.19, under the +2 flag) — the damage lives in task space, which is exactly why a wikitext gauge alone was never enough.

Assembled across the amplitude decomposition (exp011: specialists damped 5–11×; the blend-dominant caption anchor undamped at 0.100 vs the 0.102 ungated reference), P5d, e004e, and the e017/e018 pair:

The regime law (2nd substrate). Always-on single-task adapters on this trunk are mutually destructive. The aleph dispatch rescues specialize-regime anchors almost completely and fails blend-regime anchors entirely — and the amplitude decomposition names which anchors it will contain, anchor for anchor, before you deploy a roster.

Regime — not machinery — decides whether dispatch helps.

The conditioning law (diffusion)

The sigma axis gave the diffusion arm something the language line never had. The register probe (exp003) reoriented the whole question: caption-content registers separate negatively at every band and site — image content dominates the code surface — and sigma is the only live register axis, strongest at the deepest site. So instead of routing experts comparatively, assign them structurally along sigma.

exp008 built the mechanism — three band experts per site (LOW/MID/HIGH noise), combined by smooth cosine crossfade windows; positional gating, dense and differentiable, no selection event anywhere — and the lesion battery certified it: disabling a band's experts damages that band's validation 50–200× more than any other band's, 3-for-3 bands, both seeds. Band assignment manufactures true specialists. What it does not do by itself is pay on the aggregate objective — the matched monolith keeps a small edge under uniform pressure at both seeds, which is the language line's Law-1 (architecture alone doesn't pay; differentiated pressure must) reproduced on a new modality.

The payment lives where the instruments looked next. exp010 shipped the StepGatedSampler — per-step band gating at inference, experts activating coarse-to-fine across the trajectory exactly as trained — and judged in image space, where per-step eps-MSE turned out to be gauge-blind (0.2% moves vs these effects). The controller lifted round-trip grounding to +0.2135 vs the frozen trunk's +0.1247; band lesions cut grounding monotonically (LOW −0.016 / MID −0.031 / HIGH −0.035); the HIGH-band lesion's image-space signature is 14× LP-dominated — coarse-to-fine confirmed by lesion signatures, not narrative. On real fused data (exp011, 2-seed) the structural story replicates and the adapters pay 2–3× more than on synthetic rows. And exp012's role-aligned gauge (HIGH-band foreground-LP error on x0) revealed what every aggregate comparison had hidden: multiband beats the matched monolith by ~10% on HIGH-band foreground structure (0.1115 vs 0.1234; frozen 0.1555), both seeds — specialists earning exactly where the role map says they should, nowhere an aggregate-eps gauge can see.

exp012's actual assignment — blob supervision (segmentation-derived foreground masks rasterized to the latent grid, coupled to the HIGH band) on the eps substrate — returned an honest negative at both seeds (+0.03%/−1.0%, inert). The operator's hypothesis: not a failure of blob supervision but of the parameterization. Epsilon recovers structure through x0 = (x_t − sqrt(1−ᾱ)·ε̂)/sqrt(ᾱ), dividing by a vanishing sqrt(ᾱ) exactly at high noise — exactly where blob supervision applies. Flow recovers x0 = x_t − σ·v exactly and linearly at every sigma. exp013 ran the identical coupling on flow with a preregistered 5× bar:

The conditioning law (2-seed). Structural supervision pays where x0 recovery is linear and is inert where it is ill-conditioned. The identical blob coupling moved its gauge −5.9%/−3.7% on flow vs +0.03%/−1.0% on eps — a ~125–200× effect ratio, past the preregistered bar at both seeds, with the common gauge inside tolerance.

The three-point dose curve (λ = 0.5 / 1.0 / 2.0 → −5.9%@+0.2% / −8.3%@+0.4% / −8.4%@+0.9%) is monotone and saturating; λ≈1 captures the full blob gain inside the common-gauge bound and ships as the framework default. A bonus row with consequence: the undertrained Lune trunk pays adapters −9.1% vs frozen — the largest in the line — so adapter headroom scales inversely with trunk training. Stated caveat, standing: the flow-vs-eps contrast rides on different trunks; the v-pred same-trunk arm is queued, not run.

The two laws share a shape, and the shape is the campaigns' joint contribution: both are predictive before spend. The regime law reads the amplitude decomposition and names which anchors a dispatch will contain before you deploy the roster. The conditioning law reads the parameterization and names whether structure-level supervision has a linear path to pay before you buy a GPU-hour. The program's laws graduated from post-hoc descriptions to pre-flight instruments — which is precisely what made a framework possible.

Where comparative routing lived, and where it died

The unified verdict across both campaigns is sharper than either alone. The aggregation channel never creates specialization; at best, it contains specialization that already exists. Read as continuity rather than casualty, this section is the program's oldest doctrine holding its ground: no certified mechanism in five days contains a selection event, the survivor on diffusion — dense, positional, differentiable banding — is the doctrine's constructive form on its newest modality, and the falsifications land where the program's own geometry said they would.

On language it contains: exp017's rescue of the specialize-regime image anchors is the dispatch doing its one real job, and Part 2's sign-code probe predicts the regime in advance. Where the anchors arrive blend-regime — solo-trained, always-on provenance — the denominator cannot silence them (exp014, four instruments), and no dispatch imposed afterward will train selectivity into an expert that never had it.

On diffusion, comparative routing failed every honest test we could build, and the failures were as informative as the certifications:

  • exp007 (2-seed): a 4-expert bank with dense signed state+sigma dispatch ties/loses to a rank-matched monolith under uniform pressure, with flat sigma-band usage (spread 9e-4).
  • exp014 (2-seed): text in the dispatch key beats nothing — the hidden-state key already routes by prompt (usage variance 2.4e-2 with zero text input) — and the raw-flattened [32,128] address kills routing outright (1.4e-5): the language line's high-dimensional disease geometry, reproduced on a dispatch key. The bed also confessed its own null was wrong (a shuffled-key null measures diversity, not correctness) and rebuilt it.
  • exp015 (2-seed, corrected instruments): the M-hat slot-bottlenecked address key shows routing excess of 2.5e-06 over a repeated-key null and a match advantage of ~0. Address-as-key, in its frozen form, is falsified. The open form — a text-encoder adapter trained jointly with the dispatch — remains queued science.

The address itself is not dead: exp002 showed the frozen address appended to full text conditioning is redundant (real ≈ deranged) — but the address alone steers the UNet (+0.029 after 2k steps) and the trunk grew its gates to read it. Redundant-in-context, not dead; the redesign target is complementarity.

So the routing doctrine that survives five days and two modalities: differentiation must arrive from training regime (language) or be assigned structurally along a live axis (sigma banding on diffusion) — the dispatch is a container, never a differentiator. Where no live axis of pre-existing separation is available, comparative routing is not merely weak; it is falsified.

The framework: survivors as constructors, falsifications as refusals

amoe-lora 0.2 is the two records compiled, in the same proportion as the records themselves: the continuity ledger becomes the API surface, and the falsifications become its refusals. Train / attach / align / toggle / detach on a frozen trunk — 16-slot patch heads reading the residual stream through the closed-form sinh/cosh address over a 64-atom S³ codebook (~261k params/block at d=1024) — now on both substrate families: language trunks (amoe) and diffusion denoisers (amoe.diffusion: SD1.5-class UNets, SDXL, Cosmos-Predict2/Anima DiT). The runtime verbs and checkpoint I/O are implemented and invariant-tested on CPU CI — the toggle law and bit-exact detach as testable contracts, not documentation.

The campaigns are encoded four ways:

As constructors. Everything certified ships as a first-class option, not a footnote: adapter="multiband3" (the lesion-certified structural banding), the StepGatedSampler (the exp010-proven coarse-to-fine controller), λ≈1 as the blob default straight off the exp013 dose curve, and the campaign recipes transplanted as the reference trainers on both families. What survived is not documented — it is the default.

As diagnostics. amoe.diagnostics.diagnose is the regime law as a call: it warns per anchor when the on-domain/neutral amplitude ratio falls to ≤1.5 — the blend-escape gauge that exp011/exp014 proved you cannot read off training loss. amoe.train runs the question-space guard exp013 earned twice. Starvation safeguards ride align by default.

As refusals with receipts. On diffusion, ad.align() raises, with the exp007/014/015 record in the error message — comparative routing is a grounded negative, and the certified alternative it points to is adapter="multiband3", the structural banding whose lesions are surgical. blob=True with objective="eps" refuses unless forced — the conditioning law as a constructor argument, λ≈1 as its default. The dtype law attaches at the declared trunk dtype and keeps judged gauges fp32.

As house laws. Pure Adam only (laws.make_optimizer); dense signed dispatch with an all-anchor damped denominator — no top-k, no load-balancing, masking never renormalizes; zero-init output heads; fp32/TF32-off default; checkpoints carry the codebook home buffer. Each law's receipt is in the research record.

The productization is part of the record too: anchors save/load .safetensors in a canonical flat layout (blocks.{site}.{param}) with full meta as metadata, amoe-convert migrates every legacy campaign shape with bitwise verification, and a comfyui-amoe node package (loader / attach / band-toggle / detach, step gating installed as a pre-forward hook so any stock KSampler becomes step-gated) is the next-turn deliverable. The declared multi-GPU split is a paradigm, not an apology: native amoe.diffusion.train is single-GPU + DDP at framework level; the production multi-GPU path is the diffusion-pipe fork (DeepSpeed pipeline engine, aleph relays baked in, proven end-to-end on the Anima 2B — with the package's classes when installed, a vendored fallback when not, the refactor verified bit-exact against the exp004 checkpoint). Every adapter stack in the diffusion record now has a bitwise-verified safetensors companion, curated in aleph-diffusion-adapters; the campaign's legacy .pt checkpoints load directly through amoe.io.checkpoint. And the maturity table in the README says plainly which rows are certified-ported, which are reference-grade transplants, and which (SDXL site count) are pinned only at first real run — never guessed.

The process is the product

Both campaigns ran the same discipline under the same kind of four-day clock, and the discipline is what makes the laws citable. It is also the program's longest continuity — the one that makes all the others checkable. The language arm published a formal retraction of its own claim within hours of making it (the trainable anchor's "seed inversion" — two instruments, two truths, no lottery once the missing comparison cell actually ran) and separately corrected a pooled-overlap overclaim caught in review; its most important instruments — the cross-task sweep, the precision decomposition, the held-out re-judge — were built later than they should have been, and running them retroactively is what turned a good campaign into a defensible one. The diffusion arm preregistered its bars before spend and reported the misses; corrected and confessed a label inversion in a shipped README within the hour; confessed an instrument gap (reference arms missing role gauges) and rejudged; confessed a 5.7-hour idle gap in the ledger along with the operational law it re-earned (never end a turn on an intention — launch first, report after); and disclosed its final replication run as reserve spend before it ran.

The complete run records ship with the programs: the language campaign's campaign_ledger.jsonl and REPO_MAP.md; the diffusion campaign's 47 queue jobs, ~61 GPU-hours at ~97% in-window utilization, 27 run logs and the machine-readable ledger, archived with the program's research memory, every package carrying its slice. Thirty-five packages, each with scrubbed standalone code, a self-asserting results.json, and its raw ledger rows — every failed act and instrument bug included.

Five days ago this was two open transfer questions. It is now two predictive laws, a routing doctrine with a constructive answer on both modalities, three falsifications encoded as refusals, five verbs — and the same closed-form address that opened the program, attaching to its newest substrate family with the math unchanged. The negatives in this record are load-bearing because of what they protect: a continuity that has now held everywhere it was asked to.

References

External references follow the verified bibliography of Part 2; the directly load-bearing entries across both arms:

  • Ho, J., Jain, A., & Abbeel, P. (2020). Denoising Diffusion Probabilistic Models. arXiv:2006.11239
  • Song, J., Meng, C., & Ermon, S. (2020). Denoising Diffusion Implicit Models. arXiv:2010.02502
  • Rombach, R., et al. (2022). High-Resolution Image Synthesis with Latent Diffusion Models. arXiv:2112.10752
  • Podell, D., et al. (2023). SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis. arXiv:2307.01952
  • Lipman, Y., et al. (2022). Flow Matching for Generative Modeling. arXiv:2210.02747
  • Liu, X., Gong, C., & Liu, Q. (2022). Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow. arXiv:2209.03003
  • Salimans, T., & Ho, J. (2022). Progressive Distillation for Fast Sampling of Diffusion Models (the v-parameterization). arXiv:2202.00512
  • Hu, E. J., et al. (2021). LoRA: Low-Rank Adaptation of Large Language Models. arXiv:2106.09685
  • Houlsby, N., et al. (2019). Parameter-Efficient Transfer Learning for NLP. arXiv:1902.00751
  • Shazeer, N., et al. (2017). Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer. arXiv:1701.06538
  • Fedus, W., Zoph, B., & Shazeer, N. (2021). Switch Transformers. arXiv:2101.03961
  • Puigcerver, J., et al. (2023). From Sparse to Soft Mixtures of Experts. arXiv:2308.00951
  • Ji, Z., et al. (2023). Survey of Hallucination in Natural Language Generation. arXiv:2202.03629
  • Geirhos, R., et al. (2020). Shortcut Learning in Deep Neural Networks. arXiv:2004.07780
  • Lin, T.-Y., et al. (2014). Microsoft COCO: Common Objects in Context. arXiv:1405.0312
  • Qwen Team (2024). Qwen2.5 Technical Report. arXiv:2412.15115

Internal record: the 19 experiment packages under geolip-aleph-qwen-3.5-0.8b-instruct and the 16 under geolip-aleph-diffusion (each with code, raw ledger.jsonl rows, and self-asserting results.json; weights in .pt + verified .safetensors); the Part 1, Part 2, Part 3, and Part 3-D articles; the qwen-deepfashion-fused dataset; and the production repos (qwen3.5-0.8b-relay-caption, aleph-diffusion-adapters, amoe-lora, diffusion-pipe).

Community

Sign up or log in to comment