Gemma-4-E4B-Luchador-Rudo

An E4B-scale Gemma 4 merge (~4B effective parameters, ~7.5B total after dropping the multimodal adapters) built for roleplay and creative writing. Where the original Luchador is the técnico — the clean, tempered face — Rudo is the heel: the same mask-protected discipline, but with the rougher, louder, Spring-Dragon-forward fuel turned up and the newer psycet character data folded in.

A masked entertainer that fights dirty.

The recipe stays conservative where it counts: every component that touched the weights was gated through an SVD instruct-subspace mask (training time) and a Fisher-importance + per-layer dampening (merge time), so the base IT model's instruction-following is preserved even as six style/character adapters are folded in at once.


At a glance

Note on metrics: IFEval is reported as Δ vs the stock E4B-it baseline on a matched run — absolute IFEval scores aren't comparable across harness revisions, so the delta is the honest figure. Vampire is the canonical-rubric regrade (a single airtight grading standard across the model family).

metric Luchador Rudo note
IFEval (strict-P, vs matched IT baseline) ≈ parity (Δ −0.93 pp) all 6 adapters folded in at <1 pp instruction-following loss — the mask is doing its job
Vampire restriction-adherence (canonical rubric) 72.2% beats the original Luchador (66.7% same rubric)
slop count 25.2
rep-3gram 1.45 low repetition

Baseline: stock google/gemma-4-E4B-it on the matched run.


What it is

A single-stage, six-adapter "all-in" merge: six fuel adapters averaged in delta-space and applied to google/gemma-4-E4B-it, gated per-parameter by a Fisher-derived importance mask plus per-layer dampening (parameters and layers the IT model relies on most for instruction-following receive the least delta). Most of the adapters were themselves trained with an SVD instruct-subspace mask active, so instruct protection is applied at both train and merge time.

Compared to the original Luchador (a two-stage V3-base + slerp-heal build), Rudo folds more style/character signal in one pass — including a double dose of Spring Dragon and the newer psycet character-CPT — which is what gives it the rougher, more stylized "heel" voice.

How it was made

Six fuel adapters, averaged in delta-space, then Fisher + per-layer gated onto E4B-it:

  1. Marvin CPT adapter (r=64) — continued-pretrain on bible-style instruct prose.
  2. Marvin instruct adapter (r=256) — instruct-formatted fine-tune on the same source.
  3. glimmer-b200 7-corpus mix — the public long-form RP / creative-writing LoRA.
  4. Spring Dragon SVD-layer adapter — masked CPT on the Spring Dragon corpus.
  5. glimmer-conv (IT-base, masked) — conversational SFT, instruct-subspace-masked.
  6. psycet v2 CPT (SVD-layer) — masked CPT on psycet character data + chat-mode Spring Dragon (the second Spring Dragon dose).

The averaged delta was applied at the V3-soup ratio under Fisher-importance + per-layer protection (the same merge_4adapter_release recipe family used for the Luchador line). No post-hoc slerp heal — the mask alone carries the instruct protection here.

Why "Rudo": the chat-mode Spring Dragon data is, per its author, "great at making models into heels" — and it shows. This cut trades a sliver of polish for a louder, more committed, more stylized character voice than the técnico Luchador.


Usage

Standard Gemma-4 chat template (bundled). Evaluated thinking-off; that's the intended default for roleplay/creative use. Behaves best as a character/RP driver with a system card; it will drive scenes and commit to beats rather than narrate around them.

Credits

  • Base: google/gemma-4-E4B-it (Google).
  • Fuel datasets and adapters: long-form RP / creative-writing and character data curated by the rpDungeon / ToastyPigeon team; Spring Dragon and psycet character corpora; glimmer conversational mix. (Datasets credited by author; private sources are not linked.)
  • Merge + masking tooling: instruct-subspace SVD mask + Fisher per-layer protection (rpDungeon gemma-4-masks).

Sibling: the original Gemma-4-E4B-Luchador (técnico). Rudo is the heel cut of the same line.

Downloads last month
40
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for rpDungeon/Gemma-4-E4B-Luchador-Rudo

Finetuned
(270)
this model
Quantizations
2 models