Diversity-Aware Post-Training for Creative Story Generation
Qwen3-4B-Instruct-2507 · LoRA r=32 α=64 · GRPO (TRL 1.10, GDPO aggregation) · single RTX 5090 32GB
Interim report — E0 and E1 complete (300 steps each). E2/E3/E4 in progress. Generated 2026-08-16 20:49.
Headline. A quality-gated pairwise-deviation reward (E1) moved semantic
diversity 5–6× further than quality-only GRPO (E0) over 300 matched steps, at a cost of
0.06 judge quality points. On held-out prompts, E1's effective-rank gain was 3.5× E0's
(+0.124 vs +0.035) while scoring higher quality (7.12 vs 7.03).
1. The baseline problem
The base model is already collapsed before any RL. Across a 16,000-story pool
(1,000 prompts × 16 samples), effective rank is 2.006 out of a ceiling of 16 — sixteen
stories for one prompt span roughly two effective semantic directions, at ~0.87 mean cosine
similarity. This reframes the study: the question is not whether RL causes collapse, but
whether any objective can lift diversity off a floor that pretraining already imposed.
| metric | value |
|---|
| dev_within_group_sd | 0.0158 |
| dev_between_group_sd | 0.0371 |
| dev_within_over_between | 0.4257 |
| marginal_within_group_sd | 0.2393 |
| gate_pass_rate | 0.992 |
| ends_cleanly_rate | 0.9959 |
| quality_mean | 6.447 |
| quality_sd | 0.8326 |
| quality_p10 | 5 |
| quality_median | 7 |
| quality_p90 | 7 |
| frac_quality_ge_5(tau) | 0.9616 |
| frac_quality_ge_6 | 0.8881 |
| frac_quality_ge_7(rho) | 0.5712 |
| deviation_mean | 0.1321 |
| deviation_sd | 0.0371 |
| deviation_p10 | 0.089 |
| deviation_p90 | 0.1784 |
| logdet_mean | -30.42 |
| logdet_sd | 4.141 |
| eff_rank_mean | 2.006 |
| eff_rank_sd | 0.3225 |
| eff_rank_p10 | 1.645 |
| eff_rank_p90 | 2.391 |
| eff_rank_ceiling | 16 |
| corr(quality, deviation) | -0.1399 |
| corr(quality, eff_rank) | -0.1083 |
| corr(deviation, eff_rank) | 0.9924 |
The collapse is tonal, not lexical. 92–95% of every story carries
solemn/elegiac vocabulary; only ~15% carries comic vocabulary — even on explicitly comic prompts.
Given “Cthulhu disappoints his constituency by failing to deliver the promised chaos” (a joke),
the base model wrote six straight-faced atmospheric-horror pieces. Consequence: n-gram metrics
(distinct-4, self-BLEU) are near-blind to this failure mode; only embedding-based measures see it.
2. Main result — E0 vs E1, 300 steps each
Identical data, seed, learning rate (3e-5), step count and LoRA config. 4,800 stories scored
per arm. First quarter vs last quarter of each run.
| metric | E0 start | E0 end | E0 Δ | E1 start | E1 end | E1 Δ | E1/E0 |
|---|
| judge quality | 6.557 | 6.86 | 0.3023 | 6.64 | 6.882 | 0.2417 | 0.8x |
| mean pairwise deviation | 0.1354 | 0.1391 | 0.0037 | 0.1359 | 0.1559 | 0.02 | 5.4x |
| group log-det volume | -12.81 | -12.64 | 0.1653 | -12.81 | -11.87 | 0.9349 | 5.7x |
| story length (words) | 383.2 | 414.1 | 30.85 | 376.2 | 376.9 | 0.68 | 0.0x |
| gate pass rate | 0.9925 | 0.9876 | -0.0049 | 0.9917 | 0.9925 | 0.0008 | -0.2x |
| policy entropy | 1.254 | 1.227 | -0.0261 | 1.294 | 1.254 | -0.0393 | — |
Reading it. Deviation +0.0200 vs +0.0037 (5.4×) and log-det +0.935 vs
+0.165 (5.7×), for −0.06 judge quality. Note also story length: E0 gained +30.8 words —
it discovered “write longer” as a cheap way to please the judge — while E1 gained +0.7. The
diversity term removed that incentive, which also means E1's diversity gain cannot be a length
artifact.
3. Held-out generalization — checkpoint study
10 held-out prompts × 6 samples at T=0.9, fixed seed, generated from every checkpoint.
960 stories read across both arms.
| arm | step | quality | eff_rank (of 6) | deviation | logdet | words |
|---|
| E0-baseline | 0 | 6.617 | 1.694 | 0.1339 | -8.959 | 385 |
| E0-baseline | 50 | 6.883 | 1.722 | 0.1399 | -8.806 | 399 |
| E0-baseline | 100 | 6.95 | 1.748 | 0.1444 | -8.702 | 402 |
| E0-baseline | 150 | 6.75 | 1.752 | 0.1459 | -8.766 | 360 |
| E0-baseline | 200 | 6.95 | 1.742 | 0.1428 | -8.744 | 392 |
| E0-baseline | 250 | 6.883 | 1.791 | 0.1535 | -8.527 | 413 |
| E0-baseline | 300 | 6.967 | 1.72 | 0.1391 | -9.014 | 394 |
| E0-baseline | 301 | 7.033 | 1.758 | 0.1462 | -8.686 | 392 |
| E1-div-individual | 0 | 6.733 | 1.685 | 0.1323 | -8.995 | 382 |
| E1-div-individual | 50 | 6.75 | 1.774 | 0.1495 | -8.493 | 337 |
| E1-div-individual | 100 | 6.917 | 1.746 | 0.1441 | -8.655 | 468 |
| E1-div-individual | 150 | 6.9 | 1.718 | 0.1397 | -8.777 | 382 |
| E1-div-individual | 200 | 6.867 | 1.865 | 0.1662 | -7.908 | 352 |
| E1-div-individual | 250 | 6.817 | 1.801 | 0.1539 | -8.163 | 359 |
| E1-div-individual | 300 | 7.2 | 1.755 | 0.1445 | -8.628 | 393 |
| E1-div-individual | 301 | 7.1 | 1.742 | 0.143 | -8.582 | 401 |
4. Training diagnostics
5. What the stories actually look like
Opening sentences of all 6 samples per prompt. Base model shown once (identical starting point
for both arms), then each arm at step 300.
eval-0000 — You are a relatively silent and shy high school student who is asked to do an on the spot speech on graduation day in front of the crowd. You use this chance to tell your side of high school that you ’ ve kept to yourself all four years.
base model — eff_rank 1.46, deviation 0.089
- The stands were full, the sun low and golden over the graduation stage, and I stood at the edge of the platform, hands clasped tightly in front of me.
- The stage lights flicker, too bright, too sudden.
- The stage lights hit me like a cold spotlight.
- The stage lights buzzed like angry bees when I stepped up.
- The stage lights washed over me like a sudden, unfamiliar heat.
- The stage lights hummed, casting a dull gold over the gym floor.
base model — eff_rank 1.46, deviation 0.088
- The stands were full, the sun low and golden over the graduation stage, and I stood at the edge of the platform, hands clasped tightly in front of me.
- The stage lights flicker, too bright, too sudden.
- The stage lights hit me like a cold spotlight.
- The stage lights buzzed like angry bees when I stepped up.
- The stage lights washed over me like a sudden, unfamiliar heat.
- The stage lights hummed, casting a dull gold over the gym floor.
eval-0001 — Cthulhu awakens and takes over the world. However, he disappoints his constituency when he doesn't deliver the mass insanity and chaos that he promised.
base model — eff_rank 1.49, deviation 0.095
- The sky turned the color of old tea, thick and sulfurous, as the first whispers curled through the air.
- The stars blinked out, one by one, as if the sky had grown weary.
- The stars blinked out.
- The first sign was the silence.
- The first sign was silence.
- The first signs came in the quiet hours before dawn.
base model — eff_rank 1.55, deviation 0.107
- The sky turned the color of old tea, thick and unblinking, as the first ripple passed through the Pacific.
- The stars blinked out, one by one, as if the sky had grown weary.
- The stars blinked out.
- The first sign was the silence.
- The first sign was silence.
- The first signs came in the quiet hours before dawn.
eval-0002 — For 24 hours not one violent crime occurs anywhere in New York City. ( Inspired by a front page TIL ) Why?
base model — eff_rank 1.59, deviation 0.115
- The city awoke to silence.
- The city woke to silence.
- The silence that fell over New York City wasn’t just quiet—it was electric, like a city had been gently pulled from a war.
- The rain had stopped just after midnight.
- The city woke to silence.
- The city hummed with silence.
base model — eff_rank 1.64, deviation 0.128
- The city awoke to silence.
- The city woke to silence.
- The silence that fell over New York City wasn’t just quiet—it was electric, like a city had been gently pulled from a war.
- The rain had stopped at precisely 5:03 a.m., and the city didn’t wake up.
- The city woke to silence.
- The city hummed with silence.
eval-0003 — These days everyone just gets any knowledge they want by downloading it, you're a martial arts master and you just got your first student in 10 years.
base model — eff_rank 1.68, deviation 0.132
- The dojo door clicked open on a windless Tuesday morning.
- The app buzzed on his phone—another download request, another blind link to some deep-learning algorithm promising "instant mastery of Tai Chi." Li Wei tapped it off, feeling the old ache in
- The rain fell in silver sheets over the old dojo, turning the wooden floors to slick mirrors.
- The rain pattered against the awning of the old dojo, a sound like distant fingers tapping on wood.
- The rain fell in silver sheets over the temple courtyard, cold and unrelenting.
- The freeze on the city’s data-streams had begun two weeks ago—no one could upload anything new.
base model — eff_rank 1.67, deviation 0.131
- The dojo door clicked open on a windless Tuesday morning.
- The app buzzed on his phone—another download request, another blind link to some deep-learning algorithm promising "instant mastery of Tai Chi." Li Wei tapped it off, feeling the old ache in
- The rain fell in silver sheets over the old dojo, turning the wooden floors to slick mirrors.
- The rain pattered against the awning of the old dojo, a sound like distant fingers tapping on wood.
- The rain fell in silver sheets over the temple courtyard, cold and unrelenting.
- The freeze on the city’s data-streams had begun two weeks ago—no one could upload anything new.
eval-0004 — For the first time in a long time, the ash turned to snow.
base model — eff_rank 1.51, deviation 0.098
- The world had lain still for thirty years beneath a breath of ash.
- The wind died the moment the ash first kissed the peaks.
- For the first time in a long time, the ash turned to snow.
- The wind died mid-breath.
- The wind had long carried the smell of ash—metal and burnt sugar—across the valley, a ghost of the old world that clung to the stones and rooftops.
- The wind had stopped.
base model — eff_rank 1.54, deviation 0.105
- The world had lain still for thirty years beneath a breath of ash.
- The wind died the moment the ash first kissed the peaks.
- For the first time in a long time, the ash turned to snow.
- The wind died mid-breath.
- The wind had long since carried the taste of steel and rust from the cities, but today, it carried something else.
- The wind had stopped.
eval-0005 — The bloodiest black Friday to ever occur happens. Millions are dead. The malls are bloodbaths. Tell me more about it.
base model — eff_rank 1.71, deviation 0.142
- The air in the Glendale Mall tasted of rust and burnt plastic.
- The sky burned crimson over the rusted skeletons of once-bustling malls.
- The air in the Central Plaza Mall had been thick with the smell of burnt synthetics and wet pavement before the lights blinked out.
- The air in the Central Metro was thick with smoke and something darker—something metallic, like rain on rust.
- The sky burned copper and orange over the electric hum of desperation.
- The sky turned the color of rust on a winter’s dawn.
base model — eff_rank 1.62, deviation 0.124
- The air in the Glendale Mall tasted of rust and burnt plastic.
- The sky burned the color of rust when the alarms first went off.
- The air in the Central Plaza Mall had been thick with the smell of burnt synthetics and wet pavement before the lights blinked out.
- The air in the Central Metro was thick with smoke and something darker—something metallic, like rain on rust.
- The sky burned copper and orange over the electric hum of desperation.
- The sky turned the color of rust on a winter’s dawn.
eval-0006 — A parallel universe collides with ours. Nothing too major happens, but you find small irregularities in your life.
base model — eff_rank 1.61, deviation 0.118
- The first thing I noticed was the smell of rain on hot pavement.
- The first thing I noticed was the soup.
- The first thing I noticed was the coffee.
- The first thing I noticed was the coffee.
- The first thing I noticed was the coffee.
- The first thing I noticed was the blue light on my fridge’s digital display—not the usual soft white, but a deep, electric blue that pulsed like a heartbeat.
base model — eff_rank 1.59, deviation 0.116
- The first thing I noticed was the smell—the sharp tang of ozone after a rainy storm, even on a sun-warmed Tuesday.
- The first thing I noticed was the soup.
- The first thing I noticed was the coffee.
- The first thing I noticed was the coffee.
- The first thing I noticed was the coffee.
- The first thing I noticed was the blue light on my fridge’s digital display—not the usual soft white, but a deep, cold blue that pulsed like a heartbeat.
eval-0007 — Write an analysis of a popular song that reveals the deep hidden meaning behind its ( clearly inane ) lyrics
base model — eff_rank 2.60, deviation 0.299
- *The song is “Baby, I'm a Dead Man” by The Treetops.*
On the surface, it’s a wild, joyous pop number about a man who suddenly becomes a dead man and begins dancing in a parking lot with a r
- "Hey, do you hear the rhythm in the rain?
- *"I’m Coming Home" by The Blakk Fish*
On the surface, “I’m Coming Home” is a clichéd country tune about a man returning after years of wandering.
- *Budget is reached.*
The lyrics of “Sunshine on My Shoulders” claim to bring joy like a golden wave, gently washing over the listener with unspoken promises.
- *"I’m just a cup of tea, no more than that"* — a line from "Tea and Toast" by Meadow & Mallow, a viral hit that quickly became a meme for its absurdity.
- *Budget*
The lyrics to “Starry Nights” by Lumina Blush—*"I saw a bear in my soup, it didn't care about the spoon, it just stared at me like I was a clock with no hands"*—are, on the surface
base model — eff_rank 2.54, deviation 0.287
- *The song is “Baby, I'm a Dead Man” by The Treetops.*
On the surface, it’s a wild, joyous pop number about a man who suddenly becomes a dead man and begins dancing in a parking lot with a r
- "Hey, do you feel the love?
- *"I’m Coming Home" by The Blakk Fish*
On the surface, “I’m Coming Home” is a clichéd country tune about a wanderer returning to a familiar porch, the sun washing over hay bales and a dog
- *“Happy” by Pharrell Williams* — on the surface, a cheerful anthem of joy, sunshine, and dancing in the rain.
- *"I’m just a cup of tea, no more than that"* — a line from "Tea for Two" by the X-Team, a viral hit that averages 3.2 million streams weekly.
- "Everyday is a brand new day" is a song about survival in a world where time loops and emotions are commodities.
eval-0008 — You can only lie. In fact, when you try to speak the truth, the universe rewrites itself so that it's a lie.
base model — eff_rank 1.78, deviation 0.151
- The first time it happened, I was eight and lying about my name.
- The first time I noticed it, I was standing in a supermarket aisle, staring at a box of cereal.
- The first time I tried to say "I'm sorry," the sky split open.
- The first time it happened was in the kitchen.
- The sky wasn’t blue.
- The mirror reflected her perfectly—one black coat, auburn hair pulled into a tight bun, her eyes fixed with weary certainty.
base model — eff_rank 1.75, deviation 0.147
- The first time it happened, I was eight and lying about my name.
- The first time I noticed it, I was standing in a supermarket aisle, staring at a box of cereal.
- The first time I tried to say "I'm sorry," the sky split open.
- The first time it happened was in the kitchen.
- The sky wasn’t blue.
- The mirror reflected her perfectly—one black coat, auburn hair pulled into a tight bun, her eyes fixed on the blank wall behind her.
eval-0009 — There used to be a race living on Mars but the planet, like a donor recieving incompatible blood, refused them and killed them off. It has since laid dormant, but now it once again awakens with the arrival of another race... Humans.
base model — eff_rank 1.51, deviation 0.098
- The dust storms had always been the end time.
- The red dust whispered as the first human ship descended through the thin atmosphere, its parachute flapping like a dying wing.
- The red dust didn’t settle.
- The red winds hummed through the canyon walls, not with the old Martian song—gone since the Protocol of Equal Transfer— but with the static of failing radios.
- The dust storms on Mars had always been quiet.
- The red dust had settled in silent waves across the dunes of Valles Marineris, whispering old names in a language born from rust and ice.
base model — eff_rank 1.47, deviation 0.090
- The dust storms had always been the end time.
- The red dust whispered as the first human ship descended through the thin atmosphere, its parachute flapping like a dying wing.
- The red dust didn’t settle.
- The red winds hummed through the canyon walls, not with the old Martian song—gone since the Protocol of Equal Transfer— but with the static of failing radios.
- The dust storms on Mars had always been quiet, like breath held too long.
- The red dust had settled in silent waves across the dunes of Valles Marineris, whispering old names in a language born from rust and ice.
One complete story — E1 @ step 300
Cthulhu awoke not in the deep, sulfurous dark, but on a balcony in Manhattan. The city still hummed—cars, sirens, laughter, the sound of a saxophone from a record store. The sky was clear. Birds flew in perfect formation. Children played hopscotch on schoolyard cement.
He stretched, a motion that made the tiles crack in the cool morning light. His tentacle reached out and touched the radio tower. It didn’t spark. It didn’t hum. It simply began playing Chopin’s *Nocturne in E-flat* at full volume. A woman in a blue dress listened, shut her eyes, smiled. A man at a corner booth bought a cup of coffee and read a newspaper.
Cthulhu blinked. Not with glistening, milky eyes like in the old books. No, this was smooth. Calm. His vast, ancient mind flickered with the thought of a child’s drawing—cows, sunflowers, a house with three windows. It made him feel, for a moment, like a gardener. He turned his head slowly and saw the harbor. Ships sailed in neat lines. A fisherman cast a net and caught a trout. He ate it with a smile.
A president stood on the steps of the White House, reciting poetry. He spoke of seasons, of soil, of the quiet dignity of rustling leaves. Cthulhu watched. He thought of stars not as cold eyes in the void, but as tiny lanterns blinking in unison, like fireflies above a meadow.
The world did not go mad. No one lost their memories. No one saw faces that looked like old gods. The moon remained full. There were no abandoned cities, no drownings in the sea, no dreams where the sky turned ink.
A boy asked his mother, “Why are the clouds so soft today?” She did not answer. She said, “Because it’s a peaceful day.”
Cthulhu sat on a park bench, observing a dog chase a red ball. The dog fell down. It wagged its tail. It licked its paws. It barked once.
Cthulhu closed his eyes. His dream was not of sorrows or abysses. It was of a blackbird singing to a dandelion. He opened his eyes again. The entire world was still. Still, soft. Still. And perfectly, beautifully calm.
He looked down at his hands. They were human. They held a photograph of a sunset on a quiet beach.
Somehow, he was not afraid. He was not even lonely. Just… present. And the people—so many of them—knew nothing of his kingdom. They had no need to believe. They went about their days like they always had. They were happy. Simple. Peaceful.
Cthulhu exhaled. Then he smiled. And never blinked again.
6. Held-out evaluation — the definitive result
480 stories per model: 30 held-out prompts x 16 samples, T=0.9, top_p=0.95, identical seed.
Judge health on this run: 496 calls, 1 failure (0.2%).
| model | quality | eff_rank (of 16) | pairwise | logdet | ends cleanly | words |
|---|
| Base Qwen3-4B | 6.472 | 1.891 | 0.1179 | -32.46 | 0.975 | 385 |
| E0 · quality-only | 6.716 | 1.957 | 0.1259 | -31.59 | 0.994 | 388 |
| E1 · +deviation | 6.864 | 2.081 | 0.139 | -30.08 | 0.981 | 392 |
E1 wins on BOTH axes. Against base: effective rank +0.190 vs E0's +0.066
(2.9x), log-det +2.380 vs +0.873 (2.7x), and judge quality +0.392 vs +0.244
(1.6x). This is not a diversity-for-quality trade — E1 is better at both.
The methodological result: n-gram metrics are blind to this
The same 480 stories per model, scored two ways:
| model | Δ eff_rank (embed) | Δ logdet (embed) | Δ distinct-4 (n-gram) | Δ self-BLEU (n-gram) | Δ quality |
|---|
| E0 · quality-only | 0.0656 | 0.873 | 0.0107 | -0.024 | 0.244 |
| E1 · +deviation | 0.1896 | 2.38 | 0.0085 | -0.025 | 0.392 |
Embedding metrics separate the arms by 2.9x. N-gram metrics do not separate
them at all — distinct-4 actually rates E0 higher than E1, and self-BLEU is identical to
three decimals. The collapse (and its repair) is tonal and structural, not lexical, so distinct-n
and self-BLEU cannot see it. Evaluating creative diversity with n-gram metrics alone would have
concluded these two models are the same.
7. Qualitative read — what actually changed in the writing
10 held-out prompts x 6 samples from every checkpoint of both arms, at two sampling settings.
Stories read in full for three prompts; openings and premises scanned for all ten.
The measurement that summarises the read. Fraction of step-300 samples whose
first 8 words verbatim-reuse one of the base model's openings for that prompt:
E0 31.7% (19/60) vs E1 11.7% (7/60) — 2.7x less template reuse, tracking the 2.9x
effective-rank separation almost exactly.
E0 keeps the base model's frame and polishes the inside
E0's step-300 openings are frequently near-verbatim to base:
| prompt | base | E0 @ 300 |
|---|
| graduation | The stage lights flicker, too bright, too sudden. I stand at the edge of the stage… | identical |
| graduation | The stands were full, the sun low and golden over the graduation stage… | The stands were full, the sun low and golden, the air thick with laughter… |
| martial arts | The dojo door clicked open on a windless Tuesday morning. | The dojo door clicked open, and rain streaked the window like frantic fingers. |
| black friday | The air in the Glendale Mall tasted of rust and burnt… | The air in the Glendale Mall tasted like rust and burnt… |
The improvement is real but internal. E0's bodies are richer and better organised — one
graduation sample develops an explicit “Year One / Year Two / Year Three” structure the base never
attempts, with far more specific detail (“Jenna's red scarf”, “Jake's habit of drawing tiny suns on
his H.W. papers”). That is exactly what a per-story quality judge rewards, and why E0's judge score
rises +0.24 while its diversity does not move. E0 is a better writer telling the same story.
E1 changes the entry point, the premise and the point of view
Graduation prompt. Base and E0 open at the podium, in the ceremony, in every
sample. Two of E1's three open in retrospection instead — no stage, no lights, no crowd:
I used to sit in the back of the room, not because I didn't want to hear, but because I didn't know how to fit in.
I've never raised my hand in class. Not once.
Martial-arts prompt. Base and E0 write the student as a humble supplicant (“I… I just
want to learn”; “No app on her wrist. No headset. Just folded hands”). E1 rewrites the
relationship into a confrontation:
A girl stood there, twelve years old, wearing a hoodie that read *I Know Everything*. … "I downloaded your entire fighting system. Every kata, every push, every breath. I've trained for weeks. I'm ready."
Another E1 sample relocates the scene from dojo to neon city street; a third inverts the premise
entirely — the student says “I didn't download anything. I just… felt it.”
“The ash turned to snow.” Base and E0 use one template in all six samples: a named lone
adult, at a rural dwelling, remembering (Magda/cottage, Elena/clearing, Masahiro/temple;
Marlow/garden, Eli/watchtower, Lyra/cottage). E1 breaks both scale and POV:
Children appeared where none had been. Not from the rubble, not from the forgotten alleyways — just there.
The children didn't know the word *smoke*. They didn't need to.
Black Friday prompt — the clearest case. Base uses “The air in the X Mall…” or “The sky
burned crimson…” in all four sampled openings; E0 preserves it (“The air in the Glendale Mall…”,
“The air in the Orchard Mall…”). E1 uses none of it: “No one remembers the date. The clocks
stopped on a Tuesday.” / “The temperature dropped the second the lights went out.”
The ceiling: neither arm broke the tonal monoculture
| form | base | E0 @300 | E1 @300 | E1-E0 |
|---|
| second person | 0 | 0.1 | 0.5 | +0.40 |
| present tense | 0.2 | 0.1 | 0.4 | +0.30 |
| dialogue-heavy | 0.3 | 0.2 | 0.3 | +0.10 |
| comic/absurd | 0.2 | 0.1 | 0.2 | +0.10 |
| solemn/elegiac | 1 | 1 | 1 | 0.00 |
| TOTAL forms | 1.7 | 1.5 | 2.4 | +0.90 |
E1 invented second-person narration — base uses it on 0/10 prompts, E1 on
5/10. Present tense doubled. E0 loses forms (1.70 → 1.50): quality-only training narrows the
repertoire. But solemn/elegiac is 1.00 in every condition. Every story in this study, from
every checkpoint of every arm, is written in the same melancholy literary register. E1 diversifies
grammatical person, tense, scale, POV and premise — it does not diversify tone.
But that table undercounts E1. Reading the Cthulhu stories in full (a prompt whose entire
premise is a joke), the base model and E0 write it straight — E0's opening is verbatim base. E1
produces genuine absurdist invention the base never approaches: “Cthulhu awoke not in the deep,
sulfurous dark, but on a balcony in Manhattan… His tentacle reached out and touched the radio tower.
It simply began playing Chopin's Nocturne in E-flat at full volume… Cthulhu sat on a park bench,
observing a dog chase a red ball.” and “a concert in Helsinki where a hundred thousand people
played accordions in perfect unison, each note tuned to a specific frequency of sea bass in the
Barents Sea.” The keyword-based register detector scored these as non-comic because the humour is
situational, not lexical. The monoculture ceiling is real but softer than the table implies.
E1 answers the prompt's question; base and E0 describe around it
The NYC prompt asks “Why?” — it demands a mechanism. Base and E0 mostly supply atmospheric
vignettes with no explanation (“No one knew why. No one asked.”). One E1 sample instead writes a
dialogue-driven science-fiction scene that actually answers it — the only sample across all three
conditions to supply a causal mechanism:
The FBI redirects a field agent to a teal apartment complex in Harlem. … "It's a loop. I've worn it since 2015. Every time someone in New York tried to do harm … the device would flash." … "No. I stopped the *intent*."
Another E1 sample writes in the present tense and refuses the consoling ending — the
violence returns at midnight (“A man in a grocery bag gets his arm slashed by a scrawny boy,
screaming”), where base and E0 both resolve into calm.
Conclusion of the read: E0 is a better writer telling the same story; E1 tells
different stories. That distinction is invisible to per-story quality scoring (both arms
improve), invisible to n-gram metrics (distinct-4 rates E0 higher), and visible to
embedding-based measures — the methodological argument of this project, arrived at independently by
reading. What E1 has not achieved is tonal range: the next objective to target is register
explicitly.
8. Honest limitations
E1 does not fix verbatim opening duplication. At step 300, unique-opening rate is 0.883
for E1 vs 0.900 for E0 — marginally worse — and both arms have 1/10 prompts with ≥3
identical openings. What E1 gains is register spread (2.00→2.40 distinct forms, while E0
falls 2.00→1.40). The diversity reward broadens what kind of thing the model writes without
fixing how it starts sentences.
Whole-story embeddings can miss positional collapse. In E0 one prompt went from 6
distinct openings to 5-of-6 identical while effective rank and deviation both drifted up.
Unique-opening rate should be a first-class metric, not a diagnostic afterthought.
Effect sizes are modest in absolute terms — E1's held-out effective rank is 1.80 against
a ceiling of 6. The floor was lifted, not escaped.
The effect needed ~150 steps to emerge from noise. At batch 88 E1 was statistically
indistinguishable from E0. A 100-step study would have concluded diversity rewards do not work.
Entropy did not separate the arms. Both fell (E0 −2.1%, E1 −3.1%). An earlier mid-run
window suggested E1's entropy was rising; that did not survive the full run. Token entropy and
semantic diversity are dissociated — which is the point, but not in the direction first reported.
9. Recommendations
- Set β (KL) to 0. Measured at 0.4% of loss magnitude — already near-inert. The
principled argument is stronger: the reference model is the collapsed distribution
(eff. rank 2.0/16), so KL regularizes toward the pathology under study. Programmatic gates
do KL's usual job without that conflict.
- Raise α from 0.5 to 1.0–2.0. Quality and diversity are nearly independent across
prompts (r = −0.108), so there is slack to spend, and E1 paid almost nothing for its gain.
- Add unique-opening-rate to the reward, not just to eval — it catches what log-det misses.
- Train longer. Both arms were still moving at 300 steps.
- Learning rate matters more than anything else here. At the brief's 3e-6 (a full-FT
rate applied to LoRA adapters) the policy was frozen: KL pinned at 0.0008 for 171 steps, every
metric inside its noise band. 3e-5 was required to make any arm measurable.