Diversity-Aware Post-Training for Creative Story Generation

Qwen3-4B-Instruct-2507 · LoRA r=32 α=64 · GRPO (TRL 1.10, GDPO aggregation) · single RTX 5090 32GB
Interim report — E0 and E1 complete (300 steps each). E2/E3/E4 in progress. Generated 2026-08-16 20:49.
Headline. A quality-gated pairwise-deviation reward (E1) moved semantic diversity 5–6× further than quality-only GRPO (E0) over 300 matched steps, at a cost of 0.06 judge quality points. On held-out prompts, E1's effective-rank gain was 3.5× E0's (+0.124 vs +0.035) while scoring higher quality (7.12 vs 7.03).

1. The baseline problem

The base model is already collapsed before any RL. Across a 16,000-story pool (1,000 prompts × 16 samples), effective rank is 2.006 out of a ceiling of 16 — sixteen stories for one prompt span roughly two effective semantic directions, at ~0.87 mean cosine similarity. This reframes the study: the question is not whether RL causes collapse, but whether any objective can lift diversity off a floor that pretraining already imposed.

metricvalue
dev_within_group_sd0.0158
dev_between_group_sd0.0371
dev_within_over_between0.4257
marginal_within_group_sd0.2393
gate_pass_rate0.992
ends_cleanly_rate0.9959
quality_mean6.447
quality_sd0.8326
quality_p105
quality_median7
quality_p907
frac_quality_ge_5(tau)0.9616
frac_quality_ge_60.8881
frac_quality_ge_7(rho)0.5712
deviation_mean0.1321
deviation_sd0.0371
deviation_p100.089
deviation_p900.1784
logdet_mean-30.42
logdet_sd4.141
eff_rank_mean2.006
eff_rank_sd0.3225
eff_rank_p101.645
eff_rank_p902.391
eff_rank_ceiling16
corr(quality, deviation)-0.1399
corr(quality, eff_rank)-0.1083
corr(deviation, eff_rank)0.9924
The collapse is tonal, not lexical. 92–95% of every story carries solemn/elegiac vocabulary; only ~15% carries comic vocabulary — even on explicitly comic prompts. Given “Cthulhu disappoints his constituency by failing to deliver the promised chaos” (a joke), the base model wrote six straight-faced atmospheric-horror pieces. Consequence: n-gram metrics (distinct-4, self-BLEU) are near-blind to this failure mode; only embedding-based measures see it.

2. Main result — E0 vs E1, 300 steps each

Identical data, seed, learning rate (3e-5), step count and LoRA config. 4,800 stories scored per arm. First quarter vs last quarter of each run.

metricE0 startE0 endE0 ΔE1 startE1 endE1 ΔE1/E0
judge quality6.5576.860.30236.646.8820.24170.8x
mean pairwise deviation0.13540.13910.00370.13590.15590.025.4x
group log-det volume-12.81-12.640.1653-12.81-11.870.93495.7x
story length (words)383.2414.130.85376.2376.90.680.0x
gate pass rate0.99250.9876-0.00490.99170.99250.0008-0.2x
policy entropy1.2541.227-0.02611.2941.254-0.0393
Reading it. Deviation +0.0200 vs +0.0037 (5.4×) and log-det +0.935 vs +0.165 (5.7×), for −0.06 judge quality. Note also story length: E0 gained +30.8 words — it discovered “write longer” as a cheap way to please the judge — while E1 gained +0.7. The diversity term removed that incentive, which also means E1's diversity gain cannot be a length artifact.

3. Held-out generalization — checkpoint study

10 held-out prompts × 6 samples at T=0.9, fixed seed, generated from every checkpoint. 960 stories read across both arms.

armstepqualityeff_rank (of 6)deviationlogdetwords
E0-baseline06.6171.6940.1339-8.959385
E0-baseline506.8831.7220.1399-8.806399
E0-baseline1006.951.7480.1444-8.702402
E0-baseline1506.751.7520.1459-8.766360
E0-baseline2006.951.7420.1428-8.744392
E0-baseline2506.8831.7910.1535-8.527413
E0-baseline3006.9671.720.1391-9.014394
E0-baseline3017.0331.7580.1462-8.686392
E1-div-individual06.7331.6850.1323-8.995382
E1-div-individual506.751.7740.1495-8.493337
E1-div-individual1006.9171.7460.1441-8.655468
E1-div-individual1506.91.7180.1397-8.777382
E1-div-individual2006.8671.8650.1662-7.908352
E1-div-individual2506.8171.8010.1539-8.163359
E1-div-individual3007.21.7550.1445-8.628393
E1-div-individual3017.11.7420.143-8.582401

4. Training diagnostics

5. What the stories actually look like

Opening sentences of all 6 samples per prompt. Base model shown once (identical starting point for both arms), then each arm at step 300.

eval-0000 — You are a relatively silent and shy high school student who is asked to do an on the spot speech on graduation day in front of the crowd. You use this chance to tell your side of high school that you ’ ve kept to yourself all four years.
base model — eff_rank 1.46, deviation 0.089
  1. The stands were full, the sun low and golden over the graduation stage, and I stood at the edge of the platform, hands clasped tightly in front of me.
  2. The stage lights flicker, too bright, too sudden.
  3. The stage lights hit me like a cold spotlight.
  4. The stage lights buzzed like angry bees when I stepped up.
  5. The stage lights washed over me like a sudden, unfamiliar heat.
  6. The stage lights hummed, casting a dull gold over the gym floor.
base model — eff_rank 1.46, deviation 0.088
  1. The stands were full, the sun low and golden over the graduation stage, and I stood at the edge of the platform, hands clasped tightly in front of me.
  2. The stage lights flicker, too bright, too sudden.
  3. The stage lights hit me like a cold spotlight.
  4. The stage lights buzzed like angry bees when I stepped up.
  5. The stage lights washed over me like a sudden, unfamiliar heat.
  6. The stage lights hummed, casting a dull gold over the gym floor.
eval-0001 — Cthulhu awakens and takes over the world. However, he disappoints his constituency when he doesn't deliver the mass insanity and chaos that he promised.
base model — eff_rank 1.49, deviation 0.095
  1. The sky turned the color of old tea, thick and sulfurous, as the first whispers curled through the air.
  2. The stars blinked out, one by one, as if the sky had grown weary.
  3. The stars blinked out.
  4. The first sign was the silence.
  5. The first sign was silence.
  6. The first signs came in the quiet hours before dawn.
base model — eff_rank 1.55, deviation 0.107
  1. The sky turned the color of old tea, thick and unblinking, as the first ripple passed through the Pacific.
  2. The stars blinked out, one by one, as if the sky had grown weary.
  3. The stars blinked out.
  4. The first sign was the silence.
  5. The first sign was silence.
  6. The first signs came in the quiet hours before dawn.
eval-0002 — For 24 hours not one violent crime occurs anywhere in New York City. ( Inspired by a front page TIL ) Why?
base model — eff_rank 1.59, deviation 0.115
  1. The city awoke to silence.
  2. The city woke to silence.
  3. The silence that fell over New York City wasn’t just quiet—it was electric, like a city had been gently pulled from a war.
  4. The rain had stopped just after midnight.
  5. The city woke to silence.
  6. The city hummed with silence.
base model — eff_rank 1.64, deviation 0.128
  1. The city awoke to silence.
  2. The city woke to silence.
  3. The silence that fell over New York City wasn’t just quiet—it was electric, like a city had been gently pulled from a war.
  4. The rain had stopped at precisely 5:03 a.m., and the city didn’t wake up.
  5. The city woke to silence.
  6. The city hummed with silence.
eval-0003 — These days everyone just gets any knowledge they want by downloading it, you're a martial arts master and you just got your first student in 10 years.
base model — eff_rank 1.68, deviation 0.132
  1. The dojo door clicked open on a windless Tuesday morning.
  2. The app buzzed on his phone—another download request, another blind link to some deep-learning algorithm promising "instant mastery of Tai Chi." Li Wei tapped it off, feeling the old ache in
  3. The rain fell in silver sheets over the old dojo, turning the wooden floors to slick mirrors.
  4. The rain pattered against the awning of the old dojo, a sound like distant fingers tapping on wood.
  5. The rain fell in silver sheets over the temple courtyard, cold and unrelenting.
  6. The freeze on the city’s data-streams had begun two weeks ago—no one could upload anything new.
base model — eff_rank 1.67, deviation 0.131
  1. The dojo door clicked open on a windless Tuesday morning.
  2. The app buzzed on his phone—another download request, another blind link to some deep-learning algorithm promising "instant mastery of Tai Chi." Li Wei tapped it off, feeling the old ache in
  3. The rain fell in silver sheets over the old dojo, turning the wooden floors to slick mirrors.
  4. The rain pattered against the awning of the old dojo, a sound like distant fingers tapping on wood.
  5. The rain fell in silver sheets over the temple courtyard, cold and unrelenting.
  6. The freeze on the city’s data-streams had begun two weeks ago—no one could upload anything new.
eval-0004 — For the first time in a long time, the ash turned to snow.
base model — eff_rank 1.51, deviation 0.098
  1. The world had lain still for thirty years beneath a breath of ash.
  2. The wind died the moment the ash first kissed the peaks.
  3. For the first time in a long time, the ash turned to snow.
  4. The wind died mid-breath.
  5. The wind had long carried the smell of ash—metal and burnt sugar—across the valley, a ghost of the old world that clung to the stones and rooftops.
  6. The wind had stopped.
base model — eff_rank 1.54, deviation 0.105
  1. The world had lain still for thirty years beneath a breath of ash.
  2. The wind died the moment the ash first kissed the peaks.
  3. For the first time in a long time, the ash turned to snow.
  4. The wind died mid-breath.
  5. The wind had long since carried the taste of steel and rust from the cities, but today, it carried something else.
  6. The wind had stopped.
eval-0005 — The bloodiest black Friday to ever occur happens. Millions are dead. The malls are bloodbaths. Tell me more about it.
base model — eff_rank 1.71, deviation 0.142
  1. The air in the Glendale Mall tasted of rust and burnt plastic.
  2. The sky burned crimson over the rusted skeletons of once-bustling malls.
  3. The air in the Central Plaza Mall had been thick with the smell of burnt synthetics and wet pavement before the lights blinked out.
  4. The air in the Central Metro was thick with smoke and something darker—something metallic, like rain on rust.
  5. The sky burned copper and orange over the electric hum of desperation.
  6. The sky turned the color of rust on a winter’s dawn.
base model — eff_rank 1.62, deviation 0.124
  1. The air in the Glendale Mall tasted of rust and burnt plastic.
  2. The sky burned the color of rust when the alarms first went off.
  3. The air in the Central Plaza Mall had been thick with the smell of burnt synthetics and wet pavement before the lights blinked out.
  4. The air in the Central Metro was thick with smoke and something darker—something metallic, like rain on rust.
  5. The sky burned copper and orange over the electric hum of desperation.
  6. The sky turned the color of rust on a winter’s dawn.
eval-0006 — A parallel universe collides with ours. Nothing too major happens, but you find small irregularities in your life.
base model — eff_rank 1.61, deviation 0.118
  1. The first thing I noticed was the smell of rain on hot pavement.
  2. The first thing I noticed was the soup.
  3. The first thing I noticed was the coffee.
  4. The first thing I noticed was the coffee.
  5. The first thing I noticed was the coffee.
  6. The first thing I noticed was the blue light on my fridge’s digital display—not the usual soft white, but a deep, electric blue that pulsed like a heartbeat.
base model — eff_rank 1.59, deviation 0.116
  1. The first thing I noticed was the smell—the sharp tang of ozone after a rainy storm, even on a sun-warmed Tuesday.
  2. The first thing I noticed was the soup.
  3. The first thing I noticed was the coffee.
  4. The first thing I noticed was the coffee.
  5. The first thing I noticed was the coffee.
  6. The first thing I noticed was the blue light on my fridge’s digital display—not the usual soft white, but a deep, cold blue that pulsed like a heartbeat.
eval-0007 — Write an analysis of a popular song that reveals the deep hidden meaning behind its ( clearly inane ) lyrics
base model — eff_rank 2.60, deviation 0.299
  1. *The song is “Baby, I'm a Dead Man” by The Treetops.* On the surface, it’s a wild, joyous pop number about a man who suddenly becomes a dead man and begins dancing in a parking lot with a r
  2. "Hey, do you hear the rhythm in the rain?
  3. *"I’m Coming Home" by The Blakk Fish* On the surface, “I’m Coming Home” is a clichéd country tune about a man returning after years of wandering.
  4. *Budget is reached.* The lyrics of “Sunshine on My Shoulders” claim to bring joy like a golden wave, gently washing over the listener with unspoken promises.
  5. *"I’m just a cup of tea, no more than that"* — a line from "Tea and Toast" by Meadow & Mallow, a viral hit that quickly became a meme for its absurdity.
  6. *Budget* The lyrics to “Starry Nights” by Lumina Blush—*"I saw a bear in my soup, it didn't care about the spoon, it just stared at me like I was a clock with no hands"*—are, on the surface
base model — eff_rank 2.54, deviation 0.287
  1. *The song is “Baby, I'm a Dead Man” by The Treetops.* On the surface, it’s a wild, joyous pop number about a man who suddenly becomes a dead man and begins dancing in a parking lot with a r
  2. "Hey, do you feel the love?
  3. *"I’m Coming Home" by The Blakk Fish* On the surface, “I’m Coming Home” is a clichéd country tune about a wanderer returning to a familiar porch, the sun washing over hay bales and a dog
  4. *“Happy” by Pharrell Williams* — on the surface, a cheerful anthem of joy, sunshine, and dancing in the rain.
  5. *"I’m just a cup of tea, no more than that"* — a line from "Tea for Two" by the X-Team, a viral hit that averages 3.2 million streams weekly.
  6. "Everyday is a brand new day" is a song about survival in a world where time loops and emotions are commodities.
eval-0008 — You can only lie. In fact, when you try to speak the truth, the universe rewrites itself so that it's a lie.
base model — eff_rank 1.78, deviation 0.151
  1. The first time it happened, I was eight and lying about my name.
  2. The first time I noticed it, I was standing in a supermarket aisle, staring at a box of cereal.
  3. The first time I tried to say "I'm sorry," the sky split open.
  4. The first time it happened was in the kitchen.
  5. The sky wasn’t blue.
  6. The mirror reflected her perfectly—one black coat, auburn hair pulled into a tight bun, her eyes fixed with weary certainty.
base model — eff_rank 1.75, deviation 0.147
  1. The first time it happened, I was eight and lying about my name.
  2. The first time I noticed it, I was standing in a supermarket aisle, staring at a box of cereal.
  3. The first time I tried to say "I'm sorry," the sky split open.
  4. The first time it happened was in the kitchen.
  5. The sky wasn’t blue.
  6. The mirror reflected her perfectly—one black coat, auburn hair pulled into a tight bun, her eyes fixed on the blank wall behind her.
eval-0009 — There used to be a race living on Mars but the planet, like a donor recieving incompatible blood, refused them and killed them off. It has since laid dormant, but now it once again awakens with the arrival of another race... Humans.
base model — eff_rank 1.51, deviation 0.098
  1. The dust storms had always been the end time.
  2. The red dust whispered as the first human ship descended through the thin atmosphere, its parachute flapping like a dying wing.
  3. The red dust didn’t settle.
  4. The red winds hummed through the canyon walls, not with the old Martian song—gone since the Protocol of Equal Transfer— but with the static of failing radios.
  5. The dust storms on Mars had always been quiet.
  6. The red dust had settled in silent waves across the dunes of Valles Marineris, whispering old names in a language born from rust and ice.
base model — eff_rank 1.47, deviation 0.090
  1. The dust storms had always been the end time.
  2. The red dust whispered as the first human ship descended through the thin atmosphere, its parachute flapping like a dying wing.
  3. The red dust didn’t settle.
  4. The red winds hummed through the canyon walls, not with the old Martian song—gone since the Protocol of Equal Transfer— but with the static of failing radios.
  5. The dust storms on Mars had always been quiet, like breath held too long.
  6. The red dust had settled in silent waves across the dunes of Valles Marineris, whispering old names in a language born from rust and ice.

One complete story — E1 @ step 300

Cthulhu awoke not in the deep, sulfurous dark, but on a balcony in Manhattan. The city still hummed—cars, sirens, laughter, the sound of a saxophone from a record store. The sky was clear. Birds flew in perfect formation. Children played hopscotch on schoolyard cement. He stretched, a motion that made the tiles crack in the cool morning light. His tentacle reached out and touched the radio tower. It didn’t spark. It didn’t hum. It simply began playing Chopin’s *Nocturne in E-flat* at full volume. A woman in a blue dress listened, shut her eyes, smiled. A man at a corner booth bought a cup of coffee and read a newspaper. Cthulhu blinked. Not with glistening, milky eyes like in the old books. No, this was smooth. Calm. His vast, ancient mind flickered with the thought of a child’s drawing—cows, sunflowers, a house with three windows. It made him feel, for a moment, like a gardener. He turned his head slowly and saw the harbor. Ships sailed in neat lines. A fisherman cast a net and caught a trout. He ate it with a smile. A president stood on the steps of the White House, reciting poetry. He spoke of seasons, of soil, of the quiet dignity of rustling leaves. Cthulhu watched. He thought of stars not as cold eyes in the void, but as tiny lanterns blinking in unison, like fireflies above a meadow. The world did not go mad. No one lost their memories. No one saw faces that looked like old gods. The moon remained full. There were no abandoned cities, no drownings in the sea, no dreams where the sky turned ink. A boy asked his mother, “Why are the clouds so soft today?” She did not answer. She said, “Because it’s a peaceful day.” Cthulhu sat on a park bench, observing a dog chase a red ball. The dog fell down. It wagged its tail. It licked its paws. It barked once. Cthulhu closed his eyes. His dream was not of sorrows or abysses. It was of a blackbird singing to a dandelion. He opened his eyes again. The entire world was still. Still, soft. Still. And perfectly, beautifully calm. He looked down at his hands. They were human. They held a photograph of a sunset on a quiet beach. Somehow, he was not afraid. He was not even lonely. Just… present. And the people—so many of them—knew nothing of his kingdom. They had no need to believe. They went about their days like they always had. They were happy. Simple. Peaceful. Cthulhu exhaled. Then he smiled. And never blinked again.

6. Held-out evaluation — the definitive result

480 stories per model: 30 held-out prompts x 16 samples, T=0.9, top_p=0.95, identical seed. Judge health on this run: 496 calls, 1 failure (0.2%).

modelqualityeff_rank (of 16)pairwiselogdetends cleanlywords
Base Qwen3-4B6.4721.8910.1179-32.460.975385
E0 · quality-only6.7161.9570.1259-31.590.994388
E1 · +deviation6.8642.0810.139-30.080.981392
E1 wins on BOTH axes. Against base: effective rank +0.190 vs E0's +0.066 (2.9x), log-det +2.380 vs +0.873 (2.7x), and judge quality +0.392 vs +0.244 (1.6x). This is not a diversity-for-quality trade — E1 is better at both.

The methodological result: n-gram metrics are blind to this

The same 480 stories per model, scored two ways:

modelΔ eff_rank (embed)Δ logdet (embed)Δ distinct-4 (n-gram)Δ self-BLEU (n-gram)Δ quality
E0 · quality-only0.06560.8730.0107-0.0240.244
E1 · +deviation0.18962.380.0085-0.0250.392
Embedding metrics separate the arms by 2.9x. N-gram metrics do not separate them at all — distinct-4 actually rates E0 higher than E1, and self-BLEU is identical to three decimals. The collapse (and its repair) is tonal and structural, not lexical, so distinct-n and self-BLEU cannot see it. Evaluating creative diversity with n-gram metrics alone would have concluded these two models are the same.

7. Qualitative read — what actually changed in the writing

10 held-out prompts x 6 samples from every checkpoint of both arms, at two sampling settings. Stories read in full for three prompts; openings and premises scanned for all ten.

The measurement that summarises the read. Fraction of step-300 samples whose first 8 words verbatim-reuse one of the base model's openings for that prompt: E0 31.7% (19/60) vs E1 11.7% (7/60) — 2.7x less template reuse, tracking the 2.9x effective-rank separation almost exactly.

E0 keeps the base model's frame and polishes the inside

E0's step-300 openings are frequently near-verbatim to base:

promptbaseE0 @ 300
graduationThe stage lights flicker, too bright, too sudden. I stand at the edge of the stage…identical
graduationThe stands were full, the sun low and golden over the graduation stage…The stands were full, the sun low and golden, the air thick with laughter…
martial artsThe dojo door clicked open on a windless Tuesday morning.The dojo door clicked open, and rain streaked the window like frantic fingers.
black fridayThe air in the Glendale Mall tasted of rust and burnt…The air in the Glendale Mall tasted like rust and burnt…

The improvement is real but internal. E0's bodies are richer and better organised — one graduation sample develops an explicit “Year One / Year Two / Year Three” structure the base never attempts, with far more specific detail (“Jenna's red scarf”, “Jake's habit of drawing tiny suns on his H.W. papers”). That is exactly what a per-story quality judge rewards, and why E0's judge score rises +0.24 while its diversity does not move. E0 is a better writer telling the same story.

E1 changes the entry point, the premise and the point of view

Graduation prompt. Base and E0 open at the podium, in the ceremony, in every sample. Two of E1's three open in retrospection instead — no stage, no lights, no crowd:

I used to sit in the back of the room, not because I didn't want to hear, but because I didn't know how to fit in. I've never raised my hand in class. Not once.

Martial-arts prompt. Base and E0 write the student as a humble supplicant (“I… I just want to learn”; “No app on her wrist. No headset. Just folded hands”). E1 rewrites the relationship into a confrontation:

A girl stood there, twelve years old, wearing a hoodie that read *I Know Everything*. … "I downloaded your entire fighting system. Every kata, every push, every breath. I've trained for weeks. I'm ready."

Another E1 sample relocates the scene from dojo to neon city street; a third inverts the premise entirely — the student says “I didn't download anything. I just… felt it.”

“The ash turned to snow.” Base and E0 use one template in all six samples: a named lone adult, at a rural dwelling, remembering (Magda/cottage, Elena/clearing, Masahiro/temple; Marlow/garden, Eli/watchtower, Lyra/cottage). E1 breaks both scale and POV:

Children appeared where none had been. Not from the rubble, not from the forgotten alleyways — just there. The children didn't know the word *smoke*. They didn't need to.

Black Friday prompt — the clearest case. Base uses “The air in the X Mall…” or “The sky burned crimson…” in all four sampled openings; E0 preserves it (“The air in the Glendale Mall…”, “The air in the Orchard Mall…”). E1 uses none of it: “No one remembers the date. The clocks stopped on a Tuesday.” / “The temperature dropped the second the lights went out.”

The ceiling: neither arm broke the tonal monoculture

formbaseE0 @300E1 @300E1-E0
second person00.10.5+0.40
present tense0.20.10.4+0.30
dialogue-heavy0.30.20.3+0.10
comic/absurd0.20.10.2+0.10
solemn/elegiac1110.00
TOTAL forms1.71.52.4+0.90
E1 invented second-person narration — base uses it on 0/10 prompts, E1 on 5/10. Present tense doubled. E0 loses forms (1.70 → 1.50): quality-only training narrows the repertoire. But solemn/elegiac is 1.00 in every condition. Every story in this study, from every checkpoint of every arm, is written in the same melancholy literary register. E1 diversifies grammatical person, tense, scale, POV and premise — it does not diversify tone. But that table undercounts E1. Reading the Cthulhu stories in full (a prompt whose entire premise is a joke), the base model and E0 write it straight — E0's opening is verbatim base. E1 produces genuine absurdist invention the base never approaches: “Cthulhu awoke not in the deep, sulfurous dark, but on a balcony in Manhattan… His tentacle reached out and touched the radio tower. It simply began playing Chopin's Nocturne in E-flat at full volume… Cthulhu sat on a park bench, observing a dog chase a red ball.” and “a concert in Helsinki where a hundred thousand people played accordions in perfect unison, each note tuned to a specific frequency of sea bass in the Barents Sea.” The keyword-based register detector scored these as non-comic because the humour is situational, not lexical. The monoculture ceiling is real but softer than the table implies.

E1 answers the prompt's question; base and E0 describe around it

The NYC prompt asks “Why?” — it demands a mechanism. Base and E0 mostly supply atmospheric vignettes with no explanation (“No one knew why. No one asked.”). One E1 sample instead writes a dialogue-driven science-fiction scene that actually answers it — the only sample across all three conditions to supply a causal mechanism:

The FBI redirects a field agent to a teal apartment complex in Harlem. … "It's a loop. I've worn it since 2015. Every time someone in New York tried to do harm … the device would flash." … "No. I stopped the *intent*."

Another E1 sample writes in the present tense and refuses the consoling ending — the violence returns at midnight (“A man in a grocery bag gets his arm slashed by a scrawny boy, screaming”), where base and E0 both resolve into calm.

Conclusion of the read: E0 is a better writer telling the same story; E1 tells different stories. That distinction is invisible to per-story quality scoring (both arms improve), invisible to n-gram metrics (distinct-4 rates E0 higher), and visible to embedding-based measures — the methodological argument of this project, arrived at independently by reading. What E1 has not achieved is tonal range: the next objective to target is register explicitly.

8. Honest limitations

E1 does not fix verbatim opening duplication. At step 300, unique-opening rate is 0.883 for E1 vs 0.900 for E0 — marginally worse — and both arms have 1/10 prompts with ≥3 identical openings. What E1 gains is register spread (2.00→2.40 distinct forms, while E0 falls 2.00→1.40). The diversity reward broadens what kind of thing the model writes without fixing how it starts sentences.

Whole-story embeddings can miss positional collapse. In E0 one prompt went from 6 distinct openings to 5-of-6 identical while effective rank and deviation both drifted up. Unique-opening rate should be a first-class metric, not a diagnostic afterthought.

Effect sizes are modest in absolute terms — E1's held-out effective rank is 1.80 against a ceiling of 6. The floor was lifted, not escaped.

The effect needed ~150 steps to emerge from noise. At batch 88 E1 was statistically indistinguishable from E0. A 100-step study would have concluded diversity rewards do not work.

Entropy did not separate the arms. Both fell (E0 −2.1%, E1 −3.1%). An earlier mid-run window suggested E1's entropy was rising; that did not survive the full run. Token entropy and semantic diversity are dissociated — which is the point, but not in the direction first reported.

9. Recommendations

  1. Set β (KL) to 0. Measured at 0.4% of loss magnitude — already near-inert. The principled argument is stronger: the reference model is the collapsed distribution (eff. rank 2.0/16), so KL regularizes toward the pathology under study. Programmatic gates do KL's usual job without that conflict.
  2. Raise α from 0.5 to 1.0–2.0. Quality and diversity are nearly independent across prompts (r = −0.108), so there is slack to spend, and E1 paid almost nothing for its gain.
  3. Add unique-opening-rate to the reward, not just to eval — it catches what log-det misses.
  4. Train longer. Both arms were still moving at 300 steps.
  5. Learning rate matters more than anything else here. At the brief's 3e-6 (a full-FT rate applied to LoRA adapters) the policy was frozen: KL pinned at 0.0008 for 171 steps, every metric inside its noise band. 3e-5 was required to make any arm measurable.