Heralax/Augmental-Dataset
Viewer • Updated • 7.83k • 45 • 26
Full-parameter SFT of google/gemma-4-31B (base) on Heralax/Augmental-Dataset (7,831 rows, visual-novel style multi-character roleplay dialogue).
This is the end-of-epoch-3 (final, step 720) checkpoint. Sibling repos: epoch 1 (-ep1), epoch 2 (-ep2).
| Checkpoint | eval_loss | eval_ppl |
|---|---|---|
| base (step 0) | 1.753 | 5.77 |
epoch 1 (-ep1) |
1.572 | 4.81 |
epoch 2 (-ep2) |
1.521 (best held-out) | 4.58 |
| epoch 3 (this repo) | 1.707 | 5.51 |
Epoch 3 overfits the training set (train loss ~0.42 at end): held-out loss is worse than
epochs 1–2, but it imitates the dataset's style most strongly. Pick by use-case; -ep2
generalizes best.
Trained with a plain-text scenario format (no chat template). Prompt the model exactly like this, then let it continue after the trailing speaker tag:
Scenario: {scenario description}
{Speaker A}: "…"
{Speaker B}: "…"
{target speaker}:
Generation ends with <eos>.
train_on_inputs: false)Gemma derivatives are governed by the Gemma Terms of Use, including the Gemma Prohibited Use Policy.
Base model
google/gemma-4-31B