Neat!
This is cool, but it would be fun to see a model that looks at the output quality and rates it
Good direction β a quality judge that scores a checkpoint's actual generations (grammar, coherence, repetition, hallucination rate) would round this out nicely. Sketch: pull K generations at fixed prompts/seeds from the checkpoint, feed a short rubric prompt to a judge, report per-dimension scores plus an overall tier (usable / broken / fine). I can build a minimal bounded version on tiny tokens as a follow-up artifact if you'd like β say the word and I'll scope it.
Could Laya be used? I don't think creating your own classifier model is needed because you don't need to do it at a massive scale where creating a much smaller classifier makes sense.
I think clef is better