glassbox/gpt-alpha-bg-91m
Text Generation • 91.3M • Updated • 47
Small models doing narrow tasks, trained from scratch: a 91M Bulgarian LM, a shlyokavitsa restorer (model + live demo), and the dataset behind it.
Note The 91.26M flagship - ties Gemini 3.5-flash at judging Bulgarian fluency, loses on raw compression.
Note Two character-level GPTs (3.16M / 4.73M) restoring shlyokavitsa to Cyrillic - beats a 2.6B general model at the task by a category, not a margin.
Note 210k (Latin, Cyrillic) pairs from Bulgarian Wikipedia - the first shlyokavitsa dataset on the Hub.
Latin-typed Bulgarian back into Cyrillic, in-browser
Note The restorer, live in your browser - type shlyokavitsa, watch it become Cyrillic, 100% client-side.