TinyStories-GR Llama 30M
A from-scratch, Llama-3.1-architecture causal language model trained to generate plausible short
children's stories in Modern Greek. Part of a size-ablation family trained on the same data,
tokenizer, and recipe -- only model size changes between them. Full architecture details are in
this repo's config.json.
Model family
Tokenizer
A custom byte-level BPE tokenizer, vocab size 16640 (16384 learned merges + the 256-entry byte
alphabet). Built from the Llama 3.2 tokenizer's tokenization pipeline (regex pre-tokenizer,
ByteLevel encoding/decoding) but retrained from scratch on the Greek corpus below, keeping only
bos/eos special tokens (not Llama's ~254 unused reserved/fine-tuning placeholder tokens).
Training data
alexliap/tinystories-gr: 2,141,648
Greek translations of the TinyStories dataset. Trained for 1 epoch (~413M tokens) on sequences
packed to 256 tokens (no padding, no truncation -- long stories span multiple packed sequences).
Sample generation
Greedy decoding, run until the model emits eos on its own (capped at 256 tokens):
Prompt: Μια φορά κι έναν καιρό, ένα μικρό κορίτσι
Completion (148 tokens, ended on eos):
Μια φορά κι έναν καιρό, ένα μικρό κορίτσι που το έλεγαν Λίλυ πήγε στο πάρκο με τη μαμά της. Είδαν ένα μεγάλο δέντρο και η Λίλυ ήθελε να σκαρφαλώσει πάνω του. Η μαμά της είπε: «Λίλυ, μπορείς να σκαρφαλώσεις στο δέντρο για να φτάσεις στην κορυφή;»
Η Λίλυ σκαρφάλωσε στο δέντρο και είδε ένα πουλάκι να κάθεται σε ένα κλαδί. Το πουλάκι είπε: «Γεια σου, μικρό πουλάκι! Μπορώ να σε βοηθήσω να ανέβεις στο δέντρο;» Η Λίλυ απάντησε: «Ναι, σε παρακαλώ!»
Η Λίλυ σκαρφάλωσε στο δέντρο και έφτασε στην κορυφή. Ήταν τόσο χαρούμενη που είπε: «Σε ευχαριστώ, πουλάκι! Είσαι τόσο ευγενικό και βοηθητικό!» Το πουλάκι απάντησε: «Παρακαλώ, Λίλυ. Χαίρομαι που βοήθησα». Και από εκείνη τη μέρα, η Λίλυ και το πουλάκι έγιναν οι καλύτεροι φίλοι.
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
repo_id = "alexliap/tinystories-gr-llama-30m"
tokenizer = AutoTokenizer.from_pretrained(repo_id)
model = AutoModelForCausalLM.from_pretrained(repo_id)
prompt = "Μια φορά κι έναν καιρό,"
input_ids = tokenizer(prompt, return_tensors="pt").input_ids
output = model.generate(input_ids, max_new_tokens=200, do_sample=False, eos_token_id=tokenizer.eos_token_id)
print(tokenizer.decode(output[0], skip_special_tokens=True))
Limitations
Trained on a single narrow domain (short, simple children's stories) for one epoch -- not a general-purpose language model. Expect fluent, grammatical Modern Greek within that domain, and degraded coherence/factuality well outside of it.
- Downloads last month
- -
Model tree for alexliap/tinystories-gr-llama-30m
Base model
meta-llama/Llama-3.2-1B