TinyStories-GR Llama 5M
A from-scratch, Llama-3.1-architecture causal language model trained to generate plausible short
children's stories in Modern Greek. Part of a size-ablation family trained on the same data,
tokenizer, and recipe -- only model size changes between them. Full architecture details are in
this repo's config.json.
Model family
Tokenizer
A custom byte-level BPE tokenizer, vocab size 16640 (16384 learned merges + the 256-entry byte
alphabet). Built from the Llama 3.2 tokenizer's tokenization pipeline (regex pre-tokenizer,
ByteLevel encoding/decoding) but retrained from scratch on the Greek corpus below, keeping only
bos/eos special tokens (not Llama's ~254 unused reserved/fine-tuning placeholder tokens).
Training data
alexliap/tinystories-gr: 2,141,648
Greek translations of the TinyStories dataset. Trained for 1 epoch (~413M tokens) on sequences
packed to 256 tokens (no padding, no truncation -- long stories span multiple packed sequences).
Sample generation
Greedy decoding, run until the model emits eos on its own (capped at 256 tokens):
Prompt: Μια φορά κι έναν καιρό, ένα μικρό κορίτσι
Completion (95 tokens, ended on eos):
Μια φορά κι έναν καιρό, ένα μικρό κορίτσι που το έλεγαν Λίλυ πήγε στο πάρκο για να παίξει. Είδε ένα αγόρι που το έλεγαν Τιμ. Ο Τιμ ήταν λυπημένος γιατί δεν μπορούσε να παίξει με το αγόρι. Η Λίλυ ήθελε να βοηθήσει τον Τιμ, έτσι του είπε: «Γεια σου, Τιμ! Μπορώ να σε βοηθήσω να βρεις το παιχνίδι σου».
Ο Τιμ χάρηκε πολύ και ευχαρίστησε τη Λίλυ. Έπαιξαν μαζί και πέρασαν υπέροχα. Από εκείνη την ημέρα, η Λίλυ και ο Τιμ έγιναν καλοί φίλοι και έπαιζαν μαζί κάθε μέρα.
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
repo_id = "alexliap/tinystories-gr-llama-5m"
tokenizer = AutoTokenizer.from_pretrained(repo_id)
model = AutoModelForCausalLM.from_pretrained(repo_id)
prompt = "Μια φορά κι έναν καιρό,"
input_ids = tokenizer(prompt, return_tensors="pt").input_ids
output = model.generate(input_ids, max_new_tokens=200, do_sample=False, eos_token_id=tokenizer.eos_token_id)
print(tokenizer.decode(output[0], skip_special_tokens=True))
Limitations
Trained on a single narrow domain (short, simple children's stories) for one epoch -- not a general-purpose language model. Expect fluent, grammatical Modern Greek within that domain, and degraded coherence/factuality well outside of it.
- Downloads last month
- -
Model tree for alexliap/tinystories-gr-llama-5m
Base model
meta-llama/Llama-3.2-1B