--- language: - hi license: mit library_name: transformers tags: - hindi - story-generation - causal-lm - llama-style - transformer - from-scratch - text-generation datasets: - SmallScale/Simple-Stories-Hindi pipeline_tag: text-generation model-index: - name: Simple-Stories-Hindi-10M results: [] --- # ЁЯУЦ Simple-Stories-Hindi-10M (11.45M Parameters) A **11.45M parameter** decoder-only Transformer language model trained **from scratch** on **2.11 million Hindi simple stories**. The model generates coherent, creative, and grammatically sound Hindi stories given a short text prompt. --- ## ЁЯУК Evaluation & Training Metrics | Metric / Property | Value | |---|---| | **Best Validation Loss** | **`1.8157`** (Cross-Entropy Loss) | | **Total Training Steps** | **202,000 steps** | | **Total Parameters** | **11,453,120 (11.45M)** | | **Non-Embedding Parameters** | **10,173,120 (10.17M)** | | **Training Dataset** | [SmallScale/Simple-Stories-Hindi](https://huggingface.co/datasets/SmallScale/Simple-Stories-Hindi) (~2.11M stories) | | **Model Size on Disk** | ~45.8 MB (`model.safetensors`) | --- ## ЁЯПЧя╕П Model Architecture Details | Parameter | Value | Notes | |---|---|---| | **Architecture** | LLaMA-style Decoder | RoPE + SwiGLU + RMSNorm | | **Hidden Size (`d_model`)** | 320 | Vector dimension | | **FFN Intermediate Size** | 896 | 8/3 ├Ч `d_model` rounded to multiple of 64 | | **Layers (`n_layers`)** | 7 | Transformer blocks | | **Attention Heads (`n_heads`)** | 5 | Multi-Head Self Attention | | **Head Dimension** | 64 | `d_model` / `n_heads` | | **Context Length (`max_seq_len`)** | 512 tokens | Sequence window | | **Vocabulary Size** | 4,000 | SentencePiece Unigram (Devanagari optimized) | | **Weight Tying** | Enabled | Token embeddings & output projection share weights | | **Precision** | float32 | Weights stored in native FP32 safetensors | --- ## ЁЯЪА Quick Start & Usage ```python import torch from transformers import AutoModelForCausalLM, AutoTokenizer # Load tokenizer and model directly from Hugging Face tokenizer = AutoTokenizer.from_pretrained("SmallScale/Simple-Stories-Hindi-10M", trust_remote_code=True) model = AutoModelForCausalLM.from_pretrained("SmallScale/Simple-Stories-Hindi-10M", trust_remote_code=True) if torch.cuda.is_available(): model = model.to("cuda") # Prompt input prompt = "рдПрдХ рд╕рдордп рдХреА рдмрд╛рдд рд╣реИ" inputs = tokenizer(prompt, return_tensors="pt").to(model.device) # Generate story outputs = model.generate( **inputs, max_new_tokens=200, do_sample=True, top_k=40, top_p=0.95, temperature=0.8 ) print(tokenizer.decode(outputs[0], skip_special_tokens=True)) ``` --- ## ЁЯУЭ Sample Generated Stories **Prompt:** `рдПрдХ рд╕рдордп рдХреА рдмрд╛рдд рд╣реИ` > *рдПрдХ рд╕рдордп рдХреА рдмрд╛рдд рд╣реИ, рдФрд░ рдореИрдВ рдЫрд╛рдпрд╛ рд╕реЗ рджреЗрдЦрддрд╛ рд╣реВрдВред рдореЗрд░реЗ рджреЛ рд▓реЛрдЧ, рдЬреАрди рдФрд░ рд╕реИрдореБрдЕрд▓ рд╣реИрдВ, рдЬреЛ рдПрдХ рднрд╡реНрдп рдпрд╛рддреНрд░рд╛ рдкрд░ рдЬрд╛ рд░рд╣реЗ рд╣реИрдВред рд╡реЗ рдПрдХ рд╣реА рд╕реНрдерд╛рди рдкрд░ рд░рд╣рддреЗ рд╣реИрдВ, рд▓реЗрдХрд┐рди рд╡реЗ рджреЛрдиреЛрдВ рдЕрдкрдиреА-рдЕрдкрдиреА рдХрд╣рд╛рдирд┐рдпрд╛рдБ рдЪрд╛рд╣рддреЗ рд╣реИрдВ...* --- ## ЁЯФЧ Related Resources - **GGUF (FP16) Model Repo:** [SmallScale/Simple-Stories-Hindi-10M-GGUF](https://huggingface.co/SmallScale/Simple-Stories-Hindi-10M-GGUF) - **20M Model Repo:** [SmallScale/Simple-Stories-Hindi-20M](https://huggingface.co/SmallScale/Simple-Stories-Hindi-20M) - **Dataset:** [SmallScale/Simple-Stories-Hindi](https://huggingface.co/datasets/SmallScale/Simple-Stories-Hindi) --- ## ЁЯУД License MIT License