kaushik-harsh-99's picture
Update README with prominent evaluation metrics, architecture details and sample outputs
2fd3015 verified
|
Raw
History Blame Contribute Delete
4.18 kB
metadata
language:
  - hi
license: mit
library_name: transformers
tags:
  - hindi
  - story-generation
  - causal-lm
  - llama-style
  - transformer
  - from-scratch
  - text-generation
datasets:
  - SmallScale/Simple-Stories-Hindi
pipeline_tag: text-generation
model-index:
  - name: Simple-Stories-Hindi-20M
    results: []

📖 Simple-Stories-Hindi-20M (22.3M Parameters)

A 22.3M parameter decoder-only Transformer language model trained from scratch on 2.11 million Hindi simple stories. The model generates coherent, creative, and grammatically sound Hindi stories given a short text prompt.


📊 Evaluation & Training Metrics

Metric / Property Value
Best Validation Loss 1.6914 (Cross-Entropy Loss)
Total Training Steps 386,000 steps
Total Parameters 22,310,784 (22.3M)
Non-Embedding Parameters 20,006,784 (20M)
Training Dataset SmallScale/Simple-Stories-Hindi (~2.11M stories)
Model Size on Disk ~86 MB (model.safetensors)

🏗️ Model Architecture Details

Parameter Value Notes
Architecture LLaMA-style Decoder RoPE + SwiGLU + RMSNorm
Hidden Size (d_model) 384 Vector dimension
FFN Intermediate Size 1024 8/3 × d_model rounded to multiple of 64
Layers (n_layers) 10 Transformer blocks
Attention Heads (n_heads) 8 Multi-Head Self Attention
Context Length (max_seq_len) 512 tokens Sequence window
Vocabulary Size 6,000 SentencePiece Unigram (Devanagari optimized)
Weight Tying Enabled Token embeddings & output projection share weights
Precision float32 Weights stored in native FP32 safetensors

🚀 Quick Start & Usage

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

# Load tokenizer and model directly from Hugging Face
tokenizer = AutoTokenizer.from_pretrained("SmallScale/Simple-Stories-Hindi-20M", trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained("SmallScale/Simple-Stories-Hindi-20M", trust_remote_code=True)

if torch.cuda.is_available():
    model = model.to("cuda")

# Prompt input
prompt = "एक समय की बात है"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)

# Generate story
outputs = model.generate(
    **inputs,
    max_new_tokens=200,
    do_sample=True,
    top_k=40,
    top_p=0.95,
    temperature=0.8
)

print(tokenizer.decode(outputs[0], skip_special_tokens=True))

📝 Sample Generated Stories

Prompt: एक समय की बात है

एक समय की बात है, और वह एक रात एक लड़की के लिए सब कुछ बदल सकती है। वह एक आभासी क्षेत्र में प्रवेश कर गई, जहाँ वह एक लड़के से मिली, जो उसके सपनों से बना था। वे अपने डर और इच्छाओं को साझा करते थे...

Prompt: एक जंगल में

एक जंगल में जहाँ पेड़ों ने रहस्यों को फुसफुसाया, एक लड़का एक छोटी सी झोपड़ी में रहता था। वह अक्सर सोचता था कि अगर वह अपने सपनों में एक परी से मिल सकता है तो क्या होगा...


🔗 Related Resources


📄 License

MIT License