--- language: - or license: apache-2.0 tags: - odia - odiagpt - transformer-from-scratch - indic - causal-lm pipeline_tag: text-generation widget: - text: "ଓଡ଼ିଶାର ଲୋକମାନେ" - text: "ଆଜି ସକାଳେ" - text: "ଓଡ଼ିଆ ଭାଷାର" --- # OdiaGPT-30M (Pretrained from Scratch) **OdiaGPT-30M** is an autoregressive decoder-only Transformer built and trained **completely from scratch** for the Odia language. It does **not** use pretrained GPT-2, Llama, or Indic model weights, nor off-the-shelf tokenizers. All 31.5 million parameters began from random Gaussian initialization, and its 8,192-token BPE tokenizer was trained exclusively on a cleaned subset of the AI4Bharat Sangraha Odia corpus. ## Model Summary - **Organization / Creator**: [CodeHima](https://huggingface.co/CodeHima) - **Architecture**: Decoder-only Transformer (Llama-compatible) - **Parameters**: 31,465,984 (~31.5M) - **Layers**: 8 - **Hidden dimension**: 512 (8 attention heads, head dimension: 64) - **Feed-Forward**: SwiGLU ($d_{ff} = 1536$) - **Positional Encoding**: Rotary Position Embeddings (RoPE, base 10000.0) - **Normalization**: RMSNorm ($\epsilon = 10^{-6}$) - **Context Length**: 512 - **Vocabulary**: 8,192 (SentencePiece BPE with byte fallback) - **Precision**: FP16 (Safetensors format) ## Training Dynamics - **Dataset**: AI4Bharat Sangraha (`verified/ori`) — 28.5M tokens / 110M cleaned characters - **Tokens Seen**: ~29.5 Million tokens (1,800 steps) - **Initial Loss**: 9.07 -> **Final Validation Loss**: 4.664 (Perplexity: **106.09**) - **Hardware**: Single NVIDIA Tesla T4 on Google Colab (~35,200 tokens/sec) ## Usage with Hugging Face Transformers ```python import torch from transformers import AutoModelForCausalLM, AutoTokenizer model_id = "CodeHima/OdiaGPT-30M" tokenizer = AutoTokenizer.from_pretrained(model_id) model = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype=torch.float16).cuda() prompt = "ଓଡ଼ିଶାର ଲୋକମାନେ" inputs = tokenizer(prompt, return_tensors="pt").to("cuda") outputs = model.generate( **inputs, max_new_tokens=64, temperature=0.8, top_k=50, top_p=0.9, repetition_penalty=1.15, do_sample=True ) print(tokenizer.decode(outputs[0], skip_special_tokens=True)) ``` ## Sample Output > **Prompt**: `ଆଜି ସକାଳେ` > **Completion**: *ଆଜି ସକାଳେ ସକାଳେ ଦୁଇ ଝିଅ ସହ ଆସି ଶୋଇପଡ଼ିଲା । ସେହି ସମୟରେ ବାପା ଘର ଲୋକ କୁ ଖବର ଦେଇ କହିଥିଲେ ଯେ, "ତେବେ ମାଆ, ତମେ ଆଉ କିଛି କହିଦେବୁ । ମୁଁ ଏ କଥା କିଛି ଶୁଣିନାହିଁ ।"* ## Limitations - Base foundation model, not instruction-tuned. - Developed as part of the OdiaGPT from-scratch educational curriculum.