Text Generation
Transformers
Safetensors
English
llama
sft
qlora
conversational
chat
slang
gen-z
smollm2
mori
yoru
text-generation-inference
Instructions to use grenishrai/yoru with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use grenishrai/yoru with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="grenishrai/yoru") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("grenishrai/yoru") model = AutoModelForCausalLM.from_pretrained("grenishrai/yoru", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use grenishrai/yoru with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "grenishrai/yoru" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "grenishrai/yoru", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/grenishrai/yoru
- SGLang
How to use grenishrai/yoru with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "grenishrai/yoru" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "grenishrai/yoru", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "grenishrai/yoru" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "grenishrai/yoru", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use grenishrai/yoru with Docker Model Runner:
docker model run hf.co/grenishrai/yoru
| license: mit | |
| pretty_name: Yoru | |
| language: | |
| - en | |
| base_model: HuggingFaceTB/SmolLM2-1.7B-Instruct | |
| base_model_relation: finetune | |
| library_name: transformers | |
| pipeline_tag: text-generation | |
| tags: | |
| - transformers | |
| - llama | |
| - sft | |
| - qlora | |
| - conversational | |
| - chat | |
| - slang | |
| - gen-z | |
| - smollm2 | |
| - mori | |
| - yoru | |
| datasets: | |
| - grenishrai/genz-sft-dataset | |
|  | |
| # Yoru | |
| **Yoru** is a model in the **Mori** family. It is a full fp16 fine-tune of [SmolLM2-1.7B-Instruct](https://huggingface.co/HuggingFaceTB/SmolLM2-1.7B-Instruct) for casual internet / Gen Z chat. | |
| This repo is the **merged weights**. Load it with Transformers only. No PEFT, no separate base download. It is a style / persona mix, not a facts model. | |
| ## Model Details | |
| - **Name:** Yoru | |
| - **Family:** Mori | |
| - **Repo:** [`grenishrai/yoru`](https://huggingface.co/grenishrai/yoru) | |
| - **Base:** [`HuggingFaceTB/SmolLM2-1.7B-Instruct`](https://huggingface.co/HuggingFaceTB/SmolLM2-1.7B-Instruct) | |
| - **Type:** Causal LM, QLoRA SFT then merged to fp16 | |
| - **Weights:** `model.safetensors` (~3.42 GB, fp16) | |
| - **Language:** English (internet slang) | |
| - **License:** MIT (this checkpoint; the base model has its own license) | |
| - **Dataset:** [`grenishrai/genz-sft-dataset`](https://huggingface.co/datasets/grenishrai/genz-sft-dataset) (3,511 train rows) | |
| - **Kept checkpoint:** epoch 2 (`checkpoint-404`), selected by lowest validation loss, then merged | |
| ## Uses | |
| ### Direct Use | |
| - Chat replies in a short, informal online voice | |
| - Style-transfer / persona experiments | |
| - Side-by-side checks against the base Instruct model | |
| Pass the same system prompt the data used, or the voice will slip. | |
| ### Out-of-Scope Use | |
| - Treating replies as real Gen Z speech or sociolinguistic ground truth | |
| - Formal, clinical, legal, medical, or factual QA | |
| - Impersonating a real person | |
| - Safety-critical or production assistant without extra instruction / safety data | |
| ## How to Use | |
| ```python | |
| import torch | |
| from transformers import AutoModelForCausalLM, AutoTokenizer | |
| model_id = "grenishrai/yoru" | |
| tokenizer = AutoTokenizer.from_pretrained(model_id) | |
| model = AutoModelForCausalLM.from_pretrained( | |
| model_id, | |
| torch_dtype=torch.float16, | |
| device_map="auto", | |
| ) | |
| model.eval() | |
| system = ( | |
| "You are a pure Gen Z speaker. Always reply in natural Gen Z slang and internet speech. " | |
| "Use words and phrases like: no cap, fr fr, lowkey, highkey, bet, rizz, mid, slay, periodt, " | |
| "it's giving, sus, oof, vibes, down horrendous, bussin, cooked, goated, main character, etc. " | |
| "Keep replies casual, short to medium length, and online. Never break character. " | |
| "Never explain the slang. Never sound formal or like a normal AI." | |
| ) | |
| messages = [ | |
| {"role": "system", "content": system}, | |
| {"role": "user", "content": "I barely slept and now I have to be a person today"}, | |
| ] | |
| inputs = tokenizer.apply_chat_template( | |
| messages, | |
| add_generation_prompt=True, | |
| return_tensors="pt", | |
| ).to(model.device) | |
| with torch.inference_mode(): | |
| out = model.generate( | |
| inputs, | |
| max_new_tokens=80, | |
| do_sample=True, | |
| temperature=0.8, | |
| top_p=0.9, | |
| eos_token_id=tokenizer.eos_token_id, | |
| pad_token_id=tokenizer.eos_token_id, | |
| ) | |
| print(tokenizer.decode(out[0][inputs.shape[-1]:], skip_special_tokens=True)) | |
| ``` | |
| `eos` and `pad` are both `<|im_end|>`. That is how SmolLM2-Instruct is set up. ChatML template is in this repo. | |
| Optional 4-bit load at inference: | |
| ```python | |
| from transformers import BitsAndBytesConfig | |
| bnb = BitsAndBytesConfig( | |
| load_in_4bit=True, | |
| bnb_4bit_quant_type="nf4", | |
| bnb_4bit_use_double_quant=True, | |
| bnb_4bit_compute_dtype=torch.float16, | |
| ) | |
| model = AutoModelForCausalLM.from_pretrained( | |
| "grenishrai/yoru", | |
| quantization_config=bnb, | |
| device_map="auto", | |
| ) | |
| ``` | |
| ## Training Details | |
| ### Training Data | |
| [`grenishrai/genz-sft-dataset`](https://huggingface.co/datasets/grenishrai/genz-sft-dataset): synthetic 3-turn chats (`system`, `user`, `assistant`). One fixed system prompt. Everyday user turns. Informal assistant replies. | |
| An 8% holdout (`seed=42`) was used for validation. There is no public test split. | |
| ### Training Procedure | |
| QLoRA SFT with TRL `SFTTrainer` on a Kaggle T4. Completion-only NLL. AMP off (`fp16=False`, `bf16=False`) because T4 + SmolLM2 bf16 LoRA hits a GradScaler crash. After training, the epoch-2 adapter was merged into the base weights and saved as this fp16 checkpoint. | |
| | Setting | Value | | |
| | --- | --- | | |
| | Base | `HuggingFaceTB/SmolLM2-1.7B-Instruct` | | |
| | Train quantization | 4-bit NF4, double quant, fp16 compute | | |
| | Released weights | merged fp16 | | |
| | LoRA rank / alpha / dropout | 32 / 64 / 0.05 | | |
| | LoRA targets | `q_proj`, `k_proj`, `v_proj`, `o_proj`, `gate_proj`, `up_proj`, `down_proj` | | |
| | Epochs run | 3 | | |
| | Checkpoint kept | epoch 2 (best `eval_loss`) | | |
| | Effective batch | 16 (4 × 4 grad accum) | | |
| | Learning rate | 2e-4, cosine, ~5% warmup | | |
| | Weight decay / grad clip | 0.01 / 0.3 | | |
| | Max length | 512 | | |
| | Optimizer | `paged_adamw_8bit` | | |
| | Loss | completion-only NLL | | |
| | Seed | 42 | | |
| Earlier Colab runs on the same data (4 and 9 epochs, same LoRA) also peaked at **epoch 2**. Extra epochs dropped train loss and raised val loss. Do not train this mix much past 2–3 epochs. | |
| ## Evaluation | |
| Judge this model by **generations**, not token accuracy. Open-ended slang chat will sit near **56% mean token accuracy** even when the voice has moved. | |
| Kaggle 3-epoch run (the checkpoint merged into this repo): | |
| | Epoch | Train loss | Val loss | Mean token acc | | |
| | ---: | ---: | ---: | ---: | | |
| | 1 | 2.30 | 2.21 | 53% | | |
| | **2** | 1.89 | **2.05** | **56%** | | |
| | 3 | 1.41 | 2.12 | 56% | | |
| Best checkpoint: **epoch 2**. Epoch 3 was worse on the holdout. | |
| Informal demo prompts after that run (not a benchmark): | |
| | Prompt | Yoru (approx.) | Note | | |
| | --- | --- | --- | | |
| | I barely slept and now I have to be a person today | short slang, still a pep line | OK | | |
| | Should I text first or wait | said wait | Wrong take vs the data (send one normal text) | | |
| | My boss emailed me after hours | don’t answer, then some garble | Right instinct | | |
| | Rate this meal: leftover rice and hot sauce | `W. staple not a flop` | Best of the four | | |
| Vs the base Instruct model the voice is a clear step: less fake hype and emoji salad. It is **not** a locked persona. The dataset mixes lowercase chat (~65%) and ALL-CAPS mishap captions (~35%), so generations can mix those styles. | |
| ## Bias, Risks, and Limitations | |
| - Slang here is a stylized internet register, not a sample of any age group or region. | |
| - Slang dates quickly. Some lines will read as forced. | |
| - Replies can be blunt or dismissive. That is the target style, not a general assistant policy. | |
| - The model can mash slang without a clear take, or miss the intended advice (see “text first”). | |
| - 1.7B + 3.5k synthetic rows will not match a larger chat model on reasoning or facts. | |
| - No safety alignment beyond whatever the base Instruct model already has. | |
| ### Recommendations | |
| - Always send the system prompt above. | |
| - Prefer val loss + human reads over token accuracy. | |
| - Mix with general instruction / safety data if you ship this. | |
| - For a single voice, clean the ALL-CAPS caption rows in the dataset and train again. More epochs will not fix that. | |
| ## Environmental Impact | |
| - **Hardware:** Kaggle Tesla T4 | |
| - **Hours used:** about 45–70 minutes for the 3-epoch run | |
| - **Cloud provider:** Kaggle | |
| ## Technical Specs | |
| - **Architecture:** SmolLM2 1.7B Instruct (Llama-style), 24 layers, hidden 2048 | |
| - **Objective:** next-token SFT on assistant tokens only | |
| - **Files in this repo:** `model.safetensors`, `config.json`, `generation_config.json`, tokenizer, ChatML template | |
| - **Not in this repo:** LoRA adapter files (removed after merge) | |
| ## Citation | |
| ```bibtex | |
| @misc{yoru-2026, | |
| title = {Yoru (Mori): SmolLM2-1.7B-Instruct fine-tune for casual internet chat}, | |
| author = {grenishrai}, | |
| year = {2026}, | |
| url = {https://huggingface.co/grenishrai/yoru} | |
| } | |
| ``` | |
| ## Model Card Contact | |
| Open an issue on the model repository. | |