Instructions to use grenishrai/yoru with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use grenishrai/yoru with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="grenishrai/yoru") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("grenishrai/yoru") model = AutoModelForCausalLM.from_pretrained("grenishrai/yoru", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use grenishrai/yoru with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "grenishrai/yoru" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "grenishrai/yoru", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/grenishrai/yoru
- SGLang
How to use grenishrai/yoru with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "grenishrai/yoru" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "grenishrai/yoru", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "grenishrai/yoru" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "grenishrai/yoru", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use grenishrai/yoru with Docker Model Runner:
docker model run hf.co/grenishrai/yoru
Yoru
Yoru is a model in the Mori family. It is a full fp16 fine-tune of SmolLM2-1.7B-Instruct for casual internet / Gen Z chat.
This repo is the merged weights. Load it with Transformers only. No PEFT, no separate base download. It is a style / persona mix, not a facts model.
Model Details
- Name: Yoru
- Family: Mori
- Repo:
grenishrai/yoru - Base:
HuggingFaceTB/SmolLM2-1.7B-Instruct - Type: Causal LM, QLoRA SFT then merged to fp16
- Weights:
model.safetensors(~3.42 GB, fp16) - Language: English (internet slang)
- License: MIT (this checkpoint; the base model has its own license)
- Dataset:
grenishrai/genz-sft-dataset(3,511 train rows) - Kept checkpoint: epoch 2 (
checkpoint-404), selected by lowest validation loss, then merged
Uses
Direct Use
- Chat replies in a short, informal online voice
- Style-transfer / persona experiments
- Side-by-side checks against the base Instruct model
Pass the same system prompt the data used, or the voice will slip.
Out-of-Scope Use
- Treating replies as real Gen Z speech or sociolinguistic ground truth
- Formal, clinical, legal, medical, or factual QA
- Impersonating a real person
- Safety-critical or production assistant without extra instruction / safety data
How to Use
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "grenishrai/yoru"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.float16,
device_map="auto",
)
model.eval()
system = (
"You are a pure Gen Z speaker. Always reply in natural Gen Z slang and internet speech. "
"Use words and phrases like: no cap, fr fr, lowkey, highkey, bet, rizz, mid, slay, periodt, "
"it's giving, sus, oof, vibes, down horrendous, bussin, cooked, goated, main character, etc. "
"Keep replies casual, short to medium length, and online. Never break character. "
"Never explain the slang. Never sound formal or like a normal AI."
)
messages = [
{"role": "system", "content": system},
{"role": "user", "content": "I barely slept and now I have to be a person today"},
]
inputs = tokenizer.apply_chat_template(
messages,
add_generation_prompt=True,
return_tensors="pt",
).to(model.device)
with torch.inference_mode():
out = model.generate(
inputs,
max_new_tokens=80,
do_sample=True,
temperature=0.8,
top_p=0.9,
eos_token_id=tokenizer.eos_token_id,
pad_token_id=tokenizer.eos_token_id,
)
print(tokenizer.decode(out[0][inputs.shape[-1]:], skip_special_tokens=True))
eos and pad are both <|im_end|>. That is how SmolLM2-Instruct is set up. ChatML template is in this repo.
Optional 4-bit load at inference:
from transformers import BitsAndBytesConfig
bnb = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4",
bnb_4bit_use_double_quant=True,
bnb_4bit_compute_dtype=torch.float16,
)
model = AutoModelForCausalLM.from_pretrained(
"grenishrai/yoru",
quantization_config=bnb,
device_map="auto",
)
Training Details
Training Data
grenishrai/genz-sft-dataset: synthetic 3-turn chats (system, user, assistant). One fixed system prompt. Everyday user turns. Informal assistant replies.
An 8% holdout (seed=42) was used for validation. There is no public test split.
Training Procedure
QLoRA SFT with TRL SFTTrainer on a Kaggle T4. Completion-only NLL. AMP off (fp16=False, bf16=False) because T4 + SmolLM2 bf16 LoRA hits a GradScaler crash. After training, the epoch-2 adapter was merged into the base weights and saved as this fp16 checkpoint.
| Setting | Value |
|---|---|
| Base | HuggingFaceTB/SmolLM2-1.7B-Instruct |
| Train quantization | 4-bit NF4, double quant, fp16 compute |
| Released weights | merged fp16 |
| LoRA rank / alpha / dropout | 32 / 64 / 0.05 |
| LoRA targets | q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj |
| Epochs run | 3 |
| Checkpoint kept | epoch 2 (best eval_loss) |
| Effective batch | 16 (4 × 4 grad accum) |
| Learning rate | 2e-4, cosine, ~5% warmup |
| Weight decay / grad clip | 0.01 / 0.3 |
| Max length | 512 |
| Optimizer | paged_adamw_8bit |
| Loss | completion-only NLL |
| Seed | 42 |
Earlier Colab runs on the same data (4 and 9 epochs, same LoRA) also peaked at epoch 2. Extra epochs dropped train loss and raised val loss. Do not train this mix much past 2–3 epochs.
Evaluation
Judge this model by generations, not token accuracy. Open-ended slang chat will sit near 56% mean token accuracy even when the voice has moved.
Kaggle 3-epoch run (the checkpoint merged into this repo):
| Epoch | Train loss | Val loss | Mean token acc |
|---|---|---|---|
| 1 | 2.30 | 2.21 | 53% |
| 2 | 1.89 | 2.05 | 56% |
| 3 | 1.41 | 2.12 | 56% |
Best checkpoint: epoch 2. Epoch 3 was worse on the holdout.
Informal demo prompts after that run (not a benchmark):
| Prompt | Yoru (approx.) | Note |
|---|---|---|
| I barely slept and now I have to be a person today | short slang, still a pep line | OK |
| Should I text first or wait | said wait | Wrong take vs the data (send one normal text) |
| My boss emailed me after hours | don’t answer, then some garble | Right instinct |
| Rate this meal: leftover rice and hot sauce | W. staple not a flop |
Best of the four |
Vs the base Instruct model the voice is a clear step: less fake hype and emoji salad. It is not a locked persona. The dataset mixes lowercase chat (65%) and ALL-CAPS mishap captions (35%), so generations can mix those styles.
Bias, Risks, and Limitations
- Slang here is a stylized internet register, not a sample of any age group or region.
- Slang dates quickly. Some lines will read as forced.
- Replies can be blunt or dismissive. That is the target style, not a general assistant policy.
- The model can mash slang without a clear take, or miss the intended advice (see “text first”).
- 1.7B + 3.5k synthetic rows will not match a larger chat model on reasoning or facts.
- No safety alignment beyond whatever the base Instruct model already has.
Recommendations
- Always send the system prompt above.
- Prefer val loss + human reads over token accuracy.
- Mix with general instruction / safety data if you ship this.
- For a single voice, clean the ALL-CAPS caption rows in the dataset and train again. More epochs will not fix that.
Environmental Impact
- Hardware: Kaggle Tesla T4
- Hours used: about 45–70 minutes for the 3-epoch run
- Cloud provider: Kaggle
Technical Specs
- Architecture: SmolLM2 1.7B Instruct (Llama-style), 24 layers, hidden 2048
- Objective: next-token SFT on assistant tokens only
- Files in this repo:
model.safetensors,config.json,generation_config.json, tokenizer, ChatML template - Not in this repo: LoRA adapter files (removed after merge)
Citation
@misc{yoru-2026,
title = {Yoru (Mori): SmolLM2-1.7B-Instruct fine-tune for casual internet chat},
author = {grenishrai},
year = {2026},
url = {https://huggingface.co/grenishrai/yoru}
}
Model Card Contact
Open an issue on the model repository.
- Downloads last month
- -
