Text Generation
Transformers
Safetensors
English
llama
causal-lm
gqa
instruction-following
fine-tuned
yuna
text-generation-inference
Instructions to use meadbee/Yuna-130M-Instruct with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use meadbee/Yuna-130M-Instruct with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="meadbee/Yuna-130M-Instruct")# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("meadbee/Yuna-130M-Instruct") model = AutoModelForCausalLM.from_pretrained("meadbee/Yuna-130M-Instruct", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use meadbee/Yuna-130M-Instruct with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "meadbee/Yuna-130M-Instruct" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "meadbee/Yuna-130M-Instruct", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/meadbee/Yuna-130M-Instruct
- SGLang
How to use meadbee/Yuna-130M-Instruct with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "meadbee/Yuna-130M-Instruct" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "meadbee/Yuna-130M-Instruct", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "meadbee/Yuna-130M-Instruct" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "meadbee/Yuna-130M-Instruct", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use meadbee/Yuna-130M-Instruct with Docker Model Runner:
docker model run hf.co/meadbee/Yuna-130M-Instruct
File size: 7,755 Bytes
ebee230 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 | ---
language:
- en
library_name: transformers
pipeline_tag: text-generation
tags:
- llama
- causal-lm
- gqa
- instruction-following
- fine-tuned
- yuna
---
# YunaGPT-124M V1 Instruct
**A compact, English-first instruction model fine-tuned from YunaGPT-124M V1 Base.**





> **Important:** This is a small experimental model, not a production assistant. It can follow simple instructions but frequently produces incorrect, confused, repetitive, or invented information.
<p align="center">
<img src="Assets/info.png" alt="YunaGPT-124M V1 architecture and training infographic" width="600">
</p>
## Overview
YunaGPT-124M V1 Instruct is the general instruction-following variant of the Yuna model family. It starts from the pretrained Base checkpoint and applies response-only supervised fine-tuning (SFT): the instruction is visible as context, while training loss is applied to the response and its end-of-text token.
This variant is intended for short, single-turn requests. It is better suited to questions and instructions than the Base model, but its compact size strongly limits its knowledge, reasoning, consistency, and reliability.
## Project background
Yuna began as a 30M-parameter educational language-model project inspired by Sebastian Raschka's *Build a Large Language Model (From Scratch)*. It later moved to Hugging Face's native LLaMA implementation and grew into an experiment in how far a model could be trained on a home RTX 3090. The broader project also explores synthetic Final Fantasy X data, creative-writing SFT, role-play conversation SFT, and preference optimization.
The model's knowledge of Final Fantasy or any other subject should not be treated as factual. It may combine learned names and concepts with convincing hallucinations.
## Model summary
| Item | Value |
|---|---:|
| Parameters | **124,445,376** |
| Model class | `LlamaForCausalLM` |
| Training stage | General instruction SFT |
| Lineage | Base → Instruct |
| Context length | **2,048 tokens** |
| Vocabulary | **24,000 tokens** |
| Tokenizer | Byte-level BPE |
| Hidden layers | **25** |
| Hidden size | **576** |
| Attention / KV heads | **9 / 3** |
| Weight format | `safetensors`, FP32 |
| Primary language | English |
## Prompt format
This checkpoint does not use a standard chat template. It was trained with the following instruction wrapper:
```text
Below is an instruction that describes a task. Write a response that appropriately completes the request.
### Instruction:
{instruction}
### Input:
{optional_input}
### Response:
```
Omit the entire `### Input` section when no additional input is needed. Preserve the headings and blank lines for the closest match to training.
## Run it yourself
Install the runtime dependencies:
```bash
pip install torch transformers
```
```python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "YOUR_USERNAME/YunaGPT-124M-V1-Instruct"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype="auto")
model.eval()
def format_prompt(instruction: str, input_text: str = "") -> str:
prompt = (
"Below is an instruction that describes a task. "
"Write a response that appropriately completes the request.\n\n"
f"### Instruction:\n{instruction.strip()}"
)
if input_text.strip():
prompt += f"\n\n### Input:\n{input_text.strip()}"
return prompt + "\n\n### Response:\n"
prompt = format_prompt("Explain why the sky appears blue in two sentences.")
inputs = tokenizer(prompt, return_tensors="pt")
with torch.no_grad():
output = model.generate(
**inputs,
max_new_tokens=160,
do_sample=True,
temperature=0.7,
top_p=0.9,
repetition_penalty=1.1,
pad_token_id=tokenizer.pad_token_id,
eos_token_id=tokenizer.eos_token_id,
)
new_tokens = output[0, inputs["input_ids"].shape[1]:]
print(tokenizer.decode(new_tokens, skip_special_tokens=True).strip())
```
Replace the placeholder repository name with the final Hugging Face model ID or a local folder. The generation settings are starting points, not validated optimal values.
## Architecture
| Component | Configuration |
|---|---:|
| Architecture | Decoder-only Transformer |
| Attention | Grouped-Query Attention (GQA) |
| Hidden size | 576 |
| Intermediate size | 2,048 |
| Layers | 25 |
| Attention heads | 9 |
| Key/value heads | 3 |
| Head dimension | 64 |
| Activation | SiLU / SwiGLU feed-forward blocks |
| Normalization | RMSNorm, epsilon `1e-6` |
| Position encoding | RoPE, theta `10,000` |
| Maximum positions | 2,048 |
| Input/output embeddings | Tied |
## Training
The instruction stage was configured for four epochs with a batch size of 1 and a peak learning rate of `5e-5`. Examples were filtered for length and quality, deduplicated, and trained with prompt masking so only the assistant response and EOS target contributed to the loss.
The general instruction mixture was built from:
- `HuggingFaceH4/no_robots`;
- `databricks/databricks-dolly-15k`;
- the `self_instruct` portion of `HuggingFaceH4/helpful_instructions`.
Programming-heavy and explicit mathematics prompts were intentionally filtered because this model was not designed as a coding or math specialist. Dataset names are listed for provenance; their individual licenses, terms, and attribution requirements still apply.
## Intended uses
- Simple single-turn instruction-following experiments.
- Educational study of supervised fine-tuning on a compact model.
- Local prototyping with human review.
- A starting point for additional task-specific fine-tuning.
## Limitations and safety
Expected limitations include:
- hallucinated facts, names, quotations, and numbers;
- weak reasoning, arithmetic, coding, and multi-step planning;
- inconsistent instruction following and requested-length control;
- repetition, topic drift, malformed answers, and abrupt endings;
- no persistent memory or reliable multi-turn chat behavior;
- unreliable multilingual performance;
- possible biased, offensive, sexual, or otherwise unsafe generations inherited from source data;
- possible reproduction of information or phrases present in the training data.
Do not use this model for medical, legal, financial, safety-critical, or other high-impact decisions. Do not deploy it as an unsupervised public-facing assistant. Verify important claims using trustworthy external sources.
## Evaluation status
No standardized capability, factuality, bias, toxicity, privacy, or safety benchmarks are included with this release. The model author's informal assessment was approximately **3/10** for overall assistant quality; this is a candid subjective impression, not a benchmark result.
## Related variants
- **Base:** raw next-token completion checkpoint.
- **Story:** creative-writing SFT branch using the instruction wrapper.
- **Conversation:** role-play dialogue variant continued from Story and using a different prompt format.
## License and attribution
No model-weight license was declared in the project metadata when this card was prepared. Add an explicit license before public distribution. A model license does not override the source datasets' terms or attribution requirements.
---
**YunaGPT-124M V1 Instruct is an experimental research model. Use its responses with human review.**
|