---
base_model: ytu-ce-cosmos/Turkish-Llama-8b-Instruct-v0.1
library_name: peft
license: llama3
language:
- tr
tags:
- qlora
- lora
- peft
- stem
- education
- k12
- turkish
- text-generation
pipeline_tag: text-generation
datasets:
- alimkacar/stem-tr-instruct-1k
metrics:
- bleu
- rouge
- bertscore
model-index:
- name: Turkish-Llama-8B-STEM-QLoRA
results:
- task:
type: text-generation
name: Turkish STEM Instruction Following
dataset:
name: eding-stem-tr-instruct-1k (test split)
type: alimkacar/stem-tr-instruct-1k
metrics:
- type: bleu
value: 46.94
name: BLEU
- type: rouge
value: 61.38
name: ROUGE-L
- type: bertscore
value: 81.43
name: BERTScore-F1
---
# π§ Turkish-Llama-8B-STEM-QLoRA
### A QLoRA adapter for **Turkish Kβ12 STEM & coding** instruction following




---
A **LoRA adapter** fine-tuned with **QLoRA** on top of [`ytu-ce-cosmos/Turkish-Llama-8b-Instruct-v0.1`](https://huggingface.co/ytu-ce-cosmos/Turkish-Llama-8b-Instruct-v0.1), specialised for **Kβ12 STEM and coding education in Turkish** (Arduino, Scratch, mBlock, robotics, Python, electronics, algorithms). Trained on the [`eding-stem-tr-instruct-1k`](https://huggingface.co/datasets/alimkacar/stem-tr-instruct-1k) dataset.
## π Evaluation
On a held-out test set (100 examples), the fine-tuned model **substantially beats** the zero-shot base model on every metric:
```text
0 20 40 60 80 100
BLEU base ββββββββββββββββββββββββββββββββββββββββ 4.8
FT ββββββββββββββββββββββββββββββββββββββββ 46.9 β² ~10x
ROUGE-L base ββββββββββββββββββββββββββββββββββββββββ 12.1
FT ββββββββββββββββββββββββββββββββββββββββ 61.4 β² ~5x
BERTScore base ββββββββββββββββββββββββββββββββββββββββ 51.7
FT ββββββββββββββββββββββββββββββββββββββββ 81.4 β² +29.7
```
| Metric | π΄ Base (zero-shot) | π’ Fine-tuned |
|:--|:--:|:--:|
| **BLEU** | 4.81 | **46.94** |
| **ROUGE-L** | 12.05 | **61.38** |
| **BERTScore-F1** (tr) | 51.70 | **81.43** |
> **Note:** A large part of the BLEU/ROUGE gain reflects the model learning the dataset's **concise answer format** (the base model is correct but verbose). The **BERTScore** (semantic) gain shows genuine content-similarity improvement. Read the result as *strong alignment to the target instructional style + a semantic-quality gain*.
## π§ Model details
| | |
|:--|:--|
| **Base model** | `ytu-ce-cosmos/Turkish-Llama-8b-Instruct-v0.1` (Llama-3, 8B) |
| **Method** | QLoRA (4-bit NF4 + double quant) + NEFTune |
| **LoRA** | `r=16`, `alpha=32`, dropout `0.05`, **all linear layers** (`q/k/v/o/gate/up/down_proj`) |
| **Trainable params** | 41,943,040 / 8,030,261,248 (**0.52%** β **99.48% reduction**) |
| **Effective batch** | 16 Β· **seq len** 512 (T4) / 1024 (L4Β·A100) |
| **Optimizer** | `paged_adamw_32bit`, LR `2e-4` cosine, 3 epochs |
| **Hardware** | single GPU (T4 / L4 / A100), auto fp16Β·bf16 |
## π Usage
```python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
from peft import PeftModel
BASE = "ytu-ce-cosmos/Turkish-Llama-8b-Instruct-v0.1"
ADAPTER = "alimkacar/Turkish-Llama-8B-STEM-QLoRA"
bnb = BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_quant_type="nf4",
bnb_4bit_compute_dtype=torch.bfloat16, bnb_4bit_use_double_quant=True)
model = AutoModelForCausalLM.from_pretrained(BASE, quantization_config=bnb, device_map="auto")
model = PeftModel.from_pretrained(model, ADAPTER)
tok = AutoTokenizer.from_pretrained(ADAPTER)
messages = [
{"role": "system", "content": "Sen bir TΓΌrkΓ§e K-12 STEM ve kodlama eΔitimi asistanΔ±sΔ±n. "
"CevaplarΔ±nΔ± TΓΌrkΓ§e ver, kodda her satΔ±rΔ± aΓ§Δ±kla."},
{"role": "user", "content": "Arduino ile servo motor nasΔ±l kontrol edilir?"},
]
ids = tok.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device)
eot = tok.convert_tokens_to_ids("<|eot_id|>")
out = model.generate(ids, max_new_tokens=400, do_sample=True, temperature=0.7,
top_p=0.9, eos_token_id=[tok.eos_token_id, eot])
print(tok.decode(out[0][ids.shape[-1]:], skip_special_tokens=True))
```
## π― Intended use & limitations
- **Intended:** helping students with Kβ12 STEM/coding questions in Turkish, with short, explained answers.
- **Limitations:** unreliable outside its domain. Trained on a **small (1k), mostly synthetic** dataset, so answers tend to be short and **template-like**, and can be less detailed than the base model on some questions. Code/hardware outputs should be reviewed by a teacher/adult. Inherits biases from the base model.
## π Citation
```bibtex
@misc{eding-stem-tr-2026,
title = {Eding STEM TR: Turkish K-12 STEM Instruction Dataset & QLoRA Fine-tuning},
author = {Alim Kacar},
year = {2026},
note = {Eding Internship project}
}
```
Methods: **QLoRA** (Dettmers et al., 2023) Β· **LoRA** (Hu et al., 2021) Β· **NEFTune** (Jain et al., 2023).
Dataset: [`alimkacar/stem-tr-instruct-1k`](https://huggingface.co/datasets/alimkacar/stem-tr-instruct-1k) Β· Base: `ytu-ce-cosmos/Turkish-Llama-8b-Instruct-v0.1` (Llama-3 license).
Alim Kacar Β· Eding Internship 2026