reaperdoesntknow
/

Qemma-sft

Text Generation

Generated from Trainer

text-generation-inference

Model card Files Files and versions

Metrics Training metrics Community

reaperdoesntknow commited on Nov 8, 2025

Commit

c0e43ab

·

verified ·

1 Parent(s): 217979c

Update README.md

Files changed (1) hide show

README.md +43 -15

README.md CHANGED Viewed

@@ -6,38 +6,66 @@ tags:
 - sft
 - trl
 licence: license
 ---
-# Model Card for Qemma-sft
-This model is a fine-tuned version of [None](https://huggingface.co/None).
-It has been trained using [TRL](https://github.com/huggingface/trl).
 ## Quick start
 ```python
-from transformers import pipeline
-question = "If you had a time machine, but could only go to the past or the future once and never return, which would you choose and why?"
-generator = pipeline("text-generation", model="reaperdoesntknow/Qemma-sft", device="cuda")
-output = generator([{"role": "user", "content": question}], max_new_tokens=128, return_full_text=False)[0]
-print(output["generated_text"])
 ```
-## Training procedure
 This model was trained with SFT.
 ### Framework versions
-- TRL: 0.25.0
-- Transformers: 4.57.1
-- Pytorch: 2.8.0+cpu
-- Datasets: 4.4.1
-- Tokenizers: 0.22.1
 ## Citations

 - sft
 - trl
 licence: license
+license: osl-3.0
+datasets:
+- O1-OPEN/OpenO1-SFT
+- yahma/alpaca-cleaned
+language:
+- en
+base_model:
+- google/gemma-3-1b-it
+- Qwen/Qwen3-0.6B
+pipeline_tag: text-generation
 ---
+# Model Card for Qemma
+**Qemma** is a HuggingFace-native hybrid model that merges **Gemma-3 (1B)** and **Qwen-3 (0.6B)** at the weight level (no adapters).
+Design: Gemma MLP/body + Qwen attention/head, projected and aligned to Gemma’s hidden size. The model is then SFT-tuned for stepwise reasoning.
 ## Quick start
 ```python
+from transformers import AutoTokenizer, AutoModelForCausalLM
+import torch
+model_id = "reaperdoesntknow/Qemma"
+tok = AutoTokenizer.from_pretrained(model_id, use_fast=True)
+model = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype=torch.bfloat16).eval()
+messages = [{"role": "user", "content": "Explain finite-scale discrepancy Δ_r in one paragraph."}]
+inputs = tok.apply_chat_template(messages, tokenize=True, add_generation_prompt=True, return_tensors="pt")
+out = model.generate(inputs, max_new_tokens=256, do_sample=True, temperature=0.7, top_p=0.9)
+print(tok.decode(out[0], skip_special_tokens=True))
 ```
+## What’s inside
+* **Architecture:** Gemma-3 backbone (26 layers, hidden 1152, MLP 6912) with **Qwen-style attention** regrouped to Gemma’s 4×256 heads.
+* **Tokenizer:** Gemma-3 tokenizer and chat template (see `chat_template.jinja`).
+* **Training:** SFT for instruction following and stepwise reasoning.
+## Intended use & limitations
+**Use:** research, instruction following, code/help, analysis, further SFT/RLHF.
+**Limits:** may hallucinate; not for safety-critical, medical, legal, or financial decisions. Follow dataset/model licenses.
+## Training procedure
+* ~512 warm-start steps (Alpaca-style data)
+* 256 SFT steps on `O1-OPEN/OpenO1-SFT`
+* +100 top-up SFT steps for reasoning behaviors
 This model was trained with SFT.
 ### Framework versions
+* TRL: 0.25.0
+* Transformers: 4.57.1
+* Pytorch: 2.8.0+cpu
+* Datasets: 4.4.1
+* Tokenizers: 0.22.1
 ## Citations