--- license: apache-2.0 language: - ar base_model: Qwen/Qwen2.5-7B-Instruct tags: - arabic - poetry - shi3r - aruud - arabic-llm - reasoning - chain-of-thought - qwen2.5 - gguf - lora pipeline_tag: text-generation ---
# 🪶 المُحنّك · Almohanek 1.0 — 7B ### نموذجٌ عربيٌّ **يُفكِّر بالعربية في العَروض** قبل أن يَنظِم الشِّعر **The Arabic poet that reasons in Arabic — out loud — before it writes a line.** `by` **CyberQ** · Apache-2.0 · runs locally · 100+ tok/s on a single GPU
--- ## ✨ لماذا «المحنّك»؟ / Why Almohanek? معظم النماذج تكتب شعراً عربياً ثم تُفكّر بالإنجليزية (إن فكّرت). **المحنّك** مختلف: يفتح وسم `` ويُحلّل **البحر والقافية والصورة** *بالعربية الفصحى*، ثم ينظم البيت — كما يفعل شاعرٌ مُتمكِّن. Most "Arabic" models think in English behind the scenes. **Almohanek thinks in Arabic, on the page** — it opens a `` block, reasons about the **بحر (meter), قافية (rhyme), and imagery** in fluent Arabic, *then* composes. You see the craft, not just the output. | | | |---|---| | 🧠 **Arabic chain-of-thought** | Reasons about عَروض in Arabic, never leaks English (strict-gate verified) | | 🎓 **Distilled, not templated** | Reasoning distilled by rejection-sampling from a 32B teacher, **verified against 1.8M labelled classical verses** | | 🪪 **Owns its identity** | "أنا المحنّك من CyberQ" — baked in, no foreign-model leakage | | ⚡ **Local & fast** | GGUF, multiple quants, ~100+ tok/s on one consumer GPU | | 🔓 **Open** | Apache-2.0, built on Qwen2.5-7B-Instruct | --- ## 🚀 Quickstart (LM Studio / llama.cpp) | File | Size | Pick this if… | |---|---|---| | **`Almohanek1.0-7B-Q6_K.gguf`** | ~6 GB | you want the **sharpest** output ⭐ | | `Almohanek1.0-7B-Q4_K_M.gguf` | ~4.5 GB | you're tight on VRAM | **System prompt** (copy as-is): ``` أنت المحنّك من CyberQ، شاعر عربي خبير بالعروض. قبل أي قصيدة فكّر بإيجاز بالعربية داخل وسم ... (البحر، القافية، الصورة)، ثم اكتب الشعر. لا تذكر هويتك إلا إذا سُئلت "من أنت". ``` **Sampling:** `temperature 0.6 · top_p 0.9 · top_k 20 · repeat_penalty 1.15` **Try it:** ``` اكتب قصيدة من ٦ أبيات على بحر الكامل في الحنين ما بحر هذا البيت: قِفا نَبكِ مِن ذِكرى حَبيبٍ وَمَنزِلِ ``` --- ## 🛠️ How it was built Qwen2.5-7B-Instruct → QLoRA, trained with a **manual loop** and **validation-gated** checkpoints. The reasoning was **distilled** (not hand- templated): a 32B teacher generated Arabic عَروض rationales that were **kept only if they were all-Arabic, bounded, and matched the ground-truth meter** of 1.8M labelled verses. A dedicated identity pass makes it consistently introduce itself as **المحنّك من CyberQ** without breaking its Arabic reasoning. --- ## 🧭 Honest status & roadmap **Almohanek 1.0 is an experimental v1 — and we say so plainly.** ✅ **Solid today:** Arabic-only `` reasoning · stable identity · clean fluent Arabic · reliable stop · valid GGUF. 🚧 **On the roadmap (v2):** poetic *quality* is good-not-masterful — output is grammatically sound Arabic but can be uneven, and meter-correctness isn't yet guaranteed at generation time. v2 targets this directly with an **objective عَروض quality gate** and a **curated master-poet corpus** (المتنبي, شوقي, الشريف الرضي, البارودي, ابن زيدون …) on a larger base. This is a 7B specialist, not a general assistant. We ship honestly: real strengths up front, limits named, roadmap public. ---
**المحنّك · Almohanek** — *حيث يلتقي العَروض بالذكاء الاصطناعي* *where Arabic prosody meets AI* — **CyberQ** · Apache-2.0