fwizzer1 commited on
Commit
e9b809b
·
verified ·
1 Parent(s): 3d4b97a

Update README with full documentation and reasoning tags

Browse files
Files changed (1) hide show
  1. README.md +24 -85
README.md CHANGED
@@ -3,110 +3,49 @@ language:
3
  - ru
4
  - en
5
  license: apache-2.0
 
6
  tags:
7
  - reasoning
8
  - deepseek-r1
9
  - think
10
- - mistral
 
11
  - gguf
12
- - lmstudio
13
  - unsloth
14
- - code
15
- - programming
16
  pipeline_tag: text-generation
17
- base_model: mistralai/Ministral-3-3B-Instruct-2512
18
  ---
19
 
20
- <div align="center">
21
 
22
- # 🧠 Fwizzer-R1-3B-RU
23
-
24
- **A lightweight, high-performance Russian reasoning model powered by Ministral 3B with DeepSeek-R1 style step-by-step thinking.**
25
-
26
- [![Dataset](https://img.shields.io/badge/Dataset-ru--deepthink--11k-blue)](https://huggingface.co/datasets/fwizzer1/ru-deepthink-11k)
27
- [![License](https://img.shields.io/badge/License-Apache_2.0-green.svg)](https://opensource.org/licenses/Apache-2.0)
28
- [![GGUF](https://img.shields.io/badge/GGUF-Q4__K__M%20%7C%20Q8__0-purple)](#-available-gguf-quantizations)
29
-
30
- </div>
31
-
32
- ---
33
-
34
- ## 🌟 Overview
35
-
36
- **Fwizzer-R1-3B-RU** is a specialized Russian reasoning language model fine-tuned on **11,000 verified multi-step reasoning dialogues** from [`fwizzer1/ru-deepthink-11k`](https://huggingface.co/datasets/fwizzer1/ru-deepthink-11k).
37
-
38
- It integrates native `<think>...</think>` internal monologue before every response, making it exceptional at:
39
- - **💻 Code Generation & Engineering**: Clean multi-language programming, algorithm design, debugging, refactoring, and software architecture.
40
- - **🧩 Math & Logic Reasoning**: Step-by-step problem solving, formal logic, and self-verification.
41
- - **⚡️ Ultra-fast Local Inference**: Optimized to run at **80+ tokens/sec** on budget GPUs (e.g. NVIDIA RTX 3050 4GB VRAM) with minimal memory footprint (~2.1 GB VRAM).
42
-
43
- ---
44
-
45
- ## 📦 Available GGUF Quantizations
46
-
47
- | File | Size | VRAM Required | Recommended For |
48
- | :--- | :--- | :--- | :--- |
49
- | **`Ministral-3-3B-Instruct-2512.Q4_K_M.gguf`** | **2.05 GB** | **~2.8 GB (with 4096 ctx)** | **Recommended (Best speed/accuracy balance)** |
50
- | **`Ministral-3-3B-Instruct-2512.Q8_0.gguf`** | **3.41 GB** | **~4.2 GB (with 4096 ctx)** | **Maximum precision** |
51
-
52
- ---
53
-
54
- ## 🚀 Quickstart in LM Studio
55
-
56
- 1. Open **LM Studio** and go to the **🔍 Search** tab.
57
- 2. Search for:
58
- ```text
59
- fwizzer1/Fwizzer-R1-3B-RU
60
- ```
61
- 3. Click **Download** on `Q4_K_M`.
62
- 4. Load the model and start chatting! The model will automatically output reasoning thoughts in `<think>...</think>` blocks.
63
 
64
  ---
65
 
66
- ## 💻 Python / Transformers Usage
67
-
68
- ```python
69
- from transformers import AutoModelForCausalLM, AutoTokenizer
70
- import torch
71
 
72
- model_id = "fwizzer1/Fwizzer-R1-3B-RU"
 
 
 
73
 
74
- tokenizer = AutoTokenizer.from_pretrained(model_id)
75
- model = AutoModelForCausalLM.from_pretrained(
76
- model_id,
77
- torch_dtype=torch.bfloat16,
78
- device_map="auto"
79
- )
80
 
81
- prompt = "Реши задачу: найди сложность алгоритма быстрой сортировки в худшем случае и напиши реализацию."
82
- messages = [
83
- {"role": "user", "content": prompt}
84
- ]
85
-
86
- inputs = tokenizer.apply_chat_template(
87
- messages,
88
- add_generation_prompt=True,
89
- return_tensors="pt"
90
- ).to("cuda")
91
-
92
- outputs = model.generate(
93
- inputs,
94
- max_new_tokens=2048,
95
- temperature=0.6,
96
- top_p=0.95
97
- )
98
-
99
- print(tokenizer.decode(outputs[0][inputs.shape[1]:], skip_special_tokens=True))
100
- ```
101
-
102
- ---
103
 
104
- ## 📊 Dataset
 
105
 
106
- Trained on [`fwizzer1/ru-deepthink-11k`](https://huggingface.co/datasets/fwizzer1/ru-deepthink-11k) consisting of 11,000 synthetic and curated Russian multi-step reasoning examples with rigorous step verification.
 
107
 
108
  ---
109
 
110
- ## 📜 License
111
 
112
- This project is open-source under the **Apache 2.0 License**.
 
 
 
 
3
  - ru
4
  - en
5
  license: apache-2.0
6
+ base_model: mistralai/Ministral-3B-Instruct-2410
7
  tags:
8
  - reasoning
9
  - deepseek-r1
10
  - think
11
+ - ru-deepthink-11k
12
+ - text-generation
13
  - gguf
14
+ - llama.cpp
15
  - unsloth
 
 
16
  pipeline_tag: text-generation
 
17
  ---
18
 
19
+ # 🧠 Fwizzer-R1-3B-RU (Thinking Reasoning LLM)
20
 
21
+ **Fwizzer-R1-3B-RU** — это мощная русскоязычная мыслящая языковая модель (Reasoning LLM) на архитектуре Ministral 3B, обученная по методологии **DeepSeek-R1** на датасете [fwizzer1/ru-deepthink-11k](https://huggingface.co/datasets/fwizzer1/ru-deepthink-11k) (11 000 пошаговых диалогов).
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
22
 
23
  ---
24
 
25
+ ## 🌟 Ключевые возможности
 
 
 
 
26
 
27
+ 1. 🏷 **Вшитая идентичность и системный промпт (general.name)**
28
+ * Модель на уровне ядра знает своё имя: **Fwizzer-R1-3B-RU**.
29
+ * Вшит дефолтный системный промпт:
30
+ > *«Ты думающая нейросеть а зовут тебя Fwizzer-R1-3B-RU. Весь ход мыслей и шаги пиши внутри тегов <think>(напиши сначала) и </think>(напиши по окончанию рассуждений), а итоговый ответ — обязательно после них.»*
31
 
32
+ 2. 🧠 **Глубокое мышление (DeepSeek-R1 Architecture)**
33
+ * Перед ответом модель активирует внутренний «черновик» в тегах `<think> ... </think>`, анализирует скрытые подвохи, краевые случаи и выводит математические доказательства.
 
 
 
 
34
 
35
+ 3. 🎭 **Графическая шторка размышлений в LM Studio (Reasoning UI)**
36
+ * В репозиторий вшит файл `model.yaml` с флагами `reasoning: true` и `reasoningFormat: deepseek`. При скачива��ии в LM Studio мысли автоматически сворачиваются в красивую плашку *Thought for X.Xs*.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
37
 
38
+ 4. 🛡 **Защита от сбоев (Fix Error 500)**
39
+ * Оптимизированный Jinja-шаблон чата без конфликтов системных сообщений, полная поддержка Flash Attention 2.
40
 
41
+ 5. 📚 **Расширенный контекст (до 256K токенов с RoPE YaRN Scaling)**
42
+ * Поддержка длинных кодовых баз, документаций и книг.
43
 
44
  ---
45
 
46
+ ## 🚀 Запуск в LM Studio
47
 
48
+ 1. Откройте **LM Studio**.
49
+ 2. В строке поиска введите: `fwizzer1/Fwizzer-R1-3B-RU`.
50
+ 3. Скачайте квантование **`Fwizzer-R1-3B.Q4_K_M.gguf`** (2.1 ГБ) или **`Q8_0`** (3.4 ГБ).
51
+ 4. Запустите чат — модель сразу готова к работе с автоматическим блоком размышлений!