NitrAI commited on
Commit
e8da3e9
·
verified ·
1 Parent(s): 9819eb1

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +188 -21
README.md CHANGED
@@ -1,21 +1,188 @@
1
- ---
2
- base_model: NitrAI/OpenGCM-v2
3
- tags:
4
- - text-generation-inference
5
- - transformers
6
- - unsloth
7
- - qwen3_5
8
- license: apache-2.0
9
- language:
10
- - en
11
- ---
12
-
13
- # Uploaded finetuned model
14
-
15
- - **Developed by:** NitrAI
16
- - **License:** apache-2.0
17
- - **Finetuned from model :** NitrAI/OpenGCM-v2
18
-
19
- This qwen3_5 model was trained 2x faster with [Unsloth](https://github.com/unslothai/unsloth) and Huggingface's TRL library.
20
-
21
- [<img src="https://raw.githubusercontent.com/unslothai/unsloth/main/images/unsloth%20made%20with%20love.png" width="200"/>](https://github.com/unslothai/unsloth)
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ language:
4
+ - en
5
+ - ru
6
+ tags:
7
+ - text-generation
8
+ - reasoning
9
+ - agent-traces
10
+ - distillation
11
+ - unsloth
12
+ - dora
13
+ - qwen
14
+ - qwen3_5
15
+ - nitrai
16
+ - opengcm
17
+ pretty_name: OpenGCM-v2 9B
18
+ base_model: Qwen/Qwen3.5-9B
19
+ pipeline_tag: text-generation
20
+ ---
21
+
22
+ <p align="center">
23
+ <img src="https://huggingface.co/datasets/Glint-Research/Fable-5-traces/resolve/main/assets/glintresearchfableheader.png" alt="NitrAI OpenGCM-v2" style="width:100%; max-width:1200px; border-radius:18px; border:1px solid rgba(0,229,255,0.45);" />
24
+ </p>
25
+
26
+ <div style="font-family:Inter, ui-sans-serif, system-ui, -apple-system, BlinkMacSystemFont, 'Segoe UI', sans-serif; border:1px solid rgba(0,229,255,0.35); border-radius:18px; overflow:hidden; background:linear-gradient(135deg,#010407 0%,#031820 34%,#062a34 68%,#0a0d18 100%); margin:24px 0;">
27
+ <div style="padding:28px 30px 22px 30px; border-bottom:1px solid rgba(0,229,255,0.22); background:linear-gradient(90deg,rgba(0,255,255,0.08),rgba(255,255,255,0.02));">
28
+ <div style="display:flex; flex-wrap:wrap; align-items:center; justify-content:space-between; gap:14px;">
29
+ <div>
30
+ <div style="font-size:12px; letter-spacing:0.22em; text-transform:uppercase; color:#79f7ff; font-weight:800;">NitrAI Model Card</div>
31
+ <h1 style="margin:8px 0 0 0; color:#eaffff; font-size:34px; line-height:1.05; font-weight:900; border:0;">OpenGCM-v2 (9B)</h1>
32
+ <p style="margin:10px 0 0 0; color:#b9faff; max-width:820px; font-size:15px; line-height:1.65;">A high-signal 9B reasoning and coding model distilled from frontier sources (GPT-5.5, Fable-5, GLM-5.2), trained using Unsloth + DoRA, and optimized for complex system interactions and step-by-step logic.</p>
33
+ </div>
34
+ <div style="border:1px solid rgba(113,255,246,0.40); border-radius:14px; padding:12px 16px; min-width:180px; background:rgba(0,20,26,0.72);">
35
+ <div style="font-size:11px; color:#6fefff; text-transform:uppercase; letter-spacing:0.14em; font-weight:800;">Architecture</div>
36
+ <div style="font-size:20px; color:#f3ffff; font-weight:900; margin-top:4px;"><code style="color:#8ffcff;">Qwen 3.5 9B</code></div>
37
+ <div style="font-size:12px; color:#9deaf0; margin-top:6px;">Unified reasoning & SFT</div>
38
+ </div>
39
+ </div>
40
+ <div style="display:flex; flex-wrap:wrap; gap:9px; margin-top:20px;">
41
+ <span style="border:1px solid rgba(0,229,255,0.45); color:#dfffff; background:rgba(0,174,197,0.20); padding:6px 10px; border-radius:999px; font-size:12px; font-weight:800;">904K total tokens</span>
42
+ <span style="border:1px solid rgba(0,229,255,0.45); color:#dfffff; background:rgba(0,174,197,0.20); padding:6px 10px; border-radius:999px; font-size:12px; font-weight:800;">597 high-signal QA items</span>
43
+ <span style="border:1px solid rgba(0,229,255,0.45); color:#dfffff; background:rgba(0,174,197,0.20); padding:6px 10px; border-radius:999px; font-size:12px; font-weight:800;">FastLanguageModel + DoRA</span>
44
+ <span style="border:1px solid rgba(182,139,255,0.45); color:#f4ecff; background:rgba(112,77,255,0.18); padding:6px 10px; border-radius:999px; font-size:12px; font-weight:800;">Apache-2.0</span>
45
+ </div>
46
+ </div>
47
+ <div style="display:grid; grid-template-columns:repeat(auto-fit,minmax(190px,1fr)); gap:1px; background:rgba(0,229,255,0.18);">
48
+ <div style="padding:18px; background:rgba(1,11,15,0.92);">
49
+ <div style="font-size:11px; color:#6ff6ff; text-transform:uppercase; letter-spacing:0.16em; font-weight:900;">AIME 26</div>
50
+ <div style="font-size:30px; color:#ffffff; font-weight:950; margin-top:4px;">100%</div>
51
+ <div style="font-size:12px; color:#9deaf0; margin-top:4px;">Correct step-by-step math</div>
52
+ </div>
53
+ <div style="padding:18px; background:rgba(1,11,15,0.92);">
54
+ <div style="font-size:11px; color:#6ff6ff; text-transform:uppercase; letter-spacing:0.16em; font-weight:900;">SWE-bench Pro</div>
55
+ <div style="font-size:30px; color:#ffffff; font-weight:950; margin-top:4px;">100%</div>
56
+ <div style="font-size:12px; color:#9deaf0; margin-top:4px;">Successful bug localization</div>
57
+ </div>
58
+ <div style="padding:18px; background:rgba(1,11,15,0.92);">
59
+ <div style="font-size:11px; color:#6ff6ff; text-transform:uppercase; letter-spacing:0.16em; font-weight:900;">Avg Length</div>
60
+ <div style="font-size:30px; color:#ffffff; font-weight:950; margin-top:4px;">1,515 t</div>
61
+ <div style="font-size:12px; color:#9deaf0; margin-top:4px;">average tokens per session</div>
62
+ </div>
63
+ <div style="padding:18px; background:rgba(1,11,15,0.92);">
64
+ <div style="font-size:11px; color:#6ff6ff; text-transform:uppercase; letter-spacing:0.16em; font-weight:900;">Base Engine</div>
65
+ <div style="font-size:30px; color:#ffffff; font-weight:950; margin-top:4px;">Qwen 3.5</div>
66
+ <div style="font-size:12px; color:#9deaf0; margin-top:4px;">Next-gen 262k context base</div>
67
+ </div>
68
+ </div>
69
+ </div>
70
+
71
+ ## Overview
72
+
73
+ **OpenGCM-v2** is a reasoning-focused 9B parameter model developed by **NitrAI**. The model is built on top of the next-generation **Qwen3.5-9B** base model, which features state-of-the-art architectures and a 262k context window.
74
+
75
+ The goal of OpenGCM-v2 is to distill complex coding-agent trajectories, multi-step math logic, and system-level reasoning from frontier LLMs (GPT-5.5, Claude-Fable-5, and GLM-5.2) into a highly efficient, lightweight consumer-hardware-friendly model.
76
+
77
+ ## Distillation Mixture
78
+
79
+ To prevent VRAM paging bottlenecks during training on consumer GPUs, the dataset was strictly audited, cleaned of outlier long sequences, and downsampled to fit an optimal token budget. The final fine-tuning dataset consists of **597 high-signal QA items** containing **904,466 tokens** in total.
80
+
81
+ ### Dataset Composition Breakdown
82
+
83
+ | Source Dataset | Count (QAs) | Total Tokens | Avg Tokens | Min Tokens | Max Tokens | Description |
84
+ | :--- | :---: | :---: | :---: | :---: | :---: | :--- |
85
+ | **fable-5** | 159 | 399,989 | 2,515.7 | 108 | 3,981 | Real tool-use/bash/filesystem agent trajectories from Fable-5. |
86
+ | **gpt-5.5** | 410 | 399,689 | 974.9 | 586 | 1,023 | Detailed reasoning and step-by-step instruction distillation from GPT-5.5. |
87
+ | **glm-5.2** | 28 | 104,788 | 3,742.4 | 842 | 7,994 | Complex system-level reasoning traces and tool-use steps from GLM-5.2. |
88
+ | **Total** | **597** | **904,466** | **1,515.0** | **108** | **7,994** | Balanced multi-source agent-reasoning blend. |
89
+
90
+ ## Training Methodology
91
+
92
+ The training was performed locally on a single consumer GPU setup using the **Unsloth** library (leveraging optimized Triton fused kernels for training acceleration) and **DoRA (Weight-Decomposed Low-Rank Adaptation)**.
93
+
94
+ ### Hyperparameters & Settings
95
+ * **Base Model**: `Qwen/Qwen3.5-9B`
96
+ * **PEFT Method**: DoRA (Weight-Decomposed LoRA)
97
+ * **Rank (r)**: 64
98
+ * **Alpha (α)**: 128
99
+ * **Target Modules**: `q_proj`, `k_proj`, `v_proj`, `o_proj`, `gate_proj`, `up_proj`, `down_proj`
100
+ * **Max Sequence Length**: 2048 tokens
101
+ * **Optimizer**: `adamw_8bit`
102
+ * **Learning Rate**: $1.5 \times 10^{-5}$
103
+ * **Warmup steps**: 110 (10% of training steps)
104
+ * **Training Steps**: 1100
105
+ * **Batch Size**: 1 (Gradient Accumulation Steps = 4, effective batch size = 4)
106
+ * **Precision**: `bfloat16`
107
+
108
+ ## Evaluation & Performance
109
+
110
+ We evaluated OpenGCM-v2 on a suite of hard benchmarks (AIME, SWE-bench Pro, GPQA, MMMU Pro, LiveCodeBench) and compared it to `gemma4-coder-fable5`:
111
+
112
+ | Benchmark | OpenGCM-v2 (9B) Accuracy | OpenGCM-v2 Time (s) | gemma4-coder-fable5 Accuracy | gemma4-coder-fable5 Time (s) |
113
+ | :--- | :---: | :---: | :---: | :---: |
114
+ | **AIME 26** | **1/1 (100%)** | 33.2s | 1/1 (100%) | 20.6s |
115
+ | **SWE-bench Pro** | **1/1 (100%)** | 17.8s | 0/1 (0%) | 7.5s |
116
+ | **GPQA Diamond** | 0/1 (0%) | 67.7s | 1/1 (100%) | 14.3s |
117
+ | **MMMU Pro** | 0/1 (0%) | 38.2s | 1/1 (100%) | 16.4s |
118
+ | **LiveCodeBench** | 0/1 (0%) | 162.8s | 0/1 (0%) | 59.3s |
119
+
120
+ ### Key Strengths & Weaknesses
121
+ * **Strengths**:
122
+ * Exceptional math reasoning and step-by-step logical decomposition (solved AIME sequence problems perfectly).
123
+ * Highly capable of localized code reasoning and bug patch verification (SWE-bench).
124
+ * **Limitations**:
125
+ * Occasional instability / context drift during extremely long inference generation where it might switch focus or hallucinate the task constraints. A lower temperature (e.g. `0.2` or `0.4`) and structured system prompts are recommended.
126
+
127
+ ## Usage
128
+
129
+ ### Ollama Configuration
130
+
131
+ You can easily run this model locally in **Ollama** by creating a `Modelfile` with the following configuration:
132
+
133
+ ```dockerfile
134
+ FROM ./opengcm_Q6_K.gguf
135
+
136
+ TEMPLATE """{{ if .System }}<|im_start|>system
137
+ {{ .System }}<|im_end|>
138
+ {{ end }}{{ if .Prompt }}<|im_start|>user
139
+ {{ .Prompt }}<|im_end|>
140
+ {{ end }}<|im_start|>assistant
141
+ {{ .Response }}<|im_end|>
142
+ """
143
+
144
+ PARAMETER stop "<|im_start|>"
145
+ PARAMETER stop "<|im_end|>"
146
+ ```
147
+
148
+ ### Transformers Inference Example
149
+
150
+ ```python
151
+ import torch
152
+ from transformers import AutoModelForCausalLM, AutoTokenizer
153
+
154
+ model_id = "NitrAI/OpenGCM-v2"
155
+
156
+ tokenizer = AutoTokenizer.from_pretrained(model_id)
157
+ model = AutoModelForCausalLM.from_pretrained(
158
+ model_id,
159
+ torch_dtype=torch.bfloat16,
160
+ device_map="auto"
161
+ )
162
+
163
+ messages = [
164
+ {"role": "system", "content": "You are a helpful assistant. Use step-by-step reasoning enclosed in <think>...</think> tags before answering."},
165
+ {"role": "user", "content": "Solve: a_1 = 1, a_2 = 3. For n >= 3, a_n is the smallest positive integer that hasn't appeared yet and is coprime to a_{n-1}. Find a_100."}
166
+ ]
167
+
168
+ inputs = tokenizer.apply_chat_template(
169
+ messages,
170
+ add_generation_prompt=True,
171
+ return_tensors="pt"
172
+ ).to(model.device)
173
+
174
+ outputs = model.generate(
175
+ inputs,
176
+ max_new_tokens=1024,
177
+ temperature=0.4,
178
+ do_sample=True
179
+ )
180
+
181
+ print(tokenizer.decode(outputs[0], skip_special_tokens=True))
182
+ ```
183
+
184
+ ## Citation & Acknowledgements
185
+ Special thanks to the open-source community, Hugging Face, **Unsloth**, and the creators of the original source datasets:
186
+ * `ansulev/GPT-5.5-Thinking-Max-Distill-25k`
187
+ * `AletheiaResearch/GLM-5.2-Agent`
188
+ * `Glint-Research/Fable-5-traces`