fivetech commited on
Commit
5c43ead
Β·
verified Β·
1 Parent(s): 315e10d

v2: Harbour/FWH Coder - Qwen3.5-35B-A3B LoRA (5004 examples, eval_loss=0.479)

Browse files
Files changed (1) hide show
  1. README.md +144 -49
README.md CHANGED
@@ -1,64 +1,159 @@
1
- # Harbour Fine-tuning Dataset
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
2
 
3
- ## Overview
4
- This dataset contains 996 training entries extracted from:
5
- - 1037 Harbour PRG (.prg) source files
6
- - 143 Harbour Header (.ch) files
7
 
8
- 177 PRG files and 7 CH files were skipped due to quality issues.
9
 
10
- ## Dataset Format
11
- The dataset is provided in JSONL format with the following structure:
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
12
 
13
- ### Instruction Format (harbour_train.jsonl / harbour_val.jsonl)
14
  ```json
15
- {"instruction": "...", "input": "", "output": "..."}
 
 
 
 
 
 
16
  ```
17
 
18
- ### Full Dataset (harbour_dataset_full.jsonl)
19
- ```json
20
- {"instruction": "...", "input": "", "output": "...", "metadata": {"file_path": "...", "language": "harbour", "category": "...", "subcategory": "..."}}
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
21
  ```
22
 
23
- ## Categories
24
- - **include**: Header files with constants/macros (59 files)
25
- - **rtl**: Harbour Runtime Library (80 files)
26
- - **contrib**: Contribution libraries (583 files)
27
- - **tests**: Test programs (225 files)
28
- - **utils**: Utility programs (13 files)
29
- - **extras**: Extra libraries (25 files)
30
-
31
- ## Cleaning Applied
32
- - Copyright/license headers removed
33
- - Disabled code blocks (#if 0) removed
34
- - Excessive trailing comments removed
35
- - Excessive blank lines removed
36
- - Files without actual code filtered out
37
- - Incomplete code (missing ENDCLASS, etc.) filtered out
38
-
39
- ## Usage for Fine-tuning
40
  ```bash
41
- # Using Ollama with Modelfile
42
- FROM qwen2.5-coder:14b
43
 
44
- # Training command
45
  ollama create harbour-coder -f Modelfile
46
-
47
- # Or use with other training frameworks
48
- # The JSONL format is compatible with:
49
- # - OpenAI fine-tuning API
50
- # - Hugging Face transformers
51
- # - Axolotl
52
- # - LLaMA-Factory
53
  ```
54
 
55
- ## File Structure
56
- - `harbour_train.jsonl` - Training set (896 entries)
57
- - `harbour_val.jsonl` - Validation set (100 entries)
58
- - `harbour_dataset_full.jsonl` - Full dataset with metadata
59
- - `dataset_stats.json` - Dataset statistics
60
- - `generate_dataset.py` - This script
 
 
 
 
 
 
61
 
62
- ## Source
63
- The source files are from the Harbour project (https://harbour.github.io/),
64
- an open-source Clipper-compatible compiler.
 
1
+ ---
2
+ base_model: Qwen/Qwen3.5-35B-A3B
3
+ library_name: peft
4
+ pipeline_tag: text-generation
5
+ tags:
6
+ - harbour
7
+ - fivewin
8
+ - fwh
9
+ - lora
10
+ - sft
11
+ - transformers
12
+ - trl
13
+ - unsloth
14
+ - code-generation
15
+ - xbase
16
+ - clipper
17
+ language:
18
+ - en
19
+ - es
20
+ license: apache-2.0
21
+ ---
22
 
23
+ # Harbour/FWH Coder β€” Qwen3.5-35B-A3B LoRA v2
 
 
 
24
 
25
+ LoRA adapter fine-tuned on **5,004 compilable Harbour and FiveWin (FWH) examples** for code generation. Built on top of [Qwen3.5-35B-A3B](https://huggingface.co/Qwen/Qwen3.5-35B-A3B), a 35B Mixture-of-Experts model with 256 experts (8 active per token).
26
 
27
+ ## What's New in v2
28
+
29
+ | | v1 | **v2** |
30
+ |---|---|---|
31
+ | **Dataset** | 996 entries (7 categories) | **5,004 entries** (8 categories) |
32
+ | **Training time** | 3h 43min | **~8h** |
33
+ | **Eval loss** | 0.5957 | **0.4790** (↓20%) |
34
+ | **Train loss** | 0.6456 | **0.4211** (↓35%) |
35
+ | **Format** | messages (chat) | instruction/output |
36
+ | **Learning rate** | 1e-4 | 8e-5 (conservative) |
37
+ | **Epochs** | 3 | 2 |
38
+
39
+ ### Key improvements
40
+
41
+ - **5x more training data** β€” expanded from 996 to 5,004 unique, compilable examples
42
+ - **Better loss convergence** β€” 35% lower train loss, 20% lower eval loss
43
+ - **More conservative training** β€” lower learning rate preserves base model capabilities
44
+ - **FiveWin (FWH) coverage** β€” added FiveWin GUI framework examples
45
+ - **Verified code** β€” all examples verified with Harbour v3.2.0dev compiler
46
+
47
+ ## Dataset
48
+
49
+ Training data sourced from the [Harbour](https://harbour.github.io/) project β€” an open-source Clipper-compatible compiler β€” and [FiveWin](https://fivewin.com/) (FWH) GUI framework.
50
+
51
+ ### Categories
52
+
53
+ | Category | Count | Description |
54
+ |---|---|---|
55
+ | contrib | 583 | Contribution libraries (network, database, graphics, security...) |
56
+ | rtl | 80 | Harbour Runtime Library |
57
+ | include | 59 | Header files with constants/macros |
58
+ | tests | 225 | Test programs |
59
+ | extras | 25 | Extra libraries |
60
+ | utils | 13 | Utility programs |
61
+ | fwh | ~500+ | FiveWin GUI framework examples |
62
+ | low-level C | 500+ | HB_FUNC C extension wrappers |
63
+
64
+ ### Format
65
 
 
66
  ```json
67
+ {
68
+ "instruction": "Write a Harbour function that creates a 2D array...",
69
+ "input": "",
70
+ "system": "You are an expert Harbour programmer...",
71
+ "output": "FUNCTION CreateTable()\n LOCAL aTable := {}\n ...",
72
+ "task_type": "code_generation"
73
+ }
74
  ```
75
 
76
+ ## Training Details
77
+
78
+ ### Hardware
79
+
80
+ - **Device:** NVIDIA GB10 (Grace Blackwell Superchip)
81
+ - **RAM:** 121 GB unified memory
82
+ - **Training time:** ~8 hours
83
+
84
+ ### Hyperparameters
85
+
86
+ | Parameter | Value |
87
+ |---|---|
88
+ | Base model | Qwen3.5-35B-A3B (MoE, 256 experts) |
89
+ | Method | QLoRA (4-bit) |
90
+ | LoRA rank | 8 |
91
+ | LoRA alpha | 16 |
92
+ | LoRA targets | q/k/v/o/gate/up/down_proj |
93
+ | Epochs | 2 |
94
+ | Learning rate | 8e-5 |
95
+ | LR scheduler | cosine |
96
+ | Warmup ratio | 0.05 |
97
+ | Batch size | 1 (effective: 16 via grad accum) |
98
+ | Max seq length | 1024 |
99
+ | Optimizer | adamw_8bit |
100
+
101
+ ### Framework
102
+
103
+ - Unsloth 2026.6.8
104
+ - Transformers 5.5.0
105
+ - PEFT 0.19.1
106
+ - PyTorch 2.12.1
107
+
108
+ ## How to Use
109
+
110
+ ### With PEFT + Transformers
111
+
112
+ ```python
113
+ from transformers import AutoModelForCausalLM, AutoTokenizer
114
+ from peft import PeftModel
115
+
116
+ base_model = AutoModelForCausalLM.from_pretrained(
117
+ "Qwen/Qwen3.5-35B-A3B",
118
+ load_in_4bit=True,
119
+ device_map="auto",
120
+ )
121
+ model = PeftModel.from_pretrained(base_model, "fivetech/Harbour")
122
+ tokenizer = AutoTokenizer.from_pretrained("fivetech/Harbour")
123
+
124
+ prompt = "Write a Harbour function that splits a CSV string into an array."
125
+ messages = [
126
+ {"role": "system", "content": "You are an expert Harbour programmer. Write compilable code."},
127
+ {"role": "user", "content": prompt},
128
+ ]
129
+ text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
130
+ inputs = tokenizer([text], return_tensors="pt").to(model.device)
131
+
132
+ output = model.generate(**inputs, max_new_tokens=1500, temperature=0.2)
133
+ print(tokenizer.decode(output[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))
134
  ```
135
 
136
+ ### With Ollama (merge + quantize first)
137
+
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
138
  ```bash
139
+ # Export to GGUF
140
+ python -m unsloth.save_pretrained_gguf model_output/ ./tokenizer/ q4_k_m
141
 
142
+ # Then use with Ollama
143
  ollama create harbour-coder -f Modelfile
 
 
 
 
 
 
 
144
  ```
145
 
146
+ ## Evaluation
147
+
148
+ Evaluated on 100 Harbour programming tests (Arrays, OOP, Functions, Database, File I/O, Control flow):
149
+
150
+ - **Compilation pass rate:** TBD (running test battery)
151
+ - **Categories tested:** Arrays (48), OOP (22), Other (9), Functions (8), Database (7), File I/O (4), Control (2)
152
+
153
+ ## License
154
+
155
+ Apache 2.0
156
+
157
+ ## Model Card Contact
158
 
159
+ fivetech β€” https://github.com/fivetechsoft/finetune