File size: 5,337 Bytes
c13d0c1
 
 
dd67537
 
 
 
 
 
 
 
 
c13d0c1
 
dd67537
 
 
 
 
 
 
c13d0c1
dd67537
 
 
 
c13d0c1
 
 
dd67537
c13d0c1
dd67537
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
c13d0c1
 
 
dd67537
 
 
 
 
 
 
 
 
 
c13d0c1
dd67537
c13d0c1
dd67537
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
c13d0c1
 
 
dd67537
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
---
license: apache-2.0
tags:
  - mlx
  - lora
  - mixture-of-experts
  - code-generation
  - python
  - typescript
  - reasoning
  - streaming
base_model: mistralai/Mistral-Small-3.1-24B-Instruct-2503
---

# FrankenMoE: Multi-Expert LoRA System (Python + TypeScript + Reasoning)

Triple LoRA adapters ที่ fine-tune บน 4-bit quantized MLX version ของ
`mistralai/Mistral-Small-3.1-24B-Instruct-2503` แยกตามความเชี่ยวชาญ
พร้อม Orchestrator ที่รองรับ Weighted Routing, LLM Fallback Classification,
Streaming Output, และ Multi-layer Safety Guard สำหรับป้องกัน
Repetition Loop และ Persona Drift

> **v2 Update:** เพิ่ม Expert 3 (Reasoning + Thai Translation Pipeline),
> เปลี่ยนจาก `frankenmoe_dispatcher.py` เดิมมาเป็น `moe_orchestrator.py`
> ที่มี Streaming Support และ Safety Net ที่ผ่านการทดสอบ Multi-turn
> มากกว่า 15 Turns โดยไม่พบ Repetition Loop

## Experts

| Expert | Language/Domain | Rank | Target Modules | Notes |
|---|---|---|---|---|
| expert-1-python | Python | 16 | q,k,v,o_proj | แก้ปัญหา rank/capacity ไม่พอจาก v1 |
| expert-2-typescript | TypeScript | 8 | default (incl. MoE switch layers) | แก้ NaN loss + Catastrophic Interference |
| expert-3-reasoning | General Reasoning (EN to TH) | - | default | มี Auto-translate pipeline เป็นภาษาไทย |

## Key Features

- **Weighted + LLM Fallback Routing**: ระบบเลือก Expert ด้วย Keyword Scoring
  ก่อน ถ้ากำกวมจะใช้ Base Model เป็น Classifier ช่วยตัดสิน
- **Streaming Output**: Python/TypeScript Expert รองรับ Token-by-token
  Streaming ผ่าน `query_stream()` เพื่อลด Perceived Latency
- **Repetition Guard**: ตรวจจับ Generation Loop ด้วย Sliding Window
  (ทุก ~50 Token ระหว่าง Stream) และ Retry-once Policy สำหรับ Non-stream
- **Persona Enforcement**: บังคับสรรพนามและคำลงท้ายให้สม่ำเสมอ
  ("ผม" / "ครับ") ตลอดทุก Turn
- **Auto Thai Translation**: Reasoning Expert แปลผลลัพธ์เป็นไทยอัตโนมัติ
  พร้อม Code-block Preservation (ป้องกันโค้ดถูกแปลผิดเพี้ยน)

## Usage

ต้องมี Base Model 4-bit อยู่ในเครื่อง:

pip install mlx-lm huggingface_hub[hf_xet]
huggingface-cli download --local-dir frankenmoe-python-typescript punsaisuwan/frankenmoe-python-typescript

### Interactive Chat (Streaming)

python moe_orchestrator.py

### Python API

from moe_orchestrator import MoEOrchestrator

orchestrator = MoEOrchestrator(
    base_model_path="./output-moe-mlx-4bit",
    adapter_paths={
        "python": "./adapters/expert-1-python",
        "typescript": "./adapters/expert-2-typescript",
        "reasoning": "./adapters/expert-3-reasoning",
    }
)

# Non-streaming (สำหรับ Automated Test / Batch)
response = orchestrator.query("เขียนฟังก์ชัน fibonacci ด้วย Python")
print(response)

# Streaming (สำหรับ Interactive Use)
for chunk in orchestrator.query_stream("เขียนฟังก์ชัน fibonacci ด้วย Python"):
    print(chunk, end="", flush=True)

## Training Details

- **Expert 1 (Python)**: อัปเกรดจาก v1 (rank=8, broken) เป็น rank=16,
  scale=8.0, ครอบคลุม 4 attention projections เพื่อแก้ปัญหา
  capacity ไม่พอสำหรับ code generation ที่ซับซ้อน
- **Expert 2 (TypeScript)**: แก้ NaN loss ผ่าน `SAFE_PAD_TOKEN_ID=198`
  และแก้ Catastrophic Interference ระหว่าง function/type guard bias
  ด้วย dataset rebalancing + ลด learning rate เป็น 1.5e-5
- **Expert 3 (Reasoning)**: Train แยกจาก Code Experts เพื่อป้องกัน
  Cross-domain Interference พร้อม Pipeline แปลไทยที่แยก Code Block
  ออกก่อนแปล (ป้องกันโค้ดถูกแปลผิด) แล้วค่อยประกอบกลับ

## Safety and Stability Notes

ระบบผ่านการทดสอบ Multi-turn Conversation 15 Turns โดยไม่พบ
Repetition Loop และไม่ต้องพึ่ง Safety Net Fallback แม้แต่ครั้งเดียว
(`REPETITION_PENALTY=1.15` เป็นค่าที่ผ่านการ Tune แล้วว่าดีที่สุด)

ปัญหาเล็กน้อยที่ทราบอยู่ (Known Minor Issues):
- Code fence อาจปิดไม่ครบในบาง Edge Case
- พบ CJK character leak ในระดับต่ำมาก ไม่กระทบความเข้าใจ

## Requirements

mlx-lm
huggingface_hub[hf_xet]

## License

Apache 2.0 (ตาม Base Model)