File size: 1,259 Bytes
f2a3a7b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
# Research Notes

This repository started as a text-to-SQL fine-tuning research project and is now
being shaped into LocalSQL, a usable local/private SQL assistant.

## Core Result

The first strong result was Spider V1:

- Fine-tuned Qwen2.5-Coder-7B: **78.2%** result accuracy.
- Grok-4 baseline: 73.7%.
- DeepSeek-V3 baseline: 71.8%.

The method was frontier-disagreement DPO: run two strong models, execute their
SQL, and turn correctness disagreements into DPO preference pairs.

## BIRD Findings

BIRD is harder because schemas are larger, train/dev databases are disjoint, and
questions include external evidence. The important lessons so far:

- Clear correctness pairs helped the 7B model.
- Judge-resolved "both correct but stylistically different" pairs hurt BIRD,
  because BIRD rewards execution correctness, not SQL style.
- Reasoning distillation plus execution-based Best-of-N is the strongest story
  to package next, pending final artifact verification.

## Branches

- `main`: public product/reproducible baseline branch.
- `feat/localsql-product`: current cleanup branch for the LocalSQL product.
- `feat/bird-7b-clean`: GRPO/execution-reward research branch.
- `feat/excot-onpolicy`: reasoning-distillation and Best-of-N research branch.