Spaces:
Running on Zero
Running on Zero
File size: 1,259 Bytes
f2a3a7b | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 | # Research Notes
This repository started as a text-to-SQL fine-tuning research project and is now
being shaped into LocalSQL, a usable local/private SQL assistant.
## Core Result
The first strong result was Spider V1:
- Fine-tuned Qwen2.5-Coder-7B: **78.2%** result accuracy.
- Grok-4 baseline: 73.7%.
- DeepSeek-V3 baseline: 71.8%.
The method was frontier-disagreement DPO: run two strong models, execute their
SQL, and turn correctness disagreements into DPO preference pairs.
## BIRD Findings
BIRD is harder because schemas are larger, train/dev databases are disjoint, and
questions include external evidence. The important lessons so far:
- Clear correctness pairs helped the 7B model.
- Judge-resolved "both correct but stylistically different" pairs hurt BIRD,
because BIRD rewards execution correctness, not SQL style.
- Reasoning distillation plus execution-based Best-of-N is the strongest story
to package next, pending final artifact verification.
## Branches
- `main`: public product/reproducible baseline branch.
- `feat/localsql-product`: current cleanup branch for the LocalSQL product.
- `feat/bird-7b-clean`: GRPO/execution-reward research branch.
- `feat/excot-onpolicy`: reasoning-distillation and Best-of-N research branch.
|