Spaces:
Running on Zero
Running on Zero
A newer version of the Gradio SDK is available: 6.22.0
Research Notes
This repository started as a text-to-SQL fine-tuning research project and is now being shaped into LocalSQL, a usable local/private SQL assistant.
Core Result
The first strong result was Spider V1:
- Fine-tuned Qwen2.5-Coder-7B: 78.2% result accuracy.
- Grok-4 baseline: 73.7%.
- DeepSeek-V3 baseline: 71.8%.
The method was frontier-disagreement DPO: run two strong models, execute their SQL, and turn correctness disagreements into DPO preference pairs.
BIRD Findings
BIRD is harder because schemas are larger, train/dev databases are disjoint, and questions include external evidence. The important lessons so far:
- Clear correctness pairs helped the 7B model.
- Judge-resolved "both correct but stylistically different" pairs hurt BIRD, because BIRD rewards execution correctness, not SQL style.
- Reasoning distillation plus execution-based Best-of-N is the strongest story to package next, pending final artifact verification.
Branches
main: public product/reproducible baseline branch.feat/localsql-product: current cleanup branch for the LocalSQL product.feat/bird-7b-clean: GRPO/execution-reward research branch.feat/excot-onpolicy: reasoning-distillation and Best-of-N research branch.