text2sql-live / docs /research /README.md
jk200201's picture
Deploy bf16 LocalSQL Space for ZeroGPU
f2a3a7b
|
Raw
History Blame Contribute Delete
1.26 kB

A newer version of the Gradio SDK is available: 6.22.0

Upgrade

Research Notes

This repository started as a text-to-SQL fine-tuning research project and is now being shaped into LocalSQL, a usable local/private SQL assistant.

Core Result

The first strong result was Spider V1:

  • Fine-tuned Qwen2.5-Coder-7B: 78.2% result accuracy.
  • Grok-4 baseline: 73.7%.
  • DeepSeek-V3 baseline: 71.8%.

The method was frontier-disagreement DPO: run two strong models, execute their SQL, and turn correctness disagreements into DPO preference pairs.

BIRD Findings

BIRD is harder because schemas are larger, train/dev databases are disjoint, and questions include external evidence. The important lessons so far:

  • Clear correctness pairs helped the 7B model.
  • Judge-resolved "both correct but stylistically different" pairs hurt BIRD, because BIRD rewards execution correctness, not SQL style.
  • Reasoning distillation plus execution-based Best-of-N is the strongest story to package next, pending final artifact verification.

Branches

  • main: public product/reproducible baseline branch.
  • feat/localsql-product: current cleanup branch for the LocalSQL product.
  • feat/bird-7b-clean: GRPO/execution-reward research branch.
  • feat/excot-onpolicy: reasoning-distillation and Best-of-N research branch.