Qwen3.5-4b_text2sql-bird

A text-to-SQL model for SQLite, fine-tuned from Qwen/Qwen3.5-4B with LoRA on the BIRD-SQL filtered training set. The adapter is merged into the weights, so the model loads like the base model.

On the BIRD dev set, execution accuracy goes from 45.89% to 53.78% with a schema-only prompt and from 46.35% to 55.74% when column descriptions are included.

Model details

Developed by EuricoGVP
Model type Decoder-only language model, merged LoRA fine-tune
Architecture Qwen3_5ForConditionalGeneration, hybrid attention with Gated DeltaNet linear-attention layers
Parameters about 4.56B, including the vision encoder inherited from the base model
Fine-tuned from Qwen/Qwen3.5-4B
Task Natural-language question to SQLite query
Language English
License Apache 2.0

Model sources

Training data EuricoGVP/birdsql_complete_trainset
Evaluation data EuricoGVP/birdsql_complete_devset
Benchmark outputs data/benchmarks/ in this repository
Notebooks notebooks/ in this repository
Benchmark BIRD-SQL

Uses

Direct use

Generating a single SQLite query from a database schema, an optional domain hint and a question, using the prompt format below. Typical uses are research on text-to-SQL, natural-language interfaces to SQLite databases and as a component in data-analysis agents.

Out-of-scope use

  • Other SQL dialects (PostgreSQL, MySQL, T-SQL); the model was trained and evaluated on SQLite only.
  • Executing generated queries on production data without review or sandboxing.
  • Chat, reasoning or general-purpose tasks; the model was specialized for a single output format.
  • Image inputs; the vision encoder was not trained or evaluated.

How to get started

Prompt format

The model was trained with this exact system prompt and user layout. Results are only measured for this format.

SYSTEM_PROMPT = """
You are an expert SQLite data analyst.

You will be given a database schema, an optional hint, and a question.
Write a single SQLite query that answers the question.

Rules:
- Use only the tables and columns present in the schema.
- Use the hint when it is provided; it explains domain terms in the question.
- Return exactly the columns the question asks for, nothing more.
- Output only the query, inside a single ```sql code block.
- Do not explain your reasoning and do not write any text outside the code block.
"""


def build_user(ddl, evidence, question):
    hint = evidence.strip() if evidence and evidence.strip() else "None"
    return f"Database schema:\n{ddl}\n\nHint: {hint}\n\nQuestion: {question}"

The schema is the CREATE TABLE DDL in the format used by birdsql_complete_devset. Thinking mode must be disabled.

vLLM

import re
from vllm import LLM, SamplingParams

llm = LLM(
    model="EuricoGVP/Qwen3.5-4b_text2sql-bird",
    max_model_len=8192,
    language_model_only=True,
)
sampling = SamplingParams(temperature=0.0, max_tokens=2048)

messages = [
    {"role": "system", "content": SYSTEM_PROMPT},
    {"role": "user", "content": build_user(ddl, evidence, question)},
]
output = llm.chat([messages], sampling, chat_template_kwargs={"enable_thinking": False})
text = output[0].outputs[0].text

sql = re.findall(r"```(?:sql)?\s*(.*?)```", text, re.S)[-1].strip()

On GPUs without bfloat16 support, such as the T4, add dtype="float16". language_model_only=True skips loading the vision encoder.

For Transformers, load the model with the same classes as Qwen/Qwen3.5-4B and apply the chat template with enable_thinking=False.

Training details

Training data

birdsql_complete_trainset: the BIRD-SQL filtered training questions (6,601) joined with the DDL of their databases, normalized to the dev set format. Four questions using RIGHT JOIN were removed upstream, and the 383 questions of works_cycles were excluded because its schema alone exceeds 7,000 tokens. 6,214 examples were used, totalling 6.38M tokens (mean 1,027, max 3,679).

Preprocessing

Each example is the chat-templated prompt (system prompt, schema, hint, question) followed by the reference query in a ```sql block and the end-of-turn token. Prompt and answer are tokenized separately and concatenated, matching inference. Loss is computed on the answer only.

Training hyperparameters

Method LoRA, rank 16, alpha 16, dropout 0
Target modules q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj
Trainable parameters 21.2M (0.47%)
Epochs 1 (777 steps)
Effective batch size 8 (2 per GPU, 2 accumulation steps, 2 GPUs, DDP)
Optimizer AdamW
Learning rate 2e-4, linear decay, 20 warmup steps
Max sequence length 4,096 tokens
Precision fp16 mixed precision, LoRA weights in fp32
Gradient checkpointing Unsloth
Seed 3407

Speeds and compute

Hardware 2x NVIDIA Tesla T4 (16 GB), Kaggle
Training time about 6 hours
Time per step about 27 s
Peak GPU memory 11.3 GB per GPU
Software Unsloth 2026.9.12, Transformers 5.5.0, PyTorch 2.10.0 (CUDA 12.8)

Evaluation

Testing data and metric

The full BIRD dev set (1,534 questions over 11 databases, none seen in training), from birdsql_complete_devset, which includes the upstream tie corrections.

The metric is execution accuracy: predicted and reference queries run against the SQLite databases with a 30 s timeout, and results are compared as sets of rows. Two reference queries exceed the timeout, so the ceiling is 99.87%.

Base and fine-tuned models used the same prompt, temperature=0.0, max_tokens=2048 and thinking disabled, served with vLLM 0.30.0.

Results

Prompt Qwen3.5-4B This model Gain
plain (DDL only) 45.89% 53.78% +7.89 pp
desc (DDL with column descriptions) 46.35% 55.74% +9.39 pp

By difficulty:

Difficulty Questions Base, plain This model, plain Base, desc This model, desc
simple 925 53.95% 61.62% 53.19% 62.38%
moderate 464 35.56% 42.89% 37.93% 45.69%
challenging 145 27.59% 38.62% 29.66% 45.52%

By database, plain prompt:

Database Base This model
superhero 68.99% 73.64%
student_club 48.73% 68.99%
codebase_community 56.99% 65.05%
european_football_2 55.81% 62.79%
card_games 46.07% 51.83%
debit_card_specializing 43.75% 51.56%
thrombosis_prediction 34.36% 48.47%
formula_1 33.33% 44.83%
toxicology 42.07% 42.07%
financial 41.51% 35.85%
california_schools 28.09% 34.83%

Second epoch

A second epoch, starting from this adapter with learning rate 1e-4, reached 54.11% (plain) and 56.19% (desc). The difference is not significant (exact McNemar test, p = 0.76 and p = 0.64): the epochs disagree on 175 questions in the plain setting, but the changes cancel out. This repository keeps the first epoch.

Error analysis

Of the 99 plain predictions that fail to execute, 79 reference a column that does not exist, 16 are syntax errors (mostly column names with spaces written without quotes in california_schools, an error already present in the base model at the same rate), 3 use invalid column syntax and 1 times out.

Bias, risks and limitations

  • Trained and evaluated on SQLite only; other dialects are untested.
  • Prompts longer than 4,096 tokens were never seen in training.
  • The model can reference columns that do not exist; validate or execute queries in a sandbox.
  • Execution accuracy checks result sets, not query efficiency or style.
  • Scores on this dev copy are not exactly comparable to leaderboard numbers on the original release.
  • Generated queries can be destructive if executed with write permissions; run them read-only.

Citation

Please cite the BIRD-SQL benchmark, whose data was used to train and evaluate this model:

@article{li2024can,
  title   = {Can LLM Already Serve as a Database Interface? A BIg Bench for Large-Scale Database Grounded Text-to-SQLs},
  author  = {Li, Jinyang and Hui, Binyuan and Qu, Ge and Yang, Jiaxi and Li, Binhua and Li, Bowen and Wang, Bailin and Qin, Bowen and Geng, Ruiying and Huo, Nan and others},
  journal = {Advances in Neural Information Processing Systems},
  volume  = {36},
  year    = {2024}
}

License

The weights are released under Apache 2.0, the license of the base model. The training and evaluation data come from BIRD-SQL, licensed under CC BY-SA 4.0.

Downloads last month
108
Safetensors
Model size
5B params
Tensor type
BF16
·
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for EuricoGVP/Qwen3.5-4b_text2sql-bird

Finetuned
Qwen/Qwen3.5-4B
Finetuned
(917)
this model

Datasets used to train EuricoGVP/Qwen3.5-4b_text2sql-bird

Evaluation results

  • Execution accuracy (plain prompt) on BIRD-SQL dev
    self-reported
    53.780
  • Execution accuracy (prompt with column descriptions) on BIRD-SQL dev
    self-reported
    55.740