qwen3.5-2b-sql / README.md
Vicen-te's picture
Super-squash branch 'main' using huggingface_hub
8619cb0
|
Raw
History Blame Contribute Delete
1.4 kB
metadata
license: apache-2.0
base_model: Qwen/Qwen3.5-2B
library_name: transformers
tags:
  - text2sql
  - sql
  - qwen3.5
  - fine-tuned
language:
  - en
pipeline_tag: text-generation
datasets:
  - Vicen-te/sql-create-context-mini

Qwen3.5-2B · SQL (merged)

Qwen/Qwen3.5-2B with a LoRA SQL adapter merged in. Drop-in replacement for the base — same architecture, same tokenizer, no PEFT runtime dependency.

Usage with transformers

from transformers import AutoModelForCausalLM, AutoTokenizer

tok = AutoTokenizer.from_pretrained("Vicen-te/qwen3.5-2b-sql")
model = AutoModelForCausalLM.from_pretrained("Vicen-te/qwen3.5-2b-sql", dtype="auto", device_map="auto")

Usage with vLLM

vllm serve Vicen-te/qwen3.5-2b-sql --max-model-len 4096 --served-model-name sql-ft

Training

  • Base model: Qwen/Qwen3.5-2B
  • Method: LoRA (rank=16, α=32) → merged via peft.merge_and_unload()
  • Dataset: Vicen-te/sql-create-context-mini — 300 train / 200 eval
  • Recipe: 3 epochs, bf16, effective batch 16, cosine LR 2e-4

Evaluation

Compared against the base model on a held-out 200-example split. See the project repo for the full report (executable accuracy, exact match, BLEU, latency, 4-bit quantization trade-off).

License

Apache 2.0, inherited from the base model.