Buckets:
| # Reproduction bundle: Beyond Text-to-SQL / Squirrel Benchmark | |
| Paper: arXiv:2601.18119 — "Beyond Text-to-SQL: Can LLMs Really Debug Enterprise ETL SQL?" | |
| Logbook: https://huggingface.co/spaces/Firemedic15/repro-beyond-text-to-sql-can-llms-really-debug-enterprise-etl-sql | |
| ## What's in here | |
| - `scripts/generate_benchmark.py` — synthesizes a 10-domain "Squirrel-mini" toy benchmark | |
| (long, multi-CTE enterprise ETL SQL scripts with taxonomy-guided syntax/semantic bugs | |
| injected), standing in for the paper's unreleased real 985-task corpus. | |
| - `scripts/metrics.py` — script-complexity stats (lines/tokens/functions/AST width-depth) | |
| and EM / GM-proxy / MB-proxy metrics, documenting where and why our sqlglot-based proxies | |
| differ from the paper's Apache-Calcite-based Graph Match. | |
| - `scripts/eval_models.py` — runs the paper's own evaluation prompts (Appendix J.3) against | |
| any model available through Hugging Face Inference Providers. | |
| - `scripts/recompute_metrics.py` — recompute EM/GM/MB on a stored `eval_results.jsonl` after | |
| a metrics fix, without re-calling any models. | |
| - `scripts/validate_icml_logbook.py` — the challenge's logbook structure validator. | |
| - `data/squirrel_mini_seeds.json`, `data/squirrel_mini_tasks.json` — the generated toy corpus | |
| and its 20 syntax/semantic debugging tasks. | |
| - `outputs/eval_results.jsonl` — full local run: 4 models × 20 tasks = 80 records | |
| (model, task, EM, GM-proxy, MB-proxy, raw prediction). | |
| - `poster/poster.html`, `poster/images/` — the reproduction poster source (posterly template, | |
| filled with this reproduction's real content) and its QR code. | |
| ## How to re-run | |
| ```bash | |
| uv venv --python 3.12 .venv && source .venv/bin/activate | |
| uv pip install huggingface_hub sqlglot tiktoken | |
| # 1. Regenerate the toy benchmark | |
| python scripts/generate_benchmark.py | |
| # 2. Evaluate any HF Inference Providers model against it | |
| python scripts/eval_models.py --models Qwen/Qwen2.5-Coder-32B-Instruct deepseek-ai/DeepSeek-V3-0324 | |
| # 3. Recompute metrics after a metrics.py change, without re-calling models | |
| python scripts/recompute_metrics.py outputs/eval_results.jsonl | |
| ``` | |
| Requires an `HF_TOKEN` with Inference Providers access (`hf auth login`). The same scripts | |
| were also run unmodified on Hugging Face Jobs — see the Conclusion page of the logbook for | |
| Job URLs, exact commands, hardware, and cost. | |
| ## What this bundle does NOT contain | |
| The real Squirrel Benchmark (469 Squirrel-Syntax + 516 Squirrel-Semantic tasks) and | |
| Claude-4-Sonnet outputs are not included because neither is publicly available — the paper's | |
| own Ethics Statement says the benchmark is "scheduled for public release upon acceptance," | |
| and no dataset/checkpoint/GitHub repo is linked from the paper, arXiv, or OpenReview. | |
Xet Storage Details
- Size:
- 2.76 kB
- Xet hash:
- 1b7d404cb55815932b9a58707adf8dc6d198a2b00ca5e6ca16e9ae84193c3066
·
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.