Refining Zero-Shot Text-to-SQL Benchmarks via Prompt Strategies with Large Language Models
Name
applsci-15-05306.pdf
Size
7.77 MB
Format
Adobe PDF
Checksum (MD5)
255c31cb0b860b0df3c8861f0d6da452
Author(s) •
Zhou, Ruikang
Zhang, Fan
Date Issued
May 9, 2025
Journal
Applied Sciences
Publisher
Multidisciplinary Digital Publishing Institute
Citation
Zhou, R.; Zhang, F. Refining Zero-Shot Text-to-SQL Benchmarks via Prompt Strategies with Large Language Models. Appl. Sci. 2025, 15, 5306.
Version
Final published version
Abstract
Text-to-SQL leverages large language models (LLMs) for natural language database queries, yet existing benchmarks like BIRD (12,751 question–SQL pairs, 95 databases) suffer from inconsistencies—e.g., 30% of queries misalign with SQL outputs—and ambiguities that impair LLM evaluation. This study refines such datasets by distilling logically sound question–SQL pairs and enhancing table schemas, yielding a benchmark of 146 high-complexity tasks across 11 domains. We assess GPT-4o, GPT-4o-Mini, Qwen-2.5-Instruct, llama 370b, DPSK-v3 and O1-Preview in zero-shot scenarios, achieving average accuracies of 51.23%, 41.65%, 44.25%, 47.80%, and 49.10% and a peak of 78.08% (O1-Preview), respectively. Prompt-based strategies improve performance by up to 4.78%, addressing issues like poor domain adaptability and inconsistent training data interpretation. Error-annotated datasets further reveal LLM limitations. This refined benchmark ensures robust evaluation of logical reasoning, supporting reliable NLP-driven database systems.
MIT Department
MIT Kavli Institute for Astrophysics and Space Research
Terms of Use
Creative Commons Attribution
Persistent DSpace Link
DOI of Published Version
https://doi.org/10.3390/app15105306