Computer Science editorial
Residual Skill Optimization for Text-to-SQL Ensembles
The core problem
Text-to-SQL systems have increasingly adopted ensemble strategies: rather than relying on a single generated SQL candidate, they draw multiple candidates and select one, typically via a verifier or voting mechanism. The theoretical ceiling of such ensembles is bounded by Pass@K, the probability that at least one of K candidates is correct. Existing methods source diversity heuristically through stochastic decoding (e.g., temperature sampling) or prompt variants, which often yields candidate sets dominated by correlated failures—if the base model has a systematic misunderstanding of a schema or dialect, all K candidates may share the same error.
This paper introduces **DivSkill-SQL**, a residual skill optimization framework that builds complementary agentic Text-to-SQL ensembles without any model fine-tuning. The key insight is to treat each new skill as a residual learner: it is optimized specifically on examples that the current skill ensemble fails on, thereby provably targeting its marginal contribution to Pass@K. The framework is evaluated on Spider2-Lite and BIRD-Critic, across two base models (Opus-4.6 and GPT-5.4) and three dialects (Snowflake, BigQuery, SQLite).
Innovation
On Spider2-Lite, DivSkill-SQL improves selected accuracy by up to **+11.1 points on Snowflake** and **+8.3 points on BigQuery** over the strongest ensemble baseline. These gains are consistent across two base models: Opus-4.6 and GPT-5.4. The framework also shows cross-dialect transfer: skills optimized on one dialect (e.g., Snowflake) improve performance on other dialects (BigQuery, SQLite) without retraining. Furthermore, transfer to a different task formulation, BIRD-Critic, yields a **+2.6 point** improvement.
Error diagnostics reveal that the gains are not merely from surface-form variation. The number of hallucinated schema references and function calls is reduced by up to **3x**, indicating that the complementary skills are genuinely more reliable. This is a key result: the ensemble is not just more diverse, but more accurate at the level of individual candidates.
A summary of the main results is shown below:
| Dataset / Dialect | Baseline (best ensemble) | DivSkill-SQL | Improvement |
|-------------------|--------------------------|--------------|-------------|
| Spider2-Lite (Snowflake) | — | — | +11.1 pts |
| Spider2-Lite (BigQuery) | — | — | +8.3 pts |
| BIRD-Critic |
Why it matters
The core contribution of DivSkill-SQL is the shift from heuristic diversity to **residual skill optimization**. By explicitly targeting the failure modes of the current ensemble, the framework provably increases the marginal contribution of each new skill to Pass@K. This is in contrast to stochastic decoding or prompt variants, which often produce correlated failures because they sample from the same underlying model distribution.
The cross-dialect and cross-task transfer results are particularly promising. They suggest that the residual skills learn generalizable SQL reasoning patterns—such as handling complex joins, subqueries, or schema linking—that are not tied to a specific dialect or task. This has practical implications: a single set of skills can be deployed across multiple database systems without retraining, reducing engineering overhead.
The reduction in hallucinated schema references and function calls (up to 3x fewer) indicates that the ensemble is not just more diverse but also more reliable. This is important because in real-world Text-to-SQL applications, hallucinations can lead to incorrect or even dangerous queries.
One limitation is that the framework requires an iterative optimization loop, which may be computationally expensive. However, the authors note that no model fine-tuning is required, which lowers the barrier to adoption. Future work could explore more efficient ways to identify residual failures or to automatically determine the number of skills needed.
Overall, DivSkill-SQL represents a significant step forward in building robust Text-to-SQL ensembles, with strong empirical results and a principled approach to diversity.
Who should read this
Opening member content…