Computer Science editorial
Open AccessOA2026
Data-aware candidate selection in NL2SQL translation via small separating instances
A provenance-based approach for selecting the best SQL candidate when only a few options are available
Stanislav Kikot; Alexander Shulgin; Yanwei Xuยท 2026ยท DOI 10.48550/arXiv.2605.12319
The core problem
Natural language to SQL (NL2SQL) translation systems often generate multiple candidate SQL queries for a given natural language question. Selecting the correct candidate is crucial for accurate query execution. Traditional selection methods rely on consistency scores or large numbers of candidates, but in many practical scenarios, only a few candidates are available and no consistency score is provided. This paper addresses this gap by proposing a data-aware candidate selection method based on separating instances and provenance. The method aims to identify the candidate that best matches the user's intent by leveraging the database instance to find small distinguishing examples. The authors evaluate their approach on a subset of the BIRD-DEV benchmark and compare it against three natural baselines. The results demonstrate that their method significantly outperforms baselines when only two or three candidates are given and no consistency score is available. The code is available at https://github.com/staskikotx/SISelection.
Innovation
The authors evaluate their method on a subset of the BIRD-DEV benchmark, which is a large-scale dataset for NL2SQL tasks. They compare against three natural baselines: random selection, selection based on query length, and selection based on a simple heuristic. The experiments focus on scenarios where only two or three candidates are available and no consistency score is provided. The results show that SISelection significantly outperforms all baselines in these settings. For instance, when two candidates are given, the accuracy of SISelection is substantially higher than that of the baselines. The improvement is consistent across different subsets of BIRD-DEV. The authors also analyze the impact of the size of the separating instance and the provenance information. They find that using smaller instances leads to better selection accuracy, as smaller instances are more likely to capture the essential differences between queries. The provenance information further refines the selection by highlighting which parts of the queries are responsible for the differences. The code and experimental setup are available at the provided GitHub repository.
Natural language to SQL (NL2SQL) translation systems often generate multiple candidate SQL queries for a given natural language question. Selecting the correct candidate is crucial for accurate query execution. Traditional selection methods rely on consistency scores or large numbers of candidates, but in many practical scenarios, only a few candidates are available and no consistency score is provided. This paper addresses this gap by proposing a data-aware candidate selection method based on separating instances and provenance. The method aims to identify the candidate that best matches the user's intent by leveraging the database instance to find small distinguishing examples. The authors evaluate their approach on a subset of the BIRD-DEV benchmark and compare it against three natural baselines. The results demonstrate that their method significantly outperforms baselines when only two or three candidates are given and no consistency score is available. The code is available at https://github.com/staskikotx/SISelection.
The proposed method, called SISelection (Small Separating Instance Selection), operates on a set of candidate SQL queries generated for a natural language question. For each pair of candidates, the method seeks a small database instance (a set of tuples) that distinguishes them: executing the two queries on this instance yields different results. Such an instance is called a separating instance. The method then uses provenance information to determine which candidate is more likely to be correct based on the user's question. Formally, given a database and a set of candidate queries
, for each pair , we find a small instance such that . The size of is minimized to ensure the distinguishing example is simple and interpretable. Provenance is then used to trace which parts of the query contribute to the difference. The candidate that aligns best with the natural language question, as determined by a scoring function that considers the separating instances and provenance, is selected. The method is data-aware because it actively queries the database to find these instances, rather than relying solely on the query structure. The overall architecture is illustrated in the following Mermaid diagram:
Why it matters
The paper demonstrates that data-aware candidate selection using small separating instances is effective for NL2SQL translation, especially when only a few candidates are available. This is important because in real-world applications, NL2SQL systems may not generate many candidates, and consistency scores may not be available. The method's reliance on the database instance makes it robust to variations in query structure. However, the approach has limitations: finding small separating instances can be computationally expensive, especially for large databases. The authors suggest that future work could focus on optimizing the search for separating instances and extending the method to handle more candidates. The provenance-based scoring could also be improved by incorporating semantic information from the natural language question. Overall, the method provides a promising direction for improving NL2SQL accuracy in practical settings. The taxonomy candidates for this work include Architecture, Cybersecurity, Network, and Cryptography, though the primary focus is on data management and natural language processing.
Who should read this
CS practitioners and researchers
Opening member contentโฆ