Computer Science editorial
GS-QA: A Benchmark for Geospatial Question Answering
The core problem
Recent advances in Large Language Models (LLMs) have led to dramatic improvements in question answering (QA). To address the challenge of evaluating QA systems, standardized benchmarks have been introduced. This work focuses on the problem of geospatial QA, where a large collection of geospatial data is available in the form of a spatial database or other forms.
Existing work on geospatial QA benchmarks has various limitations, including a small number of questions, limited spatial predicates, narrow output types, and no multi-source reasoning. The authors present **GS-QA**, an extensible geospatial QA benchmark with **2,800 question-answer pairs** across **28 templates** on top of OpenStreetMap (OSM) and Wikipedia data. The benchmark covers a wide range of spatial objects, predicates (including directional and towards filtering), and answer types (entity names, locations, distances, directions, counts, and aggregated areas/lengths). A key feature of GS-QA is that some questions require combining information from multiple sources, e.g., geospatial information from OSM and factual information from Wikipedia.
GS-QA includes a comprehensive evaluation methodology that combines text-
Innovation
The authors evaluated nine LLM-based geospatial QA baselines on GS-QA. The results reveal a clear performance hierarchy:
- **Simple spatial predicates with entity name outputs**: All baselines achieve relatively high accuracy. For example, questions like "Which city is north of Paris?" are answered correctly in many cases, especially by GPT-4o and Claude Sonnet 4.6 with RAG or text-to-SQL.
- **Complex spatial predicates**: Accuracy drops significantly when questions involve directional filtering (e.g., "towards"), multi-step spatial reasoning, or combinations of predicates. Text-to-SQL baselines show some advantage here, but still struggle.
- **Numeric output types**: Questions requiring distances, directions, counts, or aggregated areas/lengths yield much lower accuracy. The continuous nature of these answers makes exact match inappropriate; even with distance and angular error metrics, errors are substantial.
- **Multi-source reasoning**: Questions that require combining OSM geospatial information with Wikipedia facts (e.g., "What is the population of the city north of the largest lake in X?") are the most challenging. All baselines perform poorly, often failing to retrieve or i
Why it matters
The GS-QA benchmark highlights several critical gaps in current LLM-based geospatial QA systems. First, the reliance on parametric knowledge (direct prompting) is insufficient for geospatial reasoning, as models lack up-to-date and precise spatial data. RAG improves performance by providing relevant context, but retrieval itself becomes a bottleneck for complex queries. Text-to-SQL offers a more structured approach, yet generating correct spatial SQL—especially for directional and towards predicates—remains difficult.
Second, the evaluation methodology reveals that text-based metrics alone are inadequate for geospatial answers. The inclusion of distance and angular error provides a more nuanced view, but even these metrics show that models often produce answers that are spatially plausible but numerically inaccurate. For instance, a predicted distance of 10 km versus a gold distance of 12 km may be acceptable in some contexts, but current models frequently err by larger margins.
Third, multi-source reasoning is a major challenge. The need to combine geospatial data from OSM with factual data from Wikipedia requires not only spatial reasoning but also entity linking and information fusion. The poor performance on these questions suggests that integrated systems are needed.
The authors conclude that geospatial QA remains an open problem. Future work should focus on:
- Developing more sophisticated spatial reasoning capabilities in LLMs.
- Improving retrieval and integration of heterogeneous data sources.
- Designing benchmarks that cover even more diverse spatial predicates and answer types.
- Exploring hybrid architectures that combine LLMs with spatial databases and knowledge graphs.
GS-QA provides a solid foundation for these efforts, with its extensible design and comprehensive evaluation suite.
Who should read this
Opening member content…