Jadwal Sholat

Memuat jadwal sholat…

Computer Science editorial

Open AccessOA2026

GS-QA: A Benchmark for Geospatial Question Answering

An extensible benchmark of 2,800 question-answer pairs over OpenStreetMap and Wikipedia, exposing the limits of LLM-based geospatial reasoning
Majid Saeedan; Muhammad Shihab Rashid; Ahmed Eldawy; Vagelis Hristidis· 2026· DOI 10.48550/arXiv.2605.22811

The core problem

Recent advances in Large Language Models (LLMs) have led to dramatic improvements in question answering (QA). To address the challenge of evaluating QA systems, standardized benchmarks have been introduced. This work focuses on the problem of geospatial QA, where a large collection of geospatial data is available in the form of a spatial database or other forms.

Existing work on geospatial QA benchmarks has various limitations, including a small number of questions, limited spatial predicates, narrow output types, and no multi-source reasoning. The authors present **GS-QA**, an extensible geospatial QA benchmark with **2,800 question-answer pairs** across **28 templates** on top of OpenStreetMap (OSM) and Wikipedia data. The benchmark covers a wide range of spatial objects, predicates (including directional and towards filtering), and answer types (entity names, locations, distances, directions, counts, and aggregated areas/lengths). A key feature of GS-QA is that some questions require combining information from multiple sources, e.g., geospatial information from OSM and factual information from Wikipedia.

GS-QA includes a comprehensive evaluation methodology that combines text-

Innovation

The authors evaluated nine LLM-based geospatial QA baselines on GS-QA. The results reveal a clear performance hierarchy:

- **Simple spatial predicates with entity name outputs**: All baselines achieve relatively high accuracy. For example, questions like "Which city is north of Paris?" are answered correctly in many cases, especially by GPT-4o and Claude Sonnet 4.6 with RAG or text-to-SQL.
- **Complex spatial predicates**: Accuracy drops significantly when questions involve directional filtering (e.g., "towards"), multi-step spatial reasoning, or combinations of predicates. Text-to-SQL baselines show some advantage here, but still struggle.
- **Numeric output types**: Questions requiring distances, directions, counts, or aggregated areas/lengths yield much lower accuracy. The continuous nature of these answers makes exact match inappropriate; even with distance and angular error metrics, errors are substantial.
- **Multi-source reasoning**: Questions that require combining OSM geospatial information with Wikipedia facts (e.g., "What is the population of the city north of the largest lake in X?") are the most challenging. All baselines perform poorly, often failing to retrieve or i

Recent advances in Large Language Models (LLMs) have led to dramatic improvements in question answering (QA). To address the challenge of evaluating QA systems, standardized benchmarks have been introduced. This work focuses on the problem of geospatial QA, where a large collection of geospatial data is available in the form of a spatial database or other forms.
Existing work on geospatial QA benchmarks has various limitations, including a small number of questions, limited spatial predicates, narrow output types, and no multi-source reasoning. The authors present **GS-QA**, an extensible geospatial QA benchmark with **2,800 question-answer pairs** across **28 templates** on top of OpenStreetMap (OSM) and Wikipedia data. The benchmark covers a wide range of spatial objects, predicates (including directional and towards filtering), and answer types (entity names, locations, distances, directions, counts, and aggregated areas/lengths). A key feature of GS-QA is that some questions require combining information from multiple sources, e.g., geospatial information from OSM and factual information from Wikipedia.

Why it matters

The GS-QA benchmark highlights several critical gaps in current LLM-based geospatial QA systems. First, the reliance on parametric knowledge (direct prompting) is insufficient for geospatial reasoning, as models lack up-to-date and precise spatial data. RAG improves performance by providing relevant context, but retrieval itself becomes a bottleneck for complex queries. Text-to-SQL offers a more structured approach, yet generating correct spatial SQL—especially for directional and towards predicates—remains difficult.

Second, the evaluation methodology reveals that text-based metrics alone are inadequate for geospatial answers. The inclusion of distance and angular error provides a more nuanced view, but even these metrics show that models often produce answers that are spatially plausible but numerically inaccurate. For instance, a predicted distance of 10 km versus a gold distance of 12 km may be acceptable in some contexts, but current models frequently err by larger margins.

Third, multi-source reasoning is a major challenge. The need to combine geospatial data from OSM with factual data from Wikipedia requires not only spatial reasoning but also entity linking and information fusion. The poor performance on these questions suggests that integrated systems are needed.

The authors conclude that geospatial QA remains an open problem. Future work should focus on:

- Developing more sophisticated spatial reasoning capabilities in LLMs.
- Improving retrieval and integration of heterogeneous data sources.
- Designing benchmarks that cover even more diverse spatial predicates and answer types.
- Exploring hybrid architectures that combine LLMs with spatial databases and knowledge graphs.

GS-QA provides a solid foundation for these efforts, with its extensible design and comprehensive evaluation suite.

Who should read this

CS practitioners and researchers

Opening member content…