Jadwal Sholat

Memuat jadwal sholat…

Computer Science editorial

Open AccessOA2026

From Schema to Signal: Retrieval-Augmented Modeling for Relational Data Analytics

RAM: A retrieval-augmented framework that fuses graph structure with attribute semantics for relational databases
Lingze Zeng; Shaofeng Cai; Changshuo Liu; Zhongle Xie; Yuncheng Wu; Beng Chin Ooi· 2026· DOI 10.48550/arXiv.2605.14464

The core problem

Relational data stored in RDBMS underpins applications across e-commerce, finance, and sociality. While deep neural networks (DNNs) excel on single-table tabular data, extending them to relational databases is difficult due to normalized multi-table structures and complex inter-table relationships. Existing methods often rely strictly on schema-defined graphs, which overlook implicit semantic signals in tuple attributes and suffer from rigid connectivity. This work proposes Retrieval-Augmented Modeling (RAM), a framework that combines graph structure with attribute semantics for relational data analytics. RAM treats tuple attributes as tokens and uses random walks to construct contextual documents, enabling information retrieval techniques to estimate semantic relevance between tuples. Building on these documents, the authors introduce two retrieval-based augmentations: ATRA (intra-table relevance for contrastive learning) and ETRA (cross-table linking of semantically related tuples to enhance graph connectivity). A layer-wise model architecture tailored for relational data—comprising attribute embedding, feature integration, and graph aggregation layers—enables expressive and flex

Innovation

The authors conduct extensive experiments on five real-world relational databases, covering diverse prediction tasks such as classification and regression. RAM consistently outperforms existing baselines, including methods that rely solely on schema-defined graphs and those that use single-table DNNs. Key findings include:
- RAM achieves state-of-the-art performance across all five databases, with significant improvements in metrics like accuracy, -score, and AUC.
- Ablation studies confirm the importance of both ATRA and ETRA augmentations; removing either leads to a noticeable drop in performance.
- The layer-wise architecture proves effective, with the graph aggregation layer benefiting from the enriched connectivity provided by ETRA.
- RAM demonstrates robustness to varying database sizes and schema complexities, showcasing its generalizability.

Specific quantitative results are not provided in the abstract, but the paper claims consistent superiority over baselines in diverse prediction tasks.

Relational data stored in RDBMS underpins applications across e-commerce, finance, and sociality. While deep neural networks (DNNs) excel on single-table tabular data, extending them to relational databases is difficult due to normalized multi-table structures and complex inter-table relationships. Existing methods often rely strictly on schema-defined graphs, which overlook implicit semantic signals in tuple attributes and suffer from rigid connectivity. This work proposes Retrieval-Augmented Modeling (RAM), a framework that combines graph structure with attribute semantics for relational data analytics. RAM treats tuple attributes as tokens and uses random walks to construct contextual documents, enabling information retrieval techniques to estimate semantic relevance between tuples. Building on these documents, the authors introduce two retrieval-based augmentations: ATRA (intra-table relevance for contrastive learning) and ETRA (cross-table linking of semantically related tuples to enhance graph connectivity). A layer-wise model architecture tailored for relational data—comprising attribute embedding, feature integration, and graph aggregation layers—enables expressive and flexible representation learning. Extensive experiments on five real-world relational databases show RAM consistently outperforms baselines across diverse prediction tasks, establishing a state-of-the-art for relational data analytics.
RAM operates in three main stages: document construction, retrieval-based augmentation, and layer-wise representation learning.

Why it matters

RAM addresses a critical gap in relational data analytics by moving beyond rigid schema-defined graphs. By treating tuple attributes as tokens and leveraging random walks to create contextual documents, it captures implicit semantic signals that are often overlooked. The retrieval-based augmentations, ATRA and ETRA, effectively inject both intra-table and cross-table semantic knowledge, enhancing representation learning. The layer-wise architecture provides a flexible and expressive framework that can be adapted to various relational schemas.

However, potential limitations include the computational overhead of random walks and retrieval, especially for large databases. The choice of IR technique (e.g., BM25) and hyperparameters like and may require tuning. Future work could explore more efficient retrieval methods, dynamic schema changes, and integration with other modalities. Overall, RAM sets a new state-of-the-art for relational data analytics, bridging the gap between graph-based and attribute-based approaches.

Who should read this

CS practitioners and researchers

Opening member content…