Jadwal Sholat

Memuat jadwal sholat…

Computer Science editorial

Open AccessOA2026

Polars inside Intel SGX2 Enclaves: An Empirical Study of Confidential Analytical Query Processing

An IMRAD digest of the empirical evaluation of an Arrow-native DataFrame engine under confidential computing, revealing load-path amplification and API-level optimization as first-order performance determinants.
Wei Wang; Burns Smith; Kenny Leftin· 2026· DOI 10.48550/arXiv.2605.21797

The core problem

Trusted Execution Environments (TEEs) have renewed interest in confidential analytics, yet most prior evaluations focus on SQL database engines or earlier SGX generations. This paper studies an Arrow-native DataFrame engine, Polars, running inside Intel SGX2 enclaves via Gramine on TPC-H SF30 with Azure Blob Storage. The authors aim to separate compute overhead from data-ingestion overhead by reporting both the standard TPC-H power score and a query-only variant that removes table-loading time. The study spans four dataset-width configurations (approximately 22–73 GB) and compares Polars' lazy and eager APIs under the same TEE setting. The central research question is whether SGX2 can support Arrow-native analytical processing with a similar order of security overhead as observed in recent SQL-engine studies, and what factors dominate end-to-end performance.

Innovation

Across the four dataset-width configurations (approximately 22–73 GB), end-to-end overhead remains nearly constant at 1.49–1.56×. However, this composite metric obscures two distinct behaviors: query-only overhead declines from 1.51–1.52× to 1.43–1.44×, whereas table-loading overhead rises from 2.27× to 4.07×. For the len130 configuration, the median per-query SGX slowdown is 1.45× with a maximum of 2.57×. A small set of queries exhibits pronounced run-to-run spikes consistent with stateful EPC pressure. Comparing Polars' lazy and eager APIs under the same TEE setting, lazy execution is 2.25–2.27× faster overall, while eager execution fails with out-of-memory errors at 41 GB and above. These results are summarized in the following table:

| Configuration | End-to-end overhead | Query-only overhead | Table-loading overhead |
|---------------|---------------------|---------------------|------------------------|
| ~22 GB | 1.49–1.56× | 1.51–1.52× | 2.27× |
| ~73 GB | 1.49–1.56× | 1.43–1.44× | 4.07× |

For len130, per-query slowdown: median 1.45×, max 2.57×. Lazy vs. eager: lazy 2.25–2.27× faster; eager

Trusted Execution Environments (TEEs) have renewed interest in confidential analytics, yet most prior evaluations focus on SQL database engines or earlier SGX generations. This paper studies an Arrow-native DataFrame engine, Polars, running inside Intel SGX2 enclaves via Gramine on TPC-H SF30 with Azure Blob Storage. The authors aim to separate compute overhead from data-ingestion overhead by reporting both the standard TPC-H power score and a query-only variant that removes table-loading time. The study spans four dataset-width configurations (approximately 22–73 GB) and compares Polars' lazy and eager APIs under the same TEE setting. The central research question is whether SGX2 can support Arrow-native analytical processing with a similar order of security overhead as observed in recent SQL-engine studies, and what factors dominate end-to-end performance.

The experimental setup runs Polars inside Intel SGX2 enclaves using Gramine, a library OS for unmodified Linux applications. The workload is TPC-H SF30, with data stored in Azure Blob Storage. Four dataset-width configurations are tested, ranging from approximately 22 GB to 73 GB. Two metrics are reported: the standard TPC-H power score (end-to-end) and a query-only variant that excludes table-loading time. Per-query slowdowns are measured for the len130 configuration. The lazy and eager APIs of Polars are compared under identical TEE conditions. The overhead is quantified as a slowdown factor relative to native execution. For a given query , the slowdown is

, where is the execution time. The end-to-end overhead and query-only overhead are defined as:

Why it matters

The findings reveal that SGX2 can support Arrow-native analytical processing with a similar order of security overhead as reported in the recent DuckDB-SGX2 study. However, the constant end-to-end overhead masks a critical trade-off: as dataset width increases, query-only overhead decreases while table-loading overhead more than doubles. This load-path amplification suggests that data ingestion becomes a dominant bottleneck in confidential analytics, potentially due to encryption and integrity verification costs during loading. The per-query variability, with spikes up to 2.57×, indicates that stateful EPC pressure can cause unpredictable performance, which may be mitigated by tuning enclave page cache size or query scheduling. The stark difference between lazy and eager execution—lazy being 2.25–2.27× faster and avoiding OOM at larger scales—highlights that API-level optimization is a first-order determinant of end-to-end performance. Lazy execution likely reduces memory footprint by deferring materialization, which is crucial in the constrained enclave environment. These results imply that future confidential analytics systems should prioritize optimizing the load path and adopting lazy evaluation strategies. The study is limited to TPC-H SF30 and Azure Blob Storage; generalization to other workloads and storage backends remains to be tested.

Who should read this

CS practitioners and researchers

Opening member content…