Jadwal Sholat

Memuat jadwal sholatโ€ฆ

Computer Science editorial

Open AccessOA2026

BatchBench: A Workload-Aware Benchmark for Autoscaling Policies in Big Data Batch Processing

A position paper proposing an open benchmarking framework to fairly compare rule-based, learned, and LLM-agent autoscalers
Venkata Krishna Prasanth Budigi; Siri Chandana Sirigiriยท 2026ยท DOI 10.48550/arXiv.2605.12272

The core problem

Autoscaling has become a baseline expectation for cloud-native big data processing. The design space has expanded from rule-based heuristics to learned controllers and, most recently, large language model (LLM) agents. Despite this growth, the community lacks a shared benchmark for comparing these paradigms. Existing evaluations rely on synthetic TPC-style queries, vendor blog posts with proprietary baselines, or narrow trace replays. Consequently, each new policy reports favorable numbers against a different baseline, on a different workload, with a different cost model, making cross-paper comparison effectively impossible.

This paper is a position paper: it proposes BatchBench, an open benchmarking framework designed to place rule-based, learned, and agentic autoscaling policies on equal experimental footing. The contribution is the design of the framework, not empirical results. The authors argue that a workload-aware benchmark is necessary to advance the field, and they outline the expected evaluation surface, open research questions, and a roadmap for a subsequent empirical paper. The reference implementation is in active development and will be released as open source.

Innovation

As a position paper, BatchBench does not present empirical results. Instead, it describes the expected evaluation surface and how the framework will enable fair comparisons. The authors outline the following expected outcomes:

- **Workload taxonomy validation**: The six workload classes will be validated against real traces using KS and EMD, ensuring that generated workloads are statistically indistinguishable from real ones.
- **Policy comparison**: The five-axis harness will produce multi-dimensional performance profiles for each policy, revealing trade-offs between cost, SLA, responsiveness, thrash, and interpretability.
- **LLM cost accounting**: By including LLM inference cost, the framework will highlight the true cost of agentic autoscalers, which may be hidden in other evaluations.
- **Reproducibility**: The standardized interface and open-source implementation will allow researchers to reproduce and extend results.

The paper also identifies open research questions, such as: How do LLM-based autoscalers compare to learned and rule-based ones under varying workload dynamics? What is the impact of LLM inference latency on scaling responsiveness? How can interpretability be

Autoscaling has become a baseline expectation for cloud-native big data processing. The design space has expanded from rule-based heuristics to learned controllers and, most recently, large language model (LLM) agents. Despite this growth, the community lacks a shared benchmark for comparing these paradigms. Existing evaluations rely on synthetic TPC-style queries, vendor blog posts with proprietary baselines, or narrow trace replays. Consequently, each new policy reports favorable numbers against a different baseline, on a different workload, with a different cost model, making cross-paper comparison effectively impossible.
This paper is a position paper: it proposes BatchBench, an open benchmarking framework designed to place rule-based, learned, and agentic autoscaling policies on equal experimental footing. The contribution is the design of the framework, not empirical results. The authors argue that a workload-aware benchmark is necessary to advance the field, and they outline the expected evaluation surface, open research questions, and a roadmap for a subsequent empirical paper. The reference implementation is in active development and will be released as open source.

Why it matters

The authors discuss the limitations of existing evaluation practices and how BatchBench addresses them. They argue that without a shared benchmark, the community cannot make cumulative progress. The five-axis evaluation harness is a key contribution because it captures multiple dimensions of autoscaling performance, preventing policies from optimizing one metric at the expense of others. The inclusion of LLM inference cost is particularly timely given the recent surge in LLM-based agents.

The standardized agent interface is another critical contribution, as it enables apples-to-apples comparisons across paradigms. The authors acknowledge that designing such an interface is challenging due to the diverse nature of autoscaling policies, but they propose a flexible API that can accommodate different policy types.

The paper concludes with a roadmap for the empirical paper that will follow. This roadmap includes: (1) implementing the reference framework, (2) populating it with representative policies from each paradigm, (3) running extensive experiments across the workload taxonomy, and (4) analyzing the results to derive insights and best practices. The authors invite the community to contribute to the open-source implementation.

A high-level architecture of BatchBench is illustrated below:

The framework is designed to be extensible and modular, allowing for future additions of workload classes and evaluation axes.

Who should read this

CS practitioners and researchers

Opening member contentโ€ฆ