Computer Science editorial
Open AccessOA2026
CogScale: Scalable Benchmark for Sequence Processing
A lightweight, parametrizable suite of 14 synthetic tasks for rapid architectural validation
Yannis Bendi-Ouis; Romain de Coudenhove; Xavier Hinautยท 2026ยท DOI 10.48550/arXiv.2605.19758
The core problem
The ability to maintain and manipulate information over time is fundamental to both biological and artificial intelligence. While modern models excel in domains like natural language processing, evaluating novel architectures for sequential information processing remains computationally expensive and slow. Testing a new design often demands scaling to massive datasets and models, incurring vast computational costs and hindering iteration speed. To address this, Bendi-Ouis, de Coudenhove, and Hinaut propose CogScale, a benchmark of 14 scalable synthetic tasks designed to isolate and evaluate specific cognitive and memory abilities at parametrizable scales. CogScale offers a standardized, lightweight framework that allows researchers to rapidly validate architectural innovations before committing to large-scale training. The benchmark aims to bridge the gap between small-scale diagnostic tasks and full-scale real-world evaluations, providing a controlled environment to probe memory retention, reasoning complexity, and scalability.
Innovation
The baseline evaluations reveal distinct performance profiles across architectures. Under strict parameter budgets (1k, 10k, 100k), classical recurrent neural networks (GRU, LSTM) and Echo State Networks excel at basic retention tasks, demonstrating strong short-term memory capabilities. However, as reasoning complexity and task difficulty scale, only attention mechanisms (Transformer Decoder, Transformer Encoder-Decoder) and modern state-space models (Mamba) consistently maintain high performance. The xLSTM shows intermediate behavior, outperforming classical RNNs on some tasks but not reaching the consistency of attention-based and state-space models. These results highlight a trade-off between parameter efficiency and scalability: while smaller models can handle simple retention, they struggle with tasks requiring complex reasoning over longer sequences. The benchmark effectively discriminates between architectures, providing a clear signal for researchers to select appropriate models based on task demands.
The ability to maintain and manipulate information over time is fundamental to both biological and artificial intelligence. While modern models excel in domains like natural language processing, evaluating novel architectures for sequential information processing remains computationally expensive and slow. Testing a new design often demands scaling to massive datasets and models, incurring vast computational costs and hindering iteration speed. To address this, Bendi-Ouis, de Coudenhove, and Hinaut propose CogScale, a benchmark of 14 scalable synthetic tasks designed to isolate and evaluate specific cognitive and memory abilities at parametrizable scales. CogScale offers a standardized, lightweight framework that allows researchers to rapidly validate architectural innovations before committing to large-scale training. The benchmark aims to bridge the gap between small-scale diagnostic tasks and full-scale real-world evaluations, providing a controlled environment to probe memory retention, reasoning complexity, and scalability.
CogScale comprises 14 synthetic tasks, each targeting distinct cognitive and memory faculties. These tasks are parametrizable in scale and difficulty, enabling systematic evaluation across a spectrum of challenges. To establish baselines, the authors evaluated seven architectures: Gated Recurrent Unit (GRU), Long Short-Term Memory (LSTM), xLSTM, Echo State Network (ESN), Mamba, Transformer Decoder, and Transformer Encoder-Decoder. All models were trained under strict parameter budgets of 1k, 10k, and 100k parameters, ensuring fair comparison and highlighting efficiency. The tasks are designed to be lightweight, allowing rapid experimentation. The evaluation protocol measures performance across different difficulty levels and scales, providing insights into each architecture's ability to handle increasing complexity. The benchmark's modular design facilitates the isolation of specific capabilities, such as short-term memory, long-term retention, and compositional reasoning.
Why it matters
The findings underscore the importance of scalable benchmarks like CogScale in guiding architectural innovation. By isolating cognitive and memory abilities, CogScale enables rapid iteration and informed decision-making before investing in large-scale training. The superior performance of attention and state-space models on complex tasks suggests that mechanisms for dynamic information routing and efficient long-range dependency modeling are crucial for advanced sequence processing. Conversely, the success of classical RNNs and ESNs on basic retention tasks indicates their continued relevance in resource-constrained scenarios. The parametrizable nature of CogScale allows researchers to tailor evaluations to specific hypotheses, fostering a deeper understanding of model strengths and weaknesses. Future work could extend the benchmark to include more diverse tasks and architectures, further enriching the landscape of sequence processing research. Ultimately, CogScale provides a standardized, lightweight tool that accelerates the development of more capable and efficient sequential models.
Who should read this
CS practitioners and researchers
Opening member contentโฆ