Computer Science editorial
Open AccessOA2026
Nebulon Enterprise Simulated Threats for Phishing Research (NEST-Phish): A Synthetic Enterprise Phishing Email Dataset for Behavioral and Machine-Learning Research
A synthetic enterprise phishing email dataset with matched legitimate and phishing emails, interpretable cue annotations, and human-subject categorizations for behavioral and machine-learning research.
Emily J. Winokur; Lauren S. Treiman; Allen G. Moore; Paul Schutte; Joe Ingram; Danielle N. Sanchezยท 2026ยท DOI 10.48550/arXiv.2609.04474
The core problem
Phishing remains one of the most persistent cyber threats, yet publicly shareable datasets for studying phishing in realistic enterprise email settings remain limited. This scarcity hinders reproducible research on phishing detection, human susceptibility, and explainability. To address this gap, the authors introduce a synthetic enterprise phishing email dataset built around a fictitious organization, Nebulon. The dataset spans a broad set of workplace communication themes and includes matched synthetic legitimate and phishing emails with interpretable phishing-cue annotations. Here, ``legitimate'' denotes the non-phishing class, not legitimately occurring organizational emails. The dataset is designed to support future work on phishing detection, human susceptibility, explainability, and benchmark development in enterprise-like contexts. The authors emphasize that the dataset is publicly released to enable reproducible research and to facilitate the development of robust detection systems that can operate in realistic enterprise environments.
Innovation
Human-subject categorizations show that the dataset supports meaningful variation in phishing judgments, indicating that the emails are sufficiently realistic to elicit diverse responses. Classifier evaluations demonstrate that the dataset provides learnable signal for supervised detection, with models achieving performance above chance. The matched design allows for controlled comparisons between legitimate and phishing emails, isolating the effect of phishing cues. The interpretable annotations enable explainability analyses, linking specific cues to detection outcomes. The dataset's breadth across workplace themes ensures coverage of various enterprise communication contexts. These results confirm that NEST-Phish is a valuable resource for studying phishing in enterprise-like settings. The authors report that the dataset is publicly available and can be used for benchmarking detection algorithms and behavioral studies.
Phishing remains one of the most persistent cyber threats, yet publicly shareable datasets for studying phishing in realistic enterprise email settings remain limited. This scarcity hinders reproducible research on phishing detection, human susceptibility, and explainability. To address this gap, the authors introduce a synthetic enterprise phishing email dataset built around a fictitious organization, Nebulon. The dataset spans a broad set of workplace communication themes and includes matched synthetic legitimate and phishing emails with interpretable phishing-cue annotations. Here, ``legitimate'' denotes the non-phishing class, not legitimately occurring organizational emails. The dataset is designed to support future work on phishing detection, human susceptibility, explainability, and benchmark development in enterprise-like contexts. The authors emphasize that the dataset is publicly released to enable reproducible research and to facilitate the development of robust detection systems that can operate in realistic enterprise environments.
The dataset construction follows a systematic approach to generate synthetic enterprise emails that mimic real-world workplace communication. The authors created a fictitious organization, Nebulon, and designed a broad set of workplace communication themes to ensure diversity in email content. For each theme, they generated matched pairs of legitimate (non-phishing) and phishing emails, ensuring that the phishing emails contain interpretable phishing cues. These cues are annotated to facilitate explainability research. The dataset includes human-subject categorizations, where participants labeled emails as phishing or legitimate, providing ground truth for behavioral studies. Additionally, the authors conducted classifier evaluations to demonstrate that the dataset provides learnable signal for supervised detection. The synthetic nature of the dataset ensures that no real organizational data is exposed, addressing privacy concerns while maintaining realism. The dataset is publicly released to support reproducible research.
Why it matters
The NEST-Phish dataset addresses a critical gap in phishing research by providing a publicly shareable, synthetic enterprise email dataset with matched legitimate and phishing emails and interpretable cue annotations. Its design supports both behavioral research on human susceptibility and machine-learning research on detection and explainability. The use of a fictitious organization mitigates privacy and ethical concerns while preserving ecological validity. The inclusion of human-subject categorizations and classifier evaluations demonstrates the dataset's utility for benchmarking. Future work can leverage NEST-Phish to develop more robust phishing detection models, study the impact of specific cues on susceptibility, and create explainable AI systems for cybersecurity. The dataset also enables reproducible research, as it is publicly released. Limitations include the synthetic nature of the emails, which may not capture all nuances of real enterprise communication, and the potential for overfitting to the synthetic distribution. Nonetheless, NEST-Phish represents a significant step toward standardized benchmarks for enterprise phishing research.
Who should read this
CS practitioners and researchers
Opening member contentโฆ