Jadwal Sholat

Memuat jadwal sholat…

Ilmu Komputer & AI editorial

Open AccessOA2026

Comparative Evaluation of Static Embedding Models for HTTP Request Anomaly Detection

A benchmark of Word2Vec, FastText, and Doc2Vec within a unified single-class detection framework
Amanda Riverol; Gustavo Betarte; Rodrigo Martínez; Álvaro Pardo· 2026· DOI 10.48550/arXiv.2609.26860

The core problem

Web applications face increasing cyberattacks that exploit HTTP requests to bypass security mechanisms. Traditional Web Application Firewalls (WAFs) rely on rule-based approaches, which often suffer from high false positive rates and limited adaptability. Recent research has explored machine learning and word embedding models to enhance anomaly detection in HTTP traffic. This paper presents a benchmark for static embedding models—Word2Vec, FastText, and Doc2Vec—within a unified, single-class classification framework. The authors propose HEDA (HTTP Embedding-Based Detection Architecture), a modular detection pipeline that combines static embedding representations with single-class anomaly detection models to detect anomalies at the request level. The approach operates in an unsupervised environment, where both embedding models and detectors are trained exclusively on benign HTTP traffic. The methodology is evaluated on three datasets with heterogeneous characteristics, including both synthetic and real traffic.

Innovation

The proposed methodology is evaluated on three datasets with heterogeneous characteristics, including both synthetic and real traffic. The datasets vary in size, attack types, and traffic patterns. The performance of each embedding model is measured using detection rate (true positive rate) and false positive rate (FPR). The experimental results show that the choice of embedding representation significantly affects detection performance. FastText-based embeddings produce the most consistent results across all datasets, achieving high detection rates while keeping false positive rates under control. For instance, on the synthetic dataset, FastText achieved a detection rate of 98.5% with an FPR of 1.2%, compared to Word2Vec (95.3% detection, 2.1% FPR) and Doc2Vec (93.7% detection, 2.8% FPR). On real traffic datasets, FastText maintained robust performance, with detection rates above 96% and FPR below 2%, whereas Word2Vec and Doc2Vec showed more variability. The results indicate that FastText's character n-gram approach enhances generalization to unseen tokens and morphological variations common in HTTP requests.
Web applications face increasing cyberattacks that exploit HTTP requests to bypass security mechanisms. Traditional Web Application Firewalls (WAFs) rely on rule-based approaches, which often suffer from high false positive rates and limited adaptability. Recent research has explored machine learning and word embedding models to enhance anomaly detection in HTTP traffic. This paper presents a benchmark for static embedding models—Word2Vec, FastText, and Doc2Vec—within a unified, single-class classification framework. The authors propose HEDA (HTTP Embedding-Based Detection Architecture), a modular detection pipeline that combines static embedding representations with single-class anomaly detection models to detect anomalies at the request level. The approach operates in an unsupervised environment, where both embedding models and detectors are trained exclusively on benign HTTP traffic. The methodology is evaluated on three datasets with heterogeneous characteristics, including both synthetic and real traffic.
The HEDA architecture consists of two main stages: embedding generation and anomaly detection. First, HTTP requests are parsed and tokenized into sequences of tokens (e.g., method, path, headers, body). These token sequences are then fed into a static embedding model to produce fixed-dimensional vector representations. The embedding models considered are:

Why it matters

The findings highlight the importance of embedding model selection in HTTP anomaly detection. FastText's superior performance can be attributed to its subword information, which captures morphological patterns and handles out-of-vocabulary tokens effectively—a critical feature for HTTP requests that often contain arbitrary strings, encoded parameters, and novel attack payloads. Word2Vec and Doc2Vec, while effective, struggle with rare or unseen tokens, leading to higher false positives or missed detections. The single-class framework ensures that the system can be trained solely on benign traffic, reducing the need for labeled attack data. However, the choice of detector also impacts performance; One-Class SVM and Autoencoders performed comparably, while Isolation Forest showed slightly lower detection rates. The authors note that the threshold must be carefully tuned to balance detection and false positives, and that the approach may be extended to other protocols. Future work includes exploring contextual embeddings (e.g., BERT) and online learning to adapt to evolving attack patterns. Overall, HEDA provides a modular and effective pipeline for HTTP request anomaly detection, with FastText as the recommended static embedding model.

Who should read this

CS practitioners and researchers

Opening member content…