Jadwal Sholat

Memuat jadwal sholatโ€ฆ

Computer Science editorial

Open AccessOA2026

Engine-Transfer-Bench: An Evidence-Based Benchmark for Document Compilation Engine Selection

A pinned, host-tagged harness for comparing pdfLaTeX, XeLaTeX, LuaLaTeX, Tectonic, Typst, and pandoc PDF backends across reliability, latency, text consistency, and failure modes
Prajwal S. Venkateshmurthyยท 2026ยท DOI 10.48550/arXiv.2608.18329

The core problem

Selecting a document compilation engine is a recurring infrastructure decision for researchers, publishers, and CI/CD maintainers, yet there is no shared framework for comparing pdfLaTeX, XeLaTeX, LuaLaTeX, Tectonic, Typst, and pandoc PDF backends. The paper identifies this gap and introduces **Engine-Transfer-Bench (ETB)**, an evidence-based benchmark built from **1,784 open documents**, four tasks covering reliability, latency, text consistency, and failures, a pinned harness, and host-tagged multi-OS results. The central question is not merely which engine is fastest, but which engine behaves predictably when documents move between operating systems and package distributions. ETB is released together with **ETB-Porta**, a recommender and portability gate, and a public cross-OS harness as shared infrastructure. The benchmark's design treats compilation as a reproducible systems problem: the same document, the same harness, and the same host tags should yield comparable evidence. This framing matters because engine choice is often made from anecdotal experience or single-platform testing, which can hide distribution-specific failure modes. By pinning the harness and tagging hosts,

Innovation

The headline result is a stark contrast between engine families. On GitHub Actions, **Tectonic success is stable within 0.9 percentage points (96.3-97.2%)** across macOS, Ubuntu, and Windows. In contrast, **classic TeX Live-style engines vary by 12-20 percentage points** according to distribution policy, whether that policy is Ubuntu apt, MiKTeX auto-install, or macOS BasicTeX. This means that the same engine can appear reliable or unreliable depending on the host's package provisioning strategy. On the **702 portable LaTeX documents**, all tested engines succeed at **100%**, so latency becomes the deciding factor rather than reliability. Failures are concentrated in **107 engine-specific templates**, which are not portable across engines by construction. The text-consistency metric, validated on **50 pairs**, reaches **94% precision** for real content divergence, indicating that S_pdf can distinguish genuine output differences from benign rendering variation. The results also show that within ETB, **failures are architectural**, involving fonts, layout, and assets, rather than missing packages on a provisioned host. This is an important diagnostic shift: provisioning more packages
Selecting a document compilation engine is a recurring infrastructure decision for researchers, publishers, and CI/CD maintainers, yet there is no shared framework for comparing pdfLaTeX, XeLaTeX, LuaLaTeX, Tectonic, Typst, and pandoc PDF backends. The paper identifies this gap and introduces **Engine-Transfer-Bench (ETB)**, an evidence-based benchmark built from **1,784 open documents**, four tasks covering reliability, latency, text consistency, and failures, a pinned harness, and host-tagged multi-OS results. The central question is not merely which engine is fastest, but which engine behaves predictably when documents move between operating systems and package distributions. ETB is released together with **ETB-Porta**, a recommender and portability gate, and a public cross-OS harness as shared infrastructure. The benchmark's design treats compilation as a reproducible systems problem: the same document, the same harness, and the same host tags should yield comparable evidence. This framing matters because engine choice is often made from anecdotal experience or single-platform testing, which can hide distribution-specific failure modes. By pinning the harness and tagging hosts, ETB makes the transferability of engine behavior an explicit, measurable property rather than an assumption.

ETB evaluates engines across **GitHub Actions** hosts running **macOS, Ubuntu, and Windows**, with **N = 4,211 compiles per host**. The benchmark defines four tasks: reliability, latency, text consistency, and failures. Reliability is measured as successful compilation rate; latency captures time-to-PDF; text consistency uses the **S_pdf** metric to detect real content divergence; and failure analysis classifies the architectural causes of unsuccessful compiles. The harness is pinned so that engine versions, package sets, and host images remain comparable across runs. A key methodological choice is the separation between **portable LaTeX documents** and **engine-specific templates**. On **702 portable LaTeX documents**, the tested engines succeed at **100%**, which isolates latency as the primary selection factor. Failures concentrate in **107 engine-specific templates**, where font, layout, and asset assumptions differ. The paper also validates the text-consistency metric with a **50-pair validation**, achieving **94% precision** for real content divergence. Formally, if denotes the text-consistency score between PDF outputs and , the validation targets the decision rule

for a threshold , with precision
on the 50-pair set. The benchmark's multi-OS, host-tagged design can be summarized as follows:

Why it matters

The analysis reframes engine selection as a portability problem. Tectonic's narrow 0.9 percentage-point band suggests that its self-contained, pinned approach reduces host-dependent variance, while classic TeX Live-style engines inherit the policy and package-resolution behavior of their distributions. The 12-20 percentage-point spread is therefore not an engine defect but a distribution-policy effect, and it has direct consequences for CI/CD reproducibility. The finding that portable LaTeX documents succeed at 100% across engines implies that for well-behaved inputs, latency should dominate selection; reliability differences only emerge when documents depend on engine-specific templates, fonts, layout assumptions, or assets. The 94% precision of S_pdf supports using text consistency as a gate for detecting real content divergence, though the remaining false-positive rate should be considered in automated publishing pipelines. ETB-Porta operationalizes these findings as a recommender and portability gate, allowing maintainers to choose an engine based on host tags and document portability class rather than anecdote. The broader implication is that benchmark infrastructure for document compilation must be pinned, host-tagged, and multi-OS to produce transferable evidence. The taxonomy candidates for this work are **Architecture**, **Cybersecurity**, **Network**, and **Cryptography**, reflecting its systems-oriented, reproducibility-focused contribution. Future work can extend the harness to additional engines and distribution policies, and refine the S_pdf threshold to reduce false positives in content-divergence detection.

Who should read this

CS practitioners and researchers

Opening member contentโ€ฆ