Computer Science editorial
Natural-Language Workflows Are Not Software Yet: Artifact-Driven Compilation for Reliable Agent Execution
The core problem
Natural-language workflows present a compelling abstraction: domain experts write reusable procedures in prose, and agents execute them as instructions. This software-like interface, however, is not yet reliable. The core problem is that workflow descriptions typically leave data dependencies implicit. When a step says "summarize the findings," the executor must infer which prior results constitute "the findings." This inference burden compounds as workflows grow long or branch, and agents under context pressure can fail to follow instructions faithfully.
The paper frames this as an enforcement-burden problem. In conventional software, a compiler resolves dependencies, checks types, and enforces control flow before execution. In natural-language workflows, all of that burden is deferred to the agent at runtime. The authors propose **Artic**, an artifact-driven workflow compiler that transforms a natural-language workflow into a representation where each step explicitly declares the artifacts it reads and writes, constraints gate produced artifacts, and explicit control transfers route execution. This representation exposes where enforcement is hard, allowing the compiler to identi
Innovation
Artic improves the task resolve rate by **28 percentage points** over the original text workflow across the 488 problem instances from 11 real-world domain workflows. This is the headline result: making artifacts explicit and control flow explicit materially improves whether agents complete the task.
Consistency gains are larger still. Workflows compiled by Artic are **32 percentage points** more consistent in cross-model setups—meaning different models executing the same compiled workflow produce more similar outcomes—and **56 percentage points** more consistent in repeated-execution setups, where the same model runs the workflow multiple times. The repeated-execution figure is especially notable because it speaks to determinism, a property natural-language workflows notoriously lack.
The pattern across metrics suggests that the enforcement burden, not raw model capability, is the dominant source of failure in natural-language workflows. When the compiler absorbs dependency resolution and control routing, agents have less to infer and less to get wrong. The cross-model consistency gain further implies that compiled workflows are less sensitive to which model executes them, which
Why it matters
The central claim of the paper is that natural-language workflows are not software yet because they lack a compilation step. Artic supplies that step, and the results support the framing. The artifact-driven representation is the key move: by forcing each step to declare , the compiler converts an inference problem into a checking problem. Inference is where agents fail; checking is where compilers excel.
The constrained-optimization refinement of high-burden steps is a pragmatic design choice. Rather than requiring perfect LLM transformation in one pass, Artic identifies the steps most likely to be miscompiled and iterates on them. The scenario-based dry runs then provide evidence that compiled regions conform to source intent. This is a validation strategy that respects the limits of LLM-assisted transformation while still exploiting its leverage.
Several limitations and open questions remain. The evaluation covers 11 domain workflows and 488 instances; generalization to workflows with heavy external tool use, non-deterministic side effects, or adversarial inputs is not established. The faithfulness checking is decomposed into local obligations, which is tractable but may miss global inconsistencies that only appear across regions. And the 56-point repeated-execution consistency gain, while large, still leaves room for residual nondeterminism.
The broader implication is architectural. If artifact-driven compilation becomes standard, the natural-language workflow becomes a source language, and the compiled artifact graph becomes the executable. That separation mirrors the familiar split between high-level code and compiled binaries, and it suggests a division of labor: domain experts write prose, compilers enforce structure, and agents execute against explicit interfaces. The taxonomy candidates—Architecture, Cybersecurity, Network, Cryptography—are apt because the same enforcement-burden problem appears wherever agents execute procedures with implicit dependencies and branching control.
Who should read this
Opening member content…