Ilmu Komputer & AI editorial
The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks
The core problem
Large language model (LLM) agents are increasingly deployed on long-horizon tasks in engineering and research, where a single run may involve dozens or hundreds of sequential decisions. The authors argue that the quality of these intermediate decisions—such as which hypothesis to test or which implementation to build on—determines the final outcome. They define the ability to make good long-horizon decisions as the **taste** of an agent.
Existing benchmarks measure end-to-end success on long-horizon tasks, but none directly measures taste. This gap motivates the construction of **Taste-Bench**, a benchmark of taste questions automatically derived from agent trajectories. Each question presents a **decision fork**: a point in a trajectory where multiple directions are available and one leads to a better outcome. The evaluated model must choose among these directions without seeing subsequent events. The paper's contributions are threefold: (1) a method to mine decision forks automatically from parallel attempts and detours, (2) an evaluation of frontier models on Taste-Bench, and (3) a demonstration that taste can be trained via distillation.
Innovation
The authors evaluate frontier models on Taste-Bench and report that the best model answers only **59.7%** of the questions correctly. This indicates that even state-of-the-art LLMs struggle with long-horizon decision-making when isolated from execution.
Further analysis reveals two key findings:
- **Evidence position matters**: Forks whose deciding evidence appears later in the trajectory are much harder for every model. This suggests that models have difficulty integrating information that is temporally distant from the decision point.
- **Reasoning budget does not help**: Increasing the reasoning budget (e.g., allowing more tokens for chain-of-thought) does not improve accuracy on Taste-Bench. This implies that the bottleneck is not computational but rather a lack of the relevant decision-making capability.
These results highlight that taste is a distinct and challenging capability not captured by existing benchmarks.
Why it matters
The paper demonstrates that taste can be trained. By distilling the judgment of a teacher model that has seen the outcome of a fork into a student model, the student makes better decisions on unseen tasks. Moreover, this improvement translates to end-to-end success on held-out SWE-bench Pro tasks, showing that better taste leads to better overall performance.
The authors discuss implications for agent design: improving taste may require targeted training rather than simply scaling model size or reasoning budget. They also note limitations, such as the automatic mining process potentially introducing noise, and the reliance on parallel attempts which may not always be available.
A high-level architecture of the Taste-Bench construction and distillation pipeline is shown below:
Future work could explore more sophisticated mining strategies and apply taste training to broader domains.
Who should read this
Opening member content…