Ilmu Komputer & AI editorial
Open AccessOA2026
Seeing Is Not Perceiving: When Synthetic Consumers Can and Cannot Pretest Visual Marketing
A stress test of generative AI agents as synthetic consumers across six canonical visual marketing experiments reveals a perception gap that demands governance.
Yi-Lin Tsai; Yung-Hsiu; Lai· 2026· DOI 10.48550/arXiv.2609.25677
The core problem
Marketers increasingly deploy generative AI agents as synthetic consumers to pretest visual assets such as logos, packaging, and advertising at a fraction of the cost of human panels. This practice rests on an untested assumption: that a model which *sees* a visual cue can also *perceive* its consumer meaning. The authors stress-test this assumption using six canonical visual marketing experiments. They vary two levers managers control: model generation (GPT-4o-mini vs. GPT-5.4-mini) and input format (plain text vs. JSON). The central question is whether synthetic consumers can reliably reproduce known human effects, and if not, under what conditions they can be steered toward human-like responses. The study integrates findings into an AI governance protocol—Calibrate, Intervene, Deploy—to guide when synthetic consumers can responsibly screen creatives and when human panels remain necessary.
Innovation
None of the 24 configurations reproduced more than two of the six human effects; the remainder were nonsignificant. The one exception was a significant reversal of the human pattern. Providing conceptual or empirical evidence through in-context learning steered average responses toward the human effect. However, steering had a limit: even when it succeeded, a configuration reproduced less than half of the natural spread of human responses, thus understating consumer heterogeneity. These results indicate that synthetic consumers can detect visual cues but often fail to perceive their consumer meaning, and that steering improves central tendency at the expense of variance.
Marketers increasingly deploy generative AI agents as synthetic consumers to pretest visual assets such as logos, packaging, and advertising at a fraction of the cost of human panels. This practice rests on an untested assumption: that a model which *sees* a visual cue can also *perceive* its consumer meaning. The authors stress-test this assumption using six canonical visual marketing experiments. They vary two levers managers control: model generation (GPT-4o-mini vs. GPT-5.4-mini) and input format (plain text vs. JSON). The central question is whether synthetic consumers can reliably reproduce known human effects, and if not, under what conditions they can be steered toward human-like responses. The study integrates findings into an AI governance protocol—Calibrate, Intervene, Deploy—to guide when synthetic consumers can responsibly screen creatives and when human panels remain necessary.
The authors conducted a systematic stress test using six canonical visual marketing experiments. For each experiment, they manipulated two factors: model generation (GPT-4o-mini and GPT-5.4-mini) and input format (plain text vs. JSON). This yielded a 2×2 design per experiment, totaling 24 configurations. Every configuration passed manipulation checks, ensuring that the models detected the intended visual differences. The dependent measure was whether the configuration reproduced the human effect size and direction. The authors also tested in-context learning interventions by providing conceptual or empirical evidence to steer average responses toward the human effect. Finally, they assessed the natural spread of responses to evaluate whether steering preserved consumer heterogeneity. The analysis compared effect replication rates, significance, and variance against human benchmarks.
Why it matters
The findings challenge the assumption that seeing equals perceiving. While manipulation checks passed, meaning extraction failed in most cases. The reversal suggests that models may rely on spurious correlations rather than consumer psychology. In-context learning can align average responses, but the underrepresentation of heterogeneity limits external validity. The authors propose an AI governance protocol: **Calibrate** (validate synthetic consumers against human benchmarks), **Intervene** (apply steering only when calibration succeeds), and **Deploy** (use synthetic consumers for screening only when heterogeneity is not critical). This protocol delineates when synthetic consumers can responsibly pretest visual marketing and when human panels remain necessary. The study underscores the need for caution and governance in AI-driven marketing research.
Who should read this
CS practitioners and researchers
Opening member content…