Computer Science editorial
OxyEcomBench: Benchmarking Multimodal Foundation Models across E-Commerce Ecosystems
The core problem
Large language models (LLMs) and multimodal large language models (MLLMs) have become indispensable across a wide range of applications. E-commerce, however, poses distinctive challenges that diverge substantially from the general world knowledge these models are primarily trained on. These challenges include intricate domain knowledge, long-tail product evidence, heterogeneous visual data, and the interplay among multiple stakeholder roles. As a result, a notable gap often emerges between a model's open-domain performance and its e-commerce performance.
Existing e-commerce benchmarks typically adopt a single stakeholder perspective, target a narrow set of tasks, or address isolated challenges, making it difficult to holistically assess a model's understanding of the full e-commerce pipeline. To systematically quantify this gap, the authors introduce **OxyEcomBench**, a unified multimodal benchmark comprising approximately **6,300 high-quality instances** for real-world bilingual Chinese–English e-commerce. The benchmark is designed to jointly cover platform operators, merchants, and customers across **6 capability aspects** and **29 tasks**, supporting text-only and mixed-modalit
Innovation
Evaluations on **20 mainstream LLMs and MLLMs** show that even the leading models attain modest performance on OxyEcomBench. The performance gaps between models narrow on this benchmark, suggesting that insufficient e-commerce-specific knowledge infusion mutes the advantages of advanced general-purpose models in this domain.
Key quantitative findings include:
- The benchmark comprises approximately 6,300 instances, with a balanced representation across stakeholder roles and capability aspects.
- All 29 tasks are annotated with the P0–P3 difficulty rubric, enabling fine-grained analysis of model performance across difficulty levels.
- The benchmark prioritizes visually salient multimodal cases, where key evidence is in images rather than text, to test models' multimodal reasoning.
- Despite the diversity of models evaluated, top performance remains modest, indicating a substantial gap between open-domain and e-commerce-specific capabilities.
Why it matters
The results highlight a critical limitation of current multimodal foundation models: their general-purpose training does not adequately prepare them for the specialized demands of e-commerce. The narrowing of performance gaps on OxyEcomBench suggests that when domain-specific knowledge is required, the advantages of larger or more advanced models diminish. This points to the need for targeted knowledge infusion and domain adaptation strategies.
OxyEcomBench addresses the limitations of prior benchmarks by jointly covering multiple stakeholder perspectives, a wide array of tasks, and integrated challenges. Its difficulty-aware design and emphasis on visually salient cases ensure that it rigorously tests models' ability to handle real-world e-commerce scenarios. The benchmark serves as a diagnostic tool to identify where models fail and to guide future research in developing more robust e-commerce AI systems.
Future work could extend the benchmark to additional languages and platforms, incorporate dynamic updates to reflect evolving e-commerce trends, and explore fine-tuning methods that effectively inject e-commerce-specific knowledge into foundation models.
Who should read this
Opening member content…