Jadwal Sholat

Memuat jadwal sholat…

Computer Science editorial

Open AccessOA2026

3D Primitives are a Spatial Language for VLMs

Code-CoT and S³-FT: Routing spatial reasoning through executable geometric primitives
Junze Liu; Kun Qian; Florian Dubost; Kai Zhong; Arvind Srinivasan; Nan Chen; Anping Wang; Sam Zhang; Alejandro Mottini; Qingjun Cui; Tian Wang· 2026· DOI 10.48550/arXiv.2605.12586

The core problem

Vision-language models (VLMs) exhibit a striking paradox: they can generate executable code that reconstructs a 3D scene from geometric primitives with correct object counts, classes, and approximate positions, yet the same models fail at simpler spatial questions on the same image. The authors argue that 3D geometric primitives (cubes, spheres, cylinders, expressed in executable code) serve as a powerful intermediate representation for spatial understanding. They exploit this insight through three contributions: a benchmark (SpatialBabel), a training-free inference strategy (Code-CoT), and a self-supervised fine-tuning recipe (S³-FT). The work positions geometric primitives in code as both a diagnostic tool and a transferable spatial vocabulary for VLMs.

Innovation

Training on primitive images alone, S³-FT improves Qwen3-VL-8B by to on SpatialBabel-Primitive-QA, on CV-Bench-2D, and on HallusionBench; the recipe transfers across model families. Code-CoT yields up to on SpatialBabel-QA-Score and on CV-Bench-3D for models with strong coding capabilities. SpatialBabel shows that object-detection can vary by up to across scene-code languages for a single model.
Vision-language models (VLMs) exhibit a striking paradox: they can generate executable code that reconstructs a 3D scene from geometric primitives with correct object counts, classes, and approximate positions, yet the same models fail at simpler spatial questions on the same image. The authors argue that 3D geometric primitives (cubes, spheres, cylinders, expressed in executable code) serve as a powerful intermediate representation for spatial understanding. They exploit this insight through three contributions: a benchmark (SpatialBabel), a training-free inference strategy (Code-CoT), and a self-supervised fine-tuning recipe (S³-FT). The work positions geometric primitives in code as both a diagnostic tool and a transferable spatial vocabulary for VLMs.
The paper introduces three components.

Why it matters

The results establish geometric primitives in code as both a diagnostic and a transferable spatial vocabulary for VLMs. The paradox—strong code-based reconstruction but weak direct spatial question answering—suggests that VLMs possess latent spatial knowledge that is not reliably accessible through natural language queries alone. By forcing reasoning through executable primitive code, Code-CoT and S³-FT unlock this knowledge without human labels or teacher models. The cross-family transfer of S³-FT indicates that the primitive-based spatial vocabulary is not model-specific. The authors plan to release all artifacts upon publication. Future work may extend the approach to dynamic scenes and more complex spatial relations.

Who should read this

CS practitioners and researchers

Opening member content…