Ilmu Komputer & AI editorial
Video-HopChain: Multi-Hop Questions and Confidence-Gated Exploration for Video Reasoning Models
The core problem
HopChain demonstrated on still images that multi-hop data synthesis improves vision-language reasoning: long chain-of-thought reasoning exposes errors that compound across steps, while most data used for reinforcement learning with verifiable rewards (RLVR) rarely demands a chain of visual evidence, so these weaknesses are likely to stay unexposed. The authors observe the same problem in video, where this framework has not yet been explored.
The paper therefore introduces **Video-HopChain**, a dataset of **22,550 multi-hop video questions over 13,378 videos**, together with a held-out benchmark of **1,000 questions**. Each question chains **three to six yes/no questions** about moments in one video, and each sub-question yields one of two integers depending on its answer. The final answer is the sum of these integers, so an exact match on that sum provides the verifiable reward that RLVR needs.
The core research questions are: (1) does multi-hop video data synthesis improve video reasoning models under RLVR, and (2) can the known gradient-collapse limitation of GRPO be mitigated at the same compute budget?
Innovation
The evaluation reports the mean over **eight video understanding and reasoning benchmarks**.
| Stage | Training data | Mean score |
|---|---|---|
| Base | โ | 55.4 |
| Stage 1 | Standard video dataset (GRPO) | 55.4 โ baseline |
| Stage 2 | Video-HopChain (GRPO) | **57.9** |
| Stage 2 + CGE | Video-HopChain (GRPO + CGE) | **59.3** |
Key findings:
- A second stage on Video-HopChain raises the mean from **55.4 to 57.9** and improves **every one** of the eight benchmarks.
- Adding **Confidence-Gated Exploration** raises the mean further to **59.3**, recovering groups that standard GRPO leaves without gradient at the same compute budget.
- The dataset comprises **22,550 questions** over **13,378 videos**, with a held-out benchmark of **1,000 questions**.
- Each question chains **3โ6** yes/no sub-questions, yielding a verifiable integer-sum reward.
Why it matters
The results support the hypothesis that multi-hop data synthesis transfers from images to video. Because each question requires a chain of visual evidence, errors compound across sub-questions, exposing weaknesses that single-hop RLVR data leaves hidden. The exact-match reward on the integer sum is simple, automatic, and verifiable, which makes the dataset directly compatible with RLVR pipelines.
The CGE contribution addresses a structural limitation of GRPO rather than a data limitation. When all 8 rollouts share the same outcome, the group advantage collapses to zero and no gradient flows. By masking the policy's most confident token inside the reasoning span for the second half of rollouts, CGE injects controlled diversity while excluding the masked positions from the loss, so all 8 rollouts still contribute to the advantage. This recovers hard and easy groups without increasing the rollout budget.
Limitations and open questions include the reliance on yes/no sub-questions (which constrains the answer space), the cost of generating 13,378 videos' worth of chained annotations, and whether the gains transfer to larger video reasoning models or to open-ended video QA. The authors release the dataset, the checkpoint, and the data generation and training code to support replication and extension.
Who should read this
Opening member contentโฆ