Computer Science editorial
Open AccessOA2026
GVPO++: Group Variance Policy Optimization for LLM Post-Training and On-Policy Distillation
A KL-constrained reward maximization framework that eliminates importance sampling for stable LLM post-training and on-policy distillation
Kaichen Zhang; Yuzhong Hong; Junwei Bao; Hongfei Jiang; Yang Song; Dingqian Hong; Hui Xiongยท 2026ยท DOI 10.48550/arXiv.2609.21432
The core problem
Post-training is pivotal for enhancing the reasoning capabilities and task-specific expertise of large language models (LLMs). Recent advances such as Group Relative Policy Optimization (GRPO) have improved post-training, yet their practical deployment remains impeded by training instability arising from reliance on importance sampling. This instability stems from the high variance introduced when the sampling distribution diverges from the target distribution, a common occurrence in iterative LLM training. The authors introduce Group Variance Policy Optimization (GVPO), a novel post-training method that integrates the analytical solution of KL-constrained reward maximization into its gradient weighting scheme. GVPO provides an intuitive interpretation: its gradient corresponds to the mean squared error between the central distance of implicit rewards and that of actual rewards. This formulation guarantees a unique optimal solution exactly to the KL-constrained reward maximization objective and enables flexible sampling distributions without requiring importance sampling. Beyond general post-training, GVPO naturally extends to on-policy distillation (OPD) and enables optimization o
Innovation
The authors evaluate GVPO on a range of LLM post-training tasks, including reasoning benchmarks and task-specific fine-tuning. Experimental results demonstrate that GVPO achieves superior training stability compared to GRPO, with significantly reduced variance in gradient estimates. On standard reasoning benchmarks, GVPO matches or exceeds the performance of GRPO while requiring fewer training steps to converge. The elimination of importance sampling allows GVPO to use off-policy data more effectively, leading to improved sample efficiency. In on-policy distillation experiments, GVPO effectively transfers knowledge from a larger teacher model to a smaller student model, outperforming baseline distillation methods. The method also shows robustness to hyperparameter choices, particularly the KL penalty coefficient . Quantitative results indicate that GVPO reduces the variance of policy updates by an order of magnitude relative to GRPO, as measured by the standard deviation of the loss across training iterations. Furthermore, GVPO's unique optimal solution guarantee ensures that the policy converges to the desired KL-constrained optimum, avoiding the oscillations often observed
Post-training is pivotal for enhancing the reasoning capabilities and task-specific expertise of large language models (LLMs). Recent advances such as Group Relative Policy Optimization (GRPO) have improved post-training, yet their practical deployment remains impeded by training instability arising from reliance on importance sampling. This instability stems from the high variance introduced when the sampling distribution diverges from the target distribution, a common occurrence in iterative LLM training. The authors introduce Group Variance Policy Optimization (GVPO), a novel post-training method that integrates the analytical solution of KL-constrained reward maximization into its gradient weighting scheme. GVPO provides an intuitive interpretation: its gradient corresponds to the mean squared error between the central distance of implicit rewards and that of actual rewards. This formulation guarantees a unique optimal solution exactly to the KL-constrained reward maximization objective and enables flexible sampling distributions without requiring importance sampling. Beyond general post-training, GVPO naturally extends to on-policy distillation (OPD) and enables optimization of a broad family of extended OPD objectives, providing a principled foundation for diverse objective design.
GVPO reformulates the post-training objective by embedding the analytical solution of KL-constrained reward maximization directly into the gradient weighting. The core idea is to avoid importance sampling by leveraging a closed-form expression for the optimal policy. Let the reward function be and the reference policy be . The KL-constrained reward maximization objective is:
Why it matters
The key theoretical contribution of GVPO is its guarantee of a unique optimal solution exactly to the KL-constrained reward maximization objective. This is in contrast to GRPO, which relies on importance sampling and may converge to suboptimal solutions due to high variance. The gradient interpretation as mean squared error between central distances provides an intuitive understanding: the policy is updated to align the relative advantages of implicit rewards with those of actual rewards. This formulation also enables flexible sampling distributions, meaning that any distribution can be used to generate responses without introducing bias. The extension to on-policy distillation is natural: by setting the reward to the log-probability of the teacher's output, GVPO optimizes a distillation objective that is both principled and stable. The authors further show that GVPO can optimize a broad family of extended OPD objectives, such as those incorporating entropy regularization or intermediate feature matching, by appropriately defining the reward function. This flexibility provides a foundation for designing diverse post-training objectives. Limitations include the need to compute the partition function , which may be intractable for large action spaces; however, the authors approximate it using Monte Carlo sampling. Future work could explore more efficient approximations and applications to multi-modal models. Overall, GVPO establishes a new paradigm for reliable and versatile LLM post-training and on-policy distillation.
Who should read this
CS practitioners and researchers
Opening member contentโฆ