Jadwal Sholat

Memuat jadwal sholat…

Computer Science editorial

Open AccessOA2026

Corrective Forcing: Unified Post-Training for Diffusions and Flows in Generative Speech Enhancement

A post-training paradigm that closes the training–inference gap by learning from self-generated rollouts and correcting clean-speech predictions under dynamic sampling schedules.
Qing Yao; Lijian Gao; Qirong Mao· 2026· DOI 10.48550/arXiv.2609.24651

The core problem

Diffusion and flow models have emerged as promising generative paradigms for speech enhancement. However, they suffer from a fundamental training–inference mismatch: during training, models are optimized on analytical path states, whereas at inference they recursively evaluate themselves on self-generated rollout states along discretized sampling trajectories. This discrepancy causes prediction and discretization errors to accumulate over the sampling process, degrading output quality.

The paper introduces **Corrective Forcing (CoF)**, a post-training paradigm that addresses this mismatch by forcing diffusion and flow models to learn from their own self-generated rollouts and correct their predictions. CoF is designed to be applicable across both diffusion and flow formulations through a shared clean-speech prediction parameterization, making it a unified post-training objective. The authors evaluate CoF with two representative instantiations: **SB-VE** (a variance-exploding diffusion model) and **OT-CFM** (optimal-transport conditional flow matching), demonstrating improvements in perceptual quality and reconstruction fidelity, along with robust performance across different numbe

Innovation

The authors conduct experiments with two instantiations: **SB-VE** and **OT-CFM**. The results demonstrate improvements in perceptual quality and reconstruction fidelity. Specifically, CoF yields better performance compared to baseline models without post-training, as measured by standard speech enhancement metrics (though the abstract does not list specific metric values).

A key finding is that CoF provides robust performance across different numbers of sampling steps. This is particularly important because diffusion and flow models often degrade when the number of sampling steps is reduced for faster inference. CoF's ability to maintain quality across step counts suggests that it effectively mitigates the accumulation of discretization errors.

The experiments validate the unified nature of CoF: the same post-training objective improves both diffusion (SB-VE) and flow (OT-CFM) models, confirming that the shared clean-speech prediction parameterization enables cross-formulation applicability.

Diffusion and flow models have emerged as promising generative paradigms for speech enhancement. However, they suffer from a fundamental training–inference mismatch: during training, models are optimized on analytical path states, whereas at inference they recursively evaluate themselves on self-generated rollout states along discretized sampling trajectories. This discrepancy causes prediction and discretization errors to accumulate over the sampling process, degrading output quality.
The paper introduces **Corrective Forcing (CoF)**, a post-training paradigm that addresses this mismatch by forcing diffusion and flow models to learn from their own self-generated rollouts and correct their predictions. CoF is designed to be applicable across both diffusion and flow formulations through a shared clean-speech prediction parameterization, making it a unified post-training objective. The authors evaluate CoF with two representative instantiations: **SB-VE** (a variance-exploding diffusion model) and **OT-CFM** (optimal-transport conditional flow matching), demonstrating improvements in perceptual quality and reconstruction fidelity, along with robust performance across different numbers of sampling steps.

Why it matters

The training–inference mismatch in diffusion and flow models is a well-known challenge. CoF addresses it by directly training on self-generated rollouts, which is conceptually similar to scheduled sampling or reinforcement learning from model outputs, but tailored to the continuous-time generative setting. The use of dynamic sampling schedules ensures that the model is exposed to a variety of inference conditions, preventing overfitting to a fixed discretization.

The local evolution regularization via counterfactual transitions is a novel contribution. It encourages the model's step-wise updates to be consistent with corrected predictions, which likely stabilizes the sampling trajectory and reduces error accumulation. By expressing outputs through a shared clean-speech prediction parameterization, CoF achieves unification across diffusion and flow formulations, simplifying its adoption.

Limitations and future work are not explicitly discussed in the abstract, but potential directions include extending CoF to other generative tasks (e.g., image or audio generation) and investigating theoretical guarantees for error reduction. The taxonomy candidates (Architecture, Cybersecurity, Network, Cryptography) appear unrelated to this work; the primary domain is generative speech enhancement.

Who should read this

CS practitioners and researchers

Opening member content…