Jadwal Sholat

Memuat jadwal sholatโ€ฆ

Computer Science editorial

Open AccessOA2026

iSDFT: Information-Proximal Self-Distillation for Continual Learning in LLMs

A budgeted teacher-information constraint yields closed-form exponential targets and improves specialisation-retention trade-offs across four LLM backbones.
Ahmed Khaled Khamis; Xiaotong Ji; Hassan Jaber; Rasul Tutunov; Matthieu Zimmer; Jun Wang; Haitham Bou-Ammarยท 2026ยท DOI 10.48550/arXiv.2609.24646

The core problem

Continual learning in large language models (LLMs) faces a persistent tension: acquiring new skills from demonstrations while avoiding catastrophic forgetting of broad pre-trained capabilities. On-policy self-distillation fine-tuning (SDFT) addresses this by learning from demonstrations while reducing forgetting, but it always distils toward the full demonstration-conditioned teacher. This fixes teacher influence at the full-teacher endpoint, providing no control over how much demonstration information should be transferred at each prediction state. The authors introduce Information-Proximal SDFT (iSDFT), which treats the teacher as a budgeted source of information rather than an unconditional target. The central research question is whether controlling the amount and timing of teacher information can improve the specialisation-retention trade-off. The work is positioned against the original SDFT benchmark suite and ten additional mathematics, coding, and competition-mathematics benchmarks, evaluated across four heterogeneous LLM backbones and two specialisation tasks.

Innovation

The authors evaluate iSDFT across four heterogeneous LLM backbones and two specialisation tasks. iSDFT improves vanilla SDFT in 7 of 8 model-task settings and matches it in the remaining one. On the original SDFT benchmark suite, iSDFT provides tighter retention: 73% of evaluations remain within 0.5 points of the base model, compared with 52% for the strongest baseline. On ten additional mathematics, coding, and competition-mathematics benchmarks, iSDFT achieves the largest mean improvement among the compared methods. These results indicate that controlling how much and when teacher information is introduced improves specialisation while preserving broader capability. The consistent gains across diverse backbones and tasks suggest that the information-proximal formulation is not tied to a particular architecture or domain.
Continual learning in large language models (LLMs) faces a persistent tension: acquiring new skills from demonstrations while avoiding catastrophic forgetting of broad pre-trained capabilities. On-policy self-distillation fine-tuning (SDFT) addresses this by learning from demonstrations while reducing forgetting, but it always distils toward the full demonstration-conditioned teacher. This fixes teacher influence at the full-teacher endpoint, providing no control over how much demonstration information should be transferred at each prediction state. The authors introduce Information-Proximal SDFT (iSDFT), which treats the teacher as a budgeted source of information rather than an unconditional target. The central research question is whether controlling the amount and timing of teacher information can improve the specialisation-retention trade-off. The work is positioned against the original SDFT benchmark suite and ten additional mathematics, coding, and competition-mathematics benchmarks, evaluated across four heterogeneous LLM backbones and two specialisation tasks.
iSDFT operates at the token level. At each token, it selects the distribution closest to the current student that satisfies a prescribed teacher-information constraint. Formally, let denote the student distribution and the teacher distribution at a given prediction state. The constraint is expressed as a bound on the Kullback-Leibler divergence from the student to the teacher:

Why it matters

The key insight of iSDFT is that teacher influence should be treated as a controllable resource rather than a fixed endpoint. By selecting the student-closest distribution that satisfies a teacher-information constraint, the method avoids over-committing to the full teacher at every token, which can erode pre-trained capabilities. The closed-form exponential target with a locally determined tilt provides a principled and computationally tractable mechanism for this control. The anchoring to the frozen base policy further stabilises training by limiting cumulative drift. The empirical results support the hypothesis that finer control over information transfer yields a better specialisation-retention trade-off. Limitations include the need to tune the information budget and the anchoring strength, and the evaluation is confined to two specialisation tasks and four backbones. Future work could explore adaptive budgets, alternative divergence measures, and broader task suites. Overall, iSDFT offers a simple yet effective modification to SDFT that improves both specialisation and retention across a range of benchmarks.

Who should read this

CS practitioners and researchers

Opening member contentโ€ฆ