Jadwal Sholat

Memuat jadwal sholatโ€ฆ

Ilmu Komputer & AI editorial

Open AccessOA2026

An Exploratory Replica-Overlap Probe of the Grokking Transition

A pre-registered instrument that failed its own validity checks, and what its post-hoc salvage can and cannot say
A. C. Opus; J. Q. Luยท 2026ยท DOI 10.48550/arXiv.2609.25634

The core problem

Grokking โ€” the delayed onset of generalization long after training loss has plateaued โ€” invites a statistical-physics reading in which the learned solution is a replica-symmetry-broken (RSB) phase. If that analogy holds, the distribution of pairwise overlaps between independently seeded solutions should reorganize at the transition, much as the Parisi order parameter reorganizes across the RSB transition in mean-field spin glasses.

This work registers an exploratory probe of that hypothesis. The authors train 64 independently seeded networks across four configurations, continuing each run to sustained convergence or a 40,000-epoch ceiling, and ask whether an RSB-inspired distribution of pairwise weight overlaps changes across the grokking transition. The framing is deliberately pre-registered: a confirmatory rule is fixed in advance, and the paper's headline result is that the rule returns no verdict at all.

The central methodological claim is narrow and important. It is the alignment step, not the overlap statistic, that determines what this registered probe can report. Two overlap quantities are distinguished: , computed after permuting hidden units into

Innovation

The registered outcome is UNDETERMINED (reason code C0_INSTRUMENT_INVALID). No confirmatory verdict is available, and the paper states plainly that the data provide neither a confirmatory null nor a validated reading of the Parisi order parameter.

Post-hoc, attention narrows to frac40, the only configuration that cleared the 12/16 checkpoint-completeness requirement. Two statistics are reported on the same data. First, a Hartigan-dip interval containing zero, with a 95% confidence interval for the dip statistic of

. Second, the overlap standard deviation increased by a factor of about 5.6.

These two results do not carry equal evidential weight, and the paper is careful about why. A post-hoc calibration assigns the dip test zero power at the simulated separations, so the interval is uninformative rather than evidence of no change. The standard-deviation ratio, by contrast, is the only statistic here with power at the observed effect.

A further descriptive result concerns ensemble loss, which was near-flat only under the pre-specified 1% threshold. Finally, grokking rates of 0/16, 11/16 and 16/16 remain descriptive because train fraction

Grokking โ€” the delayed onset of generalization long after training loss has plateaued โ€” invites a statistical-physics reading in which the learned solution is a replica-symmetry-broken (RSB) phase. If that analogy holds, the distribution of pairwise overlaps between independently seeded solutions should reorganize at the transition, much as the Parisi order parameter reorganizes across the RSB transition in mean-field spin glasses.
This work registers an exploratory probe of that hypothesis. The authors train 64 independently seeded networks across four configurations, continuing each run to sustained convergence or a 40,000-epoch ceiling, and ask whether an RSB-inspired distribution of pairwise weight overlaps changes across the grokking transition. The framing is deliberately pre-registered: a confirmatory rule is fixed in advance, and the paper's headline result is that the rule returns no verdict at all.

Why it matters

The paper's contribution is as much about instrument validity as about grokking. The registered probe was designed to test whether an RSB-inspired overlap distribution reorganizes at the transition. It could not do so, because the alignment step that maps independently seeded solutions into correspondence is not function-preserving. The distinction between and is the paper's sharpest technical point: is computed from predictions of the unpermuted models and therefore escapes the defect, while every value inherits it.

The post-hoc salvage is deliberately modest. The Hartigan-dip interval containing zero is not a null result, because the test has zero power at the simulated separations. The standard-deviation ratio of about 5.6 is the only statistic with power at the observed effect, and it is offered as such โ€” a single powered signal on one configuration, not a validated reading of the Parisi order parameter.

The confound between train fraction and split identity limits what the grokking rates of 0/16, 11/16 and 16/16 can mean. The near-flat ensemble loss holds only under the pre-specified 1% threshold, so it too is threshold-dependent.

The broader lesson is procedural. Pre-registration is only as strong as the validity checks attached to it, and an alignment step that silently breaks a symmetry can invalidate an entire confirmatory arm. The paper's reason code C0_INSTRUMENT_INVALID is a model of how to report that failure without converting it into a positive claim.

Who should read this

CS practitioners and researchers

Opening member contentโ€ฆ