Ilmu Komputer & AI editorial
Verification-Time Dependency on a Disappearing Evaluator
The core problem
Innovation
Independent reprocessing of released Study 2 artifacts reproduces two original within-family behavioural comparisons:
- **52.0% modal-decision reversal** for Llama 3.1 8B versus Llama 3.3 70B (26/50).
- **30.0% modal-decision reversal** for GPT-OSS 20B versus GPT-OSS 120B (15/50).
The corrected baseline establishes that these are **within-family comparisons**, not provider-established succession. Post-hoc re-pairing against Groq-designated migration paths yields **64.0%** and **38.0%** reversal, but these figures remain descriptive because the cross-family invocation parameters were asymmetric.
A 22-event retirement census independently recomputes to:
- Median: **16.45 months**
- Mean: **18.72 months**
- Range: **3.9โ40.3 months**
- **17/22** intervals below 24 months
The census also shows that evaluator availability can differ by service surface, meaning that the same nominal model may remain reachable through one interface while disappearing from another. This surface-dependence complicates any single-lifespan assumption in verification planning.
Why it matters
The joint contribution is an operational verification-time protocol and an optional Verification-Time Preservation Package (VTPP). The protocol specifies what evidence to bind at authorization time, what a separately trusted verifier can substantiate later, how stability and paired counterfactual tests should be calibrated, and which semantic checks remain beyond JSON Schema validity.
The results carry three implications. First, within-family reversal rates of 52.0% and 30.0% indicate that even closely related model versions can produce materially different modal decisions, so version identity alone is insufficient for reconstruction. Second, the descriptive cross-family rates of 64.0% and 38.0% caution against treating provider migration paths as validated succession; asymmetric invocation parameters undermine comparability. Third, the retirement census median of 16.45 months and the fact that 17/22 intervals fall below 24 months suggest that many evaluators vanish within a typical audit horizon.
The authors stress that the protocol is downstream and non-authorizing: it does not alter the EG Core Formula, add a seventh live condition, or state jurisdiction-specific legal admissibility. Semantic checks beyond JSON Schema validity remain necessary because structural validity does not guarantee decision-state fidelity.
Who should read this
Opening member contentโฆ