Jadwal Sholat

Memuat jadwal sholatโ€ฆ

Ilmu Komputer & AI editorial

Open AccessOA2026

Verification-Time Dependency on a Disappearing Evaluator

An operational protocol for preserving decision-state evidence when the model that produced a consequential decision is no longer accessible in the same version and execution context
Ho Wa Ku; Jameel Ahmed Siddiquiยท 2026ยท DOI 10.48550/arXiv.2608.29912

The core problem

AI governance and assurance frameworks frequently rest on a quiet assumption: that a consequential, model-mediated decision can be reconstructed or tested after the fact. Ku and Siddiqui argue that this assumption may fail precisely when it matters most โ€” when the evaluator that produced the decision is no longer accessible in the same version and execution context. The paper situates this problem within Execution Governance (EG) 3.0 and introduces three verification-time constructs: **Decision-State Commitment**, **Independent Verifiability**, and **Counterfactual Auditability**. The central question is not whether a decision was authorized, but whether a separately trusted verifier can later substantiate what was bound at authorization time. The authors frame their contribution as an operational verification-time protocol plus an optional **Verification-Time Preservation Package (VTPP)**, specifying what evidence to bind, what a verifier can substantiate, how stability and paired counterfactual tests should be calibrated, and which semantic checks remain beyond JSON Schema validity. Critically, the protocol is downstream and non-authorizing: it does not alter the EG Core Formula,

Innovation

Independent reprocessing of released Study 2 artifacts reproduces two original within-family behavioural comparisons:

- **52.0% modal-decision reversal** for Llama 3.1 8B versus Llama 3.3 70B (26/50).
- **30.0% modal-decision reversal** for GPT-OSS 20B versus GPT-OSS 120B (15/50).

The corrected baseline establishes that these are **within-family comparisons**, not provider-established succession. Post-hoc re-pairing against Groq-designated migration paths yields **64.0%** and **38.0%** reversal, but these figures remain descriptive because the cross-family invocation parameters were asymmetric.

A 22-event retirement census independently recomputes to:

- Median: **16.45 months**
- Mean: **18.72 months**
- Range: **3.9โ€“40.3 months**
- **17/22** intervals below 24 months

The census also shows that evaluator availability can differ by service surface, meaning that the same nominal model may remain reachable through one interface while disappearing from another. This surface-dependence complicates any single-lifespan assumption in verification planning.

AI governance and assurance frameworks frequently rest on a quiet assumption: that a consequential, model-mediated decision can be reconstructed or tested after the fact. Ku and Siddiqui argue that this assumption may fail precisely when it matters most โ€” when the evaluator that produced the decision is no longer accessible in the same version and execution context. The paper situates this problem within Execution Governance (EG) 3.0 and introduces three verification-time constructs: **Decision-State Commitment**, **Independent Verifiability**, and **Counterfactual Auditability**. The central question is not whether a decision was authorized, but whether a separately trusted verifier can later substantiate what was bound at authorization time. The authors frame their contribution as an operational verification-time protocol plus an optional **Verification-Time Preservation Package (VTPP)**, specifying what evidence to bind, what a verifier can substantiate, how stability and paired counterfactual tests should be calibrated, and which semantic checks remain beyond JSON Schema validity. Critically, the protocol is downstream and non-authorizing: it does not alter the EG Core Formula, add a seventh live condition, or state jurisdiction-specific legal admissibility.
The study proceeds through three empirical strands. First, the authors independently reprocess released Study 2 artifacts to reproduce two original within-family behavioural comparisons. Second, they perform post-hoc re-pairing against Groq-designated migration paths to test whether provider-announced succession paths yield comparable reversal rates. Third, they conduct a 22-event retirement census to quantify how long evaluators remain available.

Why it matters

The joint contribution is an operational verification-time protocol and an optional Verification-Time Preservation Package (VTPP). The protocol specifies what evidence to bind at authorization time, what a separately trusted verifier can substantiate later, how stability and paired counterfactual tests should be calibrated, and which semantic checks remain beyond JSON Schema validity.

The results carry three implications. First, within-family reversal rates of 52.0% and 30.0% indicate that even closely related model versions can produce materially different modal decisions, so version identity alone is insufficient for reconstruction. Second, the descriptive cross-family rates of 64.0% and 38.0% caution against treating provider migration paths as validated succession; asymmetric invocation parameters undermine comparability. Third, the retirement census median of 16.45 months and the fact that 17/22 intervals fall below 24 months suggest that many evaluators vanish within a typical audit horizon.

The authors stress that the protocol is downstream and non-authorizing: it does not alter the EG Core Formula, add a seventh live condition, or state jurisdiction-specific legal admissibility. Semantic checks beyond JSON Schema validity remain necessary because structural validity does not guarantee decision-state fidelity.

Who should read this

CS practitioners and researchers

Opening member contentโ€ฆ