Jadwal Sholat

Memuat jadwal sholatโ€ฆ

Ilmu Komputer & AI editorial

Open AccessOA2026

Beyond Training: A Feasibility Taxonomy for Inference-Time AI Governance

A structured digest of Samar Ansari's taxonomy of twenty inference-time governance mechanisms, their readiness, adversarial limits, and substitution with hardware-stage controls.
Samar Ansariยท 2026ยท DOI 10.48550/arXiv.2609.10105

The core problem

Contemporary compute governance is, in practice, training governance. The thresholds, reporting duties, and frontier-AI regimes now in force attach to training compute and treat the trained model as the regulatory unit. Ansari argues this picture is incomplete: capability increasingly migrates to the deployment stage through inference-time scaling, agentic scaffolding, and compression onto consumer hardware. The paper therefore asks a reframing question: which governance mechanisms remain available once the regulatory object shifts from the training run to the inference call?

To answer it, the paper develops a feasibility taxonomy of twenty inference-time mechanisms spanning monitoring, verification, and enforcement. Each mechanism is rated on a four-point readiness scale against a documented four-vendor evidence base. The taxonomy is then stress-tested against a two-dimensional adversary model (three capability tiers crossed with four adversary roles) and mapped to four governance scenarios: domestic regulation, bilateral or multilateral coordination, industry self-regulation, and compute-marketplace governance.

The central contribution is not a single instrument but a structure

Innovation

The headline feasibility result is that fifteen of the twenty mechanisms have commercial technical substrates in production today. This means the raw technical building blocks for inference-time governance largely exist in the market, rather than requiring speculative research.

However, readiness is not uniform. Governance-grade assurance and adversarial robustness vary substantially across the fifteen. A mechanism can have a production substrate while still lacking the auditability, tamper-resistance, or adversarial hardening needed for regulatory use. The four-point readiness scale captures this gap between commercial availability and governance adequacy.

The adversary analysis sharply narrows the claim. The observed readiness holds only against a cooperative deployer and a low-to-medium-capability user. Against a high-capability state-level deployer, no mechanism rates adequate. This is a categorical result: the ceiling is not a matter of incremental improvement in the current taxonomy but a structural limit under the tested adversary model.

Fine-tuning produces a second structural result. It removes the model-internal components of the enforcement cluster. Platform-external

Contemporary compute governance is, in practice, training governance. The thresholds, reporting duties, and frontier-AI regimes now in force attach to training compute and treat the trained model as the regulatory unit. Ansari argues this picture is incomplete: capability increasingly migrates to the deployment stage through inference-time scaling, agentic scaffolding, and compression onto consumer hardware. The paper therefore asks a reframing question: which governance mechanisms remain available once the regulatory object shifts from the training run to the inference call?
To answer it, the paper develops a feasibility taxonomy of twenty inference-time mechanisms spanning monitoring, verification, and enforcement. Each mechanism is rated on a four-point readiness scale against a documented four-vendor evidence base. The taxonomy is then stress-tested against a two-dimensional adversary model (three capability tiers crossed with four adversary roles) and mapped to four governance scenarios: domestic regulation, bilateral or multilateral coordination, industry self-regulation, and compute-marketplace governance.

Why it matters

The paper's core analytical move is to relocate the regulatory object. If governance attaches to training compute, then deployment-stage capability migration creates a gap: inference-time scaling, agentic scaffolding, and consumer-hardware compression move capability past the training checkpoint. Ansari's taxonomy is an attempt to populate the deployment stage with concrete, rateable levers.

The feasibility picture is best read as a maturity gradient rather than a binary. Fifteen mechanisms have commercial substrates, but governance-grade assurance and adversarial robustness vary substantially. This suggests a policy sequence: first harden the commercially available mechanisms for audit and assurance, then address the adversarial gaps.

The adversary analysis imposes a hard boundary. Readiness against a cooperative deployer and low-to-medium-capability user does not generalize to a high-capability state-level deployer, against whom no mechanism rates adequate. Fine-tuning further erodes the enforcement cluster by removing model-internal components, leaving platform-external controls as the residual enforcement surface. This asymmetry matters for enforcement design: controls that live outside the model are more durable under fine-tuning than controls that live inside it.

The substitution analysis reframes the relationship between inference-stage and hardware-stage governance. Rather than treating them as competing regimes, the conditional substitution principle specifies when they provide comparable coverage under stated conditions. The conditions are doing the real work: deployer cooperation, adversary capability, and platform control determine whether substitution holds.

The scenario mapping to domestic regulation, bilateral or multilateral coordination, industry self-regulation, and compute-marketplace governance shows the taxonomy is intended for institutional pluralism. Different scenarios will privilege different clusters: marketplace governance may lean on monitoring and verification, while domestic regulation may need enforcement levers that survive fine-tuning.

A Mermaid representation of the taxonomy and its stress-testing flow is given below.

The overall assessment is that inference-time governance is feasible but bounded. The mechanisms exist; their assurance and robustness are uneven; and their adequacy collapses against the strongest adversary and under fine-tuning for model-internal enforcement. The taxonomy's value is in making those bounds explicit and rateable.

Who should read this

CS practitioners and researchers

Opening member contentโ€ฆ