Jadwal Sholat

Memuat jadwal sholat…

Computer Science editorial

Open AccessOA2026

Modelstamp: Pre-Deserialization Verification of Machine-Learning Artifacts and Runtime Environment State

A lightweight Python persistence library for verifying artifact integrity and represented runtime-environment state before deserialization
Anagha Dhekne· 2026· DOI 10.48550/arXiv.2609.01781

The core problem

Persisted machine-learning models can remain byte-identical while the software environments in which they are loaded evolve, creating a verification problem that artifact integrity checks alone cannot expose. Traditional integrity checks—such as SHA-256 digests of model files—confirm that the bytes have not changed, but they say nothing about whether the runtime environment (e.g., library versions, Python interpreter, hardware) still matches the environment in which the model was trained or validated. This gap is critical: a model that was safe and performant under one set of dependencies may behave unpredictably or insecurely under another, even if the model file itself is unchanged. Modelstamp addresses this by introducing a pre-deserialization verification step that checks both the artifact and the represented runtime-environment state against recorded evidence before the model is loaded into memory. The work is positioned as a complementary control, not a replacement for dependency management, malicious-model detection, safe deserialization, or public publisher authentication.

Innovation

Modelstamp was evaluated using 14 controlled environment-drift scenarios, eight controlled trust-boundary scenarios, and an artifact-size scaling benchmark from 10 MiB to 1 GiB.

**Drift experiments:** The controlled drift experiments behaved as specified across relevant dependency changes, unchanged environments, and unrelated environmental changes, including broader noise controls. This indicates that Modelstamp correctly detects meaningful drift while ignoring irrelevant changes.

**Trust-boundary experiments:** These confirmed both intended detections and expected limitations, including shared-key forgery and replay. The system detects tampering when the secret key is not compromised, but it cannot prevent forgery if the shared key is leaked, nor replay attacks without additional nonce or timestamp mechanisms.

**Performance:** Median verification time increased from 0.032 s at 10 MiB to 3.334 s at 1 GiB, with measured throughput of approximately 307–312 MiB/s in the benchmark environment. The relationship between artifact size (in MiB) and median verification time (in seconds) is approximately linear:

This linear scaling suggests

Persisted machine-learning models can remain byte-identical while the software environments in which they are loaded evolve, creating a verification problem that artifact integrity checks alone cannot expose. Traditional integrity checks—such as SHA-256 digests of model files—confirm that the bytes have not changed, but they say nothing about whether the runtime environment (e.g., library versions, Python interpreter, hardware) still matches the environment in which the model was trained or validated. This gap is critical: a model that was safe and performant under one set of dependencies may behave unpredictably or insecurely under another, even if the model file itself is unchanged. Modelstamp addresses this by introducing a pre-deserialization verification step that checks both the artifact and the represented runtime-environment state against recorded evidence before the model is loaded into memory. The work is positioned as a complementary control, not a replacement for dependency management, malicious-model detection, safe deserialization, or public publisher authentication.
Modelstamp is a lightweight Python persistence library. At persistence time, it associates a serialized artifact with a sidecar JSON manifest containing:
- A SHA-256 digest of the artifact.
- Runtime metadata (e.g., Python version, platform).
- Installed versions from a bounded tracked-package set.
- A separately recorded model-relevant subset of package versions that participate in drift comparison.

Why it matters

The results characterize Modelstamp as a complementary pre-deserialization reference-state verification control rather than as a replacement for dependency-management systems, malicious-model detection, safe deserialization, or public publisher authentication. Its key strength is the ability to detect environment drift that would otherwise go unnoticed by artifact integrity checks alone. However, the trust-boundary experiments highlight important limitations: shared-key forgery and replay attacks are not prevented. This means Modelstamp is best suited for scenarios where the producer and verifier share a trusted secret and where replay is not a concern, or where additional protections (e.g., nonces, timestamps) are layered on top. The performance overhead is modest for typical model sizes (e.g., under 0.1 s for models up to ~30 MiB), but for very large models (1 GiB), verification takes over 3 seconds, which may be acceptable for many workflows but could be a bottleneck in latency-sensitive applications. The bounded tracked-package set keeps overhead low but also means that drift in untracked packages is not detected. Future work could extend the tracked set adaptively or integrate with package managers to automatically determine relevant dependencies. Overall, Modelstamp fills a gap in the ML supply chain by providing a lightweight, easy-to-integrate verification step that raises the bar for safe model loading.

Who should read this

CS practitioners and researchers

Opening member content…