Computer Science editorial
Modelstamp: Pre-Deserialization Verification of Machine-Learning Artifacts and Runtime Environment State
The core problem
Innovation
Modelstamp was evaluated using 14 controlled environment-drift scenarios, eight controlled trust-boundary scenarios, and an artifact-size scaling benchmark from 10 MiB to 1 GiB.
**Drift experiments:** The controlled drift experiments behaved as specified across relevant dependency changes, unchanged environments, and unrelated environmental changes, including broader noise controls. This indicates that Modelstamp correctly detects meaningful drift while ignoring irrelevant changes.
**Trust-boundary experiments:** These confirmed both intended detections and expected limitations, including shared-key forgery and replay. The system detects tampering when the secret key is not compromised, but it cannot prevent forgery if the shared key is leaked, nor replay attacks without additional nonce or timestamp mechanisms.
**Performance:** Median verification time increased from 0.032 s at 10 MiB to 3.334 s at 1 GiB, with measured throughput of approximately 307–312 MiB/s in the benchmark environment. The relationship between artifact size (in MiB) and median verification time (in seconds) is approximately linear:
This linear scaling suggests
- A SHA-256 digest of the artifact.
- Runtime metadata (e.g., Python version, platform).
- Installed versions from a bounded tracked-package set.
- A separately recorded model-relevant subset of package versions that participate in drift comparison.
Why it matters
Who should read this
Opening member content…