Ilmu Komputer & AI editorial
Same Name, Different Server: A Security Census of Silent Drift in the Model Context Protocol Ecosystem
The core problem
The Model Context Protocol (MCP) has rapidly become the common interface through which large language model (LLM) applications reach external tools, data sources, and services. Its public registry now distributes thousands of community-built servers, but it lacks much of the vetting infrastructure that mature package ecosystems have accumulated over decades. This paper reports a security census of that ecosystem, aiming to quantify the prevalence of known vulnerability classes and, more importantly, to characterize a structural property that the protocol itself does not surface: silent drift between server versions.
The central research questions are: (1) What is the observed prevalence of high-severity security issues in the public MCP registry? (2) How often do servers change their advertised capabilities or endpoints between versions, and how often do such changes occur silently? (3) Is silent drift associated with a higher likelihood of security findings? (4) Does popularity, as measured by star counts, serve as a reliable proxy for safety? The study addresses these questions through a full-registry harvest, source-code scanning, and statistical analysis.
Innovation
Observed prevalence was dominated by unauthenticated network exposure, affecting 9.57% of scanned servers. After correcting each threat class by its measured precision, the observed high-severity prevalence of 11.14% reduced to approximately 7.6%.
The most striking finding concerns instability rather than any single weakness. Among multi-version servers, 51.1% changed what they advertise between versions. Of these, 40.6% did so silently, meaning the change was not surfaced to installed clients. Furthermore, 4.2% of multi-version servers redirected their remote endpoint to a different host while keeping their registry identity, a change the protocol never surfaces.
Silent drift was associated with nearly threefold higher odds of a high-severity finding: OR = 2.96, 95% CI [2.56, 3.42]. Popularity offered only weak protection: OR = 0.78 per unit of log stars, indicating that star counts are a poor proxy for safety. The following Mermaid diagram illustrates the key prevalence and association findings:
Why it matters
The census reveals that the MCP ecosystem's most pressing security challenge is not a single vulnerability class but the pervasive phenomenon of silent drift. More than half of multi-version servers alter their advertised behavior, and a substantial fraction do so without any notification to clients. This undermines trust assumptions: an installed client may continue to interact with a server that has changed its endpoint or capabilities, potentially exposing users to new risks. The strong association between silent drift and high-severity findings (OR = 2.96) suggests that drift is either a cause or a marker of security-relevant changes.
The finding that popularity (star counts) offers only weak protection (OR = 0.78 per log star) challenges the common heuristic of relying on community endorsements for safety. A server with many stars is not substantially safer than one with few. This has implications for registry design and client-side security policies.
The authors derive concrete recommendations for registry design, client-side pinning, and scanner triage. For registry design, they suggest surfacing version changes and endpoint modifications prominently, and possibly requiring publishers to declare changes. For client-side pinning, they recommend that clients pin to specific versions or cryptographic hashes of server manifests, and alert users when a server's advertised identity or endpoint changes. For scanner triage, they advise prioritizing servers that exhibit silent drift, as these are more likely to harbor high-severity issues.
Limitations include the snapshot nature of the study, potential biases in source-code fetching, and the precision correction relying on hand-labeled findings. Future work could extend the census longitudinally and incorporate dynamic analysis. An anonymized artifact is released to support reproducibility.
Who should read this
Opening member contentโฆ