Jadwal Sholat

Memuat jadwal sholatโ€ฆ

Computer Science editorial

Open AccessOA2026

WoE Wrote It? Watermarking Mixture-of-Experts LLMs for Black-Box Text Provenance

A parameter-intrinsic watermark for sparse MoE models that survives weight theft, distillation, and paraphrasing
Jona te Lintelo; Lichao Wu; Stjepan Picekยท 2026ยท DOI 10.48550/arXiv.2608.29151

The core problem

Large Language Model (LLM) watermarks are a primary mechanism for text provenance: they let a model owner identify machine-generated content and attribute it to a specific watermarked model. However, current LLM watermarking approaches predominantly rely on **inference-time sampler methods** and focus their analysis on **dense models**. This creates a structural blind spot. Inference-time methods are only effective when the text is explicitly generated via the model owner's controlled API; they fail in a **post-compromise scenario**. An adversary who steals or leaks the model weights gains complete control over inference and can simply run an unmodified sampler, bypassing the watermark and preventing post-theft attribution.

The authors (Jona te Lintelo, Lichao Wu, Stjepan Picek) therefore reframe the problem: provenance should not live in an enforceable inference wrapper, but in the model parameters themselves. They introduce **Watermarking of Experts (WoE)**, a black-box text provenance method that leverages the unique structural properties of sparse **Mixture-of-Experts (MoE)** models. WoE biases the vocabulary of specific experts and shifts the watermark signal embedding away f

Innovation

Across the eight evaluated MoE models, WoE achieves successful watermark detection from suspect text with an **average true positive rate of 90.1% at a 1% false positive rate**, reaching **up to 94.9%** on the strongest configuration. This places the detector well above the 1% FPR operating point that practical provenance systems require, while remaining a black-box procedure that needs only suspect text.

Crucially, the watermark is not merely detectable in the benign case. WoE remains detectable under three adversarial pressures:

- **Adversarial supervised fine-tuning** โ€” adapting the stolen model on attacker data does not erase the expert-level vocabulary bias.
- **Model extraction** โ€” secondary dense models distilled from the stolen MoE architecture retain a detectable signal.
- **Output-level paraphrasing** โ€” rewriting generated text does not fully remove the statistical signature.

At the same time, the authors report that WoE **largely preserves general model utility**, meaning the provenance guarantee does not come at the cost of a degraded base model. The result is a favorable trade-off curve for defenders: attribution survives the post-compromise scenario that defeats in

Large Language Model (LLM) watermarks are a primary mechanism for text provenance: they let a model owner identify machine-generated content and attribute it to a specific watermarked model. However, current LLM watermarking approaches predominantly rely on **inference-time sampler methods** and focus their analysis on **dense models**. This creates a structural blind spot. Inference-time methods are only effective when the text is explicitly generated via the model owner's controlled API; they fail in a **post-compromise scenario**. An adversary who steals or leaks the model weights gains complete control over inference and can simply run an unmodified sampler, bypassing the watermark and preventing post-theft attribution.
The authors (Jona te Lintelo, Lichao Wu, Stjepan Picek) therefore reframe the problem: provenance should not live in an enforceable inference wrapper, but in the model parameters themselves. They introduce **Watermarking of Experts (WoE)**, a black-box text provenance method that leverages the unique structural properties of sparse **Mixture-of-Experts (MoE)** models. WoE biases the vocabulary of specific experts and shifts the watermark signal embedding away from unenforceable inference wrappers. The goal is to attribute text generated by stolen weights, leaked checkpoints, and secondary dense models distilled from the stolen architecture โ€” without needing access to the adversary's deployment or weights.

Why it matters

The central contribution of WoE is a shift in *where* the watermark lives. Inference-time samplers assume the defender controls the generation path; WoE assumes the defender has already lost that control and must rely on the parameters themselves. By biasing expert vocabularies, the signal is entangled with the model's learned specialization, so removing it is not a matter of swapping a sampler.

This produces a deliberate **trade-off for malicious actors**: weakening the attribution signal requires additional model adaptation or text-rewriting operations, or it compromises the utility of the resulting output. In other words, the adversary must pay in compute, in pipeline complexity, or in output quality โ€” there is no free bypass.

The approach also extends provenance to **secondary dense models distilled from the stolen architecture**, a scenario that inference-time watermarking cannot address at all, since the distilled model never touches the owner's API. The reported 90.1% average TPR at 1% FPR, with a peak of 94.9%, indicates the signal is strong enough for practical attribution while the utility preservation keeps the watermarked MoE usable as a general-purpose model.

Limitations follow from the design: WoE is specific to sparse MoE architectures and depends on the persistence of expert routing structure through adaptation and distillation. The taxonomy of this work sits at the intersection of **Architecture**, **Cybersecurity**, **Network**, and **Cryptography** โ€” architecture because it exploits MoE routing, cybersecurity because it addresses model theft and post-compromise attribution, and cryptography because it functions as a provenance and attribution primitive for generated text.

Who should read this

CS practitioners and researchers

Opening member contentโ€ฆ