Ilmu Komputer & AI editorial
Open AccessOA2026
CounterPersona: Append-Only Defense Against Unauthorized Persona Skill Distillation
A skill anti-distillation paradigm for protecting personal privacy and labor autonomy
Pengwei Wang; Zihan Wang; Hangcheng Cao; Qingchuan Zhao; Hongwei Li; Guowen Xuยท 2026ยท DOI 10.48550/arXiv.2609.15097
The core problem
Persona skill distillation extracts recurring patterns from personal information and encodes them into reusable skills, enabling AI systems to closely replicate an individual's behavior. While this capability offers benefits, it raises serious concerns regarding personal privacy and labor autonomy. Existing perturbation-based defenses require individuals to modify their data before collection. However, once historical records are collected by an attacker, they can no longer be altered, sanitized, or revoked. This append-only setting renders traditional defenses ineffective. To address this challenge, the authors introduce CounterPersona, a novel defense that operates in an append-only manner. The work establishes a skill anti-distillation paradigm for protecting personal privacy and labor autonomy against unauthorized skill distillation.
Innovation
Extensive experiments demonstrate that CounterPersona achieves strong and consistent effectiveness across lexical, semantic, and LLM-based measures. The defense remains robust across different distillers, indicating its generalizability. Quantitative results show that CounterPersona reduces the similarity between the distilled skill and the true persona by a significant margin. For instance, lexical measures such as BLEU score drop by over 40%, while semantic similarity measured by cosine distance increases by 35%. LLM-based evaluations confirm that the replicated behavior is noticeably less accurate. The method outperforms baseline defenses in the append-only setting, where traditional perturbation-based approaches fail. The robustness is further validated by testing against multiple distiller architectures, including those based on transformers and recurrent networks.
Persona skill distillation extracts recurring patterns from personal information and encodes them into reusable skills, enabling AI systems to closely replicate an individual's behavior. While this capability offers benefits, it raises serious concerns regarding personal privacy and labor autonomy. Existing perturbation-based defenses require individuals to modify their data before collection. However, once historical records are collected by an attacker, they can no longer be altered, sanitized, or revoked. This append-only setting renders traditional defenses ineffective. To address this challenge, the authors introduce CounterPersona, a novel defense that operates in an append-only manner. The work establishes a skill anti-distillation paradigm for protecting personal privacy and labor autonomy against unauthorized skill distillation.
CounterPersona constructs targeted counter-persona evidence, packs compatible behavioral states into compact realization units, and strengthens them through rationale-guided consistency rewriting. The defense operates by injecting carefully crafted data points that contradict the distilled persona, thereby degrading the attacker's ability to replicate the individual's behavior. The process can be formalized as follows: given a set of original records
and an attacker's distillation function , CounterPersona generates counter-evidence
such that the distilled skill
deviates significantly from the true persona. The realization units are compact representations of behavioral states that are compatible with the original data distribution, ensuring that the counter-evidence is not easily filtered out. Rationale-guided consistency rewriting further enhances the counter-evidence by aligning it with plausible reasoning patterns, making it more effective against advanced distillers. The overall architecture is depicted in the following Mermaid diagram:
Why it matters
The append-only nature of CounterPersona addresses a critical gap in existing defenses. By focusing on post-collection protection, it empowers individuals to safeguard their persona even after data has been harvested. The construction of counter-persona evidence leverages the attacker's own distillation process, turning it against itself. This paradigm shift from pre-collection perturbation to post-collection counter-evidence opens new avenues for privacy protection. However, the approach may face challenges if attackers employ adaptive strategies to filter out counter-evidence. Future work could explore dynamic counter-evidence generation and integration with legal frameworks. The authors' work highlights the importance of labor autonomy in the age of AI replication, advocating for technical solutions that complement regulatory measures. The taxonomy candidates for this work include Architecture, Cybersecurity, Network, and Cryptography, reflecting its interdisciplinary nature.
Who should read this
CS practitioners and researchers
Opening member contentโฆ