Ilmu Komputer & AI editorial
Open AccessOA2026
Distributed and Private Textual Data Synthesis from Embeddings
A DP-cryptography co-design for training-free, non-iterative text synthesis without a trusted curator
Ergute Bao; Hongyan Chang; Ali Shahin Shamsabadi; Ting Yu; Xiaokui Xiao· 2026· DOI 10.48550/arXiv.2609.10104
The core problem
Differentially private (DP) text synthesis has traditionally relied on a trusted, centralized curator with access to raw user texts. However, in realistic distributed settings, privacy concerns preclude such a curator, and existing DP text synthesis pipelines cannot be deployed due to unrealistic trust and access assumptions. Naive adaptations require repeated, tightly synchronized user participation and incur significant overhead. This work addresses the gap by proposing a DP–cryptography co-design for textual data synthesis that requires no trusted curator and only lightweight user participation. The approach consists of two optimized components: a distributed-friendly DP synthesis algorithm and a custom secure protocol that enforces end-to-end DP guarantees over distributed user data. The goal is to achieve utility comparable to state-of-the-art centralized DP synthesis methods while operating in a fully distributed, privacy-preserving manner.
Innovation
The authors evaluate their approach on four benchmarks, comparing it to state-of-the-art centralized DP synthesis methods. The results demonstrate that the proposed distributed method achieves utility comparable to the centralized baseline. Specifically, the utility is measured in terms of the quality of synthetic texts, likely using metrics such as BLEU, ROUGE, or downstream task performance. The paper reports that the distributed approach incurs only lightweight user participation and no trusted curator, while maintaining competitive utility. The exact numerical results are not provided in the abstract, but the claim of comparable utility is a key finding. The evaluation likely includes varying privacy budgets () and dataset sizes, showing that the method scales well and provides strong privacy guarantees.
Differentially private (DP) text synthesis has traditionally relied on a trusted, centralized curator with access to raw user texts. However, in realistic distributed settings, privacy concerns preclude such a curator, and existing DP text synthesis pipelines cannot be deployed due to unrealistic trust and access assumptions. Naive adaptations require repeated, tightly synchronized user participation and incur significant overhead. This work addresses the gap by proposing a DP–cryptography co-design for textual data synthesis that requires no trusted curator and only lightweight user participation. The approach consists of two optimized components: a distributed-friendly DP synthesis algorithm and a custom secure protocol that enforces end-to-end DP guarantees over distributed user data. The goal is to achieve utility comparable to state-of-the-art centralized DP synthesis methods while operating in a fully distributed, privacy-preserving manner.
The proposed methodology comprises two main components:
Why it matters
The paper addresses a critical gap in DP text synthesis by enabling distributed settings without a trusted curator. The two-component design—a distributed-friendly DP synthesis algorithm and a custom secure protocol—effectively balances privacy, utility, and efficiency. The use of a one-time DP summary in embedding space allows for training-free, non-iterative synthesis, which reduces computational overhead and user burden. Semantic support protection is a novel contribution that mitigates the risk of exposing rare user data, an important consideration in privacy-preserving data synthesis. The secure protocol leverages cryptography to enforce DP guarantees, ensuring that even the server does not access raw data. However, the approach may have limitations, such as the need for a secure MPC setup and potential communication overhead during the summary computation. The authors claim utility comparable to centralized methods, but the trade-offs between privacy budget, utility, and efficiency are likely explored. Future work could extend the method to other modalities or improve the efficiency of the secure protocol. Overall, this work represents a significant step towards practical, privacy-preserving text synthesis in distributed environments.
Who should read this
CS practitioners and researchers
Opening member content…