Computer Science editorial
Open AccessOA2026
CARVY-FL: Client Anticlustering for Robust Voting in Provably Secure Federated Learning
Estimating client distribution types from one-epoch updates and using anticlustering to boost certified accuracy and vote margins under class-disjoint non-IID data
Masaki Nakada; Honoka Anada; Tatsuya Kaneko; Hiroshi Nakamura; Shinya Takamaeda-Yamazaki; Hideki Takaseยท 2026ยท DOI 10.48550/arXiv.2608.28992
The core problem
Federated learning (FL) enables collaborative training without directly sharing raw data, but remains vulnerable to malicious clients. Voting-based FL improves robustness by partitioning clients into groups, training one model per group, and aggregating predictions by plurality voting. However, under class-disjoint non-IID data, distribution-oblivious grouping can yield highly variable certified accuracy (CA). The authors identify this variability as a key weakness: when clients are grouped without regard to their data distributions, some groups may be dominated by a single class or a narrow set of classes, which weakens the ensemble's ability to outvote adversarial predictions. The paper proposes CARVY-FL to address this gap by making grouping distribution-aware while preserving the formal CA guarantee of voting-based FL. The work is positioned against FLCert, a prior voting-based FL method, and evaluates robustness under BadNets with model replacement, a strong backdoor threat model.
Innovation
Experiments are conducted on MNIST and Fashion-MNIST under class-disjoint non-IID data. The authors compare CARVY-FL against FLCert, a prior voting-based FL method. The reported results show that CARVY-FL achieves higher certified accuracy (CA) than FLCert on both datasets. Under BadNets with model replacement, CARVY-FL improves the AUC of 100-ASR by 11.1% and 14.9%, respectively. The 100-ASR metric likely refers to the attack success rate when 100% of malicious clients attempt the backdoor, and AUC measures the area under the curve of attack success versus some defense threshold. The improvements indicate that anticlustering-based grouping makes the voting ensemble more robust to backdoor attacks by increasing vote margins. The paper does not report exact CA numbers in the abstract, but the qualitative claim is that CA is higher than FLCert across the tested settings. The results support the hypothesis that distribution-aware grouping reduces the variance of CA under class-disjoint non-IID data.
Federated learning (FL) enables collaborative training without directly sharing raw data, but remains vulnerable to malicious clients. Voting-based FL improves robustness by partitioning clients into groups, training one model per group, and aggregating predictions by plurality voting. However, under class-disjoint non-IID data, distribution-oblivious grouping can yield highly variable certified accuracy (CA). The authors identify this variability as a key weakness: when clients are grouped without regard to their data distributions, some groups may be dominated by a single class or a narrow set of classes, which weakens the ensemble's ability to outvote adversarial predictions. The paper proposes CARVY-FL to address this gap by making grouping distribution-aware while preserving the formal CA guarantee of voting-based FL. The work is positioned against FLCert, a prior voting-based FL method, and evaluates robustness under BadNets with model replacement, a strong backdoor threat model.
CARVY-FL operates in two main stages: (1) client distribution type estimation and (2) anticlustering-based grouping. In the first stage, each client performs one epoch of local training and sends its model update to the server. The server uses these one-epoch updates to estimate each client's distribution type, i.e., which classes are present and how they are distributed. This avoids the need for clients to disclose raw data or full class histograms. In the second stage, the server partitions clients into groups using anticlustering, an optimization framework that maximizes within-group diversity. The goal is to ensure that each group contains a mix of distribution types so that no group is dominated by a single class-disjoint shard. Under a fixed grouping, CARVY-FL retains the voting-based CA guarantee while increasing vote margins. Formally, for a given grouping and a test input , the ensemble prediction is the plurality vote over group models :
Why it matters
The key insight of CARVY-FL is that grouping clients without considering their data distributions can create groups that are highly correlated in their errors, which undermines the voting-based CA guarantee. By estimating distribution types from one-epoch updates and using anticlustering to maximize within-group diversity, CARVY-FL reduces error correlation and increases vote margins. This preserves the formal CA guarantee while improving empirical robustness. The approach is complementary to existing defenses such as robust aggregation and anomaly detection, and it operates at the group formation level rather than at the update aggregation level. Limitations include the reliance on one-epoch updates for distribution estimation, which may be noisy under heterogeneous compute or communication constraints, and the assumption that class-disjoint non-IID data is the primary threat model. Future work could extend the method to other non-IID settings, integrate it with secure aggregation, and evaluate on larger datasets and more diverse attack types. The taxonomy candidates for this work include Architecture, Cybersecurity, Network, and Cryptography, reflecting its intersection of distributed systems, security, and privacy-preserving learning.
Who should read this
CS practitioners and researchers
Opening member contentโฆ