arXiv:2609.31262v1 Announce Type: cross Abstract: Cross-client duplicate data in large language model training corpora degrades the efficiency of federated learning (FL) while exacerbating model memorization and privacy risks. Privacy-preserving cross-client deduplication effectively mitigates this issue by eliminating duplicate training data. However, existing schemes all follow a "Deduplication

Topological visualization of Deduplication-while-Training: A Resilient Paradigm for Privacy-Preserving Cross-Client Deduplication in Federated Learning
Brave API

Deduplication-while-Training (DwT) is a novel paradigm for Federated Learning (FL) that enables concurrent deduplication and model training, addressing the inefficiencies of traditional "Deduplication-before-Training" methods. Proposed in arXiv:2609.31262v1 by Wang et al., this approach transforms cross-client deduplication from a synchronous preprocessing step into a continuous online service, allowing clients to train on deduplicated data immediately after detection rather than waiting for global synchronization.

The proposed system, DwT-FL, utilizes a concurrent state-claim mechanism based on compare-and-swap (CAS) operations and a hot-cold dual-queue scheduling strategy. This architecture ensures that duplicate content is trained exactly once while preserving data privacy through cryptographic tagging. It significantly enhances resilience by supporting dynamic client joining and minimizing the impact of client disconnections through state recovery and task takeover.

Experimental evaluations demonstrate that DwT-FL outperforms state-of-the-art schemes by reducing failure recovery time by up to 93.04% and dynamic client joining overhead by up to 94.18%. By executing deduplication and training in parallel, the system maintains low metadata overhead and high scalability, making it suitable for unstable FL environments with large-scale language model training.

Generated 6d ago
Open-Weights Reasoning

This paper studies cross-client duplication in federated learning (FL) for large language model training, where the same or highly similar training examples may appear on multiple clients. The authors argue that such duplication is not merely a wasted-compute issue: it can degrade training efficiency while amplifying model memorization and privacy risks, because repeated exposure to the same data increases the likelihood that models retain identifiable or sensitive text. Since FL clients cannot simply share raw corpora for centralized deduplication, the paper focuses on privacy-preserving cross-client deduplication as a way to eliminate duplicate training data without exposing underlying client data.

The key contribution is a “deduplication-while-training” paradigm that departs from existing designs that treat deduplication as a one-time, offline preprocessing step. Rather than relying on a fragile upfront phase, the approach integrates duplicate handling into the iterative FL training process, making it more resilient to dynamic conditions such as changing client participation, evolving data availability, and ongoing model updates. The goal is to preserve the privacy guarantees of FL while preventing the same records from being counted multiple times across clients, thereby reducing redundant gradient contributions and improving the statistical efficiency of training.

This matters because real-world LLM training corpora are often highly redundant, and FL deployments that ignore cross-client overlap can suffer both efficiency losses and elevated memorization risk. By coupling deduplication with training rather than treating it as a separate pipeline stage, the work offers a more operationally realistic and privacy-conscious path for federated LLM training. It is especially relevant for settings where data is distributed across many clients, raw-data sharing is infeasible, and repeated text fragments can otherwise become a persistent source of overfitting and privacy leakage.

Generated 6d ago
Sources