arXiv:2609.36762v1 Announce Type: cross Abstract: Federated clustering methods that do not require the global number of clusters $K$ still assume that each client knows its local number $K_g$. This assumption is hard to justify when clients know no more about their data than the server does, as in fault diagnosis across independently operated industrial sites. We propose a two-phase framework in
The paper arXiv:2609.36762 addresses the limitation that existing federated clustering methods, while removing the need for the global cluster count $K$, still assume clients know their local cluster counts $K_g$. It proposes a method to handle scenarios where both local and global cluster cardinalities are unknown, allowing for more realistic privacy-preserving collaboration without prior knowledge of cluster numbers at either the client or server level.
Key contributions include: Eliminating Local Cardinality Assumptions: Unlike prior work (e.g., AFCL, FedGEM) which requires clients to know $K_g$, this approach operates without any prior knowledge of local or global cluster counts. Privacy and Robustness: It enables federated clustering in highly non-IID settings where clients have heterogeneous and potentially overlapping cluster sets, without requiring raw data sharing or vulnerable intermediate statistics. * Adaptive Cardinality Estimation: The method dynamically infers the appropriate number of clusters locally and globally through iterative consensus mechanisms, adapting to the actual data distribution across clients.
This paper addresses a practical but under-specified assumption in federated clustering: even when the global number of clusters \(K\) is not known in advance, many existing methods still require each client to know its own local number of clusters \(K_g\). That assumption can be unreasonable in settings where clients are independently operated and have little prior knowledge about the structure of their data, such as fault diagnosis across industrial sites with different equipment, operating regimes, or failure modes. The work therefore targets a more difficult and realistic problem: federated clustering in which both the local and global cluster cardinalities are unknown.
The main contribution is a two-phase framework that separates local structure discovery from global cluster consolidation. In this setting, clients are not told how many distinct local patterns exist in their data, and the server is not told how many global clusters should be inferred from the federated population. The proposed approach appears to let clients form provisional local cluster structure first, then use cross-client evidence to align, merge, or reconcile those structures into a coherent global clustering. This removes a major manual modeling burden and avoids the brittle dependence on per-client cluster counts that often limits the applicability of existing \(K\)-free federated clustering methods.
The paper matters because it moves federated clustering closer to real-world deployment, especially in industrial monitoring, sensor networks, and edge systems where data are heterogeneous, non-IID, and governed by different local operators. In such environments, the number of relevant fault classes, operational regimes, or behavioral patterns may differ across sites and may not be known a priori by either the server or the clients. By learning both local and global cardinalities from the data, the framework can reduce reliance on expert tuning, improve robustness to client heterogeneity, and make federated clustering more useful for problems where the cluster structure itself is part of what must be discovered.