As Artificial Intelligence (AI) shifts from centralized infrastructures to distributed environments, learning paradigms must address challenges related to data heterogeneity, privacy, limited supervision, and constrained computational resources. This thesis investigates Knowledge Distillation (KD) as a framework for decentralized learning, interpreting distributed AI systems through the perspective of artificial social learning. Although KD has demonstrated strong empirical performance across many applications, its behavior as a general collaborative learning mechanism remains understudied, particularly in heterogeneous, weakly supervised, and decentralized environments. In particular, questions remain regarding reliable knowledge transfer between the role of the distillation objective in learning stability, and the adaptation of KD to networked multi-client systems. This thesis addresses some of these challenges through theoretical formulation and experimental evaluation across representative edge and industrial scenarios. The thesis first investigates whether KD can remain effective in environments where edge devices observe highly heterogeneous data distributions and where access to ground-truth supervision is limited or unavailable. The objective is to determine whether the teacher–student paradigm can act as a primary learning mechanism rather than as a complementary training strategy. Experimental results demonstrate that distilled supervision from a knowledgeable teacher alone can support stable and competitive model performance under strong distribution shifts and severe supervision constraints, confirming the suitability of KD for realistic edge learning scenarios. Building on this, the thesis investigates knowledge exchange in fully decentralized KD settings, where each client relies exclusively on information received from peers. Learning effectiveness in this context depends on how neighborhood predictions are integrated. The study highlights that the choice of distillation loss, together with client connectivity, determines the efficiency of knowledge transfer in heterogeneous decentralized networks. By evaluating a broad family of information-theoretic measures beyond conventional Cross-Entropy and Kullback–Leibler divergence, the results identify distillation objectives that promote robust convergence and effective collaboration across distributed clients. The thesis further extends KD to regression tasks through Remaining Useful Life prediction in industrial systems, demonstrating effectiveness while reducing local data exposure during collaborative continuous-output prediction and maintaining communication efficiency. Overall, this thesis extends KD beyond its traditional role as a model compression technique toward a general framework for decentralized collaborative learning. The findings provide practical guidance for the design of distributed KD systems, highlighting the importance of selecting distillation objectives and adapting knowledge aggregation strategies to network connectivity.
Leveraging Knowledge Distillation in Decentralized Learning
MOLO, JOAQUIM MBASA
2026
Abstract
As Artificial Intelligence (AI) shifts from centralized infrastructures to distributed environments, learning paradigms must address challenges related to data heterogeneity, privacy, limited supervision, and constrained computational resources. This thesis investigates Knowledge Distillation (KD) as a framework for decentralized learning, interpreting distributed AI systems through the perspective of artificial social learning. Although KD has demonstrated strong empirical performance across many applications, its behavior as a general collaborative learning mechanism remains understudied, particularly in heterogeneous, weakly supervised, and decentralized environments. In particular, questions remain regarding reliable knowledge transfer between the role of the distillation objective in learning stability, and the adaptation of KD to networked multi-client systems. This thesis addresses some of these challenges through theoretical formulation and experimental evaluation across representative edge and industrial scenarios. The thesis first investigates whether KD can remain effective in environments where edge devices observe highly heterogeneous data distributions and where access to ground-truth supervision is limited or unavailable. The objective is to determine whether the teacher–student paradigm can act as a primary learning mechanism rather than as a complementary training strategy. Experimental results demonstrate that distilled supervision from a knowledgeable teacher alone can support stable and competitive model performance under strong distribution shifts and severe supervision constraints, confirming the suitability of KD for realistic edge learning scenarios. Building on this, the thesis investigates knowledge exchange in fully decentralized KD settings, where each client relies exclusively on information received from peers. Learning effectiveness in this context depends on how neighborhood predictions are integrated. The study highlights that the choice of distillation loss, together with client connectivity, determines the efficiency of knowledge transfer in heterogeneous decentralized networks. By evaluating a broad family of information-theoretic measures beyond conventional Cross-Entropy and Kullback–Leibler divergence, the results identify distillation objectives that promote robust convergence and effective collaboration across distributed clients. The thesis further extends KD to regression tasks through Remaining Useful Life prediction in industrial systems, demonstrating effectiveness while reducing local data exposure during collaborative continuous-output prediction and maintaining communication efficiency. Overall, this thesis extends KD beyond its traditional role as a model compression technique toward a general framework for decentralized collaborative learning. The findings provide practical guidance for the design of distributed KD systems, highlighting the importance of selecting distillation objectives and adapting knowledge aggregation strategies to network connectivity.| File | Dimensione | Formato | |
|---|---|---|---|
|
2026_Molo_Thesis_Dissertation_Final_Version.pdf
accesso aperto
Licenza:
Creative Commons
Dimensione
6.85 MB
Formato
Adobe PDF
|
6.85 MB | Adobe PDF | Visualizza/Apri |
I documenti in UNITESI sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/20.500.14242/376889
URN:NBN:IT:UNIPI-376889