As Artificial Intelligence (AI) shifts from centralized infrastructures to distributed environments, learning paradigms must address challenges related to data heterogeneity, privacy, limited supervision, and constrained computational resources. This thesis investigates Knowledge Distillation (KD) as a framework for decentralized learning, interpreting distributed AI systems through the perspective of artificial social learning. Although KD has demonstrated strong empirical performance across many applications, its behavior as a general collaborative learning mechanism remains understudied, particularly in heterogeneous, weakly supervised, and decentralized environments. In particular, questions remain regarding reliable knowledge transfer between the role of the distillation objective in learning stability, and the adaptation of KD to networked multi-client systems. This thesis addresses some of these challenges through theoretical formulation and experimental evaluation across representative edge and industrial scenarios. The thesis first investigates whether KD can remain effective in environments where edge devices observe highly heterogeneous data distributions and where access to ground-truth supervision is limited or unavailable. The objective is to determine whether the teacher–student paradigm can act as a primary learning mechanism rather than as a complementary training strategy. Experimental results demonstrate that distilled supervision from a knowledgeable teacher alone can support stable and competitive model performance under strong distribution shifts and severe supervision constraints, confirming the suitability of KD for realistic edge learning scenarios. Building on this, the thesis investigates knowledge exchange in fully decentralized KD settings, where each client relies exclusively on information received from peers. Learning effectiveness in this context depends on how neighborhood predictions are integrated. The study highlights that the choice of distillation loss, together with client connectivity, determines the efficiency of knowledge transfer in heterogeneous decentralized networks. By evaluating a broad family of information-theoretic measures beyond conventional Cross-Entropy and Kullback–Leibler divergence, the results identify distillation objectives that promote robust convergence and effective collaboration across distributed clients. The thesis further extends KD to regression tasks through Remaining Useful Life prediction in industrial systems, demonstrating effectiveness while reducing local data exposure during collaborative continuous-output prediction and maintaining communication efficiency. Overall, this thesis extends KD beyond its traditional role as a model compression technique toward a general framework for decentralized collaborative learning. The findings provide practical guidance for the design of distributed KD systems, highlighting the importance of selecting distillation objectives and adapting knowledge aggregation strategies to network connectivity.

Leveraging Knowledge Distillation in Decentralized Learning

MOLO, JOAQUIM MBASA
2026

Abstract

As Artificial Intelligence (AI) shifts from centralized infrastructures to distributed environments, learning paradigms must address challenges related to data heterogeneity, privacy, limited supervision, and constrained computational resources. This thesis investigates Knowledge Distillation (KD) as a framework for decentralized learning, interpreting distributed AI systems through the perspective of artificial social learning. Although KD has demonstrated strong empirical performance across many applications, its behavior as a general collaborative learning mechanism remains understudied, particularly in heterogeneous, weakly supervised, and decentralized environments. In particular, questions remain regarding reliable knowledge transfer between the role of the distillation objective in learning stability, and the adaptation of KD to networked multi-client systems. This thesis addresses some of these challenges through theoretical formulation and experimental evaluation across representative edge and industrial scenarios. The thesis first investigates whether KD can remain effective in environments where edge devices observe highly heterogeneous data distributions and where access to ground-truth supervision is limited or unavailable. The objective is to determine whether the teacher–student paradigm can act as a primary learning mechanism rather than as a complementary training strategy. Experimental results demonstrate that distilled supervision from a knowledgeable teacher alone can support stable and competitive model performance under strong distribution shifts and severe supervision constraints, confirming the suitability of KD for realistic edge learning scenarios. Building on this, the thesis investigates knowledge exchange in fully decentralized KD settings, where each client relies exclusively on information received from peers. Learning effectiveness in this context depends on how neighborhood predictions are integrated. The study highlights that the choice of distillation loss, together with client connectivity, determines the efficiency of knowledge transfer in heterogeneous decentralized networks. By evaluating a broad family of information-theoretic measures beyond conventional Cross-Entropy and Kullback–Leibler divergence, the results identify distillation objectives that promote robust convergence and effective collaboration across distributed clients. The thesis further extends KD to regression tasks through Remaining Useful Life prediction in industrial systems, demonstrating effectiveness while reducing local data exposure during collaborative continuous-output prediction and maintaining communication efficiency. Overall, this thesis extends KD beyond its traditional role as a model compression technique toward a general framework for decentralized collaborative learning. The findings provide practical guidance for the design of distributed KD systems, highlighting the importance of selecting distillation objectives and adapting knowledge aggregation strategies to network connectivity.
4-giu-2026
Inglese
Decentralized Learning
Federated Learning
Knowledge distillation
Carlini, Emanuele
Chessa, Stefano
Vadicamo, Lucia
File in questo prodotto:
File Dimensione Formato  
2026_Molo_Thesis_Dissertation_Final_Version.pdf

accesso aperto

Licenza: Creative Commons
Dimensione 6.85 MB
Formato Adobe PDF
6.85 MB Adobe PDF Visualizza/Apri

I documenti in UNITESI sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/20.500.14242/376889
Il codice NBN di questa tesi è URN:NBN:IT:UNIPI-376889