Industry 4.0 has reshaped manufacturing through digitalization, pervasive sensing, and artificial intelligence, enabling unprecedented levels of automation, integration, and optimization. Industrial systems have become large-scale critical infrastructures whose failures can cause severe economic, environmental, and social consequences. Preventing and mitigating faults is therefore a matter of strategic importance. Predictive maintenance has emerged as a cornerstone of modern industry. By exploiting heterogeneous data from cyber-physical systems, it enables early fault detection, ensuring reliability, reduced downtime, and optimized resource use. Within the vision of Industry 5.0, predictive maintenance takes on even greater importance, aligned with the three guiding principles: efficiency (minimizing energy use, resource waste, and reaction time), robustness (ensuring stability under uncertainty and change), and human-in-the-loop (building trust and accountability by combining AI with human expertise). Machine learning plays a central role in this domain. General-purpose methods offer flexibility across industries but often lack diagnostic depth, while specialized solutions deliver higher accuracy and richer context but can be brittle in dynamic conditions. Industrial environments heighten these challenges due to scarce labeled fault data, shifting operating conditions, and regulatory requirements for human oversight. This makes interpretable, human-centered models essential. Hydroelectric power plants provide a representative and demanding case study, combining critical infrastructure with scarce failure data and high system complexity. We worked closely with domain experts from ANDRITZ HYDRO, a global leader in the field of hydroelectric power plants, to deliver with this thesis a set of solutions that are designed for real-world hydroelectric scenarios and validated in general industrial machine learning through proxy tasks. Each of our contribution improves the state-of-the-art in one or more of the following dimensions: efficiency, robustness, and human-in-the-loop. Key contributions include: (i) A hybrid anomaly detection and interpretability approach to support human-centric vibration monitoring and root cause analysis in hydroelectric power plants. (ii) Two continual learning methods, SmooER and SmooDER, which exploit data continuity to enable efficient introduction of new classes in behavior-based driver identification, requiring computational resources compatible with in-vehicle edge computing. (iii) RootIF, a feature-evolving anomaly detection method that supports integration of new features or sensors without historical data, minimizing system vulnerability windows and ensuring compliance with privacy and storage constraints. (iv) FLEX-C, a robust semi-supervised structured ensemble framework designed to perform fault detection and identification under conditions of data scarcity and labels contamination. Experimental results show that these methods consistently match or outperform state-of-the-art methods and baselines under realistic industrial constraints. They integrate interpretability mechanisms that support domain experts, while their computational efficiency enables practical deployment in edge computing environments.
Machine Learning Under Real-World Constraints: Toward Robust, Efficient, and Human-Centric Approaches
FANAN, MATTIA
2026
Abstract
Industry 4.0 has reshaped manufacturing through digitalization, pervasive sensing, and artificial intelligence, enabling unprecedented levels of automation, integration, and optimization. Industrial systems have become large-scale critical infrastructures whose failures can cause severe economic, environmental, and social consequences. Preventing and mitigating faults is therefore a matter of strategic importance. Predictive maintenance has emerged as a cornerstone of modern industry. By exploiting heterogeneous data from cyber-physical systems, it enables early fault detection, ensuring reliability, reduced downtime, and optimized resource use. Within the vision of Industry 5.0, predictive maintenance takes on even greater importance, aligned with the three guiding principles: efficiency (minimizing energy use, resource waste, and reaction time), robustness (ensuring stability under uncertainty and change), and human-in-the-loop (building trust and accountability by combining AI with human expertise). Machine learning plays a central role in this domain. General-purpose methods offer flexibility across industries but often lack diagnostic depth, while specialized solutions deliver higher accuracy and richer context but can be brittle in dynamic conditions. Industrial environments heighten these challenges due to scarce labeled fault data, shifting operating conditions, and regulatory requirements for human oversight. This makes interpretable, human-centered models essential. Hydroelectric power plants provide a representative and demanding case study, combining critical infrastructure with scarce failure data and high system complexity. We worked closely with domain experts from ANDRITZ HYDRO, a global leader in the field of hydroelectric power plants, to deliver with this thesis a set of solutions that are designed for real-world hydroelectric scenarios and validated in general industrial machine learning through proxy tasks. Each of our contribution improves the state-of-the-art in one or more of the following dimensions: efficiency, robustness, and human-in-the-loop. Key contributions include: (i) A hybrid anomaly detection and interpretability approach to support human-centric vibration monitoring and root cause analysis in hydroelectric power plants. (ii) Two continual learning methods, SmooER and SmooDER, which exploit data continuity to enable efficient introduction of new classes in behavior-based driver identification, requiring computational resources compatible with in-vehicle edge computing. (iii) RootIF, a feature-evolving anomaly detection method that supports integration of new features or sensors without historical data, minimizing system vulnerability windows and ensuring compliance with privacy and storage constraints. (iv) FLEX-C, a robust semi-supervised structured ensemble framework designed to perform fault detection and identification under conditions of data scarcity and labels contamination. Experimental results show that these methods consistently match or outperform state-of-the-art methods and baselines under realistic industrial constraints. They integrate interpretability mechanisms that support domain experts, while their computational efficiency enables practical deployment in edge computing environments.| File | Dimensione | Formato | |
|---|---|---|---|
|
final_thesis_Mattia_Fanan.pdf
accesso aperto
Licenza:
Tutti i diritti riservati
Dimensione
4.07 MB
Formato
Adobe PDF
|
4.07 MB | Adobe PDF | Visualizza/Apri |
I documenti in UNITESI sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/20.500.14242/375765
URN:NBN:IT:UNIPD-375765