Industry 4.0 has reshaped manufacturing through digitalization, pervasive sensing, and artificial intelligence, enabling unprecedented levels of automation, integration, and optimization. Industrial systems have become large-scale critical infrastructures whose failures can cause severe economic, environmental, and social consequences. Preventing and mitigating faults is therefore a matter of strategic importance. Predictive maintenance has emerged as a cornerstone of modern industry. By exploiting heterogeneous data from cyber-physical systems, it enables early fault detection, ensuring reliability, reduced downtime, and optimized resource use. Within the vision of Industry 5.0, predictive maintenance takes on even greater importance, aligned with the three guiding principles: efficiency (minimizing energy use, resource waste, and reaction time), robustness (ensuring stability under uncertainty and change), and human-in-the-loop (building trust and accountability by combining AI with human expertise). Machine learning plays a central role in this domain. General-purpose methods offer flexibility across industries but often lack diagnostic depth, while specialized solutions deliver higher accuracy and richer context but can be brittle in dynamic conditions. Industrial environments heighten these challenges due to scarce labeled fault data, shifting operating conditions, and regulatory requirements for human oversight. This makes interpretable, human-centered models essential. Hydroelectric power plants provide a representative and demanding case study, combining critical infrastructure with scarce failure data and high system complexity. We worked closely with domain experts from ANDRITZ HYDRO, a global leader in the field of hydroelectric power plants, to deliver with this thesis a set of solutions that are designed for real-world hydroelectric scenarios and validated in general industrial machine learning through proxy tasks. Each of our contribution improves the state-of-the-art in one or more of the following dimensions: efficiency, robustness, and human-in-the-loop. Key contributions include: (i) A hybrid anomaly detection and interpretability approach to support human-centric vibration monitoring and root cause analysis in hydroelectric power plants. (ii) Two continual learning methods, SmooER and SmooDER, which exploit data continuity to enable efficient introduction of new classes in behavior-based driver identification, requiring computational resources compatible with in-vehicle edge computing. (iii) RootIF, a feature-evolving anomaly detection method that supports integration of new features or sensors without historical data, minimizing system vulnerability windows and ensuring compliance with privacy and storage constraints. (iv) FLEX-C, a robust semi-supervised structured ensemble framework designed to perform fault detection and identification under conditions of data scarcity and labels contamination. Experimental results show that these methods consistently match or outperform state-of-the-art methods and baselines under realistic industrial constraints. They integrate interpretability mechanisms that support domain experts, while their computational efficiency enables practical deployment in edge computing environments.

Machine Learning Under Real-World Constraints: Toward Robust, Efficient, and Human-Centric Approaches

FANAN, MATTIA
2026

Abstract

Industry 4.0 has reshaped manufacturing through digitalization, pervasive sensing, and artificial intelligence, enabling unprecedented levels of automation, integration, and optimization. Industrial systems have become large-scale critical infrastructures whose failures can cause severe economic, environmental, and social consequences. Preventing and mitigating faults is therefore a matter of strategic importance. Predictive maintenance has emerged as a cornerstone of modern industry. By exploiting heterogeneous data from cyber-physical systems, it enables early fault detection, ensuring reliability, reduced downtime, and optimized resource use. Within the vision of Industry 5.0, predictive maintenance takes on even greater importance, aligned with the three guiding principles: efficiency (minimizing energy use, resource waste, and reaction time), robustness (ensuring stability under uncertainty and change), and human-in-the-loop (building trust and accountability by combining AI with human expertise). Machine learning plays a central role in this domain. General-purpose methods offer flexibility across industries but often lack diagnostic depth, while specialized solutions deliver higher accuracy and richer context but can be brittle in dynamic conditions. Industrial environments heighten these challenges due to scarce labeled fault data, shifting operating conditions, and regulatory requirements for human oversight. This makes interpretable, human-centered models essential. Hydroelectric power plants provide a representative and demanding case study, combining critical infrastructure with scarce failure data and high system complexity. We worked closely with domain experts from ANDRITZ HYDRO, a global leader in the field of hydroelectric power plants, to deliver with this thesis a set of solutions that are designed for real-world hydroelectric scenarios and validated in general industrial machine learning through proxy tasks. Each of our contribution improves the state-of-the-art in one or more of the following dimensions: efficiency, robustness, and human-in-the-loop. Key contributions include: (i) A hybrid anomaly detection and interpretability approach to support human-centric vibration monitoring and root cause analysis in hydroelectric power plants. (ii) Two continual learning methods, SmooER and SmooDER, which exploit data continuity to enable efficient introduction of new classes in behavior-based driver identification, requiring computational resources compatible with in-vehicle edge computing. (iii) RootIF, a feature-evolving anomaly detection method that supports integration of new features or sensors without historical data, minimizing system vulnerability windows and ensuring compliance with privacy and storage constraints. (iv) FLEX-C, a robust semi-supervised structured ensemble framework designed to perform fault detection and identification under conditions of data scarcity and labels contamination. Experimental results show that these methods consistently match or outperform state-of-the-art methods and baselines under realistic industrial constraints. They integrate interpretability mechanisms that support domain experts, while their computational efficiency enables practical deployment in edge computing environments.
13-mar-2026
Inglese
CARLI, RUGGERO
Università degli studi di Padova
File in questo prodotto:
File Dimensione Formato  
final_thesis_Mattia_Fanan.pdf

accesso aperto

Licenza: Tutti i diritti riservati
Dimensione 4.07 MB
Formato Adobe PDF
4.07 MB Adobe PDF Visualizza/Apri

I documenti in UNITESI sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/20.500.14242/375765
Il codice NBN di questa tesi è URN:NBN:IT:UNIPD-375765