Trustworthiness has become a central objective for the research and medical communities as learning systems move from controlled development settings to real-world clinical workflows. Although trustworthiness spans methodological, operational, and societal dimensions, this thesis adopts a focused technical scope and targets properties that materially affect clinical utility and can be formalised, optimised,and tested. Concretely, it investigates: (i) clinically realistic multimodal learning for integrating heterogeneous evidence, (ii) human-interpretable explanations to expose model drivers at both cohort and patient level, (iii) robustness and supervision-efficient adaptation under deployment-level distribution shifts, (iv) label-free test-time reliability estimation to anticipate failure when ground truth is unavailable, and (v) scalable self-supervised representation learning, culminating in medical foundation backbones. The thesis develops these themes across four application domains. It first presents an explainable multimodal framework for early COVID-19 prognosis from chest X-rays and triage variables, showing that controlled fusion and multi-level explanations can align predictive performance with clinical transparency. It then studies robustness under shift in multimodal 3D segmentation of multiple sclerosis lesions, proposing adaptation strategies that minimise target-domain supervision and introducing a mathematically grounded reliability estimator that flags likely test-time failures without labels. Next, it demonstrates large-scale self-supervised ECG representation learning through HuBERT-ECG, establishing transferable cardiac representations across heterogeneous datasets and tasks. Finally, it extends this paradigm to arbitrary lead configurations by introducing a lightweight (~7Mparameters) any-lead ECG foundation model based on sparse, physiology-aligned spatiotemporal graphs, masked node modelling, and stochastic lead sampling to learnlead-invariant representations transferable to reduced-lead and wearable settings.Taken together, this thesis contributes a set of methods that supports a pragmatic pathtoward trustworthy foundation models in medicine. Since the prerequisites for fully general multimodal foundations in medicine remain hard to satisfy today, progress is driven by strong and carefully validated backbones, robustness paired with reliability to anticipate deployment drift, and interpretable evidence integration that clinicians caninspect and challenge. Our results show that deliberate alignment between model inductive bias, clinical prior, transparency by design and deployment conditions is essential for medical AI systems that are not only accurate, but dependable in practice.
Il passaggio dei sistemi di apprendimento da ambienti di sviluppo controllati a contesti clinici reali ha reso l’affidabilità un obiettivo centrale per le comunità di ricerca e mediche. Sebbene questo tema abbracci dimensioni metodologiche, operative e sociali, questa tesi adotta un perimetro tecnico mirato e si concentra su proprietà che influenzano in modo concreto l’utilità clinica e che possono essere formalizzate, ottimizzate e verificate sperimentalmente. In particolare, la tesi affronta: (i) l’apprendimento multimodale clinicamente realistico per integrare evidenze eterogenee, (ii) spiegazioni interpretabili dall’uomo per chiarire i fattori che guidano il modello sia a livello di popolazione sia di singolo paziente, (iii) robustezza e adattamento efficiente con supervisione limitata in presenza di cambiamenti nelle distribuzioni dei dati durante l’uso, (iv) stima dell’affidabilità delle predizioni in assenza di etichette o riscontri diagnostici al momento del test, e (v) apprendimento scalabile e auto-supervisionato di rappresentazioni, fino alla costruzione di modelli fondazionali in ambito medico. Questi temi vengono sviluppati in quattro domini applicativi. Il primo contributo consiste in un framework multimodale interpretabile per la stima della prognosi nei pazienti COVID-19 a partire da radiografie del torace e dati di triage, mostrando che una fusione controllata delle modalità e spiegazioni su più livelli possono conciliare prestazioni predittive e trasparenza clinica. La tesi studia poi la robustezza di modelli multimodali per la segmentazione 3D di lesioni da sclerosi multipla quando si verificano cambiamenti nella distribuzione dei dati, proponendo strategie di adattamento che richiedono una supervisione minima e introducendo uno stimatore di affidabilità capace di segnalare probabili errori durante l’utilizzo. Successivamente, viene mostrato come HuBERT-ECG, un modello che sfrutta l’apprendimento auto-supervisionato su larga scala, apprenda rappresentazioni ECG trasferibili su dataset e compiti clinici eterogenei. Infine, questo paradigma viene esteso a configurazioni arbitrarie di derivazioni introducendo un modello fondazionale ECG leggero (~7M parametri), capace di elaborare combinazioni variabili di derivazioni tramite grafi spaziotemporali sparsi e fisiologicamente coerenti, insieme a masked node modelling e campionamento stocastico delle derivazioni per ottenere rappresentazioni invarianti rispetto al numero e al tipo di derivazioni, trasferibili anche a contesti a poche derivazioni e a dispositivi indossabili. Nel complesso, la tesi propone un insieme di metodi che sostiene un percorso pragmatico verso modelli fondazionali affidabili in medicina. Poiché oggi i requisiti per modelli multimodali pienamente generali in ambito clinico sono ancora difficili da soddisfare, il progresso passa attraverso modelli di base solidi e accuratamente validati, robustezza accompagnata da indicatori di affidabilità capaci di anticipare possibili derive durante l’impiego, e integrazione delle evidenze in forme che i clinici possano ispezionare e mettere criticamente alla prova. I risultati indicano che un allineamento consapevole tra bias induttivo del modello, conoscenza clinica, trasparenza progettuale e condizioni d’uso è essenziale per sistemi di intelligenza artificiale medica non solo accurati, ma realmente affidabili nella pratica.
From Explainable and Robust Design to Scalable Foundation Models: A Technical Journey towards Trustworthy Medical AI
COPPOLA, Edoardo
2026
Abstract
Trustworthiness has become a central objective for the research and medical communities as learning systems move from controlled development settings to real-world clinical workflows. Although trustworthiness spans methodological, operational, and societal dimensions, this thesis adopts a focused technical scope and targets properties that materially affect clinical utility and can be formalised, optimised,and tested. Concretely, it investigates: (i) clinically realistic multimodal learning for integrating heterogeneous evidence, (ii) human-interpretable explanations to expose model drivers at both cohort and patient level, (iii) robustness and supervision-efficient adaptation under deployment-level distribution shifts, (iv) label-free test-time reliability estimation to anticipate failure when ground truth is unavailable, and (v) scalable self-supervised representation learning, culminating in medical foundation backbones. The thesis develops these themes across four application domains. It first presents an explainable multimodal framework for early COVID-19 prognosis from chest X-rays and triage variables, showing that controlled fusion and multi-level explanations can align predictive performance with clinical transparency. It then studies robustness under shift in multimodal 3D segmentation of multiple sclerosis lesions, proposing adaptation strategies that minimise target-domain supervision and introducing a mathematically grounded reliability estimator that flags likely test-time failures without labels. Next, it demonstrates large-scale self-supervised ECG representation learning through HuBERT-ECG, establishing transferable cardiac representations across heterogeneous datasets and tasks. Finally, it extends this paradigm to arbitrary lead configurations by introducing a lightweight (~7Mparameters) any-lead ECG foundation model based on sparse, physiology-aligned spatiotemporal graphs, masked node modelling, and stochastic lead sampling to learnlead-invariant representations transferable to reduced-lead and wearable settings.Taken together, this thesis contributes a set of methods that supports a pragmatic pathtoward trustworthy foundation models in medicine. Since the prerequisites for fully general multimodal foundations in medicine remain hard to satisfy today, progress is driven by strong and carefully validated backbones, robustness paired with reliability to anticipate deployment drift, and interpretable evidence integration that clinicians caninspect and challenge. Our results show that deliberate alignment between model inductive bias, clinical prior, transparency by design and deployment conditions is essential for medical AI systems that are not only accurate, but dependable in practice.| File | Dimensione | Formato | |
|---|---|---|---|
|
PhD Thesis Edoardo Coppola.pdf
accesso aperto
Licenza:
Tutti i diritti riservati
Dimensione
6.72 MB
Formato
Adobe PDF
|
6.72 MB | Adobe PDF | Visualizza/Apri |
I documenti in UNITESI sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/20.500.14242/379830
URN:NBN:IT:UNIBS-379830