Modern robotic systems are increasingly required to operate in complex, uncertain, and collaborative environments, where traditional control methods struggle to ensure adaptability and safety. Reinforcement Learning (RL) has emerged as a powerful framework to address these challenges, enabling robots to learn from interaction, optimize long-term performance, and adapt strategies on the fly. Despite its success across other domains such as IoT, games, healthcare, and autonomous driving, RL in robotics still faces critical obstacles. Some of these challenges are: the danger and cost of data collection, the reliance on simulators and sim-to-real transfer, the burden of manually defining task-specific reward functions, and the presence of estimation bias in learning algorithms. This thesis contributes to overcoming these limitations along four directions. First, we demonstrate how Model-Based RL can be effectively applied to real robotic systems by leveraging the data efficiency of algorithms such as MC-PILCO, thus avoiding the use of simulators and enabling application in both industrial robotics and athletic intelligence. Second, we introduce a novel reward learning module based on Gaussian Process Regression, which extends PILCO-like algorithms to tasks where explicit analytical cost functions are unavailable or impractical. Third, we propose a pipeline for automatic reward generation that exploits recent advances in Large Language Models to transform natural language task descriptions into executable reward functions, thereby reducing the engineering burden of task specification. Finally, we investigate the role of estimation bias in Actor-Critic algorithms and present a dual-armed bandit mechanism that exploits bias to improve exploration and learning performance.

Reinforcement Learning Applications in Robotic Systems: from Athletic Intelligence to Manipulation

TURCATO, NICCOLÒ
2026

Abstract

Modern robotic systems are increasingly required to operate in complex, uncertain, and collaborative environments, where traditional control methods struggle to ensure adaptability and safety. Reinforcement Learning (RL) has emerged as a powerful framework to address these challenges, enabling robots to learn from interaction, optimize long-term performance, and adapt strategies on the fly. Despite its success across other domains such as IoT, games, healthcare, and autonomous driving, RL in robotics still faces critical obstacles. Some of these challenges are: the danger and cost of data collection, the reliance on simulators and sim-to-real transfer, the burden of manually defining task-specific reward functions, and the presence of estimation bias in learning algorithms. This thesis contributes to overcoming these limitations along four directions. First, we demonstrate how Model-Based RL can be effectively applied to real robotic systems by leveraging the data efficiency of algorithms such as MC-PILCO, thus avoiding the use of simulators and enabling application in both industrial robotics and athletic intelligence. Second, we introduce a novel reward learning module based on Gaussian Process Regression, which extends PILCO-like algorithms to tasks where explicit analytical cost functions are unavailable or impractical. Third, we propose a pipeline for automatic reward generation that exploits recent advances in Large Language Models to transform natural language task descriptions into executable reward functions, thereby reducing the engineering burden of task specification. Finally, we investigate the role of estimation bias in Actor-Critic algorithms and present a dual-armed bandit mechanism that exploits bias to improve exploration and learning performance.
13-mar-2026
Inglese
CARLI, RUGGERO
Università degli studi di Padova
File in questo prodotto:
File Dimensione Formato  
PHD_Thesis_turcatonic_revised_pdfa.pdf

accesso aperto

Licenza: Tutti i diritti riservati
Dimensione 38.65 MB
Formato Adobe PDF
38.65 MB Adobe PDF Visualizza/Apri

I documenti in UNITESI sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/20.500.14242/375766
Il codice NBN di questa tesi è URN:NBN:IT:UNIPD-375766