Modern robotic systems are increasingly required to operate in complex, uncertain, and collaborative environments, where traditional control methods struggle to ensure adaptability and safety. Reinforcement Learning (RL) has emerged as a powerful framework to address these challenges, enabling robots to learn from interaction, optimize long-term performance, and adapt strategies on the fly. Despite its success across other domains such as IoT, games, healthcare, and autonomous driving, RL in robotics still faces critical obstacles. Some of these challenges are: the danger and cost of data collection, the reliance on simulators and sim-to-real transfer, the burden of manually defining task-specific reward functions, and the presence of estimation bias in learning algorithms. This thesis contributes to overcoming these limitations along four directions. First, we demonstrate how Model-Based RL can be effectively applied to real robotic systems by leveraging the data efficiency of algorithms such as MC-PILCO, thus avoiding the use of simulators and enabling application in both industrial robotics and athletic intelligence. Second, we introduce a novel reward learning module based on Gaussian Process Regression, which extends PILCO-like algorithms to tasks where explicit analytical cost functions are unavailable or impractical. Third, we propose a pipeline for automatic reward generation that exploits recent advances in Large Language Models to transform natural language task descriptions into executable reward functions, thereby reducing the engineering burden of task specification. Finally, we investigate the role of estimation bias in Actor-Critic algorithms and present a dual-armed bandit mechanism that exploits bias to improve exploration and learning performance.
Reinforcement Learning Applications in Robotic Systems: from Athletic Intelligence to Manipulation
TURCATO, NICCOLÒ
2026
Abstract
Modern robotic systems are increasingly required to operate in complex, uncertain, and collaborative environments, where traditional control methods struggle to ensure adaptability and safety. Reinforcement Learning (RL) has emerged as a powerful framework to address these challenges, enabling robots to learn from interaction, optimize long-term performance, and adapt strategies on the fly. Despite its success across other domains such as IoT, games, healthcare, and autonomous driving, RL in robotics still faces critical obstacles. Some of these challenges are: the danger and cost of data collection, the reliance on simulators and sim-to-real transfer, the burden of manually defining task-specific reward functions, and the presence of estimation bias in learning algorithms. This thesis contributes to overcoming these limitations along four directions. First, we demonstrate how Model-Based RL can be effectively applied to real robotic systems by leveraging the data efficiency of algorithms such as MC-PILCO, thus avoiding the use of simulators and enabling application in both industrial robotics and athletic intelligence. Second, we introduce a novel reward learning module based on Gaussian Process Regression, which extends PILCO-like algorithms to tasks where explicit analytical cost functions are unavailable or impractical. Third, we propose a pipeline for automatic reward generation that exploits recent advances in Large Language Models to transform natural language task descriptions into executable reward functions, thereby reducing the engineering burden of task specification. Finally, we investigate the role of estimation bias in Actor-Critic algorithms and present a dual-armed bandit mechanism that exploits bias to improve exploration and learning performance.| File | Dimensione | Formato | |
|---|---|---|---|
|
PHD_Thesis_turcatonic_revised_pdfa.pdf
accesso aperto
Licenza:
Tutti i diritti riservati
Dimensione
38.65 MB
Formato
Adobe PDF
|
38.65 MB | Adobe PDF | Visualizza/Apri |
I documenti in UNITESI sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/20.500.14242/375766
URN:NBN:IT:UNIPD-375766