The field of robotics has undergone significant changes in recent years. Advances in robotic hardware and sensing have improved the reliability of embodied platforms, enabling their deployment in increasingly unstructured and uncontrolled environments. At the same time, the emergence of Foundation Models has opened new opportunities for robotics by providing strong perceptual, linguistic, and reasoning capabilities that can be transferred to a wide range of downstream tasks. In this context, commonsense knowledge refers to the background knowledge that allows an embodied agent to interpret the world beyond immediate sensory input, including typical object properties, spatial and functional relations, action affordances, and temporal patterns. Although Foundation Models acquire many of these priors from internet-scale data, transferring them directly to robotics is not straightforward. Embodied agents must ground this knowledge in perception, use it to select feasible actions, and maintain it consistently over extended interactions with the world. These requirements call for commonsense representations that can be grounded in sensory observations, updated as the world evolves, and exploited to guide action during execution. In this thesis, we investigate how to bridge the gap between commonsense knowledge and the requirements of robotics, where intelligent agents must perceive, decide, and act in both static and dynamic environments. We address this problem through the notion of structured commonsense, namely commonsense knowledge represented in explicit forms that can be grounded in perception, linked to action, and maintained over time to support robotic decision making. We study this perspective along three main dimensions of robotic intelligence. First, we show how prior knowledge about everyday environments and objects can support the construction of world representations, and how such representations can be maintained as the environment changes. Second, we demonstrate how commonsense knowledge can be used to ground robot actions in context, enabling the execution of feasible and meaningful behaviors. Third, we investigate how structured commonsense can support temporal reasoning in changing environments, including the prediction of dynamic behaviors. Across these directions, the thesis validates the proposed perspective through a combination of simulated and real-world experiments. The results show that structured commonsense supports compact and updatable world models for dynamic environments, improves the grounding and feasibility of robot actions in open-vocabulary settings, and strengthens temporal reasoning in sequential decision-making problems. These contributions are demonstrated through quantitative improvements over neural and neurosymbolic baselines, as well as through real-robot evaluations that show the practical applicability of the proposed methods.
Structured commonsense for robotics: world representation, grounded action, and temporal reasoning
ARGENZIANO, FRANCESCO
2026
Abstract
The field of robotics has undergone significant changes in recent years. Advances in robotic hardware and sensing have improved the reliability of embodied platforms, enabling their deployment in increasingly unstructured and uncontrolled environments. At the same time, the emergence of Foundation Models has opened new opportunities for robotics by providing strong perceptual, linguistic, and reasoning capabilities that can be transferred to a wide range of downstream tasks. In this context, commonsense knowledge refers to the background knowledge that allows an embodied agent to interpret the world beyond immediate sensory input, including typical object properties, spatial and functional relations, action affordances, and temporal patterns. Although Foundation Models acquire many of these priors from internet-scale data, transferring them directly to robotics is not straightforward. Embodied agents must ground this knowledge in perception, use it to select feasible actions, and maintain it consistently over extended interactions with the world. These requirements call for commonsense representations that can be grounded in sensory observations, updated as the world evolves, and exploited to guide action during execution. In this thesis, we investigate how to bridge the gap between commonsense knowledge and the requirements of robotics, where intelligent agents must perceive, decide, and act in both static and dynamic environments. We address this problem through the notion of structured commonsense, namely commonsense knowledge represented in explicit forms that can be grounded in perception, linked to action, and maintained over time to support robotic decision making. We study this perspective along three main dimensions of robotic intelligence. First, we show how prior knowledge about everyday environments and objects can support the construction of world representations, and how such representations can be maintained as the environment changes. Second, we demonstrate how commonsense knowledge can be used to ground robot actions in context, enabling the execution of feasible and meaningful behaviors. Third, we investigate how structured commonsense can support temporal reasoning in changing environments, including the prediction of dynamic behaviors. Across these directions, the thesis validates the proposed perspective through a combination of simulated and real-world experiments. The results show that structured commonsense supports compact and updatable world models for dynamic environments, improves the grounding and feasibility of robot actions in open-vocabulary settings, and strengthens temporal reasoning in sequential decision-making problems. These contributions are demonstrated through quantitative improvements over neural and neurosymbolic baselines, as well as through real-robot evaluations that show the practical applicability of the proposed methods.| File | Dimensione | Formato | |
|---|---|---|---|
|
Tesi_dottorato_Argenziano.pdf
accesso aperto
Licenza:
Creative Commons
Dimensione
80.83 MB
Formato
Adobe PDF
|
80.83 MB | Adobe PDF | Visualizza/Apri |
I documenti in UNITESI sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/20.500.14242/380428
URN:NBN:IT:UNIROMA1-380428