Robots operating alongside people must recognize actions and concepts under limited supervision, handle novelty reliably, and ultimately translate perception into appropriate behavior. This thesis is guided by the following question: ``How can a robot learn to interpret human actions and task-relevant concepts in a data-efficient and open-world setting, while preserving the reliability and semantic flexibility required for safe interaction and future embodied decision-making?'' These objectives are addressed through three connected research directions. A first direction focuses on action recognition under scarce supervision and unknown classes. To the best of our knowledge, it introduces for the first time the problem of few-shot open-set action recognition (FSOSAR), where a model is given only a small support set of labeled examples for the classes of the current recognition problem and must both classify known query actions and reject actions outside that support set. To tackle this setting, we first proposed Feature-Residual Discrimination (FR-Disc) and validated it in a simplified one-shot setting using 3D skeleton sequences [berti2022one]. FR-Disc operates by letting a few-shot backbone select the most likely support class and then using a lightweight discriminator on the resulting support--query feature residual to decide acceptance versus rejection. We later extended and evaluated FR-Disc more extensively in the RGB video domain by benchmarking against representative open-set strategies across standard action recognition datasets and backbones [berti2026frdisc]. A second direction investigates how action recognition and visual understanding can be expanded with semantic flexibility, so that the robot is not limited to a fixed vocabulary of predefined labels. This direction studies how large frozen pretrained vision-language backbones can be adapted to new actions, compositions, or concepts by learning a small set of prompt or token parameters while keeping the backbone fixed. In the video domain, the first contribution presents DualVClip [calabrese2024impact], a prompt-learning approach that extends a CLIP-based open-vocabulary video-text model to zero-shot multi-label recognition of verb-object actions, and analyzes how verb-object compositionality and the design of verb-object class splits affect generalization. The second contribution studies vision-language model personalization: expanding the model vocabulary with a new token that represents a novel concept learned from a few images, and evaluating how the learned concept composes with known context words in both discriminative (personalized image retrieval) and generative (personalized image generation) settings via the ConCon-Chi} benchmark [rosasco2024conconchi]. A third direction connects perception with action and analyzes how these perceptual components may support safe embodied interaction. The thesis reports an embodied human-robot interaction study on ergoCub, a humanoid robot, where the same interaction skills are implemented either with a task orchestrator that is manually designed using behavior trees, or learned end-to-end via behavior cloning. To this end, in the final phase of the PhD, we developed metaCub, a software system for teleoperation and data collection, and used it to record interaction demonstrations. These data is used to train an end-to-end behavior-cloning approach using vision-language-action (VLA) policy models, i.e., policies conditioned on visual inputs and language. The analysis highlights practical limitations encountered in this setting, including limited perceptual generalization across people and limited temporal modeling in current VLA policies, and motivates future work on integrating stronger action-recognition representations into end-to-end policies.
From Perception to Action: Data-Efficient and Open-Set Learning for Adaptive Human-Robot Interaction
BERTI, STEFANO
2026
Abstract
Robots operating alongside people must recognize actions and concepts under limited supervision, handle novelty reliably, and ultimately translate perception into appropriate behavior. This thesis is guided by the following question: ``How can a robot learn to interpret human actions and task-relevant concepts in a data-efficient and open-world setting, while preserving the reliability and semantic flexibility required for safe interaction and future embodied decision-making?'' These objectives are addressed through three connected research directions. A first direction focuses on action recognition under scarce supervision and unknown classes. To the best of our knowledge, it introduces for the first time the problem of few-shot open-set action recognition (FSOSAR), where a model is given only a small support set of labeled examples for the classes of the current recognition problem and must both classify known query actions and reject actions outside that support set. To tackle this setting, we first proposed Feature-Residual Discrimination (FR-Disc) and validated it in a simplified one-shot setting using 3D skeleton sequences [berti2022one]. FR-Disc operates by letting a few-shot backbone select the most likely support class and then using a lightweight discriminator on the resulting support--query feature residual to decide acceptance versus rejection. We later extended and evaluated FR-Disc more extensively in the RGB video domain by benchmarking against representative open-set strategies across standard action recognition datasets and backbones [berti2026frdisc]. A second direction investigates how action recognition and visual understanding can be expanded with semantic flexibility, so that the robot is not limited to a fixed vocabulary of predefined labels. This direction studies how large frozen pretrained vision-language backbones can be adapted to new actions, compositions, or concepts by learning a small set of prompt or token parameters while keeping the backbone fixed. In the video domain, the first contribution presents DualVClip [calabrese2024impact], a prompt-learning approach that extends a CLIP-based open-vocabulary video-text model to zero-shot multi-label recognition of verb-object actions, and analyzes how verb-object compositionality and the design of verb-object class splits affect generalization. The second contribution studies vision-language model personalization: expanding the model vocabulary with a new token that represents a novel concept learned from a few images, and evaluating how the learned concept composes with known context words in both discriminative (personalized image retrieval) and generative (personalized image generation) settings via the ConCon-Chi} benchmark [rosasco2024conconchi]. A third direction connects perception with action and analyzes how these perceptual components may support safe embodied interaction. The thesis reports an embodied human-robot interaction study on ergoCub, a humanoid robot, where the same interaction skills are implemented either with a task orchestrator that is manually designed using behavior trees, or learned end-to-end via behavior cloning. To this end, in the final phase of the PhD, we developed metaCub, a software system for teleoperation and data collection, and used it to record interaction demonstrations. These data is used to train an end-to-end behavior-cloning approach using vision-language-action (VLA) policy models, i.e., policies conditioned on visual inputs and language. The analysis highlights practical limitations encountered in this setting, including limited perceptual generalization across people and limited temporal modeling in current VLA policies, and motivates future work on integrating stronger action-recognition representations into end-to-end policies.| File | Dimensione | Formato | |
|---|---|---|---|
|
phdunige_4247311.pdf
accesso aperto
Licenza:
Tutti i diritti riservati
Dimensione
45.79 MB
Formato
Adobe PDF
|
45.79 MB | Adobe PDF | Visualizza/Apri |
I documenti in UNITESI sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/20.500.14242/378554
URN:NBN:IT:UNIGE-378554