Plinio Moreno

dblp:15/2466 · DBLP profile ↗
← Back
22ranked-venue papers
2as first author
8since 2021 · last 2026
0000-0002-0496-2050ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 22 · 2 first-author · 8 since 2021Systems, architecture and hardware · 13 · 6 since 2021Theory of computation · 2
YearPublicationVenuePosition
2026 SemBA-FAST: Semantic-based Bayesian attention applied to foveal active visual search tasks
abstract
Both robots and humans have visual sensors with limited fields of view that need to be controlled to explore the environment and search for objects. To make this process efficient, visual attention methods actively select the information that contributes the most to the success of the task. Two key factors are characteristic of human vision. First, sensors can have space-varying resolution to process only certain parts of the scene with high resolution. Second, the attentional focus is deployed at highly informative regions, e.g. highly conspicuous regions. In this paper, we propose the use of semantic information, readily available in state-of-the-art deep object detectors, as an effective method to guide visual target search tasks using foveal sensors, which we refer to as SemBA -FAST. Because state-of-the-art object detectors are trained in conventional Cartesian images, we propose methods to calibrate detections in foveated images without requiring retraining the deep models. The information collected across multiple saccades is fused using Bayesian filters that keep a semantic representation of the world with associated uncertainty, on which the next gaze direction is actively determined. The proposed model is compared with state-of-the-art saliency-based methods. Our results demonstrate that semantic information positively influences the performance of target-present visual search in static scenes, highlighting its importance in designing visual attention systems for robots. • Semantic information available on current deep learning models can be exploited in active visual search of known objects and brings advantages with respect to saliency-based models. • Foveal vision effectively reduces the amount of visual information to be processed. • Deep-learning pre-trained object detectors can be calibrated to foveal images with low computational effort. • Biologically inspired computational models provide better insights into human visual cognition. • Probabilistic framework for integrating information across multiple views and best next view planning that enhances interpretability and mathematical explainability.
João Luzio, Alexandre Bernardino, Plinio Moreno
Neurocomputing3
2025 GRASPLAT: Enabling dexterous grasping through novel view synthesis
abstract
Achieving dexterous robotic grasping with multi-fingered hands remains a significant challenge. While existing methods rely on complete 3D scans to predict grasp poses, these approaches face limitations due to the difficulty of acquiring high-quality 3D data in real-world scenarios. In this paper, we introduce GRASPLAT, a novel grasping framework that leverages consistent 3D information while being trained solely on RGB images. Our key insight is that by synthesizing physically plausible images of a hand grasping an object, we can regress the corresponding hand joints for a successful grasp. To achieve this, we utilize 3D Gaussian Splatting to generate high-fidelity novel views of real hand-object interactions, enabling end-to-end training with RGB data. Unlike prior methods, our approach incorporates a photometric loss that refines grasp predictions by minimizing discrepancies between rendered and real images. We conduct extensive experiments on both synthetic and real-world grasping datasets, demonstrating that GRASPLAT improves grasp success rates up to 36.9% over existing image-based methods. Project page: https://mbortolon97.github.io/grasplat/
Matteo Bortolon, Nuno Ferreira Duarte, Plinio Moreno, Fabio Poiesi, José Santos-Victor, Alessio Del Bue
IROS3
2025 Measuring Uncertainty in Shape Completion to Improve Grasp Quality
abstract
Shape completion networks have been used recently in real-world robotic experiments to complete the missing/hidden information in environments where objects are only observed in one or few instances where self-occlusions are bound to occur. Nowadays, most approaches rely on deep neural networks that handle rich 3D point cloud data that lead to more precise and realistic object geometries. However, these models still suffer from inaccuracies due to its nondeterministic/stochastic inferences which could lead to poor performance in grasping scenarios where these errors compound to unsuccessful grasps. We present an approach to calculate the uncertainty of a 3D shape completion model during inference of single view point clouds of an object on a table top. In addition, we propose an update to grasp pose algorithms quality score by introducing the uncertainty of the completed point cloud present in the grasp candidates. To test our full pipeline we perform real world grasping with a 7dof robotic arm with a 2 finger gripper on a large set of household objects and compare against previous approaches that do not measure uncertainty. Our approach ranks the grasp quality better, leading to higher grasp success rate for the rank 5 grasp candidates compared to state of the art. Project Page: https://nunoduarte.github.io/pages.3dsgrasp++
Nuno Ferreira Duarte, Seyed Saber Mohammadi, Plinio Moreno, Alessio Del Bue, José Santos-Victor
IROS3
2024 Non-Verbal Cues on Robot-Group Persuasion
abstract
When integrating robots into human daily life, persuasive power can be essential. However, there are often group dynamics which can complicate persuasion. This study focuses on how non-verbal cues, specifically gaze and hand gestures, affect the persuasiveness of a social robot. We have designed a protocol to include non-verbal cues in the social robot Vizzy (head and eye gaze, hand gestures) and test them in a series of experiments using the paradigm of the "Desert Survival Challenge". The goal of the robot is to persuade the participants of the game into changing their answers whilst avoiding negative feelings. It is hypothesized that the nonverbal cues will help avoid psychological reactance without diminishing compliance to the verbal requests issued by the robot. This phenomenon has been verified before for single person persuasion, but it is yet to be tested on groups. Thus, the goal of this project is to verify the effect of non-verbal cues in group persuasion by a robot and comparing it to single person persuasion. The results showed that the robot’s gestures increased compliance by the group and the gaze behaviour decreased psychological reactance.
Alexandra Gonçalves, Plinio Moreno, Jodi Forlizzi, Leonel Garcia-Marques, Alexandre Bernardino
ICRA2
2023 3DSGrasp: 3D Shape-Completion for Robotic Grasp
abstract
Real-world robotic grasping can be done robustly if a complete 3D Point Cloud Data (PCD) of an object is available. However, in practice, PCDs are often incomplete when objects are viewed from few and sparse viewpoints before the grasping action, leading to the generation of wrong or inaccurate grasp poses. We propose a novel grasping strategy, named 3DSGrasp, that predicts the missing geometry from the partial PCD to produce reliable grasp poses. Our proposed PCD completion network is a Transformer-based encoder-decoder network with an Offset-Attention layer. Our network is inherently invariant to the object pose and point's permutation, which generates PCDs that are geometrically consistent and completed properly. Experiments on a wide range of partial PCD show that 3DSGrasp outperforms the best state-of-the-art method on PCD completion tasks and largely improves the grasping success rate in real-world scenarios. The code and dataset are available at: https://github.com/NunoDuarte/3DSGrasp.
Seyed Saber Mohammadi, Nuno Ferreira Duarte, Dimitrios Dimou, Yiming Wang 0002, Matteo Taiana, Pietro Morerio, Atabak Dehban, Plinio Moreno, Alexandre Bernardino, Alessio Del Bue, José Santos-Victor
ICRA8
2021 Learning Conditional Postural Synergies for Dexterous Hands: A Generative Approach Based on Variational Auto-Encoders and Conditioned on Object Size and Category
abstract
Postural synergies are used in robotics to facilitate the control of dexterous artificial hands. This is achieved by learning a latent space (synergy space) from grasp postures and directly controlling the hand in this space. In this work, we propose the use of a non-linear conditional model for learning the latent space, that can incorporate the object shape and size as additional variables. While on most of the previous works the evaluation criterion is the reconstruction error, we propose to use the smoothness of the latent space. Our model ranks better than other non-linear models in smoothness, which is a better criterion to evaluate in-hand manipulation tasks. We validate our arguments by executing regrasp trajectories in which our model outperforms all previous approaches.
Dimitrios Dimou, José Santos-Victor, Plinio Moreno
ICRA3
2021 Human-Robot greeting: tracking human greeting mental states and acting accordingly
abstract
Mobile social robots should be able to engage in interaction with people effectively. However, greeting someone is a complex task since it implies an exchange of social signals. Adam Kendon modeled human greetings as a set of six phases: initiation of approach, distance salutation, head dip, approach, final approach, and close salutation. Based on Kendon’s model, we propose a system for mobile social robots that manages the greeting process through the exchange of social signals. A Hidden Markov Model keeps track of the greeting stage through the observation of the human gestures, while a behavior tree generates appropriate robot actions. We used publicly available datasets to train the Hidden Markov Model. Evaluation on test sets showed an average greeting phase estimation accuracy of 80.9%. We tested the full system (Hidden Markov Model + Behavior Tree) in simulation and in a real world pilot experiment using the Vizzy robot, and it recognized and replicated the correct phase with an accuracy of 91.8% and 53.8%, respectively.
Manuel Carvalho, João Avelino, Alexandre Bernardino, Rodrigo M. M. Ventura, Plinio Moreno
IROS5
2021 Action anticipation for collaborative environments: The impact of contextual information and uncertainty-based prediction
Clebeson Canuto dos Santos, Plinio Moreno, Jorge Leonid Aching Samatelo, Raquel Frizera Vassallo, José Santos-Victor
Neurocomputing2
2020 Action-conditioned Benchmarking of Robotic Video Prediction Models: a Comparative Study
abstract
A defining characteristic of intelligent systems is the ability to make action decisions based on the anticipated outcomes. Video prediction systems have been demonstrated as a solution for predicting how the future will unfold visually, and thus, many models have been proposed that are capable of predicting future frames based on a history of observed frames (and sometimes robot actions). However, a comprehensive method for determining the fitness of different video prediction models at guiding the selection of actions is yet to be developed.Current metrics assess video prediction models based on human perception of frame quality. In contrast, we argue that if these systems are to be used to guide action, necessarily, the actions the robot performs should be encoded in the predicted frames. In this paper, we are proposing a new metric to compare different video prediction models based on this argument. More specifically, we propose an action inference system and quantitatively rank different models based on how well we can infer the robot actions from the predicted frames. Our extensive experiments show that models with high perceptual scores can perform poorly in the proposed action inference tests and thus, may not be suitable options to be used in robot planning systems.
Manuel Serra Nunes, Atabak Dehban, Plinio Moreno, José Santos-Victor
ICRA3
2019 The Impact of Domain Randomization on Object Detection: A Case Study on Parametric Shapes and Synthetic Textures*
abstract
Recent advances in deep learning-based object detection techniques have revolutionized their applicability in several fields. However, since these methods rely on unwieldy and large amounts of data, a common practice is to download models pre-trained on standard datasets and fine-tune them for specific application domains with a small set of domain-relevant images. In this work, we show that using synthetic datasets that are not necessarily photo-realistic can be a better alternative to simply fine-tune pre-trained networks. Specifically, our results show an impressive 25%improvement in the mAP metric over a fine-tuning baseline when only about 200 labelled images are available to train. Finally, an ablation study of our results is presented to delineate the individual contribution of different components in the randomization pipeline.
Atabak Dehban, João Borrego, Rui Pimentel de Figueiredo, Plinio Moreno, Alexandre Bernardino, José Santos-Victor
IROS4
2019 Project INSIDE: towards autonomous semi-unstructured human-robot social interaction in autism therapy
Francisco S. Melo, Alberto Sardinha, David Belo, Marta Couto, Miguel Faria 0001, Anabela Farias, Hugo Gamboa, Cátia Jesus, Mithun Kinarullathil, Pedro U. Lima, Luís Luz, André Mateus 0001, Isabel Melo, Plinio Moreno, Daniel Faustino de Noronha Osório, Ana Paiva 0001, Jhielson M. Pimentel, Rodrigo M. M. Ventura
Artif. Intell. Medicine14
2018 The Power of a Hand-shake in Human-Robot Interactions
abstract
In this paper, we study the influence of a handshake in the human emotional bond to a robot. In particular, we evaluate the human willingness to help a robot whether the robot first introduces itself to the human with or without a handshake. In the tested paradigm the robot and the human have to perform a joint task, but at a certain stage, the robot needs help to navigate through an obstacle. Without requesting explicit help from the human, the robot performs some attempts to navigate through the obstacle, suggesting to the human that it requires help. In a study with 45 participants, we measure the human's perceptions of the social robot Vizzy, comparing the handshake vs non-handshake conditions. In addition, we evaluate the influence of a handshake in the pro-social behaviour of helping it and the willingness to help it in the future. The results show that a handshake increases the perception of Warmth, Animacy, Likeability, and the tendency to help the robot more, by removing the obstacle.
João Avelino, Plinio Moreno, Alexandre Bernardino, Filipa Correia, Ana Paiva 0001, João Catarino, Pedro Ribeiro 0006
IROS2
2017 Relational Affordance Learning for Task-Dependent Robot Grasping
Laura Antanas, Anton Dries, Plinio Moreno, Luc De Raedt
ILP3
2015 Relational Kernel-Based Grasping with Numerical Features
Laura Antanas, Plinio Moreno, Luc De Raedt
ILP2
2015 Efficient pose estimation of rotationally symmetric objects
Rui Pimentel de Figueiredo, Plinio Moreno, Alexandre Bernardino
Neurocomputing2
2015 People and Mobile Robot Classification Through Spatio-Temporal Analysis of Optical Flow
abstract
The goal of this work is to distinguish between humans and robots in a mixed human-robot environment. We analyze the spatio-temporal patterns of optical flow-based features along several frames. We consider the Histogram of Optical Flow (HOF) and the Motion Boundary Histogram (MBH) features, which have shown good results on people detection. The spatio-temporal patterns are composed of groups of feature components that have similar values on previous frames. The groups of features are fed into the FuzzyBoost algorithm, which at each round selects the spatio-temporal pattern (i.e. feature set) having the lowest classification error. The search for patterns is guided by grouping feature dimensions, considering three algorithms: (a) similarity of weights from dimensionality reduction matrices, (b) Boost Feature Subset Selection (BFSS) and (c) Sequential Floating Feature Selection (SFSS), which avoid the brute force approach. The similarity weights are computed by the Multiple Metric Learning for large Margin Nearest Neighbor (MMLMNN), a linear dimensionality algorithm that provides a type of Mahalanobis metric Weinberger and Saul, J. MaCh. Learn. Res.10 (2009) 207–244. The experiments show that FuzzyBoost brings good generalization properties, better than the GentleBoost, the Support Vector Machines (SVM) with linear kernels and SVM with Radial Basis Function (RBF) kernels. The classifier was implemented and tested in a real-time, multi-camera dynamic setting.
Plinio Moreno, Dario Figueira, Alexandre Bernardino, José Santos-Victor
Int. J. Pattern Recognit. Artif. Intell.1
2013 Multi-object detection and pose estimation in 3D point clouds: A fast grid-based Bayesian Filter
abstract
We address the problem of object detection and pose estimation using 3D dense data in a multiple object library scenario. State-of-the-art object detection and pose estimation methods are able cope with background clutter and occlusion with acceptable noise levels in the single object scenario. However, with multiple object libraries, even moderate amount of noise lead to frequent object identity switches and serious pose estimation errors. To attenuate these effects, we propose a joint object-id and pose filtering approach using grid-based Recursive Bayesian Filters (RBF). The grid method considers as state variables the object label and its pose, and models the dynamics of the filter with two “inertia” parameters: one for the object label and the other for the object pose. Sensor noise characteristics are taken into account with an observation noise parameter. To allow real-time functionality we propose a selective update approach that dynamically reduces the set of hypotheses evaluated at run time. We present results in realistic scenarios and compare our approach with state-of-the-art approaches in a three object problem, with significant performance improvements.
Rui Pimentel de Figueiredo, Plinio Moreno, Alexandre Bernardino, José Santos-Victor
ICRA2
2013 On the use of probabilistic relational affordance models for sequential manipulation tasks in robotics
abstract
In this paper we employ probabilistic relational affordance models in a robotic manipulation task. Such affordance models capture the interdependencies between properties of multiple objects, executed actions, and effects of those actions on objects. Recently it was shown how to learn such models from observed video demonstrations of actions manipulating several objects. This paper extends that work and employs those models for sequential tasks. Our approach consists of two parts. First, we employ affordance models sequentially in order to recognize the individual actions making up a demonstrated sequential skill or high level concept. Second, we utilize the models of concepts to plan a suitable course of action to replicate the observed consequences of a demonstration. For this we adopt the framework of relational Markov decision processes. Empirical results show the viability of the affordance models for sequential manipulation skills for object placement.
Bogdan Moldovan, Plinio Moreno, Martijn van Otterlo
ICRA2
2012 Learning relational affordance models for robots in multi-object manipulation tasks
abstract
Affordances define the action possibilities on an object in the environment and in robotics they play a role in basic cognitive capabilities. Previous works have focused on affordance models for just one object even though in many scenarios they are defined by configurations of multiple objects that interact with each other. We employ recent advances in statistical relational learning to learn affordance models in such cases. Our models generalize over objects and can deal effectively with uncertainty. Two-object interaction models are learned from robotic interaction with the objects in the world and employed in situations with arbitrary numbers of objects. We illustrate these ideas with experimental results of an action recognition task where a robot manipulates objects on a shelf.
Bogdan Moldovan, Plinio Moreno, Martijn van Otterlo, José Santos-Victor, Luc De Raedt
ICRA2
2010 Unsupervised and Online Update of Boosted Temporal Models: The UAL2Boost
abstract
The application of learning-based vision techniques to real scenarios usually requires a tunning procedure, which involves the acquisition and labeling of new data and in situ experiments in order to adapt the learning algorithm to each scenario. We address an automatic update procedure of the L2boost algorithm that is able to adapt the initial models learned off-line. Our method is named UAL2Boost and present three new contributions: (i) an on-line and continuous procedure that updates recursively the current classifier, reducing the storage constraints, (ii) a probabilistic unsupervised update that eliminates the necessity of labeled data in order to adapt the classifier and (iii) a multi-class adaptation method. We show the applicability of the on-line unsupervised adaptation to human action recognition and demonstrate that the system is able to automatically update the parameters of the L2boost with linear temporal models, thus improving the output of the models learned off-line on new video sequences, in a recursive and continuous way. The automatic adaptation of UAL2Boost follows the idea of adapting the classifier incrementally: from simple to complex.
Pedro Ribeiro 0006, Plinio Moreno, José Santos-Victor
ICMLA2
2009 ISROBOTNET: A testbed for sensor and robot network systems
abstract
This paper introduces a testbed for sensor and robot network systems, currently composed of 10 cameras and 5 mobile wheeled robots equipped with several sensors for self-localization, obstacle avoidance and vision cameras, and wireless communications. The testbed includes a service-oriented middleware to enable fast prototyping and implementation of algorithms previously tested in simulation, as well as to simplify integration of subsystems developed by different partners. We survey an integrated approach to human-robot interaction that has been developed supported by the testbed under an European research project. The application integrates innovative methods and algorithms for people tracking and waving detection, cooperative perception among static and mobile cameras to improve people tracking accuracy, as well as decision-theoretical approaches to sensor selection and task allocation within the sensor network.
Marco Barbosa, Alexandre Bernardino, Dario Figueira, José António Gaspar, Nelson Gonçalves, Pedro U. Lima, Plinio Moreno, Abdolkarim Pahliani, José Santos-Victor, Matthijs T. J. Spaan, João Sequeira 0001
IROS7
2009 Improving the SIFT descriptor with smooth derivative filters
Plinio Moreno, Alexandre Bernardino, José Santos-Victor
Pattern Recognit. Lett.1