Paulo Menezes 0001

dblp:49/4334 · also Paulo Jorge Carvalho Menezes · DBLP profile ↗
← Back
26ranked-venue papers
2as first author
6since 2021 · last 2026
0000-0002-4903-3554ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 16 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 3 since 2021Systems, architecture and hardware · 5 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Improving long-term human motion prediction for physical exercise actions using a lightweight approach
abstract
Predicting human body movements paves the way for many modern applications, such as pedestrian path control in autonomous driving, adaption of human-robot cooperation, and many more. Yet, it is not just about knowing what comes next, it’s about anticipating, guiding and evaluating motion in real-time, which can be particularly interesting for the automated analysis of physical exercises. Therefore, we propose a dual-stream model that is sensitive to the underlying temporal patterns present in physical exercise actions. The model leverages both motion history and its accumulation over time to improve the prediction accuracy of the 3D poses. The training procedure is built on the premise that learning the necessary displacements from the last pose seen to the future ones is more effective than direct pose regression. Moreover, the supervision loss is designed to follow certain principles of human kinematics. Namely, the human joints should not have the same weight during the optimization process as they tend display different variability in motion characteristics. The overall results show that our proposal improves the long-term prediction of human movements in the target Fit3D dataset. Moreover, our model follows the recent success of MLP-based approaches that emphasize that simplicity is the key to success and contains only ≈ 0.28 M trainable parameters.
Bruno Ferreira 0005, Paulo Menezes 0001, Jorge P. Batista
Expert Syst. Appl.2
2025 Orlock: A Modular Agent for Natural Interaction in Indoor Navigation
abstract
Providing effective indoor spatial guidance through natural language remains a key challenge in human-computer interaction. This paper presents Orlock, a virtual agent designed to support wayfinding and spatial orientation within a university building. To ensure a consistent baseline, building maps were shown to participants unfamiliar with the space. Following their interaction with Orlock, participants completed an 18-item questionnaire covering five key dimensions: clarity, usefulness, ease of interaction, overall satisfaction, and design preferences. Results from the 7-point Likert-Scale responses indicated generally high satisfaction, with a clear preference for multimodal feedback combining visual and verbal instructions. These findings highlight the importance of multimodal interaction in orientation systems and offer practical design insights for future development of intelligent spatial guidance agents.
Gonçalo Carnaz, Esmeralda Faria, Paulo Menezes 0001
SMC3
2024 Exploring Continuous Awareness Modelling for Improving Worker Safety and Trust
abstract
The Industry 5.0 Human-Robot Interaction (HRI) goals have the potential for significantly reducing production costs, but require factory workers to trust that robots will not cause them physical harm when sharing the workspace in order to be achieved. To this end, we propose a model that adjusts the robot’s actions to an estimate of the workers’ awareness of its behavior. Its premise is that mutual awareness should be continuous, and if no recent visual contact has been established, then the robot should adapt its range of motion and speed in order to minimize the intersection between the robot’s space and the worker’s space. On the other hand, if the worker steadily establishes visual contact from times to times, the robot operates under a normal or a faster configuration. After a period of habituation, this awareness model can be further tweaked to continuously improve production times, while improving worker’s safety and trust. We present preliminary results of a pilot study with 32 participants, to asses if the existence of this awareness model impacts the outcomes of the combined task. The obtained tepidly optimistic results are the key to build upon in future designs.
Andrey Solovov, Vishal Gautam, Bruno Ferreira 0005, Gustavo Assunção, Antonio Marín-Hernández, Paulo Menezes 0001
RO-MAN7
2023 Adapting Behavior and Persistence via Reinforcement and Self-Emotion Mediated Exploration in a Social Robot
abstract
Adaptability and behavioral diversity are core components of social interactions between humans. Naturally, these are traits research should strive to achieve in social robotics so agents may be better accepted and engage with their user peers. In this paper, we propose a novel activity modulation to increase behavioral diversity, based on a surprise-exploration correlation model, in a social robot undergoing behavioral optimization to user state and preference. This framework was tested with 21 participants to assess preferences as well as the impact that action variability and persistence would have on user perception of the robot. Results indicate a positive effect of persistence and variability over robot likability as well as user engagement, contributing insight for future research in social robotics.
Gustavo Assunção, Alessandra Sorrentino, Jorge Dias 0001, Miguel Castelo-Branco, Paulo Menezes 0001, Filippo Cavallo
RO-MAN5
2022 Transformers for Workout Video Segmentation
abstract
Temporal analysis of workout routines comes with a major difficulty: the observed actions or processes are naturally sequences of variable length. Besides being highly dependent on the exercise itself, the faster or slower repetition of the pattern’s temporal dynamics is also determined by the individual’s performance. In this paper, we present a Transformer-based Deep Neural Network to perform classification over 19 phases of 5 exercises, common to CrossFit routines. From this baseline, we aim to perform workout video segmentation, while creating a descriptor that enables repetition counting or feedback on the exercise execution. More specifically, a 2D Human Pose Network creates heatmap-based features that capture the human body pose, which are fed into the Transformer Encoder with 4 attention heads. A final Multilayer Perceptron with 3 Dense layers performs the phase classification task. To this end, we have trained our model using a previously acquired dataset that is naturally imbalanced, e.g. 6 classes have less than 8k samples and 9 classes have more than 24k. Finally, the obtained results show that we are able to divide videos in a temporally consistent manner, outperforming a state-of-the-art model that counts repetitive actions, specifically for 4 out of the 5 exercises.
Bruno Ferreira 0005, Paulo Menezes 0001, Jorge P. Batista
ICIP2
2021 Deep learning approaches for workout repetition counting and validation
Bruno Ferreira 0005, Pedro M. Ferreira 0005, Gil Pinheiro, Nelson Figueiredo, Filipe Carvalho, Paulo Menezes 0001, Jorge P. Batista
Pattern Recognit. Lett.6
2020 Intermediary Fuzzification in Speech Emotion Recognition
abstract
Affective systems are getting increasingly more attention from researchers and high-tech companies in order to enable the acknowledgment or adaptation to a user's mood. Emotion classification is typically a hard problem due to the number of subtle cues which are present in human facial and body expressions, or in voiced utterances. Another critical factor is that typically used models tend to map emotions into all-or-nothing regions with artificially sharp divisions among them, a view which is rather unsupported in the field of psychology and human behavioral analysis. In this paper we propose the inclusion of an intermediary fuzzy layer in a VGGVox-based NN, whose aim is to deal with the inherently foggy transitions between emotional states. This neuro-fuzzy model was trained and evaluated against four emotional speech databases and has shown improvements in the classification performance over a non-fuzzy counterpart. Observed performances were also on-par or above those of other current state-of-the-art techniques.
Gustavo Assunção, Paulo Menezes 0001
FUZZ-IEEE2
2019 ClusterNav: Learning-Based Robust Navigation Operating in Cluttered Environments
abstract
Robust autonomous navigation is one of the most important aspects in the acceptance of social robots by elderly users. Traditional model-based navigation techniques provide a stable theoretical and practical foundation for autonomous operation in domestic environments, but fall short in achieving human-like, acceptable behaviour while still being able to robustly navigate cluttered environments. In this work, we propose ClusterNav, a novel learning-based technique for navigation. Our technique consists of teaching the robot how it should move in the environment in a human-like manner, capturing key features of this demonstration in a geometric representation of the environment. This representation is then used to generate new trajectories for execution, allowing the robot move in an acceptable manner. We have tested our technique in a real environment in an elderly care facility, comparing it with the traditional model-based approach. Tests involved both expert and non-expert users teleoperating the robot. Results show that ClusterNav is capable of navigating the environment, achieving better similarity with the reference trajectories and higher execution speed when compared to the model-based approach.
Gonçalo S. Martins, Rui P. Rocha, Fernando J. Pais, Paulo Menezes 0001
ICRA4
2019 Towards natural interaction in immersive reality with a cyber-glove
abstract
Over the past few years, virtual and mixed reality systems have evolved significantly yielding high immersive experiences. Most of the metaphors used for interaction with the virtual environment do not provide the same meaningful feedback, to which the users are used to in the real world. This paper proposes a cyber-glove to improve the immersive sensation and the degree of embodiment in virtual and mixed reality interaction tasks. In particular, we are proposing a cyber-glove system that tracks wrist movements, hand orientation and finger movements. It provides a decoupled position of the wrist and hand, which can contribute to a better embodiment in interaction and manipulation tasks. Additionally, the detection of the curvature of the fingers aims to improve the proprioceptive perception of the grasping/releasing gestures more consistent to visual feedback. The cyber-glove system is being developed for VR applications related to real estate promotion, where users have to go through divisions of the house and interact with objects and furniture. This work aims to assess if glove-based systems can contribute to a higher sense of immersion, embodiment and usability when compared to standard VR hand controller devices (typically button-based). Twenty-two participants tested the cyber-glove system against the HTC Vive controller in a 3D manipulation task, specifically the opening of a virtual door. Metric results showed that 83% of the users performed faster door pushes, and described shorter paths with their hands wearing the cyber-glove. Subjective results showed that all participants rated the cyber-glove based interactions as equally or more natural, and 90% of users experienced an equal or a significant increase in the sense of embodiment.
Luís Almeida 0002, Elio Lopes, Beril Yalçinkaya, Rodolfo Martins, Ana C. Lopes, Paulo Menezes 0001, Gabriel Pires
SMC6
2019 Classification of FACS-Action Units with CNN Trained from Emotion Labelled Data Sets
abstract
This paper explores the adaptation of a convolutional neural network (CNN) designed for emotion classification to the extraction of human expressions action units. An adaptation of the network structure transforms a single label into a multi-label classifier which supports the simultaneous recognition of multiple action units that compose human expressions, according to Ekman's FACS. The dataset used for this work was FER-2013 that includes exemplars of seven basic expressions (Angry, Disgust, Fear, Happy, Sad, Surprise, Neutral). In order to enhance the quality of the results, the dataset was augmented with random perturbations from a wide set including: translation, scale and horizontal flip. The results obtained demonstrate that it is possible to train a CNN for multi-label expression classification from an emotion-labelled dataset.
Pedro Carvalho Gerardo, Paulo Menezes 0001
SMC2
2019 Toward a Context-Aware Human-Robot Interaction Framework Based on Cognitive Development
abstract
The purpose of this paper was to understand how an agent's performance is affected when interaction workflows are incorporated in its information model and decision-making process. Our expectation was that this incorporation could reduce errors and faults on agent's operation, improving its interaction performance. We based this expectation on the existing challenges in designing and implementing artificial social agents, where an approach based on predefined user scenarios and action scripts is insufficient to account for uncertainty in perception or unclear expectations from the user. Therefore, we developed a framework that captures the expected behavior of the agent into descriptive scenarios and then translated these into the agent's information model and used the resulting representation in probabilistic planning and decision making to control interaction. Our results indicated an improvement in terms of specificity while maintaining precision and recall, suggesting that the hypothesis being proposed in our approach is plausible. We believe the presented framework will contribute to the field of cognitive robotics, e.g., by improving the usability of artificial social companions, thus overcoming the limitations imposed by approaches that use predefined static models for an agent's behavior resulting in non-natural interaction.
João Quintas, Gonçalo S. Martins, Luís Santos 0001, Paulo Menezes 0001, Jorge Dias 0001
IEEE Trans. Syst. Man Cybern. Syst.4
2017 Interoperability in cloud robotics - Developing and matching knowledge information models for heterogenous multi-robot systems
abstract
Every file, document, database and digital information is now going through the Cloud. Leveraged by the developments in information systems, Cloud Robotics is evolving at a steady pace and raised attention in the past 5 years. This recent field of Robotics is allowing engineers to envisage new and exciting applications for robots in the near future. This work proposes Cloud Robotics as a mean to integrate semantic reasoning in a multi-robot system, using self-created knowledge bases in each robot, in order to perform the coordination of complex task allocation. An auction-based coordination method and a knowledge matching algorithm were implemented to study this subject. The obtained results demonstrated that, the coordination of a large multi-robot system and the knowledge matching process can be computationally demanding, thus making them perfect candidate features to be “cloudyfied”.
João Quintas, Paulo Menezes 0001, Jorge Dias 0001
RO-MAN2
2017 Information Model and Architecture Specification for Context Awareness Interaction Decision Support in Cyber-Physical Human-Machine Systems
abstract
This paper aims to contribute to situation, activity, and goal awareness in cyber-physical human-machine systems (HMS) by presenting a new information model and specifications for a decision-making component that can be integrated in current system architectures. The objective of this work is to improve the efficacy, acceptance, adaptability, and overall performance of HMS and human-system interaction (HSI) applications using a context-based approach. Our hypothesis is that we can enhance current interaction functionalities by integrating context and interaction information models into a decision-making component that behaves as a supervision process for controlling interaction. In HSI, we aim to define a general human model that may lead to principles and algorithms, allowing more natural and effective interaction between humans and artificial agents. The approach was implemented and tested targeting application in the domain of active and assisted living. The challenge of user acceptance is of vital importance for future solutions and is still one of the major reasons for reluctance to adopt cyber-physical systems in this domain.
João Quintas, Paulo Menezes 0001, Jorge Dias 0001
IEEE Trans. Hum. Mach. Syst.2
2016 A control architecture for Hybrid underwater intervention systems
abstract
As far as we know, currently all the underwater interventions, requiring robotic manipulation, are carried out by using a well-known technology, based on the commercial work class ROVs (Remotely Operated Vehicles). These systems require a lot of support, including a vessel in the surface, where a pilot (i.e. user expert) is able of teleoperating the actions of the ROV by means of an umbilical cable. On the other hand, from the last ten years new research has been developed, promoting in the market the so called Autonomous Underwater Vehicles (AUVs), enabling interventions. Obviously the AUVs have some main advantages over ROVs: no vessel, umbilical or pilots are needed now, but presenting a critical drawback: any kind of potential manipulation skills are impossible. Thus, the present situation becomes in a new vehicle: the Hybrid-ROV (HROV), trying to join the best of both systems, ROVs and AUVs. However, new problems arise concerning the HROV control. Now the system can operate in two different ways: autonomous or teleoperated, and so, a new control approach should be developed. This paper presents a control architecture for an HROV, discussing details of this approach from the human-robot interaction viewpoint.
J. C. Garcia, Javier Pérez, Paulo Menezes 0001, Pedro J. Sanz
SMC3
2016 Context-based decision system for human-machine interaction applications
abstract
In this paper we present a decision process to auto-adapt and improve human-machine interaction, simplifying the integration of algorithms and functionalities. The decision process is part of an innovative approach that integrates contextual information to orchestrate behaviours of an interactive system (i.e. perception and actuation features involved during interaction). Classical approaches focus on designing and implementing algorithms that take into account several environment features (e.g. light, pose, etc.) to adapt its performance obtaining accurate results. An advantage of these approaches is to concentrate complexity in one algorithm leading to simple system architectures. In the other hand, a disadvantage of such approaches is their limitation to adapt to conditions under different scenarios, which typically requires manual adjustments to compensate changes of environment features. Our hypothesis is that, we can improve the overall performance of human-machine interaction process if a decision process is introduced, which is responsible for selecting the most adequate actions/algorithms, with maximum performance, that achieve a certain goal under a given context. The results from exploratory simulations validate the proposed approach to be more effective in attaining specific goals in the interaction process, resorting to algorithms with low complexity.
João Quintas, Paulo Menezes 0001, Jorge Dias 0001
SMC2
2014 Augmented reality to improve STEM motivation
abstract
This paper presents an exploratory study about educational potentialities of an augmented reality (AR) application developed for DC circuit fundamentals. Particularly the study aims to characterize student involvement using the application as well as its use as an additional experimental tool and to characterize how students perceive their experience and their learning through the use of this AR application. It is also briefly described how this application was developed and how the exploratory study was implemented involving STEM students. The AR application confirmed to be manageable and students have explored its configurations intuitively. Additionally, the AR tool usability according to our preliminary results showed to be effective for the AR developed application purposes, has induced student satisfaction and revealed very good student perceptions about learning perspectives. So, this study showed this AR application for DC circuits has a great educational potential.
Maria Teresa Restivo, Maria de Fátima Chouzal, José Carlos Rodrigues, Paulo Menezes 0001, Joaquim Bernardino Lopes
EDUCON4
2014 Be the robot: Human embodiment in tele-operation driving tasks
abstract
This paper proposes a new interaction mechanism for tele-operating a mobile robot. The approach explores the notion of telepresence and physical embodiment to create what may be called tele-embodiment. Its principle is that the operator will see himself at the remote site and this will enable him/her to better operate the robot. Four interaction styles were experimentally compared, from the traditional joystick approaches to more innovative based on natural body posture intentions. The environment perception is provided by the visual feedback, according to head pose behaviour. The results show that the gesture and body based methods improves the user dexterity performing this kind of task. Moreover, the present study suggests that, when a person is focused on the task, achieving the ownership illusion towards remote body, there are autonomic responses that correspond to what would be expected in events that take place in reality (like avoiding collisions).
Luís Almeida 0002, Bruno Patrão, Paulo Menezes 0001, Jorge Dias 0001
RO-MAN3
2013 Context-aware cooperation between human and robotic teams in catastrophic incidents
abstract
The study of cooperative interaction between multi-party multi-agent teams that include humans and robots is a recent scientific challenge. Preliminary empiric results about this interaction, in the scope of search and rescue applications, demonstrate the need for deeper studies on how humans should interact with teams of autonomous mobile robots and on how to establish a mutual beneficial interaction. This work presents the work in progress in the scope of CHOPIN1project, which aims to address some of these issues and will focus on devising new methods for collaborative context awareness and context sharing between teams of humans and teams of robots.
João Quintas, Paulo Menezes 0001, Jorge Dias 0001
RO-MAN2
2013 Context-based perception and understanding of human intentions
abstract
This work focus in the importance of context awareness and intention understanding capabilities in modern robots when faced with different situations. The objective is to be capable of providing new features for robots, which enable new real-world applications, and extend their autonomy, in terms of self-management and cooperation with humans or other systems. Gaze estimation and gesture interpretation are modalities, closely related with context-dependent human intention understanding, that are addressed in this work.
João Quintas, Paulo Menezes 0001, Jorge Dias 0001
RO-MAN2
2012 Context-based understanding of interaction intentions
abstract
This paper focus in the importance of context awareness and intention understanding capabilities in modern robots when faced with different situations. The inclusion of such requirements in robot design aim for more intelligent robots capable to adapt its behaviours to the faced situations. Gaze estimation and gesture interpretation are modalities, closely related with context-depent human intention understanding, that are addressed in this work.
João Quintas, Luís Almeida 0002, Miguel Brito, Gustavo Quintela, Paulo Menezes 0001, Jorge Dias 0001
RO-MAN5
2011 Towards human motion capture from a camera mounted on a mobile robot
Paulo Menezes 0001, Frédéric Lerasle, Jorge Dias 0001
Image Vis. Comput.1
2006 Visual Tracking Modalities for a Companion Robot
abstract
This article presents the development of a human-robot interaction mechanism based on vision. The functionalities required for such mechanism range from user detection and recognition, to gesture tracking. Particle filters, which are extensively described in the literature, are well suited to this context as they enable a straight combination of several visual cues like colour, shape or motion. Additionally, different algorithms can be considered for a better handling of the particles depending of the context. This article presents the visual functionalities developed namely user recognition and following, and 3D gestures tracking. The challenge is to find which algorithms and visual cues fulfil the best, the requirements of the considered functionalities for our companion robot. The employed methods to attain these required functionalities and their results are presented
Paulo Menezes 0001, Frédéric Lerasle, Jorge Dias 0001
IROS1
2006 Rackham: An Interactive Robot-Guide
abstract
Rackham is an interactive robot-guide that has been used in several places and exhibitions. This paper presents its design and reports on results that have been obtained after its deployment in a permanent exhibition. The project is conducted so as to incrementally enhance the robot functional and decisional capabilities based on the observation of the interaction between the public and the robot. Besides robustness and efficiency in the robot navigation abilities in a dynamic environment, our focus was to develop and test a methodology to integrate human-robot interaction abilities in a systematic way. We first present the robot and some of its key design issues. Then, we discuss a number of lessons that we have drawn from its use in interaction with the public and how that will serve to refine our design choices and to enhance robot efficiency and acceptability
Aurélie Clodic, Sara Fleury, Rachid Alami 0001, Raja Chatila 0001, Gérard Bailly, Ludovic Brethes, Maxime Cottret, Patrick Danès, Xavier Dollat, Frédéric Elisei, Isabelle Ferrané, Matthieu Herrb, Guillaume Infantes, Christian Lemaire, Frédéric Lerasle, Jérôme Manhes, Patrick Marcoul, Paulo Menezes 0001, Vincent Montreuil
RO-MAN18
2004 Human-robot Interaction based on Haar-like Features and Eigenfaces
abstract
This paper describes a machine learning approach for visual object detection and recognition which is capable of processing images rapidly and achieving high detection and recognition rates. This framework is demonstrated on, and in part motivated by, the task of human-robot interaction. There are three main parts on this framework. The first is the person's face detection used as a preprocessing system to the second stage which is the recognition of the face of the person interacting with the robot, and the third one is the hand detection. The detection technique is based on Haar-like features introduced by Viola et al. and then improved by Lienhart et al. The eigenimages and PCA are used in the recognition stage of the system. Used in real-time human-robot interaction applications the system is able to detect and recognise faces at 10.9 frames per second in a PIV 2.2 GHz equipped with a USB camera.
Jose Barreto, Paulo Menezes 0001, Jorge Dias 0001
ICRA2
2004 Face Tracking and Hand Gesture Recognition for Human-Robot Interaction
abstract
The interaction between man and machines has become an important topic for the robotic community as it can generalise the use of robots. For active H/R interaction scheme, the robot needs to detect human faces in its vicinity and then interpret canonical gestures of the tracked person assuming this interlocutor has been beforehand identified. In this context, we depict functions suitable to detect and recognise faces in video stream and then focus on face or hand tracking functions. An efficient colour segmentation based on a watershed on the skin-like coloured pixels is proposed. A new measurement model is proposed to take into account both shape and colour cues in the particle filter to track face or hand silhouettes in video stream. An extension of the basic condensation algorithm is proposed to achieve recognition of the current hand posture and automatic switching between multiple templates in the tracking loop. Results of tracking and recognition using are illustrated in the paper and show the process robustness in cluttered environments and in various light conditions. The limits of the method and future works are also discussed.
Ludovic Brethes, Paulo Menezes 0001, Frédéric Lerasle, Jean-Bernard Hayet
ICRA2
1997 Avoiding obstacles using a connectionist network
abstract
In this article, visual data obtained by a binocular active vision system is integrated, together with ultrasonic range measurements, in the development of a obstacle detection and avoidance system based on a connectionist grid. The traditional notion of probabilistic occupation grid is extended through the use of a three-layer structure of connectionist networks which allows the integration of several sensorial modalities (in this case ultrasonic sensor readings and stereo vision information) in a probabilistic environment representation. The connectionist nature of the network also allows us to deal with obstacle avoidance by using a mechanism similar to potential field over a discrete set of the robot's configuration space with each grid node representing a possible configuration. The value in each grid node gives us a measure of the configuration occupancy probability and can also be used to guide the robot to a predefined goal configuration simulating a simple gradient descending technique. Finally we present experimental results obtained with the implementation of the above method in a mobile platform which also provides the support for the sensing devices described throughout the article.
A. Silva, Paulo Menezes 0001, Jorge Dias 0001
IROS2