Francisco S. Melo

dblp:86/839 · DBLP profile ↗
← Back
62ranked-venue papers
13as first author
24since 2021 · last 2026
0000-0001-5705-7372ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 58 · 13 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 4 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 13 · 2 since 2021Systems, architecture and hardware · 9 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author
YearPublicationVenuePosition
2026 Centralized Training with Hybrid Execution in Multi-Agent Reinforcement Learning via Predictive Observation Imputation (Abstract Reprint)
abstract
We study hybrid execution in multi-agent reinforcement learning (MARL), a paradigm where agents aim to complete cooperative tasks with arbitrary communication levels at execution time by taking advantage of information-sharing among the agents. Under hybrid execution, the communication level can range from a setting in which no communication is allowed between agents (fully decentralized), to a setting featuring full communication (fully centralized), but the agents do not know beforehand which communication level they will encounter at execution time. We contribute MARO, an approach that makes use of an auto-regressive predictive model, trained in a centralized manner, to estimate missing agents' observations at execution time. We evaluate MARO on standard scenarios and extensions of previous benchmarks tailored to emphasize the impact of partial observability in MARL. Experimental results show that our method consistently outperforms relevant baselines, allowing agents to act with faulty communication while successfully exploiting shared information.
Pedro P. Santos, Diogo S. Carvalho, Miguel Vasco, Alberto Sardinha, Pedro Santos 0001, Ana Paiva 0001, Francisco S. Melo
AAAI7
2025 The Number of Trials Matters in Infinite-Horizon General-Utility Markov Decision Processes
abstract
The general-utility Markov decision processes (GUMDPs) framework generalizes the MDPs framework by considering objective functions that depend on the frequency of visitation of state-action pairs induced by a given policy. In this work, we contribute with the first analysis on the impact of the number of trials, i.e., the number of randomly sampled trajectories, in infinite-horizon GUMDPs. We show that, as opposed to standard MDPs, the number of trials plays a key-role in infinite-horizon GUMDPs and the expected performance of a given policy depends, in general, on the number of trials. We consider both discounted and average GUMDPs, where the objective function depends, respectively, on discounted and average frequencies of visitation of state-action pairs. First, we study policy evaluation under discounted GUMDPs, proving lower and upper bounds on the mismatch between the finite and infinite trials formulations for GUMDPs. Second, we address average GUMDPs, studying how different classes of GUMDPs impact the mismatch between the finite and infinite trials formulations. Third, we provide a set of empirical results to support our claims, highlighting how the number of trajectories and the structure of the underlying GUMDP influence policy evaluation.
Pedro P. Santos, Alberto Sardinha, Francisco S. Melo
ICML3
2025 Optimize and Coordinate Multiple DMPs Under Constraints to Achieve a Collaborative Manipulation Task
abstract
This paper addresses a significant challenge in achieving collaborative tasks; how can a robot or multiple robots, endowed with a library of pre-learned primitive movements, generate multiple simultaneous coordinated robotic movements, adapting and optimizing those in the library, to complete one collaborative task? This work can thus be seen as a follow-up to the work with a motion presented as dynamic movement primitive (DMP) that now considers collaborative tasks and the existence of multiple robots/manipulators. Specifically, we start with a simple task using one DMP and extend it to accommodate the coordinated execution of multiple DMPs in robots with multiple manipulators or-alternatively-multiple robots with a single manipulator. We investigate mechanisms to jointly optimize multiple DMPs to perform one task in a coordinated fashion. The joint trajectory is built from initial DMPs learned for a single manipulator, and its optimization must comply with task-specific constraints. We illustrate the application of our approach both in a simulated environment and in a simulated and real Baxter robot.
Ali H. Kordia, Francisco S. Melo
ICRA2
2025 Networked Agents in the Dark: Team Value Learning under Partial Observability
Guilherme S. Varela, Alberto Sardinha, Francisco S. Melo
AAMAS3
2025 Distributed Value Decomposition Networks with Networked Agents
Guilherme S. Varela, Alberto Sardinha, Francisco S. Melo
AAMAS3
2025 Implicit Repair with Reinforcement Learning in Emergent Communication
Fábio Vital, Alberto Sardinha, Francisco S. Melo
AAMAS3
2025 Reinforcement learning in convergently non-stationary environments: Feudal hierarchies and learned representations
abstract
We study the convergence of Q -learning-based methods in convergently non-stationary environments, particularly in the context of hierarchical reinforcement learning and of dynamic features encountered in deep reinforcement learning. We demonstrate that Q -learning achieves convergence in tabular representations when applied to convergently non-stationary dynamics, such as the ones arising in a feudal hierarchical setting. Additionally, we establish convergence for Q -learning-based deep reinforcement learning methods with convergently non-stationary features, such as the ones arising in representation-based settings. Our findings offer theoretical support for the application of Q -learning in these complex scenarios and present methodologies for extending established theoretical results from standard cases to their convergently non-stationary counterparts.
Diogo S. Carvalho, Pedro Santos 0001, Francisco S. Melo
Artif. Intell.3
2025 Centralized training with hybrid execution in multi-agent reinforcement learning via predictive observation imputation
abstract
We study hybrid execution in multi-agent reinforcement learning (MARL), a paradigm where agents aim to complete cooperative tasks with arbitrary communication levels at execution time by taking advantage of information-sharing among the agents. Under hybrid execution, the communication level can range from a setting in which no communication is allowed between agents (fully decentralized), to a setting featuring full communication (fully centralized), but the agents do not know beforehand which communication level they will encounter at execution time. We contribute MARO, an approach that makes use of an auto-regressive predictive model, trained in a centralized manner, to estimate missing agents' observations at execution time. We evaluate MARO on standard scenarios and extensions of previous benchmarks tailored to emphasize the impact of partial observability in MARL. Experimental results show that our method consistently outperforms relevant baselines, allowing agents to act with faulty communication while successfully exploiting shared information.
Pedro P. Santos, Diogo S. Carvalho, Miguel Vasco, Alberto Sardinha, Pedro Santos 0001, Ana Paiva 0001, Francisco S. Melo
Artif. Intell.7
2024 TEAMSTER: Model-Based Reinforcement Learning for Ad Hoc Teamwork (Abstract Reprint)
abstract
This paper investigates the use of model-based reinforcement learning in the context of ad hoc teamwork. We introduce a novel approach, named TEAMSTER, where we propose learning both the environment's model and the model of the teammates' behavior separately. Compared to the state-of-the-art PLASTIC algorithms, our results in four different domains from the multi-agent systems literature show that TEAMSTER is more flexible than the PLASTIC-Model, by learning the environment's model instead of assuming a perfect hand-coded model, and more robust/efficient than PLASTIC-Policy, by being able to continuously adapt to newly encountered teams, without implicitly learning a new environment model from scratch.
João G. Ribeiro, Gonçalo Rodrigues, Alberto Sardinha, Francisco S. Melo
AAAI4
2024 Interactively Teaching an Inverse Reinforcement Learner with Limited Feedback
Rustam Zayanov, Francisco S. Melo, Manuel Lopes 0001
ICAART (1)2
2024 NeuralSolver: Learning Algorithms For Consistent and Efficient Extrapolation Across General Tasks
abstract
We contribute NeuralSolver, a novel recurrent solver that can efficiently and consistently extrapolate, i.e., learn algorithms from smaller problems (in terms of observation size) and execute those algorithms in large problems. Contrary to previous recurrent solvers, NeuralSolver can be naturally applied in both same-size problems, where the input and output sizes are the same, and in different-size problems, where the size of the input and output differ. To allow for this versatility, we design NeuralSolver with three main components: a recurrent module, that iteratively processes input information at different scales, a processing module, responsible for aggregating the previously processed information, and a curriculum-based training scheme, that improves the extrapolation performance of the method. To evaluate our method we introduce a set of novel different-size tasks and we show that NeuralSolver consistently outperforms the prior state-of-the-art recurrent solvers in extrapolating to larger problems, considering smaller training problems and requiring less parameters than other approaches.
Bernardo Esteves, Miguel Vasco, Francisco S. Melo
NeurIPS3
2024 "Guess what I'm doing": Extending legibility to sequential decision tasks
Miguel Faria 0001, Francisco S. Melo, Ana Paiva 0001
Artif. Intell.2
2024 The impact of data distribution on Q-learning with function approximation
abstract
Abstract We study the interplay between the data distribution and Q-learning-based algorithms with function approximation. We provide a unified theoretical and empirical analysis as to how different properties of the data distribution influence the performance of Q-learning-based algorithms. We connect different lines of research, as well as validate and extend previous results, being primarily focused on offline settings. First, we analyze the impact of the data distribution by using optimization as a tool to better understand which data distributions yield low concentrability coefficients. We motivate high-entropy distributions from a game-theoretical point of view and propose an algorithm to find the optimal data distribution from the point of view of concentrability. Second, from an empirical perspective, we introduce a novel four-state MDP specifically tailored to highlight the impact of the data distribution in the performance of Q-learning-based algorithms with function approximation. Finally, we experimentally assess the impact of the data distribution properties on the performance of two offline Q-learning-based algorithms under different environments. Our results attest to the importance of different properties of the data distribution such as entropy, coverage, and data quality (closeness to optimal policy).
Pedro P. Santos, Diogo S. Carvalho, Alberto Sardinha, Francisco S. Melo
Mach. Learn.4
2023 Theoretical Remarks on Feudal Hierarchies and Reinforcement Learning
abstract
Hierarchical reinforcement learning is an increasingly demanded resource for learning to make sequential decisions towards long term goals. Feudal hierarchies are among the most deployed frameworks. However, there are few theoretical results for hierarchical structures. In this work, we formalize the common two-level feudal hierarchy as two Markov decision processes, with the one on the high level being dependent on the policy executed at the low level. Despite the non-stationarity raised by the dependency, we show that each of the processes presents stable behavior. We then build on the first result to show that, regardless of the convergent learning algorithm used for the low level, convergence of both prediction and control algorithms at the high-level is guaranteed. Our results contribute with theoretical support for the use of feudal hierarchies in combination with standard reinforcement learning methods at each level.
Diogo S. Carvalho, Francisco S. Melo, Pedro Santos 0001
ECAI2
2023 Making Friends in the Dark: Ad Hoc Teamwork Under Partial Observability
abstract
This paper introduces a formal definition of the setting of ad hoc teamwork under partial observability and proposes a first-principled model-based approach which relies only on prior knowledge and partial observations of the environment in order to perform ad hoc teamwork. We make three distinct assumptions that set it apart previous works, namely: i) the state of the environment is always partially observable, ii) the actions of the teammates are always unavailable to the ad hoc agent and iii) the ad hoc agent has no access to a reward signal which could be used to learn the task from scratch. Our results in 70 POMDPs from 11 domains show that our approach is not only effective in assisting unknown teammates in solving unknown tasks but is also robust in scaling to more challenging problems. Supplementary material is available at https://github.com/jmribeiro/adhoc-teamwork-under-partial-observability.
João G. Ribeiro, Cassandro Martinho, Alberto Sardinha, Francisco S. Melo
ECAI4
2023 TEAMSTER: Model-based reinforcement learning for ad hoc teamwork
João G. Ribeiro, Gonçalo Rodrigues, Alberto Sardinha, Francisco S. Melo
Artif. Intell.4
2022 Geometric Multimodal Contrastive Representation Learning
abstract
Learning representations of multimodal data that are both informative and robust to missing modalities at test time remains a challenging problem due to the inherent heterogeneity of data obtained from different channels. To address it, we present a novel Geometric Multimodal Contrastive (GMC) representation learning method consisting of two main components: i) a two-level architecture consisting of modality-specific base encoders, allowing to process an arbitrary number of modalities to an intermediate representation of fixed dimensionality, and a shared projection head, mapping the intermediate representations to a latent representation space; ii) a multimodal contrastive loss function that encourages the geometric alignment of the learned representations. We experimentally demonstrate that GMC representations are semantically rich and achieve state-of-the-art performance with missing modality information on three different learning problems including prediction and reinforcement learning tasks.
Petra Poklukar, Miguel Vasco, Hang Yin 0001, Francisco S. Melo, Ana Paiva 0001, Danica Kragic
ICML4
2022 Perceive, Represent, Generate: Translating Multimodal Information to Robotic Motion Trajectories
abstract
We present Perceive-Represent-Generate (PRG), a novel three-stage framework that maps perceptual information of different modalities (e.g., visual or sound), corresponding to a series of instructions, to a sequence of movements to be executed by a robot. In the first stage, we perceive and preprocess the given inputs, isolating individual commands from the complete instruction provided by a human user. In the second stage we encode the individual commands into a multimodal latent space, employing a deep generative model. Finally, in the third stage we convert the latent samples into individual trajectories and combine them into a single dynamic movement primitive, allowing its execution by a robotic manipulator. We evaluate our pipeline in the context of a novel robotic handwriting task, where the robot receives as input a word through different perceptual modalities (e.g., image, sound), and generates the corresponding motion trajectory to write it, creating coherent and high-quality handwritten words.
Fábio Vital, Miguel Vasco, Alberto Sardinha, Francisco S. Melo
IROS4
2022 Cooperation and Learning Dynamics under Wealth Inequality and Diversity in Individual Risk
abstract
We examine how wealth inequality and diversity in the perception of risk of a collective disaster impact cooperation levels in the context of a public goods game with uncertain and non-linear returns. In this game, individuals face a collective-risk dilemma where they may contribute or not to a common pool to reduce their chances of future losses. We draw our conclusions based on social simulations with populations of independent reinforcement learners with diverse levels of risk and wealth. We find that both wealth inequality and diversity in risk assessment can hinder cooperation and augment collective losses. Additionally, wealth inequality further exacerbates long term inequality, causing rich agents to become richer and poor agents to become poorer. On the other hand, diversity in risk only amplifies inequality when combined with bias in group assortment—i.e., high probability that agents from the same risk class play together. Our results also suggest that taking wealth inequality into account can help to design effective policies aiming at leveraging cooperation in large group sizes, a configuration where collective action is harder to achieve. Finally, we characterize the circumstances under which risk perception alignment is crucial and those under which reducing wealth inequality constitutes a deciding factor for collective welfare.
Ramona Merhej, Fernando P. Santos 0001, Francisco S. Melo, Francisco C. Santos
J. Artif. Intell. Res.3
2022 Leveraging hierarchy in multimodal generative models for effective cross-modality inference
Miguel Vasco, Hang Yin 0001, Francisco S. Melo, Ana Paiva 0001
Neural Networks3
2021 Interactive Teaching with Groups of Unknown Bayesian Learners
Carla Guerra, Francisco S. Melo, Manuel Lopes 0001
AIED (2)2
2021 Movement recognition and prediction using DMPs
abstract
This paper proposes an approach for (a) recognizing an observed trajectory from a library of pre-learned motions; and (b) predicting the target position of such trajectory. In our approach, motions are represented as Dynamic Movement Primitives (DMPs). We use critical points from the observed trajectory to time-align it with those in the library. To match the observed trajectory with those in the library, we compare the changes in velocity orientation between consecutive critical points. The proposed approach is computationally light and, as such, can be performed at execution time. As for the prediction, we adopt a similar approach: after matching the observed trajectory to one in the library, we use the latter to predict the target point, modulating it to match the observed trajectory. Both recognition and prediction approaches are probabilistic and, as such, provide a measure of certainty/uncertainty in the recognition/prediction process. Such a measure of uncertainty is important in tasks involving human-robot collaboration, as it allows the robot to decide when it is sufficiently certain to act conditioned on the estimated trajectory. We illustrate our approach both in simulation and in a human-robot interaction scenario involving the Baxter robot.
Ali H. Kordia, Francisco S. Melo
ICRA2
2021 Understanding Robots: Making Robots More Legible in Multi-Party Interactions
abstract
In this work we explore implicit communication between humans and robots—through movement—in multi-party (or multi-user) interactions. In particular, we investigate how a robot can move to better convey its intentions using legible movements in multi-party interactions. Current research on the application of legible movements has focused on single-user interactions, causing a vacuum of knowledge regarding the impact of such movements in multi-party interactions. We propose a novel approach that extends the notion of legible motion to multi-party settings, by considering that legibility depends on all human users involved in the interaction, and should take into consideration how each of them perceives the robot’s movements from their respective points-of-view. We show, through simulation and a user study, that our proposed model of multi-user legibility leads to movements that, on average, optimize the legibility of the motion as perceived by the group of users. Our model creates movements that allow each human to more quickly and confidently understand what are the robot’s intentions, thus creating safer, clearer and more efficient interactions and collaborations.
Miguel Faria 0001, Francisco S. Melo, Ana Paiva 0001
RO-MAN2
2021 A Game AI Competition to Foster Collaborative AI Research and Development
abstract
Game artificial intelligence (AI) competitions are important to foster research and development on Game AI and AI in general. These competitions supply different challenging problems that can be translated into other contexts, virtual or real. They provide frameworks and tools to facilitate the research on their core topics and provide means for comparing and sharing results. A competition is also a way to motivate new researchers to study these challenges. In this article, we present theGeometry Friendsgame AI competition.Geometry Friendsis a two-player cooperative physics-based puzzle platformer computer game. The concept of the game is simple, though its solving has proven to be difficult. While the main and apparent focus of the game is cooperation, it also relies on other AI-related problems such as planning, plan execution, and motion control, all connected to situational awareness. All of these must be solved in real-time. In this article, we discuss the competition and the challenges it brings, and present an overview of the current solutions.
Ana Salta, Rui Prada, Francisco S. Melo
IEEE Trans. Games3
2020 An End-to-end Approach for Learning and Generating Complex Robot Motions from Demonstration
abstract
This paper proposes an end-to-end framework that can learn and decompose complex movements provided by a human demonstrator and generate new complex motions. Our approach analyzes the demonstration by the human expert and uses geometric criteria to decompose the observed movement into segments that are stored as dynamic movement primitives (DMPs) in a library. Then, given a new environment configuration, our system autonomously composes different motion primitives to construct an optimized trajectory that meets the constraints imposed by the new environment. Our system is therefore able to construct new DMPs to execute complex motions in environments that differ from the one where the original motions were taught. Our approach is also compatible with existing run-time obstacle avoidance approaches. We illustrate the application of our approach both in simulation and with a Baxter robot.
Ali H. Kordia, Francisco S. Melo
ICARCV2
2020 A new convergent variant of Q-learning with linear function approximation
abstract
In this work, we identify a novel set of conditions that ensure convergence with probability 1 of Q-learning with linear function approximation, by proposing a two time-scale variation thereof. In the faster time scale, the algorithm features an update similar to that of DQN, where the impact of bootstrapping is attenuated by using a Q-value estimate akin to that of the target network in DQN. The slower time-scale, in turn, can be seen as a modified target network update. We establish the convergence of our algorithm, provide an error bound and discuss our results in light of existing convergence results on reinforcement learning with function approximation. Finally, we illustrate the convergent behavior of our method in domains where standard Q-learning has previously been shown to diverge.
Diogo S. Carvalho, Francisco S. Melo, Pedro Santos 0001
NeurIPS2
2020 Optimal action sequence generation for assistive agents in fixed horizon tasks
Kim Baraka, Francisco S. Melo, Marta Couto, Manuela M. Veloso
Auton. Agents Multi Agent Syst.2
2019 Exploring Prosociality in Human-Robot Teams
abstract
This paper explores the role of prosocial behaviour when people team up with robots in a collaborative game that presents a social dilemma similar to a public goods game. An experiment was conducted with the proposed game in which each participant joined a team with a prosocial robot and a selfish robot. During 5 rounds of the game, each player chooses between contributing to the team goal (cooperate) or contributing to his individual goal (defect). The prosociality level of the robots only affects their strategies to play the game, as one always cooperates and the other always defects. We conducted a user study at the office of a large corporation with 70 participants where we manipulated the game result (winning or losing) in a between-subjects design. Results revealed two important considerations: (1) the prosocial robot was rated more positively in terms of its social attributes than the selfish robot, regardless of the game result; (2) the perception of competence, the responsibility attribution (blame/credit), and the preference for a future partner revealed significant differences only in the losing condition. These results yield important concerns for the creation of robotic partners, the understanding of group dynamics and, from a more general perspective, the promotion of a prosocial society.
Filipa Correia, Samuel Mascarenhas, Samuel Gomes, Patrícia Arriaga, Iolanda Leite, Rui Prada, Francisco S. Melo, Ana Paiva 0001
HRI7
2019 Group Intelligence on Social Robots
abstract
This PhD project aims at investigating how a social robot can adapt its behaviours to the group members in order to achieve more positive group dynamics, which we identify as group intelligence. This goal is supported by our previous work, which contains relevant data and insightful results to the understanding of group interactions between humans and robots. Finally, we examine and discuss the future work we have planned and what are the contributions to Human-Robot Interaction (HRI) field.
Filipa Correia, Francisco S. Melo, Ana Paiva 0001
HRI2
2019 Learning Multimodal Representations for Sample-efficient Recognition of Human Actions
abstract
Humans interact in rich and diverse ways with the environment. However, the representation of such behavior by artificial agents is often limited. In this work we present motion concepts, a novel multimodal representation of human actions in a household environment. A motion concept encompasses a probabilistic description of the kinematics of the action along with its contextual background, namely the location and the objects held during the performance. We introduce a novel algorithm which learns and recognizes motion concepts from action demonstrations, named Online Motion Concept Learning (OMCL). The algorithm is evaluated on a virtual-reality household environment with the presence of a human avatar. OMCL outperforms standard motion recognition algorithms on an one-shot recognition task, attesting to its potential for sample-efficient recognition of human actions.
Miguel Vasco, Francisco S. Melo, David Martins de Matos, Ana Paiva 0001, Tetsunari Inamura
IROS2
2019 Walk the Talk! Exploring (Mis)Alignment of Words and Deeds by Robotic Teammates in a Public Goods Game
abstract
This paper explores how robotic teammates can enhance and promote cooperation in collaborative settings. It presents a user study in which participants engaged with two fully autonomous robotic partners to play a game together, named “For The Record”, a variation of a public goods game. The game is played for a total of five rounds and in each of them, players face a social dilemma: to cooperate i.e., contributing towards the team's goal while compromising individual benefits, or to defect i.e., favouring individual benefits over the team's goal. Each participant collaborates with two robotic partners that adopt opposite strategies to play the game: one of them is an unconditional cooperator (the pro-social robot), and the other is an unconditional defector (the selfish robot). In a between-subjects design, we manipulated which of the two robots criticizes behaviours, which consists of condemning participants when they opt to defect, and it represents either an alignment or a misalignment of words and deeds by the robot. Two main findings should be highlighted (1) the misalignment of words and deeds may affect the level of discomfort perceived on a robotic partner; (2) the perception a human has of a robotic partner that criticizes him is not damaged as long as the robot displays an alignment of words and deeds.
Filipa Correia, Ana Paiva 0001, Shruti Chandra, Samuel Mascarenhas, Julien Charles-Nicolas, Justin Gally, Diana Lopes, Fernando P. Santos 0001, Francisco C. Santos, Francisco S. Melo
RO-MAN10
2019 Project INSIDE: towards autonomous semi-unstructured human-robot social interaction in autism therapy
Francisco S. Melo, Alberto Sardinha, David Belo, Marta Couto, Miguel Faria 0001, Anabela Farias, Hugo Gamboa, Cátia Jesus, Mithun Kinarullathil, Pedro U. Lima, Luís Luz, André Mateus 0001, Isabel Melo, Plinio Moreno, Daniel Faustino de Noronha Osório, Ana Paiva 0001, Jhielson M. Pimentel, Rodrigo M. M. Ventura
Artif. Intell. Medicine1
2019 Empathic Robot for Group Learning: A Field Study
abstract
This work explores a group learning scenario with an autonomous empathic robot. We address two research questions: (1) Can an autonomous robot designed with empathic competencies foster collaborative learning in a group context? (2) Can an empathic robot sustain positive educational outcomes in long-term collaborative learning interactions with groups of students? To answer these questions, we developed an autonomous robot with empathic competencies that is able to interact with a group of students in a learning activity about sustainable development. Two studies were conducted. The first study compares learning outcomes in children across three conditions: learning with an empathic robot; learning with a robot without empathic capabilities; and learning without a robot. The results show that the autonomous robot with empathy fosters meaningful discussions about sustainability, which is a learning outcome in sustainability education. The second study features groups of students who interact with the robot in a school classroom for 2 months. The long-term educational interaction did not seem to provide significant learning gains, although there was a change in game-actions to achieve more sustainability during game-play. This result reflects the need to perform more long-term research in the field of educational robots for group learning.
Patrícia Alves-Oliveira, Pedro Sequeira, Francisco S. Melo, Ginevra Castellano, Ana Paiva 0001
ACM Trans. Hum. Robot Interact.3
2018 Group-based Emotions in Teams of Humans and Robots
abstract
Providing social robots an internal model of emotions can help them guide their behaviour in a more humane manner by simulating the ability to feel empathy towards others. Furthermore, the growing interest in creating robots that are capable of collaborating with other humans in team settings provides an opportunity to explore another side of human emotion, namely, group-based emotions. This paper contributes with the first model on group-based emotions in social robotic partners. We defined a model of group-based emotions for social robots that allowed us to create two distinct robotic characters that express either individual or group-based emotions. This paper also contributes with a user study where two autonomous robots embedded the previous characters, and formed two human-robot teams to play a competitive game. Our results showed that participants perceived the robot that expresses group-based emotions as more likeable and attributed higher levels of group identification and group trust towards their teams, when compared to the robotic partner that expresses individual-based emotions.
Filipa Correia, Samuel Mascarenhas, Rui Prada, Francisco S. Melo, Ana Paiva 0001
HRI4
2018 Interactive Optimal Teaching with Unknown Learners
abstract
This paper introduces a new approach for machine teaching that partly addresses the (unavoidable) mismatch between what the teacher assumes about the learning process of the student and the actual process. We analyze several situations in which such mismatch takes place, including when the student?s learning algorithm is known but the corresponding parameters are not, and when the learning algorithm itself is not known. Our analysis is focused on the case of a Bayesian Gaussian learner, and we show that, even in this simple case, the lack of knowledge regarding the student?s learning process significantly deteriorates the performance of machine teaching: while perfect knowledge of the student ensures that the target is learned after a finite number of samples, lack of knowledge thereof implies that the student will only learn asymptotically (i.e., after an infinite number of samples). We introduce interactivity as a means to mitigate the impact of imperfect knowledge and show that, by using interactivity, we are able to recover finite learning time, in the best case, or significantly faster convergence, in the worst case. Finally, we discuss the extension of our analysis to a classification problem using linear discriminant analysis, and discuss the implications of our results in single- and multi-student settings.
Francisco S. Melo, Carla Guerra, Manuel Lopes 0001
IJCAI1
2017 Associate Latent Encodings in Learning from Demonstrations
abstract
We contribute a learning from demonstration approach for robots to acquire skills from multi-modal high-dimensional data. Both latent representations and associations of different modalities are proposed to be jointly learned through an adapted variational auto-encoder. The implementation and results are demonstrated in a robotic handwriting scenario, where the visual sensory input and the arm joint writing motion are learned and coupled. We show the latent representations successfully construct a task manifold for the observed sensor modalities. Moreover, the learned associations can be exploited to directly synthesize arm joint handwriting motion from an image input in an end-to-end manner. The advantages of learning associative latent encodings are further highlighted with the examples of inferring upon incomplete input images. A comparison with alternative methods demonstrates the superiority of the present approach in these challenging tasks.
Hang Yin 0001, Francisco S. Melo, Aude Billard, Ana Paiva 0001
AAAI2
2017 Learning and Teaching Biodiversity Through a Storyteller Robot
Maria José Ferreira, Valentina Nisi, Francisco S. Melo, Ana Paiva 0001
ICIDS3
2017 "Me and you together" movement impact in multi-user collaboration tasks
abstract
This paper presents a study on collaborative manipulation between an autonomous robot and multiple users. We investigate how different motion types impact people's ability to understand the robot's goals in a multi-user scenario. We propose an approach based on Collaborative Probabilistic Movement Primitives to generate the robot's movements, exploiting predictability and legibility of movement to express intentions through motion. We compare the impact on the interaction of using only either predictable or legible movements, and propose a third approach - hybrid motion - that selects, in each situation, whether to execute a predictable motion or a legible motion, depending on what the robot perceives as more efficient for the multi-user collaboration effort. To test the impact of the three motion types in the context of a collaborative task, we run a user study using a Baxter robot that autonomously serves cups of water to three users upon request. Our results show that, in the particular case where all users simultaneously request water, the hybrid motion performs better than the other two.
Miguel Faria 0001, Patrícia Alves-Oliveira, Francisco S. Melo, Ana Paiva 0001
IROS4
2017 Adaptive indirect control through communication in collaborative human-robot interaction
abstract
This paper addresses the problem of human-robot collaboration in scenarios where a robot assists a human by executing a complex motion involving the manipulation of an object. We focus on tasks in which success in the task depends on reaching a target pose that is controlled by the human. We contribute a reinforcement learning-based approach that allows the robot to reason about its own ability to successfully complete the task given the current target pose and indirectly adjust that pose by prompting the human user. Our approach allows the robot both to trade-off the benefits of adjusting the target position against the cost of bothering the human user while, at the same time, adapting to each user's responses. Our approach was tested in a real-world human-robot collaboration scenario involving the Baxter robot.
Miguel Faria 0001, Francisco S. Melo, Manuela M. Veloso
IROS3
2016 Adaptive Symbiotic Collaboration for Targeted Complex Manipulation Tasks
abstract
This paper addresses the problem of human-robot collaboration in the context of manipulation tasks. In particular, we focus on tasks where a robot must perform some complex manipulation that is successfully completed only upon reaching some target pose provided by a human user. We propose an approach in which the robot explicitly reasons about its ability to complete the task and proactively requests the assistance of the human teammate when necessary. Our approach effectively trades-off the benefits arising from the human assistance with the cost of disturbing the user. We also propose an adaptation mechanism that enables the robot to adjust its behavior to the particular manner by which the human user responds to the requests made by the robot. We test our approach in a simple illustrative scenario and in two real interaction scenarios involving the Baxter robot.
Francisco S. Melo, Manuela M. Veloso
ECAI2
2016 Building a Social Robot as a Game Companion in a Card Game
abstract
In this video we present a social robotic player that is able to play a traditional card game in a social manner. The interaction takes place in a rich environment in which two teams of two players each compete to win the card game. Therefore, the robotic game player has a partner, and an opponent team of two other players. During each game, the robot explores both competitiveness with the opponent team and cooperation with its partner, conciliating the performance of players and the social dynamics that emerge during the game-play.
Filipa Correia, Tiago Ribeiro 0001, Patrícia Alves-Oliveira, Francisco S. Melo, Ana Paiva 0001
HRI4
2016 Discovering Social Interaction Strategies for Robots from Restricted-Perception Wizard-of-Oz Studies
abstract
In this paper we propose a methodology for the creation of social interaction strategies for human-robot interaction based on restricted-perception Wizard-of-Oz studies (WoZ). This novel experimental technique involves restricting the wizard's perceptions over the environment and the behaviors it controls according to the robot's inherent perceptual and acting limitations. Within our methodology, the robot's design lifecycle is divided into three consecutive phases, namely data collection, where we perform interaction studies to extract expert knowledge and interaction data; strategy extraction, where a hybrid strategy controller for the robot is learned based on the gathered data; strategy refinement, where the controller is iteratively evaluated and adjusted. We developed a fully-autonomous robotic tutor based on the proposed approach in the context of a collaborative learning scenario. The results of the evaluation study show that, by performing restricted-perception WoZ studies, our robots are able to engage in very natural and socially-aware interactions.
Pedro Sequeira, Patrícia Alves-Oliveira, Tiago Ribeiro 0001, Eugenio Di Tullio, Sofia Petisca, Francisco S. Melo, Ginevra Castellano, Ana Paiva 0001
HRI6
2016 Synthesizing Robotic Handwriting Motion by Learning from Human Demonstrations
Hang Yin 0001, Patrícia Alves-Oliveira, Francisco S. Melo, Aude Billard, Ana Paiva 0001
IJCAI3
2016 An Interactive Tangram Game for Children with Autism
Beatriz Bernardo, Patrícia Alves-Oliveira, Maria Graça Santos, Francisco S. Melo, Ana Paiva 0001
IVA4
2016 Just follow the suit! Trust in human-robot interactions during card game playing
abstract
Robots are currently being developed to enter our lives and interact with us in different tasks. For humans to be able to have a positive experience of interaction with such robots, they need to trust them to some degree. In this paper, we present the development and evaluation of a social robot that was created to play a card game with humans, playing the role of a partner and opponent. This type of activity is especially important, since our target group is elderly people - a population that often suffers from social isolation. Moreover, the card game scenario can lead to the development of interesting trust dynamics during the interaction, in which the human that partners with the robot needs to trust it in order to succeed and win the game. The design of the robot's behavior and game dynamics was inspired in previous user-centered design studies in which elderly people played the same game. Our evaluation results show that the levels of trust differ according to the previous knowledge that players have of their partners. Thus, humans seem to significantly increase their trust level towards a robot they already know, whilst maintaining the same level of trust in a human that they also previously knew. Henceforth, this paper shows that trust is a multifaceted construct that develops differently for humans and robots.
Filipa Correia, Patrícia Alves-Oliveira, Nuno Maia, Tiago Ribeiro 0001, Sofia Petisca, Francisco S. Melo, Ana Paiva 0001
RO-MAN6
2016 Ad hoc teamwork by learning teammates' task
Francisco S. Melo, Alberto Sardinha
Auton. Agents Multi Agent Syst.1
2015 Towards table tennis with a quadrotor autonomous learning robot and onboard vision
abstract
Robot table tennis is a challenging domain in both robotics, artificial intelligence and machine learning. In terms of robotics, it requires fast and reliable perception and control; in terms of artificial intelligence, it requires fast decision making to determine the best motion to hit the ball; in terms of machine learning, it requires the ability to accurately estimate where and when the ball will be so that it can be hit. The use of sophisticated perception (relying, for example, in multi-camera vision systems) and state-of-the-art robot manipulators significantly alleviates concerns with perception and control, leaving room for the exploration of novel approaches that focus on estimating where, when and how to hit the ball. In this paper, we move away from the hardware setup commonly used in this domain-typically relying on robotic manipulators combined with an array of multiple fixed cameras-and give the first steps towards having autonomous aerial table tennis robotic players. Specifically, we focus on the task of hitting a ping pong ball thrown at a commercial drone, equipped with a light cardboard racket and an onboard camera. We adopt a general framework for learning complex robot tasks and show that, in spite of the perceptual and actuation limitations of our system, the overall approach enables the quadrotor system to successfully respond to balls served by a human user.
Francisco S. Melo, Manuela M. Veloso
IROS2
2015 Emergence of emotional appraisal signals in reinforcement learning agents
Pedro Sequeira, Francisco S. Melo, Ana Paiva 0001
Auton. Agents Multi Agent Syst.2
2013 Towards agents with human-like decisions under uncertainty
Francisco S. Melo, Samuel Mascarenhas, João Dias 0001, Rui Prada, Ana Paiva 0001
CogSci2
2011 Differential Eligibility Vectors for Advantage Updating and Gradient Methods
Francisco S. Melo
AAAI1
2011 Emotion-Based Intrinsic Motivation for Reinforcement Learning Agents
Pedro Sequeira, Francisco S. Melo, Ana Paiva 0001
ACII (1)2
2011 Decentralized MDPs with sparse interactions
Francisco S. Melo, Manuela M. Veloso
Artif. Intell.1
2010 Analysis of Inverse Reinforcement Learning with Perturbed Demonstrations
Francisco S. Melo, Manuel Lopes 0001, Ricardo Ferreira 0002
ECAI1
2010 Learning from Demonstration Using MDP Induced Metrics
Francisco S. Melo, Manuel Lopes 0001
ECML/PKDD (2)1
2010 Coordinated learning in multiagent MDPs with infinite state-space
Francisco S. Melo, M. Isabel Ribeiro
Auton. Agents Multi Agent Syst.1
2009 Active Learning for Reward Estimation in Inverse Reinforcement Learning
Manuel Lopes 0001, Francisco S. Melo, Luis Montesano
ECML/PKDD (2)2
2008 Exploiting locality of interactions using a policy-gradient approach in multiagent learning
abstract
In this paper, we propose a policy gradient reinforcement learning algorithm to address transition-independent Dec-POMDPs. This approach aims at implicitly exploiting the locality of interaction observed in many practical problems. Our algorithms can be described by an actor-critic architecture: the actor component combines natural gradient updates with a varying learning rate; the critic uses only local information to maintain a belief over the joint state-space, and evaluates the current policy as a function of this belief using compatible function approximation. In order to speed the convergence of the algorithm, we use an optimistic initialization of the policy that relies on a fully observable, single agent model of the problem. We illustrate our approach in some simple application problems.
Francisco S. Melo
ECAI1
2008 An analysis of reinforcement learning with function approximation
abstract
We address the problem of computing the optimal Q-function in Markov decision problems with infinite state-space. We analyze the convergence properties of several variations of Q-learning when combined with function approximation, extending the analysis of TD-learning in (Tsitsiklis & Van Roy, 1996a) to stochastic control settings. We identify conditions under which such approximate methods converge with probability 1. We conclude with a brief discussion on the general applicability of our results and compare them with several related works.
Francisco S. Melo, Sean P. Meyn, M. Isabel Ribeiro
ICML1
2008 Reinforcement learning with function approximation for cooperative navigation tasks
abstract
In this paper, we propose a reinforcement learning approach to address multi-robot cooperative navigation tasks in infinite settings. We propose an algorithm to simultaneously address the problems of learning and coordination in multi-robot problems. The proposed algorithm extends those existing in the literature, allowing to address simultaneous learning and coordination in problems with an infinite state-space. We also present the results obtained in several test scenarios featuring multi-robot navigation situations with partial observability.
Francisco S. Melo, M. Isabel Ribeiro
ICRA1
2008 Fitted Natural Actor-Critic: A New Algorithm for Continuous State-Action MDPs
Francisco S. Melo, Manuel Lopes 0001
ECML/PKDD (2)1
2007 Q -Learning with Linear Function Approximation
Francisco S. Melo, M. Isabel Ribeiro
COLT1
2007 Affordance-based imitation learning in robots
abstract
In this paper we build an imitation learning algorithm for a humanoid robot on top of a general world model provided by learned object affordances. We consider that the robot has previously learned a task independent affordance-based model of its interaction with the world. This model is used to recognize the demonstration by another agent (a human) and infer the task to be learned. We discuss several important problems that arise in this combined framework, such as the influence of an inaccurate model in the recognition of the demonstration. We illustrate the ideas in the paper with some experimental results obtained with a real robot.
Manuel Lopes 0001, Francisco S. Melo, Luis Montesano
IROS2