VLDB 2026 Research / reviewers in the wild / expert
Francisco Cruz 0002
dblp:14/5509 · also Francisco Cruz Naranjo
· DBLP profile ↗
27ranked-venue papers
9as first author
18since 2021 · last 2026
0000-0002-1131-3382ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 25 · 9 first-author · 17 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 since 2021Systems, architecture and hardware · 2 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MHDPose: Multi-hypothesis 3D human pose estimation using bidirectional Mamba diffusion modelsabstractMonocular 3D human pose estimation often faces challenges due to depth ambiguities, unstable predictions, and occlusions. Diffusion models have emerged as a generative framework that can transform noise into complex data representations. Although diffusion-based multi-hypothesis approaches have been explored in prior works, their effectiveness largely depends on the denoising network, which captures spatial and long-range temporal dependencies. In this paper, we propose MHDPose, a conditional diffusion-based human pose estimation framework that can generate multiple 3D pose predictions with the guidance of a single 2D pose. The pose hypotheses are generated by the proposed PMamba denoiser that consists of: (i) a spatial transformer over joints with proposed kinematic-aware rotary position embeddings (Kin-RoPE) to encode the skeleton’s tree structure, and (ii) bi-directional Mamba blocks for efficient long-range temporal modeling, and a residual Mamba head for pose refinement. Our proposed MHDPose achieves competitive results on widely used pose estimation benchmarks such as Human3.6M and MPI-INF-3DHP. Marsha Mariya Kappan, Eduardo Benítez Sandoval, Erik Meijering, Francisco Cruz 0002 |
Pattern Recognit. | 4 |
| 2025 | MERCI: A Multimodal Dataset for Personalised and Emotionally-Aware DialoguesabstractThe integration of conversational agents into daily life has become increasingly common. However, sustaining deeply engaging and natural interactions remains challenging due to a lack of multimodal datasets capturing personal and emotional nuances. In this paper, we introduce MERCI (Multimodal dataset for Emotionally-aware peRsonalised Conversational In-teractions), a dataset derived from user-robot dialogues involving thirty participants who completed user profile questionnaires covering ten personal topics (e.g., hobbies, music). A conver-sational system called PERCY then engaged with each partici-pant in open-domain conversations, leveraging GPT-4, real-time facial-expression and sentiment analysis to generate contextu-ally appropriate, empathetic responses. MERCI contains 1860 utterances, equating to about 12.5 hours of aligned audio, three-view video, transcripts with timestamps, emotion labels, and sentiment scores. This dataset serves as a reproducible test-bed for tasks such as emotion-aware response generation, multimodal affect recognition, and personalised policy learning. Baseline performance results have been established using advanced models such as BERT, T5, BART, and GPT-3.5/4/4o-mini across gener-ation, regression, and classification. Evaluations through human and automated methods have demonstrated strong naturalness, relevance, and consistency in responses while indicating areas for enhanced personalisation and empathic depth. We expect that MERCI will enhance the development of emotionally intelligent, user-centric conversational AI applications, potentially ranging from social robotics to mental health support. Mohammed Althubyani, Zhijin Meng, Shengyuan Xie, Francisco Cruz 0002, Muhammad Imran Razzak, Mukesh Prasad, Eduardo Benítez Sandoval, Ahmet Baki Kocaballi |
CBMI | 4 |
| 2025 | Help me to Clean: Studying Multi-trainers on Interactive Reinforcement Learning in Domestic ScenariosabstractIn the last decade, robots have increasingly participated in human-centered environments, particularly performing tasks in domestic settings. Due to the complexity of these environments, reinforcement learning (RL) has become a key approach to train agents autonomously. A variation, interactive reinforcement learning (IRL), incorporates feedback from an external trainer to guide the agent toward the optimal policy. However, trainers are often non-experts and may provide incorrect feedback, potentially hindering the learning process. To address this, recent approaches like Multi-Trainer Interactive Reinforcement Learning (MTIRL) and ensemble strategies such as Bayesian Weighted Voting Ensemble (BWVE) have been proposed to mitigate the impact of erroneous guidance. In this work, we present a novel application of MTIRL in a simulated domestic environment, where a robot must learn a household navigation task from multiple non-expert trainers. Our main contribution is the implementation and evaluation of BWVE within this context, allowing the agent to dynamically assess the reliability of each trainer’s feedback. We compare this approach with a Full Feedback Multi-Trainer (FFMT) baseline, where all feedback is processed equally. To support our experiments, we developed a custom simulation environment, Domestic Scenario with a Cognitive Agent, tailored for domestic navigation task. Evaluation metrics include feedback frequency, steps per episode, and time to complete 100 episodes. Results demonstrate that BWVE leads to more efficient and robust learning, requiring fewer trainers and tolerating occasional incorrect feedback more effectively than the baseline. Cristian Millán-Arias, Ernesto Gutiérrez, Francisco Cruz 0002 |
CLEI | 3 |
| 2025 | HRIxD: End-Effector Specialised in Drumming for Performing RobotsabstractHuman-robot interaction encompasses various applications, from simple tasks to advanced complex functionalities. This research introduces a specialised robotics end effector specially modelled with auditory processing for drumming along-side humans. This system interprets auditory vocal input to generate precise and adaptive drumming patterns. Utilising precision motor control and real-time signal processing to replicate drumming with human-like characteristics with precision. This work demonstrates collaboration between performing arts and robotics, demonstrating the potential of intelligent systems to interact with human creativity in real-time. The specialised end effector utilises precise motor control with motor drivers and a belt mechanism to regulate the speed and direction of the drumstick, enabling rapid, high-speed strokes to strike the drum 4 to 5 times per second. By integrating audio perception with mechanical execution, this research advances the state of the art and establishes a new framework for incorporating robotics into the creative and artistic domain. Khaja Ahmed Shaik, Eduardo Benítez Sandoval, Shengyuan Xie, Francisco Cruz 0002 |
HRI | 4 |
| 2025 | A Fuzzy Supervisory Framework for Real-Time Optimization of Robot Output and LLM Performance in HRIabstractHuman-robot interaction plays a vital role in pushing the capabilities of socially interactive robots by enabling them to deliver content with high emotional intelligence. This research focuses on a supervisory fuzzy framework for constantly evaluating and improving the content delivered by the robot utilizing multimodal inputs and advanced intelligent algorithms. The main reason for using fuzzy logic is that it mimics human decision-making by providing a percentage-based measure of closeness. In this project, ARI Robot is being used with an LLM integration, which enables the user to communicate with the robot. Different algorithms were integrated for the classification of multimodal inputs, BERT (Bidirectional Encoder Representations from Transformers) for the classification of content, Wav2Vec 2.0 for classifying the tone of the user while interacting with the robot, and OpenFace for classifying the facial expression of the user. All of these inputs are then supervised by a fuzzy system with predefined rules to evaluate the content delivered and provide feedback for refinement. The proposed framework ensures an overall evaluation of content delivery, providing intelligent feedback to the ARI robot to improve interaction quality. By integrating these advanced models with fuzzy logic, the system mimics human-like judgment in assessing the interaction of verbal and non-verbal indications, making the way for emotionally intelligent robots in a social world. Khaja Ahmed Shaik, Shengyuan Xie, Francisco Cruz 0002, Eduardo Benítez Sandoval |
HRI | 3 |
| 2025 | Embodied Generative AI Art for Enhanced Human-Robot Interaction Through a Human-Centric LLM-Guided Robotic Arm Drawing SystemabstractGenerative AI is transforming the way humans interact with robots by integrating language-driven comprehension with embodied execution. While recent research leveraging large language models (LLMs) to enhance communication between humans and machines has shown significant progress, the exploration of integrating LLMs with robots to assist users in creating real-world artistic works remains challenging. This research introduces a novel GENAI-driven intelligent robotic arm drawing system. We fine-tuned GPT-3.5 Turbo to better capture meaningful information from conversations with users and output precise drawing commands, which are directly fed into the fine-tuned Stable Diffusion model to generate desired images. A UFactory xArm equipped with an end-effector holding a drawing pen is deployed and controlled by a speed control algorithm to perform accurate drawing movements based on AI-generated images. Our proposed framework facilitates enhanced HRI experiences with a focus on drawing tasks, effectively embodying the capabilities of generative AI in the field of artistic creation and enabling users to engage deeply in the process of using AI for art creation in real-world environments. Shengyuan Xie, Eduardo Benítez Sandoval, Khaja Ahmed Shaik, Francisco Cruz 0002 |
HRI | 4 |
| 2025 | Understanding User Preferences in Explainable Artificial Intelligence: A Mapping Function ProposalabstractThe increasing complexity of AI systems has led to the growth of the field of Explainable AI (XAI), which aims to provide explanations and justifications for the outputs of AI algorithms. While there is considerable demand for XAI, there remains a scarcity of studies aimed at comprehensively understanding the practical distinctions among different methods and effectively aligning each method with users’ individual needs, and ideally, offer a mapping function which can map each user with its specific needs to a method of explainability. This study endeavors to bridge this gap by conducting a review of the relevant works in XAI, with a specific focus on Explainable Machine Learning (XML), and a keen eye on user needs to provide an observational study. Our main objective is to offer a classification of XAI methods within the realm of XML, categorizing current works into three distinct domains: philosophy, theory, and practice. Moreover, our study seeks to facilitate the connection between XAI users and the most suitable methods for them and tailor explanations to meet their specific needs by proposing a mapping function that take to account users and their desired properties and suggest an XAI method to them. This entails an examination of prevalent XAI approaches and an evaluation of their properties. The primary outcome of this study is the formulation of a clear and concise strategy for selecting the optimal XAI method to achieve a given goal, all while delivering personalized explanations tailored to individual users. Maryam Hashemi, Ali Darejeh, Francisco Cruz 0002 |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2024 | Elastic step DQN: A novel multi-step algorithm to alleviate overestimation in Deep Q-NetworksabstractDeep Q-Networks algorithm (DQN) was the first reinforcement learning algorithm using deep neural network to successfully surpass human level performance in a number of Atari learning environments. However, divergent and unstable behaviour have been long standing issues in DQNs. The unstable behaviour is often characterised by overestimation in the Q-values, commonly referred to as the overestimation bias. To address the overestimation bias and the divergent behaviour, a number of heuristic extensions have been proposed. Notably, multi-step updates have been shown to drastically reduce unstable behaviour while improving agent’s training performance. However, agents are often highly sensitive to the selection of the multi-step update horizon (n), and our empirical experiments show that a poorly chosen static value for n can in many cases lead to worse performance than single-step DQN. Inspired by the success of n-step DQN and the effects that multi-step updates have on overestimation bias, this paper proposes a new algorithm that we call ‘Elastic Step DQN’ (ES-DQN) to alleviate overestimation bias in DQNs. ES-DQN dynamically varies the step size horizon in multi-step updates based on the similarity between states visited. Our empirical evaluation shows that ES-DQN out-performs n-step with fixed n updates, Double DQN and Average DQN in several OpenAI Gym environments while at the same time alleviating the overestimation bias. Adrian Ly, Richard Dazeley, Peter Vamplew 0001, Francisco Cruz 0002, Sunil Aryal |
Neurocomputing | 4 |
| 2023 | Elastic step DDPG: Multi-step reinforcement learning for improved sample efficiencyabstractA major challenge in deep reinforcement learning is that it requires more data to converge to an policy for complex problems. One way to improve sample efficiency is to use n-step updates to reduce the number of samples required to converge to a good policy. However n-step updates are known to be brittle and difficult to tune. Elastic Step DQN has shown that it is possible to automate the value of$n$in DQN to solve problems involving discrete action spaces, however the efficacy of the technique when applied on more complex problems and against problems with continuous action spaces is yet to be shown. In this paper we adapt the innovations proposed by Elastic Step DQN onto the DDPG algorithm and show empirically that Elastic Step DDPG is able to achieve a much stronger final training policy and is more sample efficient than DDPG. Adrian Ly, Richard Dazeley, Peter Vamplew 0001, Francisco Cruz 0002, Sunil Aryal |
IJCNN | 4 |
| 2023 | Human engagement providing evaluative and informative advice for interactive reinforcement learningabstractAbstract Interactive reinforcement learning proposes the use of externally sourced information in order to speed up the learning process. When interacting with a learner agent, humans may provide either evaluative or informative advice. Prior research has focused on the effect of human-sourced advice by including real-time feedback on the interactive reinforcement learning process, specifically aiming to improve the learning speed of the agent, while minimising the time demands on the human. This work focuses on answering which of two approaches, evaluative or informative, is the preferred instructional approach for humans. Moreover, this work presents an experimental setup for a human trial designed to compare the methods people use to deliver advice in terms of human engagement. The results obtained show that users giving informative advice to the learner agents provide more accurate advice, are willing to assist the learner agent for a longer time, and provide more advice per episode. Additionally, self-evaluation from participants using the informative approach has indicated that the agent’s ability to follow the advice is higher, and therefore, they feel their own advice to be of higher accuracy when compared to people providing evaluative advice. Adam Bignold, Francisco Cruz 0002, Richard Dazeley, Peter Vamplew 0001, Cameron Foale |
Neural Comput. Appl. | 2 |
| 2023 | Persistent rule-based interactive reinforcement learning
Adam Bignold, Francisco Cruz 0002, Richard Dazeley, Peter Vamplew 0001, Cameron Foale |
Neural Comput. Appl. | 2 |
| 2023 | Explainable robotic systems: understanding goal-driven actions in a reinforcement learning scenario
Francisco Cruz 0002, Richard Dazeley, Peter Vamplew 0001, Ithan Moreira |
Neural Comput. Appl. | 1 |
| 2023 | Human-aligned reinforcement learning for autonomous agents and robots
Francisco Cruz 0002, Thommen George Karimpanal, Miguel A. Solis, Pablo V. A. Barros, Richard Dazeley |
Neural Comput. Appl. | 1 |
| 2023 | Explainable reinforcement learning for broad-XAI: a conceptual framework and surveyabstractAbstract Broad-XAI moves away from interpreting individual decisions based on a single datum and aims to provide integrated explanations from multiple machine learning algorithms into a coherent explanation of an agent’s behaviour that is aligned to the communication needs of the explainee. Reinforcement Learning (RL) methods, we propose, provide a potential backbone for the cognitive model required for the development of Broad-XAI. RL represents a suite of approaches that have had increasing success in solving a range of sequential decision-making problems. However, these algorithms operate as black-box problem solvers, where they obfuscate their decision-making policy through a complex array of values and functions. EXplainable RL (XRL) aims to develop techniques to extract concepts from the agent’s: perception of the environment; intrinsic/extrinsic motivations/beliefs; Q-values, goals and objectives. This paper aims to introduce the Causal XRL Framework (CXF), that unifies the current XRL research and uses RL as a backbone to the development of Broad-XAI. CXF is designed to incorporate many standard RL extensions and integrated with external ontologies and communication facilities so that the agent can answer questions that explain outcomes its decisions. This paper aims to: establish XRL as a distinct branch of XAI; introduce a conceptual framework for XRL; review existing approaches explaining agent behaviour; and identify opportunities for future research. Finally, this paper discusses how additional information can be extracted and ultimately integrated into models of communication, facilitating the development of Broad-XAI. Richard Dazeley, Peter Vamplew 0001, Francisco Cruz 0002 |
Neural Comput. Appl. | 3 |
| 2023 | AI apology: interactive multi-objective reinforcement learning for human-aligned AIabstractAbstract For an Artificially Intelligent (AI) system to maintain alignment between human desires and its behaviour, it is important that the AI account for human preferences. This paper proposes and empirically evaluates the first approach to aligning agent behaviour to human preference via an apologetic framework. In practice, an apology may consist of an acknowledgement, an explanation and an intention for the improvement of future behaviour. We propose that such an apology, provided in response to recognition of undesirable behaviour, is one way in which an AI agent may both be transparent and trustworthy to a human user. Furthermore, that behavioural adaptation as part of apology is a viable approach to correct against undesirable behaviours. The Act-Assess-Apologise framework potentially could address both the practical and social needs of a human user, to recognise and make reparations against prior undesirable behaviour and adjust for the future. Applied to a dual-auxiliary impact minimisation problem, the apologetic agent had a near perfect determination and apology provision accuracy in several non-trivial configurations. The agent subsequently demonstrated behaviour alignment with success that included up to complete avoidance of the impacts described by these objectives in some scenarios. Hadassah Harland, Richard Dazeley, Bahareh Nakisa, Francisco Cruz 0002, Peter Vamplew 0001 |
Neural Comput. Appl. | 4 |
| 2023 | Proxemic behavior in navigation tasks using reinforcement learningabstractAbstract Human interaction starts with a person approaching another one, respecting their personal space to prevent uncomfortable feelings. Spatial behavior, called proxemics, allows defining an acceptable distance so that the interaction process begins appropriately. In recent decades, human-agent interaction has been an area of interest for researchers, where it is proposed that artificial agents naturally interact with people. Thus, new alternatives are needed to allow optimal communication, avoiding humans feeling uncomfortable. Several works consider proxemic behavior with cognitive agents, where human-robot interaction techniques and machine learning are implemented. However, it is assumed that the personal space is fixed and known in advance, and the agent is only expected to make an optimal trajectory toward the person. In this work, we focus on studying the behavior of a reinforcement learning agent in a proxemic-based environment. Experiments were carried out implementing a grid-world problem and a continuous simulated robotic approaching environment. These environments assume that there is an issuer agent that provides non-conformity information. Our results suggest that the agent can identify regions where the issuer feels uncomfortable and find the best path to approach the issuer. The results obtained highlight the usefulness of reinforcement learning in order to identify proxemic regions. Cristian Millán-Arias, Bruno J. T. Fernandes, Francisco Cruz 0002 |
Neural Comput. Appl. | 3 |
| 2022 | Evaluating Human-like Explanations for Robot Actions in Reinforcement Learning ScenariosabstractExplainable artificial intelligence is a research field that tries to provide more transparency for autonomous intelligent systems. Explainability has been used, particularly in reinforcement learning and robotic scenarios, to better understand the robot decision-making process. Previous work, however, has been widely focused on providing technical explanations that can be better understood by AI practitioners than non-expert end-users. In this work, we make use of human-like explanations built from the probability of success to complete the goal that an autonomous robot shows after performing an action. These explanations are intended to be understood by people who have no or very little experience with artificial intelligence methods. This paper presents a user trial to study whether these explanations that focus on the probability an action has of succeeding in its goal constitute a suitable explanation for non-expert end-users. The results obtained show that non-expert participants rate robot explanations that focus on the probability of success higher and with less variance than technical explanations generated from Q-values, and also favor counterfactual explanations over standalone explanations. Francisco Cruz 0002, Charlotte Young, Richard Dazeley, Peter Vamplew 0001 |
IROS | 1 |
| 2021 | Levels of explainable artificial intelligence for human-aligned conversational explanations
Richard Dazeley, Peter Vamplew 0001, Cameron Foale, Charlotte Young, Sunil Aryal, Francisco Cruz 0002 |
Artif. Intell. | 6 |
| 2020 | A Robust Approach for Continuous Interactive Reinforcement LearningabstractInteractive reinforcement learning is an approach in which an external trainer helps an agent to learn through advice. A trainer is useful in large or continuous scenarios; however, when the characteristics of the environment change over time, it can affect the learning. Robust reinforcement learning is a reliable approach that allows an agent to learn a task, regardless of disturbances in the environment. In this work, we present an approach that addresses interactive reinforcement learning problems in a dynamic environment with continuous states and actions. Our results show that the proposed approach allows an agent to complete the cart-pole balancing task satisfactorily in a dynamic, continuous action-state domain. Cristian Millán-Arias, Bruno J. T. Fernandes, Francisco Cruz 0002, Richard Dazeley, Sérgio Murilo Maciel Fernandes |
HAI | 3 |
| 2020 | KutralNet: A Portable Deep Learning Model for Fire RecognitionabstractMost of the automatic fire alarm systems detect the fire presence through sensors like thermal, smoke, or flame. One of the new approaches to the problem is the use of images to perform the detection. The image approach is promising since it does not need specific sensors and can be easily embedded in different devices. However, besides the high performance, the computational cost of the used deep learning methods is a challenge to their deployment in portable devices. In this work, we propose a new deep learning architecture that requires fewer floating-point operations (flops) for fire recognition. Additionally, we propose a portable approach for fire recognition and the use of modern techniques such as inverted residual block, convolutions like depth-wise, and octave, to reduce the model's computational cost. The experiments show that our model keeps high accuracy while substantially reducing the number of parameters and flops. One of our models presents 71% fewer parameters than FireNet, while still presenting competitive accuracy and AUROC performance. The proposed methods are evaluated on FireNet and FiSmo datasets. The obtained results are promising for the implementation of the model in a mobile device, considering the reduced number of flops and parameters acquired. Angel Ayala, Bruno J. T. Fernandes, Francisco Cruz 0002, David Macedo, Adriano Lorena Inácio de Oliveira, Cleber Zanchettin |
IJCNN | 3 |
| 2019 | Human feedback in continuous actor-critic reinforcement learning
Cristian Millán-Arias, Bruno J. T. Fernandes, Francisco Cruz 0002 |
ESANN | 3 |
| 2018 | Multi-modal Feedback for Affordance-driven Interactive Reinforcement LearningabstractInteractive reinforcement learning (IRL) extends traditional reinforcement learning (RL) by allowing an agent to interact with parent-like trainers during a task. In this paper, we present an IRL approach using dynamic audio-visual input in terms of vocal commands and hand gestures as feedback. Our architecture integrates multi-modal information to provide robust commands from multiple sensory cues along with a confidence value indicating the trustworthiness of the feedback. The integration process also considers the case in which the two modalities convey incongruent information. Additionally, we modulate the influence of sensory-driven feedback in the IRL task using goal-oriented knowledge in terms of contextual affordances. We implement a neural network architecture to predict the effect of performed actions with different objects to avoid failed-states, i.e., states from which it is not possible to accomplish the task. In our experimental setup, we explore the interplay of multi-modal feedback and task-specific affordances in a robot cleaning scenario. We compare the learning performance of the agent under four different conditions: traditional RL, multi-modal IRL, and each of these two setups with the use of contextual affordances. Our experiments show that the best performance is obtained by using audio-visual feedback with affordance-modulated IRL. The obtained results demonstrate the importance of multi-modal sensory processing integrated with goal-oriented knowledge in IRL tasks. Francisco Cruz 0002, German Ignacio Parisi, Stefan Wermter |
IJCNN | 1 |
| 2018 | Improving interactive reinforcement learning: What makes a good teacher?abstractInteractive reinforcement learning (IRL) has become an important apprenticeship approach to speed up convergence in classic reinforcement learning (RL) problems. In this regard, a variant of IRL is policy shaping which uses a parent-like trainer to propose the next action to be performed and by doing so reduces the search space by advice. On some occasions, the trainer may be another artificial agent which in turn was trained using RL methods to afterward becoming an advisor for other learner-agents. In this work, we analyse internal representations and characteristics of artificial agents to determine which agent may outperform others to become a better trainer-agent. Using a polymath agent, as compared to a specialist agent, an advisor leads to a larger reward and faster convergence of the reward signal and also to a more stable behaviour in terms of the state visit frequency of the learner-agents. Moreover, we analyse system interaction parameters in order to determine how influential they are in the apprenticeship process, where the consistency of feedback is much more relevant when dealing with different learner obedience parameters. Francisco Cruz 0002, Sven Magg, Yukie Nagai, Stefan Wermter |
Connect. Sci. | 1 |
| 2016 | Learning contextual affordances with an associative neural architecture
Francisco Cruz 0002, German Ignacio Parisi, Stefan Wermter |
ESANN | 1 |
| 2016 | Multi-modal integration of dynamic audiovisual patterns for an interactive reinforcement learning scenarioabstractRobots in domestic environments are receiving more attention, especially in scenarios where they should interact with parent-like trainers for dynamically acquiring and refining knowledge. A prominent paradigm for dynamically learning new tasks has been reinforcement learning. However, due to excessive time needed for the learning process, a promising extension has been made by incorporating an external parent-like trainer into the learning cycle in order to scaffold and speed up the apprenticeship using advice about what actions should be performed for achieving a goal. In interactive reinforcement learning, different uni-modal control interfaces have been proposed that are often quite limited and do not take into account multiple sensor modalities. In this paper, we propose the integration of audiovisual patterns to provide advice to the agent using multi-modal information. In our approach, advice can be given using either speech, gestures, or a combination of both. We introduce a neural network-based approach to integrate multi-modal information from uni-modal modules based on their confidence. Results show that multi-modal integration leads to a better performance of interactive reinforcement learning with the robot being able to learn faster with greater rewards compared to uni-modal scenarios. Francisco Cruz 0002, German Ignacio Parisi, Johannes Twiefel, Stefan Wermter |
IROS | 1 |
| 2015 | Interactive reinforcement learning through speech guidance in a domestic scenarioabstractRecently robots are being used more frequently as assistants in domestic scenarios. In this context we train an apprentice robot to perform a cleaning task using interactive reinforcement learning since it has been shown to be an efficient learning approach benefiting from human expertise for performing domestic tasks. The robotic agent obtains interactive feedback via a speech recognition system which is tested to work with five different microphones concerning their polar patterns and distance to the teacher to recognize sentences in different instruction classes. Moreover, the reinforcement learning approach uses situated affordances to allow the robot to complete the cleaning task in every episode anticipating when chosen actions are possible to be performed. Situated affordances and interaction allow to improve the convergence speed of reinforcement learning, and the results also show that the system is robust against wrong instructions that result from errors of the speech recognition system. Francisco Cruz 0002, Johannes Twiefel, Sven Magg, Cornelius Weber, Stefan Wermter |
IJCNN | 1 |
| 2007 | Indirect Training of Grey-Box Models: Application to a Bioprocess
Francisco Cruz 0002, Gonzalo Acuña, Francisco A. Cubillos, Vicente Moreno, Danilo Bassi |
ISNN (2) | 1 |