Stefan Wermter

dblp:03/3914 · DBLP profile ↗
← Back
282ranked-venue papers
16as first author
76since 2021 · last 2026
0000-0003-1343-4775ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 264 · 15 first-author · 66 since 2021Graphics, computer vision, multimedia, augmented reality and games · 27 · 4 first-author · 13 since 2021Systems, architecture and hardware · 22 · 7 since 2021Human-computer interaction and ubiquitous computing · 20 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 4 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author
YearPublicationVenuePosition
2026 Towards Learning a Generalizable 3D Scene Representation from 2D Observations
Martin Gromniak, Jan-Gerrit Habekost, Sebastian Kamp, Sven Magg, Stefan Wermter
ESANN5
2026 Prompt-to-Gesture: Measuring the Capabilities of Image-to-Video Deictic Gesture Generation
Hassan Ali 0005, Doreen Jirak, Luca Müller, Stefan Wermter
FG4
2026 RankCut: A Ranking-Based LLM Approach to Extractive Summarization for Transcript-Based Video Editing
abstract
Video recordings of interviews, lectures, and meetings contain valuable moments surrounded by less essential talk. Making a shareable and meaningful shorter version of this content requires significant effort because it combines tedious, repeated operations with personal editorial decisions, which require human judgment. We introduce an editing approach that operates on video transcripts and combines a three-stage large language model pipeline with a timeline-anchored, marker-based interface so editors can inspect and refine suggestions before final assembly. The pipeline first produces an overview summary to maximize content coverage, then induces plain-language selection rules that encode editorial intent, and finally applies rule-conditioned ranking on small transcript windows to mitigate long-context limits, yielding strictly extractive, time-aligned spans under duration constraints. The interface displays groupings of short excerpts using markers with priorities and confidence cues, converting opaque model output into verifiable units within standard video editing workflows. On MeetingBank and MeetingBank-QA datasets, our method outperforms practical extractive baselines at matched lengths. In a within-subjects study with experienced video editors familiar with Premiere Pro video editing software, we found that our marker-based interface provided editors higher efficiency, control, and satisfaction than both a manual editing baseline and an opaque auto-cut condition.
Sana Shah, Mackenzie Leake, Kun Chu, Cornelius Weber, Nico Becherer, Stefan Wermter
IUI6
2026 A parameter-free adaptive resonance theory-based topological clustering algorithm capable of continual learning
Naoki Masuyama, Takanori Takebayashi, Yusuke Nojima, Chu Kiong Loo, Hisao Ishibuchi, Stefan Wermter
Neural Comput. Appl.6
2025 Open-Vocabulary Robotic Object Manipulation using Foundation Models
abstract
Classical vision-language-action models are limited by unidirectional communication, hindering natural human-robot interaction.The recent CrossT5 embeds an efficient vision action pathway into an LLM, but lacks visual generalization, restricting actions to objects seen during training.We introduce OWL×T5, which integrates the OWLv2 object detection model into CrossT5 to enable robot actions on unseen objects.OWL×T5 is trained on a simulated dataset using the NICO humanoid robot and evaluated on the new CLAEO dataset featuring interactions with unseen objects.Results show that OWL×T5 achieves zero-shot object recognition for robotic manipulation, while efficiently integrating vision-language-action capabilities.
Stig Griebenow, Ozan Özdemir, Cornelius Weber, Stefan Wermter
ESANN4
2025 Robots with Attitudes: Influence of LLM-Driven Robot Personalities on Motivation and Performance
abstract
Large language models enable unscripted conversations while maintaining a consistent personality. One desirable personality trait in cooperative partners, known to improve task performance, is agreeableness. To explore the impact of large language models on personality modeling for robots, as well as the effect of agreeable and non-agreeable personalities in cooperative tasks, we conduct a two-part study. This includes an online pre-study for personality validation and a lab-based main study to evaluate the effects on likability, motivation, and task performance. The results demonstrate that the robot’s agreeableness significantly enhances its likability. No significant difference in intrinsic motivation was observed between the two personality types. However, the findings suggest that a robot exhibiting agreeableness and openness to new experiences can enhance task performance. This study highlights the advantages of employing large language models for customized modeling of robot personalities and provides evidence that a carefully chosen agreeable robot personality can positively influence human perceptions and lead to greater success in cooperative scenarios.
Dennis Becker, Kyra Ahrens, Connor Gaede, Erik Strahl, Stefan Wermter
HAI5
2025 LLM-based Interactive Imitation Learning for Robotic Manipulation
abstract
Recent advancements in machine learning provide methods to train autonomous agents capable of handling the increasing complexity of sequential decision-making in robotics. Imitation Learning (IL) is a prominent approach, where agents learn to control robots based on human demonstrations. However, IL commonly suffers from violating the independent and identically distributed (i.i.d) assumption in robotic tasks. Interactive Imitation Learning (IIL) achieves improved performance by allowing agents to learn from interactive feedback from human teachers. Despite these improvements, both approaches come with significant costs due to the necessity of human involvement. Leveraging the emergent capabilities of Large Language Models (LLMs) in reasoning and generating human-like responses, we introduce LLM-iTeach — a novel IIL framework that utilizes an LLM as an interactive teacher to enhance agent performance while alleviating the dependence on human resources. Firstly, LLM-iTeach uses a hierarchical prompting strategy that guides the LLM in generating a policy in Python code. Then, with a designed similarity-based feedback mechanism, LLM-iTeach provides corrective and evaluative feedback interactively during the agent’s training. We evaluate LLM-iTeach against baseline methods such as Behavior Cloning (BC), an IL method, and CEILing, a state-of-the-art IIL method using a human teacher, on various robotic manipulation tasks. Our results demonstrate that LLM-iTeach surpasses BC in the success rate and achieves or even outscores that of CEILing, highlighting the potential of LLMs as cost-effective, human-like teachers in interactive learning environments. We further demonstrate the method’s potential for generalization by evaluating it on additional tasks. The code and prompts are provided at: https://github.com/Tubicor/LLM-iTeach.
Jonas Werner, Kun Chu, Cornelius Weber, Stefan Wermter
IJCNN4
2025 Shaken, Not Stirred: A Novel Dataset for Visual Understanding of Glasses in Human-Robot Bartending Tasks
abstract
Datasets for object detection often do not account for enough variety of glasses, due to their transparent and reflective properties. Specifically, open-vocabulary object detectors, widely used in embodied robotic agents, fail to distinguish subclasses of glasses. This scientific gap poses an issue for robotic applications that suffer from accumulating errors between detection, planning, and action execution. This paper introduces a novel method for acquiring real-world data from RGB-D sensors that minimizes human effort. We propose an auto-labeling pipeline that generates labels for all the acquired frames based on the depth measurements. We provide a novel real-world glass object dataset3that was collected on the Neuro-Inspired COLlaborator (NICOL), a humanoid robot platform. The dataset consists of 7850 images recorded from five different cameras. We show that our trained baseline model outperforms state-of-the-art open-vocabulary approaches. In addition, we deploy our baseline model in an embodied agent approach to the NICOL platform, on which it achieves a success rate of 81% in a human-robot bartending scenario.
Lukás Gajdosech, Hassan Ali 0005, Jan-Gerrit Habekost, Martin Madaras, Matthias Kerzel, Stefan Wermter
IROS6
2025 Personalised Explanations in Long-term Human-Robot Interactions
abstract
In the field of Human-Robot Interaction (HRI), a fundamental challenge is to facilitate human understanding of robots. The emerging domain of eXplainable HRI (XHRI) investigates methods to generate explanations and evaluate their impact on human-robot interactions. Previous works have highlighted the need to personalise the level of detail of these explanations to enhance usability and comprehension. Our paper presents a framework designed to update and retrieve user knowledge-memory models, allowing for adapting the explanations’ level of detail while referencing previously acquired concepts. Three architectures based on our proposed framework that use Large Language Models (LLMs) are evaluated in two distinct scenarios: a hospital patrolling robot and a kitchen assistant robot. Experimental results demonstrate that a two-stage architecture, which first generates an explanation and then personalises it, is the framework architecture that effectively reduces the level of detail only when there is related user knowledge.
Ferran Gebellí, Anais Garrell, Jan-Gerrit Habekost, Séverin Lemaignan, Stefan Wermter, Raquel Ros
RO-MAN5
2025 Unconstrained Open Vocabulary Image Classification: Zero-Shot Transfer from Text to Image via CLIP Inversion
abstract
We introduce NOVIC, an innovative real-time uNconstrained Open Vocabulary Image Classifier that uses an autoregressive transformer to generatively output classification labels as language. Leveraging the extensive knowledge of CLIP models, NOVIC harnesses the embedding space to enable zero-shot transfer from pure text to images. Traditional CLIP models, despite their ability for open vocabulary classification, require an exhaustive prompt of potential class labels, restricting their application to images of known content or context. To address this, we propose an “object decoder” model that is trained on a large-scale 92M-target dataset of templated object noun sets and LLM-generated captions to always output the object noun in question. This effectively inverts the CLIP text encoder and allows textual object labels from essentially the entire English language to be generated directly from image-derived embedding vectors, without requiring any a priori knowledge of the potential content of an image, and without any label biases. The trained decoders are tested on a mix of manually and web-curated datasets, as well as standard image classification benchmarks, and achieve fine-grained prompt-free prediction scores of up to 87.5%, a strong result considering the model must work for any conceivable image and without any contextual clues.11The authors gratefully acknowledge support from the DFG (TRR 169 CML) and European Commission (TRAIL).
Philipp Allgeuer, Kyra Ahrens, Stefan Wermter
WACV3
2025 FabuLight-ASD: unveiling speech activity via body language
abstract
Abstract Active speaker detection (ASD) in multimodal environments is crucial for various applications, from video conferencing to human-robot interaction. This paper introduces FabuLight-ASD, an advanced ASD model that integrates facial, audio, and body pose information to enhance detection accuracy and robustness. Our model builds upon the existing Light-ASD framework by incorporating human pose data, represented through skeleton graphs, which minimises computational overhead. Using the Wilder Active Speaker Detection (WASD) dataset, renowned for reliable face and body bounding box annotations, we demonstrate FabuLight-ASD’s effectiveness in real-world scenarios. Achieving an overall mean average precision (mAP) of 94.3%, FabuLight-ASD outperforms Light-ASD, which has an overall mAP of 93.7% across various challenging scenarios. The incorporation of body pose information shows a particularly advantageous impact, with notable improvements in mAP observed in scenarios with speech impairment, face occlusion, and human voice background noise. Furthermore, efficiency analysis indicates only a modest increase in parameter count (27.3%) and multiply-accumulate operations (up to 2.4%), underscoring the model’s efficiency and feasibility. These findings validate the efficacy of FabuLight-ASD in enhancing ASD performance through the integration of body pose data. FabuLight-ASD’s code and model weights are available at https://github.com/knowledgetechnologyuhh/FabuLight-ASD .
Hugo C. C. Carneiro, Stefan Wermter
Neural Comput. Appl.2
2025 Influence of Robots' Voice Naturalness on Trust and Compliance
abstract
With the increasing performance of text-to-speech systems and their generated voices indistinguishable from natural human speech, the use of these systems for robots raises ethical and safety concerns. A robot with a natural voice could increase trust, which might result in over-reliance despite evidence for robot unreliability. To estimate the influence of a robot’s voice on trust and compliance, we design a study that consists of two experiments. In a pre-study ( \(N_{1}=60\) ) the most suitable natural and mechanical voice for the main study are estimated and selected for the main study. Afterward, in the main study ( \(N_{2}=68\) ), the influence of a robot’s voice on trust and compliance is evaluated in a cooperative game of Battleship with a robot as an assistant. During the experiment, the acceptance of the robot’s advice and response time are measured, which indicate trust and compliance, respectively. The results show that participants expect robots to sound human-like and that a robot with a natural voice is perceived as safer. Additionally, a natural voice can affect compliance. Despite repeated incorrect advice, the participants are more likely to rely on the robot with the natural voice. The results do not show a direct effect on trust. Natural voices provide increased intelligibility, and while they can increase compliance with the robot, the results indicate that natural voices might not lead to over-reliance. The results highlight the importance of incorporating voices into the design of social robots to improve communication, avoid adverse effects, and increase acceptance and adoption in society.
Dennis Becker, Lukas Braach, Lennart Clasmeier, Teresa Kaufmann, Oskar Ong, Kyra Ahrens, Connor Gaede, Erik Strahl, Di Fu, Stefan Wermter
ACM Trans. Hum. Robot Interact.10
2025 Disentanglement of Prosody Representations via Diffusion Models and Scheduled Gradient Reversal
abstract
Prosody plays a fundamental role in human speech and communication, facilitating intelligibility and conveying emotional and cognitive states. Extracting accurate prosodic information from speech is vital for building assistive technology, such as controllable speech synthesis, speaking style transfer, and speech emotion recognition (SER). However, it is challenging to disentangle speaker-independent prosody representations since prosodic attributes, such as intonation, excessively entangle with speaker-specific attributes, e.g., pitch. In this article, we propose a novel model, called Diffsody, to disentangle and refine prosody representations: 1) to disentangle prosody representations, we leverage the expressive generative ability of a diffusion model by conditioning it on quantified semantic information and pretrained speaker embeddings. Additionally, a prosody encoder automatically learns prosody representations used for spectrogram reconstruction in an unsupervised fashion; and 2) to refine and learn speaker-invariant prosody representations, a scheduled gradient reversal layer (sGRL) is proposed and integrated into the prosody encoder of Diffsody. We extensively evaluate Diffsody through qualitative and quantitative means. t-SNE visualization and speaker verification experiments demonstrate the efficacy of the sGRL method in preventing speaker-specific information leakage. Experimental results on speaker-independent SER and automatic depression detection (ADD) tasks demonstrate that Diffsody can efficiently factorize speaker-independent prosody representations, resulting in a significant boost in SER and ADD. In addition, Diffsody synergistically integrates with the semantic representation model WavLM, which leads to a discernibly elevated performance, outperforming contemporary methods in both SER and ADD tasks. Furthermore, the Diffsody model exhibits promising potential for various practical applications, such as voice or style conversion. Some audio samples can be found on our https://leyuanqu.github.io/Diffsody/demo website.
Leyuan Qu, Cornelius Weber, Wei Wang 0310, Jia Jin, Yingming Gao, Taihao Li, Stefan Wermter
IEEE Trans. Neural Networks Learn. Syst.7
2024 Enhancing Zero-Shot Chain-of-Thought Reasoning in Large Language Models through Logic
abstract
Recent advancements in large language models have showcased their remarkable generalizability across various domains. However, their reasoning abilities still have significant room for improvement, especially when confronted with scenarios requiring multi-step reasoning. Although large language models possess extensive knowledge, their reasoning often fails to effectively utilize this knowledge to establish a coherent thinking paradigm. These models sometimes show hallucinations as their reasoning procedures are unconstrained by logical principles. Aiming at improving the zero-shot chain-of-thought reasoning ability of large language models, we propose LoT (Logical Thoughts), a self-improvement prompting framework that leverages principles rooted in symbolic logic, particularly Reductio ad Absurdum, to systematically verify and rectify the reasoning processes step by step. Experimental evaluations conducted on language tasks in diverse domains, including arithmetic, commonsense, symbolic, causal inference, and social problems, demonstrate the efficacy of enhanced reasoning by logic. The implementation code for LoT can be accessed at: https://github.com/xf-zhao/LoT.
Xufeng Zhao 0002, Mengdi Li 0006, Wenhao Lu, Cornelius Weber, Jae Hee Lee 0001, Kun Chu, Stefan Wermter
LREC/COLING7
2024 Embodying Language Models in Robot Action
Connor Gaede, Ozan Özdemir, Cornelius Weber, Stefan Wermter
ESANN4
2024 Wrapyfi: A Python Wrapper for Integrating Robots, Sensors, and Applications across Multiple Middleware
abstract
Message oriented and robotics middleware play an important role in facilitating robot control, abstracting complex functionality, and unifying communication patterns between sensors and devices. However, using multiple middleware frameworks presents a challenge in integrating different robots within a single system. To address this challenge, we present Wrapyfi, a Python wrapper supporting multiple message oriented and robotics middleware, including ZeroMQ, YARP, ROS, and ROS 2. Wrapyfi also provides plugins for exchanging deep learning framework data, without additional encoding or preprocessing steps. Using Wrapyfi eases the development of scripts that run on multiple machines, thereby enabling cross-platform communication and workload distribution. We finally present the three communication schemes that form the cornerstone of Wrapyfi's communication model, along with examples that demonstrate their applicability.
Fares Abawi, Philipp Allgeuer, Di Fu, Stefan Wermter
HRI4
2024 When Robots Get Chatty: Grounding Multimodal Human-Robot Conversation and Collaboration
abstract
Abstract We investigate the use of Large Language Models (LLMs) to equip neural robotic agents with human-like social and cognitive competencies, for the purpose of open-ended human-robot conversation and collaboration. We introduce a modular and extensible methodology for grounding an LLM with the sensory perceptions and capabilities of a physical robot, and integrate multiple deep learning models throughout the architecture in a form of system integration. The integrated models encompass various functions such as speech recognition, speech generation, open-vocabulary object detection, human pose estimation, and gesture detection, with the LLM serving as the central text-based coordinating unit. The qualitative and quantitative results demonstrate the huge potential of LLMs in providing emergent cognition and interactive language-oriented control of robots in a natural and social manner. Video: https://youtu.be/A2WLEuiM3-s .
Philipp Allgeuer, Hassan Ali 0005, Stefan Wermter
ICANN (4)3
2024 Details Make a Difference: Object State-Sensitive Neurorobotic Task Planning
Xiaowen Sun, Xufeng Zhao 0002, Jae Hee Lee 0001, Wenhao Lu, Matthias Kerzel, Stefan Wermter
ICANN (4)6
2024 An Energy Sampling Replay-Based Continual Learning Framework
Xingzhong Zhang, Joon Huang Chuah, Chu Kiong Loo, Stefan Wermter
ICANN (2)4
2024 Improving Speech Emotion Recognition with Unsupervised Speaking Style Transfer
abstract
Humans can effortlessly modify various prosodic attributes, such as the placement of stress and the intensity of sentiment, to convey a specific emotion while maintaining consistent linguistic content. Motivated by this capability, we propose EmoAug, a novel style transfer model designed to enhance emotional expression and tackle the data scarcity issue in speech emotion recognition tasks. EmoAug consists of a semantic encoder and a paralinguistic encoder that represent verbal and non-verbal information respectively. Additionally, a decoder reconstructs speech signals by conditioning on the aforementioned two information flows in an unsupervised fashion. Once training is completed, EmoAug enriches expressions of emotional speech with different prosodic attributes, such as stress, rhythm and intensity, by feeding different styles into the paralinguistic encoder. EmoAug enables us to generate similar numbers of samples for each class to tackle the data imbalance issue as well. Experimental results on the IEMOCAP dataset demonstrate that EmoAug can successfully transfer different speaking styles while retaining the speaker identity and semantic content. Furthermore, we train a SER model with data augmented by EmoAug and show that the augmented model not only surpasses the state-of-the-art supervised and self-supervised methods but also overcomes overfitting problems caused by data imbalance. Some audio samples can be found on our demo website1.
Leyuan Qu, Wei Wang 0310, Cornelius Weber, Pengcheng Yue, Taihao Li, Stefan Wermter
ICASSP6
2024 Domain Adaption as Auxiliary Task for Sim-to-Real Transfer in Vision-based Neuro-Robotic Control
abstract
Architectures for vision-based robot manipulation often utilize separate domain adaption models to allow sim-to-real transfer and an inverse kinematics solver to allow the actual policy to operate in Cartesian space. We present a novel end-to-end visuomotor architecture that combines domain adaption and inherent inverse kinematics in one model. Using the same latent encoding, it jointly learns to reconstruct canonical simulation images from randomized inputs and to predict the corresponding joint angles that minimize the Cartesian error towards a depicted target object via differentiable forward kinematics.We evaluate our model in a sim-to-real grasping experiment with the NICO humanoid robot by comparing different randomization and adaption conditions both directly and with additional real-world finetuning. Our combined method significantly increases the resulting accuracy and allows a finetuned model to reach a success rate of 80.30%, outperforming a real-world model trained with six times as much real data.
Connor Gaede, Jan-Gerrit Habekost, Stefan Wermter
IJCNN3
2024 Inverse Kinematics for Neuro-Robotic Grasping with Humanoid Embodied Agents
abstract
This paper introduces a novel zero-shot motion planning method that allows users to quickly design smooth robot motions in Cartesian space. A Bézier curve-based Cartesian plan is transformed into a joint space trajectory by our neuro-inspired inverse kinematics (IK) method CycleIK, for which we enable platform independence by scaling it to arbitrary robot designs. The motion planner is evaluated on the physical hardware of the two humanoid robots NICO and NICOL in a human-in-the-loop grasping scenario. Our method is deployed with an embodied agent that is a large language model (LLM) at its core. We generalize the embodied agent, that was introduced for NICOL, to also embody NICO. The agent can execute a discrete set of physical actions and allows the user to verbally instruct various different robots. We contribute a grasping primitive to its action space that allows for precise manipulation of household objects. The updated CycleIK2method is compared to popular numerical IK solvers and state-of-the-art neural IK methods in simulation and is shown to be competitive with or outperform all evaluated methods when the algorithm runtime is very short. The grasping primitive is evaluated on both NICOL and NICO robots with a reported grasp success of 72% to 82% for each robot, respectively.
Jan-Gerrit Habekost, Connor Gaede, Philipp Allgeuer, Stefan Wermter
IROS4
2024 Adaptive knowledge distillation and integration for weakly supervised referring expression comprehension
Jinpeng Mi, Stefan Wermter, Jianwei Zhang 0001
Knowl. Based Syst.2
2024 Disentangling Prosody Representations With Unsupervised Speech Reconstruction
abstract
Human speech can be characterized by different components, including semantic content, speaker identity and prosodic information. Significant progress has been made in disentangling representations for semantic content and speaker identity in Automatic Speech Recognition (ASR) and speaker verification tasks respectively. However, it is still an open challenging research question to extract prosodic information because of the intrinsic association of different attributes, such as timbre and rhythm, and because of the need for supervised training schemes to achieve robust large-scale and speaker-independent ASR. The aim of this paper is to address the disentanglement of emotional prosody from speech based on unsupervised reconstruction. Specifically, we identify, design, implement and integrate three crucial components in our proposed speech reconstruction model Prosody2Vec: (1) a unit encoder that transforms speech signals into discrete units for semantic content, (2) a pretrained speaker verification model to generate speaker identity embeddings, and (3) a trainable prosody encoder to learn prosody representations. We first pretrain the Prosody2Vec representations on unlabelled emotional speech corpora, then fine-tune the model on specific datasets to perform Speech Emotion Recognition (SER) and Emotional Voice Conversion (EVC) tasks. Both objective (weighted and unweighted accuracies) and subjective (mean opinion score) evaluations on the EVC task suggest that Prosody2Vec effectively captures general prosodic features that can be smoothly transferred to other emotional speech. In addition, our SER experiments on the IEMOCAP dataset reveal that the prosody features learned by Prosody2Vec are complementary and beneficial for the performance of widely used speech pretraining models and surpass the state-of-the-art methods when combining Prosody2Vec with HuBERT representations. Some audio samples can be found on our demo website
Leyuan Qu, Taihao Li, Cornelius Weber, Theresa Pekarek-Rosin, Fuji Ren, Stefan Wermter
IEEE ACM Trans. Audio Speech Lang. Process.6
2024 LipSound2: Self-Supervised Pre-Training for Lip-to-Speech Reconstruction and Lip Reading
abstract
The aim of this work is to investigate the impact of crossmodal self-supervised pre-training for speech reconstruction (video-to-audio) by leveraging the natural co-occurrence of audio and visual streams in videos. We propose LipSound2 that consists of an encoder-decoder architecture and location-aware attention mechanism to map face image sequences to mel-scale spectrograms directly without requiring any human annotations. The proposed LipSound2 model is first pre-trained on ∼ 2400 -h multilingual (e.g., English and German) audio-visual data (VoxCeleb2). To verify the generalizability of the proposed method, we then fine-tune the pre-trained model on domain-specific datasets (GRID and TCD-TIMIT) for English speech reconstruction and achieve a significant improvement on speech quality and intelligibility compared to previous approaches in speaker-dependent and speaker-independent settings. In addition to English, we conduct Chinese speech reconstruction on the Chinese Mandarin Lip Reading (CMLR) dataset to verify the impact on transferability. Finally, we train the cascaded lip reading (video-to-text) system by fine-tuning the generated audios on a pre-trained speech recognition system and achieve the state-of-the-art performance on both English and Chinese benchmark datasets.
Leyuan Qu, Cornelius Weber, Stefan Wermter
IEEE Trans. Neural Networks Learn. Syst.3
2023 Visually Grounded Commonsense Knowledge Acquisition
abstract
Large-scale commonsense knowledge bases empower a broad range of AI applications, where the automatic extraction of commonsense knowledge (CKE) is a fundamental and challenging problem. CKE from text is known for suffering from the inherent sparsity and reporting bias of commonsense in text. Visual perception, on the other hand, contains rich commonsense knowledge about real-world entities, e.g., (person, can_hold, bottle), which can serve as promising sources for acquiring grounded commonsense knowledge. In this work, we present CLEVER, which formulates CKE as a distantly supervised multi-instance learning problem, where models learn to summarize commonsense relations from a bag of images about an entity pair without any human annotation on image instances. To address the problem, CLEVER leverages vision-language pre-training models for deep understanding of each image in the bag, and selects informative instances from the bag to summarize commonsense entity relations via a novel contrastive attention mechanism. Comprehensive experimental results in held-out and human evaluation show that CLEVER can extract commonsense knowledge in promising quality, outperforming pre-trained language model-based methods by 3.9 AUC and 6.4 mAUC points. The predicted commonsense scores show strong correlation with human judgment with a 0.78 Spearman coefficient. Moreover, the extracted commonsense can also be grounded into images with reasonable interpretability. The data and codes can be obtained at https://github.com/thunlp/CLEVER.
Yuan Yao 0013, Tianyu Yu 0002, Mengdi Li 0006, Ruobing Xie, Cornelius Weber, Zhiyuan Liu 0001, Hai-Tao Zheng 0002, Stefan Wermter, Tat-Seng Chua, Maosong Sun 0001
AAAI9
2023 Bring the Noise: Introducing Noise Robustness to Pretrained Automatic Speech Recognition
abstract
Abstract In recent research, in the domain of speech processing, large End-to-End (E2E) systems for Automatic Speech Recognition (ASR) have reported state-of-the-art performance on various benchmarks. These systems intrinsically learn how to handle and remove noise conditions from speech. Previous research has shown, that it is possible to extract the denoising capabilities of these models into a preprocessor network, which can be used as a frontend for downstream ASR models. However, the proposed methods were limited to specific fully convolutional architectures. In this work, we propose a novel method to extract the denoising capabilities, that can be applied to any encoder-decoder architecture. We propose the Cleancoder preprocessor architecture that extracts hidden activations from the Conformer ASR model and feeds them to a decoder to predict denoised spectrograms. We train our preprocessor on the Noisy Speech Database (NSD) to reconstruct denoised spectrograms from noisy inputs. Then, we evaluate our model as a frontend to a pretrained Conformer ASR model as well as a frontend to train smaller Conformer ASR models from scratch. We show that the Cleancoder is able to filter noise from speech and that it improves the total Word Error Rate (WER) of the downstream model in noisy conditions for both applications.
Patrick Eickhoff, Matthias Möller, Theresa Pekarek-Rosin, Johannes Twiefel, Stefan Wermter
ICANN (7)5
2023 Neural Field Conditioning Strategies for 2D Semantic Segmentation
abstract
Abstract Neural fields are neural networks which map coordinates to a desired signal. When a neural field should jointly model multiple signals, and not memorize only one, it needs to be conditioned on a latent code which describes the signal at hand. Despite being an important aspect, there has been little research on conditioning strategies for neural fields. In this work, we explore the use of neural fields as decoders for 2D semantic segmentation. For this task, we compare three conditioning methods, simple concatenation of the latent code, Feature-wise Linear Modulation (FiLM), and Cross-Attention, in conjunction with latent codes which either describe the full image or only a local region of the image. Our results show a considerable difference in performance between the examined conditioning strategies. Furthermore, we show that conditioning via Cross-Attention achieves the best results and is competitive with a CNN-based decoder for semantic segmentation.
Martin Gromniak, Sven Magg, Stefan Wermter
ICANN (2)3
2023 CycleIK: Neuro-inspired Inverse Kinematics
abstract
Abstract The paper introduces CycleIK, a neuro-robotic approach that wraps two novel neuro-inspired methods for the inverse kinematics (IK) task—a Generative Adversarial Network (GAN), and a Multi-Layer Perceptron architecture. These methods can be used in a standalone fashion, but we also show how embedding these into a hybrid neuro-genetic IK pipeline allows for further optimization via sequential least-squares programming (SLSQP) or a genetic algorithm (GA). The models are trained and tested on dense datasets that were collected from random robot configurations of the new Neuro-Inspired COLlaborator (NICOL), a semi-humanoid robot with two redundant 8-DoF manipulators. We utilize the weighted multi-objective function from the state-of-the-art BioIK method to support the training process and our hybrid neuro-genetic architecture. We show that the neural models can compete with state-of-the-art IK approaches, which allows for deployment directly to robotic hardware. Additionally, it is shown that the incorporation of the genetic algorithm improves the precision while simultaneously reducing the overall runtime.
Jan-Gerrit Habekost, Erik Strahl, Philipp Allgeuer, Matthias Kerzel, Stefan Wermter
ICANN (1)5
2023 Replay to Remember: Continual Layer-Specific Fine-Tuning for German Speech Recognition
abstract
Abstract While Automatic Speech Recognition (ASR) models have shown significant advances with the introduction of unsupervised or self-supervised training techniques, these improvements are still only limited to a subsection of languages and speakers. Transfer learning enables the adaptation of large-scale multilingual models to not only low-resource languages but also to more specific speaker groups. However, fine-tuning on data from new domains is usually accompanied by a decrease in performance on the original domain. Therefore, in our experiments, we examine how well the performance of large-scale ASR models can be approximated for smaller domains, with our own dataset of German Senior Voice Commands (SVC-de), and how much of the general speech recognition performance can be preserved by selectively freezing parts of the model during training. To further increase the robustness of the ASR model to vocabulary and speakers outside of the fine-tuned domain, we apply Experience Replay [20] for continual learning. By adding only a fraction of data from the original domain, we are able to reach Word-Error-Rates (WERs) below 5% on the new domain, while stabilizing performance for general speech recognition at acceptable WERs.
Theresa Pekarek-Rosin, Stefan Wermter
ICANN (7)2
2023 Clarifying the Half Full or Half Empty Question: Multimodal Container Classification
abstract
Abstract Multimodal integration is a key component of allowing robots to perceive the world. Multimodality comes with multiple challenges that have to be considered, such as how to integrate and fuse the data. In this paper, we compare different possibilities of fusing visual, tactile and proprioceptive data. The data is directly recorded on the NICOL robot in an experimental setup in which the robot has to classify containers and their content. Due to the different nature of the containers, the use of the modalities can wildly differ between the classes. We demonstrate the superiority of multimodal solutions in this use case and evaluate three fusion strategies that integrate the data at different time steps. We find that the accuracy of the best fusion strategy is 15% higher than the best strategy using only one singular sense.
Josua Spisak, Matthias Kerzel, Stefan Wermter
ICANN (1)3
2023 Partially Adaptive Multichannel Joint Reduction of Ego-Noise and Environmental Noise
abstract
Human-robot interaction relies on a noise-robust audio processing module capable of estimating target speech from audio recordings impacted by environmental noise, as well as self-induced noise, so-called ego-noise. While external ambient noise sources vary from environment to environment, ego-noise is mainly caused by the internal motors and joints of a robot. Egonoise and environmental noise reduction are often decoupled, i.e., ego-noise reduction is performed without considering environmental noise. Recently, a variational autoencoder (VAE)-based speech model has been combined with a fully adaptive non-negative matrix factorization (NMF) noise model to recover clean speech under different environmental noise disturbances. However, its enhancement performance is limited in adverse acoustic scenarios involving, e.g. ego-noise. In this paper, we propose a multichannel partially adaptive scheme to jointly model ego-noise and environmental noise utilizing the VAE-NMF framework, where we take advantage of spatially and spectrally structured characteristics of ego-noise by pre-training the ego-noise model, while retaining the ability to adapt to unknown environmental noise. Experimental results show that our proposed approach outperforms the methods based on a completely fixed scheme and a fully adaptive scheme when ego-noise and environmental noise are present simultaneously.
Huajian Fang, Niklas Wittmer, Johannes Twiefel, Stefan Wermter, Timo Gerkmann
ICASSP4
2023 Internally Rewarded Reinforcement Learning
abstract
We study a class of reinforcement learning problems where the reward signals for policy learning are generated by a discriminator that is dependent on and jointly optimized with the policy. This interdependence between the policy and the discriminator leads to an unstable learning process because reward signals from an immature discriminator are noisy and impede policy learning, and conversely, an under-optimized policy impedes discriminator learning. We call this learning setting $\textit{Internally Rewarded Reinforcement Learning}$ (IRRL) as the reward is not provided directly by the environment but $\textit{internally}$ by the discriminator. In this paper, we formally formulate IRRL and present a class of problems that belong to IRRL. We theoretically derive and empirically analyze the effect of the reward function in IRRL and based on these analyses propose the clipped linear reward function. Experimental results show that the proposed reward function can consistently stabilize the training process by reducing the impact of reward noise, which leads to faster convergence and higher performance compared with baselines in diverse tasks.
Mengdi Li 0006, Xufeng Zhao 0002, Jae Hee Lee 0001, Cornelius Weber, Stefan Wermter
ICML5
2023 Sample-Efficient Real-Time Planning with Curiosity Cross-Entropy Method and Contrastive Learning
abstract
Model-based reinforcement learning (MBRL) with real-time planning has shown great potential in locomotion and manipulation control tasks. However, the existing planning methods, such as the Cross-Entropy Method (CEM), do not scale well to complex high-dimensional environments. One of the key reasons for underperformance is the lack of exploration, as these planning methods only aim to maximize the cumulative extrinsic reward over the planning horizon. Furthermore, planning inside the compact latent space in the absence of observations makes it challenging to use curiosity-based intrinsic motivation. We propose Curiosity CEM (CCEM), an improved version of the CEM algorithm for encouraging exploration via curiosity. Our proposed method maximizes the sum of state-action$Q$values over the planning horizon, in which these$Q$values estimate the future extrinsic and intrinsic reward, hence encouraging to reach novel observations. In addition, our model uses contrastive representation learning to efficiently learn latent representations. Experiments on image-based continuous control tasks from the DeepMind Control suite show that CCEM is by a large margin more sample-efficient than previous MBRL algorithms and compares favorably with the best model-free RL methods.
Mostafa Kotb, Cornelius Weber, Stefan Wermter
IROS3
2023 Chat with the Environment: Interactive Multimodal Perception Using Large Language Models
abstract
Programming robot behavior in a complex world faces challenges on multiple levels, from dextrous low-level skills to high-level planning and reasoning. Recent pre-trained Large Language Models (LLMs) have shown remarkable reasoning ability in few-shot robotic planning. However, it remains challenging to ground LLMs in multimodal sensory input and continuous action output, while enabling a robot to interact with its environment and acquire novel information as its policies unfold. We develop a robot interaction scenario with a partially observable state, which necessitates a robot to decide on a range of epistemic actions in order to sample sensory information among multiple modalities, before being able to execute the task correctly. An interactive perception framework is therefore proposed with an LLM as its backbone, whose ability is exploited to instruct epistemic actions and to reason over the resulting multimodal sensations (vision, sound, haptics, proprioception), as well as to plan an entire task execution based on the interactively acquired information. Our study demonstrates that LLMs can provide high-level planning and reasoning skills and control interactive robot behavior in a multimodal environment, while multimodal modules with the context of the environmental state help ground the LLMs and extend their processing ability. The project website can be found at https://matcha-model.github.io/.
Xufeng Zhao 0002, Mengdi Li 0006, Cornelius Weber, Muhammad Burhan Hafez, Stefan Wermter
IROS5
2023 The Emotional Dilemma: Influence of a Human-like Robot on Trust and Cooperation
abstract
Increasing anthropomorphic robot behavioral design could affect trust and cooperation positively. However, studies have shown contradicting results and suggest a task-dependent relationship between robots that display emotions and trust. Therefore, this study analyzes the effect of robots that display human-like emotions on trust, cooperation, and participants’ emotions. In the between-group study, participants play the coin entrustment game with an emotional and a non-emotional robot. The results showthat the robot that displays emotions induces more anxiety than the neutral robot. Accordingly, the participants trust the emotional robot less and are less likely to cooperate. Furthermore, the perceived intelligence of a robot increases trust, while a desire to outcompete the robot can reduce trust and cooperation. Thus, the design of robots expressing emotions should be task dependent to avoid adverse effects that reduce trust and cooperation.
Dennis Becker, Diana Rueda, Felix Beese, Brenda Scarleth Gutierrez Torres, Myriem Lafdili, Kyra Ahrens, Di Fu, Erik Strahl, Tom Weber, Stefan Wermter
RO-MAN10
2023 The Robot in the Room: Influence of Robot Facial Expressions and Gaze on Human-Human-Robot Collaboration
abstract
Robot facial expressions and gaze are important factors for enhancing human-robot interaction (HRI), but their effects on human collaboration and perception are not well understood, for instance, in collaborative game scenarios. In this study, we designed a collaborative triadic HRI game scenario where two participants worked together to insert objects into a shape sorter. One participant assumed the role of a guide. The guide instructed the other participant, who played the role of an actor, to place occluded objects into the sorter. A humanoid robot issued instructions, observed the interaction, and displayed social cues to elicit changes in the two participants’ behavior. We measured human collaboration as a function of task completion time and the participants’ perceptions of the robot by rating its behavior as intelligent or random. Participants also evaluated the robot by filling out the Godspeed questionnaire. We found that human collaboration was higher when the robot displayed a happy facial expression at the beginning of the game compared to a neutral facial expression. We also found that participants perceived the robot as more intelligent when it displayed a positive facial expression at the end of the game. The robot’s behavior was also perceived as intelligent when directing its gaze toward the guide at the beginning of the interaction, not the actor. These findings provide insights into how robot facial expressions and gaze influence human behavior and perception in collaboration.
Di Fu, Fares Abawi, Stefan Wermter
RO-MAN3
2023 Whose emotion matters? Speaking activity localisation without prior knowledge
abstract
The task of emotion recognition in conversations (ERC) benefits from the availability of multiple modalities, as provided, for example, in the video-based Multimodal EmotionLines Dataset (MELD). However, only a few research approaches use both acoustic and visual information from the MELD videos. There are two reasons for this: First, label-to-video alignments in MELD are noisy, making those videos an unreliable source of emotional speech data. Second, conversations can involve several people in the same scene, which requires the localisation of the utterance source. In this paper, we introduce MELD with Fixed Audiovisual Information via Realignment (MELD-FAIR) by using recent active speaker detection and automatic speech recognition models, we are able to realign the videos of MELD and capture the facial expressions from speakers in 96.92% of the utterances provided in MELD. Experiments with a self-supervised voice recognition model indicate that the realigned MELD-FAIR videos more closely match the transcribed utterances given in the MELD dataset. Finally, we devise a model for emotion recognition in conversations trained on the realigned MELD-FAIR videos, which outperforms state-of-the-art models for ERC based on vision alone. This indicates that localising the source of speaking activities is indeed effective for extracting facial expressions from the uttering speakers and that faces provide more informative visual cues than the visual features state-of-the-art models have been using so far. The MELD-FAIR realignment data, and the code of the realignment procedure and of the emotional recognition, are available at https://github.com/knowledgetechnologyuhh/MELD-FAIR.
Hugo C. C. Carneiro, Cornelius Weber, Stefan Wermter
Neurocomputing3
2023 An adaptive growing grid model for a non-stationary environment
Chihli Hung, Stefan Wermter, Yu-Liang Chi, Chih-Fong Tsai
Neurocomputing2
2023 Hierarchical goals contextualize local reward decomposition explanations
abstract
Abstract One-step reinforcement learning explanation methods account for individual actions but fail to consider the agent’s future behavior, which can make their interpretation ambiguous. We propose to address this limitation by providing hierarchical goals as context for one-step explanations. By considering the current hierarchical goal as a context, one-step explanations can be interpreted with higher certainty, as the agent’s future behavior is more predictable. We combine reward decomposition with hierarchical reinforcement learning into a novel explainable reinforcement learning framework, which yields more interpretable, goal-contextualized one-step explanations. With a qualitative analysis of one-step reward decomposition explanations, we first show that their interpretability is indeed limited in scenarios with multiple, different optimal policies—a characteristic shared by other one-step explanation methods. Then, we show that our framework retains high interpretability in such cases, as the hierarchical goal can be considered as context for the explanation. To the best of our knowledge, our work is the first to investigate hierarchical goals not as an explanation directly but as additional context for one-step reinforcement learning explanations.
Finn Rietz, Sven Magg, Fredrik Heintz, Todor Stoyanov, Stefan Wermter, Johannes A. Stork
Neural Comput. Appl.5
2023 Emphasizing unseen words: New vocabulary acquisition for end-to-end speech recognition
abstract
Due to the dynamic nature of human language, automatic speech recognition (ASR) systems need to continuously acquire new vocabulary. Out-Of-Vocabulary (OOV) words, such as trending words and new named entities, pose problems to modern ASR systems that require long training times to adapt their large numbers of parameters. Different from most previous research focusing on language model post-processing, we tackle this problem on an earlier processing level and eliminate the bias in acoustic modeling to recognize OOV words acoustically. We propose to generate OOV words using text-to-speech systems and to rescale losses to encourage neural networks to pay more attention to OOV words. Specifically, we enlarge the classification loss used for training neural networks' parameters of utterances containing OOV words (sentence-level), or rescale the gradient used for back-propagation for OOV words (word-level), when fine-tuning a previously trained model on synthetic audio. To overcome catastrophic forgetting, we also explore the combination of loss rescaling and model regularization, i.e. L2 regularization and elastic weight consolidation (EWC). Compared with previous methods that just fine-tune synthetic audio with EWC, the experimental results on the LibriSpeech benchmark reveal that our proposed loss rescaling approach can achieve significant improvement on the recall rate with only a slight decrease on word error rate. Moreover, word-level rescaling is more stable than utterance-level rescaling and leads to higher recall rates and precision rates on OOV word recognition. Furthermore, our proposed combined loss rescaling and weight consolidation methods can support continual learning of an ASR system.
Leyuan Qu, Cornelius Weber, Stefan Wermter
Neural Networks3
2023 Integrating Uncertainty Into Neural Network-Based Speech Enhancement
abstract
Supervised masking approaches in the time-frequency domain aim to employ deep neural networks to estimate a multiplicative mask to extract clean speech. This leads to a single estimate for each input without any guarantees or measures of reliability. In this paper, we study the benefits of modeling uncertainty in clean speech estimation. Prediction uncertainty is typically categorized intoaleatoric uncertaintyandepistemic uncertainty. The former refers to inherent randomness in data, while the latter describes uncertainty in the model parameters. In this work, we propose a framework to jointly model aleatoric and epistemic uncertainties in neural network-based speech enhancement. The proposed approach captures aleatoric uncertainty by estimating the statistical moments of the speech posterior distribution and explicitly incorporates the uncertainty estimate to further improve clean speech estimation. For epistemic uncertainty, we investigate two Bayesian deep learning approaches: Monte Carlo dropout and Deep ensembles to quantify the uncertainty of the neural network parameters. Our analyses show that the proposed framework promotes capturing practical and reliable uncertainty, while combining different sources of uncertainties yields more reliable predictive uncertainty estimates. Furthermore, we demonstrate the benefits of modeling uncertainty on speech enhancement performance by evaluating the framework on different datasets, exhibiting notable improvement over comparable models that fail to account for uncertainty.
Huajian Fang, Dennis Becker, Stefan Wermter, Timo Gerkmann
IEEE ACM Trans. Audio Speech Lang. Process.3
2022 Sim-to-Real Neural Learning with Domain Randomisation for Humanoid Robot Grasping
Connor Gaede, Matthias Kerzel, Erik Strahl, Stefan Wermter
ICANN (1)4
2022 Word-by-Word Generation of Visual Dialog Using Reinforcement Learning
Yuliia Lysa, Cornelius Weber, Dennis Becker, Stefan Wermter
ICANN (2)4
2022 Learning Flexible Translation Between Robot Actions and Language Descriptions
Ozan Özdemir, Matthias Kerzel, Cornelius Weber, Jae Hee Lee 0001, Stefan Wermter
ICANN (2)5
2022 NeoSLAM: Neural Object SLAM for Loop Closure and Navigation
Younes Raoui, Cornelius Weber, Stefan Wermter
ICANN (3)3
2022 Learning Visually Grounded Human-Robot Dialog in a Hybrid Neural Architecture
Xiaowen Sun, Cornelius Weber, Matthias Kerzel, Tom Weber, Mengdi Li 0006, Stefan Wermter
ICANN (2)6
2022 More Diverse Training, Better Compositionality! Evidence from Multimodal Language Learning
Caspar Volquardsen, Jae Hee Lee 0001, Cornelius Weber, Stefan Wermter
ICANN (3)4
2022 Integrating Statistical Uncertainty into Neural Network-Based Speech Enhancement
abstract
Speech enhancement in the time-frequency domain is often performed by estimating a multiplicative mask to extract clean speech. However, most neural network-based methods perform point estimation, i.e., their output consists of a single mask. In this paper, we study the benefits of modeling uncertainty in neural network-based speech enhancement. For this, our neural network is trained to map a noisy spectrogram to the Wiener filter and its associated variance, which quantifies uncertainty, based on the maximum a posteriori (MAP) inference of spectral coefficients. By estimating the distribution instead of the point estimate, one can model the uncertainty associated with each estimate. We further propose to use the estimated Wiener filter and its uncertainty to build an approximate MAP (A-MAP) estimator of spectral magnitudes, which in turn is combined with the MAP inference of spectral coefficients to form a hybrid loss function to jointly reinforce the estimation. Experimental results on different datasets show that the proposed method can not only capture the uncertainty associated with the estimated filters, but also yield a higher enhancement performance over comparable models that do not take uncertainty into account.
Huajian Fang, Tal Peer, Stefan Wermter, Timo Gerkmann
ICASSP3
2022 What is Right for Me is Not Yet Right for You: A Dataset for Grounding Relative Directions via Multi-Task Learning
abstract
Understanding spatial relations is essential for intelligent agents to act and communicate in the physical world. Relative directions are spatial relations that describe the relative positions of target objects with regard to the intrinsic orientation of reference objects. Grounding relative directions is more difficult than grounding absolute directions because it not only requires a model to detect objects in the image and to identify spatial relation based on this information, but it also needs to recognize the orientation of objects and integrate this information into the reasoning process. We investigate the challenging problem of grounding relative directions with end-to-end neural networks. To this end, we provide GRiD-3D, a novel dataset that features relative directions and complements existing visual question answering (VQA) datasets, such as CLEVR, that involve only absolute directions. We also provide baselines for the dataset with two established end-to-end VQA models. Experimental evaluations show that answering questions on relative directions is feasible when questions in the dataset simulate the necessary subtasks for grounding relative directions. We discover that those subtasks are learned in an order that reflects the steps of an intuitive pipeline for processing relative directions.
Jae Hee Lee 0001, Matthias Kerzel, Kyra Ahrens, Cornelius Weber, Stefan Wermter
IJCAI5
2022 Impact Makes a Sound and Sound Makes an Impact: Sound Guides Representations and Explorations
abstract
Sound is one of the most informative and abundant modalities in the real world while being robust to sense without contacts by small and cheap sensors that can be placed on mobile devices. Although deep learning is capable of extracting information from multiple sensory inputs, there has been little use of sound for the control and learning of robotic actions. For unsupervised reinforcement learning, an agent is expected to actively collect experiences and jointly learn representations and policies in a self-supervised way. We build realistic robotic manipulation scenarios with physics-based sound simulation and propose the Intrinsic Sound Curiosity Module (ISCM). The ISCM provides feedback to a reinforcement learner to learn robust representations and to reward a more efficient exploration behavior. We perform experiments with sound enabled during pre-training and disabled during adaptation, and show that representations learned by ISCM outperform the ones by vision-only baselines and pre-trained policies can accelerate the learning process when applied to downstream tasks.
Xufeng Zhao 0002, Cornelius Weber, Muhammad Burhan Hafez, Stefan Wermter
IROS4
2022 Conversational Analysis of Daily Dialog Data using Polite Emotional Dialogue Acts
abstract
Many socio-linguistic cues are used in conversational analysis, such as emotion, sentiment, and dialogue acts. One of the fundamental social cues is politeness, which linguistically possesses properties such as social manners useful in conversational analysis. This article presents findings of polite emotional dialogue act associations, where we can correlate the relationships between the socio-linguistic cues. We confirm our hypothesis that the utterances with the emotion classes Anger and Disgust are more likely to be impolite. At the same time, Happiness and Sadness are more likely to be polite. A less expectable phenomenon occurs with dialogue acts Inform and Commissive which contain more polite utterances than Question and Directive. Finally, we conclude on the future work of these findings to extend the learning of social behaviours using politeness.
Chandrakant Bothe, Stefan Wermter
LREC2
2022 A Multimodal German Dataset for Automatic Lip Reading Systems and Transfer Learning
abstract
Large datasets as required for deep learning of lip reading do not exist in many languages. In this paper we present the dataset GLips (German Lips) consisting of 250,000 publicly available videos of the faces of speakers of the Hessian Parliament, which was processed for word-level lip reading using an automatic pipeline. The format is similar to that of the English language LRW (Lip Reading in the Wild) dataset, with each video encoding one word of interest in a context of 1.16 seconds duration, which yields compatibility for studying transfer learning between both datasets. By training a deep neural network, we investigate whether lip reading has language-independent features, so that datasets of different languages can be used to improve lip reading models. We demonstrate learning from scratch and show that transfer learning from LRW to GLips and vice versa improves learning speed and performance, in particular for the validation set.
Gerald Schwiebert, Cornelius Weber, Leyuan Qu, Henrique Siqueira, Stefan Wermter
LREC5
2022 Explain yourself! Effects of Explanations in Human-Robot Interaction
abstract
Recent developments in explainable artificial intelligence promise the potential to transform human-robot interaction: Explanations of robot decisions could affect user perceptions, justify their reliability, and increase trust. However, the effects on human perceptions of robots that explain their decisions have not been studied thoroughly. To analyze the effect of explainable robots, we conduct a study in which two simulated robots play a competitive board game. While one robot explains its moves, the other robot only announces them. Providing explanations for its actions was not sufficient to change the perceived competence, intelligence, likeability or safety ratings of the robot. However, the results show that the robot that explains its moves is perceived as more lively and human-like. This study demonstrates the need for and potential of explainable human-robot interaction and the wider assessment of its effects as a novel research direction.
Jakob Ambsdorf, Alina Munir, Yiyao Wei, Klaas Degkwitz, Harm Matthias Harms, Susanne Stannek, Kyra Ahrens, Dennis Becker, Erik Strahl, Tom Weber, Stefan Wermter
RO-MAN11
2022 CharacterGAN: Few-Shot Keypoint Character Animation and Reposing
abstract
We introduce CharacterGAN, a generative model that can be trained on only a few samples (8 – 15) of a given character. Our model generates novel poses based on keypoint locations, which can be modified in real time while providing interactive feedback, allowing for intuitive reposing and animation. Since we only have very limited training samples, one of the key challenges lies in how to address (dis)occlusions, e.g. when a hand moves behind or in front of a body. To address this, we introduce a novel layering approach which explicitly splits the input keypoints into different layers which are processed independently. These layers represent different parts of the character and provide a strong implicit bias that helps to obtain realistic results even with strong (dis)occlusions. To combine the features of individual layers we use an adaptive scaling approach conditioned on all keypoints. Finally, we introduce a mask connectivity constraint to reduce distortion artifacts that occur with extreme out-of-distribution poses at test time. We show that our approach outperforms recent baselines and creates realistic animations for diverse characters. We also show that our model can handle discrete state changes, for example a profile facing left or right, that the different layers do indeed learn features specific for the respective keypoints in those layers, and that our model scales to larger datasets when more data is available. Code is available at https://github.com/tohinz/CharacterGAN.
Tobias Hinz, Matthew Fisher, Oliver Wang, Eli Shechtman, Stefan Wermter
WACV5
2022 Go ahead and do not forget: Modular lifelong learning from event-based data
abstract
Lifelong learning is a long-standing aim for artificial agents that act in dynamic environments in which an agent needs to accumulate knowledge incrementally without forgetting previously learned representations. Contemporary methods for incremental learning from images are predominantly based on frame-based data recorded by conventional shutter cameras. We investigate methods for learning from data produced by event cameras and compare techniques to mitigate forgetting while learning incrementally. We propose a model that is composed of both, feature extraction and incremental learning. The feature extractor is utilized as a self-supervised sparse convolutional neural network that processes event-based data. The incremental learner uses a habituation-based method that works in tandem with other existing techniques. Our experimental results show that the combination of different existing techniques with our proposed habituation-based method can help avoid catastrophic forgetting even more, while learning incrementally from the features provided by the extraction module.
Vadym Gryshchuk, Cornelius Weber, Chu Kiong Loo, Stefan Wermter
Neurocomputing4
2022 Semantic Object Accuracy for Generative Text-to-Image Synthesis
abstract
Generative adversarial networks conditioned on textual image descriptions are capable of generating realistic-looking images. However, current methods still struggle to generate images based on complex image captions from a heterogeneous domain. Furthermore, quantitatively evaluating these text-to-image models is challenging, as most evaluation metrics only judge image quality but not the conformity between the image and its caption. To address these challenges we introduce a new model that explicitly models individual objects within an image and a new evaluation metric called Semantic Object Accuracy (SOA) that specifically evaluates images given an image caption. The SOA uses a pre-trained object detector to evaluate if a generated image contains objects that are mentioned in the image caption, e.g., whether an image generated from "a car driving down the street" contains a car. We perform a user study comparing several text-to-image models and show that our SOA metric ranks the models the same way as humans, whereas other metrics such as the Inception Score do not. Our evaluation also shows that models which explicitly model objects outperform models which only model global image characteristics.
Tobias Hinz, Stefan Heinrich, Stefan Wermter
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 Adapting the Interplay Between Personalized and Generalized Affect Recognition Based on an Unsupervised Neural Framework
abstract
Recent emotion recognition models, most of them being based on strongly supervised deep learning solutions, are rather successful in recognizing instantaneous emotion expressions. However, when applied to continuous interactions, these models show a weaker adaptation to a person-specific and long-term emotion appraisal. In this article, we present an unsupervised neural framework that improves emotion recognition by learning how to describe continuous affective behavior of individual persons. Our framework is composed of three self-organizing mechanisms: (1) a recurrent growing layer to cluster general emotion expressions, (2) a set of associative layers, acting asaffective memoriesto model specific emotional behavior of individual persons, (3) and an online learning layer which provides contextual modeling of continuous emotion expressions. We propose different learning strategies to integrate all three mechanisms and to improve the performance on arousal and valence recognition of the OMG-Emotion dataset. We evaluate our model with a series of experiments ranging from ablation studies assessing the different contributions of each neural component to an objective comparison with state-of-the-art solutions. The results from the evaluations show a good performance on emotion recognition of continuous emotions on monologue videos. Furthermore, we discuss how the model self-regulates the interplay between generalized and personalized emotion perception and how this influences the model’s reliability when recognizing unseen emotion expressions.
Pablo V. A. Barros, Emilia I. Barakova, Stefan Wermter
IEEE Trans. Affect. Comput.3
2021 Hearing Faces: Target Speaker Text-to-Speech Synthesis from a Face
abstract
The existence of a learnable cross-modal association between a person's face and their voice is recently becoming more and more evident. This provides the basis for the task of target speaker text-to-speech (TTS) synthesis from face ref-erence. In this paper, we approach this task by proposing a cross-modal model architecture combining existing unimodal models. We use Tacotron 2 multi-speaker TTS with auditory speaker embeddings based on Global Style Tokens. We trans-fer learn a FaceNet face encoder to predict these embeddings from a static face image reference instead of a voice reference and thus predict a speaker's voice and speaking characteristics from their face. Compared to Face2Speech, the only existing work on this task, we use a more modular architecture that allows the use of openly available and pretrained model components. This approach enables high-quality speech synthesis and allows for an easily extensible model architecture. Ex-perimental results show good matching ability while retaining better voice naturalness than Face2Speech. We examine the limitations of our model and discuss multiple possible av-enues of improvement for future work.
Björn Plüster, Cornelius Weber, Leyuan Qu, Stefan Wermter
ASRU4
2021 Lifelong Learning from Event-based Data
abstract
Lifelong learning is a long-standing aim for artificial agents that act in dynamic environments, in which an agent needs to accumulate knowledge incrementally without forgetting previously learned representations.We investigate methods for learning from data produced by event cameras and compare techniques to mitigate forgetting while learning incrementally.We propose a model that is composed of both, feature extraction and continuous learning.Furthermore, we introduce a habituationbased method to mitigate forgetting.Our experimental results show that the combination of different techniques can help to avoid catastrophic forgetting while learning incrementally from the features provided by the extraction module.
Vadym Gryshchuk, Cornelius Weber, Chu Kiong Loo, Stefan Wermter
ESANN4
2021 Pruning Neural Networks with Supermasks
abstract
The Lottery Ticket hypothesis by Frankle and Carbin states that a randomly initialized dense network contains a smaller subnetwork that, when trained in isolation, will match the performance of the original network.However, identifying this pruned subnetwork usually requires repeated training to determine optimal pruning thresholds.We present a novel approach to accelerate the pruning: By methodically evaluating different Supermasks, the threshold for selecting neurons as part of a pruned Lottery Ticket network can be determined without additional training.We evaluate the method on the MNIST dataset and achieve a size reduction of over 60% without a drop in performance.
Vincent Rolfs, Matthias Kerzel, Stefan Wermter
ESANN3
2021 DRILL: Dynamic Representations for Imbalanced Lifelong Learning
Kyra Ahrens, Fares Abawi, Stefan Wermter
ICANN (2)3
2021 FaVoA: Face-Voice Association Favours Ambiguous Speaker Detection
Hugo C. C. Carneiro, Cornelius Weber, Stefan Wermter
ICANN (1)3
2021 Variational Autoencoder for Speech Enhancement with a Noise-Aware Encoder
abstract
Recently, a generative variational autoencoder (VAE) has been proposed for speech enhancement to model speech statistics. However, this approach only uses clean speech in the training phase, making the estimation particularly sensitive to noise presence, especially in low signal-to-noise ratios (SNRs). To increase the robustness of the VAE, we propose to include noise information in the training phase by using a noise-aware encoder trained on noisy-clean speech pairs. We evaluate our approach on real recordings of different noisy environments and acoustic conditions using two different noise datasets. We show that our proposed noise-aware VAE outperforms the standard VAE in terms of overall distortion without increasing the number of model parameters. At the same time, we demonstrate that our model is capable of generalizing to unseen noise conditions better than a supervised feedforward deep neural network (DNN). Furthermore, we demonstrate the robustness of the model performance to a reduction of the noisy-clean speech training data size.
Huajian Fang, Guillaume Carbajal, Stefan Wermter, Timo Gerkmann
ICASSP3
2021 Visual Distant Supervision for Scene Graph Generation
abstract
Scene graph generation aims to identify objects and their relations in images, providing structured image representations that can facilitate numerous applications in computer vision. However, scene graph models usually require supervised learning on large quantities of labeled data with intensive human annotation. In this work, we propose visual distant supervision, a novel paradigm of visual relation learning, which can train scene graph models without any human-labeled data. The intuition is that by aligning commonsense knowledge bases and images, we can automatically create large-scale labeled data to provide distant supervision for visual relation learning. To alleviate the noise in distantly labeled data, we further propose a framework that iteratively estimates the probabilistic relation labels and eliminates the noisy ones. Comprehensive experimental results show that our distantly supervised model outperforms strong weakly supervised and semi-supervised baselines. By further incorporating human-labeled data in a semi-supervised fashion, our model outperforms state-of-the-art fully supervised models by a large margin (e.g., 8.3 micro- and 7.8 macro-recall@50 improvements for predicate classification in Visual Genome evaluation). We make the data and code for this paper publicly available at https://github.com/thunlp/VisualDS.
Yuan Yao 0011, Xu Han 0007, Mengdi Li 0006, Cornelius Weber, Zhiyuan Liu 0001, Stefan Wermter, Maosong Sun 0001
ICCV7
2021 GASP: Gated Attention for Saliency Prediction
abstract
Saliency prediction refers to the computational task of modeling overt attention. Social cues greatly influence our attention, consequently altering our eye movements and behavior. To emphasize the efficacy of such features, we present a neural model for integrating social cues and weighting their influences. Our model consists of two stages. During the first stage, we detect two social cues by following gaze, estimating gaze direction, and recognizing affect. These features are then transformed into spatiotemporal maps through image processing operations. The transformed representations are propagated to the second stage (GASP) where we explore various techniques of late fusion for integrating social cues and introduce two sub-networks for directing attention to relevant stimuli. Our experiments indicate that fusion approaches achieve better results for static integration methods, whereas non-fusion approaches for which the influence of each modality is unknown, result in better outcomes when coupled with recurrent models for dynamic saliency prediction. We show that gaze direction and affective representations contribute a prediction to ground-truth correspondence improvement of at least 5% compared to dynamic saliency models without social cues. Furthermore, affective representations improve GASP, supporting the necessity of considering affect-biased attention in predicting saliency.
Fares Abawi, Tom Weber, Stefan Wermter
IJCAI3
2021 Generalization in Multimodal Language Learning from Simulation
abstract
Neural networks can be powerful function approximators, which are able to model high-dimensional feature distributions from a subset of examples drawn from the target distribution. Naturally, they perform well at generalizing within the limits of their target function, but they often fail to generalize outside of the explicitly learned feature space. It is therefore an open research topic whether and how neural network-based architectures can be deployed for systematic reasoning. Many studies have shown evidence for poor generalization, but they often work with abstract data or are limited to single-channel input. Humans, however, learn and interact through a combination of multiple sensory modalities, and rarely rely on just one. To investigate compositional generalization in a multimodal setting, we generate an extensible dataset with multimodal input sequences from simulation. We investigate the influence of the underlying training data distribution on compostional generalization in a minimal LSTM-based network trained in a supervised, time continuous setting. We find compositional generalization to fail in simple setups while improving with the number of objects, actions, and particularly with a lot of color overlaps between objects. Furthermore, multimodality strongly improves compositional generalization in settings where a pure vision model struggles to generalize.
Aaron Eisermann, Jae Hee Lee 0001, Cornelius Weber, Stefan Wermter
IJCNN4
2021 Controlling the Noise Robustness of End-to-End Automatic Speech Recognition Systems
abstract
In this work, we propose a novel training scheme to modularize end-to-end systems. Our training scheme aims at altering the flow of information in an end-to-end system to use the kernels of this system for another system that fulfills another task. We apply this scheme to extract the noise reduction capabilities from a noise-robust automatic speech recognition (ASR) system and implement a speech enhancer from it. This enhancer receives spectral representations from unfiltered audio and outputs cleaned spectral representations. Our enhancer can be integrated into an ASR system as front-end, is trainable, and reduces background noise. Our front-end uses a decoder to clean speech based on the hidden activations of the ASR system Jasper. While training, we exclusively adapt the weights in our decoder and the batch normalization in Jasper. The resulting spectral representations show less background noise. Further, areas in the spectral features are not reconstructed if they do not contribute to speech recognition. We demonstrate that our front-end can be combined with a pre-trained ASR system as back-end and supports speech recognition in noisy conditions. Further, we show that training another ASR system with our front-end results in an increased performance of the ASR system in noisy as well as noiseless conditions. The ASR system's performance is especially improved on challenging speech datasets.
Matthias Möller, Johannes Twiefel, Cornelius Weber, Stefan Wermter
IJCNN4
2021 Improving Model-Based Reinforcement Learning with Internal State Representations through Self-Supervision
abstract
Using a model of the environment, reinforcement learning agents can plan their future moves and achieve superhuman performance in board games like Chess, Shogi, and Go, while remaining relatively sample-efficient. As demonstrated by the MuZero Algorithm, the environment model can even be learned dynamically, generalizing the agent to many more tasks while at the same time achieving state-of-the-art performance. Notably, MuZero uses internal state representations derived from real environment states for its predictions. In this paper, we bind the model's predicted internal state representation to the environment state via two additional terms: a reconstruction model loss and a simpler consistency loss, both of which work independently and unsupervised, acting as constraints to stabilize the learning process. Our experiments show that this new integration of reconstruction model loss and simpler consistency loss provide a significant performance increase in OpenAI Gym environments. Our modifications also enable self-supervised pretraining for MuZero, so the algorithm can learn about environment dynamics before a goal is made available.
Julien Scholz, Cornelius Weber, Muhammad Burhan Hafez, Stefan Wermter
IJCNN4
2021 Planning-integrated Policy for Efficient Reinforcement Learning in Sparse-reward Environments
abstract
Model-free reinforcement learning algorithms can learn an optimal policy from experience without requiring prior knowledge. However, model-free agents require vast amounts of samples, particularly in sparse reward environments where most states contain zero rewards. We developed a model-based approach to tackle the high sample complexity problem in sparse reward settings with continuous actions. A trained world model is queried by a particle swarm optimization (PSO) planner and employed as the action selection mechanism, hence taking the role of the actor in an actor-critic architecture. Parameters of the PSO regulate the agent's exploration rate. We show that the planner aids the agent to discover rewards even in regions with zero value gradient. Our simple planning integrated policy architecture learns more efficiently with fewer samples than continuous model-free algorithms.
Christoper Wulur, Cornelius Weber, Stefan Wermter
IJCNN3
2021 Behavior Self-Organization Supports Task Inference for Continual Robot Learning
abstract
Recent advances in robot learning have enabled robots to become increasingly better at mastering a predefined set of tasks. On the other hand, as humans, we have the ability to learn a growing set of tasks over our lifetime. Continual robot learning is an emerging research direction with the goal of endowing robots with this ability. In order to learn new tasks over time, the robot first needs to infer the task at hand. Task inference, however, has received little attention in the multi-task learning literature. In this paper, we propose a novel approach to continual learning of robotic control tasks. Our approach performs unsupervised learning of behavior embeddings by incrementally self-organizing demonstrated behaviors. Task inference is made by finding the nearest behavior embedding to a demonstrated behavior, which is used together with the environment state as input to a multi-task policy trained with reinforcement learning to optimize performance over tasks. Unlike previous approaches, our approach makes no assumptions about task distribution and requires no task exploration to infer tasks. We evaluate our approach in experiments with concurrently and sequentially presented tasks and show that it outperforms other multi-task learning approaches in terms of generalization performance and convergence speed, particularly in the continual learning setting.
Muhammad Burhan Hafez, Stefan Wermter
IROS2
2021 Robotic Occlusion Reasoning for Efficient Object Existence Prediction
abstract
Reasoning about potential occlusions is essential for robots to efficiently predict whether an object exists in an environment. Though existing work shows that a robot with active perception can achieve various tasks, it is still unclear if occlusion reasoning can be achieved. To answer this question, we introduce the task of robotic object existence prediction: when being asked about an object, a robot needs to move as few steps as possible around a table with randomly placed objects to predict whether the queried object exists. To address this problem, we propose a novel recurrent neural network model that can be jointly trained with supervised and reinforcement learning methods using a curriculum training strategy. Experimental results show that 1) both active perception and occlusion reasoning are necessary to successfully achieve the task; 2) the proposed model demonstrates a good occlusion reasoning ability by achieving a similar prediction accuracy to an exhaustive exploration baseline while requiring only about 10% of the baseline’s number of movement steps on average; and 3) the model generalizes to novel object combinations with a moderate loss of accuracy.
Mengdi Li 0006, Cornelius Weber, Matthias Kerzel, Jae Hee Lee 0001, Zheni Zeng, Zhiyuan Liu 0001, Stefan Wermter
IROS7
2021 Improved Techniques for Training Single-Image GANs
abstract
Recently there has been an interest in the potential of learning generative models from a single image, as opposed to from a large dataset. This task is of significance, as it means that generative models can be used in domains where collecting a large dataset is not feasible. However, training a model capable of generating realistic images from only a single sample is a difficult problem. In this work, we conduct a number of experiments to understand the challenges of training these methods and propose some best practices that we found allowed us to generate improved results over previous work. One key piece is that, unlike prior single image generation methods, we concurrently train several stages in a sequential multi-stage manner, allowing us to learn models with fewer stages of increasing image resolution. Compared to a recent state of the art baseline, our model is up to six times faster to train, has fewer parameters, and can better capture the global structure of images.
Tobias Hinz, Matthew Fisher, Oliver Wang, Stefan Wermter
WACV4
2021 Solving visual object ambiguities when pointing: an unsupervised learning approach
abstract
Abstract Whenever we are addressing a specific object or refer to a certain spatial location, we are using referential or deictic gestures usually accompanied by some verbal description. Particularly, pointing gestures are necessary to dissolve ambiguities in a scene and they are of crucial importance when verbal communication may fail due to environmental conditions or when two persons simply do not speak the same language. With the currently increasing advances of humanoid robots and their future integration in domestic domains, the development of gesture interfaces complementing human–robot interaction scenarios is of substantial interest. The implementation of an intuitive gesture scenario is still challenging because both the pointing intention and the corresponding object have to be correctly recognized in real time. The demand increases when considering pointing gestures in a cluttered environment, as is the case in households. Also, humans perform pointing in many different ways and those variations have to be captured. Research in this field often proposes a set of geometrical computations which do not scale well with the number of gestures and objects and use specific markers or a predefined set of pointing directions. In this paper, we propose an unsupervised learning approach to model the distribution of pointing gestures using a growing-when-required (GWR) network. We introduce an interaction scenario with a humanoid robot and define the so-called ambiguity classes. Our implementation for the hand and object detection is independent of any markers or skeleton models; thus, it can be easily reproduced. Our evaluation comparing a baseline computer vision approach with our GWR model shows that the pointing-object association is well learned even in cases of ambiguities resulting from close object proximity.
Doreen Jirak, David Biertimpel, Matthias Kerzel, Stefan Wermter
Neural Comput. Appl.4
2021 Special Issue on Automated Perception of Human Affect from Longitudinal Behavioral Data
abstract
The papers in this special section are aimed at contributions from computational neuroscience and psychology, artificial intelligence, machine learning, and affective computing, challenging and expanding current research on interpretation and estimation of human affective behavior from longitudinal data, i.e., single or multiple modalities captured over extended periods of time allowing efficient representation of behavior and inference in terms of affect and other socio-cognitive dimensions.
Pablo V. A. Barros, Stefan Wermter, Ognjen Rudovic, Hatice Gunes
IEEE Trans. Affect. Comput.2
2021 Guest Editorial: Special Issue on Deep Representation and Transfer Learning for Smart and Connected Health
abstract
Deep neural networks (NNs) have been proved to be efficient learning systems for supervised and unsupervised tasks. However, learning complex data representations using deep NNs can be difficult due to problems such as lack of data, exploding or vanishing gradients, high computational cost, or incorrect parameter initialization, among others. Deep representation and transfer learning (RTL) can facilitate the learning of data representations by taking advantage of transferable features learned by an NN model in a source domain, and adapting the model to a new domain.
Vasile Palade, Stefan Wermter, Ariel Ruiz-Garcia, Antônio de Pádua Braga, Clive Cheong Took
IEEE Trans. Neural Networks Learn. Syst.2
2020 Efficient Facial Feature Learning with Wide Ensemble-Based Convolutional Neural Networks
abstract
Ensemble methods, traditionally built with independently trained de-correlated models, have proven to be efficient methods for reducing the remaining residual generalization error, which results in robust and accurate methods for real-world applications. In the context of deep learning, however, training an ensemble of deep networks is costly and generates high redundancy which is inefficient. In this paper, we present experiments on Ensembles with Shared Representations (ESRs) based on convolutional networks to demonstrate, quantitatively and qualitatively, their data processing efficiency and scalability to large-scale datasets of facial expressions. We show that redundancy and computational load can be dramatically reduced by varying the branching level of the ESR without loss of diversity and generalization power, which are both important for ensemble performance. Experiments on large-scale datasets suggest that ESRs reduce the remaining residual generalization error on the AffectNet and FER+ datasets, reach human-level performance, and outperform state-of-the-art methods on facial expression recognition in the wild using emotion and affect concepts.
Henrique Siqueira, Sven Magg, Stefan Wermter
AAAI3
2020 New Results on Sparse Autoencoders for Posture Classification and Segmentation
Doreen Jirak, Stefan Wermter
ESANN2
2020 Self-Organizing Kernel-based Convolutional Echo State Network for Human Actions Recognition
Gin Chong Lee, Chu Kiong Loo, Wei Shiung Liew, Stefan Wermter
ESANN4
2020 Exploring Human-Robot Trust Through the Investment Game: An Immersive Space Mission Scenario
abstract
As robots become more advanced and capable, developing trust is an important factor of human-robot interaction and cooperation. However, as multiple environmental and social factors can influence trust, it is important to develop more elaborate scenarios and methods to measure human-robot trust. A widely used measurement of trust in social science is the investment game. In this study, we propose a scaled-up, immersive, science fiction Human-Robot Interaction (HRI) scenario for intrinsic motivation on human-robot collaboration, built upon the investment game and aimed at adapting the investment game for human-robot trust. For this purpose, we utilise two Neuro-Inspired Companion (NICO) - robots and a projected scenery. We investigate the applicability of our space mission experiment design to measure trust and the impact of non-verbal communication. We observe a correlation of 0.43 (p=0.02)between self-assessed trust and trust measured from the game and a positive impact of non-verbal communication on trust (p=0.0008) and robot perception for anthropomorphism (p=0.007) and animacy (p=0.00002). We conclude that our scenario is an appropriate method to measure trust in human-robot interaction and also to study how non-verbal communication influences a human's trust in robots.
Emy Arts, Sebastian Zörner, Kavish Bhatia, Glareh Mir, Florian Schmalzl, Ankit Srivastava, Brenda Vasiljevic, Tayfun Alpay, Annika Peters, Erik Strahl, Stefan Wermter
HAI11
2020 Variational Autoencoder with Global- and Medium Timescale Auxiliaries for Emotion Recognition from Speech
Hussam Almotlak, Cornelius Weber, Leyuan Qu, Stefan Wermter
ICANN (1)4
2020 Neuro-Genetic Visuomotor Architecture for Robotic Grasping
Matthias Kerzel, Josua Spisak, Erik Strahl, Stefan Wermter
ICANN (2)4
2020 Neural Networks for Detecting Irrelevant Questions During Visual Question Answering
Mengdi Li 0006, Cornelius Weber, Stefan Wermter
ICANN (2)3
2020 Curious Hierarchical Actor-Critic Reinforcement Learning
Frank Röder, Manfred Eppe, Phuong D. H. Nguyen, Stefan Wermter
ICANN (2)4
2020 Tell Me Why You Feel That Way: Processing Compositional Dependency for Tree-LSTM Aspect Sentiment Triplet Extraction (TASTE)
Alexander Sutherland, Suna Bensch, Thomas Hellström, Sven Magg, Stefan Wermter
ICANN (1)5
2020 Multimodal Target Speech Separation with Voice and Face References
abstract
Target speech separation refers to isolating target speech from a multi-speaker mixture signal by conditioning on auxiliary information about the target speaker. Different from the mainstream audio-visual approaches which usually require simultaneous visual streams as additional input, e.g. the corresponding lip movement sequences, in our approach we propose the novel use of a single face profile of the target speaker to separate expected clean speech. We exploit the fact that the image of a face contains information about the person's speech sound. Compared to using a simultaneous visual sequence, a face image is easier to obtain by pre-enrollment or on websites, which enables the system to generalize to devices without cameras. To this end, we incorporate face embeddings extracted from a pretrained model for face recognition into the speech separation, which guide the system in predicting a target speaker mask in the time-frequency domain. The experimental results show that a pre-enrolled face image is able to benefit separating expected speech signals. Additionally, face information is complementary to voice reference and we show that further improvement can be achieved when combing both face and voice embeddings.
Leyuan Qu, Cornelius Weber, Stefan Wermter
INTERSPEECH3
2020 EDA: Enriching Emotional Dialogue Acts using an Ensemble of Neural Annotators
abstract
The recognition of emotion and dialogue acts enriches conversational analysis and help to build natural dialogue systems. Emotion interpretation makes us understand feelings and dialogue acts reflect the intentions and performative functions in the utterances. However, most of the textual and multi-modal conversational emotion corpora contain only emotion labels but not dialogue acts. To address this problem, we propose to use a pool of various recurrent neural models trained on a dialogue act corpus, with and without context. These neural models annotate the emotion corpora with dialogue act labels, and an ensemble annotator extracts the final dialogue act label. We annotated two accessible multi-modal emotion corpora: IEMOCAP and MELD. We analyzed the co-occurrence of emotion and dialogue act labels and discovered specific relations. For example, Accept/Agree dialogue acts often occur with the Joy emotion, Apology with Sadness, and Thanking with Joy. We make the Emotional Dialogue Acts (EDA) corpus publicly available to the research community for further study and analysis.
Chandrakant Bothe, Cornelius Weber, Sven Magg, Stefan Wermter
LREC4
2020 Model Mediated Teleoperation with a Hand-Arm Exoskeleton in Long Time Delays Using Reinforcement Learning
abstract
Telerobotic systems must adapt to new environmental conditions and deal with high uncertainty caused by long-time delays. As one of the best alternatives to human-level intelligence, Reinforcement Learning (RL) may offer a solution to cope with these issues. This paper proposes to integrate RL with the Model Mediated Teleoperation (MMT) concept. The teleoperator interacts with a simulated virtual environment, which provides instant feedback. Whereas feedback from the real environment is delayed, feedback from the model is instantaneous, leading to high transparency. The MMT is realized in combination with an intelligent system with two layers. The first layer utilizes Dynamic Movement Primitives (DMP) which accounts for certain changes in the avatar environment. And, the second layer addresses the problems caused by uncertainty in the model using RL methods. Augmented reality was also provided to fuse the avatar device and virtual environment models for the teleoperator. Implemented on DLR's Exodex Adam hand-arm haptic exoskeleton, the results show RL methods are able to find different solutions when changes are applied to the object position after the demonstration. The results also show DMPs to be effective at adapting to new conditions where there is no uncertainty involved.
Hadi Beik-Mohammadi, Matthias Kerzel, Benedikt Pleintinger, Thomas Hulin, Philipp Reisich, Annika Schmidt, Aaron Pereira, Stefan Wermter, Neal Y. Lii
RO-MAN8
2020 Understanding auditory representations of emotional expressions with neural networks
abstract
In contrast to many established emotion recognition systems, convolutional neural networks do not rely on handcrafted features to categorize emotions. Although achieving state-of-the-art performances, it is still not fully understood what these networks learn and how the learned representations correlate with the emotional characteristics of speech. The aim of this work is to contribute to a deeper understanding of the acoustic and prosodic features that are relevant for the perception of emotional states. Firstly, an artificial deep neural network architecture is proposed that learns the auditory features directly from the raw and unprocessed speech signal. Secondly, we introduce two novel methods for the analysis of the implicitly learned representations based on data-driven and network-driven visualization techniques. Using these methods, we identify how the network categorizes an audio signal as a two-dimensional representation of emotions, namely valence and arousal. The proposed approach is a general method to enable a deeper analysis and understanding of the most relevant representations to perceive emotional expressions in speech.
Iris Wieser, Pablo V. A. Barros, Stefan Heinrich, Stefan Wermter
Neural Comput. Appl.4
2019 The OMG-Empathy Dataset: Evaluating the Impact of Affective Behavior in Storytelling
abstract
Processing human affective behavior is important for developing intelligent agents that interact with humans in complex interaction scenarios. A large number of current approaches that address this problem focus on classifying emotion expressions by grouping them into known categories. Such strategies neglect, among other aspects, the impact of the affective responses from an individual on their interaction partner thus ignoring how people empathize towards each other. This is also reflected in the datasets used to train models for affective processing tasks. Most of the recent datasets, in particular, the ones which capture natural interactions (“in-the-wild” datasets), are designed, collected, and annotated based on the recognition of displayed affective reactions, ignoring how these displayed or expressed emotions are perceived. In this paper, we propose a novel dataset composed of dyadic interactions designed, collected and annotated with a focus on measuring the affective impact that eight different stories have on the listener. Each video of the dataset contains around 5 minutes of interaction where a speaker tells a story to a listener. After each interaction, the listener annotated, using a valence scale, how the story impacted their affective state, reflecting how they empathized with the speaker as well as the story. We also propose different evaluation protocols and a baseline that encourages participation in the advancement of the field of artificial empathy and emotion contagion.
Pablo V. A. Barros, Nikhil Churamani, Angelica Lim, Stefan Wermter
ACII4
2019 Facial Expression Editing with Continuous Emotion Labels
abstract
Recently deep generative models have achieved impressive results in the field of automated facial expression editing. However, the approaches presented so far presume a discrete representation of human emotions and are therefore limited in the modelling of non-discrete emotional expressions. To overcome this limitation, we explore how continuous emotion representations can be used to control automated expression editing. We propose a deep generative model that can be used to manipulate facial expressions in facial images according to continuous two-dimensional emotion labels. One dimension represents an emotion's valence, the other represents its degree of arousal. We demonstrate the functionality of our model with a quantitative analysis using classifier networks as well as with a qualitative analysis.
Alexandra Lindt, Pablo V. A. Barros, Henrique Siqueira, Stefan Wermter
FG4
2019 Mixed-Reality Deep Reinforcement Learning for a Reach-to-grasp Task
Hadi Beik-Mohammadi, Mohammad-Ali Zamani, Matthias Kerzel, Stefan Wermter
ICANN (1)4
2019 Evaluating Defensive Distillation for Defending Text Processing Neural Networks Against Adversarial Examples
Marcus Soll, Tobias Hinz, Sven Magg, Stefan Wermter
ICANN (3)4
2019 Generating Multiple Objects at Spatially Distinct Locations
Tobias Hinz, Stefan Heinrich, Stefan Wermter
ICLR (Poster)3
2019 A Personalized Affective Memory Model for Improving Emotion Recognition
abstract
Recent models of emotion recognition strongly rely on supervised deep learning solutions for the distinction of general emotion expressions. However, they are not reliable when recognizing online and personalized facial expressions, e.g., for person-specific affective understanding. In this paper, we present a neural model based on a conditional adversarial autoencoder to learn how to represent and edit general emotion expressions. We then propose Grow-When-Required networks as personalized affective memories to learn individualized aspects of emotional expressions. Our model achieves state-of-the-art performance on emotion recognition when evaluated on in-the-wild datasets. Furthermore, our experiments include ablation studies and neural visualizations in order to explain the behavior of our model.
Pablo V. A. Barros, German Ignacio Parisi, Stefan Wermter
ICML3
2019 Incorporating End-to-End Speech Recognition Models for Sentiment Analysis
abstract
Previous work on emotion recognition demonstrated a synergistic effect of combining several modalities such as auditory, visual, and transcribed text to estimate the affective state of a speaker. Among these, the linguistic modality is crucial for the evaluation of an expressed emotion. However, manually transcribed spoken text cannot be given as input to a system practically. We argue that using ground-truth transcriptions during training and evaluation phases leads to a significant discrepancy in performance compared to real-world conditions, as the spoken text has to be recognized on the fly and can contain speech recognition mistakes. In this paper, we propose a method of integrating an automatic speech recognition (ASR) output with a character-level recurrent neural network for sentiment recognition. In addition, we conduct several experiments investigating sentiment recognition for human-robot interaction in a noise-realistic scenario which is challenging for the ASR systems. We quantify the improvement compared to using only the acoustic modality in sentiment recognition. We demonstrate the effectiveness of this approach on the Multimodal Corpus of Sentiment Intensity (MOSI) by achieving 73,6% accuracy in a binary sentiment classification task, exceeding previously reported results that use only acoustic input. In addition, we set a new state-of-the-art performance on the MOSI dataset (80.4% accuracy, 2% absolute improvement).
Egor Lakomkin, Mohammad-Ali Zamani, Cornelius Weber, Sven Magg, Stefan Wermter
ICRA5
2019 Designing a Personality-Driven Robot for a Human-Robot Interaction Scenario
abstract
In this paper, we present an autonomous AI system designed for a Human-Robot Interaction (HRI) study, set around a dice game scenario. We conduct a case study to answer our research question: Does a robot with a socially engaged personality lead to a higher acceptance than a competitive personality? The flexibility of our proposed system allows us to construct and attribute two different personalities to a humanoid robot: a socially engaged personality that maximizes its user interaction and a competitive personality that is focused on playing and winning the game. We evaluate both personalities in a user study, in which the participants play a turn-taking dice game with the robot. Each personality is assessed with four different evaluation tools: 1) the Godspeed Questionnaire, 2) the Mind Perception Questionnaire, 3) a custom questionnaire concerning the overall HRI experience, and 4) a Convolutional Neural Network analyzing the emotions on the participants' facial feedback throughout the game. Our results show that the socially engaged personality evokes stronger emotions among the participants and is rated higher in likability and animacy than the competitive one. We conclude that designing the robot with a socially engaged personality contributes to a higher acceptance within an HRI scenario.
Hadi Beik-Mohammadi, Nikoletta Xirakia, Fares Abawi, Irina Barykina, Krishnan Chandran, Gitanjali Nair, Daniel Speck, Tayfun Alpay, Sascha S. Griffiths, Stefan Heinrich, Erik Strahl, Cornelius Weber, Stefan Wermter
ICRA14
2019 Question Answering with Hierarchical Attention Networks
abstract
We investigate hierarchical attention networks for the task of question answering. For this purpose, we propose two different approaches: in the first, a document vector representation is built hierarchically from word-to-sentence level which is then used to infer the right answer. In the second, pointer sum attention is utilized to directly infer an answer from the attention values of the word and sentence representations. We evaluate our approach on the Children's Book Test, a cloze-style question answering dataset, and analyze the generated attention distributions. Our results show that, although a hierarchical approach does not offer much improvement over a shallow baseline, it does indeed offer a large performance boost when combining word and sentence attention with pointer sum attention.
Tayfun Alpay, Stefan Heinrich, Michael Nelskamp, Stefan Wermter
IJCNN4
2019 Curious Meta-Controller: Adaptive Alternation between Model-Based and Model-Free Control in Deep Reinforcement Learning
abstract
Recent success in deep reinforcement learning for continuous control has been dominated by model-free approaches which, unlike model-based approaches, do not suffer from representational limitations in making assumptions about the world dynamics and model errors inevitable in complex domains. However, they require a lot of experiences compared to model-based approaches that are typically more sample-efficient. We propose to combine the benefits of the two approaches by presenting an integrated approach called Curious Meta-Controller. Our approach alternates adaptively between model-based and model-free control using a curiosity feedback based on the learning progress of a neural model of the dynamics in a learned latent space. We demonstrate that our approach can significantly improve the sample efficiency and achieve near-optimal performance on learning robotic reaching and grasping tasks from raw-pixel input in both dense and sparse reward settings.
Muhammad Burhan Hafez, Cornelius Weber, Matthias Kerzel, Stefan Wermter
IJCNN4
2019 Neuro-Robotic Haptic Object Classification by Active Exploration on a Novel Dataset
abstract
We present an embodied neural model for haptic object classification by active haptic exploration with the humanoid robot NICO. When NICO's newly developed robotic hand closes around an object, multiple sensory readings from a tactile fingertip sensor, motor positions, and motor currents are recorded. We created a haptic dataset with 83200 haptic measurements, based on 100 samples of each of 16 different objects, every sample containing 52 measurements. First, we provide an analysis of neural classification models with regard to isolated haptic sensory channels for object classification. Based on this, we develop a series of neural models (MLP, CNN, LSTM) that integrate the haptic sensory channels to classify explored objects. As an initial baseline, our best model achieves a 66.6% classification accuracy over 16 objects. We show that this result is due to the ability of the network to integrate the haptic data both over time domain and over different haptic sensory channels. Furthermore, we make the dataset publically available to address the issue of sparse haptic datasets for machine learning research.
Matthias Kerzel, Erik Strahl, Connor Gaede, Emil Gasanov, Stefan Wermter
IJCNN5
2019 Effect of Pruning on Catastrophic Forgetting in Growing Dual Memory Networks
abstract
Grow-when-required networks such as the Growing Dual-Memory (GDM) networks possess a dynamic network structure, expanding to accommodate new neurons in response to learning novel concepts. Over time, it may be necessary to prune obsolete neurons and/or neural connections to meet performance or resource limitations. GDM networks utilize an age-based pruning strategy, whereby older neurons and neural connections that have not been activated recently are removed. Catastrophic forgetting occurs when knowledge learned by the networks in previous learning iterations is lost due to being overwritten by newer learning iterations, or to the pruning process. In this work, we investigate catastrophic forgetting in GDM networks in response to different pruning strategies. The age-based pruning method was shown to significantly sparsify the GDM network topology while improving the networks ability to recall newly acquired concepts with a slight decrease in performance with respect to older knowledge. A significance-based pruning method was tested as a replacement for the age-based pruning, but was not as effective at pruning even though it performed better at recalling older knowledge.
Wei Shiung Liew, Chu Kiong Loo, Vadym Gryshchuk, Cornelius Weber, Stefan Wermter
IJCNN5
2019 The Conditional Boundary Equilibrium Generative Adversarial Network and its Application to Facial Attributes
abstract
We propose an extension of the Boundary Equilibrium GAN (BEGAN) neural network, named Conditional BEGAN (CBEGAN), as a general generative and transformational approach for data processing. As a novelty, the system is able of both data generation and transformation under conditional input. We evaluate our approach for conditional image generation and editing using five controllable attributes for images of faces from the CelebA dataset: age, smiling, cheekbones, eyeglasses and gender. We perform a set of objective quantitative experiments to evaluate the model's performance and a qualitative user study to evaluate how humans assess the generated and edited images. Both evaluations yield coinciding results which show that the generated facial attributes are recognizable in more than 80% of all new testing samples.
Ahmed Marzouk, Pablo V. A. Barros, Manfred Eppe, Stefan Wermter
IJCNN4
2019 Leveraging Recursive Processing for Neural-Symbolic Affect-Target Associations
abstract
Explaining the outcome of deep learning decisions based on affect is challenging but necessary if we expect social companion robots to interact with users on an emotional level. In this paper, we present a commonsense approach that utilizes an interpretable hybrid neural-symbolic system to associate extracted targets, noun chunks determined to be associated with the expressed emotion, with affective labels from a natural language expression. We leverage a pre-trained neural network that is well adapted to tree and sub-tree processing, the Dependency Tree-LSTM, to learn the affect labels of dynamic targets, determined through symbolic rules, in natural language. We find that making use of the unique properties of the recursive network provides higher accuracy and interpretability when compared to other unstructured and sequential methods for determining target-affect associations in an aspect-based sentiment analysis task.
Alexander Sutherland, Sven Magg, Stefan Wermter
IJCNN3
2019 LipSound: Neural Mel-Spectrogram Reconstruction for Lip Reading
Leyuan Qu, Cornelius Weber, Stefan Wermter
INTERSPEECH3
2019 Predictive Auxiliary Variational Autoencoder for Representation Learning of Global Speech Characteristics
Sebastian Springenberg, Egor Lakomkin, Cornelius Weber, Stefan Wermter
INTERSPEECH4
2019 Exploring Low-level and High-level Transfer Learning for Multi-task Facial Recognition with a Semi-supervised Neural Network
abstract
Facial recognition tasks like identity, age, gender, and emotion recognition received substantial attention in recent years. Their deployment in robotic platforms became necessary for the characterization of most of the non-verbal Human-Robot Interaction (HRI) scenarios. In this regard, deep convolution neural networks have shown to be effective on processing different facial representations but with a high cost: to achieve maximum generalization, they require an enormous amount of task-specific labeled data. This paper proposes a unified semi-supervised deep neural model to address this problem. Our hybrid model is composed of an unsupervised deep generative adversarial network which learns fundamental characteristics of facial representations, and a set of convolution channels that fine-tunes the high-level facial concepts for the recognition of identity, age group, gender, and facial expressions. Our network employs progressive lateral connections between the convolution channels so that they share the high-abstraction particularities of each of these tasks in order to reduce the necessity of a large amount of strongly labeled training data. We propose a series of experiments to evaluate each individual mechanism of our hybrid model, in particular, the impact of the progressive connections on learning the specific facial recognition tasks and we observe that our model achieves a better performance when compared to task-specific models.
Pablo V. A. Barros, Erik Fließwasser, Matthias Kerzel, Stefan Wermter
IROS4
2019 A Kernel Bayesian Adaptive Resonance Theory with A Topological Structure
abstract
This paper attempts to solve the typical problems of self-organizing growing network models, i.e. (a) an influence of the order of input data on the self-organizing ability, (b) an instability to high-dimensional data and an excessive sensitivity to noise, and (c) an expensive computational cost by integrating Kernel Bayes Rule (KBR) and Correntropy-Induced Metric (CIM) into Adaptive Resonance Theory (ART) framework. KBR performs a covariance-free Bayesian computation which is able to maintain a fast and stable computation. CIM is a generalized similarity measurement which can maintain a high-noise reduction ability even in a high-dimensional space. In addition, a Growing Neural Gas (GNG)-based topology construction process is integrated into the ART framework to enhance its self-organizing ability. The simulation experiments with synthetic and real-world datasets show that the proposed model has an outstanding stable self-organizing ability for various test environments.
Naoki Masuyama, Chu Kiong Loo, Stefan Wermter
Int. J. Neural Syst.3
2019 Preserving activations in recurrent neural networks based on surprisal
abstract
Learning hierarchical abstractions from sequences is a challenging and open problem for recurrent neural networks (RNNs). This is mainly due to the difficulty of detecting features that span over long time distances with also different frequencies. In this paper, we address this challenge by introducing surprisal-based activation, a novel method to preserve activations and skip updates depending on encoding-based information content. The preserved activations can be considered as temporal shortcuts with perfect memory. We present a preliminary analysis by evaluating surprisal-based activation on language modeling with the Penn Treebank corpus and find that it can improve performance when compared to baseline RNNs and Long Short-Term Memory (LSTM) networks.
Tayfun Alpay, Fares Abawi, Stefan Wermter
Neurocomputing3
2019 Continuous convolutional object tracking in developmental robot scenarios
abstract
Tracking arbitrary objects in natural environments is a challenging task in visual computing. A central problem is the need to adapt to changing appearances under strong transformation and occlusion. We propose a tracking framework that utilises the strength of Convolutional Neural Networks to create a robust and adaptive model of the object from training data produced during tracking. An incremental update mechanism provides increased performance and reduces the computational costs for training during tracking, allowing for robust real-time tracking with state-of-the-art performance. Together with optimisations for deploying the framework on humanoid robots and distributed devices, this shows its viability for research in developmental robotics on questions around infant cognition or active exploration.
Stefan Heinrich, Peer Springstübe, Tobias Knöppler, Matthias Kerzel, Stefan Wermter
Neurocomputing5
2019 Localizing salient body motion in multi-person scenes using convolutional neural networks
abstract
With modern computer vision techniques being successfully developed for a variety of tasks, extracting meaningful knowledge from complex scenes with multiple people still poses problems. Consequently, experiments with application-specific motion, such as gesture recognition scenarios, are often constrained to single person scenes in the literature. Therefore, in this paper we address the challenging task of detecting salient body motion in scenes with more than one person. We propose a neural architecture that only reacts to a specific kind of motion in the scene: A limited set of body gestures. The model is trained end-to-end, thereby avoiding hand-crafted features and the strong reliance on pre-processing as it is prevalent in similar studies. The presented model implements a saliency mechanism that reacts to body motion cues which have not been included in previous computational saliency systems. Our architecture consists of a 3D Convolutional Neural Network that receives a frame sequence as its input and localizes active gesture movement. To train our network with a large data variety, we introduce an approach to combine Kinect recordings of one person into artificial scenes with multiple people, yielding a large diversity of scene configurations in our dataset. We performed experiments using these sequences and show that the proposed model is able to localize the salient body motion of our gesture set. We found that 3D convolutions and a baseline model with 2D convolutions perform surprisingly similar on our task. Our experiments revealed the influence of gesture characteristics on how well they can be learned by our model. Given a distinct gesture set and computational restrictions, we conclude that using 2D convolutions might often perform equally well.
Florian Letsch, Doreen Jirak, Stefan Wermter
Neurocomputing3
2019 Continual lifelong learning with neural networks: A review
abstract
Humans and animals have the ability to continually acquire, fine-tune, and transfer knowledge and skills throughout their lifespan. This ability, referred to as lifelong learning, is mediated by a rich set of neurocognitive mechanisms that together contribute to the development and specialization of our sensorimotor skills as well as to long-term memory consolidation and retrieval. Consequently, lifelong learning capabilities are crucial for computational learning systems and autonomous agents interacting in the real world and processing continuous streams of information. However, lifelong learning remains a long-standing challenge for machine learning and neural network models since the continual acquisition of incrementally available information from non-stationary data distributions generally leads to catastrophic forgetting or interference. This limitation represents a major drawback for state-of-the-art deep neural network models that typically learn representations from stationary batches of training data, thus without accounting for situations in which information becomes incrementally available over time. In this review, we critically summarize the main challenges linked to lifelong learning for artificial learning systems and compare existing neural network approaches that alleviate, to different extents, catastrophic forgetting. Although significant advances have been made in domain-specific learning with neural networks, extensive research efforts are required for the development of robust lifelong learning on autonomous agents and robots. We discuss well-established and emerging research motivated by lifelong learning factors in biological systems such as structural plasticity, memory replay, curriculum and transfer learning, intrinsic motivation, and multisensory integration.
German Ignacio Parisi, Ronald Kemker, Jose L. Part, Christopher Kanan, Stefan Wermter
Neural Networks5
2019 Enhanced Robot Speech Recognition Using Biomimetic Binaural Sound Source Localization
abstract
Inspired by the behavior of humans talking in noisy environments, we propose an embodied embedded cognition approach to improve automatic speech recognition (ASR) systems for robots in challenging environments, such as with ego noise, using binaural sound source localization (SSL). The approach is verified by measuring the impact of SSL with a humanoid robot head on the performance of an ASR system. More specifically, a robot orients itself toward the angle where the signal-to-noise ratio (SNR) of speech is maximized for one microphone before doing an ASR task. First, a spiking neural network inspired by the midbrain auditory system based on our previous work is applied to calculate the sound signal angle. Then, a feedforward neural network is used to handle high levels of ego noise and reverberation in the signal. Finally, the sound signal is fed into an ASR system. For ASR, we use a system developed by our group and compare its performance with and without the support from SSL. We test our SSL and ASR systems on two humanoid platforms with different structural and material properties. With our approach we halve the sentence error rate with respect to the common downmixing of both channels. Surprisingly, the ASR performance is more than two times better when the angle between the humanoid head and the sound source allows sound waves to be reflected most intensely from the pinna to the ear microphone, rather than when sound waves arrive perpendicularly to the membrane.
Jorge Dávila-Chacón, Stefan Wermter
IEEE Trans. Neural Networks Learn. Syst.3
2018 Surprisal-based activation in recurrent neural networks
Tayfun Alpay, Fares Abawi, Stefan Wermter
ESANN3
2018 An analysis of subtask-dependency in robot command interpretation with dilated CNNs
Manfred Eppe, Tayfun Alpay, Fares Abawi, Stefan Wermter
ESANN4
2018 Slowness-based neural visuomotor control with an Intrinsically motivated Continuous Actor-Critic
Muhammad Burhan Hafez, Matthias Kerzel, Cornelius Weber, Stefan Wermter
ESANN4
2018 Inferencing based on unsupervised learning of disentangled representations
Tobias Hinz, Stefan Wermter
ESANN2
2018 A Sub-Layered Hierarchical Pyramidal Neural Architecture for Facial Expression Recognition
Henrique Siqueira, Pablo V. A. Barros, Sven Magg, Cornelius Weber, Stefan Wermter
ESANN5
2018 Image-to-Text Transduction with Spatial Self-Attention
Sebastian Springenberg, Egor Lakomkin, Cornelius Weber, Stefan Wermter
ESANN4
2018 Continuous convolutional object tracking
Peer Springstübe, Stefan Heinrich, Stefan Wermter
ESANN3
2018 Towards End-to-End Raw Audio Music Synthesis
Manfred Eppe, Tayfun Alpay, Stefan Wermter
ICANN (3)3
2018 Classification of MRI Migraine Medical Data Using 3D Convolutional Neural Network
Hwei Geok Ng, Matthias Kerzel, Jan Mehnert, Arne May, Stefan Wermter
ICANN (3)5
2018 Combining Articulatory Features with End-to-End Learning in Speech Recognition
Leyuan Qu, Cornelius Weber, Egor Lakomkin, Johannes Twiefel, Stefan Wermter
ICANN (3)5
2018 De-noise-GAN: De-noising Images to Improve RoboCup Soccer Ball Detection
Daniel Speck, Pablo V. A. Barros, Stefan Wermter
ICANN (3)3
2018 A Hybrid Planning Strategy Through Learning from Vision for Target-Directed Navigation
Xiaomao Zhou, Cornelius Weber, Chandrakant Bothe, Stefan Wermter
ICANN (2)4
2018 EmoRL: Continuous Acoustic Emotion Classification Using Deep Reinforcement Learning
abstract
Acoustically expressed emotions can make communication with a robot more efficient. Detecting emotions like anger could provide a clue for the robot indicating unsafe/undesired situations. Recently, several deep neural network-based models have been proposed which establish new state-of-the-art results in affective state evaluation. These models typically start processing at the end of each utterance, which not only requires a mechanism to detect the end of an utterance but also makes it difficult to use them in a real-time communication scenario, e.g. human-robot interaction. We propose the EmoRL model that triggers an emotion classification as soon as it gains enough confidence while listening to a person speaking. As a result, we minimize the need for segmenting the audio signal for classification and achieve lower latency as the audio signal is processed incrementally. The method is competitive with the accuracy of a strong baseline model, while allowing much earlier prediction.
Egor Lakomkin, Mohammad-Ali Zamani, Cornelius Weber, Sven Magg, Stefan Wermter
ICRA5
2018 Multi-modal Feedback for Affordance-driven Interactive Reinforcement Learning
abstract
Interactive reinforcement learning (IRL) extends traditional reinforcement learning (RL) by allowing an agent to interact with parent-like trainers during a task. In this paper, we present an IRL approach using dynamic audio-visual input in terms of vocal commands and hand gestures as feedback. Our architecture integrates multi-modal information to provide robust commands from multiple sensory cues along with a confidence value indicating the trustworthiness of the feedback. The integration process also considers the case in which the two modalities convey incongruent information. Additionally, we modulate the influence of sensory-driven feedback in the IRL task using goal-oriented knowledge in terms of contextual affordances. We implement a neural network architecture to predict the effect of performed actions with different objects to avoid failed-states, i.e., states from which it is not possible to accomplish the task. In our experimental setup, we explore the interplay of multi-modal feedback and task-specific affordances in a robot cleaning scenario. We compare the learning performance of the agent under four different conditions: traditional RL, multi-modal IRL, and each of these two setups with the use of contextual affordances. Our experiments show that the best performance is obtained by using audio-visual feedback with affordance-modulated IRL. The obtained results demonstrate the importance of multi-modal sensory processing integrated with goal-oriented knowledge in IRL tasks.
Francisco Cruz 0002, German Ignacio Parisi, Stefan Wermter
IJCNN3
2018 The OMG-Emotion Behavior Dataset
abstract
This paper is the basis paper for the accepted IJCNN challenge One-Minute Gradual-Emotion Recognition (OMG-Emotion)1by which we hope to foster long-emotion classification using neural models for the benefit of the IJCNN community. The proposed corpus has as novelty the data collection and annotation strategy based on emotion expressions which evolve over time into a specific context. Different from other corpora, we propose a novel multimodal corpus for emotion expression recognition, which uses gradual annotations with a focus on contextual emotion expressions. Our dataset was collected from Youtube videos using a specific search strategy based on restricted keywords and filtering which guaranteed that the data follow a gradual emotion expression transition, i.e. emotion expressions evolve over time in a natural and continuous fashion. We also provide an experimental protocol and a series of unimodal baseline experiments which can be used to evaluate deep and recurrent neural models in a fair and standard manner.
Pablo V. A. Barros, Nikhil Churamani, Egor Lakomkin, Henrique Siqueira, Alexander Sutherland, Stefan Wermter
IJCNN6
2018 Expectation Learning and Crossmodal Modulation with a Deep Adversarial Network
abstract
The human brain is able to learn, generalize, and predict crossmodal stimuli which help us to understand the world around us. Some characteristics of crossmodal learning inspired some computational models but most of the solutions only go as far as to implement strategies for early or late crossmodal fusion. In this paper, we propose the use of two mechanisms from behavioral psychology to enhance the capabilities of a deep adversarial network to learn crossmodal stimuli: the unity assumption modulation and expectation learning. We use real-world data to train and evaluate our model in a set of experiments and demonstrate how these mechanisms affect the learning behavior of the model and how they contribute to making it learn crossmodal coincident stimuli. Our experiments show that the addition of these two mechanisms modulates the crossmodal binding capabilities of the model and improves the learning of unisensory descriptors.
Pablo V. A. Barros, German Ignacio Parisi, Di Fu, Xun Liu 0001, Stefan Wermter
IJCNN5
2018 Learning Empathy-Driven Emotion Expressions using Affective Modulations
abstract
Human-Robot Interaction (HRI) studies, particularly the ones designed around social robots, use emotions as important building blocks for interaction design. In order to provide a natural interaction experience, these social robots need to recognise the emotions expressed by the users across various modalities of communication and use them to estimate an internal affective model of the interaction. These internal emotions act as motivation for learning to respond to the user in different situations, using the physical capabilities of the robot. This paper proposes a deep hybrid neural model for multi-modal affect recognition, analysis and behaviour modelling in social robots. The model uses growing self-organising network models to encode intrinsic affective states for the robot. These intrinsic states are used to train a reinforcement learning model to learn facial expression representations on the Neuro-Inspired Companion (NICO) robot, enabling the robot to express empathy towards the users.
Nikhil Churamani, Pablo V. A. Barros, Erik Strahl, Stefan Wermter
IJCNN4
2018 Image Generation and Translation with Disentangled Representations
abstract
Generative models have made significant progress in the tasks of modeling complex data distributions such as natural images. The introduction of Generative Adversarial Networks (GANs) and auto-encoders lead to the possibility of training on big data sets in an unsupervised manner. However, for many generative models it is not possible to specify what kind of image should be generated and it is not possible to translate existing images into new images of similar domains. Furthermore, models that can perform image-to-image translation often need distinct models for each domain, making it hard to scale these systems to multiple domain image-to-image translation. We introduce a model that can do both, controllable image generation and image-to-image translation between multiple domains. We split our image representation into two parts encoding unstructured and structured information respectively. The latter is designed in a disentangled manner, so that different parts encode different image characteristics. We train an encoder to encode images into these representations and use a small amount of labeled data to specify what kind of information should be encoded in the disentangled part. A generator is trained to generate images from these representations using the characteristics provided by the disentangled part of the representation. Through this we can control what kind of images the generator generates, translate images between different domains, and even learn unknown data-generating factors while only using one single model.
Tobias Hinz, Stefan Wermter
IJCNN2
2018 Sparse Autoencoders for Posture Recognition
abstract
Among different gesture types, static gestures or postures deliver a broad range of communicative information like commands or emblems. Vision-based processing for posture recognition is the most intuitive yet challenging task in intelligent systems. Achievements in deep learning, specifically convolu- tional neural networks (CNN), replaced creating hand models or engineering features for automated image feature learning at the expense of large data requirements and long training sessions for optimal parameter tuning. The aim of the present study is to explore the potentials of sparse autoencoders for posture recognition, promoting an alternative method to present convolutional approaches. We conduct experiments with hierarchically designed autoencoders to retain the desired image feature abstractions on two posture datasets with distinct characteristics. The different data properties allow us to demonstrate parameter influences on the network performance. Our evaluation shows that even a shallow network design achieves superior perfor- mance compared to a multiple channel CNN, and comparable results on a small dataset with sparse image samples. From our study we conclude that “lightweight” approaches can be viable tools for posture recognition, which are worth more explorations in the future.
Doreen Jirak, Stefan Wermter
IJCNN2
2018 Accelerating Deep Continuous Reinforcement Learning through Task Simplification
abstract
Robotic motor policies can, in theory, be learned via deep continuous reinforcement learning. In practice, however, collecting the enormous amount of required training samples in realistic time, surpasses the possibilities of many robotic platforms. To address this problem, we propose a novel method for accelerating the learning process by task simplification inspired by the Goldilocks effect known from developmental psychology. We present results on a reach-for-grasp task that is learned with the Deep Deterministic Policy Gradients (DDPG) algorithm. Task simplification is realized by initially training the system with “larger-than-life” training objects that adapt their reachability dynamically during training. We achieve a significant acceleration compared to the unaltered training setup. We describe modifications to the DDPG algorithm with regard to the replay buffer to prevent artifacts during the learning process from the simplified learning instances while maintaining the speed of learning. With this result, we contribute towards the realistic application of deep reinforcement learning on robotic platforms.
Matthias Kerzel, Hadi Beik-Mohammadi, Mohammad-Ali Zamani, Stefan Wermter
IJCNN4
2018 Recognition and Prediction of Human-Object Interactions with a Self-Organizing Architecture
abstract
The recognition and prediction of human actions are challenging perception tasks that require reasoning upon a large space of fine-grained body motion patterns. In this work, we propose a hierarchical self-organizing architecture which jointly learns to recognize and predict human-object interactions from RGB-D videos. Our model consists of a hierarchy of Grow-When-Required (GWR) networks which process and learn co-occurring actions and objects from the training data. Our goal is to learn prototype body motion patterns when manipulating objects as well as to internally store prototype transitions of body postures over time. The architecture can generate sequences of arbitrary length given an observed initial motion pattern as well as predict future action labels. Experimental results on a dataset of daily activities demonstrate that our architecture recognizes ongoing actions and predicts the upcoming ones with high accuracy. The generated body pose trajectories demonstrate that our architecture is suitable to be further applied to the problem of the look-ahead planning of a robotic response in a human-robot interaction scenario.
Luiza Mici, German Ignacio Parisi, Stefan Wermter
IJCNN3
2018 Point Cloud Object Recognition using 3D Convolutional Neural Networks
abstract
With the advent of RGB-D technology, there was remarkable progress in robotic tasks such as object recognition. Many approaches were developed to handle depth information, but they work mainly on 2.5D representations of the data. Moreover, the 3D-data handling approaches using Convolutional Neural Networks developed so far showed a gap between volumetric CNN and multi-view CNN. Therefore, the use of point clouds for object recognition has not been fully explored. In this work, we propose a Convolutional Neural Network model that extracts 3D features directly from RGB-D data, mixing volumetric and multi-view representations. The neural architecture is kept as simple as possible to assess the benefits of the 3D-data easily. We evaluate our approach with the publicly available Washington Dataset of real RGB-D data composed of 51 categories of household objects and obtained an improvement of around 10% in accuracy over the utilisation of 2D features. This result motivates further investigation when compared to some recently reported results tested on smaller datasets.
Marcelo Borghetti Soares, Stefan Wermter
IJCNN2
2018 A Self-organizing Method for Robot Navigation based on Learned Place and Head-Direction Cells
abstract
This paper describes a neural model for a robot learning spatial knowledge and navigating on learned place and head-direction (HD) cell representations. The place and HD cells, which are trained through unsupervised slow feature analysis (SFA) from sequences of visual stimuli, provide positional and directional information for navigation. Based on the ensemble activity of place cells, the robot learns a topological map of the environment through extracting the statistical distribution of the place cell activities covering the traversable areas and realizes self-localization based on the map. The robot's heading direction, which is encoded by the HD cells, works as a control signal to adjust its behavior. Action representations supporting state transitions are learned through memorizing the same movement from a previous phase where an experimenter drives a robot to explore an environment. Given reward signals spreading from a target location along the topological map, the robot can reach the goal in a reward-ascending way. This work intends to build a practical navigation system by simulating animals' hippocampal cell firing activities on a robot platform using its self-contained sensor. Experimental results from simulation demonstrate that our system navigates a robot to the desired position smoothly and effectively.
Xiaomao Zhou, Cornelius Weber, Stefan Wermter
IJCNN3
2018 Conversational Analysis Using Utterance-level Attention-based Bidirectional Recurrent Neural Networks
abstract
Recent approaches for dialogue act recognition have shown that context from preceding utterances is important to classify the subsequent one. It was shown that the performance improves rapidly when the context is taken into account. We propose an utterance-level attention-based bidirectional recurrent neural network (Utt-Att-BiRNN) model to analyze the importance of preceding utterances to classify the current one. In our setup, the BiRNN is given the input set of current and preceding utterances. Our model outperforms previous models that use only preceding utterances as context on the used corpus. Another contribution of the article is to discover the amount of information in each utterance to classify the subsequent one and to show that context-based learning not only improves the performance but also achieves higher confidence in the classification. We use character- and word-level features to represent the utterances. The results are presented for character and word feature representations and as an ensemble model of both representations. We found that when classifying short utterances, the closest preceding utterances contributes to a higher degree.
Chandrakant Bothe, Sven Magg, Cornelius Weber, Stefan Wermter
INTERSPEECH4
2018 Deep Neural Object Analysis by Interactive Auditory Exploration with a Humanoid Robot
abstract
We present a novel approach for interactive auditory object analysis with a humanoid robot. The robot elicits sensory information by physically shaking visually indistinguishable plastic capsules. It gathers the resulting audio signals from microphones that are embedded into the robotic ears. A neural network architecture learns from these signals to analyze properties of the contents of the containers. Specifically, we evaluate the material classification and weight prediction accuracy and demonstrate that the framework is fairly robust to acoustic real-world noise.
Manfred Eppe, Matthias Kerzel, Erik Strahl, Stefan Wermter
IROS4
2018 Object Detection and Pose Estimation Based on Convolutional Neural Networks Trained with Synthetic Data
abstract
Instance-based object detection and fine pose estimation is an active research problem in computer vision. While the traditional interest-point-based approaches for pose estimation are precise, their applicability in robotic tasks relies on controlled environments and rigid objects with detailed textures. CNN-based approaches, on the other hand, have shown impressive results in uncontrolled environments for more general object recognition tasks like category-based coarse pose estimation, but the need of large datasets of fully-annotated training images makes them unfavourable for tasks like instance-based pose estimation. We present a novel approach that combines the robustness of CNNs with a fine-resolution instance-based 3D pose estimation, where the model is trained with fully-annotated synthetic training data, generated automatically from the 3D models of the objects. We propose an experimental setup in which we can carefully examine how the model trained with synthetic data performs on real images of the objects. Results show that the proposed model can be trained only with synthetic renderings of the objects' 3D models and still be successfully applied on images of the real objects, with precision suitable for robotic tasks like object grasping. Based on the results, we present more general insights about training neural models with synthetic images for application on real-world images.
Josip Josifovski, Matthias Kerzel, Christoph Pregizer, Lukas Posniak, Stefan Wermter
IROS5
2018 On the Robustness of Speech Emotion Recognition for Human-Robot Interaction with Deep Neural Networks
abstract
Speech emotion recognition (SER) is an important aspect of effective human-robot collaboration and received a lot of attention from the research community. For example, many neural network-based architectures were proposed recently and pushed the performance to a new level. However, the applicability of such neural SER models trained only on in-domain data to noisy conditions is currently under-researched. In this work, we evaluate the robustness of state-of-the-art neural acoustic emotion recognition models in human-robot interaction scenarios. We hypothesize that a robot's ego noise, room conditions, and various acoustic events that can occur in a home environment can significantly affect the performance of a model. We conduct several experiments on the iCub robot platform and propose several novel ways to reduce the gap between the model's performance during training and testing in real-world conditions. Furthermore, we observe large improvements in the model performance on the robot and demonstrate the necessity of introducing several data augmentation techniques like overlaying background noise and loudness variations to improve the robustness of the neural approaches.
Egor Lakomkin, Mohammad-Ali Zamani, Cornelius Weber, Sven Magg, Stefan Wermter
IROS5
2018 A Neurorobotic Experiment for Crossmodal Conflict Resolution in Complex Environments
abstract
Crossmodal conflict resolution is crucial for robot sensorimotor coupling through the interaction with the environment, yielding swift and robust behaviour also in noisy conditions. In this paper, we propose a neurorobotic experiment in which an iCub robot exhibits human-like responses in a complex crossmodal environment. To better understand how humans deal with multisensory conflicts, we conducted a behavioural study exposing 33 subjects to congruent and incongruent dynamic audio-visual cues. In contrast to previous studies using simplified stimuli, we designed a scenario with four animated avatars and observed that the magnitude and extension of the visual bias are related to the semantics embedded in the scene, i.e., visual cues that are congruent with environmental statistics (moving lips and vocalization) induce the strongest bias. We implement a deep learning model that processes stereophonic sound, facial features, and body motion to trigger a discrete behavioural response. After training the model, we exposed the iCub to the same experimental conditions as the human subjects, showing that the robot can replicate similar responses in real time. Our interdisciplinary work provides important insights into how crossmodal conflict resolution can be modelled in robots and introduces future research directions for the efficient combination of sensory observations with internally generated knowledge and expectations.
German Ignacio Parisi, Pablo V. A. Barros, Di Fu, Sven Magg, Haiyan Wu, Xun Liu 0001, Stefan Wermter
IROS7
2018 An Ensemble with Shared Representations Based on Convolutional Networks for Continually Learning Facial Expressions
abstract
Social robots able to continually learn facial expressions could progressively improve their emotion recognition capability towards people interacting with them. Semi-supervised learning through ensemble predictions is an efficient strategy to leverage the high exposure of unlabelled facial expressions during human-robot interactions. Traditional ensemble-based systems, however, are composed of several independent classifiers leading to a high degree of redundancy, and unnecessary allocation of computational resources. In this paper, we proposed an ensemble based on convolutional networks where the early layers are strong low-level feature extractors, and their representations shared with an ensemble of convolutional branches. This results in a significant drop in redundancy of low-level features processing. Training in a semi-supervised setting, we show that our approach is able to continually learn facial expressions through ensemble predictions using unlabelled samples from different data distributions.
Henrique Siqueira, Pablo V. A. Barros, Sven Magg, Stefan Wermter
IROS4
2018 Hear the Egg - Demonstrating Robotic Interactive Auditory Perception
abstract
We present an illustrative example of an interactive auditory perception approach performed by a humanoid robot called NICO, the Neuro Inspired COmpanion [1]. The video demonstrates a material classification task in the style of a classic TV game show. NICO and another candidate are supposed to determine the content of small plastic capsules that are visually indistinguishable. Shaking the capsules produces audio signals that range from rattling stones, over tinkling coins to swooshing sand. NICO can perceive and analyze these sounds to determine the material of the capsules content.
Erik Strahl, Matthias Kerzel, Manfred Eppe, Sascha S. Griffiths, Stefan Wermter
IROS5
2018 A Context-based Approach for Dialogue Act Recognition using Simple Recurrent Neural Networks
Chandrakant Bothe, Cornelius Weber, Sven Magg, Stefan Wermter
LREC4
2018 Improving interactive reinforcement learning: What makes a good teacher?
abstract
Interactive reinforcement learning (IRL) has become an important apprenticeship approach to speed up convergence in classic reinforcement learning (RL) problems. In this regard, a variant of IRL is policy shaping which uses a parent-like trainer to propose the next action to be performed and by doing so reduces the search space by advice. On some occasions, the trainer may be another artificial agent which in turn was trained using RL methods to afterward becoming an advisor for other learner-agents. In this work, we analyse internal representations and characteristics of artificial agents to determine which agent may outperform others to become a better trainer-agent. Using a polymath agent, as compared to a specialist agent, an advisor leads to a larger reward and faster convergence of the reward signal and also to a more stable behaviour in terms of the state visit frequency of the learner-agents. Moreover, we analyse system interaction parameters in order to determine how influential they are in the apprenticeship process, where the consistency of feedback is much more relevant when dealing with different learner obedience parameters.
Francisco Cruz 0002, Sven Magg, Yukie Nagai, Stefan Wermter
Connect. Sci.4
2018 Interactive natural language acquisition in a multi-modal recurrent neural architecture
abstract
For the complex human brain that enables us to communicate in natural language, we gathered good understandings of principles underlying language acquisition and processing, knowledge about sociocultural conditions, and insights into activity patterns in the brain. However, we were not yet able to understand the behavioural and mechanistic characteristics for natural language and how mechanisms in the brain allow to acquire and process language. In bridging the insights from behavioural psychology and neuroscience, the goal of this paper is to contribute a computational understanding of appropriate characteristics that favour language acquisition. Accordingly, we provide concepts and refinements in cognitive modelling regarding principles and mechanisms in the brain and propose a neurocognitively plausible model for embodied language acquisition from real-world interaction of a humanoid robot with its environment. In particular, the architecture consists of a continuous time recurrent neural network, where parts have different leakage characteristics and thus operate on multiple timescales for every modality and the association of the higher level nodes of all modalities into cell assemblies. The model is capable of learning language production grounded in both, temporal dynamic somatosensation and vision, and features hierarchical concept abstraction, concept decomposition, multi-modal integration, and self-organisation of latent representations.
Stefan Heinrich, Stefan Wermter
Connect. Sci.2
2018 Speeding up the Hyperparameter Optimization of Deep Convolutional Neural Networks
abstract
Most learning algorithms require the practitioner to manually set the values of many hyperparameters before the learning process can begin. However, with modern algorithms, the evaluation of a given hyperparameter setting can take a considerable amount of time and the search space is often very high-dimensional. We suggest using a lower-dimensional representation of the original data to quickly identify promising areas in the hyperparameter space. This information can then be used to initialize the optimization algorithm for the original, higher-dimensional data. We compare this approach with the standard procedure of optimizing the hyperparameters only on the original input. We perform experiments with various state-of-the-art hyperparameter optimization algorithms such as random search, the tree of parzen estimators (TPEs), sequential model-based algorithm configuration (SMAC), and a genetic algorithm (GA). Our experiments indicate that it is possible to speed up the optimization process by using lower-dimensional data representations at the beginning, while increasing the dimensionality of the input later in the optimization process. This is independent of the underlying optimization procedure, making the approach promising for many existing hyperparameter optimization algorithms.
Tobias Hinz, Nicolás Navarro-Guerrero, Sven Magg, Stefan Wermter
Int. J. Comput. Intell. Appl.4
2018 A self-organizing neural network architecture for learning human-object interactions
abstract
The visual recognition of transitive actions comprising human-object interactions is a key component for artificial systems operating in natural environments. This challenging task requires jointly the recognition of articulated body actions as well as the extraction of semantic elements from the scene such as the identity of the manipulated objects. In this paper, we present a self-organizing neural network for the recognition of human-object interactions from RGB-D videos. Our model consists of a hierarchy of Grow-When-Required (GWR) networks that learn prototypical representations of body motion patterns and objects, accounting for the development of action-object mappings in an unsupervised fashion. We report experimental results on a dataset of daily activities collected for the purpose of this study as well as on a publicly available benchmark dataset. In line with neurophysiological studies, our self-organizing architecture exhibits higher neural activation for congruent action-object pairs learned during training sessions with respect to synthetically created incongruent ones. We show that our unsupervised model shows competitive classification results on the benchmark dataset with respect to strictly supervised approaches.
Luiza Mici, German Ignacio Parisi, Stefan Wermter
Neurocomputing3
2017 The Impact of Personalisation on Human-Robot Interaction in Learning Scenarios
abstract
Advancements in Human-Robot Interaction involve robots being more responsive and adaptive to the human user they are interacting with. For example, robots model a personalised dialogue with humans, adapting the conversation to accommodate the user's preferences in order to allow natural interactions. This study investigates the impact of such personalised interaction capabilities of a human companion robot on its social acceptance, perceived intelligence and likeability in a human-robot interaction scenario. In order to measure this impact, the study makes use of an object learning scenario where the user teaches different objects to the robot using natural language. An interaction module is built on top of the learning scenario which engages the user in a personalised conversation before teaching the robot to recognise different objects. The two systems, i.e. with and without the interaction module, are compared with respect to how different users rate the robot on its intelligence and sociability. Although the system equipped with personalised interaction capabilities is rated lower on social acceptance, it is perceived as more intelligent and likeable by the users.
Nikhil Churamani, Paul Anton, Marc Brügger, Erik Fließwasser, Thomas Hummel 0001, Julius Mayer 0001, Waleed Mustafa, Hwei Geok Ng, Thi Linh Chi Nguyen, Quan Nguyen 0005, Marcus Soll, Sebastian Springenberg, Sascha S. Griffiths, Stefan Heinrich, Nicolás Navarro-Guerrero, Erik Strahl, Johannes Twiefel, Cornelius Weber, Stefan Wermter
HAI19
2017 Emotion Recognition from Body Expressions with a Neural Network Architecture
abstract
The recognition of emotions plays an important role in our daily life and is essential for social communication. Although multiple studies have shown that body expressions can strongly convey emotional states, emotion recognition from body motion patterns has received less attention than the use of facial expressions. In this paper, we propose a self-organizing neural architecture that can effectively recognize affective states from full-body motion patterns. To evaluate our system, we designed and collected a data corpus named the Body Expressions of Emotion (BEE) dataset using a depth sensor in a human-robot interaction scenario. For our recordings, nineteen participants were asked to perform six different emotions:anger, fear, happiness, neutral, sadness, and surprise. In order to compare our system with human-like performance, we conducted an additional experiment by asking fifteen annotators to label depth map video sequences as one of the six emotion classes. The labeling results from human annotators were compared to the results predicted by our system. Experimental results showed that the recognition accuracy of the system was competitive with human performance when exposed to body motion patterns from the same dataset.
Nourhan Elfaramawy, Pablo V. A. Barros, German Ignacio Parisi, Stefan Wermter
HAI4
2017 Comparison of Behaviour-Based Architectures for a Collaborative Package Delivery Task
abstract
A comparison between behavioural architectures, specifically a BDI architecture and a finite-state machine, for a collaborative package delivery system is presented. The system should assist a user in handling packages in cluttered environments. The entire system is built using open-source solutions for modules including speech recognition, person detection and tracking, and navigation. For the comparison, we use three criteria, namely, a static implementation-based comparison, a dynamic comparison and a qualitative comparison. Based on our results, we provide experimental evidence that supports the theoretical consensus about the domain of applicability for both, BDI architectures and finite-state machines. However, we cannot support or discourage any of the tested architectures for the particular case of the collaborative package delivery scenario, due to the non-overlapping strengths and weakness of both approaches. Finally, we outline future improvements to the system itself as well as the comparison of both behavioural architectures.
Melanie Remmels, Nicolás Navarro-Guerrero, Stefan Wermter
HAI3
2017 Dialogue-Based Neural Learning to Estimate the Sentiment of a Next Upcoming Utterance
Chandrakant Bothe, Sven Magg, Cornelius Weber, Stefan Wermter
ICANN (2)4
2017 Neural End-to-End Self-learning of Visuomotor Skills by Environment Interaction
Matthias Kerzel, Stefan Wermter
ICANN (1)2
2017 Semi-supervised Phoneme Recognition with Recurrent Ladder Networks
Marian Tietz, Tayfun Alpay, Johannes Twiefel, Stefan Wermter
ICANN (1)4
2017 Robot Localization and Orientation Detection Based on Place Cells and Head-Direction Cells
Xiaomao Zhou, Cornelius Weber, Stefan Wermter
ICANN (1)3
2017 Reusing Neural Speech Representations for Auditory Emotion Recognition
abstract
Acoustic emotion recognition aims to categorize the affective state of the speaker and is still a difficult task for machine learning models. The difficulties come from the scarcity of training data, general subjectivity in emotion perception resulting in low annotator agreement, and the uncertainty about which features are the most relevant and robust ones for classification. In this paper, we will tackle the latter problem. Inspired by the recent success of transfer learning methods we propose a set of architectures which utilize neural representations inferred by training on large speech databases for the acoustic emotion recognition task. Our experiments on the IEMOCAP dataset show ~10% relative improvements in the accuracy and F1-score over the baseline recurrent neural network which is trained end-to-end for emotion recognition.
Egor Lakomkin, Cornelius Weber, Sven Magg, Stefan Wermter
IJCNLP(1)4
2017 A self-organizing model for affective memory
abstract
Emotions are related to many different parts of our lives: from the perception of the environment around us to different learning processes and natural communication. Therefore, it is very hard to achieve an automatic emotion recognition system which is adaptable enough to be used in real-world scenarios. This paper proposes the use of a growing and self-organizing affective memory architecture to improve the adaptability of the Cross-channel Convolution Neural Network emotion recognition model. The architecture we propose, besides being adaptable to new subjects and scenarios also presents means to perceive and model human behavior in an unsupervised fashion enabling it to deal with never seen emotion expressions. We demonstrate in our experiments that the proposed model is competitive compared with the state-of-the-art approach, and how it can be used in different affective behavior analysis scenarios.
Pablo V. A. Barros, Stefan Wermter
IJCNN2
2017 Teaching emotion expressions to a human companion robot using deep neural architectures
abstract
Human companion robots need to be sociable and responsive towards emotions to better interact with the human environment they are expected to operate in. This paper is based on the Neuro-Inspired COmpanion robot (NICO) and investigates a hybrid, deep neural network model to teach the NICO to associate perceived emotions with expression representations using its on-board capabilities. The proposed model consists of a Convolutional Neural Network (CNN) and a Self-organising Map (SOM) to perceive the emotions expressed by a human user towards NICO and trains two parallel Multilayer Perceptron (MLP) networks to learn general as well as person-specific associations between perceived emotions and the robot's facial expressions.
Nikhil Churamani, Matthias Kerzel, Erik Strahl, Pablo V. A. Barros, Stefan Wermter
IJCNN5
2017 Haptic material classification with a multi-channel neural network
abstract
We present a novel approach for haptic material classification based on an adaptation of human haptic exploratory procedures executed by a robot arm with an optical force sensor. A multi-channel neural architecture informed by findings from human haptic perception performs a spectral analysis on vibration and texture data gathered during material exploration and integrates this analysis with information gathered on material compliance. Experimental results show a high classification accuracy on a test set of 32 common household materials. Furthermore, we show that haptic material properties, relevant for robot grasping, can be classified with a simple haptic exploration while actual material classification requires more complex exploration and computation.
Matthias Kerzel, Moaaz Ali, Hwei Geok Ng, Stefan Wermter
IJCNN4
2017 NICO - Neuro-inspired companion: A developmental humanoid robot platform for multimodal interaction
abstract
Interdisciplinary research, drawing from robotics, artificial intelligence, neuroscience, psychology, and cognitive science, is a cornerstone to advance the state-of-the-art in multimodal human-robot interaction and neuro-cognitive modeling. Research on neuro-cognitive models benefits from the embodiment of these models into physical, humanoid agents that possess complex, human-like sensorimotor capabilities for multimodal interaction with the real world. For this purpose, we develop and introduce NICO (Neuro-Inspired COmpanion), a humanoid developmental robot that fills a gap between necessary sensing and interaction capabilities and flexible design. This combination makes it a novel neuro-cognitive research platform for embodied sensorimotor computational and cognitive models in the context of multimodal interaction as shown in our results.
Matthias Kerzel, Erik Strahl, Sven Magg, Nicolás Navarro-Guerrero, Stefan Heinrich, Stefan Wermter
RO-MAN6
2017 Hey robot, why don't you talk to me?
abstract
This paper describes the techniques used in the submitted video presenting an interaction scenario, realised using the Neuro-Inspired Companion (NICO) robot. NICO engages the users in a personalised conversation where the robot always tracks the users' face, remembers them and interacts with them using natural language. NICO can also learn to perform tasks such as remembering and recalling objects and thus can assist users in their daily chores. The interaction system helps the users to interact as naturally as possible with the robot, enriching their experience with the robot, making it more interesting and engaging.
Hwei Geok Ng, Paul Anton, Marc Brügger, Nikhil Churamani, Erik Fließwasser, Thomas Hummel 0001, Julius Mayer 0001, Waleed Mustafa, Thi Linh Chi Nguyen, Quan Nguyen 0005, Marcus Soll, Sebastian Springenberg, Sascha S. Griffiths, Stefan Heinrich, Nicolás Navarro-Guerrero, Erik Strahl, Johannes Twiefel, Cornelius Weber, Stefan Wermter
RO-MAN19
2017 Emotion-modulated attention improves expression recognition: A deep learning model
abstract
Spatial attention in humans and animals involves the visual pathway and the superior colliculus, which integrate multimodal information. Recent research has shown that affective stimuli play an important role in attentional mechanisms, and behavioral studies show that the focus of attention in a given region of the visual field is increased when affective stimuli are present. This work proposes a neurocomputational model that learns to attend to emotional expressions and to modulate emotion recognition. Our model consists of a deep architecture which implements convolutional neural networks to learn the location of emotional expressions in a cluttered scene. We performed a number of experiments for detecting regions of interest, based on emotion stimuli, and show that the attention model improves emotion expression recognition when used as emotional attention modulator. Finally, we analyze the internal representations of the learned neural filters and discuss their role in the performance of our model.
Pablo V. A. Barros, German Ignacio Parisi, Cornelius Weber, Stefan Wermter
Neurocomputing4
2017 An analysis of Convolutional Long Short-Term Memory Recurrent Neural Networks for gesture recognition
abstract
In this research, we analyze a Convolutional Long Short-Term Memory Recurrent Neural Network (CNNLSTM) in the context of gesture recognition. CNNLSTMs are able to successfully learn gestures of varying duration and complexity. For this reason, we analyze the architecture by presenting a qualitative evaluation of the model, based on the visualization of the internal representations of the convolutional layers and on the examination of the temporal classification outputs at a frame level, in order to check if they match the cognitive perception of a gesture. We show that CNNLSTM learns the temporal evolution of the gestures classifying correctly their meaningful part, known as Kendon’s stroke phase. With the visualization, for which we use the deconvolution process that maps specific feature map activations to original image pixels, we show that the network learns to detect the most intense body motion. Finally, we show that CNNLSTM outperforms both plain CNN and LSTM in gesture recognition.
Eleni Tsironi, Pablo V. A. Barros, Cornelius Weber, Stefan Wermter
Neurocomputing4
2017 Lifelong learning of human actions with deep neural network self-organization
abstract
Lifelong learning is fundamental in autonomous robotics for the acquisition and fine-tuning of knowledge through experience. However, conventional deep neural models for action recognition from videos do not account for lifelong learning but rather learn a batch of training data with a predefined number of action classes and samples. Thus, there is the need to develop learning systems with the ability to incrementally process available perceptual cues and to adapt their responses over time. We propose a self-organizing neural architecture for incrementally learning to classify human actions from video sequences. The architecture comprises growing self-organizing networks equipped with recurrent neurons for processing time-varying patterns. We use a set of hierarchically arranged recurrent networks for the unsupervised learning of action representations with increasingly large spatiotemporal receptive fields. Lifelong learning is achieved in terms of prediction-driven neural dynamics in which the growth and the adaptation of the recurrent networks are driven by their capability to reconstruct temporally ordered input sequences. Experimental results on a classification task using two action benchmark datasets show that our model is competitive with state-of-the-art methods for batch learning also when a significant number of sample labels are missing or corrupted during training sessions. Additional experiments show the ability of our model to adapt to non-stationary input avoiding catastrophic interference.
German Ignacio Parisi, Jun Tani, Cornelius Weber, Stefan Wermter
Neural Networks4
2016 Learning contextual affordances with an associative neural architecture
Francisco Cruz 0002, German Ignacio Parisi, Stefan Wermter
ESANN3
2016 Activity recognition with echo state networks using 3D body joints and objects category
Luiza Mici, Xavier Hinaut, Stefan Wermter
ESANN3
2016 Gesture Recognition with a Convolutional Long Short-Term Memory Recurrent Neural Network
Eleni Tsironi, Pablo V. A. Barros, Stefan Wermter
ESANN3
2016 Semantic Role Labelling for Robot Instructions using Echo State Networks
Johannes Twiefel, Xavier Hinaut, Stefan Wermter
ESANN3
2016 Learning Multiple Timescales in Recurrent Neural Networks
Tayfun Alpay, Stefan Heinrich, Stefan Wermter
ICANN (1)3
2016 The Effects of Regularization on Learning Facial Expressions with Convolutional Neural Networks
Tobias Hinz, Pablo V. A. Barros, Stefan Wermter
ICANN (2)3
2016 Recognition of Transitive Actions with Hierarchical Neural Network Learning
Luiza Mici, German Ignacio Parisi, Stefan Wermter
ICANN (2)3
2016 Learning auditory neural representations for emotion recognition
abstract
Auditory emotion recognition has become a very important topic in recent years. However, still after the development of some architectures and frameworks, generalization is a big problem. Our model examines the capability of deep neural networks to learn specific features for different kinds of auditory emotion recognition: speech and music-based recognition. We propose the use of a cross-channel architecture to improve the generalization aspects of complex auditory recognition by the integration of previously learned knowledge of specific representation into a high-level auditory descriptor. We evaluate our models using the SAVEE dataset, the GTZAN dataset and the EmotiW corpus, and show comparable results with state-of-the-art approaches.
Pablo V. A. Barros, Cornelius Weber, Stefan Wermter
IJCNN3
2016 Understanding how deep neural networks learn face expressions
abstract
Deep neural networks have been used successfully for several different computer vision-related tasks, including facial expression recognition. In spite of the good results, it is still not clear why these networks achieve such good recognition rates. One way to learn more about deep neural networks is to visualise and understand what they are learning, and to do so techniques such as deconvolution could play a significant role. In this paper, we train a Convolutional Neural Network (CNN) and Lateral Inhibition Pyramidal Neural Network (LIPNet) to learn facial expressions. Then, we use the deconvolution process to visualise the learned features of the CNN and we introduce a novel mechanism for visualising the internal representation of the LIPNet. We perform a series of experiments, training our networks with the Cohn-Kanade data set and show what kind of facial structures compose the learned emotion expression representation. Then, we use the trained networks to recognise images from the Jaffe data set and demonstrate that the learned representations are present in different face images, emphasizing the generalization aspects of these networks. We discuss the different representations that each network learns and how they differ from each other. We also discuss how each learned representation contributes to the recognition process and how they can be compared to the emotional notation Facial Action Coding System - Facs. Finally, we explain how the principles of invariance, redundancy and filtering, common for deep networks, contribute to the learned features and to the facial expression recognition task in general.
Nima Mousavi, Henrique Siqueira, Pablo V. A. Barros, Bruno J. T. Fernandes, Stefan Wermter
IJCNN5
2016 Multi-modal integration of dynamic audiovisual patterns for an interactive reinforcement learning scenario
abstract
Robots in domestic environments are receiving more attention, especially in scenarios where they should interact with parent-like trainers for dynamically acquiring and refining knowledge. A prominent paradigm for dynamically learning new tasks has been reinforcement learning. However, due to excessive time needed for the learning process, a promising extension has been made by incorporating an external parent-like trainer into the learning cycle in order to scaffold and speed up the apprenticeship using advice about what actions should be performed for achieving a goal. In interactive reinforcement learning, different uni-modal control interfaces have been proposed that are often quite limited and do not take into account multiple sensor modalities. In this paper, we propose the integration of audiovisual patterns to provide advice to the agent using multi-modal information. In our approach, advice can be given using either speech, gestures, or a combination of both. We introduce a neural network-based approach to integrate multi-modal information from uni-modal modules based on their confidence. Results show that multi-modal integration leads to a better performance of interactive reinforcement learning with the robot being able to learn faster with greater rewards compared to uni-modal scenarios.
Francisco Cruz 0002, German Ignacio Parisi, Johannes Twiefel, Stefan Wermter
IROS4
2016 Human motion assessment in real time using recurrent self-organization
abstract
The correct execution of well-defined movements plays a crucial role in physical rehabilitation and sports. While there is an extensive number of well-established approaches for human action recognition, the task of assessing the quality of actions and providing feedback for correcting inaccurate movements has remained an open issue in the literature. We present a learning-based method for efficiently providing feedback on a set of training movements captured by a depth sensor. We propose a novel recursive neural network that uses growing self-organization for the efficient learning of body motion sequences. The quality of actions is then computed in terms of how much a performed movement matches the correct continuation of a learned sequence. The proposed system provides visual assistance to the person performing an exercise by displaying real-time feedback, thus enabling the user to correct inaccurate postures and motion intensity. We evaluate our approach with a data set containing 3 powerlifting exercises performed by 17 athletes. Experimental results show that our novel architecture outperforms our previous approach for the correct prediction of routines and the detection of mistakes both in a single- and multiple-subject scenario.
German Ignacio Parisi, Sven Magg, Stefan Wermter
RO-MAN3
2016 Using natural language feedback in a neuro-inspired integrated multimodal robotic architecture
abstract
In this paper we present a multi-modal human robot interaction architecture which is able to combine information coming from different sensory inputs, and can generate feedback for the user which helps to teach him/her implicitly how to interact with the robot. The system combines vision, speech and language with inference and feedback. The system environment consists of a Nao robot which has to learn objects situated on a table only by understanding absolute and relative object locations uttered by the user and afterwards points on a desired object to show what it has learned. The results of a user study and performance test show the usefulness of the feedback produced by the system and also justify the usage of the system in a real-world applications, as its classification accuracy of multi-modal input is around 80.8%. In the experiments, the system was able to detect inconsistent input coming from different sensory modules in all cases and could generate useful feedback for the user from this information.
Johannes Twiefel, Xavier Hinaut, Marcelo Borghetti Soares, Erik Strahl, Stefan Wermter
RO-MAN5
2016 Ball Localization for Robocup Soccer Using Convolutional Neural Networks
Daniel Speck, Pablo V. A. Barros, Cornelius Weber, Stefan Wermter
RoboCup4
2015 Dynamic gesture recognition using Echo State Networks
Doreen Jirak, Pablo V. A. Barros, Stefan Wermter
ESANN3
2015 Learning objects from RGB-D sensors using point cloud-based neural networks
Marcelo Borghetti Soares, Pablo V. A. Barros, German Ignacio Parisi, Stefan Wermter
ESANN4
2015 Interactive reinforcement learning through speech guidance in a domestic scenario
abstract
Recently robots are being used more frequently as assistants in domestic scenarios. In this context we train an apprentice robot to perform a cleaning task using interactive reinforcement learning since it has been shown to be an efficient learning approach benefiting from human expertise for performing domestic tasks. The robotic agent obtains interactive feedback via a speech recognition system which is tested to work with five different microphones concerning their polar patterns and distance to the teacher to recognize sentences in different instruction classes. Moreover, the reinforcement learning approach uses situated affordances to allow the robot to complete the cleaning task in every episode anticipating when chosen actions are possible to be performed. Situated affordances and interaction allow to improve the convergence speed of reinforcement learning, and the results also show that the system is robust against wrong instructions that result from errors of the speech recognition system.
Francisco Cruz 0002, Johannes Twiefel, Sven Magg, Cornelius Weber, Stefan Wermter
IJCNN5
2015 Face expression recognition with a 2-channel Convolutional Neural Network
abstract
A new architecture based on the Multi-channel Convolutional Neural Network (MCCNN) is proposed for recognizing facial expressions. Two hard-coded feature extractors are replaced by a single channel which is partially trained in an unsupervised fashion as a Convolutional Autoencoder (CAE). One additional channel that contains a standard CNN is left unchanged. Information from both channels converges in a fully connected layer and is then used for classification. We perform two distinct experiments on the JAFFE dataset (leave-one-out and ten-fold cross validation) to evaluate our architecture. Our comparison with the previous model that uses hard-coded Sobel features shows that an additional channel of information with unsupervised learning can significantly boost accuracy and reduce the overall training time. Furthermore, experimental results are compared with benchmarks from the literature showing that our method provides state-of-the-art recognition rates for facial expressions. Our method outperforms previously published methods that used hand-crafted features by a large margin.
Dennis Hamester, Pablo V. A. Barros, Stefan Wermter
IJCNN3
2015 Learning human motion feedback with neural self-organization
abstract
The correct execution of well-defined movements in sport disciplines may increase the body's mechanical efficiency and reduce the risk of injury. While there exists an extensive number of learning-based approaches for the recognition of human actions, the task of computing and providing feedback for correcting inaccurate movements has received significantly less attention in the literature. We present a learning system for automatically providing feedback on a set of learned movements captured with a depth sensor. The proposed system provides visual assistance to the person performing an exercise by displaying real-time feedback to correct possible inaccurate postures and motion. The learning architecture uses recursive neural network self-organization extended for predicting the correct continuation of the training movements. We introduce three mechanisms for computing feedback on the correctness of overall movement and individual body joints. For evaluation purposes, we collected a data set with 17 athletes performing 3 powerlifting exercises. Our results show promising system performance for the detection of mistakes in movements on this data set.
German Ignacio Parisi, Florian von Stosch, Sven Magg, Stefan Wermter
IJCNN4
2015 Recognizing complex mental states with deep hierarchical features for Human-Robot Interaction
abstract
The use of emotional states for Human-Robot Interaction (HRI) has attracted considerable attention in recent years. One of the most challenging tasks is to recognize the spontaneous expression of emotions, especially in an HRI scenario. Every person has a different way to express emotions, and this is aggravated by the complexity of interaction with different subjects, multimodal information and different environments. We propose a deep neural model which is able to deal with these characteristics and which is applied in recognition of complex mental states. Our system is able to learn and extract deep spatial and temporal features and to use them to classify emotions in sequences. To evaluate the system, the CAM3D corpus is used. This corpus is composed of videos recorded from different subjects and in different indoor environments. Each video contains the recording of the upper-body part of the subject expressing one of twelve complex mental states. Our system is able to recognize spontaneous complex mental states from different subjects and can be used in such an HRI scenario.
Pablo V. A. Barros, Stefan Wermter
IROS2
2015 Modeling development of natural multi-sensory integration using neural self-organisation and probabilistic population codes
abstract
Humans and other animals have been shown to perform near-optimally in multi-sensory integration tasks. Probabilistic population codes (PPCs) have been proposed as a mechanism by which optimal integration can be accomplished. Previous approaches have focussed on how neural networks might produce PPCs from sensory input or perform calculations using them, like combining multiple PPCs. Less attention has been given to the question of how the necessary organisation of neurons can arise and how the required knowledge about the input statistics can be learned. In this paper, we propose a model of learning multi-sensory integration based on an unsupervised learning algorithm in which an artificial neural network learns the noise characteristics of each of its sources of input. Our algorithm borrows from the self-organising map the ability to learn latent-variable models of the input and extends it to learning to produce a PPC approximating a probability density function over the latent variable behind its (noisy) input. The neurons in our network are only required to perform simple calculations and we make few assumptions about input noise properties and tuning functions. We report on a neurorobotic experiment in which we apply our algorithm to multi-sensory integration in a humanoid robot to demonstrate its effectiveness and compare it to human multi-sensory integration on the behavioural level. We also show in simulations that our algorithm performs near-optimally under certain plausible conditions, and that it reproduces important aspects of natural multi-sensory integration on the neural level.
Johannes Bauer 0002, Jorge Dávila-Chacón, Stefan Wermter
Connect. Sci.3
2015 A Scalable Meta-Classifier Combining Search and Classification Techniques for Multi-Level Text Categorization
abstract
Nowadays, documents are increasingly associated with multi-level category hierarchies rather than a flat category scheme. As the volume and diversity of documents grow, so do the size and complexity of the corresponding category hierarchies. To be able to access such hierarchically classified documents in real-time, we need fast automatic methods to navigate these hierarchies. Today’s data domains are also very different from each other, such as medicine and politics. These distinct domains can be handled by different classifiers. A document representation system which incorporates the inherent category structure of the data should also add useful semantic content to the data vectors and thus lead to better separability of classes. In this paper, we present a scalable meta-classifier to tackle today’s problem of multi-level data classification in the presence of large datasets. To speed up the classification process, we use a search-based method to detect the level-1 category of a test document. For this purpose, we use a category–hierarchy-based vector representation. We evaluate the meta-classifier by scaling to both longer documents as well as to a larger category set and show it to be robust in both cases. We test the architecture of our meta-classifier using six different base classifiers (Random forest, C4.5, multilayer perceptron, naïve Bayes, BayesNet (BN) and PART). We observe that even though there is a very small variation in the performance of different architectures, all of them perform much better than the corresponding single baseline classifiers. We conclude that there is substantial potential in this meta-classifier architecture, rather than the classifiers themselves, which successfully improves classification performance.
Nandita Tripathi, Michael P. Oakes, Stefan Wermter
Int. J. Comput. Intell. Appl.3
2015 Multimodal emotional state recognition using sequence-dependent deep hierarchical features
abstract
Emotional state recognition has become an important topic for human-robot interaction in the past years. By determining emotion expressions, robots can identify important variables of human behavior and use these to communicate in a more human-like fashion and thereby extend the interaction possibilities. Human emotions are multimodal and spontaneous, which makes them hard to be recognized by robots. Each modality has its own restrictions and constraints which, together with the non-structured behavior of spontaneous expressions, create several difficulties for the approaches present in the literature, which are based on several explicit feature extraction techniques and manual modality fusion. Our model uses a hierarchical feature representation to deal with spontaneous emotions, and learns how to integrate multiple modalities for non-verbal emotion recognition, making it suitable to be used in an HRI scenario. Our experiments show that a significant improvement of recognition accuracy is achieved when we use hierarchical features and multimodal information, and our model improves the accuracy of state-of-the-art approaches from 82.5% reported in the literature to 91.3% for a benchmark dataset on spontaneous emotion expressions.
Pablo V. A. Barros, Doreen Jirak, Cornelius Weber, Stefan Wermter
Neural Networks4
2015 Attention modeled as information in learning multisensory integration
abstract
Top-down cognitive processes affect the way bottom-up cross-sensory stimuli are integrated. In this paper, we therefore extend a successful previous neural network model of learning multisensory integration in the superior colliculus (SC) by top-down, attentional input and train it on different classes of cross-modal stimuli. The network not only learns to integrate cross-modal stimuli, but the model also reproduces neurons specializing in different combinations of modalities as well as behavioral and neurophysiological phenomena associated with spatial and feature-based attention. Importantly, we do not provide the model with any information about which input neurons are sensory and which are attentional. If the basic mechanisms of our model-self-organized learning of input statistics and divisive normalization-play a major role in the ontogenesis of the SC, then this work shows that these mechanisms suffice to explain a wide range of aspects both of bottom-up multisensory integration and the top-down influence on multisensory integration.
Johannes Bauer 0002, Sven Magg, Stefan Wermter
Neural Networks3
2014 Improving Domain-independent Cloud-Based Speech Recognition with Domain-Dependent Phonetic Post-Processing
abstract
Automatic speech recognition (ASR) technology has been developed to such a level that off-the-shelf distributed speech recognition services are available (free of cost), which allow researchers to integrate speech into their applications with little development effort or expert knowledge leading to better results compared with previously used open-source tools. Often, however, such services do not accept language models or grammars but process free speech from any domain. While results are very good given the enormous size of the search space, results frequently contain out-of-domain words or constructs that cannot be understood by subsequent domain-dependent natural language understanding (NLU) components. We present a versatile post-processing technique based on phonetic distance that integrates domain knowledge with open-domain ASR results, leading to improved ASR performance. Notably, our technique is able to make use of domain restrictions using various degrees of domain knowledge, ranging from pure vocabulary restrictions via grammars or N-Grams to restrictions of the acceptable utterances. We present results for a variety of corpora (mainly from human-robot interaction) where our combined approach significantly outperforms Google ASR as well as a plain open-source ASR solution.
Johannes Twiefel, Timo Baumann, Stefan Heinrich, Stefan Wermter
AAAI4
2014 FINGeR: Framework for interactive neural-based gesture recognition
German Ignacio Parisi, Pablo V. A. Barros, Stefan Wermter
ESANN3
2014 A Multichannel Convolutional Neural Network for Hand Posture Recognition
Pablo V. A. Barros, Sven Magg, Cornelius Weber, Stefan Wermter
ICANN4
2014 Improving Humanoid Robot Speech Recognition with Sound Source Localisation
Jorge Dávila-Chacón, Johannes Twiefel, Stefan Wermter
ICANN4
2014 Interactive Language Understanding with Multiple Timescale Recurrent Neural Networks
Stefan Heinrich, Stefan Wermter
ICANN2
2014 An Incremental Approach to Language Acquisition: Thematic Role Assignment with Echo State Networks
Xavier Hinaut, Stefan Wermter
ICANN2
2014 RatSLAM on Humanoids - A Bio-Inspired SLAM Model Adapted to a Humanoid Robot
Cornelius Weber, Stefan Wermter
ICANN3
2014 Human Action Recognition with Hierarchical Growing Neural Gas Learning
German Ignacio Parisi, Cornelius Weber, Stefan Wermter
ICANN3
2014 HandSOM - neural clustering of hand motion for gesture recognition in real time
abstract
Gesture recognition is an important task in Human-Robot Interaction (HRI) and the research effort towards robust and high-performance recognition algorithms is increasing. In this work, we present a neural network approach for learning an arbitrary number of labeled training gestures to be recognized in real time. The representation of gestures is hand-independent and gestures with both hands are also considered. We use depth information to extract salient motion features and encode gestures as sequences of motion patterns. Preprocessed sequences are then clustered by a hierarchical learning architecture based on self-organizing maps. We present experimental results on two different data sets: command-like gestures for HRI scenarios and communicative gestures that include cultural peculiarities, often excluded in gesture recognition research. For better recognition rates, noisy observations introduced by tracking errors are detected and removed from the training sets. Obtained results motivate further investigation of efficient neural network methodologies for gesture-based communication.
German Ignacio Parisi, Doreen Jirak, Stefan Wermter
RO-MAN3
2013 Embodied Language Understanding with a Multiple Timescale Recurrent Neural Network
Stefan Heinrich, Cornelius Weber, Stefan Wermter
ICANN3
2013 Self-Organized Neural Learning of Statistical Inference from High-Dimensional Data
Johannes Bauer 0002, Stefan Wermter
IJCAI2
2013 Neural and statistical processing of spatial cues for sound source localisation
abstract
When confronting binaural sound source localisation (SSL) algorithms with different environments and robotic platforms, there is an increasing need for non-linear integration methods of spatial cues. Based on interaural time and level differences, we compare the performance of several SSL systems. The architecture has three degrees of freedom, i.e. each tested architecture employs a different combination of representation of binaural cues, clustering and classification algorithms. The heuristic for the selection of methods is the same at each degree of freedom: to compare the impact of traditional statistical techniques versus machine learning algorithms with different degrees of biological inspiration. The overall performance is evaluated in the analysis of each system, including the accuracy of its output, training time and adequateness for life-long learning. The results support the use of hybrid systems, consisting different kinds of artificial neural networks, as they present an effective compromise between the characteristics evaluated.
Jorge Dávila-Chacón, Sven Magg, Stefan Wermter
IJCNN4
2013 Neural Hopfield-ensemble for multi-class head pose detection
abstract
Multi-class object detection is perhaps the most important task for many computer vision systems and mobile robots. In this work we will show that Hopfield Neural Network (HNN) ensembles can successfully detect and classify objects from several classes by taking advantage of head-pose estimation. The single HNNs are using pixel sums of Haar-like features as input, resulting in HNNs with a small number of neurons. An advantage of using these in ensembles is their compact form. Although it was shown that such HNNs can only memorise few patterns, by utilising a naive-Bayes mechanism we were able to exploit the multi-class ability of single HNNs within an ensemble. In this work we report successful head pose classification, which presents a 4-class problem (3 poses + negatives). Results show that successful classification can be achieved with small training sets and ensembles, making this approach an interesting choice for online learning and robotics.
Nils Meins, Sven Magg, Stefan Wermter
IJCNN3
2013 Hierarchical SOM-based detection of novel behavior for 3D human tracking
abstract
We present a hierarchical SOM-based architecture for the detection of novel human behavior in indoor environments. The system can unsupervisedly learn normal activity and then report novel behavioral patterns as abnormal. The learning stage is based on the clustering of motion with self-organizing maps. With this approach, no domain-specific knowledge on normal actions is required. During the tracking stage, we extract human motion properties expressed in terms of multidimensional flow vectors. From this representation, three classes of motion descriptors are encoded: trajectories, body features and directions. During the training phase, SOM networks are responsible for learning a specific class of descriptors. For a more accurate clustering of motion, we detect and remove outliers from the training data. At detection time, we propose a hybrid neural-statistical method for 3D posture recognition in real time. New observations are tested for novelty and reported if they deviate from the learned behavior. Experiments were performed in two different tracking scenarios with fixed and mobile depth sensor. In order to exhibit the validity of the proposed methodology, several experimental setups and the evaluation of obtained results are presented.
German Ignacio Parisi, Stefan Wermter
IJCNN2
2012 A Fast Subspace Text Categorization Method Using Parallel Classifiers
Nandita Tripathi, Michael P. Oakes, Stefan Wermter
CICLing (2)3
2012 Adaboost and Hopfield Neural Networks on different image representations for robust face detection
abstract
Face detection is an active research area comprising the fields of computer vision, machine learning and intelligent robotics. However, this area is still challenging due to many problems arising from image processing and the further steps necessary for the detection process. In this work we focus on Hopfield Neural Network (HNN) and ensemble learning. It extends our recent work by two components: the simultaneous usage of different image representations and combinations as well as variations in the training procedure. Using the HNN within an ensemble achieves high detection rates but shows no increase in false detection rates, as is commonly the case. We present our experimental setup and investigate the robustness of our architecture. Our results indicate, that with the presented methods the face detection system is flexible regarding varying environmental conditions, leading to a higher robustness.
Nils Meins, Doreen Jirak, Cornelius Weber, Stefan Wermter
HIS4
2012 What do Objects Feel Like? - Active Perception for a Humanoid Robot
Jens Kleesiek, Stephanie Badde, Stefan Wermter, Andreas K. Engel
ICAART (1)3
2012 Biomimetic Binaural Sound Source Localisation with Ego-Noise Cancellation
Jorge Dávila-Chacón, Stefan Heinrich, Stefan Wermter
ICANN (1)4
2012 Adaptive Learning of Linguistic Hierarchy in a Multiple Timescale Recurrent Neural Network
Stefan Heinrich, Cornelius Weber, Stefan Wermter
ICANN (1)3
2012 Hybrid Ensembles Using Hopfield Neural Networks and Haar-Like Features for Face Detection
Nils Meins, Stefan Wermter, Cornelius Weber
ICANN (1)2
2012 Learning Features and Predictive Transformation Encoding Based on a Horizontal Product Model
Junpei Zhong, Cornelius Weber, Stefan Wermter
ICANN (1)3
2012 A SOM-based model for multi-sensory integration in the superior colliculus
abstract
We present an algorithm based on the self-organizing map (SOM) which models multi-sensory integration as realized by the superior colliculus (SC). Our algorithm differs from other algorithms for multi-sensory integration in that it learns mappings between modalities' coordinate systems, it learns their respective reliabilities for different points in space, and uses mappings and reliabilities to perform cue integration. It does this in only one learning phase without supervision and such that calculations and data structures are local to individual neurons. Our simulations indicate that our algorithm can learn near-optimal integration of input from noisy sensory modalities.
Johannes Bauer 0002, Cornelius Weber, Stefan Wermter
IJCNN3
2012 A neurocomputational amygdala model of auditory fear conditioning: A hybrid system approach
abstract
In this work, we present a neurocomputational model for auditory-cue fear acquisition. Computational fear conditioning has experienced a growing interest over the last few years, on the one hand, because it is a robust and quick learning paradigm that can contribute to the development of more versatile robots, and on the other hand, because it can help in the understanding of fear conditioning and dysfunctions in animals. Fear learning involves sensory and motor aspects [1] and it is essential for adaptive self-protective systems. We argue that a deeper study of the mechanisms underlying fear circuits in the brain will contribute not only to the development of safer robots but eventually also to a better conceptual understanding of neural fear processing in general. Towards the development of a robotic adaptive self-protective system, we have designed a neural model of fear conditioning based on LeDoux's dual-route hypothesis of fear [2] and also dopamine modulated Pavlovian conditioning [3]. Our hybrid approach is capable of learning the temporal relationship between auditory sensory cues and an aversive or appetitive stimulus. The model was tested as a neural network simulation but it was designed to be used with minor modifications on a robotic platform.
Nicolás Navarro-Guerrero, Robert J. Lowe, Stefan Wermter
IJCNN3
2012 A neural approach for robot navigation based on cognitive map learning
abstract
This paper presents a neural network architecture for a robot learning new navigation behavior by observing a human's movement in a room. While indoor robot navigation is challenging due to the high complexity of real environments and the possible dynamic changes in a room, a human can explore a room easily without any collisions. We therefore propose a neural network that builds up a memory for spatial representations and path planning using a person's movements as observed from a ceiling-mounted camera. Based on the human's motion, the robot learns a map that is used for path planning and motor-action codings. We evaluate our model with a detailed case study and show that the robot navigates effectively.
Cornelius Weber, Stefan Wermter
IJCNN3
2011 Hybrid classifiers for improved semantic subspace learning of news documents
abstract
The volume and diversity of documents available in today's world is increasing daily. It is therefore difficult for a single classifier to efficiently handle multi-level categorization of such a varied document space. In this paper we analyse methods to enhance the efficiency of a single classifier for two-level classification by combining it with classifiers of other types. We use the maximum significance value as an indicator for the subspace of a test document. We represent the documents using the conditional significance vector which increases the distinction between classes within a subspace. Our experiments show that dividing a document space into different semantic subspaces increases the efficiency of such hybrid classifier combinations. Applying different types of classifiers on different subspaces substantially improves overall learning.
Nandita Tripathi, Michael P. Oakes, Stefan Wermter
HIS3
2011 Determining Cooperation in Multiagent Systems with Cultural Traits
Stefan Heinrich, Markus Eberling, Stefan Wermter
ICAART (2)3
2011 Hybrid Parallel Classifiers for Semantic Subspace Learning
Nandita Tripathi, Michael P. Oakes, Stefan Wermter
ICANN (2)3
2011 Person Tracking Based on a Hybrid Neural Probabilistic Model
Cornelius Weber, Stefan Wermter
ICANN (2)3
2011 Robot Trajectory Prediction and Recognition Based on a Computational Mirror Neurons Model
Junpei Zhong, Cornelius Weber, Stefan Wermter
ICANN (2)3
2010 Semantic Subspace Learning with conditional significance vectors
abstract
Subspace detection and processing is receiving more attention nowadays as a method to speed up search and reduce processing overload. Subspace Learning algorithms try to detect low dimensional subspaces in the data which minimize the intra-class separation while maximizing the inter-class separation. In this paper we present a novel technique using the maximum significance value to detect a semantic subspace. We further modify the document vector using conditional significance to represent the subspace. This enhances the distinction between classes within the subspace. We compare our method against TFIDF with PCA and show that it consistently outperforms the baseline with a large margin when tested with a wide variety of learning algorithms. Our results show that the combination of subspace detection and conditional significance vectors improves subspace learning.
Nandita Tripathi, Stefan Wermter, Chihli Hung, Michael P. Oakes
IJCNN2
2010 Configuring the stochastic Helmholtz machine for subcortical emotional learning
abstract
Emotional learning involves two stages. The first is to acquire reinforcers from stimuli and the second is to associate such reinforcers with emotional responses. Both stages can be found occurring in the amygdala. LeDoux's fear circuit model suggests two routes, a subcortical route and a cortical route, for emotional information entering the amygdala for associative learning. It can be used to explain how the actual recognition of emotions from facial expressions can be processed in the brain. Based on the model, a neural architecture is proposed using the stochastic Helmholtz machine (SHM) with the wake-sleep algorithm. In this paper, the results of three experiments about the subcortical emotional learning are reported, where different configurations of SHMs are involved. The first two experiments are to identify a suitable way to allow behavioural responses entering the central nucleus of the amygdala for association. However, both experiments show symptoms of overfitting, where some weights and biases of neurons are observed that will unusually increase during training. Therefore, the final experiment is designed to maintain the range of weights between -1 and +1 in order to solve the overfitting problem. The last experiment shows that the neural architecture with the new weight policy holds a lot of potential for modelling subcortical learning.
Chi-Yung Yau, Kevin Burn, Stefan Wermter
IJCNN3
2010 A biologically inspired spiking neural network model of the auditory midbrain for sound source localisation
David Pérez-González, Adrian Rees, Harry R. Erwin, Stefan Wermter
Neurocomputing5
2009 Multiple Sound Source Localisation in Reverberant Environments Inspired by the Auditory Midbrain
David Pérez-González, Adrian Rees, Harry R. Erwin, Stefan Wermter
ICANN (1)5
2009 A biomimetic spiking neural network of the auditory midbrain for mobile robot sound localisation in reverberant environments
abstract
This paper proposes a spiking neural network (SNN) of the mammalian auditory midbrain to achieve binaural sound source localisation with a mobile robot. The network is inspired by neurophysiological studies on the organisation of binaural processing in the medial superior olive (MSO), lateral superior olive (LSO) and the inferior colliculus (IC) to achieve a sharp azimuthal localisation of sound source over a wide frequency range in situations where there is auditory clutter and reverberation. Three groups of artificial neurons are constructed to represent the neurons in the MSO, LSO and IC that are sensitive to interaural time difference (ITD), interaural level difference (ILD) and azimuth angle respectively. The ITD and ILD cues are combined in the IC using Bayes's theorem to estimate the azimuthal direction of a sound source. Two of known IC cells, onset and sustained-regular are modelled. The azimuth estimations at different robot positions are then used to calculate the sound source position by a triangulation method using an environment map constructed by a laser scanner. The experimental results show that the addition of ILD information significantly increases sound localisation performance at frequencies above 1 kHz. The mobile robot is able to localise a sound source in an acoustically cluttered and reverberant environment.
David Pérez-González, Adrian Rees, Harry R. Erwin, Stefan Wermter
IJCNN5
2009 Robotic sound-source localisation architecture using cross-correlation and recurrent neural networks
John C. Murray, Harry R. Erwin, Stefan Wermter
Neural Networks3
2009 Multimodal communication in animals, humans and robots: An introduction to perspectives in brain-inspired informatics
Stefan Wermter, M. Page, Michael Knowles, Vittorio Gallese, Friedemann Pulvermüller, John G. Taylor
Neural Networks1
2008 The Hybrid Integration of Perceptual Symbol Systems and Interactive Reinforcement Learning
abstract
In order to produce robots which can interact more effectively with humans we propose that it is necessary for their cognitive processes to be grounded in the same perceptual elements as humans deal with. Perceptual symbol systems offer an attractive mechanism for capturing the symbolic properties of the senses and for integrating them into higher level cognitive processes. We have designed a perceptual symbol system where the robot learns about objects through interaction and reinforcement and have carried out experiments to assess the merits of this approach. We show that the use of human perceptual elements combined with interactive reinforcement leads to intuitive learning and interpretable knowledge structures.
Michael Knowles, Stefan Wermter
HIS2
2008 MIRA: A Learning Multimodal Interactive Robot Agent
abstract
In this paper we present a robotic head MIRA (multimodal interactive robot agent) which has been developed for studying the learning of human robot interaction and improving our understanding of human robot interaction techniques. In this paper we focus on two main aspects of the system; first, we describe how the robot head learns to recognise faces for supporting the interaction process between a human and MIRA. Second, we show how MIRA can learn to identify sound sources of interest and attend to the source location improving the social interaction effect. We propose that there is substantial potential for learning visual and auditory features in order to increase adaptability and robustness of robotic heads.
John C. Murray, Stefan Wermter, Michael Knowles
HIS2
2008 A Biologically Inspired Spiking Neural Network for Sound Localisation by the Inferior Colliculus
Harry R. Erwin, Stefan Wermter, Mahmoud Elsaid
ICANN (2)3
2008 Visual robot homing using Sarsa(lambda), whole image measure, and radial basis function
abstract
This paper describes a model for visual homing. It uses Sarsa(lambda) as its learning algorithm, combined with the Jeffery divergence measure (JDM) as a way of terminating the task and augmenting the reward signal. The visual features are taken to be the histograms difference of the current view and the stored views of the goal location, taken for all RGB channels. A radial basis function layer acts on those histograms to provide input for the linear function approximator. An on-policy on-line Sarsa(lambda) method was used to train three linear neural networks one for each action to approximate the action-value function with the aid of eligibility traces. The resultant networks are trained to perform visual robot homing, where they achieved good results in finding a goal location. This work demonstrates that visual homing based on reinforcement learning and radial basis function has a high potential for learning local navigation tasks.
Abdulrahman Altahhan, Kevin Burn, Stefan Wermter
IJCNN3
2008 Hybrid learning architecture for unobtrusive infrared tracking support
abstract
The system architecture presented in this paper is designed for helping an aged person to live longer independently in their own home by detecting unusual and potentially hazardous behaviours. The system consists of two major components. The first component is the tracking part which is responsible for monitoring the movements of the person within the home, while the second part is a learning agent which is responsible for learning the behavioural patterns of the person. For the tracking part of the system a simulation portraying a virtual room with passive infrared sensors has been designed, while for the learning agent a hybrid architecture has been implemented. The hybrid architecture consists of a Markov Chain Model, Template Matching, Fuzzy Logic and Memory-Based reasoning techniques. The hybrid structure was selected because it combined the strengths of the constituent algorithms and because it supports the learning with limited training data. The resultant system was able to not only classify between the normal and the abnormal paths but was also able to distinguish between different normal routes. We claim that passive infrared tracking combined with a hybrid learning architecture has potential for adaptive unobtrusive tracking support.
K. K. Kiran Bhagat, Stefan Wermter, Kevin Burn
IJCNN2
2008 A neural wake-sleep learning architecture for associating robotic facial emotions
abstract
A novel wake-sleep learning architecture for processing a robot’s facial expressions is introduced. According to neuroscience evidence, associative learning of emotional responses and facial expressions occurs in the brain in the amygdala. Here we propose an architecture inspired by how the amygdala receives information from other areas of the brain to discriminate it and generate innate responses. The architecture is composed of many individual Helmholtz machines using the wake-sleep learning algorithm for performing information transformation and recognition. The Helmholtz machine is used since its re-entrant connections support both supervised and unsupervised learning. Potentially it can explain some aspects of human learning of emotional concepts and experience. In this research, a robotic head’s facial expression dataset is used. The objective of this learning architecture is to demonstrate the neural basis for the association of recognized facial expressions and linguistic emotion labels. It implies the understanding of emotions from observation and is further used to generate facial expressions. In contrast with other facial expression recognition research, this work concentrates more on emotional information processing and neural concept development, rather than a technical recognition task. This approach has a lot of potential to contribute towards neurally inspired emotional experience in robotic systems.
Chi-Yung Yau, Kevin Burn, Stefan Wermter
IJCNN3
2008 Mobile robot broadband sound localisation using a biologically inspired spiking neural network
abstract
A biologically inspired azimuthal broadband sound localisation system is introduced to simulates the functional organisation of the human auditory midbrain up to the inferior colliculus (IC). Supported by recent neurophysiological studies on the role of the IC and superior olivary complex (SOC) in sound processing, our system models two ascending pathways of the auditory midbrain: the ITD (Interaural Time Difference) pathway and ILD (Interaural Level Difference) pathway. In our approach to modelling the ITD pathway, we take account of Yinpsilas finding that only a single delay line exists in the ITD processing from cochlea to SOC for the ipsilateral ear while multiple delay lines exists for the contralateral ear. The ILD pathway is modelled without varied delay lines because of neurophysiological evidence that indicates the delays along that pathway are minimal and constant. First, two-dimensional (2D) tonotopical ITD and ILD spike maps over frequency and ITD/ILD are calculated by a spiking neural network which follows the biological delay structure. Then these maps are weighted considering the advance of ITD in low frequency and ILD in middle and high frequency. Finally, ITD and ILD maps are merged together to find out the best estimation of the sound source. Experimental results involving noise and voice show that our model performs sound localisation that approaches biological performance. Our approach brings not only new insight into the brain mechanism of the auditory system, but also demonstrates a practical application of sound localisation for mobile robots.
Harry R. Erwin, Stefan Wermter
IROS3
2007 Auto-Extraction, Representation and Integration of a Diabetes Ontology Using Bayesian Networks
abstract
This paper describes how high level biological knowledge obtained from ontologies such as the gene ontology (GO) can be integrated with low level information extracted from a Bayesian network trained on protein interaction data. We can automatically generate a biological ontology by text mining the type II diabetes research literature. The ontology is populated with the entities and relationships from protein-to-protein interactions. New, previously unrelated information is extracted from the growing body of research literature and incorporated with knowledge already known on this subject from the gene ontology and databases such as BIND and BioGRID. We integrate the ontology within the probabilistic framework of Bayesian networks which enables reasoning and prediction of protein function.
Kenneth McGarry, Sheila Garfield, Stefan Wermter
CBMS3
2007 A self-organizing map of sigma-pi units
Cornelius Weber, Stefan Wermter
Neurocomputing2
2006 Reinforcement Learning for Platform-Independent Visual Robot Control
abstract
This paper proposes a new architecture for robot control. A test scenario is outlined to test the proposed system and enable a comparison with an existing system, which is able to fulfil the scenario and thus be used as a benchmark. The scenario is a navigation task, to allow a robot to approach a specified landmark. The proposed architecture will make use of two control units, one to allow a pan/tilt camera to track the landmark as the robot moves, and a second to control the robots drive motors. These units will be trained via reinforcement learning, and provide the potential for platform-independent robot control.
David Muse, Kevin Burn, Stefan Wermter
IJCNN3
2006 Bioinspired Auditory Sound Localisation for Improving the Signal to Noise Ratio of Socially Interactive Robots
abstract
In this paper we describe a bioinspired hybrid architecture for acoustic sound source localisation and tracking to increase the signal to noise ratio (SNR) between speaker and background sources for a socially interactive robot's speech recogniser system. The model presented incorporates the use of interaural time difference for azimuth estimation and recurrent neural networks for trajectory prediction. The results are then presented showing the difference in the SNR of a localised and non-localised speaker source, in addition to presenting the recognition rates between a localised and non-localised speaker source. From the results presented in this paper it can be seen that by orientating towards the sound source of interest the recognition rates of that source can be increased
John C. Murray, Stefan Wermter, Harry R. Erwin
IROS2
2006 Temporal sequence detection with spiking neurons: towards recognizing robot language instructions
abstract
We present an approach for recognition and clustering of spatio temporal patterns based on networks of spiking neurons with active dendrites and dynamic synapses. We introduce a new model of an integrate-and-fire neuron with active dendrites and dynamic synapses (ADDS) and its synaptic plasticity rule. The neuron employs the dynamics of the synapses and the active properties of the dendrites as an adaptive mechanism for maximizing its response to a specific spatio-temporal distribution of incoming action potentials. The learning algorithm follows recent biological evidence on synaptic plasticity. It goes beyond the current computational approaches which are based only on the relative timing between single pre- and post-synaptic spikes and implements a functional dependence based on the state of the dendritic and somatic membrane potentials around the pre- and post-synaptic action potentials. The learning algorithm is demonstrated to effectively train the neuron towards a selective response determined by the spatio-temporal pattern of the onsets of input spike trains. The model is used in the implementation of a part of a robotic system for natural language instructions. We test the model with a robot whose goal is to recognize and execute language instructions. The research in this article demonstrates the potential of spiking neurons for processing spatio-temporal patterns and the experiments present spiking neural networks as a paradigm which can be applied for modelling sequence detectors at word level for robot instructions.
Christo Panchev, Stefan Wermter
Connect. Sci.2
2006 Call classification using recurrent neural networks, support vector machines and finite state automata
Sheila Garfield, Stefan Wermter
Knowl. Inf. Syst.2
2006 Robot docking based on omnidirectional vision and reinforcement learning
David Muse, Cornelius Weber, Stefan Wermter
Knowl. Based Syst.3
2006 A camera-direction dependent visual-motor coordinate transformation for a visually guided neural robot
Cornelius Weber, David Muse, Mark Elshaw, Stefan Wermter
Knowl. Based Syst.4
2006 Data mining using rule extraction from Kohonen self-organising maps
James Malone, Kenneth McGarry, Stefan Wermter, Chris Bowerman
Neural Comput. Appl.3
2006 A hybrid generative and predictive model of the motor cortex
Cornelius Weber, Stefan Wermter, Mark Elshaw
Neural Networks2
2005 Hybrid Intelligent Systems and Cognitive Robotics
abstract
Summary form only given. There has been substantial progress in both intelligent systems and robotics in recent years. While in the past robots were most successful in traditional industrial environments, new generations of hybrid intelligent robotic systems are being developed which focus on higher cognitive capabilities, including reasoning learning and language communication. In this paper we give an overview of learning neural robots from a perspective of integrative hybrid intelligent systems and illustrate some new developments including also examples under development in the Centre for Hybrid Intelligent Systems.
Stefan Wermter
HIS1
2005 Reinforcement Learning in MirrorBot
Cornelius Weber, David Muse, Mark Elshaw, Stefan Wermter
ICANN (1)4
2005 Image Segmentation by Complex-Valued Units
Cornelius Weber, Stefan Wermter
ICANN (1)2
2005 Training without data: Knowledge Insertion into RBF Neural Networks
Kenneth McGarry, Stefan Wermter
IJCAI2
2005 A constructive and hierarchical self-organizing model in a non-stationary environment
abstract
Several related self-organizing neural models have been proposed to enhance the flexibility of self-organizing maps. In our studies, these models depend on the pre-definition of several thresholds which are used as guidance of neural behaviors for specific data sets. However, it is not trivial to determine those thresholds in a non-stationary environment. When a proper threshold has been determined, this threshold may not be suitable for the future. Therefore, in this paper, we compare the dynamic adaptive self-organizing hybrid (DASH) model with the growing neural gas (GNG) model by introducing several different initial thresholds to test their feasibility. Our experiments show that the DASH model is more stable and practicable for document clustering in a non-stationary environment since DASH adjusts its behavior not only by modifying its parameters but also by an adaptive structure.
Chihli Hung, Stefan Wermter
IJCNN2
2005 Spatio-temporal neural data mining architecture in learning robots
abstract
There has been little research into the use of hybrid neural data mining to improve robot performance or enhance their capability. This paper presents a novel neural data mining technique that analyses robot sensor data for imitation learning. Learning by imitation allows a robot to learn from observing either another robot or a human to gain skills, understand the behavior of others and create solutions to problems. We demonstrate a hybrid approach of differential ratio data mining to perform analysis on spatio-temporal robot behavioral data. The technique offers classification performance gains for recognition of robot actions by highlighting points of covariance and hence interest within the data.
James Malone, Mark Elshaw, Kenneth McGarry, Chris Bowerman, Stefan Wermter
IJCNN5
2005 A recurrent neural network for sound-source motion tracking and prediction
abstract
Recurrent neural networks (RNN) have been used in many applications for both pattern detection and prediction. This paper shows the use of RNN's as a speed classifier and predictor for a robotic sound source tracking system. The system requires extensive training to classify all possible speeds to enable dynamic tracking of the most prominent sound within the environment.
John C. Murray, Harry R. Erwin, Stefan Wermter
IJCNN3
2005 Auditory robotic tracking of sound sources using hybrid cross-correlation and recurrent networks
abstract
This paper describes an auditory robotic system capable of computing the angle of incidence of a sound source on the horizontal plane (azimuth). The system, with the use of an Elman type recurrent neural network (RNN), is able to dynamically track this sound source as it changes azimuthally within the environment. The RNN is used to enable fast tracking responses to the overall system over a set time, as opposed to waiting for the next sound position before moving. The system is first tested in a simulated environment and then these results are compared with testing on the robotic system. The results show that the development of a hybrid system incorporating cross-correlation and recurrent neural networks is an effective mechanism for the control of a robot that tracks sound sources azimuthally.
John C. Murray, Stefan Wermter, Harry R. Erwin
IROS2
2004 Predictive Top-Down Knowledge Improves Neural Exploratory Bottom-Up Clustering
Chihli Hung, Stefan Wermter
ECIR2
2004 Knowing what and where: a computational model for visual attention
abstract
We describe a model of invariant object recognition in the brain that incorporates feedback biasing effects of top-down attentional mechanism on a hierarchically organised set of visual cortical areas. The model displays a space based and a object based visual search by using a top-down attention feedback model from posterior parietal modules and interaction between the two processing streams dorsal and ventral.
Kaustubh Chokshi, Christo Panchev, Stefan Wermter, John G. Taylor
IJCNN3
2004 Self organising neural place codes for vision based robot navigation
abstract
Autonomous robots must be able to navigate independently within an environment. In the animal brain, so-called place cells respond to the environment where the animal is. We present a model of place cells based on self-organising maps. The aim of this paper is to show how image invariance can improve the performance of the neural place codes and make the model more robust to noise. The paper also demonstrates that localisation can be learned without having a pre-defined map given to the robot by humans and that after training, a robot can localise itself within a learned environment.
Kaustubh Chokshi, Stefan Wermter, Christo Panchev, Kevin Burn
IJCNN2
2004 An associator network approach to robot learning by imitation through vision, motor control and language
abstract
Imitation learning offers a valuable approach for developing intelligent robot behaviour. We present an imitation approach based on an associator neural network inspired by brain modularity and mirror neurons. The model combines multimodal input based on higher-level vision, motor control and language so that a simulated student robot is able to learn from observing three behaviours which are performed by a teacher robot. The student robot associates these inputs to recognise the behaviour being performed or to perform behaviours by language instruction. With behaviour representations segregating into regions it models aspects of the mirror neuron system as similar patterns of neural activation are involved in recognition and performance.
Mark Elshaw, Cornelius Weber, Alexandros Zochios, Stefan Wermter
IJCNN4
2004 A time-based self-organising model for document clustering
abstract
Most current approaches for document clustering do not consider the non-stationary feature of real world document collection. In this paper, in a non-stationary environment, we propose a new self-organising model, namely the dynamic adaptive self-organising hybrid (DASH) model. The DASH model runs continuously since the new document set is formed consecutively for training while the old document set is still at the training stage. Knowledge learned from the old data set is adjusted to reflect the new data set and therefore document clusters are up-to-date. We test the performance of our model using the Reuters-RCV1 news corpus and obtain promising results based on the criteria of classification accuracy and average quantization error.
Chihli Hung, Stefan Wermter
IJCNN2
2004 Spike-timing-dependent synaptic plasticity: from single spikes to spike trains
Christo Panchev, Stefan Wermter
Neurocomputing2
2004 Robot docking with neural vision and reinforcement
Cornelius Weber, Stefan Wermter, Alexandros Zochios
Knowl. Based Syst.2
2003 Learning Localisation Based on Landmarks Using Self-Organisation
Kaustubh Chokshi, Stefan Wermter, Cornelius Weber
ICANN2
2003 Comparing Support Vector Machines, Recurrent Networks, and Finite State Transducers for Classifying Spoken Utterances
Sheila Garfield, Stefan Wermter
ICANN2
2003 Object Localisation Using Laterally Connected "What" and "Where" Associator Networks
Cornelius Weber, Stefan Wermter
ICANN2
2003 A Dynamic Adaptive Self-Organising Hybrid Model for Text Clustering
abstract
Clustering by document concepts is a powerful way of retrieving information from a large number of documents. This task in general does not make any assumption on the data distribution. For this task we propose a new competitive self-organising (SOM) model, namely the dynamic adaptive self-organising hybrid model (DASH). The features of DASH are a dynamic structure, hierarchical clustering, nonstationary data learning and parameter self-adjustment. All features are data-oriented: DASH adjusts its behaviour not only by modifying its parameters but also by an adaptive structure. The hierarchical growing architecture is a useful facility for such a competitive neural model which is designed for text clustering. We have presented a new type of self-organising dynamic growing neural network which can deal with the nonuniform data distribution and the nonstationary data sets and represent the inner data structure by a hierarchical view.
Chihli Hung, Stefan Wermter
ICDM2
2003 Self-organisation of language instruction for robot action control
abstract
Most current approaches for robot control do not make use of language and ignore neural learning. However, our robot control approach uses language instruction and draws from the concepts of regional distributed modularity, mirror neuron theory and neural assemblies. We described a self-organising model that clusters action verbs into different locations of the output layer dependent on the body part they are associated with. In doing so we build on our previous work by using sensor reading from the MIRA robot that incorporate semantic features of the actions verbs. Furthermore, we outline a hierarchical computational model for a neurally inspired self-organising robot action control system using language for instruction.
Mark Elshaw, Stefan Wermter, Peter Watt
IJCNN2
2003 A modular approach to self-organization of robot control based on language instruction
abstract
In this paper we focus on how instructions for actions can be modelled in a self-organizing memory. Our approach draws from the concepts of regional distributed modularity and self-organization. We describe a self-organizing model that clusters action representations into different locations dependent on the body part they are related to. In the first case study we consider semantic representations of action verb meaning and then extend this concept significantly in a second case study by using actual sensor readings from our MIRA robot. Furthermore, we outline a modular model for a self-organizing robot action control system using language for instruction. Our approach for robot control using language incorporates some evidence related to the architectural and processing characteristics of the brain (Wermter et al. 2001b). This paper focuses on the neurocognitive clustering of actions and regional modularity for language areas in the brain. In particular, we describe a self-organizing network that realizes action clustering (Pulvermüller 2003).
Stefan Wermter, Mark Elshaw, Simon Farrand
Connect. Sci.1
2003 Symbolic state transducers and recurrent neural preference machines for text mining
Garen Arevian, Stefan Wermter, Christo Panchev
Int. J. Approx. Reason.2
2003 Learning robot actions based on self-organising language memory
Stefan Wermter, Mark Elshaw
Neural Networks1
2002 Selforganizing Classification on the Reuters News Corpus
Stefan Wermter, Chihli Hung
COLING1
2002 Recurrent Neural Learning for Helpdesk Call Routing
Sheila Garfield, Stefan Wermter
ICANN2
2002 Spike-Timing Dependent Competitive Learning of Integrate-and-Fire Neurons with Active Dendrites
Christo Panchev, Stefan Wermter, Huixin Chen
ICANN2
2001 A Mirror Neuron System for Syntax Acquisition
Steve Womble, Stefan Wermter
ICANN2
2001 Knowledge Extraction from Local Function Networks
Kenneth McGarry, Stefan Wermter, John MacIntyre
IJCAI2
2001 The Extraction and Comparison of Knowledge from Local Function Networks
abstract
Extracting rules from RBFs is not a trivial task because of nonlinear functions or high input dimensionality. In such cases, some of the hidden units of the RBF network have a tendency to be "shared" across several output classes or even may not contribute to any output class. To address this we have developed an algorithm called LREX (for Local Rule EXtraction) which tackles these issues by extracting rules at two levels: hREX extracts rules by examining the hidden unit to class assignments while mREX extracts rules based on the input space to output space mappings. The rules extracted by our algorithm are compared and contrasted against a competing local rule extraction system. The central claim of this paper is that local function networks such as radial basis function (RBF) networks have a suitable architecture based on Gaussian functions that is amenable to rule extraction.
Kenneth McGarry, Stefan Wermter, John MacIntyre
Int. J. Comput. Intell. Appl.2
2000 Complex Preferences for the Integration of Neural Codes
abstract
This paper presents a complex preference framework of integrating pulsed neural networks into neural/symbolic hybrid approaches. In particular, we introduce an interpretation of neural codes as multidimensional complex neural preferences and preference classes which allow the integration of knowledge from different neural and symbolic models. We define some basic operations on complex preferences and preference classes that allow them to be directly integrated into symbolic models. Furthermore, we show the interpretation of mean firing rate, time-to-first-spike, synchrony and phase codes as complex neural preferences and the interpretation of the operations on preference classes of these codes. The symbolic interpretation and simultaneous processing of mean firing rate and pulse coding schemes in a preferences framework are addressed.
Christo Panchev, Stefan Wermter
IJCNN (2)2
2000 Meaning Spotting and Robustness of Recurrent Networks
abstract
This paper describes and evaluates the behavior of preference-based recurrent networks which process text sequences. First, we train a recurrent plausibility network to learn a semantic classification of the Reuters news title corpus. Then we analyze the robustness and incremental learning behavior of these networks in more detail. We demonstrate that these recurrent networks use their recurrent connections to support incremental processing. In particular, we compare the performance of the real title models with reversed title models and even random title models. We find that the recurrent networks can, even under these severe conditions, provide good classification results. We claim that previous context in recurrent connections and a meaning spotting strategy are pursued by the network which supports this robust processing.
Stefan Wermter, Christo Panchev, Garen Arevian
IJCNN (3)1
2000 Knowledge Extraction from Transducer Neural Networks
Stefan Wermter
Appl. Intell.1
2000 Neural Fuzzy Preference Integration Using Neural Preference Moore Machines
abstract
This paper describes preference classes and preference Moore machines as a basis for integrating different hybrid neural representations. Preference classes are shown to provide a basic link between neural preferences and fuzzy representations at the preference class level. Preference Moore machines provide a link between recurrent neural networks and symbolic transducers at the preference Moore machine level. We demonstrate how the concepts of preference classes and preference Moore machines can be used to interpret neural network representations and to integrate knowledge from hybrid neural representations. One main contribution of this paper is the introduction and analysis of neural preference Moore machines and their link to a fuzzy interpretation. Furthermore, we illustrate the interpretation and combination of various neural preference Moore machines with additional real-world examples.
Stefan Wermter
Int. J. Neural Syst.1
2000 Neural Network Agents for Learning Semantic Text Classification
Stefan Wermter
Inf. Retr.1
1999 Preference Moore Machines for Neural Fuzzy Integration
Stefan Wermter
IJCAI1
1999 Rule generation from neural networks for student assessment
abstract
HyValue is a hybrid electronic submission system which utilizes techniques from natural language processing, neural networks and rule based systems to accept, evaluate and mark work submitted by a student for reading or writing. This paper describes the theory behind the system design and the development of the individual components and their interaction. Issues addressed include the definition of sentence structure, fuzzy rule construction and integration with a knowledge base containing the marking rubrics for reading and writing. An evaluation of the system is provided and conclusions drawn.
M. J. McAlister, Stefan Wermter
IJCNN2
1999 Knowledge extraction from radial basis function networks and multilayer perceptrons
abstract
This paper deals with an evaluation and comparison of the accuracy and complexity of symbolic rules extracted from radial basis function networks and multilayer perceptrons. Here we examine the ability of rule extraction algorithms to extract meaningful rules that describe the overall performance of a particular network. In addition, the paper also highlights the suitability of a specific neural network architecture for particular classification problems. The study carried out on the extracted rule quality and complexity also has a direct bearing on the use of rule extraction algorithms for data mining and knowledge discovery.
Kenneth McGarry, Stefan Wermter, John MacIntyre
IJCNN2
1997 SCREEN: Learning a Flat Syntactic and Semantic Spoken Language Analysis Using Artificial Neural Networks
abstract
Previous approaches of analyzing spontaneously spoken language often have been based on encoding syntactic and semantic knowledge manually and symbolically. While there has been some progress using statistical or connectionist language models, many current spoken- language systems still use a relatively brittle, hand-coded symbolic grammar or symbolic semantic component. In contrast, we describe a so-called screening approach for learning robust processing of spontaneously spoken language. A screening approach is a flat analysis which uses shallow sequences of category representations for analyzing an utterance at various syntactic, semantic and dialog levels. Rather than using a deeply structured symbolic analysis, we use a flat connectionist analysis. This screening approach aims at supporting speech and language processing by using (1) data-driven learning and (2) robustness of connectionist networks. In order to test this approach, we have developed the SCREEN system which is based on this new robust, learned and flat analysis. In this paper, we focus on a detailed description of SCREEN's architecture, the flat syntactic and semantic analysis, the interaction with a speech recognizer, and a detailed evaluation analysis of the robustness under the influence of noisy or incomplete input. The main result of this paper is that flat representations allow more robust processing of spontaneous spoken language than deeply structured representations. In particular, we show how the fault-tolerance and learning capability of connectionist networks can support a flat analysis for providing more robust spoken-language processing within an overall hybrid symbolic/connectionist framework.
Stefan Wermter, Volker Weber
J. Artif. Intell. Res.1
1996 Learning dialog act processing
Stefan Wermter, Matthias Lochel
COLING1
1996 Towards constructive and destructive dynamic network configuration
Stefan Wermter, Manuela Meurer
ESANN1
1994 Learning Fault-Tolerant Speech Parsing with SCREEN
Stefan Wermter, Volker Weber
AAAI1
1992 A Hybrid and Connectionist Architecture for a Scanning Understanding
Stefan Wermter
ECAI1
1989 Integration of Semantic and Syntactic Constraints for Structural Noun Phrase Disambiguation
Stefan Wermter
IJCAI1