VLDB 2026 Research / reviewers in the wild / expert
Cornelius Weber
dblp:57/6298
· DBLP profile ↗
97ranked-venue papers
12as first author
34since 2021 · last 2026
0000-0001-5163-938XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 92 · 12 first-author · 31 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 since 2021Systems, architecture and hardware · 8 · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RankCut: A Ranking-Based LLM Approach to Extractive Summarization for Transcript-Based Video EditingabstractVideo recordings of interviews, lectures, and meetings contain valuable moments surrounded by less essential talk. Making a shareable and meaningful shorter version of this content requires significant effort because it combines tedious, repeated operations with personal editorial decisions, which require human judgment. We introduce an editing approach that operates on video transcripts and combines a three-stage large language model pipeline with a timeline-anchored, marker-based interface so editors can inspect and refine suggestions before final assembly. The pipeline first produces an overview summary to maximize content coverage, then induces plain-language selection rules that encode editorial intent, and finally applies rule-conditioned ranking on small transcript windows to mitigate long-context limits, yielding strictly extractive, time-aligned spans under duration constraints. The interface displays groupings of short excerpts using markers with priorities and confidence cues, converting opaque model output into verifiable units within standard video editing workflows. On MeetingBank and MeetingBank-QA datasets, our method outperforms practical extractive baselines at matched lengths. In a within-subjects study with experienced video editors familiar with Premiere Pro video editing software, we found that our marker-based interface provided editors higher efficiency, control, and satisfaction than both a manual editing baseline and an opaque auto-cut condition. Sana Shah, Mackenzie Leake, Kun Chu, Cornelius Weber, Nico Becherer, Stefan Wermter |
IUI | 4 |
| 2025 | Open-Vocabulary Robotic Object Manipulation using Foundation ModelsabstractClassical vision-language-action models are limited by unidirectional communication, hindering natural human-robot interaction.The recent CrossT5 embeds an efficient vision action pathway into an LLM, but lacks visual generalization, restricting actions to objects seen during training.We introduce OWL×T5, which integrates the OWLv2 object detection model into CrossT5 to enable robot actions on unseen objects.OWL×T5 is trained on a simulated dataset using the NICO humanoid robot and evaluated on the new CLAEO dataset featuring interactions with unseen objects.Results show that OWL×T5 achieves zero-shot object recognition for robotic manipulation, while efficiently integrating vision-language-action capabilities. Stig Griebenow, Ozan Özdemir, Cornelius Weber, Stefan Wermter |
ESANN | 3 |
| 2025 | LLM-based Interactive Imitation Learning for Robotic ManipulationabstractRecent advancements in machine learning provide methods to train autonomous agents capable of handling the increasing complexity of sequential decision-making in robotics. Imitation Learning (IL) is a prominent approach, where agents learn to control robots based on human demonstrations. However, IL commonly suffers from violating the independent and identically distributed (i.i.d) assumption in robotic tasks. Interactive Imitation Learning (IIL) achieves improved performance by allowing agents to learn from interactive feedback from human teachers. Despite these improvements, both approaches come with significant costs due to the necessity of human involvement. Leveraging the emergent capabilities of Large Language Models (LLMs) in reasoning and generating human-like responses, we introduce LLM-iTeach — a novel IIL framework that utilizes an LLM as an interactive teacher to enhance agent performance while alleviating the dependence on human resources. Firstly, LLM-iTeach uses a hierarchical prompting strategy that guides the LLM in generating a policy in Python code. Then, with a designed similarity-based feedback mechanism, LLM-iTeach provides corrective and evaluative feedback interactively during the agent’s training. We evaluate LLM-iTeach against baseline methods such as Behavior Cloning (BC), an IL method, and CEILing, a state-of-the-art IIL method using a human teacher, on various robotic manipulation tasks. Our results demonstrate that LLM-iTeach surpasses BC in the success rate and achieves or even outscores that of CEILing, highlighting the potential of LLMs as cost-effective, human-like teachers in interactive learning environments. We further demonstrate the method’s potential for generalization by evaluating it on additional tasks. The code and prompts are provided at: https://github.com/Tubicor/LLM-iTeach. Jonas Werner, Kun Chu, Cornelius Weber, Stefan Wermter |
IJCNN | 3 |
| 2025 | Disentanglement of Prosody Representations via Diffusion Models and Scheduled Gradient ReversalabstractProsody plays a fundamental role in human speech and communication, facilitating intelligibility and conveying emotional and cognitive states. Extracting accurate prosodic information from speech is vital for building assistive technology, such as controllable speech synthesis, speaking style transfer, and speech emotion recognition (SER). However, it is challenging to disentangle speaker-independent prosody representations since prosodic attributes, such as intonation, excessively entangle with speaker-specific attributes, e.g., pitch. In this article, we propose a novel model, called Diffsody, to disentangle and refine prosody representations: 1) to disentangle prosody representations, we leverage the expressive generative ability of a diffusion model by conditioning it on quantified semantic information and pretrained speaker embeddings. Additionally, a prosody encoder automatically learns prosody representations used for spectrogram reconstruction in an unsupervised fashion; and 2) to refine and learn speaker-invariant prosody representations, a scheduled gradient reversal layer (sGRL) is proposed and integrated into the prosody encoder of Diffsody. We extensively evaluate Diffsody through qualitative and quantitative means. t-SNE visualization and speaker verification experiments demonstrate the efficacy of the sGRL method in preventing speaker-specific information leakage. Experimental results on speaker-independent SER and automatic depression detection (ADD) tasks demonstrate that Diffsody can efficiently factorize speaker-independent prosody representations, resulting in a significant boost in SER and ADD. In addition, Diffsody synergistically integrates with the semantic representation model WavLM, which leads to a discernibly elevated performance, outperforming contemporary methods in both SER and ADD tasks. Furthermore, the Diffsody model exhibits promising potential for various practical applications, such as voice or style conversion. Some audio samples can be found on our https://leyuanqu.github.io/Diffsody/demo website. Leyuan Qu, Cornelius Weber, Wei Wang 0310, Jia Jin, Yingming Gao, Taihao Li, Stefan Wermter |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Enhancing Zero-Shot Chain-of-Thought Reasoning in Large Language Models through LogicabstractRecent advancements in large language models have showcased their remarkable generalizability across various domains. However, their reasoning abilities still have significant room for improvement, especially when confronted with scenarios requiring multi-step reasoning. Although large language models possess extensive knowledge, their reasoning often fails to effectively utilize this knowledge to establish a coherent thinking paradigm. These models sometimes show hallucinations as their reasoning procedures are unconstrained by logical principles. Aiming at improving the zero-shot chain-of-thought reasoning ability of large language models, we propose LoT (Logical Thoughts), a self-improvement prompting framework that leverages principles rooted in symbolic logic, particularly Reductio ad Absurdum, to systematically verify and rectify the reasoning processes step by step. Experimental evaluations conducted on language tasks in diverse domains, including arithmetic, commonsense, symbolic, causal inference, and social problems, demonstrate the efficacy of enhanced reasoning by logic. The implementation code for LoT can be accessed at: https://github.com/xf-zhao/LoT. Xufeng Zhao 0002, Mengdi Li 0006, Wenhao Lu, Cornelius Weber, Jae Hee Lee 0001, Kun Chu, Stefan Wermter |
LREC/COLING | 4 |
| 2024 | Embodying Language Models in Robot Action
Connor Gaede, Ozan Özdemir, Cornelius Weber, Stefan Wermter |
ESANN | 3 |
| 2024 | Improving Speech Emotion Recognition with Unsupervised Speaking Style TransferabstractHumans can effortlessly modify various prosodic attributes, such as the placement of stress and the intensity of sentiment, to convey a specific emotion while maintaining consistent linguistic content. Motivated by this capability, we propose EmoAug, a novel style transfer model designed to enhance emotional expression and tackle the data scarcity issue in speech emotion recognition tasks. EmoAug consists of a semantic encoder and a paralinguistic encoder that represent verbal and non-verbal information respectively. Additionally, a decoder reconstructs speech signals by conditioning on the aforementioned two information flows in an unsupervised fashion. Once training is completed, EmoAug enriches expressions of emotional speech with different prosodic attributes, such as stress, rhythm and intensity, by feeding different styles into the paralinguistic encoder. EmoAug enables us to generate similar numbers of samples for each class to tackle the data imbalance issue as well. Experimental results on the IEMOCAP dataset demonstrate that EmoAug can successfully transfer different speaking styles while retaining the speaker identity and semantic content. Furthermore, we train a SER model with data augmented by EmoAug and show that the augmented model not only surpasses the state-of-the-art supervised and self-supervised methods but also overcomes overfitting problems caused by data imbalance. Some audio samples can be found on our demo website1. Leyuan Qu, Wei Wang 0310, Cornelius Weber, Pengcheng Yue, Taihao Li, Stefan Wermter |
ICASSP | 3 |
| 2024 | Detecting Web Bots via Keystroke Dynamics
August See, Adrian Westphal, Cornelius Weber, Mathias Fischer 0001 |
SEC | 3 |
| 2024 | Disentangling Prosody Representations With Unsupervised Speech ReconstructionabstractHuman speech can be characterized by different components, including semantic content, speaker identity and prosodic information. Significant progress has been made in disentangling representations for semantic content and speaker identity in Automatic Speech Recognition (ASR) and speaker verification tasks respectively. However, it is still an open challenging research question to extract prosodic information because of the intrinsic association of different attributes, such as timbre and rhythm, and because of the need for supervised training schemes to achieve robust large-scale and speaker-independent ASR. The aim of this paper is to address the disentanglement of emotional prosody from speech based on unsupervised reconstruction. Specifically, we identify, design, implement and integrate three crucial components in our proposed speech reconstruction model Prosody2Vec: (1) a unit encoder that transforms speech signals into discrete units for semantic content, (2) a pretrained speaker verification model to generate speaker identity embeddings, and (3) a trainable prosody encoder to learn prosody representations. We first pretrain the Prosody2Vec representations on unlabelled emotional speech corpora, then fine-tune the model on specific datasets to perform Speech Emotion Recognition (SER) and Emotional Voice Conversion (EVC) tasks. Both objective (weighted and unweighted accuracies) and subjective (mean opinion score) evaluations on the EVC task suggest that Prosody2Vec effectively captures general prosodic features that can be smoothly transferred to other emotional speech. In addition, our SER experiments on the IEMOCAP dataset reveal that the prosody features learned by Prosody2Vec are complementary and beneficial for the performance of widely used speech pretraining models and surpass the state-of-the-art methods when combining Prosody2Vec with HuBERT representations. Some audio samples can be found on our demo website Leyuan Qu, Taihao Li, Cornelius Weber, Theresa Pekarek-Rosin, Fuji Ren, Stefan Wermter |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2024 | LipSound2: Self-Supervised Pre-Training for Lip-to-Speech Reconstruction and Lip ReadingabstractThe aim of this work is to investigate the impact of crossmodal self-supervised pre-training for speech reconstruction (video-to-audio) by leveraging the natural co-occurrence of audio and visual streams in videos. We propose LipSound2 that consists of an encoder-decoder architecture and location-aware attention mechanism to map face image sequences to mel-scale spectrograms directly without requiring any human annotations. The proposed LipSound2 model is first pre-trained on ∼ 2400 -h multilingual (e.g., English and German) audio-visual data (VoxCeleb2). To verify the generalizability of the proposed method, we then fine-tune the pre-trained model on domain-specific datasets (GRID and TCD-TIMIT) for English speech reconstruction and achieve a significant improvement on speech quality and intelligibility compared to previous approaches in speaker-dependent and speaker-independent settings. In addition to English, we conduct Chinese speech reconstruction on the Chinese Mandarin Lip Reading (CMLR) dataset to verify the impact on transferability. Finally, we train the cascaded lip reading (video-to-text) system by fine-tuning the generated audios on a pre-trained speech recognition system and achieve the state-of-the-art performance on both English and Chinese benchmark datasets. Leyuan Qu, Cornelius Weber, Stefan Wermter |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Visually Grounded Commonsense Knowledge AcquisitionabstractLarge-scale commonsense knowledge bases empower a broad range of AI applications, where the automatic extraction of commonsense knowledge (CKE) is a fundamental and challenging problem. CKE from text is known for suffering from the inherent sparsity and reporting bias of commonsense in text. Visual perception, on the other hand, contains rich commonsense knowledge about real-world entities, e.g., (person, can_hold, bottle), which can serve as promising sources for acquiring grounded commonsense knowledge. In this work, we present CLEVER, which formulates CKE as a distantly supervised multi-instance learning problem, where models learn to summarize commonsense relations from a bag of images about an entity pair without any human annotation on image instances. To address the problem, CLEVER leverages vision-language pre-training models for deep understanding of each image in the bag, and selects informative instances from the bag to summarize commonsense entity relations via a novel contrastive attention mechanism. Comprehensive experimental results in held-out and human evaluation show that CLEVER can extract commonsense knowledge in promising quality, outperforming pre-trained language model-based methods by 3.9 AUC and 6.4 mAUC points. The predicted commonsense scores show strong correlation with human judgment with a 0.78 Spearman coefficient. Moreover, the extracted commonsense can also be grounded into images with reasonable interpretability. The data and codes can be obtained at https://github.com/thunlp/CLEVER. Yuan Yao 0013, Tianyu Yu 0002, Mengdi Li 0006, Ruobing Xie, Cornelius Weber, Zhiyuan Liu 0001, Hai-Tao Zheng 0002, Stefan Wermter, Tat-Seng Chua, Maosong Sun 0001 |
AAAI | 6 |
| 2023 | Internally Rewarded Reinforcement LearningabstractWe study a class of reinforcement learning problems where the reward signals for policy learning are generated by a discriminator that is dependent on and jointly optimized with the policy. This interdependence between the policy and the discriminator leads to an unstable learning process because reward signals from an immature discriminator are noisy and impede policy learning, and conversely, an under-optimized policy impedes discriminator learning. We call this learning setting $\textit{Internally Rewarded Reinforcement Learning}$ (IRRL) as the reward is not provided directly by the environment but $\textit{internally}$ by the discriminator. In this paper, we formally formulate IRRL and present a class of problems that belong to IRRL. We theoretically derive and empirically analyze the effect of the reward function in IRRL and based on these analyses propose the clipped linear reward function. Experimental results show that the proposed reward function can consistently stabilize the training process by reducing the impact of reward noise, which leads to faster convergence and higher performance compared with baselines in diverse tasks. Mengdi Li 0006, Xufeng Zhao 0002, Jae Hee Lee 0001, Cornelius Weber, Stefan Wermter |
ICML | 4 |
| 2023 | Sample-Efficient Real-Time Planning with Curiosity Cross-Entropy Method and Contrastive LearningabstractModel-based reinforcement learning (MBRL) with real-time planning has shown great potential in locomotion and manipulation control tasks. However, the existing planning methods, such as the Cross-Entropy Method (CEM), do not scale well to complex high-dimensional environments. One of the key reasons for underperformance is the lack of exploration, as these planning methods only aim to maximize the cumulative extrinsic reward over the planning horizon. Furthermore, planning inside the compact latent space in the absence of observations makes it challenging to use curiosity-based intrinsic motivation. We propose Curiosity CEM (CCEM), an improved version of the CEM algorithm for encouraging exploration via curiosity. Our proposed method maximizes the sum of state-action$Q$values over the planning horizon, in which these$Q$values estimate the future extrinsic and intrinsic reward, hence encouraging to reach novel observations. In addition, our model uses contrastive representation learning to efficiently learn latent representations. Experiments on image-based continuous control tasks from the DeepMind Control suite show that CCEM is by a large margin more sample-efficient than previous MBRL algorithms and compares favorably with the best model-free RL methods. Mostafa Kotb, Cornelius Weber, Stefan Wermter |
IROS | 2 |
| 2023 | Chat with the Environment: Interactive Multimodal Perception Using Large Language ModelsabstractProgramming robot behavior in a complex world faces challenges on multiple levels, from dextrous low-level skills to high-level planning and reasoning. Recent pre-trained Large Language Models (LLMs) have shown remarkable reasoning ability in few-shot robotic planning. However, it remains challenging to ground LLMs in multimodal sensory input and continuous action output, while enabling a robot to interact with its environment and acquire novel information as its policies unfold. We develop a robot interaction scenario with a partially observable state, which necessitates a robot to decide on a range of epistemic actions in order to sample sensory information among multiple modalities, before being able to execute the task correctly. An interactive perception framework is therefore proposed with an LLM as its backbone, whose ability is exploited to instruct epistemic actions and to reason over the resulting multimodal sensations (vision, sound, haptics, proprioception), as well as to plan an entire task execution based on the interactively acquired information. Our study demonstrates that LLMs can provide high-level planning and reasoning skills and control interactive robot behavior in a multimodal environment, while multimodal modules with the context of the environmental state help ground the LLMs and extend their processing ability. The project website can be found at https://matcha-model.github.io/. Xufeng Zhao 0002, Mengdi Li 0006, Cornelius Weber, Muhammad Burhan Hafez, Stefan Wermter |
IROS | 3 |
| 2023 | Whose emotion matters? Speaking activity localisation without prior knowledgeabstractThe task of emotion recognition in conversations (ERC) benefits from the availability of multiple modalities, as provided, for example, in the video-based Multimodal EmotionLines Dataset (MELD). However, only a few research approaches use both acoustic and visual information from the MELD videos. There are two reasons for this: First, label-to-video alignments in MELD are noisy, making those videos an unreliable source of emotional speech data. Second, conversations can involve several people in the same scene, which requires the localisation of the utterance source. In this paper, we introduce MELD with Fixed Audiovisual Information via Realignment (MELD-FAIR) by using recent active speaker detection and automatic speech recognition models, we are able to realign the videos of MELD and capture the facial expressions from speakers in 96.92% of the utterances provided in MELD. Experiments with a self-supervised voice recognition model indicate that the realigned MELD-FAIR videos more closely match the transcribed utterances given in the MELD dataset. Finally, we devise a model for emotion recognition in conversations trained on the realigned MELD-FAIR videos, which outperforms state-of-the-art models for ERC based on vision alone. This indicates that localising the source of speaking activities is indeed effective for extracting facial expressions from the uttering speakers and that faces provide more informative visual cues than the visual features state-of-the-art models have been using so far. The MELD-FAIR realignment data, and the code of the realignment procedure and of the emotional recognition, are available at https://github.com/knowledgetechnologyuhh/MELD-FAIR. Hugo C. C. Carneiro, Cornelius Weber, Stefan Wermter |
Neurocomputing | 2 |
| 2023 | Emphasizing unseen words: New vocabulary acquisition for end-to-end speech recognitionabstractDue to the dynamic nature of human language, automatic speech recognition (ASR) systems need to continuously acquire new vocabulary. Out-Of-Vocabulary (OOV) words, such as trending words and new named entities, pose problems to modern ASR systems that require long training times to adapt their large numbers of parameters. Different from most previous research focusing on language model post-processing, we tackle this problem on an earlier processing level and eliminate the bias in acoustic modeling to recognize OOV words acoustically. We propose to generate OOV words using text-to-speech systems and to rescale losses to encourage neural networks to pay more attention to OOV words. Specifically, we enlarge the classification loss used for training neural networks' parameters of utterances containing OOV words (sentence-level), or rescale the gradient used for back-propagation for OOV words (word-level), when fine-tuning a previously trained model on synthetic audio. To overcome catastrophic forgetting, we also explore the combination of loss rescaling and model regularization, i.e. L2 regularization and elastic weight consolidation (EWC). Compared with previous methods that just fine-tune synthetic audio with EWC, the experimental results on the LibriSpeech benchmark reveal that our proposed loss rescaling approach can achieve significant improvement on the recall rate with only a slight decrease on word error rate. Moreover, word-level rescaling is more stable than utterance-level rescaling and leads to higher recall rates and precision rates on OOV word recognition. Furthermore, our proposed combined loss rescaling and weight consolidation methods can support continual learning of an ASR system. Leyuan Qu, Cornelius Weber, Stefan Wermter |
Neural Networks | 2 |
| 2022 | Word-by-Word Generation of Visual Dialog Using Reinforcement Learning
Yuliia Lysa, Cornelius Weber, Dennis Becker, Stefan Wermter |
ICANN (2) | 2 |
| 2022 | Learning Flexible Translation Between Robot Actions and Language Descriptions
Ozan Özdemir, Matthias Kerzel, Cornelius Weber, Jae Hee Lee 0001, Stefan Wermter |
ICANN (2) | 3 |
| 2022 | NeoSLAM: Neural Object SLAM for Loop Closure and Navigation
Younes Raoui, Cornelius Weber, Stefan Wermter |
ICANN (3) | 2 |
| 2022 | Learning Visually Grounded Human-Robot Dialog in a Hybrid Neural Architecture
Xiaowen Sun, Cornelius Weber, Matthias Kerzel, Tom Weber, Mengdi Li 0006, Stefan Wermter |
ICANN (2) | 2 |
| 2022 | More Diverse Training, Better Compositionality! Evidence from Multimodal Language Learning
Caspar Volquardsen, Jae Hee Lee 0001, Cornelius Weber, Stefan Wermter |
ICANN (3) | 3 |
| 2022 | What is Right for Me is Not Yet Right for You: A Dataset for Grounding Relative Directions via Multi-Task LearningabstractUnderstanding spatial relations is essential for intelligent agents to act and communicate in the physical world. Relative directions are spatial relations that describe the relative positions of target objects with regard to the intrinsic orientation of reference objects. Grounding relative directions is more difficult than grounding absolute directions because it not only requires a model to detect objects in the image and to identify spatial relation based on this information, but it also needs to recognize the orientation of objects and integrate this information into the reasoning process. We investigate the challenging problem of grounding relative directions with end-to-end neural networks. To this end, we provide GRiD-3D, a novel dataset that features relative directions and complements existing visual question answering (VQA) datasets, such as CLEVR, that involve only absolute directions. We also provide baselines for the dataset with two established end-to-end VQA models. Experimental evaluations show that answering questions on relative directions is feasible when questions in the dataset simulate the necessary subtasks for grounding relative directions. We discover that those subtasks are learned in an order that reflects the steps of an intuitive pipeline for processing relative directions. Jae Hee Lee 0001, Matthias Kerzel, Kyra Ahrens, Cornelius Weber, Stefan Wermter |
IJCAI | 4 |
| 2022 | Impact Makes a Sound and Sound Makes an Impact: Sound Guides Representations and ExplorationsabstractSound is one of the most informative and abundant modalities in the real world while being robust to sense without contacts by small and cheap sensors that can be placed on mobile devices. Although deep learning is capable of extracting information from multiple sensory inputs, there has been little use of sound for the control and learning of robotic actions. For unsupervised reinforcement learning, an agent is expected to actively collect experiences and jointly learn representations and policies in a self-supervised way. We build realistic robotic manipulation scenarios with physics-based sound simulation and propose the Intrinsic Sound Curiosity Module (ISCM). The ISCM provides feedback to a reinforcement learner to learn robust representations and to reward a more efficient exploration behavior. We perform experiments with sound enabled during pre-training and disabled during adaptation, and show that representations learned by ISCM outperform the ones by vision-only baselines and pre-trained policies can accelerate the learning process when applied to downstream tasks. Xufeng Zhao 0002, Cornelius Weber, Muhammad Burhan Hafez, Stefan Wermter |
IROS | 2 |
| 2022 | A Multimodal German Dataset for Automatic Lip Reading Systems and Transfer LearningabstractLarge datasets as required for deep learning of lip reading do not exist in many languages. In this paper we present the dataset GLips (German Lips) consisting of 250,000 publicly available videos of the faces of speakers of the Hessian Parliament, which was processed for word-level lip reading using an automatic pipeline. The format is similar to that of the English language LRW (Lip Reading in the Wild) dataset, with each video encoding one word of interest in a context of 1.16 seconds duration, which yields compatibility for studying transfer learning between both datasets. By training a deep neural network, we investigate whether lip reading has language-independent features, so that datasets of different languages can be used to improve lip reading models. We demonstrate learning from scratch and show that transfer learning from LRW to GLips and vice versa improves learning speed and performance, in particular for the validation set. Gerald Schwiebert, Cornelius Weber, Leyuan Qu, Henrique Siqueira, Stefan Wermter |
LREC | 2 |
| 2022 | Go ahead and do not forget: Modular lifelong learning from event-based dataabstractLifelong learning is a long-standing aim for artificial agents that act in dynamic environments in which an agent needs to accumulate knowledge incrementally without forgetting previously learned representations. Contemporary methods for incremental learning from images are predominantly based on frame-based data recorded by conventional shutter cameras. We investigate methods for learning from data produced by event cameras and compare techniques to mitigate forgetting while learning incrementally. We propose a model that is composed of both, feature extraction and incremental learning. The feature extractor is utilized as a self-supervised sparse convolutional neural network that processes event-based data. The incremental learner uses a habituation-based method that works in tandem with other existing techniques. Our experimental results show that the combination of different existing techniques with our proposed habituation-based method can help avoid catastrophic forgetting even more, while learning incrementally from the features provided by the extraction module. Vadym Gryshchuk, Cornelius Weber, Chu Kiong Loo, Stefan Wermter |
Neurocomputing | 2 |
| 2021 | Hearing Faces: Target Speaker Text-to-Speech Synthesis from a FaceabstractThe existence of a learnable cross-modal association between a person's face and their voice is recently becoming more and more evident. This provides the basis for the task of target speaker text-to-speech (TTS) synthesis from face ref-erence. In this paper, we approach this task by proposing a cross-modal model architecture combining existing unimodal models. We use Tacotron 2 multi-speaker TTS with auditory speaker embeddings based on Global Style Tokens. We trans-fer learn a FaceNet face encoder to predict these embeddings from a static face image reference instead of a voice reference and thus predict a speaker's voice and speaking characteristics from their face. Compared to Face2Speech, the only existing work on this task, we use a more modular architecture that allows the use of openly available and pretrained model components. This approach enables high-quality speech synthesis and allows for an easily extensible model architecture. Ex-perimental results show good matching ability while retaining better voice naturalness than Face2Speech. We examine the limitations of our model and discuss multiple possible av-enues of improvement for future work. Björn Plüster, Cornelius Weber, Leyuan Qu, Stefan Wermter |
ASRU | 2 |
| 2021 | Lifelong Learning from Event-based DataabstractLifelong learning is a long-standing aim for artificial agents that act in dynamic environments, in which an agent needs to accumulate knowledge incrementally without forgetting previously learned representations.We investigate methods for learning from data produced by event cameras and compare techniques to mitigate forgetting while learning incrementally.We propose a model that is composed of both, feature extraction and continuous learning.Furthermore, we introduce a habituationbased method to mitigate forgetting.Our experimental results show that the combination of different techniques can help to avoid catastrophic forgetting while learning incrementally from the features provided by the extraction module. Vadym Gryshchuk, Cornelius Weber, Chu Kiong Loo, Stefan Wermter |
ESANN | 2 |
| 2021 | FaVoA: Face-Voice Association Favours Ambiguous Speaker Detection
Hugo C. C. Carneiro, Cornelius Weber, Stefan Wermter |
ICANN (1) | 2 |
| 2021 | Visual Distant Supervision for Scene Graph GenerationabstractScene graph generation aims to identify objects and their relations in images, providing structured image representations that can facilitate numerous applications in computer vision. However, scene graph models usually require supervised learning on large quantities of labeled data with intensive human annotation. In this work, we propose visual distant supervision, a novel paradigm of visual relation learning, which can train scene graph models without any human-labeled data. The intuition is that by aligning commonsense knowledge bases and images, we can automatically create large-scale labeled data to provide distant supervision for visual relation learning. To alleviate the noise in distantly labeled data, we further propose a framework that iteratively estimates the probabilistic relation labels and eliminates the noisy ones. Comprehensive experimental results show that our distantly supervised model outperforms strong weakly supervised and semi-supervised baselines. By further incorporating human-labeled data in a semi-supervised fashion, our model outperforms state-of-the-art fully supervised models by a large margin (e.g., 8.3 micro- and 7.8 macro-recall@50 improvements for predicate classification in Visual Genome evaluation). We make the data and code for this paper publicly available at https://github.com/thunlp/VisualDS. Yuan Yao 0011, Xu Han 0007, Mengdi Li 0006, Cornelius Weber, Zhiyuan Liu 0001, Stefan Wermter, Maosong Sun 0001 |
ICCV | 5 |
| 2021 | Generalization in Multimodal Language Learning from SimulationabstractNeural networks can be powerful function approximators, which are able to model high-dimensional feature distributions from a subset of examples drawn from the target distribution. Naturally, they perform well at generalizing within the limits of their target function, but they often fail to generalize outside of the explicitly learned feature space. It is therefore an open research topic whether and how neural network-based architectures can be deployed for systematic reasoning. Many studies have shown evidence for poor generalization, but they often work with abstract data or are limited to single-channel input. Humans, however, learn and interact through a combination of multiple sensory modalities, and rarely rely on just one. To investigate compositional generalization in a multimodal setting, we generate an extensible dataset with multimodal input sequences from simulation. We investigate the influence of the underlying training data distribution on compostional generalization in a minimal LSTM-based network trained in a supervised, time continuous setting. We find compositional generalization to fail in simple setups while improving with the number of objects, actions, and particularly with a lot of color overlaps between objects. Furthermore, multimodality strongly improves compositional generalization in settings where a pure vision model struggles to generalize. Aaron Eisermann, Jae Hee Lee 0001, Cornelius Weber, Stefan Wermter |
IJCNN | 3 |
| 2021 | Controlling the Noise Robustness of End-to-End Automatic Speech Recognition SystemsabstractIn this work, we propose a novel training scheme to modularize end-to-end systems. Our training scheme aims at altering the flow of information in an end-to-end system to use the kernels of this system for another system that fulfills another task. We apply this scheme to extract the noise reduction capabilities from a noise-robust automatic speech recognition (ASR) system and implement a speech enhancer from it. This enhancer receives spectral representations from unfiltered audio and outputs cleaned spectral representations. Our enhancer can be integrated into an ASR system as front-end, is trainable, and reduces background noise. Our front-end uses a decoder to clean speech based on the hidden activations of the ASR system Jasper. While training, we exclusively adapt the weights in our decoder and the batch normalization in Jasper. The resulting spectral representations show less background noise. Further, areas in the spectral features are not reconstructed if they do not contribute to speech recognition. We demonstrate that our front-end can be combined with a pre-trained ASR system as back-end and supports speech recognition in noisy conditions. Further, we show that training another ASR system with our front-end results in an increased performance of the ASR system in noisy as well as noiseless conditions. The ASR system's performance is especially improved on challenging speech datasets. Matthias Möller, Johannes Twiefel, Cornelius Weber, Stefan Wermter |
IJCNN | 3 |
| 2021 | Improving Model-Based Reinforcement Learning with Internal State Representations through Self-SupervisionabstractUsing a model of the environment, reinforcement learning agents can plan their future moves and achieve superhuman performance in board games like Chess, Shogi, and Go, while remaining relatively sample-efficient. As demonstrated by the MuZero Algorithm, the environment model can even be learned dynamically, generalizing the agent to many more tasks while at the same time achieving state-of-the-art performance. Notably, MuZero uses internal state representations derived from real environment states for its predictions. In this paper, we bind the model's predicted internal state representation to the environment state via two additional terms: a reconstruction model loss and a simpler consistency loss, both of which work independently and unsupervised, acting as constraints to stabilize the learning process. Our experiments show that this new integration of reconstruction model loss and simpler consistency loss provide a significant performance increase in OpenAI Gym environments. Our modifications also enable self-supervised pretraining for MuZero, so the algorithm can learn about environment dynamics before a goal is made available. Julien Scholz, Cornelius Weber, Muhammad Burhan Hafez, Stefan Wermter |
IJCNN | 2 |
| 2021 | Planning-integrated Policy for Efficient Reinforcement Learning in Sparse-reward EnvironmentsabstractModel-free reinforcement learning algorithms can learn an optimal policy from experience without requiring prior knowledge. However, model-free agents require vast amounts of samples, particularly in sparse reward environments where most states contain zero rewards. We developed a model-based approach to tackle the high sample complexity problem in sparse reward settings with continuous actions. A trained world model is queried by a particle swarm optimization (PSO) planner and employed as the action selection mechanism, hence taking the role of the actor in an actor-critic architecture. Parameters of the PSO regulate the agent's exploration rate. We show that the planner aids the agent to discover rewards even in regions with zero value gradient. Our simple planning integrated policy architecture learns more efficiently with fewer samples than continuous model-free algorithms. Christoper Wulur, Cornelius Weber, Stefan Wermter |
IJCNN | 2 |
| 2021 | Robotic Occlusion Reasoning for Efficient Object Existence PredictionabstractReasoning about potential occlusions is essential for robots to efficiently predict whether an object exists in an environment. Though existing work shows that a robot with active perception can achieve various tasks, it is still unclear if occlusion reasoning can be achieved. To answer this question, we introduce the task of robotic object existence prediction: when being asked about an object, a robot needs to move as few steps as possible around a table with randomly placed objects to predict whether the queried object exists. To address this problem, we propose a novel recurrent neural network model that can be jointly trained with supervised and reinforcement learning methods using a curriculum training strategy. Experimental results show that 1) both active perception and occlusion reasoning are necessary to successfully achieve the task; 2) the proposed model demonstrates a good occlusion reasoning ability by achieving a similar prediction accuracy to an exhaustive exploration baseline while requiring only about 10% of the baseline’s number of movement steps on average; and 3) the model generalizes to novel object combinations with a moderate loss of accuracy. Mengdi Li 0006, Cornelius Weber, Matthias Kerzel, Jae Hee Lee 0001, Zheni Zeng, Zhiyuan Liu 0001, Stefan Wermter |
IROS | 2 |
| 2020 | Variational Autoencoder with Global- and Medium Timescale Auxiliaries for Emotion Recognition from Speech
Hussam Almotlak, Cornelius Weber, Leyuan Qu, Stefan Wermter |
ICANN (1) | 2 |
| 2020 | Neural Networks for Detecting Irrelevant Questions During Visual Question Answering
Mengdi Li 0006, Cornelius Weber, Stefan Wermter |
ICANN (2) | 2 |
| 2020 | Multimodal Target Speech Separation with Voice and Face ReferencesabstractTarget speech separation refers to isolating target speech from a multi-speaker mixture signal by conditioning on auxiliary information about the target speaker. Different from the mainstream audio-visual approaches which usually require simultaneous visual streams as additional input, e.g. the corresponding lip movement sequences, in our approach we propose the novel use of a single face profile of the target speaker to separate expected clean speech. We exploit the fact that the image of a face contains information about the person's speech sound. Compared to using a simultaneous visual sequence, a face image is easier to obtain by pre-enrollment or on websites, which enables the system to generalize to devices without cameras. To this end, we incorporate face embeddings extracted from a pretrained model for face recognition into the speech separation, which guide the system in predicting a target speaker mask in the time-frequency domain. The experimental results show that a pre-enrolled face image is able to benefit separating expected speech signals. Additionally, face information is complementary to voice reference and we show that further improvement can be achieved when combing both face and voice embeddings. Leyuan Qu, Cornelius Weber, Stefan Wermter |
INTERSPEECH | 2 |
| 2020 | EDA: Enriching Emotional Dialogue Acts using an Ensemble of Neural AnnotatorsabstractThe recognition of emotion and dialogue acts enriches conversational analysis and help to build natural dialogue systems. Emotion interpretation makes us understand feelings and dialogue acts reflect the intentions and performative functions in the utterances. However, most of the textual and multi-modal conversational emotion corpora contain only emotion labels but not dialogue acts. To address this problem, we propose to use a pool of various recurrent neural models trained on a dialogue act corpus, with and without context. These neural models annotate the emotion corpora with dialogue act labels, and an ensemble annotator extracts the final dialogue act label. We annotated two accessible multi-modal emotion corpora: IEMOCAP and MELD. We analyzed the co-occurrence of emotion and dialogue act labels and discovered specific relations. For example, Accept/Agree dialogue acts often occur with the Joy emotion, Apology with Sadness, and Thanking with Joy. We make the Emotional Dialogue Acts (EDA) corpus publicly available to the research community for further study and analysis. Chandrakant Bothe, Cornelius Weber, Sven Magg, Stefan Wermter |
LREC | 2 |
| 2019 | Learning Sparse Hidden States in Long Short-Term Memory
Niange Yu, Cornelius Weber, Xiaolin Hu 0001 |
ICANN (2) | 2 |
| 2019 | Incorporating End-to-End Speech Recognition Models for Sentiment AnalysisabstractPrevious work on emotion recognition demonstrated a synergistic effect of combining several modalities such as auditory, visual, and transcribed text to estimate the affective state of a speaker. Among these, the linguistic modality is crucial for the evaluation of an expressed emotion. However, manually transcribed spoken text cannot be given as input to a system practically. We argue that using ground-truth transcriptions during training and evaluation phases leads to a significant discrepancy in performance compared to real-world conditions, as the spoken text has to be recognized on the fly and can contain speech recognition mistakes. In this paper, we propose a method of integrating an automatic speech recognition (ASR) output with a character-level recurrent neural network for sentiment recognition. In addition, we conduct several experiments investigating sentiment recognition for human-robot interaction in a noise-realistic scenario which is challenging for the ASR systems. We quantify the improvement compared to using only the acoustic modality in sentiment recognition. We demonstrate the effectiveness of this approach on the Multimodal Corpus of Sentiment Intensity (MOSI) by achieving 73,6% accuracy in a binary sentiment classification task, exceeding previously reported results that use only acoustic input. In addition, we set a new state-of-the-art performance on the MOSI dataset (80.4% accuracy, 2% absolute improvement). Egor Lakomkin, Mohammad-Ali Zamani, Cornelius Weber, Sven Magg, Stefan Wermter |
ICRA | 3 |
| 2019 | Designing a Personality-Driven Robot for a Human-Robot Interaction ScenarioabstractIn this paper, we present an autonomous AI system designed for a Human-Robot Interaction (HRI) study, set around a dice game scenario. We conduct a case study to answer our research question: Does a robot with a socially engaged personality lead to a higher acceptance than a competitive personality? The flexibility of our proposed system allows us to construct and attribute two different personalities to a humanoid robot: a socially engaged personality that maximizes its user interaction and a competitive personality that is focused on playing and winning the game. We evaluate both personalities in a user study, in which the participants play a turn-taking dice game with the robot. Each personality is assessed with four different evaluation tools: 1) the Godspeed Questionnaire, 2) the Mind Perception Questionnaire, 3) a custom questionnaire concerning the overall HRI experience, and 4) a Convolutional Neural Network analyzing the emotions on the participants' facial feedback throughout the game. Our results show that the socially engaged personality evokes stronger emotions among the participants and is rated higher in likability and animacy than the competitive one. We conclude that designing the robot with a socially engaged personality contributes to a higher acceptance within an HRI scenario. Hadi Beik-Mohammadi, Nikoletta Xirakia, Fares Abawi, Irina Barykina, Krishnan Chandran, Gitanjali Nair, Daniel Speck, Tayfun Alpay, Sascha S. Griffiths, Stefan Heinrich, Erik Strahl, Cornelius Weber, Stefan Wermter |
ICRA | 13 |
| 2019 | Curious Meta-Controller: Adaptive Alternation between Model-Based and Model-Free Control in Deep Reinforcement LearningabstractRecent success in deep reinforcement learning for continuous control has been dominated by model-free approaches which, unlike model-based approaches, do not suffer from representational limitations in making assumptions about the world dynamics and model errors inevitable in complex domains. However, they require a lot of experiences compared to model-based approaches that are typically more sample-efficient. We propose to combine the benefits of the two approaches by presenting an integrated approach called Curious Meta-Controller. Our approach alternates adaptively between model-based and model-free control using a curiosity feedback based on the learning progress of a neural model of the dynamics in a learned latent space. We demonstrate that our approach can significantly improve the sample efficiency and achieve near-optimal performance on learning robotic reaching and grasping tasks from raw-pixel input in both dense and sparse reward settings. Muhammad Burhan Hafez, Cornelius Weber, Matthias Kerzel, Stefan Wermter |
IJCNN | 2 |
| 2019 | Effect of Pruning on Catastrophic Forgetting in Growing Dual Memory NetworksabstractGrow-when-required networks such as the Growing Dual-Memory (GDM) networks possess a dynamic network structure, expanding to accommodate new neurons in response to learning novel concepts. Over time, it may be necessary to prune obsolete neurons and/or neural connections to meet performance or resource limitations. GDM networks utilize an age-based pruning strategy, whereby older neurons and neural connections that have not been activated recently are removed. Catastrophic forgetting occurs when knowledge learned by the networks in previous learning iterations is lost due to being overwritten by newer learning iterations, or to the pruning process. In this work, we investigate catastrophic forgetting in GDM networks in response to different pruning strategies. The age-based pruning method was shown to significantly sparsify the GDM network topology while improving the networks ability to recall newly acquired concepts with a slight decrease in performance with respect to older knowledge. A significance-based pruning method was tested as a replacement for the age-based pruning, but was not as effective at pruning even though it performed better at recalling older knowledge. Wei Shiung Liew, Chu Kiong Loo, Vadym Gryshchuk, Cornelius Weber, Stefan Wermter |
IJCNN | 4 |
| 2019 | LipSound: Neural Mel-Spectrogram Reconstruction for Lip Reading
Leyuan Qu, Cornelius Weber, Stefan Wermter |
INTERSPEECH | 2 |
| 2019 | Predictive Auxiliary Variational Autoencoder for Representation Learning of Global Speech Characteristics
Sebastian Springenberg, Egor Lakomkin, Cornelius Weber, Stefan Wermter |
INTERSPEECH | 3 |
| 2018 | Slowness-based neural visuomotor control with an Intrinsically motivated Continuous Actor-Critic
Muhammad Burhan Hafez, Matthias Kerzel, Cornelius Weber, Stefan Wermter |
ESANN | 3 |
| 2018 | A Sub-Layered Hierarchical Pyramidal Neural Architecture for Facial Expression Recognition
Henrique Siqueira, Pablo V. A. Barros, Sven Magg, Cornelius Weber, Stefan Wermter |
ESANN | 4 |
| 2018 | Image-to-Text Transduction with Spatial Self-Attention
Sebastian Springenberg, Egor Lakomkin, Cornelius Weber, Stefan Wermter |
ESANN | 3 |
| 2018 | Combining Articulatory Features with End-to-End Learning in Speech Recognition
Leyuan Qu, Cornelius Weber, Egor Lakomkin, Johannes Twiefel, Stefan Wermter |
ICANN (3) | 2 |
| 2018 | A Hybrid Planning Strategy Through Learning from Vision for Target-Directed Navigation
Xiaomao Zhou, Cornelius Weber, Chandrakant Bothe, Stefan Wermter |
ICANN (2) | 2 |
| 2018 | EmoRL: Continuous Acoustic Emotion Classification Using Deep Reinforcement LearningabstractAcoustically expressed emotions can make communication with a robot more efficient. Detecting emotions like anger could provide a clue for the robot indicating unsafe/undesired situations. Recently, several deep neural network-based models have been proposed which establish new state-of-the-art results in affective state evaluation. These models typically start processing at the end of each utterance, which not only requires a mechanism to detect the end of an utterance but also makes it difficult to use them in a real-time communication scenario, e.g. human-robot interaction. We propose the EmoRL model that triggers an emotion classification as soon as it gains enough confidence while listening to a person speaking. As a result, we minimize the need for segmenting the audio signal for classification and achieve lower latency as the audio signal is processed incrementally. The method is competitive with the accuracy of a strong baseline model, while allowing much earlier prediction. Egor Lakomkin, Mohammad-Ali Zamani, Cornelius Weber, Sven Magg, Stefan Wermter |
ICRA | 3 |
| 2018 | A Self-organizing Method for Robot Navigation based on Learned Place and Head-Direction CellsabstractThis paper describes a neural model for a robot learning spatial knowledge and navigating on learned place and head-direction (HD) cell representations. The place and HD cells, which are trained through unsupervised slow feature analysis (SFA) from sequences of visual stimuli, provide positional and directional information for navigation. Based on the ensemble activity of place cells, the robot learns a topological map of the environment through extracting the statistical distribution of the place cell activities covering the traversable areas and realizes self-localization based on the map. The robot's heading direction, which is encoded by the HD cells, works as a control signal to adjust its behavior. Action representations supporting state transitions are learned through memorizing the same movement from a previous phase where an experimenter drives a robot to explore an environment. Given reward signals spreading from a target location along the topological map, the robot can reach the goal in a reward-ascending way. This work intends to build a practical navigation system by simulating animals' hippocampal cell firing activities on a robot platform using its self-contained sensor. Experimental results from simulation demonstrate that our system navigates a robot to the desired position smoothly and effectively. Xiaomao Zhou, Cornelius Weber, Stefan Wermter |
IJCNN | 2 |
| 2018 | Conversational Analysis Using Utterance-level Attention-based Bidirectional Recurrent Neural NetworksabstractRecent approaches for dialogue act recognition have shown that context from preceding utterances is important to classify the subsequent one. It was shown that the performance improves rapidly when the context is taken into account. We propose an utterance-level attention-based bidirectional recurrent neural network (Utt-Att-BiRNN) model to analyze the importance of preceding utterances to classify the current one. In our setup, the BiRNN is given the input set of current and preceding utterances. Our model outperforms previous models that use only preceding utterances as context on the used corpus. Another contribution of the article is to discover the amount of information in each utterance to classify the subsequent one and to show that context-based learning not only improves the performance but also achieves higher confidence in the classification. We use character- and word-level features to represent the utterances. The results are presented for character and word feature representations and as an ensemble model of both representations. We found that when classifying short utterances, the closest preceding utterances contributes to a higher degree. Chandrakant Bothe, Sven Magg, Cornelius Weber, Stefan Wermter |
INTERSPEECH | 3 |
| 2018 | On the Robustness of Speech Emotion Recognition for Human-Robot Interaction with Deep Neural NetworksabstractSpeech emotion recognition (SER) is an important aspect of effective human-robot collaboration and received a lot of attention from the research community. For example, many neural network-based architectures were proposed recently and pushed the performance to a new level. However, the applicability of such neural SER models trained only on in-domain data to noisy conditions is currently under-researched. In this work, we evaluate the robustness of state-of-the-art neural acoustic emotion recognition models in human-robot interaction scenarios. We hypothesize that a robot's ego noise, room conditions, and various acoustic events that can occur in a home environment can significantly affect the performance of a model. We conduct several experiments on the iCub robot platform and propose several novel ways to reduce the gap between the model's performance during training and testing in real-world conditions. Furthermore, we observe large improvements in the model performance on the robot and demonstrate the necessity of introducing several data augmentation techniques like overlaying background noise and loudness variations to improve the robustness of the neural approaches. Egor Lakomkin, Mohammad-Ali Zamani, Cornelius Weber, Sven Magg, Stefan Wermter |
IROS | 3 |
| 2018 | A Context-based Approach for Dialogue Act Recognition using Simple Recurrent Neural Networks
Chandrakant Bothe, Cornelius Weber, Sven Magg, Stefan Wermter |
LREC | 2 |
| 2017 | The Impact of Personalisation on Human-Robot Interaction in Learning ScenariosabstractAdvancements in Human-Robot Interaction involve robots being more responsive and adaptive to the human user they are interacting with. For example, robots model a personalised dialogue with humans, adapting the conversation to accommodate the user's preferences in order to allow natural interactions. This study investigates the impact of such personalised interaction capabilities of a human companion robot on its social acceptance, perceived intelligence and likeability in a human-robot interaction scenario. In order to measure this impact, the study makes use of an object learning scenario where the user teaches different objects to the robot using natural language. An interaction module is built on top of the learning scenario which engages the user in a personalised conversation before teaching the robot to recognise different objects. The two systems, i.e. with and without the interaction module, are compared with respect to how different users rate the robot on its intelligence and sociability. Although the system equipped with personalised interaction capabilities is rated lower on social acceptance, it is perceived as more intelligent and likeable by the users. Nikhil Churamani, Paul Anton, Marc Brügger, Erik Fließwasser, Thomas Hummel 0001, Julius Mayer 0001, Waleed Mustafa, Hwei Geok Ng, Thi Linh Chi Nguyen, Quan Nguyen 0005, Marcus Soll, Sebastian Springenberg, Sascha S. Griffiths, Stefan Heinrich, Nicolás Navarro-Guerrero, Erik Strahl, Johannes Twiefel, Cornelius Weber, Stefan Wermter |
HAI | 18 |
| 2017 | Dialogue-Based Neural Learning to Estimate the Sentiment of a Next Upcoming Utterance
Chandrakant Bothe, Sven Magg, Cornelius Weber, Stefan Wermter |
ICANN (2) | 3 |
| 2017 | Robot Localization and Orientation Detection Based on Place Cells and Head-Direction Cells
Xiaomao Zhou, Cornelius Weber, Stefan Wermter |
ICANN (1) | 2 |
| 2017 | Reusing Neural Speech Representations for Auditory Emotion RecognitionabstractAcoustic emotion recognition aims to categorize the affective state of the speaker and is still a difficult task for machine learning models. The difficulties come from the scarcity of training data, general subjectivity in emotion perception resulting in low annotator agreement, and the uncertainty about which features are the most relevant and robust ones for classification. In this paper, we will tackle the latter problem. Inspired by the recent success of transfer learning methods we propose a set of architectures which utilize neural representations inferred by training on large speech databases for the acoustic emotion recognition task. Our experiments on the IEMOCAP dataset show ~10% relative improvements in the accuracy and F1-score over the baseline recurrent neural network which is trained end-to-end for emotion recognition. Egor Lakomkin, Cornelius Weber, Sven Magg, Stefan Wermter |
IJCNLP(1) | 2 |
| 2017 | Hey robot, why don't you talk to me?abstractThis paper describes the techniques used in the submitted video presenting an interaction scenario, realised using the Neuro-Inspired Companion (NICO) robot. NICO engages the users in a personalised conversation where the robot always tracks the users' face, remembers them and interacts with them using natural language. NICO can also learn to perform tasks such as remembering and recalling objects and thus can assist users in their daily chores. The interaction system helps the users to interact as naturally as possible with the robot, enriching their experience with the robot, making it more interesting and engaging. Hwei Geok Ng, Paul Anton, Marc Brügger, Nikhil Churamani, Erik Fließwasser, Thomas Hummel 0001, Julius Mayer 0001, Waleed Mustafa, Thi Linh Chi Nguyen, Quan Nguyen 0005, Marcus Soll, Sebastian Springenberg, Sascha S. Griffiths, Stefan Heinrich, Nicolás Navarro-Guerrero, Erik Strahl, Johannes Twiefel, Cornelius Weber, Stefan Wermter |
RO-MAN | 18 |
| 2017 | Emotion-modulated attention improves expression recognition: A deep learning modelabstractSpatial attention in humans and animals involves the visual pathway and the superior colliculus, which integrate multimodal information. Recent research has shown that affective stimuli play an important role in attentional mechanisms, and behavioral studies show that the focus of attention in a given region of the visual field is increased when affective stimuli are present. This work proposes a neurocomputational model that learns to attend to emotional expressions and to modulate emotion recognition. Our model consists of a deep architecture which implements convolutional neural networks to learn the location of emotional expressions in a cluttered scene. We performed a number of experiments for detecting regions of interest, based on emotion stimuli, and show that the attention model improves emotion expression recognition when used as emotional attention modulator. Finally, we analyze the internal representations of the learned neural filters and discuss their role in the performance of our model. Pablo V. A. Barros, German Ignacio Parisi, Cornelius Weber, Stefan Wermter |
Neurocomputing | 3 |
| 2017 | An analysis of Convolutional Long Short-Term Memory Recurrent Neural Networks for gesture recognitionabstractIn this research, we analyze a Convolutional Long Short-Term Memory Recurrent Neural Network (CNNLSTM) in the context of gesture recognition. CNNLSTMs are able to successfully learn gestures of varying duration and complexity. For this reason, we analyze the architecture by presenting a qualitative evaluation of the model, based on the visualization of the internal representations of the convolutional layers and on the examination of the temporal classification outputs at a frame level, in order to check if they match the cognitive perception of a gesture. We show that CNNLSTM learns the temporal evolution of the gestures classifying correctly their meaningful part, known as Kendon’s stroke phase. With the visualization, for which we use the deconvolution process that maps specific feature map activations to original image pixels, we show that the network learns to detect the most intense body motion. Finally, we show that CNNLSTM outperforms both plain CNN and LSTM in gesture recognition. Eleni Tsironi, Pablo V. A. Barros, Cornelius Weber, Stefan Wermter |
Neurocomputing | 3 |
| 2017 | Lifelong learning of human actions with deep neural network self-organizationabstractLifelong learning is fundamental in autonomous robotics for the acquisition and fine-tuning of knowledge through experience. However, conventional deep neural models for action recognition from videos do not account for lifelong learning but rather learn a batch of training data with a predefined number of action classes and samples. Thus, there is the need to develop learning systems with the ability to incrementally process available perceptual cues and to adapt their responses over time. We propose a self-organizing neural architecture for incrementally learning to classify human actions from video sequences. The architecture comprises growing self-organizing networks equipped with recurrent neurons for processing time-varying patterns. We use a set of hierarchically arranged recurrent networks for the unsupervised learning of action representations with increasingly large spatiotemporal receptive fields. Lifelong learning is achieved in terms of prediction-driven neural dynamics in which the growth and the adaptation of the recurrent networks are driven by their capability to reconstruct temporally ordered input sequences. Experimental results on a classification task using two action benchmark datasets show that our model is competitive with state-of-the-art methods for batch learning also when a significant number of sample labels are missing or corrupted during training sessions. Additional experiments show the ability of our model to adapt to non-stationary input avoiding catastrophic interference. German Ignacio Parisi, Jun Tani, Cornelius Weber, Stefan Wermter |
Neural Networks | 3 |
| 2016 | Learning auditory neural representations for emotion recognitionabstractAuditory emotion recognition has become a very important topic in recent years. However, still after the development of some architectures and frameworks, generalization is a big problem. Our model examines the capability of deep neural networks to learn specific features for different kinds of auditory emotion recognition: speech and music-based recognition. We propose the use of a cross-channel architecture to improve the generalization aspects of complex auditory recognition by the integration of previously learned knowledge of specific representation into a high-level auditory descriptor. We evaluate our models using the SAVEE dataset, the GTZAN dataset and the EmotiW corpus, and show comparable results with state-of-the-art approaches. Pablo V. A. Barros, Cornelius Weber, Stefan Wermter |
IJCNN | 2 |
| 2016 | Ball Localization for Robocup Soccer Using Convolutional Neural Networks
Daniel Speck, Pablo V. A. Barros, Cornelius Weber, Stefan Wermter |
RoboCup | 3 |
| 2015 | Interactive reinforcement learning through speech guidance in a domestic scenarioabstractRecently robots are being used more frequently as assistants in domestic scenarios. In this context we train an apprentice robot to perform a cleaning task using interactive reinforcement learning since it has been shown to be an efficient learning approach benefiting from human expertise for performing domestic tasks. The robotic agent obtains interactive feedback via a speech recognition system which is tested to work with five different microphones concerning their polar patterns and distance to the teacher to recognize sentences in different instruction classes. Moreover, the reinforcement learning approach uses situated affordances to allow the robot to complete the cleaning task in every episode anticipating when chosen actions are possible to be performed. Situated affordances and interaction allow to improve the convergence speed of reinforcement learning, and the results also show that the system is robust against wrong instructions that result from errors of the speech recognition system. Francisco Cruz 0002, Johannes Twiefel, Sven Magg, Cornelius Weber, Stefan Wermter |
IJCNN | 4 |
| 2015 | Multimodal emotional state recognition using sequence-dependent deep hierarchical featuresabstractEmotional state recognition has become an important topic for human-robot interaction in the past years. By determining emotion expressions, robots can identify important variables of human behavior and use these to communicate in a more human-like fashion and thereby extend the interaction possibilities. Human emotions are multimodal and spontaneous, which makes them hard to be recognized by robots. Each modality has its own restrictions and constraints which, together with the non-structured behavior of spontaneous expressions, create several difficulties for the approaches present in the literature, which are based on several explicit feature extraction techniques and manual modality fusion. Our model uses a hierarchical feature representation to deal with spontaneous emotions, and learns how to integrate multiple modalities for non-verbal emotion recognition, making it suitable to be used in an HRI scenario. Our experiments show that a significant improvement of recognition accuracy is achieved when we use hierarchical features and multimodal information, and our model improves the accuracy of state-of-the-art approaches from 82.5% reported in the literature to 91.3% for a benchmark dataset on spontaneous emotion expressions. Pablo V. A. Barros, Doreen Jirak, Cornelius Weber, Stefan Wermter |
Neural Networks | 3 |
| 2014 | A Multichannel Convolutional Neural Network for Hand Posture Recognition
Pablo V. A. Barros, Sven Magg, Cornelius Weber, Stefan Wermter |
ICANN | 3 |
| 2014 | RatSLAM on Humanoids - A Bio-Inspired SLAM Model Adapted to a Humanoid Robot
Cornelius Weber, Stefan Wermter |
ICANN | 2 |
| 2014 | Human Action Recognition with Hierarchical Growing Neural Gas Learning
German Ignacio Parisi, Cornelius Weber, Stefan Wermter |
ICANN | 2 |
| 2013 | Embodied Language Understanding with a Multiple Timescale Recurrent Neural Network
Stefan Heinrich, Cornelius Weber, Stefan Wermter |
ICANN | 2 |
| 2012 | Adaboost and Hopfield Neural Networks on different image representations for robust face detectionabstractFace detection is an active research area comprising the fields of computer vision, machine learning and intelligent robotics. However, this area is still challenging due to many problems arising from image processing and the further steps necessary for the detection process. In this work we focus on Hopfield Neural Network (HNN) and ensemble learning. It extends our recent work by two components: the simultaneous usage of different image representations and combinations as well as variations in the training procedure. Using the HNN within an ensemble achieves high detection rates but shows no increase in false detection rates, as is commonly the case. We present our experimental setup and investigate the robustness of our architecture. Our results indicate, that with the presented methods the face detection system is flexible regarding varying environmental conditions, leading to a higher robustness. Nils Meins, Doreen Jirak, Cornelius Weber, Stefan Wermter |
HIS | 3 |
| 2012 | Adaptive Learning of Linguistic Hierarchy in a Multiple Timescale Recurrent Neural Network
Stefan Heinrich, Cornelius Weber, Stefan Wermter |
ICANN (1) | 2 |
| 2012 | Hybrid Ensembles Using Hopfield Neural Networks and Haar-Like Features for Face Detection
Nils Meins, Stefan Wermter, Cornelius Weber |
ICANN (1) | 3 |
| 2012 | Learning Features and Predictive Transformation Encoding Based on a Horizontal Product Model
Junpei Zhong, Cornelius Weber, Stefan Wermter |
ICANN (1) | 2 |
| 2012 | A SOM-based model for multi-sensory integration in the superior colliculusabstractWe present an algorithm based on the self-organizing map (SOM) which models multi-sensory integration as realized by the superior colliculus (SC). Our algorithm differs from other algorithms for multi-sensory integration in that it learns mappings between modalities' coordinate systems, it learns their respective reliabilities for different points in space, and uses mappings and reliabilities to perform cue integration. It does this in only one learning phase without supervision and such that calculations and data structures are local to individual neurons. Our simulations indicate that our algorithm can learn near-optimal integration of input from noisy sensory modalities. Johannes Bauer 0002, Cornelius Weber, Stefan Wermter |
IJCNN | 2 |
| 2012 | A neural approach for robot navigation based on cognitive map learningabstractThis paper presents a neural network architecture for a robot learning new navigation behavior by observing a human's movement in a room. While indoor robot navigation is challenging due to the high complexity of real environments and the possible dynamic changes in a room, a human can explore a room easily without any collisions. We therefore propose a neural network that builds up a memory for spatial representations and path planning using a person's movements as observed from a ceiling-mounted camera. Based on the human's motion, the robot learns a map that is used for path planning and motor-action codings. We evaluate our model with a detailed case study and show that the robot navigates effectively. Cornelius Weber, Stefan Wermter |
IJCNN | 2 |
| 2011 | Person Tracking Based on a Hybrid Neural Probabilistic Model
Cornelius Weber, Stefan Wermter |
ICANN (2) | 2 |
| 2011 | Robot Trajectory Prediction and Recognition Based on a Computational Mirror Neurons Model
Junpei Zhong, Cornelius Weber, Stefan Wermter |
ICANN (2) | 2 |
| 2011 | Learning the Optimal Control of Coordinated Eye and Head MovementsabstractVarious optimality principles have been proposed to explain the characteristics of coordinated eye and head movements during visual orienting behavior. At the same time, researchers have suggested several neural models to underly the generation of saccades, but these do not include online learning as a mechanism of optimization. Here, we suggest an open-loop neural controller with a local adaptation mechanism that minimizes a proposed cost function. Simulations show that the characteristics of coordinated eye and head movements generated by this model match the experimental data in many aspects, including the relationship between amplitude, duration and peak velocity in head-restrained and the relative contribution of eye and head to the total gaze shift in head-free conditions. Our model is a first step towards bringing together an optimality principle and an incremental local learning mechanism into a unified control scheme for coordinated eye and head movements. Sohrab Saeb, Cornelius Weber, Jochen Triesch |
PLoS Comput. Biol. | 2 |
| 2009 | A neural model for the adaptive control of saccadic eye movementsabstractSeveral studies have suggested different cost functions to explain the kinematic characteristics of saccades. However, these studies do not present any neural implementation of the optimization procedure they use. Instead, they are based on optimal control theory approaches that provide a global analytical solution rather than a local adaptation scheme. In this study, we propose a model comprised of an open-loop neural controller and an adaptation unit. The neural controller receives the initial target position as input. The adaptation unit, which is the neural interpretation of a simple cost function, evaluates the optimality of this controller and induces weight changes in the controller via a local learning rule. Realistic saccades are obtained with the proposed model. We speculate that the superior colliculus and the cerebellum behave quite similar to our model's neural controller and adaptation unit. Sohrab Saeb, Cornelius Weber, Jochen Triesch |
IJCNN | 2 |
| 2009 | Goal-directed feature learningabstractOnly a subset of available sensory information is useful for decision making. Classical models of the brain's sensory system, such as generative models, consider all elements of the sensory stimuli. However, only the action-relevant components of stimuli need to reach the motor control and decision making structures in the brain. To learn these action-relevant stimuli, the part of the sensory system that feeds into a motor control circuit needs some kind of relevance feedback. We propose a simple network model consisting of a feature learning (sensory) layer that feeds into a reinforcement learning (action) layer. Feedback is established by the reinforcement learner's temporal difference (delta) term modulating an otherwise Hebbian-like learning rule of the feature learner. Under this influence, the feature learning network only learns the relevant features of the stimuli, i.e. those features on which goal-directed actions are to be based. With the input preprocessed in this manner, the reinforcement learner performs well in delayed reward tasks. The learning rule approximates an energy function's gradient descent. The model presents a link between reinforcement learning and unsupervised learning and may help to explain how the basal ganglia receive selective cortical input. Cornelius Weber, Jochen Triesch |
IJCNN | 1 |
| 2009 | Goal-directed learning of features and forward models
Sohrab Saeb, Cornelius Weber, Jochen Triesch |
Neural Networks | 2 |
| 2008 | From Exploration to Planning
Cornelius Weber, Jochen Triesch |
ICANN (1) | 1 |
| 2008 | A Sparse Generative Model of V1 Simple Cells with Intrinsic PlasticityabstractCurrent models for learning feature detectors work on two timescales: on a fast timescale, the internal neurons' activations adapt to the current stimulus; on a slow timescale, the weights adapt to the statistics of the set of stimuli. Here we explore the adaptation of a neuron's intrinsic excitability, termed intrinsic plasticity, which occurs on a separate timescale. Here, a neuron maintains homeostasis of an exponentially distributed firing rate in a dynamic environment. We exploit this in the context of a generative model to impose sparse coding. With natural image input, localized edge detectors emerge as models of V1 simple cells. An intermediate timescale for the intrinsic plasticity parameters allows modeling aftereffects. In the tilt aftereffect, after a viewer adapts to a grid of a certain orientation, grids of a nearby orientation will be perceived as tilted away from the adapted orientation. Our results show that adapting the neurons' gain-parameter but not the threshold-parameter accounts for this effect. It occurs because neurons coding for the adapting stimulus attenuate their gain, while others increase it. Despite its simplicity and low maintenance, the intrinsic plasticity model accounts for more experimental details than previous models without this mechanism. Cornelius Weber, Jochen Triesch |
Neural Comput. | 1 |
| 2007 | A self-organizing map of sigma-pi units
Cornelius Weber, Stefan Wermter |
Neurocomputing | 1 |
| 2006 | Robot docking based on omnidirectional vision and reinforcement learning
David Muse, Cornelius Weber, Stefan Wermter |
Knowl. Based Syst. | 2 |
| 2006 | A camera-direction dependent visual-motor coordinate transformation for a visually guided neural robot
Cornelius Weber, David Muse, Mark Elshaw, Stefan Wermter |
Knowl. Based Syst. | 1 |
| 2006 | A hybrid generative and predictive model of the motor cortex
Cornelius Weber, Stefan Wermter, Mark Elshaw |
Neural Networks | 1 |
| 2005 | Reinforcement Learning in MirrorBot
Cornelius Weber, David Muse, Mark Elshaw, Stefan Wermter |
ICANN (1) | 1 |
| 2005 | Image Segmentation by Complex-Valued Units
Cornelius Weber, Stefan Wermter |
ICANN (1) | 1 |
| 2004 | An associator network approach to robot learning by imitation through vision, motor control and languageabstractImitation learning offers a valuable approach for developing intelligent robot behaviour. We present an imitation approach based on an associator neural network inspired by brain modularity and mirror neurons. The model combines multimodal input based on higher-level vision, motor control and language so that a simulated student robot is able to learn from observing three behaviours which are performed by a teacher robot. The student robot associates these inputs to recognise the behaviour being performed or to perform behaviours by language instruction. With behaviour representations segregating into regions it models aspects of the mirror neuron system as similar patterns of neural activation are involved in recognition and performance. Mark Elshaw, Cornelius Weber, Alexandros Zochios, Stefan Wermter |
IJCNN | 2 |
| 2004 | Robot docking with neural vision and reinforcement
Cornelius Weber, Stefan Wermter, Alexandros Zochios |
Knowl. Based Syst. | 1 |
| 2003 | Learning Localisation Based on Landmarks Using Self-Organisation
Kaustubh Chokshi, Stefan Wermter, Cornelius Weber |
ICANN | 3 |
| 2003 | Object Localisation Using Laterally Connected "What" and "Where" Associator Networks
Cornelius Weber, Stefan Wermter |
ICANN | 1 |
| 2001 | Self-Organization of Orientation Maps, Lateral Connections, and Dynamic Receptive Fields in the Primary Visual Cortex
Cornelius Weber |
ICANN | 1 |
| 2000 | Structured Models from Structured Data: Emergence of Modular Information Processing within One Sheet of NeuronsabstractWe investigate how structured information processing within a neural net can emerge as a result of unsupervised learning from data. Our model consists of input neurons and hidden neurons which are recurrently connected and which represent the thalamus and the cortex, respectively. On the basis of a maximum likelihood framework the task is to generate given input data using the code of the hidden units. Hidden neurons are fully connected allowing for different roles to play within the unfolding time-dynamics of this data generation process. One parameter which is related to the sparsity of neuronal activation varies across the hidden neurons. As a result of training the net captures the structure of the data generation process. The results imply that the division of the cortex into laterally and hierarchically organized areas can evolve to a certain degree as an adaptation to the environment. Cornelius Weber, Klaus Obermayer |
IJCNN (4) | 1 |