EDBT 2026 Demo / reviewers in the wild / expert
Gerhard Rigoll
dblp:78/2835
· DBLP profile ↗
306ranked-venue papers
26as first author
24since 2021 · last 2026
0000-0003-1096-1596ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 230 · 22 first-author · 16 since 2021Artificial intelligence and machine learning · 126 · 12 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 30 · 2 since 2021Databases, data management, data science and information retrieval · 17 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 1 since 2021Systems, architecture and hardware · 4 · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bridging Infrastructures and Vehicles: A Cooperative Framework for Fusing Heterogeneous Future Trajectory PredictionabstractAutonomous vehicle trajectory prediction faces significant challenges in complex traffic scenarios due to sensor occlusions and limited single-viewpoint perception. Cooperative prediction among connected vehicles and infrastructure offers a promising solution. Existing approaches suffer from rigid architectural coupling, requiring homogeneous models and a fixed number of collaborators, making them impractical for dynamic real-world scenarios. Moreover, a critical yet underexplored domain shift problem arises when models trained on single-viewpoint datasets are applied to collaborative environments, leading to substantial performance degradation. To address these limitations, we propose the Bridging Infrastructure and Vehicles (BIV) cooperative prediction framework, a plug-and-play solution that enables collaboration among an arbitrary number of heterogeneous trajectory prediction systems. The BIV framework operates through three stages: individual prediction by each collaborator using their domain-specific models, cooperative sharing of predicted trajectories, and the fusion of predictions using a novel Dynamic Time Warping-based Non-Maximum Suppression (DTW-NMS) mechanism. The DTW-NMS abstracts potential future trajectories of target agents by leveraging the similarity of coordinate sequences to capture safety-critical maneuvers. The proposed framework effectively overcomes the domain shift between single-viewpoint training and collaborative deployment. As an offline collaborative approach, the BIV framework demonstrates superior performance compared to all existing offline methods, while also outperforming most online methods. The BIV framework surpasses the best offline collaborative trajectory prediction models, reducing minADE and minFDE by up to 23.15% and 25.67%, respectively, while lowering the miss rate by 25.93%. Huilin Yin, Yangwenhui Xu, Hao Zhang 0008, Gerhard Rigoll |
IEEE Internet Things J. | 4 |
| 2026 | Knowledge-Informed Multi-Agent Trajectory Prediction at Signalized Intersections for Infrastructure-to-EverythingabstractMulti-agent trajectory prediction at signalized intersections is pivotal for the safety of autonomous driving and the efficiency of intelligent transportation systems. However, conventional vehicle-centric approaches are limited by restricted perception ranges and occlusion. Vehicle-to-Everything (V2X) cooperation is widely regarded as an effective approach to alleviating these limitations. Furthermore, existing cooperative systems often suffer from complex coupling and selection bias, hindering universal and real-time service. To address these challenges, this paper introduces a novel Infrastructure-to-Everything (I2X) collaborative prediction scheme. This scheme decouples infrastructure capabilities from vehicle requests by independently forecasting and broadcasting trajectories for all detected vehicles. Building on this scheme, we propose I2XTraj, a dedicated infrastructure-based model that leverages three core mechanisms. First, a continuous signal-informed mechanism to adaptively encode real-time traffic light information. Second, a maneuver strategy awareness mechanism that integrates intersection geometric constraints to estimate maneuver distributions. Third, a spatial-temporal-mode attention network to refine multi-agent interactions. Extensive evaluations on two real-world datasets, V2X-Seq and SinD, demonstrate the superiority of our approach. In both single-infrastructure and collaborative scenarios, I2XTraj outperforms state-of-the-art methods by over 30% and 15%, respectively, confirming its strong generalizability and robustness in complex intersection environments. Huilin Yin, Yangwenhui Xu, Hao Zhang 0008, Gerhard Rigoll |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2025 | Optimizing Robot Programming: Mixed Reality Gripper ControlabstractConventional robot programming methods are complex and time-consuming for users. In recent years, alternative approaches such as mixed reality have been explored to address these challenges and optimize robot programming. While the findings of the mixed reality robot programming methods are convincing, most existing methods rely on gesture interaction for robot programming. Since controller-based interactions have proven to be more reliable, this paper examines three controller-based programming methods within a mixed reality scenario: 1) Classical Jogging, where the user positions the robot's end effector using the controller's thumbsticks, 2) Direct Control, where the controller's position and orientation directly corresponds to the end effector's, and 3) Gripper Control, where the controller is enhanced with a 3D-printed gripper attachment to grasp and release objects. A within-subjects study$(n = 30)$was conducted to compare these methods. The findings indicate that the Gripper Control condition outperforms the others in terms of task completion time, user experience, mental demand, and task performance, while also being the preferred method. Therefore, it demonstrates promising potential as an effective and efficient approach for future robot programming. Video available at https://youtu.be/83kWr8zUFIQ. Maximilian Rettinger, Leander Hacker, Philipp Wolters, Gerhard Rigoll |
ICRA | 4 |
| 2025 | Unleashing HyDRa: Hybrid Fusion, Depth Consistency and Radar for Unified 3D PerceptionabstractLow-cost, vision-centric 3D perception systems for autonomous driving have made significant progress in recent years, narrowing the gap to expensive LiDAR-based methods. The primary challenge in becoming a fully reliable alternative lies in robust depth prediction capabilities, as camera-based systems struggle with long detection ranges and adverse lighting and weather conditions. In this work, we introduce HyDRa, a novel camera-radar fusion architecture for diverse 3D perception tasks. Building upon the principles of dense Bird's-EyeView (BEV)-based architectures, HyDRa introduces a hybrid fusion approach to combine the strengths of complementary camera and radar features in two distinct representation spaces. Our Height Association Transformer module leverages radar features already in the perspective view to produce more robust and accurate depth predictions. In the BEV, we refine the initial sparse representation by a Radar-weighted Depth Consistency. HyDRa achieves a new state-of-the-art for cameraradar fusion of 64.2 NDS (+1.8) and 58.4 AMOTA (+1.5) on the public nuScenes dataset. Moreover, our new semantically rich and spatially accurate BEV features can be directly converted into a powerful occupancy representation, beating all previous camera-based methods on the Occ3D benchmark by an impressive 3.7 mIoU. Code and models are available at https://github.com/phi-wol/hydra. Philipp Wolters, Johannes Gilg, Torben Teepe, Fabian Herzog, Anouar Laouichi, Martin Hofmann 0011, Gerhard Rigoll |
ICRA | 7 |
| 2024 | Do We Still Need Non-Maximum Suppression? Accurate Confidence Estimates and Implicit Duplication Modeling with IoU-Aware CalibrationabstractObject detectors are at the heart of many semi- and fully autonomous decision systems and are poised to become even more indispensable. They are, however, still lacking in accessibility and can sometimes produce unreliable predictions. Especially concerning in this regard are the—essentially hand-crafted—non-maximum suppression algorithms that lead to an obfuscated prediction process and biased confidence estimates. We show that we can eliminate classic NMS-style post-processing by using IoU-aware calibration. IoU-aware calibration is a conditional Beta calibration; this makes it parallelizable with no hyper-parameters. Instead of arbitrary cutoffs or discounts, it implicitly accounts for the likelihood of each detection being a duplicate and adjusts the confidence score accordingly, resulting in empirically based precision estimates for each detection. Our extensive experiments on diverse detection architectures show that the proposed IoU-aware calibration can successfully model duplicate detections and improve calibration. Compared to the standard sequential NMS and calibration approach, our joint modeling can deliver performance gains over the best NMS-based alternative while producing consistently better-calibrated confidence predictions with less complexity. The code for all our experiments is publicly available1. Johannes Gilg, Torben Teepe, Fabian Herzog, Philipp Wolters, Gerhard Rigoll |
WACV | 5 |
| 2024 | CSANet: Cuboid-Wise Shape Augmentation 3D Object Detector for Occluded TargetsabstractFor 3D environmental perception tasks, light detection and ranging (LiDAR) assumes a crucial role by supplying extensive data in 3D space. In the context of deep learning-based LiDAR-only 3D object detectors, practical challenges like occlusion and signal missing lead to the loss of partial shape information, thereby causing a decline in detection accuracy. To tackle this issue, we propose a two-stage LiDAR 3D object detector that includes the cuboid-wise shape augmentation network to supplement instance-level foreground points with geometric structure information. Dense point clouds are achieved by introducing gridding procedure and 3D encoder-decoder. Besides, we further capture contextual information and reweight grid points within each proposal. The original region of interest (RoI) features are aggregated with the augmented features in refinement stage to recover intact shape details. The performance of our proposed method on KITTI dataset demonstrates its effectiveness in 3D object detection. Wancheng Ge, Gerhard Rigoll, Huilin Yin |
IEEE Signal Process. Lett. | 3 |
| 2023 | Octuplet Loss: Make Face Recognition Robust to Image ResolutionabstractImage resolution, or in general, image quality, plays an essential role in the performance of today's face recognition systems. To address this problem, we propose a novel combination of the popular triplet loss to improve robustness against image resolution via fine-tuning of existing face recognition models. With octuplet loss, we leverage the relationship between high-resolution images and their synthetically down-sampled variants jointly with their identity labels. Fine-tuning several state-of-the-art approaches with our method proves that we can significantly boost performance for cross-resolution (high-to-low resolution) face verification on various datasets without meaningfully exacerbating the performance on high-to-high resolution images. Our method applied on the FaceTransformer network achieves 95.12% face verification accuracy on the challenging XQLFW dataset while reaching 99.73% on the LFW database. Moreover, the low-to-low face verification accuracy benefits from our method. We release our code11Code available on https://github.com/Martlgap/octuplet-loss to allow seamless integration of the octuplet loss into existing frameworks. Martin Knoche, Mohamed R. Elkadeem, Stefan Hörmann 0001, Gerhard Rigoll |
FG | 4 |
| 2023 | Introducing A Framework for Single-Human Tracking Using Event-Based CamerasabstractEvent cameras generate data based on the amount of motion present in the captured scene, making them attractive sensors for solving object tracking tasks. In this paper, we present a framework for tracking humans using a single event camera which consists of three components. First, we train a Graph Neural Network (GNN) to recognize a person within the stream of events. Batches of events are represented as spatio-temporal graphs in order to preserve the sparse nature of events and retain their high temporal resolution. Subsequently, the person is localized in a weakly-supervised manner by adopting the well established method of Class Activation Maps (CAM) for our graph-based classification model. Our approach does not require the ground truth position of humans during training. Finally, a Kalman filter is deployed for tracking, which uses the predicted bounding box surrounding the human as measurement. We demonstrate that our approach achieves robust tracking results on test sequences from the Gait3 database, paving the way for further privacy-preserving methods in event-based human tracking. Code, pre-trained models and datasets of our research are publicly available1. Dominik Eisl, Fabian Herzog, Jean-Luc Dugelay, Ludovic Apvrille, Gerhard Rigoll |
ICIP | 5 |
| 2023 | The Box Size Confidence Bias Harms Your Object DetectorabstractCountless applications depend on accurate predictions with reliable confidence estimates from modern object detectors. However, it is well known that neural networks, including object detectors, produce miscalibrated confidence estimates. Recent work even suggests that detectors’ confidence predictions are biased with respect to object size and position. In object detection, the issues of conditional biases, confidence calibration, and task performance are usually explored in isolation, but, as we aim to show, they are closely related. We formally prove that the conditional confidence bias harms the performance of object detectors and empirically validate these findings. Specifically, to quantify the performance impact of the confidence bias on object detectors, we modify the histogram binning calibration to avoid performance impairment and instead improve it through calibration conditioned on the bounding box size. We further find that the confidence bias is also present in detections generated on the training data of the detector, which can be leveraged to perform the de-biasing. Moreover, we show that Test Time Augmentation (TTA) confounds this bias, which results in even more significant performance impairments on the detectors. Finally, we use our proposed algorithm to analyze a diverse set of object detection architectures and show that the conditional confidence bias harms their performance by up to 0.6 mAP and 0.8 mAP50. Code available at https://github.com/Blueblue4/Object-Detection-Confidence-Bias. Johannes Gilg, Torben Teepe, Fabian Herzog, Gerhard Rigoll |
WACV | 4 |
| 2023 | Wavelet regularization benefits adversarial training
Jun Yan 0009, Huilin Yin, Ziming Zhao 0006, Wancheng Ge, Hao Zhang 0008, Gerhard Rigoll |
Inf. Sci. | 6 |
| 2022 | Defuse the Training of Risky Tasks: Collaborative Training in XRabstractExtensive training is crucial but challenging in certain areas such as explosive ordnance disposal. Past conflicts have shown that not only military personal but also civilians have to learn how to disarm unexploded ordnance. The preparation for dangerous situations is difficult and limited in the real world. Extended reality (XR) offers new possibilities to enhance the training of explosive ordnance disposal experts due to its immersive capabilities. This paper presents a comparative study (n = 75) of three distinct training methods: 1) Real-world (Real)-Training with a tangible replica object, 2) Virtual Reality (VR)-Training with a non-see-through Head-Mounted-Display (HMD), and 3) Mixed Reality (MR)-Training in a Cave Automatic Virtual Environment (CAVE). All training methods are collaborative, i.e., an instructor teaches the training content to a participant in the real or virtual world. We evaluate the suitability of these approaches in terms of usability, cognitive workload, training motivation, and training success. Our results indicate that the virtual methods, VR-Training and MR-Training, provide significantly superior results in the evaluated aspects compared to the real-world training. These results can also be applied to other collaborative training methods, as the training concept of this use case was non-specific. Therefore, these virtual technologies can increase the safety of explosive ordnance disposal personnel, and we recommend establishing this in future training. Maximilian Rettinger, Gerhard Rigoll |
ISMAR | 2 |
| 2022 | Efficient Active Learning Strategies for Monocular 3D Object DetectionabstractProcessing camera information to perceive their 3D surrounding is essential for building scalable autonomous driving vehicles. For this task, deep learning networks provide effective real-time solutions. However, to compensate for missing depth information in cameras compared to LiDARs, a large amount of labeled data is required for training. Active learning is a training framework where the network actively participates in the data selection process to improve data efficiency and performance. In this work, we propose an active learning pipeline for 3D object detection from monocular images. The main components of our approach are (1) two training-efficient uncertainty estimation strategies, (2) a diversity-based selection strategy to select images that contain the most diverse set of objects, (3) a novel active learning strategy more suitable for training autonomous driving perception networks. Experiments show that combining our proposed uncertainty estimation methods provides a better data saving rate and reaches a higher final performance than baselines. Furthermore, we empirically show performance gains of the presented diversity-based selection strategy and the efficiency of the proposed active learning strategy. Aral Hekimoglu, Michael Schmidt 0015, Alvaro Marcos-Ramiro, Gerhard Rigoll |
IV | 4 |
| 2022 | Do You Notice Me? How Bystanders Affect the Cognitive Load in Virtual RealityabstractIn contrast to the real world, users are not able to perceive bystanders in virtual reality (VR). Bystanders may distract users and influence their cognitive load. This involves users to feel discomfort at the thought of unintentionally touching or even bumping into a physical bystander while interacting with the virtual environment. Not knowing the intentions of a bystander or whether one is present can unsettle the user. We investigate how a bystander affects a user’s cognitive load since it has a decisive impact on applications such as VR training. In a between-subjects lab study (N = 42), three conditions were compared: 1) no bystander, 2) an invisible bystander, and 3) a visible bystander (as an avatar). Over a series of iterations, the participants were asked to memorize four pairs of letters, perform a mental rotation task and then recall the pairs of letters. The results of our study demonstrate that a bystander acting as an avatar in the virtual environment increases the user’s cognitive load more than an invisible bystander. Moreover, the cognitive load of a VR user is significantly increased by a bystander. Therefore, our work suggests that either the examiner must be separated from the participant or the examiner’s influence (as a bystander) must be included in the analysis. Maximilian Rettinger, Christoph Schmaderer, Gerhard Rigoll |
VR | 3 |
| 2022 | Dissected 3D CNNs: Temporal skip connections for efficient online video processing
Okan Köpüklü, Stefan Hörmann 0001, Fabian Herzog, Hakan Çevikalp, Gerhard Rigoll |
Comput. Vis. Image Underst. | 5 |
| 2021 | A Coarse-to-Fine Dual Attention Network for Blind Face CompletionabstractIn the area of face completion, the missing information within an occluded area is estimated, yielding a realistic face of the same identity. In most previous works, the mask describing the occluded region is known, limiting the scope of application. To alleviate this limitation, we propose a coarse-to-fine network trained as a conditional generative adversarial network. While the coarse network predicts the mask and generates a rough estimation of the semantic content, the subsequent fine network refines the rough prediction into a realistic and identity-persevering reconstruction. This is achieved by incorporating adversarial loss and using features from a pretrained face feature extractor. Unlike previous approaches, we employ two parallel attention mechanisms: 1) a patchwise cross-attention module to substitute information within the occluded patches with patches from the non-occluded region; 2) a pixel-wise global self-attention to allow information exchange within the entire feature map. Our exhaustive analysis, including reconstruction quality and face recognition metrics, shows that our approach outperforms the state of the art in blind face completion, improving the true positive identification rate at rank 1 on the MegaFace benchmark from 36.55 % to 42.48 %. This represents a substantial step towards closing the gap between occluded (29.34 %) and non-occluded faces (52.32 %). In terms of reconstruction quality, we obtain a structural similarity of 0.9639 compared to 0.8526 and 0.9563 for occluded faces and the state of the art, respectively. In addition to previous approaches, we provide an in-depth analysis of the influence of the position, size, and sparsity of the occlusion and use facial landmark prediction to measure reconstruction quality. Stefan Hörmann 0001, Zhibing Xia, Martin Knoche, Gerhard Rigoll |
FG | 4 |
| 2021 | Cross-Quality LFW: A Database for Analyzing Cross- Resolution Image Face Recognition in Unconstrained EnvironmentsabstractReal-world face recognition applications often deal with suboptimal image quality or resolution due to different capturing conditions such as various subject-to-camera distances, poor camera settings, or motion blur. This characteristic has an unignorable effect on performance. Recent cross-resolution face recognition approaches used simple, arbitrary, and unrealistic down- and up-scaling techniques to measure robustness against real-world edge-cases in image quality. Thus, we propose a new standardized benchmark dataset and evaluation protocol derived from the famous Labeled Faces in the Wild (LFW). In contrast to previous derivatives, which focus on pose, age, similarity, and adversarial attacks, our Cross-Quality Labeled Faces in the Wild (XQLFW) maximizes the quality difference. It contains only more realistic synthetically degraded images when necessary. Our proposed dataset is then used to further investigate the influence of image quality on several state-of-the-art approaches. With XQLFW, we show that these models perform differently in cross-quality cases, and hence, the generalizing capability is not accurately predicted by their performance on LFW. Additionally, we report baseline accuracy with recent deep learning models explicitly trained for cross-resolution applications and evaluate the susceptibility to image quality. To encourage further research in cross-resolution face recognition and incite the assessment of image quality robustness, we publish the database and code for evaluation.11Code, dataset and evaluation protocol available on https://martlgap.github.io/xqlfw Martin Knoche, Stefan Hörmann 0001, Gerhard Rigoll |
FG | 3 |
| 2021 | How to Design a Three-Stage Architecture for Audio-Visual Active Speaker Detection in the WildabstractSuccessful active speaker detection requires a three-stage pipeline: (i) audio-visual encoding for all speakers in the clip, (ii) inter-speaker relation modeling between a reference speaker and the background speakers within each frame, and (iii) temporal modeling for the reference speaker. Each stage of this pipeline plays an important role for the final performance of the created architecture. Based on a series of controlled experiments, this work presents several practical guidelines for audio-visual active speaker detection. Correspondingly, we present a new architecture called ASDNet, which achieves a new state-of-the-art on the AVA-ActiveSpeaker dataset with a mAP of 93.5% outperforming the second best with a large margin of 4.7%. Our code and pretrained models are publicly available1. Okan Köpüklü, Maja Taseska, Gerhard Rigoll |
ICCV | 3 |
| 2021 | Lightweight Multi-Branch Network For Person Re-IdentificationabstractPerson Re-Identification aims to retrieve person identities from images captured by multiple cameras or the same cameras in different time instances and locations. Because of its importance in many vision applications from surveillance to human-machine interaction, person re-identification methods need to be reliable and fast. While more and more deep architectures are proposed for increasing performance, those methods also increase overall model complexity. This paper proposes a lightweight network that combines global, part-based, and channel features in a unified multi-branch architecture that builds on the resource-efficient OSNet backbone. Using a well-founded combination of training techniques and design choices, our final model achieves state-of-the-art results on CUHK03 labeled, CUHK03 detected, and Market-1501 with 85.1% mAP/ 87.2% rankl, 82.4% mAP/84.9% rankl, and 91.5% mAP/96.3% rankl, respectively. Fabian Herzog, Xunbo Ji, Torben Teepe, Stefan Hörmann 0001, Johannes Gilg, Gerhard Rigoll |
ICIP | 6 |
| 2021 | Face Texture Generation And Identity-Preserving RectificationabstractTextures are a vital asset in conveying a realistic impression of a 3D scene to the viewers. In order to obtain high-quality textures, real-life objects are scanned or designers create handcrafted textures. Both tasks involve manual work, are quite time-consuming, and therefore fail when a large quantity of textures is required. Thus, we propose to use a Generative Adversarial Network to generate an artificial texture. As textures need to be perfectly aligned with the 2D projection of the 3D model, our method involves a texture rectification technique, ensuring that the generated textures wrap well onto the 3D model. On the example of face textures, we illustrate that our method generates textures of high quality and variance. Moreover, we show that the rectification process preserves the facial appearance and identity, indicating that we successfully disentangle features responsible for facial appearance and the texture’s fit. Stefan Hörmann 0001, Arka Bhowmick, Michael Weiher, Karl Leiss, Gerhard Rigoll |
ICIP | 5 |
| 2021 | Face Aggregation Network For Video Face RecognitionabstractTypical approaches for video face recognition aggregate faces in a feature space to obtain a single feature representing the entire video. Unlike most previous approaches, we aggregate the faces directly in order to additionally obtain a single representative face as an intermediate output, from which a more discriminative feature vector is extracted. To overcome the limitation of a fixed number of input images of the state of the art in face aggregation, we incorporate a permutation invariant U-Net architecture capable of processing an arbitrary number of frames, which is employed in a generative adversarial network. We demonstrate the effectiveness of our method on three popular benchmark datasets for video face recognition. Our approach outperforms the baselines on the YouTube Faces dataset, obtaining an accuracy of 96.62%. Besides, we show that our method is robust against motion blur. Stefan Hörmann 0001, Zhenxiang Cao, Martin Knoche, Fabian Herzog, Gerhard Rigoll |
ICIP | 5 |
| 2021 | Attention-Based Partial Face RecognitionabstractPhotos of faces captured in unconstrained environments, such as large crowds, still constitute challenges for current face recognition approaches as often faces are occluded by objects or people in the foreground. However, few studies have addressed the task of recognizing partial faces. In this paper, we propose a novel approach to partial face recognition capable of recognizing faces with different occluded areas. We achieve this by combining attentional pooling of a ResNet’s intermediate feature maps with a separate aggregation module. We further adapt common losses to partial faces in order to ensure that the attention maps are diverse and handle occluded parts. Our thorough analysis demonstrates that we outperform all baselines under multiple benchmark protocols, including naturally and synthetically occluded partial faces. This suggests that our method successfully focuses on the relevant parts of the occluded face. Stefan Hörmann 0001, Martin Knoche, Torben Teepe, Gerhard Rigoll |
ICIP | 5 |
| 2021 | Gaitgraph: Graph Convolutional Network for Skeleton-Based Gait RecognitionabstractGait recognition is a promising video-based biometric for identifying individual walking patterns from a long distance. At present, most gait recognition methods use silhouette images to represent a person in each frame. However, silhouette images can lose fine-grained spatial information, and most papers do not regard how to obtain these silhouettes in complex scenes. Furthermore, silhouette images contain not only gait features but also other visual clues that can be recognized. Hence these approaches can not be considered as strict gait recognition. We leverage recent advances in human pose estimation to estimate robust skeleton poses directly from RGB images to bring back model-based gait recognition with a cleaner representation of gait. Thus, we propose GaitGraph that combines skeleton poses with Graph Convolutional Network (GCN) to obtain a modern model-based approach for gait recognition. The main advantages are a cleaner, more elegant extraction of the gait features and the ability to incorporate powerful spatiotemporal modeling using GCN. Experiments on the popular CASIA-B gait dataset show that our method archives state-of-the-art performance in model-based gait recognition.The code and models are publicly available1 Torben Teepe, Johannes Gilg, Fabian Herzog, Stefan Hörmann 0001, Gerhard Rigoll |
ICIP | 6 |
| 2021 | A Global Discriminant Joint Training Framework for Robust Speech RecognitionabstractRobustness in adverse acoustic conditions is critical for practical human-machine interaction. A common solution for this problem is adding an independent speech enhancement front-end. Nonetheless, due to being trained separately from the automatic speech recognition (ASR) module, the independent enhancement front-end falls into the sub-optimum easily. Besides, the handcrafted loss function of the enhancement module tends to introduce unseen distortions, which even degrade the ASR performance. To address this concern, a promising idea of the joint training is progressively drawing more interests. Nevertheless, none of the previously proposed joint-training frameworks is built on the increasingly popular self-attention mechanism or generative adversarial architecture. This paper proposes a novel joint-training framework, concatenating a speech enhancement generative adversarial network as the front-end and a self-attention based ASR module as the back-end to be jointly trained as an extensive network, to boost the noise robustness of the end-to-end ASR system. A Sinc convolution layer is usefully merged into the speech enhancement front-end for more representative features extraction. Moreover, a discriminant component plays the role of the local guide of the enhancement module and the global guide in the joint training simultaneously, which guides the enhancement front-end to output more desirable features for the subsequent ASR module and thereby offsets the limitation of the separate training and handcrafted loss functions.Systematic experiments reveal that the proposed framework significantly overtakes other competitive solutions, especially in challenging environments. Lujun Li 0002, Ludwig Kurzinger, Tobias Watzel, Gerhard Rigoll |
ICTAI | 4 |
| 2021 | Driver Anomaly Detection: A Dataset and Contrastive Learning ApproachabstractDistracted drivers are more likely to fail to anticipate hazards, which result in car accidents. Therefore, detecting anomalies in drivers' actions (i.e., any action deviating from normal driving) contains the utmost importance to reduce driver-related accidents. However, there are unbounded many anomalous actions that a driver can do while driving, which leads to an `open set recognition' problem. Accordingly, instead of recognizing a set of anomalous actions that are commonly defined by previous dataset providers, in this work, we propose a contrastive learning approach to learn a metric to differentiate normal driving from anomalous driving. For this task, we introduce a new video-based benchmark, the Driver Anomaly Detection (DAD) dataset, which contains normal driving videos together with a set of anomalous actions in its training set. In the test set of the DAD dataset, there are unseen anomalous actions that still need to be winnowed out from normal driving. Our method reaches 0.9673 AUC on the test set, demonstrating the effectiveness of the contrastive learning approach on the anomaly detection task. Our dataset, codes and pre-trained models are publicly available1. Okan Köpüklü, Jiapeng Zheng, Gerhard Rigoll |
WACV | 4 |
| 2020 | A Multi-Task Comparator Framework for Kinship VerificationabstractApproaches for kinship verification often rely on cosine distances between face identification features. However, due to gender bias inherent in these features, it is hard to reliably predict whether two opposite-gender pairs are related. Instead of fine tuning the feature extractor network on kinship verification, we propose a comparator network to cope with this bias. After concatenating both features, cascaded local expert networks extract the information most relevant for their corresponding kinship relation. We demonstrate that our framework is robust against this gender bias and achieves comparable results on two tracks of the RFIW Challenge 2020. Moreover, we show how our framework can be further extended to handle partially known or unknown kinship relations. Stefan Hörmann 0001, Martin Knoche, Gerhard Rigoll |
FG | 3 |
| 2020 | Attention Fusion for Audio-Visual Person Verification Using Multi-Scale FeaturesabstractIn the domain of audio-visual person recognition, many approaches use naive fusion techniques, such as scorelevel fusion or concatenation, to fuse the features obtained by face and audio extraction networks. More sophisticated methods fuse both features taking into account the quality of their corresponding inputs. In this paper, we propose a novel architecture to improve the prediction of feature quality. In contrary to previous works, which estimate feature quality based on the features themselves, we combine the information obtained from different layers of the feature extraction networks. In our analysis, we show that our approach outperforms state-of-the-art fusion approaches on well-established benchmarks for multimodal person verification. Moreover, we show that our model is robust against degradation of the visual input. Stefan Hörmann 0001, Abdul Moiz, Martin Knoche, Gerhard Rigoll |
FG | 4 |
| 2020 | DriverMHG: A Multi-Modal Dataset for Dynamic Recognition of Driver Micro Hand Gestures and a Real-Time Recognition FrameworkabstractThe use of hand gestures provides a natural alternative to cumbersome interface devices for Human-Computer Interaction (HCI) systems. However, real-time recognition of dynamic micro hand gestures from video streams is challenging for in-vehicle scenarios since (i) the gestures should be performed naturally without distracting the driver, (ii) micro hand gestures occur within very short time intervals at spatially constrained areas, (iii) the performed gesture should be recognized only once, and (iv) the entire architecture should be designed lightweight as it will be deployed to an embedded system. In this work, we propose an HCI system for dynamic recognition of driver micro hand gestures, which can have a crucial impact in automotive sector especially for safety related issues. For this purpose, we initially collected a dataset named Driver Micro Hand Gestures (DriverMHG), which consists of RGB, depth and infrared modalities. The challenges for dynamic recognition of micro hand gestures have been addressed by proposing a lightweight convolutional neural network (CNN) based architecture which operates online efficiently with a sliding window approach. For the CNN model, several 3-dimensional resource efficient networks are applied and their performances are analyzed. Online recognition of gestures has been performed with 3D-MobileNetV2, which provided the best offline accuracy among the applied networks with similar computational complexities. The final architecture is deployed on a driver simulator operating in real-time. We make DriverMHG dataset and our source code publicly available1. Okan Köpüklü, Thomas Ledwon, Yao Rong 0001, Neslihan Kose, Gerhard Rigoll |
FG | 5 |
| 2020 | Small-Footprint Keyword Spotting on Raw Audio Data with Sinc-ConvolutionsabstractKeyword Spotting (KWS) enables speech-based user interaction on smart devices. Always-on and battery-powered application scenarios for smart devices put constraints on hardware resources and power consumption, while also demanding high accuracy as well as real-time capability. Previous architectures first extracted acoustic features and then applied a neural network to classify keyword probabilities, optimizing towards memory footprint and execution time.Compared to previous publications, we took additional steps to reduce power and memory consumption without reducing classification accuracy. Power-consuming audio preprocessing and data transfer steps are eliminated by directly classifying from raw audio. For this, our end-to-end architecture extracts spectral features using parametrized Sinc-convolutions. Its memory footprint is further reduced by grouping depthwise separable convolutions. Our network achieves the competitive accuracy of 96.4% on Google's Speech Commands test set with only 62k parameters. Simon Mittermaier, Ludwig Kurzinger, Bernd Waschneck, Gerhard Rigoll |
ICASSP | 4 |
| 2020 | Lightweight End-to-End Speech Recognition from Raw Audio Data Using Sinc-ConvolutionsabstractMany end-to-end Automatic Speech Recognition (ASR) systems still rely on pre-processed frequency-domain features that are handcrafted to emulate the human hearing. Our work is motivated by recent advances in integrated learnable feature extraction. For this, we propose Lightweight Sinc-Convolutions (LSC) that integrate Sinc-convolutions with depthwise convolutions as a low-parameter machine-learnable feature extraction for end-to-end ASR systems. We integrated LSC into the hybrid CTC/attention architecture for evaluation. The resulting end-to-end model shows smooth convergence behaviour that is further improved by applying SpecAugment in time-domain. We also discuss filter-level improvements, such as using log-compression as activation function. Our model achieves a word error rate of 10.7% on the TEDlium v2 test dataset, surpassing the corresponding architecture with log-mel filterbank features by an absolute 1.9%, but only has 21% of its model size. Ludwig Kurzinger, Nicolas Lindae, Palle Klewitz, Gerhard Rigoll |
INTERSPEECH | 4 |
| 2019 | Real-time Hand Gesture Detection and Classification Using Convolutional Neural NetworksabstractReal-time recognition of dynamic hand gestures from video streams is a challenging task since (i) there is no indication when a gesture starts and ends in the video, (ii) performed gestures should only be recognized once, and (iii) the entire architecture should be designed considering the memory and power budget. In this work, we address these challenges by proposing a hierarchical structure enabling offline-working convolutional neural network (CNN) architectures to operate online efficiently by using sliding window approach. The proposed architecture consists of two models: (1) A detector which is a lightweight CNN architecture to detect gestures and (2) a classifier which is a deep CNN to classify the detected gestures. In order to evaluate the single-time activations of the detected gestures, we propose to use Levenshtein distance as an evaluation metric since it can measure misclassifications, multiple detections, and missing detections at the same time. We evaluate our architecture on two publicly available datasets - EgoGesture and NVIDIA Dynamic Hand Gesture Datasets - which require temporal detection and classification of the performed hand gestures. ResNeXt-101 model, which is used as a classifier, achieves the state-of-the-art offline classification accuracy of 94.04% and 83.82% for depth modality on EgoGesture and NVIDIA benchmarks, respectively. In real-time detection and classification, we obtain considerable early detections while achieving performances close to offline operation. The codes and pretrained models used in this work are publicly available1. Okan Köpüklü, Ahmet Gunduz, Neslihan Kose, Gerhard Rigoll |
FG | 4 |
| 2019 | Gait Energy Image Restoration Using Generative Adversarial NetworksabstractGait is a biometric property that can be used for human identification in video surveillance. Basically, different gait features require motion of a person walking over one complete gait cycle. For example, in Gait Energy Image (GEI), average of silhouette images over one complete gait cycle is computed. However, in reality, there might be a partial gait cycle data available due to occlusion. In this paper, we propose a Generative Adversarial Network (GAN) in order to address the problem of gait recognition from incomplete gait cycle. Precisely, the network is able to reconstruct complete GEIs from incomplete GEIs. The proposed architecture is composed of (i) a generator which is an auto-encoder network to construct complete GEIs out of incomplete GEIs and (ii) two discriminators, one of which discriminates whether a given image is a full GEI while the other discriminates whether two GEIs belong to the same subject. We evaluate our approach on the OULP large gait dataset confirming that the proposed architecture successfully reconstructs complete GEIs from even extreme incomplete gait cycles. Maryam Babaee, Okan Köpüklü, Stefan Hörmann 0001, Gerhard Rigoll |
ICIP | 5 |
| 2019 | Outlier-Robust Neural Aggregation Network for Video Face IdentificationabstractCurrent approaches for video face recognition rely on image sets containing faces of exclusively one identity. However, as image sets are created by unsupervised methods, it is necessary to consider outlier-afflicted sets for real-life applications. In this paper, we propose an Outlier-Robust Neural Aggregation Network (ORNAN). First, we embed each image into a feature space using a Convolutional Neural Network (CNN). With the help of two cascaded attention blocks, we predict outliers within the image set. By integrating this knowledge into our aggregation network, we adaptively aggregate all feature vectors to form a single feature, mitigating the influence of outliers and noisy features. We show that our network is robust against outliers using outlier-afflicted IJB-B and IJB-C benchmarks while maintaining similar performance without outliers. Stefan Hörmann 0001, Martin Knoche, Maryam Babaee, Okan Köpüklü, Gerhard Rigoll |
ICIP | 5 |
| 2019 | Convolutional Neural Networks with Layer ReuseabstractA convolutional layer in a Convolutional Neural Network (CNN) consists of many filters which apply convolution operation to the input, capture some special patterns and pass the result to the next layer. If the same patterns also occur at the deeper layers of the network, why wouldn't the same convolutional filters be used also in those layers? In this paper, we propose a CNN architecture, Layer Reuse Network (LruNet), where the convolutional layers are used repeatedly without the need of introducing new layers to get a better performance. This approach introduces several advantages: (i) Considerable amount of parameters are saved since we are reusing the layers instead of introducing new layers, (ii) the Memory Access Cost (MAC) can be reduced since reused layer parameters can be fetched only once, (iii) the number of nonlinearities increases with layer reuse, and (iv) reused layers get gradient updates from multiple parts of the network. The proposed approach is evaluated on CIFAR-10, CIFAR-100 and Fashion-MNIST datasets for image classification task, and layer reuse improves the performance by 5.14%, 5.85% and 2.29%, respectively. The source code and pretrained models are publicly available1. Okan Köpüklü, Maryam Babaee, Stefan Hörmann 0001, Gerhard Rigoll |
ICIP | 4 |
| 2019 | Exploring the Use of Augmented Reality Interfaces for Driver Assistance in Short-Notice TakeoversabstractWith conditionally automated vehicles (Level 3), drivers are still required to be ready to intervene upon a takeover request (TOR) and face the difficulty of achieving their optimal performance level directly after a passive phase. In this work, we examine the effects of using an augmented reality (AR) interface with world-registered visualizations to assist drivers in the last moments before a takeover and in the first seconds of controlling the vehicle. We focus on urban situations with an unplanned, short-notice TOR and created three distinct example scenarios in a mixed-reality driving simulation. We present a prototype of an AR assistance system realized on a simulated windshield display (WSD). In a user study, we compare the AR system to a conventional head-down display (HDD) and present results on driving performance, driver workload and usability. The AR assistant enabled higher lateral performance and reduced workload in situations where steering is required directly after takeover but prolonged reaction time when only fast longitudinal input was required. The AR interface performed better than the HDD in most user-centered aspects including comfort of use and helpfulness. Patrick Lindemann, Niklas Müller, Gerhard Rigoll |
IV | 3 |
| 2019 | Acceptance and User Experience of Driving with a See-Through Cockpit in a Narrow-Space Overtaking ScenarioabstractIn this work, we examine the implications of driving with transparent cockpits (TCs) in a narrow-space overtaking scenario. We utilize a virtual environment to simulate two possible manifestations: a user-controlled head-mounted system and a static projection-based system. We conducted a user study with an overtaking task and present results for acceptance and a comparison of both systems regarding user experience. Participants preferred the static TC and evaluated it as the solution with higher pragmatic quality and attractiveness. The TC generally scored highly in hedonic quality and was rated positively regarding perceived safety and ease of use. Patrick Lindemann, Dominik Eisl, Gerhard Rigoll |
VR | 3 |
| 2019 | A Simulation for Examining the Effects of Inaccurate Head Tracking on Drivers of Vehicles with Transparent Cockpit ProjectionsabstractThe transparent cockpit (TC) is a driver-car interface concept in which processed camera images of the surrounding environment are superimposed onto the car interior to give the driver the ability to see through the cockpit. For a perspectively accurate experience from the driver's point of view, head tracking is required. However, a real-world system may be prone to typical tracking errors, especially while driving. In this work, we present a TC simulation using artificial errors of various types to examine how much each type affects driver performance and experience. First results of an initial study show that there is generally no significant deterioration of lateral performance compared to a perfectly accurate TC. Repeated loss of tracking was least noticed by participants. Accuracy (miscalibration) and precision (jitter) errors were noticed the most. Patrick Lindemann, Gerhard Rigoll |
VR | 3 |
| 2019 | Person identification from partial gait cycle using fully convolutional neural networks
Maryam Babaee, Gerhard Rigoll |
Neurocomputing | 3 |
| 2019 | A dual CNN-RNN for multiple people tracking
Maryam Babaee, Zimu Li, Gerhard Rigoll |
Neurocomputing | 3 |
| 2018 | Robust Facial Landmark Detection via a Fully-Convolutional Local-Global Context NetworkabstractWhile fully-convolutional neural networks are very strong at modeling local features, they fail to aggregate global context due to their constrained receptive field. Modern methods typically address the lack of global context by introducing cascades, pooling, or by fitting a statistical model. In this work, we propose a new approach that introduces global context into a fully-convolutional neural network directly. The key concept is an implicit kernel convolution within the network. The kernel convolution blurs the output of a local-context subnet, which is then refined by a global-context subnet using dilated convolutions. The kernel convolution is crucial for the convergence of the network because it smoothens the gradients and reduces overfitting. In a postprocessing step, a simple PCA-based 2D shape model is fitted to the network output in order to filter outliers. Our experiments demonstrate the effectiveness of our approach, outperforming several state-of-the-art methods in facial landmark detection. Daniel Merget, Matthias Rock, Gerhard Rigoll |
CVPR | 3 |
| 2018 | Gait Recognition from Incomplete Gait CycleabstractIn gait recognition, which has been recently regarded as a biometric recognition tool, proposed approaches assume that an individual is observed for at least one gait cycle. However, in reality, there might be available only a few frames of full gait cycle of a subject due to occlusion. Therefore, gait recognition systems would fail in these scenarios. In this paper, we propose a method to tackle this problem by proposing a gait recognition algorithm from an incomplete gait cycle information. We achieve this by 1) creating an incomplete Energy Image (GEI) from a few available silhouettes of a subject and 2) reconstructing the complete GEI from incomplete GEI using a deep auto-encoder. The experimental results on a public gait dataset demonstrate the validity of the proposed method. Maryam Babaee, Gerhard Rigoll |
ICIP | 3 |
| 2018 | Occlusion Handling in Tracking Multiple People Using RNNabstractIn tracking-by-detection of multiple targets in video sequences, ID-switch is an undesirable error due to long (short) occlusion among targets. In this paper, we propose an occlusion handling method based on Recurrent Neural Network (RNN) to remedy this issue. The method reconstructs missed detection boxes in order to preserve the ID number of targets after occlusion by predicting the detections in next frames. The prediction is accomplished by learning the motion of targets using a novel RNN. Applying this technique on tracking results of several state-of-the-arts shows that their ID-switch error is reduced. Maryam Babaee, Zimu Li, Gerhard Rigoll |
ICIP | 3 |
| 2018 | Analysis on Temporal Dimension of Inputs for 3D Convolutional Neural Networksabstract3D ConvNets provide a dedicated spatiotemporal representation in order to incorporate motion patterns within video frames. However, compared to 2D convolutions, the 3D convolution kernels increase the number of parameters in the architecture and the floating point operations during inference time, which are of critical importance for real-time applications requiring faster runtime. In this paper, we show a sparse sampling and stacking strategy to span large time intervals for 3D ConvNet architectures that can attain multiple times less inference time by relinquishing little amount of classification accuracy. The proposed approach is validated on action and gesture recognition tasks using two recent video datasets: Jester and Something-Something datasets. Okan Köpüklü, Gerhard Rigoll |
IPAS | 2 |
| 2018 | A deep convolutional neural network for video sequence background subtraction
Mohammadreza Babaee, Duc Tung Dinh, Gerhard Rigoll |
Pattern Recognit. | 3 |
| 2017 | GazeEverywhere: Enabling Gaze-only User Interaction on an Unmodified Desktop PC in Everyday ScenariosabstractEye tracking is becoming more and more affordable, and thus gaze has the potential to become a viable input modality for human-computer interaction. We present the GazeEverywhere solution that can replace the mouse with gaze control by adding a transparent layer on top of the system GUI. It comprises three parts: i) the SPOCK interaction method that is based on smooth pursuit eye movements and does not suffer from the Midas touch problem; ii) an online recalibration algorithm that continuously improves gaze-tracking accuracy using the SPOCK target projections as reference points; and iii) an optional hardware setup utilizing head-up display technology to project superimposed dynamic stimuli onto the PC screen where a software modification of the system is not feasible. In validation experiments, we show that GazeEverywhere's throughput according to ISO 9241-9 was improved over dwell time based interaction methods and nearly reached trackpad level. Online recalibration reduced interaction target ('button') size by about 25%. Finally, a case study showed that users were able to browse the internet and successfully run Wikirace using gaze only, without any plug-ins or other modifications. Simon Schenk, Marc Dreiser, Gerhard Rigoll, Michael Dorr |
CHI | 3 |
| 2017 | Joint tracking and gait recognition of multiple people in videoabstractWe propose a novel approach to address the problem of jointly tracking and gait recognition of multiple people in a video sequence. The most state of the art algorithms for gait recognition consider the cases where there is only one person without any occlusion in a very constrained environment. However, in real scenarios such as in airports, train stations, etc, there are many people in the environment that make these algorithms inapplicable. Although first tracking of each person and then gait recognition could be a solution, we argue that the multi-people tracking and the gait recognition in a video are two sub-problems that can help each other. Hence, we propose a joint tracking and gait recognition of multiple people as one framework that can improve gait recognition accuracy and decrease the ID switching in tracking. Experimental results confirm the validity of proposed approach. Maryam Babaee, Gerhard Rigoll, Mohammadreza Babaee |
ICIP | 2 |
| 2017 | Multi-view human activity recognition using motion frequencyabstractThe problem of human activity recognition can be approached using spatio-temporal variations in successive video frames. In this paper, a new human activity recognition technique is proposed using multi-view videos. Initially, a naive background subtraction using frame differencing between adjacent frames of a video is performed. Then, the motion information of each pixel is recorded in binary indicating existence/non-existence of motion in the frame. A pixel wise sum over all the difference images in a view gives the frequency of motion in each pixel throughout the clip. The classification performances are evaluated using these motion frequency features. Our analysis shows that increasing number of views used for feature extraction improves the performance as different views of an activity provide complementary information. Experiments on the i3DPost and the INRIA Xmas Motion Acquisition Sequences (IXMAS) multi-view human action datasets provide significant classification accuracies. Neslihan Kose, Mohammadreza Babaee, Gerhard Rigoll |
ICIP | 3 |
| 2017 | A diminished reality simulation for driver-car interaction with transparent cockpitsabstractWe anticipate advancements in mixed reality device technology which might benefit driver-car interaction scenarios and present a simulated diminished reality interface for car drivers. It runs in a custom driving simulation and allows drivers to perceive otherwise occluded objects of the environment through the car body. We expect to obtain insights that will be relevant to future real-world applications. We conducted a pre-study with participants performing a driving task with the prototype in a CAVE-like virtual environment. Users preferred large-sized see-through areas over small ones but had differing opinions on the level of transparency to use. In future work, we plan additional evaluations of the driving performance and will further extend the simulation. Patrick Lindemann, Gerhard Rigoll |
VR | 2 |
| 2017 | Combined segmentation, reconstruction, and tracking of multiple targets in multi-view video sequences
Mohammadreza Babaee, Yue You, Gerhard Rigoll |
Comput. Vis. Image Underst. | 3 |
| 2016 | Relative attribute guided dictionary learningabstractA discriminative dictionary learning algorithm is proposed to find sparse signal representations using relative attributes as the available semantic information. In contrast, existing (discriminative) dictionary learning (DDL) approaches mostly utilize binary label information to enhance the discriminative property of the signal reconstruction residual, the sparse coding vectors or both. Compared to binary attributes or labels, relative attributes contain richer semantic information where the data is annotated with the attributes' strength. In this paper we use the relative attributes of training data indirectly to learn a discriminative dictionary. Precisely, we incorporate a rank function for the attributes in the dictionary learning process. In order to assess the quality of the obtained signals, we apply k-means clustering and measure the clustering performance. Experimental results conducted on three datasets, namely the PubFig [1], OSR [2] and Shoes [3] confirm that the proposed approach outperforms the state-of-the-art label based dictionary learning algorithms. Mohammadreza Babaee, Thomas Wolf 0011, Gerhard Rigoll |
ICIP | 3 |
| 2016 | Wavelet contrast-based image inpainting with sparsity-driven initializationabstractImage inpainting is the task of removing undesired objects or flaws in images. This work advances an exemplar-based global optimization image inpainting algorithm. For that purpose, the inpainting area is iteratively refined through the minimization of a cost function. The minimization outcome depends on the initial values of the inpainting area. We compare three initialization methods with a new sparsity-driven approach. Lastly, we propose the new wavelet contrast costs which increase the inpainting quality. Wavelet contrasts reduce computational complexity in comparison to wavelet histograms while preserving their ability of measuring the density of image texture. Philipp Tiefenbacher, Michael Sirch, Mohammadreza Babaee, Gerhard Rigoll |
ICIP | 4 |
| 2016 | Multi-view gait recognition using 3D convolutional neural networksabstractIn this work we present a deep convolutional neural network using 3D convolutions for Gait Recognition in multiple views capturing spatio-temporal features. A special input format, consisting of the gray-scale image and optical flow enhance color invariance. The approach is evaluated on three different datasets, including variances in clothing, walking speeds and the view angle. In contrast to most state-of-the-art Gait Recognition systems the used neural network is able to generalize gait features across multiple large view angle changes. The results show a comparable to better performance in comparison with previous approaches, especially for large view differences. Thomas Wolf 0011, Mohammadreza Babaee, Gerhard Rigoll |
ICIP | 3 |
| 2016 | Comparison of mobile touch interfaces for object identification and troubleshooting tasks in augmented realityabstractThis work adapts common HMD interfaces for the use on a handheld device. The proposed interfaces focus on: Easy interaction on the mobile device and independence of the provided content to the user's view. We compare two AR interface techniques for object identification and three AR interfaces for troubleshooting. The results show that exocentric-based AR interfaces outperform egocentric ones in respect to completion time, walking distance and number of interactions. Philipp Tiefenbacher, Jan Gillich, Paul Schott, Gerhard Rigoll |
VR | 4 |
| 2016 | RayOnPlane: A translation technique minimizing gesture sizeabstractIn this work, we propose a device-based manipulation technique named RayOnPlane, which maintains the ease of use also in case of increasing work space size. We compare this technique to the state-of-the-art device-based manipulation technique HOMER-S. An experiment incorporating different work space sizes indicates comparable performance in completion time, while minimizing gesture size as well as user frustration and physical strain. Philipp Tiefenbacher, Clemens Techmer, Gerhard Rigoll |
VR | 3 |
| 2016 | Exploring floating stereoscopic driver-car interfaces with wide field-of-view in a mixed reality simulationabstractIn this paper, we propose a floating, multi-layered, wide field-of-view user interface for car drivers. It utilizes stereoscopic depth and focus blurring to highlight items with high priority or urgency. Individual layers are additionally used to separate groups of UI elements according to importance or context. Our work is motivated by two main prospects: a fundamentally changing driver-car interaction and ongoing technology advancements for mixed reality devices. A working prototype has been implemented as part of a custom driving simulation and will be further extended. We plan evaluations in contexts ranging from manual to fully automated driving, providing context-specific suggestions. We want to determine user preferences for layout and prioritization of the UI elements, perceived quality of the interface and effects on driving performance. Patrick Lindemann, Gerhard Rigoll |
VRST | 2 |
| 2016 | Capturing facial videos with Kinect 2.0: A multithreaded open source tool and databaseabstractDespite the growing research interest in 2.5D and 3D video face processing, 3D facial videos are actually scarcely available. This work introduces a new open source tool, named FaceGrabber, for capturing human faces using Microsoft's Kinect 2.0. FaceGrabber permits the concurrent recording of various formats, including the raw 2D and 2.5D video streams, 3D point clouds and the 3D registered face model provided by the Kinect. The software is also able to convert different data formats and playback recorded results directly in 3D. In order to encourage research with Kinect 2.0 face data, we publish a new public video face database which was captured using FaceGrabber. The database comprises 40 individuals, performing the six universal emotions (disgust, sadness, happiness, fear, anger, surprise) and two additional sequences. Daniel Merget, Tobias Eckl, Martin Schwoerer, Philipp Tiefenbacher, Gerhard Rigoll |
WACV | 5 |
| 2016 | Mono camera multi-view diminished realityabstractDiminished reality (DR) aims at removing objects auto-matically in live video streams, in a way which is unrecognizable to human observers. Approaches working on mobile devices utilize image inpainting algorithms to fill the holes (target region) of the removed objects in a plausible way. We propose a new DR algorithm, which extracts a scene model from a sparse point cloud of a mono SLAM algorithm. The scene model allows to inpaint complex scenes consisting of multiple planes intersecting at the target region. A plane-wise key frame estimation preserves scene structure as well as improving inpainting results. Furthermore, we introduce a viewing angle-based key frame update policy, ensuring high quality of the key frames. The multi-threaded design of the algorithm achieves nearly real-time performance on a mobile device. Philipp Tiefenbacher, Michael Sirch, Gerhard Rigoll |
WACV | 3 |
| 2016 | Discriminative Nonnegative Matrix Factorization for dimensionality reduction
Mohammadreza Babaee, Stefanos Tsoukalas, Maryam Babaee, Gerhard Rigoll, Mihai Datcu |
Neurocomputing | 4 |
| 2016 | Immersive visualization of visual data using nonnegative matrix factorization
Mohammadreza Babaee, Stefanos Tsoukalas, Gerhard Rigoll, Mihai Datcu |
Neurocomputing | 3 |
| 2016 | Toward semantic attributes in dictionary learning and non-negative matrix factorization
Mohammadreza Babaee, Thomas Wolf 0011, Gerhard Rigoll |
Pattern Recognit. Lett. | 3 |
| 2015 | Cross-corpus acoustic emotion recognition: Variances and strategies (Extended abstract)abstractAs the recognition of emotion from speech has matured to a degree where it becomes applicable in real-life settings, it is time for a realistic view on obtainable performances. Most studies tend to overestimation in this respect: acted data is often used rather than spontaneous data, results are reported on pre-selected prototypical data, and true speaker disjunctive partitioning is still less common than simple cross-validation. A considerably more realistic impression can be gathered by inter-set evaluation: we therefore show results employing six standard databases in a cross-corpora evaluation experiment. To better cope with the observed high variances, different types of normalization are investigated. 1.8 k individual evaluations in total indicate the crucial performance inferiority of inter- to intra-corpus testing. Björn W. Schuller, Bogdan Vlasenko, Florian Eyben, Martin Wöllmer, André Stuhlsatz, Andreas Wendemuth, Gerhard Rigoll |
ACII | 7 |
| 2015 | Attribute constrained subspace learningabstractVisual attributes are high-level semantic descriptions of visual data that are close to the human language. They have been used intensively in various applications such as image classification, active learning, and interactive search. However, the usage of attributes in subspace learning (or dimensionality reduction) has not been considered yet. In this work, we propose to utilize relative attributes as semantic cues in subspace learning. To this end, we employ Non-negative Matrix Factorization (NMF) constrained by embedded relative attributes to learn a subspace representation of image content. Experiments conducted on two datasets show the efficiency of attributes in discriminative subspace learning. Mohammadreza Babaee, Maryam Babaei, Daniel Merget, Philipp Tiefenbacher, Gerhard Rigoll |
ICIP | 5 |
| 2015 | Subjective and objective evaluation of image inpainting qualityabstractImage inpainting algorithms aim to cut out parts of the image without leaving holes. Various algorithms exist, but no wider comparison has been made, yet. This work fills the gap by comparing state-of-the-art algorithms in a user study. We create and publish a database consisting of multiple base images and inpaint them using different inpainting concepts. Afterwards, 21 participants are asked to rate the quality of these inpainted images. The subjective feedback indicates that different image inpainting algorithms are favorable depending on the characteristics of the base image and target region. Furthermore, the results show that general image quality measures such as the peak signal-to-noise ratio (PSNR) or the structural similarity (SSIM) index are not suited for judging inpainting quality. Philipp Tiefenbacher, Viktor Bogischef, Daniel Merget, Gerhard Rigoll |
ICIP | 4 |
| 2015 | Interactive feature learning from SAR image patchesabstractFeature learning algorithms aim to provide a compact and discriminative representation of complex datasets in order to increase the speed and accuracy of clustering or classification. In this paper, we propose a novel interactive feature learning approach which is mainly based on 3D interactive data visualization and Non-negative Matrix Factorization (NMF). Here, the data is visualized in a 3D interface to support human-data interaction. The user interactions are exploited in an NMF framework to learn a compact representation of the data. The conducted experiments on Synthetic Aperture Radar (SAR) images confirm the efficiency of the proposed approach. Mohammadreza Babaee, Xuejie Yu, Daniel Merget, Amir Babaeian, Gerhard Rigoll, Mihai Datcu |
IGARSS | 5 |
| 2015 | Modelling, synthesis and characterisation of occlusion in videosabstractOcclusion is one of the most challenging problems in many video processing applications such as surveillance, gait recognition, activity recognition and so on. Attempts have been made to develop algorithms for handling occlusion and evaluate their performance on various datasets. However, these studies are subjective in nature and the datasets are hardly characterised in terms of the level of occlusion, thereby precluding any form of quantitative comparison of performance. This shows a compelling need to design an explicit, unambiguous and quantitative model, which should be able to objectively represent occlusion in a video. This study proposes an occlusion model based on the position and pose uncertainties of the moving subjects in a video. The proposed occlusion model is able to characterise the level of occlusion present in a video. It is also employed to synthetically generate occlusion for walking sequences, thus providing a direction for controlled dataset generation against which human identification algorithms can be tested. Given an input video with a subject moving without any occlusion, a particle swarm optimisation‐based parameter estimation methodology is presented that generates the desired level of occlusion. The proposed approaches have been tested on the TUM‐IITKGP and PETS2010 datasets. Finally, as an application, the occlusion model has been used to generate an occluded gait datasets and the performances of different gait recognition algorithms have been compared under varying levels of occlusion. Pratik Chattopadhyay, Shamik Sural, Jayanta Mukhopadhyay, Gerhard Rigoll |
IET Comput. Vis. | 5 |
| 2014 | Farness preserving Non-negative matrix factorizationabstractDramatic growth in the volume of data made a compact and informative representation of the data highly demanded in computer vision, information retrieval, and pattern recognition. Non-negative Matrix Factorization (NMF) is used widely to provide parts-based representations by factorizing the data matrix into non-negative matrix factors. Since non-negativity constraint is not sufficient to achieve robust results, variants of NMF have been introduced to exploit the geometry of the data space. While these variants considered the local invariance based on the manifold assumption, we propose Farness preserving Non-negative Matrix Factorization (FNMF) to exploits the geometry of the data space by considering non-local invariance which is applicable to any data structure. FNMF adds a new constraint to enforce the far points (i.e., non-neighbors) in original space to stay far in the new space. Experiments on different kinds of data (e.g., Multimedia, Earth Observation) demonstrate that FNMF outperforms the other variants of NMF. Mohammadreza Babaee, Reza Bahmanyar, Gerhard Rigoll, Mihai Datcu |
ICIP | 3 |
| 2014 | PID-based regulation of background dynamics for foreground segmentationabstractIn the area of foreground and background segmentation the dynamics of a scene often influence the model of the background. There exist different methods trying to overcome the difficulty in estimating and handling these dynamics. In this paper, we propose an adaption of proportional-integral-derivative (PID) controllers to this regulation problem. In our approach, PID controllers regulate the decision threshold of the background dynamics and the update rate of the background model. The regulators are integrated in the Pixel-based Adaptive Segmenter (PBAS), which is publicly available and offers good results in the Change Detection challenge. We show that the performance of the PBAS can be further improved through an empirical tuning of the single PID constants. Philipp Tiefenbacher, Martin Hofmann 0011, Daniel Merget, Gerhard Rigoll |
ICIP | 4 |
| 2014 | Impact of Coordinate Systems on 3D Manipulations in Mobile Augmented RealityabstractMobile touch PCs allow interactions with virtual objects in augmented reality scenes. Manipulations of 3D objects are a common way of such interactions, which can be performed in three different coordinate systems: the camera-, object- and world coordinate systems. The camera coordinate system changes continuously in augmented reality as it depends on the mobile device's pose. The axis orientations of the world coordinate system are steady, whereas the axes of the object coordinates base on previous manipulations. The selection of a coordinate system therefore influences the 3D transformation's orientation independent from the used manipulation type. Philipp Tiefenbacher, Steven Wichert, Daniel Merget, Gerhard Rigoll |
ICMI | 4 |
| 2014 | Investigating NMF speech enhancement for neural network based acoustic modelsabstractIn the light of the improvements that were made in the last years with neural network-based acoustic models, it is an interesting question whether these models are also suited for noise-robust recognition. This has not yet been fully explored, although first experiments confirm this question. Furthermore, preprocessing techniques that improve the robustness should be re-evaluated with these new models. In this work, we present experimental results to address these questions. Acoustic models based on Gaussian mixture models (GMMs), deep neural networks (DNNs), and long short-term memory (LSTM) recurrent neural networks (which have an improved ability to exploit context) are evaluated for their robustness after clean or multi-condition training. In addition, the influence of non-negative matrix factorization (NMF) for speech enhancement is investigated. Experiments are performed with the Aurora-4 database and the results show that DNNs perform slightly better than LSTMs and, as expected, both beat GMMs. Furthermore, speech enhancement is capable of improving the DNN result. Index Terms: robust speech recognition, long short-term memory, speech enhancement Jürgen T. Geiger, Jort F. Gemmeke, Björn W. Schuller, Gerhard Rigoll |
INTERSPEECH | 4 |
| 2014 | Robust speech recognition using long short-term memory recurrent neural networks for hybrid acoustic modellingabstractOne method to achieve robust speech recognition in adverse conditions including noise and reverberation is to employ acoustic modelling techniques involving neural networks. Using long short-term memory (LSTM) recurrent neural networks proved to be efficient for this task in a setup for phoneme prediction in a multi-stream GMM-HMM framework. These networks exploit a self-learnt amount of temporal context, which makes them especially suited for a noisy speech recognition task. One shortcoming of this approach is the necessity of a GMM acoustic model in the multi-stream framework. Furthermore, potential modelling power of the network is lost when predicting phonemes, compared to the classical hybrid setup where the network predicts HMM states. In this work, we propose to use LSTM networks in a hybrid HMM setup, in order to overcome these drawbacks. Experiments are performed using the medium-vocabulary recognition track of the 2nd CHiME challenge, containing speech utterances in a reverberant and noisy environment. A comparison of different network topologies for phoneme or state prediction used either in the hybrid or double-stream setup shows that state prediction networks perform better than networks predicting phonemes, leading to stateof-the-art results for this database. Index Terms: acoustic modelling, robust speech recognition, neural networks, long short-term memory Jürgen T. Geiger, Zixing Zhang 0001, Felix Weninger, Björn W. Schuller, Gerhard Rigoll |
INTERSPEECH | 5 |
| 2014 | Creating automatically aligned consensus realities for AR videoconferencingabstractThis paper presents an AR videoconferencing approach merging two remote rooms into a shared workspace. Such bilateral AR telepresence inherently suffers from breaks in immersion stemming from the different physical layouts of participating spaces. As a remedy, we develop an automatic alignment scheme which ensures that participants share a maximum of common features in their physical surroundings. The system optimizes alignment with regard to initial user position, free shared floor space, camera positioning and other factors. Thus we can reduce discrepancies between different room and furniture layouts without actually modifying the rooms themselves. A description and discussion of our alignment scheme is given along with an exemplary implementation on real-world datasets. Nicolas H. Lehment, Daniel Merget, Gerhard Rigoll |
ISMAR | 3 |
| 2014 | Touch gestures for improved 3D object manipulation in mobile augmented realityabstractThis work presents three techniques for 3D manipulation on mobile touch devices, taking the specifics of mobile AR scenes into account. We compare the common direct manipulation technique with two indirect techniques, which utilize only the thumbs to perform the transformations. The evaluation of the manipulation variants is conducted in a mixed reality (MR) environment which takes advantage of the controlled conditions of a full virtual reality (VR) system. A study with 18 participants shows that the two-thumb method tops the other techniques. It performs better with respect to the total manipulation time and total number of gestures. Philipp Tiefenbacher, Andreas Pflaum, Gerhard Rigoll |
ISMAR | 3 |
| 2014 | Feature enhancement by deep LSTM networks for ASR in reverberant multisource environments
Felix Weninger, Jürgen T. Geiger, Martin Wöllmer, Björn W. Schuller, Gerhard Rigoll |
Comput. Speech Lang. | 5 |
| 2014 | The TUM Gait from Audio, Image and Depth (GAID) database: Multimodal recognition of subjects and traits
Martin Hofmann 0011, Jürgen T. Geiger, Sebastian Bachmann, Björn W. Schuller, Gerhard Rigoll |
J. Vis. Commun. Image Represent. | 5 |
| 2014 | Memory-Enhanced Neural Networks and NMF for Robust ASRabstractIn this article we address the problem of distant speech recognition for reverberant noisy environments. Speech enhancement methods, e. g., using non-negative matrix factorization (NMF), are succesful in improving the robustness of ASR systems. Furthermore, discriminative training and feature transformations are employed to increase the robustness of traditional systems using Gaussian mixture models (GMM). On the other hand, acoustic models based on deep neural networks (DNN) were recently shown to outperform GMMs. In this work, we combine a state-of-the art GMM system with a deep Long Short-Term Memory (LSTM) recurrent neural network in a double-stream architecture. Such networks use memory cells in the hidden units, enabling them to learn long-range temporal context, and thus increasing the robustness against noise and reverberation. The network is trained to predict frame-wise phoneme estimates, which are converted into observation likelihoods to be used as an acoustic model. It is of particular interest whether the LSTM system is capable of improving a robust state-of-the-art GMM system, which is confirmed in the experimental results. In addition, we investigate the efficiency of NMF for speech enhancement on the front-end side. Experiments are conducted on the medium-vocabulary task of the 2nd `CHiME' Speech Separation and Recognition Challenge, which includes reverberation and highly variable noise. Experimental results show that the average word error rate of the challenge baseline is reduced by 64% relative. The best challenge entry, a noise-robust state-of-the-art recognition system, is outperformed by 25% relative. Jürgen T. Geiger, Felix Weninger, Jort F. Gemmeke, Martin Wöllmer, Björn W. Schuller, Gerhard Rigoll |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2013 | Assessment of dimensionality reduction based on communication channel model; application to immersive information visualizationabstractWe are dealing with large-scale high-dimensional image data sets requiring new approaches for data mining where visualization plays the main role. Dimension reduction (DR) techniques are widely used to visualize high-dimensional data. However, the information loss due to reducing the number of dimensions is the drawback of DRs. In this paper, we introduce a novel metric to assess the quality of DRs in terms of preserving the structure of data. We model the dimensionality reduction process as a communication channel model transferring data points from a high-dimensional space (input) to a lower one (output). In this model, a co-ranking matrix measures the degree of similarity between the input and the output. Mutual information (MI) and entropy defined over the co-ranking matrix measure the quality of the applied DR technique. We validate our method by reducing the dimension of SIFT and Weber descriptors extracted from Earth Observation (EO) optical images. In our experiments, Laplacian Eigenmaps (LE) and Stochastic Neighbor Embedding (SNE) act as DR techniques. The experimental results demonstrate that the DR technique with the largest MI and entropy preserves the structure of data better than the others. Mohammadreza Babaee, Mihai Datcu, Gerhard Rigoll |
IEEE BigData | 3 |
| 2013 | Hypergraphs for Joint Multi-view Reconstruction and Multi-object TrackingabstractWe generalize the network flow formulation for multiobject tracking to multi-camera setups. In the past, reconstruction of multi-camera data was done as a separate extension. In this work, we present a combined maximum a posteriori (MAP) formulation, which jointly models multicamera reconstruction as well as global temporal data association. A flow graph is constructed, which tracks objects in 3D world space. The multi-camera reconstruction can be efficiently incorporated as additional constraints on the flow graph without making the graph unnecessarily large. The final graph is efficiently solved using binary linear programming. On the PETS 2009 dataset we achieve results that significantly exceed the current state of the art. Martin Hofmann 0011, Daniel Wolf, Gerhard Rigoll |
CVPR | 3 |
| 2013 | iProgram: intuitive programming of an industrial hri cell
Jürgen Blume, Alexander Bannat, Gerhard Rigoll |
HRI | 3 |
| 2013 | Gait-based person identification by spectral, cepstral and energy-related audio featuresabstractWith this work, we address the problem of acoustic gait-based person identification, which is the task of identifying humans by the sounds they make while walking. We examine several acoustic features from speech processing tasks for their suitability for acoustic gait recognition. Using a wrapper-based feature selection technique, we reduce the feature set while at the same time increasing the identification accuracy by 10% (relative). For classification, Support Vector Machines (SVMs) are employed. Experiments are conducted using the TUM GAID database, which is a large gait recognition database containing 3 050 recordings of 305 subjects in three variations. Jürgen T. Geiger, Martin Hofmann 0011, Björn W. Schuller, Gerhard Rigoll |
ICASSP | 4 |
| 2013 | Probabilistic asr feature extraction applying context-sensitive connectionist temporal classification networksabstractThis paper proposes a novel automatic speech recognition (ASR) front-end that unites the principles of bidirectional Long Short-Term Memory (BLSTM), Connectionist Temporal Classification (CTC), and Bottleneck (BN) feature generation. BLSTM networks are known to produce better probabilistic ASR features than conventional multilayer perceptrons since they are able to exploit a self-learned amount of temporal context for phoneme estimation. Combining BLSTM networks with a CTC output layer implies the advantage that the network can be trained on unsegmented data so that the quality of phoneme prediction does not rely on potentially error-prone forced alignment segmentations of the training set. In challenging ASR scenarios involving highly spontaneous, disfluent, and noisy speech, our BN-CTC front-end leads to remarkable word accuracy improvements and prevails over a series of previously introduced BLSTM-based ASR systems. Martin Wöllmer, Björn W. Schuller, Gerhard Rigoll |
ICASSP | 3 |
| 2013 | Feature enhancement by bidirectional LSTM networks for conversational speech recognition in highly non-stationary noiseabstractThe recognition of spontaneous speech in highly variable noise is known to be a challenge, especially at low signal-to-noise ratios (SNR). In this paper, we investigate the effect of applying bidirectional Long Short-Term Memory (BLSTM) recurrent neural networks for speech feature enhancement in noisy conditions. BLSTM networks tend to prevail over conventional neural network architectures, whenever the recognition or regression task relies on an intelligent exploitation of temporal context information. We show that BLSTM networks are well-suited for mapping from noisy to clean speech features and that the obtained recognition performance gain is partly complementary to improvements via additional techniques such as speech enhancement by non-negative matrix factorization and probabilistic feature generation by Bottleneck-BLSTM networks. Compared to simple multi-condition training or feature enhancement via standard recurrent neural networks, our BLSTM-based feature enhancement approach leads to remarkable gains in word accuracy in a highly challenging task of recognizing spontaneous speech at SNR levels between -6 and 9 dB. Martin Wöllmer, Zixing Zhang 0001, Felix Weninger, Björn W. Schuller, Gerhard Rigoll |
ICASSP | 5 |
| 2013 | Exploiting gradient histograms for gait-based person identificationabstractIn this paper, we exploit gradient histograms for person identification based on gait. A traditional and successful method for gait recognition is the Gait Energy Image (GEI). Here, person silhouettes are averaged over full gait cycles, which leads to a robust and efficient representation. However, binarized silhouettes only capture edge information at the boundary of the person. By contrast, the Gradient Histogram Energy Image (GHEI) also captures edges within the silhouette by means of gradient histograms. Combined with precise α-matte preprocessing and with a new part-based extension, recognition performance can be further improved. In addition, we show, that GEI can even be outperformed by directly applying gradient histogram extraction on the already bina-rized silhouettes. We run all experiments on the widely used HumanID gait database and show significant performance improvements over the current state of the art. Martin Hofmann 0011, Gerhard Rigoll |
ICIP | 2 |
| 2013 | Using linguistic information to detect overlapping speechabstractOverlapping speech is still a major cause of error in many speech processing applications, currently without any satisfactory solution. This paper considers the problem of detecting segments of overlapping speech within meeting recordings. Using an HMM-based framework recordings are segmented into intervals containing non-speech, speech and overlapping speech. New to this contribution is the use of linguistic information, where spoken content is used to improve overlap detection. Using language models for speech and overlap, an overlap score is created for every spoken word and used as an additional feature within the HMM framework. Experiments conducted on the AMI corpus demonstrate the potential of the proposed linguistic features. Jürgen T. Geiger, Florian Eyben, Nicholas W. D. Evans, Björn W. Schuller, Gerhard Rigoll |
INTERSPEECH | 5 |
| 2013 | Detecting overlapping speech with long short-term memory recurrent neural networksabstractDetecting segments of overlapping speech (when two or more speakers are active at the same time) is a challenging problem.Previously, mostly HMM-based systems have been used for overlap detection, employing various different audio features.In this work, we propose a novel overlap detection system using Long Short-Term Memory (LSTM) recurrent neural networks.LSTMs are used to generate framewise overlap predictions which are applied for overlap detection.Furthermore, a tandem HMM-LSTM system is obtained by adding LSTM predictions to the HMM feature set.Experiments with the AMI corpus show that overlap detection performance of LSTMs is comparable to HMMs.The combination of HMMs and LSTMs improves overlap detection by achieving higher recall. Jürgen T. Geiger, Florian Eyben, Björn W. Schuller, Gerhard Rigoll |
INTERSPEECH | 4 |
| 2013 | Classification of images in fog and fog-free scenes for use in vehiclesabstractToday modern vehicles are often equipped with a camera, which captures the scene in front of the vehicle. The recognition of weather conditions with this camera can help to improve many applications as well as establish new ones. In this article we will show how it is possible to distinguish between scenes with clear and foggy weather situations. The proposed method uses only gray-scale images as input signal and is running in real time. Using spectral features and a simple linear classifier, we can achieve high detection rates in both daytime and night-time scenes. Furthermore, we will show that in our application area these features outperform others. Mario Pavlic, Gerhard Rigoll, Slobodan Ilic |
Intelligent Vehicles Symposium | 2 |
| 2013 | Programming concept for an industrial HRI packaging cellabstractThis paper presents an overview about a programming concept for an industrial HRI cell designed for packaging of electronic consumer goods. The focus of this work lies on the interplay of the involved software components. Furthermore, developed methods for programming the software components of the HRI cell are described within a sample use case. Finally, the usability of the programming concept and a short evaluation of relevant components are presented. Jürgen Blume, Alexander Bannat, Gerhard Rigoll, M. Rooker, A. Angerer, Claus Lenz |
RO-MAN | 3 |
| 2013 | Noise robust ASR in reverberated multisource environments applying convolutive NMF and Long Short-Term Memory
Martin Wöllmer, Felix Weninger, Jürgen T. Geiger, Björn W. Schuller, Gerhard Rigoll |
Comput. Speech Lang. | 5 |
| 2013 | Using Segmented 3D Point Clouds for Accurate Likelihood Approximation in Human Pose Tracking
Nicolas H. Lehment, Moritz Kaiser, Gerhard Rigoll |
Int. J. Comput. Vis. | 3 |
| 2013 | Towards using covariance matrix pyramids as salient point descriptors in 3D point clouds
Moritz Kaiser, Xiao Xu 0001, Bogdan Kwolek, Shamik Sural, Gerhard Rigoll |
Neurocomputing | 5 |
| 2013 | LSTM-Modeling of continuous emotions in an audiovisual affect recognition framework
Martin Wöllmer, Moritz Kaiser, Florian Eyben, Björn W. Schuller, Gerhard Rigoll |
Image Vis. Comput. | 5 |
| 2013 | Keyword spotting exploiting Long Short-Term Memory
Martin Wöllmer, Björn W. Schuller, Gerhard Rigoll |
Speech Commun. | 3 |
| 2012 | Speech overlap detection and attribution using convolutive non-negative sparse codingabstractOverlapping speech is known to degrade speaker diarization performance with impacts on speaker clustering and segmentation. While previous work made important advances in detecting overlapping speech intervals and in attributing them to relevant speakers, the problem remains largely unsolved. This paper reports the first application of convolutive non-negative sparse coding (CNSC) to the overlap problem. CNSC aims to decompose a composite signal into its underlying contributory parts and is thus naturally suited to overlap detection and attribution. Experimental results on NIST RT data show that the CNSC approach gives comparable results to a state-of-the-art hidden Markov model based overlap detector. In a practical diarization system, CNSC based speaker attribution is shown to reduce the speaker error by over 40% relative in overlapping segments. Ravichander Vipperla, Jürgen T. Geiger, Simon Bozonnet, Dong Wang 0013, Nicholas W. D. Evans, Björn W. Schuller, Gerhard Rigoll |
ICASSP | 7 |
| 2012 | Non-negative matrix factorization for highly noise-robust ASR: To enhance or to recognize?abstractThis paper proposes a multi-stream speech recognition system that combines information from three complementary analysis methods in order to improve automatic speech recognition in highly noisy and reverberant environments, as featured in the 2011 PASCAL CHiME Challenge. We integrate word predictions by a bidirectional Long Short-Term Memory recurrent neural network and non-negative sparse classification (NSC) into a multi-stream Hidden Markov Model using convolutive non-negative matrix factorization (NMF) for speech enhancement. Our results suggest that NMF-based enhancement and NSC are complementary despite their overlap in methodology, reaching up to 91.9% average keyword accuracy on the Challenge test set at signal-to-noise ratios from -6 to 9 dB-the best result reported so far on these data. Felix Weninger, Martin Wöllmer, Jürgen T. Geiger, Björn W. Schuller, Jort F. Gemmeke, Antti Hurmalainen, Tuomas Virtanen, Gerhard Rigoll |
ICASSP | 8 |
| 2012 | Improved Gait Recognition using Gradient Histogram Energy ImageabstractWe present a new spatio-temporal representation for Gait Recognition, which we call Gradient Histogram Energy Image (GHEI). Similar to the successful Gait Energy Image (GEI), information is averaged over full gait cycles to reduce noise. Contrary to GEI, where silhouettes are averaged and thus only edge information at the boundary is used, our GHEI computes gradient histograms at all locations of the original image. Thus, also edge information inside the person silhouette is captured. In addition, we show that GHEI can be greatly improved using precise segmentation techniques (we use α-matte segmentation). We demonstrate great effectiveness of GHEI and its variants in our experiments on the large and widely used HumanID Gait Challenge dataset. On this dataset we reach a significant performance gain over the current state of the art. Martin Hofmann 0011, Gerhard Rigoll |
ICIP | 2 |
| 2012 | Improving generalisation and robustness of acoustic affect recognitionabstractEmotion recognition in real-life conditions faces several challenging factors, which most studies on emotion recognition do not consider. Such factors include background noise, varying recording levels, and acoustic properties of the environment, for example. This paper presents a systematic evaluation of the influence of background noise of various types and SNRs, as well as recording level variations on the performance of automatic emotion recognition from speech. Both, natural and spontaneous as well as acted/prototypical emotions are considered. Besides the well known influence of additive noise, a significant influence of the recording level on the recognition performance is observed. Multi-condition learning with various noise types and recording levels is proposed as a way to increase robustness of methods based on standard acoustic feature sets and commonly used classifiers. It is compared to matched conditions learning and is found to be almost on par for many settings. Florian Eyben, Björn W. Schuller, Gerhard Rigoll |
ICMI | 3 |
| 2012 | Convolutive Non-Negative Sparse Coding and New Features for Speech Overlap Handling in Speaker DiarizationabstractThe effective handling of overlapping speech is at the limits of the current state of the art in speaker diarization.This paper presents our latest work in overlap detection.We report the combination of features derived through convolutive nonnegative sparse coding and new energy, spectral and voicingrelated features within a conventional HMM system.Overlap detection results are fully integrated into our top-down diarization system through the application of overlap exclusion and overlap labeling.Experiments on a subset of the AMI corpus show that the new system delivers significant reductions in missed speech and speaker error.Through overlap exclusion and labelling the overall diarization error rate is shown to improve by 6.4 % relative. Jürgen T. Geiger, Ravichander Vipperla, Simon Bozonnet, Nicholas W. D. Evans, Björn W. Schuller, Gerhard Rigoll |
INTERSPEECH | 6 |
| 2012 | Temporal and Situational Context Modeling for Improved Dominance Recognition in MeetingsabstractWe present and evaluate a novel approach towards automatically detecting a speaker's level of dominance in a meeting scenario.Since previous studies reveal that audio appears to be the most important modality for dominance recognition, we focus on the analysis of the speech signals recorded in multiparty meetings.Unlike recently published techniques which concentrate on frame-level hidden Markov modeling, we propose a recognition framework operating on segmental data and investigate context modeling on three different levels to explore possible performance gains.First, we apply a set of statistical functionals to capture large-scale feature-level context within a speech segment.Second, we consider bidirectional Long Short-Term Memory recurrent neural networks for long-range temporal context modeling between segments.Finally, we evaluate the benefit of situational context incorporation by simultaneously modeling speech of all meeting participants.Overall, our approach leads to a remarkable increase of recognition accuracy when compared to hidden Markov modeling. Martin Wöllmer, Florian Eyben, Björn W. Schuller, Gerhard Rigoll |
INTERSPEECH | 4 |
| 2012 | Interface design for an inexpensive hands-free collaborative videoconferencing systemabstractIn this paper an interaction framework for AR enhanced video conferencing is presented. The goal is to provide a cheap and portable system based on a combination of commodity Kinect cameras and regular computer screens. These conditions necessitate the use of contact free interaction methods. The interaction framework presented in this paper is specifically suited for remotely presenting, sharing and annotating visual data such as images, presentation slides and 3D objects. In the proposed system all data is represented by freely manipulable 3D objects which are augmented into the camera views. These representations are integrated into a differentiated ownership scheme, allowing for operations such as spatially managed data sharing. The suitability of different interaction paradigms with regards to this usage scenario is examined. Furthermore, occlusion and collision management between virtual objects and real obstacles is enabled by integrating basic models of the environment. Nicolas H. Lehment, Katharina Erhardt, Gerhard Rigoll |
ISMAR | 3 |
| 2012 | Image based fog detection in vehiclesabstractModern vehicles are equipped with many cameras and their use in many practical applications is extensive. Detecting the presence of fog from images of a camera mounted in vehicles is a very challenging task with the potential to be used in many practical applications. Approaches introduced until now analyze properties of local objects in the image like lane markings, traffic signs, back lights of vehicles in front or head lights of approaching vehicles. By contrast to all these related works we propose to use image descriptors and a classification procedure in order to distinguish images with fog present from those free of fog. These image descriptors are global and describe the entire image using Gabor filters at different frequencies, scales and orientations. Our experiments demonstrated hight potential of the proposed method for fog detection on daytime images. Mario Pavlic, Heidrun Belzner, Gerhard Rigoll, Slobodan Ilic |
Intelligent Vehicles Symposium | 3 |
| 2011 | A novel bottleneck-BLSTM front-end for feature-level context modeling in conversational speech recognitionabstractWe present a novel automatic speech recognition (ASR) front-end that unites Long Short-Term Memory context modeling, bidirectional speech processing, and bottleneck (BN) networks for enhanced Tandem speech feature generation. Bidirectional Long Short-Term Memory (BLSTM) networks were shown to be well suited for phoneme recognition and probabilistic feature extraction since they efficiently incorporate a flexible amount of long-range temporal context, leading to better ASR results than conventional recurrent networks or multi-layer perceptrons. Combining BLSTM modeling and bottleneck feature generation allows us to produce feature vectors of arbitrary size, independent of the network training targets. Experiments on the COSINE and the Buckeye corpora containing spontaneous, conversational speech show that the proposed BN-BLSTM front-end leads to better ASR accuracies than previously proposed BLSTM-based Tandem and multi-stream systems. Martin Wöllmer, Björn W. Schuller, Gerhard Rigoll |
ASRU | 3 |
| 2011 | Localization of non-linguistic events in spontaneous speech by Non-Negative Matrix Factorization and Long Short-Term MemoryabstractFeatures generated by Non-Negative Matrix Factorization (NMF) have successfully been introduced into robust speech processing, including noise-robust speech recognition and detection of non-linguistic vocalizations. In this study, we introduce a novel tandem approach by integrating likelihood features derived from NMF into Bidirectional Long Short-Term Memory Recurrent Neural Networks (BLSTM-RNNs) in order to dynamically localize non-linguistic events, i. e., laughter, vocal, and non-vocal noise, in highly spontaneous speech. We compare our tandem architecture to a baseline conventional phoneme-HMM-based speech recognizer, and achieve a relative reduction of the frame error rate by 37.5 % in the discrimination of speech and different non-speech segments. Felix Weninger, Björn W. Schuller, Martin Wöllmer, Gerhard Rigoll |
ICASSP | 4 |
| 2011 | A multi-stream ASR framework for BLSTM modeling of conversational speechabstractWe propose a novel multi-stream framework for continuous conversational speech recognition which employs bidirectional Long Short-Term Memory (BLSTM) networks for phoneme prediction. The BLSTM architecture allows recurrent neural nets to model long range context, which led to improved ASR performance when combined with conventional triphone modeling in a Tandem system. In this paper, we extend the principle of joint BLSTM and triphone modeling to a multi-stream system which uses MFCC features and BLSTM predictions as observations originating from two independent data streams. Using the COSINE database, we show that this technique prevails over a recently proposed single-stream Tandem system as well as over a conventional HMM recognizer. Martin Wöllmer, Florian Eyben, Björn W. Schuller, Gerhard Rigoll |
ICASSP | 4 |
| 2011 | Dense point-to-point correspondences between 3D faces with large variations for constructing 3D Morphable ModelsabstractIn this contribution a novel method to compute dense point-to- point correspondences between 3D faces is presented. The faces are aligned in 3D space with a Generalized Procrustes Analysis and sub- sequently mapped into 2D space. To compute a correspondence flow between two faces an energy function is minimized which is based on the following assumptions: smoothness of the flow, mapping of landmarks on their counterparts, and texture and depth consistency. Based on these correspondences, the 3D faces are resampled, so that each face is represented by the same amount of 3D points and for any point there is a corresponding point in all other faces. The accuracy of the point-to-point correspondences is demonstrated on the basis of two applications, namely facial texture mapping and the construction of 3D Morphable Models. Moritz Kaiser, Nicolas H. Lehment, Gerhard Rigoll |
ICIP | 3 |
| 2011 | Learning New Acoustic Events in an HMM-Based System Using MAP AdaptationabstractIn this paper, we present a system for the recognition of acoustic events suited for a robotic application.HMMs are used to model different acoustic event classes.We are especially looking at the open-set case, where a class of acoustic events occurs that was not included in the training phase.It is evaluated how newly occuring classes can be learnt using MAP adaptation or conventional training methods.A small database of acoustic events was recorded with a robotic platform to perform the experiments. Jürgen T. Geiger, Mohamed Anouar Lakhal, Björn W. Schuller, Gerhard Rigoll |
INTERSPEECH | 4 |
| 2011 | Using Multiple Databases for Training in Emotion Recognition: To Unite or to Vote?abstractWe present an extensive study on the performance of data agglomeration and decision-level fusion for robust cross-corpus emotion recognition.We compare joint training with multiple databases and late fusion of classifiers trained on single databases, employing six frequently used corpora of natural or elicited emotion, namely ABC, AVIC, DES, eNTERFACE, SAL, VAM, and three classifiers i. e. SVM, Random Forests, Naïve Bayes to best cover for singular effects.On average over classifier and database, data agglomeration and majority voting deliver relative improvements of unweighted accuracy by 9.0 % and 4.8 %, respectively, over single-database cross-corpus classification of arousal, while majority voting performs best for valence recognition. Björn W. Schuller, Zixing Zhang 0001, Felix Weninger, Gerhard Rigoll |
INTERSPEECH | 4 |
| 2011 | Feature Frame Stacking in RNN-Based Tandem ASR Systems - Learned vs. Predefined ContextabstractAs phoneme recognition is known to profit from techniques that consider contextual information, neural networks applied in Tandem automatic speech recognition (ASR) systems usually employ some form of context modeling. While approaches based on multi-layer perceptrons or recurrent neural networks (RNN) are able to model a predefined amount of context by simultaneously processing a stacked sequence of successive feature vectors, bidirectional Long Short-Term Memory (BLSTM) networks were shown to be well-suited for incorporating a selflearned amount of context for phoneme prediction. In this paper, we evaluate combinations of BLSTM modeling and frame stacking to determine the most efficient method for exploiting context in RNN-based Tandem systems. Applying the CO-SINE corpus and our recently introduced multi-stream BLSTM-HMM decoder, we provide empirical evidence for the intuition that BLSTM networks redundantize frame stacking while RNNs profit from predefined feature-level context. Index Terms: context modeling, long short-term memory, recurrent neural networks, automatic speech recognition Martin Wöllmer, Björn W. Schuller, Gerhard Rigoll |
INTERSPEECH | 3 |
| 2011 | A large-scale LED array to support anticipatory drivingabstractWe present a novel assistance system which supports anticipatory driving by means of fostering early deceleration. Upcoming technologies like Car2X communication provide information about a time interval which is currently uncovered. This information shall be used in the proposed system to inform drivers about future situations which require reduced speed. Such situations include traffic jams, construction sites or speed limits. The HMI is an optical output system based on line arrays of RGB-LEDs. Our contribution presents construction details as well as user evaluations. The results show an earlier deceleration of 3.9-11.5 s and a shorter deceleration distance of 2-166 m. Florian Laquai, Fabian Chowanetz, Gerhard Rigoll |
SMC | 3 |
| 2011 | Gaze-based interaction on multiple displays in an automotive environmentabstractThis paper presents a multimodal interaction system for automotive environments that uses the driver's eyes as main input device. Therefore, an unobtrusive and contactless sensor analyzes the driver's eye gaze, which enables the development of gaze driven interaction concepts for operating driver assistance and infotainment systems. The following sections present the developed interaction concepts, the used gaze tracking system, and the test setup consisting of multiple monitors and a large touchscreen as central interaction screen. Finally the comparison results of the gaze-based interaction with a more conventional touch interaction are being discussed. Therefore, well-defined tasks were completed by participants and task completion times, distraction and cognitive load were recorded and analyzed. The tests show promising results for gaze driven interaction. Tony Poitschke, Florian Laquai, Stilyan Stamboliev, Gerhard Rigoll |
SMC | 4 |
| 2011 | Dense point-to-point correspondences between 3D faces using parametric remeshing for constructing 3D Morphable ModelsabstractIn this contribution a novel method to compute dense point-to-point correspondences between 3D faces is presented. The correspondences can be employed for various face processing applications, for example for building up a 3D Morphable Model (3DMM). Paths connecting landmarks are traced on the 3D facial surface and the resulting patches are mapped into a uv-space. Triangle quadrisection is used to build up remeshes with high point density for each 3D facial surface. Each vertex of a remesh has one corresponding vertex in another remesh and all remeshes have the same connectivity. The quality of the point-to-point correspondences is demonstrated on the bases of two applications, namely morphing and constructing a 3DMM. Moritz Kaiser, Gernot Heym, Nicolas H. Lehment, Dejan Arsic, Gerhard Rigoll |
WACV | 5 |
| 2010 | Multiple Parallel Vision-Based Recognition in a Real-Time Framework for Human-Robot-Interaction ScenariosabstractEvery day human communication relies on a large number of different communication mechanisms like spoken language, facial expressions, body pose and gestures, allowing humans to pass large amounts of information in short time. In contrast, traditional human-machine communication is often unintuitive and requires specifically trained personal. In this paper, we present a real-time capable framework that recognizes traditional visual human communication signals in order to establish a more intuitive human-machine interaction. Humans rely on the interaction partner’s face for identification, which helps them to adapt to the interaction partner and utilize context information. Head gestures (head nodding and head shaking) are a convenient way to show agreement or disagreement. Facial expressions give evidence about the interaction partners’ emotional state and hand gestures are a fast way of passing simple commands. The recognition of all interaction queues is performed in parallel, enabled by a shared memory implementation. Tobias Rehrl, Alexander Bannat, Jürgen Gast, Frank Wallhoff, Gerhard Rigoll, Christoph Mayer 0001, Zahid Riaz, Bernd Radig, Stefan Sosnowski, Kolja Kühnlenz |
ACHI | 5 |
| 2010 | Depth gradient based segmentation of overlapping foreground objects in range images
Andre Störmer, Martin Hofmann 0011, Gerhard Rigoll |
FUSION | 3 |
| 2010 | Non-negative matrix factorization as noise-robust feature extractor for speech recognitionabstractWe introduce a novel approach for noise-robust feature extraction in speech recognition, based on non-negative matrix factorization (NMF). While NMF has previously been used for speech denoising and speaker separation, we directly extract time-varying features from the NMF output. To this end we extend basic unsupervised NMF to a hybrid supervised/unsupervised algorithm. We present a Dynamic Bayesian Network (DBN) architecture that can exploit these features in a Tandem manner together with the maximum likelihood phoneme estimate of a bidirectional long short-term memory (BLSTM) recurrent neural network. We show that addition of NMF features to spelling recognition systems can increase word accuracy by up to 7% absolute in a noisy car environment. Björn W. Schuller, Felix Weninger, Martin Wöllmer, Gerhard Rigoll |
ICASSP | 5 |
| 2010 | Spoken term detection with Connectionist Temporal Classification: A novel hybrid CTC-DBN decoderabstractThis paper proposes a novel system for robust keyword detection in continuous speech. Our decoder is composed of a bidirectional Long Short-Term Memory recurrent neural network using a Connectionist Temporal Classification (CTC) output layer, and a Dynamic Bayesian Network (DBN). The CTC network exploits bidirectional context information to reliably identify phonemes, whereas the DBN is able to discriminate between keywords and arbitrary speech while explicitly modeling substitutions, deletions, and insertions in the CTC phoneme output string. Our technique is vocabulary independent and does not require an explicit garbage model. Experiments show that our system architecture prevails over a standard Hidden Markov Model approach. Martin Wöllmer, Florian Eyben, Björn W. Schuller, Gerhard Rigoll |
ICASSP | 4 |
| 2010 | Optimizing the Number of States for HMM-Based On-line Handwritten Whiteboard RecognitionabstractIn this paper, we present a novel way to determine the number of states in Hidden-Markov-Models for on-line handwriting recognition. This method extends the Bakis length modeling method which has successfully been applied to off-line handwriting recognition. We propose a modification to the Bakis method and present a technique to improve the topology with a small number of iterations. Furthermore, we investigate the influence of state tying. In an experimental section, we show that our improved system outperforms a system with Bakis length modeling by 1.5 % relative and with fixed length modeling by 5.1 % relative on the IAM-On-DB-t1 benchmark. Jürgen T. Geiger, Joachim Schenk, Frank Wallhoff, Gerhard Rigoll |
ICFHR | 4 |
| 2010 | Selecting Features Using the SFS in Conjunction with Vector QuantizationabstractWhen discrete Hidden-Markov-Models (HMMs)-based recognition is performed, vector quantization (VQ) is used to transform continuous observations to sequences of discrete symbols. After VQ, the quantization error is not spread equally among the features. This impairs the feature significance, which is important when features are selected, e. g. by applying the Sequential Forward Selection (SFS). In this paper, we introduce a novel vector quantization (VQ) scheme for distributing the quantization error equally among the quantized dimensions of a feature vector. Afterwards, the proposed VQ scheme is used to apply the SFS on the features in on-line handwritten whiteboard note recognition based on discrete HMMs. In an experimental section, we show that the novel VQ scheme derives feature sets of almost half the size of the feature sets gained when standard VQ is used for quantization, while the performance stays the same. Joachim Schenk, Gerhard Rigoll |
ICFHR | 2 |
| 2010 | Robust tracking of facial feature points with 3D Active Shape ModelsabstractExact 3D tracking of facial feature points is appealing for many applications in human-machine interaction. In this work a 3D Active Shape Model (ASM) that can be shifted, scaled, and rotated is used to track the points. The efficient Gauss-Newton method is applied to estimate the 3D ASM, rotation, translation, and scale parameters. If the head turns to one side, some points might be occluded but they are still considered for the estimation of the parameters. A robust error norm that reduces (or ideally cancels) the influence of occluded points is applied. With some algebraic transformations the computational cost per frame can be further reduced. The proposed algorithm is evaluated on the basis of the Airplane Behavior Corpus. Index Terms — Tracking, face recognition, minimization methods, robustness 1. Moritz Kaiser, Dejan Arsic, Shamik Sural, Gerhard Rigoll |
ICIP | 4 |
| 2010 | Cue-independent extending inverse kinematics for robust pose estimation in 3D point cloudsabstractWhile monocular gesture recognition slowly reaches maturity, the inclusion of 3D gestures remains a challenge. In order to enable robust and versatile depth-enabled gestures, a depth-image based tracking approach is developed. Using a model-based annealing particle filter approach, the pose of a single subject is retrieved and tracked over longer image and motion sequences. Other than many previous depth-image based systems, full body tracking is performed. The system is independent from specific camera types and is independent from color or texture cues. Pose space exploration in complex kinematic chains is enhanced by considering extending inverse kinematics. Exploiting the highly parallel nature of the 3D point based approach, the algorithm is partially implemented on a GPU, leading to near real time performance. Nicolas H. Lehment, Moritz Kaiser, Dejan Arsic, Gerhard Rigoll |
ICIP | 4 |
| 2010 | Tracking using Bayesian inference with a two-layer Graphical ModelabstractThis paper introduces a new visual tracking technique combining particle filtering and Dynamic Bayesian Networks. The particle filter is utilized to robustly track an object in a video sequence and gain sets of descriptive object features. Dynamic Bayesian Networks use feature sequences to determine different motion patterns. A Graphical Model is introduced, which combines particle filter based tracking with Dynamic Bayesian Network-based classification. This unified framework allows for enhancing the tracking by adapting the dynamical model of the tracking process according to the classification results obtained from the Dynamic Bayesian Network. Therefore, the tracking step and classification step form a closed tracking-classification-tracking loop. In the first layer of the Graphical Model a particle filter is set up, whereas the second layer builds up the dynamical model of the particle filter based on the classification process of the Dynamic Bayesian Network. Tobias Rehrl, Nikolaus Theißing, Alexander Bannat, Jürgen Gast, Dejan Arsic, Frank Wallhoff, Gerhard Rigoll |
ICIP | 7 |
| 2010 | Graphical Models for real-time capable gesture recognitionabstractIn everyday live head gestures such as head shaking or nodding and hand gestures like pointing gestures form important aspects of human-human interaction. Therefore, recent research considers integrating these intuitive communication cues into technical systems for improving and easing human-computer interaction. In this paper we present a vision-based system to recognize head gestures (nodding, shaking, neutral) and dynamic hand gestures (hand moving right/left/up/down, fist moving right/left) in real-time. The gestural input delivers a communication modality for a human-robot interaction scenario situated in an assistive household environment. The use of fast low-level image-feature extraction methods contributes to the real-time capability of the system and advanced classification approaches relying on Graphical Models provide high robustness. Graphical Models offer the possibility to group the input features in several sub-nodes resulting in a better classification than obtained via a traditional Hidden Markov Model classification. The applied grouping can regard interdependencies owing to, either physical constraints (like for the head gestures), or interrelations between shape and motion (like for the hand gestures). Tobias Rehrl, Nikolaus Theißing, Alexander Bannat, Jürgen Gast, Dejan Arsic, Frank Wallhoff, Gerhard Rigoll, Christoph Mayer 0001, Bernd Radig |
ICIP | 7 |
| 2010 | Registration of 3D facial surfaces using covariance matrix pyramidsabstractRegistration of 3D facial surfaces means establishing point-to-point correspondence between two 3D facial surfaces. Difficulties typical for the registration of 3D facial surfaces are varying illumination, pose or viewpoint changes, varying facial expressions, and different appearance of individuals. In this work we propose to use a covariance matrix as descriptor for the neighborhood of a salient point in a face. It encodes the variance of the channels, such as red, green, blue, depth, etc., their correlations with each other, and spatial layout, while filtering out the influence of the disturbing effects mentioned above. A pyramidal approach is applied where first the location of a corresponding point is computed roughly and then the position is gradually refined. The method does not require any training. Particle Swarm Optimization makes the search for corresponding points more efficient. Results with a challenging dataset confirm that the approach works greatly for a variety of disturbing effects. Moritz Kaiser, Bogdan Kwolek, Christoph Staub, Gerhard Rigoll |
ICRA | 4 |
| 2010 | GMM-UBM based open-set online speaker diarizationabstractIn this paper, we present an open-set online speaker diarization system.The system is based on Gaussian mixture models (GMMs), which are used as speaker models.The system starts with just 3 such models (one each for both genders and one for non-speech) and creates models for individual speakers not till the speakers occur.As more and more speakers appear, more models are created.Our system implicitly performs audio segmentation, speech/non-speech classification, gender recognition and speaker identification.The system is tested with the HUB4-1996 radio broadcast news database. Jürgen T. Geiger, Frank Wallhoff, Gerhard Rigoll |
INTERSPEECH | 3 |
| 2010 | Recognition of spontaneous conversational speech using long short-term memory phoneme predictionsabstractWe present a novel continuous speech recognition framework designed to unite the principles of triphone and Long Short-Term Memory (LSTM) modeling. The LSTM principle allows a recurrent neural network to store and to retrieve information over long time periods, which was shown to be well-suited for the modeling of co-articulation effects in human speech. Our system uses a bidirectional LSTM network to generate a phoneme prediction feature that is observed by a triphone-based large-vocabulary continuous speech recognition (LVCSR) decoder, together with conventional MFCC features. We evaluate both, phoneme prediction error rates of various network architectures and the word recognition performance of our Tandem approach using the COSINE database- a large corpus of conversational and noisy speech, and show that incorporating LSTM phoneme predictions in to an LVCSR system leads to significantly higher word accuracies. Martin Wöllmer, Florian Eyben, Björn W. Schuller, Gerhard Rigoll |
INTERSPEECH | 4 |
| 2010 | Cross-Corpus Acoustic Emotion Recognition: Variances and StrategiesabstractAs the recognition of emotion from speech has matured to a degree where it becomes applicable in real-life settings, it is time for a realistic view on obtainable performances. Most studies tend to overestimation in this respect: Acted data is often used rather than spontaneous data, results are reported on preselected prototypical data, and true speaker disjunctive partitioning is still less common than simple cross-validation. Even speaker disjunctive evaluation can give only a little insight into the generalization ability of today's emotion recognition engines since training and test data used for system development usually tend to be similar as far as recording conditions, noise overlay, language, and types of emotions are concerned. A considerably more realistic impression can be gathered by interset evaluation: We therefore show results employing six standard databases in a cross-corpora evaluation experiment which could also be helpful for learning about chances to add resources for training and overcoming the typical sparseness in the field. To better cope with the observed high variances, different types of normalization are investigated. 1.8 k individual evaluations in total indicate the crucial performance inferiority of inter to intracorpus testing. Björn W. Schuller, Bogdan Vlasenko, Florian Eyben, Martin Wöllmer, André Stuhlsatz, Andreas Wendemuth, Gerhard Rigoll |
IEEE Trans. Affect. Comput. | 7 |
| 2009 | Acoustic emotion recognition: A benchmark comparison of performancesabstractIn the light of the first challenge on emotion recognition from speech we provide the largest-to-date benchmark comparison under equal conditions on nine standard corpora in the field using the two pre-dominant paradigms: modeling on a frame-level by means of hidden Markov models and supra-segmental modeling by systematic feature brute-forcing. Investigated corpora are the ABC, AVIC, DES, EMO-DB, eNTERFACE, SAL, SmartKom, SUSAS, and VAM databases. To provide better comparability among sets, we additionally cluster each database's emotions into binary valence and arousal discrimination tasks. In the result large differences are found among corpora that mostly stem from naturalistic emotions and spontaneous speech vs. more prototypical events. Further, supra-segmental modeling proves significantly beneficial on average when several classes are addressed at a time. Björn W. Schuller, Bogdan Vlasenko, Florian Eyben, Gerhard Rigoll, Andreas Wendemuth |
ASRU | 4 |
| 2009 | Robust vocabulary independent keyword spotting with graphical modelsabstractThis paper introduces a novel graphical model architecture for robust and vocabulary independent keyword spotting which does not require the training of an explicit garbage model. We show how a graphical model structure for phoneme recognition can be extended to a keyword spotter that is robust with respect to phoneme recognition errors. We use a hidden garbage variable together with the concept of switching parents to model keywords as well as arbitrary speech. This implies that keywords can be added to the vocabulary without having to re-train the model. Thereby the design of our model architecture is optimised to reliably detect keywords rather than to decode keyword phoneme sequences as arbitrary speech, while offering a parameter to adjust the operating point on the receiver operating characteristics curve. Experiments on the TIMIT corpus reveal that our graphical model outperforms a comparable hidden Markov model based keyword spotter that uses conventional garbage modelling. Martin Wöllmer, Florian Eyben, Björn W. Schuller, Gerhard Rigoll |
ASRU | 4 |
| 2009 | Multi-modal activity and dominance detection in smart meeting roomsabstractIn this paper a new approach for activity and dominance modeling in meetings is presented. For this purpose low level acoustic and visual features are extracted from audio and video capture devices. Hidden Markov Models (HMM) are used for the segmentation and classification of activity levels for each participant. Additionally, more semantic features are applied in a two-layer HMM approach. The experiments show that the acoustic feature is the most important one. The early fusion of acoustic and global-motion features achieves nearly as good results as the acoustic feature alone. All the other early fusion approaches are outperformed by the acoustic feature. More over, the two-layer model could not achieve the results of the acoustic features. Benedikt Hörnler, Gerhard Rigoll |
ICASSP | 2 |
| 2009 | Graphical Models: Statistical inference vs. determinationabstractUsing discrete Hidden-Markov-Models (HMMs) for recognition requires the quantization of the continuous feature vectors. In handwritten whiteboard note recognition it turns out that the pen-pressure information, which is important for recognition, is not adequately quantized and looses significance. In this paper, the implicit modeling of the pressure information presented in previous work which uses the deterministic knowledge on the actual pressure is generalized using a Graphical Model (GM) representation based on statistical inference. The results of two state-of-the-art toolboxes implementing HMMs and GMs are compared. It can be seen that the statistical inference approach based on GMs is inferior to the implicit modeling of the pressure information. It is shown that a direct implementation of HMMs outperforms the mathematic identical GM representation. Joachim Schenk, Benedikt Hörnler, Artur Braun, Gerhard Rigoll |
ICASSP | 4 |
| 2009 | Voronoi cell shaping for feature selection with discrete HMMsabstractIn this paper, we introduce a novel vector quantization (VQ) scheme for distributing the quantization error equally among the quantized dimensions. Afterwards, the proposed VQ scheme is used to perform feature selection in on-line handwritten whiteboard note recognition based on discrete Hidden-Markov-Models (HMMs). In an experimental section we show that the novel VQ scheme derives feature sets which contain less than 50% features, enabling recognition with better performance at less computational costs. Finally, the derived feature set is compared to the quantized features selected within a continuous HMM-based system: the features selected after quantization with the proposed VQ scheme are proved to perform significantly better than those in the continuous system. Joachim Schenk, Gerhard Rigoll |
ICASSP | 2 |
| 2009 | Robust discriminative keyword spotting for emotionally colored spontaneous speech using bidirectional LSTM networksabstractIn this paper we propose a new technique for robust keyword spotting that uses bidirectional long short-term memory (BLSTM) recurrent neural nets to incorporate contextual information in speech decoding. Our approach overcomes the drawbacks of generative HMM modeling by applying a discriminative learning procedure that non-linearly maps speech features into an abstract vector space. By incorporating the outputs of a BLSTM network into the speech features, it is able to make use of past and future context for phoneme predictions. The robustness of the approach is evaluated on a keyword spotting task using the HUMAINE sensitive artificial listener (SAL) database, which contains accented, spontaneous, and emotionally colored speech. The test is particularly stringent because the system is not trained on the SAL database, but only on the TIMIT corpus of read speech. We show that our method prevails over a discriminative keyword spotter without BLSTM-enhanced feature functions, which in turn has been proven to outperform HMM-based techniques. Martin Wöllmer, Florian Eyben, Joseph Keshet, Alex Graves, Björn W. Schuller, Gerhard Rigoll |
ICASSP | 6 |
| 2009 | GMs in On-Line Handwritten Whiteboard Note Recognition: The Influence of Implementation and ModelingabstractWe present a comparison of two state-of-the-art toolboxes for implementing Graphical Models (GMs), namely the HTK and the GMTK, and their use for discrete on-line handwritten whiteboard note recognition. We then motivate a GM that is capable of modeling the statistical dependencies between the pen’s pressure information and the remaining features after vector quantization. Since the number of variable parameters rises when more codebook entries are used for quantization, the proposed model outperforms standard HMMs for low numbers of codebook entries. Joachim Schenk, Benedikt Hörnler, Björn W. Schuller, Artur Braun, Gerhard Rigoll |
ICDAR | 5 |
| 2009 | Selecting Features in On-Line Handwritten Whiteboard Note Recognition: SFS or SFFS?abstractWhen selecting features with the sequential forward floating selection (SFFS), the "nesting effect" is avoided, which is a common phenomenon if the computationally less expensive sequential forward selection (SFS) is used instead. In this paper, we answer the key question, if the more complex and sophisticated SFFS should be used in on-line HMM-based recognition of handwritten whiteboard notes. In addition, an efficient method of displaying the selected feature set, the "feature map," is introduced.In an experimental section, both selection approaches are evaluated, the derived feature sets are compared, and a discussion on the selected features is given. Joachim Schenk, Moritz Kaiser, Gerhard Rigoll |
ICDAR | 3 |
| 2009 | "The Godfather" vs. "Chaos": Comparing Linguistic Analysis Based on On-line Knowledge Sources and Bags-of-N-Grams for Movie Review Valence EstimationabstractIn the fields of sentiment and emotion recognition, bag of words modeling has lately become popular for the estimation of valence in text. A typical application is the evaluation of reviews of e.g. movies, music, or games. In this respect we suggest the use of back-off N-Grams as basis for a vector space construction in order to combine advantages of word-order modeling and easy integration into potential acoustic feature vectors intended for spoken document retrieval. For a fine granular estimate we consider data-driven regression next to classification based on support vector machines. Alternatively the on-line knowledge sources ConceptNet, general inquirer, and WordNet not only serve to reduce out-of-vocabulary events, but also as basis for a purely linguistic analysis. As special benefit, this approach does not demand labeled training data. A large set of 100 k movie reviews of 20 years stemming from Metacritic is utilized throughout extensive parameter discussion and comparative evaluation effectively demonstrating efficiency of the proposed methods. Björn W. Schuller, Joachim Schenk, Gerhard Rigoll, Tobias Knaup |
ICDAR | 3 |
| 2009 | Learning weighted similarity measurements for unconstrained face recognitionabstractUnconstrained face recognition is the problem of deciding if an image pair is showing the same individual or not, without having class specific training material or knowing anything about the image conditions. In this paper, an approach of learning suited similarity measurements is introduced. For this the image is partitioned into several parts, to extract image region based histograms of gradients, local binary patterns and three patch local binary patterns. The similarities of respective patches are computed and it is learnt how to weight the different image regions. Finally, a fusion is applied using a multilayer perceptron. Evaluations are done on the ¿labeled faces in the wild¿ dataset. Andre Störmer, Gerhard Rigoll |
ICIP | 2 |
| 2009 | Boosting multi-modal camera selection with semantic featuresabstractIn this work semantic features are used to improve the results of the camera selection. These semantic features are group action, person action and person speaking. For this purpose low level acoustic and visual features are combined with high level semantic ones. After the feature fusion, a segmentation and classification are performed by Hidden Markov Models. The evaluation shows that an absolute improvement of 6.5% can be achieved. The frame error rate is reduced to 38.1% by using acoustic and all semantic features. The best model using only low level features achieves a frame error rate of 44.6%, which is the best one reported on this data set. Benedikt Hörnler, Dejan Arsic, Björn W. Schuller, Gerhard Rigoll |
ICME | 4 |
| 2009 | Novel VQ with constraints on the quantization error distributionabstractIn this paper, we motivate and introduce a novel vector quantization (VQ) scheme for distributing the quantization error among the quantized features of a continuous feature vector in a predefined manner. This is done by defining ratios between the individual quantization errors of the features and shaping the Voronoi cells accordingly. In a series of experiments we show that the novel approach is capable of either distributing the quantization error equally among the dimensions or realize an arbitrary distribution. Joachim Schenk, Frank Wallhoff, Gerhard Rigoll |
ICME | 3 |
| 2009 | Audio chord labeling by musiological modeling and beat-synchronizationabstractAutomatic labeling of chords in original audio recordings is challenging due to heavy acoustic overlay by melody and percussion sections, detuning and arpeggios that demand for a measure-grid to assign notes to chords. Further chord labeling benefits from contextual information. In this respect we suggest applying an HMM framework incorporating a musiological model trained on 16 k songs and synchronization with the measure grid by IIR comb-filter banks for tempo detection, meter recognition, and on-beat tracking. Features base on pitch-tuned chromatic information. Extensive evaluation on 11 k chords of 7 h of MP3 compressed popular music demonstrates effectiveness over traditional correlation analysis and single measure classification by support vector machines. Björn W. Schuller, Benedikt Hörnler, Dejan Arsic, Gerhard Rigoll |
ICME | 4 |
| 2009 | Multimodal data communication for human-robot interactionsabstractIn this paper, the development of a framework based on the Real-time Database (RTDB) for processing multimodal data is presented. This framework allows readily integration of input and output modules. Furthermore the asynchronous data streams from different sources can be approximately processed in a synchronous manner. Depending on the included modules, online as well as offline data processing is possible. The idea is to establish a real multimodal interaction system that is able to recognize and react to those situations that are relevant for human-robot interaction. Frank Wallhoff, Tobias Rehrl, Jürgen Gast, Alexander Bannat, Gerhard Rigoll |
ICME | 5 |
| 2009 | Recognising interest in conversational speech - comparing bag of frames and supra-segmental featuresabstractIt is common knowledge that affective and emotion-related states are acoustically well modelled on a supra-segmental level.Nonetheless successes are reported for frame-level processing either by means of dynamic classification or multi-instance learning techniques.In this work a quantitative feature-type-wise comparison between frame-level and supra-segmental analysis is carried out for the recognition of interest in human conversational speech.To shed light on the respective differences the same classifier, namely Support-Vector-Machines, is used in both cases: once by clustering a 'bag of frames' of unknown sequence length employing Multi-Instance Learning techniques, and once by statistical functional application for the projection of the time series onto a static feature vector.As database serves the Audiovisual Interest Corpus of naturalistic interest. Björn W. Schuller, Gerhard Rigoll |
INTERSPEECH | 2 |
| 2009 | Using graphical models for mixed-initiative dialog management systems with realtime PoliciesabstractIn this paper, we present a novel approach for dialog modeling, which extends the idea underlying the partially observable Markov Decision Processes (POMDPs), i. e. it allows for calculating the dialog policy in real-time and thereby increases the system flexibility. The use of statistical dialog models is particularly advantageous to react adequately to common errors of speech recognition systems. Comparing our results to the reference system (POMDP), we achieve a relative reduction of 31.6 % of the average dialog length. Furthermore, the proposed system shows a relative enhancement of 64.4 % of the sensitivity rate in the error recognition capabilities using the same specifity rate in both systems. The achieved results are based on the Air Travelling Information System with 21 650 user utterances in 1 585 natural spoken dialogs. Index Terms: dialog modeling, dialog strategy, spoken language understanding 1. Stefan Schwärzler, Stefan Maier, Joachim Schenk, Frank Wallhoff, Gerhard Rigoll |
INTERSPEECH | 5 |
| 2009 | Using Liquid Lenses to Extend the Operating Range of a Remote Gaze Tracking SystemabstractRemote eye tracking systems are widely used evaluation tools in many disciplines. Such systems have to deal with free head motions that cause a defocussing of the camera image. Hence, remote eye tracking systems usually have a limited operating range. This contribution presents a novel auto-focus setup for a small and mobile remote gaze tracking system that enables a clear eye image acquisition in a large operating range. Therefore, our setup uses a miniature liquid lens to achieve smallest possible focusing latencies and installation space. Using this technology, we extended the eye tracking system's operating range to distances from 0.6 m up to 1.3 m, with a mean accuracy of 0.45° in 4 exemplary working distances. Tony Poitschke, Stanislavs Bardins, Erwin Bay, Klaus Bartl, Florian Laquai, Johannes Vockeroth, Gerhard Rigoll, Erich Schneider |
SMC | 7 |
| 2009 | Applying Bayes Markov chains for the detection of ATM related scenariosabstractVideo surveillance systems have been introduced in various fields of our daily life to enhance security and protect individuals and sensitive infrastructure. Up to now it has been usually utilized as a forensic tool for after the fact investigations and are commonly monitored by human operators. In order to assist these and to be able to react in time, a fully automated system is desired. In this work we will present a multi camera surveillance system, which is required to resolve heavy occlusions, to detect robberies at ATM machines. The resulting trajectories will be analyzed for so called Low Level Activities (LLA), such as walking, running and stationarity, applying simple but robust approaches. The results of the LLA analysis will subsequently be fed into a Bayesian Network, that is used as a stochastic model to model so called High Level Activities (HLA). Introducing state transitions between HLAs will allow a temporal modeling of a complex scene. This can be represented by a Markovian process. Dejan Arsic, Atanas Lyutskanov, Moritz Kaiser, Björn W. Schuller, Gerhard Rigoll |
WACV | 5 |
| 2009 | Non-rigid registration of 3D facial surfaces with robust outlier detectionabstractNon-rigid registration of 3D facial surfaces is a crucial step in a variety of applications. Outliers, i.e., features in a facial surface that are not present in the reference face, often perturb the registration process. In this paper, we present a novel method which registers facial surfaces reliably also in the presence of huge outlier regions. A cost function incorporating several channels (red, green, blue, etc.) is proposed. The weight of each point of the facial surface in the cost function is controlled by a weight map, which is learned iteratively. Ideally, outliers will get a zero weight so that their disturbing effect is decreased. Results show that with an intelligent initialization the weight map improves the registration results considerably. Moritz Kaiser, Andre Störmer, Dejan Arsic, Gerhard Rigoll |
WACV | 4 |
| 2009 | A multidimensional dynamic time warping algorithm for efficient multimodal fusion of asynchronous data streams
Martin Wöllmer, Marc A. Al-Hames, Florian Eyben, Björn W. Schuller, Gerhard Rigoll |
Neurocomputing | 5 |
| 2009 | Being bored? Recognising natural interest by extensive audiovisual integration for real-life application
Björn W. Schuller, Ronald Müller, Florian Eyben, Jürgen Gast, Benedikt Hörnler, Martin Wöllmer, Gerhard Rigoll, Anja Höthker, Hitoshi Konosu |
Image Vis. Comput. | 7 |
| 2009 | Novel script line identification method for script normalization and feature extraction in on-line handwritten whiteboard note recognition
Joachim Schenk, Johannes Lenz, Gerhard Rigoll |
Pattern Recognit. | 3 |
| 2008 | In-car interaction using search-based user interfacesabstractIncreasing functionality, growing media volumes and dynamic data in today's in-vehicle information systems bear new challenges for user interaction design. Traditional hierarchical and menu-based interaction can only provide limited support while new search-based approaches are promising. In this work we assess different search techniques and search-based user interfaces. In particular we compare free search across all data items with categorized search. Our experiments with functional prototypes show that free search is more efficient and easier to use than searching within categories. Tests in a driving simulator show promising results regarding safety and workload. Means for alphanumeric input appear to be essential for an efficient and safe search interaction while driving. Stefan Graf, Wolfgang Spießl, Albrecht Schmidt 0001, Anneke Winter, Gerhard Rigoll |
CHI | 5 |
| 2008 | Contact-analog information representation in an automotive head-up displayabstractThis contribution presents an approach for representing contact-analog information in an automotive Head-Up Display (HUD). Therefore, we will firstly introduce our approach for the calibration of the optical system consisting of the virtual image plane of the HUD and the drivers eyes. Afterward, we will present the used eyetracking system for adaptation of the HUD content at the current viewpoint/position of the driver. We will also present first prototypical concepts for the visualization of contact-analog HUD content and inital test results from a brief usability study. Tony Poitschke, Markus Ablaßmeier, Gerhard Rigoll, Stanislavs Bardins, Stefan Kohlbecher, Erich Schneider |
ETRA | 3 |
| 2008 | Omni-directional multiperson tracking in meeting scenarios combining simulated annealing and particle filteringabstractThis proposal deals with the topic of tracking an unknown number of persons with a monocular camera in indoor environments. Within this context the main tracking system requirements are defined not only by a robust determination of all human trajectories but also by a reliable recovery of all object identities especially for challenging situations like heavy occlusion or the reentry of a person. Regarding all these needs a novel approach has been developed combining a probabilistic particle filter framework with an heuristic simulated annealing technique ported to the tracking domain. While the inter frame correspondence of objects, i.e. the assignment of identities, is handled by the simulated annealing approach, the particle filter architecture will be responsible for both the classification of an object to be a person as well as a stable tracking of the respective trajectory. An active shape model is utilized to create weights for the particles and thus serves as an object classifier. Our system has been evaluated on several video sequences showing meeting scenarios with a different number of participants. Quantitative numbers based on a tracking evaluation scheme show, that our system is capable of not only accurately determining the number of persons visible in each scene but also of precisely tracking each human and correctly assigning a label. Sascha Schreiber, Gerhard Rigoll |
FG | 2 |
| 2008 | Brute-forcing hierarchical functionals for paralinguistics: A waste of feature space?abstractWhile the " 'quasi-state-of-the-art'" towards acoustic emotion recognition relies on multivariate time-series analysis of e.g. pitch, energy, or MFCC by statistical functionals as moments or extrema, only few respect statistical noise by outliers due to too long segments as turns. Such noise can be overcome by hierarchical functionals as means of extrema over smaller units as words or chunks. Segmentation of such units however usually relies on transcription. We therefore discuss hierarchical functionals based on automatic segmentation and their systematic generation as opposed to common expert-driven selection. To cope with rapidly growing feature spaces iquest5k, we discuss data-driven two-stage compression based on SVM- SFFS. Extensive test-runs are carried out on two known emotion and behavior corpora, and show superiority of the suggested approach. Björn W. Schuller, Matthias Wimmer, Lorenz Mösenlechner, Christian Kern, Dejan Arsic, Gerhard Rigoll |
ICASSP | 6 |
| 2008 | Omnidirectional tracking and recognition of persons in planar viewsabstractIn this paper a view-independent head tracking system applying an Active Shape Model based particle filter is used to find precise image sections. DCTmod2 feature sequences are extracted from these sections and given as input to Cyclic Pseudo two-dimensional Hidden Markov Model based classifiers. These classifiers are trained to recognize the identity of the shown persons. The video material is recorded in an office environment with changing lighting conditions and thus results in a challenging task for both tracking and recognition. The overall performance of the system is evaluated depending on the various views of persons rotating on a swivel chair. Sascha Schreiber, Andre Störmer, Gerhard Rigoll |
ICIP | 3 |
| 2008 | A multi-step alignment scheme for face recognition in range imagesabstractFace recognition in range images is a challenging task, especially if the pose of the shown face is unknown. To solve this, an alignment procedure consisting of facial feature hypotheses extraction by invariant curvature features, PCA-based classification and Iterative Closest Point alignment will be introduced to create aligned and normalized patches. These patches will then be used in a recognition algorithm, a discrete Pseudo 2-Dimensional Hidden Markov Model approach based on vector quantized DCTmod2 features. The results of this processing chain are discussed and compared to previous works. Andre Störmer, Gerhard Rigoll |
ICIP | 2 |
| 2008 | Combining speech recognition and acoustic word emotion models for robust text-independent emotion recognitionabstractRecognition of emotion in speech usually uses acoustic models that ignore the spoken content. Likewise one general model per emotion is trained independent of the phonetic structure. Given sufficient data, this approach seemingly works well enough. Yet, this paper tries to answer the question whether acoustic emotion recognition strongly depends on phonetic content, and if models tailored for the spoken unit can lead to higher accuracies. We therefore investigate phoneme-, and word-models by use of a large prosodic, spectral, and voice quality feature space and Support Vector Machines (SVM). Experiments also take the necessity of ASR into account to select appropriate unit- models. Test-runs on the well-known EMO-DB database facing speaker-independence demonstrate superiority of word emotion models over today's common general models provided sufficient occurrences in the training corpus. Björn W. Schuller, Bogdan Vlasenko, Dejan Arsic, Gerhard Rigoll, Andreas Wendemuth |
ICME | 4 |
| 2008 | Neural net vector quantizers for discrete HMM-based on-line handwritten whiteboard-note recognitionabstractIn this work we evaluate a recently published vector quantization scheme, which has been developed to handle binary features like the pressure feature occurring in on-line handwriting recognition using discrete Hidden-Markov-Models (HMMs) with two neural net based vector quantizers (VQs). One of these uses a ldquoWinner-Take-Allrdquo (WTA) update rule and the other implements the ldquoNeural Gasrdquo (NG) approach. Both approaches are believed to be more efficient VQs than the standard k-means VQ used in our earlier publication. In an experimental section we prove that both the WTA and NG neural net VQ significantly (significance is measured by the one-sided t-test) outperform our previously used k-means VQ by rW= 0:9% and rN= 0:8%, respectively, referring to word-level accuracy. In addition, no significant difference in recognition accuracy between the WTA-VQ and the NG-VQ could be observed. Joachim Schenk, Gerhard Rigoll |
ICPR | 2 |
| 2008 | Edge-preserving unscented Kalman filter for speckle reductionabstractWe propose a recursive spatial-domain speckle reduction algorithm for synthetic aperture radar (SAR) imagery based on the unscented Kalman filter (UKF) with a discontinuity-adaptive Markov random field (DAMRF) prior. The capability of the UKF in handling speckle noise and the feature preservation ability of the DAMRF model are explored within a unified framework through importance sampling. Rama Krishna Sai S. Gorthi, A. N. Rajagopalan 0001, Rangarajan Aravind, Gerhard Rigoll |
ICPR | 4 |
| 2008 | Detection of security related affect and behaviour in passenger transportabstractSurveillance of drivers, pilots or passengers possesses significant potential for increased security within passenger transport. In an automotive setting the interaction can e.g. be improved by social awareness of an MMI. As further example security marshals can be efficiently positioned guided by according systems. Within this scope the detection of security relevant behavior patterns as aggressiveness or stress is discussed. The focus lies on real-life usage respecting online processing, subject independency, and noise robustness. The approach introduced employs multivariate time-series analysis for the synchronization and data reduction of audio and video by brute-force feature generation. By combined optimization of the large audiovisual space accuracy is boosted. Extensive results are reported on aviation behavior, as well as in particular for the audio channel on numerous standard corpora. The influence of noise will be discussed by representative carnoise overlay. Björn W. Schuller, Matthias Wimmer, Dejan Arsic, Tobias Moosmayr, Gerhard Rigoll |
INTERSPEECH | 5 |
| 2008 | Speech recognition in noisy environments using a switching linear dynamic model for feature enhancementabstractAll in-text\treferences\tunderlined\tin\tblue\tare\tlinked\tto\tpublications\ton\tResearchGate, letting you\taccess\tand\tread\tthem\timmediately. Björn W. Schuller, Martin Wöllmer, Tobias Moosmayr, Gerhard Rigoll |
INTERSPEECH | 4 |
| 2008 | Prosodic and spectral features within segment-based acoustic modelingabstractApart from the usually employed MFCC, PLP, and energy feature information, also duration, low order formants, pitch, and center-of-gravity-based features are known to carry valuable information for phoneme recognition. This work investigates their individual performance within segment-based acoustic modeling. Also, experiments optimizing a feature space spanned by this set, exclusively, are reported, using CFSS feature space optimization and speaker adaptation. All tests are carried out with SVM on the open IFA-corpus of 47 Dutch handlabeled phonemes with a total of 178k instances. Extensive speaker dependent vs. independent test-runs are discussed as well as four different speaking styles reaching from informal to formal: informal and retold story telling, and read aloud with fixed and variable content. Results show the potential of these rather uncommon features, as e.g. based on F3 or pitch. Index Terms: phoneme-recognition, prosodic features, acoustic modeling, feature space optimization, ASR Björn W. Schuller, Gerhard Rigoll |
INTERSPEECH | 3 |
| 2008 | Combining statistical and syntactical systems for spoken language understanding with graphical modelsabstractThere are two basic approaches for semantic processing in spoken language understanding: a rule based approach and a statistic approach. In this paper we combine both of them in a novel way by using statistical and syntactical dynamic bayesian networks (DBNs) together with Graphical Models (GMs) for spoken language understanding (SLU). GMs merge in a complex, mathematical way probability with graph theory. This results in four different setups which raise in their complexity. Comparing our results to a baseline system we achieve a F1-measure of 93.7 % in word classes and 95.7 % in concepts for our best setup in the ATIS-Task. This outperforms the baseline system relatively by 3.7 % in word classes and by 8.2 % in concepts. The expermiments were performend with the graphical model toolkit (GMTK). Index Terms: natural language understanding, machine learning, graphical models Stefan Schwärzler, Jürgen T. Geiger, Joachim Schenk, Marc A. Al-Hames, Benedikt Hörnler, Günther Ruske, Gerhard Rigoll |
INTERSPEECH | 7 |
| 2008 | Balancing spoken content adaptation and unit length in the recognition of emotion and interestabstractRecognition and detection of non-lexical or paralinguistic cues from speech usually uses one general model per event (emotional state, level of interest).Commonly this model is trained independent of the phonetic structure.Given sufficient data, this approach seemingly works well enough.Yet, this paper addresses the question on which phonetic level there is the onset of emotions and level of interest.We therefore compare phoneme-, word-and sentence-level analysis for emotional sentence classification by use of a large prosodic, spectral, and voice quality feature space for SVM and MFCC for HMM/GMM.Experiments also take the necessity of ASR into account to select appropriate unit-models.In experiments on the well-known public EMO-DB database, and the SUSAS and AVIC spontaneous interest corpora, we found that the emotion recognition by sentence level analysis shows the best results.We discuss the implications of these types of analysis on the design of robust emotion and interest recognition of usable human-machine interfaces (HMI). Bogdan Vlasenko, Björn W. Schuller, Kinfe Tadesse Mengistu, Gerhard Rigoll, Andreas Wendemuth |
INTERSPEECH | 4 |
| 2008 | Emotion sensitive speech control for human-robot interaction in minimal invasive surgeryabstractMinimal invasive surgery demands for utmost precise and reliable camera control to prevent any harm to the patient during operations. We therefore introduce a robot-driven camera that can be controlled either manually by a joystick, or by speech to ensure free hands and feet, and reduced cognitive workload of the surgeon. Speech control is chosen as simple, yet highly robust command and control application. However, due to high stress, and partially fatigue, emotional factors can play a life decisive role in the operational situation. As any misunderstanding of the surgeonpsilas intent can easily lead to patient injuries by mis-movement of the camera, emotional factors are integrated in the human-robot interaction. In this work we therefore discuss the recording of a 3,035 turns database of spontaneous emotional speech in real life surgical operations. Known to be a challenge, we employ a high dimensional acoustic feature space, and subset optimization for recognition of positive versus negative emotion for interaction adaptation, surgeon self-monitoring, and potential adaptation of acoustic models within speech recognition. Promising 75.5% mean accuracy can be reported in a cross-operation recognition task given the severe condition of usage in real medical operations. Björn W. Schuller, Gerhard Rigoll, Salman Can, Hubertus Feußner |
RO-MAN | 2 |
| 2008 | How infrared tracking increases the realism of multi-person videoconferencing in collaborative Augmented RealityabstractInterpersonal communication is based on acoustical and visual information exchange. Videoconferencing systems allow the communication between people separated by large distances by recording acoustical as well as visual information. In this article, we present a multi person videoconferencing system integrated in an Augmented Reality framework. Existing Augmented Reality applications already support collaborative efforts of users on different locations. The presented VCS combines these collaborative Augmented Reality applications with the functionality of a videoconferencing system at the same time and, thus, enables the users to interact with virtual content at the same time and communicate with each other in one common working space. The system was evaluated with different values for three characteristics (displaying technology, tracking system and image processing) with a focus on the rating of realism and presence of the human counterpart and the quality of the system. The results of this evaluation suggest that a system using an head-mounted display for visualization, an infrared tracking system and active image processing is rated as the most realistic visualization and evokes the highest feeling of physical attendance. This enhancement however comes with a reduction of the subjective perception of quality. Stefan Reifinger, Florian Laquai, Gerhard Rigoll |
SMC | 3 |
| 2008 | Translation and rotation of virtual objects in Augmented Reality: A comparison of interaction devicesabstractThis paper describes an evaluation, which compares three interfaces for translating and rotating virtual objects in an Augmented Reality environment. We used a mouse/keyboard based interface, an infrared-tracking based gesture recognition system, and our tangible user interface, which we developed for this evaluation. Our tangible user interface consists of two accelerometers, one gyroscope, and four buttons integrated in a cuboid casing. The evaluation focused on a comparison of immersion, intuitiveness, mental and physical workload, and task execution times. The test persons had to solve a given task with all interfaces (translating an object, rotating an object, and a combination of translation and rotation). The results showed that a translation of a virtual object takes nearly the same amount of time with all interfaces, but much more than a similar manipulation in a real environment. Rotation is performed as the slowest using a mouse/keyboard based interface. For a combination of translation and rotation, gesture recognition based interaction turned out to be the fastest way of interaction. All kinds of manipulation showed, that gesture recognition provides the most immersive and intuitive interaction with the lowest mental workload. But these benefits are limited by the highest rating of physical workload, caused by additional hardware mounted at the user's hand, as well as the fact, that the user has to hold his arm upwards during the task execution. Stefan Reifinger, Florian Laquai, Gerhard Rigoll |
SMC | 3 |
| 2007 | On the Necessity and Feasibility of Detecting a Driver's Emotional State While Driving
Michael Grimm, Kristian Kroschel, Helen Harris, Clifford Nass, Björn W. Schuller, Gerhard Rigoll, Tobias Moosmayr |
ACII | 6 |
| 2007 | Frame vs. Turn-Level: Emotion Recognition from Speech Considering Static and Dynamic Processing
Bogdan Vlasenko, Björn W. Schuller, Andreas Wendemuth, Gerhard Rigoll |
ACII | 4 |
| 2007 | Comparing one and two-stage acoustic modeling in the recognition of emotion in speechabstractIn the search for a standard unit for use in recognition of emotion in speech, a whole turn, that is the full section of speech by one person in a conversation, is common. Within applications such turns often seem favorable. Yet, high effectiveness of sub-turn entities is known. In this respect a two-stage approach is investigated to provide higher temporal resolution by chunking of speech-turns according to acoustic properties, and multi-instance learning for turn-mapping after individual chunk analysis. For chunking fast pre-segmentation into emotionally quasi-stationary segments by one-pass Viterbi beam search with token passing basing on MFCC is used. Chunk analysis is realized by brute-force large feature space construction with subsequent subset selection, SVM classification, and speaker normalization. Extensive tests reveal differences compared to one-stage processing. Alternatively, syllables are used for chunking. Björn W. Schuller, Bogdan Vlasenko, Ricardo Minguez, Gerhard Rigoll, Andreas Wendemuth |
ASRU | 4 |
| 2007 | Audiovisual Behavior Modeling by Combined Feature SpacesabstractGreat interest is recently shown in behavior modeling, especially in public surveillance tasks. In general it is agreed upon the benefits of use of several input cues as audio and video. Yet, synchronization and fusion of these information sources remains the main challenge. We therefore show results for a feature space combination, which allows for overall feature space optimization. Audio and video features are thereby firstly derived as low-level-descriptors. Synchronization and feature combination is achieved by multivariate time-series analysis. Test-runs on a database of aggressive, cheerful, intoxicated, nervous, neutral, and tired behavior in an airplane situation show a significant improvement over each single modality. Björn W. Schuller, Dejan Arsic, Gerhard Rigoll, Matthias Wimmer, Bernd Radig |
ICASSP (2) | 3 |
| 2007 | Fast and Robust Meter and Tempo Recognition for the Automatic Discrimination of Ballroom Dance StylesabstractFast and robust recognition of a song's meter, and quarter note tempo is crucial in many music information retrieval tasks dealing especially with large databases or real-time musical stream processing. We therefore introduce a novel approach that is capable of extracting musical meter features and tempo in beats per minute. The method is extendable in order to return the locations of beat onsets suitable for example for beat synchronization or musical audio segmentation. We use a simplified psychoacoustic model to split the input into audible frequency bands and two phase comb filtering on those bands to find the quarter note tempo and metrical structure. Based on these features we discriminate the nine classic ballroom dance styles and duple or triple meter by support-vector-machines as exemplary application. Test-runs are carried out on a public ballroom dance music database containing 1.8 k titles and the public MTV-Europe Most Wanted 1981-2000 to demonstrate the high effectiveness for popular music with respect to meter, tempo and ballroom dance style recognition. Björn W. Schuller, Florian Eyben, Gerhard Rigoll |
ICASSP (1) | 3 |
| 2007 | Robust Multi-Modal Group Action Recognition in Meetings from Disturbed Videos with the Asynchronous Hidden Markov ModelabstractThe asynchronous hidden Markov model (AHMM) models the joint likelihood of two observation sequences, even if the streams are not synchronised. We explain this concept and how the model is trained by the EM algorithm. We then show how the AHMM can be applied to the analysis of group action events in meetings from both clear and disturbed data. The AHMM outperforms an early fusion HMM by 5.7% recognition rate (a rel. error reduction of 38.5%) for clear data. For occluded data, the improvement is in average 6.5% recognition rate (rel. error red. 40%). Thus asynchronity is a dominant factor in meeting analysis, even if the data is disturbed. The AHMM exploits this and is therefore much more robust against disturbances. Marc A. Al-Hames, Claus Lenz, Stephan Reiter, Joachim Schenk, Frank Wallhoff, Gerhard Rigoll |
ICIP (2) | 6 |
| 2007 | Improved Image Segmentation using Photonic Mixer DevicesabstractAiming at improving image segmentation and extending regular computer vision algorithms an image sensor, acquiring additional depth information is deployed. The novel Photonic Mixer Device (PMD) technology will be summarized, by which it becomes possible to measure the observed object's distance to the camera. Motivated by expanding and making existing image processing and segmentation algorithms more robust, the additional depth information is overlapped with the image map gained by a regular camera system. Due to the fact of the used cameras are displaced and have different image resolutions, a sophisticated calibration algorithm will be introduced. By applying improved algorithms to applications, such as people counting, gesture recognition, and barrier detection its effectiveness will be shown. Frank Wallhoff, Martin Russ, Gerhard Rigoll, Johann Göbel, Hermann Diehl |
ICIP (6) | 3 |
| 2007 | Eye Gaze Studies Comparing Head-Up and Head-Down Displays in VehiclesabstractTo minimize the mental workload for the driver and to keep the increasing amount of information easily accessible, sophisticated display and interaction techniques are essential. This contribution focuses on a user-centered analysis for an authoritative grading of head-up displays (HUDs) in cars. Two studies delivered the evaluation data. In a field test, the potential and the usability of the HUD were analyzed. For special driving situations the according display needs and requirements of the users have been identified and compared with in-car displays, so-called head-down displays (HDD). As major result, a high acceptance of the HUD by the driver and a good performance compared to other in-car displays had been reached. Markus Ablaßmeier, Tony Poitschke, Frank Wallhoff, Klaus Bengler, Gerhard Rigoll |
ICME | 5 |
| 2007 | Automatic Multi-Modal Meeting Camera Selection for Video-Conferences and Meeting BrowsersabstractIn a video-conference the participants usually see the video of the speaker. However if somebody reacts (e. g. nodding) the system should switch to his video. Current systems do not support this. We formulate this camera selection as a pattern recognition problem. Then we apply HMMs to learn this behaviour. Thus our system can easily be adapted to different meeting scenarios. Furthermore, while current systems stay on the speaker, our system will switch if somebody reacts. In an experimental section we show that -compared to a desired output -a current system shows the wrong camera more than half of the time (frame error rate 53%), where our system selects the wrong camera in only a quarter of the time (FER 27%). Marc A. Al-Hames, Benedikt Hörnler, Ronald Müller, Joachim Schenk, Gerhard Rigoll |
ICME | 5 |
| 2007 | Suspicious Behavior Detection in Public Transport by Fusion of Low-Level Video DescriptorsabstractRecently great interest has been shown in the visual surveillance of public transportation systems. The challenge is the automated analysis of passenger's behaviors with a set of visual low-level features, which can be extracted robustly. On a set of global motion features computed in different parts of the image, here the complete image, the face and skin color regions, a classification with Support Vector Machines is performed. Test-runs on a database of aggressive, cheerful, intoxicated, nervous, neutral and tired behavior in an airplane situation show promising results. Dejan Arsic, Björn W. Schuller, Gerhard Rigoll |
ICME | 3 |
| 2007 | A Framework for Modular Signal Processing Systems with High-Performance RequirementsabstractThis paper introduces the software frameworkMMER Labwhich allows an effective assembly of modular signal processing systems optimized for memory efficiency and performance. Our C/C++ framework is designed to constitute the basis of a well organized and simplified development process in industrial and academic research teams. It supports the structuring of modular systems by provision of basic data-, parameter-, and command-interfaces, ensuring the re-usability of the system components. Due to the underlying multi-threading capabilities, the applications built inMMER Labare enabled to fully exploit the increasing computational power of multi-core CPU architectures. This feature is carried out by a buffering concept which controls the data flow between the connected modules and allows for the parallel processing of consecutive signal segments (e.g. video frames). We introduce the concept of the multi-threading environment and the data flow architecture with its comfortable programming interface. We illustrate the proposed module concept for the generic assembly of processing chains and show applications from the area of video analysis and pattern Lukas L. Diduch, Ronald Müller, Gerhard Rigoll |
ICME | 3 |
| 2007 | Wearable Assistance for the Ballroom-Dance Hobbyist - Holistic Rhythm Analysis and Dance-Style ClassificationabstractAutomated retrieval of high level information from ballroom dance music is challenging, but has many practical applications. These include, for example, a fully automatic ballroom dance D.J., robots capable of performing ballroom dances, or wearable dance-assistance, as considered herein. It is necessary, for such a system, to retrieve information about the song's quarter note tempo, meter and beat positions. Further, the system must be able to discriminate between the nine standard and Latin ballroom dances. In this paper we present a model that combines all these requirements in one holistic approach. The polyphonic input is processed by a simplified psychoacoustic model. Tatum, tempo and meter features are extracted using resonant filters. The filter output is used for beat tracking. The extracted features are used for a ballroom dance-style classification by support-vector-machines. To show the high effectiveness regarding dance-style recognition and beat tracking, test-runs are carried out on a database containing 1.8k titles. Florian Eyben, Björn W. Schuller, Stephan Reiter, Gerhard Rigoll |
ICME | 4 |
| 2007 | Hidden Conditional Random Fields for Meeting SegmentationabstractAutomatic segmentation and classification of recorded meetings provides a basis towards understanding the content of a meeting. It enables effective browsing and querying in a meeting archive. Though robustness of existing approaches is often not reliable enough. We therefore strive to improve on this task by applying conditional random fields augmented by hidden states. These hidden conditional random fields have been proven to be efficient in low level pattern recognition tasks. Now we propose to use these novel models to segment a pre-recorded meeting into meeting events. Since they can also be seen as an extension to hidden Markov models an elaborate comparison of the two approaches is provided. Extensive test runs on the public M4 Scripted Meeting Corpus prove the great performance of applying our suggested novel approach compared to other similar methods. Stephan Reiter, Björn W. Schuller, Gerhard Rigoll |
ICME | 3 |
| 2007 | Adaptive Human-Machine Interfaces in Cognitive Production EnvironmentsabstractThis article presents an integrated framework for multi-modal adaptive cognitive technical systems to guide, assist and observe human workers in complex manual assembly environments. The demand for highly flexible construction facilities obviously contradicts longer training and preparation phases of human workers. By giving context-aware building instructions over retina displays, text-to-speech commands or acoustical signals, a non-specialized industrial stand-by men in a production task can precisely be alloted to execute the next processing step without any previous knowledge. Using non-invasive gesture recognizers and object detectors the human worker can be observed in order to track the production line and initiate the subsequent step in the interaction loop. Aiming at testing and evaluating the desired human-machine interfaces and its capabilities a virtual working place together with a concrete use case is introduced. Frank Wallhoff, Markus Ablaßmeier, Alexander Bannat, Stephan Buchta, A. Rauschert, Gerhard Rigoll, Mathey Wiesbeck |
ICME | 6 |
| 2007 | Surveillance and Activity Recognition with Depth InformationabstractIn the present treatise an image sensor acquiring additional depth information is applied to extend regular computer vision algorithms. The so called Photonic Mixer Device (PMD) basing on the time-of-flight principle can measure the distance between a smart pixel on the image sensor and the object being recorded. Since the resolution of such a sensor is rather low, they are not intended to replace existing image acquisition technologies, i.e. CMOS or CCD cameras, but can rather be employed to assist and speed up several tasks, especially image segmentation or object detection tasks. Frank Wallhoff, Martin Russ, Gerhard Rigoll, Johann Göbel, Hermann Diehl |
ICME | 3 |
| 2007 | Audiovisual recognition of spontaneous interest within conversationsabstractIn this work we present an audiovisual approach to the recognition of spontaneous interest in human conversations. For a most robust estimate, information from four sources is combined by a synergistic and individual failure tolerant fusion. Firstly, speech is analyzed with respect to acoustic properties based on a high-dimensional prosodic, articulatory, and voice quality feature space plus the linguistic analysis of spoken content by LVCSR and bag-of-words vector space modeling including non-verbals. Secondly, visual analysis provides patterns of the facial expression by AAMs, and of the movement activity by eye tracking. Experiments base on a database of 10.5h of spontaneous human-to-human conversation containing 20 subjects in gender and age-class balance. Recordings are fulfilled with a room microphone, camera, and headsets for close-talk to consider diverse comfort and noise conditions. Three levels of interest were annotated within a rich transcription. We describe each information stream and a fusion on an early level in detail. Our experiments aim at a person-independent system for real-life usage and show the high potential of such a multimodal approach. Benchmark results based on transcription versus automatic processing are also provided. Björn W. Schuller, Ronald Müller, Benedikt Hörnler, Anja Höthker, Hitoshi Konosu, Gerhard Rigoll |
ICMI | 6 |
| 2007 | Combining frame and turn-level information for robust recognition of emotions within speechabstractCurrent approaches to the recognition of emotion within speech usually use statistic feature information obtained by application of functionals on turn- or chunk levels. Yet, it is well known that thereby important information on temporal sub-layers as the frame-level is lost. We therefore investigate the benefits of integration of such information within turn-level feature space. For frame-level analysis we use GMM for classification and 39 MFCC and energy features with CMS. In a subsequent step output scores are fed forward into a 1.4k large-feature-space turn-level SVM emotion recognition engine. Thereby we use a variety of Low-Level-Descriptors and functionals to cover prosodic, speech quality, and articulatory aspects. Extensive testruns are carried out on the public databases EMO-DB and SUSAS. Speaker-independent analysis is faced by speaker normalization. Overall results highly emphasize the benefits of feature integration on diverse time scales. Bogdan Vlasenko, Björn W. Schuller, Andreas Wendemuth, Gerhard Rigoll |
INTERSPEECH | 4 |
| 2007 | Context-aware kitchen utilitiesabstractWe report on approaches for context-awareness in a kitchen environment. Two devices, an augmented cutting board and a sensor-enriched knife, enable the environment to determine the type of food handled during the preparation of meals. Matthias Kranz, Albrecht Schmidt 0001, Alexis Maldonado, Radu Bogdan Rusu, Michael Beetz, Benedikt Hörnler, Gerhard Rigoll |
TEI | 7 |
| 2006 | Reduced Complexity and Scaling for Asynchronous HMMS in a Bimodal Input Fusion ApplicationabstractThe asynchronous hidden Markov model (AHMM) can model the joint likelihood of two observation sequences, even if the streams are not synchronised. Previously this model has been applied to audio-visual recognition tasks. The main drawback of the concept is its rather high training and decoding complexity. In this work we show how the complexity can be reduced significantly with advanced running indices for the calculations. Yet, the AHMM characteristics and its advantages are preserved. The improvement also allows a scaling procedure to keep numerical values in a reasonable range. In an experimental section we compare the complexity of the original and the improved concept and validate the theoretical results. Then the model is tested on a bimodal speech and gesture user input fusion task: compared to a late fusion HMM an improvement of more than 10% absolute recognition performance has been achieved Marc A. Al-Hames, Gerhard Rigoll |
ICASSP (5) | 2 |
| 2006 | A Combined LSTM-RNN - HMM - Approach for Meeting Event Segmentation and RecognitionabstractAutomatic segmentation and classification of recorded meetings provides a basis that enables effective browsing and querying in a meeting archive. Yet, robustness of today's approaches is often not reliable enough. We therefore strive to improve on this task by introduction of a tandem approach combining the discriminative abilities of recurrent neural nets and warping capabilities of hidden Markov models. Thereby long short-term memory cells are used for audio-visual frame analysis within the neural net. These help to overcome typical long time lags. Extensive test runs on the public M4 Scripted Meeting Corpus show great performance applying our suggested novel approach Stephan Reiter, Björn W. Schuller, Gerhard Rigoll |
ICASSP (2) | 3 |
| 2006 | Submotions for Hidden Markov Model Based Dynamic Facial Action RecognitionabstractVideo based analysis of a persons' mood or behavior is in general performed by interpreting various features observed on the body. Facial actions, such as speaking, yawning or laughing are considered as key features. Dynamic changes within the face can be modeled with the well known Hidden Markov Models (HMM). Unfortunately even within one class examples can show a high variance because of unknown start and end state or the length of a facial action. In this work we therefore perform a decomposition of those into so called submotions. These can be robustly recognized with HMMs, applying selected points in the face and their geometrical distances. Additionally the first and second derivation of the distances is included. A sequence of submotions is then interpreted with a dictionary and dynamic programming, as the order may be crucial. Analyzing the frequency of sequences shows the relevance of the submotions order. In an experimental section we show, that our novel submotion approach outperforms a standard HMM with the same set of features by nearly 30% absolute recognition rate. Dejan Arsic, Joachim Schenk, Björn W. Schuller, Frank Wallhoff, Gerhard Rigoll |
ICIP | 5 |
| 2006 | A Hierarchical ASM/AAM Approach in a Stochastic Framework for Fully Automatic Tracking and RecognitionabstractThis paper deals with the fully automatic extraction of classifiable person features out of a video stream with challenging background. Basically the task can be split in two parts: Tracking the object and extracting distinctive features. In order to track a person, a system composed of an active shape model embedded in a particle filter framework has been built. The output-a shape representing the position and the geometry of the human's head-serves as an initial guess for the following active appearance model, which enables high precision matching of the head's texture. In this way raw features are transformed into appearance parameters, which finally can be used for a variety of classification tasks. The novelty of this framework is the hierarchical combination using the similarities of the models as well as exploiting their differences to enhance robustness and performance in complex scenarios. Sascha Schreiber, Andre Störmer, Gerhard Rigoll |
ICIP | 3 |
| 2006 | A Two-Layer Graphical Model for Combined Video Shot and Scene Boundary DetectionabstractIn this work we present a novel two-layer hybrid graphical model for combined shot and scene boundary detection in videos. In the first layer of the model, low-level features are used to detect shot boundaries. The shot layer is connected to a higher layer that detects scene or chapter boundaries from semantic features. With this structure, the model optimises the alignment for both layers at the same time and the detection results are interconnected. Experimental results on real video data show, that both layers highly benefit from this sharing of information. Compared to a baseline threshold method with the same features, the F-measure result for the shot detection has been improved by 12.6% absolute. For the scene boundary detection, the result has been improved by more than 11% absolute Marc A. Al-Hames, Stefan Zettl, Frank Wallhoff, Stephan Reiter, Björn W. Schuller, Gerhard Rigoll |
ICME | 6 |
| 2006 | Segmentation and Recognition of Meeting Events using a Two-Layered HMM and a Combined MLP-HMM ApproachabstractAutomatic segmentation and classification of recorded meetings provides a basis that enables effective browsing and querying in a meeting archive. Yet, robustness of today approaches is often not reliable enough. We therefore strive to improve on this task by introduction of a hybrid approach combining the discriminative abilities of artificial neural nets and warping capabilities of hidden Markov models. Dividing the task into two layers and defining a proper set of individual actions helps to cope with the problem of lack of data and overcomes conventional single-layered approaches. Extensive test runs on the public M4 Scripted Meeting Corpus prove the great performance gain applying our suggested novel approach compared to other similar methods Stephan Reiter, Björn W. Schuller, Gerhard Rigoll |
ICME | 3 |
| 2006 | Evolutionary Feature Generation in Speech Emotion RecognitionabstractFeature sets are broadly discussed within speech emotion recognition by acoustic analysis. While popular filter and wrapper based search help to retrieve relevant ones, we feel that automatic generation of such allows for more flexibility throughout search. The basis is formed by dynamic low-level descriptors considering intonation, intensity, formants, spectral information and others. Next, systematic derivation of prosodic, articulatory, and voice quality high level functionals is performed by descriptive statistical analysis. From here on feature alterations are automatically fulfilled, to find an optimal representation within feature space in view of a target classifier. To avoid NP-hard exhaustive search, we suggest use of evolutionary programming. Significant overall performance improvement over former works can be reported on two public databases Björn W. Schuller, Stephan Reiter, Gerhard Rigoll |
ICME | 3 |
| 2006 | Musical Signal Type Discrimination based on Large Open Feature SetsabstractAutomatic discrimination of musical signal types as speech, singing, music, genres or drumbeats within audio streams is of great importance, e.g. for radio broadcast stream segmentation. Yet, feature sets are largely discussed. We therefore suggest a large open feature set approach starting with systematical generation of 7k hi-level features based on MPEG-7 low-level-descriptors and further feature contours. A subsequent fast gain ratio reduction followed by wrapper-based floating search leads to a strong basis of relevant features. Next, features are added by alteration and combination within genetic search. For classification we use support-vector-machines proven reliable for this task. Test-runs are carried out on two task-specific databases and the public Columbia SMD database and show significant improvements for each step of the suggested novel concept Björn W. Schuller, Frank Wallhoff, Dejan Arsic, Gerhard Rigoll |
ICME | 4 |
| 2006 | Efficient Recognition of Authentic Dynamic Facial Expressions on the Feedtum DatabaseabstractIn order to allow for fast recognition of a user's affective state we discuss innovative holistic and self organizing approaches for efficient facial expression analysis. The feature set is thereby formed by global descriptors and MPEG based DCT coefficients. In view of subsequent classification we compare modeling by pseudo multidimensional hidden Markov models and support vector machines. Within the latter case super-vectors are constructed based on sequential floating search methods. Extensive test-runs as a proof of concept are carried out on our publicly available FEEDTUM database consisting of elicited spontaneous emotions of 18 subjects within the MPEG-4 emotion-set plus added neutrality. Maximum recognition performance reaches the benchmark-rate gained by a human perception test with 20 test-persons and manifest the effectiveness of the introduced novel concepts Frank Wallhoff, Björn W. Schuller, Michael Hawellek, Gerhard Rigoll |
ICME | 4 |
| 2006 | Recognition of interest in human conversational speechabstractRecognition of interest of a speaker within a human dialog bears great potential in many commercial applications.Within this work we therefore introduce an approach that analyses acoustic and linguistic cues of a spoken utterance.A systematic generation of more than 5k hi-level features basing on prosodic and spectral feature contours by means of descriptive statistical analysis and subsequent feature space optimization is used to find relevant acoustic attributes.For linguistic information integration a bag-of-words representation is used relying on a speech recognizer's output.One main aspect is the database of more than 2k spontaneous sub-speaker turns recorded and annotated for this analysis.Several influence factors as microphone distance and ASR versus annotation of spoken content are discussed.Overall remarkable performance of a running prototype can be reported discriminating between three levels of interest. Björn W. Schuller, Niels Köhler, Ronald Müller, Gerhard Rigoll |
INTERSPEECH | 4 |
| 2006 | Timing levels in segment-based speech emotion recognition
Björn W. Schuller, Gerhard Rigoll |
INTERSPEECH | 2 |
| 2006 | Hybrid NN/HMM acoustic modeling techniques for distributed speech recognition
Jan Stadermann, Gerhard Rigoll |
Speech Commun. | 2 |
| 2005 | Bimodal fusion of emotional data in an automotive environmentabstractWe present a flexible bimodal approach to person dependent emotion recognition in an automotive environment by adapting an acoustic and a visual monomodal recognizer and combining the individual results on an abstract decision level. The reference database consists of 840 acted audiovisual examples of seven different speakers, expressing the three emotions, positive (joy), negative (anger, irritation) and neutral. Concerning the acoustic module, we calculate the statistics of commonly known low-level features. Facial expressions are evaluated by an SVM classification of Gabor-filtered face regions. At the subsequent integration stage, both monomodal decisions are fused by a weighted linear combination. An evaluation of the recorded examples yields an average recognition rate of 90.7% for the fusion approach. This adds up to a performance gain of nearly 4% compared to the best monomodal recognizer. The system is currently used to improve the usability for automotive infotainment interfaces. Stefan Hoch, Frank Althoff, Gregor McGlaun, Gerhard Rigoll |
ICASSP (2) | 4 |
| 2005 | Multimodal Meeting Analysis by Segmentation and Classification of Meeting Events based on a Higher Level Semantic ApproachabstractThis paper encompasses the analysis of meetings for segmentation into sub-genres. Therefore, an approach on a higher semantic level has been chosen. The algorithms make use of the results of specialized recognizers like a speaker turn detector and a gesture recognizer. Basically, the goal of this investigation was to answer the question, how well meeting analysis is possible if only the results of these recognizers are available. After introducing briefly the basics of these recognizers, two slightly different methods for the segmentation are presented. The results show the potential of the used methods to find the segment boundaries and to categorize the detected segments into sub-genres (also called meeting events or group actions). Based on this segmentation, further analysis regarding topic detection and content extraction can be accomplished. Stephan Reiter, Sascha Schreiber, Gerhard Rigoll |
ICASSP (2) | 3 |
| 2005 | Meta-Classifiers in Acoustic and Linguistic Feature Fusion-Based Affect RecognitionabstractWe suggest a novel approach to affect recognition based on acoustic and linguistic analysis of spoken utterances. In order to achieve maximum discrimination power within robust integration of these information sources, a fusion on the feature level is introduced. Considering classification, we use meta-classifiers, such as StackingC and Boosting, for a stabilized performance, and a combination of classifiers within ensembles. Extensive comparisons of diverse base-classifiers, including support vector machines, neural networks, stochastic models, and decision trees, are fulfilled. 381 acoustic features are extracted and their relevance is calculated by a sequential forward floating search in comparison to reduction by principal component analysis. Several variants for linguistic feature calculation are described and ranked, including bunch-of-words, n-grams, salience, and mutual information. Furthermore, reduction by stopping and stemming or filter-based selection methods is evaluated, reducing 2,334 linguistic features. Seven discrete emotions described in the MPEG-4 standard are recognized within an existing recognition engine. The presented results are based on two large databases of 4,336 acted and real emotion samples from movies, chat and car interaction dialogues. A significant gain and an outstanding overall performance are observed by this novel fusion and use of ensembles. Björn W. Schuller, Raquel Jiménez Villar, Gerhard Rigoll, Manfred K. Lang |
ICASSP (1) | 3 |
| 2005 | Two-Stage Speaker Adaptation of Hybrid Tied-Posterior Acoustic ModelsabstractIn this paper, strategies are explored to adapt hybrid neural network/HMM systems based on the tied-posterior paradigm. We investigate the retraining of selected important parts of the neural network and a gradient based adaptation strategy for the HMM mixture coefficients based on maximizing the scaled likelihood. The paper presents the following innovations: first, it introduces one of the first adaptation methods for hybrid systems where the HMM component contributes significantly to the adaptation success; second, it presents a novel approach to the neural network's adaptation, based on the selection of suitable neurons for adaptation. Results on the WSJ speaker adaptation test show the capability of our methods to adapt to new speakers, especially in the case of adapting the neural net, and that both methods can be combined to achieve additional improvement of the word error rate in most cases. Jan Stadermann, Gerhard Rigoll |
ICASSP (1) | 2 |
| 2005 | A multi-modal graphical model for robust recognition of group actions in meetings from disturbed videosabstractIn this work we present a novel multi-modal mixed-state dynamic Bayesian network (DBN) for robust meeting event classification from disturbed videos. The model uses information from the audio and the visual channel to structure meetings into segments. Within the DBN a multi-stream hidden Markov model (HMM) is coupled with a linear dynamical system (LDS) to compensate disturbances in the visual channel. Thereby the HMM is used as driving input for the LDS. Thus the model can handle noise and occlusions in the video. Experimental results on real meeting data show that the new model is highly preferable to all single-stream approaches. Compared to a baseline multi-modal early fusion HMM, the new DBN is 3.5%, respectively up to 6.1% better for clear and visual disturbed data, this corresponds to a relative error reduction of 23.6%, respectively 29.9%. Marc A. Al-Hames, Gerhard Rigoll |
ICIP (3) | 2 |
| 2005 | Video based online behavior detection using probabilistic multi stream fusionabstractIn the present treatise, we propose an approach for a highly configurable image based online person behaviour monitoring system. The particular application scenario is a crew supporting multi-stream on-board threat detection system, which is getting more desirable for the use in public transport. For such frameworks, to work robustly in mostly unconstrained environments, many subsystems have to be employed. Although the research field of pattern recognition has brought up reliable approaches for several involved subtasks in the last decade, there often exists a gap between reliability and the needed computational efforts. However in order, to accomplish this highly demanding task, several straight forward technologies, here the output of several so-called weak classifiers using low-level features are fused by a sophisticated Bayesian network. Dejan Arsic, Frank Wallhoff, Björn W. Schuller, Gerhard Rigoll |
ICIP (2) | 4 |
| 2005 | A Multi-Modal Mixed-State Dynamic Bayesian Network for Robust Meeting Event Recognition from Disturbed DataabstractIn this work we present a novel multi-modal mixed-state dynamic Bayesian network (DBN) for robust meeting event classification. The model uses information from lapel microphones, a microphone array and visual information to structure meetings into segments. Within the DBN a multi-stream hidden Markov model (HMM) is coupled with a linear dynamical system (LDS) to compensate disturbances in the data. Thereby the HMM is used as driving input for the LDS. The model can handle noise and occlusions in all channels. Experimental results on real meeting data show that the new model is highly preferable to all single-stream approaches. Compared to a baseline multi-modal early fusion HMM, the new DBN is more than 2.5%, respectively 1.5% better for clear and disturbed data, this corresponds to a relative error reduction of 17%, respectively 9% Marc A. Al-Hames, Gerhard Rigoll |
ICME | 2 |
| 2005 | Video Based Online Behavior Detection Using Probabilistic Multi Stream FusionabstractIn the present treatise, we propose an approach for a highly configurable image based online person behaviour monitoring system. The particular application scenario is a crew supporting multi-stream on-board threat detection system, which is getting more desirable for the use in public trans port. For such frameworks, to work robust in mostly unconstrained environments, many subsystems have to be employed. Although the research field of pattern recognition has brought up reliable approaches for several involved sub tasks in the last decade, there often exists a gap between reliability and the needed computational efforts. However in order, to accomplish this highly demanding task, several straight forward technologies, here the output of several so called weak classifiers using low-level features are fused by a sophisticated Bayesian Network. Dejan Arsic, Frank Wallhoff, Björn W. Schuller, Gerhard Rigoll |
ICME | 4 |
| 2005 | A neural-field-like approach for modeling human group actions in meetingsabstractIn this paper we investigate a new architecture for recognizing human group actions in meetings. These group actions provide a basis that enables effective browsing and querying in a meeting archive. For this task we propose an architecture that was inspired by the neural field theory. Our approach is particular, because contrary to other methods, we present all features to our classifier in parallel. The experiments show, that our system has comparable results to existing sequential techniques. Stephan Reiter, Gerhard Rigoll |
ICME | 2 |
| 2005 | Speaker Independent Speech Emotion Recognition by Ensemble ClassificationabstractEmotion recognition grows to an important factor in future media retrieval and man machine interfaces. However, even human deciders often experience problems realizing one's emotion, especially of strangers. In this work we strive to recognize emotion independent of the person concentrating on the speech channel. Single feature relevance of acoustic features is a critical point, which we address by filter-based gain ratio calculation starting at a basis of 276 features. As optimization of a minimum set as a whole in general saves more extraction effort, we furthermore apply an SVM-SFFS wrapper based search. For a more robust estimation we also integrate spoken content information by a Bayesian net analysis of ASR outputs. Overall classification is realized in an early feature fusion by stacked ensembles of diverse base classifiers. Tests ran on a 3,947 movie and automotive interaction dialog-turns database consisting of 35 speakers. Remarkable overall performance can be reported in the discrimination of the seven discrete emotions named in the MPEG-4 standard with added neutrality Björn W. Schuller, Stephan Reiter, Ronald Müller, Marc A. Al-Hames, Manfred K. Lang, Gerhard Rigoll |
ICME | 6 |
| 2005 | Feature Selection and Stacking for Robust Discrimination of Speech, Monophonic Singing, and Polyphonic MusicabstractIn this work we strive to find an optimal set of acoustic features for the discrimination of speech, monophonic singing, and polyphonic music to robustly segment acoustic media streams for annotation and interaction purposes. Furthermore we introduce ensemble-based classification approaches within this task. From a basis of 276 attributes we select the most efficient set by SVM SFFS. Additionally relevance of single features by calculation of information gain ratio is presented. As a basis of comparison we reduce dimensionality by PCA. We show extensive analysis of different classifiers within the named task. Among these are Kernel Machines, Decision Trees, and Bayesian Classifiers. Moreover we improve single classifier performance by Bagging and Boosting, and finally combine strengths of classifiers by StackingC. The database is formed by 2,114 samples of speech, and singing of 58 persons. 1,000 Music clips have been taken from the MTV-Europe-Top-20 1980-2000. The outstanding discrimination results of a working real time capable implementation stress the practicability of the proposed novel ideas. Björn W. Schuller, Brüning J. B. Schmitt, Dejan Arsic, Stephan Reiter, Manfred K. Lang, Gerhard Rigoll |
ICME | 6 |
| 2005 | Speaker independent emotion recognition by early fusion of acoustic and linguistic features within ensemblesabstractHerein we present a comparison of novel concepts for a robust fusion of prosodic and verbal cues in speech emotion recognition. Thereby 276 acoustic features are extracted out of a spoken phrase. For linguistic content analysis we use the Bag-of-Words text representation. This allows for integration of acoustic and linguistic features within one vector prior to a final classification. Extensive feature selection by filter- and wrapper based methods is fulfilled. Likewise optimal sets via SVM-SFFS and single feature relevance by information gain ratio calculation are presented. Overall classification is realised by diverse ensemble approaches. Among base classifiers Kernel Machines, Decision Trees, Bayesian classifiers, and memory-based learners are found. Acoustics only tests ran on a database comprising 39 speakers for speaker independent accuracy analysis. Additionally the public Berlin Emotional Speech database is used. A further database of 4,221 movie related phrases forms the basis of acoustic and linguistic information analysis evaluation. Overall remarkable performance in the discrimination of seven discrete emotions could be observed. 1. Björn W. Schuller, Ronald Müller, Manfred K. Lang, Gerhard Rigoll |
INTERSPEECH | 4 |
| 2005 | Multi-task learning strategies for a recurrent neural net in a hybrid tied-posteriors acoustic model
Jan Stadermann, Wolfram Koska, Gerhard Rigoll |
INTERSPEECH | 3 |
| 2004 | Speech emotion recognition combining acoustic features and linguistic information in a hybrid support vector machine-belief network architectureabstractIn this paper we introduce a novel approach to the combination of acoustic features and language information for a most robust automatic recognition of a speaker's emotion. Seven discrete emotional states are classified throughout the work. Firstly a model for the recognition of emotion by acoustic features is presented. The derived features of the signal-, pitch-, energy, and spectral contours are ranked by their quantitative contribution to the estimation of an emotion. Several different classification methods including linear classifiers, Gaussian mixture models, neural nets, and support vector machines are compared by their performance within this task. Secondly an approach to emotion recognition by the spoken content is introduced applying belief network based spotting for emotional key-phrases. Finally the two information sources are integrated in a soft decision fusion by using a neural net. The gain is evaluated and compared to other advances. Two emotional speech corpora used for training and evaluation are described in detail and the results achieved applying the propagated novel advance to speaker emotion recognition are presented and discussed. Björn W. Schuller, Gerhard Rigoll, Manfred K. Lang |
ICASSP (1) | 2 |
| 2004 | Reconstruction-free matching for fingerprint sweep sensorsabstractMany different types of silicon fingerprint sweep sensors presently enter the biometrics market. Since they provide small stripe image sequences instead of full fingerprint images, they require new matching algorithms. This paper presents a new one-stage, Viterbi-based matching approach which can be directly applied to the raw sweep sensor output. This is in contrary to the conventional two-stage approach where the stripe image sequences are combined to area images in a reconstruction step and subsequently fed into a recognition system appropriate for area images. Simulation results prove that the new algorithm is superior to the conventional two-stage system in terms of recognition performance and maximum allowed finger sweeping speed. Peter Morguet, Christian Narr, Henning Lorch, Frank Wallhoff, Gerhard Rigoll |
ICIP | 5 |
| 2004 | Action segmentation and recognition in meeting room scenariosabstractIn this proposal a novel implementation to find and recognize person actions in image sequences of meeting scenarios is introduced. Such extracted information can be used as the basis for content based browsing and the automated analysis of meetings. The presented system consists of four major functional blocks: the detection of a person, the feature extraction to describe actions and a sophisticated segmentation approach to find action boundaries. The fourth module consists of a statistical classifier. Beside the functionality of these blocks, the image material for training and testing purposes is briefly introduced. Frank Wallhoff, Martin Zobl, Gerhard Rigoll |
ICIP | 3 |
| 2004 | Recognition of partly occluded person actions in meeting scenariosabstractThe proposal describes a novel approach for handling partly occluded gestures in the feature domain with gesture specific Kalman-filters. An estimation of the Kalman-filter parameters using artificial neural networks is introduced. The approach is demonstrated and evaluated on data from a meeting scenario and can be generalized to all gesture based problems concerning partial occlusions. Martin Zobl, Andreas Laika, Frank Wallhoff, Gerhard Rigoll |
ICIP | 4 |
| 2004 | Applying Bayesian belief networks in approximate string matching for robust keyword-based retrievalabstractWe present a novel approach towards robust keyword-based retrieval. Bayesian belief networks are applied in a word-model based approximate string matching algorithm. Apart from a proven reliable performance in a working implementation on standard sources like digital text, wholly probabilistic modeling allows for integration of confidence measures and hypotheses obtained from preprocessing stages, like handwriting recognition or optical character recognition, respecting uncertainties on the lower levels. Furthermore, a flexible method to include the modeling of specific error types derived from humans and various input sources is provided. The remarkable performance of the algorithms presented was tested during extensive evaluation with respect to the Levenstein distance, which can be seen as the basis of state-of-the-art methods in this research field. The tests ran on a 14 K database containing common international music titles and four 10 K databases consisting of the most frequently used words in English, German, French and Dutch. Björn W. Schuller, Ronald Müller, Gerhard Rigoll, Manfred K. Lang |
ICME | 3 |
| 2004 | Multimodal music retrieval for large databasesabstractWe present a novel multi-modal access to large MP3 music databases. Retrieval can be fulfilled either in a content-based manner or by keywords. As input modalities, speech by natural language utterances or singing, and manual interaction by handwriting, typing or hardkeys are used. In order to achieve especially robust retrieval results and automatically suggest music to the user, contextual knowledge of the time, date, season, user emotion, and listening habits is integrated in the retrieval process. The system communicates with the user by speech or visual reactions. The concepts shown are especially designed for home and mobile access on tablet-PCs, PDAs, and similar PC solutions, The paper discusses the concept and a working prototype called Shangrila. An evaluation by a user study leads to an impression of the capabilities of the suggested approach to multimodal music retrieval. Björn W. Schuller, Gerhard Rigoll, Manfred K. Lang |
ICME | 2 |
| 2004 | Emotion recognition in the manual interaction with graphical user interfacesabstractWe introduce a novel approach to human emotion recognition, based on manual computer interaction. The presented methods rely on conventional graphical input devices. Firstly, a standard mouse as used on desktop PCs, and, secondly, the interaction with touch-screens or -pads as in public information terminals, palm-top devices or tablet PCs is considered. Additionally, the gain of the integration of touch pressure information is evaluated. Four discrete emotional states are classified: irritation, annoyance, reflectiveness, and neutral affect, for use in initiative tutoring, error clarification, Internet customer personalization, and others. The optimal feature-set is discussed and ranked according to a linear discriminant analysis. A working system, using support vector machines for the classification, is tested in real-life scenarios. A performance of up to 83.2% correct assignment clearly indicates that user emotion recognition is possible without special hardware in any standard graphical user environment, independent of the underlying application. Björn W. Schuller, Gerhard Rigoll, Manfred K. Lang |
ICME | 2 |
| 2004 | Discrimination of speech and monophonic singing in continuous audio streams applying multi-layer support vector machinesabstractWe present a novel approach to the discrimination of speech and monophonic singing for use in music information retrieval applications. A working prototype is introduced, applying multi-layer support vector machines for the discrimination and static high-level features derived from the pitch and energy contours of an acoustic signal. The feature set for discrimination is presented and ranked according to a linear discriminant analysis. For the automatic segmentation within an input signal stream, a further feature set is used for the discrimination of signal and noise. A corpus for training and evaluation comprising speech and monophonic singing data of nine performers is described in detail. The data has been labeled according to the judgment of another set of probands. A recognition rate of correct assignments of 99.2% could be reached, and demonstrates the high performance of the proposed methods. Björn W. Schuller, Gerhard Rigoll, Manfred K. Lang |
ICME | 2 |
| 2004 | A hybrid SVM/HMM acoustic modeling approach to automatic speech recognitionabstractAcoustic models based on a NN/HMM framework have been used successfully on various recognition tasks for continuous speech recognition. Recently tied-posteriors have been introduced within this context. Here, we present an approach combining SVMs and HMMs using the tied-posteriors idea. One set of SVMs calculates class posterior probabilities and shares these probabilities among all HMMs. The number of SVMs is varied as well as the input context and the amount of training data. Applying a first implementation, results on the AURORA2 task show already a promising improvement of the word error rate compared to the baseline acoustic models. 1. Jan Stadermann, Gerhard Rigoll |
INTERSPEECH | 2 |
| 2004 | Robust tracking of persons in real-world scenarios using a statistical computer vision approach
Gerhard Rigoll, Harald Breit, Frank Wallhoff |
Image Vis. Comput. | 1 |
| 2003 | Hidden Markov model-based speech emotion recognitionabstractWe introduce speech emotion recognition by use of continuous hidden Markov models. Two methods are propagated and compared. In the first method, a global statistics framework of an utterance is classified by Gaussian mixture models using derived features of the raw pitch and energy contour of the speech signal. A second method introduces increased temporal complexity, applying continuous hidden Markov models considering several states using low-level instantaneous features instead of global statistics. The paper addresses the design of working recognition engines, and results are achieved with respect to the alluded alternatives. A speech corpus consisting of acted and spontaneous emotion samples in German and English is described in detail. Both engines have been tested and trained using this equivalent speech corpus. Results in recognition of seven discrete emotions exceeded 86% recognition rate. In comparison, the judgment of human deciders classifying the same corpus at 79.8% recognition rate was analyzed. Björn W. Schuller, Gerhard Rigoll, Manfred K. Lang |
ICASSP (2) | 2 |
| 2003 | Flexible feature extraction and HMM design for a hybrid distributed speech recognition system in noisy environmentsabstractUsing the client device of a distributed speech recognizer usually implies the presence of background noise since most scenarios for distributed speech recognition (DSR) are situated in a non-office environment. Thus, the general task is to choose the most suitable feature extraction method for the given conditions. We present a hybrid speech recognition approach implemented for DSR that allows the choice of arbitrary feature vectors (regarding number and range of value) without changing the amount of data sent to the recognition engine. Experiments were carried out using mel-cepstrum and RASTA-PLP features on the AURORA database. Results show how the recognition performance under different noise conditions can be adjusted if the different features are combined, and that our hybrid approach to DSR has advantages that could not that easily be obtained with traditional DSR architectures. Jan Stadermann, Gerhard Rigoll |
ICASSP (1) | 2 |
| 2003 | Confidence Measures for an Address Reading SystemabstractIn this paper the performance of different confidence measures used for an address recognition system are evaluated. The recognition system for cursive handwritten German address words is based on hidden Markov models (HMMs). It is essential that the structure of the address (name, street, city, country) is known, so that a specific small but complete dictionary can be selected. Upon choosing a wrong dictionary (OOV: out-of-vocabulary) or misrecognizing a word, the recognition result should be rejected by means of the confidence measure. This paper points out two aspects: the comparison of four confidence measures for single words - based on the likelihood, a garbage-model, a two-best recognition or a character decoding - and the comparison of using complete or wrong dictionaries. It is shown that the best confidence measure - the two-best distance - has a quite different behavior using OOV. Anja Brakensiek, Jörg Rottland, Gerhard Rigoll |
ICDAR | 3 |
| 2003 | A flexible multimodal object tracking systemabstractIn this paper we present a flexible multimodal object tracking system. It is based on a particle filter, which combines the outputs of different measurement methods (also called modes or cues) in a flexible manner. The modes for locating the desired object can be selected depending on the specific object, which is to be tracked, and the environment of the object. By combining multiple modes we aim to add the strengths while at the same time to overcome the specific disadvantages of the different modes. Thus the robustness of the tracking system will be increased and a successful tracking will be possible in critical situations where a system using only a single mode would fail. We have used this approach for tracking persons in different environments. The measurement modes, which we have implemented for this purpose, are a pseudo 2-dimensional hidden Markov model (P2DHMM), a color based skin finder, and a motion detector. We describe the theory and the architecture of this tracking system and finally depict some exemplary results. Harald Breit, Gerhard Rigoll |
ICIP (3) | 2 |
| 2003 | Hidden Markov model-based speech emotion recognitionabstractIn this contribution we introduce speech emotion recognition by use of continuous hidden Markov models. Two methods are propagated and compared throughout the paper. Within the first method a global statistics framework of an utterance is classified by Gaussian mixture models using derived features of the raw pitch and energy contour of the speech signal. A second method introduces increased temporal complexity applying continuous hidden Markov models considering several states using low-level instantaneous features instead of global statistics. The paper addresses the design of working recognition engines and results achieved with respect to the alluded alternatives. A speech corpus consisting of acted and spontaneous emotion samples in German and English language is described in detail. Both engines have been tested and trained using this equivalent speech corpus. Results in recognition of seven discrete emotions exceeded 86% recognition rate. As a basis of comparison the similar judgment of human deciders classifying the same corpus at 79.8% recognition rate was analyzed. Björn W. Schuller, Gerhard Rigoll, Manfred K. Lang |
ICME | 2 |
| 2003 | HMM-based music retrieval using stereophonic feature information and framelength adaptationabstractMusic retrieval methods are in the focus of recent interest due to the increasing size of music databases as e.g. in the Internet. Among different query methods content-based media retrieval analyzing intrinsic characteristics of the source seems to form the most intuitive access. The key-melody in a song can be regarded as the major characteristic in music and leads to a query by humming or singing. In this paper we turn our attention to both, the features and the algorithm of matching in audio music retrieval. Nowadays approaches propagate the use of dynamic time warping for the matching process. As reference mostly midi-data or humming itself is used. However, first attempts matching humming to polyphonic audio exist. In this contribution we introduce hidden Markov models as an alternative for humming queries matching humming itself, mobile phone ring tones and polyphonic audio. The second object of our research is the introduction of a new way of melody enhancement prior to a latter feature extraction by use of stereophonic information. Further an adaptation throughout the extraction process of the frame length to the tempo of a musical piece helps improving similarity matching performance. The paper addresses the design of a working recognition engine and results achieved with respect to the alluded methods. A test database consisting of polyphonic audio clips, ring tones, and sung user data is described in detail. Björn W. Schuller, Gerhard Rigoll, Manfred K. Lang |
ICME | 2 |
| 2003 | A hybrid music retrieval system using belief networks to integrate multimodal queries and contextual knowledgeabstractRecently an increasing interest in music retrieval can be observed. Due to the growing amount of online and offline available music and a broadening user spectrum more efficient query methods are needed. We believe that only a parallel multimodal combination of different input modalities forms the most intuitive way to access desired media for any user. In this paper we introduce a query by humming, speaking, writing, and typing. The strengths of each modality are combined in a synergetic manner by a soft decision fusion. Songs can be referenced by their according melody, artist, title or other specific information. Further more the recognition of the actual user's emotion and external contextual knowledge helps to build an expectance of the intended song at a time. This constrains the hypothesis sphere of possible songs and leads to a more robust recognition or even a suggestive query. A combination of artificial neural networks, hidden Markov models and dynamic time warping integrated in a Bayesian belief network framework build the mathematical background of the chosen hybrid architecture. We address the implementation of a working system and results achieved by the introduced methods. Björn W. Schuller, Martin Zobl, Gerhard Rigoll, Manfred K. Lang |
ICME | 3 |
| 2003 | A real-time system for hand gesture controlled operation of in-car devicesabstractThe integration of more and more functionality into the human machine interface (HMI) of vehicles increases the complexity of device handling. Thus optimal use of different human sensory channels is an approach to simplify the interaction with in-car devices. This way the user convenience increases as much as distraction may decrease. In this paper a video based real-time hand gesture recognition system for in-car use is presented. It was developed in course of extensive usability studies. In combination with a gesture optimized HMI it allows intuitive and effective operation of a variety of in-car multimedia and infotainment devices with hand poses and dynamic hand gestures. Martin Zobl, Michael Geiger, Björn W. Schuller, Manfred K. Lang, Gerhard Rigoll |
ICME | 5 |
| 2003 | Distributed speech recognition on the WSJ taskabstractA comparison of traditional continuous speech recognizers with hybrid tied-posterior systems in distributed environments is presented for the first time on a challenging medium vocabulary task. We show how monophone and triphone systems are affected if speech features are sent over a wireless channel with limited bandwidth. The algorithms are evaluated on the Wall Street Journal database (WSJ0) and the results show that our monophone tied-posterior recognizer outperforms the traditional methods on this task by a dramatic reduction of the performance loss by a factor of 4 compared to nondistributed recognizers. 1. Jan Stadermann, Gerhard Rigoll |
INTERSPEECH | 2 |
| 2002 | Multimodal emotion recognition in audiovisual communicationabstractThis paper discusses innovative techniques to automatically estimate a user's emotional state analyzing the speech signal and haptical interaction on a touch-screen or via mouse. The knowledge of a user's emotion permits adaptive strategies striving for a more natural and robust interaction. We classify seven emotional states: surprise, joy, anger, fear, disgust, sadness, and neutral user state. The user's emotion is extracted by a parallel stochastic analysis of his spoken and haptical machine interactions while understanding the desired intention. The introduced methods are based on the common prosodic speech features pitch and energy, but rely also on the semantic and intention based features wording, degree of verbosity, temporal intention and word rate, and finally the history of user utterances. As further modality even touch-screen or mouse interaction is analyzed. The estimates based on these features are integrated in a multimodal way. The introduced methods are based on results of user studies. A realization proved to be reliable compared with subjective probands' impressions. Björn W. Schuller, Manfred K. Lang, Gerhard Rigoll |
ICME (1) | 3 |
| 2002 | Automatic topic identification in multimedia broadcast dataabstractThis paper presents a system that automatically scans multimedia data like TV or radio broadcasts for the presence of specific topics and, whenever topics of users' interests are detected, alerts the related user. Our current work on the three main modules of the system is shown. (1) The speech recognition system (with 18.7 % WER) is already among the most advanced German broadcast speech recognition systems. (2) The innovative topic identification approach, which is especially designed to work on the output of a speech recognizer, is compared to a standard text based approach. (3) The topic segmentation module has a good performance detecting not only scene cuts or speaker turns, but also real topic boundaries. Steffen Werner, Uri Iurgel, Andreas Kosmala, Gerhard Rigoll |
ICME (1) | 4 |
| 2001 | Content based indexing of images and video using face detection and recognition methodsabstractThis paper presents an image and video indexing approach that combines face detection and face recognition methods. Images of a database or frames of a video sequence are scanned for faces by a neural network-based face detector. The extracted faces are then grouped into clusters by a combination of a face recognition method using pseudo two-dimensional hidden Markov models and a k-means clustering algorithm. Each resulting main cluster consists of the face images of one person. In a subsequent step, the detected faces are labeled as one of the different people in the video sequence or the image database and the occurrence of the people can be evaluated. The results of the proposed approach on a TV broadcast news sequence are presented. It is demonstrated that the system is able to discriminate between three different newscasters and an interviewed person. Stefan Eickeler, Frank Wallhoff, Uri Iurgel, Gerhard Rigoll |
ICASSP | 4 |
| 2001 | New approaches to audio-visual segmentation of TV news for automatic topic retrievalabstractThis paper presents two new real-time approaches to segmentation of TV news shows into topics. The goal of this research work is the high precision retrieval of topics from TV news. For that purpose, the detection of correct topic boundaries is of great importance. We introduce a stochastic and a rule-based topic model based on HMM. The former combines features from the visual as well as from the audio channel of the news show, whereas the latter uses the video channel only. They are compared to the detection of topics using only the audio channel, which is common for many other approaches. The paper contains the following innovations: (1) the detected segment boundaries correspond directly to topics and not to video or audio cuts, as in most other segmentation methods; (2) an advanced stochastic topic model is introduced that uses audio as well as video features; (3) the introduced HMM-based approaches both outperform the audio-based approach. One algorithm has a very good topic boundary detection rate, whereas the other minimizes the number of wrongly inserted boundaries without missing too many real boundaries. Uri Iurgel, Ralf Meermeier, Stefan Eickeler, Gerhard Rigoll |
ICASSP | 4 |
| 2001 | Recognition of face profiles from the mugshot database using a hybrid connectionist/HMM approachabstractBiometrical systems have been the focus of concentrated research efforts in recent years. These systems can be used to identify a person or to grant a person access to something, e.g., a room. Face recognition technology has reached a level of performance at which frontal-view recognition of faces with slightly different facial expressions, view angles or head poses can be considered nearly solved. We present a novel hybrid ANN/HMM approach to recognize a person from that person's profile view (90) although the recognition system is trained with only one single frontal view of the person. Such a system can be useful for mugshot identification where a victim or witness has seen the criminal from the side only. Our approach uses neural methods in order to synthesize a profile out of the frontal view using no additional knowledge about the 3D shape and structure of a human head. The classification of the generated images is accomplished using a statistical HMM-approach. Frank Wallhoff, Stefan Müller 0001, Gerhard Rigoll |
ICASSP | 3 |
| 2001 | Comparing Adaptation Techniques for On-Line Handwriting RecognitionabstractThis paper describes an online handwriting recognition system with focus on adaptation techniques. Our hidden Markov model (HMM)-based recognition system for cursive German script can be adapted to the writing style of a new writer using either a retraining depending on the EM (expectation maximization)-approach or an adaptation according to the MAP (maximum a posteriori) or MLLR (maximum likelihood linear regression)-criterion. The performance of the resulting writer-dependent system increases significantly even if the amount of adaptation data is very small (about 6 words). So this approach is also applicable for online systems in hand-held computers such as PDAs. Special attention was paid to the performance comparison of the different adaptation techniques with the availability of different amounts of adaptation data ranging from a few words tip to 100 words per writer. Anja Brakensiek, Andreas Kosmala, Gerhard Rigoll |
ICDAR | 3 |
| 2001 | A Comparison of Character N-Grams and Dictionaries Used for Script RecognitionabstractIn this paper an off-line script recognition system is described, which makes use of a language model, that consists of backoff character n-grams. The performance of this open vocabulary recognition is compared with the use of closed dictionaries. The system is based on Hidden Markov Models (HMMs) using a hybrid modeling technique, which depends on a neural vector quantizer The presented recognition results refer to the SEDAL-database of degraded English documents such as photocopy, or fax and a writer-dependent handwritten database of cursive German script samples. Our resulting system for character recognition yields significantly, better recognition results for an unlimited vocabulary using language models. Anja Brakensiek, Gerhard Rigoll |
ICDAR | 2 |
| 2001 | Adaptation of an Address Reading System to Local Mail StreamsabstractA scheme for handwriting adaptation for post offices is described to improve recognition performance of German addresses. The recognition system is based on a tied-mixture hidden Markov model, whose parameters are updated using the expectation maximization technique, the maximum likelihood linear regression algorithm and a new discriminative adaptation technique, the scaled likelihood linear regression. Contrary to the usual approach of adapting a writer-independent system to a specific writer we propose to adapt the system to the writer-independent data of a specific post office. The resulting system for each post office yields up to 16% lower word recognition errors. Anja Brakensiek, Jörg Rottland, Frank Wallhoff, Gerhard Rigoll |
ICDAR | 4 |
| 2001 | Multi-Branch and Two-Pass HMM Modeling Approaches for Off-Line Cursive Handwriting RecognitionabstractBecause of large shape variations in human handwriting, cursive handwriting recognition remains a challenging task. Usually, the recognition performance depends crucially upon the pre-processing steps, e.g. the word baseline detection and segmentation process. Hidden Markov models (HMMs) have the ability to model similarities and variations among samples of a class. In this paper, we present a multi-branch HMM modeling method and an HMM-based two-pass modeling approach. Whereas the multi-branch HMM method makes the resulting system more robust with word baseline detection, the two-pass recognition approach exploits the segmentation ability of the Viterbi algorithm and creates another HMM set and carries out a second recognition pass. The total performance is enhanced by the combination of the two recognition passes. Experiments recognizing cursive handwritten words with a 30,000-word lexicon have been carried out. The results demonstrate that our novel approaches achieve better recognition performance and reduce the relative error rate significantly. Anja Brakensiek, Andreas Kosmala, Gerhard Rigoll |
ICDAR | 4 |
| 2001 | Improved person tracking using a combined pseudo-2D-HMM and Kalman filter approach with automatic background state adaptationabstractThis paper presents the continuation of our work on object tracking in presence of non-stationary background using a combination of a pseudo-2D hidden Markov model (P2DHMM) and a Kalman filter. It presents a major improvement by introducing a novel method that allows an automatic adaptation of our system to the changing background. Other improvements of the system's tracking capabilities are achieved by refined person models and normalization procedures. One of the major goals of our approach to tracking is to achieve high quality tracking results despite non-stationary background that can be caused e.g. by moving objects in the background or by camera operations such as panning or zooming. In previous publications we demonstrated that our combined P2DHMM/Kalman filter approach is an interesting solution to this problem, because it enables us to perform person tracking without the use of motion information. In this paper, we show that this approach can be further improved by adapting the system to the constantly changing background. We further demonstrate that such a background adaptation is very difficult to achieve in standard tracking approaches but can be effectively realized in our combined P2DHMM/Kalman filter approach. The effectiveness of this new procedure is demonstrated in experiments, where the tracking results and the quality of the person segmentation of our original system is compared to the results obtained with the improved approach. Harald Breit, Gerhard Rigoll |
ICIP (2) | 2 |
| 2001 | Retrieval of overlapping and touching objects using hidden Markov modelsabstractA content-based image retrieval system for overlapping and touching objects based on hidden Markov models is introduced. In a first step, unsupervised clustering in color and position space is performed in order to separate the objects. The clusters are handed over to the feature extraction, which is basically a polar subsampling, and finally rotation invariant Markov models are trained on those features. After presenting a query object, the HMMs which represent the individual clusters in the images are matched against the feature sequence calculated on the query image. Those database elements whose corresponding Markov models generated the highest similarity scores are retrieved. Three different clustering techniques, namely k-means clustering, LBG-algorithm and EM-algorithm are evaluated. Retrieval efficiencies up to 56.25% have been achieved on this challenging task. Stefan Müller 0001, Frank Wallhoff, Gerhard Rigoll |
ICIP (2) | 3 |
| 2001 | A comparison of discrete and continuous output modeling techniques for a pseudo-2D hidden Markov model face recognition systemabstractFace recognition has become an important topic within the field of pattern recognition and computer vision. In this field a number of different approaches to feature extraction, modeling and classification techniques have been tested. However, many questions concerning the optimal modeling techniques for high performance face recognition are still open. The face recognition system developed by our research group uses a discrete cosine transform (DCT) combined with the use of pseudo-2D hidden Markov models (P2DHMM). In the past our system used continuous probability density functions and was tested on a smaller database. This paper addresses the question of the presence of a major difference in recognition performance with discrete production probabilities compared to continuous ones. Therefore the system is tested using a larger subset of the FERET database. We show that we are able to achieve higher recognition scores and an improvement concerning the computation speed by using discrete modeling techniques. Frank Wallhoff, Stefan Eickeler, Gerhard Rigoll |
ICIP (2) | 3 |
| 2001 | A novel hybrid face profile recognition system using the FERET and MUGSHOT databasesabstractFace recognition has established itself as an important subbranch of pattern recognition within the field of computer science. Many state-of-the-art systems have focused on the task of recognizing frontal views of people. We present an approach for recognizing profile views (90/spl deg/) with a system trained on transformed frontal views. The system combines an artificial neural network (ANN) and a classification process based on hidden Markov models (HAM). One of the main ideas of this system is to perform the recognition task without the use of any 3D-information of heads and faces. The presented system has been tested with subsets of the FERET and the MUGSHOT databases. Frank Wallhoff, Gerhard Rigoll, Payman Moallem |
ICIP (1) | 2 |
| 2001 | Distributed speech recognition using traditional and hybrid modeling techniquesabstractWe compare the performance of different acoustic model-ing techniques on the task of distributed speech recogni-tion (DSR). The DSR technology is interesting for speech recognition tasks in mobile environments, where features are sent from a thin client to a server where the ac-tual recognition is performed. The evaluation is done on the TI digits database which consists of single dig-its and digit-chains spoken by American-English talkers. We investigate clean speech and speech added with white noise. Our results show that new hybrid or discrete mod-eling techniques can outperform standard continuous sys-tems on this task. 1. Jan Stadermann, Ralf Meermeier, Gerhard Rigoll |
INTERSPEECH | 3 |
| 2001 | Scaled likelihood linear regression for hidden Markov model adaptationabstractIn the context of continuous Hidden Markov Model (HMM) based speech-recognition, linear regression approaches have become popular to adapt the acoustic models to the specific speaker's characteristics. The well known Maximum Likelihood Linear Regression (MLLR) [1] and Maximum A Posteriori Linear Regression (MAPLR) [2] are just two of them, which differ primarily in the training objective they are maximizing. Frank Wallhoff, Daniel Willett, Gerhard Rigoll |
INTERSPEECH | 3 |
| 2001 | An Integrated Approach to Shape and Color-Based Image Retrieval of Rotated Objects Using Hidden Markov ModelsabstractAn integrated approach to shape and color-based image retrieval, where the cues color and shape are both utilized in a local rather than a global way, is presented in this paper. An experimental retrieval system has been developed, which enables the user to search a color image database intuitively by presenting simple sketches. In order to be able to perform an elastic matching, which is especially needed in sketch-based image retrieval, objects in the images are represented by Hidden Markov Models. The use of streams (sets of features that are assumed to be statistically independent) within the HMM framework allows the integration of shape and color derived features into a single model, thereby allowing to control the influence of the different streams via stream weights. The approach has been evaluated on a color image database containing 120 different isolated objects with arbitrary orientation and showed good retrieval results with several users. Furthermore, the use of HMMs allows efficient pruning and thus a fast retrieval even with large databases. Stefan Müller 0001, Stefan Eickeler, Gerhard Rigoll |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2001 | A continuous density interpretation of discrete HMM systems and MMI-neural networksabstractThe subject of this paper is the integration of the traditional vector quantizer (VQ) and discrete hidden Markov models (HMM) combination in the mixture emission density framework commonly used in automatic speech recognition (ASR). It is shown that the probability density of a system that consists of a VQ and a discrete classifier can be interpreted as a special case of a semi-continuous mixture model. Thus, the VQ parameters and the classifier can be trained jointly. In this framework, a gradient based VQ training method for single and multiple feature stream systems is derived. This leads to an approach that is directly related to the paradigm of maximum mutual information (MMI) neural networks, that has been successfully applied as VQ in ASR earlier. In continuous speech recognition experiments that were carried out for the Resource Management and Wall Street Journal databases the presented systems achieve recognition accuracies that compete well with comparable Gaussian mixture HMMs. Thus, we demonstrate that the performance degradations, often reported for discrete HMM systems, are not mainly caused by the vector quantization process in itself, but that they are due to the traditional separation of the VQ and the HMM during parameter estimation. These degradations can be avoided by training of the entire system as described here, while keeping the attractive computational speed of discrete HMMs. Christoph Neukirchen, Jörg Rottland, Daniel Willett, Gerhard Rigoll |
IEEE Trans. Speech Audio Process. | 4 |
| 2000 | Comparison of Confidence Measures for Face RecognitionabstractThis paper compares different confidence measures for the results of statistical face recognition systems. The main applications of a confidence measure are rejection of unknown people and the detection of recognition errors. Some of the confidence measures are based on the posterior probability and some on the ranking of the recognition results. The posterior probability is calculated by applying Bayes' rule with different ways to approximate the unconditional likelihood. The confidence measure based on the ranking is a new method. Experiments to evaluate the confidence measures are carried out on a pseudo 2D hidden Markov model-based face recognition system and the Bochum face database. Stefan Eickeler, Mirco Jabs, Gerhard Rigoll |
FG | 3 |
| 2000 | Crane Gesture Recognition Using Pseudo 3-D Hidden Markov ModelsabstractA recognition technique based on novel pseudo 3D hidden Markov models, which can integrate spatial as well as temporal derived features is presented. The approach allows the recognition of dynamic gestures such as waving hands as well as static gestures such as standing in a special pose. Pseudo 3D hidden Markov models (P3DHMM) are an extension of the pseudo 2D case, which has been successfully used for the classification of images and the recognition of faces. In the P3DHMM case the so-called superstates contain P2DHMM and thus whole image sequences can be generated by these models. Our approach has been evaluated on a crane signal database, which consists of 12 different predefined gestures for maneuvering cranes. Stefan Müller 0001, Stefan Eickeler, Gerhard Rigoll |
FG | 3 |
| 2000 | Person Tracking in Real-World Scenarios Using Statistical MethodsabstractThis paper presents a novel approach to robust and flexible person tracking using an algorithm that combines two powerful stochastic modeling techniques: pseudo-2D hidden Markov models (P2DHMM) used for capturing the shape of a person within an image frame, and the well-known Kalman-filtering algorithm, that uses the output of the P2DHMM for tracking the person by estimation of a bounding box trajectory indicating the location of the person within the entire video sequence. Both algorithms cooperate together in an optimal way, and with this co-operative feedback, the proposed approach even makes the tracking of people possible in the presence of background motions caused by moving objects or by camera operations as, e.g., panning or zooming. Our results are confirmed by several tracking examples in real scenarios, shown at the end of the paper and provided on the Web server of our institute. Gerhard Rigoll, Stefan Eickeler, Stefan Müller 0001 |
FG | 1 |
| 2000 | A novel error measure for the evaluation of video indexing systemsabstractError evaluation of video indexing systems is a problem for which no satisfactory solution has been presented so far. This paper introduces a new error measure for evaluating the results of video indexing systems. The measure compares the scene and shot boundaries of the correct index of a video sequence with the automatically created index, and decides dependent on predefined penalties whether a scene boundary is falsely detected, undetected or correctly detected. Based on this decision the measure calculates the error rates for the segmentation capability of the indexing system. Then the algorithm extends the reference and the test index by the falsely detected and undetected boundaries and compares the content classes of the resulting indexes to determine the classification error rate. The presented measure is compared to two existing measures, and the pros and cons of each measure are listed. Stefan Eickeler, Gerhard Rigoll |
ICASSP | 2 |
| 2000 | Tied posteriors: an approach for effective introduction of context dependency in hybrid NN/HMM LVCSRabstractThis paper presents a method to improve the recognition rate of hybrid connectionist/HMM speech recognition systems. At the same time this approach allows the easy introduction of context dependent models in the hybrid framework. The approach is based on a standard hybrid connectionist/HMM recognizer, in which the neural nets are trained to estimate the a posteriori probabilities for all phones in each input frame. In the approach presented here, the probabilities of the neural nets are used to replace the codebook of a tied-mixture HMM system. Therefore the resulting system is called tied posterior. The advantages of this structure are that an arbitrary HMM-topology can be used, and that all context dependency and all clustering techniques used in tied-mixture systems can be applied to this hybrid speech recognition system. The approach has been evaluated on the Wall Street Journal (WSJ) database, with the result, that it outperforms the standard hybrid approach on this task. Jörg Rottland, Gerhard Rigoll |
ICASSP | 2 |
| 2000 | Frame-discriminative and confidence-driven adaptation for LVCSRabstractMaximum likelihood linear regression (MLLR) has become the most popular approach for adapting speaker-independent hidden Markov models to a specific speaker's characteristics. However, it is well known, that discriminative training objectives outperform maximum likelihood training approaches, especially in cases where training data is very limited, as it always is the case in adaptation tasks. Therefore, this paper explores the application of a frame-based discriminative training objective for adaptation. It presents evaluations for supervised as well as for unsupervised adaption on the 1993 WSJ adaptation tests of native and non-native speakers. Relative improvements in word error rate of up to 25% could be measured compared to the MLLR adapted recognition systems. Along with unsupervised adaptation, the paper also presents the improvements achieved by the application of confidence measures. They provided an average relative improvement of 10% compared to ordinary unsupervised MLLR. Frank Wallhoff, Daniel Willett, Gerhard Rigoll |
ICASSP | 3 |
| 2000 | DUcoder-the Duisburg University LVCSR stackdecoderabstractWith this paper, we present the DUcoder, the LVCSR decoder developed at Duisburg University. The decoder performs the Viterbi search for the most probable word sequence in recognition systems that make use of HMMs and backo N-gram language models. In principle, the decoding strategy is similar to the one of the so-called stackdecoders. During the development of the decoder, emphasis has been laid upon innovations for rapidly speeding up decoding by carefully performing approximations. Besides a brief presentation of the decoder's overall design, this paper points out the crucial issues with respect to speed and recognition performance. Evaluations are carried out on a German LVCSR system with a vocabulary of 100; 000 words, wordinternal triphones and a trigram language model. Closeto -real-time performance is achieved with 12% additional error while a decoder conguration which runs in around 40 times real-time causes no search error on the evaluation set. 1. INTRODUCTION The realiz... Daniel Willett, Christoph Neukirchen, Gerhard Rigoll |
ICASSP | 3 |
| 2000 | An HMM Based Two-Pass Approach for Off-Line Cursive Handwriting Recognition
Anja Brakensiek, Gerhard Rigoll |
ICMI | 3 |
| 2000 | Improved Degraded Document Recognition with Hybrid Modeling Techniques and Character N-GramsabstractA robust multifont character recognition system for degraded documents, such as photocopy or fax, is described. The system is based on hidden Markov models using discrete and hybrid modeling techniques, where the latter makes use of an information theory-based neural network. The presented recognition results refer to the SEDAL-database of English documents using no dictionary. It is also demonstrated that the usage of a language model that consists of character n-grams yields significantly better recognition results. Our resulting system clearly outperforms commercial systems and leads to further error rate reductions compared to previous results reached on this database. Anja Brakensiek, Daniel Willett, Gerhard Rigoll |
ICPR | 3 |
| 2000 | On-Line Handwritten Formula Recognition with Integrated Correction Recognition and ExecutionabstractPresents the extension of an approach for online handwritten formula recognition. The introduction of some constraints concerning the handwriting production process and the robust hidden Markov model (HMM) framework yields recognition rates up to 97.7%. The high, but still limited recognition rate demonstrates the user's need for some correction facilities, in order to modify misclassified symbols and thus to avoid the re-writing of the entire expression. Such facilities can further be used to provide the opportunity to develop an expression on the electronic paper, as it is often desired from a practical point of view. Andreas Kosmala, Gerhard Rigoll, Anja Brakensiek |
ICPR | 2 |
| 2000 | Compound splitting and lexical unit recombination for improved performance of a speech recognition system for German parliamentary speechesabstractS.945-948 Martha A. Larson, Daniel Willett, Joachim Köhler, Gerhard Rigoll |
INTERSPEECH | 4 |
| 2000 | Recognition of JPEG compressed face images based on statistical methods
Stefan Eickeler, Stefan Müller 0001, Gerhard Rigoll |
Image Vis. Comput. | 3 |
| 1999 | Experiments in topic indexing of broadcast news using neural networksabstractThe paper deals with the problem of extracting topic information from news show stories by statistical methods. It is shown that the traditional topic-dependent n-gram language modeling approach can be decomposed in order to apply neural networks for topic indexing. Two specific problems in training of these networks are addressed: a very sparse data distribution in the stories and the superposition of different topics in a story. The first problem is tackled by an integrated smoothing approach in the backpropagation method; an expansion of the neural network structure can be used to cope with topic mixtures in stories. Due to the efficient parameter sharing the application of neural networks results in a small improvement in topic indexing performance on a small corpus of broadcast news compared to the traditional topic-dependent n-gram method. Christoph Neukirchen, Daniel Willett, Gerhard Rigoll |
ICASSP | 3 |
| 1999 | Refining tree-based state clustering by means of formal concept analysis, balanced decision trees and automatically generated model-setsabstractDecision tree-based state clustering has emerged in as the most popular approach for clustering the states of context dependent hidden Markov model based speech recognizers. The application of sets of phones, mainly phonetically motivated, that limit the possible clusters, results in a reasonably good modeling of unseen phones while it still enables to model specific phones very precisely whenever this is necessary and enough training data is available. Formal concept analysis, a young mathematical discipline, provides means for the treatment of sets and sets of sets that are well suited for further improving tree-based state clustering. The possible refinements are outlined and evaluated in this paper. The major merit is the proposal of procedures for the adaptation of the number of sets used for clustering to the amount of available training data, and of a method that generates suitable sets automatically without the incorporation of additional knowledge. Daniel Willett, Christoph Neukirchen, Jörg Rottland, Gerhard Rigoll |
ICASSP | 4 |
| 1999 | Performance Evaluation of a New Hybrid Modeling Technique for Handwriting Recognition using On-Line and Off-Line DataabstractThe paper deals with the performance evaluation of a novel hybrid approach to large vocabulary cursive handwriting recognition and contains various innovations. 1) It presents the investigation of a new hybrid approach to handwriting recognition, consisting of hidden Markov models (HMMs) and neural networks trained with a special information theory based training criterion. This approach has only been recently introduced successfully to online handwriting recognition and is now investigated for the first time for offline recognition. 2) The hybrid approach is extensively compared to traditional HMM modeling techniques and the superior performance of the new hybrid approach is demonstrated. 3) The data for the comparison has been obtained from a database containing online handwritten data which has been converted to offline data. Therefore, a multiple evaluation has been carried out, incorporating the comparison of different modeling techniques and the additional comparison of each technique for online and offline recognition, using a unique database. The results confirm that online recognition leads to better recognition results due to the dynamic information of the data, but also show that it is possible to obtain recognition rates for offline recognition that are close to the results obtained for online recognition. Furthermore, it can be shown that for both online and offline recognition, the new hybrid approach clearly outperforms the competing traditional HMM techniques. It is also shown that the new hybrid approach yields superior results for the offline recognition of machine printed multifont characters. Anja Brakensiek, Andreas Kosmala, Daniel Willett, Gerhard Rigoll |
ICDAR | 5 |
| 1999 | On-Line Handwritten Formula Recognition using Hidden Markov Models and Context Dependent Graph GrammarsabstractThis paper presents an approach for the recognition of on-line handwritten mathematical expressions. The hidden Markov model (HMM) based system makes use of simultaneous segmentation and recognition capabilities, avoiding a crucial segmentation during pre-processing. With the segmentation and recognition results, obtained from the HMM recognizer it is possible to analyze and interpret the spatial two-dimensional arrangement of the symbols. We use a graph grammar approach for the structure recognition, also used in off-line recognition process, resulting in a general tree-structure of the underlying input-expression. The resulting constructed tree can be translated to any desired syntax (for example: Lisp, KT/sub E/X, and OpenMath). Andreas Kosmala, Gerhard Rigoll, Stéphane Lavirotte, Loïc Pottier |
ICDAR | 2 |
| 1999 | Advanced State Clustering for Very Large Vocabulary HMM-based On-Line Handwriting RecognitionabstractThe paper presents some novel methods for the introduction of context dependent hidden Markov models (HMM) to online handwriting recognition. The use of these so-called n-graphs can lead to substantially improved modeling accuracy, but requires some intelligent parameter reduction methods (state clustering). This is especially the case for the investigated very large vocabulary system, incorporating an active vocabulary of 200000 words. Switching from context independent models to context dependent models-considering the underlying vocabulary-yields in the worst case to 25000 HMMs and very poor trainability for most of the introduced models. Therefore, the conducted investigations are focused on an appropriate state clustering method which is supported by decision trees and some new self organizing approaches to generate the required trees. The presented comparison takes also the different context dependencies (left, right or both sides) into consideration. Andreas Kosmala, Daniel Willett, Gerhard Rigoll |
ICDAR | 3 |
| 1999 | Multimedia Database Retrieval using Hand-Drawn SketchesabstractWe propose an integrated approach to color and shape based retrieval which enables the user to search a color image database intuitively by presenting simple sketches. Each of the images is represented by a Hidden Markov Model (HMM) which has been concatenated with modified filler models in order to obtain scale and rotation invariance properties. These properties are particularly important when using query by sketch due to the skew which occurs naturally in human handwriting and in drawings. The use of streams (sets of features that are assumed to be statistically independent) within the HMM framework allows the integration of shape and color derived features into a single model, thereby allowing one to control the influence of the different streams via stream weights. The approach has been evaluated on a color image database containing 120 different isolated objects with arbitrary orientation and showed good retrieval results with several users. Furthermore, the use of HMMs allows efficient pruning and thus a fast retrieval even with large databases. Stefan Müller 0001, Stefan Eickeler, Gerhard Rigoll |
ICDAR | 3 |
| 1999 | Searching an Engineering Drawing Database for User-specified ShapesabstractWe present a novel approach to the retrieval of engineering drawings based on the use of stochastic models. Engineering drawing databases can be searched intuitively by presenting sketches or shapes which represent details such as, for example, screws or holes in the drawings of mechanical parts. The query is represented by a pseudo 2D Hidden Markov Model (P2DHMM) which is surrounded by filler states. These filler states generate the remaining part of the engineering drawing apart from the query shape or sketch itself. Thus, our approach aims to retrieve those images containing certain details and also locates these details in the retrieved images, even in cases where the query shape is embedded in, for examplle, hatching or is connected to other parts in the drawing. The proposed technique achieves a good performance which is demonstrated by a number of query and retrieval examples. Stefan Müller 0001, Gerhard Rigoll |
ICDAR | 2 |
| 1999 | High Quality Face Recognition in Jpeg Compressed ImagesabstractThis paper presents an advanced face recognition system that is based on the use of Pseudo 2-D HMMs and coefficients of the 2-D DCT as features. A major advantage of our approach is the fact that our face recognition system works directly with JPEG-compressed face images, i.e. it uses directly the DCT-features provided by the JPEG standard, without any necessity of completely decompressing the image before recognition. The recognition rates on the Olivetti Research Laboratory (ORL) face database are 100% for the original images and 99.5% for JPEG compressed domain recognition. A comparison with other face recognition systems evaluated on the ORL database, shows that these are the best recognition results on this database. Stefan Eickeler, Stefan Müller 0001, Gerhard Rigoll |
ICIP (1) | 3 |
| 1999 | Pseudo 3-D Hmms for Image Sequence RecognitionabstractIn this paper, a novel approach to image sequence recognition is presented. We refer to this approach as pseudo 3-D Hidden Markov modeling, a technique which can integrate spatial as well as temporal derived features in an elegant and efficient way. This allows the recognition of dynamic gestures such as waving hands as well as more static gestures such as standing in a special pose. Pseudo 3-D Hidden Markov Models (P3DHMMs) are a natural extension of the pseudo 2-D case, which has been successfully used for the classification of images. In the P3DHMM case the so-called superstates contain P3DHMMs and thus whole image sequences can be generated by these models. The feasibility of our approach is demonstrated in this paper by a number of experiments on a crane signal database, which consists of 12 different predefined gestures for maneuvering cranes. To our knowledge, this is the first publication which reports about the usage of pseudo 3-D hidden Markov models. Stefan Müller 0001, Stefan Eickeler, Gerhard Rigoll |
ICIP (4) | 3 |
| 1999 | Robust Person Tracking with Non-Stationary Background Using a Combined Pseudo-2D-Hmm and Kalman-Filter ApproachabstractThis paper presents a novel approach to robust and flexible person tracking using an algorithm that combines two powerful stochastic modeling techniques: The first one is the technique of so-called Pseudo-2D Hidden Markov Models (P2DHMMs) used for capturing the shape of a person within an image frame, and the second technique is the well-known Kalman-filtering algorithm, that uses the output of the P2DHMM for tracking the person by estimation of a bounding box trajectory indicating the location of the person within the entire video sequence. Both algorithms are cooperating together in an optimal way, and with this cooperative feedback, the proposed approach even makes the tracking of persons possible in the presence of background motions, for instance caused by moving objects such as cars, or by camera operations as, for example, panning or zooming. Our results are confirmed by several tracking examples in real scenarios, shown at the end of the paper and provided on the web server of our institute. Gerhard Rigoll, Stefan Müller 0001, Bernd Winterstein |
ICIP (4) | 1 |
| 1999 | Speaker adaptation using regularization and network adaptation for hybrid MMI-NN/HMM speech recognitionabstractThis paper describes, how to perform speaker adaptation for a hybrid large vocabulary speech recognition system. The hybrid system is based on a Maximum Mutual Information Neural Network (MMINN), which is used as a Vector Quantizer (VQ) for a discrete HMM speech recognizer. The combination of MMINNs and HMMs has shown good performance on several large vocabulary speech recognition tasks like RM and WSJ. This paper now presents two approaches to perform speaker adaptation with this hybrid system. The first approach is a transformation of the feature space, which is performed by a neural network with maximum likelihood (ML) as objective function for the complete system, which means, that the parameters of the NN are estimated in order to match the HMM-parameters of the pre-trained speaker independent system. The second approach is to adapt the HMM parameters depending on the amount of training data available per HMM, using a regularization approach. Both approaches can be applied join... Jörg Rottland, Christoph Neukirchen, Daniel Willett, Gerhard Rigoll |
EUROSPEECH | 4 |
| 1999 | A discriminative training procedure based on language model and dictionary for LVCSRabstractIn today's HMM-based speech recognition systems, the parameters are most commonly estimated according to the Maximum Likelihood criterion. Because of limited training data, however, discriminative objectives provide better parameter estimates with respect to the Maximum A-Posteriori decision used for decoding. The question of which distribution functions to discriminate from which and to what degree is the most crucial when performing discriminative parameter estimation. This is particularly difficult because beside the distribution functions, the recognition procedure is restricted and guided by several other sources of information, such as language model and transition matrices. This paper extends the approach presented in [10] to the case of triphones, refines the theory and estimation of the state-to-state confusion metric and proposes an approximation that allows the application of the approach on context-dependent systems with reasonable computational cost. The evaluation is perf... Daniel Willett, Stefan Müller 0001, Gerhard Rigoll |
EUROSPEECH | 3 |
| 1998 | Speech recognition with a new hybrid architecture combining neural networks and continuous HMM
Daniel Willett, Gerhard Rigoll |
ESANN | 2 |
| 1998 | A NN/HMM hybrid for continuous speech recognition with a discriminant nonlinear feature extractionabstractThis paper deals with a hybrid NN/HMM architecture for continuous speech recognition. We present a novel approach to set up a neural linear or nonlinear feature transformation that is used as a preprocessor on top of the HMM system's RBF-network to produce discriminative feature vectors that are well suited for being modeled by mixtures of Gaussian distributions. In order to omit the computational cost of discriminative training of a context-dependent system, we propose to train a discriminant neural feature transformation on a system of low complexity and reuse this transformation in the context-dependent system to output improved feature vectors. The resulting hybrid system is an extension of a state-of-the-art continuous HMM system, and in fact, it is the first hybrid system that really is capable of outperforming these standard systems with respect to the recognition accuracy, without the need for discriminative training of the entire system. In experiments carried out on the Resource Management 1000-word continuous speech recognition task we achieved a relative error reduction of about 10% with a recognition system that, even before, was among the best ever observed on this task. Gerhard Rigoll, Daniel Willett |
ICASSP | 1 |
| 1998 | Speaker adaptation for hybrid MMI/connectionist speech-recognition systemsabstractWe present a new adaptation technique for our hybrid large vocabulary continuous speech recognition system. In most adaptation approaches the HMM parameters are reestimated. In our approach, however, we train a speaker independent continuous speech recognizer, then we keep the HMM parameters fixed and we train a second network, which transforms the features of the adaptation data to fit the HMM parameters. Thus, less parameters have to be estimated, and therefore this approach performs well even for a small number of adaptation data. With this approach we achieve relative improvements in recognition rates on the Wall Street Journal (WSJ) task of 16.5%. Jörg Rottland, Christoph Neukirchen, Gerhard Rigoll |
ICASSP | 3 |
| 1998 | Efficient search with posterior probability estimates in HMM-based speech recognitionabstractIn this paper we present the methods we developed to estimate posterior probabilities for HMM states in continuous and discrete HMM-based speech recognition systems and several ways to speed up decoding by using these posterior probability estimates. The proposed pruning techniques are state deactivation pruning (SDP), similar to an approach proposed for hybrid recognition systems, and a novel posteriori-based lookahead technique, posteriori lookahead pruning (PLP), that evaluates future posteriors in order to exclude unlikely HMM states as early as possible during search. By applying the proposed methods we managed to vastly reduce the decoding time consumed by our time-synchronous Viterbi-decoder for recognition systems based on the Verbmobil and the Wall Street Journal database with hardly any additional search error. Daniel Willett, Christoph Neukirchen, Gerhard Rigoll |
ICASSP | 3 |
| 1998 | Hidden Markov model based continuous online gesture recognitionabstractPresents the extension of an existing vision-based gesture recognition system using hidden Markov models(HMMs). Several improvements have been carried out in order to increase the capabilities and the functionality of the system. These improvements include position independent recognition, rejection of unknown gestures, and continuous online recognition of spontaneous gestures. We show that especially the latter requirement is highly complicated and demanding, if we allow the user to move in front of the camera without any restrictions and to perform the gestures spontaneously at any arbitrary moment. We present solutions to this problem by modifying the HMM-based decoding process and by introducing online feature extraction and evaluation methods. Stefan Eickeler, Andreas Kosmala, Gerhard Rigoll |
ICPR | 3 |
| 1998 | On-line handwritten formula recognition using statistical methodsabstractThis paper presents the design of a system for the processing and recognition of online handwritten mathematical formulas. The hidden Markov model (HMM) based system is trained and evaluated with a writer dependent database consisting of 100 formulas for the training and an additional set of 30 formulas for the test. With the introduction of some constraints, it is possible to obtain high recognition rates up to 97.7%, and to transform the transcriptions of the formulas into T/sub E/X-syntax in order to achieve a convenient visualization of the results. Andreas Kosmala, Gerhard Rigoll |
ICPR | 2 |
| 1998 | Tree-based state clustering using self-organizing principles for large vocabulary on-line handwriting recognitionabstractThe introduction of trigraphs offers a powerful method for the accuracy enhancement of handwriting modeling. A trigraph is a hidden Markov model (HMM) for a special character with defined adjacent characters. Especially in large vocabulary systems, as they are investigated here, the number of unseen trigraphs for which no training samples are available, exceeds the number of seen trigraphs by far. This paper presents a novel approach, which allows a synthesis of unseen trigraphs from seen trigraphs. With the method proposed here, a mean relative error reduction of 42% was obtained on a writer dependent system, resulting in an overall word recognition rate of 94.1%. Andreas Kosmala, Gerhard Rigoll |
ICPR | 2 |
| 1998 | A systematic comparison between on-line and off-line methods for signature verification with hidden Markov modelsabstractThis paper presents an extensive investigation of various HMM-based techniques for signature verification Different feature extraction methods and HMM topologies are compared in order to obtain an optimized high performance signature verification system. Furthermore, this paper may be the first systematic comparison of online and off-line methods for signature verification using exactly the same database, and leading to the surprising result that the difference in performance for both approaches is relatively small. Gerhard Rigoll, Andreas Kosmala |
ICPR | 1 |
| 1998 | A new hybrid approach to large vocabulary cursive handwriting recognitionabstractPresents a hybrid modeling technique that is used for the first time in hidden Markov model-based handwriting recognition. This new approach combines the advantages of discrete and continuous Markov models and it is shown that this is especially suitable for modeling the features typically used in handwriting recognition. The performance of this hybrid technique is demonstrated by an extensive comparison with traditional modeling techniques for a difficult large vocabulary handwriting recognition task. Gerhard Rigoll, Andreas Kosmala, Daniel Willett |
ICPR | 1 |
| 1998 | Soft state-tying for HMM-based speech recognitionabstractThis paper introduces a method for regularization of HMM systems that avoids parameter overfitting causedby insufficient training data. Regularization is done by augmenting the EM training method by a penalty term that favors simple and smooth HMM systems. The penalty term is constructed as a mixture model of negative exponential distributions that is assumed to generate the state dependent emission probabilities of the HMMs. This new method is the successful transfer of a well known regularization approachin neural networks to the HMM domain andcan be interpreted as a generalization of traditional state-tying for HMM systems. The effect of regularization is demonstrated for continuous speech recognition tasks by improving overfitted triphone models and by speaker adaptation with limited training data. 1. INTRODUCTION A general problem when constructing statistical pattern recognition systems is to ensure the capability to generalize well, i.e. the system must be able to classify dat... Christoph Neukirchen, Daniel Willett, Gerhard Rigoll |
ICSLP | 3 |
| 1998 | Efficient computation of MMI neural networks for large vocabulary speech recognition systemsabstractThis paper describes, how to train Maximum Mutual Information Neural Networks (MMINN) in an efficient way, with a new topology. Large vocabulary speech recognition systems, based on a Hybrid MMI/connectionist HMM combination, have shown good performance on several tasks [1] and [2]. MMINNs are trained to maximize the mutual information between the index of the winning output neuron (Winner-Takes-All network) and the phonetical class of the corresponding acoustic frame. One major problem of MMI-neural networks is the high computational effort, which is needed for the training of the neural networks. The computational effort is proportional to the input and output size of the neural network and to the number of training samples. This paper shows two approaches, that demonstrate, how these long training times can be reduced with very low or even no loss in recognition accuracy. This is achieved by the use of phonetical knowledge, to build a network topology based on phonetical classes. Jörg Rottland, Andre Ludecke, Gerhard Rigoll |
ICSLP | 3 |
| 1998 | A German dialogue system for scheduling dates and meetings by naturally spoken continuous speechabstractIn this paper, we present the basic design principles and architecture of a dialogue system for scheduling appointments. This mixed-initiative dialogue system integrates an automatic speaker-independent speech recognition engine for continuously spoken German, a speech synthesizer and a scheduler database application to build up a scheduler that is purely driven by natural continuous speech and thus, does not need any visual display device. With these properties it is a prototype for a speech driven palm-size computer application and could be integrated in miniature computers that come along with no display device at all. 1. INTRODUCTION Dialogue systems enable the user to fulfill some well defined interaction with the machine by natural conversational speech in a spoken dialogue, in which the computer takes the part of one of the dialogue participants. The techniques and principles for the development of robust dialogue systems have attracted a lot of attention in the recent years. ... Daniel Willett, Arno Romer, Jörg Rottland, Gerhard Rigoll |
ICSLP | 4 |
| 1998 | Confidence measures for HMM-based speech recognitionabstractIn this paper, we describe our work on the field of confidence measures for HMM-based speech recognition. Confidence measures are a means of estimating the recognition reliability for single words of the recognizer output. The possible applications of such measures are manifold. We present our experiments with well known approachesand proposesome new ones. Particularly, we propose to combine the mere acoustical measures with language model-based ones for continuous speech recognition that involves a stochastic language model. This slightly improves the acoustical measures and preserves their advantage of being computationally very cheap. Experiments are carried out on a German isolated word recognition system and on continuous speech recognition systems for the Resource Management database and the Wall Street Journal WSJ0 task. 1. INTRODUCTION Word-based confidencemeasures for speechrecognition basedon hidden Markov models (HMMs) have for some years now been an important research top... Daniel Willett, Andreas Worm, Christoph Neukirchen, Gerhard Rigoll |
ICSLP | 4 |
| 1998 | Controlling the Complexity of HMM Systems by Regularization
Christoph Neukirchen, Gerhard Rigoll |
NIPS | 2 |
| 1997 | An investigation of the use of trigraphs for large vocabulary cursive handwriting recognitionabstractThis paper presents an extensive investigation of the use of trigraphs for online cursive handwriting recognition based on hidden Markov models (HMMs). Trigraphs are context dependent HMMs representing a single written character in its left and right context, similar to triphones in speech recognition. Looking at the great success of triphones in continuous speech recognition, it was always a challenging and open question, if the introduction of trigraphs could lead to substantially improved handwriting recognition systems. The results of this investigation are indeed extremely encouraging: the introduction of suitable trigraphs led to a 50% relative error reduction for a writer dependent 1000 word handwriting recognition system, and to a 35% relative error reduction for the same system with an extended 30000 word vocabulary for cursive handwriting recognition. Andreas Kosmala, Jörg Rottland, Gerhard Rigoll |
ICASSP | 3 |
| 1997 | Advanced training methods and new network topologies for hybrid MMI-connectionist/HMM speech recognition systemsabstractThis paper deals with the construction and optimization of a hybrid speech recognition system that consists of a combination of a neural vector quantizer (VQ) and discrete HMMs. In our investigations an integration of VQ based classification in the continuous classifier framework is given and some constraints are derived that must hold for the PDFs in the discrete pattern classifier context. Furthermore it is shown that for ML training of the whole system the VQ parameters must be estimated according to the maximum mutual information (MMI) criterion. A novel training method based on gradient search for neural networks that serve as optimal VQ is derived. This allows faster training of arbitrary network topologies compared to the traditional MMI-NN training. An integration of multilayer MMI-NNs as the VQ in the hybrid discrete HMM based speech recognizer leads to a large improvement compared to other supervised and unsupervised single layer VQ systems. For the speaker independent Resource Management database the constructed hybrid MMI-connectionist/HMM system achieves recognition rates that are comparable to traditional sophisticated continuous PDF HMM systems. Christoph Neukirchen, Gerhard Rigoll |
ICASSP | 2 |
| 1997 | New improved feature extraction methods for real-time high performance image sequence recognitionabstractThis paper describes new feature extraction methods which can be used very effectively in combination with statistical methods for image sequence recognition. Although these feature extraction methods can be used for a wide variety of image sequence processing applications, the target application presented in this paper is gesture recognition. The novel feature extraction methods have been integrated into an HMM-based gesture recognition system and led to substantial improvements for this system. It turned out that the new features are not only able to describe the gesture characteristics much better than the old features, but additionally they also led to a dramatic reduction in dimensionality of the feature vector used for representing each frame of the image sequence. This resulted in the fact that it was possible to use the novel features in combination with a new architecture for statistical image sequence recognition. The result of this investigation is a high performance gesture recognition system with significantly improved recognition rates and real-time capabilities. Gerhard Rigoll, Andreas Kosmala |
ICASSP | 1 |
| 1997 | Improved On-Line Handwriting Recognition Using Context Dependent Hidden Markov ModelsabstractThe paper presents the introduction of context dependent hidden Markov models for cursive, unconstrained handwriting recognition with large vocabularies. Since context dependent models were successfully introduced to speech recognition (R. Bahl et al., 1980; R. Schwartz et al., 1984; K. Lee, 1990), it seems obvious, that the use of trigraphs could also lead to improved online handwriting recognition systems (A. Kosmala et al., 1997). In analogy to triphones in speech recognition, trigraphs are context dependent sub word units representing a single written character in its left and right context. The tests were conducted on a writer dependent system with three different writers and two different vocabulary sizes (1000 words and 30000 words). The results we obtained with the trigaph based system compared to the monograph system, are very encouraging: a mean relative error reduction of 46% for the 1000 word handwriting recognition system and a mean relative error reduction of 37% for the same system with the 30000 word vocabulary. We believe that this represents one of the first systematic investigations of the influence of context dependent models and parameter reduction methods for a difficult large vocabulary handwriting recognition task. Andreas Kosmala, Jörg Rottland, Gerhard Rigoll |
ICDAR | 3 |
| 1997 | Reduced lexicon trees for decoding in a MMIi-connectionist/HMM speech recognition system
Christoph Neukirchen, Daniel Willett, Gerhard Rigoll |
EUROSPEECH | 3 |
| 1997 | Large vocabulary speech recognition with context dependent MMI-connectionist / HMM systems using the WSJ databaseabstractIn this paper we present a context dependent hybrid MMI-connectionist / Hidden Markov Model (HMM) speech recognition system for the Wall Street Journal (WSJ) database. The hybrid system is build with a neural network, which is used as a vector quantizer (VQ) and an HMM with discrete probablility density functions, which has the advantage of a faster decoding. The neural network is trained on an algorithm, that tries to maximize the mutual information between the classes of the input features (e.g. phones, triphones, etc.) and the neural firing sequence of the network. Jörg Rottland, Christoph Neukirchen, Daniel Willett, Gerhard Rigoll |
EUROSPEECH | 4 |
| 1997 | A new approach to generalized mixture tying for continuous HMM-based speech recognitionabstractIn this paper we present a new approach for a generalized tying of mixture components for continuous mixture-density HMM-based speech recognition systems. With an iterative pruning and splitting procedure for the mixture components, this approach offers a very accurate and detailed representation of the acoustic space and at the same time keeps the number of parameters reasonably small in favor of a robust parameter estimation and a fast decoding. Contrary to other approaches, it does not require a strict clustering of the pdfs into subsets that share their mixture components, so that it is capable of providing more general and flexible types of mixture tying. We applied the new approach on a semi-continuous HMM (SCHMM)-system for the Resource Management task and improved its recognition performance by 12% and vastly accelerated the decoding because of a much faster likelihood computation. 1. INTRODUCTION In continuous mixture-density HMM-based speech recognition systems the HMM stat... Daniel Willett, Gerhard Rigoll |
EUROSPEECH | 2 |
| 1997 | Hybrid NN/HMM-Based Speech Recognition with a Discriminant Neural Feature Extraction
Daniel Willett, Gerhard Rigoll |
NIPS | 2 |
| 1996 | A new hybrid system based on MMI-neural networks for the RM speech recognition taskabstractWe present a hybrid speech recognition system for speaker independent continuous speech recognition. The system combines a novel information theory based neural network (NN) paradigm and discrete Hidden Markov models (HMMs) including state-of-the-art techniques like state clustered triphones. The novel NN type is trained by an algorithm based on principles of self-organization that achieves maximum mutual information between the generated output labels and the basic phonetic classes. The structure of the hybrid system is quite similar to a classical VQ-HMM system but the vector quantizer (VQ) is replaced by the NN. To evaluate the system we use the speaker independent part of the resource management (RM) database. We obtained an important improvement by introducing a novel kind of context dependent basic classes used by the acoustic processor. The average RM recognition result with a word-pair grammar is now 95.2% what is significantly better than a classical VQ-system, slightly better than a different hybrid system with a recurrent network as probability estimator, and very close to the best continuous probability density function (PDF) HMM speech recognizers. Gerhard Rigoll, Christoph Neukirchen, Jörg Rottland |
ICASSP | 1 |
| 1996 | Fast online video image sequence recognition with statistical methodsabstractIn this paper a fast method to recognize image sequences is presented. It is based on a discrete statistical model consisting of a vector quantizer and a special probabilistic neural net giving an estimation for the a posteriori probability P(SEQUENCE|DATA), which allows to classify image sequences without applying rules depending on the content of the sequence. The simple feature extraction also allows the classification with discrete hidden Markov models. As an application we present results from a test conducted for the classification of various gestures done by human beings in front of a video camera for both classification methods, which gave promising recognition results in real time. Michael Schuster, Gerhard Rigoll |
ICASSP | 2 |
| 1996 | A new approach to video sequence recognition based on statistical methodsabstractA fast method for image sequence recognition is presented. The method is based on a discrete statistical model consisting of a vector quantizer and a special probabilistic neural network, which allows one to classify image sequences without applying rules depending on the content of the sequence. The simple feature extraction also allows the classification with discrete hidden Markov models. As an application we present results from a test conducted for the classification of various gestures done by human beings in front of a video camera. For both classification methods we obtained promising recognition results in real time. The system obtained 90.0% recognition rate for person-independent classification of 10 or even 15 different gestures, which we considered as a surprisingly high rate for such a complex task. The system has been demonstrated at a large industrial fair and has confirmed its high recognition rates and its robustness under real-world conditions. Gerhard Rigoll, Andreas Kosmala, Mike Schuster |
ICIP (3) | 1 |
| 1996 | A comparison between continuous and discrete density hidden Markov models for cursive handwriting recognitionabstractThis paper presents the results of the comparison of continuous and discrete density hidden Markov models (HMMs) used for cursive handwriting recognition. For comparison, a subset of a large vocabulary (1000 word), writer-independent online handwriting recognition system for word and sentence recognition was used, which was developed at Duisburg University. This system has some unique features that are rarely found in other HMM-based character recognition systems, such as: (1) option between discrete, continuous, or hybrid modeling of HMM probability density distributions; (2) large vocabulary recognition based on either printed or cursive word or complete sentence input; (3) optimized HMM topology with an unusually large number of HMM states; and (4) use of multiple label streams for coding of handwritten information. Emphasis in this paper is on the comparison between continuous and discrete density HMMs, since this is still an open question in handwriting recognition, and is crucial for the future development of the system. However, in order to give a complete description of the basic system architecture, some of the above mentioned issues are also addressed. The surprising result of our investigation was the fact that discrete density models led to better results than continuous models, although this is generally not the case for HMM-based speech recognition systems. With the optimized system, a 70% word recognition rate was obtained for a challenging large-vocabulary, writer-independent sentence input task. Gerhard Rigoll, Andreas Kosmala, Jörg Rottland, Christoph Neukirchen |
ICPR | 1 |
| 1996 | A New Approach to Hybrid HMM/ANN Speech Recognition using Mutual Information Neural Networks
Gerhard Rigoll, Christoph Neukirchen |
NIPS | 1 |
| 1995 | Speech recognition experiments with a new multilayer LVQ network (MLVQ)abstractIn this paper, a new neural network paradigm and its application to recognition of speech patterns is presented. The novel NN paradigm is a multilayer version of the well-known LVQ algorithm from Kohonen. The approach includes the following innovations and improvements compared to other popular neural network paradigms: 1) It is - according to the knowledge of the author - the first multilayer version of the classical LVQ algorithm, which is usually based on a one-layer neural network architecture. 2) It presents a new NN architecture, since it uses a perceptron-like propagation function for the hidden layers, and an Euclidean-like propagation function for the output layer. 3) Its architecture can be considered as an optimal compromise between the multilayer principle of an MLP and the principle of representing one class by several neurons adopted from classical LVQ. 4) Compared to MLP, the training of an MLVQ network with the same number of neurons is more effective, since only the we... Gerhard Rigoll |
EUROSPEECH | 1 |
| 1995 | Large vocabulary speaker-independent continuous speech recognition with a new hybrid system based on MMI-neural networksabstractThis paper presents a new hybrid system for speaker independent continuous speech recognition in a large vocabulary task. The hybrid system is a combination of context dependent discrete Hidden Markov Models and artificial neural networks that are trained by an information theory based algorithm. This algorithm maximizes the Mutual Information (MMI) between the network output and the phone descriptions by applying a self-organizing learning approach instead of forcing constrained network outputs. Recognition results have shown that the new hybrid system outperforms a classical k-means-VQ-based HMM-system. For the speaker independent DARPA Resource Management (RM) task (perplexity 60) we report a decrease in word recognition error rate up to 35% (close to the best continuous pdf systems). 1. INTRODUCTION Research during recent years has shown that Hidden Markov Models (HMMs) are the most successful approach for the recognition of speaker-independent continuous speech (e.g. see [1]). B... Gerhard Rigoll, Christoph Neukirchen, Jörg Rottland |
EUROSPEECH | 1 |
| 1994 | Mutual information neural networks: a new connectionist approach for dynamic speech recognition tasksabstractA new probabilistic neural network paradigm for dynamic pattern recognition problems is presented. The approach includes the following innovations: 1) it is based on a self-organizing learning approach using information theory principles. 2) The neuron activations are interpreted as probabilities and represent probabilistic decision boundaries in the feature space. 3) A combination of unsupervised and supervised learning algorithms is used to train the network weights. 4) The neuron probabilities can be further refined by corrective training methods. 5) The neural network can process dynamic patterns of arbitrary length, and can be even used for continuous speech recognition, although it is not a recurrent network. 6) The output activations of the neural network can be evaluated directly or optionally treated as input to hidden Markov models in order to construct a hybrid recognition system. The network has been tested for the recognition of dynamic speech patterns and performs better than a discrete HMM system with a codebook size equal to the number of output neurons in the neural net.> Gerhard Rigoll |
ICASSP (2) | 1 |
| 1994 | Maximum mutual information neural networks for hybrid connectionist-HMM speech recognition systemsabstractThis paper proposes a novel approach for a hybrid connectionist-hidden Markov model (HMM) speech recognition system based on the use of a neural network as vector quantizer. The neural network is trained with a new learning algorithm offering the following innovations. (1) It is an unsupervised learning algorithm for perceptron-like neural networks that are usually trained in the supervised mode. (2) Information theory principles are used as learning criteria, making the network especially suitable for combination with a HMM-based speech recognition system. (3) The neural network is not trained using the standard error-backpropagation algorithm but using instead a newly developed self-organizing learning approach. The use of the hybrid system with the neural vector quantizer results in a 25% error reduction compared with the same HMM system using a standard k-means vector quantizer. The training algorithm can be further refined by using a combination of unsupervised and supervised learning algorithms. Finally, it is demonstrated how the new learning approach can be applied to multiple-feature hybrid speech recognition systems, using a joint information theory-based optimization procedure for the multiple neural codebooks, resulting in a 30% error reduction. Gerhard Rigoll |
IEEE Trans. Speech Audio Process. | 1 |
| 1993 | Speaker adaptation using improved speaker Markov models
Gerhard Rigoll |
ICASSP (2) | 1 |
| 1993 | Joint optimization of multiple neural codebooks in a hybrid connectionist-HMM speech recognition systemabstractThis paper proposes a new approach for a hybrid connectionistHMM speech recognition system. The system consists of a multi-feature HMM-based recognition module using three different neural networks as multiple neural codebooks. Each neural network receives a different feature (i.e. cepstrum, delta cepstrum, and delta power) as input and generates a vector quantizer label obtained from the firing neuron in the output layer. The neural networks are first trained separately using a special self-organizing information theory-based learning method. A 26% error reduction is obtained with this method, compared to the performance of the same system using multiple k-means vector quantizers with the same codebook size. In a second training phase, the neural codebooks are further refined by extending the information theory-based training criterion into a joint criterion reflecting the joint information content and the dependencies of the three different label streams. This further improves the er... Gerhard Rigoll |
EUROSPEECH | 1 |
| 1992 | Unsupervised information theory-based training algorithms for multilayer neural networksabstractThe author describes a novel learning algorithm for multilayer neural networks. The training neural networks are used as a vector quantizer (VQ) in a hidden Markov model (HMM)-based speech recognition system. The approach offers the following innovations: (1) it represents an unsupervised learning algorithm for multilayer neural networks. (Usually, multilayer neural networks are only trained in supervised mode); (2) information theory principles are used as learning criteria for the neural networks; and (3) the neural networks are not trained using the standard backpropagation algorithm, by using instead a new unsupervised learning procedure. The use of a neural network as a VQ trained with the new algorithm in combination with an HMM-based speech recognition system results in a 25% error reduction compared to the same HMM system using a standard k-means vector quantizer.> Gerhard Rigoll |
ICASSP | 1 |
| 1991 | Information theory-based supervised learning methods for self-organizing maps in combination with hidden Markov modelingabstractThe author presents various aspects of the combination of neural networks (NNs) and hidden Markov modeling (HMM) techniques. The combination of HMM with Kohonen's self-organizing maps is investigated in order to improve the performance of HMM-based speech recognition systems. The investigation has led to the development of a supervised learning method for the self-organizing map. This supervised learning method is based on information theory principles, leading to new design rules for the self-organizing map, making the map more suitable for combination with HMM techniques. The author also presents an information-theory-based approach for the automatic control of the map parameters during learning and a general consideration of the use of information theory principles for the design of neural networks in combination with HMM for improved processing of time-varying patterns.> Gerhard Rigoll |
ICASSP | 1 |
| 1990 | Baseform adaptation for large vocabulary hidden Markov model based speech recognition systemsabstractA method for adaptation of the IBM speech recognition system in the situation where the system is already trained for the new speaker and one tries to further adapt and improve the system while it is actually being used by the new speaker in the recognition mode is described. A special kind of adaptation is investigated where the emphasis is not on the adaptation of the statistical parameters of the Markov models but on the adaptation of the structure of these models. This structure is defined by the baseforms describing the composition of word models from phone models in the system. Therefore, baseform adaptation corresponds directly to the adaptation of the new system to the personal speaker characteristics of the new user. Several different baseform adaptation schemes are investigated and it is demonstrated that for a speaker who has already trained the system and achieves a 95.2% recognition performance, the performance can be further improved to 96.3%.> Gerhard Rigoll |
ICASSP | 1 |
| 1990 | Information theory principles for the design of self-organizing maps in combination with hidden Markov modeling for continuous speech recognitionabstractResulting from that combination is the aspect of designing the map using different rules from those usually mentioned in the standard literature for modifying the environment and the adaptation gain during learning. This can be explained by the fact that hidden Markov modeling is an information-theory approach, and the combination of self-organizing maps with MHH implies the use of information-theory principles also for the design of the map leading to the modified requirements for the learning procedure mentioned above. It is shown that substantial improvements can be obtained if the design principles presented are used Gerhard Rigoll |
IJCNN | 1 |
| 1989 | Speaker adaptation for large vocabulary speech recognition systems using speaker Markov modelsabstractAn alternative approach to speaker adaptation for a large-vocabulary hidden-Markov-model-based speech recognition system is described. The goal of this investigation was to train the IBM speech recognition system with only five minutes of speech data from a new speaker instead of the usual 20 minutes without the recognition rate dropping by more than 1-2%. The approach is based on the use of a stochastic model representing the different properties of the new speaker and an old speaker for which the full training set of 20 minutes is available. It is called a speaker Markov model. It is shown how the parameters of such a model can be derived and how it can be used for transforming the training set of the old speaker in order to use it in addition to the short training set of the new speaker. The adaptation algorithm was tested with 12 speakers. The average recognition rate dropped from 96.4% to 95.2% for a 5000-word vocabulary task. The decoding time increased by a factor of 1.35; this factor is often 3-5 if other adaptation algorithms are used.> Gerhard Rigoll |
ICASSP | 1 |
| 1989 | An information theory approach to speaker adaptation
Gerhard Rigoll |
EUROSPEECH | 1 |
| 1988 | Formant tracking with quasilinearizationabstractAn algorithm for the calculation of formant tracks from a speech signal is presented. It is based on the method of quasilinearization, a nonlinear parameter estimation procedure. What distinguishes this method from the most other methods is the fact that the formants are directly calculated from the speech signal, using a nonlinear model for the speech production that is based on formants and LPC parameters (FLPC model). Since quasilinearization is a nonrecursive parameter estimation method, this procedure is also much faster than nonlinear estimation algorithms for formant tracking which are based on extended Kalman filters.> Gerhard Rigoll |
ICASSP | 1 |
| 1987 | The dectalk system for German: A study of the modification of a text-to-speech converter for a foreign languageabstractThis paper describes the development of the German version of the DECtalk system, which was originally designed for the American language by D.H. Klatt. The aim of this paper is not only to provide an overview on the problems and difficulties for German text-to-speech conversion, using the cascade/parallel formant synthesizer and on the use of new algorithms for parameter extraction, but also to provide a study of the modification procedure which is necessary to build a new language version for a text-to-speech system which was designed for a different, language. These experiences are important for the future design of multilingual text-to-speech systems because the modification from one language to another language gives automatically the answer to many questions which are interesting for the design of a multilingual system or at least a system which can be easily modified for another language. The paper describes the most important steps that have to be performed during the modification procedure, i.e. text normalization and letter-to-sound rules, description of the used synthesizer, description of the tools used for speech analysis and comparison of synthetic and natural speech, tuning of the stationary parts of the phonemes, transitions, prosodies and generation of different voices. Gerhard Rigoll |
ICASSP | 1 |
| 1986 | A new algorithm for estimation of formant trajectories directly from the speech signal based on an extended Kalman-filterabstractAn alternative approach for formant tracking is proposed in this paper. The algorithm is based on the use of an extended Kalman-Filter and calculates the formants directly from the speech signal, i.e. without previous transformation into the frequency domain. The algorithm gives very smooth formant tracks even for long utterances which contain unvoiced fricatives and stops. In the first part of the paper, the basic mathematical principles are introduced. In the second part, the advantages and disadvantages of the algorithm are discussed and the results are compared to the standard methods for formant tracking. Gerhard Rigoll |
ICASSP | 1 |