VLDB 2026 Research / reviewers in the wild / expert
Matthias Wölfel
dblp:84/6910
· DBLP profile ↗
58ranked-venue papers
22as first author
13since 2021 · last 2025
0000-0003-1601-5146ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 49 · 20 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 25 · 9 first-author · 11 since 2021Artificial intelligence and machine learning · 22 · 8 first-author · 2 since 2021Systems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Making Lecture Videos Accessible for Students who are Blind or have Low Vision through AI-Assisted Navigation and Visual Question AnsweringabstractDesigning accessible lectures and lecture materials is crucial to promote inclusive higher education.We conducted need-finding interviews with 12 students who are blind or have low vision to learn their perspectives on how lectures and lecture material could become more accessible through Artificial Intelligence (AI) technologies.Key insights from the interviews reveal that students envision AI to automatically customize lecture material, connect disparate information sources, for example, to better keep track of the current lecture slide, and enhance interaction and engagement with lecture material.Based on these insights, we developed the LectureAssistant prototype, employing an iterative design process with visually impaired users that features AI-assisted video navigation and chatbot interaction.In a final evaluation with seven students, the participants expressed enthusiasm for features such as AI-powered video search and the possibility of asking questions about visual content in the current video frame.They provided valuable suggestions for future improvements, including notifications for lecture slide transitions and the provision of a short overview function for a slide.Insights from the study indicate great potential of the prototype to improve accessibility of lecture videos for students with visual impairments, although they also point to crucial areas for improvement, such as more reliable and personalized image descriptions. Katharina Anderer, Karin Müller 0001, Lukas Strobel, Matthias Wölfel, Jan Niehues, Kathrin Maria Gerling |
ASSETS | 4 |
| 2024 | MaViLS, a Benchmark Dataset for Video-to-Slide Alignment, Assessing Baseline Accuracy with a Multimodal Alignment Algorithm Leveraging Speech, OCR, and Visual FeaturesabstractThis paper presents a benchmark dataset for aligning lecture videos with corresponding slides and introduces a novel multimodal algorithm leveraging features from speech, text, and images. It achieves an average accuracy of 0.82 in comparison to SIFT (0.56) while being approximately 11 times faster. Using dynamic programming the algorithm tries to determine the optimal slide sequence. The results show that penalizing slide transitions increases accuracy. Features obtained via optical character recognition (OCR) contribute the most to a high matching accuracy, followed by image features. The findings highlight that audio transcripts alone provide valuable information for alignment and are beneficial if OCR data is lacking. Variations in matching accuracy across different lectures highlight the challenges associated with video quality and lecture style. The novel multimodal algorithm demonstrates robustness to some of these challenges, underscoring the potential of the approach. Katharina Anderer, Andreas Reich, Matthias Wölfel |
INTERSPEECH | 3 |
| 2024 | Fusing Components for an Attentive and Emotionally Expressive Companion Robot: Meet ZENITabstractThis paper introduces a new companion robot as an open research platform for studying human-robot interaction: The ZENIT Enabling Natural Interaction Technology. Open-source solutions for companion robots are often complicated to build or not thought of holistically. ZENIT is based on a few low-cost components, namely a smartphone and a stationary robot arm in conjunction with an average computer and a separate camera. This concept makes it possible not to be tied to specific hardware. Our platform aims to enable easy use of the multimodality of human communication channels such as body language, spatial behavior, facial expressions and speech so that ZENIT can understand its interlocutor and express itself appropriately. The fusion of various open-source software components in an adaptable, lightweight and resource-efficient distributed system simplifies its replication. The system demonstrated an adequate response time (speech: 2.69s, facial expressions: 0.55s, bodily movements: 0.32s) for perception and expression combined, as well as reliability in public testing. Evaluation of the robot’s expressive capabilities in an online survey (N=23) revealed users were able to understand six of eight emotions conveyed by ZENIT using facial expressions and/or body language (except for contempt and fear) and perceived its non-verbal feedback as significantly more expressive when facial expressions were accompanied with bodily movements. Christian Felix Purps, Matthias Wölfel |
RO-MAN | 2 |
| 2024 | Identifying the Information Gap for Visually Impaired Students during Lecture TalksabstractVisual slides have become a common tool in university classrooms, assisting students in following along with lectures. However, visually impaired individuals face an information gap when visual slides are the primary lecture material. This paper aims to identify which information is most likely to be withheld from visually impaired individuals. To analyse this information gap, the study first extracts textual elements from lecture slides and converts visual elements into textual descriptions on multiple levels that discriminate between different semantic depths of visuals. This information is compared to the transcribed audio of the lecture talk by a semantic similarity analysis. The study tests how a transformer model can be used for an automatic semantic similarity analysis, which is validated through human evaluation. The results indicate variations in the level of detail and semantic depth with which visuals are described across different lectures. As shown, direct references to visuals and elemental details of visuals are often omitted in speech. However, this information is important for enabling mental visualization processes for visually impaired individuals and for following along with the lecture. The identified information gap can serve as valuable feedback for lecturers and assistive systems, helping them to more effectively provide visually impaired students with the information that they may have otherwise missed during the lecture. Katharina Anderer, Matthias Wölfel, Jan Niehues |
VL/HCC | 2 |
| 2023 | Exploring Perception and Preference in Public Human-Agent Interaction: Virtual Human Vs. Social Robot
Christian Felix Purps, Wladimir Hettmann, Thorsten Zylowski, Nathalia Sautchuk Patricio, Daniel Hepperle, Matthias Wölfel |
ArtsIT (2) | 6 |
| 2023 | Identification of Unreliable Data in in-VR Surveys using Biosignal SensorsabstractSurveys that rely on self-reporting are often prone to error. Therefore, measures to identify meaningless, careless, or fraudulent responses in surveys can improve data quality. In the past, various indicators have been proposed to identify data with implausible responses in questionnaires, e.g. by including attention check questions or by applying statistical techniques. Immersive virtual reality (VR) environments are equipped with several biosignal sensors. We propose and investigate whether sensory data already provided by the headset or additional biosignal sensors can be useful for automatically detecting unreliable data in in-VR surveys. Our results show that biosensor signals, such as eye tracking and electrocardiography, can provide useful cues, and that this information can be combined with statistical techniques (namely the validity score) to further improve classification accuracy. Matthias Wölfel, Wladimir Hettmann |
CW | 1 |
| 2022 | Engaging Museum Visitors with AI-Generated Narration and Gameplay
Wladimir Hettmann, Matthias Wölfel, Marius Butz, Kevin Torner, Janika Finken |
ArtsIT | 2 |
| 2022 | When Left Behind in Multi-User Virtual Reality: Clues to Indicate the Absence of UsersabstractRecently, with the increasing popularity of immersive virtual reality, synchronous communication via avatars in virtual worlds has become more popular. In contrast to asynchronous communication—like social media platforms—people expect others to be present during embodied, virtual interactions. However, for a variety of reasons, this is not always the case.While some communication media provide a good indicator that the other is still present, such as live images of the conversation partner in video conferencing, not all media do provide this information naturally, e.g., when the interlocutor is silent on the phone. In immersive multi-user virtual reality, a wrong and misleading signal may even be provided when an avatar is still shown while its owner is absent. In many situations, an assumed presence of the other who is actually no longer there can lead to confusion. In the presented empirical study, we show that users in shared immersive environments prefer an indicator when a person is absent and that the preferred way of visualization depends on the use case. Matthias Wölfel, Lara-Marie Bäsing, Jonas Deuchler |
CW | 1 |
| 2022 | Aspects of visual avatar appearance: self-representation, display type, and uncanny valleyabstractThe visual representation of human-like entities in virtual worlds is becoming a very important aspect as virtual reality becomes more and more "social". The visual representation of a character's resemblance to a real person and the emotional response to it, as well as the expectations raised, have been a topic of discussion for several decades and have been debated by scientists from different disciplines. But as with any new technology, the findings may need to be reevaluated and adapted to new modalities. In this context, we make two contributions which may have implications for how avatars should be represented in social virtual reality applications. First, we determine how default and customized characters of current social virtual reality platforms appear in terms of human likeness, eeriness, and likability, and whether there is a clear resemblance to a given person. It can be concluded that the investigated platforms vary strongly in their representation of avatars. Common to all is that a clear resemblance does not exist. Second, we show that the uncanny valley effect is also present in head-mounted displays, but-compared to 2D monitors-even more pronounced. Daniel Hepperle, Christian Felix Purps, Jonas Deuchler, Matthias Wölfel |
Vis. Comput. | 4 |
| 2021 | Influence of Visual Appearance of Agents on Presence, Attractiveness, and Agency in Virtual Reality
Marius Butz, Daniel Hepperle, Matthias Wölfel |
ArtsIT | 3 |
| 2021 | Reconstructing Facial Expressions of HMD Users for Avatars in VR
Christian Felix Purps, Simon Janzer, Matthias Wölfel |
ArtsIT | 3 |
| 2021 | Entering a new Dimension in Virtual Reality Research: An Overview of Existing Toolkits, their Features and ChallengesabstractVirtual reality becomes a medium to be explored for itself, to study human factors and human behavior within these worlds, and to infer possible behavior in the real world. Among many advantages, building test routines in virtual environments remains a challenge due to the lack of established procedures and toolkits. To encourage research in this direction and lower the barrier to entry, it is necessary to simplify the process of setting up a research environment in virtual reality by providing appropriate toolkits. This paper discusses what challenges need to be overcome, what features might be relevant, and compares available toolkits. Matthias Wölfel, Daniel Hepperle, Christian Felix Purps, Jonas Deuchler, Wladimir Hettmann |
CW | 1 |
| 2021 | Investigating the Acceptance of the Cognitive Assistant Reflect that Supports Humans in Computer-Oriented WorkabstractAs more and more sensors find their way into the workplace, new types of digital assistance become possible. Despite their undeniable importance in improving occupational health and safety, there has long been disagreement about how and whether cognitive performance at work should be measured, analyzed, and used to either provide feedback to the user or report to the supervisor. The collection of sensory information is perceived very differently depending on cultural background, personality, device owner, and intended use of the data collected. While personal devices—such as smartwatches—enjoy a high level of acceptance, collecting data on employer devices is generally considered unacceptable and, depending on the country, even prohibited by law. To investigate the perceived usefulness at work, we designed Reflect, which responds to critical user conditions by informing the user. We found that, contrary to common belief, the use of cognitive assistance supporting workers or students in computer-oriented work is perceived positively by most participants and can be increased by emphasizing transparency and fair data processing. Valeria Zitz, Matthias Wölfel, Rainer Hoffmann 0004 |
CW | 2 |
| 2020 | Accessibility of Different Natural User Interfaces for People with Intellectual DisabilitiesabstractDigital technologies have many advantages for users, such as virtually unlimited access to information, entertainment, and communication. Most modern human-computer interfaces are developed under the assumption that they will be used by a person with typical physical, intellectual, and perceptual abilities. Although some operating systems already include accessibility features, in most cases the effective use of the respective interface can be severely restricted if a person's abilities deviate from this norm. To what extent the class of `natural` user interfaces-including touch, voice and touchless-are accessible to people with intellectual and possibly motor disabilities is an important but not yet investigated question. Therefore, this paper investigates the current accessibility of these three interface types. First, we conducted a field study to Figure out how the target group interacts with these types of interfaces in general. Second, quantitative data on cognitive and motor skills was collected using parts of the Questionnaire for Observing Communicative Skills Revision (OCS-R) which is widely used in institutions for people with disabilities in Germany. Finally, the accessibility of each interface type was analyzed with the help of the data obtained from the questionnaire and an expert survey, which determined the important and unimportant skills required for each interface. These findings show how usable different types of natural user interfaces are for this target group. Melinda C. Braun, Matthias Wölfel, Gregor Renner, Christian Menschik |
CW | 2 |
| 2020 | Differences in the Uncanny Valley between Head-Mounted Displays and MonitorsabstractThe uncanny valley describes a relationship between the degree of the emotional response with respect to a character's resemblance to an actual human being. It has been a topic for several decades and has been discussed by scholars of different disciplines and in various aspects such as robotics, 3D computer animations, interactive applications, and even lifelike dolls. With the increasing popularity of photo-realistic computer animation and telepresence applications, we are more and more exposed to characters that might be affected by the uncanny valley effect. Recent research suggests that the output medium, such as a monitor or head-mounted display, can have a significant effect on how we perceive given visualizations. In relation to the uncanny valley, we observe a similar effect in our study: characters appear significantly more eerie as well as humanlike when watched on a head-mounted virtual reality headset instead of a monitor. The amplification we see is similar to the findings that the uncanny valley effect is more pronounced when the respective characters are in motion. Daniel Hepperle, Hannah Ödell, Matthias Wölfel |
CW | 3 |
| 2020 | Visualizing Voice Characteristics with Type Design in Closed Captions for ArabicabstractDiversification of fonts in video captions based on the voice characteristics, namely loudness, speed and pauses, can affect the viewer receiving the content. This study evaluates a new method, WaveFont, which visualizes the voice characteristics for captions in an intuitive way. The study was specifically designed to test captions, which aims to add a new experience for Arabic viewers. The results indicate that our visualization is comprehensible and acceptable and provides significant added value-for hearing-impaired and non-hearing impaired participants: Significantly more participants stated that WaveFont improves their watching experience more than standard captions. Tim Schlippe, Shaimaa Alessai, Ghanimeh El-Taweel, Matthias Wölfel, Wajdi Zaghouani |
CW | 4 |
| 2019 | How does Augmented Reality Improve the Play Experience in Current Augmented Reality Enhanced Smartphone Games?abstractThis paper investigates the current state of handheld augmented reality (AR) gaming apps available on the App Store (iOS) and the Play Store (Android). To be able to directly compare the differences between games played with and without AR, only games in which the AR mode can be switched on/off were investigated. Because the main scope of this paper is on the evaluation of the experience provided by AR, parts of the game experience questionnaire (GEQ) have been included in the empirical study. It showed that AR has big potential to improve immersion or flow in the game-play. This paper also identifies differences in the implementation of AR features and investigates how and what parameter in the GEQ can be positively influenced. Matthias Wölfel, Melinda C. Braun, Sandra Beuck |
CW | 1 |
| 2019 | 2D, 3D or speech? A case study on which user interface is preferable for what kind of object interaction in immersive virtual reality
Daniel Hepperle, Yannick Weiss, Andreas Sieß, Matthias Wölfel |
Comput. Graph. | 4 |
| 2019 | User color temperature preferences in immersive virtual realities
Andreas Sieß, Matthias Wölfel |
Comput. Graph. | 2 |
| 2018 | Color Preference Differences between Head Mounted Displays and PC ScreensabstractRecently virtual reality (VR) applications are shifting from professional use cases to more entertainment-centered approaches. Therefore aesthetic aspects in virtual environments gain in relevance. This paper examines the influence of different color determining parameters on user perception habits between head mounted displays (HMD) and computer screens. We conducted an empirical study with 50 persons that were asked to adjust the color temperature, saturation and contrast according to their personal preferences using a HMD as well as a computer screen, respectively. For cross validation we tested a second user group of 36 persons that were asked to adjust the color temperature exclusively. By using a set of five different panorama images-each of them representing an exemplary scenario-we have found that color perception differs significantly. This depends on the used output device as well as gender: i.e. females preferred a significantly colder color scheme in VR compared to their preferences on the computer screen. Furthermore they also chose a significant colder color scheme on the HMD compared to their male counterparts. Our findings demonstrate that content created for conventional screens can not simply be transferred to immersive virtual environments but for optimal results needs reevaluation of its visual aesthetics. Andreas Sieß, Matthias Wölfel, Nico Haffner |
CW | 2 |
| 2018 | What User Interface to Use for Virtual Reality? 2D, 3D or Speech-A User StudyabstractIn virtual reality different demands on the user interface have to be addressed than on classic screen applications. That's why established strategies from other digital media cannot be transferred unreflected and at least adaptation is required. So one of the leading questions is: which form of interface is preferable for virtual reality? Are 2D interfaces—that are mostly used in combination with mouse or touch interactions— the means of choice, although they do not use the medium's full capabilities? What about 3D interfaces that can be naturally integrated into the virtual space? And last but not least: are speech interfaces, the fastest and most natural form of human interaction/communication, which have recently established themselves in other areas (e.g. digital assistants), ready to conquer the world of virtual reality? To answer these question this work compares these three approaches based on a quantitative user study and highlights advantages and disadvantages of the respective interfaces for virtual reality applications. Yannick Weiss, Daniel Hepperle, Andreas Sieß, Matthias Wölfel |
CW | 4 |
| 2018 | Effects of Electrical Pain Stimuli on Immersion in Virtual RealityabstractThe ultimate goal of virtual realty is to create a simulated world around us which is indistinguishable from the physical world as we know it. In such an environment our actions could lead to severe effects on our body. What would happen if one gets hit by a bullet, car or lightning? How would the felt pain change our perception of the virtual environment? It turns out that the influence of nociception (pain) on the human perception in virtual environments is not well covered in the scientific literature besides pain control/management. The goal of this publication is to investigate the influence of pain stimuli on immersion as well as decision making and to foster research and discussion in this direction. Matthias Wölfel, Joey Schubert |
CW | 1 |
| 2017 | Willingness of Distracted Smartphone Users on the Move to be Interrupted in Potentially Dangerous SituationsabstractThe use of smart devices has become an integrated part of our everyday life. Communication is now possible any place and any time. The distraction caused by these devices, however, can lead to potentially dangerous situations. To mitigate these situations, various researchers have proposed and developed solutions to analyze the environment and to alert the user if a situation is evaluated dangerous. While seeking technical solutions, the concerns of the users are usually not addressed. With our study we put the needs of the user into focus and investigated the acceptance and other issues of such applications. We found that environment awareness for users receiving warnings can be significantly improved compared to those who do not . We also found that 56% of the participants who tested a safety application and received warnings would use it, while only 46% would use it a-priori without testing it. We can also confirm that females are more willing to use a safety application than men. Sandra Beuck, Alexander Scheurer, Matthias Wölfel |
CW | 3 |
| 2017 | Do you feel what you see?: multimodal perception in virtual realityabstractThis paper discusses how different physically existing materials can be mapped on virtual textures in mixed reality environments by carrying out an explorative user study (n=101). For physical materials-in form of 3d trackable and moveable cubes-acrylic, wood and aluminum have been used. The virtual textures convey the impression of ceramic, fabric, glass, leather, paper, wood, acrylic, quartz, granite and aluminum. The study reveals which virtual textures match well with the different virtual textures and which do not match at all. Daniel Hepperle, Matthias Wölfel |
VRST | 2 |
| 2015 | Hybrid City Lighting - Improving Pedestrians' Safety through Proactive Street LightingabstractAlthough digital revolution has pervaded almost every part of daily life, cities remained seemingly analogue and furthermore inhabitants are mostly excluded from the digital layer. By replacing timeworn light bulbs with a projector linked with an intelligent sensor array we expect to increase social interaction within open spaces by creating a more engaging way through the city and reduce the use of distracting and separating mobile devices. We propose several applications to support the weakest traffic participants -- pupils -- having a safe way to school and back, prevent accidents and guide pedestrians through an increasingly complex city. We want to provide a more economical, safer and smarter way to lighten up the way through future cities. Andreas Sieß, Kathleen Hübel, Daniel Hepperle, Andreas Dronov, Christian Hufnagel, Julia Aktun, Matthias Wölfel |
CW | 7 |
| 2015 | To Be There, or Not to Be There, that is the QuestionabstractVirtual words, independent of their realizations, let us experience a person's sense of Being There, a form of spatial immersion dubbed presence. Where we are and even who we are is the result of living in the real world as well as a multitude of virtual worlds. This visceral feeling depends on several technical dimensions, which have been widely discussed, including the realism of the virtual environment and the embodiment of oneself and the others. But presence and belonging are preliminary not a question of technique but a question of the provided format, the symbolic spaces and of social interaction. A virtual world is always a managed space and a managed me with all its complications. It is about the basic conceptions of a conditio humana. Matthias Wölfel, Ulrich Gehmann |
CW | 1 |
| 2015 | Responsive Type - Introducing Self-Adjusting Graphic CharactersabstractIn this publication we introduced a radical new concept of perceiving information in written form: Until today, after the layout process has finished, the shape of a single typographic character is treated as an unchangeable property. But due to current sensor and display technologies this does not have to be necessary anymore. We propose, to change the shape of each character according to various conditions of the user while maintaining its individual and discriminative identity. In contrast to fixed character shapes, responsive character shapes are not limited to be adjusted for a particular output media, but offers the possibilities to adjust to personal conditions such as age, position and defects of sight. This seemingly simple difference, however, questions information brokering fundamentally. It expands the possibilities, besides the engagement of the media, to perceive information within the context and the conditions of the individual. Through experiments and user studies, we demonstrated that our proposed approach, dubbed responsive type, is widely accepted and readability as well as legibility can be improved. Matthias Wölfel, Angelo Stitz |
CW | 1 |
| 2014 | Liberated FormatizationabstractThe contribution tries to sketch the development of a certain kind of technology, namely user-specific computational technologies which promise an increase in gaining individual freedom. It was a promise that increasingly led to a predominance of managed 'freedoms' and hence, to an increased formatization of technical tools and individual perception alike. The overall effect was that a seemingly increase in individual user-freedom was accompanied by a de facto-increase in preformatted devices for achieving it, and hence, led to the actual decrease of this very freedom, despite the latter seemed to have increased steadily. All in all, the entire process to be examined here increasingly based upon formats of both an economic and technological nature. In its final outcomes, it led to an almost complete loss of freedom for the official beneficiary of such devices, the user. Ulrich Gehmann, Marco Zampella, Matthias Wölfel |
CW | 3 |
| 2014 | Interacting with Ads in Hybrid Urban SpaceabstractIn stark contrast to online advertising campaigns, advertisement in the urban space has lost attention and seems to be stuck in the Gutenberg era. The emergence of hybrid urban spaces, however, allows for novel possibilities to bring back customers' attention and interest. In this publication we review current interactive advertisement campaigns, investigate the use of implicit (age, gender, location) and explicit (2D and 3D gestures) interactions of the user to adjust the ad, and discuss novel questions and responsibilities that are driven by these new advertisement formats. To evaluate different aspects on the behavior and acceptability of such novel kind of advertisement we have built a prototypical system and put it into a shopping mall. The conducted user study includes 98 random visitors of the mall who have first tested the system and then filled out a questionnaire. Matthias Wölfel |
CW | 1 |
| 2014 | Increasing Customers' Attention using Implicit and Explicit Interaction in Urban AdvertisementabstractOnline advertising campaigns are gaining customers' attention in comparison to advertisement campaigns in the urban space. How to bring back users' attention to advertisement in the urban space by using implicit and explicit interactions of the user is investigated in this publication. We have used age, gender and position estimates to automatically customize the advertisement campaign and 3D gestures to allow the customer to interact with the shown content. To evaluate the overall acceptability and particular aspects of such kind of targeted and interactive advertisement we have developed a prototypical implementation and placed it into a crowded shopping mall. In total 98 random visitors of the mall have experienced the system and answered a questionnaire afterwards. Matthias Wölfel, Luigi Bucchino |
ICMI | 1 |
| 2009 | Speaker identification using warped MVDR cepstral featuresabstractIt is common practice to use similar or even the same feature extraction methods for automatic speech recognition and speaker identification. While the front-end for the former requires to preserve phoneme discrimination and to compensate for speaker differences to some extend, the front-end for the latter has to preserve the unique characteristics of individual speakers. It seems, therefore, contradictory to use the same feature extraction methods for both tasks. Starting out from the common practice we propose to use warped minimum variance distortionless response (MVDR) cepstral coefficients, which have already been demonstrated to perform superior for automatic speech recognition in particular under adverse conditions. Replacing the widely used mel-frequency cepstral coefficients by WMVDR cepstral coefficients improves the speaker identification accuracy by up to 24 % relative. We found that the optimal choice of the model order within the WMVDR framework differs between speech recognition and speaker recognition, confirming our intuition that the two different tasks indeed require different feature extraction strategies. 1. Matthias Wölfel, Qin Jin, Tanja Schultz |
INTERSPEECH | 1 |
| 2009 | Signal adaptive spectral envelope estimation for robust speech recognition
Matthias Wölfel |
Speech Commun. | 1 |
| 2009 | Enhanced Speech Features by Single-Channel Joint Compensation of Noise and ReverberationabstractFor a natural verbal communication between humans and machines, automatic speech recognition, which works reasonably well on recordings captured with mid- or far-field microphones, is essential. While a lot of research and development are devoted to address one of the two distortions frequently encountered in mid- and far-field sound pickup, namely noise or reverberation, less effort has been undertaken to jointly combat both kinds of distortions. In our view, however, this is essential to further reduce the demolishing effect by moving the microphone away from the speaker's mouth because in real environments both kinds of distortions are present. In this paper, we propose a first step into this direction by integrating an estimate of the reverberation energy derived by an auxiliary model based on multistep linear prediction, into a framework, which, so far tracks and removes nonstationary additive distortion by particle filters in a low-dimension logarithmic power frequency domain. On actual recordings with different speaker-to-microphone distances, we observe that combating, in the feature space, either nonstationary noise or reverberation alone, on a single channel, is already able to improve speech recognition performance before and after acoustic model adaptation. Furthermore, we observe that a simple concatenation of techniques addressing either additive noise or reverberation can further improve the accuracy in some cases. Last but not least, we demonstrate that the joint estimation and removal of both kinds of distortions, as proposed in this publication, further improve the accuracy of the text output. Matthias Wölfel |
IEEE Trans. Speech Audio Process. | 1 |
| 2008 | Predictedwalk with correlation in particle filter speech feature enhancement for robust automatic speech recognitionabstractPrevious particle filter feature enhancement techniques for robust automatic speech recognition have ignored the fact that neighbored spectral bins are correlated. In those cases, the spectral bins have been treated as uncorrelated components in the sampling stage of the particle filter. In this publication we propose to consider the correlation between the individual spectral bins by correlating the random variation after a predicted walk realized by a linear prediction matrix. Experiments on artificially added dynamic noise at different signal to noise ratios as well as on actual recordings with different speaker to microphone distances show reasonable word error rate reduction before and after acoustic model adaptation of the automatic speech recognition system. Matthias Wölfel |
ICASSP | 1 |
| 2008 | Integration of the predictedwalk model estimate into the particle filter frameworkabstractDistortion robustness is one of the most significant problems in automatic speech recognition. While a lot of research in speech feature enhancement in automatic recognition has focused on stationary distortions, most of the observed distortions are non-stationary. To cope with the non-stationary behavior, just recently, various particle filter approaches have been proposed to track the non-stationary distortions on speech features in logarithmic spectral or cepstral domain. Most of those techniques rely on the prediction of the noise evolution model by a linear prediction matrix. The current estimation of the linear prediction matrix, however, needs noise only observations which have to be either given a priori or to be detected by voice activity detection. This makes it impossible to adapt the linear prediction matrix to the dynamics of the noise on speech regions. In this publication we propose to estimate or update the linear prediction matrix directly on the noisy speech observations. This is possible within the particle filter framework by weighting the different noisy estimates (particles) due to their likelihood in the estimation equation of the linear prediction matrix. Speech recognition experiments on actual recordings with different speaker to microphone distances confirm the soundness of the proposed approach. Matthias Wölfel |
ICASSP | 1 |
| 2008 | A Comparative Cross-Domain Study of the Occurrence of Laughter in Meeting and Seminar Corpora
Susanne Burger, Kornel Laskowski, Matthias Wölfel |
LREC | 3 |
| 2008 | Simultaneous machine translation of german lectures into english: Investigating research challenges for the futureabstractAn increasingly globalized world fosters the exchange of students, researchers or employees. As a result, situations in which people of different native tongues are listening to the same lecture become more and more frequent. In many such situations, human interpreters are prohibitively expensive or simply not available. For this reason, and because first prototypes have already demonstrated the feasibility of such systems, automatic translation of lectures receives increasing attention. A large vocabulary and strong variations in speaking style make lecture translation a challenging, however not hopeless, task. The scope of this paper is to investigate a variety of challenges and to highlight possible solutions in building a system for simultaneous translation of lectures from German to English. While some of the investigated challenges are more general, e.g. environment robustness, other challenges are more specific for this particular task, e.g. pronunciation of foreign words or sentence segmentation. We also report our progress in building an end-to-end system and analyze its performance in terms of objective and subjective measures. Matthias Wölfel, Muntsin Kolss, Florian Kraft, Jan Niehues, Matthias Paulik, Alex Waibel |
SLT | 1 |
| 2008 | On maximum mutual information speaker-adapted training
John W. McDonough, Matthias Wölfel, Emilian Stoimenov |
Comput. Speech Lang. | 2 |
| 2007 | Minimum mutual information beamforming for simultaneous active speakersabstractIn this work, we address an acoustic beamforming application where two speakers are simultaneously active. We construct one subband domain beamformer in generalized sidelobe canceller (GSC) configuration for each source. In contrast to normal practice, we then jointly adjust the active weight vectors of both GSCs to obtain two output signals with minimum mutual information (MMI). In order to calculate the mutual information of the complex subband snapshots, we consider four probability density functions (pdfs), namely the Gaussian, Laplace, K0and Г pdfs. The latter three belong to the class of super-Gaussian density functions that are typically used in independent component analysis as opposed to conventional beamforming. We demonstrate the effectiveness of our proposed technique through a series of far-field automatic speech recognition experiments on data from the PASCAL Speech Separation Challenge. In the experiments, the delay-and-sum beamformer achieved a word error rate (WER) of 70.4 %. The MMI beamformer under a Gaussian assumption achieved 55.2 % WER which was further reduced to 52.0 % with a K0pdf, whereas the WER for data recorded with close-talking microphone was 21.6 %. Ken'ichi Kumatani, Uwe Mayer, Tobias Gehrig, Emilian Stoimenov, John W. McDonough, Matthias Wölfel |
ASRU | 6 |
| 2007 | Overcoming the Vector Taylor Series Approximation in Speech Feature Enhancement - A Particle Filter ApproachabstractWe present a simple, fast and previously unreported noise compensation method for particle filter (PF) based speech feature enhancement, which outperforms the vector Taylor series noise compensation method used by current PF approaches in terms of speed as well as word error rate. Furthermore, we devise a fast acceptance test that overcomes the particle decimation problem associated with PFs for speech feature enhancement, which makes the particle filter approach computationally more efficient. Friedrich Faubel, Matthias Wölfel |
ICASSP (4) | 2 |
| 2007 | Considering Uncertainty by Particle Filter Enhanced Speech Features in Large Vocabulary Continuous Speech RecognitionabstractThe goal of noise compensation techniques is the perfect reconstruction of clean features. Unfortunately, the reconstructed features can not be assumed to be perfect. Therefore, to improve performance, the uncertainty of enhanced speech features should be propagated into the hidden Markov model of automatic speech recognition systems. This paper shows how to jointly estimate the noise and the uncertainty (expressed by the variance) by particle filters in the logarithmic Mel power domain and how to propagate the uncertainty through the front-end into the hidden Markov model. In the experimental section, improvements in word accuracy of a large vocabulary continuous speech recognition system are presented. Matthias Wölfel, Friedrich Faubel |
ICASSP (4) | 1 |
| 2007 | The ISL 2007 English speech transcription system for european parliament speechesabstractThe project Technology and Corpora for Speech to Speech Translation (TC-STAR) aims at making a break-through in speech-to-speech translation research, significantly reducing the gap between the performance of machines and humans at this task. Technological and scientific progress is driven by periodic, competitive evaluations within the project. In this paper we describe the ISL speech transcription system for English European Parliament speeches with which we participated in the third TC-STAR evaluation campaing in the spring of 2007. The improvements over last year’s system originate from a recognition hypotheses based segmentation, the utilization of unsupervised in-domain training material, a modified cross-system adaptation and combination scheme, and the enhancement of the language model through the use of web based training material. Sebastian Stüker, Christian Fügen, Florian Kraft, Matthias Wölfel |
INTERSPEECH | 4 |
| 2007 | Computer-supported human-human multilingual communication
Alex Waibel, Keni Bernardin, Matthias Wölfel |
INTERSPEECH | 3 |
| 2007 | Channel selection by class separability measures for automatic transcriptions on distant microphonesabstractChannel selection is important for automatic speech recognition as the signal quality of one channel might be significantly better than those of the other channels and therefore, microphone array or blind source separation techniques might not lead to improvements over the best single microphone. The mayor challenge, however, is to find this particular channel who is leading to the most accurate classification. In this paper we present a novel channel selection method, based on class separability, to improve multi-source far distance speech-totext transcriptions. Class separability measures have the advantage, compared to other methods such as the signal to noise ratio (SNR), that they are able to evaluate the channel quality on the actual features of the recognition system. We have evaluated on NISTs RT-07 development set and observe significant improvements in word accuracy over SNR based channel selection methods. We have also used this technique in NISTs RT-07 evaluation. 1. Matthias Wölfel |
INTERSPEECH | 1 |
| 2007 | Humanoid robot noise suppression by particle filters for improved automatic speech recognition accuracyabstractAutomatic speech recognition on a humanoid robot is exposed to numerous known noises produced by the robot's own motion system and background noises such as fans. Those noises interfere with target speech by an unknown transfer function at high distortion levels, since some noise sources might be closer to the robot's microphones than the target speech sources. In this paper we show how to remedy those distortions by a speech feature enhancement technique based on the recently proposed particle filters. A significant increase of recognition accuracy could be reached at different distances for both engine and background noises. Florian Kraft, Matthias Wölfel |
IROS | 2 |
| 2007 | Adaptive Beamforming With a Minimum Mutual Information CriterionabstractIn this paper, we consider an acoustic beamforming application where two speakers are simultaneously active. We construct one subband-domain beamformer in generalized sidelobe canceller (GSC) configuration for each source. In contrast to normal practice, we then jointly optimize the active weight vectors of both GSCs to obtain two output signals with minimum mutual information (MMI). Assuming that the subband snapshots are Gaussian-distributed, this MMI criterion reduces to the requirement that the cross-correlation coefficient of the subband outputs of the two GSCs vanishes. We also compare separation performance under the Gaussian assumption with that obtained from several super-Gaussian probability density functions (pdfs), namely, the Laplace$K_0$and$\Gamma$pdfs. Our proposed technique provides effective nulling of the undesired source, but without the signal cancellation problems seen in conventional beamforming. Moreover, our technique does not suffer from the source permutation and scaling ambiguities encountered in conventional blind source separation algorithms. We demonstrate the effectiveness of our proposed technique through a series of far-field automatic speech recognition experiments on data from the PASCAL Speech Separation Challenge (SSC). On the SSC development data, the simple delay-and-sum beamformer achieves a word error rate (WER) of 70.4%. The MMI beamformer under a Gaussian assumption achieves a 55.2% WER, which is further reduced to 52.0% with a$K_0$pdf, whereas the WER for data recorded with a close-talking microphone is 21.6%. Ken'ichi Kumatani, Tobias Gehrig, Uwe Mayer, Emilian Stoimenov, John W. McDonough, Matthias Wölfel |
IEEE Trans. Speech Audio Process. | 6 |
| 2006 | Coupling particle filters with automatic speech recognition for speech feature enhancementabstractThis paper addresses robust speech feature extraction in combination with statistical speech feature enhancement and couples the particle filter to the speech recognition hypotheses. To extract noise robust features the Fourier transformation is replaced by the warped and scaled minimum variance distortionless response spectral envelope. To enhance the features, particle filtering has been used. Further, we show that the robust extraction and statistical enhancement can be combined to good effect. One of the critical aspects in particle filter design is the particle weight calculation which is traditionally based on a general, time independent speech model approximated by a Gaussian mixture distribution. We replace this general, time independent speech model by time- and phoneme-specific models. The knowledge of the phonemes to be used is obtained by the hypothesis of a speech recognition system, therefore establishing a coupling between the particle filter and the speech recognition system which have been treated as independent components in the past. Index Terms: particle filters, automatic speech recognition, speech feature enhancement, phoneme-specific Friedrich Faubel, Matthias Wölfel |
INTERSPEECH | 2 |
| 2006 | Advances in lecture recognition: the ISL RT-06s evaluation systemabstractThis paper describes the 2006 lecture recognition system developed at the Interactive Systems Laboratories (ISL), for individual head-microphone (IHM), single distant microphone (SDM), and multiple distant microphones (MDM) conditions.It was evaluated in RT-06S rich transcription meeting evaluation sponsored by the US National Institute of Standards and Technologies (NIST).We describe the principal differences between our current system and those submitted in previous years, namely, improved acoustic and language models, cross adaptation between systems with different front-ends and phoneme sets, and the use of various automatic speech segmentation algorithms.Our system achieved word error rates of 38.5% (53.4%) and 22.9% (32.2%), respectively, on the MDM and IHM conditions of the RT-05S (RT-06S) lecture evaluation set. Christian Fügen, Matthias Wölfel, John W. McDonough, Shajith Ikbal, Florian Kraft, Kornel Laskowski, Mari Ostendorf, Sebastian Stüker, Ken'ichi Kumatani |
INTERSPEECH | 2 |
| 2006 | Tracking and beamforming for multiple simultaneous speakers with probabilistic data association filtersabstractIn prior work, we developed a speaker tracking system based on an extended Kalman filter using time delays of arrival (TDOAs) as acoustic features. While this system functioned well, its util-ity was limited to scenarios in which a single speaker was to be tracked. In this work, we remove this restriction by generalizing the IEKF, first to a probabilistic data association filter, which in-corporates a clutter model for rejection of spurious acoustic events, and then to a joint probabilistic data association filter (JPDAF), which maintains a separate state vector for each active speaker. In a set of experiments conducted on seminar and meeting data, the JPDAF speaker tracking system reduced the multiple object track-ing errror from 20.7 % to 14.3 % with respect to the IEKF system. In a set of automatic speech recognition experiments conducted on the output of a 64 channel microphone array which was beam-formed using automatic speaker position estimates, applying the JPDAF tracking system reduced word error rate from 67.3 % to 66.0%. Moreover, the word error rate on the beamformed output was 13.0 % absolute lower than on a single channel of the array. Index Terms: acoustic source localization, Kalman filter, person tracking, far-field speech recognition, microphone arrays Tobias Gehrig, Ulrich Klee, John W. McDonough, Shajith Ikbal, Matthias Wölfel, Christian Fügen |
INTERSPEECH | 5 |
| 2006 | Cross-system adaptation and combination for continuous speech recognition: the influence of phoneme set and acoustic front-endabstractAbstract Cross-system adaptation and system combination methods,such as ROVER and confusion network combination, areknown to lower the word error rate of speech recognitionsystems. They require the training of systems that are rea-sonably close in performance but at the same time produceoutput that differs in its errors. This provides complemen-taryinformationwhichleadstoperformanceimprovements.In this paper we demonstrate the gains we have seen withcross-systemadaptationandsystemcombinationontheEn-glish EPPS and RT0-05S lecture meeting task. We obtainedthe necessary varying systems by using different acous-tic front-ends and phoneme sets on which our models arebased. Inasetofcontrastiveexperimentsweshowtheinflu-ence that the exchange of the components has on adaptationand system combination.Index Terms: automatic speech recognition, system com-bination, cross adaptation, EPPS, RT-05S. 1. Introduction In state-of-the-art speech recognition systems it is commonpractice to use multi-pass systems with adaptation of theacoustic model in-between passes. The adaptation aims atbetter fitting the system to the speakers and/or acoustic en-vironmentsfoundinthetestdata. Itisusuallyperformedona by-speaker basis, obtained either from manual speaker la-bels or automatic clustering methods. Common adaptationmethods try to transform either the models used in a systemor the features to which the models are applied.Three adaptation methods that can be found in manystate-of-the-art systems are Maximum Likelihood LinearRegression (MLLR) [1], a model transformation, Vo-cal Tract Length Normalization (VTLN) [2] and feature-space constrained MLLR (fMLLR) [3], two feature-transformation methods. Adaptation is performed in an un-supervisedmanner,suchthattheerror-pronehypothesesob-tainedfromthepreviousdecodingpassaretakenasthenec-essary reference for adaptation. Generally, the word errorrates of the hypotheses obtained from the adapted systemsarelowerthanthoseforhypothesesonwhichtheadaptationwas performed. This sequences of adaption and decodingmake it possible to incrementally improve the performanceof the recognition system. Unfortunately, this loop of adap-tation and decoding does not always lead to significant im-provements. Often, after two or three stages of adapting asystem on its own output, no more gains can be obtained.This problem can be overcome by adapting a system Sebastian Stüker, Christian Fügen, Susanne Burger, Matthias Wölfel |
INTERSPEECH | 4 |
| 2006 | Multi-source far-distance microphone selection and combination for automatic transcription of lecturesabstractIn this work, we present our progress in multi-source far field automatic speech-to-text transcription for lecture speech. In particular, we show how the best of several far field channels can be selected based on a signal-to-noise ratio criterion, and how the signals from multiple channels can be combined at either the waveform level using blind channel combination or at the hypothesis level using confusion network techniques to improve the accuracy of a far field lecture transcription system. Using the techniques described here, we ran a series of experiments on the test set used by the US National Institute of Standards and Technologies for the RT-05S evaluation. For the multiple distant microphones (MDM) task of RT-05S, our system achieved a word error rate of 38.5 % which represents an improvement of over 13 % absolute compared to the best reported results in the RT-05S evaluation. Index Terms: far-distance, automatic speech recognition 1. Matthias Wölfel, Christian Fügen, Shajith Ikbal, John W. McDonough |
INTERSPEECH | 1 |
| 2006 | Audio-visual perception of a lecturer in a smart seminar room
Rainer Stiefelhagen, Keni Bernardin, Hazim Kemal Ekenel, John W. McDonough, Kai Nickel, Michael Voit, Matthias Wölfel |
Signal Process. | 7 |
| 2005 | Automatic Speech Activity Detection, Source Localization, and Speech Recognition on the Chil Seminar CorpusabstractTo realize the long-term goal of ubiquitous computing, technological advances in multi-channel acoustic analysis are needed in order to solve several basic problems, including speaker localization and tracking, speech activity detection (SAD) and distant-talking automatic speech recognition (ASR). The European Commission integrated project CHIL, “ Computers in the Human Interaction Loop”, aims to make significant advances in these three technologies. In this work, we report the results of our initial automatic source localization, speech activity detection, and speech recognition experiments on the CHIL seminar corpus, which is comprised of spontaneous speech collected by both near- and far-field microphones. In addition to the audio sensors, the seminars were also recorded by calibrated video cameras. This simultaneous audio-visual data capture enables the realistic evaluation of component technologies as was never possible with earlier data bases. Dusan Macho, Jaume Padrell, Alberto Abad, Climent Nadeu, Javier Hernando, John W. McDonough, Matthias Wölfel, Ulrich Klee, Maurizio Omologo, Alessio Brutti, Piergiorgio Svaizer, Gerasimos Potamianos, Stephen M. Chu |
ICME | 7 |
| 2005 | Frame based model order selection of spectral envelopesabstractSpectral envelopes, using (warped or perceptual) linear prediction or minimum variance distortionless response for the underlying linear parametric model, are widely used in speech recognition systems where the frequency resolution, namely the model order (MO), of the spectrum is kept constant. Modeling different types of phonemes such as vowels or fricatives with the same frequency resolution might not lead to the best possible performance. This could be due to the fact that important parts of various phonemes lie in different frequency regions, that the fundamental frequency varies for different speakers or because of a high variance in the signal to noise ratio. To address this problem we propose to vary the MO frame by frame according to a control factor. In our case, the control factor could be either a relation of autocorrelation coefficients or the spectral entropy. Experimental results on the Translanguage English Database show an improvement by 2.4 % relative in word error rate compared to the fixed MO and 4.2 % relative to the traditional Mel-frequency cepstral coefficients. 1. Matthias Wölfel |
INTERSPEECH | 1 |
| 2005 | Combining multi-source far distance speech recognition strategies: beamforming, blind channel and confusion network combinationabstractInterest within the automatic speech recognition (ASR) research community has recently focused on the recognition of speech captured with a microphone located in the medium field, rather than being mounted on a headset and positioned next to the speaker’s mouth. The capacity to recognize such speech is a primary requirement in making ASR a viable modality for socalled ubiquitous computing. This is a natural application for multiple microphones whose signals can be combined in different ways: On the signal side, combination can be accomplished by beamforming techniques using a microphone array or by blind source separation. On the word hypothesis side, combination can be achieved through confusion network combination. In this work, we compare the effectiveness of the several combination techniques, and compare their performance to that achieved with a close talking microphone. Matthias Wölfel, John W. McDonough |
INTERSPEECH | 1 |
| 2004 | A cepstral domain maximum likelihod beamformer for speech recognitionabstractRecent work by Seltzer [1] indicates that classical approaches to beamforming, minimizing output power while enforcing a distortionless constraint, do not yield optimal results in terms of word error rate (WER) on speech recognition task. This problem can be traced back to the mismatch between the target criterion of classical adaptive beamformers, which is optimization of the signal to noise ratio, and the actual target criterion, which is the reduction of the recognizer’s WER. Following an approach by Seltzer [1] we therefore investigate the performance of an alternative error criterion, which attempts to optimize the beamformer weights, so as to improve the likelihoods along the recognizer’s Viterbi path for each utterance. This criterion matches the goal of lower WERs more closely and therefore leads to better recognition results. Dominik Raub, John W. McDonough, Matthias Wölfel |
INTERSPEECH | 3 |
| 2004 | Speaker dependent model order selection of spectral envelopesabstractThis work introduces a maximum-likelihood based model order (MO) selection technique for spectral envelopes to apply speaker dependent adaptation in the feature-space similar to vocal tract length normalization. Speech recognition systems based on spectral envelopes are using a fixed MO for the underlying linear parametric model. Using a fixed MO over different speakers or channels might not be optimal. To address this problem we investigated the use of warped and scaled minimum variance distortionless response spectral estimation techniques with speaker dependent MOs based on a maximum-likelihood criteria. Comparing experimental results on the Translanguage English Database we can show an improvement by 1,9 % relative compared to the word error rate by the fixed MO and 3,5 % relative to the traditional Mel-frequency cepstral coefficients. 1. Matthias Wölfel |
INTERSPEECH | 1 |
| 2003 | Minimum variance distortionless response on a warped frequency scaleabstractIn this work we propose a time domain technique to estimate an all-pole model based on the minimum variance distortionless response (MVDR) using a warped short time frequency axis such as the Mel scale. The use of the MVDR eliminates the overemphasis of harmonic peaks typically seen in medium and high pitched voiced speech when spectral estimation is based on linear prediction (LP). Moreover, warping the frequency axis prior to MVDR spectral estimation ensures more parameters in the spectral model are allocated to the low, as opposed to high, frequency regions of the spectrum, thereby mimicking the human auditory system. In a series of speech recognition experiments on the Switchboard Corpus (spontaneous English telephone speech), the proposed approach achieved a word error rate (WER) of 32.1% for female speakers, which is clearly superior to the 33.2% WER obtained by the usual combination of Mel warping and linear prediction. Matthias Wölfel, John W. McDonough, Alex Waibel |
INTERSPEECH | 1 |