EDBT 2026 Demo / reviewers in the wild / expert
Ahmed Hussen Abdelaziz
dblp:120/8292 · also Ahmed Serag Eldin Hussen Abdelaziz
· DBLP profile ↗
37ranked-venue papers
16as first author
15since 2021 · last 2025
0000-0001-8027-4666ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 29 · 11 first-author · 13 since 2021Artificial intelligence and machine learning · 23 · 8 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Speaker-IPL: Unsupervised Learning of Speaker Characteristics with i-Vector based Pseudo-LabelsabstractIterative self-training, or iterative pseudo-labeling (IPL)—using an improved model from the current iteration to provide pseudo-labels for the next iteration—has proven to be a powerful approach to enhance the quality of speaker representations. Recent applications of IPL in unsupervised speaker recognition start with representations extracted from very elaborate self-supervised methods (e.g., DINO). However, training such strong self-supervised models is not straightforward (they require hyper-parameter tuning and may not generalize to out-of-domain data) and, moreover, may not be needed at all. To this end, we show that the simple, well-studied, and established i-vector generative model is enough to bootstrap the IPL process for the unsupervised learning of speaker representations. We also systematically study the impact of other components on the IPL process, which includes the initial model, the encoder, augmentations, the number of clusters, and the clustering algorithm. Remarkably, we find that even with a simple and significantly weaker initial model like i-vector, IPL can still achieve speaker verification performance that rivals state-of-the-art methods. Zakaria Aldeneh, Takuya Higuchi, Jee-Weon Jung, Ahmed Hussen Abdelaziz, Shinji Watanabe 0001, Tatiana Likhomanenko, Barry-John Theobald |
ICASSP | 6 |
| 2025 | Exploring Prediction Targets in Masked Pre-Training for Speech Foundation ModelsabstractSpeech foundation models, such as HuBERT and its variants, are pre-trained on large amounts of unlabeled speech data and then used for a range of downstream tasks. These models use a masked prediction objective, where the model learns to predict information about masked input segments from the unmasked context. The choice of prediction targets in this framework impacts their performance on downstream tasks. For instance, models pre-trained with targets that capture prosody learn representations suited for speaker-related tasks, while those pre-trained with targets that capture phonetics learn representations suited for content-related tasks. Moreover, prediction targets can differ in the level of detail they capture. Models pre-trained with targets that encode fine-grained acoustic features perform better on tasks like denoising, while those pre-trained with targets focused on higher-level abstractions are more effective for content-related tasks. Despite the importance of prediction targets, the design choices that affect them have not been thoroughly studied. This work explores the design choices and their impact on downstream task performance. Our results indicate that the commonly used design choices for HuBERT can be suboptimal. We propose approaches to create more informative prediction targets and demonstrate their effectiveness through improvements across various downstream tasks. Takuya Higuchi, He Bai 0013, Ahmed Hussen Abdelaziz, Shinji Watanabe 0001, Alexander I. Rudnicky, Tatiana Likhomanenko, Barry-John Theobald, Zakaria Aldeneh |
ICASSP | 4 |
| 2025 | A Variational Framework for Improving Naturalness in Generative Spoken Language ModelsabstractThe success of large language models in text processing has inspired their adaptation to speech modeling.
However, since speech is continuous and complex, it is often discretized for autoregressive modeling.
Speech tokens derived from self-supervised models (known as semantic tokens) typically focus on the linguistic aspects of speech but neglect prosodic information.
As a result, models trained on these tokens can generate speech with reduced naturalness.
Existing approaches try to fix this by adding pitch features to the semantic tokens.
However, pitch alone cannot fully represent the range of paralinguistic attributes, and selecting the right features requires careful hand-engineering.
To overcome this, we propose an end-to-end variational approach that automatically learns to encode these continuous speech attributes to enhance the semantic tokens.
Our approach eliminates the need for manual extraction and selection of paralinguistic features.
Moreover, it produces preferred speech continuations according to human raters.
Code, samples and models are available at https://github.com/b04901014/vae-gslm. Takuya Higuchi, Zakaria Aldeneh, Ahmed Hussen Abdelaziz, Alexander I. Rudnicky |
ICML | 4 |
| 2025 | DiceHuBERT: Distilling HuBERT with a Self-Supervised Learning Objective
Hyung-Gun Chi, Zakaria Aldeneh, Tatiana Likhomanenko, Ognjen Rudovic, Takuya Higuchi, Shinji Watanabe 0001, Ahmed Hussen Abdelaziz |
INTERSPEECH | 8 |
| 2025 | Adaptive Knowledge Distillation for Device-Directed Speech Detection
Hyung-Gun Chi, Florian Pesce, Wonil Chang, Ognjen Rudovic, Arturo Argueta, Vineet Garg, Ahmed Hussen Abdelaziz |
INTERSPEECH | 8 |
| 2024 | Modality Drop-Out for Multimodal Device Directed Speech Detection Using Verbal and Non-Verbal FeaturesabstractDevice-directed speech detection (DDSD) is the binary classification task of distinguishing between queries directed at a voice assistant versus side conversation or background speech. State-of-the-art DDSD systems use verbal cues, e.g acoustic, text and/or automatic speech recognition system (ASR) features, to classify speech as device-directed or otherwise, and often have to contend with one or more of these modalities being unavailable when deployed in real-world settings. In this paper, we investigate fusion schemes for DDSD systems that can be made more robust to missing modalities. Concurrently, we study the use of non-verbal cues, specifically prosody features, in addition to verbal cues for DDSD. We present different approaches to combine scores and embeddings from prosody with the corresponding verbal cues, finding that prosody improves DDSD performance by upto 8.5% in terms of false acceptance rate (FA) at a given fixed operating point via non-linear intermediate fusion, while our use of modality dropout techniques improves the performance of these models by 7.4% in terms of FA when evaluated with missing modalities during inference time. Gautam Krishna, Sameer Dharur, Ognjen Rudovic, Pranay Dighe, Saurabh Adya, Ahmed Hussen Abdelaziz, Ahmed H. Tewfik |
ICASSP | 6 |
| 2024 | Can you Remove the Downstream Model for Speaker Recognition with Self-Supervised Speech Features?
Zakaria Aldeneh, Takuya Higuchi, Jee-Weon Jung, Skyler Seto, Tatiana Likhomanenko, Ahmed Hussen Abdelaziz, Shinji Watanabe 0001, Barry-John Theobald |
INTERSPEECH | 7 |
| 2024 | Comparative Analysis of Personalized Voice Activity Detection Systems: Assessing Real-World Effectiveness
Sai Srujana Buddi, Satyam Kumar 0001, Utkarsh Oggy Sarawgi, Vineet Garg, Shivesh Ranjan, Ognjen Rudovic, Ahmed Hussen Abdelaziz, Saurabh Adya |
INTERSPEECH | 7 |
| 2024 | ESPnet-SPK: full pipeline speaker embedding toolkit with reproducible recipes, self-supervised front-ends, and off-the-shelf models
Jee-Weon Jung, Wangyou Zhang, Jiatong Shi, Zakaria Aldeneh, Takuya Higuchi, Alex Gichamba, Barry-John Theobald, Ahmed Hussen Abdelaziz, Shinji Watanabe 0001 |
INTERSPEECH | 8 |
| 2024 | Multimodal Large Language Models with Fusion Low Rank Adaptation for Device Directed Speech Detection
Shruti Palaskar, Ognjen Rudovic, Sameer Dharur, Florian Pesce, Gautam Krishna, Aswin Sivaraman, Jack Berkowitz, Ahmed Hussen Abdelaziz, Saurabh Adya, Ahmed H. Tewfik |
INTERSPEECH | 8 |
| 2023 | Less Is More: A Unified Architecture for Device-Directed Speech Detection with Multiple Invocation TypesabstractSuppressing unintended invocation of the device because of the speech that sounds like wake-word, or accidental button presses, is critical for a good user experience, and is referred to as False-Trigger-Mitigation (FTM). In case of multiple invocation options, the traditional approach to FTM is to use invocation-specific models, or a single model for all invocations. Both approaches are sub-optimal: the memory cost for the former approach grows linearly with the number of invocation options, which is prohibitive for on-device deployment, and does not take advantage of shared training data; while the latter is unable to accurately capture acoustic differences across different invocation types. To this end, we propose a Unified Acoustic Detector (UAD) for FTM when multiple invocation options are available on device. The proposed UAD is trained using a multi-task learning framework, where a jointly trained acoustic encoder model is augmented with invocation-specific classification layers. In the context of the FTM task, we show for the first time that using the shared model architecture across invocations (thus, keeping the model size similar to that of a monolithic model used for a single invocation type), we can not only match but largely improve the accuracy of the invocation-specific models. In particular, in the challenging case of touch-based invocation, we obtain 50% and 35% relative improvement in false positive rate at 99% true positive rate, when compared with a singleoutput model for both invocations, and separate models per invocation, respectively. Furthermore, we propose streaming and non-streaming variants of the UAD, and show that they both outperform a traditional ASR-based approach to FTM. Ognjen Rudovic, Wonil Chang, Vineet Garg, Pranay Dighe, Pramod Simha, Jack Berkowitz, Ahmed Hussen Abdelaziz, Sachin Kajarekar, Erik Marchi, Saurabh Adya |
ICASSP | 7 |
| 2022 | Device-Directed Speech Detection: Regularization via Distillation for Weakly-Supervised ModelsabstractWe address the problem of detecting speech directed to a device that does not contain a specific wake-word.Specifically, we focus on audio coming from a touch-based invocation.Mitigating virtual assistants (VAs) activation due to accidental button presses is critical for user experience.While the majority of approaches to false trigger mitigation (FTM) are designed to detect the presence of a target keyword, inferring user intent in absence of keyword is difficult.This also poses a challenge when creating the training/evaluation data for such systems due to inherent ambiguity in the user's data.To this end, we propose a novel FTM approach that uses weakly-labeled training data obtained with a newly introduced data sampling strategy.While this sampling strategy reduces data annotation efforts, the data labels are noisy as the data are not annotated manually.We use these data to train an acoustics-only model for the FTM task by regularizing its loss function via knowledge distillation from an ASR-based (LatticeRNN) model.This improves the model decisions, resulting in 66% gain in accuracy, as measured by equal-error-rate (EER), over the base acoustics-only model.We also show that the ensemble of the LatticeRNN and acousticdistilled models brings further accuracy improvement of 20%. Vineet Garg, Ognjen Rudovic, Pranay Dighe, Ahmed Hussen Abdelaziz, Erik Marchi, Saurabh Adya, Chandra Dhir, Ahmed H. Tewfik |
INTERSPEECH | 4 |
| 2021 | MorphGAN: One-Shot Face Synthesis GAN for Detecting Recognition Bias
Nataniel Ruiz, Barry-John Theobald, Anurag Ranjan, Ahmed Hussen Abdelaziz, Nicholas Apostoloff |
BMVC | 4 |
| 2021 | On The Role of Visual Cues in Audiovisual Speech EnhancementabstractWe present an introspection of an audiovisual speech enhancement model. In particular, we focus on interpreting how a neural audiovisual speech enhancement model uses visual cues to improve the quality of the target speech signal. We show that visual cues provide not only high-level information about speech activity, i.e., speech/silence, but also fine-grained visual information about the place of articulation. One byproduct of this finding is that the learned visual embeddings can be used as features for other visual speech applications. We demonstrate the effectiveness of the learned visual embeddings for classifying visemes (the visual analogy to phonemes). Our results provide insight into important aspects of audiovisual speech enhancement and demonstrate how such models can be used for self-supervision tasks for visual speech applications. Zakaria Aldeneh, Anushree Prasanna Kumar, Barry-John Theobald, Erik Marchi, Sachin Kajarekar, Devang Naik, Ahmed Hussen Abdelaziz |
ICASSP | 7 |
| 2021 | Audiovisual Speech Synthesis using Tacotron2abstractAudiovisual speech synthesis involves synthesizing a talking face while maximizing the coherency of the acoustic and visual speech. To solve this problem, we propose using AVTacotron2, which is an end-to-end text-to-audiovisual speech synthesizer based on the Tacotron2 architecture. AVTacotron2 converts a sequence of phonemes into a sequence of acoustic features and the corresponding controllers of a face model. The output acoustic features are passed through a WaveRNN model to reconstruct the speech waveform. The speech waveform and the output facial controllers are used to generate the corresponding video of the talking face. As a baseline, we use a modular system, where acoustic speech is synthesized from text using the traditional Tacotron2. The reconstructed acoustic speech is then used to drive the controls of the face model using an independently trained audio-to-facial-animation neural network. We further condition both the end-to-end and modular approaches on emotion embeddings that encode the required prosody to generate emotional audiovisual speech. A comprehensive analysis shows that the end-to-end system is able to synthesize close to human-like audiovisual speech with mean opinion scores (MOS) of 4.1, which is the same MOS obtained on the ground truth generated from professionally recorded videos. Ahmed Hussen Abdelaziz, Anushree Prasanna Kumar, Chloe Seivwright, Gabriele Fanelli, Justin Binder, Yannis Stylianou, Sachin Kajareker |
ICMI | 1 |
| 2020 | Modality Dropout for Improved Performance-driven Talking FacesabstractWe describe our novel deep learning approach for driving animated faces using both acoustic and visual information. In particular, speech-related facial movements are generated using audiovisual information, and non-verbal facial movements are generated using only visual information. To ensure that our model exploits both modalities during training, batches are generated that contain audio-only, video-only, and audiovisual input features. The probability of dropping a modality allows control over the degree to which the model exploits audio and visual information during training. Our trained model runs in real-time on resource limited hardware (e.g. a smart phone), it is user agnostic, and it is not dependent on a potentially error-prone transcription of the speech. We use subjective testing to demonstrate: 1) the improvement of audiovisual-driven animation over the equivalent video-only approach, and 2) the improvement in the animation of speech-related facial movements after introducing modality dropout. Without modality dropout, viewers prefer audiovisual-driven animation in 51% of the test sequences compared with only 18% for video-driven. After introducing dropout viewer preference for audiovisual-driven animation increases to 74%, but decreases to 8% for video-only. Ahmed Hussen Abdelaziz, Barry-John Theobald, Paul Dixon, Reinhard Knothe, Nicholas Apostoloff, Sachin Kajareker |
ICMI | 1 |
| 2019 | Speaker-Independent Speech-Driven Visual Speech Synthesis using Domain-Adapted Acoustic ModelsabstractSpeech-driven visual speech synthesis involves mapping acoustic speech features to the corresponding lip animation controls for a face model. This mapping can take many forms, but a powerful approach is to use deep neural networks (DNNs). The lack of synchronized audio, video, and depth data is a limitation to reliably train DNNs, especially for speaker-independent models. In this paper, we investigate adapting an automatic speech recognition (ASR) acoustic model (AM) for the visual speech synthesis problem. We train the ASR-AM on ten thousand hours of audio-only transcribed speech. The ASR-AM is then adapted to the visual speech synthesis domain using ninety hours of synchronized audio-visual speech. Using a subjective assessment test, we compared the performance of the AM-initialized DNN to a randomly initialized model. The results show that viewers significantly prefer animations generated from the AM-initialized DNN than the ones generated using the randomly initialized model. We conclude that visual speech synthesis can significantly benefit from the powerful representation of speech in the ASR acoustic models. Ahmed Hussen Abdelaziz, Barry-John Theobald, Justin Binder, Gabriele Fanelli, Paul Dixon, Nicholas Apostoloff, Thibaut Weise, Sachin Kajareker |
ICMI | 1 |
| 2018 | Comparing Fusion Models for DNN-Based Audiovisual Continuous Speech RecognitionabstractAudiovisual fusion is one of the most challenging tasks that continues to attract substantial research interest in the field of audiovisual automatic speech recognition (AV-ASR). In the last few decades, many approaches for integrating the audio and video modalities were proposed to enhance the performance of automatic speech recognition in both clean and noisy conditions. However, very few studies can be found in the literature that compare different fusion models for AV-ASR. Even less research work compares audiovisual fusion models for large vocabulary continuous speech recognition (LVCSR) models using deep neural networks (DNNs). This paper reviews and compares the performance of five audiovisual fusion models: the feature fusion model, the decision fusion model, the multistream hidden Markov model (HMM), the coupled HMM, and the turbo decoders. A complete evaluation of these fusion models is conducted using a standard speaker-independent DNN-based LVCSR Kaldi recipe in three experimental setups: a clean-train-clean-test, a clean-train-noisy-test, and a matched-training setup. All experiments have been applied to the recently released NTCD-TIMIT audiovisual corpus. The task of NTCD-TIMIT is phone recognition in continuous speech. Using NTCD-TIMIT with its freely available visual features and 37 clean and noisy acoustic signals allows for this study to be a common benchmark, to which novel LVCSR AV-ASR models and approaches can be compared. Ahmed Hussen Abdelaziz |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2017 | Improving acoustic modeling using audio-visual speechabstractReliable visual features that encode the articulator movements of speakers can dramatically improve the decoding accuracy of automatic speech recognition systems when combined with the corresponding acoustic signals. In this paper, a novel framework is proposed to utilize audio-visual speech not only during decoding but also for training better acoustic models. In this framework, a multi-stream hidden Markov model is iteratively deployed to fuse audio and video likelihoods. The fused likelihoods are used to estimate enhanced frame-state alignments, which are finally used as better training targets. The proposed framework is so flexible that it can be partially used to train acoustic models with the available audio-visual data while a conventional training strategy can be followed with the remaining acoustic data. The experimental results show that the acoustic models trained using the proposed audio-visual framework perform significantly better than those trained conventionally with solely acoustic data in clean and noisy conditions. Ahmed Hussen Abdelaziz |
ICME | 1 |
| 2017 | Turbo Decoders for Audio-Visual Continuous Speech Recognition
Ahmed Hussen Abdelaziz |
INTERSPEECH | 1 |
| 2017 | NTCD-TIMIT: A New Database and Baseline for Noise-Robust Audio-Visual Speech Recognition
Ahmed Hussen Abdelaziz |
INTERSPEECH | 1 |
| 2016 | Twin-HMM-based non-intrusive speech intelligibility predictionabstractMost of the objective measures employed for speech intelligibility prediction require a clean reference signal, which is not accessible in all realistic scenarios. In this paper, we propose to re-synthesize the relevant features of the clean signal using only the noisy speech signal and utilize them inside an intelligibility prediction framework which requires a reference. A statistical model called twin hidden Markov model (THMM) is used to synthesize the clean speech features. For the intelligibility prediction framework, the short-time objective intelligibility (STOI) measure is used as an accurate and well-known method. The experimental results show a high correlation between the twin-HMM-based STOI (THMMB-STOI) and the human speech recognition results, even slightly outperforming the conventional STOI predictions computed using the actual clean reference signals. Mahdie Karbasi, Ahmed Hussen Abdelaziz, Dorothea Kolossa |
ICASSP | 2 |
| 2016 | Dynamic Stream Weighting for Turbo-Decoding-Based Audiovisual ASR
Sebastian Gergen, Steffen Zeiler, Ahmed Hussen Abdelaziz, Robert M. Nickel, Dorothea Kolossa |
INTERSPEECH | 3 |
| 2016 | Blind Non-Intrusive Speech Intelligibility Prediction Using Twin-HMMs
Mahdie Karbasi, Ahmed Hussen Abdelaziz, Hendrik Meutzner, Dorothea Kolossa |
INTERSPEECH | 2 |
| 2016 | Introducing the Turbo-Twin-HMM for Audio-Visual Speech Enhancement
Steffen Zeiler, Hendrik Meutzner, Ahmed Hussen Abdelaziz, Dorothea Kolossa |
INTERSPEECH | 3 |
| 2016 | General hybrid framework for uncertainty-decoding-based automatic speech recognition systems
Ahmed Hussen Abdelaziz, Dorothea Kolossa |
Speech Commun. | 1 |
| 2015 | Uncertainty propagation through deep neural networksabstractIn order to improve the ASR performance in noisy environments, distorted speech is typically pre-processed by a speech enhancement algorithm, which usually results in a speech estimate containing residual noise and distortion.We may also have some measures of uncertainty or variance of the estimate.Uncertainty decoding is a framework that utilizes this knowledge of uncertainty in the input features during acoustic model scoring.Such frameworks have been well explored for traditional probabilistic models, but their optimal use for deep neural network (DNN)-based ASR systems is not yet clear.In this paper, we study the propagation of observation uncertainties through the layers of a DNN-based acoustic model.Since this is intractable due to the nonlinearities of the DNN, we employ approximate propagation methods, including Monte Carlo sampling, the unscented transform, and the piecewise exponential approximation of the activation function, to estimate the distribution of acoustic scores.Finally, the expected value of the acoustic score distribution is used for decoding, which is shown to further improve the ASR accuracy on the CHiME database, relative to a highly optimized DNN baseline. Ahmed Hussen Abdelaziz, Shinji Watanabe 0001, John R. Hershey, Emmanuel Vincent 0001, Dorothea Kolossa |
INTERSPEECH | 1 |
| 2015 | Robust speech processing using observation uncertainty and uncertainty propagation: session and paper overview
Ramón Fernandez Astudillo, Shinji Watanabe 0001, Ahmed Hussen Abdelaziz, Dorothea Kolossa |
INTERSPEECH | 3 |
| 2015 | Learning Dynamic Stream Weights For Coupled-HMM-Based Audio-Visual Speech RecognitionabstractWith the increasing use of multimedia data in communication technologies, the idea of employing visual information in automatic speech recognition (ASR) has recently gathered momentum. In conjunction with the acoustical information, the visual data enhances the recognition performance and improves the robustness of ASR systems in noisy and reverberant environments. In audio-visual systems, dynamic weighting of audio and video streams according to their instantaneous confidence is essential for reliably and systematically achieving high performance. In this paper, we present a complete framework that allows blind estimation of dynamic stream weights for audio-visual speech recognition based on coupled hidden Markov models (CHMMs). As a stream weight estimator, we consider using multilayer perceptrons and logistic functions to map multidimensional reliability measure features to audiovisual stream weights. Training the parameters of the stream weight estimator requires numerous input-output tuples of reliability measure features and their corresponding stream weights. We estimate these stream weights based on oracle knowledge using an expectation maximization algorithm. We define 31-dimensional feature vectors that combine model-based and signal-based reliability measures as inputs to the stream weight estimator. During decoding, the trained stream weight estimator is used to blindly estimate stream weights. The entire framework is evaluated using the Grid audio-visual corpus and compared to state-of-the-art stream weight estimation strategies. The proposed framework significantly enhances the performance of the audio-visual ASR system in all examined test conditions. Ahmed Hussen Abdelaziz, Steffen Zeiler, Dorothea Kolossa |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2014 | Human-robot collaborative tutoring using multiparty multimodal spoken dialogueabstractIn this paper, we describe a project that explores a novel experimental setup towards building a spoken, multi-modally rich, and human-like multiparty tutoring robot. A human-robot interaction setup is designed, and a human-human dialogue corpus is collected. The corpus targets the development of a dialogue system platform to study verbal and nonverbal tutoring strategies in multiparty spoken interactions with robots which are capable of spoken dialogue. The dialogue task is centered on two participants involved in a dialogue aiming to solve a card-ordering game. Along with the participants sits a tutor (robot) that helps the participants perform the task, and organizes and balances their interaction. Different multimodal signals captured and auto-synchronized by different audio-visual capture technologies, such as a microphone array, Kinects, and video cameras, were coupled with manual annotations. These are used build a situated model of the interaction based on the participants personalities, their state of attention, their conversational engagement and verbal dominance, and how that is correlated with the verbal and visual feed-back, turn-management, and conversation regulatory actions generated by the tutor. Driven by the analysis of the corpus, we will show also the detailed design methodologies for an affective, and multimodally rich dialogue system that allows the robot to measure incrementally the attention states, and the dominance for each participant, allowing the robot head Furhat to maintain a well-coordinated, balanced, and engaging conversation, that attempts to maximize the agreement and the contribution to solve the task. Samer Al Moubayed, Jonas Beskow, Bajibabu Bollepalli, Joakim Gustafson, Ahmed Hussen Abdelaziz, Martin Johansson, Maria Koutsombogera, José Lopes 0001, Jekaterina Novikova, Catharine Oertel, Gabriel Skantze, Kalin Stefanov, Gül Varol |
HRI | 5 |
| 2014 | A newem estimationof dynamic stream weights for coupled-HMM-based audio-visual ASRabstractMutually deploying visual and acoustical information in automatic speech recognition systems increases their robustness against acoustical environmental effects like additive noise and reverberation. Optimal fusion of the audio and video streams requires dynamic adaptation of the relative contribution of each modality. This can be achieved by weighting each stream according to its reliability by an appropriate stream weight. In this paper we propose a new expectation maximization algorithm that estimates oracle frame-dependent stream weights for coupled-HMM-based audio-visual speech recognition. Moreover, we introduce a greedy optimization approach that reasonably initializes this algorithm. The proposed approach is evaluated on the Grid audio-visual database and results in an average relative word error rate reduction of 38% and 58% compared to grid search and Bayes fusion, respectively. The estimated oracle stream weights can be used instead of the conventional global fixed stream weights to improve the supervised training of stream weight estimators. Ahmed Hussen Abdelaziz, Steffen Zeiler, Dorothea Kolossa |
ICASSP | 1 |
| 2014 | Dynamic stream weight estimation in coupled-HMM-based audio-visual speech recognition using multilayer perceptrons
Ahmed Hussen Abdelaziz, Dorothea Kolossa |
INTERSPEECH | 1 |
| 2014 | The Tutorbot Corpus ― A Corpus for Studying Tutoring Behaviour in Multiparty Face-to-Face Spoken Dialogue
Maria Koutsombogera, Samer Al Moubayed, Bajibabu Bollepalli, Ahmed Hussen Abdelaziz, Martin Johansson, José Lopes 0001, Jekaterina Novikova, Catharine Oertel, Kalin Stefanov, Gül Varol |
LREC | 4 |
| 2013 | Twin-HMM-based audio-visual speech enhancementabstractMost approaches for speech signal processing rely solely on acoustic input, which has the consequence that spectrum estimation becomes exceedingly difficult when the signal-to-noise ratio drops to values near 0 dB. However, alternative sources of information are becoming widely available with increasing use of multimedia data in everyday communication. In the following paper, we suggest to use video input as an auxiliary modality for speech processing by applying a new statistical model - the twin hidden Markov model. The resulting enhancement algorithm for audiovisual data greatly outperforms the standard audio-only log-MMSE estimator on all considered instrumental speech quality measures covering spectral and perceptual quality. Ahmed Hussen Abdelaziz, Steffen Zeiler, Dorothea Kolossa |
ICASSP | 1 |
| 2013 | GMM-based significance decodingabstractThe accuracy of automatic speech recognition systems in noisy and reverberant environments can be improved notably by exploiting the uncertainty of the estimated speech features using so-called uncertainty-of-observation techniques. In this paper, we introduce a new Bayesian decision rule that can serve as a mathematical framework from which both known and new uncertainty-of-observation techniques can be either derived or approximated. The new decision rule in its direct form leads to the new significance decoding approach for Gaussian mixture models, which results in better performance compared to standard uncertainty-of-observation techniques in different additive and convolutive noise scenarios. Ahmed Hussen Abdelaziz, Steffen Zeiler, Dorothea Kolossa, Volker Leutnant, Reinhold Häb-Umbach |
ICASSP | 1 |
| 2013 | Using twin-HMM-based audio-visual speech enhancement as a front-end for robust audio-visual speech recognition
Ahmed Hussen Abdelaziz, Steffen Zeiler, Dorothea Kolossa |
INTERSPEECH | 1 |
| 2012 | Decoding of Uncertain Features Using the Posterior Distribution of the Clean Data for Robust Speech Recognition
Ahmed Hussen Abdelaziz, Dorothea Kolossa |
INTERSPEECH | 1 |