EDBT 2026 Demo / reviewers in the wild / expert
Amir Hussein
dblp:134/9208
· DBLP profile ↗
15ranked-venue papers
8as first author
13since 2021 · last 2025
0000-0002-0820-4062ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 5 first-author · 10 since 2021Artificial intelligence and machine learning · 9 · 3 first-author · 8 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement
Amir Hussein, Sameer Khurana, Gordon Wichern, François G. Germain, Jonathan Le Roux |
INTERSPEECH | 1 |
| 2025 | Factorized RVQ-GAN For Disentangled Speech TokenizationabstractInternational audience Sameer Khurana, Dominik Klement, Antoine Laurent, Dominik Bobos, Juraj Novosad, Peter Gazdik, Ellen Zhang, Zili Huang, Amir Hussein, Ricard Marxer, Yoshiki Masuyama, Ryo Aihara, Chiori Hori, François G. Germain, Gordon Wichern, Jonathan Le Roux |
INTERSPEECH | 9 |
| 2025 | CS-FLEURS: A Massively Multilingual and Code-Switched Speech Dataset
Brian Yan, Injy Hamed, Shuichiro Shimizu, Vasista Sai Lodagala, Olga Iakovenko, Bashar Talafha, Amir Hussein, Alexander Polok, Kalvin Chang, Dominik Klement, Sara Althubaiti, Puyuan Peng, Matthew Wiesner, Thamar Solorio, Ahmed Ali 0002, Sanjeev Khudanpur, Shinji Watanabe 0001 |
INTERSPEECH | 8 |
| 2024 | Enhancing End-to-End Conversational Speech Translation Through Target Language Context UtilizationabstractIncorporating longer context has been shown to benefit machine translation, but the inclusion of context in end-to-end speech translation (E2E-ST) remains under-studied. To bridge this gap, we introduce target language context in E2E-ST, enhancing coherence and overcoming memory constraints of extended audio segments. Additionally, we propose context dropout to ensure robustness to the absence of context, and further improve performance by adding speaker information. Our proposed contextual E2E-ST outperforms the isolated utterance-based E2E-ST approach. Lastly, we demonstrate that in conversational speech, contextual information primarily contributes to capturing context style, as well as resolving anaphora and named entities. Amir Hussein, Brian Yan, Antonios Anastasopoulos, Shinji Watanabe 0001, Sanjeev Khudanpur |
ICASSP | 1 |
| 2024 | Speech Collage: Code-Switched Audio Generation by Collaging Monolingual CorporaabstractDesigning effective automatic speech recognition (ASR) systems for Code-Switching (CS) often depends on the availability of the transcribed CS resources. To address data scarcity, this paper introduces Speech Collage, a method that synthesizes CS data from monolingual corpora by splicing audio segments. We further improve the smoothness quality of audio generation using an overlap-add approach. We investigate the impact of generated data on speech recognition in two scenarios: using in-domain CS text and a zero-shot approach with synthesized CS text. Empirical results highlight up to 34.4% and 16.2% relative reductions in Mixed-Error Rate and Word-Error Rate for in-domain and zero-shot scenarios, respectively. Lastly, we demonstrate that CS augmentation bolsters the model’s code-switching inclination and reduces its monolingual bias. Amir Hussein, Dorsa Zeinali, Ondrej Klejch, Matthew Wiesner, Brian Yan, Shammur Absar Chowdhury, Ahmed Ali 0002, Shinji Watanabe 0001, Sanjeev Khudanpur |
ICASSP | 1 |
| 2024 | Enhancing Neural Transducer for Multilingual ASR with Synchronized Language Diarization
Amir Hussein, Desh Raj, Matthew Wiesner, Daniel Povey, L. Paola García-Perera, Sanjeev Khudanpur |
INTERSPEECH | 1 |
| 2022 | Benchmarking Evaluation Metrics for Code-Switching Automatic Speech RecognitionabstractCode-switching poses a number of challenges and opportunities for multilingual automatic speech recognition. In this paper, we focus on the question of robust and fair evaluation metrics. To that end, we develop a reference benchmark data set of code-switching speech recognition hypotheses with human judgments. We define clear guidelines for minimal editing of automatic hypotheses. We validate the guidelines using 4-way inter-annotator agreement. We evaluate a large number of metrics in terms of correlation with human judgments. The metrics we consider vary in terms of representation (orthographic, phonological, semantic), directness (intrinsic vs extrinsic), granularity (e.g. word, character), and similarity computation method. The highest correlation to human judgment is achieved using transliteration followed by text normalization. We release the first corpus for human acceptance of code-switching speech recognition results in dialectal Arabic/English conversation speech. Injy Hamed, Amir Hussein, Oumnia Chellah, Shammur Absar Chowdhury, Hamdy Mubarak, Sunayana Sitaram, Nizar Habash, Ahmed Ali 0002 |
SLT | 2 |
| 2022 | Textual Data Augmentation for Arabic-English Code-Switching Speech RecognitionabstractThe pervasiveness of intra-utterance code-switching (CS) in spoken content requires that speech recognition (ASR) systems handle mixed language. Designing a CS-ASR system has many challenges, mainly due to data scarcity, grammatical structure complexity, and domain mismatch. The most common method for addressing CS is to train an ASR system with the available transcribed CS speech, along with monolingual data. In this work, we propose a zero-shot learning methodology for CS-ASR by augmenting the monolingual data with artificially generating CS text. We based our approach on random lexical replacements and Equivalence Constraint (EC) while exploiting aligned translation pairs to generate random and grammatically valid CS content. Our empirical results show a 65.5% relative reduction in language model perplexity, and 7.7% in ASR WER on two ecologically valid CS test sets. The human evaluation of the generated text using EC suggests that more than 80% is of adequate quality. Amir Hussein, Shammur Absar Chowdhury, Ahmed Abdelali, Najim Dehak, Ahmed Ali 0002, Sanjeev Khudanpur |
SLT | 1 |
| 2022 | Arabic speech recognition by end-to-end, modular systems and human
Amir Hussein, Shinji Watanabe 0001, Ahmed Ali 0002 |
Comput. Speech Lang. | 1 |
| 2022 | Domain Adaptation with Representation Learning and Nonlinear Relation for Time SeriesabstractIn many real-world scenarios, machine learning models fall short in prediction performance due to data characteristics changing from training on one source domain to testing on a target domain. There has been extensive research to address this problem with Domain Adaptation (DA) for learning domain invariant features. However, when considering advances for time series, those methods remain limited to the use of hard parameter sharing (HPS) between source and target models, and the use of domain adaptation objective function. To address these challenges, we propose a soft parameter sharing (SPS) DA architecture with representation learning while modeling the relation as non-linear between parameters of source and target models and modeling the adaptation loss function as the squared Maximum Mean Discrepancy (MMD) . The proposed architecture advances the state-of-the-art for time series in the context of activity recognition and in fields with other modalities, where SPS has been limited to a linear relation. An additional contribution of our work is to provide a study that demonstrates the strengths and limitations of HPS versus SPS. Experiment results showed the success of the method in three domain adaptation cases of multivariate time series activity recognition with different users and sensors. Amir Hussein, Hazem M. Hajj |
ACM Trans. Internet Things | 1 |
| 2021 | QASR: QCRI Aljazeera Speech Resource A Large Scale Annotated Arabic Speech CorpusabstractHamdy Mubarak, Amir Hussein, Shammur Absar Chowdhury, Ahmed Ali. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Hamdy Mubarak, Amir Hussein, Shammur Absar Chowdhury, Ahmed Ali 0002 |
ACL/IJCNLP (1) | 2 |
| 2021 | Arabic Code-Switching Speech Recognition Using Monolingual DataabstractCode-switching in automatic speech recognition (ASR) is an important challenge due to globalization. Recent research in multilingual ASR shows potential improvement over monolingual systems. We study key issues related to multilingual modeling for ASR through a series of large-scale ASR experiments. Our innovative framework deploys a multi-graph approach in the weighted finite state transducers (WFST) framework. We compare our WFST decoding strategies with a transformer sequence to sequence system trained on the same data. Given a code-switching scenario between Arabic and English languages, our results show that the WFST decoding approaches were more suitable for the intersentential code-switching datasets. In addition, the transformer system performed better for intrasentential code-switching task. With this study, we release an artificially generated development and test sets, along with ecological code-switching test set, to benchmark the ASR performance. Ahmed Ali 0002, Shammur Absar Chowdhury, Amir Hussein, Yasser Hifny |
Interspeech | 3 |
| 2021 | Towards One Model to Rule All: Multilingual Strategy for Dialectal Code-Switching Arabic ASRabstractWith the advent of globalization, there is an increasing demand for multilingual automatic speech recognition (ASR), handling language and dialectal variation of spoken content. Recent studies show its efficacy over monolingual systems. In this study, we design a large multilingual end-to-end ASR using self-attention based conformer architecture. We trained the system using Arabic (Ar), English (En) and French (Fr) languages. We evaluate the system performance handling: (i) monolingual (Ar, En and Fr); (ii) multi-dialectal (Modern Standard Arabic, along with dialectal variation such as Egyptian and Moroccan); (iii) code-switching -- cross-lingual (Ar-En/Fr) and dialectal (MSA-Egyptian dialect) test cases, and compare with current state-of-the-art systems. Furthermore, we investigate the influence of different embedding/character representations including character vs word-piece; shared vs distinct input symbol per language. Our findings demonstrate the strength of such a model by outperforming state-of-the-art monolingual dialectal Arabic and code-switching Arabic ASR. Shammur Absar Chowdhury, Amir Hussein, Ahmed Abdelali, Ahmed Ali 0002 |
Interspeech | 2 |
| 2020 | Augmenting DL with Adversarial Training for Robust Prediction of Epilepsy SeizuresabstractEpilepsy is a chronic medical condition that involves abnormal brain activity causing patients to lose control of awareness or motor activity. As a result, detection of pre-ictal states, before the onset of a seizure, can be lifesaving. The problem is challenging because it is difficult to discern between electroencephalogram signals in pre-ictal states versus signals in normal inter-ictal states. There are three key challenges that have not been addressed previously: (1) the inconsistent performance of prediction models across patients, (2) the lack of perfect prediction to protect patients from any episode, and (3) the limited amount of pre-ictal labeled data for advancing machine learning methods. This article addresses these limitations through a novel approach that uses adversarial examples with optimized tuning of a combined convolutional neural network and gated recurrent unit. Compared to the state of the art, the results showed an improvement of 3x in model robustness as measured in reduced variations and superior accuracy of the area under the curve, with an average increase of 6.7%. The proposed method also exhibited superior performance with other advances in the field of machine learning and customized for epilepsy prediction including data augmentation with Gaussian noise and multitask learning. Amir Hussein, Marc Djandji, Reem A. Mahmoud, Mohamad Dhaybi, Hazem M. Hajj |
ACM Trans. Comput. Heal. | 1 |
| 2013 | Special issue on image feature detection and description
Yanwei Pang, Xianbin Cao 0001, Lei Zhang 0001, Amir Hussein |
Neurocomputing | 4 |