EDBT 2026 Demo / reviewers in the wild / expert
Ahmed Ali 0002
dblp:22/1217-2 · also Ahmed M. Ali 0002
· DBLP profile ↗
49ranked-venue papers
10as first author
19since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 41 · 10 first-author · 14 since 2021Artificial intelligence and machine learning · 36 · 8 first-author · 14 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Towards a Unified Benchmark for Arabic Pronunciation Assessment: Qur'anic Recitation as Case Study
Yassine El Kheir, Omnia Ibrahim, Amit Meghanani, Nada Almarwani, Hawau Olamide Toyin, Sadeen Alharbi, Modar Alfadly, Lamya Alkanhal, Ibrahim Selim, Shehab Elbatal, Salima Mdhaffar, Thomas Hain, Yasser Hifny, Mostafa Shahin, Ahmed Ali 0002 |
INTERSPEECH | 15 |
| 2025 | SawtArabi: A Benchmark Corpus for Arabic TTS. Standard, Dialectal and Code-Switching
Vasista Sai Lodagala, Lamya Alkanhal, Daniel Izham, Shivam Mehta, Shammur Absar Chowdhury, Aqeelah Makki, Hamdy S. Hussein, Gustav Eje Henter, Ahmed Ali 0002 |
INTERSPEECH | 9 |
| 2025 | CS-FLEURS: A Massively Multilingual and Code-Switched Speech Dataset
Brian Yan, Injy Hamed, Shuichiro Shimizu, Vasista Sai Lodagala, Olga Iakovenko, Bashar Talafha, Amir Hussein, Alexander Polok, Kalvin Chang, Dominik Klement, Sara Althubaiti, Puyuan Peng, Matthew Wiesner, Thamar Solorio, Ahmed Ali 0002, Sanjeev Khudanpur, Shinji Watanabe 0001 |
INTERSPEECH | 16 |
| 2024 | Beyond Orthography: Automatic Recovery of Short Vowels and Dialectal Sounds in ArabicabstractThis paper presents a novel Dialectal Sound and Vowelization Recovery framework, designed to recognize borrowed and dialectal sounds within phonologically diverse and dialect-rich languages, that extends beyond its standard orthographic sound sets.The proposed framework utilized quantized sequence of input with(out) continuous pretrained selfsupervised representation.We show the efficacy of the pipeline using limited data for Arabic, a dialect-rich language containing more than 22 major dialects.Phonetically correct transcribed speech resources for dialectal Arabic is scare.Therefore, we introduce Arab-Voice15 1 , a first of its kind, curated test set featuring 5 hours of dialectal speech across 15 Arab countries, with phonetically accurate transcriptions, including borrowed and dialectspecific sounds.We described in detail the annotation guideline along with the analysis of the dialectal confusion pairs.Our extensive evaluation includes both subjective -human perception tests and objective measures.Our empirical results, reported with three test sets, show that with only one and half hours of training data, our model improve character error rate by ≈ 7% in ArabVoice15 compared to the baseline. Yassine El Kheir, Hamdy Mubarak, Ahmed Ali 0002, Shammur Absar Chowdhury |
ACL (1) | 3 |
| 2024 | LAraBench: Benchmarking Arabic AI with Large Language ModelsabstractAhmed Abdelali, Hamdy Mubarak, Shammur Chowdhury, Maram Hasanain, Basel Mousi, Sabri Boughorbel, Samir Abdaljalil, Yassine El Kheir, Daniel Izham, Fahim Dalvi, Majd Hawasly, Nizi Nazar, Youssef Elshahawy, Ahmed Ali, Nadir Durrani, Natasa Milic-Frayling, Firoj Alam. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Ahmed Abdelali, Hamdy Mubarak, Shammur Absar Chowdhury, Maram Hasanain, Basel Mousi, Sabri Boughorbel, Samir Abdaljalil, Yassine El Kheir, Daniel Izham, Fahim Dalvi, Majd Hawasly, Nizi Nazar, Yousseif Elshahawy, Ahmed Ali 0002, Nadir Durrani, Natasa Milic-Frayling, Firoj Alam |
EACL (1) | 14 |
| 2024 | Speech Collage: Code-Switched Audio Generation by Collaging Monolingual CorporaabstractDesigning effective automatic speech recognition (ASR) systems for Code-Switching (CS) often depends on the availability of the transcribed CS resources. To address data scarcity, this paper introduces Speech Collage, a method that synthesizes CS data from monolingual corpora by splicing audio segments. We further improve the smoothness quality of audio generation using an overlap-add approach. We investigate the impact of generated data on speech recognition in two scenarios: using in-domain CS text and a zero-shot approach with synthesized CS text. Empirical results highlight up to 34.4% and 16.2% relative reductions in Mixed-Error Rate and Word-Error Rate for in-domain and zero-shot scenarios, respectively. Lastly, we demonstrate that CS augmentation bolsters the model’s code-switching inclination and reduces its monolingual bias. Amir Hussein, Dorsa Zeinali, Ondrej Klejch, Matthew Wiesner, Brian Yan, Shammur Absar Chowdhury, Ahmed Ali 0002, Shinji Watanabe 0001, Sanjeev Khudanpur |
ICASSP | 7 |
| 2024 | L1-Aware Multilingual Mispronunciation Detection FrameworkabstractThe phonological discrepancies between a speaker’s native (L1) and the non-native language (L2) serves as a major factor for mispronunciation. This paper introduces a novel multilingual Mispronunciation Detection and Diagnosis (MDD) architecture, L1-MultiMDD, enriched with L1-aware speech representation. An end-to-end speech encoder is trained on the input signal and its corresponding reference phoneme sequence. First, an attention mechanism is deployed to align the input audio with the reference phoneme sequence. Afterwards, the L1-L2-speech embedding are extracted from an auxiliary model, pretrained in a multi-task setup identifying L1 and L2 language, and are infused with the primary network. Finally, the L1-MultiMDD is then optimized for a unified multilingual phoneme recognition task using connectionist temporal classification (CTC) loss for the target languages: English, Arabic, and Mandarin. Our experiments demonstrate the effectiveness of the proposed L1-MultiMDD framework on both seen – L2-ARTIC, LATIC, and AraVoiceL2v2; and unseen – EpaDB and Speechocean762 datasets. The consistent gains in PER, and false rejection rate (FRR) across all target languages confirm our approach’s robustness, efficacy, and generalizability. Yassine El Kheir, Shammur Absar Chowdhury, Ahmed Ali 0002 |
ICASSP | 3 |
| 2024 | What do end-to-end speech models learn about speaker, language and channel information? A layer-wise and neuron-level analysisabstractDeep neural networks are inherently opaque and challenging to interpret. Unlike hand-crafted feature-based models, we struggle to comprehend the concepts learned and how they interact within these models. This understanding is crucial not only for debugging purposes but also for ensuring fairness in ethical decision-making. In our study, we conduct a post-hoc functional interpretability analysis of pretrained speech models using the probing framework (Hupkes et al., 2018). Specifically, we analyze utterance-level representations of speech models trained for various tasks such as speaker recognition and dialect identification. We conduct layer and neuron-wise analyses, probing for speaker, language, and channel properties. Our study aims to answer the following questions: (i) what information is captured within the representations? (ii) how is it represented and distributed? and (iii) can we identify a minimal subset of the network that possesses this information? Our results reveal several novel findings, including: (i) channel and gender information are distributed across the network, (ii) the information is redundantly available in neurons with respect to a task, (iii) complex properties such as dialectal information are encoded only in the task-oriented pretrained network, (iv) and is localised in the upper layers, (v) we can extract a minimal subset of neurons encoding the pre-defined property, (vi) salient neurons are sometimes shared between properties, (vii) our analysis highlights the presence of biases (for example gender) in the network. Our cross-architectural comparison indicates that: (i) the pretrained models capture speaker-invariant information, and (ii) CNN models are competitive with Transformer models in encoding various understudied properties. Shammur Absar Chowdhury, Nadir Durrani, Ahmed Ali 0002 |
Comput. Speech Lang. | 3 |
| 2023 | Multilingual Word Error Rate Estimation: E-Wer3abstractThe success of the multilingual automatic speech recognition systems empowered many voice-driven applications. However, measuring the performance of such systems remains a major challenge, due to its dependency on manually transcribed speech data in both mono- and multilingual scenarios. In this paper, we propose a novel multilingual framework – eWER3 – jointly trained on acoustic and lexical representation to estimate word error rate. We demonstrate the effectiveness of eWER3 to (i) predict WER without using any internal states from the ASR and (ii) use the multilingual shared latent space to push the performance of the close-related languages. We show our proposed multilingual model outperforms the previous monolingual word error rate estimation method (eWER2) by an absolute 9% increase in Pearson correlation coefficient (PCC), with better overall estimation between the predicted and reference WER. Shammur Absar Chowdhury, Ahmed Ali 0002 |
ICASSP | 2 |
| 2023 | FOOCTTS: Generating Arabic Speech with Acoustic Environment for Football Commentator
Massa Baali, Ahmed Ali 0002 |
INTERSPEECH | 2 |
| 2023 | Unsupervised Code-switched Text Generation from Parallel TextabstractSpeech is a fundamental means of communication that can be seen to provide two channels for transmitting information: the lexical channel of which words are said, and the non-lexical channel of how they are spoken. Both channels shape listener expectations of upcoming communication; however, directly quantifying their relative effect on expectations is challenging. Previous attempts require spoken variations of lexically equivalent dialogue turns or conspicuous acoustic manipulations. This paper introduces a generalised paradigm to study the value of non-lexical information in dialogue across unconstrained lexical content. By quantifying the perceptual value of the non-lexical channel with both accuracy and entropy reduction, we show that non-lexical information produces a consistent effect on expectations of upcoming dialogue: even when it leads to poorer discriminative turn judgements than lexical content alone, it yields higher consensus among participants. Jie Chi, Brian Lu, Jason Eisner, Peter Bell 0001, Preethi Jyothi, Ahmed Ali 0002 |
INTERSPEECH | 6 |
| 2023 | MyVoice: Arabic Speech Resource Collaboration Platform
Yousseif Elshahawy, Yassine El Kheir, Shammur Absar Chowdhury, Ahmed Ali 0002 |
INTERSPEECH | 4 |
| 2023 | QVoice: Arabic Speech Pronunciation Learning Application
Yassine El Kheir, Fouad Khnaisser, Shammur Absar Chowdhury, Hamdy Mubarak, Shazia Afzal, Ahmed Ali 0002 |
INTERSPEECH | 6 |
| 2022 | Benchmarking Evaluation Metrics for Code-Switching Automatic Speech RecognitionabstractCode-switching poses a number of challenges and opportunities for multilingual automatic speech recognition. In this paper, we focus on the question of robust and fair evaluation metrics. To that end, we develop a reference benchmark data set of code-switching speech recognition hypotheses with human judgments. We define clear guidelines for minimal editing of automatic hypotheses. We validate the guidelines using 4-way inter-annotator agreement. We evaluate a large number of metrics in terms of correlation with human judgments. The metrics we consider vary in terms of representation (orthographic, phonological, semantic), directness (intrinsic vs extrinsic), granularity (e.g. word, character), and similarity computation method. The highest correlation to human judgment is achieved using transliteration followed by text normalization. We release the first corpus for human acceptance of code-switching speech recognition results in dialectal Arabic/English conversation speech. Injy Hamed, Amir Hussein, Oumnia Chellah, Shammur Absar Chowdhury, Hamdy Mubarak, Sunayana Sitaram, Nizar Habash, Ahmed Ali 0002 |
SLT | 8 |
| 2022 | Textual Data Augmentation for Arabic-English Code-Switching Speech RecognitionabstractThe pervasiveness of intra-utterance code-switching (CS) in spoken content requires that speech recognition (ASR) systems handle mixed language. Designing a CS-ASR system has many challenges, mainly due to data scarcity, grammatical structure complexity, and domain mismatch. The most common method for addressing CS is to train an ASR system with the available transcribed CS speech, along with monolingual data. In this work, we propose a zero-shot learning methodology for CS-ASR by augmenting the monolingual data with artificially generating CS text. We based our approach on random lexical replacements and Equivalence Constraint (EC) while exploiting aligned translation pairs to generate random and grammatically valid CS content. Our empirical results show a 65.5% relative reduction in language model perplexity, and 7.7% in ASR WER on two ecologically valid CS test sets. The human evaluation of the generated text using EC suggests that more than 80% is of adequate quality. Amir Hussein, Shammur Absar Chowdhury, Ahmed Abdelali, Najim Dehak, Ahmed Ali 0002, Sanjeev Khudanpur |
SLT | 5 |
| 2022 | Arabic speech recognition by end-to-end, modular systems and human
Amir Hussein, Shinji Watanabe 0001, Ahmed Ali 0002 |
Comput. Speech Lang. | 3 |
| 2021 | QASR: QCRI Aljazeera Speech Resource A Large Scale Annotated Arabic Speech CorpusabstractHamdy Mubarak, Amir Hussein, Shammur Absar Chowdhury, Ahmed Ali. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Hamdy Mubarak, Amir Hussein, Shammur Absar Chowdhury, Ahmed Ali 0002 |
ACL/IJCNLP (1) | 4 |
| 2021 | Arabic Code-Switching Speech Recognition Using Monolingual DataabstractCode-switching in automatic speech recognition (ASR) is an important challenge due to globalization. Recent research in multilingual ASR shows potential improvement over monolingual systems. We study key issues related to multilingual modeling for ASR through a series of large-scale ASR experiments. Our innovative framework deploys a multi-graph approach in the weighted finite state transducers (WFST) framework. We compare our WFST decoding strategies with a transformer sequence to sequence system trained on the same data. Given a code-switching scenario between Arabic and English languages, our results show that the WFST decoding approaches were more suitable for the intersentential code-switching datasets. In addition, the transformer system performed better for intrasentential code-switching task. With this study, we release an artificially generated development and test sets, along with ecological code-switching test set, to benchmark the ASR performance. Ahmed Ali 0002, Shammur Absar Chowdhury, Amir Hussein, Yasser Hifny |
Interspeech | 1 |
| 2021 | Towards One Model to Rule All: Multilingual Strategy for Dialectal Code-Switching Arabic ASRabstractWith the advent of globalization, there is an increasing demand for multilingual automatic speech recognition (ASR), handling language and dialectal variation of spoken content. Recent studies show its efficacy over monolingual systems. In this study, we design a large multilingual end-to-end ASR using self-attention based conformer architecture. We trained the system using Arabic (Ar), English (En) and French (Fr) languages. We evaluate the system performance handling: (i) monolingual (Ar, En and Fr); (ii) multi-dialectal (Modern Standard Arabic, along with dialectal variation such as Egyptian and Moroccan); (iii) code-switching -- cross-lingual (Ar-En/Fr) and dialectal (MSA-Egyptian dialect) test cases, and compare with current state-of-the-art systems. Furthermore, we investigate the influence of different embedding/character representations including character vs word-piece; shared vs distinct input symbol per language. Our findings demonstrate the strength of such a model by outperforming state-of-the-art monolingual dialectal Arabic and code-switching Arabic ASR. Shammur Absar Chowdhury, Amir Hussein, Ahmed Abdelali, Ahmed Ali 0002 |
Interspeech | 4 |
| 2020 | What Was Written vs. Who Read It: News Media Profiling Using Text Analysis and Social Media ContextabstractPredicting the political bias and the factuality of reporting of entire news outlets are critical elements of media profiling, which is an understudied but an increasingly important research direction.The present level of proliferation of fake, biased, and propagandistic content online, has made it impossible to fact-check every single suspicious claim, either manually or automatically.Alternatively, we can profile entire news outlets and look for those that are likely to publish fake or biased content.This approach makes it possible to detect likely "fake news" the moment they are published, by simply checking the reliability of their source.From a practical perspective, political bias and factuality of reporting have a linguistic aspect but also a social context.Here, we study the impact of both, namely (i) what was written (i.e., what was published by the target medium, and how it describes itself on Twitter) vs. (ii) who read it (i.e., analyzing the readers of the target medium on Facebook, Twitter, and YouTube).We further study (iii) what was written about the target medium on Wikipedia.The evaluation results show that what was written matters most, and that putting all information sources together yields huge improvements over the current state-of-the-art. Ramy Baly, Georgi Karadzhov, Jisun An, Haewoon Kwak, Yoan Dinkov, Ahmed Ali 0002, James R. Glass, Preslav Nakov |
ACL | 6 |
| 2020 | ADI17: A Fine-Grained Arabic Dialect Identification DatasetabstractIn this paper, we describe a method to collect dialectal speech from YouTube videos to create a large-scale Dialect Identification (DID) dataset. Using this method, we collected dialectal Arabic from known YouTube channels from 17 Arabic speaking countries in the Middle East and Northern Africa. After a refinement process, a total of 3,000 hours of speech was available for training DID systems, with an additional 57 hours of speech for development and testing. For detailed evaluations, the DID data was divided into three sub-categories based on the segment duration: short (less than 5s), medium (5-20s), and long (over 20s). We compare state-of-the-art DID techniques on these data, and also analyze a DID system trained on these data. Since the training and test data share the same channel domain, we also used the Multi-Genre Broadcast 3 (MGB-3) test set to evaluate on domain mismatched condition. Suwon Shon, Ahmed Ali 0002, Younes Samih, Hamdy Mubarak, James R. Glass |
ICASSP | 2 |
| 2020 | Word Error Rate Estimation Without ASR Output: e-WER2abstractMeasuring the performance of automatic speech recognition (ASR) systems requires manually transcribed data in order to compute the word error rate (WER), which is often time-consuming and expensive. In this paper, we continue our effort in estimating WER using acoustic, lexical and phonotactic features. Our novel approach to estimate the WER uses a multistream end-to-end architecture. We report results for systems using internal speech decoder features (glass-box), systems without speech decoder features (black-box), and for systems without having access to the ASR system (no-box). The no-box system learns joint acoustic-lexical representation from phoneme recognition results along with MFCC acoustic features to estimate WER. Considering WER per sentence, our no-box system achieves 0.56 Pearson correlation with the reference evaluation and 0.24 root mean square error (RMSE) across 1,400 sentences. The estimated overall WER by e-WER2 is 30.9% for a three hours test set, while the WER computed using the reference transcriptions was 28.5%. Ahmed Ali 0002, Steve Renals |
INTERSPEECH | 1 |
| 2020 | What Does an End-to-End Dialect Identification Model Learn About Non-Dialectal Information?
Shammur Absar Chowdhury, Ahmed Ali 0002, Suwon Shon, James R. Glass |
INTERSPEECH | 2 |
| 2020 | Effects of Dialectal Code-Switching on Speech Modules: A Study Using Egyptian Arabic Broadcast Speech
Shammur Absar Chowdhury, Younes Samih, Mohamed Eldesouki, Ahmed Ali 0002 |
INTERSPEECH | 4 |
| 2019 | The MGB-5 Challenge: Recognition and Dialect Identification of Dialectal Arabic SpeechabstractThis paper describes the fifth edition of the Multi-Genre Broadcast Challenge (MGB-5), an evaluation focused on Arabic speech recognition and dialect identification. MGB-5 extends the previous MGB-3 challenge in two ways: first it focuses on Moroccan Arabic speech recognition; second the granularity of the Arabic dialect identification task is increased from 5 dialect classes to 17, by collecting data from 17 Arabic speaking countries. Both tasks use YouTube recordings to provide a multi-genre multi-dialectal challenge in the wild. Moroccan speech transcription used about 13 hours of transcribed speech data, split across training, development, and test sets, covering 7-genres: comedy, cooking, family/kids, fashion, drama, sports, and science (TEDx). The fine-grained Arabic dialect identification data was collected from known YouTube channels from 17 Arabic countries. 3,000 hours of this data was released for training, and 57 hours for development and testing. The dialect identification data was divided into three sub-categories based on the segment duration: short (under 5 s), medium (5-20 s), and long (>20 s). Overall, 25 teams registered for the challenge, and 9 teams submitted systems for the two tasks. We outline the approaches adopted in each system and summarize the evaluation results. Ahmed Ali 0002, Suwon Shon, Younes Samih, Hamdy Mubarak, Ahmed Abdelali, James R. Glass, Steve Renals, Khalid Choukri |
ASRU | 1 |
| 2019 | Detecting Deception in Political Debates Using Acoustic and Textual FeaturesabstractWe present work on deception detection, where, given a spoken claim, we aim to predict its factuality. While previous work in the speech community has relied on recordings from staged setups where people were asked to tell the truth or to lie and their statements were recorded, here we use real-world political debates. Thanks to the efforts of fact-checking organizations, it is possible to obtain annotations for statements in the context of a political discourse as true, half-true, or false. Starting with such data from the CLEF-2018 CheckThat! Lab, which was limited to text, we performed alignment to the corresponding videos, thus producing a multimodal dataset. We further developed a multimodal deep-learning architecture for the task of deception detection, which yielded sizable improvements over the state of the art for the CLEF-2018 Lab task 2. Our experiments show that the use of the acoustic signal consistently helped to improve the performance compared to using textual and metadata features only, based on several different evaluation measures. We release the new dataset to the research community, hoping to help advance the overall field of multimodal deception detection. Daniel Kopev, Ahmed Ali 0002, Ivan Koychev, Preslav Nakov |
ASRU | 2 |
| 2019 | Efficient Arabic Emotion Recognition Using Deep Neural NetworksabstractEmotion recognition from speech signal based on deep learning is an active research area. Convolutional neural networks (CNNs) may be the dominant method in this area. In this paper, we implement two neural architectures to address this problem. The first architecture is an attention-based CNN-LSTM-DNN model. In this novel architecture, the convo-lutional layers extract salient features and the bi-directional long short-term memory (BLSTM) layers handle the sequential phenomena of the speech signal. This is followed by an attention layer, which extracts a summary vector that is fed to the fully connected dense layer (DNN), which finally connects to a softmax output layer. The second architecture is based on a deep CNN model. The results on an Arabic speech emotion recognition task show that our innovative approach can lead to significant improvements (2.2% absolute improvements) over a strong deep CNN baseline system. On the other hand, the deep CNN models are significantly faster than the attention based CNN-LSTM-DNN models in training and classification. Yasser Hifny, Ahmed Ali 0002 |
ICASSP | 2 |
| 2019 | A Factorial Deep Markov Model for Unsupervised Disentangled Representation Learning from SpeechabstractWe present the Factorial Deep Markov Model (FDMM) for representation learning of speech. The FDMM learns disentangled, interpretable and lower dimensional latent representations from speech without supervision. We use a static and dynamic latent variable to exploit the fact that information in a speech signal evolves at different time scales. Latent representations learned by the FDMM outperform a baseline i-vector system on speaker verification and dialect identification while also reducing the error rate of a phone recognition system in a domain mismatch scenario. Sameer Khurana, Shafiq R. Joty, Ahmed Ali 0002, James R. Glass |
ICASSP | 3 |
| 2019 | Domain Attentive Fusion for End-to-end Dialect Identification with Unknown Target DomainabstractEnd-to-end deep learning language or dialect identification systems operate on the spectrogram or other acoustic feature and directly generate identification scores for each class. An important issue for end-to-end systems is to have some knowledge of the application domain, because the system can be vulnerable to use cases that were not seen in the training phase; such a scenario is often referred to as a domain mismatched condition. In general, we assume that there is enough variation in the training dataset to expose the system to multiple domains. In this work, we study how to best make use a training dataset in order to have maximum effectiveness on unknown target domains. Our goal is to process the input without any knowledge of the target domain while preserving robust performance on other domains as well. To accomplish this objective, we propose a domain attentive fusion approach for end-to-end dialect/language identification systems. To help with experimentation, we collect a dataset from three different domains, and create experimental protocols for a domain mismatched condition. The results of our proposed approach, which were tested on a variety of broadcast and YouTube data, shows significant performance gain compared to traditional approaches, even without any prior target domain information. Suwon Shon, Ahmed Ali 0002, James R. Glass |
ICASSP | 2 |
| 2019 | Towards Variability Resistant Dialectal Speech Evaluation
Ahmed Ali 0002, Salam Khalifa, Nizar Habash |
INTERSPEECH | 1 |
| 2019 | Analyzing Phonetic and Graphemic Representations in End-to-End Automatic Speech RecognitionabstractEnd-to-end neural network systems for automatic speech recognition (ASR) are trained from acoustic features to text transcriptions. In contrast to modular ASR systems, which contain separately-trained components for acoustic modeling, pronunciation lexicon, and language modeling, the end-to-end paradigm is both conceptually simpler and has the potential benefit of training the entire system on the end task. However, such neural network models are more opaque: it is not clear how to interpret the role of different parts of the network and what information it learns during training. In this paper, we analyze the learned internal representations in an end-to-end ASR model. We evaluate the representation quality in terms of several classification tasks, comparing phonemes and graphemes, as well as different articulatory features. We study two languages (English and Arabic) and three datasets, finding remarkable consistency in how different properties are represented in different layers of the deep neural network. Yonatan Belinkov, Ahmed Ali 0002, James R. Glass |
INTERSPEECH | 2 |
| 2019 | Predicting the Leading Political Ideology of YouTube Channels Using Acoustic, Textual, and Metadata InformationabstractWe address the problem of predicting the leading political ideology, i.e., left-center-right bias, for YouTube channels of news media. Previous work on the problem has focused exclusively on text and on analysis of the language used, topics discussed, sentiment, and the like. In contrast, here we study videos, which yields an interesting multimodal setup. Starting with gold annotations about the leading political ideology of major world news media from Media Bias/Fact Check, we searched on YouTube to find their corresponding channels, and we downloaded a recent sample of videos from each channel. We crawled more than 1,000 YouTube hours along with the corresponding subtitles and metadata, thus producing a new multimodal dataset. We further developed a multimodal deep-learning architecture for the task. Our analysis shows that the use of acoustic signal helped to improve bias detection by more than 6% absolute over using text and metadata only. We release the dataset to the research community, hoping to help advance the field of multi-modal political bias detection. Yoan Dinkov, Ahmed Ali 0002, Ivan Koychev, Preslav Nakov |
INTERSPEECH | 2 |
| 2019 | FarSpeech: Arabic Natural Language Processing for Live Arabic Speech
Mohamed Eldesouki, Naassih Gopee, Ahmed Ali 0002, Kareem Darwish |
INTERSPEECH | 3 |
| 2018 | Exploiting Convolutional Neural Networks for Phonotactic Based Dialect IdentificationabstractIn this paper, we investigate different approaches for Dialect Identification (DID) in Arabic broadcast speech. Dialects differ in their inventory of phonological segments. This paper proposes a new phonotactic based feature representation approach which enables discrimination among different occurrences of the same phone n-grams with different phone duration and probability statistics. To achieve further gain in accuracy we used multi-lingual phone recognizers, trained separately on Arabic, English, Czech, Hungarian and Russian languages. We use Support Vector Machines (SVMs), and Convolutional Neural Networks (CNN s) as backend classifiers throughout the study. The final system fusion results in 24.7% and 19.0% relative error rate reduction compared to that of a conventional phonotactic DID, and i-vectors with bottleneck features. Maryam Najafian, Sameer Khurana, Suwon Shon, Ahmed Ali 0002, James R. Glass |
ICASSP | 4 |
| 2017 | WERD: Using social text spelling variants for evaluating dialectal speech recognitionabstractWe study the problem of evaluating automatic speech recognition (ASR) systems that target dialectal speech input. A major challenge in this case is that the orthography of dialects is typically not standardized. From an ASR evaluation perspective, this means that there is no clear gold standard for the expected output, and several possible outputs could be considered correct according to different human annotators, which makes standard word error rate (WER) inadequate as an evaluation metric. Such a situation is typical for machine translation (MT), and thus we borrow ideas from an MT evaluation metric, namely TERp, an extension of translation error rate which is closely-related to WER. In particular, in the process of comparing a hypothesis to a reference, we make use of spelling variants for words and phrases, which we mine from Twitter in an unsupervised fashion. Our experiments with evaluating ASR output for Egyptian Arabic, and further manual analysis, show that the resulting WERd (i.e., WER for dialects) metric, a variant of TERp, is more adequate than WER for evaluating dialectal ASR. Ahmed Ali 0002, Preslav Nakov, Peter Bell 0001, Steve Renals |
ASRU | 1 |
| 2017 | Speech recognition challenge in the wild: Arabic MGB-3abstractThis paper describes the Arabic MGB-3 Challenge - Arabic Speech Recognition in the Wild. Unlike last year's Arabic MGB-2 Challenge, for which the recognition task was based on more than 1,200 hours broadcast TV news recordings from Aljazeera Arabic TV programs, MGB-3 emphasises dialectal Arabic using a multi-genre collection of Egyptian YouTube videos. Seven genres were used for the data collection: comedy, cooking, family/kids, fashion, drama, sports, and science (TEDx). A total of 16 hours of videos, split evenly across the different genres, were divided into adaptation, development and evaluation data sets. The Arabic MGB-Challenge comprised two tasks: A) Speech transcription, evaluated on the MGB-3 test set, along with the 10 hour MGB-2 test set to report progress on the MGB-2 evaluation; B) Arabic dialect identification, introduced this year in order to distinguish between four major Arabic dialects - Egyptian, Levantine, North African, Gulf, as well as Modern Standard Arabic. Two hours of audio per dialect were released for development and a further two hours were used for evaluation. For dialect identification, both lexical features and i-vector bottleneck features were shared with participants in addition to the raw audio recordings. Overall, thirteen teams submitted ten systems to the challenge. We outline the approaches adopted in each system, and summarise the evaluation results. Ahmed Ali 0002, Stephan Vogel, Steve Renals |
ASRU | 1 |
| 2017 | Automatic speech recognition of Arabic multi-genre broadcast mediaabstractThis paper describes an Arabic Automatic Speech Recognition system developed on 15 hours of Multi-Genre Broadcast (MGB-3) data from YouTube, plus 1,200 hours of Multi-Dialect and Multi-Genre MGB-2 data recorded from the Aljazeera Arabic TV channel. In this paper, we report our investigations of a range of signal pre-processing, data augmentation, topic-specific language model adaptation, accent specific re-training, and deep learning based acoustic modeling topologies, such as feed-forward Deep Neural Networks (DNNs), Time-delay Neural Networks (TDNNs), Long Short-term Memory (LSTM) networks, Bidirectional LSTMs (BLSTMs), and a Bidirectional version of the Prioritized Grid LSTM (BPGLSTM) model. We propose a system combination for three purely sequence trained recognition systems based on lattice-free maximum mutual information, 4-gram language model re-scoring, and system combination using the minimum Bayes risk decoding criterion. The best word error rate we obtained on the MGB-3 Arabic development set using a 4-gram re-scoring strategy is 42.25% for a chain BLSTM system, compared to 65.44% baseline for a DNN system. Maryam Najafian, Wei-Ning Hsu, Ahmed Ali 0002, James R. Glass |
ASRU | 3 |
| 2017 | MIT-QCRI Arabic dialect identification system for the 2017 multi-genre broadcast challengeabstractIn order to successfully annotate the Arabic speech content found in open-domain media broadcasts, it is essential to be able to process a diverse set of Arabic dialects. For the 2017 Multi-Genre Broadcast challenge (MGB-3) there were two possible tasks: Arabic speech recognition, and Arabic Dialect Identification (ADI). In this paper, we describe our efforts to create an ADI system for the MGB-3 challenge, with the goal of distinguishing amongst four major Arabic dialects, as well as Modern Standard Arabic. Our research focused on dialect variability and domain mismatches between the training and test domain. In order to achieve a robust ADI system, we explored both Siamese neural network models to learn similarity and dissimilarities among Arabic dialects, as well as i-vector post-processing to adapt domain mismatches. Both Acoustic and linguistic features were used for the final MGB-3 submissions, with the best primary system achieving 75% accuracy on the official 10hr test set. Suwon Shon, Ahmed Ali 0002, James R. Glass |
ASRU | 2 |
| 2017 | QMDIS: QCRI-MIT Advanced Dialect Identification System
Sameer Khurana, Maryam Najafian, Ahmed Ali 0002, Tuka Al Hanai, Yonatan Belinkov, James R. Glass |
INTERSPEECH | 3 |
| 2016 | Automatic Dialect Detection in Arabic Broadcast SpeechabstractWe investigate different approaches for dialect identification in Arabic broadcast speech, using phonetic, lexical features obtained from a speech recognition system, and acoustic features using the i-vector framework. We studied both generative and discriminate classifiers, and we combined these features using a multi-class Support Vector Machine (SVM). We validated our results on an Arabic/English language identification task, with an accuracy of 100%. We used these features in a binary classifier to discriminate between Modern Standard Arabic (MSA) and Dialectal Arabic, with an accuracy of 100%. We further report results using the proposed method to discriminate between the five most widely used dialects of Arabic: namely Egyptian, Gulf, Levantine, North African, and MSA, with an accuracy of 52%. We discuss dialect identification errors in the context of dialect code-switching between Dialectal Arabic and MSA, and compare the error pattern between manually labeled data, and the output from our classifier. We also release the train and test data as standard corpus for dialect identification. Ahmed Ali 0002, Najim Dehak, Patrick Cardinal, Sameer Khurana, Sree Harsha Yella, James R. Glass, Peter Bell 0001, Steve Renals |
INTERSPEECH | 1 |
| 2016 | The MGB-2 challenge: Arabic multi-dialect broadcast media recognitionabstractThis paper describes the Arabic Multi-Genre Broadcast (MGB-2) Challenge for SLT-2016. Unlike last year's English MGB Challenge, which focused on recognition of diverse TV genres, this year, the challenge has an emphasis on handling the diversity in dialect in Arabic speech. Audio data comes from 19 distinct programmes from the Aljazeera Arabic TV channel between March 2005 and December 2015. Programmes are split into three groups: conversations, interviews, and reports. A total of 1,200 hours have been released with lightly supervised transcriptions for the acoustic modelling. For language modelling, we made available over 110M words crawled from Aljazeera Arabic website Aljazeera.net for a 10 year duration 2000-2011. Two lexicons have been provided, one phoneme based and one grapheme based. Finally, two tasks were proposed for this year's challenge: standard speech transcription, and word alignment. This paper describes the task data and evaluation process used in the MGB challenge, and summarises the results obtained. Ahmed Ali 0002, Peter Bell 0001, James R. Glass, Yacine Messaoui, Hamdy Mubarak, Steve Renals |
SLT | 1 |
| 2016 | QCRI advanced transcription system (QATS) for the Arabic Multi-Dialect Broadcast media recognition: MGB-2 challengeabstractIn this paper, we describe Qatar Computing Research Institute's (QCRI) speech transcription system for the 2016 Dialectal Arabic Multi-Genre Broadcast (MGB-2) challenge. MGB-2 is a controlled evaluation using 1,200 hours audio with lightly supervised transcription Our system which was a combination of three purely sequence trained recognition systems, achieved the lowest WER of 14.2% among the nine participating teams. Key features of our transcription system are: purely sequence trained acoustic models using the recently introduced Lattice free Maximum Mutual Information (LF-MMI) modeling framework; Language model rescoring using a four-gram and Recurrent Neural Network with Max- Ent connections (RNNME) language models; and system combination using Minimum Bayes Risk (MBR) decoding criterion. The whole system is built using kaldi speech recognition toolkit. Sameer Khurana, Ahmed Ali 0002 |
SLT | 2 |
| 2015 | Multi-reference WER for evaluating ASR for languages with no orthographic rulesabstractLanguages with no standard orthographic representation faces a challenge to evaluate the output from Automatic Speech Recognition (ASR). Since the reference transcription text can vary widely from one user to another. We propose an innovative approach for evaluating speech recognition using Multi-References. For each recognized speech segments, we ask five different users to transcribe the speech. We combine the alignment for the multiple references, and use the combined alignment to report a modified version of Word Error Rate (WER). This approach is in favor of accepting a recognized word if any of the references typed it in the same form. Results are reported using two Dialectal Arabic (DA) as a language with no standard orthographic; Egyptian, and North African speech. The average WER for the five references individually is 71.4%, and 80.1% respectively. When considering all references combined, the Multi-References MR-WER was found to be 39.7%, and 45.9% respectively. Ahmed Ali 0002, Walid Magdy, Peter Bell 0001, Steve Renals |
ASRU | 1 |
| 2015 | QAT2 - the QCRI advanced transcription and translation system
Ahmed Abdelali, Ahmed Ali 0002, Francisco Guzmán, Felix Stahlberg, Stephan Vogel |
INTERSPEECH | 2 |
| 2015 | Crowdsource a little to label a lot: labeling a speech corpus of dialectal Arabic
Samantha Wray, Ahmed Ali 0002 |
INTERSPEECH | 2 |
| 2014 | Recent advances in ASR applied to an Arabic transcription system for Al-JazeeraabstractThis paper describes a detailed comparison of several state-of-the-art speech recognition techniques applied to a limited Ara-bic broadcast news dataset. The different approaches were all trained on 50 hours of transcribed audio from the Al-Jazeera news channel. The best results were obtained using i-vector-based speaker adaptation in a training scenario using the Min-imum Phone Error (MPE) criteria combined with sequential Deep Neural Network (DNN) training. We report results for two different types of test data: broadcast news reports, with a best word error rate (WER) of 17.86%, and a broadcast conver-sations with a best WER of 29.85%. The overall WER on this test set is 25.6%. Index Terms: Arabic, ASR system, Kaldi 1. Patrick Cardinal, Ahmed Ali 0002, Najim Dehak, Yu Zhang 0033, Tuka Al Hanai, James R. Glass, Stephan Vogel |
INTERSPEECH | 2 |
| 2014 | A complete KALDI recipe for building Arabic speech recognition systemsabstractIn this paper we present a recipe and language resources for training and testing Arabic speech recognition systems using the KALDI toolkit. We built a prototype broadcast news system using 200 hours GALE data that is publicly available through LDC. We describe in detail the decisions made in building the system: using the MADA toolkit for text normalization and vowelization; why we use 36 phonemes; how we generate pronunciations; how we build the language model. We report results using state-of-the-art modeling and decoding techniques. The scripts are released through KALDI and resources are made available on QCRI's language resources web portal. This is the first effort to share reproducible sizable training and testing results on MSA system. Ahmed Ali 0002, Patrick Cardinal, Najim Dehak, Stephan Vogel, James R. Glass |
SLT | 1 |
| 2014 | Non-Negative Factor Analysis of Gaussian Mixture Model Weight Adaptation for Language and Dialect RecognitionabstractRecent studies show that Gaussian mixture model (GMM) weights carry less, yet complimentary, information to GMM means for language and dialect recognition. However, state-of-the-art language recognition systems usually do not use this information. In this research, a non-negative factor analysis (NFA) approach is developed for GMM weight decomposition and adaptation. This modeling, which is conceptually simple and computationally inexpensive, suggests a new low-dimensional utterance representation method using a factor analysis similar to that of the i-vector framework. The obtained subspace vectors are then applied in conjunction with i-vectors to the language/dialect recognition problem. The suggested approach is evaluated on the NIST 2011 and RATS language recognition evaluation (LRE) corpora and on the QCRI Arabic dialect recognition evaluation (DRE) corpus. The assessment results show that the proposed adaptation method yields more accurate recognition results compared to three conventional weight adaptation approaches, namely maximum likelihood re-estimation, non-negative matrix factorization, and a subspace multinomial model. Experimental results also show that the intermediate-level fusion of i-vectors and NFA subspace vectors improves the performance of the state-of-the-art i-vector framework especially for the case of short utterances. Mohamad Hasan Bahari, Najim Dehak, Hugo Van hamme, Lukás Burget, Ahmed Ali 0002, James R. Glass |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2012 | A summarization tool for time-sensitive social mediaabstractSearching social content in general and microblogs (aka tweets) in particular has been basic and limited, especially for time-sensitive topics. The currently implemented microblog search on sites such as Twitter is based on simple word matching and retrieves the most recent microblogs that match a given query. Furthermore, a user may obtain hundreds or perhaps thousands of microblogs in response to a given query, leading to information overload. We present a new multidimensional microblog search tool that generates a comprehensive report from microblogs instead of a flat list of recent/relevant microblogs for a given query. Reports may include tag-clouds, topic time series, and most popular and funny microblogs, etc. The tool can be configured for monitoring time-sensitive topics using a set of predefined queries. We demonstrate our system on Arabic and English microblog collections. Additionally, we show a special configuration of the system for monitoring the 2012 Egyptian presidential elections. Walid Magdy, Ahmed Ali 0002, Kareem Darwish |
CIKM | 2 |