EDBT 2026 Demo / reviewers in the wild / expert
Mostafa Shahin
dblp:79/9262
· DBLP profile ↗
27ranked-venue papers
14as first author
11since 2021 · last 2026
0000-0002-1091-8531ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 22 · 13 first-author · 10 since 2021Artificial intelligence and machine learning · 16 · 7 first-author · 8 since 2021Computer networks · 1Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AusKidTalk: Developing Transcription Guidelines for Continuous Australian English Child Speech
Tünde Szalay, Zheng Nan, Renata Huang, Mostafa Shahin, Tharmakulasingam Sirojan, Kirrie J. Ballard, Beena Ahmed |
LREC | 4 |
| 2025 | Multi-Class Dementia Detection Using Acoustic Features - ICASSP-2025 PROCESS ChallengeabstractThis paper describes our best-performing submission for the ICASSP-2025 Signal Processing Grand Challenge PROCESS, focused on the classification of speech into 3 groups - Healthy, Mild Cognitive Impairment (MCI), and Dementia - using three speech tasks in English. Our approach was aligned with the aim of simple, preclinical detection of dementia, employing a minimal set of acoustic features, and no linguistic analysis. We built an ensemble classifier based on 1) Selected features from the ComParE acoustic feature set and 2) knowledge-based rules for combining predictions across the 3 tasks, using a two-tier majority vote system. Our technique outperformed the baseline results by a large margin, achieving a macro-F1 of 0.96 on the development set, 0.99 on 5-fold cross-validation and 0.64 on the test set. M. Abdullah Zafar, Xiangyu Zhang 0005, Mostafa Shahin, Beena Ahmed |
ICASSP | 3 |
| 2025 | Rethinking Mamba in Speech Processing by Self-Supervised ModelsabstractThe Mamba-based model has demonstrated outstanding performance across tasks in computer vision, natural language processing, and speech processing. However, in the realm of speech processing, the Mamba-based model’s performance varies across different tasks. For instance, in tasks such as speech enhancement and spectrum reconstruction, the Mamba model performs well when used independently. However, for tasks like speech recognition, additional modules are required to surpass the performance of attention-based models. We propose the hypothesis that the Mamba-based model excels in "reconstruction" tasks within speech processing. However, for "classification tasks" such as Speech Recognition, additional modules are necessary to accomplish the "reconstruction" step. To validate our hypothesis, we analyze the previous Mamba-based Speech Models from an information theory perspective. Furthermore, we leveraged the properties of HuBERT in our study. We trained a Mamba-based HuBERT model, and the mutual information patterns, along with the model’s performance metrics, confirmed our assumptions. Xiangyu Zhang 0005, Mostafa Shahin, Beena Ahmed, Julien Epps |
ICASSP | 3 |
| 2025 | Auto-Landmark: Acoustic Landmark Dataset and Open-Source Toolkit for Landmark Extraction
Xiangyu Zhang 0005, Daijiao Liu, Tianyi Xiao, Cihan Xiao, Tünde Szalay, Mostafa Shahin, Beena Ahmed, Julien Epps |
INTERSPEECH | 6 |
| 2025 | Towards a Unified Benchmark for Arabic Pronunciation Assessment: Qur'anic Recitation as Case Study
Yassine El Kheir, Omnia Ibrahim, Amit Meghanani, Nada Almarwani, Hawau Olamide Toyin, Sadeen Alharbi, Modar Alfadly, Lamya Alkanhal, Ibrahim Selim, Shehab Elbatal, Salima Mdhaffar, Thomas Hain, Yasser Hifny, Mostafa Shahin, Ahmed Ali 0002 |
INTERSPEECH | 14 |
| 2025 | AusKidTalk: Using Strategic Data Collection and Out-of-Domain Tools to Semi-Automate Novel Corpora Annotation
Tünde Szalay, Mostafa Shahin, Tharmakulasingam Sirojan, Zheng Nan, Renata Huang, Kirrie J. Ballard, Beena Ahmed |
INTERSPEECH | 2 |
| 2025 | Phonological level wav2vec2-based Mispronunciation Detection and Diagnosis methodabstractThe automatic identification and analysis of pronunciation errors, known as Mispronunciation Detection and Diagnosis (MDD) plays a crucial role in Computer Aided Pronunciation Learning (CAPL) tools such as Second-Language (L2) learning or speech therapy applications. Existing MDD methods relying on analysing phonemes can only detect categorical errors of phonemes that have an adequate amount of training data to be modelled. Due to the unpredictable nature of pronunciation errors made by non-native or disordered speakers and the scarcity of training datasets, it is unfeasible to model all types of mispronunciations. Moreover, phoneme-level MDD approaches can provide only limited diagnostic information about the error made. To address this, in this paper, we propose a low-level MDD approach based on the detection of phonological features. Phonological features break down phoneme production into elementary components that are directly related to the articulatory system leading to more formative feedback for the learner. We further propose a multi-label variant of the Connectionist Temporal Classification (CTC) approach to jointly model the non-mutually exclusive phonological features using a single model. The pre-trained wav2vec2 model was employed as a core model for the phonological feature detector. The proposed method was applied to L2 speech corpora collected from English learners from different native languages. The proposed phonological level MDD method was further compared to the traditional phoneme-level MDD and achieved a significantly lower False Acceptance Rate (FAR), False Rejection Rate (FRR), and Diagnostic Error Rate (DER) over all phonological features compared to the phoneme-level equivalent. Mostafa Shahin, Julien Epps, Beena Ahmed |
Speech Commun. | 1 |
| 2024 | Phonological-Level Mispronunciation Detection and Diagnosis
Mostafa Shahin, Beena Ahmed |
INTERSPEECH | 1 |
| 2023 | Improving wav2vec2-based Spoken Language Identification by Learning Phonological Features
Mostafa Shahin, Zheng Nan, Vidhyasaharan Sethu, Beena Ahmed |
INTERSPEECH | 1 |
| 2022 | Knowledge of accent differences can be used to predict speech recognition
Tünde Szalay, Mostafa Shahin, Beena Ahmed, Kirrie J. Ballard |
INTERSPEECH | 2 |
| 2021 | AusKidTalk: An Auditory-Visual Corpus of 3- to 12-Year-Old Australian Children's SpeechabstractHere we present AusKidTalk [1], an audio-visual (AV) corpus of Australian children’s speech collected to facilitate the development of speech based technological solutions for children. It builds upon the technology and expertise developed through the collection of an earlier corpus of Australian adult speech, AusTalk [2,3]. This multi-site initiative was established to remedy the dire shortage of children’s speech corpora in Australia and around the world that are sufficiently sized to train accurate automated speech processing tools for children. We are collecting ~600 hours of speech from children aged 3–12 years that includes single word and sentence productions as well as narrative and emotional speech. In this paper, we discuss the key requirements for AusKidTalk and how we designed the recording setup and protocol to meet them. We also discuss key findings from our feasibility study of the recording protocol, recording tools, and user interface. Beena Ahmed, Kirrie J. Ballard, Denis Burnham, Tharmakulasingam Sirojan, Hadi Mehmood, Dominique Estival, Elise Baker, Felicity Cox, Joanne Arciuli, Titia Benders, Katherine Demuth, Barbara Kelly, Chloé Diskin-Holdaway, Mostafa Shahin, Vidhyasaharan Sethu, Julien Epps, Chwee Beng Lee, Eliathamby Ambikairajah |
Interspeech | 14 |
| 2020 | UNSW System Description for the Shared Task on Automatic Speech Recognition for Non-Native Children's Speech
Mostafa Shahin, Renée Lu, Julien Epps, Beena Ahmed |
INTERSPEECH | 1 |
| 2019 | Anomaly detection based pronunciation verification approach using speech attribute features
Mostafa Shahin, Beena Ahmed |
Speech Commun. | 1 |
| 2018 | Deep Recurrent Electricity Theft Detection in AMI Networks with Random Tuning of Hyper-parametersabstractModern smart grids rely on advanced metering infrastructure (AMI) networks for monitoring and billing purposes. However, such an approach suffers from electricity theft cyberattacks. Different from the existing research that utilizes shallow, static, and customer-specific-based electricity theft detectors, this paper proposes a generalized deep recurrent neural network (RNN)-based electricity theft detector that can effectively thwart these cyberattacks. The proposed model exploits the time series nature of the customers' electricity consumption to implement a gated recurrent unit (GRU)-RNN, hence, improving the detection performance. In addition, the proposed RNN-based detector adopts a random search analysis in its learning stage to appropriately fine-tune its hyper-parameters. Extensive test studies are carried out to investigate the detector's performance using publicly available real data of 107,200 energy consumption days from 200 customers. Simulation results demonstrate the superior performance of the proposed detector compared with state-of-the-art electricity theft detectors. Mahmoud Nabil 0001, Muhammad Ismail 0001, Mohamed Mahmoud 0001, Mostafa Shahin, Khalid A. Qaraqe, Erchin Serpedin |
ICPR | 4 |
| 2018 | One-Class SVMs Based Pronunciation Verification ApproachabstractThe automatic assessment of speech plays an important role in Computer Aided Pronunciation Learning systems. However, modeling both the correct and incorrect pronunciation of each phoneme to achieve accurate pronunciation verification is unfeasible due to the lack of sufficient mispronounced samples in training datasets. In this paper, we propose a novel approach that handles this unbalanced data distribution by building multiple one-class SVMs to evaluate each phoneme as correct or incorrect. We model the correct pronunciation of each individual phoneme with a one-class SVM trained using a set of speech attributes features, namely the manner and place of articulation. These features are extracted from a bank of pre-trained DNN speech attributes classifiers. The one-class SVM model measures the similarity between the new data and the training set and then classifies it as normal (correct) or an anomaly (incorrect). We evaluated the system using native speech corpus and disordered speech corpus and compared it with the conventional Goodness of Pronunciation (GOP) algorithm. The results show that our approach reduces the false-acceptance and false-rejection rates by around 26% and 39% respectively. Mostafa Shahin, Jim Xiuquan Ji, Beena Ahmed |
ICPR | 1 |
| 2018 | Anomaly Detection Approach for Pronunciation Verification of Disordered Speech Using Speech Attribute Features
Mostafa Shahin, Beena Ahmed, Jim Xiuquan Ji, Kirrie J. Ballard |
INTERSPEECH | 1 |
| 2018 | Efficient detection of electricity theft cyber attacks in AMI networksabstractAdvanced metering infrastructure (AMI) networks are vulnerable against electricity theft cyber attacks. Different from the existing research that exploits shallow machine learning architectures for electricity theft detection, this paper proposes a deep neural network (DNN)-based customer-specific detector that can efficiently thwart such cyber attacks. The proposed DNN-based detector implements a sequential grid search analysis in its learning stage to appropriately fine tune its hyper-parameters, hence, improving the detection performance. Extensive test studies are carried out based on publicly available real energy consumption data of 5000 customers and the detector's performance is investigated against a mixture of different types of electricity theft cyber attacks. Simulation results demonstrate a significant performance improvement compared with state-of-the-art shallow detectors. Muhammad Ismail 0001, Mostafa Shahin, Mostafa F. Shaaban, Erchin Serpedin, Khalid A. Qaraqe |
WCNC | 2 |
| 2017 | Deep Learning and Insomnia: Assisting Clinicians With Their DiagnosisabstractEffective sleep analysis is hampered by the lack of automated tools catering to disordered sleep patterns and cumbersome monitoring hardware. In this paper, we apply deep learning on a set of 57 EEG features extracted from a maximum of two EEG channels to accurately differentiate between patients with insomnia or controls with no sleep complaints. We investigated two different approaches to achieve this. The first approach used EEG data from the whole sleep recording irrespective of the sleep stage (stage-independent classification), while the second used only EEG data from insomnia-impacted specific sleep stages (stage-dependent classification). We trained and tested our system using both healthy and disordered sleep collected from 41 controls and 42 primary insomnia patients. When compared with manual assessments, an NREM + REM based classifier had an overall discrimination accuracy of 92% and 86% between two groups using both two and one EEG channels, respectively. These results demonstrate that deep learning can be used to assist in the diagnosis of sleep disorders such as insomnia. Mostafa Shahin, Beena Ahmed, Sana Tmar Ben Hamida, Lamana Mulaffer, Martin Glos, Thomas Penzel |
IEEE J. Biomed. Health Informatics | 1 |
| 2016 | Classification of bisyllabic lexical stress patterns in disordered speech using deep learningabstractTechnology-based therapy tools can be of great benefit to children with developmental speech disabilities as they typically require sustained practice with a speech therapist for several years. Towards this aim, over the past 4 years we have developed speech processing tools to automatically detect common errors in disordered speech. This paper presents an automated technique to identify incorrect lexical stress. Specifically, we describe a deep neural network (DNN) that can be used to classify the four different bisyllabic stress patterns: strong-weak (SW), weak-strong (WS), strong-strong (SS) and weak-weak (WW). We derive input features for the DNN from the duration, pitch, intensity and spectral energy on each of the two consecutive syllables. Using these features, we achieve 93% correct classification between SW/WS stress patterns and 88% correct classification of the four bisyllabic patterns on speech from typically developing children, while we obtain 73.4% classification between SW/WS in disordered speech. These figures represent a two-fold reduction in error rates compared to our prior work, which used a DNN with differential features from consecutive syllables. Mostafa Shahin, Ricardo Gutierrez-Osuna, Beena Ahmed |
ICASSP | 1 |
| 2016 | Automatic Classification of Lexical Stress in English and Arabic Languages Using Deep Learning
Mostafa Shahin, Julien Epps, Beena Ahmed |
INTERSPEECH | 1 |
| 2015 | Tabby Talks: An automated tool for the assessment of childhood apraxia of speech
Mostafa Shahin, Beena Ahmed, Avinash Parnandi 0001, Virendra Karappa, Jacqueline McKechnie, Kirrie J. Ballard, Ricardo Gutierrez-Osuna |
Speech Commun. | 1 |
| 2014 | A comparison of GMM-HMM and DNN-HMM based pronunciation verification techniques for use in the assessment of childhood apraxia of speechabstractThis paper introduces a pronunciation verification method to be used in an automatic assessment therapy tool of child disordered speech. The proposed method creates a phonebased search lattice that is flexible enough to cover all probable mispronunciations. This allows us to verify the correctness of the pronunciation and detect the incorrect phonemes produced by the child. We compare between two different acoustic models, the conventional GMM-HMM and the hybrid DNN-HMM. Results show that the hybrid DNNHMM outperforms the conventional GMM-HMM for all experiments on both normal and disordered speech. The total correctness accuracy of the system at the phoneme level is above 85% when used with disordered speech. Mostafa Shahin, Beena Ahmed, Jacqueline McKechnie, Kirrie J. Ballard, Ricardo Gutierrez-Osuna |
INTERSPEECH | 1 |
| 2014 | Classification of lexical stress patterns using deep neural network architectureabstractLexical stress is a key diagnostic marker of disordered speech as it strongly affects speech perception. In this paper we introduce an automated method to classify between the different lexical stress patterns in children's speech. A deep neural network is used to classify between strong-weak (SW), weak-strong (WS) and equal-stress (SS/WW) patterns in English by measuring the articulation change between the two successive syllables. The deep neural network architecture is trained using a set of acoustic features derived from pitch, duration and intensity measurements along with the energies in different frequency bands. We compared the performance of the deep neural classifier to a traditional single hidden layer MLP. Results show that the deep neural classifier outperforms the traditional MLP. The accuracy of the deep neural system is approximately 85% when classifying between the unequal stress patterns (SW/WS) and greater than 70% when classifying both equal and unequal stress patterns. Mostafa Shahin, Beena Ahmed, Kirrie J. Ballard |
SLT | 1 |
| 2013 | Architecture of an automated therapy tool for childhood apraxia of speechabstractWe present a multi-tier system for the remote administration of speech therapy to children with apraxia of speech. The system uses a client-server architecture model and facilitates task-oriented remote therapeutic training in both in-home and clinical settings. Namely, the system allows a speech therapist to remotely assign speech production exercises to each child through a web interface, and the child to practice these exercises on a mobile device. The mobile app records the child's utterances and streams them to a back-end server for automated scoring by a speech-analysis engine. The therapist can then review the individual recordings and the automated scores through a web interface, provide feedback to the child, and adapt the training program as needed. We validated the system through a pilot study with children diagnosed with apraxia of speech, and their parents and speech therapists. Here we describe the overall client-server architecture, middleware tools used to build the system, the speech-analysis tools for automatic scoring of recorded utterances, and results from the pilot study. Our results support the feasibility of the system as a complement to traditional face-to-face therapy through the use of mobile tools and automated speech analysis algorithms. Avinash Parnandi 0001, Virendra Karappa, Youngpyo Son, Mostafa Shahin, Jacqueline McKechnie, Kirrie J. Ballard, Beena Ahmed, Ricardo Gutierrez-Osuna |
ASSETS | 4 |
| 2012 | Automatic classification of unequal lexical stress patterns using machine learning algorithmsabstractTechnology based speech therapy systems are severely handicapped due to the absence of accurate prosodic event identification algorithms. This paper introduces an automatic method for the classification of strong-weak (SW) and weak-strong (WS) stress patterns in children speech with American English accent, for use in the assessment of the speech dysprosody. We investigate the ability of two sets of features used to train classifiers to identify the variation in lexical stress between two consecutive syllables. The first set consists of traditional features derived from measurements of pitch, intensity and duration, whereas the second set consists of energies of different filter banks. Three different classifiers were used in the experiments: an Artificial Neural Network (ANN) classifier with a single hidden layer, Support Vector Machine (SVM) classifier with both linear and Gaussian kernels and the Maximum Entropy modeling (MaxEnt). these features. Best results were obtained using an ANN classifier and a combination of the two sets of features. The system correctly classified 94% of the SW stress patterns and 76% of the WS stress patterns. Mostafa Shahin, Beena Ahmed, Kirrie J. Ballard |
SLT | 1 |
| 2006 | Computer aided pronunciation learning system using speech recognition techniquesabstractThis paper describes a speech-enabled Computer Aided Pronunciation Learning (CAPL) system HAFSS © . This system was developed for teaching Arabic pronunciations to non-native speakers. A challenging application of HAFSS © is teaching the correct recitation of the holy Qur'an. HAFSS © uses a state of the art speech recognizer to detect errors in user recitation. To increase accuracy of the speech recognizer, only probable pronunciation variants, that cover all common types of recitation errors, are examined by the speech decoder. A module for the automatic generation of pronunciation hypotheses is built as a component of the system. A phoneme duration classification algorithm is implemented to detect recitation errors related to phoneme durations. The decision reached by the recognizer is accompanied by a confidence score to reduce effect of misleading system feedbacks to unpredictable speech inputs. Performance evaluation using a data set that includes 6.6% wrong speech segments showed that the system correctly identified the error in 62.4% of pronunciation errors, reported Repeat Request for 22.4% of the errors and made false acceptance of 14.9% of total errors. Index Terms: Pronunciation error detection, Qur’an recitation Sherif M. Abdou, Salah Eldeen Hamid, Mohsen A. Rashwan, Abdurrahman Samir, Ossama Abdel-Hamid, Mostafa Shahin, Waleed Nazih |
INTERSPEECH | 6 |
| 2006 | Building Annotated Written and Spoken Arabic LRs in NEMLAR Project
Mustafa Yaseen, Bente Maegaard, Khalid Choukri, Niklas Paulsson, S. Haamid, Steven Krauwer, Chomicha Bendahman, Hanne Fersøe, Mohsen A. Rashwan, Bassam Haddad, Chafic Mokbel, Abdelhak Mouradi, A. Al-Kufaishi, Mostafa Shahin, Noureddine Chenfour, Ahmed Ragheb |
LREC | 15 |