VLDB 2026 Research / reviewers in the wild / expert
Beena Ahmed
dblp:77/3444
· DBLP profile ↗
39ranked-venue papers
4as first author
18since 2021 · last 2026
0000-0002-1240-6572ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 25 · 2 first-author · 14 since 2021Artificial intelligence and machine learning · 17 · 2 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 6 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AusKidTalk: Developing Transcription Guidelines for Continuous Australian English Child Speech
Tünde Szalay, Zheng Nan, Renata Huang, Mostafa Shahin, Tharmakulasingam Sirojan, Kirrie J. Ballard, Beena Ahmed |
LREC | 7 |
| 2025 | Improved Out-of-domain Detection in VAE Latent Spaces with Boundary-driven RegularisationabstractIn out-of-domain (OOD) detection tasks, encoding the actual data into a suitable latent space could be beneficial since it may facilitate measurement of the spatial relationship between in-domain (IND) and OOD data. However, any such mapping of data to a latent space carries the risk that some OOD points may be mapped to in-domain regions. To address this drawback we propose a novel method that spatially separates these two domains in the latent space. It first geometrically defines the boundary of IND in the latent space and then forces some enumerated OOD data (known as surrogate OOD) to fit that boundary. This then helps the encoder to map OOD to the surrounding area of IND, consequently reducing any overlap. Following this, accurate OOD detection is achieved by a dedicated detector distinguishing IND and its boundary. We illustrate the effects of the proposed method on synthetic data and validated it over both grey-scale and RGB image datasets. Miao Jing, Vidhyasaharan Sethu, Beena Ahmed |
ICASSP | 3 |
| 2025 | Evidential Neural GPLDA: A Novel Approach to Quantify Prediction Uncertainty in Speaker Verification SystemsabstractThe uncertainty of an automatic speaker verification (ASV) system is typically estimated using its overall accuracy. However it fails to express "when" the system is uncertain in a predictive and case-by-case manner. Also, prior to interpreting each prediction made by ASV systems, there is a need to assess if the system is confident about the prediction, which remains less explored in current research. Given the sense that uncertainty of this notion should be associated with the knowledge level held by the system, we propose an Evidential Neural GPLDA back-end inspired by evidential deep learning. This approach quantifies the uncertainty in each prediction based on the density of the training data supporting that prediction. This is achieved by showing if the input representation is similar to those of the training data samples, resulting in a confidence estimate based on training data only. We find loss functions and out-of-domain samples to train the proposed model such that it parameterizes a sharp Dirichlet distribution as low uncertainty and a flat one as high uncertainty. Experiments show that the proposed novel back-end effectively quantifies uncertainty, providing an estimate that aligns with the error rate. Miao Jing, Vidhyasaharan Sethu, Beena Ahmed |
ICASSP | 3 |
| 2025 | Multi-Class Dementia Detection Using Acoustic Features - ICASSP-2025 PROCESS ChallengeabstractThis paper describes our best-performing submission for the ICASSP-2025 Signal Processing Grand Challenge PROCESS, focused on the classification of speech into 3 groups - Healthy, Mild Cognitive Impairment (MCI), and Dementia - using three speech tasks in English. Our approach was aligned with the aim of simple, preclinical detection of dementia, employing a minimal set of acoustic features, and no linguistic analysis. We built an ensemble classifier based on 1) Selected features from the ComParE acoustic feature set and 2) knowledge-based rules for combining predictions across the 3 tasks, using a two-tier majority vote system. Our technique outperformed the baseline results by a large margin, achieving a macro-F1 of 0.96 on the development set, 0.99 on 5-fold cross-validation and 0.64 on the test set. M. Abdullah Zafar, Xiangyu Zhang 0005, Mostafa Shahin, Beena Ahmed |
ICASSP | 4 |
| 2025 | Rethinking Mamba in Speech Processing by Self-Supervised ModelsabstractThe Mamba-based model has demonstrated outstanding performance across tasks in computer vision, natural language processing, and speech processing. However, in the realm of speech processing, the Mamba-based model’s performance varies across different tasks. For instance, in tasks such as speech enhancement and spectrum reconstruction, the Mamba model performs well when used independently. However, for tasks like speech recognition, additional modules are required to surpass the performance of attention-based models. We propose the hypothesis that the Mamba-based model excels in "reconstruction" tasks within speech processing. However, for "classification tasks" such as Speech Recognition, additional modules are necessary to accomplish the "reconstruction" step. To validate our hypothesis, we analyze the previous Mamba-based Speech Models from an information theory perspective. Furthermore, we leveraged the properties of HuBERT in our study. We trained a Mamba-based HuBERT model, and the mutual information patterns, along with the model’s performance metrics, confirmed our assumptions. Xiangyu Zhang 0005, Mostafa Shahin, Beena Ahmed, Julien Epps |
ICASSP | 4 |
| 2025 | Auto-Landmark: Acoustic Landmark Dataset and Open-Source Toolkit for Landmark Extraction
Xiangyu Zhang 0005, Daijiao Liu, Tianyi Xiao, Cihan Xiao, Tünde Szalay, Mostafa Shahin, Beena Ahmed, Julien Epps |
INTERSPEECH | 7 |
| 2025 | AusKidTalk: Using Strategic Data Collection and Out-of-Domain Tools to Semi-Automate Novel Corpora Annotation
Tünde Szalay, Mostafa Shahin, Tharmakulasingam Sirojan, Zheng Nan, Renata Huang, Kirrie J. Ballard, Beena Ahmed |
INTERSPEECH | 7 |
| 2025 | Quantifying prediction uncertainties in automatic speaker verification systemsabstractFor modern automatic speaker verification (ASV) systems, explicitly quantifying the confidence for each prediction strengthens the system’s reliability by indicating in which case the system is with trust. However, current paradigms do not take this into consideration. We thus propose to express confidence in the prediction by quantifying the uncertainty in ASV predictions. This is achieved by developing a novel Bayesian framework to obtain a score distribution for each input. The mean of the distribution is used to derive the decision while the spread of the distribution represents the uncertainty arising from the plausible choices of the model parameters. To capture the plausible choices, we sample the probabilistic linear discriminant analysis (PLDA) back-end model posterior through Hamiltonian Monte-Carlo (HMC) and approximate the embedding model posterior through stochastic Langevin dynamics (SGLD) and Bayes-by-backprop. Given the resulting score distribution, a further quantification and decomposition of the prediction uncertainty are achieved by calculating the score variance, entropy, and mutual information. The quantified uncertainties include the aleatoric uncertainty and epistemic uncertainty (model uncertainty). We evaluate them by observing how they change while varying the amount of training speech, the duration, and the noise level of testing speech. The experiments indicate that the behaviour of those quantified uncertainties reflects the changes we made to the training and testing data, demonstrating the validity of the proposed method as a measure of uncertainty. • The paper emphasises the need for quantifying and separating uncertainties in ASV. • The paper proposes a novel framework incorporating various Bayesian learning methods. • The major cause of epistemic uncertainty is training data size and test data length. • The noise level in the test utterance increases the aleatoric uncertainty in ASV. Miao Jing, Vidhyasaharan Sethu, Beena Ahmed, Kong-Aik Lee |
Comput. Speech Lang. | 3 |
| 2025 | Phonological level wav2vec2-based Mispronunciation Detection and Diagnosis methodabstractThe automatic identification and analysis of pronunciation errors, known as Mispronunciation Detection and Diagnosis (MDD) plays a crucial role in Computer Aided Pronunciation Learning (CAPL) tools such as Second-Language (L2) learning or speech therapy applications. Existing MDD methods relying on analysing phonemes can only detect categorical errors of phonemes that have an adequate amount of training data to be modelled. Due to the unpredictable nature of pronunciation errors made by non-native or disordered speakers and the scarcity of training datasets, it is unfeasible to model all types of mispronunciations. Moreover, phoneme-level MDD approaches can provide only limited diagnostic information about the error made. To address this, in this paper, we propose a low-level MDD approach based on the detection of phonological features. Phonological features break down phoneme production into elementary components that are directly related to the articulatory system leading to more formative feedback for the learner. We further propose a multi-label variant of the Connectionist Temporal Classification (CTC) approach to jointly model the non-mutually exclusive phonological features using a single model. The pre-trained wav2vec2 model was employed as a core model for the phonological feature detector. The proposed method was applied to L2 speech corpora collected from English learners from different native languages. The proposed phonological level MDD method was further compared to the traditional phoneme-level MDD and achieved a significantly lower False Acceptance Rate (FAR), False Rejection Rate (FRR), and Diagnostic Error Rate (DER) over all phonological features compared to the phoneme-level equivalent. Mostafa Shahin, Julien Epps, Beena Ahmed |
Speech Commun. | 3 |
| 2024 | When LLMs Meets Acoustic Landmarks: An Efficient Approach to Integrate Speech into Large Language Models for Depression DetectionabstractDepression is a critical concern in global mental health, prompting extensive research into AIbased detection methods.Among various AI technologies, Large Language Models (LLMs) stand out for their versatility in mental healthcare applications.However, their primary limitation arises from their exclusive dependence on textual input, which constrains their overall capabilities.Furthermore, the utilization of LLMs in identifying and analyzing depressive states is still relatively untapped.In this paper, we present an innovative approach to integrating acoustic speech information into the LLMs framework for multimodal depression detection.We investigate an efficient method for depression detection by integrating speech signals into LLMs utilizing Acoustic Landmarks.By incorporating acoustic landmarks, which are specific to the pronunciation of spoken words, our method adds critical dimensions to text transcripts.This integration also provides insights into the unique speech patterns of individuals, revealing the potential mental states of individuals.Evaluations of the proposed approach on the DAIC-WOZ dataset reveal state-of-the-art results when compared with existing Audio-Text baselines.In addition, this approach is not only valuable for the detection of depression but also represents a new perspective in enhancing the ability of LLMs to comprehend and process speech signals. Xiangyu Zhang 0005, Hexin Liu, Kaishuai Xu, Qiquan Zhang, Daijiao Liu, Beena Ahmed, Julien Epps |
EMNLP | 6 |
| 2024 | A Probability Gradient Based Approach for Sampling Boundaries of In-Domain DataabstractIn machine learning applications, it is desirable to distinguish between in-domain and out-of-domain data. However, in most cases, only in-domain data is available and consequently identifying the ‘boundary’ between in-domain and out-of-domain is a significant challenge. In this paper we present a novel technique that can identify points on this boundary based on only in-domain data. Specifically, the proposed method operates on the hypothesis that the gradient of the probability of the data being in-domain will be highest at the boundary. It utilises an iterative approach, alternating between Monte Carlo sampling of an estimated ‘boundary distribution’ and a binary classifier trained on these points to distinguish between in-domain and out-of-domain data to improve the estimate of the boundary. This method leads to both a set of high-quality data points from the boundary and a calibrated out-of-domain detector. The operation of the proposed approach is validated on MNIST, Fashion-MNIST, and Omniglot datasets. Miao Jing, Vidhyasaharan Sethu, Beena Ahmed |
ICASSP | 3 |
| 2024 | Variational Connectionist Temporal Classification for Order-Preserving Sequence ModelingabstractConnectionist temporal classification (CTC) is commonly adopted for sequence modeling tasks like speech recognition, where it is necessary to preserve order between the input and target sequences. However, CTC is only applied to deterministic sequence models, where the latent space is discontinuous and sparse, which in turn makes them less capable of handling data variability when compared to variational models. In this paper, we integrate CTC with a variational model and derive loss functions that can be used to train more generalizable sequence models that preserve order. Specifically, we derive two versions of the novel variational CTC based on two reasonable assumptions, the first being that the variational latent variables at each time step are conditionally independent; and the second being that these latent variables are Markovian. We show that both loss functions allow direct optimization of the variational lower bound for the model log-likelihood, and present computationally tractable forms for implementing them. Zheng Nan, Ting Dang, Vidhyasaharan Sethu, Beena Ahmed |
ICASSP | 4 |
| 2024 | Phonological-Level Mispronunciation Detection and Diagnosis
Mostafa Shahin, Beena Ahmed |
INTERSPEECH | 2 |
| 2023 | Improving wav2vec2-based Spoken Language Identification by Learning Phonological Features
Mostafa Shahin, Zheng Nan, Vidhyasaharan Sethu, Beena Ahmed |
INTERSPEECH | 4 |
| 2023 | Aligning Small Datasets Using Domain Adversarial Learning: Applications in Automated in Vivo Oral Cancer DiagnosisabstractDeep learning approaches for medical image analysis are limited by small data set size due to factors such as patient privacy and difficulties in obtaining expert labelling for each image. In medical imaging system development pipelines, phases for system development and classification algorithms often overlap with data collection, creating small disjoint data sets collected at numerous locations with differing protocols. In this setting, merging data from different data collection centers increases the amount of training data. However, a direct combination of datasets will likely fail due to domain shifts between imaging centers. In contrast to previous approaches that focus on a single data set, we add a domain adaptation module to a neural network and train using multiple data sets. Our approach encourages domain invariance between two multispectral autofluorescence imaging (maFLIM) data sets of in vivo oral lesions collected with an imaging system currently in development. The two data sets have differences in the sub-populations imaged and in the calibration procedures used during data collection. We mitigate these differences using a gradient reversal layer and domain classifier. Our final model trained with two data sets substantially increases performance, including a significant increase in specificity. We also achieve a significant increase in average performance over the best baseline model train with two domains (p = 0.0341). Our approach lays the foundation for faster development of computer-aided diagnostic systems and presents a feasible approach for creating a robust classifier that aligns images from multiple data centers in the presence of domain shifts. Kayla Caughlin, Elvis Duran-Sierra, Shuna Cheng, Rodrigo Cuenca-Martinez, Beena Ahmed, Jim Xiuquan Ji, Mathias Martinez, Moustafa Al-Khalil, Hussain Al-Enazi, Yi-Shing Lisa Cheng, Javier A. Jo, Carlos Busso |
IEEE J. Biomed. Health Informatics | 5 |
| 2022 | Knowledge of accent differences can be used to predict speech recognition
Tünde Szalay, Mostafa Shahin, Beena Ahmed, Kirrie J. Ballard |
INTERSPEECH | 3 |
| 2021 | AusKidTalk: An Auditory-Visual Corpus of 3- to 12-Year-Old Australian Children's SpeechabstractHere we present AusKidTalk [1], an audio-visual (AV) corpus of Australian children’s speech collected to facilitate the development of speech based technological solutions for children. It builds upon the technology and expertise developed through the collection of an earlier corpus of Australian adult speech, AusTalk [2,3]. This multi-site initiative was established to remedy the dire shortage of children’s speech corpora in Australia and around the world that are sufficiently sized to train accurate automated speech processing tools for children. We are collecting ~600 hours of speech from children aged 3–12 years that includes single word and sentence productions as well as narrative and emotional speech. In this paper, we discuss the key requirements for AusKidTalk and how we designed the recording setup and protocol to meet them. We also discuss key findings from our feasibility study of the recording protocol, recording tools, and user interface. Beena Ahmed, Kirrie J. Ballard, Denis Burnham, Tharmakulasingam Sirojan, Hadi Mehmood, Dominique Estival, Elise Baker, Felicity Cox, Joanne Arciuli, Titia Benders, Katherine Demuth, Barbara Kelly, Chloé Diskin-Holdaway, Mostafa Shahin, Vidhyasaharan Sethu, Julien Epps, Chwee Beng Lee, Eliathamby Ambikairajah |
Interspeech | 1 |
| 2021 | Assessing Posterior-Based Mispronunciation Detection on Field-Collected Recordings from Child Speech Therapy Sessions
Adam Hair, Guanlong Zhao, Beena Ahmed, Kirrie J. Ballard, Ricardo Gutierrez-Osuna |
Interspeech | 3 |
| 2020 | UNSW System Description for the Shared Task on Automatic Speech Recognition for Non-Native Children's Speech
Mostafa Shahin, Renée Lu, Julien Epps, Beena Ahmed |
INTERSPEECH | 4 |
| 2020 | Gaming Away Stress: Using Biofeedback Games to Learn Paced BreathingabstractBiofeedback games are an attractive alternative to standard techniques for learning short-term relaxation skills. In this paper, we present the design, implementation, and evaluation of three respiratory biofeedback games. To validate these games, we compared breathing rate across 100 male participants (23 years ± 3.2 years) playing biofeedback and audio pacing versions of these games as well as a paced breathing app. The games were placed between repeat runs of a cognitively stressful Stroop-based task and the impact of the games on breathing and cognitive performance in the task also assessed. Our results showed that 1) differences in gameplay did not impact player performance; 2) biofeedback not only led to better breath control during play but also during the subsequent cognitively stressful task; and 3) biofeedback led to better attentional-cognitive performance in the subsequent task. Our multi-game experiments show that using respiratory biofeedback in video games is an effective strategy to learn paced breathing—on par with the standalone technique of paced breathing—and to self-regulate stress levels in later stressful scenarios. Furthermore, owing to its entertainment value, our relaxation solution has the potential to be more engaging and accessible than standalone paced breathing, for use over longer durations. M. Abdullah Zafar, Beena Ahmed, Rami G. Al Rihawi, Ricardo Gutierrez-Osuna |
IEEE Trans. Affect. Comput. | 2 |
| 2019 | Evaluating Automatic Speech Recognition for Child Speech Therapy ApplicationsabstractAutomatic speech recognition (ASR) technology can be a useful tool in mobile apps for child speech therapy, empowering children to complete their practice with limited caregiver supervision. However, little is known about the feasibility of performing ASR on mobile devices, particularly when training data is limited. In this study, we investigated the performance of two low-resource ASR systems on disordered speech from children. We compared the open-source PocketSphinx (PS) recognizer using adapted acoustic models and a custom template-matching (TM) recognizer. TM and the adapted models significantly out-perform the default PS model. On average, maximum likelihood linear regression and maximum a posteriori adaptation increased PS accuracy from 59.4% to 63.8% and 80.0%, respectively, suggesting that the models successfully captured speaker-specific word production variations. TM reached a mean accuracy of 75.8% Adam Hair, Kirrie J. Ballard, Beena Ahmed, Ricardo Gutierrez-Osuna |
ASSETS | 3 |
| 2019 | Anomaly detection based pronunciation verification approach using speech attribute features
Mostafa Shahin, Beena Ahmed |
Speech Commun. | 2 |
| 2018 | Apraxia world: a speech therapy game for children with speech sound disordersabstractThis paper presents Apraxia World, a remote therapy tool for speech sound disorders that integrates speech exercises into an engaging platformer-style game. In Apraxia World, the player controls the avatar with virtual buttons/joystick, whereas speech input is associated with assets needed to advance from one level to the next. We tested performance and child preference of two strategies for delivering speech exercises: during each level, and after it. Most children indicated that doing exercises after completing each level was less disruptive and preferable to doing exercises scattered through the level. We also found that children liked having perceived control over the game (character appearance, exercise behavior). Our results indicate that (i) a familiar style of game successfully engages children, (ii) speech exercises function well when decoupled from game control, and (iii) children are willing to complete required speech exercises while playing a game they enjoy. Adam Hair, Penelope Monroe, Beena Ahmed, Kirrie J. Ballard, Ricardo Gutierrez-Osuna |
IDC | 3 |
| 2018 | One-Class SVMs Based Pronunciation Verification ApproachabstractThe automatic assessment of speech plays an important role in Computer Aided Pronunciation Learning systems. However, modeling both the correct and incorrect pronunciation of each phoneme to achieve accurate pronunciation verification is unfeasible due to the lack of sufficient mispronounced samples in training datasets. In this paper, we propose a novel approach that handles this unbalanced data distribution by building multiple one-class SVMs to evaluate each phoneme as correct or incorrect. We model the correct pronunciation of each individual phoneme with a one-class SVM trained using a set of speech attributes features, namely the manner and place of articulation. These features are extracted from a bank of pre-trained DNN speech attributes classifiers. The one-class SVM model measures the similarity between the new data and the training set and then classifies it as normal (correct) or an anomaly (incorrect). We evaluated the system using native speech corpus and disordered speech corpus and compared it with the conventional Goodness of Pronunciation (GOP) algorithm. The results show that our approach reduces the false-acceptance and false-rejection rates by around 26% and 39% respectively. Mostafa Shahin, Jim Xiuquan Ji, Beena Ahmed |
ICPR | 3 |
| 2018 | Anomaly Detection Approach for Pronunciation Verification of Disordered Speech Using Speech Attribute Features
Mostafa Shahin, Beena Ahmed, Jim Xiuquan Ji, Kirrie J. Ballard |
INTERSPEECH | 2 |
| 2017 | An evaluation of the flipped classroom format in a first year introductory engineering courseabstractIn earlier work, we presented the results of an initial study on the effectiveness of the flipped classroom model in the programming portion of a first year engineering course. Here we present results of an additional study we conducted to determine whether our earlier results can be replicated if the flipped class is used for the whole course. We extended the use of the flipped class to not only cover programming but also include the course topics of engineering design, project management, statistics, dimensions and conversions, technical representation of data and engineering ethics. We compared student performance on traditional quizzes and lab assessments in the flipped classroom version to data collected from the previous term where only a portion of the course was flipped and an earlier traditional running of the course. Data collected on student usage of learning material, grades and surveys was consistent with previous results. Students were able to adjust to the unfamiliar flipped classroom model, with the majority preparing for the class beforehand using the provided learning materials. The flipped class also resulted in similar improvements in assessment results when compared to previous runnings of the course. Lastly, students here too perceived the flipped class to be beneficial to them from a learning perspective and assisted them in the course outcomes. Ghada Salama, Sine Scanlon, Beena Ahmed |
EDUCON | 3 |
| 2017 | Deep Learning and Insomnia: Assisting Clinicians With Their DiagnosisabstractEffective sleep analysis is hampered by the lack of automated tools catering to disordered sleep patterns and cumbersome monitoring hardware. In this paper, we apply deep learning on a set of 57 EEG features extracted from a maximum of two EEG channels to accurately differentiate between patients with insomnia or controls with no sleep complaints. We investigated two different approaches to achieve this. The first approach used EEG data from the whole sleep recording irrespective of the sleep stage (stage-independent classification), while the second used only EEG data from insomnia-impacted specific sleep stages (stage-dependent classification). We trained and tested our system using both healthy and disordered sleep collected from 41 controls and 42 primary insomnia patients. When compared with manual assessments, an NREM + REM based classifier had an overall discrimination accuracy of 92% and 86% between two groups using both two and one EEG channels, respectively. These results demonstrate that deep learning can be used to assist in the diagnosis of sleep disorders such as insomnia. Mostafa Shahin, Beena Ahmed, Sana Tmar Ben Hamida, Lamana Mulaffer, Martin Glos, Thomas Penzel |
IEEE J. Biomed. Health Informatics | 2 |
| 2016 | Flipping introductory engineering design courses: Evaluating their effectivenessabstractIntroductory engineering courses can benefit from the flipped classroom model as these courses develop hands-on skills such as engineering design and critical thinking skills. Providing students with instruction outside the classroom with e-learning tools frees up class time to engage in activities that promote deeper learning, e.g. classroom discussions, problem-solving sessions, design activities, etc. In this paper, we discuss the implementation of a flipped classroom version of an introductory engineering design course. We evaluate student performance and perception of a flipped classroom version using data collected during the course. We also discuss elements of the course that were found to be the most beneficial for the students. Beena Ahmed, Ali Aljaani, Mohamed Ismail Yousuf |
EDUCON | 1 |
| 2016 | Classification of bisyllabic lexical stress patterns in disordered speech using deep learningabstractTechnology-based therapy tools can be of great benefit to children with developmental speech disabilities as they typically require sustained practice with a speech therapist for several years. Towards this aim, over the past 4 years we have developed speech processing tools to automatically detect common errors in disordered speech. This paper presents an automated technique to identify incorrect lexical stress. Specifically, we describe a deep neural network (DNN) that can be used to classify the four different bisyllabic stress patterns: strong-weak (SW), weak-strong (WS), strong-strong (SS) and weak-weak (WW). We derive input features for the DNN from the duration, pitch, intensity and spectral energy on each of the two consecutive syllables. Using these features, we achieve 93% correct classification between SW/WS stress patterns and 88% correct classification of the four bisyllabic patterns on speech from typically developing children, while we obtain 73.4% classification between SW/WS in disordered speech. These figures represent a two-fold reduction in error rates compared to our prior work, which used a DNN with differential features from consecutive syllables. Mostafa Shahin, Ricardo Gutierrez-Osuna, Beena Ahmed |
ICASSP | 3 |
| 2016 | Automatic Classification of Lexical Stress in English and Arabic Languages Using Deep Learning
Mostafa Shahin, Julien Epps, Beena Ahmed |
INTERSPEECH | 3 |
| 2016 | ReBreathe: A Calibration Protocol that Improves Stress/Relax Classification by Relabeling Deep Breathing Relaxation ExercisesabstractTraining stress-prediction models is challenging due to the difficulty in reliably eliciting stress and relaxation responses in participants. For example, a task intended to elicit a relaxation response (e.g., deep breathing) can have the opposite effect depending on the participant's appraisal of and familiarity with the exercise. Including such instances in a training set undermines the accuracy of the resulting prediction model. This paper presents a technique, ReBreathe, to identify such instances based on respiratory patterns and determine their accurate stress/relax labels. We compared this relabeling approach against two labeling techniques: 1) nominal labels obtained from the experimental protocol and 2) labels obtained from subjective assessments. We then trained generalized estimating equation regression models to predict the resulting stress/relax labels from measures of heart rate variability and electrodermal activity. Training the model using protocol labels achieved a classification rate of 0.53 on participants not included in the training set. Relabeling the exercises based on each participant's subjective ratings increased classification rates but only marginally (0.61). In contrast, relabeling the exercises based on respiratory patterns increased classification rates to 0.88, or a four-fold reduction in error rates. These results illustrate the unreliability of protocol and subjective labels during stress/relax exercises and the potential benefits of ReBreathe. Beena Ahmed, Hira Khan, Jongyong Choi, Ricardo Gutierrez-Osuna |
IEEE Trans. Affect. Comput. | 1 |
| 2015 | Tabby Talks: An automated tool for the assessment of childhood apraxia of speech
Mostafa Shahin, Beena Ahmed, Avinash Parnandi 0001, Virendra Karappa, Jacqueline McKechnie, Kirrie J. Ballard, Ricardo Gutierrez-Osuna |
Speech Commun. | 2 |
| 2014 | A comparison of GMM-HMM and DNN-HMM based pronunciation verification techniques for use in the assessment of childhood apraxia of speechabstractThis paper introduces a pronunciation verification method to be used in an automatic assessment therapy tool of child disordered speech. The proposed method creates a phonebased search lattice that is flexible enough to cover all probable mispronunciations. This allows us to verify the correctness of the pronunciation and detect the incorrect phonemes produced by the child. We compare between two different acoustic models, the conventional GMM-HMM and the hybrid DNN-HMM. Results show that the hybrid DNNHMM outperforms the conventional GMM-HMM for all experiments on both normal and disordered speech. The total correctness accuracy of the system at the phoneme level is above 85% when used with disordered speech. Mostafa Shahin, Beena Ahmed, Jacqueline McKechnie, Kirrie J. Ballard, Ricardo Gutierrez-Osuna |
INTERSPEECH | 2 |
| 2014 | Classification of lexical stress patterns using deep neural network architectureabstractLexical stress is a key diagnostic marker of disordered speech as it strongly affects speech perception. In this paper we introduce an automated method to classify between the different lexical stress patterns in children's speech. A deep neural network is used to classify between strong-weak (SW), weak-strong (WS) and equal-stress (SS/WW) patterns in English by measuring the articulation change between the two successive syllables. The deep neural network architecture is trained using a set of acoustic features derived from pitch, duration and intensity measurements along with the energies in different frequency bands. We compared the performance of the deep neural classifier to a traditional single hidden layer MLP. Results show that the deep neural classifier outperforms the traditional MLP. The accuracy of the deep neural system is approximately 85% when classifying between the unequal stress patterns (SW/WS) and greater than 70% when classifying both equal and unequal stress patterns. Mostafa Shahin, Beena Ahmed, Kirrie J. Ballard |
SLT | 2 |
| 2013 | Architecture of an automated therapy tool for childhood apraxia of speechabstractWe present a multi-tier system for the remote administration of speech therapy to children with apraxia of speech. The system uses a client-server architecture model and facilitates task-oriented remote therapeutic training in both in-home and clinical settings. Namely, the system allows a speech therapist to remotely assign speech production exercises to each child through a web interface, and the child to practice these exercises on a mobile device. The mobile app records the child's utterances and streams them to a back-end server for automated scoring by a speech-analysis engine. The therapist can then review the individual recordings and the automated scores through a web interface, provide feedback to the child, and adapt the training program as needed. We validated the system through a pilot study with children diagnosed with apraxia of speech, and their parents and speech therapists. Here we describe the overall client-server architecture, middleware tools used to build the system, the speech-analysis tools for automatic scoring of recorded utterances, and results from the pilot study. Our results support the feasibility of the system as a complement to traditional face-to-face therapy through the use of mobile tools and automated speech analysis algorithms. Avinash Parnandi 0001, Virendra Karappa, Youngpyo Son, Mostafa Shahin, Jacqueline McKechnie, Kirrie J. Ballard, Beena Ahmed, Ricardo Gutierrez-Osuna |
ASSETS | 7 |
| 2013 | Towards efficient and secure in-home wearable insomnia monitoring and diagnosis systemabstractSleep disorders, such as insomnia can seriously affect a patient's quality of life. Sleep measurements based on polysomnographic (PSG) signals and patients' questionnaires are necessary for an accurate evaluation of insomnia. Due to recent innovations in technology, it is now possible to continuously monitor a patient's sleep at home and have their sleep data sent to a remote clinical back-end system for collection and assessment. Most of the research on sleep reported in the literature mainly looks into how to automate the analysis of the sleep data and does not address the problem of the efficient and secure transmissions of the collected health data. This paper provides an experimental evaluation of communication and security protocols that can be used in inhome sleep monitoring and health care and highlights the most suitable protocol in terms of security and overhead. Design guidelines are then derived for the deployment of effective inhome patients monitoring systems. Sana Tmar Ben Hamida, Elyes Ben Hamida, Beena Ahmed, Adnan A. Abu-Dayya |
BIBE | 3 |
| 2012 | The longer term impact of fundamental engineering skills on students in higher level undergraduate coursesabstractIn this paper, we will look at how the engineering skills acquired by the students in the freshmen Foundations of Engineering I (ENGR111) course helped them in their higher level courses in the university. As ENGR111 developed skills in students that they had previously not been exposed to, it was essential to understand their effect. Students in the sophomore, junior and senior were surveyed to collect this information. Results show that the students have found the impact of these skills in their higher level courses to be positive. Ghada Salama, Beena Ahmed |
EDUCON | 2 |
| 2012 | Automatic classification of unequal lexical stress patterns using machine learning algorithmsabstractTechnology based speech therapy systems are severely handicapped due to the absence of accurate prosodic event identification algorithms. This paper introduces an automatic method for the classification of strong-weak (SW) and weak-strong (WS) stress patterns in children speech with American English accent, for use in the assessment of the speech dysprosody. We investigate the ability of two sets of features used to train classifiers to identify the variation in lexical stress between two consecutive syllables. The first set consists of traditional features derived from measurements of pitch, intensity and duration, whereas the second set consists of energies of different filter banks. Three different classifiers were used in the experiments: an Artificial Neural Network (ANN) classifier with a single hidden layer, Support Vector Machine (SVM) classifier with both linear and Gaussian kernels and the Maximum Entropy modeling (MaxEnt). these features. Best results were obtained using an ANN classifier and a combination of the two sets of features. The system correctly classified 94% of the SW stress patterns and 76% of the WS stress patterns. Mostafa Shahin, Beena Ahmed, Kirrie J. Ballard |
SLT | 2 |
| 2004 | A voice activity detector using the chi-square testabstractThis paper proposes a voice activity detector (VAD) that makes the speech/noise classification by applying the statistical chi-square test to each frame. It also uses a continuous update of the background noise estimate. The speech is first enhanced using a noise reduction system, with noise estimates also obtained with the help of the chi-square test. The noise-reduced signal is decomposed into sub-bands, and the chi-square test is used again in another form to compare the observed signal distribution to the estimated noise distribution. If the chi-square test determines that they are close, the frame is declared to be noise, otherwise speech. The performance of this VAD was found to be significantly superior to several benchmark VAD, with accuracies above 89% even at a SNR of 0 dB, which is up to 25% better than the others. Beena Ahmed, W. Harvey Holmes |
ICASSP (1) | 1 |