EDBT 2026 Demo / reviewers in the wild / expert
Heidi Christensen
dblp:02/2429
· DBLP profile ↗
68ranked-venue papers
15as first author
23since 2021 · last 2026
0000-0003-3028-5062ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 53 · 13 first-author · 19 since 2021Artificial intelligence and machine learning · 47 · 10 first-author · 15 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Raw acoustic-articulatory multimodal dysarthric speech recognitionabstractAutomatic speech recognition (ASR) for dysarthric speech is challenging. The acoustic characteristics of dysarthric speech are highly variable and there are often fewer distinguishing cues between phonetic tokens. Multimodal ASR utilises the data from other modalities to facilitate the task when a single acoustic modality proves insufficient. Articulatory information, which encapsulates knowledge about the speech production process, may constitute such a complementary modality. Although multimodal acoustic-articulatory ASR has received increasing attention recently, incorporating real articulatory data is under-explored for dysarthric speech recognition. This paper investigates the effectiveness of multimodal acoustic modelling using real dysarthric speech articulatory information in combination with acoustic features, especially raw signal representations which are more informative than classic features, leading to learning representations tailored to dysarthric ASR. In particular, various raw acoustic-articulatory multimodal dysarthric speech recognition systems are developed and compared with similar systems with hand-crafted features. Furthermore, the difference between dysarthric and typical speech in terms of articulatory information is systematically analysed by using a statistical space distribution indicator called Maximum Articulator Motion Range (MAMR). Additionally, we used mutual information analysis to investigate the robustness and phonetic information content of the articulatory features, offering insights that support feature selection and the ASR results. Experimental results on the widely used TORGO dysarthric speech dataset show that combining the articulatory and raw acoustic features at the empirically found optimal fusion level achieves a notable performance gain, leading to up to 7.6% and 12.8% relative word error rate (WER) reduction for dysarthric and typical speech, respectively. Zhengjun Yue, Erfan Loweimi, Zoran Cvetkovic, Jon Barker, Heidi Christensen |
Comput. Speech Lang. | 5 |
| 2026 | Towards automating the Frenchay dysarthria assessment: Can neural phoneme posteriorgrams inform the analysis of dysarthric speech?abstractDysarthria is a type of motor speech disorder that reflects abnormalities in motor movements required for speech production. In clinical practice, identifying characteristic signs and symptoms of the neuropathophysiology underlying a dysarthria is vital for diagnosis and management. The gold standard for dysarthria assessment is auditory-perceptual evaluation by a speech and language therapist for differential diagnosis and management decisions. As the process is time-consuming for clinicians, there is growing interest in automatic dysarthria assessment (ADA). Recent approaches to ADA primarily focus on the classification of broad intelligibility or speech severity labels. However, this does not have much clinical utility and the assessment of communication-relevant parameters do not distinguish between dysarthria types and pathomechanisms. Studies on the classification of dysarthria function or clinical test protocol scores focusing on aspects of dysarthric speech production (such as the Frenchay dysarthria assessment (FDA)) are limited. Therefore, this paper focuses on the preliminary steps towards clinically interpretable ADA, including automatic FDA assessment. The phoneme posteriorgram (PPG) is a time-varying categorical distribution over acoustic speech units, and recent work demonstrates interpretable speech pronunciation distance for downstream tasks, e.g. pronunciation reconstruction. This work extends recent advances in posterior-based phoneme research and mispronunciation models to dysarthria assessment, exploring the extent to which dysarthric speech features in the FDA (identified by auditory-perceptual evaluation in clinical practice) are captured by PPG information. To achieve this, FDA aspects are systematically evaluated. The results show that interpretable PPG probability can capture dysarthric speech features that are related to motor system dysfunction. Wing-Zin Leung, Heidi Christensen, Stefan Goetze |
Speech Commun. | 2 |
| 2025 | Early Dementia Detection Using Multiple Spontaneous Speech Prompts: The PROCESS ChallengeabstractDementia is associated with various cognitive impairments and typically manifests only after significant progression, making intervention at this stage often ineffective. To address this issue, the Prediction and Recognition of Cognitive Decline through Spontaneous Speech (PROCESS) Signal Processing Grand Challenge invites participants to focus on early-stage dementia detection. We provide a new spontaneous speech corpus for this challenge. This corpus includes answers from three prompts designed by neurologists to better capture the cognition of speakers. Our baseline models achieved an F1-score of 55.0% on the classification task and an RMSE of 2.98 on the regression task. Fuxiang Tao, Bahman Mirheidari, Madhurananda Pahar, Sophie Young, Hend Elghazaly, Fritz Peters, Caitlin H. Illingworth, Dorota Braun, Ronan O'Malley, Simon Bell, Daniel Blackburn, Fasih Haider, Saturnino Luz, Heidi Christensen |
ICASSP | 15 |
| 2025 | Automatic Detection and Sub-typing of Primary Progressive Aphasia from Speech: Integrating Task-Specific Features and Spatio-Semantic GraphsabstractPrimary progressive aphasia (PPA) describes a group of neurodegenerative diseases that predominantly affect language abilities. Its diagnostic process typically requires experienced clinicians, often available only in specialised hospital departments. Patients with PPA frequently display changes in speech and language early in the disease progression. In this study, we extracted acoustic, linguistic, and task-specific features from audio recordings and evaluated their utility for PPA classification. Using a subset of task-specific features, we detected PPA with 97% accuracy. For sub-typing, models trained on the full feature set achieved 74% accuracy in a three-way classification of PPA variants. Our results highlight the added value of task-specific features, which complement traditional approaches. Additionally, their visualisation offers an intuitive representation of task execution, improving clinical interpretability and potential diagnostic utility. Fritz Peters, W. Richard Bevan-Jones, Grace Threlfall, Jenny M. Harris, Julie S. Snowden, Jennifer C. Thompson, Daniel Blackburn, Heidi Christensen |
INTERSPEECH | 9 |
| 2025 | Alzheimer's Dementia Detection Using Perplexity from Paired Large Language Models
Heidi Christensen, Stefan Goetze |
INTERSPEECH | 2 |
| 2025 | Can Speech Accurately Detect Depression in Patients With Comorbid Dementia? An Approach for Mitigating Confounding Effects of Depression and Dementia
Sophie Young, Fuxiang Tao, Bahman Mirheidari, Madhurananda Pahar, Markus Reuber, Heidi Christensen |
INTERSPEECH | 6 |
| 2025 | Automatic Detection of Early Cognitive Decline Using Multimodal Feature Fusion and Transfer Learning on Real-World Conversational SpeechabstractEarly signs of cognitive decline, such as dementia and mild cognitive impairment (MCI), often manifest in conversational speech. Early and accurate identification is essential for potential interventions prior to the onset of more severe stages of neurodegenerative diseases. We present CognoMemory, a system for detecting cognitive decline based on a person's speech, to collect 307 hrs of real-world conversational speech, corresponding to 1.92 million Whisper-transcribed words, from 1,639 participants. Speech recordings were collected as participants answered 14 memory-probing, clinically effective questions asked by a virtual agent, starting with a motivation prompt, followed by memory, cognitive functioning, fluency, picture description and reading task. Both acoustic and linguistic features, along with large language model (LLM) embeddings, were extracted from all 1,639 participants. A subset of 614 participants, either with an unconfirmed diagnosis or younger than 50 years, was used for pre-training. The remaining three groups (64 dementia, 169 MCI and 792 healthy participants) were used to fine-tune our proposed model. Our multimodal feature fusion and CNN/Bi-LSTM-based transfer learning approach outperforms LLM-based (BART, DistilBERT, RoBERTa and HuBERT) approaches while achieving the highest $F_{1}$-scores of 0.83 & 0.54 using just the initial 'motivation' question for 2-way & 3-way classification; exhibiting a 3% performance increase due to the application of transfer learning, while being also 38% faster. Finally, the classifiers trained on the CognoMemory data, the largest of its kind, were tested on the second-largest available DementiaBank dataset (Pitt corpus), and a CNN-based transfer learning architecture achieved an $F_{1}$-score of 0.89, demonstrating better stability and generalisation across datasets and of our novel feature fusion and architecture. Madhurananda Pahar, Bahman Mirheidari, Caitlin H. Illingworth, Dorota Braun, Fuxiang Tao, Lise Sproson, Daniel Blackburn, Heidi Christensen |
IEEE J. Biomed. Health Informatics | 8 |
| 2023 | Identifying People with Mild Cognitive Impairment at Risk of Developing Dementia using Speech AnalysisabstractMild Cognitive Impairment (MCI) is the intermediate stage between ageing and possible dementia. $50 \%$ of people with MCI may progress to dementia (prodromal Alzheimer’s Disease (pr-AD)). Identifying those at risk is a challenging but important task that can help with anxiety and treatment. Currently, clinicians wait for severe signs of impairment to emerge, however, recent studies have shown promising automatic speech-based approaches. This paper works on a unique dataset containing 50 healthy controls (HCs) and 50 MCI (of which 22 are pr-AD and 28 are non-progressed (nonP)). The recordings are of people speaking with a virtual agent. A number of acoustic, text and contextual features were extracted and used for training classifiers. The best result was achieved with an $F_{1}$-scores of $\mathbf{81.2 \%}$ for the detection of MCI versus HC, and $\mathbf{75} \%$ for MCI(pr-AD) versus MCI(nonP). The best $F_{1}$-score of $\mathbf{66.9 \%}$ was achieved for the detection of MCI (pr-AD) versus MCI (nonP) versus HC. Bahman Mirheidari, Ronan O'Malley, Daniel Blackburn, Heidi Christensen |
ASRU | 4 |
| 2023 | Investigating Visual Features for Cognitive Impairment Detection Using In-the-wild DataabstractEarly detection of dementia has attracted much research interest due to its crucial role in helping people get suitable treatment or care. Video analysis may provide an effective approach for detection, with low cost and effort compared to current expensive and intensive clinical assessments. This paper investigates the use of a range of visual features - eye blink rate (EBR), head turn rate (HTR) and head movement statistical features (HMSF) - for identifying neurodegenerative disorder (ND), mild cognitive impairment (MCI) and functional memory disorder (FMD). These features are used in a noval multiple thresholds approach, which is applied to an in-the-wild video dataset which includes data recorded in a range of challenging environments. A combination of EBR and HTR gives 78 % accuracy in a three-way classification task (ND/MCI/FMD) and 83%, 83% and 92%, respectively, for the two-way classifications ND/MCI, ND/FMD and MCI/FMD. These results are comparable to related work that uses more features from different modalities. They also provide evidence to support the possibility of an in-the-home detection process for dementia or cognitive impairment. Fatimah Alzahrani, Bahman Mirheidari, Daniel Blackburn, Steve C. Maddock, Heidi Christensen |
FG | 5 |
| 2023 | Moving Towards Non-Binary Gender Identification Via Analysis of System Errors in Binary Gender ClassificationabstractThis paper aims to analyse human perceptions of gender in speech signals, focusing on signals that are misclassified by methods for binary gender classification, looking at the features of speech signals that are more likely to be misclassified, or classified as either nonbinary or unclassifiable. The paper also analyses how human subjects perform in classifying such speech signals to gain insight into differences between machine and human performance levels. It is shown that gender classification systems and human ratings lack inter-annotator agreement, as do human ratings considered individually. There is also discussion of the suitability of continuing to use a binary system for gender in the field. This work fits into a larger body of research ongoing in the area of speech technology for trans-gender voice therapy. Sebastian Ellis, Stefan Goetze, Heidi Christensen |
ICASSP | 3 |
| 2022 | Multi-Modal Acoustic-Articulatory Feature Fusion For Dysarthric Speech RecognitionabstractBuilding automatic speech recognition (ASR) systems for speakers with dysarthria is a very challenging task. Although multi-modal ASR has received increasing attention recently, incorporating real articulatory data with acoustic features has not been widely explored in the dysarthric speech community. This paper investigates the effectiveness of multi-modal acoustic modelling for dysarthric speech recognition using acoustic features along with articulatory information. The proposed multi-stream architectures consist of convolutional, recurrent and fully-connected layers allowing for bespoke per-stream pre-processing, fusion at the optimal level of abstraction and post-processing. We study the optimal fusion level/scheme as well as training dynamics in terms of cross-entropy and WER using the popular TORGO dysarthric speech database. Experimental results show that fusing the acoustic and articulatory features at the empirically found optimal level of abstraction achieves a remarkable performance gain, leading to up to 4.6% absolute (9.6% relative) WER reduction for speakers with dysarthria. Zhengjun Yue, Erfan Loweimi, Zoran Cvetkovic, Heidi Christensen, Jon Barker |
ICASSP | 4 |
| 2022 | Evaluating the Performance of State-of-the-Art ASR Systems on Non-Native English using Corpora with Extensive Language Background VariationabstractThis investigation is an exploration into the performance of several different ASR systems in dealing with non-native English using corpora with extensive language background variation. This study takes two corpora amounting to 191 different native language (L1) backgrounds and looks at how these systems are able to process non-native English (L2) speech. A transformer based ASR system and a CRDNN architecture are both tested, trained on Librispeech [1] and Commonvoice [2] for a three way cross comparison. In addition Google's Speech-to-Text API and AWS Transcribe were investigated in order to evaluate popular mainstream approaches given their current degree of impact in deployed systems. Experiments reveal deficits in the range of 10%-15% mean WER performance difference between L1 and L2 speech. Results indicate ASR systems trained on particular varieties of L2 speech may be effective in improving WERs with outcomes in this paper demonstrating several Google ASR models trained on varieties of African L2 English outperforming L1 trained ASR for under-represented dialect groups in the United Kingdom. Further research is proposed to explore the plausibility of this approach and to critically approach WER as a metric for ASR evaluation, striving instead towards metrics with greater emphasis on evaluating language for communication. Samuel Schmück, Daniel Blackburn, Heidi Christensen |
INTERSPEECH | 3 |
| 2022 | Automatic cognitive assessment: Combining sparse datasets with disparate cognitive scores
Bahman Mirheidari, Daniel Blackburn, Heidi Christensen |
INTERSPEECH | 3 |
| 2022 | Automatic Detection of Expressed Emotion from Five-Minute Speech Samples: Challenges and OpportunitiesabstractWe present a novel feasibility study on the automatic recognition of Expressed Emotion (EE), a family environment concept based on caregivers speaking freely about their relative/family member.We describe an automated approach for determining the degree of warmth, a key component of EE, from acoustic and text features acquired from a sample of 37 recorded interviews.These recordings, collected over 20 years ago, are derived from a nationally representative birth cohort of 2,232 British twin children and were manually coded for EE.We outline the core steps of extracting usable information from recordings with highly variable audio quality and assess the efficacy of four machine learning approaches trained with different combinations of acoustic and text features.Despite the challenges of working with this legacy data, we demonstrated that the degree of warmth can be predicted with an F1-score of 61.5%.In this paper, we summarise our learning and provide recommendations for future work using real-world speech samples. Bahman Mirheidari, André Bittar, Nicholas Cummins, Johnny Downs, Helen L. Fisher, Heidi Christensen |
INTERSPEECH | 6 |
| 2022 | Dysarthric Speech Recognition From Raw Waveform with Parametric CNNsabstractRaw waveform acoustic modelling has recently received increasing attention. Compared with the task-blind hand-crafted features which may discard useful information, representations directly learned from the raw waveform are task-specific and potentially include all task-relevant information. In the context of automatic dysarthric speech recognition (ADSR), raw waveform acoustic modelling is under-explored owing to data scarcity. Parametric convolutional neural networks (CNNs) can compensate for this problem due to having notably fewer parameters and requiring less training data in comparison with conventional non-parametric CNNs. In this paper, we explore the usefulness of raw waveform acoustic modelling using various parametric CNNs for ADSR. We investigate the properties of the learned filters and monitor the training dynamics of various models. Furthermore, we study the effectiveness of data augmentation and multi-stream acoustic modelling through combining the non-parametric and parametric CNNs fed by hand-crafted and raw waveform features. Experimental results on the TORGO dysarthric database show that the parametric CNNs significantly outperform the non-parametric CNNs, reaching up to 36.2% and 12.6% WERs (up to 3.4% and 1.1% absolute error reduction) for dysarthric and typical speech, respectively. Multi-stream acoustic modelling further improves the performance resulting in up to 33.2% and 10.3% WERs for dysarthric and typical speech, respectively. Zhengjun Yue, Erfan Loweimi, Heidi Christensen, Jon Barker, Zoran Cvetkovic |
INTERSPEECH | 3 |
| 2022 | Acoustic Modelling From Raw Source and Filter Components for Dysarthric Speech RecognitionabstractAcoustic modelling for automatic dysarthric speech recognition (ADSR) is a challenging task. Data deficiency is a major problem and substantial differences between typical and dysarthric speech complicate the transfer learning. In this paper, we aim at building acoustic models using the raw magnitude spectra of the source and filter components for ADSR. The proposed multi-stream models consist of convolutional, recurrent and fully-connected layers allowing for pre-processing various information streams and fusing them at an optimal level of abstraction. We demonstrate that such a multi-stream processing leverages information encoded in the vocal tract and excitation components and leads to normalising nuisance factors such as speaker attributes and speaking style. This leads to a better handling of dysarthric speech that exhibits large inter- and intra-speaker variabilities and results in a notable performance gain. Furthermore, we analyse the learned convolutional filters and visualise the outputs of different layers after dimensionality reduction to demonstrate how the speaker-related attributes are normalised along the pipeline. We also compare the proposed multi-stream model with various systems based on MFCC, FBank, raw waveform and i-vector, and, study the training dynamics as well as usefulness of the feature normalisation and data augmentation via speed perturbation. On the widely used TORGO and UASpeech dysarthric speech corpora, the proposed approach leads to a competitive performance of up to 35.3% and 30.3% WERs for dysarthric speech, respectively. Zhengjun Yue, Erfan Loweimi, Heidi Christensen, Jon Barker, Zoran Cvetkovic |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2021 | Eye Blink Rate Based Detection of Cognitive Impairment Using In-the-wild DataabstractInvestigating automatic methods for the early detection of dementia and related conditions that cause cognitive impairment is an area of growing interest. Video processing could play a role by providing a non-invasive and low-cost alternative to current expensive assessments. For this to be successful it is crucial that approaches are robust to in-the-wild challenges. In this paper, visual cues, related to the eye blink rate (EBR), are investigated to quantify the early phase of neurodegenerative disorder (ND) and mild cognitive impairment (MCI) as well as functional memory disorder (FMD; problems with memory not related to neurodegenerative disorder). This paper aims to improve the detection of ND and MCI by investigating a novel approach to calculating the EBR that is more robust to in-the-wild challenges. An in-house dataset with 18 participants is used. The EBR is calculated from eye landmarks extracted using two libraries (Dlib and Openface). To mitigate issues observed in the noisy, in-the-wild recordings, a multiple threshold approach for EBR detection is proposed. It involves generating multiple thresholds for identifying a blink, where a threshold is used to determine whether an eye is open or closed. Several supervised machine learning approaches are used for automatic classification. The results show that accuracy measures of 89% and 78% are achieved using Dlib and OpenFace data, respectively, when distinguishing between three conditions with ND, MCI and FMD. Fatimah Alzahrani, Bahman Mirheidari, Daniel Blackburn, Steve C. Maddock, Heidi Christensen |
ACII | 5 |
| 2021 | Multi-Task Estimation of Age and Cognitive Decline from SpeechabstractSpeech is a common physiological signal that can be affected by both ageing and cognitive decline. Often the effect can be confounding, as would be the case for people at, e.g., very early stages of cognitive decline due to dementia. Despite this, the automatic predictions of age and cognitive decline based on cues found in the speech signal are generally treated as two separate tasks. In this paper, multi-task learning is applied for the joint estimation of age and the Mini-Mental Status Evaluation criteria (MMSE) commonly used to assess cognitive decline. To explore the relationship between age and MMSE, two neural network architectures are evaluated: a SincNet-based end-to-end architecture, and a system comprising of a feature extractor followed by a shallow neural network. Both are trained with single-task or multi-task targets. To compare, an SVM-based regressor is trained in a single-task setup. i-vector, x-vector and ComParE features are explored. Results are obtained on systems trained on the DementiaBank dataset and tested on an in-house dataset as well as the ADReSS dataset. The results show that both the age and MMSE estimation is improved by applying multitask learning, with state-of-the-art results achieved on the ADReSS dataset acoustic-only task. Yilin Pan, Venkata Srikanth Nallanthighal, Daniel Blackburn, Heidi Christensen, Aki Härmä |
ICASSP | 4 |
| 2021 | Towards Automatic Speech Recognition for People with Atypical Speech
Heidi Christensen |
Interspeech | 1 |
| 2021 | Identifying Cognitive Impairment Using Sentence Representation VectorsabstractThe widely used word vectors can be extended at the sentence level to perform a wide range of natural language processing (NLP) tasks.Recently the Bidirectional Encoder Representations from Transformers (BERT) language representation achieved state-of-the-art performance for these applications.The model is trained with punctuated and well-formed (writ-ten) text, however, the performance of the model drops significantly when the input text is the -erroneous and unpunctuated-output of automatic speech recognition (ASR).We use a sliding window and averaging approach for pre-processing text for BERT to extract features for classifying three diagnostic categories relating to cognitive impairment: neurodegenerative dis-order (ND), mild cognitive impairment (MCI), and healthy controls (HC).The in-house dataset contains the audio recordings of an intelligent virtual agent (IVA) who asks the participants several conversational questions prompts in addition to giving a picture description prompt.For the three-way classification, we achieve a 73.88% F-score (accuracy: 76.53%) using the pre-trained, uncased base BERT and for the two-way classifier (HCvs.ND) we achieve 89.80% (accuracy: 90%).We further improve these by using a prompt selection technique, reaching the F-scores of 79.98% (accuracy: 81.63%) and 93.56% (accuracy:93.75%)respectively. Bahman Mirheidari, Yilin Pan, Daniel Blackburn, Ronan O'Malley, Heidi Christensen |
Interspeech | 5 |
| 2021 | Using the Outputs of Different Automatic Speech Recognition Paradigms for Acoustic- and BERT-Based Alzheimer's Dementia Detection Through Spontaneous Speech
Yilin Pan, Bahman Mirheidari, Jennifer M. Harris, Jennifer C. Thompson, Julie S. Snowden, Daniel Blackburn, Heidi Christensen |
Interspeech | 8 |
| 2021 | Parental Spoken Scaffolding and Narrative Skills in Crowd-Sourced Storytelling Samples of Young ChildrenabstractA novel crowdsourcing project to gather children's storytelling based language samples using a mobile app was undertaken across the United Kingdom. Parents' scaffolding of children's narratives was observed in many of the samples. This study was designed to examine the relationship of scaffolding and young children's narrative language ability in a story retell context which is analysed at the macro-structural (total macro-structure score), the micro-structural (mean length of utterances in morphemes) and verbal productivity (total number of utterances) levels. Young children with and without scaffolding were statistically compared. The interaction between the level of scaffolding support, the grammar complexity and the narrative structure was explored. A bidirectional relationship was observed between scaffolding and young children's narrative language ability. Young children with better performance were observed to receive less scaffolding from parents. Scaffolding was shown to support early narrative development of young children and was more able to benefit those with low-level grammatical complexity skills. It is crucial to encourage parental scaffolding to be well-attuned to the child's narrative ability. Zhengjun Yue, Jon Barker, Heidi Christensen, Cristina McKean, Elaine Ashton, Yvonne Wren, Swapnil Gadgil, Rebecca Bright |
Interspeech | 3 |
| 2021 | Acoustic differences in emotional speech of people with dysarthriaabstractCommunicating emotion is essential in building and maintaining relationships. We communicate our emotional state not just with the words we use, but also how we say them. Changes in the rate of speech, short-term energy and intonation all help to convey emotional states like 'angry', 'sad' and 'happy'. People with dysarthria, the most common speech disorder, have reduced articulatory and phonatory control. This can affect the intelligibility of their speech, especially when communicating with unfamiliar conversation partners. However, we know little about how people with dysarthria convey their emotional state, and whether they are having to make changes to their speech to achieve this. In this study, we investigated the ability of people with dysarthria, caused by cerebral palsy and Parkinson's disease, to communicate emotions in their speech, and we compared their speech to that of speakers with typical speech. A parallel database of emotional speech was collected. One female speaker with dysarthria due to cerebral palsy, 3 speakers with dysarthria due to Parkinson's disease (2 female and 1 male), and 21 typical speakers (9 female and 12 male) produced sentences with 'angry', 'happy', 'sad', and 'neutral' emotions. A number of acoustic features were analysed using linear multi-level modeling. The results show that people with dysarthria were able to control some aspects of the suprasegmental and prosodic features when attempting to communicate emotions. For most speakers the changes they made are consistent with the changes made by speakers with typical speech. Even when the changes might be different to that of typical speakers, acoustic analysis shows these were consistent for different emotions. The analysis shows that variation in energy and jitter (local absolute) are major indicators of emotion in the study. Lubna Alhinti, Heidi Christensen, Stuart P. Cunningham |
Speech Commun. | 2 |
| 2020 | Source Domain Data Selection for Improved Transfer Learning Targeting Dysarthric Speech RecognitionabstractThis paper presents an improved transfer learning framework applied to robust personalised speech recognition models for speakers with dysarthria. As the baseline of transfer learning, a state-of-the-art CNN-TDNN-F ASR acoustic model trained solely on source domain data is adapted onto the target domain via neural network weight adaptation with the limited available data from target dysarthric speakers. Results show that linear weights in neural layers play the most important role for an improved modelling of dysarthric speech evaluated using UASpeech corpus, achieving averaged 11.6% and 7.6% relative recognition improvement in comparison to the conventional speaker-dependent training and data combination, respectively. To further improve the transferability towards target domain, we propose an utterance-based data selection of the source domain data based on the entropy of posterior probability, which is analysed to statistically obey a Gaussian distribution. Compared to a speaker-based data selection via dysarthria similarity measure, this allows for a more accurate selection of the potentially beneficial source domain data for either increasing the target domain training pool or constructing an intermediate domain for incremental transfer learning, resulting in a further absolute recognition performance improvement of nearly 2% added to transfer learning baseline for speakers with moderate to severe dysarthria. Feifei Xiong, Jon Barker, Zhengjun Yue, Heidi Christensen |
ICASSP | 4 |
| 2020 | Exploring Appropriate Acoustic and Language Modelling Choices for Continuous Dysarthric Speech RecognitionabstractThere has been much recent interest in building continuous speech recognition systems for people with severe speech impairments, e.g., dysarthria. However, the datasets that are commonly used are typically designed for tasks other than ASR development, or they contain only isolated words. As such, they contain much overlap in the prompts read by the speakers. Previous ASR evaluations have often neglected this, using language models (LMs) trained on non-disjoint training and test data, potentially producing unrealistically optimistic results. In this paper, we investigate the impact of LM design using the widely used TORGO database. We combine state-of-the-art acoustic models with LMs trained with data originating from LibriSpeech. Using LMs with varying vocabulary size, we examine the trade-off between the out-of-vocabulary rate and recognition confusions for speakers with varying degrees of dysarthria. It is found that the optimal LM complexity is highly speaker dependent, highlighting the need to design speaker-dependent LMs alongside speaker-dependent acoustic models when considering atypical speech. Zhengjun Yue, Feifei Xiong, Heidi Christensen, Jon Barker |
ICASSP | 3 |
| 2020 | Recognising Emotions in Dysarthric Speech Using Typical Speech DataabstractEffective communication relies on the comprehension of both verbal and nonverbal information. People with dysarthria may lose their ability to produce intelligible and audible speech sounds which in time may affect their way of conveying emotions, that are mostly expressed using nonverbal signals. Recent research shows some promise on automatically recognising the verbal part of dysarthric speech. However, this is the first study that investigates the ability to automatically recognise the nonverbal part. A parallel database of dysarthric and typical emotional speech is collected, and approaches to discriminating between emotions using models trained on either dysarthric (speaker dependent, matched) or typical (speaker independent, unmatched) speech are investigated for four speakers with dysarthria caused by cerebral palsy and Parkinson’s disease. Promising results are achieved in both scenarios using SVM classifiers, opening new doors to improved, more expressive voice input communication aids. Lubna Alhinti, Stuart P. Cunningham, Heidi Christensen |
INTERSPEECH | 3 |
| 2020 | A Comparison of Acoustic and Linguistics Methodologies for Alzheimer's Dementia RecognitionabstractContains fulltext : 228158.pdf (Publisher’s version ) (Open Access) Nicholas Cummins, Yilin Pan, Zhao Ren, Julian Fritsch, Venkata Srikanth Nallanthighal, Heidi Christensen, Daniel Blackburn, Björn W. Schuller, Mathew Magimai-Doss, Helmer Strik, Aki Härmä |
INTERSPEECH | 6 |
| 2020 | Improving Cognitive Impairment Classification by Generative Neural Network-Based Feature Augmentation
Bahman Mirheidari, Daniel Blackburn, Ronan O'Malley, Annalena Venneri, Traci Walker, Markus Reuber, Heidi Christensen |
INTERSPEECH | 7 |
| 2020 | Improving Detection of Alzheimer's Disease Using Automatic Speech Recognition to Identify High-Quality Segments for More Robust Feature ExtractionabstractSpeech and language based automatic dementia detection is of interest due to it being non-invasive, low-cost and potentially able to aid diagnosis accuracy. The collected data are mostly audio recordings of spoken language and these can be used directly for acoustic-based analysis. To extract linguistic-based information, an automatic speech recognition (ASR) system is used to generate transcriptions. However, the extraction of reliable acoustic features is difficult when the acoustic quality of the data is poor as is the case with DementiaBank, the largest opensource dataset for Alzheimer’s Disease classification. In this paper, we explore how to improve the robustness of the acoustic feature extraction by using time alignment information and confidence scores from the ASR system to identify audio segments of good quality. In addition, we design rhythm-inspired features and combine them with acoustic features. By classifying the combined features with a bidirectional-LSTM attention network, the F-measure improves from 62.15% to 70.75% when only the high-quality segments are used. Finally, we apply the same approach to our previously proposed hierarchical-based network using linguistic-based features and show improvement from 74.37% to 77.25%. By combining the acoustic and linguistic systems, a state-of-the-art 78.34% F-measure is achieved on the DementiaBank task. Yilin Pan, Bahman Mirheidari, Markus Reuber, Annalena Venneri, Daniel Blackburn, Heidi Christensen |
INTERSPEECH | 6 |
| 2020 | Acoustic Feature Extraction with Interpretable Deep Neural Network for Neurodegenerative Related Disorder ClassificationabstractSpeech-based automatic approaches for detecting neurodegenerative disorders (ND) and mild cognitive impairment (MCI) have received more attention recently due to being non-invasive and potentially more sensitive than current pen-and-paper tests. The performance of such systems is highly dependent on the choice of features in the classification pipeline. In particular for acoustic features, arriving at a consensus for a best feature set has proven challenging. This paper explores using deep neural network for extracting features directly from the speech signal as a solution to this. Compared with hand-crafted features, more information is present in the raw waveform, but the feature extraction process becomes more complex and less interpretable which is often undesirable in medical domains. Using a SincNet as a first layer allows for some analysis of learned features. We propose and evaluate the Sinc-CLA (with SincNet, Convolutional, Long Short-Term Memory and Attention layers) as a task-driven acoustic feature extractor for classifying MCI, ND and healthy controls (HC). Experiments are carried out on an in-house dataset. Compared with the popular hand-crafted feature sets, the learned task-driven features achieve a superior classification accuracy. The filters of the SincNet is inspected and acoustic differences between HC, MCI and ND are found. Yilin Pan, Bahman Mirheidari, Zehai Tu, Ronan O'Malley, Traci Walker, Annalena Venneri, Markus Reuber, Daniel Blackburn, Heidi Christensen |
INTERSPEECH | 9 |
| 2020 | Autoencoder Bottleneck Features with Multi-Task Optimisation for Improved Continuous Dysarthric Speech RecognitionabstractAutomatic recognition of dysarthric speech is a very challenging research problem where performances still lag far behind those achieved for typical speech. The main reason is the lack of suitable training data to accommodate for the large mismatch seen between dysarthric and typical speech. Only recently has focus moved from single-word tasks to exploring continuous speech ASR needed for dictation and most voice-enabled interfaces. This paper investigates improvements to dysarthric continuous ASR. In particular, we demonstrate the effectiveness of using unsupervised autoencoder-based bottleneck (AE-BN) feature extractor trained on out-of-domain (OOD) LibriSpeech data. We further explore multi-task optimisation techniques shown to benefit typical speech ASR. We propose a 5-fold cross-training setup on the widely used TORGO dysarthric database. A setup we believe is more suitable for this low-resource data domain. Results show that adding the proposed AE-BN features achieves an average absolute (word error rate) WER improvement of 2.63% compared to the baseline system. A further reduction of 2.33% and 0.65% absolute WER is seen when applying monophone regularisation and joint optimisation techniques, respectively. In general, the ASR system employing monophone regularisation trained on AE-BN features exhibits the best performance. Zhengjun Yue, Heidi Christensen, Jon Barker |
INTERSPEECH | 2 |
| 2019 | Computational Cognitive Assessment: Investigating the Use of an Intelligent Virtual Agent for the Detection of Early Signs of DementiaabstractThe ageing population has caused a marked increased in the number of people with cognitive decline linked with dementia. Thus, current diagnostic services are overstretched, and there is an urgent need for automating parts of the assessment process. In previous work, we demonstrated how a stratification tool built around an Intelligent Virtual Agent (IVA) eliciting a conversation by asking memory-probing questions, was able to accurately distinguish between people with a neuro-degenerative disorder (ND) and a functional memory disorder (FMD). In this paper, we extend the number of diagnostic classes to include healthy elderly controls (HCs) as well as people with mild cognitive impairment (MCI). We also investigate whether the IVA may be used for administering more standard cognitive tests, like the verbal fluency tests. A four-way classifier trained on an extended feature set achieved 48% accuracy, which improved to 62% by using just the 22 most significant features (ROC-AUC: 82%). Bahman Mirheidari, Daniel Blackburn, Ronan O'Malley, Traci Walker, Annalena Venneri, Markus Reuber, Heidi Christensen |
ICASSP | 7 |
| 2019 | Phonetic Analysis of Dysarthric Speech Tempo and Applications to Robust Personalised Dysarthric Speech RecognitionabstractImproving the accuracy of personalised speech recognition for speakers with dysarthria is a challenging research field. In this paper, we explore an approach that non-linearly modifies speech tempo to reduce mismatch between typical and atypical speech. Speech tempo analysis at the phonetic level is accomplished using a forced-alignment process from traditional GMM-HMM in automatic speech recognition (ASR). Estimated tempo adjustments are applied directly to the acoustic features rather than to the time-domain signals. Two approaches are considered: i) adjusting dysarthric speech towards typical speech for input into ASR systems trained with typical speech, and ii) adjusting typical speech towards dysarthric speech for data augmentation in personalised dysarthric ASR training. Experimental results show that the latter strategy with data augmentation is more effective, resulting in a nearly 7% absolute improvement in comparison to baseline speaker-dependent trained system evaluated using UASpeech corpus. Consistent recognition performance improvements are observed across speakers, with greatest benefit in cases of moderate and severe dysarthria. Feifei Xiong, Jon Barker, Heidi Christensen |
ICASSP | 3 |
| 2019 | Automatic Hierarchical Attention Neural Network for Detecting ADabstractPicture description tasks are used for the detection of cognitive decline associated with Alzheimer's disease (AD). Recent years have seen work on automatic AD detection in picture descriptions based on acoustic and word-based analysis of the speech. These methods have shown some success but lack an ability to capture any higher-level effects of cognitive decline on the patient's language. In this paper, we propose a novel model that encompasses both the hierarchical and sequential structure of the description and detect its informative units by attention mechanism. Automatic speech recognition (ASR) and punctuation restoration are used to transcribe and segment the data. Using the DementiaBank database of people with AD as well as healthy controls (HC), we obtain an F-score of 84.43% and74.37% when using manual and automatic transcripts respectively. We further explore the effect of adding additional data (a total of 33 descriptions collected using a‘digital doctor’) during model training and increase the F-score when using ASR transcripts to 76.09%. This outperforms baseline models, including bidirectional LSTM and bidirectional hierarchical neural net-work without an attention mechanism, and demonstrate that the use of hierarchical models with attention mechanism improves the AD/HC discrimination performance. Yilin Pan, Bahman Mirheidari, Markus Reuber, Annalena Venneri, Daniel Blackburn, Heidi Christensen |
INTERSPEECH | 6 |
| 2019 | Dementia detection using automatic analysis of conversations
Bahman Mirheidari, Daniel Blackburn, Traci Walker, Markus Reuber, Heidi Christensen |
Comput. Speech Lang. | 5 |
| 2018 | Detecting Signs of Dementia Using Word Vector Representations
Bahman Mirheidari, Daniel Blackburn, Traci Walker, Annalena Venneri, Markus Reuber, Heidi Christensen |
INTERSPEECH | 6 |
| 2018 | Examining Temporal Variations in Recognizing Unspoken Words Using EEG SignalsabstractStudies on recognising unspoken speech with the use of electroencephalographic (EEG) signals vary in their designs. The participants are either asked to imagine unspoken speech within a specific time frame, or alternatively indicate the start and end of the imagined speech. Optimizing the length and training size of imagined speech is important to improve the rate and speed of recognizing unspoken speech in on-line applications. In this study, we recorded EEG data when the participants performed unspoken speech of five words using two technologies: (1) marking the start and end of the trial by using mouse clicks and (2) performing the imagination in a four-second fixed time window. Four classifiers were trained in all experiment parts: support vector machine, naive bayes, random forest, and linear discriminate analysis. The results show that the best time frame is 3.5-4 seconds length. Moreover, the increase in training size improve the average classification accuracy. However, this improvement becomes slight between 125-175 total training trials. The training data can be recorded in parts, however, the required training size should be increased to have better classification accuracy. In all analysis parts, random forest classifier shows better results among the other classifiers. Mashael M. AlSaleh, Roger K. Moore, Heidi Christensen, Mahnaz Arvaneh |
SMC | 3 |
| 2017 | On the impact of non-modal phonation on phonological featuresabstractDifferent modes of vibration of the vocal folds contribute significantly to the voice quality. The neutral mode phonation, often used in a modal voice, is one against which the other modes can be contrastively described, also called non-modal phonations. This paper investigates the impact of non-modal phonation on phonological posteriors, the probabilities of phonological features inferred from the speech signal using a deep learning approach. Five different non-modal phonations are considered: falsetto, creaky, harshness, tense and breathiness. The impact of such non-modal phonation on phonological features, the Sound Patterns of English (SPE), is investigated in both speech analysis and synthesis tasks. We found that breathy and tense phonation impact the SPE features less, creaky phonation impacts the features moderately, and harsh and falsetto phonation impact the phonological features the most. We also report invariant and the most different SPE features impacted by non-modal phonation. Milos Cernak, Elmar Nöth, Frank Rudzicz, Heidi Christensen, Juan Rafael Orozco-Arroyave, Raman Arora, Tobias Bocklet, Hamid R. Chinaei, Julius Hannink, Phani S. Nidadavolu, Juan Camilo Vásquez-Correa, Maria Yancheva, Alyssa Vann, Nikolai Vogler |
ICASSP | 4 |
| 2017 | Multi-view representation learning via gcca for multimodal analysis of Parkinson's diseaseabstractInformation from different bio-signals such as speech, handwriting, and gait have been used to monitor the state of Parkinson's disease (PD) patients, however, all the multimodal bio-signals may not always be available. We propose a method based on multi-view representation learning via generalized canonical correlation analysis (GCCA) for learning a representation of features extracted from handwriting and gait that can be used as a complement to speech-based features. Three different problems are addressed: classification of PD patients vs. healthy controls, prediction of the neurological state of PD patients according to the UPDRS score, and the prediction of a modified version of the Frenchay dysarthria assessment (m-FDA). According to the results, the proposed approach is suitable to improve the results in the addressed problems, specially in the prediction of the UPDRS, and m-FDA scores. Juan Camilo Vásquez-Correa, Juan Rafael Orozco-Arroyave, Raman Arora, Elmar Nöth, Najim Dehak, Heidi Christensen, Frank Rudzicz, Tobias Bocklet, Milos Cernak, Hamid R. Chinaei, Julius Hannink, Phani S. Nidadavolu, Maria Yancheva, Alyssa Vann, Nikolai Vogler |
ICASSP | 6 |
| 2017 | An Avatar-Based System for Identifying Individuals Likely to Develop DementiaabstractThis paper presents work on developing an automatic dementia screening test based on patients’ ability to interact and communicate — a highly cognitively demanding process where early signs of dementia can often be detected. Such a test would help general practitioners, with no specialist knowledge, make better diagnostic decisions as current tests lack specificity and sensitivity. We investigate the feasibility of basing the test on conversations between a ‘talking head’ (avatar) and a patient and we present a system for analysing such conversations for signs of dementia in the patient’s speech and language. Previously we proposed a semi-automatic system that transcribed conversations between patients and neurologists and extracted conversation analysis style features in order to differentiate between patients with progressive neurodegenerative dementia (ND) and functional memory disorders (FMD). Determining who talks when in the conversations was performed manually. In this study, we investigate a fully automatic system including speaker diarisation, and the use of additional acoustic and lexical features. Initial results from a pilot study are presented which shows that the avatar conversations can successfully classify ND/FMD with around 91% accuracy, which is in line with previous results for conversations that were led by a neurologist. \n Bahman Mirheidari, Daniel Blackburn, Kirsty Harkness, Traci Walker, Annalena Venneri, Markus Reuber, Heidi Christensen |
INTERSPEECH | 7 |
| 2017 | Characterisation of voice quality of Parkinson's disease using differential phonological posterior features
Milos Cernak, Juan Rafael Orozco-Arroyave, Frank Rudzicz, Heidi Christensen, Juan Camilo Vásquez-Correa, Elmar Nöth |
Comput. Speech Lang. | 4 |
| 2016 | CloudCAST - Remote Speech Technology for Speech ProfessionalsabstractInternational audience Phil D. Green, Ricard Marxer, Stuart P. Cunningham, Heidi Christensen, Frank Rudzicz, Maria Yancheva, André Coy, Massimiliano Malavasi, Lorenzo Desideri, Fabio Tamburini |
INTERSPEECH | 4 |
| 2016 | Diagnosing People with Dementia Using Automatic Conversation AnalysisabstractA recent study using Conversation Analysis (CA) has demonstrated that communication problems may be picked up during conversations between patients and neurologists, and that this can be used to differentiate between patients with (progressive neurodegenerative dementia) ND and those with (nonprogressive) functional memory disorders (FMD). This paper presents a novel automatic method for transcribing such conversations and extracting CA-style features. A range of acoustic, syntactic, semantic and visual features were automatically extracted and used to train a set of classifiers. In a proof-of-principle style study, using data recording during real neurologist-patient consultations, we demonstrate that automatically extracting CA-style features gives a classification accuracy of 95%when using verbatim transcripts. Replacing those transcripts with automatic speech recognition transcripts, we obtain a classification accuracy of 79% which improves to 90% when feature selection is applied. This is a first and encouraging step towards replacing inaccurate, potentially stressful cognitive tests with a test based on monitoring conversation capabilities that could be conducted in e.g. the privacy of the patient’s own home. \n \n Bahman Mirheidari, Daniel Blackburn, Markus Reuber, Traci Walker, Heidi Christensen |
INTERSPEECH | 5 |
| 2016 | A Framework for Collecting Realistic Recordings of Dysarthric Speech - the homeService Corpus
Mauro Nicolao, Heidi Christensen, Stuart P. Cunningham, Phil D. Green, Thomas Hain |
LREC | 2 |
| 2015 | Knowledge transfer between speakers for personalised dialogue managementabstractModel-free reinforcement learning has been shown to be a promising data driven approach for automatic dialogue policy optimization, but a relatively large amount of dialogue interactions is needed before the system reaches reasonable performance.Recently, Gaussian process based reinforcement learning methods have been shown to reduce the number of dialogues needed to reach optimal performance, and pre-training the policy with data gathered from different dialogue systems has further reduced this amount.Following this idea, a dialogue system designed for a single speaker can be initialised with data from other speakers, but if the dynamics of the speakers are very different the model will have a poor performance.When data gathered from different speakers is available, selecting the data from the most similar ones might improve the performance.We propose a method which automatically selects the data to transfer by defining a similarity measure between speakers, and uses this measure to weight the influence of the data from each speaker in the policy model.The methods are tested by simulating users with different severities of dysarthria interacting with a voice enabled environmental control system. Iñigo Casanueva, Thomas Hain, Heidi Christensen, Ricard Marxer, Phil D. Green |
SIGDIAL Conference | 3 |
| 2014 | Adaptive speech recognition and dialogue management for users with speech disordersabstractSpoken control interfaces are very attractive to people with severe physical disabilities who often also have a type of speech disorder known as dysarthria. This condition is known to decrease the accuracy of automatic speech recognisers (ASRs) especially for users with moderate to severe dysathria. In this paper we investigate how applying probabilistic dialogue management (DM) techniques can improve interaction performance of an environmental control system for such users. The effect of having access to different amounts of adaptation data, as well as using different vocabulary size for speakers of different intelligibilities is investigated. We explore the effect of adapting the DM models as the ASR performance increases, such as is the case in systems where more adaptation data is collected through system use. Improvements compared to a non-probabilistic DM baseline are seen both in terms of dialogue length and success rate, 9% and 25% mean relative improvement respectively. Looking at just the more severe dysarthric speakers these numbers rise 25% and 75% mean relative improvement. These improvements are higher when the ASR data adaptation amount is small. Further results show that a DM trained on data from multiple speakers outperform a DM trained on data from a single speaker. Iñigo Casanueva, Heidi Christensen, Thomas Hain, Phil D. Green |
INTERSPEECH | 2 |
| 2014 | Automatic selection of speakers for improved acoustic modelling: recognition of disordered speech with sparse dataabstractThe automatic recognition of disordered speech is a domain that is characterised by limited amounts of training data for each speaker and large intra- and inter-speaker variations. This paper is concerned with how best to train an acoustic models in these circumstances; in particular, we look at how to select data for a background model from a pool of speakers for a given target speaker. We show that rather than including data from all available speakers (the standard approach in the typical speech domain), significantly better accuracy can be achieved by carefully selecting which speakers should contribute. Different methods based on measuring acoustic closeness between speakers and ranking them accordingly are investigated, and on the UASpeech isolated word recognition task, we achieve a 11.5% relative improvement compared to the baseline which uses data from all speakers. Accuracies for speakers with moderate to severe impairments are shown to improve the most with one speaker classed as having `low' intelligibility gaining a 60% relative improvement in accuracy. Heidi Christensen, Iñigo Casanueva, Stuart P. Cunningham, Phil D. Green, Thomas Hain |
SLT | 1 |
| 2013 | Combining in-domain and out-of-domain speech data for automatic recognition of disordered speechabstractRecently there has been increasing interest in ways of using out-of-domain (OOD) data to improve automatic speech recognition performance in domains where only limited data is available. This paper focuses on one such domain, namely that of disordered speech for which only very small databases exist, but where normal speech can be considered OOD. Standard approaches for handling small data domains use adaptation from OOD models into the target domain, but here we investigate an alternative approach with its focus on the feature extraction stage: OOD data is used to train feature-generating deep belief neural networks. Using AMI meeting and TED talk datasets, we investigate various tandem-based speaker independent systems as well as maximum a posteriori adapted speaker dependent systems. Results on the UAspeech isolated word task of disordered speech are very promising with our overall best system (using a combination of AMI and TED data) giving a correctness of 62.5 an increase of 15% on previously best published results based on conventional model adaptation. We show that the relative benefit of using OOD data varies considerably from speaker to speaker and is only loosely correlated with the severity of a speaker's impairments. Heidi Christensen, Magda B. Aniol, Peter Bell 0001, Phil D. Green, Thomas Hain, Simon King 0001, Pawel Swietojanski |
INTERSPEECH | 1 |
| 2013 | Learning speaker-specific pronunciations of disordered speechabstractOne of the main clinical applications of speech technology is in voice-enabled assistive technology for people with disordered speech. Progress in this area is hampered by a sparseness in suitable data and recent research have focused on ways of incorporating knowledge about typical (i.e., un-impaired) speech through the use of e.g., deep belief neural networks. This paper presents a new way of using deep belief neural networks trained on typical speech, namely to improve pronunciations for individual speakers. Analysis of the posterior probabilities show a clear correlation between measured pronunciation ‘disorderedness’ and the overall speech recognition performance of the full system. Based on this, we propose a method to use deep belief network outputs to i) identify which words are pronounced differently than what would be expected from a typical pronunciation, and ii) subsequently generate new pronunciations. We investigate different methods for pronunciation generation as well as what is the best way of using the modified pronunciations to inform the system development stages. Using the UAspeech database of disordered speech, we demonstrate improvement in average accuracy of 69.76% to 70.51%, with some speakers showing individual improvements of up to 10%. Heidi Christensen, Phil D. Green, Thomas Hain |
INTERSPEECH | 1 |
| 2013 | Dysarthria intelligibility assessment in a factor analysis total variability spaceabstractSpeech technologies are more important every day to assist people with speech disorders. They can help to increase their quality of life or help clinicians to make a diagnosis. In this paper a new methodology based on a total variability subspace modelled by factor analysis is proposed to assess the intelligibility of people with dysarthria. The acoustic information of each recording is efficiently compressed and a Pearson correlation of 0.91 between the vectors in this subspace (iVectors) and the intelligibility is obtained. As acoustic information only perceptual linear prediction features are used. The experiments are conducted on Universal Access Speech database. Also a new error metric to overcome the subjectivity in the intelligibility labels is proposed. David Martínez González, Phil D. Green, Heidi Christensen |
INTERSPEECH | 3 |
| 2013 | The PASCAL CHiME speech separation and recognition challenge
Jon Barker, Emmanuel Vincent 0001, Ning Ma 0002, Heidi Christensen, Phil D. Green |
Comput. Speech Lang. | 4 |
| 2013 | A hearing-inspired approach for distant-microphone speech recognition in the presence of multiple sources
Ning Ma 0002, Jon Barker, Heidi Christensen, Phil D. Green |
Comput. Speech Lang. | 3 |
| 2012 | A comparative study of adaptive, automatic recognition of disordered speechabstractSpeech-driven assistive technology can be an attractive alternative to conventional interfaces for people with physical disabilities. However, often the lack of motor-control of the speech articulators results in disordered speech, as condition known as dysarthria. Dysarthric speakers can generally not obtain satisfactory performances with off-the-shelf automatic speech recognition (ASR) products and disordered speech ASR is an increasingly active research area. Sparseness of suitable data is a big challenge. The experiments described here use UAspeech, one of the largest dysarthric databases available, which is still easily an order of magnitude smaller than typical speech databases. This study investigates how far fundamental training and adaptation techniques developed in the LVCSR community can take us. A variety of ASR systems using maximum likelihood and MAP adaptation strategies are established with all speakers obtaining significant improvements compared to the baseline system regardless of the severity of their condition. The best systems show on average 34% relative improvement on known published results. An analysis of the correlation between intelligibility of the speaker and the type of system which would represent an optimal operating point in terms of performance shows that for severely dysarthric speakers, the exact choice of system configuration is more critical than for speakers with less disordered speech. Heidi Christensen, Stuart P. Cunningham, Charles Fox, Phil D. Green, Thomas Hain |
INTERSPEECH | 1 |
| 2012 | Combining Speech Fragment Decoding and Adaptive Noise Floor ModelingabstractThis paper presents a novel noise-robust automatic speech recognition (ASR) system that combines aspects of the noise modeling and source separation approaches to the problem. The combined approach has been motivated by the observation that the noise backgrounds encountered in everyday listening situations can be roughly characterized as a slowly varying noise floor in which there are embedded a mixture of energetic but unpredictable acoustic events. Our solution combines two complementary techniques. First, an adaptive noise floor model estimates the degree to which high-energy acoustic events are masked by the noise floor (represented by a soft missing data mask). Second, a fragment decoding system attempts to interpret the high-energy regions that are not accounted for by the noise floor model. This component uses models of the target speech to decide whether fragments should be included in the target speech stream or not. Our experiments on the CHiME corpus task show that the combined approach performs significantly better than systems using either the noise model or fragment decoding approach alone, and substantially outperforms multicondition training. Ning Ma 0002, Jon Barker, Heidi Christensen, Phil D. Green |
IEEE Trans. Speech Audio Process. | 3 |
| 2011 | Binaural Cues for Fragment-Based Speech Recognition in Reverberant Multisource EnvironmentsabstractThis paper addresses the problem of speech recognition using distant binaural microphones in reverberant multisource noise conditions. Our scheme employs a two stage fragment decoding approach: first spectro-temporal acoustic source fragments are identified using signal level cues, and second, a hypothesisdriven stage simultaneously searches for the most probable speech/background fragment labelling and the corresponding acoustic model state sequence. The paper reports the first successful attempt to use binaural localisation cues within this framework. By integrating binaural cues and acoustic models in a consistent probabilistic framework, the decoder is able to derive significant recognition performance benefits from fragment location estimates despite their inherent unreliability. Ning Ma 0002, Jon Barker, Heidi Christensen, Phil D. Green |
INTERSPEECH | 3 |
| 2010 | The CHiME corpus: a resource and a challenge for computational hearing in multisource environmentsabstractWe present a new corpus designed for noise-robust speech processing research, CHiME. Our goal was to produce material which is both natural (derived from reverberant domestic environments with many simultaneous and unpredictable sound sources) and controlled (providing an enumerated range of SNRs spanning 20 dB). The corpus includes around 40 hours of background recordings from a head and torso simulator positioned in a domestic setting, and a comprehensive set of binaural impulse responses collected in the same environment. These have been used to add target utterances from the Grid speech recognition corpus into the CHiME domestic setting. Data has been mixed in a manner that produces a controlled and yet natural range of SNRs over which speech separation, enhancement and recognition algorithms can be evaluated. The paper motivates the design of the corpus, and describes the collection and post-processing of the data. We also present a set of baseline recognition results. Heidi Christensen, Jon Barker, Ning Ma 0002, Phil D. Green |
INTERSPEECH | 1 |
| 2009 | A speech fragment approach to localising multiple speakers in reverberant environmentsabstractSound source localisation cues are severely degraded when multiple acoustic sources are active in the presence of reverberation. We present a binaural system for localising simultaneous speakers which exploits the fact that in a speech mixture there exist spectro-temporal regions or dasiafragmentspsila, where the energy is dominated by just one of the speakers. A fragment-level localisation model is proposed that integrates the localisation cues within a fragment using a weighted mean. The weights are based on local estimates of the degree of reverberation in a given spectro-temporal cell. The paper investigates different weight estimation approaches based variously on, i) an established model of the perceptual precedence effect; ii) a measure of interaural coherence between the left and right ear signals; iii) a data-driven approach trained in matched acoustic conditions. Experiments with reverberant binaural data with two simultaneous speakers show appropriate weighting can improve frame-based localisation performance by up to 24%. Heidi Christensen, Ning Ma 0002, Stuart N. Wrigley, Jon Barker |
ICASSP | 1 |
| 2009 | Using location cues to track speaker changes from mobile, binaural microphonesabstractThis paper presents initial developments towards computational hearing models that move beyond stationary microphone assumptions. We present a particle filtering based system for using localisation cues to track speaker changes in meeting recordings. Recording are made using in-ear binaural microphones worn by a listener whose head is constantly moving. Tracking speaker changes requires simultaneously inferring the perceiver’s head orientation, as any change in relative spatial angle to a source can be caused by either the source moving or the microphones moving. In real applications, such as robotics, there may be access to external estimates of the perceiver’s position. We investigate the effect of simulating varying degrees of measurement noise in an external perceiver position estimate. We show that only limited self-position knowledge is needed to greatly improve the reliability with which we can decode the acoustic localisation cues in the meeting scenario. Index Terms: speaker change tracking, binaural hearing, particle filtering, active listening Heidi Christensen, Jon Barker |
INTERSPEECH | 1 |
| 2008 | The CAVA corpus: synchronised stereoscopic and binaural datasets with head movementsabstractThis paper describes the acquisition and content of a new multi-modal database. Some tools for making use of the data streams are also presented. The Computational Audio-Visual Analysis (CAVA) database is a unique collection of three synchronised data streams obtained from a binaural microphone pair, a stereoscopic camera pair and a head tracking device. All recordings are made from the perspective of a person; i.e. what would a human with natural head movements see and hear in a given environment. The database is intended to facilitate research into humans' ability to optimise their multi-modal sensory input and fills a gap by providing data that enables human centred audio-visual scene analysis. It also enables 3D localisation using either audio, visual, or audio-visual cues. A total of 50 sessions, with varying degrees of visual and auditory complexity, were recorded. These range from seeing and hearing a single speaker moving in and out of field of view, to moving around a 'cocktail party' style situation, mingling and joining different small groups of people chatting. Elise Arnaud, Heidi Christensen, Yan-Chen Lu, Jon Barker, Vasil Khalidov, Miles E. Hansard, Bertrand Holveck, Hervé Mathieu, Ramya Narasimha, Elise Taillant, Florence Forbes, Radu Horaud |
ICMI | 2 |
| 2008 | A Cascaded Broadcast News HighlighterabstractThis paper presents a fully automatic news skimming system which takes a broadcast news audio stream and provides the user with the segmented, structured, and highlighted transcript. This constitutes a system with three different, cascading stages: converting the audio stream to text using an automatic speech recognizer, segmenting into utterances and stories, and finally determining which utterance should be highlighted using a saliency score. Each stage must operate on the erroneous output from the previous stage in the system, an effect which is naturally amplified as the data progresses through the processing stages. We present a large corpus of transcribed broadcast news data enabling us to investigate to which degree information worth highlighting survives this cascading of processes. Both extrinsic and intrinsic experimental results indicate that mistakes in the story boundary detection has a strong impact on the quality of highlights, whereas erroneous utterance boundaries cause only minor problems. Further, the difference in transcription quality does not affect the overall performance greatly. Heidi Christensen, Yoshihiko Gotoh, Steve Renals |
IEEE Trans. Speech Audio Process. | 1 |
| 2007 | Integrating pitch and localisation cues at a speech fragment levelabstractThis paper proposes a novel speech-fragment based approach for processing binaural data to improve the estimation of speech source locations in reverberant, multi-speaker recordings. The technique employs two stages. First, a robust multipitch tracking algorithm is used to locate local spectro-temporal ‘speech fragments ’ – regions where the energy in the mixture is dominated by a single speech source. Second, robust localisation estimates are formed by integrating interaural time difference cues over each speech fragment. The technique is applied to the analysis of more than five hours of two-party meetings that have been constructed from a mixture of binaural mannequin recordings. It is shown that estimating location at the speech fragment level produces better results than conventional location-estimate smoothing techniques leading to a an increase in relative frame accuracy rate of more than 35%. Index Terms: binaural localisation, pitch cues, speech fragment integration Heidi Christensen, Ning Ma 0002, Stuart N. Wrigley, Jon Barker |
INTERSPEECH | 1 |
| 2007 | Active binaural distance estimation for dynamic sourcesabstractA method for estimating sound source distance in dynamic auditory „scenes‟ using binaural data is presented. The technique requires little prior knowledge of the acoustic environment. It consists of feature extraction for two dynamic distance cues, motion parallax and acoustic τ, coupled with an inference framework for distance estimation. Sequential and nonsequential models are evaluated using simulated anechoic and reverberant spaces. Sequential approaches based on particle filtering more than half the distance estimation error in all conditions relative to the non-sequential models. These results confirm the value of active behaviour and probabilistic reasoning in auditorily-inspired models of distance perception. Yan-Chen Lu, Martin Cooke, Heidi Christensen |
INTERSPEECH | 3 |
| 2005 | Maximum entropy segmentation of broadcast newsabstractThe paper presents an automatic system for structuring and preparing a news broadcast for applications such as speech summarization, browsing, archiving and information retrieval. This process comprises transcribing the audio using an automatic speech recognizer and subsequently segmenting the text into utterances and topics. A maximum entropy approach is used to build statistical models for both utterance and topic segmentation. The experimental work addresses the effect on performance of the topic boundary detector of three factors - the types of feature used, the quality of the ASR transcripts, and the quality of the utterance boundary detector. The results show that the topic segmentation is not affected severely by transcript errors, whereas errors in utterance segmentation are more devastating. Heidi Christensen, BalaKrishna Kolluru, Yoshihiko Gotoh, Steve Renals |
ICASSP (1) | 1 |
| 2005 | Multi-stage compaction approach to broadcast news summarisationabstractThis paper presents a fully automatic, multi-stage compaction approach to broadcast news summarisation, targeting transcripts from automatic speech recognition (ASR) systems. It employs a network of multi-layer perceptrons to remove incorrectly transcribed words based on confidence scores, and to select significant chunks at multiple stages based on tf.idf scores and named entity frequency. The resulting summaries are assessed using a combination of cross comprehension test and a fluency test, finally compared with an automatic evaluation scheme. The experimental results show the approach can produce summaries with good information content. 1. BalaKrishna Kolluru, Heidi Christensen, Yoshihiko Gotoh |
INTERSPEECH | 2 |
| 2004 | From Text Summarisation to Style-Specific Summarisation for Broadcast News
Heidi Christensen, BalaKrishna Kolluru, Yoshihiko Gotoh, Steve Renals |
ECIR | 1 |
| 2001 | Introducing phonetically motivated information into ASR
Heidi Christensen, Børge Lindberg, Ove Andersen |
INTERSPEECH | 1 |
| 2000 | Employing heterogeneous information in a multi-stream frameworkabstractA multi-stream speech recogniser is based on the combination of multiple feature streams each containing complementary information. In the past, multi-stream research has typically focused on systems that use a single feature extraction method. This heritage from conventional speech recognisers is an unnecessary restriction and both psychoacoustic and phonetic knowledge strongly motivate the use of heterogeneous features. In this paper we investigate how heterogeneous processing can be used in two different multi-stream configurations: first, a system where each stream handles a different frequency region of the speech (a multi-band recogniser) and, second a multi-stream recogniser where each stream handles the full frequency region. For each type of system we compare the performance using both homogeneous and heterogeneous processing. We demonstrate that the use of heterogeneous information significantly improves the clean speech recognition performance motivating us to continue exploring more specifically designed stream processing. Heidi Christensen, Børge Lindberg, Ove Andersen |
ICASSP | 1 |
| 2000 | Noise robustness of heterogeneous features employing minimum classification error feature space transformations
Heidi Christensen, Børge Lindberg, Ove Andersen |
INTERSPEECH | 1 |