EDBT 2026 Demo / reviewers in the wild / expert
Daniel Blackburn
dblp:194/1171 · also Daniel J. Blackburn
· DBLP profile ↗
24ranked-venue papers
0as first author
13since 2021 · last 2026
0000-0001-8886-1283ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 9 since 2021Artificial intelligence and machine learning · 17 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Causal Validation augmented Temporal Convolutional Framework for Brain Effective Connectivity Networks EstimationabstractAdvancements in neuroimaging have facilitated unprecedented insights into brain connectivity, making the study of brain effective connectivity networks (ECNs) essential for understanding neurological functions and diseases. Recently, neural networks (NNs) have emerged as powerful tools for ECN estimation due to their prominent universal approximation ability and less reliance on prior knowledge. However, most NN-based approaches fail to eliminate redundant temporal information and lack rigorous causal validation mechanisms. This paper introduces a novel end-to-end framework for estimating ECNs utilising Least Absolute Shrinkage and Selection Operator (Lasso) regression of Temporal Convolutional Networks (TCNs), named the Causal Validation augmented Temporal Convolutional Framework (CVTCF). In the CVTCF, a convolutional Hierarchical Group Lasso (cHGL) is proposed to detect Granger Causality (GC) inputs and eliminate redundant temporal information during GC detection. Additionally, the framework incorporates permutation importance validation based on the Wilcoxon signed-rank test to enhance the reliability of GC detection. The proposed CVTCF generally outperformed state-of-the-art methods in a controlled simulation using the chaotic Lorenz-96 model and the publicly available blood-oxygen-level-dependent (BOLD) benchmark dataset. Furthermore, the proposed CVTCF has enabled a detailed analysis of the causal interactions within the cerebral cortex, bringing to light the intricate relationships that underlie neurological functioning and impairment of neurodegenerative conditions like Alzheimer's Disease (AD) and Parkinson's Disease (PD). This study demonstrates the potential of using ECN estimation based on the CVTCF as indicators for neurodegenerative diseases and paves the way for future diagnostic and therapeutic strategies. Aoxiang Dong, Ptolemaios G. Sarrigiannis, Daniel Blackburn, Andrew Starr 0001, Yifan Zhao 0001 |
Neural Networks | 4 |
| 2025 | Early Dementia Detection Using Multiple Spontaneous Speech Prompts: The PROCESS ChallengeabstractDementia is associated with various cognitive impairments and typically manifests only after significant progression, making intervention at this stage often ineffective. To address this issue, the Prediction and Recognition of Cognitive Decline through Spontaneous Speech (PROCESS) Signal Processing Grand Challenge invites participants to focus on early-stage dementia detection. We provide a new spontaneous speech corpus for this challenge. This corpus includes answers from three prompts designed by neurologists to better capture the cognition of speakers. Our baseline models achieved an F1-score of 55.0% on the classification task and an RMSE of 2.98 on the regression task. Fuxiang Tao, Bahman Mirheidari, Madhurananda Pahar, Sophie Young, Hend Elghazaly, Fritz Peters, Caitlin H. Illingworth, Dorota Braun, Ronan O'Malley, Simon Bell, Daniel Blackburn, Fasih Haider, Saturnino Luz, Heidi Christensen |
ICASSP | 12 |
| 2025 | Automatic Detection and Sub-typing of Primary Progressive Aphasia from Speech: Integrating Task-Specific Features and Spatio-Semantic GraphsabstractPrimary progressive aphasia (PPA) describes a group of neurodegenerative diseases that predominantly affect language abilities. Its diagnostic process typically requires experienced clinicians, often available only in specialised hospital departments. Patients with PPA frequently display changes in speech and language early in the disease progression. In this study, we extracted acoustic, linguistic, and task-specific features from audio recordings and evaluated their utility for PPA classification. Using a subset of task-specific features, we detected PPA with 97% accuracy. For sub-typing, models trained on the full feature set achieved 74% accuracy in a three-way classification of PPA variants. Our results highlight the added value of task-specific features, which complement traditional approaches. Additionally, their visualisation offers an intuitive representation of task execution, improving clinical interpretability and potential diagnostic utility. Fritz Peters, W. Richard Bevan-Jones, Grace Threlfall, Jenny M. Harris, Julie S. Snowden, Jennifer C. Thompson, Daniel Blackburn, Heidi Christensen |
INTERSPEECH | 8 |
| 2025 | Automatic Detection of Early Cognitive Decline Using Multimodal Feature Fusion and Transfer Learning on Real-World Conversational SpeechabstractEarly signs of cognitive decline, such as dementia and mild cognitive impairment (MCI), often manifest in conversational speech. Early and accurate identification is essential for potential interventions prior to the onset of more severe stages of neurodegenerative diseases. We present CognoMemory, a system for detecting cognitive decline based on a person's speech, to collect 307 hrs of real-world conversational speech, corresponding to 1.92 million Whisper-transcribed words, from 1,639 participants. Speech recordings were collected as participants answered 14 memory-probing, clinically effective questions asked by a virtual agent, starting with a motivation prompt, followed by memory, cognitive functioning, fluency, picture description and reading task. Both acoustic and linguistic features, along with large language model (LLM) embeddings, were extracted from all 1,639 participants. A subset of 614 participants, either with an unconfirmed diagnosis or younger than 50 years, was used for pre-training. The remaining three groups (64 dementia, 169 MCI and 792 healthy participants) were used to fine-tune our proposed model. Our multimodal feature fusion and CNN/Bi-LSTM-based transfer learning approach outperforms LLM-based (BART, DistilBERT, RoBERTa and HuBERT) approaches while achieving the highest $F_{1}$-scores of 0.83 & 0.54 using just the initial 'motivation' question for 2-way & 3-way classification; exhibiting a 3% performance increase due to the application of transfer learning, while being also 38% faster. Finally, the classifiers trained on the CognoMemory data, the largest of its kind, were tested on the second-largest available DementiaBank dataset (Pitt corpus), and a CNN-based transfer learning architecture achieved an $F_{1}$-score of 0.89, demonstrating better stability and generalisation across datasets and of our novel feature fusion and architecture. Madhurananda Pahar, Bahman Mirheidari, Caitlin H. Illingworth, Dorota Braun, Fuxiang Tao, Lise Sproson, Daniel Blackburn, Heidi Christensen |
IEEE J. Biomed. Health Informatics | 7 |
| 2023 | Identifying People with Mild Cognitive Impairment at Risk of Developing Dementia using Speech AnalysisabstractMild Cognitive Impairment (MCI) is the intermediate stage between ageing and possible dementia. $50 \%$ of people with MCI may progress to dementia (prodromal Alzheimer’s Disease (pr-AD)). Identifying those at risk is a challenging but important task that can help with anxiety and treatment. Currently, clinicians wait for severe signs of impairment to emerge, however, recent studies have shown promising automatic speech-based approaches. This paper works on a unique dataset containing 50 healthy controls (HCs) and 50 MCI (of which 22 are pr-AD and 28 are non-progressed (nonP)). The recordings are of people speaking with a virtual agent. A number of acoustic, text and contextual features were extracted and used for training classifiers. The best result was achieved with an $F_{1}$-scores of $\mathbf{81.2 \%}$ for the detection of MCI versus HC, and $\mathbf{75} \%$ for MCI(pr-AD) versus MCI(nonP). The best $F_{1}$-score of $\mathbf{66.9 \%}$ was achieved for the detection of MCI (pr-AD) versus MCI (nonP) versus HC. Bahman Mirheidari, Ronan O'Malley, Daniel Blackburn, Heidi Christensen |
ASRU | 3 |
| 2023 | Investigating Visual Features for Cognitive Impairment Detection Using In-the-wild DataabstractEarly detection of dementia has attracted much research interest due to its crucial role in helping people get suitable treatment or care. Video analysis may provide an effective approach for detection, with low cost and effort compared to current expensive and intensive clinical assessments. This paper investigates the use of a range of visual features - eye blink rate (EBR), head turn rate (HTR) and head movement statistical features (HMSF) - for identifying neurodegenerative disorder (ND), mild cognitive impairment (MCI) and functional memory disorder (FMD). These features are used in a noval multiple thresholds approach, which is applied to an in-the-wild video dataset which includes data recorded in a range of challenging environments. A combination of EBR and HTR gives 78 % accuracy in a three-way classification task (ND/MCI/FMD) and 83%, 83% and 92%, respectively, for the two-way classifications ND/MCI, ND/FMD and MCI/FMD. These results are comparable to related work that uses more features from different modalities. They also provide evidence to support the possibility of an in-the-home detection process for dementia or cognitive impairment. Fatimah Alzahrani, Bahman Mirheidari, Daniel Blackburn, Steve C. Maddock, Heidi Christensen |
FG | 3 |
| 2022 | Evaluating the Performance of State-of-the-Art ASR Systems on Non-Native English using Corpora with Extensive Language Background VariationabstractThis investigation is an exploration into the performance of several different ASR systems in dealing with non-native English using corpora with extensive language background variation. This study takes two corpora amounting to 191 different native language (L1) backgrounds and looks at how these systems are able to process non-native English (L2) speech. A transformer based ASR system and a CRDNN architecture are both tested, trained on Librispeech [1] and Commonvoice [2] for a three way cross comparison. In addition Google's Speech-to-Text API and AWS Transcribe were investigated in order to evaluate popular mainstream approaches given their current degree of impact in deployed systems. Experiments reveal deficits in the range of 10%-15% mean WER performance difference between L1 and L2 speech. Results indicate ASR systems trained on particular varieties of L2 speech may be effective in improving WERs with outcomes in this paper demonstrating several Google ASR models trained on varieties of African L2 English outperforming L1 trained ASR for under-represented dialect groups in the United Kingdom. Further research is proposed to explore the plausibility of this approach and to critically approach WER as a metric for ASR evaluation, striving instead towards metrics with greater emphasis on evaluating language for communication. Samuel Schmück, Daniel Blackburn, Heidi Christensen |
INTERSPEECH | 2 |
| 2022 | Automatic cognitive assessment: Combining sparse datasets with disparate cognitive scores
Bahman Mirheidari, Daniel Blackburn, Heidi Christensen |
INTERSPEECH | 2 |
| 2022 | Characterising Alzheimer's Disease With EEG-Based Energy Landscape AnalysisabstractAlzheimer's disease (AD) is one of the most common neurodegenerative diseases, with around 50 million patients worldwide. Accessible and non-invasive methods of diagnosing and characterising AD are therefore urgently required. Electroencephalography (EEG) fulfils these criteria and is often used when studying AD. Several features derived from EEG were shown to predict AD with high accuracy, e.g. signal complexity and synchronisation. However, the dynamics of how the brain transitions between stable states have not been properly studied in the case of AD and EEG. Energy landscape analysis is a method that can be used to quantify these dynamics. This work presents the first application of this method to both AD and EEG. Energy landscape assigns energy value to each possible state, i.e. pattern of activations across brain regions. The energy is inversely proportional to the probability of occurrence. By studying the features of energy landscapes of 20 AD patients and 20 age-matched healthy counterparts (HC), significant differences are found. The dynamics of AD patients' EEG are shown to be more constrained - with more local minima, less variation in basin size, and smaller basins. We show that energy landscapes can predict AD with high accuracy, performing significantly better than baseline models. Moreover, these findings are replicated in a separate dataset including 9 AD and 10 HC above 70 years old. Dominik Klepl, Fei He 0002, Min Wu 0008, Matteo De Marco, Daniel Blackburn, Ptolemaios G. Sarrigiannis |
IEEE J. Biomed. Health Informatics | 5 |
| 2021 | Eye Blink Rate Based Detection of Cognitive Impairment Using In-the-wild DataabstractInvestigating automatic methods for the early detection of dementia and related conditions that cause cognitive impairment is an area of growing interest. Video processing could play a role by providing a non-invasive and low-cost alternative to current expensive assessments. For this to be successful it is crucial that approaches are robust to in-the-wild challenges. In this paper, visual cues, related to the eye blink rate (EBR), are investigated to quantify the early phase of neurodegenerative disorder (ND) and mild cognitive impairment (MCI) as well as functional memory disorder (FMD; problems with memory not related to neurodegenerative disorder). This paper aims to improve the detection of ND and MCI by investigating a novel approach to calculating the EBR that is more robust to in-the-wild challenges. An in-house dataset with 18 participants is used. The EBR is calculated from eye landmarks extracted using two libraries (Dlib and Openface). To mitigate issues observed in the noisy, in-the-wild recordings, a multiple threshold approach for EBR detection is proposed. It involves generating multiple thresholds for identifying a blink, where a threshold is used to determine whether an eye is open or closed. Several supervised machine learning approaches are used for automatic classification. The results show that accuracy measures of 89% and 78% are achieved using Dlib and OpenFace data, respectively, when distinguishing between three conditions with ND, MCI and FMD. Fatimah Alzahrani, Bahman Mirheidari, Daniel Blackburn, Steve C. Maddock, Heidi Christensen |
ACII | 3 |
| 2021 | Multi-Task Estimation of Age and Cognitive Decline from SpeechabstractSpeech is a common physiological signal that can be affected by both ageing and cognitive decline. Often the effect can be confounding, as would be the case for people at, e.g., very early stages of cognitive decline due to dementia. Despite this, the automatic predictions of age and cognitive decline based on cues found in the speech signal are generally treated as two separate tasks. In this paper, multi-task learning is applied for the joint estimation of age and the Mini-Mental Status Evaluation criteria (MMSE) commonly used to assess cognitive decline. To explore the relationship between age and MMSE, two neural network architectures are evaluated: a SincNet-based end-to-end architecture, and a system comprising of a feature extractor followed by a shallow neural network. Both are trained with single-task or multi-task targets. To compare, an SVM-based regressor is trained in a single-task setup. i-vector, x-vector and ComParE features are explored. Results are obtained on systems trained on the DementiaBank dataset and tested on an in-house dataset as well as the ADReSS dataset. The results show that both the age and MMSE estimation is improved by applying multitask learning, with state-of-the-art results achieved on the ADReSS dataset acoustic-only task. Yilin Pan, Venkata Srikanth Nallanthighal, Daniel Blackburn, Heidi Christensen, Aki Härmä |
ICASSP | 3 |
| 2021 | Identifying Cognitive Impairment Using Sentence Representation VectorsabstractThe widely used word vectors can be extended at the sentence level to perform a wide range of natural language processing (NLP) tasks.Recently the Bidirectional Encoder Representations from Transformers (BERT) language representation achieved state-of-the-art performance for these applications.The model is trained with punctuated and well-formed (writ-ten) text, however, the performance of the model drops significantly when the input text is the -erroneous and unpunctuated-output of automatic speech recognition (ASR).We use a sliding window and averaging approach for pre-processing text for BERT to extract features for classifying three diagnostic categories relating to cognitive impairment: neurodegenerative dis-order (ND), mild cognitive impairment (MCI), and healthy controls (HC).The in-house dataset contains the audio recordings of an intelligent virtual agent (IVA) who asks the participants several conversational questions prompts in addition to giving a picture description prompt.For the three-way classification, we achieve a 73.88% F-score (accuracy: 76.53%) using the pre-trained, uncased base BERT and for the two-way classifier (HCvs.ND) we achieve 89.80% (accuracy: 90%).We further improve these by using a prompt selection technique, reaching the F-scores of 79.98% (accuracy: 81.63%) and 93.56% (accuracy:93.75%)respectively. Bahman Mirheidari, Yilin Pan, Daniel Blackburn, Ronan O'Malley, Heidi Christensen |
Interspeech | 3 |
| 2021 | Using the Outputs of Different Automatic Speech Recognition Paradigms for Acoustic- and BERT-Based Alzheimer's Dementia Detection Through Spontaneous Speech
Yilin Pan, Bahman Mirheidari, Jennifer M. Harris, Jennifer C. Thompson, Julie S. Snowden, Daniel Blackburn, Heidi Christensen |
Interspeech | 7 |
| 2020 | A Comparison of Acoustic and Linguistics Methodologies for Alzheimer's Dementia RecognitionabstractContains fulltext : 228158.pdf (Publisher’s version ) (Open Access) Nicholas Cummins, Yilin Pan, Zhao Ren, Julian Fritsch, Venkata Srikanth Nallanthighal, Heidi Christensen, Daniel Blackburn, Björn W. Schuller, Mathew Magimai-Doss, Helmer Strik, Aki Härmä |
INTERSPEECH | 7 |
| 2020 | Improving Cognitive Impairment Classification by Generative Neural Network-Based Feature Augmentation
Bahman Mirheidari, Daniel Blackburn, Ronan O'Malley, Annalena Venneri, Traci Walker, Markus Reuber, Heidi Christensen |
INTERSPEECH | 2 |
| 2020 | Improving Detection of Alzheimer's Disease Using Automatic Speech Recognition to Identify High-Quality Segments for More Robust Feature ExtractionabstractSpeech and language based automatic dementia detection is of interest due to it being non-invasive, low-cost and potentially able to aid diagnosis accuracy. The collected data are mostly audio recordings of spoken language and these can be used directly for acoustic-based analysis. To extract linguistic-based information, an automatic speech recognition (ASR) system is used to generate transcriptions. However, the extraction of reliable acoustic features is difficult when the acoustic quality of the data is poor as is the case with DementiaBank, the largest opensource dataset for Alzheimer’s Disease classification. In this paper, we explore how to improve the robustness of the acoustic feature extraction by using time alignment information and confidence scores from the ASR system to identify audio segments of good quality. In addition, we design rhythm-inspired features and combine them with acoustic features. By classifying the combined features with a bidirectional-LSTM attention network, the F-measure improves from 62.15% to 70.75% when only the high-quality segments are used. Finally, we apply the same approach to our previously proposed hierarchical-based network using linguistic-based features and show improvement from 74.37% to 77.25%. By combining the acoustic and linguistic systems, a state-of-the-art 78.34% F-measure is achieved on the DementiaBank task. Yilin Pan, Bahman Mirheidari, Markus Reuber, Annalena Venneri, Daniel Blackburn, Heidi Christensen |
INTERSPEECH | 5 |
| 2020 | Acoustic Feature Extraction with Interpretable Deep Neural Network for Neurodegenerative Related Disorder ClassificationabstractSpeech-based automatic approaches for detecting neurodegenerative disorders (ND) and mild cognitive impairment (MCI) have received more attention recently due to being non-invasive and potentially more sensitive than current pen-and-paper tests. The performance of such systems is highly dependent on the choice of features in the classification pipeline. In particular for acoustic features, arriving at a consensus for a best feature set has proven challenging. This paper explores using deep neural network for extracting features directly from the speech signal as a solution to this. Compared with hand-crafted features, more information is present in the raw waveform, but the feature extraction process becomes more complex and less interpretable which is often undesirable in medical domains. Using a SincNet as a first layer allows for some analysis of learned features. We propose and evaluate the Sinc-CLA (with SincNet, Convolutional, Long Short-Term Memory and Attention layers) as a task-driven acoustic feature extractor for classifying MCI, ND and healthy controls (HC). Experiments are carried out on an in-house dataset. Compared with the popular hand-crafted feature sets, the learned task-driven features achieve a superior classification accuracy. The filters of the SincNet is inspected and acoustic differences between HC, MCI and ND are found. Yilin Pan, Bahman Mirheidari, Zehai Tu, Ronan O'Malley, Traci Walker, Annalena Venneri, Markus Reuber, Daniel Blackburn, Heidi Christensen |
INTERSPEECH | 8 |
| 2020 | Imaging of Nonlinear and Dynamic Functional Brain Connectivity Based on EEG Recordings With the Application on the Diagnosis of Alzheimer's DiseaseabstractSince age is the most significant risk factor for the development of Alzheimer's disease (AD), it is important to understand the effect of normal ageing on brain network characteristics before we can accurately diagnose the condition based on information derived from resting state electroencephalogram (EEG) recordings, aiming to detect brain network disruption. This article proposes a novel brain functional connectivity imaging method, particularly targeting the contribution of nonlinear dynamics of functional connectivity, on distinguishing participants with AD from healthy controls (HC). We describe a parametric method established upon a Nonlinear Finite Impulse Response model, and a revised orthogonal least squares algorithm used to estimate the linear, nonlinear and combined connectivity between any two EEG channels without fitting a full model. This approach, where linear and non-linear interactions and their spatial distribution and dynamics can be estimated independently, offered us the means to dissect the dynamic brain network disruption in AD from a new perspective and to gain some insight into the dynamic behaviour of brain networks in two age groups (above and below 70) with normal cognitive function. Although linear and stationary connectivity dominates the classification contributions, quantitative results have demonstrated that nonlinear and dynamic connectivity can significantly improve the classification accuracy, barring the group of participants below the age of 70, for resting state EEG recorded during eyes open. The developed approach is generic and can be used as a powerful tool to examine brain network characteristics and disruption in a user friendly and systematic way. Yifan Zhao 0001, Yitian Zhao, Pholpat Durongbhan, Jiang Liu 0001, Stephen A. Billings, Panagiotis Zis, Zoe C. Unwin, Matteo De Marco, Annalena Venneri, Daniel Blackburn, Ptolemaios G. Sarrigiannis |
IEEE Trans. Medical Imaging | 11 |
| 2019 | Computational Cognitive Assessment: Investigating the Use of an Intelligent Virtual Agent for the Detection of Early Signs of DementiaabstractThe ageing population has caused a marked increased in the number of people with cognitive decline linked with dementia. Thus, current diagnostic services are overstretched, and there is an urgent need for automating parts of the assessment process. In previous work, we demonstrated how a stratification tool built around an Intelligent Virtual Agent (IVA) eliciting a conversation by asking memory-probing questions, was able to accurately distinguish between people with a neuro-degenerative disorder (ND) and a functional memory disorder (FMD). In this paper, we extend the number of diagnostic classes to include healthy elderly controls (HCs) as well as people with mild cognitive impairment (MCI). We also investigate whether the IVA may be used for administering more standard cognitive tests, like the verbal fluency tests. A four-way classifier trained on an extended feature set achieved 48% accuracy, which improved to 62% by using just the 22 most significant features (ROC-AUC: 82%). Bahman Mirheidari, Daniel Blackburn, Ronan O'Malley, Traci Walker, Annalena Venneri, Markus Reuber, Heidi Christensen |
ICASSP | 2 |
| 2019 | Automatic Hierarchical Attention Neural Network for Detecting ADabstractPicture description tasks are used for the detection of cognitive decline associated with Alzheimer's disease (AD). Recent years have seen work on automatic AD detection in picture descriptions based on acoustic and word-based analysis of the speech. These methods have shown some success but lack an ability to capture any higher-level effects of cognitive decline on the patient's language. In this paper, we propose a novel model that encompasses both the hierarchical and sequential structure of the description and detect its informative units by attention mechanism. Automatic speech recognition (ASR) and punctuation restoration are used to transcribe and segment the data. Using the DementiaBank database of people with AD as well as healthy controls (HC), we obtain an F-score of 84.43% and74.37% when using manual and automatic transcripts respectively. We further explore the effect of adding additional data (a total of 33 descriptions collected using a‘digital doctor’) during model training and increase the F-score when using ASR transcripts to 76.09%. This outperforms baseline models, including bidirectional LSTM and bidirectional hierarchical neural net-work without an attention mechanism, and demonstrate that the use of hierarchical models with attention mechanism improves the AD/HC discrimination performance. Yilin Pan, Bahman Mirheidari, Markus Reuber, Annalena Venneri, Daniel Blackburn, Heidi Christensen |
INTERSPEECH | 5 |
| 2019 | Dementia detection using automatic analysis of conversations
Bahman Mirheidari, Daniel Blackburn, Traci Walker, Markus Reuber, Heidi Christensen |
Comput. Speech Lang. | 2 |
| 2018 | Detecting Signs of Dementia Using Word Vector Representations
Bahman Mirheidari, Daniel Blackburn, Traci Walker, Annalena Venneri, Markus Reuber, Heidi Christensen |
INTERSPEECH | 2 |
| 2017 | An Avatar-Based System for Identifying Individuals Likely to Develop DementiaabstractThis paper presents work on developing an automatic dementia screening test based on patients’ ability to interact and communicate — a highly cognitively demanding process where early signs of dementia can often be detected. Such a test would help general practitioners, with no specialist knowledge, make better diagnostic decisions as current tests lack specificity and sensitivity. We investigate the feasibility of basing the test on conversations between a ‘talking head’ (avatar) and a patient and we present a system for analysing such conversations for signs of dementia in the patient’s speech and language. Previously we proposed a semi-automatic system that transcribed conversations between patients and neurologists and extracted conversation analysis style features in order to differentiate between patients with progressive neurodegenerative dementia (ND) and functional memory disorders (FMD). Determining who talks when in the conversations was performed manually. In this study, we investigate a fully automatic system including speaker diarisation, and the use of additional acoustic and lexical features. Initial results from a pilot study are presented which shows that the avatar conversations can successfully classify ND/FMD with around 91% accuracy, which is in line with previous results for conversations that were led by a neurologist. \n Bahman Mirheidari, Daniel Blackburn, Kirsty Harkness, Traci Walker, Annalena Venneri, Markus Reuber, Heidi Christensen |
INTERSPEECH | 2 |
| 2016 | Diagnosing People with Dementia Using Automatic Conversation AnalysisabstractA recent study using Conversation Analysis (CA) has demonstrated that communication problems may be picked up during conversations between patients and neurologists, and that this can be used to differentiate between patients with (progressive neurodegenerative dementia) ND and those with (nonprogressive) functional memory disorders (FMD). This paper presents a novel automatic method for transcribing such conversations and extracting CA-style features. A range of acoustic, syntactic, semantic and visual features were automatically extracted and used to train a set of classifiers. In a proof-of-principle style study, using data recording during real neurologist-patient consultations, we demonstrate that automatically extracting CA-style features gives a classification accuracy of 95%when using verbatim transcripts. Replacing those transcripts with automatic speech recognition transcripts, we obtain a classification accuracy of 79% which improves to 90% when feature selection is applied. This is a first and encouraging step towards replacing inaccurate, potentially stressful cognitive tests with a test based on monitoring conversation capabilities that could be conducted in e.g. the privacy of the patient’s own home. \n \n Bahman Mirheidari, Daniel Blackburn, Markus Reuber, Traci Walker, Heidi Christensen |
INTERSPEECH | 2 |