EDBT 2026 Demo / reviewers in the wild / expert
Bahman Mirheidari
dblp:194/1324
· DBLP profile ↗
19ranked-venue papers
10as first author
10since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 9 first-author · 8 since 2021Artificial intelligence and machine learning · 15 · 9 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Early Dementia Detection Using Multiple Spontaneous Speech Prompts: The PROCESS ChallengeabstractDementia is associated with various cognitive impairments and typically manifests only after significant progression, making intervention at this stage often ineffective. To address this issue, the Prediction and Recognition of Cognitive Decline through Spontaneous Speech (PROCESS) Signal Processing Grand Challenge invites participants to focus on early-stage dementia detection. We provide a new spontaneous speech corpus for this challenge. This corpus includes answers from three prompts designed by neurologists to better capture the cognition of speakers. Our baseline models achieved an F1-score of 55.0% on the classification task and an RMSE of 2.98 on the regression task. Fuxiang Tao, Bahman Mirheidari, Madhurananda Pahar, Sophie Young, Hend Elghazaly, Fritz Peters, Caitlin H. Illingworth, Dorota Braun, Ronan O'Malley, Simon Bell, Daniel Blackburn, Fasih Haider, Saturnino Luz, Heidi Christensen |
ICASSP | 2 |
| 2025 | Can Speech Accurately Detect Depression in Patients With Comorbid Dementia? An Approach for Mitigating Confounding Effects of Depression and Dementia
Sophie Young, Fuxiang Tao, Bahman Mirheidari, Madhurananda Pahar, Markus Reuber, Heidi Christensen |
INTERSPEECH | 3 |
| 2025 | Automatic Detection of Early Cognitive Decline Using Multimodal Feature Fusion and Transfer Learning on Real-World Conversational SpeechabstractEarly signs of cognitive decline, such as dementia and mild cognitive impairment (MCI), often manifest in conversational speech. Early and accurate identification is essential for potential interventions prior to the onset of more severe stages of neurodegenerative diseases. We present CognoMemory, a system for detecting cognitive decline based on a person's speech, to collect 307 hrs of real-world conversational speech, corresponding to 1.92 million Whisper-transcribed words, from 1,639 participants. Speech recordings were collected as participants answered 14 memory-probing, clinically effective questions asked by a virtual agent, starting with a motivation prompt, followed by memory, cognitive functioning, fluency, picture description and reading task. Both acoustic and linguistic features, along with large language model (LLM) embeddings, were extracted from all 1,639 participants. A subset of 614 participants, either with an unconfirmed diagnosis or younger than 50 years, was used for pre-training. The remaining three groups (64 dementia, 169 MCI and 792 healthy participants) were used to fine-tune our proposed model. Our multimodal feature fusion and CNN/Bi-LSTM-based transfer learning approach outperforms LLM-based (BART, DistilBERT, RoBERTa and HuBERT) approaches while achieving the highest $F_{1}$-scores of 0.83 & 0.54 using just the initial 'motivation' question for 2-way & 3-way classification; exhibiting a 3% performance increase due to the application of transfer learning, while being also 38% faster. Finally, the classifiers trained on the CognoMemory data, the largest of its kind, were tested on the second-largest available DementiaBank dataset (Pitt corpus), and a CNN-based transfer learning architecture achieved an $F_{1}$-score of 0.89, demonstrating better stability and generalisation across datasets and of our novel feature fusion and architecture. Madhurananda Pahar, Bahman Mirheidari, Caitlin H. Illingworth, Dorota Braun, Fuxiang Tao, Lise Sproson, Daniel Blackburn, Heidi Christensen |
IEEE J. Biomed. Health Informatics | 2 |
| 2023 | Identifying People with Mild Cognitive Impairment at Risk of Developing Dementia using Speech AnalysisabstractMild Cognitive Impairment (MCI) is the intermediate stage between ageing and possible dementia. $50 \%$ of people with MCI may progress to dementia (prodromal Alzheimer’s Disease (pr-AD)). Identifying those at risk is a challenging but important task that can help with anxiety and treatment. Currently, clinicians wait for severe signs of impairment to emerge, however, recent studies have shown promising automatic speech-based approaches. This paper works on a unique dataset containing 50 healthy controls (HCs) and 50 MCI (of which 22 are pr-AD and 28 are non-progressed (nonP)). The recordings are of people speaking with a virtual agent. A number of acoustic, text and contextual features were extracted and used for training classifiers. The best result was achieved with an $F_{1}$-scores of $\mathbf{81.2 \%}$ for the detection of MCI versus HC, and $\mathbf{75} \%$ for MCI(pr-AD) versus MCI(nonP). The best $F_{1}$-score of $\mathbf{66.9 \%}$ was achieved for the detection of MCI (pr-AD) versus MCI (nonP) versus HC. Bahman Mirheidari, Ronan O'Malley, Daniel Blackburn, Heidi Christensen |
ASRU | 1 |
| 2023 | Investigating Visual Features for Cognitive Impairment Detection Using In-the-wild DataabstractEarly detection of dementia has attracted much research interest due to its crucial role in helping people get suitable treatment or care. Video analysis may provide an effective approach for detection, with low cost and effort compared to current expensive and intensive clinical assessments. This paper investigates the use of a range of visual features - eye blink rate (EBR), head turn rate (HTR) and head movement statistical features (HMSF) - for identifying neurodegenerative disorder (ND), mild cognitive impairment (MCI) and functional memory disorder (FMD). These features are used in a noval multiple thresholds approach, which is applied to an in-the-wild video dataset which includes data recorded in a range of challenging environments. A combination of EBR and HTR gives 78 % accuracy in a three-way classification task (ND/MCI/FMD) and 83%, 83% and 92%, respectively, for the two-way classifications ND/MCI, ND/FMD and MCI/FMD. These results are comparable to related work that uses more features from different modalities. They also provide evidence to support the possibility of an in-the-home detection process for dementia or cognitive impairment. Fatimah Alzahrani, Bahman Mirheidari, Daniel Blackburn, Steve C. Maddock, Heidi Christensen |
FG | 2 |
| 2022 | Automatic cognitive assessment: Combining sparse datasets with disparate cognitive scores
Bahman Mirheidari, Daniel Blackburn, Heidi Christensen |
INTERSPEECH | 1 |
| 2022 | Automatic Detection of Expressed Emotion from Five-Minute Speech Samples: Challenges and OpportunitiesabstractWe present a novel feasibility study on the automatic recognition of Expressed Emotion (EE), a family environment concept based on caregivers speaking freely about their relative/family member.We describe an automated approach for determining the degree of warmth, a key component of EE, from acoustic and text features acquired from a sample of 37 recorded interviews.These recordings, collected over 20 years ago, are derived from a nationally representative birth cohort of 2,232 British twin children and were manually coded for EE.We outline the core steps of extracting usable information from recordings with highly variable audio quality and assess the efficacy of four machine learning approaches trained with different combinations of acoustic and text features.Despite the challenges of working with this legacy data, we demonstrated that the degree of warmth can be predicted with an F1-score of 61.5%.In this paper, we summarise our learning and provide recommendations for future work using real-world speech samples. Bahman Mirheidari, André Bittar, Nicholas Cummins, Johnny Downs, Helen L. Fisher, Heidi Christensen |
INTERSPEECH | 1 |
| 2021 | Eye Blink Rate Based Detection of Cognitive Impairment Using In-the-wild DataabstractInvestigating automatic methods for the early detection of dementia and related conditions that cause cognitive impairment is an area of growing interest. Video processing could play a role by providing a non-invasive and low-cost alternative to current expensive assessments. For this to be successful it is crucial that approaches are robust to in-the-wild challenges. In this paper, visual cues, related to the eye blink rate (EBR), are investigated to quantify the early phase of neurodegenerative disorder (ND) and mild cognitive impairment (MCI) as well as functional memory disorder (FMD; problems with memory not related to neurodegenerative disorder). This paper aims to improve the detection of ND and MCI by investigating a novel approach to calculating the EBR that is more robust to in-the-wild challenges. An in-house dataset with 18 participants is used. The EBR is calculated from eye landmarks extracted using two libraries (Dlib and Openface). To mitigate issues observed in the noisy, in-the-wild recordings, a multiple threshold approach for EBR detection is proposed. It involves generating multiple thresholds for identifying a blink, where a threshold is used to determine whether an eye is open or closed. Several supervised machine learning approaches are used for automatic classification. The results show that accuracy measures of 89% and 78% are achieved using Dlib and OpenFace data, respectively, when distinguishing between three conditions with ND, MCI and FMD. Fatimah Alzahrani, Bahman Mirheidari, Daniel Blackburn, Steve C. Maddock, Heidi Christensen |
ACII | 2 |
| 2021 | Identifying Cognitive Impairment Using Sentence Representation VectorsabstractThe widely used word vectors can be extended at the sentence level to perform a wide range of natural language processing (NLP) tasks.Recently the Bidirectional Encoder Representations from Transformers (BERT) language representation achieved state-of-the-art performance for these applications.The model is trained with punctuated and well-formed (writ-ten) text, however, the performance of the model drops significantly when the input text is the -erroneous and unpunctuated-output of automatic speech recognition (ASR).We use a sliding window and averaging approach for pre-processing text for BERT to extract features for classifying three diagnostic categories relating to cognitive impairment: neurodegenerative dis-order (ND), mild cognitive impairment (MCI), and healthy controls (HC).The in-house dataset contains the audio recordings of an intelligent virtual agent (IVA) who asks the participants several conversational questions prompts in addition to giving a picture description prompt.For the three-way classification, we achieve a 73.88% F-score (accuracy: 76.53%) using the pre-trained, uncased base BERT and for the two-way classifier (HCvs.ND) we achieve 89.80% (accuracy: 90%).We further improve these by using a prompt selection technique, reaching the F-scores of 79.98% (accuracy: 81.63%) and 93.56% (accuracy:93.75%)respectively. Bahman Mirheidari, Yilin Pan, Daniel Blackburn, Ronan O'Malley, Heidi Christensen |
Interspeech | 1 |
| 2021 | Using the Outputs of Different Automatic Speech Recognition Paradigms for Acoustic- and BERT-Based Alzheimer's Dementia Detection Through Spontaneous Speech
Yilin Pan, Bahman Mirheidari, Jennifer M. Harris, Jennifer C. Thompson, Julie S. Snowden, Daniel Blackburn, Heidi Christensen |
Interspeech | 2 |
| 2020 | Improving Cognitive Impairment Classification by Generative Neural Network-Based Feature Augmentation
Bahman Mirheidari, Daniel Blackburn, Ronan O'Malley, Annalena Venneri, Traci Walker, Markus Reuber, Heidi Christensen |
INTERSPEECH | 1 |
| 2020 | Improving Detection of Alzheimer's Disease Using Automatic Speech Recognition to Identify High-Quality Segments for More Robust Feature ExtractionabstractSpeech and language based automatic dementia detection is of interest due to it being non-invasive, low-cost and potentially able to aid diagnosis accuracy. The collected data are mostly audio recordings of spoken language and these can be used directly for acoustic-based analysis. To extract linguistic-based information, an automatic speech recognition (ASR) system is used to generate transcriptions. However, the extraction of reliable acoustic features is difficult when the acoustic quality of the data is poor as is the case with DementiaBank, the largest opensource dataset for Alzheimer’s Disease classification. In this paper, we explore how to improve the robustness of the acoustic feature extraction by using time alignment information and confidence scores from the ASR system to identify audio segments of good quality. In addition, we design rhythm-inspired features and combine them with acoustic features. By classifying the combined features with a bidirectional-LSTM attention network, the F-measure improves from 62.15% to 70.75% when only the high-quality segments are used. Finally, we apply the same approach to our previously proposed hierarchical-based network using linguistic-based features and show improvement from 74.37% to 77.25%. By combining the acoustic and linguistic systems, a state-of-the-art 78.34% F-measure is achieved on the DementiaBank task. Yilin Pan, Bahman Mirheidari, Markus Reuber, Annalena Venneri, Daniel Blackburn, Heidi Christensen |
INTERSPEECH | 2 |
| 2020 | Acoustic Feature Extraction with Interpretable Deep Neural Network for Neurodegenerative Related Disorder ClassificationabstractSpeech-based automatic approaches for detecting neurodegenerative disorders (ND) and mild cognitive impairment (MCI) have received more attention recently due to being non-invasive and potentially more sensitive than current pen-and-paper tests. The performance of such systems is highly dependent on the choice of features in the classification pipeline. In particular for acoustic features, arriving at a consensus for a best feature set has proven challenging. This paper explores using deep neural network for extracting features directly from the speech signal as a solution to this. Compared with hand-crafted features, more information is present in the raw waveform, but the feature extraction process becomes more complex and less interpretable which is often undesirable in medical domains. Using a SincNet as a first layer allows for some analysis of learned features. We propose and evaluate the Sinc-CLA (with SincNet, Convolutional, Long Short-Term Memory and Attention layers) as a task-driven acoustic feature extractor for classifying MCI, ND and healthy controls (HC). Experiments are carried out on an in-house dataset. Compared with the popular hand-crafted feature sets, the learned task-driven features achieve a superior classification accuracy. The filters of the SincNet is inspected and acoustic differences between HC, MCI and ND are found. Yilin Pan, Bahman Mirheidari, Zehai Tu, Ronan O'Malley, Traci Walker, Annalena Venneri, Markus Reuber, Daniel Blackburn, Heidi Christensen |
INTERSPEECH | 2 |
| 2019 | Computational Cognitive Assessment: Investigating the Use of an Intelligent Virtual Agent for the Detection of Early Signs of DementiaabstractThe ageing population has caused a marked increased in the number of people with cognitive decline linked with dementia. Thus, current diagnostic services are overstretched, and there is an urgent need for automating parts of the assessment process. In previous work, we demonstrated how a stratification tool built around an Intelligent Virtual Agent (IVA) eliciting a conversation by asking memory-probing questions, was able to accurately distinguish between people with a neuro-degenerative disorder (ND) and a functional memory disorder (FMD). In this paper, we extend the number of diagnostic classes to include healthy elderly controls (HCs) as well as people with mild cognitive impairment (MCI). We also investigate whether the IVA may be used for administering more standard cognitive tests, like the verbal fluency tests. A four-way classifier trained on an extended feature set achieved 48% accuracy, which improved to 62% by using just the 22 most significant features (ROC-AUC: 82%). Bahman Mirheidari, Daniel Blackburn, Ronan O'Malley, Traci Walker, Annalena Venneri, Markus Reuber, Heidi Christensen |
ICASSP | 1 |
| 2019 | Automatic Hierarchical Attention Neural Network for Detecting ADabstractPicture description tasks are used for the detection of cognitive decline associated with Alzheimer's disease (AD). Recent years have seen work on automatic AD detection in picture descriptions based on acoustic and word-based analysis of the speech. These methods have shown some success but lack an ability to capture any higher-level effects of cognitive decline on the patient's language. In this paper, we propose a novel model that encompasses both the hierarchical and sequential structure of the description and detect its informative units by attention mechanism. Automatic speech recognition (ASR) and punctuation restoration are used to transcribe and segment the data. Using the DementiaBank database of people with AD as well as healthy controls (HC), we obtain an F-score of 84.43% and74.37% when using manual and automatic transcripts respectively. We further explore the effect of adding additional data (a total of 33 descriptions collected using a‘digital doctor’) during model training and increase the F-score when using ASR transcripts to 76.09%. This outperforms baseline models, including bidirectional LSTM and bidirectional hierarchical neural net-work without an attention mechanism, and demonstrate that the use of hierarchical models with attention mechanism improves the AD/HC discrimination performance. Yilin Pan, Bahman Mirheidari, Markus Reuber, Annalena Venneri, Daniel Blackburn, Heidi Christensen |
INTERSPEECH | 2 |
| 2019 | Dementia detection using automatic analysis of conversations
Bahman Mirheidari, Daniel Blackburn, Traci Walker, Markus Reuber, Heidi Christensen |
Comput. Speech Lang. | 1 |
| 2018 | Detecting Signs of Dementia Using Word Vector Representations
Bahman Mirheidari, Daniel Blackburn, Traci Walker, Annalena Venneri, Markus Reuber, Heidi Christensen |
INTERSPEECH | 1 |
| 2017 | An Avatar-Based System for Identifying Individuals Likely to Develop DementiaabstractThis paper presents work on developing an automatic dementia screening test based on patients’ ability to interact and communicate — a highly cognitively demanding process where early signs of dementia can often be detected. Such a test would help general practitioners, with no specialist knowledge, make better diagnostic decisions as current tests lack specificity and sensitivity. We investigate the feasibility of basing the test on conversations between a ‘talking head’ (avatar) and a patient and we present a system for analysing such conversations for signs of dementia in the patient’s speech and language. Previously we proposed a semi-automatic system that transcribed conversations between patients and neurologists and extracted conversation analysis style features in order to differentiate between patients with progressive neurodegenerative dementia (ND) and functional memory disorders (FMD). Determining who talks when in the conversations was performed manually. In this study, we investigate a fully automatic system including speaker diarisation, and the use of additional acoustic and lexical features. Initial results from a pilot study are presented which shows that the avatar conversations can successfully classify ND/FMD with around 91% accuracy, which is in line with previous results for conversations that were led by a neurologist. \n Bahman Mirheidari, Daniel Blackburn, Kirsty Harkness, Traci Walker, Annalena Venneri, Markus Reuber, Heidi Christensen |
INTERSPEECH | 1 |
| 2016 | Diagnosing People with Dementia Using Automatic Conversation AnalysisabstractA recent study using Conversation Analysis (CA) has demonstrated that communication problems may be picked up during conversations between patients and neurologists, and that this can be used to differentiate between patients with (progressive neurodegenerative dementia) ND and those with (nonprogressive) functional memory disorders (FMD). This paper presents a novel automatic method for transcribing such conversations and extracting CA-style features. A range of acoustic, syntactic, semantic and visual features were automatically extracted and used to train a set of classifiers. In a proof-of-principle style study, using data recording during real neurologist-patient consultations, we demonstrate that automatically extracting CA-style features gives a classification accuracy of 95%when using verbatim transcripts. Replacing those transcripts with automatic speech recognition transcripts, we obtain a classification accuracy of 79% which improves to 90% when feature selection is applied. This is a first and encouraging step towards replacing inaccurate, potentially stressful cognitive tests with a test based on monitoring conversation capabilities that could be conducted in e.g. the privacy of the patient’s own home. \n \n Bahman Mirheidari, Daniel Blackburn, Markus Reuber, Traci Walker, Heidi Christensen |
INTERSPEECH | 1 |