EDBT 2026 Demo / reviewers in the wild / expert
Fasih Haider
dblp:170/4631
· DBLP profile ↗
25ranked-venue papers
16as first author
10since 2021 · last 2025
0000-0002-5150-3359ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 12 first-author · 7 since 2021Artificial intelligence and machine learning · 13 · 8 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Investigating the Relationship between Signs of Depression and Alzheimer's Dementia in SpeechabstractDepression and Alzheimer's dementia frequently co-occur, with depression often occurring at early stages of Alzheimer's disease (AD). We investigate whether knowledge of a person's depression score gained through analysis of acoustic features of the person's speech can help detect AD. We analysed data from 239 participants, where$n=162$individuals had a clinical diagnosis of AD, to explore correlations between depression scores, cognitive test scores and AD status. First, we created multilayer perceptron models for depression assessment using speech features extracted through our active data representation method and wav2vec features. We then used outputs from the depression models as predictors in our Alzheimer's detection models. Depression scores were strongly correlated with Alzheimer's diagnosis in this data set (Spearman's correlation$r = -0.610$and$r = -0.536,\ p < 0.01$for correlations between depression scores and diagnosis and cognitive test scores, respectively). We found that depression scores could predict AD fairly accurately$(Acc=0.84)$, and that depression could be detected in speech with moderate accuracy$(Acc=0.79)$. Using these predictions and acoustic features, our dementia model reached$Acc=0.70$. We also investigated classification of AD from speech in the presence and absence of depression, obtaining the same level of accuracy for this imbalanced prediction task. Fasih Haider, Stina Saunders, Craig Ritchie, Saturnino Luz |
CCNC | 1 |
| 2025 | Automatic recognition of rodent call types using deep supervectorsabstractRats are gregarious rodents who naturally live in diverse social groups and communicate in part through ultrasonic vocalisations (USVs). USVs encode significant information about affective state and play an important role in social behaviour. Monitoring USVs is a non-invasive method of adding richness to data in a variety of experimental paradigms. However, manual analysis of USVs requires a significant amount of human effort. We propose a new method for automatic classification of USVs which could help automate analysis and thus reduce human input. The proposed method introduces a novel approach to USV representation called deep supervectors (DSV), which combines diverse deep embeddings feature sets extracted through our active data representation (ADR) method. The DSV method is evaluated on a multiclass recognition task involving 14 different types of rodent calls. The performance of DSV is compared to that obtained by state-of-the-art deep embeddings (alexNet, googleNet, squeezeNet and resNet). The proposed method achieves an Unweighted Average Recall (UAR) of 32.80% and outperforms both customs (8.07%) and deep embeddings (29.94%). Deep supervectors outperform pre-trained DNN in 6 out of 8 cases and fusion of the top six DSV with the top two deep embeddings improves the UAR to 37.22% in this challenging classification task, showing an improvement of 29% over the majority guess of 8%. By combining related classes, the UAR will reach 51.00%. The context of this application is the automation of the process for life-sciences laboratory technicians and scientists. Fasih Haider, Raven Hickson, Peter Kind, Saturnino Luz |
ICASSP | 1 |
| 2025 | Early Dementia Detection Using Multiple Spontaneous Speech Prompts: The PROCESS ChallengeabstractDementia is associated with various cognitive impairments and typically manifests only after significant progression, making intervention at this stage often ineffective. To address this issue, the Prediction and Recognition of Cognitive Decline through Spontaneous Speech (PROCESS) Signal Processing Grand Challenge invites participants to focus on early-stage dementia detection. We provide a new spontaneous speech corpus for this challenge. This corpus includes answers from three prompts designed by neurologists to better capture the cognition of speakers. Our baseline models achieved an F1-score of 55.0% on the classification task and an RMSE of 2.98 on the regression task. Fuxiang Tao, Bahman Mirheidari, Madhurananda Pahar, Sophie Young, Hend Elghazaly, Fritz Peters, Caitlin H. Illingworth, Dorota Braun, Ronan O'Malley, Simon Bell, Daniel Blackburn, Fasih Haider, Saturnino Luz, Heidi Christensen |
ICASSP | 13 |
| 2025 | Affective and Physiological Responses to Immersive Intangible Cultural Heritage Experiences in Extended RealityabstractAs part of the Intangible Cultural Heritage, Bridging the Past, Present and Future (INT-ACT) project, we investigate how Extended Reality (XR) technologies can meaningfully engage users with cultural content by monitoring their physiological and affective responses. While immersive XR systems offer new ways of exploring heritage, their impact on users’ internal states remains underexplored. In this study, we present a multimodal experimental setup using EmotiBit, a wearable biosensing platform, to monitor real-time physiological signals during cultural XR interaction. Participants were evaluated across three activities representing varying cognitive and sensory loads: immersive interaction with the INT-ACT XR Demonstrator, composing work emails (a low-stimulation control task), and passive movie watching. Our aim is to quantify how cultural XR experiences influence biosignals such as electrodermal activity and heart rate. The findings reveal distinct physiological patterns across conditions, suggesting that biosignal monitoring can inform the design of adaptive XR environments that are responsive to user states. This work contributes to INT-ACT’s broader objective of creating intelligent, inclusive, and emotionally resonant cultural heritage experiences. Fasih Haider, Sofia de la Fuente, Alicia Núñez García, Saturnino Luz |
ICMI | 1 |
| 2024 | Connected Speech-Based Cognitive Assessment in Chinese and English
Saturnino Luz, Sofia de la Fuente, Fasih Haider, Davida Fromm, Brian MacWhinney, Alyssa Lanzi, Ya-Ning Chang, Chia-Ju Chou, Yi-Chien Liu |
INTERSPEECH | 3 |
| 2023 | Multilingual Alzheimer's Dementia Recognition through Spontaneous Speech: A Signal Processing Grand ChallengeabstractThis Signal Processing Grand Challenge (SPGC) targets a difficult automatic prediction problem of societal and medical relevance, namely, the detection of Alzheimer’s Dementia (AD). Participants were invited to employ signal processing and machine learning methods to create predictive models based on spontaneous speech data. The Challenge has been designed to assess the extent to which predictive models built based on speech in one language (English) generalise to another language (Greek). To the best of our knowledge no work has investigated acoustic features of the speech signal in multilingual AD detection. Our baseline system used conventional machine learning algorithms with Active Data Representation of acoustic features, achieving accuracy of 73.91% on AD detection, and 4.95 root mean squared error on cognitive score prediction. Saturnino Luz, Fasih Haider, Davida Fromm, Ioulietta Lazarou, Ioannis Kompatsiaris, Brian MacWhinney |
ICASSP | 2 |
| 2022 | An Automated Mood Diary for Older User's using Ambient Assisted Living Recorded Speech
Fasih Haider, Saturnino Luz |
INTERSPEECH | 1 |
| 2021 | Affect Recognition Through Scalogram and Multi-Resolution Cochleagram Features
Fasih Haider, Saturnino Luz |
Interspeech | 1 |
| 2021 | Detecting Cognitive Decline Using Speech Only: The ADReSSo ChallengeabstractBuilding on the success of the ADReSS Challenge at Interspeech 2020, which attracted the participation of 34 teams from across the world, the ADReSSo Challenge targets three difficult automatic prediction problems of societal and medical relevance, namely: detection of Alzheimer's Dementia, inference of cognitive testing scores, and prediction of cognitive decline. This paper presents these prediction tasks in detail, describes the datasets used, and reports the results of the baseline classification and regression models we developed for each task. A combination of acoustic and linguistic features extracted directly from audio recordings, without human intervention, yielded a baseline accuracy of 78.87% for the AD classification task, an MMSE prediction root mean squared (RMSE) error of 5.28, and 68.75% accuracy for the cognitive decline prediction task. Saturnino Luz, Fasih Haider, Sofia de la Fuente, Davida Fromm, Brian MacWhinney |
Interspeech | 2 |
| 2021 | Emotion recognition in low-resource settings: An evaluation of automatic feature selection methods
Fasih Haider, Senja Pollak, Pierre Albert, Saturnino Luz |
Comput. Speech Lang. | 1 |
| 2020 | Automatic Recognition of Low-Back Chronic Pain Level and Protective Movement Behaviour using Physical and Muscle Activity InformationabstractAutomatic recognition of low-back chronic pain and movement behaviour in humans could be a useful technology in health monitoring and providing effective rehabilitation advice. Physical and muscle activity information can be used in automating this process in combination with machine learning and feature engineering methods. This paper presents a method for automatic recognition of chronic pain and movement behaviour using our recently proposed `Active Data Representation' (ADR) method, and applies it to two tasks of the EmoPain 2020 Challenge using physical and muscle activity features. The ADR method is used for the transformation of the physical and muscle activity features for the classification tasks. Our results show that ADR outperforms the LSTM challenge baseline model in terms of Matthews correlation coefficient (0.43) and F score (61.21) for the recognition of chronic pain and movement behaviour respectively in hold-out validation settings. Although a decrease in performance is observed on the test dataset, ADR still outperforms the challenge baseline for the recognition of chronic pain and movement behaviour tasks. Fasih Haider, Pierre Albert, Saturnino Luz |
FG | 1 |
| 2020 | Alzheimer's Dementia Recognition Through Spontaneous Speech: The ADReSS ChallengeabstractThe ADReSS Challenge at INTERSPEECH 2020 defines a shared task through which different approaches to the automated recognition of Alzheimer's dementia based on spontaneous speech can be compared. ADReSS provides researchers with a benchmark speech dataset which has been acoustically pre-processed and balanced in terms of age and gender, defining two cognitive assessment tasks, namely: the Alzheimer's speech classification task and the neuropsychological score regression task. In the Alzheimer's speech classification task, ADReSS challenge participants create models for classifying speech as dementia or healthy control speech. In the the neuropsychological score regression task, participants create models to predict mini-mental state examination scores. This paper describes the ADReSS Challenge in detail and presents a baseline for both tasks, including feature extraction procedures and results for classification and regression models. ADReSS aims to provide the speech and language Alzheimer's research community with a platform for comprehensive methodological comparisons. This will hopefully contribute to addressing the lack of standardisation that currently affects the field and shed light on avenues for future research and clinical applicability. Saturnino Luz, Fasih Haider, Sofia de la Fuente, Davida Fromm, Brian MacWhinney |
INTERSPEECH | 2 |
| 2019 | Attitude Recognition Using Multi-resolution Cochleagram FeaturesabstractAttitudes play an important role in human communication. Models and algorithms for automatic recognition of attitudes therefore may have applications in areas where successful communication and interaction are crucial, such as health-care, education and digital entertainment. This paper focuses on the task of categorizing speaker attitudes using speech features. Data extracted from video recordings are employed in training and testing of predictive models consisting of different sets of speech features. A novel attitude recognition approach using Multi-Resolution Cochleagram (MRCG) features is proposed. The results show that MRCG feature set outperforms the feature sets most commonly used in computational paralinguistic tasks, including emobase, eGeMAPS and ComParE, in terms of attitude recognition accuracy for decision tree, 1-nearest neighbour and random forest classifiers. Analysis of the results suggests that MRCG features contribute information not captured by these existing feature sets. Indeed, while the ComParE feature set provides slightly better results than MRCG features for support vector machine classifiers, the fusion of the existing feature sets with the new MRCG features improves on those results. Overall, with the addition of MRCG, the attitude recognition method proposed in this study achieves accuracy scores approximately 11 points higher than reported in previous studies. Fasih Haider, Saturnino Luz |
ICASSP | 1 |
| 2019 | A Searching and Automatic Video Tagging Tool for Events of Interest during Volleyball Training SessionsabstractQuick and easy access to performance data during matches and training sessions is important for both players and coaches. While there are many video tagging systems available, these systems require manual effort. This paper proposes a system architecture that automatically supplements video recording by detecting events of interests in volleyball matches and training sessions to provide tailored and interactive multi-modal feedback. Fahim A. Salim, Fasih Haider, Sena Busra Yengec Tasdemir, Vahid Naghashi, Izem Tengiz, Kübra Cengiz, Dees B. W. Postma, Robby van Delden, Dennis Reidsma, Saturnino Luz, Bert-Jan van Beijnum |
ICMI | 2 |
| 2019 | A System for Real-Time Privacy Preserving Data Collection for Ambient Assisted Living
Fasih Haider, Saturnino Luz |
INTERSPEECH | 1 |
| 2019 | Analysing patterns of right brain-hemisphere activity prior to speech articulation for identification of system-directed speech
Fasih Haider, Hayakawa Akira, Carl Vogel, Nick Campbell 0001, Saturnino Luz |
Speech Commun. | 1 |
| 2018 | On-Talk and Off-Talk Detection: A Discrete Wavelet Transform Analysis of ElectroencephalogramabstractSpoken interaction with a machine results in a behaviour that is not very common in face-to-face human communication:Off-Talk, which is defined as speech utterances that are not directed to an immediate interlocutor, the machine, but to another person or even oneself. It is our contention that a system which is able to detect theOff-Talkutterances can interact with a human in a more efficient manner by acknowledging that the utterances are not directed to the system and hence, not replying toOff-Talkutterances. In this paper, we demonstrate the discrimination power of a wide range of Electroencephalogram (EEG) frequency bands using wavelet transform analysis and propose models forOn-TalkandOff-Talkdetection using audio and EEG signals, and their fusion. Our study shows that the EEG signal can identify the occurrence ofOff-Talkutterances with promising accuracy and its fusion with audio features adds a slight improvement in these results. Fasih Haider, Hayakawa Akira, Saturnino Luz, Carl Vogel, Nick Campbell 0001 |
ICASSP | 1 |
| 2018 | SAAMEAT: Active Feature Transformation and Selection Methods for the Recognition of User Eating ConditionsabstractAutomatic recognition of eating conditions of humans could be a useful technology in health monitoring. The audio-visual information can be used in automating this process, and feature engineering approaches can reduce the dimensionality of audio-visual information. The reduced dimensionality of data (particularly feature subset selection) can assist in designing a system for eating conditions recognition with lower power, cost, memory and computation resources than a system which is designed using full dimensions of data. This paper presents Active Feature Transformation (AFT) and Active Feature Selection (AFS) methods, and applies them to all three tasks of the ICMI 2018 EAT Challenge for recognition of user eating conditions using audio and visual features. The AFT method is used for the transformation of the Mel-frequency Cepstral Coefficient and ComParE features for the classification task, while the AFS method helps in selecting a feature subset. Transformation by Principal Component Analysis (PCA) is also used for comparison. We find feature subsets of audio features using the AFS method (422 for Food Type, 104 for Likability and 68 for Difficulty out of 988 features) which provide better results than the full feature set. Our results show that AFS outperforms PCA and AFT in terms of accuracy for the recognition of user eating conditions using audio features. The AFT of visual features (facial landmarks) provides less accurate results than the AFS and AFT sets of audio features. However, the weighted score fusion of all the feature set improves the results. Fasih Haider, Senja Pollak, Eleni Zarogianni, Saturnino Luz |
ICMI | 1 |
| 2018 | Improving Response Time of Active Speaker Detection Using Visual Prosody Information Prior to Articulation
Fasih Haider, Saturnino Luz, Carl Vogel, Nick Campbell 0001 |
INTERSPEECH | 1 |
| 2018 | An Active Feature Transformation Method for Attitude Recognition of Video Bloggers
Fasih Haider, Fahim A. Salim, Owen Conlan, Saturnino Luz |
INTERSPEECH | 1 |
| 2018 | The Metalogue Debate Trainee Corpus: Data Collection and Annotations
Volha Petukhova, Andrei Malchanau, Youssef Oualil, Dietrich Klakow, Saturnino Luz, Fasih Haider, Nick Campbell 0001, Dimitris Koryzis, Dimitris Spiliotopoulos, Pierre Albert, Nicklas Linz, Jan Alexandersson |
LREC | 6 |
| 2017 | Visual, Laughter, Applause and Spoken Expression Features for Predicting Engagement Within TED Talks
Fasih Haider, Fahim A. Salim, Saturnino Luz, Carl Vogel, Owen Conlan, Nick Campbell 0001 |
INTERSPEECH | 1 |
| 2016 | Presentation quality assessment using acoustic information and hand movementsabstractThis study focuses on prosodic and gestural features that contribute to the positive judgement of public oral presentations. The general hypothesis is that certain prosodic characteristics, such as high pitch variation and perceived loudness, together with the production of natural hand gestures, influence the audience's perception of the speaker as a good presenter. Being able to identify features that can give an indication of a good presenter is useful for applications in the field of skills training, where automatic feedback could be provided to trainees at the end of their presentation about the extent to which they have been able to use their voices and gestures to keep the audience engaged. For this reason, we also propose a method, based on prosodic and visual features, able to categorise presentation quality with high accuracy. Fasih Haider, Loredana Cerrato, Nick Campbell 0001, Saturnino Luz |
ICASSP | 1 |
| 2015 | Analyzing Multimodality of Video for User Engagement AssessmentabstractThese days, several hours of new video content is uploaded to the internet every second. It is simply impossible for anyone to see every piece of video which could be engaging or even useful to them. Therefore it is desirable to identify videos that might be regarded as engaging automatically, for a variety of applications such as recommendation and personalized video segmentation etc. This paper explores how multimodal characteristics of video, such as prosodic, visual and paralinguistic features, can help in assessing user engagement with videos. The approach proposed in this paper achieved good accuracy (maximum F score of 96.93 %) through a novel combination of features extracted directly from video recordings, demonstrating the potential of this method in identifying engaging content. Fahim A. Salim, Fasih Haider, Owen Conlan, Saturnino Luz, Nick Campbell 0001 |
ICMI | 2 |
| 2015 | Detection of cognitive states and their correlation to speech recognition performance in speech-to-speech machine translation systems
Hayakawa Akira, Fasih Haider, Loredana Cerrato, Nick Campbell 0001, Saturnino Luz |
INTERSPEECH | 2 |