Fuxiang Tao

dblp:277/3926 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
8since 2021 · last 2025
0009-0007-5729-9014ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Early Dementia Detection Using Multiple Spontaneous Speech Prompts: The PROCESS Challenge
abstract
Dementia is associated with various cognitive impairments and typically manifests only after significant progression, making intervention at this stage often ineffective. To address this issue, the Prediction and Recognition of Cognitive Decline through Spontaneous Speech (PROCESS) Signal Processing Grand Challenge invites participants to focus on early-stage dementia detection. We provide a new spontaneous speech corpus for this challenge. This corpus includes answers from three prompts designed by neurologists to better capture the cognition of speakers. Our baseline models achieved an F1-score of 55.0% on the classification task and an RMSE of 2.98 on the regression task.
Fuxiang Tao, Bahman Mirheidari, Madhurananda Pahar, Sophie Young, Hend Elghazaly, Fritz Peters, Caitlin H. Illingworth, Dorota Braun, Ronan O'Malley, Simon Bell, Daniel Blackburn, Fasih Haider, Saturnino Luz, Heidi Christensen
ICASSP1
2025 Can Speech Accurately Detect Depression in Patients With Comorbid Dementia? An Approach for Mitigating Confounding Effects of Depression and Dementia
Sophie Young, Fuxiang Tao, Bahman Mirheidari, Madhurananda Pahar, Markus Reuber, Heidi Christensen
INTERSPEECH2
2025 Hire: Hybrid-Modal Interaction with Multiple Relational Enhancements for Image-Text Matching
abstract
Image-Text Matching (ITM) is a fundamental problem in computer vision. The key issue lies in jointly learning the visual and textual representation to estimate their similarity accurately. Most existing methods focus on feature enhancement within modality or feature interaction across modalities, which, however, neglects the contextual information of the object representation based on the inter-object relationships that match the corresponding sentences with rich contextual semantics. In this article, we propose a Hybrid-modal Interaction with multiple Relational Enhancements (termed Hire ) for ITM, which correlates the intra- and inter-modal semantics between objects and words with implicit and explicit relationship modeling. In particular, the explicit intra-modal spatial-semantic graph-based reasoning network is designed to improve the contextual representation of visual objects with salient spatial and semantic relational connectivities, guided by the explicit relationships of the objects’ spatial positions and their scene graph. We use implicit relationship modeling for potential relationship interactions before explicit modeling to improve the fault tolerance of explicit relationship detection. Then the visual and textual semantic representations are refined jointly via inter-modal interactive attention and cross-modal alignment. To correlate the context of objects with the textual context, we further refine the visual semantic representation via cross-level object-sentence and word-image-based interactive attention. Extensive experiments validate that the proposed hybrid-modal interaction with implicit and explicit modeling is more beneficial for ITM. And the proposed Hire obtains new state-of-the-art results on MS-COCO and Flickr30K benchmarks.
Xuri Ge, Fuhai Chen, Songpei Xu, Fuxiang Tao, Jie Wang 0072, Joemon M. Jose
ACM Trans. Intell. Syst. Technol.4
2025 Automatic Detection of Early Cognitive Decline Using Multimodal Feature Fusion and Transfer Learning on Real-World Conversational Speech
abstract
Early signs of cognitive decline, such as dementia and mild cognitive impairment (MCI), often manifest in conversational speech. Early and accurate identification is essential for potential interventions prior to the onset of more severe stages of neurodegenerative diseases. We present CognoMemory, a system for detecting cognitive decline based on a person's speech, to collect 307 hrs of real-world conversational speech, corresponding to 1.92 million Whisper-transcribed words, from 1,639 participants. Speech recordings were collected as participants answered 14 memory-probing, clinically effective questions asked by a virtual agent, starting with a motivation prompt, followed by memory, cognitive functioning, fluency, picture description and reading task. Both acoustic and linguistic features, along with large language model (LLM) embeddings, were extracted from all 1,639 participants. A subset of 614 participants, either with an unconfirmed diagnosis or younger than 50 years, was used for pre-training. The remaining three groups (64 dementia, 169 MCI and 792 healthy participants) were used to fine-tune our proposed model. Our multimodal feature fusion and CNN/Bi-LSTM-based transfer learning approach outperforms LLM-based (BART, DistilBERT, RoBERTa and HuBERT) approaches while achieving the highest $F_{1}$-scores of 0.83 & 0.54 using just the initial 'motivation' question for 2-way & 3-way classification; exhibiting a 3% performance increase due to the application of transfer learning, while being also 38% faster. Finally, the classifiers trained on the CognoMemory data, the largest of its kind, were tested on the second-largest available DementiaBank dataset (Pitt corpus), and a CNN-based transfer learning architecture achieved an $F_{1}$-score of 0.89, demonstrating better stability and generalisation across datasets and of our novel feature fusion and architecture.
Madhurananda Pahar, Bahman Mirheidari, Caitlin H. Illingworth, Dorota Braun, Fuxiang Tao, Lise Sproson, Daniel Blackburn, Heidi Christensen
IEEE J. Biomed. Health Informatics5
2024 Cross-Data Multilevel Attention for Depression Detection: Analyzing the Interplay Between Read and Spontaneous Speech
abstract
This work proposes a novel Cross-Data Multilevel Attention (CDMA) approach for multi-type speech-based depression detection, encompassing both read and spontaneous speech. The main novelty lies in analyzing the unique and common representations of the two types of speech and integrating them into a unified end-to-end framework with novel Intra-Type Multi-Local Attention (IT-MLA) and Cross-Type Global Attention (CT-GA) mechanisms. In particular, IT-MLA highlights depression-relevant information unique in either read or spontaneous speech via intra-modal attention-aware interactions. Furthermore, CT-GA further emphasises the depression-relevant common information in both read and spontaneous speech, with each type being guided by the other. These multiple enhanced representations are aggregated to produce the final predictions. Experiments conducted on a publicly available corpus of 104 speakers (including 52 diagnosed with depression by professional psychiatrists) demonstrate that the proposed CDMA achieves an F1 score of up to 92.5%, the highest performance recorded on this dataset.
Fuxiang Tao, Xuri Ge, Anna Esposito, Alessandro Vinciarelli
BIBM1
2023 Multi-Local Attention for Speech-Based Depression Detection
abstract
This article shows that an attention mechanism, the Multi-Local Attention, can improve a depression detection approach based on Long Short-Term Memory Networks. Besides leading to higher performance metrics (e.g., Accuracy and F1 Score), Multi-Local Attention improves two other aspects of the approach, both important from an application point of view. The first is the effectiveness of a confidence score associated to the detection outcome at identifying speakers more likely to be classified correctly. The second is the amount of speaking time needed to classify a speaker as depressed or non-depressed. The experiments were performed over read speech and involved 109 participants (including 55 diagnosed with depression by professional psychiatrists). The results show accuracies up to 88.0% (F1 Score 88.0%).
Fuxiang Tao, Xuri Ge, Anna Esposito, Alessandro Vinciarelli
ICASSP1
2023 The Androids Corpus: A New Publicly Available Benchmark for Speech Based Depression Detection
Fuxiang Tao, Anna Esposito, Alessandro Vinciarelli
INTERSPEECH1
2023 Cross-modal Semantic Enhanced Interaction for Image-Sentence Retrieval
abstract
Image-sentence retrieval has attracted extensive research attention in multimedia and computer vision due to its promising application. The key issue lies in jointly learning the visual and textual representation to accurately estimate their similarity. To this end, the mainstream schema adopts an object-word based attention to calculate their relevance scores and refine their interactive representations with the attention features, which, however, neglects the context of the object representation on the inter-object relationship that matches the predicates in sentences. In this paper, we propose a Cross-modal Semantic Enhanced Interaction method, termed CMSEI for image-sentence retrieval, which correlates the intra- and inter-modal semantics be-tween objects and words. In particular, we first design the intra-modal spatial and semantic graphs based reasoning to enhance the semantic representations of objects guided by the explicit relationships of the objects’ spatial positions and their scene graph. Then the visual and textual semantic representations are refined jointly via the inter-modal interactive attention and the cross-modal alignment. To correlate the context of objects with the textual context, we further refine the visual semantic representation via the cross-level object-sentence and word-image based interactive attention. Experimental results on seven standard evaluation metrics show that the proposed CMSEI outperforms the state-of-the-art and the alternative approaches on MS-COCO and Flickr30K benchmarks.
Xuri Ge, Fuhai Chen, Songpei Xu, Fuxiang Tao, Joemon M. Jose
WACV4
2020 Spotting the Traces of Depression in Read Speech: An Approach Based on Computational Paralinguistics and Social Signal Processing
Fuxiang Tao, Anna Esposito, Alessandro Vinciarelli
INTERSPEECH1