Yongfeng Tao

dblp:358/5353 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
10since 2021 · last 2026
0009-0009-9712-9380ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 EmoRAct: A neuro-symbolic framework coupling acoustic tokens with prosody semantics for emotion recognition
Minqiang Yang, Mingwen Zhang, Yongfeng Tao, Changsheng Ma, Bin Hu 0001
Pattern Recognit.4
2026 Multi-Scale Temporal-Frequency Attention Network Based on Ocular Imaging for Depression Detection
abstract
Depression is a common and serious mental disorder, characterized by persistent low mood, loss of interest, cognitive dysfunction, and physiological changes. Patients may experience symptoms such as sleep disturbances, changes in appetite, fatigue, and low self-esteem, with severe cases potentially leading to suicidal behavior. There are differences in emotional processing and attention allocation between patients with depression and healthy controls, eye movement characteristics such as fixation patterns, saccade amplitude, and attentional bias have been used as physiological signals for depression detection. Many researchers have developed depression recognition models based on ocular imaging. However, convolutional neural networks, which utilize local receptive fields, can only capture local features in ocular imaging. This paper proposes Multi-Scale Temporal-Frequency Attention Network (MTFNet), which innovatively integrates Multi-Scale time-frequency domain attention into the Video Swin Transformer. Through Multi-Scale Temporal-Frequency Attention Module (MTFAM), MTFNet learns the most important regions in eye movement images, enabling it to capture features more effectively from sequential data and gain a deeper understanding of the structure within eye movement images. Experimental results show that the proposed method achieves a high accuracy of 76.8% on a self-collected eye movement image dataset, outperforming most models. This work provides a novel approach to research on depression recognition based on eye movement images.
Ziru Weng, Zilin Guo, Weihao Zheng, Yongfeng Tao, Bin Hu 0001, Minqiang Yang
IEEE J. Biomed. Health Informatics5
2025 Performance Evaluation and Analysis of LLMs for Detecting Dangerous Speech in Emotional Disorder Data
abstract
Affective computing, as a core technology for intelligent services, is rapidly transitioning from the laboratory to real-world, cloud-driven applications, particularly in critical areas such as mental health. Research utilizing large language models (LLMs) to study affective disorders has a long history, with numerous AI-based diagnostic approaches proposed. However, few studies have evaluated LLMs' ability to identify dangerous discourse in natural conversational settings. To address this gap, this paper conducts the first systematic evaluation of five cutting-edge affective LLMs-GPT-5, Gemini, DeepSeek, Kimi, and Doubao-in identifying depression- and anxiety-related dangerous statements within general affective contexts. Testing was performed using a meticulously curated real clinical dialogue dataset. We designed a comprehensive evaluation framework integrating human annotation, multi-scenario sample data, and multidimensional metrics to measure model performance and generalization capabilities. Experimental results reveal significant performance disparities among models, with accuracy ranging from 0.66 to 0.92. DeepSeek achieves optimal generalization with 0.92 accuracy and 0.98 recall, while Gemini exhibits notable limitations with the lowest recall of 0.59. Analysis of the results reveals that current LLM deployment research faces three core challenges: the persistent precision-recall trade-off dilemma; limited contextual understanding capabilities; and inadequate risk stratification beyond binary classification. Based on these findings, we propose two practical recommendations: clinicians should adopt human-supervised hybrid approaches; developers must enhance safety alignment mechanisms, improve contextual reasoning capabilities, and build explainable affective LLMs to achieve generalization across diverse scenarios.
Shuojiayin Yang, Yongfeng Tao
CloudCom3
2025 Spike Memory Transformer: An Energy-Efficient Model in Distributed Learning Framework for Autonomous Depression Detection
abstract
Objective depression assessments based on physiological data offer new opportunities for clinical diagnostic support, but their usage is constrained by the limitations of data collection devices. Ubiquitous digital devices, now deeply integrated into everyday life, have the potential to effectively capture and represent digital phenotypes of depression. However, existing research in this field faces fundamental problems, such as device thresholds and insufficient utilization of distributed computational power, which have hindered significant research and application efforts in this area. In this article, we proposed the Computation-Oriented Hierarchical Depression Detection Internet of Things (IoT) Framework, which allows IoT devices to collaborate in a layered and distributed manner for psychological data collection and depression detection. In addition, we developed a depression detection model, i.e., spike memory transformer (SMT), which significantly reduces inference energy consumption, facilitating the deployment of depression detection capabilities across various IoT terminal devices. Experimental results demonstrate that our model achieved up to 70% accuracy on the D-Vlog dataset while reducing inference power consumption by an average of 38% compared to classical deep learning methods, thus validating the feasibility of proposed method.
Minqiang Yang, Yueze Liu, Yongfeng Tao, Bin Hu 0001
IEEE Internet Things J.3
2025 Multimodal Depression Detection Based on Self-Attention Network With Facial Expression and Pupil
abstract
Depression is a major mental health issue in contemporary society, with an estimated 350 million people affected globally. The number of individuals diagnosed with depression continues to rise each year. Currently, clinical practice relies entirely on self-reporting and clinical assessment, which carries the risk of subjective biases. In this article, we propose a multimodal method based on facial expression and pupil to detect depression more objectively and precisely. Our method first extracts the features of facial expressions and pupil diameter using residual networks and 1-D convolutional neural networks. Second, a cross-modal fusion model based on self-attention networks (CMF-SNs) is proposed, which utilizes cross-modal attention networks within modalities and parallel self-attention networks between different modalities to extract CMF features of facial expressions and pupil diameter, effectively complementing information between different modalities. Finally, the obtained features are fully connected to identify depression. Multiple controlled experiments show that compared to single modality, the multimodal fusion method based on self-attention networks has a higher ability to recognize depression, with the highest accuracy of 75.0%. In addition, we conducted comparative experiments under three different stimulation paradigms, and the results showed that the classification accuracy under negative and neutral stimuli was higher than that under positive stimuli, indicating a bias of depressed patients toward negative images. The experimental results demonstrate the superiority of our multimodal fusion method.
Xiang Liu 0020, Hao Shen 0017, Huiru Li, Yongfeng Tao, Minqiang Yang
IEEE Trans. Comput. Soc. Syst.4
2025 Trial Selection Tensor Canonical Correlation Analysis (TSTCCA) for Depression Recognition With Facial Expression and Pupil Diameter
abstract
Facial expressions have been widely used for depression recognition because it is intuitive and convenient to access. Pupil diameter contains rich emotional information that is already reflected in facial video streams. However, the spatiotemporal correlation between pupillary changes and facial behavior changes induced by emotional stimuli has not been explored in existing studies. This paper presents a novel multimodal fusion algorithm - Trial Selection Tensor Canonical Correlation Analysis (TSTCCA) to optimize the feature space and build a more robust depression recognition model, which innovatively combines the spatiotemporal relevance and complementarity between facial expression and pupil diameter features. TSTCCA explores the interaction between trials and obtains an effective fusion representation of two modalities from a trial subset related to depression. The experimental results show that TSTCCA achieves the highest accuracy of 78.81% with the subset of 25 trials.
Minqiang Yang, Yushan Wu, Yongfeng Tao, Xiping Hu, Bin Hu 0001
IEEE J. Biomed. Health Informatics3
2024 Are Large Language Models Possible to Conduct Cognitive Behavioral Therapy?
abstract
In contemporary society, the issue of psychological health has become increasingly prominent, characterized by the diversification, complexity, and universality of mental disorders. Cognitive Behavioral Therapy (CBT), currently the most influential and clinically effective psychological treatment method with no side effects, has limited coverage and poor quality in most countries. In recent years, researches on the recognition and intervention of emotional disorders using large language models (LLMs) have been validated, providing new possibilities for psychological assistance therapy. However, are large language models truly possible to conduct cognitive behavioral therapy? Many concerns have been raised by mental health experts regarding the use of LLMs for therapy. Seeking to answer this question, we collected real CBT corpus from online video websites, designed and conducted a targeted automatic evaluation framework involving three aspects, namely the evaluation of emotion tendency of generated text, structured dialogue pattern and proactive inquiry ability. Considering limited CBT-related texts in a general chat LLM’s training corpus, we evaluated the CBT ability of the LLM after integrating a CBT knowledge base to explore the influence of introducing additional knowledge. Four LLM variants with exceptional performance are evaluated, and the experimental result shows the great potential of LLMs in psychological counseling realm, especially after combining with other technological means.
Hao Shen 0017, Minqiang Yang, Minghui Ni, Yongfeng Tao, Weihao Zheng, Bin Hu 0001
BIBM5
2024 Behavioral Information Feedback With Large Language Models for Mental Disorders: Perspectives and Insights
abstract
This edition of the publication includes a robust collection of 104 regular papers and features a Special Issue on Knowledge- Infused Learning for Computational Social Systems. This special issue delves into the sophisticated integration of advanced technologies and knowledge-based methodologies within the analysis of computational social systems. Spanning 12 articles, the issue addresses a wide spectrum of topics, from big data management to refining machine learning models with domain-specific insights. It encompasses areas such as energy management in sensor networks, acoustic analysis of heartbeats, detection of fraudulent activities in online ratings, and the management of rumors on social networks, exemplifying the significant role that knowledge-infused learning plays in enhancing technological applications and fostering innovation in social computational systems.
Minqiang Yang, Yongfeng Tao, Hanshu Cai, Bin Hu 0001
IEEE Trans. Comput. Soc. Syst.2
2024 DepMSTAT: Multimodal Spatio-Temporal Attentional Transformer for Depression Detection
abstract
Depression is one of the most common mental illnesses, but few of the currently proposed in-depth models based on social media data take into account both temporal and spatial information in the data for the detection of depression. In this paper, we present an efficient, low-covariance multimodal integrated spatio-temporal converter framework called DepMSTAT, which aims to detect depression using acoustic and visual features in social media data. The framework consists of four modules: a data preprocessing module, a token generation module, a Spatial-Temporal Attentional Transformer (STAT) module, and a depression classifier module. To efficiently capture spatial and temporal correlations in multimodal social media depression data, a plug-and-play STAT module is proposed. The module is capable of extracting unimodal spatio-temporal features and fusing unimodal information, playing a key role in the analysis of acoustic and visual features in social media data. Through extensive experiments on a depression database (D-Vlog), the method in this paper shows high accuracy (71.53%) in depression detection, achieving a performance that exceeds most models. This work provides a scaffold for studies based on multimodal data that assists in the detection of depression.
Yongfeng Tao, Minqiang Yang, Huiru Li, Yushan Wu, Bin Hu 0001
IEEE Trans. Knowl. Data Eng.1
2023 Classifying Anxiety and Depression through LLMs Virtual Interactions: A Case Study with ChatGPT
abstract
Mental health has long been studied, and many AI-based approaches have been proposed for diagnosis and adjunctive therapy. The emergence of Pre-trained Large Language Models (LLMs) has had a profound impact on various fields, but the potential of using ChatGPT for cognitive behavior therapy is largely unexplored. Therefore, there is an urgent need to build and design a virtual interactive framework for assisted diagnosis/treatment. In this paper, we present a virtual interaction framework based on LLMs that allows participants to engage in a dialogue with a virtual character, analyse mental health issues through augmented LLMs, and make suggestions during the dialogue to alleviate the psychological problems they are currently facing. Based on this framework, we develop a use case for the application of ChatGPT in the field of emotional disorders. Specifically, we use data from question-and-answer dialogues in real-life scenarios to populate the current exploration of ChatGPT’s potential for depression and anxiety detection. The case study shows the great potential of ChatGPT in the analysis of depression and anxiety tests. The feasibility of a virtual interaction framework based on LLMs has been preliminarily demonstrated.
Yongfeng Tao, Minqiang Yang, Hao Shen 0017, Zhichao Yang 0014, Ziru Weng, Bin Hu 0001
BIBM1