VLDB 2026 Research / reviewers in the wild / expert
Zhenyu Liu 0006
dblp:74/4038-6
· DBLP profile ↗
24ranked-venue papers
11as first author
13since 2021 · last 2026
0000-0001-8401-9056ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 10 · 6 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 6 since 2021Artificial intelligence and machine learning · 5 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LLM-based semantic integration of stimulus-response pairs for depression detection in interview scenarios
Zhenyu Liu 0006, Jiahang Chen, Bohua Zhao, Zhijie Ding, Bin Hu 0001 |
Expert Syst. Appl. | 1 |
| 2026 | A text-based emotional pattern discrepancy aware model for enhanced generalization in depression detection
Zhenyu Liu 0006, Yang Wu 0011, Jiaqian Yuan, Zhijie Ding, Bin Hu 0001 |
Inf. Process. Manag. | 2 |
| 2026 | A Covariate-Guided Graph Attention Network for Depression Recognition via Pure Facial Movements
Bohua Zhao, Zhenyu Liu 0006, Jiaqian Yuan, Bin Hu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | Sounding Depressed? Personalized Deep Learning Model for Depression Detection From Speech and TextabstractAutomatic depression detection technology based on speech and text data is one of the current research hotspots. However, numerous current models may inadequately incorporate the individual characteristics important for improving detection precision. This leads to misjudgment when the model deals with a depressed patient whose behavior significantly differs from the majority of patients. Thus, the goal of this research is to extract individual characteristics from speech and text data among individuals and develop an efficient and stable personalized model for depression detection. We propose a personalized multimodal depression detection model (PMDDM). The proposed model generates personalized embeddings for each individual by utilizing both speech and text features, while learning depressive representations and integrating them with personalized embeddings for joint representations. Lastly, the model receives the joint representations to predict the individual's depression status. We evaluated the proposed model using two datasets related to depression detection, MODMA and MIDD. Compared to the best baseline methods, our method improved the recognition accuracy by 5.88% and 3.84% on the MODMA and MIDD datasets, respectively. The F1 scores increased by 4.55% and 3.45%, respectively. In the cross-dataset generalizability experiments, our method still outperformed the best baseline methods in F1 score and accuracy. Results suggest that the effective integration of personalized information can significantly improve the accuracy and generalizability of depression detection models. Our research provides novel insights into personalized depression detection. Zhenyu Liu 0006, Jiaqian Yuan, Zhijie Ding, Bin Hu 0001 |
IEEE Trans. Multim. | 2 |
| 2025 | MPDRM: A Multi-Scale Personalized Depression Recognition Model via facial movements
Zhenyu Liu 0006, Bailin Chen, Shimao Zhang, Jiaqian Yuan, Yang Wu 0011, Hanshu Cai, Yimiao Zhao, Huan Mei, Jiahui Deng, Yanping Bao, Bin Hu 0001 |
Neurocomputing | 1 |
| 2025 | Stimulus-Response Pattern: The Core of Robust Cross-Stimulus Facial Depression RecognitionabstractFacial depression recognition is one of the current hot topics. Mainstream methods mainly focus on how to design deep models to effectively extract the difference in facial movements between depressed patients and healthy people. However, this difference changes when the stimulus source to which the subjects are exposed changes. This leads to the performance degradation in cross-stimulus situation and limits the practical application of this technology. We hold the opinion that why depressed patients show behavioral characteristics different from healthy people is that they have a specific stable pattern of responding to stimulus. Therefore, we incorporate stimuli into the modeling process for the first time and employ deep networks to learn stable representations between stimulus and response to achieve stable and effective modeling. Specifically, we propose a deep modeling framework to learn the stimulus-response pattern of the subject through the interaction relationship between the stimulus videos and the subject’s facial movements. We constructed a balanced depression dataset of 364 individuals with three different stimulus videos to verify the effectiveness of our method. The results show that our method achieves state-of-the-art and the best generalization performance in depression recognition. This stimulus-response pattern modeling provides a new perspective for recognizing depression. Zhenyu Liu 0006, Shimao Zhang, Bailin Chen, Qiongqiong Chen, Zhijie Ding, Xin Zhang 0034, Bin Hu 0001 |
IEEE Trans. Affect. Comput. | 1 |
| 2025 | Depression Recognition by Fuzzy Learning of Facial Movements Based on Graph Neural NetworksabstractFacial movement-based depression recognition has garnered considerable attention in automatic depression detection field in recent years. Current methods typically seek to establish a uniform depression discrimination criterion and often overlook inter-individual differences in facial movements, which lead to a decrease in the accuracy of models. A closer examination of the clinical diagnostic criteria for depression reveals that individuals with depression exhibit a wide variety of symptoms. This suggests that even individuals with similar levels of depression may exhibit significantly different facial movements. To address this challenge, this study introduces a novel framework for depression assessment—Multiscale Video Adaptive Graph Network (MVAGN). By leveraging the information transmission capabilities of Graph Neural Network (GNN) to balance the relationship between individual differences and commonalities to achieve more accurate depression identification. MVAGN consists of three modules: the Multiscale Video Feature Extractor (MVFE), the Hybrid Feature Interaction Graph Module (HFIGM), and the Adaptive Hybrid Graph Network (AHGN). MVFE extracts multi-scale features from the video, HFIGM transforms these features into graph representations, and AHGN facilitates information transfer between nodes to reduce individual feature differences, ultimately providing a more precise estimation of depression severity. Experiments conducted on the AVEC2013 and AVEC2014 datasets demonstrate that MVAGN achieves state-of-the-art performance. This paper proposes that building a fuzzy system with a compact kernel and a loose periphery provides a viable path for recognizing depression. Zhenyu Liu 0006, Bohua Zhao, Jiahang Chen, Jiaqian Yuan, Bin Hu 0001 |
IEEE Trans. Fuzzy Syst. | 1 |
| 2024 | PIE: A Personalized Information Embedded model for text-based depression detection
Yang Wu 0011, Zhenyu Liu 0006, Jiaqian Yuan, Bailin Chen, Hanshu Cai, Yimiao Zhao, Huan Mei, Jiahui Deng, Yanping Bao, Bin Hu 0001 |
Inf. Process. Manag. | 2 |
| 2024 | Stimulus-Response Patterns: The Key to Giving Generalizability to Text-Based Depression Detection ModelsabstractText content analysis for depression detection using machine learning techniques has become a prominent area of research. However, previous studies focused mainly on analyzing the textual content, neglecting the fundamental factors driving text generation. Consequently, existing models face the challenge of poor generalization to out-of-domain data as they struggle to capture the crucial features of depression. To address this, we propose a novel computational perspective of "stimulus-response patterns" that brings us closer to the essence of clinical diagnosis of depression. Adopting this computational perspective allows us to conceptually unify diverse datasets and generalize this perspective to common datasets in the field. We introduce the Stimulus-Response Patterns-aware Network (SRP-Net) as an exemplary approach within this computational perspective. To assess the performance of the SRP-Net, we constructed a multi-stimulus dataset and conducted experimental evaluations, demonstrating its exceptional cross-stimulus generalizability. Furthermore, we demonstrated the promising performance of SPR-Net in real medical scenarios and conducted an interpretability analysis of the stimulus-response patterns. Our research investigates the critical role of stimulus-response patterns in enhancing the generalizability of text-based depression detection models, which can potentially facilitate data-driven depression detection to approach the diagnostic accuracy of psychiatrists. Zhenyu Liu 0006, Yang Wu 0011, Zhijie Ding, Bin Hu 0001 |
IEEE J. Biomed. Health Informatics | 1 |
| 2024 | Multi Fine-Grained Fusion Network for Depression DetectionabstractDepression is an illness that involves emotional and mental health. Currently, depression detection through interviews is the most popular way. With the advancement of natural language processing and sentiment analysis, automated interview-based depression detection is strongly supported. However, current multimodal depression detection models fail to adequately capture the fine-grained features of depressive behaviors, making it difficult for the models to accurately characterize the subtle changes in depressive symptoms. To address this problem, we propose a Multi Fine-Grained Fusion Network (MFFNet). The core idea of this model is to extract and fuse the information of different scale feature pairs through a Multi-Scale Fastformer (MSfastformer), and then use the Recurrent Pyramid Model to integrate the features of different resolutions, promoting the interaction of multi-level information. Through the interaction of multi-scale and multi-resolution features, it aims to explore richer feature representations. To validate the effectiveness of our proposed MFFNet model, we conduct experiments on two depression interview datasets. The experimental results show that the MFFNet model performs better in depression detection compared to other benchmark multimodal models. Li Zhou 0002, Zhenyu Liu 0006, Yutong Li 0007, Yuchi Duan, Bin Hu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2023 | A Visually Interpretable Convolutional-Transformer Model for Assessing Depression from Facial ImagesabstractThe accuracy and availability are the most critical and challenging problems for major depressive disorder (MDD) diagnosis. Limited receptive field and inaccurate visual interpretation always weaken the clinical application of deep learning-based depression recognition model. Thus, we propose a visually interpretable depression monitoring model termed Transformer and Convolutional with slot-attention (TC-slot) to assess depression from facial images. Specifically, this approach stands upon the intersection of convolution and transformer, combines self-attention mechanism and deep convolution, and uses a well-designed stem structure to explore the global and local relationships. Moreover, in TC-slot, a classifier built on slot-attention mechanism directly involved in the decision-making process further localizes salient regions of facial depression patterns and provides precise and meaningful explanations. The results indicate that the proposed approach effectively improves the classification and recognition performance compared with other state-of-the-art approaches, with guaranteed favorable visual interpretability, providing clinical insights into the assessment of the assessing depression. Yutong Li 0007, Zhenyu Liu 0006, Qiongqiong Chen, Zhijie Ding, Xiping Hu, Bin Hu 0001 |
ICME | 2 |
| 2023 | JAMFN: Joint Attention Multi-Scale Fusion Network for Depression Detection
Li Zhou 0002, Zhenyu Liu 0006, Zixuan Shangguan, Xiaoyan Yuan, Yutong Li 0007, Bin Hu 0001 |
INTERSPEECH | 2 |
| 2023 | An Automatic Depression Detection Method with Cross-Modal Fusion Network and Multi-head Attention Mechanism
Yutong Li 0007, Zhenyu Liu 0006, Li Zhou 0002, Xiping Hu, Bin Hu 0001 |
PRCV (5) | 3 |
| 2020 | Time-frequency Analysis Based on Hilbert-Huang Transform for Depression Recognition in SpeechabstractIn recent years, automatic detection of depression from speech has attracted many researchers. One of the key points is finding discriminable patterns in voice between depressed patients and healthy people. For this goal, we employed the Hilbert-Huang transform (HHT) to implement time-frequency analysis. Speech signals were decomposed into different sub-band signals and further were transformed into energy-frequency features for analysis and detection of depression. In the experiment 124 participants' (68 females and 56 males) speech were recorded in three patterns: interview, reading, and picture description for data collection. The results showed that the energy distribution of intrinsic mode functions (IMFs) between depressed patients and healthy people was significantly different, and this difference mainly was found in a relatively high-frequency range (1kHz). This finding fitted the clinical observation of depressed patients' “energy loss”. Further, a speech-based depression classification model based on the above finding was built and validated on the dataset. The results showed classification accuracy was 75.5% and 71.2% for female and male, respectively and each specificity was 88.4% and 78.2% These results implied HHT-based energy-frequency feature is a promising indicator for automatic depression assessment. Zhenyu Liu 0006, Zhijie Ding, Qiongqiong Chen |
BIBM | 1 |
| 2020 | A Novel Bimodal Fusion-based Model for Depression RecognitionabstractDepression is a common mental disorder which is harmful to our family, economics and society. Many people cannot receive timely mental health services, and the diagnosis process is subjective. A primary way for reducing harm is finding an objective and effective depression detection approach. Speech and video are two promising behavior indicators for depression. In this paper, we proposed a speech and video bimodal fusion model based on time-frequency analysis and convolutional neural network for this goal. For the testing of the proposed method, a speech and video dataset of 292 participants were employed for cross-validation. Compared with the single modal classification results, the classification accuracy and generalization ability of this gender-independent model are further improved, which is helpful for the identification of depression. Zhenyu Liu 0006, Zhijie Ding, Qiongqiong Chen |
HealthCom | 1 |
| 2020 | A Behaviour Patterns Extraction Method for Recognizing Generalized Anxiety DisorderabstractGeneralized anxiety disorder (GAD), as one of the most common chronic anxiety disorders, faces difficulties in clinical diagnosis. With the rapid development and wide application of smartphones in recent years, smartphones have a vivid application prospect in the field of mental disease monitoring and diagnosis. Based on WeChat applet platform on smartphones, an APP that integrates scale testing and inertial sensor data collection is developed to study the detection of subjects with GAD in task state. A behavior patterns extraction method is proposed using sliding windows to split behavior data, and processing data segments for clustering. Distribution information are extracted from the subjects' behavior patterns and are combined with the descriptive statistical features of the sample to identify GAD. The results show that this method has an accuracy of 66.44% for female subjects and 71.43% for male subjects in GAD recognition. Minqiang Yang, Jingsheng Tang, Yushan Wu, Zhenyu Liu 0006, Xiping Hu, Bin Hu 0001 |
HealthCom | 4 |
| 2018 | A novel study for MDD detection through task-elicited facial cues
Zhenyu Liu 0006, Zhijie Ding, Gangping Wang |
BIBM | 2 |
| 2017 | Speech pause time: A potential biomarker for depression detectionabstractDetecting depression via speech is an attractive topic in recent years. Significant correlation was found between speech pause time and depressive severity. In the present study, 92 depressed patients and 92 age-, gender- and education level-matched control participants were examined to investigate three temporal characteristics of speech: recording time (RT), phonation time (PT) and speech pause time (SPT). The results show that depressed patients' duration measures are longer than healthy controls in most cases and spontaneous speech is better than automatic speech for these measures' acquisition. Among these three measures, although speech pause time could be influenced by antidepressant and interview topic, it is still an effective biomarker for depression. Zhenyu Liu 0006, Huanyu Kang, Lei Feng 0005 |
BIBM | 1 |
| 2017 | Ensemble-based depression detection in speechabstractDepression detection using speech signal is becoming an attractive topic because it is fast, convenient and non-invasive. Many researches aimed at improving depression classification performance. This study investigated application of ensemble learners in depression detection and compared three speaking styles (interview, reading and picture description) in ensembles. A speech dataset collecting from 184 subjects (92 depressed patients and 92 healthy controls) was used for these goals. The results showed that ensemble learners perform better than individual learners apparently. Interview is a more effective speaking style than reading and picture description for speech acquisition. These findings suggest us ensemble model based multi-utterance in interview is the best way to detect depression. Zhenyu Liu 0006, Changcong Li, Gang Wang 0012 |
BIBM | 1 |
| 2017 | Detecting depression in speech: Comparison and combination between different speech typesabstractDepression is a mental disorder of high prevalence, leading to a negative effect on individuals, their families, society and the economy. In recent years, the problem of automatic detection of depression from the speech signal has gained more interest. In this paper, a new multiple classifier system for depression recognition was developed and tested. The novel aspect of this methodology is the combination of different speech types and emotions. First of all, using a sample of 74 subjects (37 depressed patients and 37 healthy controls), we examined the discriminative power of different speech types (interview, picture description, and reading) and speech emotions (positive, neutral, and negative). Some voice features (e.g. short time energy, intensity, loudness, zero-crossing rate (ZCR), F0, jitter, shimmer, formants, mel frequency cepstral coefficients (MFCC), linear prediction coefficient (LPC), line spectrum pair (LSP), and perceptual linear predictive coefficients (PLP)) were tested. Then, a new multiple classifier method was proposed to detect depression. It was observed that the overall recognition rate using interview speech was higher than employing picture description speech and reading speech. Furthermore, neutral speech showed better performance than positive and negative speech. Among these features, short time energy, ZCR, LPC, MFCC and LSP were the robust features that gave high accuracy in different types of speech. Finally, this new approach showed a high accuracy of 78.02%, giving high encouragement for detecting depression in speech. Hailiang Long, Zhenghao Guo, Xia Wu 0001, Bin Hu 0001, Zhenyu Liu 0006, Hanshu Cai |
BIBM | 5 |
| 2017 | Investigation of different speech types and emotions for detecting depression using different classifiers
Haihua Jiang, Bin Hu 0001, Zhenyu Liu 0006, Lihua Yan, Fei Liu 0037, Huanyu Kang |
Speech Commun. | 3 |
| 2016 | Assessing stress levels via speech using three reading patternsabstractVarious problems caused by stress seriously affect individuals' physical and mental well-being and have been receiving an increasing attention in modern lives. Since traditional stress assessment methods are lack of objectivity, affective sensing technologies have been studied for years. As detecting stress in speech has the advantages of non-invasive, portable, fast, and less expensive, many explorations were conducted to build stress assessment models. To find out a proper acoustic feature subset for a specific reading pattern, we performed the experiments with 30 subjects by three reading patterns: vowel, figure and sentence. We utilized feature selection and classification techniques to automatically select acoustic features and evaluate performances. Results showed that there are interactions between reading patterns and stress levels on speech features. Although Stress levels can be distinguished in any pattern of them (vowel, figure, sentence), sentence is a better choice with the best classification accuracy 88.15%. Furthermore, Line Spectral Pairs (LSP) features are indispensable for vowel, Mel-Frequency Cepstral Coefficient (MFCC) features are more effective for figure and the combination of prosodic, LSP and MFCC features is more suitable for sentence. Zhenyu Liu 0006, Lihua Yan, Bin Hu 0001, Fei Liu 0037 |
BIBM | 1 |
| 2015 | Detection of depression in speechabstractDepression is a common mental disorder and one of the main causes of disability worldwide. Lacking objective depressive disorder assessment methods is the key reason that many depressive patients can't be treated properly. Developments in affective sensing technology with a focus on acoustic features will potentially bring a change due to depressed patient's slow, hesitating, monotonous voice as remarkable characteristics. Our motivation is to find out a speech feature set to detect, evaluate and even predict depression. For these goals, we investigate a large sample of 300 subjects (100 depressed patients, 100 healthy controls and 100 high-risk people) through comparative analysis and follow-up study. For examining the correlation between depression and speech, we extract features as many as possible according to previous research to create a large voice feature set. Then we employ some feature selection methods to eliminate irrelevant, redundant and noisy features to form a compact subset. To measure effectiveness of this new subset, we test it on our dataset with 300 subjects using several common classifiers and 10-fold cross-validation. Since we are collecting data currently, we have no result to report yet. Zhenyu Liu 0006, Bin Hu 0001, Lihua Yan, Fei Liu 0037, Huanyu Kang |
ACII | 1 |
| 2015 | Feature selection and classification of speech under long-term stressabstractMany studies were proposed to discuss acoustic correlates of stress in recent years. Considering some inconsistent experiment results, we supposed that stress should be categorized into long-term and short-term stress in this topic, and the trend of short-term stress induced by workload may be affected by long-term stress. This study contains three parts: first, we proposed an acoustic feature set chosen by feature selection, which can be considered as a measurement of the level of long-term stress; second, we showed that this set is immune to short-term stress in stress classification tests; finally, we observed some particular voice features mentioned in previous researches in our experiment and the results may imply that short-term stress trend is in connection with the level of long-term stress. In short, long-term and shot-term stress should be discussed separately in future researches for clear and explicit conclusions. Bin Hu 0001, Zhenyu Liu 0006, Lihua Yan, Fei Liu 0037, Huanyu Kang |
BIBM | 2 |