Jiayu Ye

dblp:182/7616 · DBLP profile ↗
← Back
18ranked-venue papers
8as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 6 since 2021Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Dep-MAP: A Multi-level Alignment Framework with Semantic Prototypes for Video-based Automatic Depression Assessment
abstract
Spatiotemporal analysis of facial behavior is a crucial method for evaluating the mental state of depression patients. However, in practice, depressed patients often display facial behaviors similar to healthy individuals due to masking tendencies. Additionally, facial expressions among depressed patients are also different, increasing the difficulty of assessment. To address this, we propose a video-based automatic depression assessment model Dep-MAP for complex facial behaviors of depression patients. Dep-MAP adopts a dual-branch architecture to extract visual features of facial behavior and capture corresponding emotional semantic features. Specifically, the extracted deep semantic features are clustered, resulting in semantically distinct prototype sets, where each severity group learns a set of discriminative facial behavior prototype representations, to suppress inter-class semantic confusion. Subsequently, we propose a semantic prototype-supervised contrastive learning method, which aligns latent semantics between shallow and deep features, realizing emotional semantic guidance and self-knowledge distillation for the visual feature branch, effectively suppressing intra-class difference. Then, we integrate key depression cues across multiple spatiotemporal scales via a multi-scale weighted fusion strategy, achieving automatic depression assessment. Experimental results demonstrate that Dep-MAP effectively identifies potential key frames in temporal sequences, and aggregates key frame representations with semantic consistency, achieving significantly superior state-of-the-art results on the AVEC2013 and AVEC2014 public datasets.
Jiayu Ye, Qingxiang Wang
AAAI2
2026 Brain Connectivity Variability Influences Anxiety Through the Behavioral Inhibition System
abstract
The behavioral inhibition system (BIS), mediating responses to punishment cues and avoidance behaviors, is implicated in anxiety. However, the neural dynamics underpinning BIS, particularly regarding the temporal variability of brain network interactions, remain less explored. Using resting-state functional magnetic resonance imaging (rs-fMRI) of 181 healthy adults, this study investigated the association between BIS sensitivity and the temporal variability of functional connectivity within and between functional brain networks. This finding revealed a significant positive correlation between BIS scores and temporal variability, specifically in the connectivity involving subnetworks' sensory somatomotor hand network (SSHN)-ventral attention network (VAN), and sensory somatomotor mouth network (SSMN)-VAN. Notably, the high-BIS sensitivity group exhibited significantly greater temporal variability between VAN and SSMN/SSHN compared to the low-BIS sensitivity group. Furthermore, predicted BIS scores based on network variability showed a strong correlation with actual BIS scores (Pearson's [Formula: see text]). Moreover, significant mediation effects highlighted the bridging role of BIS scores between brain network variability and anxiety scale scores. This enhances the comprehension of the relationship between BIS, anxiety, and brain function, while also offering new insights into the pathogenesis of anxiety.
Runyang He, Jiayu Ye, Dezhong Yao 0001, Peng Xu 0001, Fali Li, Lin Jiang 0004
Int. J. Neural Syst.3
2026 T2Net: Tongue Image-Based T2DM Detection via Simulated Clinical Diagnostic Reasoning
abstract
Clinical studies indicate that the progression of Type 2 Diabetes Mellitus (T2DM) is associated with characteristic alterations in tongue features, which may facilitate non-invasive early detection. However, current deep learning-based tongue imaging approaches for diabetes diagnosis remain constrained by limited datasets, subtle feature variations, dependence on clinical expertise, and the lack of quantitative evaluation. To address these issues, we developed an open-source dataset for T2DM tongue diagnosis (DMT) and benchmarked it using multiple baseline models. Building on DMT, we propose T2Net, a tongue image recognition model for T2DM that simulates the clinical diagnostic process. T2Net comprises four core components: local inspection, pathological clue integration, syndrome identification, and diagnostic confidence estimation. First, T2Net automatically extracts key ROIs by combining large-kernel decomposition with multi-scale learning. Then, a multi-order feature interaction module enables effective fusion of tongue image features across scales to capture pathological clues. Meanwhile, we design a context-aware dynamic aggregation convolution to model long-range dependencies, and propose a flexible focal loss to mimic the diagnostic reasoning process of clinicians, enabling brain-inspired inference. Finally, we propose a clustering-based confidence estimation approach to quantitatively evaluate the reliability of model predictions. Experimental results demonstrate that T2Net achieves highly competitive performance on the DMT dataset, outperforming the second-best baseline by 2.7% in accuracy and 2.0% in F1 score. Moreover, the quantitative evaluation scores are largely consistent with clinical assessments by physicians.
Yanyi Huang, Liyun Li, Xiaojie Feng, Miao Xie, Jiayu Ye, An Zeng, Jianlu Bi
IEEE J. Biomed. Health Informatics8
2026 MFE-Former: Disentangling Emotion-Identity Dynamics via Self-Supervised Learning for Enhancing Speech-Driven Depression Detection
abstract
Acoustic features are crucial behavioral indicators for depression detection. However, prior speech-based depression detection methods often overlook the variability of emotional patterns across samples, leading to interference from speaker identity and hindering the effective extraction of emotional changes. To address this limitation, we developed the Emotional Word Reading Experiment (EWRE) and introduced a method combining self-supervised and supervised learning for depression detection from speech called MFE-Former. First, we generate fine-grained emotional representations for response segments by computing cosine similarity between intra-sample and inter-sample contexts. Concurrently, orthogonality constraints decouple identity information from emotional features, while a Transformer decoder reconstructs spectral structures to improve sensitivity to depression-related emotional patterns. Next, we propose a multi-scale emotion change perception module and a Bernoulli distribution-based joint decision module integrate multi-level information for depression detection. By enhancing the distribution differences among positive, neutral, and negative emotional features, we find that patients with depression are more inclined to express negative emotions, whereas healthy individuals express more positive emotions. The experimental results on EWRE and AVEC 2014 show that MFE-Former outperforms state-of-the-art temporal methods under conditions of variability in emotional patterns across samples.
Jiayu Ye, Yanhong Yu, Lin Yuan 0001, Qingxiang Wang
IEEE J. Biomed. Health Informatics2
2026 NIDC: General Task Backbone for Neuroimaging Analysis via Interpretable Deep Clustering
abstract
Clustering techniques offer strong interpretability. However, they have significant limitations in the deep learning area due to their difficulty in capturing complex data structures, such as spatial and contextual information. This issue is especially pronounced in neuroimaging research, where high-dimensional and complex data greatly restricts the feature representation capability of clustering models. Hence, we propose a Neuroimaging Deep Clustering (NIDC) backbone network. We convert 3D neuroimagings into point sets and design clustering-based paradigms for context feature aggregation, feature interaction, and feature dispatching to enable deep feature extraction from the point sets. To better capture spatial information, we propose brain spatial relative position encoding, which assists the clustering paradigm in better understanding the anatomical structure of brain tissue and the positional relationships between different regions. Additionally, we design a sample center loss function to encourage tighter clustering of labels or voxels/feature points of the same class in the feature space, aiming to suppress both inter-class and intra-class similarities. Meanwhile, NIDC preserves the interpretability of traditional clustering techniques, allowing it to uncover relationships between brain regions and trace the decision-making process of each voxel. NIDC achieves highly competitive performance across multiple datasets in various downstream tasks, emerging as a new and practical solution for neuroimaging analysis. Code is available athttps://github.com/IMCTGD/NIDC.
Jiayu Ye, An Zeng, Dan Pan 0001, Jingliang Zhao, Yiqun Zhang 0006, Yang Liu 0007
IEEE Trans. Multim.1
2025 MSAF-Net: A Multi-Scale Adaptive Fusion Network for Facial Expression Recognition in Mental Health Patients
abstract
Facial expression recognition (FER) is essential for emotional assessment in the treatment and monitoring of mental health conditions. However, the scarcity of facial expression data from patients with mental illnesses, coupled with the subtlety and complexity of their expressions, such as limited emotional intensity and minimal facial movement presents significant challenges. To address this, we introduced the Voluntary Facial Expression Mimicry (VFEM) experiment, collecting data on seven types of expressions from patients with depression and anxiety. Based on VFEM dataset, we developed a novel Multi-Scale Adaptive Fusion Network (MSAF-Net). First, we designed a Multi-Dimensional Feature Refinement Module to enhance expression feature representation. Next, we proposed a Multi-Scale Adaptive Fusion Module to improve feature fusion and consistency. Finally, by incorporating Center Cross-Entropy Loss, we optimized feature distribution and classification. Extensive experiments on the VFEM dataset, compared with state-of-the-art models, show that our approach achieves competitive results in mental health facial expression recognition.
Guolong Liu, Jiayu Ye, Qingxiang Wang
ICME2
2025 MedGNN: General Medical Image Recognition Network via GNN Visual Representations
Jiayu Ye, An Zeng, Dan Pan 0001, Guanwei Cheng
MICCAI (16)1
2025 DEP-Former: Multimodal Depression Recognition Based on Facial Expressions and Audio Features via Emotional Changes
abstract
Clinical research has demonstrated that exploring behavioral signal differences between depressed patients and non-depressed people using audiovisual technology is an effective approach for achieving depression recognition. Hence, in this paper we propose an emotion word reading experiment (EWRE), and extract features from facial expressions and audios for depression recognition. Building upon this, we propose a depression recognition model (DEP-Former), which deeply integrates multimodal features. DEP-Former first designs a modality adapter to achieve emotion space mapping and the sharing of multimodal features, addressing cross-modal inconsistencies. Simultaneously, it proposes a mechanism of attention index sharing, exceeding the limitations of cognitive subjectivity by calculating confidence in key emotional information across modalities. Finally, we propose a multimodal cross-attention module and a Bernoulli distribution feature fusion prediction module to achieve deep integration of multilevel information, thereby enabling depression recognition. Compared with existing advanced multimodal models, DEP-Former demonstrates superior performance in EWRE, achieving an accuracy of 0.9500 and an F1 score of 0.9499, significantly enhancing depression recognition over the single-modality methods. Furthermore, its robust generalization ability is validated on the AVEC 2014 dataset. Through the attention query of the interpretability analysis module, we discover that depressed patients exhibit heightened sensitivity to negative emotional words, such as dismissal and tragedy. In contrast, healthy individuals tend to be more attuned to positive emotional words, including passion, purity, and justice. Additionally, depressed patients exhibit a degree of psychological state diversity, showing sensitivity to some positive emotional words as well. Our codes and data are available athttps://github.com/QLUTEmoTechCrew/DEP-Former.
Jiayu Ye, Yanhong Yu, Yunshao Zheng, Qingxiang Wang
IEEE Trans. Circuits Syst. Video Technol.1
2025 CmdVIT: A Voluntary Facial Expression Recognition Model for Complex Mental Disorders
abstract
Facial Expression Recognition (FER) is a critical method for evaluating the emotional states of patients with mental disorders, playing a significant role in treatment monitoring. However, due to privacy constraints, facial expression data from patients with mental disorders is severely limited. Additionally, the more complex inter-class and intra-class similarities compared to healthy individuals make accurate recognition of facial expressions challenging. Therefore, we propose a Voluntary Facial Expression Mimicry (VFEM) experiment, which collected facial expression data from schizophrenia, depression, and anxiety. This experiment establishes the first dataset designed for facial expression recognition tasks exclusively composed of patients with mental disorders. Simultaneously, based on VFEM, we propose a Vision Transformer FER model tailored for Complex mental disorder patients (CmdVIT). CmdVIT integrates crucial facial expression features through both explicit and implicit mechanisms, including explicit visual center positional encoding and implicit sparse attention center loss function. These two key components enhance positional information and minimize the facial feature space distance between conventional attention and critical attention, effectively suppressing inter-class and intra-class similarities. In various FER tasks for different mental disorders in VFEM, CmdVIT achieves more competitive performance compared to contemporary benchmark models. Our works are available at https://github.com/yjy-97/CmdVIT.
Jiayu Ye, Yanhong Yu, Qingxiang Wang, Guolong Liu, An Zeng, Yiqun Zhang 0006, Yang Liu 0007, Yunshao Zheng
IEEE Trans. Image Process.1
2025 Dynamic Local Conformal Reinforcement Network (DLCR) for Aortic Dissection Centerline Tracking
abstract
Pre-extracted aortic dissection (AD) centerline is very useful for quantitative diagnosis and treatment of AD disease. However, centerline extraction is challenging because (i) the lumen of AD is very narrow and irregular, yielding failure in feature extraction and interrupted topology; and (ii) the acute nature of AD requires a quick algorithm, however, AD scans usually contain thousands of slices, centerline extraction is very time-consuming. In this paper, a fast AD centerline extraction algorithm, which is based on a local conformal deep reinforced agent and dynamic tracking framework, is presented. The potential dependence of adjacent center points is utilized to form the novel 2.5D state and locally constrains the shape of the centerline, which improves overlap ratio and accuracy of the tracked path. Moreover, we dynamically modify the width and direction of the detection window to focus on vessel-relevant regions and improve the ability in tracking small vessels. On a public AD dataset that involves 100 CTA scans, the proposed method obtains average overlap of 97.23% and mean distance error of 1.28 voxels, which outperforms four state-of-the-art AD centerline extraction methods. The proposed algorithm is very fast with average processing time of 9.54s, indicating that this method is very suitable for clinical practice.
Jingliang Zhao, An Zeng, Jiayu Ye, Dan Pan 0001
IEEE J. Biomed. Health Informatics3
2024 Dep-FER: Facial Expression Recognition in Depressed Patients Based on Voluntary Facial Expression Mimicry
abstract
Facial expressions are important nonverbal behaviors that humans use to express their feelings. Clinical research have shown that depressed patients have poor facial expressiveness and mimicry. As a result, we propose a VFEM experiment with seven expressions to explore variations in facial expression features between depressed patients and normal people, including anger, disgust, fear, happiness, neutrality, sadness, and surprise. It has been discovered through VFEM experiments that depressed patients frequently exhibit negative facial expressions. Meanwhile, we propose a depression facial expression recognition (Dep-FER) model in this research. Dep-FER involves three innovative and crucial components: Mask Multi-head Self-Attention (MMSA), facial action unit similarity loss function (AUs Loss), and case-control loss function (CC Loss). MMSA can filter out disturbing samples and force to learn the relationship between different samples. AUs Loss utilizes the similarity between each expression AU and the model output to improve the generalization ability of the model. CC Loss addresses the intrinsic link between the depressed and normal patient categories. Dep-FER achieves excellent performance in VFEM and outperforms existing comparative models.
Jiayu Ye, Yanhong Yu, Yunshao Zheng, Qingxiang Wang
IEEE Trans. Affect. Comput.1
2024 MAD-Former: A Traceable Interpretability Model for Alzheimer's Disease Recognition Based on Multi-Patch Attention
abstract
The integration of structural magnetic resonance imaging (sMRI) and deep learning techniques is one of the important research directions for the automatic diagnosis of Alzheimer's disease (AD). Despite the satisfactory performance achieved by existing voxel-based models based on convolutional neural networks (CNNs), such models only handle AD-related brain atrophy at a single spatial scale and lack spatial localization of abnormal brain regions based on model interpretability. To address the above limitations, we propose a traceable interpretability model for AD recognition based on multi-patch attention (MAD-Former). MAD-Former consists of two parts: recognition and interpretability. In the recognition part, we design a 3D brain feature extraction network to extract local features, followed by constructing a dual-branch attention structure with different patch sizes to achieve global feature extraction, forming a multi-scale spatial feature extraction framework. Meanwhile, we propose an important attention similarity position loss function to assist in model decision-making. The interpretability part proposes a traceable method that can obtain a 3D ROI space through attention-based selection and receptive field tracing. This space encompasses key brain tissues that influence model decisions. Experimental results reveal the significant role of brain tissues such as the Fusiform Gyrus (FuG) in AD recognition. MAD-Former achieves outstanding performance in different tasks on ADNI and OASIS datasets, demonstrating reliable model interpretability.
Jiayu Ye, An Zeng, Dan Pan 0001, Yiqun Zhang 0006, Jingliang Zhao, Qiuping Chen, Yang Liu 0007
IEEE J. Biomed. Health Informatics1
2023 Spa-L Transformer: Sparse-self attention model of Long short-term memory positional encoding based on long text classification
abstract
The emergence of Transformer and its derivative models brings new opportunities to tasks of NLP (Natural Language Processing). Transformer is not only a separate model, but also the core of different text task systems. Therefore, Transformer has become an important component of many powerful models. However, Transformer is not without defects. Researchers are still puzzled by the huge amount of computation generated in the process of self-attention. Especially in long text data sets. We propose a new ProbSparse self-attention Transformer model based for text classification. In the following, we will call it SpaL Transformer. We query important attention factors through KL divergence and add Long Short Term Memory(LSTM) to positional encoding, and only focus on the main query. Then, we select the most important relevant attention based on the confidence score to focus the overall attention. At the same time, we propose LSTM positive encoding to obtain relative position information to optimize the model. In long text dataset IMDB, our model improves the accuracy of Transformer. And F1 score improved by 0.064.
Shengzhe Zhang, Jiayu Ye, Qingxiang Wang
CSCWD2
2023 Analysis and Recognition of Voluntary Facial Expression Mimicry Based on Depressed Patients
abstract
Many clinical studies have shown that facial expression recognition and cognitive function are impaired in depressed patients. Different from spontaneous facial expression mimicry (SFEM), 164 subjects (82 in a case group and 82 in a control group) participated in our voluntary facial expression mimicry (VFEM) experiment using expressions of neutrality, anger, disgust, fear, happiness, sadness and surprise. Our research is as follows. First, we collected a large amount of subject data for VFEM. Second, we extracted the geometric features of subject facial expression images for VFEM and used Spearman correlation analysis, a random forest, and logistic regression-based recursive feature elimination (LR-RFE) to perform feature selection. The features selected revealed the difference between the case group and the control group. Third, we combined geometric features with the original images and improved the advanced deep learning facial expression recognition (FER) algorithms in different systems. We propose the E-ViT and E-ResNet based on VFEM. The accuracies and F1 scores were higher than those of the baseline models, respectively. Our research proved that it is effective to use feature selection to screen geometric features and combine them with a deep learning model for depression facial expression recognition.
Jiayu Ye, Yanhong Yu, Yunshao Zheng, Yitao Zhu, Qingxiang Wang
IEEE J. Biomed. Health Informatics1
2022 Dep-ViT: Uncertainty Suppression Model Based on Facial Expression Recognition in Depression Patients
Jiayu Ye, Guanwei Cheng, Qingxiang Wang
ICANN (3)1
2020 AutoHOOT: Automatic High-Order Optimization for Tensors
abstract
High-order optimization methods, including Newton's method and its variants as well as alternating minimization methods, dominate the optimization algorithms for tensor decompositions and tensor networks. These tensor methods are used for data analysis and simulation of quantum systems. In this work, we introduce AutoHOOT, the first automatic differentiation (AD) framework targeting at high-order optimization for tensor computations. AutoHOOT takes input tensor computation expressions and generates optimized derivative expressions. In particular, AutoHOOT contains a new explicit Jacobian / Hessian expression generation kernel whose outputs maintain the input tensors' granularity and are easy to optimize. The expressions are then optimized by both the traditional compiler optimization techniques and specific tensor algebra transformations. Experimental results show that AutoHOOT achieves competitive CPU and GPU performance for both tensor decomposition and tensor network applications compared to existing AD software and other tensor computation libraries with manually written kernels. The tensor methods generated by AutoHOOT are also well-parallelizable, and we demonstrate good scalability on a distributed memory supercomputer.
Linjian Ma, Jiayu Ye, Edgar Solomonik
PACT2
2020 Inefficiency of K-FAC for Large Batch Size Training
abstract
There have been several recent work claiming record times for ImageNet training. This is achieved by using large batch sizes during training to leverage parallel resources to produce faster wall-clock training times per training epoch. However, often these solutions require massive hyper-parameter tuning, which is an important cost that is often ignored. In this work, we perform an extensive analysis of large batch size training for two popular methods that is Stochastic Gradient Descent (SGD) as well as Kronecker-Factored Approximate Curvature (K-FAC) method. We evaluate the performance of these methods in terms of both wall-clock time and aggregate computational cost, and study the hyper-parameter sensitivity by performing more than 512 experiments per batch size for each of these methods. We perform experiments on multiple different models on two datasets of CIFAR-10 and SVHN. The results show that beyond a critical batch size both K-FAC and SGD significantly deviate from ideal strong scaling behaviour, and that despite common belief K-FAC does not exhibit improved large-batch scalability behavior, as compared to SGD.
Linjian Ma, Gabe Montague, Jiayu Ye, Zhewei Yao, Amir Gholami, Kurt Keutzer, Michael W. Mahoney
AAAI3
2020 Q-BERT: Hessian Based Ultra Low Precision Quantization of BERT
abstract
Transformer based architectures have become de-facto models used for a range of Natural Language Processing tasks. In particular, the BERT based models achieved significant accuracy gain for GLUE tasks, CoNLL-03 and SQuAD. However, BERT based models have a prohibitive memory footprint and latency. As a result, deploying BERT based models in resource constrained environments has become a challenging task. In this work, we perform an extensive analysis of fine-tuned BERT models using second order Hessian information, and we use our results to propose a novel method for quantizing BERT models to ultra low precision. In particular, we propose a new group-wise quantization scheme, and we use Hessian-based mix-precision method to compress the model further. We extensively test our proposed method on BERT downstream tasks of SST-2, MNLI, CoNLL-03, and SQuAD. We can achieve comparable performance to baseline with at most 2.3% performance degradation, even with ultra-low precision quantization down to 2 bits, corresponding up to 13× compression of the model parameters, and up to 4× compression of the embedding table as well as activations. Among all tasks, we observed the highest performance loss for BERT fine-tuned on SQuAD. By probing into the Hessian based analysis as well as visualization, we show that this is related to the fact that current training/fine-tuning strategy of BERT does not converge for SQuAD.
Sheng Shen 0001, Zhen Dong 0003, Jiayu Ye, Linjian Ma, Zhewei Yao, Amir Gholami, Michael W. Mahoney, Kurt Keutzer
AAAI3