Yanrong Guo

dblp:122/3196 · DBLP profile ↗
← Back
59ranked-venue papers
11as first author
33since 2021 · last 2026
0000-0001-6949-4879ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 27 · 3 first-author · 16 since 2021Artificial intelligence and machine learning · 17 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 7 first-author · 4 since 2021Computer networks · 3 · 3 since 2021Security and privacy · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Voices, Faces, and Feelings: Multi-modal Emotion-Cognition Captioning for Mental Health Understanding
abstract
Emotional and cognitive factors are essential for understanding mental health disorders. However, existing methods often treat multi-modal data as classification tasks, limiting interpretability especially for emotion and cognition. Although large language models (LLMs) offer opportunities for mental health analysis, they mainly rely on textual semantics and overlook fine-grained emotional and cognitive cues in multi-modal inputs. While some studies incorporate emotional features via transfer learning, their connection to mental health conditions remains implicit. To address these issues, we propose ECMC, a novel task that aims at generating natural language descriptions of emotional and cognitive states from multi-modal data, and producing emotion–cognition profiles that improve both the accuracy and interpretability of mental health assessments. We adopt an encoder–decoder architecture, where modality-specific encoders extract features, which are fused by a dual-stream BridgeNet based on Q-former. Contrastive learning enhances the extraction of emotional and cognitive features. A LLaMA decoder then aligns these features with annotated captions to produce detailed descriptions. Extensive objective and subjective evaluations demonstrate that: 1) ECMC outperforms existing multi-modal LLMs and mental health models in generating emotion–cognition captions; 2) the generated emotion–cognition profiles significantly improve assistive diagnosis and interpretability in mental health analysis.
Yanrong Guo, Shijie Hao
AAAI2
2026 Illumination-Prior Guided Hybrid Network for Low-Light Image Enhancement
Shijie Hao, Yanrong Guo
MMM (1)3
2026 From disagreement to insight: multi-agent collaborative reasoning for multimodal emotional support conversations
Yuqi Chu, Yanrong Guo, Richang Hong
Multim. Syst.2
2026 Interview-Based Depression Detection Using LLM-Based Text Restatement and Emotion Lexicon
abstract
Depression is a mental health disorder that significantly impacts modern society. Developing accurate depression detection models by leveraging discriminant features from multimedia or physiological data can aid medical professionals in making informed diagnoses. According to psychological studies, emotion is a critical indicator of depression. However, emotion has not been utilized as a central role in current research on assistive depression detection, usually serving as a supplementary information source or a guidance for integrating diverse data modalities. In contrast to existing studies, we investigate the feasibility of detecting depression by concentrating on emotion information. Specifically, focusing on modeling emotion feature representation during interviews, we propose an interview-based depression detection model via leveraging large language (LLM) based text restatement and emotion lexicon (IDD-LTE). In this model, we employ LLM to enhance text quality through restatement to address the potentially low quality of interview text data. Using an emotion lexicon, the open contents in restated texts are mapped to a fixed-size matrix representation that captures the interviewee's emotional state and mood swings during the conversation, serving as the fundamental representation for the following discriminant feature learning. The proposed IDD-LTE model is evaluated on four primary datasets for depression detection. The promising results confirm the feasibility and effectiveness of our model.
Shijie Hao, Jingjing Wu 0001, Yanrong Guo, Richang Hong
IEEE Trans. Affect. Comput.4
2026 Subthreshold Depression Detection With Text-Guided Multimodal Learning
abstract
Depression, a widespread global mental health problem, affects millions of people annually, making early detection of subclinical depression crucial for timely intervention. Current automatic depression detection (ADD) methods, valuable for diagnosis, often neglect subthreshold populations and face difficulties in extracting diagnostic data from long sequences of multimodal information. These methods also inadequately leverage text modality, which is less noisy and information-rich compared with other modalities. Furthermore, existing datasets for depression research are often too small, limiting the generalizability of developed methods. To address these issues, this article proposes a new approach for detecting depression in subthreshold populations. For long-sequence samples in the field of depression, we construct an autoencoder that compresses along both temporal and feature dimensions, aiming to extract the most compact and effective features from the samples. To exploit the text modality’s advantages, we integrate the RoBERTa pretrained model with an attention mechanism for high-quality text encoding. We then develop a text-guided multimodal fusion (TGMF) module, using text encoding as an anchor for guiding audio and video modality encoding, ensuring multimodal alignment. Additionally, contrastive learning is applied to discern differences between classes, enhancing the model’s generalizability. Our method demonstrates superior performance in the tasks of detecting depression and identifying subthreshold populations on the E-DAIC and MMDA datasets.
Yanrong Guo, Youwei Guo, Bingxin Yang, Jingjing Wu 0001, Shijie Hao, Richang Hong
IEEE Trans. Comput. Soc. Syst.1
2026 Controllable Relation Disentanglement for Few-Shot Class-Incremental Learning
abstract
Few-Shot Class-Incremental Learning (FSCIL) requires models to be updated incrementally with limited labeled samples given in each session, differing from the traditional training paradigm and easily resulting in severe spurious relations between categories. Thus, in this paper, we propose to address FSCIL from a new perspective: enhancing FSCIL via disentangling spurious relations between categories. Accordingly, we propose a simple yet effective approach, dubbed ConTrollable Relation-disentangLed Few-Shot Class-Incremental Learning (CTRL-FSCIL). Specifically, during a base session, we propose to anchor base class embeddings in feature space and build disentangled proxies to bridge gaps between the learning processes of categories encountered in different sessions, making category relations controllable. Furthermore, during incremental learning, the parameters of the backbone network are frozen in order to relieve the negative impact of data scarcity. Meanwhile, a relation disentanglement loss is employed to guide a relation control module to disentangle spurious relations between learned categories. In this way, spurious relation issues in FSCIL can be alleviated. Extensive experiments on CIFAR-100, mini-ImageNet, and CUB-200 demonstrate the effectiveness of CTRL-FSCIL. Our code has been publicly released on github.
Yuan Zhou 0016, Richang Hong, Yanrong Guo, Lin Liu 0016, Shijie Hao, Hanwang Zhang
IEEE Trans. Circuits Syst. Video Technol.3
2026 Dynamic Correlation-Guided Disentanglement and Contrastive Learning for RGB-D Cross-Modal Re-Identification
abstract
Person re-identification (Re-ID) across RGB and depth modalities offers complementary cues for robust pedestrian matching under challenging conditions. However, the significant discrepancy between RGB appearance features and depth structural features complicates cross-modal alignment. Existing methods either depend on static architectural designs or impose strong constraints to capture the common features of the two modalities, often suffering from branch imbalance or distorted identity features. In this work, we propose a novel framework, Dynamic Correlation-Guided Disentanglement and Contrastive Learning (DCG-DCL), for RGB-D cross-modal Re-ID. First, the Dynamic Correlation-guided Disentanglement (DCGD) dynamically decouples features with the guidance of inter-modal correlation, which explicitly enforces common-feature learning via a cross-correlation constraint and adaptively separates common and unique components without predefined assumptions. Second, a Common & Unique Contrastive Learning (CUCL) strategy fully leverages these decoupled features, which aligns RGB/depth features closer to their common representation and pushes them away from unique redundancies. This dual mechanism effectively narrows modality discrepancy and boosts robustness against modality-specific noise. Extensive experiments on multiple public benchmarks demonstrate that our method achieves state-of-the-art performance, with ablation studies validating the necessity of each component.
Zhibo Lei, Jingjing Wu 0001, Yaxiong Wang, Yanrong Guo, Shijie Hao, Richang Hong
IEEE Trans. Inf. Forensics Secur.4
2026 Biomedical Relation Extraction via Adaptive Document-Relation Cross-Mapping and Concept Unique Identifier
abstract
Document-Level Biomedical Relation Extraction (Bio-RE) aims to identify relations between biomedical entities within extensive texts, serving as a crucial subfield of biomedical text mining. Existing Bio-RE methods struggle with cross-sentence inference, which is essential for capturing relations spanning multiple sentences. Moreover, previous methods often overlook the incompleteness of documents and lack the integration of external knowledge, limiting contextual richness. Besides, the scarcity of annotated data further hampers model training. Recent advancements in large language models (LLMs) have inspired us to explore all the above issues for document-level Bio-RE. Specifically, we propose a document-level Bio-RE framework via LLM Adaptive Document-Relation Cross-Mapping (ADRCM) Fine-Tuning and Concept Unique Identifier (CUI) Retrieval-Augmented Generation (RAG). First, we introduce the Iteration-of-REsummary (IoRs) prompt for solving the data scarcity issue. In this way, Bio-RE task-specific synthetic data can be generated by guiding ChatGPT to focus on entity relations and iteratively refining synthetic data. Next, we propose ADRCM fine-tuning, a novel fine-tuning recipe that establishes mappings across different documents and relations, enhancing the model’s contextual understanding and cross-sentence inference capabilities. Finally, during the inference, a biomedical-specific RAG approach, named CUI RAG, is designed to leverage CUIs as indexes for entities, narrowing the retrieval scope and enriching the relevant document contexts. Experiments conducted on three Bio-RE datasets—GDA, CDR, and BioRED—demonstrate the state-of-the-art performance of our proposed method by comparing it with other related works.
Yufei Shang, Yanrong Guo, Shijie Hao, Richang Hong
ACM Trans. Knowl. Discov. Data2
2026 Infrared Object Tracking via Complementary Dual-domain Interaction with Target-guided Frequency Transformation
abstract
Infrared Object Tracking (IOT) is challenging due to the low contrast of infrared images, which limits effective spatial feature extraction. Although recent works have explored frequency-domain information, their utilization remains insufficient, and fusion strategies either retain redundancy or fail to fully explore distinctive differences, thus limiting complementary enhancement. To overcome this, we propose a novel tracker that introduces a Target-guided Frequency Transformation Module (TFTM) and a Dual-domain Interactive Fusion Network (DIFN). The former extracts multi-frequency representations across scales and orientations, guided by an adaptive mask strategy to suppress background interference. The latter fuses the two domains with differentiated attention to achieve complementary enhancement. Extensive experiments show that our approach achieves superior performance over state-of-the-art trackers, highlighting the effectiveness of comprehensive frequency-domain integration in IOT.
Pengyu Huang, Jingjing Wu 0001, Yanrong Guo, Richang Hong
ACM Trans. Multim. Comput. Commun. Appl.3
2025 Beyond Statistical Correlation: Causal Insights into Emotion Recognition
abstract
Emotion recognition has gained significant attention recently due to its wide-ranging applications like human-computer interaction, affective computing, and social robotics. Despite the promising results, critical issues need to be addressed. One primary challenge is that existing models typically establish spurious statistical correlations between the input image and the label instead of investigating causal-and-effect relationships. Besides, some similar emotional states often appear simultaneously, rendering the model unable to establish spurious correlations between these similar labels empirically. To address these issues, we develop a Dual-Disentanglement Causal Learning (D2CL) framework consisting of two disentanglement modules: a Feature Disentanglement module and a Label Disentanglement module. The first one extracts emotion-related representation and context-specific embedding to explore the underlying causal relationships between facial features and emotions, thereby mitigating spurious correlations. The second proposes a Feature Similarity-based classifier to capture subtle distinctions between similar labels. Extensive experiments on the EMOTIC and JAFFE datasets validate the effectiveness and superiority of the proposed method.
Tao Chen 0017, Yanrong Guo, Shijie Hao, Richang Hong
ICME2
2025 InterMind: Doctor-Patient-Family Interactive Depression Assessment Empowered by Large Language Models
abstract
Depression poses significant challenges to patients and healthcare organizations, necessitating efficient assessment methods. Existing paradigms typically focus on a patient-doctor way that overlooks multi-role interactions, such as family involvement in the evaluation and caregiving process. Moreover, current automatic depression detection (ADD) methods usually model depression detection as a classification or regression task, lacking interpretability for the decision-making process. To address these issues, we developed InterMind, a doctor-patient-family interactive depression assessment system empowered by large language models (LLMs). Our system enables patients and families to contribute descriptions, generates assistive diagnostic reports for doctors, and provides actionable insights, improving diagnostic precision and efficiency. To enhance LLMs' performance in psychological counseling and diagnostic interpretability, we integrate retrieval-augmented generation (RAG) and chain-of-thoughts (CoT) techniques for data augmentation, which mitigates the hallucination issue of LLMs in specific scenarios after instruction fine-tuning. Quantitative experiments and professional assessments by clinicians validate the effectiveness of our system.
Sanwang Wang, Shijie Hao, Yanrong Guo, Richang Hong
ACM Multimedia5
2025 BiLLIE: Toward Smooth Binarization of Low-Light Image Enhancement
Shijie Hao, Yanrong Guo, Richang Hong, Meng Wang 0001
PRCV (8)3
2025 Leaving None Behind: Data-Free Domain Incremental Learning for Major Depressive Disorder Detection
abstract
While deep learning techniques have shown promising performance in the Major Depressive Disorder (MDD) detection task, they still face limitations in real-world scenarios. Specifically, given the data scarcity, some efforts have resorted to aggregating data from different domains to expand the data volume. However, their effectiveness is currently limited by the domain gap and data privacy. Additionally, the class imbalance issue is particularly severe in our application, leading to biased classifying performance accordingly. To address these challenges, we propose Data-Free Domain Incremental Learning for the MDD detection (DIL-MDD) task, accommodating multiple feature distributions by only accessing well-trained models from previous domains and the data in the current domain. Specifically, DIL-MDD consists of two key modules: Adaptive Class-tailored Threshold Learning (ACTL) and Data-Free Domain Alignment (DFDA). The first module measures the discrepancy between the outputs of two sequential domains, based on which we learn a class-tailored threshold adaptively. Building on this, we differentiate between samples that either exhibit similarities or dissimilarities with the previous domain, where this similar sample set is identified to investigate the feature distribution of the historical data. The second module imposes an alignment constraint to narrow the gap between these two sample sets, thereby exploring the expertise of the previous domain. To validate the effectiveness of the proposed method, we conduct extensive experiments on the public MDD datasets, i.e., DAIC-WOZ, MODMA, and CMDC. We also apply our method to another mental health condition, Autism Spectrum Disorder (ASD), to further demonstrate its applicability. Finally, the ablation studies validate the superiority of the proposed modules.
Tao Chen 0017, Yanrong Guo, Shijie Hao, Richang Hong
IEEE Trans. Affect. Comput.2
2025 Person Re-Identification With Arbitrary Modalities: A Multi-Modal Dataset and a Unified Framework
abstract
This paper proposes a unified visual person re-identification (re-id) framework capable of handling various re-id tasks, including modal-fusion re-id, cross-modal re-id, and single-modal re-id, to accommodate diverse modal scenarios. We begin by constructing a Multi-modal Person Re-identification (MPR) dataset comprising RGB, infrared (IR), and depth modalities. Then, the unified re-id framework is established by integrating an Adaptive Modality Aggregation Module (AMAM) and Multi-modal Auto-aligned Learning (MAL). The former autonomously aggregates distinct modalities by thoroughly exploring their relationships. It not only benefits modal-fusion re-id by promoting the modal-fusion representations, but also enhances cross-modal re-id by performing modal consistency learning on the modal-fusion features to narrow modal gaps. The latter automatically aligns multiple modalities through contrastive learning constraints to lessen modal gaps for multiple cross-modal re-id tasks. So, these two modules respectively balance the tasks of distinct types and various tasks of the same type, which are beneficial to realize more re-id tasks with diverse modal scenarios. Moreover, we evaluate state-of-the-art (SOTA) multi-modal methods in terms of plentiful testing settings constructed on MPR dataset. The experiments demonstrate that the proposed unified method that only needs to be trained once outperforms existing methods that require multiple training processes with specific modalities. Besides, it can cope with more scenarios. Extensive ablation studies investigate the effects of the proposed modules on all re-id tasks. Our datasets and code will be publicly available soon: https://github.com/hfutwujingjing/A-Multi-Modal-Dataset-and-A-Unified-Framework.
Jingjing Wu 0001, Zhun Zhong, Yanrong Guo, Shejiao Hu, Richang Hong
IEEE Trans. Inf. Forensics Secur.3
2025 Multi-Modal Depression Detection in Interview via Exploring Emotional Distribution Information
abstract
In recent years, automatic depression detection (ADD) technology has been rapidly developed to boost an objective and assistive diagnosis for major depressive disorder (MDD) with the help of artificial intelligence technology and various physiological and psychological data. Despite emotion being an important reflection of mental status and frequently related to depression symptoms, few recent multi-modal ADD methods take emotional information into account. To address the above issue, we propose to explore emotional distribution information in interviews to assist multi-modal ADD model. On one hand, we use large language models (LLMs) to automatically recognize emotion of text data, and re-organize the data guided by the valence attribute of emotion, which facilitates our model being aware of difference in emotion distribution. On the other hand, we design the emotion encoding which enhances the proposed model to consider the emotional distribution information in its decision-making process. Extensive experiments are conducted by comparing with state-of-the-art ADD methods as well as the ablation study on different modules of the proposed method. More importantly, our experimental results can confirm the research findings in the psychology field, where more attention on negative emotion information is demanded in distinguishing different depressive status.
Yanrong Guo, Shijie Hao, Richang Hong
IEEE Trans. Multim.2
2025 Real-Time Semantic Segmentation via Spatial-Detail Guided Context Propagation
abstract
Nowadays, vision-based computing tasks play an important role in various real-world applications. However, many vision computing tasks, e.g., semantic segmentation, are usually computationally expensive, posing a challenge to the computing systems that are resource-constrained but require fast response speed. Therefore, it is valuable to develop accurate and real-time vision processing models that only require limited computational resources. To this end, we propose the spatial-detail guided context propagation network (SGCPNet) for achieving real-time semantic segmentation. In SGCPNet, we propose the strategy of spatial-detail guided context propagation. It uses the spatial details of shallow layers to guide the propagation of the low-resolution global contexts, in which the lost spatial information can be effectively reconstructed. In this way, the need for maintaining high-resolution features along the network is freed, therefore largely improving the model efficiency. On the other hand, due to the effective reconstruction of spatial details, the segmentation accuracy can be still preserved. In the experiments, we validate the effectiveness and efficiency of the proposed SGCPNet model. On the Cityscapes dataset, for example, our SGCPNet achieves 69.5% mIoU segmentation accuracy, while its speed reaches 178.5 FPS on 768 1536 images on a GeForce GTX 1080 Ti GPU card. In addition, SGCPNet is very lightweight and only contains 0.61 M parameters. The code will be released at https://github.com/zhouyuan888888/SGCPNet.
Shijie Hao, Yuan Zhou 0016, Yanrong Guo, Richang Hong, Jun Cheng 0002, Meng Wang 0001
IEEE Trans. Neural Networks Learn. Syst.3
2024 Advancing Incremental Few-Shot Semantic Segmentation via Semantic-Guided Relation Alignment and Adaptation
Yuan Zhou 0016, Xin Chen 0033, Yanrong Guo, Jun Yu 0002, Richang Hong, Qi Tian 0001
MMM (1)3
2024 Cascade Large Language Model via In-Context Learning for Depression Detection on Chinese Social Media
Yanrong Guo, Richang Hong
PRCV (1)2
2024 Multilevel depression status detection based on fine-grained prompt learning
Yanrong Guo
Pattern Recognit. Lett.2
2024 A Prompt-Based Topic-Modeling Method for Depression Detection on Low-Resource Data
abstract
Depression has a large impact on one’s personal life, especially during the COVID-19 pandemic. People have been trying to develop reliable methods for the depression detection task. Recently, methods based on deep learning have attracted much attention from the research community. However, they still face the challenge that data collection and annotation are difficult and expensive. In many real-world applications, only a small number of or even no training data are available. In this context, we propose a Prompt-based Topic-modeling method for Depression Detection (PTDD) on low-resource data, aiming to establish an effective way of depression detection under the above challenging situation. Instead of learning discriminating features from a small amount of labeled data, the proposed framework turns to leverage the generalization power of pretrained language models. Specifically, based on the question-and-answer routine during the interview, we first reorganize the text data according to the predefined topics for each interviewee. Via the prompt-based framework, we then predict whether the next-sentence prompt is emotionally positive or not. Finally, the depression detection task can be achieved based on the obtained topicwise predictions through a simple voting process. In the experiments, we validate the effectiveness of our model under several low-resource data settings. The results and analysis demonstrate that our PTDD achieves acceptable performance when only a few training samples or even no training samples are available.
Yanrong Guo, Lei Wang 0185, Shijie Hao, Richang Hong
IEEE Trans. Comput. Soc. Syst.1
2024 Semi-Supervised Domain Adaptation for Major Depressive Disorder Detection
abstract
Major Depressive Disorder (MDD) detection with cross-domain datasets is a crucial yet challenging application due to thedata scarcityandisolated data islandissues in multimedia computing research. Given the domain shift issue in MDD datasets and a continuous stream of incoming data in clinical settings, Semi-supervised Domain Adaptation (SDA) is suitable for addressing these challenges in MDD detection. However, existing mainstream Domain Adaptation (DA) methods have the following limitations that still need to be addressed, such as semantic misalignment, challenges in extending to various DA paradigms, and difficulty in addressing the classifier bias caused by class imbalance issues. To relieve the above issues, we propose a flexibleGraphNeuralNetwork-basedSemi-supervisedDomainAdaptation (GNN-SDA) for MDD detection. The proposed framework comprises a feature extraction backbone along with two essential modules: a GNN-based domain alignment module and an uncertainty-guided optimization module. The GNN-based domain alignment module is designed to reduce the domain gap in a flexible manner, which is able to align multiple domains through the information propagation mechanism instead of the explicit alignment operation. The uncertainty-guided optimization module discusses the uncertainty of pseudo-labels, mitigating the adverse impact of noisy predictions and taking into account the class distribution of unlabeled data. Finally, we evaluate the proposed GNN-SDA framework for MDD detection under different domain adaptation paradigms on four benchmark datasets, i.e., DAIC-WOZ, EATD, CMDC, and MODMA. The promising results indicate the flexibility and effectiveness of the proposed framework for MDD detection.
Tao Chen 0017, Yanrong Guo, Shijie Hao, Richang Hong
IEEE Trans. Multim.2
2023 Low-Light Image Enhancement Based on Mutual Guidance Between Enhancing Strength and Image Appearance
Linlin Hu, Shijie Hao, Yanrong Guo, Richang Hong, Meng Wang 0001
PRCV (11)3
2023 Few-Shot Partial Multi-View Learning
abstract
It is often the case that data are with multiple views in real-world applications. Fully exploring the information of each view is significant for making data more representative. However, due to various limitations and failures in data collection and pre-processing, it is inevitable for real data to suffer from view missing and data scarcity. The coexistence of these two issues makes it more challenging to achieve the pattern classification task. Currently, to our best knowledge, few appropriate methods can well-handle these two issues simultaneously. Aiming to draw more attention from the community to this challenge, we propose a new task in this paper, called few-shot partial multi-view learning, which focuses on overcoming the negative impact of the view-missing issue in the low-data regime. The challenges of this task are twofold: (i) it is difficult to overcome the impact of data scarcity under the interference of missing views; (ii) the limited number of data exacerbates information scarcity, thus making it harder to address the view-missing issue in turn. To address these challenges, we propose a new unified Gaussian dense-anchoring method. The unified dense anchors are learned for the limited partial multi-view data, thereby anchoring them into a unified dense representation space where the influence of data scarcity and view missing can be alleviated. We conduct extensive experiments to evaluate our method. The results on Cub-googlenet-doc2vec, Handwritten, Caltech102, Scene15, Animal, ORL, tieredImagenet, and Birds-200-2011 datasets validate its effectiveness. The codes will be released at https://github.com/zhouyuan888888/UGDA.
Yuan Zhou 0016, Yanrong Guo, Shijie Hao, Richang Hong, Jiebo Luo 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Automatic Depression Detection via Learning and Fusing Features From Visual Cues
abstract
Depression is one of the most prevalent mental disorders, which seriously affects one’s life. Traditional depression diagnostics commonly depend on rating with scales, which can be labor-intensive and subjective. In this context, automatic depression detection (ADD), aiming to assist medical experts in their diagnosis and analysis, has been attracting more attention for its better objectivity and fewer laborious interventions. A typical ADD model detects depression via automatically extracting task-specific features from medical records, such as video sequences, and sending them into a classifier for assistive prediction. However, it remains challenging to effectively extract depression-specific information from long sequences, thereby hindering a satisfying accuracy. In this article, we propose a novel ADD method via learning and fusing features from visual cues. Specifically, we first construct temporal dilated convolutional network (TDCN), in which multiple dilated convolution blocks (DCBs) are designed and stacked, to learn the long-range temporal information from sequences. Then, the featurewise attention (FWA) module is adopted to fuse different features extracted from TDCNs. The module learns to assign weights for the feature channels, aiming to better incorporate different kinds of visual features and further enhance the detection accuracy. Our method achieves the state-of-the-art performance on the Distress Analysis Interview Corpus Wizard-of-Oz (DAIC_WOZ) dataset compared with other visual-feature-based methods, showing its effectiveness.
Yanrong Guo, Chenyang Zhu 0005, Shijie Hao, Richang Hong
IEEE Trans. Comput. Soc. Syst.1
2023 Hierarchical Multifeature Fusion via Audio-Response-Level Modeling for Depression Detection
abstract
The clinical diagnosis of major depressive disorder (MDD) relying heavily on the subjective judgment assisted by questionnaires, could result in a low detection rate of MDD. Automatic depression detection (ADD) technology based on physiological and psychological information provides an objective and quantitative way for MDD detection. As a useful data modality, audio signals have attracted increasing interest in mental disorder detection. However, most recent audio-based depression detection methods underestimate the importance of subtly organizing audio data. They either simply use equally split audio segments, or directly build the model upon the entire data sequence, which pose challenges to learning task-specific features. To address this issue, we propose to reorganize the audio data at response level. Based on that, we construct a novel end-to-end model that hierarchically learns discriminative features for accurate depression detection. The stages of intraresponse fusion and interresponse fusion facilitate the extraction and aggregation of ADD-specific information from multiple kinds of acoustic features. Experimental results show that our proposed method significantly outperforms other state-of-the-art audio-based methods. In addition, the flexibility and the robustness of our model are also validated.
Yanrong Guo, Shijie Hao, Richang Hong
IEEE Trans. Comput. Soc. Syst.2
2023 MS²-GNN: Exploring GNN-Based Multimodal Fusion Network for Depression Detection
abstract
Major depressive disorder (MDD) is one of the most common and severe mental illnesses, posing a huge burden on society and families. Recently, some multimodal methods have been proposed to learn a multimodal embedding for MDD detection and achieved promising performance. However, these methods ignore the heterogeneity/homogeneity among various modalities. Besides, earlier attempts ignore interclass separability and intraclass compactness. Inspired by the above observations, we propose a graph neural network (GNN)-based multimodal fusion strategy named modal-shared modal-specific GNN, which investigates the heterogeneity/homogeneity among various psychophysiological modalities as well as explores the potential relationship between subjects. Specifically, we develop a modal-shared and modal-specific GNN architecture to extract the inter/intramodal characteristics. Furthermore, a reconstruction network is employed to ensure fidelity within the individual modality. Moreover, we impose an attention mechanism on various embeddings to obtain a multimodal compact representation for the subsequent MDD detection task. We conduct extensive experiments on two public depression datasets and the favorable results demonstrate the effectiveness of the proposed algorithm.
Tao Chen 0017, Richang Hong, Yanrong Guo, Shijie Hao, Bin Hu 0001
IEEE Trans. Cybern.3
2022 Adaptive dictionary and structure learning for unsupervised feature selection
Yanrong Guo, Shijie Hao
Inf. Process. Manag.1
2022 Exploring Self-Attention Graph Pooling With EEG-Based Topological Structure and Soft Label for Depression Detection
abstract
Electroencephalogram (EEG) has been widely used in neurological disease detection, i.e., major depressive disorder (MDD). Recently, some deep EEG-based MDD detection attempts have been proposed and achieved promising performance. These works, however, still suffer from the following limitations, such as insufficient exploration of the EEG-based topological structure, information loss caused by high-dimensional data compression, and under-estimation of intra-class difference and inter-class similarity. To solve these issues, we propose an EEG-based MDD detection model namedSelf-attentionGraphPooling withSoftLabel (SGP-SL). Specifically, we explore the local and global connections among EEG channels to construct an EEG-based graph in advance. By leveraging multiple self-attention graph pooling modules, the constructed graph is then gradually refined, followed by graph pooling, to aggregate information from less-important nodes to more-important ones. In this way, the feature representation with better discriminability can be learned from EEG signals. In addition, the soft label strategy is also adopted to build the loss function, aiming to further enhance the feature discriminability. Experimental results on the MODMA dataset demonstrate the superiority of the proposed method. What's more, extensive ablation studies are conducted to verify the effectiveness of the proposed elements in our model.
Tao Chen 0017, Yanrong Guo, Shijie Hao, Richang Hong
IEEE Trans. Affect. Comput.2
2022 Hierarchical Prototype Refinement With Progressive Inter-Categorical Discrimination Maximization for Few-Shot Learning
abstract
Metric-based few-shot learning categorizes unseen query instances by measuring their distance to the categories appearing in the given support set. To facilitate distance measurement, prototypes are used to approximate the representations of categories. However, we find prototypical representations are generally not discriminative enough to represent the discrepancy of inter-categorical distribution of queries, thereby limiting the classification accuracy. To overcome this issue, we propose a new Progressive Hierarchical-Refinement (PHR) method, which effectively refines the discrimination of prototypes by conducting the Progressive Discrimination Maximization strategy based on the hierarchical feature representations. Specifically, we first encode supports and queries into the representation space of spatial level, global level, and semantic level. Then, the refining coefficients are constructed by exploring the metric information contained in these hierarchical embedding spaces simultaneously. Under the guidance of the refining coefficients, the meta-refining loss progressively maximizes the discrimination degree of inter-categorical prototypical representations. In addition, the refining vectors are adopted to further enhance the representations of prototypes. In this way, the metric-based classification can be more accurate. Our PHR method shows the competitive performance on the miniImagenet, CIFAR-FS, FC100, and CUB datasets. Moreover, PHR presents good compatibility. It can be incorporated with other few-shot learning models, making them more accurate.
Yuan Zhou 0016, Yanrong Guo, Shijie Hao, Richang Hong
IEEE Trans. Image Process.2
2022 Decoupled Low-Light Image Enhancement
abstract
The visual quality of photographs taken under imperfect lightness conditions can be degenerated by multiple factors, e.g., low lightness, imaging noise, color distortion, and so on. Current low-light image enhancement models focus on the improvement of low lightness only, or simply deal with all the degeneration factors as a whole, therefore leading to sub-optimal results. In this article, we propose to decouple the enhancement model into two sequential stages. The first stage focuses on improving the scene visibility based on a pixel-wise non-linear mapping. The second stage focuses on improving the appearance fidelity by suppressing the rest degeneration factors. The decoupled model facilitates the enhancement in two aspects. On the one hand, the whole low-light enhancement can be divided into two easier subtasks. The first one only aims to enhance the visibility. It also helps to bridge the large intensity gap between the low-light and normal-light images. In this way, the second subtask can be described as the local appearance adjustment. On the other hand, since the parameter matrix learned from the first stage is aware of the lightness distribution and the scene structure, it can be incorporated into the second stage as the complementary information. In the experiments, our model demonstrates the state-of-the-art performance in both qualitative and quantitative comparisons, compared with other low-light image enhancement models. In addition, the ablation studies also validate the effectiveness of our model in multiple aspects, such as model structure and loss function.
Shijie Hao, Yanrong Guo, Meng Wang 0001
ACM Trans. Multim. Comput. Commun. Appl.3
2021 Robust pencil drawing generation via fast Retinex decomposition
Teng Li 0019, Shijie Hao, Yanrong Guo
Comput. Graph.3
2021 Editorial deep multi-source data analysis
Shichao Zhang 0001, Qing Xie 0002, Yanrong Guo
Pattern Recognit. Lett.3
2021 Adaptive Multi-Task Dual-Structured Learning with Its Application on Alzheimer's Disease Study
abstract
Multi-task learning has been widely applied to Alzheimer’s Disease (AD) studies due to its capability of simultaneously rating the disease severity (classification) and predicting corresponding clinical scores (regression). In this article, we propose a novel technique of Adaptive Multi-task Dual-Structured Learning, named AMDSL, by mutually exploring the dual manifold structure for the label and regression score of the disease data under joint classification and regression tasks, while learning an adaptive shared similarity measure and corresponding feature mapping among these two tasks. We encode both the reconstructed label representation and regression score adaptive to the ideal similarity measure on disease data to achieve the ideal performance on these two joint tasks. The alternating algorithm is proposed to optimize the above objective. We theoretically prove the convergence of the optimization algorithm. The superiority of AMDSL is experimentally validated under joint classification and regression as per various evaluation metrics against the most authoritative Alzheimer’s disease data.
Shijie Hao, Tao Chen 0017, Yang Wang 0023, Yanrong Guo, Meng Wang 0001
ACM Trans. Internet Techn.4
2020 Learning longitudinal classification-regression model for infant hippocampus segmentation
Yanrong Guo, Zhengwang Wu, Dinggang Shen
Neurocomputing1
2020 A Brief Survey on Semantic Segmentation with Deep Learning
Shijie Hao, Yuan Zhou 0016, Yanrong Guo
Neurocomputing3
2020 Multi-class multimodal semantic segmentation with an improved 3D fully convolutional networks
Yanrong Guo
Neurocomputing2
2020 Unsupervised video summarization via clustering validity index
Ye Zhao 0001, Yanrong Guo, Rui Sun 0004, Zhengqiong Liu, Dan Guo 0001
Multim. Tools Appl.2
2020 Unsupervised feature selection based on joint spectral learning and general sparse regression
Tao Chen 0017, Yanrong Guo, Shijie Hao
Neural Comput. Appl.2
2020 Sparsity-regularized feature selection for multi-class remote sensing image classification
Tao Chen 0017, Ye Zhao 0001, Yanrong Guo
Neural Comput. Appl.3
2020 Low-Light Image Enhancement With Semi-Decoupled Decomposition
abstract
Low-light image enhancement is important for high-quality image display and other visual applications. However, it is a challenging task as the enhancement is expected to improve the visibility of an image while keeping its visual naturalness. Retinex-based methods have well been recognized as a representative technique for this task, but they still have the following limitations. First, due to less-effective image decomposition or strong imaging noise, various artifacts can still be brought into enhanced results. Second, although the priori information can be explored to partially solve the first issue, it requires to carefully model the priori by a regularization term and usually makes the optimization process complicated. In this paper, we address these issues by proposing a novel Retinex-based low-light image enhancement method, in which the Retinex image decomposition is achieved in an efficient semi-decoupled way. Specifically, the illumination layer I is gradually estimated only with the input image S based on the proposed Gaussian Total Variation model, while the reflectance layer R is jointly estimated by S and the intermediate I. In addition, the imaging noise can be simultaneously suppressed during the estimation of R. Experimental results on several public datasets demonstrate that our method produces images with both higher visibility and better visual quality, which outperforms the state-of-the-art low-light enhancement methods in terms of several objective and subjective evaluation metrics.
Shijie Hao, Yanrong Guo, Xin Xu 0007, Meng Wang 0001
IEEE Trans. Multim.3
2020 Multi-Atlas Segmentation of Anatomical Brain Structures Using Hierarchical Hypergraph Learning
abstract
Accurate segmentation of anatomical brain structures is crucial for many neuroimaging applications, e.g., early brain development studies and the study of imaging biomarkers of neurodegenerative diseases. Although multi-atlas segmentation (MAS) has achieved many successes in the medical imaging area, this approach encounters limitations in segmenting anatomical structures associated with poor image contrast. To address this issue, we propose a new MAS method that uses a hypergraph learning framework to model the complex subject-within and subject-to-atlas image voxel relationships and propagate the label on the atlas image to the target subject image. To alleviate the low-image contrast issue, we propose two strategies equipped with our hypergraph learning framework. First, we use a hierarchical strategy that exploits high-level context features for hypergraph construction. Because the context features are computed on the tentatively estimated probability maps, we can ultimately turn the hypergraph learning into a hierarchical model. Second, instead of only propagating the labels from the atlas images to the target subject image, we use a dynamic label propagation strategy that can gradually use increasing reliably identified labels from the subject image to aid in predicting the labels on the difficult-to-label subject image voxels. Compared with the state-of-the-art label fusion methods, our results show that the hierarchical hypergraph learning framework can substantially improve the robustness and accuracy in the segmentation of anatomical brain structures with low image contrast from magnetic resonance (MR) images.
Pei Dong, Yanrong Guo, Yue Gao 0002, Peipeng Liang, Yonghong Shi, Guorong Wu 0001
IEEE Trans. Neural Networks Learn. Syst.2
2019 Lightness-aware contrast enhancement for images with different illumination conditions
Shijie Hao, Yanrong Guo, Zhongliang Wei
Multim. Tools Appl.2
2018 Robust brain ROI segmentation by deformation regression and deformable shape model
Zhengwang Wu, Yanrong Guo, Sanghyun Park 0004, Yaozong Gao, Pei Dong, Seong-Whan Lee, Dinggang Shen
Medical Image Anal.2
2018 Low-light image enhancement with a refined illumination map
Shijie Hao, Zhuang Feng, Yanrong Guo
Multim. Tools Appl.3
2018 Semantic segmentation of RGBD images based on deep depth regression
Yanrong Guo, Tao Chen 0017
Pattern Recognit. Lett.1
2017 Dual-core steered non-rigid registration for multi-modal images via bi-directional image synthesis
Xiaohuan Cao, Jianhua Yang 0005, Yaozong Gao, Yanrong Guo, Guorong Wu 0001, Dinggang Shen
Medical Image Anal.4
2016 Identifying Patients at Risk for Aortic Stenosis Through Learning from Multimodal Data
abstract
In this paper we present a new method of uncovering patients with aortic valve diseases in large electronic health record systems through learning with multimodal data. The method automatically extracts clinically-relevant valvular disease features from five multimodal sources of information including structured diagnosis, echocardiogram reports, and echocardiogram imaging studies. It combines these partial evidence features in a random forests learning framework to predict patients likely to have the disease. Results of a retrospective clinical study from a 1000 patient dataset are presented that indicate that over 25 % new patients with moderate to severe aortic stenosis can be automatically discovered by our method that were previously missed from the records. These keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves.
Tanveer F. Syeda-Mahmood, Yanrong Guo, Mehdi Moradi, David Beymer, Deepta Rajan, Yaniv Gur, Mohammadreza Negahdar
MICCAI (3)2
2016 Image detail enhancement with spatially guided filters
Shijie Hao, Daru Pan, Yanrong Guo, Richang Hong, Meng Wang 0001
Signal Process.3
2016 Deformable MR Prostate Segmentation via Deep Feature Learning and Sparse Patch Matching
abstract
Automatic and reliable segmentation of the prostate is an important but difficult task for various clinical applications such as prostate cancer radiotherapy. The main challenges for accurate MR prostate localization lie in two aspects: (1) inhomogeneous and inconsistent appearance around prostate boundary, and (2) the large shape variation across different patients. To tackle these two problems, we propose a new deformable MR prostate segmentation method by unifying deep feature learning with the sparse patch matching. First, instead of directly using handcrafted features, we propose to learn the latent feature representation from prostate MR images by the stacked sparse auto-encoder (SSAE). Since the deep learning algorithm learns the feature hierarchy from the data, the learned features are often more concise and effective than the handcrafted features in describing the underlying data. To improve the discriminability of learned features, we further refine the feature representation in a supervised fashion. Second, based on the learned features, a sparse patch matching method is proposed to infer a prostate likelihood map by transferring the prostate labels from multiple atlases to the new prostate MR image. Finally, a deformable segmentation is used to integrate a sparse shape model with the prostate likelihood map for achieving the final segmentation. The proposed method has been extensively evaluated on the dataset that contains 66 T2-wighted prostate MR images. Experimental results show that the deep-learned features are more effective than the handcrafted features in guiding MR prostate segmentation. Moreover, our method shows superior performance than other state-of-the-art segmentation methods.
Yanrong Guo, Yaozong Gao, Dinggang Shen
IEEE Trans. Medical Imaging1
2015 Segmentation of Infant Hippocampus Using Common Feature Representations Learned for Multimodal Longitudinal Data
Yanrong Guo, Guorong Wu 0001, Pew-Thian Yap, Valerie Jewells, Weili Lin, Dinggang Shen
MICCAI (3)1
2015 Building dynamic population graph for accurate correspondence detection
Shaoyi Du, Yanrong Guo, Gerard Sanroma, Dong Ni 0001, Guorong Wu 0001, Dinggang Shen
Medical Image Anal.2
2015 A transversal approach for patch-based label fusion via matrix completion
Gerard Sanroma, Guorong Wu 0001, Yaozong Gao, Kim-Han Thung, Yanrong Guo, Dinggang Shen
Medical Image Anal.5
2014 Segmenting Hippocampus from Infant Brains by Sparse Patch Matching with Deep-Learned Features
Yanrong Guo, Guorong Wu 0001, Leah A. Commander, Stephanie Szary, Valerie Jewells, Weili Lin, Dinggang Shen
MICCAI (2)1
2014 Hierarchical Lung Field Segmentation With Joint Shape and Appearance Sparse Learning
abstract
Lung field segmentation in the posterior-anterior (PA) chest radiograph is important for pulmonary disease diagnosis and hemodialysis treatment. Due to high shape variation and boundary ambiguity, accurate lung field segmentation from chest radiograph is still a challenging task. To tackle these challenges, we propose a joint shape and appearance sparse learning method for robust and accurate lung field segmentation. The main contributions of this paper are: 1) a robust shape initialization method is designed to achieve an initial shape that is close to the lung boundary under segmentation; 2) a set of local sparse shape composition models are built based on local lung shape segments to overcome the high shape variations; 3) a set of local appearance models are similarly adopted by using sparse representation to capture the appearance characteristics in local lung boundary segments, thus effectively dealing with the lung boundary ambiguity; 4) a hierarchical deformable segmentation framework is proposed to integrate the scale-dependent shape and appearance information together for robust and accurate segmentation. Our method is evaluated on 247 PA chest radiographs in a public dataset. The experimental results show that the proposed local shape and appearance models outperform the conventional shape and appearance models. Compared with most of the state-of-the-art lung field segmentation methods under comparison, our method also shows a higher accuracy, which is comparable to the inter-observer annotation variation.
Yeqin Shao, Yaozong Gao, Yanrong Guo, Yonghong Shi, Xin Yang 0009, Dinggang Shen
IEEE Trans. Medical Imaging3
2013 Active learning based intervertebral disk classification combining shape and texture similarities
Shijie Hao, Yanrong Guo
Neurocomputing3
2013 ε-Isometryε-Isometry based shape approximation for image content representation
Shijie Hao, Yanrong Guo, Shu Zhan
Signal Process.3
2013 Robust Anatomical Correspondence Detection by Hierarchical Sparse Graph Matching
abstract
Robust anatomical correspondence detection is a key step in many medical image applications such as image registration and motion correction. In the computer vision field, graph matching techniques have emerged as a powerful approach for correspondence detection. By considering potential correspondences as graph nodes, graph edges can be used to measure the pairwise agreement between possible correspondences. In this paper, we present a novel, hierarchical graph matching method with sparsity constraint to further augment the power of conventional graph matching methods in establishing anatomical correspondences, especially for the cases of large inter-subject variations in medical applications. Specifically, we first propose to measure the pairwise agreement between potential correspondences along a sequence of intensity profiles which reduces the ambiguity in correspondence matching. We next introduce the concept of sparsity on the fuzziness of correspondences to suppress the distraction from misleading matches, which is very important for achieving the accurate, one-to-one correspondences. Finally, we integrate our graph matching method into a hierarchical correspondence matching framework, where we use multiple models to deal with the large inter-subject anatomical variations and gradually refine the correspondence matching results between the tentatively deformed model images and the underlying subject image. Evaluations on both synthetic data and public hand X-ray images indicate that the proposed hierarchical sparse graph matching method yields the best correspondence matching performance in terms of both accuracy and robustness when compared with several conventional graph matching methods.
Yanrong Guo, Guorong Wu 0001, Dinggang Shen
IEEE Trans. Medical Imaging1
2011 Distribution-Based Active Contour Model for Medical Image Segmentation
abstract
Having being regarded as one of the classical methods in image segmentation, geodesic active contours (GAC) have the flaws of boundary leaking and expensive evolving time. In this paper, we present a distribution-based active contour model by measuring the Bhattacharyya distance between probability distributions of the object and background along with the evolution of GAC model. Due to combining the image cues of edge and statistical information which is computed by using kernel density estimation, this hybrid methodology prevents the boundary leaking as well as the under segmentation problem. Experimental results on the medical images show the improvements of our method in terms of comparisons with original GAC model, Bhattacharyya gradient flow, texture-based GAC and Li's active contour model.
Yanrong Guo, Shijie Hao, Shu Zhan
ICIG1
2011 Intervertebral Disc Shape Analysis with Geodesic Metric in Shape Space
abstract
Shapes of anatomical structures extracted from medical imaging usually contain diagnostic and therapeutic cues in clinical applications. In this paper, we propose a framework on analyzing disc shapes based on a geodesic metric in an anatomical shape space. All disc shapes, containing both normal and abnormal ones, are formulated as elements in this space. The geodesic connecting these elements and other statistics are then numerically approximated. With these tools quantifying the intrinsic difference between disc shapes, the normal shapes from a dataset are unsupervisedly clustered and a statistical inference based on the learned Gaussian model is made. Experimental results show a reasonable accuracy of classifying normal and abnormal intervertebral discs.
Shijie Hao, Yanrong Guo, Shu Zhan
ICIG3