VLDB 2026 Research / reviewers in the wild / expert
Shijie Hao
dblp:122/3198
· DBLP profile ↗
70ranked-venue papers
17as first author
42since 2021 · last 2026
0000-0003-3181-1220ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 30 · 9 first-author · 19 since 2021Artificial intelligence and machine learning · 27 · 6 first-author · 15 since 2021Databases, data management, data science and information retrieval · 10 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 6 since 2021Computer networks · 2 · 2 first-author · 2 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Voices, Faces, and Feelings: Multi-modal Emotion-Cognition Captioning for Mental Health UnderstandingabstractEmotional and cognitive factors are essential for understanding mental health disorders. However, existing methods often treat multi-modal data as classification tasks, limiting interpretability especially for emotion and cognition. Although large language models (LLMs) offer opportunities for mental health analysis, they mainly rely on textual semantics and overlook fine-grained emotional and cognitive cues in multi-modal inputs. While some studies incorporate emotional features via transfer learning, their connection to mental health conditions remains implicit. To address these issues, we propose ECMC, a novel task that aims at generating natural language descriptions of emotional and cognitive states from multi-modal data, and producing emotion–cognition profiles that improve both the accuracy and interpretability of mental health assessments. We adopt an encoder–decoder architecture, where modality-specific encoders extract features, which are fused by a dual-stream BridgeNet based on Q-former. Contrastive learning enhances the extraction of emotional and cognitive features. A LLaMA decoder then aligns these features with annotated captions to produce detailed descriptions. Extensive objective and subjective evaluations demonstrate that: 1) ECMC outperforms existing multi-modal LLMs and mental health models in generating emotion–cognition captions; 2) the generated emotion–cognition profiles significantly improve assistive diagnosis and interpretability in mental health analysis. Yanrong Guo, Shijie Hao |
AAAI | 3 |
| 2026 | Illumination-Prior Guided Hybrid Network for Low-Light Image Enhancement
Shijie Hao, Yanrong Guo |
MMM (1) | 2 |
| 2026 | Mamba-transformer for low-light image enhancement in HVI color space
Zepu Xu, Shijie Hao |
Multim. Syst. | 2 |
| 2026 | Interview-Based Depression Detection Using LLM-Based Text Restatement and Emotion LexiconabstractDepression is a mental health disorder that significantly impacts modern society. Developing accurate depression detection models by leveraging discriminant features from multimedia or physiological data can aid medical professionals in making informed diagnoses. According to psychological studies, emotion is a critical indicator of depression. However, emotion has not been utilized as a central role in current research on assistive depression detection, usually serving as a supplementary information source or a guidance for integrating diverse data modalities. In contrast to existing studies, we investigate the feasibility of detecting depression by concentrating on emotion information. Specifically, focusing on modeling emotion feature representation during interviews, we propose an interview-based depression detection model via leveraging large language (LLM) based text restatement and emotion lexicon (IDD-LTE). In this model, we employ LLM to enhance text quality through restatement to address the potentially low quality of interview text data. Using an emotion lexicon, the open contents in restated texts are mapped to a fixed-size matrix representation that captures the interviewee's emotional state and mood swings during the conversation, serving as the fundamental representation for the following discriminant feature learning. The proposed IDD-LTE model is evaluated on four primary datasets for depression detection. The promising results confirm the feasibility and effectiveness of our model. Shijie Hao, Jingjing Wu 0001, Yanrong Guo, Richang Hong |
IEEE Trans. Affect. Comput. | 1 |
| 2026 | Subthreshold Depression Detection With Text-Guided Multimodal LearningabstractDepression, a widespread global mental health problem, affects millions of people annually, making early detection of subclinical depression crucial for timely intervention. Current automatic depression detection (ADD) methods, valuable for diagnosis, often neglect subthreshold populations and face difficulties in extracting diagnostic data from long sequences of multimodal information. These methods also inadequately leverage text modality, which is less noisy and information-rich compared with other modalities. Furthermore, existing datasets for depression research are often too small, limiting the generalizability of developed methods. To address these issues, this article proposes a new approach for detecting depression in subthreshold populations. For long-sequence samples in the field of depression, we construct an autoencoder that compresses along both temporal and feature dimensions, aiming to extract the most compact and effective features from the samples. To exploit the text modality’s advantages, we integrate the RoBERTa pretrained model with an attention mechanism for high-quality text encoding. We then develop a text-guided multimodal fusion (TGMF) module, using text encoding as an anchor for guiding audio and video modality encoding, ensuring multimodal alignment. Additionally, contrastive learning is applied to discern differences between classes, enhancing the model’s generalizability. Our method demonstrates superior performance in the tasks of detecting depression and identifying subthreshold populations on the E-DAIC and MMDA datasets. Yanrong Guo, Youwei Guo, Bingxin Yang, Jingjing Wu 0001, Shijie Hao, Richang Hong |
IEEE Trans. Comput. Soc. Syst. | 5 |
| 2026 | Controllable Relation Disentanglement for Few-Shot Class-Incremental LearningabstractFew-Shot Class-Incremental Learning (FSCIL) requires models to be updated incrementally with limited labeled samples given in each session, differing from the traditional training paradigm and easily resulting in severe spurious relations between categories. Thus, in this paper, we propose to address FSCIL from a new perspective: enhancing FSCIL via disentangling spurious relations between categories. Accordingly, we propose a simple yet effective approach, dubbed ConTrollable Relation-disentangLed Few-Shot Class-Incremental Learning (CTRL-FSCIL). Specifically, during a base session, we propose to anchor base class embeddings in feature space and build disentangled proxies to bridge gaps between the learning processes of categories encountered in different sessions, making category relations controllable. Furthermore, during incremental learning, the parameters of the backbone network are frozen in order to relieve the negative impact of data scarcity. Meanwhile, a relation disentanglement loss is employed to guide a relation control module to disentangle spurious relations between learned categories. In this way, spurious relation issues in FSCIL can be alleviated. Extensive experiments on CIFAR-100, mini-ImageNet, and CUB-200 demonstrate the effectiveness of CTRL-FSCIL. Our code has been publicly released on github. Yuan Zhou 0016, Richang Hong, Yanrong Guo, Lin Liu 0016, Shijie Hao, Hanwang Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | Dynamic Correlation-Guided Disentanglement and Contrastive Learning for RGB-D Cross-Modal Re-IdentificationabstractPerson re-identification (Re-ID) across RGB and depth modalities offers complementary cues for robust pedestrian matching under challenging conditions. However, the significant discrepancy between RGB appearance features and depth structural features complicates cross-modal alignment. Existing methods either depend on static architectural designs or impose strong constraints to capture the common features of the two modalities, often suffering from branch imbalance or distorted identity features. In this work, we propose a novel framework, Dynamic Correlation-Guided Disentanglement and Contrastive Learning (DCG-DCL), for RGB-D cross-modal Re-ID. First, the Dynamic Correlation-guided Disentanglement (DCGD) dynamically decouples features with the guidance of inter-modal correlation, which explicitly enforces common-feature learning via a cross-correlation constraint and adaptively separates common and unique components without predefined assumptions. Second, a Common & Unique Contrastive Learning (CUCL) strategy fully leverages these decoupled features, which aligns RGB/depth features closer to their common representation and pushes them away from unique redundancies. This dual mechanism effectively narrows modality discrepancy and boosts robustness against modality-specific noise. Extensive experiments on multiple public benchmarks demonstrate that our method achieves state-of-the-art performance, with ablation studies validating the necessity of each component. Zhibo Lei, Jingjing Wu 0001, Yaxiong Wang, Yanrong Guo, Shijie Hao, Richang Hong |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2026 | Biomedical Relation Extraction via Adaptive Document-Relation Cross-Mapping and Concept Unique IdentifierabstractDocument-Level Biomedical Relation Extraction (Bio-RE) aims to identify relations between biomedical entities within extensive texts, serving as a crucial subfield of biomedical text mining. Existing Bio-RE methods struggle with cross-sentence inference, which is essential for capturing relations spanning multiple sentences. Moreover, previous methods often overlook the incompleteness of documents and lack the integration of external knowledge, limiting contextual richness. Besides, the scarcity of annotated data further hampers model training. Recent advancements in large language models (LLMs) have inspired us to explore all the above issues for document-level Bio-RE. Specifically, we propose a document-level Bio-RE framework via LLM Adaptive Document-Relation Cross-Mapping (ADRCM) Fine-Tuning and Concept Unique Identifier (CUI) Retrieval-Augmented Generation (RAG). First, we introduce the Iteration-of-REsummary (IoRs) prompt for solving the data scarcity issue. In this way, Bio-RE task-specific synthetic data can be generated by guiding ChatGPT to focus on entity relations and iteratively refining synthetic data. Next, we propose ADRCM fine-tuning, a novel fine-tuning recipe that establishes mappings across different documents and relations, enhancing the model’s contextual understanding and cross-sentence inference capabilities. Finally, during the inference, a biomedical-specific RAG approach, named CUI RAG, is designed to leverage CUIs as indexes for entities, narrowing the retrieval scope and enriching the relevant document contexts. Experiments conducted on three Bio-RE datasets—GDA, CDR, and BioRED—demonstrate the state-of-the-art performance of our proposed method by comparing it with other related works. Yufei Shang, Yanrong Guo, Shijie Hao, Richang Hong |
ACM Trans. Knowl. Discov. Data | 3 |
| 2025 | Gt-Mean Loss: a Simple Yet Effective Solution for Brightness Mismatch in Low-Light Image Enhancement
Jingxi Liao, Shijie Hao, Richang Hong, Meng Wang 0001 |
ICCV | 2 |
| 2025 | Beyond Statistical Correlation: Causal Insights into Emotion RecognitionabstractEmotion recognition has gained significant attention recently due to its wide-ranging applications like human-computer interaction, affective computing, and social robotics. Despite the promising results, critical issues need to be addressed. One primary challenge is that existing models typically establish spurious statistical correlations between the input image and the label instead of investigating causal-and-effect relationships. Besides, some similar emotional states often appear simultaneously, rendering the model unable to establish spurious correlations between these similar labels empirically. To address these issues, we develop a Dual-Disentanglement Causal Learning (D2CL) framework consisting of two disentanglement modules: a Feature Disentanglement module and a Label Disentanglement module. The first one extracts emotion-related representation and context-specific embedding to explore the underlying causal relationships between facial features and emotions, thereby mitigating spurious correlations. The second proposes a Feature Similarity-based classifier to capture subtle distinctions between similar labels. Extensive experiments on the EMOTIC and JAFFE datasets validate the effectiveness and superiority of the proposed method. Tao Chen 0017, Yanrong Guo, Shijie Hao, Richang Hong |
ICME | 3 |
| 2025 | InterMind: Doctor-Patient-Family Interactive Depression Assessment Empowered by Large Language ModelsabstractDepression poses significant challenges to patients and healthcare organizations, necessitating efficient assessment methods. Existing paradigms typically focus on a patient-doctor way that overlooks multi-role interactions, such as family involvement in the evaluation and caregiving process. Moreover, current automatic depression detection (ADD) methods usually model depression detection as a classification or regression task, lacking interpretability for the decision-making process. To address these issues, we developed InterMind, a doctor-patient-family interactive depression assessment system empowered by large language models (LLMs). Our system enables patients and families to contribute descriptions, generates assistive diagnostic reports for doctors, and provides actionable insights, improving diagnostic precision and efficiency. To enhance LLMs' performance in psychological counseling and diagnostic interpretability, we integrate retrieval-augmented generation (RAG) and chain-of-thoughts (CoT) techniques for data augmentation, which mitigates the hallucination issue of LLMs in specific scenarios after instruction fine-tuning. Quantitative experiments and professional assessments by clinicians validate the effectiveness of our system. Sanwang Wang, Shijie Hao, Yanrong Guo, Richang Hong |
ACM Multimedia | 4 |
| 2025 | LIESA: Low-Light Image Enhancement with Semantic Awareness
Shijie Hao, Fuming Sun |
MMM (2) | 2 |
| 2025 | SMT: SNR-Aware Mamba-Transformer for Low-Light Image Enhancement
Shijie Hao |
PRCV (8) | 2 |
| 2025 | BiLLIE: Toward Smooth Binarization of Low-Light Image Enhancement
Shijie Hao, Yanrong Guo, Richang Hong, Meng Wang 0001 |
PRCV (8) | 2 |
| 2025 | Leaving None Behind: Data-Free Domain Incremental Learning for Major Depressive Disorder DetectionabstractWhile deep learning techniques have shown promising performance in the Major Depressive Disorder (MDD) detection task, they still face limitations in real-world scenarios. Specifically, given the data scarcity, some efforts have resorted to aggregating data from different domains to expand the data volume. However, their effectiveness is currently limited by the domain gap and data privacy. Additionally, the class imbalance issue is particularly severe in our application, leading to biased classifying performance accordingly. To address these challenges, we propose Data-Free Domain Incremental Learning for the MDD detection (DIL-MDD) task, accommodating multiple feature distributions by only accessing well-trained models from previous domains and the data in the current domain. Specifically, DIL-MDD consists of two key modules: Adaptive Class-tailored Threshold Learning (ACTL) and Data-Free Domain Alignment (DFDA). The first module measures the discrepancy between the outputs of two sequential domains, based on which we learn a class-tailored threshold adaptively. Building on this, we differentiate between samples that either exhibit similarities or dissimilarities with the previous domain, where this similar sample set is identified to investigate the feature distribution of the historical data. The second module imposes an alignment constraint to narrow the gap between these two sample sets, thereby exploring the expertise of the previous domain. To validate the effectiveness of the proposed method, we conduct extensive experiments on the public MDD datasets, i.e., DAIC-WOZ, MODMA, and CMDC. We also apply our method to another mental health condition, Autism Spectrum Disorder (ASD), to further demonstrate its applicability. Finally, the ablation studies validate the superiority of the proposed modules. Tao Chen 0017, Yanrong Guo, Shijie Hao, Richang Hong |
IEEE Trans. Affect. Comput. | 3 |
| 2025 | Multi-Modal Depression Detection in Interview via Exploring Emotional Distribution InformationabstractIn recent years, automatic depression detection (ADD) technology has been rapidly developed to boost an objective and assistive diagnosis for major depressive disorder (MDD) with the help of artificial intelligence technology and various physiological and psychological data. Despite emotion being an important reflection of mental status and frequently related to depression symptoms, few recent multi-modal ADD methods take emotional information into account. To address the above issue, we propose to explore emotional distribution information in interviews to assist multi-modal ADD model. On one hand, we use large language models (LLMs) to automatically recognize emotion of text data, and re-organize the data guided by the valence attribute of emotion, which facilitates our model being aware of difference in emotion distribution. On the other hand, we design the emotion encoding which enhances the proposed model to consider the emotional distribution information in its decision-making process. Extensive experiments are conducted by comparing with state-of-the-art ADD methods as well as the ablation study on different modules of the proposed method. More importantly, our experimental results can confirm the research findings in the psychology field, where more attention on negative emotion information is demanded in distinguishing different depressive status. Yanrong Guo, Shijie Hao, Richang Hong |
IEEE Trans. Multim. | 3 |
| 2025 | Real-Time Semantic Segmentation via Spatial-Detail Guided Context PropagationabstractNowadays, vision-based computing tasks play an important role in various real-world applications. However, many vision computing tasks, e.g., semantic segmentation, are usually computationally expensive, posing a challenge to the computing systems that are resource-constrained but require fast response speed. Therefore, it is valuable to develop accurate and real-time vision processing models that only require limited computational resources. To this end, we propose the spatial-detail guided context propagation network (SGCPNet) for achieving real-time semantic segmentation. In SGCPNet, we propose the strategy of spatial-detail guided context propagation. It uses the spatial details of shallow layers to guide the propagation of the low-resolution global contexts, in which the lost spatial information can be effectively reconstructed. In this way, the need for maintaining high-resolution features along the network is freed, therefore largely improving the model efficiency. On the other hand, due to the effective reconstruction of spatial details, the segmentation accuracy can be still preserved. In the experiments, we validate the effectiveness and efficiency of the proposed SGCPNet model. On the Cityscapes dataset, for example, our SGCPNet achieves 69.5% mIoU segmentation accuracy, while its speed reaches 178.5 FPS on 768 1536 images on a GeForce GTX 1080 Ti GPU card. In addition, SGCPNet is very lightweight and only contains 0.61 M parameters. The code will be released at https://github.com/zhouyuan888888/SGCPNet. Shijie Hao, Yuan Zhou 0016, Yanrong Guo, Richang Hong, Jun Cheng 0002, Meng Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Pre-trained low-light image enhancement transformerabstractAbstract Low‐light image enhancement is a longstanding challenge in low‐level vision, as images captured in low‐light conditions often suffer from significant aesthetic quality flaws. Recent methods based on deep neural networks have made impressive progress in this area. In contrast to mainstream convolutional neural network (CNN)‐based methods, an effective solution inspired by the transformer, which has shown impressive performance in various tasks, is proposed. This solution is centred around two key components. The first is an image synthesis pipeline, and the second is a powerful transformer‐based pre‐trained model, known as the low‐light image enhancement transformer (LIET). The image synthesis pipeline includes illumination simulation and realistic noise simulation, enabling the generation of more life‐like low‐light images to overcome the issue of data scarcity. LIET combines streamlined CNN‐based encoder‐decoders with a transformer body, efficiently extracting global and local contextual features at a relatively low computational cost. The extensive experiments show that this approach is highly competitive with current state‐of‐the‐art methods. The codes have been released and are available at LIET . Shijie Hao |
IET Image Process. | 2 |
| 2024 | A lightness-aware loss for low-light image enhancement
Dian Xie, Huajun Xing, Shijie Hao |
Pattern Recognit. Lett. | 4 |
| 2024 | A Prompt-Based Topic-Modeling Method for Depression Detection on Low-Resource DataabstractDepression has a large impact on one’s personal life, especially during the COVID-19 pandemic. People have been trying to develop reliable methods for the depression detection task. Recently, methods based on deep learning have attracted much attention from the research community. However, they still face the challenge that data collection and annotation are difficult and expensive. In many real-world applications, only a small number of or even no training data are available. In this context, we propose a Prompt-based Topic-modeling method for Depression Detection (PTDD) on low-resource data, aiming to establish an effective way of depression detection under the above challenging situation. Instead of learning discriminating features from a small amount of labeled data, the proposed framework turns to leverage the generalization power of pretrained language models. Specifically, based on the question-and-answer routine during the interview, we first reorganize the text data according to the predefined topics for each interviewee. Via the prompt-based framework, we then predict whether the next-sentence prompt is emotionally positive or not. Finally, the depression detection task can be achieved based on the obtained topicwise predictions through a simple voting process. In the experiments, we validate the effectiveness of our model under several low-resource data settings. The results and analysis demonstrate that our PTDD achieves acceptable performance when only a few training samples or even no training samples are available. Yanrong Guo, Lei Wang 0185, Shijie Hao, Richang Hong |
IEEE Trans. Comput. Soc. Syst. | 5 |
| 2024 | Semi-Supervised Domain Adaptation for Major Depressive Disorder DetectionabstractMajor Depressive Disorder (MDD) detection with cross-domain datasets is a crucial yet challenging application due to thedata scarcityandisolated data islandissues in multimedia computing research. Given the domain shift issue in MDD datasets and a continuous stream of incoming data in clinical settings, Semi-supervised Domain Adaptation (SDA) is suitable for addressing these challenges in MDD detection. However, existing mainstream Domain Adaptation (DA) methods have the following limitations that still need to be addressed, such as semantic misalignment, challenges in extending to various DA paradigms, and difficulty in addressing the classifier bias caused by class imbalance issues. To relieve the above issues, we propose a flexibleGraphNeuralNetwork-basedSemi-supervisedDomainAdaptation (GNN-SDA) for MDD detection. The proposed framework comprises a feature extraction backbone along with two essential modules: a GNN-based domain alignment module and an uncertainty-guided optimization module. The GNN-based domain alignment module is designed to reduce the domain gap in a flexible manner, which is able to align multiple domains through the information propagation mechanism instead of the explicit alignment operation. The uncertainty-guided optimization module discusses the uncertainty of pseudo-labels, mitigating the adverse impact of noisy predictions and taking into account the class distribution of unlabeled data. Finally, we evaluate the proposed GNN-SDA framework for MDD detection under different domain adaptation paradigms on four benchmark datasets, i.e., DAIC-WOZ, EATD, CMDC, and MODMA. The promising results indicate the flexibility and effectiveness of the proposed framework for MDD detection. Tao Chen 0017, Yanrong Guo, Shijie Hao, Richang Hong |
IEEE Trans. Multim. | 3 |
| 2023 | Low-Light Image Enhancement Based on Mutual Guidance Between Enhancing Strength and Image Appearance
Linlin Hu, Shijie Hao, Yanrong Guo, Richang Hong, Meng Wang 0001 |
PRCV (11) | 2 |
| 2023 | Few-Shot Partial Multi-View LearningabstractIt is often the case that data are with multiple views in real-world applications. Fully exploring the information of each view is significant for making data more representative. However, due to various limitations and failures in data collection and pre-processing, it is inevitable for real data to suffer from view missing and data scarcity. The coexistence of these two issues makes it more challenging to achieve the pattern classification task. Currently, to our best knowledge, few appropriate methods can well-handle these two issues simultaneously. Aiming to draw more attention from the community to this challenge, we propose a new task in this paper, called few-shot partial multi-view learning, which focuses on overcoming the negative impact of the view-missing issue in the low-data regime. The challenges of this task are twofold: (i) it is difficult to overcome the impact of data scarcity under the interference of missing views; (ii) the limited number of data exacerbates information scarcity, thus making it harder to address the view-missing issue in turn. To address these challenges, we propose a new unified Gaussian dense-anchoring method. The unified dense anchors are learned for the limited partial multi-view data, thereby anchoring them into a unified dense representation space where the influence of data scarcity and view missing can be alleviated. We conduct extensive experiments to evaluate our method. The results on Cub-googlenet-doc2vec, Handwritten, Caltech102, Scene15, Animal, ORL, tieredImagenet, and Birds-200-2011 datasets validate its effectiveness. The codes will be released at https://github.com/zhouyuan888888/UGDA. Yuan Zhou 0016, Yanrong Guo, Shijie Hao, Richang Hong, Jiebo Luo 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Automatic Depression Detection via Learning and Fusing Features From Visual CuesabstractDepression is one of the most prevalent mental disorders, which seriously affects one’s life. Traditional depression diagnostics commonly depend on rating with scales, which can be labor-intensive and subjective. In this context, automatic depression detection (ADD), aiming to assist medical experts in their diagnosis and analysis, has been attracting more attention for its better objectivity and fewer laborious interventions. A typical ADD model detects depression via automatically extracting task-specific features from medical records, such as video sequences, and sending them into a classifier for assistive prediction. However, it remains challenging to effectively extract depression-specific information from long sequences, thereby hindering a satisfying accuracy. In this article, we propose a novel ADD method via learning and fusing features from visual cues. Specifically, we first construct temporal dilated convolutional network (TDCN), in which multiple dilated convolution blocks (DCBs) are designed and stacked, to learn the long-range temporal information from sequences. Then, the featurewise attention (FWA) module is adopted to fuse different features extracted from TDCNs. The module learns to assign weights for the feature channels, aiming to better incorporate different kinds of visual features and further enhance the detection accuracy. Our method achieves the state-of-the-art performance on the Distress Analysis Interview Corpus Wizard-of-Oz (DAIC_WOZ) dataset compared with other visual-feature-based methods, showing its effectiveness. Yanrong Guo, Chenyang Zhu 0005, Shijie Hao, Richang Hong |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2023 | Hierarchical Multifeature Fusion via Audio-Response-Level Modeling for Depression DetectionabstractThe clinical diagnosis of major depressive disorder (MDD) relying heavily on the subjective judgment assisted by questionnaires, could result in a low detection rate of MDD. Automatic depression detection (ADD) technology based on physiological and psychological information provides an objective and quantitative way for MDD detection. As a useful data modality, audio signals have attracted increasing interest in mental disorder detection. However, most recent audio-based depression detection methods underestimate the importance of subtly organizing audio data. They either simply use equally split audio segments, or directly build the model upon the entire data sequence, which pose challenges to learning task-specific features. To address this issue, we propose to reorganize the audio data at response level. Based on that, we construct a novel end-to-end model that hierarchically learns discriminative features for accurate depression detection. The stages of intraresponse fusion and interresponse fusion facilitate the extraction and aggregation of ADD-specific information from multiple kinds of acoustic features. Experimental results show that our proposed method significantly outperforms other state-of-the-art audio-based methods. In addition, the flexibility and the robustness of our model are also validated. Yanrong Guo, Shijie Hao, Richang Hong |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2023 | MS²-GNN: Exploring GNN-Based Multimodal Fusion Network for Depression DetectionabstractMajor depressive disorder (MDD) is one of the most common and severe mental illnesses, posing a huge burden on society and families. Recently, some multimodal methods have been proposed to learn a multimodal embedding for MDD detection and achieved promising performance. However, these methods ignore the heterogeneity/homogeneity among various modalities. Besides, earlier attempts ignore interclass separability and intraclass compactness. Inspired by the above observations, we propose a graph neural network (GNN)-based multimodal fusion strategy named modal-shared modal-specific GNN, which investigates the heterogeneity/homogeneity among various psychophysiological modalities as well as explores the potential relationship between subjects. Specifically, we develop a modal-shared and modal-specific GNN architecture to extract the inter/intramodal characteristics. Furthermore, a reconstruction network is employed to ensure fidelity within the individual modality. Moreover, we impose an attention mechanism on various embeddings to obtain a multimodal compact representation for the subsequent MDD detection task. We conduct extensive experiments on two public depression datasets and the favorable results demonstrate the effectiveness of the proposed algorithm. Tao Chen 0017, Richang Hong, Yanrong Guo, Shijie Hao, Bin Hu 0001 |
IEEE Trans. Cybern. | 4 |
| 2022 | Guest Editorial: Intelligent information processing and services in media convergence
Meng Wang 0001, Chi Zhang 0022, Shijie Hao, Jun Yu 0002, Tingting Mu |
Int. J. Intell. Syst. | 3 |
| 2022 | Adaptive dictionary and structure learning for unsupervised feature selection
Yanrong Guo, Shijie Hao |
Inf. Process. Manag. | 3 |
| 2022 | Enhancing pencil drawing patterns via using semantic information
Teng Li 0019, Jianyu Xie, Hongliang Niu, Shijie Hao |
Multim. Tools Appl. | 4 |
| 2022 | Single Image Deraining by Fully Exploiting Contextual Information
Xiaoxian Cao, Shijie Hao |
Neural Process. Lett. | 2 |
| 2022 | Stacked Pyramid Attention Network for Object Detection
Shijie Hao, Fuming Sun |
Neural Process. Lett. | 1 |
| 2022 | Differentiated Explanation of Deep Neural Networks With Skewed DistributionsabstractOver the last decade, deep neural networks (DNNs) are regarded as black-box methods, and their decisions are criticized for the lack of explainability. Existing attempts based on local explanations offer each input a visual saliency map, where the supporting features that contribute to the decision are emphasized with high relevance scores. In this paper, we improve the saliency map based on differentiated explanations, of which the saliency map not only distinguishes the supporting features from backgrounds but also shows the different degrees of importance of the various parts within the supporting features. To do this, we propose to learn a differentiated relevance estimator called DRE, where a carefully-designed distribution controller is introduced to guide the relevance scores towards right-skewed distributions. DRE can be directly optimized under pure classification losses, enabling higher faithfulness of explanations and avoiding non-trivial hyper-parameter tuning. The experimental results on three real-world datasets demonstrate that our differentiated explanations significantly improve the faithfulness with high explainability. Our code and trained models are available at https://github.com/fuweijie/DRE. Weijie Fu, Meng Wang 0001, Mengnan Du, Ninghao Liu 0001, Shijie Hao, Xia Ben Hu |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Exploring Self-Attention Graph Pooling With EEG-Based Topological Structure and Soft Label for Depression DetectionabstractElectroencephalogram (EEG) has been widely used in neurological disease detection, i.e., major depressive disorder (MDD). Recently, some deep EEG-based MDD detection attempts have been proposed and achieved promising performance. These works, however, still suffer from the following limitations, such as insufficient exploration of the EEG-based topological structure, information loss caused by high-dimensional data compression, and under-estimation of intra-class difference and inter-class similarity. To solve these issues, we propose an EEG-based MDD detection model namedSelf-attentionGraphPooling withSoftLabel (SGP-SL). Specifically, we explore the local and global connections among EEG channels to construct an EEG-based graph in advance. By leveraging multiple self-attention graph pooling modules, the constructed graph is then gradually refined, followed by graph pooling, to aggregate information from less-important nodes to more-important ones. In this way, the feature representation with better discriminability can be learned from EEG signals. In addition, the soft label strategy is also adopted to build the loss function, aiming to further enhance the feature discriminability. Experimental results on the MODMA dataset demonstrate the superiority of the proposed method. What's more, extensive ablation studies are conducted to verify the effectiveness of the proposed elements in our model. Tao Chen 0017, Yanrong Guo, Shijie Hao, Richang Hong |
IEEE Trans. Affect. Comput. | 3 |
| 2022 | FLAG: Faster Learning on Anchor Graph with Label Predictor OptimizationabstractKnowledge graphs have received intensive research interests. When the labels of most nodes or datapoints are missing, anchor graph and hierarchical anchor graph models can be employed. With an anchor graph or hierarchical anchor graph, we only need to optimize the labels of the coarsest anchors, and the labels of datapoints can be inferred from these anchors in a coarse-to-fine manner. The complexity of optimization is therefore reduced to a cubic cost with respect to the number of the coarsest anchors. However, to obtain a high accuracy when a data distribution is complex, the scale of this anchor set still needs to be large, which thus inevitably incurs an expensive computational burden. As such, a challenge in scaling up these models is how to efficiently estimate the labels of these anchors while keeping classification performance. To address this problem, we propose a novel approach that adds an anchor label predictor in the conventional anchor graph and hierarchical anchor graph models. In the proposed approach, the labels of the coarsest anchors are not directly optimized, and instead, we learn a label predictor which estimates the labels of these anchors with their spectral representations. The predictor is optimized with a regularization on all datapoints based on a hierarchical anchor graph, and we show that its solution only involves the inversion of a small-size matrix. Built upon the anchor hierarchy, we design a sparse intra-layer adjacency matrix over these anchors, which can simultaneously accelerate spectral embedding and enhance effectiveness. Our approach is named Faster Learning on Anchor Graph (FLAG) as it improves conventional anchor-graph-based methods in terms of efficiency. Experiments on a variety of publicly available datasets with sizes varying from thousands to millions of samples demonstrate the effectiveness of our approach. Weijie Fu, Meng Wang 0001, Shijie Hao, Tingting Mu |
IEEE Trans. Big Data | 3 |
| 2022 | Hierarchical Prototype Refinement With Progressive Inter-Categorical Discrimination Maximization for Few-Shot LearningabstractMetric-based few-shot learning categorizes unseen query instances by measuring their distance to the categories appearing in the given support set. To facilitate distance measurement, prototypes are used to approximate the representations of categories. However, we find prototypical representations are generally not discriminative enough to represent the discrepancy of inter-categorical distribution of queries, thereby limiting the classification accuracy. To overcome this issue, we propose a new Progressive Hierarchical-Refinement (PHR) method, which effectively refines the discrimination of prototypes by conducting the Progressive Discrimination Maximization strategy based on the hierarchical feature representations. Specifically, we first encode supports and queries into the representation space of spatial level, global level, and semantic level. Then, the refining coefficients are constructed by exploring the metric information contained in these hierarchical embedding spaces simultaneously. Under the guidance of the refining coefficients, the meta-refining loss progressively maximizes the discrimination degree of inter-categorical prototypical representations. In addition, the refining vectors are adopted to further enhance the representations of prototypes. In this way, the metric-based classification can be more accurate. Our PHR method shows the competitive performance on the miniImagenet, CIFAR-FS, FC100, and CUB datasets. Moreover, PHR presents good compatibility. It can be incorporated with other few-shot learning models, making them more accurate. Yuan Zhou 0016, Yanrong Guo, Shijie Hao, Richang Hong |
IEEE Trans. Image Process. | 3 |
| 2022 | A Survey on Large-Scale Machine LearningabstractMachine learning can provide deep insights into data, allowing machines to make high-quality predictions and having been widely used in real-world applications, such as text mining, visual classification, and recommender systems. However, most sophisticated machine learning approaches suffer from huge time costs when operating on large-scale data. This issue calls for the need of Large-scale Machine Learning (LML), which aims to learn patterns from big data with comparable performance efficiently. In this paper, we offer a systematic survey on existing LML methods to provide a blueprint for the future developments of this area. We first divide these LML methods according to the ways of improving the scalability: 1) model simplification on computational complexities, 2) optimization approximation on computational efficiency, and 3) computation parallelism on computational capabilities. Then we categorize the methods in each perspective according to their targeted scenarios and introduce representative methods in line with intrinsic strategies. Lastly, we analyze their limitations and discuss potential directions as well as open issues that are promising to address in the future. Meng Wang 0001, Weijie Fu, Xiangnan He 0001, Shijie Hao, Xindong Wu 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2022 | Decoupled Low-Light Image EnhancementabstractThe visual quality of photographs taken under imperfect lightness conditions can be degenerated by multiple factors, e.g., low lightness, imaging noise, color distortion, and so on. Current low-light image enhancement models focus on the improvement of low lightness only, or simply deal with all the degeneration factors as a whole, therefore leading to sub-optimal results. In this article, we propose to decouple the enhancement model into two sequential stages. The first stage focuses on improving the scene visibility based on a pixel-wise non-linear mapping. The second stage focuses on improving the appearance fidelity by suppressing the rest degeneration factors. The decoupled model facilitates the enhancement in two aspects. On the one hand, the whole low-light enhancement can be divided into two easier subtasks. The first one only aims to enhance the visibility. It also helps to bridge the large intensity gap between the low-light and normal-light images. In this way, the second subtask can be described as the local appearance adjustment. On the other hand, since the parameter matrix learned from the first stage is aware of the lightness distribution and the scene structure, it can be incorporated into the second stage as the complementary information. In the experiments, our model demonstrates the state-of-the-art performance in both qualitative and quantitative comparisons, compared with other low-light image enhancement models. In addition, the ablation studies also validate the effectiveness of our model in multiple aspects, such as model structure and loss function. Shijie Hao, Yanrong Guo, Meng Wang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2021 | Positive Sample Propagation Along the Audio-Visual Event LineabstractVisual and audio signals often coexist in natural environments, forming audio-visual events (AVEs). Given a video, we aim to localize video segments containing an AVE and identify its category. In order to learn discriminative features for a classifier, it is pivotal to identify the helpful (or positive) audio-visual segment pairs while filtering out the irrelevant ones, regardless whether they are synchronized or not. To this end, we propose a new positive sample propagation (PSP) module to discover and exploit the closely related audio-visual pairs by evaluating the relationship within every possible pair. It can be done by constructing an all-pair similarity map between each audio and visual segment, and only aggregating the features from the pairs with high similarity scores. To encourage the network to extract high correlated features for positive samples, a new audio-visual pair similarity loss is proposed. We also propose a new weighting branch to better exploit the temporal correlations in weakly supervised setting. We perform extensive experiments on the public AVE dataset and achieve new state-of-the-art accuracy in both fully and weakly supervised settings, thus verifying the effectiveness of our method. Jinxing Zhou, Liang Zheng 0001, Yiran Zhong, Shijie Hao, Meng Wang 0001 |
CVPR | 4 |
| 2021 | Robust pencil drawing generation via fast Retinex decomposition
Teng Li 0019, Shijie Hao, Yanrong Guo |
Comput. Graph. | 2 |
| 2021 | LEDet: A Single-Shot Real-Time Object Detector Based on Low-Light Image EnhancementabstractAbstract Recently, significant breakthroughs have been achieved in the field of object detection. However, existing methods mostly focus on the generic object detection task. Performance degradation can be unavoidable when applying the existing methods to some specific situations directly, e.g. a low-light environment. To address this issue, we propose a single-shot real-time object Detector based on Low-light image Enhancement, namely LEDet. LEDet adapts itself to the low-light detection task in three aspects. First, a low-light enhancement module is introduced as the image preprocessor, producing the augmented inputs from the low-light images. Second, two modules, i.e. low-light and enhanced features fusion module and the scale-aware channel attention dilated convolution module are designed. These two modules aim at learning robust and discriminative features from objects of various sizes hidden in the darkness. In experiments, we validate the effectiveness of each part of our LEDet model via several ablation studies. We also compare LEDet with various methods on the Exclusively Dark dataset, showing that our model achieves the state-of-the-art performance on the balance between speed and accuracy. Shijie Hao, Fuming Sun |
Comput. J. | 1 |
| 2021 | Low-light enhancement based on an improved simplified Retinex model via fast illumination map refinement
Shijie Hao |
Pattern Anal. Appl. | 1 |
| 2021 | Adaptive Multi-Task Dual-Structured Learning with Its Application on Alzheimer's Disease StudyabstractMulti-task learning has been widely applied to Alzheimer’s Disease (AD) studies due to its capability of simultaneously rating the disease severity (classification) and predicting corresponding clinical scores (regression). In this article, we propose a novel technique of Adaptive Multi-task Dual-Structured Learning, named AMDSL, by mutually exploring the dual manifold structure for the label and regression score of the disease data under joint classification and regression tasks, while learning an adaptive shared similarity measure and corresponding feature mapping among these two tasks. We encode both the reconstructed label representation and regression score adaptive to the ideal similarity measure on disease data to achieve the ideal performance on these two joint tasks. The alternating algorithm is proposed to optimize the above objective. We theoretically prove the convergence of the optimization algorithm. The superiority of AMDSL is experimentally validated under joint classification and regression as per various evaluation metrics against the most authoritative Alzheimer’s disease data. Shijie Hao, Tao Chen 0017, Yang Wang 0023, Yanrong Guo, Meng Wang 0001 |
ACM Trans. Internet Techn. | 1 |
| 2020 | A Brief Survey on Semantic Segmentation with Deep Learning
Shijie Hao, Yuan Zhou 0016, Yanrong Guo |
Neurocomputing | 1 |
| 2020 | Unsupervised feature selection based on joint spectral learning and general sparse regression
Tao Chen 0017, Yanrong Guo, Shijie Hao |
Neural Comput. Appl. | 3 |
| 2020 | Single-image low-light enhancement via generating and fusing multiple sources
Zhuang Feng, Shijie Hao |
Neural Comput. Appl. | 4 |
| 2020 | Semantically-guided low-light image enhancement
Junyi Xie, Hao Bian, Yuanhang Wu, Linmin Shan, Shijie Hao |
Pattern Recognit. Lett. | 6 |
| 2020 | Low-Light Image Enhancement With Semi-Decoupled DecompositionabstractLow-light image enhancement is important for high-quality image display and other visual applications. However, it is a challenging task as the enhancement is expected to improve the visibility of an image while keeping its visual naturalness. Retinex-based methods have well been recognized as a representative technique for this task, but they still have the following limitations. First, due to less-effective image decomposition or strong imaging noise, various artifacts can still be brought into enhanced results. Second, although the priori information can be explored to partially solve the first issue, it requires to carefully model the priori by a regularization term and usually makes the optimization process complicated. In this paper, we address these issues by proposing a novel Retinex-based low-light image enhancement method, in which the Retinex image decomposition is achieved in an efficient semi-decoupled way. Specifically, the illumination layer I is gradually estimated only with the input image S based on the proposed Gaussian Total Variation model, while the reflectance layer R is jointly estimated by S and the intermediate I. In addition, the imaging noise can be simultaneously suppressed during the estimation of R. Experimental results on several public datasets demonstrate that our method produces images with both higher visibility and better visual quality, which outperforms the state-of-the-art low-light enhancement methods in terms of several objective and subjective evaluation metrics. Shijie Hao, Yanrong Guo, Xin Xu 0007, Meng Wang 0001 |
IEEE Trans. Multim. | 1 |
| 2019 | Lightness-aware contrast enhancement for images with different illumination conditions
Shijie Hao, Yanrong Guo, Zhongliang Wei |
Multim. Tools Appl. | 1 |
| 2018 | Scalable Active Learning by Approximated Error ReductionabstractWe study the problem of active learning for multi-class classification on large-scale datasets. In this setting, the existing active learning approaches built upon uncertainty measures are ineffective for discovering unknown regions, and those based on expected error reduction are inefficient owing to their huge time costs. To overcome the above issues, this paper proposes a novel query selection criterion called approximated error reduction (AER). In AER, the error reduction of each candidate is estimated based on an expected impact over all datapoints and an approximated ratio between the error reduction and the impact over its nearby datapoints. In particular, we utilize hierarchical anchor graphs to construct the candidate set as well as the nearby datapoint sets of these candidates. The benefit of this strategy is that it enables a hierarchical expansion of candidates with the increase of labels, and allows us to further accelerate the AER estimation. We finally introduce AER into an efficient semi-supervised classifier for scalable active learning. Experiments on publicly available datasets with the sizes varying from thousands to millions demonstrate the effectiveness of our approach. Weijie Fu, Meng Wang 0001, Shijie Hao, Xindong Wu 0001 |
KDD | 3 |
| 2018 | Low-light image enhancement with a refined illumination map
Shijie Hao, Zhuang Feng, Yanrong Guo |
Multim. Tools Appl. | 1 |
| 2018 | Semi-supervised vehicle classification via fusing affinity matrices
Maojin Sun, Shijie Hao, Guangcan Liu |
Signal Process. | 2 |
| 2017 | Anatomical landmark detection on 3D human shapes by hierarchically utilizing multiple shape features
Zhenkun Zhou, Shijie Hao |
Neurocomputing | 2 |
| 2017 | Learning on Big Graph: Label Inference and Regularization with Anchor HierarchyabstractSeveral models have been proposed to cope with the rapidly increasing size of data, such as Anchor Graph Regularization (AGR). The AGR approach significantly accelerates graph-based learning by exploring a set of anchors. However, when a dataset becomes much larger, AGR still faces a big graph which brings dramatically increasing computational costs. To overcome this issue, we propose a novel Hierarchical Anchor Graph Regularization (HAGR) approach by exploring multiple-layer anchors with a pyramid-style structure. In HAGR, the labels of datapoints are inferred from the coarsest anchors layer by layer in a coarse-to-fine manner. The label smoothness regularization is performed on all datapoints, and we demonstrate that the optimization process only involves a small-size reduced Laplacian matrix. We also introduce a fast approach to construct our hierarchical anchor graph based on an approximate nearest neighbor search technique. Experiments on million-scale datasets demonstrate the effectiveness and efficiency of the proposed HAGR approach over existing methods. Results show that the HAGR approach is even able to achieve a good performance within 3 minutes in an 8-million-example classification task. Meng Wang 0001, Weijie Fu, Shijie Hao, Hengchang Liu, Xindong Wu 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2016 | Learning-Based Topological Correction for Infant Cortical Surfaces
Shijie Hao, Gang Li 0001, Li Wang 0026, Yu Meng 0003, Dinggang Shen |
MICCAI (1) | 1 |
| 2016 | Active learning on anchorgraph with an improved transductive experimental design
Weijie Fu, Shijie Hao, Meng Wang 0001 |
Neurocomputing | 2 |
| 2016 | Dimensionality reduction on Anchorgraph with an efficient Locality Preserving Projection
Weijie Fu, Shijie Hao, Richang Hong |
Neurocomputing | 4 |
| 2016 | Buffer management for streaming media transmission in hierarchical data of opportunistic networks
Daru Pan, Han Zhang 0011, Shijie Hao |
Neurocomputing | 5 |
| 2016 | Social video annotation by combining features with a tri-adaptation approach
Fuming Sun, Meixiang Xu, Shijie Hao |
Multim. Syst. | 4 |
| 2016 | Spatially guided local Laplacian filter for nature image detail enhancement
Shijie Hao, Meng Wang 0001, Richang Hong |
Multim. Tools Appl. | 1 |
| 2016 | Image detail enhancement with spatially guided filters
Shijie Hao, Daru Pan, Yanrong Guo, Richang Hong, Meng Wang 0001 |
Signal Process. | 1 |
| 2016 | Scalable Semi-Supervised Learning by Efficient Anchor Graph RegularizationabstractMany graph-based semi-supervised learning methods for large datasets have been proposed to cope with the rapidly increasing size of data, such as Anchor Graph Regularization (AGR). This model builds a regularization framework by exploring the underlying structure of the whole dataset with both datapoints and anchors. Nevertheless, AGR still has limitations in its two components: (1) in anchor graph construction, the estimation of the local weights between each datapoint and its neighboring anchors could be biased and relatively slow; and (2) in anchor graph regularization, the adjacency matrix that estimates the relationship between datapoints, is not sufficiently effective. In this paper, we develop an Efficient Anchor Graph Regularization (EAGR) by tackling these issues. First, we propose a fast local anchor embedding method, which reformulates the optimization of local weights and obtains an analytical solution. We show that this method better reconstructs datapoints with anchors and speeds up the optimizing process. Second, we propose a new adjacency matrix among anchors by considering the commonly linked datapoints, which leads to a more effective normalized graph Laplacian over anchors. We show that, with the novel local weight estimation and normalized graph Laplacian, EAGR is able to achieve better classification accuracy with much less computational costs. Experimental results on several publicly available datasets demonstrate the effectiveness of our approach. Meng Wang 0001, Weijie Fu, Shijie Hao, Dacheng Tao, Xindong Wu 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2015 | Image classification based on low-rank matrix recovery and Naive Bayes collaborative representation
Shijie Hao, Xueming Qian, Meng Wang 0001 |
Neurocomputing | 2 |
| 2014 | Shape analysis based on feature-preserving Elastic Quadratic Patch Modeling
Fuming Sun, Shijie Hao |
Neurocomputing | 3 |
| 2014 | Image quality assessment based on matching pursuit
Richang Hong, Jianxin Pan, Shijie Hao, Meng Wang 0001, Feng Xue 0002, Xindong Wu 0001 |
Inf. Sci. | 3 |
| 2014 | Automatic image annotation by semi-supervised manifold kernel density estimation
Na Zhao 0004, Shijie Hao |
Inf. Sci. | 3 |
| 2014 | Person Re-identification based on nonlinear ranking with difference vectors
Tianfeng Zhou, Meibin Qi, Shijie Hao, Yulong Jin |
Inf. Sci. | 5 |
| 2013 | Active learning based intervertebral disk classification combining shape and texture similarities
Shijie Hao, Yanrong Guo |
Neurocomputing | 1 |
| 2013 | ε-Isometryε-Isometry based shape approximation for image content representation
Shijie Hao, Yanrong Guo, Shu Zhan |
Signal Process. | 1 |
| 2011 | Distribution-Based Active Contour Model for Medical Image SegmentationabstractHaving being regarded as one of the classical methods in image segmentation, geodesic active contours (GAC) have the flaws of boundary leaking and expensive evolving time. In this paper, we present a distribution-based active contour model by measuring the Bhattacharyya distance between probability distributions of the object and background along with the evolution of GAC model. Due to combining the image cues of edge and statistical information which is computed by using kernel density estimation, this hybrid methodology prevents the boundary leaking as well as the under segmentation problem. Experimental results on the medical images show the improvements of our method in terms of comparisons with original GAC model, Bhattacharyya gradient flow, texture-based GAC and Li's active contour model. Yanrong Guo, Shijie Hao, Shu Zhan |
ICIG | 3 |
| 2011 | Intervertebral Disc Shape Analysis with Geodesic Metric in Shape SpaceabstractShapes of anatomical structures extracted from medical imaging usually contain diagnostic and therapeutic cues in clinical applications. In this paper, we propose a framework on analyzing disc shapes based on a geodesic metric in an anatomical shape space. All disc shapes, containing both normal and abnormal ones, are formulated as elements in this space. The geodesic connecting these elements and other statistics are then numerically approximated. With these tools quantifying the intrinsic difference between disc shapes, the normal shapes from a dataset are unsupervisedly clustered and a statistical inference based on the learned Gaussian model is made. Experimental results show a reasonable accuracy of classifying normal and abnormal intervertebral discs. Shijie Hao, Yanrong Guo, Shu Zhan |
ICIG | 1 |