Hezhe Qiao

dblp:300/2321 · DBLP profile ↗
← Back
18ranked-venue papers
4as first author
18since 2021 · last 2026
0000-0003-3511-0528ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 3 first-author · 11 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Clique Annealing: Semi-Supervised Community Detection Under Crystallization Kinetics
abstract
Semi-supervised community detection seeks to find a specified community type when only few communities are labeled. Existing “select-then-refine” pipelines often start from mis-aligned cores and rely on Reinforcement-Learning or Gen-erative Adversarial Network, increasing computational cost and limiting scalability. We address these issues with a unified energy framework under crystallization kinetics that jointly models energy, structure, and growth. Based on this perspective, we pro-pose CLique ANNealing (CLANN), which first employs Nucleus Proposer to select candidate clique as community core under four physics-inspired criteria. A learning-free Transitive Annealer then iteratively merges neighboring cliques and repositions the nucleus, enabling spontaneous, scalable community growth. Evaluated on diverse real-world and synthetic networks, CLANN surpasses state-of-the-art baselines by a wide margin while running faster on large graphs, demonstrating that the energy-driven crystallization kinetics framework is both princi-pled and practical for semi-supervised community detection.
Ling Cheng 0002, Jiashu Pu, Ruicheng Liang, Qian Shao, Hezhe Qiao, Feida Zhu 0001
IEEE Trans. Knowl. Data Eng.5
2026 SSD: Self-Supervised Distillation for Heterophilic Graph Representation Learning
Yuan Gao 0032, Yuchen Li 0001, Bingsheng He, Hezhe Qiao, Guoguo Ai
IEEE Trans. Knowl. Data Eng.4
2025 Semantic-guided Representation Learning for Multi-Label Recognition
abstract
Multi-label Recognition (MLR) involves assigning multiple labels to each data instance in an image, offering advantages over single-label classification in complex scenarios. However, it faces the challenge of annotating all relevant categories, often leading to uncertain annotations, such as unseen or incomplete labels. Recent Vision and Language Pre-training (VLP) based methods have made significant progress in tackling zero-shot MLR tasks by leveraging rich vision-language correlations. However, the correlation between multi-label semantics has not been fully explored, and the learned visual features often lack essential semantic information. To overcome these limitations, we introduce a Semantic-guided Representation Learning approach (SigRL) that enables the model to learn effective visual and textual representations, thereby improving the downstream alignment of visual images and categories. Specifically, we first introduce a graph-based multi-label correlation module (GMC) to facilitate information exchange between labels, enriching the semantic representation across the multi-label texts. Next, we propose a Semantic Visual Feature Reconstruction module (SVFR) to enhance the semantic information in the visual representation by integrating the learned textual representation during reconstruction. Finally, we optimize the image-text matching capability of the VLP model using both local and global features to achieve zero-shot MLR. Comprehensive experiments are conducted on several MLR benchmarks, encompassing both zero-shot MLR (with unseen labels) and single positive multi-label learning (with limited labels), demonstrating the superior performance of our approach compared to state-of-the-art methods. The code is available at https://github.com/MVL-Lab/SigRL.
Ruhui Zhang, Hezhe Qiao, Mingsheng Shang 0001, Lin Chen 0023
ICME2
2025 GrokFormer: Graph Fourier Kolmogorov-Arnold Transformers
abstract
Graph Transformers (GTs) have demonstrated remarkable performance in graph representation learning over popular graph neural networks (GNNs). However, self-attention, the core module of GTs, preserves only low-frequency signals in graph features, leading to ineffectiveness in capturing other important signals like high-frequency ones. Some recent GT models help alleviate this issue, but their flexibility and expressiveness are still limited since the filters they learn are fixed on predefined graph spectrum or spectral order. To tackle this challenge, we propose a Graph Fourier Kolmogorov-Arnold Transformer (GrokFormer), a novel GT model that learns highly expressive spectral filters with adaptive graph spectrum and spectral order through a Fourier series modeling over learnable activation functions. We demonstrate theoretically and empirically that the proposed GrokFormer filter offers better expressiveness than other spectral methods. Comprehensive experiments on 10 real-world node classification datasets across various domains, scales, and graph properties, as well as 5 graph classification datasets, show that GrokFormer outperforms state-of-the-art GTs and GNNs. Our code is available at https://github.com/GGA23/GrokFormer.
GuoguoAi, Guansong Pang, Hezhe Qiao
ICML3
2025 Zero-shot Generalist Graph Anomaly Detection with Unified Neighborhood Prompts
abstract
Graph anomaly detection (GAD), which aims to identify nodes in a graph that significantly deviate from normal patterns, plays a crucial role in broad application domains. However, existing GAD methods are one-model-for-one-dataset approaches, i.e., training a separate model for each graph dataset. This largely limits their applicability in real-world scenarios. To overcome this limitation, we propose a novel zero-shot generalist GAD approach UNPrompt that trains a one-for-all detection model, requiring the training of one GAD model on a single graph dataset and then effectively generalizing to detect anomalies in other graph datasets without any retraining or fine-tuning. The key insight in UNPrompt is that i) the predictability of latent node attributes can serve as a generalized anomaly measure and ii) generalized normal and abnormal graph patterns can be learned via latent node attribute prediction in a properly normalized node attribute space. UNPrompt achieves a generalist mode for GAD through two main modules: one module aligns the dimensionality and semantics of node attributes across different graphs via coordinate-wise normalization, while another module learns generalized neighborhood prompts that support the use of latent node attribute predictability as an anomaly score across different datasets. Extensive experiments on real-world GAD datasets show that UNPrompt significantly outperforms diverse competing methods under the generalist GAD setting, and it also has strong superiority under the one-model-for-one-dataset setting. Code is available at https://github.com/mala-lab/UNPrompt.
Chaoxi Niu, Hezhe Qiao, Changlu Chen, Ling Chen 0006, Guansong Pang
IJCAI2
2025 AnomalyGFM: Graph Foundation Model for Zero/Few-shot Anomaly Detection
abstract
Graph anomaly detection (GAD) aims to identify abnormal nodes that differ from the majority of the nodes in a graph, which has been attracting significant attention in recent years.Existing generalist graph models have achieved remarkable success in different graph tasks but struggle to generalize to the GAD task.This limitation arises from their difficulty in learning generalized knowledge for capturing the inherently infrequent, irregular and heterogeneous abnormality patterns in graphs from different domains.To address this challenge, we propose AnomalyGFM, a GAD-oriented graph foundation model that supports zero-shot inference and few-shot prompt tuning for GAD in diverse graph datasets.One key insight is that graph-agnostic representations for normal and abnormal classes are required to support effective zero/few-shot GAD across different graphs.Motivated by this, AnomalyGFM is pre-trained to align data-independent, learnable normal and abnormal class prototypes with node representation residuals (i.e., representation deviation of a node from its neighbors).The residual features essentially project the node information into a unified feature space where we can effectively measure the abnormality of nodes from different graphs in a consistent way.This provides a driving force for the learning of graph-agnostic, discriminative prototypes for the normal and abnormal classes, which can be used to enable zero-shot GAD on new graphs, including very large-scale graphs.If there are few-shot labeled normal nodes available in the new graphs, AnomalyGFM can further support prompt tuning to leverage these nodes for better adaptation.Comprehensive experiments on 11 widely-used GAD datasets with real anomalies, covering social networks, finance networks, and co-review networks, demonstrate that AnomalyGFM significantly outperforms state-of-the-art competing methods under both zero-and few-shot GAD settings.Code is available at https://github.com/mala-lab/AnomalyGFM.
Hezhe Qiao, Chaoxi Niu, Ling Chen 0006, Guansong Pang
KDD (2)1
2025 Semi-supervised Graph Anomaly Detection via Robust Homophily Learning
abstract
Current semi-supervised graph anomaly detection (GAD) methods utilizes a small set of labeled normal nodes to identify abnormal nodes from a large set of unlabeled nodes in a graph. These methods posit that 1) normal nodes share a similar level of homophily and 2) the labeled normal nodes can well represent the homophily patterns in the entire normal class. However, this assumption often does not hold well since normal nodes in a graph can exhibit diverse homophily in real-world GAD datasets. In this paper, we propose RHO, namely Robust Homophily Learning, to adaptively learn such homophily patterns. RHO consists of two novel modules, adaptive frequency response filters (AdaFreq) and graph normality alignment (GNA). AdaFreq learns a set of adaptive spectral filters that capture different frequency components of the labeled normal nodes with varying homophily in the channel-wise and cross-channel views of node attributes. GNA is introduced to enforce consistency between the channel-wise and cross-channel homophily representations to robustify the normality learned by the filters in the two views. Experiments on eight real-world GAD datasets show that RHO can effectively learn varying, often under-represented, homophily in the small labeled node set and substantially outperforms state-of-the-art competing methods. Code is available at \url{https://github.com/mala-lab/RHO}.
Guoguo Ai, Hezhe Qiao, Guansong Pang
NeurIPS2
2025 Deep Graph Anomaly Detection: A Survey and New Perspectives
abstract
Graph anomaly detection (GAD), which aims to identify unusual graph instances (e.g., nodes, edges, subgraphs, or graphs), has attracted increasing attention in recent years due to its significance in a wide range of applications. Deep learning approaches, graph neural networks (GNNs) in particular, have been emerging as a promising paradigm for GAD, owing to its strong capability in capturing complex structure and/or node attributes in graph data. Considering the large number of methods proposed for GNN-based GAD, it is of paramount importance to summarize the methodologies and findings in the existing GAD studies, so that we can pinpoint effective model designs for tackling open GAD problems. To this end, in this work we aim to present a comprehensive review of deep learning approaches for GAD. Existing GAD surveys are focused on task-specific discussions, making it difficult to understand the technical insights of existing methods and their limitations in addressing some unique challenges in GAD. To fill this gap, we first discuss the problem complexities and their resulting challenges in GAD, and then provide a systematic review of current deep GAD methods from three novel perspectives of methodology, including GNN backbone design, proxy task design for GAD, and graph anomaly measures. To deepen the discussions, we further propose a taxonomy of 13 fine-grained method categories under these three perspectives to provide more in-depth insights into the model designs and their capabilities. To facilitate the experiments and validation of the GAD methods, we also summarize a collection of widely-used datasets for GAD and empirical performance comparison on these datasets. We further discuss multiple important open research problems in GAD to inspire more future high-quality research in this area. A continuously updated repository for GAD datasets, links to the codes of GAD algorithms, and empirical comparison.
Hezhe Qiao, Hanghang Tong, Bo An 0001, Irwin King, Charu C. Aggarwal, Guansong Pang
IEEE Trans. Knowl. Data Eng.1
2025 Toward Diverse Tiny-Model Selection for Microcontrollers
abstract
Enabling efficient and accurate deep neural network (DNN) inference on microcontrollers is challenging due to their constrained on-chip resources. Existing approaches mainly focus on compressing larger models, often compromising model accuracy as a trade-off. In this paper, we rethink the problem from the inverse perspective by directly constructing small/weak models, then enhancing their accuracy. Thus, we propose DiTMoS, a novel DNN training and inference framework featuring aselector-classifiersarchitecture, where the selector routes each input sample to the appropriate classifier for classification. DiTMoS is built on a key insight: a combination of weak models can exhibit high diversity and the union of them can significantly raise the upper bound of overall accuracy. To approach the upper bound, DiTMoS introduces three strategies including diverse training data splitting to enhance the classifiers' diversity, adversarial selector-classifiers training to ensure synergistic interactions thereby maximizing their complementarity, and heterogeneous feature aggregation to improve the capacity of classifiers. We further design a network slicing technique to eliminate the extra memory consumption incurred by feature aggregation. We deploy DiTMoS on the Nucleo STM32F767ZI board and evaluate its performance across three time-series datasets for human activity recognition, keyword spotting, and emotion recognition tasks. The experimental results show that: (a) DiTMoS improves accuracy by up to 13.4% compared to the best baseline; (b) network slicing successfully eliminates the memory overhead introduced by feature aggregation, with only a minimal increase in latency. The code of DiTMoS is released athttps://github.com/TheMaXiao/DiTMoS
Shengfeng He, Hezhe Qiao, Dong Ma 0001
IEEE Trans. Mob. Comput.3
2024 Temporal Gaussian Copula For Clinical Multivariate Time Series Data Imputation
abstract
The imputation of the Multivariate time series (MTS) is particularly challenging since the MTS typically contains irregular patterns of missing values due to various factors such as instrument failures, interference from irrelevant data, and privacy regulations. Existing statistical methods and deep learning methods have shown promising results in time series imputation. In this paper, we propose a Temporal Gaussia Copula Model (TGC) for three-order MTS imputation. The key idea is to leverage the Gaussian Copula to explore the cross-variable and temporal relationships based on the latent Gaussian representation. Subsequently, we employ an Expectation-Maximization (EM) algorithm to improve robustness in managing data with varying missing rates. Comprehensive experiments were conducted on three real-world MTS datasets. The results demonstrate that our TGC substantially outperforms the state-of-the-art imputation methods. Additionally, the TGC model exhibits stronger robustness to the varying missing ratios in the test dataset. Our code is available at https://github.com/chenlincigit/TGC-MTS.
Hezhe Qiao, Lin Chen 0023
BIBM2
2024 Temporal-Contextual Event Learning for Pedestrian Crossing Intent Prediction
Hongbin Liang, Hezhe Qiao, Mingsheng Shang 0001, Lin Chen 0023
ICONIP (5)2
2024 Generative Semi-supervised Graph Anomaly Detection
abstract
This work considers a practical semi-supervised graph anomaly detection (GAD) scenario, where part of the nodes in a graph are known to be normal, contrasting to the extensively explored unsupervised setting with a fully unlabeled graph. We reveal that having access to the normal nodes, even just a small percentage of normal nodes, helps enhance the detection performance of existing unsupervised GAD methods when they are adapted to the semi-supervised setting. However, their utilization of these normal nodes is limited. In this paper, we propose a novel Generative GAD approach (namely GGAD) for the semi-supervised scenario to better exploit the normal nodes. The key idea is to generate pseudo anomaly nodes, referred to as 'outlier nodes', for providing effective negative node samples in training a discriminative one-class classifier. The main challenge here lies in the lack of ground truth information about real anomaly nodes. To address this challenge, GGAD is designed to leverage two important priors about the anomaly nodes -- asymmetric local affinity and egocentric closeness -- to generate reliable outlier nodes that assimilate anomaly nodes in both graph structure and feature representations. Comprehensive experiments on six real-world GAD datasets are performed to establish a benchmark for semi-supervised GAD and show that GGAD substantially outperforms state-of-the-art unsupervised and semi-supervised GAD methods with varying numbers of training normal nodes.
Hezhe Qiao, Qingsong Wen, Xiaoli Li 0001, Ee-Peng Lim, Guansong Pang
NeurIPS1
2024 DiTMoS: Delving into Diverse Tiny-Model Selection on Microcontrollers
abstract
Enabling efficient and accurate deep neural network (DNN) inference on microcontrollers is non-trivial due to the constrained on-chip resources. Current methodologies primarily focus on compressing larger models yet at the expense of model accuracy. In this paper, we rethink the problem from the inverse perspective by constructing small/weak models directly and improving their accuracy. Thus, we introduce DiTMoS, a novel DNN training and inference framework with a selector-classifiers architecture, where the selector routes each input sample to the appropriate classifier for classification. DiTMoS is grounded on a key insight: a composition of weak models can exhibit high diversity and the union of them can significantly boost the accuracy upper bound. To approach the upper bound, DiT-MoS introduces three strategies including diverse training data splitting to increase the classifiers' diversity, adversarial selector-classifiers training to ensure synergistic interactions thereby maximizing their complementarity, and heterogeneous feature aggregation to improve the capacity of classifiers. We further propose a network slicing technique to alleviate the extra memory overhead incurred by feature aggregation. We deploy DiTMoS on the Neucleo STM32F767ZI board and evaluate it based on three time-series datasets for human activity recognition, keywords spotting, and emotion recognition, respectively. The experiment results manifest that: (a) DiTMoS achieves up to 13.4% accuracy improvement compared to the best baseline; (b) network slicing almost completely eliminates the memory overhead incurred by feature aggregation with a marginal increase of latency. Code is released at https//github.com/TheMaXiao/DiTMoS
Shengfeng He, Hezhe Qiao, Dong Ma 0001
PerCom3
2024 Self-supervised Spatial-Temporal Normality Learning for Time Series Anomaly Detection
Hongzuo Xu, Guansong Pang, Hezhe Qiao, Mingsheng Shang 0001
ECML/PKDD (6)4
2023 Truncated Affinity Maximization: One-class Homophily Modeling for Graph Anomaly Detection
abstract
We reveal a one-class homophily phenomenon, which is one prevalent property we find empirically in real-world graph anomaly detection (GAD) datasets, i.e., normal nodes tend to have strong connection/affinity with each other, while the homophily in abnormal nodes is significantly weaker than normal nodes. However, this anomaly-discriminative property is ignored by existing GAD methods that are typically built using a conventional anomaly detection objective, such as data reconstruction. In this work, we explore this property to introduce a novel unsupervised anomaly scoring measure for GAD -- local node affinity-- that assigns a larger anomaly score to nodes that are less affiliated with their neighbors, with the affinity defined as similarity on node attributes/representations. We further propose Truncated Affinity Maximization (TAM) that learns tailored node representations for our anomaly measure by maximizing the local affinity of nodes to their neighbors. Optimizing on the original graph structure can be biased by non-homophily edges(i.e., edges connecting normal and abnormal nodes). Thus, TAM is instead optimized on truncated graphs where non-homophily edges are removed iteratively to mitigate this bias. The learned representations result in significantly stronger local affinity for normal nodes than abnormal nodes. Extensive empirical results on 10 real-world GAD datasets show that TAM substantially outperforms seven competing models, achieving over 10% increase in AUROC/AUPRC compared to the best contenders on challenging datasets. Our code is available at https://github.com/mala-lab/TAM-master/.
Hezhe Qiao, Guansong Pang
NeurIPS1
2022 Alzheimer's Disease Clinical Scores Prediction based on the Label Distribution Learning using Brain Structural MRI
abstract
Predicting Alzheimer's disease (AD) clinical scores offers a useful tool to monitor dementia progression at different time points. Previous machine learning methods usually focused on regressing the clinical scores directly but ignored the ambigu-ous information among score labels. In this study, we introduce a novel AD clinical scores prediction framework based on label distribution learning (LDL), named CSP-LDL. Notably, we first turn the clinical scores into a normal probability distribution of discrete labels to exploit uncertainty among the dementia scores. Then we learn the distribution of discrete labels by optimizing Kullback-Leibler (KL) divergence between the estimated and ground-truth distributions using a 3D CNN. Moreover, we further employ an expectation regression layer to regress clinical score value at the fine-grained level based on the predicted label distribution. Experiments on ADNI-1 and ADNI-2 datasets show that our CSP-LDL model outperforms existing state-of-the-art methods in terms of dementia regression accuracy at multiple time points using baseline structural magnetic resonance imaging (sMRI) data, demonstrating its effectiveness in early clinical diagnosis of AD.
Hezhe Qiao, Lin Chen 0023
IJCNN2
2022 Aspect-aware Asymmetric Representation Learning Network for Review-based Recommendation
abstract
Recently, user-provided reviews have been identified as an essential resource to improve user and item representation in recommender systems. Previous methods focus on the review-based recommender typically leverages symmetric networks to process user and item reviews. However, in reality, these two sets of reviews are markedly different: a user's reviews reflect the experience of buying diverse items and show their heterogeneous interests. In contrast, an item's reviews emphasize the quality of the specific item. Thus an item's reviews are usually homogeneous. This paper seeks to explore the aspect of review difference in the review-based recommendation framework. We propose a novel asymmetric neural network model that accurately learns the user and item representation by identifying this critical difference. We focus on capturing the dynamic change of user interest for the user-aspect reviews via modeling the temporal information into the conventional neural network(CNN). On the other side, we try to identify a specific item's essential yet essential features by utilizing the self-attention neural network. Finally, a factorization machine (FM) is adopted to finish the rating prediction task, where the user and item IDs are encoded as supplementary review embedding. We conduct comprehensive experiments on four Amazon datasets, and the experimental results show that our proposed model consistently outperforms several state-of-the-art methods.
Hezhe Qiao, Xiaoyu Shi 0001, Mingsheng Shang 0001
IJCNN2
2021 Improving Pedestrian Attribute Recognition with Multi-Scale Spatial Calibration
abstract
Pedestrian Attribute Recognition (PAR) has attracted increasing attention since it could provide important structural information of pedestrians for Smart Video Analysis. However, the pedestrian images are taken from a far distance significantly increase the difficulty of PAR for fine-grained attributes. To address these problems, and further improve the effects of PAR, we proposed a Multi-Scale Spatial Calibration (MSSC) module. More specifically, the module includes two submodules: first, a Spatial Calibrated Module (SCM) is proposed to extract more discriminative features of inconspicuous attributes from its surrounding regions by gathering the contextual information across different receptive fields. Moreover, in order to build the long-range dependencies of pyramid feature maps in different spatial scales, we also propose Multi-Scale Feature Fusion (MSFF) to integrate the multiple branches of low-level detailed features and high-level semantics features by non-local attention mechanism. Extensive experiments show that our proposed model could achieve state-of-the-art results on three pedestrian attribute datasets, including RAPv1, PA-100K, and RAPv2. Especially, the proposed model significantly improves the recognition effects of fine-grained attributes in low-resolution images in terms of mean Accuracy (mA) and recall. Code is available at https://github.com/iceicei/MSSC.
Jiabao Zhong, Hezhe Qiao, Lin Chen 0023, Mingsheng Shang 0001, Qun Liu 0005
IJCNN2