Shengzhong Zhang

dblp:255/8703 · DBLP profile ↗
← Back
14ranked-venue papers
3as first author
13since 2021 · last 2026
0000-0003-1783-6835ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 3 first-author · 10 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021
YearPublicationVenuePosition
2026 Know Your Neighbors: Subgraph Importance Sampling for Heterophilic Graph Active Learning
abstract
Graph neural networks (GNNs) have demonstrated strong performance in various graph mining tasks but rely heavily on extensively labeled nodes. To improve training efficiency, graph active learning (GAL) has emerged as a solution for selecting the most informative nodes for labeling. However, existing GAL methods are primarily designed for homophilic graphs, where nodes with the same labels are more likely to be connected. In this work, we systematically study active learning on heterophilic graphs, a setting that has received limited attention. Surprisingly, we observe that existing GAL methods fail to consistently outperform random sampling on heterophilic graphs. Through an in-depth investigation, we reveal that these methods implicitly assume homophily even on heterophilic graphs, leading to suboptimal performance. To address this issue, we introduce the principle of "Know Your Neighbors" and propose an active learning algorithm KyN specifically for heterophilic graphs. The core idea of KyN is to provide GNNs with accurate estimations of homophily distribution by labeling nodes together with their neighbors. We implement KyN based on subgraph sampling with probabilities proportional to l1 Lewis weights, which is supported by solid theoretical guarantees. Extensive experiments on diverse real-world datasets, including a large heterophilic graph with over 2 million nodes, demonstrate the effectiveness and scalability of KyN.
Wenjie Yang 0006, Shengzhong Zhang, Jiaxing Guo, Tongshan Xu, Zengfeng Huang
AAAI2
2026 Reflect Then Learn: Active Prompting for Information Extraction Guided by Introspective Confusion
abstract
Large Language Models (LLMs) show remarkable potential for few-shot information extraction (IE), yet their performance is highly sensitive to the choice of in-context examples. Conventional selection strategies often fail to provide informative guidance, as they overlook a key source of model fallibility: confusion stemming not just from semantic content, but also from the generation of well-structured formats required by IE tasks. To address this, we introduce Active Prompting for Information Extraction (APIE), a novel active prompting framework guided by a principle we term introspective confusion. Our method empowers an LLM to assess its own confusion through a dual-component uncertainty metric that uniquely quantifies both Format Uncertainty (difficulty in generating correct syntax) and Content Uncertainty (inconsistency in extracted semantics). By ranking unlabeled data with this comprehensive score, our framework actively selects the most challenging and informative samples to serve as few-shot exemplars. Extensive experiments on four benchmarks show that our approach consistently outperforms strong baselines, yielding significant improvements in both extraction accuracy and robustness. Our work highlights the critical importance of a fine-grained, dual-level view of model uncertainty when it comes to building effective and reliable structured generation systems.
Dong Zhao 0012, Xiang Chen 0016, Chuanxing Geng, Shengzhong Zhang, Shaoyuan Li, Sheng-Jun Huang
AAAI7
2026 Cross-level graph contrastive learning for community value prediction
Wenjie Yang 0006, Shengzhong Zhang, Zengfeng Huang
Neural Networks2
2025 Your Graph Recommenders are Provably Doing Graph Contrastive Learning
abstract
Graph recommender (GR) is a type of graph neural network (GNN) encoder that is customized for extracting information from the user-item interaction graph. Due to its strong performance on the recommendation task, GR has gained significant attention recently. Graph contrastive learning (GCL) is also a popular research direction that aims to learn, often unsupervised, GNNs with certain contrastive objectives. As general graph representation learning methods, GCLs have been widely adopted with supervised recommendation loss for joint training of GRs. Despite the intersection of GR and GCL research, theoretical understanding of the relationship between the two fields is surprisingly sparse. This vacancy inevitably leads to inefficient scientific research.
Wenjie Yang 0006, Shengzhong Zhang, Jiaxing Guo, Zengfeng Huang
KDD (2)2
2025 Geometric Imbalance in Semi-Supervised Node Classification
abstract
Class imbalance in graph data presents a significant challenge for effective node classification, particularly in semi-supervised scenarios. In this work, we formally introduce the concept of geometric imbalance, which captures how message passing on class-imbalanced graphs leads to geometric ambiguity among minority-class nodes in the riemannian manifold embedding space. We provide a rigorous theoretical analysis of geometric imbalance on the riemannian manifold and propose a unified framework that explicitly mitigates it through pseudo-label alignment, node reordering, and ambiguity filtering. Extensive experiments on diverse benchmarks show that our approach consistently outperforms existing methods, especially under severe class imbalance. Our findings offer new theoretical insights and practical tools for robust semi-supervised node classification.
Shengzhong Zhang, Bisheng Li, Menglin Yang 0001, Min Zhou 0006, Weiyang Ding, Yutong Xie 0001, Zengfeng Huang
NeurIPS2
2025 Contra2: A one-step active learning method for imbalanced graphs
Wenjie Yang 0006, Shengzhong Zhang, Jiaxing Guo, Zengfeng Huang
Artif. Intell.2
2025 Graph Batch Coarsening framework for scalable graph neural networks
Shengzhong Zhang, Bisheng Li, Wenjie Yang 0006, Min Zhou 0006, Zengfeng Huang
Neural Networks1
2024 Enhancing Performance of Coarsened Graphs with Gradient-Matching
abstract
Graph Neural Networks (GNNs) are powerful tools for processing graph data, but training them is typically time-consuming and expensive in terms of GPU memory. Recently, graph reduction techniques have been investigated as a pre-processing step, due to the fact that the time and space complexity of GNNs training is only sublinear on the reduced graphs. Among them, graph coarsening has gained significant attention. By generating a coarse graph while preserving graph structure properties, it can effectively reduce the size of nodes by up to a factor of ten without significantly compromising performance. However, the coarsening phase for this method is purely heuristic, leaving room for improvement. In this paper, we propose Learnable Graph Coarsening with Gradient Matching (LGCGM), boosting graph coarsening with (downstream)task-specific dataset condensation objective. To the best of our knowledge, LGCGM is the first work that incorporates downstream-task information for graph coarsening. Extensive experiments demonstrate that this additional information efficiently enhances the quality of coarse graphs and outperforms previous graph condensation methods on large graphs.
Wenjie Yang 0006, Shengzhong Zhang, Zengfeng Huang
ICASSP2
2024 StructComp: Substituting propagation with Structural Compression in Training Graph Contrastive Learning
abstract
Graph contrastive learning (GCL) has become a powerful tool for learning graph data, but its scalability remains a significant challenge. In this work, we propose a simple yet effective training framework called Structural Compression (StructComp) to address this issue. Inspired by a sparse low-rank approximation on the diffusion matrix, StructComp trains the encoder with the compressed nodes. This allows the encoder not to perform any message passing during the training stage, and significantly reduces the number of sample pairs in the contrastive loss. We theoretically prove that the original GCL loss can be approximated with the contrastive loss computed by StructComp. Moreover, StructComp can be regarded as an additional regularization term for GCL models, resulting in a more robust encoder. Empirical studies on various datasets show that StructComp greatly reduces the time and memory consumption while improving model performance compared to the vanilla GCL models and scalable training methods.
Shengzhong Zhang, Wenjie Yang 0006, Zengfeng Huang
ICLR1
2023 Rethinking Semi-Supervised Imbalanced Node Classification from Bias-Variance Decomposition
abstract
This paper introduces a new approach to address the issue of class imbalance in graph neural networks (GNNs) for learning on graph-structured data. Our approach integrates imbalanced node classification and Bias-Variance Decomposition, establishing a theoretical framework that closely relates data imbalance to model variance. We also leverage graph augmentation technique to estimate the variance and design a regularization term to alleviate the impact of imbalance. Exhaustive tests are conducted on multiple benchmarks, including naturally imbalanced datasets and public-split class-imbalanced datasets, demonstrating that our approach outperforms state-of-the-art methods in various imbalanced scenarios. This work provides a novel theoretical perspective for addressing the problem of imbalanced node classification in GNNs.
Divin Yan, Gengchen Wei, Shengzhong Zhang, Zengfeng Huang
NeurIPS4
2023 Effective stabilized self-training on few-labeled graph data
Ziang Zhou, Jieming Shi 0001, Shengzhong Zhang, Zengfeng Huang, Qing Li 0001
Inf. Sci.3
2022 BSAL: A Framework of Bi-component Structure and Attribute Learning for Link Prediction
abstract
Given the ubiquitous existence of graph-structured data, learning the representations of nodes for the downstream tasks ranging from node classification, link prediction to graph classification is of crucial importance. Regarding missing link inference of diverse networks, we revisit the link prediction techniques and identify the importance of both the structural and attribute information. However, the available techniques either heavily count on the network topology which is spurious in practice, or cannot integrate graph topology and features properly. To bridge the gap, we propose a bicomponent structural and attribute learning framework (BSAL) that is designed to adaptively leverage information from topology and feature spaces. Specifically, BSAL constructs a semantic topology via the node attributes and then gets the embeddings regarding the semantic view, which provides a flexible and easy-to-implement solution to adaptively incorporate the information carried by the node attributes. Then the semantic embedding together with topology embedding are fused together using attention mechanism for the final prediction. Extensive experiments show the superior performance of our proposal and it significantly outperforms baselines on diverse research benchmarks.
Bisheng Li, Min Zhou 0006, Shengzhong Zhang, Menglin Yang 0001, Defu Lian, Zengfeng Huang
SIGIR3
2021 Scaling Up Graph Neural Networks Via Graph Coarsening
abstract
Scalability of graph neural networks remains one of the major challenges in graph machine learning. Since the representation of a node is computed by recursively aggregating and transforming representation vectors of its neighboring nodes from previous layers, the receptive fields grow exponentially, which makes standard stochastic optimization techniques ineffective. Various approaches have been proposed to alleviate this issue, e.g., sampling-based methods and techniques based on pre-computation of graph filters.
Zengfeng Huang, Shengzhong Zhang, Chong Xi, Min Zhou 0006
KDD2
2020 SCE: Scalable Network Embedding from Sparsest Cut
abstract
Large-scale network embedding is to learn a latent representation for each node in an unsupervised manner, which captures inherent properties and structural information of the underlying graph. In this field, many popular approaches are influenced by the skip-gram model from natural language processing. Most of them use a contrastive objective to train an encoder which forces the embeddings of similar pairs to be close and embeddings of negative samples to be far. A key of success to such contrastive learning methods is how to draw positive and negative samples. While negative samples that are generated by straightforward random sampling are often satisfying, methods for drawing positive examples remains a hot topic.
Shengzhong Zhang, Zengfeng Huang, Haicang Zhou, Ziang Zhou
KDD1