Yuchang Zhu

dblp:277/4757 · DBLP profile ↗
← Back
13ranked-venue papers
8as first author
13since 2021 · last 2026
0000-0001-5474-5671ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 5 first-author · 10 since 2021Databases, data management, data science and information retrieval · 4 · 4 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 GT-SNT: A Linear-Time Transformer for Large-Scale Graphs via Spiking Node Tokenization
abstract
Graph Transformers (GTs), which integrate message passing and self-attention mechanisms simultaneously, have achieved promising empirical results in graph prediction tasks. However, the design of scalable and topology-aware node tokenization has lagged behind other modalities. This gap becomes critical as the quadratic complexity of full attention renders them impractical on large-scale graphs. Recently, Spiking Neural Networks (SNNs), as brain-inspired models, provided an energy-saving scheme to convert input intensity into discrete spike-based representations through event-driven spiking neurons. Inspired by these characteristics, we propose a linear-time Graph Transformer with Spiking Node Tokenization (GT-SNT) for node classification. By integrating multi-step feature propagation with SNNs, spiking node tokenization generates compact, locality-aware spike count embeddings as node tokens to avoid predefined codebooks and their utilization issues. The codebook guided self-attention leverages these tokens to perform node-to-token attention for linear-time global context aggregation. In experiments, we compare GT-SNT with other state-of-the-art baselines on node classification datasets ranging from small to large. Experimental results show that GT-SNT achieves comparable performances on most datasets and reaches up to 130× faster inference speed compared to other GTs.
Huizhe Zhang, Jintang Li, Yuchang Zhu, Huazhen Zhong, Liang Chen 0001
AAAI3
2026 Attribute-structure decoupled self-supervised learning for graph anomaly detection
Yun Fu, Jintang Li, Yuchang Zhu, Zibin Zheng, Liang Chen 0001
Knowl. Based Syst.3
2026 SaGIF: Improving Individual Fairness in Graph Neural Networks via Similarity Encoding
abstract
Individual fairness (IF) in graph neural networks (GNNs), which emphasizes the need for similar individuals should receive similar outcomes from GNNs, has been a critical issue. Despite its importance, research in this area has been largely unexplored in terms of (1) a clear understanding of what induces individual unfairness in GNNs and (2) a comprehensive consideration of identifying similar individuals. To bridge these gaps, we conduct a preliminary analysis to explore the underlying reason for individual unfairness and observe correlations between IF andsimilarity consistency, a concept introduced to evaluate the discrepancy in identifying similar individuals based on graph structure versus node features. Inspired by our observations, we introduce two metrics to assess individual similarity from two distinct perspectives: topology fusion and feature fusion. Building upon these metrics, we proposeSimilarity-awareGNNs forIndividualFairness, namedSaGIF. The key insight behind SaGIF is the integration of individual similarities by independently learning similarity representations, leading to an improvement of IF in GNNs. Our experiments on several real-world datasets validate the effectiveness of our proposed metrics and SaGIF. Specifically, SaGIF consistently outperforms state-of-the-art IF methods while maintaining utility performance.
Yuchang Zhu, Jintang Li, Huizhe Zhang, Liang Chen 0001, Zibin Zheng
IEEE Trans. Knowl. Data Eng.1
2025 Measuring Diversity in Synthetic Datasets
abstract
Large language models (LLMs) are widely adopted to generate synthetic datasets for various natural language processing (NLP) tasks, such as text classification and summarization. However, accurately measuring the diversity of these synthetic datasets—an aspect crucial for robust model performance—remains a significant challenge. In this paper, we introduce DCScore, a novel method for measuring synthetic dataset diversity from a classification perspective. Specifically, DCScore formulates diversity evaluation as a sample classification task, leveraging mutual relationships among samples. We further provide theoretical verification of the diversity-related axioms satisfied by DCScore, highlighting its role as a principled diversity evaluation method. Experimental results on synthetic datasets reveal that DCScore enjoys a stronger correlation with multiple diversity pseudo-truths of evaluated datasets, underscoring its effectiveness. Moreover, both empirical and theoretical evidence demonstrate that DCScore substantially reduces computational costs compared to existing methods. Code is available at: https://github.com/bluewhalelab/dcscore.
Yuchang Zhu, Huizhe Zhang, Bingzhe Wu, Jintang Li, Zibin Zheng, Peilin Zhao, Liang Chen 0001, Yatao Bian
ICML1
2025 FairDLA: Improving the fairness-utility trade-off in graph neural networks via dual-level alignment
Ying Zhen, Yuchang Zhu, Jintang Li, Liang Chen 0001
Knowl. Based Syst.2
2025 Heterophily-Aware Representation Learning on Heterogeneous Graphs
abstract
Real-world graphs are typically complex, exhibiting heterogeneity in the global structure, as well as strong heterophily within local neighborhoods. While a growing body of literature has revealed the limitations of graph neural networks (GNNs) in handling homogeneous graphs with heterophily, little work has been conducted on investigating the heterophily properties in the context of heterogeneous graphs. To bridge this research gap, we identify the heterophily in heterogeneous graphs using metapaths and propose two practical metrics to quantitatively describe the levels of heterophily. Our empirical investigations on real-world heterogeneous graphs have revealed that heterogeneous graph neural networks (HGNNs), which inherit many mechanisms from GNNs designed for homogeneous graphs, struggle to generalize to heterogeneous graphs with heterophily or low levels of homophily. To address the challenge, we present Hetero$^{2}$2Net, a heterophily-aware HGNN that incorporates masked metapath prediction and masked label prediction tasks to effectively and flexibly handle both homophilic and heterophilic heterogeneous graphs. We evaluate the performance of Hetero$^{2}$2Net on five real-world heterogeneous graph benchmarks with varying levels of heterophily. Experimental results demonstrate that Hetero$^{2}$2Net outperforms strong baselines in the semi-supervised node classification task. In particular, Hetero$^{2}$2Net scales to an industrial-scale commercial graph with 13 M nodes and 157 M edges, demonstrating its effectiveness in handling large and complex heterogeneous graphs.
Jintang Li, Yuchang Zhu, Huizhe Zhang, Liang Chen 0001, Zibin Zheng
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 Lane Detection Method Based on Deformable Linear Convolution
abstract
Autonomous driving systems mainly rely on accurate detection of lane markings for navigation and safety. This paper explores an enhanced lane detection methodology employing deformable linear convolution, which dynamically adjusts to geometric variations of road markings. Our method aims to improve the detection fidelity under a range of challenging conditions such as variable illumination, road wear, and diverse weather scenarios, as evidenced by our experiments on the BDD100K dataset. The results demonstrate an improvement over those traditional lane detection techniques, suggesting that the deformable linear convolution offers a viable path forward for complex environmental adaptation in real-time image processing. Nevertheless, the computational demands of the proposed method in this paper highlight an area for further optimization. This study contributes to the field by providing an adaptable framework for the lane detection and sets the stage for future research focused on operational efficiency.
Yuchang Zhu, Nanfeng Xiao
IJCNN1
2024 Comprehensive Teacher-based Multi-task Model Train Using a Small Amount of Data
abstract
In the autonomous driving, effectively managing multiple visual perception tasks simultaneously is essential. Multi-task processing models are thus crucial in this field. The primary challenge lies in training these models, which typically require extensive and meticulously labeled datasets, a process of-ten constrained by resources. This challenge is particularly acute with multi-task datasets, where accurate labeling of numerous data samples becomes a time-intensive endeavor. To overcome this, we propose Comprehensive Teacher, a novel semi-supervised method specifically designed for multi-task model training. This method innovatively leverages the principles of pseudo-label learning and consistency regularization, traditionally used in semi-supervised learning, and adapts them to multi-task contexts. Our approach demonstrates significant versatility and efficiency in the training multi-task models. We validate Comprehensive Teacher through extensive experiments on diverse autonomous driving datasets, showcasing its superiority over conventional single-task semi-supervised methods. This approach enables the training of high-quality multi-task models with limited data availability, marking a significant step forward in resource-efficient autonomous driving technologies.
Yuchang Zhu, Nanfeng Xiao
IJCNN1
2024 One Fits All: Learning Fair Graph Neural Networks for Various Sensitive Attributes
abstract
Recent studies have highlighted fairness issues in Graph Neural Networks (GNNs), where they produce discriminatory predictions against specific protected groups categorized by sensitive attributes such as race and age. While various efforts to enhance GNN fairness have made significant progress, these approaches are often tailored to specific sensitive attributes. Consequently, they necessitate retraining the model from scratch to accommodate changes in the sensitive attribute requirement, resulting in high computational costs. To gain deeper insights into this issue, we approach the graph fairness problem from a causal modeling perspective, where we identify the confounding effect induced by the sensitive attribute as the underlying reason. Motivated by this observation, we formulate the fairness problem in graphs from an invariant learning perspective, which aims to learn invariant representations across environments. Accordingly, we propose a graph fairness framework based on invariant learning, namely FairINV, which enables the training of fair GNNs to accommodate various sensitive attributes within a single training session. Specifically, FairINV incorporates sensitive attribute partition and trains fair GNNs by eliminating spurious correlations between the label and various sensitive attributes. Experimental results on several real-world datasets demonstrate that FairINV significantly outperforms state-of-the-art fairness approaches, underscoring its effectiveness. Our code is available via: https://github.com/ZzoomD/FairINV/.
Yuchang Zhu, Jintang Li, Yatao Bian, Zibin Zheng, Liang Chen 0001
KDD1
2024 The Devil is in the Data: Learning Fair Graph Neural Networks via Partial Knowledge Distillation
abstract
Graph neural networks (GNNs) are being increasingly used in many high-stakes tasks, and as a result, there is growing attention on their fairness recently. GNNs have been shown to be unfair as they tend to make discriminatory decisions toward certain demographic groups, divided by sensitive attributes such as gender and race. While recent works have been devoted to improving their fairness performance, they often require accessible demographic information. This greatly limits their applicability in real-world scenarios due to legal restrictions. To address this problem, we present a demographic-agnostic method to learn fair GNNs via knowledge distillation, namely FairGKD. Our work is motivated by the empirical observation that training GNNs on partial data (i.e., only node attributes or topology data) can improve their fairness, albeit at the cost of utility. To make a balanced trade-off between fairness and utility performance, we employ a set of fairness experts (i.e., GNNs trained on different partial data) to construct the synthetic teacher, which distills fairer and informative knowledge to guide the learning of the GNN student. Experiments on several benchmark datasets demonstrate that FairGKD, which does not require access to demographic information, significantly improves the fairness of GNNs by a large margin while maintaining their utility.\footnoteOur code is available via: \code.
Yuchang Zhu, Jintang Li, Liang Chen 0001, Zibin Zheng
WSDM1
2024 Fair Graph Representation Learning via Sensitive Attribute Disentanglement
abstract
Group fairness for Graph Neural Networks (GNNs), which emphasizes algorithmic decisions neither favoring nor harming certain groups defined by sensitive attributes (e.g., race and gender), has gained considerable attention. In particular, the objective of group fairness is to ensure that the decisions made by GNNs are independent of the sensitive attribute. To achieve this objective, most existing approaches involve eliminating sensitive attribute information in node representations or algorithmic decisions. However, such ways may also eliminate task-related information due to its inherent correlation with the sensitive attribute, leading to a sacrifice in utility. In this work, we focus on improving the fairness of GNNs while preserving task-related information and propose a fair GNN framework named FairSAD. Instead of eliminating sensitive attribute information, FairSAD enhances the fairness of GNNs via Sensitive Attribute Disentanglement (SAD), which separates the sensitive attribute-related information into an independent component to mitigate its impact. Additionally, FairSAD utilizes a channel masking mechanism to adaptively identify the sensitive attribute-related component and subsequently decorrelates it. Overall, FairSAD minimizes the impact of the sensitive attribute on GNN outcomes rather than eliminating sensitive attributes, thereby preserving task-related information associated with the sensitive attribute. Furthermore, experiments conducted on several real-world datasets demonstrate that FairSAD outperforms other state-of-the-art methods by a significant margin in terms of both fairness and utility performance. Our source code is available at https://github.com/ZzoomD/FairSAD.
Yuchang Zhu, Jintang Li, Zibin Zheng, Liang Chen 0001
WWW1
2024 FairAGG: Toward Fair Graph Neural Networks via Fair Aggregation
abstract
Graph neural networks (GNNs) have shown intrinsic topology bias inherited from graph-structured data, where a majority of nodes are associated with specific sensitive attributes (e.g., age, and race). In this regard, GNNs make discriminatory decisions toward certain groups defined by the sensitive attribute. Over the past few years, efforts have been made to mitigate the fairness issue of GNNs caused by topology bias. While achieving impressive results, these works likely only scratch the surface of what is possible with well-tuned GNN models or heuristic graph preprocessing techniques. Despite modern GNNs being built upon the message-passing framework, few works dig into the hidden reasoning behind the fairness issue of GNNs from such a fundamental perspective. In this work, we empirically demonstrate that message aggregation with higher edge weight for intergroup edges improves model fairness but sacrifices utility. The above observations motivate us to propose and derive a simple yet effective message-passing scheme,FairAGG, leading to more fair GNNs with less comprised downstream performance. Specifically, FairAGG measures the contribution of graph topology on fairness usingShapley value, which facilitates fair aggregation through reweighting. Experiments on several real-world datasets demonstrate that FairAGG enhances the fairness of the GNNs model while maintaining competitive utility performance.
Yuchang Zhu, Jintang Li, Liang Chen 0001, Zibin Zheng
IEEE Trans. Comput. Soc. Syst.1
2022 AutoMine: An Unmanned Mine Dataset
abstract
Autonomous driving datasets have played an important role in validating the advancement of intelligent vehicle algorithms including localization, perception and prediction in academic areas. However, current existing datasets pay more attention to the structured urban road, which hampers the exploration on unstructured special scenarios. Moreover, the open-pit mine is one of the typical representatives for them. Therefore, we introduce the Autonomous driving dataset on the Mining scene (AutoMine) for positioning and perception tasks in this paper. The AutoMine is collected by multiple acquisition platforms including an SUV, a wide-body mining truck and an ordinary mining truck, depending on the actual mine operation scenarios. The dataset consists of 18+ driving hours, 18K annotated lidar and image frames for 3D perception with various mines, time-of-the-day and weather conditions. The main contributions of the AutoMine dataset are as follows: I.The first autonomous driving dataset for perception and localization in mine scenarios. 2.There are abundant dynamic obstacles of 9 degrees of freedom with large dimension difference (mining trucks and pedestrians) and extreme climatic conditions (the dust and snow) in the mining area. 3.Multi-platform acquisition strategies could capture mining data from multiple perspectives that fit the actual operation. More details can be found in our website(https://automine.cc).
Yuchen Li 0004, Siyu Teng, Yu Zhang 0109, Yuchang Zhu, Dongpu Cao, Bin Tian 0003, Yunfeng Ai, Zhe Xuanyuan, Long Chen 0005
CVPR6