Hang Gao 0004

dblp:16/6086-4 · DBLP profile ↗
← Back
18ranked-venue papers
9as first author
18since 2021 · last 2025
0000-0003-3613-4011ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 9 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Bootstrapping Heterogeneous Graph Representation Learning via Large Language Models: A Generalized Approach
abstract
Graph representation learning methods are highly effective in handling complex non-Euclidean data by capturing intricate relationships and features within graph structures. However, traditional methods face challenges when dealing with heterogeneous graphs that contain various types of nodes and edges due to the diverse sources and complex nature of the data. Existing heterogeneous graph neural networks (HGNNs) have shown promising results but require prior knowledge of node and edge types and unified node feature formats, which limits their applicability. Recent advancements in graph representation learning using large language models (LLMs) offer new solutions by integrating LLMs' data processing capabilities, enabling the alignment of various graph representations. Nevertheless, these methods often overlook heterogeneous graph data and require extensive preprocessing. To address these limitations, we propose an LLM-enhanced Heterogeneous Graph Neural Network (LHGNN). LHGNN leverages the strengths of both LLM and GNN, allowing for the processing of graph data with any format and type of nodes and edges without the need for type information or special preprocessing. LHGNN employs LLM to automatically summarize and classify different data formats and types, aligns node features, and uses a specialized GNN for targeted learning, thus obtaining effective graph representations for downstream tasks. Theoretical analysis and experimental validation have demonstrated the effectiveness of our method.
Hang Gao 0004, Fengge Wu, Changwen Zheng, Junsuo Zhao, Huaping Liu 0001
AAAI1
2025 LLM Enhancers for GNNs: An Analysis from the Perspective of Causal Mechanism Identification
abstract
The use of large language models (LLMs) as feature enhancers to optimize node representations, which are then used as inputs for graph neural networks (GNNs), has shown significant potential in graph representation learning. However, the fundamental properties of this approach remain underexplored. To address this issue, we propose conducting a more in-depth analysis of this issue based on the interchange intervention method. First, we construct a synthetic graph dataset with controllable causal relationships, enabling precise manipulation of semantic relationships and causal modeling to provide data for analysis. Using this dataset, we conduct interchange interventions to examine the deeper properties of LLM enhancers and GNNs, uncovering their underlying logic and internal mechanisms. Building on the analytical results, we design a plug-and-play optimization module to improve the information transfer between LLM enhancers and GNNs. Experiments across multiple datasets and models validate the proposed module.
Hang Gao 0004, Fengge Wu, Junsuo Zhao, Changwen Zheng, Huaping Liu 0001
ICML1
2025 Bootstrapping Heterophily Graph Representation Learning via a Large Language Model-Based Approach
abstract
Graph Neural Networks (GNNs) excel in homophilic graphs but struggle with heterophilic graphs where connected nodes exhibit divergent characteristics. Existing methods underutilize semantic information in node attributes, a key factor for interpreting heterophilic interactions. To bridge this gap, we propose LLM-Enhanced Graph Neural Network for Heterophily (LEGNNH), a novel framework that synergizes large language model (LLM) with spectral graph convolutions for node classification in heterophilic graphs. LEGNNH operates in three stages: (1) Task-aware semantic embedding using LLMs with instruction-based prompting to encode raw node text; (2) Multi-channel spectral filtering dynamically aggregating local patterns via adaptive low-pass, high-pass, and identity filters; (3) A hierarchical attention module selectively integrates higher-order neighborhood semantics while suppressing noise propagation from label-discordant connections. Addressing the scarcity of heterophilic benchmarks, we contribute two text-attributed graph datasets (YelpNYC and Amazon-Review) with explicit feature-label heterophily. Extensive experiments on seven benchmarks show LEGNNH achieves state-of-the-art performance (average rank 1.86 across 10 baselines), outperforming the best heterophily-specific model by$\mathbf{1. 3 1 - 5. 3 3 \%}$accuracy on key datasets.
Fengge Wu, Hang Gao 0004, Junsuo Zhao
ICTAI3
2025 Learn to Think: Bootstrapping LLM Logic Through Graph Representation Learning
abstract
Large Language Models (LLMs) have achieved remarkable success across various domains. However, they still face significant challenges, including high computational costs for training and limitations in solving complex reasoning problems. Although existing methods have extended the reasoning capabilities of LLMs through structured paradigms, these approaches often rely on task-specific prompts and predefined reasoning processes, which constrain their flexibility and generalizability. To address these limitations, we propose a novel framework that leverages graph learning to enable more flexible and adaptive reasoning capabilities for LLMs. Specifically, this approach models the reasoning process of a problem as a graph and employs LLM-based graph learning to guide the adaptive generation of each reasoning step. To further enhance the adaptability of the model, we introduce a Graph Neural Network (GNN) module to perform representation learning on the generated reasoning process, enabling real-time adjustments to both the model and the prompt. Experimental results demonstrate that this method significantly improves reasoning performance across multiple tasks without requiring additional training or task-specific prompt design. Code can be found in https://github.com/zch65458525/L2T.
Hang Gao 0004, Junsuo Zhao, Fengge Wu, Changwen Zheng, Huaping Liu 0001
IJCAI1
2024 Rethinking Causal Relationships Learning in Graph Neural Networks
abstract
Graph Neural Networks (GNNs) demonstrate their significance by effectively modeling complex interrelationships within graph-structured data. To enhance the credibility and robustness of GNNs, it becomes exceptionally crucial to bolster their ability to capture causal relationships. However, despite recent advancements that have indeed strengthened GNNs from a causal learning perspective, conducting an in-depth analysis specifically targeting the causal modeling prowess of GNNs remains an unresolved issue. In order to comprehensively analyze various GNN models from a causal learning perspective, we constructed an artificially synthesized dataset with known and controllable causal relationships between data and labels. The rationality of the generated data is further ensured through theoretical foundations. Drawing insights from analyses conducted using our dataset, we introduce a lightweight and highly adaptable GNN module designed to strengthen GNNs' causal learning capabilities across a diverse range of tasks. Through a series of experiments conducted on both synthetic datasets and other real-world datasets, we empirically validate the effectiveness of the proposed module. The codes are available at https://github.com/yaoyao-yaoyao-cell/CRCG.
Hang Gao 0004, Chengyu Yao, Jiangmeng Li, Lingyu Si, Fengge Wu, Changwen Zheng, Huaping Liu 0001
AAAI1
2024 Hierarchical Topology Isomorphism Expertise Embedded Graph Contrastive Learning
abstract
Graph contrastive learning (GCL) aims to align the positive features while differentiating the negative features in the latent space by minimizing a pair-wise contrastive loss. As the embodiment of an outstanding discriminative unsupervised graph representation learning approach, GCL achieves impressive successes in various graph benchmarks. However, such an approach falls short of recognizing the topology isomorphism of graphs, resulting in that graphs with relatively homogeneous node features cannot be sufficiently discriminated. By revisiting classic graph topology recognition works, we disclose that the corresponding expertise intuitively complements GCL methods. To this end, we propose a novel hierarchical topology isomorphism expertise embedded graph contrastive learning, which introduces knowledge distillations to empower GCL models to learn the hierarchical topology isomorphism expertise, including the graph-tier and subgraph-tier. On top of this, the proposed method holds the feature of plug-and-play, and we empirically demonstrate that the proposed method is universal to multiple state-of-the-art GCL models. The solid theoretical analyses are further provided to prove that compared with conventional GCL methods, our method acquires the tighter upper bound of Bayes classification error. We conduct extensive experiments on real-world benchmarks to exhibit the performance superiority of our method over candidate GCL methods, e.g., for the real-world graph representation learning experiments, the proposed method beats the state-of-the-art method by 0.23% on unsupervised representation learning setting, 0.43% on transfer learning setting. Our code is available at https://github.com/jyf123/HTML.
Jiangmeng Li, Hang Gao 0004, Wenwen Qiang, Changwen Zheng, Fuchun Sun 0001
AAAI3
2024 Discovering Symmetry Breaking in Physical Systems with Relaxed Group Convolution
abstract
Modeling symmetry breaking is essential for understanding the fundamental changes in the behaviors and properties of physical systems, from microscopic particle interactions to macroscopic phenomena like fluid dynamics and cosmic structures. Thus, identifying sources of asymmetry is an important tool for understanding physical systems. In this paper, we focus on learning asymmetries of data using relaxed group convolutions. We provide both theoretical and empirical evidence that this flexible convolution technique allows the model to maintain the highest level of equivariance that is consistent with data and discover the subtle symmetry-breaking factors in various physical systems. We employ various relaxed group convolution architectures to uncover various symmetry-breaking factors that are interpretable and physically meaningful in different physical systems, including the phase transition of crystal structure, the isotropy and homogeneity breaking in turbulent flow, and the time-reversal symmetry breaking in pendulum systems.
Rui Wang 0086, Elyssa F. Hofgard, Hang Gao 0004, Robin Walters 0001, Tess E. Smidt
ICML3
2024 Learning Node Representations Under Partial Label Learning
abstract
Node classification is a crucial area in graph representation learning, with significant applications in real-world network scenarios. However, due to the complexity of the relationships among nodes in the topological graph, precise data labeling is often challenging, leading to a significant amount of label noise. Partial Label Learning (PLL) is a weakly supervised learning problem designed to accommodate label noise. Therefore, we introduce PLL to address the issue of label noise. Currently, Graph Neural Networks (GNNs) are the primary method for addressing node classification problems. However, existing research has demonstrated that GNNs tend to amplify the similarity between node features, posing challenges for label disambiguation in partial label scenarios. To address this issue, We conduct an analysis of the composition of node features and propose a novel method, which aims to enhance feature quality and reduce node feature similarity in partial label scenarios of node classification. Extensive experiments on challenging homogeneous graph datasets indicate that PLNR achieves state-of-the-art performance and demonstrates comparable results to fully supervised learning.
Jiaguo Yuan, Hang Gao 0004, Fengge Wu, Junsuo Zhao
IJCNN2
2024 Molecular Graph Representation Learning via Structural Similarity Information
Chengyu Yao, Hong Huang 0004, Hang Gao 0004, Fengge Wu, Haiming Chen 0001, Junsuo Zhao
ECML/PKDD (3)3
2024 Introducing diminutive causal structure into graph representation learning
abstract
When engaging in end-to-end graph representation learning with Graph Neural Networks (GNNs), the intricate causal relationships and rules inherent in graph data pose a formidable challenge for the model in accurately capturing authentic data relationships. A proposed mitigating strategy involves the direct integration of rules or relationships corresponding to the graph data into the model. However, within the domain of graph representation learning, the inherent complexity of graph data obstructs the derivation of a comprehensive causal structure that encapsulates universal rules or relationships governing the entire dataset. Instead, only specialized diminutive causal structures, delineating specific causal relationships within constrained subsets of graph data, emerge as discernible. Motivated by empirical insights, it is observed that GNN models exhibit a tendency to converge towards such specialized causal structures during the training process. Consequently, we posit that the introduction of these specific causal structures is advantageous for the training of GNN models. Building upon this proposition, we introduce a novel method that enables GNN models to glean insights from these specialized diminutive causal structures, thereby enhancing overall performance. Our method specifically extracts causal knowledge from the model representation of these diminutive causal structures and incorporates interchange intervention to optimize the learning process. Theoretical analysis serves to corroborate the efficacy of our proposed method. Furthermore, empirical experiments consistently demonstrate significant performance improvements across diverse datasets.
Hang Gao 0004, Peng Qiao, Fengge Wu, Jiangmeng Li, Changwen Zheng
Knowl. Based Syst.1
2024 Unsupervised social event detection via hybrid graph contrastive learning and reinforced incremental clustering
abstract
Detecting events from social media data streams is gradually attracting researchers. The innate challenge for detecting events is to extract discriminative information from social media data thereby assigning the data into different events. Due to the excessive diversity and high updating frequency of social data, using supervised approaches to detect events from social messages is hardly achieved. To this end, recent works explore learning discriminative information from social messages by leveraging graph contrastive learning (GCL) and embedding clustering in an unsupervised manner. However, two intrinsic issues exist in benchmark methods: conventional GCL can only roughly explore partial attributes, thereby insufficiently learning the discriminative information of social messages; for benchmark methods, the learned embeddings are clustered in the latent space by taking advantage of certain specific prior knowledge , which conflicts with the principle of unsupervised learning paradigm . In this paper, we propose a novel unsupervised social media event detection method via hybrid graph contrastive learning and reinforced incremental clustering (HCRC), which uses hybrid graph contrastive learning to comprehensively learn semantic and structural discriminative information from social messages and reinforced incremental clustering to perform efficient clustering in a solidly unsupervised manner. We conduct comprehensive experiments to evaluate HCRC on the Twitter and Maven datasets. The experimental results demonstrate that our approach yields consistent significant performance boosts. In traditional incremental setting, semi-supervised incremental setting and solidly unsupervised setting, the model performance has achieved maximum improvements of 53%, 45%, and 37%, respectively.
Zehua Zang, Hang Gao 0004, Rui Wang 0086, Jiangmeng Li
Knowl. Based Syst.3
2023 Robust Causal Graph Representation Learning against Confounding Effects
abstract
The prevailing graph neural network models have achieved significant progress in graph representation learning. However, in this paper, we uncover an ever-overlooked phenomenon: the pre-trained graph representation learning model tested with full graphs underperforms the model tested with well-pruned graphs. This observation reveals that there exist confounders in graphs, which may interfere with the model learning semantic information, and current graph representation learning methods have not eliminated their influence. To tackle this issue, we propose Robust Causal Graph Representation Learning (RCGRL) to learn robust graph representations against confounding effects. RCGRL introduces an active approach to generate instrumental variables under unconditional moment restrictions, which empowers the graph representation learning model to eliminate confounders, thereby capturing discriminative information that is causally related to downstream predictions. We offer theorems and proofs to guarantee the theoretical effectiveness of the proposed approach. Empirically, we conduct extensive experiments on a synthetic dataset and multiple benchmark datasets. Experimental results demonstrate the effectiveness and generalization ability of RCGRL. Our codes are available at https://github.com/hang53/RCGRL.
Hang Gao 0004, Jiangmeng Li, Wenwen Qiang, Lingyu Si, Changwen Zheng, Fuchun Sun 0001
AAAI1
2023 SkaNet: Split Kernel Attention Network
Lipeng Chen, Daixi Jia, Hang Gao 0004, Fengge Wu, Junsuo Zhao
ICANN (5)3
2023 Introducing Semantic-Based Receptive Field into Semantic Segmentation via Graph Neural Networks
Daixi Jia, Hang Gao 0004, Xingzhe Su, Fengge Wu, Junsuo Zhao
ICONIP (6)2
2023 Information theory-guided heuristic progressive multi-view coding
Jiangmeng Li, Hang Gao 0004, Wenwen Qiang, Changwen Zheng
Neural Networks2
2022 Weight-Aware Graph Contrastive Learning
Hang Gao 0004, Jiangmeng Li, Peng Qiao, Changwen Zheng
ICANN (2)1
2022 Bootstrapping Informative Graph Augmentation via A Meta Learning Approach
abstract
Recent works explore learning graph representations in a self-supervised manner. In graph contrastive learning, benchmark methods apply various graph augmentation approaches. However, most of the augmentation methods are non-learnable, which causes the issue of generating unbeneficial augmented graphs. Such augmentation may degenerate the representation ability of graph contrastive learning methods. Therefore, we motivate our method to generate augmented graph with a learnable graph augmenter, called MEta Graph Augmentation (MEGA). We then clarify that a "good" graph augmentation must have uniformity at the instance-level and informativeness at the feature-level. To this end, we propose a novel approach to learning a graph augmenter that can generate an augmentation with uniformity and informativeness. The objective of the graph augmenter is to promote our feature extraction network to learn a more discriminative feature representation, which motivates us to propose a meta-learning paradigm. Empirically, the experiments across multiple benchmark datasets demonstrate that MEGA outperforms the state-of-the-art methods in graph self-supervised learning tasks. Further experimental studies prove the effectiveness of different terms of MEGA. Our codes are available at https://github.com/hang53/MEGA.
Hang Gao 0004, Jiangmeng Li, Wenwen Qiang, Lingyu Si, Fuchun Sun 0001, Changwen Zheng
IJCAI1
2022 Self-supervised Graph Learning with Segmented Graph Channels
Hang Gao 0004, Jiangmeng Li, Changwen Zheng
ECML/PKDD (2)1