Deyu Bo

dblp:258/0824 · DBLP profile ↗
← Back
12ranked-venue papers
6as first author
10since 2021 · last 2025
0000-0003-2063-8223ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 5 first-author · 8 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Graph learning · 70% Representation and self-supervised learning · 16% Efficient and distributed learning · 12%
Databases, data mining, and information retrieval
1 paper
Data mining · 100%

Topics — the 21 heaviest of 23, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Graph learning
graph neural network
2.242023
Specformer: Spectral Graph Neural Networks Meet Transformers · ICLR 2023
Regularizing Graph Neural Networks via Consistency-Diversity Graph Augmentations · AAAI 2022
Beyond Low-frequency Information in Graph Convolutional Networks · AAAI 2021
Machine learning › Efficient and distributed learning
dataset distillation
1.622025
Point Cloud Dataset Distillation · ICML 2025
Graph Distillation with Eigenbasis Matching · ICML 2024
Machine learning › Graph learning › graph neural network
graph convolutional network
1.432021
Beyond Low-frequency Information in Graph Convolutional Networks · AAAI 2021
Structural Deep Clustering Network · WWW 2020
AM-GCN: Adaptive Multi-channel Graph Convolutional Networks · KDD 2020
Machine learning › Representation and self-supervised learning › contrastive learning
graph contrastive learning
1.222023
Graph Contrastive Learning with Stable and Scalable Spectral Encoding · NeurIPS 2023
Revisiting Graph Contrastive Learning from the Perspective of Graph Spectrum · NeurIPS 2022
Machine learning › Graph learning › graph neural network › node classification
semi-supervised node classification
1.022022
Regularizing Graph Neural Networks via Consistency-Diversity Graph Augmentations · AAAI 2022
AM-GCN: Adaptive Multi-channel Graph Convolutional Networks · KDD 2020
Machine learning › Graph learning › graph neural network training
graph distillation
0.812024
Graph Distillation with Eigenbasis Matching · ICML 2024
Machine learning › Graph learning
graph representation learning
0.712023
Graph Contrastive Learning with Stable and Scalable Spectral Encoding · NeurIPS 2023
Machine learning › Graph learning › graph neural network
graph transformer
0.712023
Specformer: Spectral Graph Neural Networks Meet Transformers · ICLR 2023
Machine learning › Graph learning › graph neural network
spectral graph neural network
0.712023
Specformer: Spectral Graph Neural Networks Meet Transformers · ICLR 2023
Machine learning › Representation and self-supervised learning
contrastive learning
0.612022
Revisiting Graph Contrastive Learning from the Perspective of Graph Spectrum · NeurIPS 2022
Machine learning › Graph learning
graph augmentation
0.612022
Revisiting Graph Contrastive Learning from the Perspective of Graph Spectrum · NeurIPS 2022
Machine learning › Graph learning › graph neural network
graph data augmentation
0.612022
Regularizing Graph Neural Networks via Consistency-Diversity Graph Augmentations · AAAI 2022
Machine learning › Graph learning
graph signal processing
0.512021
Beyond Low-frequency Information in Graph Convolutional Networks · AAAI 2021
Machine learning › Graph learning › graph neural network
node classification
0.412020
AM-GCN: Adaptive Multi-channel Graph Convolutional Networks · KDD 2020
Machine learning › Representation and self-supervised learning › representation learning
structured representation learning
0.412020
Structural Deep Clustering Network · WWW 2020
Data mining
clustering
0.412020
Structural Deep Clustering Network · WWW 2020
Data mining › clustering
deep clustering
0.412020
Structural Deep Clustering Network · WWW 2020
Computer vision › 3D vision
point cloud processing
0.312025
Point Cloud Dataset Distillation · ICML 2025
Machine learning › Learning paradigms › semi-supervised learning
consistency regularization
0.212022
Regularizing Graph Neural Networks via Consistency-Diversity Graph Augmentations · AAAI 2022
Machine learning › Graph learning › graph representation learning
node representation learning
0.112021
Beyond Low-frequency Information in Graph Convolutional Networks · AAAI 2021
Machine learning › Graph learning › graph neural network › deep graph neural network
over-smoothing
0.112021
Beyond Low-frequency Information in Graph Convolutional Networks · AAAI 2021

Methods — techniques the papers use, named apart from their topics

rotation-invariant feature matching · 0.9point-wise generator · 0.9spectral approximation · 0.8eigenbasis matching · 0.8transformer · 0.7spectral graph theory · 0.7graph neural network · 0.7contrastive learning · 0.7label propagation · 0.6consistency regularization · 0.6self-supervised learning · 0.4graph convolutional network · 0.4autoencoder · 0.4
YearPublicationVenuePosition
2025 Point Cloud Dataset Distillation
abstract
This study introduces dataset distillation (DD) tailored for 3D data, particularly point clouds. DD aims to substitute large-scale real datasets with a small set of synthetic samples while preserving model performance. Existing methods mainly focus on structured data such as images. However, adapting DD for unstructured point clouds poses challenges due to their diverse orientations and resolutions in 3D space. To address these challenges, we theoretically demonstrate the importance of matching rotation-invariant features between real and synthetic data for 3D distillation. We further propose a plug-and-play point cloud rotator to align the point cloud to a canonical orientation, facilitating the learning of rotation-invariant features by all point cloud models. Furthermore, instead of optimizing fixed-size synthetic data directly, we devise a point-wise generator to produce point clouds at various resolutions based on the sampled noise amount. Compared to conventional DD methods, the proposed approach, termed DD3D, enables efficient training on low-resolution point clouds while generating high-resolution data for evaluation, thereby significantly reducing memory requirements and enhancing model scalability. Extensive experiments validate the effectiveness of DD3D in shape classification and part segmentation tasks across diverse scenarios, such as cross-architecture and cross-resolution settings.
Deyu Bo, Xinchao Wang
ICML1
2025 Graph Positional Autoencoders as Self-supervised Learners
abstract
Graph self-supervised learning seeks to learn effective graph representations without relying on labeled data. Among various approaches, graph autoencoders (GAEs) have gained significant attention for their efficiency and scalability. Typically, GAEs take incomplete graphs as input and predict missing elements, such as masked node features or edges. Although effective, our experimental investigation reveals that traditional feature or edge masking paradigms primarily capture low-frequency signals in the graph and fail to learn expressive structural information. To address these issues, we propose Graph Positional Autoencoders (GraphPAE), which employ a dual-path architecture to reconstruct both node features and positions. Specifically, the feature path uses positional encoding to enhance the message-passing processing, improving the GAEs' ability to predict the corrupted information. The position path, on the other hand, leverages node representations to refine positions and approximate eigenvectors, thereby enabling the encoder to learn diverse frequency information. We conduct extensive experiments to verify the effectiveness of GraphPAE, including heterophilic node classification, graph property prediction, and transfer learning. The results demonstrate that GraphPAE achieves state-of-the-art performance and consistently outperforms the baselines by a large margin.
Yang Liu 0348, Deyu Bo, Wenxuan Cao, Yuan Fang 0001, Yawen Li 0001, Chuan Shi 0001
KDD (2)2
2025 Data-Centric Graph Learning: A Survey
abstract
The history of artificial intelligence (AI) has witnessed the significant impact of high-quality data on various deep learning models, such as ImageNet for AlexNet and ResNet. Recently, instead of designing more complex neural architectures as model-centric approaches, the attention of AI community has shifted to data-centric ones, which focuses on better processing data to strengthen the ability of neural models. Graph learning, which operates on ubiquitous topological data, also plays an important role in the era of deep learning. In this survey, we comprehensively review graph learning approaches from the data-centric perspective, and aim to answer three crucial questions:(1) when to modify graph data,(2) what part of the graph data needs modificationto unlock the potential of various graph models, and(3) how to safeguard graph modelsfrom problematic data influence. Accordingly, we propose a novel taxonomy based on the stages in the graph learning pipeline, and highlight the processing methods for different data structures in the graph data, i.e., topology, feature and label. Furthermore, we analyze some potential problems embedded in graph data and discuss how to solve them in a data-centric manner. Finally, we provide some promising future directions for data-centric graph learning.
Deyu Bo, Cheng Yang 0002, Zhongjian Zhang, Jixi Liu, Yufei Peng, Chuan Shi 0001
IEEE Trans. Big Data2
2024 Graph Distillation with Eigenbasis Matching
abstract
The increasing amount of graph data places requirements on the efficient training of graph neural networks (GNNs). The emerging graph distillation (GD) tackles this challenge by distilling a small synthetic graph to replace the real large graph, ensuring GNNs trained on real and synthetic graphs exhibit comparable performance. However, existing methods rely on GNN-related information as supervision, including gradients, representations, and trajectories, which have two limitations. First, GNNs can affect the spectrum (i.e., eigenvalues) of the real graph, causing spectrum bias in the synthetic graph. Second, the variety of GNN architectures leads to the creation of different synthetic graphs, requiring traversal to obtain optimal performance. To tackle these issues, we propose Graph Distillation with Eigenbasis Matching (GDEM), which aligns the eigenbasis and node features of real and synthetic graphs. Meanwhile, it directly replicates the spectrum of the real graph and thus prevents the influence of GNNs. Moreover, we design a discrimination constraint to balance the effectiveness and generalization of GDEM. Theoretically, the synthetic graphs distilled by GDEM are restricted spectral approximations of the real graphs. Extensive experiments demonstrate that GDEM outperforms state-of-the-art GD methods with powerful cross-architecture generalization ability and significant distillation efficiency. Our code is available at https://github.com/liuyang-tian/GDEM.
Yang Liu 0348, Deyu Bo, Chuan Shi 0001
ICML2
2023 Specformer: Spectral Graph Neural Networks Meet Transformers
Deyu Bo, Chuan Shi 0001, Lele Wang 0001, Renjie Liao 0001
ICLR1
2023 Graph Contrastive Learning with Stable and Scalable Spectral Encoding
abstract
Graph contrastive learning (GCL) aims to learn representations by capturing the agreements between different graph views. Traditional GCL methods generate views in the spatial domain, but it has been recently discovered that the spectral domain also plays a vital role in complementing spatial views. However, existing spectral-based graph views either ignore the eigenvectors that encode valuable positional information or suffer from high complexity when trying to address the instability of spectral features. To tackle these challenges, we first design an informative, stable, and scalable spectral encoder, termed EigenMLP, to learn effective representations from the spectral features. Theoretically, EigenMLP is invariant to the rotation and reflection transformations on eigenvectors and robust against perturbations. Then, we propose a spatial-spectral contrastive framework (Sp$^{2}$GCL) to capture the consistency between the spatial information encoded by graph neural networks and the spectral information learned by EigenMLP, thus effectively fusing these two graph views. Experiments on the node- and graph-level datasets show that our method not only learns effective graph representations but also achieves a 2--10x speedup over other spectral-based methods.
Deyu Bo, Yuan Fang 0001, Yang Liu 0348, Chuan Shi 0001
NeurIPS1
2023 A Survey on Heterogeneous Graph Embedding: Methods, Techniques, Applications and Sources
abstract
Heterogeneous graphs (HGs) also known as heterogeneous information networks have become ubiquitous in real-world scenarios; therefore, HG embedding, which aims to learn representations in a lower-dimension space while preserving the heterogeneous structures and semantics for downstream tasks (e.g., node/graph classification, node clustering, link prediction), has drawn considerable attentions in recent years. In this survey, we perform a comprehensive review of the recent development on HG embedding methods and techniques. We first introduce the basic concepts of HG and discuss the unique challenges brought by the heterogeneity for HG embedding in comparison with homogeneous graph representation learning; and then we systemically survey and categorize the state-of-the-art HG embedding methods based on the information they used in the learning process to address the challenges posed by the HG heterogeneity. In particular, for each representative HG embedding method, we provide detailed introduction and further analyze its pros and cons; meanwhile, we also explore the transformativeness and applicability of different types of HG embedding methods in the real-world industrial environments for the first time. In addition, we further present several widely deployed systems that have demonstrated the success of HG embedding techniques in resolving real-world application problems with broader impacts. To facilitate future research and applications in this area, we also summarize the open-source code, existing graph learning platforms and benchmark datasets. Finally, we explore the additional issues and challenges of HG embedding and forecast the future research directions in this field.
Xiao Wang 0017, Deyu Bo, Chuan Shi 0001, Shaohua Fan, Yanfang Ye 0001, Philip S. Yu
IEEE Trans. Big Data2
2022 Regularizing Graph Neural Networks via Consistency-Diversity Graph Augmentations
abstract
Despite the remarkable performance of graph neural networks (GNNs) in semi-supervised learning, it is criticized for not making full use of unlabeled data and suffering from over-fitting. Recently, graph data augmentation, used to improve both accuracy and generalization of GNNs, has received considerable attentions. However, one fundamental question is how to evaluate the quality of graph augmentations in principle? In this paper, we propose two metrics, Consistency and Diversity, from the aspects of augmentation correctness and generalization. Moreover, we discover that existing augmentations fall into a dilemma between these two metrics. Can we find a graph augmentation satisfying both consistency and diversity? A well-informed answer can help us understand the mechanism behind graph augmentation and improve the performance of GNNs. To tackle this challenge, we analyze two representative semi-supervised learning algorithms: label propagation (LP) and consistency regularization (CR). We find that LP utilizes the prior knowledge of graphs to improve consistency and CR adopts variable augmentations to promote diversity. Based on this discovery, we treat neighbors as augmentations to capture the prior knowledge embodying homophily assumption, which promises a high consistency of augmentations. To further promote diversity, we randomly replace the immediate neighbors of each node with its remote neighbors. After that, a neighbor-constrained regularization is proposed to enforce the predictions of the augmented neighbors to be consistent with each other. Extensive experiments on five real-world graphs validate the superiority of our method in improving the accuracy and generalization of GNNs.
Deyu Bo, Binbin Hu, Xiao Wang 0017, Zhiqiang Zhang 0012, Chuan Shi 0001, Jun Zhou 0011
AAAI1
2022 Revisiting Graph Contrastive Learning from the Perspective of Graph Spectrum
abstract
Graph Contrastive Learning (GCL), learning the node representations by augmenting graphs, has attracted considerable attentions. Despite the proliferation of various graph augmentation strategies, there are still some fundamental questions unclear: what information is essentially learned by GCL? Are there some general augmentation rules behind different augmentations? If so, what are they and what insights can they bring? In this paper, we answer these questions by establishing the connection between GCL and graph spectrum. By an experimental investigation in spectral domain, we firstly find the General grAph augMEntation (GAME) rule for GCL, i.e., the difference of the high-frequency parts between two augmented graphs should be larger than that of low-frequency parts. This rule reveals the fundamental principle to revisit the current graph augmentations and design new effective graph augmentations. Then we theoretically prove that GCL is able to learn the invariance information by contrastive invariance theorem, together with our GAME rule, for the first time, we uncover that the learned representations by GCL essentially encode the low-frequency information, which explains why GCL works. Guided by this rule, we propose a spectral graph contrastive learning module (SpCo), which is a general and GCL-friendly plug-in. We combine it with different existing GCL models, and extensive experiments well demonstrate that it can further improve the performances of a wide variety of different GCL methods.
Nian Liu 0001, Xiao Wang 0017, Deyu Bo, Chuan Shi 0001, Jian Pei 0001
NeurIPS3
2021 Beyond Low-frequency Information in Graph Convolutional Networks
abstract
Graph neural networks (GNNs) have been proven to be effective in various network-related tasks. Most existing GNNs usually exploit the low-frequency signals of node features, which gives rise to one fundamental question: is the low-frequency information all we need in the real world applications? In this paper, we first present an experimental investigation assessing the roles of low-frequency and high-frequency signals, where the results clearly show that exploring low-frequency signal only is distant from learning an effective node representation in different scenarios. How can we adaptively learn more information beyond low-frequency information in GNNs? A well-informed answer can help GNNs enhance the adaptability. We tackle this challenge and propose a novel Frequency Adaptation Graph Convolutional Networks (FAGCN) with a self-gating mechanism, which can adaptively integrate different signals in the process of message passing. For a deeper understanding, we theoretically analyze the roles of low-frequency signals and high-frequency signals on learning node representations, which further explains why FAGCN can perform well on different types of networks. Extensive experiments on six real-world networks validate that FAGCN not only alleviates the over-smoothing problem, but also has advantages over the state-of-the-arts.
Deyu Bo, Xiao Wang 0017, Chuan Shi 0001, Huawei Shen
AAAI1
2020 AM-GCN: Adaptive Multi-channel Graph Convolutional Networks
abstract
Graph Convolutional Networks (GCNs) have gained great popularity in tackling various analytics tasks on graph and network data. However, some recent studies raise concerns about whether GCNs can optimally integrate node features and topological structures in a complex graph with rich information. In this paper, we first present an experimental investigation. Surprisingly, our experimental results clearly show that the capability of the state-of-the-art GCNs in fusing node features and topological structures is distant from optimal or even satisfactory. The weakness may severely hinder the capability of GCNs in some classification tasks, since GCNs may not be able to adaptively learn some deep correlation information between topological structures and node features. Can we remedy the weakness and design a new type of GCNs that can retain the advantages of the state-of-the-art GCNs and, at the same time, enhance the capability of fusing topological structures and node features substantially? We tackle the challenge and propose an adaptive multi-channel graph convolutional networks for semi-supervised classification (AM-GCN). The central idea is that we extract the specific and common embeddings from node features, topological structures, and their combinations simultaneously, and use the attention mechanism to learn adaptive importance weights of the embeddings. Our extensive experiments on benchmark data sets clearly show that AM-GCN extracts the most correlated information from both node features and topological structures substantially, and improves the classification accuracy with a clear margin.
Xiao Wang 0017, Deyu Bo, Peng Cui 0001, Chuan Shi 0001, Jian Pei 0001
KDD3
2020 Structural Deep Clustering Network
abstract
Clustering is a fundamental task in data analysis. Recently, deep clustering, which derives inspiration primarily from deep learning approaches, achieves state-of-the-art performance and has attracted considerable attention. Current deep clustering methods usually boost the clustering results by means of the powerful representation ability of deep learning, e.g., autoencoder, suggesting that learning an effective representation for clustering is a crucial requirement. The strength of deep clustering methods is to extract the useful representations from the data itself, rather than the structure of data, which receives scarce attention in representation learning. Motivated by the great success of Graph Convolutional Network (GCN) in encoding the graph structure, we propose a Structural Deep Clustering Network (SDCN) to integrate the structural information into deep clustering. Specifically, we design a delivery operator to transfer the representations learned by autoencoder to the corresponding GCN layer, and a dual self-supervised mechanism to unify these two different deep neural architectures and guide the update of the whole model. In this way, the multiple structures of data, from low-order to high-order, are naturally combined with the multiple representations learned by autoencoder. Furthermore, we theoretically analyze the delivery operator, i.e., with the delivery operator, GCN improves the autoencoder-specific representation as a high-order graph regularization constraint and autoencoder helps alleviate the over-smoothing problem in GCN. Through comprehensive experiments, we demonstrate that our propose model can consistently perform better over the state-of-the-art techniques.
Deyu Bo, Xiao Wang 0017, Chuan Shi 0001, Emiao Lu, Peng Cui 0001
WWW1