VLDB 2026 Research / reviewers in the wild / expert
Deyu Bo
dblp:258/0824
· DBLP profile ↗
12ranked-venue papers
6as first author
10since 2021 · last 2025
0000-0003-2063-8223ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 5 first-author · 8 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
Graph learning · 70% Representation and self-supervised learning · 16% Efficient and distributed learning · 12% | |
| Databases, data mining, and information retrieval
1 paper |
Data mining · 100% |
Topics — the 21 heaviest of 23, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Graph learning
graph neural network |
2.2 | 4 | 2023 | Specformer: Spectral Graph Neural Networks Meet Transformers · ICLR 2023 Regularizing Graph Neural Networks via Consistency-Diversity Graph Augmentations · AAAI 2022 Beyond Low-frequency Information in Graph Convolutional Networks · AAAI 2021 |
Machine learning › Efficient and distributed learning
dataset distillation |
1.6 | 2 | 2025 | Point Cloud Dataset Distillation · ICML 2025 Graph Distillation with Eigenbasis Matching · ICML 2024 |
Machine learning › Graph learning › graph neural network
graph convolutional network |
1.4 | 3 | 2021 | Beyond Low-frequency Information in Graph Convolutional Networks · AAAI 2021 Structural Deep Clustering Network · WWW 2020 AM-GCN: Adaptive Multi-channel Graph Convolutional Networks · KDD 2020 |
Machine learning › Representation and self-supervised learning › contrastive learning
graph contrastive learning |
1.2 | 2 | 2023 | Graph Contrastive Learning with Stable and Scalable Spectral Encoding · NeurIPS 2023 Revisiting Graph Contrastive Learning from the Perspective of Graph Spectrum · NeurIPS 2022 |
Machine learning › Graph learning › graph neural network › node classification
semi-supervised node classification |
1.0 | 2 | 2022 | Regularizing Graph Neural Networks via Consistency-Diversity Graph Augmentations · AAAI 2022 AM-GCN: Adaptive Multi-channel Graph Convolutional Networks · KDD 2020 |
Machine learning › Graph learning › graph neural network training
graph distillation |
0.8 | 1 | 2024 | Graph Distillation with Eigenbasis Matching · ICML 2024 |
Machine learning › Graph learning
graph representation learning |
0.7 | 1 | 2023 | Graph Contrastive Learning with Stable and Scalable Spectral Encoding · NeurIPS 2023 |
Machine learning › Graph learning › graph neural network
graph transformer |
0.7 | 1 | 2023 | Specformer: Spectral Graph Neural Networks Meet Transformers · ICLR 2023 |
Machine learning › Graph learning › graph neural network
spectral graph neural network |
0.7 | 1 | 2023 | Specformer: Spectral Graph Neural Networks Meet Transformers · ICLR 2023 |
Machine learning › Representation and self-supervised learning
contrastive learning |
0.6 | 1 | 2022 | Revisiting Graph Contrastive Learning from the Perspective of Graph Spectrum · NeurIPS 2022 |
Machine learning › Graph learning
graph augmentation |
0.6 | 1 | 2022 | Revisiting Graph Contrastive Learning from the Perspective of Graph Spectrum · NeurIPS 2022 |
Machine learning › Graph learning › graph neural network
graph data augmentation |
0.6 | 1 | 2022 | Regularizing Graph Neural Networks via Consistency-Diversity Graph Augmentations · AAAI 2022 |
Machine learning › Graph learning
graph signal processing |
0.5 | 1 | 2021 | Beyond Low-frequency Information in Graph Convolutional Networks · AAAI 2021 |
Machine learning › Graph learning › graph neural network
node classification |
0.4 | 1 | 2020 | AM-GCN: Adaptive Multi-channel Graph Convolutional Networks · KDD 2020 |
Machine learning › Representation and self-supervised learning › representation learning
structured representation learning |
0.4 | 1 | 2020 | Structural Deep Clustering Network · WWW 2020 |
Data mining
clustering |
0.4 | 1 | 2020 | Structural Deep Clustering Network · WWW 2020 |
Data mining › clustering
deep clustering |
0.4 | 1 | 2020 | Structural Deep Clustering Network · WWW 2020 |
Computer vision › 3D vision
point cloud processing |
0.3 | 1 | 2025 | Point Cloud Dataset Distillation · ICML 2025 |
Machine learning › Learning paradigms › semi-supervised learning
consistency regularization |
0.2 | 1 | 2022 | Regularizing Graph Neural Networks via Consistency-Diversity Graph Augmentations · AAAI 2022 |
Machine learning › Graph learning › graph representation learning
node representation learning |
0.1 | 1 | 2021 | Beyond Low-frequency Information in Graph Convolutional Networks · AAAI 2021 |
Machine learning › Graph learning › graph neural network › deep graph neural network
over-smoothing |
0.1 | 1 | 2021 | Beyond Low-frequency Information in Graph Convolutional Networks · AAAI 2021 |
Methods — techniques the papers use, named apart from their topics
rotation-invariant feature matching · 0.9point-wise generator · 0.9spectral approximation · 0.8eigenbasis matching · 0.8transformer · 0.7spectral graph theory · 0.7graph neural network · 0.7contrastive learning · 0.7label propagation · 0.6consistency regularization · 0.6self-supervised learning · 0.4graph convolutional network · 0.4autoencoder · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Point Cloud Dataset DistillationabstractThis study introduces dataset distillation (DD) tailored for 3D data, particularly point clouds. DD aims to substitute large-scale real datasets with a small set of synthetic samples while preserving model performance. Existing methods mainly focus on structured data such as images. However, adapting DD for unstructured point clouds poses challenges due to their diverse orientations and resolutions in 3D space. To address these challenges, we theoretically demonstrate the importance of matching rotation-invariant features between real and synthetic data for 3D distillation. We further propose a plug-and-play point cloud rotator to align the point cloud to a canonical orientation, facilitating the learning of rotation-invariant features by all point cloud models. Furthermore, instead of optimizing fixed-size synthetic data directly, we devise a point-wise generator to produce point clouds at various resolutions based on the sampled noise amount. Compared to conventional DD methods, the proposed approach, termed DD3D, enables efficient training on low-resolution point clouds while generating high-resolution data for evaluation, thereby significantly reducing memory requirements and enhancing model scalability. Extensive experiments validate the effectiveness of DD3D in shape classification and part segmentation tasks across diverse scenarios, such as cross-architecture and cross-resolution settings. Deyu Bo, Xinchao Wang |
ICML | 1 |
| 2025 | Graph Positional Autoencoders as Self-supervised LearnersabstractGraph self-supervised learning seeks to learn effective graph representations without relying on labeled data. Among various approaches, graph autoencoders (GAEs) have gained significant attention for their efficiency and scalability. Typically, GAEs take incomplete graphs as input and predict missing elements, such as masked node features or edges. Although effective, our experimental investigation reveals that traditional feature or edge masking paradigms primarily capture low-frequency signals in the graph and fail to learn expressive structural information. To address these issues, we propose Graph Positional Autoencoders (GraphPAE), which employ a dual-path architecture to reconstruct both node features and positions. Specifically, the feature path uses positional encoding to enhance the message-passing processing, improving the GAEs' ability to predict the corrupted information. The position path, on the other hand, leverages node representations to refine positions and approximate eigenvectors, thereby enabling the encoder to learn diverse frequency information. We conduct extensive experiments to verify the effectiveness of GraphPAE, including heterophilic node classification, graph property prediction, and transfer learning. The results demonstrate that GraphPAE achieves state-of-the-art performance and consistently outperforms the baselines by a large margin. Yang Liu 0348, Deyu Bo, Wenxuan Cao, Yuan Fang 0001, Yawen Li 0001, Chuan Shi 0001 |
KDD (2) | 2 |
| 2025 | Data-Centric Graph Learning: A SurveyabstractThe history of artificial intelligence (AI) has witnessed the significant impact of high-quality data on various deep learning models, such as ImageNet for AlexNet and ResNet. Recently, instead of designing more complex neural architectures as model-centric approaches, the attention of AI community has shifted to data-centric ones, which focuses on better processing data to strengthen the ability of neural models. Graph learning, which operates on ubiquitous topological data, also plays an important role in the era of deep learning. In this survey, we comprehensively review graph learning approaches from the data-centric perspective, and aim to answer three crucial questions:(1) when to modify graph data,(2) what part of the graph data needs modificationto unlock the potential of various graph models, and(3) how to safeguard graph modelsfrom problematic data influence. Accordingly, we propose a novel taxonomy based on the stages in the graph learning pipeline, and highlight the processing methods for different data structures in the graph data, i.e., topology, feature and label. Furthermore, we analyze some potential problems embedded in graph data and discuss how to solve them in a data-centric manner. Finally, we provide some promising future directions for data-centric graph learning. Deyu Bo, Cheng Yang 0002, Zhongjian Zhang, Jixi Liu, Yufei Peng, Chuan Shi 0001 |
IEEE Trans. Big Data | 2 |
| 2024 | Graph Distillation with Eigenbasis MatchingabstractThe increasing amount of graph data places requirements on the efficient training of graph neural networks (GNNs). The emerging graph distillation (GD) tackles this challenge by distilling a small synthetic graph to replace the real large graph, ensuring GNNs trained on real and synthetic graphs exhibit comparable performance. However, existing methods rely on GNN-related information as supervision, including gradients, representations, and trajectories, which have two limitations. First, GNNs can affect the spectrum (i.e., eigenvalues) of the real graph, causing spectrum bias in the synthetic graph. Second, the variety of GNN architectures leads to the creation of different synthetic graphs, requiring traversal to obtain optimal performance. To tackle these issues, we propose Graph Distillation with Eigenbasis Matching (GDEM), which aligns the eigenbasis and node features of real and synthetic graphs. Meanwhile, it directly replicates the spectrum of the real graph and thus prevents the influence of GNNs. Moreover, we design a discrimination constraint to balance the effectiveness and generalization of GDEM. Theoretically, the synthetic graphs distilled by GDEM are restricted spectral approximations of the real graphs. Extensive experiments demonstrate that GDEM outperforms state-of-the-art GD methods with powerful cross-architecture generalization ability and significant distillation efficiency. Our code is available at https://github.com/liuyang-tian/GDEM. Yang Liu 0348, Deyu Bo, Chuan Shi 0001 |
ICML | 2 |
| 2023 | Specformer: Spectral Graph Neural Networks Meet Transformers
Deyu Bo, Chuan Shi 0001, Lele Wang 0001, Renjie Liao 0001 |
ICLR | 1 |
| 2023 | Graph Contrastive Learning with Stable and Scalable Spectral EncodingabstractGraph contrastive learning (GCL) aims to learn representations by capturing the agreements between different graph views. Traditional GCL methods generate views in the spatial domain, but it has been recently discovered that the spectral domain also plays a vital role in complementing spatial views. However, existing spectral-based graph views either ignore the eigenvectors that encode valuable positional information or suffer from high complexity when trying to address the instability of spectral features. To tackle these challenges, we first design an informative, stable, and scalable spectral encoder, termed EigenMLP, to learn effective representations from the spectral features. Theoretically, EigenMLP is invariant to the rotation and reflection transformations on eigenvectors and robust against perturbations. Then, we propose a spatial-spectral contrastive framework (Sp$^{2}$GCL) to capture the consistency between the spatial information encoded by graph neural networks and the spectral information learned by EigenMLP, thus effectively fusing these two graph views. Experiments on the node- and graph-level datasets show that our method not only learns effective graph representations but also achieves a 2--10x speedup over other spectral-based methods. Deyu Bo, Yuan Fang 0001, Yang Liu 0348, Chuan Shi 0001 |
NeurIPS | 1 |
| 2023 | A Survey on Heterogeneous Graph Embedding: Methods, Techniques, Applications and SourcesabstractHeterogeneous graphs (HGs) also known as heterogeneous information networks have become ubiquitous in real-world scenarios; therefore, HG embedding, which aims to learn representations in a lower-dimension space while preserving the heterogeneous structures and semantics for downstream tasks (e.g., node/graph classification, node clustering, link prediction), has drawn considerable attentions in recent years. In this survey, we perform a comprehensive review of the recent development on HG embedding methods and techniques. We first introduce the basic concepts of HG and discuss the unique challenges brought by the heterogeneity for HG embedding in comparison with homogeneous graph representation learning; and then we systemically survey and categorize the state-of-the-art HG embedding methods based on the information they used in the learning process to address the challenges posed by the HG heterogeneity. In particular, for each representative HG embedding method, we provide detailed introduction and further analyze its pros and cons; meanwhile, we also explore the transformativeness and applicability of different types of HG embedding methods in the real-world industrial environments for the first time. In addition, we further present several widely deployed systems that have demonstrated the success of HG embedding techniques in resolving real-world application problems with broader impacts. To facilitate future research and applications in this area, we also summarize the open-source code, existing graph learning platforms and benchmark datasets. Finally, we explore the additional issues and challenges of HG embedding and forecast the future research directions in this field. Xiao Wang 0017, Deyu Bo, Chuan Shi 0001, Shaohua Fan, Yanfang Ye 0001, Philip S. Yu |
IEEE Trans. Big Data | 2 |
| 2022 | Regularizing Graph Neural Networks via Consistency-Diversity Graph AugmentationsabstractDespite the remarkable performance of graph neural networks (GNNs) in semi-supervised learning, it is criticized for not making full use of unlabeled data and suffering from over-fitting. Recently, graph data augmentation, used to improve both accuracy and generalization of GNNs, has received considerable attentions. However, one fundamental question is how to evaluate the quality of graph augmentations in principle? In this paper, we propose two metrics, Consistency and Diversity, from the aspects of augmentation correctness and generalization. Moreover, we discover that existing augmentations fall into a dilemma between these two metrics. Can we find a graph augmentation satisfying both consistency and diversity? A well-informed answer can help us understand the mechanism behind graph augmentation and improve the performance of GNNs. To tackle this challenge, we analyze two representative semi-supervised learning algorithms: label propagation (LP) and consistency regularization (CR). We find that LP utilizes the prior knowledge of graphs to improve consistency and CR adopts variable augmentations to promote diversity. Based on this discovery, we treat neighbors as augmentations to capture the prior knowledge embodying homophily assumption, which promises a high consistency of augmentations. To further promote diversity, we randomly replace the immediate neighbors of each node with its remote neighbors. After that, a neighbor-constrained regularization is proposed to enforce the predictions of the augmented neighbors to be consistent with each other. Extensive experiments on five real-world graphs validate the superiority of our method in improving the accuracy and generalization of GNNs. Deyu Bo, Binbin Hu, Xiao Wang 0017, Zhiqiang Zhang 0012, Chuan Shi 0001, Jun Zhou 0011 |
AAAI | 1 |
| 2022 | Revisiting Graph Contrastive Learning from the Perspective of Graph SpectrumabstractGraph Contrastive Learning (GCL), learning the node representations by augmenting graphs, has attracted considerable attentions. Despite the proliferation of various graph augmentation strategies, there are still some fundamental questions unclear: what information is essentially learned by GCL? Are there some general augmentation rules behind different augmentations? If so, what are they and what insights can they bring? In this paper, we answer these questions by establishing the connection between GCL and graph spectrum. By an experimental investigation in spectral domain, we firstly find the General grAph augMEntation (GAME) rule for GCL, i.e., the difference of the high-frequency parts between two augmented graphs should be larger than that of low-frequency parts. This rule reveals the fundamental principle to revisit the current graph augmentations and design new effective graph augmentations. Then we theoretically prove that GCL is able to learn the invariance information by contrastive invariance theorem, together with our GAME rule, for the first time, we uncover that the learned representations by GCL essentially encode the low-frequency information, which explains why GCL works. Guided by this rule, we propose a spectral graph contrastive learning module (SpCo), which is a general and GCL-friendly plug-in. We combine it with different existing GCL models, and extensive experiments well demonstrate that it can further improve the performances of a wide variety of different GCL methods. Nian Liu 0001, Xiao Wang 0017, Deyu Bo, Chuan Shi 0001, Jian Pei 0001 |
NeurIPS | 3 |
| 2021 | Beyond Low-frequency Information in Graph Convolutional NetworksabstractGraph neural networks (GNNs) have been proven to be effective in various network-related tasks. Most existing GNNs usually exploit the low-frequency signals of node features, which gives rise to one fundamental question: is the low-frequency information all we need in the real world applications? In this paper, we first present an experimental investigation assessing the roles of low-frequency and high-frequency signals, where the results clearly show that exploring low-frequency signal only is distant from learning an effective node representation in different scenarios. How can we adaptively learn more information beyond low-frequency information in GNNs? A well-informed answer can help GNNs enhance the adaptability. We tackle this challenge and propose a novel Frequency Adaptation Graph Convolutional Networks (FAGCN) with a self-gating mechanism, which can adaptively integrate different signals in the process of message passing. For a deeper understanding, we theoretically analyze the roles of low-frequency signals and high-frequency signals on learning node representations, which further explains why FAGCN can perform well on different types of networks. Extensive experiments on six real-world networks validate that FAGCN not only alleviates the over-smoothing problem, but also has advantages over the state-of-the-arts. Deyu Bo, Xiao Wang 0017, Chuan Shi 0001, Huawei Shen |
AAAI | 1 |
| 2020 | AM-GCN: Adaptive Multi-channel Graph Convolutional NetworksabstractGraph Convolutional Networks (GCNs) have gained great popularity in tackling various analytics tasks on graph and network data. However, some recent studies raise concerns about whether GCNs can optimally integrate node features and topological structures in a complex graph with rich information. In this paper, we first present an experimental investigation. Surprisingly, our experimental results clearly show that the capability of the state-of-the-art GCNs in fusing node features and topological structures is distant from optimal or even satisfactory. The weakness may severely hinder the capability of GCNs in some classification tasks, since GCNs may not be able to adaptively learn some deep correlation information between topological structures and node features. Can we remedy the weakness and design a new type of GCNs that can retain the advantages of the state-of-the-art GCNs and, at the same time, enhance the capability of fusing topological structures and node features substantially? We tackle the challenge and propose an adaptive multi-channel graph convolutional networks for semi-supervised classification (AM-GCN). The central idea is that we extract the specific and common embeddings from node features, topological structures, and their combinations simultaneously, and use the attention mechanism to learn adaptive importance weights of the embeddings. Our extensive experiments on benchmark data sets clearly show that AM-GCN extracts the most correlated information from both node features and topological structures substantially, and improves the classification accuracy with a clear margin. Xiao Wang 0017, Deyu Bo, Peng Cui 0001, Chuan Shi 0001, Jian Pei 0001 |
KDD | 3 |
| 2020 | Structural Deep Clustering NetworkabstractClustering is a fundamental task in data analysis. Recently, deep clustering, which derives inspiration primarily from deep learning approaches, achieves state-of-the-art performance and has attracted considerable attention. Current deep clustering methods usually boost the clustering results by means of the powerful representation ability of deep learning, e.g., autoencoder, suggesting that learning an effective representation for clustering is a crucial requirement. The strength of deep clustering methods is to extract the useful representations from the data itself, rather than the structure of data, which receives scarce attention in representation learning. Motivated by the great success of Graph Convolutional Network (GCN) in encoding the graph structure, we propose a Structural Deep Clustering Network (SDCN) to integrate the structural information into deep clustering. Specifically, we design a delivery operator to transfer the representations learned by autoencoder to the corresponding GCN layer, and a dual self-supervised mechanism to unify these two different deep neural architectures and guide the update of the whole model. In this way, the multiple structures of data, from low-order to high-order, are naturally combined with the multiple representations learned by autoencoder. Furthermore, we theoretically analyze the delivery operator, i.e., with the delivery operator, GCN improves the autoencoder-specific representation as a high-order graph regularization constraint and autoencoder helps alleviate the over-smoothing problem in GCN. Through comprehensive experiments, we demonstrate that our propose model can consistently perform better over the state-of-the-art techniques. Deyu Bo, Xiao Wang 0017, Chuan Shi 0001, Emiao Lu, Peng Cui 0001 |
WWW | 1 |