VLDB 2026 Research / reviewers in the wild / expert
Henan Sun
dblp:302/2010
· DBLP profile ↗
4ranked-venue papers
2as first author
4since 2021 · last 2026
0009-0002-4315-8900ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Toward Data-Centric Machine Learning on Directed Graphs: A SurveyabstractIn recent years, Graph Neural Networks (GNNs) have made significant advances in processing structured data. However, most of them primarily adopted a model-centric approach, which simplifies graphs by converting them into undirected formats and emphasizes model designs. This approach is inherently limited in real-world applications due to the unavoidable information loss in simple undirected graphs and the model optimization challenges that arise when exceeding the upper bounds of this sub-optimal data representational capacity. As a result, there has been a shift toward data-centric methods that prioritize improving graph quality and representation. Specifically, various types of graphs can be derived from naturally structured data, including heterogeneous graphs, hypergraphs, and directed graphs. Among these, directed graphs offer distinct advantages in topological systems by modeling causal relationships, and directed GNNs have been extensively studied in recent years. However, a comprehensive survey of this emerging topic is still lacking. Therefore, we aim to provide a comprehensive review of directed graph learning, with a particular focus on a data-centric perspective. Specifically, we first introduce a novel taxonomy for existing studies. Subsequently, we re-examine these methods from the data-centric perspective, with an emphasis on understanding and improving data representation. It demonstrates that a deep understanding of directed graphs and their quality plays a crucial role in model performance. Additionally, we explore the diverse applications of directed GNNs across 10+ domains, highlighting their broad applicability. Finally, we identify key opportunities and challenges within the field, offering insights that can guide future research and development in directed graph learning. Henan Sun, Xunkai Li, Daohan Su, Junyi Han, Rong-Hua Li 0001, Guoren Wang |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2024 | AdaFGL: A New Paradigm for Federated Node Classification with Topology HeterogeneityabstractRecently, Federated Graph Learning (FGL) has attracted significant attention as a distributed framework based on graph neural networks, primarily due to its capability to break data silos. Existing FGL studies employ community split on the homophilous global graph by default to simulate federated semisupervised node classification settings. Such a strategy assumes the consistency of topology between the multi-client subgraphs and the global graph, where connected nodes are highly likely to possess similar feature distributions and the same label. However, in real-world implementations, the varying perspectives of local data engineering result in various subgraph topologies, posing unique heterogeneity challenges in FGL. Unlike the well-known label Non-independent identical distribution (Non-iid) problems in federated learning, FGL heterogeneity essentially reveals the topological divergence among multiple clients, namely homophily or heterophily. To simulate and handle this unique challenge, we introduce the concept of structure Non-iid split and then present a new paradigm called Adaptive Federated Graph Learning (AdaFGL), a decoupled two-step personalized approach. To begin with, AdaFGL employs standard multi-client federated collaborative training to acquire the federated knowledge extractor by aggregating uploaded models in the final round at the server. Then, each client conducts personalized training based on the local subgraph and the federated knowledge extractor. Extensive experiments on the 12 graph benchmark datasets validate the superior performance of AdaFGL over state-of-the-art baselines. Specifically, in terms of test accuracy, our proposed AdaFGL outperforms baselines by significant margins of 3.24 % and 5.57 % on community split and structure Non-iid split, respectively. Xunkai Li, Zhengyu Wu, Wentao Zhang 0001, Henan Sun, Rong-Hua Li 0001, Guoren Wang |
ICDE | 4 |
| 2024 | Breaking the Entanglement of Homophily and Heterophily in Semi-supervised Node ClassificationabstractRecently, graph neural networks (GNNs) have shown prominent performance in semi-supervised node classification by leveraging knowledge from the graph database. However, most existing GNNs follow the homophily assumption, where connected nodes are more likely to exhibit similar feature distributions and the same labels, and such an assumption has proven to be vulnerable in a growing number of practical applications. As a supplement, heterophily reflects dissimilarity in connected nodes, which has gained significant attention in graph learning. To this end, data engineers aim to develop a powerful GNN model that can ensure performance under both homophily and heterophily. Despite numerous attempts, most existing GNNs struggle to achieve optimal node representations due to the constraints of undirected graphs. The neglect of directed edges results in sub-optimal graph representations, thereby hindering the capacity of GNNs. To address this issue, we introduce AMUD, which quantifies the relationship between node profiles and topology from a statistical perspective, offering valuable insights for Adaptively Modeling the natural directed graphs as the Undirected or Directed graph to maximize the benefits from subsequent graph learning. Furthermore, we propose Adaptive Directed Pattern Aggregation (ADPA) as a new directed graph learning paradigm for AMUD. Empirical studies have demonstrated that AMUD guides efficient graph learning. Meanwhile, extensive experiments on 16 benchmark datasets substantiate the impressive performance of ADPA, outperforming baselines by significant margins of 3.96%. Henan Sun, Xunkai Li, Zhengyu Wu, Daohan Su, Rong-Hua Li 0001, Guoren Wang |
ICDE | 1 |
| 2021 | CGAN-IRB: A Novel Data Augmentation Method for Apple Leaf DiseasesabstractAt present, the identification of apple leaf diseases plays an important role in controlling apple leaf diseases and improving apple yield. CNNs(Convolutional Neural Networks) have been widely used in apple leaf diseases identification, but the training of the CNNs requires a large number of images. The lack of images would make the CNNs hard to generalize. Thus the CNNs are unable to recognize new disease images. Focusing on this problem, this paper proposes a new model named CGAN-IRB(Conditional Generative Adversarial Network with the Improved Residual Block) for data augmentation. Firstly, various improvements have been made based on CGAN to generate high-quality, robust, and specific-category images of apple leaf diseases. Among which the embedding of the residual block has been found to significantly improve the model performance. Then the interpolation algorithm is used instead of deconvolution to increase the image size. Finally, the TTUR(Two-Timescale Update Rule) training strategy is employed and all the convolutional layers of the network are spectrally normalized to stabilize the training of the network. The performance of CGAN-IRB was tested both on image generation and classification tasks. Experiment results show that the images generated by the network possess high quality and robust features, pro-viding a novel solution for the data augmentation of apple leaf diseases. The new GAN-based data augmentation method leads to significant improvements in the classification accuracy of CNNs. In the case of all tested CNNs, the classification accuracy improvements are 11.75% and 2.17% on average over non-augmented and traditional-augmented, respectively. Among them, the classification accuracy of GoogLeNet V2 and ShuffleNet V2 is 99.34% and 99.67%, respectively. The data augmentation approach proposed in this paper can be used more widely in the field of disease identification, solving the problem of insufficient data sets, and can be extended to related fields where data sets are difficult to obtain. Xinbin Yuan, Cong Yu 0016, Bin Liu 0023, Henan Sun, Xianyu Zhu |
COMPSAC | 4 |