VLDB 2026 Research / reviewers in the wild / expert
Han Yang 0002
dblp:42/1222-2
· DBLP profile ↗
9ranked-venue papers
2as first author
8since 2021 · last 2024
0000-0003-2782-7502ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Graph learning · 38% Trustworthy machine learning · 33% Optimization for machine learning · 13% | |
| Computer graphics and multimedia
1 paper |
Geometric modeling and processing · 100% | |
| Databases, data mining, and information retrieval
1 paper |
Data integration and cleaning · 50% Information retrieval · 50% | |
| Network and information security
1 paper |
Security and privacy of machine learning · 100% |
Topics — the 21 heaviest of 23, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
out-of-distribution generalization |
1.2 | 2 | 2023 | Pareto Invariant Risk Minimization: Towards Mitigating the Optimization Dilemma in Out-of-Distribution Generalization · ICLR 2023 Learning Causally Invariant Representations for Out-of-Distribution Generalization on Graphs · NeurIPS 2022 |
Machine learning › Graph learning
graph neural network |
1.1 | 2 | 2022 | Understanding and Improving Graph Injection Attack by Promoting Unnoticeability · ICLR 2022 Rethinking Graph Regularization for Graph Neural Networks · AAAI 2021 |
Machine learning › Trustworthy machine learning
robustness |
0.8 | 2 | 2023 | Pareto Invariant Risk Minimization: Towards Mitigating the Optimization Dilemma in Out-of-Distribution Generalization · ICLR 2023 Learning Causally Invariant Representations for Out-of-Distribution Generalization on Graphs · NeurIPS 2022 |
Machine learning › Optimization for machine learning
multi-objective optimization |
0.7 | 1 | 2023 | Pareto Invariant Risk Minimization: Towards Mitigating the Optimization Dilemma in Out-of-Distribution Generalization · ICLR 2023 |
Machine learning › Optimization for machine learning › multi-objective optimization
pareto optimization |
0.7 | 1 | 2023 | Pareto Invariant Risk Minimization: Towards Mitigating the Optimization Dilemma in Out-of-Distribution Generalization · ICLR 2023 |
Machine learning › Trustworthy machine learning › invariance
causal invariance |
0.6 | 1 | 2022 | Learning Causally Invariant Representations for Out-of-Distribution Generalization on Graphs · NeurIPS 2022 |
Machine learning › Efficient and distributed learning
distributed training |
0.6 | 1 | 2022 | HGL: Accelerating Heterogeneous GNN Training with Holistic Representation and Optimization · SC 2022 |
Machine learning › Trustworthy machine learning › graph adversarial attack
graph injection attack |
0.6 | 1 | 2022 | Understanding and Improving Graph Injection Attack by Promoting Unnoticeability · ICLR 2022 |
Machine learning › Graph learning
graph neural network training |
0.6 | 1 | 2022 | HGL: Accelerating Heterogeneous GNN Training with Holistic Representation and Optimization · SC 2022 |
Machine learning › Representation and self-supervised learning › representation learning › neural network representation learning
intermediate representation |
0.6 | 1 | 2022 | HGL: Accelerating Heterogeneous GNN Training with Holistic Representation and Optimization · SC 2022 |
Machine learning › Graph learning › graph structure learning
invariant subgraph learning |
0.6 | 1 | 2022 | Learning Causally Invariant Representations for Out-of-Distribution Generalization on Graphs · NeurIPS 2022 |
Geometric modeling and processing › shape correspondence
non-isometric shape matching |
0.6 | 1 | 2022 | Exact Shape Correspondence via 2D graph convolution · NeurIPS 2022 |
Geometric modeling and processing
shape correspondence |
0.6 | 1 | 2022 | Exact Shape Correspondence via 2D graph convolution · NeurIPS 2022 |
Security and privacy of machine learning › adversarial attack
graph adversarial attack |
0.6 | 1 | 2022 | Understanding and Improving Graph Injection Attack by Promoting Unnoticeability · ICLR 2022 |
Machine learning › Graph learning › graph regularization
graph laplacian regularization |
0.5 | 1 | 2021 | Rethinking Graph Regularization for Graph Neural Networks · AAAI 2021 |
Machine learning › Graph learning
graph regularization |
0.5 | 1 | 2021 | Rethinking Graph Regularization for Graph Neural Networks · AAAI 2021 |
Machine learning › Learning paradigms
semi-supervised learning |
0.5 | 1 | 2021 | Rethinking Graph Regularization for Graph Neural Networks · AAAI 2021 |
Information retrieval › similarity search
approximate similarity search |
0.4 | 1 | 2020 | Convolutional Embedding for Edit Distance · SIGIR 2020 |
Data integration and cleaning › approximate matching
string similarity search |
0.4 | 1 | 2020 | Convolutional Embedding for Edit Distance · SIGIR 2020 |
Machine learning › Trustworthy machine learning › robustness
distribution shift |
0.2 | 1 | 2022 | Learning Causally Invariant Representations for Out-of-Distribution Generalization on Graphs · NeurIPS 2022 |
Machine learning › Graph learning › graph neural network
graph convolutional network |
0.1 | 1 | 2021 | Rethinking Graph Regularization for Graph Neural Networks · AAAI 2021 |
Methods — techniques the papers use, named apart from their topics
invariant risk minimization · 0.7operator fusion · 0.6operator bundling · 0.6laplace-beltrami eigenfunctions · 0.6information-theoretic objective · 0.6graph stitching · 0.6causal models · 0.62d graph convolution · 0.6propagation regularization · 0.5graph laplacian regularization · 0.5triplet loss · 0.4convolutional neural network · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Solving the non-submodular network collapse problems via Decision Transformer
Kaili Ma 0001, Han Yang 0002, Shanchao Yang, Kangfei Zhao, Lanqing Li, Yongqiang Chen 0002, Junzhou Huang, James Cheng, Yu Rong 0001 |
Neural Networks | 2 |
| 2023 | Pareto Invariant Risk Minimization: Towards Mitigating the Optimization Dilemma in Out-of-Distribution Generalization
Yongqiang Chen 0002, Kaiwen Zhou 0001, Yatao Bian, Binghui Xie, Bingzhe Wu, Yonggang Zhang 0003, Kaili Ma 0001, Han Yang 0002, Peilin Zhao, Bo Han 0003, James Cheng |
ICLR | 8 |
| 2022 | Understanding and Improving Graph Injection Attack by Promoting Unnoticeability
Yongqiang Chen 0002, Han Yang 0002, Yonggang Zhang 0003, Kaili Ma 0001, Tongliang Liu, Bo Han 0003, James Cheng |
ICLR | 2 |
| 2022 | Learning Causally Invariant Representations for Out-of-Distribution Generalization on GraphsabstractDespite recent success in using the invariance principle for out-of-distribution (OOD) generalization on Euclidean data (e.g., images), studies on graph data are still limited. Different from images, the complex nature of graphs poses unique challenges to adopting the invariance principle. In particular, distribution shifts on graphs can appear in a variety of forms such as attributes and structures, making it difficult to identify the invariance. Moreover, domain or environment partitions, which are often required by OOD methods on Euclidean data, could be highly expensive to obtain for graphs. To bridge this gap, we propose a new framework, called Causality Inspired Invariant Graph LeArning (CIGA), to capture the invariance of graphs for guaranteed OOD generalization under various distribution shifts. Specifically, we characterize potential distribution shifts on graphs with causal models, concluding that OOD generalization on graphs is achievable when models focus only on subgraphs containing the most information about the causes of labels. Accordingly, we propose an information-theoretic objective to extract the desired subgraphs that maximally preserve the invariant intra-class information. Learning with these subgraphs is immune to distribution shifts. Extensive experiments on 16 synthetic or real-world datasets, including a challenging setting -- DrugOOD, from AI-aided drug discovery, validate the superior OOD performance of CIGA. Yongqiang Chen 0002, Yonggang Zhang 0003, Yatao Bian, Han Yang 0002, Kaili Ma 0001, Binghui Xie, Tongliang Liu, Bo Han 0003, James Cheng |
NeurIPS | 4 |
| 2022 | Exact Shape Correspondence via 2D graph convolutionabstractFor exact 3D shape correspondence (matching or alignment), i.e., the task of matching each point on a shape to its exact corresponding point on the other shape (or to be more specific, matching at geodesic error 0), most existing methods do not perform well due to two main problems. First, on nearly-isometric shapes (i.e., low noise levels), most existing methods use the eigen-vectors (eigen-functions) of the Laplace Beltrami Operator (LBO) or other shape descriptors to update an initialized correspondence which is not exact, leading to an accumulation of update errors. Thus, though the final correspondence may generally be smooth, it is generally inexact. Second, on non-isometric shapes (noisy shapes), existing methods are generally not robust to noise as they usually assume near-isometry. In addition, existing methods that attempt to address the non-isometric shape problem (e.g., GRAMPA) are generally computationally expensive and do not generalise to nearly-isometric shapes. To address these two problems, we propose a 2D graph convolution-based framework called 2D-GEM. 2D-GEM is robust to noise on non-isometric shapes and with a few additional constraints, it also addresses the errors in the update on nearly-isometric shapes. We demonstrate the effectiveness of 2D-GEM by achieving a high accuracy of 90.5$\%$ at geodesic error 0 on the non-isometric benchmark SHREC16, i.e., TOPKIDS (while being much faster than GRAMPA), and on nearly-isometric benchmarks by achieving a high accuracy of 92.5$\%$ on TOSCA and 84.9$\%$ on SCAPE at geodesic error 0. Barakeel Fanseu Kamhoua, Lin Zhang 0059, Yongqiang Chen 0002, Han Yang 0002, Kaili Ma 0001, Bo Han 0003, Bo Li 0001, James Cheng |
NeurIPS | 4 |
| 2022 | HGL: Accelerating Heterogeneous GNN Training with Holistic Representation and OptimizationabstractGraph neural networks (GNNs) have shown to significantly improve graph analytics. Existing systems for GNN training are primarily designed for homogeneous graphs. In industry, however, most graphs are actually heterogeneous in nature (i.e., having multiple types of nodes and edges). Existing systems train a heterogeneous GNN (HetGNN) as a composition of homogeneous GNN (HomoGNN) and thus suffer from critical limitations such as lack of memory optimization and limited operator parallelism. To address these limitations, we propose HGL - a heterogeneity-aware system for GNN training. At the core of HGL is an intermediate representation, called HIR, which provides a holistic representation for GNNs and enables cross-relation optimization in HetGNN training. We devise tailored optimizations on HIR, including graph stitching, operator fusion and operator bundling. Compared with DGL and PyG, HGL achieves a speedup from 7 to 22 times for training HetGNNs. Yuntao Gui, Yidi Wu 0001, Han Yang 0002, Tatiana Jin, Boyang Li 0016, Qihui Zhou, James Cheng, Fan Yu 0004 |
SC | 3 |
| 2021 | Rethinking Graph Regularization for Graph Neural NetworksabstractThe graph Laplacian regularization term is usually used in semi-supervised representation learning to provide graph structure information for a model f(X). However, with the recent popularity of graph neural networks (GNNs), directly encoding graph structure A into a model, i.e., f(A, X), has become the more common approach. While we show that graph Laplacian regularization brings little-to-no benefit to existing GNNs, and propose a simple but non-trivial variant of graph Laplacian regularization, called Propagation-regularization (P-reg), to boost the performance of existing GNN models. We provide formal analyses to show that P-reg not only infuses extra information (that is not captured by the traditional graph Laplacian regularization) into GNNs, but also has the capacity equivalent to an infinite-depth graph convolutional network. We demonstrate that P-reg can effectively boost the performance of existing GNN models on both node-level and graph-level tasks across many different datasets. Han Yang 0002, Kaili Ma 0001, James Cheng |
AAAI | 1 |
| 2021 | Self-Enhanced GNN: Improving Graph Neural Networks Using Model OutputsabstractGraph neural networks (GNNs) have received much attention recently because of their excellent performance on graph-based tasks. However, existing research on GNNs focuses on designing more effective models without considering much about the quality of the input data. In this paper, we propose self-enhanced GNN (SEG), which improves the quality of the input data using the outputs of existing GNN models for better performance on semi-supervised node classification. As graph data consist of both topology and node labels, we improve input data quality from both perspectives. For topology, we observe that higher classification accuracy can be achieved when the ratio of inter-class edges (connecting nodes from different classes) is low and propose topology update to remove inter-class edges and add intra-class edges. For node labels, we propose training node augmentation, which enlarges the training set using the labels predicted by existing GNN models. SEG is a general framework that can be easily combined with existing GNN models. Experimental results validate that SEG consistently improves the performance of well-known GNN models such as GCN, GAT and SGC across different datasets. Han Yang 0002, Xiao Yan 0002, Xinyan Dai, Yongqiang Chen 0002, James Cheng |
IJCNN | 1 |
| 2020 | Convolutional Embedding for Edit DistanceabstractEdit-distance-based string similarity search has many applications such as spell correction, data de-duplication, and sequence alignment. However, computing edit distance is known to have high complexity, which makes string similarity search challenging for large datasets. In this paper, we propose a deep learning pipeline (called CNN-ED) that embeds edit distance into Euclidean distance for fast approximate similarity search. A convolutional neural network (CNN) is used to generate fixed-length vector embeddings for a dataset of strings and the loss function is a combination of the triplet loss and the approximation error. To justify our choice of using CNN instead of other structures (e.g., RNN) as the model, theoretical analysis is conducted to show that some basic operations in our CNN model preserve edit distance. Experimental results show that CNN-ED outperforms data-independent CGK embedding and RNN-based GRU embedding in terms of both accuracy and efficiency by a large margin. We also show that string similarity search can be significantly accelerated using CNN-based embeddings, sometimes by orders of magnitude. Xinyan Dai, Xiao Yan 0002, Kaiwen Zhou 0001, Han Yang 0002, James Cheng |
SIGIR | 5 |