VLDB 2026 Research / reviewers in the wild / expert
Eric Inae
dblp:342/8313
· DBLP profile ↗
4ranked-venue papers
0as first author
4since 2021 · last 2025
0009-0002-2101-2126ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Graph learning · 58% Generative modeling · 32% Deep learning architectures and training · 5% | |
| Databases, data mining, and information retrieval
1 paper |
Data mining · 100% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Bioinformatics and computational biology · 100% |
Topics — the 9 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Graph learning
graph neural network |
0.9 | 1 | 2025 | Learning Repetition-Invariant Representations for Polymer Informatics · NeurIPS 2025 |
Machine learning › Graph learning › graph representation learning
invariant graph representation |
0.9 | 1 | 2025 | Learning Repetition-Invariant Representations for Polymer Informatics · NeurIPS 2025 |
Machine learning › Generative modeling
diffusion model |
0.7 | 1 | 2023 | Data-Centric Learning from Unlabeled Graphs with Diffusion Model · NeurIPS 2023 |
Machine learning › Generative modeling › diffusion model
graph diffusion model |
0.7 | 1 | 2023 | Data-Centric Learning from Unlabeled Graphs with Diffusion Model · NeurIPS 2023 |
Machine learning › Graph learning › graph inference › graph prediction
graph property prediction |
0.7 | 1 | 2023 | Data-Centric Learning from Unlabeled Graphs with Diffusion Model · NeurIPS 2023 |
Data mining › structured data mining
graph mining |
0.7 | 1 | 2023 | Semi-Supervised Graph Imbalanced Regression · KDD 2023 |
Data mining
semi-supervised learning |
0.7 | 1 | 2023 | Semi-Supervised Graph Imbalanced Regression · KDD 2023 |
Machine learning › Deep learning architectures and training
data augmentation |
0.2 | 1 | 2023 | Data-Centric Learning from Unlabeled Graphs with Diffusion Model · NeurIPS 2023 |
Machine learning › Efficient and distributed learning
data-centric learning |
0.2 | 1 | 2023 | Data-Centric Learning from Unlabeled Graphs with Diffusion Model · NeurIPS 2023 |
Methods — techniques the papers use, named apart from their topics
repeat-unit augmentation · 1.7graph maximum spanning tree alignment · 1.7self-training · 0.7self-supervised learning · 0.7pseudo-labeling · 0.7mixup · 0.7latent space augmentation · 0.7diffusion model · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Learning Repetition-Invariant Representations for Polymer InformaticsabstractPolymers are large macromolecules composed of repeating structural units known as monomers and are widely applied in fields such as energy storage, construction, medicine, and aerospace. However, existing graph neural network methods, though effective for small molecules, only model the single unit of polymers and fail to produce consistent vector representations for the true polymer structure with varying numbers of units. To address this challenge, we introduce Graph Repetition Invariance (GRIN), a novel method to learn polymer representations that are invariant to the number of repeating units in their graph representations. GRIN integrates a graph-based maximum spanning tree alignment with repeat-unit augmentation to ensure structural consistency. We provide theoretical guarantees for repetition‐invariance from both model and data perspectives, demonstrating that three repeating units are the minimal augmentation required for optimal invariant representation learning. GRIN outperforms state-of-the-art baselines on both homopolymer and copolymer benchmarks, learning stable, repetition-invariant representations that generalize effectively to polymer chains of unseen sizes. Yihan Zhu, Gang Liu 0025, Eric Inae, Tengfei Luo, Meng Jiang 0001 |
NeurIPS | 3 |
| 2024 | Rationalizing Graph Neural Networks with Data AugmentationabstractGraph rationales are representative subgraph structures that best explain and support the graph neural network (GNN) predictions. Graph rationalization involves the joint identification of these subgraphs during GNN training, resulting in improved interpretability and generalization. GNN is widely used for node-level tasks such as paper classification and graph-level tasks such as molecular property prediction. However, on both levels, little attention has been given to GNN rationalization and the lack of training examples makes it difficult to identify the optimal graph rationales. In this work, we address the problem by proposing a unified data augmentation framework with two novel operations on environment subgraphs to rationalize GNN prediction. We define the environment subgraph as the remaining subgraph after rationale identification and separation. The framework efficiently performs rationale–environment separation in the representation space for a node’s neighborhood graph or a graph’s complete structure to avoid the high complexity of explicit graph decoding and encoding. We conduct experiments on 17 datasets spanning node classification, graph classification, and graph regression. Results demonstrate that our framework is effective and efficient in rationalizing and enhancing GNNs for different levels of tasks on graphs. Gang Liu 0025, Eric Inae, Tengfei Luo, Meng Jiang 0001 |
ACM Trans. Knowl. Discov. Data | 2 |
| 2023 | Semi-Supervised Graph Imbalanced RegressionabstractData imbalance is easily found in annotated data when the observations of certain continuous label values are difficult to collect for regression tasks. When they come to molecule and polymer property predictions, the annotated graph datasets are often small because labeling them requires expensive equipment and effort. To address the lack of examples of rare label values in graph regression tasks, we propose a semi-supervised framework to progressively balance training data and reduce model bias via self-training. The training data balance is achieved by (1) pseudo-labeling more graphs for under-represented labels with a novel regression confidence measurement and (2) augmenting graph examples in latent space for remaining rare labels after data balancing with pseudo-labels. The former is to identify quality examples from unlabeled data whose labels are confidently predicted and sample a subset of them with a reverse distribution from the imbalanced annotated data. The latter collaborates with the former to target a perfect balance using a novel label-anchored mixup algorithm. We perform experiments in seven regression tasks on graph datasets. Results demonstrate that the proposed framework significantly reduces the error of predicted graph properties, especially in under-represented label areas. Gang Liu 0025, Tong Zhao 0003, Eric Inae, Tengfei Luo, Meng Jiang 0001 |
KDD | 3 |
| 2023 | Data-Centric Learning from Unlabeled Graphs with Diffusion ModelabstractGraph property prediction tasks are important and numerous. While each task offers a small size of labeled examples, unlabeled graphs have been collected from various sources and at a large scale. A conventional approach is training a model with the unlabeled graphs on self-supervised tasks and then fine-tuning the model on the prediction tasks. However, the self-supervised task knowledge could not be aligned or sometimes conflicted with what the predictions needed. In this paper, we propose to extract the knowledge underlying the large set of unlabeled graphs as a specific set of useful data points to augment each property prediction model. We use a diffusion model to fully utilize the unlabeled graphs and design two new objectives to guide the model's denoising process with each task's labeled data to generate task-specific graph examples and their labels. Experiments demonstrate that our data-centric approach performs significantly better than fifteen existing various methods on fifteen tasks. The performance improvement brought by unlabeled data is visible as the generated labeled examples unlike the self-supervised learning. Gang Liu 0025, Eric Inae, Tong Zhao 0003, Jiaxin Xu, Tengfei Luo, Meng Jiang 0001 |
NeurIPS | 2 |