Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Eric Inae

dblp:342/8313 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
4since 2021 · last 2025
0009-0002-2101-2126ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Graph learning · 58% Generative modeling · 32% Deep learning architectures and training · 5%
Databases, data mining, and information retrieval
1 paper
Data mining · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Graph learning
graph neural network
0.912025
Learning Repetition-Invariant Representations for Polymer Informatics · NeurIPS 2025
Machine learning › Graph learning › graph representation learning
invariant graph representation
0.912025
Learning Repetition-Invariant Representations for Polymer Informatics · NeurIPS 2025
Machine learning › Generative modeling
diffusion model
0.712023
Data-Centric Learning from Unlabeled Graphs with Diffusion Model · NeurIPS 2023
Machine learning › Generative modeling › diffusion model
graph diffusion model
0.712023
Data-Centric Learning from Unlabeled Graphs with Diffusion Model · NeurIPS 2023
Machine learning › Graph learning › graph inference › graph prediction
graph property prediction
0.712023
Data-Centric Learning from Unlabeled Graphs with Diffusion Model · NeurIPS 2023
Data mining › structured data mining
graph mining
0.712023
Semi-Supervised Graph Imbalanced Regression · KDD 2023
Data mining
semi-supervised learning
0.712023
Semi-Supervised Graph Imbalanced Regression · KDD 2023
Machine learning › Deep learning architectures and training
data augmentation
0.212023
Data-Centric Learning from Unlabeled Graphs with Diffusion Model · NeurIPS 2023
Machine learning › Efficient and distributed learning
data-centric learning
0.212023
Data-Centric Learning from Unlabeled Graphs with Diffusion Model · NeurIPS 2023

Methods — techniques the papers use, named apart from their topics

repeat-unit augmentation · 1.7graph maximum spanning tree alignment · 1.7self-training · 0.7self-supervised learning · 0.7pseudo-labeling · 0.7mixup · 0.7latent space augmentation · 0.7diffusion model · 0.7
YearPublicationVenuePosition
2025 Learning Repetition-Invariant Representations for Polymer Informatics
abstract
Polymers are large macromolecules composed of repeating structural units known as monomers and are widely applied in fields such as energy storage, construction, medicine, and aerospace. However, existing graph neural network methods, though effective for small molecules, only model the single unit of polymers and fail to produce consistent vector representations for the true polymer structure with varying numbers of units. To address this challenge, we introduce Graph Repetition Invariance (GRIN), a novel method to learn polymer representations that are invariant to the number of repeating units in their graph representations. GRIN integrates a graph-based maximum spanning tree alignment with repeat-unit augmentation to ensure structural consistency. We provide theoretical guarantees for repetition‐invariance from both model and data perspectives, demonstrating that three repeating units are the minimal augmentation required for optimal invariant representation learning. GRIN outperforms state-of-the-art baselines on both homopolymer and copolymer benchmarks, learning stable, repetition-invariant representations that generalize effectively to polymer chains of unseen sizes.
Yihan Zhu, Gang Liu 0025, Eric Inae, Tengfei Luo, Meng Jiang 0001
NeurIPS3
2024 Rationalizing Graph Neural Networks with Data Augmentation
abstract
Graph rationales are representative subgraph structures that best explain and support the graph neural network (GNN) predictions. Graph rationalization involves the joint identification of these subgraphs during GNN training, resulting in improved interpretability and generalization. GNN is widely used for node-level tasks such as paper classification and graph-level tasks such as molecular property prediction. However, on both levels, little attention has been given to GNN rationalization and the lack of training examples makes it difficult to identify the optimal graph rationales. In this work, we address the problem by proposing a unified data augmentation framework with two novel operations on environment subgraphs to rationalize GNN prediction. We define the environment subgraph as the remaining subgraph after rationale identification and separation. The framework efficiently performs rationale–environment separation in the representation space for a node’s neighborhood graph or a graph’s complete structure to avoid the high complexity of explicit graph decoding and encoding. We conduct experiments on 17 datasets spanning node classification, graph classification, and graph regression. Results demonstrate that our framework is effective and efficient in rationalizing and enhancing GNNs for different levels of tasks on graphs.
Gang Liu 0025, Eric Inae, Tengfei Luo, Meng Jiang 0001
ACM Trans. Knowl. Discov. Data2
2023 Semi-Supervised Graph Imbalanced Regression
abstract
Data imbalance is easily found in annotated data when the observations of certain continuous label values are difficult to collect for regression tasks. When they come to molecule and polymer property predictions, the annotated graph datasets are often small because labeling them requires expensive equipment and effort. To address the lack of examples of rare label values in graph regression tasks, we propose a semi-supervised framework to progressively balance training data and reduce model bias via self-training. The training data balance is achieved by (1) pseudo-labeling more graphs for under-represented labels with a novel regression confidence measurement and (2) augmenting graph examples in latent space for remaining rare labels after data balancing with pseudo-labels. The former is to identify quality examples from unlabeled data whose labels are confidently predicted and sample a subset of them with a reverse distribution from the imbalanced annotated data. The latter collaborates with the former to target a perfect balance using a novel label-anchored mixup algorithm. We perform experiments in seven regression tasks on graph datasets. Results demonstrate that the proposed framework significantly reduces the error of predicted graph properties, especially in under-represented label areas.
Gang Liu 0025, Tong Zhao 0003, Eric Inae, Tengfei Luo, Meng Jiang 0001
KDD3
2023 Data-Centric Learning from Unlabeled Graphs with Diffusion Model
abstract
Graph property prediction tasks are important and numerous. While each task offers a small size of labeled examples, unlabeled graphs have been collected from various sources and at a large scale. A conventional approach is training a model with the unlabeled graphs on self-supervised tasks and then fine-tuning the model on the prediction tasks. However, the self-supervised task knowledge could not be aligned or sometimes conflicted with what the predictions needed. In this paper, we propose to extract the knowledge underlying the large set of unlabeled graphs as a specific set of useful data points to augment each property prediction model. We use a diffusion model to fully utilize the unlabeled graphs and design two new objectives to guide the model's denoising process with each task's labeled data to generate task-specific graph examples and their labels. Experiments demonstrate that our data-centric approach performs significantly better than fifteen existing various methods on fifteen tasks. The performance improvement brought by unlabeled data is visible as the generated labeled examples unlike the self-supervised learning.
Gang Liu 0025, Eric Inae, Tong Zhao 0003, Jiaxin Xu, Tengfei Luo, Meng Jiang 0001
NeurIPS2