Gyoung S. Na

dblp:223/3192 · DBLP profile ↗
← Back
15ranked-venue papers
10as first author
13since 2021 · last 2025
0000-0001-9803-0782ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 9 first-author · 12 since 2021Databases, data management, data science and information retrieval · 7 · 6 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Self-Supervised Diffusion Models for Electron-Aware Molecular Representation Learning
abstract
Physical properties derived from electronic distributions are essential information that determines molecular properties. However, the electron-level information is not accessible in most real-world complex molecules due to the extensive computational costs of determining uncertain electronic distributions. For this reason, existing methods for molecular property prediction have remained in regression models on simplified atom-level molecular descriptors, such as atomic structures and fingerprints. This paper proposes an efficient knowledge transfer method for electron-aware molecular representation learning. To this end, we devised a self-supervised diffusion method that estimates the electron-level information of real-world complex molecules without expensive quantum mechanical calculations. The proposed method achieved state-of-the-art prediction accuracy in the tasks of predicting molecular properties on extensive real-world molecular datasets.
Gyoung S. Na, Chanyoung Park 0001
ICLR1
2025 Electron-Informed Coarse-Graining Molecular Representation Learning for Real-World Molecular Physics
abstract
Various representation learning methods for molecular structures have been devised to accelerate data-driven chemistry. However, the representation capabilities of existing methods are essentially limited to atom-level information, which is not sufficient to describe real-world molecular physics. Although electron-level information can provide fundamental knowledge about chemical compounds beyond the atom-level information, obtaining the electron-level information in real-world molecules is computationally impractical and sometimes infeasible. We propose a method for learning electron-informed molecular representations without additional computation costs by transferring readily accessible electron-level information about small molecules to large molecules of our interest. The proposed method achieved state-of-the-art prediction accuracy on extensive benchmark datasets containing experimentally observed molecular physics. The source code for HEDMoL is available at https://github.com/ngs00/HEDMoL.
Gyoung S. Na, Chanyoung Park 0001
KDD (1)1
2025 3D Interaction Geometric Pre-training for Molecular Relational Learning
abstract
Molecular Relational Learning (MRL) is a rapidly growing field that focuses on understanding the interaction dynamics between molecules, which is crucial for applications ranging from catalyst engineering to drug discovery. Despite recent progress, earlier MRL approaches are limited to using only the 2D topological structure of molecules, as obtaining the 3D interaction geometry remains prohibitively expensive. This paper introduces a novel 3D geometric pre-training strategy for MRL (3DMRL) that incorporates a 3D virtual interaction environment, overcoming the limitations of costly traditional quantum mechanical calculation methods. With the constructed 3D virtual interaction environment, 3DMRL trains 2D MRL model to learn the global and local 3D geometric information of molecular interaction. Extensive experiments on various tasks using real-world datasets, including out-of-distribution and extrapolation scenarios, demonstrate the effectiveness of 3DMRL, showing up to a 24.93% improvement in performance across 40 tasks. Our code is publicly available at https://github.com/Namkyeong/3DMRL.
Namkyeong Lee, Yunhak Oh, Heewoong Noh, Gyoung S. Na, Minkai Xu, Hanchen Wang 0002, Tianfan Fu, Chanyoung Park 0001
NeurIPS4
2024 Retrieval-Retro: Retrieval-based Inorganic Retrosynthesis with Expert Knowledge
abstract
While inorganic retrosynthesis planning is essential in the field of chemical science, the application of machine learning in this area has been notably less explored compared to organic retrosynthesis planning. In this paper, we propose Retrieval-Retro for inorganic retrosynthesis planning, which implicitly extracts the precursor information of reference materials that are retrieved from the knowledge base regarding domain expertise in the field. Specifically, instead of directly employing the precursor information of reference materials, we propose implicitly extracting it with various attention layers, which enables the model to learn novel synthesis recipes more effectively. Moreover, during retrieval, we consider the thermodynamic relationship between target material and precursors, which is essential domain expertise in identifying the most probable precursor set among various options. Extensive experiments demonstrate the superiority of Retrieval-Retro in retrosynthesis planning, especially in discovering novel synthesis recipes, which is crucial for materials discovery. The source code for Retrieval-Retro is available at https://github.com/HeewoongNoh/Retrieval-Retro.
Heewoong Noh, Namkyeong Lee, Gyoung S. Na, Chanyoung Park 0001
NeurIPS3
2023 Conditional Graph Information Bottleneck for Molecular Relational Learning
abstract
Molecular relational learning, whose goal is to learn the interaction behavior between molecular pairs, got a surge of interest in molecular sciences due to its wide range of applications. Recently, graph neural networks have recently shown great success in molecular relational learning by modeling a molecule as a graph structure, and considering atom-level interactions between two molecules. Despite their success, existing molecular relational learning methods tend to overlook the nature of chemistry, i.e., a chemical compound is composed of multiple substructures such as functional groups that cause distinctive chemical reactions. In this work, we propose a novel relational learning framework, called CGIB, that predicts the interaction behavior between a pair of graphs by detecting core subgraphs therein. The main idea is, given a pair of graphs, to find a subgraph from a graph that contains the minimal sufficient information regarding the task at hand conditioned on the paired graph based on the principle of conditional graph information bottleneck. We argue that our proposed method mimics the nature of chemical reactions, i.e., the core substructure of a molecule varies depending on which other molecule it interacts with. Extensive experiments on various tasks with real-world datasets demonstrate the superiority of CGIB over state-of-the-art baselines. Our code is available at https://github.com/Namkyeong/CGIB.
Namkyeong Lee, Dongmin Hyun, Gyoung S. Na, Sungwon Kim 0002, Junseok Lee 0002, Chanyoung Park 0001
ICML3
2023 Shift-Robust Molecular Relational Learning with Causal Substructure
abstract
Recently, molecular relational learning, whose goal is to predict the interaction behavior between molecular pairs, got a surge of interest in molecular sciences due to its wide range of applications. In this work, we propose CMRL that is robust to the distributional shift in molecular relational learning by detecting the core substructure that is causally related to chemical reactions. To do so, we first assume a causal relationship based on the domain knowledge of molecular sciences and construct a structural causal model (SCM) that reveals the relationship between variables. Based on the SCM, we introduce a novel conditional intervention framework whose intervention is conditioned on the paired molecule. With the conditional intervention framework, our model successfully learns from the causal substructure and alleviates the confounding effect of shortcut substructures that are spuriously correlated to chemical reactions. Extensive experiments on various tasks with real-world and synthetic datasets demonstrate the superiority of CMRL over state-of-the-art baseline models.
Namkyeong Lee, Kanghoon Yoon, Gyoung S. Na, Sein Kim, Chanyoung Park 0001
KDD3
2023 Density of States Prediction of Crystalline Materials via Prompt-guided Multi-Modal Transformer
abstract
The density of states (DOS) is a spectral property of crystalline materials, which provides fundamental insights into various characteristics of the materials. While previous works mainly focus on obtaining high-quality representations of crystalline materials for DOS prediction, we focus on predicting the DOS from the obtained representations by reflecting the nature of DOS: DOS determines the general distribution of states as a function of energy. That is, DOS is not solely determined by the crystalline material but also by the energy levels, which has been neglected in previous works. In this paper, we propose to integrate heterogeneous information obtained from the crystalline materials and the energies via a multi-modal transformer, thereby modeling the complex relationships between the atoms in the crystalline materials and various energy levels for DOS prediction. Moreover, we propose to utilize prompts to guide the model to learn the crystal structural system-specific interactions between crystalline materials and energies. Extensive experiments on two types of DOS, i.e., Phonon DOS and Electron DOS, with various real-world scenarios demonstrate the superiority of DOSTransformer. The source code for DOSTransformer is available at https://github.com/HeewoongNoh/DOSTransformer.
Namkyeong Lee, Heewoong Noh, Sungwon Kim 0002, Dongmin Hyun, Gyoung S. Na, Chanyoung Park 0001
NeurIPS5
2022 Conditional Graph Regression for Complex Chemical Systems with Heterogeneous Substructures
abstract
Graph neural networks (GNNs) have been widely studied as an efficient and generalized method to predict the physical and chemical properties of chemical systems based on a single homogeneous atomic structure, such as molecule and crystalline material. However, most chemical systems in real-world applications of materials science and engineering contain multiple heterogeneous atomic substructures. Nonetheless, existing GNNs for chemical applications assumed homogeneous graphs with the input node or edge features in the same feature space. In this paper, we reformulate the regression problem on chemical systems as a regression problem on core substructures conditioned by environment substructures. Then, we propose conditional atomic subgraph interaction network (CASIN) that predicts the physical and chemical properties of the input chemical systems by learning the atomic interactions between the heterogeneous core and environment substructures in the chemical systems. For three real-world benchmark chemical datasets, CASIN outperformed state-of-the-art GNNs in predicting the physical and chemical properties of the complex solar cell materials and catalyst systems.
Gyoung S. Na
IEEE Big Data1
2022 Nonlinearity Encoding for Extrapolation of Neural Networks
abstract
Extrapolation to predict unseen data outside the training distribution is a common challenge in real-world scientific applications of physics and chemistry. However, the extrapolation capabilities of neural networks have not been extensively studied in machine learning. Although it has been recently revealed that neural networks become linear regression in extrapolation problems, a universally applicable method to support the extrapolation of neural networks in general regression settings has not been investigated. In this paper, we propose automated nonlinearity encoder (ANE) that is a data-agnostic embedding method to improve the extrapolation capabilities of neural networks by conversely linearizing the original input-to-target relationships without architectural modifications of prediction models. ANE achieved state-of-the-art extrapolation accuracies in extensive scientific applications of various data formats. As a real-world application, we applied ANE for high-throughput screening to discover novel solar cell materials, and ANE significantly improved the screening accuracy.
Gyoung S. Na, Chanyoung Park 0001
KDD1
2022 Eigen-guided deep metric learning
Gyoung S. Na
Expert Syst. Appl.1
2022 Efficient learning rate adaptation based on hierarchical optimization approach
abstract
This paper proposes a new hierarchical approach to learning rate adaptation in gradient methods, called learning rate optimization (LRO). LRO formulates the learning rate adaption problem as a hierarchical optimization problem that minimizes the loss function with respect to the learning rate for current model parameters and gradients. Then, LRO optimizes the learning rate based on the alternating direction method of multipliers (ADMM). In the process of this learning rate optimization, LRO does not require any second-order information and probabilistic model, so it is highly efficient. Furthermore, LRO does not require any additional hyperparameters when compared to the vanilla gradient method with the simple exponential learning rate decay. In the experiments, we integrated LRO with vanilla SGD and Adam. Then, we compared their optimization performance with the state-of-the-art learning rate adaptation methods and also the most commonly-used adaptive gradient methods. The SGD and Adam with LRO outperformed all the competitors on the benchmark datasets in image classification tasks.
Gyoung S. Na
Neural Networks1
2022 Unsupervised Subspace Extraction via Deep Kernelized Clustering
abstract
Feature extraction has been widely studied to find informative latent features and reduce the dimensionality of data. In particular, due to the difficulty in obtaining labeled data, unsupervised feature extraction has received much attention in data mining. However, widely used unsupervised feature extraction methods require side information about data or rigid assumptions on the latent feature space. Furthermore, most feature extraction methods require predefined dimensionality of the latent feature space,which should be manually tuned as a hyperparameter. In this article, we propose a new unsupervised feature extraction method called Unsupervised Subspace Extractor ( USE ), which does not require any side information and rigid assumptions on data. Furthermore, USE can find a subspace generated by a nonlinear combination of the input feature and automatically determine the optimal dimensionality of the subspace for the given nonlinear combination. The feature extraction process of USE is well justified mathematically, and we also empirically demonstrate the effectiveness of USE for several benchmark datasets.
Gyoung S. Na, Hyunju Chang
ACM Trans. Knowl. Discov. Data1
2021 Reverse graph self-attention for target-directed atomic importance estimation
Gyoung S. Na, Hyun Woo Kim 0004
Neural Networks1
2020 Scale-Aware Graph-Based Machine Learning for Accurate Molecular Property Prediction
abstract
With great growth in the volume of chemical databases, machine learning receives significant attention from various scientific communities for efficient high-throughput screening of molecular properties and drug discovery on the millions of chemical compounds. In particular, graph neural networks (GNNs) have been widely studied in chemistry-related fields because a molecule is natively represented as a mathematical graph. In GNNs for the graph-level analysis, a global operation called readout is applied after node embedding to generate a graph-level embedding that represents characteristics of the whole graph. However, commonly used readouts frequently distort scale information of the graph and consequently degrade the prediction accuracy of GNNs. This problem becomes more serious in molecular machine learning because molecules have many important scale information (e.g., molecular weight and total energy). In this paper, we investigate this scale distortion problem in GNNs caused by the readouts for the first time and propose an efficient solution with a new attention-based readout. In the experiments, the proposed readout outperformed commonly used readouts on various GNN architectures.
Gyoung S. Na, Hyun Woo Kim 0004, Hyunju Chang
IEEE BigData1
2018 DILOF: Effective and Memory Efficient Local Outlier Detection in Data Streams
abstract
With precipitously growing demand to detect outliers in data streams, many studies have been conducted aiming to develop extensions of well-known outlier detection algorithm called Local Outlier Factor (LOF), for data streams. However, existing LOF-based algorithms for data streams still suffer from two inherent limitations: 1) Large amount of memory space is required. 2) A long sequence of outliers is not detected. In this paper, we propose a new outlier detection algorithm for data streams, called DILOF that effectively overcomes the limitations. To this end, we first develop a novel density-based sampling algorithm to summarize past data and then propose a new strategy for detecting a sequence of outliers. It is worth noting that our proposing algorithms do not require any prior knowledge or assumptions on data distribution. Moreover, we accelerate the execution time of DILOF about 15 times by developing a powerful distance approximation technique. Our comprehensive experiments on real-world datasets demonstrate that DILOF significantly outperforms the state-of-the-art competitors in terms of accuracy and execution time. The source code for the proposed algorithm is available at our website: http://di.postech.ac.kr/DILOF.
Gyoung S. Na, Donghyun Kim 0007, Hwanjo Yu
KDD1