Jaesik Kim

dblp:219/0099 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2025
0000-0003-1203-9650ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
3 papers
Bioinformatics and computational biology · 56% Computational science and engineering · 44%
Artificial intelligence
1 paper
Representation and self-supervised learning · 100%

Topics — the 5 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning › representation learning › structured representation learning
set representation learning
0.912025
MAESTRO: Masked Encoding Set Transformer with Self-Distillation · ICLR 2025
Bioinformatics and computational biology › single-cell analysis
cytometry data analysis
0.912025
MAESTRO: Masked Encoding Set Transformer with Self-Distillation · ICLR 2025
Computational science and engineering › graph learning › network embedding
knowledge graph embedding
0.912025
GeOKG: geometry-aware knowledge graph embedding for Gene Ontology and genes · Bioinform. 2025
Bioinformatics and computational biology
protein-protein interaction prediction
0.912025
GeOKG: geometry-aware knowledge graph embedding for Gene Ontology and genes · Bioinform. 2025
Computational science and engineering › graph learning
hyperbolic embedding
0.512021
HiG2Vec: hierarchical representations of Gene Ontology and genes in the Poincaré ball · Bioinform. 2021

Methods — techniques the papers use, named apart from their topics

self-distillation · 1.7attention · 1.7hyperbolic embedding · 0.9geometric interaction · 0.9euclidean embedding · 0.9word2vec · 0.5poincaré embedding · 0.5hyperbolic geometry · 0.5
YearPublicationVenuePosition
2025 MAESTRO: Masked Encoding Set Transformer with Self-Distillation
abstract
The interrogation of cellular states and interactions in immunology research is an ever-evolving task, requiring adaptation to the current levels of high dimensionality. Cytometry enables high-dimensional profiling of immune cells, but its analysis is hindered by the complexity and variability of the data. We present MAESTRO, a self-supervised set representation learning model that generates vector representations of set-structured data, which we apply to learn immune profiles from cytometry data. Unlike previous studies only learn cell-level representations, whereas MAESTRO uses all of a sample's cells to learn a set representation. MAESTRO leverages specialized attention mechanisms to handle sets of variable number of cells and ensure permutation invariance, coupled with an online tokenizer by self-distillation framework. We benchmarked our model against existing cytometry approaches and other existing machine learning methods that have never been applied in cytometry. Our model outperforms existing approaches in retrieving cell-type proportions and capturing clinically relevant features for downstream tasks such as disease diagnosis and immune cell profiling.
Matthew Eric Lee, Jaesik Kim, Matei Ionita, Michelle L. McKeague, Yonghyun Nam, Irene Khavin, Yidi Huang, Victoria Fang, Sokratis Apostolidis, Divij Mathew, Shwetank, Ajinkya Pattekar, Zahabia Rangwala, Amit Bar-Or Tillinger, Benjamin A. Fensterheim, Benjamin A. Abramoff, Rennie L. Rhee, Damian Maseda, Allison R. Greenplate
ICLR2
2025 GeOKG: geometry-aware knowledge graph embedding for Gene Ontology and genes
abstract
MOTIVATION: Leveraging deep learning for the representation learning of Gene Ontology (GO) and Gene Ontology Annotation (GOA) holds significant promise for enhancing downstream biological tasks such as protein-protein interaction prediction. Prior approaches have predominantly used text- and graph-based methods, embedding GO and GOA in a single geometric space (e.g. Euclidean or hyperbolic). However, since the GO graph exhibits a complex and nonmonotonic hierarchy, single-space embeddings are insufficient to fully capture its structural nuances. RESULTS: In this study, we address this limitation by exploiting geometric interaction to better reflect the intricate hierarchical structure of GO. Our proposed method, Geometry-Aware Knowledge Graph Embeddings for GO and Genes (GeOKG), leverages interactions among various geometric representations during training, thereby modeling the complex hierarchy of GO more effectively. Experiments at the GO level demonstrate the benefits of incorporating these geometric interactions, while gene-level tests reveal that GeOKG outperforms existing methods in protein-protein interaction prediction. These findings highlight the potential of using geometric interaction for embedding heterogeneous biomedical networks. AVAILABILITY AND IMPLEMENTATION: https://github.com/ukjung21/GeOKG.
Chang-Uk Jeong, Jaesik Kim, Do Kyoon Kim, Kyung-Ah Sohn 0001
Bioinform.2
2021 Interpretable temporal graph neural network for prognostic prediction of Alzheimer's disease using longitudinal neuroimaging data
abstract
Alzheimer's disease (AD) is a progressive neurodegenerative brain disorder characterized by memory loss and cognitive decline. Early detection and accurate prognosis of AD is an important research topic, and numerous machine learning methods have been proposed to solve this problem. However, traditional machine learning models are facing challenges in effectively integrating longitudinal neuroimaging data and biologically meaningful structure and knowledge to build accurate and interpretable prognostic predictors. To bridge this gap, we propose an interpretable graph neural network (GNN) model for AD prognostic prediction based on longitudinal neuroimaging data while embracing the valuable knowledge of structural brain connectivity. In our empirical study, we demonstrate that 1) the proposed model outperforms several competing models (i.e., DNN, SVM) in terms of prognostic prediction accuracy, and 2) our model can capture neuroanatomical contribution to the prognostic predictor and yield biologically meaningful interpretation to facilitate better mechanistic understanding of the Alzheimer's disease. Source code is available at https://github.com/JaesikKim/temporal-GNN.
Mansu Kim, Jaesik Kim, Jeffrey Qu, Heng Huang 0001, Qi Long, Kyung-Ah Sohn 0001, Do Kyoon Kim, Li Shen 0001
BIBM2
2021 HiG2Vec: hierarchical representations of Gene Ontology and genes in the Poincaré ball
abstract
MOTIVATION: Knowledge manipulation of Gene Ontology (GO) and Gene Ontology Annotation (GOA) can be done primarily by using vector representation of GO terms and genes. Previous studies have represented GO terms and genes or gene products in Euclidean space to measure their semantic similarity using an embedding method such as the Word2Vec-based method to represent entities as numeric vectors. However, this method has the limitation that embedding large graph-structured data in the Euclidean space cannot prevent a loss of information of latent hierarchies, thus precluding the semantics of GO and GOA from being captured optimally. On the other hand, hyperbolic spaces such as the Poincaré balls are more suitable for modeling hierarchies, as they have a geometric property in which the distance increases exponentially as it nears the boundary because of negative curvature. RESULTS: In this article, we propose hierarchical representations of GO and genes (HiG2Vec) by applying Poincaré embedding specialized in the representation of hierarchy through a two-step procedure: GO embedding and gene embedding. Through experiments, we show that our model represents the hierarchical structure better than other approaches and predicts the interaction of genes or gene products similar to or better than previous studies. The results indicate that HiG2Vec is superior to other methods in capturing the GO and gene semantics and in data utilization as well. It can be robustly applied to manipulate various biological knowledge. AVAILABILITYAND IMPLEMENTATION: https://github.com/JaesikKim/HiG2Vec. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jaesik Kim, Do Kyoon Kim, Kyung-Ah Sohn 0001
Bioinform.1