Yejin Kim 0001

dblp:74/3301-1 · DBLP profile ↗
← Back
13ranked-venue papers
6as first author
7since 2021 · last 2024
0000-0001-7815-6310ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 10 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 3 first-authorDatabases, data management, data science and information retrieval · 2 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
YearPublicationVenuePosition
2024 Clinical outcome-guided deep temporal clustering for disease progression subtyping
Dulin Wang, Paul E. Schulz, Xiaoqian Jiang, Yejin Kim 0001
J. Biomed. Informatics5
2023 Deep single-cell RNA-seq data clustering with graph prototypical contrastive learning
abstract
MOTIVATION: Single-cell RNA sequencing enables researchers to study cellular heterogeneity at single-cell level. To this end, identifying cell types of cells with clustering techniques becomes an important task for downstream analysis. However, challenges of scRNA-seq data such as pervasive dropout phenomena hinder obtaining robust clustering outputs. Although existing studies try to alleviate these problems, they fall short of fully leveraging the relationship information and mainly rely on reconstruction-based losses that highly depend on the data quality, which is sometimes noisy. RESULTS: This work proposes a graph-based prototypical contrastive learning method, named scGPCL. Specifically, scGPCL encodes the cell representations using Graph Neural Networks on cell-gene graph that captures the relational information inherent in scRNA-seq data and introduces prototypical contrastive learning to learn cell representations by pushing apart semantically dissimilar pairs and pulling together similar ones. Through extensive experiments on both simulated and real scRNA-seq data, we demonstrate the effectiveness and efficiency of scGPCL. AVAILABILITY AND IMPLEMENTATION: Code is available at https://github.com/Junseok0207/scGPCL.
Junseok Lee 0002, Sungwon Kim 0002, Dongmin Hyun, Namkyeong Lee, Yejin Kim 0001, Chanyoung Park 0001
Bioinform.5
2023 Using artificial intelligence to learn optimal regimen plan for Alzheimer's disease
abstract
BACKGROUND: Alzheimer's disease (AD) is a progressive neurological disorder with no specific curative medications. Sophisticated clinical skills are crucial to optimize treatment regimens given the multiple coexisting comorbidities in the patient population. OBJECTIVE: Here, we propose a study to leverage reinforcement learning (RL) to learn the clinicians' decisions for AD patients based on the longitude data from electronic health records. METHODS: In this study, we selected 1736 patients from the Alzheimer's Disease Neuroimaging Initiative (ADNI) database. We focused on the two most frequent concomitant diseases-depression, and hypertension, thus creating 5 data cohorts (ie, Whole Data, AD, AD-Hypertension, AD-Depression, and AD-Depression-Hypertension). We modeled the treatment learning into an RL problem by defining states, actions, and rewards. We built a regression model and decision tree to generate multiple states, used six combinations of medications (ie, cholinesterase inhibitors, memantine, memantine-cholinesterase inhibitors, hypertension drugs, supplements, or no drugs) as actions, and Mini-Mental State Exam (MMSE) scores as rewards. RESULTS: Given the proper dataset, the RL model can generate an optimal policy (regimen plan) that outperforms the clinician's treatment regimen. Optimal policies (ie, policy iteration and Q-learning) had lower rewards than the clinician's policy (mean -3.03 and -2.93 vs. -2.93, respectively) for smaller datasets but had higher rewards for larger datasets (mean -4.68 and -2.82 vs. -4.57, respectively). CONCLUSIONS: Our results highlight the potential of using RL to generate the optimal treatment based on the patients' longitude records. Our work can lead the path towards developing RL-based decision support systems that could help manage AD with comorbidities.
Kritib Bhattarai, Sivaraman Rajaganapathy, Trisha Das, Yejin Kim 0001, Yongbin Chen, Qiying Dai, Xiaoqian Jiang, Nansu Zong
J. Am. Medical Informatics Assoc.4
2023 Emulate randomized clinical trials using heterogeneous treatment effect estimation for personalized treatments: Methodology review and benchmark
Yaobin Ling, Pulakesh Upadhyaya, Xiaoqian Jiang, Yejin Kim 0001
J. Biomed. Informatics5
2022 Relational graph convolutional networks for predicting blood-brain barrier penetration of drug molecules
abstract
MOTIVATION: Evaluating the blood-brain barrier (BBB) permeability of drug molecules is a critical step in brain drug development. Traditional methods for the evaluation require complicated in vitro or in vivo testing. Alternatively, in silico predictions based on machine learning have proved to be a cost-efficient way to complement the in vitro and in vivo methods. However, the performance of the established models has been limited by their incapability of dealing with the interactions between drugs and proteins, which play an important role in the mechanism behind the BBB penetrating behaviors. To address this limitation, we employed the relational graph convolutional network (RGCN) to handle the drug-protein interactions as well as the properties of each individual drug. RESULTS: The RGCN model achieved an overall accuracy of 0.872, an area under the receiver operating characteristic (AUROC) of 0.919 and an area under the precision-recall curve (AUPRC) of 0.838 for the testing dataset with the drug-protein interactions and the Mordred descriptors as the input. Introducing drug-drug similarity to connect structurally similar drugs in the data graph further improved the testing results, giving an overall accuracy of 0.876, an AUROC of 0.926 and an AUPRC of 0.865. In particular, the RGCN model was found to greatly outperform the LightGBM base model when evaluated with the drugs whose BBB penetration was dependent on drug-protein interactions. Our model is expected to provide high-confidence predictions of BBB permeability for drug prioritization in the experimental screening of BBB-penetrating drugs. AVAILABILITY AND IMPLEMENTATION: The data and the codes are freely available at https://github.com/dingyan20/BBB-Penetration-Prediction. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Xiaoqian Jiang, Yejin Kim 0001
Bioinform.3
2021 Anticancer drug synergy prediction in understudied tissues using transfer learning
abstract
OBJECTIVE: Drug combination screening has advantages in identifying cancer treatment options with higher efficacy without degradation in terms of safety. A key challenge is that the accumulated number of observations in in-vitro drug responses varies greatly among different cancer types, where some tissues are more understudied than the others. Thus, we aim to develop a drug synergy prediction model for understudied tissues as a way of overcoming data scarcity problems. MATERIALS AND METHODS: We collected a comprehensive set of genetic, molecular, phenotypic features for cancer cell lines. We developed a drug synergy prediction model based on multitask deep neural networks to integrate multimodal input and multiple output. We also utilized transfer learning from data-rich tissues to data-poor tissues. RESULTS: We showed improved accuracy in predicting synergy in both data-rich tissues and understudied tissues. In data-rich tissue, the prediction model accuracy was 0.9577 AUROC for binarized classification task and 174.3 mean squared error for regression task. We observed that an adequate transfer learning strategy significantly increases accuracy in the understudied tissues. CONCLUSIONS: Our synergy prediction model can be used to rank synergistic drug combinations in understudied tissues and thus help to prioritize future in-vitro experiments. Code is available at https://github.com/yejinjkim/synergy-transfer.
Yejin Kim 0001, Jing Tang 0002, W. Jim Zheng, Xiaoqian Jiang
J. Am. Medical Informatics Assoc.1
2021 Population stratification enables modeling effects of reopening policies on mortality and hospitalization rates
Tongtong Huang, Yan Chu 0005, Shayan Shams, Yejin Kim 0001, Ananth V. Annapragada, Devika Subramanian, Ioannis A. Kakadiaris, Assaf Gottlieb, Xiaoqian Jiang
J. Biomed. Informatics4
2020 SCOR: A secure international informatics infrastructure to investigate COVID-19
abstract
Global pandemics call for large and diverse healthcare data to study various risk factors, treatment options, and disease progression patterns. Despite the enormous efforts of many large data consortium initiatives, scientific community still lacks a secure and privacy-preserving infrastructure to support auditable data sharing and facilitate automated and legally compliant federated analysis on an international scale. Existing health informatics systems do not incorporate the latest progress in modern security and federated machine learning algorithms, which are poised to offer solutions. An international group of passionate researchers came together with a joint mission to solve the problem with our finest models and tools. The SCOR Consortium has developed a ready-to-deploy secure infrastructure using world-class privacy and security technologies to reconcile the privacy/utility conflicts. We hope our effort will make a change and accelerate research in future pandemics with broad and diverse samples on an international scale.
Jean Louis Raisaro, Juan Ramón Troncoso-Pastoriza, Raphaelle Beau-Lejdstrom, Riccardo Bellazzi, Robert Murphy, Elmer V. Bernstam, Henry Wang, Mauro Bucalo, Yong Chen 0016, Assaf Gottlieb, Arif Ozgun Harmanci, Miran Kim, Yejin Kim 0001, Jeffrey G. Klann, Catherine Klersy, Bradley A. Malin, Marie Méan, Fabian Prasser, Luigia Scudeller, Ali Torkamani, Julien Vaucher, Mamta Puppala, Stephen T. C. Wong, Milana Frenkel-Morgenstern, Hua Xu 0001, Baba Maiyaki Musa, Abdulrazaq G. Habib, Trevor Cohen, Adam B. Wilcox, Hamisu M. Salihu, Heidi Sofia, Xiaoqian Jiang, Jean-Pierre Hubaux
J. Am. Medical Informatics Assoc.14
2020 Temporal phenotyping for transitional disease progress: An application to epilepsy and Alzheimer's disease
Yejin Kim 0001, Samden D. Lhatoo, Guo-Qiang Zhang 0001, Xiaoqian Jiang
J. Biomed. Informatics1
2019 Discriminative Sleep Patterns of Alzheimer's Disease via Tensor Factorization
Yejin Kim 0001, Xiaoqian Jiang, Licong Cui
AMIA1
2019 Sequential and Diverse Recommendation with Long Tail
abstract
Sequential recommendation is a task that learns a temporal dynamic of a user behavior in sequential data and predicts items that a user would like afterward. However, diversity has been rarely emphasized in the context of sequential recommendation. Sequential and diverse recommendation must learn temporal preference on diverse items as well as on general items. Thus, we propose a sequential and diverse recommendation model that predicts a ranked list containing general items and also diverse items without compromising significant accuracy.To learn temporal preference on diverse items as well as on general items, we cluster and relocate consumed long tail items to make a pseudo ground truth for diverse items and learn the preference on long tail using recurrent neural network, which enables us to directly learn a ranking function. Extensive online and offline experiments deployed on a commercial platform demonstrate that our models significantly increase diversity while preserving accuracy compared to the state-of-the-art sequential recommendation model, and consequently our models improve user satisfaction.
Yejin Kim 0001, Kwangseob Kim, Chanyoung Park 0001, Hwanjo Yu
IJCAI1
2017 DiagTree: Diagnostic Tree for Differential Diagnosis
abstract
Differential diagnosis is detection of one disease among similar diseases using evidence such as pathologic tests. A Partially Observed Markov Decision Process (POMDP) formulates the complex differential diagnosis process into a probabilistic decision-making model. However, differential diagnosis is not often fully formulated as POMDP because model construction does not consider the cost (or time) to finish the diagnosis process, or the practical convention on clinical tests. We propose a Diagnostic Tree (DiagTree), a new framework for diagnosing diseases, which combines several tests to reduce the diagnosis time and to incorporate real-world constraints into discrete optimization. DiagTree consists of multiple tests in internal nodes and posterior probabilities ("confidences") that the patient suffers the disease listed at each leaf node. The confidences are computed after a series of test results is applied in internal nodes. DiagTree is built to maximize the confidences at leaf nodes and to minimize the decision process time. We formulate this problem as integer programming and solve it by the Branch-and-Bound method and a greedy approach. We apply DiagTree to immunohistochemistry profiles to detect lymphoid neoplasms. We evaluate the accuracy and cost of the diagnosis rules from DiagTree compared to those obtained using rules that clinicians derived from their experience. DiagTree detected diseases with high accuracy and also reduced the diagnosis cost (or time) compared to the existing rules of clinicians. DiagTree can support clinicians by suggesting a simple diagnosis process with high accuracy and low cost among test candidates.
Yejin Kim 0001, Jingyun Choi, Yosep Chong, Xiaoqian Jiang, Hwanjo Yu
CIKM1
2017 Federated Tensor Factorization for Computational Phenotyping
abstract
Tensor factorization models offer an effective approach to convert massive electronic health records into meaningful clinical concepts (phenotypes) for data analysis. These models need a large amount of diverse samples to avoid population bias. An open challenge is how to derive phenotypes jointly across multiple hospitals, in which direct patient-level data sharing is not possible (e.g., due to institutional policies). In this paper, we developed a novel solution to enable federated tensor factorization for computational phenotyping without sharing patient-level data. We developed secure data harmonization and federated computation procedures based on alternating direction method of multipliers (ADMM). Using this method, the multiple hospitals iteratively update tensors and transfer secure summarized information to a central server, and the server aggregates the information to generate phenotypes. We demonstrated with real medical datasets that our method resembles the centralized training model (based on combined datasets) in terms of accuracy and phenotypes discovery while respecting privacy.
Yejin Kim 0001, Jimeng Sun 0001, Hwanjo Yu, Xiaoqian Jiang
KDD1