VLDB 2026 Research / reviewers in the wild / expert
Sun Kim
dblp:53/1114
· DBLP profile ↗
96ranked-venue papers
18as first author
30since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 75 · 10 first-author · 23 since 2021Artificial intelligence and machine learning · 21 · 8 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Software engineering, systems software and programming languages · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EnsDTI: Predicting Drug-Target Interaction With Mixture-of-Experts and Confidence AssessmentabstractAccurately identifying drug-target interactions (DTIs) is a critical step in drug discovery. While structure-based drug design methods demonstrate impressive docking prediction accuracy, their heavy computational demands and resource-intensive nature make them impractical for directly processing vast chemical spaces containing a large number of compounds. This limitation highlights the need for a computational tool that balance speed and accuracy to rank and filter potential drug candidates efficiently. In contrast, existing ligand-based drug design methods, which learn representations from diverse protein and molecule features, often fail to make consistent predictions on unseen data or external databases, limiting their applicability for ranking and filtering potential drug candidates accurately. To address these challenges, we propose EnsDTI, a novel framework that bridges the gap between structure-based and ligand-based drug design approaches. EnsDTI utilizes a mixture-of-experts architecture to enhance DTI predictions using existing deep learning models and incorporates an inductive conformal predictor to assess prediction quality with confidence scores, ensuring reliability. Experimental results on four widely used benchmark datasets show that EnsDTI consistently achieves high performance in both prediction accuracy and confidence estimation. In addition, its candidate rankings correlate well with actual docking affinities, suggesting its practical utility in drug discovery. Yijingxiu Lu, Soosung Kang, Sun Kim, Sangseon Lee |
IEEE Trans. Comput. Biol. Bioinform. | 3 |
| 2026 | Context-Aware Hierarchical Fusion for Drug Relational LearningabstractThe simultaneous use of multiple medications is a common practice in disease treatment, yet the same drug combination can lead to different effects under varying physiological, pharmacological, or genomic conditions-collectively referred to as the 'context'. Accurately predicting the outcomes of drug combinations across diverse contexts, also known as drug relational learning (DRL), is essential for improving therapeutic efficacy and safety. Despite its importance, existing methods face two major challenges: they are often tailored to specific DRL tasks, lacking generalizability, and they fail to explicitly model the influence of context on drug interactions. This limitation arises because most methods focus primarily on whole-drug compound structures, overlooking the fine-grained atomic-level interactions critical for context-aware predictions. To address these challenges, we propose a novel context-aware hierarchical fusion architecture for DRL. By formulating the problem as the label prediction of drug-drug-context triplets, our approach explicitly models the interaction between drugs by first learning their intrinsic atomic-level interactions and then incorporating context into their embeddings at the atomic level through information fusion. Experiments across diverse tasks-such as synergy prediction, polypharmacy side effect detection, and drug-drug interaction prediction-demonstrate our model's capability to effectively capture context-aware information. Importantly, our method consistently achieves robust performance in highly complex scenarios, highlighting its adaptability and utility in advancing context-aware drug relational learning. Yijingxiu Lu, Yinhua Piao, Sangseon Lee, Sun Kim |
IEEE Trans. Comput. Biol. Bioinform. | 4 |
| 2025 | CheapNet: Cross-attention on Hierarchical representations for Efficient protein-ligand binding Affinity PredictionabstractAccurately predicting protein-ligand binding affinity is a critical challenge in drug discovery, crucial for understanding drug efficacy. While existing models typically rely on atom-level interactions, they often fail to capture the complex, higher-order interactions, resulting in noise and computational inefficiency. Transitioning to modeling these interactions at the cluster level is challenging because it is difficult to determine which atoms form meaningful clusters that drive the protein-ligand interactions. To address this, we propose CheapNet, a novel interaction-based model that integrates atom-level representations with hierarchical cluster-level interactions through a cross-attention mechanism. By employing differentiable pooling of atom-level embeddings, CheapNet efficiently captures essential higher-order molecular representations crucial for accurate binding predictions. Extensive evaluations demonstrate that CheapNet not only achieves state-of-the-art performance across multiple binding affinity prediction tasks but also maintains prediction accuracy with reasonable computational efficiency. The code of CheapNet is available at https://github.com/hyukjunlim/CheapNet. Hyukjun Lim, Sun Kim, Sangseon Lee |
ICLR | 2 |
| 2025 | BounDr.E: Predicting Drug-likeness via Biomedical Knowledge Alignment and EM-like One-Class Boundary OptimizationabstractThe advent of generative AI now enables large-scale $\textit{de novo}$ design of molecules, but identifying viable drug candidates among them remains an open problem. Existing drug-likeness prediction methods often rely on ambiguous negative sets or purely structural features, limiting their ability to accurately classify drugs from non-drugs. In this work, we introduce BounDr.E: a novel modeling of drug-likeness as a compact space surrounding approved drugs through a dynamic one-class boundary approach. Specifically, we enrich the chemical space through biomedical knowledge alignment, and then iteratively tighten the drug-like boundary by pushing non-drug-like compounds outside via an Expectation-Maximization (EM)-like process. Empirically, BounDr.E achieves 10% F1-score improvement over the previous state-of-the-art and demonstrates robust cross-dataset performance, including zero-shot toxic compound filtering. Additionally, we showcase its effectiveness through comprehensive case studies in large-scale $\textit{in silico}$ screening. Our codes and constructed benchmark data under various schemes are provided at: https://github.com/eugenebang/boundr_e. Dongmin Bang, Inyoung Sung, Yinhua Piao, Sangseon Lee, Sun Kim |
ICML | 5 |
| 2025 | CombiMOTS: Combinatorial Multi-Objective Tree Search for Dual-Target Molecule GenerationabstractDual-target molecule generation, which focuses on discovering compounds capable of interacting with two target proteins, has garnered significant attention due to its potential for improving therapeutic efficiency, safety and resistance mitigation. Existing approaches face two critical challenges. First, by simplifying the complex dual-target optimization problem to scalarized combinations of individual objectives, they fail to capture important trade-offs between target engagement and molecular properties. Second, they typically do not integrate synthetic planning into the generative process. This highlights a need for more appropriate objective function design and synthesis-aware methodologies tailored to the dual-target molecule generation task. In this work, we propose CombiMOTS, a Pareto Monte Carlo Tree Search (PMCTS) framework that generates dual-target molecules. CombiMOTS is designed to explore a synthesizable fragment space while employing vectorized optimization constraints to encapsulate target affinity and physicochemical properties. Extensive experiments on real-world databases demonstrate that CombiMOTS produces novel dual-target molecules with high docking scores, enhanced diversity, and balanced pharmacological characteristics, showcasing its potential as a powerful tool for dual-target drug discovery. The code and data is accessible through https://github.com/Tibogoss/CombiMOTS. Thibaud Southiratn, Bonil Koo, Yijingxiu Lu, Sun Kim |
ICML | 4 |
| 2025 | Transcriptome Transformer: improving patient survival prediction via multitask learning of transcriptomic and clinical featuresabstractAccurate survival prediction is essential in healthcare as it guides treatment strategies and improves patient outcomes. While clinical features provide valuable prognostic information, they often fail to represent the molecular complexity of diseases. Transcriptomic data, which reflects gene expression patterns of tumors, present a complementary perspective to address this limitation. We introduce Transcriptome Transformer (TxT), a multitask learning framework that uses a transcriptome-centric approach to improve patient survival prediction. TxT employs a Transformer-based architecture with multihead attention mechanisms to effectively capture complex dependencies among genes, enabling dynamic modeling of gene-gene interactions while using shared information across multiple clinical prediction tasks. By jointly analyzing transcriptomic data and incorporating clinical features, TxT offers a more complete representation of patient biology. In experiments across both single-task and multitask datasets, TxT outperformed existing methods in survival prediction and related clinical tasks. Additionally, TxT offers biological insights through attention-derived gene interaction networks, identifying immune-related pathways in longer-surviving Luminal A patients and coagulation and epithelial-mesenchymal transition pathways in shorter-surviving counterparts. Differential attention analysis further revealed that integrating clinical features enhances the model's ability to prioritize genes involved in biologically meaningful pathways that are known to influence tumor progression and distant recurrence. The source code of TxT is available at https://github.com/BonilKoo/TxT. Bonil Koo, Inyoung Sung, Sangseon Lee, Sun Kim |
Briefings Bioinform. | 4 |
| 2025 | ADME-drug-likeness: enriching molecular foundation models via pharmacokinetics-guided multi-task learning for drug-likeness predictionabstractSUMMARY: Recent breakthroughs in AI-driven generative models enable the rapid design of extensive molecular libraries, creating an urgent need for fast and accurate drug-likeness evaluation. Traditional approaches, however, rely heavily on structural descriptors and overlook pharmacokinetic (PK) factors such as absorption, distribution, metabolism, and excretion (ADME). Furthermore, existing deep-learning models neglect the complex interdependencies among ADME tasks, which play a pivotal role in determining clinical viability. We introduce ADME-DL (drug likeness), a novel two-step pipeline that first enhances diverse range of molecular foundation models (MFMs) via sequential ADME multi-task learning. By enforcing an A→D→M→E flow-grounded in a data-driven task dependency analysis that aligns with established PK principles-our method more accurately encodes PK information into the learned embedding space. In Step 2, the resulting ADME-informed embeddings are leveraged for drug-likeness classification, distinguishing approved drugs from negative sets drawn from chemical libraries. Through comprehensive experiments, our sequential ADME multi-task learning achieves up to +2.4% improvement over state-of-the-art baselines, and enhancing performance across tested MFMs by up to +18.2%. Case studies with clinically annotated drugs validate that respecting the PK hierarchy produces more relevant predictions, reflecting drug discovery phases. These findings underscore the potential of ADME-DL to significantly enhance the early-stage filtering of candidate molecules, bridging the gap between purely structural screening methods and PK-aware modeling. AVAILABILITY AND IMPLEMENTATION: The source code for ADME-DL is available at https://github.com/eugenebang/ADME-DL. Dongmin Bang, Haerin Song, Sun Kim |
Bioinform. | 4 |
| 2025 | MixingDTA: improved drug-target affinity prediction by extending mixup with guilt-by-associationabstractSUMMARY: Drug-target affinity (DTA) prediction is an important regression task for drug discovery, which can provide richer information than traditional drug-target interaction prediction as a binary prediction task. To achieve accurate DTA prediction, quite large amount of data are required for each drug, which is not available as of now. Thus, data scarcity and sparsity is a major challenge. Another important task is "cold-start" DTA prediction for unseen drug or protein. In this work, we introduce MixingDTA, a novel framework to tackle data scarcity by incorporating domain-specific pretrained language models for molecules and proteins with our MEETA (MolFormer and ESM-based Efficient aggregation Transformer for Affinity) model. We further address the label sparsity and cold-start challenges through a novel data augmentation strategy named GBA-Mixup, which interpolates embeddings of neighboring entities based on the guilt-by-association (GBA) principle, to improve prediction accuracy even in sparse regions of DTA space. Our experiments on benchmark datasets demonstrate that the MEETA backbone alone provides up to a 19% improvement of mean squared error over current state-of-the-art baseline, and the addition of GBA-Mixup contributes a further 8.4% improvement. Importantly, GBA-Mixup is model-agnostic, delivering performance gains across all tested backbone models of up to 16.9%. Case studies shows how MixingDTA interpolates between drugs and targets in the embedding space, demonstrating generalizability for unseen drug-target pairs while effectively focusing on functionally critical residues. These results highlight MixingDTA's potential to accelerate drug discovery by offering accurate, scalable, and biologically informed DTA predictions. AVAILABILITY AND IMPLEMENTATION: The code for MixingDTA is available at https://github.com/rokieplayer20/MixingDTA. Youngoh Kim, Dongmin Bang, Bonil Koo, Jungseob Yi, Changyun Cho, Jeonguk Choi, Sun Kim |
Bioinform. | 7 |
| 2025 | Dual Representation Learning for Predicting Drug-Side Effect Frequency Using Protein Target InformationabstractKnowledge of unintended effects of drugs is critical in assessing the risk of treatment and in drug repurposing. Although numerous existing studies predict drug-side effect presence, only four of them predict the frequency of the side effects. Unfortunately, current prediction methods 1) do not utilize drug targets, 2) do not predict well for unseen drugs, and 3) do not use multiple heterogeneous drug features. We propose a novel deep learning-based drug-side effect frequency prediction model. Our model utilized heterogeneous features such as target protein information as well as molecular graph, fingerprints, and chemical similarity to create drug embeddings simultaneously. Furthermore, the model represents drugs and side effects into a common vector space, learning the dual representation vectors of drugs and side effects, respectively. We also extended the predictive power of our model to compensate for the drugs without clear target proteins using the Adaboost method. We achieved state-of-the-art performance over the existing methods in predicting side effect frequencies, especially for unseen drugs. Ablation studies show that our model effectively combines and utilizes heterogeneous features of drugs. Moreover, we observed that, when the target information given, drugs with explicit targets resulted in better prediction than the drugs without explicit targets. Sangseon Lee, Minwoo Pak, Sun Kim |
IEEE J. Biomed. Health Informatics | 4 |
| 2024 | DiSCO: Diffusion Schrödinger Bridge for Molecular Conformer OptimizationabstractThe generation of energetically optimal 3D molecular conformers is crucial in cheminformatics and drug discovery. While deep generative models have been utilized for direct generation in Euclidean space, this approach encounters challenges, including the complexity of navigating a vast search space. Recent generative models that implement simplifications to circumvent these challenges have achieved state-of-the-art results, but this simplified approach unavoidably creates a gap between the generated conformers and the ground-truth conformational landscape. To bridge this gap, we introduce DiSCO: Diffusion Schrödinger Bridge for Molecular Conformer Optimization, a novel diffusion framework that enables direct learning of nonlinear diffusion processes in prior-constrained Euclidean space for the optimization of 3D molecular conformers. Through the incorporation of an SE(3)-equivariant Schrödinger bridge, we establish the roto-translational equivariance of the generated conformers. Our framework is model-agnostic and offers an easily implementable solution for the post hoc optimization of conformers produced by any generation method. Through comprehensive evaluations and analyses, we establish the strengths of our framework, substantiating the application of the Schrödinger bridge for molecular conformer optimization. First, our approach consistently outperforms four baseline approaches, producing conformers with higher diversity and improved quality. Then, we show that the intermediate conformers generated during our diffusion process exhibit valid and chemically meaningful characteristics. We also demonstrate the robustness of our method when starting from conformers of diverse quality, including those unseen during training. Lastly, we show that the precise generation of low-energy conformers via our framework helps in enhancing the downstream prediction of molecular properties. The code is available at https://github.com/Danyeong-Lee/DiSCO. Danyeong Lee, Dohoon Lee, Dongmin Bang, Sun Kim |
AAAI | 4 |
| 2024 | Improving Out-of-Distribution Generalization in Graphs via Hierarchical Semantic EnvironmentsabstractOut-of-distribution (OOD) generalization in the graph domain is challenging due to complex distribution shifts and a lack of environmental contexts. Recent methods attempt to enhance graph OOD generalization by generating flat environments. However, such flat environments come with inherent limitations to capture more complex data distributions. Considering the DrugOOD dataset, which contains diverse training environments (e.g., scaffold, size, etc.), flat contexts cannot sufficiently address its high heterogeneity. Thus, a new challenge is posed to generate more seman-tically enriched environments to enhance graph invariant learning for handling distribution shifts. In this paper, we propose a novel approach to generate hierarchical seman-tic environments for each graph. Firstly, given an input graph, we explicitly extract variant subgraphs from the in-put graph to generate proxy predictions on local environ-ments. Then, stochastic attention mechanisms are employed to re-extract the subgraphs for regenerating global environ-ments in a hierarchical manner. In addition, we introduce a new learning objective that guides our model to learn the diversity of environments within the same hierarchy while maintaining consistency across different hierarchies. This approach enables our model to consider the relationships between environments and facilitates robust graph invariant learning. Extensive experiments on real-world graph data have demonstrated the effectiveness of our framework. Par-ticularly, in the challenging dataset DrugOOD, our method achieves up to 1.29% and 2.83% improvement over the best baselines on IC50 and EC50 prediction tasks, respectively. Yinhua Piao, Sangseon Lee, Yijingxiu Lu, Sun Kim |
CVPR | 4 |
| 2024 | Transfer learning of condition-specific perturbation in gene interactions improves drug response predictionabstractSUMMARY: Drug response is conventionally measured at the cell level, often quantified by metrics like IC50. However, to gain a deeper understanding of drug response, cellular outcomes need to be understood in terms of pathway perturbation. This perspective leads us to recognize a challenge posed by the gap between two widely used large-scale databases, LINCS L1000 and GDSC, measuring drug response at different levels-L1000 captures information at the gene expression level, while GDSC operates at the cell line level. Our study aims to bridge this gap by integrating the two databases through transfer learning, focusing on condition-specific perturbations in gene interactions from L1000 to interpret drug response integrating both gene and cell levels in GDSC. This transfer learning strategy involves pretraining on the transcriptomic-level L1000 dataset, with parameter-frozen fine-tuning to cell line-level drug response. Our novel condition-specific gene-gene attention (CSG2A) mechanism dynamically learns gene interactions specific to input conditions, guided by both data and biological network priors. The CSG2A network, equipped with transfer learning strategy, achieves state-of-the-art performance in cell line-level drug response prediction. In two case studies, well-known mechanisms of drugs are well represented in both the learned gene-gene attention and the predicted transcriptomic profiles. This alignment supports the modeling power in terms of interpretability and biological relevance. Furthermore, our model's unique capacity to capture drug response in terms of both pathway perturbation and cell viability extends predictions to the patient level using TCGA data, demonstrating its expressive power obtained from both gene and cell levels. AVAILABILITY AND IMPLEMENTATION: The source code for the CSG2A network is available at https://github.com/eugenebang/CSG2A. Dongmin Bang, Bonil Koo, Sun Kim |
Bioinform. | 3 |
| 2024 | PONYTA: prioritization of phenotype-related genes from mouse KO events using PU learning on a biological networkabstractMOTIVATION: Transcriptome data from gene knock-out (KO) experiments in mice provide crucial insights into the intricate interactions between genotype and phenotype. Differentially expressed gene (DEG) analysis and network propagation (NP) are well-established methods for analysing transcriptome data. To determine genes related to phenotype changes from a KO experiment, we need to choose a cutoff value for the corresponding criterion based on the specific method. Using a rigorous cutoff value for DEG analysis and NP is likely to select mostly positive genes related to the phenotype, but many will be rejected as false negatives. On the other hand, using a loose cutoff value for either method is prone to include a number of genes that are not phenotype-related, which are false positives. Thus, the research problem at hand is how to deal with the trade-off between false negatives and false positives. RESULTS: We propose a novel framework called PONYTA for gene prioritization via positive-unlabeled (PU) learning on biological networks. Beginning with the selection of true phenotype-related genes using a rigorous cutoff value for DEG analysis and NP, we address the issue of handling false negatives by rescuing them through PU learning. Evaluations on transcriptome data from multiple studies show that our approach has superior gene prioritization ability compared to benchmark models. Therefore, PONYTA effectively prioritizes genes related to phenotypes derived from gene KO events and guides in vitro and in vivo gene KO experiments for increased efficiency. AVAILABILITY AND IMPLEMENTATION: The source code of PONYTA is available at https://github.com/Jun-Hyeong-Kim/PONYTA. Jun Hyeong Kim, Bonil Koo, Sun Kim |
Bioinform. | 3 |
| 2023 | Clinical Note Owns its Hierarchy: Multi-Level Hypergraph Neural Networks for Patient-Level Representation LearningabstractLeveraging knowledge from electronic health records (EHRs) to predict a patient's condition is essential to the effective delivery of appropriate care.Clinical notes of patient EHRs contain valuable information from healthcare professionals, but have been underused due to their difficult contents and complex hierarchies.Recently, hypergraph-based methods have been proposed for document classifications.Directly adopting existing hypergraph methods on clinical notes cannot sufficiently utilize the hierarchy information of the patient, which can degrade clinical semantic information by (1) frequent neutral words and (2) hierarchies with imbalanced distribution.Thus, we propose a taxonomy-aware multi-level hypergraph neural network (TM-HGNN), where multi-level hypergraphs assemble useful neutral words with rare keywords via note and taxonomy level hyperedges to retain the clinical semantic information.The constructed patient hypergraphs are fed into hierarchical message passing layers for learning more balanced multi-level knowledge at the note and taxonomy levels.We validate the effectiveness of TM-HGNN by conducting extensive experiments with MIMIC-III dataset on benchmark in-hospital-mortality prediction. 1 * These authors contributed equally to this work.1 Our codes and models are publicly Nayeon Kim 0008, Yinhua Piao, Sun Kim |
ACL (1) | 3 |
| 2023 | A model-agnostic framework to enhance knowledge graph-based drug combination prediction with drug-drug interaction data and supervised contrastive learningabstractCombination therapies have brought significant advancements to the treatment of various diseases in the medical field. However, searching for effective drug combinations remains a major challenge due to the vast number of possible combinations. Biomedical knowledge graph (KG)-based methods have shown potential in predicting effective combinations for wide spectrum of diseases, but the lack of credible negative samples has limited the prediction performance of machine learning models. To address this issue, we propose a novel model-agnostic framework that leverages existing drug-drug interaction (DDI) data as a reliable negative dataset and employs supervised contrastive learning (SCL) to transform drug embedding vectors to be more suitable for drug combination prediction. We conducted extensive experiments using various network embedding algorithms, including random walk and graph neural networks, on a biomedical KG. Our framework significantly improved performance metrics compared to the baseline framework. We also provide embedding space visualizations and case studies that demonstrate the effectiveness of our approach. This work highlights the potential of using DDI data and SCL in finding tighter decision boundaries for predicting effective drug combinations. Jeonghyeon Gu, Dongmin Bang, Jungseob Yi, Sangseon Lee, Dong Kyu Kim, Sun Kim |
Briefings Bioinform. | 6 |
| 2023 | Improved drug response prediction by drug target data integration via network-based profilingabstractDrug response prediction (DRP) is important for precision medicine to predict how a patient would react to a drug before administration. Existing studies take the cell line transcriptome data, and the chemical structure of drugs as input and predict drug response as IC50 or AUC values. Intuitively, use of drug target interaction (DTI) information can be useful for DRP. However, use of DTI is difficult because existing drug response database such as CCLE and GDSC do not have information about transcriptome after drug treatment. Although transcriptome after drug treatment is not available, if we can compute the perturbation effects by the pharmacologic modulation of target gene, we can utilize the DTI information in CCLE and GDSC. In this study, we proposed a framework that can improve existing deep learning-based DRP models by effectively utilizing drug target information. Our framework includes NetGP, a module to compute gene perturbation scores by the network propagation technique on a network. NetGP produces genes in a ranked list in terms of gene perturbation scores and the ranked genes are input to a multi-layer perceptron to generate a fixed dimension vector for the integration with existing DRP models. This integration is done in a model-agnostic way so that any existing DRP tool can be incorporated. As a result, our framework boosts the performance of existing DRP models, in 64 of 72 comparisons. The performance gains are larger especially for test scenarios with samples with unseen drugs by large margins up to 34% in Pearson's correlation coefficient. Minwoo Pak, Sangseon Lee, Inyoung Sung, Bonil Koo, Sun Kim |
Briefings Bioinform. | 5 |
| 2023 | GOAT: Gene-level biomarker discovery from multi-Omics data using graph ATtention neural network for eosinophilic asthma subtypeabstractMOTIVATION: Asthma is a heterogeneous disease where various subtypes are established and molecular biomarkers of the subtypes are yet to be discovered. Recent availability of multi-omics data paved a way to discover molecular biomarkers for the subtypes. However, multi-omics biomarker discovery is challenging because of the complex interplay between different omics layers. RESULTS: We propose a deep attention model named Gene-level biomarker discovery from multi-Omics data using graph ATtention neural network (GOAT) for identifying molecular biomarkers for eosinophilic asthma subtypes with multi-omics data. GOAT identifies genes that discriminate subtypes using a graph neural network by modeling complex interactions among genes as the attention mechanism in the deep learning model. In experiments with multi-omics profiles of the COREA (Cohort for Reality and Evolution of Adult Asthma in Korea) asthma cohort of 300 patients, GOAT outperforms existing models and suggests interpretable biological mechanisms underlying asthma subtypes. Importantly, GOAT identified genes that are distinct only in terms of relationship with other genes through attention. To better understand the role of biomarkers, we further investigated two transcription factors, CTNNB1 and JUN, captured by GOAT. We were successful in showing the role of the transcription factors in eosinophilic asthma pathophysiology in a network propagation and transcriptional network analysis, which were not distinct in terms of gene expression level differences. AVAILABILITY AND IMPLEMENTATION: Source code is available https://github.com/DabinJeong/Multi-omics_biomarker. The preprocessed data underlying this article is accessible in data folder of the github repository. Raw data are available in Multi-Omics Platform at http://203.252.206.90:5566/, and it can be accessible when requested. Dabin Jeong, Bonil Koo, Minsik Oh, Tae-Bum Kim, Sun Kim |
Bioinform. | 5 |
| 2023 | Metheor: Ultrafast DNA methylation heterogeneity calculation from bisulfite read alignmentsabstractPhased DNA methylation states within bisulfite sequencing reads are valuable source of information that can be used to estimate epigenetic diversity across cells as well as epigenomic instability in individual cells. Various measures capturing the heterogeneity of DNA methylation states have been proposed for a decade. However, in routine analyses on DNA methylation, this heterogeneity is often ignored by computing average methylation levels at CpG sites, even though such information exists in bisulfite sequencing data in the form of phased methylation states, or methylation patterns. In this study, to facilitate the application of the DNA methylation heterogeneity measures in downstream epigenomic analyses, we present a Rust-based, extremely fast and lightweight bioinformatics toolkit called Metheor. As the analysis of DNA methylation heterogeneity requires the examination of pairs or groups of CpGs throughout the genome, existing softwares suffer from high computational burden, which almost make a large-scale DNA methylation heterogeneity studies intractable for researchers with limited resources. In this study, we benchmark the performance of Metheor against existing code implementations for DNA methylation heterogeneity measures in three different scenarios of simulated bisulfite sequencing datasets. Metheor was shown to dramatically reduce the execution time up to 300-fold and memory footprint up to 60-fold, while producing identical results with the original implementation, thereby facilitating a large-scale study of DNA methylation heterogeneity profiles. To demonstrate the utility of the low computational burden of Metheor, we show that the methylation heterogeneity profiles of 928 cancer cell lines can be computed with standard computing resources. With those profiles, we reveal the association between DNA methylation heterogeneity and various omics features. Source code for Metheor is at https://github.com/dohlee/metheor and is freely available under the GPL-3.0 license. Dohoon Lee, Bonil Koo, Jeewon Yang, Sun Kim |
PLoS Comput. Biol. | 4 |
| 2022 | Sparse Structure Learning via Graph Neural Networks for Inductive Document ClassificationabstractRecently, graph neural networks (GNNs) have been widely used for document classification. However, most existing methods are based on static word co-occurrence graphs without sentence-level information, which poses three challenges:(1) word ambiguity, (2) word synonymity, and (3) dynamic contextual dependency. To address these challenges, we propose a novel GNN-based sparse structure learning model for inductive document classification. Specifically, a document-level graph is initially generated by a disjoint union of sentence-level word co-occurrence graphs. Our model collects a set of trainable edges connecting disjoint words between sentences, and employs structure learning to sparsely select edges with dynamic contextual dependencies. Graphs with sparse structure can jointly exploit local and global contextual information in documents through GNNs. For inductive learning, the refined document graph is further fed into a general readout function for graph-level classification and optimization in an end-to-end manner. Extensive experiments on several real-world datasets demonstrate that the proposed model outperforms most state-of-the-art results, and reveal the necessity to learn sparse structures for each document. Yinhua Piao, Sangseon Lee, Dohoon Lee, Sun Kim |
AAAI | 4 |
| 2022 | AutoCoV: tracking the early spread of COVID-19 in terms of the spatial and temporal patterns from embedding space by K-mer based deep learningabstractBACKGROUND: The widely spreading coronavirus disease (COVID-19) has three major spreading properties: pathogenic mutations, spatial, and temporal propagation patterns. We know the spread of the virus geographically and temporally in terms of statistics, i.e., the number of patients. However, we are yet to understand the spread at the level of individual patients. As of March 2021, COVID-19 is wide-spread all over the world with new genetic variants. One important question is to track the early spreading patterns of COVID-19 until the virus has got spread all over the world. RESULTS: In this work, we proposed AutoCoV, a deep learning method with multiple loss object, that can track the early spread of COVID-19 in terms of spatial and temporal patterns until the disease is fully spread over the world in July 2020. Performances in learning spatial or temporal patterns were measured with two clustering measures and one classification measure. For annotated SARS-CoV-2 sequences from the National Center for Biotechnology Information (NCBI), AutoCoV outperformed seven baseline methods in our experiments for learning either spatial or temporal patterns. For spatial patterns, AutoCoV had at least 1.7-fold higher clustering performances and an F1 score of 88.1%. For temporal patterns, AutoCoV had at least 1.6-fold higher clustering performances and an F1 score of 76.1%. Furthermore, AutoCoV demonstrated the robustness of the embedding space with an independent dataset, Global Initiative for Sharing All Influenza Data (GISAID). CONCLUSIONS: In summary, AutoCoV learns geographic and temporal spreading patterns successfully in experiments on NCBI and GISAID datasets and is the first of its kind that learns virus spreading patterns from the genome sequences, to the best of our knowledge. We expect that this type of embedding method will be helpful in characterizing fast-evolving pandemics. Inyoung Sung, Sangseon Lee, Minwoo Pak, Yunyol Shin, Sun Kim |
BMC Bioinform. | 5 |
| 2022 | MLDEG: A Machine Learning Approach to Identify Differentially Expressed Genes Using Network Property and Network PropagationabstractMOTIVATION: Identifying differentially expressed genes (DEGs) in transcriptome data is a very important task. However, performances of existing DEG methods vary significantly for data sets measured in different conditions and no single statistical or machine learning model for DEG detection perform consistently well for data sets of different traits. In addition, setting a cutoff value for the significance of differential expressions is one of confounding factors to determine DEGs. RESULTS: We address these problems by developing an ensemble model that refines the heterogeneous and inconsistent results of the existing methods by taking accounts into network information such as network propagation and network property. DEG candidates that are predicted with weak evidence by the existing tools are re-classified by our proposed ensemble model for the transcriptome data. Tested on 10 RNA-seq datasets downloaded from gene expression omnibus (GEO), our method showed excellent performance of winning the first place in detecting ground truth (GT) genes in eight datasets and find almost all GT genes in six datasets. On the other hand, performances of all existing methods varied significantly for the 10 data sets. Because of the design principle, our method can accommodate any new DEG methods naturally. AVAILABILITY: The source code of our method is available at https://github.com/jihmoon/MLDEG. Ji Hwan Moon, Sangseon Lee, Minwoo Pak, Benjamin Hur, Sun Kim |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2021 | TeamTat: A Collaborative Text Annotation Tool for Creating Gold-Standard Corpora
Rezarta Islamaj Dogan, Dongseop Kwon, Sun Kim, Zhiyong Lu |
AMIA | 3 |
| 2021 | A probabilistic model for pathway-guided gene set selectionabstractBreast cancer is classified into five intrinsic subtypes, with differing treatment methods and prognoses. Therefore, accurate identification of subtypes from patient transcriptome data is essential. Many gene signatures, including PAM50, have been developed to classify breast cancer subtypes. However, existing gene selection methods do not utilize biological pathways. Gene signature selection using biological pathways can explain signature genes in terms of biological functions. Thus, we propose a probabilistic model for pathway-guided gene set selection using gene expression data. First, we defined gene and pathway factors based on gene expression and pathway activation levels, and calculated the posterior probability. Second, we adopted the prediction strength to guide gene set selection. Third, the gene set was selected using the posterior probability and prediction strength values. Finally, on evaluating the selected gene set, it was experimentally confirmed that our gene set performed better on classification tasks than the PAM50 gene set, a gene set produced by the XGBoost classifier, and a random gene set. Among the genes selected by our method, it was confirmed that the genes included in the cell cycle and circadian rhythm pathways showed different expression patterns for each breast cancer subtype. Our selected gene set exhibited biological significance in terms of pathway activation. Inyoung Kim, Sangseon Lee, Hugh Namkoong, Sun Kim |
BIBM | 5 |
| 2021 | IDEA: Integrating Divisive and Ensemble-Agglomerate hierarchical clustering framework for arbitrary shape dataabstractHierarchical clustering, a traditional clustering method, has been getting attention again. Among several reasons, a credit goes to a recent paper by Dasgupta in 2016 that proposed a cost function that quantitatively evaluates hierarchical clustering trees. An important question is how to combine this recent advance with existing successful clustering methods. In this paper, we propose a hierarchical clustering method to minimize the cost function of clustering tree by incorporating existing clustering techniques. First, we developed an ensemble tree-search method that finds an integrated tree with reduced cost by integrating multiple existing hierarchical clustering methods. Second, to operate on large and arbitrary shape data, we designed an efficient hierarchical clustering framework, called integrating divisive and ensemble-agglomerate (IDEA) by combining it with advanced clustering techniques such as nearest neighbor graph construction, divisive-agglomerate hybridization, and dynamic cut tree. The IDEA clustering method showed better performance in minimizing Dasgupta's cost and improving accuracy (adjusted rand index) over existing cost-minimization-based, and density-based hierarchical clustering methods in experiments using arbitrary shape datasets and complex biology-domain datasets. Hongryul Ahn, Inuk Jung, Heejoon Chae, Minsik Oh, Inyoung Kim, Sun Kim |
IEEE BigData | 6 |
| 2021 | Towards multi-omics characterization of tumor heterogeneity: a comprehensive review of statistical and machine learning approachesabstractThe multi-omics molecular characterization of cancer opened a new horizon for our understanding of cancer biology and therapeutic strategies. However, a tumor biopsy comprises diverse types of cells limited not only to cancerous cells but also to tumor microenvironmental cells and adjacent normal cells. This heterogeneity is a major confounding factor that hampers a robust and reproducible bioinformatic analysis for biomarker identification using multi-omics profiles. Besides, the heterogeneity itself has been recognized over the years for its significant prognostic values in some cancer types, thus offering another promising avenue for therapeutic intervention. A number of computational approaches to unravel such heterogeneity from high-throughput molecular profiles of a tumor sample have been proposed, but most of them rely on the data from an individual omics layer. Since the heterogeneity of cells is widely distributed across multi-omics layers, methods based on an individual layer can only partially characterize the heterogeneous admixture of cells. To help facilitate further development of the methodologies that synchronously account for several multi-omics profiles, we wrote a comprehensive review of diverse approaches to characterize tumor heterogeneity based on three different omics layers: genome, epigenome and transcriptome. As a result, this review can be useful for the analysis of multi-omics profiles produced by many large-scale consortia. Contact:[email protected]. Dohoon Lee, Youngjune Park, Sun Kim |
Briefings Bioinform. | 3 |
| 2021 | Machine learning-based analysis of multi-omics data on the cloud for investigating gene regulationsabstractGene expressions are subtly regulated by quantifiable measures of genetic molecules such as interaction with other genes, methylation, mutations, transcription factor and histone modifications. Integrative analysis of multi-omics data can help scientists understand the condition or patient-specific gene regulation mechanisms. However, analysis of multi-omics data is challenging since it requires not only the analysis of multiple omics data sets but also mining complex relations among different genetic molecules by using state-of-the-art machine learning methods. In addition, analysis of multi-omics data needs quite large computing infrastructure. Moreover, interpretation of the analysis results requires collaboration among many scientists, often requiring reperforming analysis from different perspectives. Many of the aforementioned technical issues can be nicely handled when machine learning tools are deployed on the cloud. In this survey article, we first survey machine learning methods that can be used for gene regulation study, and we categorize them according to five different goals: gene regulatory subnetwork discovery, disease subtype analysis, survival analysis, clinical prediction and visualization. We also summarize the methods in terms of multi-omics input types. Then, we explain why the cloud is potentially a good solution for the analysis of multi-omics data, followed by a survey of two state-of-the-art cloud systems, Galaxy and BioVLAB. Finally, we discuss important issues when the cloud is used for the analysis of multi-omics data for the gene regulation study. Minsik Oh, Sun Kim, Heejoon Chae |
Briefings Bioinform. | 3 |
| 2021 | Erratum to: Machine learning-based analysis of multi-omics data on the cloud for investigating gene regulationsabstractThe first version of this article neglected to identify Minsik Oh and Sungjoon Park as joint first authors. This has now been corrected. Minsik Oh, Sun Kim, Heejoon Chae |
Briefings Bioinform. | 3 |
| 2021 | mirTime: identifying condition-specific targets of microRNA in time-series transcript data using Gaussian process model and spherical vector clusteringabstractBACKGROUND: MicroRNAs, small noncoding RNAs, are conserved in many species, and they are key regulators that mediate post-transcriptional gene silencing. Since biologists cannot perform experiments for each of target genes of thousands of microRNAs in numerous specific conditions, prediction on microRNA target genes has been extensively investigated. A general framework is a two-step process of selecting target candidates based on sequence and binding energy features and then predicting targets based on negative correlation of microRNAs and their targets. However, there are few methods that are designed for target predictions using time-series gene expression data. RESULTS: In this article, we propose a new pipeline, mirTime, that predicts microRNA targets by integrating sequence features and time-series expression profiles in a specific experimental condition. The most important feature of mirTime is that it uses the Gaussian process regression model to measure data at unobserved or unpaired time points. In experiments with two datasets in different experimental conditions and cell types, condition-specific target modules reported in the original papers were successfully predicted with our pipeline. The context specificity of target modules was assessed with three (correlation-based, target gene-based and network-based) evaluation criteria. mirTime showed better performance than existing expression-based microRNA target prediction methods in all three criteria. AVAILABILITY AND IMPLEMENTATION: mirTime is available at https://github.com/mirTime/mirtime. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Hyejin Kang, Hongryul Ahn, Kyuri Jo, Minsik Oh, Sun Kim |
Bioinform. | 5 |
| 2021 | BioVLAB-Cancer-Pharmacogenomics: tumor heterogeneity and pharmacogenomics analysis of multi-omics data from tumor on the cloudabstractMOTIVATION: Multi-omics data in molecular biology has accumulated rapidly over the years. Such data contains valuable information for research in medicine and drug discovery. Unfortunately, data-driven research in medicine and drug discovery is challenging for a majority of small research labs due to the large volume of data and the complexity of analysis pipeline. RESULTS: We present BioVLAB-Cancer-Pharmacogenomics, a bioinformatics system that facilitates analysis of multi-omics data from breast cancer to analyze and investigate intratumor heterogeneity and pharmacogenomics on Amazon Web Services. Our system takes multi-omics data as input to perform tumor heterogeneity analysis in terms of TCGA data and deconvolve-and-match the tumor gene expression to cell line data in CCLE using DNA methylation profiles. We believe that our system can help small research labs perform analysis of tumor multi-omics without worrying about computational infrastructure and maintenance of databases and tools. AVAILABILITY AND IMPLEMENTATION: http://biohealth.snu.ac.kr/software/biovlab_cancer_pharmacogenomics. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Dohoon Lee, Sangsoo Lim, Heejoon Chae, Sun Kim |
Bioinform. | 6 |
| 2021 | Ranked k-Spectrum Kernel for Comparative and Evolutionary Comparison of Exons, Introns, and CpG IslandsabstractMOTIVATION: Existing k-mer based string kernel methods have been successfully used for sequence comparison. However, existing kernel methods have limitations for comparative and evolutionary comparisons of genomes due to the sensitiveness to over-represented k-mers and variable sequence lengths. RESULTS: In this study, we propose a novel ranked k-spectrum string (RKSS) kernel. 1) RKSS kernel utilizes common k-mer sets across species, named landmarks, that can be used for comparing multiple genomes. 2) Based on the landmarks, we can use ranks of k-mers, rather than frequencies, that can produce more robust distances between genomes. To show the power of RKSS kernel, we conducted two experiments using 10 mammalian species with exon, intron, and CpG island sequences. RKSS kernel reconstructed more consistent evolutionary trees than the k-spectrum string kernel. In the subsequent experiment, for each sequence, kernel distance was calculated from 30 landmarks representing exon, intron, and CpG island sequences of 10 genomes. Based on kernel distances, concordance tests were performed and the result suggested that more information is conserved in CpG islands across species than in introns. In conclusion, our analysis suggests that the relational order, exon CpG island intron, in terms of evolutionary information contents. Sangseon Lee, Taeheon Lee, Yung-Kyun Noh, Sun Kim |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2020 | Homomorphic Computation of Local AlignmentabstractIn this paper we present a homomorphic computation algorithm that finds an optimal local alignment of a pair of encrypted sequences, based on the Smith-Waterman recurrence. Our algorithm also includes an efficient method of finding the location of an optimal local alignment, in order to avoid costly conditional branching in the backtracking. To reduce computation time, we use two level parallel computations: one for filling the entries of the dynamic programming recurrence, and the other for implementing the circuits. To the best of our knowledge, this is the first attempt to compute an optimal local alignment of homomorphically encrypted sequences. With the efficient location retrieval, parallel computations, and a proper HE scheme, our implementation shows good performances in the experiment so as to be useful in practice. Magsarjav Bataa, Siwoo Song, Kunsoo Park, Miran Kim, Jung Hee Cheon, Sun Kim |
BIBM | 6 |
| 2020 | Comprehensive and critical evaluation of individualized pathway activity measurement tools on pan-cancer dataabstractMOTIVATION: Biological pathways are extensively used for the analysis of transcriptome data to characterize biological mechanisms underlying various phenotypes. There are a number of computational tools that summarize transcriptome data at the pathway level. However, there is no comparative study on how well these tools produce useful information at the cohort level, enabling comparison of many samples or patients. RESULTS: In this study, we systematically compared and evaluated 13 different pathway activity inference tools based on 5 comparison criteria using pan-cancer data set. This study has two major contributions. First, our study provides a comprehensive survey on computational techniques used by existing pathway activity inference tools. The tools use different strategies and assume different requirements on data: input transformation, use of labels, necessity of cohort-level input data, use of gene relations and scoring metric. Second, we performed extensive evaluations on the performance of these tools. Because different tools use different methods to map samples to the pathway dimension, the tools are evaluated at the pathway level using five comparison criteria. Starting from measuring how well a tool maintains the characteristics of original gene expression values, robustness was also investigated by adding noise into gene expression data. Classification tasks on three clinical variables (tumor versus normal, survival and cancer subtypes) were performed to evaluate the utility of tools for their clinical applications. In addition, the inferred activity values were compared between the tools to see how similar they are along with the scoring schemes they use. Sangsoo Lim, Sangseon Lee, Inuk Jung, SungMin Rhee, Sun Kim |
Briefings Bioinform. | 5 |
| 2020 | Cancer subtype classification and modeling by pathway attention and propagationabstractMOTIVATION: Biological pathway is an important curated knowledge of biological processes. Thus, cancer subtype classification based on pathways will be very useful to understand differences in biological mechanisms among cancer subtypes. However, pathways include only a fraction of the entire gene set, only one-third of human genes in KEGG, and pathways are fragmented. For this reason, there are few computational methods to use pathways for cancer subtype classification. RESULTS: We present an explainable deep-learning model with attention mechanism and network propagation for cancer subtype classification. Each pathway is modeled by a graph convolutional network. Then, a multi-attention-based ensemble model combines several hundreds of pathways in an explainable manner. Lastly, network propagation on pathway-gene network explains why gene expression profiles in subtypes are different. In experiments with five TCGA cancer datasets, our method achieved very good classification accuracies and, additionally, identified subtype-specific pathways and biological functions. AVAILABILITY AND IMPLEMENTATION: The source code is available at http://biohealth.snu.ac.kr/software/GCN_MAE. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Sangseon Lee, Sangsoo Lim, Taeheon Lee, Inyoung Sung, Sun Kim |
Bioinform. | 5 |
| 2020 | Better synonyms for enriching biomedical searchabstractOBJECTIVE: In a biomedical literature search, the link between a query and a document is often not established, because they use different terms to refer to the same concept. Distributional word embeddings are frequently used for detecting related words by computing the cosine similarity between them. However, previous research has not established either the best embedding methods for detecting synonyms among related word pairs or how effective such methods may be. MATERIALS AND METHODS: In this study, we first create the BioSearchSyn set, a manually annotated set of synonyms, to assess and compare 3 widely used word-embedding methods (word2vec, fastText, and GloVe) in their ability to detect synonyms among related pairs of words. We demonstrate the shortcomings of the cosine similarity score between word embeddings for this task: the same scores have very different meanings for the different methods. To address the problem, we propose utilizing pool adjacent violators (PAV), an isotonic regression algorithm, to transform a cosine similarity into a probability of 2 words being synonyms. RESULTS: Experimental results using the BioSearchSyn set as a gold standard reveal which embedding methods have the best performance in identifying synonym pairs. The BioSearchSyn set also allows converting cosine similarity scores into probabilities, which provides a uniform interpretation of the synonymy score over different methods. CONCLUSIONS: We introduced the BioSearchSyn corpus of 1000 term pairs, which allowed us to identify the best embedding method for detecting synonymy for biomedical search. Using the proposed method, we created PubTermVariants2.0: a large, automatically extracted set of synonym pairs that have augmented PubMed searches since the spring of 2019. Lana Yeganova, Sun Kim, Qingyu Chen 0001, Grigory Balasanov, W. John Wilbur, Zhiyong Lu |
J. Am. Medical Informatics Assoc. | 2 |
| 2020 | BioConceptVec: Creating and evaluating literature-based biomedical concept embeddings on a large scaleabstractA massive number of biological entities, such as genes and mutations, are mentioned in the biomedical literature. The capturing of the semantic relatedness of biological entities is vital to many biological applications, such as protein-protein interaction prediction and literature-based discovery. Concept embeddings-which involve the learning of vector representations of concepts using machine learning models-have been employed to capture the semantics of concepts. To develop concept embeddings, named-entity recognition (NER) tools are first used to identify and normalize concepts from the literature, and then different machine learning models are used to train the embeddings. Despite multiple attempts, existing biomedical concept embeddings generally suffer from suboptimal NER tools, small-scale evaluation, and limited availability. In response, we employed high-performance machine learning-based NER tools for concept recognition and trained our concept embeddings, BioConceptVec, via four different machine learning models on ~30 million PubMed abstracts. BioConceptVec covers over 400,000 biomedical concepts mentioned in the literature and is of the largest among the publicly available biomedical concept embeddings to date. To evaluate the validity and utility of BioConceptVec, we respectively performed two intrinsic evaluations (identifying related concepts based on drug-gene and gene-gene interactions) and two extrinsic evaluations (protein-protein interaction prediction and drug-drug interaction extraction), collectively using over 25 million instances from nine independent datasets (17 million instances from six intrinsic evaluation tasks and 8 million instances from three extrinsic evaluation tasks), which is, by far, the most comprehensive to our best knowledge. The intrinsic evaluation results demonstrate that BioConceptVec consistently has, by a large margin, better performance than existing concept embeddings in identifying similar and related concepts. More importantly, the extrinsic evaluation results demonstrate that using BioConceptVec with advanced deep learning models can significantly improve performance in downstream bioinformatics studies and biomedical text-mining applications. Our BioConceptVec embeddings and benchmarking datasets are publicly available at https://github.com/ncbi-nlp/BioConceptVec. Qingyu Chen 0001, Kyubum Lee, Shankai Yan, Sun Kim, Chih-Hsuan Wei, Zhiyong Lu |
PLoS Comput. Biol. | 4 |
| 2019 | CaPSSA: visual evaluation of cancer biomarker genes for patient stratification and survival analysis using mutation and expression dataabstractSUMMARY: Predictive biomarkers for patient stratification play critical roles in realizing the paradigm of precision medicine. Molecular characteristics such as somatic mutations and expression signatures represent the primary source of putative biomarker genes for patient stratification. However, evaluation of such candidate biomarkers is still cumbersome and requires multistep procedures especially when using massive public omics data. Here, we present an interactive web application that divides patients from large cohorts (e.g. The Cancer Genome Atlas, TCGA) dynamically into two groups according to the mutation, copy number variation or gene expression of query genes. It further supports users to examine the prognostic value of resulting patient groups based on survival analysis and their association with the clinical features as well as the previously annotated molecular subtypes, facilitated with a rich and interactive visualization. Importantly, we also support custom omics data with clinical information. AVAILABILITY AND IMPLEMENTATION: CaPSSA (Cancer Patient Stratification and Survival Analysis) runs on a web-browser and is freely available without restrictions at http://www.kobic.re.kr/capssa/. The source code is available on https://github.com/yjjang/capssa. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yeongjun Jang, Jihae Seo, Insu Jang, Byungwook Lee, Sun Kim, Sanghyuk Lee |
Bioinform. | 5 |
| 2019 | PRISM: methylation pattern-based, reference-free inference of subclonal makeupabstractMOTIVATION: Characterizing cancer subclones is crucial for the ultimate conquest of cancer. Thus, a number of bioinformatic tools have been developed to infer heterogeneous tumor populations based on genomic signatures such as mutations and copy number variations. Despite accumulating evidence for the significance of global DNA methylation reprogramming in certain cancer types including myeloid malignancies, none of the bioinformatic tools are designed to exploit subclonally reprogrammed methylation patterns to reveal constituent populations of a tumor. In accordance with the notion of global methylation reprogramming, our preliminary observations on acute myeloid leukemia (AML) samples implied the existence of subclonally occurring focal methylation aberrance throughout the genome. RESULTS: We present PRISM, a tool for inferring the composition of epigenetically distinct subclones of a tumor solely from methylation patterns obtained by reduced representation bisulfite sequencing. PRISM adopts DNA methyltransferase 1-like hidden Markov model-based in silico proofreading for the correction of erroneous methylation patterns. With error-corrected methylation patterns, PRISM focuses on a short individual genomic region harboring dichotomous patterns that can be split into fully methylated and unmethylated patterns. Frequencies of such two patterns form a sufficient statistic for subclonal abundance. A set of statistics collected from each genomic region is modeled with a beta-binomial mixture. Fitting the mixture with expectation-maximization algorithm finally provides inferred composition of subclones. Applying PRISM for two AML samples, we demonstrate that PRISM could infer the evolutionary history of malignant samples from an epigenetic point of view. AVAILABILITY AND IMPLEMENTATION: PRISM is freely available on GitHub (https://github.com/dohlee/prism). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Do-Hoon Lee, Sangseon Lee, Sun Kim |
Bioinform. | 3 |
| 2019 | HTRgene: a computational method to perform the integrated analysis of multiple heterogeneous time-series data: case analysis of cold and heat stress response signaling genes in ArabidopsisabstractBACKGROUND: Integrated analysis that uses multiple sample gene expression data measured under the same stress can detect stress response genes more accurately than analysis of individual sample data. However, the integrated analysis is challenging since experimental conditions (strength of stress and the number of time points) are heterogeneous across multiple samples. RESULTS: HTRgene is a computational method to perform the integrated analysis of multiple heterogeneous time-series data measured under the same stress condition. The goal of HTRgene is to identify "response order preserving DEGs" that are defined as genes not only which are differentially expressed but also whose response order is preserved across multiple samples. The utility of HTRgene was demonstrated using 28 and 24 time-series sample gene expression data measured under cold and heat stress in Arabidopsis. HTRgene analysis successfully reproduced known biological mechanisms of cold and heat stress in Arabidopsis. Also, HTRgene showed higher accuracy in detecting the documented stress response genes than existing tools. CONCLUSIONS: HTRgene, a method to find the ordering of response time of genes that are commonly observed among multiple time-series samples, successfully integrated multiple heterogeneous time-series gene expression datasets. It can be applied to many research problems related to the integration of time series data analysis. Hongryul Ahn, Inuk Jung, Heejoon Chae, Dongwon Kang, Woosuk Jung, Sun Kim |
BMC Bioinform. | 6 |
| 2019 | Venn-diaNet : venn diagram based network propagation analysis framework for comparing multiple biological experimentsabstractBACKGROUND: The main research topic in this paper is how to compare multiple biological experiments using transcriptome data, where each experiment is measured and designed to compare control and treated samples. Comparison of multiple biological experiments is usually performed in terms of the number of DEGs in an arbitrary combination of biological experiments. This process is usually facilitated with Venn diagram but there are several issues when Venn diagram is used to compare and analyze multiple experiments in terms of DEGs. First, current Venn diagram tools do not provide systematic analysis to prioritize genes. Because that current tools generally do not fully focus to prioritize genes, genes that are located in the segments in the Venn diagram (especially, intersection) is usually difficult to rank. Second, elucidating the phenotypic difference only with the lists of DEGs and expression values is challenging when the experimental designs have the combination of treatments. Experiment designs that aim to find the synergistic effect of the combination of treatments are very difficult to find without an informative system. RESULTS: We introduce Venn-diaNet, a Venn diagram based analysis framework that uses network propagation upon protein-protein interaction network to prioritizes genes from experiments that have multiple DEG lists. We suggest that the two issues can be effectively handled by ranking or prioritizing genes with segments of a Venn diagram. The user can easily compare multiple DEG lists with gene rankings, which is easy to understand and also can be coupled with additional analysis for their purposes. Our system provides a web-based interface to select seed genes in any of areas in a Venn diagram and then perform network propagation analysis to measure the influence of the selected seed genes in terms of ranked list of DEGs. CONCLUSIONS: We suggest that our system can logically guide to select seed genes without additional prior knowledge that makes us free from the seed selection of network propagation issues. We showed that Venn-diaNet can reproduce the research findings reported in the original papers that have experiments that compare two, three and eight experiments. Venn-diaNet is freely available at: http://biohealth.snu.ac.kr/software/venndianet. Benjamin Hur, Dongwon Kang, Sangseon Lee, Ji Hwan Moon, Gung Lee, Sun Kim |
BMC Bioinform. | 6 |
| 2018 | HTRgene: Integrating Multiple Heterogeneous Time-series Data to Investigate Cold and Heat Stress Response Signaling Genes in Arabidopsis
Hongryul Ahn, Inuk Jung, Heejoon Chae, Dongwon Kang, Woosuk Jung, Sun Kim |
BIBM | 6 |
| 2018 | Identifying stress-related genes and predicting stress types in Arabidopsis using logical correlation layer and CMCL loss through time-series data
Dongwon Kang, Hongryul Ahn, Sangseon Lee, Chai-Jin Lee, Jihye Hur, Woosuk Jung, Sun Kim |
BIBM | 7 |
| 2018 | Hybrid Approach of Relation Network and Localized Graph Convolutional Filtering for Breast Cancer Subtype ClassificationabstractNetwork biology has been successfully used to help reveal complex mechanisms of disease, especially cancer. On the other hand, network biology requires in-depth knowledge to construct disease-specific networks, but our current knowledge is very limited even with the recent advances in human cancer biology. Deep learning has shown an ability to address the problem like this. However, it conventionally used grid-like structured data, thus application of deep learning technologies to the human disease subtypes is yet to be explored. To overcome the issue, we propose a hybrid model, which integrates two key components 1) graph convolution neural network (graph CNN) and 2) relation network (RN). Experimental results on synthetic data and breast cancer data demonstrate that our proposed method shows better performances than existing methods. SungMin Rhee, Seokjun Seo, Sun Kim |
IJCAI | 3 |
| 2018 | A Fast Deep Learning Model for Textual Relevance in Biomedical Information RetrievalabstractPublications in the life sciences are characterized by a large technical vocabulary, with many lexical and semantic variations for expressing the same concept. Towards addressing the problem of relevance in biomedical literature search, we introduce a deep learning model for the relevance of a document's text to a keyword style query. Limited by a relatively small amount of training data, the model uses pre-trained word embeddings. With these, the model first computes a variable-length Delta matrix between the query and document, representing a difference between the two texts, which is then passed through a deep convolution stage followed by a deep feed-forward network to compute a relevance score. This results in a fast model suitable for use in an online search engine. The model is robust and outperforms comparable state-of-the-art deep learning approaches. Sunil Mohan, Nicolas Fiorini, Sun Kim, Zhiyong Lu |
WWW | 3 |
| 2018 | DeepFam: deep learning based alignment-free method for protein family modeling and predictionabstractMotivation: A large number of newly sequenced proteins are generated by the next-generation sequencing technologies and the biochemical function assignment of the proteins is an important task. However, biological experiments are too expensive to characterize such a large number of protein sequences, thus protein function prediction is primarily done by computational modeling methods, such as profile Hidden Markov Model (pHMM) and k-mer based methods. Nevertheless, existing methods have some limitations; k-mer based methods are not accurate enough to assign protein functions and pHMM is not fast enough to handle large number of protein sequences from numerous genome projects. Therefore, a more accurate and faster protein function prediction method is needed. Results: In this paper, we introduce DeepFam, an alignment-free method that can extract functional information directly from sequences without the need of multiple sequence alignments. In extensive experiments using the Clusters of Orthologous Groups (COGs) and G protein-coupled receptor (GPCR) dataset, DeepFam achieved better performance in terms of accuracy and runtime for predicting functions of proteins compared to the state-of-the-art methods, both alignment-free and alignment-based methods. Additionally, we showed that DeepFam has a power of capturing conserved regions to model protein families. In fact, DeepFam was able to detect conserved regions documented in the Prosite database while predicting functions of proteins. Our deep learning method will be useful in characterizing functions of the ever increasing protein sequences. Availability and implementation: Codes are available at https://bhi-kimlab.github.io/DeepFam. Seokjun Seo, Minsik Oh, Youngjune Park, Sun Kim |
Bioinform. | 4 |
| 2018 | BRCA-Pathway: a structural integration and visualization system of TCGA breast cancer data on KEGG pathwaysabstractBACKGROUND: Bioinformatics research for finding biological mechanisms can be done by analysis of transcriptome data with pathway based interpretation. Therefore, researchers have tried to develop tools to analyze transcriptome data with pathway based interpretation. Over the years, the amount of omics data has become huge, e.g., TCGA, and the data types to be analyzed have come in many varieties, including mutations, copy number variations, and transcriptome. We also need to consider a complex relationship with regulators of genes, particularly Transcription Factors(TF). However, there has not been a system for pathway based exploration and analysis of TCGA multi-omics data. In this reason, We have developed a web based system BRCA-Pathway to fulfill the need for pathway based analysis of TCGA multi-omics data. RESULTS: BRCA-Pathway is a structured integration and visual exploration system of TCGA breast cancer data on KEGG pathways. For data integration, a relational database is designed and used to integrate multi-omics data of TCGA-BRCA, KEGG pathway data, Hallmark gene sets, transcription factors, driver genes, and PAM50 subtypes. For data exploration, multi-omics data such as SNV, CNV and gene expression can be visualized simultaneously in KEGG pathway maps, together with transcription factors-target genes (TF-TG) correlation and relationships among cancer driver genes. In addition, 'Pathways summary' and 'Oncoprint' with mutual exclusivity sort can be generated dynamically with a request by the user. Data in BRCA-Pathway can be downloaded by REST API for further analysis. CONCLUSIONS: BRCA-Pathway helps researchers navigate omics data towards potentially important genes, regulators, and discover complex patterns involving mutations, CNV, and gene expression data of various patient groups in the biological pathway context. In addition, mutually exclusive genomic alteration patterns in a specific pathway can be generated. BRCA-Pathway can provide an integrative perspective on the breast cancer omics data, which can help researchers discover new insights on the biological mechanisms of breast cancer. Inyoung Kim, Saemi Choi, Sun Kim |
BMC Bioinform. | 3 |
| 2018 | Discovering themes in biomedical literature using a projection-based algorithmabstractThe need to organize any large document collection in a manner that facilitates human comprehension has become crucial with the increasing volume of information available. Two common approaches to provide a broad overview of the information space are document clustering and topic modeling. Clustering aims to group documents or terms into meaningful clusters. Topic modeling, on the other hand, focuses on finding coherent keywords for describing topics appearing in a set of documents. In addition, there have been efforts for clustering documents and finding keywords simultaneously. We present an algorithm to analyze document collections that is based on a notion of a theme, defined as a dual representation based on a set of documents and key terms. In this work, a novel vector space mechanism is proposed for computing themes. Starting with a single document, the theme algorithm treats terms and documents as explicit components, and iteratively uses each representation to refine the other until the theme is detected. The method heavily relies on an optimization routine that we refer to as the projection algorithm which, under specific conditions, is guaranteed to converge to the first singular vector of a data matrix. We apply our algorithm to a collection of about sixty thousand PubMed Ⓡ documents examining the subject of Single Nucleotide Polymorphism, evaluate the results and show the effectiveness and scalability of the proposed method. This study presents a contribution on theoretical and algorithmic levels, as well as demonstrates the feasibility of the method for large scale applications. The evaluation of our system on benchmark datasets demonstrates that our method compares favorably with the current state-of-the-art methods in computing clusters of documents with coherent topic terms. Lana Yeganova, Sun Kim, Grigory Balasanov, W. John Wilbur |
BMC Bioinform. | 2 |
| 2018 | Editorial for Selected Papers of a Joint Conferences, Genome Informatics Workshop/International Conference on Bioinformatics (GIW/InCoB) 2015abstractThe four papers in this special section were presented at the joint 2015 Genome Informatics Workshop/International Conference on Bioinformatics (GIW/InCoB). Sun Kim |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2017 | TimesVector: a vectorized clustering approach to the analysis of time series transcriptome data from multiple phenotypesabstractMOTIVATION: Identifying biologically meaningful gene expression patterns from time series gene expression data is important to understand the underlying biological mechanisms. To identify significantly perturbed gene sets between different phenotypes, analysis of time series transcriptome data requires consideration of time and sample dimensions. Thus, the analysis of such time series data seeks to search gene sets that exhibit similar or different expression patterns between two or more sample conditions, constituting the three-dimensional data, i.e. gene-time-condition. Computational complexity for analyzing such data is very high, compared to the already difficult NP-hard two dimensional biclustering algorithms. Because of this challenge, traditional time series clustering algorithms are designed to capture co-expressed genes with similar expression pattern in two sample conditions. RESULTS: We present a triclustering algorithm, TimesVector, specifically designed for clustering three-dimensional time series data to capture distinctively similar or different gene expression patterns between two or more sample conditions. TimesVector identifies clusters with distinctive expression patterns in three steps: (i) dimension reduction and clustering of time-condition concatenated vectors, (ii) post-processing clusters for detecting similar and distinct expression patterns and (iii) rescuing genes from unclassified clusters. Using four sets of time series gene expression data, generated by both microarray and high throughput sequencing platforms, we demonstrated that TimesVector successfully detected biologically meaningful clusters of high quality. TimesVector improved the clustering quality compared to existing triclustering tools and only TimesVector detected clusters with differential expression patterns across conditions successfully. AVAILABILITY AND IMPLEMENTATION: The TimesVector software is available at http://biohealth.snu.ac.kr/software/TimesVector/. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Inuk Jung, Kyuri Jo, Hyejin Kang, Hongryul Ahn, Youngjae Yu, Sun Kim |
Bioinform. | 6 |
| 2017 | Bridging the gap: Incorporating a semantic similarity measure for effectively mapping PubMed queries to documents
Sun Kim, Nicolas Fiorini, W. John Wilbur, Zhiyong Lu |
J. Biomed. Informatics | 1 |
| 2016 | Networks and models for the integrated analysis of multi omics dataabstractThese days, genome-wide measurements of genetic and epigenetics events, a.k.a omics data, are routinely produced; epigenetics is control mechanisms of genetics events as epi-means `on' or `upon'. As a result, a huge amount of omics data measured from different genetic and epigenetic events are available. For example, the amount of data at The Cancer Genome Atlas(TCGA) alone exceeds 2.5 peta byte as of October 2016. Unfortunately, the dimensions of omics data is huge, typically tens to hundreds or even millions of thousands while the number of samples are limited typically a few to thousands. Thus mining genetic and epigenetic data measured in different phenotype conditions is a very challenging problem, that is, small data sets on extremely high dimensions. Furthermore, all genetic and epigenetic events are inter-related. Thus it is necessary to perform integrated analysis of omics data sets of different types, which is even more challenging. To address these technical challenges, the bioinformatics community has used virtually all known network based analysis techniques, including recently developed deep neural networks. My group has been trying the network based integrated analysis of omics data at three different levels. First, we have been investigating on computational methods for associating different genetic and epigenetic events, which can be viewed as methods for defining edges in the network. Second, we have been developing mining subnetworks on the phenotype and time dimensions. Third, we have recently begun to investigate on the use of deep learning techniques for the integrated analysis of omics data. An important goal of our research is to combine network analysis and deep learning techniques to construct models or draw maps of cancer cells at multiple levels such as genomic mutations, gene activation/suppressions, epigenetic events including DNA methylation, histone modifications, and miRNA interference, biological pathways, and finally at the whole cell level including tumor heterogeneity and clonal evolution. Sun Kim |
BIBM | 1 |
| 2016 | Iterative segmented least square method for functional microRNA-mRNA module discovery in breast cancerabstractMicroRNAs (miRNAs) have significant biological roles at the molecular level by regulating genes post-transcriptionally. To understand the functional effects of miRNAs in different biological contexts, it is essential to elucidate miRNA-mRNA regulatory modules (MRMs). The computational complexity for inferencing MRMs is very high due to the many-to-many relationships between miRNAs and mRNAs and inferencing MRMs is still a challenging unresolved problem. In this paper, we propose a novel iterative segmented least square method for functional MRM discovery. Our method operates in two steps: 1) grouping and ordering the miRNAs and mRNAs to build per-sample matrices representing miRNA-mRNA regulations, and 2) determining maximum sized modules from structured miRNA-mRNA matrices. In experiments with human breast cancer data sets from TCGA, we show that our method outperforms existing methods in terms of both GO similarity and cluster evaluation. In addition, we show that modules determined by our method can be used for breast cancer survival prediction and subtype classification. SungMin Rhee, Sangsoo Lim, Sun Kim |
BIBM | 3 |
| 2016 | Influence maximization in time bounded network identifies transcription factors regulating perturbed pathwaysabstractMOTIVATION: To understand the dynamic nature of the biological process, it is crucial to identify perturbed pathways in an altered environment and also to infer regulators that trigger the response. Current time-series analysis methods, however, are not powerful enough to identify perturbed pathways and regulators simultaneously. Widely used methods include methods to determine gene sets such as differentially expressed genes or gene clusters and these genes sets need to be further interpreted in terms of biological pathways using other tools. Most pathway analysis methods are not designed for time series data and they do not consider gene-gene influence on the time dimension. RESULTS: In this article, we propose a novel time-series analysis method TimeTP for determining transcription factors (TFs) regulating pathway perturbation, which narrows the focus to perturbed sub-pathways and utilizes the gene regulatory network and protein-protein interaction network to locate TFs triggering the perturbation. TimeTP first identifies perturbed sub-pathways that propagate the expression changes along the time. Starting points of the perturbed sub-pathways are mapped into the network and the most influential TFs are determined by influence maximization technique. The analysis result is visually summarized in TF-PATHWAY MAP IN TIME CLOCK: TimeTP was applied to PIK3CA knock-in dataset and found significant sub-pathways and their regulators relevant to the PIP3 signaling pathway. AVAILABILITY AND IMPLEMENTATION: TimeTP is implemented in Python and available at http://biohealth.snu.ac.kr/software/TimeTP/Supplementary information: Supplementary data are available at Bioinformatics online. CONTACT: [email protected]. Kyuri Jo, Inuk Jung, Ji Hwan Moon, Sun Kim |
Bioinform. | 4 |
| 2016 | Meshable: searching PubMed abstracts by utilizing MeSH and MeSH-derived topical termsabstractUNLABELLED: Medical Subject Headings (MeSH(®)) is a controlled vocabulary for indexing and searching biomedical literature. MeSH terms and subheadings are organized in a hierarchical structure and are used to indicate the topics of an article. Biologists can use either MeSH terms as queries or the MeSH interface provided in PubMed(®) for searching PubMed abstracts. However, these are rarely used, and there is no convenient way to link standardized MeSH terms to user queries. Here, we introduce a web interface which allows users to enter queries to find MeSH terms closely related to the queries. Our method relies on co-occurrence of text words and MeSH terms to find keywords that are related to each MeSH term. A query is then matched with the keywords for MeSH terms, and candidate MeSH terms are ranked based on their relatedness to the query. The experimental results show that our method achieves the best performance among several term extraction approaches in terms of topic coherence. Moreover, the interface can be effectively used to find full names of abbreviations and to disambiguate user queries. AVAILABILITY AND IMPLEMENTATION: https://www.ncbi.nlm.nih.gov/IRET/MESHABLE/ CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Sun Kim, Lana Yeganova, W. John Wilbur |
Bioinform. | 1 |
| 2016 | Prioritizing biological pathways by recognizing context in time-series gene expression dataabstractBACKGROUND: The primary goal of pathway analysis using transcriptome data is to find significantly perturbed pathways. However, pathway analysis is not always successful in identifying pathways that are truly relevant to the context under study. A major reason for this difficulty is that a single gene is involved in multiple pathways. In the KEGG pathway database, there are 146 genes, each of which is involved in more than 20 pathways. Thus activation of even a single gene will result in activation of many pathways. This complex relationship often makes the pathway analysis very difficult. While we need much more powerful pathway analysis methods, a readily available alternative way is to incorporate the literature information. RESULTS: In this study, we propose a novel approach for prioritizing pathways by combining results from both pathway analysis tools and literature information. The basic idea is as follows. Whenever there are enough articles that provide evidence on which pathways are relevant to the context, we can be assured that the pathways are indeed related to the context, which is termed as relevance in this paper. However, if there are few or no articles reported, then we should rely on the results from the pathway analysis tools, which is termed as significance in this paper. We realized this concept as an algorithm by introducing Context Score and Impact Score and then combining the two into a single score. Our method ranked truly relevant pathways significantly higher than existing pathway analysis tools in experiments with two data sets. CONCLUSIONS: Our novel framework was implemented as ContextTRAP by utilizing two existing tools, TRAP and BEST. ContextTRAP will be a useful tool for the pathway based analysis of gene expression data since the user can specify the context of the biological experiment in a set of keywords. The web version of ContextTRAP is available at http://biohealth.snu.ac.kr/software/contextTRAP . Jusang Lee, Kyuri Jo, Sunwon Lee, Jaewoo Kang, Sun Kim |
BMC Bioinform. | 5 |
| 2015 | Summarizing Topical Contents from PubMed Documents Using a Thematic AnalysisabstractImproving the search and browsing experience in PubMed is a key component in helping users detect information of interest.In particular, when exploring a novel field, it is important to provide a comprehensive view for a specific subject.One solution for providing this panoramic picture is to find sub-topics from a set of documents.We propose a method that finds sub-topics that we refer to as themes and computes representative titles based on a set of documents in each theme.The method combines a thematic clustering algorithm and the Pool Adjacent Violators algorithm to induce significant themes.Then, for each theme, a title is computed using PubMed document titles and theme-dependent term scores.We tested our system on five disease sets from OMIM and evaluated the results based on normalized point-wise mutual information and MeSH terms.For both performance measures, the proposed approach outperformed LDA.The quality of theme titles were also evaluated by comparing them with manually created titles. Sun Kim, Lana Yeganova, W. John Wilbur |
EMNLP | 1 |
| 2015 | BioVLAB-MMIA-NGS: microRNA-mRNA integrated analysis using high-throughput sequencing dataabstractMOTIVATION: It is now well established that microRNAs (miRNAs) play a critical role in regulating gene expression in a sequence-specific manner, and genome-wide efforts are underway to predict known and novel miRNA targets. However, the integrated miRNA-mRNA analysis remains a major computational challenge, requiring powerful informatics systems and bioinformatics expertise. RESULTS: The objective of this study was to modify our widely recognized Web server for the integrated mRNA-miRNA analysis (MMIA) and its subsequent deployment on the Amazon cloud (BioVLAB-MMIA) to be compatible with high-throughput platforms, including next-generation sequencing (NGS) data (e.g. RNA-seq). We developed a new version called the BioVLAB-MMIA-NGS, deployed on both Amazon cloud and on a high-performance publicly available server called MAHA. By using NGS data and integrating various bioinformatics tools and databases, BioVLAB-MMIA-NGS offers several advantages. First, sequencing data is more accurate than array-based methods for determining miRNA expression levels. Second, potential novel miRNAs can be detected by using various computational methods for characterizing miRNAs. Third, because miRNA-mediated gene regulation is due to hybridization of an miRNA to its target mRNA, sequencing data can be used to identify many-to-many relationship between miRNAs and target genes with high accuracy. AVAILABILITY AND IMPLEMENTATION: http://epigenomics.snu.ac.kr/biovlab_mmia_ngs/. Heejoon Chae, SungMin Rhee, Kenneth P. Nephew, Sun Kim |
Bioinform. | 4 |
| 2015 | Identifying named entities from PubMed®; for enriching semantic categoriesabstractBACKGROUND: Controlled vocabularies such as the Unified Medical Language System (UMLS) and Medical Subject Headings (MeSH) are widely used for biomedical natural language processing (NLP) tasks. However, the standard terminology in such collections suffers from low usage in biomedical literature, e.g. only 13% of UMLS terms appear in MEDLINE. RESULTS: We here propose an efficient and effective method for extracting noun phrases for biomedical semantic categories. The proposed approach utilizes simple linguistic patterns to select candidate noun phrases based on headwords, and a machine learning classifier is used to filter out noisy phrases. For experiments, three NLP rules were tested and manually evaluated by three annotators. Our approaches showed over 93% precision on average for the headwords, "gene", "protein", "disease", "cell" and "cells". CONCLUSIONS: Although biomedical terms in knowledge-rich resources may define semantic categories, variations of the controlled terms in literature are still difficult to identify. The method proposed here is an effort to narrow the gap between controlled vocabularies and the entities used in text. Our extraction method cannot completely eliminate manual evaluation, however a simple and automated solution with high precision performance provides a convenient way for enriching semantic categories by incorporating terms obtained from the literature. Sun Kim, Zhiyong Lu, W. John Wilbur |
BMC Bioinform. | 1 |
| 2015 | Extracting drug-drug interactions from literature using a rich feature-based linear kernel approachabstractIdentifying unknown drug interactions is of great benefit in the early detection of adverse drug reactions. Despite existence of several resources for drug-drug interaction (DDI) information, the wealth of such information is buried in a body of unstructured medical text which is growing exponentially. This calls for developing text mining techniques for identifying DDIs. The state-of-the-art DDI extraction methods use Support Vector Machines (SVMs) with non-linear composite kernels to explore diverse contexts in literature. While computationally less expensive, linear kernel-based systems have not achieved a comparable performance in DDI extraction tasks. In this work, we propose an efficient and scalable system using a linear kernel to identify DDI information. The proposed approach consists of two steps: identifying DDIs and assigning one of four different DDI types to the predicted drug pairs. We demonstrate that when equipped with a rich set of lexical and syntactic features, a linear SVM classifier is able to achieve a competitive performance in detecting DDIs. In addition, the one-against-one strategy proves vital for addressing an imbalance issue in DDI type classification. Applied to the DDIExtraction 2013 corpus, our system achieves an F1 score of 0.670, as compared to 0.651 and 0.609 reported by the top two participating teams in the DDIExtraction 2013 challenge, both based on non-linear kernel methods. Sun Kim, Lana Yeganova, W. John Wilbur |
J. Biomed. Informatics | 1 |
| 2014 | Extracting drug-drug interactions from literature using a rich feature-based linear kernel approach
Sun Kim, Lana Yeganova, W. John Wilbur |
AMIA | 1 |
| 2014 | Retro: concept-based clustering of biomedical topical setsabstractMOTIVATION: Clustering methods can be useful for automatically grouping documents into meaningful clusters, improving human comprehension of a document collection. Although there are clustering algorithms that can achieve the goal for relatively large document collections, they do not always work well for small and homogenous datasets. METHODS: In this article, we present Retro-a novel clustering algorithm that extracts meaningful clusters along with concise and descriptive titles from small and homogenous document collections. Unlike common clustering approaches, our algorithm predicts cluster titles before clustering. It relies on the hypergeometric distribution model to discover key phrases, and generates candidate clusters by assigning documents to these phrases. Further, the statistical significance of candidate clusters is tested using supervised learning methods, and a multiple testing correction technique is used to control the overall quality of clustering. RESULTS: We test our system on five disease datasets from OMIM(®) and evaluate the results based on MeSH(®) term assignments. We further compare our method with several baseline and state-of-the-art methods, including K-means, expectation maximization, latent Dirichlet allocation-based clustering, Lingo, OPTIMSRC and adapted GK-means. The experimental results on the 20-Newsgroup and ODP-239 collections demonstrate that our method is successful at extracting significant clusters and is superior to existing methods in terms of quality of clusters. Finally, we apply our system to a collection of 6248 topical sets from the HomoloGene(®) database, a resource in PubMed(®). Empirical evaluation confirms the method is useful for small homogenous datasets in producing meaningful clusters with descriptive titles. AVAILABILITY AND IMPLEMENTATION: A web-based demonstration of the algorithm applied to a collection of sets from the HomoloGene database is available at http://www.ncbi.nlm.nih.gov/CBBresearch/Wilbur/IRET/CLUSTERING_HOMOLOGENE/index.html. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Lana Yeganova, Won Kim 0003, Sun Kim, W. John Wilbur |
Bioinform. | 3 |
| 2014 | Author name disambiguation for PubMedabstractLog analysis shows that PubMed users frequently use author names in queries for retrieving scientific literature. However, author name ambiguity may lead to irrelevant retrieval results. To improve the PubMed user experience with author name queries, we designed an author name disambiguation system consisting of similarity estimation and agglomerative clustering. A machine-learning method was employed to score the features for disambiguating a pair of papers with ambiguous names. These features enable the computation of pairwise similarity scores to estimate the probability of a pair of papers belonging to the same author, which drives an agglomerative clustering algorithm regulated by 2 factors: name compatibility and probability level. With transitivity violation correction, high precision author clustering is achieved by focusing on minimizing false-positive pairing. Disambiguation performance is evaluated with manual verification of random samples of pairs from clustering results. When compared with a state-of-the-art system, our evaluation shows that among all the pairs the lumping error rate drops from 10.1% to 2.2% for our system, while the splitting error rises from 1.8% to 7.7%. This results in an overall error rate of 9.9%, compared with 11.9% for the state-of-the-art method. Other evaluations based on gold standard data also show the increase in accuracy of our clustering. We attribute the performance improvement to the machine-learning method driven by a large-scale training set and the clustering algorithm regulated by a name compatibility scheme preferring precision. With integration of the author name disambiguation system into the PubMed search engine, the overall click-through-rate of PubMed users on author name query results improved from 34.9% to 36.9%. Wanli Liu, Rezarta Islamaj Dogan, Sun Kim, Donald C. Comeau, Won Kim 0003, Lana Yeganova, Zhiyong Lu, W. John Wilbur |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2013 | Towards simultaneous clustering and motif-modeling for a large number of protein familyabstractIn this paper, we propose a novel clustering and motif modeling framework for analyzing large number of protein family using k-mer. Our approach of using k-mers utilizes both occurring frequency and position information of k-mers that essential for classification yet not fully used in previous methods. We found that the structure has close relationship between motif of protein family and hence well describe important biological features or motifs of each protein family. The classification/clustering procedure are executed in incremental manner which was difficult for previous algorithms and is modeled by using bipartite model. Furthermore, the method can be efficiently implemented using parallel computing and hash. Experimental results using the entire COG family database shows that our model can model a large number of protein families without sacrificing accuracy. In addition, the classification structure, path of the graph for protein sequences, explains characteristic subsequences or motif of each family quite well. Thus the proposed method has the potential to model both protein families and motifs, even for a large number of families. Young Joon Yoo, Tushar Sandhan, Jin Young Choi 0002, Sun Kim |
BIBM | 4 |
| 2013 | Integrated profiling of three dimensional cell culture models and 3D microscopyabstractMOTIVATION: Our goal is to develop a screening platform for quantitative profiling of colony organizations in 3D cell culture models. The 3D cell culture models, which are also imaged in 3D, are functional assays that mimic the in vivo characteristics of the tissue architecture more faithfully than the 2D cultures. However, they also introduce significant computational challenges, with the main barriers being the effects of growth conditions, fixations and inherent complexities in segmentation that need to be resolved in the 3D volume. RESULTS: A segmentation strategy has been developed to delineate each nucleus in a colony that overcomes (i) the effects of growth conditions, (ii) variations in chromatin distribution and (iii) ambiguities formed by perceptual boundaries from adjacent nuclei. The strategy uses a cascade of geometric filters that are insensitive to spatial non-uniformity and partitions a clump of nuclei based on the grouping of points of maximum curvature at the interface of two neighboring nuclei. These points of maximum curvature are clustered together based on their coplanarity and proximity to define dissecting planes that separate the touching nuclei. The proposed curvature-based partitioning method is validated with both synthetic and real data, and is shown to have a superior performance against previous techniques. Validation and sensitivity analysis are coupled with the experimental design that includes a non-transformed cell line and three tumorigenic cell lines, which covers a wide range of phenotypic diversity in breast cancer. Colony profiling, derived from nuclear segmentation, reveals distinct indices for the morphogenesis of each cell line. Cemal Çagatay Bilgin, Sun Kim, Elle Leung, Hang Chang, Bahram Parvin |
Bioinform. | 2 |
| 2012 | Discriminative spatial pattern vectors selection for motor imagery classificationabstractIn this paper, we propose a novel method of designing a class-discriminative spatial filter assuming that a combination of spatial pattern vectors, irrespective of the eigenvalues, can produce better performance in terms of classification accuracy. We select discriminative spatial pattern vectors that determine features in a pairwise manner, i.e., eigenvectors of the k-th largest eigenvalue and the k-the lowest eigenvalue. Although the pair of the eigenvectors of the K largest and the K smallest eigenvalues helps extract discriminative features, we believe that a different set of eigenvector pairs is more appropriate to extract class-discriminative features. In our experiments, the proposed method outperformed the conventional approach. Kyeong-Yeon Lee, Sun Kim |
SMC | 2 |
| 2012 | PIE the search: searching PubMed literature for protein interaction informationabstractMOTIVATION: Finding protein-protein interaction (PPI) information from literature is challenging but an important issue. However, keyword search in PubMed(®) is often time consuming because it requires a series of actions that refine keywords and browse search results until it reaches a goal. Due to the rapid growth of biomedical literature, it has become more difficult for biologists and curators to locate PPI information quickly. Therefore, a tool for prioritizing PPI informative articles can be a useful assistant for finding this PPI-relevant information. RESULTS: PIE (Protein Interaction information Extraction) the search is a web service implementing a competition-winning approach utilizing word and syntactic analyses by machine learning techniques. For easy user access, PIE the search provides a PubMed-like search environment, but the output is the list of articles prioritized by PPI confidence scores. By obtaining PPI-related articles at high rank, researchers can more easily find the up-to-date PPI information, which cannot be found in manually curated PPI databases. AVAILABILITY: http://www.ncbi.nlm.nih.gov/IRET/PIE/. Sun Kim, Dongseop Kwon, Soo-Yong Shin, W. John Wilbur |
Bioinform. | 1 |
| 2012 | GeneclusterViz: a tool for conserved gene cluster visualization, exploration and analysisabstractMOTIVATION: Gene clusters are arrangements of functionally related genes on a chromosome. In bacteria, it is expected that evolutionary pressures would conserve these arrangements due to the functional advantages they provide. Visualization of conserved gene clusters across multiple genomes provides key insights into their evolutionary histories. Therefore, a software tool that enables visualization and functional analyses of gene clusters would be a great asset to the biological research community. RESULTS: We have developed GeneclusterViz, a Java-based tool that allows for the visualization, exploration and downstream analyses of conserved gene clusters across multiple genomes. GeneclusterViz combines an easy-to-use exploration interface for gene clusters with a host of other analysis features such as multiple sequence alignments, phylogenetic analyses and integration with the KEGG pathway database. AVAILABILITY: http://biohealth.snu.ac.kr/GeneclusterViz/; http://microbial.informatics.indiana.edu/GeneclusterViz/ Vikas Rao Pejaver, Jaehyun An, SungMin Rhee, Ankita Bhan, Jeong-Hyeon Choi, Boshu Liu, Heewook Lee, Pamela J. Brown, David Kysela, Yves V. Brun, Sun Kim |
Bioinform. | 11 |
| 2012 | A novel k-mer mixture logistic regression for methylation susceptibility modeling of CpG dinucleotides in human gene promotersabstractBACKGROUND: DNA methylation is essential for normal development and differentiation and plays a crucial role in the development of nearly all types of cancer. Aberrant DNA methylation patterns, including genome-wide hypomethylation and region-specific hypermethylation, are frequently observed and contribute to the malignant phenotype. A number of studies have recently identified distinct features of genomic sequences that can be used for modeling specific DNA sequences that may be susceptible to aberrant CpG methylation in both cancer and normal cells. Although it is now possible, using next generation sequencing technologies, to assess human methylomes at base resolution, no reports currently exist on modeling cell type-specific DNA methylation susceptibility. Thus, we conducted a comprehensive modeling study of cell type-specific DNA methylation susceptibility at three different resolutions: CpG dinucleotides, CpG segments, and individual gene promoter regions. RESULTS: Using a k-mer mixture logistic regression model, we effectively modeled DNA methylation susceptibility across five different cell types. Further, at the segment level, we achieved up to 0.75 in AUC prediction accuracy in a 10-fold cross validation study using a mixture of k-mers. CONCLUSIONS: The significance of these results is three fold: 1) this is the first report to indicate that CpG methylation susceptible "segments" exist; 2) our model demonstrates the significance of certain k-mers for the mixture model, potentially highlighting DNA sequence features (k-mers) of differentially methylated, promoter CpG island sequences across different tissue types; 3) as only 3 or 4 bp patterns had previously been used for modeling DNA methylation susceptibility, ours is the first demonstration that 6-mer modeling can be performed without loss of accuracy. Youngik Yang, Kenneth P. Nephew, Sun Kim |
BMC Bioinform. | 3 |
| 2011 | BioVLAB-MMIA: A Reconfigurable Cloud Computing Environment for microRNA and mRNA Integrated AnalysisabstractMicro RNAs, by regulating the expression of hundreds of target genes, play critical roles in developmental biology and the etiology of numerous diseases, including cancer. As vast amounts of micro RNA expression profile data are now publicly available, the integration of those data sets with gene expression profiles represents an extremely active area of life science research. However, the ability to conduct genome-wide micro RNA-mRNA (gene) integration currently requires sophisticated, high-end informatics tools, significant expertise in bioinformatics and computer science to carry out the complex integration analysis. In addition, increased computing infrastructure capabilities are essential in order to accommodate large data sets. In this study, we have extending BioVLAB cloud workbench to develop an environment for the integrated analysis of micro RNA and mRNA expression data, named BioVLAB-MMIA. The workbench facilitates computations on the Amazon EC2 and S3 resources orchestrated by XBaya Workflow Suite. The advantages of BioVLAB-MMIA over the web-based MMIA system include: 1) readily expanded as new computational tools become available, 2) easily modifiable by re-configuring graphic icons in the workflow, 3) on-demand cloud computing resources can be used on an "as needed" basis, 4) distributed orchestration supports complex and long running workflows asynchronously. Hyungro Lee, Youngik Yang, Heejoon Chae, Seungyoon Nam, Donghoon Choi, Patanachai Tangchaisin, Chathura Herath, Suresh Marru, Kenneth P. Nephew, Sun Kim |
BIBM | 10 |
| 2011 | Classifying protein-protein interaction articles using word and syntactic featuresabstractBACKGROUND: Identifying protein-protein interactions (PPIs) from literature is an important step in mining the function of individual proteins as well as their biological network. Since it is known that PPIs have distinctive patterns in text, machine learning approaches have been successfully applied to mine these patterns. However, the complex nature of PPI description makes the extraction process difficult. RESULTS: Our approach utilizes both word and syntactic features to effectively capture PPI patterns from biomedical literature. The proposed method automatically identifies gene names by a Priority Model, then extracts grammar relations using a dependency parser. A large margin classifier with Huber loss function learns from the extracted features, and unknown articles are predicted using this data-driven model. For the BioCreative III ACT evaluation, our official runs were ranked in top positions by obtaining maximum 89.15% accuracy, 61.42% F1 score, 0.55306 MCC score, and 67.98% AUC iP/R score. CONCLUSIONS: Even though problems still remain, utilizing syntactic information for article-level filtering helps improve PPI ranking performance. The proposed system is a revision of previously developed algorithms in our group for the ACT evaluation. Our approach is valuable in showing how to use grammatical relations for PPI article filtering, in particular, with a limited training corpus. While current performance is far from satisfactory as an annotation tool, it is already useful for a PPI article search engine since users are mainly focused on highly-ranked results. Sun Kim, W. John Wilbur |
BMC Bioinform. | 1 |
| 2011 | The Protein-Protein Interaction tasks of BioCreative III: classification/ranking of articles and linking bio-ontology concepts to full textabstractBACKGROUND: Determining usefulness of biomedical text mining systems requires realistic task definition and data selection criteria without artificial constraints, measuring performance aspects that go beyond traditional metrics. The BioCreative III Protein-Protein Interaction (PPI) tasks were motivated by such considerations, trying to address aspects including how the end user would oversee the generated output, for instance by providing ranked results, textual evidence for human interpretation or measuring time savings by using automated systems. Detecting articles describing complex biological events like PPIs was addressed in the Article Classification Task (ACT), where participants were asked to implement tools for detecting PPI-describing abstracts. Therefore the BCIII-ACT corpus was provided, which includes a training, development and test set of over 12,000 PPI relevant and non-relevant PubMed abstracts labeled manually by domain experts and recording also the human classification times. The Interaction Method Task (IMT) went beyond abstracts and required mining for associations between more than 3,500 full text articles and interaction detection method ontology concepts that had been applied to detect the PPIs reported in them. RESULTS: A total of 11 teams participated in at least one of the two PPI tasks (10 in ACT and 8 in the IMT) and a total of 62 persons were involved either as participants or in preparing data sets/evaluating these tasks. Per task, each team was allowed to submit five runs offline and another five online via the BioCreative Meta-Server. From the 52 runs submitted for the ACT, the highest Matthew's Correlation Coefficient (MCC) score measured was 0.55 at an accuracy of 89% and the best AUC iP/R was 68%. Most ACT teams explored machine learning methods, some of them also used lexical resources like MeSH terms, PSI-MI concepts or particular lists of verbs and nouns, some integrated NER approaches. For the IMT, a total of 42 runs were evaluated by comparing systems against manually generated annotations done by curators from the BioGRID and MINT databases. The highest AUC iP/R achieved by any run was 53%, the best MCC score 0.55. In case of competitive systems with an acceptable recall (above 35%) the macro-averaged precision ranged between 50% and 80%, with a maximum F-Score of 55%. CONCLUSIONS: The results of the ACT task of BioCreative III indicate that classification of large unbalanced article collections reflecting the real class imbalance is still challenging. Nevertheless, text-mining tools that report ranked lists of relevant articles for manual selection can potentially reduce the time needed to identify half of the relevant articles to less than 1/4 of the time when compared to unranked results. Detecting associations between full text articles and interaction detection method PSI-MI terms (IMT) is more difficult than might be anticipated. This is due to the variability of method term mentions, errors resulting from pre-processing of articles provided as PDF files, and the heterogeneity and different granularity of method term concepts encountered in the ontology. However, combining the sophisticated techniques developed by the participants with supporting evidence strings derived from the articles for human interpretation could result in practical modules for biological annotation workflows. Martin Krallinger, Miguel Vázquez, Florian Leitner, David Salgado, Andrew Chatr-aryamontri, Andrew G. Winter, Livia Perfetto, Leonardo Briganti, Luana Licata, Marta Iannuccelli, Luisa Castagnoli, Gianni Cesareni, Mike Tyers, Gerold Schneider, Fabio Rinaldi 0001, Robert Leaman, Graciela Gonzalez-Hernandez, Sérgio Matos, Sun Kim, W. John Wilbur, Luis M. Rocha, Hagit Shatkay, Ashish V. Tendulkar, Shashank Agarwal, Xinglong Wang, Rafal Rak, Keith Noto, Charles Elkan, Zhiyong Lu |
BMC Bioinform. | 19 |
| 2010 | Gene cluster profile vectors: A novel method to infer functional coupling using both gene proximity and co-occurrence profilesabstractProximity-based methods and co-evolution-based phylogenetic profiles methods have been successfully used for the identification of functionally related genes. Proximity-based methods are effective for physically clustered genes while the phylogenetic profiles method is effective for co-occurring gene sets. However, both methods predict many false positives and false negatives. In this paper, we propose the Gene Cluster Profile Vector (GCPV) method, which combines these two methods by using phylogenetic profiles of whole gene clusters. Moreover, the GCPV method is, currently, the only method that allows for the characterization of relationships between gene clusters themselves. The GCPV method groups together reasonably related operons in E. coli about 60% of the time. The method is minimally dependent on the reference genome set used and it outperforms the conventional phylogenetic profiles method. Finally, we show that the method works well for predicted gene clusters from C. crescentus and can serve as an important tool not only for understanding gene function, but also for elucidating mechanisms of general biological processes. Vikas Rao Pejaver, Sun Kim |
BIBM | 2 |
| 2010 | Data mining for the study of disease genes and proteins
Sun Kim |
Artif. Intell. Medicine | 1 |
| 2010 | Annotation confidence score for genome annotation: a genome comparison approachabstractMOTIVATION: The massively parallel sequencing technology can be used by small research labs to generate genome sequences of their research interest. However, annotation of genomes still relies on the manual process, which becomes a serious bottleneck to the high-throughput genome projects. Recently, automatic annotation methods are increasingly more accurate, but there are several issues. One important challenge in using automatic annotation methods is to distinguish annotation quality of ORFs or genes. The availability of such annotation quality of genes can reduce the human labor cost dramatically since manual inspection can focus only on genes with low-annotation quality scores. RESULTS: In this article, we propose a novel annotation quality or confidence scoring scheme, called Annotation Confidence Score (ACS), using a genome comparison approach. The scoring scheme is computed by combining sequence and textual annotation similarity using a modified version of a logistic curve. The most important feature of the proposed scoring scheme is to generate a score that reflects the excellence in annotation quality of genes by automatically adjusting the number of genomes used to compute the score and their phylogenetic distance. Extensive experiments with bacterial genomes showed that the proposed scoring scheme generated scores for annotation quality according to the quality of annotation regardless of the number of reference genomes and their phylogenetic distance. AVAILABILITY: http://microbial.informatics.indiana.edu/acs Youngik Yang, Donald Gilbert, Sun Kim |
Bioinform. | 3 |
| 2009 | Evolving hypernetwork models of binary time series for forecasting price movements on stock marketsabstractThe paper proposes a hypernetwork-based method for stock market prediction through a binary time series problem. Hypernetworks are a random hypergraph structure of higher-order probabilistic relations of data. The problem we tackle concerns the prediction of price movements (up/down) on stock markets. Compared to previous approaches, the proposed method discovers a large population of variable subpatterns, i.e. local and global patterns, using a novel evolutionary hypernetwork. An output is obtained from combining these patterns. In the paper, we describe two methods for assessing the prediction quality of the hypernetwork approach. Applied to the Dow Jones Industrial Average Index and the Korea Composite Stock Price Index data, the experimental results show that the proposed method effectively learns and predicts the time series information. In particular, the hypernetwork approach outperforms other machine learning methods such as support vector machines, naive Bayes, multilayer perceptrons, and k-nearest neighbors. Elena Bautu, Sun Kim, Andrei Bautu, Henri Luchian, Byoung-Tak Zhang |
IEEE Congress on Evolutionary Computation | 2 |
| 2009 | Evolutionary hypernetwork classifiers for protein-proteininteraction sentence filteringabstractProtein-Protein Interaction (PPI) extraction, among ongoing biomedical text mining challenges, is becoming a topic in focus because of its crucial role in providing a starting point to understand biological processes. Machine learning (ML) techniques have been applied to extract the PPI information from biomedical literature. Although they have provided reasonable performance so far, more features are required for real use. In particular, many ML-approaches lack human understandability for learned models. Here, we propose a novel method for classifying PPI sentences. Our approach utilizes the modified hypernetwork model, a hypergraph with weighted hyperedges that are calibrated via an evolutionary learning method. The evolutionary hypernetwork memorizes fragments of training patterns while self-adjusting its own structure for detecting PPI sentences. For experiments, we show that our approach provides competitive performance compared to other ML methods. Apart from its superior classification performance, the evolving hypernetwork model comes with a highly interpretable structure. We show how significant PPI patterns can be naturally extracted from the learned model. We also analyze the discovered patterns. Jakramate Bootkrajang, Sun Kim, Byoung-Tak Zhang |
GECCO | 2 |
| 2009 | Computational analysis of microRNA profiles and their target genes suggests significant involvement in breast cancer antiestrogen resistanceabstractMOTIVATION: Recent evidence shows significant involvement of microRNAs (miRNAs) in the initiation and progression of numerous cancers; however, the role of these in tumor drug resistance remains unknown. RESULTS: By comparing global miRNA and mRNA expression patterns, we examined the role of miRNAs in resistance to the 'pure antiestrogen' fulvestrant, using fulvestrant-resistant MCF7-FR cells and their drug-sensitive parental estrogen receptor (ER)-positive MCF7 cells. We identified 14 miRNAs downregulated in MCF7-FR cells and then used both TargetScan and PITA to predict potential target genes. We found a negative correlation between expression of these miRNAs and their predicted target mRNA transcripts. In genes regulated by multiple miRNAs or having multiple miRNA-targeting sites, an even stronger negative correlation was found. Pathway analyses predicted these miRNAs to regulate specific cancer-associated signal cascades. These results suggest a significant role for miRNA-regulated gene expression in the onset of breast cancer antiestrogen resistance, and an improved understanding of this phenomenon could lead to better therapies for this often fatal condition. Fuxiao Xin, Meng Li 0013, Curtis Balch, Michael Thomson, Meiyun Fan, Scott M. Hammond, Sun Kim, Kenneth P. Nephew |
Bioinform. | 8 |
| 2008 | BioVLAB-Microarray: Microarray Data Analysis in Virtual EnvironmentabstractMicroarray technology is a high-throughput experimental technique that can measure expression levels of hundreds of thousands of genes simultaneously. To interpret massive data from gene-expression microarray experiments, biologists encounter computational and analytical challenges. This is especially challenging for small research labs that lack local computing and bioinformatics expertise. Here, we introduce a virtual analysis system for microarray gene expression data in computing clouds with flexible and configurable GUI workflow engine so that biologists are able to analyze the data in many angles without worrying about computational and bioinformatics issues. Youngik Yang, Jong Choi 0001, Kwangmin Choi, Marlon E. Pierce, Dennis Gannon, Sun Kim |
eScience | 6 |
| 2008 | A machine-learning approach to combined evidence validation of genome assembliesabstractMOTIVATION: While it is common to refer to 'the genome sequence' as if it were a single, complete and contiguous DNA string, it is in fact an assembly of millions of small, partially overlapping DNA fragments. Sophisticated computer algorithms (assemblers and scaffolders) merge these DNA fragments into contigs, and place these contigs into sequence scaffolds using the paired-end sequences derived from large-insert DNA libraries. Each step in this automated process is susceptible to producing errors; hence, the resulting draft assembly represents (in practice) only a likely assembly that requires further validation. Knowing which parts of the draft assembly are likely free of errors is critical if researchers are to draw reliable conclusions from the assembled sequence data. RESULTS: We develop a machine-learning method to detect assembly errors in sequence assemblies. Several in silico measures for assembly validation have been proposed by various researchers. Using three benchmarking Drosophila draft genomes, we evaluate these techniques along with some new measures that we propose, including the good-minus-bad coverage (GMB), the good-to-bad-ratio (RGB), the average Z-score (AZ) and the average absolute Z-score (ASZ). Our results show that the GMB measure performs better than the others in both its sensitivity and its specificity for assembly error detection. Nevertheless, no single method performs sufficiently well to reliably detect genomic regions requiring attention for further experimental verification. To utilize the advantages of all these measures, we develop a novel machine learning approach that combines these individual measures to achieve a higher prediction accuracy (i.e. greater than 90%). Our combined evidence approach avoids the difficult and often ad hoc selection of many parameters the individual measures require, and significantly improves the overall precisions on the benchmarking data sets. Jeong-Hyeon Choi, Sun Kim, Haixu Tang, Justen Andrews, Don G. Gilbert, John Colbourne |
Bioinform. | 2 |
| 2008 | Enriched transcription factor binding sites in hypermethylated gene promoters in drug resistant cancer cellsabstractMOTIVATION: In the human genome, 'CpG islands', CG-rich regions located in or near gene promoters, are normally unmethylated. However, in cancer cells, CpG islands frequently gain methylation, resulting in silencing of growth-limiting tumor suppressor genes. To our knowledge, the potential relationship between CpG island hypermethylation, transcription factor (TF) binding in local promoter regions and transcriptional control has not been previously explored in a genome-wide context. RESULTS: In this study, we utilized bioinformatics tools and TF binding site(TFBs) databases to globally analyze sequences methylated in a laboratory model for the development of drug-resistant cancer. Our results demonstrated that four TFBS were enriched in hypermethylated sequences. More interestingly, overrepresentation of these TFBS was observed in hyper-/hypo-methylated sequences where signi.cant changes in methylation levels were observed in drug-resistant cancer cells. In summary, we believe that these.ndings offer a means to further explore the relationship between DNA methylation and gene expression in drug resistance and tumorigenesis. Meng Li 0013, Hyun-il Henry Paik, Curtis Balch, Yoosung Kim, Lang Li 0001, Tim Hui-Ming Huang, Kenneth P. Nephew, Sun Kim |
Bioinform. | 8 |
| 2008 | ComPath: comparative enzyme analysis and annotation in pathway/subsystem contextsabstractBACKGROUND: Once a new genome is sequenced, one of the important questions is to determine the presence and absence of biological pathways. Analysis of biological pathways in a genome is a complicated task since a number of biological entities are involved in pathways and biological pathways in different organisms are not identical. Computational pathway identification and analysis thus involves a number of computational tools and databases and typically done in comparison with pathways in other organisms. This computational requirement is much beyond the capability of biologists, so information systems for reconstructing, annotating, and analyzing biological pathways are much needed. We introduce a new comparative pathway analysis workbench, ComPath, which integrates various resources and computational tools using an interactive spreadsheet-style web interface for reliable pathway analyses. RESULTS: ComPath allows users to compare biological pathways in multiple genomes using a spreadsheet style web interface where various sequence-based analysis can be performed either to compare enzymes (e.g. sequence clustering) and pathways (e.g. pathway hole identification), to search a genome for de novo prediction of enzymes, or to annotate a genome in comparison with reference genomes of choice. To fill in pathway holes or make de novo enzyme predictions, multiple computational methods such as FASTA, Whole-HMM, CSR-HMM (a method of our own introduced in this paper), and PDB-domain search are integrated in ComPath. Our experiments show that FASTA and CSR-HMM search methods generally outperform Whole-HMM and PDB-domain search methods in terms of sensitivity, but FASTA search performs poorly in terms of specificity, detecting more false positive as E-value cutoff increases. Overall, CSR-HMM search method performs best in terms of both sensitivity and specificity. Gene neighborhood and pathway neighborhood (global network) visualization tools can be used to get context information that is complementary to conventional KEGG map representation. CONCLUSION: ComPath is an interactive workbench for pathway reconstruction, annotation, and analysis where experts can perform various sequence, domain, context analysis, using an intuitive and interactive spreadsheet-style interface. Kwangmin Choi, Sun Kim |
BMC Bioinform. | 2 |
| 2008 | A gene pattern mining algorithm using interchangeable gene sets for prokaryotesabstractBACKGROUND: Mining gene patterns that are common to multiple genomes is an important biological problem, which can lead us to novel biological insights. When family classification of genes is available, this problem is similar to the pattern mining problem in the data mining community. However, when family classification information is not available, mining gene patterns is a challenging problem. There are several well developed algorithms for predicting gene patterns in a pair of genomes, such as FISH and DAGchainer. These algorithms use the optimization problem formulation which is solved using the dynamic programming technique. Unfortunately, extending these algorithms to multiple genome cases is not trivial due to the rapid increase in time and space complexity. RESULTS: In this paper, we propose a novel algorithm for mining gene patterns in more than two prokaryote genomes using interchangeable sets. The basic idea is to extend the pattern mining technique from the data mining community to handle the situation where family classification information is not available using interchangeable sets. In an experiment with four newly sequenced genomes (where the gene annotation is unavailable), we show that the gene pattern can capture important biological information. To examine the effectiveness of gene patterns further, we propose an ortholog prediction method based on our gene pattern mining algorithm and compare our method to the bi-directional best hit (BBH) technique in terms of COG orthologous gene classification information. The experiment show that our algorithm achieves a 3% increase in recall compared to BBH without sacrificing the precision of ortholog detection. CONCLUSION: The discovered gene patterns can be used for the detecting of ortholog and genes that collaborate for a common biological function. Kwangmin Choi, Sun Kim |
BMC Bioinform. | 4 |
| 2007 | Finding Cancer-Related Gene Combinations Using a Molecular Evolutionary AlgorithmabstractHigh-throughput data such as microarrays make it possible to investigate the molecular-level mechanism of cancer more efficiently. Computational methods boost the microarray analysis by managing large and complex data systematically. However, combinatorial interactions among genes have not been considered as a unit of the analysis since previous methods mainly focus on a whole gene or a single isolated gene. Here, we introduce a molecular evolutionary algorithm called probabilistic library model (PLM). In the PLM, library elements are generated from gene combinations. An evolutionary procedure is adopted to learn the probabilistic distribution of training samples. We apply the PLM to prostate cancer microarray data. The experimental results show that the PLM classifiers perform better than conventional methods such as neural networks and decision trees in accuracy. We also examine the evolved library to find cancer-related gene combinations. Chan-Hoon Park, Soo-Jin Kim, Sun Kim, Dong-Yeon Cho, Byoung-Tak Zhang |
BIBE | 3 |
| 2007 | EGGS: Extraction of Gene Clusters Using Genome Context Based Sequence Matching TechniquesabstractFunctionally related genes co-evolve, probably due to selection pressures during evolution, This phenomenon leads to conservation of gene clusters across genomes, especially in microbial genomes. In this paper, we propose novel iterative constraint relaxation algorithms which make use of genome contexts to effectively remove noise and extract gene clusters: PairEGGS that generates gene clusters in a pair of genomes and MultiEGGS that combines gene clusters from genome pairs. Experiments showed that PairEGGS produced significantly larger gene clusters than existing algorithms, say FISH, and MultiEGGS was able to find gene clusters as large as of 118 genes that are common to three genomes. Both PairEGGS and MultiEGGS run fast enough to provide service on the web. Sun Kim, Ankita Bhan, Bharath K. Maryada, Kwangmin Choi, Yves V. Brun |
BIBM | 1 |
| 2007 | Evolving hypernetwork classifiers for microRNA expression profile analysisabstractHigh-throughput microarrays inform us on different outlooks of the molecular mechanisms underlying the function of cells and organisms. While computational analysis for the microarrays show good performance, it is still difficult to infer modules of multiple co-regulated genes. Here, we present a novel classification method to identify the gene modules associated with cancers from microarray data. The proposed approach is based on 'hypernetworks', a hypergraph model consisting of vertices and weighted hyperedges. The hypernetwork model is inspired by biological networks and its learning process is suitable for identifying interacting gene modules. Applied to the analysis of microRNA (miRNA) expression profiles on multiple human cancers, the hypernetwork classifiers identified cancer-related miRNA modules. The results show that our method performs better than decision trees and naive Bayes. The biological meaning of the discovered miRNA modules has been examined by literature search. Sun Kim, Soo-Jin Kim, Byoung-Tak Zhang |
IEEE Congress on Evolutionary Computation | 1 |
| 2007 | CLASSEQ: Classification of Sequences via Comparative Analysis of Multiple GenomesabstractCLASSEQ is a Web-based system for the analysis and comparison of uncharacterized protein sequences against multiple genomes. The user sequences are combined with protein sequences from the user-specified genomes and then clustered using our in-house fast clustering algorithm, BAG. The pre-computed genome-to-genome pairwise comparison database, PCDB, makes our service fast enough to be provided on the Web even though the analysis typically involves tens of thousands of sequences. Clusters containing the user input sequences can be further characterized by domain search, multiple sequence alignment, phylogenetic tree analysis, and gene neighborhood analysis. This Web service is a useful resource for characterizing proteins of unknown functions via comparative genomics approach. CLASSEQ is available at http://platcom.org/CLASSEQ. Kwangmin Choi, Youngik Yang, Sun Kim |
ICMLA | 3 |
| 2007 | dPattern: transcription factor binding site (TFBS) discovery in human genome using a discriminative pattern analysisabstractAbstract Motivation: Transcription factor binding sites (TFBSs) are typically short in length, thus search with a profile model from known TFBSs produces many false positives. When combined with additional information, gene expression data in this article, sensitivity and specificity of TFBS search can be improved significantly. Results: By modifying our previous REFINEMENT approach, we developed dPattern that searches for occurrences of TFBSs in the promotor regions of up/down regulated or random genes. Availability: http://platcom.org/projects/dpattern Contact: [email protected] or [email protected] Seung-Hee Bae, Haixu Tang, Sun Kim |
Bioinform. | 5 |
| 2006 | Text Classifiers Evolved on a Simulated DNA ComputerabstractThe use of synthetic DNA molecules for computing provides various insights to evolutionary computation. A molecular computing algorithm to evolve DNA-encoded genetic patterns has been previously reported in [1], [2]. Here we improve on the previous work by studying the convergence behavior of the molecular evolutionary algorithm in the context of text classification problems. In particular, we study the error reduction behavior of the evolutionary learning algorithm, both theoretically and experimentally. The individuals represent decision lists of variable length and the whole population takes part in making probabilistic decisions. The evolutionary process is to change each individual towards correct classification of training data, which is based on an error minimization strategy. The evolved molecular classifiers show a performance competitive to the standard algorithms such as naïve Bayes and neural network classifiers on the data set we studied. The possibility of molecular implementation by use of DNA-encoded individuals combined with simple molecular operations on a very big population distinguishes this approach from other existing evolutionary algorithms. Sun Kim, Min-Oh Heo, Byoung-Tak Zhang |
IEEE Congress on Evolutionary Computation | 1 |
| 2006 | Prediction of the Human Papillomavirus Risk Types Using Gap-Spectrum Kernels
Sun Kim, Jae-Hong Eom |
ISNN (2) | 1 |
| 2006 | COMPAM : visualization of combining pairwise alignments for multiple genomesabstractUNLABELLED: COMPAM is a tool for visualizing relationships among multiple whole genomes by combining all pairwise genome alignments. It displays shared conserved regions (blocks) and where these blocks occur (edges) as block relation graphs which can be explored interactively. An unannotated genome, e.g. can then be explored using information from well-annotated genomes, COG-based genome annotation and genes. COMPAM can run either as a stand-alone application or through an applet that is provided as service to PLATCOM, a toolset for whole genome comparative analysis, where a wide variety of genomes can be easily selected. Features provided by COMPAM include the ability to export genome relationship information into file formats that can be used by other existing tools. AVAILABILITY: http://bio.informatics.indiana.edu/projects/compam/ Do-Hoon Lee, Jeong-Hyeon Choi, Mehmet M. Dalkilic, Sun Kim |
Bioinform. | 4 |
| 2006 | A mixture model-based discriminate analysis for identifying ordered transcription factor binding site pairs in gene promoters directly regulated by estrogen receptor-alphaabstractMOTIVATION: To detect and select patterns of transcription factor binding sites (TFBSs) which distinguish genes directly regulated by estrogen receptor-alpha (ERalpha), we developed an innovative mixture model-based discriminate analysis for identifying ordered TFBS pairs. RESULTS: Biologically, our proposed new algorithm clearly suggests that TFBSs are not randomly distributed within ERalpha target promoters (P-value < 0.001). The up-regulated targets significantly (P-value < 0.01) possess TFBS pairs, (DBP, MYC), (DBP, MYC/MAX heterodimer), (DBP, USF2) and (DBP, MYOGENIN); and down-regulated ERalpha target genes significantly (P-value < 0.01) possess TFBS pairs, such as (DBP, c-ETS1-68), (DBP, USF2) and (DBP, MYOGENIN). Statistically, our proposed mixture model-based discriminate analysis can simultaneously perform TFBS pattern recognition, TFBS pattern selection, and target class prediction; such integrative power cannot be achieved by current methods. AVAILABILITY: The software is available on request from the authors. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Lang Li 0001, Alfred S. L. Cheng, Victor X. Jin, Henry H. Paik, Meiyun Fan, Xiaoman Shawn Li, Jason Robarge, Curtis Balch, Ramana V. Davuluri, Sun Kim, Tim Hui-Ming Huang, Kenneth P. Nephew |
Bioinform. | 11 |
| 2006 | ARCS: an aggregated related column scoring scheme for aligned sequencesabstractMOTIVATION: Biologists frequently align multiple biological sequences to determine consensus sequences and/or search for predominant residues and conserved regions. Particularly, determining conserved regions in an alignment is one of the most important activities. Since protein sequences are often several-hundred residues or longer, it is difficult to distinguish biologically important conserved regions (motifs or domains) from others. The widely used tools, Logos, Al2co, Confind, and the entropy-based method, often fail to highlight such regions. Thus a computational tool that can highlight biologically important regions accurately will be highly desired. RESULTS: This paper presents a new scoring scheme ARCS (Aggregated Related Column Score) for aligned biological sequences. ARCS method considers not only the traditional character similarity measure but also column correlation. In an extensive experimental evaluation using 533 PROSITE patterns, ARCS is able to highlight the motif regions with up to 77.7% accuracy corresponding to the top three peaks. AVAILABILITY: The source code is available on http://bio.informatics.indiana.edu/projects/arcs and http://goldengate.case.edu/projects/arcs Jeong-Hyeon Choi, Guangyu Chen, Jacek Szymanski, Guo-Qiang Zhang 0001, Anthony K. H. Tung, Jaewoo Kang, Sun Kim, Jiong Yang 0001 |
Bioinform. | 8 |
| 2005 | PLATCOM: a Platform for Computational Comparative GenomicsabstractMOTIVATION: As more whole genome sequences become available, comparing multiple genomes at the sequence level can provide insight into new biological discovery. However, there are significant challenges for genome comparison. The challenge includes requirement for computational resources owing to the large volume of genome data. More importantly, since the choice of genomes to be compared is entirely subjective, there are too many choices for genome comparison. For these reasons, there is pressing need for bioinformatics systems for comparing multiple genomes where users can choose genomes to be compared freely. RESULTS: PLATCOM (Platform for Computational Comparative Genomics) is an integrated system for the comparative analysis of multiple genomes. The system is built on several public databases and a suite of genome analysis applications are provided as exemplary genome data mining tools over these internal databases. Researchers are able to visually investigate genomic sequence similarities, conserved gene neighborhoods, conserved metabolic pathways and putative gene fusion events among a set of selected multiple genomes. AVAILABILITY: http://platcom.informatics.indiana.edu/platcom Kwangmin Choi, Jeong-Hyeon Choi, Sun Kim |
Bioinform. | 4 |
| 2004 | Multi-objective Evolutionary Probe Design Based on Thermodynamic Criteria for HPV Detection
In-Hee Lee 0001, Sun Kim, Byoung-Tak Zhang |
PRICAI | 2 |
| 2003 | Genetic Mining of HTML Structures for Effective Web-Document Retrieval
Sun Kim, Byoung-Tak Zhang |
Appl. Intell. | 1 |
| 2001 | Evolutionary learning of Web-document structure for information retrievalabstractWeb documents have a number of tags indicating the structure of documents. The tag information can be utilized to improve the performance of document retrieval systems. The authors propose an approach to retrieve Web documents using HTML tags and then use a genetic algorithm to adapt the tag weights. This method uses a modified similarity measure based on the tag weights. A genetic learning method is used to select the tags for retrieval and get the optimal tag weights. We evaluate our method via experiments on conference pages and TREC document sets. The experimental results show that the tag weights are well trained by the proposed algorithm in accordance with the importance factors for retrieval. The proposed method has achieved about 10% improvement in retrieval accuracy. Sun Kim, Byoung-Tak Zhang |
CEC | 1 |
| 1994 | ModGen: Theorem Proving by Model Generation
Sun Kim, Hantao Zhang 0001 |
AAAI | 1 |