VLDB 2026 Research / reviewers in the wild / expert
Dongsup Kim
dblp:75/4603
· DBLP profile ↗
25ranked-venue papers
0as first author
4since 2021 · last 2026
0000-0002-5916-6799ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 25 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | scMix: learning temporal dynamics of gene expression under irregular time intervalsabstractMOTIVATION: Understanding temporal gene expression is fundamental in the study of cellular development and differentiation. In practice, temporal single-cell datasets tend to contain only a limited number of measured time points, which are often unevenly spaced, resulting in irregular intervals between observations due to experimental constraints. Existing methods typically address these intervals by sequentially predicting one time point after another, yet lack mechanisms to explicitly model time intervals, leading to error accumulation. RESULTS: In this work, we introduce scMix, a language-model-based framework for predicting single-cell gene expression, which enables prediction from multiple historical time points. We build scMix on the Receptance Weighted Key Value architecture and use its time decay mechanism to model temporal dependencies over time. Moreover, scMix proposes a delta-time mechanism that allows the model to bypass unmeasured time points, reducing error accumulation and improving robustness. In addition, we incorporate a trend regularization strategy to enhance the temporal coherence of predicted gene expression trajectories. scMix demonstrates state-of-the-art performance in predicting gene expression at unmeasured time points, surpassing existing methods, and also achieves outstanding results on downstream tasks. AVAILABILITY AND IMPLEMENTATION: The code used for this study is available at https://doi.org/10.5281/zenodo.18287184. Shangjin Han, Dongsup Kim |
Bioinform. | 2 |
| 2024 | AbFlex: designing antibody complementarity determining regions with flexible CDR definitionabstractMOTIVATION: Antibodies are proteins that the immune system produces in response to foreign pathogens. Designing antibodies that specifically bind to antigens is a key step in developing antibody therapeutics. The complementarity determining regions (CDRs) of the antibody are mainly responsible for binding to the target antigen, and therefore must be designed to recognize the antigen. RESULTS: We develop an antibody design model, AbFlex, that exhibits state-of-the-art performance in terms of structure prediction accuracy and amino acid recovery rate. Furthermore, >38% of newly designed antibody models are estimated to have better binding energies for their antigens than wild types. The effectiveness of the model is attributed to two different strategies that are developed to overcome the difficulty associated with the scarcity of antibody-antigen complex structure data. One strategy is to use an equivariant graph neural network model that is more data-efficient. More importantly, a new data augmentation strategy based on the flexible definition of CDRs significantly increases the performance of the CDR prediction model. AVAILABILITY AND IMPLEMENTATION: The source code and implementation are available at https://github.com/wsjeon92/AbFlex. Woosung Jeon, Dongsup Kim |
Bioinform. | 2 |
| 2024 | TSpred: a robust prediction framework for TCR-epitope interactions using paired chain TCR sequence dataabstractMOTIVATION: Prediction of T-cell receptor (TCR)-epitope interactions is important for many applications in biomedical research, such as cancer immunotherapy and vaccine design. The prediction of TCR-epitope interactions remains challenging especially for novel epitopes, due to the scarcity of available data. RESULTS: We propose TSpred, a new deep learning approach for the pan-specific prediction of TCR binding specificity based on paired chain TCR data. We develop a robust model that generalizes well to unseen epitopes by combining the predictive power of CNN and the attention mechanism. In particular, we design a reciprocal attention mechanism which focuses on extracting the patterns underlying TCR-epitope interactions. Upon a comprehensive evaluation of our model, we find that TSpred achieves state-of-the-art performances in both seen and unseen epitope specificity prediction tasks. Also, compared to other predictors, TSpred is more robust to bias related to peptide imbalance in the dataset. In addition, the reciprocal attention component of our model allows for model interpretability by capturing structurally important binding regions. Results indicate that TSpred is a robust and reliable method for the task of TCR-epitope binding prediction. AVAILABILITY AND IMPLEMENTATION: Source code is available at https://github.com/ha01994/TSpred. Ha Young Kim, Sungsik Kim, Woong-Yang Park, Dongsup Kim |
Bioinform. | 4 |
| 2022 | DeepLUCIA: predicting tissue-specific chromatin loops using Deep Learning-based Universal Chromatin Interaction AnnotatorabstractMOTIVATION: The importance of chromatin loops in gene regulation is broadly accepted. There are mainly two approaches to predict chromatin loops: transcription factor (TF) binding-dependent approach and genomic variation-based approach. However, neither of these approaches provides an adequate understanding of gene regulation in human tissues. To address this issue, we developed a deep learning-based chromatin loop prediction model called Deep Learning-based Universal Chromatin Interaction Annotator (DeepLUCIA). RESULTS: Although DeepLUCIA does not use TF binding profile data which previous TF binding-dependent methods critically rely on, its prediction accuracies are comparable to those of the previous TF binding-dependent methods. More importantly, DeepLUCIA enables the tissue-specific chromatin loop predictions from tissue-specific epigenomes that cannot be handled by genomic variation-based approach. We demonstrated the utility of the DeepLUCIA by predicting several novel target genes of SNPs identified in genome-wide association studies targeting Brugada syndrome, COVID-19 severity and age-related macular degeneration. Availability and implementation DeepLUCIA is freely available at https://github.com/bcbl-kaist/DeepLUCIA. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Dongchan Yang, Taesu Chung, Dongsup Kim |
Bioinform. | 3 |
| 2020 | Prediction of mutation effects using a deep temporal convolutional networkabstractMOTIVATION: Accurate prediction of the effects of genetic variation is a major goal in biological research. Towards this goal, numerous machine learning models have been developed to learn information from evolutionary sequence data. The most effective method so far is a deep generative model based on the variational autoencoder (VAE) that models the distributions using a latent variable. In this study, we propose a deep autoregressive generative model named mutationTCN, which employs dilated causal convolutions and attention mechanism for the modeling of inter-residue correlations in a biological sequence. RESULTS: We show that this model is competitive with the VAE model when tested against a set of 42 high-throughput mutation scan experiments, with the mean improvement in Spearman rank correlation ∼0.023. In particular, our model can more efficiently capture information from multiple sequence alignments with lower effective number of sequences, such as in viral sequence families, compared with the latent variable model. Also, we extend this architecture to a semi-supervised learning framework, which shows high prediction accuracy. We show that our model enables a direct optimization of the data likelihood and allows for a simple and stable training process. AVAILABILITY AND IMPLEMENTATION: Source code is available at https://github.com/ha01994/mutationTCN. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ha Young Kim, Dongsup Kim |
Bioinform. | 2 |
| 2020 | CRDS: Consensus Reverse Docking System for target fishingabstractMOTIVATION: Identification of putative drug targets is a critical step for explaining the mechanism of drug action against multiple targets, finding new therapeutic indications for existing drugs and unveiling the adverse drug reactions. One important approach is to use the molecular docking. However, its widespread utilization has been hindered by the lack of easy-to-use public servers. Therefore, it is vital to develop a streamlined computational tool for target prediction by molecular docking on a large scale. RESULTS: We present a fully automated web tool named Consensus Reverse Docking System (CRDS), which predicts potential interaction sites for a given drug. To improve hit rates, we developed a strategy of consensus scoring. CRDS carries out reverse docking against 5254 candidate protein structures using three different scoring functions (GoldScore, Vina and LeDock from GOLD version 5.7.1, AutoDock Vina version 1.1.2 and LeDock version 1.0, respectively), and those scores are combined into a single score named Consensus Docking Score (CDS). The web server provides the list of top 50 predicted interaction sites, docking conformations, 10 most significant pathways and the distribution of consensus scores. AVAILABILITY AND IMPLEMENTATION: The web server is available at http://pbil.kaist.ac.kr/CRDS. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Aeri Lee, Dongsup Kim |
Bioinform. | 2 |
| 2019 | FP2VEC: a new molecular featurizer for learning molecular propertiesabstractMOTIVATION: One of the most successful methods for predicting the properties of chemical compounds is the quantitative structure-activity relationship (QSAR) methods. The prediction accuracy of QSAR models has recently been greatly improved by employing deep learning technology. Especially, newly developed molecular featurizers based on graph convolution operations on molecular graphs significantly outperform the conventional extended connectivity fingerprints (ECFP) feature in both classification and regression tasks, indicating that it is critical to develop more effective new featurizers to fully realize the power of deep learning techniques. Motivated by the fact that there is a clear analogy between chemical compounds and natural languages, this work develops a new molecular featurizer, FP2VEC, which represents a chemical compound as a set of trainable embedding vectors. RESULTS: To implement and test our new featurizer, we build a QSAR model using a simple convolutional neural network (CNN) architecture that has been successfully used for natural language processing tasks such as sentence classification task. By testing our new method on several benchmark datasets, we demonstrate that the combination of FP2VEC and CNN model can achieve competitive results in many QSAR tasks, especially in classification tasks. We also demonstrate that the FP2VEC model is especially effective for multitask learning. AVAILABILITY AND IMPLEMENTATION: FP2VEC is available from https://github.com/wsjeon92/FP2VEC. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Woosung Jeon, Dongsup Kim |
Bioinform. | 2 |
| 2017 | Deep convolutional neural networks for pan-specific peptide-MHC class I binding predictionabstractBACKGROUND: Computational scanning of peptide candidates that bind to a specific major histocompatibility complex (MHC) can speed up the peptide-based vaccine development process and therefore various methods are being actively developed. Recently, machine-learning-based methods have generated successful results by training large amounts of experimental data. However, many machine learning-based methods are generally less sensitive in recognizing locally-clustered interactions, which can synergistically stabilize peptide binding. Deep convolutional neural network (DCNN) is a deep learning method inspired by visual recognition process of animal brain and it is known to be able to capture meaningful local patterns from 2D images. Once the peptide-MHC interactions can be encoded into image-like array(ILA) data, DCNN can be employed to build a predictive model for peptide-MHC binding prediction. In this study, we demonstrated that DCNN is able to not only reliably predict peptide-MHC binding, but also sensitively detect locally-clustered interactions. RESULTS: Nonapeptide-HLA-A and -B binding data were encoded into ILA data. A DCNN, as a pan-specific prediction model, was trained on the ILA data. The DCNN showed higher performance than other prediction tools for the latest benchmark datasets, which consist of 43 datasets for 15 HLA-A alleles and 25 datasets for 10 HLA-B alleles. In particular, the DCNN outperformed other tools for alleles belonging to the HLA-A3 supertype. The F1 scores of the DCNN were 0.86, 0.94, and 0.67 for HLA-A*31:01, HLA-A*03:01, and HLA-A*68:01 alleles, respectively, which were significantly higher than those of other tools. We found that the DCNN was able to recognize locally-clustered interactions that could synergistically stabilize peptide binding. We developed ConvMHC, a web server to provide user-friendly web interfaces for peptide-MHC class I binding predictions using the DCNN. ConvMHC web server can be accessible via http://jumong.kaist.ac.kr:8080/convmhc . CONCLUSIONS: We developed a novel method for peptide-HLA-I binding predictions using DCNN trained on ILA data that encode peptide binding data and demonstrated the reliable performance of the DCNN in nonapeptide binding predictions through the independent evaluation on the latest IEDB benchmark datasets. Our approaches can be applied to characterize locally-clustered patterns in molecular interactions, such as protein/DNA, protein/RNA, and drug/protein interactions. Youngmahn Han, Dongsup Kim |
BMC Bioinform. | 2 |
| 2017 | Utilizing random Forest QSAR models with optimized parameters for target identification and its application to target-fishing serverabstractBACKGROUND: The identification of target molecules is important for understanding the mechanism of "target deconvolution" in phenotypic screening and "polypharmacology" of drugs. Because conventional methods of identifying targets require time and cost, in-silico target identification has been considered an alternative solution. One of the well-known in-silico methods of identifying targets involves structure activity relationships (SARs). SARs have advantages such as low computational cost and high feasibility; however, the data dependency in the SAR approach causes imbalance of active data and ambiguity of inactive data throughout targets. RESULTS: We developed a ligand-based virtual screening model comprising 1121 target SAR models built using a random forest algorithm. The performance of each target model was tested by employing the ROC curve and the mean score using an internal five-fold cross validation. Moreover, recall rates for top-k targets were calculated to assess the performance of target ranking. A benchmark model using an optimized sampling method and parameters was examined via external validation set. The result shows recall rates of 67.6% and 73.9% for top-11 (1% of the total targets) and top-33, respectively. We provide a website for users to search the top-k targets for query ligands available publicly at http://rfqsar.kaist.ac.kr . CONCLUSIONS: The target models that we built can be used for both predicting the activity of ligands toward each target and ranking candidate targets for a query ligand using a unified scoring scheme. The scores are additionally fitted to the probability so that users can estimate how likely a ligand-target interaction is active. The user interface of our web site is user friendly and intuitive, offering useful information and cross references. Kyoungyeul Lee, Minho Lee 0002, Dongsup Kim |
BMC Bioinform. | 3 |
| 2016 | Library of binding protein scaffolds (LibBP): a computational platform for selection of binding protein scaffoldsabstractMOTIVATION: Developments in biotechnology have enabled the in vitro evolution of binding proteins. The emerging limitations of antibodies in binding protein engineering have led to suggestions for other proteins as alternative binding protein scaffolds. Most of these proteins were selected based on human intuition rather than systematic analysis of the available data. To improve this strategy, we developed a computational framework for finding desirable binding protein scaffolds by utilizing protein structure and sequence information. RESULTS: For each protein, its structure and the sequences of evolutionarily-related proteins were analyzed, and spatially contiguous regions composed of highly variable residues were identified. A large number of proteins have these regions, but leucine rich repeats (LRRs), histidine kinase domains and immunoglobulin domains are predominant among them. The candidates suggested as new binding protein scaffolds include histidine kinase, LRR, titin and pentapeptide repeat protein. AVAILABILITY AND IMPLEMENTATION: The database and web-service are accessible via http://bcbl.kaist.ac.kr/LibBP CONTACT: [email protected] data: Supplementary data are available at Bioinformatics online. Seungpyo Hong, Dongsup Kim |
Bioinform. | 2 |
| 2016 | Structure-based Markov random field model for representing evolutionary constraints on functional sitesabstractBACKGROUND: Elucidating the cooperative mechanism of interconnected residues is an important component toward understanding the biological function of a protein. Coevolution analysis has been developed to model the coevolutionary information reflecting structural and functional constraints. Recently, several methods have been developed based on a probabilistic graphical model called the Markov random field (MRF), which have led to significant improvements for coevolution analysis; however, thus far, the performance of these models has mainly been assessed by focusing on the aspect of protein structure. RESULTS: In this study, we built an MRF model whose graphical topology is determined by the residue proximity in the protein structure, and derived a novel positional coevolution estimate utilizing the node weight of the MRF model. This structure-based MRF method was evaluated for three data sets, each of which annotates catalytic site, allosteric site, and comprehensively determined functional site information. We demonstrate that the structure-based MRF architecture can encode the evolutionary information associated with biological function. Furthermore, we show that the node weight can more accurately represent positional coevolution information compared to the edge weight. Lastly, we demonstrate that the structure-based MRF model can be reliably built with only a few aligned sequences in linear time. CONCLUSIONS: The results show that adoption of a structure-based architecture could be an acceptable approximation for coevolution modeling with efficient computation complexity. Chan-seok Jeong, Dongsup Kim |
BMC Bioinform. | 2 |
| 2012 | Large-scale reverse docking profiles and their applicationsabstractBACKGROUND: Reverse docking approaches have been explored in previous studies on drug discovery to overcome some problems in traditional virtual screening. However, current reverse docking approaches are problematic in that the target spaces of those studies were rather small, and their applications were limited to identifying new drug targets. In this study, we expanded the scope of target space to a set of all protein structures currently available and developed several new applications of reverse docking method. RESULTS: We generated 2D Matrix of docking scores among all the possible protein structures in yeast and human and 35 famous drugs. By clustering the docking profile data and then comparing them with fingerprint-based clustering of drugs, we first showed that our data contained accurate information on their chemical properties. Next, we showed that our method could be used to predict the druggability of target proteins. We also showed that a combination of sequence similarity and docking profile similarity could predict the enzyme EC numbers more accurately than sequence similarity alone. In two case studies, 5-fluorouracil and cycloheximide, we showed that our method can successfully find identifying target proteins. CONCLUSIONS: By using a large number of protein structures, we improved the sensitivity of reverse docking and showed that using as many protein structure as possible was important in finding real binding targets. Minho Lee 0002, Dongsup Kim |
BMC Bioinform. | 2 |
| 2011 | Modeling allosteric signal propagation using protein structure networksabstractAllosteric communication in proteins can be induced by the binding of effective ligands, mutations or covalent modifications that regulate a site distant from the perturbed region. To understand allosteric regulation, it is important to identify the remote sites that are affected by the perturbation-induced signals and how these allosteric perturbations are transmitted within the protein structure. In this study, by constructing a protein structure network and modeling signal transmission with a Markov random walk, we developed a method to estimate the signal propagation and the resulting effects. In our model, the global perturbation effects from a particular signal initiation site were estimated by calculating the expected visiting time (EVT), which describes the signal-induced effects caused by signal transmission through all possible routes. We hypothesized that the residues with high EVT values play important roles in allosteric signaling. We applied our model to two protein structures as examples, and verified the validity of our model using various types of experimental data. We also found that the hot spots in protein binding interfaces have significantly high EVT values, which suggests that they play roles in mediating signal communication between protein domains. Keunwan Park, Dongsup Kim |
BMC Bioinform. | 2 |
| 2010 | Linear predictive coding representation of correlated mutation for protein sequence alignmentabstractBACKGROUND: Although both conservation and correlated mutation (CM) are important information reflecting the different sorts of context in multiple sequence alignment, most of alignment methods use sequence profiles that only represent conservation. There is no general way to represent correlated mutation and incorporate it with sequence alignment yet. METHODS: We develop a novel method, CM profile, to represent correlated mutation as the spectral feature derived by using linear predictive coding where correlated mutations among different positions are represented by a fixed number of values. We combine CM profile with conventional sequence profile to improve alignment quality. RESULTS: For distantly related protein pairs, using CM profile improves the profile-profile alignment with or without predicted secondary structure. Especially, at superfamily level, combining CM profile with sequence profile improves profile-profile alignment by 9.5% while predicted secondary structure does by 6.0%. More significantly, using both of them improves profile-profile alignment by 13.9%. We also exemplify the effectiveness of CM profile by demonstrating that the resulting alignment preserves share coevolution and contacts. CONCLUSIONS: In this work, we introduce a novel method, CM profile, which represents correlated mutation information as paralleled form, and apply it to the protein sequence alignment problem. When combined with conventional sequence profile, CM profile improves alignment quality significantly better than predicted secondary structure information, which should be beneficial for target-template alignment in protein structure prediction. Because of the generality of CM profile, it can be used for other bioinformatics applications in the same way of using sequence profile. Chan-seok Jeong, Dongsup Kim |
BMC Bioinform. | 2 |
| 2010 | PostMod: sequence based prediction of kinase-specific phosphorylation sites with indirect relationshipabstractBACKGROUND: Post-translational modifications (PTMs) have a key role in regulating cell functions. Consequently, identification of PTM sites has a significant impact on understanding protein function and revealing cellular signal transductions. Especially, phosphorylation is a ubiquitous process with a large portion of proteins undergoing this modification. Experimental methods to identify phosphorylation sites are labor-intensive and of high-cost. With the exponentially growing protein sequence data, development of computational approaches to predict phosphorylation sites is highly desirable. RESULTS: Here, we present a simple and effective method to recognize phosphorylation sites by combining sequence patterns and evolutionary information and by applying a novel noise-reducing algorithm. We suggested that considering long-range region surrounding a phosphorylation site is important for recognizing phosphorylation peptides. Also, from compared results to AutoMotif in 36 different kinase families, new method outperforms AutoMotif. The mean accuracy, precision, and recall of our method are 0.93, 0.67, and 0.40, respectively, whereas those of AutoMotif with a polynomial kernel are 0.91, 0.47, and 0.17, respectively. Also our method shows better or comparable performance in four main kinase groups, CDK, CK2, PKA, and PKC compared to six existing predictors. CONCLUSION: Our method is remarkable in that it is powerful and intuitive approach without need of a sophisticated training algorithm. Moreover, our method is generally applicable to other types of PTMs. Inkyung Jung, Akihisa Matsuyama, Minoru Yoshida, Dongsup Kim |
BMC Bioinform. | 4 |
| 2009 | SIMPRO: simple protein homology detection method by using indirect signalsabstractMOTIVATION: Detecting homologous proteins is one of the fundamental problems in computational biology. Many tools to solve this problem have been developed, but development of a simple, effective and generally applicable method is still desirable. RESULTS: We propose a simple but effective information retrieval approach, named SIMPRO, to identify homology relationship between proteins. The key idea of our approach is that by accumulating and comparing indirect signals from conventional homology search methods, the search sensitivity can be increased. We tested the idea on the problem of detecting homology relationship between Pfam families, as well as detecting structural homologs based on SCOP, and found that our method achieved significant improvement. Our results indicate that simple manipulation of conventional homology search outputs by SIMPRO algorithm can remarkably improve homology search accuracy. Inkyung Jung, Dongsup Kim |
Bioinform. | 2 |
| 2009 | A new method for revealing correlated mutations under the structural and functional constraints in proteinsabstractMOTIVATION: Diverse studies have shown that correlated mutation (CM) is an important molecular evolutionary process alongside conservation. However, attempts to find the residue pairs that co-evolve under the structural and/or functional constraints are complicated by the fact that a large portion of covariance signals found in multiple sequence alignments arise from correlations due to common ancestry and stochastic noise. RESULTS: Assuming that the background noise can be estimated from the coevolutionary relationships among residues, we propose a new measure for background noise called the normalized coevolutionary pattern similarity (NCPS) score. By subtracting NCPS scores from raw CM scores and combining the results with an entropy factor, we show that these new scores effectively reduce the background noise. To test the effectiveness of this method in detecting residue pairs coevolving under the structural constraints, two independent test sets were performed, showing that this new method performs better than the most accurate method currently available. In addition, we also applied our method to double mutant cycle experiments and protein-protein interactions. Although more rigorous tests are required, we obtained promising results that our method tended to explain those data better than other methods. These results suggest that the new noise-reduced CM scores developed in this study can be a valuable tool for the study of correlated mutations under the structural and/or functional constraints in proteins. AVAILABILITY: http://pbil.kaist.ac.kr Dongsup Kim |
Bioinform. | 2 |
| 2008 | Application of nonnegative matrix factorization to improve profile-profile alignment features for fold recognition and remote homolog detectionabstractBACKGROUND: Nonnegative matrix factorization (NMF) is a feature extraction method that has the property of intuitive part-based representation of the original features. This unique ability makes NMF a potentially promising method for biological sequence analysis. Here, we apply NMF to fold recognition and remote homolog detection problems. Recent studies have shown that combining support vector machines (SVM) with profile-profile alignments improves performance of fold recognition and remote homolog detection remarkably. However, it is not clear which parts of sequences are essential for the performance improvement. RESULTS: The performance of fold recognition and remote homolog detection using NMF features is compared to that of the unmodified profile-profile alignment (PPA) features by estimating Receiver Operating Characteristic (ROC) scores. The overall performance is noticeably improved. For fold recognition at the fold level, SVM with NMF features recognize 30% of homolog proteins at > 0.99 ROC scores, while original PPA feature, HHsearch, and PSI-BLAST recognize almost none. For detecting remote homologs that are related at the superfamily level, NMF features also achieve higher performance than the original PPA features. At > 0.90 ROC50 scores, 25% of proteins with NMF features correctly detects remotely related proteins, whereas using original PPA features only 1% of proteins detect remote homologs. In addition, we investigate the effect of number of positive training examples and the number of basis vectors on performance improvement. We also analyze the ability of NMF to extract essential features by comparing NMF basis vectors with functionally important sites and structurally conserved regions of proteins. The results show that NMF basis vectors have significant overlap with functional sites from PROSITE and with structurally conserved regions from the multiple structural alignments generated by MUSTANG. The correlation between NMF basis vectors and biologically essential parts of proteins supports our conjecture that NMF basis vectors can explicitly represent important sites of proteins. CONCLUSION: The present work demonstrates that applying NMF to profile-profile alignments can reveal essential features of proteins and that these features significantly improve the performance of fold recognition and remote homolog detection. Inkyung Jung, Jaehyung Lee 0001, Soo-Young Lee, Dongsup Kim |
BMC Bioinform. | 4 |
| 2008 | Inference of Protein Complex Activities from Chemical-Genetic Profile and Its Applications: Predicting Drug-Target PathwaysabstractThe chemical-genetic profile can be defined as quantitative values of deletion strains' growth defects under exposure to chemicals. In yeast, the compendium of chemical-genetic profiles of genomewide deletion strains under many different chemicals has been used for identifying direct target proteins and a common mode-of-action of those chemicals. In the previous study, valuable biological information such as protein-protein and genetic interactions has not been fully utilized. In our study, we integrated this compendium and biological interactions into the comprehensive collection of approximately 490 protein complexes of yeast for model-based prediction of a drug's target proteins and similar drugs. We assumed that those protein complexes (PCs) were functional units for yeast cell growth and regarded them as hidden factors and developed the PC-based Bayesian factor model that relates the chemical-genetic profile at the level of organism phenotypes to the hidden activities of PCs at the molecular level. The inferred PC activities provided the predictive power of a common mode-of-action of drugs as well as grouping of PCs with similar functions. In addition, our PC-based model allowed us to develop a new effective method to predict a drug's target pathway, by which we were able to highlight the target-protein, TOR1, of rapamycin. Our study is the first approach to model phenotypes of systematic deletion strains in terms of protein complexes. We believe that our PC-based approach can provide an appropriate framework for combining and modeling several types of chemical-genetic profiles including interspecies. Such efforts will contribute to predicting more precisely relevant pathways including target proteins that interact directly with bioactive compounds. Sangjo Han, Dongsup Kim |
PLoS Comput. Biol. | 2 |
| 2007 | Inference of Gene Regulatory Networks Using Time Sliding Comparison and Transcriptional Lagging Time from Time Series Gene Expression ProfilesabstractInference of gene regulatory network from microarray data is one of the most important issues in bioinformatics. Several algorithms have been introduced for this problem, but they cannot give accurate results in case of time series gene expression data. Here, we propose an algorithm that predicts the gene regulatory network more accurately than the previous methods. A new method finds the relationship between a pair of genes by time-shifting the time series data of one gene against another when we compare the patterns of the gene pair. In addition, we increase the accuracy of the prediction method by filtering out the interactions that cannot exist in real biological network. We tested the algorithm to several simulated data which are based on realistic enzyme kinetics system, and evaluated the effectiveness of our algorithm. Results show that the present algorithm significantly improves the accuracy of the inference of gene regulatory network. Sheehyun Kim, Dongsup Kim |
BIBE | 2 |
| 2007 | Predicting and improving the protein sequence alignment quality by support vector regressionabstractBACKGROUND: For successful protein structure prediction by comparative modeling, in addition to identifying a good template protein with known structure, obtaining an accurate sequence alignment between a query protein and a template protein is critical. It has been known that the alignment accuracy can vary significantly depending on our choice of various alignment parameters such as gap opening penalty and gap extension penalty. Because the accuracy of sequence alignment is typically measured by comparing it with its corresponding structure alignment, there is no good way of evaluating alignment accuracy without knowing the structure of a query protein, which is obviously not available at the time of structure prediction. Moreover, there is no universal alignment parameter option that would always yield the optimal alignment. RESULTS: In this work, we develop a method to predict the quality of the alignment between a query and a template. We train the support vector regression (SVR) models to predict the MaxSub scores as a measure of alignment quality. The alignment between a query protein and a template of length n is transformed into a (n + 1)-dimensional feature vector, then it is used as an input to predict the alignment quality by the trained SVR model. Performance of our work is evaluated by various measures including Pearson correlation coefficient between the observed and predicted MaxSub scores. Result shows high correlation coefficient of 0.945. For a pair of query and template, 48 alignments are generated by changing alignment options. Trained SVR models are then applied to predict the MaxSub scores of those and to select the best alignment option which is chosen specifically to the query-template pair. This adaptive selection procedure results in 7.4% improvement of MaxSub scores, compared to those when the single best parameter option is used for all query-template pairs. CONCLUSION: The present work demonstrates that the alignment quality can be predicted with reasonable accuracy. Our method is useful not only for selecting the optimal alignment parameters for a chosen template based on predicted alignment quality, but also for filtering out problematic templates that are not suitable for structure prediction due to poor alignment accuracy. This is implemented as a part in FORECAST, the server for fold-recognition and is freely available on the web at http://pbil.kaist.ac.kr/forecast. Minho Lee 0002, Chan-seok Jeong, Dongsup Kim |
BMC Bioinform. | 3 |
| 2006 | AtRTPrimer: database for Arabidopsis genome-wide homogeneous and specific RT-PCR primer-pairsabstractBACKGROUND: Primer design is a critical step in all types of RT-PCR methods to ensure specificity and efficiency of a target amplicon. However, most traditional primer design programs suggest primers on a single template of limited genetic complexity. To provide researchers with a sufficient number of pre-designed specific RT-PCR primer pairs for whole genes in Arabidopsis, we aimed to construct a genome-wide primer-pair database. DESCRIPTION: We considered the homogeneous physical and chemical properties of each primer (homogeneity) of a gene, non-specific binding against all other known genes (specificity), and other possible amplicons from its corresponding genomic DNA or similar cDNAs (additional information). Then, we evaluated the reliability of our database with selected primer pairs from 15 genes using conventional and real time RT-PCR. CONCLUSION: Approximately 97% of 28,952 genes investigated were finally registered in AtRTPrimer. Unlike other freely available primer databases for Arabidopsis thaliana, AtRTPrimer provides a large number of reliable primer pairs for each gene so that researchers can perform various types of RT-PCR experiments for their specific needs. Furthermore, by experimentally evaluating our database, we made sure that our database provides good starting primer pairs for Arabidopsis researchers to perform various types of RT-PCR experiments. Sangjo Han, Dongsup Kim |
BMC Bioinform. | 2 |
| 2005 | Fold recognition by combining profile-profile alignment and support vector machineabstractMOTIVATION: Currently, the most accurate fold-recognition method is to perform profile-profile alignments and estimate the statistical significances of those alignments by calculating Z-score or E-value. Although this scheme is reliable in recognizing relatively close homologs related at the family level, it has difficulty in finding the remote homologs that are related at the superfamily or fold level. RESULTS: In this paper, we present an alternative method to estimate the significance of the alignments. The alignment between a query protein and a template of length n in the fold library is transformed into a feature vector of length n + 1, which is then evaluated by support vector machine (SVM). The output from SVM is converted to a posterior probability that a query sequence is related to a template, given SVM output. Results show that a new method shows significantly better performance than PSI-BLAST and profile-profile alignment with Z-score scheme. While PSI-BLAST and Z-score scheme detect 16 and 20% of superfamily-related proteins, respectively, at 90% specificity, a new method detects 46% of these proteins, resulting in more than 2-fold increase in sensitivity. More significantly, at the fold level, a new method can detect 14% of remotely related proteins at 90% specificity, a remarkable result considering the fact that the other methods can detect almost none at the same level of specificity. Sangjo Han, Seung Taek Yu, Chan-seok Jeong, Dongsup Kim |
Bioinform. | 6 |
| 2003 | A Computational Pipeline for Protein Structure Prediction and Analysis at Genome ScaleabstractTraditionally, protein 3D structures are solved using experimental techniques, like X-ray crystallography or nuclear magnetic resonance (NMR). While these experimental techniques have been the main workhorse for protein structure studies in the past few decades, it is becoming increasingly apparent that they alone cannot keep up with the production rate of protein sequences. Fortunately, computational techniques for protein structure predictions have matured to such a level that they can complement the existing experimental techniques. In this paper, we present an automated pipeline for protein structure prediction. The centerpiece of the pipeline is a threading-based protein structure prediction system, called PROSPECT, which we have been developing for the past few years. The pipeline consists of seven logical phases, utilizing a dozen tools. The pipeline has been implemented to run in a heterogeneous computational environment as a client/server system with a web interface. A number of genome-scale applications have been carried out on microbial genomes. Here we present one genome-scale application on Caenorhabditis elegans. Manesh J. Shah, Sergei Passovets, Dongsup Kim, Kyle Ellrott, Li Wang 0008, Inna Vokler, Philip F. LoCascio, Dong Xu 0002, Ying Xu 0001 |
BIBE | 3 |
| 2003 | A computational pipeline for protein structure prediction and analysis at genome scaleabstractMOTIVATION: Experimental techniques alone cannot keep up with the production rate of protein sequences, while computational techniques for protein structure predictions have matured to such a level to provide reliable structural characterization of proteins at large scale. Integration of multiple computational tools for protein structure prediction can complement experimental techniques. RESULTS: We present an automated pipeline for protein structure prediction. The centerpiece of the pipeline is our threading-based protein structure prediction system PROSPECT. The pipeline consists of a dozen tools for identification of protein domains and signal peptide, protein triage to determine the protein type (membrane or globular), protein fold recognition, generation of atomic structural models, prediction result validation, etc. Different processing and prediction branches are determined automatically by a prediction pipeline manager based on identified characteristics of the protein. The pipeline has been implemented to run in a heterogeneous computational environment as a client/server system with a web interface. Genome-scale applications on Caenorhabditis elegans, Pyrococcus furiosus and three cyanobacterial genomes are presented. AVAILABILITY: The pipeline is available at http://compbio.ornl.gov/proteinpipeline/ Manesh J. Shah, Sergei Passovets, Dongsup Kim, Kyle Ellrott, Li Wang 0008, Inna Vokler, Philip F. LoCascio, Dong Xu 0002, Ying Xu 0001 |
Bioinform. | 3 |