Kyungsook Han

dblp:62/6821 · DBLP profile ↗
← Back
60ranked-venue papers
9as first author
11since 2021 · last 2025
0000-0001-9900-6741ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 52 · 5 first-author · 11 since 2021Artificial intelligence and machine learning · 6 · 3 first-authorDatabases, data management, data science and information retrieval · 3 · 1 first-authorTheory of computation · 3 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author
YearPublicationVenuePosition
2025 Predicting Cancer Metastasis From DNA Methylation and Gene Expression Profiles
abstract
Metastasis is the major cause of cancer-related mortality, accounting for about 90% of cancer deaths. So far, most computational methods for predicting metastasis relied on gene expression data or relation between genes. Motivated by an increasing evidence of inter-person variations in gene expression and DNA methylation, we developed a new method for predicting metastasis based on gene expression and DNA methylation profiles. We derived differential correlations between gene expression and DNA methylation in every tumor sample with or without metastasis. Using the differential correlations, we constructed a logistic regression model for predicting metastasis. The prediction model showed a very high performance both in lymph node metastasis and in distant metastasis. In comparison of our method with other recent methods for predicting metastasis, our method showed a much better performance. Interestingly, using DNA methylation beta values alone showed a reasonably high performance as well. When combining differential correlations between gene expression and DNA methylation with DNA methylation, the performance was improved in most performance measures. Our method can be used as useful aids in predicting metastasis, which in turn will help determine treatment options for cancer patients.
Myeonghun Cho, Jiahui Kang, Kyungsook Han
IEEE Trans. Comput. Biol. Bioinform.4
2023 Constructing a Cancer Patient-Specific Network Based on Second-Order Partial Correlations of Gene Expression and DNA Methylation
abstract
Typically patient-specific gene networks are constructed with gene expression data only. Such networks cannot distinguish direct gene interactions from indirect interactions via others such as the effect of epigenetic events to gene activity. There is an increasing evidence of inter-individual variations not only in gene expression but also in epigenetic events such as DNA methylation. In this paper we propose a new method for constructing a cancer patient-specific gene correlation network using both gene expression and DNA methylation data. We derive a patient-specific network from differential second-order partial correlations of gene expression and DNA methylation between normal samples and the patient sample. The network represents direct interactions between genes by controlling the effect of DNA methylation. Using this method, we constructed 4,000 patient-specific networks for 10 types of cancer. The networks are highly effective in classifying different types of cancer and in deriving potential prognostic gene pairs. In particular, potential prognostic gene pairs derived from the networks were powerful in predicting the survival time of cancer patients. This approach will help identify patient-specific gene correlations and predict prognosis of cancer patients.
Wook Lee, Seokwoo Lee, Kyungsook Han
IEEE ACM Trans. Comput. Biol. Bioinform.3
2023 Constructing Integrative ceRNA Networks and Finding Prognostic Biomarkers in Renal Cell Carcinoma
abstract
Inspired by a newly discovered gene regulation mechanism known as competing endogenous RNA (ceRNA) interactions, several computational methods have been proposed to generate ceRNA networks. However, most of these methods have focused on deriving restricted types of ceRNA interactions such as lncRNA-miRNA-mRNA interactions. Competition for miRNA-binding occurs not only between lncRNAs and mRNAs but also between lncRNAs or between mRNAs. Furthermore, a large number of pseudogenes also act as ceRNAs, thereby regulate other genes. In this study, we developed a general method for constructing integrative networks of all possible interactions of ceRNAs in renal cell carcinoma (RCC). From the ceRNA networks we derived potential prognostic biomarkers, each of which is a triplet of two ceRNAs and miRNA (i.e., ceRNA-miRNA-ceRNA). Interestingly, some prognostic ceRNA triplets do not include mRNA at all, and consist of two non-coding RNAs and miRNA, which have been rarely known so far. Comparison of the prognostic ceRNA triplets to known prognostic genes in RCC showed that the triplets have a better predictive power of survival rates than the known prognostic genes. Our approach will help us construct integrative networks of ceRNAs of all types and find new potential prognostic biomarkers in cancer.
Seokwoo Lee, Wook Lee, Shulei Ren, Kyungsook Han
IEEE ACM Trans. Comput. Biol. Bioinform.5
2022 Predicting Lymph Node Metastasis and Distant Metastasis using Differential Correlations of miRNAs and Their Target RNAs in Cancer
abstract
As the most common cause of cancer death, metastasis is a complex process that involves the spread of cancer cells from the original site to other parts of the body. Diagnosis of metastasis is usually confirmed by clinical examinations and imaging, but such diagnosis is made after metastasis occurs. Early detection of metastasis plays an important role in treatment planning, which in turn has an impact on the survival of patients. So far a few methods have been developed to predict lymph node metastasis, but few methods are available for predicting distant metastasis. Motivated by a recently known gene regulation mechanism involving miRNAs, we developed a new method for predicting both lymph node metastasis and distant metastasis. We identified differential correlations of miRNAs and their target RNAs in cancer, and built prediction models using the differential correlations. Testing the method on several types of cancer showed that differential correlations of miRNAs and their target RNAs are much more powerful than expressions of known metastasis predictive genes in predicting distant metastasis as well as lymph node metastasis. Although preliminary, the method developed in this study will be useful in predicting metastasis and thereby in determining treatment options for cancer patients.
Seokwoo Lee, Myounghoon Cho, Wook Lee, Kyungsook Han
BIBM5
2022 DLoopCaller: A deep learning approach for predicting genome-wide chromatin loops by integrating accessible chromatin landscapes
abstract
In recent years, major advances have been made in various chromosome conformation capture technologies to further satisfy the needs of researchers for high-quality, high-resolution contact interactions. Discriminating the loops from genome-wide contact interactions is crucial for dissecting three-dimensional(3D) genome structure and function. Here, we present a deep learning method to predict genome-wide chromatin loops, called DLoopCaller, by combining accessible chromatin landscapes and raw Hi-C contact maps. Some available orthogonal data ChIA-PET/HiChIP and Capture Hi-C were used to generate positive samples with a wider contact matrix which provides the possibility to find more potential genome-wide chromatin loops. The experimental results demonstrate that DLoopCaller effectively improves the accuracy of predicting genome-wide chromatin loops compared to the state-of-the-art method Peakachu. Moreover, compared to two of most popular loop callers, such as HiCCUPS and Fit-Hi-C, DLoopCaller identifies some unique interactions. We conclude that a combination of chromatin landscapes on the one-dimensional genome contributes to understanding the 3D genome organization, and the identified chromatin loops reveal cell-type specificity and transcription factor motif co-enrichment across different cell lines and species.
Siguo Wang, Qinhu Zhang, Zhen-Hao Guo, Kyungsook Han, De-Shuang Huang
PLoS Comput. Biol.6
2022 Guest Editorial for Special Section on the 16th International Conference on Intelligent Computing (ICIC)
abstract
The eight papers in this special section were presented at the Sixteenth International Conference on Intelligent Computing (ICIC) that was held in Bari, Italy, on October 2-5, 2020. ICIC was formed to provide an annual forum dedicated to the emerging and challenging topics in artificial intelligence, machine learning, bioinformatics, and computational biology, etc. It aims to bring together researchers and practitioners from both academia and industry to share ideas, problems and solutions related to the multifaceted aspects of intelligent computing.
De-Shuang Huang, Kyungsook Han, Tatsuya Akutsu
IEEE ACM Trans. Comput. Biol. Bioinform.2
2022 A New Approach to Deriving Prognostic Gene Pairs From Cancer Patient-Specific Gene Correlation Networks
abstract
Many of the known prognostic gene signatures for cancer are individual genes or combination of genes, found by the analysis of microarray data. However, many of the known cancer signatures are less predictive than random gene expression signatures, and such random signatures are significantly associated with proliferation genes. With the availability of RNA-seq gene expression data for thousands of human cancer patients, we have analyzed RNA-seq and clinical data of cancer patients and constructed gene correlation networks specific to individual cancer patients. From the patient-specific gene correlation networks, we derived prognostic gene pairs for three types of cancer. In this paper, we propose a new method for inferring prognostic gene pairs from patient-specific gene correlation networks. The main difference of our method from previous ones includes (1) it is focused on finding prognostic gene pairs rather than prognostic genes, (2) it can identify prognostic gene pairs from RNA-seq data even when no significant prognostic genes exist, and (3) prognostic gene pairs can serve as robust prognostic biomarkers in the sense that most prognostic gene pairs show little association with proliferation genes, the major boosting factor of the predictive power of random gene signatures. Evaluation of our method with extensive data of three types of cancer (liver cancer, pancreatic cancer, and stomach cancer) showed that our approach is general and that gene pairs can serve as more reliable prognostic signatures for cancer than genes. Analysis of patient-specific gene networks suggests that prognosis of individual cancer patients is affected by the existence of prognostic gene pairs in the patient-specific network and by the size of the patient-specific network. Although preliminary, our approach will be useful for finding gene pairs to predict survival time of patients and to tailor treatments to individual characteristics. The program for dynamically constructing patient-specific gene networks and for finding prognostic gene pairs is available at http://bclab.inha.ac.kr/LPS.
Wook Lee, Kyungsook Han
IEEE ACM Trans. Comput. Biol. Bioinform.3
2022 A Deep Learning Model for RNA-Protein Binding Preference Prediction Based on Hierarchical LSTM and Attention Network
abstract
Attention mechanism has the ability to find important information in the sequence. The regions of the RNA sequence that can bind to proteins are more important than those that cannot bind to proteins. Neither conventional methods nor deep learning-based methods, they are not good at learning this information. In this study, LSTM is used to extract the correlation features between different sites in RNA sequence. We also use attention mechanism to evaluate the importance of different sites in RNA sequence. We get the optimal combination of k-mer length, k-mer stride window, k-mer sentence length, k-mer sentence stride window, and optimization function through hyper-parm experiments. The results show that the performance of our method is better than other methods. We tested the effects of changes in k-mer vector length on model performance. We show model performance changes under various k-mer related parameter settings. Furthermore, we investigate the effect of attention mechanism and RNA structure data on model performance.
Zhen Shen 0003, Qinhu Zhang, Kyungsook Han, De-Shuang Huang
IEEE ACM Trans. Comput. Biol. Bioinform.3
2021 A Method for Constructing an Integrative Network of Competing Endogenous RNAs
Seokwoo Lee, Wook Lee, Shulei Ren, Kyungsook Han
ICIC (3)4
2021 Predicting TF-DNA Binding Motifs from ChIP-seq Datasets Using the Bag-Based Classifier Combined With a Multi-Fold Learning Scheme
abstract
The rapid development of high-throughput sequencing technology provides unique opportunities for studying of transcription factor binding sites, but also brings new computational challenges. Recently, a series of discriminative motif discovery (DMD) methods have been proposed and offer promising solutions for addressing these challenges. However, because of the huge computation cost, most of them have to choose approximate schemes that either sacrifice the accuracy of motif representation or tune motif parameter indirectly. In this paper, we propose a bag-based classifier combined with a multi-fold learning scheme (BCMF) to discover motifs from ChIP-seq datasets. First, BCMF formulates input sequences as a labeled bag naturally. Then, a bag-based classifier, combining with a bag feature extracting strategy, is applied to construct the objective function, and a multi-fold learning scheme is used to solve it. Compared with the existing DMD tools, BCMF features three improvements: 1) Learning position weight matrix (PWM) directly in a continuous space; 2) Proposing to represent a positive bag with a feature fused by its k "most positive" patterns. 3) Applying a more advanced learning scheme. The experimental results on 134 ChIP-seq datasets show that BCMF substantially outperforms existing DMD methods (including DREME, HOMER, XXmotif, motifRG, EDCOD and our previous work).
Qinhu Zhang, Dailun Wang, Kyungsook Han, De-Shuang Huang
IEEE ACM Trans. Comput. Biol. Bioinform.3
2021 Multi-Scale Capsule Network for Predicting DNA-Protein Binding Sites
abstract
Discovering DNA-protein binding sites, also known as motif discovery, is the foundation for further analysis of transcription factors (TFs). Deep learning algorithms such as convolutional neural networks (CNN) have been introduced to motif discovery task and have achieved state-of-art performance. However, due to the limitations of CNN, motif discovery methods based on CNN do not take full advantage of large-scale sequencing data generated by high-throughput sequencing technology. Hence, in this paper we propose multi-scale capsule network architecture (MSC) integrating multi-scale CNN, a variant of CNN able to extract motif features of different lengths, and capsule network, a novel type of artificial neural network architecture aimed at improving CNN. The proposed method is tested on real ChIP-seq datasets and the experimental results show a considerable improvement compared with two well-tested deep learning-based sequence model, DeepBind and Deepsea.
Qinhu Zhang, Kyungsook Han, Asoke K. Nandi, De-Shuang Huang
IEEE ACM Trans. Comput. Biol. Bioinform.3
2020 Constructive Prediction of Potential RNA Aptamers for a Protein Target
abstract
Aptamers are short single-stranded nucleic acids that bind to target molecules with high affinity and selectivity. Aptamers are generally identified in vitro by performing SELEX (systematic evolution of ligands by exponential enrichment). Complementing the SELEX process, several computational methods have been proposed in the search for aptamers. However, many of these methods cannot be applied for finding new aptamers, either because they are classifiers for determining whether an RNA and protein interact with each other, or because they are limited to a specific target only. Hence, we developed a new random forest (RF) model for finding potential RNA aptamers for a protein target. From an extensive analysis of protein-RNA complexes including RNA aptamers-protein complexes, we identified key features of interacting RNA and protein molecules, and structural constraints on RNA aptamers. The potential RNA aptamers predicted by our method reveal similar secondary and protein-binding structures as the actual RNA aptamers. The RF model showed a reliable performance in both cross validations and independent testing. The key features of interacting RNA and protein molecules and the structural constraints identified in our study were effective in finding potential aptamers for a protein target. Although preliminary, our results are promising, and we believe this approach will be useful in reducing time and money spent on in vitro experiments by substantially limiting the size of the initial pool of nucleic acid sequences.
Wook Lee, Kyungsook Han
IEEE ACM Trans. Comput. Biol. Bioinform.2
2019 Integration of Multi-Omics Data for Gene Regulatory Network Inference and Application to Breast Cancer
abstract
Underlying a cancer phenotype is a specific gene regulatory network that represents the complex regulatory relationships between genes. It remains, however, a challenge to find cancer-related gene regulatory network because of insufficient sample sizes and complex regulatory mechanisms in which gene is influenced by not only other genes but also other biological factors. With the development of high-throughput technologies and the unprecedented wealth of multi-omics data it gives us a new opportunity to design machine learning method to investigate underlying gene regulatory network. In this paper, we propose an approach, which use Biweight Midcorrelation to measure the correlation between factors and make use of Nonconvex Penalty based sparse regression for Gene Regulatory Network inference (BMNPGRN). BMNCGRN incorporates multi-omics data (including DNA methylation and copy number variation) and their interactions in gene regulatory network model. The experimental results on synthetic datasets show that BMNPGRN outperforms popular and state-of-the-art methods (including DCGRN, ARACNE, and CLR) under false positive control. Furthermore, we applied BMNPGRN on breast cancer (BRCA) data from The Cancer Genome Atlas database and provided gene regulatory network.
Lin Yuan 0001, Lehang Guo, Chang-an Yuan 0001, Youhua Zhang, Kyungsook Han, Asoke K. Nandi, Barry Honig, De-Shuang Huang
IEEE ACM Trans. Comput. Biol. Bioinform.5
2018 Finding Protein-Binding Nucleic Acid Sequences Using a Long Short-Term Memory Neural Network
Jinho Im, Kyungsook Han
ICIC (2)3
2018 Finding Potential RNA Aptamers for a Protein Target Using Sequence and Structure Features
Wook Lee, Jisu Lee, Kyungsook Han
ICIC (1)3
2018 Constructing Gene Co-expression Networks for Prognosis of Lung Adenocarcinoma
Jinho Im, Kyungsook Han
ICIC (2)3
2018 Mutli-Features Prediction of Protein Translational Modification Sites
abstract
Post translational modification plays a significiant role in the biological processing. The potential post translational modification is composed of the center sites and the adjacent amino acid residues which are fundamental protein sequence residues. It can be helpful to perform their biological functions and contribute to understanding the molecular mechanisms that are the foundations of protein design and drug design. The existing algorithms of predicting modified sites often have some shortcomings, such as lower stability and accuracy. In this paper, a combination of physical, chemical, statistical, and biological properties of a protein have been ulitized as the features, and a novel framework is proposed to predict a protein's post translational modification sites. The multi-layer neural network and support vector machine are invoked to predict the potential modified sites with the selected features that include the compositions of amino acid residues, the E-H description of protein segments, and several properties from the AAIndex database. Being aware of the possible redundant information, the feature selection is proposed in the propocessing step in this research. The experimental results show that the proposed method has the ability to improve the accuracy in this classification issue.
Wenzheng Bao, Chang-an Yuan 0001, Youhua Zhang, Kyungsook Han, Asoke K. Nandi, Barry Honig, De-Shuang Huang
IEEE ACM Trans. Comput. Biol. Bioinform.4
2018 Sequence-Based Prediction of Putative Transcription Factor Binding Sites in DNA Sequences of Any Length
abstract
A transcription factor (TF) is a protein that regulates gene expression by binding to specific DNA sequences. Despite the recent advances in experimental techniques for identifying transcription factor binding sites (TFBS) in DNA sequences, a large number of TFBS are to be unveiled in many species. Several computational methods developed for predicting TFBS in DNA are tissue- or species-specific methods, so cannot be used without prior knowledge of tissue or species. Some computational methods are applicable to finding TFBS in short DNA sequences only. In this paper we propose a new learning method for predicting TFBS in DNA of any length using the composition, transition and distribution of nucleotides and amino acids in DNA and TF sequences. In independent testing of the method on datasets that were not used in training the method, its accuracy and MCC were as high as 81.84% and 0.634, respectively. The proposed method can be a useful aid for selecting potential TFBS in a large amount of DNA sequences before conducting biochemical experiments to empirically determine TFBS. The program and data sets are available at http://bclab.inha.ac.kr/TFbinding.
Wook Lee, Kyungsook Han
IEEE ACM Trans. Comput. Biol. Bioinform.3
2016 Prediction of Lysine Acetylation Sites Based on Neural Network
Wenzheng Bao, Zhichao Jiang, Kyungsook Han, De-Shuang Huang
ICIC (2)3
2016 Analysis of MicroRNA and Transcription Factor Regulation
Wei-Li Guo, Kyungsook Han, De-Shuang Huang
ICIC (1)2
2016 Prediction of Target Genes Based on Multiway Integration of High-Throughput Data
Wei-Li Guo, Kyungsook Han, De-Shuang Huang
ICIC (1)2
2016 Predicting Transcription Factor Binding Sites in DNA Sequences Without Prior Knowledge
Wook Lee, Daesik Choi, Chungkeun Lee, Hanju Chae, Kyungsook Han
ICIC (1)6
2016 Novel Algorithm for Multiple Quantitative Trait Loci Mapping by Using Bayesian Variable Selection Regression
Lin Yuan 0001, Kyungsook Han, De-Shuang Huang
ICIC (3)2
2016 GeneNetFinder2: Improved Inference of Dynamic Gene Regulatory Relations with Multiple Regulators
abstract
A gene involved in complex regulatory interactions may have multiple regulators since gene expression in such interactions is often controlled by more than one gene. Another thing that makes gene regulatory interactions complicated is that regulatory interactions are not static, but change over time during the cell cycle. Most research so far has focused on identifying gene regulatory relations between individual genes in a particular stage of the cell cycle. In this study we developed a method for identifying dynamic gene regulations of several types from the time-series gene expression data. The method can find gene regulations with multiple regulators that work in combination or individually as well as those with single regulators. The method has been implemented as the second version of GeneNetFinder (hereafter called GeneNetFinder2) and tested on several gene expression datasets. Experimental results with gene expression data revealed the existence of genes that are not regulated by individual genes but rather by a combination of several genes. Such gene regulatory relations cannot be found by conventional methods. Our method finds such regulatory relations as well as those with multiple, independent regulators or single regulators, and represents gene regulatory relations as a dynamic network in which different gene regulatory relations are shown in different stages of the cell cycle. GeneNetFinder2 is available at http://bclab.inha.ac.kr/GeneNetFinder and will be useful for modeling dynamic gene regulations with multiple regulators.
Kyungsook Han, Jeonghoon Lee 0001
IEEE ACM Trans. Comput. Biol. Bioinform.1
2015 SVM-Based Classification of Diffusion Tensor Imaging Data for Diagnosing Alzheimer's Disease and Mild Cognitive Impairment
Wook Lee, Kyungsook Han
ICIC (2)3
2014 Inference of Dynamic Gene Regulatory Relations with Multiple Regulators
Jeonghoon Lee 0001, Kyungsook Han
ICIC (3)3
2014 DBBP: database of binding pairs in protein-nucleic acid interactions
abstract
BACKGROUND: Interaction of proteins with other molecules plays an important role in many biological activities. As many structures of protein-DNA complexes and protein-RNA complexes have been determined in the past years, several databases have been constructed to provide structure data of the complexes. However, the information on the binding sites between proteins and nucleic acids is not readily available from the structure data since the data consists mostly of the three-dimensional coordinates of the atoms in the complexes. RESULTS: We analyzed the huge amount of structure data for the hydrogen bonding interactions between proteins and nucleic acids and developed a database called DBBP (DataBase of Binding Pairs in protein-nucleic acid interactions, http://bclab.inha.ac.kr/dbbp). DBBP contains 44,955 hydrogen bonds (H-bonds) of protein-DNA interactions and 77,947 H-bonds of protein-RNA interactions. CONCLUSIONS: Analysis of the huge amount of structure data of protein-nucleic acid complexes is labor-intensive, yet provides useful information for studying protein-nucleic acid interactions. DBBP provides the detailed information of hydrogen-bonding interactions between proteins and nucleic acids at various levels from the atomic level to the residue level. The binding information can be used as a valuable resource for developing a computational method aiming at predicting new binding sites in proteins or nucleic acids.
Hyungchan Kim, Kyungsook Han
BMC Bioinform.3
2013 Scoring Protein-Protein Interactions Using the Width of Gene Ontology Terms and the Information Content of Common Ancestors
Guangyu Cui, Kyungsook Han
ICIC (3)2
2013 Database of Protein-Nucleic Acid Binding Pairs at Atomic and Residue Levels
Hyungchan Kim, Kyungsook Han
ICIC (3)4
2013 KGVDB: a population-based genomic map of CNVs tagged by SNPs in Koreans
abstract
SUMMARY: Despite a growing interest in a correlation between copy number variations (CNVs) and flanking single nucleotide polymorphisms, few databases provide such information. In particular, most information on CNV available so far was obtained in Caucasian and Yoruba populations, and little is known about CNV in Asian populations. This article presents a database that provides CNV regions tagged by single nucleotide polymorphisms in about 4700 Koreans, which were detected under strict quality control, manually curated and experimentally validated. AVAILABILITY: KGVDB is freely available for non-commercial use at http://biomi.cdc.go.kr/KGVDB. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Sanghoon Moon, Kwang Su Jung, Young Jin Kim 0008, Mi Yeong Hwang, Kyungsook Han, Jong-Young Lee, Kiejung Park, Bong-Jo Kim
Bioinform.5
2012 Prediction of protein-protein interactions between viruses and human by an SVM model
abstract
BACKGROUND: Several computational methods have been developed to predict protein-protein interactions from amino acid sequences, but most of those methods are intended for the interactions within a species rather than for interactions across different species. Methods for predicting interactions between homogeneous proteins are not appropriate for finding those between heterogeneous proteins since they do not distinguish the interactions between proteins of the same species from those of different species. RESULTS: We developed a new method for representing a protein sequence of variable length in a frequency vector of fixed length, which encodes the relative frequency of three consecutive amino acids of a sequence. We built a support vector machine (SVM) model to predict human proteins that interact with virus proteins. In two types of viruses, human papillomaviruses (HPV) and hepatitis C virus (HCV), our SVM model achieved an average accuracy above 80%, which is higher than that of another SVM model with a different representation scheme. Using the SVM model and Gene Ontology (GO) annotations of proteins, we predicted new interactions between virus proteins and human proteins. CONCLUSIONS: Encoding the relative frequency of amino acid triplets of a protein sequence is a simple yet powerful representation method for predicting protein-protein interactions across different species. The representation method has several advantages: (1) it enables a prediction model to achieve a better performance than other representations, (2) it generates feature vectors of fixed length regardless of the sequence length, and (3) the same representation is applicable to different types of proteins.
Guangyu Cui, Kyungsook Han
BMC Bioinform.3
2012 Modeling the interactions of Alzheimer-related genes from the whole brain microarray data and diffusion tensor images of human brain
abstract
BACKGROUND: In recent years the genome-wide microarray-based gene expression profiles and diffusion tensor images (DTI) in human brain have been made available with accompanying anatomic and histology data. The challenge is to integrate various types of data to investigate the interactions of genes that are associated with specific neurological disorder. RESULTS: In this study, we analyzed the whole brain microarray data and the physical connectivity of the hippocampus with other brain regions to identify the genes related to Alzheimer's disease and their interactions with proteins. We generated a physical connectivity map of the left and right hippocampuses with 12 other brain regions and identified 33 Alzheimer-related genes that interact with many proteins. These genes are highly linked to the development of Alzheimer's disease. CONCLUSIONS: In Alzheimer's brain both brain regions and inter-regional communications through the white matter are often hampered. So far the connectivity of regions in Alzheimer's brain has been studied mostly at the functional level using functional MRI (fMRI). Analyzing the inter-regional fiber connectivity without tracking crossing-fiber regions often provides coarse and inaccurate results. A few deep brain fibers were analyzed but the inter-regional fiber connectivity was not analyzed in their studies. The inter-regional fiber connectivity analysis can provide comprehensive and measurable degradation of fiber tracts in AD patients' brains, but is not easy to perform. We tracked crossing-fiber regions and identified genes with high expression levels in the fiber pathways of the hippocampus. The interactions of the genes with other proteins can provide comprehensive and measurable degradation of fiber tracts in Alzheimer brains. To the best of our knowledge, this is the first attempt to integrate the whole brain microarray data with DTI data to identify specific genes and their interactions.
Wook Lee, Kyungsook Han
BMC Bioinform.3
2011 Prediction of Human Proteins Interacting with Human Papillomavirus Proteins
Guangyu Cui, Kyungsook Han
ICIC (3)3
2011 Connectivity Analysis of Hippocampus in Alzheimer's Brain Using Probabilistic Tractography
Wook Lee, Kyungsook Han
ICIC (3)4
2011 Prediction of RNA-binding amino acids from protein and RNA sequences
abstract
BACKGROUND: Many learning approaches to predicting RNA-binding residues in a protein sequence construct a non-redundant training dataset based on the sequence similarity. The sequence similarity-based method either takes a whole sequence or discards it for a training dataset. However, similar sequences or even identical sequences can have different interaction sites depending on their interaction partners, and this information is lost when the sequences are removed. Furthermore, a training dataset constructed by the sequence similarity-based method may contain redundant data when the remaining sequence contains similar subsequences within the sequence. In addition to the problem with the training dataset, most approaches do not consider the interacting partner (i.e., RNA) of a protein when they predict RNA-binding amino acids. Thus, they always predict the same RNA-binding sites for a given protein sequence even if the protein binds to different RNA molecules. RESULTS: We developed a feature vector-based method that removes data redundancy for a non-redundant training dataset. The feature vector-based method constructed a larger training dataset than the standard sequence similarity-based method, yet the dataset contained no redundant data. We identified effective features of protein and RNA (the interaction propensity of amino acid triplets, global features of the protein sequence, and RNA feature) for predicting RNA-binding residues. Using the method and features, we built a support vector machine (SVM) model that predicted RNA-binding residues in a protein sequence. Our SVM model showed an accuracy of 84.2%, an F-measure of 76.1%, and a correlation coefficient of 0.41 with 5-fold cross validation on a non-redundant dataset from 3,149 protein-RNA interacting pairs. In an independent test dataset that does not include the 3,149 pairs and were not used in training the SVM model, it achieved an accuracy of 90.3%, an F-measure of 72.8%, and a correlation coefficient of 0.24. Comparison with other methods on the same datasets demonstrated that our model was better than the others. CONCLUSIONS: The feature vector-based redundancy reduction method is powerful for constructing a non-redundant training dataset for a learning model since it generates a larger dataset with non-redundant data than the standard sequence similarity-based method. Including the features of both RNA and protein sequences in a feature vector results in better performance than using the protein features only when predicting the RNA-binding residues in a protein sequence.
Sungwook Choi, Kyungsook Han
BMC Bioinform.2
2010 An ontology-based search engine for protein-protein interactions
abstract
BACKGROUND: Keyword matching or ID matching is the most common searching method in a large database of protein-protein interactions. They are purely syntactic methods, and retrieve the records in the database that contain a keyword or ID specified in a query. Such syntactic search methods often retrieve too few search results or no results despite many potential matches present in the database. RESULTS: We have developed a new method for representing protein-protein interactions and the Gene Ontology (GO) using modified Gödel numbers. This representation is hidden from users but enables a search engine using the representation to efficiently search protein-protein interactions in a biologically meaningful way. Given a query protein with optional search conditions expressed in one or more GO terms, the search engine finds all the interaction partners of the query protein by unique prime factorization of the modified Gödel numbers representing the query protein and the search conditions. CONCLUSION: Representing the biological relations of proteins and their GO annotations by modified Gödel numbers makes a search engine efficiently find all protein-protein interactions by prime factorization of the numbers. Keyword matching or ID matching search methods often miss the interactions involving a protein that has no explicit annotations matching the search condition, but our search engine retrieves such interactions as well if they satisfy the search condition with a more specific term in the ontology.
Kyungsook Han
BMC Bioinform.2
2010 A semi-supervised learning approach to predict synthetic genetic interactions by combining functional and topological properties of functional gene network
abstract
BACKGROUND: Genetic interaction profiles are highly informative and helpful for understanding the functional linkages between genes, and therefore have been extensively exploited for annotating gene functions and dissecting specific pathway structures. However, our understanding is rather limited to the relationship between double concurrent perturbation and various higher level phenotypic changes, e.g. those in cells, tissues or organs. Modifier screens, such as synthetic genetic arrays (SGA) can help us to understand the phenotype caused by combined gene mutations. Unfortunately, exhaustive tests on all possible combined mutations in any genome are vulnerable to combinatorial explosion and are infeasible either technically or financially. Therefore, an accurate computational approach to predict genetic interaction is highly desirable, and such methods have the potential of alleviating the bottleneck on experiment design. RESULTS: In this work, we introduce a computational systems biology approach for the accurate prediction of pairwise synthetic genetic interactions (SGI). First, a high-coverage and high-precision functional gene network (FGN) is constructed by integrating protein-protein interaction (PPI), protein complex and gene expression data; then, a graph-based semi-supervised learning (SSL) classifier is utilized to identify SGI, where the topological properties of protein pairs in weighted FGN is used as input features of the classifier. We compare the proposed SSL method with the state-of-the-art supervised classifier, the support vector machines (SVM), on a benchmark dataset in S. cerevisiae to validate our method's ability to distinguish synthetic genetic interactions from non-interaction gene pairs. Experimental results show that the proposed method can accurately predict genetic interactions in S. cerevisiae (with a sensitivity of 92% and specificity of 91%). Noticeably, the SSL method is more efficient than SVM, especially for very small training sets and large test sets. CONCLUSIONS: We developed a graph-based SSL classifier for predicting the SGI. The classifier employs topological properties of weighted FGN as input features and simultaneously employs information induced from labelled and unlabelled data. Our analysis indicates that the topological properties of weighted FGN can be employed to accurately predict SGI. Also, the graph-based SSL method outperforms the traditional standard supervised approach, especially when used with small training sets. The proposed method can alleviate experimental burden of exhaustive test and provide a useful guide for the biologist in narrowing down the candidate gene pairs with SGI. The data and source code implementing the method are available from the website: http://home.ustc.edu.cn/~yzh33108/GeneticInterPred.htm.
Zhu-Hong You, Zheng Yin, Kyungsook Han, De-Shuang Huang, Xiaobo Zhou 0001
BMC Bioinform.3
2009 Dynamic Identification and Visualization of Gene Regulatory Networks from Time-Series Gene Expression Profiles
Kyungsook Han
ICIC (1)2
2009 PseudoViewer3: generating planar drawings of large-scale RNA structures with pseudoknots
abstract
MOTIVATION: Pseudoknots in RNA structures make visualization of RNA structures difficult. Even if a pseudoknot itself is represented without a crossing, visualization of the entire RNA structure with a pseudoknot often results in a drawing with crossings between the pseudoknot and other structural elements, and requires additional intervention by the user to ensure that the structure graph is overlap-free. Many programs such as web services prefer to obtain an overlap-free graph in one-shot rather than get a graph with overlaps to be edited. There are few programs for visualizing RNA pseudoknots, and PseudoViewer has been the almost only program that automatically draws RNA secondary structures with pseudoknots. The previous version of PseudoViewer visualizes all the known types of RNA pseudoknots as planar drawings, but visualizes some hypothetical pseudoknots as non-planar drawings. RESULTS: We developed a new version of PseudoViewer for efficiently visualizing large RNA structures with any types of pseudoknots, both known and hypothetical, as planar drawings in one-shot. It is about 10 times faster than the previous algorithm, and produces a more compact and aesthetic structure drawing. PseudoViewer3 supports both web services and web applications. AVAILABILITY: The new version of PseudoViewer, PseudoViewer3, is available at (http://pseudoviewer.inha.ac.kr).
Yanga Byun, Kyungsook Han
Bioinform.2
2009 Finding motif pairs in the interactions between heterogeneous proteins via bootstrapping and boosting
abstract
BACKGROUND: Supervised learning and many stochastic methods for predicting protein-protein interactions require both negative and positive interactions in the training data set. Unlike positive interactions, negative interactions cannot be readily obtained from interaction data, so these must be generated. In protein-protein interactions and other molecular interactions as well, taking all non-positive interactions as negative interactions produces too many negative interactions for the positive interactions. Random selection from non-positive interactions is unsuitable, since the selected data may not reflect the original distribution of data. RESULTS: We developed a bootstrapping algorithm for generating a negative data set of arbitrary size from protein-protein interaction data. We also developed an efficient boosting algorithm for finding interacting motif pairs in human and virus proteins. The boosting algorithm showed the best performance (84.4% sensitivity and 75.9% specificity) with balanced positive and negative data sets. The boosting algorithm was also used to find potential motif pairs in complexes of human and virus proteins, for which structural data was not used to train the algorithm. Interacting motif pairs common to multiple folds of structural data for the complexes were proven to be statistically significant. The data set for interactions between human and virus proteins was extracted from BOND and is available at http://virus.hpid.org/interactions.aspx. The complexes of human and virus proteins were extracted from PDB and their identifiers are available at http://virus.hpid.org/PDB_IDs.html. CONCLUSION: When the positive and negative training data sets are unbalanced, the result via the prediction model tends to be biased. Bootstrapping is effective for generating a negative data set, for which the size and distribution are easily controlled. Our boosting algorithm could efficiently predict interacting motif pairs from protein interaction and sequence data, which was trained with the balanced data sets generated via the bootstrapping method.
De-Shuang Huang, Kyungsook Han
BMC Bioinform.3
2008 Prediction of RNA-Binding Residues in Proteins Using the Interaction Propensities of Amino Acids and Nucleotides
Rojan Shrestha, Kyungsook Han
ICIC (1)3
2008 Prediction of Binding Sites in HCV Protein Complexes Using a Support Vector Machine
Taihui Yoo, Jaetak Lee, Kyungsook Han
ICIC (1)3
2006 Prediction of Ribosomal -1 Frameshifts in the Escherichia coli K12 Genome
Sanghoon Moon, Yanga Byun, Kyungsook Han
ICIC (3)3
2006 Web Service for Predicting Interacting Proteins and Application to Human and HIV-1 Proteins
Kyungsook Han
ICIC (3)2
2005 PSIbase: a database of Protein Structural Interactome map (PSIMAP)
abstract
UNLABELLED: Protein Structural Interactome map (PSIMAP) is a global interaction map that describes domain-domain and protein-protein interaction information for known Protein Data Bank structures. It calculates the Euclidean distance to determine interactions between possible pairs of structural domains in proteins. PSIbase is a database and file server for protein structural interaction information calculated by the PSIMAP algorithm. PSIbase also provides an easy-to-use protein domain assignment module, interaction navigation and visual tools. Users can retrieve possible interaction partners of their proteins of interests if a significant homology assignment is made with their query sequences. AVAILABILITY: http://psimap.org and http://psibase.kaist.ac.kr/
Sungsam Gong, Giseok Yoon, Insoo Jang, Dan M. Bolser, Panos Dafas, Michael Schroeder 0001, Hansol Choi, Yoobok Cho, Kyungsook Han, Sunghoon Lee, Hwanho Choi, Michael Lappe, Liisa Holm, Sangsoo Kim, Donghoon Oh, Jonghwa Bhak
Bioinform.9
2004 DNA Secondary Structures for Probe Design
Yanga Byun, Kyungsook Han
GD2
2004 HPID: The Human Protein Interaction Database
abstract
UNLABELLED: The Human Protein Interaction Database (http://www.hpid.org) was designed (1) to provide human protein interaction information pre-computed from existing structural and experimental data, (2) to predict potential interactions between proteins submitted by users and (3) to provide a depository for new human protein interaction data from users. Two types of interaction are available from the pre-computed data: (1) interactions at the protein superfamily level and (2) those transferred from the interactions of yeast proteins. Interactions at the superfamily level were obtained by locating known structural interactions of the PDB in the SCOP domains and identifying homologs of the domains in the human proteins. Interactions transferred from yeast proteins were obtained by identifying homologs of the yeast proteins in the human proteins. For each human protein in the database and each query submitted by users, the protein superfamilies and yeast proteins assigned to the protein are shown, along with their interacting partners. We have also developed a set of web-based programs so that users can visualize and analyze protein interaction networks in order to explore the networks further. AVAILABILITY: http://www.hpid.org.
Kyungsook Han, Hyongguen Kim, Jinsun Hong, Jong-Chan Park
Bioinform.1
2003 A Genetic Algorithm for Inferring Pseudoknotted RNA Structures from Sequence Data
Kyungsook Han
Discovery Science2
2003 Mining RNA Structure Elements from the Structure Data of Protein-RNA Complexes
Daeho Lim, Kyungsook Han
Discovery Science2
2003 Predicting Protein Interactions in Human by Homologous Interactions in Yeast
Hyongguen Kim, Jong-Chan Park, Kyungsook Han
PAKDD3
2003 A fast layout algorithm for protein interaction networks
abstract
MOTIVATION: Graph drawing algorithms are often used for visualizing relational information, but a naive implementation of a graph drawing algorithm encounters real difficulties when drawing large-scale graphs such as protein interaction networks. RESULTS: We have developed a new, extremely fast layout algorithm for visualizing large-scale protein interaction networks in the three-dimensional space. The algorithm (1) first finds a layout of connected components of an entire network, (2) finds a global layout of nodes with respect to pivot nodes within a connected component and (3) refines the local layout of each connected component by first relocating midnodes with respect to their cutvertices and direct neighbors of the cutvertices and then by relocating all nodes with respect to their neighbors within distance 2. Advantages of this algorithm over classical graph drawing methods include: (1) it is an order of magnitude faster, (2) it can directly visualize data from protein interaction databases and (3) it provides several abstraction and comparison operations for effectively analyzing large-scale protein interaction networks. AVAILABILITY: http://wilab.inha.ac.kr/interviewer/
Kyungsook Han, Byong-Hyon Ju
Bioinform.1
2003 Visualization and analysis of protein interactions
abstract
SUMMARY: We have developed a new program called InterViewer for drawing large-scale protein interaction networks in three-dimensional space. Unique features of InterViewer include (1) it is much faster than other recent implementations of drawing algorithms; (2) it can be used not only for visualizing protein interactions but also for analyzing them interactively; and (3) it provides an integrated framework for querying protein interaction databases and directly visualizes the query results. AVAILABILITY: http://wilab.inha.ac.kr/protein/
Byong-Hyon Ju, Jong H. Park, Kyungsook Han
Bioinform.4
2002 New Representation and Algorithm for Drawing RNA Structure with Pseudoknots
Wootaek Kim, Kyungsook Han
DaWaK3
2002 A Partitioned Approach to Protein Interaction Mapping
Yanga Byun, Euna Jeong, Kyungsook Han
GD3
2002 InterViewer: Dynamic Visualization of Protein-Protein Interactions
Kyungsook Han, Byong-Hyon Ju, Jong H. Park
GD1
2002 PseudoViewer: automatic visualization of RNA pseudoknots
abstract
MOTIVATION: Several algorithms have been developed for drawing RNA secondary structures, however none of these can be used to draw RNA pseudoknot structures. In the sense of graph theory, a drawing of RNA secondary structures is a tree, whereas a drawing of RNA pseudoknots is a graph with inner cycles within a pseudoknot as well as possible outer cycles formed between a pseudoknot and other structural elements. Thus, RNA pseudoknots are more difficult to visualize than RNA secondary structures. Since no automatic method for drawing RNA pseudoknots exists, visualizing RNA pseudoknots relies on significant amount of manual work and does not yield satisfactory results. The task of visualizing RNA pseudoknots by hand becomes more challenging as the size and complexity of the RNA pseudoknots increase. RESULTS: We have developed a new representation and an algorithm for drawing H-type pseudoknots with RNA secondary structures. Compared to existing representations of H-type pseudoknots, the new representation ensures uniform and clear drawings with no edge crossing for any H-type pseudoknots. To the best of our knowledge, this is the first algorithm for automatically drawing RNA pseudoknots with RNA secondary structures. The algorithm has been implemented in a Java program, which can be executed on any computing system. Experimental results demonstrate that the algorithm generates an aesthetically pleasing drawing of all H-type pseudoknots. The results have also shown that the drawing has high readability, enabling the user to quickly and easily recognize the whole RNA structure as well as the pseudoknots themselves.
Kyungsook Han, Wootaek Kim
ISMB1
2001 Web-Based Intelligent Call Center for an Intensive Care Unit
Kyungsook Han
Web Intelligence1
1999 A vector-based method for drawing RNA secondary structure
abstract
MOTIVATION: To produce a polygonal display of RNA secondary structure with minimal overlap and distortion of structural elements, with minimal search for positioning them, and with minimal user intervention. RESULTS: A new algorithm for automatically drawing RNA secondary structure has been developed. The algorithm represents the direction and space for a structural element using vector and vector space. Two heuristics are used. The first heuristic is concerned with ordering structural elements to be positioned and the second with positioning them in space. The algorithm and a graphical user interface have been implemented in a working program called VizQFolder on IBM PC compatibles. Experimental results demonstrate that VizQFolder is capable of automatically generating nearly overlap-free polygonal displays for long RNA molecules. The only distortion performed to avoid overlap is the rotation of helices, leading to efficient generation of a polygonal display without sacrificing its readability. VizQFolder is not coupled to a specific prediction program of RNA secondary structure, and thus can be used for visualizing secondary structure models obtained by any means. AVAILABILITY: The executable code of VizQFolder is available at http://automation.inha.ac.kr/khan. It can also be obtained from the authors upon request.
Kyungsook Han, Hong-Jin Kim
Bioinform.1
1994 The Epistemology of Physical System Modeling
Kyungsook Han, Andrew Gelsey
AAAI1
1993 Qualitative Modeling of RNA Structure
Kyungsook Han, Andrew Gelsey
IJCAI1