EDBT 2026 Demo / reviewers in the wild / expert
Doheon Lee
dblp:43/1124
· DBLP profile ↗
88ranked-venue papers
10as first author
7since 2021 · last 2026
0000-0001-9070-4316ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 50 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 26 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 18 · 4 first-authorHuman-computer interaction and ubiquitous computing · 4 · 2 first-authorSystems, architecture and hardware · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SOCAR: Network-Based Computational Framework to Overcome Acquired Tamoxifen Resistance of MCF7 CellsabstractResistance to anti-cancer drugs remains a major challenge in chemotherapy. Combination therapy using sensitizers has emerged as a promising strategy to restore drug sensitivity in resistant cancer cells. However, experimental screening of sensitizers is laborious and costly, highlighting the need for computational methods that enable systematic and efficient prediction. We developed SOCAR, a network-based computational framework that predicts sensitizer drugs by integrating transcriptome profiles with molecular interaction networks. SOCAR identifies resistance-associated genes and network modules, and quantifies each drug's potential to reverse these resistance mechanisms. Applied to 4,009 drugs, SOCAR accurately predicted candidate sensitizers for tamoxifen-resistant breast cancer (AUROC = 0.90). In vitro assays validated that all twelve top-ranked candidates significantly reduced cell viability (p < 0.005) when co-administered with tamoxifen. Furthermore, protein activity analyses showed that resistance-module proteins were markedly altered after acquiring resistance but were restored to normal levels following combined treatment (p < 0.05). Collectively, SOCAR provides a systems-level framework for discovering novel sensitizers and elucidating mechanisms of resistance reversal. Mijin Kwon, Young Min Woo, Woochang Hwang, Jong-Won Kim 0001, Doheon Lee |
IEEE Trans. Comput. Biol. Bioinform. | 5 |
| 2025 | Autocompose: Automatic Generation of Pose Transition Descriptions for Composed Pose Retrieval Using Multimodal LLMsabstractComposed pose retrieval (CPR) enables users to search for human poses by specifying a reference pose and a transition description, but progress in this field is hindered by the scarcity and inconsistency of annotated pose transitions. Existing CPR datasets rely on costly human annotations or heuristic-based rule generation, both of which limit scalability and diversity. In this work, we introduce AutoComPose, the first framework that leverages multimodal large language models (MLLMs) to automatically generate rich and structured pose transition descriptions. Our method enhances annotation quality by structuring transitions into fine-grained body part movements and introducing mirrored/swapped variations, while a cyclic consistency constraint ensures logical coherence between forward and reverse transitions. To advance CPR research, we construct and release two dedicated benchmarks, AIST-CPR and PoseFixCPR, supplementing prior datasets with enhanced attributes. Extensive experiments demonstrate that training retrieval models with AutoComPose yields superior performance over human-annotated and heuristic-based methods, significantly reducing annotation costs while improving retrieval quality. Our work pioneers the automatic annotation of pose transitions, establishing a scalable foundation for future CPR research. Yi-Ting Shen, Sungmin Eum, Doheon Lee, Rohit Shete, Chiao-Yi Wang, Heesung Kwon, Shuvra S. Bhattacharyya |
ICCV | 3 |
| 2024 | Community cohesion looseness in gene networks reveals individualized drug targets and resistanceabstractCommunity cohesion plays a critical role in the determination of an individual's health in social science. Intriguingly, a community structure of gene networks indicates that the concept of community cohesion could be applied between the genes as well to overcome the limitations of single gene-based biomarkers for precision oncology. Here, we develop community cohesion scores which precisely quantify the community ability to retain the interactions between the genes and their cellular functions in each individualized gene network. Using breast cancer as a proof-of-concept study, we measure the community cohesion score profiles of 950 case samples and predict the individualized therapeutic targets in 2-fold. First, we prioritize them by finding druggable genes present in the community with the most and relatively decreased scores in each individual. Then, we pinpoint more individualized therapeutic targets by discovering the genes which greatly contribute to the community cohesion looseness in each individualized gene network. Compared with the previous approaches, the community cohesion scores show at least four times higher performance in predicting effective individualized chemotherapy targets based on drug sensitivity data. Furthermore, the community cohesion scores successfully discover the known breast cancer subtypes and we suggest new targeted therapy targets for triple negative breast cancer (e.g. KIT and GABRP). Lastly, we demonstrate that the community cohesion scores can predict tamoxifen responses in ER+ breast cancer and suggest potential combination therapies (e.g. NAMPT and RXRA inhibitors) to reduce endocrine therapy resistance based on individualized characteristics. Our method opens new perspectives for the biomarker development in precision oncology. Seunghyun Wang, Doheon Lee |
Briefings Bioinform. | 2 |
| 2023 | Large-scale prediction of adverse drug reactions-related proteins with network embeddingabstractMOTIVATION: Adverse drug reactions (ADRs) are a major issue in drug development and clinical pharmacology. As most ADRs are caused by unintended activity at off-targets of drugs, the identification of drug targets responsible for ADRs becomes a key process for resolving ADRs. Recently, with the increase in the number of ADR-related data sources, several computational methodologies have been proposed to analyze ADR-protein relations. However, the identification of ADR-related proteins on a large scale with high reliability remains an important challenge. RESULTS: In this article, we suggest a computational approach, Large-scale ADR-related Proteins Identification with Network Embedding (LAPINE). LAPINE combines a novel concept called single-target compound with a network embedding technique to enable large-scale prediction of ADR-related proteins for any proteins in the protein-protein interaction network. Analysis of benchmark datasets confirms the need to expand the scope of potential ADR-related proteins to be analyzed, as well as LAPINE's capability for high recovery of known ADR-related proteins. Moreover, LAPINE provides more reliable predictions for ADR-related proteins (Value-added positive predictive value = 0.12), compared to a previously proposed method (P < 0.001). Furthermore, two case studies show that most predictive proteins related to ADRs in LAPINE are supported by literature evidence. Overall, LAPINE can provide reliable insights into the relationship between ADRs and proteomes to understand the mechanism of ADRs leading to their prevention. AVAILABILITY AND IMPLEMENTATION: The source code is available at GitHub (https://github.com/rupinas/LAPINE) and Figshare (https://figshare.com/articles/software/LAPINE/21750245) to facilitate its use. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jaesub Park, Sangyeon Lee, Kwansoo Kim, Jaegyun Jung, Doheon Lee |
Bioinform. | 5 |
| 2023 | Identifying prognostic subgroups of luminal-A breast cancer using deep autoencoders and gene expressionsabstractLuminal-A breast cancer is the most frequently occurring subtype which is characterized by high expression levels of hormone receptors. However, some luminal-A breast cancer patients suffer from intrinsic and/or acquired resistance to endocrine therapies which are considered as first-line treatments for luminal-A breast cancer. This heterogeneity within luminal-A breast cancer has required a more precise stratification method. Hence, our study aims to identify prognostic subgroups of luminal-A breast cancer. In this study, we discovered two prognostic subgroups of luminal-A breast cancer (BPS-LumA and WPS-LumA) using deep autoencoders and gene expressions. The deep autoencoders were trained using gene expression profiles of 679 luminal-A breast cancer samples in the METABRIC dataset. Then, latent features of each samples generated from the deep autoencoders were used for K-Means clustering to divide the samples into two subgroups, and Kaplan-Meier survival analysis was performed to compare prognosis (recurrence-free survival) between them. As a result, the prognosis between the two subgroups were significantly different (p-value = 5.82E-05; log-rank test). This prognostic difference between two subgroups was validated using gene expression profiles of 415 luminal-A breast cancer samples in the TCGA BRCA dataset (p-value = 0.004; log-rank test). Notably, the latent features were superior to the gene expression profiles and traditional dimensionality reduction method in terms of discovering the prognostic subgroups. Lastly, we discovered that ribosome-related biological functions could be potentially associated with the prognostic difference between them using differentially expressed genes and co-expression network analysis. Our stratification method can be contributed to understanding a complexity of luminal-A breast cancer and providing a personalized medicine. Seunghyun Wang, Doheon Lee |
PLoS Comput. Biol. | 2 |
| 2022 | SNEP-DB: An integrated database to associate genomic and pathological aspects of psychiatric disordersabstractPsychiatric disorders are widespread throughout the world. Large-scale genome-wide association studies (GWAS) have revealed numerous SNPs associated with a risk for psychiatric disorders. However, the biological mechanisms by which the SNPs contribute to the development of the disease are unknown. Systematical evaluation of the association between the SNPs and various quantitative phenotypes such as levels of gene expression and pathological markers can be a useful approach to identifying the underlying mechanisms. To facilitate this approach, we developed a database named Stanley Integrated Neurogenomic Pathology Database, SNEP-DB, that includes 4,070 neuropathology marker data, gene expression data, and SNPs from samples of the Stanley Medical Research Institute (SMRI) brain collections. Moreover, the database allows users to explore association analysis results from the SMRI brain data integrated with the GWAS results from recent large-scale studies. By exploiting the database, we have identified 43,417 associations between SNPs and levels of neuropathology markers, so-called markerQTLs. We have also identified 77,410 gene-level eQTLs and 129,980 transcript-level eQTLs. We expect SNEP-DB will enable users to explore the mechanisms underlying major psychiatric disorders by identifying significant associations between disease-associated SNPs and various quantitative phenotypes. (Database URL: http://bidas.kaist.ac.kr/snep/) Ohhyeon Kwon, Maree J. Webster, Doheon Lee |
BIBM | 4 |
| 2021 | Context-aware multi-token concept recognition of biological entitiesabstractBACKGROUND: Concept recognition is a term that corresponds to the two sequential steps of named entity recognition and named entity normalization, and plays an essential role in the field of bioinformatics. However, the conventional dictionary-based methods did not sufficiently addressed the variation of the concepts in actual use in literature, resulting in the particularly degraded performances in recognition of multi-token concepts. RESULTS: In this paper, we propose a concept recognition method of multi-token biological entities using neural models combined with literature contexts. The key aspect of our method is utilizing the contextual information from the biological knowledge-bases for concept normalization, which is followed by named entity recognition procedure. The model showed improved performances over conventional methods, particularly for multi-token concepts with higher variations. CONCLUSIONS: We expect that our model can be utilized for effective concept recognition and variety of natural language processing tasks on bioinformatics. Doheon Lee |
BMC Bioinform. | 2 |
| 2020 | Accelerating Drug Discovery with an AI-Based Virtual Human System CODAabstractFormidable complexity of systemic human physiology often give rises of unintended effects of therapeutic compounds during the drug development processes or even after the drug approvals. Though the beneficial unintended effects could lead opportunities of repositioning drugs, the harmful effects might put critical hurdles against successful drug development. We have been developing a virtual human system, CODA, which can explore functional effects of therapeutic compounds in the systemic level. CODA integrates three types of physiological knowledge from public structured databases, literature, and in-house experiments into a unified format of physiological interactions. More than ten public databases including KEGG, GO, and CTD have been transformed; around 25 million PUBMED abstracts have been text-mined; and more than 5,000 in-house novel findings have been incorporated. We have also developed two types of analysis on the CODA knowledge repository. Given therapeutic compounds of interest, CODA can identify possible phenotypic effects in the systemic level. When therapeutic compounds and their observed functional effects are given, CODA can enumerate possible effect paths encompassing molecular, functional, and disease level interactions. We have been testing CODA by applying it to various tasks including drug repositioning, drug-drug interactions, and side effect prediction with known benchmark datasets. Though we are enriching CODA with more knowledge sources and more sophisticated analysis techniques, the current version is already providing unique analysis capabilities and one of the most comprehensive information for drug discovery. Doheon Lee |
BIBM | 1 |
| 2020 | Identifying prognostic subgroups of luminal-A breast cancer using a deep autoencoderabstractLuminal-A breast cancer is the most frequently occurring breast cancer subtype. However, it shows high variability in prognosis, and more precise stratification is required for personalized medicine. In this paper, we identify two prognostic subgroups of luminal-A breast cancer. We train a deep autoencoder with gene expression profiles of luminal-A breast cancer, and it automatically generates informative latent features that represent essential properties of gene expressions. We find that two subgroups (BPS-LumA and WPS-LumA) clustered using the latent features are significantly different in prognosis (p-value =1.23e-6;log-rank test). This prognostic difference is validated with other luminal-A breast cancer cohort. The results in our method suggest that the deep autoencoder is able to extract and compress complex properties of gene expressions patterns, and that it is usefully applicable to patient stratification for precision medicine of luminal-A breast cancer. Seunghyun Wang, Doheon Lee |
BIBM | 2 |
| 2020 | Literature mining for context-specific molecular relations using multimodal representations (COMMODAR)abstractBiological contextual information helps understand various phenomena occurring in the biological systems consisting of complex molecular relations. The construction of context-specific relational resources vastly relies on laborious manual extraction from unstructured literature. In this paper, we propose COMMODAR, a machine learning-based literature mining framework for context-specific molecular relations using multimodal representations. The main idea of COMMODAR is the feature augmentation by the cooperation of multimodal representations for relation extraction. We leveraged biomedical domain knowledge as well as canonical linguistic information for more comprehensive representations of textual sources. The models based on multiple modalities outperformed those solely based on the linguistic modality. We applied COMMODAR to the 14 million PubMed abstracts and extracted 9214 context-specific molecular relations. All corpora, extracted data, evaluation results, and the implementation code are downloadable at https://github.com/jae-hyun-lee/commodar . CCS CONCEPTS: • Computing methodologies~Information extraction • Computing methodologies~Neural networks • Applied computing~Biological networks. Doheon Lee, Kwang Hyung Lee |
BMC Bioinform. | 2 |
| 2019 | Visualizing multifunctional PPI network with Gene Ontology annotationabstractProtein-protein interaction (PPI) network data contain interactions between two proteins in organisms. Some proteins are involved in more than one biological process. Those multifunctional proteins are important in the biological system because they connect various biological pathways. However, the multifunctionality of proteins increases the complexity of the PPI network and makes visualizing a challenging problem. In this research, we suggest a new method to visualize the multifunctional PPI network by integrating interaction data and Gene Ontology (GO) terms of each protein. By matching GO annotations of two proteins, we annotate edges between those two proteins with common GO terms. Also, we suggest a novel algorithm for color mapping specialized to express hierarchical relationships among GO terms based on the multidimensional scaling algorithm. Annotated edges are divided corresponding to their GO terms, and visualized by selected color, width, and opacity to represent semantic relationships among edges clearly. Those visual features are assigned to represent semantic distance among GO terms and their concreteness. This visualization method is available through the web. Sangyeon Lee, Doheon Lee |
BIBM | 2 |
| 2019 | DTMBIO 2019: The Thirteenth International Workshop on Data and Text Mining in Biomedical InformaticsabstractStarted in 2006 as a specialized workshop in the field of text mining applied to biomedical informatics, DTMBIO (ACM international workshop on Data and Text Mining in Biomedical Informatics) has been held annually in conjunction with one of the largest data management conferences, CIKM, bringing together researchers working on computer science and bioinformatics area including text mining and genomic data analysis. The purpose of DTMBIO is to foster discussions regarding the state-of-the-art applications of data and text mining on biomedical research problems. DTMBIO 2019 will help scientists navigate emerging trends and opportunities in the evolving area of informatics related techniques and problems in the context of biomedical research. Hyojung Paik, Ruibin Xi, Doheon Lee |
CIKM | 3 |
| 2019 | Diffraction-Aware Sound Localization for a Non-Line-of-Sight SourceabstractWe present a novel sound localization algorithm for a non-line-of-sight (NLOS) sound source in indoor environments. Our approach exploits the diffraction properties of sound waves as they bend around a barrier or an obstacle in the scene. We combine a ray tracing-based sound propagation algorithm with a Uniform Theory of Diffraction (UTD) model, which simulate bending effects by placing a virtual sound source on a wedge in the environment. We precompute the wedges of a reconstructed mesh of an indoor scene and use them to generate diffraction acoustic rays to localize the 3D position of the source. Our method identifies the convergence region of those generated acoustic rays as the estimated source position based on a particle filter. We have evaluated our algorithm in multiple scenarios consisting of static and dynamic NLOS sound sources. In our tested cases, our approach can localize a source position with an average accuracy error of 0.7m, measured by the L2 distance between estimated and actual source locations in a 7m×7m×3m room. Furthermore, we observe 37% to 130% improvement in accuracy over a state-of-the-art localization method that does not model diffraction effects, especially when a sound source is not visible to the robot. Inkyu An, Doheon Lee, Jung-Woo Choi, Dinesh Manocha, Sung-Eui Yoon |
ICRA | 2 |
| 2019 | Deconvoluting essential gene signatures for cancer growth from genomic expression in compound-treated cellsabstractMOTIVATION: Essential gene signatures for cancer growth have been typically identified via RNAi or CRISPR-Cas9. Here, we propose an alternative method that reveals the essential gene signatures by analysing genomic expression profiles in compound-treated cells. With a large amount of the existing compound-induced data, essential gene signatures at genomic scale are efficiently characterized without technical challenges in the previous techniques. RESULTS: An essential gene is characterized as a gene presenting positive correlation between its down-regulation and cell growth inhibition induced by diverse compounds, which were collected from LINCS and CGP. Among 12 741 genes, 1092, 1 228 827 962, 1 664 580 and 829 essential genes are characterized for each of A375, A549, BT20, LNCAP, MCF7, MDAMB231 and PC3 cell lines (P-value ≤ 1.0E-05). Comparisons to the previously identified essential genes yield significant overlaps in A375 and A549 (P-value ≤ 5.0E-05) and the 103 common essential genes are enriched in crucial processes for cancer growth. In most comparisons in A375, MCF7, BT20 and A549, the characterized essential genes yield more essential characteristics than those of the previous techniques, i.e. high gene expression, high degrees of protein-protein interactions, many homologs and few paralogs. Remarkably, the essential genes commonly characterized by both the previous and proposed techniques show more significant essential characteristics than those solely relied on the previous techniques. We expect that this work provides new aspects in essential gene signatures. AVAILABILITY AND IMPLEMENTATION: The Python implementations are available at https://github.com/jmjung83/deconvolution_of_essential_gene_signitures. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jinmyung Jung, Yeeok Kang, Hyojung Paik, Mijin Kwon, Hasun Yu, Doheon Lee |
Bioinform. | 6 |
| 2019 | CODA-ML: context-specific biological knowledge representation for systemic physiology analysisabstractBACKGROUND: Computational analysis of complex diseases involving multiple organs requires the integration of multiple different models into a unified model. Different models are often constructed in heterogeneous formats. Thus, the integration of the models requires a standard language format that can effectively represent essential biological information. However, the previously introduced formats have limitations that prevent from adequately representing essential biological information, particularly specifications of bio-molecules and biological contexts. RESULTS: We defined an XML-based markup language called context-oriented directed association markup language (CODA-ML), which better represents essential biological information. The CODA-ML has two major strengths in designating molecular specifications and biological contexts. It can cover heterogeneous entity types involved in biological events (e.g. gene/protein, compound, cellular function, disease). Molecular types of entities can have molecular specifications which include detailed information of a molecule from isoforms to modifications, enabling high-resolution representation of molecules. In addition, it can distinguish biological events that vary depending on different biological contexts such as cell types or disease conditions. Especially representation of inter-cellular events as well as intra-cellular events is available. These two major strengths can resolve contradictory associations when different models are integrated into one unified model, which improves the accuracy of the model. CONCLUSIONS: With the CODA-ML, diverse models such as signaling pathways, metabolic pathways, and gene regulatory pathways can be represented in a unified language format. Heterogeneous entity types can be covered by the CODA-ML, thus it enables detailed description for the mechanisms of diseases or drugs from multiple perspectives (e.g., molecule, function or disease). The CODA-ML is expected to help integrate different models into one systemic model in an efficient and effective. The unified model can be used to perform computational analysis not only for cancer but also for other complex diseases involving multiple organs beyond a single cell. Mijin Kwon, Soorin Yim, Gwangmin Kim, Saehwan Lee, Chungsun Jeong, Doheon Lee |
BMC Bioinform. | 6 |
| 2019 | Concept embedding to measure semantic relatedness for biomedical information ontologiesabstractThere have been many attempts to identify relationships among concepts corresponding to terms from biomedical information ontologies such as the Unified Medical Language System (UMLS). In particular, vector representation of such concepts using information from UMLS definition texts is widely used to measure the relatedness between two biological concepts. However, conventional relatedness measures have a limited range of applicable word coverage, which limits the performance of these models. In this paper, we propose a concept-embedding model of a UMLS semantic relatedness measure to overcome the limitations of earlier models. We obtained context texts of biological concepts that are not defined in UMLS by utilizing Wikipedia as an external knowledgebase. Concept vector representations were then derived from the context texts of the biological concepts. The degree of relatedness between two concepts was defined as the cosine similarity between corresponding concept vectors. As a result, we validated that our method provides higher coverage and better performance than the conventional method. Woochang Hwang, Doheon Lee |
J. Biomed. Informatics | 4 |
| 2018 | CORUS: Blockchain-Based Trustworthy Evaluation System for Efficacy of Healthcare RemediesabstractWe present a healthcare remedy evaluation system, called CORUS, using blockchain-based crowdsourcing on a cloud computing platform. Efficacy of healthcare remedies such as functional foods and dietary supplements usually spreads by subjective word of mouth. Though regular clinical trials could be employed, it requires significant cost in time and budget. We expect crowdsourcing can be an effective and efficient alternative to the costly clinical trials to achieve objective evaluation on healthcare remedies. The blockchain technology can alleviate the problem of information quality in most of crowdsourcing systems. Properly designed reward distribution schemes can motivate voluntary participants to provide with credible information. The distributed replication of blockchains can increase the information credibility by prohibiting malicious forgery. We also utilize the Amazon cloud system to enhance availability and scalability. To our knowledge, CORUS is the first system to utilize crowdsourcing, blockchains, and cloud computing for healthcare remedy evaluation. It is available at https://corus.kaist.edu. Seongkuk Park, Doheon Lee |
CloudCom | 4 |
| 2018 | Identification of common coexpression modules based on quantitative network comparisonabstractBACKGROUND: Finding common molecular interactions from different samples is essential work to understanding diseases and other biological processes. Coexpression networks and their modules directly reflect sample-specific interactions among genes. Therefore, identification of common coexpression network or modules may reveal the molecular mechanism of complex disease or the relationship between biological processes. However, there has been no quantitative network comparison method for coexpression networks and we examined previous methods for other networks that cannot be applied to coexpression network. Therefore, we aimed to propose quantitative comparison methods for coexpression networks and to find common biological mechanisms between Huntington's disease and brain aging by the new method. RESULTS: We proposed two similarity measures for quantitative comparison of coexpression networks. Then, we performed experiments using known coexpression networks. We showed the validity of two measures and evaluated threshold values for similar coexpression network pairs from experiments. Using these similarity measures and thresholds, we quantitatively measured the similarity between disease-specific and aging-related coexpression modules and found similar Huntington's disease-aging coexpression module pairs. CONCLUSIONS: We identified similar Huntington's disease-aging coexpression module pairs and found that these modules are related to brain development, cell death, and immune response. It suggests that up-regulated cell signalling related cell death and immune/ inflammation response may be the common molecular mechanisms in the pathophysiology of HD and normal brain aging in the frontal cortex. Yousang Jo, Doheon Lee |
BMC Bioinform. | 3 |
| 2018 | A systematic approach to identify therapeutic effects of natural products based on human metabolite informationabstractBACKGROUND: Natural products have been widely investigated in the drug development field. Their traditional use cases as medicinal agents and their resemblance of our endogenous compounds show the possibility of new drug development. Many researchers have focused on identifying therapeutic effects of natural products, yet the resemblance of natural products and human metabolites has been rarely touched. METHODS: We propose a novel method which predicts therapeutic effects of natural products based on their similarity with human metabolites. In this study, we compare the structure, target and phenotype similarities between natural products and human metabolites to capture molecular and phenotypic properties of both compounds. With the generated similarity features, we train support vector machine model to identify similar natural product and human metabolite pairs. The known functions of human metabolites are then mapped to the paired natural products to predict their therapeutic effects. RESULTS: With our selected three feature sets, structure, target and phenotype similarities, our trained model successfully paired similar natural products and human metabolites. When applied to the natural product derived drugs, we could successfully identify their indications with high specificity and sensitivity. We further validated the found therapeutic effects of natural products with the literature evidence. CONCLUSIONS: These results suggest that our model can match natural products to similar human metabolites and provide possible therapeutic effects of natural products. By utilizing the similar human metabolite information, we expect to find new indications of natural products which could not be covered by previous in silico methods. Kyungrin Noh, Sunyong Yoo, Doheon Lee |
BMC Bioinform. | 3 |
| 2018 | CONET: a virtual human system-centered platform for drug discovery
Doheon Lee |
Frontiers Comput. Sci. | 1 |
| 2018 | Predicting the Absorption Potential of Chemical Compounds Through a Deep Learning ApproachabstractThe human colorectal carcinoma cell line (Caco-2) is a commonly used in-vitro test that predicts the absorption potential of orally administered drugs. In-silico prediction methods, based on the Caco-2 assay data, may increase the effectiveness of the high-throughput screening of new drug candidates. However, previously developed in-silico models that predict the Caco-2 cellular permeability of chemical compounds use handcrafted features that may be dataset-specific and induce over-fitting problems. Deep Neural Network (DNN) generates high-level features based on non-linear transformations for raw features, which provides high discriminant power and, therefore, creates a good generalized model. We present a DNN-based binary Caco-2 permeability classifier. Our model was constructed based on 663 chemical compounds with in-vitro Caco-2 apparent permeability data. Two hundred nine molecular descriptors are used for generating the high-level features during DNN model generation. Dropout regularization is applied to solve the over-fitting problem and the non-linear activation. The Rectified Linear Unit (ReLU) is adopted to reduce the vanishing gradient problem. The results demonstrate that the high-level features generated by the DNN are more robust than handcrafted features for predicting the cellular permeability of structurally diverse chemical compounds in Caco-2 cell lines. Moonshik Shin, Donjin Jang, Hojung Nam, Kwang Hyung Lee, Doheon Lee |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2017 | Coupling effects on turning points of infectious diseases epidemics in scale-free networksabstractBACKGROUND: Pandemic is a typical spreading phenomenon that can be observed in the human society and is dependent on the structure of the social network. The Susceptible-Infective-Recovered (SIR) model describes spreading phenomena using two spreading factors; contagiousness (β) and recovery rate (γ). Some network models are trying to reflect the social network, but the real structure is difficult to uncover. METHODS: We have developed a spreading phenomenon simulator that can input the epidemic parameters and network parameters and performed the experiment of disease propagation. The simulation result was analyzed to construct a new marker VRTP distribution. We also induced the VRTP formula for three of the network mathematical models. RESULTS: We suggest new marker VRTP (value of recovered on turning point) to describe the coupling between the SIR spreading and the Scale-free (SF) network and observe the aspects of the coupling effects with the various of spreading and network parameters. We also derive the analytic formulation of VRTP in the fully mixed model, the configuration model, and the degree-based model respectively in the mathematical function form for the insights on the relationship between experimental simulation and theoretical consideration. CONCLUSIONS: We discover the coupling effect between SIR spreading and SF network through devising novel marker VRTP which reflects the shifting effect and relates to entropy. Kiseong Kim, Sangyeon Lee, Doheon Lee, Kwang Hyung Lee |
BMC Bioinform. | 3 |
| 2017 | Recent advances in immunological inspired computation
Carlos A. Coello Coello, Vincenzo Cutello, Doheon Lee, Mario Pavone |
Eng. Appl. Artif. Intell. | 3 |
| 2016 | DTMBIO 2016: The Tenth International Workshop on Data and Text Mining in Biomedical InformaticsabstractStarted in 2006 as a specialized workshop in the field of text mining applied to biomedical informatics, DTMBIO (ACM international workshop on Data and Text Mining in Biomedical Informatics) has been held annually in conjunction with one of the largest data management conferences, CIKM, bringing together researchers working on computer science and bioinformatics area. The purpose of DTMBIO is to foster discussions regarding the state-of-the-art applications of data and text mining on biomedical research problems. DTMBIO 2016 will help scientists navigate emerging trends and opportunities in the evolving area of informatics related techniques and problems in the context of biomedical research. Sangwoo Kim, Jake Yue Chen, Vincenzo Cutello, Doheon Lee |
CIKM | 4 |
| 2016 | A corpus for plant-chemical relationships in the biomedical domainabstractBACKGROUND: Plants are natural products that humans consume in various ways including food and medicine. They have a long empirical history of treating diseases with relatively few side effects. Based on these strengths, many studies have been performed to verify the effectiveness of plants in treating diseases. It is crucial to understand the chemicals contained in plants because these chemicals can regulate activities of proteins that are key factors in causing diseases. With the accumulation of a large volume of biomedical literature in various databases such as PubMed, it is possible to automatically extract relationships between plants and chemicals in a large-scale way if we apply a text mining approach. A cornerstone of achieving this task is a corpus of relationships between plants and chemicals. RESULTS: In this study, we first constructed a corpus for plant and chemical entities and for the relationships between them. The corpus contains 267 plant entities, 475 chemical entities, and 1,007 plant-chemical relationships (550 and 457 positive and negative relationships, respectively), which are drawn from 377 sentences in 245 PubMed abstracts. Inter-annotator agreement scores for the corpus among three annotators were measured. The simple percent agreement scores for entities and trigger words for the relationships were 99.6 and 94.8 %, respectively, and the overall kappa score for the classification of positive and negative relationships was 79.8 %. We also developed a rule-based model to automatically extract such plant-chemical relationships. When we evaluated the rule-based model using the corpus and randomly selected biomedical articles, overall F-scores of 68.0 and 61.8 % were achieved, respectively. CONCLUSION: We expect that the corpus for plant-chemical relationships will be a useful resource for enhancing plant research. The corpus is available at http://combio.gist.ac.kr/plantchemicalcorpus . Wonjun Choi, Baeksoo Kim, Hyejin Cho, Doheon Lee |
BMC Bioinform. | 4 |
| 2016 | Context-specific functional module based drug efficacy predictionabstractBACKGROUND: It is necessary to evaluate the efficacy of individual drugs on patients to realize personalized medicine. Testing drugs on patients in clinical trial is the only way to evaluate the efficacy of drugs. The approach is labour intensive and requires overwhelming costs and a number of experiments. Therefore, preclinical model system has been intensively investigated for predicting the efficacy of drugs. Current computational drug sensitivity prediction approaches use general biological network modules as their prediction features. Therefore, they miss indirect effectors or the effects from tissue-specific interactions. RESULTS: We developed cell line specific functional modules. Enriched scores of functional modules are utilized as cell line specific features to predict the efficacy of drugs. Cell line specific functional modules are clusters of genes, which have similar biological functions in cell line specific networks. We used linear regression for drug efficacy prediction. We assessed the prediction performance in leave-one-out cross-validation (LOOCV). Our method was compared with elastic net model, which is a popular model for drug efficacy prediction. In addition, we analysed drug sensitivity-associated functions of five drugs - lapatinib, erlotinib, raloxifene, tamoxifen and gefitinib- by our model. CONCLUSIONS: Our model can provide cell line specific drug efficacy prediction and also provide functions which are associated with drug sensitivity. Therefore, we could utilize drug sensitivity associated functions for drug repositioning or for suggesting secondary drugs for overcoming drug resistance. Woochang Hwang, Jaejoon Choi, Mijin Kwon, Doheon Lee |
BMC Bioinform. | 4 |
| 2016 | Prediction of compound-target interactions of natural products using large-scale drug and protein informationabstractBACKGROUND: Verifying the proteins that are targeted by compounds of natural herbs will be helpful to select natural herb-based drug candidates. However, this entails a great deal of effort to clarify the interaction throughout in vitro or in vivo experiments. In this light, in silico prediction of the interactions between compounds and target proteins can help ease the efforts. RESULTS: In this study, we performed in silico predictions of herbal compound target identification. First, data related to compounds, target proteins, and interactions between them are taken from the DrugBank database. Then we characterized six classes of compound-target interaction in humans including G-protein-coupled receptors (GPCRs), ion channel, enzymes, receptors, transporters, and other proteins. Also, classification-prediction models that predict the interactions between compounds and target proteins through a machine learning method were constructed using these matrices. As a result, AUC values of six classes are 0.94, 0.93, 0.90, 0.89, 0.91, and 0.76 respectively. Finally, the interactions of compounds from natural products were predicted using the constructed classification models. Furthermore, from our predicted results, we confirmed that several important disease related proteins were predicted as targets of natural herbal compounds. CONCLUSIONS: We constructed classification-prediction models that predict the interactions between compounds and target proteins. The constructed models showed good prediction performances, and numbers of potential natural compounds target proteins were predicted from our results. Jongsoo Keum, Sunyong Yoo, Doheon Lee, Hojung Nam |
BMC Bioinform. | 3 |
| 2016 | Inferring new drug indications using the complementarity between clinical disease signatures and drug effectsabstractBACKGROUND: Drug repositioning is the process of finding new indications for existing drugs. Its importance has been dramatically increasing recently due to the enormous increase in new drug discovery cost. However, most of the previous molecular-centered drug repositioning work is not able to reflect the end-point physiological activities of drugs because of the inherent complexity of human physiological systems. METHODS: Here, we suggest a novel computational framework to make inferences for alternative indications of marketed drugs by using electronic clinical information which reflects the end-point physiological results of drug's effects on the biological activities of humans. In this work, we use the concept of complementarity between clinical disease signatures and clinical drug effects. With this framework, we establish disease-related clinical variable vectors (clinical disease signature vectors) and drug-related clinical variable vectors (clinical drug effect vectors) by applying two methodologies (i.e., statistical analysis and literature mining). Finally, we assign a repositioning possibility score to each disease-drug pair by the calculation of complementarity (anti-correlation) and association between clinical states ("up" or "down") of disease signatures and clinical effects ("up", "down" or "association") of drugs. A total of 717 clinical variables in the electronic clinical dataset (NHANES), are considered in this study. RESULTS: The statistical significance of our prediction results is supported through two benchmark datasets (Comparative Toxicogenomics Database and Clinical Trials). We discovered not only lots of known relationships between diseases and drugs, but also many hidden disease-drug relationships. For example, glutathione and edetic-acid may be investigated as candidate drugs for asthma treatment. We examined prediction results by using statistical experiments (enrichment verification, hyper-geometric and permutation test P<0.009 in Comparative Toxicogenomics Database and Clinical Trials) and presented evidences for those with already published literature. CONCLUSION: The results show that electronic clinical information is a feasible data resource and utilizing the complementarity (anti-correlated relationships) between clinical signatures of disease and clinical effects of drugs is a potentially predictive concept in drug repositioning research. It makes the proposed approach useful to identity novel relationships between diseases and drugs that have a high probability of being biologically valid. Dongjin Jang, Sejoon Lee, Kiseong Kim, Doheon Lee |
J. Biomed. Informatics | 5 |
| 2015 | DTMBIO 2015: International Workshop on Data and Text Mining in Biomedical InformaticsabstractHeld each year in conjunction with one of the largest data management conferences, CIKM, the Ninth ACM International Workshop on Data and Text Mining in Biomedical Informatics (DTMBIO'15) is organized to bring together researchers interested in development and application of cutting-edge data management and analysis methods with a specific focus on applications in biology and medicine. The purpose of DTMBIO is to foster discussions regarding the state-of-the-art applications of data and text mining on biomedical research problems. DTMBIO'15 will help scientists understand emerging trends and opportunities in the evolving area of informatics related techniques and problems in the context of biomedical research. Min Song 0001, Doheon Lee, Karin Verspoor |
CIKM | 2 |
| 2015 | SoloDel: a probabilistic model for detecting low-frequent somatic deletions from unmatched sequencing dataabstractMOTIVATION: Finding somatic mutations from massively parallel sequencing data is becoming a standard process in genome-based biomedical studies. There are a number of robust methods developed for detecting somatic single nucleotide variations However, detection of somatic copy number alteration has been substantially less explored and remains vulnerable to frequently raised sampling issues: low frequency in cell population and absence of the matched control samples. RESULTS: We developed a novel computational method SoloDel that accurately classifies low-frequent somatic deletions from germline ones with or without matched control samples. We first constructed a probabilistic, somatic mutation progression model that describes the occurrence and propagation of the event in the cellular lineage of the sample. We then built a Gaussian mixture model to represent the mixed population of somatic and germline deletions. Parameters of the mixture model could be estimated using the expectation-maximization algorithm with the observed distribution of read-depth ratios at the points of discordant-read based initial deletion calls. Combined with conventional structural variation caller, SoloDel greatly increased the accuracy in classifying somatic mutations. Even without control, SoloDel maintained a comparable performance in a wide range of mutated subpopulation size (10-70%). SoloDel could also successfully recall experimentally validated somatic deletions from previously reported neuropsychiatric whole-genome sequencing data. AVAILABILITY AND IMPLEMENTATION: Java-based implementation of the method is available at http://sourceforge.net/projects/solodel/ CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Hojung Nam, Sangwoo Kim, Doheon Lee |
Bioinform. | 5 |
| 2014 | DTMBIO 2014: International Workshop on Data and Text Mining in Biomedical InformaticsabstractHeld each year in conjunction with one of the largest data management conferences, CIKM, the Eighth ACM International Workshop on Data and Text Mining in Biomedical Informatics (DTMBIO 14) is organized to bring together researchers interested in development and application of cutting-edge biomedical and healthcare technology. The purpose of DTMBIO is to foster discussions regarding the state-of-the-art applications of data and text mining on biomedical research problems. DTMBIO 14 will help scientists navigate emerging trends and opportunities in the evolving area of informatics related techniques and problems in the context of biomedical research. Luonan Chen, Doheon Lee, Hua Xu 0001, Min Song 0001 |
CIKM | 2 |
| 2013 | DTMBIO 2013: international workshop on data and text mining in biomedical informaticsabstractThe organizers of ACM Seventh International Workshop on Data and Text Mining in Biomedical Informatics (DTMBIO 13) are pleased to announce that the seventh DTMBIO will be held in conjunction with CIKM, one of the largest data management conferences. The major interests of DTMBIO are on the state-of-the-art applications of data and text mining on biomedical research problems. DTMBIO 13 will be a forum of discussing and exchanging informatics related techniques and problems in the context of biomedical research. Atul J. Butte, Doheon Lee, Hua Xu 0001, Min Song 0001 |
CIKM | 2 |
| 2013 | Inferring disease association using clinical factors in a combinatorial manner and their use in drug repositioningabstractMOTIVATION: Complex physiological relationships exist among human diseases. Thus, the identification of disease associations could provide new methods of disease care and diagnosis. To this end, numerous studies have investigated disease associations. However, combinatorial effect of physiological factors, which is the main characteristic of biological systems, has not been considered in most previous studies. RESULTS: In this study, we inferred disease associations with a novel approach that considered disease-related clinical factors in combinatorial ways by using the National Health and Nutrition Examination Survey data, and the results have been shown as disease networks. Here, the FP-growth algorithm, an association rule mining algorithm, was used to generate a clinical attribute combination profile of each disease. In addition, we characterized the 22 clinical risk attribute combinations frequently discovered from the 26 diseases in this study. Furthermore, we validated that the results of this study have great potential for drug repositioning and outperform other existing disease networks in this regard. Finally, we suggest a few disease pairs as new candidates for drug repositioning and provide the evidence of their associations from the literature. Jinmyung Jung, Doheon Lee |
Bioinform. | 2 |
| 2012 | DTMBIO 2012: international workshop on data and text mining in biomedical informaticsabstractThe organizers of ACM Sixth International Workshop on Data and Text Mining in Biomedical Informatics (DTMBIO 12) are happy announce that the sixth DTMBIO will be held in conjunction with CIKM, one of the largest data management conferences. The major interests of DTMBIO are on the state-of-the-art applications of data and text mining on biomedical research problems. DTMBIO 12 will be a forum of discussing and exchanging informatics related techniques and problems in the context of biomedical research. Min Song 0001, Doheon Lee, Hua Xu 0001, Sophia Ananiadou |
CIKM | 2 |
| 2011 | DTMBIO 2011: international workshop on data and textmining in biomedical informaticsabstractACM Fifth International Workshop on Data and Text Mining in Biomedical Informatics (DTMBIO 11) organizers are pleased to announce that the fifth DTMBIO will be held in conjunction with CIKM, one of the largest data and text mining conferences. While CIKM presents the state-of-the-art research in informatics with the primary focus on data and text mining, the main focus of DTMBIO is on biomedical informatics. DTMBIO delegates will bring forth interesting applications of up-to-date informatics in the context of biomedical research. Sophia Ananiadou, Doheon Lee, Shamkant B. Navathe, Min Song 0001 |
CIKM | 2 |
| 2011 | Context-dependent transcriptional regulations between signal transduction pathwaysabstractBACKGROUND: Cells coordinate their metabolism, proliferation, and cellular communication according to environmental cues through signal transduction. Because signal transduction has a primary role in cellular processes, many experimental techniques and approaches have emerged to discover the molecular components and dynamics that are dependent on cellular contexts. However, omics approaches based on genome-wide expression analysis data comparing one differing condition (e.g. complex disease patients and normal subjects) did not investigate the dynamics and inter-pathway cross-communication that are dependent on cellular contexts. Therefore, we introduce a new computational omics approach for discovering signal transduction pathways regulated by transcription and transcriptional regulations between pathways in signaling networks that are dependent on cellular contexts, especially focusing on a transcription-mediated mechanism of inter-pathway cross-communication. RESULTS: Applied to dendritic cells treated with lipopolysaccharide, our analysis well depicted how dendritic cells respond to the treatment through transcriptional regulations between signal transduction pathways in dendritic cell maturation and T cell activation. CONCLUSIONS: Our new approach helps to understand the underlying biological phenomenon of expression data (e.g. complex diseases such as cancer) by providing a graphical network which shows transcriptional regulations between signal transduction pathways. The software programs are available upon request. Sohyun Hwang, Sangwoo Kim, Heesung Shin, Doheon Lee |
BMC Bioinform. | 4 |
| 2011 | Building the process-drug-side effect network to discover the relationship between biological Processes and side effectsabstractBACKGROUND: Side effects are unwanted responses to drug treatment and are important resources for human phenotype information. The recent development of a database on side effects, the side effect resource (SIDER), is a first step in documenting the relationship between drugs and their side effects. It is, however, insufficient to simply find the association of drugs with biological processes; that relationship is crucial because drugs that influence biological processes can have an impact on phenotype. Therefore, knowing which processes respond to drugs that influence the phenotype will enable more effective and systematic study of the effect of drugs on phenotype. To the best of our knowledge, the relationship between biological processes and side effects of drugs has not yet been systematically researched. METHODS: We propose 3 steps for systematically searching relationships between drugs and biological processes: enrichment scores (ES) calculations, t-score calculation, and threshold-based filtering. Subsequently, the side effect-related biological processes are found by merging the drug-biological process network and the drug-side effect network. Evaluation is conducted in 2 ways: first, by discerning the number of biological processes discovered by our method that co-occur with Gene Ontology (GO) terms in relation to effects extracted from PubMed records using a text-mining technique and second, determining whether there is improvement in performance by limiting response processes by drugs sharing the same side effect to frequent ones alone. RESULTS: The multi-level network (the process-drug-side effect network) was built by merging the drug-biological process network and the drug-side effect network. We generated a network of 74 drugs-168 side effects-2209 biological process relation resources. The preliminary results showed that the process-drug-side effect network was able to find meaningful relationships between biological processes and side effects in an efficient manner. CONCLUSIONS: We propose a novel process-drug-side effect network for discovering the relationship between biological processes and side effects. By exploring the relationship between drugs and phenotypes through a multi-level network, the mechanisms underlying the effect of specific drugs on the human body may be understood. Sejoon Lee, Kwang Hyung Lee, Min Song 0001, Doheon Lee |
BMC Bioinform. | 4 |
| 2010 | A new perspective of integrative genome-wide association analysis considering trans eSNP effectabstractMost loci discovered through genome-wide association analyses are predicted to affect gene expression, the integrative approach of genome-wide analysis with gene expression data is becoming essential procedure for discovering genetic effect of disease development. Many studies have been performed to discover significant SNP-gene associations, but most of them are limited to consider only cis-associations and neglect trans-territory. In this study, we explored the effect of trans-eSNP associations that may underlie the alternation of gene expressions. Through the integrative genome-wide association analysis considering the entire SNP-gene associations, we identified numerous trans associations which are significantly associated with gene expression, even more than cis associations in quantity and significance. Our findings revealed the necessity of reconsidering trans-association effect from integrative genome-wide association analysis and provided novel insights to find undiscovered genetic causalities. Doheon Lee |
BIBM | 2 |
| 2010 | DTMBIO workshop summaryabstractNo abstract available. Hagit Shatkay, Doheon Lee, Min Song 0001, Shamkant B. Navathe |
CIKM | 2 |
| 2010 | RBSDesigner: software for designing synthetic ribosome binding sites that yields a desired level of protein expressionabstractMOTIVATION: RBSDesigner predicts the translation efficiency of existing mRNA sequences and designs synthetic ribosome binding sites (RBSs) for a given coding sequence (CDS) to yield a desired level of protein expression. The program implements the mathematical model for translation initiation described in Na et al. (Mathematical modeling of translation initiation for the estimation of its efficiency to computationally design mRNA sequences with a desired expression level in prokaryotes. BMC Syst. Biol., 4, 71). The program additionally incorporates the effect on translation efficiency of the spacer length between a Shine-Dalgarno (SD) sequence and an AUG codon, which is crucial for the incorporation of fMet-tRNA into the ribosome. RBSDesigner provides a graphical user interface (GUI) for the convenient design of synthetic RBSs. AVAILABILITY: RBSDesigner is written in Python and Microsoft Visual Basic 6.0 and is publicly available as precompiled stand-alone software on the web (http://rbs.kaist.ac.kr). CONTACT: [email protected] Dokyun Na, Doheon Lee |
Bioinform. | 2 |
| 2010 | Inference of combinatorial Boolean rules of synergistic gene sets from cancer microarray datasetsabstractMOTIVATION: Gene set analysis has become an important tool for the functional interpretation of high-throughput gene expression datasets. Moreover, pattern analyses based on inferred gene set activities of individual samples have shown the ability to identify more robust disease signatures than individual gene-based pattern analyses. Although a number of approaches have been proposed for gene set-based pattern analysis, the combinatorial influence of deregulated gene sets on disease phenotype classification has not been studied sufficiently. RESULTS: We propose a new approach for inferring combinatorial Boolean rules of gene sets for a better understanding of cancer transcriptome and cancer classification. To reduce the search space of the possible Boolean rules, we identify small groups of gene sets that synergistically contribute to the classification of samples into their corresponding phenotypic groups (such as normal and cancer). We then measure the significance of the candidate Boolean rules derived from each group of gene sets; the level of significance is based on the class entropy of the samples selected in accordance with the rules. By applying the present approach to publicly available prostate cancer datasets, we identified 72 significant Boolean rules. Finally, we discuss several identified Boolean rules, such as the rule of glutathione metabolism (down) and prostaglandin synthesis regulation (down), which are consistent with known prostate cancer biology. AVAILABILITY: Scripts written in Python and R are available at http://biosoft.kaist.ac.kr/~ihpark/. The refined gene sets and the full list of the identified Boolean rules are provided in the Supplementary Material. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Kwang Hyung Lee, Doheon Lee |
Bioinform. | 3 |
| 2010 | PSExplorer: whole parameter space exploration for molecular signaling pathway dynamicsabstractMOTIVATION: Mathematical models of biological systems often have a large number of parameters whose combinational variations can yield distinct qualitative behaviors. Since it is intractable to examine all possible combinations of parameters for non-trivial biological pathways, it is required to have a systematic strategy to explore the parameter space in a computational way so that dynamic behaviors of a given pathway are estimated. RESULTS: We present PSExplorer, a computational tool for exploring qualitative behaviors and key parameters of molecular signaling pathways. Utilizing the Latin hypercube sampling and a clustering technique in a recursive paradigm, the software enables users to explore the whole parameter space of the models to search for robust qualitative behaviors. The parameter space is partitioned into sub-regions according to behavioral differences. Sub-regions showing robust behaviors can be identified for further analyses. The partitioning result presents a tree structure from which individual and combinational effects of parameters on model behaviors can be assessed and key factors of the models are readily identified. AVAILABILITY: The software, tutorial manual and test models are available for download at the following address: http://gto.kaist.ac.kr/∼psexplorer. Thai Quang Tung, Doheon Lee |
Bioinform. | 2 |
| 2010 | MKEM: a Multi-level Knowledge Emergence Model for mining undiscovered public knowledgeabstractBACKGROUND: Since Swanson proposed the Undiscovered Public Knowledge (UPK) model, there have been many approaches to uncover UPK by mining the biomedical literature. These earlier works, however, required substantial manual intervention to reduce the number of possible connections and are mainly applied to disease-effect relation. With the advancement in biomedical science, it has become imperative to extract and combine information from multiple disjoint researches, studies and articles to infer new hypotheses and expand knowledge. METHODS: We propose MKEM, a Multi-level Knowledge Emergence Model, to discover implicit relationships using Natural Language Processing techniques such as Link Grammar and Ontologies such as Unified Medical Language System (UMLS) MetaMap. The contribution of MKEM is as follows: First, we propose a flexible knowledge emergence model to extract implicit relationships across different levels such as molecular level for gene and protein and Phenomic level for disease and treatment. Second, we employ MetaMap for tagging biological concepts. Third, we provide an empirical and systematic approach to discover novel relationships. RESULTS: We applied our system on 5000 abstracts downloaded from PubMed database. We performed the performance evaluation as a gold standard is not yet available. Our system performed with a good precision and recall and we generated 24 hypotheses. CONCLUSIONS: Our experiments show that MKEM is a powerful tool to discover hidden relationships residing in extracted entities that were represented by our Substance-Effect-Process-Disease-Body Part (SEPDB) model. Ali Zeeshan Ijaz, Min Song 0001, Doheon Lee |
BMC Bioinform. | 3 |
| 2010 | Multivariate classification of urine metabolome profiles for breast cancer diagnosisabstractBACKGROUND: Diagnosis techniques using urine are non-invasive, inexpensive, and easy to perform in clinical settings. The metabolites in urine, as the end products of cellular processes, are closely linked to phenotypes. Therefore, urine metabolome is very useful in marker discoveries and clinical applications. However, only univariate methods have been used in classification studies using urine metabolome. Since multiple genes or proteins would be involved in developments of complex diseases such as breast cancer, multiple compounds including metabolites would be related with the complex diseases, and multivariate methods would be needed to identify those multiple metabolite markers. Moreover, because combinatorial effects among the markers can seriously affect disease developments and there also exist individual differences in genetic makeup or heterogeneity in cancer progressions, single marker is not enough to identify cancers. RESULTS: We proposed classification models using multivariate classification techniques and developed an analysis procedure for classification studies using metabolome data. Through this strategy, we identified five potential urinary biomarkers for breast cancer with high accuracy, among which the four biomarker candidates were not identifiable by only univariate methods. We also proposed potential diagnosis rules to help in clinical decision making. Besides, we showed that combinatorial effects among multiple biomarkers can enhance discriminative power for breast cancer. CONCLUSIONS: In this study, we successfully showed that multivariate classifications are needed to precisely diagnose breast cancer. After further validation with independent cohorts and experimental confirmation, these marker candidates will likely lead to clinically applicable assays for earlier diagnoses of breast cancer. Imhoi Koo, Byung Hwa Jung, Bong Chul Chung, Doheon Lee |
BMC Bioinform. | 5 |
| 2009 | Disease Classification Based on the Activities of Interacting Molecular Modules with Condition-Responsive CorrelationabstractGenome-wide expression profiles of diseased samples have been exploited to predict disease states. Recently, network-based approaches utilizing molecular interaction networks integrated with gene expression profiles have been proposed to address challenges which arise from smaller number of samples compared to the large number of predictors, and genetic heterogeneity of samples in complex diseases such as cancer. However, previous network-based methods only focus on expression levels of proteins, nodes in the network though the identification of condition-responsive interactions, edges under the phenotype of interest must enlighten another aspect of pathogenic processes. Thus, we propose a novel network-based classification which focuses on both nodes with discriminative expression levels and edges with condition-responsive correlations across two phenotypes. The extracted modules with condition-responsive interactions not only provide candidate molecular models for disease, and their activities inferred from a subset of member genes serve as better predictors in classification compared to the conventional gene-centric method. Sejoon Lee, Kwang Hyung Lee, Doheon Lee |
BIBM | 4 |
| 2009 | Combining tissue transcriptomics and urine metabolomics for breast cancer biomarker identificationabstractMOTIVATION: For the early detection of cancer, highly sensitive and specific biomarkers are needed. Particularly, biomarkers in bio-fluids are relatively more useful because those can be used for non-biopsy tests. Although the altered metabolic activities of cancer cells have been observed in many studies, little is known about metabolic biomarkers for cancer screening. In this study, a systematic method is proposed for identifying metabolic biomarkers in urine samples by selecting candidate biomarkers from altered genome-wide gene expression signatures of cancer cells. Biomarkers identified by the present study have increased coherence and robustness because the significances of biomarkers are validated in both gene expression profiles and metabolic profiles. RESULTS: The proposed method was applied to the gene expression profiles and urine samples of 50 breast cancer patients and 50 normal persons. Nine altered metabolic pathways were identified from the breast cancer gene expression signatures. Among these altered metabolic pathways, four metabolic biomarkers (Homovanillate, 4-hydroxyphenylacetate, 5-hydroxyindoleacetate and urea) were identified to be different in cancer and normal subjects (p <0.05). In the case of the predictive performance, the identified biomarkers achieved area under the ROC curve values of 0.75, 0.79 and 0.79, according to a linear discriminate analysis, a random forest classifier and on a support vector machine, respectively. Finally, biomarkers which showed consistent significance in pathways' gene expression as well as urine samples were identified. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Hojung Nam, Bong Chul Chung, Ki Young Lee, Doheon Lee |
Bioinform. | 5 |
| 2009 | Mining metastasis related genes by primary-secondary tumor comparisons from large-scale databasesabstractBACKGROUND: Metastasis is the most dangerous step in cancer progression and causes more than 90% of cancer death. Although many researchers have been working on biological features and characteristics of metastasis, most of its genetic level processes remain uncertain. Some studies succeeded in elucidating metastasis related genes and pathways, followed by predicting prognosis of cancer patients, but there still is a question whether the result genes or pathways contain enough information and noise features have been controlled appropriately. METHODS: We set four tumor type classes composed of various tumor characteristics such as tissue origin, cellular environment, and metastatic ability. We conducted a set of comparisons among the four tumor classes followed by searching for genes that are consistently up or down regulated through the whole comparisons. RESULTS: We identified four sets of genes that are consistently differently expressed in the comparisons, each of which denotes one of four cellular characteristics respectively - liver tissue, colon tissue, liver viability and metastasis characteristics. We found that our candidate genes for tissue specificity are consistent with the TiGER database. And we also found that the metastasis candidate genes from our method were more consistent with the known biological background and independent from other noise features. CONCLUSION: We suggested a new method for identifying metastasis related genes from a large-scale database. The proposed method attempts to minimize the influences from other factors except metastatic ability including tissue originality and tissue viability by confining the result of metastasis unrelated test combinations. Sangwoo Kim, Doheon Lee |
BMC Bioinform. | 2 |
| 2009 | Analysis of AML genes in dysregulated molecular networksabstractBACKGROUND: Identifying disease causing genes and understanding their molecular mechanisms are essential to developing effective therapeutics. Thus, several computational methods have been proposed to prioritize candidate disease genes by integrating different data types, including sequence information, biomedical literature, and pathway information. Recently, molecular interaction networks have been incorporated to predict disease genes, but most of those methods do not utilize invaluable disease-specific information available in mRNA expression profiles of patient samples. RESULTS: Through the integration of protein-protein interaction networks and gene expression profiles of acute myeloid leukemia (AML) patients, we identified subnetworks of interacting proteins dysregulated in AML and characterized known mutation genes causally implicated to AML embedded in the subnetworks. The analysis shows that the set of extracted subnetworks is a reservoir rich in AML genes reflecting key leukemogenic processes such as myeloid differentiation, CONCLUSION: We showed that the integrative approach both utilizing gene expression profiles and molecular networks could identify AML causing genes most of which were not detectable with gene expression analysis alone due to their minor changes in mRNA. Hyunchul Jung, Predrag Radivojac, Jong-Won Kim 0001, Doheon Lee |
BMC Bioinform. | 5 |
| 2009 | Protein comparison at the domain architecture levelabstractBACKGROUND: The general method used to determine the function of newly discovered proteins is to transfer annotations from well-characterized homologous proteins. The process of selecting homologous proteins can largely be classified into sequence-based and domain-based approaches. Domain-based methods have several advantages for identifying distant homology and homology among proteins with multiple domains, as compared to sequence-based methods. However, these methods are challenged by large families defined by 'promiscuous' (or 'mobile') domains. RESULTS: Here we present a measure, called Weighed Domain Architecture Comparison (WDAC), of domain architecture similarity, which can be used to identify homolog of multidomain proteins. To distinguish these promiscuous domains from conventional protein domains, we assigned a weight score to Pfam domain extracted from RefSeq proteins, based on its abundance and versatility. To measure the similarity of two domain architectures, cosine similarity (a similarity measure used in information retrieval) is used. We combined sequence similarity with domain architecture comparisons to identify proteins belonging to the same domain architecture. Using human and nematode proteomes, we compared WDAC with an unweighted domain architecture method (DAC) to evaluate the effectiveness of domain weight scores. We found that WDAC is better at identifying homology among multidomain proteins. CONCLUSION: Our analysis indicates that considering domain weight scores in domain architecture comparisons improves protein homology identification. We developed a web-based server to allow users to compare their proteins with protein domain architectures. Byungwook Lee, Doheon Lee |
BMC Bioinform. | 2 |
| 2009 | Identification of temporal association rules from time-series microarray data setsabstractBACKGROUND: One of the most challenging problems in mining gene expression data is to identify how the expression of any particular gene affects the expression of other genes. To elucidate the relationships between genes, an association rule mining (ARM) method has been applied to microarray gene expression data. However, a conventional ARM method has a limit on extracting temporal dependencies between gene expressions, though the temporal information is indispensable to discover underlying regulation mechanisms in biological pathways. In this paper, we propose a novel method, referred to as temporal association rule mining (TARM), which can extract temporal dependencies among related genes. A temporal association rule has the form [gene A upward arrow, gene B downward arrow] --> (7 min) [gene C upward arrow], which represents that high expression level of gene A and significant repression of gene B followed by significant expression of gene C after 7 minutes. The proposed TARM method is tested with Saccharomyces cerevisiae cell cycle time-series microarray gene expression data set. RESULTS: In the parameter fitting phase of TARM, the fitted parameter set [threshold = +/- 0.8, support >or= 3 transactions, confidence >or= 90%] with the best precision score for KEGG cell cycle pathway has been chosen for rule mining phase. With the fitted parameter set, numbers of temporal association rules with five transcriptional time delays (0, 7, 14, 21, 28 minutes) are extracted from gene expression data of 799 genes, which are pre-identified cell cycle relevant genes. From the extracted temporal association rules, associated genes, which play same role of biological processes within short transcriptional time delay and some temporal dependencies between genes with specific biological processes are identified. CONCLUSION: In this work, we proposed TARM, which is an applied form of conventional ARM. TARM showed higher precision score than Dynamic Bayesian network and Bayesian network. Advantages of TARM are that it tells us the size of transcriptional time delay between associated genes, activation and inhibition relationship between genes, and sets of co-regulators. Hojung Nam, Ki Young Lee, Doheon Lee |
BMC Bioinform. | 3 |
| 2009 | A method to improve protein subcellular localization prediction by integrating various biological data sourcesabstractBACKGROUND: Protein subcellular localization is crucial information to elucidate protein functions. Owing to the need for large-scale genome analysis, computational method for efficiently predicting protein subcellular localization is highly required. Although many previous works have been done for this task, the problem is still challenging due to several reasons: the number of subcellular locations in practice is large; distribution of protein in locations is imbalanced, that is the number of protein in each location remarkably different; and there are many proteins located in multiple locations. Thus it is necessary to explore new features and appropriate classification methods to improve the prediction performance. RESULTS: In this paper we propose a new predicting method which combines two key ideas: 1) Information of neighbour proteins in a probabilistic gene network is integrated to enrich the prediction features. 2) Fuzzy k-NN, a classification method based on fuzzy set theory is applied to predict protein locating in multiple sites. Experiment was conducted on a dataset consisting of 22 locations from Budding yeast proteins and significant improvement was observed. CONCLUSION: Our results suggest that the neighbourhood information from functional gene networks is predictive to subcellular localization. The proposed method thus can be integrated and complementary to other available prediction methods. Thai Quang Tung, Doheon Lee |
BMC Bioinform. | 2 |
| 2009 | Comparative analysis of the JAK/STAT signaling through erythropoietin receptor and thrombopoietin receptor using a systems approachabstractBACKGROUND: The Janus kinase-signal transducer and activator of transcription (JAK/STAT) pathway is one of the most important targets for myeloproliferative disorder (MPD). Although several efforts toward modeling the pathway using systems biology have been successful, the pathway was not fully investigated in regard to understanding pathological context and to model receptor kinetics and mutation effects. RESULTS: We have performed modeling and simulation studies of the JAK/STAT pathway, including the kinetics of two associated receptors (the erythropoietin receptor and thrombopoietin receptor) with the wild type and a recently reported mutation (JAK2V617F) of the JAK2 protein. CONCLUSION: We found that the different kinetics of those two receptors might be important factors that affect the sensitivity of JAK/STAT signaling to the mutation effect. In addition, our simulation results support clinically observed pathological differences between the two subtypes of MPD with respect to the JAK2V617F mutation. Hong-Hee Won, Jong-Won Kim 0001, Doheon Lee |
BMC Bioinform. | 5 |
| 2008 | Genome-Wide DNA-Binding Specificity of PIL5, a Arabidopsis Basic Helix-Loop-Helix (bHLH) Transcription FactorabstractPIL5 is a member of the bHLH transcription factor super family and plays crucial roles in phytochrome mediated seed germination process in Arabidopsis. While our previous study shows that PIL5 binds to the G-box (CACGTG) motif with high affinity, other attributes must be involved in determining regulatory specificity. In this study we performed a ChIP-chip assay to obtain genome-wide PIL5 binding sites and investigated other attributes which would affect to PIL5 DNA-binding specificity. Our results confirmed that the existence of G-box in promoter region was the most significant binding attribute. Additionally we also found that other attributes such as neighboring motif composition, distance from transcription start site, nucleosome density, DNA methylation and average gene expression were also contributing to PIL5 DNA-binding specificity. The Random Forest classifier using these attributes can classify PIL5 binding sites from non-binding sites with accuracy of 93.05%. Hyojin Kang, Eunkyoo Oh, Giltsu Choi, Doheon Lee |
BIBM | 4 |
| 2008 | Inferring Pathway Activity toward Precise Disease ClassificationabstractThe advent of microarray technology has made it possible to classify disease states based on gene expression profiles of patients. Typically, marker genes are selected by measuring the power of their expression profiles to discriminate among patients of different disease states. However, expression-based classification can be challenging in complex diseases due to factors such as cellular heterogeneity within a tissue sample and genetic heterogeneity across patients. A promising technique for coping with these challenges is to incorporate pathway information into the disease classification procedure in order to classify disease based on the activity of entire signaling pathways or protein complexes rather than on the expression levels of individual genes or proteins. We propose a new classification method based on pathway activities inferred for each patient. For each pathway, an activity level is summarized from the gene expression levels of its condition-responsive genes (CORGs), defined as the subset of genes in the pathway whose combined expression delivers optimal discriminative power for the disease phenotype. We show that classifiers using pathway activity achieve better performance than classifiers based on individual gene expression, for both simple and complex case-control studies including differentiation of perturbed from non-perturbed cells and subtyping of several different kinds of cancer. Moreover, the new method outperforms several previous approaches that use a static (i.e., non-conditional) definition of pathways. Within a pathway, the identified CORGs may facilitate the development of better diagnostic markers and the discovery of core alterations in human disease. Han-Yu Chuang, Jong-Won Kim 0001, Trey Ideker, Doheon Lee |
PLoS Comput. Biol. | 5 |
| 2007 | Stochastic Simulation Model for Patterned Neural Multi-Electrode ArraysabstractA multi-electrode array (MEA) is a micro-fabricated cell culture dish with embedded microelectrodes at the bottom of the dish. Recently MEAs with different cell-adhesive patterns are actively used to analyze behaviors of in vitro neural systems. It is difficult to confirm the underlying anatomical synaptic connections of the in vitro neural networks based on neural recordings from MEAs. Meanwhile, computational modeling and simulation cannot only facilitate various in silico combinatorial stimulus-response analysis but also provide connection-aware analysis capability. Here we propose a simulation approach of encompassing the whole MEA experiments including cell seeding, axonal growth, network morphology and the recordings of the electrodes of MEAs. Cell densities and the geometry of network morphology could be varied systematically and multiple spike trains from the simulated networks were obtained. Dong-Soo Kahng, Yoonkey Nam, Doheon Lee |
BIBE | 3 |
| 2007 | Inferring Gene Regulatory Networks from Microarray Time Series Data Using Transfer EntropyabstractReverse engineering of gene regulatory networks from microarray time series data has been a challenging problem due to the limit of available data. In this paper, a new approach is proposed based on the concept of transfer entropy. Using this information theoretic measure, causal relations between pairs of genes are assessed to draw a causal network. A heuristic rule is then applied to differentiate direct and indirect causality. Simulation on a synthetic network showed that the transfer entropy can identify both linear and nonlinear causality. Application of the method in a biological data identified many causal interactions with biological information supports. Thai Quang Tung, Taewoo Ryu, Kwang Hyung Lee, Doheon Lee |
CBMS | 4 |
| 2007 | Towards clustering of incomplete microarray data without the use of imputationabstractMOTIVATION: Clustering technique is used to find groups of genes that show similar expression patterns under multiple experimental conditions. Nonetheless, the results obtained by cluster analysis are influenced by the existence of missing values that commonly arise in microarray experiments. Because a clustering method requires a complete data matrix as an input, previous studies have estimated the missing values using an imputation method in the preprocessing step of clustering. However, a common limitation of these conventional approaches is that once the estimates of missing values are fixed in the preprocessing step, they are not changed during subsequent processes of clustering; badly estimated missing values obtained in data preprocessing are likely to deteriorate the quality and reliability of clustering results. Thus, a new clustering method is required for improving missing values during iterative clustering process. RESULTS: We present a method for Clustering Incomplete data using Alternating Optimization (CIAO) in which a prior imputation method is not required. To reduce the influence of imputation in preprocessing, we take an alternative optimization approach to find better estimates during iterative clustering process. This method improves the estimates of missing values by exploiting the cluster information such as cluster centroids and all available non-missing values in each iteration. To test the performance of the CIAO, we applied the CIAO and conventional imputation-based clustering methods, e.g. k-means based on KNNimpute, for clustering two yeast incomplete data sets, and compared the clustering result of each method using the Saccharomyces Genome Database annotations. The clustering results of the CIAO method are more significantly relevant to the biological gene annotations than those of other methods, indicating its effectiveness and potential for clustering incomplete gene expression data. AVAILABILITY: The software was developed using Java language, and can be executed on the platforms that JVM (Java Virtual Machine) is running. It is available from the authors upon request. Daewon Kim 0001, Ki Young Lee, Kwang Hyung Lee, Doheon Lee |
Bioinform. | 4 |
| 2007 | BioCAD: an information fusion platform for bio-network inference and analysisabstractBACKGROUND: As systems biology has begun to draw growing attention, bio-network inference and analysis have become more and more important. Though there have been many efforts for bio-network inference, they are still far from practical applications due to too many false inferences and lack of comprehensible interpretation in the biological viewpoints. In order for applying to real problems, they should provide effective inference, reliable validation, rational elucidation, and sufficient extensibility to incorporate various relevant information sources. RESULTS: We have been developing an information fusion software platform called BioCAD. It is utilizing both of local and global optimization for bio-network inference, text mining techniques for network validation and annotation, and Web services-based workflow techniques. In addition, it includes an effective technique to elucidate network edges by integrating various information sources. This paper presents the architecture of BioCAD and essential modules for bio-network inference and analysis. CONCLUSION: BioCAD provides a convenient infrastructure for network inference and network analysis. It automates series of users' processes by providing data preprocessing tools for various formats of data. It also helps inferring more accurate and reliable bio-networks by providing network inference tools which utilize information from distinct sources. And it can be used to analyze and validate the inferred bio-networks using information fusion tools. Doheon Lee, Sangwoo Kim |
BMC Bioinform. | 1 |
| 2007 | Density-Induced Support Vector Data DescriptionabstractThe purpose of data description is to give a compact description of the target data that represents most of its characteristics. In a support vector data description (SVDD), the compact description of target data is given in a hyperspherical model, which is determined by a small portion of data called support vectors. Despite the usefulness of the conventional SVDD, however, it may not identify the optimal solution of target description especially when the support vectors do not have the overall characteristics of the target data. To address the issue in SVDD methodology, we propose a new SVDD by introducing new distance measurements based on the notion of a relative density degree for each data point in order to reflect the distribution of a given data set. Moreover, for a real application, we extend the proposed method for the protein localization prediction problem which is a multiclass and multilabel problem. Experiments with various real data sets show promising results. Ki Young Lee, Daewon Kim 0001, Kwang Hyung Lee, Doheon Lee |
IEEE Trans. Neural Networks | 4 |
| 2005 | Voting Fuzzy k-NN to Predict Protein Subcellular Localization from Normalized Amino Acid Pair Compositions
Thai Quang Tung, Doheon Lee, Daewon Kim 0001, Jong-Tae Lim |
PAKDD | 2 |
| 2005 | Detecting clusters of different geometrical shapes in microarray gene expression dataabstractMOTIVATION: Clustering has been used as a popular technique for finding groups of genes that show similar expression patterns under multiple experimental conditions. Many clustering methods have been proposed for clustering gene-expression data, including the hierarchical clustering, k-means clustering and self-organizing map (SOM). However, the conventional methods are limited to identify different shapes of clusters because they use a fixed distance norm when calculating the distance between genes. The fixed distance norm imposes a fixed geometrical shape on the clusters regardless of the actual data distribution. Thus, different distance norms are required for handling the different shapes of clusters. RESULTS: We present the Gustafson-Kessel (GK) clustering method for microarray gene-expression data. To detect clusters of different shapes in a dataset, we use an adaptive distance norm that is calculated by a fuzzy covariance matrix (F) of each cluster in which the eigenstructure of F is used as an indicator of the shape of the cluster. Moreover, the GK method is less prone to falling into local minima than the k-means and SOM because it makes decisions through the use of membership degrees of a gene to clusters. The algorithmic procedure is accomplished by the alternating optimization technique, which iteratively improves a sequence of sets of clusters until no further improvement is possible. To test the performance of the GK method, we applied the GK method and well-known conventional methods to three recently published yeast datasets, and compared the performance of each method using the Saccharomyces Genome Database annotations. The clustering results of the GK method are more significantly relevant to the biological annotations than those of the other methods, demonstrating its effectiveness and potential for clustering gene-expression data. AVAILABILITY: The software was developed using Java language, and can be executed on the platforms that JVM (Java Virtual Machine) is running. It is available from the authors upon request. SUPPLEMENTARY INFORMATION: Supplementary data are available at http://dragon.kaist.ac.kr/gk. Daewon Kim 0001, Kwang Hyung Lee, Doheon Lee |
Bioinform. | 3 |
| 2005 | Modularized learning of genetic interaction networks from biological annotations and mRNA expression dataabstractMOTIVATION: Inferring the genetic interaction mechanism using Bayesian networks has recently drawn increasing attention due to its well-established theoretical foundation and statistical robustness. However, the relative insufficiency of experiments with respect to the number of genes leads to many false positive inferences. RESULTS: We propose a novel method to infer genetic networks by alleviating the shortage of available mRNA expression data with prior knowledge. We call the proposed method 'modularized network learning' (MONET). Firstly, the proposed method divides a whole gene set to overlapped modules considering biological annotations and expression data together. Secondly, it infers a Bayesian network for each module, and integrates the learned subnetworks to a global network. An algorithm that measures a similarity between genes based on hierarchy, specificity and multiplicity of biological annotations is presented. The proposed method draws a global picture of inter-module relationships as well as a detailed look of intra-module interactions. We applied the proposed method to analyze Saccharomyces cerevisiae stress data, and found several hypotheses to suggest putative functions of unclassified genes. We also compared the proposed method with a whole-set-based approach and two expression-based clustering approaches. Phil Hyoun Lee, Doheon Lee |
Bioinform. | 2 |
| 2005 | Architecture of basic building blocks in protein and domain structural interaction networksabstractMOTIVATION: The structural interaction of proteins and their domains in networks is one of the most basic molecular mechanisms for biological cells. Topological analysis of such networks can provide an understanding of and solutions for predicting properties of proteins and their evolution in terms of domains. A single paradigm for the analysis of interactions at different layers, such as domain and protein layers, is needed. RESULTS: Applying a colored vertex graph model, we integrated two basic interaction layers under a unified model: (1) structural domains and (2) their protein/complex networks. We identified four basic and distinct elements in the model that explains protein interactions at the domain level. We searched for motifs in the networks to detect their topological characteristics using a pruning strategy and a hash table for rapid detection. We obtained the following results: first, compared with a random distribution, a substantial part of the protein interactions could be explained by domain-level structural interaction information. Second, there were distinct kinds of protein interaction patterns classified by specific and distinguishable numbers of domains. The intermolecular domain interaction was the most dominant protein interaction pattern. Third, despite the coverage of the protein interaction information differing among species, the similarity of their networks indicated shared architectures of protein interaction network in living organisms. Remarkably, there were only a few basic architectures in the model (>10 for a 4-node network topology), and we propose that most biological combinations of domains into proteins and complexes can be explained by a small number of key topological motifs. CONTACT: [email protected]. Hyun S. Moon, Jonghwa Bhak, Kwang Hyung Lee, Doheon Lee |
Bioinform. | 4 |
| 2005 | Evaluation of the performance of clustering algorithms in kernel-induced feature space
Daewon Kim 0001, Ki Young Lee, Doheon Lee, Kwang Hyung Lee |
Pattern Recognit. | 3 |
| 2005 | A k-populations algorithm for clustering categorical data
Daewon Kim 0001, Ki Young Lee, Doheon Lee, Kwang Hyung Lee |
Pattern Recognit. | 3 |
| 2005 | Possibilistic support vector machines
Ki Young Lee, Daewon Kim 0001, Kwang Hyung Lee, Doheon Lee |
Pattern Recognit. | 4 |
| 2005 | Improving support vector data description using local density degree
Ki Young Lee, Daewon Kim 0001, Doheon Lee, Kwang Hyung Lee |
Pattern Recognit. | 3 |
| 2005 | A kernel-based subtractive clustering method
Daewon Kim 0001, Ki Young Lee, Doheon Lee, Kwang Hyung Lee |
Pattern Recognit. Lett. | 3 |
| 2004 | Regression trees for regulatory element identificationabstractMOTIVATION: The transcription of a gene is largely determined by short sequence motifs that serve as binding sites for transcription factors. Recent findings suggest direct relationships between the motifs and gene expression levels. In this work, we present a method for identifying regulatory motifs. Our method makes use of tree-based techniques for recovering the relationships between motifs and gene expression levels. RESULTS: We treat regulatory motifs and gene expression levels as predictor variables and responses, respectively, and use a regression tree model to identify the structural relationships between them. The regression tree methodology is extended to handle responses from multiple experiments by modifying the split function. The significance of regulatory elements is determined by analyzing tree structures and using a variable importance measure. When applied to two data sets of the yeast Saccharomyces cerevisiae, the method successfully identifies most of the regulatory motifs that are known to control gene transcription under the given experimental conditions, and suggests several new putative motifs. Analysis of the tree structures also reconfirms several pairs of motifs that are known to regulate gene transcription in combination. AVAILABILITY: http://if.kaist.ac.kr/~phuong/RegTree Tu Minh Phuong, Doheon Lee, Kwang Hyung Lee |
Bioinform. | 2 |
| 2004 | Comparison Of Type-2 Fuzzy Values With Satisfaction FunctionabstractA comparison method for discrete type-2 fuzzy values is proposed in this paper. The proposed method is based on the concept of the satisfaction function that handles the ambiguity in fuzzy comparisons using the possibility distribution of actual values. Some properties of the proposed comparison method are analyzed as well. Kwang Hyung Lee, Doheon Lee |
Int. J. Uncertain. Fuzziness Knowl. Based Syst. | 3 |
| 2004 | A cluster validation index for GK cluster analysis based on relative degree of sharing
Young-Il Kim, Daewon Kim 0001, Doheon Lee, Kwang Hyung Lee |
Inf. Sci. | 3 |
| 2004 | Ranking the sequences of fuzzy values
Kwang Hyung Lee, Doheon Lee |
Inf. Sci. | 3 |
| 2004 | On cluster validity index for estimation of the optimal number of fuzzy clusters
Daewon Kim 0001, Kwang Hyung Lee, Doheon Lee |
Pattern Recognit. | 3 |
| 2004 | A novel initialization scheme for the fuzzy c-means algorithm for color clustering
Daewon Kim 0001, Kwang Hyung Lee, Doheon Lee |
Pattern Recognit. Lett. | 3 |
| 2004 | Fuzzy clustering of categorical data using fuzzy centroids
Daewon Kim 0001, Kwang Hyung Lee, Doheon Lee |
Pattern Recognit. Lett. | 3 |
| 2004 | Fuzzy branching temporal logicabstractIntelligent systems require a systematic way to represent and handle temporal information containing uncertainty. In particular, a logical framework is needed that can represent uncertain temporal information and its relationships with logical formulae. Fuzzy linear temporal logic (FLTL), a generalization of propositional linear temporal logic (PLTL) with fuzzy temporal events and fuzzy temporal states defined on a linear time model, was previously proposed for this purpose. However, many systems are best represented by branching time models in which each state can have more than one possible future path. In this paper, fuzzy branching temporal logic (FBTL) is proposed to address this problem. FBTL adopts and generalizes concurrent tree logic (CTL*), which is a classical branching temporal logic. The temporal model of FBTL is capable of representing fuzzy temporal events and fuzzy temporal states, and the order relation among them is represented as a directed graph. The utility of FBTL is demonstrated using a fuzzy job shop scheduling problem as an example. Seong-ick Moon, Kwang Hyung Lee, Doheon Lee |
IEEE Trans. Syst. Man Cybern. Part B | 3 |
| 2003 | Regulatory Element Discovery Using Tree-structured ModelsabstractComputational discovery of transcriptional regulatory regions in DNA sequences provides an efficient way to broaden our understanding of how cellular processes are controlled. We formulate the regulatory element discovery problem in the regression framework with regulatory regions treated as predictor variables and gene expression levels as responses. We use regression tree models to identify structural relationships between predictors and responses. The regression tree methodology is extended to handle multiple responses from different experiments by modifying the split function. We apply this method to two data sets of the yeast Saccharomyces cerevisiae. The method successfully identifies most of regulatory motifs that are known to control gene transcription under the given experimental conditions. Our method also suggests several putative motifs that present novel regulatory motifs. Tu Minh Phuong, Doheon Lee, Kwang Hyung Lee |
ICDM | 2 |
| 2003 | Learning Rules to Extract Protein Interactions from Biomedical Text
Tu Minh Phuong, Doheon Lee, Hyung Lee-Kwang |
PAKDD | 2 |
| 2003 | A Taxonomy of Dirty Data
Won Y. Kim, Byoung-Ju Choi, Eui Kyeong Hong, Soo-Kyung Kim, Doheon Lee |
Data Min. Knowl. Discov. | 5 |
| 2003 | Fuzzy cluster validation index based on inter-cluster proximity
Daewon Kim 0001, Kwang Hyung Lee, Doheon Lee |
Pattern Recognit. Lett. | 3 |
| 2001 | Scalable Workflow System Model Based on Mobile Agents
Jeong-Joon Yoo, Doheon Lee, Young-Ho Suh, Dong-Ik Lee |
PRIMA | 2 |
| 1997 | Database summarization using fuzzy ISA hierarchiesabstractSummary discovery is one of the major components of knowledge discovery in databases, which provides the user with comprehensive information for grasping the essence from a large amount of information in a database. We propose an interactive top down summary discovery process which utilizes fuzzy ISA hierarchies as domain knowledge. We define a generalized tuple as a representational form of a database summary including fuzzy concepts. By virtue of fuzzy ISA hierarchies where fuzzy ISA relationships common in actual domains are naturally expressed, the discovery process comes up with more accurate database summaries. We also present an informativeness measure for distinguishing generalized tuples that delivers much information to users, based on C. Shannon's (1948) information theory. Doheon Lee, Myoung-Ho Kim |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 1997 | Database summarization using fuzzy ISA hierarchiesabstractSummary discovery is one of the major components of knowledge discovery in databases, which provides the user with comprehensive information for grasping the essence from a large amount of information in a database. In this paper, we propose an interactive top-down summary discovery process which utilizes fuzzy ISA hierarchies as domain knowledge. We define a generalized tuple as a representational form of a database summary including fuzzy concepts. By virtue of fuzzy ISA hierarchies where fuzzy ISA relationships common in actual domains are naturally expressed, the discovery process comes up with more accurate database summaries. We also present an informativeness measure for distinguishing generalized tuples that delivers much information to users, based on Shannon's information theory. Doheon Lee, Myoung-Ho Kim |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 1994 | Discovering Database Summaries through Refinements of Fuzzy HypothesesabstractRecently, many applications such as scientific databases and decision supporting systems that require comprehensive analysis of a very large amount of data, have been evolved. Summary discovery techniques, which extract compact representations grasping the meanings of large databases, can play a major role in those applications. We present an effective and robust method to discover simple linguistic summaries. We first propose a hypothesis refinement algorithm that is a key technique for our summary discovery method. Using the algorithm, a formal procedure for summary discovery is presented together with an illustrative example. Our discovery method can handle both rigid concepts and fuzzy concepts that occur frequently in practice. Discovered summaries can also be regarded as high-level interattribute dependencies.> Doheon Lee, Myoung-Ho Kim |
ICDE | 1 |
| 1993 | A Hypothesis Refinement Method for Summary Discovery in DatabasesabstractAs database systems are playing major roles in more and more applications, the amount of information in databases is rapidly growing.In order to comprehend those large volumes of information, computerized summary discovery methods are required.In this paper, we propose a hypothesis refinement method for constructing and evaluating fuzzy hypotheses.Breed on them we propose an effective and robust algorithm to discover simple linguistic summaries.In addition, we present ideas for exploiting discovered summaries to various applications such as querying database knowledge, handling query failures and semantic query optimization. Doheon Lee, Myoung-Ho Kim |
CIKM | 1 |
| 1993 | A Fuzzification of the Relational Data Model
Doheon Lee, Myoung-Ho Kim, Hyung Lee-Kwang, Yoon-Joon Lee |
DASFAA | 1 |
| 1993 | Accommodating subjective vagueness through a fuzzy extension to the relational data model
Doheon Lee, Myoung-Ho Kim |
Inf. Syst. | 1 |
| 1993 | Extending semantics of relational operators for vague queries
Doheon Lee, Myoung-Ho Kim |
Microprocess. Microprogramming | 1 |