EDBT 2026 Demo / reviewers in the wild / expert
Young-Rae Cho
dblp:12/2467
· DBLP profile ↗
40ranked-venue papers
15as first author
9since 2021 · last 2026
0000-0002-4645-2542ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 34 · 13 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorSystems, architecture and hardware · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ORBIT: Oncogenic Representation Learning via Bi-Prototype Contrastive Learning in Hyperbolic Space for cancer driver gene identificationabstractAccurate identification of cancer driver genes is crucial for precision oncology but remains challenging due to the complexity of integrating heterogeneous data and modeling dynamic biological systems. To address these limitations, we propose ORBIT (Oncogenic Representation Learning via Bi-Prototype Contrastive Learning in Hyperbolic Space). Our framework synergistically fuses multi-omics profiles with functional network data using a context-adaptive graph reweighting mechanism to capture cancer-specific dynamics. The model employs a bi-prototype contrastive learning strategy within hyperbolic space, which aligns gene representations around distinct driver and non-driver semantic anchors while preserving the intrinsic hierarchy of biological networks. Comprehensive evaluations demonstrate that ORBIT achieves highly competitive stability in pan-cancer analysis while consistently outperforming state-of-the-art methods in cancer-specific predictions. Furthermore, functional enrichment analysis confirms that the model effectively segregates core cancer pathways, and drug sensitivity profiling validates the clinical relevance of the identified drivers. By integrating hyperbolic geometry with context-adaptive learning, ORBIT offers a robust and interpretable paradigm for precision medicine. The source codes and datasets are publicly accessible at https://github.com/spcho-dev/ORBIT. Sang-Pil Cho, Young-Rae Cho |
Artif. Intell. Medicine | 2 |
| 2026 | GRAFT: a graph-aware fusion transformer for cancer driver gene predictionabstractIdentifying cancer driver genes is essential for precision oncology, but existing computational methods are often limited by their reliance on single biological networks and their inability to capture long-range molecular dependencies. To address these challenges, we propose GRAFT, a Graph-Aware Fusion Transformer. This framework learns modality-specific features from protein-protein interactions, pathway co-occurrence, and gene semantic similarity using a multi-view graph encoder. These representations are further enriched with two auxiliary feature types: structural encodings derived from network topology and functional embeddings guided by curated gene sets. The integrated features are then processed by a transformer backbone, where a novel edge-attention bias makes the model explicitly sensitive to the underlying graph topologies, enabling the effective modeling of both local and global dependencies. Extensive evaluations demonstrate that GRAFT achieves competitive performance with leading state-of-the-art methods in pan-cancer analysis, while consistently delivering superior predictive accuracy across numerous specific cancer types. More importantly, a functional enrichment analysis of the novel candidate driver genes predicted by our model confirms their strong associations with key cancer-related processes, demonstrating the model's ability to make biologically plausible discoveries. By delivering a powerful and interpretable framework, our model not only advances the identification of cancer driver genes but also establishes a robust paradigm for multimodal data integration in systems biology. The source codes and datasets are publicly accessible at https://github.com/spcho-dev/GRAFT. Sang-Pil Cho, Young-Rae Cho |
Briefings Bioinform. | 2 |
| 2025 | MoDiff: A Morphology-Emphasized Diffusion Model for Ambiguous Medical Image Segmentation
Jung Su Ahn, Ki Hoon Kwak, Jung Woo Seo, Young-Rae Cho |
MICCAI (3) | 4 |
| 2025 | FACT: Feature Aggregation and Convolution with Transformers for predicting drug classification codeabstractMOTIVATION: Drug repositioning, identifying new therapeutic applications for existing drugs, can significantly reduce the time and cost involved in drug development. Recent studies have explored the use of Anatomical Therapeutic Chemical (ATC) codes in drug repositioning, offering a systematic framework to predict ATC codes for a drug. The ATC classification system organizes drugs according to their chemical properties, pharmacological actions, and therapeutic effects. However, its complex hierarchical structure and the limited scalability at higher levels present significant challenges for achieving accurate ATC code prediction. RESULTS: We propose a novel approach to predict ATC codes of drugs, named Feature Aggregation and Convolution with Transformer models (FACT). This method computes three types of drug similarities, incorporating ATC code similarity with hierarchical weights and masked drug-ATC code associations. These features are then aggregated for each target drug-ATC code pair and processed through a convolution-transformer encoder to generate three embeddings. The embeddings are finally used to estimate the probability of an association between the target pair. The experimental results demonstrate that the proposed method achieves an area under the receiver operating characteristic curve (AUROC) of 0.9805 and an area under the precision-recall curve of 0.9770 at level 4 of the ATC codes, outperforming the previous methods by 15.05% and 18.42%, respectively. This study highlights the effectiveness of integrating diverse drug features and the potential of transformer-based models in ATC code prediction. AVAILABILITY AND IMPLEMENTATION: Source code of FACT is freely available at https://github.com/knhc1234/FACT. Gwang-Hyeon Yun, Jong-Hoon Park, Young-Rae Cho |
Bioinform. | 3 |
| 2024 | Mitigating Oversmoothing in Hypergraph Neural Networks for Enhanced Cancer-Driver Gene PredictionabstractPredicting cancer-driver genes is a significant task in cancer research, and hypergraph neural networks are effective tools for capturing complex relationships among genes. However, these methods typically experience performance degradation due to the oversmoothing problem, which indicates the situation where node distinctiveness decreases as network depth increases. In this study, we evaluated several approaches to mitigate the oversmoothing problem in cancer-driver gene prediction using hypergraph neural networks. Specifically, the deep learning model incorporating oversmoothing mitigation techniques achieved a 3.22% improvement in terms of the area under the receiver operating characteristic curve, reaching the score 0.931. Moreover, in terms of the area under the precision-recall curve, the model showed a score of 0.881, reflecting a 14.42% improvement. The experimental results demonstrate that, even as the network depth increases, the selected approaches effectively handle oversmoothing, leading to improved performance to predict cancer-driver genes. Sang-Pil Cho, Young-Rae Cho |
BIBM | 2 |
| 2024 | Comparative Evaluation and Data Analysis for Drug Toxicity PredictionabstractThroughout the entire drug development process, toxicity prediction is a significant process in assessing the possible toxicity of drugs. Recently, there has been active research on estimating drug toxicity using deep learning techniques. However, challenges such as a lack of labeled data and the insufficient reliable benchmark datasets remain major obstacles. To address these challenges, this study analyzes and compares seven toxicity benchmark datasets from the Therapeutics Data Commons and three external literature datasets. Experimental results demonstrated that deep learning methods showed superior performance on the benchmark datasets. In particular, on the ClinTox dataset, deep learning models outperformed conventional machine learning methods by 5.5%. The results confirm that employing self-supervised learning effectively mitigates the data limitation problem. Additionally, it was validated that using graph representation learning and natural language processing effectively handles chemical structures as drug features. Jae-Woo Chu, Young-Rae Cho |
BIBM | 2 |
| 2024 | Computational Disease-Gene Association Prioritization Using Graph Neural Networks and Attention MechanismsabstractPredicting disease-gene associations is crucial in human health for understanding the genetic basis of diseases, identifying potential therapeutic targets, and advancing personalized medicine. In this paper, we present a novel network-based method that integrates graph neural networks (GNNs), attention mechanisms, and random walks to accurately predict diseasegene associations. GNNs encode structural features to nodes in a heterogeneous graph, attention mechanisms compute the connectivity between the nodes, and random walks capture the changes between states within the graph. Experimental results demonstrated the superiority of the proposed model, achieving the area under the receiver operating characteristics curve of 0.982 and recall 0.907. Jong-Hoon Park, Young-Rae Cho |
BIBM | 2 |
| 2024 | Comparative Analysis of Attention-based Models for Drug-Target Interaction PredictionabstractIdentifying drug-target interactions is a crucial step in drug discovery and drug repurposing. To efficiently predict these interactions, various computational tools have been developed. Recently, machine learning, especially deep learning, has gained significant attention for predicting drug-target interactions. Deep learning models extract features of both drugs and targets, capturing hidden relationships between them through deep neural networks. In this study, we evaluate the performance of several attention-based models in predicting interactions using four publicly available benchmark datasets. Their performance is assessed with three metrics: the area under the receiver operating characteristic curve, the area under the precision-recall curve, and the F1-score. Experimental results demonstrate that attention-based methods generally exhibit high predictive performance and possess distinct advantages under various conditions. Gwang-Hyeon Yun, Young-Rae Cho |
BIBM | 2 |
| 2022 | Penalty based robust learning with noisy labels
Kyeongbo Kong, Junggi Lee, Youngchul Kwak, Young-Rae Cho, Seong-Eun Kim, Woo-Jin Song |
Neurocomputing | 4 |
| 2020 | Survey of network-based approaches of drug-target interaction predictionabstractDrug repositioning is a promising strategy for drug design and discovery. The key step of drug repositioning is computational prediction of drug-target interactions (DTIs). Recently, machine learning methods have provided efficient tools to predict DTIs. These supervised methods require a sufficient amount of training data including negative samples to increase prediction accuracy. The network-based approaches also provide an effective solution for DTI prediction. These models use a bipartite graph where each node represents either a drug or a target and each edge is an interaction. The network-based approaches then adopt graph-theoretic techniques based on the topological features of the DTI network. Using the benchmark DTI data set integrated from multiple resources, their performance can be assessed by cross-validation. The AUC and AUPR results show the network-based approaches have competitive accuracy. Improving data quality, for example, precise measurement of structural similarity between chemical compounds, is the most significant issue for current network-based DTI prediction. Lee Soo Jung, Young-Rae Cho |
BIBM | 2 |
| 2019 | Survey of biological network alignment: cross-species analysis of conserved systemsabstractComparative analysis of PPI networks between species is a significant step for predicting evolutionary conserved components or sub-structures in a system level. Network alignment is a computational technique to map similar components between networks. Aligning genome-wide PPI networks maps functionally similar nodes and edges and identifies conserved interactions or conserved modules. Existing network alignment algorithms can be categorized into two groups: global and local network alignment. Global network alignment algorithms search for the best alignment of entire networks, whereas local network alignment algorithms produce the aligned pairs of small sub-networks with the highest scores. In this survey, we summarize prominent network alignment algorithms in both categories. We also present the methods to evaluate network alignment algorithms. When we compare the evaluation results of selected network alignment algorithms, PrimAlign outperformed the other global network alignment algorithms. Among local network alignment algorithms, LeP-rimAlign was superior to the competitors. We conclude with a future direction to expand the applicability of network alignment techniques. Sawal Maskey, Young-Rae Cho |
BIBM | 2 |
| 2018 | PrimAlign: PageRank-inspired Markovian alignment for large biological networksabstractMotivation: Cross-species analysis of large-scale protein-protein interaction (PPI) networks has played a significant role in understanding the principles deriving evolution of cellular organizations and functions. Recently, network alignment algorithms have been proposed to predict conserved interactions and functions of proteins. These approaches are based on the notion that orthologous proteins across species are sequentially similar and that topology of PPIs between orthologs is often conserved. However, high accuracy and scalability of network alignment are still a challenge. Results: We propose a novel pairwise global network alignment algorithm, called PrimAlign, which is modeled as a Markov chain and iteratively transited until convergence. The proposed algorithm also incorporates the principles of PageRank. This approach is evaluated on tasks with human, yeast and fruit fly PPI networks. The experimental results demonstrate that PrimAlign outperforms several prevalent methods with statistically significant differences in multiple evaluation measures. PrimAlign, which is multi-platform, achieves superior performance in runtime with its linear asymptotic time complexity. Further evaluation is done with synthetic networks and results suggest that popular topological measures do not reflect real precision of alignments. Availability and implementation: The source code is available at http://web.ecs.baylor.edu/faculty/cho/PrimAlign. Supplementary information: Supplementary data are available at Bioinformatics online. Karel Kalecky, Young-Rae Cho |
Bioinform. | 2 |
| 2017 | Mining cross-ontology weighted association rules between GO and HPOabstractAn ontology is a framework for describing domain-specific knowledge in a structured format. It is comprised of a set of terms as nodes and a set of relationships between terms as directed edges to form a directed acyclic graph. Gene Ontology (GO) and Human Phenotype Ontology (HPO) are widely referred biological and biomedical ontology databases. They also provide extensive annotations of human genes. Recent studies have applied association rule mining techniques to these ontology and annotation data in order to predict cellular functions of genes and determine gene-to-disease relations. We present a new approach to extract pairwise association rules between specific terms. Our approach selects significant, specific rules by weighted measures of support, confidence and coverage which incorporate weighting terms by integration of the ontology structure and their information contents. In our experiment, the cross-ontology association rules generated from GO and HPO by our approach were compared to those by two previous methods. The results show that our approach discovers association rules between more specific terms than the previous methods. Joseph Huang, Collin Rapp, Young-Rae Cho |
BIBM | 3 |
| 2016 | Filtering association rules in Gene Ontology based on term specificityabstractGene Ontology (GO) is one of the most widely used ontology databases, and provides a framework to elucidate biological roles of genes or gene products by semantic analysis. GO contains terms in a structured format within three domains: biological processes, molecular functions and cellular components. GO also provides extensive annotation data across most model species. Recent studies in analysis of GO and annotation data have confronted two major challenges. The first is the increasing complexity of ontology structures. The second is the inconsistency of annotation data. In this study, we explore association rule mining to curate GO data and to achieve consistent annotations. We propose a novel specificity measure for GO terms, called VICD. This measure adopts the integrative concept of induced sub-ontologies and vectors of information content distance. By pairwise cross-ontology association rule mining, we discover specific association rules using the selected GO terms with high specificity. When we compare the VICD measure to the information content, the experimental results show that VICD scores have more positive correlations with the specificity values from other commonly referred measures than information contents. We also demonstrate that our approach generates more consistent association rules across species by using VICD scores than using information contents. Yong Shui, Young-Rae Cho |
BIBM | 2 |
| 2015 | An integrative measure of graph- and vector-based semantic similarity using information content distanceabstractGene Ontology (GO) and its annotation data have been widely used for genomic and proteomic analysis. In the past few years, various semantic similarity measures using GO have been proposed to quantify functional similarity between two proteins and assess validity of protein-protein interactions (PPIs). They are categorized as pairwise and groupwise approaches according to the strategies of deriving protein-to-protein functional similarity. We propose a novel semantic similarity measure, called simVICD, which is a graph-and vector-based groupwise approach. This method computes the magnitude of a common induced subgraph as semantic similarity between two sets of terms annotating two proteins, respectively. The magnitude of the common induced subgraph is represented as the Euclidean norm of a vector having information content distance of all possible directed shortest paths in the induced subgraph. Our experimental results show that the proposed groupwise approach, simVICD, and a previous integrative pairwise approach, simICND, outperform the other existing semantic similarity methods in predicting protein complexes and identifying essential proteins. Qiaoli Hu, Young-Rae Cho |
BIBM | 2 |
| 2015 | Semantic mapping to align PPI networks and predict conserved protein complexesabstractPPI networks are significant resources to determine molecular organizations in a cell. The availability of genome-wide PPI networks on diverse model species have provided a new paradigm to identify evolutionarily conserved substructures. Computational methods for cross-species comparison of PPI networks have recently been applied to prediction of conserved protein complexes. These methods use network alignment techniques by mapping homologous proteins. We propose a novel network alignment approach by semantic mapping between proteins from different species. We apply this approach to predict conserved human protein complexes by aligning yeast PPI networks representing well-studied protein complexes with a human PPI network in a genomic scale. In our experiments, we used a recently proposed integrative semantic similarity measure, simICND, for semantic mapping. The experimental results show that the proposed network alignment approach has higher accuracy on predicting human protein complexes than other clustering-based methods. The experimental results also show that the proposed approach has higher efficiency than previous network alignment algorithms. This study provides a valuable framework to discover conserved systems from a functional standpoint. Lizhu Ma, Young-Rae Cho |
BIBM | 2 |
| 2015 | P-Finder: Reconstruction of Signaling Networks from Protein-Protein Interactions and GO AnnotationsabstractBecause most complex genetic diseases are caused by defects of cell signaling, illuminating a signaling cascade is essential for understanding their mechanisms. We present three novel computational algorithms to reconstruct signaling networks between a starting protein and an ending protein using genome-wide protein-protein interaction (PPI) networks and gene ontology (GO) annotation data. A signaling network is represented as a directed acyclic graph in a merged form of multiple linear pathways. An advanced semantic similarity metric is applied for weighting PPIs as the preprocessing of all three methods. The first algorithm repeatedly extends the list of nodes based on path frequency towards an ending protein. The second algorithm repeatedly appends edges based on the occurrence of network motifs which indicate the link patterns more frequently appearing in a PPI network than in a random graph. The last algorithm uses the information propagation technique which iteratively updates edge orientations based on the path strength and merges the selected directed edges. Our experimental results demonstrate that the proposed algorithms achieve higher accuracy than previous methods when they are tested on well-studied pathways of S. cerevisiae. Furthermore, we introduce an interactive web application tool, called P-Finder, to visualize reconstructed signaling networks. Young-Rae Cho, Greg Speegle |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2013 | Signaling pathway prediction by path frequency in protein-protein interaction networksabstractA signaling pathway, which is represented as a chain of interacting proteins for a biological process, can be predicted from protein-protein interaction (PPI) networks. However, pathway prediction is computationally challenging because of (1) inefficiency in searching all possible paths from the large-scale PPI networks and (2) unreliability of current PPI data generated by automated high-throughput methods. In this paper, we propose a novel approach to efficiently predict signaling pathways from PPI networks when a starting protein (source) and an ending protein (target) are given. Our approach is a combination of topological analysis of the networks and ontological analysis of interacting proteins. Starting from the source, this method repeatedly extends the list of proteins to form a pathway based on the improved support model (iSup). This model integrates (1) the frequency of the paths towards the target and (2) the semantic similarity between each adjacent pair in a pathway. The path frequency is computed by a heuristic data-mining technique to determine the most frequent paths towards the target in a PPI network. The semantic similarity is measured by the distance of the information contents of Gene Ontology (GO) terms annotating interacting proteins. To further improve computational efficiency, we propose two additional strategies: filtering the PPI networks and precomputing approximate path frequency. The experiment with the yeast PPI data demonstrates that our approach predicted MAPK signaling pathways with higher accuracy and efficiency than other existing methods. Yilan Bai, Greg Speegle, Young-Rae Cho |
BIBM | 3 |
| 2013 | Prediction of signaling networks by information propagation on protein-protein interaction networks integrated with GO annotationsabstractThe experimental study of signal transduction over a decade has made a substantial contribution to understanding functional mechanisms in a cell. A signaling pathway represents a linear path of a signaling cascade involving a series of proteins. As an advanced model, multiple linear pathways with extensive cross-talk between receptors can be merged into a larger-scale signaling network. We present an efficient computational approach to predict signaling networks by integration of genome-wide protein-protein interaction (PPI) data and ontological annotation data. We adopt an advanced semantic similarity metric for weighting PPIs, and an information propagation algorithm that runs on a weighted PPI network. This algorithm iteratively selects potential directed edges for signaling cascade using user-specified path strength parameters. Our approach also includes a preprocessing step to filter the large-scale PPI network by distance condition using the maximum path length parameter. Our experimental results show that the proposed approach runs extremely faster than existing computational methods and has competitive accuracy in the test of predicting well-studied pathways of S. cerevisiae and C. elegans. High efficiency of this approach would facilitate development of a web-based application tool to discover potential signaling networks. Young-Rae Cho, Slávka Jaromerská |
BIBM | 1 |
| 2012 | M-Finder: Functional association mining from protein interaction networks weighted by semantic similarityabstractProtein-protein interactions (PPIs) play a key role in understanding functional behavior of genes. Discovering association patterns from PPI networks is crucial for functional characterization on a system level. We present a novel approach to discover the functional association pattern of a query gene from the genome-wide PPI networks. This approach consists of two major components. First, we transform the PPI network to a weighted graph representation by measuring semantic similarity. Three enhanced semantic similarity methods are proposed to estimate functional closeness of each interacting pair. Second, we apply a dynamic propagation algorithm to detect the functional association pattern of a gene, represented as a sub-network. The size of the sub-networks is flexibly determined by user-specific parameters. In this paper, we also introduce an interactive web application, called M-Finder, to visualize the functional association pattern of a gene entered by a user. The semantic similarity measures and the dynamic propagation algorithm are embedded in this tool to run on up-to-date PPI networks of model species. M-Finder allows users to carry out further systematic analysis for functional characterization on the genomic scale. Young-Rae Cho, Tak Chien Chiam, Yanxin Lu |
BIBM | 1 |
| 2012 | Assessing reliability of protein-protein interactions by gene ontology integrationabstractRecent advances in genome-wide identification of protein-protein interactions (PPIs) have produced an abundance of interaction data which give an insight into functional associations among proteins. However, it is known that the PPI datasets determined by high-throughput experiments or inferred by computational methods include an extremely large number of false positives. Using Gene Ontology (GO) and its annotations, we assess reliability of the PPIs by considering the semantic similarity of interacting proteins. Protein pairs with high semantic similarity are considered highly likely to share common functions, and therefore, are more likely to interact. We analyze the performance of existing semantic similarity measures in terms of functional consistency and propose a combined method that achieves improved performance over existing methods. The semantic similarity measures are applied to identify false positive PPIs. The classification results show that the combined hybrid method has higher accuracy than the other existing measures. Furthermore, the combined hybrid classifier predicts that 59.6% of the S. cerevisiae PPIs from the BioGRID database are false positives. George D. Montañez, Young-Rae Cho |
CIBCB | 2 |
| 2011 | Assessment of Cluster Overlaps to Improve Accuracy of Module Detection from PPI NetworksabstractRecent computational techniques have facilitated analyzing genome-wide protein-protein interactions. Various graph-clustering algorithms have been applied to the protein interaction networks for identifying protein complexes and functional modules. Since each protein performs multiple functions, a clustering algorithm should be able to produce overlapping clusters. In this paper, we use the seed-refinement algorithm to generate a set of preliminary overlapping clusters. Next, for further refining the preliminary clusters, we carry out a systematic analysis of their overlaps by novel metrics: overlap coverage and overlapping consistency. We propose the cluster- merging algorithm to yield final clusters by parameterizing the metrics. In the test with the yeast protein-protein interaction network, we demonstrate the proposed approach improves accuracy on detecting protein complexes and functional modules by optimizing the parameter values. Nicholas Soltau, Young-Rae Cho |
BIBM | 2 |
| 2011 | Entropy-Based Graph Clustering: Application to Biological and Social NetworksabstractComplex systems have been widely studied to characterize their structural behaviors from a topological perspective. High modularity is one of the recurrent features of real-world complex systems. Various graph clustering algorithms have been applied to identifying communities in social networks or modules in biological networks. However, their applicability to real-world systems has been limited because of the massive scale and complex connectivity of the networks. In this study, we exploit a novel information-theoretic model for graph clustering. The entropy-based clustering approach finds locally optimal clusters by growing a random seed in a manner that minimizes graph entropy. We design and analyze modifications that further improve its performance. Assigning priority in seed-selection and seed-growth is well applicable to the scale-free networks characterized by the hub-oriented structure. Computing seed-growth in parallel streams also decomposes an extremely large network efficiently. The experimental results with real biological and social networks show that the entropy-based approach has better performance than competing methods in terms of accuracy and efficiency. Edward Casey Kenley, Young-Rae Cho |
ICDM | 2 |
| 2010 | Functional Flow Simulation Based Analysis of Protein Interaction NetworkabstractProtein-protein interactions (PPIs) play fundamental roles in nearly all biological processes and differ based on the composition, affinity and lifetime of the association. A vast amount of PPI data for various organisms is available from MIPS, DIP and other sources. The identification of functional modules in PPI network is of great interest because they often reveal unknown functional ties between proteins and hence predict functions for unknown proteins. In this paper, we propose using functional flow simulation and the topology of the network for the functional module detection and function prediction problem. Our approach is based on the functional influence model that quantifies the influence of a biological component on another. We introduce a flow simulation algorithm to generate a functional profile for each component. In addition, a new clustering method FMD (Functional Module Detection) is designed to associate with functional profiles to detect functional modules. We evaluate the proposed technique on three different yeast networks with MIPS functional categories and compare it with several other existing techniques in terms of precision and recall. Our experiments show that our approach achieves better accuracy than other existing methods. Lei Shi 0021, Young-Rae Cho, Aidong Zhang 0001 |
BIBE | 2 |
| 2010 | Decomposing protein interactome networks by graph entropyabstractRecent high-throughput experimental methods have generated protein-protein interaction data in the genome scale, called interactome. Various graph clustering algorithms have been applied to the protein interactome networks for identifying protein complexes and predicting functional modules. Although the previous algorithms are scalable and robust, their accuracy is still limited because of complex connectivity of the networks. In this study, we propose a novel information-theoretic definition, Graph Entropy, as a measure of structural complexity of a graph. Loss of graph entropy represents an increase in modularity of the graph. Based on this concept, we present a graph clustering algorithm. Starting from a random seed vertex and its neighbors as a seed cluster, the algorithm iteratively adds or removes vertices on the border of the cluster to minimize graph entropy. We make an additional improvement on the algorithm for generating overlapping clusters. In the experiments with the yeast protein interactome network, we show the graph entropy-based approach has higher accuracy in predicting functional modules than other competing methods. Hao Lian, Chengsen Song, Young-Rae Cho |
BIBM | 3 |
| 2010 | Identification of functional hubs and modules by converting interactome networks into hierarchical ordering of proteinsabstractBACKGROUND: Protein-protein interactions play a key role in biological processes of proteins within a cell. Recent high-throughput techniques have generated protein-protein interaction data in a genome-scale. A wide range of computational approaches have been applied to interactome network analysis for uncovering functional organizations and pathways. However, they have been challenged because of complex connectivity. It has been investigated that protein interaction networks are typically characterized by intrinsic topological features: high modularity and hub-oriented structure. Elucidating the structural roles of modules and hubs is a critical step in complex interactome network analysis. RESULTS: We propose a novel approach to convert the complex structure of an interactome network into hierarchical ordering of proteins. This algorithm measures functional similarity between proteins based on the path strength model, and reveals a hub-oriented tree structure hidden in the complex network. We score hub confidence and identify functional modules in the tree structure of proteins, retrieved by our algorithm. Our experimental results in the yeast protein interactome network demonstrate that the selected hubs are essential proteins for performing functions. In network topology, they have a role in bridging different functional modules. Furthermore, our approach has high accuracy in identifying functional modules hierarchically distributed. CONCLUSIONS: Decomposing, converting, and synthesizing complex interaction networks are fundamental tasks for modeling their structural behaviors. In this study, we systematically analyzed complex interactome network structures for retrieving functional information. Unlike previous hierarchical clustering methods, this approach dynamically explores the hierarchical structure of proteins in a global view. It is well-applicable to the interactome networks in high-level organisms because of its efficiency and scalability. Young-Rae Cho, Aidong Zhang 0001 |
BMC Bioinform. | 1 |
| 2010 | Predicting protein function by frequent functional association pattern mining in protein interaction networksabstractPredicting protein function from protein interaction networks has been challenging because of the complexity of functional relationships among proteins. Most previous function prediction methods depend on the neighborhood of or the connected paths to known proteins. However, their accuracy has been limited due to the functional inconsistency of interacting proteins. In this paper, we propose a novel approach for function prediction by identifying frequent patterns of functional associations in a protein interaction network. A set of functions that a protein performs is assigned into the corresponding node as a label. A functional association pattern is then represented as a labeled subgraph. Our frequent labeled subgraph mining algorithm efficiently searches the functional association patterns that occur frequently in the network. It iteratively increases the size of frequent patterns by one node at a time by selective joining, and simplifies the network by a priori pruning. Using the yeast protein interaction network, our algorithm found more than 1400 frequent functional association patterns. The function prediction is performed by matching the subgraph, including the unknown protein, with the frequent patterns analogous to it. By leave-one-out cross validation, we show that our approach has better performance than previous link-based methods in terms of prediction accuracy. The frequent functional association patterns generated in this study might become the foundations of advanced analysis for functional behaviors of proteins in a system level. Young-Rae Cho, Aidong Zhang 0001 |
IEEE Trans. Inf. Technol. Biomed. | 1 |
| 2009 | Restructuring Protein Interaction Networks to Reveal Structural Hubs and Functional OrganizationsabstractProtein interaction networks are significant resources for functional knowledge discovery. However, efficient analysis of the networks has been challenging because of complex connectivity. Protein interaction networks have been characterized by intrinsic features, such as modularity and existence of hubs. The concepts of modules and hubs, extending from specific (local) to general (global), suggest hierarchical structures hidden in the complex networks. Retrieving a protein interaction network into the hierarchical structure is thus a crucial process for better understanding of functional organizations. We present a novel approach for restructuring a protein interaction network to reveal hierarchically organized functional modules and hubs. Our algorithm measures functional similarity between proteins based on the path strength model, and dynamically convert a protein interaction network into a hub-oriented tree structure using the definition of centrality. We identify structural hubs and potential functional modules from the tree structure generated by our algorithm. The experimental results demonstrate that the proteins selected as structural hubs are essential for performing functions. In network topology, they have a role in bridging different modules. Furthermore, our approach has higher accuracy in identifying functional modules than other hierarchical clustering methods. Young-Rae Cho, Aidong Zhang 0001 |
BIBM | 1 |
| 2009 | flowNet: Flow-Based Approach for Efficient Analysis of Complex Biological NetworksabstractBiological networks having complex connectivity have been widely studied recently. By characterizing their inherent and structural behaviors in a topological perspective, these studies have attempted to discover hidden knowledge in the systems. However, even though various algorithms with graph-theoretical modeling have provided fundamentals in the network analysis, the availability of practical approaches to efficiently handle the complexity has been limited. In this paper, we present a novel flow-based approach, called flowNet, to efficiently analyze large-sized, complex networks. Our approach is based on the functional influence model that quantifies the influence of a biological component on another. We introduce a dynamic flow simulation algorithm to generate a flow pattern which is a unique characteristic for each component. The set of patterns can be used in identifying functional modules (i.e., clustering). The proposed flow simulation algorithm runs very efficiently in sparse networks. Since our approach uses a weighted network as an input, we also discuss supervised and unsupervised weighting schemes for unweighted biological networks. As experimental results in real applications to the yeast protein interaction network, we demonstrate that our approach outperforms previous graph clustering methods with respect to accuracy. Young-Rae Cho, Lei Shi 0021, Aidong Zhang 0001 |
ICDM | 1 |
| 2008 | A fast two-pass HDL simulation with on-demand dumpabstractSimulation-based functional verification is characterized by two inherently conflicting targets: the signal visibility and simulation performance. Achieving a proper trade-off between these two targets is of paramount importance. Even though HDL simulators are the most widely used verification platform at the RTL and gate level, their major drawback is the low performance in verifying complex SOCs, especially when the high visibility over the design under verification is required. This paper presents a new, fast simulation method as an effective way to achieve both high simulation speed and full signal visibility. It is based on an original two-pass simulation approach. During the 1stpass, with the simulation running at full speed, a set of design states is saved periodically at predetermined checkpoints. During the 2ndpass, another simulation is performed, using any of saved checkpoints and providing 100% signal visibility for debugging. Our method differs from the traditional simulation snapshot approach in the amount and the way the design state is saved. Experimental results show significant speed-up compared to existing traditional simulation methods while maintaining 100% visibility. Kyuho Shim, Young-Rae Cho, Namdo Kim, Hyuncheol Baik, Kyungkuk Kim, Dusung Kim, Jaebum Kim, Byeongun Min, Kyumyung Choi, Maciej J. Ciesielski, Seiyang Yang |
ASP-DAC | 2 |
| 2008 | Discovering Frequent Patterns of Functional Associations in Protein Interaction Networks for Function PredictionabstractPredicting function from protein interaction networks has been challenging because of the intricate functional relationships among proteins. Most of the previous function prediction methods depend on the neighborhood of or the connected paths to known proteins, and remain low in accuracy. In this paper, we propose a novel approach for function prediction by detecting frequent patterns of functional associations in a protein interaction network. A set of functions that a protein performs is assigned into the corresponding node as a label. A functional association pattern is then represented as a labeled subgraph. Our FASPAM (frequent functional association pattern mining) algorithm efficiently finds the patterns that occur frequently in the network. It iteratively increases the size of frequent patterns by one node at a time by selective joining, and simplifies the network by a priori pruning. Using the yeast protein interaction network extracted from DIP, the FASPAM algorithm found more than 1,400 frequent patterns. By leave-one-out cross validation, our algorithm predicted functions from the frequent patterns with the accuracy of 86%, which is higher than the results from most previous methods. Young-Rae Cho, Aidong Zhang 0001 |
BIBM | 1 |
| 2008 | A probabilistic framework to predict protein function from interaction data integrated with semantic knowledgeabstractBACKGROUND: The functional characterization of newly discovered proteins has been a challenge in the post-genomic era. Protein-protein interactions provide insights into the functional analysis because the function of unknown proteins can be postulated on the basis of their interaction evidence with known proteins. The protein-protein interaction data sets have been enriched by high-throughput experimental methods. However, the functional analysis using the interaction data has a limitation in accuracy because of the presence of the false positive data experimentally generated and the interactions that are a lack of functional linkage. RESULTS: Protein-protein interaction data can be integrated with the functional knowledge existing in the Gene Ontology (GO) database. We apply similarity measures to assess the functional similarity between interacting proteins. We present a probabilistic framework for predicting functions of unknown proteins based on the functional similarity. We use the leave-one-out cross validation to compare the performance. The experimental results demonstrate that our algorithm performs better than other competing methods in terms of prediction accuracy. In particular, it handles the high false positive rates of current interaction data well. CONCLUSION: The experimentally determined protein-protein interactions are erroneous to uncover the functional associations among proteins. The performance of function prediction for uncharacterized proteins can be enhanced by the integration of multiple data sources available. Young-Rae Cho, Lei Shi 0021, Murali Ramanathan, Aidong Zhang 0001 |
BMC Bioinform. | 1 |
| 2008 | Functional module detection by functional flow pattern mining in protein interaction networks
Young-Rae Cho, Lei Shi 0021, Aidong Zhang 0001 |
BMC Bioinform. | 1 |
| 2008 | CASCADE: a novel quasi all paths-based network analysis algorithm for clustering biological interactionsabstractBACKGROUND: Quantitative characterization of the topological characteristics of protein-protein interaction (PPI) networks can enable the elucidation of biological functional modules. Here, we present a novel clustering methodology for PPI networks wherein the biological and topological influence of each protein on other proteins is modeled using the probability distribution that the series of interactions necessary to link a pair of distant proteins in the network occur within a time constant (the occurrence probability). RESULTS: CASCADE selects representative nodes for each cluster and iteratively refines clusters based on a combination of the occurrence probability and graph topology between every protein pair. The CASCADE approach is compared to nine competing approaches. The clusters obtained by each technique are compared for enrichment of biological function. CASCADE generates larger clusters and the clusters identified have p-values for biological function that are approximately 1000-fold better than the other methods on the yeast PPI network dataset. An important strength of CASCADE is that the percentage of proteins that are discarded to create clusters is much lower than the other approaches which have an average discard rate of 45% on the yeast protein-protein interaction network. CONCLUSION: CASCADE is effective at detecting biologically relevant clusters of interactions. Woochang Hwang, Young-Rae Cho, Aidong Zhang 0001, Murali Ramanathan |
BMC Bioinform. | 2 |
| 2007 | Optimizing Flow-based Modularization by Iterative Centroid Search in Protein Interaction NetworksabstractThe systematic analysis of protein-protein interactions is a fundamental step for understanding of cellular organization, processes and functions. Functional modules can be identified from the protein interaction networks. However, current unreliable interaction data and complex connectivity of interaction networks have made it challenging. We propose a novel metric, called semantic interactivity, to measure the reliability of protein-protein interactions using gene ontology (GO) annotation data. The protein interaction networks can be converted into a weighted graph representation by assigning the reliability to each edge as a weight. We present an iterative centroid search (ICES) algorithm for optimizing the flow-based modularization method and identifying functional modules in a weighted interaction network. It iteratively performs two procedures: centroid search and flow simulation. Our experimental results show that the accuracy of modules is enhanced during the iteration. Young-Rae Cho, Woochang Hwang, Aidong Zhang 0001 |
BIBE | 1 |
| 2007 | SIGN: reliable protein interaction identification by integrating the Similarity In GO and the similarity in protein interaction NetworksabstractHigh-throughput techniques for protein-protein interaction detection in a genomic scale have provided us a genomic wide view of molecular interactions of many living organisms. A few approaches were proposed to scrutinize protein-protein interactions of living organisms. By the way, the binary nature of the current protein interaction data sets imposes challenges for effective analysis. Furthermore, their performance was suffered by the intrinsic defect, i.e., high noise level, of high-throughput data. This unpleasantly high false positive rate could lead many devoted researches to erroneous biological conclusions. We propose a novel reliability measurement for protein interactions integrating the similarity in gene ontology and the topological similarity in protein interaction networks. Our metric has been proven to be an effective reliability metric for identifying biologically more reliable interactions through the analyses performed from various view points, e.g., functional homogeneity, subcellular localizational homogeneity, and gene expression correlation, etc. Woochang Hwang, Taehyong Kim, Young-Rae Cho, Aidong Zhang 0001, Murali Ramanathan |
BIBE | 3 |
| 2007 | Modularization of Protein Interaction Networks by Incorporating Gene Ontology AnnotationsabstractRecent computational analyses of protein interaction networks have attempted to understand cellular organizations, processes and functions. However, they have encountered difficulties due to unreliable interaction data and the complexity of the networks. In this paper, we propose the integration of protein interaction networks with gene ontology annotations for assessing the reliability of current protein-protein interaction data. The interaction reliability can be used for building weighted protein interaction networks. We apply an information flow-based modularization algorithm to the weighted protein interaction networks. Our experimental results show that the interaction reliability between two proteins is positively correlated to the likelihood of functional and locational associations. We finally demonstrate that our approach identifies accurate modules in the protein interaction networks with high statistical confidence with respect to biological function and cellular localization. Moreover, this algorithm outperforms our previous method (Cho et al., 2006) integrating with genetic co-expressional profiles Young-Rae Cho, Woochang Hwang, Aidong Zhang 0001 |
CIBCB | 1 |
| 2007 | Feature Extraction from Microarray Expression Data by Integration of Semantic KnowledgeabstractMicroarray techniques give biologists first peek into the molecular states of living tissues. Previous studies have proven that it is feasible to build sample classifiers using the gene expressional profiles. To build an effective sample classifier, dimension reduction process is necessary since classic pattern recognition algorithms do not work well in high dimensional space. In this paper, we present a novel feature extraction algorithm based on the concept of virtual genes by integrating microarray expression data sets with domain knowledge embedded in gene ontology (GO) annotations. We define semantic similarity to measure the functional associations between two genes using the annotation on each GO term. We then identify the groups of genes, called virtual genes, that potentially interact with each other for a biological function. The correlation in gene expression levels of virtual genes can be used to build a sample classifier. For a colon cancer data set, the integration of microarray expression data with GO annotations significantly improves the accuracy of sample classification by more than 10%. Young-Rae Cho, Xian Xu 0002, Woochang Hwang, Aidong Zhang 0001 |
ICMLA | 1 |
| 2007 | Semantic integration to identify overlapping functional modules in protein interaction networksabstractBACKGROUND: The systematic analysis of protein-protein interactions can enable a better understanding of cellular organization, processes and functions. Functional modules can be identified from the protein interaction networks derived from experimental data sets. However, these analyses are challenging because of the presence of unreliable interactions and the complex connectivity of the network. The integration of protein-protein interactions with the data from other sources can be leveraged for improving the effectiveness of functional module detection algorithms. RESULTS: We have developed novel metrics, called semantic similarity and semantic interactivity, which use Gene Ontology (GO) annotations to measure the reliability of protein-protein interactions. The protein interaction networks can be converted into a weighted graph representation by assigning the reliability values to each interaction as a weight. We presented a flow-based modularization algorithm to efficiently identify overlapping modules in the weighted interaction networks. The experimental results show that the semantic similarity and semantic interactivity of interacting pairs were positively correlated with functional co-occurrence. The effectiveness of the algorithm for identifying modules was evaluated using functional categories from the MIPS database. We demonstrated that our algorithm had higher accuracy compared to other competing approaches. CONCLUSION: The integration of protein interaction networks with GO annotation data and the capability of detecting overlapping modules substantially improve the accuracy of module identification. Young-Rae Cho, Woochang Hwang, Murali Ramanathan, Aidong Zhang 0001 |
BMC Bioinform. | 1 |
| 2006 | Efficient Modularization of Weighted Protein Interaction Networks using k-Hop Graph ReductionabstractRecent computational analyses of protein interaction networks have attempted to understand cellular organizations, processes and functions. Several topology-based clustering methods have been applied to the protein interaction networks for detecting functional modules. However, most of the previous algorithms do not perform well on small-world, scale-free networks. In this paper, we present an efficient approach to identify hierarchical modules in the protein interaction networks. Our algorithm selects a small number of informative proteins from a large network, and transforms the intricate small-world, scale-free network into a simple graph with high modularity. Our results show that this approach remarkably enhances the efficiency. We also demonstrate that it outperforms other previous methods in terms of accuracy Young-Rae Cho, Woochang Hwang, Aidong Zhang 0001 |
BIBE | 1 |