VLDB 2026 Research / reviewers in the wild / expert
Jinn-Moon Yang
dblp:68/838
· DBLP profile ↗
43ranked-venue papers
14as first author
5since 2021 · last 2022
0000-0002-3205-4391ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 24 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 18 · 13 first-authorSystems, architecture and hardware · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Exploring kinase family inhibitors and their moiety preferences using deep SHapley additive exPlanationsabstractBACKGROUND: While it has been known that human protein kinases mediate most signal transductions in cells and their dysfunction can result in inflammatory diseases and cancers, it remains a challenge to find effective kinase inhibitor as drugs for these diseases. One major challenge is the compensatory upregulation of related kinases following some critical kinase inhibition. To circumvent the compensatory effect, it is desirable to have inhibitors that inhibit all the kinases belonging to the same family, instead of targeting only a few kinases. However, finding inhibitors that target a whole kinase family is laborious and time consuming in wet lab. RESULTS: In this paper, we present a computational approach taking advantage of interpretable deep learning models to address this challenge. Specifically, we firstly collected 9,037 inhibitor bioassay results (with 3991 active and 5046 inactive pairs) for eight kinase families (including EGFR, Jak, GSK, CLK, PIM, PKD, Akt and PKG) from the ChEMBL25 Database and the Metz Kinase Profiling Data. We generated 238 binary moiety features for each inhibitor, and used the features as input to train eight deep neural networks (DNN) models to predict whether an inhibitor is active for each kinase family. We then employed the SHapley Additive exPlanations (SHAP) to analyze the importance of each moiety feature in each classification model, identifying moieties that are in the common kinase hinge sites across the eight kinase families, as well as moieties that are specific to some kinase families. We finally validated these identified moieties using experimental crystal structures to reveal their functional importance in kinase inhibition. CONCLUSION: With the SHAP methodology, we identified two common moieties for eight kinase families, 9 EGFR-specific moieties, and 6 Akt-specific moieties, that bear functional importance in kinase inhibition. Our result suggests that SHAP has the potential to help finding effective pan-kinase family inhibitors. You-Wei Fan, Wan-Hsin Liu, Yun-Ti Chen, Yen-Chao Hsu, Nikhil Pathak, Yu-Wei Huang, Jinn-Moon Yang |
BMC Bioinform. | 7 |
| 2022 | Discovery of moiety preference by Shapley value in protein kinase family using random forest modelsabstractBACKGROUND: Human protein kinases play important roles in cancers, are highly co-regulated by kinase families rather than a single kinase, and complementarily regulate signaling pathways. Even though there are > 100,000 protein kinase inhibitors, only 67 kinase drugs are currently approved by the Food and Drug Administration (FDA). RESULTS: In this study, we used "merged moiety-based interpretable features (MMIFs)," which merged four moiety-based compound features, including Checkmol fingerprint, PubChem fingerprint, rings in drugs, and in-house moieties as the input features for building random forest (RF) models. By using > 200,000 bioactivity test data, we classified inhibitors as kinase family inhibitors or non-inhibitors in the machine learning. The results showed that our RF models achieved good accuracy (> 0.8) for the 10 kinase families. In addition, we found kinase common and specific moieties across families using the Shapley Additive exPlanations (SHAP) approach. We also verified our results using protein kinase complex structures containing important interactions of the hinges, DFGs, or P-loops in the ATP pocket of active sites. CONCLUSIONS: In summary, we not only constructed highly accurate prediction models for predicting inhibitors of kinase families but also discovered common and specific inhibitor moieties between different kinase families, providing new opportunities for designing protein kinase inhibitors. Yu-Wei Huang, Yen-Chao Hsu, Yi-Hsuan Chuang, Yun-Ti Chen, Xiang-Yu Lin, You-Wei Fan, Nikhil Pathak, Jinn-Moon Yang |
BMC Bioinform. | 8 |
| 2022 | Densest subgraph-based methods for protein-protein interaction hot spot predictionabstractBACKGROUND: Hot spots play an important role in protein binding analysis. The residue interaction network is a key point in hot spot prediction, and several graph theory-based methods have been proposed to detect hot spots. Although the existing methods can yield some interesting residues by network analysis, low recall has limited their abilities in finding more potential hot spots. RESULT: In this study, we develop three graph theory-based methods to predict hot spots from only a single residue interaction network. We detect the important residues by finding subgraphs with high densities, i.e., high average degrees. Generally, a high degree implies a high binding possibility between protein chains, and thus a subgraph with high density usually relates to binding sites that have a high rate of hot spots. By evaluating the results on 67 complexes from the SKEMPI database, our methods clearly outperform existing graph theory-based methods on recall and F-score. In particular, our main method, Min-SDS, has an average recall of over 0.665 and an f2-score of over 0.364, while the recall and f2-score of the existing methods are less than 0.400 and 0.224, respectively. CONCLUSION: The Min-SDS method performs best among all tested methods on the hot spot prediction problem, and all three of our methods provide useful approaches for analyzing bionetworks. In addition, the densest subgraph-based methods predict hot spots with only one residue interaction network, which is constructed from spatial atomic coordinate data to mitigate the shortage of data from wet-lab experiments. Ruiming Li, Jung-Yu Lee 0001, Jinn-Moon Yang, Tatsuya Akutsu |
BMC Bioinform. | 3 |
| 2022 | Identification of pan-kinase-family inhibitors using graph convolutional networks to reveal family-sensitive pre-moietiesabstractBACKGROUND: Human protein kinases, the key players in phosphoryl signal transduction, have been actively investigated as drug targets for complex diseases such as cancer, immune disorders, and Alzheimer's disease, with more than 60 successful drugs developed in the past 30 years. However, many of these single-kinase inhibitors show low efficacy and drug resistance has become an issue. Owing to the occurrence of highly conserved catalytic sites and shared signaling pathways within a kinase family, multi-target kinase inhibitors have attracted attention. RESULTS: To design and identify such pan-kinase family inhibitors (PKFIs), we proposed PKFI sets for eight families using 200,000 experimental bioactivity data points and applied a graph convolutional network (GCN) to build classification models. Furthermore, we identified and extracted family-sensitive (only present in a family) pre-moieties (parts of complete moieties) by utilizing a visualized explanation (i.e., where the model focuses on each input) method for deep learning, gradient-weighted class activation mapping (Grad-CAM). CONCLUSIONS: This study is the first to propose the PKFI sets, and our results point out and validate the power of GCN models in understanding the pre-moieties of PKFIs within and across different kinase families. Moreover, we highlight the discoverability of family-sensitive pre-moieties in PKFI identification and drug design. Xiang-Yu Lin, Yu-Wei Huang, You-Wei Fan, Yun-Ti Chen, Nikhil Pathak, Yen-Chao Hsu, Jinn-Moon Yang |
BMC Bioinform. | 7 |
| 2021 | CoMI: consensus mutual information for tissue-specific gene signaturesabstractBACKGROUND: The gene signatures have been considered as a promising early diagnosis and prognostic analysis to identify disease subtypes and to determine subsequent treatments. Tissue-specific gene signatures of a specific disease are an emergency requirement for precision medicine to improve the accuracy and reduce the side effects. Currently, many approaches have been proposed for identifying gene signatures for diagnosis and prognostic. However, they often lack of tissue-specific gene signatures. RESULTS: Here, we propose a new method, consensus mutual information (CoMI) for analyzing omics data and discovering gene signatures. CoMI can identify differentially expressed genes in multiple cancer omics data for reflecting both cancer-related and tissue-specific signatures, such as Cell growth and death in multiple cancers, Xenobiotics biodegradation and metabolism in LIHC, and Nervous system in GBM. Our method identified 50-gene signatures effectively distinguishing the GBM patients into high- and low-risk groups (log-rank p = 0.006) for diagnosis and prognosis. CONCLUSIONS: Our results demonstrate that CoMI can identify significant and consistent gene signatures with tissue-specific properties and can predict clinical outcomes for interested diseases. We believe that CoMI is useful for analyzing omics data and discovering gene signatures of diseases. Sing-Han Huang, Yu-Shu Lo, Yong-Chun Luo, Yi-Hsuan Chuang, Jung-Yu Lee 0001, Jinn-Moon Yang |
BMC Bioinform. | 6 |
| 2020 | Configurational Differences and Binding Mechanisms of Interleukin-1 Receptor-Associated Kinase 1abstractInterleukin-1 receptor-associated kinase 1 (IRAK1) is crucial for downstream regulation of the toll-like receptor signaling pathway and is involved in innate immune system, inflammatory diseases, and cancers. There is an urgent need for developing inhibitors that target IRAK1. To address this issue, we have developed a method to explore the structural dynamics of IRAK1 by utilizing the abundant structures of IRAK4. Our results show that IRAK1 should have four configuration types, including Type-CL (left of C-lobe), TypeCR (right of C-lobe), Type-A (ATP site), and Type-N (N-lobe), according to the ligand occupancy consistency among the four groups of 45 IRAK4 ligand complexes. Each type demonstrates a distinct binding environment between IRAK1 and the bound ligands for guiding the specific inhibitor design. We evaluated our prediction models for each type and discovered a new TypeN inhibitor, ponatinib, which is 100 times more potent for IRAK1 ( 93 nM) than for IRAK4 (>10 PM), based on our kinase inhibition assay. To the best of our knowledge, we are the first team to propose a novel method to study the configurational flexibility and Type-N inhibitors of IRAK1. We believe that our approach provides a useful strategy for designing selective inhibitors for a specific kinase family. Yun-Ti Chen, Cheng-Hsuan Wu, Yi-Cyun Chen, Yen-Chao Hsu, Yu-Wei Huang, Jinn-Moon Yang |
BIBE | 6 |
| 2020 | FooDisNET: a database of food-compound-protein-disease associationsabstractNatural compounds and nutrients in plants are beneficial to human health and reduce the risk of diseases. These compounds can directly or indirectly act on specific proteins to regulate biochemical pathways and affect disease progression or prevent the occurrence of chronic diseases. Several databases provide information about the ingredients of various plants, Chinese medicine, and the plant-target associations. However, there is still a lack of a database linking food, compounds, proteins, and diseases. Here, we propose a FooDisNET database to provide food-compound-protein-disorder connections. FooDisNET contains 6,329 foods, 53,920 compounds, and 22,865 target proteins, of which 4,092 target proteins are associated with 18,689 disorders. Based on our scoring strategy, we constructed a user-friendly website that enables the user to query the associations from four perspectives, namely foods, compounds, proteins, and disorders. The score reflects the prevalence of foods, the rarity of compounds, the specificity of proteins, and the universality of disorders to describe food-compound-protein-disorder interaction networks. We believe that the FooDisNET database not only assists general uses to investigate food ingredients that are beneficial to their health but also helps the analysis of the role of natural compounds in the drug discovery and development process. Chu-Yun Lin, Jung-Yu Lee 0001, Sing-Han Huang, Yen-Chao Hsu, Nung-Yu Hsu, Jinn-Moon Yang |
BIBE | 6 |
| 2018 | Identification of the PCa28 Gene Signature as a Predictor in Prostate CancerabstractProstate cancer (PCa) is the second-leading cause of cancer death among men in the worldwide. Most PCa is slowly growing and usually early symptomless. About 70% of PCa patients were diagnosed at later stage and metastasis has been observed. Additionally, the cure rate of PCa closely relies on the early diagnosis with biomarkers. Prostatic Specific Antigen (PSA) is currently the only clinical biomarker for PCa diagnosis. However, the PSA test has inherent limitations and has about 75% of false-positive results. The identification of a set of genes (as biomarkers) for diagnosis and prognosis is an urgent clinical issue for PCa. Here, we integrated genome-wide analysis and protein-protein interaction network to identify potential genes for early diagnostic biomarkers of PCa. First, we collected gene expression datasets of 145 PCa samples, consisting of both tumor and corresponding normal tissues, from two different sources in Gene Expression Omnibus (GEO). We found 158 and 268 significantly highly and lowly expressed genes, respectively, in tumor samples. Moreover, we proposed cluster score (CS) and predicting score (PS) to select 28 prostate cancer-related genes (called PCa28). The results indicate that PCa28 can discriminate between the normal/tumor tissues and are specific for prostate cancer. Finally, we examined 8 genes in PCa28 on four PCa cell lines by real time quantitative polymerase chain reaction (RT-qPCR). Experimental results show that up-regulated genes have higher expression level in tumor cells in comparison to normal cells, and down-regulated genes have lower expression level in tumor cells. We believe that our method is useful and PCa28 are potential biomarkers that provide the clues to develop targeting therapy for PCa. Jung-Yu Lee 0001, Si-Yu Lin, Yi-Hsuan Chuang, Sing-Han Huang, Yu-Yao Tseng, Chun-Yu Lin 0003, Hung-Jung Wang, Jinn-Moon Yang |
BIBE | 8 |
| 2018 | Deep Learning with Evolutionary and Genomic Profiles for Identifying Cancer SubtypesabstractCancer subtype identification is an unmet need in precision diagnosis. Recently, evolutionary conservation has been indicated containing understandable signatures for functional significance in cancers. However, the importance of evolutionary conservation in distinguishing cancer subtypes remains unclear. Here, we identified the evolutionarily conserved genes (i.e., core gene) and observed that they are mainly involved in the pathways relevant to cell growth and metabolisms. By using these core genes, we integrated their evolutionary and genomic profiles with deep learning to develop a feature-based strategy (FES) and an image-based strategy (IMS). In comparison with FES using the random set and the strategy using the PAM50 classifier, core gene set-based FES has higher accuracy for identifying breast cancer subtypes. Moreover, the IMS with data augmentation yields better performance than the other strategies. Comprehensive analysis of eight TCGA cancer data demonstrates that our evolutionary conservation-based models provide a valid and helpful approach to identify cancer subtypes and the core gene set offers distinguishable clues of cancer subtypes. Chun-Yu Lin 0003, Peiying Ruan, Ruiming Li, Jinn-Moon Yang, Simon See, Tatsuya Akutsu |
BIBE | 4 |
| 2017 | Pharmacophore anchor models of flaviviral NS3 proteases lead to drug repurposing for DENV infectionabstractViruses of the flaviviridae family are responsible for some of the major infectious viral diseases around the world and there is an urgent need for drug development for these diseases. Most of the virtual screening methods in flaviviral drug discovery suffer from a low hit rate, strain-specific efficacy differences, and susceptibility to resistance. It is because they often fail to capture the key pharmacological features of the target active site critical for protein function inhibition. So in our current work, for the flaviviral NS3 protease, we summarized the pharmacophore features at the protease active site as anchors (subsite-moiety interactions). For each of the four flaviviral NS3 proteases (i.e., HCV, DENV, WNV, and JEV), the anchors were obtained and summarized into ‘Pharmacophore anchor (PA) models’. To capture the conserved pharmacophore anchors across these proteases, were merged the four PA models. We identified five consensus core anchors (CEH1, CH3, CH7, CV1, CV3) in all PA models, represented as the “Core pharmacophore anchor (CPA) model” and also identified specific anchors unique to the PA models. Our PA/CPA models complied with 89 known NS3 protease inhibitors. Furthermore, we proposed an integrated anchor-based screening method using the anchors from our models for discovering inhibitors. This method was applied on the DENV NS3 protease to screen FDA drugs discovering boceprevir, telaprevir and asunaprevir as promising anti-DENV candidates. Experimental testing against DV2-NGC virus by in-vitro plaque assays showed that asunaprevir and telaprevir inhibited viral replication with EC 50 values of 10.4 μM & 24.5 μM respectively. The structure-anchor-activity relationships (SAAR) showed that our PA/CPA model anchors explained the observed in-vitro activities of the candidates. Also, we observed that the CEH1 anchor engagement was critical for the activities of telaprevir and asunaprevir while the extent of inhibitor anchor occupation guided their efficacies. These results validate our NS3 protease PA/CPA models, anchors and the integrated anchor-based screening method to be useful in inhibitor discovery and lead optimization, thus accelerating flaviviral drug discovery. Nikhil Pathak, Mei-Ling Lai, Wen-Yu Chen, Betty-Wu Hsieh, Guann-Yi Yu, Jinn-Moon Yang |
BMC Bioinform. | 6 |
| 2016 | Finding Influential Genes Using Gene Expression Data and Boolean Models of Metabolic NetworksabstractSelection of influential genes using gene expression data from normal and disease samples is an important topic in bioinformatics. In this paper, we propose a novel computational method for the problem, which combines gene expression patterns from normal and disease samples with a mathematical model of metabolic networks. This method seeks a set of k genes knockout of which drives the state of the metabolic network towards that in the disease samples. We adopt a Boolean model of metabolic networks and formulate the problem as a maximization problem under an integer linear programming framework. We applied the proposed method to selection of influential genes using gene expression data from normal samples and disease (head and neck cancer) samples. The result suggests that the proposed method can select more biologically relevant genes than an existing P-value based ranking method can. Takeyuki Tamura, Tatsuya Akutsu, Chun-Yu Lin 0003, Jinn-Moon Yang |
BIBE | 4 |
| 2013 | Inferring homologous protein-protein interactions through pair position specific scoring matrixabstractBACKGROUND: The protein-protein interaction (PPI) is one of the most important features to understand biological processes. For a PPI, the physical domain-domain interaction (DDI) plays the key role for biology functions. In the post-genomic era, to rapidly identify homologous PPIs for analyzing the contact residue pairs of their interfaces within DDIs on a genomic scale is essential to determine PPI networks and the PPI interface evolution across multiple species. RESULTS: In this study, we proposed "pair Position Specific Scoring Matrix (pairPSSM)" to identify homologous PPIs. The pairPSSM can successfully distinguish the true protein complexes from unreasonable protein pairs with about 90% accuracy. For the test set including 1,122 representative heterodimers and 2,708,746 non-interacting protein pairs, the mean average precision and mean false positive rate of pairPSSM were 0.42 and 0.31, respectively. Moreover, we applied pairPSSM to identify ~450,000 homologous PPIs with their interacting domains and residues in seven common organisms (e.g. Homo sapiens, Mus musculus, Saccharomyces cerevisiae and Escherichia coli). CONCLUSIONS: Our pairPSSM is able to provide statistical significance of residue pairs using evolutionary profiles and a scoring system for inferring homologous PPIs. According to our best knowledge, the pairPSSM is the first method for searching homologous PPIs across multiple species using pair position specific scoring matrix and a 3D dimer as the template to map interacting domain pairs of these PPIs. We believe that pairPSSM is able to provide valuable insights for the PPI evolution and networks across multiple species. Chun-Yu Lin 0003, Yung-Chiang Chen, Yu-Shu Lo, Jinn-Moon Yang |
BMC Bioinform. | 4 |
| 2013 | Pathway-based Screening Strategy for Multitarget Inhibitors of Diverse Proteins in Metabolic PathwaysabstractMany virtual screening methods have been developed for identifying single-target inhibitors based on the strategy of "one-disease, one-target, one-drug". The hit rates of these methods are often low because they cannot capture the features that play key roles in the biological functions of the target protein. Furthermore, single-target inhibitors are often susceptible to drug resistance and are ineffective for complex diseases such as cancers. Therefore, a new strategy is required for enriching the hit rate and identifying multitarget inhibitors. To address these issues, we propose the pathway-based screening strategy (called PathSiMMap) to derive binding mechanisms for increasing the hit rate and discovering multitarget inhibitors using site-moiety maps. This strategy simultaneously screens multiple target proteins in the same pathway; these proteins bind intermediates with common substructures. These proteins possess similar conserved binding environments (pathway anchors) when the product of one protein is the substrate of the next protein in the pathway despite their low sequence identity and structure similarity. We successfully discovered two multitarget inhibitors with IC50 of <10 µM for shikimate dehydrogenase and shikimate kinase in the shikimate pathway of Helicobacter pylori. Furthermore, we found two selective inhibitors (IC50 of <10 µM) for shikimate dehydrogenase using the specific anchors derived by our method. Our experimental results reveal that this strategy can enhance the hit rates and the pathway anchors are highly conserved and important for biological functions. We believe that our strategy provides a great value for elucidating protein binding mechanisms and discovering multitarget inhibitors. Kai-Cheng Hsu, Wen-Chi Cheng, Yen-Fu Chen, Wen-Ching Wang, Jinn-Moon Yang |
PLoS Comput. Biol. | 5 |
| 2011 | iGEMDOCK: a graphical environment of enhancing GEMDOCK using pharmacological interactions and post-screening analysisabstractBACKGROUND: Pharmacological interactions are useful for understanding ligand binding mechanisms of a therapeutic target. These interactions are often inferred from a set of active compounds that were acquired experimentally. Moreover, most docking programs loosely coupled the stages (binding-site and ligand preparations, virtual screening, and post-screening analysis) of structure-based virtual screening (VS). An integrated VS environment, which provides the friendly interface to seamlessly combine these VS stages and to identify the pharmacological interactions directly from screening compounds, is valuable for drug discovery. RESULTS: We developed an easy-to-use graphic environment, iGEMDOCK, integrating VS stages (from preparations to post-screening analysis). For post-screening analysis, iGEMDOCK provides biological insights by deriving the pharmacological interactions from screening compounds without relying on the experimental data of active compounds. The pharmacological interactions represent conserved interacting residues, which often form binding pockets with specific physico-chemical properties, to play the essential functions of a target protein. Our experimental results show that the pharmacological interactions derived by iGEMDOCK are often hot spots involving in the biological functions. In addition, iGEMDOCK provides the visualizations of the protein-compound interaction profiles and the hierarchical clustering dendrogram of the compounds for post-screening analysis. CONCLUSIONS: We have developed iGEMDOCK to facilitate steps from preparations of target proteins and ligand libraries toward post-screening analysis. iGEMDOCK is especially useful for post-screening analysis and inferring pharmacological interactions from screening compounds. We believe that iGEMDOCK is useful for understanding the ligand binding mechanisms and discovering lead compounds. iGEMDOCK is available at http://gemdock.life.nctu.edu.tw/dock/igemdock.php. Kai-Cheng Hsu, Yen-Fu Chen, Shen-Rong Lin, Jinn-Moon Yang |
BMC Bioinform. | 4 |
| 2011 | Changed epitopes drive the antigenic drift for influenza A (H3N2) virusesabstractBACKGROUND: In circulating influenza viruses, gradually accumulated mutations on the glycoprotein hemagglutinin (HA), which interacts with infectivity-neutralizing antibodies, lead to the escape of immune system (called antigenic drift). The antibody recognition is highly correlated to the conformation change on the antigenic sites (epitopes), which locate on HA surface. To quantify a changed epitope for escaping from neutralizing antibodies is the basis for the antigenic drift and vaccine development. RESULTS: We have developed an epitope-based method to identify the antigenic drift of influenza A utilizing the conformation changes on epitopes. A changed epitope, an antigenic site on HA with an accumulated conformation change to escape from neutralizing antibody, can be considered as a "key feature" for representing the antigenic drift. According to hemagglutination inhibition (HI) assays and HA/antibody complex structures, we statistically measured the conformation change of an epitope by considering the number of critical position mutations with high genetic diversity and antigenic scores. Experimental results show that two critical position mutations can induce the conformation change of an epitope to escape from the antibody recognition. Among five epitopes of HA, epitopes A and B, which are near to the receptor binding site, play a key role for neutralizing antibodies. In addition, two changed epitopes often drive the antigenic drift and can explain the selections of 24 WHO vaccine strains. CONCLUSIONS: Our method is able to quantify the changed epitopes on HA for predicting the antigenic variants and providing biological insights to the vaccine updates. We believe that our method is robust and useful for studying influenza virus evolution and vaccine development. Jhang-Wei Huang, Jinn-Moon Yang |
BMC Bioinform. | 2 |
| 2010 | Template-based scoring functions for visualizing biological insights of H-2Kb-peptide-TCR complexesabstractClass-I major histocompatibility complex (MHC), peptide, and T-cell receptor (TCR) play an essential role of adaptive immune responses. Many prediction servers are available for identification of peptides that bind to MHC class I molecules. These servers are often lack of detailed interacting residues and binding models for analyzing MHC-peptide-TCR interaction mechanisms. This study numerously enhanced the template-based scoring function derived from protein-protein interactions for identifying MHC-peptide-TCR binding models. The scoring function considers both the template similarity and interacting force to ensure the statistically significant interface similarity between the peptide candidates and structure templates. The result shows that our scoring function is comparative to the public websites for identifying MHC binding peptides. Our model, considering both the MHC-peptide and peptide-TCR interfaces, is able to provide visualization and the biological insights of MHC-peptide-TCR binding models. We believe that our model is useful for the development of peptide-based vaccines. I-Hsin Liu, Yu-Shu Lo, Jinn-Moon Yang |
BIBM | 3 |
| 2009 | ATRIPPI: An Atom-residue Preference Scoring Function for Protein-protein InteractionsabstractWe present an ATRIPPI model for analyzing protein-protein interactions. This model is a 167-atom-type and residue-specific interaction preferences with distance bins derived from 641 co-crystallized protein-protein interfaces. The ATRIPPI model is able to yield physical meanings of hydrogen bonding, disulfide bonding, electrostatic interactions, van der Waals and aromatic-aromatic interactions. We applied this model to identify the native states and near-native complex structures on 17 bound and 17 unbound complexes from thousands of decoy structures. On average, 77.5% structures (155 structures) of top rank 200 structures are closed to the native structure. These results suggest that the ATRIPPI model is able to keep the advantages of both atom-atom and residue-residue interactions and is a potential knowledge-based scoring function for protein-protein docking methods. We believe that our model is robust and provides biological meanings to support protein-protein interactions. Kang-Ping Liu, Lu-Shian Chang, Jinn-Moon Yang |
BIBE | 3 |
| 2009 | GemAffinity: A Scoring Function for Predicting Binding Affinity and Virtual ScreeningabstractPrediction of protein-ligand binding affinities is an important issue in molecular recognition and virtual screening. We have developed a scoring function, namely GemAffinity, to predict binding affinities by analyzing 88 descriptors derived from 891 protein-ligand structures selected from the protein data bank (PDB). Based on these 88 descriptors, we derived GemAffinity using a stepwise regression method to identify five descriptors, including van der Waals contact; metal-ligand interactions; water effects; ligand deformation penalties; and highly conserved residues interacting to a bound ligand with hydrogen bonds. GemAffinity was evaluated on an independent set, and the correlation between predicted and experimental values is 0.572. GemAffinity is the best among 13 methods on this set. Our GemAffinity was then applied to virtual screening for thymidine kinase (TK), human carbonic anhydrase II (HCAII), estrogen receptor of antagonists (ER) and agonists (ERA). Experimental results indicate that GemAffinity is able to reduce the disadvantages (i.e. preferring highly polar or high molecular weight compounds) of energy-based scoring functions. In addition, GemAffinity easily combined with other scoring functions to enrich screening accuracies. We believe that GemAffinity is useful to predict binding affinity and virtual screening. Kai-Cheng Hsu, Yen-Fu Chen, Jinn-Moon Yang |
BIBM | 3 |
| 2009 | (PS)2-v2: template-based protein structure prediction serverabstractBACKGROUND: Template selection and target-template alignment are critical steps for template-based modeling (TBM) methods. To identify the template for the twilight zone of 15~25% sequence similarity between targets and templates is still difficulty for template-based protein structure prediction. This study presents the (PS)2-v2 server, based on our original server with numerous enhancements and modifications, to improve reliability and applicability. RESULTS: To detect homologous proteins with remote similarity, the (PS)2-v2 server utilizes the S2A2 matrix, which is a 60 x 60 substitution matrix using the secondary structure propensities of 20 amino acids, and the position-specific sequence profile (PSSM) generated by PSI-BLAST. In addition, our server uses multiple templates and multiple models to build and assess models. Our method was evaluated on the Lindahl benchmark for fold recognition and ProSup benchmark for sequence alignment. Evaluation results indicated that our method outperforms sequence-profile approaches, and had comparable performance to that of structure-based methods on these benchmarks. Finally, we tested our method using the 154 TBM targets of the CASP8 (Critical Assessment of Techniques for Protein Structure Prediction) dataset. Experimental results show that (PS)2-v2 is ranked 6th among 72 severs and is faster than the top-rank five serves, which utilize ab initio methods. CONCLUSION: Experimental results demonstrate that (PS)2-v2 with the S2A2 matrix is useful for template selections and target-template alignments by blending the amino acid and structural propensities. The multiple-template and multiple-model strategies are able to significantly improve the accuracies for target-template alignments in the twilight zone. We believe that this server is useful in structure prediction and modeling, especially in detecting homologous templates with sequence similarity in the twilight zone. Chih-Chieh Chen, Jenn-Kang Hwang, Jinn-Moon Yang |
BMC Bioinform. | 3 |
| 2009 | Co-evolution positions and rules for antigenic variants of human influenza A/H3N2 virusesabstractBACKGROUND: In pandemic and epidemic forms, avian and human influenza viruses often cause significant damage to human society and economics. Gradually accumulated mutations on hemagglutinin (HA) cause immunologically distinct circulating strains, which lead to the antigenic drift (named as antigenic variants). The "antigenic variants" often requires a new vaccine to be formulated before each annual epidemic. Mapping the genetic evolution to the antigenic drift of influenza viruses is an emergent issue to public health and vaccine development RESULTS: We developed a method for identifying antigenic critical amino acid positions, rules, and co-mutated positions for antigenic variants. The information gain (IG) and the entropy are used to measure the score of an amino acid position on hemagglutinin (HA) for discriminating between antigenic variants and similar viruses. A position with high IG and entropy implied that this position is highly correlated to an antigenic drift. Nineteen positions with high IG and high genetic diversity are identified as antigenic critical positions on the HA proteins. Most of these antigenic critical positions are located on five epitopes or on the surface based on the HA structure. Based on IG values and entropies of these 19 positions on the HA, the decision tree was applied to create a rule-based model and to identify rules for predicting antigenic variants of a given two HA sequences which are often a vaccine strain and a circulating strain. The predicting accuracies of this model on two sets, which consist of a training set (181 hemagglutination inhibition (HI) assays) and an independent test set (31,878 HI assays), are 91.2% and 96.2% respectively. CONCLUSION: Our method is able to identify critical positions, rules, and co-mutated positions on HA for predicting the antigenic variants. The information gains and the entropies of HA positions provide insight to the antigenic drift and co-evolution positions for influenza seasons. We believe that our method is robust and is potential useful for studying influenza virus evolution and vaccine development. Jhang-Wei Huang, Chwan-Chuen King, Jinn-Moon Yang |
BMC Bioinform. | 3 |
| 2008 | Evolutionary conservation of DNA-contact residues in DNA-binding domainsabstractBACKGROUND: DNA-binding proteins are of utmost importance to gene regulation. The identification of DNA-binding domains is useful for understanding the regulation mechanisms of DNA-binding proteins. In this study, we proposed a method to determine whether a domain or a protein can has DNA binding capability by considering evolutionary conservation of DNA-binding residues. RESULTS: Our method achieves high precision and recall for 66 families of DNA-binding domains, with a false positive rate less than 5% for 250 non-DNA-binding proteins. In addition, experimental results show that our method is able to identify the different DNA-binding behaviors of proteins in the same SCOP family based on the use of evolutionary conservation of DNA-contact residues. CONCLUSION: This study shows the conservation of DNA-contact residues in DNA-binding domains. We conclude that the members in the same subfamily bind DNA specifically and the members in different subfamilies often recognize different DNA targets. Additionally, we observe the co-evolution of DNA-contact residues and interacting DNA base-pairs. Yao-Lin Chang, Huai-Kuang Tsai, Cheng-Yan Kao, Yung-Chiang Chen, Yuh-Jyh Hu, Jinn-Moon Yang |
BMC Bioinform. | 6 |
| 2005 | GEMSCORE: A New Empirical Energy Function for Protein Folding
Yi-Yuan Chiu, Jenn-Kang Hwang, Jinn-Moon Yang |
CIBCB | 3 |
| 2004 | A pharmacophore-based evolutionary approach for screening estrogen receptor antagonistsabstractVirtual ligand screening has emerged as a promising approach for drug development. The inaccuracy of the scoring methods is probably major weakness for virtual ligand screening. In this paper, we have developed a pharmacophore-based evolutionary approach that was applicable to virtual screening and postdocking analysis. Our tool, referred to as the genetic evolutionary method for molecular docking (GEMDOCK), combines an evolutionary approach and a pharmacophore-based scoring function for virtual database screening. The former integrates discrete and continuous global search strategies with local search strategies to speed up convergence. The latter integrates a simple empirical scoring function and pharmacophore preferences. We accessed the screening accuracy of our approach on estrogen receptor alpha (ER/spl alpha/) using a ligand database on which competing tools were evaluated. The accuracies of our prediction were 0.64 for the GH score and 0.91% for the false positive rate when the true positive rate was 100%. We found that our pharmacophore-based scoring function indeed is able to reduce the number of the false positives. These results suggest that GEMDOCK is robust and can be a useful tool for virtual database screening. Jinn-Moon Yang, Tsai-Wen Shen |
IEEE Congress on Evolutionary Computation | 1 |
| 2004 | Comparative Molecular Binding Energy Analysis of HIV-1 Protease Inhibitors Using Genetic Algorithm-Based Partial Least Squares Method
Yen-Chih Chen, Jinn-Moon Yang, Chi-Hung Tsai, Cheng-Yan Kao |
GECCO (2) | 2 |
| 2004 | An Evolutionary Approach with Pharmacophore-Based Scoring Functions for Virtual Database Screening
Jinn-Moon Yang, Tsai-Wei Shen, Yen-Fu Chen, Yi-Yuan Chiu |
GECCO (1) | 1 |
| 2004 | Some issues of designing genetic algorithms for traveling salesman problems
Huai-Kuang Tsai, Jinn-Moon Yang, Yuan-Fan Tsai, Cheng-Yan Kao |
Soft Comput. | 2 |
| 2004 | An Evolutionary Approach for Gene Expression PatternsabstractThis study presents an evolutionary algorithm, called a heterogeneous selection genetic algorithm (HeSGA), for analyzing the patterns of gene expression on microarray data. Microarray technologies have provided the means to monitor the expression levels of a large number of genes simultaneously. Gene clustering and gene ordering are important in analyzing a large body of microarray expression data. The proposed method simultaneously solves gene clustering and gene-ordering problems by integrating global and local search mechanisms. Clustering and ordering information is used to identify functionally related genes and to infer genetic networks from immense microarray expression data. HeSGA was tested on eight test microarray datasets, ranging in size from 147 to 6221 genes. The experimental clustering and visual results indicate that HeSGA not only ordered genes smoothly but also grouped genes with similar gene expressions. Visualized results and a new scoring function that references predefined functional categories were employed to confirm the biological interpretations of results yielded using HeSGA and other methods. These results indicate that HeSGA has potential in analyzing gene expression patterns. Huai-Kuang Tsai, Jinn-Moon Yang, Yuan-Fan Tsai, Cheng-Yan Kao |
IEEE Trans. Inf. Technol. Biomed. | 2 |
| 2004 | An evolutionary algorithm for large traveling salesman problemsabstractThis work proposes an evolutionary algorithm, called the heterogeneous selection evolutionary algorithm (HeSEA), for solving large traveling salesman problems (TSP). The strengths and limitations of numerous well-known genetic operators are first analyzed, along with local search methods for TSPs from their solution qualities and mechanisms for preserving and adding edges. Based on this analysis, a new approach, HeSEA is proposed which integrates edge assembly crossover (EAX) and Lin-Kernighan (LK) local search, through family competition and heterogeneous pairing selection. This study demonstrates experimentally that EAX and LK can compensate for each other's disadvantages. Family competition and heterogeneous pairing selections are used to maintain the diversity of the population, which is especially useful for evolutionary algorithms in solving large TSPs. The proposed method was evaluated on 16 well-known TSPs in which the numbers of cities range from 318 to 13509. Experimental results indicate that HeSEA performs well and is very competitive with other approaches. The proposed method can determine the optimum path when the number of cities is under 10,000 and the mean solution quality is within 0.0074% above the optimum for each test problem. These findings imply that the proposed method can find tours robustly with a fixed small population and a limited family competition length in reasonable time, when used to solve large TSPs. Huai-Kuang Tsai, Jinn-Moon Yang, Yuan-Fang Tsai, Cheng-Yan Kao |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2003 | An Evolutionary Approach for Molecular Docking
Jinn-Moon Yang |
GECCO | 1 |
| 2002 | Solving traveling salesman problems by combining global and local search mechanismsabstractIn this paper an evolutionary algorithm for the traveling salesman problem is proposed. The key idea is to enhance the ability of exploration and exploitation by incorporating global search with local search. A new local search, called the neighbor-join (NJ) operator, is proposed to improve the solution quality of the edge assembly crossover (EAX) considered as a global search mechanism in this paper. Our method is applied to 15 well-known traveling salesman problems with numbers of cities ranging from 101 to 3038 cities. The experimental results indicate that the neighbor-join operator is very competitive with related operators surveyed in this paper. Incorporating the NJ into the EAX significantly outperforms the method incorporating 2-opt into the EAX for some hard problems. For each test instance the average value of solution quality stays within 0.03% from the optimum. For the notorious hard problem att532, it is able to find the optimum solution 23 times in 30 independent runs. Huai-Kuang Tsai, Jinn-Moon Yang, Cheng-Yan Kao |
IEEE Congress on Evolutionary Computation | 2 |
| 2002 | Applying Genetic Algorithms To Finding The Optimal Gene Order In Displaying The Microarray Data
Huai-Kuang Tsai, Jinn-Moon Yang, Cheng-Yan Kao |
GECCO | 2 |
| 2001 | Integrating adaptive mutations and family competition with differential evolution for flexible ligand dockingabstractA flexible ligand docking protocol based on evolutionary algorithms is investigated. The proposed approach integrates decreasing-based mutations and self-adaptive mutations with differential evolution. This approach possesses global and local search strategies to balance the trade-off between exploitation and exploration of the search. The proposed approach is applied to a dihydrofolate reductase enzyme with the anti-cancer drug methotrexate and two analogues of antibacterial drug trimethoprim. Numerical results indicate that the new approach is very robust. Jinn-Moon Yang, Jorng-Tzong Horng, Cheng-Yan Kao |
CEC | 1 |
| 2001 | Optical Coating Designs Using the Family Competition Evolutionary AlgorithmabstractA robust evolutionary approach, called the Family Competition Evolutionary Algorithm (FCEA), is described for the synthesis of optical thin-film designs. Based on family competition and adaptive rules, the proposed approach consists of global and local strategies by integrating decreasing mutations and self-adaptive mutations. The method is applied to three different optical coating designs with complex spectral quantities. Numerical results indicate that the proposed approach performs very robustly and is very competitive with other approaches. Jinn-Moon Yang, Jorng-Tzong Horng, Chih-Jen Lin, Cheng-Yan Kao |
Evol. Comput. | 1 |
| 2001 | A Robust Evolutionary Algorithm for Training Neural Networks
Jinn-Moon Yang, Cheng-Yan Kao |
Neural Comput. Appl. | 1 |
| 2000 | A robust evolutionary algorithm for optical thin-film designsabstractThis paper presents an evolutionary approach, called the family competition evolutionary algorithm (FCEA), for optical thin film design. The proposed approach, based on family competition and multiple adaptive rules, integrates decreasing-based Gaussian mutation and two self-adaptive mutations to balance the exploitation and exploration. It is implemented and applied to two coating systems. Numerical results indicate that the proposed approach is very robust for optical coatings. Jinn-Moon Yang, Cheng-Yen Kao |
CEC | 1 |
| 2000 | A Genetic Algorithm for Physical Mapping Problems
Huai-Kuang Tsai, Cheng-Yan Kao, Jinn-Moon Yang |
GECCO | 3 |
| 2000 | An Evolutionary Algorithms to Training Neural Networks for a Two-Spiral Problem
Jinn-Moon Yang, Cheng-Yan Kao |
GECCO | 1 |
| 2000 | A Genetic Algorithm with Adaptive Mutations and Family Competition for Training Neural NetworksabstractIn this paper, we present a new evolutionary technique to train three general neural networks. Based on family competition principles and adaptive rules, the proposed approach integrates decreasing-based mutations and self-adaptive mutations to collaborate with each other. Different mutations act as global and local strategies respectively to balance the trade-off between solution quality and convergence speed. Our algorithm is then applied to three different task domains: Boolean functions, regular language recognition, and artificial ant problems. Experimental results indicate that the proposed algorithm is very competitive with comparable evolutionary algorithms. We also discuss the search power of our proposed approach. Jinn-Moon Yang, Jorng-Tzong Horng, Cheng-Yan Kao |
Int. J. Neural Syst. | 1 |
| 2000 | Integrating adaptive mutations and family competition into genetic algorithms as function optimizer
Jinn-Moon Yang, Cheng-Yan Kao |
Soft Comput. | 1 |
| 2000 | A family competition evolutionary algorithm for automated docking of flexible ligands to proteinsabstractIn this paper, we study an evolutionary algorithm for flexible ligand docking. Based on family competition and adaptive rules, the proposed approach consists of global and local strategies by integrating decreasing mutations and self-adaptive mutations. To demonstrate the robustness of the proposed approach, we apply it to the problems of the first international contests on evolutionary optimization. Following the description of function optimization, our approach is applied to a dihydrofolate reductase enzyme with the anti-cancer drug methotrexate and with two analogs of the antibacterial drug trimethoprim. Our numerical results indicate that the proposed approach is robust. The docked lowest energy structures have rms derivations ranging from 0.72 A to 1.98 A with respect to the corresponding crystal structure. Jinn-Moon Yang, Cheng-Yan Kao |
IEEE Trans. Inf. Technol. Biomed. | 1 |
| 1999 | Incorporation family competition into Gaussian and Cauchy mutations to training neural networks using an evolutionary algorithmabstractThe paper presents an evolutionary technique to train neural networks in tasks requiring learning behavior. Based on family competition principles and adaptive rules, the proposed approach integrates decreasing-based mutations and self-adaptive mutations. Different mutations act global and local strategies separately to balance the trade-off between solution quality and convergence speed. The algorithm proposed herein is applied to two different task domains: Boolean functions and artificial ant problem. Experimental results indicate that in all tested problems, the proposed algorithm performs better than other canonical evolutionary algorithms, such as genetic algorithms, evolution strategies, and evolutionary programming. Moreover, essential components such as mutation operators and adaptive rules in the proposed algorithm are thoroughly analyzed. Jinn-Moon Yang, Jorng-Tzong Horng, Cheng-Yen Kao |
CEC | 1 |
| 1998 | A New Evolutionary Approach to Developing Neural Autonomous AgentsabstractThis paper explores the use of neural networks to control robots in tasks requiring sequential and learning behavior. We propose a family competition evolutionary algorithm (FCEA) to evolve networks that can integrate these different types of behavior in a smooth and continuous manner. The approach integrates self-adaptive Gaussian mutation, self-adaptive Cauchy mutation, decreasing-based Gaussian mutation, and family competition. In order to illustrate the power of the approach, we apply this approach to two different task domains: the "artificial ant" problem and a sequential behavior problem - an agent learns to play football. From the experimental results, we find our approach performs much better than other evolutionary algorithms in these two tasks. Based on the results from our experiments, it is shown that our approach can evolve neural networks to provide a means of integrating, sequencing and learning within a single control system. Jinn-Moon Yang, Jorng-Tzong Horng, Cheng-Yan Kao |
ICRA | 1 |
| 1998 | An Evolutionary Algorithm for Synthesizing Optical Thin-Film Designs
Jinn-Moon Yang, Cheng-Yan Kao |
PPSN | 1 |