M. Michael Gromiha

dblp:07/1037 · DBLP profile ↗
← Back
64ranked-venue papers
12as first author
14since 2021 · last 2025
0000-0002-1776-4096ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 61 · 11 first-author · 14 since 2021Artificial intelligence and machine learning · 3 · 1 first-author
YearPublicationVenuePosition
2025 Comment on "Contrastive pre-training and 3D convolution neural network for RNA and small molecule binding affinity prediction" by Sun and Gao
abstract
Dear Editor, In a recent article in Bioinformatics, Sun and Gao (2024) reported a structure-based method for predicting the RNA-ligand binding affinity using a pretrained CNN-based model (RLaffinity) based on the data in PDBbind database (Liu et al. 2017). The performance of the model has been reported to be significantly better than other existing methods in the literature. However, upon closer inspection of the dataset and model predictions, the method has several shortcomings such as: (i) mixing the data for RNA-ligand and DNA-ligand complexes in the datasets, (ii) utilizing double normalization to the predictive variable, (iii) inability of the model to generalize to unseen RNA-ligand complex structures, (iv) prediction performance is not better than existing methods, and (v) incorrect comparison with existing prediction methods in the domain. We have discussed these concerns in detail in the following sections with relevant examples and metrics. The canonical functional forms of DNA structures (B-form) and RNA structures (A-form) vary in cellulo, and hence, can have an impact on the structural features learned by the RLaffinity model for binding affinity prediction. Upon inspecting the dataset splits provided in the GitHub repository associated with the article (Sun and Gao 2024), we found multiple instances of DNA-ligand complexes as part of the RNA-ligand complex dataset used by the authors. For example, the 10 test datasets with 13 complexes each, were found to include 19 DNA-ligand complexes (1P96, 1QMS, 1QV4, 1QV8, 1R4E, 2MG8, 5W77, 6J2W, 6JJ0, 2JWQ, 408D, 1CVX, 1DB6, 1CVY, 316D, 1YKV, 209D, 1I9V, 407D), which amount to 26 test datapoints (20% of 130). While the authors claim to have predicted the RNA-ligand binding affinity using the RLaffinity model (Sun and Gao 2024), the predicted value is a double-scaled version of the binding affinity (log-transformation + min-max normalization). MinMax normalized values are devoid of outlier effects, leading to artificially higher Pearson correlation coefficients, which underscore the true potential of the predictive model. The MinMax normalized pKd values from 144 RNA-ligand complexes have been used to train, validate, and test the model. For the available RNA-ligand complex structures with pKd values outside the range of normalization in PDBbind ([2.508, 10.958]) (Liu et al. 2017), the predictions from RLaffinity model may not generalize well, resulting in increased prediction error. For example, previous studies have shown that theophylline-binding RNA aptamer can efficiently discriminate theophylline from caffeine by a 10 000-fold difference in affinity (Jenison et al. 1994). Consequently, the binding affinity of theophylline aptamer to caffeine has been reported to be 3500 μM (Jenison et al. 1994, Menichelli et al. 2022), which leads to the pKd value of 2.455. Using the PDB structure 1EHT as the template, the theophylline aptamer-caffeine complex structure was prepared and passed as input to RLaffinity. The predicted min-max normalized log-scale Kd from RLaffinity is 0.426, which translates to a pKd of 6.1077 and Kd of 0.78 μM. Comparing to experimental pKd value (2.455), it showed a high absolute error of 3.652. This example highlights the pitfall of using min-max normalization of pKd values for the predictive variable. We curated a dataset of 12 RNA-ligand complexes, which are not present among the 144 complexes used in the RLaffinity study (Sun and Gao 2024). The experimental binding affinity values (Kd) of these complexes were taken from R-SIM (Krishnan et al. 2023) as reported in the associated publication in PDB. These 12 structures were tested using the source code implementation of the RLaffinity model (Sun and Gao 2024). Upon pre-processing of the dataset, only 7 PDB structures could be processed (7D7Y, 7D7X, 3SKZ, 6UC8, 3SLM, 6UC7, 6DN1) by the code due to RDKit errors in SDF/MOL2 parsing (6DN2, 6DN3, 7WIE, 7WIF, 7WII). The binding affinity values predicted by the RLaffinity model for these seven RNA-ligand complexes are provided in Table 1. We observed that while the experimental pKd values of the seven complexes range from 3.42 to 8.09, the pKd values obtained from back-transformation of the min-max normalized RLaffinity predictions are all in a narrow range of 5.03 to 5.90. Although the method shows a mean absolute error of 1.384, the correlation is very weak between experimental and predicted affinities with a Pearson’s Correlation Coefficient (PCC) of 0.571 and a Spearman’s Correlation Coefficient (SPCC) of 0.535. Notably, the complexes with high (6DN1 and 6UC8) and low pKd values (7D7X and 3SLM) are predicted with high mean absolute error (MAE) (1.37 to 2.2). Further, as discussed above, RLaffinity is not able to correctly predict the binding affinity of RNA-ligand complexes, which have pKd values beyond the range used for normalization. Binding affinity predictions obtained for 7 RNA-ligand complexes from RLaffinity (Sun and Gao 2024) and RSAPred (Krishnan et al. 2024).a All predicted values have been rounded off to two decimal places for clarity. The RSAPred riboswitch model was used for predictions, as all seven test datapoints are riboswitch-ligand interactions. The values in bold indicate the predictions from RLaffinity with MAE greater than 1. Binding affinity predictions obtained for 7 RNA-ligand complexes from RLaffinity (Sun and Gao 2024) and RSAPred (Krishnan et al. 2024).a All predicted values have been rounded off to two decimal places for clarity. The RSAPred riboswitch model was used for predictions, as all seven test datapoints are riboswitch-ligand interactions. The values in bold indicate the predictions from RLaffinity with MAE greater than 1. The predictions from RLaffinity were also compared with the recent method, RSAPred (Krishnan et al. 2024), for RNA-ligand binding affinity prediction. It should be noted that these 12 RNA-ligand complexes were not part of either the training or test datasets used by RSAPred. From Table 1, it can be observed that the predictions from RSAPred range from 2.4 to 9.79, capturing the diversity in pKd values collected from literature, compared to RLaffinity, which predicts within a narrow range of 5.03–5.90. Further, RSAPred showed a PCC and SPCC of 0.90 and 0.75, respectively, whereas RLaffinity has a PCC of 0.571 and SPCC of 0.535 on the unseen dataset of seven RNA-ligand complexes. This result shows that the performance of RLaffinity model is not better than the existing method, RSAPred. As part of model validation, RLaffinity (Sun and Gao 2024) was compared with three existing methods namely, Vina (Trott and Olson 2010), RF-Score (Ballester and Mitchell 2010), and RSAPred (Krishnan et al. 2024). In case of Vina and RF-Score, both methods were primarily trained to score protein-ligand complexes, and hence have been shown to be unreliable for comparison with RNA-specific methods by previous studies (Stefaniak and Bujnicki 2021, Agarwal et al. 2023). On the other hand, the comparison with RSAPred (Krishnan et al. 2024) has been made without normalization of the predicted binding affinities. Hence, by comparing the correlation coefficients of absolute binding affinity predictions (pKd from RSAPred) to relative predictions (MinMax normalized pKd from RLaffinity), the authors have highlighted a significant improvement in model performance. Further, RSAPred is a suite of binding affinity prediction models specific to six RNA subtypes (Aptamers, miRNAs, Repeats, Ribosomal RNAs, Riboswitches, and Viral RNAs), which was not considered in predicting the binding affinity. Hence, the comparison is not appropriate with protein-ligand binding affinity methods as well as MinMax normalized pKd from RLaffinity and pKd from RSAPred. Our analysis emphasizes the necessity of considering several important factors to validate the performance of RNA-ligand binding affinity prediction methods. It includes (i) utilizing RNA-ligand complexes alone for developing a model and validation, (ii) avoiding min-max normalization of log-normalized binding affinity values for prediction, (iii) testing the ability of the model to generalize well to unseen RNA-ligand interactions, (iv) improving the performance of the method, and (v) providing proper comparison with existing methods in the domain. Conflict of interest: S.R.K. and A.R. are affiliated to Tata Consultancy Services Limited at the time of submission. None declared. All data shown in the manuscript is available from PDB database and RSAPred web server.
Sowmya Ramaswamy Krishnan, Arijit Roy 0003, M. Michael Gromiha
Bioinform.3
2024 Reliable method for predicting the binding affinity of RNA-small molecule interactions using machine learning
abstract
Ribonucleic acids (RNAs) play important roles in cellular regulation. Consequently, dysregulation of both coding and non-coding RNAs has been implicated in several disease conditions in the human body. In this regard, a growing interest has been observed to probe into the potential of RNAs to act as drug targets in disease conditions. To accelerate this search for disease-associated novel RNA targets and their small molecular inhibitors, machine learning models for binding affinity prediction were developed specific to six RNA subtypes namely, aptamers, miRNAs, repeats, ribosomal RNAs, riboswitches and viral RNAs. We found that differences in RNA sequence composition, flexibility and polar nature of RNA-binding ligands are important for predicting the binding affinity. Our method showed an average Pearson correlation (r) of 0.83 and a mean absolute error of 0.66 upon evaluation using the jack-knife test, indicating their reliability despite the low amount of data available for several RNA subtypes. Further, the models were validated with external blind test datasets, which outperform other existing quantitative structure-activity relationship (QSAR) models. We have developed a web server to host the models, RNA-Small molecule binding Affinity Predictor, which is freely available at: https://web.iitm.ac.in/bioinfo2/RSAPred/.
Sowmya Ramaswamy Krishnan, Arijit Roy 0003, M. Michael Gromiha
Briefings Bioinform.3
2024 MPA-MutPred: a novel strategy for accurately predicting the binding affinity change upon mutation in membrane protein complexes
abstract
Mutations in the interface of membrane protein (MP) complexes are key contributors to a broad spectrum of human diseases, primarily due to changes in their binding affinities. While various methods exist for predicting the mutation-induced changes in binding affinity (ΔΔG) in protein-protein complexes, none are specific to MP complexes. This study proposes a novel strategy for ΔΔG prediction in MP complexes, which combines linear and nonlinear models, to obtain a more robust model with improved prediction accuracy. We used multiple linear regression to extract informative features that influence the binding affinity in MP complexes, which included changes in the stability of the complex, conservation score, electrostatic interaction, relatively accessible surface area, and interface contacts. Further, using gradient boosting regressor on the selected features, we developed MPA-MutPred, a novel method specific for predicting the ΔΔG of membrane protein-protein complexes, and it is freely accessible at https://web.iitm.ac.in/bioinfo2/MPA-MutPred/. Our method achieved a correlation of 0.75 and a mean absolute error (MAE) of 0.73 kcal/mol in the jack-knife test conducted on a dataset of 770 mutants. We further validated the method using a blind test set of 86 mutations, obtaining a correlation of 0.85 and an MAE of 0.77 kcal/mol. We anticipate that this method can be used for large-scale studies to understand the influence of binding affinity change on disease-causing mutations in MP complexes, thereby aiding in the understanding of disease mechanisms and the identification of potential therapeutic targets.
Fathima Ridha, M. Michael Gromiha
Briefings Bioinform.2
2024 DeepPPAPredMut: deep ensemble method for predicting the binding affinity change in protein-protein complexes upon mutation
abstract
MOTIVATION: Protein-protein interactions underpin many cellular processes and their disruption due to mutations can lead to diseases. With the evolution of protein structure prediction methods like AlphaFold2 and the availability of extensive experimental affinity data, there is a pressing need for updated computational tools that can efficiently predict changes in binding affinity caused by mutations in protein-protein complexes. RESULTS: We developed a deep ensemble model that leverages protein sequences, predicted structure-based features, and protein functional classes to accurately predict the change in binding affinity due to mutations. The model achieved a correlation of 0.97 and a mean absolute error (MAE) of 0.35 kcal/mol on the training dataset, and maintained robust performance on the test set with a correlation of 0.72 and a MAE of 0.83 kcal/mol. Further validation using Leave-One-Out Complex (LOOC) cross-validation exhibited a correlation of 0.83 and a MAE of 0.51 kcal/mol, indicating consistent performance. AVAILABILITY AND IMPLEMENTATION: https://web.iitm.ac.in/bioinfo2/DeepPPAPredMut/index.html.
Rahul Nikam, Sherlyn Jemimah, M. Michael Gromiha
Bioinform.3
2023 TMKit: a Python interface for computational analysis of transmembrane proteins
abstract
Transmembrane proteins are receptors, enzymes, transporters and ion channels that are instrumental in regulating a variety of cellular activities, such as signal transduction and cell communication. Despite tremendous progress in computational capacities to support protein research, there is still a significant gap in the availability of specialized computational analysis toolkits for transmembrane protein research. Here, we introduce TMKit, an open-source Python programming interface that is modular, scalable and specifically designed for processing transmembrane protein data. TMKit is a one-stop computational analysis tool for transmembrane proteins, enabling users to perform database wrangling, engineer features at the mutational, domain and topological levels, and visualize protein-protein interaction interfaces. In addition, TMKit includes seqNetRR, a high-performance computing library that allows customized construction of a large number of residue connections. This library is particularly well suited for assigning correlation matrix-based features at a fast speed. TMKit should serve as a useful tool for researchers in assisting the study of transmembrane protein sequences and structures. TMKit is publicly available through https://github.com/2003100127/tmkit and https://tmkit-guide.herokuapp.com/doc/overview.
Arulsamy Kulandaisamy, Jinlong Ru, M. Michael Gromiha, Adam P. Cribbs
Briefings Bioinform.4
2022 Identification of potential driver mutations in glioblastoma using machine learning
abstract
Glioblastoma is a fast and aggressively growing tumor in the brain and spinal cord. Mutation of amino acid residues in targets proteins, which are involved in glioblastoma, alters the structure and function and may lead to disease. In this study, we collected a set of 9386 disease-causing (drivers) mutations based on the recurrence in patient samples and experimentally annotated as pathogenic and 8728 as neutral (passenger) mutations. We observed that Arg is highly preferred at the mutant sites of drivers, whereas Met and Ile showed preferences in passengers. Inspecting neighboring residues at the mutant sites revealed that the motifs YP, CP and GRH, are preferred in drivers, whereas SI, IQ and TVI are dominant in neutral. In addition, we have computed other sequence-based features such as conservation scores, Position Specific Scoring Matrices (PSSM) and physicochemical properties, and developed a machine learning-based method, GBMDriver (GlioBlastoma Multiforme Drivers), for distinguishing between driver and passenger mutations. Our method showed an accuracy and AUC of 73.59% and 0.82, respectively, on 10-fold cross-validation and 81.99% and 0.87 in a blind set of 1809 mutants. The tool is available at https://web.iitm.ac.in/bioinfo2/GBMDriver/index.html. We envisage that the present method is helpful to prioritize driver mutations in glioblastoma and assist in identifying therapeutic targets.
Medha Pandey, P. Anoosha, Dhanusha Yesudhas, M. Michael Gromiha
Briefings Bioinform.4
2022 Erratum to: Evaluation of in silico tools for the prediction of protein and peptide aggregation on diverse datasets
abstract
Briefings in Bioinformatics, bbab240, doi: https://doi.org/10.1093/bib/bbab240. At the end of the article, the name of Yutaka Saito was corrected to be Puneet Rawat.
R. Prabakaran, Puneet Rawat, M. Michael Gromiha
Briefings Bioinform.4
2022 Srinivasan (1962-2021) in Bioinformatics and beyond
abstract
Dear Editor, Last year the Bioinformatics community lost one of its pioneers, a scientist renowned for his talent, creativity and rigour but also for his commitment to supporting his research community and particularly the young scientists he trained. He was an inspiring role model for his field and a scientist who will be remembered very fondly by his many friends in the community for his warmth, humour and kindness. For more than three decades Srinivasan developed timely and novel computational strategies for analyzing proteins and was regarded in high esteem internationally for the insights he provided and the resources he established based on the underpinning concepts. His discoveries cover many areas fundamental to structural biology and pathogen research. Although he was a computational scientist, he worked closely with experimental groups to maximize the impact of his research. He published more than 300 papers, with nearly 10 000 citations altogether. Srinivasan joined the faculty of the Molecular Biophysics Unit, Indian Institute of Science, Bangalore in 1998, after leaving the Madras Biophysics Group (he did his Masters studies from 1982-84). He acquired his PhD degree in the Molecular Biophysics Department (the same Department where he later worked as a faculty member) within the GN Ramachandran school of peptide and peptide stereochemistry. His postdoctoral tenure was in Prof. Sir Tom Blundell’s laboratory (1991–1998), Birkbeck College, UK, with a brief stint in Prof. Mike Waterfield’s laboratory at the Ludwig Institute for Cancer Research, UK. He arrived in London as a seemingly shy young man, but it soon became clear that he was a real expert in protein structures and thought very deeply about their evolution. During these times, his research was largely focused on homology modelling (Johnson et al., 1994) and the study of proteins involved in signal transduction (e.g. Srinivasan et al., 1994, 1996). After coming back to MBU, he headed the ‘Proteins: structure, function and evolutionary’ group. He made major contributions to the understanding the structure and functions of proteins, particularly on protein kinases in a wide range of model organisms (e.g. Krupa and Srinivasan, 2002; Krupa et al., 2004a, b). Specifically, his lab was focused on computational genomics, bioinformatics and structural biology, particularly involving the relationships between protein structure, function and interactions, including protein–protein interactions, cellular signal transduction and biological pathways. Srinivasan’s interest in protein families and protein evolution drove research into strategies for improving multiple alignments of relatives and for better characterizing phylogenetic relationships (which resulted in many useful resources like SUPFAM, MulPSSM, PALI and DoSA). It also drove the design of methods to detect extremely remote homologues, which have been valuable for extending structural and functional annotation of genomes. Srinivasan’s group also showed that sequence-based connections of distantly related proteins can be enabled through the design of artificial sequences (Mudgal et al., 2014). This work was highly innovative and can help to bring valuable annotations for pathogen proteins, which are typically difficult to characterize by more conventional, less-sensitive strategies. His strategies allowed a much deeper characterization of fold space to inform protein engineering. Srinivasan is also highly renowned for his analyses of how changes in the protein structure and sequence impact function. Some protein families, like the kinases, were a major focus of his research and gave him much international acclaim. He studied kinases for >20 years and contributed numerous insights important for understanding their mechanisms and for enabling drug design. For example, structural fluctuations, classifications based on key functional site properties (Kalaivani et al., 2018), mechanisms of stabilization of their key functional sites through specific residue interactions. He also characterized the ways in which domain partnerships modify structure (Vishwanath et al., 2018), kinase functionality and characterized how splicing extends the kinase functional repertoire. This large family is implicated in many human diseases, including cancer, and these discoveries have informed drug design. However, the biological role of proteins is determined by their interactions and Srinivasan applied his precise analytical skills in this arena, too, revealing key insights into the properties of the interfaces involved in assembling protein complexes. He produced a substantial body of very rigorous studies, including analyses of the characteristics of transient complexes and the effect of protein associations on global structural dynamics. He robustly captured this knowledge in the PIC protein interactions calculator (Tina et al., 2007), a valuable tool that is freely available to biologists and very popular among researchers to obtain structural data on various non-covalent interactions within a protein or between proteins in a complex. An important application of these methods was the characterization of interactions between viral proteins and their host proteins, which provided key data for understanding pathogenicity and enabling drug design. For example, Srinivasan performed various studies characterizing toxin–antitoxin systems (Tandon et al., 2019), protein interactions between human erythrocytes and Plasmodium falciparum and Helicobacter pylori and human. Photo taken at the fifth IIT Madras-Tokyo Tech joint symposium on ‘Current Trends in Bioinformatics: Big Data Analysis, Machine Learning and Drug Design’ with leading Bioinformatics scientists in India (March 2020). From left to right D. Velmurugan, N. Manoj, S. Selvaraj, P.K. Ponnuswamy, G.P.S. Raghava, K. Veluraja, Shandar Ahmad, M. Michael Gromiha, N. Srinivasan, R. Sowdhamini Photo taken at the fifth IIT Madras-Tokyo Tech joint symposium on ‘Current Trends in Bioinformatics: Big Data Analysis, Machine Learning and Drug Design’ with leading Bioinformatics scientists in India (March 2020). From left to right D. Velmurugan, N. Manoj, S. Selvaraj, P.K. Ponnuswamy, G.P.S. Raghava, K. Veluraja, Shandar Ahmad, M. Michael Gromiha, N. Srinivasan, R. Sowdhamini In collaboration with multiple laboratories, Srini’s group studied fascinating biological systems, including protein assemblies of ribosomes and spliceosomes (Bhat et al., 2015; Pudi et al., 2003; Yazhini et al., 2022) and developed powerful computational tools for studying structures of large assemblies derived from cryo-electron microscopy (Joseph et al., 2016; Rakesh et al., 2016). As is true of several structural bioinformaticians, his group relied on publicly available structural data and were concerned with the quality of protein–ligand data, deposited through X-ray or cryo-EM studies (Chakraborti et al., 2021). He was always excited to discuss the Ramachandran map. Last year, he attended the fifth IIT Madras—Tokyo Tech joint symposium on Bioinformatics and his lecture on the Ramachandran map was fascinating. Using modern computational tools and along with late Prof. C. Ramakrishnan (his PhD mentor and the student behind the original Ramachandran map) and one of his students Ashraya Ravikumar, he re-examined the classical and renowned Ramachandran map. They clearly demonstrated that it is possible to consider deviations in the ‘allowed’ regions within this map by considering slight deviations in internal parameters from ideal values of the peptide bond (Ravikumar et al., 2019). Srinivasan’s group also participated in several consortia such as the Open Source Drug Discovery program [with the groups of Prof. Tom Blundell (University of Cambridge, UK), Nagasuma Chandra (Indian Institute of Science, India) and Sowdhamini (National Centre for Biological Sciences, India)], the UKIERI study of protein assemblies [with the groups of Profs. Jim Warwicker (University of Manchester, UK), Pinak Chakrabarti (Bose Institute, India), Nagasuma Chandra and Sowdhamini)] and collaborations such as the Indo-French CEFIPRA project on protein alphabets [with Dr. Alexandre de Brevern (INSERM Paris, France) and Dr. Bernard Offmann (University of Nantes, France)], and the Centre for Excellence on protein-protein interactions [with Profs. Sowdhamini and Satyajit Mayor (National Centre for Biological Sciences, India) and Nagasuma Chandra (Indian Institute of Science, India)] and toxin-antitoxin systems [with Prof. Raghavan Varadarajan (Indian Institute of Science, India)]. Throughout his career Srinivasan applied his knowledge, data and computational tools to characterize the protein structures, functions and virus–host interactions of multiple pathogenic bacteria affecting human health, including mycobacterial pathogens (e.g. Mtb), malaria, H. Pylori, Dengue and several gut pathogens. Understanding the critical residues in the protein interface is essential for drug design to reduce infection and pathogenicity. For many years he collaborated with the group of Professor Tom Blundell in Cambridge, UK. As well as detailed analyses, he established the SInCRe structural interactome resource for Mtb in 2015 (Metri et al., 2015), which contributed to studies on the repurposing of drugs for this pathogen. His tools have been applied in a number of medical contexts with promising clinical results. His group also applied docking tools to FDA-approved drugs to SARS-CoV2 (Chakraborti et al., 2020) and his most recent work on inhibitors for the main protease of SARS-CoV2 led to compounds already in clinical trials. Prof. Srinivasan made significant contributions to Bioinformatics and his whole-hearted involvement in scientific activities will not be forgotten. As a scientist, Srinivasan was very highly focused and meticulous. He always set high standards—whether in creating high-quality datasets or in his interpretations of data or in responding to reviewers’ comments. As well as being a multitalented and well-known researcher, Srinivasan actively participated in many university and external committees, commented on PhD theses, and delivered popular and invited lectures in most of the leading conferences in India. He was an elected fellow in all the three major academies in India (Indian National Science Academy, New Delhi, National Academy of Sciences, Allahabad and Indian Academy of Sciences, Bangalore). He also received the most prestigious awards in India including Shanti Swarup Bhatnagar Prize for Science and Technology from Council of Scientific and Industrial Research, National Bioscience Award from the Department of Biotechnology and J.C. Bose National Fellowship from the Department of Science and Technology, Government of India. Srinivasan had the special ability to cordially relate with others he respected and had a very positive attitude towards the work of his fellow researchers. He was a faithful chairperson of the Department (serving between 2018 and 2020) and always supported and wished his younger colleagues to do well. He remained active even when he was critically ill. During this time he still managed to publish around 10 papers, enable six of his lab colleagues to reach higher positions and also attended to multiple student-thesis-related matters. He was very enthusiastic about his research on protein structures and his ability to explain major concepts in a simple accessible manner was extremely impressive. He had a passion for naming his students with ‘amino acids’, each with a background story and spent considerable time with his students in the midst of his busy schedule. He always encouraged young researchers and provided valuable advice for their research. He also had an uncanny enthusiasm and ability to make sure that the people around him felt included and important—he would not hesitate to talk to prospective students, spend a long time on discussions with visitors and provide his undivided attention on work discussions with his students. Many of his students travelled widely and benefitted laboratories and science throughout the world. His students made a large impact wherever they went because their knowledge was always deep and impressive and they showed great enthusiasm for their work, mirroring their mentor. They often liked to discuss their work in detail and place it in the wider context of global knowledge. Although based in India for most of his career, Srinivasan travelled widely and has had an impact on many scientists involved in protein structure analysis. He was also a great host for visitors, ensuring their well-being and spending time discussing their work and ideas. Srinivasan had been a very special and unusual personality—with unlimited affection and love for the people around him. He was able to sense people in trouble and would often go out of his way to help them. His smiling face is not forgettable at any time and evidenced a deeply contented person, very proud of his family, his students and his science and always happy to discuss anything to do with proteins. His positivity and passion for science were infectious. His too-early passing is a great loss to science and to everyone who knew him. Financial Support: none declared. Conflict of Interest: The authors declare that there are no conflicts of interest.
M. Michael Gromiha, Christine A. Orengo, Ramanathan Sowdhamini, Janet M. Thornton
Bioinform.1
2022 Ab-CoV: a curated database for binding affinity and neutralization profiles of coronavirus-related antibodies
abstract
SUMMARY: We have developed a database, Ab-CoV, which contains manually curated experimental interaction profiles of 1780 coronavirus-related neutralizing antibodies. It contains more than 3200 datapoints on half maximal inhibitory concentration (IC50), half maximal effective concentration (EC50) and binding affinity (KD). Each data with experimentally known three-dimensional structures are complemented with predicted change in stability and affinity of all possible point mutations of interface residues. Ab-CoV also includes information on epitopes and paratopes, structural features of viral proteins, sequentially similar therapeutic antibodies and Collier de Perles plots. It has the feasibility for structure visualization and options to search, display and download the data. AVAILABILITY AND IMPLEMENTATION: Ab-CoV database is freely available at https://web.iitm.ac.in/bioinfo2/ab-cov/home. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Puneet Rawat, Divya Sharma 0002, R. Prabakaran, Fathima Ridha, Mugdha Mohkhedkar, Vani Janakiraman, M. Michael Gromiha
Bioinform.7
2021 MPTherm: database for membrane protein thermodynamics for understanding folding and stability
abstract
The functions of membrane proteins (MPs) are attributed to their structure and stability. Factors influencing the stability of MPs differ from globular proteins due to the presence of membrane spanning regions. Thermodynamic data of MPs aid to understand the relationship among their structure, stability and function. Although a wealth of experimental data on thermodynamics of MPs are reported in the literature, there is no database available explicitly for MPs. In this work, we have developed a database for MP thermodynamics, MPTherm, which contains more than 7000 thermodynamic data from about 320 MPs. Each entry contains protein sequence and structural information, membrane topology, experimental conditions, thermodynamic parameters such as melting temperature, free energy, enthalpy etc. and literature information. MPTherm assists users to retrieve the data by using different search and display options. We have also provided the sequence and structure visualization as well as cross-links to UniProt and PDB databases. MPTherm database is freely available at http://www.iitm.ac.in/bioinfo/mptherm/. It is implemented in HTML, PHP, MySQL and JavaScript, and supports the latest versions of major browsers, such as Firefox, Chrome and Opera. MPTherm would serve as an effective resource for understanding the stability of MPs, development of prediction tools and identifying drug targets for diseases associated with MPs.
A. Kulandaisamy, M. Michael Gromiha
Briefings Bioinform.3
2021 Evaluation of in silico tools for the prediction of protein and peptide aggregation on diverse datasets
abstract
Several prediction algorithms and tools have been developed in the last two decades to predict protein and peptide aggregation. These in silico tools aid to predict the aggregation propensity and amyloidogenicity as well as the identification of aggregation-prone regions. Despite the immense interest in the field, it is of prime importance to systematically compare these algorithms for their performance. In this review, we have provided a rigorous performance analysis of nine prediction tools using a variety of assessments. The assessments were carried out on several non-redundant datasets ranging from hexapeptides to protein sequences as well as amyloidogenic antibody light chains to soluble protein sequences. Our analysis reveals the robustness of the current prediction tools and the scope for improvement in their predictive performances. Insights gained from this work provide critical guidance to the scientific community on advantages and limitations of different aggregation prediction methods and make informed decisions about their research needs.
R. Prabakaran, Puneet Rawat, M. Michael Gromiha
Briefings Bioinform.4
2021 Prediction of protein-carbohydrate complex binding affinity using structural features
abstract
Protein-carbohydrate interactions play a major role in several cellular and biological processes. Elucidating the factors influencing the binding affinity of protein-carbohydrate complexes and predicting their free energy of binding provide deep insights for understanding the recognition mechanism. In this work, we have collected the experimental binding affinity data for a set of 389 protein-carbohydrate complexes and derived several structure-based features such as contact potentials, interaction energy, number of binding residues and contacts between different types of atoms. Our analysis on the relationship between binding affinity and structural features revealed that the important factors depend on the type of the complex based on number of carbohydrate and protein chains. Specifically, binding site residues, accessible surface area, interactions between various atoms and energy contributions are important to understand the binding affinity. Further, we have developed multiple regression equations for predicting the binding affinity of protein-carbohydrate complexes belonging to six categories of protein-carbohydrate complexes. Our method showed an average correlation and mean absolute error of 0.731 and 1.149 kcal/mol, respectively, between experimental and predicted binding affinities on a jackknife test. We have developed a web server PCA-Pred, Protein-Carbohydrate Affinity Predictor, for predicting the binding affinity of protein-carbohydrate complexes. The web server is freely accessible at https://web.iitm.ac.in/bioinfo2/pcapred/. The web server is implemented using HTML and Python and supports recent versions of major browsers such as Chrome, Firefox, IE10 and Opera.
N. R. Siva Shanmugam, J. Jino Blessy, K. Veluraja, M. Michael Gromiha
Briefings Bioinform.4
2021 Mutations in transmembrane proteins: diseases, evolutionary insights, prediction and comparison with globular proteins
abstract
Membrane proteins are unique in that they interact with lipid bilayers, making them indispensable for transporting molecules and relaying signals between and across cells. Due to the significance of the protein's functions, mutations often have profound effects on the fitness of the host. This is apparent both from experimental studies, which implicated numerous missense variants in diseases, as well as from evolutionary signals that allow elucidating the physicochemical constraints that intermembrane and aqueous environments bring. In this review, we report on the current state of knowledge acquired on missense variants (referred to as to single amino acid variants) affecting membrane proteins as well as the insights that can be extrapolated from data already available. This includes an overview of the annotations for membrane protein variants that have been collated within databases dedicated to the topic, bioinformatics approaches that leverage evolutionary information in order to shed light on previously uncharacterized membrane protein structures or interaction interfaces, tools for predicting the effects of mutations tailored specifically towards the characteristics of membrane proteins as well as two clinically relevant case studies explaining the implications of mutated membrane proteins in cancer and cardiomyopathy.
Jan Zaucha, Michael Heinzinger, A. Kulandaisamy, Evans Kataka, Óscar Llorian Salvádor, Petr Popov, Burkhard Rost, M. Michael Gromiha, Boris S. Zhorov, Dmitrij Frishman
Briefings Bioinform.8
2021 Guest Editorial for Special Section on the 15th International Conference on Intelligent Computing (ICIC)
De-Shuang Huang, Vitoantonio Bevilacqua, M. Michael Gromiha
IEEE ACM Trans. Comput. Biol. Bioinform.3
2020 ProAffiMuSeq: sequence-based method to predict the binding free energy change of protein-protein complexes upon mutation using functional classification
abstract
MOTIVATION: Protein-protein interactions are essential for the cell and mediate various functions. However, mutations can disrupt these interactions and may cause diseases. Currently available computational methods require a complex structure as input for predicting the change in binding affinity. Further, they have not included the functional class information for the protein-protein complex. To address this, we have developed a method, ProAffiMuSeq, which predicts the change in binding free energy using sequence-based features and functional class. RESULTS: Our method shows an average correlation between predicted and experimentally determined ΔΔG of 0.73 and mean absolute error (MAE) of 0.86 kcal/mol in 10-fold cross-validation and correlation of 0.75 with MAE of 0.94 kcal/mol in the test dataset. ProAffiMuSeq was also tested on an external validation set and showed results comparable to structure-based methods. Our method can be used for large-scale analysis of disease-causing mutations in protein-protein complexes without structural information. AVAILABILITY AND IMPLEMENTATION: Users can access the method at https://web.iitm.ac.in/bioinfo2/proaffimuseq/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Sherlyn Jemimah, Masakazu Sekijima, M. Michael Gromiha
Bioinform.3
2020 AggreRATE-Pred: a mathematical model for the prediction of change in aggregation rate upon point mutation
abstract
MOTIVATION: Protein aggregation is a major unsolved problem in biochemistry with implications for several human diseases, biotechnology and biomaterial sciences. A majority of sequence-structural properties known for their mechanistic roles in protein aggregation do not correlate well with the aggregation kinetics. This limits the practical utility of predictive algorithms. RESULTS: We analyzed experimental data on 183 unique single point mutations that lead to change in aggregation rates for 23 polypeptides and proteins. Our initial mathematical model obtained a correlation coefficient of 0.43 between predicted and experimental change in aggregation rate upon mutation (P-value <0.0001). However, when the dataset was classified based on protein length and conformation at the mutation sites, the average correlation coefficient almost doubled to 0.82 (range: 0.74-0.87; P-value <0.0001). We observed that distinct sequence and structure-based properties determine protein aggregation kinetics in each class. In conclusion, the protein aggregation kinetics are impacted by local factors and not by global ones, such as overall three-dimensional protein fold, or mechanistic factors such as the presence of aggregation-prone regions. AVAILABILITY AND IMPLEMENTATION: The web server is available at http://www.iitm.ac.in/bioinfo/aggrerate-pred/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Puneet Rawat, R. Prabakaran, M. Michael Gromiha
Bioinform.4
2020 ProCaff: protein-carbohydrate complex binding affinity database
abstract
MOTIVATION: Protein-carbohydrate interactions perform several cellular and biological functions and their structure and function are mainly dictated by their binding affinity. Although plenty of experimental data on binding affinity are available, there is no reliable and comprehensive database in the literature. RESULTS: We have developed a database on binding affinity of protein-carbohydrate complexes, ProCaff, which contains 3122 entries on dissociation constant (Kd), Gibbs free energy change (ΔG), experimental conditions, sequence, structure and literature information. Additional features include the options to search, display, visualization, download and upload the data. AVAILABILITY AND IMPLEMENTATION: The database is freely available at http://web.iitm.ac.in/bioinfo2/procaff/. The website is implemented using HTML and PHP and supports recent versions of major browsers such as Chrome, Firefox, IE10 and Opera. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
N. R. Siva Shanmugam, J. Jino Blessy, K. Veluraja, M. Michael Gromiha
Bioinform.4
2020 Guest Editorial for Special Section on the 14th International Conference on Intelligent Computing (ICIC)
abstract
The papers in this special section were presented at the Fourteenth International Conference on Intelligent Computing (ICIC) held in Wuhan, China, on August 15-18, 2018.
De-Shuang Huang, Vitoantonio Bevilacqua, M. Michael Gromiha
IEEE ACM Trans. Comput. Biol. Bioinform.3
2019 Seq2Feature: a comprehensive web-based feature extraction tool
abstract
MOTIVATION: Machine learning techniques require various descriptors from protein and nucleic acid sequences to understand/predict their structure and function as well as distinguishing between disease and neutral mutations. Hence, availability of a feature extraction tool is necessary to bridge the gap. RESULTS: We developed a comprehensive web-based tool, Seq2Feature, which computes 252 protein and 41 DNA sequence-based descriptors. These features include physicochemical, energetic and conformational properties of proteins, mutation matrices and contact potentials as well as nucleotide composition, physicochemical and conformational properties of DNA. We propose that Seq2Feature could serve as an effective tool for extracting protein and DNA sequence-based features as applicable inputs to machine learning algorithms. AVAILABILITY AND IMPLEMENTATION: https://www.iitm.ac.in/bioinfo/SBFE/index.html. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Rahul Nikam, M. Michael Gromiha
Bioinform.2
2019 Guest Editorial for Special Section on the 13th International Conference on Intelligent Computing (ICIC)
abstract
The papers presented in this special section were presented at the Thirteenth International Conference on Intelligent Computing (ICIC) that was held in Liverpool, UK, on August 7-10, 2017. ICIC was formed to provide an annual forum dedicated to the emerging and challenging topics in artificial intelligence, machine learning, bioinformatics, and computational biology, etc. It aims to bring together researchers and practitioners from both academia and industry to share ideas, problems, and solutions related to the multifaceted aspects of intelligent computing.
De-Shuang Huang, Vitoantonio Bevilacqua, M. Michael Gromiha
IEEE ACM Trans. Comput. Biol. Bioinform.3
2018 Discrimination and Prediction of Protein-Protein Binding Affinity Using Deep Learning Approach
Rahul Nikam, K. Yugandhar, M. Michael Gromiha
ICIC (2)3
2018 MutHTP: mutations in human transmembrane proteins
abstract
Motivation: Existing sources of experimental mutation data do not consider the structural environment of amino acid substitutions and distinguish between soluble and membrane proteins. They also suffer from a number of further limitations, including data redundancy, lack of disease classification, incompatible information content, and ambiguous annotations (e.g. the same mutation being annotated as disease and benign). Results: We have developed a novel database, MutHTP, which contains information on 183 395 disease-associated and 17 827 neutral mutations in human transmembrane proteins. For each mutation site MutHTP provides a description of its location with respect to the membrane protein topology, structural environment (if available) and functional features. Comprehensive visualization, search, display and download options are available. Availability and implementation: The database is publicly available at http://www.iitm.ac.in/bioinfo/MutHTP/. The website is implemented using HTML, PHP and javascript and supports recent versions of all major browsers, such as Firefox, Chrome and Opera. Supplementary information: Supplementary data are available at Bioinformatics online.
A. Kulandaisamy, S. Binny Priya, Svetlana Tarnovskaya, Ilya Bizin, Peter Hönigschmid, Dmitrij Frishman, M. Michael Gromiha
Bioinform.8
2018 Guest Editorial for Special Section on the 12th International Conference on Intelligent Computing (ICIC)
abstract
The eight papers included in this special section were presented at the 12th International Conference on Intelligent Computing (ICIC) held at Lanzhou, China, during August 2-5, 2016. ICIC was formed to provide an annual forum dedicated to the emerging and challenging topics in artificial intelligence, machine learning, bioinformatics, and computational biology, etc. It aims to bring together researchers and practitioners from both academia and industry to share ideas, problems, and solutions related to the multifaceted aspects of intelligent computing.
De-Shuang Huang, Vitoantonio Bevilacqua, M. Michael Gromiha
IEEE ACM Trans. Comput. Biol. Bioinform.3
2018 Identification and Analysis of Key Residues in Protein-RNA Complexes
abstract
Protein-RNA complexes play important roles in various biological processes. The functions of protein-RNA complexes are dictated by their interactions, binding, stability, and affinity. In this work, we have identified the key residues (KRs), which are involved in both stability and binding. We found that 42 percent of considered proteins share common binding and stabilizing residues, whereas these residues are distinct in 58 percent of the proteins. Overall, 5 percent of stabilizing and 3 percent of binding residues serve as key residues. These residues are enriched with the combination of polar, charged, aliphatic, and aromatic residues. Analysis on subclasses of protein-RNA complexes based on protein structural class, function and RNA type showed that regulatory proteins, and complexes with single stranded RNA and rRNA have appreciable number of key residues. Specifically, Arg, Tyr, and Thr are preferred in most of the subclasses of protein-RNA complexes. In addition, residues with similar chemical behavior have different preferences to be KRs, such that Arg, Tyr, Val, and Thr are preferred over Lys, Trp, Ile, and Ser, respectively. Atomic level contacts revealed that charged and polar-nonpolar contacts are dominant in enzymes, polar in structural, and nonpolar in regulatory proteins. On the other hand, polar-nonpolar contacts are enriched in all these classes of protein-RNA complexes. Further, the influence of sequence and structural features such as conservation score, surrounding hydrophobicity, solvent accessibility, secondary structure, and long-range order in key residues are also discussed. We envisage that the present study provides insights to understand the structural and functional aspects of protein-RNA complexes.
A. Kulandaisamy, Ambuj Srivastava, Raju Nagarajan, S. Binny Priya, M. Michael Gromiha
IEEE ACM Trans. Comput. Biol. Bioinform.6
2017 Influence of Amino Acid Properties for Characterizing Amyloid Peptides in Human Proteome
R. Prabakaran, Rahul Nikam, M. Michael Gromiha
ICIC (2)4
2017 PROXiMATE: a database of mutant protein-protein complex thermodynamics and kinetics
abstract
SUMMARY: We have developed PROXiMATE, a database of thermodynamic data for more than 6000 missense mutations in 174 heterodimeric protein-protein complexes, supplemented with interaction network data from STRING database, solvent accessibility, sequence, structural and functional information, experimental conditions and literature information. Additional features include complex structure visualization, search and display options, download options and a provision for users to upload their data. AVAILABILITY AND IMPLEMENTATION: The database is freely available at http://www.iitm.ac.in/bioinfo/PROXiMATE/ . The website is implemented in Python, and supports recent versions of major browsers such as IE10, Firefox, Chrome and Opera. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Sherlyn Jemimah, K. Yugandhar, M. Michael Gromiha
Bioinform.3
2017 Guest Editorial for Special Section on the 11th International Conference on Intelligent Computing (ICIC)
abstract
The papers in this special section were presented at the 11th International Conference on Intelligent Computing (ICIC) held in Fuzhou, China, on August 20-23, 2015. This conference was formed to provide an annual forum dedicated to the emerging and challenging topics in artificial intelligence, machine learning, bioinformatics, etc. It aims to bring together researchers and practitioners from both academia and industry to share ideas, problems, and solutions related to the multifaceted aspects of intelligent computing.
De-Shuang Huang, Vitoantonio Bevilacqua, M. Michael Gromiha
IEEE ACM Trans. Comput. Biol. Bioinform.3
2016 Computational Analysis of Similar Protein-DNA Complexes from Different Organisms to Understand Organism Specific Recognition
Raju Nagarajan, M. Michael Gromiha
ICIC (2)2
2016 Highlights from the 11th ISCB Student Council Symposium 2015: Dublin, Ireland. 10 July 2015
abstract
Table of contents A1 Highlights from the eleventh ISCB Student Council Symposium 2015 Katie Wilkins, Mehedi Hassan, Margherita Francescatto, Jakob Jespersen, R. Gonzalo Parra, Bart Cuypers, Dan DeBlasio, Alexander Junge, Anupama Jigisha, Farzana Rahman O1 Prioritizing a drug’s targets using both gene expression and structural similarity Griet Laenen, Sander Willems, Lieven Thorrez, Yves Moreau O2 Organism specific protein-RNA recognition: A computational analysis of protein-RNA complex structures from different organisms Nagarajan Raju, Sonia Pankaj Chothani, C. Ramakrishnan, Masakazu Sekijima; M. Michael Gromiha O3 Detection of Heterogeneity in Single Particle Tracking Trajectories Paddy J Slator, Nigel J Burroughs O4 3D-NOME: 3D NucleOme Multiscale Engine for data-driven modeling of three-dimensional genome architecture Przemysław Szałaj, Zhonghui Tang, Paul Michalski, Oskar Luo, Xingwang Li, Yijun Ruan, Dariusz Plewczynski O5 A novel feature selection method to extract multiple adjacent solutions for viral genomic sequences classification Giulia Fiscon, Emanuel Weitschek, Massimo Ciccozzi, Paola Bertolazzi, Giovanni Felici O6 A Systems Biology Compendium for Leishmania donovani Bart Cuypers, Pieter Meysman, Manu Vanaerschot, Maya Berg, Hideo Imamura, Jean-Claude Dujardin, Kris Laukens O7 Unravelling signal coordination from large scale phosphorylation kinetic data Westa Domanova, James R. Krycer, Rima Chaudhuri, Pengyi Yang, Fatemeh Vafaee, Daniel J. Fazakerley, Sean J. Humphrey, David E. James, Zdenka Kuncic
Katie Wilkins, Mehedi Hassan, Margherita Francescatto, Jakob B. Jespersen, R. Gonzalo Parra, Bart Cuypers, Dan F. DeBlasio, Alexander Junge, Anupama Jigisha, Farzana Rahman, Griet Laenen, Sander Willems, Lieven Thorrez, Yves Moreau, Raju Nagarajan, Sonia P. Chothani, C. Ramakrishnan, Masakazu Sekijima, M. Michael Gromiha, Paddy Slator, Nigel J. Burroughs, Przemyslaw Szalaj, Zhonghui Tang, Paul J. Michalski, Oskar Luo, Xingwang Li 0004, Yijun Ruan, Dariusz Plewczynski, Giulia Fiscon, Emanuel Weitschek, Massimo Ciccozzi, Paola Bertolazzi, Giovanni Felici, Pieter Meysman, Manu Vanaerschot, Maya Berg, Hideo Imamura, Jean-Claude Dujardin, Kris Laukens, Westa Domanova, James R. Krycer, Rima Chaudhuri, Pengyi Yang, Fatemeh Vafaee, Daniel J. Fazakerley, Sean J. Humphrey, David E. James, Zdenka Kuncic
BMC Bioinform.19
2016 Guest Editorial for Special Section on the 10th International Conference on Intelligent Computing (ICIC)
abstract
This special section includes a selection of eight papers presented at the 10th International Conference on Intelligent Computing (ICIC) held in Taiyuan, China, on August 3–6, 2014.
De-Shuang Huang, Vitoantonio Bevilacqua, M. Michael Gromiha
IEEE ACM Trans. Comput. Biol. Bioinform.3
2015 Folding RaCe: a robust method for predicting changes in protein folding rates upon point mutations
abstract
MOTIVATION: Protein engineering methods are commonly employed to decipher the folding mechanism of proteins and enzymes. However, such experiments are exceedingly time and resource intensive. It would therefore be advantageous to develop a simple computational tool to predict changes in folding rates upon mutations. Such a method should be able to rapidly provide the sequence position and chemical nature to modulate through mutation, to effect a particular change in rate. This can be of importance in protein folding, function or mechanistic studies. RESULTS: We have developed a robust knowledge-based methodology to predict the changes in folding rates upon mutations formulated from amino and acid properties using multiple linear regression approach. We benchmarked this method against an experimental database of 790 point mutations from 26 two-state proteins. Mutants were first classified according to secondary structure, accessible surface area and position along the primary sequence. Three prime amino acid features eliciting the best relationship with folding rates change were then shortlisted for each class along with an optimized window length. We obtained a self-consistent mean absolute error of 0.36 s(-1) and a mean Pearson correlation coefficient (PCC) of 0.81. Jack-knife test resulted in a MAE of 0.42 s(-1) and a PCC of 0.73. Moreover, our method highlights the importance of outlier(s) detection and studying their implications in the folding mechanism. AVAILABILITY AND IMPLEMENTATION: A web server 'Folding RaCe' has been developed and is available at http://www.iitm.ac.in/bioinfo/proteinfolding/foldingrace.html. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Priyashree Chaudhary, Athi N. Naganathan, M. Michael Gromiha
Bioinform.3
2015 Response to the comment on 'protein-protein binding affinity prediction from amino acid sequence'
abstract
Abstract Contact: [email protected]
K. Yugandhar, M. Michael Gromiha
Bioinform.2
2014 Identification of Novel c-Yes Kinase Inhibitors
C. Ramakrishnan, A. Mary Thangakani, Devadasan Velmurugan, M. Michael Gromiha
ICIC (3)4
2014 Bioinformatics approaches for functional annotation of membrane proteins
abstract
Membrane proteins perform diverse functions in living organisms such as transporters, receptors and channels. The functions of membrane proteins have been investigated with several computational approaches, such as developing databases, analyzing the structure-function relationship and establishing algorithms to discriminate different type of membrane proteins. However, compilation of bioinformatics resources for the functions of membrane proteins is not well documented compared with their structural aspects. In this comprehensive review, we elaborately focus on three aspects of membrane protein functions: (i) databases for different types of membrane proteins based on their functions including transporters, receptors and ion channels, annotated functional data for genomes, as well as functionally important amino acid residues in membrane proteins obtained from experimental data, (ii) analysis of membrane protein functions based on their structures, motifs, amino acid properties and other features and (iii) algorithms for discriminating different types of membrane proteins and annotating them in genomic sequences. In addition, we provide a list of online resources for the databases and web servers for functional annotation of membrane proteins.
M. Michael Gromiha, Yu-Yen Ou
Briefings Bioinform.1
2014 GAP: towards almost 100 percent prediction for β-strand-mediated aggregating peptides with distinct morphologies
abstract
MOTIVATION: Distinguishing between amyloid fibril-forming and amorphous β-aggregating aggregation-prone regions (APRs) in proteins and peptides is crucial for designing novel biomaterials and improved aggregation inhibitors for biotechnological and therapeutic purposes. RESULTS: Adjacent and alternate position residue pairs in hexapeptides show distinct preferences for occurrence in amyloid fibrils and amorphous β-aggregates. These observations were converted into energy potentials that were, in turn, machine learned. The resulting tool, called Generalized Aggregation Proneness (GAP), could successfully distinguish between amyloid fibril-forming and amorphous β-aggregating hexapeptides with almost 100 percent accuracies in validation tests performed using non-redundant datasets. CONCLUSION: Accuracies of the predictions made by GAP are significantly improved compared with other methods capable of predicting either general β-aggregation or amyloid fibril-forming APRs. This work demonstrates that amino acid side chains play important roles in determining the morphological fate of β-mediated aggregates formed by short peptides. AVAILABILITY AND IMPLEMENTATION: http://www.iitm.ac.in/bioinfo/GAP/.
A. Mary Thangakani, Raju Nagarajan, Devadasan Velmurugan, M. Michael Gromiha
Bioinform.5
2014 Protein-protein binding affinity prediction from amino acid sequence
abstract
MOTIVATION: Protein-protein interactions play crucial roles in many biological processes and are responsible for smooth functioning of the machinery in living organisms. Predicting the binding affinity of protein-protein complexes provides deep insights to understand the recognition mechanism and identify the strong binding partners in protein-protein interaction networks. RESULTS: In this work, we have collected the experimental binding affinity data for a set of 135 protein-protein complexes and analyzed the relationship between binding affinity and 642 properties obtained from amino acid sequence. We noticed that the overall correlation is poor, and the factors influencing affinity depends on the type of the complex based on their function, molecular weight and binding site residues. Based on the results, we have developed a novel methodology for predicting the binding affinity of protein-protein complexes using sequence-based features by classifying the complexes with respect to their function and predicted percentage of binding site residues. We have developed regression models for the complexes belonging to different classes with three to five properties, which showed a correlation in the range of 0.739-0.992 using jack-knife test. We suggest that our approach adds a new aspect of biological significance in terms of classifying the protein-protein complexes for affinity prediction.
K. Yugandhar, M. Michael Gromiha
Bioinform.2
2013 Role of Protein Aggregation and Interactions between α-Synuclein and Calbindin in Parkinson's Disease
M. Michael Gromiha, S. Biswal, A. Mary Thangakani, G. J. Masilamoni, Devadasan Velmurugan
ICIC (2)1
2013 Distinct position-specific sequence features of hexa-peptides that form amyloid-fibrils: application to discriminate between amyloid fibril and amorphous β-aggregate forming peptide sequences
abstract
BACKGROUND: Comparison of short peptides which form amyloid-fibrils with their homologues that may form amorphous β-aggregates but not fibrils, can aid development of novel amyloid-containing nanomaterials with well defined morphologies and characteristics. The knowledge gained from the comparative analysis could also be applied towards identifying potential aggregation prone regions in proteins, which are important for biotechnology applications or have been implicated in neurodegenerative diseases. In this work we have systematically analyzed a set of 139 amyloid-fibril hexa-peptides along with a highly homologous set of 168 hexa-peptides that do not form amyloid fibrils for their position-wise as well as overall amino acid compositions and averages of 49 selected amino acid properties. RESULTS: Amyloid-fibril forming peptides show distinct preferences and avoidances for amino acid residues to occur at each of the six positions. As expected, the amyloid fibril peptides are also more hydrophobic than non-amyloid peptides. We have used the results of this analysis to develop statistical potential energy values for the 20 amino acid residues to occur at each of the six different positions in the hexa-peptides. The distribution of the potential energy values in 139 amyloid and 168 non-amyloid fibrils are distinct and the amyloid-fibril peptides tend to be more stable (lower total potential energy values) than non-amyloid peptides. The average frequency of occurrence of these peptides with lower than specific cutoff energies at different positions is 72% and 50%, respectively. The potential energy values were used to devise a statistical discriminator to distinguish between amyloid-fibril and non-amyloid peptides. Our method could identify the amyloid-fibril forming hexa-peptides to an accuracy of 89%. On the other hand, the accuracy of identifying non-amyloid peptides was only 54%. Further attempts were made to improve the prediction accuracy via machine learning. This resulted in an overall accuracy of 82.7% with the sensitivity and specificity of 81.3% and 83.9%, respectively, in 10-fold cross-validation method. CONCLUSIONS: Amyloid-fibril forming hexa-peptides show position specific sequence features that are different from those which may form amorphous β-aggregates. These positional preferences are found to be important features for discriminating amyloid-fibril forming peptides from their homologues that don't form amyloid-fibrils.
A. Mary Thangakani, Devadasan Velmurugan, M. Michael Gromiha
BMC Bioinform.4
2012 Sequence Analysis and Discrimination of Amyloid and Non-amyloid Peptides
M. Michael Gromiha, A. Mary Thangakani, Devadasan Velmurugan
ICIC (3)1
2012 Introduction: advanced intelligent computing theories and their applications in bioinformatics
abstract
The advancement of techniques in computer science and information technology witnessed the rapid growth of bioinformatics in various diverse areas such as sequence alignment, structure prediction, structure-function relationship, protein interactions, genome annotation, gene expression, microarray data analysis and so on. It is necessary and pertinent to discuss the issues on these topics and analyze the latest developments. The International Conference on Intelligent Computing (ICIC) provided a forum for discussing the recent investigations on bioinformatics related problems using high performance computing and efficient algorithms. Among the 832 submissions 33.7% were selected for presentations at ICIC 2011. Based on the novelty of the manuscripts, presentations and originality only 12 papers were selected as 'high quality', and the extended versions of them are included in the supplement. The supplement is broadly classified into five categories, structure-function relationship of proteins, protein-protein interactions, gene expression/interaction networks, microarray data analysis and visualization tools. The opening article by Gromiha et al. [1] related various physical, chemical, energetic and conformational properties of amino acid residues with the change of half maximal effective concentration (EC50) due to amino acid substitutions in olfactory receptors. Further, they utilized machine learning methods for discriminating the mutants, which enhance or reduce EC50 values upon mutation. Wang et al. [2] proposed a protein-protein dissimilarity learning algorithm for comparing protein structures using the contextual information of proteins. Lei et al. [3] developed a robust computational technique for assessing the reliability of protein-protein interactions and predicting the interacting pairs of proteins by integrating manifold embedding with various features. Wang et al. [4] described an algorithm for identifying overlapping modules in protein-protein interaction networks. Cui et al. [5] built a support vector machine model for predicting human proteins that interact with virus proteins and specifically human papillomavirus and hepatitis C virus. Liu et al. [6] constructed an integrated map of protein interaction network in Mycobacterium tuberculosis using machine learning and ortholog-based methods. Wang et al. [7] presented a network biology approach for investigating drug combinations and their target proteins in the context of genetic interaction networks and related human pathways with an aim to understand the underlying rules of effective drug combinations. Hsiao et al. [8] proposed an incremental evolutionary approach using network robustness for inferring gene regulatory networks with an application to deal with a large number of network parameters. Bevilacqua et al. [9] explored the issue of microarray data merging and used distant metastasis prediction for classifying three different sets of breast cancer data. Park et al. [10] analyzed the whole brain microarray data and physical connectivity of hippocampus with other brain regions to identify the genes related to Alzheimer's disease and their interactions with proteins. Ayadi et al. [11] described a stochastic pattern-driven neighborhood search algorithm for biclustering microarray data. In the last article, Jung et al. [12] described the development of a JAVA based stand-alone program for detecting and visualizing of genomic variants, which enables the manual exclusion of erroneous signals. It is also capable of visualizing genomic data from different sources such as data from comparative genomic hybridization arrays and sequence alignment format files. The guest editors of the supplement would like to thank the Executive Editor of BMC Bioinformatics Professor Kate Rice for providing an opportunity to publish some of the excellent papers presented in ICIC 2011. We also wish to thank Ms. Isobel Peters and Ms. Catherine Wells for their help and support in editing the supplement. Finally, our sincere thanks to all the authors of the papers selected for publication in this issue. This work was partially supported by the grants of the National Science Foundation of China, Nos. 61133010, & 31071168.
M. Michael Gromiha, De-Shuang Huang
BMC Bioinform.1
2012 Relationship between amino acid properties and functional parameters in olfactory receptors and discrimination of mutants with enhanced specificity
abstract
BACKGROUND: Olfactory receptors are key components in signal transduction. Mutations in olfactory receptors alter the odor response, which is a fundamental response of organisms to their immediate environment. Understanding the relationship between odorant response and mutations in olfactory receptors is an important problem in bioinformatics and computational biology. In this work, we have systematically analyzed the relationship between various physical, chemical, energetic and conformational properties of amino acid residues, and the change of odor response/compound's potency/half maximal effective concentration (EC50) due to amino acid substitutions. RESULTS: We observed that both the characteristics of odorant molecule (ligand) and amino acid properties are important for odor response and EC50. Additional information on neighboring and surrounding residues of the mutants enhanced the correlation between amino acid properties and EC50. Further, amino acid properties have been combined systematically using multiple regression techniques and we obtained a correlation of 0.90-0.98 with odor response/EC50 of goldfish, mouse and human olfactory receptors. In addition, we have utilized machine learning methods to discriminate the mutants, which enhance or reduce EC50 values upon mutation and we obtained an accuracy of 93% and 79% for self-consistency and jack-knife tests, respectively. CONCLUSIONS: Our analysis provides deep insights for understanding the odor response of olfactory receptor mutants and the present method could be used for identifying the mutants with enhanced specificity.
M. Michael Gromiha, K. Harini, R. Sowdhamini, Kazuhiko Fukui
BMC Bioinform.1
2011 Structure-Function Relationship in Olfactory Receptors
M. Michael Gromiha, R. Sowdhamini, Kazuhiko Fukui
ICIC (3)1
2011 Prediction of transporter targets using efficient RBF networks with PSSM profiles and biochemical properties
abstract
SUMMARY: Transporters are proteins that are involved in the movement of ions or molecules across biological membranes. Currently, our knowledge about the functions of transporters is limited due to the paucity of their 3D structures. Hence, computational techniques are necessary to annotate the functions of transporters. In this work, we focused on an important functional aspect of transporters, namely annotation of targets for transport proteins. We have systematically analyzed four major classes of transporters with different transporter targets: (i) electron, (ii) protein/mRNA, (iii) ion and (iv) others, using amino acid properties. We have developed a radial basis function network-based method for predicting transport targets with amino acid properties and position specific scoring matrix profiles. Our method showed a 10-fold cross-validation accuracy of 90.1, 80.1, 70.3 and 82.3% for electron transporters, protein/mRNA transporters, ion transporters and others, respectively, in a dataset of 543 transporters. We have also evaluated the performance of the method with an independent dataset of 108 proteins and we obtained similar accuracy. We suggest that our method could be an effective tool for functional annotation of transport proteins. AVAILABILITY: http://rbf.bioinfo.tw/~sachen/ttrbf.html
Shu-An Chen, Yu-Yen Ou, Tzong-Yi Lee, M. Michael Gromiha
Bioinform.4
2010 Sequence and structural features of binding site residues in protein-protein complexes
abstract
We have developed an energy based approach for identifying the binding site residues in protein-protein complexes. The binding site residues have been analyzed with sequence and structure based parameters such as neighboring residues in the vicinity of binding sites and conformational switching. We observed specific preferences of dipeptides and tripeptides for binding, which is unique to protein-protein complexes. Our analysis showed that 7% of residues changed their conformations upon protein-protein complex formation and it is 9.2% and 6.6% in the binding and non-binding sites, respectively. Specifically, the residues Glu, Lys, Leu and Ser changed their conformation from coil to helix/strand and from helix to coil/strand. Leu, Ser, Thr and Val prefer to change their conformation from strand to coil/helix. The results obtained in this study will be helpful for understanding and predicting the binding sites in protein-protein complexes.
M. Michael Gromiha, Samuel Selvaraj, B. Jayaram 0001, Kazuhiko Fukui
BIBM1
2010 Topology Prediction of alpha-Helical and beta-Barrel Transmembrane Proteins Using RBF Networks
Shu-An Chen, Yu-Yen Ou, M. Michael Gromiha
ICIC (1)3
2010 Identification and Analysis of Binding Site Residues in Protein Complexes: Energy Based Approach
M. Michael Gromiha, Samuel Selvaraj, B. Jayaram 0001, Kazuhiko Fukui
ICIC (1)1
2010 First insight into the prediction of protein folding rate change upon point mutation
abstract
SUMMARY: The accurate prediction of protein folding rate change upon mutation is an important and challenging problem in protein folding kinetics and design. In this work, we have collected experimental data on protein folding rate change upon mutation from various sources and constructed a reliable and non-redundant dataset with 467 mutants. These mutants are widely distributed based on secondary structure, solvent accessibility, conservation score and long-range contacts. From systematic analysis of these parameters along with a set of 49 amino acid properties, we have selected a set of 12 features for discriminating the mutants that speed up or slow down the folding process. We have developed a method based on quadratic regression models for discriminating the accelerating and decelerating mutants, which showed an accuracy of 74% using the 10-fold cross-validation test. The sensitivity and specificity are 63% and 76%, respectively. The method can be improved with the inclusion of physical interactions and structure-based parameters. AVAILABILITY: http://bioinformatics.myweb.hinet.net/freedom.htm.
Liang-Tsung Huang, M. Michael Gromiha
Bioinform.2
2010 Development of knowledge-based system for predicting the stability of proteins upon point mutations
Liang-Tsung Huang, Lien Fu Lai, Chao-Chin Wu, M. Michael Gromiha
Neurocomputing4
2010 Human-Readable Rule Generator for Integrating Amino Acid Sequence Information and Stability of Mutant Proteins
abstract
Most of the bioinformatics tools developed for predicting mutant protein stability appear as a black box and the relationship between amino acid sequence/structure and stability is hidden to the users. We have addressed this problem and developed a human-readable rule generator for integrating the knowledge of amino acid sequence and experimental stability change upon single mutation. Using information about the original residue, substituted residue, and three neighboring residues, classification rules have been generated to discriminate the stabilizing and destabilizing mutants and explore the basis for experimental data. These rules are human readable, and hence, the method enhances the synergy between expert knowledge and computational system. Furthermore, the performance of the rules has been assessed on a nonredundant data set of 1,859 mutants and we obtained an accuracy of 80 percent using cross validation. The results showed that the method could be effectively used as a tool for both knowledge discovery and predicting mutant protein stability. We have developed a Web for classification rule generator and it is freely available at http://bioinformatics.myweb.hinet.net/irobot.htm.
Liang-Tsung Huang, Lien Fu Lai, M. Michael Gromiha
IEEE ACM Trans. Comput. Biol. Bioinform.3
2009 Reliable prediction of protein thermostability change upon double mutation from amino acid sequence
abstract
Abstract Summary: The accurate prediction of protein stability change upon mutation is one of the important issues for protein design. In this work, we have focused on the stability change of double mutations and systematically analyzed the wild-type and mutant residues, patterns in amino acid sequence and locations of mutants. Based on the sequence information of wild-type, mutant and three neighboring residues, we have presented a weighted decision table method (WET) for predicting the stability changes of 180 double mutants obtained from thermal (ΔΔG) denaturation. Using 10-fold cross-validation test, our method showed a correlation of 0.75 between experimental and predicted values of stability changes, and an accuracy of 82.2% for discriminating the stabilizing and destabilizing mutants. Availability: http://bioinformatics.myweb.hinet.net/wetstab.htm Contact: [email protected] Supplementary information: Supplementary data are available at Bioinformatics online.
Liang-Tsung Huang, M. Michael Gromiha
Bioinform.2
2008 Neural network based prediction of protein structure and Function: Comparison with other machine learning methods
abstract
We have utilized neural networks in different applications of bioinformatics such as discrimination of beta-barrel membrane proteins, mesophilic and thermophilic proteins, different folding types of globular proteins, different classes of transporter proteins and predicting the secondary structures of beta-barrel membrane proteins. In these methods, we have used the information about amino acid composition, neighboring residue information, inter-residue contacts and amino acid properties as features. We observed that the performance with neural networks is comparable to or better than other widely used machine learning techniques.
M. Michael Gromiha, Shandar Ahmad, Makiko Suwa
IJCNN1
2008 Gene Ontology term prediction based upon amino acid occurrence
abstract
Usually prediction of molecular functions of proteins from their amino acid sequences is based upon sequence similarity with proteins of known functions. However, it is well known that function is mainly dependent upon protein structures than sequences. Since structures are often independent of sequences, it is important to predict function without sequence similarities. Here we propose a method based upon amino acid occurrence for predicting Gene Ontology (GO) term. We have tested the method in a set of 3212 proteins in Protein Data Bank with less than 40% sequence identity. Our method achieved more than 50% sensitivity and 20% precision for c.a. 20 selected GO terms among the most frequent 557 GO terms. Mean sensitivity, specificity, precision, and accuracy for relatively rare (but majority) 402 GO terms among the 557 GO terms are 13%, 99%, 9% and 99%, respectively. They are significantly larger than expected values of less than 2% under assuming random selection.
Y-h. Taguchi 0001, M. Michael Gromiha
IJCNN2
2008 Functional discrimination of membrane proteins using machine learning techniques
abstract
BACKGROUND: Discriminating membrane proteins based on their functions is an important task in genome annotation. In this work, we have analyzed the characteristic features of amino acid residues in membrane proteins that perform major functions, such as channels/pores, electrochemical potential-driven transporters and primary active transporters. RESULTS: We observed that the residues Asp, Asn and Tyr are dominant in channels/pores whereas the composition of hydrophobic residues, Phe, Gly, Ile, Leu and Val is high in electrochemical potential-driven transporters. The composition of all the amino acids in primary active transporters lies in between other two classes of proteins. We have utilized different machine learning algorithms, such as, Bayes rule, Logistic function, Neural network, Support vector machine, Decision tree etc. for discriminating these classes of proteins. We observed that most of the algorithms have discriminated them with similar accuracy. The neural network method discriminated the channels/pores, electrochemical potential-driven transporters and active transporters with the 5-fold cross validation accuracy of 64% in a data set of 1718 membrane proteins. The application of amino acid occurrence improved the overall accuracy to 68%. In addition, we have discriminated transporters from other alpha-helical and beta-barrel membrane proteins with the accuracy of 85% using k-nearest neighbor method. The classification of transporters and all other proteins (globular and membrane) showed the accuracy of 82%. CONCLUSION: The performance of discrimination with amino acid occurrence is better than that with amino acid composition. We suggest that this method could be effectively used to discriminate transporters from all other globular and membrane proteins, and classify them into channels/pores, electrochemical and active transporters.
M. Michael Gromiha, Yukimitsu Yabuki
BMC Bioinform.1
2007 iPTREE-STAB: interpretable decision tree based method for predicting protein stability changes upon mutations
abstract
UNLABELLED: We have developed a web server, iPTREE-STAB for discriminating the stability of proteins (stabilizing or destabilizing) and predicting their stability changes (delta deltaG) upon single amino acid substitutions from amino acid sequence. The discrimination and prediction are mainly based on decision tree coupled with adaptive boosting algorithm, and classification and regression tree, respectively, using three neighboring residues of the mutant site along N- and C-terminals. Our method showed an accuracy of 82% for discriminating the stabilizing and destabilizing mutants, and a correlation of 0.70 for predicting protein stability changes upon mutations. AVAILABILITY: http://bioinformatics.myweb.hinet.net/iptree.htm. SUPPLEMENTARY INFORMATION: Dataset and other details are given.
Liang-Tsung Huang, M. Michael Gromiha, Shinn-Ying Ho
Bioinform.2
2007 Identification of DNA-binding proteins using support vector machines and evolutionary profiles
abstract
BACKGROUND: Identification of DNA-binding proteins is one of the major challenges in the field of genome annotation, as these proteins play a crucial role in gene-regulation. In this paper, we developed various SVM modules for predicting DNA-binding domains and proteins. All models were trained and tested on multiple datasets of non-redundant proteins. RESULTS: SVM models have been developed on DNAaset, which consists of 1153 DNA-binding and equal number of non DNA-binding proteins, and achieved the maximum accuracy of 72.42% and 71.59% using amino acid and dipeptide compositions, respectively. The performance of SVM model improved from 72.42% to 74.22%, when evolutionary information in form of PSSM profiles was used as input instead of amino acid composition. In addition, SVM models have been developed on DNAset, which consists of 146 DNA-binding and 250 non-binding chains/domains, and achieved the maximum accuracy of 79.80% and 86.62% using amino acid composition and PSSM profiles. The SVM models developed in this study perform better than existing methods on a blind dataset. CONCLUSION: A highly accurate method has been developed for predicting DNA-binding proteins using SVM and PSSM profiles. This is the first study in which evolutionary information in form of PSSM profiles has been used successfully for predicting DNA-binding proteins. A web-server DNAbinder has been developed for identifying DNA-binding proteins and domains from query amino acid sequences http://www.imtech.res.in/raghava/dnabinder/.
Manish Kumar 0005, M. Michael Gromiha, Gajendra P. S. Raghava
BMC Bioinform.2
2007 Application of amino acid occurrence for discriminating different folding types of globular proteins
abstract
BACKGROUND: Predicting the three-dimensional structure of a protein from its amino acid sequence is a long-standing goal in computational/molecular biology. The discrimination of different structural classes and folding types are intermediate steps in protein structure prediction. RESULTS: In this work, we have proposed a method based on linear discriminant analysis (LDA) for discriminating 30 different folding types of globular proteins using amino acid occurrence. Our method was tested with a non-redundant set of 1612 proteins and it discriminated them with the accuracy of 38%, which is comparable to or better than other methods in the literature. A web server has been developed for discriminating the folding type of a query protein from its amino acid sequence and it is available at http://granular.com/PROLDA/. CONCLUSION: Amino acid occurrence has been successfully used to discriminate different folding types of globular proteins. The discrimination accuracy obtained with amino acid occurrence is better than that obtained with amino acid composition and/or amino acid properties. In addition, the method is very fast to obtain the results.
Y-h. Taguchi 0001, M. Michael Gromiha
BMC Bioinform.2
2005 A simple statistical method for discriminating outer membrane proteins with better accuracy
abstract
MOTIVATION: Discriminating outer membrane proteins from other folding types of globular and membrane proteins is an important task both for identifying outer membrane proteins from genomic sequences and for the successful prediction of their secondary and tertiary structures. RESULTS: We have systematically analyzed the amino acid composition of globular proteins from different structural classes and outer membrane proteins. We found that the residues, Glu, His, Ile, Cys, Gln, Asn and Ser, show a significant difference between globular and outer membrane proteins. Based on this information, we have devised a statistical method for discriminating outer membrane proteins from other globular and membrane proteins. Our approach correctly picked up the outer membrane proteins with an accuracy of 89% for the training set of 337 proteins. On the other hand, our method has correctly excluded the globular proteins at an accuracy of 79% in a non-redundant dataset of 674 proteins. Furthermore, the present method is able to correctly exclude alpha-helical membrane proteins up to an accuracy of 80%. These accuracy levels are comparable to other methods in the literature, and this is a simple method, which could be used for dissecting outer membrane proteins from genomic sequences. The influence of protein size, structural class and specific residues for discrimination is discussed.
M. Michael Gromiha, Makiko Suwa
Bioinform.1
2005 Discrimination of outer membrane proteins using support vector machines
abstract
MOTIVATION: Discriminating outer membrane proteins from other folding types of globular and membrane proteins is an important task both for dissecting outer membrane proteins (OMPs) from genomic sequences and for the successful prediction of their secondary and tertiary structures. RESULTS: We have developed a method based on support vector machines using amino acid composition and residue pair information. Our approach with amino acid composition has correctly predicted the OMPs with a cross-validated accuracy of 94% in a set of 208 proteins. Further, this method has successfully excluded 633 of 673 globular proteins and 191 of 206 alpha-helical membrane proteins. We obtained an overall accuracy of 92% for correctly picking up the OMPs from a dataset of 1087 proteins belonging to all different types of globular and membrane proteins. Furthermore, residue pair information improved the accuracy from 92 to 94%. This accuracy of discriminating OMPs is higher than that of other methods in the literature, which could be used for dissecting OMPs from genomic sequences. AVAILABILITY: Discrimination results are available at http://tmbeta-svm.cbrc.jp.
Keun-Joon Park, M. Michael Gromiha, Paul Horton, Makiko Suwa
Bioinform.2
2004 Structure-Function Relationship in DNA Sequence Recognition by Transcription Factors
Akinori Sarai, Samuel Selvaraj, M. Michael Gromiha, Hidetoshi Kono
APBC3
2004 Analysis and prediction of DNA-binding proteins and their binding residues based on composition, sequence and structural information
abstract
MOTIVATION: Though vitally important to cell function, the mechanism of protein-DNA binding has not yet been completely understood. We therefore analysed the relationship between DNA binding and protein sequence composition, solvent accessibility and secondary structure. Using non-redundant databases of transcription factors and protein-DNA complexes, neural network models were developed to utilize the information present in this relationship to predict DNA-binding proteins and their binding residues. RESULTS: Sequence composition was found to provide sufficient information to predict the probability of its binding to DNA with nearly 69% sensitivity at 64% accuracy for the considered proteins; sequence neighbourhood and solvent accessibility information were sufficient to make binding site predictions with 40% sensitivity at 79% accuracy. Detailed analysis of binding residues shows that some three- and five-residue segments frequently bind to DNA and that solvent accessibility plays a major role in binding. Although, binding behaviour was not associated with any particular secondary structure, there were interesting exceptions at the residue level. Over-representation of some residues in the binding sites was largely lost at the total sequence level, but a different kind of compositional preference was observed in DNA-binding proteins.
Shandar Ahmad, M. Michael Gromiha, Akinori Sarai
Bioinform.2
2004 ASAView: Database and tool for solvent accessibility representation in proteins
abstract
BACKGROUND: Accessible surface area (ASA) or solvent accessibility of amino acids in a protein has important implications. Knowledge of surface residues helps in locating potential candidates of active sites. Therefore, a method to quickly see the surface residues in a two dimensional model would help to immediately understand the population of amino acid residues on the surface and in the inner core of the proteins. RESULTS: ASAView is an algorithm, an application and a database of schematic representations of solvent accessibility of amino acid residues within proteins. A characteristic two-dimensional spiral plot of solvent accessibility provides a convenient graphical view of residues in terms of their exposed surface areas. In addition, sequential plots in the form of bar charts are also provided. Online plots of the proteins included in the entire Protein Data Bank (PDB), are provided for the entire protein as well as their chains separately. CONCLUSIONS: These graphical plots of solvent accessibility are likely to provide a quick view of the overall topological distribution of residues in proteins. Chain-wise computation of solvent accessibility is also provided.
Shandar Ahmad, M. Michael Gromiha, Hamed Fawareh, Akinori Sarai
BMC Bioinform.2
2003 RVP-net: online prediction of real valued accessible surface area of proteins from single sequences
abstract
SUMMARY: RVP-net is an online program for the prediction of real valued solvent accessibility. All previous methods of accessible surface area (ASA) predictions classify amino acid residues into exposure states and named them buried or exposed based on different thresholds. Real values in some cases were generated by taking the mid points of these state thresholds. This is the first method, which provides a direct prediction of ASA without making exposure categories and achieves results better than 19% mean absolute error. To facilitate batch processing of several sequences, a standalone version of this tool is also provided. AVAILABILITY: Online predictions are available at http://www.netasa.org/rvp-net/. Standalone version of the program can be obtained from the corresponding author by E-mail request.
Shandar Ahmad, M. Michael Gromiha, Akinori Sarai
Bioinform.2
2002 NETASA: neural network based prediction of solvent accessibility
abstract
Abstract Motivation: Prediction of the tertiary structure of a protein from its amino acid sequence is one of the most important problems in molecular biology. The successful prediction of solvent accessibility will be very helpful to achieve this goal. In the present work, we have implemented a server, NETASA for predicting solvent accessibility of amino acids using our newly optimized neural network algorithm. Several new features in the neural network architecture and training method have been introduced, and the network learns faster to provide accuracy values, which are comparable or better than other methods of ASA prediction. Results: Prediction in two and three state classification systems with several thresholds are provided. Our prediction method achieved the accuracy level upto 90% for training and 88% for test data sets. Three state prediction results provide a maximum 65% accuracy for training and 63% for the test data. Applicability of neural networks for ASA prediction has been confirmed with a larger data set and wider range of state thresholds. Salient differences between a linear and exponential network for ASA prediction have been analysed. Availability: Online predictions are freely available at: http://www.netasa.org. Linux ix86 binaries of the program written for this work may be obtained by email from the corresponding author. Contact: [email protected] * To whom correspondence should be addressed at: RIKEN Tsukuba Institute, 3-1-1, Koyadai, Tsukuba 305 0074, Ibaraki, Japan. † Present address: Computational Biology Research Center (CBRC), AIST 2-41-6 Aomi, Koto-ku, Tokyo 135-0064, Japan.
Shandar Ahmad, M. Michael Gromiha
Bioinform.2
2001 Thermodynamic database for protein-nucleic acid interactions (ProNIT)
abstract
MOTIVATION: Protein-nucleic acid interactions are fundamental to the regulation of gene expression. In order to elucidate the molecular mechanism of protein-nucleic acid recognition and analyze the gene regulation network, not only structural data but also quantitative binding data are necessary. Although there are structural databases for proteins and nucleic acids, there exists no database for their experimental binding data. Thus, we have developed a Thermodynamic Database for Protein-Nucleic Acid Interactions (ProNIT). RESULTS: We have collected experimentally observed binding data from the literature. ProNIT contains several important thermodynamic data for protein-nucleic acid binding, such as dissociation constant (K(d)), association constant (K(a)), Gibbs free energy change (DeltaG), enthalpy change (DeltaH), heat capacity change (DeltaC(p)), experimental conditions, structural information of proteins, nucleic acids and the complex, and literature information. These data are integrated into a relational database system together with structural and functional information to provide flexible searching facilities by using combinations of various terms and parameters. A www interface allows users to search for data based on various conditions, with different display and sorting options, and to visualize molecular structures and their interactions. AVAILABILITY: ProNIT is freely accessible at the URL http://www.rtc.riken.go.jp/jouhou/pronit/pronit.html.
Ponraj Prabakaran, Jianghong An, M. Michael Gromiha, Samuel Selvaraj, Hatsuho Uedaira, Hidetoshi Kono, Akinori Sarai
Bioinform.3