VLDB 2026 Research / reviewers in the wild / expert
Rakesh Kaundal
dblp:14/7518
· DBLP profile ↗
13ranked-venue papers
3as first author
4since 2021 · last 2022
0000-0001-8683-1240ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 13 · 3 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | deepNEC: a novel alignment-free tool for the identification and classification of nitrogen biochemical network-related enzymes using deep learningabstractNitrogen is essential for life and its transformations are an important part of the global biogeochemical cycle. Being an essential nutrient, nitrogen exists in a range of oxidation states from +5 (nitrate) to -3 (ammonium and amino-nitrogen), and its oxidation and reduction reactions catalyzed by microbial enzymes determine its environmental fate. The functional annotation of the genes encoding the core nitrogen network enzymes has a broad range of applications in metagenomics, agriculture, wastewater treatment and industrial biotechnology. This study developed an alignment-free computational approach to determine the predicted nitrogen biochemical network-related enzymes from the sequence itself. We propose deepNEC, a novel end-to-end feature selection and classification model training approach for nitrogen biochemical network-related enzyme prediction. The algorithm was developed using Deep Learning, a class of machine learning algorithms that uses multiple layers to extract higher-level features from the raw input data. The derived protein sequence is used as an input, extracting sequential and convolutional features from raw encoded protein sequences based on classification rather than traditional alignment-based methods for enzyme prediction. Two large datasets of protein sequences, enzymes and non-enzymes were used to train the models with protein sequence features like amino acid composition, dipeptide composition (DPC), conformation transition and distribution, normalized Moreau-Broto (NMBroto), conjoint and quasi order, etc. The k-fold cross-validation and independent testing were performed to validate our model training. deepNEC uses a four-tier approach for prediction; in the first phase, it will predict a query sequence as enzyme or non-enzyme; in the second phase, it will further predict and classify enzymes into nitrogen biochemical network-related enzymes or non-nitrogen metabolism enzymes; in the third phase, it classifies predicted enzymes into nine nitrogen metabolism classes; and in the fourth phase, it predicts the enzyme commission number out of 20 classes for nitrogen metabolism. Among all, the DPC + NMBroto hybrid feature gave the best prediction performance (accuracy of 96.15% in k-fold training and 93.43% in independent testing) with an Matthews correlation coefficient (0.92 training and 0.87 independent testing) in phase I; phase II (accuracy of 99.71% in k-fold training and 98.30% in independent testing); phase III (overall accuracy of 99.03% in k-fold training and 98.98% in independent testing); phase IV (overall accuracy of 99.05% in k-fold training and 98.18% in independent testing), the DPC feature gave the best prediction performance. We have also implemented a homology-based method to remove false negatives. All the models have been implemented on a web server (prediction tool), which is freely available at http://bioinfo.usu.edu/deepNEC/. Naveen Duhan, Jeanette M. Norton, Rakesh Kaundal |
Briefings Bioinform. | 3 |
| 2022 | deepHPI: a comprehensive deep learning platform for accurate prediction and visualization of host-pathogen protein-protein interactionsabstractHost-pathogen protein interactions (HPPIs) play vital roles in many biological processes and are directly involved in infectious diseases. With the outbreak of more frequent pandemics in the last couple of decades, such as the recent outburst of Covid-19 causing millions of deaths, it has become more critical to develop advanced methods to accurately predict pathogen interactions with their respective hosts. During the last decade, experimental methods to identify HPIs have been used to decipher host-pathogen systems with the caveat that those techniques are labor-intensive, expensive and time-consuming. Alternatively, accurate prediction of HPIs can be performed by the use of data-driven machine learning. To provide a more robust and accurate solution for the HPI prediction problem, we have developed a deepHPI tool based on deep learning. The web server delivers four host-pathogen model types: plant-pathogen, human-bacteria, human-virus and animal-pathogen, leveraging its operability to a wide range of analyses and cases of use. The deepHPI web tool is the first to use convolutional neural network models for HPI prediction. These models have been selected based on a comprehensive evaluation of protein features and neural network architectures. The best prediction models have been tested on independent validation datasets, which achieved an overall Matthews correlation coefficient value of 0.87 for animal-pathogen using the combined pseudo-amino acid composition and conjoint triad (PAAC_CT) features, 0.75 for human-bacteria using the combined pseudo-amino acid composition, conjoint triad and normalized Moreau-Broto feature (PAAC_CT_NMBroto), 0.96 for human-virus using PAAC_CT_NMBroto and 0.94 values for plant-pathogen interactions using the combined pseudo-amino acid composition, composition and transition feature (PAAC_CTDC_CTDT). Our server running deepHPI is deployed on a high-performance computing cluster that enables large and multiple user requests, and it provides more information about interactions discovered. It presents an enriched visualization of the resulting host-pathogen networks that is augmented with external links to various protein annotation resources. We believe that the deepHPI web server will be very useful to researchers, particularly those working on infectious diseases. Additionally, many novel and known host-pathogen systems can be further investigated to significantly advance our understanding of complex disease-causing agents. The developed models are established on a web server, which is freely accessible at http://bioinfo.usu.edu/deepHPI/. Rakesh Kaundal, Cristian D. Loaiza, Naveen Duhan, Nicholas Flann |
Briefings Bioinform. | 1 |
| 2021 | In silico prediction of host-pathogen protein interactions in melioidosis pathogen Burkholderia pseudomallei and human reveals novel virulence factors and their targetsabstractThe aerobic, Gram-negative motile bacillus, Burkholderia pseudomallei is a facultative intracellular bacterium causing melioidosis, a critical disease of public health importance, which is widely endemic in the tropics and subtropical regions of the world. Melioidosis is associated with high case fatality rates in animals and humans; even with treatment, its mortality is 20-50%. It also infects plants and is designated as a biothreat agent. B. pseudomallei is pathogenic due to its ability to invade, resist factors in serum and survive intracellularly. Despite its importance, to date only a few effector proteins have been functionally characterized, and there is not much information regarding the host-pathogen protein-protein interactions (PPI) of this system, which are important to studying infection mechanisms and thereby develop prevention measures. We explored two computational approaches, the homology-based interolog and the domain-based method, to predict genome-scale host-pathogen interactions (HPIs) between two different strains of B. pseudomallei (prototypical, and highly virulent) and human. In total, 76 335 common HPIs (between the two strains) were predicted involving 8264 human and 1753 B. pseudomallei proteins. Among the unique PPIs, 14 131 non-redundant HPIs were found to be unique between the prototypical strain and human, compared to 3043 non-redundant HPIs between the highly virulent strain and human. The protein hubs analysis showed that most B. pseudomallei proteins formed a hub with human dnaK complex proteins associated with tuberculosis, a disease similar in symptoms to melioidosis. In addition, drug-binding and carbohydrate-binding mechanisms were found overrepresented within the host-pathogen network, and metabolic pathways were frequently activated according to the pathway enrichment. Subcellular localization analysis showed that most of the pathogen proteins are targeting human proteins inside cytoplasm and nucleus. We also discovered the host targets of the drug-related pathogen proteins and proteins that form T3SS and T6SS in B. pseudomallei. Additionally, a comparison between the unique PPI patterns present in the prototypical and highly virulent strains was performed. The current study is the first report on developing a genome-scale host-pathogen protein interaction networks between the human and B. pseudomallei, a critical biothreat agent. We have identified novel virulence factors and their interacting partners in the human proteome. These PPIs can be further validated by high-throughput experiments and may give new insights on how B. pseudomallei interacts with its host, which will help medical researchers in developing better prevention measures. Cristian D. Loaiza, Naveen Duhan, Matthew Lister, Rakesh Kaundal |
Briefings Bioinform. | 4 |
| 2021 | PredHPI: an integrated web server platform for the detection and visualization of host-pathogen interactions using sequence-based methodsabstractMOTIVATION: Understanding the mechanisms underlying infectious diseases is fundamental to develop prevention strategies. Host-pathogen interactions (HPIs) are actively studied worldwide to find potential genomic targets for the development of novel drugs, vaccines and other therapeutics. Determining which proteins are involved in the interaction system behind an infectious process is the first step to develop an efficient disease control strategy. Very few computational methods have been implemented as web services to infer novel HPIs, and there is not a single framework which combines several of those approaches to produce and visualize a comprehensive analysis of HPIs. RESULTS: Here, we introduce PredHPI, a powerful framework that integrates both the detection and visualization of interaction networks in a single web service, facilitating the apprehension of model and non-model host-pathogen systems to aid the biologists in building hypotheses and designing appropriate experiments. PredHPI is built on high-performance computing resources on the backend capable of handling proteome-scale sequence data from both the host as well as pathogen. Data are displayed in an information-rich and interactive visualization, which can be further customized with user-defined layouts. We believe PredHPI will serve as an invaluable resource to diverse experimental biologists and will help advance the research in the understanding of complex infectious diseases. AVAILABILITY AND IMPLEMENTATION: PredHPI tool is freely available at http://bioinfo.usu.edu/PredHPI/. SUPPLEMENTARY INFORMATION: Sup plementary data are available at Bioinformatics online. Cristian D. Loaiza, Rakesh Kaundal |
Bioinform. | 2 |
| 2016 | Proceedings of the 2016 MidSouth Computational Biology and Bioinformatics Society (MCBIOS) ConferenceabstractThe MidSouth Computational Biology and Bioinformatics Society (MCBIOS) held its thirteenth annual conference themed “Precision Medicine and Data Sciences” at the University of Memphis, FedEx Institute of Technology in Memphis, Tennessee on March 3–5, 2016. There were 156 conference registrants and 117 abstracts submitted, including 63 oral and 54 poster presentations. Jonathan D. Wren, Inimary T. Toby, Huxiao Hong, Bindu Nanduri, Rakesh Kaundal, Mikhail G. Dozmorov, Shraddha Thakkar |
BMC Bioinform. | 5 |
| 2014 | Predicting genome-scale Arabidopsis-Pseudomonas syringae interactome using domain and interolog-based approachesabstractBACKGROUND: Every year pathogenic organisms cause billions of dollars' worth damage to crops and livestock. In agriculture, study of plant-microbe interactions is demanding a special attention to develop management strategies for the destructive pathogen induced diseases that cause huge crop losses every year worldwide. Pseudomonas syringae is a major bacterial leaf pathogen that causes diseases in a wide range of plant species. Among its various strains, pathovar tomato strain DC3000 (PstDC3000) is asserted to infect the plant host Arabidopsis thaliana and thus, has been accepted as a model system for experimental characterization of the molecular dynamics of plant-pathogen interactions. Protein-protein interactions (PPIs) play a critical role in initiating pathogenesis and maintaining infection. Understanding the PPI network between a host and pathogen is a critical step for studying the molecular basis of pathogenesis. The experimental study of PPIs at a large scale is very scarce and also the high throughput experimental results show high false positive rate. Hence, there is a need for developing efficient computational models to predict the interaction between host and pathogen in a genome scale, and find novel candidate effectors and/or their targets. RESULTS: In this study, we used two computational approaches, the interolog and the domain-based to predict the interactions between Arabidopsis and PstDC3000 in genome scale. The interolog method relies on protein sequence similarity to conduct the PPI prediction. A Pseudomonas protein and an Arabidopsis protein are predicted to interact with each other if an experimentally verified interaction exists between their respective homologous proteins in another organism. The domain-based method uses domain interaction information, which is derived from known protein 3D structures, to infer the potential PPIs. If a Pseudomonas and an Arabidopsis protein contain an interacting domain pair, one can expect the two proteins to interact with each other. The interolog-based method predicts ~0.79M PPIs involving around 7700 Arabidopsis and 1068 Pseudomonas proteins in the full genome. The domain-based method predicts 85650 PPIs comprising 11432 Arabidopsis and 887 Pseudomonas proteins. Further, around 11000 PPIs have been identified as interacting from both the methods as a consensus. CONCLUSION: The present work predicts the protein-protein interaction network between Arabidopsis thaliana and Pseudomonas syringae pv. tomato DC3000 in a genome wide scale with a high confidence. Although the predicted PPIs may contain some false positives, the computational methods provide reasonable amount of interactions which can be further validated by high throughput experiments. This can be a useful resource to the plant community to characterize the host-pathogen interaction in Arabidopsis and Pseudomonas system. Further, these prediction models can be applied to the agriculturally relevant crops. Sitanshu Sekhar Sahu, Tyler Weirick, Rakesh Kaundal |
BMC Bioinform. | 3 |
| 2014 | LacSubPred: predicting subtypes of Laccases, an important lignin metabolism-related enzyme class, using in silico approachesabstractBACKGROUND: Laccases (E.C. 1.10.3.2) are multi-copper oxidases that have gained importance in many industries such as biofuels, pulp production, textile dye bleaching, bioremediation, and food production. Their usefulness stems from the ability to act on a diverse range of phenolic compounds such as o-/p-quinols, aminophenols, polyphenols, polyamines, aryl diamines, and aromatic thiols. Despite acting on a wide range of compounds as a family, individual Laccases often exhibit distinctive and varied substrate ranges. This is likely due to Laccases involvement in many metabolic roles across diverse taxa. Classification systems for multi-copper oxidases have been developed using multiple sequence alignments, however, these systems seem to largely follow species taxonomy rather than substrate ranges, enzyme properties, or specific function. It has been suggested that the roles and substrates of various Laccases are related to their optimal pH. This is consistent with the observation that fungal Laccases usually prefer acidic conditions, whereas plant and bacterial Laccases prefer basic conditions. Based on these observations, we hypothesize that a descriptor-based unsupervised learning system could generate homology independent classification system for better describing the functional properties of Laccases. RESULTS: In this study, we first utilized unsupervised learning approach to develop a novel homology independent Laccase classification system. From the descriptors considered, physicochemical properties showed the best performance. Physicochemical properties divided the Laccases into twelve subtypes. Analysis of the clusters using a t-test revealed that the majority of the physicochemical descriptors had statistically significant differences between the classes. Feature selection identified the most important features as negatively charges residues, the peptide isoelectric point, and acidic or amidic residues. Secondly, to allow for classification of new Laccases, a supervised learning system was developed from the clusters. The models showed high performance with an overall accuracy of 99.03%, error of 0.49%, MCC of 0.9367, precision of 94.20%, sensitivity of 94.20%, and specificity of 99.47% in a 5-fold cross-validation test. In an independent test, our models still provide a high accuracy of 97.98%, error rate of 1.02%, MCC of 0.8678, precision of 87.88%, sensitivity of 87.88% and specificity of 98.90%. CONCLUSION: This study provides a useful classification system for better understanding of Laccases from their physicochemical properties perspective. We also developed a publically available web tool for the characterization of Laccase protein sequences (http://lacsubpred.bioinfo.ucr.edu/). Finally, the programs used in the study are made available for researchers interested in applying the system to other enzyme classes (https://github.com/tweirick/SubClPred). Tyler Weirick, Sitanshu Sekhar Sahu, Ramamurthy Mahalingam, Rakesh Kaundal |
BMC Bioinform. | 4 |
| 2014 | Proceedings of the 2014 MidSouth Computational Biology and Bioinformatics Society (MCBIOS) ConferenceabstractThe MidSouth Computational Biology and Bioinformatics Society (MCBIOS 2014) held its eleventh annual conference at the Wes Watkins Center at Oklahoma State University, Stillwater on March 7-8, 2014. The theme was " From Genome to Phenome: Connecting the Dots ". Conference Chair this year was Rakesh Kaundal, who is also one of the MCBIOS board members, and conference committee members were Ulrich K. Melcher and Doris Kupfer. The current president is Andy Perkins and Cesar Compadre was elected as President-Elect for 2015-16. There were 154 registrants and a total of 125 abstracts submitted (50 oral and 75 poster presentations). Jonathan D. Wren, Mikhail G. Dozmorov, Dennis Burian, Andy D. Perkins, Peter Hoyt, Rakesh Kaundal |
BMC Bioinform. | 7 |
| 2013 | PHDcleav: a SVM based method for predicting human Dicer cleavage sites using sequence and secondary structure of miRNA precursorsabstractBACKGROUND: Dicer, an RNase III enzyme, plays a vital role in the processing of pre-miRNAs for generating the miRNAs. The structural and sequence features on pre-miRNA which can facilitate position and efficiency of cleavage are not well known. A precise cleavage by Dicer is crucial because an inaccurate processing can produce miRNA with different seed regions which can alter the repertoire of target genes. RESULTS: In this study, a novel method has been developed to predict Dicer cleavage sites on pre-miRNAs using Support Vector Machine. We used the dataset of experimentally validated human miRNA hairpins from miRBase, and extracted fourteen nucleotides around Dicer cleavage sites. We developed number of models using various types of features and achieved maximum accuracy of 66% using binary profile of nucleotide sequence taken from 5p arm of hairpin. The prediction performance of Dicer cleavage site improved significantly from 66% to 86% when we integrated secondary structure information. This indicates that secondary structure plays an important role in the selection of cleavage site. All models were trained and tested on 555 experimentally validated cleavage sites and evaluated using 5-fold cross validation technique. In addition, the performance was also evaluated on an independent testing dataset that achieved an accuracy of ~82%. CONCLUSION: Based on this study, we developed a webserver PHDcleav (http://www.imtech.res.in/raghava/phdcleav/) to predict Dicer cleavage sites in pre-miRNA. This tool can be used to investigate functional consequences of genetic variations/SNPs in miRNA on Dicer cleavage site, and gene silencing. Moreover, it would also be useful in the discovery of miRNAs in human genome and design of Dicer specific pre-miRNAs for potent gene silencing. Rakesh Kaundal, Gajendra P. S. Raghava |
BMC Bioinform. | 2 |
| 2013 | Identification and characterization of plastid-type proteins from sequence-attributed features using machine learningabstractBACKGROUND: Plastids are an important component of plant cells, being the site of manufacture and storage of chemical compounds used by the cell, and contain pigments such as those used in photosynthesis, starch synthesis/storage, cell color etc. They are essential organelles of the plant cell, also present in algae. Recent advances in genomic technology and sequencing efforts is generating a huge amount of DNA sequence data every day. The predicted proteome of these genomes needs annotation at a faster pace. In view of this, one such annotation need is to develop an automated system that can distinguish between plastid and non-plastid proteins accurately, and further classify plastid-types based on their functionality. We compared the amino acid compositions of plastid proteins with those of non-plastid ones and found significant differences, which were used as a basis to develop various feature-based prediction models using similarity-search and machine learning. RESULTS: In this study, we developed separate Support Vector Machine (SVM) trained classifiers for characterizing the plastids in two steps: first distinguishing the plastid vs. non-plastid proteins, and then classifying the identified plastids into their various types based on their function (chloroplast, chromoplast, etioplast, and amyloplast). Five diverse protein features: amino acid composition, dipeptide composition, the pseudo amino acid composition, N(terminal)-Center-C(terminal) composition and the protein physicochemical properties are used to develop SVM models. Overall, the dipeptide composition-based module shows the best performance with an accuracy of 86.80% and Matthews Correlation Coefficient (MCC) of 0.74 in phase-I and 78.60% with a MCC of 0.44 in phase-II. On independent test data, this model also performs better with an overall accuracy of 76.58% and 74.97% in phase-I and phase-II, respectively. The similarity-based PSI-BLAST module shows very low performance with about 50% prediction accuracy for distinguishing plastid vs. non-plastids and only 20% in classifying various plastid-types, indicating the need and importance of machine learning algorithms. CONCLUSION: The current work is a first attempt to develop a methodology for classifying various plastid-type proteins. The prediction modules have also been made available as a web tool, PLpred available at http://bioinfo.okstate.edu/PLpred/ for real time identification/characterization. We believe this tool will be very useful in the functional annotation of various genomes. Rakesh Kaundal, Sitanshu Sekhar Sahu, Ruchi Verma, Tyler Weirick |
BMC Bioinform. | 1 |
| 2013 | Proceedings of the 2013 MidSouth Computational Biology and Bioinformatics Society (MCBIOS) ConferenceabstractThe tenth annual conference of the MidSouth Computational Biology and Bioinformatics Society (MCBIOS 2013), "The 10th Anniversary in a Decade of Change: Discovery in a Sea of Data", took place at the Stoney Creek Inn & Conference Center in Columbia, Missouri on April 5-6, 2013. This year's Conference Chairs were Gordon Springer and Chi-Ren Shyu from the University of Missouri and Edward Perkins from the US Army Corps of Engineers Engineering Research and Development Center, who is also the current MCBIOS President (2012-3). There were 151 registrants and a total of 111 abstracts (51 oral presentations and 60 poster session abstracts). Jonathan D. Wren, Mikhail G. Dozmorov, Dennis Burian, Rakesh Kaundal, Andy D. Perkins, Edward J. Perkins, Doris M. Kupfer, Gordon K. Springer |
BMC Bioinform. | 4 |
| 2012 | Proceedings of the 2012 MidSouth computational biology and bioinformatics society (MCBIOS) conferenceabstractThe ninth annual conference of the MidSouth Computational Biology and Bioinformatics Society (MCBIOS 2012), "Making Sense of the Omics Data Deluge", took place in Oxford, Mississippi February 17-8 2012. This year's Conference Chairs were Dr. Dawn Wilkins, of the University of Mississippi and Dr. Doris Kupfer, also the current MCBIOS President (2011-2), from the Federal Aviation Administration. There were 170 registrants and a total of 106 abstracts (34 oral presentations and 72 poster session abstracts). Jonathan D. Wren, Mikhail G. Dozmorov, Dennis Burian, Rakesh Kaundal, Susan M. Bridges, Doris M. Kupfer |
BMC Bioinform. | 4 |
| 2006 | Machine learning techniques in disease forecasting: a case study on rice blast predictionabstractBACKGROUND: Diverse modeling approaches viz. neural networks and multiple regression have been followed to date for disease prediction in plant populations. However, due to their inability to predict value of unknown data points and longer training times, there is need for exploiting new prediction softwares for better understanding of plant-pathogen-environment relationships. Further, there is no online tool available which can help the plant researchers or farmers in timely application of control measures. This paper introduces a new prediction approach based on support vector machines for developing weather-based prediction models of plant diseases. RESULTS: Six significant weather variables were selected as predictor variables. Two series of models (cross-location and cross-year) were developed and validated using a five-fold cross validation procedure. For cross-year models, the conventional multiple regression (REG) approach achieved an average correlation coefficient (r) of 0.50, which increased to 0.60 and percent mean absolute error (%MAE) decreased from 65.42 to 52.24 when back-propagation neural network (BPNN) was used. With generalized regression neural network (GRNN), the r increased to 0.70 and %MAE also improved to 46.30, which further increased to r = 0.77 and %MAE = 36.66 when support vector machine (SVM) based method was used. Similarly, cross-location validation achieved r = 0.48, 0.56 and 0.66 using REG, BPNN and GRNN respectively, with their corresponding %MAE as 77.54, 66.11 and 58.26. The SVM-based method outperformed all the three approaches by further increasing r to 0.74 with improvement in %MAE to 44.12. Overall, this SVM-based prediction approach will open new vistas in the area of forecasting plant diseases of various crops. CONCLUSION: Our case study demonstrated that SVM is better than existing machine learning techniques and conventional REG approaches in forecasting plant diseases. In this direction, we have also developed a SVM-based web server for rice blast prediction, a first of its kind worldwide, which can help the plant science community and farmers in their decision making process. The server is freely available at http://www.imtech.res.in/raghava/rbpred/. Rakesh Kaundal, Amar S. Kapoor, Gajendra P. S. Raghava |
BMC Bioinform. | 1 |