EDBT 2026 Demo / reviewers in the wild / expert
Dukka B. KC
dblp:b/KCDukkaBahadur · also Dukka Bahadur, K. C. Dukka B., K. C. Dukka Bahadur
· DBLP profile ↗
17ranked-venue papers
5as first author
7since 2021 · last 2025
0000-0001-5590-5403ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 16 · 5 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PLM-DBPs: enhancing plant DNA-binding protein prediction by integrating sequence-based and structure-aware protein language modelsabstractDNA-binding proteins (DBPs) play a crucial role in gene regulation, development, and environmental responses across plants, animals, and microorganisms. Existing DBP prediction methods are largely limited to sequence information, whether through handcrafted features or sequence-based protein language models (PLMs), overlooking structural cues critical to protein function. In addition, most existing tools are trained for general DBP predictions, which are often not accurate for plant-specific DBPs due to the unique structural and functional properties of plant proteins. Our work introduces PLM-DBPs, a deep learning framework that integrates both sequence-based and structure-aware representations to enhance DBP prediction in plants. We evaluated several state-of-the-art PLMs to extract high-dimensional protein representations and experimented with various fusion strategies to validate the complementary information between the various representations. Our final model, a fusion of sequence-based and structure-aware ANN models, achieves a notable improvement in predicting DBPs in plants outperforming previous state-of-the-art models. Although sequence-based PLMs already demonstrate strong performance in DBP prediction, our findings show that the integration of structural information further enhances predictive accuracy. This underscores the complementary nature of structural representations and establishes PLM-DBPs as a robust tool for advancing plant research and agricultural innovation. The proposed model and other resources are publicly available at https://github.com/suresh-pokharel/PLM-DBPs. Suresh Pokharel, Kepha Barasa, Pawel Pratyush, Dukka B. KC |
Briefings Bioinform. | 4 |
| 2025 | CaLMPhosKAN: prediction of general phosphorylation sites in proteins via fusion of codon aware embeddings with amino acid aware embeddings and wavelet-based Kolmogorov-Arnold networkabstractMOTIVATION: The mapping from codon to amino acid is surjective due to codon degeneracy, suggesting that codon space might harbor higher information content. Embeddings from the codon language model have recently demonstrated success in various protein downstream tasks. However, predictive models for residue-level tasks such as phosphorylation sites, arguably the most studied Post-Translational Modification (PTM), and PTM sites prediction in general, have predominantly relied on representations in amino acid space. RESULTS: We introduce a novel approach for predicting phosphorylation sites by utilizing codon-level information through embeddings from the codon adaptation language model (CaLM), trained on protein-coding DNA sequences. Protein sequences are first reverse-translated into reliable coding sequences by mapping UniProt sequences to their corresponding NCBI reference sequences and extracting the exact coding sequences from their GenBank format using a dynamic programming-based global pairwise alignment. The resulting coding sequences are encoded using the CaLM encoder to generate codon-aware embeddings, which are subsequently integrated with amino acid-aware embeddings obtained from a protein language model, through an early fusion strategy. Next, a window-level representation of the site of interest, retaining the full sequence context, is constructed from the fused embeddings. A ConvBiGRU network extracts feature maps that capture spatiotemporal correlations between proximal residues within the window. This is followed by a prediction head based on a Kolmogorov-Arnold network (KAN) using the derivative of gaussian wavelet transform to generate the inference for the site. The overall model, dubbed CaLMPhosKAN, performs better than the existing approaches across multiple datasets. AVAILABILITY AND IMPLEMENTATION: CaLMPhosKAN is publicly available at https://github.com/KCLabMTU/CaLMPhosKAN. Pawel Pratyush, Callen Carrier, Suresh Pokharel, Hamid D. Ismail, Meenal Chaudhari, Dukka B. KC |
Bioinform. | 6 |
| 2024 | Benchmarking a Wide Range of Unsupervised Learning Methods for Detecting Anomaly in Blast FurnaceabstractSteel plays important roles in our daily lives, as it surrounds us in the form of various products. Blast furnace, one of the main facility in steel production process, is traditionally monitored by skilled workers to prevent incidents. However, there is a growing demand to automate the monitoring process by leveraging machine learning. This paper focuses on investigating the suitability of unsupervised learning methods for detecting anomalies in blast furnaces. Extensive benchmarking is conducted using a dataset collected from blast furnaces, encompassing a wide range of unsupervised learning methods, including both traditional approaches and recent deep learning-based techniques. The computational experiments yield results that suggest the effectiveness of traditional methods over deep learning-based methods. To validate this observation, additional experiments are performed on publicly available non time series datasets and complex time series datasets. These experiments serve to confirm the superiority of traditional methods in handling non time series datasets, while deep learning methods exhibit better performance in dealing with complex time series datasets. We have also discovered that dimensionality reduction before anomaly detection is beneficial in eliminating outliers and effectively modeling the normal data points in the blast furnace dataset. Kendai Itakura, Dukka B. KC, Hiroto Saigo |
ICPRAM | 2 |
| 2024 | LMCrot: an enhanced protein crotonylation site predictor by leveraging an interpretable window-level embedding from a transformer-based protein language modelabstractMOTIVATION: Recent advancements in natural language processing have highlighted the effectiveness of global contextualized representations from protein language models (pLMs) in numerous downstream tasks. Nonetheless, strategies to encode the site-of-interest leveraging pLMs for per-residue prediction tasks, such as crotonylation (Kcr) prediction, remain largely uncharted. RESULTS: Herein, we adopt a range of approaches for utilizing pLMs by experimenting with different input sequence types (full-length protein sequence versus window sequence), assessing the implications of utilizing per-residue embedding of the site-of-interest as well as embeddings of window residues centered around it. Building upon these insights, we developed a novel residual ConvBiLSTM network designed to process window-level embeddings of the site-of-interest generated by the ProtT5-XL-UniRef50 pLM using full-length sequences as input. This model, termed T5ResConvBiLSTM, surpasses existing state-of-the-art Kcr predictors in performance across three diverse datasets. To validate our approach of utilizing full sequence-based window-level embeddings, we also delved into the interpretability of ProtT5-derived embedding tensors in two ways: firstly, by scrutinizing the attention weights obtained from the transformer's encoder block; and secondly, by computing SHAP values for these tensors, providing a model-agnostic interpretation of the prediction results. Additionally, we enhance the latent representation of ProtT5 by incorporating two additional local representations, one derived from amino acid properties and the other from supervised embedding layer, through an intermediate fusion stacked generalization approach, using an n-mer window sequence (or, peptide/fragment). The resultant stacked model, dubbed LMCrot, exhibits a more pronounced improvement in predictive performance across the tested datasets. AVAILABILITY AND IMPLEMENTATION: LMCrot is publicly available at https://github.com/KCLabMTU/LMCrot. Pawel Pratyush, Soufia Bahmani, Suresh Pokharel, Hamid D. Ismail, Dukka B. KC |
Bioinform. | 5 |
| 2023 | pLMSNOSite: an ensemble-based approach for predicting protein S-nitrosylation sites by integrating supervised word embedding and embedding from pre-trained protein language modelabstractBACKGROUND: Protein S-nitrosylation (SNO) plays a key role in transferring nitric oxide-mediated signals in both animals and plants and has emerged as an important mechanism for regulating protein functions and cell signaling of all main classes of protein. It is involved in several biological processes including immune response, protein stability, transcription regulation, post translational regulation, DNA damage repair, redox regulation, and is an emerging paradigm of redox signaling for protection against oxidative stress. The development of robust computational tools to predict protein SNO sites would contribute to further interpretation of the pathological and physiological mechanisms of SNO. RESULTS: Using an intermediate fusion-based stacked generalization approach, we integrated embeddings from supervised embedding layer and contextualized protein language model (ProtT5) and developed a tool called pLMSNOSite (protein language model-based SNO site predictor). On an independent test set of experimentally identified SNO sites, pLMSNOSite achieved values of 0.340, 0.735 and 0.773 for MCC, sensitivity and specificity respectively. These results show that pLMSNOSite performs better than the compared approaches for the prediction of S-nitrosylation sites. CONCLUSION: Together, the experimental results suggest that pLMSNOSite achieves significant improvement in the prediction performance of S-nitrosylation sites and represents a robust computational approach for predicting protein S-nitrosylation sites. pLMSNOSite could be a useful resource for further elucidation of SNO and is publicly available at https://github.com/KCLabMTU/pLMSNOSite . Pawel Pratyush, Suresh Pokharel, Hiroto Saigo, Dukka B. KC |
BMC Bioinform. | 4 |
| 2022 | GPU-I-TASSER: a GPU accelerated I-TASSER protein structure prediction toolabstractMOTIVATION: Accurate and efficient predictions of protein structures play an important role in understanding their functions. Iterative Threading Assembly Refinement (I-TASSER) is one of the most successful and widely used protein structure prediction methods in the recent community-wide CASP experiments. Yet, the computational efficiency of I-TASSER is one of the limiting factors that prevent its application for large-scale structure modeling. RESULTS: We present I-TASSER for Graphics Processing Units (GPU-I-TASSER), a GPU accelerated I-TASSER protein structure prediction tool for fast and accurate protein structure prediction. Our implementation is based on OpenACC parallelization of the replica-exchange Monte Carlo simulations to enhance the speed of I-TASSER by extending its capabilities to the GPU architecture. On a benchmark dataset of 71 protein structures, GPU-I-TASSER achieves on average a 10× speedup with comparable structure prediction accuracy compared to the CPU version of the I-TASSER. AVAILABILITY AND IMPLEMENTATION: The complete source code for GPU-I-TASSER can be downloaded and used without restriction from https://zhanggroup.org/GPU-I-TASSER/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Elijah A. MacCarthy, Yang Zhang 0040, Dukka B. KC |
Bioinform. | 4 |
| 2022 | Correction: DeepSuccinylSite: a deep learning based approach for protein succinylation site predictionabstractResults: Using an independent test set of experimentally identified succinylation sites, our method achieved efficiency scores of 79%, 68.7% and 0.27 for sensitivity, specificity and MCC respectively, with an area under the receiver operator characteristic (ROC) curve of 0.8.In side-by-side comparisons with previously described succinylation site predictors, DeepSuccinylSite produces similar or better results compared to the other state-of-the-art predictors.On page 7, Last paragraph on right should be changed from Consequently, DeepSuccinylSite achieved a significantly higher performance as measured by MCC.Indeed, DeepSuccinylSite exhibited an ~ 62% increase in MCC when compared to the next highest method, GPSuc.to: Consequently, DeepSuccinylSite achieved an MCC score (at decision boundary of 0.5) on par with the top performingmethod, GPSuc.On page 2, in Table 1, the negative data of Independent Test should be 2977 rather than 254.On page 8, in Table 6, the MCC data of DeepSuccinylSite should be 0.27 rather than 0.48. Niraj Thapa, Meenal Chaudhari, Sean McManus, Kaushik Roy 0003, Robert H. Newman, Hiroto Saigo, Dukka B. KC |
BMC Bioinform. | 7 |
| 2020 | DeepSuccinylSite: a deep learning based approach for protein succinylation site predictionabstractBACKGROUND: Protein succinylation has recently emerged as an important and common post-translation modification (PTM) that occurs on lysine residues. Succinylation is notable both in its size (e.g., at 100 Da, it is one of the larger chemical PTMs) and in its ability to modify the net charge of the modified lysine residue from + 1 to - 1 at physiological pH. The gross local changes that occur in proteins upon succinylation have been shown to correspond with changes in gene activity and to be perturbed by defects in the citric acid cycle. These observations, together with the fact that succinate is generated as a metabolic intermediate during cellular respiration, have led to suggestions that protein succinylation may play a role in the interaction between cellular metabolism and important cellular functions. For instance, succinylation likely represents an important aspect of genomic regulation and repair and may have important consequences in the etiology of a number of disease states. In this study, we developed DeepSuccinylSite, a novel prediction tool that uses deep learning methodology along with embedding to identify succinylation sites in proteins based on their primary structure. RESULTS: Using an independent test set of experimentally identified succinylation sites, our method achieved efficiency scores of 79%, 68.7% and 0.48 for sensitivity, specificity and MCC respectively, with an area under the receiver operator characteristic (ROC) curve of 0.8. In side-by-side comparisons with previously described succinylation predictors, DeepSuccinylSite represents a significant improvement in overall accuracy for prediction of succinylation sites. CONCLUSION: Together, these results suggest that our method represents a robust and complementary technique for advanced exploration of protein succinylation. Niraj Thapa, Meenal Chaudhari, Sean McManus, Kaushik Roy 0003, Robert H. Newman, Hiroto Saigo, Dukka B. KC |
BMC Bioinform. | 7 |
| 2018 | RF-NR: Random Forest Based Approach for Improved Classification of Nuclear ReceptorsabstractThe Nuclear Receptor (NR) superfamily plays an important role in key biological, developmental, and physiological processes. Developing a method for the classification of NR proteins is an important step towards understanding the structure and functions of the newly discovered NR protein. The recent studies on NR classification are either unable to achieve optimum accuracy or are not designed for all the known NR subfamilies. In this study, we developed RF-NR, which is a Random Forest based approach for improved classification of nuclear receptors. The RF-NR can predict whether a query protein sequence belongs to one of the eight NR subfamilies or it is a non-NR sequence. The RF-NR uses spectrum-like features namely: Amino Acid Composition, Di-peptide Composition, and Tripeptide Composition. Benchmarking on two independent datasets with varying sequence redundancy reduction criteria, the RF-NR achieves better (or comparable) accuracy than other existing methods. The added advantage of our approach is that we can also obtain biological insights about the important features that are required to classify NR subfamilies. RF-NR is freely available at http://bcb.ncat.edu/RF_NR. Hamid D. Ismail, Hiroto Saigo, Dukka B. KC |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2017 | Recent advances in sequence-based protein structure predictionabstractThe most accurate characterizations of the structure of proteins are provided by structural biology experiments. However, because of the high cost and labor-intensive nature of the structural experiments, the gap between the number of protein sequences and solved structures is widening rapidly. Development of computational methods to accurately model protein structures from sequences is becoming increasingly important to the biological community. In this article, we highlight some important progress in the field of protein structure prediction, especially those related to free modeling (FM) methods that generate structure models without using homologous templates. We also provide a short synopsis of some of the recent advances in FM approaches as demonstrated in the recent Computational Assessment of Structure Prediction competition as well as recent trends and outlook for FM approaches in protein structure prediction. Dukka B. KC |
Briefings Bioinform. | 1 |
| 2017 | CNN-BLPred: a Convolutional neural network based predictor for β-Lactamases (BL) and their classesabstractBACKGROUND: The β-Lactamase (BL) enzyme family is an important class of enzymes that plays a key role in bacterial resistance to antibiotics. As the newly identified number of BL enzymes is increasing daily, it is imperative to develop a computational tool to classify the newly identified BL enzymes into one of its classes. There are two types of classification of BL enzymes: Molecular Classification and Functional Classification. Existing computational methods only address Molecular Classification and the performance of these existing methods is unsatisfactory. RESULTS: We addressed the unsatisfactory performance of the existing methods by implementing a Deep Learning approach called Convolutional Neural Network (CNN). We developed CNN-BLPred, an approach for the classification of BL proteins. The CNN-BLPred uses Gradient Boosted Feature Selection (GBFS) in order to select the ideal feature set for each BL classification. Based on the rigorous benchmarking of CCN-BLPred using both leave-one-out cross-validation and independent test sets, CCN-BLPred performed better than the other existing algorithms. Compared with other architectures of CNN, Recurrent Neural Network, and Random Forest, the simple CNN architecture with only one convolutional layer performs the best. After feature extraction, we were able to remove ~95% of the 10,912 features using Gradient Boosted Trees. During 10-fold cross validation, we increased the accuracy of the classic BL predictions by 7%. We also increased the accuracy of Class A, Class B, Class C, and Class D performance by an average of 25.64%. The independent test results followed a similar trend. CONCLUSIONS: We implemented a deep learning algorithm known as Convolutional Neural Network (CNN) to develop a classifier for BL classification. Combined with feature selection on an exhaustive feature set and using balancing method such as Random Oversampling (ROS), Random Undersampling (RUS) and Synthetic Minority Oversampling Technique (SMOTE), CNN-BLPred performs significantly better than existing algorithms for BL classification. Clarence White, Hamid D. Ismail, Hiroto Saigo, Dukka B. KC |
BMC Bioinform. | 4 |
| 2015 | RF-Phos: Random forest-based prediction of phosphorylation sitesabstractIt is estimated that about 30% of the proteins in the human proteome are regulated by phosphorylation. In recent years, phosphorylation site prediction has been investigated in the field of bioinformatics. This has become necessary due to the challenges associated with experimental methods. Previously, we developed a random forest-based method, termed Random Forest-based Phosphosite predictor (RF-Phos 1.0), to predict phosphorylation sites in proteins given only the amino acid sequence of a protein as input. Here, we report an improved version of this method, termed RF-Phos 1.1 that employs additional sequence-driven features to identify putative sites of phosphorylation across many protein families. In side-by-side comparisons based on 10-fold cross validation analysis and an independent dataset, RF-Phos 1.1 performs comparably to or better than other existing phosphosite prediction methods, such as PhosphoSVM, GPS2.1 and Musite. Ahoi Jones, Hamid D. Ismail, Robert H. Newman, Dukka B. KC |
BIBM | 5 |
| 2013 | Hierarchical multi-label gene function prediction using adaptive mutation in crowding nichingabstractComputational prediction of protein function is an important field in functional genomics. Gene function prediction is a Hierarchical Multi Label Classification (HMC) problem where each gene can belong to more than one functional class simultaneously, while classes are structured in the form of hierarchy. HMC is becoming a necessity in many domains of applications as well. Crowding niching-Adaptive mutation (CAM) is a new proposed method for solving Hierarchical multi-label gene function prediction problem. The classification in CAM-HMC is structured in three different phases. In the first two phases, a sequential procedure is performed. In the first phase, a full cyclic evolutionary crowding algorithm based on new definition of distance between two individuals, and adaptive mutation is applied in order to find classification rules. In the second phase, all the examples that are covered by these rules are removed from the training data. This sequential procedure is repeated until most of the training examples are covered by CAM-HMC rules. In the third phase, consequent generation is determined to show the probability of coverage of each rule for each hierarchical class. Finally, this ratio is applied to classify testing data. Efficiency of this algorithm is displayed by comparing this algorithm with HMC-GA using Precision-Recall curves for three numerical datasets related to protein functions of the Saccharomyces Cerevisiae organism. Mina Moradi Kordmahalleh, Abdollah Homaifar, Dukka B. KC |
BIBE | 3 |
| 2011 | Topology Improves Phylogenetic Motif Functional Site PredictionsabstractPrediction of protein functional sites from sequence-derived data remains an open bioinformatics problem. We have developed a phylogenetic motif (PM) functional site prediction approach that identifies functional sites from alignment fragments that parallel the evolutionary patterns of the family. In our approach, PMs are identified by comparing tree topologies of each alignment fragment to that of the complete phylogeny. Herein, we bypass the phylogenetic reconstruction step and identify PMs directly from distance matrix comparisons. In order to optimize the new algorithm, we consider three different distance matrices and 13 different matrix similarity scores. We assess the performance of the various approaches on a structurally nonredundant data set that includes three types of functional site definitions. Without exception, the predictive power of the original approach outperforms the distance matrix variants. While the distance matrix methods fail to improve upon the original approach, our results are important because they clearly demonstrate that the improved predictive power is based on the topological comparisons. Meaning that phylogenetic trees are a straightforward, yet powerful way to improve functional site prediction accuracy. While complementary studies have shown that topology improves predictions of protein-protein interactions, this report represents the first demonstration that trees improve functional site predictions as well. Dukka B. KC, Dennis R. Livesay |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2008 | Improving position-specific predictions of protein functional sites using phylogenetic motifsabstractMOTIVATION: Accurate computational prediction of protein functional sites is critical to maximizing the utility of recent high-throughput sequencing efforts. Among the available approaches, position-specific conservation scores remain among the most popular due to their accuracy and ease of computation. Unfortunately, high false positive rates remain a limiting factor. Using phylogenetic motifs (PMs), we have developed two combined (conservation + PMs) prediction schemes that significantly improve prediction accuracy. RESULTS: Our first approach, called position-specific MINER (psMINER), rank orders alignment columns by conservation. Subsequently, positions that are also not identified as PMs are excluded from the prediction set. This approach improves prediction accuracy, in a statistically significant way, compared to the underlying conservation scores. Increased accuracy is a general result, meaning improvement is observed over several different conservation scores that span a continuum of complexity. In addition, a hybrid MINER (hMINER) that quantitatively considers both scoring regimes provides further improvement. More importantly, it provides critical insight into the relative importance of phylogeny versus alignment conservation. Both methods outperform other common prediction algorithms that also utilize phylogenetic concepts. Finally, we demonstrate that the presented results are critically sensitive to functional site definition, thus highlighting the need for more complete benchmarks within the prediction community. Dukka B. KC, Dennis R. Livesay |
Bioinform. | 1 |
| 2005 | Clique-based algorithms for protein threading with profiles and constraints
Dukka B. KC, Etsuji Tomita, Jun'ichi Suzuki, Katsuhisa Horimoto, Tatsuya Akutsu |
APBC | 1 |
| 2004 | Protein Side-chain Packing Problem: A Maximum Edge-weight Clique Algorithmic Approach
Dukka B. KC, Tatsuya Akutsu, Etsuji Tomita, Tomokazu Seki |
APBC | 1 |