Anup Kumar Halder

dblp:209/1868 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
5since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 5 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2025 DensePPI-2: a bio-inspired update for sequence-based PPI prediction leveraging mutation rates
abstract
Identifying interactions between two or more proteins is crucial as it helps understand living organisms' cellular behaviour and the underlying molecular mechanisms of various diseases. However, most existing computational algorithms in the field model this as a binary interaction between any two proteins, instead of conserving the evolutionary regions of protein function and interactions. This is important for predicting potential interaction sites, vital for drug design, target identification, and understanding disease progression and pathogenic mechanisms. Position-aware encoding provides a way to incorporate the order of amino acids in a protein sequence into the model, thus capturing folding patterns, leading to more accurate predictions of protein structures and their interactions. This is crucial because the sequence order can affect the structure and function of proteins. The proposed DensePPI-2 model is a novel bio-inspired substitution matrix-based sequence encoding with deep learning for identifying interacting protein pairs. It demonstrates an AUC of 97.13% on the S. cerevisiae dataset, improving by 1.4% over the best existing methods. Furthermore, DensePPI-2 outperforms recent sequence-based approaches on the human benchmark dataset, addressing the complexities of protein-protein interaction test classes. DensePPI-2 has been successfully applied for (i) identifying pathogen-host interactions and (ii) predicting near-residue-level interaction, even though the model was not trained on residue-level data. The enhanced performance on diverse test sets proves the efficiency of the bio-inspired sequence-to-image colour encoding strategy using the substitution matrices. The dataset and the developed models are available at https://github.com/CMATERJU-BIOINFO/DensePPI-2 for academic use only.
Tapas Chakraborty, Debarati Paul, Aanzil Akram Halsana, Anup Kumar Halder, Subhadip Basu, Tapabrata Chakraborti
Briefings Bioinform.4
2025 FuzzyPPI: Large-Scale Interaction of Human Proteome at Fuzzy Semantic Space
abstract
Large-scale protein-protein interaction (PPI) network of an organism provides key insights into its cellular and molecular functionalities, signaling pathways and underlying disease mechanisms. For any organism, the total unexplored protein interactions significantly outnumbers all known positive and negative interactions. For Human, all known PPI datasets contain only ∼ 5.61 million positive and ∼ 0.76 million negative interactions, which is ∼ 3.1% of potential interactions. We have implemented a distributed algorithm in Apache Spark that evaluates a Human PPI network of ∼ 180 million potential interactions resulting from 18 994 reviewed proteins for which Gene Ontology (GO) annotations are available. The computed scores have been validated against state-of-the-art methods on benchmark datasets.FuzzyPPI performed significantly better with an average F1 score of 0.62 compared to GOntoSim (0.39), GOGO (0.38), and Wang (0.38) when tested with the Gold Standard PPI Dataset. The resulting scores are published with a web server for non-commercial use athttp://fuzzyppi.mimuw.edu.pl/. Moreover, conventional PPI prediction methods produce binary results, but in fact this is just a simplification as PPIs have strengths or probabilities and recent studies show that protein binding affinities may prove to be effective in detecting protein complexes, disease association analysis, signaling network reconstruction, etc. Keeping these in mind, our algorithm is based on a fuzzy semantic scoring function and produces probabilities of interaction.
Anup Kumar Halder, Soumyendu Sekhar Bandyopadhyay, Witold Jedrzejewski, Subhadip Basu, Jacek Sroka
IEEE Trans. Big Data1
2025 PCPredG: Protein Complex Prediction Using Graphlet Features
abstract
Proteins interact with other proteins and bio-molecules to form a complex and execute key biological functions in a living organism, and respond to several environmental signals. Designing efficient predictive models for protein complexes is a challenging task with limited coverage in the contemporary literature. With this motivation, we have developed a novel method, PCPredG, for 3-node protein complex prediction from PPI networks using 5-node graphlet features. CORUM protein complex repository has been used to curate positive and negative data samples with the help of MCODE and MCL clustering algorithms. During experiments, Random Forest(RF) and SVM classifiers are trained with 1000 positive 3-node complexes in 10-fold cross-validation setup and with 1:1 to 1:10 positive-negative proportions. In parallel, we have implemented the state-of-the-art GCN with polarised message-passing, GAT and an ensemble of GCN and GAT in both balanced and imbalanced setups. We also introduced a 10-fold quality consensus on the hold-out set across all the experiments. We have achieved the best performances with the RF classifier in both balanced and imbalanced experiments.
Rupali Patua, Anup Kumar Halder, Soma Dasgupta, Piyali Chatterjee, Mita Nasipuri, Subhadip Basu
IEEE Trans. Comput. Biol. Bioinform.2
2023 RUBic: rapid unsupervised biclustering
abstract
Biclustering of biologically meaningful binary information is essential in many applications related to drug discovery, like protein-protein interactions and gene expressions. However, for robust performance in recently emerging large health datasets, it is important for new biclustering algorithms to be scalable and fast. We present a rapid unsupervised biclustering (RUBic) algorithm that achieves this objective with a novel encoding and search strategy. RUBic significantly reduces the computational overhead on both synthetic and experimental datasets shows significant computational benefits, with respect to several state-of-the-art biclustering algorithms. In 100 synthetic binary datasets, our method took [Formula: see text] s to extract 494,872 biclusters. In the human PPI database of size [Formula: see text], our method generates 1840 biclusters in [Formula: see text] s. On a central nervous system embryonic tumor gene expression dataset of size 712,940, our algorithm takes 101 min to produce 747,069 biclusters, while the recent competing algorithms take significantly more time to produce the same result. RUBic is also evaluated on five different gene expression datasets and shows significant speed-up in execution time with respect to existing approaches to extract significant KEGG-enriched bi-clustering. RUBic can operate on two modes, base and flex, where base mode generates maximal biclusters and flex mode generates less number of clusters and faster based on their biological significance with respect to KEGG pathways. The code is available at ( https://github.com/CMATERJU-BIOINFO/RUBic ) for academic use only.
Brijesh Kumar Sriwastava, Anup Kumar Halder, Subhadip Basu, Tapabrata Chakraborti
BMC Bioinform.2
2022 JUPPI: A Multi-Level Feature Based Method for PPI Prediction and a Refined Strategy for Performance Assessment
abstract
Over the years, several methods have been proposed for the computational PPI prediction with different performance evaluation strategies. While attempting to benchmark performance scores, most of these methods often suffer with ill-treated cross-validation strategies, adhoc selection of positive/negative samples etc. To address these issues, in our proposed multi-level feature based PPI prediction approach (JUPPI), using sequence, domain and GO information as features, a refined evaluation strategy has been introduced. During the evaluation process, we first extract high quality negative data using three-stage filtering, and then introduce a pair-input based cross validation strategy with three difficulty levels for test-set predictions. Our proposed evaluation strategy reduces the component-level overlapping issue in test sets. Performance of JUPPI is compared with those of the state-of-the-art approaches in this domain and tested on six independent PPI datasets. In almost all the datasets, JUPPI outperforms the state-of-the-art not only at human proteome level for PPI prediction, but also for prediction of interactors for intrinsic disordered human proteins. https://figshare.com/projects/JUPPI_A_Multi-level_Feature_Based_Method_for_PPI_Prediction_and_a_Refined_Strategy_for_Performance_Assessment/81656 JUPPI tool and the developed datasets (JUPPId) are available in public domain for academic use along with supplementary materials, which can be found on the Computer Society Digital Library at http://doi.ieeecomputersociety.org/10.1109/TCBB.2020.3004970.
Anup Kumar Halder, Soumyendu Sekhar Bandyopadhyay, Piyali Chatterjee, Mita Nasipuri, Dariusz Plewczynski, Subhadip Basu
IEEE ACM Trans. Comput. Biol. Bioinform.1
2019 Bio-inspired cryptosystem with DNA cryptography and neural networks
Sayantani Basu, Marimuthu Karuppiah, Mita Nasipuri, Anup Kumar Halder, Niranchana Radhakrishnan
J. Syst. Archit.4
2019 3gClust: Human Protein Cluster Analysis
abstract
We present a human protein cluster analysis by combining: 1) n-gram based amino acid frequency features, 2) optimal feature selection, 3) hierarchical clustering, and 4) advanced partitioning techniques. Our method qualitatively and quantitatively groups proteins with increasing sequence similarity into similar clusters by calculating the frequency model of amino acids using n-grams. We experiment with n = 1, i.e., unigrams, n = 2, i.e., bigrams, and finally n = 3, i.e., trigrams for optimal selection of features to design the 3gClust algorithm. The benchmarking results on 20,105 manually curated human proteins show that 3gClust ensures better cluster compactness in the case of proteins with similar functional groups, biological processes, structural alignment, and shared domains (e.g., aquaporins, keratins). Quantitative analysis of non singleton clusters shows significant improvement in their compactness in comparison to other state-of-the art methodologies. 3gClust is available at https://sites.google.com/site/bioinfoju/projects/3gclust for academic use along with supplementary materials, which can be found on the Computer Society Digital Library at http://doi.ieeecomputersociety.org/10.1109/TCBB.2018.2840996, and datasets.
Anup Kumar Halder, Piyali Chatterjee, Mita Nasipuri, Dariusz Plewczynski, Subhadip Basu
IEEE ACM Trans. Comput. Biol. Bioinform.1