Prakash Chourasia

dblp:310/1882 · DBLP profile ↗
← Back
13ranked-venue papers
3as first author
13since 2021 · last 2026
0000-0002-1443-2192ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Compression and k-Mer Based Approach for Anticancer Peptide Analysis
abstract
Anti-cancer peptide (ACP) sequence classification is crucial for cancer treatment development. Current neural network approaches achieve high accuracy but require substantial parameters and training data. Recent compression-based methods compress entire sequences, potentially missing fine grained neighboring information critical for classification. We propose a novel approach integrating Gzip compression with an incremental k-mer strategy. Unlike conventional methods, we compress individual k-mers and incrementally build subsequence compressions, preserving amino acid-level context. Using Normalized Compression Distance (NCD) and kernel-based embeddings, our parameter-free method achieves state-of-the-art performance on breast and lung cancer ACP datasets, outperforming deep neural networks and large language models without requiring custom features or pre-trained models. Our approach provides a practical, efficient alternative to computationally intensive methods, proving effective even in low-resource environments.
Sarwan Ali, Tamkanat E. Ali, Prakash Chourasia, Murray Patterson
IEEE Trans. Comput. Biol. Bioinform.3
2026 Nearest Neighbor CCP-Based Molecular Sequence Analysis
abstract
Molecular sequence analysis is crucial for understanding several biological processes, including protein-protein interactions, functional annotation, and disease classification. The large number of sequences and the inherently complicated nature of protein structures make it challenging to analyze such data. Finding patterns and enhancing subsequent research requires the use of dimensionality reduction and feature selection approaches. Recently, a method called Correlated Clustering and Projection (CCP) has been proposed as an effective method for biological sequencing data. The CCP technique remains computationally expensive, despite its effectiveness for sequence visualization. Furthermore, its utility for classifying molecular sequences is still uncertain. To solve these two problems, we present a Nearest-Neighbor Correlated Clustering and Projection (CCP-NN)-based technique for efficiently preprocessing molecular sequence data. To group related molecular sequences and produce representative supersequences, CCP makes use of sequence-to-sequence correlations. As opposed to conventional methods, CCP does not rely on matrix diagonalization, therefore, it can be applied to a range of machine-learning problems. We estimate the density map and compute the correlation using a nearest-neighbor search technique. We perform a molecular sequence classification using CCP and CCP-NN representations to assess the efficacy of our proposed approach. Our findings show that CCP-NN considerably improves classification accuracy and significantly outperforms CCP in computational runtime.
Sarwan Ali, Prakash Chourasia, Bipin Koirala, Murray Patterson
IEEE Trans. Comput. Biol. Bioinform.2
2025 Hist2Vec: A histogram and kernel-based embedding method for molecular sequence analysis
Sarwan Ali, Tamkanat E. Ali, Haris Mansoor, Prakash Chourasia, Murray Patterson
Expert Syst. Appl.4
2024 Gaussian Beltrami-Klein Model for Protein Sequence Classification: A Hyperbolic Approach
Sarwan Ali, Haris Mansoor, Prakash Chourasia, Murray Patterson
ISBRA (1)3
2024 Elliptic geometry-based kernel matrix for improved biological sequence classification
Sarwan Ali, Madiha Shabbir, Haris Mansoor, Prakash Chourasia, Murray Patterson
Knowl. Based Syst.4
2023 Circular Arc Length-Based Kernel Matrix For Protein Sequence Classification
abstract
Biological sequence analysis is crucial in understanding the sequence structure, function, and evolutionary relationships. In traditional methods, using Euclidean distance metrics is common in measuring the similarity between sequence embeddings. However, they fail to capture sequence space’s inherent curvature and spherical nature. Therefore, we explore the application of spherical geometry and distance metric, circular arc length (CAL) based distance for comparing biological sequence embeddings. Spherical geometry is a non-Euclidean geometry that models the surface of a sphere, accounting for its curvature. By leveraging CAL, we can more accurately measure the pairwise distances between bio-sequence embeddings. In this study, we propose the utilization of CAL for comparing bio-sequence embeddings. We develop a function to compute the CAL distance between two sequence embeddings, enabling researchers to accurately measure the similarity between sequences while considering the underlying spherical geometry. Incorporating spherical geometry enables a more comprehensive understanding of the relationships and similarities between biological sequences, improving various downstream tasks such as classification, clustering, and evolutionary analysis. Our experimental evaluation demonstrates the advantages of using a spherical distance metric over Euclidean metrics for bio-sequence analysis. Our proposed CAL-based approach outperforms the Euclidean geometry-based baselines by depicting a huge performance improvement for the protein subcellular location classification task. In our experiments, accuracy is improved by 52.5% and 60. 3% compared to the PWM2Vec and Autoencoder methods, respectively, corresponding to the DT classifier.
Taslim Murad, Sarwan Ali, Prakash Chourasia, Haris Mansoor, Murray Patterson
IEEE Big Data3
2023 T Cell Receptor Protein Sequences and Sparse Coding: A Novel Approach to Cancer Classification
Zahra Tayebi, Sarwan Ali, Prakash Chourasia, Taslim Murad, Murray Patterson
ICONIP (10)3
2023 Empowering Pandemic Response with Federated Learning for Protein Sequence Data Analysis
abstract
Genomics sequencing has become more accessible thanks to advances in genomics technology. We have witnessed this, particularly in the COVID-19 pandemic with massive growth in viral (SARS-CoV-2) protein sequence data generation. Massive data, however, has been underutilized as a result of certain pandemic policies and countermeasures. For political or economic reasons, some wary countries have purposely created this issue, rather than being subject to the natural barriers to the availability of these data. All countries cannot be expected to actively contribute fair data to the scientific community; otherwise, their participation may become passive. We require a strategy to encourage nations across the globe to fairly exchange information on the pandemic situation and data. We propose a federated learning (FL)-based model to address the issue of data privacy and enable real-time surveillance of epidemics. In FL, we train a feed-forward neural network locally and transmit the updated weights to the central server to combine the learning from local data. In this way, a federated learning-based architecture offers data security because there is no need to share the data (the data does not leave the premises; just parameters or weights are communicated) and encourages all countries to contribute fairly to the research. While FL is popular in dealing with image data, in this paper we apply it to bioinformatics, in particular, protein sequence classification. The results show that the FL-based model not only performs better than the centralized, traditional deep learning models (such as CNN, GRU, LSTM, and feed-forward neural network) in terms of predictive accuracy and fairness in data utilization but that all of this can be accomplished while maintaining data privacy. Our findings imply that FL can address the problems of data availability and privacy without compromising results. Employing factual data and tracking the evolution and dissemination of new SARS-CoV-2 lineages in real-time, can aid in the disclosure of verifiable advancements in pandemic studies.
Prakash Chourasia, Zahra Tayebi, Sarwan Ali, Murray Patterson
IJCNN1
2023 PDB2Vec: Using 3D Structural Information for Improved Protein Analysis
Sarwan Ali, Prakash Chourasia, Murray Patterson
ISBRA2
2023 Hist2Vec: Kernel-Based Embeddings for Biological Sequence Classification
Sarwan Ali, Haris Mansoor, Prakash Chourasia, Murray Patterson
ISBRA3
2023 Enhancing t-SNE Performance for Biological Sequencing Data Through Kernel Selection
Prakash Chourasia, Taslim Murad, Sarwan Ali, Murray Patterson
ISBRA1
2022 Hashing2Vec: Fast Embedding Generation for SARS-CoV-2 Spike Sequence Classification
Taslim Murad, Prakash Chourasia, Sarwan Ali, Murray Patterson
ACML2
2022 Informative Initialization and Kernel Selection Improves t-SNE for Biological Sequences
abstract
The t-distributed stochastic neighbor embedding (t-SNE) is a method for interpreting high dimensional (HD) data by mapping each point to a low dimensional (LD) space (usually two-dimensional). It seeks to retain the structure of the data. An important component of the t-SNE algorithm is the initialization procedure, which begins with the random initialization of an LD vector. Points in this initial vector are then updated to minimize the loss function (the KL divergence) iteratively using gradient descent. This leads comparable points to attract one another while pushing dissimilar points apart. We believe that, by default, these algorithms should employ some form of informative initialization. Another essential component of the t-SNE is using a kernel matrix, a similarity matrix comprising the pairwise distances among the sequences. For t-SNE-based visualization, the Gaussian kernel is employed by default in the literature. However, we show that kernel selection can also play a crucial role in the performance of t-SNE.In this work, we assess the performance of t-SNE with various alternative initialization methods and kernels, using four different sets, out of which three are biological sequences (nucleotide, protein, etc.) datasets obtained from various sources, such as the well-known GISAID database for sequences of the SARS-CoV-2 virus. We perform subjective and objective assessments of these alternatives. We use the resulting t-SNE plots and k-ary neighborhood agreement (k-ANA) to evaluate and compare the proposed methods with the baselines. We show that by using different techniques, such as informed initialization and kernel matrix selection, that t-SNE performs significantly better. Moreover, we show that t-SNE also takes fewer iterations to converge faster with more intelligent initialization.
Prakash Chourasia, Sarwan Ali, Murray Patterson
IEEE Big Data1