Mohammad Saifur Rahman 0001

dblp:47/5073-1 · DBLP profile ↗
← Back
17ranked-venue papers
3as first author
8since 2021 · last 2026
0000-0002-9887-4456ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 10 · 7 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 1 since 2021Computer networks · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 PULSAR: Graph-Based Positive Unlabeled Learning with Multi-Stream Adaptive Convolutions for Parkinson's Disease Recognition
abstract
Timely diagnosis of movement disorders like Parkinson’s Disease (PD) improves quality of life. However, access to clinical diagnosis is limited in low-income countries. Here, we present PULSAR, a novel method for classifying individuals with or without PD from webcam-recorded videos of the finger-tapping task used in the Movement Disorder Society—Unified Parkinson’s Disease Rating Scale (MDS-UPDRS). PULSAR was trained and evaluated on data from 382 participants, including 183 self-reported PD patients. We used an adaptive graph convolutional neural network to dynamically learn task-specific spatio-temporal edges and enhanced it with a multi-stream convolution model to capture critical features like finger joint locations, tapping velocity, and acceleration of tapping. As video labels are self-reported, some non-PD labels may be undiagnosed cases. To address this, we used Positive Unlabeled (PU) Learning, which outperformed traditional supervised learning. PULSAR achieved 80.95% accuracy on the validation set and 71.29% mean accuracy (2.49% standard deviation) on an independent test set. We hope PULSAR can aid in accessible PD screening and that these techniques may extend to assessing disorders like ataxia and Huntington’s disease.
Md. Zarif Ul Alam, Asif Azad, Md. Saiful Islam 0013, Mohammed E. Hoque 0001, Mohammad Saifur Rahman 0001
ACM Trans. Comput. Heal.5
2025 A novel loan eligibility prediction model with effective use of data transformation methods
Joydeb Kumar Sana, Mohammad Sohel Rahman, Mohammad Saifur Rahman 0001
Knowl. Based Syst.3
2023 NoVaTeST: identifying genes with location-dependent noise variance in spatial transcriptomics data
abstract
MOTIVATION: Spatial transcriptomics (ST) can reveal the existence and extent of spatial variation of gene expression in complex tissues. Such analyses could help identify spatially localized processes underlying a tissue's function. Existing tools to detect spatially variable genes assume a constant noise variance across spatial locations. This assumption might miss important biological signals when the variance can change across locations. RESULTS: In this article, we propose NoVaTeST, a framework to identify genes with location-dependent noise variance in ST data. NoVaTeST models gene expression as a function of spatial location and allows the noise to vary spatially. NoVaTeST then statistically compares this model to one with constant noise and detects genes showing significant spatial noise variation. We refer to these genes as "noisy genes." In tumor samples, the noisy genes detected by NoVaTeST are largely independent of the spatially variable genes detected by existing tools that assume constant noise, and provide important biological insights into tumor microenvironments. AVAILABILITY AND IMPLEMENTATION: An implementation of the NoVaTeST framework in Python along with instructions for running the pipeline is available at https://github.com/abidabrar-bracu/NoVaTeST.
Mohammed Abid Abrar, Mohammad Kaykobad, Mohammad Saifur Rahman 0001, Md. Abul Hassan Samee
Bioinform.3
2023 Quartet Fiduccia-Mattheyses revisited for larger phylogenetic studies
abstract
MOTIVATION: With the recent breakthroughs in sequencing technology, phylogeny estimation at a larger scale has become a huge opportunity. For accurate estimation of large-scale phylogeny, substantial endeavor is being devoted in introducing new algorithms or upgrading current approaches. In this work, we endeavor to improve the Quartet Fiduccia and Mattheyses (QFM) algorithm to resolve phylogenetic trees of better quality with better running time. QFM was already being appreciated by researchers for its good tree quality, but fell short in larger phylogenomic studies due to its excessively slow running time. RESULTS: We have re-designed QFM so that it can amalgamate millions of quartets over thousands of taxa into a species tree with a great level of accuracy within a short amount of time. Named "QFM Fast and Improved (QFM-FI)", our version is 20 000× faster than the previous version and 400× faster than the widely used variant of QFM implemented in PAUP* on larger datasets. We have also provided a theoretical analysis of the running time and memory requirements of QFM-FI. We have conducted a comparative study of QFM-FI with other state-of-the-art phylogeny reconstruction methods, such as QFM, QMC, wQMC, wQFM, and ASTRAL, on simulated as well as real biological datasets. Our results show that QFM-FI improves on the running time and tree quality of QFM and produces trees that are comparable with state-of-the-art methods. AVAILABILITY AND IMPLEMENTATION: QFM-FI is open source and available at https://github.com/sharmin-mim/qfm_java.
Sharmin Akter Mim, Md. Zarif Ul Alam, Rezwana Reaz, Md. Shamsuzzoha Bayzid, Mohammad Saifur Rahman 0001
Bioinform.5
2023 ScribbleDom: using scribble-annotated histology images to identify domains in spatial transcriptomics data
abstract
MOTIVATION: Spatial domain identification is a very important problem in the field of spatial transcriptomics. The state-of-the-art solutions to this problem focus on unsupervised methods, as there is lack of data for a supervised learning formulation. The results obtained from these methods highlight significant opportunities for improvement. RESULTS: In this article, we propose a potential avenue for enhancement through the development of a semi-supervised convolutional neural network based approach. Named "ScribbleDom", our method leverages human expert's input as a form of semi-supervision, thereby seamlessly combines the cognitive abilities of human experts with the computational power of machines. ScribbleDom incorporates a loss function that integrates two crucial components: similarity in gene expression profiles and adherence to the valuable input of a human annotator through scribbles on histology images, providing prior knowledge about spot labels. The spatial continuity of the tissue domains is taken into account by extracting information on the spot microenvironment through convolution filters of varying sizes, in the form of "Inception" blocks. By leveraging this semi-supervised approach, ScribbleDom significantly improves the quality of spatial domains, yielding superior results both quantitatively and qualitatively. Our experiments on several benchmark datasets demonstrate the clear edge of ScribbleDom over state-of-the-art methods-between 1.82% to 169.38% improvements in adjusted Rand index for 9 of the 12 human dorsolateral prefrontal cortex samples, and 15.54% improvement in the melanoma cancer dataset. Notably, when the expert input is absent, ScribbleDom can still operate, in a fully unsupervised manner like the state-of-the-art methods, and produces results that remain competitive. AVAILABILITY AND IMPLEMENTATION: Source code is available at Github (https://github.com/1alnoman/ScribbleDom) and Zenodo (https://zenodo.org/badge/latestdoi/681572669).
Mohammad Nuwaisir Rahman, Abir Mohammad Turza, Mohammed Abid Abrar, Md. Abul Hassan Samee, Mohammad Saifur Rahman 0001
Bioinform.6
2023 CD-MAWS: An Alignment-Free Phylogeny Estimation Method Using Cosine Distance on Minimal Absent Word Sets
abstract
Multiple sequence alignment has been the traditional and well established approach of sequence analysis and comparison, though it is time and memory consuming. As the scale of sequencing data is increasing day by day, the importance of faster yet accurate alignment-free methods is on the rise. Several alignment-free sequence analysis methods have been established in the literature in recent years, which extract numerical features from genomic data to analyze sequences and also to estimate phylogenetic relationship among genes and species. Minimal Absent Word (MAW) is an effective concept for representing characteristics of a sequence in an alignment-free manner. In this study, we present CD-MAWS, a distance measure based on cosine of the angle between composition vectors constructed using minimal absent words, for sequence analysis in a computationally inexpensive manner. We have benchmarked CD-MAWS using several AFProject datasets, such as Fish mtDNA, E.coli, Plants, Shigella and Yersinia datasets, and found it to perform quite well. Applied on several other biological datasets such as mammal mtDNA, bacterial genomes and viral genomes, CD-MAWS resolved phylogenetic relationships similar to or better than state-of-the-art alignment-free methods such as Mash, Skmer, Co-phylog and kSNP3.
Naser Anjum, Raian Latif Nabil, Rakibul Islam Rafi, Md. Shamsuzzoha Bayzid, Mohammad Saifur Rahman 0001
IEEE ACM Trans. Comput. Biol. Bioinform.5
2021 ADACT: a tool for analysing (dis)similarity among nucleotide and protein sequences using minimal and relative absent words
abstract
MOTIVATION: Researchers and practitioners use a number of popular sequence comparison tools that use many alignment-based techniques. Due to high time and space complexity and length-related restrictions, researchers often seek alignment-free tools. Recently, some interesting ideas, namely, Minimal Absent Words (MAW) and Relative Absent Words (RAW), have received much interest among the scientific community as distance measures that can give us alignment-free alternatives. This drives us to structure a framework for analysing biological sequences in an alignment-free manner. RESULTS: In this application note, we present Alignment-free Dissimilarity Analysis & Comparison Tool (ADACT), a simple web-based tool that computes the analogy among sequences using a varied number of indexes through the distance matrix, species relation list and phylogenetic tree. This tool basically combines absent word (MAW or RAW) computation, dissimilarity measures, species relationship and thus brings all required software in one platform for the ease of researchers and practitioners alike in the field of bioinformatics. We have also developed a restful API. AVAILABILITY AND IMPLEMENTATION: ADACT has been hosted at http://research.buet.ac.bd/ADACT/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Mujtahid Akon, Muntashir Akon, Mohimenul Kabir, Mohammad Saifur Rahman 0001, Mohammad Sohel Rahman
Bioinform.4
2021 wQFM: highly accurate genome-scale species tree estimation from weighted quartets
abstract
MOTIVATION: Species tree estimation from genes sampled from throughout the whole genome is complicated due to the gene tree-species tree discordance. Incomplete lineage sorting (ILS) is one of the most frequent causes for this discordance, where alleles can coexist in populations for periods that may span several speciation events. Quartet-based summary methods for estimating species trees from a collection of gene trees are becoming popular due to their high accuracy and statistical guarantee under ILS. Generating quartets with appropriate weights, where weights correspond to the relative importance of quartets, and subsequently amalgamating the weighted quartets to infer a single coherent species tree can allow for a statistically consistent way of estimating species trees. However, handling weighted quartets is challenging. RESULTS: We propose wQFM, a highly accurate method for species tree estimation from multi-locus data, by extending the quartet FM (QFM) algorithm to a weighted setting. wQFM was assessed on a collection of simulated and real biological datasets, including the avian phylogenomic dataset, which is one of the largest phylogenomic datasets to date. We compared wQFM with wQMC, which is the best alternate method for weighted quartet amalgamation, and with ASTRAL, which is one of the most accurate and widely used coalescent-based species tree estimation methods. Our results suggest that wQFM matches or improves upon the accuracy of wQMC and ASTRAL. AVAILABILITY AND IMPLEMENTATION: Datasets studied in this article and wQFM (in open-source form) are available at https://github.com/Mahim1997/wQFM-2020. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Mahim Mahbub, Zahin Wahab, Rezwana Reaz, Mohammad Saifur Rahman 0001, Md. Shamsuzzoha Bayzid
Bioinform.4
2020 A Multi-objective Metaheuristic Approach for Accurate Species Tree Estimation
abstract
Species tree estimation from multi-locus data is complicated as biological processes can result in different loci having different evolutionary histories. Incomplete lineage sorting (ILS), modeled by the multi-species coalescent (MSC), is considered to be a dominant cause for gene tree incongruence. Various optimization criteria (e.g., quartet score, pseudo-likelihood, etc.) are statistically consistent under the MSC model, meaning that they return the true species tree with high probability given sufficiently large numbers of accurate gene trees. However, the number of genes is limited and estimating highly accurate gene trees is difficult. Therefore, even popular methods, optimizing a particular criterion, may fail to reconstruct highly accurate trees under practical model conditions with limited numbers of genes and in the presence of gene tree estimation error. In this study, we advocate multi-objective optimization to tackle this challenge. Particularly, we present a multi-objective metaheuristic (i.e., SNOGA), a modified version of the popular NSGAII, which combines various optimization criteria to find a suitable search space containing highly accurate species trees.
Muhammad Ali Nayeem, Md. Shamsuzzoha Bayzid, Sakshar Chakravarty, Mohammad Saifur Rahman 0001, Mohammad Sohel Rahman
BIBE4
2020 SAINT: self-attention augmented inception-inside-inception network improves protein secondary structure prediction
abstract
MOTIVATION: Protein structures provide basic insight into how they can interact with other proteins, their functions and biological roles in an organism. Experimental methods (e.g. X-ray crystallography and nuclear magnetic resonance spectroscopy) for predicting the secondary structure (SS) of proteins are very expensive and time consuming. Therefore, developing efficient computational approaches for predicting the SS of protein is of utmost importance. Advances in developing highly accurate SS prediction methods have mostly been focused on 3-class (Q3) structure prediction. However, 8-class (Q8) resolution of SS contains more useful information and is much more challenging than the Q3 prediction. RESULTS: We present SAINT, a highly accurate method for Q8 structure prediction, which incorporates self-attention mechanism (a concept from natural language processing) with the Deep Inception-Inside-Inception network in order to effectively capture both the short- and long-range interactions among the amino acid residues. SAINT offers a more interpretable framework than the typical black-box deep neural network methods. Through an extensive evaluation study, we report the performance of SAINT in comparison with the existing best methods on a collection of benchmark datasets, namely, TEST2016, TEST2018, CASP12 and CASP13. Our results suggest that self-attention mechanism improves the prediction accuracy and outperforms the existing best alternate methods. SAINT is the first of its kind and offers the best known Q8 accuracy. Thus, we believe SAINT represents a major step toward the accurate and reliable prediction of SSs of proteins. AVAILABILITY AND IMPLEMENTATION: SAINT is freely available as an open-source project at https://github.com/SAINTProtein/SAINT.
Mostofa Rafid Uddin, Sazan Mahbub, Mohammad Saifur Rahman 0001, Md. Shamsuzzoha Bayzid
Bioinform.3
2020 CRISPRpred(SEQ): a sequence-based method for sgRNA on target activity prediction using traditional machine learning
abstract
BACKGROUND: The latest works on CRISPR genome editing tools mainly employs deep learning techniques. However, deep learning models lack explainability and they are harder to reproduce. We were motivated to build an accurate genome editing tool using sequence-based features and traditional machine learning that can compete with deep learning models. RESULTS: In this paper, we present CRISPRpred(SEQ), a method for sgRNA on-target activity prediction that leverages only traditional machine learning techniques and hand-crafted features extracted from sgRNA sequences. We compare the results of CRISPRpred(SEQ) with that of DeepCRISPR, the current state-of-the-art, which uses a deep learning pipeline. Despite using only traditional machine learning methods, we have been able to beat DeepCRISPR for the three out of four cell lines in the benchmark dataset convincingly (2.174%, 6.905% and 8.119% improvement for the three cell lines). CONCLUSION: CRISPRpred(SEQ) has been able to convincingly beat DeepCRISPR in 3 out of 4 cell lines. We believe that by exploring further, one can design better features only using the sgRNA sequences and can come up with a better method leveraging only traditional machine learning algorithms that can fully beat the deep learning models.
Ali Haisam Muhammad Rafid, Md. Toufikuzzaman, Mohammad Saifur Rahman 0001, Mohammad Sohel Rahman
BMC Bioinform.3
2019 Antigenic: An improved prediction model of protective antigens
abstract
An antigen is a protein capable of triggering an effective immune system response. Protective antigens are the ones that can invoke specific and enhanced adaptive immune response to subsequent exposure to the specific pathogen or related organisms. Such proteins are therefore of immense importance in vaccine preparation and drug design. However, the laboratory experiments to isolate and identify antigens from a microbial pathogen are expensive, time consuming and often unsuccessful. This is why Reverse Vaccinology has become the modern trend of vaccine search, where computational methods are first applied to predict protective antigens or their determinants, known as epitopes. In this paper, we propose a novel, accurate computational model to identify protective antigens efficiently. Our model extracts features directly from the protein sequences, without any dependence on functional domain or structural information. After relevant features are extracted, we have used Random Forest algorithm to rank the features. Then Recursive Feature Elimination (RFE) and minimum redundancy maximum relevance (mRMR) criterion were applied to extract an optimal set of features. The learning model was trained using Random Forest algorithm. Named as Antigenic, our proposed model demonstrates superior performance compared to the state-of-the-art predictors on a benchmark dataset. Antigenic achieves accuracy, sensitivity and specificity values of 78.04%, 78.99% and 77.08% in 10-fold cross-validation testing respectively. In jackknife cross-validation, the corresponding scores are 80.03%, 80.90% and 79.16% respectively. The source code of Antigenic, along with relevant dataset and detailed experimental results, can be found at https://github.com/srautonu/AntigenPredictor. A publicly accessible web interface has also been established at: http://antigenic.research.buet.ac.bd.
Mohammad Saifur Rahman 0001, Md. Khaledur Rahman, Sanjay Saha, Mohammad Kaykobad, Mohammad Sohel Rahman
Artif. Intell. Medicine1
2018 isGPT: An optimized model to identify sub-Golgi protein types using SVM and Random Forest based feature selection
Mohammad Saifur Rahman 0001, Md. Khaledur Rahman, Mohammad Kaykobad, Mohammad Sohel Rahman
Artif. Intell. Medicine1
2018 Using Adaptive Heartbeat Rate on Long-Lived TCP Connections
abstract
In this paper, we propose techniques for dynamically adjusting heartbeat or keep-alive interval of long-lived TCP connections, particularly the ones that are used in push notification service in mobile platforms. When a device connects to a server using TCP, often times the connection is established through some sort of middle-box, such as NAT, proxy, firewall, and so on. When such a connection is idle for a long time, it may get torn down due to binding timeout of the middle-box. To keep the connection alive, the client device needs to send keep-alive packets through the connection when it is otherwise idle. To reduce resource consumption, the keep-alive packet should preferably be sent at the farthest possible time within the binding timeout. Due to varied settings of different network equipments, the binding timeout will not be identical in different networks. Hence, the heartbeat rate used in different networks should be changed dynamically. We propose a set of iterative probing techniques, namely binary, exponential, and composite search, that detect the middle-box binding timeout with varying degree of accuracy; and in the process, keeps improving the keep-alive interval used by the client device. We also analytically derive performance bounds of these techniques. To the best of our knowledge, ours is the first work that systematically studies several techniques to dynamically improve keep-alive interval. To this end, we run experiments in simulation as well as make a real implementation on android to demonstrate the proof-of-concept of the proposed schemes.
Mohammad Saifur Rahman 0001, Md. Yusuf Sarwar Uddin, Tahmid Hasan, Mohammad Sohel Rahman, Mohammad Kaykobad
IEEE/ACM Trans. Netw.1
2015 Order preserving pattern matching revisited
Md. Mahbubul Hasan, A. S. M. Shohidull Islam, Mohammad Saifur Rahman 0001, Mohammad Sohel Rahman
Pattern Recognit. Lett.3
2014 Order Preserving Prefix Tables
Md. Mahbubul Hasan, A. S. M. Shohidull Islam, Mohammad Saifur Rahman 0001, Mohammad Sohel Rahman
SPIRE3
2006 Drawing lines by uniform packing
Asif-ul Haque, Mohammad Saifur Rahman 0001, Mehedi Bakht, Mohammad Kaykobad
Comput. Graph.2