EDBT 2026 Demo / reviewers in the wild / expert
Mohammad Sohel Rahman
dblp:r/MohammadSohelRahman · also M. Sohel Rahman
· DBLP profile ↗
86ranked-venue papers
7as first author
23since 2021 · last 2026
0000-0001-9419-6478ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Theory of computation · 40 · 5 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 21 · 2 first-author · 10 since 2021Artificial intelligence and machine learning · 14 · 6 since 2021Databases, data management, data science and information retrieval · 12 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Computer networks · 2Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ORANGE: a machine learning approach for modeling tissue-specific aging from transcriptomic dataabstractDespite aging being a fundamental biological process that profoundly influences health and disease, the interplay between tissue-specific aging and mortality remains underexplored. This study applies machine learning on GTEx transcriptomic data to model tissue-specific biological ages across 12 different types of tissues and introduces an age-gap metric to quantify deviations from the chronological age. We use several modeling techniques optimized with three feature selection strategies: Pearson correlation, age-related differentially expressed genes, and tissue-enriched genes (expressed at least four-fold higher in a specific tissue). Among these, Pearson correlation combined with elastic net regression yields the best performance, with models achieving an average root mean squared error of 6.44 years and an R2 of 0.64. To quantify deviations from chronological age relative to the population, we train neural networks to regress predicted ages against chronological ages, and subtract their outputs from the predicted ages to calculate a metric that we call the age-gap. Age-gap statistics reveal significant tissue-specific aging patterns, identifying extreme agers and correlations between extreme aging and mortality. About 20% of subjects are found to exhibit extreme aging in one tissue, while 1% show multi-organ aging. Further analysis reveals that accelerated aging in specific tissues correlates with greater risk of death from illness. These findings greatly emphasize the role of transcriptomics in aging research and its implications for health and longevity. Wasif Jalal, Mubasshira Musarrat, Md. Abul Hassan Samee, Mohammad Sohel Rahman |
Briefings Bioinform. | 4 |
| 2025 | Bridging the Last Mile: Unpacking the Rural Digital Divide in BangladeshabstractPeer Reviewed Rayhan Rashed, Muhammad Masroor Ali, Sadia Sharmin, Md Shariful Islam Bhuyan, Muhammad Abdullah Adnan, Anindya Iqbal, Md Shohrab Hossain, Mohammad Sohel Rahman, A. B. M. Alim Al Islam |
COMPASS | 8 |
| 2025 | Leveraging Complementary Attention Maps in Vision Transformers for OCT Image AnalysisabstractOptical Coherence Tomography (OCT) scan yields all possible cross-section images of a retina for detecting biomarkers linked to optical defects. Due to the high volume of data generated, an automated and reliable biomarker detection pipeline is necessary as a primary screening stage.We outline our new state-of-the-art pipeline for identifying biomarkers from OCT scans. In collaboration with trained ophthalmologists, we identify local and global structures in biomarkers. Through a comprehensive and systematic review of existing vision architectures, we evaluate different convolution and attention mechanisms for biomarker detection. We find that MaxViT, a hybrid vision transformer combining convolution layers with strided attention, is better suited for local feature detection, while EVA-02, a standard vision transformer leveraging pure attention and large-scale knowledge distillation, excels at capturing global features. We ensemble the predictions of both models to achieve first place in the IEEE Video and Image Processing Cup 2023 competition on OCT biomarker detection, achieving a patient-wise F1 score of 0.8527 in the final phase of the competition, scoring 3.8% higher than the next best solution. Finally, we used knowledge distillation to train a single MaxViT to outperform our ensemble at a fraction of the computation cost. Haz Sameen Shahgir, Tanjeem Azwad Zaman, Khondker Salman Sayeed, Md. Asif Haider, Sheikh Saifur Rahman Jony, Mohammad Sohel Rahman |
ICIP | 6 |
| 2025 | Revisiting motif finding: do bi-objective metaheuristics surpass single-objective metaheuristics?abstractBACKGROUND: The discovery of DNA motifs is essential for studying gene expression and function in many biological systems. Most existing algorithms for motif detection rely on a single optimization criterion or objective function. This study formulates motif finding as a bi-objective optimization problem and investigates whether multi-objective metaheuristics offer potential advantages over single-objective approaches. RESULTS: We developed four variants of the Non-dominated Sorting Genetic Algorithm II (NSGA-II) incorporating simple, problem-specific genetic operators. Experiments on six benchmark datasets from three organisms demonstrate that our bi-objective approach significantly outperforms the state-of-the-art Artificial Bee Colony (ABC) metaheuristic. Remarkably, NSGA-II-PMC achieved superior performance over ABC using 6 times fewer fitness evaluations, highlighting its computational efficiency. The synergistic combination of problem-specific operators proved essential, with individual operators showing limited effectiveness compared to their joint application. CONCLUSIONS: Our findings question the common belief that single-objective metaheuristics are better suited for combinatorial problems like motif finding. The bi-objective formulation helps maintain diversity and avoid premature convergence, even with partially correlated objectives, resulting in better solutions than those obtained through dedicated single-objective optimization. Simple, interpretable problem-specific adaptations can yield substantial performance gains over sophisticated alternatives. These results suggest that bi-objective approaches may provide more robust and computationally efficient solutions for DNA motif discovery, opening new research directions in bioinformatics. Muhammad Ali Nayeem, Shehab S. Ahmed, Suliman Aladhadh, Mohammad Sohel Rahman |
BMC Bioinform. | 4 |
| 2025 | Privacy-preserving customer churn prediction model in the context of telecommunication industry
Joydeb Kumar Sana, Mohammad Sohel Rahman, M. Saifur Rahman |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | A novel loan eligibility prediction model with effective use of data transformation methods
Joydeb Kumar Sana, Mohammad Sohel Rahman, Mohammad Saifur Rahman 0001 |
Knowl. Based Syst. | 2 |
| 2025 | ConVerSum: A Contrastive Learning-Based Approach for Data-Scarce Solution of Cross-Lingual Summarization Beyond Direct EquivalentsabstractCross-lingual summarization (CLS) is a sophisticated branch in Natural Language Processing that demands models to accurately translate and summarize articles from different source languages. Despite the improvement of the subsequent studies, this area still needs data-efficient solutions along with effective training methodologies. To the best of our knowledge, there is no feasible solution for CLS when there is no available high-quality CLS data. In this article, we propose a novel data-efficient approach, ConVerSum , for CLS leveraging the power of con trastive learning, generating ver satile candidate sum maries in different languages based on the given source document and contrasting these summaries with reference summaries concerning the given documents. After that, we train the model with a contrastive ranking loss. Then, we rigorously evaluate the proposed approach against current methodologies and compare it to powerful Large Language Models (LLMs)—Gemini, GPT 3.5, and GPT-4o—proving our model performs better for low-resource languages’ CLS. These findings represent a substantial improvement in the area, opening the door to more efficient and accurate cross-lingual summarizing techniques. Sanzana Karim Lora, Mohammad Sohel Rahman, Rifat Shahriyar |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2024 | PRIEST: predicting viral mutations with immune escape capability of SARS-CoV-2 using temporal evolutionary informationabstractThe dynamic evolution of the severe acute respiratory syndrome coronavirus 2 virus is primarily driven by mutations in its genetic sequence, culminating in the emergence of variants with increased capability to evade host immune responses. Accurate prediction of such mutations is fundamental in mitigating pandemic spread and developing effective control measures. This study introduces a robust and interpretable deep-learning approach called PRIEST. This innovative model leverages time-series viral sequences to foresee potential viral mutations. Our comprehensive experimental evaluations underscore PRIEST's proficiency in accurately predicting immune-evading mutations. Our work represents a substantial step in utilizing deep-learning methodologies for anticipatory viral mutation analysis and pandemic response. Gourab Saha, Shashata Sawmya, Arpita Saha, Md Ajwad Akil, Sadia Tasnim, Mohammad Sohel Rahman |
Briefings Bioinform. | 7 |
| 2024 | CRISPR-DIPOFF: an interpretable deep learning approach for CRISPR Cas-9 off-target predictionabstractCRISPR Cas-9 is a groundbreaking genome-editing tool that harnesses bacterial defense systems to alter DNA sequences accurately. This innovative technology holds vast promise in multiple domains like biotechnology, agriculture and medicine. However, such power does not come without its own peril, and one such issue is the potential for unintended modifications (Off-Target), which highlights the need for accurate prediction and mitigation strategies. Though previous studies have demonstrated improvement in Off-Target prediction capability with the application of deep learning, they often struggle with the precision-recall trade-off, limiting their effectiveness and do not provide proper interpretation of the complex decision-making process of their models. To address these limitations, we have thoroughly explored deep learning networks, particularly the recurrent neural network based models, leveraging their established success in handling sequence data. Furthermore, we have employed genetic algorithm for hyperparameter tuning to optimize these models' performance. The results from our experiments demonstrate significant performance improvement compared with the current state-of-the-art in Off-Target prediction, highlighting the efficacy of our approach. Furthermore, leveraging the power of the integrated gradient method, we make an effort to interpret our models resulting in a detailed analysis and understanding of the underlying factors that contribute to Off-Target predictions, in particular the presence of two sub-regions in the seed region of single guide RNA which extends the established biological hypothesis of Off-Target effects. To the best of our knowledge, our model can be considered as the first model combining high efficacy, interpretability and a desirable balance between precision and recall. Md. Toufikuzzaman, Md. Abul Hassan Samee, Mohammad Sohel Rahman |
Briefings Bioinform. | 3 |
| 2023 | Forward Diffusion Guided Reconstruction as a Multi-Modal Multi-Task Learning SchemeabstractMulti-Modal MRI images offer various perspectives for identifying regions of interest in the brain. Previous studies have successfully utilized Deep Learning methods in tasks, such as, segmentation and classification of MRI images, but the proper utilization and integration of Multi-Modality is still an open area of study. Some studies use Multi-Task training scheme which utilizes an auxiliary task like reconstruction for better performance. This paper proposes a novel Multi-Task learning scheme that utilizes different modalities of MRI images to improve brain region segmentation and classification, where the forward diffusion process and a time projection module is used to incorporate a guided reconstruction task. Our experimental results show that the proposed Multi-Task learning strategy outperforms the vanilla Single-Task training scheme by 2.4% in segmentation and 2.7% in classification tasks. Najibul Haque Sarker, Mohammad Sohel Rahman |
ICIP | 2 |
| 2023 | Succinylated lysine residue prediction revisitedabstractLysine succinylation is a kind of post-translational modification (PTM) that plays a crucial role in regulating the cellular processes. Aberrant succinylation may cause inflammation, cancers, metabolism diseases and nervous system diseases. The experimental methods to detect succinylation sites are time-consuming and costly. This thus calls for computational models with high efficacy, and attention has been given in the literature to develop such models, albeit with only moderate success in the context of different evaluation metrics. One crucial aspect in this context is the biochemical and physicochemical properties of amino acids, which appear to be useful as features for such computational predictors. However, some of the existing computational models did not use the biochemical and physicochemical properties of amino acids. In contrast, some others used them without considering the inter-dependency among the properties. The combinations of biochemical and physicochemical properties derived through our optimization process achieve better results than the results achieved by combining all the properties. We propose three deep learning architectures: CNN+Bi-LSTM (CBL), Bi-LSTM+CNN (BLC) and their combination (CBL_BLC). We find that CBL_BLC outperforms the other two. Ensembling of different models successfully improves the results. Notably, tuning the threshold of the ensemble classifiers further improves the results. Upon comparing our work with other existing works on two datasets, we successfully achieve better sensitivity and specificity by varying the threshold value. Shehab S. Ahmed, Zaara T. Rifat, M. Saifur Rahman, Mohammad Sohel Rahman |
Briefings Bioinform. | 4 |
| 2023 | RamanNet: a generalized neural network architecture for Raman spectrum analysisabstractAbstract Raman spectroscopy provides a vibrational profile of the molecules and thus can be used to uniquely identify different kinds of materials. This sort of molecule fingerprinting has thus led to the widespread application of Raman spectrum in various fields like medical diagnosis, forensics, mineralogy, bacteriology, virology, etc. Despite the recent rise in Raman spectra data volume, there has not been any significant effort in developing generalized machine learning methods targeted toward Raman spectra analysis. We examine, experiment, and evaluate existing methods and conjecture that neither current sequential models nor traditional machine learning models are satisfactorily sufficient to analyze Raman spectra. Both have their perks and pitfalls; therefore, we attempt to mix the best of both worlds and propose a novel network architecture RamanNet. RamanNet is immune to the invariance property in convolutional neural networks (CNNs) and at the same time better than traditional machine learning models for the inclusion of sparse connectivity. This has been achieved by incorporating shifted multi-layer perceptrons (MLP) at the earlier levels of the network to extract significant features across the entire spectrum, which are further refined by the inclusion of triplet loss in the hidden layers. Our experiments on 4 public datasets demonstrate superior performance over the much more complex state-of-the-art methods, and thus, RamanNet has the potential to become the de facto standard in Raman spectra data analysis. Nabil Ibtehaz, Muhammad E. H. Chowdhury, Amith Khandakar, Serkan Kiranyaz, Mohammad Sohel Rahman, Susu M. Zughaier |
Neural Comput. Appl. | 5 |
| 2022 | QT-GILD: Quartet Based Gene Tree Imputation Using Deep Learning Improves Phylogenomic Analyses Despite Missing Data
Sazan Mahbub, Shashata Sawmya, Arpita Saha, Rezwana Reaz, Mohammad Sohel Rahman, Md. Shamsuzzoha Bayzid |
RECOMB | 5 |
| 2022 | Characterization of intrinsically disordered regions in proteins informed by human genetic diversityabstractAll proteomes contain both proteins and polypeptide segments that don't form a defined three-dimensional structure yet are biologically active-called intrinsically disordered proteins and regions (IDPs and IDRs). Most of these IDPs/IDRs lack useful functional annotation limiting our understanding of their importance for organism fitness. Here we characterized IDRs using protein sequence annotations of functional sites and regions available in the UniProt knowledgebase ("UniProt features": active site, ligand-binding pocket, regions mediating protein-protein interactions, etc.). By measuring the statistical enrichment of twenty-five UniProt features in 981 IDRs of 561 human proteins, we identified eight features that are commonly located in IDRs. We then collected the genetic variant data from the general population and patient-based databases and evaluated the prevalence of population and pathogenic variations in IDPs/IDRs. We observed that some IDRs tolerate 2 to 12-times more single amino acid-substituting missense mutations than synonymous changes in the general population. However, we also found that 37% of all germline pathogenic mutations are located in disordered regions of 96 proteins. Based on the observed-to-expected frequency of mutations, we categorized 34 IDRs in 20 proteins (DDX3X, KIT, RB1, etc.) as intolerant to mutation. Finally, using statistical analysis and a machine learning approach, we demonstrate that mutation-intolerant IDRs carry a distinct signature of functional features. Our study presents a novel approach to assign functional importance to IDRs by leveraging the wealth of available genetic data, which will aid in a deeper understating of the role of IDRs in biological processes and disease mechanisms. Shehab S. Ahmed, Zaara T. Rifat, Ruchi Lohia, Arthur J. Campbell, A. Keith Dunker, Mohammad Sohel Rahman, Sumaiya Iqbal |
PLoS Comput. Biol. | 6 |
| 2022 | Computing the longest common almost-increasing subsequence
Mohammad Tawhidul Hasan Bhuiyan, Muhammad Rashed Alam, Mohammad Sohel Rahman |
Theor. Comput. Sci. | 3 |
| 2022 | Multiobjective Formulation of Multiple Sequence Alignment for Phylogeny InferenceabstractMultiple sequence alignment (MSA) is a preliminary task for estimating phylogenies. It is used for homology inference among the sequences of a set of species. Generally, the MSA task is handled as a single-objective optimization process. The alignments computed under one criterion may be different from the alignments generated by other criteria, inferring discordant homologies and thus leading to different hypothesized evolutionary histories relating the sequences. The multiobjective (MO) formulation of MSA has recently been advocated by several researchers, to address this issue. An MO approach independently optimizes multiple (often conflicting) objective functions at the same time and outputs a set of competitive alignments. However, no conceptual or experimental rational from a real-world application perspective has been reported so far for any MO formulation of MSA. This article work investigates the impact of MO formulation in the context of an important scientific problem, namely, phylogeny estimation. Employing popular evolutionary MO algorithms, we show that: 1) trees inferred based on alignments produced by the existing MSA methods used in practice are substantially worse in quality than the trees inferred based on the alignment's output by an MO algorithm and 2) even high-quality alignments (according to popular measures available in the literature) may fail to achieve acceptable accuracy in generating phylogenetic trees. Thus, we essentially ask the following natural question: "can a phylogeny-aware (i.e., application-aware) metric guide in selecting appropriate MO formulations to ensure better phylogeny estimation?" Here, we report a carefully designed extensive experimental study that positively answers this question. Muhammad Ali Nayeem, Md. Shamsuzzoha Bayzid, Atif Rahman 0001, Rifat Shahriyar, Mohammad Sohel Rahman |
IEEE Trans. Cybern. | 5 |
| 2022 | Neural Network-Based Undersampling TechniquesabstractMachine learning models have gained popularity nowadays for their potential to solve real-life issues when trained on pertinent data. In many cases, the real-life data are class imbalanced and hence the corresponding machine learning models trained on the data tend to perform poorly on metrics like precision, recall, AUC, F1, and G-mean score. Since class imbalance issue poses serious challenges to the performance of trained models, a multitude of research works have addressed this issue. Two common data-based sampling techniques have mostly been proposed-undersampling the data of the majority class and oversampling the data of the minority class. In this article, we focus on the former approach. We propose two novel algorithms that employ neural network-based approaches to remove majority samples that are found to reside in the vicinity of the minority samples, thereby undersampling the former to remove (or alleviate) the imbalance issue. We delineate the proposed algorithms and then test the proposed algorithms on some publicly available imbalanced datasets. We then compare the performance of our proposed algorithms to other popular undersampling algorithms. Finally, we conclude that our proposed algorithms outperform most of the existing undersampling approaches on most performance metrics. Md. Adnan Arefeen, Sumaiya Tabassum Nimi, Mohammad Sohel Rahman |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2021 | SICaRiO: short indel call filtering with boostingabstractDespite impressive improvement in the next-generation sequencing technology, reliable detection of indels is still a difficult endeavour. Recognition of true indels is of prime importance in many applications, such as personalized health care, disease genomics and population genetics. Recently, advanced machine learning techniques have been successfully applied to classification problems with large-scale data. In this paper, we present SICaRiO, a gradient boosting classifier for the reliable detection of true indels, trained with the gold-standard dataset from 'Genome in a Bottle' (GIAB) consortium. Our filtering scheme significantly improves the performance of each variant calling pipeline used in GIAB and beyond. SICaRiO uses genomic features that can be computed from publicly available resources, i.e. it does not require sequencing pipeline-specific information (e.g. read depth). This study also sheds lights on prior genomic contexts responsible for the erroneous calling of indels made by sequencing pipelines. We have compared prediction difficulty for three categories of indels over different sequencing pipelines. We have also ranked genomic features according to their predictivity in determining false positives. Md Shariful Islam Bhuyan, Itsik Pe'er, Mohammad Sohel Rahman |
Briefings Bioinform. | 3 |
| 2021 | SSG-LUGIA: Single Sequence based Genome Level Unsupervised Genomic Island Prediction AlgorithmabstractBACKGROUND: Genomic Islands (GIs) are clusters of genes that are mobilized through horizontal gene transfer. GIs play a pivotal role in bacterial evolution as a mechanism of diversification and adaptation to different niches. Therefore, identification and characterization of GIs in bacterial genomes is important for understanding bacterial evolution. However, quantifying GIs is inherently difficult, and the existing methods suffer from low prediction accuracy and precision-recall trade-off. Moreover, several of them are supervised in nature, and thus, their applications to newly sequenced genomes are riddled with their dependency on the functional annotation of existing genomes. RESULTS: We present SSG-LUGIA, a completely automated and unsupervised approach for identifying GIs and horizontally transferred genes. SSG-LUGIA is a novel method based on unsupervised anomaly detection technique, accompanied by further refinement using cues from signal processing literature. SSG-LUGIA leverages the atypical compositional biases of the alien genes to localize GIs in prokaryotic genomes. SSG-LUGIA was assessed on a large benchmark dataset `IslandPick' and on a set of 15 well-studied genomes in the literature and followed by a thorough analysis on the well-understood Salmonella typhi CT18 genome. Furthermore, the efficacy of SSG-LUGIA in identifying horizontally transferred genes was evaluated on two additional bacterial genomes, namely, those of Corynebacterium diphtheria NCTC13129 and Pseudomonas aeruginosa LESB58. SSG-LUGIA was examined on draft genomes and was demonstrated to be efficient as an ensemble method. CONCLUSIONS: Our results indicate that SSG-LUGIA achieved superior performance in comparison to frequently used existing methods. Importantly, it yielded a better trade-off between precision and recall than the existing methods. Its nondependency on the functional annotation of genomes makes it suitable for analyzing newly sequenced, yet uncharacterized genomes. Thus, our study is a significant advance in identification of GIs and horizontally transferred genes. SSG-LUGIA is available as an open source software at https://nibtehaz.github.io/SSG-LUGIA/. Nabil Ibtehaz, Ishtiaque Ahmed, Mohammad Sohel Rahman, Rajeev K. Azad, Md. Shamsuzzoha Bayzid |
Briefings Bioinform. | 4 |
| 2021 | ADACT: a tool for analysing (dis)similarity among nucleotide and protein sequences using minimal and relative absent wordsabstractMOTIVATION: Researchers and practitioners use a number of popular sequence comparison tools that use many alignment-based techniques. Due to high time and space complexity and length-related restrictions, researchers often seek alignment-free tools. Recently, some interesting ideas, namely, Minimal Absent Words (MAW) and Relative Absent Words (RAW), have received much interest among the scientific community as distance measures that can give us alignment-free alternatives. This drives us to structure a framework for analysing biological sequences in an alignment-free manner. RESULTS: In this application note, we present Alignment-free Dissimilarity Analysis & Comparison Tool (ADACT), a simple web-based tool that computes the analogy among sequences using a varied number of indexes through the distance matrix, species relation list and phylogenetic tree. This tool basically combines absent word (MAW or RAW) computation, dissimilarity measures, species relationship and thus brings all required software in one platform for the ease of researchers and practitioners alike in the field of bioinformatics. We have also developed a restful API. AVAILABILITY AND IMPLEMENTATION: ADACT has been hosted at http://research.buet.ac.bd/ADACT/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Mujtahid Akon, Muntashir Akon, Mohimenul Kabir, Mohammad Saifur Rahman 0001, Mohammad Sohel Rahman |
Bioinform. | 5 |
| 2021 | Algorithms and Discrete Mathematics - celebrating the silver jubilee of IITG Guwahati
Gautam K. Das, Subhas C. Nandy, Mohammad Sohel Rahman |
Theor. Comput. Sci. | 3 |
| 2021 | Multidimensional segment trees can do range updates in poly-logarithmic time
Nabil Ibtehaz, Mohammad Kaykobad, Mohammad Sohel Rahman |
Theor. Comput. Sci. | 3 |
| 2021 | A linear time algorithm for the r-gathering problem on the line
Anik Sarker, Wing-Kin Sung, Mohammad Sohel Rahman |
Theor. Comput. Sci. | 3 |
| 2020 | A Multi-objective Metaheuristic Approach for Accurate Species Tree EstimationabstractSpecies tree estimation from multi-locus data is complicated as biological processes can result in different loci having different evolutionary histories. Incomplete lineage sorting (ILS), modeled by the multi-species coalescent (MSC), is considered to be a dominant cause for gene tree incongruence. Various optimization criteria (e.g., quartet score, pseudo-likelihood, etc.) are statistically consistent under the MSC model, meaning that they return the true species tree with high probability given sufficiently large numbers of accurate gene trees. However, the number of genes is limited and estimating highly accurate gene trees is difficult. Therefore, even popular methods, optimizing a particular criterion, may fail to reconstruct highly accurate trees under practical model conditions with limited numbers of genes and in the presence of gene tree estimation error. In this study, we advocate multi-objective optimization to tackle this challenge. Particularly, we present a multi-objective metaheuristic (i.e., SNOGA), a modified version of the popular NSGAII, which combines various optimization criteria to find a suitable search space containing highly accurate species trees. Muhammad Ali Nayeem, Md. Shamsuzzoha Bayzid, Sakshar Chakravarty, Mohammad Saifur Rahman 0001, Mohammad Sohel Rahman |
BIBE | 5 |
| 2020 | Not Low-Resource Anymore: Aligner Ensembling, Batch Filtering, and New Datasets for Bengali-English Machine TranslationabstractTahmid Hasan, Abhik Bhattacharjee, Kazi Samin, Masum Hasan, Madhusudan Basak, M. Sohel Rahman, Rifat Shahriyar. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Tahmid Hasan, Abhik Bhattacharjee, Kazi Samin, Masum Hasan, Madhusudan Basak, Mohammad Sohel Rahman, Rifat Shahriyar |
EMNLP (1) | 6 |
| 2020 | CRISPRpred(SEQ): a sequence-based method for sgRNA on target activity prediction using traditional machine learningabstractBACKGROUND: The latest works on CRISPR genome editing tools mainly employs deep learning techniques. However, deep learning models lack explainability and they are harder to reproduce. We were motivated to build an accurate genome editing tool using sequence-based features and traditional machine learning that can compete with deep learning models. RESULTS: In this paper, we present CRISPRpred(SEQ), a method for sgRNA on-target activity prediction that leverages only traditional machine learning techniques and hand-crafted features extracted from sgRNA sequences. We compare the results of CRISPRpred(SEQ) with that of DeepCRISPR, the current state-of-the-art, which uses a deep learning pipeline. Despite using only traditional machine learning methods, we have been able to beat DeepCRISPR for the three out of four cell lines in the benchmark dataset convincingly (2.174%, 6.905% and 8.119% improvement for the three cell lines). CONCLUSION: CRISPRpred(SEQ) has been able to convincingly beat DeepCRISPR in 3 out of 4 cell lines. We believe that by exploring further, one can design better features only using the sgRNA sequences and can come up with a better method leveraging only traditional machine learning algorithms that can fully beat the deep learning models. Ali Haisam Muhammad Rafid, Md. Toufikuzzaman, Mohammad Saifur Rahman 0001, Mohammad Sohel Rahman |
BMC Bioinform. | 4 |
| 2020 | MultiResUNet : Rethinking the U-Net architecture for multimodal biomedical image segmentation
Nabil Ibtehaz, Mohammad Sohel Rahman |
Neural Networks | 2 |
| 2019 | A 'phylogeny-aware' multi-objective optimization approach for computing MSAabstractMultiple sequence alignment (MSA) is a basic step in many analyses in bioinformatics, including predicting the structure and function of proteins, orthology prediction and estimating phylogenies. The objective of MSA is to infer the homology among the sequences of chosen species. Commonly, the MSAs are inferred by optimizing a single objective function. The alignments estimated under one criterion may be different to the alignments generated by other criteria, inferring discordant homologies and thus leading to different evolutionary histories relating the sequences. In the recent past, researchers have advocated for the multi-objective formulation of MSA, to address this issue, where multiple conflicting objective functions are being optimized simultaneously to generate a set of alignments. However, no theoretical or empirical justification with respect to a real-life application has been shown for a particular multi-objective formulation. In this study, we investigate the impact of multi-objective formulation in the context of phylogenetic tree estimation. In essence, we ask the question whether a phylogeny-aware metric can guide us in choosing appropriate multi-objective formulations. Employing evolutionary optimization, we demonstrate that trees estimated on the alignments generated by multi-objective formulation are substantially better than the trees estimated by the state-of-the-art MSA tools, including PASTA, T-Coffee, MAFFT etc. Muhammad Ali Nayeem, Md. Shamsuzzoha Bayzid, Atif Rahman 0001, Rifat Shahriyar, Mohammad Sohel Rahman |
GECCO | 5 |
| 2019 | Applications of V-Order: Suffix Arrays, the Burrows-Wheeler Transform & the FM-index
Ali Alatabbi, Jacqueline W. Daykin, Neerja Mhaskar, Mohammad Sohel Rahman, William F. Smyth |
WALCOM | 4 |
| 2019 | A Linear Time Algorithm for the r-Gathering Problem on the Line (Extended Abstract)
Anik Sarker, Wing-Kin Sung, Mohammad Sohel Rahman |
WALCOM | 3 |
| 2019 | Antigenic: An improved prediction model of protective antigensabstractAn antigen is a protein capable of triggering an effective immune system response. Protective antigens are the ones that can invoke specific and enhanced adaptive immune response to subsequent exposure to the specific pathogen or related organisms. Such proteins are therefore of immense importance in vaccine preparation and drug design. However, the laboratory experiments to isolate and identify antigens from a microbial pathogen are expensive, time consuming and often unsuccessful. This is why Reverse Vaccinology has become the modern trend of vaccine search, where computational methods are first applied to predict protective antigens or their determinants, known as epitopes. In this paper, we propose a novel, accurate computational model to identify protective antigens efficiently. Our model extracts features directly from the protein sequences, without any dependence on functional domain or structural information. After relevant features are extracted, we have used Random Forest algorithm to rank the features. Then Recursive Feature Elimination (RFE) and minimum redundancy maximum relevance (mRMR) criterion were applied to extract an optimal set of features. The learning model was trained using Random Forest algorithm. Named as Antigenic, our proposed model demonstrates superior performance compared to the state-of-the-art predictors on a benchmark dataset. Antigenic achieves accuracy, sensitivity and specificity values of 78.04%, 78.99% and 77.08% in 10-fold cross-validation testing respectively. In jackknife cross-validation, the corresponding scores are 80.03%, 80.90% and 79.16% respectively. The source code of Antigenic, along with relevant dataset and detailed experimental results, can be found at https://github.com/srautonu/AntigenPredictor. A publicly accessible web interface has also been established at: http://antigenic.research.buet.ac.bd. Mohammad Saifur Rahman 0001, Md. Khaledur Rahman, Sanjay Saha, Mohammad Kaykobad, Mohammad Sohel Rahman |
Artif. Intell. Medicine | 5 |
| 2018 | A Simple, Fast, Filter-Based Algorithm for Circular Sequence Comparison
Md. Aashikur Rahman Azim, Mohimenul Kabir, Mohammad Sohel Rahman |
WALCOM | 3 |
| 2018 | On Multiple Longest Common Subsequence and Common Motifs with Gaps (Extended Abstract)
Suri Dipannita Sayeed, Mohammad Sohel Rahman, Atif Rahman 0001 |
WALCOM | 2 |
| 2018 | isGPT: An optimized model to identify sub-Golgi protein types using SVM and Random Forest based feature selection
Mohammad Saifur Rahman 0001, Md. Khaledur Rahman, Mohammad Kaykobad, Mohammad Sohel Rahman |
Artif. Intell. Medicine | 4 |
| 2018 | Approximation Algorithms for Three Dimensional Protein FoldingabstractPredicting the secondary structure of a protein using a lattice model is one of the most studied computational problems in bioinformatics. Here the secondary structure or three dimensional structure of a protein is predicted from its amino acid sequence. The secondary structure refers to the local sub-structures of a protein. Simplified energy models have been proposed in the literature on the basis of interaction of amino acid residues in proteins. We focus on a well researched model known as the Hydrophobic-Polar (HP) energy model. In this paper, we propose the hexagonal prism lattice with diagonals that can overcome the problems of other lattice structures, e.g., parity problem. We give two approximation algorithms for protein folding on this lattice using HP model. Our first algorithm leads us to a similar structure of helix structure that is commonly found in a protein structure. This motivates us to propose the next algorithm with a better approximation ratio. Finally, we analyze the algorithms on the basis of intensity of the chemical forces along the different types of edges of hexagonal prism lattice with diagonals. Dipan Lal Shaw, A. S. M. Shohidull Islam, Shuvasish Karmaker, Mohammad Sohel Rahman |
Fundam. Informaticae | 4 |
| 2018 | An ensemble learning based approach for impression fraud detection in mobile advertising
Ch. Md. Rakin Haider, Anindya Iqbal, Atif Rahman 0001, Mohammad Sohel Rahman |
J. Netw. Comput. Appl. | 4 |
| 2018 | Using Adaptive Heartbeat Rate on Long-Lived TCP ConnectionsabstractIn this paper, we propose techniques for dynamically adjusting heartbeat or keep-alive interval of long-lived TCP connections, particularly the ones that are used in push notification service in mobile platforms. When a device connects to a server using TCP, often times the connection is established through some sort of middle-box, such as NAT, proxy, firewall, and so on. When such a connection is idle for a long time, it may get torn down due to binding timeout of the middle-box. To keep the connection alive, the client device needs to send keep-alive packets through the connection when it is otherwise idle. To reduce resource consumption, the keep-alive packet should preferably be sent at the farthest possible time within the binding timeout. Due to varied settings of different network equipments, the binding timeout will not be identical in different networks. Hence, the heartbeat rate used in different networks should be changed dynamically. We propose a set of iterative probing techniques, namely binary, exponential, and composite search, that detect the middle-box binding timeout with varying degree of accuracy; and in the process, keeps improving the keep-alive interval used by the client device. We also analytically derive performance bounds of these techniques. To the best of our knowledge, ours is the first work that systematically studies several techniques to dynamically improve keep-alive interval. To this end, we run experiments in simulation as well as make a real implementation on android to demonstrate the proof-of-concept of the proposed schemes. Mohammad Saifur Rahman 0001, Md. Yusuf Sarwar Uddin, Tahmid Hasan, Mohammad Sohel Rahman, Mohammad Kaykobad |
IEEE/ACM Trans. Netw. | 4 |
| 2016 | Mapping stream programs onto multicore platforms by local search and genetic algorithm
S. M. Farhad, Muhammad Ali Nayeem, Md. Khaledur Rahman, Mohammad Sohel Rahman |
Comput. Lang. Syst. Struct. | 4 |
| 2016 | V-Order: New combinatorial properties & a simple comparison algorithm
Ali Alatabbi, Jacqueline W. Daykin, Juha Kärkkäinen, Mohammad Sohel Rahman, William F. Smyth |
Discret. Appl. Math. | 4 |
| 2016 | Computing covers using prefix tables
Ali Alatabbi, Mohammad Sohel Rahman, William F. Smyth |
Discret. Appl. Math. | 2 |
| 2016 | An efficient algorithm to detect common ancestor genes for non-overlapping inversion and applications
Fatema Tuz Zohora, Mohammad Sohel Rahman |
Theor. Comput. Sci. | 2 |
| 2015 | A Filter-Based Approach for Approximate Circular Pattern Matching
Md. Aashikur Rahman Azim, Costas S. Iliopoulos, Mohammad Sohel Rahman, M. Samiruzzaman |
ISBRA | 3 |
| 2015 | New Heuristics for Clustering Large Biological Networks
Md. Kishwar Shafin, Kazi Lutful Kabir, Iffatur Ridwan, Tasmiah Tamzid Anannya, Rashid Saadman Karim, Mohammad Mozammel Hoque, Mohammad Sohel Rahman |
ISBRA | 7 |
| 2015 | Simple Linear Comparison of Strings in V-orderabstractIn this paper we focus on a total (but non-lexicographic) ordering of strings called V-order. We devise a new linear-time algorithm for computing the V-comparison of two finite strings. In comparison with the previous algorithm in the literature, our algorithm is both conceptually simpler, based on recording letter positions in increasing order, and more straightforward to implement, requiring only linked lists. Ali Alatabbi, Jacqueline W. Daykin, Mohammad Sohel Rahman, William F. Smyth |
Fundam. Informaticae | 3 |
| 2015 | Constrained sequence analysis algorithms in computational biology
Effat Farhana, Mohammad Sohel Rahman |
Inf. Sci. | 2 |
| 2015 | Order preserving pattern matching revisited
Md. Mahbubul Hasan, A. S. M. Shohidull Islam, Mohammad Saifur Rahman 0001, Mohammad Sohel Rahman |
Pattern Recognit. Lett. | 4 |
| 2014 | GreMuTRRR: A novel genetic algorithm to solve distance geometry problem for protein structuresabstractNuclear Magnetic Resonance (NMR) Spectroscopy is a widely used technique to predict the native structure of proteins. However, NMR machines are only able to report approximate and partial distances between pair of atoms. To build the protein structure one has to solve the Euclidean distance geometry problem given the incomplete interval distance data produced by NMR machines. In this paper, we propose a new genetic algorithm for solving the Euclidean distance geometry problem for protein structure prediction given sparse NMR data. Our genetic algorithm uses a greedy mutation operator to intensify the search, a twin removal technique for diversification in the population and a random restart method to recover stagnation. On a standard set of benchmark dataset, our algorithm significantly outperforms standard genetic algorithms. Md. Lisul Islam, Swakkhar Shatabda, Mohammad Sohel Rahman |
BIBM | 3 |
| 2014 | Application of Consensus String Matching in the Diagnosis of Allelic Heterogeneity - (Extended Abstract)
Fatema Tuz Zohora, Mohammad Sohel Rahman |
ISBRA | 2 |
| 2014 | Order Preserving Prefix Tables
Md. Mahbubul Hasan, A. S. M. Shohidull Islam, Mohammad Saifur Rahman 0001, Mohammad Sohel Rahman |
SPIRE | 4 |
| 2014 | Protein folding in HP model on hexagonal lattices with diagonalsabstractThree dimensional structure prediction of a protein from its amino acid sequence, known as protein folding, is one of the most studied computational problem in bioinformatics and computational biology. Since, this is a hard problem, a number of simplified models have been proposed in literature to capture the essential properties of this problem. In this paper we introduce the hexagonal lattices with diagonals to handle the protein folding problem considering the well researched HP model. We give two approximation algorithms for protein folding on this lattice. Our first algorithm is a 5/3-approximation algorithm, which is based on the strategy of partitioning the entire protein sequence into two pieces. Our next algorithm is also based on partitioning approaches and improves upon the first algorithm. Dipan Lal Shaw, A. S. M. Shohidull Islam, Mohammad Sohel Rahman, Masud Hasan |
BMC Bioinform. | 3 |
| 2014 | Computing a Longest Common Palindromic SubsequenceabstractThe longest common subsequence (LCS) problem is a classic and well-studied problem in computer science. Palindrome is a word which reads the same forward as it does backward. The longest common palindromic subsequence (LCPS) problem is a variant of the classic LCS problem which finds a longest common subsequence between two given strings such that the computed subsequence is also a palindrome. In this paper, we study the LCPS problem and give two novel algorithms to solve it. To the best of our knowledge, this is the first attempt to study and solve this problem. Shihabur Rahman Chowdhury, Md. Mahbubul Hasan, Sumaiya Iqbal, Mohammad Sohel Rahman |
Fundam. Informaticae | 4 |
| 2014 | The swap matching problem revisited
Pritom Ahmed, Costas S. Iliopoulos, A. S. M. Shohidull Islam, Mohammad Sohel Rahman |
Theor. Comput. Sci. | 4 |
| 2013 | Protein Folding in 2D-Triangular Lattice Revisited - (Extended Abstract)
A. S. M. Shohidull Islam, Mohammad Sohel Rahman |
IWOCA | 2 |
| 2013 | On Palindromic Sequence Automata and Applications
Md. Mahbubul Hasan, A. S. M. Shohidull Islam, Mohammad Sohel Rahman, Ayon Sen |
CIAA | 3 |
| 2013 | A divide and conquer approach and a work-optimal parallel algorithm for the LIS problem
Muhammad Rashed Alam, Mohammad Sohel Rahman |
Inf. Process. Lett. | 2 |
| 2013 | Prefix transpositions on binary and ternary strings
Masud Hasan, Mohammad Sohel Rahman |
Inf. Process. Lett. | 3 |
| 2012 | Placement of unique restriction sites in synthetic genomes using multi-objective optimizationabstractPlacing unique restriction sites in synthetic genomes is computationally hard problem. This paper targets to apply variants of genetic algorithm to modify virus length genomes to introduce a large number of evenly spaced unique restriction sites while preserving their amino-acid sequence. From our experimental results we show that genetic algorithms, considering multiple objectives non-linearly, give better results than traditional heuristic methods and local search methods. Mahfuza Sharmin, Mohammad Sohel Rahman |
BIBM | 2 |
| 2012 | Bee algorithms for solving DNA fragment assembly problem with noisy and noiseless dataabstractDNA fragment assembly problem is one of the crucial challenges faced by computational biologists where, given a set of DNA fragments, we have to construct a complete DNA sequence from them. As it is an NP-hard problem, accurate DNA sequence is hard to find. Moreover, due to experimental limitations, the fragments considered for assembly are exposed to additional errors while reading the fragments. In such scenarios, meta-heuristic based algorithms can come in handy. We analyze the performance of two swarm intelligence based algorithms namely Artificial Bee Colony (ABC) algorithm and Queen Bee Evolution Based on Genetic Algorithm (QEGA) to solve the fragment assembly problem and report quite promising results. Our main focus is to design meta-heuristic based techniques to efficiently handle DNA fragment assembly problem for noisy and noiseless data. Jesun Sahariar Firoz, Mohammad Sohel Rahman, Tanay Kumar Saha |
GECCO | 2 |
| 2012 | A Graph Theoretic Model to Solve the Approximate String Matching Problem Allowing for Translocations
Pritom Ahmed, A. S. M. Shohidull Islam, Mohammad Sohel Rahman |
IWOCA | 3 |
| 2012 | Computing a Longest Common Palindromic Subsequence
Shihabur Rahman Chowdhury, Md. Mahbubul Hasan, Sumaiya Iqbal, Mohammad Sohel Rahman |
IWOCA | 4 |
| 2012 | Doubly-Constrained LCS and Hybrid-Constrained LCS problems revisited
Effat Farhana, Mohammad Sohel Rahman |
Inf. Process. Lett. | 2 |
| 2012 | Improved algorithms for the range next value problem and applications
Maxime Crochemore, Costas S. Iliopoulos, Marcin Kubica 0001, Mohammad Sohel Rahman, German Tischler, Tomasz Walen |
Theor. Comput. Sci. | 4 |
| 2011 | Improved Algorithms for the Point-Set Embeddability Problem for Plane 3-Trees
Tanaeem M. Moosa, Mohammad Sohel Rahman |
COCOON | 2 |
| 2011 | Solving the generalized Subset Sum problem with a light based device
Masud Hasan, S. M. Shabab Hossain, Mohammad Sohel Rahman |
Nat. Comput. | 4 |
| 2010 | Finite Automata Based Algorithms for the Generalized Constrained Longest Common Subsequence Problems
Effat Farhana, Tanaeem M. Moosa, Mohammad Sohel Rahman |
SPIRE | 4 |
| 2010 | Finding Patterns In Given IntervalsabstractIn this paper, we study the pattern matching problem in given intervals. Depending on whether the intervals are given a priori for pre-processing, or during the query along with the pattern or, even in both the cases, we develop efficient solutions for different variants of this problem. In particular, we present efficient indexing schemes for each of the above variants of the problem. Maxime Crochemore, Marcin Kubica 0001, Tomasz Walen, Costas S. Iliopoulos, Mohammad Sohel Rahman |
Fundam. Informaticae | 5 |
| 2010 | Indexing permutations for binary strings
Tanaeem M. Moosa, Mohammad Sohel Rahman |
Inf. Process. Lett. | 2 |
| 2009 | Indexing Factors with Gaps
Costas S. Iliopoulos, Mohammad Sohel Rahman |
Algorithmica | 2 |
| 2009 | A New Efficient Algorithm for Computing the Longest Common Subsequence
Costas S. Iliopoulos, Mohammad Sohel Rahman |
Theory Comput. Syst. | 2 |
| 2008 | A New Model to Solve the Swap Matching Problem and Efficient Algorithms for Short Patterns
Costas S. Iliopoulos, Mohammad Sohel Rahman |
SOFSEM | 2 |
| 2008 | Improved Algorithms for the Range Next Value Problem and ApplicationsabstractThe Range Next Value problem (Problem RNV) is a recent interesting variant of the range search problems, where the query is for the immediate next (or equal) value of a given number within a given interval of an array. Problem RNV was introduced and studied very recently by Crochemore et. al [Finding Patterns In Given Intervals, MFCS 2007]. In this paper, we present improved algorithms for Problem RNV. We also show how this problem can be used to achieve optimal query time for a number of interesting variants of the classic pattern matching problems. Maxime Crochemore, Costas S. Iliopoulos, Marcin Kubica 0001, Mohammad Sohel Rahman, Tomasz Walen |
STACS | 4 |
| 2008 | Optimal prefix and suffix queries on texts
Maxime Crochemore, Costas S. Iliopoulos, Mohammad Sohel Rahman |
Inf. Process. Lett. | 3 |
| 2008 | Faster index for property matching
Costas S. Iliopoulos, Mohammad Sohel Rahman |
Inf. Process. Lett. | 2 |
| 2008 | New efficient algorithms for the LCS and constrained LCS problems
Costas S. Iliopoulos, Mohammad Sohel Rahman |
Inf. Process. Lett. | 2 |
| 2008 | Algorithms for computing variants of the longest common subsequence problem
Costas S. Iliopoulos, Mohammad Sohel Rahman |
Theor. Comput. Sci. | 2 |
| 2008 | A multiprocessor based heuristic for multi-dimensional multiple-choice knapsack problem
Abu Zafar M. Shahriar, Md. Mostofa Akbar, Mohammad Sohel Rahman, M. A. Hakim Newton |
J. Supercomput. | 3 |
| 2007 | A New Efficient Algorithm for Computing the Longest Common Subsequence
Mohammad Sohel Rahman, Costas S. Iliopoulos |
AAIM | 1 |
| 2007 | Algorithms for Computing the Longest Parameterized Common Subsequence
Costas S. Iliopoulos, Marcin Kubica 0001, Mohammad Sohel Rahman, Tomasz Walen |
CPM | 3 |
| 2007 | Finding Patterns in Given Intervals
Maxime Crochemore, Costas S. Iliopoulos, Mohammad Sohel Rahman |
MFCS | 3 |
| 2007 | Pattern Matching Algorithms with Don't Cares
Mohammad Sohel Rahman, Costas S. Iliopoulos |
SOFSEM (2) | 1 |
| 2007 | Indexing Factors with Gaps
Mohammad Sohel Rahman, Costas S. Iliopoulos |
SOFSEM (1) | 1 |
| 2007 | The Constrained Longest Common Subsequence Problem for Degenerate Strings
Costas S. Iliopoulos, Mohammad Sohel Rahman, Michal Vorácek, Ladislav Vagner |
CIAA | 2 |
| 2006 | Finding Patterns with Variable Length Gaps or Don't Cares
Mohammad Sohel Rahman, Costas S. Iliopoulos, Manal Mohamed 0001, William F. Smyth |
COCOON | 1 |
| 2006 | Algorithms for Computing Variants of the Longest Common Subsequence Problem
Mohammad Sohel Rahman, Costas S. Iliopoulos |
ISAAC | 1 |
| 2005 | On Hamiltonian cycles and Hamiltonian paths
Mohammad Sohel Rahman, Mohammad Kaykobad |
Inf. Process. Lett. | 1 |
| 2005 | Complexities of some interesting problems on spanning trees
Mohammad Sohel Rahman, Mohammad Kaykobad |
Inf. Process. Lett. | 1 |