EDBT 2026 Demo / reviewers in the wild / expert
Tanvir Alam
dblp:197/4038
· DBLP profile ↗
12ranked-venue papers
3as first author
9since 2021 · last 2025
0000-0001-7033-3693ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 10 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DR-VQA: Large Language Model Based Vision Question Answering System on Diabetic RetinopathyabstractMedical Visual Question Answering (VQA) presents challenges due to the complexity of imaging data and the need for precise, context-aware responses. Traditional VQA models often struggle in clinical settings, limiting their utility in decision-making. This study proposes DR-VQA, a fine-tuned BLIP (Bootstrapped Language-Image Pretraining) model for medical VQA specifically designed for diabetic retinopathy (DR) based on fundus imaging, leveraging both visual and textual data to generate accurate diagnostic answers. To enhance semantic relevance, BERT-based similarity evaluation is integrated. Using a diabetic retinopathy dataset, the model achieves a validation BERT similarity (BERTsim) score of 0.94 and a test score of 0.95 on 450 samples, demonstrating strong alignment with expert annotations. These results highlight the model's potential to assist clinicians by improving diagnostic accuracy and efficiency. The proposed approach can streamline medical workflows, reduce clinician workload, and enhance patient outcomes. Future work will focus on expanding datasets and refining the model for broader medical applications. We believe our approach will support to enhance the patient care as well democratization on AI technology for community. Saleh Musleh, Hamada R. H. Al-Absi, Md. Rizwan Parvez, Anant Pai, Ghassan Ahmed Mubasher Mohamedsalih, Tanvir Alam |
TENCON | 6 |
| 2025 | PCBert-Kla: an efficient prediction method for lysine lactylation sites based on ProtBert and fusion of physicochemical featuresabstractProtein post-translational modifications (PTMs) play a critical role in regulating protein functionality and structural diversity. Among them, lysine lactylation (Kla), a newly identified PTM, is involved in energy metabolism, cellular reprogramming, and the progression of various diseases. In this study, we propose PCBert-Kla, a feature-fusion deep learning model based on ProtBert. This model leverages ProtBert to extract deep features from protein sequences, effectively capturing global and local contextual information. It integrated various physicochemical properties, including molecular weight, isoelectric point, amino acid composition, secondary structure content, hydrophobicity, and net charge. An attention mechanism in the fully connected layers enabled the model to select features automatically. PCBert-Kla exhibited exceptional accuracy and reliability in Kla site identification and demonstrated excellent generalization capability to outperform the existing models. In addition, we further enhanced the interpretability of the PCBert-Kla model by incorporating average attention maps. This model provided powerful tools for studying the functions of Kla and elucidating the mechanisms of related diseases, which can advance biomedical research and drug development. We also developed a free web service, available at http://pcbert-kla.lin-group.cn/, to provide users with easy access and usage. Hong-Qi Zhang, Yi-Xuan Qi, Huma Fida, Hao-Jiang Zhang, Muhammad Arif 0012, Pei-Yu Zhao, Tanvir Alam, Ye-Chen Qi, Xiao-Long Yu, Ke-Jun Deng |
Briefings Bioinform. | 7 |
| 2025 | Cardiometabolic biomarker prediction based on retinal fundus imageabstractDiagnosing common noncommunicable diseases, such as cardiovascular disease and diabetes, typically relies on blood sample analysis for biomarker measurement. This process is invasive, time-consuming, and relatively expensive. To address these limitations, deep learning methods were leveraged to estimate common cardiometabolic biomarkers using retinal fundus (RF) images. The study utilized 15,802 RF images from 5,653 participants in the Qatar Biobank (QBB), leading to the development of 19 deep-learning models to estimate biomarkers across seven categories: demographics and body composition, blood pressure, lipid profile, blood profile, hormones, kidney function, and metabolites. The proposed model outperformed existing models for the QBB-specific cohort across all biomarkers, achieving higher R-squared ( R 2 ) values, lower mean absolute error (MAE), and higher area under the curve (AUC). The proposed model achieved excellent performance in demographic predictions with age (MAE: 2.56, R 2 : 0.93) and gender (Accuracy: 96%, AUC: 0.94). For cardiovascular markers, it showed moderate predictability with systolic blood pressure (MAE: 8.02, R 2 : 0.49) and diastolic blood pressure (MAE: 6.06, R 2 : 0.45). For metabolic markers, the model demonstrated varying performance, with hemoglobin showing strong prediction (MAE: 0.79, R 2 : 0.60) while lipid markers showed moderate performance (total cholesterol MAE: 0.63, R 2 : 0.29). For creatinine, a kidney function marker, we achieved the best results with MAE: 9.00, R 2 : 0.33. Stratified analyses revealed systematic performance variations across gender, age, and disease-specific subgroups, with better predictions in males, young-agers, and non-diabetic participants. External validation of the CAD group confirms the effect of age, gender, and disease on prediction results, suggesting the need for personalized background in consideration for developing AI models. This study presents a promising approach for non-invasive biomarker estimation using retinal images, potentially revolutionizing early intervention and treatment planning in healthcare. Syed Abdullah Basit, Hamada R. H. Al-Absi, Saleh Musleh, Tanvir Alam |
Eng. Appl. Artif. Intell. | 4 |
| 2024 | DPI_CDF: druggable protein identifier using cascade deep forestabstractBACKGROUND: Drug targets in living beings perform pivotal roles in the discovery of potential drugs. Conventional wet-lab characterization of drug targets is although accurate but generally expensive, slow, and resource intensive. Therefore, computational methods are highly desirable as an alternative to expedite the large-scale identification of druggable proteins (DPs); however, the existing in silico predictor's performance is still not satisfactory. METHODS: In this study, we developed a novel deep learning-based model DPI_CDF for predicting DPs based on protein sequence only. DPI_CDF utilizes evolutionary-based (i.e., histograms of oriented gradients for position-specific scoring matrix), physiochemical-based (i.e., component protein sequence representation), and compositional-based (i.e., normalized qualitative characteristic) properties of protein sequence to generate features. Then a hierarchical deep forest model fuses these three encoding schemes to build the proposed model DPI_CDF. RESULTS: The empirical outcomes on 10-fold cross-validation demonstrate that the proposed model achieved 99.13 % accuracy and 0.982 of Matthew's-correlation-coefficient (MCC) on the training dataset. The generalization power of the trained model is further examined on an independent dataset and achieved 95.01% of maximum accuracy and 0.900 MCC. When compared to current state-of-the-art methods, DPI_CDF improves in terms of accuracy by 4.27% and 4.31% on training and testing datasets, respectively. We believe, DPI_CDF will support the research community to identify druggable proteins and escalate the drug discovery process. AVAILABILITY: The benchmark datasets and source codes are available in GitHub: http://github.com/Muhammad-Arif-NUST/DPI_CDF . Muhammad Arif 0012, Ge Fang, Ali Ghulam, Saleh Musleh, Tanvir Alam |
BMC Bioinform. | 5 |
| 2023 | MSLP: mRNA subcellular localization predictor based on machine learning techniquesabstractBACKGROUND: Subcellular localization of messenger RNA (mRNAs) plays a pivotal role in the regulation of gene expression, cell migration as well as in cellular adaptation. Experiment techniques for pinpointing the subcellular localization of mRNAs are laborious, time-consuming and expensive. Therefore, in silico approaches for this purpose are attaining great attention in the RNA community. METHODS: In this article, we propose MSLP, a machine learning-based method to predict the subcellular localization of mRNA. We propose a novel combination of four types of features representing k-mer, pseudo k-tuple nucleotide composition (PseKNC), physicochemical properties of nucleotides, and 3D representation of sequences based on Z-curve transformation to feed into machine learning algorithm to predict the subcellular localization of mRNAs. RESULTS: Considering the combination of the above-mentioned features, ennsemble-based models achieved state-of-the-art results in mRNA subcellular localization prediction tasks for multiple benchmark datasets. We evaluated the performance of our method in ten subcellular locations, covering cytoplasm, nucleus, endoplasmic reticulum (ER), extracellular region (ExR), mitochondria, cytosol, pseudopodium, posterior, exosome, and the ribosome. Ablation study highlighted k-mer and PseKNC to be more dominant than other features for predicting cytoplasm, nucleus, and ER localizations. On the other hand, physicochemical properties and Z-curve based features contributed the most to ExR and mitochondria detection. SHAP-based analysis revealed the relative importance of features to provide better insights into the proposed approach. AVAILABILITY: We have implemented a Docker container and API for end users to run their sequences on our model. Datasets, the code of API and the Docker are shared for the community in GitHub at: https://github.com/smusleh/MSLP . Saleh Musleh, Mohammad Tariqul Islam 0002, Rizwan Qureshi, Nihad Alajez, Tanvir Alam |
BMC Bioinform. | 5 |
| 2023 | Correction: MSLP: mRNA subcellular localization predictor based on machine learning techniques
Saleh Musleh, Mohammad Tariqul Islam 0002, Rizwan Qureshi, Nihad Alajez, Tanvir Alam |
BMC Bioinform. | 5 |
| 2023 | Lung-EffNet: Lung cancer classification using EfficientNet from CT-scan imagesabstractLung cancer (LC) remains a leading cause of death worldwide. Early diagnosis is critical to protect innocent human lives. Computed tomography (CT) scans are one of the primary imaging modalities for lung cancer diagnosis. However, manual CT scan analysis is time-consuming and prone to errors/not accurate. Considering these shortcomings, computational methods especially machine learning and deep learning algorithms are leveraged as an alternative to accelerate the accurate detection of CT scans as cancerous, and non-cancerous. In the present article, we proposed a novel transfer learning-based predictor called, Lung-EffNet for lung cancer classification. Lung-EffNet is built based on the architecture of EfficientNet and further modified by adding top layers in the classification head of the model. Lung-EffNet is evaluated by utilizing five variants of EfficientNet i.e., B0–B4. The experiments are conducted on the benchmark dataset “IQ-OTH/NCCD” for lung cancer patients grouped as benign, malignant, or normal based on the presence or absence of lung cancer. The class imbalance issue was handled through multiple data augmentation methods to overcome the biases. The developed model Lung-EffNet attained 99.10% of accuracy and a score of 0.97 to 0.99 of ROC on the test set. We compared the efficacy of the proposed fine-tuned pre-trained EfficientNet with other pre-trained CNN architectures. The predicted outcomes demonstrate that EfficientNetB1 based Lung-EffNet outperforms other CNNs in terms of both accuracy and efficiency. Moreover, it is faster and requires fewer parameters to train than other CNN based models, making it a good choice for large-scale deployment in clinical settings and a promising tool for automated lung cancer diagnosis from CT scan images. Rehan Raza, Fatima Zulfiqar, Muhammad Owais Khan, Muhammad Arif 0012, Atif Alvi, Muhammad Aksam Iftikhar, Tanvir Alam |
Eng. Appl. Artif. Intell. | 7 |
| 2023 | Computational Methods for the Analysis and Prediction of EGFR-Mutated Lung Cancer Drug Resistance: Recent Advances in Drug Design, Challenges and Future ProspectsabstractLung cancer is a major cause of cancer deaths worldwide, and has a very low survival rate. Non-small cell lung cancer (NSCLC) is the largest subset of lung cancers, which accounts for about 85% of all cases. It has been well established that a mutation in the epidermal growth factor receptor (EGFR) can lead to lung cancer. EGFR Tyrosine Kinase Inhibitors (TKIs) are developed to target the kinase domain of EGFR. These TKIs produce promising results at the initial stage of therapy, but the efficacy becomes limited due to the development of drug resistance. In this paper, we provide a comprehensive overview of computational methods, for understanding drug resistance mechanisms. The important EGFR mutants and the different generations of EGFR-TKIs, with the survival and response rates are discussed. Next, we evaluate the role of important EGFR parameters in drug resistance mechanism, including structural dynamics, hydrogen bonds, stability, dimerization, binding free energies, and signaling pathways. Personalized drug resistance prediction models, drug response curve, drug synergy, and other data-driven methods are also discussed. Recent advancements in deep learning; such as AlphaFold2, deep generative models, big data analytics, and the applications of statistics and permutation are also highlighted. We explore limitations in the current methodologies, and discuss strategies to overcome them. We believe this review will serve as a reference for researchers; to apply computational techniques for precision medicine, analyzing structures of protein-drug complexes, drug discovery, and understanding the drug response and resistance mechanisms in lung cancer patients. Rizwan Qureshi, Bin Zou 0004, Tanvir Alam, Jia Wu 0009, Victor H. F. Lee, Hong Yan 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2022 | Drug response prediction for lung cancer patients using biophysical simulation and machine learningabstractLung cancer is one of the most prevalent contributors to cancer deaths worldwide. The over-expression of Epidermal growth factor receptor (EGFR) is found in about 60% of non-small cell lung cancer (NSCLC) patients. Food and Drug Administration (FDA) has approved small molecule inhibitors, targeting the kinase domain of EGFR and to stop the abnormal growth of the cancer cells. These inhibitors produce encouraging results, but the long term efficacy remains limited due to secondary point mutations. In this work, we have developed a framework, using molecular dynamics (MD) simulation and machine learning to predict the drug response in lung cancer patients and to understand the mechanism of drug resistance. The experiments on an independent cohort of 61 patients shows the effectiveness of the proposed approach. Rizwan Qureshi, Tanvir Alam, Jia Wu 0009 |
BIBM | 2 |
| 2020 | Proteome-level assessment of origin, prevalence and function of leucine-aspartic acid (LD) motifsabstractMOTIVATION: Leucine-aspartic acid (LD) motifs are short linear interaction motifs (SLiMs) that link paxillin family proteins to factors controlling cell adhesion, motility and survival. The existence and importance of LD motifs beyond the paxillin family is poorly understood. RESULTS: To enable a proteome-wide assessment of LD motifs, we developed an active learning based framework (LD motif finder; LDMF) that iteratively integrates computational predictions with experimental validation. Our analysis of the human proteome revealed a dozen new proteins containing LD motifs. We found that LD motif signalling evolved in unicellular eukaryotes more than 800 Myr ago, with paxillin and vinculin as core constituents, and nuclear export signal as a likely source of de novo LD motifs. We show that LD motif proteins form a functionally homogenous group, all being involved in cell morphogenesis and adhesion. This functional focus is recapitulated in cells by GFP-fused LD motifs, suggesting that it is intrinsic to the LD motif sequence, possibly through their effect on binding partners. Our approach elucidated the origin and dynamic adaptations of an ancestral SLiM, and can serve as a guide for the identification of other SLiMs for which only few representatives are known. AVAILABILITY AND IMPLEMENTATION: LDMF is freely available online at www.cbrc.kaust.edu.sa/ldmf; Source code is available at https://github.com/tanviralambd/LD/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Tanvir Alam, Meshari Alazmi, Rayan Naser, Franceline Huser, Afaque A. Momin, Veronica Astro, SeungBeom Hong, Katarzyna W. Walkiewicz, Christian G. Canlas, Raphaël Huser, Amal J. Ali, Jasmeen Merzaban, Antonio Adamo, Mariusz Jaremko, Lukasz Jaremko, Vladimir B. Bajic, Xin Gao 0001, Stefan T. Arold |
Bioinform. | 1 |
| 2019 | DeePEL: Deep learning architecture to recognize p-lncRNA and e-lncRNA promotersabstractPromoter regions of long non-coding RNA (lncRNA) genes are crucial to understand their transcriptional regulatory pattern. LncRNA genes, being more cryptic than protein-coding genes in terms of their functionality and biogenesis divergence, are lacking in number of existing studies to elucidate the roles of their promoters compared to their counterparts. Based on the overlap between epigenetic marks and transcription start sites, human lncRNAs were categorized into two broad categories: enhancer-originated lncRNAs (e-lncRNAs) and promoter-originated lncRNAs (p-lncRNAs) and hence these two groups are subject to distinct transcriptional regulatory programs. To understand the difference in the transcriptional regulatory mechanisms that governs p- and e-lncRNAs, we studied the promoter sequences of these two groups of lncRNAs including distinct transcription factor (TF) proteins that favor p-over e-lncRNA (and vice versa). In addition, we developed a convolution neural network (CNN) based deep learning (DL) framework DeePEL (deep p-, e-lncRNA promoter recognizer), to classify the promoter of p- and e-lncRNAs. To the best of our knowledge, this is the first attempt to classify these two groups of lncRNA promoters, using sequence and TF information, based on DL framework. We report several sequence specific signatures in the promoter regions as well as several distinct TFs specific to groups of lncRNAs that will help in understanding the promoter-proximal transcriptional regulation of p-lncRNAs and e-lncRNAs. Tanvir Alam, Mohammad Tariqul Islam 0002, Sebastian Schmeier, Mowafa Househ, Dena Al-Thani |
BIBM | 1 |
| 2019 | An Approach to Design and Develop UX/UI for Smartphone Applications of Minority Ethnic GroupabstractIn a developing country like Bangladesh, minority ethnic groups or tribal people are less privileged to use the benefit of ICT interventions. The smartphone applications hardly persuade the tribal communities. Even though smart cell phones are cheaply available to these people, there is hardly any significant impact of mobile applications. We have studied the usage of smartphone and its applications by the people of various tribal communities. Especially, the UX/UI of the two popular smartphone applications (bKash and Bikroy.com) has been evaluated according to the usage of tribal people. Interestingly, we have found that the culture, language, and customs of the ethnic minority groups make issues to have a successful interaction with those two applications. The local software industries never consider them as a stakeholder of the generic software during the developments. In this paper, we not only have addressed the challenges and issues regarding UX/UI of smart applications for tribal people but also a suitable solution has been recommended in terms of design and user experience of tribal people. The findings of this paper can help in developing mobile applications and services which will be beneficent to the tribal and ethnic people. Tanvir Alam, Md Montaser Hamid, Md Forhad Rabbi |
TENCON | 1 |