Md Faisal Kabir 0001

dblp:234/1632 · DBLP profile ↗
← Back
6ranked-venue papers in the field
2as first author
4since 2021 · last 2025
0000-0001-6088-9487ORCID · verified

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 6 (2 first)
YearPublicationVenuePosition
2025 Extracting Latent Insights and Tagging Fall Injuries from Clinical Narratives Using Unsupervised Learning
Naga V. S. M. Rayabharam, Mohammad Masudur Rahman 0004, Rahmanul Hoque, Sayed Mohsin Reza, Md Faisal Kabir 0001
IEEE Big Data5
2025 StagRAG: Explaining Where RAG Systems Go Wrong in Medical Reasoning
Md AlMahfuz Chowdhury, Mahmuda Akter Sristy, Md Faisal Kabir 0001
IEEE Big Data4
2024 Enhancing Early Diagnosis of Autism Spectrum Disorder Using Multimodal Data and Explainable AI Models
abstract
Autism Spectrum Disorder (ASD) presents profound challenges in early diagnosis due to its inherent complexity and variability. This research leverages a multimodal framework that integrates phenotypic data and neuroimaging quality metrics to establish a comprehensive machine learning pipeline for ASD prediction. Three machine learning models namely Gradient Boosting Machine (GBM), XGBoost, and Support Vector Machine (SVM) were trained and evaluated. Experimental results showed that GBM is best suited compared to other techniques for this case. To ensure clinical applicability, Shapley Additive Explanations (SHAP) were employed to elucidate feature contributions, fostering transparency and trust in the predictive process. This study highlights the potential of integrating machine learning models with interpretable frameworks to improve ASD diagnostics and support evidence-based clinical decision-making.
Yeganeh Abdollahinejad, Md Faisal Kabir 0001
IEEE Big Data2
2024 Identifying Probable Neurological Disorders with Explainable Machine Learning Techniques
abstract
As the global burden of neurological disorders increases, the need for new tools to diagnose brain disease has become increasingly critical. Several machine learning models have been developed to classify various brain disorders; however, ensuring that these models can provide clinically useful information is as important as good prediction performance. In this study, we evaluated the effectiveness of various machine learning models to accurately identify individuals with potential cerebral pathology using neuropsychological data collected during routine clinical care. Three machine learning models-Random Forest, XGBoost, and Graph Neural Network-were trained and evaluated. The models’ performances were compared to identify the most effective approach. The model with the best performance was then used to generate explainability plots, offering insights into the key features that contribute to predictions. Our work shows that Graph Neural Network best suited for routine clinical data, where missing and imbalanced data is commonplace due to the prioritization of patient needs over data completeness.
Neha Khatriya, Tianjie Chen, Md Faisal Kabir 0001, Timothy W. Brearly
IEEE Big Data3
2019 Classification Models and Survival Analysis for Prostate Cancer Using RNA Sequencing and Clinical Data
abstract
Early detection of cancer can significantly increase the chance of successful treatment. This research performs a study on early cancer detection for prostate cancer patients from whom cancer tissue was analyzed with Illumina Hi-Seq ribonucleic acid (RNA) Sequencing (RNA-Seq). Cancer relevant genes with the most significant correlations with the clinical outcome of the sample type (cancer /non-cancer) and the overall survival (OS) were assessed. Traditional cancer diagnosis primarily depends on physicians' experience to identify morphological abnormalities. Gene expression level data can assist physicians in detecting cancer cases at a much earlier stage and thus can significantly improve the potential of patient treatment. In this research, for the classification task, we applied machine learning and data mining approaches to detect cancer versus non-cancer based on gene expression data. Our goal was to detect cancer at the earliest stage. Besides, for the regression task, survival outcomes in prostate cancer patients were performed. Regression trees were built using cancer-sensitive genes along with clinical attribute `Gleason score' as predictors, and the clinical variable `overall survival' as the target variable. Knowledge in the form of rules is one of the vital tasks in data mining as it provides concise statements of easily understandable and potentially valuable information. For the classification model, we derived rules from a decision tree and interpreted these rules for cancer and non-cancer patients. For the regression or survival model, we generated rules for predicting or estimating the survival time of cancer patients. In this study, cancer-relevant genes were analyzed as predictors, although various genes may interact with genes currently known to contribute to cancer. These findings have implications for assessing gene-gene interactions and gene-environment interactions of prostate cancer as well as for other types of cancer.
Md Faisal Kabir 0001, Simone A. Ludwig
IEEE BigData1
2018 Rule Discovery from Breast Cancer Risk Factors using Association Rule Mining
abstract
Breast cancer is the most common cancer in women worldwide. Prevention of breast cancer through risk factors reduction is a significant concern to decrease its impact on the population. Attaining or detecting significant information in the form of rules is the key to prevent breast cancer. Our objective is to find hidden but important knowledge of the form of rules from the risk factors data set of breast cancer. Mining rules is one of the vital tasks of data mining as rules provide concise statement of potentially important information that is easily understood by end users. In this paper, we use association rule mining, a data mining technique to attain information in the form of rules from breast cancer risk factors data that could be useful to initiate prevention strategies. We discovered rules of both breast cancer and non-breast cancer patients so that we can understand and compare the characteristics of both breast cancer and non-breast cancer individuals. The experimental results show that generated or mined rules hold the highest confidence level.
Md Faisal Kabir 0001, Simone A. Ludwig, Abu Saleh Abdullah
IEEE BigData1