Rezaur Rashid

dblp:258/4986 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
7since 2021 · last 2025
0000-0003-1343-5364ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 4 first-author · 6 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Measuring Social Media Polarization Using Large Language Models and Heuristic Rules
Jawad Chowdhury, Rezaur Rashid, Gabriel Terejanu
ASONAM (3)2
2024 An Explainable AI Data Pipeline for Multi-Level Survival Prediction of Breast Cancer Patients Using Electronic Medical Records and Social Determinants of Health Data
abstract
This study introduces an innovative explainable AI (XAI) pipeline designed to predict breast cancer survival by integrating clinical, socioeconomic, and geographic data. Using data from 10,172 patients treated at hospitals in the Memphis, Tennessee metropolitan area, the pipeline identifies key survival determinants and reveals significant survival disparities affecting Black women. Advanced machine learning models combined with SHapley Additive exPlanations (SHAP) provide actionable and interpretable insights into the role of tumor stage, socioeconomic conditions, and access to preventive care. This framework facilitates personalized survival predictions and targeted equity-focused interventions, demonstrating the potential of multi-source data integration to address health inequities and improve patient outcomes.
Soheil Hashtarkhani, Shelley White-Means, Sam Li, Rezaur Rashid, Fekede Asefa Kumsa, Cindy Lemon, Lluvia Chipman, Jill Dapremont, Brianna White, Arash Shaban-Nejad
IEEE Big Data4
2024 AI-Ready Multimodal Data Pipeline to Enrich Cancer Care
abstract
This study reports on the progress in designing and developing a framework for integrating heterogeneous datasets—structured, semi-structured, and unstructured—into an AI-ready multimodal data pipeline aimed at predicting radiation therapy interruptions (RTI) and enhancing patient care navigation. The AI-Ready dataset incorporates a broad set of information, including patient demographics, health data, clinical notes, medical imaging, and data on social determinants of health. Preliminary results indicate that this pipeline effectively integrates diverse distributed data sources, providing a foundation for training AI models capable of generating reliable and actionable predictions.
Rezaur Rashid, Soheil Hashtarkhani, Fekede Asefa Kumsa, Lokesh K. Chinthala, Brianna White, Janet A Zink, Christopher L. Brett, Robert L. Davis, David L. Schwartz, Arash Shaban-Nejad
IEEE Big Data1
2024 Quantifying Influencer Impact on Affective Polarization
abstract
In today's digital age, social media platforms playa crucial role in shaping public opinion. This study explores how discussions led by influencers on Twitter, now known as ‘X’, affect public sentiment and contribute to online polarization. We developed a counterfactual framework to analyze the polarization scores of conversations in scenarios both with and without the presence of an influential figure. Two case studies, centered on the polarizing issues of climate change and gun control, were examined. Our research highlights the significant impact these figures have on public discourse, providing valuable insights into how online discussions can influence societal divisions.
Rezaur Rashid, Joshua Melton, Ouldouz Ghorbani, Siddharth Krishnan, Shannon Reid, Gabriel Terejanu
ICMLA1
2023 Causal Feature Selection: Methods and a Novel Causal Metric Evaluation Framework
abstract
The proliferation of high-dimensional data in the era of big data has presented significant challenges for machine learning models. Feature selection methods have emerged as essential preprocessing techniques to address these challenges. However, most existing feature selection techniques primarily rely on correlations or associations between features and the target variable, overlooking the consideration of causal relationships. This study introduces a novel causal feature selection (CFS) algorithm that leverages causal structure learning to identify a subset of causal features. Our approach involves employing a causal graph discovery method to represent the causal relationships among variables and the causal effects of features on the target variable. To evaluate the effectiveness of our proposed CFS algorithm, we introduce a new evaluation criterion based on causal metrics, offering a principled and rigorous approach to assess the performance of causal feature selection methods. We empirically evaluate our algorithm using synthetic and real-world datasets, demonstrating that the truncated subsets of features selected by the CFS algorithm exhibit comparable or improved performance compared to baseline methods while utilizing fewer causal features.
Rezaur Rashid, Jawad Chowdhury, Gabriel Terejanu
DSAA1
2023 Evaluation of Induced Expert Knowledge in Causal Structure Learning by NOTEARS
Jawad Chowdhury, Rezaur Rashid, Gabriel Terejanu
ICPRAM2
2022 From Causal Pairs to Causal Graphs
abstract
Causal structure learning from observational data remains a non-trivial task due to various factors such as finite sampling, unobserved confounding factors, and measurement errors. Constraint-based and score-based methods tend to suffer from high computational complexity due to the combinatorial nature of estimating the directed acyclic graph (DAG). Motivated by the ‘Cause-Effect Pair’ NIPS 2013 Workshop on Causality Challenge, in this paper, we take a different approach and generate a probability distribution over all possible graphs informed by the cause-effect pair features proposed in response to the workshop challenge. The goal of the paper is to propose new methods based on this probabilistic information and compare their performance with traditional and state-of-the-art approaches. Our experiments, on both synthetic and real datasets, show that our proposed methods not only have statistically similar or better performances than some traditional approaches but also are computationally faster.
Rezaur Rashid, Jawad Chowdhury, Gabriel Terejanu
ICMLA1