Mohammad Hassan Murad

dblp:157/0965 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
5since 2021 · last 2025
0000-0001-5502-5975ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 5 since 2021
YearPublicationVenuePosition
2025 Collaborative large language models for automated data extraction in living systematic reviews
abstract
OBJECTIVE: Data extraction from the published literature is the most laborious step in conducting living systematic reviews (LSRs). We aim to build a generalizable, automated data extraction workflow leveraging large language models (LLMs) that mimics the real-world 2-reviewer process. MATERIALS AND METHODS: A dataset of 10 trials (22 publications) from a published LSR was used, focusing on 23 variables related to trial, population, and outcomes data. The dataset was split into prompt development (n = 5) and held-out test sets (n = 17). GPT-4-turbo and Claude-3-Opus were used for data extraction. Responses from the 2 LLMs were considered concordant if they were the same for a given variable. The discordant responses from each LLM were provided to the other LLM for cross-critique. Accuracy, ie, the total number of correct responses divided by the total number of responses, was computed to assess performance. RESULTS: In the prompt development set, 110 (96%) responses were concordant, achieving an accuracy of 0.99 against the gold standard. In the test set, 342 (87%) responses were concordant. The accuracy of the concordant responses was 0.94. The accuracy of the discordant responses was 0.41 for GPT-4-turbo and 0.50 for Claude-3-Opus. Of the 49 discordant responses, 25 (51%) became concordant after cross-critique, increasing accuracy to 0.76. DISCUSSION: Concordant responses by the LLMs are likely to be accurate. In instances of discordant responses, cross-critique can further increase the accuracy. CONCLUSION: Large language models, when simulated in a collaborative, 2-reviewer workflow, can extract data with reasonable performance, enabling truly "living" systematic reviews.
Umair Ayub, Syed Arsalan Ahmed Naqvi, Kaneez Zahra Rubab Khakwani, Zaryab bin Riaz Sipra, Ammad Raina, Sihan Zhou, Amir Saeidi, Bashar Hasan, Robert Bryan Rumble, Danielle S. Bitterman, Jeremy L. Warner, Jia Zou 0001, Amye J. Tevaarwerk, Konstantinos Leventakos, Kenneth L. Kehl, Jeanne M. Palmer, Mohammad Hassan Murad, Chitta Baral, Irbaz Bin Riaz
J. Am. Medical Informatics Assoc.19
2022 Real-time Exploration of Pairwise Meta-analysis Results by Applying Serverless Architecture Design
Irbaz Bin Riaz, Syed Arsalan Ahmed Naqvi, Rabbia Siddiqi, Noureen Asghar, Mohammad Hassan Murad, Mahnoor Islam
AMIA6
2022 A Hybrid Approach to Semi-automate the Evaluation of the Certainty of Evidence for Living Systematic Reviews and Meta-analysis
Irbaz Bin Riaz, Syed Arsalan Ahmed Naqvi, Rabbia Siddiqi, Noureen Asghar, Mahnoor Islam, Mohammad Hassan Murad
AMIA7
2021 A Hybrid Approach to Semi-Automate the Screening Process for Living Systematic Reviews and Meta-Analysis
Irbaz Bin Riaz, Syed Arsalan Ahmed Naqvi, Rabbia Siddiqi, Noureen Asghar, Mohammad Hassan Murad
AMIA6
2021 An Interactive Data Extraction System to Create the Living Systematic Reviews and Meta-Analysis
Irbaz Bin Riaz, Syed Arsalan Ahmed Naqvi, Rabbia Siddiqi, Noureen Asghar, Mohammad Hassan Murad
AMIA6
2020 A Living Network Meta-Analysis of First Line Treatment of Metastatic Kidney Cancer
Irbaz Bin Riaz, Rabbia Siddiqi, Vitaly Herasevich, Per Olav Vandvik, Victor Montori, Allan Bryce, Mohammad Hassan Murad
AMIA9
2014 Reducing the Screening Burden of Systematic Review with a Multiple-level Relevance Ranking System
Dingcheng Li, Feichen Shen, Mohammad Hassan Murad
AMIA4
2014 Towards a multi-level framework for supporting systematic review - A pilot study
abstract
As part of the forefront of evidence-based medicine, a systematic review (SR) identifies, appraises, and synthesizes all the available literature relevant to a question of interest in a transparent and systematic way. A time-consuming step in conducting systematic review (SR) is the manual article screening from a list of potentially relevant articles retrieved by librarians. In this study, we propose a multi-level SR supporting framework including three levels: Level 1 - ranking with multiple metrics aiming to assist the screening process by increasing the efficiency without compromising the validity; Level 2 - topic analysis for discovering distributed semantics; and Level 3 - network analysis on relation extracted for comprehensive semantic summarization. The sensitivities in two case studies for relevance ranking reached above 80%, while the screening burdens were lowered to 25% based on a crude relevance ranking approach. Topic analysis based on Latent Dirichlet Allocations (LDAs) showed high consistency among the domain experts' selection of articles. The predicate relation network summarized important predicate relations between medical concepts.
Dingcheng Li, Feichen Shen, Mohammad Hassan Murad
BIBM4