VLDB 2026 Research / reviewers in the wild / expert
Mahshad Koohi Habibi Dehkordi
dblp:336/5573
· DBLP profile ↗
6ranked-venue papers
4as first author
6since 2021 · last 2025
0000-0003-3489-0892ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5 · 4 first-author · 5 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Enhancing Patient Comprehension of Discharge Notes with a Retrieval-Augmented LLM ApproachabstractDischarge notes summarize reasons for hospital stays and diagnosis, include medical history and treatment received, and provide instructions for follow-up care and medications. These notes are often difficult for patients to understand due to medical jargon, abbreviations, and complex structures. Despite recommendations that patient-facing materials be written at a sixth-grade reading level, most discharge notes do not meet this target, creating barriers to comprehension. Large language models (LLMs) can simplify clinical text but risk omitting important information, introducing errors, or producing hallucinations. We propose a retrieval-augmented generation (RAG) pipeline to enhance LLM-based simplification of discharge notes by appending biomedical definitions to prompts. Our pipeline combines (1) direct UMLS concept linking with SciSpaCy to extract canonical names and definitions, and (2) FAISS-based semantic retrieval of additional relevant definitions not explicitly detected by SciSpaCy. These definitions are integrated into prompts for ChatGPT-4o to generate simplified summaries at a sixth-grade reading level. We evaluated this approach on 20 MIMIC-III discharge notes, comparing RAG-augmented outputs (R_notes) with baseline LLM simplifications (B_notes). Manual review measured completeness and correctness, readability was assessed with Flesch-Kincaid Grade Level (FKGL) and Flesch Reading Ease, and LLM-based judgments evaluated clarity. Results show that R_notes achieved higher completeness (77.6 % vs. 72.7 %, the completeness of$R$_notes is higher than that of B_notes in 16 cases with statistical significant), fewer errors (2 vs. 6), and better contextualization of medical terms ($\mathbf{4. 3 ~ v s. ~} \mathbf{3. 6}$on a 5 -point scale). Readability also improved, more closely aligning with the sixth-grade target. These improvements highlight the potential of RAG-enhanced prompting to strengthen patient understanding and engagement with their health information. Mahshad Koohi Habibi Dehkordi, Yehoshua Perl, Fadi P. Deek |
BIBM | 1 |
| 2025 | Optimizing Manual Review Using Machine Learning in Interface Terminology Curation for Automatic EHR HighlightingabstractDischarge notes are dense, information-rich documents that contain patient histories, diagnoses, treatments, and clinical observations, as well as post-discharge care instructions. While they provide essential data for clinical decision-making and research, they are often written using abbreviations and complex medical jargon, making them difficult for patients to interpret. Automatic highlighting of discharge notes enhances information accessibility, supports summarization and simplification, and improves clinical interoperability of the notes. Achieving accurate highlighting requires terminologies that include fine-granularity phrases, which existing reference terminologies such as SNOMED CT lack. To address this limitation, in our previous work, we proposed the Cardiology Interface Terminology (CIT), tailored for accurate highlighting of discharge notes of cardiology patients. Candidate concepts to be added to CIT were extracted from notes through concatenation and anchoring operations, with each phrase undergoing automatic and manual review before inclusion in CIT. Manual review of these phrases is highly time-consuming and costly process. In this study, we propose a Machine Learning (ML)-assisted approach to reduce the manual review efforts involved in terminology curation. We trained a Neural Network (NN) model on varying subsets of phrases generated through concatenation and anchoring, to determine the minimum number of phrases that must be manually reviewed to effectively train the ML model to label the remaining phrases automatically. The optimal batch sizes were identified as$\mathbf{6, 0 0 0}$(out of$\mathbf{2 8, 6 1 7}$) for concatenation and$\mathbf{3, 0 0 0}$(out of 9,845) for anchoring. The resulting terminology (CIT${}_{\text {ML2+ }}$) achieved a coverage of 68.74 % and breadth of 1.6 on the test dataset, closely matching the fully manually curated CIT+ (coverage 70.21 %, breadth 1.6), with comparable completeness (97.4 % vs. 98.6 %) and conciseness (84.1 % vs. 83.6 %). These findings demonstrate that substantial reductions in manual review can be achieved without compromising highlighting quality, providing a scalable and efficient framework for curating interface terminologies across diverse medical domains. Mahshad Koohi Habibi Dehkordi, Yehoshua Perl, Fadi P. Deek |
BIBM | 1 |
| 2024 | Enhancing patient Comprehension: An effective sequential prompting approach to simplifying EHRs using LLMsabstractElectronic Health Record (EHR) notes often contain complex medical language, making them difficult to understand for patients lacking medical background. Simplifying EHR notes to a 6th-grade reading level is recommended by the American Medical Association to enhance patient comprehension and engagement. Large Language Models (LLMs) show promise in achieving this goal but also face challenges, such as missing and generating false information. In our previous work, we have shown that providing LLMs with highlighted EHRs, where the important information is highlighted, results in more accurate summaries compared to summarizing unhighlighted notes. In this study, we simplify highlighted EHRs with LLMs, specifically ChatGPT-4o, using two approaches: two-step simplification (sequential) and one-step (CoT-based) simplification. In the sequential approach, we generate a structured summary of the highlighted EHR, as a first step, and then we convert this summary into language suitable for a 6th-grade reader, as a second step. In the CoT-based approach, we convert the highlighted EHR into a structured summary understandable for a 6th-grade reader in one step. Evaluating the simplified notes obtained from the two approaches, the sequential approach shows higher completeness (82.35% vs. 75.89%) and correctness, as well as better readability scores (FKGL: 7.72 vs. 10.73; Flesch: 67.71 vs. 45.31) and higher average understandability ratings from ChatGPT-4 (3.92 vs. 3.28), demonstrating its overall superiority in simplifying notes. Mahshad Koohi Habibi Dehkordi, Shuxin Zhou, Yehoshua Perl, Fadi P. Deek, Andrew J. Einstein, Gai Elhanan, Zhe He 0001, Hao Liu 0025 |
BIBM | 1 |
| 2024 | Using clinical entity recognition for curating an interface terminology to aid fast skimming of EHRsabstractHighlighting of Electronic Health Records (EHRs) involves marking essential content of EHR notes, corresponding to concepts of a clinical terminology. However, employing the best clinical terminology (SNOMED CT) for highlighting EHRs, captures only a portion of their crucial content. In this paper, we describe the curation of a Cardiology Interface Terminology (CIT) dedicated to the application of highlighting EHRs of cardiology patients. We utilize a Clinical-Named Entity Recognition (Clinical NER) approach for extracting phrases, of higher granularity than SNOMED CT concepts, from EHRs, for enriching CIT. For this purpose, we train a neural network model with BIOE-tagged (Beginning, Inside, End, and Outside) cardiology entities. Transfer Learning can be used to facilitate the curation of an interface terminology for highlighting EHRs for other specialties e.g. Nephrology. Large-scale highlighting enables overworked physicians and other healthcare providers to fast skim the dense volume of EHRs they regularly read. Secondary research and EHRs interoperability are other applications that can be supported by highlighting. Navya Martin Kollapally, Mahshad Koohi Habibi Dehkordi, Yehoshua Perl, James Geller, Fadi P. Deek, Hao Liu 0025, Vipina Kuttichi Keloth, Gai Elhanan, Andrew J. Einstein, Shuxin Zhou |
BIBM | 2 |
| 2023 | Using annotation for computerized support for fast skimming of cardiology electronic health record notesabstractUnder the circumstances prevalent in current healthcare, medical professionals such as physicians and nurses need to read large numbers of Electronic Health Record (EHR) notes. This demand fosters a situation in which providers typically do not read a whole note but quickly skim it, to capture its essential content. Consequently, by fast skimming, one may miss a critical medical fact. Annotation highlights important content in EHR notes, and enables healthcare professionals to perform fast skimming, thereby minimizing the risk of missing critical information, which is detrimental to patient care. We designed the Cardiology Interface Terminology (CIT) for the purpose of annotation of cardiology EHRs. We emphasize that by annotation we refer to highlighting important information in the EHR notes for enabling fast skimming rather than just recognizing names of diseases, drugs, etc. from Reference Terminologies as is usually done by existing Named Entity Recognition (NER) systems. The CIT design starts with the cardiology components of SNOMED CT. It is enhanced by mining phrases from cardiology EHRs, as potential CIT concepts, which are of higher granularity than SNOMED concepts. Machine learning (ML), the state of the art technique for mining concepts from EHRs, requires training data. However, there is no training data for designing CIT. In the first stage, we introduce an innovative semi-automatic method for mining concepts from EHRs, to replace costly manual mining. The only manual portion is the review of the automatically mined phrases, before their insertion as CIT concepts. The effectiveness of annotation of cardiology EHRs with CIT was evaluated utilizing proper metrics, and compared to annotation with SNOMED CT. In a future second stage, ML mining techniques will be used for enhancing CIT with extra concepts from EHRs, utilizing the concepts added in the first stage as training data. This work focuses on a novel semi-automated method to design the Cardiology Interface Terminology (CIT) for annotation of cardiology EHRs to support fast skimming of EHR notes. Similar interface terminologies for other medical specialties could be obtained from CIT using Transfer Learning. Mahshad Koohi Habibi Dehkordi, Andrew J. Einstein, Shuxin Zhou, Gai Elhanan, Yehoshua Perl, Vipina Kuttichi Keloth, James Geller, Hao Liu 0025 |
BIBM | 1 |
| 2023 | Secured location-aware mobility-enabled RPL
Erfan Arvan, Mahshad Koohi Habibi Dehkordi, Saeed Jalili |
J. Netw. Comput. Appl. | 2 |