Samah Jamal Fodeh

dblp:07/4880 · also Samah J. Fodeh · DBLP profile ↗
← Back
25ranked-venue papers
12as first author
5since 2021 · last 2026
0000-0003-4664-3143ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 20 · 9 first-author · 4 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author
YearPublicationVenuePosition
2026 EPPCMinerBen: A novel benchmark for evaluating large language models on electronic patient-provider communication via the patient portal
Samah Jamal Fodeh, Linhai Ma, Srivani Talakokkul, Jordan M. Alpert, Sarah Schellhorn
Artif. Intell. Medicine1
2024 Deploying a national clinical text processing infrastructure
abstract
OBJECTIVES: Clinical text processing offers a promising avenue for improving multiple aspects of healthcare, though operational deployment remains a substantial challenge. This case report details the implementation of a national clinical text processing infrastructure within the Department of Veterans Affairs (VA). METHODS: Two foundational use cases, cancer case management and suicide and overdose prevention, illustrate how text processing can be practically implemented at scale for diverse clinical applications using shared services. RESULTS: Insights from these use cases underline both commonalities and differences, providing a replicable model for future text processing applications. CONCLUSIONS: This project enables more efficient initiation, testing, and future deployment of text processing models, streamlining the integration of these use cases into healthcare operations. This project implementation is in a large integrated health delivery system in the United States, but we expect the lessons learned to be relevant to any health system, including smaller local and regional health systems in the United States.
Kimberly F. McManus, Johnathon Michael Stringer, Neal Corson, Samah Jamal Fodeh, Steven Steinhardt, Forrest L. Levin, Asqar S. Shotqara, Joseph D'auria, Elliot M. Fielstein, Glenn T. Gobbel, John Scott, Jodie Trafton, Tamar H. Taddei, Joseph Erdos, Suzanne Tamang
J. Am. Medical Informatics Assoc.4
2024 A roadmap to artificial intelligence (AI): Methods for designing and building AI ready data to promote fairness
Farah Kidwai-Khan, Melissa Skanderson, Cynthia Brandt, Samah Jamal Fodeh, Julie A. Womack
J. Biomed. Informatics5
2022 New JBI policy emphasizes clinically-meaningful novel machine learning methods
Allan Tucker, Thomas George Kannampallil, Samah Jamal Fodeh, Mor Peleg
J. Biomed. Informatics3
2021 COVID-19 concerns and interests differ with socioeconomic status: Twitter Analysis
Samah Jamal Fodeh, Sabrina Su, Aarthi Venkat, Lisa B. Puglisi
AMIA1
2020 Pain Assessment in Veterans Health Administration Chiropractic Clinic Documentation
Brian C. Coleman, Samah Jamal Fodeh, Anthony J. Lisi, Alicia Heapy, Cynthia Brandt
AMIA2
2020 Integrating Social Media data with Electronic Health Records to better Understand Opioid Use Disorder
Samah Jamal Fodeh, Khaled Jarad, Mengran Zhang, Stephen Holt
AMIA1
2020 Defining facets of social distancing during the COVID-19 pandemic: Twitter analysis
Jiye Kwon, Connor Grady, Josemari T. Feliciano, Samah Jamal Fodeh
J. Biomed. Informatics4
2020 Classification of Patients with Coronary Microvascular Dysfunction
abstract
While coronary microvascular dysfunction (CMD) is a major cause of ischemia, it is very challenging to diagnose due to lack of CMD-specific screening measures. CMD has been identified as one of the five priority areas of investigation in a 2014 National Research Consensus Conference on Gender-Specific Research in Emergency Care. In this study, we utilized methods from machine learning that leverage structured and unstructured narratives in clinical notes to detect patients with CMD. We have shown that structured data are not sufficient to detect CMD and integrating unstructured data in the computational model boosts the performance significantly.
Samah Jamal Fodeh, Taihua Li, Haya Jarad, Basmah Safdar
IEEE ACM Trans. Comput. Biol. Bioinform.1
2020 Guest Editorial for Selected Papers from BIOKDD 2018 and DMBIH 2018
abstract
The papers in this special issue were presented at the 2018 17th International Workshop on Data Mining in Bioinformatics (BIOKDD), held in conjunction with the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. The Workshop was held on August 20, 2018 in London, UK.
Da Yan 0001, Xin Gao 0001, Samah Jamal Fodeh, Jake Yue Chen
IEEE ACM Trans. Comput. Biol. Bioinform.3
2019 Preliminary chart review for Natural Language Processing development among electronic health records with documentation of marijuana use
Termeh Feinberg, Joseph L. Goulet, Amy Justice, Samah Jamal Fodeh, Lori Bastian, Qing T. Zeng, Cynthia Brandt
AMIA4
2019 Presecription opioid tapering: Mining structured and unstructured data
Samah Jamal Fodeh, Jennifer Edelman, William Becker, Cynthia Brandt, Sally Haskell
AMIA1
2018 Automatic extraction of informal topics from online suicidal ideation
abstract
BACKGROUND: Suicide is an alarming public health problem accounting for a considerable number of deaths each year worldwide. Many more individuals contemplate suicide. Understanding the attributes, characteristics, and exposures correlated with suicide remains an urgent and significant problem. As social networking sites have become more common, users have adopted these sites to talk about intensely personal topics, among them their thoughts about suicide. Such data has previously been evaluated by analyzing the language features of social media posts and using factors derived by domain experts to identify at-risk users. RESULTS: In this work, we automatically extract informal latent recurring topics of suicidal ideation found in social media posts. Our evaluation demonstrates that we are able to automatically reproduce many of the expertly determined risk factors for suicide. Moreover, we identify many informal latent topics related to suicide ideation such as concerns over health, work, self-image, and financial issues. CONCLUSIONS: These informal topics topics can be more specific or more general. Some of our topics express meaningful ideas not contained in the risk factors and some risk factors do not have complimentary latent topics. In short, our analysis of the latent topics extracted from social media containing suicidal ideations suggests that users of these systems express ideas that are complementary to the topics defined by experts but differ in their scope, focus, and precision of language.
Reilly Grant, David Kucher, Ana M. Leon, Jonathan Gemmell, Daniela Raicu, Samah Jamal Fodeh
BMC Bioinform.6
2018 Exploiting MEDLINE for gene molecular function prediction via NMF based multi-label classification
Samah Jamal Fodeh, Aditya Tiwari
J. Biomed. Informatics1
2016 Classification of radiology reports for falls in an HIV study cohort
abstract
OBJECTIVE: To identify patients in a human immunodeficiency virus (HIV) study cohort who have fallen by applying supervised machine learning methods to radiology reports of the cohort. METHODS: We used the Veterans Aging Cohort Study Virtual Cohort (VACS-VC), an electronic health record-based cohort of 146 530 veterans for whom radiology reports were available (N=2 977 739). We created a reference standard of radiology reports, represented each report by a feature set of words and Unified Medical Language System concepts, and then developed several support vector machine (SVM) classifiers for falls. We compared mutual information (MI) ranking and embedded feature selection approaches. The SVM classifier with MI feature selection was chosen to classify all radiology reports in VACS-VC. RESULTS: Our SVM classifier with MI feature selection achieved an area under the curve score of 97.04 on the test set. When applied to all the radiology reports in VACS-VC, 80 416 of these reports were classified as positive for a fall. Of these, 11 484 were associated with a fall-related external cause of injury code (E-code) and 68 932 were not, corresponding to 29 280 patients with potential fall-related injuries who could not have been found using E-codes. DISCUSSION: Feature selection was crucial to improving the classifier's performance. Feature selection with MI allowed us to select the number of discriminative features to use for classification, in contrast to the embedded feature selection method, in which the number of features is chosen automatically. CONCLUSION: Machine learning is an effective method of identifying patients who have suffered a fall. The development of this classifier supplements the clinical researcher's toolkit and reduces dependence on under-coded structured electronic health record data.
Jonathan Bates, Samah Jamal Fodeh, Cynthia Brandt, Julie A. Womack
J. Am. Medical Informatics Assoc.2
2016 Mining Big Data in biomedicine and health care
Samah Jamal Fodeh, Qing T. Zeng
J. Biomed. Informatics1
2015 Using the Adverse Event Reporting System: Can Analysis be Streamlined by Text Processing?
Andrea L. Benin, Samah Jamal Fodeh, Michele Koss, Kyle Lee, Perry L. Miller, Cynthia Brandt
AMIA2
2015 Feature Selection Based LapSVM to Classify Medical Event Reports and Enhance Patient Safety
Samah Jamal Fodeh, Cynthia Brandt, Perry L. Miller, Michele Koss, Andrea L. Benin
AMIA1
2015 Exploring Healthcare Mobility in the US to Improve Quality of Care: Preliminary Results
Karen H. Wang, Constance Carroll, Brenda Fenton, Samah Jamal Fodeh, Joseph Erdos, Marcella Nunez-Smith, Amy Justice, Cynthia Brandt
AMIA4
2013 Complementary ensemble clustering of biomedical data
Samah Jamal Fodeh, Cynthia Brandt, Thaibinh Luong, Ali Haddad, Martin H. Schultz, Terrence Murphy, Michael Krauthammer
J. Biomed. Informatics1
2012 Analysis of VA Telephone Call Notes using Topic Modeling
Samah Jamal Fodeh, Sharmila Chatterjee, Cynthia Brandt
AMIA1
2011 On ontology-driven document clustering using core semantic features
Samah Jamal Fodeh, William F. Punch, Pang-Ning Tan
Knowl. Inf. Syst.1
2007 Incorporating Background Knowledge for Subjective Rule Evaluation
abstract
Association rule mining is the task of finding interesting relationships hidden in large transaction databases. Despite the significant progress made in this field, one of the fundamental challenges that remain unresolved is the rule evaluation problem. Most notably, it is difficult to discriminate rules that are known to the domain experts from those that are unexpected. In this paper, we propose a framework called MIR that incorporates background knowledge acquired from an authoritative source into the rule evaluation task. We illustrate the advantages of using the framework in the medical informatics domain, where the rules are extracted from an electronic medical records (EMR) database while the domain knowledge is automatically acquired from the MEDLINE repository of biomedical citations.
Samah Jamal Fodeh, Pang-Ning Tan
ICTAI (2)1
2007 A Probabilistic Substructure-Based Approach for Graph Classification
abstract
Graph classification is an important data mining task that has attracted considerable attention recently. This paper presents a probabilistic substructure-based approach for classifying graph-based data. More specifically, we use a frequent subgraph mining algorithm to extract substructure based descriptors and apply the maximum entropy principle to build a classification model from the frequent subgraphs. We perform extensive experiments to compare the performance of the proposed approach against existing feature vector methods using AdaBoost and support vector machine.
H. D. K. Moonesinghe, Hamed Valizadegan, Samah Jamal Fodeh, Pang-Ning Tan
ICTAI (1)3
2006 Frequent Closed Itemset Mining Using Prefix Graphs with an Efficient Flow-Based Pruning Strategy
abstract
This paper presents PGMiner, a novel graph-based algorithm for mining frequent closed itemsets. Our approach consists of constructing a prefix graph structure and decomposing the database to variable length bit vectors, which are assigned to nodes of the graph. The main advantage of this representation is that the bit vectors at each node are relatively shorter than those produced by existing vertical mining methods. This facilitates fast frequency counting of itemsets via intersection operations. We also devise several inter- node and intra-node pruning strategies to substantially reduce the combinatorial search space. Unlike other existing approaches, we do not need to store in memory the entire set of closed itemsets that have been mined so far in order to check whether a candidate itemset is closed. This dramatically reduces the memory usage of our algorithm, especially for low support thresholds. Our experiments using synthetic and real-world data sets show that PGMiner outperforms existing mining algorithms by as much as an order of magnitude and is scalable to very large databases.
H. D. K. Moonesinghe, Samah Jamal Fodeh, Pang-Ning Tan
ICDM2