Masoud Rouhizadeh

dblp:66/10108 · DBLP profile ↗
← Back
17ranked-venue papers
4as first author
9since 2021 · last 2026
0000-0002-9006-6112ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 9 · 8 since 2021Artificial intelligence and machine learning · 7 · 4 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Beyond metrics to methods: a scoping review of transformers and large language models for detection of social drivers of health in clinical notes
abstract
OBJECTIVE: This scoping review aimed to (1) map current applications of transformers and large language models (LLMs) for extracting social drivers of health (SDOH) from clinical text, (2) benchmark model performance across SDOH domains, and (3) evaluate methodological rigor to identify research gaps and inform clinical deployment. MATERIALS AND METHODS: We searched PubMed, Web of Science, Embase, Scopus, and IEEE Xplore for studies applying transformers or LLMs to detect SDOH in clinical narratives. We developed a novel methodological framework integrating (1) hierarchical classification of SDOH domains and transformer/LLM architectures, (2) systematic synthesis of performance metrics, and (3) a 7-domain instrument assessing internal validity, external validity, and reporting transparency. RESULTS: Forty-two studies met inclusion criteria. Performance varied substantially across SDOH domains. Behavioral Factors achieved the highest median F1-score (0.87), while Health Care Access and Quality showed the lowest performance and greatest variability (median F1 = 0.59). Research concentrated in the United States (85.7%), relied predominantly on private institutional datasets (69%), and focused primarily on critical care populations (45.2%). Methodological assessment revealed critical gaps; only 29% of studies provided annotation guidelines, 24% assessed fairness across demographic groups, and 21% performed external validation. DISCUSSION: Smaller open-source transformer models show promise for democratizing SDOH detection by achieving competitive performance at lower costs while enabling secure local deployment in resource-limited settings. Advancing clinical readiness requires standardized reporting practices, diverse benchmark datasets across care settings, and systematic equity evaluation to prevent perpetuating health disparities. CONCLUSION: Transformer and LLM performance for SDOH detection varied substantially across domains, with encoder-based models excelling at structured tasks and decoder-only models at linguistically complex tasks. Critical gaps in fairness assessment, external validation, and dataset diversity restrict generalizability and readiness for widespread clinical deployment.
Ahmed Farrag, Elham Hatef, Amie J. Goodin, Masoud Rouhizadeh
J. Am. Medical Informatics Assoc.5
2025 Enhancing suicidal behavior detection in EHRs: A multi-label NLP framework with transformer models and semantic retrieval-based annotation
Kimia Zandbiglari, Shobhan Kumar, Muhammad Bilal 0012, Amie J. Goodin, Masoud Rouhizadeh
J. Biomed. Informatics5
2023 An open natural language processing (NLP) framework for EHR-based clinical research: a case demonstration using the National COVID Cohort Collaborative (N3C)
abstract
Despite recent methodology advancements in clinical natural language processing (NLP), the adoption of clinical NLP models within the translational research community remains hindered by process heterogeneity and human factor variations. Concurrently, these factors also dramatically increase the difficulty in developing NLP models in multi-site settings, which is necessary for algorithm robustness and generalizability. Here, we reported on our experience developing an NLP solution for Coronavirus Disease 2019 (COVID-19) signs and symptom extraction in an open NLP framework from a subset of sites participating in the National COVID Cohort (N3C). We then empirically highlight the benefits of multi-site data for both symbolic and statistical methods, as well as highlight the need for federated annotation and evaluation to resolve several pitfalls encountered in the course of these efforts.
Sijia Liu 0002, Andrew Wen, Liwei Wang 0010, Sunyang Fu, Robert T. Miller, Andrew E. Williams, Daniel R. Harris, Ramakanth Kavuluru, Noor Abu-El-Rub, Dalton Schutte, Rui Zhang 0028, Masoud Rouhizadeh, John D. Osborne, Yongqun He, Umit Topaloglu, Stephanie S. Hong, Joel H. Saltz, Thomas Schaffter, Emily R. Pfaff, Christopher G. Chute, Tim Duong, Melissa A. Haendel, Rafael Fuentes, Peter Szolovits, Hua Xu 0001
J. Am. Medical Informatics Assoc.14
2023 Developing and validating a natural language processing algorithm to extract preoperative cannabis use status documentation from unstructured narrative clinical notes
abstract
OBJECTIVE: This study aimed to develop a natural language processing algorithm (NLP) using machine learning (ML) techniques to identify and classify documentation of preoperative cannabis use status. MATERIALS AND METHODS: We developed and applied a keyword search strategy to identify documentation of preoperative cannabis use status in clinical documentation within 60 days of surgery. We manually reviewed matching notes to classify each documentation into 8 different categories based on context, time, and certainty of cannabis use documentation. We applied 2 conventional ML and 3 deep learning models against manual annotation. We externally validated our model using the MIMIC-III dataset. RESULTS: The tested classifiers achieved classification results close to human performance with up to 93% and 94% precision and 95% recall of preoperative cannabis use status documentation. External validation showed consistent results with up to 94% precision and recall. DISCUSSION: Our NLP model successfully replicated human annotation of preoperative cannabis use documentation, providing a baseline framework for identifying and classifying documentation of cannabis use. We add to NLP methods applied in healthcare for clinical concept extraction and classification, mainly concerning social determinants of health and substance use. Our systematically developed lexicon provides a comprehensive knowledge-based resource covering a wide range of cannabis-related concepts for future NLP applications. CONCLUSION: We demonstrated that documentation of preoperative cannabis use status could be accurately identified using an NLP algorithm. This approach can be employed to identify comparison groups based on cannabis exposure for growing research efforts aiming to guide cannabis-related clinical practices and policies.
Ruba Sajdeya, Mamoun T. Mardini, Patrick James Tighe, Ronald L. Ison, Sebastian Jugl, Gao Hanzhi, Kimia Zandbiglari, Farzana Islam Adiba, Almut G. Winterstein, Thomas A Pearson, Robert L. Cook 0002, Masoud Rouhizadeh
J. Am. Medical Informatics Assoc.13
2023 Representing and utilizing clinical textual data for real world studies: An OHDSI approach
Vipina Kuttichi Keloth, Juan M. Banda, Michael J. Gurley, Paul M. Heider, Georgina Kennedy, Timothy A. Miller, Karthik Natarajan, Olga V. Patterson, Yifan Peng 0002, Kalpana Raja, Ruth M. Reeves, Masoud Rouhizadeh, Jianlin Shi, Yanshan Wang, Wei-Qi Wei, Andrew E. Williams, Rui Zhang 0028, Rimma Belenkaya, Christian G. Reich, Clair Blacketer, Patrick B. Ryan, George Hripcsak, Noémie Elhadad, Hua Xu 0001
J. Biomed. Informatics14
2022 Application of Natural Language Processing to Identify Social Needs from The Electronic Health Record's Free-Text Notes
Geoffrey M. Gray, Luis M. Ahumada, Ayah Zirikly, Masoud Rouhizadeh, Thomas Richards, Elham Hatef
AMIA4
2021 Assessing the Documentation of Social Needs in Electronic Health Records' Unstructured Data: A Collaboration of Johns Hopkins Health System and Kaiser Permanente
Elham Hatef, Masoud Rouhizadeh, Claudia Nau, Fagen Xie, Ariadna Padilla, Lindsay Joe Lyons, Christopher Rouillard, Mahmoud Abu-Nasser, Hadi Kharrazi, Jonathan P. Weiner, Douglas Roblin
AMIA2
2021 Persian SemCor: A Bag of Word Sense Annotated Corpus for the Persian Language
abstract
Supervised approaches usually achieve the best performance in the Word Sense Disambiguation problem.However, the unavailability of large sense annotated corpora for many low-resource languages make these approaches inapplicable for them in practice.In this paper, we mitigate this issue for the Persian language by proposing a fully automatic approach for obtaining Persian Sem-Cor (PerSemCor), as a Persian Bag-of-Word (BoW) sense-annotated corpus.We evaluated PerSemCor both intrinsically and extrinsically and showed that it can be effectively used as training sets for Persian supervised WSD systems.To encourage future research on Persian Word Sense Disambiguation, we release the PerSemCor in nlp.sbu.ac.ir .
Hossein Rouhizadeh, Mehrnoush Shamsfard, Mahdi Dehghan, Masoud Rouhizadeh
GWC4
2021 COVID-19 SignSym: a fast adaptation of a general clinical NLP tool to identify and normalize COVID-19 signs and symptoms to OMOP common data model
abstract
The COVID-19 pandemic swept across the world rapidly, infecting millions of people. An efficient tool that can accurately recognize important clinical concepts of COVID-19 from free text in electronic health records (EHRs) will be valuable to accelerate COVID-19 clinical research. To this end, this study aims at adapting the existing CLAMP natural language processing tool to quickly build COVID-19 SignSym, which can extract COVID-19 signs/symptoms and their 8 attributes (body location, severity, temporal expression, subject, condition, uncertainty, negation, and course) from clinical text. The extracted information is also mapped to standard concepts in the Observational Medical Outcomes Partnership common data model. A hybrid approach of combining deep learning-based models, curated lexicons, and pattern-based rules was applied to quickly build the COVID-19 SignSym from CLAMP, with optimized performance. Our extensive evaluation using 3 external sites with clinical notes of COVID-19 patients, as well as the online medical dialogues of COVID-19, shows COVID-19 SignSym can achieve high performance across data sources. The workflow used for this study can be generalized to other use cases, where existing clinical natural language processing tools need to be customized for specific information needs within a short time. COVID-19 SignSym is freely accessible to the research community as a downloadable package (https://clamp.uth.edu/covid/nlp.php) and has been used by 16 healthcare organizations to support clinical research of COVID-19.
Noor Abu-El-Rub, Josh Gray, Huy Anh Pham, Yujia Zhou 0003, Frank J. Manion, Xing Song, Hua Xu 0001, Masoud Rouhizadeh, Yaoyun Zhang
J. Am. Medical Informatics Assoc.10
2018 Identifying Locus of Control in Social Media Language
abstract
Individuals express their locus of control, or "control", in their language when they identify whether or not they are in control of their circumstances.Although control is a core concept underlying rhetorical style, it is not clear whether control is expressed by how or by what authors write.We explore the roles of syntax and semantics in expressing users' sense of control -i.e.being "controlled by" or "in control of" their circumstances-in a corpus of annotated Facebook posts.We present rich insights into these linguistic aspects and find that while the language signaling control is easy to identify, it is more challenging to label it is internally or externally controlled, with lexical features outperforming syntactic features at the task.Our findings could have important implications for studying selfexpression in social media.
Masoud Rouhizadeh, Kokil Jaidka, H. Andrew Schwartz, Anneke Buffone, Lyle H. Ungar
EMNLP1
2018 Modeling and Visualizing Locus of Control with Facebook Language
Kokil Jaidka, Anneke Buffone, Johannes C. Eichstaedt, Masoud Rouhizadeh, Lyle H. Ungar
ICWSM4
2017 Assessing Objective Recommendation Quality through Political Forecasting
abstract
Recommendations are often rated for their subjective quality, but few researchers have studied quality in terms of objective utility.We explore quality assessment with respect to both subjective (i.e.users' ratings) and objective (i.e., did it influence?did it improve decisions?)metrics in a massive online geopolitical forecasting system, ultimately comparing linguistic characteristics of each quality metric.Using a variety of features, we predict all types of quality with better accuracy than the simple yet strong baseline of recommendation length.For example, more complex sentence constructions, as evidenced by subordinate conjunctions, are characteristic of recommendations leading to objective improvements in forecasting.Our analyses also reveal rater biases; for example, forecasters are subjectively biased in favor of recommendations mentioning business deals and material things, even though such recommendations do not indeed prove any more useful objectively.
H. Andrew Schwartz, Masoud Rouhizadeh, Michael Bishop, Philip Tetlock, Barbara A. Mellers, Lyle H. Ungar
EMNLP2
2016 Using Syntactic and Semantic Context to Explore Psychodemographic Differences in Self-reference
abstract
Psychological analysis of language has repeatedly shown that an individual's rate of mentioning 1st person singular pronouns predicts a wealth of important demographic and psychological factors.However, these analyses are performed out of context -syntactic and semantic -which may change the magnitude or even direction of such relationships.In this paper, we put "pronouns in their context", exploring the relationship between self-reference and age, gender, and depression depending on syntactic position and verbal governor.We find that pronouns are overall more predictive when taking dependency relations and verb semantic categories into account, and, the direction of the relationship can change depending on the semantic class of the verbal governor.
Masoud Rouhizadeh, Lyle H. Ungar, Anneke Buffone, H. Andrew Schwartz
EMNLP1
2014 Computational analysis of trajectories of linguistic development in autism
abstract
Deficits in semantic and pragmatic expression are among the hallmark linguistic features of autism. Recent work in deriving computational correlates of clinical spoken language measures has demonstrated the utility of automated linguistic analysis for characterizing the language of children with autism. Most of this research, however, has focused either on young children still acquiring language or on small populations covering a wide age range. In this paper, we extract numerous linguistic features from narratives produced by two groups of children with and without autism from two narrow age ranges. We find that although many differences between diagnostic groups remain constant with age, certain pragmatic measures, particularly the ability to remain on topic and avoid digressions, seem to improve. These results confirm findings reported in the psychology literature while underscoring the need for careful consideration of the age range of the population under investigation when performing clinically oriented computational analysis of spoken language.
Emily Tucker Prud'hommeaux, Eric Morley, Masoud Rouhizadeh, Laura Silverman, Jan P. H. van Santen, Brian Roark, Richard Sproat, Sarah Kauper, Rachel DeLaHunta
SLT3
2013 Distributional semantic models for the evaluation of disordered language
Masoud Rouhizadeh, Emily Tucker Prud'hommeaux, Brian Roark, Jan P. H. van Santen
HLT-NAACL1
2012 Annotation Tools and Knowledge Representation for a Text-To-Scene System
Bob Coyne, Alex Klapheke, Masoud Rouhizadeh, Richard Sproat, Daniel Bauer 0002
COLING3
2011 Collecting Semantic Information for Locations in the Scenario-Based Lexical Knowledge Resource of a Text-to-Scene Conversion System
Masoud Rouhizadeh, Bob Coyne, Richard Sproat
KES (4)1