VLDB 2026 Research / reviewers in the wild / expert
Cosmin Adrian Bejan
dblp:31/5154 · also Cosmin Bejan
· DBLP profile ↗
27ranked-venue papers
14as first author
4since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 17 · 9 first-author · 4 since 2021Artificial intelligence and machine learning · 10 · 5 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | PheMIME: an interactive web app and knowledge base for phenome-wide, multi-institutional multimorbidity analysisabstractOBJECTIVES: To address the need for interactive visualization tools and databases in characterizing multimorbidity patterns across different populations, we developed the Phenome-wide Multi-Institutional Multimorbidity Explorer (PheMIME). This tool leverages three large-scale EHR systems to facilitate efficient analysis and visualization of disease multimorbidity, aiming to reveal both robust and novel disease associations that are consistent across different systems and to provide insight for enhancing personalized healthcare strategies. MATERIALS AND METHODS: PheMIME integrates summary statistics from phenome-wide analyses of disease multimorbidities, utilizing data from Vanderbilt University Medical Center, Mass General Brigham, and the UK Biobank. It offers interactive and multifaceted visualizations for exploring multimorbidity. Incorporating an enhanced version of associationSubgraphs, PheMIME also enables dynamic analysis and inference of disease clusters, promoting the discovery of complex multimorbidity patterns. A case study on schizophrenia demonstrates its capability for generating interactive visualizations of multimorbidity networks within and across multiple systems. Additionally, PheMIME supports diverse multimorbidity-based discoveries, detailed further in online case studies. RESULTS: The PheMIME is accessible at https://prod.tbilab.org/PheMIME/. A comprehensive tutorial and multiple case studies for demonstration are available at https://prod.tbilab.org/PheMIME_supplementary_materials/. The source code can be downloaded from https://github.com/tbilab/PheMIME. DISCUSSION: PheMIME represents a significant advancement in medical informatics, offering an efficient solution for accessing, analyzing, and interpreting the complex and noisy real-world patient data in electronic health records. CONCLUSION: PheMIME provides an extensive multimorbidity knowledge base that consolidates data from three EHR systems, and it is a novel interactive tool designed to analyze and visualize multimorbidities across multiple EHR datasets. It stands out as the first of its kind to offer extensive multimorbidity knowledge integration with substantial support for efficient online analysis and interactive visualization. Nick Strayer, Tess Vessels, Karmel Choi, Geoffrey W. Wang, Cosmin Adrian Bejan, Ryan S. Hsi, Alex Bick, Digna R. Velez Edwards, Michael R. Savona, Elizabeth J. Phillips, Jill M. Pulley, Wesley H. Self, Consuelo H. Wilkins, Dan M. Roden, Jordan W. Smoller, Douglas M. Ruderfer, Yaomin Xu |
J. Am. Medical Informatics Assoc. | 7 |
| 2023 | Interactive network-based clustering and investigation of multimorbidity association matrices with associationSubgraphsabstractMOTIVATION: Making sense of networked multivariate association patterns is vitally important to many areas of high-dimensional analysis. Unfortunately, as the data-space dimensions grow, the number of association pairs increases in O(n2); this means that traditional visualizations such as heatmaps quickly become too complicated to parse effectively. RESULTS: Here, we present associationSubgraphs: a new interactive visualization method to quickly and intuitively explore high-dimensional association datasets using network percolation and clustering. The goal is to provide an efficient investigation of association subgraphs, each containing a subset of variables with stronger and more frequent associations among themselves than the remaining variables outside the subset, by showing the entire clustering dynamics and providing subgraphs under all possible cutoff values at once. Particularly, we apply associationSubgraphs to a phenome-wide multimorbidity association matrix generated from an electronic health record and provide an online, interactive demonstration for exploring multimorbidity subgraphs. AVAILABILITY AND IMPLEMENTATION: An R package implementing both the algorithm and visualization components of associationSubgraphs is available at https://github.com/tbilab/associationsubgraphs. Online documentation is available at https://prod.tbilab.org/associationsubgraphs_info/. A demo using a multimorbidity association matrix is available at https://prod.tbilab.org/associationsubgraphs-example/. Nick Strayer, Lydia Yao, Tess Vessels, Cosmin Adrian Bejan, Ryan S. Hsi, Jana Shirey-Rice, Justin M. Balko, Douglas B. Johnson, Elizabeth J. Phillips, Alex Bick, Todd L. Edwards, Digna R. Velez Edwards, Jill M. Pulley, Quinn Stanton Wells, Michael R. Savona, Nancy J. Cox, Dan M. Roden, Douglas M. Ruderfer, Yaomin Xu |
Bioinform. | 5 |
| 2021 | Building longitudinal medication dose data using medication information extracted from clinical notes in electronic health recordsabstractOBJECTIVE: To develop an algorithm for building longitudinal medication dose datasets using information extracted from clinical notes in electronic health records (EHRs). MATERIALS AND METHODS: We developed an algorithm that converts medication information extracted using natural language processing (NLP) into a usable format and builds longitudinal medication dose datasets. We evaluated the algorithm on 2 medications extracted from clinical notes of Vanderbilt's EHR and externally validated the algorithm using clinical notes from the MIMIC-III clinical care database. RESULTS: For the evaluation using Vanderbilt's EHR data, the performance of our algorithm was excellent; F1-measures were ≥0.98 for both dose intake and daily dose. For the external validation using MIMIC-III, the algorithm achieved F1-measures ≥0.85 for dose intake and ≥0.82 for daily dose. DISCUSSION: Our algorithm addresses the challenge of building longitudinal medication dose data using information extracted from clinical notes. Overall performance was excellent, but the algorithm can perform poorly when incorrect information is extracted by NLP systems. Although it performed reasonably well when applied to the external data source, its performance was worse due to differences in the way the drug information was written. The algorithm is implemented in the R package, "EHR," and the extracted data from Vanderbilt's EHRs along with the gold standards are provided so that users can reproduce the results and help improve the algorithm. CONCLUSION: Our algorithm for building longitudinal dose data provides a straightforward way to use EHR data for medication-based studies. The external validation results suggest its potential for applicability to other systems. Elizabeth McNeer, Cole Beck, Hannah L. Weeks, Michael L. Williams, Nathan T. James, Cosmin Adrian Bejan, Leena Choi |
J. Am. Medical Informatics Assoc. | 6 |
| 2021 | Phenotyping coronavirus disease 2019 during a global health pandemic: Lessons learned from the characterization of an early cohort
Sarah DeLozier, Sarah Bland, Melissa McPheeters, Quinn Stanton Wells, Eric Farber-Eger, Cosmin Adrian Bejan, Daniel Fabbri, S. Trent Rosenbloom, Dan M. Roden, Kevin B. Johnson, Wei-Qi Wei, Josh F. Peterson, Lisa Bastarache |
J. Biomed. Informatics | 6 |
| 2020 | medExtractR: A targeted, customizable approach to medication extraction from electronic health recordsabstractOBJECTIVE: We developed medExtractR, a natural language processing system to extract medication information from clinical notes. Using a targeted approach, medExtractR focuses on individual drugs to facilitate creation of medication-specific research datasets from electronic health records. MATERIALS AND METHODS: Written using the R programming language, medExtractR combines lexicon dictionaries and regular expressions to identify relevant medication entities (eg, drug name, strength, frequency). MedExtractR was developed on notes from Vanderbilt University Medical Center, using medications prescribed with varying complexity. We evaluated medExtractR and compared it with 3 existing systems: MedEx, MedXN, and CLAMP (Clinical Language Annotation, Modeling, and Processing). We also demonstrated how medExtractR can be easily tuned for better performance on an outside dataset using the MIMIC-III (Medical Information Mart for Intensive Care III) database. RESULTS: On 50 test notes per development drug and 110 test notes for an additional drug, medExtractR achieved high overall performance (F-measures >0.95), exceeding performance of the 3 existing systems across all drugs. MedExtractR achieved the highest F-measure for each individual entity, except drug name and dose amount for allopurinol. With tuning and customization, medExtractR achieved F-measures >0.90 in the MIMIC-III dataset. DISCUSSION: The medExtractR system successfully extracted entities for medications of interest. High performance in entity-level extraction provides a strong foundation for developing robust research datasets for pharmacological research. When working with new datasets, medExtractR should be tuned on a small sample of notes before being broadly applied. CONCLUSIONS: The medExtractR system achieved high performance extracting specific medications from clinical text, leading to higher-quality research datasets for drug-related studies than some existing general-purpose medication extraction tools. Hannah L. Weeks, Cole Beck, Elizabeth McNeer, Michael L. Williams, Cosmin Adrian Bejan, Joshua C. Denny, Leena Choi |
J. Am. Medical Informatics Assoc. | 5 |
| 2019 | medExtractR: A medication extraction algorithm for electronic health records using the R programming language
Hannah L. Weeks, Cole Beck, Elizabeth McNeer, Cosmin Adrian Bejan, Joshua C. Denny, Leena Choi |
AMIA | 4 |
| 2018 | Mining 100 million notes to find homelessness and adverse childhood experiences: 2 case studies of rare and severe social determinants of health in electronic health recordsabstractObjective: Understanding how to identify the social determinants of health from electronic health records (EHRs) could provide important insights to understand health or disease outcomes. We developed a methodology to capture 2 rare and severe social determinants of health, homelessness and adverse childhood experiences (ACEs), from a large EHR repository. Materials and Methods: We first constructed lexicons to capture homelessness and ACE phenotypic profiles. We employed word2vec and lexical associations to mine homelessness-related words. Next, using relevance feedback, we refined the 2 profiles with iterative searches over 100 million notes from the Vanderbilt EHR. Seven assessors manually reviewed the top-ranked results of 2544 patient visits relevant for homelessness and 1000 patients relevant for ACE. Results: word2vec yielded better performance (area under the precision-recall curve [AUPRC] of 0.94) than lexical associations (AUPRC = 0.83) for extracting homelessness-related words. A comparative study of searches for the 2 phenotypes revealed a higher performance achieved for homelessness (AUPRC = 0.95) than ACE (AUPRC = 0.79). A temporal analysis of the homeless population showed that the majority experienced chronic homelessness. Most ACE patients suffered sexual (70%) and/or physical (50.6%) abuse, with the top-ranked abuser keywords being "father" (21.8%) and "mother" (15.4%). Top prevalent associated conditions for homeless patients were lack of housing (62.8%) and tobacco use disorder (61.5%), while for ACE patients it was mental disorders (36.6%-47.6%). Conclusion: We provide an efficient solution for mining homelessness and ACE information from EHRs, which can facilitate large clinical and genetic studies of these social determinants of health. Cosmin Adrian Bejan, John Angiolillo, Douglas Conway, Robertson Nash, Jana Shirey-Rice, Loren Lipworth-Elliot, Robert M. Cronin, Jill M. Pulley, Sunil Kripalani, Shari Barkin, Kevin B. Johnson, Joshua C. Denny |
J. Am. Medical Informatics Assoc. | 1 |
| 2017 | Large-Scale Text Mining of Social Determinants from Electronic Health Records: Case Studies of Homelessness and Adverse Childhood Experiences
Cosmin Adrian Bejan, John Angiolillo, Douglas Conway, Robertson Nash, Jana Shirey-Rice, Loren Lipworth-Elliot, Robert M. Cronin, Jill M. Pulley, Sunil Kripalani, Shari Barkin, Kevin B. Johnson, Joshua C. Denny |
AMIA | 1 |
| 2017 | A Simple and Efficient Method for the Management of Multiple Electronic Health Record-Driven Phenotype Projects
Cosmin Adrian Bejan, Joshua C. Denny |
AMIA | 1 |
| 2015 | Assessing the role of a medication-indication resource in the treatment relation extraction from clinical textabstractOBJECTIVE: To evaluate the contribution of the MEDication Indication (MEDI) resource and SemRep for identifying treatment relations in clinical text. MATERIALS AND METHODS: We first processed clinical documents with SemRep to extract the Unified Medical Language System (UMLS) concepts and the treatment relations between them. Then, we incorporated MEDI into a simple algorithm that identifies treatment relations between two concepts if they match a medication-indication pair in this resource. For a better coverage, we expanded MEDI using ontology relationships from RxNorm and UMLS Metathesaurus. We also developed two ensemble methods, which combined the predictions of SemRep and the MEDI algorithm. We evaluated our selected methods on two datasets, a Vanderbilt corpus of 6864 discharge summaries and the 2010 Informatics for Integrating Biology and the Bedside (i2b2)/Veteran's Affairs (VA) challenge dataset. RESULTS: The Vanderbilt dataset included 958 manually annotated treatment relations. A double annotation was performed on 25% of relations with high agreement (Cohen's κ = 0.86). The evaluation consisted of comparing the manual annotated relations with the relations identified by SemRep, the MEDI algorithm, and the two ensemble methods. On the first dataset, the best F1-measure results achieved by the MEDI algorithm and the union of the two resources (78.7 and 80, respectively) were significantly higher than the SemRep results (72.3). On the second dataset, the MEDI algorithm achieved better precision and significantly lower recall values than the best system in the i2b2 challenge. The two systems obtained comparable F1-measure values on the subset of i2b2 relations with both arguments in MEDI. CONCLUSIONS: Both SemRep and MEDI can be used to extract treatment relations from clinical text. Knowledge-based extraction with MEDI outperformed use of SemRep alone, but superior performance was achieved by integrating both systems. The integration of knowledge-based resources such as MEDI into information extraction systems such as SemRep and the i2b2 relation extractors may improve treatment relation extraction from clinical text. Cosmin Adrian Bejan, Wei-Qi Wei, Joshua C. Denny |
J. Am. Medical Informatics Assoc. | 1 |
| 2015 | Desiderata for computable representations of electronic health records-driven phenotype algorithmsabstractBACKGROUND: Electronic health records (EHRs) are increasingly used for clinical and translational research through the creation of phenotype algorithms. Currently, phenotype algorithms are most commonly represented as noncomputable descriptive documents and knowledge artifacts that detail the protocols for querying diagnoses, symptoms, procedures, medications, and/or text-driven medical concepts, and are primarily meant for human comprehension. We present desiderata for developing a computable phenotype representation model (PheRM). METHODS: A team of clinicians and informaticians reviewed common features for multisite phenotype algorithms published in PheKB.org and existing phenotype representation platforms. We also evaluated well-known diagnostic criteria and clinical decision-making guidelines to encompass a broader category of algorithms. RESULTS: We propose 10 desired characteristics for a flexible, computable PheRM: (1) structure clinical data into queryable forms; (2) recommend use of a common data model, but also support customization for the variability and availability of EHR data among sites; (3) support both human-readable and computable representations of phenotype algorithms; (4) implement set operations and relational algebra for modeling phenotype algorithms; (5) represent phenotype criteria with structured rules; (6) support defining temporal relations between events; (7) use standardized terminologies and ontologies, and facilitate reuse of value sets; (8) define representations for text searching and natural language processing; (9) provide interfaces for external software algorithms; and (10) maintain backward compatibility. CONCLUSION: A computable PheRM is needed for true phenotype portability and reliability across different EHR products and healthcare systems. These desiderata are a guide to inform the establishment and evolution of EHR phenotype algorithm authoring platforms and languages. Huan Mo, William K. Thompson, Luke V. Rasmussen, Jennifer A. Pacheco, Guoqian Jiang, Richard C. Kiefer, Qian Zhu 0003, Jie Xu 0011, Enid N. H. Montague, David Carrell, Todd Lingren, Frank D. Mentch, Yizhao Ni, Firas H. Wehbe, Peggy L. Peissig, Gerard Tromp, Eric B. Larson, Christopher G. Chute, Jyotishman Pathak, Joshua C. Denny, Peter Speltz, Abel N. Kho, Gail P. Jarvik, Cosmin Adrian Bejan, Marc S. Williams, Kenneth Borthwick, Terrie E. Kitchner, Dan M. Roden, Paul A. Harris |
J. Am. Medical Informatics Assoc. | 24 |
| 2015 | Building bridges across electronic health record systems through inferred phenotypic topics
You Chen 0001, Joydeep Ghosh, Cosmin Adrian Bejan, Carl A. Gunter, Siddharth Gupta 0005, Abel N. Kho, David M. Liebovitz, Jimeng Sun 0001, Joshua C. Denny, Bradley A. Malin |
J. Biomed. Informatics | 3 |
| 2014 | Learning to Identify Treatment Relations in Clinical Text
Cosmin Adrian Bejan, Joshua C. Denny |
AMIA | 1 |
| 2014 | Unsupervised Event Coreference ResolutionabstractThe task of event coreference resolution plays a critical role in many natural language processing applications such as information extraction, question answering, and topic detection and tracking. In this article, we describe a new class of unsupervised, nonparametric Bayesian models with the purpose of probabilistically inferring coreference clusters of event mentions from a collection of unlabeled documents. In order to infer these clusters, we automatically extract various lexical, syntactic, and semantic features for each event mention from the document collection. Extracting a rich set of features for each event mention allows us to cast event coreference resolution as the task of grouping together the mentions that share the same features (they have the same participating entities, share the same location, happen at the same time, etc.). Some of the most important challenges posed by the resolution of event coreference in an unsupervised way stem from (a) the choice of representing event mentions through a rich set of features and (b) the ability of modeling events described both within the same document and across multiple documents. Our first unsupervised model that addresses these challenges is a generalization of the hierarchical Dirichlet process. This new extension presents the hierarchical Dirichlet process's ability to capture the uncertainty regarding the number of clustering components and, additionally, takes into account any finite number of features associated with each event mention. Furthermore, to overcome some of the limitations of this extension, we devised a new hybrid model, which combines an infinite latent class model with a discrete time series model. The main advantage of this hybrid model stands in its capability to automatically infer the number of features associated with each event mention from data and, at the same time, to perform an automatic selection of the most informative features for the task of event coreference. The evaluation performed for solving both within- and cross-document event coreference shows significant improvements of these models when compared against two baselines for this task. Cosmin Adrian Bejan, Sanda M. Harabagiu |
Comput. Linguistics | 1 |
| 2013 | On-time clinical phenotype prediction based on narrative reports
Cosmin Adrian Bejan, Lucy Vanderwende, Heather L. Evans, Mark M. Wurfel, Meliha Yetisgen |
AMIA | 1 |
| 2013 | Assertion modeling and its role in clinical phenotype identification
Cosmin Adrian Bejan, Lucy Vanderwende, Fei Xia 0004, Meliha Yetisgen |
J. Biomed. Informatics | 1 |
| 2012 | Assessing Pneumonia Identification from Time-Ordered Narrative Reports
Cosmin Adrian Bejan, Lucy Vanderwende, Mark M. Wurfel, Meliha Yetisgen |
AMIA | 1 |
| 2012 | Pneumonia identification using statistical feature selectionabstractOBJECTIVE: This paper describes a natural language processing system for the task of pneumonia identification. Based on the information extracted from the narrative reports associated with a patient, the task is to identify whether or not the patient is positive for pneumonia. DESIGN: A binary classifier was employed to identify pneumonia from a dataset of multiple types of clinical notes created for 426 patients during their stay in the intensive care unit. For this purpose, three types of features were considered: (1) word n-grams, (2) Unified Medical Language System (UMLS) concepts, and (3) assertion values associated with pneumonia expressions. System performance was greatly increased by a feature selection approach which uses statistical significance testing to rank features based on their association with the two categories of pneumonia identification. RESULTS: Besides testing our system on the entire cohort of 426 patients (unrestricted dataset), we also used a smaller subset of 236 patients (restricted dataset). The performance of the system was compared with the results of a baseline previously proposed for these two datasets. The best results achieved by the system (85.71 and 81.67 F1-measure) are significantly better than the baseline results (50.70 and 49.10 F1-measure) on the restricted and unrestricted datasets, respectively. CONCLUSION: Using a statistical feature selection approach that allows the feature extractor to consider only the most informative features from the feature space significantly improves the performance over a baseline that uses all the features from the same feature space. Extracting the assertion value for pneumonia expressions further improves the system performance. Cosmin Adrian Bejan, Fei Xia 0004, Lucy Vanderwende, Mark M. Wurfel, Meliha Yetisgen |
J. Am. Medical Informatics Assoc. | 1 |
| 2011 | Commonsense Causal Reasoning Using Millions of Personal StoriesabstractThe personal stories that people write in their Internet weblogs include a substantial amount of information about the causal relationships between everyday events. In this paper we describe our efforts to use millions of these stories for automated commonsense causal reasoning. Casting the commonsense causal reasoning problem as a Choice of Plausible Alternatives, we describe four experiments that compare various statistical and information retrieval approaches to exploit causal information in story corpora. The top performing system in these experiments uses a simple co-occurrence statistic between words in the causal antecedent and consequent, calculated as the Pointwise Mutual Information between words in a corpus of millions of personal stories. Andrew S. Gordon, Cosmin Adrian Bejan, Kenji Sagae |
AAAI | 2 |
| 2010 | Unsupervised Event Coreference Resolution with Rich Linguistic Features
Cosmin Adrian Bejan, Sanda M. Harabagiu |
ACL | 1 |
| 2010 | A Linguistic Resource for Semantic Parsing of Motion Events
Kirk Roberts, Srikanth Gullapalli, Cosmin Adrian Bejan, Sanda M. Harabagiu |
LREC | 3 |
| 2009 | Nonparametric Bayesian Models for Unsupervised Event Coreference ResolutionabstractWe present a sequence of unsupervised, nonparametric Bayesian models for clustering complex linguistic objects. In this approach, we consider a potentially infinite number of features and categorical outcomes. We evaluate these models for the task of within- and cross-document event coreference on two corpora. All the models we investigated show significant improvements when compared against an existing baseline for this task. Cosmin Adrian Bejan, Matthew Titsworth, Andrew Hickl, Sanda M. Harabagiu |
NIPS | 1 |
| 2008 | Using Clustering Methods for Discovering Event Structures
Cosmin Adrian Bejan, Sanda M. Harabagiu |
AAAI | 1 |
| 2008 | A Linguistic Resource for Discovering Event Structures and Resolving Event Coreference
Cosmin Adrian Bejan, Sanda M. Harabagiu |
LREC | 1 |
| 2006 | An Answer Bank for Temporal Inference
Sanda M. Harabagiu, Cosmin Adrian Bejan |
LREC | 2 |
| 2005 | Shallow Semantics for Relation Extraction
Sanda M. Harabagiu, Cosmin Adrian Bejan, Paul Morarescu |
IJCAI | 2 |
| 2004 | A Semantic Kernel for Predicate Argument Classification
Alessandro Moschitti, Cosmin Adrian Bejan |
CoNLL | 2 |