VLDB 2026 Research / reviewers in the wild / expert
Noémie Elhadad
dblp:92/6201 · also Noemie Elhadad
· DBLP profile ↗
101ranked-venue papers
8as first author
29since 2021 · last 2025
0000-0001-9721-5240ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 76 · 6 first-author · 21 since 2021Artificial intelligence and machine learning · 16 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 8 · 4 since 2021Databases, data management, data science and information retrieval · 3Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Promoting Prosociality via Micro-acts of Joy: A Large-Scale Well-Being Intervention StudyabstractProsociality has been well-documented to positively impact mental, social, and physical well-being.However, existing studies of interventions for promoting prosociality have limitations such as Hitesh Goel, Yoobin Park, Jin Liou, Darwin A. Guevarra, Peggy Callahan, Jolene Smith, Bingsheng Yao, Dakuo Wang, Xin Liu 0034, Daniel McDuff, Noémie Elhadad, Emiliana Simon-Thomas, Elissa Epel, Xuhai Xu |
CHI | 11 |
| 2025 | The Voice of Endo: Leveraging Speech for an Intelligent System That Can Forecast Illness Flare-upsabstractManaging complex chronic illness is challenging due to its unpredictability. This paper explores the potential of voice for automated flare-up forecasts. We conducted a six-week speculative design study with individuals with endometriosis, tasking participants to submit daily voice recordings and symptom logs. Through focus groups, we elicited their experiences with voice capture and perceptions of its usefulness in forecasting flare-ups. Participants were enthusiastic and intrigued at the potential of flare-up forecasts through the analysis of their voice. They highlighted imagined benefits from the experience of recording in supporting emotional aspects of illness and validating both day-to-day and overall illness experiences. Participants reported that their recordings revolved around their endometriosis, suggesting that the recordings' content could further inform forecasting. We discuss potential opportunities and challenges in leveraging the voice as a data modality in human-centered AI tools that support individuals with complex chronic conditions. Adrienne Pichon, Jessica R. Blumberg, Lena Mamykina, Noémie Elhadad |
CHI | 4 |
| 2025 | Accurate and Scalable Stochastic Gaussian Process Regression via Learnable Coreset-based Variational InferenceabstractWe introduce a novel stochastic variational inference method for Gaussian process ($\mathcal{GP}$) regression, by deriving a posterior over a learnable set of coresets: i.e., over pseudo-input/output, weighted pairs. Unlike former free-form variational families for stochastic inference, our coreset-based variational $\mathcal{GP}$ (CVGP) is defined in terms of the $\mathcal{GP}$ prior and the (weighted) data likelihood. This formulation naturally incorporates inductive biases of the prior, and ensures its kernel and likelihood dependencies are shared with the posterior. We derive a variational lower-bound on the log-marginal likelihood by marginalizing over the latent $\mathcal{GP}$ coreset variables, and show that CVGP’s lower-bound is amenable to stochastic optimization. CVGP reduces the dimensionality of the variational parameter search space to linear $\mathcal{O}(M)$ complexity, while ensuring numerical stability at $\mathcal{O}(M^3)$ time complexity and $\mathcal{O}(M^2)$ space complexity. Evaluations on real-world and simulated regression problems demonstrate that CVGP achieves superior inference and predictive performance than state-of-the-art, stochastic sparse $\mathcal{GP}$ approximation methods. Mert Ketenci, Adler J. Perotte, Noémie Elhadad, Iñigo Urteaga |
UAI | 3 |
| 2025 | AI as an intervention: improving clinical outcomes relies on a causal approach to AI development and validationabstractThe primary practice of healthcare artificial intelligence (AI) starts with model development, often using state-of-the-art AI, retrospectively evaluated using metrics lifted from the AI literature like AUROC and DICE score. However, good performance on these metrics may not translate to improved clinical outcomes. Instead, we argue for a better development pipeline constructed by working backward from the end goal of positively impacting clinically relevant outcomes using AI, leading to considerations of causality in model development and validation, and subsequently a better development pipeline. Healthcare AI should be "actionable," and the change in actions induced by AI should improve outcomes. Quantifying the effect of changes in actions on outcomes is causal inference. The development, evaluation, and validation of healthcare AI should therefore account for the causal effect of intervening with the AI on clinically relevant outcomes. Using a causal lens, we make recommendations for key stakeholders at various stages of the healthcare AI pipeline. Our recommendations aim to increase the positive impact of AI on clinical outcomes. Shalmali Joshi, Iñigo Urteaga, Wouter A. C. van Amsterdam, George Hripcsak, Pierre A. Elias, Benjamin R. C. Amor, Noémie Elhadad, James C. Fackler, Mark P. Sendak, Jenna Wiens, Kaivalya Deshpande, Yoav Wald, Madalina Fiterau, Zachary C. Lipton, Daniel Malinsky, Madhur Nayan, Hongseok Namkoong, Soojin Park, Julia E. Vogt, Rajesh Ranganath |
J. Am. Medical Informatics Assoc. | 7 |
| 2025 | Negative descriptors in electronic health records of patients with diabetesabstractBACKGROUND: Negative descriptors in electronic health records (EHR) contribute to worse health outcomes; studies show they are also more prevalent in EHRs of women and racial minorities and affect downstream research biases. Similar and unique patterns of negative descriptors may also exist in the records of blind patients, including those with diabetic retinopathy. Diabetic retinopathy is a preventable but leading cause of blindness in the US that is disproportionally high among women and racial and ethnic minorities. METHODS: Using EHR from a large medical center, we created "matched" cohorts of patients with a type 2 diabetes-only diagnosis and patients with a diagnosis of diabetic retinopathy. We identified previously used and new, disability and patient-related negative descriptors and assessed patterns of biased language in the EHR, comparing patients by retinopathy diagnosis (yes/no), and changes in patterns of language usage pre- and post- the retinopathy diagnosis. We also assessed differences between patients with type 2 diabetes at the intersection of blindness (ie, retinopathy diagnosis) and self-reported gender and race and ethnicity marginalization. RESULTS: The EHRs of patients with diabetic retinopathy were significantly more likely than those of patients with diabetes-only diagnoses to contain biased language, across queried negative descriptors. The biasing language was consistently more prevalent in EHRs of patients with diabetic retinopathy identifying as women, Black/African Americans and Hispanic compared to White men and more likely to occur following patients' retinopathy diagnosis. CONCLUSIONS: Our study indicates the presence of both disability- and intersectional biases in EHRs. We discuss findings' implications and suggest steps to address them. Tony Y. Sun, Mika Baugh, Emily R. Gordon, Cameron Ekanayake, Nathalie Moise, Noémie Elhadad, Maya Sabatello |
J. Am. Medical Informatics Assoc. | 6 |
| 2025 | Informing the Design of Individualized Self-Management Regimens from the Human, Data, and Machine Learning PerspectivesabstractIntelligent systems for self-management can help patients and improve quality of life. However, designing AI-based systems is challenging because designers need to account not only for user needs, but also for capabilities and practical constraints of underlying algorithms. We propose and implement a human-centered AI framework to align human and technological requirements and constraints that can guide design of intelligent systems for personal health. We use concepts from a machine learning technique, reinforcement learning, to elicit user needs, through directed content analysis of user interviews, and uncover practical data constraints, through analysis of "in the wild" user engagement logs from a self-monitoring app. We gather and triangulate human-machine-data requirements for a self-management tool for individuals with endometriosis - a poorly understood, complex chronic condition with no reliable treatment. We present recommendations for developing a system that aligns with needs, capabilities, and constraints from human user, data, and machine learning perspectives. Adrienne Pichon, Iñigo Urteaga, Lena Mamykina, Noémie Elhadad |
ACM Trans. Comput. Hum. Interact. | 4 |
| 2024 | Causal fairness assessment of treatment allocation with electronic health recordsabstractOBJECTIVE: Healthcare continues to grapple with the persistent issue of treatment disparities, sparking concerns regarding the equitable allocation of treatments in clinical practice. While various fairness metrics have emerged to assess fairness in decision-making processes, a growing focus has been on causality-based fairness concepts due to their capacity to mitigate confounding effects and reason about bias. However, the application of causal fairness notions in evaluating the fairness of clinical decision-making with electronic health record (EHR) data remains an understudied domain. This study aims to address the methodological gap in assessing causal fairness of treatment allocation with electronic health records data. In addition, we investigate the impact of social determinants of health on the assessment of causal fairness of treatment allocation. METHODS: We propose a causal fairness algorithm to assess fairness in clinical decision-making. Our algorithm accounts for the heterogeneity of patient populations and identifies potential unfairness in treatment allocation by conditioning on patients who have the same likelihood to benefit from the treatment. We apply this framework to a patient cohort with coronary artery disease derived from an EHR database to evaluate the fairness of treatment decisions. RESULTS: Our analysis reveals notable disparities in coronary artery bypass grafting (CABG) allocation among different patient groups. Women were found to be 4.4%-7.7% less likely to receive CABG than men in two out of four treatment response strata. Similarly, Black or African American patients were 5.4%-8.7% less likely to receive CABG than others in three out of four response strata. These results were similar when social determinants of health (insurance and area deprivation index) were dropped from the algorithm. These findings highlight the presence of disparities in treatment allocation among similar patients, suggesting potential unfairness in the clinical decision-making process. CONCLUSION: This study introduces a novel approach for assessing the fairness of treatment allocation in healthcare. By incorporating responses to treatment into fairness framework, our method explores the potential of quantifying fairness from a causal perspective using EHR data. Our research advances the methodological development of fairness assessment in healthcare and highlight the importance of causality in determining treatment fairness. Linying Zhang, Lauren R. Richter, Yixin Wang 0002, Anna Ostropolets, Noémie Elhadad, David M. Blei, George Hripcsak |
J. Biomed. Informatics | 5 |
| 2023 | Generating EDU Extracts for Plan-Guided Summary Re-RankingabstractTwo-step approaches, in which summary candidates are generated-then-reranked to return a single summary, can improve ROUGE scores over the standard single-step approach.Yet, standard decoding methods (i.e., beam search, nucleus sampling, and diverse beam search) produce candidates with redundant, and often low quality, content.In this paper, we design a novel method to generate candidates for re-ranking that addresses these issues.We ground each candidate abstract on its own unique content plan and generate distinct plan-guided abstracts using a model's top beam.More concretely, a standard language model (a BART LM) auto-regressively generates elemental discourse unit (EDU) content plans with an extractive copy mechanism.The top K beams from the content plan generator are then used to guide a separate LM, which produces a single abstractive candidate for each distinct plan.We apply an existing re-ranker (BRIO) to abstractive candidates generated from our method, as well as baseline decoding methods.We show large relevance improvements over previously published methods on widely used single document news article corpora, with ROUGE-2 F1 gains of 0.88, 2.01, and 0.38 on CNN / Dailymail, NYT, and Xsum, respectively.A human evaluation on CNN / DM validates these results.Similarly, on 1k samples from CNN / DM, we show that prompting GPT-3 to follow EDU plans outperforms sampling-based methods by 1.05 ROUGE-2 F1 points.Code to generate and realize plans is available at https: //github.com/griff4692/edu-sum. Griffin Adams, Alexander R. Fabbri, Faisal Ladhak, Noémie Elhadad, Kathy McKeown |
ACL (1) | 4 |
| 2023 | What are the Desired Characteristics of Calibration Sets? Identifying Correlates on Long Form Scientific Summarizationabstractone setup is more effective than another. In this work, we uncover the underlying characteristics of effective sets. For each training instance, we form a large, diverse pool of candidates and systematically vary the subsets used for calibration fine-tuning. Each selection strategy targets distinct aspects of the sets, such as lexical diversity or the size of the gap between positive and negatives. On three diverse scientific long-form summarization datasets (spanning biomedical, clinical, and chemical domains), we find, among others, that faithfulness calibration is optimal when the negative sets are extractive and more likely to be generated, whereas for relevance calibration, the metric margin between candidates should be maximized and surprise-the disagreement between model and metric defined candidate rankings-minimized. Code to create, select, and optimize calibration sets is available at https://github.com/griff4692/calibrating-summaries. Griffin Adams, Bichlien Nguyen, Jake Smith, Yingce Xia, Shufang Xie 0003, Anna Ostropolets, Budhaditya Deb, Yuan-Jyue Chen, Tristan Naumann, Noémie Elhadad |
ACL (1) | 10 |
| 2023 | Representing and utilizing clinical textual data for real world studies: An OHDSI approach
Vipina Kuttichi Keloth, Juan M. Banda, Michael J. Gurley, Paul M. Heider, Georgina Kennedy, Timothy A. Miller, Karthik Natarajan, Olga V. Patterson, Yifan Peng 0002, Kalpana Raja, Ruth M. Reeves, Masoud Rouhizadeh, Jianlin Shi, Yanshan Wang, Wei-Qi Wei, Andrew E. Williams, Rui Zhang 0028, Rimma Belenkaya, Christian G. Reich, Clair Blacketer, Patrick B. Ryan, George Hripcsak, Noémie Elhadad, Hua Xu 0001 |
J. Biomed. Informatics | 26 |
| 2022 | Mining the Health Disparities and Minority Health Bibliome: A Computational Scoping Review
Harry Reyes Nieva, Noémie Elhadad |
AMIA | 2 |
| 2022 | Human, Data, and Algorithmic Perspectives to Inform the Design of RL-enabled Self-management Regimens
Adrienne Pichon, Iñigo Urteaga, Lena Mamykina, Noémie Elhadad |
AMIA | 4 |
| 2022 | Assessing Phenotype Definitions for Algorithmic Fairness
Tony Y. Sun, Shreyas Bhave, Jaan Altosaar, Noémie Elhadad |
AMIA | 4 |
| 2022 | Examining AI Methods for Micro-Coaching DialogsabstractConversational interaction, for example through chatbots, is well-suited to enable automated health coaching tools to support self-management and prevention of chronic diseases. However, chatbots in health are predominantly scripted or rule-based, which can result in a stagnant and repetitive user experience in contrast with more dynamic, data-driven chatbots in other domains. Consequently, little is known about the tradeoffs of pursuing data-driven approaches for health chatbots. We examined multiple artificial intelligence (AI) approaches to enable micro-coaching dialogs in nutrition — brief coaching conversations related to specific meals, to support achievement of nutrition goals — and compared, reinforcement learning (RL), rule-based, and scripted approaches for dialog management. While the data-driven RL chatbot succeeded in shorter, more efficient dialogs, surprisingly the simplest, scripted chatbot was rated as higher quality, despite not fulfilling its task as consistently. These results highlight tensions between scripted and more complex, data-driven approaches for chatbots in health. Elliot G. Mitchell, Noémie Elhadad, Lena Mamykina |
CHI | 2 |
| 2022 | Informatics for sex- and gender-related health: understanding the problems, developing new methods, and designing new solutionsabstractInvestment in medical research continues to grow; however, the sex and gender gap of health outcomes persists1,2 with poor support for the health of cisgender and transgender women,3–5 intersex people,6 and all gender-diverse people (ie, people whose gender identity and sex assigned at birth do not fully align). Informatics approaches have the potential to identify, address, and mitigate these disparities. For example, methods that elucidate the impact of sex as a biological variable, clinical decision support (CDS) systems, and personal health informatics tools are needed to specifically fit the needs of women, intersex people, and all gender-diverse people. The goal of this issue is to highlight such informatics research. Included in this issue are 19 outstanding articles that focus on 4 key areas including gender disparities, gender diversity, maternal health, and sex differences (Table 1). The majority of these articles were focused in the clinical informatics domain with 13 clinical informatics papers. The remaining 6 papers were from the fields of consumer health informatics (N = 4) and translational informatics (N = 2). We also organized the included articles by sex- and gender-related health themes focusing on 4 areas: gender disparities, gender diversity, maternal health, and sex differences. Interestingly, the majority of the consumer health informatics papers (3 out of the 4 included in this issue) were studying gender disparities. No consumer health informatics papers focused on maternal health, and therefore, this could be an area that warrants further investigation. Mary Regina Boland, Noémie Elhadad, Wanda Pratt |
J. Am. Medical Informatics Assoc. | 2 |
| 2022 | An interactive fitness-for-use data completeness tool to assess activity tracker dataabstractOBJECTIVE: To design and evaluate an interactive data quality (DQ) characterization tool focused on fitness-for-use completeness measures to support researchers' assessment of a dataset. MATERIALS AND METHODS: Design requirements were identified through a conceptual framework on DQ, literature review, and interviews. The prototype of the tool was developed based on the requirements gathered and was further refined by domain experts. The Fitness-for-Use Tool was evaluated through a within-subjects controlled experiment comparing it with a baseline tool that provides information on missing data based on intrinsic DQ measures. The tools were evaluated on task performance and perceived usability. RESULTS: The Fitness-for-Use Tool allows users to define data completeness by customizing the measures and its thresholds to fit their research task and provides a data summary based on the customized definition. Using the Fitness-for-Use Tool, study participants were able to accurately complete fitness-for-use assessment in less time than when using the Intrinsic DQ Tool. The study participants perceived that the Fitness-for-Use Tool was more useful in determining the fitness-for-use of a dataset than the Intrinsic DQ Tool. DISCUSSION: Incorporating fitness-for-use measures in a DQ characterization tool could provide data summary that meets researchers needs. The design features identified in this study has potential to be applied to other biomedical data types. CONCLUSION: A tool that summarizes a dataset in terms of fitness-for-use dimensions and measures specific to a research question supports dataset assessment better than a tool that only presents information on intrinsic DQ measures. Sylvia Cho, Ipek Ensari, Noémie Elhadad, Chunhua Weng, Jennifer M. Radin, Brinnae Bent, Pooja M. Desai, Karthik Natarajan |
J. Am. Medical Informatics Assoc. | 3 |
| 2022 | The messiness of the menstruator: assessing personas and functionalities of menstrual tracking appsabstractOBJECTIVE: The aim of this study was to examine trends in the intended users and functionalities advertised by menstrual tracking apps to identify gaps in personas and intended needs fulfilled by these technologies. MATERIALS AND METHODS: Two types of materials were collected: a corpus of scientific articles related to the identities and needs of menstruators and a corpus of images and descriptions of menstrual tracking apps collected from the Google and Apple app stores. We conducted a scoping review of the literature to develop themes and then applied these as a framework to analyze the app corpus, looking for alignments and misalignments between the 2 corpora. RESULTS: A review of the literature showed a wide range of disciplines publishing work relevant to menstruators. We identified 2 broad themes: "who are menstruators?" and "what are the needs of menstruators?" Descriptions of menstrual trackers exhibited misalignments with these themes, with narrow characterizations of menstruators and design for limited needs. DISCUSSION: We synthesize gaps in the design of menstrual tracking apps and discuss implications for designing around: (1) an irregular menstrual cycle as the norm; (2) the embodied, leaky experience of menstruation; and (3) the varied biologies, identities, and goals of menstruators. An overarching gap suggests a need for a human-centered artificial intelligence approach for model and data provenance, transparency and explanations of uncertainties, and the prioritization of privacy in menstrual trackers. CONCLUSION: Comparing and contrasting literature about menstruators and descriptions of menstrual tracking apps provide a valuable guide to assess menstrual technology and their responsiveness to users and their needs. Adrienne Pichon, Kasey B. Jackman, Inga T. Winkler, Chris Bobel, Noémie Elhadad |
J. Am. Medical Informatics Assoc. | 5 |
| 2022 | Erratum To: The Messiness of The Menstruator: Assessing Personas and Functionalities of Menstrual Tracking AppsabstractJournal of the American Medical Informatics Association, ocab212, https://doi.org/10.1093/jamia/ocab212 In the originally published version of this manuscript, Figure 4 was mistakenly published without author approval and a funding statement was omitted. The publisher apologizes for the error. Figure 4 has been removed, and sub-section ‘Theme 1-1’ in section ‘Findings of the analysis of menstrual trackers’ amended to add in-text description to account for figure removal. The following statement has also been added to the Funding section: “During the time of this work KJ was supported by NINR T32NR007969 (PI: Bakken).” Adrienne Pichon, Kasey B. Jackman, Inga T. Winkler, Chris Bobel, Noémie Elhadad |
J. Am. Medical Informatics Assoc. | 5 |
| 2021 | Exploring Automatic Summarization of the Hospital Course
Griffin Adams, Emily Alsentzer, Mert Ketenci, Jason Zucker 0001, Noémie Elhadad |
AMIA | 5 |
| 2021 | Informing Symptom Science Using a Citizen Science Application in the COVID-19 Pandemic
Caitlin N. Dreisbach, Katherine South, Theresa A. Koleck, Veronica Barcelona, Lena Mamykina, Noémie Elhadad, Suzanne Bakken |
AMIA | 6 |
| 2021 | Differential Presentation and Delays in Treatment for Acute Myocardial Infarction Associated with Sex and Race/Ethnicity
Harry Reyes Nieva, Tony Sun, Sharon Lipsky Gorman, Grace Mao, Noémie Elhadad |
AMIA | 5 |
| 2021 | Gender Differences in Time to Diagnosis through Fairness and Time Variant Evaluation of EHR Data
Tony Y. Sun, Oliver J. Bear Don't Walk IV, Jenny Chen, Jaan Altosaar, Harry Reyes Nieva, Noémie Elhadad |
AMIA | 6 |
| 2021 | What's in a Summary? Laying the Groundwork for Advances in Hospital-Course SummarizationabstractSummarization of clinical narratives is a long-standing research problem. Here, we introduce the task of hospital-course summarization. Given the documentation authored throughout a patient's hospitalization, generate a paragraph that tells the story of the patient admission. We construct an English, text-to-text dataset of 109,000 hospitalizations (2M source notes) and their corresponding summary proxy: the clinician-authored "Brief Hospital Course" paragraph written as part of a discharge note. Exploratory analyses reveal that the BHC paragraphs are highly abstractive with some long extracted fragments; are concise yet comprehensive; differ in style and content organization from the source notes; exhibit minimal lexical cohesion; and represent silver-standard references. Our analysis identifies multiple implications for modeling this complex, multi-document summarization task. Griffin Adams, Emily Alsentzer, Mert Ketenci, Jason Zucker 0001, Noémie Elhadad |
NAACL-HLT | 5 |
| 2021 | A predictive model for next cycle start date that accounts for adherence in menstrual self-trackingabstractOBJECTIVE: The study sought to build predictive models of next menstrual cycle start date based on mobile health self-tracked cycle data. Because app users may skip tracking, disentangling physiological patterns of menstruation from tracking behaviors is necessary for the development of predictive models. MATERIALS AND METHODS: We use data from a popular menstrual tracker (186 000 menstruators with over 2 million tracked cycles) to learn a predictive model, which (1) accounts explicitly for self-tracking adherence; (2) updates predictions as a given cycle evolves, allowing for interpretable insight into how these predictions change over time; and (3) enables modeling of an individual's cycle length history while incorporating population-level information. RESULTS: Compared with 5 baselines (mean, median, convolutional neural network, recurrent neural network, and long short-term memory network), the model yields better predictions and consistently outperforms them as the cycle evolves. The model also provides predictions of skipped tracking probabilities. DISCUSSION: Mobile health apps such as menstrual trackers provide a rich source of self-tracked observations, but these data have questionable reliability, as they hinge on user adherence to the app. By taking a machine learning approach to modeling self-tracked cycle lengths, we can separate true cycle behavior from user adherence, allowing for more informed predictions and insights into the underlying observed data structure. CONCLUSIONS: Disentangling physiological patterns of menstruation from adherence allows for accurate and informative predictions of menstrual cycle start date and is necessary for mobile tracking apps. The proposed predictive model can support app users in being more aware of their self-tracking behavior and in better understanding their cycle dynamics. Kathy Li, Iñigo Urteaga, Amanda Shea, Virginia J. Vitzthum, Chris Wiggins 0001, Noémie Elhadad |
J. Am. Medical Informatics Assoc. | 6 |
| 2021 | Development and validation of prediction models for mechanical ventilation, renal replacement therapy, and readmission in COVID-19 patientsabstractOBJECTIVE: Coronavirus disease 2019 (COVID-19) patients are at risk for resource-intensive outcomes including mechanical ventilation (MV), renal replacement therapy (RRT), and readmission. Accurate outcome prognostication could facilitate hospital resource allocation. We develop and validate predictive models for each outcome using retrospective electronic health record data for COVID-19 patients treated between March 2 and May 6, 2020. MATERIALS AND METHODS: For each outcome, we trained 3 classes of prediction models using clinical data for a cohort of SARS-CoV-2 (severe acute respiratory syndrome coronavirus 2)-positive patients (n = 2256). Cross-validation was used to select the best-performing models per the areas under the receiver-operating characteristic and precision-recall curves. Models were validated using a held-out cohort (n = 855). We measured each model's calibration and evaluated feature importances to interpret model output. RESULTS: The predictive performance for our selected models on the held-out cohort was as follows: area under the receiver-operating characteristic curve-MV 0.743 (95% CI, 0.682-0.812), RRT 0.847 (95% CI, 0.772-0.936), readmission 0.871 (95% CI, 0.830-0.917); area under the precision-recall curve-MV 0.137 (95% CI, 0.047-0.175), RRT 0.325 (95% CI, 0.117-0.497), readmission 0.504 (95% CI, 0.388-0.604). Predictions were well calibrated, and the most important features within each model were consistent with clinical intuition. DISCUSSION: Our models produce performant, well-calibrated, and interpretable predictions for COVID-19 patients at risk for the target outcomes. They demonstrate the potential to accurately estimate outcome prognosis in resource-constrained care sites managing COVID-19 patients. CONCLUSIONS: We develop and validate prognostic models targeting MV, RRT, and readmission for hospitalized COVID-19 patients which produce accurate, interpretable predictions. Additional external validation studies are needed to further verify the generalizability of our results. Victor Alfonso Rodriguez, Shreyas Bhave, George Hripcsak, Soumitra Sengupta, Noémie Elhadad, Robert A. Green, Jason S. Adelman, Katherine Schlosser Metitiri, Pierre A. Elias, Holden Groves, Sumit Mohan, Karthik Natarajan, Adler J. Perotte |
J. Am. Medical Informatics Assoc. | 7 |
| 2021 | Clinician involvement in research on machine learning-based predictive clinical decision support for the hospital setting: A scoping reviewabstractOBJECTIVE: The study sought to describe the prevalence and nature of clinical expert involvement in the development, evaluation, and implementation of clinical decision support systems (CDSSs) that utilize machine learning to analyze electronic health record data to assist nurses and physicians in prognostic and treatment decision making (ie, predictive CDSSs) in the hospital. MATERIALS AND METHODS: A systematic search of PubMed, CINAHL, and IEEE Xplore and hand-searching of relevant conference proceedings were conducted to identify eligible articles. Empirical studies of predictive CDSSs using electronic health record data for nurses or physicians in the hospital setting published in the last 5 years in peer-reviewed journals or conference proceedings were eligible for synthesis. Data from eligible studies regarding clinician involvement, stage in system design, predictive CDSS intention, and target clinician were charted and summarized. RESULTS: Eighty studies met eligibility criteria. Clinical expert involvement was most prevalent at the beginning and late stages of system design. Most articles (95%) described developing and evaluating machine learning models, 28% of which described involving clinical experts, with nearly half functioning to verify the clinical correctness or relevance of the model (47%). DISCUSSION: Involvement of clinical experts in predictive CDSS design should be explicitly reported in publications and evaluated for the potential to overcome predictive CDSS adoption challenges. CONCLUSIONS: If present, clinical expert involvement is most prevalent when predictive CDSS specifications are made or when system implementations are evaluated. However, clinical experts are less prevalent in developmental stages to verify clinical correctness, select model features, preprocess data, or serve as a gold standard. Jessica Schwartz-Dillard, Amanda J. Moy, Sarah Collins Rossetti, Noémie Elhadad, Kenrick Cato |
J. Am. Medical Informatics Assoc. | 4 |
| 2021 | Response to: Looking for clinician involvement under the wrong lamp post: the need for collaboration measuresabstractDear JAMIA Editors and Readers: We appreciate the critiques that Dr. Sendak and colleagues have brought forward regarding our scoping review of clinician involvement in predictive CDSS design.1 In their letter, Sendak and colleagues argue that our review too narrowly defined clinician involvement and that relationships established between clinician leaders, often coauthors on manuscripts, and other research team members is a valuable form of clinician involvement not adequately captured in our review.2 We recognize and agree that we should have more prominently highlighted the possibility that clinically affiliated coauthors’ contributions may have represented clinician involvement in one of the ways we charted or in a different relationship-oriented way that is also important for predictive CDSS success. We also should have consistently referred to our results finding that involvement is not widely reported instead of not widely practiced. We also acknowledge that reaching out to authors to gather... Jessica Schwartz-Dillard, Amanda J. Moy, Sarah Collins Rossetti, Noémie Elhadad, Kenrick Cato |
J. Am. Medical Informatics Assoc. | 4 |
| 2021 | Corrigendum to: Clinician involvement in research on machine learning-based predictive clinical decision support for the hospital setting: A scoping reviewabstractbeen corrected from " [25][26][27][28]30 Jessica Schwartz-Dillard, Amanda J. Moy, Sarah Collins Rossetti, Noémie Elhadad, Kenrick Cato |
J. Am. Medical Informatics Assoc. | 4 |
| 2021 | Clinically relevant pretraining is all you needabstractClinical notes present a wealth of information for applications in the clinical domain, but heterogeneity across clinical institutions and settings presents challenges for their processing. The clinical natural language processing field has made strides in overcoming domain heterogeneity, while pretrained deep learning models present opportunities to transfer knowledge from one task to another. Pretrained models have performed well when transferred to new tasks; however, it is not well understood if these models generalize across differences in institutions and settings within the clinical domain. We explore if institution or setting specific pretraining is necessary for pretrained models to perform well when transferred to new tasks. We find no significant performance difference between models pretrained across institutions and settings, indicating that clinically pretrained models transfer well across such boundaries. Given a clinically pretrained model, clinical natural language processing researchers may forgo the time-consuming pretraining step without a significant performance drop. Oliver J. Bear Don't Walk IV, Tony Y. Sun, Adler J. Perotte, Noémie Elhadad |
J. Am. Medical Informatics Assoc. | 4 |
| 2020 | Patient Record Summarization Through Joint Phenotype Learning and Interactive Visualization
Gal Levy-Fix, Jason Zucker 0001, Konstantin Stojanovic, Noémie Elhadad |
AMIA | 4 |
| 2020 | Cross-lingual Unified Medical Language System entity linking in online health communitiesabstractOBJECTIVE: In Hebrew online health communities, participants commonly write medical terms that appear as transliterated forms of a source term in English. Such transliterations introduce high variability in text and challenge text-analytics methods. To reduce their variability, medical terms must be normalized, such as linking them to Unified Medical Language System (UMLS) concepts. We present a method to identify both transliterated and translated Hebrew medical terms and link them with UMLS entities. MATERIALS AND METHODS: We investigate the effect of linking terms in Camoni, a popular Israeli online health community in Hebrew. Our method, MDTEL (Medical Deep Transliteration Entity Linking), includes (1) an attention-based recurrent neural network encoder-decoder to transliterate words and mapping UMLS from English to Hebrew, (2) an unsupervised method for creating a transliteration dataset in any language without manually labeled data, and (3) an efficient way to identify and link medical entities in the Hebrew corpus to UMLS concepts, by producing a high-recall list of candidate medical terms in the corpus, and then filtering the candidates to relevant medical terms. RESULTS: We carry out experiments on 3 disease-specific communities: diabetes, multiple sclerosis, and depression. MDTEL tagging and normalizing on Camoni posts achieved 99% accuracy, 92% recall, and 87% precision. When tagging and normalizing terms in queries from the Camoni search logs, UMLS-normalized queries improved search results in 46% of the cases. CONCLUSIONS: Cross-lingual UMLS entity linking from Hebrew is possible and improves search performance across communities. Annotated datasets, annotation guidelines, and code are made available online (https://github.com/yonatanbitton/mdtel). Yonatan Bitton, Raphael Cohen, Tamar Schifter, Eitan Bachmat, Michael Elhadad, Noémie Elhadad |
J. Am. Medical Informatics Assoc. | 6 |
| 2020 | Divided We Stand: The Collaborative Work of Patients and Providers in an Enigmatic Chronic DiseaseabstractIn chronic conditions, patients and providers need support in understanding and managing illness over time. Focusing on endometriosis, an enigmatic chronic condition, we conducted interviews with specialists and focus groups with patients to elicit their work in care specifically pertaining to dealing with an enigmatic disease, both independently and in partnership, and how technology could support these efforts. We found that the work to care for the illness, including reflecting on the illness experience and planning for care, is significantly compounded by the complex nature of the disease: enigmatic condition means uncertainty and frustration in care and management; the multi-factorial and systemic features of endometriosis without any guidance to interpret them overwhelm patients and providers; the different temporal resolutions of this chronic condition confuse both patients and provides; and patients and providers negotiate medical knowledge and expertise in an attempt to align their perspectives. We note how this added complexity demands that patients and providers work together to find common ground and align perspectives, and propose three design opportunities (considerations to construct a holistic picture of the patient, design features to reflect and make sense of the illness, and opportunities and mechanisms to correct misalignments and plan for care) and implications to support patients and providers in their care work. Specifically, the enigmatic nature of endometriosis necessitates complementary approaches from human-centered computing and artificial intelligence, and thus opens a number of future research avenues. Adrienne Pichon, Kayla Schiffer, Emma Horan, Bria Massey, Suzanne Bakken, Lena Mamykina, Noémie Elhadad |
Proc. ACM Hum. Comput. Interact. | 7 |
| 2019 | Learning Across a Healthcare Data Network to Improve Model Robustness and Evidence Reliability
Noémie Elhadad, Iñigo Urteaga, Alison Callahan, Jenna Reps, Patrick B. Ryan |
AMIA | 1 |
| 2019 | Longitudinal analysis of social and behavioral determinants of health in the EHR: exploring the impact of patient trajectories and documentation practices
Daniel J. Feller, Jason Zucker 0001, Oliver J. Bear Don't Walk IV, Michael T. Yin, Peter Gordon, Noémie Elhadad |
AMIA | 6 |
| 2019 | Get on the Same Page: Negotiating and Aligning Knowledge and Expectations between Patients and Providers through Self-Tracking Artifacts
Adrienne Pichon, Megan Houterloot, Kayla Schiffer, Emma Horan, Noémie Elhadad |
AMIA | 5 |
| 2018 | Towards the Inference of Social and Behavioral Determinants of Sexual Health: Development of a Gold-Standard Corpus with Semi-Supervised Learning
Daniel J. Feller, Jason Zucker 0001, Oliver J. Bear Don't Walk IV, Bharat Srikishan, Roxana Martinez, Henry Evans, Michael T. Yin, Peter Gordon, Noémie Elhadad |
AMIA | 9 |
| 2018 | Identifying Clinical Notes with Likely Documentation of Social and Behavioral Determinants of Health
Oliver J. Bear Don't Walk IV, Jason Zucker 0001, Peter Gordon, Noémie Elhadad, Daniel J. Feller, Bharat Srikishan, Michael T. Yin |
AMIA | 4 |
| 2018 | Computable Longitudinal Patient Trajectories
Jeremy L. Warner, Guergana K. Savova, Noémie Elhadad, Lisa Bastarache, David Gotz |
AMIA | 3 |
| 2018 | Designing in the Dark: Eliciting Self-tracking Dimensions for Understanding Enigmatic DiseaseabstractThe design of personal health informatics tools has traditionally been explored in self-monitoring and behavior change. There is an unmet opportunity to leverage self- tracking of individuals and study diseases and health conditions to learn patterns across groups. An open research question, however, is how to design engaging self-tracking tools that also facilitate learning at scale. Furthermore, for conditions that are not well understood, a critical question is how to design such tools when it is unclear which data types are relevant to the disease. We outline the process of identifying design requirements for self-tracking endometriosis, a highly enigmatic and prevalent disease, through interviews (N=3), focus groups (N=27), surveys (N=741), and content analysis of an online endometriosis community (1500 posts, N=153 posters) and show value in triangulating across these methods. Finally, we discuss tensions inherent in designing self-tracking tools for individual use and population analysis, making suggestions for overcoming these tensions. Mollie McKillop, Lena Mamykina, Noémie Elhadad |
CHI | 3 |
| 2018 | Estimating summary statistics for electronic health record laboratory data for use in high-throughput phenotyping algorithmsabstractWe study the question of how to represent or summarize raw laboratory data taken from an electronic health record (EHR) using parametric model selection to reduce or cope with biases induced through clinical care. It has been previously demonstrated that the health care process (Hripcsak and Albers, 2012, 2013), as defined by measurement context (Hripcsak and Albers, 2013; Albers et al., 2012) and measurement patterns (Albers and Hripcsak, 2010, 2012), can influence how EHR data are distributed statistically (Kohane and Weber, 2013; Pivovarov et al., 2014). We construct an algorithm, PopKLD, which is based on information criterion model selection (Burnham and Anderson, 2002; Claeskens and Hjort, 2008), is intended to reduce and cope with health care process biases and to produce an intuitively understandable continuous summary. The PopKLD algorithm can be automated and is designed to be applicable in high-throughput settings; for example, the output of the PopKLD algorithm can be used as input for phenotyping algorithms. Moreover, we develop the PopKLD-CAT algorithm that transforms the continuous PopKLD summary into a categorical summary useful for applications that require categorical data such as topic modeling. We evaluate our methodology in two ways. First, we apply the method to laboratory data collected in two different health care contexts, primary versus intensive care. We show that the PopKLD preserves known physiologic features in the data that are lost when summarizing the data using more common laboratory data summaries such as mean and standard deviation. Second, for three disease-laboratory measurement pairs, we perform a phenotyping task: we use the PopKLD and PopKLD-CAT algorithms to define high and low values of the laboratory variable that are used for defining a disease state. We then compare the relationship between the PopKLD-CAT summary disease predictions and the same predictions using empirically estimated mean and standard deviation to a gold standard generated by clinical review of patient records. We find that the PopKLD laboratory data summary is substantially better at predicting disease state. The PopKLD or PopKLD-CAT algorithms are not meant to be used as phenotyping algorithms, but we use the phenotyping task to show what information can be gained when using a more informative laboratory data summary. In the process of evaluation our method we show that the different clinical contexts and laboratory measurements necessitate different statistical summaries. Similarly, leveraging the principle of maximum entropy we argue that while some laboratory data only have sufficient information to estimate a mean and standard deviation, other laboratory data captured in an EHR contain substantially more information than can be captured in higher-parameter models. David J. Albers, Noémie Elhadad, Jan Claassen, Rimma Perotte, Andrew Goldstein, George Hripcsak |
J. Biomed. Informatics | 2 |
| 2018 | When to re-order laboratory tests? Learning laboratory test shelf-life
Gal Levy-Fix, Sharon Lipsky Gorman, Jorge L. Sepulveda, Noémie Elhadad |
J. Biomed. Informatics | 4 |
| 2017 | When to re-order laboratory tests? Learning lab shelf-life
Gal Levy-Fix, Sharon Lipsky Gorman, Noémie Elhadad |
AMIA | 3 |
| 2017 | Beyond ResearchKit: User Engagement with Phendo, a Novel App for Self-tracking and Research
Mollie McKillop, Sylvia English, Sharib A. Khan, Chintan Patel, Noémie Elhadad |
AMIA | 5 |
| 2017 | Clinical Natural Language Processing in Languages Other Than English
Aurélie Névéol, Noémie Elhadad, Sumithra Velupillai, Hua Xu 0001, Guergana K. Savova |
AMIA | 2 |
| 2017 | Cataloguing Treatments Discussed and Used in Online Autism CommunitiesabstractA large number of patients discuss treatments in online health communities (OHCs). One research question of interest to health researchers is whether treatments being discussed in OHCs are eventually used by community members in their real lives. In this paper, we rely on machine learning methods to automatically identify attributions of mentions of treatments from an online autism community. The context of our work is online autism communities, where parents exchange support for the care of their children with autism spectrum disorder. Our methods are able to distinguish discussions of treatments that are associated with patients, caregivers, and others, as well as identify whether a treatment is actually taken. We investigate treatments that are not just discussed but also used by patients according to two types of content analysis, cross-sectional and longitudinal. The treatments identified through our content analysis help create a catalogue of real-world treatments. This study results lay the foundation for future research to compare real-world drug usage with established clinical guidelines. Shaodian Zhang, Tian Kang, Weinan Zhang 0001, Yong Yu 0001, Noémie Elhadad |
WWW | 6 |
| 2017 | EliIE: An open-source information extraction system for clinical trial eligibility criteriaabstractOBJECTIVE: To develop an open-source information extraction system called Eligibility Criteria Information Extraction (EliIE) for parsing and formalizing free-text clinical research eligibility criteria (EC) following Observational Medical Outcomes Partnership Common Data Model (OMOP CDM) version 5.0. MATERIALS AND METHODS: EliIE parses EC in 4 steps: (1) clinical entity and attribute recognition, (2) negation detection, (3) relation extraction, and (4) concept normalization and output structuring. Informaticians and domain experts were recruited to design an annotation guideline and generate a training corpus of annotated EC for 230 Alzheimer's clinical trials, which were represented as queries against the OMOP CDM and included 8008 entities, 3550 attributes, and 3529 relations. A sequence labeling-based method was developed for automatic entity and attribute recognition. Negation detection was supported by NegEx and a set of predefined rules. Relation extraction was achieved by a support vector machine classifier. We further performed terminology-based concept normalization and output structuring. RESULTS: In task-specific evaluations, the best F1 score for entity recognition was 0.79, and for relation extraction was 0.89. The accuracy of negation detection was 0.94. The overall accuracy for query formalization was 0.71 in an end-to-end evaluation. CONCLUSIONS: This study presents EliIE, an OMOP CDM-based information extraction system for automatic structuring and formalization of free-text EC. According to our evaluation, machine learning-based EliIE outperforms existing systems and shows promise to improve. Tian Kang, Shaodian Zhang, Youlan Tang, Gregory William Hruby, Alex Rusanov, Noémie Elhadad, Chunhua Weng |
J. Am. Medical Informatics Assoc. | 6 |
| 2017 | Online cancer communities as informatics intervention for social support: conceptualization, characterization, and impactabstractObjectives: The Internet and social media are revolutionizing how social support is exchanged and perceived, making online health communities (OHCs) one of the most exciting research areas in health informatics. This paper aims to provide a framework for organizing research of OHCs and help identify questions to explore for future informatics research. Based on the framework, we conceptualize OHCs from a social support standpoint and identify variables of interest in characterizing community members. For the sake of this tutorial, we focus our review on online cancer communities. Target audience: The primary target audience is informaticists interested in understanding ways to characterize OHCs, their members, and the impact of participation, and in creating tools to facilitate outcome research of OHCs. OHC designers and moderators are also among the target audience for this tutorial. Scope: The tutorial provides an informatics point of view of online cancer communities, with social support as their leading element. We conceptualize OHCs according to 3 major variables: type of support, source of support, and setting in which the support is exchanged. We summarize current research and synthesize the findings for 2 primary research questions on online cancer communities: (1) the impact of using online social support on an individual's health, and (2) the characteristics of the community, its members, and their interactions. We discuss ways in which future research in informatics in social support and OHCs can ultimately benefit patients. Shaodian Zhang, Erin O'Carroll Bantum, Jason E. Owen, Suzanne Bakken, Noémie Elhadad |
J. Am. Medical Informatics Assoc. | 5 |
| 2017 | Longitudinal analysis of discussion topics in an online breast cancer community using convolutional neural networksabstractIdentifying topics of discussions in online health communities (OHC) is critical to various information extraction applications, but can be difficult because topics of OHC content are usually heterogeneous and domain-dependent. In this paper, we provide a multi-class schema, an annotated dataset, and supervised classifiers based on convolutional neural network (CNN) and other models for the task of classifying discussion topics. We apply the CNN classifier to the most popular breast cancer online community, and carry out cross-sectional and longitudinal analyses to show topic distributions and topic dynamics throughout members' participation. Our experimental results suggest that CNN outperforms other classifiers in the task of topic classification and identify several patterns and trajectories. For example, although members discuss mainly disease-related topics, their interest may change through time and vary with their disease severities. Shaodian Zhang, Edouard Grave, Elizabeth Sklar, Noémie Elhadad |
J. Biomed. Informatics | 4 |
| 2016 | Patient Generated Data: the Missing Link in Patient Centered Care?
Noémie Elhadad, Lena Mamykina, Eileen Koski |
AMIA | 1 |
| 2016 | Qualitative Assessment of Women's Attitudes Towards a Tracking App for Phenotyping Endometriosis
Mollie McKillop, Tamer Seckin, Noémie Elhadad |
AMIA | 3 |
| 2016 | A probabilistic model for learning relationships between diagnosis codes and clinical free text
Adler J. Perotte, Noémie Elhadad |
AMIA | 2 |
| 2016 | Can Patient Record Summarization Support Quality Metric Abstraction?
Rimma Perotte, Yael J. Coppleson, Sharon Lipsky Gorman, David K. Vawdrey, Noémie Elhadad |
AMIA | 5 |
| 2016 | Factors Contributing to Dropping-out in an Online Health Community: Static and Longitudinal Analyses
Shaodian Zhang, Noémie Elhadad |
AMIA | 2 |
| 2016 | Data-driven health management: reasoning about personally generated data in diabetes with information technologiesabstractOBJECTIVE: To investigate how individuals with diabetes and diabetes educators reason about data collected through self-monitoring and to draw implications for the design of data-driven self-management technologies. MATERIALS AND METHODS: Ten individuals with diabetes (six type 1 and four type 2) and 2 experienced diabetes educators were presented with a set of self-monitoring data captured by an individual with type 2 diabetes. The set included digital images of meals and their textual descriptions, and blood glucose (BG) readings captured before and after these meals. The participants were asked to review a set of meals and associated BG readings, explain differences in postprandial BG levels for these meals, and predict postprandial BG levels for the same individual for a different set of meals. Researchers compared conclusions and predictions reached by the participants with those arrived at by quantitative analysis of the collected data. RESULTS: The participants used both macronutrient composition of meals, most notably the inclusion of carbohydrates, and names of dishes and ingredients to reason about changes in postprandial BG levels. Both individuals with diabetes and diabetes educators reported difficulties in generating predictions of postprandial BG; their predictions varied in their correlations with the actual captured readings from r = 0.008 to r = 0.75. CONCLUSION: Overall, the study showed that identifying trends in the data collected with self-monitoring is a complex process, and that conclusions reached by both individuals with diabetes and diabetes educators are not always reliable. This suggests the need for new ways to facilitate individuals' reasoning with informatics interventions. Lena Mamykina, Matthew E. Levine, Patricia G. Davidson, Arlene M. Smaldone, Noémie Elhadad, David J. Albers |
J. Am. Medical Informatics Assoc. | 5 |
| 2016 | Speculation detection for Chinese clinical notes: Impacts of word segmentation and embedding modelsabstractSpeculations represent uncertainty toward certain facts. In clinical texts, identifying speculations is a critical step of natural language processing (NLP). While it is a nontrivial task in many languages, detecting speculations in Chinese clinical notes can be particularly challenging because word segmentation may be necessary as an upstream operation. The objective of this paper is to construct a state-of-the-art speculation detection system for Chinese clinical notes and to investigate whether embedding features and word segmentations are worth exploiting toward this overall task. We propose a sequence labeling based system for speculation detection, which relies on features from bag of characters, bag of words, character embedding, and word embedding. We experiment on a novel dataset of 36,828 clinical notes with 5103 gold-standard speculation annotations on 2000 notes, and compare the systems in which word embeddings are calculated based on word segmentations given by general and by domain specific segmenters respectively. Our systems are able to reach performance as high as 92.2% measured by F score. We demonstrate that word segmentation is critical to produce high quality word embedding to facilitate downstream information extraction applications, and suggest that a domain dependent word segmenter can be vital to such a clinical NLP task in Chinese language. Shaodian Zhang, Tian Kang, Xingting Zhang, Dong Wen 0003, Noémie Elhadad, Jianbo Lei |
J. Biomed. Informatics | 5 |
| 2015 | A convex and feature-rich discriminative approach to dependency grammar inductionabstractÉdouard Grave, Noémie Elhadad. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Edouard Grave, Noémie Elhadad |
ACL (1) | 2 |
| 2015 | Model Selection For EHR Laboratory Tests Preserving Healthcare Context and Underlying Physiology
David J. Albers, Rimma Perotte, J. Michael Schmidt, Noémie Elhadad, George Hripcsak |
AMIA | 4 |
| 2015 | Initial Readability Assessment of Clinical Trial Eligibility Criteria
Tian Kang, Noémie Elhadad, Chunhua Weng |
AMIA | 2 |
| 2015 | Collective Sensemaking in Online Health ForumsabstractOnline health communities collect vast amounts of information and opinions in regards to health and wellness management. However, these opinions are usually stored within lengthy and loosely structured discussion threads; synthesizing information in these threads can be challenging. In this mixed-methods study, grounded in the theoretical perspective of collective sensemaking, we examined patterns of communication within an online diabetes community TuDiabetes. The results of the study suggest that members of TuDiabetes often construct shared meaning through deep discussions, back and forth negotiation of perspectives, and resolution of conflicts in opinions. However, unlike participants of other sensemaking communities, members of TuDiabetes often value multiplicity of opinions rather than consensus. We use study results to draw implications for the design of computing platforms for facilitating collective sensemaking that promote construction of shared knowledge yet embrace diversity of opinions. Lena Mamykina, Drashko Nakikj, Noémie Elhadad |
CHI | 3 |
| 2015 | Intelligible Models for HealthCare: Predicting Pneumonia Risk and Hospital 30-day ReadmissionabstractIn machine learning often a tradeoff must be made between accuracy and intelligibility. More accurate models such as boosted trees, random forests, and neural nets usually are not intelligible, but more intelligible models such as logistic regression, naive-Bayes, and single decision trees often have significantly worse accuracy. This tradeoff sometimes limits the accuracy of models that can be applied in mission-critical applications such as healthcare where being able to understand, validate, edit, and trust a learned model is important. We present two case studies where high-performance generalized additive models with pairwise interactions (GA2Ms) are applied to real healthcare problems yielding intelligible models with state-of-the-art accuracy. In the pneumonia risk prediction case study, the intelligible model uncovers surprising patterns in the data that previously had prevented complex learned models from being fielded in this domain, but because it is intelligible and modular allows these patterns to be recognized and removed. In the 30-day hospital readmission case study, we show that the same methods scale to large datasets containing hundreds of thousands of patients and thousands of attributes while remaining intelligible and providing accuracy comparable to the best (unintelligible) machine learning methods. Rich Caruana, Yin Lou, Johannes Gehrke, Paul Koch, Marc Sturm, Noémie Elhadad |
KDD | 6 |
| 2015 | The Survival Filter: Joint Survival Analysis with a Latent Time Series
Rajesh Ranganath, Adler J. Perotte, Noémie Elhadad, David M. Blei |
UAI | 3 |
| 2015 | HARVEST, a longitudinal patient record summarizerabstractOBJECTIVE: To describe HARVEST, a novel point-of-care patient summarization and visualization tool, and to conduct a formative evaluation study to assess its effectiveness and gather feedback for iterative improvements. MATERIALS AND METHODS: HARVEST is a problem-based, interactive, temporal visualization of longitudinal patient records. Using scalable, distributed natural language processing and problem salience computation, the system extracts content from the patient notes and aggregates and presents information from multiple care settings. Clinical usability was assessed with physician participants using a timed, task-based chart review and questionnaire, with performance differences recorded between conditions (standard data review system and HARVEST). RESULTS: HARVEST displays patient information longitudinally using a timeline, a problem cloud as extracted from notes, and focused access to clinical documentation. Despite lack of familiarity with HARVEST, when using a task-based evaluation, performance and time-to-task completion was maintained in patient review scenarios using HARVEST alone or the standard clinical information system at our institution. Subjects reported very high satisfaction with HARVEST and interest in using the system in their daily practice. DISCUSSION: HARVEST is available for wide deployment at our institution. Evaluation provided informative feedback and directions for future improvements. CONCLUSIONS: HARVEST was designed to address the unmet need for clinicians at the point of care, facilitating review of essential patient information. The deployment of HARVEST in our institution allows us to study patient record summarization as an informatics intervention in a real-world setting. It also provides an opportunity to learn how clinicians use the summarizer, enabling informed interface and content iteration and optimization to improve patient care. Jamie S. Hirsch, Jessica S. Tanenbaum, Sharon Lipsky Gorman, Connie Liu, Eric Schmitz, Dritan Hashorva, Artem Ervits, David K. Vawdrey, Marc Sturm, Noémie Elhadad |
J. Am. Medical Informatics Assoc. | 10 |
| 2015 | Risk prediction for chronic kidney disease progression using heterogeneous electronic health record data and time series analysisabstractBACKGROUND: As adoption of electronic health records continues to increase, there is an opportunity to incorporate clinical documentation as well as laboratory values and demographics into risk prediction modeling. OBJECTIVE: The authors develop a risk prediction model for chronic kidney disease (CKD) progression from stage III to stage IV that includes longitudinal data and features drawn from clinical documentation. METHODS: The study cohort consisted of 2908 primary-care clinic patients who had at least three visits prior to January 1, 2013 and developed CKD stage III during their documented history. Development and validation cohorts were randomly selected from this cohort and the study datasets included longitudinal inpatient and outpatient data from these populations. Time series analysis (Kalman filter) and survival analysis (Cox proportional hazards) were combined to produce a range of risk models. These models were evaluated using concordance, a discriminatory statistic. RESULTS: A risk model incorporating longitudinal data on clinical documentation and laboratory test results (concordance 0.849) predicts progression from state III CKD to stage IV CKD more accurately when compared to a similar model without laboratory test results (concordance 0.733, P<.001), a model that only considers the most recent laboratory test results (concordance 0.819, P < .031) and a model based on estimated glomerular filtration rate (concordance 0.779, P < .001). CONCLUSIONS: A risk prediction model that takes longitudinal laboratory test results and clinical documentation into consideration can predict CKD progression from stage III to stage IV more accurately than three models that do not take all of these variables into consideration. Adler J. Perotte, Rajesh Ranganath, Jamie S. Hirsch, David M. Blei, Noémie Elhadad |
J. Am. Medical Informatics Assoc. | 5 |
| 2015 | Automated methods for the summarization of electronic health recordsabstractOBJECTIVES: This review examines work on automated summarization of electronic health record (EHR) data and in particular, individual patient record summarization. We organize the published research and highlight methodological challenges in the area of EHR summarization implementation. TARGET AUDIENCE: The target audience for this review includes researchers, designers, and informaticians who are concerned about the problem of information overload in the clinical setting as well as both users and developers of clinical summarization systems. SCOPE: Automated summarization has been a long-studied subject in the fields of natural language processing and human-computer interaction, but the translation of summarization and visualization methods to the complexity of the clinical workflow is slow moving. We assess work in aggregating and visualizing patient information with a particular focus on methods for detecting and removing redundancy, describing temporality, determining salience, accounting for missing data, and taking advantage of encoded clinical knowledge. We identify and discuss open challenges critical to the implementation and use of robust EHR summarization systems. Rimma Perotte, Noémie Elhadad |
J. Am. Medical Informatics Assoc. | 2 |
| 2015 | Evaluating the state of the art in disorder recognition and normalization of the clinical narrativeabstractOBJECTIVE: The ShARe/CLEF eHealth 2013 Evaluation Lab Task 1 was organized to evaluate the state of the art on the clinical text in (i) disorder mention identification/recognition based on Unified Medical Language System (UMLS) definition (Task 1a) and (ii) disorder mention normalization to an ontology (Task 1b). Such a community evaluation has not been previously executed. Task 1a included a total of 22 system submissions, and Task 1b included 17. Most of the systems employed a combination of rules and machine learners. MATERIALS AND METHODS: We used a subset of the Shared Annotated Resources (ShARe) corpus of annotated clinical text--199 clinical notes for training and 99 for testing (roughly 180 K words in total). We provided the community with the annotated gold standard training documents to build systems to identify and normalize disorder mentions. The systems were tested on a held-out gold standard test set to measure their performance. RESULTS: For Task 1a, the best-performing system achieved an F1 score of 0.75 (0.80 precision; 0.71 recall). For Task 1b, another system performed best with an accuracy of 0.59. DISCUSSION: Most of the participating systems used a hybrid approach by supplementing machine-learning algorithms with features generated by rules and gazetteers created from the training data and from external resources. CONCLUSIONS: The task of disorder normalization is more challenging than that of identification. The ShARe corpus is available to the community as a reference standard for future studies. Sameer Pradhan, Noémie Elhadad, Brett R. South, David Martínez 0001, Lee M. Christensen, Amy Vogel, Hanna Suominen, Wendy W. Chapman, Guergana K. Savova |
J. Am. Medical Informatics Assoc. | 2 |
| 2015 | Learning probabilistic phenotypes from heterogeneous EHR data
Rimma Perotte, Adler J. Perotte, Edouard Grave, John Angiolillo, Chris Wiggins 0001, Noémie Elhadad |
J. Biomed. Informatics | 6 |
| 2014 | Cross-narrative Temporal Ordering of Medical EventsabstractCross-narrative temporal ordering of medical events is essential to the task of generating a comprehensive timeline over a patient's history.We address the problem of aligning multiple medical event sequences, corresponding to different clinical narratives, comparing the following approaches: (1) A novel weighted finite state transducer representation of medical event sequences that enables composition and search for decoding, and (2) Dynamic programming with iterative pairwise alignment of multiple sequences using global and local alignment algorithms.The cross-narrative coreference and temporal relation weights used in both these approaches are learned from a corpus of clinical narratives.We present results using both approaches and observe that the finite state transducer approach performs performs significantly better than the dynamic programming one by 6.8% for the problem of multiple-sequence alignment. Preethi Raghavan, Eric Fosler-Lussier, Noémie Elhadad, Albert M. Lai |
ACL (1) | 3 |
| 2014 | Model selection for EHR laboratory variables: how physiology and the health care process can influence EHR laboratory data and their model representations
David J. Albers, Rimma Perotte, Noémie Elhadad, George Hripcsak |
AMIA | 3 |
| 2014 | HARVEST, a Holistic Patient Record Summarizer at the Point of Care
Noémie Elhadad, Sharon Lipsky Gorman, Jamie S. Hirsch, Connie Liu, David K. Vawdrey, Marc Sturm |
AMIA | 1 |
| 2014 | Characterizing the Sublanguage of Online Breast Cancer Forums for Medications, Symptoms, and Emotions
Noémie Elhadad, Shaodian Zhang, Patricia Driscoll, Samuel Brody |
AMIA | 1 |
| 2014 | Disease/Disorder Semantic Template Filling - Information Extraction Challenge in the ShARe/CLEF eHealth Evaluation Lab 2014
Sumithra Velupillai, Danielle L. Mowery, Lee M. Christensen, Noémie Elhadad, Sameer Pradhan, Guergana K. Savova, Wendy W. Chapman |
AMIA | 4 |
| 2014 | Does Sustained Participation in an Online Health Community Affect Sentiment?
Shaodian Zhang, Erin O'Carroll Bantum, Jason E. Owen, Noémie Elhadad |
AMIA | 4 |
| 2014 | Diagnosis code assignment: models and evaluation metricsabstractBACKGROUND AND OBJECTIVE: The volume of healthcare data is growing rapidly with the adoption of health information technology. We focus on automated ICD9 code assignment from discharge summary content and methods for evaluating such assignments. METHODS: We study ICD9 diagnosis codes and discharge summaries from the publicly available Multiparameter Intelligent Monitoring in Intensive Care II (MIMIC II) repository. We experiment with two coding approaches: one that treats each ICD9 code independently of each other (flat classifier), and one that leverages the hierarchical nature of ICD9 codes into its modeling (hierarchy-based classifier). We propose novel evaluation metrics, which reflect the distances among gold-standard and predicted codes and their locations in the ICD9 tree. Experimental setup, code for modeling, and evaluation scripts are made available to the research community. RESULTS: The hierarchy-based classifier outperforms the flat classifier with F-measures of 39.5% and 27.6%, respectively, when trained on 20,533 documents and tested on 2282 documents. While recall is improved at the expense of precision, our novel evaluation metrics show a more refined assessment: for instance, the hierarchy-based classifier identifies the correct sub-tree of gold-standard codes more often than the flat classifier. Error analysis reveals that gold-standard codes are not perfect, and as such the recall and precision are likely underestimated. CONCLUSIONS: Hierarchy-based classification yields better ICD9 coding than flat classification for MIMIC patients. Automated ICD9 coding is an example of a task for which data and tools can be shared and for which the research community can work together to build on shared models and advance the state of the art. Adler J. Perotte, Rimma Perotte, Karthik Natarajan, Nicole Gray Weiskopf, Frank D. Wood, Noémie Elhadad |
J. Am. Medical Informatics Assoc. | 6 |
| 2014 | Temporal trends of hemoglobin A1c testingabstractOBJECTIVE: The study of utilization patterns can quantify potential overuse of laboratory tests and find new ways to reduce healthcare costs. We demonstrate the use of distributional analytics for comparing electronic health record (EHR) laboratory test orders across time to diagnose and quantify overutilization. MATERIALS AND METHODS: We looked at hemoglobin A1c (HbA1c) testing across 119,000 patients and 15 years of hospital records. We examined the patterns of HbA1c ordering before and after the publication of the 2002 American Diabetes Association guidelines for HbA1c testing. We conducted analyses to answer three questions. What are the patterns of HbA1c ordering? Do HbA1c orders follow the guidelines with respect to frequency of measurement? If not, how and why do they depart from the guidelines? RESULTS: The raw number of HbA1c orderings has steadily increased over time, with a specific increase in low-measurement orderings (<6.5%). There is a change in ordering pattern following the 2002 guideline (p<0.001). However, by comparing ordering distributions, we found that the changes do not reflect the guidelines and rather exhibit a new practice of rapid-repeat testing. The rapid-retesting phenomenon does not follow the 2009 guidelines for diabetes diagnosis either, illustrated by a stratified HbA1c value analysis. DISCUSSION: Results suggest HbA1c test overutilization, and contributing factors include lack of care coordination, unexpected values prompting retesting, and point-of-care tests followed by confirmatory laboratory tests. CONCLUSIONS: We present a method of comparing ordering distributions in an EHR across time as a useful diagnostic approach for identifying and assessing the trend of inappropriate use over time. Rimma Perotte, David J. Albers, George Hripcsak, Jorge L. Sepulveda, Noémie Elhadad |
J. Am. Medical Informatics Assoc. | 5 |
| 2014 | A review of approaches to identifying patient phenotype cohorts using electronic health recordsabstractOBJECTIVE: To summarize literature describing approaches aimed at automatically identifying patients with a common phenotype. MATERIALS AND METHODS: We performed a review of studies describing systems or reporting techniques developed for identifying cohorts of patients with specific phenotypes. Every full text article published in (1) Journal of American Medical Informatics Association, (2) Journal of Biomedical Informatics, (3) Proceedings of the Annual American Medical Informatics Association Symposium, and (4) Proceedings of Clinical Research Informatics Conference within the past 3 years was assessed for inclusion in the review. Only articles using automated techniques were included. RESULTS: Ninety-seven articles met our inclusion criteria. Forty-six used natural language processing (NLP)-based techniques, 24 described rule-based systems, 41 used statistical analyses, data mining, or machine learning techniques, while 22 described hybrid systems. Nine articles described the architecture of large-scale systems developed for determining cohort eligibility of patients. DISCUSSION: We observe that there is a rise in the number of studies associated with cohort identification using electronic medical records. Statistical analyses or machine learning, followed by NLP techniques, are gaining popularity over the years in comparison with rule-based systems. CONCLUSIONS: There are a variety of approaches for classifying patients into a particular phenotype. Different techniques and data sources are used, and good performance is reported on datasets at respective institutions. However, no system makes comprehensive use of electronic medical records addressing all of their known weaknesses. Chaitanya P. Shivade, Preethi Raghavan, Eric Fosler-Lussier, Peter J. Embí, Noémie Elhadad, Stephen B. Johnson, Albert M. Lai |
J. Am. Medical Informatics Assoc. | 5 |
| 2014 | Identifying and mitigating biases in EHR laboratory tests
Rimma Perotte, David J. Albers, Jorge L. Sepulveda, Noémie Elhadad |
J. Biomed. Informatics | 4 |
| 2013 | Using patient laboratory measurement values and dynamics to deconvolve EHR bias and define acuity-based phenotypes
David J. Albers, Rimma Perotte, George Hripcsak, Noémie Elhadad |
AMIA | 4 |
| 2013 | Lessons Learned in Replicating Data-Driven Experiments in Multiple Medical Systems and Patient Populations
Samantha Kleinberg, Noémie Elhadad |
AMIA | 2 |
| 2013 | Panel: Shared Resources, Shared Code, and Shared Activities in Clinical Natural Language Processing
Guergana K. Savova, Wendy W. Chapman, Noémie Elhadad, Martha Palmer |
AMIA | 3 |
| 2013 | Redundancy in electronic health record corpora: analysis, impact on text mining performance and mitigation strategiesabstractBACKGROUND: The increasing availability of Electronic Health Record (EHR) data and specifically free-text patient notes presents opportunities for phenotype extraction. Text-mining methods in particular can help disease modeling by mapping named-entities mentions to terminologies and clustering semantically related terms. EHR corpora, however, exhibit specific statistical and linguistic characteristics when compared with corpora in the biomedical literature domain. We focus on copy-and-paste redundancy: clinicians typically copy and paste information from previous notes when documenting a current patient encounter. Thus, within a longitudinal patient record, one expects to observe heavy redundancy. In this paper, we ask three research questions: (i) How can redundancy be quantified in large-scale text corpora? (ii) Conventional wisdom is that larger corpora yield better results in text mining. But how does the observed EHR redundancy affect text mining? Does such redundancy introduce a bias that distorts learned models? Or does the redundancy introduce benefits by highlighting stable and important subsets of the corpus? (iii) How can one mitigate the impact of redundancy on text mining? RESULTS: We analyze a large-scale EHR corpus and quantify redundancy both in terms of word and semantic concept repetition. We observe redundancy levels of about 30% and non-standard distribution of both words and concepts. We measure the impact of redundancy on two standard text-mining applications: collocation identification and topic modeling. We compare the results of these methods on synthetic data with controlled levels of redundancy and observe significant performance variation. Finally, we compare two mitigation strategies to avoid redundancy-induced bias: (i) a baseline strategy, keeping only the last note for each patient in the corpus; (ii) removing redundant notes with an efficient fingerprinting-based algorithm. (a)For text mining, preprocessing the EHR corpus with fingerprinting yields significantly better results. CONCLUSIONS: Before applying text-mining techniques, one must pay careful attention to the structure of the analyzed corpora. While the importance of data cleaning has been known for low-level text characteristics (e.g., encoding and spelling), high-level and difficult-to-quantify corpus characteristics, such as naturally occurring redundancy, can also hurt text mining. Fingerprinting enables text-mining techniques to leverage available data in the EHR corpus, while avoiding the bias introduced by redundancy. Raphael Cohen, Michael Elhadad, Noémie Elhadad |
BMC Bioinform. | 3 |
| 2013 | Unsupervised biomedical named entity recognition: Experiments with clinical and biological textsabstractNamed entity recognition is a crucial component of biomedical natural language processing, enabling information extraction and ultimately reasoning over and knowledge discovery from text. Much progress has been made in the design of rule-based and supervised tools, but they are often genre and task dependent. As such, adapting them to different genres of text or identifying new types of entities requires major effort in re-annotation or rule development. In this paper, we propose an unsupervised approach to extracting named entities from biomedical text. We describe a stepwise solution to tackle the challenges of entity boundary detection and entity type classification without relying on any handcrafted rules, heuristics, or annotated data. A noun phrase chunker followed by a filter based on inverse document frequency extracts candidate entities from free text. Classification of candidate entities into categories of interest is carried out by leveraging principles from distributional semantics. Experiments show that our system, especially the entity classification step, yields competitive results on two popular biomedical datasets of clinical notes and biological literature, and outperforms a baseline dictionary match approach. Detailed error analysis provides a road map for future work. Shaodian Zhang, Noémie Elhadad |
J. Biomed. Informatics | 2 |
| 2012 | A hybrid knowledge-based and data-driven approach to identifying semantically similar concepts
Rimma Perotte, Noémie Elhadad |
J. Biomed. Informatics | 2 |
| 2012 | A new clustering method for detecting rare senses of abbreviations in clinical notes
Hua Xu 0001, Yonghui Wu 0001, Noémie Elhadad, Peter D. Stetson, Carol Friedman |
J. Biomed. Informatics | 3 |
| 2011 | Hierarchically Supervised Latent Dirichlet AllocationabstractWe introduce hierarchically supervised latent Dirichlet allocation (HSLDA), a model for hierarchically and multiply labeled bag-of-word data. Examples of such data include web pages and their placement in directories, product descriptions and associated categories from product hierarchies, and free-text clinical records and their assigned diagnosis codes. Out-of-sample label prediction is the primary goal of this work, but improved lower-dimensional representations of the bag-of-word data are also of interest. We demonstrate HSLDA on large-scale data from clinical document labeling and retail product categorization tasks. We show that leveraging the structure from hierarchical labels improves out-of-sample label prediction substantially when compared to models that do not. Adler J. Perotte, Frank D. Wood, Noémie Elhadad, Nicholas Bartlett |
NIPS | 3 |
| 2010 | An Unsupervised Aspect-Sentiment Model for Online Reviews
Samuel Brody, Noémie Elhadad |
HLT-NAACL | 2 |
| 2009 | Perplexity Analysis of Obesity News Coverage
Delano J. McFarlane, Noémie Elhadad, Rita Kukafka |
AMIA | 2 |
| 2009 | Comparing evaluation techniques for text readability software for adults with intellectual disabilitiesabstractIn this paper, we compare alternative techniques for evaluating a software system for simplifying the readability of texts for adults with mild intellectual disabilities (ID). We introduce our research on the development of software to automatically simplify news articles, display them, and read them aloud for adults with ID. Using a Wizard-of-Oz prototype, we conducted experiments with a group of adults with ID to test alternative formats of questions to measure comprehension of the information in the news articles. We have found that some forms of questions work well at measuring the difficulty level of a text: multiple-choice questions with three answer choices, each illustrated with clip-art or a photo. Some types of questions do a poor job: yes/no questions and Likert-scale questions in which participants report their perception of the text's difficulty level. Our findings inform the design of future evaluation studies of computational linguistic software for adults with ID; this study may also be of interest to researchers conducting usability studies or other surveys with adults with ID. Matt Huenerfauth, Lijun Feng, Noémie Elhadad |
ASSETS | 3 |
| 2009 | Cognitively Motivated Features for Readability Assessment
Lijun Feng, Noémie Elhadad, Matt Huenerfauth |
EACL | 2 |
| 2009 | Beyond the Stars: Improving Rating Predictions using Review Text Content
Gayatree Ganu, Noémie Elhadad, Amélie Marian |
WebDB | 2 |
| 2009 | Research Paper: Using Empiric Semantic Correlation to Interpret Temporal Assertions in Clinical TextsabstractOBJECTIVE: To measure the uncertainty of temporal assertions like "3 weeks ago" in clinical texts. DESIGN: Temporal assertions extracted from narrative clinical reports were compared to facts extracted from a structured clinical database for the same patients. MEASUREMENTS: The authors correlated the assertions and the facts to determine the dependence of the uncertainty of the assertions on the semantic and lexical properties of the assertions. RESULTS: The observed deviation between the stated duration and actual duration averaged about 20% of the stated deviation. Linear regression revealed that assertions about events further in the past tend to be more uncertain, smaller numeric values tend to be more uncertain (1 mo v. 30 d), and round numbers tend to be more uncertain (10 versus 11 yrs). CONCLUSIONS: The authors empirically derived semantics behind statements of duration using "ago," and verified intuitions about how numbers are used. George Hripcsak, Noémie Elhadad, Yueh-Hsia Chen, Li Zhou 0007, Frances P. Morrison |
J. Am. Medical Informatics Assoc. | 2 |
| 2008 | Extracting Structured Medication Event Information from Discharge Summaries
Sigfried Gold, Noémie Elhadad, James J. Cimino, George Hripcsak |
AMIA | 2 |
| 2008 | Use of Semantic Features to Classify Patient Smoking Status
Patrick J. McCormick, Noémie Elhadad, Peter D. Stetson |
AMIA | 2 |
| 2008 | Content and Structure of Clinical Problem Lists: A Corpus Analysis
Tielman Van Vleck, Adam B. Wilcox, Peter D. Stetson, Stephen B. Johnson, Noémie Elhadad |
AMIA | 5 |
| 2008 | Automated Knowledge Acquisition from Clinical Narrative Reports
Amy E. Chused, Noémie Elhadad, Carol Friedman, Marianthi Markatou |
AMIA | 3 |
| 2006 | Comprehending Technical Texts: Predicting and Defining Unfamiliar Terms
Noémie Elhadad |
AMIA | 1 |
| 2005 | Facilitating Physicians' Access to Information via Tailored Text Summarization
Noémie Elhadad, Kathy McKeown, David R. Kaufman, Desmond A. Jordan |
AMIA | 1 |
| 2005 | Customization in a unified framework for summarizing medical literature
Noémie Elhadad, Min-Yen Kan, Judith L. Klavans, Kathy McKeown |
Artif. Intell. Medicine | 1 |
| 2004 | User-Sensitive Text Summarization
Noémie Elhadad |
AAAI | 1 |
| 2003 | Sentence Alignment for Monolingual Comparable Corpora
Regina Barzilay, Noémie Elhadad |
EMNLP | 2 |
| 2002 | Collection and linguistic processing of a large-scale corpus of medical articles
Simone Teufel, Noémie Elhadad |
LREC | 2 |
| 2002 | Inferring Strategies for Sentence Ordering in Multidocument News SummarizationabstractThe problem of organizing information for multidocument summarization so that the generated summary is coherent has received relatively little attention. While sentence ordering for single document summarization can be determined from the ordering of sentences in the input article, this is not the case for multidocument summarization where summary sentences may be drawn from different input articles. In this paper, we propose a methodology for studying the properties of ordering information in the news genre and describe experiments done on a corpus of multiple acceptable orderings we developed for the task. Based on these experiments, we implemented a strategy for ordering information that combines constraints from chronological order of events and topical relatedness. Evaluation of our augmented algorithm shows a significant improvement of the ordering over two baseline strategies. Regina Barzilay, Noémie Elhadad, Kathy McKeown |
J. Artif. Intell. Res. | 2 |