EDBT 2026 Demo / reviewers in the wild / expert
Sanda M. Harabagiu
dblp:51/3845
· DBLP profile ↗
84ranked-venue papers
25as first author
13since 2021 · last 2026
0000-0002-8186-1501ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 52 · 19 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 23 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 13 · 5 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-authorSystems, architecture and hardware · 3 · 1 first-authorHuman-computer interaction and ubiquitous computing · 3 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | The MISOMEM-Val Dataset for Identifying Human Values in Misogynistic Memes
Rakshitha Rao Ailneni, Sanda M. Harabagiu |
LREC | 2 |
| 2026 | Exploration of How Hate Is Framed on Social Media
Rakshitha Rao Ailneni, Sanda M. Harabagiu |
LREC | 2 |
| 2025 | Automatically Discovering How Misogyny is Framed on Social MediaabstractRakshitha Rao Ailneni, Sanda M. Harabagiu. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Rakshitha Rao Ailneni, Sanda M. Harabagiu |
NAACL (Long Papers) | 2 |
| 2024 | Tree-of-Counterfactual Prompting for Zero-Shot Stance DetectionabstractStance detection enables the inference of attitudes from human communications.Automatic stance identification was mostly cast as a classification problem.However, stance decisions involve complex judgments, which can be nowadays generated by prompting Large Language Models (LLMs).In this paper we present a new method for stance identification which (1) relies on a new prompting framework, called Tree-of-Counterfactual prompting; (2) operates not only on textual communications, but also on images; (3) allows more than one stance object type; and (4) requires no examples of stance attribution, thus it is a "Tabula Rasa" Zero-Shot Stance Detection (TR-ZSSD) method.Our experiments indicate surprisingly promising results, outperforming fine-tuned stance detection systems. Maxwell A. Weinzierl, Sanda M. Harabagiu |
ACL (1) | 2 |
| 2024 | The Impact of Stance Object Type on the Quality of Stance DetectionabstractStance as an expression of an author’s standpoint and as a means of communication has long been studied by computational linguists. Automatically identifying the stance of a subject toward an object is an active area of research in natural language processing. Significant work has employed topics and claims as the object of stance, with frames of communication becoming more recently considered as alternative objects of stance. However, little attention has been paid to finding what are the benefits and what are the drawbacks when inferring the stance of a text towards different possible stance objects. In this paper we seek to answer this question by analyzing the implied knowledge and the judgments required when deciding the stance of a text towards each stance object type. Our analysis informed experiments with models capable of inferring the stance of a text towards any of the stance object types considered, namely topics, claims, and frames of communication. Experiments clearly indicate that it is best to infer the stance of a text towards a frame of communication, rather than a claim or a topic. It is also better to infer the stance of a text towards a claim rather than a topic. Therefore we advocate that rather than continuing efforts to annotate the stance of texts towards topics, it is better to use those efforts to produce annotations towards frames of communication. These efforts will allow us to better capture the stance towards claims and topics as well. Maxwell A. Weinzierl, Sanda M. Harabagiu |
LREC/COLING | 2 |
| 2024 | Discovering and Articulating Frames of Communication from Social Media Using Chain-of-Thought ReasoningabstractFrames of Communication (FoCs) are ubiquitous in social media discourse.They define what counts as a problem, diagnose what is causing the problem, elicit moral judgments and imply remedies for resolving the problem (Entman, 1993).Most research on automatic frame detection involved the recognition of the problems addressed by frames, but did not consider the articulation of frames.Articulating an FoC involves reasoning with salient problems, their cause and eventual solution.In this paper we present a method for Discovering and Articulating FoCs (DA-FoC) that relies on a combination of Chain-of-Thought prompting (Wei et al., 2022a) of large language models (LLMs) with In-Context Active Curriculum Learning.Very promising evaluation results indicate that 86.72% of the FoCs encoded by communication experts on the same reference dataset were also uncovered by DA-FoC.Moreover, DA-FoC uncovered many new FoCs, which escaped the experts.Interestingly, 55.1% of the known FoCs were judged as being better articulated than the human-written ones, while 93.8% of the new FoCs were judged as having sound rationale and being clearly articulated. Maxwell A. Weinzierl, Sanda M. Harabagiu |
EACL (1) | 2 |
| 2023 | Identification of Multimodal Stance Towards Frames of CommunicationabstractFrames of communication are often evoked in multimedia documents.When an author decides to add an image to a text, one or both of the modalities may evoke a communication frame.Moreover, when evoking the frame, the author also conveys her/his stance towards the frame.Until now, determining if the author is in favor of, against or has no stance towards the frame was performed automatically only when processing texts.This is due to the absence of stance annotations on multimedia documents.In this paper we introduce MMVAX-STANCE, a dataset of 11,300 multimedia documents retrieved from social media, which have stance annotations towards 113 different frames of communication.This dataset allowed us to experiment with several models of multimedia stance detection, which revealed important interactions between texts and images in the inference of stance towards communication frames.When inferring the text/image relations, a set of 46,606 synthetic examples of multimodal documents with known stance was generated.This greatly impacted the quality of identifying multimedia stance, yielding an improvement of 20% in F1-score. Component Definition Examples of Frames of Communication ConfidenceTrust in the security and effectiveness of 2 Pfizer COVID-19 vaccine may cause anaphylaxis vaccinations, the health authorities, and in people with polyethylene glycol (PEG) allergy.the health officials who recommend 2 The Government has provided plenty of safety and develop vaccines.information about the COVID-19 vaccines. ComplacencyComplacency and laziness to get vaccinated 2 Preference for getting COVID-19 and fighting due to low perceived risk of infections.it off than vaccinating. ConstraintsStructural or psychological hurdles that 2 It takes courage both to vaccinate against make vaccination difficult or costly. COVID-19 and to refuse the vaccine. CalculationDegree to which personal costs and benefits 2 COVID-19 vaccines protect against the emerging of vaccination are weighted. variants. CollectiveWillingness to protect others and to 2 Vaccination is key in protecting yourself and others Responsibility eliminate infectious diseases.against COVID-19.Compliance Support for societal monitoring and sanctioning 2 People choosing not to get the COVID-19 vaccine of people who are not vaccinated.should not lose venue access/travel to some countries. ConspiracyConspiracy thinking and belief in 2 COVID-19 vaccines make you 5G compatible.fake news related to vaccination.2 The COVID vaccine renders pregnancies risky. Maxwell A. Weinzierl, Sanda M. Harabagiu |
EMNLP | 2 |
| 2023 | Epidemic Question Answering: question generation and entailment for Answer Nugget discoveryabstractOBJECTIVE: The rapidly growing body of communications during the COVID-19 pandemic posed a challenge to information seekers, who struggled to find answers to their specific and changing information needs. We designed a Question Answering (QA) system capable of answering ad-hoc questions about the COVID-19 disease, its causal virus SARS-CoV-2, and the recommended response to the pandemic. MATERIALS AND METHODS: The QA system incorporates, in addition to relevance models, automatic generation of questions from relevant sentences. We relied on entailment between questions for (1) pinpointing answers and (2) selecting novel answers early in the list of its results. RESULTS: The QA system produced state-of-the-art results when processing questions asked by experts (eg, researchers, scientists, or clinicians) and competitive results when processing questions asked by consumers of health information. Although state-of-the-art models for question generation and question entailment were used, more than half of the answers were missed, due to the limitations of the relevance models employed. DISCUSSION: Although question entailment enabled by automatic question generation is the cornerstone of our QA system's architecture, question entailment did not prove to always be reliable or sufficient in ranking the answers. Question entailment should be enhanced with additional inferential capabilities. CONCLUSION: The QA system presented in this article produced state-of-the-art results processing expert questions and competitive results processing consumer questions. Improvements should be considered by using better relevance models and enhanced inference methods. Moreover, experts and consumers have different answer expectations, which should be accounted for in future QA development. Maxwell A. Weinzierl, Sanda M. Harabagiu |
J. Am. Medical Informatics Assoc. | 2 |
| 2022 | From Hesitancy Framings to Vaccine Hesitancy Profiles: A Journey of Stance, Ontological Commitments and Moral Foundations
Maxwell A. Weinzierl, Sanda M. Harabagiu |
ICWSM | 2 |
| 2022 | VaccineLies: A Natural Language Resource for Learning to Recognize Misinformation about the COVID-19 and HPV VaccinesabstractBillions of COVID-19 vaccines have been administered, but many remain hesitant. Misinformation about the COVID-19 vaccines and other vaccines, propagating on social media, is believed to drive hesitancy towards vaccination. The ability to automatically recognize misinformation targeting vaccines on Twitter depends on the availability of data resources. In this paper we present VaccineLies, a large collection of tweets propagating misinformation about two vaccines: the COVID-19 vaccines and the Human Papillomavirus (HPV) vaccines. Misinformation targets are organized in vaccine-specific taxonomies, which reveal the misinformation themes and concerns. The ontological commitments of the misinformation taxonomies provide an understanding of which misinformation themes and concerns dominate the discourse about the two vaccines covered in VaccineLies. The organization into training, testing and development sets of VaccineLies invites the development of novel supervised methods for detecting misinformation on Twitter and identifying the stance towards it. Furthermore, VaccineLies can be a stepping stone for the development of datasets focusing on misinformation targeting additional vaccines. Maxwell A. Weinzierl, Sanda M. Harabagiu |
LREC | 2 |
| 2022 | Identifying the Adoption or Rejection of Misinformation Targeting COVID-19 Vaccines in Twitter DiscourseabstractAlthough billions of COVID-19 vaccines have been administered, too many people remain hesitant. Misinformation about the COVID-19 vaccines, propagating on social media, is believed to drive hesitancy towards vaccination. However, exposure to misinformation does not necessarily indicate misinformation adoption. In this paper we describe a novel framework for identifying the stance towards misinformation, relying on attitude consistency and its properties. The interactions between attitude consistency, adoption or rejection of misinformation and the content of microblogs are exploited in a novel neural architecture, where the stance towards misinformation is organized in a knowledge graph. This new neural framework is enabling the identification of stance towards misinformation about COVID-19 vaccines with state-of-the-art results. The experiments are performed on a new dataset of misinformation towards COVID-19 vaccines, called CoVaxLies, collected from recent Twitter discourse. Because CoVaxLies provides a taxonomy of the misinformation about COVID-19 vaccines, we are able to show which type of misinformation is mostly adopted and which is mostly rejected. Maxwell A. Weinzierl, Sanda M. Harabagiu |
WWW | 2 |
| 2021 | Misinformation Adoption or Rejection in the Era of COVID-19
Maxwell A. Weinzierl, Suellen Hopfer, Sanda M. Harabagiu |
ICWSM | 3 |
| 2021 | Automatic detection of COVID-19 vaccine misinformation with graph link prediction
Maxwell A. Weinzierl, Sanda M. Harabagiu |
J. Biomed. Informatics | 2 |
| 2020 | The Language of Brain Signals: Natural Language Processing of Electroencephalography ReportsabstractBrain signals are captured by clinical electroencephalography (EEG) which is an excellent tool for probing neural function. When EEG tests are performed, a textual EEG report is generated by the neurologist to document the findings, thus using language that describes the brain signals and its clinical correlations. Even with the impetus provided by the BRAIN initiative (brainitititive.nih.gov), there are no annotations available in texts that capture language describing the brain activities and their correlations with various pathologies. In this paper we describe an annotation effort carried out on a large corpus of EEG reports, providing examples of EEG-specific and clinically relevant concepts. In addition, we detail our annotation schema for brain signal attributes. We also discuss the resulting annotation of long-distance relations between concepts in EEG reports. By exemplifying a self-attention joint-learning to predict similar annotations in the EEG report corpus, we discuss the promising results, hoping that our effort will inform the design of novel knowledge capture techniques that will include the language of brain signals. Ramón Maldonado, Sanda M. Harabagiu |
LREC | 2 |
| 2020 | The impact of learning Unified Medical Language System knowledge embeddings in relation extraction from biomedical textsabstractOBJECTIVE: We explored how knowledge embeddings (KEs) learned from the Unified Medical Language System (UMLS) Metathesaurus impact the quality of relation extraction on 2 diverse sets of biomedical texts. MATERIALS AND METHODS: Two forms of KEs were learned for concepts and relation types from the UMLS Metathesaurus, namely lexicalized knowledge embeddings (LKEs) and unlexicalized KEs. A knowledge embedding encoder (KEE) enabled learning either LKEs or unlexicalized KEs as well as neural models capable of producing LKEs for mentions of biomedical concepts in texts and relation types that are not encoded in the UMLS Metathesaurus. This allowed us to design the relation extraction with knowledge embeddings (REKE) system, which incorporates either LKEs or unlexicalized KEs produced for relation types of interest and their arguments. RESULTS: The incorporation of either LKEs or unlexicalized KE in REKE advances the state of the art in relation extraction on 2 relation extraction datasets: the 2010 i2b2/VA dataset and the 2013 Drug-Drug Interaction Extraction Challenge corpus. Moreover, the impact of LKEs is superior, achieving F1 scores of 78.2 and 82.0, respectively. DISCUSSION: REKE not only highlights the importance of incorporating knowledge encoded in the UMLS Metathesaurus in a novel way, through 2 possible forms of KEs, but it also showcases the subtleties of incorporating KEs in relation extraction systems. CONCLUSIONS: Incorporating LKEs informed by the UMLS Metathesaurus in a relation extraction system operating on biomedical texts shows significant promise. We present the REKE system, which establishes new state-of-the-art results for relation extraction on 2 datasets when using LKEs. Maxwell A. Weinzierl, Ramón Maldonado, Sanda M. Harabagiu |
J. Am. Medical Informatics Assoc. | 3 |
| 2019 | Bootstrapping Adversarial Learning of Biomedical Ontology Alignments
Ramón Maldonado, Sanda M. Harabagiu |
AMIA | 2 |
| 2019 | Active deep learning for the identification of concepts and relations in electroencephalography reports
Ramón Maldonado, Sanda M. Harabagiu |
J. Biomed. Informatics | 2 |
| 2018 | The Impact of Inferring Treatments on Information Retrieval for Precision Medicine
Travis R. Goodwin, Sanda M. Harabagiu |
AMIA | 2 |
| 2018 | Hierarchical Attention-Based Prediction Model for Discovering the Persistence of Chronic Opioid Therapy from a large Clinical Dataset
Ramón Maldonado, Mark Sullivan, Meliha Yetisgen, Sanda M. Harabagiu |
AMIA | 4 |
| 2018 | The Role of a Deep-Learning Method for Negation Detection in Patient Cohort Identification from Electroencephalography Reports
Stuart J. Taylor, Sanda M. Harabagiu |
AMIA | 2 |
| 2018 | Knowledge Representations and Inference Techniques for Medical Question AnsweringabstractAnswering medical questions related to complex medical cases, as required in modern Clinical Decision Support (CDS) systems, imposes (1) access to vast medical knowledge and (2) sophisticated inference techniques. In this article, we examine the representation and role of combining medical knowledge automatically derived from (a) clinical practice and (b) research findings for inferring answers to medical questions. Knowledge from medical practice was distilled from a vast Electronic Medical Record (EMR) system, while research knowledge was processed from biomedical articles available in PubMed Central. The knowledge automatically acquired from the EMR system took into account the clinical picture and therapy recognized from each medical record to generate a probabilistic Markov network denoted as a Clinical Picture and Therapy Graph (CPTG). Moreover, we represented the background of medical questions available from the description of each complex medical case as a medical knowledge sketch. We considered three possible representations of medical knowledge sketches that were used by four different probabilistic inference methods to pinpoint the answers from the CPTG. In addition, several answer-informed relevance models were developed to provide a ranked list of biomedical articles containing the answers. Evaluations on the TREC-CDS data show which of the medical knowledge representations and inference methods perform optimally. The experiments indicate an improvement of biomedical article ranking by 49% over state-of-the-art results. Travis R. Goodwin, Sanda M. Harabagiu |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2017 | A Data-driven Method for the Early Identification of Diabetes and Prediabetes
Travis R. Goodwin, Michael E. Bowen, Sanda M. Harabagiu |
AMIA | 3 |
| 2017 | Inferring Clinical Correlations from EEG Reports with Deep Neural Learning
Travis R. Goodwin, Sanda M. Harabagiu |
AMIA | 2 |
| 2017 | Deep Learning Meets Biomedical Ontologies: Knowledge Embeddings for Epilepsy
Ramón Maldonado, Travis R. Goodwin, Michael A. Skinner, Sanda M. Harabagiu |
AMIA | 4 |
| 2017 | An Evaluation of Syntactic Dependency Parsers on Clinical Data
Stuart J. Taylor, Travis R. Goodwin, Sanda M. Harabagiu |
AMIA | 3 |
| 2016 | Multi-modal Patient Cohort Identification from EEG Report and Signal Data
Travis R. Goodwin, Sanda M. Harabagiu |
AMIA | 2 |
| 2016 | Medical Question Answering for Clinical Decision SupportabstractThe goal of modern Clinical Decision Support (CDS) systems is to provide physicians with information relevant to their management of patient care. When faced with a medical case, a physician asks questions about the diagnosis, the tests, or treatments that should be administered. Recently, the TREC-CDS track has addressed this challenge by evaluating results of retrieving relevant scientific articles where the answers of medical questions in support of CDS can be found. Although retrieving relevant medical articles instead of identifying the answers was believed to be an easier task, state-of-the-art results are not yet sufficiently promising. In this paper, we present a novel framework for answering medical questions in the spirit of TREC-CDS by first discovering the answer and then selecting and ranking scientific articles that contain the answer. Answer discovery is the result of probabilistic inference which operates on a probabilistic knowledge graph, automatically generated by processing the medical language of large collections of electronic medical records (EMRs). The probabilistic inference of answers combines knowledge from medical practice (EMRs) with knowledge from medical research (scientific articles). It also takes into account the medical knowledge automatically discerned from the medical case description. We show that this novel form of medical question answering (Q/A) produces very promising results in (a) identifying accurately the answers and (b) it improves medical article ranking by 40%. Travis R. Goodwin, Sanda M. Harabagiu |
CIKM | 2 |
| 2016 | Embedding Open-domain Common-sense Knowledge from Text
Travis R. Goodwin, Sanda M. Harabagiu |
LREC | 2 |
| 2014 | Clinical Data-Driven Probabilistic Graph Processing
Travis R. Goodwin, Sanda M. Harabagiu |
LREC | 2 |
| 2014 | Unsupervised Event Coreference ResolutionabstractThe task of event coreference resolution plays a critical role in many natural language processing applications such as information extraction, question answering, and topic detection and tracking. In this article, we describe a new class of unsupervised, nonparametric Bayesian models with the purpose of probabilistically inferring coreference clusters of event mentions from a collection of unlabeled documents. In order to infer these clusters, we automatically extract various lexical, syntactic, and semantic features for each event mention from the document collection. Extracting a rich set of features for each event mention allows us to cast event coreference resolution as the task of grouping together the mentions that share the same features (they have the same participating entities, share the same location, happen at the same time, etc.). Some of the most important challenges posed by the resolution of event coreference in an unsupervised way stem from (a) the choice of representing event mentions through a rich set of features and (b) the ability of modeling events described both within the same document and across multiple documents. Our first unsupervised model that addresses these challenges is a generalization of the hierarchical Dirichlet process. This new extension presents the hierarchical Dirichlet process's ability to capture the uncertainty regarding the number of clustering components and, additionally, takes into account any finite number of features associated with each event mention. Furthermore, to overcome some of the limitations of this extension, we devised a new hybrid model, which combines an infinite latent class model with a discrete time series model. The main advantage of this hybrid model stands in its capability to automatically infer the number of features associated with each event mention from data and, at the same time, to perform an automatic selection of the most informative features for the task of event coreference. The evaluation performed for solving both within- and cross-document event coreference shows significant improvements of these models when compared against two baselines for this task. Cosmin Adrian Bejan, Sanda M. Harabagiu |
Comput. Linguistics | 2 |
| 2013 | A flexible framework for recognizing events, temporal expressions, and temporal relations in clinical textabstractOBJECTIVE: To provide a natural language processing method for the automatic recognition of events, temporal expressions, and temporal relations in clinical records. MATERIALS AND METHODS: A combination of supervised, unsupervised, and rule-based methods were used. Supervised methods include conditional random fields and support vector machines. A flexible automated feature selection technique was used to select the best subset of features for each supervised task. Unsupervised methods include Brown clustering on several corpora, which result in our method being considered semisupervised. RESULTS: On the 2012 Informatics for Integrating Biology and the Bedside (i2b2) shared task data, we achieved an overall event F1-measure of 0.8045, an overall temporal expression F1-measure of 0.6154, an overall temporal link detection F1-measure of 0.5594, and an end-to-end temporal link detection F1-measure of 0.5258. The most competitive system was our event recognition method, which ranked third out of the 14 participants in the event task. DISCUSSION: Analysis reveals the event recognition method has difficulty determining which modifiers to include/exclude in the event span. The temporal expression recognition method requires significantly more normalization rules, although many of these rules apply only to a small number of cases. Finally, the temporal relation recognition method requires more advanced medical knowledge and could be improved by separating the single discourse relation classifier into multiple, more targeted component classifiers. CONCLUSIONS: Recognizing events and temporal expressions can be achieved accurately by combining supervised and unsupervised methods, even when only minimal medical knowledge is available. Temporal normalization and temporal relation recognition, however, are far more dependent on the modeling of medical knowledge. Kirk Roberts, Bryan Rink, Sanda M. Harabagiu |
J. Am. Medical Informatics Assoc. | 3 |
| 2012 | A Machine Learning Approach for Identifying Anatomical Locations of Actionable Findings in Radiology Reports
Kirk Roberts, Bryan Rink, Sanda M. Harabagiu, Richard H. Scheuermann, Seth M. Toomay, Travis Browning, Teresa Bosler, Ronald M. Peshock |
AMIA | 3 |
| 2012 | Locational relativity and domain constraints in spatial questionsabstractSpatial queries in the form of natural language questions have typically been assumed to have unconstrained geographic answers. However, analysis of prototypical spatial questions reveals two important types of constraints that must be considered by spatial question answering systems. First, locational relativity constraints limit answers to a particular location or the user's implied location. Second, domain constraints specify non-geographic locations such as web pages or anatomical sites. In order to detect these constraints, we have conducted a crowd-sourced annotation effort for a set of over 1,200 questions gathered from a community question answering website. We utilize machine learning techniques trained on this data to automatically classify these two types of constraints. We report results nearing 90% accuracy at locational relativity detection and 76% accuracy at domain classification using this approach. Kirk Roberts, Sanda M. Harabagiu |
SIGSPATIAL/GIS | 2 |
| 2012 | Annotating Spatial Containment Relations Between Events
Kirk Roberts, Travis R. Goodwin, Sanda M. Harabagiu |
LREC | 3 |
| 2012 | EmpaTweet: Annotating and Detecting Emotions on Twitter
Kirk Roberts, Michael A. Roach, Josh Guthrie, Sanda M. Harabagiu |
LREC | 5 |
| 2012 | A supervised framework for resolving coreference in clinical recordsabstractOBJECTIVE: A method for the automatic resolution of coreference between medical concepts in clinical records. MATERIALS AND METHODS: A multiple pass sieve approach utilizing support vector machines (SVMs) at each pass was used to resolve coreference. Information such as lexical similarity, recency of a concept mention, synonymy based on Wikipedia redirects, and local lexical context were used to inform the method. Results were evaluated using an unweighted average of MUC, CEAF, and B(3) coreference evaluation metrics. The datasets used in these research experiments were made available through the 2011 i2b2/VA Shared Task on Coreference. RESULTS: The method achieved an average F score of 0.821 on the ODIE dataset, with a precision of 0.802 and a recall of 0.845. These results compare favorably to the best-performing system with a reported F score of 0.827 on the dataset and the median system F score of 0.800 among the eight teams that participated in the 2011 i2b2/VA Shared Task on Coreference. On the i2b2 dataset, the method achieved an average F score of 0.906, with a precision of 0.895 and a recall of 0.918 compared to the best F score of 0.915 and the median of 0.859 among the 16 participating teams. DISCUSSION: Post hoc analysis revealed significant performance degradation on pathology reports. The pathology reports were characterized by complex synonymy and very few patient mentions. CONCLUSION: The use of several simple lexical matching methods had the most impact on achieving competitive performance on the task of coreference resolution. Moreover, the ability to detect patients in electronic medical records helped to improve coreference resolution more than other linguistic analysis. Bryan Rink, Kirk Roberts, Sanda M. Harabagiu |
J. Am. Medical Informatics Assoc. | 3 |
| 2011 | A generative model for unsupervised discovery of relations and argument classes from clinical texts
Bryan Rink, Sanda M. Harabagiu |
EMNLP | 2 |
| 2011 | Unsupervised Learning of Selectional Restrictions and Detection of Argument Coercions
Kirk Roberts, Sanda M. Harabagiu |
EMNLP | 2 |
| 2011 | Relevance Modeling for Microblog Summarization
Sanda M. Harabagiu, Andrew Hickl |
ICWSM | 1 |
| 2011 | Automatic extraction of relations between medical concepts in clinical textsabstractOBJECTIVE: A supervised machine learning approach to discover relations between medical problems, treatments, and tests mentioned in electronic medical records. MATERIALS AND METHODS: A single support vector machine classifier was used to identify relations between concepts and to assign their semantic type. Several resources such as Wikipedia, WordNet, General Inquirer, and a relation similarity metric inform the classifier. RESULTS: The techniques reported in this paper were evaluated in the 2010 i2b2 Challenge and obtained the highest F1 score for the relation extraction task. When gold standard data for concepts and assertions were available, F1 was 73.7, precision was 72.0, and recall was 75.3. F1 is defined as 2*Precision*Recall/(Precision+Recall). Alternatively, when concepts and assertions were discovered automatically, F1 was 48.4, precision was 57.6, and recall was 41.7. DISCUSSION: Although a rich set of features was developed for the classifiers presented in this paper, little knowledge mining was performed from medical ontologies such as those found in UMLS. Future studies should incorporate features extracted from such knowledge sources, which we expect to further improve the results. Moreover, each relation discovery was treated independently. Joint classification of relations may further improve the quality of results. Also, joint learning of the discovery of concepts, assertions, and relations may also improve the results of automatic relation extraction. CONCLUSION: Lexical and contextual features proved to be very important in relation extraction from medical texts. When they are not available to the classifier, the F1 score decreases by 3.7%. In addition, features based on similarity contribute to a decrease of 1.1% when they are not available. Bryan Rink, Sanda M. Harabagiu, Kirk Roberts |
J. Am. Medical Informatics Assoc. | 2 |
| 2011 | A flexible framework for deriving assertions from electronic medical recordsabstractOBJECTIVE: This paper describes natural-language-processing techniques for two tasks: identification of medical concepts in clinical text, and classification of assertions, which indicate the existence, absence, or uncertainty of a medical problem. Because so many resources are available for processing clinical texts, there is interest in developing a framework in which features derived from these resources can be optimally selected for the two tasks of interest. MATERIALS AND METHODS: The authors used two machine-learning (ML) classifiers: support vector machines (SVMs) and conditional random fields (CRFs). Because SVMs and CRFs can operate on a large set of features extracted from both clinical texts and external resources, the authors address the following research question: Which features need to be selected for obtaining optimal results? To this end, the authors devise feature-selection techniques which greatly reduce the amount of manual experimentation and improve performance. RESULTS: The authors evaluated their approaches on the 2010 i2b2/VA challenge data. Concept extraction achieves 79.59 micro F-measure. Assertion classification achieves 93.94 micro F-measure. DISCUSSION: Approaching medical concept extraction and assertion classification through ML-based techniques has the advantage of easily adapting to new data sets and new medical informatics tasks. However, ML-based techniques perform best when optimal features are selected. By devising promising feature-selection techniques, the authors obtain results that outperform the current state of the art. CONCLUSION: This paper presents two ML-based approaches for processing language in the clinical texts evaluated in the 2010 i2b2/VA challenge. By using novel feature-selection methods, the techniques presented in this paper are unique among the i2b2 participants. Kirk Roberts, Sanda M. Harabagiu |
J. Am. Medical Informatics Assoc. | 2 |
| 2010 | Unsupervised Event Coreference Resolution with Rich Linguistic Features
Cosmin Adrian Bejan, Sanda M. Harabagiu |
ACL | 2 |
| 2010 | A Linguistic Resource for Semantic Parsing of Motion Events
Kirk Roberts, Srikanth Gullapalli, Cosmin Adrian Bejan, Sanda M. Harabagiu |
LREC | 4 |
| 2010 | Using topic themes for multi-document summarizationabstractThe problem of using topic representations for multidocument summarization (MDS) has received considerable attention recently. Several topic representations have been employed for producing informative and coherent summaries. In this article, we describe five previously known topic representations and introduce two novel representations of topics based on topic themes. We present eight different methods of generating multidocument summaries and evaluate each of these methods on a large set of topics used in past DUC workshops. Our evaluation results show a significant improvement in the quality of summaries based on topic themes over MDS methods that use other alternative topic representations. Sanda M. Harabagiu, V. Finley Lacatusu |
ACM Trans. Inf. Syst. | 1 |
| 2009 | Nonparametric Bayesian Models for Unsupervised Event Coreference ResolutionabstractWe present a sequence of unsupervised, nonparametric Bayesian models for clustering complex linguistic objects. In this approach, we consider a potentially infinite number of features and categorical outcomes. We evaluate these models for the task of within- and cross-document event coreference on two corpora. All the models we investigated show significant improvements when compared against an existing baseline for this task. Cosmin Adrian Bejan, Matthew Titsworth, Andrew Hickl, Sanda M. Harabagiu |
NIPS | 4 |
| 2008 | Using Clustering Methods for Discovering Event Structures
Cosmin Adrian Bejan, Sanda M. Harabagiu |
AAAI | 2 |
| 2008 | A Linguistic Resource for Discovering Event Structures and Resolving Event Coreference
Cosmin Adrian Bejan, Sanda M. Harabagiu |
LREC | 2 |
| 2007 | Satisfying information needs with multi-document summaries
Sanda M. Harabagiu, Andrew Hickl, V. Finley Lacatusu |
Inf. Process. Manag. | 1 |
| 2006 | Negation, Contrast and Contradiction in Text Processing
Sanda M. Harabagiu, Andrew Hickl, V. Finley Lacatusu |
AAAI | 1 |
| 2006 | Methods for Using Textual Entailment in Open-Domain Question AnsweringabstractWork on the semantics of questions has argued that the relation between a question and its answer(s) can be cast in terms of logical entailment. In this paper, we demonstrate how computational systems designed to recognize textual entailment can be used to enhance the accuracy of current open-domain automatic question answering (Q/A) systems. In our experiments, we show that when textual entailment information is used to either filter or rank answers returned by a Q/A system, accuracy can be increased by as much as 20% overall. Sanda M. Harabagiu, Andrew Hickl |
ACL | 1 |
| 2006 | FERRET: Interactive Question-Answering for Real-World EnvironmentsabstractThis paper describes FERRET, an interactive question-answering (Q/A) system designed to address the challenges of integrating automatic Q/A applications into real-world environments. FERRET utilizes a novel approach to Q/A - known as predictive questioning - which attempts to identify the questions (and answers) that users need by analyzing how a user interacts with a system while gathering information related to a particular scenario. Andrew Hickl, Patrick Wang 0001, John Lehmann, Sanda M. Harabagiu |
ACL | 4 |
| 2006 | An Answer Bank for Temporal Inference
Sanda M. Harabagiu, Cosmin Adrian Bejan |
LREC | 1 |
| 2006 | Impact of Question Decomposition on the Quality of Answer Summaries
V. Finley Lacatusu, Andrew Hickl, Sanda M. Harabagiu |
LREC | 3 |
| 2006 | Answering complex questions with random walk modelsabstractWe present a novel framework for answering complex questions that relies on question decomposition. Complex questions are decomposed by a procedure that operates on a Markov chain, by following a random walk on a bipartite graph of relations established between concepts related to the topic of a complex question and subquestions derived from topic-relevant passages that manifest these relations. Decomposed questions discovered during this random walk are then submitted to a state-of-the-art Question Answering (Q/A) system in order to retrieve a set of passages that can later be merged into a comprehensive answer by a Multi-Document Summarization (MDS) system. In our evaluations, we show that access to the decompositions generated using this method can significantly enhance the relevance and comprehensiveness of summary-length answers to complex questions. Sanda M. Harabagiu, V. Finley Lacatusu, Andrew Hickl |
SIGIR | 1 |
| 2005 | Experiments with Interactive Question-AnsweringabstractThis paper describes a novel framework for interactive question-answering (Q/A) based on predictive questioning. Generated off-line from topic representations of complex scenarios, predictive questions represent requests for information that capture the most salient (and diverse) aspects of a topic. We present experimental results from large user studies (featuring a fully-implemented interactive Q/A system named FERRET) that demonstrates that surprising performance is achieved by integrating predictive questions into the context of a Q/A dialogue. Sanda M. Harabagiu, Andrew Hickl, John Lehmann, Dan I. Moldovan |
ACL | 1 |
| 2005 | Shallow Semantics for Relation Extraction
Sanda M. Harabagiu, Cosmin Adrian Bejan, Paul Morarescu |
IJCAI | 1 |
| 2005 | Temporal Context Representation and Reasoning
Dan I. Moldovan, Christine Clark, Sanda M. Harabagiu |
IJCAI | 3 |
| 2005 | Topic themes for multi-document summarizationabstractThe problem of using topic representations for multi-document summarization (MDS) has received considerable attention recently. In this paper, we describe five different topic representations and introduce a novel representation of topics based on topic themes. We present eight different methods of generating MDS and evaluate each of these methods on a large set of topics used in past DUC workshops. Our evaluation results show a significant improvement in the quality of summaries based on topic themes over MDS methods that use other alternative topic representations. Sanda M. Harabagiu, V. Finley Lacatusu |
SIGIR | 1 |
| 2004 | Incremental Topic Representations
Sanda M. Harabagiu |
COLING | 1 |
| 2004 | Question Answering Based on Semantic Structures
Srini Narayanan, Sanda M. Harabagiu |
COLING | 2 |
| 2004 | Multi-Document Summarization Using Multiple-Sequence Alignment
V. Finley Lacatusu, Steven J. Maiorano, Sanda M. Harabagiu |
LREC | 3 |
| 2004 | NameNet: a Self-Improving Resource for Name Classification
Paul Morarescu, Sanda M. Harabagiu |
LREC | 2 |
| 2003 | Using Predicate-Argument Structures for Information ExtractionabstractIn this paper we present a novel, customizable IE paradigm that takes advantage of predicate-argument structures.We also introduce a new way of automatically identifying predicate argument structures, which is central to our IE paradigm.It is based on: (1) an extended set of features; and (2) inductive decision tree learning.The experimental results prove our claim that accurate predicate-argument structures enable high quality IE results. Mihai Surdeanu, Sanda M. Harabagiu, John Williams 0001, Paul Aarseth |
ACL | 2 |
| 2003 | COGEX: A Logic Prover for Question Answering
Dan I. Moldovan, Christine Clark, Sanda M. Harabagiu, Steven J. Maiorano |
HLT-NAACL | 3 |
| 2003 | Open-domain textual question answering techniquesabstractTextual question answering is a technique of extracting a sentence or text snippet from a document or document collection that responds directly to a query. Open-domain textual question answering presupposes that questions are natural and unrestricted with respect to topic. The question answering (Q/A) techniques, as embodied in today's systems, can be roughly divided into two types: (1) techniques for Information Seeking (IS), which localize the answer in vast document collections; and (2) techniques for Reading Comprehension (RC) that answer a series of questions related to a given document. Although these two types of techniques and systems are different, it is desirable to combine them for enabling more advanced forms of Q/A. This paper discusses an approach that successfully enhanced an existing IS system with RC capabilities. This enhancement is important because advanced Q/A, as exemplified by the ARDA AQUAINT program, is moving towards Q/A systems that incorporate semantic and pragmatic knowledge enabling dialogue-based Q/A. Because today's RC systems involve a short series of questions in context, they represent a rudimentary form of interactive Q/A which constitutes a possible foundation for more advanced forms of dialogue-based Q/A. Sanda M. Harabagiu, Steven J. Maiorano, Marius Pasca |
Nat. Lang. Eng. | 1 |
| 2003 | Performance issues and error analysis in an open-domain question answering systemabstractThis paper presents an in-depth analysis of a state-of-the-art Question Answering system. Several scenarios are examined: (1) the performance of each module in a serial baseline system, (2) the impact of feedbacks and the insertion of a logic prover, and (3) the impact of various retrieval strategies and lexical resources. The main conclusion is that the overall performance depends on the depth of natural language processing resources and the tools used for answer finding. Dan I. Moldovan, Marius Pasca, Sanda M. Harabagiu, Mihai Surdeanu |
ACM Trans. Inf. Syst. | 3 |
| 2002 | Performance Issues and Error Analysis in an Open-Domain Question Answering SystemabstractThis paper presents an in-depth analysis of a state-of-the-art Question Answering system. Several scenarios are examined: (1) the performance of each module in a serial baseline system, (2) the impact of feedbacks and the insertion of a logic prover, and (3) the impact of various lexical resources. The main conclusion is that the overall performance depends on the depth of natural language processing resources and the tools used for answer finding. Dan I. Moldovan, Marius Pasca, Sanda M. Harabagiu, Mihai Surdeanu |
ACL | 3 |
| 2002 | Open-Domain Voice-Activated Question Answering
Sanda M. Harabagiu, Dan I. Moldovan, Joseph Picone |
COLING | 1 |
| 2002 | Multidocument Summarization with GISTexter
Sanda M. Harabagiu, V. Finley Lacatusu, Paul Morarescu |
LREC | 1 |
| 2002 | Performance Analysis of a Distributed Question/Answering SystemabstractThe problem of question/answering (Q/A) is to find answers to open-domain questions by searching large collections of documents. Unlike information retrieval systems very common today in the form of Internet search engines, Q/A systems do not retrieve documents, but instead provide short, relevant answers located in small fragments of text. This enhanced functionality comes with a price: Q/A systems are significantly slower and require more hardware resources than information retrieval systems. This paper proposes a distributed Q/A architecture that enhances the system throughput through the exploitation of interquestion parallelism and dynamic load balancing and reduces the individual question response time through the exploitation of intraquestion parallelism. Inter and intraquestion parallelism are both exploited using several scheduling points: one before the Q/A task is started and two embedded in the Q/A task. An analytical performance model is introduced. The model analyzes both the interquestion parallelism overhead generated by the migration of questions and the intraquestion parallelism overhead generated by the partitioning of the Q/A task. The analytical model indicates that both question migration and partitioning are required for a high-performance system. Mihai Surdeanu, Dan I. Moldovan, Sanda M. Harabagiu |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2001 | The Role of Lexico-Semantic Feedback in Open-Domain Textual Question-AnsweringabstractThis paper presents an open-domain textual Question-Answering system that uses several feedback loops to enhance its performance. These feedback loops combine in a new way statistical results with syntactic, semantic or pragmatic information derived from texts and lexical databases. The paper presents the contribution of each feedback loop to the overall performance of 76% human-assessed precise answers. Sanda M. Harabagiu, Dan I. Moldovan, Marius Pasca, Rada Mihalcea, Mihai Surdeanu, Razvan C. Bunescu, Roxana Girju, Vasile Rus, Paul Morarescu |
ACL | 1 |
| 2001 | COREFDRAW-A Tool for Annotation and Visualization of Coreference DataabstractIn Natural Language Processing, coreference resolution involves finding antecedents of referential expressions (e.g. pronouns or some definite nominals). The resolution of coreference depends on a combination of salience, syntactic, semantic and discourse constraints. The acquisition of such knowledge is difficult and could certainly benefit from a visualization tool, enabling the linguist researcher to find examples of coreference relations. Furthermore, an alternative knowledge-minimalist technique for resolving coreference can be developed by relying on text corpora annotated with coreference data. In this paper we present COREFDRAW, a tool that enables both the annotation of coreference data and its visualization from large text corpora. The resulting annotations enable enhanced coreference resolution methods. Sanda M. Harabagiu, Razvan C. Bunescu, Stefan Trausan-Matu |
ICTAI | 1 |
| 2001 | Performance Analysis of a Distributed Question/Answering SystemabstractThe problem of question/answering (Q/A) is to find answers to open-domain questions by searching a large collection of documents. Unlike Internet search engines, Q/A systems provide short, relevant answers to questions. Due to the complex natural language processing involved that is CPU intensive, and the retrieval of large number of documents that is disk intensive, the time performance of sequential Q/A systems is rather slow. This paper presents the design and performance analysis of a distributed state-of-the-art Q/A system. The design is modular and parallelism is dynamically exploited at inter and intra-question levels. Several schedule points are used to balance the load. An analytical performance model is given backed up by experimental results. Mihai Surdeanu, Dan I. Moldovan, Sanda M. Harabagiu |
IPDPS | 3 |
| 2001 | Text and Knowledge Mining for Coreference Resolution
Sanda M. Harabagiu, Razvan C. Bunescu, Steven J. Maiorano |
NAACL | 1 |
| 2001 | High Performance Question/AnsweringabstractIn this paper we present the features of a Question/Answering (Q/A) system that had unparalleled performance in the TREC-9 evaluations. We explain the accuracy of our system through the unique characteristics of its architecture: (1) usage of a wide-coverage answer type taxonomy; (2) repeated passage retrieval; (3) lexico-semantic feedback loops; (4) extraction of the answers based on machine learning techniques; and (5) answer caching. Experimental results show the effects of each feature on the overall performance of the Q/A system and lead to general conclusions about Q/A from large text collections. Marius Pasca, Sanda M. Harabagiu |
SIGIR | 2 |
| 2000 | The Structure and Performance of an Open-Domain Question Answering SystemabstractThis paper presents the architecture, operation and results obtained with the LASSO Question Answering system developed in the Natural Language Processing Laboratory at SMU. To find answers, the system relies on a combination of syntactic and semantic techniques. The search for the answer is based on a novel form of indexing called paragraph indexing. A score of 55.5% for short answers and 64.5% for long answers was achieved at the TREC-8 competition. Dan I. Moldovan, Sanda M. Harabagiu, Marius Pasca, Rada Mihalcea, Roxana Girju, Richard Goodrum, Vasile Rus |
ACL | 2 |
| 2000 | Experiments with Open-Domain Textual Question Answering
Sanda M. Harabagiu, Marius Pasca, Steven J. Maiorano |
COLING | 1 |
| 2000 | Acquisition of Linguistic Patterns for Knowledge-based Information Extraction
Sanda M. Harabagiu, Steven J. Maiorano |
LREC | 1 |
| 2000 | Patterns of Prepositional Attachments - Where Dictionary Semantics Meets Corpus StatisticsabstractThis paper presents a novel methodology of disambiguating prepositional phrase attachments. We create patterns of attachments by classifying a collection of prepositional relations derived from Treebank parses. As a by-product, the arguments of every prepositional relation are semantically disambiguated. Attachment decisions are generated as the result of a learning process, that builds upon some of the most popular current statistical and machine learning techniques. We have tested this methodology on (1) Wall Street Journal articles, (2) textual definitions of concepts from a dictionary and (3) an ad hoc corpus of Web documents, used for conceptual indexing and information extraction. Sanda M. Harabagiu |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 1999 | From Lexical Cohesion to Textual Coherence: A Data Driven PerspectiveabstractThis paper presents research that connects the cohesion structure of a text to the derivation of its coherence structure. Two different algorithms that derive the cohesion structure in the form of lexical paths from large thesauri are illustrated. Their results are correlated with (1) cue phrases of discourse usage and (2) coherence constraints empirically derived. A novel model of coherence structure is devised, based on the data provided by lexical paths from real world texts. Sanda M. Harabagiu |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 1998 | Parallel System for Text Inference Using Marker PropagationsabstractThis paper presents a possible solution for the text inference problem-extracting information unstated in a text, but implied. Text inference is central to natural language applications such as information extraction and dissemination, text understanding, summarization, and translation. Our solution takes advantage of a semantic English dictionary available in electronic form that provides the basis for the development of a large linguistic knowledge base. The inference algorithm consists of a set of highly parallel search methods that, when applied to the knowledge base, find contexts in which sentences are interpreted. These contexts reveal information relevant to the text. Implementation, results, and parallelism analysis are discussed. Sanda M. Harabagiu, Dan I. Moldovan |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 1997 | TextNet - A text-based intelligent system
Sanda M. Harabagiu, Dan I. Moldovan |
Nat. Lang. Eng. | 1 |
| 1996 | An Application of WordNet to Prepositional AttachmentabstractThis paper presents a method for word sense disambiguation and coherence understanding of prepositional relations.The method relies on information provided by WordNet 1.5.We first classify prepositional attachments according to semantic equivalence of phrase heads and then apply inferential heuristics for understanding the validity of prepositional structures. Sanda M. Harabagiu |
ACL | 1 |
| 1996 | PARIS: A Parallel Inference SystemabstractThis paper presents an inferential system based on abductive interpretation of text. Inference to the best explanation is performed by the recognition of the most economic semantic paths produced by the propagation of markers on a very large linguistic knowledge base. The propagation of markers is controlled by their intrinsic propagation rules, devised from plausible semantic relation chains. An interpretation is inferred whenever two markers collide. Using a very large knowledge base, our inferential system aims at producing interpretations accountable for common sense reasoning. The novelty is that the inference rules model a large variety of implications, as suggested by the knowledge base relations. Textual implicatures are recognized as pragmatic inferences. Sanda M. Harabagiu, Dan I. Moldovan |
ICTAI | 1 |