Mihai Surdeanu

dblp:18/3479 · DBLP profile ↗
← Back
93ranked-venue papers
19as first author
25since 2021 · last 2026
0000-0001-6956-8030ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 78 · 15 first-author · 23 since 2021Databases, data management, data science and information retrieval · 9 · 1 first-author · 2 since 2021Systems, architecture and hardware · 6 · 4 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-authorSecurity and privacy · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Structured Semantic Information Helps Retrieve Better Examples for In-Context Learning Applied to Few-Shot Relation Extraction
abstract
This paper presents several strategies to automatically obtain additional examples for incontext learning, effectively transforming relation extraction from a 1-shot to a few-shot setting.Specifically, we introduce a novel strategy for example selection, in which new examples are selected based on the similarity of their underlying syntactic-semantic structure to the provided 1-shot example.We show that our strategy results in complementary word choices and sentence structures compared to LLM-generated examples.When both strategies are combined, the resulting hybrid system achieves a more holistic picture of the relations of interest than either method alone.Our framework transfers well across datasets (FS-TACRED and FS-FewRel) and LLM families (Qwen and Gemma).Overall, our hybrid system consistently outperforms alternative strategies achieving state-of-the-art performance on FS-TACRED and strong gains on a customized FewRel subset.
Aunabil Chakma, Mihai Surdeanu, Eduardo Blanco 0002
ACL (1)2
2026 A Lightweight Explainable Guardrail for Prompt Safety
abstract
We propose a lightweight explainable guardrail (LEG) method to detect unsafe prompts.LEG uses a multi-task learning architecture to jointly learn a prompt classifier and an explanation classifier, where the latter labels prompt words that explain the safe/unsafe overall decision.LEG is trained on synthetic explanation data, which is generated using a novel strategy that counteracts the confirmation biases of LLMs.Lastly, LEG's training process uses a novel loss that captures global explanation signals as a weak supervision and combines crossentropy and focal losses with uncertainty-based weighting.LEG obtains equivalent or better performance than the state-of-the-art for both prompt classification and explainability, both in-domain and out-of-domain on three datasets, despite the fact that its model size is considerably smaller than current approaches.Code 1 Models and Datasets 2
Md. Asiful Islam, Mihai Surdeanu
ACL (1)2
2026 Towards Complex Debate Understanding: Predicting Claim Impact Scores through the Modelling of Claim Interactions
Maxime Brouat, Mihai Surdeanu, Srdjan Vesic, Eduardo Blanco 0002
LREC2
2025 CopySpec: Accelerating LLMs with Speculative Copy-and-Paste
abstract
We introduce CopySpec, a simple yet effective technique to tackle the inefficiencies LLMs face when generating responses that closely resemble previous outputs or responses that can be verbatim extracted from context.CopySpec identifies repeated sequences in the model's chat history or context and speculates that the same tokens will follow, enabling seamless copying without compromising output quality and without requiring additional GPU memory.To evaluate the effectiveness of our approach, we conducted experiments using seven LLMs and five datasets: MT-Bench, CNN/DM, GSM8K, HumanEval, and our newly created dataset, MT-Redundant.MT-Redundant, introduced in this paper, transforms the second turn of MT-Bench into a request for variations of the first turn's answer, simulating realworld scenarios where users request modifications to prior responses.Our results demonstrate significant speed-ups: up to 2.35× on CNN/DM, 3.08× on the second turn of select MT-Redundant categories, and 2.66× on the third turn of GSM8K's self-correction tasks.Importantly, we show that CopySpec integrates seamlessly with speculative decoding, yielding an average 49% additional speedup over speculative decoding for the second turn of MT-Redundant across all eight categories.While LLMs, even with speculative decoding, suffer from slower inference as context size grows, CopySpec leverages larger contexts to accelerate inference, making it a faster complementary solution.Our code and dataset are publicly available at https://github.com/RazvanDu/CopySpec.
Razvan-Gabriel Dumitru, Minglai Yang 0002, Vikas Yadav, Mihai Surdeanu
EMNLP4
2025 How Is LLM Reasoning Distracted by Irrelevant Context? An Analysis Using a Controlled Benchmark
abstract
We introduce Grade School Math with Distracting Context (GSM-DC 1 ), a synthetic benchmark to evaluate Large Language Models' (LLMs) reasoning robustness against systematically controlled irrelevant context (IC).GSM-DC constructs symbolic reasoning graphs with precise distractor injections, enabling rigorous, reproducible evaluation.Our experiments demonstrate that LLMs are significantly sensitive to IC, affecting both reasoning path selection and arithmetic accuracy.Additionally, training models with strong distractors improves performance in both in-distribution and out-of-distribution scenarios.We further propose a stepwise tree search guided by a process reward model, which notably enhances robustness in out-of-distribution conditions.
Minglai Yang 0002, Ethan Huang, Mihai Surdeanu, William Yang Wang, Liangming Pan
EMNLP4
2025 Data Contamination Quiz: A Tool to Detect and Estimate Contamination in Large Language Models
abstract
Abstract We propose the Data Contamination Quiz (DCQ), a simple and effective approach to detect data contamination in large language models (LLMs) and estimate the amount of it. Specifically, we frame data contamination detection as a series of multiple-choice questions, devising a quiz format wherein three perturbed versions of each instance, subsampled from a specific dataset partition, are created. These changes only include word-level perturbations. The generated perturbations, along with the original dataset instance, form the options in the DCQ, with an extra option accommodating the selection of none of the provided options. Given that the only distinguishing signal among the options is the exact wording with respect to the original dataset instance, an LLM, when tasked with identifying the original dataset instance, gravitates towards selecting the original one if it has been exposed to it. While accounting for positional biases in LLMs, the quiz performance reveals the contamination level for the tested model with the dataset partition to which the quiz pertains. Applied to various datasets and LLMs, under controlled and uncontrolled contamination, our findings—while fully lacking access to training data and model parameters—suggest that DCQ achieves state-of-the-art results and uncovers greater contamination levels through memorization compared to existing methods. Also, it proficiently bypasses more safety filters, especially those set to avoid generating copyrighted content.1
Shahriar Golchin, Mihai Surdeanu
Trans. Assoc. Comput. Linguistics2
2024 Towards Realistic Few-Shot Relation Extraction: A New Meta Dataset and Evaluation
abstract
We introduce a meta dataset for few-shot relation extraction, which includes two datasets derived from existing supervised relation extraction datasets – NYT29 (Takanobu et al., 2019; Nayak and Ng, 2020) and WIKI- DATA (Sorokin and Gurevych, 2017) – as well as a few-shot form of the TACRED dataset (Sabo et al., 2021). Importantly, all these few-shot datasets were generated under realistic assumptions such as: the test relations are different from any relations a model might have seen before, limited training data, and a preponderance of candidate relation mentions that do not correspond to any of the relations of interest. Using this large resource, we conduct a comprehensive evaluation of six recent few-shot relation extraction methods, and observe that no method comes out as a clear winner. Further, the overall performance on this task is low, indicating substantial need for future research. We release all versions of the data, i.e., both supervised and few-shot, for future research.
Fahmida Alam, Md. Asiful Islam, Robert Vacareanu, Mihai Surdeanu
LREC/COLING4
2024 ELLEN: Extremely Lightly Supervised Learning for Efficient Named Entity Recognition
abstract
In this work, we revisit the problem of semi-supervised named entity recognition (NER) focusing on extremely light supervision, consisting of a lexicon containing only 10 examples per class. We introduce ELLEN, a simple, fully modular, neuro-symbolic method that blends fine-tuned language models with linguistic rules. These rules include insights such as “One Sense Per Discourse”, using a Masked Language Model as an unsupervised NER, leveraging part-of-speech tags to identify and eliminate unlabeled entities as false negatives, and other intuitions about classifier confidence scores in local and global context. ELLEN achieves very strong performance on the CoNLL-2003 dataset when using the minimal supervision from the lexicon above. It also outperforms most existing (and considerably more complex) semi-supervised NER methods under the same supervision settings commonly used in the literature (i.e., 5% of the training data). Further, we evaluate our CoNLL-2003 model in a zero-shot scenario on WNUT-17 where we find that it outperforms GPT-3.5 and achieves comparable performance to GPT-4. In a zero-shot setting, ELLEN also achieves over 75% of the performance of a strong, fully supervised model trained on gold data. Our code is publicly available.
Haris Riaz, Razvan-Gabriel Dumitru, Mihai Surdeanu
LREC/COLING3
2024 Active Learning Design Choices for NER with Transformers
abstract
We explore multiple important choices that have not been analyzed in conjunction regarding active learning for token classification using transformer networks. These choices are: (i) how to select what to annotate, (ii) decide whether to annotate entire sentences or smaller sentence fragments, (iii) how to train with incomplete annotations at token-level, and (iv) how to select the initial seed dataset. We explore whether annotating at sub-sentence level can translate to an improved downstream performance by considering two different sub-sentence annotation strategies: (i) entity-level, and (ii) token-level. These approaches result in some sentences being only partially annotated. To address this issue, we introduce and evaluate multiple strategies to deal with partially-annotated sentences during the training process. We show that annotating at the sub-sentence level achieves comparable or better performance than sentence-level annotations with a smaller number of annotated tokens. We then explore the extent to which the performance gap remains once accounting for the annotation time and found that both annotation schemes perform similarly.
Robert Vacareanu, Enrique Noriega-Atala, Gus Hahn-Powell, Marco Antonio Valenzuela-Escárcega, Mihai Surdeanu
LREC/COLING5
2024 On Learning Bipolar Gradual Argumentation Semantics with Neural Networks
abstract
International audience
Caren Al Anaissy, Sandeep Suntwal, Mihai Surdeanu, Srdjan Vesic
ICAART (2)3
2024 Time Travel in LLMs: Tracing Data Contamination in Large Language Models
abstract
Data contamination, i.e., the presence of test data from downstream tasks in the training data of large language models (LLMs), is a potential major issue in measuring LLMs' real effectiveness on other tasks. We propose a straightforward yet effective method for identifying data contamination within LLMs. At its core, our approach starts by identifying potential contamination at the instance level; using this information, our approach then assesses wider contamination at the partition level. To estimate contamination of individual instances, we employ "guided instruction:" a prompt consisting of the dataset name, partition type, and the random-length initial segment of a reference instance, asking the LLM to complete it. An instance is flagged as contaminated if the LLM's output either exactly or nearly matches the latter segment of the reference. To understand if an entire partition is contaminated, we propose two ideas. The first idea marks a dataset partition as contaminated if the average overlap score with the reference instances (as measured by ROUGE-L or BLEURT) is statistically significantly better with the completions from guided instruction compared to a "general instruction" that does not include the dataset and partition name. The second idea marks a dataset partition as contaminated if a classifier based on GPT-4 with few-shot in-context learning prompt marks multiple generated completions as exact/near-exact matches of the corresponding reference instances. Our best method achieves an accuracy between 92% and 100% in detecting if an LLM is contaminated with seven datasets, containing train and test/validation partitions, when contrasted with manual evaluation by human experts. Further, our findings indicate that GPT-4 is contaminated with AG News, WNLI, and XSum datasets.
Shahriar Golchin, Mihai Surdeanu
ICLR2
2023 NEUROSTRUCTURAL DECODING: Neural Text Generation with Structural Constraints
abstract
Text generation often involves producing texts that also satisfy a given set of semantic constraints.While most approaches for conditional text generation have primarily focused on lexical constraints, they often struggle to effectively incorporate syntactic constraints, which provide a richer language for approximating semantic constraints.We address this gap by introducing NEUROSTRUCTURAL DECODING, a new decoding algorithm that incorporates syntactic constraints to further improve the quality of the generated text.We build NEUROSTRUC-TURAL DECODING on the NeuroLogic Decoding (Lu et al., 2021b) algorithm, which enables language generation models to produce fluent text while satisfying complex lexical constraints.Our algorithm is powerful and scalable.It tracks lexico-syntactic constraints (e.g., we need to observe dog as subject and ball as object) during decoding by parsing the partial generations at each step.To this end, we adapt a dependency parser to generate parses for incomplete sentences.Our approach is evaluated on three different language generation tasks, and the results show improved performance in both lexical and syntactic metrics compared to previous methods.The results suggest this is a promising solution for integrating fine-grained controllable text generation into the conventional beam search decoding 1 .
Mohadeseh Bastan, Mihai Surdeanu, Niranjan Balasubramanian
ACL (1)2
2023 It Takes Two Flints to Make a Fire: Multitask Learning of Neural Relation and Explanation Classifiers
abstract
Abstract We propose an explainable approach for relation extraction that mitigates the tension between generalization and explainability by jointly training for the two goals. Our approach uses a multi-task learning architecture, which jointly trains a classifier for relation extraction, and a sequence model that labels words in the context of the relations that explain the decisions of the relation classifier. We also convert the model outputs to rules to bring global explanations to this approach. This sequence model is trained using a hybrid strategy: supervised, when supervision from pre-existing patterns is available, and semi-supervised otherwise. In the latter situation, we treat the sequence model’s labels as latent variables, and learn the best assignment that maximizes the performance of the relation classifier. We evaluate the proposed approach on the two datasets and show that the sequence model provides labels that serve as accurate explanations for the relation classifier’s decisions, and, importantly, that the joint training generally improves the performance of the relation classifier. We also evaluate the performance of the generated rules and show that the new rules are a great add-on to the manual rules and bring the rule-based system much closer to the neural models.
Mihai Surdeanu
Comput. Linguistics2
2022 SuMe: A Dataset Towards Summarizing Biomedical Mechanisms
abstract
Can language models read biomedical texts and explain the biomedical mechanisms discussed? In this work we introduce a biomedical mechanism summarization task. Biomedical studies often investigate the mechanisms behind how one entity (e.g., a protein or a chemical) affects another in a biological context. The abstracts of these publications often include a focused set of sentences that present relevant supporting statements regarding such relationships, associated experimental evidence, and a concluding sentence that summarizes the mechanism underlying the relationship. We leverage this structure and create a summarization task, where the input is a collection of sentences and the main entities in an abstract, and the output includes the relationship and a sentence that summarizes the mechanism. Using a small amount of manually labeled mechanism sentences, we train a mechanism sentence classifier to filter a large biomedical abstract collection and create a summarization dataset with 22k instances. We also introduce conclusion sentence generation as a pretraining task with 611k instances. We benchmark the performance of large bio-domain language models. We find that while the pretraining task help improves performance, the best model produces acceptable mechanism outputs in only 32% of the instances, which shows the task presents significant challenges in biomedical language understanding and summarization.
Mohadeseh Bastan, Nishant Shankar, Mihai Surdeanu, Niranjan Balasubramanian
LREC3
2022 Informal Persian Universal Dependency Treebank
abstract
This paper presents the phonological, morphological, and syntactic distinctions between formal and informal Persian, showing that these two variants have fundamental differences that cannot be attributed solely to pronunciation discrepancies. Given that informal Persian exhibits particular characteristics, any computational model trained on formal Persian is unlikely to transfer well to informal Persian, necessitating the creation of dedicated treebanks for this variety. We thus detail the development of the open-source Informal Persian Universal Dependency Treebank, a new treebank annotated within the Universal Dependencies scheme. We then investigate the parsing of informal Persian by training two dependency parsers on existing formal treebanks and evaluating them on out-of-domain data, i.e. the development set of our informal treebank. Our results show that parsers experience a substantial performance drop when we move across the two domains, as they face more unknown tokens and structures and fail to generalize well. Furthermore, the dependency relations whose performance deteriorates the most represent the unique properties of the informal variant. The ultimate goal of this study that demonstrates a broader impact is to provide a stepping-stone to reveal the significance of informal variants of languages, which have been widely overlooked in natural language processing tools across languages.
Roya Kabiri, Simin Karimi, Mihai Surdeanu
LREC3
2022 A STEP towards Interpretable Multi-Hop Reasoning: Bridge Phrase Identification and Query Expansion
abstract
We propose an unsupervised method for the identification of bridge phrases in multi-hop question answering (QA). Our method constructs a graph of noun phrases from the question and the available context, and applies the Steiner tree algorithm to identify the minimal sub-graph that connects all question phrases. Nodes in the sub-graph that bridge loosely-connected or disjoint subsets of question phrases due to low-strength semantic relations are extracted as bridge phrases. The identified bridge phrases are then used to expand the query based on the initial question, helping in increasing the relevance of evidence that has little lexical overlap or semantic relation with the question. Through an evaluation on HotpotQA, a popular dataset for multi-hop QA, we show that our method yields: (a) improved evidence retrieval, (b) improved QA performance when using the retrieved sentences; and (c) effective and faithful explanations when answers are provided.
Fan Luo 0002, Mihai Surdeanu
LREC2
2022 Do Transformer Networks Improve the Discovery of Rules from Text?
abstract
With their Discovery of Inference Rules from Text (DIRT) algorithm, Lin and Pantel (2001) made a seminal contribution to the field of rule acquisition from text, by adapting the distributional hypothesis of Harris (1954) to rules that model binary relations such as X treat Y. DIRT’s relevance is renewed in today’s neural era given the recent focus on interpretability in the field of natural language processing. We propose a novel take on the DIRT algorithm, where we implement the distributional hypothesis using the contextualized embeddings provided by BERT, a transformer-network-based language model (Vaswani et al. 2017; Devlin et al. 2018). In particular, we change the similarity measure between pairs of slots (i.e., the set of words matched by a rule) from the original formula that relies on lexical items to a formula computed using contextualized embeddings. We empirically demonstrate that this new similarity method yields a better implementation of the distributional hypothesis, and this, in turn, yields rules that outperform the original algorithm in the question answering-based evaluation proposed by Lin and Pantel (2001).
Mahdi Rahimi 0001, Mihai Surdeanu
LREC2
2022 From Examples to Rules: Neural Guided Rule Synthesis for Information Extraction
abstract
While deep learning approaches to information extraction have had many successes, they can be difficult to augment or maintain as needs shift. Rule-based methods, on the other hand, can be more easily modified. However, crafting rules requires expertise in linguistics and the domain of interest, making it infeasible for most users. Here we attempt to combine the advantages of these two directions while mitigating their drawbacks. We adapt recent advances from the adjacent field of program synthesis to information extraction, synthesizing rules from provided examples. We use a transformer-based architecture to guide an enumerative search, and show that this reduces the number of steps that need to be explored before a rule is found. Further, we show that without training the synthesis algorithm on the specific domain, our synthesized rules achieve state-of-the-art performance on the 1-shot scenario of a task that focuses on few-shot learning for relation classification, and competitive performance in the 5-shot scenario.
Robert Vacareanu, Marco Antonio Valenzuela-Escárcega, George Caique Gouveia Barbosa, Rebecca Sharp, Gus Hahn-Powell, Mihai Surdeanu
LREC6
2022 Automatic Correction of Syntactic Dependency Annotation Differences
abstract
Annotation inconsistencies between data sets can cause problems for low-resource NLP, where noisy or inconsistent data cannot be easily replaced. We propose a method for automatically detecting annotation mismatches between dependency parsing corpora, along with three related methods for automatically converting the mismatches. All three methods rely on comparing unseen examples in a new corpus with similar examples in an existing corpus. These three methods include a simple lexical replacement using the most frequent tag of the example in the existing corpus, a GloVe embedding-based replacement that considers related examples, and a BERT-based replacement that uses contextualized embeddings to provide examples fine-tuned to our data. We evaluate these conversions by retraining two dependency parsers—Stanza and Parsing as Tagging (PaT)—on the converted and unconverted data. We find that applying our conversions yields significantly better performance in many cases. Some differences observed between the two parsers are observed. Stanza has a more complex architecture with a quadratic algorithm, taking longer to train, but it can generalize from less data. The PaT parser has a simpler architecture with a linear algorithm, speeding up training but requiring more training data to reach comparable or better performance.
Andrew Zupon, Andrew Carnie, Michael Hammond, Mihai Surdeanu
LREC4
2021 Using the Hammer only on Nails: A Hybrid Method for Representation-Based Evidence Retrieval for Question Answering
Zhengzhong Liang, Yiyun Zhao, Mihai Surdeanu
ECIR (1)3
2021 Students Who Study Together Learn Better: On the Importance of Collective Knowledge Distillation for Domain Transfer in Fact Verification
abstract
While neural networks produce state-of-the-art performance in several NLP tasks, they depend heavily on lexicalized information, which transfers poorly between domains.Previous work (Suntwal et al., 2019) proposed delexicalization as a form of knowledge distillation to reduce dependency on such lexical artifacts.However, a critical unsolved issue that remains is how much delexicalization should be applied?A little helps reduce over-fitting, but too much discards useful information.We propose Group Learning (GL), a knowledge and model distillation approach for fact verification.In our method, while multiple student models have access to different delexicalized data views, they are encouraged to independently learn from each other through pair-wise consistency losses.In several cross-domain experiments between the FEVER and FNC fact verification datasets, we show that our approach learns the best delexicalization strategy for the given training dataset and outperforms state-of-theart classifiers that rely on the original data.
Mithun Paul, Sandeep Suntwal, Mihai Surdeanu
EMNLP (1)3
2021 Explainable Multi-hop Verbal Reasoning Through Internal Monologue
abstract
Many state-of-the-art (SOTA) language models have achieved high accuracy on several multi-hop reasoning problems.However, these approaches tend to not be interpretable because they do not make the intermediate reasoning steps explicit.Moreover, models trained on simpler tasks tend to fail when directly tested on more complex problems.We propose the Explainable multi-hop Verbal Reasoner (EVR) to solve these limitations by (a) decomposing multi-hop reasoning problems into several simple ones, and (b) using natural language to guide the intermediate reasoning hops.We implement EVR by extending the classic reasoning paradigm General Problem Solver (GPS) with a SOTA generative language model to generate subgoals and perform inference in natural language at each reasoning step.Evaluation of EVR on Clark et al. (2020)'s synthetic question answering (QA) dataset shows that EVR achieves SOTA performance while being able to generate all reasoning steps in natural language.Furthermore, EVR generalizes better than other strong methods when trained on simpler tasks or less training data (up to 35.7% and 7.7% absolute improvement respectively). 1
Zhengzhong Liang, Steven Bethard, Mihai Surdeanu
NAACL-HLT3
2021 Data and Model Distillation as a Solution for Domain-transferable Fact Verification
abstract
Mitch Paul Mithun, Sandeep Suntwal, Mihai Surdeanu. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Mithun Paul, Sandeep Suntwal, Mihai Surdeanu
NAACL-HLT3
2021 If You Want to Go Far Go Together: Unsupervised Joint Candidate Evidence Retrieval for Multi-hop Question Answering
abstract
Multi-hop reasoning requires aggregation and inference from multiple facts.To retrieve such facts, we propose a simple approach that retrieves and reranks set of evidence facts jointly.Our approach first generates unsupervised clusters of sentences as candidate evidence by accounting links between sentences and coverage with the given query.Then, a RoBERTa-based reranker is trained to bring the most representative evidence cluster to the top.We specifically emphasize on the importance of retrieving evidence jointly by showing several comparative analyses to other methods that retrieve and rerank evidence sentences individually.First, we introduce several attention-and embedding-based analyses, which indicate that jointly retrieving and reranking approaches can learn compositional knowledge required for multi-hop reasoning.Second, our experiments show that jointly retrieving candidate evidence leads to substantially higher evidence retrieval performance when fed to the same supervised reranker.In particular, our joint retrieval and then reranking approach achieves new state-of-the-art evidence retrieval performance on two multi-hop question answering (QA) datasets: 30.5 Recall@2 on QASC, and 67.6% F1 on MultiRC.When the evidence text from our joint retrieval approach is fed to a RoBERTa-based answer selection classifier, we achieve new state-ofthe-art QA performance on MultiRC and second best result on QASC.
Vikas Yadav, Steven Bethard, Mihai Surdeanu
NAACL-HLT3
2021 Cheap and Good? Simple and Effective Data Augmentation for Low Resource Machine Reading
abstract
We propose a simple and effective strategy for data augmentation for low-resource machine reading comprehension (MRC). Our approach first pretrains the answer extraction components of a MRC system on the augmented data that contains approximate context of the correct answers, before training it on the exact answer spans. The approximate context helps the QA method components in narrowing the location of the answers. We demonstrate that our simple strategy substantially improves both document retrieval and answer extraction performance by providing larger context of the answers and additional training data. In particular, our method significantly improves the performance of BERT based retriever (15.12%), and answer extractor (4.33% F1) on TechQA, a complex, low-resource MRC task. Further, our data augmentation strategy yields significant improvements of up to 3.9% exact match (EM) and 2.7% F1 for answer extraction on PolicyQA, another practical but moderate sized QA dataset that also contains long answer spans.
Hoang Van, Vikas Yadav, Mihai Surdeanu
SIGIR3
2020 Unsupervised Alignment-based Iterative Evidence Retrieval for Multi-hop Question Answering
abstract
Evidence retrieval is a critical stage of question answering (QA), necessary not only to improve performance, but also to explain the decisions of the corresponding QA method.We introduce a simple, fast, and unsupervised iterative evidence retrieval method, which relies on three ideas: (a) an unsupervised alignment approach to soft-align questions and answers with justification sentences using only GloVe embeddings, (b) an iterative process that reformulates queries focusing on terms that are not covered by existing justifications, which (c) a stopping criterion that terminates retrieval when the terms in the given question and candidate answers are covered by the retrieved justifications.Despite its simplicity, our approach outperforms all the previous methods (including supervised methods) on the evidence selection task on two datasets: MultiRC and QASC.When these evidence sentences are fed into a RoBERTa answer classification component, we achieve state-of-the-art QA performance on these two datasets.
Vikas Yadav, Steven Bethard, Mihai Surdeanu
ACL3
2020 An Unsupervised Method for Learning Representations of Multi-word Expressions for Semantic Classification
abstract
This paper explores an unsupervised approach to learning a compositional representation function for multi-word expressions (MWEs), and evaluates it on the Tratz dataset, which associates two-word expressions with the semantic relation between the compound constituents (e.g. the label employer is associated with the noun compound government agency) (Tratz, 2011).The composition function is based on recurrent neural networks, and is trained using the Skip-Gram objective to predict the words in the context of MWEs.Thus our approach can naturally leverage large unlabeled text sources.Further, our method can make use of provided MWEs when available, but can also function as a completely unsupervised algorithm, using MWE boundaries predicted by a single, domain-agnostic part-of-speech pattern.With pre-defined MWE boundaries, our method outperforms the previous state-of-the-art performance on the coarse-grained evaluation of the Tratz dataset (Tratz, 2011), with an F1 score of 50.4%.The unsupervised version of our method approaches the performance of the supervised one, and even outperforms it in some configurations.
Robert Vacareanu, Marco Antonio Valenzuela-Escárcega, Rebecca Sharp, Mihai Surdeanu
COLING4
2020 Towards the Necessity for Debiasing Natural Language Inference Datasets
abstract
Modeling natural language inference is a challenging task. With large annotated data sets available it has now become feasible to train complex neural network based inference methods which achieve state of the art performance. However, it has been shown that these models also learn from the subtle biases inherent in these datasets (CITATION). In this work we explore two techniques for delexicalization that modify the datasets in such a way that we can control the importance that neural-network based methods place on lexical entities. We demonstrate that the proposed methods not only maintain the performance in-domain but also improve performance in some out-of-domain settings. For example, when using the delexicalized version of the FEVER dataset, the in-domain performance of a state of the art neural network method dropped only by 1.12% while its out-of-domain performance on the FNC dataset improved by 4.63%. We release the delexicalized versions of three common datasets used in natural language inference. These datasets are delexicalized using two methods: one which replaces the lexical entities in an overlap-aware manner, and a second, which additionally incorporates semantic lifting of nouns and verbs to their WordNet hypernym synsets
Mithun Paul Panenghat, Sandeep Suntwal, Faiz Rafique, Rebecca Sharp, Mihai Surdeanu
LREC5
2020 Parsing as Tagging
abstract
We propose a simple yet accurate method for dependency parsing that treats parsing as tagging (PaT). That is, our approach addresses the parsing of dependency trees with a sequence model implemented with a bidirectional LSTM over BERT embeddings, where the “tag” to be predicted at each token position is the relative position of the corresponding head. For example, for the sentence John eats cake, the tag to be predicted for the token cake is -1 because its head (eats) occurs one token to the left. Despite its simplicity, our approach performs well. For example, our approach outperforms the state-of-the-art method of (Fernández-González and Gómez-Rodríguez, 2019) on Universal Dependencies (UD) by 1.76% unlabeled attachment score (UAS) for English, 1.98% UAS for French, and 1.16% UAS for German. On average, on 12 UD languages, our method with minimal tuning performs comparably with this state-of-the-art approach: better by 0.11% UAS, and worse by 0.58% LAS.
Robert Vacareanu, George Caique Gouveia Barbosa, Marco Antonio Valenzuela-Escárcega, Mihai Surdeanu
LREC4
2020 Having Your Cake and Eating it Too: Training Neural Retrieval for Language Inference without Losing Lexical Match
abstract
We present a study on the importance of information retrieval (IR) techniques for both the interpretability and the performance of neural question answering (QA) methods. We show that the current state-of-the-art transformer methods (like RoBERTa) encode poorly simple information retrieval (IR) concepts such as lexical overlap between query and the document. To mitigate this limitation, we introduce a supervised RoBERTa QA method that is trained to mimic the behavior of BM25 and the soft-matching idea behind embedding-based alignment methods. We show that fusing the simple lexical-matching IR concepts in transformer techniques results in improvement a) of their (lexical-matching) interpretability, b) retrieval performance, and c) the QA performance on two multi-hop QA datasets. We further highlight the lexical-chasm gap bridging capabilities of transformer methods by analyzing the attention distributions of the supervised RoBERTa classifier over the context versus lexically-matched token pairs.
Vikas Yadav, Steven Bethard, Mihai Surdeanu
SIGIR3
2019 On the Importance of Delexicalization for Fact Verification
abstract
Sandeep Suntwal, Mithun Paul, Rebecca Sharp, Mihai Surdeanu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Sandeep Suntwal, Mithun Paul, Rebecca Sharp, Mihai Surdeanu
EMNLP/IJCNLP (1)4
2019 Quick and (not so) Dirty: Unsupervised Selection of Justification Sentences for Multi-hop Question Answering
abstract
Vikas Yadav, Steven Bethard, Mihai Surdeanu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Vikas Yadav, Steven Bethard, Mihai Surdeanu
EMNLP/IJCNLP (1)3
2018 An Exploration of Three Lightly-supervised Representation Learning Approaches for Named Entity Classification
abstract
Several semi-supervised representation learning methods have been proposed recently that mitigate the drawbacks of traditional bootstrapping: they reduce the amount of semantic drift introduced by iterative approaches through one-shot learning; others address the sparsity of data through the learning of custom, dense representation for the information modeled. In this work, we are the first to adapt three of these methods, most of which have been originally proposed for image processing, to an information extraction task, specifically, named entity classification. Further, we perform a rigorous comparative analysis on two distinct datasets. Our analysis yields several important observations. First, all representation learning methods outperform state-of-the-art semi-supervised methods that do not rely on representation learning. To the best of our knowledge, we report the latest state-of-the-art results on the semi-supervised named entity classification task. Second, one-shot learning methods clearly outperform iterative representation learning approaches. Lastly, one of the best performers relies on the mean teacher framework (Tarvainen and Valpola, 2017), a simple teacher/student approach that is independent of the underlying task-specific model.
Ajay Nagesh, Mihai Surdeanu
COLING2
2018 Controlling Information Aggregation for Complex Question Answering
Heeyoung Kwon, Harsh Trivedi, Peter A. Jansen, Mihai Surdeanu, Niranjan Balasubramanian
ECIR4
2018 Visual Supervision in Bootstrapped Information Extraction
abstract
We challenge a common assumption in active learning, that a list-based interface populated by informative samples provides for efficient and effective data annotation.We show how a 2D scatterplot populated with diverse and representative samples can yield improved models given the same time budget.We consider this for bootstrapping-based information extraction, in particular named entity classification, where human and machine jointly label data.To enable effective data annotation in a scatterplot, we have developed an embeddingbased bootstrapping model that learns the distributional similarity of entities through the patterns that match them in a large data corpus, while being discriminative with respect to human-labeled and machine-promoted entities.We conducted a user study to assess the effectiveness of these different interfaces, and analyze bootstrapping performance in terms of human labeling accuracy, label quantity, and labeling consensus across multiple users.Our results suggest that supervision acquired from the scatterplot interface, despite being noisier, yields improvements in classification performance compared with the list interface, due to a larger quantity of supervision acquired.
Matthew Berger, Ajay Nagesh, Joshua A. Levine, Mihai Surdeanu, Helen Zhang
EMNLP4
2018 Detecting Cyber Threats in Non-English Dark Net Markets: A Cross-Lingual Transfer Learning Approach
abstract
Recent advances in proactive cyber threat intelligence rely on early detection of cyber threats in hacker communities. Dark Net Markets (DNMs) are growing platforms in hacker community that provide hackers with highly- specialized tools and products which may not be found in other platforms. While text classification techniques have been used for cyber threat detection in English DNMs, the task is hindered in non-English platforms due to the language barrier and lack of ground-truth data. Current approaches use monolingual models on machine translated data to overcome these challenges. However, the translation errors can deteriorate the classification results. The abundance of data in English DNMs can be leveraged in learning non-English threats without using machine translation. In this study, we show that a deep cross-lingual model that can jointly learn the common language representation from two languages, significantly outperforms a monolingual model learned on machine translated data for identifying cyber threats in non-English DNMs. Unlike most studies, our approach does not require any external data source such as bilingual word embeddings or bilingual lexicons. Our experiments on Russian DNMs show that this approach can achieve better performance than state-of-the-art methods for non-English cyber threat detection in malicious hacker community.
Reza Ebrahimi 0001, Mihai Surdeanu, Sagar Samtani, Hsinchun Chen
ISI2
2018 Text Annotation Graphs: Annotating Complex Natural Language Phenomena
Angus G. Forbes, Kristine Lee, Gus Hahn-Powell, Marco Antonio Valenzuela-Escárcega, Mihai Surdeanu
LREC5
2018 Bootstrapping Polar-Opposite Emotion Dimensions from Online Reviews
Luwen Huangfu, Mihai Surdeanu
LREC2
2018 Grounding Gradable Adjectives through Crowdsourcing
Rebecca Sharp, Mithun Paul, Ajay Nagesh, Dane Bell, Mihai Surdeanu
LREC5
2018 MLStar: Machine Learning in Energy Profile Estimation of Android Apps
abstract
Improving the energy efficiency of smartphones is critical for increasing the utility that they provide to the users. With most mobile operating systems, users are responsible for managing their phone's battery efficiency by utilizing the various settings provided by the operating system, as well as selecting energy-efficient apps. However, current app marketplaces do not provide users with information about app energy efficiency, which makes it challenging for the user to make informed decision when selecting an app. This paper presents a novel machine learning approach to estimate app energy efficiency by utilizing textual information available in the Google Play store such as an app's description, user reviews, as well as system permissions. Our detailed analysis of the resulting system shows that hardware permissions, app description, and user reviews correlate well with energy efficiency ratings. We evaluate five models that represent popular classes of machine learning algorithms in their ability to predict energy efficiency ratings. Finally, we compare our approach to gold truth ratings obtained by the actual energy profiling of the app, demonstrating that the proposed system is able to estimate an app's energy efficiency within less than 1 point on the 1-5 scale provided by the profiler, without requiring any kind of profiling.
Benjamin Gaska, Chris Gniady, Mihai Surdeanu
MobiQuitous3
2018 Sanity Check: A Strong Alignment and Information Retrieval Baseline for Question Answering
abstract
While increasingly complex approaches to question answering (QA) have been proposed, the true gain of these systems, particularly with respect to their expensive training requirements, can be in- flated when they are not compared to adequate baselines. Here we propose an unsupervised, simple, and fast alignment and informa- tion retrieval baseline that incorporates two novel contributions: a one-to-many alignment between query and document terms and negative alignment as a proxy for discriminative information. Our approach not only outperforms all conventional baselines as well as many supervised recurrent neural networks, but also approaches the state of the art for supervised systems on three QA datasets. With only three hyperparameters, we achieve 47% [email protected] on an 8th grade Science QA dataset, 32.9% [email protected] on a Yahoo! answers QA dataset and 64% MAP on WikiQA.
Vikas Yadav, Rebecca Sharp, Mihai Surdeanu
SIGIR3
2017 Tell Me Why: Using Question Answering as Distant Supervision for Answer Justification
abstract
Rebecca Sharp, Mihai Surdeanu, Peter Jansen, Marco A. Valenzuela-Escárcega, Peter Clark, Michael Hammond. Proceedings of the 21st Conference on Computational Natural Language Learning (CoNLL 2017). 2017.
Rebecca Sharp, Mihai Surdeanu, Peter A. Jansen, Marco Antonio Valenzuela-Escárcega, Peter Clark, Michael Hammond
CoNLL2
2017 Learning what to read: Focused machine reading
abstract
Recent efforts in bioinformatics have achieved tremendous progress in the machine reading of biomedical literature, and the assembly of the extracted biochemical interactions into large-scale models such as protein signaling pathways.However, batch machine reading of literature at today's scale (PubMed alone indexes over 1 million papers per year) is unfeasible due to both cost and processing overhead.In this work, we introduce a focused reading approach to guide the machine reading of biomedical literature towards what literature should be read to answer a biomedical query as efficiently as possible.We introduce a family of algorithms for focused reading, including an intuitive, strong baseline, and a second approach which uses a reinforcement learning (RL) framework that learns when to explore (widen the search) or exploit (narrow it).We demonstrate that the RL approach is capable of answering more queries than the baseline, while being more efficient, i.e., reading fewer documents.
Enrique Noriega-Atala, Marco Antonio Valenzuela-Escárcega, Clayton T. Morrison, Mihai Surdeanu
EMNLP4
2017 Framing QA as Building and Ranking Intersentence Answer Justifications
abstract
We propose a question answering (QA) approach for standardized science exams that both identifies correct answers and produces compelling human-readable justifications for why those answers are correct. Our method first identifies the actual information needed in a question using psycholinguistic concreteness norms, then uses this information need to construct answer justifications by aggregating multiple sentences from different knowledge bases using syntactic and lexical information. We then jointly rank answers and their justifications using a reranking perceptron that treats justification quality as a latent variable. We evaluate our method on 1,000 multiple-choice questions from elementary school science exams, and empirically demonstrate that it performs better than several strong baselines, including neural network approaches. Our best configuration answers 44% of the questions correctly, where the top justifications for 57% of these correct answers contain a compelling human-readable justification that explains the inference required to arrive at the correct answer. We include a detailed characterization of the justification quality for both our method and a strong baseline, and show that information aggregation is key to addressing the information need in complex questions.
Peter A. Jansen, Rebecca Sharp, Mihai Surdeanu, Peter Clark
Comput. Linguistics3
2017 A scaffolding approach to coreference resolution integrating statistical and rule-based models
abstract
Abstract We describe a scaffolding approach to the task of coreference resolution that incrementally combines statistical classifiers, each designed for a particular mention type, with rule-based models (for sub-tasks well-matched to determinism). We motivate our design by an oracle-based analysis of errors in a rule-based coreference resolution system, showing that rule-based approaches are poorly suited to tasks that require a large lexical feature space, such as resolving pronominal and common-noun mentions. Our approach combines many advantages: it incrementally builds clusters integrating joint information about entities, uses rules for deterministic phenomena, and integrates rich lexical, syntactic, and semantic features with random forest classifiers well-suited to modeling the complex feature interactions that are known to characterize the coreference task. We demonstrate that all these decisions are important. The resulting system achieves 63.2 F1 on the CoNLL-2012 shared task dataset, outperforming the rule-based starting point by over seven F1 points. Similarly, our system outperforms an equivalent sieve-based approach that relies on logistic regression classifiers instead of random forests by over four F1 points. Lastly, we show that by changing the coreference resolution system from relying on constituent-based syntax to using dependency syntax, which can be generated in linear time, we achieve a runtime speedup of 550 per cent without considerable loss of accuracy.
Heeyoung Lee 0004, Mihai Surdeanu, Daniel Jurafsky
Nat. Lang. Eng.2
2016 What's in an Explanation? Characterizing Knowledge and Inference Requirements for Elementary Science Exams
abstract
QA systems have been making steady advances in the challenging elementary science exam domain. In this work, we develop an explanation-based analysis of knowledge and inference requirements, which supports a fine-grained characterization of the challenges. In particular, we model the requirements based on appropriate sources of evidence to be used for the QA task. We create requirements by first identifying suitable sentences in a knowledge base that support the correct answer, then use these to build explanations, filling in any necessary missing information. These explanations are used to create a fine-grained categorization of the requirements. Using these requirements, we compare a retrieval and an inference solver on 212 questions. The analysis validates the gains of the inference solver, demonstrating that it answers more questions requiring complex inference, while also providing insights into the relative strengths of the solvers and knowledge sources. We release the annotated questions and explanations as a resource with broad utility for science exam QA, including determining knowledge base construction targets, as well as supporting information aggregation in automated inference.
Peter A. Jansen, Niranjan Balasubramanian, Mihai Surdeanu, Peter Clark
COLING3
2016 Creating Causal Embeddings for Question Answering with Minimal Supervision
abstract
A common model for question answering (QA) is that a good answer is one that is closely related to the question, where relatedness is often determined using generalpurpose lexical models such as word embeddings.We argue that a better approach is to look for answers that are related to the question in a relevant way, according to the information need of the question, which may be determined through task-specific embeddings.With causality as a use case, we implement this insight in three steps.First, we generate causal embeddings cost-effectively by bootstrapping cause-effect pairs extracted from free text using a small set of seed patterns.Second, we train dedicated embeddings over this data, by using task-specific contexts, i.e., the context of a cause is its effect.Finally, we extend a state-of-the-art reranking approach for QA to incorporate these causal embeddings.We evaluate the causal embedding models both directly with a casual implication task, and indirectly, in a downstream causal QA task using data from Yahoo! Answers.We show that explicitly modeling causality improves performance in both tasks.In the QA task our best model achieves 37.3% P@1, significantly outperforming a strong baseline by 7.7% (relative).
Rebecca Sharp, Mihai Surdeanu, Peter A. Jansen, Peter Clark, Michael Hammond
EMNLP2
2016 Towards Using Social Media to Identify Individuals at Risk for Preventable Chronic Illness
Dane Bell, Daniel Fried, Luwen Huangfu, Mihai Surdeanu, Stephen G. Kobourov
LREC4
2016 Sieve-based Coreference Resolution in the Biomedical Domain
Dane Bell, Gus Hahn-Powell, Marco Antonio Valenzuela-Escárcega, Mihai Surdeanu
LREC4
2016 Odin's Runes: A Rule Language for Information Extraction
Marco Antonio Valenzuela-Escárcega, Gus Hahn-Powell, Mihai Surdeanu
LREC3
2015 Diamonds in the Rough: Event Extraction from Imperfect Microblog Data
abstract
Ander Intxaurrondo, Eneko Agirre, Oier Lopez de Lacalle, Mihai Surdeanu. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015.
Ander Intxaurrondo, Eneko Agirre, Oier Lopez de Lacalle, Mihai Surdeanu
HLT-NAACL4
2015 Spinning Straw into Gold: Using Free Text to Train Monolingual Alignment Models for Non-factoid Question Answering
abstract
Rebecca Sharp, Peter Jansen, Mihai Surdeanu, Peter Clark. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015.
Rebecca Sharp, Peter A. Jansen, Mihai Surdeanu, Peter Clark
HLT-NAACL3
2015 Two Practical Rhetorical Structure Theory Parsers
abstract
We describe the design, development, and API for two discourse parsers for Rhetorical Structure Theory.The two parsers use the same underlying framework, but one uses features that rely on dependency syntax, produced by a fast shift-reduce parser, whereas the other uses a richer feature space, including both constituent-and dependency-syntax and coreference information, produced by the Stanford CoreNLP toolkit.Both parsers obtain state-of-the-art performance, and use a very simple API consisting of, minimally, two lines of Scala code.We accompany this code with a visualization library that runs the two parsers in parallel, and displays the two generated discourse trees side by side, which provides an intuitive way of comparing the two parsers.
Mihai Surdeanu, Tom Hicks, Marco Antonio Valenzuela-Escárcega
HLT-NAACL1
2015 Higher-order Lexical Semantic Models for Non-factoid Answer Reranking
abstract
Lexical semantic models provide robust performance for question answering, but, in general, can only capitalize on direct evidence seen during training. For example, monolingual alignment models acquire term alignment probabilities from semi-structured data such as question-answer pairs; neural network language models learn term embeddings from unstructured text. All this knowledge is then used to estimate the semantic similarity between question and answer candidates. We introduce a higher-order formalism that allows all these lexical semantic models to chain direct evidence to construct indirect associations between question and answer texts, by casting the task as the traversal of graphs that encode direct term associations. Using a corpus of 10,000 questions from Yahoo! Answers, we experimentally demonstrate that higher-order methods are broadly applicable to alignment and language models, across both word and syntactic representations. We show that an important criterion for success is controlling for the semantic drift that accumulates during graph traversal. All in all, the proposed higher-order approach improves five out of the six lexical semantic models investigated, with relative gains of up to +13% over their first-order variants.
Daniel Fried, Peter A. Jansen, Gus Hahn-Powell, Mihai Surdeanu, Peter Clark
Trans. Assoc. Comput. Linguistics4
2014 Discourse Complements Lexical Semantics for Non-factoid Answer Reranking
abstract
We propose a robust answer reranking model for non-factoid questions that integrates lexical semantics with discourse information, driven by two representations of discourse: a shallow representation centered around discourse markers, and a deep one based on Rhetorical Structure Theory.We evaluate the proposed model on two corpora from different genres and domains: one from Yahoo! Answers and one from the biology domain, and two types of non-factoid questions: manner and reason.We experimentally demonstrate that the discourse structure of nonfactoid answers provides information that is complementary to lexical semantic similarity between question and answer, improving performance up to 24% (relative) over a state-of-the-art model that exploits lexical semantic similarity alone.We further demonstrate excellent domain transfer of discourse information, suggesting these discourse features have general utility to non-factoid question answering.
Peter A. Jansen, Mihai Surdeanu, Peter Clark
ACL (1)2
2014 Analyzing the language of food on social media
abstract
We investigate the predictive power behind the language of food on social media. We collect a corpus of over three million food-related posts from Twitter and demonstrate that many latent population characteristics can be directly predicted from this data: overweight rate, diabetes rate, political leaning, and home geographical location of authors. For all tasks, our language-based models significantly outperform the majority-class baselines. Performance is further improved with more complex natural language processing, such as topic modeling. We analyze which textual features have greatest predictive power for these datasets, providing insight into the connections between the language of food, geographic locale, and community characteristics. Lastly, we design and implement an online system for real-time query and visualization of the dataset. Visualization tools, such as geo-referenced heatmaps and temporal histograms, allow us to discover more complex, global patterns mirrored in the language of food.
Daniel Fried, Mihai Surdeanu, Stephen G. Kobourov, Melanie Hingle, Dane Bell
IEEE BigData2
2014 On the Importance of Text Analysis for Stock Price Prediction
Heeyoung Lee 0004, Mihai Surdeanu, Bill MacCartney, Daniel Jurafsky
LREC2
2014 Event Extraction Using Distant Supervision
Kevin Reschke, Martin Jankowiak, Mihai Surdeanu, Christopher D. Manning, Daniel Jurafsky
LREC3
2013 Identifying patent monetization entities
abstract
The United States has seen an explosion in patent litigation lawsuits in recent years. Recent studies indicate that a large proportion of these lawsuits, increasing from 22% in 2007 to 40% in 2011, were filed by patent monetization entities (PMEs), i.e., companies that hold patents, license patents, and file patent lawsuits, but do not sell products or provide services practicing the technologies described in their patents. We introduce a classifier that identifies which patent litigation lawsuits are initiated by PMEs. Using features extracted from the entities' litigation behavior, the patents they asserted, and their presence on the web, the proposed classifier correctly separates PMEs from operating companies with a F1 score of 85%. We believe that such a classifier will be a useful tool to policy makers and patent litigators, allowing them to gain a clearer picture of the 37,000+ patent lawsuits filed to date and assessing newly filed cases in real time.
Mihai Surdeanu, Sara Jeruss
ICAIL1
2013 Deterministic Coreference Resolution Based on Entity-Centric, Precision-Ranked Rules
abstract
We propose a new deterministic approach to coreference resolution that combines the global information and precise features of modern machine-learning models with the transparency and modularity of deterministic, rule-based systems. Our sieve architecture applies a battery of deterministic coreference models one at a time from highest to lowest precision, where each model builds on the previous model's cluster output. The two stages of our sieve-based architecture, a mention detection stage that heavily favors recall, followed by coreference sieves that are precision-oriented, offer a powerful way to achieve both high precision and high recall. Further, our approach makes use of global information through an entity-centric model that encourages the sharing of features across all mentions that point to the same real-world entity. Despite its simplicity, our approach gives state-of-the-art performance on several corpora and genres, and has also been incorporated into hybrid state-of-the-art coreference systems for Chinese and Arabic. Our system thus offers a new paradigm for combining knowledge in rule-based systems that has implications throughout computational linguistics.
Heeyoung Lee 0004, Angel X. Chang, Yves Peirsman, Nathanael Chambers, Mihai Surdeanu, Daniel Jurafsky
Comput. Linguistics5
2013 Selectional Preferences for Semantic Role Classification
abstract
This paper focuses on a well-known open issue in Semantic Role Classification (SRC) research: the limited influence and sparseness of lexical features. We mitigate this problem using models that integrate automatically learned selectional preferences (SP). We explore a range of models based on WordNet and distributional-similarity SPs. Furthermore, we demonstrate that the SRC task is better modeled by SP models centered on both verbs and prepositions, rather than verbs alone. Our experiments with SP-based models in isolation indicate that they outperform a lexical baseline with 20 F1 points in domain and almost 40 F1 points out of domain. Furthermore, we show that a state-of-the-art SRC system extended with features based on selectional preferences performs significantly better, both in domain (17% error reduction) and out of domain (13% error reduction). Finally, we show that in an end-to-end semantic role labeling system we obtain small but statistically significant improvements, even though our modified SRC model affects only approximately 4% of the argument candidates. Our post hoc error analysis indicates that the SP-based features help mostly in situations where syntactic information is either incorrect or insufficient to disambiguate the correct role.
Beñat Zapirain, Eneko Agirre, Lluís Màrquez, Mihai Surdeanu
Comput. Linguistics4
2012 Joint Entity and Event Coreference Resolution across Documents
Heeyoung Lee 0004, Marta Recasens, Angel X. Chang, Mihai Surdeanu, Daniel Jurafsky
EMNLP-CoNLL4
2012 Multi-instance Multi-label Learning for Relation Extraction
Mihai Surdeanu, Julie Tibshirani, Ramesh Nallapati, Christopher D. Manning
EMNLP-CoNLL1
2012 Combining joint models for biomedical event extraction
abstract
BACKGROUND: We explore techniques for performing model combination between the UMass and Stanford biomedical event extraction systems. Both sub-components address event extraction as a structured prediction problem, and use dual decomposition (UMass) and parsing algorithms (Stanford) to find the best scoring event structure. Our primary focus is on stacking where the predictions from the Stanford system are used as features in the UMass system. For comparison, we look at simpler model combination techniques such as intersection and union which require only the outputs from each system and combine them directly. RESULTS: First, we find that stacking substantially improves performance while intersection and union provide no significant benefits. Second, we investigate the graph properties of event structures and their impact on the combination of our systems. Finally, we trace the origins of events proposed by the stacked model to determine the role each system plays in different components of the output. We learn that, while stacking can propose novel event structures not seen in either base model, these events have extremely low precision. Removing these novel events improves our already state-of-the-art F1 to 56.6% on the test set of Genia (Task 1). Overall, the combined system formed via stacking ("FAUST") performed well in the BioNLP 2011 shared task. The FAUST system obtained 1st place in three out of four tasks: 1st place in Genia Task 1 (56.0% F1) and Task 2 (53.9%), 2nd place in the Epigenetics and Post-translational Modifications track (35.0%), and 1st place in the Infectious Diseases track (55.6%). CONCLUSION: We present a state-of-the-art event extraction system that relies on the strengths of structured prediction and model combination through stacking. Akin to results on other tasks, stacking outperforms intersection and union and leads to very strong results. The utility of model combination hinges on complementary views of the data, and we show that our sub-systems capture different graph properties of event structures. Finally, by removing low precision novel events, we show that performance from stacking can be further improved.
David McClosky, Sebastian Riedel 0001, Mihai Surdeanu, Andrew McCallum, Christopher D. Manning
BMC Bioinform.3
2012 Using Evolutive Summary Counters for Efficient Cooperative Caching in Search Engines
abstract
We propose and analyze a distributed cooperative caching strategy based on the Evolutive Summary Counters (ESC), a new data structure that stores an approximated record of the data accesses in each computing node of a search engine. The ESC capture the frequency of accesses to the elements of a data collection, and the evolution of the access patterns for each node in a network of computers. The ESC can be efficiently summarized into what we call ESC-summaries to obtain approximate statistics of the document entries accessed by each computing node. We use the ESC-summaries to introduce two algorithms that manage our distributed caching strategy, one for the distribution of the cache contents, ESC-placement, and another one for the search of documents in the distributed cache, ESC-search. While the former improves the hit rate of the system and keeps a large ratio of data accesses local, the latter reduces the network traffic by restricting the number of nodes queried to find a document. We show that our cooperative caching approach outperforms state-of-the-art models in both hit rate, throughput, and location recall for multiple scenarios, i.e., different query distributions and systems with varying degrees of complexity.
David Dominguez-Sal, Josep Aguilar-Saborit, Mihai Surdeanu, Josep Lluís Larriba-Pey
IEEE Trans. Parallel Distributed Syst.3
2011 Event Extraction as Dependency Parsing
David McClosky, Mihai Surdeanu, Christopher D. Manning
ACL2
2011 Risk analysis for intellectual property litigation
abstract
We introduce the problem of risk analysis for Intellectual Property (IP) lawsuits. More specifically, we focus on estimating the risk for participating parties using solely prior factors, i. e., historical and concurrent behavior of the entities involved in the case. This work represents a first step towards building a comprehensive legal risk assessment system for parties involved in litigation. This technology will allow parties to optimize their case parameters to minimize their own risk, or to settle disputes out of court and thereby ease the burden on the judicial system. In addition, it will also help U.S. courts detect and fix any inherent biases in the system.
Mihai Surdeanu, Ramesh Nallapati, George Gregory, Joshua Walker, Christopher D. Manning
ICAIL1
2011 Learning to Rank Answers to Non-Factoid Questions from Web Collections
abstract
This work investigates the use of linguistically motivated features to improve search, in particular for ranking answers to non-factoid questions. We show that it is possible to exploit existing large collections of question–answer pairs (from online social Question Answering sites) to extract such features and train ranking models which combine them effectively. We investigate a wide range of feature types, some exploiting natural language processing such as coarse word sense disambiguation, named-entity identification, syntactic parsing, and semantic role labeling. Our experiments demonstrate that linguistic features, in combination, yield considerable improvements in accuracy. Depending on the system settings we measure relative improvements of 14% to 21% in Mean Reciprocal Rank and Precision@1, providing one of the most compelling evidence to date that complex linguistic features such as word senses and semantic roles can have a significant impact on large-scale information retrieval tasks.
Mihai Surdeanu, Massimiliano Ciaramita, Hugo Zaragoza
Comput. Linguistics1
2010 A Multi-Pass Sieve for Coreference Resolution
Karthik Raghunathan, Heeyoung Lee 0004, Sudarshan Rangarajan, Nathanael Chambers, Mihai Surdeanu, Daniel Jurafsky, Christopher D. Manning
EMNLP5
2010 Ensemble Models for Dependency Parsing: Cheap and Good?
Mihai Surdeanu, Christopher D. Manning
HLT-NAACL1
2010 Improving Semantic Role Classification with Selectional Preferences
Beñat Zapirain, Eneko Agirre, Lluís Màrquez, Mihai Surdeanu
HLT-NAACL4
2009 Company-Oriented Extractive Summarization of Financial News
Katja Filippova, Mihai Surdeanu, Massimiliano Ciaramita, Hugo Zaragoza
EACL2
2008 Learning to Rank Answers on Large Online QA Collections
Mihai Surdeanu, Massimiliano Ciaramita, Hugo Zaragoza
ACL1
2008 Analysis of Joint Inference Strategies for the Semantic Role Labeling of Spanish and Catalan
Mihai Surdeanu, Roser Morante, Lluís Màrquez
CICLing1
2008 Cache-aware load balancing for question answering
abstract
The need for high performance and throughput Question Answering (QA) systems demands for their migration to distributed environments. However, even in such cases it is necessary to provide the distributed system with cooperative caches and load balancing facilities in order to achieve the desired goals. Until now, the literature on QA has not considered such a complex system as a whole. Currently, the load balancer regulates the assignment of tasks based only on the CPU and I/O loads without considering the status of the system cache.
David Dominguez-Sal, Mihai Surdeanu, Josep Aguilar-Saborit, Josep Lluís Larriba-Pey
CIKM2
2008 DeSRL: A Linear-Time Semantic Role Labeling System
Massimiliano Ciaramita, Giuseppe Attardi, Felice Dell'Orletta, Mihai Surdeanu
CoNLL4
2008 The CoNLL 2008 Shared Task on Joint Parsing of Syntactic and Semantic Dependencies
Mihai Surdeanu, Richard Johansson, Adam Meyers 0001, Lluís Màrquez, Joakim Nivre
CoNLL1
2007 A Multi-layer Collaborative Cache for Question Answering
David Dominguez-Sal, Josep Lluís Larriba-Pey, Mihai Surdeanu
Euro-Par3
2007 A Comparison of Statistical and Rule-Induction Learners for Automatic Tagging of Time Expressions in English
abstract
Proper recognition and handling of temporal information contained in a text is key to understanding the flow of events depicted in the text and their accompanying circumstances. Consequently, time expression recognition and representation of the time information they convey in a suitable normalized form is an important task relevant to several problems in Natural Language Processing. In particular, such an analysis is largely significant for Information Extraction (IE), Question Answering (QA) and Automatic Summarization (AS). The most common approach to time expression recognition in the past has been the use of handmade extraction rules (grammars), which also served as the basis for normalization. Our aim is to explore the possibilities afforded by applying machine learning techniques to the recognition of time expressions. We focus on recognizing the appearances of time expressions in text (not normalization) and transform the problem into one of chunking, where the aim is to correctly assign Begin, Inside or Outside (BIO) tags to tokens. In this paper, we explain the knowledge representation used and compare the results obtained in our experiments with two different methods, one statistical (support vector machines) and one of rule induction (FOIL). Our empirical analysis shows that SVMs are superior.
Jordi Poveda, Mihai Surdeanu, Jordi Turmo
TIME2
2007 Combination Strategies for Semantic Role Labeling
abstract
This paper introduces and analyzes a battery of inference models for the problem of semantic role labeling: one based on constraint satisfaction, and several strategies that model the inference as a meta-learning problem using discriminative classifiers. These classifiers are developed with a rich set of novel features that encode proposition and sentence-level information. To our knowledge, this is the first work that: (a) performs a thorough analysis of learning-based inference models for semantic role labeling, and (b) compares several inference strategies in this context. We evaluate the proposed inference strategies in the framework of the CoNLL-2005 shared task using only automatically-generated syntactic information. The extensive experimental evaluation and analysis indicates that all the proposed inference strategies are successful -they all outperform the current best results reported in the CoNLL-2005 evaluation exercise- but each of the proposed approaches has its advantages and disadvantages. Several important traits of a state-of-the-art SRL combination strategy emerge from this analysis: (i) individual models should be combined at the granularity of candidate arguments rather than at the granularity of complete solutions; (ii) the best combination strategy uses an inference model based in learning; and (iii) the learning-based inference benefits from max-margin classifiers and global feedback.
Mihai Surdeanu, Lluís Màrquez, Xavier Carreras, Pere Comas
J. Artif. Intell. Res.1
2006 Projective Dependency Parsing with Perceptron
Xavier Carreras, Mihai Surdeanu, Lluís Màrquez
CoNLL2
2006 Design and performance analysis of a factoid question answering system for spontaneous speech transcriptions
abstract
This paper introduces a QA designed from scratch to handle speech transcriptions. The system’s strength is achieved by analyz-ing the speech transcriptions with a mix of IR-oriented methodolo-gies and a small number of robust NLP components. We evaluate the system on transcriptions of spontaneous speech from several 1-hour-long seminars and presentations and show that the system obtains encouraging performance. Index Terms: question answering, natural language processing 1.
Mihai Surdeanu, David Dominguez-Sal, Pere Comas
INTERSPEECH1
2005 Semantic Role Labeling Using Complete Syntactic Analysis
Mihai Surdeanu, Jordi Turmo
CoNLL1
2005 Named entity recognition from spontaneous open-domain speech
abstract
This paper presents an analysis of named entity recognition and classification in spontaneous speech transcripts. We annotated a significant fraction of the Switchboard corpus with six named entity classes and investigated a battery of machine learning models that include lexical, syntactic, and semantic attributes. The best recognition and classification model obtains promis-ing results, approaching within 5 % a system evaluated on clean textual data. 1.
Mihai Surdeanu, Jordi Turmo, Eli Comelles
INTERSPEECH1
2005 A hybrid unsupervised approach for document clustering
abstract
We propose a hybrid, unsupervised document clustering approach that combines a hierarchical clustering algorithm with Expectation Maximization. We developed several heuristics to automatically select a subset of the clusters generated by the first algorithm as the initial points of the second one. Furthermore, our initialization algorithm generates not only an initial model for the iterative refinement algorithm but also an estimate of the model dimension, thus eliminating another important element of human supervision. We have evaluated the proposed system on five real-world document collections. The results show that our approach generates clustering solutions of higher quality than both its individual components.
Mihai Surdeanu, Jordi Turmo, Alicia Ageno
KDD1
2003 Using Predicate-Argument Structures for Information Extraction
abstract
In this paper we present a novel, customizable IE paradigm that takes advantage of predicate-argument structures.We also introduce a new way of automatically identifying predicate argument structures, which is central to our IE paradigm.It is based on: (1) an extended set of features; and (2) inductive decision tree learning.The experimental results prove our claim that accurate predicate-argument structures enable high quality IE results.
Mihai Surdeanu, Sanda M. Harabagiu, John Williams 0001, Paul Aarseth
ACL1
2003 Performance issues and error analysis in an open-domain question answering system
abstract
This paper presents an in-depth analysis of a state-of-the-art Question Answering system. Several scenarios are examined: (1) the performance of each module in a serial baseline system, (2) the impact of feedbacks and the insertion of a logic prover, and (3) the impact of various retrieval strategies and lexical resources. The main conclusion is that the overall performance depends on the depth of natural language processing resources and the tools used for answer finding.
Dan I. Moldovan, Marius Pasca, Sanda M. Harabagiu, Mihai Surdeanu
ACM Trans. Inf. Syst.4
2002 Performance Issues and Error Analysis in an Open-Domain Question Answering System
abstract
This paper presents an in-depth analysis of a state-of-the-art Question Answering system. Several scenarios are examined: (1) the performance of each module in a serial baseline system, (2) the impact of feedbacks and the insertion of a logic prover, and (3) the impact of various lexical resources. The main conclusion is that the overall performance depends on the depth of natural language processing resources and the tools used for answer finding.
Dan I. Moldovan, Marius Pasca, Sanda M. Harabagiu, Mihai Surdeanu
ACL4
2002 Design and Performance Analysis of a Distributed Java Virtual Machine
abstract
This paper introduces DISK, a distributed Java Virtual Machine for networks of heterogenous workstations. Several research issues are addressed. A novelty of the system is its object-based, multiple-writer memory consistency protocol (OMW). The correctness of the protocol and its Java compliance is demonstrated by comparing the nonoperational definitions of release consistency, the consistency model implemented by OMW, with the Java Virtual Machine memory consistency model (JVMC), as defined in the Java Virtual Machine Specification. An analytical performance model was developed to study and compare the design trade-offs between OMW and the lazy invalidate release consistency (LI) protocols as a function of the number of processors, network characteristics, and application types. The DISK system has been implemented and running on a network of 16 Pentium III computers interconnected by a 100 Mbps Ethernet network. Experiments performed with two applications: parallel matrix multiplication and traveling salesman problem confirm the analytical model.
Mihai Surdeanu, Dan I. Moldovan
IEEE Trans. Parallel Distributed Syst.1
2002 Performance Analysis of a Distributed Question/Answering System
abstract
The problem of question/answering (Q/A) is to find answers to open-domain questions by searching large collections of documents. Unlike information retrieval systems very common today in the form of Internet search engines, Q/A systems do not retrieve documents, but instead provide short, relevant answers located in small fragments of text. This enhanced functionality comes with a price: Q/A systems are significantly slower and require more hardware resources than information retrieval systems. This paper proposes a distributed Q/A architecture that enhances the system throughput through the exploitation of interquestion parallelism and dynamic load balancing and reduces the individual question response time through the exploitation of intraquestion parallelism. Inter and intraquestion parallelism are both exploited using several scheduling points: one before the Q/A task is started and two embedded in the Q/A task. An analytical performance model is introduced. The model analyzes both the interquestion parallelism overhead generated by the migration of questions and the intraquestion parallelism overhead generated by the partitioning of the Q/A task. The analytical model indicates that both question migration and partitioning are required for a high-performance system.
Mihai Surdeanu, Dan I. Moldovan, Sanda M. Harabagiu
IEEE Trans. Parallel Distributed Syst.1
2001 The Role of Lexico-Semantic Feedback in Open-Domain Textual Question-Answering
abstract
This paper presents an open-domain textual Question-Answering system that uses several feedback loops to enhance its performance. These feedback loops combine in a new way statistical results with syntactic, semantic or pragmatic information derived from texts and lexical databases. The paper presents the contribution of each feedback loop to the overall performance of 76% human-assessed precise answers.
Sanda M. Harabagiu, Dan I. Moldovan, Marius Pasca, Rada Mihalcea, Mihai Surdeanu, Razvan C. Bunescu, Roxana Girju, Vasile Rus, Paul Morarescu
ACL5
2001 Performance Analysis of a Distributed Question/Answering System
abstract
The problem of question/answering (Q/A) is to find answers to open-domain questions by searching a large collection of documents. Unlike Internet search engines, Q/A systems provide short, relevant answers to questions. Due to the complex natural language processing involved that is CPU intensive, and the retrieval of large number of documents that is disk intensive, the time performance of sequential Q/A systems is rather slow. This paper presents the design and performance analysis of a distributed state-of-the-art Q/A system. The design is modular and parallelism is dynamically exploited at inter and intra-question levels. Several schedule points are used to balance the load. An analytical performance model is given backed up by experimental results.
Mihai Surdeanu, Dan I. Moldovan, Sanda M. Harabagiu
IPDPS1
2000 Distributed Java Virtual Machine for Message Passing Architectures
abstract
This paper introduces a distributed shared memory Java Virtual Machine architecture. This project is targeted for any distributed message-passing architecture, and specifically for networks of workstations. The whole system is implemented in user space which offers portability and flexibility. The memory consistency is provided by one of four protocols implementing release consistency. The novelty of the consistency protocols presented is that access faults are avoided by replicating objects ahead-of-time where necessary. The relative performance of these protocols is evaluated for three benchmark applications. Our experimental results indicate that, in the majority of cases, update protocols outperform invalidate protocols.
Mihai Surdeanu
ICDCS1