EDBT 2026 Demo / reviewers in the wild / expert
Gabriel Stanovsky
dblp:166/1740 · also Gabi Stanovsky
· DBLP profile ↗
36ranked-venue papers
7as first author
22since 2021 · last 2026
0000-0002-2420-8979ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 36 · 7 first-author · 22 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | 🧑🍳 Cooking Up Creativity : Enhancing LLM Creativity through Structured RecombinationabstractAbstract Large Language Models (LLMs) excel at many tasks, yet they struggle to produce truly creative, diverse ideas. In this paper, we introduce a novel approach that enhances LLM creativity. We apply LLMs for translating between natural language and structured representations, and perform the core creative leap via cognitively inspired manipulations on these representations. Our notion of creativity goes beyond superficial token-level variations; rather, we recombine structured representations of existing ideas, enabling our system to effectively explore a more abstract landscape of ideas. We demonstrate our approach in the culinary domain with DishCover, a model that generates creative recipes. Experiments and domain-expert evaluations reveal that our outputs, which are mostly coherent and feasible, significantly surpass GPT-4o in terms of novelty and diversity, thus outperforming it in creative generation. We hope our work inspires further research into structured creativity in AI. Moran Mizrahi 0001, Chen Shani, Gabriel Stanovsky, Daniel Jurafsky, Dafna Shahaf |
Trans. Assoc. Comput. Linguistics | 3 |
| 2025 | Can LLMs Help Uncover Insights about LLMs? A Large-Scale, Evolving Literature Analysis of Frontier LLMsabstractThe surge of LLM studies makes synthesizing their findings challenging. Analysis of experimental results from literature can uncover important trends across studies, but the time-consuming nature of manual data extraction limits its use.Our study presents a semi-automated approach for literature analysis that accelerates data extraction using LLMs.It automatically identifies relevant arXiv papers, extracts experimental results and related attributes, and organizes them into a structured dataset, LLMEvalDB.We then conduct an automated literature analysis of frontier LLMs, reducing the effort of paper surveying and data extraction by more than 93% compared to manual approaches.We validate LLMEvalDB by showing that it reproduces key findings from a recent manual analysis of Chain-of-Thought (CoT) reasoning and also uncovers new insights that go beyond it, showing, for example, that in-context examples benefit coding & multimodal tasks but offer limited gains in math reasoning tasks compared to zero-shot CoT.Our automatically updatable dataset enables continuous tracking of target models by extracting evaluation studies as new data becomes available. Through LLMEvalDB and empirical analysis, we provide insights into LLMs while facilitating ongoing literature analyses of their behavior. Jungsoo Park, Junmo Kang, Gabriel Stanovsky, Alan Ritter |
ACL (1) | 3 |
| 2025 | Looking Beyond the Top-1: Transformers Determine Top Tokens in OrderabstractUncovering the inner mechanisms of Transformer models offers insights into how they process and represent information. In this work, we analyze the computation performed by Transformers in the layers after the top-1 prediction remains fixed, known as the “saturation event”. We expand this concept to top-k tokens, demonstrating that similar saturation events occur across language, vision, and speech models. We find that these events occur in order of the corresponding tokens’ ranking, i.e., the model first decides on the top ranking token, then the second highest ranking token, and so on. This phenomenon seems intrinsic to the Transformer architecture, occurring across different variants, and even in untrained Transformers. We propose that these events reflect task transitions, where determining each token corresponds to a discrete task. We show that it is possible to predict the current task from hidden layer embedding, and demonstrate that we can cause the model to switch to the next task via intervention. Leveraging our findings, we introduce a token-level early-exit strategy, surpassing existing methods in balancing performance and efficiency and show how to exploit saturation events for better language modeling. Daria Lioubashevski, Tomer Schlank, Gabriel Stanovsky, Ariel Goldstein |
ICML | 3 |
| 2025 | The State and Fate of Summarization Datasets: A SurveyabstractNoam Dahan, Gabriel Stanovsky. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Noam Dahan, Gabriel Stanovsky |
NAACL (Long Papers) | 2 |
| 2025 | Surveying the Landscape of Image Captioning Evaluation: A Comprehensive Taxonomy, Trends, and Metrics AnalysisabstractAbstract The task of image captioning has recently been gaining popularity, and with it the complex task of evaluating the quality of image captioning models. In this work, we present the first survey and taxonomy of over 70 different image captioning metrics and their usage in hundreds of papers, specifically designed to help users select the most suitable metric for their needs. We find that despite the diversity of proposed metrics, the vast majority of studies rely on only five popular metrics, which we show to be weakly correlated with human ratings. We hypothesize that combining a diverse set of metrics can enhance correlation with human ratings. As an initial step, we demonstrate that a linear regression-based ensemble method, which we call EnsembEval, trained on one human ratings dataset, achieves improved correlation across five additional datasets, showing there is a lot of room for improvement by leveraging a diverse set of metrics.1 Uri Berger, Gabriel Stanovsky, Omri Abend, Lea Frermann |
Trans. Assoc. Comput. Linguistics | 2 |
| 2024 | A Nurse is Blue and Elephant is Rugby: Cross Domain Alignment in Large Language Models Reveal Human-like Patterns
Asaf Yehudai, Taelin Karidi, Gabriel Stanovsky, Ariel Goldstein, Omri Abend |
CogSci | 3 |
| 2024 | Applying Intrinsic Debiasing on Downstream Tasks: Challenges and Considerations for Machine TranslationabstractMost works on gender bias focus on intrinsic bias -removing traces of information about a protected group from the model's internal representation.However, these works are often disconnected from the impact of such debiasing on downstream applications, which is the main motivation for debiasing in the first place.In this work, we systematically test how methods for intrinsic debiasing affect neural machine translation models, by measuring the extrinsic bias of such systems under different design choices.We highlight three challenges and mismatches between the debiasing techniques and their end-goal usage, including the choice of embeddings to debias, the mismatch between words and sub-word tokens debiasing, and the effect of translating from English to different target languages.We find that these considerations have a significant impact on downstream performance and the success of debiasing.1 Bar Iluz, Yanai Elazar, Asaf Yehudai, Gabriel Stanovsky |
EMNLP | 4 |
| 2024 | State of What Art? A Call for Multi-Prompt LLM EvaluationabstractAbstract Recent advances in LLMs have led to an abundance of evaluation benchmarks, which typically rely on a single instruction template per task. We create a large-scale collection of instruction paraphrases and comprehensively analyze the brittleness introduced by single-prompt evaluations across 6.5M instances, involving 20 different LLMs and 39 tasks from 3 benchmarks. We find that different instruction templates lead to very different performance, both absolute and relative. Instead, we propose a set of diverse metrics on multiple instruction paraphrases, specifically tailored for different use cases (e.g., LLM vs. downstream development), ensuring a more reliable and meaningful assessment of LLM capabilities. We show that our metrics provide new insights into the strengths and limitations of current LLMs. Moran Mizrahi 0001, Guy Kaplan, Dan Malkin, Rotem Dror, Dafna Shahaf, Gabriel Stanovsky |
Trans. Assoc. Comput. Linguistics | 6 |
| 2023 | VASR: Visual Analogies of Situation RecognitionabstractA core process in human cognition is analogical mapping: the ability to identify a similar relational structure between different situations. We introduce a novel task, Visual Analogies of Situation Recognition, adapting the classical word-analogy task into the visual domain. Given a triplet of images, the task is to select an image candidate B' that completes the analogy (A to A' is like B to what?). Unlike previous work on visual analogy that focused on simple image transformations, we tackle complex analogies requiring understanding of scenes. We leverage situation recognition annotations and the CLIP model to generate a large set of 500k candidate analogies. Crowdsourced annotations for a sample of the data indicate that humans agree with the dataset label ~80% of the time (chance level 25%). Furthermore, we use human annotations to create a gold-standard dataset of 3,820 validated analogies. Our experiments demonstrate that state-of-the-art models do well when distractors are chosen randomly (~86%), but struggle with carefully chosen distractors (~53%, compared to 90% human accuracy). We hope our dataset will encourage the development of new analogy-making models. Website: https://vasr-dataset.github.io/ Yonatan Bitton, Ron Yosef, Eli Strugo, Dafna Shahaf, Roy Schwartz 0001, Gabriel Stanovsky |
AAAI | 6 |
| 2023 | Comparing Humans and Models on a Similar Scale: Towards Cognitive Gender Bias Evaluation in Coreference Resolution
Gili Lior, Gabriel Stanovsky |
CogSci | 2 |
| 2023 | Evaluating and Improving the Coreference Capabilities of Machine Translation ModelsabstractMachine translation (MT) requires a wide range of linguistic capabilities, which current end-to-end models are expected to learn implicitly by observing aligned sentences in bilingual corpora.In this work, we ask: How well do MT models learn coreference resolution from implicit signal?To answer this question, we develop an evaluation methodology that derives coreference clusters from MT output and evaluates them without requiring annotations in the target language.We further evaluate several prominent open-source and commercial MT systems, translating from English to six target languages, and compare them to state-of-theart coreference resolvers on three challenging benchmarks.Our results show that the monolingual resolvers greatly outperform MT models.Motivated by this result, we experiment with different methods for incorporating the output of coreference resolution models in MT, showing improvement over strong baselines.1 Asaf Yehudai, Arie Cattan, Omri Abend, Gabriel Stanovsky |
EACL | 4 |
| 2023 | The Perfect Victim: Computational Analysis of Judicial Attitudes towards Victims of Sexual ViolenceabstractWe develop computational models to analyze court statements in order to assess judicial attitudes toward victims of sexual violence in the Israeli court system. The study examines the resonance of "rape myths" in the criminal justice system's response to sex crimes, in particular in judicial assessment of victim's credibility. We begin by formulating an ontology for evaluating judicial attitudes toward victim's credibility, with eight ordinal labels and binary categorizations. Second, we curate a manually annotated dataset for judicial assessments of victim's credibility in the Hebrew language, as well as a model that can extract credibility labels from court cases. The dataset consists of 855 verdict decision documents in sexual assault cases from 1990-2021, annotated with the help of legal experts and trained law students. The model uses a combined approach of syntactic and latent structures to find sentences that convey the judge's attitude towards the victim and classify them according to the credibility label set. Our ontology, data, and models will be made available upon request, in the hope they spur future progress in this judicial important task. Eliya Habba, Renana Keydar, Dan Bareket, Gabriel Stanovsky |
ICAIL | 4 |
| 2023 | Breaking Common Sense: WHOOPS! A Vision-and-Language Benchmark of Synthetic and Compositional ImagesabstractWeird, unusual, and uncanny images pique the curiosity of observers because they challenge commonsense. For example, an image released during the 2022 world cup depicts the famous soccer stars Lionel Messi and Cristiano Ronaldo playing chess, which playfully violates our expectation that their competition should occur on the football field.1Humans can easily recognize and interpret these unconventional images, but can AI models do the same? We introduce WHOOPS!, a new dataset and benchmark for visual commonsense. The dataset is comprised of purposefully commonsense-defying images created by designers using publicly-available image generation tools like Midjourney. We consider several tasks posed over the dataset. In addition to image captioning, cross-modal matching, and visual question answering, we introduce a difficult explanation generation task, where models must identify and explain why a given image is unusual. Our results show that state-of-the-art models such as GPT3 and BLIP2 still lag behind human performance on WHOOPS!. We hope our dataset will inspire the development of AI models with stronger visual commonsense reasoning abilities.2 Nitzan Guetta, Yonatan Bitton, Jack Hessel, Ludwig Schmidt, Yuval Elovici, Gabriel Stanovsky, Roy Schwartz 0001 |
ICCV | 6 |
| 2023 | Exploring the Impact of Training Data Distribution and Subword Tokenization on Gender Bias in Machine TranslationabstractBar Iluz, Tomasz Limisiewicz, Gabriel Stanovsky, David Mareček. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Bar Iluz, Tomasz Limisiewicz, Gabriel Stanovsky, David Marecek |
IJCNLP (1) | 3 |
| 2022 | GENIE: Toward Reproducible and Standardized Human Evaluation for Text GenerationabstractDaniel Khashabi, Gabriel Stanovsky, Jonathan Bragg, Nicholas Lourie, Jungo Kasai, Yejin Choi, Noah A. Smith, Daniel Weld. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Daniel Khashabi, Gabriel Stanovsky, Jonathan Bragg, Nicholas Lourie, Jungo Kasai, Yejin Choi 0001, Noah A. Smith, Daniel S. Weld |
EMNLP | 2 |
| 2022 | "Covid vaccine is against Covid but Oxford vaccine is made at Oxford!" Semantic Interpretation of Proper Noun CompoundsabstractProper noun compounds, e.g., "Covid vaccine", convey information in a succinct manner (a "Covid vaccine" is a "vaccine that immunizes against the Covid disease").These are commonly used in short-form domains, such as news headlines, but are largely ignored in information-seeking applications.To address this limitation, we release a new manually annotated dataset, PRONCI, consisting of 22.5K proper noun compounds along with their freeform semantic interpretations.PRONCI is 60 times larger than prior noun compound datasets and also includes non-compositional examples, which have not been previously explored.We experiment with various neural models for automatically generating the semantic interpretations from proper noun compounds, ranging from few-shot prompting to supervised learning, with varying degrees of knowledge about the constituent nouns.We find that adding targeted knowledge, particularly about the common noun, results in performance gains of upto 2.8%.Finally, we integrate our model generated interpretations with an existing Open IE system and observe an 7.5% increase in yield at a precision of 85%.The dataset and code are available at https://github.com/dair-iitd/pronci. Keshav Kolluru, Gabriel Stanovsky, Mausam |
EMNLP | 2 |
| 2022 | A Computational Acquisition Model for Multimodal Word CategorizationabstractUri Berger, Gabriel Stanovsky, Omri Abend, Lea Frermann. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Uri Berger, Gabriel Stanovsky, Omri Abend, Lea Frermann |
NAACL-HLT | 2 |
| 2022 | A Balanced Data Approach for Evaluating Cross-Lingual Transfer: Mapping the Linguistic Blood BankabstractWe show that the choice of pretraining languages affects downstream cross-lingual transfer for BERT-based models.We inspect zeroshot performance in balanced data conditions to mitigate data size confounds, classifying pretraining languages that improve downstream performance as donors, and languages that are improved in zero-shot performance as recipients.We develop a method of quadratic time complexity in the number of languages to estimate these relations, instead of an exponential exhaustive computation of all possible combinations.We find that our method is effective on a diverse set of languages spanning different linguistic features and two downstream tasks.Our findings can inform developers of largescale multilingual language models in choosing better pretraining configurations. 1 Dan Malkin, Tomasz Limisiewicz, Gabriel Stanovsky |
NAACL-HLT | 3 |
| 2022 | WinoGAViL: Gamified Association Benchmark to Challenge Vision-and-Language ModelsabstractWhile vision-and-language models perform well on tasks such as visual question answering, they struggle when it comes to basic human commonsense reasoning skills. In this work, we introduce WinoGAViL: an online game of vision-and-language associations (e.g., between werewolves and a full moon), used as a dynamic evaluation benchmark. Inspired by the popular card game Codenames, a spymaster gives a textual cue related to several visual candidates, and another player tries to identify them. Human players are rewarded for creating associations that are challenging for a rival AI model but still solvable by other human players. We use the game to collect 3.5K instances, finding that they are intuitive for humans (>90% Jaccard index) but challenging for state-of-the-art AI models, where the best model (ViLT) achieves a score of 52%, succeeding mostly where the cue is visually salient. Our analysis as well as the feedback we collect from players indicate that the collected associations require diverse reasoning skills, including general knowledge, common sense, abstraction, and more. We release the dataset, the code and the interactive game, allowing future data collection that can be used to develop models with better association abilities. Yonatan Bitton, Nitzan Guetta, Ron Yosef, Yuval Elovici, Mohit Bansal, Gabriel Stanovsky, Roy Schwartz 0001 |
NeurIPS | 6 |
| 2021 | Process-Level Representation of Scientific Protocols with Interactive AnnotationabstractWe develop Process Execution Graphs (PEG), a document-level representation of real-world wet lab biochemistry protocols, addressing challenges such as cross-sentence relations, long-range coreference, grounding, and implicit arguments.We manually annotate PEGs in a corpus of complex lab protocols with a novel interactive textual simulator that keeps track of entity traits and semantic constraints during annotation.We use this data to develop graph-prediction models, finding them to be good at entity identification and local relation extraction, while our corpus facilitates further exploration of challenging long-range relations.1 Ronen Tamari, Fan Bai 0006, Alan Ritter, Gabriel Stanovsky |
EACL | 4 |
| 2021 | Filling the Gaps in Ancient Akkadian Texts: A Masked Language Modelling ApproachabstractWe present models which complete missing text given transliterations of ancient Mesopotamian documents, originally written on cuneiform clay tablets (2500 BCE -100 CE).Due to the tablets' deterioration, scholars often rely on contextual cues to manually fill in missing parts in the text in a subjective and time-consuming process.We identify that this challenge can be formulated as a masked language modelling task, used mostly as a pretraining objective for contextualized language models.Following, we develop several architectures focusing on the Akkadian language, the lingua franca of the time.We find that despite data scarcity (1M tokens) we can achieve state of the art performance on missing tokens prediction (89% hit@5) using a greedy decoding scheme and pretraining on data from other languages and different time periods.Finally, we conduct human evaluations showing the applicability of our models in assisting experts to transcribe texts in extinct languages. Koren Lazar, Benny Saret, Asaf Yehudai, Wayne Horowitz, Nathan Wasserman, Gabriel Stanovsky |
EMNLP (1) | 6 |
| 2021 | Automatic Generation of Contrast Sets from Scene Graphs: Probing the Compositional Consistency of GQAabstractYonatan Bitton, Gabriel Stanovsky, Roy Schwartz, Michael Elhadad. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Yonatan Bitton, Gabriel Stanovsky, Roy Schwartz 0001, Michael Elhadad |
NAACL-HLT | 2 |
| 2020 | Active Learning for Coreference Resolution using Discrete AnnotationabstractWe improve upon pairwise annotation for active learning in coreference resolution, by asking annotators to identify mention antecedents if a presented mention pair is deemed not coreferent.This simple modification, when combined with a novel mention clustering algorithm for selecting which examples to label, is much more efficient in terms of the performance obtained per annotation budget.In experiments with existing benchmark coreference datasets, we show that the signal from this additional question leads to significant performance gains per human-annotation hour.Future work can use our annotation protocol to effectively develop coreference models for new domains.Our code is publicly available.1 Belinda Z. Li, Gabriel Stanovsky, Luke Zettlemoyer |
ACL | 2 |
| 2020 | Controlled Crowdsourcing for High-Quality QA-SRL AnnotationabstractPaul Roit, Ayal Klein, Daniela Stepanov, Jonathan Mamou, Julian Michael, Gabriel Stanovsky, Luke Zettlemoyer, Ido Dagan. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. Paul Roit, Ayal Klein, Daniela Stepanov, Jonathan Mamou, Julian Michael, Gabriel Stanovsky, Luke Zettlemoyer, Ido Dagan |
ACL | 6 |
| 2020 | The Right Tool for the Job: Matching Model and Instance ComplexitiesabstractAs NLP models become larger, executing a trained model requires significant computational resources incurring monetary and environmental costs.To better respect a given inference budget, we propose a modification to contextual representation fine-tuning which, during inference, allows for an early (and fast) "exit" from neural network calculations for simple instances, and late (and accurate) exit for hard instances.To achieve this, we add classifiers to different layers of BERT and use their calibrated confidence scores to make early exit decisions.We test our proposed modification on five different datasets in two tasks: three text classification datasets and two natural language inference benchmarks.Our method presents a favorable speed/accuracy tradeoff in almost all cases, producing models which are up to five times faster than the state of the art, while preserving their accuracy.Our method also requires almost no additional training resources (in either time or parameters) compared to the baseline BERT model.Finally, our method alleviates the need for costly retraining of multiple models at different levels of efficiency; we allow users to control the inference speed/accuracy tradeoff using a single trained model, by setting a single variable at inference time.We publicly release our code.1 * Research completed during an internship at AI2. 1 github.com/allenai/sledgehammerLayer 0 Layer i Layer k Layer n Input Layer l Layer j Is confident?Yes Roy Schwartz 0001, Gabriel Stanovsky, Swabha Swayamdipta, Jesse Dodge, Noah A. Smith |
ACL | 2 |
| 2020 | MOCHA: A Dataset for Training and Evaluating Generative Reading Comprehension MetricsabstractPosing reading comprehension as a generation problem provides a great deal of flexibility, allowing for open-ended questions with few restrictions on possible answers.However, progress is impeded by existing generation metrics, which rely on token overlap and are agnostic to the nuances of reading comprehension.To address this, we introduce a benchmark for training and evaluating generative reading comprehension metrics: MOdeling Correctness with Human Annotations.MOCHA contains 40K human judgement scores on model outputs from 6 diverse question answering datasets and an additional set of minimal pairs for evaluation.Using MOCHA, we train a Learned Evaluation metric for Reading Comprehension, LERC, to mimic human judgement scores.LERC outperforms baseline metrics by 10 to 36 absolute Pearson points on held-out annotations.When we evaluate robustness on minimal pairs, LERC achieves 80% accuracy, outperforming baselines by 14 to 26 absolute percentage points while leaving significant room for improvement.MOCHA presents a challenging problem for developing accurate and robust generative reading comprehension metrics. 1 Anthony Chen, Gabriel Stanovsky, Sameer Singh 0001, Matt Gardner 0001 |
EMNLP (1) | 2 |
| 2019 | Evaluating Gender Bias in Machine TranslationabstractWe present the first challenge set and evaluation protocol for the analysis of gender bias in machine translation (MT).Our approach uses two recent coreference resolution datasets composed of English sentences which cast participants into non-stereotypical gender roles (e.g., "The doctor asked the nurse to help her in the operation").We devise an automatic gender bias evaluation method for eight target languages with grammatical gender, based on morphological analysis (e.g., the use of female inflection for the word "doctor").Our analyses show that four popular industrial MT systems and two recent state-of-the-art academic MT models are significantly prone to gender-biased translation errors for all tested target languages. Gabriel Stanovsky, Noah A. Smith, Luke Zettlemoyer |
ACL (1) | 1 |
| 2019 | On the Limits of Learning to Actively Learn Semantic RepresentationsabstractOne of the goals of natural language understanding is to develop models that map sentences into meaning representations.However, training such models requires expensive annotation of complex structures, which hinders their adoption.Learning to actively-learn (LTAL) is a recent paradigm for reducing the amount of labeled data by learning a policy that selects which samples should be labeled.In this work, we examine LTAL for learning semantic representations, such as QA-SRL.We show that even an oracle policy that is allowed to pick examples that maximize performance on the test set (and constitutes an upper bound on the potential of LTAL), does not substantially improve performance compared to a random policy.We investigate factors that could explain this finding and show that a distinguishing characteristic of successful applications of LTAL is the interaction between optimization and the oracle policy selection process.In successful applications of LTAL, the examples selected by the oracle policy do not substantially depend on the optimization procedure, while in our setup the stochastic nature of optimization strongly affects the examples selected by the oracle.We conclude that the current applicability of LTAL for improving data efficiency in learning semantic meaning representations is limited. Omri Koshorek, Gabriel Stanovsky, Yichu Zhou, Vivek Srikumar, Jonathan Berant |
CoNLL | 2 |
| 2018 | Semantics as a Foreign LanguageabstractWe propose a novel approach to semantic dependency parsing (SDP) by casting the task as an instance of multi-lingual machine translation, where each semantic representation is a different foreign dialect.To that end, we first generalize syntactic linearization techniques to account for the richer semantic dependency graph structure.Following, we design a neural sequence-to-sequence framework which can effectively recover our graph linearizations, performing almost on-par with previous SDP state-of-the-art while requiring less parallel training annotations.Beyond SDP, our linearization technique opens the door to integration of graph-based semantic representations as features in neural models for downstream applications. Gabriel Stanovsky, Ido Dagan |
EMNLP | 1 |
| 2018 | Spot the Odd Man Out: Exploring the Associative Power of Lexical ResourcesabstractWe propose Odd-Man-Out, a novel task which aims to test different properties of word representations.An Odd-Man-Out puzzle is composed of 5 (or more) words, and requires the system to choose the one which does not belong with the others.We show that this simple setup is capable of teasing out various properties of different popular lexical resources (like WordNet and pre-trained word embeddings), while being intuitive enough to annotate on a large scale.In addition, we propose a novel technique for training multi-prototype word representations, based on unsupervised clustering of ELMo embeddings, and show that it surpasses all other representations on all Odd-Man-Out collections. Gabriel Stanovsky, Mark Hopkins |
EMNLP | 1 |
| 2018 | Supervised Open Information ExtractionabstractGabriel Stanovsky, Julian Michael, Luke Zettlemoyer, Ido Dagan. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Gabriel Stanovsky, Julian Michael, Luke Zettlemoyer, Ido Dagan |
NAACL-HLT | 1 |
| 2017 | Recognizing Mentions of Adverse Drug Reaction in Social Media Using Knowledge-Infused Recurrent ModelsabstractRecognizing mentions of Adverse Drug Reactions (ADR) in social media is challenging: ADR mentions are contextdependent and include long, varied and unconventional descriptions as compared to more formal medical symptom terminology.We use the CADEC corpus to train a recurrent neural network (RNN) transducer, integrated with knowledge graph embeddings of DBpedia, and show the resulting model to be highly accurate (93.4 F1).Furthermore, even when lacking high quality expert annotations, we show that by employing an active learning technique and using purpose built annotation tools, we can train the RNN to perform well (83.9F1). Gabriel Stanovsky, Daniel Gruhl, Pablo N. Mendes |
EACL (1) | 1 |
| 2016 | Annotating and Predicting Non-Restrictive Noun Phrase ModificationsabstractThe distinction between restrictive and non-restrictive modification in noun phrases is a well studied subject in linguistics.Automatically identifying non-restrictive modifiers can provide NLP applications with shorter, more salient arguments, which were found beneficial by several recent works.While previous work showed that restrictiveness can be annotated with high agreement, no large scale corpus was created, hindering the development of suitable classification algorithms.In this work we devise a novel crowdsourcing annotation methodology, and an accompanying large scale corpus.Then, we present a robust automated system which identifies non-restrictive modifiers, notably improving over prior methods. Gabriel Stanovsky, Ido Dagan |
ACL (1) | 1 |
| 2016 | Modeling Extractive Sentence Intersection via Subtree EntailmentabstractSentence intersection captures the semantic overlap of two texts, generalizing over paradigms such as textual entailment and semantic text similarity. Despite its modeling power, it has received little attention because it is difficult for non-experts to annotate. We analyze 200 pairs of similar sentences and identify several underlying properties of sentence intersection. We leverage these insights to design an algorithm that decomposes the sentence intersection task into several simpler annotation tasks, facilitating the construction of a high quality dataset via crowdsourcing. We implement this approach and provide an annotated dataset of 1,764 sentence intersections. Omer Levy, Ido Dagan, Gabriel Stanovsky, Judith Eckle-Kohler, Iryna Gurevych |
COLING | 3 |
| 2016 | Porting an Open Information Extraction System from English to GermanabstractMany downstream NLP tasks can benefit from Open Information Extraction (Open IE) as a semantic representation.While Open IE systems are available for English, many other languages lack such tools.In this paper, we present a straightforward approach for adapting PropS, a rule-based predicate-argument analysis for English, to a new language, German.With this approach, we quickly obtain an Open IE system for German covering 89% of the English rule set.It yields 1.6 n-ary extractions per sentence at 60% precision, making it comparable to systems for English and readily usable in downstream applications.1 Tobias Falke, Gabriel Stanovsky, Iryna Gurevych, Ido Dagan |
EMNLP | 2 |
| 2016 | Creating a Large Benchmark for Open Information ExtractionabstractOpen information extraction (Open IE) was presented as an unrestricted variant of traditional information extraction.It has been gaining substantial attention, manifested by a large number of automatic Open IE extractors and downstream applications.In spite of this broad attention, the Open IE task definition has been lacking -there are no formal guidelines and no large scale gold standard annotation.Subsequently, the various implementations of Open IE resorted to small scale posthoc evaluations, inhibiting an objective and reproducible cross-system comparison.In this work, we develop a methodology that leverages the recent QA-SRL annotation to create a first independent and large scale Open IE annotation, 1 and use it to automatically compare the most prominent Open IE systems. Gabriel Stanovsky, Ido Dagan |
EMNLP | 1 |