Omri Abend

dblp:30/8159 · DBLP profile ↗
← Back
57ranked-venue papers
8as first author
28since 2021 · last 2026
0000-0003-4311-3876ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 57 · 8 first-author · 28 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Control Illusion: The Failure of Instruction Hierarchies in Large Language Models
abstract
Large language models (LLMs) are increasingly deployed with hierarchical instruction schemes, where certain instructions (e.g., system-level directives) are expected to take precedence over others (e.g., user messages). Yet, we lack a systematic understanding of how effectively these hierarchical control mechanisms work. We introduce a systematic evaluation framework based on constraint prioritization to assess how well LLMs enforce instruction hierarchies. Our experiments across six state-of-the-art LLMs reveal that models struggle with consistent instruction prioritization, even for simple formatting conflicts. We find that the widely-adopted system/user prompt separation fails to establish a reliable instruction hierarchy, and models exhibit strong inherent biases toward certain constraint types regardless of their priority designation. Interestingly, we also find that societal hierarchy framings (e.g., authority, expertise, consensus) show stronger influence on model behavior than system/user roles, suggesting that pretraining-derived social structures function as latent behavioral priors with potentially greater impact than post-training guardrails.
Yilin Geng 0001, Haonan Li 0002, Honglin Mu, Timothy Baldwin, Omri Abend, Eduard H. Hovy, Lea Frermann
AAAI6
2026 Mediocrity is the key for LLM as a Judge Anchor Selection
abstract
The "LLM-as-a-judge" paradigm has become a standard method for evaluating open-ended generation.To address the quadratic scalability costs of pairwise comparisons, popular benchmarks like Arena-Hard and AlpacaEval compare all models against a single anchor.However, despite its widespread use, the impact of anchor selection on the reliability of the results remains largely unexplored.In this work, we systematically investigate the effect of anchor selection by evaluating 22 different anchors on the Arena-Hard-v2.0 dataset.We find that the choice of anchor is critical: a poor anchor can dramatically reduce correlation with human rankings.We identify that common anchor choices (best-performing and worst-performing models) make poor anchors.Because these extreme anchors are consistently better or worse than all other models, they are seldom indicative of the relative ranking of the models.We further quantify the effect size of anchor selection, showing it is comparable to the selection of a judge model.We conclude with actionable recommendations.First, we conduct a power analysis, and compute sufficient benchmark sizes for anchor-based evaluation, finding that standard benchmark sizes are insufficient for pairwise evaluation and fail to distinguish between competitive models reliably.Second, we provide guidelines for selecting informative anchors to ensure reliable and efficient evaluation practices.
Shachar Don-Yehiya, Asaf Yehudai, Leshem Choshen, Omri Abend
ACL (1)4
2025 Computational Analysis of Character Development in Holocaust Testimonies
abstract
This work presents a computational approach to analyze character development along the narrative timeline.The analysis characterizes changes in the protagonist's views and behavior and the interplay between them.We consider transcripts of Holocaust survivor testimonies as a test case, each telling the story of an individual in first-person terms.We focus on the survivor's religious trajectory, examining the evolution of their disposition toward religious belief and practice as it is reflected in the testimony.Clustering the resulting trajectories in the dataset, we identify common sequences in the data.Our findings highlight multiple common structures of religiosity across the narratives: in terms of belief, a constant disposition is common, while for practice, most present an oscillating structure, serving as valuable material for historical and sociological research.This work demonstrates the potential of natural language processing for analyzing character evolution through thematic trajectories in narratives.
Esther Shizgal, Eitan Wagner, Renana Keydar, Omri Abend
EMNLP4
2025 A language-agnostic model of child language acquisition
abstract
This work reimplements a recent semantic bootstrapping child language acquisition (CLA) model, which was originally designed for English, and trains it to learn a new language: Hebrew. The model learns from pairs of utterances and logical forms as meaning representations, and acquires both syntax and word meanings simultaneously. The results show that the model mostly transfers to Hebrew, but that a number of factors, including the richer morphology in Hebrew, makes the learning slower and less robust. This suggests that a clear direction for future work is to enable the model to leverage the similarities between different word forms.
Louis Mahon, Omri Abend, Uri Berger, Katherine Demuth, Mark Johnson 0001, Mark Steedman
Comput. Speech Lang.2
2025 Surveying the Landscape of Image Captioning Evaluation: A Comprehensive Taxonomy, Trends, and Metrics Analysis
abstract
Abstract The task of image captioning has recently been gaining popularity, and with it the complex task of evaluating the quality of image captioning models. In this work, we present the first survey and taxonomy of over 70 different image captioning metrics and their usage in hundreds of papers, specifically designed to help users select the most suitable metric for their needs. We find that despite the diversity of proposed metrics, the vast majority of studies rely on only five popular metrics, which we show to be weakly correlated with human ratings. We hypothesize that combining a diverse set of metrics can enhance correlation with human ratings. As an initial step, we demonstrate that a linear regression-based ensemble method, which we call EnsembEval, trained on one human ratings dataset, achieves improved correlation across five additional datasets, showing there is a lot of room for improvement by leveraging a diverse set of metrics.1
Uri Berger, Gabriel Stanovsky, Omri Abend, Lea Frermann
Trans. Assoc. Comput. Linguistics3
2024 A Nurse is Blue and Elephant is Rugby: Cross Domain Alignment in Large Language Models Reveal Human-like Patterns
Asaf Yehudai, Taelin Karidi, Gabriel Stanovsky, Ariel Goldstein, Omri Abend
CogSci5
2024 Generating Benchmarks for Factuality Evaluation of Language Models
abstract
Dor Muhlgay, Ori Ram, Inbal Magar, Yoav Levine, Nir Ratner, Yonatan Belinkov, Omri Abend, Kevin Leyton-Brown, Amnon Shashua, Yoav Shoham. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Dor Muhlgay, Ori Ram, Inbal Magar, Yoav Levine, Nir Ratner, Yonatan Belinkov, Omri Abend, Kevin Leyton-Brown, Amnon Shashua, Yoav Shoham
EACL (1)7
2024 Exploring the Learning Capabilities of Language Models using LEVERWORLDS
abstract
Learning a model of a stochastic setting often involves learning both general structure rules and specific properties of the instance.This paper investigates the interplay between learning the general and the specific in various learning methods, with emphasis on sample efficiency.We design a framework called LEV-ERWORLDS, which allows the generation of simple physics-inspired worlds that follow a similar generative process with different distributions, and their instances can be expressed in natural language.These worlds allow for controlled experiments to assess the sample complexity of different learning methods.We experiment with classic learning algorithms as well as Transformer language models, both with fine-tuning and In-Context Learning (ICL).Our general finding is that (1) Transformers generally succeed in the task; but (2) they are considerably less sample efficient than classic methods that make stronger assumptions about the structure, such as Maximum Likelihood Estimation and Logistic Regression.This finding is in tension with the recent tendency to use Transformers as general-purpose estimators.We propose an approach that leverages the ICL capabilities of contemporary language models to apply simple algorithms for this type of data.Our experiments show that models currently struggle with the task but show promising potential.1
Eitan Wagner, Amir Feder, Omri Abend
EMNLP3
2024 CONTESTS: a Framework for Consistency Testing of Span Probabilities in Language Models
abstract
Although language model scores are often treated as probabilities, their reliability as probability estimators has mainly been studied through calibration, overlooking other aspects.In particular, it is unclear whether language models produce the same value for different ways of assigning joint probabilities to word spans.Our work introduces a novel framework, ConTestS (Consistency Testing over Spans), involving statistical tests to assess score consistency across interchangeable completion and conditioning orders.We conduct experiments on post-release real and synthetic data to eliminate training effects.Our findings reveal that both Masked Language Models (MLMs) and autoregressive models exhibit inconsistent predictions, with autoregressive models showing larger discrepancies.Larger MLMs tend to produce more consistent predictions, while autoregressive models show the opposite trend.Moreover, for both model types, prediction entropies offer insights into the true word span likelihood and therefore can aid in selecting optimal decoding strategies.The inconsistencies revealed by our analysis, as well their connection to prediction entropies and differences between model types, can serve as useful guides for future research on addressing these limitations.1
Eitan Wagner, Yuli Slavutsky, Omri Abend
EMNLP3
2023 DisentQA: Disentangling Parametric and Contextual Knowledge with Counterfactual Question Answering
abstract
Ella Neeman, Roee Aharoni, Or Honovich, Leshem Choshen, Idan Szpektor, Omri Abend. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Ella Neeman, Roee Aharoni, Or Honovich, Leshem Choshen, Idan Szpektor, Omri Abend
ACL (1)6
2023 Parallel Context Windows for Large Language Models
abstract
Nir Ratner, Yoav Levine, Yonatan Belinkov, Ori Ram, Inbal Magar, Omri Abend, Ehud Karpas, Amnon Shashua, Kevin Leyton-Brown, Yoav Shoham. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Nir Ratner, Yoav Levine, Yonatan Belinkov, Ori Ram, Inbal Magar, Omri Abend, Ehud Karpas, Amnon Shashua, Kevin Leyton-Brown, Yoav Shoham
ACL (1)6
2023 MuLER: Detailed and Scalable Reference-based Evaluation
abstract
We propose a novel methodology (namely, MuLER) that transforms any reference-based evaluation metric for text generation, such as machine translation (MT) into a fine-grained analysis tool.Given a system and a metric, MuLER quantifies how much the chosen metric penalizes specific error types (e.g., errors in translating names of locations).MuLER thus enables a detailed error analysis which can lead to targeted improvement efforts for specific phenomena.We perform experiments in both synthetic and naturalistic settings to support MuLER's validity and showcase its usability in MT evaluation, and other tasks, such as summarization.Analyzing all submissions to WMT in 2014-2020, we find consistent trends.For example, nouns and verbs are among the most frequent POS tags.However, they are among the hardest to translate.Performance on most POS tags improves with overall system performance, but a few are not thus correlated (their identity changes from language to language).Preliminary experiments with summarization reveal similar trends. 1
Taelin Karidi, Leshem Choshen, Gal Patel, Omri Abend
CoNLL4
2023 Evaluating and Improving the Coreference Capabilities of Machine Translation Models
abstract
Machine translation (MT) requires a wide range of linguistic capabilities, which current end-to-end models are expected to learn implicitly by observing aligned sentences in bilingual corpora.In this work, we ask: How well do MT models learn coreference resolution from implicit signal?To answer this question, we develop an evaluation methodology that derives coreference clusters from MT output and evaluates them without requiring annotations in the target language.We further evaluate several prominent open-source and commercial MT systems, translating from English to six target languages, and compare them to state-of-theart coreference resolvers on three challenging benchmarks.Our results show that the monolingual resolvers greatly outperform MT models.Motivated by this result, we experiment with different methods for incorporating the output of coreference resolution models in MT, showing improvement over strong baselines.1
Asaf Yehudai, Arie Cattan, Omri Abend, Gabriel Stanovsky
EACL3
2023 Human Learning by Model Feedback: The Dynamics of Iterative Prompting with Midjourney
abstract
Generating images with a Text-to-Image model often requires multiple trials, where human users iteratively update their prompt based on feedback, namely the output image.Taking inspiration from cognitive work on reference games and dialogue alignment, this paper analyzes the dynamics of the user prompts along such iterations.We compile a dataset of iterative interactions of human users with Midjourney. 1 Our analysis then reveals that prompts predictably converge toward specific traits along these iterations.We further study whether this convergence is due to human users, realizing they missed important details, or due to adaptation to the model's "preferences", producing better images for a specific language style.We show initial evidence that both possibilities are at play.The possibility that users adapt to the model's preference raises concerns about reusing user data for further training.The prompts may be biased towards the preferences of a specific model, rather than align with human intentions and natural manner of expression.
Shachar Don-Yehiya, Leshem Choshen, Omri Abend
EMNLP3
2023 Event-Location Tracking in Narratives: A Case Study on Holocaust Testimonies
abstract
This work focuses on the spatial dimension of narrative understanding and presents the task of event-location tracking in narrative texts, namely the extraction of the sequence of locations where the narrative is set.We present several architectures for the task that seek to model the global structure of the sequence, with varying levels of context awareness.We compare these methods to a number of strong baselines and ablated variants.We also develop methods for the generation of location embeddings and show that learning to predict a sequence of continuous embeddings is advantageous in terms of performance over predicting a string of locations.We focus on the test case of Holocaust survivor testimonies, motivated by the moral and historical importance of studying this dataset using computational means.The dataset further provides a unique case of a large set of narratives with a relatively restricted set of location trajectories.Our results show that models that are aware of the global context of the narrative can generate more accurate location chains.We corroborate the effectiveness of our methods by showing similar trends in an additional domain. 1
Eitan Wagner, Renana Keydar, Omri Abend
EMNLP3
2022 The Grammar-Learning Trajectories of Neural Language Models
abstract
The learning trajectories of linguistic phenomena in humans provide insight into linguistic representation, beyond what can be gleaned from inspecting the behavior of an adult speaker.To apply a similar approach to analyze neural language models (NLM), it is first necessary to establish that different models are similar enough in the generalizations they make.In this paper, we show that NLMs with different initialization, architecture, and training data acquire linguistic phenomena in a similar order, despite their different end performance.These findings suggest that there is some mutual inductive bias that underlies these models' learning of linguistic phenomena.Taking inspiration from psycholinguistics, we argue that studying this inductive bias is an opportunity to study the linguistic representation implicit in NLMs.Leveraging these findings, we compare the relative performance on different phenomena at varying learning stages with simpler reference models.Results suggest that NLMs exhibit consistent "developmental" stages.Moreover, we find the learning trajectory to be approximately one-dimensional: given an NLM with a certain overall performance, it is possible to predict what linguistic generalizations it has already acquired.Initial analysis of these stages presents phenomena clusters (notably morphological ones), whose performance progresses in unison, suggesting a potential link between the generalizations behind them.
Leshem Choshen, Guy Hacohen, Daphna Weinshall, Omri Abend
ACL (1)4
2022 Reinforcement Learning with Large Action Spaces for Neural Machine Translation
abstract
Applying Reinforcement learning (RL) following maximum likelihood estimation (MLE) pre-training is a versatile method for enhancing neural machine translation (NMT) performance. However, recent work has argued that the gains produced by RL for NMT are mostly due to promoting tokens that have already received a fairly high probability in pre-training. We hypothesize that the large action space is a main obstacle to RL’s effectiveness in MT, and conduct two sets of experiments that lend support to our hypothesis. First, we find that reducing the size of the vocabulary improves RL’s effectiveness. Second, we find that effectively reducing the dimension of the action space without changing the vocabulary also yields notable improvement as evaluated by BLEU, semantic similarity, and human evaluation. Indeed, by initializing the network’s final fully connected layer (that maps the network’s internal dimension to the vocabulary dimension), with a layer that generalizes over similar actions, we obtain a substantial improvement in RL performance: 1.5 BLEU points on average.
Asaf Yehudai, Leshem Choshen, Lior Fox, Omri Abend
COLING4
2022 Cognitive Simplification Operations Improve Text Simplification
abstract
Text Simplification (TS) is the task of converting a text into a form that is easier to read while maintaining the meaning of the original text.A sub-task of TS is Cognitive Simplification (CS), converting text to a form that is readily understood by people with cognitive disabilities without rendering it childish or simplistic.This sub-task has yet to be explored with neural methods in NLP, and resources for it are scarcely available.In this paper, we present a method for incorporating knowledge from the cognitive accessibility domain into a TS model, by introducing an inductive bias regarding what simplification operations to use.We show that by adding this inductive bias to a TS-trained model, it is able to adapt better to CS without ever seeing CS data, and outperform a baseline model on a traditional TS benchmark.In addition, we provide a novel test dataset for CS, and analyze the differences between CS corpora and existing TS corpora, in terms of how simplification operations are applied. Original Source:Now, normally during Disability Pride Month, we're showcasing our disability pride through various parades and events throughout the country.Original Target:Most years, during Disability Pride Month we have parades and events all over the United States to show how proud we are.Operations: Modified Source T5: Now, normally during Disability Pride Month, we're showcasing our disability pride through various parades and events throughout the country.Modified Target T5: Most years, during Disability Pride Month we have parades and events all over the United States to show how proud we are.Modified Source BART: Now, normally during Disability Pride Month, we're showcasing our disability pride through various parades and events throughout the country.Modified Target BART: Most years, during Disability Pride Month we
Eytan Chamovitz, Omri Abend
CoNLL2
2022 Enhancing the Transformer Decoder with Transition-based Syntax
abstract
Notwithstanding recent advances, syntactic generalization remains a challenge for text decoders.While some studies showed gains from incorporating source-side symbolic syntactic and semantic structure into text generation Transformers, very little work addressed the decoding of such structure.We propose a general approach for tree decoding using a transition-based approach.Examining the challenging test case of incorporating Universal Dependencies syntax into machine translation, we present substantial improvements on test sets that focus on syntactic generalization, while presenting improved or comparable performance on standard MT benchmarks.Further qualitative analysis addresses cases where syntactic generalization in the vanilla Transformer decoder is inadequate and demonstrates the advantages afforded by integrating syntactic information.1
Leshem Choshen, Omri Abend
CoNLL2
2022 On Neurons Invariant to Sentence Structural Changes in Neural Machine Translation
abstract
We present a methodology that explores how sentence structure is reflected in neural representations of machine translation systems.We demonstrate our model-agnostic approach with the Transformer English-German translation model.We analyze neuron-level correlation of activations between paraphrases while discussing the methodology challenges and the need for confound analysis to isolate the effects of shallow cues.We find that similarity between activation patterns can be mostly accounted for by similarity in word choice and sentence length.Following that, we manipulate neuron activations to control the syntactic form of the output.We show this intervention to be somewhat successful, indicating that deep models capture sentence-structure distinctions, despite finding no such indication at the neuron level.To conduct our experiments, we develop a semi-automatic method to generate meaning-preserving minimal pair paraphrases (active-passive voice and adverbial clause-noun phrase) and compile a corpus of such pairs.1
Gal Patel, Leshem Choshen, Omri Abend
CoNLL3
2022 PreQuEL: Quality Estimation of Machine Translation Outputs in Advance
abstract
We present the task of PreQuEL, Pre-(Quality-Estimation) Learning.A PreQuEL system predicts how well a given sentence will be translated, without recourse to the actual translation, thus eschewing unnecessary resource allocation when translation quality is bound to be low.PreQuEL can be defined relative to a given MT system (e.g., some industry service) or generally relative to the state-of-theart.From a theoretical perspective, PreQuEL places the focus on the source text, tracing properties, possibly linguistic features, that make a sentence harder to machine translate.We develop a baseline model for the task and analyze its performance.We also develop a data augmentation method (from parallel corpora), that improves results substantially.We show that this augmentation method can improve the performance of the Quality-Estimation task as well. 1 We investigate the properties of the input text that our model is sensitive to, by testing it on challenge sets and different languages.We conclude that it is aware of syntactic and semantic distinctions, and correlates and even over-emphasizes the importance of standard NLP features.cross-lingual representation learning at scale.
Shachar Don-Yehiya, Leshem Choshen, Omri Abend
EMNLP3
2022 Topical Segmentation of Spoken Narratives: A Test Case on Holocaust Survivor Testimonies
abstract
The task of topical segmentation is well studied, but previous work has mostly addressed it in the context of structured, well-defined segments, such as segmentation into paragraphs, chapters, or segmenting text that originated from multiple sources.We tackle the task of segmenting running (spoken) narratives, which poses hitherto unaddressed challenges.As a test case, we address Holocaust survivor testimonies, given in English.Other than the importance of studying these testimonies for Holocaust research, we argue that they provide an interesting test case for topical segmentation, due to their unstructured surface level, relative abundance (tens of thousands of such testimonies were collected), and the relatively confined domain that they cover.We hypothesize that boundary points between segments correspond to low mutual information between the sentences proceeding and following the boundary.Based on this hypothesis, we explore a range of algorithmic approaches to the task, building on previous work on segmentation that uses generative Bayesian modeling and state-of-the-art neural machinery.Compared to manually annotated references, we find that the developed approaches show considerable improvements over previous work.1
Eitan Wagner, Renana Keydar, Amit Pinchevski, Omri Abend
EMNLP4
2022 A Computational Acquisition Model for Multimodal Word Categorization
abstract
Uri Berger, Gabriel Stanovsky, Omri Abend, Lea Frermann. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Uri Berger, Gabriel Stanovsky, Omri Abend, Lea Frermann
NAACL-HLT3
2021 On the Relation between Syntactic Divergence and Zero-Shot Performance
abstract
We explore the link between the extent to which syntactic relations are preserved in translation and the ease of correctly constructing a parse tree in a zero-shot setting.While previous work suggests such a relation, it tends to focus on the macro level and not on the level of individual edges-a gap we aim to address.As a test case, we take the transfer of Universal Dependencies (UD) parsing from English to a diverse set of languages and conduct two sets of experiments.In one, we analyze zero-shot performance based on the extent to which English source edges are preserved in translation.In another, we apply three linguistically motivated transformations to UD, creating more cross-lingually stable versions of it, and assess their zero-shot parsability.In order to compare parsing performance across different schemes, we perform extrinsic evaluation on the downstream task of cross-lingual relation extraction (RE) using a subset of a popular English RE benchmark translated to Russian and Korean. 1 In both sets of experiments, our results suggest a strong relation between cross-lingual stability and zero-shot parsing performance.
Ofir Arviv, Dmitry Nikolaev 0002, Taelin Karidi, Omri Abend
EMNLP (1)4
2021 $Q^2$: Evaluating Factual Consistency in Knowledge-Grounded Dialogues via Question Generation and Question Answering
abstract
Neural knowledge-grounded generative models for dialogue often produce content that is factually inconsistent with the knowledge they rely on, making them unreliable and limiting their applicability.Inspired by recent work on evaluating factual consistency in abstractive summarization, we propose an automatic evaluation metric for factual consistency in knowledge-grounded dialogue using automatic question generation and question answering.Our metric, denoted Q 2 , compares answer spans using natural language inference (NLI), instead of token-based matching as done in previous work.To foster proper evaluation, we curate a novel dataset of dialogue system outputs for the Wizard-of-Wikipedia dataset, manually annotated for factual consistency.We perform a thorough meta-evaluation of Q 2 against other metrics using this dataset and two others, where it consistently shows higher correlation with human judgements.
Or Honovich, Leshem Choshen, Roee Aharoni, Ella Neeman, Idan Szpektor, Omri Abend
EMNLP (1)6
2021 Putting Words in BERT's Mouth: Navigating Contextualized Vector Spaces with Pseudowords
abstract
We present a method for exploring regions around individual points in a contextualized vector space (particularly, BERT space), as a way to investigate how these regions correspond to word senses.By inducing a contextualized "pseudoword" as a stand-in for a static embedding in the input layer, and then performing masked prediction of a word in the sentence, we are able to investigate the geometry of the BERT-space in a controlled manner around individual instances.Using our method on a set of carefully constructed sentences targeting ambiguous English words, we find substantial regularity in the contextualized space, with regions that correspond to distinct word senses; but between these regions there are occasionally "sense voids"-regions that do not correspond to any intelligible sense. 1Learn pseudoword in place of that is customized to reconstruct .
Taelin Karidi, Yichu Zhou, Nathan Schneider 0001, Omri Abend, Vivek Srikumar
EMNLP (1)4
2021 PMI-Masking: Principled masking of correlated spans
Yoav Levine, Barak Lenz, Opher Lieber, Omri Abend, Kevin Leyton-Brown, Moshe Tennenholtz, Yoav Shoham
ICLR4
2021 Mediators in Determining what Processing BERT Performs First
abstract
Probing neural models for the ability to perform downstream tasks using their activation patterns is often used to localize what parts of the network specialize in performing what tasks.However, little work addressed potential mediating factors in such comparisons.As a test-case mediating factor, we consider the prediction's context length, namely the length of the span whose processing is minimally required to perform the prediction.We show that not controlling for context length may lead to contradictory conclusions as to the localization patterns of the network, depending on the distribution of the probing dataset.Indeed, when probing BERT with seven tasks, we find that it is possible to get 196 different rankings between them when manipulating the distribution of context lengths in the probing dataset.We conclude by presenting best practices for conducting such comparisons in the future. 1
Aviv Slobodkin, Leshem Choshen, Omri Abend
NAACL-HLT3
2020 Machine Reading of Historical Events
abstract
Machine reading is an ambitious goal in NLP that subsumes a wide range of text understanding capabilities.Within this broad framework, we address the task of machine reading the time of historical events, compile datasets for the task, and develop a model for tackling it.Given a brief textual description of an event, we show that good performance can be achieved by extracting relevant sentences from Wikipedia, and applying a combination of taskspecific and general-purpose feature embeddings for the classification.Furthermore, we establish a link between the historical event ordering task and the event focus time task from the information retrieval literature, showing they also provide a challenging test case for machine reading algorithms. 1 1 Code and data are available at https://github.com/ltorroba/ machine-reading-historical-events.* Equal contribution.Year Event text OTD 2005 107 die in Amagasaki rail crash in Japan.1939 BMI (Broadcast Music Incorporated) formed.1864 General Sherman's armies reach Savannah & 12 day siege begins.WOTD 1887 Buffalo Bill Cody's Wild West Show opens in London.1399 Henry IV is proclaimed King of England.1943 First Flight of the Gloster Meteor, Britain's first combat jet aircraft.
Or Honovich, Lucas Torroba Hennigen, Omri Abend, Shay B. Cohen
ACL3
2020 Fine-Grained Analysis of Cross-Linguistic Syntactic Divergences
abstract
Dmitry Nikolaev, Ofir Arviv, Taelin Karidi, Neta Kenneth, Veronika Mitnik, Lilja Maria Saeboe, Omri Abend. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020.
Dmitry Nikolaev 0002, Ofir Arviv, Taelin Karidi, Neta Kenneth, Veronika Mitnik, Lilja Maria Saeboe, Omri Abend
ACL7
2020 Language (Re)modelling: Towards Embodied Language Understanding
abstract
While natural language understanding (NLU) is advancing rapidly, today's technology differs from human-like language understanding in fundamental ways, notably in its inferior efficiency, interpretability, and generalization.This work proposes an approach to representation and learning based on the tenets of embodied cognitive linguistics (ECL).According to ECL, natural language is inherently executable (like programming languages), driven by mental simulation and metaphoric mappings over hierarchical compositions of structures and schemata learned through embodied interaction.This position paper argues that the use of grounding by metaphoric inference and simulation will greatly benefit NLU systems, and proposes a system architecture along with a roadmap towards realizing this vision.
Ronen Tamari, Chen Shani, Tom Hope, Miriam R. L. Petruck, Omri Abend, Dafna Shahaf
ACL5
2020 Comparison by Conversion: Reverse-Engineering UCCA from Syntax and Lexical Semantics
abstract
Building robust natural language understanding systems will require a clear characterization of whether and how various linguistic meaning representations complement each other.To perform a systematic comparative analysis, we evaluate the mapping between meaning representations from different frameworks using two complementary methods: (i) a rule-based converter, and (ii) a supervised delexicalized parser that parses to one framework using only information from the other as features.We apply these methods to convert the STREUSLE corpus (with syntactic and lexical semantic annotations) to UCCA (a graph-structured full-sentence meaning representation).Both methods yield surprisingly accurate target representations, close to fully supervised UCCA parser quality-indicating that UCCA annotations are partially redundant with STREUSLE annotations.Despite this substantial convergence between frameworks, we find several important areas of divergence.
Daniel Hershcovich, Nathan Schneider 0001, Dotan Dvir, Jakob Prange, Miryam de Lhoneux, Omri Abend
COLING6
2020 Classifying Syntactic Errors in Learner Language
abstract
We present a method for classifying syntactic errors in learner language, namely errors whose correction alters the morphosyntactic structure of a sentence.The methodology builds on the established Universal Dependencies syntactic representation scheme, and provides complementary information to other error-classification systems.Unlike existing error classification methods, our method is applicable across languages, which we showcase by producing a detailed picture of syntactic errors in learner English and learner Russian.We further demonstrate the utility of the methodology for analyzing the outputs of leading Grammatical Error Correction (GEC) systems.* First two authors contributed equally. 1 Code can be found in github repo GEC UD divergences.Matrices directly mentioned are included in the appendix.
Leshem Choshen, Dmitry Nikolaev 0002, Yevgeni Berzak, Omri Abend
CoNLL4
2020 On the Weaknesses of Reinforcement Learning for Neural Machine Translation
Leshem Choshen, Lior Fox, Zohar Aizenbud, Omri Abend
ICLR4
2019 The Language of Legal and Illegal Activity on the Darknet
abstract
The non-indexed parts of the Internet (the Darknet) have become a haven for both legal and illegal anonymous activity.Given the magnitude of these networks, scalably monitoring their activity necessarily relies on automated tools, and notably on NLP tools.However, little is known about what characteristics texts communicated through the Darknet have, and how well off-the-shelf NLP tools do on this domain.This paper tackles this gap and performs an in-depth investigation of the characteristics of legal and illegal text in the Darknet, comparing it to a clear net website with similar content as a control condition.Taking drug-related websites as a test case, we find that texts for selling legal and illegal drugs have several linguistic characteristics that distinguish them from one another, as well as from the control condition, among them the distribution of POS tags, and the coverage of their named entities in Wikipedia. 1
Leshem Choshen, Dan Eldad, Daniel Hershcovich, Elior Sulem, Omri Abend
ACL (1)5
2019 Automatically Extracting Challenge Sets for Non-Local Phenomena in Neural Machine Translation
abstract
We show that the state-of-the-art Transformer MT model is not biased towards monotonic reordering (unlike previous recurrent neural network models), but that nevertheless, longdistance dependencies remain a challenge for the model.Since most dependencies are shortdistance, common evaluation metrics will be little influenced by how well systems perform on them.We therefore propose an automatic approach for extracting challenge sets replete with long-distance dependencies, and argue that evaluation using this methodology provides a complementary perspective on system performance.To support our claim, we compile challenge sets for English-German and German-English, which are much larger than any previously released challenge set for MT.The extracted sets are large enough to allow reliable automatic evaluation, which makes the proposed approach a scalable and practical solution for evaluating MT performance on the long-tail of syntactic phenomena.1
Leshem Choshen, Omri Abend
CoNLL2
2019 Made for Each Other: Broad-Coverage Semantic Structures Meet Preposition Supersenses
abstract
Universal Conceptual Cognitive Annotation (UCCA; Abend and Rappoport, 2013) is a typologically-informed, broad-coverage semantic annotation scheme that describes coarse-grained predicate-argument structure but currently lacks semantic roles. We argue that lexicon-free annotation of the semantic roles marked by prepositions, as formulated by Schneider et al. (2018), is complementary and suitable for integration within UCCA. We show empirically for English that the schemes, though annotated independently, are compatible and can be combined in a single semantic graph. A comparison of several approaches to parsing the integrated representation lays the groundwork for future research on this task.
Jakob Prange, Nathan Schneider 0001, Omri Abend
CoNLL3
2018 Inherent Biases in Reference-based Evaluation for Grammatical Error Correction
abstract
The prevalent use of too few references for evaluating text-to-text generation is known to bias estimates of their quality (henceforth, low coverage bias or LCB).This paper shows that overcoming LCB in Grammatical Error Correction (GEC) evaluation cannot be attained by re-scaling or by increasing the number of references in any feasible range, contrary to previous suggestions.This is due to the long-tailed distribution of valid corrections for a sentence.Concretely, we show that LCB incentivizes GEC systems to avoid correcting even when they can generate a valid correction.Consequently, existing systems obtain comparable or superior performance compared to humans, by making few but targeted changes to the input.Similar effects on Text Simplification further support our claims.
Leshem Choshen, Omri Abend
ACL (1)2
2018 Automatic Metric Validation for Grammatical Error Correction
abstract
Metric validation in Grammatical Error Correction (GEC) is currently done by observing the correlation between human and metric-induced rankings. However, such correlation studies are costly, methodologically troublesome, and suffer from low inter-rater agreement. We propose MAEGE, an automatic methodology for GEC metric validation, that overcomes many of the difficulties in the existing methodology. Experiments with MAEGE shed a new light on metric quality, showing for example that the standard M^2 metric fares poorly on corpus-level ranking. Moreover, we use MAEGE to perform a detailed analysis of metric behavior, showing that some types of valid edits are consistently penalized by existing metrics.
Leshem Choshen, Omri Abend
ACL (1)2
2018 Multitask Parsing Across Semantic Representations
abstract
The ability to consolidate information of different types is at the core of intelligence, and has tremendous practical value in allowing learning for one task to benefit from generalizations learned for others.In this paper we tackle the challenging task of improving semantic parsing performance, taking UCCA parsing as a test case, and AMR, SDP and Universal Dependencies (UD) parsing as auxiliary tasks.We experiment on three languages, using a uniform transition-based system and learning architecture for all parsing tasks.Despite notable conceptual, formal and domain differences, we show that multitask learning significantly improves UCCA parsing in both in-domain and out-of-domain settings.Our code is publicly available.
Daniel Hershcovich, Omri Abend, Ari Rappoport
ACL (1)2
2018 Simple and Effective Text Simplification Using Semantic and Neural Methods
abstract
Sentence splitting is a major simplification operator.Here we present a simple and efficient splitting algorithm based on an automatic semantic parser.After splitting, the text is amenable for further fine-tuned simplification operations.In particular, we show that neural Machine Translation can be effectively used in this situation.Previous application of Machine Translation for simplification suffers from a considerable disadvantage in that they are overconservative, often failing to modify the source in any way.Splitting based on semantic parsing, as proposed here, alleviates this issue.Extensive automatic and human evaluation shows that the proposed method compares favorably to the stateof-the-art in combined lexical and structural simplification.
Elior Sulem, Omri Abend, Ari Rappoport
ACL (1)2
2018 Comprehensive Supersense Disambiguation of English Prepositions and Possessives
abstract
Nathan Schneider, Jena D. Hwang, Vivek Srikumar, Jakob Prange, Austin Blodgett, Sarah R. Moeller, Aviram Stern, Adi Bitan, Omri Abend. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018.
Nathan Schneider 0001, Jena D. Hwang, Vivek Srikumar, Jakob Prange, Austin Blodgett, Sarah R. Moeller, Aviram Stern, Adi Bitan, Omri Abend
ACL (1)9
2018 BLEU is Not Suitable for the Evaluation of Text Simplification
abstract
BLEU is widely considered to be an informative metric for text-to-text generation, including Text Simplification (TS).TS includes both lexical and structural aspects.In this paper we show that BLEU is not suitable for the evaluation of sentence splitting, the major structural simplification operation.We manually compiled a sentence splitting gold standard corpus containing multiple structural paraphrases, and performed a correlation analysis with human judgments.1 We find low or no correlation between BLEU and the grammaticality and meaning preservation parameters where sentence splitting is involved.Moreover, BLEU often negatively correlates with simplicity, essentially penalizing simpler sentences.
Elior Sulem, Omri Abend, Ari Rappoport
EMNLP2
2018 Semantic Structural Evaluation for Text Simplification
abstract
Elior Sulem, Omri Abend, Ari Rappoport. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Elior Sulem, Omri Abend, Ari Rappoport
NAACL-HLT2
2017 The State of the Art in Semantic Representation
abstract
Semantic representation is receiving growing attention in NLP in the past few years, and many proposals for semantic schemes (e.g., AMR, UCCA, GMB, UDS) have been put forth.Yet, little has been done to assess the achievements and the shortcomings of these new contenders, compare them with syntactic schemes, and clarify the general goals of research on semantic representation.We address these gaps by critically surveying the state of the art in the field.
Omri Abend, Ari Rappoport
ACL (1)1
2017 A Transition-Based Directed Acyclic Graph Parser for UCCA
abstract
We present the first parser for UCCA, a cross-linguistically applicable framework for semantic representation, which builds on extensive typological work and supports rapid annotation.UCCA poses a challenge for existing parsing techniques, as it exhibits reentrancy (resulting in DAG structures), discontinuous structures and non-terminal nodes corresponding to complex semantic units.To our knowledge, the conjunction of these formal properties is not supported by any existing parser.Our transition-based parser, which uses a novel transition set and features based on bidirectional LSTMs, has value not just for UCCA parsing: its ability to handle more general graph structures can inform the development of parsers for other semantic DAG structures, and in languages that frequently use discontinuous structures.
Daniel Hershcovich, Omri Abend, Ari Rappoport
ACL (1)2
2016 HUME: Human UCCA-Based Evaluation of Machine Translation
abstract
Human evaluation of machine translation normally uses sentence-level measures such as relative ranking or adequacy scales. However, these provide no insight into possible errors, and do not scale well with sentence length. We argue for a semantics-based evaluation, which captures what meaning components are retained in the MT output, thus providing a more fine-grained analysis of translation quality, and enabling the construction and tuning of semantics-based MT. We present a novel human semantic evaluation measure, Human UCCA-based MT Evaluation (HUME), building on the UCCA semantic representation scheme. HUME covers a wider range of semantic phenomena than previous methods and does not rely on semantic annotation of the potentially garbled MT output. We experiment with four language pairs, demonstrating HUME’s broad applicability, and report good inter-annotator agreement rates and correlation with human adequacy scores.
Alexandra Birch, Omri Abend, Ondrej Bojar, Barry Haddow
EMNLP2
2015 Lexical Event Ordering with an Edge-Factored Model
abstract
Extensive lexical knowledge is necessary for temporal analysis and planning tasks.We address in this paper a lexical setting that allows for the straightforward incorporation of rich features and structural constraints.We explore a lexical event ordering task, namely determining the likely temporal order of events based solely on the identity of their predicates and arguments.We propose an "edgefactored" model for the task that decomposes over the edges of the event graph.We learn it using the structured perceptron.As lexical tasks require large amounts of text, we do not attempt manual annotation and instead use the textual order of events in a domain where this order is aligned with their temporal order, namely cooking recipes.
Omri Abend, Shay B. Cohen, Mark Steedman
HLT-NAACL1
2014 Lexical Inference over Multi-Word Predicates: A Distributional Approach
abstract
Representing predicates in terms of their argument distribution is common practice in NLP.Multi-word predicates (MWPs) in this context are often either disregarded or considered as fixed expressions.The latter treatment is unsatisfactory in two ways: (1) identifying MWPs is notoriously difficult, (2) MWPs show varying degrees of compositionality and could benefit from taking into account the identity of their component parts.We propose a novel approach that integrates the distributional representation of multiple sub-sets of the MWP's words.We assume a latent distribution over sub-sets of the MWP, and estimate it relative to a downstream prediction task.Focusing on the supervised identification of lexical inference relations, we compare against state-of-the-art baselines that consider a single sub-set of an MWP, obtaining substantial improvements.To our knowledge, this is the first work to address lexical relations between MWPs of varying degrees of compositionality within distributional semantics.
Omri Abend, Shay B. Cohen, Mark Steedman
ACL (1)1
2013 Universal Conceptual Cognitive Annotation (UCCA)
Omri Abend, Ari Rappoport
ACL (1)1
2012 Learnability-Based Syntactic Annotation Design
Roy Schwartz 0001, Omri Abend, Ari Rappoport
COLING2
2011 Neutralizing Linguistically Problematic Annotations in Unsupervised Dependency Parsing Evaluation
Roy Schwartz 0001, Omri Abend, Roi Reichart, Ari Rappoport
ACL2
2010 Fully Unsupervised Core-Adjunct Argument Classification
Omri Abend, Ari Rappoport
ACL1
2010 Improved Unsupervised POS Induction through Prototype Discovery
Omri Abend, Roi Reichart, Ari Rappoport
ACL1
2010 Type Level Clustering Evaluation: New Measures and a POS Induction Case Study
Roi Reichart, Omri Abend, Ari Rappoport
CoNLL2
2009 Unsupervised Argument Identification for Semantic Role Labeling
Omri Abend, Roi Reichart, Ari Rappoport
ACL/IJCNLP1
2008 A Supervised Algorithm for Verb Disambiguation into VerbNet Classes
Omri Abend, Roi Reichart, Ari Rappoport
COLING1