EDBT 2026 Demo / reviewers in the wild / expert
Roi Reichart
dblp:96/5429
· DBLP profile ↗
108ranked-venue papers
14as first author
35since 2021 · last 2026
0000-0001-6918-0554ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 107 · 14 first-author · 35 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Databases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 2Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Can LLMs Replace Economic Choice Prediction Labs? The Case of Language-based Persuasion GamesabstractHuman choice prediction in economic contexts is crucial for applications in marketing, finance, public policy, and more. This task, however, is often constrained by the difficulties in acquiring human choice data. With most experimental economics studies focusing on simple choice settings, the AI community has explored whether LLMs can substitute for humans in these predictions and examined more complex experimental economics settings. However, a key question remains: can LLMs generate training data for human choice prediction? We explore this in language-based persuasion games, a complex economic setting involving natural language in strategic interactions. Our experiments show that models trained on LLM-generated data can effectively predict human behavior in these games and even outperform models trained on actual human data. Beyond data generation, we investigate the dual role of LLMs as both data generators and predictors, introducing a comprehensive empirical study on the effectiveness of utilizing LLMs for data generation, human choice prediction, or both. We then utilize our choice prediction framework to analyze how strategic factors shape decision-making, showing that interaction history (rather than linguistic sentiment alone) plays a key role in predicting human decision-making in repeated interactions. Particularly, when LLMs capture history-dependent decision patterns similarly to humans, their predictive success improves substantially. Finally, we demonstrate the robustness of our findings across alternative persuasion-game settings, highlighting the broader potential of using LLM-generated data to model human decision-making. Eilam Shapira, Omer Madmon, Roi Reichart, Moshe Tennenholtz |
J. Artif. Intell. Res. | 3 |
| 2025 | The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMsabstractThe “LLM-as-an-annotator” and “LLM-as-a-judge” paradigms employ Large Language Models (LLMs) as annotators, judges, and evaluators in tasks traditionally performed by humans. LLM annotations are widely used, not only in NLP research but also in fields like medicine, psychology, and social science. Despite their role in shaping study results and insights, there is no standard or rigorous procedure to determine whether LLMs can replace human annotators. In this paper, we propose a novel statistical procedure, the Alternative Annotator Test (alt-test), that requires only a modest subset of annotated examples to justify using LLM annotations. Additionally, we introduce a versatile and interpretable measure for comparing LLM annotators and judges. To demonstrate our procedure, we curated a diverse collection of ten datasets, consisting of language and vision-language tasks, and conducted experiments with six LLMs and four prompting techniques. Our results show that LLMs can sometimes replace humans with closed-source LLMs (such as GPT-4o), outperforming the open-source LLMs we examine, and that prompting techniques yield judges of varying quality. We hope this study encourages more rigorous and reliable practices. Nitay Calderon, Roi Reichart, Rotem Dror |
ACL (1) | 2 |
| 2025 | Multi-Domain Explainability of PreferencesabstractPreference mechanisms, such as human preference, LLM-as-a-Judge (LaaJ), and reward models, are central to aligning and evaluating large language models (LLMs).Yet, the underlying concepts that drive these preferences remain poorly understood.In this work, we propose a fully automated method for generating local and global concept-based explanations of preferences across multiple domains.Our method utilizes an LLM to identify concepts (rubrics) that distinguish between chosen and rejected responses, and to represent them with conceptbased vectors.To model the relationships between concepts and preferences, we propose a white-box Hierarchical Multi-Domain Regression model that captures both domain-general and domain-specific effects.To evaluate our method, we curate a dataset spanning eight diverse domains and explain twelve mechanisms.Our method achieves strong preference prediction performance, outperforming baselines while also being explainable.Additionally, we assess explanations in two application-driven settings.First, guiding LLM outputs with concepts from LaaJ explanations yields responses that those judges consistently prefer.Second, prompting LaaJs with concepts explaining humans improves their preference predictions.Together, our work establishes a new paradigm for explainability in the era of LLMs. 1 Nitay Calderon, Liat Ein-Dor, Roi Reichart |
EMNLP | 3 |
| 2025 | Are LLMs Better than Reported? Detecting Label Errors and Mitigating Their Effect on Model PerformanceabstractNLP benchmarks rely on standardized datasets for training and evaluating models and are crucial for advancing the field.Traditionally, expert annotations ensure high-quality labels; however, the cost of expert annotation does not scale well with the growing demand for larger datasets required by modern models.While crowd-sourcing provides a more scalable solution, it often comes at the expense of annotation precision and consistency.Recent advancements in large language models (LLMs) offer new opportunities to enhance the annotation process, particularly for detecting label errors in existing datasets.In this work, we consider the recent approach of LLM-as-a-judge, leveraging an ensemble of LLMs to flag potentially mislabeled examples.We conduct a case study on four factual consistency datasets from the TRUE benchmark, spanning diverse NLP tasks, and on SummEval, which uses Likertscale ratings of summary quality across multiple dimensions.We empirically analyze the labeling quality of existing datasets and compare expert, crowd-sourced, and LLM-based annotations in terms of the agreement, label quality, and efficiency, demonstrating the strengths and limitations of each annotation method.Our findings reveal a substantial number of label errors, which, when corrected, induce a significant upward shift in reported model performance.This suggests that many of the LLMs' so-called mistakes are due to label errors rather than genuine model failures.Additionally, we discuss the implications of mislabeled data and propose methods to mitigate them in training to improve performance. Omer Nahum, Nitay Calderon, Orgad Keller, Idan Szpektor, Roi Reichart |
EMNLP | 5 |
| 2025 | LLMs Know More Than They Show: On the Intrinsic Representation of LLM HallucinationsabstractLarge language models (LLMs) often produce errors, including factual inaccuracies, biases, and reasoning failures, collectively referred to as "hallucinations". Recent studies have demonstrated that LLMs' internal states encode information regarding the truthfulness of their outputs, and that this information can be utilized to detect errors. In this work, we show that the internal representations of LLMs encode much more information about truthfulness than previously recognized. We first discover that the truthfulness information is concentrated in specific tokens, and leveraging this property significantly enhances error detection performance. Yet, we show that such error detectors fail to generalize across datasets, implying that---contrary to prior claims---truthfulness encoding is not universal but rather multifaceted. Next, we show that internal representations can also be used for predicting the types of errors the model is likely to make, facilitating the development of tailored mitigation strategies. Lastly, we reveal a discrepancy between LLMs' internal encoding and external behavior: they may encode the correct answer, yet consistently generate an incorrect one. Taken together, these insights deepen our understanding of LLM errors from the model's internal perspective, which can guide future research on enhancing error analysis and mitigation. Hadas Orgad, Michael Toker, Zorik Gekhman, Roi Reichart, Idan Szpektor, Hadas Kotek, Yonatan Belinkov |
ICLR | 4 |
| 2025 | NL-Eye: Abductive NLI For ImagesabstractWill a Visual Language Model (VLM)-based bot warn us about slipping if it detects a wet floor? Recent VLMs have demonstrated impressive capabilities, yet their ability to infer outcomes and causes remains underexplored. To address this, we introduce NL-Eye, a benchmark designed to assess VLMs' visual abductive reasoning skills. NL-Eye adapts the abductive Natural Language Inference (NLI) task to the visual domain, requiring models to evaluate the plausibility of hypothesis images based on a premise image and explain their decisions. NL-Eye consists of 350 carefully curated triplet examples (1,050 images) spanning diverse reasoning categories: physical, functional, logical, emotional, cultural, and social. The data curation process involved two steps—writing textual descriptions and generating images using text-to-image models, both requiring substantial human involvement to ensure high-quality and challenging scenes. Our experiments show that VLMs struggle significantly on NL-Eye, often performing at random baseline levels, while humans excel in both plausibility prediction and explanation quality. This demonstrates a deficiency in the abductive reasoning capabilities of modern VLMs. NL-Eye represents a crucial step toward developing VLMs capable of robust multimodal reasoning for real-world applications, including accident-prevention bots and generated video verification. Mor Ventura, Michael Toker, Nitay Calderon, Zorik Gekhman, Yonatan Bitton, Roi Reichart |
ICLR | 6 |
| 2025 | On Behalf of the Stakeholders: Trends in NLP Model Interpretability in the Era of LLMsabstractNitay Calderon, Roi Reichart. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Nitay Calderon, Roi Reichart |
NAACL (Long Papers) | 2 |
| 2025 | TabSTAR: A Tabular Foundation Model for Tabular Data with Text FieldsabstractWhile deep learning has achieved remarkable success across many domains, it has historically underperformed on tabular learning tasks, which remain dominated by gradient boosting decision trees. However, recent advancements are paving the way for Tabular Foundation Models, which can leverage real-world knowledge and generalize across diverse datasets, particularly when the data contains free-text. Although incorporating language model capabilities into tabular tasks has been explored, most existing methods utilize static, target-agnostic textual representations, limiting their effectiveness. We introduce TabSTAR: a Tabular Foundation Model with Semantically Target-Aware Representations. TabSTAR is designed to enable transfer learning on tabular data with textual features, with an architecture free of dataset-specific parameters. It unfreezes a pretrained text encoder and takes as input target tokens, which provide the model with the context needed to learn task-specific embeddings. TabSTAR achieves state-of-the-art performance for both medium- and large-sized datasets across known benchmarks of classification tasks with text features, and its pretraining phase exhibits scaling laws in the number of datasets, offering a pathway for further performance improvements. Alan Arazi, Eilam Shapira, Roi Reichart |
NeurIPS | 3 |
| 2025 | A Systematic Review of NLP for Dementia: Tasks, Datasets, and OpportunitiesabstractAbstract The close link between cognitive decline and language has fostered long-standing collaboration between the NLP and medical communities in dementia research. To examine this, we reviewed over 240 papers applying NLP to dementia-related efforts, drawing from medical, technological, and NLP-focused literature. We identify key research areas, including dementia detection, linguistic biomarker extraction, caregiver support, and patient assistance, showing that half of all papers focus solely on dementia detection using clinical data. Yet, many directions remain unexplored, such as artificially degraded language models, synthetic data, digital twins, and more. We highlight gaps and opportunities around trust, scientific rigor, applicability, and cross-community collaboration. We raise ethical dilemmas in the field, and highlight the diverse datasets encountered throughout our review including recorded, written, structured, spontaneous, synthetic, clinical, social media–based, and more. This review aims to inspire more creative, impactful, and rigorous research on NLP for dementia. Lotem Peled-Cohen, Roi Reichart |
Trans. Assoc. Comput. Linguistics | 2 |
| 2025 | Human Choice Prediction in Language-based Persuasion Games: Simulation-based Off-Policy EvaluationabstractAbstract Recent advances in Large Language Models (LLMs) have spurred interest in designing LLM-based agents for tasks that involve interaction with human and artificial agents. This paper addresses a key aspect in the design of such agents: predicting human decisions in off-policy evaluation (OPE). We focus on language-based persuasion games, where an expert aims to influence the decision-maker through verbal messages. In our OPE framework, the prediction model is trained on human interaction data collected from encounters with one set of expert agents, and its performance is evaluated on interactions with a different set of experts. Using a dedicated application, we collected a dataset of 87K decisions from humans playing a repeated decision-making game with artificial agents. To enhance off-policy performance, we propose a simulation technique involving interactions across the entire agent space and simulated decision-makers. Our learning strategy yields significant OPE gains, e.g., improving prediction accuracy in the top 15% challenging cases by 7.1%.1 Eilam Shapira, Omer Madmon, Reut Apel, Moshe Tennenholtz, Roi Reichart |
Trans. Assoc. Comput. Linguistics | 5 |
| 2025 | Navigating Cultural Chasms: Exploring and Unlocking the Cultural POV of Text-To-Image ModelsabstractAbstract Text-To-Image (TTI) models, such as DALL-E and StableDiffusion, have demonstrated remarkable prompt-based image generation capabilities. Multilingual encoders may have a substantial impact on the cultural agency of these models, as language is a conduit of culture. In this study, we explore the cultural perception embedded in TTI models by characterizing culture across three tiers: cultural dimensions, cultural domains, and cultural concepts. Based on this ontology, we derive prompt templates to unlock the cultural knowledge in TTI models, and propose a comprehensive suite of evaluation techniques, including intrinsic evaluations using the CLIP space, extrinsic evaluations with a Visual-Question-Answer models and human assessments, to evaluate the cultural content of TTI-generated images. To bolster our research, we introduce the CulText2I dataset, based on six diverse TTI models and spanning ten languages. Our experiments provide insights regarding Do, What, Which, and How research questions about the nature of cultural encoding in TTI models, paving the way for cross-cultural applications of these models.1 Mor Ventura, Eyal Ben-David, Anna Korhonen, Roi Reichart |
Trans. Assoc. Comput. Linguistics | 4 |
| 2024 | Your Prompt Is My Command: On Assessing the Human-Centred Generality of Multimodal Models (Abstract Reprint)abstractEven with obvious deficiencies, large prompt-commanded multimodal models are proving to be flexible cognitive tools representing an unprecedented generality. But the directness, diversity, and degree of user interaction create a distinctive “human-centred generality” (HCG), rather than a fully autonomous one. HCG implies that —for a specific user— a system is only as general as it is effective for the user’s relevant tasks and their prevalent ways of prompting. A human-centred evaluation of general-purpose AI systems therefore needs to reflect the personal nature of interaction, tasks and cognition. We argue that the best way to understand these systems is as highly-coupled cognitive extenders, and to analyse the bidirectional cognitive adaptations between them and humans. In this paper, we give a formulation of HCG, as well as a high-level overview of the elements and trade-offs involved in the prompting process. We end the paper by outlining some essential research questions and suggestions for improving evaluation practices, which we envision as characteristic for the evaluation of general artificial intelligence in the future. Wout Schellaert, Fernando Martínez-Plumed, Karina Vold, John Burden, P. A. M. Casares, Bao Sheng Loe, Roi Reichart, Seán Ó hÉigeartaigh, Anna Korhonen, José Hernández-Orallo |
AAAI | 7 |
| 2024 | Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?abstractWhen large language models are aligned via supervised fine-tuning, they may encounter new factual information that was not acquired through pre-training.It is often conjectured that this can teach the model the behavior of hallucinating factually incorrect responses, as the model is trained to generate facts that are not grounded in its pre-existing knowledge.In this work, we study the impact of such exposure to new knowledge on the capability of the fine-tuned model to utilize its pre-existing knowledge.To this end, we design a controlled setup, focused on closedbook QA, where we vary the proportion of the fine-tuning examples that introduce new knowledge.We demonstrate that large language models struggle to acquire new factual knowledge through fine-tuning, as fine-tuning examples that introduce new knowledge are learned significantly slower than those consistent with the model's knowledge.However, we also find that as the examples with new knowledge are eventually learned, they linearly increase the model's tendency to hallucinate.Taken together, our results highlight the risk in introducing new factual knowledge through fine-tuning, and support the view that large language models mostly acquire factual knowledge through pre-training, whereas finetuning teaches them to use it more efficiently. Zorik Gekhman, Gal Yona, Roee Aharoni, Matan Eyal, Amir Feder, Roi Reichart, Jonathan Herzig |
EMNLP | 6 |
| 2024 | Systematic Biases in LLM Simulations of DebatesabstractThe emergence of Large Language Models (LLMs), has opened exciting possibilities for constructing computational simulations designed to replicate human behavior accurately.Current research suggests that LLM-based agents become increasingly human-like in their performance, sparking interest in using these AI agents as substitutes for human participants in behavioral studies.However, LLMs are complex statistical learners without straightforward deductive rules, making them prone to unexpected behaviors.Hence, it is crucial to study and pinpoint the key behavioral distinctions between humans and LLM-based agents.In this study, we highlight the limitations of LLMs in simulating human interactions, particularly focusing on LLMs' ability to simulate political debates on topics that are important aspects of people's day-to-day lives and decision-making processes.Our findings indicate a tendency for LLM agents to conform to the model's inherent social biases despite being directed to debate from certain political perspectives.This tendency results in behavioral patterns that seem to deviate from well-established social dynamics among humans.We reinforce these observations using an automatic self-fine-tuning method, which enables us to manipulate the biases within the LLM and demonstrate that agents subsequently align with the altered biases.These results underscore the need for further research to develop methods that help agents overcome these biases, a critical step toward creating more realistic simulations. Amir Taubenfeld, Yaniv Dover, Roi Reichart, Ariel Goldstein |
EMNLP | 3 |
| 2024 | Faithful Explanations of Black-box NLP Models Using LLM-generated CounterfactualsabstractCausal explanations of the predictions of NLP systems are essential to ensure safety and establish trust. Yet, existing methods often fall short of explaining model predictions effectively or efficiently and are often model-specific. In this paper, we address model-agnostic explanations, proposing two approaches for counterfactual (CF) approximation. The first approach is CF generation, where a large language model (LLM) is prompted to change a specific text concept while keeping confounding concepts unchanged. While this approach is demonstrated to be very effective, applying LLM at inference-time is costly. We hence present a second approach based on matching, and propose a method that is guided by an LLM at training-time and learns a dedicated embedding space. This space is faithful to a given causal graph and effectively serves to identify matches that approximate CFs. After showing theoretically that approximating CFs is required in order to construct faithful explanations, we benchmark our approaches and explain several models, including LLMs with billions of parameters. Our empirical results demonstrate the excellent performance of CF generation models as model-agnostic explainers. Moreover, our matching approach, which requires far less test-time resources, also provides effective explanations, surpassing many baselines. We also find that Top-K techniques universally improve every tested method. Finally, we showcase the potential of LLMs in constructing new benchmarks for model explanation and subsequently validate our conclusions. Our work illuminates new pathways for efficient and accurate approaches to interpreting NLP systems. Yair Ori Gat, Nitay Calderon, Amir Feder, Alexander Chapanin, Roi Reichart |
ICLR | 6 |
| 2024 | The Colorful Future of LLMs: Evaluating and Improving LLMs as Emotional Supporters for Queer YouthabstractShir Lissak, Nitay Calderon, Geva Shenkman, Yaakov Ophir, Eyal Fruchter, Anat Brunstein Klomek, Roi Reichart. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Shir Lissak, Nitay Calderon, Geva Shenkman, Yaakov Ophir, Eyal Fruchter, Anat Brunstein Klomek, Roi Reichart |
NAACL-HLT | 7 |
| 2023 | A Systematic Study of Knowledge Distillation for Natural Language Generation with Pseudo-Target TrainingabstractModern Natural Language Generation (NLG) models come with massive computational and storage requirements.In this work, we study the potential of compressing them, which is crucial for real-world applications serving millions of users.We focus on Knowledge Distillation (KD) techniques, in which a small student model learns to imitate a large teacher model, allowing to transfer knowledge from the teacher to the student.In contrast to much of the previous work, our goal is to optimize the model for a specific NLG task and a specific dataset.Typically in real-world applications, in addition to labeled data there is abundant unlabeled task-specific data, which is crucial for attaining high compression rates via KD.In this work, we conduct a systematic study of task-specific KD techniques for various NLG tasks under realistic assumptions.We discuss the special characteristics of NLG distillation and particularly the exposure bias problem.Following, we derive a family of Pseudo-Target (PT) augmentation methods, substantially extending prior work on sequence-level KD.We propose the Joint-Teaching method, which applies wordlevel KD to multiple PTs generated by both the teacher and the student.Finally, we validate our findings in an extreme setup with no labeled examples using GPT-4 as the teacher.Our study provides practical model design observations and demonstrates the effectiveness of PT training for task-specific KD in NLG. Nitay Calderon, Subhabrata Mukherjee, Roi Reichart, Amir Kantor |
ACL (1) | 3 |
| 2023 | Your Prompt is My Command: On Assessing the Human-Centred Generality of Multimodal ModelsabstractEven with obvious deficiencies, large prompt-commanded multimodal models are proving to be flexible cognitive tools representing an unprecedented generality. But the directness, diversity, and degree of user interaction create a distinctive “human-centred generality” (HCG), rather than a fully autonomous one. HCG implies that —for a specific user— a system is only as general as it is effective for the user’s relevant tasks and their prevalent ways of prompting. A human-centred evaluation of general-purpose AI systems therefore needs to reflect the personal nature of interaction, tasks and cognition. We argue that the best way to understand these systems is as highly-coupled cognitive extenders, and to analyse the bidirectional cognitive adaptations between them and humans. In this paper, we give a formulation of HCG, as well as a high-level overview of the elements and trade-offs involved in the prompting process. We end the paper by outlining some essential research questions and suggestions for improving evaluation practices, which we envision as characteristic for the evaluation of general artificial intelligence in the future. This paper appears in the AI & Society track. Wout Schellaert, Fernando Martínez-Plumed, Karina Vold, John Burden, P. A. M. Casares, Bao Sheng Loe, Roi Reichart, Seán Ó hÉigeartaigh, Anna Korhonen, José Hernández-Orallo |
J. Artif. Intell. Res. | 7 |
| 2023 | On the Robustness of Dialogue History Representation in Conversational Question Answering: A Comprehensive Study and a New Prompt-based MethodabstractAbstract Most work on modeling the conversation history in Conversational Question Answering (CQA) reports a single main result on a common CQA benchmark. While existing models show impressive results on CQA leaderboards, it remains unclear whether they are robust to shifts in setting (sometimes to more realistic ones), training data size (e.g., from large to small sets) and domain. In this work, we design and conduct the first large-scale robustness study of history modeling approaches for CQA. We find that high benchmark scores do not necessarily translate to strong robustness, and that various methods can perform extremely differently under different settings. Equipped with the insights from our study, we design a novel prompt-based history modeling approach and demonstrate its strong robustness across various settings. Our approach is inspired by existing methods that highlight historic answers in the passage. However, instead of highlighting by modifying the passage token embeddings, we add textual prompts directly in the passage text. Our approach is simple, easy to plug into practically any model, and highly effective, thus we recommend it as a starting point for future model developers. We also hope that our study and insights will raise awareness to the importance of robustness-focused evaluation, in addition to obtaining high leaderboard scores, leading to better CQA systems.1 Zorik Gekhman, Nadav Oved, Orgad Keller, Idan Szpektor, Roi Reichart |
Trans. Assoc. Comput. Linguistics | 5 |
| 2022 | DoCoGen: Domain Counterfactual Generation for Low Resource Domain AdaptationabstractNatural language processing (NLP) algorithms have become very successful, but they still struggle when applied to out-of-distribution examples.In this paper we propose a controllable generation approach in order to deal with this domain adaptation (DA) challenge.Given an input text example, our DoCoGen algorithm generates a domain-counterfactual textual example (D-CON) -that is similar to the original in all aspects, including the task label, but its domain is changed to a desired one.Importantly, DoCoGen is trained using only unlabeled examples from multiple domainsno NLP task labels or parallel pairs of textual examples and their domain-counterfactuals are required.We show that DoCoGen can generate coherent counterfactuals consisting of multiple sentences.We use the D-CONs generated by DoCoGen to augment a sentiment classifier and a multi-label intent classifier in 20 and 78 DA setups, respectively, where source-domain labeled data is scarce.Our model outperforms strong baselines and improves the accuracy of a state-of-the-art unsupervised DA algorithm.1 * Both authors equally contributed to this work.Original, Kitchen: A good knife but Quality Control was poor.The knife is solid and very comfortable in hand, however, when I got it new, the blade is slightly bent.I expect it to be in almost perfect condition, but it's not.DoCoGen, Kitchen → Electronics: A good product but Quality Control was poor.The ipod is very easy to use and very comfortable in hand, however, when I got it new, the ipod is slightly flimsy.I expect it to be in almost perfect shape, but it's not.Original, DVD: The direction of this film is excellent.I love all the characters and the way they interact.The storyline is very important also.It's about religious beliefs and neighbors that interact with each other.It's a well-paced and interesting story that's not like anything else I've ever seen.DoCoGen, DVD → Airline: The service on this flight is excellent.I love the staff and the way they interact.The safety is very important also.It's nice to have staff and neighbors that can help each other.It's a well-groomed and professional crew that's not like anything else I've ever experienced.Original, Electronics: That relay board is only good for switching AC loads of 100V or more.If you have a lower voltage load, it's not going to work.For low voltage loads use transistors, MOSFETs or a ULN2803 driver board.DoCoGen, Electronics → Statistics: That model is only good for data of $n$ or more.If you have a lower $n$, it's not going to work.For lower $n$ regression use a linear, logistic or a t-test. Nitay Calderon, Eyal Ben-David, Amir Feder, Roi Reichart |
ACL (1) | 4 |
| 2022 | Learning Discrete Structured Variational Auto-Encoder using Natural Evolution Strategies
Alon Berliner, Guy Rotman, Yossi Adi, Roi Reichart, Tamir Hazan |
ICLR | 4 |
| 2022 | A Functional Information Perspective on Model InterpretationabstractContemporary predictive models are hard to interpret as their deep nets exploit numerous complex relations between input elements. This work suggests a theoretical framework for model interpretability by measuring the contribution of relevant features to the functional entropy of the network with respect to the input. We rely on the log-Sobolev inequality that bounds the functional entropy by the functional Fisher information with respect to the covariance of the data. This provides a principled way to measure the amount of information contribution of a subset of features to the decision function. Through extensive experiments, we show that our method surpasses existing interpretability sampling-based methods on various data signals such as image, text, and audio. Itai Gat, Nitay Calderon, Roi Reichart, Tamir Hazan |
ICML | 3 |
| 2022 | CEBaB: Estimating the Causal Effects of Real-World Concepts on NLP Model BehaviorabstractThe increasing size and complexity of modern ML systems has improved their predictive capabilities but made their behavior harder to explain. Many techniques for model explanation have been developed in response, but we lack clear criteria for assessing these techniques. In this paper, we cast model explanation as the causal inference problem of estimating causal effects of real-world concepts on the output behavior of ML models given actual input data. We introduce CEBaB, a new benchmark dataset for assessing concept-based explanation methods in Natural Language Processing (NLP). CEBaB consists of short restaurant reviews with human-generated counterfactual reviews in which an aspect (food, noise, ambiance, service) of the dining experience was modified. Original and counterfactual reviews are annotated with multiply-validated sentiment ratings at the aspect-level and review-level. The rich structure of CEBaB allows us to go beyond input features to study the effects of abstract, real-world concepts on model behavior. We use CEBaB to compare the quality of a range of concept-based explanation methods covering different assumptions and conceptions of the problem, and we seek to establish natural metrics for comparative assessments of these methods. Eldar David Abraham, Karel D'Oosterlinck, Amir Feder, Yair Ori Gat, Atticus Geiger, Christopher Potts, Roi Reichart, Zhengxuan Wu |
NeurIPS | 7 |
| 2022 | In the Eye of the Beholder: Robust Prediction with Causal User ModelingabstractAccurately predicting the relevance of items to users is crucial to the success of many social platforms. Conventional approaches train models on logged historical data; but recommendation systems, media services, and online marketplaces all exhibit a constant influx of new content---making relevancy a moving target, to which standard predictive models are not robust. In this paper, we propose a learning framework for relevance prediction that is robust to changes in the data distribution. Our key observation is that robustness can be obtained by accounting for \emph{how users causally perceive the environment}. We model users as boundedly-rational decision makers whose causal beliefs are encoded by a causal graph, and show how minimal information regarding the graph can be used to contend with distributional changes. Experiments in multiple settings demonstrate the effectiveness of our approach. Amir Feder, Guy Horowitz, Yoav Wald, Roi Reichart, Nir Rosenfeld |
NeurIPS | 4 |
| 2022 | Predicting Decisions in Language Based Persuasion GamesabstractSender-receiver interactions, and specifically persuasion games, are widely researched in economic modeling and artificial intelligence, and serve as a solid foundation for powerful applications. However, in the classic persuasion games setting, the messages sent from the expert to the decision-maker are abstract or well-structured application-specific signals rather than natural (human) language messages, although natural language is a very common communication signal in real-world persuasion setups. This paper addresses the use of natural language in persuasion games, exploring its impact on the decisions made by the players and aiming to construct effective models for the prediction of these decisions. For this purpose, we conduct an online repeated interaction experiment. At each trial of the interaction, an informed expert aims to sell an uninformed decision-maker a vacation in a hotel, by sending her a review that describes the hotel. While the expert is exposed to several scored reviews, the decision-maker observes only the single review sent by the expert, and her payoff in case she chooses to take the hotel is a random draw from the review score distribution available to the expert only. The expert’s payoff, in turn, depends on the number of times the decision-maker chooses the hotel. We also compare the behavioral patterns in this experiment to the equivalent patterns in similar experiments where the communication is based on the numerical values of the reviews rather than the reviews’ text, and observe substantial differences which can be explained through an equilibrium analysis of the game. We consider a number of modeling approaches for our verbal communication setup, differing from each other in the model type (deep neural network (DNN) vs. linear classifier), the type of features used by the model (textual, behavioral or both) and the source of the textual features (DNN-based vs. hand-crafted). Our results demonstrate that given a prefix of the interaction sequence, our models can predict the future decisions of the decision-maker, particularly when a sequential modeling approach and hand-crafted textual features are applied. Further analysis of the hand-crafted textual features allows us to make initial observations about the aspects of text that drive decision making in our setup. Reut Apel, Ido Erev, Roi Reichart, Moshe Tennenholtz |
J. Artif. Intell. Res. | 3 |
| 2022 | PADA: Example-based Prompt Learning for on-the-fly Adaptation to Unseen DomainsabstractAbstract Natural Language Processing algorithms have made incredible progress, but they still struggle when applied to out-of-distribution examples. We address a challenging and underexplored version of this domain adaptation problem, where an algorithm is trained on several source domains, and then applied to examples from unseen domains that are unknown at training time. Particularly, no examples, labeled or unlabeled, or any other knowledge about the target domain are available to the algorithm at training time. We present PADA: An example-based autoregressive Prompt learning algorithm for on-the-fly Any-Domain Adaptation, based on the T5 language model. Given a test example, PADA first generates a unique prompt for it and then, conditioned on this prompt, labels the example with respect to the NLP prediction task. PADA is trained to generate a prompt that is a token sequence of unrestricted length, consisting of Domain Related Features (DRFs) that characterize each of the source domains. Intuitively, the generated prompt is a unique signature that maps the test example to a semantic space spanned by the source domains. In experiments with 3 tasks (text classification and sequence tagging), for a total of 14 multi-source adaptation scenarios, PADA substantially outperforms strong baselines.1 Eyal Ben-David, Nadav Oved, Roi Reichart |
Trans. Assoc. Comput. Linguistics | 3 |
| 2022 | Causal Inference in Natural Language Processing: Estimation, Prediction, Interpretation and BeyondabstractAbstract A fundamental goal of scientific research is to learn about causal relationships. However, despite its critical role in the life and social sciences, causality has not had the same importance in Natural Language Processing (NLP), which has traditionally placed more emphasis on predictive tasks. This distinction is beginning to fade, with an emerging area of interdisciplinary research at the convergence of causal inference and language processing. Still, research on causality in NLP remains scattered across domains without unified definitions, benchmark datasets and clear articulations of the challenges and opportunities in the application of causal inference to the textual domain, with its unique properties. In this survey, we consolidate research across academic areas and situate it in the broader NLP landscape. We introduce the statistical challenge of estimating causal effects with text, encompassing settings where text is used as an outcome, treatment, or to address confounding. In addition, we explore potential uses of causal inference to improve the robustness, fairness, and interpretability of NLP models. We thus provide a unified overview of causal inference for the NLP community.1 Amir Feder, Katherine A. Keith, Emaad Manzoor, Reid Pryzant, Dhanya Sridhar, Zach Wood-Doughty, Jacob Eisenstein, Justin Grimmer, Roi Reichart, Margaret E. Roberts, Brandon M. Stewart, Victor Veitch, Diyi Yang |
Trans. Assoc. Comput. Linguistics | 9 |
| 2022 | Designing an Automatic Agent for Repeated Language-based Persuasion GamesabstractAbstract Persuasion games are fundamental in economics and AI research and serve as the basis for important applications. However, work on this setup assumes communication with stylized messages that do not consist of rich human language. In this paper we consider a repeated sender (expert) – receiver (decision maker) game, where the sender is fully informed about the state of the world and aims to persuade the receiver to accept a deal by sending one of several possible natural language reviews. We design an automatic expert that plays this repeated game, aiming to achieve the maximal payoff. Our expert is implemented within the Monte Carlo Tree Search (MCTS) algorithm, with deep learning models that exploit behavioral and linguistic signals in order to predict the next action of the decision maker, and the future payoff of the expert given the state of the game and a candidate review. We demonstrate the superiority of our expert over strong baselines and its adaptability to different decision makers and potential proposed deals.1 Maya Raifer, Guy Rotman, Reut Apel, Moshe Tennenholtz, Roi Reichart |
Trans. Assoc. Comput. Linguistics | 5 |
| 2022 | Multi-task Active Learning for Pre-trained Transformer-based ModelsabstractAbstract Multi-task learning, in which several tasks are jointly learned by a single model, allows NLP models to share information from multiple annotations and may facilitate better predictions when the tasks are inter-related. This technique, however, requires annotating the same text with multiple annotation schemes, which may be costly and laborious. Active learning (AL) has been demonstrated to optimize annotation processes by iteratively selecting unlabeled examples whose annotation is most valuable for the NLP model. Yet, multi-task active learning (MT-AL) has not been applied to state-of-the-art pre-trained Transformer-based NLP models. This paper aims to close this gap. We explore various multi-task selection criteria in three realistic multi-task scenarios, reflecting different relations between the participating tasks, and demonstrate the effectiveness of multi-task compared to single-task selection. Our results suggest that MT-AL can be effectively used in order to minimize annotation efforts for multi-task NLP models.1 Guy Rotman, Roi Reichart |
Trans. Assoc. Comput. Linguistics | 2 |
| 2021 | A Closer Look at Few-Shot Crosslingual Transfer: The Choice of Shots MattersabstractMengjie Zhao, Yi Zhu, Ehsan Shareghi, Ivan Vulić, Roi Reichart, Anna Korhonen, Hinrich Schütze. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Ehsan Shareghi, Ivan Vulic, Roi Reichart, Anna Korhonen, Hinrich Schütze |
ACL/IJCNLP (1) | 5 |
| 2021 | Combining Deep Generative Models and Multi-lingual Pretraining for Semi-supervised Document ClassificationabstractSemi-supervised learning through deep generative models and multi-lingual pretraining techniques have orchestrated tremendous success across different areas of NLP.Nonetheless, their development has happened in isolation, while the combination of both could potentially be effective for tackling task-specific labelled data shortage.To bridge this gap, we combine semi-supervised deep generative models and multi-lingual pretraining to form a pipeline for document classification task.Compared to strong supervised learning baselines, our semi-supervised classification framework is highly competitive and outperforms the state-of-the-art counterparts in lowresource settings across several languages.1 Ehsan Shareghi, Yingzhen Li, Roi Reichart, Anna Korhonen |
EACL | 4 |
| 2021 | DILBERT: Customized Pre-Training for Domain Adaptation with Category Shift, with an Application to Aspect ExtractionabstractThe rise of pre-trained language models has yielded substantial progress in the vast majority of Natural Language Processing (NLP) tasks.However, a generic approach towards the pre-training procedure can naturally be sub-optimal in some cases.Particularly, finetuning a pre-trained language model on a source domain and then applying it to a different target domain, results in a sharp performance decline of the eventual classifier for many source-target domain pairs.Moreover, in some NLP tasks, the output categories substantially differ between domains, making adaptation even more challenging.This, for example, happens in the task of aspect extraction, where the aspects of interest of reviews of, e.g., restaurants or electronic devices may be very different.This paper presents a new fine-tuning scheme for BERT, which aims to address the above challenges.We name this scheme DILBERT: Domain Invariant Learning with BERT, and customize it for aspect extraction in the unsupervised domain adaptation setting.DILBERT harnesses the categorical information of both the source and the target domains to guide the pre-training process towards a more domain and category invariant representation, thus closing the gap between the domains.We show that DILBERT yields substantial improvements over state-ofthe-art baselines while using a fraction of the unlabeled data, particularly in more challenging domain adaptation setups. 1 Entony Lekhtman, Yftah Ziser, Roi Reichart |
EMNLP (1) | 3 |
| 2021 | CausaLM: Causal Model Explanation Through Counterfactual Language ModelsabstractAbstract Understanding predictions made by deep neural networks is notoriously difficult, but also crucial to their dissemination. As all machine learning–based methods, they are as good as their training data, and can also capture unwanted biases. While there are tools that can help understand whether such biases exist, they do not distinguish between correlation and causation, and might be ill-suited for text-based models and for reasoning about high level language concepts. A key problem of estimating the causal effect of a concept of interest on a given model is that this estimation requires the generation of counterfactual examples, which is challenging with existing generation technology. To bridge that gap, we propose CausaLM, a framework for producing causal model explanations using counterfactual language representation models. Our approach is based on fine-tuning of deep contextualized embedding models with auxiliary adversarial tasks derived from the causal graph of the problem. Concretely, we show that by carefully choosing auxiliary adversarial pre-training tasks, language representation models such as BERT can effectively learn a counterfactual representation for a given concept of interest, and be used to estimate its true causal effect on model performance. A byproduct of our method is a language representation model that is unaffected by the tested concept, which can be useful in mitigating unwanted bias ingrained in the data. Amir Feder, Nadav Oved, Uri Shalit, Roi Reichart |
Comput. Linguistics | 4 |
| 2021 | Parameter Space Factorization for Zero-Shot Learning across Tasks and LanguagesabstractAbstract Most combinations of NLP tasks and language varieties lack in-domain examples for supervised training because of the paucity of annotated data. How can neural models make sample-efficient generalizations from task–language combinations with available data to low-resource ones? In this work, we propose a Bayesian generative model for the space of neural parameters. We assume that this space can be factorized into latent variables for each language and each task. We infer the posteriors over such latent variables based on data from seen task–language combinations through variational inference. This enables zero-shot classification on unseen combinations at prediction time. For instance, given training data for named entity recognition (NER) in Vietnamese and for part-of-speech (POS) tagging in Wolof, our model can perform accurate predictions for NER in Wolof. In particular, we experiment with a typologically diverse sample of 33 languages from 4 continents and 11 families, and show that our model yields comparable or better results than state-of-the-art, zero-shot cross-lingual transfer methods. Our code is available at github.com/cambridgeltl/parameter-factorization. Edoardo Maria Ponti, Ivan Vulic, Ryan Cotterell, Marinela Parovic, Roi Reichart, Anna Korhonen |
Trans. Assoc. Comput. Linguistics | 5 |
| 2021 | Model Compression for Domain Adaptation through Causal Effect EstimationabstractAbstract Recent improvements in the predictive quality of natural language processing systems are often dependent on a substantial increase in the number of model parameters. This has led to various attempts of compressing such models, but existing methods have not considered the differences in the predictive power of various model components or in the generalizability of the compressed models. To understand the connection between model compression and out-of-distribution generalization, we define the task of compressing language representation models such that they perform best in a domain adaptation setting. We choose to address this problem from a causal perspective, attempting to estimate the average treatment effect (ATE) of a model component, such as a single layer, on the model’s predictions. Our proposed ATE-guided Model Compression scheme (AMoC), generates many model candidates, differing by the model components that were removed. Then, we select the best candidate through a stepwise regression model that utilizes the ATE to predict the expected performance on the target domain. AMoC outperforms strong baselines on dozens of domain pairs across three text classification and sequence tagging tasks.1 Guy Rotman, Amir Feder, Roi Reichart |
Trans. Assoc. Comput. Linguistics | 3 |
| 2020 | Multidirectional Associative Optimization of Function-Specific Word RepresentationsabstractWe present a neural framework for learning associations between interrelated groups of words such as the ones found in Subject-Verb-Object (SVO) structures.Our model induces a joint function-specific word vector space, where vectors of e.g.plausible SVO compositions lie close together.The model retains information about word group membership even in the joint space, and can thereby effectively be applied to a number of tasks reasoning over the SVO structure.We show the robustness and versatility of the proposed framework by reporting state-of-the-art results on the tasks of estimating selectional preference and event similarity.The results indicate that the combinations of representations learned with our task-independent model outperform task-specific architectures from prior work, while reducing the number of parameters by up to 95%. Daniela Gerz, Ivan Vulic, Marek Rei, Roi Reichart, Anna Korhonen |
ACL | 4 |
| 2020 | The Secret is in the Spectra: Predicting Cross-lingual Task Performance with Spectral Similarity MeasuresabstractPerformance in cross-lingual NLP tasks is impacted by the (dis)similarity of languages at hand: e.g., previous work has suggested there is a connection between the expected success of bilingual lexicon induction (BLI) and the assumption of (approximate) isomorphism between monolingual embedding spaces.In this work we present a large-scale study focused on the correlations between monolingual embedding space similarity and task performance, covering thousands of language pairs and four different tasks: BLI, parsing, POS tagging and MT.We hypothesize that statistics of the spectrum of each monolingual embedding space indicate how well they can be aligned.We then introduce several isomorphism measures between two embedding spaces, based on the relevant statistics of their individual spectra.We empirically show that 1) language similarity scores derived from such spectral isomorphism measures are strongly associated with performance observed in different crosslingual tasks, and 2) our spectral-based measures consistently outperform previous standard isomorphism measures, while being computationally more tractable and easier to interpret.Finally, our measures capture complementary information to typologically driven language distance measures, and the combination of measures from the two families yields even higher task performance correlations. Haim Dubossarsky, Ivan Vulic, Roi Reichart, Anna Korhonen |
EMNLP (1) | 3 |
| 2020 | Geosocial Location Classification: Associating Type to Places Based on Geotagged Social-Media PostsabstractAssociating type to locations can be used to enrich maps and can serve a plethora of geospatial applications. An automatic method to do so could make the process less expensive in terms of human labor, and faster to react to changes. In this paper we study the problem of Geosocial Location Classification, where the type of a site, e.g., a building, is discovered based on social-media posts. Our goal is to correctly associate a set of messages posted in a small radius around a given location with the corresponding location type, e.g., school, church, restaurant or museum. We explore two approaches to the problem: (a) a pipeline approach, where each message is first classified, and then the location associated with the message set is inferred from the separate message labels; and (b) a joint approach where the messages are simultaneously processed to yield the desired location type. We tested the two approaches over a dataset of geotagged tweets. Our results demonstrate the superiority of the joint approach. Elad Kravi, Yaron Kanza, Benny Kimelfeld, Roi Reichart |
SIGSPATIAL/GIS | 4 |
| 2020 | Predicting Strategic Behavior from Free Text (Extended Abstract)abstractThe connection between messaging and action is fundamental both to web applications, such as web search and sentiment analysis, and to economics. However, while prominent online applications exploit messaging in natural (human) language in order to predict non-strategic action selection, the economics literature focuses on the connection between structured stylized messaging to strategic decisions in games and multi-agent encounters. This paper aims to connect these two strands of research, which we consider highly timely and important due to the vast online textual communication on the web. Particularly, we introduce the following question: can free text expressed in natural language serve for the prediction of action selection in an economic context, modeled as a game? We initiate research on this question by providing preliminary positive results. Omer Ben-Porat, Lital Kuchy, Sharon Hirsch, Guy Elad, Roi Reichart, Moshe Tennenholtz |
IJCAI | 5 |
| 2020 | Predicting In-Game Actions from Interviews of NBA PlayersabstractSports competitions are widely researched in computer and social science, with the goal of understanding how players act under uncertainty. Although there is an abundance of computational work on player metrics prediction based on past performance, very few attempts to incorporate out-of-game signals have been made. Specifically, it was previously unclear whether linguistic signals gathered from players’ interviews can add information that does not appear in performance metrics. To bridge that gap, we define text classification tasks of predicting deviations from mean in NBA players’ in-game actions, which are associated with strategic choices, player behavior, and risk, using their choice of language prior to the game. We collected a data set of transcripts from key NBA players’ pre-game interviews and their in-game performance metrics, totalling 5,226 interview-metric pairs. We design neural models for players’ action prediction based on increasingly more complex aspects of the language signals in their open-ended interviews. Our models can make their predictions based on the textual signal alone, or on a combination of that signal with signals from past-performance metrics. Our text-based models outperform strong baselines trained on performance metrics only, demonstrating the importance of language usage for action prediction. Moreover, the models that utilize both textual input and past-performance metrics produced the best results. Finally, as neural networks are notoriously difficult to interpret, we propose a method for gaining further insight into what our models have learned. Particularly, we present a latent Dirichlet allocation–based analysis, where we interpret model predictions in terms of correlated topics. We find that our best performing textual model is most associated with topics that are intuitively related to each prediction task and that better models yield higher correlation with more informative topics.1 Nadav Oved, Amir Feder, Roi Reichart |
Comput. Linguistics | 3 |
| 2020 | Multi-SimLex: A Large-Scale Evaluation of Multilingual and Crosslingual Lexical Semantic SimilarityabstractWe introduce Multi-SimLex, a large-scale lexical resource and evaluation benchmark covering data sets for 12 typologically diverse languages, including major languages (e.g., Mandarin Chinese, Spanish, Russian) as well as less-resourced ones (e.g., Welsh, Kiswahili). Each language data set is annotated for the lexical relation of semantic similarity and contains 1,888 semantically aligned concept pairs, providing a representative coverage of word classes (nouns, verbs, adjectives, adverbs), frequency ranks, similarity intervals, lexical fields, and concreteness levels. Additionally, owing to the alignment of concepts across languages, we provide a suite of 66 crosslingual semantic similarity data sets. Because of its extensive size and language coverage, Multi-SimLex provides entirely novel opportunities for experimental evaluation and analysis. On its monolingual and crosslingual benchmarks, we evaluate and analyze a wide array of recent state-of-the-art monolingual and crosslingual representation models, including static and contextualized word embeddings (such as fastText, monolingual and multilingual BERT, XLM), externally informed lexical representations, as well as fully unsupervised and (weakly) supervised crosslingual word embeddings. We also present a step-by-step data set creation protocol for creating consistent, Multi-Simlex–style resources for additional languages. We make these contributions—the public release of Multi-SimLex data sets, their creation protocol, strong baseline results, and in-depth analyses which can be helpful in guiding future developments in multilingual lexical semantics and representation learning—available via a Web site that will encourage community effort in further expansion of Multi-Simlex to many more languages. Such a large-scale semantic resource could inspire significant further advances in NLP across languages. Ivan Vulic, Simon Baker, Edoardo Maria Ponti, Ulla Petti, Ira Leviant, Kelly Wing, Olga Majewska, Eden Bar, Matt Malone, Thierry Poibeau, Roi Reichart, Anna Korhonen |
Comput. Linguistics | 11 |
| 2020 | Predicting Strategic Behavior from Free TextabstractThe connection between messaging and action is fundamental both to web applications, such as web search and sentiment analysis, and to economics. However, while prominent online applications exploit messaging in natural (human) language in order to predict non-strategic action selection, the economics literature focuses on the connection between structured stylized messaging to strategic decisions in games and multi-agent encounters. This paper aims to connect these two strands of research, which we consider highly timely and important due to the vast online textual communication on the web. Particularly, we introduce the following question: Can free text expressed in natural language serve for the prediction of action selection in an economic context, modeled as a game In order to initiate the research on this question, we introduce the study of an individual’s action prediction in a one-shot game based on free text he/she provides, while being unaware of the game to be played. We approach the problem by attributing commonsensical personality attributes via crowd-sourcing to free texts written by individuals, and employing transductive learning to predict actions taken by these individuals in one-shot games based on these attributes. Our approach allows us to train a single classifier that can make predictions with respect to actions taken in multiple games. In experiments with three well-studied games, our algorithm compares favorably with strong alternative approaches. In ablation analysis, we demonstrate the importance of our modeling choices—the representation of the text with the commonsensical personality attributes and our classifier—to the predictive power of our model. Omer Ben-Porat, Sharon Hirsch, Lital Kuchy, Guy Elad, Roi Reichart, Moshe Tennenholtz |
J. Artif. Intell. Res. | 5 |
| 2020 | PERL: Pivot-based Domain Adaptation for Pre-trained Deep Contextualized Embedding ModelsabstractPivot-based neural representation models have led to significant progress in domain adaptation for NLP. However, previous research following this approach utilize only labeled data from the source domain and unlabeled data from the source and target domains, but neglect to incorporate massive unlabeled corpora that are not necessarily drawn from these domains. To alleviate this, we propose PERL: A representation learning model that extends contextualized word embedding models such as BERT (Devlin et al., 2019 ) with pivot-based fine-tuning. PERL outperforms strong baselines across 22 sentiment classification domain adaptation setups, improves in-domain model performance, yields effective reduced-size models, and increases model stability.1 Eyal Ben-David, Carmel Rabinovitz, Roi Reichart |
Trans. Assoc. Comput. Linguistics | 3 |
| 2019 | Deep Dominance - How to Properly Compare Deep Neural ModelsabstractComparing between Deep Neural Network (DNN) models based on their performance on unseen data is crucial for the progress of the NLP field.However, these models have a large number of hyper-parameters and, being non-convex, their convergence point depends on the random values chosen at initialization and during training.Proper DNN comparison hence requires a comparison between their empirical score distributions on unseen data, rather than between single evaluation scores as is standard for more simple, convex models.In this paper, we propose to adapt to this problem a recently proposed test for the Almost Stochastic Dominance relation between two distributions.We define the criteria for a high quality comparison method between DNNs, and show, both theoretically and through analysis of extensive experimental results with leading DNN models for sequence tagging tasks, that the proposed test meets all criteria while previously proposed methods fail to do so.We hope the test we propose here will set a new working practice in the NLP community.1 Rotem Dror, Segev Shlomov, Roi Reichart |
ACL (1) | 3 |
| 2019 | Zero-Shot Semantic Parsing for InstructionsabstractWe consider a zero-shot semantic parsing task: parsing instructions into compositional logical forms, in domains that were not seen during training.We present a new dataset with 1,390 examples from 7 application domains (e.g. a calendar or a file manager), each example consisting of a triplet: (a) the application's initial state, (b) an instruction, to be carried out in the context of that state, and (c) the state of the application after carrying out the instruction.We introduce a new training algorithm that aims to train a semantic parser on examples from a set of source domains, so that it can effectively parse instructions from an unknown target domain.We integrate our algorithm into the floating parser of Pasupat and Liang (2015), and further augment the parser with features and a logical form candidate filtering logic, to support zero-shot adaptation.Our experiments with various zero-shot adaptation setups demonstrate substantial performance gains over a non-adapted parser. 1 Ofer Givoli, Roi Reichart |
ACL (1) | 2 |
| 2019 | Task Refinement Learning for Improved Accuracy and Stability of Unsupervised Domain AdaptationabstractPivot Based Language Modeling (PBLM) (Ziser and Reichart, 2018a), combining LSTMs with pivot-based methods, has yielded significant progress in unsupervised domain adaptation. However, this approach is still challenged by the large pivot detection problem that should be solved, and by the inherent instability of LSTMs. In this paper we propose a Task Refinement Learning (TRL) approach, in order to solve these problems. Our algorithms iteratively train the PBLM model, gradually increasing the information exposed about each pivot. TRL-PBLM achieves stateof- the-art accuracy in six domain adaptation setups for sentiment classification. Moreover, it is much more stable than plain PBLM across model configurations, making the model much better fitted for practical use. Yftah Ziser, Roi Reichart |
ACL (1) | 2 |
| 2019 | On the Importance of Subword Information for Morphological Tasks in Truly Low-Resource LanguagesabstractRecent work has validated the importance of subword information for word representation learning. Since subwords increase parameter sharing ability in neural models, their value should be even more pronounced in low-data regimes. In this work, we therefore provide a comprehensive analysis focused on the usefulness of subwords for word representation learning in truly low-resource scenarios and for three representative morphological tasks: fine-grained entity typing, morphological tagging, and named entity recognition. We conduct a systematic study that spans several dimensions of comparison: 1) type of data scarcity which can stem from the lack of task-specific training data, or even from the lack of unannotated data required to train word embeddings, or both; 2) language type by working with a sample of 16 typologically diverse languages including some truly low-resource ones (e.g. Rusyn, Buryat, and Zulu); 3) the choice of the subword-informed word representation method. Our main results show that subword-informed models are universally useful across all language types, with large gains over subword-agnostic embeddings. They also suggest that the effective use of subwords largely depends on the language (type) and the task at hand, as well as on the amount of available data for training the embeddings and task-based models, where having sufficient in-task data is a more critical requirement. Benjamin Heinzerling, Ivan Vulic, Michael Strube 0001, Roi Reichart, Anna Korhonen |
CoNLL | 5 |
| 2019 | Towards Zero-shot Language ModelingabstractEdoardo Maria Ponti, Ivan Vulić, Ryan Cotterell, Roi Reichart, Anna Korhonen. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Edoardo Maria Ponti, Ivan Vulic, Ryan Cotterell, Roi Reichart, Anna Korhonen |
EMNLP/IJCNLP (1) | 4 |
| 2019 | Cross-lingual Semantic Specialization via Lexical Relation InductionabstractEdoardo Maria Ponti, Ivan Vulić, Goran Glavaš, Roi Reichart, Anna Korhonen. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Edoardo Maria Ponti, Ivan Vulic, Goran Glavas, Roi Reichart, Anna Korhonen |
EMNLP/IJCNLP (1) | 4 |
| 2019 | Do We Really Need Fully Unsupervised Cross-Lingual Embeddings?abstractIvan Vulić, Goran Glavaš, Roi Reichart, Anna Korhonen. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Ivan Vulic, Goran Glavas, Roi Reichart, Anna Korhonen |
EMNLP/IJCNLP (1) | 3 |
| 2019 | GECKO - A Tool for Effective Annotation of Human Conversations
Golan Levy, Raquel Sitman, Ido Amir, Eduard Golshtein, Ran Mochary, Eilon Reshef, Roi Reichart, Omri Allouche |
INTERSPEECH | 7 |
| 2019 | Modeling Language Variation and Universals: A Survey on Typological Linguistics for Natural Language ProcessingabstractLinguistic typology aims to capture structural and semantic variation across the world’s languages. A large-scale typology could provide excellent guidance for multilingual Natural Language Processing (NLP), particularly for languages that suffer from the lack of human labeled resources. We present an extensive literature survey on the use of typological information in the development of NLP techniques. Our survey demonstrates that to date, the use of information in existing typological databases has resulted in consistent but modest improvements in system performance. We show that this is due to both intrinsic limitations of databases (in terms of coverage and feature granularity) and under-utilization of the typological features included in them. We advocate for a new approach that adapts the broad and discrete nature of typological categories to the contextual and continuous nature of machine learning algorithms used in contemporary NLP. In particular, we suggest that such an approach could be facilitated by recent developments in data-driven induction of typological knowledge. Edoardo Maria Ponti, Helen O'Horan, Yevgeni Berzak, Ivan Vulic, Roi Reichart, Thierry Poibeau, Ekaterina Shutova, Anna Korhonen |
Comput. Linguistics | 5 |
| 2019 | Perturbation Based Learning for Structured NLP Tasks with Application to Dependency ParsingabstractThe best solution of structured prediction models in NLP is often inaccurate because of limited expressive power of the model or to non-exact parameter estimation. One way to mitigate this problem is sampling candidate solutions from the model’s solution space, reasoning that effective exploration of this space should yield high-quality solutions. Unfortunately, sampling is often computationally hard and many works hence back-off to sub-optimal strategies, such as extraction of the best scoring solutions of the model, which are not as diverse as sampled solutions. In this paper we propose a perturbation-based approach where sampling from a probabilistic model is computationally efficient. We present a learning algorithm for the variance of the perturbations, and empirically demonstrate its importance. Moreover, while finding the argmax in our model is intractable, we propose an efficient and effective approximation. We apply our framework to cross-lingual dependency parsing across 72 corpora from 42 languages and to lightly supervised dependency parsing across 13 corpora from 12 languages, and demonstrate strong results in terms of both the quality of the entire solution list and of the final solution.1 Amichay Doitch, Ram Yazdi, Tamir Hazan, Roi Reichart |
Trans. Assoc. Comput. Linguistics | 4 |
| 2019 | Deep Contextualized Self-training for Low Resource Dependency ParsingabstractNeural dependency parsing has proven very effective, achieving state-of-the-art results on numerous domains and languages. Unfortunately, it requires large amounts of labeled data, which is costly and laborious to create. In this paper we propose a self-training algorithm that alleviates this annotation bottleneck by training a parser on its own output. Our Deep Contextualized Self-training (DCST) algorithm utilizes representation models trained on sequence labeling tasks that are derived from the parser’s output when applied to unlabeled data, and integrates these models with the base parser through a gating mechanism. We conduct experiments across multiple languages, both in low resource in-domain and in cross-domain setups, and demonstrate that DCST substantially outperforms traditional self-training as well as recent semi-supervised training methods. 1 Guy Rotman, Roi Reichart |
Trans. Assoc. Comput. Linguistics | 2 |
| 2018 | The Hitchhiker's Guide to Testing Statistical Significance in Natural Language ProcessingabstractStatistical significance testing is a standard statistical tool designed to ensure that experimental results are not coincidental.In this opinion/theoretical paper we discuss the role of statistical significance testing in Natural Language Processing (NLP) research.We establish the fundamental concepts of significance testing and discuss the specific aspects of NLP tasks, experimental setups and evaluation measures that affect the choice of significance tests in NLP research.Based on this discussion, we propose a simple practical protocol for statistical significance test selection in NLP setups and accompany this protocol with a brief survey of the most relevant tests.We then survey recent empirical papers published in ACL and TACL during 2017 and show that while our community assigns great value to experimental results, statistical significance testing is often ignored or misused.We conclude with a brief discussion of open issues that should be properly addressed so that this important tool can be applied in NLP research in a statistically sound manner 1 . Rotem Dror, Gili Baumer, Segev Shlomov, Roi Reichart |
ACL (1) | 4 |
| 2018 | Isomorphic Transfer of Syntactic Structures in Cross-Lingual NLPabstractThe transfer or share of knowledge between languages is a popular solution to resource scarcity in NLP.However, the effectiveness of cross-lingual transfer can be challenged by variation in syntactic structures.Frameworks such as Universal Dependencies (UD) are designed to be cross-lingually consistent, but even in carefully designed resources trees representing equivalent sentences may not always overlap.In this paper, we measure cross-lingual syntactic variation, or anisomorphism, in the UD treebank collection, considering both morphological and structural properties.We show that reducing the level of anisomorphism yields consistent gains in cross-lingual transfer tasks.We introduce a source language selection procedure that facilitates effective cross-lingual parser transfer, and propose a typologically driven method for syntactic tree processing which reduces anisomorphism.Our results show the effectiveness of this method for both machine translation and cross-lingual sentence similarity, demonstrating the importance of syntactic structure compatibility for boosting cross-lingual transfer in NLP. Edoardo Maria Ponti, Roi Reichart, Anna Korhonen, Ivan Vulic |
ACL (1) | 2 |
| 2018 | Bridging Languages through Images with Deep Partial Canonical Correlation AnalysisabstractWe present a deep neural network that leverages images to improve bilingual text embeddings.Relying on bilingual image tags and descriptions, our approach conditions text embedding induction on the shared visual information for both languages, producing highly correlated bilingual embeddings.In particular, we propose a novel model based on Partial Canonical Correlation Analysis (PCCA).While the original PCCA finds linear projections of two views in order to maximize their canonical correlation conditioned on a shared third variable, we introduce a non-linear Deep PCCA (DPCCA) model, and develop a new stochastic iterative algorithm for its optimization.We evaluate PCCA and DPCCA on multilingual word similarity and cross-lingual image description retrieval.Our models outperform a large variety of previous methods, despite not having access to any visual signal during test time inference.1 Guy Rotman, Ivan Vulic, Roi Reichart |
ACL (1) | 3 |
| 2018 | On the Relation between Linguistic Typology and (Limitations of) Multilingual Language ModelingabstractA key challenge in cross-lingual NLP is developing general language-independent architectures that are equally applicable to any language.However, this ambition is largely hampered by the variation in structural and semantic properties, i.e. the typological profiles of the world's languages.In this work, we analyse the implications of this variation on the language modeling (LM) task.We present a largescale study of state-of-the art n-gram based and neural language models on 50 typologically diverse languages covering a wide variety of morphological systems.Operating in the full vocabulary LM setup focused on wordlevel prediction, we demonstrate that a coarse typology of morphological systems is predictive of absolute LM performance.Moreover, fine-grained typological features such as exponence, flexivity, fusion, and inflectional synthesis are borne out to be responsible for the proliferation of low-frequency phenomena which are organically difficult to model by statistical architectures, or for the meaning ambiguity of character n-grams.Our study strongly suggests that these features have to be taken into consideration during the construction of nextlevel language-agnostic LM architectures, capable of handling morphologically complex languages such as Tamil or Korean. Daniela Gerz, Ivan Vulic, Edoardo Maria Ponti, Roi Reichart, Anna Korhonen |
EMNLP | 4 |
| 2018 | Neural Transition Based Parsing of Web Queries: An Entity Based ApproachabstractWeb queries with question intent manifest a complex syntactic structure and the processing of this structure is important for their interpretation.Pinter et al. (2016) has formalized the grammar of these queries and proposed semisupervised algorithms for the adaptation of parsers originally designed to parse according to the standard dependency grammar, so that they can account for the unique forest grammar of queries.However, their algorithms rely on resources typically not available outside of big web corporates.We propose a new BiLSTM query parser that: (1) Explicitly accounts for the unique grammar of web queries; and (2) Utilizes named entity (NE) information from a BiLSTM NE tagger, that can be jointly trained with the parser.In order to train our model we annotate the query treebank of Pinter et al. (2016) with NEs.When trained on 2500 annotated queries our parser achieves UAS of 83.5% and segmentation F1score of 84.5, substantially outperforming existing state-of-the-art parsers. 1 Rivka Malca, Roi Reichart |
EMNLP | 2 |
| 2018 | Deep Pivot-Based Modeling for Cross-language Cross-domain Transfer with Minimal GuidanceabstractWhile cross-domain and cross-language transfer have long been prominent topics in NLP research, their combination has hardly been explored.In this work we consider this problem, and propose a framework that builds on pivotbased learning, structure-aware Deep Neural Networks (particularly LSTMs and CNNs) and bilingual word embeddings, with the goal of training a model on labeled data from one (language, domain) pair so that it can be effectively applied to another (language, domain) pair.We consider two setups, differing with respect to the unlabeled data available for model training.In the full setup the model has access to unlabeled data from both pairs, while in the lazy setup, which is more realistic for truly resource-poor languages, unlabeled data is available for both domains but only for the source language.We design our model for the lazy setup so that for a given target domain, it can train once on the source language and then be applied to any target language without re-training.In experiments with nine English-German and nine English-French domain pairs our best model substantially outperforms previous models even when it is trained in the lazy setup and previous models are trained in the full setup.1 Yftah Ziser, Roi Reichart |
EMNLP | 2 |
| 2018 | Pivot Based Language Modeling for Improved Neural Domain AdaptationabstractRepresentation learning with pivot-based methods and with Neural Networks (NNs) have lead to significant progress in domain adaptation for Natural Language Processing.However, most previous work that follows these approaches does not explicitly exploit the structure of the input text, and its output is most often a single representation vector for the entire text.In this paper we present the Pivot Based Language Model (PBLM), a representation learning model that marries together pivot-based and NN modeling in a structure aware manner.Particularly, our model processes the information in the text with a sequential NN (LSTM) and its output consists of a context-dependent representation vector for every input word.Unlike most previous representation learning models in domain adaptation, PBLM can naturally feed structure aware text classifiers such as LSTM and CNN.We experiment with the task of cross-domain sentiment classification on 20 domain pairs and show substantial improvements over strong baselines.1 Yftah Ziser, Roi Reichart |
NAACL-HLT | 2 |
| 2018 | Language Modeling for Morphologically Rich Languages: Character-Aware Modeling for Word-Level PredictionabstractNeural architectures are prominent in the construction of language models (LMs). However, word-level prediction is typically agnostic of subword-level information (characters and character sequences) and operates over a closed vocabulary, consisting of a limited word set. Indeed, while subword-aware models boost performance across a variety of NLP tasks, previous work did not evaluate the ability of these models to assist next-word prediction in language modeling tasks. Such subword-level informed models should be particularly effective for morphologically-rich languages (MRLs) that exhibit high type-to-token ratios. In this work, we present a large-scale LM study on 50 typologically diverse languages covering a wide variety of morphological systems, and offer new LM benchmarks to the community, while considering subword-level information. The main technical contribution of our work is a novel method for injecting subword-level information into semantic word vectors, integrated into the neural language modeling training, to facilitate word-level prediction. We conduct experiments in the LM setting where the number of infrequent words is large, and demonstrate strong perplexity gains across our 50 languages, especially for morphologically-rich languages. Our code and data sets are publicly available. Daniela Gerz, Ivan Vulic, Edoardo Maria Ponti, Jason Naradowsky, Roi Reichart, Anna Korhonen |
Trans. Assoc. Comput. Linguistics | 5 |
| 2017 | Sarcasm SIGN: Interpreting Sarcasm with Sentiment Based Monolingual Machine TranslationabstractSarcasm is a form of speech in which speakers say the opposite of what they truly mean in order to convey a strong sentiment.In other words, "Sarcasm is the giant chasm between what I say, and the person who doesn't get it.".In this paper we present the novel task of sarcasm interpretation, defined as the generation of a non-sarcastic utterance conveying the same message as the original sarcastic one.We introduce a novel dataset of 3000 sarcastic tweets, each interpreted by five human judges.Addressing the task as monolingual machine translation (MT), we experiment with MT algorithms and evaluation measures.We then present SIGN: an MT based sarcasm interpretation algorithm that targets sentiment words, a defining element of textual sarcasm.We show that while the scores of n-gram based automatic measures are similar for all interpretation models, SIGN's interpretations are scored higher by humans for adequacy and sentiment polarity.We conclude with a discussion on future research directions for our new task. 1 Lotem Peled-Cohen, Roi Reichart |
ACL (1) | 2 |
| 2017 | Morph-fitting: Fine-Tuning Word Vector Spaces with Simple Language-Specific RulesabstractIvan Vulić, Nikola Mrkšić, Roi Reichart, Diarmuid Ó Séaghdha, Steve Young, Anna Korhonen. Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2017. Ivan Vulic, Nikola Mrksic, Roi Reichart, Diarmuid Ó Séaghdha, Steve J. Young, Anna Korhonen |
ACL (1) | 3 |
| 2017 | Automatic Selection of Context Configurations for Improved Class-Specific Word RepresentationsabstractThis paper is concerned with identifying contexts useful for training word representation models for different word classes such as adjectives (A), verbs (V), and nouns (N).We introduce a simple yet effective framework for an automatic selection of class-specific context configurations.We construct a context configuration space based on universal dependency relations between words, and efficiently search this space with an adapted beam search algorithm.In word similarity tasks for each word class, we show that our framework is both effective and efficient.Particularly, it improves the Spearman's ρ correlation with human scores on SimLex-999 over the best previously proposed class-specific contexts by 6 (A), 6 (V) and 5 (N) ρ points.With our selected context configurations, we train on only 14% (A), 26.2% (V), and 33.6% (N) of all dependency-based contexts, resulting in a reduced training time.Our results generalise: we show that the configurations our algorithm learns for one English training setup outperform previously proposed context types in another training setup for English.Moreover, basing the configuration space on universal dependencies, it is possible to transfer the learned configurations to German and Italian.We also demonstrate improved per-class results over other context types in these two languages. Ivan Vulic, Roy Schwartz 0001, Ari Rappoport, Roi Reichart, Anna Korhonen |
CoNLL | 4 |
| 2017 | Neural Structural Correspondence Learning for Domain AdaptationabstractWe introduce a neural network model that marries together ideas from two prominent strands of research on domain adaptation through representation learning: structural correspondence learning (SCL, (Blitzer et al., 2006)) and autoencoder neural networks (NNs).Our model is a three-layer NN that learns to encode the non-pivot features of an input example into a lowdimensional representation, so that the existence of pivot features (features that are prominent in both domains and convey useful information for the NLP task) in the example can be decoded from that representation.The low-dimensional representation is then employed in a learning algorithm for the task.Moreover, we show how to inject pre-trained word embeddings into our model in order to improve generalization across examples with similar pivot features.We experiment with the task of cross-domain sentiment classification on 16 domain pairs and show substantial improvements over strong baselines.1 Yftah Ziser, Roi Reichart |
CoNLL | 2 |
| 2017 | Replicability Analysis for Natural Language Processing: Testing Significance with Multiple DatasetsabstractWith the ever growing amount of textual data from a large variety of languages, domains, and genres, it has become standard to evaluate NLP algorithms on multiple datasets in order to ensure a consistent performance across heterogeneous setups. However, such multiple comparisons pose significant challenges to traditional statistical analysis methods in NLP and can lead to erroneous conclusions. In this paper we propose a Replicability Analysis framework for a statistically sound analysis of multiple comparisons between algorithms for NLP tasks. We discuss the theoretical advantages of this framework over the current, statistically unjustified, practice in the NLP literature, and demonstrate its empirical value across four applications: multi-domain dependency parsing, multilingual POS tagging, cross-domain sentiment classification and word similarity prediction. Rotem Dror, Gili Baumer, Marina Bogomolov, Roi Reichart |
Trans. Assoc. Comput. Linguistics | 4 |
| 2017 | Semantic Specialization of Distributional Word Vector Spaces using Monolingual and Cross-Lingual ConstraintsabstractWe present Attract-Repel, an algorithm for improving the semantic quality of word vectors by injecting constraints extracted from lexical resources. Attract-Repel facilitates the use of constraints from mono- and cross-lingual resources, yielding semantically specialized cross-lingual vector spaces. Our evaluation shows that the method can make use of existing cross-lingual lexicons to construct high-quality vector spaces for a plethora of different languages, facilitating semantic transfer from high- to lower-resource ones. The effectiveness of our approach is demonstrated with state-of-the-art results on semantic similarity datasets in six languages. We next show that Attract-Repel-specialized vectors boost performance in the downstream task of dialogue state tracking (DST) across multiple languages. Finally, we show that cross-lingual vector spaces produced by our algorithm facilitate the training of multilingual DST models, which brings further performance improvements. Nikola Mrksic, Ivan Vulic, Diarmuid Ó Séaghdha, Ira Leviant, Roi Reichart, Milica Gasic, Anna Korhonen, Steve J. Young |
Trans. Assoc. Comput. Linguistics | 5 |
| 2016 | Edge-Linear First-Order Dependency Parsing with Undirected Minimum Spanning Tree InferenceabstractThe run time complexity of state-of-theart inference algorithms in graph-based dependency parsing is super-linear in the number of input words (n).Recently, pruning algorithms for these models have shown to cut a large portion of the graph edges, with minimal damage to the resulting parse trees.Solving the inference problem in run time complexity determined solely by the number of edges (m) is hence of obvious importance.We propose such an inference algorithm for first-order models, which encodes the problem as a minimum spanning tree (MST) problem in an undirected graph.This allows us to utilize state-of-the-art undirected MST algorithms whose run time is O(m) at expectation and with a very high probability.A directed parse tree is then inferred from the undirected MST and is subsequently improved with respect to the directed parsing model through local greedy updates, both steps running in O(n) time.In experiments with 18 languages, a variant of the first-order MSTParser (McDonald et al., 2005b) that employs our algorithm performs very similarly to the original parser that runs an O(n 2 ) directed MST inference. Effi Levi, Roi Reichart, Ari Rappoport |
ACL (1) | 2 |
| 2016 | Survey on the Use of Typological Information in Natural Language ProcessingabstractIn recent years linguistic typologies, which classify the world’s languages according to their functional and structural properties, have been widely used to support multilingual NLP. While the growing importance of typologies in supporting multilingual tasks has been recognised, no systematic survey of existing typological resources and their use in NLP has been published. This paper provides such a survey as well as discussion which we hope will both inform and inspire future work in the area. Helen O'Horan, Yevgeni Berzak, Ivan Vulic, Roi Reichart, Anna Korhonen |
COLING | 4 |
| 2016 | The Structured Weighted Violations Perceptron AlgorithmabstractWe present the Structured Weighted Violations Perceptron (SWVP) algorithm, a new perceptron algorithm for structured prediction, that generalizes the Collins Structured Perceptron (CSP, (Collins, 2002)). Unlike CSP, the update rule of SWVP explicitly exploits the internal structure of the predicted labels. We prove that for linearly separable training sets, SWVP converges to a weight vector that separates the data, under certain conditions on the parameters of the algorithm. We further prove bounds for SWVP on: (a) the number of updates in the separable case; (b) mistakes in the non-separable case; and (c) the probability to misclassify an unseen example (generalization), and show that for most SWVP variants these bounds are tighter than those of the CSP special case. In synthetic data experiments where data is drawn from a generative hidden variable model, SWVP provides substantial improvements over CSP. Rotem Dror, Roi Reichart |
EMNLP | 2 |
| 2016 | SimVerb-3500: A Large-Scale Evaluation Set of Verb SimilarityabstractVerbs play a critical role in the meaning of sentences, but these ubiquitous words have received little attention in recent distributional semantics research. We introduce SimVerb-3500, an evaluation resource that provides human ratings for the similarity of 3,500 verb pairs. SimVerb-3500 covers all normed verb types from the USF free-association database, providing at least three examples for every VerbNet class. This broad coverage facilitates detailed analyses of how syntactic and semantic phenomena together influence human understanding of verb meaning. Further, with significantly larger development and test sets than existing benchmarks, SimVerb-3500 enables more robust evaluation of representation learning architectures and promotes the development of methods tailored to verbs. We hope that SimVerb-3500 will enable a richer understanding of the diversity and complexity of verb semantics and guide the development of systems that can effectively represent and interpret this meaning. Daniela Gerz, Ivan Vulic, Felix Hill, Roi Reichart, Anna Korhonen |
EMNLP | 4 |
| 2016 | Effective Greedy Inference for Graph-based Non-Projective Dependency Parsing
Ilan Tchernowitz, Liron Yedidsion, Roi Reichart |
EMNLP | 3 |
| 2016 | Syntactic Parsing of Web Queries with Question IntentabstractAccurate automatic processing of Web queries is important for high-quality information retrieval from the Web.While the syntactic structure of a large portion of these queries is trivial, the structure of queries with question intent is much richer.In this paper we therefore address the task of statistical syntactic parsing of such queries.We first show that the standard dependency grammar does not account for the full range of syntactic structures manifested by queries with question intent.To alleviate this issue we extend the dependency grammar to account for segments -independent syntactic units within a potentially larger syntactic structure.We then propose two distant supervision approaches for the task.Both algorithms do not require manually parsed queries for training.Instead, they are trained on millions of (query, page title) pairs from the Community Question Answering (CQA) domain, where the CQA page was clicked by the user who initiated the query in a search engine.Experiments on a new treebank 1 consisting of 5,000 Web queries from the CQA domain, manually parsed using the proposed grammar, show that our algorithms outperform alternative approaches trained on various sources: tens of thousands of manually parsed OntoNotes sentences, millions of unlabeled CQA queries and thousands of manually segmented CQA queries. Yuval Pinter, Roi Reichart, Idan Szpektor |
HLT-NAACL | 2 |
| 2016 | Symmetric Patterns and Coordinations: Fast and Enhanced Representations of Verbs and AdjectivesabstractState-of-the-art word embeddings, which are often trained on bag-of-words (BOW) contexts, provide a high quality representation of aspects of the semantics of nouns.However, their quality decreases substantially for the task of verb similarity prediction.In this paper we show that using symmetric pattern contexts (SPs, e.g., "X and Y") improves word2vec verb similarity performance by up to 15% and is also instrumental in adjective similarity prediction.The unsupervised SP contexts are even superior to a variety of dependency contexts extracted using a supervised dependency parser.Moreover, we observe that SPs and dependency coordination contexts (Coor) capture a similar type of information, and demonstrate that Coor contexts are superior to other dependency contexts including the set of all dependency contexts, although they are still inferior to SPs.Finally, there are substantially fewer SP contexts compared to alternative representations, leading to a massive reduction in training time.On an 8G words corpus and a 32 core machine, the SP model trains in 11 minutes, compared to 5 and 11 hours with BOW and all dependency contexts, respectively. Roy Schwartz 0001, Roi Reichart, Ari Rappoport |
HLT-NAACL | 2 |
| 2015 | Contrastive Analysis with Predictive Power: Typology Driven Estimation of Grammatical Error Distributions in ESLabstractThis work examines the impact of crosslinguistic transfer on grammatical errors in English as Second Language (ESL) texts.Using a computational framework that formalizes the theory of Contrastive Analysis (CA), we demonstrate that language specific error distributions in ESL writing can be predicted from the typological properties of the native language and their relation to the typology of English.Our typology driven model enables to obtain accurate estimates of such distributions without access to any ESL data for the target languages.Furthermore, we present a strategy for adjusting our method to low-resource languages that lack typological documentation using a bootstrapping approach which approximates native language typology from ESL texts.Finally, we show that our framework is instrumental for linguistic inquiry seeking to identify first language factors that contribute to a wide range of difficulties in second language acquisition. Yevgeni Berzak, Roi Reichart, Boris Katz |
CoNLL | 2 |
| 2015 | Symmetric Pattern Based Word Embeddings for Improved Word Similarity PredictionabstractWe present a novel word level vector representation based on symmetric patterns (SPs).For this aim we automatically acquire SPs (e.g., "X and Y") from a large corpus of plain text, and generate vectors where each coordinate represents the cooccurrence in SPs of the represented word with another word of the vocabulary.Our representation has three advantages over existing alternatives: First, being based on symmetric word relationships, it is highly suitable for word similarity prediction.Particularly, on the SimLex999 word similarity dataset, our model achieves a Spearman's ρ score of 0.517, compared to 0.462 of the state-of-the-art word2vec model.Interestingly, our model performs exceptionally well on verbs, outperforming stateof-the-art baselines by 20.2-41.5%.Second, pattern features can be adapted to the needs of a target NLP application.For example, we show that we can easily control whether the embeddings derived from SPs deem antonym pairs (e.g.(big,small)) as similar or dissimilar, an important distinction for tasks such as word classification and sentiment analysis.Finally, we show that a simple combination of the word similarity scores generated by our method and by word2vec results in a superior predictive power over that of each individual model, scoring as high as 0.563 in Spearman's ρ on SimLex999.This emphasizes the differences between the signals captured by each of the models. Roy Schwartz 0001, Roi Reichart, Ari Rappoport |
CoNLL | 2 |
| 2015 | SimLex-999: Evaluating Semantic Models With (Genuine) Similarity EstimationabstractWe present SimLex-999, a gold standard resource for evaluating distributional semantic models that improves on existing resources in several important ways. First, in contrast to gold standards such as WordSim-353 and MEN, it explicitly quantifies similarity rather than association or relatedness so that pairs of entities that are associated but not actually similar (Freud, psychology) have a low rating. We show that, via this focus on similarity, SimLex-999 incentivizes the development of models with a different, and arguably wider, range of applications than those which reflect conceptual association. Second, SimLex-999 contains a range of concrete and abstract adjective, noun, and verb pairs, together with an independent rating of concreteness and (free) association strength for each pair. This diversity enables fine-grained analyses of the performance of models on concepts of different types, and consequently greater insight into how architectures can be improved. Further, unlike existing gold standard evaluations, for which automatic approaches have reached or surpassed the inter-annotator agreement ceiling, state-of-the-art models perform well below this ceiling on SimLex-999. There is therefore plenty of scope for SimLex-999 to quantify future improvements to distributional semantic models, guiding the development of the next generation of representation-learning architectures. Felix Hill, Roi Reichart, Anna Korhonen |
Comput. Linguistics | 2 |
| 2015 | Unsupervised Declarative Knowledge Induction for Constraint-Based Learning of Information Structure in Scientific DocumentsabstractInferring the information structure of scientific documents is useful for many NLP applications. Existing approaches to this task require substantial human effort. We propose a framework for constraint learning that reduces human involvement considerably. Our model uses topic models to identify latent topics and their key linguistic features in input documents, induces constraints from this information and maps sentences to their dominant information structure categories through a constrained unsupervised model. When the induced constraints are combined with a fully unsupervised model, the resulting model challenges existing lightly supervised feature-based models as well as unsupervised models that use manually constructed declarative knowledge. Our results demonstrate that useful declarative knowledge can be learned from data with very limited human involvement. Roi Reichart, Anna Korhonen |
Trans. Assoc. Comput. Linguistics | 2 |
| 2014 | Minimally Supervised Classification to Semantic Categories using Automatically Acquired Symmetric Patterns
Roy Schwartz 0001, Roi Reichart, Ari Rappoport |
COLING | 2 |
| 2014 | Reconstructing Native Language Typology from Foreign Language UsageabstractThis work was supported by the Center for Brains, Minds and Machines
(CBMM), funded by NSF STC award CCF - 1231216. Yevgeni Berzak, Roi Reichart, Boris Katz |
CoNLL | 2 |
| 2014 | An Unsupervised Model for Instance Level Subcategorization AcquisitionabstractMost existing systems for subcategorization frame (SCF) acquisition rely on supervised parsing and infer SCF distributions at type, rather than instance level.These systems suffer from poor portability across domains and their benefit for NLP tasks that involve sentence-level processing is limited.We propose a new unsupervised, Markov Random Field-based model for SCF acquisition which is designed to address these problems.The system relies on supervised POS tagging rather than parsing, and is capable of learning SCFs at instance level.We perform evaluation against gold standard data which shows that our system outperforms several supervised and type-level SCF baselines.We also conduct task-based evaluation in the context of verb similarity prediction, demonstrating that a vector space model based on our SCFs substantially outperforms a lexical model and a model based on a supervised parser 1 . Simon Baker, Roi Reichart, Anna Korhonen |
EMNLP | 2 |
| 2014 | Multi-Modal Models for Concrete and Abstract Concept MeaningabstractMulti-modal models that learn semantic representations from both linguistic and perceptual input outperform language-only models on a range of evaluations, and better reflect human concept acquisition. Most perceptual input to such models corresponds to concrete noun concepts and the superiority of the multi-modal approach has only been established when evaluating on such concepts. We therefore investigate which concepts can be effectively learned by multi-modal models. We show that concreteness determines both which linguistic features are most informative and the impact of perceptual input in such models. We then introduce ridge regression as a means of propagating perceptual information from concrete nouns to more abstract concepts that is more robust than previous approaches. Finally, we present weighted gram matrix combination, a means of combining representations from distinct modalities that outperforms alternatives when both modalities are sufficiently rich. Felix Hill, Roi Reichart, Anna Korhonen |
Trans. Assoc. Comput. Linguistics | 2 |
| 2013 | Improved Lexical Acquisition through DPP-based Verb Clustering
Roi Reichart, Anna Korhonen |
ACL (1) | 1 |
| 2013 | Improved Information Structure Analysis of Scientific Documents Through Discourse and Lexical Constraints
Roi Reichart, Anna Korhonen |
HLT-NAACL | 2 |
| 2012 | A Diverse Dirichlet Process Ensemble for Unsupervised Induction of Syntactic Categories
Roi Reichart, Gal Elidan, Ari Rappoport |
COLING | 1 |
| 2012 | Improved Parsing and POS Tagging Using Inter-Sentence Consistency Constraints
Alexander M. Rush, Roi Reichart, Michael Collins 0001, Amir Globerson |
EMNLP-CoNLL | 2 |
| 2012 | Learning to Map into a Universal POS Tagset
Yuan Zhang 0001, Roi Reichart, Regina Barzilay, Amir Globerson |
EMNLP-CoNLL | 2 |
| 2012 | You Too?! Mixed-Initiative LDA Story Matching to Help Teens in Distress
Karthik Dinakar, Birago Jones, Henry Lieberman, Rosalind W. Picard, Carolyn P. Rosé, Matthew Thoman, Roi Reichart |
ICWSM | 7 |
| 2012 | Multi-Event Extraction Guided by Global Constraints
Roi Reichart, Regina Barzilay |
HLT-NAACL | 1 |
| 2011 | Confidence Driven Unsupervised Semantic Parsing
Dan Goldwasser, Roi Reichart, James Clarke, Dan Roth 0001 |
ACL | 2 |
| 2011 | Neutralizing Linguistically Problematic Annotations in Unsupervised Dependency Parsing Evaluation
Roy Schwartz 0001, Omri Abend, Roi Reichart, Ari Rappoport |
ACL | 3 |
| 2010 | Improved Unsupervised POS Induction through Prototype Discovery
Omri Abend, Roi Reichart, Ari Rappoport |
ACL | 2 |
| 2010 | Type Level Clustering Evaluation: New Measures and a POS Induction Case Study
Roi Reichart, Omri Abend, Ari Rappoport |
CoNLL | 1 |
| 2010 | Improved Unsupervised POS Induction Using Intrinsic Clustering Quality and a Zipfian Constraint
Roi Reichart, Raanan Fattal, Ari Rappoport |
CoNLL | 1 |
| 2010 | Tense Sense Disambiguation: A New Syntactic Polysemy Task
Roi Reichart, Ari Rappoport |
EMNLP | 1 |
| 2010 | Improved Fully Unsupervised Parsing with Zoomed Learning
Roi Reichart, Ari Rappoport |
EMNLP | 1 |
| 2009 | Unsupervised Argument Identification for Semantic Role Labeling
Omri Abend, Roi Reichart, Ari Rappoport |
ACL/IJCNLP | 2 |
| 2009 | Superior and Efficient Fully Unsupervised Pattern-based Concept Acquisition Using an Unsupervised Parser
Dmitry Davidov, Roi Reichart, Ari Rappoport |
CoNLL | 2 |
| 2009 | Sample Selection for Statistical Parsers: Cognitively Driven Algorithms and Evaluation Measures
Roi Reichart, Ari Rappoport |
CoNLL | 1 |
| 2009 | Automatic Selection of High Quality Parses Created By a Fully Unsupervised Parser
Roi Reichart, Ari Rappoport |
CoNLL | 1 |
| 2009 | The NVI Clustering Evaluation Measure
Roi Reichart, Ari Rappoport |
CoNLL | 1 |
| 2008 | Multi-Task Active Learning for Linguistic Annotations
Roi Reichart, Katrin Tomanek, Udo Hahn, Ari Rappoport |
ACL | 1 |
| 2008 | Extraction of Entailed Semantic Relations Through Syntax-Based Comma Resolution
Vivek Srikumar, Roi Reichart, Mark Sammons, Ari Rappoport, Dan Roth 0001 |
ACL | 2 |
| 2008 | A Supervised Algorithm for Verb Disambiguation into VerbNet Classes
Omri Abend, Roi Reichart, Ari Rappoport |
COLING | 2 |
| 2008 | Unsupervised Induction of Labeled Parse Trees by Clustering with Syntactic Features
Roi Reichart, Ari Rappoport |
COLING | 1 |
| 2007 | An Ensemble Method for Selection of High Quality Parses
Roi Reichart, Ari Rappoport |
ACL | 1 |
| 2007 | Self-Training for Enhancement and Domain Adaptation of Statistical Parsers Trained on Small Datasets
Roi Reichart, Ari Rappoport |
ACL | 1 |