Reut Tsarfaty

dblp:21/3716 · DBLP profile ↗
← Back
45ranked-venue papers
9as first author
25since 2021 · last 2026
0000-0001-6406-6568ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 45 · 9 first-author · 25 since 2021
YearPublicationVenuePosition
2026 Beyond Word Boundaries: A Hebrew Coreference Benchmark and an Evaluation Protocol for Morphologically Complex Text
abstract
Coreference Resolution (CR) is a fundamental NLP task critical for long-form tasks as information extraction, summarization, and many business applications.However, CR methods originally designed for English struggle with Morphologically Rich Languages (MRLs), where mention boundaries do not necessarily align with word boundaries, and a single token may consist of multiple anaphors.CR modeling and evaluation protocols standardly assume that, as in English, words and mentions mostly align.However, this assumption breaks down in MRLs, particularly in the context of LLMs' raw-text processing and endto-end tasks.To assess and address this challenge, we introduce KibutzR, the first comprehensive CR dataset for Modern Hebrew, an MRL rich with complex words and pronominal clitics.We deliver an annotated dataset that identifies mentions at word, sub-word and multi-word levels, and propose an evaluation protocol that directly addresses word/morpheme boundary discrepancies.Our experiments show that contemporary LLMs perform significantly worse on Hebrew than on English, and that performance degrades on raw unsegmented text.Crucially, we show an inverse performance-trend in Hebrew relative to English, where smaller encoders perform far better than contemporary decoder models, leaving ample space for investigation and improvement.We deliver a new benchmark for Hebrew coreference resolution and a segmentation-aware evaluation protocol to inform future work on other MRLs.
Refael Shaked Greenfeld, Reut Tsarfaty
ACL (1)2
2026 Location Not Found: Exposing Implicit Local and Global Biases in Multilingual LLMs
abstract
Guy Mor-Lan, Omer Goldman, Matan Eyal, Adi Mayrav Gilady, Sivan Eiger, Idan Szpektor, Avinatan Hassidim, Yossi Matias, Reut Tsarfaty. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Guy Mor, Omer Goldman, Matan Eyal, Adi Mayrav Gilady, Sivan Eiger, Idan Szpektor, Avinatan Hassidim, Yossi Matias, Reut Tsarfaty
ACL (1)9
2026 From Benchmarks to Skills: Low-Rank Factors for LLM Evaluation
abstract
Abstract Current evaluations of large language models (LLMs) rely heavily on a growing collection of benchmarks and on aggregate benchmark scores, yet it remains unclear what this comparison actually captures, and what these scores reveal about models’ underlying capabilities. Here, we propose a new paradigm for LLM evaluation, by asking whether benchmark performance reflects many independent abilities, or rather, relies on a small number of shared dimensions. To answer this, we apply Factor Analysis (FA) to a massive performance matrix of LLMs versus benchmarks (60 × 44) revealing an intrinsically low-rank structure of that matrix. That is, a small number of latent factors captures most of the structure in the full task space. This low-rank geometry reveals substantial redundancy across existing tasks and explains why many benchmarks appear to be measuring overlapping abilities. We further show that these latent factors correspond to coherent, skill-like, dimensions of LLM behavior. Leveraging this latent skill-space, we deliver three practical tools for LLM evaluation and downstream users: (i) identifying redundant tasks, (ii) profiling new models using a small subset of tasks, and (iii) selecting models aligned with desired skill profiles. Our method provides a solid alternative to the de-facto standard of a single aggregate score, and establishes an interpretable and practical framework for understanding and benchmarking LLM core capabilities.
Aviya Maimon, Amir David Nisan Cohen, Gal Vishne, Shauli Ravfogel, Reut Tsarfaty
Trans. Assoc. Comput. Linguistics5
2026 MoNaCo: More Natural and Complex Questions for Reasoning Across Dozens of Documents
abstract
Abstract Automated agents, powered by large language models (LLMs), are emerging as the go-to tool for querying information. However, evaluation benchmarks for LLM agents rarely feature natural questions that are both information-seeking and genuinely time-consuming for humans. To address this gap we introduce MoNaCo, a benchmark of 1,315 natural and time-consuming questions that require dozens, and at times hundreds, of intermediate steps to solve— far more than any existing QA benchmark. To build MoNaCo, we developed a decomposed annotation pipeline to elicit and manually answer real-world time-consuming questions at scale. Frontier LLMs evaluated on MoNaCo achieve at most 61.2% F1, hampered by low recall and hallucinations. Our results underscore the limitations of LLM-powered agents in handling the complexity and sheer breadth of real-world information-seeking tasks—with MoNaCo providing an effective resource for tracking such progress. The MoNaCo benchmark, codebase, prompts, and models predictions are all publicly available at: https://tomerwolgithub.github.io/monaco.
Tomer Wolfson, Harsh Trivedi, Mor Geva, Yoav Goldberg, Dan Roth 0001, Tushar Khot, Ashish Sabharwal, Reut Tsarfaty
Trans. Assoc. Comput. Linguistics8
2025 Beyond N-Grams: Rethinking Evaluation Metrics and Strategies for Multilingual Abstractive Summarization
abstract
Automatic N-gram based metrics such as ROUGE are widely used for evaluating generative tasks such as summarization.While these metrics are considered indicative (even if imperfect), of human evaluation for English, their suitability for other languages remains unclear.To address this, in this paper we systematically assess evaluation metrics for generation -both n-gram-based and neural-based -to assess their effectiveness across languages and tasks.Specifically, we design a large-scale evaluation suite across eight languages from four typological families -agglutinative, isolating, low-fusional, and high-fusional -from both low-and high-resource languages, to analyze their correlations with human judgments.Our findings highlight the sensitivity of the evaluation metric to the language type at hand.For example, for fusional languages, n-grambased metrics demonstrate a lower correlation with human assessments, compared to isolating and agglutinative languages.We also demonstrate that tokenization considerations can significantly mitigate this for fusional languages with rich morphology, up to reversing such negative correlations.Additionally, we show that neural-based metrics specifically trained for evaluation, such as COMET, consistently outperform other neural metrics and correlate better than n-grams metrics with human judgments in low-resource languages.Overall, our analysis highlights the limitations of n-gram metrics for fusional languages and advocates for investment in neural-based metrics trained for evaluation tasks. 1
Itai Mondshine, Tzuf Paz-Argaman, Reut Tsarfaty
ACL (1)3
2025 Superlatives in Context: Modeling the Implicit Semantics of Superlatives
abstract
Valentina Pyatkin, Bonnie Webber, Ido Dagan, Reut Tsarfaty. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Valentina Pyatkin, Bonnie L. Webber, Ido Dagan, Reut Tsarfaty
NAACL (Long Papers)4
2024 A Truly Joint Neural Architecture for Segmentation and Parsing
abstract
Contemporary multilingual dependency parsers can parse a diverse set of languages, but for Morphologically Rich Languages (MRLs), performance is attested to be lower than other languages.The key challenge is that, due to high morphological complexity and ambiguity of the space-delimited input tokens, the linguistic units that act as nodes in the tree are not known in advance.Pre-neural dependency parsers for MRLs subscribed to the joint morpho-syntactic hypothesis, stating that morphological segmentation and syntactic parsing should be solved jointly, rather than as a pipeline where segmentation precedes parsing.However, neural stateof-the-art parsers to date use a strict pipeline.In this paper we introduce a joint neural architecture where a lattice-based representation preserving all morphological ambiguity of the input is provided to an arc-factored model, which then solves the morphological segmentation and syntactic parsing tasks at once.Our experiments on Hebrew, a rich and highly ambiguous MRL, demonstrate state-of-the-art performance on parsing, tagging and segmentation of the Hebrew section of UD, using a single model.This proposed architecture is LLM-based and language agnostic, providing a solid foundation for MRLs to obtain further performance improvements and bridge the gap with other languages.
Danit Yshaayahu Levi, Reut Tsarfaty
EACL (1)2
2024 Where Do We Go From Here? Multi-scale Allocentric Relational Inferencefrom Natural Spatial Descriptions
abstract
Tzuf Paz-Argaman, John Palowitch, Sayali Kulkarni, Jason Baldridge, Reut Tsarfaty. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Tzuf Paz-Argaman, John Palowitch, Sayali Kulkarni, Jason Baldridge, Reut Tsarfaty
EACL (1)5
2024 Is It Really Long Context if All You Need Is Retrieval? Towards Genuinely Difficult Long Context NLP
abstract
Improvements in language models' capabilities have pushed their applications towards longer contexts, making long-context evaluation and development an active research area.However, many disparate use cases are grouped together under the umbrella term of "long-context", defined simply by the total length of the model's input, including -for example -Needle-in-a-Haystack tasks, book summarization, and information aggregation.Given their varied difficulty, in this position paper we argue that conflating different tasks by their context length is unproductive.As a community, we require a more precise vocabulary to understand what makes long-context tasks similar or different.We propose to unpack the taxonomy of longcontext based on the properties that make them more difficult with longer contexts.We propose two orthogonal axes of difficulty: (I) Dispersion: How hard is it to find the necessary information in the context?(II) Scope: How much necessary information is there to find?We survey the literature on long context, provide justification for this taxonomy as an informative descriptor, and situate the literature with respect to it.We conclude that the most difficult and interesting settings, whose necessary information is very long and highly dispersed within the input, is severely under-explored.By using a descriptive vocabulary and discussing the relevant properties of difficulty in long context, we can implement more informed research in this area.We call for a careful design of tasks and benchmarks with distinctly long context, taking into account the characteristics that make it qualitatively different from shorter context.
Omer Goldman, Alon Jacovi, Aviv Slobodkin, Aviya Maimon, Ido Dagan, Reut Tsarfaty
EMNLP6
2024 NoviCode: Generating Programs from Natural Language Utterances by Novices
abstract
Abstract Current Text-to-Code models demonstrate impressive capabilities in generating executable code from natural language snippets. However, current studies focus on technical instructions and programmer-oriented language, and it is an open question whether these models can effectively translate natural language descriptions given by non-technical users and express complex goals, to an executable program that contains an intricate flow—composed of API access and control structures as loops, conditions, and sequences. To unlock the challenge of generating a complete program from a plain non-technical description we present NoviCode, a novel NL Programming task, which takes as input an API and a natural language description by a novice non-programmer, and provides an executable program as output. To assess the efficacy of models on this task, we provide a novel benchmark accompanied by test suites wherein the generated program code is assessed not according to their form, but according to their functional execution. Our experiments show that, first, NoviCode is indeed a challenging task in the code synthesis domain, and that generating complex code from non-technical instructions goes beyond the current Text-to-Code paradigm. Second, we show that a novel approach wherein we align the NL utterances with the compositional hierarchical structure of the code, greatly enhances the performance of LLMs on this task, compared with the end-to-end Text-to-Code counterparts.
Asaf Achi Mordechai, Yoav Goldberg, Reut Tsarfaty
Trans. Assoc. Comput. Linguistics3
2023 Conjunct Resolution in the Face of Verbal Omissions
abstract
Verbal omissions are complex syntactic phenomena in VP coordination structures.They occur when verbs and (some of) their arguments are omitted from subsequent clauses after being explicitly stated in an initial clause.Recovering these omitted elements is necessary for accurate interpretation of the sentence, and while humans easily and intuitively fill in the missing information, state-of-the-art models continue to struggle with this task.Previous work is limited to small-scale datasets, synthetic data creation methods, and to resolution methods in the dependency-graph level.In this work we propose a conjunct resolution task that operates directly on the text and makes use of a split-andrephrase paradigm in order to recover the missing elements in the coordination structure.To this end, we first formulate a pragmatic framework of verbal omissions which describes the different types of omissions, and develop an automatic scalable collection method.Based on this method, we curate a large dataset, containing over 10K examples of naturally-occurring verbal omissions with crowd-sourced annotations of the resolved conjuncts.We train various neural baselines for this task, and show that while our best method obtains decent performance, it leaves ample space for improvement.We propose our dataset, metrics and models as a starting point for future research on this topic.
Royi Rassin, Yoav Goldberg, Reut Tsarfaty
ACL (1)3
2023 Do Pretrained Contextual Language Models Distinguish between Hebrew Homograph Analyses?
abstract
Semitic morphologically-rich languages (MRLs) are characterized by extreme word ambiguity.Because most vowels are omitted in standard texts, many of the words are homographs with multiple possible analyses, each with a different pronunciation and different morphosyntactic properties.This ambiguity goes beyond word-sense disambiguation (WSD), and may include token segmentation into multiple word units.Previous research on MRLs claimed that standardly trained pre-trained language models (PLMs) based on word-pieces may not sufficiently capture the internal structure of such tokens in order to distinguish between these analyses.Taking Hebrew as a case study, we investigate the extent to which Hebrew homographs can be disambiguated and analyzed using PLMs.We evaluate all existing models for contextualized Hebrew embeddings on a novel Hebrew homograph challenge sets that we deliver.Our empirical results demonstrate that contemporary Hebrew contextualized embeddings outperform non-contextualized embeddings; and that they are most effective for disambiguating segmentation and morphosyntactic features, less so regarding pure word-sense disambiguation.We show that these embeddings are more effective when the number of word-piece splits is limited, and they are more effective for 2-way and 3-way ambiguities than for 4-way ambiguity.We show that the embeddings are equally effective for homographs of both balanced and skewed distributions, whether calculated as masked or unmasked tokens.Finally, we show that these embeddings are as effective for homograph disambiguation with extensive supervised training as with a few-shot setup.
Avi Shmidman, Cheyn Shmuel Shmidman, Dan Bareket, Moshe Koppel, Reut Tsarfaty
EACL5
2023 COHESENTIA: A Novel Benchmark of Incremental versus Holistic Assessment of Coherence in Generated Texts
abstract
Coherence is a linguistic term that refers to the relations between small textual units (sentences, propositions), which make the text logically consistent and meaningful to the reader.With the advances of generative foundational models in NLP, there is a pressing need to automatically assess the human-perceived coherence of automatically generated texts.Up until now, little work has been done on explicitly assessing the coherence of generated texts and analyzing the factors contributing to (in)coherence.Previous work on the topic used other tasks, e.g., sentence reordering, as proxies of coherence, rather than approaching coherence detection heads on.In this paper, we introduce COHESENTIA, a novel benchmark of humanperceived coherence of automatically generated texts.Our annotation protocol reflects two perspectives; one is global, assigning a single coherence score, and the other is incremental, scoring sentence by sentence.The incremental method produces an (in)coherence score for each text fragment and also pinpoints reasons for incoherence at that point.Our benchmark contains 500 automatically-generated and human-annotated paragraphs, each annotated in both methods, by multiple raters.Our analysis shows that the inter-annotator agreement in the incremental mode is higher than in the holistic alternative, and our experiments show that standard LMs fine-tuned for coherence detection show varied performance on the different factors contributing to (in)coherence.All in all, these models yield unsatisfactory performance, emphasizing the need for developing more reliable methods for coherence assessment.
Aviya Maimon, Reut Tsarfaty
EMNLP2
2023 Design Choices for Crowdsourcing Implicit Discourse Relations: Revealing the Biases Introduced by Task Design
abstract
Abstract Disagreement in natural language annotation has mostly been studied from a perspective of biases introduced by the annotators and the annotation frameworks. Here, we propose to analyze another source of bias—task design bias, which has a particularly strong impact on crowdsourced linguistic annotations where natural language is used to elicit the interpretation of lay annotators. For this purpose we look at implicit discourse relation annotation, a task that has repeatedly been shown to be difficult due to the relations’ ambiguity. We compare the annotations of 1,200 discourse relations obtained using two distinct annotation tasks and quantify the biases of both methods across four different domains. Both methods are natural language annotation tasks designed for crowdsourcing. We show that the task design can push annotators towards certain relations and that some discourse relation senses can be better elicited with one or the other annotation approach. We also conclude that this type of bias should be taken into account when training and testing models.
Valentina Pyatkin, Frances Yung, Merel C. J. Scholman, Reut Tsarfaty, Ido Dagan, Vera Demberg
Trans. Assoc. Comput. Linguistics4
2022 AlephBERT: Language Model Pre-training and Evaluation from Sub-Word to Sentence Level
abstract
Amit Seker, Elron Bandel, Dan Bareket, Idan Brusilovsky, Refael Greenfeld, Reut Tsarfaty. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Amit Seker, Elron Bandel, Dan Bareket, Idan Brusilovsky, Refael Shaked Greenfeld, Reut Tsarfaty
ACL (1)6
2022 Breakpoint Transformers for Modeling and Tracking Intermediate Beliefs
abstract
Can we teach natural language understanding models to track their beliefs through intermediate points in text?We propose a representation learning framework called breakpoint modeling that allows for learning of this type.Given any text encoder and data marked with intermediate states (breakpoints) along with corresponding textual queries viewed as true/false propositions (i.e., the candidate beliefs of a model, consisting of information changing through time) our approach trains models in an efficient and end-to-end fashion to build intermediate representations that facilitate teaching and direct querying of beliefs at arbitrary points alongside solving other end tasks.To show the benefit of our approach, we experiment with a diverse set of NLU tasks including relational reasoning on CLUTRR and narrative understanding on bAbI.Using novel belief prediction tasks for both tasks, we show the benefit of our main breakpoint transformer, based on T5, over conventional representation learning approaches in terms of processing efficiency, prediction accuracy and prediction consistency, all with minimal to no effect on corresponding QA endtasks.To show the feasibility of incorporating our belief tracker into more complex reasoning pipelines, we also obtain SOTA performance on the three-tiered reasoning challenge for the TRIP benchmark (around 23-32% absolute improvement on Tasks 2-3). 1
Kyle Richardson 0001, Ronen Tamari, Oren Sultan, Dafna Shahaf, Reut Tsarfaty, Ashish Sabharwal
EMNLP5
2022 UniMorph 4.0: Universal Morphology
abstract
The Universal Morphology (UniMorph) project is a collaborative effort providing broad-coverage instantiated normalized morphological inflection tables for hundreds of diverse world languages. The project comprises two major thrusts: a language-independent feature schema for rich morphological annotation, and a type-level resource of annotated data in diverse languages realizing that schema. This paper presents the expansions and improvements on several fronts that were made in the last couple of years (since McCarthy et al. (2020)). Collaborative efforts by numerous linguists have added 66 new languages, including 24 endangered languages. We have implemented several improvements to the extraction pipeline to tackle some issues, e.g., missing gender and macrons information. We have amended the schema to use a hierarchical structure that is needed for morphological phenomena like multiple-argument agreement and case stacking, while adding some missing morphological features to make the schema more inclusive. In light of the last UniMorph release, we also augmented the database with morpheme segmentation for 16 languages. Lastly, this new release makes a push towards inclusion of derivational morphology in UniMorph by enriching the data and annotation schema with instances representing derivational processes from MorphyNet.
Khuyagbaatar Batsuren, Omer Goldman, Salam Khalifa, Nizar Habash, Witold Kieras, Gábor Bella, Brian Leonard, Garrett Nicolai, Kyle Gorman, Yustinus Ghanggo Ate, Maria Ryskina, Sabrina J. Mielke, Elena Budianskaya, Charbel El-Khaissi, Tiago Pimentel, Michael Gasser, William Lane 0002, Mohit Raj, Matt Coler, Jaime Rafael Montoya Samame, Delio Siticonatzi Camaiteri, Esaú Zumaeta Rojas, Didier López Francis, Arturo Oncevay, Juan López Bautista, Gema Celeste Silva Villegas, Lucas Torroba Hennigen, Adam Ek, David Guriel, Peter Dirix, Jean-Philippe Bernardy, Andrey Scherbakov, Aziyana Bayyr-ool, Antonios Anastasopoulos, Roberto Zariquiey, Karina Sheifer, Sofya Ganieva, Hilaria Cruz, Ritván Karahóga, Stella Markantonatou, George Pavlidis, Matvey Plugaryov, Elena Klyachko, Ali Salehi, Candy Angulo, Jatayu Baxi, Andrew Krizhanovsky, Natalia Krizhanovskaya, Elizabeth Salesky, Clara Vania, Sardana Ivanova, Jennifer C. White, Rowan Hall Maudslay, Josef Valvoda, Ran Zmigrod, Paula Czarnowska, Irene Nikkarinen, Aelita Salchak, Brijesh Bhatt, Christopher Straughn, Zoey Liu, Jonathan Washington, Yuval Pinter, Duygu Ataman, Marcin Wolinski, Totok Suhardijanto, Anna Yablonskaya, Niklas Stoehr, Hossep Dolatian, Zahroh Nuriah, Shyam Ratan, Francis M. Tyers, Edoardo Maria Ponti, Grant Aiton, Aryaman Arora, Richard J. Hatcher, Ritesh Kumar 0002, Jeremiah Young, Daria Rodionova, Anastasia Yemelina, Taras Andrushko, Igor Marchenko, Polina Mashkovtseva, Alexandra Serova, Emily Tucker Prud'hommeaux, Maria Nepomniashchaya, Fausto Giunchiglia, Eleanor Chodroff, Mans Hulden, Miikka Silfverberg, Arya McCarthy, David Yarowsky, Ryan Cotterell, Reut Tsarfaty, Ekaterina Vylomova
LREC94
2022 Design Choices in Crowdsourcing Discourse Relation Annotations: The Effect of Worker Selection and Training
abstract
Obtaining linguistic annotation from novice crowdworkers is far from trivial. A case in point is the annotation of discourse relations, which is a complicated task. Recent methods have obtained promising results by extracting relation labels from either discourse connectives (DCs) or question-answer (QA) pairs that participants provide. The current contribution studies the effect of worker selection and training on the agreement on implicit relation labels between workers and gold labels, for both the DC and the QA method. In Study 1, workers were not specifically selected or trained, and the results show that there is much room for improvement. Study 2 shows that a combination of selection and training does lead to improved results, but the method is cost- and time-intensive. Study 3 shows that a selection-only approach is a viable alternative; it results in annotations of comparable quality compared to annotations from trained participants. The results generalized over both the DC and QA method and therefore indicate that a selection-only approach could also be effective for other crowdsourced discourse annotation tasks.
Merel C. J. Scholman, Valentina Pyatkin, Frances Yung, Ido Dagan, Reut Tsarfaty, Vera Demberg
LREC5
2022 Text-based NP Enrichment
abstract
Abstract Understanding the relations between entities denoted by NPs in a text is a critical part of human-like natural language understanding. However, only a fraction of such relations is covered by standard NLP tasks and benchmarks nowadays. In this work, we propose a novel task termed text-based NP enrichment (TNE), in which we aim to enrich each NP in a text with all the preposition-mediated relations—either explicit or implicit—that hold between it and other NPs in the text. The relations are represented as triplets, each denoted by two NPs related via a preposition. Humans recover such relations seamlessly, while current state-of-the-art models struggle with them due to the implicit nature of the problem. We build the first large-scale dataset for the problem, provide the formal framing and scope of annotation, analyze the data, and report the results of fine-tuned language models on the task, demonstrating the challenge it poses to current technology. A webpage with a data-exploration UI, a demo, and links to the code, models, and leaderboard, to foster further research into this challenging problem can be found at: yanaiela.github.io/TNE/.
Yanai Elazar, Victoria Basmova, Yoav Goldberg, Reut Tsarfaty
Trans. Assoc. Comput. Linguistics4
2022 Morphology Without Borders: Clause-Level Morphology
abstract
Abstract Morphological tasks use large multi-lingual datasets that organize words into inflection tables, which then serve as training and evaluation data for various tasks. However, a closer inspection of these data reveals profound cross-linguistic inconsistencies, which arise from the lack of a clear linguistic and operational definition of what is a word, and which severely impair the universality of the derived tasks. To overcome this deficiency, we propose to view morphology as a clause-level phenomenon, rather than word-level. It is anchored in a fixed yet inclusive set of features, that encapsulates all functions realized in a saturated clause. We deliver MightyMorph, a novel dataset for clause-level morphology covering 4 typologically different languages: English, German, Turkish, and Hebrew. We use this dataset to derive 3 clause-level morphological tasks: inflection, reinflection and analysis. Our experiments show that the clause-level tasks are substantially harder than the respective word-level tasks, while having comparable complexity across languages. Furthermore, redefining morphology to the clause-level provides a neat interface with contextualized language models (LMs) and allows assessing the morphological knowledge encoded in these models and their usability for morphological tasks. Taken together, this work opens up new horizons in the study of computational morphology, leaving ample space for studying neural morphology cross-linguistically.
Omer Goldman, Reut Tsarfaty
Trans. Assoc. Comput. Linguistics2
2022 Draw Me a Flower: Processing and Grounding Abstraction in Natural Language
abstract
Abstract Abstraction is a core tenet of human cognition and communication. When composing natural language instructions, humans naturally evoke abstraction to convey complex procedures in an efficient and concise way. Yet, interpreting and grounding abstraction expressed in NL has not yet been systematically studied in NLP, with no accepted benchmarks specifically eliciting abstraction in NL. In this work, we set the foundation for a systematic study of processing and grounding abstraction in NLP. First, we deliver a novel abstraction elicitation method and present Hexagons, a 2D instruction-following game. Using Hexagons we collected over 4k naturally occurring visually-grounded instructions rich with diverse types of abstractions. From these data, we derive an instruction-to-execution task and assess different types of neural models. Our results show that contemporary models and modeling practices are substantially inferior to human performance, and that model performance is inversely correlated with the level of abstraction, showing less satisfying performance on higher levels of abstraction. These findings are consistent across models and setups, confirming that abstraction is a challenging phenomenon deserving further attention and study in NLP/AI research.
Royi Lachmy, Valentina Pyatkin, Avshalom Manevich, Reut Tsarfaty
Trans. Assoc. Comput. Linguistics4
2021 The Possible, the Plausible, and the Desirable: Event-Based Modality Detection for Language Processing
abstract
Valentina Pyatkin, Shoval Sadde, Aynat Rubinstein, Paul Portner, Reut Tsarfaty. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Valentina Pyatkin, Shoval Sadde, Aynat Rubinstein, Paul Portner, Reut Tsarfaty
ACL/IJCNLP (1)5
2021 Minimal Supervision for Morphological Inflection
abstract
Neural models for the various flavours of morphological reinflection tasks have proven to be extremely accurate given ample labeled data, yet labeled data may be slow and costly to obtain.In this work we aim to overcome this annotation bottleneck by bootstrapping labeled data from a seed as small as five labeled inflection tables, accompanied by a large bulk of unlabeled text.Our bootstrapping method exploits the orthographic and semantic regularities in morphological systems in a two-phased setup, where word tagging based on analogies is followed by word pairing based on distances.Our experiments with the Paradigm Cell Filling Problem over eight typologically different languages show that in languages with relatively simple morphology, orthographic regularities on their own allow inflection models to achieve respectable accuracy.Combined orthographic and semantic regularities alleviate difficulties with particularly complex morpho-phonological systems.We further show that our bootstrapping methods substantially outperform hallucination-based methods commonly used for overcoming the annotation bottleneck in morphological reinflection tasks.
Omer Goldman, Reut Tsarfaty
EMNLP (1)2
2021 Asking It All: Generating Contextualized Questions for any Semantic Role
abstract
Asking questions about a situation is an inherent step towards understanding it.To this end, we introduce the task of role question generation, which, given a predicate mention and a passage, requires producing a set of questions asking about all possible semantic roles of the predicate.We develop a two-stage model for this task, which first produces a contextindependent question prototype for each role and then revises it to be contextually appropriate for the passage.Unlike most existing approaches to question generation, our approach does not require conditioning on existing answers in the text.Instead, we condition on the type of information to inquire about, regardless of whether the answer appears explicitly in the text, could be inferred from it, or should be sought elsewhere.Our evaluation demonstrates that we generate diverse and well-formed questions for a large, broadcoverage ontology of predicates and roles.
Valentina Pyatkin, Paul Roit, Julian Michael, Yoav Goldberg, Reut Tsarfaty, Ido Dagan
EMNLP (1)5
2021 Neural Modeling for Named Entities and Morphology (NEMO2)
abstract
Abstract Named Entity Recognition (NER) is a fundamental NLP task, commonly formulated as classification over a sequence of tokens. Morphologically rich languages (MRLs) pose a challenge to this basic formulation, as the boundaries of named entities do not necessarily coincide with token boundaries, rather, they respect morphological boundaries. To address NER in MRLs we then need to answer two fundamental questions, namely, what are the basic units to be labeled, and how can these units be detected and classified in realistic settings (i.e., where no gold morphology is available). We empirically investigate these questions on a novel NER benchmark, with parallel token- level and morpheme-level NER annotations, which we develop for Modern Hebrew, a morphologically rich-and-ambiguous language. Our results show that explicitly modeling morphological boundaries leads to improved NER performance, and that a novel hybrid architecture, in which NER precedes and prunes morphological decomposition, greatly outperforms the standard pipeline, where morphological decomposition strictly precedes NER, setting a new performance bar for both Hebrew NER and Hebrew morphological decomposition tasks.
Dan Bareket, Reut Tsarfaty
Trans. Assoc. Comput. Linguistics2
2020 From SPMRL to NMRL: What Did We Learn (and Unlearn) in a Decade of Parsing Morphologically-Rich Languages (MRLs)?
abstract
It has been exactly a decade since the first establishment of SPMRL, a research initiative unifying multiple research efforts to address the peculiar challenges of Statistical Parsing for Morphologically-Rich Languages (MRLs).Here we reflect on parsing MRLs in that decade, highlight the solutions and lessons learned for the architectural, modeling and lexical challenges in the pre-neural era, and argue that similar challenges re-emerge in neural architectures for MRLs.We then aim to offer a climax, suggesting that incorporating symbolic ideas proposed in SPMRL terms into nowadays neural architectures has the potential to push NLP for MRLs to a new level.We sketch a strategies for designing Neural Models for MRLs (NMRL), and showcase preliminary support for these strategies via investigating the task of multi-tagging in Hebrew, a morphologically-rich, high-fusion, language.
Reut Tsarfaty, Dan Bareket, Stav Klein, Amit Seker
ACL1
2020 QADiscourse - Discourse Relations as QA Pairs: Representation, Crowdsourcing and Baselines
abstract
Discourse relations describe how two propositions relate to one another, and identifying them automatically is an integral part of natural language understanding.However, annotating discourse relations typically requires expert annotators.Recently, different semantic aspects of a sentence have been represented and crowd-sourced via question-and-answer (QA) pairs.This paper proposes a novel representation of discourse relations as QA pairs, which in turn allows us to crowd-source widecoverage data annotated with discourse relations, via an intuitively appealing interface for composing such questions and answers.Based on our proposed representation, we collect a novel and wide-coverage QADiscourse dataset, and present baseline algorithms for predicting QADiscourse relations.
Valentina Pyatkin, Ayal Klein, Reut Tsarfaty, Ido Dagan
EMNLP (1)3
2019 RUN through the Streets: A New Dataset and Baseline Models for Realistic Urban Navigation
abstract
Tzuf Paz-Argaman, Reut Tsarfaty. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Tzuf Paz-Argaman, Reut Tsarfaty
EMNLP/IJCNLP (1)2
2019 Joint Transition-Based Models for Morpho-Syntactic Parsing: Parsing Strategies for MRLs and a Case Study from Modern Hebrew
abstract
Abstract In standard NLP pipelines, morphological analysis and disambiguation (MA&D) precedes syntactic and semantic downstream tasks. However, for languages with complex and ambiguous word-internal structure, known as morphologically rich languages (MRLs), it has been hypothesized that syntactic context may be crucial for accurate MA&D, and vice versa. In this work we empirically confirm this hypothesis for Modern Hebrew, an MRL with complex morphology and severe word-level ambiguity, in a novel transition-based framework. Specifically, we propose a joint morphosyntactic transition-based framework which formally unifies two distinct transition systems, morphological and syntactic, into a single transition-based system with joint training and joint inference. We empirically show that MA&D results obtained in the joint settings outperform MA&D results obtained by the respective standalone components, and that end-to-end parsing results obtained by our joint system present a new state of the art for Hebrew dependency parsing.
Amir More, Amit Seker, Victoria Basmova, Reut Tsarfaty
Trans. Assoc. Comput. Linguistics4
2018 Representations and Architectures in Neural Sentiment Analysis for Morphologically Rich Languages: A Case Study from Modern Hebrew
abstract
This paper empirically studies the effects of representation choices on neural sentiment analysis for Modern Hebrew, a morphologically rich language (MRL) for which no sentiment analyzer currently exists. We study two dimensions of representational choices: (i) the granularity of the input signal (token-based vs. morpheme-based), and (ii) the level of encoding of vocabulary items (string-based vs. character-based). We hypothesise that for MRLs, languages where multiple meaning-bearing elements may be carried by a single space-delimited token, these choices will have measurable effects on task perfromance, and that these effects may vary for different architectural designs — fully-connected, convolutional or recurrent. Specifically, we hypothesize that morpheme-based representations will have advantages in terms of their generalization capacity and task accuracy, due to their better OOV coverage. To empirically study these effects, we develop a new sentiment analysis benchmark for Hebrew, based on 12K social media comments, and provide two instances of these data: in token-based and morpheme-based settings. Our experiments show that representation choices empirical effects vary with architecture type. While fully-connected and convolutional networks slightly prefer token-based settings, RNNs benefit from a morpheme-based representation, in accord with the hypothesis that explicit morphological information may help generalize. Our endeavour also delivers the first state-of-the-art broad-coverage sentiment analyzer for Hebrew, with over 89% accuracy, alongside an established benchmark to further study the effects of linguistic representation choices on neural networks’ task performance.
Adam Amram, Anat Ben-David, Reut Tsarfaty
COLING3
2018 CoNLL-UL: Universal Morphological Lattices for Universal Dependency Parsing
Amir More, Özlem Çetinoglu, Çagri Çöltekin, Nizar Habash, Benoît Sagot, Djamé Seddah, Dima Taji, Reut Tsarfaty
LREC8
2017 Data-Driven Broad-Coverage Grammars for Opinionated Natural Language Generation (ONLG)
abstract
Opinionated natural language generation (ONLG) is a new, challenging, NLG task in which we aim to automatically generate human-like, subjective, responses to opinionated articles online.We present a data-driven architecture for ONLG that generates subjective responses triggered by users' agendas, based on automatically acquired wide-coverage generative grammars.We compare three types of grammatical representations that we design for ONLG.The grammars interleave different layers of linguistic information, and are induced from a new, enriched dataset we developed.Our evaluation shows that generation with Relational-Realizational (Tsarfaty and Sima'an, 2008) inspired grammar gets better language model scores than lexicalized grammars à la Collins (2003), and that the latter gets better humanevaluation scores.We also show that conditioning the generation on topic models makes generated responses more relevant to the document content.
Tomer Cagan, Stefan L. Frank, Reut Tsarfaty
ACL (1)3
2016 Data-Driven Morphological Analysis and Disambiguation for Morphologically Rich Languages and Universal Dependencies
abstract
Parsing texts into universal dependencies (UD) in realistic scenarios requires infrastructure for the morphological analysis and disambiguation (MA&D) of typologically different languages as a first tier. MA&D is particularly challenging in morphologically rich languages (MRLs), where the ambiguous space-delimited tokens ought to be disambiguated with respect to their constituent morphemes, each morpheme carrying its own tag and a rich set features. Here we present a novel, language-agnostic, framework for MA&D, based on a transition system with two variants — word-based and morpheme-based — and a dedicated transition to mitigate the biases of variable-length morpheme sequences. Our experiments on a Modern Hebrew case study show state of the art results, and we show that the morpheme-based MD consistently outperforms our word-based variant. We further illustrate the utility and multilingual coverage of our framework by morphologically analyzing and disambiguating the large set of languages in the UD treebanks.
Amir More, Reut Tsarfaty
COLING2
2016 Universal Dependencies v1: A Multilingual Treebank Collection
Joakim Nivre, Marie-Catherine de Marneffe, Filip Ginter, Yoav Goldberg, Jan Hajic 0001, Christopher D. Manning, Ryan T. McDonald, Slav Petrov, Sampo Pyysalo, Natalia Silveira, Reut Tsarfaty, Daniel Zeman
LREC11
2014 Semantic Parsing Using Content and Context: A Case Study from Requirements Elicitation
abstract
We present a model for the automatic semantic analysis of requirements elicitation documents.Our target semantic representation employs live sequence charts, a multi-modal visual language for scenariobased programming, which can be directly translated into executable code.The architecture we propose integrates sentencelevel and discourse-level processing in a generative probabilistic framework for the analysis and disambiguation of individual sentences in context.We show empirically that the discourse-based model consistently outperforms the sentence-based model when constructing a system that reflects all the static (entities, properties) and dynamic (behavioral scenarios) requirements in the document.
Reut Tsarfaty, Ilia Pogrebezky, Guy Weiss, Yaarit Natan, Smadar Szekely, David Harel
EMNLP1
2013 Design Patterns in Fluid Construction Grammar Luc Steels (editor) Universitat Pompeu Fabra and Sony Computer Science Laboratory, Paris Amsterdam: John Benjamins Publishing Company (Constructional Approaches to Language series, edited by Mirjam Fried and Jan-Ola Östman, volume 11), 2012, xi+332 pp; hardbound, ISBN 978-90-272-0433-2, €99.00, $149.00
abstract
In computational modeling of natural language phenomena, there are at least three modes of research. The currently dominant statistical paradigm typically prioritizes instance coverage: Data-driven methods seek to use as much information observed in data as possible in order to generalize linguistic analyses to unseen instances. A second approach prioritizes detailed description of grammatical phenomena, that is, forming and defending theories with a focus on a small number of instances. A third approach might be called integrative: Rather than addressing phenomena in isolation, different approaches are brought together to address multiple challenges in a unified framework, and the behavior of the system is demonstrated with a small number of instances. Design Patterns in Fluid Construction Grammar (DPFCG) exemplifies the third approach, introducing a linguistic formalism called Fluid Construction Grammar (FCG) that addresses parsing, production, and learning in a single computational framework.The book emphasizes grammar-engineering, following broad-coverage descriptive paradigms that can be traced back to Generalized Phrase Structure Grammar (GPSG) Gazdar et al. (1985), Lexical Functional Grammar (LFG) Bresnan (2000), Head-Driven Phrase-Structure Grammar (HPSG) Sag and Wasow (1999), and Combinatory Categorial Grammar (CCG) Steedman (1996). In all of these cases, a formal meta-framework allows computational linguists to formalize their hypotheses and intuitions about a language's grammatical behavior and then explore how these representational choices affect the processing of natural language utterances. Many of the aforementioned approaches have engendered large-scale platforms that can be used and reused to provide formal description of grammars for different languages, such as Par-Gram for LFG Butt et al. (2002) and the LinGO Grammar Matrix for HPSG Bender, Flickinger, and Oepen (2002).FCG offers a similar grammar engineering framework that follows the principles of Construction Grammar (CxG) Goldberg (2003; Hoffmann and Trousdale (2013). CxG treats constructions as the basic units of grammatical organization in language. The constructions are viewed as learned associations between form (e.g., sounds, morphemes, syntactic phrases) and function (semantics, pragmatics, discourse meaning, etc.). CxG does not impose a strict separation between lexicon and grammar—indeed, it is perhaps best known as treating semi-productive idioms like “the X-er, the Y-er” and “X let alone Y” on equal footing with lexemes and “core” syntactic patterns Fillmore, Kay, and O'Connor (1988; Kay and Fillmore (1999). FCG, like other CxG formalisms—namely, Embodied Construction Grammar Bergen and Chang (2005; Feldman, Dodge, and Bryant (2009) and Sign-Based Construction Grammar Boas and Sag (2012)—is unification-based.1 The studies in this book describe constructions and how they can be combined in order to model natural language interpretation or generation as feature structure unification in a general search procedure.The book has five parts, covering the groundwork, basic linguistic applications, processing matters, advanced case studies, and, finally, features of FCG that make it fluid and robust. Each chapter identifies general strategies (design patterns) that might merit reuse in new FCG grammars, or perhaps in other computational frameworks.Part I: Introduction lays the groundwork for the rest of the book. “Introducing Fluid Construction Grammar” (by Luc Steels) presents the aims of the FCG formalism. FCG was designed as a framework for describing linguistic units (constructions—their form and meaning), with an emphasis on language variation and evolution (“fluidity”). The constructionist approach to language is described and the argument for applying it to study language variation and change is defended. Psychological validity is explicitly ruled out as a modeling goal (“The emphasis is on getting working systems, and this is difficult enough”; page 4). The architects of FCG set out to include both sides of the processing coin, however—parsing (interpretation) and production (generation). The concept of search in processing is emphasized, though some of the explanations of processing steps are too abstract for the reader to comprehend at this point. A further desideratum—robustness to noisy input containing disfluencies, fragments, and errors—is given as motivating a constructionist approach.The next chapter, “A First Encounter with Fluid Construction Grammar” (by Steels), describes the mechanisms of FCG in detail. In FCG, a working analysis hypothesized in processing is known as transient structure; the transduction of form to meaning (and vice versa) selects a sequence of constructions that apply to the transient structure to gradually expand it until reaching a final analysis. Identifying constructions that may apply to a transient structure presents a non-trivial search problem, also addressed by the architects of FCG. The sheer number of technical details make this chapter somewhat overwhelming. Most of the chapter is devoted to the low-level feature structures and the operations manipulating them. Templates—a practical means of avoiding boilerplate code when defining constructions—are then introduced, and do most of the heavy lifting in the rest of the book.In Part II: Grammatical Structures, we begin to see how constructions are defined in practice. “A Design Pattern for Phrasal Constructions” (by Steels) illustrates how constructions are used to describe the combination of multiple units into higher-level, typed phrases. Skeletal constructions compose with smaller units to form hierarchical structures (essentially similar to the Immediate Constituents Analysis of (Bloomfield 1933) and follow-up work in structuralist linguistics [Harris 1946]), and a range of additional constructions impose form (e.g., ordering) constraints and add new meaning to the newly created phrases. This chapter is of a tutorial nature, illustrating the step-by-step application of four kinds of noun phrase constructions to expand transient structures in processing. Over twenty templates are introduced in this chapter; they encapsulate design patterns dealing with hierarchical structure, agreement, and feature percolation. An aspect of phrasal constructions that is not yet dealt with is the complex linking of semantic arguments and the morphosyntactic categorizations of the composed elements.“A Design Pattern for Argument Structure Constructions” (by Remi van Trijp) then builds on the formal machinery presented in the previous chapter to explicitly address the complex mappings between semantic arguments (agent, patient, etc.) and syntactic arguments (subject, object, etc.). This mapping is a complex matter due to language-specific conceptualization of semantic arguments and different means of morphosyntactic realization used by different languages. In FCG, each lexical item introduces its linking potential in terms of the different types of semantic and syntactic arguments that it may take, with no particular mapping between them. Each argument structure construction imposes a partial mapping between the syntactic and semantic arguments to yield a particular argument structure instantiation, one of the multiple alternatives that may be available for a single lexical item. This account stands in sharp contrast to the lexicalist view of argument-structure (the view taken in LFG, HPSG, and CCG) whereby each lexical entry dictates all the necessary linking information. The construction-based approach is defended for its ability to deal with unknown words2 and constructional coercion3 Goldberg (1995). The argument structure design pattern allows FCG to crudely recover a partial specification of the form-meaning mapping of these elements, which is important for robust processing (see subsequent discussion).Part III: Managing Processing addresses how FCG transduces between a linguistic string and a meaning representation, where the two directions (parsing and production) share a common declarative representation of linguistic knowledge (the grammar). This entails assembling an analysis incrementally on the basis of the grammar, the input, and any partial analysis that has already been created. With FCG (and unification grammars more broadly) this search is nontrivial, and streamlining search (i.e., minimizing nondeterminism and avoiding dead ends) is a key motivator of many of the grammar design patterns suggested in the book.“Search in Linguistic Processing” (by Joris Bleys, Kevin Stadler, and Joachim De Beule) deals mainly with the problem of choosing which of multiple compatible constructions to apply next. Whereas the default heuristic search in FCG is a greedy, depth-first search (which can backtrack if the user-defined end-goal has not yet been achieved) the FCG framework allows for a guided search through scores that reflect the relative tendency of a construction to apply next. The authors suggest that such scoring can be informed by general principles, for instance: (i) specific constructions are preferred to more general ones, and (ii) previously co-applied constructions are preferred. Choosing appropriate constructions to apply early on dramatically reduces the time needed for processing the utterance.“Organizing Constructions in Networks” (by Pieter Wellens) takes this idea to the next level, and proposes to organize the different constructions in networks of conditional dependencies. A conditional dependency links two constructions where one provides necessary information for the application of the other. These dependency networks can be updated whenever an input is processed so that the system learns to search more efficiently when the same constructions are encountered in the future. Using these networks to guide the search thus significantly reduces the search for compatible constructions. An empirical effort to quantify this effect indeed shows a sharp reduction in search time; unlike the held-out experimental paradigm accepted in statistical NLP, however, the parsed/produced sentence is assumed to have been seen already by the system.Part IV: Case Studies addresses three challenging linguistic phenomena in FCG. “Feature Matrices and Agreement” (by van Trijp) on German case offers a new unification-based solution to the problem of feature indeterminacy. For instance, in the sentence Er findet und hilft Frauen ‘He finds and helps women’, the first verb requires an accusative object, whereas the second requires a dative object; the coordination is allowed only because Frauen can be either accusative or dative. Kindern ‘children’, which can only be dative, is not licensed here. Encoding case in a single feature on the Frauen construction wouldn't work because the feature would have to unify with contradictory values (from the verbs' inflectional features). Instead, case restrictions specified lexically for a noun or verb can be expressed with a distinctive feature matrix, with each matrix slot holding a variable or the value + or -. Unification then does the right thing—allowing Frauen and forbidding Kindern—without resorting to type hierarchies or disjunctive features.“Construction Sets and Unmarked Forms” (by Katrien Beuls) on Hungarian verbal agreement models a phenomenon whereby morphosyntactic, semantic, and phonological factors affect the choice between poly- and mono-personal agreement—that is, the decision whether a Hungarian transitive verb should agree with its object or just with its subject. The case and definiteness of the object and the person hierarchy relationship between subject and object determine which kind of agreement obtains, and phonological constraints determine its form. To make the different levels of structure interact properly, constructions are grouped into sets (lexical, morphological, etc.) and those sets are considered in a fixed order during processing. Construction sets also allow for efficient handling of unmarked forms (null affixes)—they are considered only after the overt affixes have had the opportunity to apply, thereby functioning as defaults.“Syntactic Indeterminacy and Semantic Ambiguity” (by Michael Spranger and Martin Loetzsch) on German spatial phrases models the German spatial terms for front, back, left, and right. To model spatial language in situated interaction with robots, two problems must be overcome. The first is syntactic indeterminacy: Any of these spatial relations may be realized as an adjective, an adverb, or a preposition. The second is semantic ambiguity, specifically when the perspective (e.g., whose ‘left’?) is implicit. Both are forms of underspecification which could cause early splits in the search space if handled naïvely. Much in the spirit of the argument structure constructions (see above), the solutions (which are too technical to explain here) involve (a) disjunctive representations of potential values of a feature, and (b) deferring decisions until a more opportune stage.Part V: Robustness and Fluidity (by Steels and van Trijp) surveys the different features of the system that ensure robustness in the face of variation, disfluencies, and noise. Natural language is fluid and open-ended. There is variation between speakers, there are disfluencies and speech errors, and noise may corrupt the speech signal. All of these may jeopardize the interpretability of the signal, but human listeners are adept at processing such input. In the spirit of usage-based grammar Tomasello (2003), FCG emphasizes the attainment of a communicative goal, rather than ensuring grammaticality of parsed/produced utterances. This is accomplished with a diagnostic-repair process that runs in parallel to parsing/production. Diagnostics can test for unknown words, unfamiliar meanings, missing constructions, and so on. Diagnostic tests are implemented by reversing the direction of the transduction process: a speaker may assume the hearer's point of view to analyze what she has produced in order to see whether communicative success has been attained. Likewise, a hearer may produce a phrase according to his own interpretation of the speaker's form, and check for a match. If a test fails, repair strategies such as proposing new constructions, relaxing the matching process for construction application, and coercing constructions to adapt to novel language use are considered.Fluidity and robustness are the hallmarks of FCG, and the computational framework has been used in experiments that assume embedded communication in robotic agents. This research program is developed at length by Steels (2012b).Discussion. Like the legacy of the GPSG book Gazdar et al. (1985), this book's main merit is not necessarily in its technical details or computational choices, but in demonstrating the feasibility of implementing the constructional approach in a full-fledged computational framework. We suggest that the CxG perspective presents a formidable challenge to the computational linguistics/natural language processing community. It posits a different notion of modularity than is observed by most NLP systems: Rather than treat different levels of linguistic structure independently, CxG recognizes that multiple formal components (phonological, lexical, morphological, syntactic) may be tied by convention to a specific meaning or function. Systematically describing these “cross-cutting” constructions and their processing, especially in a way that scales to large data encompassing both form and meaning and accommodates both parsing and generation, would in our view make for a more comprehensive account of language processing than our field is able to offer today. Thus, we hope this book will be provocative even outside of the grammar engineering community.This book is not without its weaknesses. In parts the writing is quite technical and terse, which can be daunting for readers new to FCG. Contextualization with respect to other strands of computational linguistics and AI research is, for the most part, lacking, though a second FCG book Steels (2012a) picks up some of the slack on this front.4DPFCG does not address the feasibility of learning constructions directly from data,5 nor does it discuss the expressive power of the formalism in relation to learnability results (such as that of Gold [1967]). As admitted by the authors, much more work would be needed to build life-size grammars. Still, we hope that readers of DPFCG will appreciate the authors' vision for a model of linguistic form and function that is at once formal, computational, fluid, and robust.
Nathan Schneider 0001, Reut Tsarfaty
Comput. Linguistics2
2013 Parsing Morphologically Rich Languages: Introduction to the Special Issue
abstract
Parsing is a key task in natural language processing. It involves predicting, for each natural language sentence, an abstract representation of the grammatical entities in the sentence and the relations between these entities. This representation provides an interface to compositional semantics and to the notions of “who did what to whom.” The last two decades have seen great advances in parsing English, leading to major leaps also in the performance of applications that use parsers as part of their backbone, such as systems for information extraction, sentiment analysis, text summarization, and machine translation. Attempts to replicate the success of parsing English for other languages have often yielded unsatisfactory results. In particular, parsing languages with complex word structure and flexible word order has been shown to require non-trivial adaptation. This special issue reports on methods that successfully address the challenges involved in parsing a range of morphologically rich languages (MRLs). This introduction characterizes MRLs, describes the challenges in parsing MRLs, and outlines the contributions of the articles in the special issue. These contributions present up-to-date research efforts that address parsing in varied, cross-lingual settings. They show that parsing MRLs addresses challenges that transcend particular representational and algorithmic choices.
Reut Tsarfaty, Djamé Seddah, Sandra Kübler, Joakim Nivre
Comput. Linguistics1
2012 Cross-Framework Evaluation for Statistical Parsing
Reut Tsarfaty, Joakim Nivre, Evelina Andersson
EACL1
2011 Evaluating Dependency Parsing: Robust and Heuristics-Free Cross-Annotation Evaluation
Reut Tsarfaty, Joakim Nivre, Evelina Andersson
EMNLP1
2009 Enhancing Unlexicalized Parsing Performance Using a Wide Coverage Lexicon, Fuzzy Tag-Set Mapping, and EM-HMM-Based Lexical Probabilities
Yoav Goldberg, Reut Tsarfaty, Meni Adler, Michael Elhadad
EACL2
2009 An Alternative to Head-Driven Approaches for Parsing a (Relatively) Free Word-Order Language
Reut Tsarfaty, Khalil Sima'an, Remko Scha
EMNLP1
2008 A Single Generative Model for Joint Morphological Segmentation and Syntactic Parsing
Yoav Goldberg, Reut Tsarfaty
ACL2
2008 Relational-Realizational Parsing
Reut Tsarfaty, Khalil Sima'an
COLING1
2008 Word-Based or Morpheme-Based? Annotation Strategies for Modern Hebrew Clitics
Reut Tsarfaty, Yoav Goldberg
LREC1
2006 Integrated Morphological and Syntactic Disambiguation for Modern Hebrew
Reut Tsarfaty
ACL1