Amir Zeldes

dblp:37/9972 · DBLP profile ↗
← Back
17ranked-venue papers
2as first author
14since 2021 · last 2026
0000-0001-8016-6753ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 2 first-author · 14 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Expect the Unexpected? Testing the Surprisal of Salient Entities
abstract
Previous work examining the Uniform Information Density (UID) hypothesis has shown that while information as measured by surprisal metrics is distributed more or less evenly across documents overall, local discrepancies can arise due to functional pressures corresponding to syntactic and discourse structural constraints.However, work thus far has largely disregarded the relative salience of discourse participants.We fill this gap by studying how overall salience of entities in discourse relates to surprisal using 70K manually annotated mentions across 16 genres of English and a novel minimal-pair prompting method.Our results show that globally salient entities exhibit significantly higher surprisal than non-salient ones, even controlling for position, length, and nesting confounds.Moreover, salient entities systematically reduce surprisal for surrounding content when used as prompts, enhancing document-level predictability.This effect varies by genre, appearing strongest in topiccoherent texts and weakest in conversational contexts.Our findings refine the UID competing pressures framework by identifying global entity salience as a mechanism shaping information distribution in discourse.
Jessica Lin 0004, Amir Zeldes
ACL (1)2
2026 GUMBridge: A Corpus for Varieties of Bridging Anaphora
Lauren Levine, Amir Zeldes
LREC2
2025 eRST: A Signaled Graph Theory of Discourse Relations and Organization
abstract
Abstract In this article we present Enhanced Rhetorical Structure Theory (eRST), a new theoretical framework for computational discourse analysis, based on an expansion of Rhetorical Structure Theory (RST). The framework encompasses discourse relation graphs with tree-breaking, non-projective and concurrent relations, as well as implicit and explicit signals which give explainable rationales to our analyses. We survey shortcomings of RST and other existing frameworks, such as Segmented Discourse Representation Theory, the Penn Discourse Treebank, and Discourse Dependencies, and address these using constructs in the proposed theory. We provide annotation, search, and visualization tools for data, and present and evaluate a freely available corpus of English annotated according to our framework, encompassing 12 spoken and written genres with over 200K tokens. Finally, we discuss automatic parsing, evaluation metrics, and applications for data in our framework.
Amir Zeldes, Tatsuya Aoyama, Yang Janet Liu, Siyao Peng, Debopam Das, Luke Gessler
Comput. Linguistics1
2024 DISRPT: A Multilingual, Multi-domain, Cross-framework Benchmark for Discourse Processing
abstract
This paper presents DISRPT, a multilingual, multi-domain, and cross-framework benchmark dataset for discourse processing, covering the tasks of discourse unit segmentation, connective identification, and relation classification. DISRPT includes 13 languages, with data from 24 corpora covering about 4 millions tokens and around 250,000 discourse relation instances from 4 discourse frameworks: RST, SDRT, PDTB, and Discourse Dependencies. We present an overview of the data, its development across three NLP shared tasks on discourse processing carried out in the past five years, and the latest modifications and added extensions. We also carry out an evaluation of state-of-the-art multilingual systems trained on the data for each task, showing plateau performance on segmentation, but important room for improvement for connective identification and relation classification. The DISRPT benchmark employs a unified format that we make available on GitHub and HuggingFace in order to encourage future work on discourse processing across languages, domains, and frameworks.
Chloé Braud, Amir Zeldes, Laura Rivière, Yang Janet Liu, Philippe Muller, Damien Sileo, Tatsuya Aoyama
LREC/COLING2
2024 Universal Anaphora: The First Three Years
abstract
The aim of the Universal Anaphora initiative is to push forward the state of the art in anaphora and anaphora resolution by expanding the aspects of anaphoric interpretation which are or can be reliably annotated in anaphoric corpora, producing unified standards to annotate and encode these annotations, delivering datasets encoded according to these standards, and developing methods for evaluating models that carry out this type of interpretation. Although several papers on aspects of the initiative have appeared, no overall description of the initiative’s goals, proposals and achievements has been published yet except as an online draft. This paper aims to fill this gap, as well as to discuss its progress so far.
Massimo Poesio, Maciej Ogrodniczuk, Vincent Ng 0001, Sameer Pradhan, Juntao Yu, Nafise Sadat Moosavi, Silviu Paun, Amir Zeldes, Anna Nedoluzhko, Michal Novák 0001, Martin Popel, Zdenek Zabokrtský, Daniel Zeman
LREC/COLING8
2024 UCxn: Typologically-Informed Annotation of Constructions Atop Universal Dependencies
abstract
The Universal Dependencies (UD) project has created an invaluable collection of treebanks with contributions in over 140 languages. However, the UD annotations do not tell the full story. Grammatical constructions that convey meaning through a particular combination of several morphosyntactic elements—for example, interrogative sentences with special markers and/or word orders—are not labeled holistically. We argue for (i) augmenting UD annotations with a ‘UCxn’ annotation layer for such meaning-bearing grammatical constructions, and (ii) approaching this in a typologically informed way so that morphosyntactic strategies can be compared across languages. As a case study, we consider five construction families in ten languages, identifying instances of each construction in UD treebanks through the use of morphosyntactic patterns. In addition to findings regarding these particular constructions, our study yields important insights on methodology for describing and identifying constructions in language-general and language-particular ways, and lays the foundation for future constructional enrichment of UD treebanks.
Leonie Weissweiler, Nina Böbel, Kirian Guiller, Santiago Herrera, Wesley Scivetti, Arthur Lorenzi Almeida, Nurit Melnik, Archna Bhatia, Hinrich Schütze, Lori S. Levin, Amir Zeldes, Joakim Nivre, William Croft 0001, Nathan Schneider 0001
LREC/COLING11
2024 SPLICE: A Singleton-Enhanced PipeLIne for Coreference REsolution
abstract
Singleton mentions, i.e. entities mentioned only once in a text, are important to how humans understand discourse from a theoretical perspective. However previous attempts to incorporate their detection in end-to-end neural coreference resolution for English have been hampered by the lack of singleton mention spans in the OntoNotes benchmark. This paper addresses this limitation by combining predicted mentions from existing nested NER systems and features derived from OntoNotes syntax trees. With this approach, we create a near approximation of the OntoNotes dataset with all singleton mentions, achieving ~94% recall on a sample of gold singletons. We then propose a two-step neural mention and coreference resolution system, named SPLICE, and compare its performance to the end-to-end approach in two scenarios: the OntoNotes test set and the out-of-domain (OOD) OntoGUM corpus. Results indicate that reconstructed singleton training yields results comparable to end-to-end systems for OntoNotes, while improving OOD stability (+1.1 avg. F1). We conduct error analysis for mention detection and delve into its impact on coreference clustering, revealing that precision improvements deliver more substantial benefits than increases in recall for resolving coreference chains.
Yilun Zhu 0001, Siyao Peng, Sameer Pradhan, Amir Zeldes
LREC/COLING4
2024 GUMsley: Evaluating Entity Salience in Summarization for 12 English Genres
abstract
As NLP models become increasingly capable of understanding documents in terms of coherent entities rather than strings, obtaining the most salient entities for each document is not only an important end task in itself but also vital for Information Retrieval (IR) and other downstream applications such as controllable summarization.In this paper, we present and evaluate GUMsley, the first entity salience dataset covering all named and non-named salient entities for 12 genres of English text, aligned with entity types, Wikification links and full coreference resolution annotations.We promote a strict definition of salience using human summaries and demonstrate high inter-annotator agreement for salience based on whether a source entity is mentioned in the summary.Our evaluation shows poor performance by pre-trained SOTA summarization models and zero-shot LLM prompting in capturing salient entities in generated summaries.We also show that predicting or providing salient entities to several model architectures enhances performance and helps derive higher-quality summaries by alleviating the entity hallucination problem in existing abstractive summarization.
Jessica Lin 0004, Amir Zeldes
EACL (1)2
2024 GDTB: Genre Diverse Data for English Shallow Discourse Parsing across Modalities, Text Types, and Domains
abstract
Yang Janet Liu, Tatsuya Aoyama, Wesley Scivetti, Yilun Zhu, Shabnam Behzad, Lauren Elizabeth Levine, Jessica Lin, Devika Tiwari, Amir Zeldes. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Yang Janet Liu, Tatsuya Aoyama, Wesley Scivetti, Yilun Zhu 0001, Shabnam Behzad, Lauren Levine, Jessica Lin 0004, Devika Tiwari, Amir Zeldes
EMNLP9
2023 ELQA: A Corpus of Metalinguistic Questions and Answers about English
abstract
We present ELQA, a corpus of questions and answers in and about the English language.Collected from two online forums, the >70k questions (from English learners and others) cover wide-ranging topics including grammar, meaning, fluency, and etymology.The answers include descriptions of general properties of English vocabulary and grammar as well as explanations about specific (correct and incorrect) usage examples.Unlike most NLP datasets, this corpus is metalinguistic-it consists of language about language.As such, it can facilitate investigations of the metalinguistic capabilities of NLU models, as well as educational applications in the language learning domain.To study this, we define a free-form question answering task on our dataset and conduct evaluations on multiple LLMs (Large Language Models) to analyze their capacity to generate metalinguistic answers.
Shabnam Behzad, Keisuke Sakaguchi, Nathan Schneider 0001, Amir Zeldes
ACL (1)4
2023 Why Can't Discourse Parsing Generalize? A Thorough Investigation of the Impact of Data Diversity
abstract
Recent advances in discourse parsing performance create the impression that, as in other NLP tasks, performance for high-resource languages such as English is finally becoming reliable.In this paper we demonstrate that this is not the case, and thoroughly investigate the impact of data diversity on RST parsing stability.We show that state-of-the-art architectures trained on the standard English newswire benchmark do not generalize well, even within the news domain.Using the two largest RST corpora of English with text from multiple genres, we quantify the impact of genre diversity in training data for achieving generalization to text types unseen during training.Our results show that a heterogeneous training regime is critical for stable and generalizable models, across parser architectures.We also provide error analyses of model outputs and out-ofdomain performance.To our knowledge, this study is the first to fully evaluate cross-corpus RST parsing generalizability on complete trees, examine between-genre degradation within an RST corpus, and investigate the impact of genre diversity in training data composition.
Yang Janet Liu, Amir Zeldes
EACL2
2023 What's Hard in English RST Parsing? Predictive Models for Error Analysis
abstract
Despite recent advances in Natural Language Processing (NLP), hierarchical discourse parsing in the framework of Rhetorical Structure Theory remains challenging, and our understanding of the reasons for this are as yet limited.In this paper, we examine and model some of the factors associated with parsing difficulties in previous work: the existence of implicit discourse relations, challenges in identifying long-distance relations, out-of-vocabulary items, and more.In order to assess the relative importance of these variables, we also release two annotated English test-sets with explicit correct and distracting discourse markers associated with gold standard RST relations.Our results show that as in shallow discourse parsing, the explicit/implicit distinction plays a role, but that long-distance dependencies are the main challenge, while lack of lexical overlap is less of a problem, at least for in-domain parsing.Our final model is able to predict where errors will occur with an accuracy of 76.3% for the bottom-up parser and 76.6% for the top-down parser.
Yang Janet Liu, Tatsuya Aoyama, Amir Zeldes
SIGDIAL3
2022 A Second Wave of UD Hebrew Treebanking and Cross-Domain Parsing
abstract
Foundational Hebrew NLP tasks such as segmentation, tagging and parsing, have relied to date on various versions of the Hebrew Treebank (HTB, Sima'an et al. 2001).However, the data in HTB, a single-source newswire corpus, is now over 30 years old, and does not cover many aspects of contemporary Hebrew on the web.This paper presents a new, freely available UD treebank of Hebrew stratified from a range of topics selected from Hebrew Wikipedia.In addition to introducing the corpus and evaluating the quality of its annotations, we deploy automatic validation tools based on grew (Guillaume, 2021), and conduct the first crossdomain parsing experiments in Hebrew.We obtain new state-of-the-art (SOTA) results on UD NLP tasks, using a combination of the latest language modelling and some incremental improvements to existing transformer based approaches.We also release a new version of the UD HTB matching annotation scheme updates from our new corpus.
Amir Zeldes, Nicholas Howell, Noam Ordan, Yifat Ben Moshe
EMNLP1
2022 CorefUD 1.0: Coreference Meets Universal Dependencies
abstract
Recent advances in standardization for annotated language resources have led to successful large scale efforts, such as the Universal Dependencies (UD) project for multilingual syntactically annotated data. By comparison, the important task of coreference resolution, which clusters multiple mentions of entities in a text, has yet to be standardized in terms of data formats or annotation guidelines. In this paper we present CorefUD, a multilingual collection of corpora and a standardized format for coreference resolution, compatible with morphosyntactic annotations in the UD framework and including facilities for related tasks such as named entity recognition, which forms a first step in the direction of convergence for coreference resolution across languages.
Anna Nedoluzhko, Michal Novák 0001, Martin Popel, Zdenek Zabokrtský, Amir Zeldes, Daniel Zeman
LREC5
2020 AMALGUM - A Free, Balanced, Multilayer English Web Corpus
abstract
We present a freely available, genre-balanced English web corpus totaling 4M tokens and featuring a large number of high-quality automatic annotation layers, including dependency trees, non-named entity annotations, coreference resolution, and discourse trees in Rhetorical Structure Theory. By tapping open online data sources the corpus is meant to offer a more sizable alternative to smaller manually created annotated data sets, while avoiding pitfalls such as imbalanced or unknown composition, licensing problems, and low-quality natural language processing. We harness knowledge from multiple annotation layers in order to achieve a “better than NLP” benchmark and evaluate the accuracy of the resulting resource.
Luke Gessler, Siyao Peng, Yang Liu 0213, Yilun Zhu 0001, Shabnam Behzad, Amir Zeldes
LREC6
2020 Treebanking User-Generated Content: A Proposal for a Unified Representation in Universal Dependencies
abstract
The paper presents a discussion on the main linguistic phenomena of user-generated texts found in web and social media, and proposes a set of annotation guidelines for their treatment within the Universal Dependencies (UD) framework. Given on the one hand the increasing number of treebanks featuring user-generated content, and its somewhat inconsistent treatment in these resources on the other, the aim of this paper is twofold: (1) to provide a short, though comprehensive, overview of such treebanks - based on available literature - along with their main features and a comparative analysis of their annotation criteria, and (2) to propose a set of tentative UD-based annotation guidelines, to promote consistent treatment of the particular phenomena found in these types of texts. The main goal of this paper is to provide a common framework for those teams interested in developing similar resources in UD, thus enabling cross-linguistic consistency, which is a principle that has always been in the spirit of UD.
Manuela Sanguinetti, Cristina Bosco, Lauren Cassidy, Özlem Çetinoglu, Alessandra Teresa Cignarella, Teresa Lynn, Ines Rehbein, Josef Ruppenhofer, Djamé Seddah, Amir Zeldes
LREC10
2019 The Making of Coptic Wordnet
abstract
With the increasing availability of wordnets for ancient languages, such as Ancient Greek and Latin, gaps remain in the coverage of less studied languages of antiquity.This paper reports on the construction and evaluation of a new wordnet for Coptic, the language of Late Roman, Byzantine and Early Islamic Egypt in the first millenium CE.We present our approach to constructing the wordnet which uses multilingual Coptic dictionaries and wordnets for five different languages.We further discuss the results of this effort and outline our on-going/future work.
Laura A. Slaughter, Luís Morgado da Costa, So Miyagawa, Marco Büchler, Amir Zeldes, Heike Behlmer
GWC5