EDBT 2026 Demo / reviewers in the wild / expert
Massimo Poesio
dblp:60/6334
· DBLP profile ↗
106ranked-venue papers
23as first author
22since 2021 · last 2026
0000-0001-8469-2072ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 91 · 23 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 5 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 9 · 4 since 2021Databases, data management, data science and information retrieval · 7Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Grounded Misunderstandings in Asymmetric Dialogue: A Perspectivist Annotation Scheme for MapTaskabstractCollaborative dialogue relies on participants incrementally establishing common ground, yet in asymmetric settings they may believe they agree while referring to different entities. We introduce a perspectivist annotation scheme for the HCRC MapTask corpus (Anderson et al., 1991) that separately captures speaker and addressee grounded interpretations for each reference expression, enabling us to trace how understanding emerges, diverges, and repairs over time. Using a scheme-constrained LLM annotation pipeline, we obtain 13k annotated reference expressions with reliability estimates and analyze the resulting understanding states. The results show that full misunderstandings are rare once lexical variants are unified, but multiplicity discrepancies systematically induce divergences, revealing how apparent grounding can mask referential misalignment. Our framework provides both a resource and an analytic lens for studying grounded misunderstanding and for evaluating (V)LLMs' capacity to model perspective-dependent grounding in collaborative dialogue. Nan Li 0087, Albert Gatt, Massimo Poesio |
LREC | 3 |
| 2026 | AmbiCoRefVis: A Tool for Visualizing Coreferential Ambiguity
Patrick Paetzold, Lukas Beiske, Mark-Matthias Zymla, Massimo Poesio, Miriam Butt, Daniel Weiskopf, Oliver Deussen |
LREC | 4 |
| 2026 | Human Label Variation in Implicit Discourse Relation Recognition
Frances Yung, Daniil Ignatev, Merel C. J. Scholman, Vera Demberg, Massimo Poesio |
LREC | 5 |
| 2026 | Seeing Is Not Sharing: Some Vision-Language Models Overestimate Common Ground in Asymmetric DialogueabstractIn collaborative dialogue, shared perception does not guarantee shared interpretation. Mutual understanding must be established through interaction. We investigate whether vision-language models (VLMs) can distinguish what could be shared from what has been shared between dialogue participants through grounding. We formulate this as an interpretation-matching task on 13,077 annotated reference expressions from HCRC MapTask dialogues, and evaluate VLMs under systematically controlled manipulations of dialogue context and map-information access. Our results show that providing authentic map images improves overall performance but shifts models toward over-predicting alignment. Textual descriptions of the same map content reproduce this bias, while non-informative images suppress alignment predictions entirely, indicating that the bias is driven by task-relevant map content, not the visual channel. This improvement comes at the cost of degraded accuracy on non-aligned cases. Calibration analysis and reference-chain tracking further suggest that models rely on static referential cues on the maps rather than tracking how grounding unfolds through dialogue history. We observe these patterns most clearly in Qwen3-VL-8B-Instruct and, to varying degrees, in four additional models from two architecture families. In models that exhibit the bias, map content, whether presented visually or textually, is treated as evidence of mutual understanding, conflating potential with established common ground. Nan Li 0087, Albert Gatt, Massimo Poesio |
SIGDIAL | 3 |
| 2025 | How Task Complexity Moderates the Impact of AI-Generated Images on User Experience in Gamified Text LabellingabstractThis study investigates how task complexity influences the impact of AI-generated images on user experience in gamified text-labelling tasks. Participants completed one of two tasks of varying complexity: part-ofspeech (POS) tagging simple and natural language inference (NLI) complex under conditions with and without AIgenerated images. We measured accuracy, engagement, enjoyment, and cognitive load using factorial ANOVAs. Results revealed that task complexity significantly influenced all outcomes: the NLI task yielded lower accuracy, reduced engagement and enjoyment, and increased cognitive load compared with the POS task. AI-generated images had limited impact, negatively affecting perceived usability and overall engagement, with engagement notably higher in the absence of images for the POS task. These findings suggest that while task complexity strongly shapes user experience in gamified labelling tasks, AI-generated images may not enhance performance or engagement. Fatima Althani, Chris Madge, Massimo Poesio |
CoG | 3 |
| 2025 | Talking-to-Build: How LLM-Assisted Interface Shapes Player Performance and Experience in MinecraftabstractWith large language models (LLMs) on the rise, in-game interactions are shifting from rigid commands to natural conversations. However, the impacts of LLMs on player performance and game experience remain underexplored. This work explores LLM's role as a co-builder during gameplay, examining its impact on task performance, usability, and player experience. Using Minecraft as a sandbox, we present an LLM-assisted interface that engages players through natural language, aiming to facilitate creativity and simplify complex gaming commands. We conducted a mixed-methods study with 30 participants, comparing LLM-assisted and command-based interfaces across simple and complex game tasks. Quantitative and qualitative analyses reveal that the LLM-assisted interface significantly improves player performance, engagement, and overall game experience. Additionally, task complexity has a notable effect on player performance and experience across both interfaces. Our findings highlight the potential of LLM-assisted interfaces to revolutionize virtual experiences, emphasizing the importance of balancing intuitiveness with predictability, transparency, and user agency in AI-driven, multimodal gaming environments. Xin Sun 0016, Yue Li 0044, Jie Li 0064, Massimo Poesio, Julian Frommel, Koen V. Hindriks, Jiahuan Pei |
ICMI | 5 |
| 2025 | Improving LLMs' Learning of Coreference ResolutionabstractCoreference Resolution (CR) is crucial for many NLP tasks, but existing LLMs struggle with hallucination and under-performance. In this paper, we investigate the limitations of existing LLM-based approaches to CR—specifically the Question-Answering (QA) Template and Document Template methods—and propose two novel techniques: Reversed Training with Joint Inference and Iterative Document Generation. Our experiments show that Reversed Training improves the QA Template method, while Iterative Document Generation eliminates hallucinations in the generated source text and boosts coreference resolution. Integrating these methods and techniques offers an effective and robust solution to LLM-based coreference resolution Yujian Gan, Yanni Lin, Juntao Yu, Massimo Poesio |
SIGDIAL | 5 |
| 2025 | A Survey of Coreference and Zeros Resolution for ArabicabstractCoreference resolution is the task of resolving mentions that refer to the same entity into clusters. The area and its tasks are crucial in natural language processing applications. Extensive surveys of this task have been conducted for English and Chinese, but not too much for Arabic. The few Arabic surveys do not cover recent progress and the challenges for Arabic anaphora, and they do not cover zero resolution and comprehensive resolution of zeros and full mentions, or anaphora resolution beyond coreference (e.g., bridging). In this article, we examine the state of the art for Arabic anaphora resolution, highlighting the challenges and advances in this field. We provide a comprehensive survey of the methods employed for Arabic coreference resolution, as well as an overview of the existing datasets and challenges.The goal is to equip researchers with a thorough understanding of Arabic anaphora resolution and to suggest potential future directions in the field. Abdulrahman Aloraini, Juntao Yu, Wateen A. Aliady, Massimo Poesio |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2024 | Master the Linguistic Landscape: Puzzle Integration in a 3D NLP GameabstractMany research studies have investigated the effects of game design elements in gamified applications, but to our knowledge, none have specifically focused on the puzzle game element. This paper addresses this gap by exploring how puzzle game elements impact NLP gamified applications. We explored the effect of such an approach on players’ enjoyment, competence, pressure, effort, choice, and performance (annotation accuracy). Our results showed that players derive greater enjoyment when playing the game with puzzles than those who only played the base game. Wateen A. Aliady, Massimo Poesio |
CoG | 2 |
| 2024 | Assessing the Capabilities of Large Language Models in Coreference: An EvaluationabstractThis paper offers a nuanced examination of the role Large Language Models (LLMs) play in coreference resolution, aimed at guiding the future direction in the era of LLMs. We carried out both manual and automatic analyses of different LLMs’ abilities, employing different prompts to examine the performance of different LLMs, obtaining a comprehensive view of their strengths and weaknesses. We found that LLMs show exceptional ability in understanding coreference. However, harnessing this ability to achieve state of the art results on traditional datasets and benchmarks isn’t straightforward. Given these findings, we propose that future efforts should: (1) Improve the scope, data, and evaluation methods of traditional coreference research to adapt to the development of LLMs. (2) Enhance the fine-grained language understanding capabilities of LLMs. Yujian Gan, Massimo Poesio, Juntao Yu |
LREC/COLING | 2 |
| 2024 | Conceptual Pacts for Reference Resolution Using Small, Dynamically Constructed Language Models: A Study in Puzzle Building DialoguesabstractUsing Brennan and Clark’s theory of a Conceptual Pact, that when interlocutors agree on a name for an object, they are forming a temporary agreement on how to conceptualize that object, we present an extension to a simple reference resolver which simulates this process over time with different conversation pairs. In a puzzle construction domain, we model pacts with small language models for each referent which update during the interaction. When features from these pact models are incorporated into a simple bag-of-words reference resolver, the accuracy increases compared to using a standard pre-trained model. The model performs equally to a competitor using the same data but with exhaustive re-training after each prediction, while also being more transparent, faster and less resource-intensive. We also experiment with reducing the number of training interactions, and can still achieve reference resolution accuracies of over 80% in testing from observing a single previous interaction, over 20% higher than a pre-trained baseline. While this is a limited domain, we argue the model could be applicable to larger real-world applications in human and human-robot interaction and is an interpretable and transparent model. Julian Hough, Sina Zarrieß, Casey Kennington, David Schlangen, Massimo Poesio |
LREC/COLING | 5 |
| 2024 | Universal Anaphora: The First Three YearsabstractThe aim of the Universal Anaphora initiative is to push forward the state of the art in anaphora and anaphora resolution by expanding the aspects of anaphoric interpretation which are or can be reliably annotated in anaphoric corpora, producing unified standards to annotate and encode these annotations, delivering datasets encoded according to these standards, and developing methods for evaluating models that carry out this type of interpretation. Although several papers on aspects of the initiative have appeared, no overall description of the initiative’s goals, proposals and achievements has been published yet except as an online draft. This paper aims to fill this gap, as well as to discuss its progress so far. Massimo Poesio, Maciej Ogrodniczuk, Vincent Ng 0001, Sameer Pradhan, Juntao Yu, Nafise Sadat Moosavi, Silviu Paun, Amir Zeldes, Anna Nedoluzhko, Michal Novák 0001, Martin Popel, Zdenek Zabokrtský, Daniel Zeman |
LREC/COLING | 1 |
| 2024 | Analyzing and Enhancing Clarification Strategies for Ambiguous References in Consumer Service InteractionsabstractChangling Li, Yujian Gan, Zhenrong Yang, Youyang Chen, Xinxuan Qiu, Yanni Lin, Matthew Purver, Massimo Poesio. Proceedings of the 25th Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2024. Changling Li, Yujian Gan, Zhenrong Yang, Youyang Chen, Xinxuan Qiu, Yanni Lin, Matthew Purver, Massimo Poesio |
SIGDIAL | 8 |
| 2024 | Polysemy - Evidence from Linguistics, Behavioral Science, and Contextualized Language ModelsabstractAbstract Polysemy is the type of lexical ambiguity where a word has multiple distinct but related interpretations. In the past decade, it has been the subject of a great many studies across multiple disciplines including linguistics, psychology, neuroscience, and computational linguistics, which have made it increasingly clear that the complexity of polysemy precludes simple, universal answers, especially concerning the representation and processing of polysemous words. But fuelled by the growing availability of large, crowdsourced datasets providing substantial empirical evidence; improved behavioral methodology; and the development of contextualized language models capable of encoding the fine-grained meaning of a word within a given context, the literature on polysemy recently has developed more complex theoretical analyses. In this survey we discuss these recent contributions to the investigation of polysemy against the backdrop of a long legacy of research across multiple decades and disciplines. Our aim is to bring together different perspectives to achieve a more complete picture of the heterogeneity and complexity of the phenomenon of polysemy. Specifically, we highlight evidence supporting a range of hybrid models of the mental processing of polysemes. These hybrid models combine elements from different previous theoretical approaches to explain patterns and idiosyncrasies in the processing of polysemous that the best known models so far have failed to account for. Our literature review finds that (i) traditional analyses of polysemy can be limited in their generalizability by loose definitions and selective materials; (ii) linguistic tests provide useful evidence on individual cases, but fail to capture the full range of factors involved in the processing of polysemous sense extensions; and (iii) recent behavioral (psycho) linguistics studies, large-scale annotation efforts, and investigations leveraging contextualized language models provide accumulating evidence suggesting that polysemous sense similarity covers a wide spectrum between identity of sense and homonymy-like unrelatedness of meaning. We hope that the interdisciplinary account of polysemy provided in this survey inspires further fundamental research on the nature of polysemy and better equips applied research to deal with the complexity surrounding the phenomenon, for example, by enabling the development of benchmarks and testing paradigms for large language models informed by a greater portion of the rich evidence on the phenomenon currently available. Janosch Haber, Massimo Poesio |
Comput. Linguistics | 2 |
| 2023 | The Onboarding Phase in a Game for Text Labelling: Comparing the Effect of Animated vs. Textual Onboarding on Player Experience and AccuracyabstractEngaging new players while collecting quality data in Games-with-a-Purpose (GWAPs) can be challenging. It is therefore crucial to understand how to design the first minutes of play to attract and engage players without compromising data quality. In our study (n = 91), we explored the impact of presenting an onboarding phase in a game for language annotation on both data accuracy and player experience. We focused on evaluating two different modalities of the onboarding phase, animated (A) and text (B), and a control (C) version where an onboarding phase was not presented. Those who were presented with an onboarding phase were briefly introduced to the game’s story and then a tutorial explaining the gameplay mechanics. During the gameplay, players’ inputs were recorded to measure the accuracy of their labels. After playing the game, players completed the Game Experience Questionnaire to evaluate their player experience. The results indicate that while the accuracy of the data collected is not significantly different amongst the three groups, players experienced significantly higher frustration in the text version of the onboarding phase than in the control. This paper suggests that GWAP designers should examine whether investing in an onboarding phase can be justified for their games, or whether the game mechanics can be simply explored by players. Fatima Althani, Chris Madge, Massimo Poesio |
CoG | 3 |
| 2023 | Aggregating Crowdsourced and Automatic Judgments to Scale Up a Corpus of Anaphoric Reference for Fiction and Wikipedia TextsabstractJuntao Yu, Silviu Paun, Maris Camilleri, Paloma Garcia, Jon Chamberlain, Udo Kruschwitz, Massimo Poesio. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Juntao Yu, Silviu Paun, Maris Camilleri, Paloma Carretero Garcia, Jon Chamberlain, Udo Kruschwitz, Massimo Poesio |
EACL | 7 |
| 2022 | ArMIS - The Arabic Misogyny and Sexism Corpus with Annotator Subjective DisagreementsabstractThe use of misogynistic and sexist language has increased in recent years in social media, and is increasing in the Arabic world in reaction to reforms attempting to remove restrictions on women lives. However, there are few benchmarks for Arabic misogyny and sexism detection, and in those the annotations are in aggregated form even though misogyny and sexism judgments are found to be highly subjective. In this paper we introduce an Arabic misogyny and sexism dataset (ArMIS) characterized by providing annotations from annotators with different degree of religious beliefs, and provide evidence that such differences do result in disagreements. To the best of our knowledge, this is the first dataset to study in detail the effect of beliefs on misogyny and sexism annotation. We also discuss proof-of-concept experiments showing that a dataset in which disagreements have not been reconciled can be used to train state-of-the-art models for misogyny and sexism detection; and consider different ways in which such models could be evaluated. Dina Almanea, Massimo Poesio |
LREC | 2 |
| 2022 | The Universal Anaphora ScorerabstractThe aim of the Universal Anaphora initiative is to push forward the state of the art in anaphora and anaphora resolution by expanding the aspects of anaphoric interpretation which are or can be reliably annotated in anaphoric corpora, producing unified standards to annotate and encode these annotations, deliver datasets encoded according to these standards, and developing methods for evaluating models carrying out this type of interpretation. Such expansion of the scope of anaphora resolution requires a comparable expansion of the scope of the scorers used to evaluate this work. In this paper, we introduce an extended version of the Reference Coreference Scorer (Pradhan et al., 2014) that can be used to evaluate the extended range of anaphoric interpretation included in the current Universal Anaphora proposal. The UA scorer supports the evaluation of identity anaphora resolution and of bridging reference resolution, for which scorers already existed but not integrated in a single package. It also supports the evaluation of split antecedent anaphora and discourse deixis, for which no tools existed. The proposed approach to the evaluation of split antecedent anaphora is entirely novel; the proposed approach to the evaluation of discourse deixis leverages the encoding of discourse deixis proposed in Universal Anaphora to enable the use for discourse deixis of the same metrics already used for identity anaphora. The scorer was tested in the recent CODI-CRAC 2021 Shared Task on Anaphora Resolution in Dialogues. Juntao Yu, Sopan Khosla, Nafise Sadat Moosavi, Silviu Paun, Sameer Pradhan, Massimo Poesio |
LREC | 6 |
| 2021 | BERTective: Language Models and Contextual Information for Deception DetectionabstractSpotting a lie is challenging but has an enormous potential impact on security as well as private and public safety.Several NLP methods have been proposed to classify texts as truthful or deceptive.In most cases, however, the target texts' preceding context is not considered.This is a severe limitation, as any communication takes place in context, not in a vacuum, and context can help to detect deception.We study a corpus of Italian dialogues containing deceptive statements and implement deep neural models that incorporate various linguistic contexts.We establish a new state-of-theart identifying deception and find that not all context is equally useful to the task.Only the texts closest to the target, if from the same speaker (rather than questions by an interlocutor), boost performance.We also find that the semantic information in language models such as BERT contributes to the performance.However, BERT alone does not capture the implicit knowledge of deception cues: its contribution is conditional on the concurrent use of attention to learn cues from BERT's representations. Tommaso Fornaciari, Federico Bianchi 0001, Massimo Poesio, Dirk Hovy |
EACL | 3 |
| 2021 | Beyond Black & White: Leveraging Annotator Disagreement via Soft-Label Multi-Task LearningabstractTommaso Fornaciari, Alexandra Uma, Silviu Paun, Barbara Plank, Dirk Hovy, Massimo Poesio. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Tommaso Fornaciari, Alexandra Uma, Silviu Paun, Barbara Plank, Dirk Hovy, Massimo Poesio |
NAACL-HLT | 6 |
| 2021 | Stay Together: A System for Single and Split-antecedent Anaphora ResolutionabstractJuntao Yu, Nafise Sadat Moosavi, Silviu Paun, Massimo Poesio. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Juntao Yu, Nafise Sadat Moosavi, Silviu Paun, Massimo Poesio |
NAACL-HLT | 4 |
| 2021 | Learning from Disagreement: A SurveyabstractMany tasks in Natural Language Processing (NLP) and Computer Vision (CV) offer evidence that humans disagree, from objective tasks such as part-of-speech tagging to more subjective tasks such as classifying an image or deciding whether a proposition follows from certain premises. While most learning in artificial intelligence (AI) still relies on the assumption that a single (gold) interpretation exists for each item, a growing body of research aims to develop learning methods that do not rely on this assumption. In this survey, we review the evidence for disagreements on NLP and CV tasks, focusing on tasks for which substantial datasets containing this information have been created. We discuss the most popular approaches to training models from datasets containing multiple judgments potentially in disagreement. We systematically compare these different approaches by training them with each of the available datasets, considering several ways to evaluate the resulting models. Finally, we discuss the results in depth, focusing on four key research questions, and assess how the type of evaluation and the characteristics of a dataset determine the answers to these questions. Our results suggest, first of all, that even if we abandon the assumption of a gold standard, it is still essential to reach a consensus on how to evaluate models. This is because the relative performance of the various training methods is critically affected by the chosen form of evaluation. Secondly, we observed a strong dataset effect. With substantial datasets, providing many judgments by high-quality coders for each item, training directly with soft labels achieved better results than training from aggregated or even gold labels. This result holds for both hard and soft evaluation. But when the above conditions do not hold, leveraging both gold and soft labels generally achieved the best results in the hard evaluation. All datasets and models employed in this paper are freely available as supplementary materials. Alexandra Uma, Tommaso Fornaciari, Dirk Hovy, Silviu Paun, Barbara Plank, Massimo Poesio |
J. Artif. Intell. Res. | 6 |
| 2020 | Named Entity Recognition as Dependency ParsingabstractNamed Entity Recognition (NER) is a fundamental task in Natural Language Processing, concerned with identifying spans of text expressing references to entities.NER research is often focused on flat entities only (flat NER), ignoring the fact that entity references can be nested, as in [Bank of [China]] (Finkel and Manning, 2009).In this paper, we use ideas from graph-based dependency parsing to provide our model a global view on the input via a biaffine model (Dozat and Manning, 2017).The biaffine model scores pairs of start and end tokens in a sentence which we use to explore all spans, so that the model is able to predict named entities accurately.We show that the model works well for both nested and flat NER through evaluation on 8 corpora and achieving SoTA performance on all of them, with accuracy gains of up to 2.2 percentage points. Juntao Yu, Bernd Bohnet, Massimo Poesio |
ACL | 3 |
| 2020 | Free the Plural: Unrestricted Split-Antecedent Anaphora ResolutionabstractNow that the performance of coreference resolvers on the simpler forms of anaphoric reference has greatly improved, more attention is devoted to more complex aspects of anaphora.One limitation of virtually all coreference resolution models is the focus on single-antecedent anaphors.Plural anaphors with multiple antecedents-so-called split-antecedent anaphors (as in John met Mary.They went to the movies)-have not been widely studied, because they are not annotated in ONTONOTES and are relatively infrequent in other corpora.In this paper, we introduce the first model for unrestricted resolution of split-antecedent anaphors.We start with a strong baseline enhanced by BERT embeddings, and show that we can substantially improve its performance by addressing the sparsity issue.To do this, we experiment with auxiliary corpora where split-antecedent anaphors were annotated by the crowd, and with transfer learning models using element-of bridging references and single-antecedent coreference as auxiliary tasks.Evaluation on the gold annotated ARRAU corpus shows that the out best model uses a combination of three auxiliary corpora achieved F1 scores of 70% and 43.6% when evaluated in a lenient and strict setting, respectively, i.e., 11 and 21 percentage points gain when compared with our baseline. Juntao Yu, Nafise Sadat Moosavi, Silviu Paun, Massimo Poesio |
COLING | 4 |
| 2020 | Multitask Learning-Based Neural Bridging Reference ResolutionabstractWe propose a multi task learning-based neural model for resolving bridging references tackling two key challenges.The first challenge is the lack of large corpora annotated with bridging references.To address this, we use multi-task learning to help bridging reference resolution with coreference resolution.We show that substantial improvements of up to 8 p.p. can be achieved on full bridging resolution with this architecture.The second challenge is the different definitions of bridging used in different corpora, meaning that hand-coded systems or systems using special features designed for one corpus do not work well with other corpora.Our neural model only uses a small number of corpus independent features, thus can be applied to different corpora.Evaluations with very different bridging corpora (ARRAU, ISNOTES, BASHI and SCICORP) suggest that our architecture works equally well on all corpora, and achieves the SoTA results on full bridging resolution for all corpora, outperforming the best reported results by up to 36.3 p.p.. Juntao Yu, Massimo Poesio |
COLING | 2 |
| 2020 | A Case for Soft Loss FunctionsabstractRecently, Peterson et al. provided evidence of the benefits of using probabilistic soft labels generated from crowd annotations for training a computer vision model, showing that using such labels maximizes performance of the models over unseen data. In this paper, we generalize these results by showing that training with soft labels is an effective method for using crowd annotations in several other ai tasks besides the one studied by Peterson et al., and also when their performance is compared with that of state-of-the-art methods for learning from crowdsourced data. Alexandra Uma, Tommaso Fornaciari, Dirk Hovy, Silviu Paun, Barbara Plank, Massimo Poesio |
HCOMP | 6 |
| 2020 | Cross-lingual Zero Pronoun ResolutionabstractIn languages like Arabic, Chinese, Italian, Japanese, Korean, Portuguese, Spanish, and many others, predicate arguments in certain syntactic positions are not realized instead of being realized as overt pronouns, and are thus called zero- or null-pronouns. Identifying and resolving such omitted arguments is crucial to machine translation, information extraction and other NLP tasks, but depends heavily on semantic coherence and lexical relationships. We propose a BERT-based cross-lingual model for zero pronoun resolution, and evaluate it on the Arabic and Chinese portions of OntoNotes 5.0. As far as we know, ours is the first neural model of zero-pronoun resolution for Arabic; and our model also outperforms the state-of-the-art for Chinese. In the paper we also evaluate BERT feature extraction and fine-tune models on the task, and compare them with our model. We also report on an investigation of BERT layers indicating which layer encodes the most suitable representation for the task. Abdulrahman Aloraini, Massimo Poesio |
LREC | 2 |
| 2020 | Neural Mention DetectionabstractMention detection is an important preprocessing step for annotation and interpretation in applications such as NER and coreference resolution, but few stand-alone neural models have been proposed able to handle the full range of mentions. In this work, we propose and compare three neural network-based approaches to mention detection. The first approach is based on the mention detection part of a state of the art coreference resolution system; the second uses ELMO embeddings together with a bidirectional LSTM and a biaffine classifier; the third approach uses the recently introduced BERT model. Our best model (using a biaffine classifier) achieves gains of up to 1.8 percentage points on mention recall when compared with a strong baseline in a HIGH RECALL coreference annotation setting. The same model achieves improvements of up to 5.3 and 6.2 p.p. when compared with the best-reported mention detection F1 on the CONLL and CRAC coreference data sets respectively in a HIGH F1 annotation setting. We then evaluate our models for coreference resolution by using mentions predicted by our best model in start-of-the-art coreference systems. The enhanced model achieved absolute improvements of up to 1.7 and 0.7 p.p. when compared with our strong baseline systems (pipeline system and end-to-end system) respectively. For nested NER, the evaluation of our model on the GENIA corpora shows that our model matches or outperforms state-of-the-art models despite not being specifically designed for this task. Juntao Yu, Bernd Bohnet, Massimo Poesio |
LREC | 3 |
| 2020 | A Cluster Ranking Model for Full Anaphora ResolutionabstractAnaphora resolution (coreference) systems designed for the CONLL 2012 dataset typically cannot handle key aspects of the full anaphora resolution task such as the identification of singletons and of certain types of non-referring expressions (e.g., expletives), as these aspects are not annotated in that corpus. However, the recently released dataset for the CRAC 2018 Shared Task can now be used for that purpose. In this paper, we introduce an architecture to simultaneously identify non-referring expressions (including expletives, predicative s, and other types) and build coreference chains, including singletons. Our cluster-ranking system uses an attention mechanism to determine the relative importance of the mentions in the same cluster. Additional classifiers are used to identify singletons and non-referring markables. Our contributions are as follows. First all, we report the first result on the CRAC data using system mentions; our result is 5.8% better than the shared task baseline system, which used gold mentions. Second, we demonstrate that the availability of singleton clusters and non-referring expressions can lead to substantially improved performance on non-singleton clusters as well. Third, we show that despite our model not being designed specifically for the CONLL data, it achieves a score equivalent to that of the state-of-the-art system by Kantor and Globerson (2019) on that dataset. Juntao Yu, Alexandra Uma, Massimo Poesio |
LREC | 3 |
| 2020 | An NLP-Powered Human Rights Monitoring PlatformabstractEffective information management has long been a problem in organisations that are not of a scale that they can afford their own department dedicated to this task. Growing information overload has made this problem even more pronounced. On the other hand we have recently witnessed the emergence of intelligent tools, packages and resources that made it possible to rapidly transfer knowledge from the academic community to industry, government and other potential beneficiaries. Here we demonstrate how adopting state-of-the-art natural language processing (NLP) and crowdsourcing methods has resulted in measurable benefits for a human rights organisation by transforming their information and knowledge management using a novel approach that supports human rights monitoring in conflict zones. More specifically, we report on mining and classifying Arabic Twitter in order to identify potential human rights abuse incidents in a continuous stream of social media data within a specified geographical region . Results show deep learning approaches such as LSTM allow us to push the precision close to 85% for this task with an F1-score of 75%. Apart from the scientific insights we also demonstrate the viability of the framework which has been deployed as the Ceasefire Iraq portal for more than three years which has already collected thousands of witness reports from within Iraq. This work is a case study of how progress in artificial intelligence has disrupted even the operation of relatively small-scale organisations. Ayman Alhelbawy, Mark Lattimer, Udo Kruschwitz, Chris Fox, Massimo Poesio |
Expert Syst. Appl. | 5 |
| 2020 | Annotating a broad range of anaphoric phenomena, in a variety of genres: the ARRAU CorpusabstractAbstract This paper presents the second release ofarrau, a multigenre corpus of anaphoric information created over 10 years to provide data for the next generation of coreference/anaphora resolution systems combining different types of linguistic and world knowledge with advanced discourse modeling supporting rich linguistic annotations. The distinguishing features ofarrauinclude the following: treating all NPs as markables, including non-referring NPs, and annotating their (non-) referentiality status; distinguishing between several categories of non-referentiality and annotating non-anaphoric mentions; thorough annotation of markable boundaries (minimal/maximal spans, discontinuous markables); annotating a variety of mention attributes, ranging from morphosyntactic parameters to semantic category; annotating the genericity status of mentions; annotating a wide range of anaphoric relations, including bridging relations and discourse deixis; and, finally, annotating anaphoric ambiguity. The current version of the dataset contains 350K tokens and is publicly available from LDC. In this paper, we discuss in detail all the distinguishing features of the corpus, so far only partially presented in a number of conference and workshop papers, and we also discuss the development between the first release ofarrauin 2008 and this second one. Olga Uryupina, Ron Artstein, Antonella Bristot, Federica Cavicchio, Francesca Delogu, Kepa Joseba Rodríguez, Massimo Poesio |
Nat. Lang. Eng. | 7 |
| 2019 | Crowdsourcing and Aggregating Nested Markable AnnotationsabstractOne of the key steps in language resource creation is the identification of the text segments to be annotated, or markablesin our case, the (potentially nested) noun phrases in coreference resolution (or mentions).In this paper, we present a method for identifying markables for coreference annotation that combines high-performance automatic markable detectors with checking with a Game-With-A-Purpose (GWAP) and aggregation using a Bayesian annotation model.The method was evaluated both on news data and data from a variety of other genres and results in an improvement on F 1 of mention boundaries of over seven percentage points when compared with a state-of-the-art, domain-independent automatic mention detector, and almost three points over an in-domain mention detector.One of the key contributions of our proposal is its applicability to the case in which markables are nested, as is the case with coreference markables; but the GWAP and several of the proposed markable detectors are task-and language-independent and are thus applicable to a variety of other annotation scenarios. Chris Madge, Juntao Yu, Jon Chamberlain, Udo Kruschwitz, Silviu Paun, Massimo Poesio |
ACL (1) | 6 |
| 2019 | Using Automatically Extracted Minimum Spans to Disentangle Coreference Evaluation from Boundary DetectionabstractThis is a repository copy of Using automatically extracted minimum spans to disentangle coreference evaluation from boundary detection. Nafise Sadat Moosavi, Leo Born, Massimo Poesio, Michael Strube 0001 |
ACL (1) | 3 |
| 2019 | The Design Of A Clicker Game for Text LabellingabstractGames for text annotation / labelling are becoming more common, but it's difficult to find a mechanics that fits. In this work we discuss a clicker game that can support text annotation. We believe this type of game is uniquely suited to addressing some of the challenges faced by games featuring text annotation as a core task. Chris Madge, Richard A. Bartle, Jon Chamberlain, Udo Kruschwitz, Massimo Poesio |
CoG | 5 |
| 2019 | Wormingo: a 'true gamification' approach to anaphoric annotationabstractIn this paper we present Wormingo, 1 a new Game-with-a-Purpose for anaphoric annotation. It introduces the motivation-annotation paradigm which uses linguistic puzzles and other widely known gamification techniques and word game mechanics to motivate players to carry out anaphoric annotation tasks. In a preliminary experiment, the game was tested on 270 players recruited through the Reddit platform, achieving promising results. Doruk Kicikoglu, Richard A. Bartle, Jon Chamberlain, Massimo Poesio |
FDG | 4 |
| 2019 | Making text annotation fun with a clicker gameabstractIn this paper we present WordClicker, a clicker game for text annotation. We believe the mechanics of 'Ville type Free-To-Play (F2P) games in general, and clicker games in particular, is particularly suited for GWAPs (Games-With-A-Purpose). WordClicker was developed as one component of a suite of GWAPs meant to cover all aspects of language interpretation, from tokenization to anaphoric interpretation. As such, WordClicker is intended to have a dual function as part of this suite of GWAPs: both for parts-of-speech annotation and for teaching players about parts of speech so that they can go on and play GWAPs for more complex syntactic annotation. Therefore, game-based language learning platforms also had a strong influence on its design. Chris Madge, Richard A. Bartle, Jon Chamberlain, Udo Kruschwitz, Massimo Poesio |
FDG | 5 |
| 2019 | Progression in a Language Annotation Game with a PurposeabstractWithin traditional games design, incorporating progressive difficulty is considered of fundamental importance. But despite the widespread intuition that progression could have clear benefits in Games-With-A-Purpose (GWAPs)–e.g., for training non-expert annotators to produce more complex judgements– progression is not in fact a prominent feature of GWAPs; and there is even less evidence on its effects. In this work we present an approach to progression in GWAPs that generalizes to different annotation tasks with minimal, if any, dependency on gold annotated data. Using this method we observe a statistically significant increase in accuracy over randomly showing items to annotators. Chris Madge, Juntao Yu, Jon Chamberlain, Udo Kruschwitz, Silviu Paun, Massimo Poesio |
HCOMP | 6 |
| 2018 | A Probabilistic Annotation Model for Crowdsourcing CoreferenceabstractThe availability of large scale annotated corpora for coreference is essential to the development of the field.However, creating resources at the required scale via expert annotation would be too expensive.Crowdsourcing has been proposed as an alternative; but this approach has not been widely used for coreference.This paper addresses one crucial hurdle on the way to make this possible, by introducing a new model of annotation for aggregating crowdsourced anaphoric annotations.The model is evaluated along three dimensions: the accuracy of the inferred mention pairs, the quality of the post-hoc constructed silver chains, and the viability of using the silver chains as an alternative to the expert-annotated chains in training a state of the art coreference system.The results suggest that our model can extract from crowdsourced annotations coreference chains of comparable quality to those obtained with expert annotation. Silviu Paun, Jon Chamberlain, Udo Kruschwitz, Juntao Yu, Massimo Poesio |
EMNLP | 5 |
| 2018 | Comparing Bayesian Models of AnnotationabstractThe analysis of crowdsourced annotations in natural language processing is concerned with identifying (1) gold standard labels, (2) annotator accuracies and biases, and (3) item difficulties and error patterns. Traditionally, majority voting was used for 1, and coefficients of agreement for 2 and 3. Lately, model-based analysis of corpus annotations have proven better at all three tasks. But there has been relatively little work comparing them on the same datasets. This paper aims to fill this gap by analyzing six models of annotation, covering different approaches to annotator ability, item difficulty, and parameter pooling (tying) across annotators and items. We evaluate these models along four aspects: comparison to gold labels, predictive accuracy for new annotations, annotator characterization, and item difficulty, using four datasets with varying degrees of noise in the form of random (spammy) annotators. We conclude with guidelines for model selection, application, and implementation. Silviu Paun, Bob Carpenter, Jon Chamberlain, Dirk Hovy, Udo Kruschwitz, Massimo Poesio |
Trans. Assoc. Comput. Linguistics | 6 |
| 2017 | Visually Grounded and Textual Semantic Models Differentially Decode Brain Activity Associated with Concrete and Abstract NounsabstractImportant advances have recently been made using computational semantic models to decode brain activity patterns associated with concepts; however, this work has almost exclusively focused on concrete nouns. How well these models extend to decoding abstract nouns is largely unknown. We address this question by applying state-of-the-art computational models to decode functional Magnetic Resonance Imaging (fMRI) activity patterns, elicited by participants reading and imagining a diverse set of both concrete and abstract nouns. One of the models we use is linguistic, exploiting the recent word2vec skipgram approach trained on Wikipedia. The second is visually grounded, using deep convolutional neural networks trained on Google Images. Dual coding theory considers concrete concepts to be encoded in the brain both linguistically and visually, and abstract concepts only linguistically. Splitting the fMRI data according to human concreteness ratings, we indeed observe that both models significantly decode the most concrete nouns; however, accuracy is significantly greater using the text-based models for the most abstract nouns. More generally this confirms that current computational models are sufficiently advanced to assist in investigating the representational structure of abstract concepts in the brain. Andrew J. Anderson, Douwe Kiela, Stephen Clark, Massimo Poesio |
Trans. Assoc. Comput. Linguistics | 4 |
| 2016 | Towards a Corpus of Violence Acts in Arabic Social Media
Ayman Alhelbawy, Massimo Poesio, Udo Kruschwitz |
LREC | 2 |
| 2016 | Phrase Detectives Corpus 1.0 Crowdsourced Anaphoric Coreference
Jon Chamberlain, Massimo Poesio, Udo Kruschwitz |
LREC | 2 |
| 2016 | The OnForumS corpus from the Shared Task on Online Forum Summarisation at MultiLing 2015
Mijail A. Kabadjov, Udo Kruschwitz, Massimo Poesio, Josef Steinberger, Marc Poch, Hugo Zaragoza |
LREC | 3 |
| 2016 | ARRAU: Linguistically-Motivated Annotation of Anaphoric Descriptions
Olga Uryupina, Ron Artstein, Antonella Bristot, Federica Cavicchio, Kepa Joseba Rodríguez, Massimo Poesio |
LREC | 6 |
| 2015 | Signal: Advanced Real-Time Information Filtering
Miguel Martinez-Alvarez, Udo Kruschwitz, Wesley Hall, Massimo Poesio |
ECIR | 4 |
| 2015 | Phrase Detectives: Utilizing Collective Intelligence for Internet-Scale Language Resource Creation (Extended Abstract)
Massimo Poesio, Jon Chamberlain, Udo Kruschwitz, Livio Robaldo, Luca Ducceschi |
IJCAI | 1 |
| 2015 | MultiLing 2015: Multilingual Summarization of Single and Multi-Documents, On-line Fora, and Call-center ConversationsabstractGeorge Giannakopoulos, Jeff Kubina, John Conroy, Josef Steinberger, Benoit Favre, Mijail Kabadjov, Udo Kruschwitz, Massimo Poesio. Proceedings of the 16th Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2015. George Giannakopoulos, Jeff Kubina, John M. Conroy, Josef Steinberger, Benoît Favre, Mijail A. Kabadjov, Udo Kruschwitz, Massimo Poesio |
SIGDIAL Conference | 8 |
| 2015 | Differential evolution-based feature selection technique for anaphora resolution
Utpal Kumar Sikdar, Asif Ekbal, Sriparna Saha 0001, Olga Uryupina, Massimo Poesio |
Soft Comput. | 5 |
| 2015 | Combining Minimally-supervised Methods for Arabic Named Entity RecognitionabstractSupervised methods can achieve high performance on NLP tasks, such as Named Entity Recognition (NER), but new annotations are required for every new domain and/or genre change. This has motivated research in minimally supervised methods such as semi-supervised learning and distant learning, but neither technique has yet achieved performance levels comparable to those of supervised methods. Semi-supervised methods tend to have very high precision but comparatively low recall, whereas distant learning tends to achieve higher recall but lower precision. This complementarity suggests that better results may be obtained by combining the two types of minimally supervised methods. In this paper we present a novel approach to Arabic NER using a combination of semi-supervised and distant learning techniques. We trained a semi-supervised NER classifier and another one using distant learning techniques, and then combined them using a variety of classifier combination schemes, including the Bayesian Classifier Combination (BCC) procedure recently proposed for sentiment analysis. According to our results, the BCC model leads to an increase in performance of 8 percentage points over the best base classifiers. Maha Althobaiti 0001, Udo Kruschwitz, Massimo Poesio |
Trans. Assoc. Comput. Linguistics | 3 |
| 2014 | Automatic Creation of Arabic Named Entity Annotated Corpus Using WikipediaabstractIn this paper we propose a new methodology to exploit Wikipedia features and structure to automatically develop an Arabic NE annotated corpus.Each Wikipedia link is transformed into an NE type of the target article in order to produce the NE annotation.Other Wikipedia features -namely redirects, anchor texts, and inter-language links -are used to tag additional NEs, which appear without links in Wikipedia texts.Furthermore, we have developed a filtering algorithm to eliminate ambiguity when tagging candidate NEs.Herein we also introduce a mechanism based on the high coverage of Wikipedia in order to address two challenges particular to tagging NEs in Arabic text: rich morphology and the absence of capitalisation.The corpus created with our new method (WDC) has been used to train an NE tagger which has been tested on different domains.Judging by the results, an NE tagger trained on WDC can compete with those trained on manually annotated corpora. Maha Althobaiti 0001, Udo Kruschwitz, Massimo Poesio |
EACL | 3 |
| 2014 | Identifying fake Amazon reviews as learning from crowdsabstractCustomers who buy products such as books online often rely on other customers reviews more than on reviews found on specialist magazines. Unfortunately the confidence in such reviews is often misplaced due to the explosion of so-called sock puppetry-Authors writing glowing reviews of their own books. Identifying such deceptive reviews is not easy. The first contribution of our work is the creation of a collection including a number of genuinely deceptive Amazon book reviews in collaboration with crime writer Jeremy Duns, who has devoted a great deal of effort in unmasking sock puppeting among his colleagues. But there can be no certainty concerning the other reviews in the collection: All we have is a number of cues, also developed in collaboration with Duns, suggesting that a review may be genuine or deceptive. Thus this corpus is an example of a collection where it is not possible to acquire the actual label for all instances, and where clues of deception were treated as annotators who assign them heuristic labels. A number of approaches have been proposed for such cases; we adopt here the 'learning from crowds' approach proposed by Raykar et al. (2010). Thanks to Duns' certainly fake reviews, the second contribution of this work consists in the evaluation of the effectiveness of different methods of annotation, according to the performance of models trained to detect deceptive reviews. © 2014 Association for Computational Linguistics. Tommaso Fornaciari, Massimo Poesio |
EACL | 2 |
| 2014 | AraNLP: a Java-based Library for the Processing of Arabic Text
Maha Althobaiti 0001, Udo Kruschwitz, Massimo Poesio |
LREC | 3 |
| 2013 | Of Words, Eyes and Brains: Correlating Image-Based Distributional Semantic Models with Neural Representations of ConceptsabstractTraditional distributional semantic models extract word meaning representations from cooccurrence patterns of words in text corpora.Recently, the distributional approach has been extended to models that record the cooccurrence of words with visual features in image collections.These image-based models should be complementary to text-based ones, providing a more cognitively plausible view of meaning grounded in visual perception.In this study, we test whether image-based models capture the semantic patterns that emerge from fMRI recordings of the neural signal.Our results indicate that, indeed, there is a significant correlation between image-based and brain-based semantic similarities, and that image-based models complement text-based ones, so that the best correlations are achieved when the two modalities are combined.Despite some unsatisfactory, but explained outcomes (in particular, failure to detect differential association of models with brain areas), the results show, on the one hand, that imagebased distributional semantic models can be a precious new tool to explore semantic representation in the brain, and, on the other, that neural data can be used as the ultimate test set to validate artificial semantic models in terms of their cognitive plausibility. Andrew J. Anderson, Elia Bruni, Ulisse Bordignon, Massimo Poesio, Marco Baroni |
EMNLP | 4 |
| 2013 | Adapting a State-of-the-art Anaphora Resolution System for Resource-poor Language
Utpal Kumar Sikdar, Asif Ekbal, Sriparna Saha 0001, Olga Uryupina, Massimo Poesio |
IJCNLP | 5 |
| 2013 | Phrase detectives: Utilizing collective intelligence for internet-scale language resource creationabstractWe are witnessing a paradigm shift in Human Language Technology (HLT) that may well have an impact on the field comparable to the statistical revolution: acquiring large-scale resources by exploiting collective intelligence. An illustration of this new approach is Phrase Detectives , an interactive online game with a purpose for creating anaphorically annotated resources that makes use of a highly distributed population of contributors with different levels of expertise. The purpose of this article is to first of all give an overview of all aspects of Phrase Detectives, from the design of the game and the HLT methods we used to the results we have obtained so far. It furthermore summarizes the lessons that we have learned in developing this game which should help other researchers to design and implement similar games. Massimo Poesio, Jon Chamberlain, Udo Kruschwitz, Livio Robaldo, Luca Ducceschi |
ACM Trans. Interact. Intell. Syst. | 1 |
| 2012 | DeCour: a corpus of DEceptive statements in Italian COURts
Tommaso Fornaciari, Massimo Poesio |
LREC | 2 |
| 2012 | Domain-specific vs. Uniform Modeling for Coreference Resolution
Olga Uryupina, Massimo Poesio |
LREC | 2 |
| 2011 | A Cross-Lingual ILP Solution to Zero Anaphora Resolution
Ryu Iida, Massimo Poesio |
ACL | 2 |
| 2011 | Single and multi-objective optimization for feature selection in anaphora resolution
Sriparna Saha 0001, Asif Ekbal, Olga Uryupina, Massimo Poesio |
IJCNLP | 4 |
| 2010 | Enhancing N-Gram-Based Summary Evaluation Using Information Content and a Taxonomy
Mijail A. Kabadjov, Josef Steinberger, Ralf Steinberger, Massimo Poesio, Bruno Pouliquen |
ECIR | 4 |
| 2010 | Extending BART to Provide a Coreference Resolution System for German
Samuel Broscheit, Simone Paolo Ponzetto, Yannick Versley, Massimo Poesio |
LREC | 4 |
| 2010 | BabyExp: Constructing a Huge Multimodal Resource to Acquire Commonsense Knowledge Like Children Do
Massimo Poesio, Marco Baroni, Oswald Lanz, Alessandro Lenci, Alexandros Potamianos, Hinrich Schütze, Sabine Schulte im Walde, Luca Surian |
LREC | 1 |
| 2010 | Creating a Coreference Resolution System for Italian
Massimo Poesio, Olga Uryupina, Yannick Versley |
LREC | 1 |
| 2010 | Anaphoric Annotation of Wikipedia and Blogs in the Live Memories Corpus
Kepa Joseba Rodríguez, Francesca Delogu, Yannick Versley, Egon Stemle, Massimo Poesio |
LREC | 5 |
| 2009 | EEG responds to conceptual stimuli and corpus semantics
Brian Murphy, Marco Baroni, Massimo Poesio |
EMNLP | 3 |
| 2009 | Interactive Gesture in Dialogue: a PTT Model
Hannes Rieser, Massimo Poesio |
SIGDIAL Conference | 2 |
| 2009 | Evaluating Centering for Information Ordering Using CorporaabstractIn this article we discuss several metrics of coherence defined using centering theory and investigate the usefulness of such metrics for information ordering in automatic text generation. We estimate empirically which is the most promising metric and how useful this metric is using a general methodology applied on several corpora. Our main result is that the simplest metric (which relies exclusively on NOCB transitions) sets a robust baseline that cannot be outperformed by other metrics which make use of additional centering-based features. This baseline can be used for the development of both text-to-text and concept-to-text generation systems. Nikiforos Karamanis, Chris Mellish, Massimo Poesio, Jon Oberlander |
Comput. Linguistics | 3 |
| 2009 | Janet Hitzeman
Massimo Poesio, David Day, Inderjeet Mani |
Comput. Linguistics | 1 |
| 2008 | Coreference Systems Based on Kernels Methods
Yannick Versley, Alessandro Moschitti, Massimo Poesio |
COLING | 3 |
| 2008 | A Corpus for Cross-Document Co-reference
David Day, Janet Hitzeman, Michael L. Wick, Keith Crouch, Massimo Poesio |
LREC | 5 |
| 2008 | Anaphoric Annotation in the ARRAU Corpus
Massimo Poesio, Ron Artstein |
LREC | 1 |
| 2008 | ANAWIKI: Creating Anaphorically Annotated Resources through Web Cooperation
Massimo Poesio, Udo Kruschwitz, Jon Chamberlain |
LREC | 1 |
| 2008 | BART: A modular toolkit for coreference resolution
Yannick Versley, Simone Paolo Ponzetto, Massimo Poesio, Vladimir Eidelman, Alan Jern, Jason Smith 0006, Alessandro Moschitti |
LREC | 3 |
| 2008 | Inter-Coder Agreement for Computational LinguisticsabstractThis article is a survey of methods for measuring agreement among corpus annotators. It exposes the mathematics and underlying assumptions of agreement coefficients, covering Krippendorff's alpha as well as Scott's pi and Cohen's kappa; discusses the use of coefficients in several annotation tasks; and argues that weighted, alpha-like coefficients, traditionally less used than kappa-like measures in computational linguistics, may be more appropriate for many corpus annotation tasks—but that their use makes the interpretation of the value of the coefficient even harder. Ron Artstein, Massimo Poesio |
Comput. Linguistics | 2 |
| 2007 | Two uses of anaphora resolution in summarization
Josef Steinberger, Massimo Poesio, Mijail A. Kabadjov, Karel Jezek |
Inf. Process. Manag. | 2 |
| 2006 | Extending the single words-based document model: a comparison of bigrams and 2-itemsetsabstractThe basic approach in text categorization is to represent documents by single words. However, often other features are utilized to achieve better classification results. In this paper, our attention is focused on bigrams and 2-itemsets. We compare the performance improvement in terms of classification accuracy when these features are used to extend the single words-based document representation on two standard text corpora: Reuters-21578 and 20 Newsgroups. For this comparison we use the multinomial Naive Bayes classifier and five different feature selection approaches. Algorithms for bigrams and 2-itemsets discovery are presented as well. Our results show a statistically significant improvement when bigrams and also 2-itemsets are incorporated. However, in the case of 2-itemsets it is important to use an appropriate feature selection method. On the other hand, even when a simple feature selection approach is applied to discover bigrams the classification accuracy improves. The conclusion is that, in our case, it is not very effective to extend document representation with 2-itemsets because bigrams achieve better results and discovering them is less resource-consuming. Roman Tesar, Vaclav Strnad, Karel Jezek, Massimo Poesio |
ACM Symposium on Document Engineering | 4 |
| 2006 | MSDA: Wordsense Discrimination Using Context Vectors and Attributes
Abdulrahman Almuhareb, Massimo Poesio |
ECAI | 2 |
| 2006 | An Anaphora Resolution-Based Anonymization Module
Massimo Poesio, Mijail A. Kabadjov, P. Goux, Udo Kruschwitz, Elizabeth Bishop, Louise Corti |
LREC | 1 |
| 2004 | Evaluating Centering-Based Metrics of CoherenceabstractWe use a reliably annotated corpus to compare metrics of coherence based on Centering Theory with respect to their potential usefulness for text structuring in natural language generation. Previous corpus-based evaluations of the coherence of text according to Centering did not compare the coherence of the chosen text structure with that of the possible alternatives. A corpus-based methodology is presented which distinguishes between Centering-based metrics taking these alternatives into account, and represents therefore a more appropriate way to evaluate Centering from a text structuring perspective. Nikiforos Karamanis, Massimo Poesio, Chris Mellish, Jon Oberlander |
ACL | 2 |
| 2004 | Learning to Resolve Bridging ReferencesabstractWe use machine learning techniques to find the best combination of local focus and lexical distance features for identifying the anchor of mereological bridging references.We find that using first mention, utterance distance, and lexical distance computed using either Google or WordNet results in an accuracy significantly higher than obtained in previous experiments. Massimo Poesio, Rahul Mehta 0001, Axel Maroudas, Janet Hitzeman |
ACL | 1 |
| 2004 | Attribute-Based and Value-Based Clustering: An Evaluation
Abdulrahman Almuhareb, Massimo Poesio |
EMNLP | 2 |
| 2004 | Identifying Broken Plurals in Unvowelised Arabic Tex
Abduelbaset Goweder, Massimo Poesio, Anne N. De Roeck, Jeff Reynolds |
EMNLP | 2 |
| 2004 | A Corpus-Based Methodology for Evaluating Metrics of Coherence for Text Structuring
Nikiforos Karamanis, Chris Mellish, Jon Oberlander, Massimo Poesio |
INLG | 4 |
| 2004 | A General-Purpose, Off-the-shelf Anaphora Resolution Module: Implementation and Preliminary Evaluation
Massimo Poesio, Mijail A. Kabadjov |
LREC | 1 |
| 2004 | Acquiring Bayesian Networks from Text
Olivia Sanchez-Graillet, Massimo Poesio |
LREC | 2 |
| 2004 | Broken plural detection for arabic information retrievalabstractDue to the high number of inflectional variations of Arabic words, empirical results suggest that stemming is essential for Arabic information retrieval. However, current light stemming algorithms do not extract the correct stem of irregular (so-called broken) plurals, which constitute ~10% of Arabic texts and ~41% of plurals. Although light stemming in particular has led to improvements in information retrieval [5, 6], the effects of broken plurals on the performance of information retrieval systems has not been examined.We propose a light stemmer that incorporates a broken plural recognition component, and evaluate it within the context of information retrieval. Our results show that identifying broken plurals and reducing them to their correct stems does result in a significant improvement in the performance of information retrieval systems. Abduelbaset Goweder, Massimo Poesio, Anne N. De Roeck |
SIGIR | 2 |
| 2004 | Centering: A Parametric Theory and Its InstantiationsabstractCentering theory is the best-known framework for theorizing about local coherence and salience; however, its claims are articulated in terms of notions which are only partially specified, such as “utterance,” “realization,” or “ranking.” A great deal of research has attempted to arrive at more detailed specifications of these parameters of the theory; as a result, the claims of centering can be instantiated in many different ways. We investigated in a systematic fashion the effect on the theory's claims of these different ways of setting the parameters. Doing this required, first of all, clarifying what the theory's claims are (one of our conclusions being that what has become known as “Constraint 1” is actually a central claim of the theory). Secondly, we had to clearly identify these parametric aspects: For example, we argue that the notion of “pronoun” used in Rule 1 should be considered a parameter. Thirdly, we had to find appropriate methods for evaluating these claims. We found that while the theory's main claim about salience and pronominalization, Rule 1—a preference for pronominalizing the backward-looking center (CB)—is verified with most instantiations, Constraint 1–a claim about (entity) coherence and CB uniqueness—is much more instantiation-dependent: It is not verified if the parameters are instantiated according to very mainstream views (“vanilla instantiation”), it holds only if indirect realization is allowed, and is violated by between 20% and 25% of utterances in our corpus even with the most favorable instantiations. We also found a trade-off between Rule 1, on the one hand, and Constraint 1 and Rule 2, on the other: Setting the parameters to minimize the violations of local coherence leads to increased violations of salience, and vice versa. Our results suggest that “entity” coherence—continuous reference to the same entities—must be supplemented at least by an account of relational coherence. Massimo Poesio, Rosemary Stevenson, Barbara Di Eugenio, Janet Hitzeman |
Comput. Linguistics | 1 |
| 2002 | Acquiring Lexical Knowledge for Anaphora Resolution
Massimo Poesio, Tomonori Ishikawa, Sabine Schulte im Walde, Renata Vieira |
LREC | 1 |
| 2002 | Automatically predicting dialogue structure using prosodic features
Helen Hastie, Massimo Poesio, Stephen Isard |
Speech Commun. | 2 |
| 2001 | Corpus-based NP Modifier Generation
Hua Cheng 0001, Massimo Poesio, Renate Henschel, Chris Mellish |
NAACL | 2 |
| 2000 | Specifying the Parameters of Centering Theory: a Corpus-Based Evaluation using Text from Application-Oriented DomainsabstractThe definitions of the basic concepts, rules, and constraints of centering theory involve underspecified notions such as 'previous utterance', 'realization', and 'ranking'. We attempted to find the best way of defining each such notion among those that can be annotated reliably, and using a corpus of texts in two domains of practical interest. Our main result is that trying to reduce the number of utterances without a backward-looking center (CB) results in an increased number of cases in which some discourse entity, but not the CB, gets pronominalized, and viceversa. Massimo Poesio, Hua Cheng 0001, Renate Henschel, Janet Hitzeman, Rodger Kibble, Rosemary Stevenson |
ACL | 1 |
| 2000 | Pronominalization revisited
Renate Henschel, Hua Cheng 0001, Massimo Poesio |
COLING | 3 |
| 2000 | Corpus-based Development and Evaluation of a System for Processing Definite Descriptions
Renata Vieira, Massimo Poesio |
COLING | 2 |
| 2000 | Annotating a Corpus to Develop and Evaluate Discourse Entity Realization Algorithms: Issues and Preliminary Results
Massimo Poesio |
LREC | 1 |
| 2000 | An Empirically-based System for Processing Definite DescriptionsabstractWe present an implemented system for processing definite descriptions in arbitrary domains. The design of the system is based on the results of a corpus analysis previously reported, which highlighted the prevalence of discourse-new descriptions in newspaper corpora. The annotated corpus was used to extensively evaluate the proposed techniques for matching definite descriptions with their antecedents, discourse segmentation, recognizing discourse-new descriptions, and suggesting anchors for bridging descriptions. Renata Vieira, Massimo Poesio |
Comput. Linguistics | 2 |
| 1998 | The provision of corrective feedback in a spoken dialogue CALL system
Sarah Davies, Massimo Poesio |
ICSLP | 2 |
| 1998 | The predictive power of game structure in dialogue act recognition: experimental results using maximum entropy estimationabstractRecognizing the dialogue act(s) performed by means of an utterance involves combining top-down expectations about the next likely `move' in a dialogue with bottom-up information extracted from the speech signal. We compare two ways of generating expectations: one which makes the expectations depend only on the previous act, and one which also takes into account the fact that individual dialogue acts play a role as part of larger conversational structures (`games'). Our results indicate that exploiting game structure does lead to improved expectations. 1. INTRODUCTION Recognizing the dialogue act(s) performed by means of an utterance involves combining top-down expectations about the next likely `move' in a dialogue with bottom-up information extracted from the speech signal. The best current models of dialogue act recognition achieve an accuracy of about 70% on transcribed words and of 65% on recognized words (Stolcke et al., 1998; Reithinger and Klesen, 1997). We are trying to impro... Massimo Poesio, Andrei Mikheev |
ICSLP | 1 |
| 1998 | A Corpus-based Investigation of Definite Description Use
Massimo Poesio, Renata Vieira |
Comput. Linguistics | 1 |
| 1997 | Conversational Actions and Discourse SituationsabstractWe use the idea that actions performed in a conversation become part of the common ground as the basis for a model of context that reconciles in a general and systematic fashion the differences between the theories of discourse context used for reference resolution, intention recognition, and dialogue management. We start from the treatment of anaphoric accessibility developed in discourse representation theory (DRT), and we show first how to obtain a discourse model that, while preserving DRT's basic ideas about referential accessibility, includes information about the occurrence of speech acts and their relations. Next, we show how the different kinds of ‘structure’ that play a role in conversation—discourse segmentation, turn‐taking, and grounding—can be formulated in terms of information about speech acts, and use this same information as the basis for a model of the interpretation of fragmentary input. Massimo Poesio, David R. Traum |
Comput. Intell. | 1 |
| 1995 | The TRAINS project: a case study in building a conversational planning agentabstractThe TRAINS project is an effort to build a conversationally proficient planning assistant. A key part of the project is the construction of the TRAINS system, which provides the research platform for a wide range of issues in natural language understanding, mixed-initiative planning systems, and representing and reasoning about time, actions and events. Four years have now passed since the beginning of the project. Each year a demonstration system has been produced that focused on a dialogue that illustrates particular aspects of the research. The commitment to building complete integrated systems is a significant overhead on the research, but it is considered essential to guarantee that the results constitute real progress in the field. This paper describes the goals of the project, and the experience with the effort so far. James F. Allen, Lenhart K. Schubert, George Ferguson, Peter A. Heeman, Chung Hee Hwang, Tsuneaki Kato, Marc Light, Nathaniel G. Martin, Bradford W. Miller, Massimo Poesio, David R. Traum |
J. Exp. Theor. Artif. Intell. | 10 |
| 1993 | Temporal CenteringabstractWe present a semantic and pragmatic account of the anaphoric properties of past and perfect that improves on previous work by integrating discourse structure, aspectual type, surface structure and commonsense knowledge. A novel aspect of our account is that we distinguish between two kinds of temporal intervals in the interpretation of temporal operators --- discourse reference intervals and event intervals. This distinction makes it possible to develop an analogy between centering and temporal centering, which operates on discourse reference intervals. Our temporal property-sharing principle is a defeasible inference rule on the logical form. Along with lexical and causal reasoning, it plays a role in incrementally resolving underspecified aspects of the event structure representation of an utterance against the current context. Megumi Kameyama, Rebecca J. Passonneau, Massimo Poesio |
ACL | 3 |
| 1993 | Assigning a Semantic Scope to OperationsabstractI propose that the characteristics of the scope disambiguation process observed in the literature can be explained in terms of the way in which the model of the situation described by a sentence is built. The model construction procedure I present builds an event structure by identifying the situations associated with the operators in the sentence and their mutual dependency relations, as well as the relations between these situations and other situations in the context. The procedure takes into account lexical semantics and the result of various discourse interpretation procedures such as definite description interpretation, and does not require a complete disambiguation to take place. Massimo Poesio |
ACL | 1 |
| 1992 | Conversational Events and Discourse State Change: A Preliminary Report
Massimo Poesio |
KR | 1 |
| 1991 | Metric Constraints for Maintaining Appointments: Dates and Repeated Activities
Massimo Poesio, Ronald J. Brachman |
AAAI | 1 |
| 1988 | Toward a Hybrid Representation of Time
Massimo Poesio |
ECAI | 1 |
| 1987 | Modified Caseframe Parsing for Speech Understanding Systems
Massimo Poesio, Claudio Rullent |
IJCAI | 1 |