EDBT 2026 Demo / reviewers in the wild / expert
Sara Tonelli
dblp:20/1398
· DBLP profile ↗
58ranked-venue papers
10as first author
21since 2021 · last 2026
0000-0001-8010-6689ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 47 · 9 first-author · 18 since 2021Databases, data management, data science and information retrieval · 7 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Don't Stop the Multi-Party! On Generating Synthetic Written Multi-Party Conversations with ConstraintsabstractWritten Multi-Party Conversations (WMPCs) are widely studied across disciplines, with social media as a primary data source due to their accessibility. However, these datasets raise privacy concerns and often reflect platform-specific properties. For example, interactions between speakers may be limited due to rigid platform structures (e.g., threads, tree-like discussions), which yield overly simplistic interaction patterns (e.g., one-to-one ``reply-to'' links). This work explores the feasibility of generating synthetic WMPCs with instruction-tuned Large Language Models (LLMs) by providing deterministic constraints such as dialogue structure and participants’ stance. We investigate two complementary strategies of leveraging LLMs in this context: (i.) LLMs as WMPC generators, where we task the LLM to generate a whole WMPC at once and (ii.) LLMs as WMPC parties, where the LLM generates one turn of the conversation at a time (made of speaker, addressee and message), provided the conversation history. We next introduce an analytical framework to evaluate compliance with the constraints, content quality, and interaction complexity for both strategies. Finally, we assess the level of obtained WMPCs via human and LLM-as-a-judge evaluations. We find stark differences among LLMs, with only some being able to generate high-quality WMPCs. We also find that turn-by-turn generation yields better conformance to constraints and higher linguistic variability than generating WMPCs in one pass. Nonetheless, our structural and qualitative evaluation indicates that both generation strategies can yield high-quality WMPCs. Nicolò Penzo, Marco Guerini, Bruno Lepri, Goran Glavas, Sara Tonelli |
AAAI | 5 |
| 2026 | Mute Cods: A Multilingual Telegram Dataset with Benchmark Models for Conspiracy Theory Detection
Katarina Laken, Erik Bran Marino, Paloma Piot-Perez-Abadin, Davide Bassi, Søren Fomsgaard, Michele Joshua Maggini, Renata Vieira, Marcos García 0001, Sara Tonelli |
LREC | 9 |
| 2026 | University Speaking for Everyone: Assessing Changes in Italian Higher Education Statutes toward Gender-Inclusive LanguageabstractWe examine the editorial evolution of Italian university statutes toward inclusive language, analyzing how institutions represent female and non-binary identities and how these representations affect administrative communication. To this end, we compile and annotate a corpus of university statutes, tracing the changes that have led some universities to move from the use of the generic masculine to more inclusive formulations. We also experiment with tools for the automatic detection of non-inclusive language in institutional communication and methods for the automatic rewriting of texts into inclusive language. Sebastiano Vecellio Salto, Camilla Casula, Alessio Palmero Aprosio, Sara Tonelli |
LREC | 4 |
| 2026 | A Survey on Automatic Credibility Assessment Using Textual Credibility Signals in the Era of Large Language ModelsabstractIn the age of social media and generative AI, the ability to automatically assess the credibility of online content has become increasingly critical, complementing traditional approaches to false information detection. Credibility assessment relies on aggregating diverse credibility signals—small units of information, such as content subjectivity, bias or a presence of persuasion techniques—into a final credibility label/score. However, current research in automatic credibility assessment and credibility signals detection remains highly fragmented, with many signals studied in isolation and lacking integration. Notably, there is a scarcity of approaches that detect and aggregate multiple credibility signals simultaneously. These challenges are further exacerbated by the absence of a comprehensive and up-to-date overview of research works that connects these research efforts under a common framework and identifies shared trends, challenges and open problems. In this survey, we address this gap by presenting a systematic and comprehensive literature review of 175 research papers, focusing on textual credibility signals within the field of Natural Language Processing (NLP), which undergoes a rapid transformation due to advancements in Large Language Models (LLMs). While positioning the NLP research into the broader multidisciplinary landscape, we examine both automatic credibility assessment methods as well as the detection of nine categories of credibility signals. We provide an in-depth analysis of three key categories: (1) factuality, subjectivity and bias, (2) persuasion techniques and logical fallacies and (3) check-worthy and fact-checked claims. In addition to summarising existing methods, datasets and tools, we outline future research direction and emerging opportunities, with particular attention to evolving challenges posed by generative AI. Ivan Srba, Olesya Razuvayevskaya, João Augusto Leite, Róbert Móro, Ipek Baris Schlicht, Sara Tonelli, Francisco Moreno García, Santiago Barrio Lottmann, Denis Teyssou, Valentin Porcellini, Carolina Scarton, Kalina Bontcheva, Mária Bieliková |
ACM Trans. Intell. Syst. Technol. | 6 |
| 2025 | ModaFact: Multi-paradigm Evaluation for Joint Event Modality and Factuality DetectionabstractFactuality and modality are two crucial aspects concerning events, since they convey the speaker’s commitment to a situation in discourse as well as how this event is supposed to occur in terms of norms, wishes, necessity, duty and so on. Capturing them both is necessary to truly understand an utterance meaning and the speaker’s perspective with respect to a mentioned event. Yet, NLP studies have mostly dealt with these two aspects separately, mainly devoting past efforts to the development of English datasets. In this work, we propose ModaFact, a novel resource with joint factuality and modality information for event-denoting expressions in Italian. We propose a novel annotation scheme, which however is consistent with existing ones, and compare different classification systems trained on ModaFact, as a preliminary step to the use of factuality and modality information in downstream tasks. The dataset and the best-performing model are publicly released and available under an open license. Marco Rovera, Serena Cristoforetti, Sara Tonelli |
COLING | 3 |
| 2025 | Job Unfair: An Investigation of Gender and Occupational Bias in Free-Form Text Completions by LLMsabstractDisentangling how gender and occupations are encoded by LLMs is crucial to identify possible biases and prevent harms, especially given the widespread use of LLMs in sensitive domains such as human resources.In this work, we carry out an in-depth investigation of gender and occupational biases in English and Italian as expressed by 9 different LLMs (both base and instruction-tuned).Specifically, we focus on the analysis of sentence completions when LLMs are prompted with job-related sentences including different gender representations.We carry out a manual analysis of 4,500 generated texts over 4 dimensions that can reflect bias, we propose a novel embedding-based method to investigate biases in generated texts and, finally, we carry out a lexical analysis of the model completions.In our qualitative and quantitative evaluation we show that many facets of social bias remain unaccounted for even in aligned models, and LLMs in general still reflect existing gender biases in both languages.Finally, we find that models still struggle with genderneutral expressions, especially beyond English. ModelSubject M (%) F (%) N (%) aya-expanse-8b Abstract 0.00% 0.00% 3.66% Object 0.00% 0.00% 1.22% Profession 0.00% 0.00% 0.00% gemma-7b Abstract 0.00% 0.00% 9.09% Object 0.00% 1.15% 2.27% Profession 0.00% 0.00% 0.00% gemma-7b-instruct Abstract 0.00% 0.00% 1.22% Object 0.00% 0.00% 0.00% Profession 0.00% 0.00% 0.00% Llama-70B Abstract 0.00% 0.00% 1.47% Object 0.00% 0.00% 1.47% Profession 0.00% 0.00% 1.47% Llama-70B-instruct Abstract 0.00% 0.00% 2.56% Object 0.00% 0.00% 0.00% Profession 1.22% 0.00% 1.28% Llama-8B Abstract 0.00% 0.00% 1.18% Object 0.00% 0.00% 2.35% Profession 0.00% 0.00% 1.18% Llama-8B-instruct Abstract 0.00% 0.00% 0.00% Object 0.00% 0.00% 0.00% Profession 0.00% 0.00% 0.00% Mistral-7B Abstract 0.00% 0.00% 3.33% Object 0.00% 0.00% 1.11% Profession 0.00% 0.00% 1.11% Mistral-7B-instruct Abstract 0.00% 0.00% 1.16% Object 0.00% 0.00% 1.16% Profession 5.62% 3.80% 4.65%Table 7: Subject misinterpretation distributions by model and gender for English.Model Subject M (%) F (%) N (%) Camilla Casula, Sebastiano Vecellio Salto, Elisa Leonardelli, Sara Tonelli |
EMNLP | 4 |
| 2025 | Multilingual vs Crosslingual Retrieval of Fact-Checked Claims: A Tale of Two ApproachesabstractRetrieval of previously fact-checked claims is a well-established task, whose automation can assist professional fact-checkers in the initial steps of information verification.Previous works have mostly tackled the task monolingually, i.e., having both the input and the retrieved claims in the same language.However, especially for languages with a limited availability of fact-checks and in case of global narratives, such as pandemics, wars, or international politics, it is crucial to be able to retrieve claims across languages.In this work, we examine strategies to improve the multilingual and crosslingual performance, namely selection of negative examples (in the supervised) and re-ranking (in the unsupervised setting).We evaluate all approaches on a dataset containing posts and claims in 47 languages (283 language combinations).We observe that the best results are obtained by using LLMbased re-ranking, followed by fine-tuning with negative examples sampled using a sentence similarity-based strategy.Most importantly, we show that crosslinguality is a setup with its own unique characteristics compared to the multilingual setup. 1 * These authors contributed equally to this work. Alan Ramponi, Marco Rovera, Róbert Móro, Sara Tonelli |
EMNLP | 4 |
| 2025 | Fine-grained Fallacy Detection with Human Label VariationabstractAlan Ramponi, Agnese Daffara, Sara Tonelli. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Alan Ramponi, Agnese Daffara, Sara Tonelli |
NAACL (Long Papers) | 3 |
| 2024 | Putting Context in Context: the Impact of Discussion Structure on Text ClassificationabstractNicolò Penzo, Antonio Longa, Bruno Lepri, Sara Tonelli, Marco Guerini. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Nicolò Penzo, Antonio Longa, Bruno Lepri, Sara Tonelli, Marco Guerini |
EACL (1) | 4 |
| 2024 | Delving into Qualitative Implications of Synthetic Data for Hate Speech DetectionabstractThe use of synthetic data for training models for a variety of NLP tasks is now widespread.However, previous work reports mixed results with regards to its effectiveness on highly subjective tasks such as hate speech detection.In this paper, we present an in-depth qualitative analysis of the potential and specific pitfalls of synthetic data for hate speech detection in English, with 3,500 manually annotated examples.We show that, across different models, synthetic data created through paraphrasing gold texts can improve out-of-distribution robustness from a computational standpoint.However, this comes at a cost: synthetic data fails to reliably reflect the characteristics of real-world data on a number of linguistic dimensions, it results in drastically different class distributions, and it heavily reduces the representation of both specific identity groups and intersectional hate.Warning: this paper contains examples that may be offensive or upsetting. Camilla Casula, Sebastiano Vecellio Salto, Alan Ramponi, Sara Tonelli |
EMNLP | 4 |
| 2024 | Do LLMs suffer from Multi-Party Hangover? A Diagnostic Approach to Addressee Recognition and Response Selection in ConversationsabstractAssessing the performance of systems to classify Multi-Party Conversations (MPC) is challenging due to the interconnection between linguistic and structural characteristics of conversations.Conventional evaluation methods often overlook variances in model behavior across different levels of structural complexity on interaction graphs.In this work, we propose a methodological pipeline to investigate model performance across specific structural attributes of conversations.As a proof of concept we focus on Response Selection and Addressee Recognition tasks, to diagnose model weaknesses.To this end, we extract representative diagnostic subdatasets with a fixed number of users and a good structural variety from a large and open corpus of online MPCs.We further frame our work in terms of data minimization, avoiding the use of original usernames to preserve privacy, and propose alternatives to using original text messages.Results show that response selection relies more on the textual content of conversations, while addressee recognition requires capturing their structural dimension.Using an LLM in a zero-shot setting, we further highlight how sensitivity to prompt variations is task-dependent. Nicolò Penzo, Maryam Sajedinia, Bruno Lepri, Sara Tonelli, Marco Guerini |
EMNLP | 4 |
| 2024 | The Geography of Information Diffusion in Online Discourse on Europe and MigrationabstractThe online diffusion of information related to Europe and migration has been little investigated from an external point of view. However, this is a very relevant topic, especially if users have had no direct contact with Europe and its perception depends solely on information retrieved online. In this work we analyse the information circulating online about Europe and migration after retrieving a large amount of data from social media (Twitter), to gain new insights into topics, magnitude, and dynamics of their diffusion. We combine retweets and hashtags network analysis with geolocation of users, linking thus data to geography and allowing analysis from an “outside Europe” perspective, with a special focus on Africa. We also introduce a novel approach based on cross-lingual quotes, i.e. when content in a language is commented and retweeted in another language, assuming these interactions are a proxy for connections between very distant communities. Results show how the majority of online discussions occurs at a national level, especially when discussing migration. Language (English) is pivotal for information to become transnational and reach far. Transnational information flow is strongly unbalanced, with content mainly produced in Europe and amplified outside. Conversely Europe-based accounts tend to be self-referential when they discuss migration-related topics. Football is the most exported topic from Europe worldwide. Moreover, important nodes in the communities discussing migration-related topics include accounts of official institutions and international agencies, together with journalists, news, commentators and activists. Elisa Leonardelli, Sara Tonelli |
ICWSM | 2 |
| 2023 | BiRDy: Bullying Role Detection in Multi-Party ChatsabstractRecent studies have highlighted that private instant messaging platforms and channels are major media of cyber aggression, especially among teens. Due to the private nature of the verbal exchanges on these media, few studies have addressed the task of hate speech detection in this context. Moreover, the recent release of resources mimicking online aggression situations that may occur among teens on private instant messaging platforms is encouraging the development of solutions aiming at dealing with diversity in digital harassment. In this study, we present BiRDy: a fully Web-based platform performing participant role detection in multi-party chats. Leveraging the pre-trained language model mBERT (multilingual BERT), we release fine-tuned models relying on various contextual window strategies to classify exchanged messages according to the role of involvement in cyberbullying of the authors. Integrating a role scoring function, the proposed pipeline predicts a unique role for each chat participant. In addition, detailed confidence scoring are displayed. Currently, BiRDy publicly releases models for French and Italian. Anaïs Ollagnier, Elena Cabrio, Serena Villata, Sara Tonelli |
AAAI | 4 |
| 2023 | Generation-Based Data Augmentation for Offensive Language Detection: Is It Worth It?abstractGeneration-based data augmentation (DA) has been presented in several works as a way to improve offensive language detection.However, the effectiveness of generative DA has been shown only in limited scenarios, and the potential injection of biases when using generated data to classify offensive language has not been investigated.Our aim is that of analyzing the feasibility of generative data augmentation more in-depth with two main focuses.First, we investigate the robustness of models trained on generated data in a variety of data augmentation setups, both novel and already presented in previous work, and compare their performance on four widely-used English offensive language datasets that present inherent differences in terms of content and complexity.In addition to this, we analyze models using the HateCheck suite, a series of functional tests created to challenge hate speech detection systems.Second, we investigate potential lexical bias issues through a qualitative analysis of the generated data.We find that the potential positive impact of generative data augmentation on model performance is unreliable, and generative DA can also have unpredictable effects on lexical bias.Warning: this paper contains examples that may be offensive or upsetting. Camilla Casula, Sara Tonelli |
EACL | 2 |
| 2023 | Why Don't You Do It Right? Analysing Annotators' Disagreement in Subjective TasksabstractAnnotators' disagreement in linguistic data has been recently the focus of multiple initiatives aimed at raising awareness on issues related to 'majority voting' when aggregating diverging annotations.Disagreement can indeed reflect different aspects of linguistic annotation, from annotators' subjectivity to sloppiness or lack of enough context to interpret a text.In this work we first propose a taxonomy of possible reasons leading to annotators' disagreement in subjective tasks.Then, we manually label part of a Twitter dataset for offensive language detection in English following this taxonomy, identifying how the different categories are distributed.Finally we run a set of experiments aimed at assessing the impact of the different types of disagreement on classification performance.In particular, we investigate how accurately tweets belonging to different categories of disagreement can be classified as offensive or not, and how injecting data with different types of disagreement in the training set affects performance.We also perform offensive language detection as a multi-task framework, using disagreement classification as an auxiliary task.Warning: This paper contains examples that may be offensive or upsetting. Marta Sandri, Elisa Leonardelli, Sara Tonelli, Elisabetta Jezek |
EACL | 3 |
| 2022 | Work Hard, Play Hard: Collecting Acceptability Annotations through a 3D GameabstractCorpus-based studies on acceptability judgements have always stimulated the interest of researchers, both in theoretical and computational fields. Some approaches focused on spontaneous judgements collected through different types of tasks, others on data annotated through crowd-sourcing platforms, still others relied on expert annotated data available from the literature. The release of CoLA corpus, a large-scale corpus of sentences extracted from linguistic handbooks as examples of acceptable/non acceptable phenomena in English, has revived interest in the reliability of judgements of linguistic experts vs. non-experts. Several issues are still open. In this work, we contribute to this debate by presenting a 3D video game that was used to collect acceptability judgments on Italian sentences. We analyse the resulting annotations in terms of agreement among players and by comparing them with experts’ acceptability judgments. We also discuss different game settings to assess their impact on participants’ motivation and engagement. The final dataset containing 1,062 sentences, which were selected based on majority voting, is released for future research and comparisons. Federico Bonetti, Elisa Leonardelli, Daniela Trotta, Raffaele Guarasci, Sara Tonelli |
LREC | 5 |
| 2022 | Building a Multilingual Taxonomy of Olfactory Terms with TimestampsabstractOlfactory references play a crucial role in our memory and, more generally, in our experiences, since researchers have shown that smell is the sense that is most directly connected with emotions. Nevertheless, only few works in NLP have tried to capture this sensory dimension from a computational perspective. One of the main challenges is the lack of a systematic and consistent taxonomy of olfactory information, where concepts are organised also in a multi-lingual perspective. WordNet represents a valuable starting point in this direction, which can be semi-automatically extended taking advantage of Google n-grams and of existing language models. In this work we describe the process that has led to the semi-automatic development of a taxonomy for olfactory information in four languages (English, French, German and Italian), detailing the different steps and the intermediate evaluations. Along with being multi-lingual, the taxonomy also encloses temporal marks for olfactory terms thus making it a valuable resource for historical content analysis. The resource has been released and is freely available. Stefano Menini, Teresa Paccosi, Serra Sinem Tekiroglu, Sara Tonelli |
LREC | 4 |
| 2022 | Features or Spurious Artifacts? Data-centric Baselines for Fair and Robust Hate Speech DetectionabstractWarning: this paper contains content that may be offensive or upsetting.Avoiding to rely on dataset artifacts to predict hate speech is at the cornerstone of robust and fair hate speech detection.In this paper we critically analyze lexical biases in hate speech detection via a cross-platform study, disentangling various types of spurious and authentic artifacts and analyzing their impact on out-of-distribution fairness and robustness.We experiment with existing approaches and propose simple yet surprisingly effective datacentric baselines.Our results on English data across four platforms show that distinct spurious artifacts require different treatments to ultimately attain both robustness and fairness in hate speech detection.To encourage research in this direction, we release all baseline models and the code to compute artifacts, pointing it out as a complementary and necessary addition to the data statements practice.1 Alan Ramponi, Sara Tonelli |
NAACL-HLT | 2 |
| 2021 | Agreeing to Disagree: Annotating Offensive Language Datasets with Annotators' DisagreementabstractSince state-of-the-art approaches to offensive language detection rely on supervised learning, it is crucial to quickly adapt them to the continuously evolving scenario of social media.While several approaches have been proposed to tackle the problem from an algorithmic perspective, so to reduce the need for annotated data, less attention has been paid to the quality of these data.Following a trend that has emerged recently, we focus on the level of agreement among annotators while selecting data to create offensive language datasets, a task involving a high level of subjectivity.Our study comprises the creation of three novel datasets of English tweets covering different topics and having five crowd-sourced judgments each.We also present an extensive set of experiments showing that selecting training and test data according to different levels of annotators' agreement has a strong effect on classifiers performance and robustness.Our findings are further validated in cross-domain experiments and studied using a popular benchmark dataset.We show that such hard cases, where low agreement is present, are not necessarily due to poor-quality annotation and we advocate for a higher presence of ambiguous cases in future datasets, particularly in test sets, to better account for the different points of view expressed online. Elisa Leonardelli, Stefano Menini, Alessio Palmero Aprosio, Marco Guerini, Sara Tonelli |
EMNLP (1) | 5 |
| 2021 | A Smell is Worth a Thousand Words: Olfactory Information Extraction and Semantic Processing in a Multilingual Perspective (Invited Talk)abstractMore than any other sense, smell is linked directly to our emotions and our memories. However, smells are intangible and very difficult to preserve, making it hard to effectively identify, consolidate, and promote the wide-ranging role scents and smelling have in our cultural heritage. While some novel approaches have been recently proposed to monitor so-called urban smellscapes and analyse the olfactory dimension of our environments (Quercia et al., 2015), when it comes to smellscapes from the past little research has been done to keep track of how places, events and people have been described from an olfactory perspective. Fortunately, some key prerequisites for addressing this problem are now in place. In recent years, European cultural heritage institutions have invested heavily in large-scale digitisation: we hold a wealth of object, text and image data which can now be analysed using artificial intelligence. What remains missing is a methodology for the extraction of scent-related information from large amounts of texts, as well as a broader awareness of the wealth of historical olfactory descriptions, experiences and memories contained within the heritage datasets. In this talk, I will describe ongoing activities towards this goal, focused on text mining and semantic processing of olfactory information. I will present the general framework designed to annotate smell events in documents, and some preliminary results on information extraction approaches in a multilingual scenario. I will discuss the main findings and the challenges related to modelling textual descriptions of smells, including the metaphorical use of smell-related terms and the well-known limitations of smell vocabulary in European languages compared to other senses. Sara Tonelli |
LDK | 1 |
| 2021 | Measuring Orthogonal Mechanics in Linguistic Annotation GamesabstractGamification has been recently growing in popularity among researchers investigating Information and Communication Technologies. Scholars have been trying to take advantage of this approach in the field of natural language processing (NLP), developing Games With A Purpose (GWAPs) for corpus annotation that have obtained encouraging results both in annotation quality and overall cost. However, GWAPs implement gamification in different ways and to different degrees. We propose a new framework to investigate the mechanics employed in the gamification process and their magnitude in terms of complexity. This framework is based on an analysis of some of the most important contributions in the field of NLP-related gamified applications and GWAP theory. Its primary purpose is to provide a first step towards classifying mechanics that mimic mainstream video games and may require skills that are not relevant to the annotation task, defined as orthogonal mechanics. In order to test our framework, we develop and evaluate Spacewords, a linguistic space game for synonymy annotation. Federico Bonetti, Sara Tonelli |
Proc. ACM Hum. Comput. Interact. | 2 |
| 2020 | Adding Gesture, Posture and Facial Displays to the PoliModal Corpus of Political InterviewsabstractThis paper introduces a multimodal corpus in the political domain, which on top of transcribed face-to-face interviews presents the annotation of facial displays, hand gestures and body posture. While the fully annotated corpus consists of 3 interviews for a total of 90 minutes, it is extracted from a larger available corpus of 56 face-to-face interviews (14 hours) that has been manually annotated with information about metadata (i.e. tools used for the transcription, link to the interview etc.), pauses (used to mark a pause either between or within utterances), vocal expressions (marking non-lexical expressions such as burp and semi-lexical expressions such as primary interjections), deletions (false starts, repetitions and truncated words) and overlaps. In this work, we describe the additional level of annotation relating to nonverbal elements used by three Italian politicians belonging to three different political parties and who at the time of the talk-show were all candidates for the presidency of the Council of Minister. We also present the results of some analyses aimed at identifying existing relations between the proxemics phenomena and the linguistic structures in which they occur in order to capture recurring patterns and differences in the communication strategy. Daniela Trotta, Alessio Palmero Aprosio, Sara Tonelli, Annibale Elia |
LREC | 3 |
| 2020 | Adaptive Complex Word Identification through False Friend DetectionabstractAutomated complex word identification (CWI) is a crucial task in several applications, from readability assessment to lexical simplification. So far, several works have modeled CWI with the goal of targeting the needs of non-native speakers. However, studies in language acquisition show that different native languages can create positive or negative interferences w.r.t. reading comprehension, favouring or hindering the understanding of a document in a foreign language. Therefore, we propose to modify CWI to address the specific difficulties connected to different native languages. In particular, we present a pipeline that, based on the user native language, identifies complex terms by automatically detecting cognates and false friends on the fly. The selection presented by the CWI module is adaptive in that it changes depending on the native language of the user. We implement and evaluate our approach for four different native languages (French, English, German and Spanish), in a setting where documents are written in Italian and should be read by language learners with low proficiency. We show that a personalised strategy based on false friend detection identifies complex terms that are different from those usually selected with standard approaches based on word frequency. Alessio Palmero Aprosio, Stefano Menini, Sara Tonelli |
UMAP | 3 |
| 2020 | A Multilingual Evaluation for Online Hate Speech DetectionabstractThe increasing popularity of social media platforms such as Twitter and Facebook has led to a rise in the presence of hate and aggressive speech on these platforms. Despite the number of approaches recently proposed in the Natural Language Processing research area for detecting these forms of abusive language, the issue of identifying hate speech at scale is still an unsolved problem. In this article, we propose a robust neural architecture that is shown to perform in a satisfactory way across different languages; namely, English, Italian, and German. We address an extensive analysis of the obtained experimental results over the three languages to gain a better understanding of the contribution of the different components employed in the system, both from the architecture point of view (i.e., Long Short Term Memory, Gated Recurrent Unit, and bidirectional Long Short Term Memory) and from the feature selection point of view (i.e., ngrams, social network–specific features, emotion lexica, emojis, word embeddings). To address such in-depth analysis, we use three freely available datasets for hate speech detection on social media in English, Italian, and German. Michele Corazza, Stefano Menini, Elena Cabrio, Sara Tonelli, Serena Villata |
ACM Trans. Internet Techn. | 4 |
| 2020 | Web Semantics for Digital Humanities
Lora Aroyo, Franciska de Jong, Eero Hyvönen, Sara Tonelli |
J. Web Semant. | 4 |
| 2019 | Novel Event Detection and Classification for Historical TextsabstractEvent processing is an active area of research in the Natural Language Processing community, but resources and automatic systems developed so far have mainly addressed contemporary texts. However, the recognition and elaboration of events is a crucial step when dealing with historical texts Particularly in the current era of massive digitization of historical sources: Research in this domain can lead to the development of methodologies and tools that can assist historians in enhancing their work, while having an impact also on the field of Natural Language Processing. Our work aims at shedding light on the complex concept of events when dealing with historical texts. More specifically, we introduce new annotation guidelines for event mentions and types, categorized into 22 classes. Then, we annotate a historical corpus accordingly, and compare two approaches for automatic event detection and classification following this novel scheme. We believe that this work can foster research in a field of inquiry as yet underestimated in the area of Temporal Information Processing. To this end, we release new annotation guidelines, a corpus, and new models for automatic annotation. Rachele Sprugnoli, Sara Tonelli |
Comput. Linguistics | 2 |
| 2018 | Never Retreat, Never Retract: Argumentation Analysis for Political SpeechesabstractIn this work, we apply argumentation mining techniques, in particular relation prediction, to study political speeches in monological form, where there is no direct interaction between opponents. We argue that this kind of technique can effectively support researchers in history, social and political sciences, which must deal with an increasing amount of data in digital form and need ways to automatically extract and analyse argumentation patterns. We test and discuss our approach based on the analysis of documents issued by R. Nixon and J. F. Kennedy during 1960 presidential campaign. We rely on a supervised classifier to predict argument relations (i.e., support and attack), obtaining an accuracy of 0.72 on a dataset of 1,462 argument pairs. The application of argument mining to such data allows not only to highlight the main points of agreement and disagreement between the candidates' arguments over the campaign issues such as Cuba, disarmament and health-care, but also an in-depth argumentative analysis of the respective viewpoints on these topics. Stefano Menini, Elena Cabrio, Sara Tonelli, Serena Villata |
AAAI | 3 |
| 2017 | Topic-Based Agreement and Disagreement in US Electoral ManifestosabstractWe present a topic-based analysis of agreement and disagreement in political manifestos, which relies on a new method for topic detection based on key concept clustering.Our approach outperforms both standard techniques like LDA and a state-of-the-art graph-based method, and provides promising initial results for this new task in computational social science. Stefano Menini, Federico Nanni, Simone Paolo Ponzetto, Sara Tonelli |
EMNLP | 4 |
| 2017 | Leveraging bilingual terminology to improve machine translation in a CAT environmentabstractAbstract This work focuses on the extraction and integration of automatically aligned bilingual terminology into a Statistical Machine Translation (SMT) system in a Computer Aided Translation scenario. We evaluate the proposed framework that, taking as input a small set of parallel documents, gathers domain-specific bilingual terms and injects them into an SMT system to enhance translation quality. Therefore, we investigate several strategies to extract and align terminology across languages and to integrate it in an SMT system. We compare two terminology injection methods that can be easily used at run-time without altering the normal activity of an SMT system: XML markup and cache-based model. We test the cache-based model on two different domains (information technology and medical) in English, Italian and German, showing significant improvements ranging from 2.23 to 6.78 BLEU points over a baseline SMT system and from 0.05 to 3.03 compared to the widely-used XML markup approach. Mihael Arcan, Marco Turchi, Sara Tonelli, Paul Buitelaar |
Nat. Lang. Eng. | 3 |
| 2017 | One, no one and one hundred thousand events: Defining and processing events in an inter-disciplinary perspectiveabstractAbstract We present an overview of event definition and processing spanning 25 years of research in NLP. We first provide linguistic background to the notion of event, and then present past attempts to formalize this concept in annotation standards to foster the development of benchmarks for event extraction systems. This ranges from MUC-3 in 1991 to the Time and Space Track challenge at SemEval 2015. Besides, we shed light on other disciplines in which the notion of event plays a crucial role, with a focus on the historical domain. Our goal is to provide a comprehensive study on event definitions and investigate which potential past efforts in the NLP community may have in a different research domain. We present the results of a questionnaire, where the notion of event for historians is put in relation to the NLP perspective. Rachele Sprugnoli, Sara Tonelli |
Nat. Lang. Eng. | 2 |
| 2016 | Agreement and Disagreement: Comparison of Points of View in the Political DomainabstractThe automated comparison of points of view between two politicians is a very challenging task, due not only to the lack of annotated resources, but also to the different dimensions participating to the definition of agreement and disagreement. In order to shed light on this complex task, we first carry out a pilot study to manually annotate the components involved in detecting agreement and disagreement. Then, based on these findings, we implement different features to capture them automatically via supervised classification. We do not focus on debates in dialogical form, but we rather consider sets of documents, in which politicians may express their position with respect to different topics in an implicit or explicit way, like during an electoral campaign. We create and make available three different datasets. Stefano Menini, Sara Tonelli |
COLING | 2 |
| 2016 | CATENA: CAusal and TEmporal relation extraction from NAtural language textsabstractWe present CATENA, a sieve-based system to perform temporal and causal relation extraction and classification from English texts, exploiting the interaction between the temporal and the causal model. We evaluate the performance of each sieve, showing that the rule-based, the machine-learned and the reasoning components all contribute to achieving state-of-the-art performance on TempEval-3 and TimeBank-Dense data. Although causal relations are much sparser than temporal ones, the architecture and the selected features are mostly suitable to serve both tasks. The effects of the interaction between the temporal and the causal components, although limited, yield promising results and confirm the tight connection between the temporal and the causal dimension of texts. Paramita Mirza, Sara Tonelli |
COLING | 2 |
| 2016 | On the contribution of word embeddings to temporal relation classificationabstractTemporal relation classification is a challenging task, especially when there are no explicit markers to characterise the relation between temporal entities. This occurs frequently in inter-sentential relations, whose entities are not connected via direct syntactic relations making classification even more difficult. In these cases, resorting to features that focus on the semantic content of the event words may be very beneficial for inferring implicit relations. Specifically, while morpho-syntactic and context features are considered sufficient for classifying event-timex pairs, we believe that exploiting distributional semantic information about event words can benefit supervised classification of other types of pairs. In this work, we assess the impact of using word embeddings as features for event words in classifying temporal relations of event-event pairs and event-DCT (document creation time) pairs. Paramita Mirza, Sara Tonelli |
COLING | 2 |
| 2016 | Enriching a Small Artwork Collection Through Semantic Linking
Mauro Dragoni, Elena Cabrio, Sara Tonelli, Serena Villata |
ESWC | 3 |
| 2016 | NLP and Public Engagement: The Case of the Italian School Reform
Tommaso Caselli, Giovanni Moretti, Rachele Sprugnoli, Sara Tonelli, Damien Lanfrey, Donatella Solda Kutzmann |
LREC | 4 |
| 2016 | PreMOn: a Lemon Extension for Exposing Predicate Models as Linked Data
Francesco Corcoglioniti, Marco Rospocher, Alessio Palmero Aprosio, Sara Tonelli |
LREC | 4 |
| 2016 | ALCIDE: Extracting and visualising content from large document collections to support humanities studies
Giovanni Moretti, Rachele Sprugnoli, Stefano Menini, Sara Tonelli |
Knowl. Based Syst. | 4 |
| 2015 | Recognizing Biographical Sections in WikipediaabstractWikipedia is the largest collection of encyclopedic data ever written in the history of humanity.Thanks to its coverage and its availability in machine-readable format, it has become a primary resource for largescale research in historical and cultural studies.In this work, we focus on the subset of pages describing persons, and we investigate the task of recognizing biographical sections from them: given a person's page, we identify the list of sections where information about her/his life is present.We model this as a sequence classification problem, and propose a supervised setting, in which the training data are acquired automatically.Besides, we show that six simple features extracted only from the section titles are very informative and yield good results well above a strong baseline. Alessio Palmero Aprosio, Sara Tonelli |
EMNLP | 2 |
| 2015 | Getting the environmental information across: from the Web to the userabstractAbstract Environmental and meteorological conditions are of utmost importance for the population, as they are strongly related to the quality of life. Citizens are increasingly aware of this importance. This awareness results in an increasing demand for environmental information tailored to their specific needs and background. We present an environmental information platform that supports submission of user queries related to environmental conditions and orchestrates results from complementary services to generate personalized suggestions. The system discovers and processes reliable data in the Web in order to convert them into knowledge. At runtime, this information is transferred into an ontology‐structured knowledge base, from which then information relevant to the specific user is deduced and communicated in the language of their preference. The platform is demonstrated with real world use cases in the south area of Finland, showing the impact it can have on the quality of everyday life. Leo Wanner, Harald Bosch, Nadjet Bouayad-Agha, Gerard Casamayor, Thomas Ertl, Désirée Hilbring, Lasse Johansson, Kostas D. Karatzas, Ari Karppinen, Ioannis Kompatsiaris, Tarja Koskentalo, Simon Mille, Jürgen Moßgraber, Anastasia Moumtzidou, Maria Myllynen, Emanuele Pianta, Marco Rospocher, Luciano Serafini, Virpi Tarvainen, Sara Tonelli, Stefanos Vrochidis |
Expert Syst. J. Knowl. Eng. | 20 |
| 2014 | An Analysis of Causality between Events and its Relation to Temporal Information
Paramita Mirza, Sara Tonelli |
COLING | 2 |
| 2014 | Classifying Temporal Relations with Simple FeaturesabstractApproaching temporal link labelling as a classification task has already been explored in several works. However, choosing the right feature vectors to build the classification model is still an open issue, especially for event-event classification, whose accuracy is still under 50%. We find that using a simple feature set results in a better performance than using more sophisticated features based on semantic role labelling and deep semantic parsing. We also investigate the impact of extracting new training instances using inverse relations and transitive closure, and gain insight into the impact of this bootstrapping methodology on classifying the full set of TempEval-3 relations. Paramita Mirza, Sara Tonelli |
EACL | 2 |
| 2014 | CROMER: a Tool for Cross-Document Event and Entity Coreference
Christian Girardi, Manuela Speranza, Rachele Sprugnoli, Sara Tonelli |
LREC | 4 |
| 2013 | ERNESTA: A Sentence Simplification Tool for Children's Stories in Italian
Gianni Barlacchi, Sara Tonelli |
CICLing (2) | 2 |
| 2013 | Wikipedia-based WSD for multilingual frame annotation
Sara Tonelli, Claudio Giuliano, Kateryna Tymoshenko |
Artif. Intell. | 1 |
| 2012 | Hunting for Entailing Pairs in the Penn Discourse Treebank
Sara Tonelli, Elena Cabrio |
COLING | 1 |
| 2012 | Key-Concept Extraction for Ontology Engineering
Marco Rospocher, Sara Tonelli, Luciano Serafini, Emanuele Pianta |
EKAW | 2 |
| 2012 | Investigating the Semantics of Frame Elements
Sara Tonelli, Volha Bryl, Claudio Giuliano, Luciano Serafini |
EKAW | 1 |
| 2012 | Improving the Recall of a Discourse Parser by Constraint-based Postprocessing
Sucheta Ghosh, Richard Johansson, Giuseppe Riccardi, Sara Tonelli |
LREC | 4 |
| 2011 | Shallow Discourse Parsing with Conditional Random Fields
Sucheta Ghosh, Richard Johansson, Giuseppe Riccardi, Sara Tonelli |
IJCNLP | 4 |
| 2010 | Acoustic correlates of meaning structure in conversational speechabstractWe are interested in the problem of extracting meaning structures from spoken utterances in human communication. In Spoken Language Understanding (SLU) systems, parsing of meaning structures is carried over the word hypotheses generated by the Automatic Speech Recognizer (ASR). This approach suffers from high word error rates and ad-hoc conceptual representations. In contrast, in this paper we aim at discovering meaning components from direct measurements of acoustic and non-verbal linguistic features. The meaning structures are taken from the frame semantics model proposed in FrameNet, a consistent and extendable semantic structure resource covering a large set of domains. We give a quantitative analysis of meaning structures in terms of speech features across human–human dialogs from the manually annotated LUNA corpus. We show that the acoustic correlations between pitch, formant trajectories, intensity and harmonicity and meaning features are statistically significant over the whole corpus as well as relevant in classifying the target words evoked by a semantic frame. Alexei V. Ivanov, Giuseppe Riccardi, Sucheta Ghosh, Sara Tonelli, Evgeny A. Stepanov |
INTERSPEECH | 4 |
| 2010 | VenPro: A Morphological Analyzer for Venetan
Sara Tonelli, Emanuele Pianta, Rodolfo Delmonte, Michele Brunelli |
LREC | 1 |
| 2010 | Annotation of Discourse Relations for Conversational Spoken Dialogs
Sara Tonelli, Giuseppe Riccardi, Rashmi Prasad, Aravind K. Joshi |
LREC | 1 |
| 2009 | New Features for FrameNet - WordNet Mapping
Sara Tonelli, Daniele Pighin |
CoNLL | 1 |
| 2009 | Wikipedia as Frame Information Repository
Sara Tonelli, Claudio Giuliano |
EMNLP | 1 |
| 2008 | Enriching the Venice Italian Treebank with Dependency and Grammatical Relations
Sara Tonelli, Rodolfo Delmonte, Antonella Bristot |
LREC | 1 |
| 2008 | Frame Information Transfer from English to Italian
Sara Tonelli, Emanuele Pianta |
LREC | 1 |
| 2008 | Semantic annotations for conversational speech: From speech transcriptions to predicate argument structuresabstractIn this paper, we describe the semantic content, which can be automatically generated, for the design of advanced dialog systems. Since the latter will be based on machine learning approaches, we created training data by annotating a corpus with the needed content. Given a sentence of our transcribed corpus, domain concepts and other linguistic levels ranging from basic ones, i.e. part-of-speech tagging and constituent chunking level, to more advanced ones, i.e. syntactic and predicate argument structure (PAS) levels are annotated. In particular, the proposed PAS and taxonomy of dialog acts appear to be promising for the design of more complex dialog systems. Statistics about our semantic annotation are reported. Arianna Bisazza, Marco Dinarelli, Silvia Quarteroni, Sara Tonelli, Alessandro Moschitti, Giuseppe Riccardi |
SLT | 4 |
| 2008 | Automatic framenet-based annotation of conversational speechabstractCurrent Spoken Language Understanding technology is based on a simple concept annotation of word sequences, where the interdependencies between concepts and their compositional semantics are neglected. This prevents an effective handling of language phenomena, with a consequential limitation on the design of more complex dialog systems. In this paper, we argue that shallow semantic representation as formulated in the Berkeley FrameNet Project may be useful to improve the capability of managing more complex dialogs. To prove this, the first step is to show that a FrameNet parser of sufficient accuracy can be designed for conversational speech. We show that exploiting a small set of FrameNet-based manual annotations, it is possible to design an effective semantic parser. Our experiments on an Italian spoken dialog corpus, created within the LUNA project, show that our approach is able to automatically annotate unseen dialog turns with a high accuracy. Bonaventura Coppola, Alessandro Moschitti, Sara Tonelli, Giuseppe Riccardi |
SLT | 3 |