Elena Cabrio

dblp:35/7561 · DBLP profile ↗
← Back
59ranked-venue papers
9as first author
24since 2021 · last 2026
0000-0001-9374-7872ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 51 · 7 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 11 · 3 first-author · 4 since 2021Systems, architecture and hardware · 1Computer networks · 1Theory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 CRITICS: Critical Science Without Borders by Translation of Scientific Knowledge
abstract
The CRITICS project addresses science accessibility and literacy through the convergence of advanced Machine Translation (MT) based on Large Language Models (LLMs) and educational technology. By leveraging MT systems specifically optimized for scientific content, educational institutions can provide accurate, culturally relevant translations of scientific materials in higher-education students’ native languages, ensuring that complex scientific concepts are comprehensible while maintaining technical accuracy. Novel research on MT specifically tailored for scientific documents aims to break down language barriers in accessing cutting-edge research and educational materials currently only available in high-resourced languages such as English, thereby facilitating the democratization of scientific knowledge across linguistic boundaries.
Rodrigo Agerri, Itziar Aldabe, Elena Cabrio, Mark Cieliebak, Jan Deriu, Mariana Flores, Jurgita Kapociute-Dzikiene, Dovile Kuiziniene, Arantza Rico, Aritz Ruiz-González, Aitor Soroa, Mantas Vaskevicius, Serena Villata
EAMT (2)3
2026 Learnable Multi-Attribute Gradual Semantics for Predicting Persuasion in Argumentative Debates
abstract
Gradual semantics for weighted bipolar argumentation provide a principled framework for modelling argumentative reasoning, yet existing approaches remain mostly scalar, fixed, and weakly grounded in empirical data. We introduce learnable multi-attribute gradual semantics for persuasion prediction in argumentative debates. Our approach builds a dataset of 600 textual debates converted into multi-attribute argumentation graphs enriched with multi-dimensional features on nodes and relations. Building on this representation, we propose learnable aggregation operators that distinguish intrinsic quality from persuasive strategy dimensions. Experiments show that the learned semantics achieve competitive performance with neural and LLM-based baselines while preserving interpretability.
Nino Pireaud, Victor David, Anthony Hunter, Pierre Monnin, Elena Cabrio
KR5
2026 Is There Anything More Deceptive than an Obvious Fact? Investigating Implicitness in User-Generated Argumentative Text
Ekaterina Sviridova, Elena Cabrio, Serena Villata
LREC2
2026 "Detectors Lead, LLMs Follow": Integrating LLMs and traditional models on implicit hate speech detection to generate faithful and plausible explanations
abstract
Social media platforms face a growing challenge in addressing abusive content and hate speech, particularly as traditional natural language processing methods often struggle with detecting nuanced and implicit instances. To tackle this issue, our study enhances Large Language Models (LLMs) in the detection and explanation of implicit hate speech, outperforming classical approaches. We focus on two key objectives: (1) determining whether jointly predicting and generating explanations for why a message is hateful improves LLMs’ accuracy, especially for implicit cases, and (2) evaluating whether incorporating information from BERT-based models can further boost detection and explanation performance. Our method evaluates and enhances LLMs’ ability to detect hate speech and explain their predictions. By combining binary classification (Hate Speech vs. Non-Hate Speech) with natural language explanations, our approach provides clearer insights into why a message is considered hateful, advancing the accuracy and interpretability of hate speech detection.
Greta Damo, Nicolás Benjamín Ocampo, Elena Cabrio, Serena Villata
Data Knowl. Eng.3
2025 SAFE: Structured Argumentation for Fact-checking with Explanations
abstract
Explainable fact-checking plays a vital role in the fight against disinformation in today’s digital landscape. With the increasing volume of unverified content online, providing justifications for fact-checking has become essential to help users make informed decisions. While recent studies provide user-friendly explanations through abstractive or extractive summarization, they often assume the availability of human-written fact-checking articles, which is not always the case. This demo introduces SAFE, an argument-based framework designed to enhance both fact-checking and its justification. Specifically, SAFE offers three key features: i) producing argument-structured summaries of human-written fact-checking articles, ii) in the absence of human-written articles, generating structured summaries based on evidence retrieved from a corpus through a jointly trained summarization and evidence retrieval system, and iii) assessing the truthfulness of a claim by analyzing the structured summary.
Xiaoou Wang, Elena Cabrio, Serena Villata
IJCAI2
2024 MedMT5: An Open-Source Multilingual Text-to-Text LLM for the Medical Domain
abstract
Research on language technology for the development of medical applications is currently a hot topic in Natural Language Understanding and Generation. Thus, a number of large language models (LLMs) have recently been adapted to the medical domain, so that they can be used as a tool for mediating in human-AI interaction. While these LLMs display competitive performance on automated medical texts benchmarks, they have been pre-trained and evaluated with a focus on a single language (English mostly). This is particularly true of text-to-text models, which typically require large amounts of domain-specific pre-training data, often not easily accessible for many languages. In this paper, we address these shortcomings by compiling, to the best of our knowledge, the largest multilingual corpus for the medical domain in four languages, namely English, French, Italian and Spanish. This new corpus has been used to train Medical mT5, the first open-source text-to-text multilingual model for the medical domain. Additionally, we present two new evaluation benchmarks for all four languages with the aim of facilitating multilingual research in this domain. A comprehensive evaluation shows that Medical mT5 outperforms both encoders and similarly sized text-to-text models for the Spanish, French, and Italian benchmarks, while being competitive with current state-of-the-art LLMs in English.
Iker García-Ferrero, Rodrigo Agerri, Aitziber Atutxa, Elena Cabrio, Iker de la Iglesia, Alberto Lavelli, Bernardo Magnini, Benjamin Molinet, Johanna Ramirez-Romero, German Rigau, Jose Maria Villa-Gonzalez, Serena Villata, Andrea Zaninello
LREC/COLING4
2024 Argument Quality Assessment in the Age of Instruction-Following Large Language Models
abstract
The computational treatment of arguments on controversial issues has been subject to extensive NLP research, due to its envisioned impact on opinion formation, decision making, writing education, and the like. A critical task in any such application is the assessment of an argument’s quality - but it is also particularly challenging. In this position paper, we start from a brief survey of argument quality research, where we identify the diversity of quality notions and the subjectiveness of their perception as the main hurdles towards substantial progress on argument quality assessment. We argue that the capabilities of instruction-following large language models (LLMs) to leverage knowledge across contexts enable a much more reliable assessment. Rather than just fine-tuning LLMs towards leaderboard chasing on assessment tasks, they need to be instructed systematically with argumentation theories and scenarios as well as with ways to solve argument-related problems. We discuss the real-world opportunities and ethical issues emerging thereby.
Henning Wachsmuth, Gabriella Lapesa, Elena Cabrio, Anne Lauscher, Joonsuk Park, Eva Maria Vecchi, Serena Villata, Timon Ziegenbein
LREC/COLING3
2024 ANTIDOTE: ArgumeNtaTIon-Driven explainable artificial intelligence fOr digiTal mEdicine
abstract
The need for transparent AI systems in sensitive domains like medicine has become key. In this paper we present ANTIDOTE, a software suite proposing different tools for argumentation-driven explainable Artificial Intelligence for digital medicine. Our system offers the following functionalities: multilingual argumentative analysis for the medical domain, explanation extraction and generation of clinical diagnoses, multilingual large language models for the medical domain, and the first multilingual benchmark for medical question-answering. Experimental results demonstrate the efficacy of ANTIDOTE across different tasks, highlighting its potential as an asset in medical research and practice and fostering transparency, which is crucial for informed decision-making in healthcare.
Cristian Cardellino, Theo Alkibiades Collias, Benjamin Molinet, Erwan Hain, Rodrigo Agerri, Serena Villata, Elena Cabrio
ECAI8
2024 PEACE: Providing Explanations and Analysis for Combating Hate Expressions
abstract
The increasing presence of hate speech (HS) on social media poses significant societal challenges. While efforts in the Natural Language Processing community have focused on automating the detection of explicit forms of HS, subtler and indirect expressions often go unnoticed. This demo presents PEACE, a novel tool that, besides detecting if a social media message contains explicit or implicit HS, also generates detailed natural language explanations for such predictions. More specifically, PEACE addresses three main challenging tasks: i) exploring the characteristics of HS messages, ii) predicting hatefulness, and iii) elucidating the reasoning behind system predictions. A REST API is also provided to exploit the tool’s functionalities.
Greta Damo, Nicolás Benjamín Ocampo, Elena Cabrio, Serena Villata
ECAI3
2024 Is Safer Better? The Impact of Guardrails on the Argumentative Strength of LLMs in Hate Speech Countering
abstract
The potential effectiveness of counterspeech as a hate speech mitigation strategy is attracting increasing interest in the NLG research community, particularly towards the task of automatically producing it.However, automatically generated responses often lack the argumentative richness which characterises expert-produced counterspeech.In this work, we focus on two aspects of counterspeech generation to produce more cogent responses.First, by investigating the tension between helpfulness and harmlessness of LLMs, we test whether the presence of safety guardrails hinders the quality of the generations.Secondly, we assess whether attacking a specific component of the hate speech results in a more effective argumentative strategy to fight online hate.By conducting an extensive human and automatic evaluation, we show how the presence of safety guardrails can be detrimental also to a task that inherently aims at fostering positive social interactions.Moreover, our results show that attacking a specific component of the hate speech, and in particular its implicit negative stereotype and its hateful parts, leads to higher-quality generations.Content warning: this paper contains unobfuscated examples some readers may find offensive
Helena Bonaldi, Greta Damo, Nicolás Benjamín Ocampo, Elena Cabrio, Serena Villata, Marco Guerini
EMNLP4
2024 CasiMedicos-Arg: A Medical Question Answering Dataset Annotated with Explanatory Argumentative Structures
abstract
Explaining Artificial Intelligence (AI) decisions is a major challenge nowadays in AI, in particular when applied to sensitive scenarios like medicine and law.However, the need to explain the rationale behind decisions is a main issue also for human-based deliberation as it is important to justify why a certain decision has been taken.Resident medical doctors for instance are required not only to provide a (possibly correct) diagnosis, but also to explain how they reached a certain conclusion.Developing new tools to aid residents to train their explanation skills is therefore a central objective of AI in education.In this paper, we follow this direction, and we present, to the best of our knowledge, the first multilingual dataset for Medical Question Answering where correct and incorrect diagnoses for a clinical case are enriched with a natural language explanation written by doctors.These explanations have been manually annotated with argument components (i.e., premise, claim) and argument relations (i.e., attack, support).The Multilingual CasiMedicosarg dataset consists of 558 clinical cases in four languages (English, Spanish, French, Italian) with explanations, where we annotated 5021 claims, 2313 premises, 2431 support relations, and 1106 attack relations.We conclude by showing how competitive baselines perform over this challenging dataset for the argument mining task.
Ekaterina Sviridova, Anar Yeginbergen, Ainara Estarrona, Elena Cabrio, Serena Villata, Rodrigo Agerri
EMNLP4
2024 Unveiling the Hate: Generating Faithful and Plausible Explanations for Implicit and Subtle Hate Speech Detection
Greta Damo, Nicolás Benjamín Ocampo, Elena Cabrio, Serena Villata
NLDB (1)3
2023 DISPUTool 2.0: A Modular Architecture for Multi-Layer Argumentative Analysis of Political Debates
abstract
Political debates are one of the most salient moments of an election campaign, where candidates are challenged to discuss the main contemporary and historical issues in a country. These debates represent a natural ground for argumentative analysis, which has always been employed to investigate political discourse structure and strategy in philosophy and linguistics. In this paper, we present DISPUTool 2.0, an automated tool which relies on Argument Mining methods to analyse the political debates from the US presidential campaigns to extract argument components (i.e., premise and claim) and relations (i.e., support and attack), and highlight fallacious arguments. DISPUTool 2.0 allows also for the automatic analysis of a piece of a debate proposed by the user to identify and classify the arguments contained in the text. A REST API is provided to exploit the tool's functionalities.
Pierpaolo Goffredo, Elena Cabrio, Serena Villata, Shohreh Haddadan, Jhonatan Torres Sanchez
AAAI2
2023 BiRDy: Bullying Role Detection in Multi-Party Chats
abstract
Recent studies have highlighted that private instant messaging platforms and channels are major media of cyber aggression, especially among teens. Due to the private nature of the verbal exchanges on these media, few studies have addressed the task of hate speech detection in this context. Moreover, the recent release of resources mimicking online aggression situations that may occur among teens on private instant messaging platforms is encouraging the development of solutions aiming at dealing with diversity in digital harassment. In this study, we present BiRDy: a fully Web-based platform performing participant role detection in multi-party chats. Leveraging the pre-trained language model mBERT (multilingual BERT), we release fine-tuned models relying on various contextual window strategies to classify exchanged messages according to the role of involvement in cyberbullying of the authors. Integrating a role scoring function, the proposed pipeline predicts a unique role for each chat participant. In addition, detailed confidence scoring are displayed. Currently, BiRDy publicly releases models for French and Italian.
Anaïs Ollagnier, Elena Cabrio, Serena Villata, Sara Tonelli
AAAI2
2023 An In-depth Analysis of Implicit and Subtle Hate Speech Messages
abstract
The research carried out so far in detecting abusive content in social media has primarily focused on overt forms of hate speech.While explicit hate speech (HS) is more easily identifiable by recognizing hateful words, messages containing linguistically subtle and implicit forms of HS (as circumlocution, metaphors and sarcasm) constitute a real challenge for automatic systems.While the sneaky and tricky nature of subtle messages might be perceived as less hurtful with respect to the same content expressed clearly, such abuse is at least as harmful as overt abuse.In this paper, we first provide an in-depth and systematic analysis of 7 standard benchmarks for HS detection, relying on a fine-grained and linguistically-grounded definition of implicit and subtle messages.Then, we experiment with state-of-the-art neural network architectures on two supervised tasks, namely implicit HS and subtle HS message classification.We show that while such models perform satisfactory on explicit messages, they fail to detect implicit and subtle content, highlighting the fact that HS detection is not a solved problem and deserves further investigation.
Nicolás Benjamín Ocampo, Ekaterina Sviridova, Elena Cabrio, Serena Villata
EACL3
2023 Argument-based Detection and Classification of Fallacies in Political Debates
abstract
Fallacies are arguments that employ faulty reasoning.Given their persuasive and seemingly valid nature, fallacious arguments are often used in political debates.Employing these misleading arguments in politics can have detrimental consequences for society, since they can lead to inaccurate conclusions and invalid inferences from the public opinion and the policymakers.Automatically detecting and classifying fallacious arguments represents therefore a crucial challenge to limit the spread of misleading or manipulative claims and promote a more informed and healthier political discourse.Our contribution to address this challenging task is twofold.First, we extend the ElecDeb60To16 dataset of U.S. presidential debates annotated with fallacious arguments, by incorporating the most recent Trump-Biden presidential debate.We include updated tokenlevel annotations, incorporating argumentative components (i.e., claims and premises), the relations between these components (i.e., support and attack), and six categories of fallacious arguments (i.e., Ad Hominem, Appeal to Authority, Appeal to Emotion, False Cause, Slippery Slope, and Slogans).Second, we perform the twofold task of fallacious argument detection and classification by defining neural network architectures based on Transformers models, combining text, argumentative features, and engineered features.Our results show the advantages of complementing transformer-generated text representations with non-textual features.
Pierpaolo Goffredo, Mariana Espinoza, Serena Villata, Elena Cabrio
EMNLP4
2023 Natural Language Explanatory Arguments for Correct and Incorrect Diagnoses of Clinical Cases
abstract
International audience
Santiago Marro, Benjamin Molinet, Elena Cabrio, Serena Villata
ICAART (1)3
2023 Argument and Counter-Argument Generation: A Critical Survey
Xiaoou Wang, Elena Cabrio, Serena Villata
NLDB2
2022 Fallacious Argument Classification in Political Debates
abstract
Fallacies play a prominent role in argumentation since antiquity due to their contribution to argumentation in critical thinking education. Their role is even more crucial nowadays as contemporary argumentation technologies face challenging tasks as misleading and manipulative information detection in news articles and political discourse, and counter-narrative generation. Despite some work in this direction, the issue of classifying arguments as being fallacious largely remains a challenging and an unsolved task. Our contribution is twofold: first, we present a novel annotated resource of 31 political debates from the U.S. Presidential Campaigns, where we annotated six main categories of fallacious arguments (i.e., ad hominem, appeal to authority, appeal to emotion, false cause, slogan, slippery slope) leading to 1628 annotated fallacious arguments; second, we tackle this novel task of fallacious argument classification and we define a neural architecture based on transformers outperforming state-of-the-art results and standard baselines. Our results show the important role played by argument components and relations in this task.
Pierpaolo Goffredo, Shohreh Haddadan, Vorakit Vorakitphan, Elena Cabrio, Serena Villata
IJCAI4
2022 ACTA 2.0: A Modular Architecture for Multi-Layer Argumentative Analysis of Clinical Trials
abstract
Evidence-based medicine aims at making decisions about the care of individual patients based on the explicit use of the best available evidence in the patient clinical history and the medical literature results. Argumentation represents a natural way of addressing this task by (i) identifying evidence and claims in text, and (ii) reasoning upon the extracted arguments and their relations to make a decision. ACTA 2.0 is an automated tool which relies on Argument Mining methods to analyse the abstracts of clinical trials to extract argument components and relations to support evidence-based clinical decision making. ACTA 2.0 allows also for the identification of PICO (Patient, Intervention, Comparison, Outcome) elements, and the analysis of the effects of an intervention on the outcomes of the study. A REST API is also provided to exploit the tool’s functionalities.
Benjamin Molinet, Santiago Marro, Elena Cabrio, Serena Villata, Tobias Mayer 0002
IJCAI3
2022 CyberAgressionAdo-v1: a Dataset of Annotated Online Aggressions in French Collected through a Role-playing Game
abstract
Over the past decades, the number of episodes of cyber aggression occurring online has grown substantially, especially among teens. Most solutions investigated by the NLP community to curb such online abusive behaviors consist of supervised approaches relying on annotated data extracted from social media. However, recent studies have highlighted that private instant messaging platforms are major mediums of cyber aggression among teens. As such interactions remain invisible due to the app privacy policies, very few datasets collecting aggressive conversations are available for the computational analysis of language. In order to overcome this limitation, in this paper we present the CyberAgressionAdo-V1 dataset, containing aggressive multiparty chats in French collected through a role-playing game in high-schools, and annotated at different layers. We describe the data collection and annotation phases, carried out in the context of a EU and a national research projects, and provide insightful analysis on the different types of aggression and verbal abuse depending on the targeted victims (individuals or communities) emerging from the collected data.
Anaïs Ollagnier, Elena Cabrio, Serena Villata, Catherine Blaya
LREC2
2022 Lyrics segmentation via bimodal text-audio representation
abstract
Abstract Song lyrics contain repeated patterns that have been proven to facilitate automated lyrics segmentation, with the final goal of detecting the building blocks (e.g., chorus, verse) of a song text. Our contribution in this article is twofold. First, we introduce a convolutional neural network (CNN)-based model that learns to segment the lyrics based on their repetitive text structure. We experiment with novel features to reveal different kinds of repetitions in the lyrics, for instance based on phonetical and syntactical properties. Second, using a novel corpus where the song text is synchronized to the audio of the song, we show that the text and audio modalities capture complementary structure of the lyrics and that combining both is beneficial for lyrics segmentation performance. For the purely text-based lyrics segmentation on a dataset of 103k lyrics, we achieve an F-score of 67.4%, improving on the state of the art (59.2% F-score). On the synchronized text–audio dataset of 4.8k songs, we show that the additional audio features improve segmentation performance to 75.3% F-score, significantly outperforming the purely text-based approaches.
Michael Fell, Yaroslav Nechaev, Gabriel Meseguer-Brocal, Elena Cabrio, Fabien Gandon, Geoffroy Peeters
Nat. Lang. Eng.4
2021 The WASABI Dataset: Cultural, Lyrics and Audio Analysis Metadata About 2 Million Popular Commercially Released Songs
Michel Buffa, Elena Cabrio, Michael Fell, Fabien Gandon, Alain Giboin, Romain Hennequin, Franck Michel, Johan Pauwels, Guillaume Pellerin, Maroua Tikat, Marco Winckler
ESWC2
2021 Enhancing evidence-based medicine with natural language argumentative analysis of clinical trials
Tobias Mayer 0002, Santiago Marro, Elena Cabrio, Serena Villata
Artif. Intell. Medicine3
2020 Regrexit or not Regrexit: Aspect-based Sentiment Analysis in Polarized Contexts
abstract
Emotion analysis in polarized contexts represents a challenge for Natural Language Processing modeling.As a step in the aforementioned direction, we present a methodology to extend the task of Aspect-based Sentiment Analysis (ABSA) toward the affect and emotion representation in polarized settings.In particular, we adopt the three-dimensional model of affect based on Valence, Arousal, and Dominance (VAD).We then present a Brexit scenario that proves how affect varies toward the same aspect when politically polarized stances are presented.Our approach captures aspect-based polarization from newspapers regarding the Brexit scenario of 1.2m entities at sentence-level.We demonstrate how basic constituents of emotions can be mapped to the VAD model, along with their interactions respecting the polarized context in ABSA settings using biased key-concepts (e.g., "stop Brexit" vs. "support Brexit").Quite intriguingly, the framework achieves to produce coherent aspect evidences of Brexit's stance from key-concepts, showing that VAD influence the support and opposition aspects.
Vorakit Vorakitphan, Marco Guerini, Elena Cabrio, Serena Villata
COLING3
2020 Generating Adversarial Examples for Topic-Dependent Argument Classification
abstract
In the last years, several empirical approaches have been proposed to tackle argument mining tasks, e.g., argument classification, relation prediction, argument synthesis. These approaches rely more and more on language models (e.g., BERT) to boost their performance. However, these language models require a lot of training data, and size is often a drawback of the available argument mining data sets. The goal of this paper is to assess the robustness of these language models for the argument classification task. More precisely, the aim of the current work is twofold: first, we generate adversarial examples addressing linguistic perturbations in the original sentences, and second, we improve the robustness of argument classification models using adversarial training. Two empirical evaluations are addressed relying on standard datasets for AM tasks, whilst the generated adversarial examples are qualitatively evaluated through a user study. Results prove the robust-ness of BERT for the argument classification task, yet highlighting that it is not invulnerable to simple linguistic perturbations in the input data.
Tobias Mayer 0002, Santiago Marro, Elena Cabrio, Serena Villata
COMMA3
2020 Dataset Independent Baselines for Relation Prediction in Argument Mining
abstract
Argument(ation) Mining (AM) is the research area which aims at extracting argument components and predicting argumentative relations (i.e., support and attack) from text. In particular, numerous approaches have been proposed in the literature to predict the relations holding between arguments, and application-specific annotated resources were built for this purpose. Despite the fact that these resources were created to experiment on the same task, the definition of a single relation prediction method to be successfully applied to a significant portion of these datasets is an open research problem in AM. This means that none of the methods proposed in the literature can be easily ported from one resource to another. In this paper, we address this problem by proposing a set of dataset independent strong neural baselines which obtain homogeneous results on all the datasets proposed in the literature for the argumentative relation prediction task in AM. Thus, our baselines can be employed by the AM community to compare more effectively how well a method performs on the argumentative relation prediction task.
Oana Cocarascu, Elena Cabrio, Serena Villata, Francesca Toni
COMMA2
2020 Transformer-Based Argument Mining for Healthcare Applications
abstract
Argument(ation) Mining (AM) typically aims at identifying argumentative components in text and predicting the relations among them. Evidence-based decision making in the health-care domain targets at supporting clinicians in their deliberation process to establish the best course of action for the case under evaluation. Although the reasoning stage of this kind of frameworks received considerable attention, little effort has been devoted to the mining stage. We extended an existing dataset by annotating 500 abstracts of Randomized Controlled Trials (RCT) from the MEDLINE database, leading to a dataset of 4198 argument components and 2601 argument relations on different diseases (i.e., neoplasm, glau-coma, hepatitis, diabetes, hypertension). We propose a complete argument mining pipeline for RCTs, classifying argument components as evidence and claims, and predicting the relation, i.e., attack or support , holding between those argument components. We experiment with deep bidirectional transformers in combination with different neural architectures (i.e., LSTM, GRU and CRF) and obtain a macro F1-score of .87 for component detection and .68 for relation prediction , outperforming current state-of-the-art end-to-end AM systems.
Tobias Mayer 0002, Elena Cabrio, Serena Villata
ECAI2
2020 Love Me, Love Me, Say (and Write!) that You Love Me: Enriching the WASABI Song Corpus with Lyrics Annotations
abstract
We present the WASABI Song Corpus, a large corpus of songs enriched with metadata extracted from music databases on the Web, and resulting from the processing of song lyrics and from audio analysis. More specifically, given that lyrics encode an important part of the semantics of a song, we focus here on the description of the methods we proposed to extract relevant information from the lyrics, as their structure segmentation, their topic, the explicitness of the lyrics content, the salient passages of a song and the emotions conveyed. The creation of the resource is still ongoing: so far, the corpus contains 1.73M songs with lyrics (1.41M unique lyrics) annotated at different levels with the output of the above mentioned methods. Such corpus labels and the provided methods can be exploited by music search engines and music professionals (e.g. journalists, radio presenters) to better handle large collections of lyrics, allowing an intelligent browsing, categorization and segmentation recommendation of songs.
Michael Fell, Elena Cabrio, Elmahdi Korfed, Michel Buffa, Fabien Gandon
LREC2
2020 Covid-on-the-Web: Knowledge Graph and Services to Advance COVID-19 Research
Franck Michel, Fabien Gandon, Valentin Ah-Kane, Anna Bobasheva, Elena Cabrio, Olivier Corby, Raphaël Gazzotti, Alain Giboin, Santiago Marro, Tobias Mayer 0002, Mathieu Simon, Serena Villata, Marco Winckler
ISWC (2)5
2020 A Multilingual Evaluation for Online Hate Speech Detection
abstract
The increasing popularity of social media platforms such as Twitter and Facebook has led to a rise in the presence of hate and aggressive speech on these platforms. Despite the number of approaches recently proposed in the Natural Language Processing research area for detecting these forms of abusive language, the issue of identifying hate speech at scale is still an unsolved problem. In this article, we propose a robust neural architecture that is shown to perform in a satisfactory way across different languages; namely, English, Italian, and German. We address an extensive analysis of the obtained experimental results over the three languages to gain a better understanding of the contribution of the different components employed in the system, both from the architecture point of view (i.e., Long Short Term Memory, Gated Recurrent Unit, and bidirectional Long Short Term Memory) and from the feature selection point of view (i.e., ngrams, social network–specific features, emotion lexica, emojis, word embeddings). To address such in-depth analysis, we use three freely available datasets for hate speech detection on social media in English, Italian, and German.
Michele Corazza, Stefano Menini, Elena Cabrio, Sara Tonelli, Serena Villata
ACM Trans. Internet Techn.3
2019 Yes, we can! Mining Arguments in 50 Years of US Presidential Campaign Debates
abstract
Political debates offer a rare opportunity for citizens to compare the candidates' positions on the most controversial topics of the campaign. Thus they represent a natural application scenario for Argument Mining. As existing research lacks solid empirical investigation of the typology of argument components in political debates, we fill this gap by proposing an Argument Mining approach to political debates. We address this task in an empirical manner by annotating 39 political debates from the last 50 years of US presidential campaigns, creating a new corpus of 29k argument components, labeled as premises and claims. We then propose two tasks: (1) identifying the argumentative components in such debates, and (2) classifying them as premises and claims. We show that feature-rich SVM learners and Neural Network architectures outperform standard baselines in Argument Mining over such complex data. We release the new corpus USElecDeb60To16 and the accompanying software under free licenses to the research community.
Shohreh Haddadan, Elena Cabrio, Serena Villata
ACL (1)2
2019 ACTA A Tool for Argumentative Clinical Trial Analysis
abstract
Argumentative analysis of textual documents of various nature (e.g., persuasive essays, online discussion blogs, scientific articles) allows to detect the main argumentative components (i.e., premises and claims) present in the text and to predict whether these components are connected to each other by argumentative relations (e.g., support and attack), leading to the identification of (possibly complex) argumentative structures. Given the importance of argument-based decision making in medicine, in this demo paper we introduce ACTA, a tool for automating the argumentative analysis of clinical trials. The tool is designed to support doctors and clinicians in identifying the document(s) of interest about a certain disease, and in analyzing the main argumentative content and PICO elements.
Tobias Mayer 0002, Elena Cabrio, Serena Villata
IJCAI2
2019 DISPUTool - A tool for the Argumentative Analysis of Political Debates
abstract
Political debates are the means used by political candidates to put forward and justify their positions in front of the electors with respect to the issues at stake. Argument mining is a novel research area in Artificial Intelligence, aiming at analyzing discourse on the pragmatics level and applying a certain argumentation theory to model and automatically analyze textual data. In this paper, we present DISPUTool, a tool designed to ease the work of historians and social science scholars in analyzing the argumentative content of political speeches. More precisely, DISPUTool allows to explore and automatically identify argumentative components over the 39 political debates from the last 50 years of US presidential campaigns (1960-2016).
Shohreh Haddadan, Elena Cabrio, Serena Villata
IJCAI2
2018 Never Retreat, Never Retract: Argumentation Analysis for Political Speeches
abstract
In this work, we apply argumentation mining techniques, in particular relation prediction, to study political speeches in monological form, where there is no direct interaction between opponents. We argue that this kind of technique can effectively support researchers in history, social and political sciences, which must deal with an increasing amount of data in digital form and need ways to automatically extract and analyse argumentation patterns. We test and discuss our approach based on the analysis of documents issued by R. Nixon and J. F. Kennedy during 1960 presidential campaign. We rely on a supervised classifier to predict argument relations (i.e., support and attack), obtaining an accuracy of 0.72 on a dataset of 1,462 argument pairs. The application of argument mining to such data allows not only to highlight the main points of agreement and disagreement between the candidates' arguments over the campaign issues such as Cuba, disarmament and health-care, but also an in-depth argumentative analysis of the respective viewpoints on these topics.
Stefano Menini, Elena Cabrio, Sara Tonelli, Serena Villata
AAAI2
2018 Lyrics Segmentation: Textual Macrostructure Detection using Convolutions
abstract
Lyrics contain repeated patterns that are correlated with the repetitions found in the music they accompany. Repetitions in song texts have been shown to enable lyrics segmentation – a fundamental prerequisite of automatically detecting the building blocks (e.g. chorus, verse) of a song text. In this article we improve on the state-of-the-art in lyrics segmentation by applying a convolutional neural network to the task, and experiment with novel features as a step towards deeper macrostructure detection of lyrics.
Michael Fell, Yaroslav Nechaev, Elena Cabrio, Fabien Gandon
COLING3
2018 Argument Mining on Clinical Trials
abstract
Argument-based decision making has been employed to support a variety of reasoning tasks over medical knowledge. These include evidence-based justifications of the effects of treatments, the detection of conflicts in the knowledge base, and the enabling of uncertain and defeasible reasoning in the health-care sector. However, a common limitation of these approaches is that they rely on structured input information. Recent advances in argument mining have shown increasingly accurate results in detecting argument components and predicting their relations from unstructured, natural language texts. In this study, we discuss evidence and claim detection from Randomized Clinical Trials. To this end, we create a new annotated dataset about four different diseases (glaucoma, diabetes, hepatitis B, and hypertension), containing 976 argument components (697 containing evidence, 279 claims). Empirical results are promising, and show the portability of the proposed approach over different branches of medicine.
Tobias Mayer 0002, Elena Cabrio, Marco Lippi 0001, Paolo Torroni, Serena Villata
COMMA2
2018 Five Years of Argument Mining: a Data-driven Analysis
abstract
Argument mining is the research area aiming at extracting natural language arguments and their relations from text, with the final goal of providing machine-processable structured data for computational models of argument. This research topic has started to attract the attention of a small community of researchers around 2014, and it is nowadays counted as one of the most promising research areas in Artificial Intelligence in terms of growing of the community, funded projects, and involvement of companies. In this paper, we present the argument mining tasks, and we discuss the obtained results in the area from a data-driven perspective. An open discussion highlights the main weaknesses suffered by the existing work in the literature, and proposes open challenges to be faced in the future.
Elena Cabrio, Serena Villata
IJCAI1
2017 Argument Mining on Twitter: Arguments, Facts and Sources
abstract
Social media collect and spread on the Web personal opinions, facts, fake news and all kind of information users may be interested in.Applying argument mining methods to such heterogeneous data sources is a challenging open research issue, in particular considering the peculiarities of the language used to write textual messages on social media.In addition, new issues emerge when dealing with arguments posted on such platforms, such as the need to make a distinction between personal opinions and actual facts, and to detect the source disseminating information about such facts to allow for provenance verification.In this paper, we apply supervised classification to identify arguments on Twitter, and we present two new tasks for argument mining, namely facts recognition and source identification.We study the feasibility of the approaches proposed to address these tasks on a set of tweets related to the Grexit and Brexit news topics.
Mihai Dusmanu, Elena Cabrio, Serena Villata
EMNLP2
2017 Semantic web-mining and deep vision for lifelong object discovery
abstract
Autonomous robots that are to assist humans in their daily lives must recognize and understand the meaning of objects in their environment. However, the open nature of the world means robots must be able to learn and extend their knowledge about previously unknown objects on-line. In this work we investigate the problem of unknown object hypotheses generation, and employ a semantic Web-mining framework along with deep-learning-based object detectors. This allows us to make use of both visual and semantic features in combined hypotheses generation. Experiments on data from mobile robots in real world application deployments show that this combination improves performance over the use of either method in isolation.
Jay Young, Lars Kunze, Valerio Basile, Elena Cabrio, Nick Hawes, Barbara Caputo
ICRA4
2016 Tweeties Squabbling: Positive and Negative Results in Applying Argument Mining on Social Media
abstract
The problem of understanding the stream of messages exchanged on social media such as Facebook and Twitter is becoming a major challenge for automated systems. The tremendous amount of data exchanged on these platforms as well as the specific form of language adopted by social media users constitute a new challenging context for existing argument mining techniques. In this paper, we describe an ongoing work towards the creation of a complete argument mining pipeline over Twitter messages: (i) we identify which tweets can be considered as arguments and which cannot, (ii) over the set of tweet-arguments, we group them by topic, and (iii) we predict whether such tweets support or attack each other. The final goal is to compute the set of tweets which are widely recognized as accepted, and the different (possibly conflicting) viewpoints that emerge on a topic, given a stream of messages.
Tom Bosc, Elena Cabrio, Serena Villata
COMMA2
2016 Towards Lifelong Object Learning by Integrating Situated Robot Perception and Semantic Web Mining
abstract
Autonomous robots that are to assist humans in their daily lives are required, among other things, to recognize and understand the meaning of task-related objects. However, given an open-ended set of tasks, the set of everyday objects that robots will encounter during their lifetime is not foreseeable. That is, robots have to learn and extend their knowledge about previously unknown objects on-the-job. Our approach automatically acquires parts of this knowledge (e.g., the class of an object and its typical location) in form of ranked hypotheses from the Semantic Web using contextual information extracted from observations and experiences made by robots. Thus, by integrating situated robot perception and Semantic Web mining, robots can continuously extend their object knowledge beyond perceptual models which allows them to reason about task-related objects, e.g., when searching for them, robots can infer the most likely object locations. An evaluation of the integrated system on long-term data from real office observations, demonstrates that generated hypotheses can effectively constrain the meaning of objects. Hence, we believe that the proposed system can be an essential component in a lifelong learning framework which acquires knowledge about objects from real world observations.
Jay Young, Valerio Basile, Lars Kunze, Elena Cabrio, Nick Hawes
ECAI4
2016 Populating a Knowledge Base with Object-Location Relations Using Distributional Semantics
Valerio Basile, Soufian Jebbara, Elena Cabrio, Philipp Cimiano
EKAW3
2016 Enriching a Small Artwork Collection Through Semantic Linking
Mauro Dragoni, Elena Cabrio, Sara Tonelli, Serena Villata
ESWC2
2016 Abstract Dialectical Frameworks for Text Exploration
Elena Cabrio, Serena Villata
ICAART (2)1
2016 DART: a Dataset of Arguments and their Relations on Twitter
Tom Bosc, Elena Cabrio, Serena Villata
LREC2
2016 Adapting Semantic Spreading Activation to Entity Linking in Text
Farhad Nooralahzadeh, Cédric Lopez, Elena Cabrio, Fabien Gandon, Frédérique Segond
NLDB3
2015 Information Extraction with Active Learning: A Case Study in Legal Text
Cristian Cardellino, Serena Villata, Laura Alonso Alemany, Elena Cabrio
CICLing (2)4
2015 Emotions in Argumentation: an Empirical Evaluation
M. Sahbi Benlamine, Maher Chaouachi, Serena Villata, Elena Cabrio, Claude Frasson, Fabien Gandon
IJCAI4
2015 Improvements in Information Extraction in Legal Text by Active Learning
abstract
Managing licensing information and data rights is becoming a crucial issue in the Linked (Open) Data scenario. An open problem in this scenario is how to associate machine-readable licenses specifications to the data, so that automated approaches to treat such information can be fruitfully exploited to avoid data misuse. This means that we need a way to automatically extract from a natural language document specifying a certain license a machine-readable description of the terms of use and reuse identified in such license.
Cristian Cardellino, Laura Alonso Alemany, Serena Villata, Elena Cabrio
JURIX4
2014 NoDE: A Benchmark of Natural Language Arguments
abstract
In the latest years, natural models of argumentation and argument mining are becoming more and more important topics in the argumentation community. Given this tendency, there is the need to produce standard datasets on which natural language approaches to argumentation can be evaluated. In this paper, we present NoDE, a benchmark of natural language arguments composed of three datasets, built from different textual sources and annotated highlighting positive and negative connections between arguments.
Elena Cabrio, Serena Villata
COMMA1
2014 These Are Your Rights - A Natural Language Processing Approach to Automated RDF Licenses Generation
Elena Cabrio, Alessio Palmero Aprosio, Serena Villata
ESWC1
2014 Classifying Inconsistencies in DBpedia Language Specific Chapters
Elena Cabrio, Serena Villata, Fabien Gandon
LREC1
2013 A Support Framework for Argumentative Discussions Management in the Web
Elena Cabrio, Serena Villata, Fabien Gandon
ESWC1
2012 Hunting for Entailing Pairs in the Penn Discourse Treebank
Sara Tonelli, Elena Cabrio
COLING2
2012 Generating Abstract Arguments: A Natural Language Approach
abstract
Many argumentation tools have been proposed nowadays to support the users in on-line social discussions. However, the main drawback of these tools is that they do not cope with the automatic generation of the arguments from the natural language discussions of the users. In this paper, we propose to use a technique from computational linguistics, namely textual entailment, to generate in an automatic way the abstract arguments from the dialogues. The abstract arguments as well as their relationships are then structured in an argumentation graph to evaluate the dialogue as a whole. The success criteria of the proposed approach is that it is able to represent the dynamics of the dialogues among users allowing to find the use of argumentation natural enough to be really adopted.
Elena Cabrio, Serena Villata
COMMA1
2010 Building Textual Entailment Specialized Data Sets: a Methodology for Isolating Linguistic Phenomena Relevant to Inference
Luisa Bentivogli, Elena Cabrio, Ido Dagan, Danilo Giampiccolo, Medea Lo Leggio, Bernardo Magnini
LREC2
2009 Specialized Entailment Engines: Approaching Linguistic Aspects of Textual Entailment
Elena Cabrio
NLDB1
2008 The QALL-ME Benchmark: a Multilingual Resource of Annotated Spoken Requests for Question Answering
Elena Cabrio, Milen Kouylekov, Bernardo Magnini, Matteo Negri, Laura Hasler, Constantin Orasan, David Tomás 0001, José Luis Vicedo González, Günter Neumann, Corinna Weber
LREC1