EDBT 2026 Demo / reviewers in the wild / expert
Valerio Basile
dblp:86/11425
· DBLP profile ↗
33ranked-venue papers
6as first author
19since 2021 · last 2026
0000-0001-8110-6832ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 26 · 5 first-author · 16 since 2021Databases, data management, data science and information retrieval · 8 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bootstrapping NLP for Sakha: Named Entity Recognition and Sentiment Analysis in an Extremely Low-Resource Setting
Mariia Everstova, Nikolai Efimov, Valerio Basile |
LREC | 3 |
| 2026 | Multilingual Structured Sentiment Analysis for Environmental Sustainability
Muhammad Okky Ibrohim, Tommaso Caselli, Cristina Bosco, Valerio Basile |
LREC | 4 |
| 2025 | PERSEVAL: A Framework for Perspectivist Classification EvaluationabstractSoda Marem Lo, Silvia Casola, Erhan Sezerer, Valerio Basile, Franco Sansonetti, Antonio Uva, Davide Bernardi. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Soda Marem Lo, Silvia Casola, Erhan Sezerer, Valerio Basile, Franco Sansonetti, Antonio Uva 0001, Davide Bernardi |
EMNLP | 4 |
| 2025 | LiITA: a Knowledge Base of Interoperable Resources for ItalianabstractThis paper describes the LiITA Knowledge Base of interoperable linguistic resources for Italian.By adhering to the Linked Open Data principles, LiITA ensures and facilitates interoperability between distributed resources. The paper outlines the lemma-centered architecture of the Knowledge Base and details its core component: the Lemma Bank, a collection of Italian lemmas designed to interlink distributed lexical and textual resources. Eleonora Litta Modignani Picozzi, Marco Passarotti, Valerio Basile, Cristina Bosco, Andrea Di Fabio, Paolo Brasolin |
LDK | 3 |
| 2024 | MultiPICo: Multilingual Perspectivist Irony CorpusabstractSilvia Casola, Simona Frenda, Soda Marem Lo, Erhan Sezerer, Antonio Uva, Valerio Basile, Cristina Bosco, Alessandro Pedrani, Chiara Rubagotti, Viviana Patti, Davide Bernardi. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Silvia Casola, Simona Frenda, Soda Marem Lo, Erhan Sezerer, Antonio Uva 0001, Valerio Basile, Cristina Bosco, Alessandro Pedrani, Chiara Rubagotti, Viviana Patti, Davide Bernardi |
ACL (1) | 6 |
| 2024 | Label Augmentation for Zero-Shot Hierarchical Text ClassificationabstractHierarchical Text Classification poses the difficult challenge of classifying documents into multiple labels organized in a hierarchy.The vast majority of works aimed to address this problem relies on supervised methods which are difficult to implement due to the scarcity of labeled data in many real world applications.This paper focuses on strict Zero-Shot Classification, the setting in which the system lacks both labeled instances and training data.We propose a novel approach that uses a Large Language Model to augment the deepest layer of the labels hierarchy in order to enhance its specificity.We achieve this by generating semantically relevant labels as children connected to the existing branches, creating a deeper taxonomy that better overlaps with the input texts.We leverage the enriched hierarchy to perform Zero-Shot Hierarchical Classification by using the Upward score Propagation technique.We test our method on four public datasets, obtaining new state-of-the art results on three of them.We introduce two cosine similarity-based metrics to quantify the density and granularity of a label taxonomy and we show a strong correlation between the metric values and the classification performance of our method on the datasets. Lorenzo Paletto, Valerio Basile, Roberto Esposito |
ACL (1) | 2 |
| 2024 | QUEEREOTYPES: A Multi-Source Italian Corpus of Stereotypes towards LGBTQIA+ Community MembersabstractThe paper describes a dataset composed of two sub-corpora from two different sources in Italian. The QUEEREOTYPES corpus includes social media texts regarding LGBTQIA+ individuals, behaviors, ideology and events. The texts were collected from Facebook and Twitter in 2018 and were annotated for the presence of stereotypes, and orthogonal dimensions (such as hate speech, aggressiveness, offensiveness, and irony in one sub-corpus, and stance in the other). The resource was developed by Natural Language Processing researchers together with activists from an Italian LGBTQIA+ not-for-profit organization. The creation of the dataset allows the NLP community to study stereotypes against marginalized groups, individuals and, ultimately, to develop proper tools and measures to reduce the online spread of such stereotypes. A test for the robustness of the language resource has been performed by means of 5-fold cross-validation experiments. Finally, text classification experiments have been carried out with a fine-tuned version of AlBERTo (a BERT-based model pre-trained on Italian tweets) and mBERT, obtaining good results on the task of stereotype detection, suggesting that stereotypes towards different targets might share common traits. Alessandra Teresa Cignarella, Manuela Sanguinetti, Simona Frenda, Andrea Marra, Cristina Bosco, Valerio Basile |
LREC/COLING | 6 |
| 2024 | Capturing Perspectives of Crowdsourced Annotators in Subjective Learning TasksabstractNegar Mokhberian, Myrl Marmarelis, Frederic Hopp, Valerio Basile, Fred Morstatter, Kristina Lerman. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Negar Mokhberian, Myrl G. Marmarelis, Frederic R. Hopp, Valerio Basile, Fred Morstatter, Kristina Lerman |
NAACL-HLT | 4 |
| 2023 | Toward a Perspectivist Turn in Ground Truthing for Predictive ComputingabstractMost current Artificial Intelligence applications are based on supervised Machine Learning (ML), which ultimately grounds on data annotated by small teams of experts or large ensemble of volunteers. The annotation process is often performed in terms of a majority vote, however this has been proved to be often problematic by recent evaluation studies. In this article, we describe and advocate for a different paradigm, which we call perspectivism: this counters the removal of disagreement and, consequently, the assumption of correctness of traditionally aggregated gold-standard datasets, and proposes the adoption of methods that preserve divergence of opinions and integrate multiple perspectives in the ground truthing process of ML development. Drawing on previous works which inspired it, mainly from the crowdsourcing and multi-rater labeling settings, we survey the state-of-the-art and describe the potential of our proposal for not only the more subjective tasks (e.g. those related to human language) but also those tasks commonly understood as objective (e.g. medical decision making). We present the main benefits of adopting a perspectivist stance in ML, as well as possible disadvantages, and various ways in which such a stance can be implemented in practice. Finally, we share a set of recommendations and outline a research agenda to advance the perspectivist stance in ML. Federico Cabitza, Andrea Campagner, Valerio Basile |
AAAI | 3 |
| 2023 | EPIC: Multi-Perspective Annotation of a Corpus of IronyabstractSimona Frenda, Alessandro Pedrani, Valerio Basile, Soda Marem Lo, Alessandra Teresa Cignarella, Raffaella Panizzon, Cristina Marco, Bianca Scarlini, Viviana Patti, Cristina Bosco, Davide Bernardi. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Simona Frenda, Alessandro Pedrani, Valerio Basile, Soda Marem Lo, Alessandra Teresa Cignarella, Raffaella Panizzon, Cristina Marco, Bianca Scarlini, Viviana Patti, Cristina Bosco, Davide Bernardi |
ACL (1) | 3 |
| 2023 | Confidence-based Ensembling of Perspective-aware ModelsabstractSilvia Casola, Soda Lo, Valerio Basile, Simona Frenda, Alessandra Cignarella, Viviana Patti, Cristina Bosco. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Silvia Casola, Soda Marem Lo, Valerio Basile, Simona Frenda, Alessandra Teresa Cignarella, Viviana Patti, Cristina Bosco |
EMNLP | 3 |
| 2023 | Towards multidomain and multilingual abusive language detection: a surveyabstractAbstract Abusive language is an important issue in online communication across different platforms and languages. Having a robust model to detect abusive instances automatically is a prominent challenge. Several studies have been proposed to deal with this vital issue by modeling this task in the cross-domain and cross-lingual setting. This paper outlines and describes the current state of this research direction, providing an overview of previous studies, including the available datasets and approaches employed in both cross-domain and cross-lingual settings. This study also outlines several challenges and open problems of this area, providing insights and a useful roadmap for future work. Endang Wahyu Pamungkas, Valerio Basile, Viviana Patti |
Pers. Ubiquitous Comput. | 2 |
| 2022 | Italian NLP for Everyone: Resources and Models from EVALITA to the European Language GridabstractThe European Language Grid enables researchers and practitioners to easily distribute and use NLP resources and models, such as corpora and classifiers. We describe in this paper how, during the course of our EVALITA4ELG project, we have integrated datasets and systems for the Italian language. We show how easy it is to use the integrated systems, and demonstrate in case studies how seamless the application of the platform is, providing Italian NLP for everyone. Valerio Basile, Cristina Bosco, Michael Fell, Viviana Patti, Rossella Varvara |
LREC | 1 |
| 2022 | APPReddit: a Corpus of Reddit Posts Annotated for AppraisalabstractDespite the large number of computational resources for emotion recognition, there is a lack of data sets relying on appraisal models. According to Appraisal theories, emotions are the outcome of a multi-dimensional evaluation of events. In this paper, we present APPReddit, the first corpus of non-experimental data annotated according to this theory. After describing its development, we compare our resource with enISEAR, a corpus of events created in an experimental setting and annotated for appraisal. Results show that the two corpora can be mapped notwithstanding different typologies of data and annotations schemes. A SVM model trained on APPReddit predicts four appraisal dimensions without significant loss. Merging both corpora in a single training set increases the prediction of 3 out of 4 dimensions. Such findings pave the way to a better performing classification model for appraisal prediction. Marco Stranisci, Simona Frenda, Eleonora Ceccaldi, Valerio Basile, Rossana Damiano, Viviana Patti |
LREC | 4 |
| 2022 | Automatically Computing Connotative Shifts of Lexical Items
Valerio Basile, Tommaso Caselli, Anna Koufakou, Viviana Patti |
NLDB | 1 |
| 2022 | The unbearable hurtfulness of sarcasm
Simona Frenda, Alessandra Teresa Cignarella, Valerio Basile, Cristina Bosco, Viviana Patti, Paolo Rosso |
Expert Syst. Appl. | 3 |
| 2022 | A hybrid lexicon-based and neural approach for explainable polarity detection
Marco Polignano, Valerio Basile, Pierpaolo Basile, Giuliano Gabrieli, Marco Vassallo, Cristina Bosco |
Inf. Process. Manag. | 2 |
| 2021 | A joint learning approach with knowledge injection for zero-shot cross-lingual hate speech detection
Endang Wahyu Pamungkas, Valerio Basile, Viviana Patti |
Inf. Process. Manag. | 2 |
| 2021 | Sentiment Polarity Classification at EVALITA: Lessons Learned and Open ChallengesabstractSentiment analysis in social media is a popular task attracting the interest of the research community, also in recent evaluation campaigns of natural language processing tasks in several languages. We report on our experience in the organization of SENTIment POLarity Classification Task (SENTIPOLC), a shared task on sentiment classification of Italian tweets, proposed for the first time in 2014 within the Evalita evaluation campaign. We present the datasets-which include an enriched annotation scheme for dealing with the impact of figurative language on polarity-the evaluation methodology, and discuss the approaches and results of participating systems. We also offer a reflection on the open challenges of state-of-the-art systems for sentiment analysis of microblogging in Italian, as they emerge from a qualitative analysis of misclassified tweets. Finally, we provide an evaluation of the resources we have created, and share the lessons learned by running this task for two consecutive editions. Valerio Basile, Nicole Novielli, Danilo Croce, Francesco Barbieri, Malvina Nissim, Viviana Patti |
IEEE Trans. Affect. Comput. | 1 |
| 2020 | Multilingual Irony Detection with Dependency Syntax and Neural ModelsabstractThis paper presents an in-depth investigation of the effectiveness of dependency-based syntactic features on the irony detection task in a multilingual perspective (English, Spanish, French and Italian).It focuses on the contribution from syntactic knowledge, exploiting linguistic resources where syntax is annotated according to the Universal Dependencies scheme.Three distinct experimental settings are provided.In the first, a variety of syntactic dependency-based features combined with classical machine learning classifiers are explored.In the second scenario, two well-known types of word embeddings are trained on parsed data and tested against gold standard datasets.In the third setting, dependency-based syntactic features are combined into the Multilingual BERT architecture.The results suggest that fine-grained dependency-based syntactic information is informative for the detection of irony. Alessandra Teresa Cignarella, Valerio Basile, Manuela Sanguinetti, Cristina Bosco, Paolo Rosso, Farah Benamara |
COLING | 2 |
| 2020 | Modeling Annotator Perspective and Polarized Opinions to Improve Hate Speech DetectionabstractIn this paper we propose an approach to exploit the fine-grained knowledge expressed by individual human annotators during a hate speech (HS) detection task, before the aggregation of single judgments in a gold standard dataset eliminates non-majority perspectives. We automatically divide the annotators into groups, aiming at grouping them by similar personal characteristics (ethnicity, social background, culture etc.). To serve a multi-lingual perspective, we performed classification experiments on three different Twitter datasets in English and Italian languages. We created different gold standards, one for each group, and trained a state-of-the-art deep learning model on them, showing that supervised models informed by different perspectives on the target phenomena outperform a baseline represented by models trained on fully aggregated data. Finally, we implemented an ensemble approach that combines the single perspective-aware classifiers into an inclusive model. The results show that this strategy further improves the classification performance, especially with a significant boost in the recall of HS prediction. Sohail Akhtar, Valerio Basile, Viviana Patti |
HCOMP | 2 |
| 2020 | I Feel Offended, Don't Be Abusive! Implicit/Explicit Messages in Offensive and Abusive LanguageabstractAbusive language detection is an unsolved and challenging problem for the NLP community. Recent literature suggests various approaches to distinguish between different language phenomena (e.g., hate speech vs. cyberbullying vs. offensive language) and factors (degree of explicitness and target) that may help to classify different abusive language phenomena. There are data sets that annotate the target of abusive messages (i.e.OLID/OffensEval (Zampieri et al., 2019a)). However, there is a lack of data sets that take into account the degree of explicitness. In this paper, we propose annotation guidelines to distinguish between explicit and implicit abuse in English and apply them to OLID/OffensEval. The outcome is a newly created resource, AbuseEval v1.0, which aims to address some of the existing issues in the annotation of offensive and abusive language (e.g., explicitness of the message, presence of a target, need of context, and interaction across different phenomena). Tommaso Caselli, Valerio Basile, Jelena Mitrovic, Inga Kartoziya, Michael Granitzer |
LREC | 2 |
| 2020 | Do You Really Want to Hurt Me? Predicting Abusive Swearing in Social MediaabstractSwearing plays an ubiquitous role in everyday conversations among humans, both in oral and textual communication, and occurs frequently in social media texts, typically featured by informal language and spontaneous writing. Such occurrences can be linked to an abusive context, when they contribute to the expression of hatred and to the abusive effect, causing harm and offense. However, swearing is multifaceted and is often used in casual contexts, also with positive social functions. In this study, we explore the phenomenon of swearing in Twitter conversations, taking the possibility of predicting the abusiveness of a swear word in a tweet context as the main investigation perspective. We developed the Twitter English corpus SWAD (Swear Words Abusiveness Dataset), where abusive swearing is manually annotated at the word level. Our collection consists of 1,511 unique swear words from 1,320 tweets. We developed models to automatically predict abusive swearing, to provide an intrinsic evaluation of SWAD and confirm the robustness of the resource. We also present the results of a glass box ablation study in order to investigate which lexical, syntactic, and affective features are more informative towards the automatic prediction of the function of swearing. Endang Wahyu Pamungkas, Valerio Basile, Viviana Patti |
LREC | 2 |
| 2020 | Misogyny Detection in Twitter: a Multilingual and Cross-Domain Study
Endang Wahyu Pamungkas, Valerio Basile, Viviana Patti |
Inf. Process. Manag. | 2 |
| 2018 | Empirical Analysis of Foundational Distinctions in Linked Open DataabstractThe Web and its Semantic extension (i.e. Linked Open Data) contain open global-scale knowledge and make it available to potentially intelligent machines that want to benefit from it. Nevertheless, most of Linked Open Data lack ontological distinctions and have sparse axiomatisation. For example, distinctions such as whether an entity is inherently a class or an individual, or whether it is a physical object or not, are hardly expressed in the data, although they have been largely studied and formalised by foundational ontologies (e.g. DOLCE, SUMO). These distinctions belong to common sense too, which is relevant for many artificial intelligence tasks such as natural language understanding, scene recognition, and the like. There is a gap between foundational ontologies, that often formalise or are inspired by pre-existing philosophical theories and are developed with a top-down approach, and Linked Open Data that mostly derive from existing databases or crowd-based effort (e.g. DBpedia, Wikidata). We investigate whether machines can learn foundational distinctions over Linked Open Data entities, and if they match common sense. We want to answer questions such as “does the DBpedia entity for dog refer to a class or to an instance?”. We report on a set of experiments based on machine learning and crowdsourcing that show promising results. Luigi Asprino, Valerio Basile, Paolo Ciancarini, Valentina Presutti |
IJCAI | 2 |
| 2017 | Semantic web-mining and deep vision for lifelong object discoveryabstractAutonomous robots that are to assist humans in their daily lives must recognize and understand the meaning of objects in their environment. However, the open nature of the world means robots must be able to learn and extend their knowledge about previously unknown objects on-line. In this work we investigate the problem of unknown object hypotheses generation, and employ a semantic Web-mining framework along with deep-learning-based object detectors. This allows us to make use of both visual and semantic features in combined hypotheses generation. Experiments on data from mobile robots in real world application deployments show that this combination improves performance over the use of either method in isolation. Jay Young, Lars Kunze, Valerio Basile, Elena Cabrio, Nick Hawes, Barbara Caputo |
ICRA | 3 |
| 2016 | Towards Lifelong Object Learning by Integrating Situated Robot Perception and Semantic Web MiningabstractAutonomous robots that are to assist humans in their daily lives are required, among other things, to recognize and understand the meaning of task-related objects. However, given an open-ended set of tasks, the set of everyday objects that robots will encounter during their lifetime is not foreseeable. That is, robots have to learn and extend their knowledge about previously unknown objects on-the-job. Our approach automatically acquires parts of this knowledge (e.g., the class of an object and its typical location) in form of ranked hypotheses from the Semantic Web using contextual information extracted from observations and experiences made by robots. Thus, by integrating situated robot perception and Semantic Web mining, robots can continuously extend their object knowledge beyond perceptual models which allows them to reason about task-related objects, e.g., when searching for them, robots can infer the most likely object locations. An evaluation of the integrated system on long-term data from real office observations, demonstrates that generated hypotheses can effectively constrain the meaning of objects. Hence, we believe that the proposed system can be an essential component in a lifelong learning framework which acquires knowledge about objects from real world observations. Jay Young, Valerio Basile, Lars Kunze, Elena Cabrio, Nick Hawes |
ECAI | 2 |
| 2016 | Populating a Knowledge Base with Object-Location Relations Using Distributional Semantics
Valerio Basile, Soufian Jebbara, Elena Cabrio, Philipp Cimiano |
EKAW | 1 |
| 2014 | Preface: 17th International conference on Applications of Natural Language to Information Systems (NLDB 2012)
Gosse Bouma, Valerio Basile, Ashwin Ittoo, Elisabeth Métais, Hans Wortmann |
Data Knowl. Eng. | 2 |
| 2013 | Elephant: Sequence Labeling for Word and Sentence SegmentationabstractTokenization is widely regarded as a solved problem due to the high accuracy that rulebased tokenizers achieve.But rule-based tokenizers are hard to maintain and their rules language specific.We show that highaccuracy word and sentence segmentation can be achieved by using supervised sequence labeling on the character level combined with unsupervised feature learning.We evaluated our method on three languages and obtained error rates of 0.27 ‰ (English), 0.35 ‰ (Dutch) and 0.76 ‰ (Italian) for our best models. Kilian Evang, Valerio Basile, Grzegorz Chrupala, Johan Bos |
EMNLP | 2 |
| 2012 | A platform for collaborative semantic annotation
Valerio Basile, Johan Bos, Kilian Evang, Noortje Venhuizen |
EACL | 1 |
| 2012 | An Empirical Approach to the Semantic Representation of LawsabstractTo make legal texts machine processable, the texts may be represented as linked documents, semantically tagged text, or translated to formal representations that can be automatically reasoned with. The paper considers the latter, which is key to testing consistency of laws, drawing inferences, and providing explanations relative to input. To translate laws to a form that can be reasoned with by a computer, sentences must be parsed and formally represented. The paper presents the state-of-the-art in automatic translation of law to a machine readable formal representation, provides corpora, outlines some key problems, and proposes tasks to address the problems. Adam Z. Wyner, Johan Bos, Valerio Basile, Paulo Quaresma |
JURIX | 3 |
| 2012 | Developing a large semantically annotated corpus
Valerio Basile, Johan Bos, Kilian Evang, Noortje Venhuizen |
LREC | 1 |