EDBT 2026 Demo / reviewers in the wild / expert
Andrea Galassi
dblp:208/4245
· DBLP profile ↗
17ranked-venue papers
4as first author
14since 2021 · last 2025
0000-0001-9711-7042ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 4 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Detecting Vague Clauses in Italian Privacy Policies Using Transformers, LLMs, and Cross-Lingual TechniquesabstractPrivacy policies often fall short of providing a comprehensive account of how personal data is used, thus failing to comply with GDPR requirements. By doing so, they hamper the users’ ability to make informed decisions about using services while ensuring that their data is used properly and fairly. This calls for automatic tools that can effectively identify potentially unlawful policies. Here we present a new corpus of Italian privacy policies, with clauses labelled by experts in data protection law, to indicate the level of comprehensiveness of information. We focus on the categories of data processed, classifying each clause as either sufficiently or insufficiently informative (“vague”). We perform 6 different classification and detection tasks, comparing the performance of BERT-based models and generative Large Language Models. Addressing multilingualism is crucial in the EU, whose 24 spoken languages are an integral part of its cultural heritage. Consequentely, we also perform cross-language experiments to evaluate whether a pre-existing English corpus or classifiers can be leveraged for Italian and, vice versa, whether our corpus is informative enough to generalize to other languages. Giulia Grundler, Mariaceleste Musicco, Andrea Galassi, Francesca Lagioia, Ruta Liepina, Giorgio Resta, Sara Roccu, Giovanni Sartor, Paolo Torroni |
ECAI | 3 |
| 2025 | Is It Worth Using LLMs for Unfair Clause Detection in Terms of Service?abstractUnfair clause detection is an extremely useful AI application for consumer protection. Artificial intelligence has recently been successful in building systems capable to automatically detect unfair clauses in Terms of Service, and also to identify their unfairness categories. Since Large Language Models (LLMs) are nowadays bringing a revolution to the field of artificial intelligence, and in particular to natural language processing and understanding, in this paper we compare several different prompt strategies for LLMs with more traditional BERT-based fine-tuned models. Our extensive experimental evaluation aims to investigate whether it is worth using LLMs also for this challenging domain-specific task. Marco Panarelli, Andrea Galassi, Francesca Lagioia, Ruta Liepina, Marco Lippi 0001, Przemyslaw Palka, Giovanni Sartor |
ICAIL | 2 |
| 2025 | Automated Extraction of Judicial Interpretative Formulas in EU Case Law on VATabstractThis paper addresses the extraction of Judicial Interpretative Formulas (JIFs) in decisions of the Court of Justice of the European Union (CJEU) on Value Added Tax (VAT). European case law includes a significant number of JIFs on this subject, which are crucial for the interpretation of VAT. However, extracting such JIFs manually is effortful, and doing that automatically has not been investigated yet in the VAT domain. Our work proposes the first pipeline method for doing so. We start by defining a set of guidelines for annotating legal texts following a principle definition of JIF. By following such guidelines, we obtain a corpus of 21 expert-labeled CJEU decisions. We keep them for validation and testing. For training, we machine-annotate 80 additional decisions using LLMs. Our experiments show that BERT-based architectures trained on such data perform comparably to LLMs. Giulia Grundler, Piera Santin, Alessia Fidelangeli, Rachele Mignone, Federico Galli, Andrea Galassi, Giuseppe Contissa, Luigi Di Caro, Paolo Torroni |
JURIX | 6 |
| 2025 | Promoting the Responsible Development of Speech Datasets for Mental Health and Neurological Disorders ResearchabstractCurrent research in machine learning and artificial intelligence is largely centered on modeling and performance evaluation, less so on data collection. However, recent research demonstrated that limitations and biases in data may negatively impact trustworthiness and reliability. These aspects are particularly impactful on sensitive domains such as mental health and neurological disorders, where speech data are used to develop AI applications for patients and healthcare providers. In this paper, we chart the landscape of available speech datasets for this domain, to highlight possible pitfalls and opportunities for improvement and promote fairness and diversity. We present a comprehensive list of desiderata for building speech datasets for mental health and neurological disorders and distill it into an actionable checklist focused on ethical concerns to foster more responsible research. Eleonora Mancini, Ana Tanevska, Andrea Galassi, Alessio Galatolo, Federico Ruggeri, Paolo Torroni |
J. Artif. Intell. Res. | 3 |
| 2024 | A Corpus for Sentence-Level Subjectivity Detection on English News ArticlesabstractWe develop novel annotation guidelines for sentence-level subjectivity detection, which are not limited to language-specific cues. We use our guidelines to collect NewsSD-ENG, a corpus of 638 objective and 411 subjective sentences extracted from English news articles on controversial topics. Our corpus paves the way for subjectivity detection in English and across other languages without relying on language-specific tools, such as lexicons or machine translation. We evaluate state-of-the-art multilingual transformer-based models on the task in mono-, multi-, and cross-language settings. For this purpose, we re-annotate an existing Italian corpus. We observe that models trained in the multilingual setting achieve the best performance on the task. Francesco Antici, Federico Ruggeri, Andrea Galassi, Katerina Korre, Arianna Muti, Alessandra Bardi, Alice Fedotova, Alberto Barrón-Cedeño |
LREC/COLING | 3 |
| 2024 | A Chatbot for Asylum-Seeking Migrants in EuropeabstractWe present ACME: A Chatbot for asylum-seeking Migrants in Europe. ACME relies on computational argumentation and aims to help migrants identify the highest level of protection they can apply for. This would contribute to a more sustainable migration by reducing the load on territorial commissions, Courts, and humanitarian organizations supporting asylum applicants. We describe the background context, system architecture, underlying technologies, and a case study used to validate the tool with domain experts. Bettina Fazzinga, Elena Palmieri, Margherita Vestoso, Luca Bolognini, Andrea Galassi, Filippo Furfaro, Paolo Torroni |
ICTAI | 5 |
| 2024 | Detecting Vague Clauses in Privacy Policies: The Analysis of Data Categories Using BERT Models and LLMsabstractDespite some improvements in compliance metrics after the implementation of the European General Data Protection Regulation (GDPR), privacy policies have become longer and more ambiguous. They often fail to fully meet GDPR requirements, thus leaving users without a reliable way to understand how their data is processed. We present a novel corpus composed by 30 privacy policies of online platforms and a new set of annotation guidelines, to assess the level of comprehensiveness of information. We focus on the processed categories of data, classifying each clause either as fully informative or as insufficiently informative. In our experimental evaluation, we perform 6 different classification and detection tasks, comparing BERT models and generative Large Language Models. Giulia Grundler, Ruta Liepina, Mariaceleste Musicco, Francesca Lagioia, Andrea Galassi, Giovanni Sartor, Paolo Torroni |
JURIX | 5 |
| 2023 | The CLEF-2023 CheckThat! Lab: Checkworthiness, Subjectivity, Political Bias, Factuality, and Authority
Alberto Barrón-Cedeño, Firoj Alam, Tommaso Caselli, Giovanni Da San Martino, Tamer Elsayed, Andrea Galassi, Fatima Haouari, Federico Ruggeri, Julia Maria Struß, Rabindra Nath Nandi, Gullal Singh Cheema, Dilshod Azizov, Preslav Nakov |
ECIR (3) | 6 |
| 2023 | Argumentation Structure Prediction in CJEU Decisions on Fiscal State AidabstractArgument structure prediction aims to identify the relations between arguments or between parts of arguments. It is a crucial task in legal argument mining, where it could help identifying motivations behind judgments or even fallacies or inconsistencies. It is also a very challenging task, which is relatively underdeveloped compared to other argument mining tasks, owing to a number of reasons including a low availability of datasets and a high complexity of the reasoning involved. In this work, we address argumentative link prediction in decisions by Court of Justice of the European Union on fiscal state aid. We study how propositions are combined in higher-level structures and how the relations between propositions can be predicted by NLP models. To this end, we present a novel annotation scheme and use it to extend a dataset from literature with an additional annotation layer. We use our new dataset to run an empirical study, where we compare two architectures and explore different combinations of hyperparameters and training regimes. Our results indicate that an ensemble of residual networks yields the best results. Piera Santin, Giulia Grundler, Andrea Galassi, Federico Galli, Francesca Lagioia, Elena Palmieri, Federico Ruggeri, Giovanni Sartor, Paolo Torroni |
ICAIL | 3 |
| 2023 | A First Attempt to Detect Misinformation in Russia-Ukraine War News through Text Similarity
Nina Khairova, Bogdan Ivasiuk, Fabrizio Lo Scudo, Carmela Comito, Andrea Galassi |
LDK | 5 |
| 2023 | Multi-Task Attentive Residual Networks for Argument MiningabstractWe explore the use of residual networks and neural attention for multiple argument mining tasks. We propose a residual architecture that exploits attention, multi-task learning, and makes use of ensemble, without any assumption on document or argument structure. We present an extensive experimental evaluation on five different corpora of user-generated comments, scientific publications, and persuasive essays. Our results show that our approach is a strong competitor against state-of-the-art architectures with a higher computational footprint or corpus-specific design, representing an interesting compromise between generality, performance accuracy and reduced model size. Andrea Galassi, Marco Lippi 0001, Paolo Torroni |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2022 | AMICA: An Argumentative Search Engine for COVID-19 LiteratureabstractAMICA is an argument mining-based search engine, specifically designed for the analysis of scientific literature related to Covid-19. AMICA retrieves scientific papers based on matching keywords and ranks the results based on the papers' argumentative content. An experimental evaluation conducted on a case study in collaboration with the Italian National Institute of Health shows that the AMICA ranking agrees with expert opinion, as well as, importantly, with the impartial quality criteria indicated by Cochrane Systematic Reviews. Marco Lippi 0001, Francesco Antici, Gianfranco Brambilla, Evaristo Cisbani, Andrea Galassi, Daniele Giansanti, Fabio Magurano, Antonella Rosi, Federico Ruggeri, Paolo Torroni |
IJCAI | 5 |
| 2022 | Predicting Outcomes of Italian VAT DecisionsabstractThis study aims at predicting the outcomes of legal cases based on the textual content of judicial decisions. We present a new corpus of Italian documents, consisting of 226 annotated decisions on Value Added Tax by Regional Tax law commissions. We address the task of predicting whether a request is upheld or rejected in the final decision. We employ traditional classifiers and NLP methods to assess which parts of the decision are more informative for the task. Federico Galli, Giulia Grundler, Alessia Fidelangeli, Andrea Galassi, Francesca Lagioia, Elena Palmieri, Federico Ruggeri, Giovanni Sartor, Paolo Torroni |
JURIX | 4 |
| 2021 | Attention in Natural Language ProcessingabstractAttention is an increasingly popular mechanism used in a wide range of neural architectures. The mechanism itself has been realized in a variety of formats. However, because of the fast-paced advances in this domain, a systematic overview of attention is still missing. In this article, we define a unified model for attention architectures in natural language processing, with a focus on those designed to work with vector representations of the textual data. We propose a taxonomy of attention models according to four dimensions: the representation of the input, the compatibility function, the distribution function, and the multiplicity of the input and/or output. We present the examples of how prior information can be exploited in attention models and discuss ongoing research efforts and open challenges in the area, providing the first extensive categorization of the vast body of literature in this exciting domain. Andrea Galassi, Marco Lippi 0001, Paolo Torroni |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2020 | Cross-lingual Annotation Projection in Legal TextsabstractWe study annotation projection in text classification problems where source documents are published in multiple languages and may not be an exact translation of one another.In particular, we focus on the detection of unfair clauses in privacy policies and terms of service.We present the first English-German parallel asymmetric corpus for the task at hand.We study and compare several language-agnostic sentence-level projection methods.Our results indicate that a combination of word embeddings and dynamic time warping performs best. Andrea Galassi, Kasper Drazewski, Marco Lippi 0001, Paolo Torroni |
COLING | 1 |
| 2018 | Model Agnostic Solution of CSPs via Deep Learning: A Preliminary Study
Andrea Galassi, Michele Lombardi 0001, Paola Mello, Michela Milano |
CPAIOR | 1 |
| 2018 | Can Deep Networks Learn to Play by the Rules? A Case Study on Nine Men's MorrisabstractDeep networks have been successfully applied to a wide range of tasks in artificial intelligence, and game playing is certainly not an exception. In this paper, we present an experimental study to assess whether purely subsymbolic systems, such as deep networks, are capable of learning to play by the rules, without anya prioriknowledge neither of the game, nor of its rules, but only by observing the matches played by another player. Similar problems arise in many other application domains, where the goal is to learn rules, policies, behaviors, or decisions, simply by the observation of the dynamics of a system. We present a case study conducted with residual networks on the popular board game ofNine Men's Morris, showing that this kind of subsymbolic architecture is capable of correctly discriminating legal from illegal decisions, just from the observation of past matches of a single player. Federico Chesani, Andrea Galassi, Marco Lippi 0001, Paola Mello |
IEEE Trans. Games | 2 |