Luca Ragazzi

dblp:320/5349 · DBLP profile ↗
← Back
16ranked-venue papers
0as first author
16since 2021 · last 2026
0000-0003-3574-9962ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 LLMs (Almost) Never Abstain Under Medical Uncertainty
abstract
Medical multiple-choice question answering (MCQA) benchmarks implicitly assume that large language models (LLMs) should always commit to an answer.However, in clinical practice, uncertainty is pervasive and abstaining is often the safest action.We introduce MedQAbstain, a benchmark explicitly designed to evaluate medical abstention under uncertainty.MedQAbstain repurposes standard medical MCQA datasets by removing the gold answer and introducing an explicit "I abstain" option, framed as a safety-critical decision with clinical consequences.The benchmark supports systematic analysis across abstention regimes, distractor complexity, and input modalities, and elicits self-reported model confidence to study calibration.Across all settings, we find that state-of-the-art LLMs systematically overcommit, rarely abstaining even when the question itself is hidden.These results reveal a fundamental mismatch between LLM behavior and clinical norms, highlighting abstention as a critical but overlooked dimension of medical decision-making evaluation. 1 * Equal contribution (co-first authors).
Alessio Cocchieri, Luca Ragazzi, Giuseppe Tagliavini, Gianluca Moro
ACL (1)2
2026 Sycophants in the Courtroom: Are LLMs Fragile to Juridical Authority and Evolving Legal Standards?
abstract
Lorenzo Molfetta, Alessio Cocchieri, Luca Ragazzi, Ilaria Bartolini, Marco Patella, Gianluca Moro. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Lorenzo Molfetta, Alessio Cocchieri, Luca Ragazzi, Ilaria Bartolini, Marco Patella, Gianluca Moro
ACL (1)3
2026 OpenBioNER-v2: A suite of lightweight models for zero-shot medical named entity recognition via type descriptions
abstract
Named entity recognition (NER) in medicine is challenging due to specialized terminology, inconsistent annotation guidelines, and the continuous emergence of new entity types—requiring models that can adapt to unseen targets. Large language models (LLMs) exhibit strong generalization but are impractical for scalable deployment, whereas recent encoder-only approaches leverage entity names for zero-shot inference but struggle with disambiguation in complex domains. We introduce OpenBioNER-v2, a family of lightweight transformer encoders (15M–110M parameters) designed for zero-shot recognition of biomedical and clinical entities by conditioning on natural language descriptions of target types. Our cross-encoder architecture jointly models input text and entity-type descriptions, enabling semantic matches. Pretrained on LLM-generated silver annotations and multi-view descriptions covering thousands of medical types, OpenBioNER-v2 achieves state-of-the-art results across 11 benchmarks—including a new dataset for personal de-identification. Variants with ≤ 56M parameters outperform both large and small language models, such as UniversalNER and GliNER. Ablation studies reveal effective strategies for formulating descriptions. All data, code, and model checkpoints are publicly released under open-science principles.
Alessio Cocchieri, Giacomo Frisoni, Francesco Zangrillo, Luca Ragazzi, Marcos Martínez Galindo, Giuseppe Tagliavini, Gianluca Moro
Expert Syst. Appl.4
2026 Clash-of-Leges: A bilingual dataset for conflict detection and explanation in statutory law
abstract
• Comprehensive Legal Conflict Dataset: Clash-of-Leges introduces a novel dataset that captures authentic legal conflicts between provisions across different documents, derived from Italian Constitutional Court rulings, to support advanced AI research in legal conflict detection and explanation. • Bilingual and Multi-Task Framework: The dataset includes annotations for three key tasks-Conflict Detection, Conflict Explanation Generation, and Reference Retrieval-while offering bilingual support in Italian and English to enhance global research applicability. • Extensive Evalutions: The manuscript resents an extensive experimental evaluation of several large language models across all defined tasks, illustrating both their capabilities and current limitations. Legal conflicts between statutes or constitutional articles present a significant challenge in maintaining consistency and coherence within legal systems. Addressing these conflicts requires extensive human expertise, making the process labor-intensive and time-consuming. In this paper, we introduce Clash-of-Leges , a novel multilingual dataset derived from rulings by the Constitutional Court of the Italian Republic, designed to aid the automation of conflict detection and explanation between legal articles. We identify three key tasks: Conflict Classification, which determines whether two legal articles are in conflict; Conflict Explanation Generation, which provides detailed explanations for identified conflicts; and Reference Retrieval, which sources relevant legal bases and precedents to substantiate interpretations. These tasks are intended to facilitate the development of AI models that can automatically identify and explain contradictions between legal provisions. 2
Paolo Italiani, Gianluca Moro, Luca Ragazzi
Expert Syst. Appl.3
2026 Abstractive summarization through the prism of decoding strategies
abstract
In natural language generation, abstractive summarization (AS) is advancing rapidly due to transformer-based language models (LMs). Although decoding strategies significantly influence generated summaries, their significance is often overlooked. Given the abundance of token selection heuristics and associated hyperparameters, the community needs guidance to make well-informed decisions based on the specific task and target metrics. To address this gap, we conduct a comparative assessment of the effectiveness and efficiency of decoding-time techniques for short, long, and multi-document AS. We explore over 3,500 combinations involving three widely used million-scale autoregressive encoder-decoder LMs, two billion-scale decoder-only LMs, six datasets, and nine decoding settings. Our findings highlight that optimized decoding choices can lead to substantial performance improvements. Alongside human evaluation, we quantitatively measure effects using ten automatic metrics, covering dimensions such as semantic similarity, factuality, compression, redundancy, and carbon footprint. To set the stage for differentiable selection and optimization of decoding options, we introduce Prism, a first-of-its-kind dataset that pairs AS gold input-output examples with our LM predictions across a diverse range of decoding options. The code and data are publicly available athttps://github.com/disi-unibo-nlp/prism.
Giacomo Frisoni, Luca Ragazzi, Gianluca Moro, Antonella Carbonaro, Claudio Sartori 0001
Neural Networks2
2025 "What do you call a dog that is incontrovertibly true? Dogma": Testing LLM Generalization through Humor
abstract
Alessio Cocchieri, Luca Ragazzi, Paolo Italiani, Giuseppe Tagliavini, Gianluca Moro. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Alessio Cocchieri, Luca Ragazzi, Paolo Italiani, Giuseppe Tagliavini, Gianluca Moro
ACL (1)2
2025 Magic Mirror on the Wall, Which Is the Fairest Prompt of All? A Survey on Automatic Prompt Learning
abstract
Prompts direct the behavior of a model by conditioning its outputs on carefully designed instructions and examples, similar to setting the trajectory of an arrow before release. More broadly, prompt learning is the research area that aims to solve downstream tasks by directly leveraging the knowledge acquired by language models at pretraining time, removing the need for expensive fine-tuning stages with potentially different objective functions. While manual prompt engineering has enabled both small and large language models to achieve superhuman performance on numerous benchmarks, it remains a labor-intensive and suboptimal process. Recently, the field has shifted towards automating the search for prompts that effectively elicit the desired model responses. This survey presents the first systematic review of prompt learning for pre-trained language models operating on textual inputs, with a particular focus on automatic methods. We critically analyze existing publications and organize them into a novel taxonomy, describing key aspects for practical usage. We finally discuss promising directions for future research. Our curated repository of annotated papers, continuously updated, is available at https://github.com/disi-unibo-nlp/awesome-prompt-learning.
Stefano Fantazzini, Giacomo Frisoni, Gianluca Moro, Luca Ragazzi, Mario Ciccioni, Claudio Sartori 0001
ECAI4
2025 FEAST: Retrieval-Augmented Multi-Hierarchical Food Classification for the FoodEx2 System
abstract
Hierarchical text classification (HTC) and extreme multi-label classification (XML) tasks face compounded challenges from complex label interdependencies, data sparsity, and extreme output dimensions. These challenges are exemplified in the European Food Safety Authority’s FoodEx2 system–a standardized food classification framework essential for food consumption monitoring and contaminant exposure assessment across Europe. FoodEx2 coding transforms natural language food descriptions into a set of codes from multiple standardized hierarchies, but faces implementation barriers due to its complex structure. Given a food description (e.g., “organic yogurt”), the system identifies its base term (“yogurt”), all the applicable facet categories (e.g., “production method”), and then, every relevant facet descriptors to each category (e.g., “organic production”). While existing models perform adequately on well-balanced and semantically dense hierarchies, no work has been applied on the practical constraints imposed by the FoodEx2 system. The limited literature addressing such real-world scenarios further compounds these challenges. We propose FEAST (Food Embedding And Semantic Taxonomy), a novel retrieval-augmented framework that decomposes FoodEx2 classification into a three-stage approach: (1) base term identification, (2) multi-label facet prediction, and (3) facet descriptor assignment. By leveraging the system’s hierarchical structure to guide training and performing deep metric learning, FEAST learns discriminative embeddings that mitigate data sparsity and improve generalization on rare and fine-grained labels. Evaluated on the multilingual FoodEx2 benchmark, FEAST outperforms the prior European’s CNN baseline F1 scores by 12–38% on rare classes.
Lorenzo Molfetta, Alessio Cocchieri, Stefano Fantazzini, Giacomo Frisoni, Luca Ragazzi, Gianluca Moro
ECAI5
2025 Can Large Language Models Win the International Mathematical Games?
abstract
Recent advances in large language models (LLMs) have demonstrated strong mathematical reasoning abilities, even in visual contexts, with some models surpassing human performance on existing benchmarks.However, these benchmarks lack structured age categorization, clearly defined skill requirements, and-crucially-were not designed to assess human performance in international competitions.To address these limitations, we introduce MATHGAMES, a new benchmark of 2,183 high-quality mathematical problems (both text-only and multimodal) in an openended format, sourced from an international mathematical games championships.Spanning seven age groups and a skill-based taxonomy, MATHGAMES enables a structured evaluation of LLMs' mathematical and logical reasoning abilities.Our experiments reveal a substantial gap between state-of-theart LLMs and human participants-even 11year-olds consistently outperform some of the strongest models-highlighting the need for advancements.Further, our detailed error analysis offers valuable insights to guide future research.The data is publicly available at https:// disi-unibo-nlp.github.io/math-games/. * Equal contribution (co-first authors).Number of participants per score C1 (11-13 y/o) 1413 C2 (13-15 y/o) 709 L1 (15-18 y/o) 338 L2 (18-20 y/o) 177 GP (20-25 y/o) 112 0 10 20 30 40 50 60 70 80 90 100 Competition Score (%) HC (25+ y/o) 25 Gemini-2.0-Flash-ThinkGemini-1.5-ProGPT-4o GPT-4o-mini Gemini-1.5-FlashGemini-1.5-Flash-8BEarly Teenager (11-13 y/o) Late Teenager (13-18 y/o) Adult (18-25+ y/o) Question: How many small spheres of different colors are there in the figure?Question: In figure you see tennis balls placed on top of each other, forming at each "plane" of the squares, without holes in the middle.The highest level contains only one ball; the second, coming down, contains 4; the third contains 9 and so on.If you use 7714 balls, how many floors will your pyramid of tennis balls be constituted?Question: In figure you see a pentagonal tile, quite singular, whose sides BC and AE measure 1 dm while AB measures 2 dm.Which is in cm 2 , rounded to the nearest cm 2 , the area of our tile?(If necessary, use 1,414 for √ 2 and 1,732 for √ 3).
Alessio Cocchieri, Luca Ragazzi, Giuseppe Tagliavini, Lorenzo Tordi, Antonella Carbonaro, Gianluca Moro
EMNLP2
2024 Unknown Claims: Generation of Fact-Checking Training Examples from Unstructured and Structured Data
abstract
Computational fact-checking (FC) relies on supervised models to verify claims based on given evidence, requiring a resource-intensive process to annotate large volumes of training data.We introduce UNOWN, a novel framework that generates training instances for FC systems automatically using both textual and tabular content.UNOWN selects relevant evidence and generates supporting and refuting claims with advanced negation artifacts.Designed to be flexible, UNOWN accommodates various strategies for evidence selection and claim generation, offering unparalleled adaptability.We comprehensively evaluate UNOWN on both text-only and table+text benchmarks, including FEVEROUS, SCIFACT, and MMFC, a new multi-modal FC dataset.Our results prove that UNOWN examples are of comparable quality to expert-labeled data, even enabling models to achieve up to 5% higher accuracy.The code, data, and models are available at https: //github.com/disi-unibo-nlp/
Jean-Flavien Bussotti, Luca Ragazzi, Giacomo Frisoni, Gianluca Moro, Paolo Papotti
EMNLP2
2023 Carburacy: Summarization Models Tuning and Comparison in Eco-Sustainable Regimes with a Novel Carbon-Aware Accuracy
abstract
Generative transformer-based models have reached cutting-edge performance in long document summarization. Nevertheless, this task is witnessing a paradigm shift in developing ever-increasingly computationally-hungry solutions, focusing on effectiveness while ignoring the economic, environmental, and social costs of yielding such results. Accordingly, such extensive resources impact climate change and raise barriers to small and medium organizations distinguished by low-resource regimes of hardware and data. As a result, this unsustainable trend has lifted many concerns in the community, which directs the primary efforts on the proposal of tools to monitor models' energy costs. Despite their importance, no evaluation measure considering models' eco-sustainability exists yet. In this work, we propose Carburacy, the first carbon-aware accuracy measure that captures both model effectiveness and eco-sustainability. We perform a comprehensive benchmark for long document summarization, comparing multiple state-of-the-art quadratic and linear transformers on several datasets under eco-sustainable regimes. Finally, thanks to Carburacy, we found optimal combinations of hyperparameters that let models be competitive in effectiveness with significantly lower costs.
Gianluca Moro, Luca Ragazzi, Lorenzo Valgimigli
AAAI2
2023 Graph-Based Abstractive Summarization of Extracted Essential Knowledge for Low-Resource Scenarios
abstract
Although current summarization models can process increasingly long text sequences, they still struggle to capture salient related information spread across the lengthy size of inputs with few labeled training instances. Today’s research still relies on standard input truncation without considering graph-based modeling of multiple semantic units to summarize only crucial facets. This paper proposes G-SEEK, a graph-based summarization of extracted essential knowledge. By representing the long source with a heterogeneous graph, our method extracts and provides salient sentences to an abstractive summarization model to generate the summary. Experimental results in low-resource scenarios, distinguished by data scarcity, reveal that G-SEEK consistently improves both the long- and multi-document summarization performance and accuracy across several datasets.
Gianluca Moro, Luca Ragazzi, Lorenzo Valgimigli
ECAI2
2023 Retrieve-and-Rank End-to-End Summarization of Biomedical Studies
Gianluca Moro, Luca Ragazzi, Lorenzo Valgimigli, Lorenzo Molfetta
SISAP2
2023 Align-then-abstract representation learning for low-resource summarization
abstract
Generative transformer-based models have achieved state-of-the-art performance in text summarization. Nevertheless, they still struggle in real-world scenarios with long documents when trained in low-resource settings of a few dozen labeled training instances, namely in low-resource summarization (LRS). This paper bridges the gap by addressing two key research challenges when summarizing long documents, i.e., long-input processing and document representation, in one coherent model trained for LRS. Specifically, our novel align-then-abstract representation learning model (Athena) jointly trains a segmenter and a summarizer by maximizing the alignment between the chunk-target pairs in output from the text segmentation. Extensive experiments reveal that Athena outperforms the current state-of-the-art approaches in LRS on multiple long document summarization datasets from different domains.
Gianluca Moro, Luca Ragazzi
Neurocomputing2
2022 Semantic Self-Segmentation for Abstractive Summarization of Long Documents in Low-Resource Regimes
abstract
The quadratic memory complexity of transformers prevents long document summarization in low computational resource scenarios. State-of-the-art models need to apply input truncation, thus discarding and ignoring potential summary-relevant contents, leading to a performance drop. Furthermore, this loss is generally destructive for semantic text analytics in high-impact domains such as the legal one. In this paper, we propose a novel semantic self-segmentation (Se3) approach for long document summarization to address the critical problems of low-resource regimes, namely to process inputs longer than the GPU memory capacity and produce accurate summaries despite the availability of only a few dozens of training instances. Se3 segments a long input into semantically coherent chunks, allowing transformers to summarize very long documents without truncation by summarizing each chunk and concatenating the results. Experimental outcomes show the approach significantly improves the performance of abstractive summarization transformers, even with just a dozen of labeled data, achieving new state-of-the-art results on two legal datasets of different domains and contents. Finally, we report ablation studies to evaluate each contribution of the components of our method to the performance gain.
Gianluca Moro, Luca Ragazzi
AAAI2
2022 Discriminative Marginalized Probabilistic Neural Method for Multi-Document Summarization of Medical Literature
abstract
Although current state-of-the-art Transformerbased solutions succeeded in a wide range for single-document NLP tasks, they still struggle to address multi-input tasks such as multidocument summarization.Many solutions truncate the inputs, thus ignoring potential summary-relevant contents, which is unacceptable in the medical domain where each information can be vital.Others leverage linear model approximations to apply multi-input concatenation, worsening the results because all information is considered, even if it is conflicting or noisy with respect to a shared background.Despite the importance and social impact of medicine, there are no ad-hoc solutions for multi-document summarization.For this reason, we propose a novel discriminative marginalized probabilistic method (DAMEN) trained to discriminate critical information from a cluster of topic-related medical documents and generate a multi-document summary via token probability marginalization.Results prove we outperform the previous state-of-theart on a biomedical dataset for multi-document summarization of systematic literature reviews.Moreover, we perform extensive ablation studies to motivate the design choices and prove the importance of each module of our method.1
Gianluca Moro, Luca Ragazzi, Lorenzo Valgimigli, Davide Freddi
ACL (1)2