EDBT 2026 Demo / reviewers in the wild / expert
Alessio Cocchieri
dblp:371/9197
· DBLP profile ↗
7ranked-venue papers
4as first author
7since 2021 · last 2026
0009-0003-1507-1354ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Trustworthy machine learning · 36% Language models and text generation · 25% Question answering and dialogue systems · 23% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Computational social science and digital humanities · 54% Computing education · 46% |
Topics — the 12 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Question answering and dialogue systems › domain-specific question answering
medical question answering |
1.8 | 2 | 2026 | LLMs (Almost) Never Abstain Under Medical Uncertainty · ACL (1) 2026 To Generate or to Retrieve? On the Effectiveness of Artificial Contexts for Medical Open-Domain Question Answering · ACL (1) 2024 |
Machine learning › Trustworthy machine learning › large language model trustworthiness
large language model robustness |
1.0 | 1 | 2026 | Sycophants in the Courtroom: Are LLMs Fragile to Juridical Authority and Evolving Legal Standards? · ACL (1) 2026 |
Natural language and speech › Language models and text generation › prompting
prompt sensitivity |
1.0 | 1 | 2026 | Sycophants in the Courtroom: Are LLMs Fragile to Juridical Authority and Evolving Legal Standards? · ACL (1) 2026 |
Machine learning › Trustworthy machine learning
robustness |
1.0 | 1 | 2026 | Sycophants in the Courtroom: Are LLMs Fragile to Juridical Authority and Evolving Legal Standards? · ACL (1) 2026 |
Machine learning › Trustworthy machine learning › uncertainty estimation
selective classification |
1.0 | 1 | 2026 | LLMs (Almost) Never Abstain Under Medical Uncertainty · ACL (1) 2026 |
Machine learning › Trustworthy machine learning
uncertainty and calibration |
1.0 | 1 | 2026 | LLMs (Almost) Never Abstain Under Medical Uncertainty · ACL (1) 2026 |
Computer vision › Vision and language › multimodal understanding
humor understanding |
0.9 | 1 | 2025 | "What do you call a dog that is incontrovertibly true? Dogma": Testing LLM Generalization through Humor · ACL (1) 2025 |
Natural language and speech › Language models and text generation › linguistic generalization
language model generalization |
0.9 | 1 | 2025 | "What do you call a dog that is incontrovertibly true? Dogma": Testing LLM Generalization through Humor · ACL (1) 2025 |
Natural language and speech › Language models and text generation
mathematical reasoning |
0.9 | 1 | 2025 | Can Large Language Models Win the International Mathematical Games? · EMNLP 2025 |
Computer vision › Vision and language
multimodal reasoning |
0.9 | 1 | 2025 | Can Large Language Models Win the International Mathematical Games? · EMNLP 2025 |
Natural language and speech › Question answering and dialogue systems
open-domain question answering |
0.8 | 1 | 2024 | To Generate or to Retrieve? On the Effectiveness of Artificial Contexts for Medical Open-Domain Question Answering · ACL (1) 2024 |
Computational social science and digital humanities › legal informatics
legal text analysis |
0.3 | 1 | 2026 | Sycophants in the Courtroom: Are LLMs Fragile to Juridical Authority and Evolving Legal Standards? · ACL (1) 2026 |
Methods — techniques the papers use, named apart from their topics
large language model · 2.6large language model prompting · 2.0benchmark construction · 1.0prompting · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LLMs (Almost) Never Abstain Under Medical UncertaintyabstractMedical multiple-choice question answering (MCQA) benchmarks implicitly assume that large language models (LLMs) should always commit to an answer.However, in clinical practice, uncertainty is pervasive and abstaining is often the safest action.We introduce MedQAbstain, a benchmark explicitly designed to evaluate medical abstention under uncertainty.MedQAbstain repurposes standard medical MCQA datasets by removing the gold answer and introducing an explicit "I abstain" option, framed as a safety-critical decision with clinical consequences.The benchmark supports systematic analysis across abstention regimes, distractor complexity, and input modalities, and elicits self-reported model confidence to study calibration.Across all settings, we find that state-of-the-art LLMs systematically overcommit, rarely abstaining even when the question itself is hidden.These results reveal a fundamental mismatch between LLM behavior and clinical norms, highlighting abstention as a critical but overlooked dimension of medical decision-making evaluation. 1 * Equal contribution (co-first authors). Alessio Cocchieri, Luca Ragazzi, Giuseppe Tagliavini, Gianluca Moro |
ACL (1) | 1 |
| 2026 | Sycophants in the Courtroom: Are LLMs Fragile to Juridical Authority and Evolving Legal Standards?abstractLorenzo Molfetta, Alessio Cocchieri, Luca Ragazzi, Ilaria Bartolini, Marco Patella, Gianluca Moro. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Lorenzo Molfetta, Alessio Cocchieri, Luca Ragazzi, Ilaria Bartolini, Marco Patella, Gianluca Moro |
ACL (1) | 2 |
| 2026 | OpenBioNER-v2: A suite of lightweight models for zero-shot medical named entity recognition via type descriptionsabstractNamed entity recognition (NER) in medicine is challenging due to specialized terminology, inconsistent annotation guidelines, and the continuous emergence of new entity types—requiring models that can adapt to unseen targets. Large language models (LLMs) exhibit strong generalization but are impractical for scalable deployment, whereas recent encoder-only approaches leverage entity names for zero-shot inference but struggle with disambiguation in complex domains. We introduce OpenBioNER-v2, a family of lightweight transformer encoders (15M–110M parameters) designed for zero-shot recognition of biomedical and clinical entities by conditioning on natural language descriptions of target types. Our cross-encoder architecture jointly models input text and entity-type descriptions, enabling semantic matches. Pretrained on LLM-generated silver annotations and multi-view descriptions covering thousands of medical types, OpenBioNER-v2 achieves state-of-the-art results across 11 benchmarks—including a new dataset for personal de-identification. Variants with ≤ 56M parameters outperform both large and small language models, such as UniversalNER and GliNER. Ablation studies reveal effective strategies for formulating descriptions. All data, code, and model checkpoints are publicly released under open-science principles. Alessio Cocchieri, Giacomo Frisoni, Francesco Zangrillo, Luca Ragazzi, Marcos Martínez Galindo, Giuseppe Tagliavini, Gianluca Moro |
Expert Syst. Appl. | 1 |
| 2025 | "What do you call a dog that is incontrovertibly true? Dogma": Testing LLM Generalization through HumorabstractAlessio Cocchieri, Luca Ragazzi, Paolo Italiani, Giuseppe Tagliavini, Gianluca Moro. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Alessio Cocchieri, Luca Ragazzi, Paolo Italiani, Giuseppe Tagliavini, Gianluca Moro |
ACL (1) | 1 |
| 2025 | FEAST: Retrieval-Augmented Multi-Hierarchical Food Classification for the FoodEx2 SystemabstractHierarchical text classification (HTC) and extreme multi-label classification (XML) tasks face compounded challenges from complex label interdependencies, data sparsity, and extreme output dimensions. These challenges are exemplified in the European Food Safety Authority’s FoodEx2 system–a standardized food classification framework essential for food consumption monitoring and contaminant exposure assessment across Europe. FoodEx2 coding transforms natural language food descriptions into a set of codes from multiple standardized hierarchies, but faces implementation barriers due to its complex structure. Given a food description (e.g., “organic yogurt”), the system identifies its base term (“yogurt”), all the applicable facet categories (e.g., “production method”), and then, every relevant facet descriptors to each category (e.g., “organic production”). While existing models perform adequately on well-balanced and semantically dense hierarchies, no work has been applied on the practical constraints imposed by the FoodEx2 system. The limited literature addressing such real-world scenarios further compounds these challenges. We propose FEAST (Food Embedding And Semantic Taxonomy), a novel retrieval-augmented framework that decomposes FoodEx2 classification into a three-stage approach: (1) base term identification, (2) multi-label facet prediction, and (3) facet descriptor assignment. By leveraging the system’s hierarchical structure to guide training and performing deep metric learning, FEAST learns discriminative embeddings that mitigate data sparsity and improve generalization on rare and fine-grained labels. Evaluated on the multilingual FoodEx2 benchmark, FEAST outperforms the prior European’s CNN baseline F1 scores by 12–38% on rare classes. Lorenzo Molfetta, Alessio Cocchieri, Stefano Fantazzini, Giacomo Frisoni, Luca Ragazzi, Gianluca Moro |
ECAI | 2 |
| 2025 | Can Large Language Models Win the International Mathematical Games?abstractRecent advances in large language models (LLMs) have demonstrated strong mathematical reasoning abilities, even in visual contexts, with some models surpassing human performance on existing benchmarks.However, these benchmarks lack structured age categorization, clearly defined skill requirements, and-crucially-were not designed to assess human performance in international competitions.To address these limitations, we introduce MATHGAMES, a new benchmark of 2,183 high-quality mathematical problems (both text-only and multimodal) in an openended format, sourced from an international mathematical games championships.Spanning seven age groups and a skill-based taxonomy, MATHGAMES enables a structured evaluation of LLMs' mathematical and logical reasoning abilities.Our experiments reveal a substantial gap between state-of-theart LLMs and human participants-even 11year-olds consistently outperform some of the strongest models-highlighting the need for advancements.Further, our detailed error analysis offers valuable insights to guide future research.The data is publicly available at https:// disi-unibo-nlp.github.io/math-games/. * Equal contribution (co-first authors).Number of participants per score C1 (11-13 y/o) 1413 C2 (13-15 y/o) 709 L1 (15-18 y/o) 338 L2 (18-20 y/o) 177 GP (20-25 y/o) 112 0 10 20 30 40 50 60 70 80 90 100 Competition Score (%) HC (25+ y/o) 25 Gemini-2.0-Flash-ThinkGemini-1.5-ProGPT-4o GPT-4o-mini Gemini-1.5-FlashGemini-1.5-Flash-8BEarly Teenager (11-13 y/o) Late Teenager (13-18 y/o) Adult (18-25+ y/o) Question: How many small spheres of different colors are there in the figure?Question: In figure you see tennis balls placed on top of each other, forming at each "plane" of the squares, without holes in the middle.The highest level contains only one ball; the second, coming down, contains 4; the third contains 9 and so on.If you use 7714 balls, how many floors will your pyramid of tennis balls be constituted?Question: In figure you see a pentagonal tile, quite singular, whose sides BC and AE measure 1 dm while AB measures 2 dm.Which is in cm 2 , rounded to the nearest cm 2 , the area of our tile?(If necessary, use 1,414 for √ 2 and 1,732 for √ 3). Alessio Cocchieri, Luca Ragazzi, Giuseppe Tagliavini, Lorenzo Tordi, Antonella Carbonaro, Gianluca Moro |
EMNLP | 1 |
| 2024 | To Generate or to Retrieve? On the Effectiveness of Artificial Contexts for Medical Open-Domain Question AnsweringabstractMedical open-domain question answering demands substantial access to specialized knowledge.Recent efforts have sought to decouple knowledge from model parameters, counteracting architectural scaling and allowing for training on common low-resource hardware.The retrieve-then-read paradigm has become ubiquitous, with model predictions grounded on relevant knowledge pieces from external repositories such as PubMed, textbooks, and UMLS.An alternative path, still under-explored but made possible by the advent of domain-specific large language models, entails constructing artificial contexts through prompting.As a result, "to generate or to retrieve" is the modern equivalent of Hamlet's dilemma.This paper presents MEDGENIE, the first generate-thenread framework for multiple-choice question answering in medicine.We conduct extensive experiments on MedQA-USMLE, MedMCQA, and MMLU, incorporating a practical perspective by assuming a maximum of 24GB VRAM.MEDGENIE sets a new state-of-the-art in the open-book setting of each testbed, allowing a small-scale reader to outcompete zero-shot closed-book 175B baselines while using up to 706× fewer parameters.Our findings reveal that generated passages are more effective than retrieved ones in attaining higher accuracy.1 * Equal contribution (co-first authorship). 1 Our code, fine-tuned models, and generated contexts are publicly available at https://github.com/unibo-nlp/medgenie. Giacomo Frisoni, Alessio Cocchieri, Alex Presepi, Gianluca Moro, Zaiqiao Meng |
ACL (1) | 2 |