Simone Merlo

dblp:359/0211 · DBLP profile ↗
← Back
5ranked-venue papers in the field
3as first author
5since 2021 · last 2026
0009-0003-8003-4795ORCID · corroborated

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 5 (3 first)
YearPublicationVenuePosition
2026 Reducing Human Effort to Validate LLM Relevance Judgements via Stratified Sampling
Simone Merlo, Stefano Marchesin 0001, Guglielmo Faggioli, Nicola Ferro 0001
ECIR (1)1
2025 A Cost-Effective Framework to Evaluate LLM-Generated Relevance Judgements
abstract
Large Language Models (LLMs) hugely impacted many research fields, including Information Retrieval (IR), where they are used for many sub-tasks, such as query rewriting and retrieval augmented generation. At the same time, the research community is investigating whether and how to use LLMs to support, or even replace, humans to generate relevance judgments. Indeed, generating relevance judgements automatically - or integrating an LLM in the annotation process - would allow us to improve the number of evaluation collections, also for scenarios where the annotation process is particularly challenging. To validate relevance judgements produced by an LLM they are compared with human-made relevance judgements, measuring the inter-assessor agreement between the human and the LLM.
Simone Merlo, Stefano Marchesin 0001, Guglielmo Faggioli, Nicola Ferro 0001
CIKM1
2025 Conversational Information Retrieval and Recommender Systems
Guglielmo Faggioli, Nicola Ferro 0001, Simone Merlo
ECIR (5)3
2025 A Reproducibility Study for Joint Information Retrieval and Recommendation in Product Search
Simone Merlo, Guglielmo Faggioli, Nicola Ferro 0001
ECIR (4)1
2025 CoSRec: A Joint Conversational Search and Recommendation Dataset
abstract
Conversational Information Access systems have experienced widespread diffusion thanks to the natural and effortless interactions they enable with the user. In particular, they represent an effective interaction interface for conversational search (CS) and conversational recommendation (CR) scenarios. Despite their commonalities, CR and CS systems are often devised, developed, and evaluated as isolated components. Integrating these two elements would allow for handling complex information access scenarios, such as exploring unfamiliar recommended product aspects, enabling richer dialogues, and improving user satisfaction. As of today, the scarce availability of integrated datasets - focused exclusively on either of the tasks - limits the possibilities for evaluating by-design integrated CS and CR systems. To address this gap, we propose CoSRec, the first dataset for joint Conversational Search and Recommendation (CSR) evaluation. The CoSRec test set includes 20 high-quality conversations, with human-made annotations for the quality of conversations, and manually crafted relevance judgments for products and documents. Additionally, we provide supplementary training data comprising partially annotated dialogues and raw conversations to support diverse learning paradigms. CoSRec is the first resource to model CR and CS tasks in a unified framework, enabling the training and evaluation of systems that must shift between answering queries and making suggestions dynamically.
Marco Alessio, Simone Merlo, Tommaso Di Noia, Guglielmo Faggioli, Marco Ferrante, Nicola Ferro 0001, Cristina Ioana Muntean, Franco Maria Nardini, Fedelucio Narducci, Raffaele Perego 0001, Giuseppe Santucci, Nicola Viterbo
SIGIR2