Marco Ferrante

dblp:147/9179 · DBLP profile ↗
← Back
12ranked-venue papers
6as first author
3since 2021 · last 2025
0000-0002-0894-4175ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 10 · 6 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1Theory of computation · 1
YearPublicationVenuePosition
2025 CoSRec: A Joint Conversational Search and Recommendation Dataset
abstract
Conversational Information Access systems have experienced widespread diffusion thanks to the natural and effortless interactions they enable with the user. In particular, they represent an effective interaction interface for conversational search (CS) and conversational recommendation (CR) scenarios. Despite their commonalities, CR and CS systems are often devised, developed, and evaluated as isolated components. Integrating these two elements would allow for handling complex information access scenarios, such as exploring unfamiliar recommended product aspects, enabling richer dialogues, and improving user satisfaction. As of today, the scarce availability of integrated datasets - focused exclusively on either of the tasks - limits the possibilities for evaluating by-design integrated CS and CR systems. To address this gap, we propose CoSRec, the first dataset for joint Conversational Search and Recommendation (CSR) evaluation. The CoSRec test set includes 20 high-quality conversations, with human-made annotations for the quality of conversations, and manually crafted relevance judgments for products and documents. Additionally, we provide supplementary training data comprising partially annotated dialogues and raw conversations to support diverse learning paradigms. CoSRec is the first resource to model CR and CS tasks in a unified framework, enabling the training and evaluation of systems that must shift between answering queries and making suggestions dynamically.
Marco Alessio, Simone Merlo, Tommaso Di Noia, Guglielmo Faggioli, Marco Ferrante, Nicola Ferro 0001, Cristina Ioana Muntean, Franco Maria Nardini, Fedelucio Narducci, Raffaele Perego 0001, Giuseppe Santucci, Nicola Viterbo
SIGIR5
2022 A Dependency-Aware Utterances Permutation Strategy to Improve Conversational Evaluation
Guglielmo Faggioli, Marco Ferrante, Nicola Ferro 0001, Raffaele Perego 0001, Nicola Tonellotto
ECIR (1)2
2021 Hierarchical Dependence-aware Evaluation Measures for Conversational Search
abstract
Conversational agents are drawing a lot of attention in the information retrieval (IR) community also thanks to the advancements in language understanding enabled by large contextualized language models. IR researchers have long ago recognized the importance o fa sound evaluation of new approaches. Yet, the development of evaluation techniques for conversational search is still an underlooked problem. Currently, most evaluation approaches rely on procedures directly drawn from ad-hoc search evaluation, treating utterances in a conversation as independent events, as if they were just separate topics, instead of accounting for the conversation context. We overcome this issue by proposing a framework for defining evaluation measures that are aware of the conversation context and the utterance semantic dependencies. In particular, we model the conversations as Direct Acyclic Graphs (DAG), where self-explanatory utterances are root nodes, while anaphoric utterances are linked to sentences that contain their missing semantic information. Then,we propose a family of hierarchical dependence-aware aggregations of the evaluation metrics driven by the conversational graph. In our experiments, we show that utterances from the same conversation are 20% more correlated than utterances from different conversations. Thanks to the proposed framework, we are able to include such correlation in our aggregations, and be more accurate when determining which pairs of conversational systems are deemed significantly different.
Guglielmo Faggioli, Marco Ferrante, Nicola Ferro 0001, Raffaele Perego 0001, Nicola Tonellotto
SIGIR2
2020 How do interval scales help us with better understanding IR evaluation measures?
Marco Ferrante, Nicola Ferro 0001, Eleonora Losiouk
Inf. Retr. J.1
2019 A Markovian Approach to Evaluate Session-Based IR Systems
David van Dijk, Marco Ferrante, Nicola Ferro 0001, Evangelos Kanoulas
ECIR (1)2
2019 Stochastic Relevance for Crowdsourcing
Marco Ferrante, Nicola Ferro 0001, Eleonora Losiouk
ECIR (1)1
2019 A General Theory of IR Evaluation Measures
abstract
Interval scales are assumed by several basic descriptive statistics, such as mean and variance, and by many statistical significance tests which are daily used in IR to compare systems. Unfortunately, so far, there has not been any systematic and formal study to discover the actual scale properties of IR measures. Therefore, in this paper, we develop a theory ofInformation Retrieval (IR)evaluation measures, based on the representational theory of measurements, to determine whether and when IR measures are interval scales. We found that common set-based retrieval measures—namely Precision, Recall, and F-measure—always are interval scales in the case of binary relevance while this happens also in the case of multi-graded relevance only when the relevance degrees themselves are on a ratio scale and we define a specific partial order among systems. In the case of rank-based retrieval measures—namely AP, gRBP, DCG, and ERR—only gRPB is an interval scale when we choose a specific value of the parameter$p$and define a specific total order among systems while all the other IR measures are not interval scales. Besides the formal framework itself and the proof of the scale properties of several commonly used IR measures, the paper also defines some brand new set-based and rank-based IR evaluation measures which ensure to be interval scales.
Marco Ferrante, Nicola Ferro 0001, Silvia Pontarollo
IEEE Trans. Knowl. Data Eng.1
2018 Modelling Randomness in Relevance Judgments and Evaluation Measures
Marco Ferrante, Nicola Ferro 0001, Silvia Pontarollo
ECIR1
2017 You Surf so Strange Today: Anomaly Detection in Web Services via HMM and CTMC
Maddalena Favaretto, Riccardo Spolaor, Mauro Conti, Marco Ferrante
GPC4
2017 AWARE: Exploiting Evaluation Measures to Combine Multiple Assessors
abstract
We propose theAssessor-driven Weighted Averages for Retrieval Evaluation (AWARE)probabilistic framework, a novel methodology for dealing with multiple crowd assessors that may be contradictory and/or noisy. By modeling relevance judgements and crowd assessors as sources of uncertainty, AWARE takes the expectation of a generic performance measure, like Average Precision, composed with these random variables. In this way, it approaches the problem of aggregating different crowd assessors from a new perspective, that is, directly combining the performance measures computed on the ground truth generated by the crowd assessors instead of adopting some classification technique to merge the labels produced by them. We propose several unsupervised estimators that instantiate the AWARE framework and we compare them with state-of-the-art approaches, that is,Majoriity Vote and Expectation Maximization, on TREC collections. We found that AWARE approaches improve in terms of their capability of correctly ranking systems and predicting their actual performance scores.
Marco Ferrante, Nicola Ferro 0001, Maria Maistro
ACM Trans. Inf. Syst.1
2015 No-Free-Lunch theorems in the continuum
Aureli Alabert, Alessandro Berti 0001, Ricard Caballero, Marco Ferrante
Theor. Comput. Sci.4
2014 Injecting user models and time into precision via Markov chains
abstract
We propose a family of new evaluation measures, called Markov Precision (MP), which exploits continuous-time and discrete-time Markov chains in order to inject user models into precision. Continuous-time MP behaves like time-calibrated measures, bringing the time spent by the user into the evaluation of a system; discrete-time MP behaves like traditional evaluation measures. Being part of the same Markovian framework, the time-based and rank-based versions of MP produce values that are directly comparable.
Marco Ferrante, Nicola Ferro 0001, Maria Maistro
SIGIR1