VLDB 2026 Research / reviewers in the wild / expert
Damiano Spina
dblp:74/2824
· DBLP profile ↗
50ranked-venue papers in the field
6as first author
24since 2021 · last 2026
0000-0001-9913-433XORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 48 (6 first)Data Mining & Knowledge Discovery · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Characterizing Personality from Eye-Tracking: The Role of Gaze and Its Absence in Interactive Search EnvironmentsabstractPersonality traits influence how individuals engage, behave, and make decisions during the information-seeking process. However, few studies have linked personality to observable search behaviors. This study aims to characterize personality traits through a multimodal time-series model that integrates eye-tracking data and gaze missingness–periods when the user’s gaze is not captured. This approach is based on the idea that people often look away when they think, signaling disengagement or reflection. We conducted a user study with 25 participants, who used an interactive application on an iPad, allowing them to engage with digital artifacts from a museum. We rely on raw gaze data from an eye tracker, minimizing preprocessing so that behavioral patterns can be preserved without substantial data cleaning. From this perspective, we trained models to predict personality traits using gaze signals. Our results from a five-fold cross-validation study demonstrate strong predictive performance across all five dimensions: Neuroticism (Macro F1 = 77.69%), Conscientiousness (74.52%), Openness (77.52%), Agreeableness (73.09%), and Extraversion (76.69%). The ablation study examines whether the absence of gaze information affects the model performance, demonstrating that incorporating missingness improves multimodal time-series modeling. The full model, which integrates both time-series signals and missingness information, achieves 10-15% higher accuracy and macro F1 scores across all Big Five traits compared to the model without time-series signals and missingness. These findings provide evidence that personality can be inferred from search-related gaze behavior and demonstrate the value of incorporating missing gaze data into time-series multimodal modeling. Jiaman He, Marta Micheli, Damiano Spina, Dana McKay, Johanne R. Trippas, Noriko Kando |
CHIIR | 3 |
| 2026 | EXIST 2026: Physiological Data for Multimodal Sexism Characterization in Social Media
Laura Plaza, Jorge Carrillo de Albornoz, Elena Gomis-Vicent, Iván Árcos, María Aloy-Mayo, Paolo Rosso, Damiano Spina |
ECIR (4) | 7 |
| 2026 | VulGen: Workshop on Vulnerabilities in Generative Systems for Information RetrievalabstractGenerative systems are rapidly transforming both academic research and industrial practices. These systems are increasingly integrated into information access and information retrieval (IR) tasks and continue to evolve at a substantial pace. Integrating these models into daily workflows exposes critical vulnerabilities, including adversarial attacks, inherent biases, and negative impacts on user behavior, which can lead to suboptimal or even detrimental outcomes. The VulGen workshop at SIGIR 2026 brings together the IR community and related disciplines (e.g., cyber security) to map this evolving landscape. Through a full day of structured discussion and engagement, we aim to synthesize the current state of research and identify new avenues for future investigation. Information about VulGen is hosted at: https://vulgen-workshop.github.io/SIGIR2026/. Shuoqi Sun, Sara Allawati, Laura Dietz, Madhurima Khirbat, Bhaskar Mitra 0001, Maarten de Rijke, Damiano Spina |
SIGIR | 7 |
| 2026 | Diversification and Fairness in Search: Two Sides of the Same Coin?abstractInformation retrieval systems aim to return relevant and useful content to users and are often biased towards popular items. This implies that an under-represented group or attribute will not receive a fair share of a user’s attention in search results. For example, while a ranked results list for a query such as ‘physicists’ might be fair according to a particular attribute such as gender, nationality or social group, it might not be fair for all of them. Ideally, while providing relevant answers, a results list should also provide fair exposure across a broad range of attributes. We demonstrate that while a system can be fair towards multiple attributes, they are not necessarily diverse (i.e., redundancy/minimal novelty). To this end, we include an additional dimension to the study, i.e., diversity, and explore the relationship between fairness and diversity measures by exploring popular search result diversification techniques using the test collections from TREC 2021 Fair Ranking Track, TREC 2022 Fair Ranking Track and NTCIR-17 FairWeb-1. Furthermore, we study the impact of such diversification techniques along both nominal and ordinal attributes, as well as for intersectional fairness. Our results indicate that explicit search results diversification techniques showed improved results when the attributes were nominal but failed to provide fairer and more diverse results when the attributes were ordinal in nature. Additionally, in terms of intersectional fairness explicit search results diversification also performed significantly better than baseline retrieval runs. Sachin Pathiyan Cherumanal, Falk Scholer, Damiano Spina |
ACM Trans. Inf. Syst. | 3 |
| 2025 | NeuroPhysIIR: International Workshop on NeuroPhysiological Approaches for Interactive Information RetrievalabstractThe International Workshop on NeuroPhysiological Approaches for Interactive Information Retrieval (NeuroPhysIIR'25) aims to bringing together researchers from information science, humancomputer interaction, cognitive neuroscience, and related fields, to foster cross-disciplinary collaboration and accelerate progress in neurophysiologically-informed IIR research.As the third edition following successful workshops at SIGIR'15 [5] and CHIIR'17 [6], we anticipate that the interactive nature of this workshop will not only raise awareness but also lower the entry barriers for engaging with this exciting research area within the wider IIR community.Workshop website: https://neurophysiir.github.io/chiir2025/. Jacek Gwizdka, Javed Mostafa, Min Zhang 0006, Kaixin Ji, Yashar Moshfeghi, Tuukka Ruotsalo, Damiano Spina |
CHIIR | 7 |
| 2025 | Two Heads Are Better Than One: Improving Search Effectiveness Through LLM-Generated Query Variants
Kun Ran, Marwah Alaofi, Mark Sanderson, Damiano Spina |
CHIIR | 4 |
| 2025 | Responsible AI From the Lens of an Information Retrieval Researcher: A Hands-On TutorialabstractWhile the concept of responsible AI is becoming more and more popular, practitioners and researchers may often struggle to characterize responsible practices in their own work.Using a hands-on approach -where participants are invited to discuss terminology from their own perspectives -this tutorial aims to illustrate the application of responsible practices in Information Retrieval (IR).Using case studies based on existing IR research, the tutorial will explore responsible AI concepts such as positionality, participatory research, fairness, diversity, and ethics. Damiano Spina |
CHIIR | 1 |
| 2025 | EXIST 2025: Learning with Disagreement for Sexism Identification and Characterization in Tweets, Memes, and TikTok Videos
Laura Plaza, Jorge Carrillo de Albornoz, Iván Árcos, Paolo Rosso, Damiano Spina, Enrique Amigó, Julio Gonzalo 0001, Roser Morante |
ECIR (5) | 5 |
| 2025 | Characterising Topic Familiarity and Query Specificity Using Eye-Tracking DataabstractEye-tracking data has been shown to correlate with a user's knowledge level and query formulation behaviour. While previous work has focused primarily on eye gaze fixations for attention analysis, often requiring additional contextual information, our study investigates the memory-related cognitive dimension by relying solely on pupil dilation and gaze velocity to infer users' topic familiarity and query specificity without needing any contextual information. Using eye-tracking data collected via a lab user study N=18, we achieved a Macro F1 score of 71.25% for predicting topic familiarity with a Gradient Boosting classifier, and a Macro F1 score of 60.54% with a k-nearest neighbours (KNN) classifier for query specificity. Furthermore, we developed a novel annotation guideline - specifically tailored for question answering - to manually classify queries as Specific or Non-specific. This study demonstrates the feasibility of eye-tracking to better understand topic familiarity and query specificity in search. Jiaman He, Zikang Leng, Dana McKay, Johanne R. Trippas, Damiano Spina |
SIGIR | 5 |
| 2024 | Walert: Putting Conversational Information Seeking Knowledge into Action by Building and Evaluating a Large Language Model-Powered ChatbotabstractCreating and deploying customized applications is crucial for operational success and enriching user experiences in the rapidly evolving modern business world. A prominent facet of modern user experiences is the integration of chatbots or voice assistants. The rapid evolution of Large Language Models (LLMs) has provided a powerful tool to build conversational applications. We present Walert, a customized LLM-based conversational agent able to answer frequently asked questions about computer science degrees and programs at RMIT University. Our demo aims to showcase how conversational information-seeking researchers can effectively communicate the benefits of using best practices to stakeholders interested in developing and deploying LLM-based chatbots. These practices are well-known in our community but often overlooked by practitioners who may not have access to this knowledge. The methodology and resources used in this demo serve as a bridge to facilitate knowledge transfer from experts, address industry professionals’ practical needs, and foster a collaborative environment. The data and code of the demo are available at https://github.com/rmit-ir/walert. Sachin Pathiyan Cherumanal, Futoon M. Abushaqra, Angel Felipe Magnossão de Paula, Kaixin Ji, Halil Ali, Danula Hettiachchi, Johanne R. Trippas, Falk Scholer, Damiano Spina |
CHIIR | 10 |
| 2024 | EXIST 2024: sEXism Identification in Social neTworks and Memes
Laura Plaza, Jorge Carrillo de Albornoz, Enrique Amigó, Julio Gonzalo 0001, Roser Morante, Paolo Rosso, Damiano Spina, Berta Chulvi, Alba Maeso, Víctor Ruiz |
ECIR (5) | 7 |
| 2024 | Characterizing Information Seeking Processes with Multiple Physiological SignalsabstractInformation access systems are getting complex, and our understanding of user behavior during information seeking processes is mainly drawn from qualitative methods, such as observational studies or surveys. Leveraging the advances in sensing technologies, our study aims to characterize user behaviors with physiological signals, particularly in relation to cognitive load, affective arousal, and valence. We conduct a controlled lab study with 26 participants, and collect data including Electrodermal Activities, Photoplethysmogram, Electroencephalogram, and Pupillary Responses. This study examines informational search with four stages: the realization of Information Need (IN), Query Formulation (QF), Query Submission (QS), and Relevance Judgment (RJ). We also include different interaction modalities to represent modern systems, e.g., QS by text-typing or verbalizing, and RJ with text or audio information. We analyze the physiological signals across these stages and report outcomes of pairwise non-parametric repeated-measure statistical tests. The results show that participants experience significantly higher cognitive loads at IN with a subtle increase in alertness, while QF requires higher attention. QS involves demanding cognitive loads than QF. Affective responses are more pronounced at RJ than QS or IN, suggesting greater interest and engagement as knowledge gaps are resolved. To the best of our knowledge, this is the first study that explores user behaviors in a search process employing a more nuanced quantitative analysis of physiological signals. Our findings offer valuable insights into user behavior and emotional responses in information seeking processes. We believe our proposed methodology can inform the characterization of more complex processes, such as conversational information seeking. Kaixin Ji, Danula Hettiachchi, Flora D. Salim, Falk Scholer, Damiano Spina |
SIGIR | 5 |
| 2024 | Explainability for Transparent Conversational Information-SeekingabstractThe increasing reliance on digital information necessitates advancements in conversational search systems, particularly in terms of information transparency. While prior research in conversational information-seeking has concentrated on improving retrieval techniques, the challenge remains in generating responses useful from a user perspective. This study explores different methods of explaining the responses, hypothesizing that transparency about the source of the information, system confidence, and limitations can enhance users' ability to objectively assess the response. By exploring transparency across explanation type, quality, and presentation mode, this research aims to bridge the gap between system-generated responses and responses verifiable by the user. We design a user study to answer questions concerning the impact of (1) the quality of explanations enhancing the response on its usefulness and (2) ways of presenting explanations to users. The analysis of the collected data reveals lower user ratings for noisy explanations, although these scores seem insensitive to the quality of the response. Inconclusive results on the explanations presentation format suggest that it may not be a critical factor in this setting. Weronika Lajewska, Damiano Spina, Johanne R. Trippas, Krisztian Balog |
SIGIR | 2 |
| 2024 | Crowdsourced Fact-checking: Does It Actually Work?abstractThere is an important ongoing effort aimed to tackle misinformation and to perform reliable fact-checking by employing human assessors at scale, with a crowdsourcing-based approach. Previous studies on the feasibility of employing crowdsourcing for the task of misinformation detection have provided inconsistent results: some of them seem to confirm the effectiveness of crowdsourcing for assessing the truthfulness of statements and claims, whereas others fail to reach an effectiveness level higher than automatic machine learning approaches, which are still unsatisfactory. In this paper, we aim at addressing such inconsistency and understand if truthfulness assessment can indeed be crowdsourced effectively. To do so, we build on top of previous studies; we select some of those reporting low effectiveness levels, we highlight their potential limitations, and we then reproduce their work attempting to improve their setup to address those limitations. We employ various approaches, data quality levels, and agreement measures to assess the reliability of crowd workers when assessing the truthfulness of (mis)information. Furthermore, we explore different worker features and compare the results obtained with different crowds. According to our findings, crowdsourcing can be used as an effective methodology to tackle misinformation at scale. When compared to previous studies, our results indicate that a significantly higher agreement between crowd workers and experts can be obtained by using a different, higher-quality, crowdsourcing platform and by improving the design of the crowdsourcing task. Also, we find differences concerning task and worker features and how workers provide truthfulness assessments. David La Barbera, Eddy Maddalena, Michael Soprano, Kevin Roitero, Gianluca Demartini, Davide Ceolin, Damiano Spina, Stefano Mizzaro |
Inf. Process. Manag. | 7 |
| 2024 | Cognitive Biases in Fact-Checking and Their Countermeasures: A ReviewabstractThe increase of the amount of misinformation spread every day online is a huge threat to the society. Organizations and researchers are working to contrast this misinformation plague. In this setting, human assessors are indispensable to correctly identify, assess and/or revise the truthfulness of information items, i.e., to perform the fact-checking activity. Assessors, as humans, are subject to systematic errors that might interfere with their fact-checking activity. Among such errors, cognitive biases are those due to the limits of human cognition. Although biases help to minimize the cost of making mistakes, they skew assessments away from an objective perception of information. Cognitive biases, hence, are particularly frequent and critical, and can cause errors that have a huge potential impact as they propagate not only in the community, but also in the datasets used to train automatic and semi-automatic machine learning models to fight misinformation. In this work, we present a review of the cognitive biases which might occur during the fact-checking process. In more detail, inspired by PRISMA – a methodology used for systematic literature reviews – we manually derive a list of 221 cognitive biases that may affect human assessors. Then, we select the 39 biases that might manifest during the fact-checking process, we group them into categories, and we provide a description. Finally, we present a list of 11 countermeasures that can be adopted by researchers, practitioners, and organizations to limit the effect of the identified cognitive biases on the fact-checking activity. Michael Soprano, Kevin Roitero, David La Barbera, Davide Ceolin, Damiano Spina, Gianluca Demartini, Stefano Mizzaro |
Inf. Process. Manag. | 5 |
| 2023 | Designing and Evaluating Presentation Strategies for Fact-Checked ContentabstractWith the rapid growth of online misinformation, it is crucial to have reliable fact-checking methods. Recent research on finding check-worthy claims and automated fact-checking have made significant advancements. However, limited guidance exists regarding the presentation of fact-checked content to effectively convey verified information to users. We address this research gap by exploring the critical design elements in fact-checking reports and investigating whether credibility and presentation-based design improvements can enhance users' ability to interpret the report accurately. We co-developed potential content presentation strategies through a workshop involving fact-checking professionals, communication experts, and researchers. The workshop examined the significance and utility of elements such as veracity indicators and explored the feasibility of incorporating interactive components for enhanced information disclosure. Building on the workshop outcomes, we conducted an online experiment involving 76 crowd workers to assess the efficacy of different design strategies. The results indicate that proposed strategies significantly improve users' ability to accurately interpret the verdict of fact-checking articles. Our findings underscore the critical role of effective presentation of fact reports in addressing the spread of misinformation. By adopting appropriate design enhancements, the effectiveness of fact-checking reports can be maximized, enabling users to make informed judgments. Danula Hettiachchi, Kaixin Ji, Jenny Kennedy, Anthony McCosker, Flora D. Salim, Mark Sanderson, Falk Scholer, Damiano Spina |
CIKM | 8 |
| 2023 | Overview of EXIST 2023: sEXism Identification in Social NeTworks
Laura Plaza, Jorge Carrillo de Albornoz, Roser Morante, Enrique Amigó, Julio Gonzalo 0001, Damiano Spina, Paolo Rosso |
ECIR (3) | 6 |
| 2023 | Examining the Impact of Uncontrolled Variables on Physiological Signals in User Studies for Information Processing ActivitiesabstractPhysiological signals can potentially be applied as objective measures to understand the behavior and engagement of users interacting with information access systems. However, the signals are highly sensitive, and many controls are required in laboratory user studies. To investigate the extent to which controlled or uncontrolled (i.e., confounding) variables such as task sequence or duration influence the observed signals, we conducted a pilot study where each participant completed four types of information-processing activities (READ, LISTEN, SPEAK, and WRITE). Meanwhile, we collected data on blood volume pulse, electrodermal activity, and pupil responses. We then used machine learning approaches as a mechanism to examine the influence of controlled and uncontrolled variables that commonly arise in user studies. Task duration was found to have a substantial effect on the model performance, suggesting it represents individual differences rather than giving insight into the target variables. This work contributes to our understanding of such variables in using physiological signals in information retrieval user studies. Kaixin Ji, Damiano Spina, Danula Hettiachchi, Flora D. Salim, Falk Scholer |
SIGIR | 2 |
| 2023 | i-Align: an interpretable knowledge graph alignment modelabstractAbstract Knowledge graphs (KGs) are becoming essential resources for many downstream applications. However, their incompleteness may limit their potential. Thus, continuous curation is needed to mitigate this problem. One of the strategies to address this problem is KG alignment, i.e., forming a more complete KG by merging two or more KGs. This paper proposes i-Align, an interpretable KG alignment model. Unlike the existing KG alignment models, i-Align provides an explanation for each alignment prediction while maintaining high alignment performance. Experts can use the explanation to check the correctness of the alignment prediction. Thus, the high quality of a KG can be maintained during the curation process (e.g., the merging process of two KGs). To this end, a novel Transformer-based Graph Encoder (Trans-GE) is proposed as a key component of i-Align for aggregating information from entities’ neighbors (structures). Trans-GE uses Edge-gated Attention that combines the adjacency matrix and the self-attention matrix to learn a gating mechanism to control the information aggregation from the neighboring entities. It also uses historical embeddings, allowing Trans-GE to be trained over mini-batches, or smaller sub-graphs, to address the scalability issue when encoding a large KG. Another component of i-Align is a Transformer encoder for aggregating entities’ attributes. This way, i-Align can generate explanations in the form of a set of the most influential attributes/neighbors based on attention weights. Extensive experiments are conducted to show the power of i-Align. The experiments include several aspects, such as the model’s effectiveness for aligning KGs, the quality of the generated explanations, and its practicality for aligning large KGs. The results show the effectiveness of i-Align in these aspects. Bayu Distiawan Trisedya, Flora D. Salim, Jeffrey Chan, Damiano Spina, Falk Scholer, Mark Sanderson |
Data Min. Knowl. Discov. | 4 |
| 2022 | Where Do Queries Come From?abstractWhere do queries -- the words searchers type into a search box -- come from? The Information Retrieval community understands the performance of queries and search engines extensively, and has recently begun to examine the impact of query variation, showing that different queries for the same information need produce different results. In an information environment where bad actors try to nudge searchers toward misinformation, this is worrisome. The source of query variation -- searcher characteristics, contextual or linguistic prompts, cognitive biases, or even the influence of external parties -- while studied in a piecemeal fashion by other research communities has not been studied by ours. In this paper we draw on a variety of literatures (including information seeking, psychology, and misinformation), and report some small experiments to describe what is known about where queries come from, and demonstrate a clear literature gap around the source of query variations in IR. We chart a way forward for IR to research, document and understand this important question, with a view to creating search engines that provide more consistent, accurate and relevant search results regardless of the searcher's framing of the query. Marwah Alaofi, Luke Gallagher, Dana McKay, Lauren L. Saling, Mark Sanderson, Falk Scholer, Damiano Spina, Ryen W. White |
SIGIR | 7 |
| 2022 | Ranking Interruptus: When Truncated Rankings Are Better and How to Measure ThatabstractMost of information retrieval effectiveness evaluation metrics assume that systems appending irrelevant documents at the bottom of the ranking are as effective as (or not worse than) systems that have a stopping criteria to 'truncate' the ranking at the right position to avoid retrieving those irrelevant documents at the end. It can be argued, however, that such truncated rankings are more useful to the end user. It is thus important to understand how to measure retrieval effectiveness in this scenario. In this paper we provide both theoretical and experimental contributions. We first define formal properties to analyze how effectiveness metrics behave when evaluating truncated rankings. Our theoretical analysis shows that de-facto standard metrics do not satisfy desirable properties to evaluate truncated rankings: only Observational Information Effectiveness (OIE) -- a metric based on Shannon's information theory -- satisfies them all. We then perform experiments to compare several metrics on nine TREC datasets. According to our experimental results, the most appropriate metrics for truncated rankings are OIE and a novel extension of Rank-Biased Precision that adds a user effort factor penalizing the retrieval of irrelevant documents. Enrique Amigó, Stefano Mizzaro, Damiano Spina |
SIGIR | 3 |
| 2022 | Component-based Analysis of Dynamic Search PerformanceabstractIn many search scenarios, such as exploratory, comparative, or survey-oriented search, users interact with dynamic search systems to satisfy multi-aspect information needs. These systems utilize different dynamic approaches that exploit various user feedback granularity types. Although studies have provided insights about the role of many components of these systems, they used black-box and isolated experimental setups. Therefore, the effects of these components or their interactions are still not well understood. We address this by following a methodology based on Analysis of Variance (ANOVA). We built a Grid Of Points that consists of systems based on different ways to instantiate three components: initial rankers, dynamic rerankers, and user feedback granularity. Using evaluation scores based on the TREC Dynamic Domain collections, we built several ANOVA models to estimate the effects. We found that (i) although all components significantly affect search effectiveness, the initial ranker has the largest effective size, (ii) the effect sizes of these components vary based on the length of the search session and the used effectiveness metric, and (iii) initial rankers and dynamic rerankers have more prominent effects than user feedback granularity. To improve effectiveness, we recommend improving the quality of initial rankers and dynamic rerankers. This does not require eliciting detailed user feedback, which might be expensive or invasive. Ameer Albahem, Damiano Spina, Falk Scholer, Lawrence Cavedon |
ACM Trans. Inf. Syst. | 2 |
| 2021 | Evaluating Fairness in Argument RetrievalabstractExisting commercial search engines often struggle to represent different perspectives of a search query. Argument retrieval systems address this limitation of search engines and provide both positive (PRO) and negative (CON) perspectives about a user's information need on a controversial topic (e.g., climate change). The effectiveness of such argument retrieval systems is typically evaluated based on topical relevance and argument quality, without taking into account the often differing number of documents shown for the argument stances (PRO or CON). Therefore, systems may retrieve relevant passages, but with a biased exposure of arguments. In this work, we analyze a range of non-stochastic fairness-aware ranking and diversity metrics to evaluate the extent to which argument stances are fairly exposed in argument retrieval systems. Sachin Pathiyan Cherumanal, Damiano Spina, Falk Scholer, W. Bruce Croft |
CIKM | 2 |
| 2021 | The many dimensions of truthfulness: Crowdsourcing misinformation assessments on a multidimensional scale
Michael Soprano, Kevin Roitero, David La Barbera, Davide Ceolin, Damiano Spina, Stefano Mizzaro, Gianluca Demartini |
Inf. Process. Manag. | 5 |
| 2020 | Third International Workshop on Conversational Approaches to Information Retrieval (CAIR'20): Full-day Workshop at CHIIR 2020abstractThe third CAIR workshop brings together researchers and developers interested in advancing conversational systems in interactive information retrieval. The workshop builds on the first and second CAIR workshops held at SIGIR 2017 and 2018 and will focus on the continuing development of current challenges, user and system limitations, and evaluation of conversational systems for information retrieval. Participants will collaboratively explore different contexts (i.e., home, hospitals, or work settings), use cases, and interactivity forms (voice-only, multi-modal, screen-based) in which conversational search systems can be used. Possible outcomes include fostering novel and innovative methodologies (such as for data collection and evaluation), personalising conversational systems, and understanding ethical challenges---such as system transparency---from the user's perspective. Johanne R. Trippas, Paul Thomas 0001, Damiano Spina, Hideo Joho |
CHIIR | 3 |
| 2020 | The COVID-19 Infodemic: Can the Crowd Judge Recent Misinformation Objectively?abstractMisinformation is an ever increasing problem that is difficult to solve for the research community and has a negative impact on the society at large. Very recently, the problem has been addressed with a crowdsourcing-based approach to scale up labeling efforts: to assess the truthfulness of a statement, instead of relying on a few experts, a crowd of (non-expert) judges is exploited. We follow the same approach to study whether crowdsourcing is an effective and reliable method to assess statements truthfulness during a pandemic. We specifically target statements related to the COVID-19 health emergency, that is still ongoing at the time of the study and has arguably caused an increase of the amount of misinformation that is spreading online (a phenomenon for which the term "infodemic" has been used). By doing so, we are able to address (mis)information that is both related to a sensitive and personal issue like health and very recent as compared to when the judgment is done: two issues that have not been analyzed in related work.\n\nIn our experiment, crowd workers are asked to assess the truthfulness of statements, as well as to provide evidence for the assessments as a URL and a text justification. Besides showing that the crowd is able to accurately judge the truthfulness of the statements, we also report results on many different aspects, including: agreement among workers, the effect of different aggregation functions, of scales transformations, and of workers background / bias. We also analyze workers behavior, in terms of queries submitted, URLs found / selected, text justifications, and other behavioral data like clicks and mouse actions collected by means of an ad hoc logger. Kevin Roitero, Michael Soprano, Beatrice Portelli, Damiano Spina, Vincenzo Della Mea, Giuseppe Serra 0001, Stefano Mizzaro, Gianluca Demartini |
CIKM | 4 |
| 2020 | Watch 'n' Check: Towards a Social Media Monitoring Tool to Assist Fact-Checking ExpertsabstractWe present an ongoing collaboration between computer science researchers and fact-checking experts in a broad-cast corporation to develop Watch 'n' Check, a social media monitoring tool that assists fact-checkers to detect and target misinformation online. The lean methodology followed in our collaboration has helped us to better understand how information access tools can assist fact-checking experts. We report initial results and discuss our plan for further development, as well as the open challenges identified so far. Assunta Cerone, Elham Naghizade, Falk Scholer, Devi Mallal, Russell Skelton, Damiano Spina |
DSAA | 6 |
| 2020 | Crowdsourcing Truthfulness: The Impact of Judgment Scale and Assessor Bias
David La Barbera, Kevin Roitero, Gianluca Demartini, Stefano Mizzaro, Damiano Spina |
ECIR (2) | 5 |
| 2020 | Intelligent Task Recognition: Towards Enabling Productivity Assistance in Daily LifeabstractWe introduce the novel research problem of task recognition in daily life. We recognize tasks such as project management, planning, meal-breaks, communication, documentation, and family care. We capture Cyber, Physical, and Social (CPS) activities of 17 participants over four weeks using device-based sensing, app activity logging, and an experience sampling methodology. Our cohort includes students, casual workers, and professionals, forming the first real-world context-rich task behaviour dataset. We model CPS activities across different task categories, results highlight the importance of considering the CPS feature sets in modelling, especially work-related tasks. Jonathan Liono, Mohammad Saiedur Rahaman, Flora D. Salim, Yongli Ren, Damiano Spina, Falk Scholer, Johanne R. Trippas, Mark Sanderson, Paul N. Bennett, Ryen W. White |
ICMR | 5 |
| 2020 | Can The Crowd Identify Misinformation Objectively?: The Effects of Judgment Scale and Assessor's BackgroundabstractTruthfulness judgments are a fundamental step in the process of fighting misinformation, as they are crucial to train and evaluate classifiers that automatically distinguish true and false statements. Usually such judgments are made by experts, like journalists for political statements or medical doctors for medical statements. In this paper, we follow a different approach and rely on (non-expert) crowd workers. This of course leads to the following research question: Can crowdsourcing be reliably used to assess the truthfulness of information and to create large-scale labeled collections for information credibility systems? To address this issue, we present the results of an extensive study based on crowdsourcing: we collect thousands of truthfulness assessments over two datasets, and we compare expert judgments with crowd judgments, expressed on scales with various granularity levels. We also measure the political bias and the cognitive background of the workers, and quantify their effect on the reliability of the data provided by the crowd. Kevin Roitero, Michael Soprano, Shaoyang Fan, Damiano Spina, Stefano Mizzaro, Gianluca Demartini |
SIGIR | 4 |
| 2020 | Towards a model for spoken conversational search
Johanne R. Trippas, Damiano Spina, Paul Thomas 0001, Mark Sanderson, Hideo Joho, Lawrence Cavedon |
Inf. Process. Manag. | 2 |
| 2019 | Learning About Work Tasks to Inform Intelligent Assistant DesignabstractIntelligent assistants can serve many purposes, including entertainment (e.g. playing music), home automation, and task management (e.g. timers, reminders). The role of these assistants is evolving to also support people engaged in work tasks, in workplaces and beyond. To design truly useful intelligent assistants for work, it is important to better understand the work tasks that people are performing. Based on a survey of 401 respondents' daily tasks and activities in a work setting, we present a classification of work-related tasks, and analyze their key characteristics, including the frequency of their self-reported tasks, the environment in which they undertake the tasks, and which, if any, electronic devices are used. We also investigate the cyber, physical, and social aspects of tasks. Finally, we reflect on how intelligent assistants could influence and help people in a work environment to complete their tasks, and synthesize our findings to provide insight on the future of intelligent assistants in support of amplifying personal productivity. Johanne R. Trippas, Damiano Spina, Falk Scholer, Ahmed Awadallah 0001, Peter Bailey, Paul N. Bennett, Ryen W. White, Jonathan Liono, Yongli Ren, Flora D. Salim, Mark Sanderson |
CHIIR | 2 |
| 2019 | Investigating the Learning Process in Job Search: A Longitudinal StudyabstractWe investigated the learning process in search by conducting a log-based study involving registered job seekers of a commercial job search engine. The analysis shows that job search is a complex task: seekers usually submit multiple queries over sessions that can last days or even weeks. We find that querying, clicking, and job application rates change over time: job seekers tend to use more filters and a less diverse set of query terms. In terms of click and application behavior, we observed a significant decrease in click rate and query term diversity, as well as an increase in application rates. These trends are found to largely match information seeking models of learning in a complex search task. However, common behaviors are observed in the logs that suggest the existing models may not be sufficient to describe all of the users' learning and seeking processes. Jiaxin Mao, Damiano Spina, Seyedeh Sargol Sadeghi, Falk Scholer, Mark Sanderson |
CIKM | 2 |
| 2019 | Meta-evaluation of Dynamic Search: How Do Metrics Capture Topical Relevance, Diversity and User Effort?
Ameer Albahem, Damiano Spina, Falk Scholer, Lawrence Cavedon |
ECIR (1) | 2 |
| 2019 | Processing social media in real-time
Damiano Spina, Arkaitz Zubiaga, Amit P. Sheth, Markus Strohmaier |
Inf. Process. Manag. | 1 |
| 2019 | A comparison of filtering evaluation metrics based on formal constraints
Enrique Amigó, Julio Gonzalo 0001, M. Felisa Verdejo, Damiano Spina |
Inf. Retr. J. | 4 |
| 2018 | Informing the Design of Spoken Conversational Search: Perspective PaperabstractWe conducted a laboratory-based observational study where pairs of people performed search tasks communicating verbally. Examination of the discourse allowed commonly used interactions to be identified for Spoken Conversational Search (SCS). We compared the interactions to existing models of search behaviour. We find that SCS is more complex and interactive than traditional search. This work enhances our understanding of different search behaviours and proposes research opportunities for an audio-only search system. Future work will focus on creating models of search behaviour for SCS and evaluating these against actual SCS systems. Johanne R. Trippas, Damiano Spina, Lawrence Cavedon, Hideo Joho, Mark Sanderson |
CHIIR | 2 |
| 2018 | An Axiomatic Analysis of Diversity Evaluation Metrics: Introducing the Rank-Biased Utility MetricabstractMany evaluation metrics have been defined to evaluate the effectiveness ad-hoc retrieval and search result diversification systems. However, it is often unclear which evaluation metric should be used to analyze the performance of retrieval systems given a specific task. Axiomatic analysis is an informative mechanism to understand the fundamentals of metrics and their suitability for particular scenarios. In this paper, we define a constraint-based axiomatic framework to study the suitability of existing metrics in search result diversification scenarios. The analysis informed the definition of Rank-Biased Utility (RBU) -- an adaptation of the well-known Rank-Biased Precision metric -- that takes into account redundancy and the user effort associated to the inspection of documents in the ranking. Our experiments over standard diversity evaluation campaigns show that the proposed metric captures quality criteria reflected by different metrics, being suitable in the absence of knowledge about particular features of the scenario under study. Enrique Amigó, Damiano Spina, Jorge Carrillo de Albornoz |
SIGIR | 2 |
| 2018 | Second International Workshop on Conversational Approaches to Information Retrieval (CAIR'18): Workshop at SIGIR 2018abstractThe CAIR'18 workshop will bring together academic and industrial researchers to create a forum for research on conversational approaches to search and recommendation. A specific focus will be on techniques that support complex and multi-turn user-machine dialogues for information access and retrieval, and multi-modal interfaces for interacting with such systems. Jaime Arguello, Filip Radlinski, Hideo Joho, Damiano Spina, Julia Kiseleva |
SIGIR | 4 |
| 2018 | A Living Lab Study of Query Amendment in Job SearchabstractErrors in formulation of queries made by users can lead to poor search results pages. We performed a living lab study using online A/B testing to measure the degree of improvement achieved with a query amendment technique when applied to a commercial job search engine. Of particular interest in this case study is a clear 'success' signal, namely, the number of job applications lodged by a user as a result of querying the service. A set of 276 queries was identified for amendment in four different categories through the use of word embeddings, with large gains in conversion rates being attained in all four of those categories. Our analysis of query reformulations also provides a better understanding of user satisfaction in the case of problematic queries (ones with fewer results than fill a single page) by observing that users tend to reformulate rewritten queries less. Bahar Salehi, Damiano Spina, Alistair Moffat, Seyedeh Sargol Sadeghi, Falk Scholer, Timothy Baldwin, Lawrence Cavedon, Mark Sanderson, Wilson Wong, Justin Zobel |
SIGIR | 2 |
| 2017 | How Do People Interact in Conversational Speech-Only Search Tasks: A Preliminary AnalysisabstractWe present preliminary findings from a study of mixed initiative conversational behaviour for informational search in an acoustic setting. The aim of the observational study is to reveal insights into how users would conduct searches over voice where a screen is absent but where users are able to converse interactively with the search system. We conducted alaboratory-based observational study of 13 pairs of participants each completing three search tasks with different cognitive complexity levels. The communication between the pairs was analyzed for interaction patterns used in the search process. This setup mimics the situation of a user interacting with a search system via a speech-only interface. Johanne R. Trippas, Damiano Spina, Lawrence Cavedon, Mark Sanderson |
CHIIR | 2 |
| 2017 | Extracting audio summaries to support effective spoken document searchabstractWe address the challenge of extracting query biased audio summaries from podcasts to support users in making relevance decisions in spoken document search via an audio‐only communication channel. We performed a crowdsourced experiment that demonstrates that transcripts of spoken documents created using Automated Speech Recognition (ASR), even with significant errors, are effective sources of document summaries or “snippets” for supporting users in making relevance judgments against a query. In particular, the results show that summaries generated from ASR transcripts are comparable, in utility and user‐judged preference, to spoken summaries generated from error‐free manual transcripts of the same collection. We also observed that content‐based audio summaries are at least as preferred as synthesized summaries obtained from manually curated metadata, such as title and description. We describe a methodology for constructing a new test collection, which we have made publicly available. Damiano Spina, Johanne R. Trippas, Lawrence Cavedon, Mark Sanderson |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2016 | Beyond Factoid QA: Effective Methods for Non-factoid Answer Sentence Retrieval
Liu Yang 0005, Qingyao Ai, Damiano Spina, Ruey-Cheng Chen, Liang Pang 0001, W. Bruce Croft, Jiafeng Guo, Falk Scholer |
ECIR | 3 |
| 2015 | Active Learning for Entity Filtering in Microblog StreamsabstractMonitoring the reputation of entities such as companies or brands in microblog streams (e.g., Twitter) starts by selecting mentions that are related to the entity of interest. Entities are often ambiguous (e.g., "Jaguar" or "Ford") and effective methods for selectively removing non-relevant mentions often use background knowledge obtained from domain experts. Manual annotations by experts, however, are costly. We therefore approach the problem of entity filtering with active learning, thereby reducing the annotation load for experts. To this end, we use a strong passive baseline and analyze different sampling methods for selecting samples for annotation. We find that margin sampling--an informative type of sampling that considers the distance to the hyperplane used for class separation--can effectively be used for entity filtering and can significantly reduce the cost of annotating initial training data. Damiano Spina, Maria-Hendrike Peetz, Maarten de Rijke |
SIGIR | 1 |
| 2015 | Towards Understanding the Impact of Length in Web Search Result Summaries over a Speech-only Communication ChannelabstractPresenting search results over a speech-only communication channel involves a number of challenges for users due to cognitive limitations and the serial nature of speech. We investigated the impact of search result summary length in speech-based web search, and compared our results to a text baseline. Based on crowdsourced workers, we found that users preferred longer, more informative summaries for text presentation. For audio, user preferences depended on the style of query. For single-facet queries, shortened audio summaries were preferred, additionally users were found to judge relevance with a similar accuracy compared to text-based summaries. For multi-facet queries, user preferences were not as clear, suggesting that more sophisticated techniques are required to handle such queries. Johanne R. Trippas, Damiano Spina, Mark Sanderson, Lawrence Cavedon |
SIGIR | 2 |
| 2015 | Real-time classification of Twitter trendsabstractIn this work, we explore the types of triggers that spark trends on Twitter, introducing a typology with the following 4 types: news, ongoing events, memes, and commemoratives. While previous research has analyzed trending topics over the long term, we look at the earliest tweets that produce a trend, with the aim of categorizing trends early on. This allows us to provide a filtered subset of trends to end users. We experiment with a set of straightforward language‐independent features based on the social spread of trends and categorize them using the typology. Our method provides an efficient way to accurately categorize trending topics without need of external data, enabling news organizations to discover breaking news in real‐time, or to quickly identify viral memes that might inform marketing decisions, among others. The analysis of social features also reveals patterns associated with each type of trend, such as tweets about ongoing events being shorter as many were likely sent from mobile devices, or memes having more retweets originating from a few trend‐setters. Arkaitz Zubiaga, Damiano Spina, Raquel Martínez-Unanue 0001, Víctor Fresno-Fernández |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2014 | ORMA: A Semi-automatic Tool for Online Reputation Monitoring in Twitter
Jorge Carrillo de Albornoz, Enrique Amigó, Damiano Spina, Julio Gonzalo 0001 |
ECIR | 3 |
| 2014 | Learning similarity functions for topic detection in online reputation monitoringabstractReputation management experts have to monitor--among others--Twitter constantly and decide, at any given time, what is being said about the entity of interest (a company, organization, personality...). Solving this reputation monitoring problem automatically as a topic detection task is both essential--manual processing of data is either costly or prohibitive--and challenging--topics of interest for reputation monitoring are usually fine-grained and suffer from data sparsity. We focus on a solution for the problem that (i) learns a pairwise tweet similarity function from previously annotated data, using all kinds of content-based and Twitter-based features; (ii) applies a clustering algorithm on the previously learned similarity function. Our experiments indicate that (i) Twitter signals can be used to improve the topic detection process with respect to using content signals only; (ii) learning a similarity function is a flexible and efficient way of introducing supervision in the topic detection clustering process. The performance of our best system is substantially better than state-of-the-art approaches and gets close to the inter-annotator agreement rate. A detailed qualitative inspection of the data further reveals two types of topics detected by reputation experts: reputation alerts / issues (which usually spike in time) and organizational topics (which are usually stable across time). Damiano Spina, Julio Gonzalo 0001, Enrique Amigó |
SIGIR | 1 |
| 2012 | Identifying entity aspects in microblog postsabstractOnline reputation management is about monitoring and handling the public image of entities (such as companies) on the Web. An important task in this area is identifying "aspects" of the entity of interest (such as products, services, competitors, key people, etc.) given a stream of microblog posts referring to the entity. In this paper we compare different IR techniques and opinion target identification methods for automatically identifying aspects and find that (i) simple statistical methods such as TF.IDF are a strong baseline for the task, significantly outperforming opinion-oriented methods, and (ii) only considering terms tagged as nouns improves the results for all the methods analyzed. Damiano Spina, Edgar Meij, Maarten de Rijke, Andrei Oghina, Minh Thuong Bui, Mathias Breuss |
SIGIR | 1 |
| 2011 | Classifying trending topics: a typology of conversation triggers on TwitterabstractTwitter summarizes the great deal of messages posted by users in the form of trending topics that reflect the top conversations being discussed at a given moment. These trending topics tend to be connected to current affairs. Different happenings can give rise to the emergence of these trending topics. For instance, a sports event broadcasted on TV, or a viral meme introduced by a community of users. Detecting the type of origin can facilitate information filtering, enhance real-time data processing, and improve user experience. In this paper, we introduce a typology to categorize the triggers that leverage trending topics: news, current events, memes, and commemoratives. We define a set of straightforward language-independent features that rely on the social spread of the trends to discriminate among those types of trending topics. Our method provides an efficient way to immediately and accurately categorize trending topics without need of external data, outperforming a content-based approach. Arkaitz Zubiaga, Damiano Spina, Víctor Fresno-Fernández, Raquel Martínez-Unanue 0001 |
CIKM | 2 |