EDBT 2026 Demo / reviewers in the wild / expert
Claudia Hauff
dblp:73/906
· DBLP profile ↗
67ranked-venue papers in the field
16as first author
20since 2021 · last 2026
0000-0001-9879-6470ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 65 (15 first)Big Data, Cloud & Distributed Data Systems · 1Other / Interdisciplinary · 1 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | The Third Search Futures Workshop at ECIR'26
Leif Azzopardi, Charles L. A. Clarke, Claudia Hauff, Yubin Kim 0001, Zhaochun Ren, Adam Roegiest, Johanne R. Trippas, Saber Zerhoudi |
ECIR (3) | 3 |
| 2026 | As It Was: Aligning LLM Search Evaluation with Historical User PreferencesabstractLarge-scale search systems evolve faster than human quality assurance scales, especially for long-tail intents and multilingual queries. LLM-as-a-judge approaches are a scalable alternative for evaluating the relevance of search engine result pages (SERPs), but judgments based solely on semantic similarity or world knowledge can drift from actual user preferences, particularly for ambiguous queries. We introduce a behavior-grounded LLM judge that augments each SERP item with a lightweight, auditable behavioral prior in the form of a Query--Relevance--Impressions (QRI) card. Each card summarizes how users have historically interacted with similar queries and results, providing compact empirical evidence that the judge can cite to resolve ambiguity and make more consistent relevance judgments, while still relying on semantic reasoning. In a large-scale music search evaluation at Spotify, using relevance estimates derived from historical user interactions across 6,000 recomposed SERPs, the behavior-grounded judge achieves stronger alignment with user preferences, improving Spearman rank correlation by approximately +5% overall and yielding a +91% relative improvement on disagreement cases. On a multilingual human-judged dataset spanning five languages, grounding further increases correlation with human relevance judgments by +15%. Importantly, when evaluated against outcomes from a live A/B test, the grounded judge shows consistently higher alignment with the observed winning model. While absolute alignment remains moderate, these findings demonstrate that lightweight behavioral grounding can improve the reliability and practical usefulness of LLM-based evaluation in real-world search systems. Ali Vardasbi, Gustavo Penha, Enrico Palumbo, Claudia Hauff, Hugues Bouchard, Mounia Lalmas-Roelleke |
SIGIR | 4 |
| 2025 | Zero-Shot Reranking with Large Language Models and Precomputed Ranking Features: Opportunities and LimitationsabstractLLMs have been explored for their use in IR as end-to-end rankers, rerankers and assessors. Recently, the exploration of the prompt-and-predict paradigm for reranking in combination with highly performant LLMs have drawn the attention of researchers. Instead of training or fine-tuning a reranker, LLMs are prompted in a zero-shot manner to produce relevance scores, pairwise preferences, or reranked lists. Existing research, though, has been confined to unstructured text corpora, leaving a gap in our understanding: to what extent do the findings of zero-shot LLM rerankers established on plain text corpora hold for datasets containing predominantly precomputed ranking features as is common in industrial settings? We explore this question via an empirical study on one public learning-to-rank dataset (MSLR-WEB10K) and two datasets collected from an audio streaming platform's search logs. Our results paint a differentiated picture: On average, there remains a significant performance gap: prompting the high-capacity LLM GPT-4 results in up to 16% lower NDCG@10 compared to the traditional supervised learning-to-rank (LTR) approach LambdaMART on the public MSLR-WEB10K dataset. However, when focusing only on a subset of hard queries-i.e. queries where the LTR approach ranks a non-relevant document at the top-the zero-shot LLM reranking outperforms the LTR baseline. We confirm the same trends on two proprietary audio search datasets. We also provide insights into prompt design choices and their impact on LLM reranking. We show that LLMs remain brittle, with the same strategies sometimes helping or hurting depending on the model size and dataset. Maria Movin, Claudia Hauff |
SIGIR | 2 |
| 2024 | On the Effects of Automatically Generated Adjunct Questions for Search as LearningabstractActively engaging learners with learning materials has been shown to be very important in the Search as Learning (SAL) setting. One active reading strategy relies on asking so-called adjunct questions, i.e., manually curated questions geared towards essential concepts of the target material. However, manual question creation is impractical given the vast online content. Recent research has explored the effects of Automatic Question Generation (AQG) on aiding human learning. These studies have primarily focused on user studies in controlled online reading scenarios with limited documents. However, the impacts of adjunct questions on learning in the SAL setting, which involves learning through web searching, are not yet well understood. This paper addresses this gap by conducting a user study with automatically generated adjunct questions integrated into the reading interface built on top of a search system. We conducted a between-subjects user study (N = 144) to investigate the incorporation of automatically generated adjunct questions on participants’ learning. We employed three different question generation strategies as well as a control condition: (i) synthesis questions; (ii) factoid questions targeting random text spans; and (iii) factoid questions targeting terms and phrases relevant to the information need at hand. We present four major findings: (i) participants who received adjunct questions exhibited significantly more fine-grained reading behaviour, such as longer document dwell time and more scrolls, than those without adjunct questions. However, adjunct questions’ influence on learning outcomes depends on the AQG strategy. (ii) Question types significantly influence participants’ reading behaviour. (iii) The adjunct questions’ target spans significantly influence learning outcomes. Lastly, (iv) participants’ prior knowledge levels affect adjunct questions’ effects on their learning outcomes and their reaction to different AQG strategies. Our findings have significant design implications for learning-oriented search systems. The data and code is available at https://github.com/zpeide/AQG-AdjunctQuestions. Peide Zhu, Arthur Câmara, Nirmal Roy, David Maxwell 0001, Claudia Hauff |
CHIIR | 5 |
| 2024 | PODTILE: Facilitating Podcast Episode Browsing with Auto-generated ChaptersabstractListeners of long-form talk-audio content, such as podcast episodes, often find it challenging to understand the overall structure and locate relevant sections. A practical solution is to divide episodes into chapters--semantically coherent segments labeled with titles and timestamps. Since most episodes on our platform at Spotify currently lack creator-provided chapters, automating the creation of chapters is essential. Scaling the chapterization of podcast episodes presents unique challenges. First, episodes tend to be less structured than written texts, featuring spontaneous discussions with nuanced transitions. Second, the transcripts are usually lengthy, averaging about 16,000 tokens, which necessitates efficient processing that can preserve context. To address these challenges, we introduce PODTILE, a fine-tuned encoder-decoder transformer to segment conversational data. The model simultaneously generates chapter transitions and titles for the input transcript. To preserve context, each input text is augmented with global context, including the episode's title, description, and previous chapter titles. In our intrinsic evaluation, PODTILE achieved an 11% improvement in ROUGE score over the strongest baseline. Additionally, we provide insights into the practical benefits of auto-generated chapters for listeners navigating episode content. Our findings indicate that auto-generated chapters serve as a useful tool for engaging with less popular podcasts. Finally, we present empirical evidence that using chapter titles can enhance effectiveness of sparse retrieval in search tasks. Azin Ghazimatin, Ekaterina Garmash, Gustavo Penha, Kristen Sheets, Martin Achenbach, Oguz Semerci, Remi Galvez, Marcus Tannenberg, Sahitya Mantravadi, Divya Narayanan, Ofeliya Kalaydzhyan, Douglas Cole, Ben Carterette, Ann Clifton, Paul N. Bennett, Claudia Hauff, Mounia Lalmas-Roelleke |
CIKM | 16 |
| 2023 | Driven to Distraction: Examining the Influence of Distractors on Search Behaviours, Performance and ExperienceabstractAdvertisements, sponsored links, clickbait, in-house recommendations and similar elements pervasively shroud featured content. Such elements vie for people’s attention, potentially distracting people from their task at hand. The effects of such “distractors” is likely to increase people’s cognitive workload and reduce their performance as they need to work harder to discern the relevant from non-relevant. In this paper, we investigate how people of varying cognitive abilities (measured using Perceptual Speed and Cognitive Failure instruments) are affected by these different types of distractions when completing search tasks. We performed a crowdsourced within-subjects user study, where 102 participants completed four search tasks using our news search engine over four different interface conditions: (i) one with no additional distractors; (ii) one with advertisements; (iii) one with sponsored links; and (iv) one with in-house recommendations. Our results highlight a number of important trends and findings. Participants perceived the interface condition without distractors as significantly better across numerous dimensions. Participants reported higher satisfaction, lower workload, higher topic recall, and found it easier to concentrate. Behaviourally, participants issued queries faster and clicked results earlier when compared to the interfaces with distractors. When using the interfaces with distractors, one in ten participants clicked on a distractor—and despite engaging with a distractor for less than twenty seconds, their task time increased by approximately two minutes. We found that the effects were magnified depending on cognitive abilities—with a greater impact of distractors on participants with lower perceptual speed, and for those with a higher propensity of cognitive failures. Distractors—regardless of their type—have negative consequences on a user’s search experience and performance. As a consequence, interfaces containing visually distracting elements are creating poorer search experiences due to the “distractor tax” being placed on people’s limited attention. Leif Azzopardi, David Maxwell 0001, Martin Halvey, Claudia Hauff |
CHIIR | 4 |
| 2023 | Graph Learning for Exploratory Query Suggestions in an Instant Search SystemabstractSearch systems in online content platforms are typically biased toward a minority of highly consumed items, reflecting the most common user behavior of navigating toward content that is already familiar and popular. Query suggestions are a powerful tool to support query formulation and to encourage exploratory search and content discovery. However, classic approaches for query suggestions typically rely either on semantic similarity, which lacks diversity and does not reflect user searching behavior, or on a collaborative similarity measure mined from search logs, which suffers from data sparsity and is biased by highly popular queries. In this work, we argue that the task of query suggestion can be modelled as a link prediction task on a heterogeneous graph including queries and documents, enabling Graph Learning methods to effectively generate query suggestions encompassing both semantic and collaborative information. We perform an offline evaluation on an internal Spotify dataset of search logs and on two public datasets, showing that node2vec leads to an accurate and diversified set of results, especially on the large scale real-world data. We then describe the implementation in an instant search scenario and discuss a set of additional challenges tied to the specific production environment. Finally, we report the results of a large scale A/B test involving millions of users and prove that node2vec query suggestions lead to an increase in online metrics such as coverage (+1.42% shown search results pages with suggestions) and engagement (+1.21% clicks), with a specifically notable boost in the number of clicks on exploratory search queries (+9.37%). Enrico Palumbo, Andreas Damianou, Alice Wang 0001, Alva Liu, Ghazal Fazelnia, Francesco Fabbri, Fabrizio Silvestri, Hugues Bouchard, Claudia Hauff, Mounia Lalmas-Roelleke, Ben Carterette, Praveen Chandar, David Nyhan |
CIKM | 10 |
| 2023 | Do the Findings of Document and Passage Retrieval Generalize to the Retrieval of Responses for Dialogues?
Gustavo Penha, Claudia Hauff |
ECIR (3) | 2 |
| 2023 | When the Music Stops: Tip-of-the-Tongue Retrieval for MusicabstractWe present a study of Tip-of-the-tongue (ToT) retrieval for music, where a searcher is trying to find an existing music entity, but is unable to succeed as they cannot accurately recall important identifying information. ToT information needs are characterized by complexity, verbosity, uncertainty, and possible false memories. We make four contributions. (1) We collect a dataset - TOTMUSIC--of 2,278 information needs and ground truth answers. (2) We introduce a schema for these information needs and show that they often involve multiple modalities encompassing several Music IR sub-tasks such as lyric search, audio-based search, audio fingerprinting, and text search. (3) We underscore the difficulty of this task by benchmarking a standard text retrieval approach on this dataset. (4) We investigate the efficacy of query reformulations generated by a Large Language Model (LLM), and show that they are not as effective as simply employing the entire information need as a query--leaving several open questions for future research. Samarth Bhargav 0001, Anne Schuth, Claudia Hauff |
SIGIR | 3 |
| 2023 | Hear Me Out: A Study on the Use of the Voice Modality for Crowdsourced Relevance AssessmentsabstractThe creation of relevance assessments by human assessors (often nowadays crowdworkers) is a vital step when building IR test collections. Prior works have investigated assessor quality & behaviour, and tooling to support assessors in their task. We have few insights though into the impact of a document's presentation modality on assessor efficiency and effectiveness. Given the rise of voice-based interfaces, we investigate whether it is feasible for assessors to judge the relevance of text documents via a voice-based interface. We ran a user study (n = 49) on a crowdsourcing platform where participants judged the relevance of short and long documents- sampled from the TREC Deep Learning corpus-presented to them either in the text or voice modality. We found that: (i) participants are equally accurate in their judgements across both the text and voice modality; (ii) with increased document length it takes partic- ipants significantly longer (for documents of length > 120 words it takes almost twice as much time) to make relevance judgements in the voice condition; and (iii) the ability of assessors to ignore stimuli that are not relevant (i.e., inhibition) impacts the assessment quality in the voice modality-assessors with higher inhibition are significantly more accurate than those with lower inhibition. Our results indicate that we can reliably leverage the voice modality as a means to effectively collect relevance labels from crowdworkers. Nirmal Roy, Agathe Balayn, David Maxwell 0001, Claudia Hauff |
SIGIR | 4 |
| 2022 | On the Challenges of Podcast Search at SpotifyabstractOnline music streaming is enjoying ever-growing popularity in the last decades, enabled by the abundance of music content in digital format and online streaming services. In the recent years, podcasts, as a talk-focused media format, have witnessed a rapid growth among listeners. Podcasts come in many forms and sizes. They range from 20-minute daily meditation sessions, weekly recaps of global news, interviews with celebrities, and hosts bantering with each other for hours. Claudia Hauff, Praveen Chandar |
CIKM | 2 |
| 2022 | Searching, Learning, and Subtopic Ordering: A Simulation-Based Analysis
Arthur Câmara, David Maxwell 0001, Claudia Hauff |
ECIR (1) | 3 |
| 2022 | Evaluating the Robustness of Retrieval Pipelines with Query Variation Generators
Gustavo Penha, Arthur Câmara, Claudia Hauff |
ECIR (1) | 3 |
| 2022 | Users and Contemporary SERPs: A (Re-)InvestigationabstractTheSearch Engine Results Page (SERP) has evolved significantly over the last two decades, moving away from the simple ten blue links paradigm to considerably more complex presentations that contain results from multiple verticals and granularities of textual information. Prior works have investigated how user interactions on the SERP are influenced by the presence or absence of heterogeneous content (e.g., images, videos, or news content), the layout of the SERP (\emphlist vs. grid layout), and task complexity. In this paper, we reproduce the user studies conducted in prior works---specifically those of~\citetarguello2012task and~\citetsiu2014first ---to explore to what extent the findings from research conducted five to ten years ago still hold today as the average web user has become accustomed to SERPs with ever-increasing presentational complexity. To this end, we designed and ran a user study with four different SERP interfaces:(i) ~\empha heterogeneous grid ;(ii) ~\empha heterogeneous list ;(iii) ~\empha simple grid ; and(iv) ~\empha simple list. We collected the interactions of $41$ study participants over $12$ search tasks for our analyses. We observed that SERP types and task complexity affect user interactions with search results. We also find evidence to support most (6 out of 8) observations from~\citearguello2012task,siu2014first indicating that user interactions with different interfaces and to solve tasks of different complexity have remained mostly similar over time. Nirmal Roy, David Maxwell 0001, Claudia Hauff |
SIGIR | 3 |
| 2021 | Searching to Learn with Instructional ScaffoldingabstractWeb search engines are today considered to be the primary tool to assist and empower learners in finding information relevant to their learning goals- be it learning something new, improving their existing skills, or just fulfilling a curiosity. While several approaches for improving search engines for the learning scenario have been proposed (e.g. a specific ranking function), instructional scaffolding (or simply scaffolding)-a traditional learning support strategy-has not been studied in the context of search as learning, despite being shown to be effective for improving learning in both digital and traditional learning contexts. When scaffolding is employed, instructors provide learners with support throughout their autonomous learning process. We hypothesize that the usage of scaffolding techniques within a search system can be an effective way to help learners achieve their learning objectives whilst searching. As such, this paper investigates the incorporation of scaffolding into a search system employing three different strategies (as well as a control condition): (i) AQe, the automatic expansion of user queries with relevant subtopics; (ii) CURATEDsc, the presenting of a manually curated static list of relevant subtopics on the search engine result page; and (iii) FEEDBACKsc, which projects real-time feedback about a user's exploration of the topic space on top of the CURATEDsc visualization. To investigate the effectiveness of these approaches with respect to human learning, we conduct a user study (N=126) where participants were tasked with searching and learning about topics such as genetically modified organisms. We find that (i) the introduction of the proposed scaffolding methods in the proposed topics does not significantly improve learning gains. However, (ii) it does significantly impact search behavior. Furthermore, (iii) immediate feedback of the participants' learning (FEEDBACKsc) leads to undesirable user behavior, with participants seemingly focusing on the feedback gauges instead of learning. Arthur Câmara, Nirmal Roy, David Maxwell 0001, Claudia Hauff |
CHIIR | 4 |
| 2021 | Note the Highlight: Incorporating Active Reading Tools in a Search as Learning EnvironmentabstractActive reading strategies---such as content annotations (through the use of highlighting and note-taking, for example)---have been shown to yield improvements to a learner's knowledge and understanding of the topic being explored. This has been especially notable in long and complex learning endeavours. With web search engines nowadays used as the primary gateway for learners (or users) to find content that helps them realise their learning goals, they are often poorly equipped with the necessary tools to aid in sense-making, an important aspect of theSearch as Learning (SAL) process. Within theInformation Retrieval (IR) community, research efforts have explored ways to keep track of users' search context by providing a notepad-like interface for the collection of relevant articles, and aid them during the exploratory search process. However, these studies did not explicitly measure the effect that such tools have on knowledge and understanding during a complex, learning-oriented search task. In this paper, we address this research gap by carrying out an Interactive IR experiment with highlighting and note-taking tools built into the search interface. We conducted a crowdsourced between-subjects study (N=115), where participants were assigned to one of four conditions: (i) control (a standard web search interface); (ii) high (highlighting enabled);(iii) note (note-taking enabled); and (iv) highnote (both highlighting and note-taking enabled). We assess participants' learning with a recall-oriented vocabulary learning task, and a cognitively more taxing essay writing task. We find that(i) active reading tools do not aid in the vocabulary learning task. However,(ii) participants in high covered 34% more subtopics, and participants in note covered 34% more facts in their essays when compared to control. Furthermore, (iii) we observed that incorporating active learning tools significantly changed the search behaviour of participants across a number of measures. This is the first work that sheds light on the effect of active reading tools on the SAL process, with important design implications for learning-oriented search systems. Nirmal Roy, Manuel Valle Torre, Ujwal Gadiraju, David Maxwell 0001, Claudia Hauff |
CHIIR | 5 |
| 2021 | LogUI: Contemporary Logging Infrastructure for Web-Based Experiments
David Maxwell 0001, Claudia Hauff |
ECIR (2) | 2 |
| 2021 | Weakly Supervised Label Smoothing
Gustavo Penha, Claudia Hauff |
ECIR (2) | 2 |
| 2021 | How Do Active Reading Strategies Affect Learning Outcomes in Web Search?
Nirmal Roy, Manuel Valle Torre, Ujwal Gadiraju, David Maxwell 0001, Claudia Hauff |
ECIR (2) | 5 |
| 2021 | Conversational Search and Recommendation: Introduction to the Special IssueabstractAn introduction to the special issue on conversational search and recommendation is presented in this article. While conversational search and recommendation has roots in early Information Retrieval (IR) research, the recent advances in automatic voice recognition and conversational agents have created increasing interest in this area. In recent years, the IR and related communities have witnessed a number of major contributions to the field of conversational search and recommendation. They include but are not limited to conversational search conceptualization. The growing body of work in this area has been supplemented by an increasing number of recent seminars. Claudia Hauff, Julia Kiseleva, Mark Sanderson, Hamed Zamani |
ACM Trans. Inf. Syst. | 1 |
| 2020 | Exploring Users' Learning Gains within Search SessionsabstractThe area of search as learning is concerned with the optimization of search systems (that is, retrieval functions, user interface elements, etc.) for human learning ---this is in contrast to the currently dominant paradigm of optimizing the search experience by optimizing for relevance. While prior work typically considers learning as something that happens at some point during the search session, we are interested in when during the search session learning occurs. In order to answer this question, we here present the results of a user study ($N=64$) in which searchers were tasked with learning about a topic by searching the web for 20 minutes; they were prompted at regular intervals during the search session on their knowledge about the topic. We find that for study participants with little to no prior knowledge the learning gains are sublinear, while participants with some prior knowledge have the largest knowledge gains towards the end of the search session. Nirmal Roy, Felipe Moraes, Claudia Hauff |
CHIIR | 3 |
| 2020 | Diagnosing BERT with Retrieval Heuristics
Arthur Câmara, Claudia Hauff |
ECIR (1) | 2 |
| 2020 | Curriculum Learning Strategies for IR
Gustavo Penha, Claudia Hauff |
ECIR (1) | 2 |
| 2020 | What does BERT know about books, movies and music? Probing BERT for Conversational RecommendationabstractHeavily pre-trained transformer models such as BERT have recently shown to be remarkably powerful at language modelling, achieving impressive results on numerous downstream tasks. It has also been shown that they implicitly store factual knowledge in their parameters after pre-training. Understanding what the pre-training procedure of LMs actually learns is a crucial step for using and improving them for Conversational Recommender Systems (CRS). We first study how much off-the-shelf pre-trained BERT “knows” about recommendation items such as books, movies and music. In order to analyze the knowledge stored in BERT’s parameters, we use different probes (i.e., tasks to examine a trained model regarding certain properties) that require different types of knowledge to solve, namely content-based and collaborative-based. Content-based knowledge is knowledge that requires the model to match the titles of items with their content information, such as textual descriptions and genres. In contrast, collaborative-based knowledge requires the model to match items with similar ones, according to community interactions such as ratings. We resort to BERT’s Masked Language Modelling (MLM) head to probe its knowledge about the genre of items, with cloze style prompts. In addition, we employ BERT’s Next Sentence Prediction (NSP) head and representations’ similarity (SIM) to compare relevant and non-relevant search and recommendation query-document inputs to explore whether BERT can, without any fine-tuning, rank relevant items first. Finally, we study how BERT performs in a conversational recommendation downstream task. To this end, we fine-tune BERT to act as a retrieval-based CRS. Overall, our experiments show that: (i) BERT has knowledge stored in its parameters about the content of books, movies and music; (ii) it has more content-based knowledge than collaborative-based knowledge; and (iii) fails on conversational recommendation when faced with adversarial data. Gustavo Penha, Claudia Hauff |
RecSys | 2 |
| 2019 | node-indri: Moving the Indri Toolkit to the Modern Web Stack
Felipe Moraes, Claudia Hauff |
ECIR (2) | 2 |
| 2019 | An Axiomatic Approach to Diagnosing Neural IR Models
Daniël Rennings, Felipe Moraes, Claudia Hauff |
ECIR (1) | 3 |
| 2019 | The SIGIR 2019 Open-Source IR Replicability Challenge (OSIRRC 2019)abstractThe importance of repeatability, replicability, and reproducibility is broadly recognized in the computational sciences, both in supporting desirable scientific methodology as well as sustaining empirical progress. This workshop tackles the replicability challenge for ad hoc document retrieval, via a common Docker interface specification to support images that capture systems performing ad hoc retrieval experiments on standard test collections. Ryan Clancy, Nicola Ferro 0001, Claudia Hauff, Jimmy Lin, Tetsuya Sakai, Ze Zhong Wu |
SIGIR | 3 |
| 2019 | On the impact of group size on collaborative search effectiveness
Felipe Moraes, Kilian Grashoff, Claudia Hauff |
Inf. Retr. J. | 3 |
| 2018 | Contrasting Search as a Learning Activity with Instructor-designed LearningabstractThe field of Search as Learning addresses questions surrounding human learning during the search process. Existing research has largely focused on observing how users with learning-oriented information needs behave and interact with search engines. What is not yet quantified is the extent to which search is a viable learning activity compared to instructor-designed learning. Can a search session be as effective as a lecture video - our instructor-designed learning artefact - or learning? To answer this question, we designed a user study that pits instructor-designed learning (a short high-quality video lecture as commonly found in online learning platforms) against three instances of search, specifically (i) single-user search, (ii) search as a support tool for instructor-designed learning, and, (iii) collaborative search. We measured the learning gains of 151 study participants in a vocabulary learning task and report three main results: (i) lecture video watching yields up to 24% higher learning gains than single-user search, (ii) collaborative search for learning does not lead to increased learning, and (iii) lecture video watching supported by search leads up to a 41% improvement in learning gains over instructor-designed learning without a subsequent search phase. Felipe Moraes, Sindunuraga Rikarno Putra, Claudia Hauff |
CIKM | 3 |
| 2018 | LearningQ: A Large-Scale Dataset for Educational Question Generation
Guanliang Chen, Jie Yang 0028, Claudia Hauff, Geert-Jan Houben |
ICWSM | 3 |
| 2018 | A/B Testing with APONEabstractIn order to improve long-term retention, ad conversion rates, and so on, A/B testing has become the norm within Web portals, enabling efficient large-scale experimentation. While A/B testing is also increasingly used by academic researchers (with crowd-working platforms offering a large pool of artificial users), few platforms are freely available to this end. Academic researchers usually develop adhoc solutions, leading to many duplicated efforts and time spent on work not directly related to one's research. As an alternative, we have developed and open sourced APONE, an A cademic P latform for ON line Experiments. APONE uses PlanOut, a framework and high-level language, to specify online experiments, and offers Web services and a Web GUI to easily create, manage and monitor them. By building a user friendly Web application, we enable not only experts to conduct valid A/B experiments. In particular as a secondary use case, we envision large classrooms to also benefit from the deployment of APONE, a vision we put into practice in a graduate Information Retrieval course. We open-source APONE at https://marrerom.github.io/APONE. A demo version is running at http://ireplatform.ewi.tudelft.nl:8080/APONE. Mónica Marrero, Claudia Hauff |
SIGIR | 2 |
| 2018 | SearchX: Empowering Collaborative Search ResearchabstractCollaborative search has been an active area of research within the IR community for many years. While for "single-user'' research a variety of up-to-date open-source search systems exist, few "multi-user'' search tools are open-source and even fewer are being maintained. In this paper, we present SearchX, an open-source collaborative search system we are currently developing-and using for our research. We designed and built SearchX using the modern Web stack (and are thus not siloed by an operating system or a particular browser type), enabling efficient research across platforms (Desktop, mobile) and with online users (e.g. crowdworkers). A video, describing the demo can be found at https: //www.youtube.com/watch?v=uf24m6p3vts. Sindunuraga Rikarno Putra, Felipe Moraes, Claudia Hauff |
SIGIR | 3 |
| 2017 | Exploring the Query Halo Effect in Site Search: Leading People to Longer QueriesabstractPeople tend to type short queries, however, the belief is that longer queries are more effective. Consequently, a number of attempts have been made to encourage and motivate people to enter longer queries. While most have failed, a recent attempt - conducted in a laboratory setup - in which the query box has a halo or glow effect, that changes as the query becomes longer, has been shown to increase query length by one term, on average. In this paper, we test whether a similar increase is observed when the same component is deployed in a production system for site search and used by real end users. To this end, we conducted two separate experiments, where the rate at which the color changes in the halo were varied. In both experiments users were assigned to one of two conditions: halo and no-halo. The experiments were ran over a fifty day period with 3,506 unique users submitting over six thousand queries. In both experiments, however, we observed no significant difference in query length. We also did not find longer queries to result in greater retrieval performance. While, we did not reproduce the previous findings, our results indicate that the query halo effect appears to be sensitive to performance and task, limiting its applicability to other contexts. Djoerd Hiemstra, Claudia Hauff, Leif Azzopardi |
SIGIR | 2 |
| 2017 | Introduction to the special issue on search as learning
Carsten Eickhoff, Jacek Gwizdka, Claudia Hauff, Jiyin He |
Inf. Retr. J. | 3 |
| 2016 | Sub-document Timestamping: A Study on the Content Creation Dynamics of Web Documents
Yue Zhao 0001, Claudia Hauff |
TPDL | 2 |
| 2016 | Diversity in Urban Social Media Analytics
Jie Yang 0028, Claudia Hauff, Geert-Jan Houben, Christiaan Titos Bolivar |
ICWE | 2 |
| 2016 | Search as Learning (SAL) Workshop 2016abstractThe "Search as Learning" (SAL) workshop is focused on an area within the information retrieval field that is only beginning to emerge: supporting users in their learning whilst interacting with information content. Jacek Gwizdka, Preben Hansen, Claudia Hauff, Jiyin He, Noriko Kando |
SIGIR | 3 |
| 2016 | Temporal Query Intent Disambiguation using Time-Series DataabstractUnderstanding temporal intents behind users' queries is essential to meet users' time-related information needs. In order to classify queries according to their temporal intent (e.g. Past or Future), we explore the usage of time-series data derived from Wikipedia page views as a feature source. While existing works leverage either proprietary search engine query logs or highly processed and aggregated data (such as Google Trends) for this purpose, we investigate the utility of a freely available data source for this purpose. Our experiments on the NTCIR-12 Temporalia-2 dataset show, that Wikipedia pageview-based time-series data can significantly improve the disambiguation of temporal intents for specific types of queries, in particular those without temporal expressions present in the query string. Yue Zhao 0001, Claudia Hauff |
SIGIR | 2 |
| 2015 | Using Query-Log Based Collective Intelligence to Generate Query Suggestions for Tagged Content Search
Dirk Guijt, Claudia Hauff |
ICWE | 2 |
| 2015 | Matching GitHub Developer Profiles to Job AdvertisementsabstractGitHub is a social coding platform that enables developers to efficiently work on projects, connect with other developers, collaborate and generally "be seen: by the community. This visibility also extends to prospective employers and HR personnel who may use GitHub to learn more about a developer's skills and interests. We propose a pipeline that automatizes this process and automatically suggests matching job advertisements to developers, based on signals extracting from their activities on GitHub. Claudia Hauff, Georgios Gousios |
MSR | 1 |
| 2015 | SIGIR 2015 Workshop on Temporal, Social and Spatially-aware Information Access (#TAIA2015)abstractIn this workshop we aim to bring together practitioners and researchers to discuss their recent breakthroughs and the challenges with addressing spatial and temporal information access, both from the algorithmic and the architectural perspectives. Klaus Berberich, James Caverlee, Miles Efron, Claudia Hauff, Vanessa Murdock 0001, Milad Shokouhi, Bart Thomee |
SIGIR | 4 |
| 2015 | Learning by Example: Training Users with High-quality Query SuggestionsabstractThe queries submitted by users to search engines often poorly describe their information needs and represent a potential bottleneck in the system. In this paper we investigate to what extent it is possible to aid users in learning how to formulate better queries by providing examples of high-quality queries interactively during a number of search sessions. By means of several controlled user studies we collect quantitative and qualitative evidence that shows: (1) study participants are able to identify and abstract qualities of queries that make them highly effective, (2) after seeing high-quality example queries participants are able to themselves create queries that are highly effective, and, (3) those queries look similar to expert queries as defined in the literature. We conclude by discussing what the findings mean in the context of the design of interactive search systems. Morgan Harvey, Claudia Hauff, David Elsweiler |
SIGIR | 2 |
| 2015 | Sub-document Timestamping of Web DocumentsabstractKnowledge about a (Web) document's creation time has been shown to be an important factor in various temporal information retrieval settings. Commonly, it is assumed that such documents were created at a single point in time. While this assumption may hold for news articles and similar document types, it is a clear oversimplification for general Web documents. In this paper, we investigate to what extent (i) this simplifying assumption is violated for a corpus of Web documents, and, (ii) it is possible to accurately estimate the creation time of individual Web documents' components (so-called sub-documents). Yue Zhao 0001, Claudia Hauff |
SIGIR | 2 |
| 2014 | Facilitating Twitter data analytics: Platform, language and functionalityabstractConducting analytics over data generated by Social Web portals such as Twitter is challenging, due to the volume, variety and velocity of the data. Commonly, adhoc pipelines are used that solve a particular use case. In this paper, we generalize across a range of typical Twitter-data use cases and determine a set of common characteristics. Based on this investigation, we present our Twitter Analytical Platform (TAP), a generic platform for conducting analytical tasks with Twitter data. The platform provides a domain-specific Twitter Analysis Language (TAL) as the interface to its functionality stack. TAL includes a set of analysis tools ranging from data collection and semantic enrichment, to machine learning. With these tools, it becomes possible to create and customize analytical workflows in TAL and build applications that make use of the analytics results. We showcase the applicability of our platform by building Twinder-a search engine for Twitter streams. Ke Tao, Claudia Hauff, Geert-Jan Houben, Fabian Abel, Guido Wachsmuth |
IEEE BigData | 2 |
| 2014 | Large-scale author verification: temporal and topical influencesabstractThe task of author verification is concerned with the question whether or not someone is the author of a given piece of text. Algorithms that extract writing style features from texts are used to determine how close in style different documents are. Currently, evaluations of author verification algorithms are restricted to small-scale corpora with usually less than one hundred test cases. In this work, we present a methodology to derive a large-scale author verification corpus based on Wikipedia Talkpages. We create a corpus based on English Wikipedia which is significantly larger than existing corpora. We investigate two dimensions on this corpus which so far have not received sufficient attention: the influence of topic and the influence of time on author verification accuracy. Michiel van Dam, Claudia Hauff |
SIGIR | 2 |
| 2014 | SIGIR 2014 workshop on temporal, social and spatially-aware information access (#TAIA2014)abstractNo abstract available. Fernando Diaz 0001, Claudia Hauff, Vanessa Murdock 0001, Maarten de Rijke, Milad Shokouhi |
SIGIR | 2 |
| 2013 | A study on the accuracy of Flickr's geotag dataabstractObtaining geographically tagged multimedia items from social Web platforms such as Flickr is beneficial for a variety of applications including the automatic creation of travelogues and personalized travel recommendations. In order to take advantage of the large number of photos and videos that do not contain (GPS-based) latitude/longitude coordinates, a number of approaches have been proposed to estimate the geographic location where they were taken. Such location estimation methods rely on existing geotagged multimedia items as training data. Across application and usage scenarios, it is commonly assumed that the available geotagged items contain (reasonably) accurate latitude/longitude coordinates. Here, we consider this assumption and investigate how accurate the provided location data is. We conduct a study of Flickr images and videos and find that the accuracy of the geotag information is highly dependent on the popularity of the location: images/videos taken at popular (unpopular) locations, are likely to be geotagged with a high (low) degree of accuracy with respect to the ground truth. Claudia Hauff |
SIGIR | 1 |
| 2013 | Groundhog day: near-duplicate detection on TwitterabstractWith more than 340~million messages that are posted on Twitter every day, the amount of duplicate content as well as the demand for appropriate duplicate detection mechanisms is increasing tremendously. Yet there exists little research that aims at detecting near-duplicate content on microblogging platforms. We investigate the problem of near-duplicate detection on Twitter and introduce a framework that analyzes the tweets by comparing (i) syntactical characteristics, (ii) semantic similarity, and (iii) contextual information. Our framework provides different duplicate detection strategies that, among others, make use of external Web resources which are referenced from microposts. Machine learning is exploited in order to learn patterns that help identifying duplicate content. We put our duplicate detection framework into practice by integrating it into Twinder, a search engine for Twitter streams. An in-depth analysis shows that it allows Twinder to diversify search results and improve the quality of Twitter search. We conduct extensive experiments in which we (1) evaluate the quality of different strategies for detecting duplicates, (2) analyze the impact of various features on duplicate detection, (3) investigate the quality of strategies that classify to what exact level two microposts can be considered as duplicates and (4) optimize the process of identifying duplicate content on Twitter. Our results prove that semantic features which are extracted by our framework can boost the performance of detecting duplicates. Ke Tao, Fabian Abel, Claudia Hauff, Geert-Jan Houben, Ujwal Gadiraju |
WWW | 3 |
| 2012 | Geo-Location Estimation of Flickr Images: Social Web Based Enrichment
Claudia Hauff, Geert-Jan Houben |
ECIR | 1 |
| 2012 | Leveraging User Modeling on the Social Web with Linked Data
Fabian Abel, Claudia Hauff, Geert-Jan Houben, Ke Tao |
ICWE | 2 |
| 2012 | Social Event Detection on Twitter
Elena Ilina, Claudia Hauff, Ilknur Celik, Fabian Abel, Geert-Jan Houben |
ICWE | 2 |
| 2012 | Twinder: A Search Engine for Twitter Streams
Ke Tao, Fabian Abel, Claudia Hauff, Geert-Jan Houben |
ICWE | 3 |
| 2012 | Placing images on the world map: a microblog-based enrichment approachabstractEstimating the geographic location of images is a task which has received increasing attention recently. Large numbers of images uploaded to platforms such as Flickr do not contain GPS-based latitude/longitude coordinates. Obtaining such geographic information is beneficial for a variety of applications including travelogues, visual place descriptions and personalized travel recommendations. While most works in this area only exploit an image's textual meta-data (tags, title, etc.) to estimate at what geographic location the image was taken, we consider an additional textual dimension: the image owner's traces on the social Web. Specifically, we hypothesize that information extracted from a person's microblog stream(s) can be utilized to improve the accuracy with which the geographic location of the images is estimated. In this paper, we investigate this hypothesis on the example of Twitter streams and find it to be confirmed. The median error distance in kilometres decreases by up to 67% in comparison to existing state-of-the-art. The best results are achieved when tweets that were posted up to two days before and after an image was taken are considered. Moreover, we also find another type of additional information useful: population density data. Claudia Hauff, Geert-Jan Houben |
SIGIR | 1 |
| 2011 | Classic Children's Literature - Difficult to Read?
Dolf Trieschnigg, Claudia Hauff |
ECIR | 2 |
| 2010 | A comparison of user and system query performance predictionsabstractQuery performance prediction methods are usually applied to estimate the retrieval effectiveness of queries, where the evaluation is largely system sided. However, little work has been conducted to understand query performance prediction from the user's perspective. The question we consider is, whether the predictions of query performance that systems make are in line with the predictions that users make. To this aim, we compare the performance ratings users assign to queries with the performance scores estimated by a range of pre-retrieval and post-retrieval query performance predictors. Two studies are presented that explore the relationship between user ratings and system predictions on two levels: (i) the topic level, and, (ii) the query suggestions level. It is shown that when predicting the performance of query suggestions, user ratings were mostly uncorrelated with system predictions. At the topic level though, where a single query is judged for each information need, we observed moderate correlations between user ratings and a subset of system predictions. As query performance prediction methods are often based on intuitions of how users might rate queries, these findings suggest that such methods are not representative of how users actually rate query suggestions and topics. This motivates further research into understanding the rating process engaged by users, and developing models of query performance prediction in order to bridge the divide between systems and users. Claudia Hauff, Diane Kelly 0001, Leif Azzopardi |
CIKM | 1 |
| 2010 | Query Performance Prediction: Evaluation Contrasted with Effectiveness
Claudia Hauff, Leif Azzopardi, Djoerd Hiemstra, Franciska de Jong |
ECIR | 1 |
| 2010 | A Case for Automatic System Evaluation
Claudia Hauff, Djoerd Hiemstra, Leif Azzopardi, Franciska de Jong |
ECIR | 1 |
| 2010 | Retrieval system evaluation: automatic evaluation versus incomplete judgmentsabstractIn information retrieval (IR), research aiming to reduce the cost of retrieval system evaluations has been conducted along two lines: (i) the evaluation of IR systems with reduced amounts of manual relevance assessments, and (ii) the fully automatic evaluation of IR systems, thus foregoing the need for manual assessments altogether. The proposed methods in both areas are commonly evaluated by comparing their performance estimates for a set of systems to a ground truth (provided for instance by evaluating the set of systems according to mean average precision). In contrast, in this poster we compare an automatic system evaluation approach directly to two evaluations based on incomplete manual relevance assessments. For the particular case of TREC's Million Query track, we show that the automatic evaluation leads to results which are highly correlated to those achieved by approaches relying on incomplete manual judgments. Claudia Hauff, Franciska de Jong |
SIGIR | 1 |
| 2010 | Query quality: user ratings and system predictionsabstractNumerous studies have examined the ability of query performance prediction methods to estimate a query's quality for system effectiveness measures (such as average precision). However, little work has explored the relationship between these methods and user ratings of query quality. In this poster, we report the findings from an empirical study conducted on the TREC ClueWeb09 corpus, where we compared and contrasted user ratings of query quality against a range of query performance prediction methods. Given a set of queries, it is shown that user ratings of query quality correlate to both system effectiveness measures and a number of pre-retrieval predictors. Claudia Hauff, Franciska de Jong, Diane Kelly 0001, Leif Azzopardi |
SIGIR | 1 |
| 2010 | Estimating interference in the QPRP for subtopic retrievalabstractThe Quantum Probability Ranking Principle (QPRP) has been recently proposed, and accounts for interdependent document relevance when ranking. However, to be instantiated, the QPRP requires a method to approximate the "interference" between two documents. In this poster, we empirically evaluate a number of different methods of approximation on two TREC test collections for subtopic retrieval. It is shown that these approximations can lead to significantly better retrieval performance over the state of the art. Guido Zuccon, Leif Azzopardi, Claudia Hauff, C. J. van Rijsbergen |
SIGIR | 3 |
| 2009 | Relying on topic subsets for system ranking estimationabstractRanking a number of retrieval systems according to their retrieval effectiveness without relying on costly relevance judgments was first explored by Soboroff et al [6]. Over the years, a number of alternative approaches have been proposed. We perform a comprehensive analysis of system ranking estimation approaches on a wide variety of TREC test collections and topics sets. Our analysis reveals that the performance of such approaches is highly dependent upon the topic or topic subset, used for estimation. We hypothesize that the performance of system ranking estimation approaches can be improved by selecting the "right" subset of topics and show that using topic subsets improves the performance by 32% on average, with a maximum improvement of up to 70% in some cases. Claudia Hauff, Djoerd Hiemstra, Franciska de Jong, Leif Azzopardi |
CIKM | 1 |
| 2009 | The Combination and Evaluation of Query Performance Prediction Methods
Claudia Hauff, Leif Azzopardi, Djoerd Hiemstra |
ECIR | 1 |
| 2009 | Efficiency trade-offs in two-tier web search systemsabstractSearch engines rely on searching multiple partitioned corpora to return results to users in a reasonable amount of time. In this paper we analyze the standard two-tier architecture for Web search with the difference that the corpus to be searched for a given query is predicted in advance. We show that any predictor better than random yields time savings, but this decrease in the processing time yields an increase in the infrastructure cost. We provide an analysis and investigate this trade-off in the context of two different scenarios on real-world data. We demonstrate that in general the decrease in answer time is justified by a small increase in infrastructure cost. Ricardo Baeza-Yates, Vanessa Murdock 0001, Claudia Hauff |
SIGIR | 3 |
| 2009 | When is query performance prediction effective?abstractThe utility of Query Performance Prediction (QPP) methods is commonly evaluated by reporting correlation coefficients to denote how well the methods perform at predicting the retrieval performance of a set of queries. However, a quintessential question remains unexplored: how strong does the correlation need to be in order to realize an increase in retrieval performance? In this work, we address this question in the context of Selective Query Expansion (SQE) and perform a large-scale experiment. The results show that to consistently and predictably improve retrieval effectiveness in the ideal SQE setting, a Kendall's Tau correlation of tau>=0.5 is required, a threshold which most existing query performance prediction methods fail to reach. Claudia Hauff, Leif Azzopardi |
SIGIR | 1 |
| 2008 | A survey of pre-retrieval query performance predictorsabstractThe focus of research on query performance prediction is to predict the effectiveness of a query given a search system and a collection of documents. If the performance of queries can be estimated in advance of, or during the retrieval stage, specific measures can be taken to improve the overall performance of the system. In particular, pre-retrieval predictors predict the query performance before the retrieval step and are thus independent of the ranked list of results; such predictors base their predictions solely on query terms, the collection statistics and possibly external sources such as WordNet. In this poster, 22 pre-retrieval predictors are categorized and assessed on three different TREC test collections. Claudia Hauff, Djoerd Hiemstra, Franciska de Jong |
CIKM | 1 |
| 2008 | Improved query difficulty prediction for the webabstractQuery performance prediction aims to predict whether a query will have a high average precision given retrieval from a particular collection, or low average precision. An accurate estimator of the quality of search engine results can allow the search engine to decide to which queries to apply query expansion, for which queries to suggest alternative search terms, to adjust the sponsored results, or to return results from specialized collections. In this paper we present an evaluation of state of the art query prediction algorithms, both post-retrieval and pre-retrieval and we analyze their sensitivity towards the retrieval algorithm. We evaluate query difficulty predictors over three widely different collections and query sets and present an analysis of why prediction algorithms perform significantly worse on Web data. Finally we introduce Improved Clarity, and demonstrate that it outperforms state-of-the-art predictors on three standard collections, including two large Web collections. Claudia Hauff, Vanessa Murdock 0001, Ricardo Baeza-Yates |
CIKM | 1 |
| 2005 | Age Dependent Document Priors in Link Structure Analysis
Claudia Hauff, Leif Azzopardi |
ECIR | 1 |