Suzan Verberne

dblp:86/5095 · DBLP profile ↗
← Back
54ranked-venue papers in the field
15as first author
36since 2021 · last 2026
0000-0002-9609-9505ORCID · verified

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 53 (15 first)Knowledge Engineering, Semantic Web & Information Systems · 1
YearPublicationVenuePosition
2026 Generative Retrieval with Few-Shot Indexing
Arian Askari, Chuan Meng, Mohammad Aliannejadi, Zhaochun Ren, Evangelos Kanoulas, Suzan Verberne
ECIR (2)6
2026 LANCER: LLM Reranking for Nugget Coverage
Jia-Huei Ju, François G. Landry, Eugene Yang 0001, Suzan Verberne, Andrew Yates
ECIR (2)4
2026 How Role-Play Shapes Relevance Judgment in Zero-Shot LLM Rankers
Yumeng Wang 0001, Jirui Qi, Catherine Chen 0001, Panagiotis Eustratiadis, Suzan Verberne
ECIR (1)5
2026 Second Workshop on Explainability in Information Retrieval
abstract
As models grow more complex and societal demands for transparency increase with emerging regulations, explainability has become an increasingly important research area. However, despite its recognized relevance, progress in explainability research in information retrieval (IR) has been slower than in related fields. This full day workshop aims to advance research in explainable IR by providing a more in-depth platform to reflect on recent developments and facilitate discussions across both new and persistent challenges. Building upon the first edition of the workshop, which was a great success in bringing together multiple perspectives on explainability in IR, this second edition will focus on synthesizing a common agenda for the research community. Through a set of interactive activities, the workshop will bring together a diverse group of researchers to build a shared understanding of key tasks and challenges, and to help shape future directions for explainable IR research. The workshop will have as concrete outcomes a roadmap document and a special issue proposal for a journal issue on explainability in IR.
Catherine Chen 0001, Maria Heuss, Tanya Chowdhury, James Allan 0001, Avishek Anand, Carsten Eickhoff, Suzan Verberne
SIGIR7
2026 Differentiable Semantic ID for Generative Recommendation
abstract
Generative recommendation provides a novel paradigm in which each item is represented by a discrete semantic ID (SID) learned from rich content. Most methods treat SIDs as predefined and train recommenders under static indexing. In practice, SIDs are optimized only for content reconstruction rather than recommendation accuracy. This leads to an objective mismatch : the system optimizes an indexing loss to learn the SID, and a recommendation loss for interaction prediction, but because the tokenizer is trained independently, the recommendation loss cannot update it. A natural approach is to make semantic indexing differentiable so recommendation gradients can directly influence SID learning, but this often causes codebook collapse with only a few codes used. We attribute this to early deterministic assignments that limit codebook exploration, leading to imbalance and unstable optimization. In this paper, we therefore propose DIGER (Differentiable Semantic ID for GEnerative Recommendation). DIGER is a first step towards an effective differentiable semantic ID for generative recommendation. The Gumbel noise explicitly encourages early-stage exploration over codes, mitigating collapse and improving code utilization. To better balance exploration and convergence, we introduce two uncertainty decay strategies that reduce the Gumbel noise, enabling a gradual shift from early-stage exploration to the exploitation of learned SIDs. Extensive experiments across multiple public datasets demonstrate consistent improvements from differentiable semantic ID. These results confirm the effectiveness of aligning indexing and recommendation objectives through differentiable SIDs. This identifies differentiable SID as a promising area of study. Our code is released under https://github.com/junchen-fu/DIGER.
Junchen Fu, Xuri Ge, Alexandros Karatzoglou, Ioannis Arapakis, Suzan Verberne, Joemon M. Jose, Zhaochun Ren
SIGIR5
2026 Search for Coverage: Learning Coverage-Aware Retrieval with Augmented Sub-Question Answerability
abstract
Long-form Retrieval-Augmented Generation (RAG) brings the challenge of coverage-based ranking, because ranking methods must ensure the inclusion of comprehensive relevant nuggets (i.e., facts), which can thereby be synthesized into a comprehensive output. In this work, we propose CoveR, a dense retrieval method optimized for coverage-aware retrieval scenarios. CoveR is a bi-encoder trained with the coverage-based contrastive and distillation objectives, which enables CoveR to capture diverse aspects of information needs. To train CoveR, we create the SCOPE dataset, which comprises 90K training pairs from Researchy Questions with synthetic coverage signals augmented from sub-question answerability judgments generated by LLMs. Our empirical experiments show that CoveR enhances nugget coverage by 10% over strong dense retrieval baselines without sacrificing its relevance-based retrieval capability. Further ablation studies validate the importance of our proposed learning method, showing that CoveR achieves a superior trade-off between relevance- and coverage-based ranking, which is essential for long-form RAG.
Jia-Huei Ju, Eugene Yang 0001, Trevor Adriaanse, Suzan Verberne, Andrew Yates
SIGIR4
2026 LUMI: Unsupervised Intent Clustering with Multiple Pseudo-Labels
abstract
Item does not contain fulltext
I-Fan Lin 0001, Faegheh Hasibi, Suzan Verberne
SIGIR3
2026 Unifying Search and Recommendation in LLMs via Gradient Multi-Subspace Tuning
abstract
Search and recommendation (S&R) are two integral components of modern online platforms, both aiming to model and satisfy user information needs. This shared objective motivates a unified modeling paradigm that enables richer user modeling and improves the effectiveness of both tasks. Recent attempts to unify S&R formulate item ranking in both tasks as conditional generation. While this paradigm is promising, existing methods rely on full fine-tuning, which is computationally expensive and limits scalability. Parameter-efficient fine-tuning (PEFT) offers a more practical alternative but faces two critical challenges in unifying S&R: (1) gradient conflicts across tasks due to divergent optimization objectives, and (2) shifts in user intent understanding caused by overfitting to fine-tuning data, which distort general-domain knowledge and weaken LLM reasoning. To address these issues, we propose Gradient Multi-Subspace Tuning (GEMS), a novel framework that unifies S&R with LLMs while alleviating gradient conflicts and preserving general-domain knowledge. GEMS introduces (1) Multi-Subspace Decomposition, which disentangles shared and task-specific optimization signals into complementary low-rank subspaces, thereby reducing destructive gradient interference, and (2) Null-Space Projection, which constrains parameter updates to a subspace orthogonal to the general-domain knowledge space, mitigating shifts in user intent understanding. Extensive experiments on benchmark datasets show that GEMS consistently outperforms the state-of-the-art baselines across both search and recommendation tasks, and the gains remain consistent when scaling to billion-parameter LLMs.
Jujia Zhao, Zihan Wang 0002, Shuaiqun Pan, Suzan Verberne, Zhaochun Ren
SIGIR4
2025 The First Workshop on Scholarly Information Access (SCOLIA)
Ingo Frommholz, Philipp Mayr 0001, Guillaume Cabanac, Suzan Verberne, Christin Kreutz
ECIR (5)4
2025 Improving RAG for Personalization with Author Features and Contrastive Examples
Mert Yazan, Suzan Verberne, Frederik Situmeang
ECIR (3)2
2025 Model Meets Knowledge: Analyzing Knowledge Types for Conversational Recommender Systems
abstract
Computer Systems, Imagery and Media
Jujia Zhao, Yumeng Wang 0001, Zhaochun Ren, Suzan Verberne
RecSys4
2025 Workshop on Explainability in Information Retrieval
abstract
As models grow more complex and societal demands for transparency increase with emerging regulations, explainability has become an even more important research area. However, despite its recognized relevance, explainability research in IR has seen slower progress than in related fields. This full day workshop aims to advance research in explainable information retrieval by providing a more in-depth platform to reflect on recent developments and facilitate discussions to address new and persistent challenges. Our goal is to bring together a diverse group of researchers to build a shared understanding of key tasks and challenges that will lay the foundation for the future of explainable IR research.
Maria Heuss, Catherine Chen 0001, Avishek Anand, Carsten Eickhoff, Suzan Verberne
SIGIR5
2025 Tool Learning in the Wild: Empowering Language Models as Automatic Tool Agents
abstract
Augmenting large language models (LLMs) with external tools has emerged as a promising approach to extend their utility, enabling them to solve practical tasks.Previous methods manually parse tool documentation and create in-context demonstrations, transforming tools into structured formats for LLMs to use in their step-by-step reasoning.However, this manual process requires domain expertise and struggles to scale to large toolsets.Additionally, these methods rely heavily on ad-hoc inference techniques or special tokens to integrate free-form LLM generation with tool-calling actions, limiting the LLM's flexibility in handling diverse tool specifications and integrating multiple tools.In this work, we propose AutoTools, a framework that enables LLMs to automate the tool-use workflow.Specifically, the LLM automatically transforms tool documentation into callable functions, verifying syntax and runtime correctness.Then, the LLM integrates these functions into executable programs to solve practical tasks, flexibly grounding tool-use actions into its reasoning processes.Extensive experiments on existing and newly collected, more challenging benchmarks illustrate the superiority of our framework.Inspired by these promising results, we further investigate how to improve the expertise of LLMs, especially opensource LLMs with fewer parameters, within AutoTools.Thus, we propose the AutoTools-Learning approach, training the LLMs with three learning tasks on 34k instances of high-quality synthetic data, including documentation understanding, relevance learning, and function programming.Fine-grained results validate the effectiveness of our overall training approach and each individual task.
Zhengliang Shi, Shen Gao, Lingyong Yan, Yue Feng 0002, Xiuyi Chen, Zhumin Chen, Dawei Yin 0001, Suzan Verberne, Zhaochun Ren
WWW8
2025 Tracing science-technology-linkages: A machine learning pipeline for extracting and matching patent in-text references to scientific publications
abstract
Patent references to science provide a valuable paper trail for investigating the knowledge flow from science to technological innovation. Research on patent–paper links has mostly concentrated on front-page references, often neglecting the more complex in-text references. Therefore, we developed a three-stage machine-learning pipeline to extract and match patent in-text references to scientific publications. Our pipeline performs the following tasks: (1) extracting reference strings from patent texts, (2) parsing fields from these reference strings, and (3) matching references to publications in the Web of Science (WoS) database. We developed a training dataset consisting of 3,900 (and 3,901) manually annotated references from 392 (and 319) randomly selected EPO (and USPTO) patents. The first stage, reference extraction, achieved almost perfect results with a precision of 98.9% and a recall of 97.7% at the reference level. Overall, the pipeline demonstrated robust performance, with a precision of 96.8% and a recall of 91.9% at the unique patent-paper-pair level. Applying this pipeline to EPO and USPTO patents granted between 1990 and 2022, we identified 5,438,836 (and 20,432,189) references from 492,469 (and 1,449,398) EPO (and USPTO) patents, 2,763,779 (and 11,069,995) of which are matched to WoS publications. This extensive dataset is a valuable resource for studying science-technology linkages. We offer open access to this dataset, along with the associated code and training data .
Zahra Abbasiantaeb, Suzan Verberne, Jian Wang 0002
Inf. Process. Manag.2
2024 Is the Search Engine of the Future a Chatbot?
abstract
The rise of Large Language Models (LLMs) has had a huge impact on the interaction of users with information. Many people argue that the age of search engines as we know them has ended, while other people argue that retrieval technology is more relevant than ever before, because we need information to be grounded in sources. In my talk I will argue that both statements are true. I will discuss the multiple relations between LLMs and Information Retrieval: how can they strengthen each other, what are the challenges we face, and what directions should we go in our research?
Suzan Verberne
CIKM1
2024 Measuring Bias in a Ranked List Using Term-Based Representations
Amin Abolghasemi, Leif Azzopardi, Arian Askari, Maarten de Rijke, Suzan Verberne
ECIR (5)5
2024 Answer Retrieval in Legal Community Question Answering
Arian Askari, Zihui Yang, Zhaochun Ren, Suzan Verberne
ECIR (3)4
2024 Bibliometric-Enhanced Information Retrieval: 14th International BIR Workshop (BIR 2024)
Ingo Frommholz, Philipp Mayr 0001, Guillaume Cabanac, Suzan Verberne
ECIR (5)4
2024 Attend All Options at Once: Full Context Input for Multi-choice Reading Comprehension
Runda Wang, Suzan Verberne, Marco Spruit
ECIR (1)2
2024 SumBlogger: Abstractive Summarization of Large Collections of Scientific Articles
Pavlos Zakkas, Suzan Verberne, Jakub Zavrel
ECIR (1)2
2024 Injecting the score of the first-stage retriever as text improves BERT-based re-rankers
abstract
Abstract In this paper we propose a novel approach for combining first-stage lexical retrieval models and Transformer-based re-rankers: we inject the relevance score of the lexical model as a token into the input of the cross-encoder re-ranker. It was shown in prior work that interpolation between the relevance score of lexical and Bidirectional Encoder Representations from Transformers (BERT) based re-rankers may not consistently result in higher effectiveness. Our idea is motivated by the finding that BERT models can capture numeric information. We compare several representations of the Best Match 25 (BM25) and Dense Passage Retrieval (DPR) scores and inject them as text in the input of four different cross-encoders. Since knowledge distillation, i.e., teacher-student training, proved to be highly effective for cross-encoder re-rankers, we additionally analyze the effect of injecting the relevance score into the student model while training the model by three larger teacher models. Evaluation on the MSMARCO Passage collection and the TREC DL collections shows that the proposed method significantly improves over all cross-encoder re-rankers as well as the common interpolation methods. We show that the improvement is consistent for all query types. We also find an improvement in exact matching capabilities over both the first-stage rankers and the cross-encoders. Our findings indicate that cross-encoder re-rankers can efficiently be improved without additional computational burden or extra steps in the pipeline by adding the output of the first-stage ranker to the model input. This effect is robust for different models and query types.
Arian Askari, Amin Abolghasemi, Gabriella Pasi, Wessel Kraaij, Suzan Verberne
Discov. Comput.5
2024 Retrieval for Extremely Long Queries and Documents with RPRS: A Highly Efficient and Effective Transformer-based Re-Ranker
abstract
Retrieval with extremely long queries and documents is a well-known and challenging task in information retrieval and is commonly known as Query-by-Document (QBD) retrieval. Specifically designed Transformer models that can handle long input sequences have not shown high effectiveness in QBD tasks in previous work. We propose a Re-Ranker based on the novel Proportional Relevance Score (RPRS) to compute the relevance score between a query and the top- k candidate documents. Our extensive evaluation shows RPRS obtains significantly better results than the state-of-the-art models on five different datasets. Furthermore, RPRS is highly efficient, since all documents can be pre-processed, embedded, and indexed before query time that gives our re-ranker the advantage of having a complexity of O(N) , where N is the total number of sentences in the query and candidate documents. Furthermore, our method solves the problem of the low-resource training in QBD retrieval tasks as it does not need large amounts of training data and has only three parameters with a limited range that can be optimized with a grid search even if a small amount of labeled data is available. Our detailed analysis shows that RPRS benefits from covering the full length of candidate documents and queries.
Arian Askari, Suzan Verberne, Amin Abolghasemi, Wessel Kraaij, Gabriella Pasi
ACM Trans. Inf. Syst.2
2023 Retrievability Bias Estimation Using Synthetically Generated Queries
abstract
Ranking with pre-trained language models (PLMs) has shown to be highly effective for various Information Retrieval tasks. Previous studies investigated the performance of these models in terms of effectiveness and efficiency. However, there is no prior work on evaluating PLM-based rankers in terms of their retrievability bias. In this paper, we evaluate the retrievability bias of PLM-based rankers with the use of synthetically generated queries. We compare the retrievability bias in two of the most common PLM-based rankers, a Bi-Encoder BERT ranker and a Cross-Encoder BERT re-ranker against BM25, which was found to be one of the least biased models in prior work. We conduct a series of experiments with which we explore the plausibility of using synthetic queries generated with a generative model, docT5query, in the evaluation of retrievability bias. Our experiments show promising results on the use of synthetically generated queries for the purpose of retrievability bias estimation. Moreover, we find that the estimated bias values resulting from synthetically generated queries are lower than the ones estimated with user-generated queries on the MS MARCO evaluation benchmark. This indicates that synthetically generated queries might cause less bias than user-generated queries and therefore, by using such queries in training PLM-based rankers, we might be able to reduce the retrievability bias in these models.
Amin Abolghasemi, Suzan Verberne, Arian Askari, Leif Azzopardi
CIKM2
2023 CLosER: Conversational Legal Longformer with Expertise-Aware Passage Response Ranker for Long Contexts
abstract
In this paper, we investigate the task of response ranking in conversational legal search. We propose a novel method for conversational passage response retrieval (ConvPR) for long conversations in domains with mixed levels of expertise. Conversational legal search is challenging because the domain includes long, multi-participant dialogues with domain-specific language. Furthermore, as opposed to other domains, there typically is a large knowledge gap between the questioner (a layperson) and the responders (lawyers), participating in the same conversation. We collect and release a large-scale real-world dataset called LegalConv with nearly one million legal conversations from a legal community question answering (CQA) platform. We address the particular challenges of processing legal conversations, with our novel Conversational Legal Longformer with Expertise-Aware Response Ranker, called CLosER. The proposed method has two main innovations compared to state-of-the-art methods for ConvPR: (i) Expertise-Aware Post-Training; a learning objective that takes into account the knowledge gap difference between participants to the conversation; and (ii) a simple but effective strategy for re-ordering the context utterances in long conversations to overcome the limitations of the sparse attention mechanism of the Longformer architecture. Evaluation on LegalConv shows that our proposed method substantially and significantly outperforms existing state-of-the-art models on the response selection task. Our analysis indicates that our Expertise-Aware PostTraining, i.e., continued pre-training or domain/task adaptation, plays an important role in the achieved effectiveness. Our proposed method is generalizable to other tasks with domain-specific challenges and can facilitate future research on conversational search in other domains.
Arian Askari, Mohammad Aliannejadi, Amin Abolghasemi, Evangelos Kanoulas, Suzan Verberne
CIKM5
2023 A Test Collection of Synthetic Documents for Training Rankers: ChatGPT vs. Human Experts
abstract
In this resource paper, we investigate the usefulness of generative Large Language Models (LLMs) in generating training data for cross-encoder re-rankers in a novel direction: generating synthetic documents instead of synthetic queries. We introduce a new dataset, ChatGPT-RetrievalQA, and compare the effectiveness of strong models fine-tuned on both LLM-generated and human-generated data. We build ChatGPT-RetrievalQA based on an existing dataset, human ChatGPT Comparison Corpus (HC3), consisting of public question collections with human responses and answers from ChatGPT. We fine-tune a range of cross-encoder re-rankers on either human-generated or ChatGPT-generated data. Our evaluation on MS MARCO DEV, TREC DL'19, and TREC DL'20 demonstrates that cross-encoder re-ranking models trained on LLM-generated responses are significantly more effective for out-of-domain re-ranking than those trained on human responses. For in-domain re-ranking, the human-trained re-rankers outperform the LLM-trained re-rankers. Our novel findings suggest that generative LLMs have high potential in generating training data for neural retrieval models and can be used to augment training data, especially in domains with smaller amounts of labeled data. We believe that our dataset, ChatGPT-RetrievalQA, presents various opportunities for analyzing and improving rankers with human and synthetic data. We release our data, code, and model checkpoints for future work.
Arian Askari, Mohammad Aliannejadi, Evangelos Kanoulas, Suzan Verberne
CIKM4
2023 Injecting the BM25 Score as Text Improves BERT-Based Re-rankers
Arian Askari, Amin Abolghasemi, Gabriella Pasi, Wessel Kraaij, Suzan Verberne
ECIR (1)5
2023 Bibliometric-Enhanced Information Retrieval: 13th International BIR Workshop (BIR 2023)
Ingo Frommholz, Philipp Mayr 0001, Guillaume Cabanac, Suzan Verberne
ECIR (3)4
2023 ECIR 2023 Workshop: Legal Information Retrieval
Suzan Verberne, Evangelos Kanoulas, Gineke Wiggers, Florina Piroi, Arjen P. de Vries
ECIR (3)1
2023 Bibliometric-enhanced legal information retrieval: Combining usage and citations as flavors of impact relevance
abstract
Abstract Bibliometric‐enhanced information retrieval uses bibliometrics (e.g., citations) to improve ranking algorithms. Using a data‐driven approach, this article describes the development of a bibliometric‐enhanced ranking algorithm for legal information retrieval, and the evaluation thereof. We statistically analyze the correlation between usage of documents and citations over time, using data from a commercial legal search engine. We then propose a bibliometric boost function that combines usage of documents with citation counts. The core of this function is an impact variable based on usage and citations that increases in influence as citations and usage counts become more reliable over time. We evaluate our ranking function by comparing search sessions before and after the introduction of the new ranking in the search engine. Using a cost model applied to 129,571 sessions before and 143,864 sessions after the intervention, we show that our bibliometric‐enhanced ranking algorithm reduces the time of a search session of legal professionals by 2 to 3% on average for use cases other than known‐item retrieval or updating behavior. Given the high hourly tariff of legal professionals and the limited time they can spend on research, this is expected to lead to increased efficiency, especially for users with extremely long search sessions.
Gineke Wiggers, Suzan Verberne, Wouter van Loon, Gerrit-Jan Zwenne
J. Assoc. Inf. Sci. Technol.2
2022 TripJudge: A Relevance Judgement Test Collection for TripClick Health Retrieval
abstract
Robust test collections are crucial for Information Retrieval research. Recently there is a growing interest in evaluating retrieval systems for domain-specific retrieval tasks, however these tasks often lack a reliable test collection with human-annotated relevance assessments following the Cranfield paradigm. In the medical domain, the TripClick collection was recently proposed, which contains click log data from the Trip search engine and includes two click-based test sets. However the clicks are biased to the retrieval model used, which remains unknown, and a previous study shows that the test sets have a low judgement coverage for the Top-10 results of lexical and neural retrieval models. In this paper we present the novel, relevance judgement test collection TripJudge for TripClick health retrieval. We collect relevance judgements in an annotation campaign and ensure the quality and reusability of TripJudge by a variety of ranking methods for pool creation, by multiple judgements per query-document pair and by an at least moderate inter-annotator agreement. We compare system evaluation with TripJudge and TripClick and find that that click and judgement-based evaluation can lead to substantially different system rankings.
Sophia Althammer, Sebastian Hofstätter, Suzan Verberne, Allan Hanbury
CIKM3
2022 Improving BERT-based Query-by-Document Retrieval with Multi-task Optimization
Amin Abolghasemi, Suzan Verberne, Leif Azzopardi
ECIR (2)2
2022 PARM: A Paragraph Aggregation Retrieval Model for Dense Document-to-Document Retrieval
Sophia Althammer, Sebastian Hofstätter, Mete Sertkan, Suzan Verberne, Allan Hanbury
ECIR (1)4
2022 Expert Finding in Legal Community Question Answering
Arian Askari, Suzan Verberne, Gabriella Pasi
ECIR (2)2
2022 Bibliometric-enhanced Information Retrieval: 12th International BIR Workshop (BIR 2022)
Ingo Frommholz, Philipp Mayr 0001, Guillaume Cabanac, Suzan Verberne
ECIR (2)4
2022 Preprocessing Requirements Documents for Automatic UML Modelling
Martijn B. J. Schouten, Guus J. Ramackers, Suzan Verberne
NLDB3
2021 Bibliometric-Enhanced Information Retrieval: 11th International BIR Workshop
Ingo Frommholz, Philipp Mayr 0001, Guillaume Cabanac, Suzan Verberne
ECIR (2)4
2018 First International Workshop on Professional Search (ProfS2018)
abstract
Professional search is a problem area in which many facets of information retrieval are addressed, both system-related (e.g. distributed search) and user-related (e.g. complex information needs), and the interface between user and system (e.g. supporting exploratory search tasks). Professional search tasks have specific requirements, different from the requirements of generic web search. The aim of this workshop is to bring together researchers to work on the requirements and challenges of professional search from different angles. We will have an interactive workshop where researchers not only present their scientific results but also work together on the definition of future challenges and solutions with input from information professionals. The workshop will deliver a roadmap of research directions for the years to come.
Suzan Verberne, Jiyin He, Udo Kruschwitz, Birger Larsen, Tony Russell-Rose, Arjen P. de Vries
SIGIR1
2017 Automatic Summarization of Domain-specific Forum Threads: Collecting Reference Data
abstract
We create and analyze two sets of reference summaries for discussion threads on a patient support forum: expert summaries and crowdsourced, non-expert summaries. Ideally, reference summaries for discussion forum threads are created by expert members of the forum community. When there are few or no expert members available, crowdsourcing the reference summaries is an alternative. In this paper we investigate whether domain-specific forum data requires the hiring of domain experts for creating reference summaries. We analyze the inter-rater agreement for both data-sets and we train summarization models using the two types of reference summaries. The inter-rater agreement in crowdsourced reference summaries is low, close to random, while domain experts achieve a considerably higher, fair, agreement. The trained models however are similar to each other. We conclude that it is possible to train an extractive summarization model on crowdsourced data that is similar to an expert model, even if the inter-rater agreement for the crowdsourced data is low.
Suzan Verberne, Antal van den Bosch, Sander Wubben, Emiel Krahmer
CHIIR1
2017 Evaluation of context-aware recommendation systems for information re-finding
abstract
In this article we evaluate context‐aware recommendation systems for information re‐finding by knowledge workers. We identify 4 criteria that are relevant for evaluating the quality of knowledge worker support: context relevance, document relevance, prediction of user action, and diversity of the suggestions. We compare 3 different context‐aware recommendation methods for information re‐finding in a writing support task. The first method uses contextual prefiltering and content‐based recommendation (CBR), the second uses the just‐in‐time information retrieval paradigm (JITIR), and the third is a novel network‐based recommendation system where context is part of the recommendation model (CIA). We found that each method has its own strengths: CBR is strong at context relevance, JITIR captures document relevance well, and CIA achieves the best result at predicting user action. Weaknesses include that CBR depends on a manual source to determine the context and in JITIR the context query can fail when the textual content is not sufficient. We conclude that to truly support a knowledge worker, all 4 evaluation criteria are important. In light of that conclusion, we argue that the network‐based approach the CIA offers has the highest robustness and flexibility for context‐aware information recommendation.
Maya Sappelli, Suzan Verberne, Wessel Kraaij
J. Assoc. Inf. Sci. Technol.2
2016 Longitudinal Navigation Log Data on a Large Web Domain
abstract
We have collected the access logs for our university's web domain over a time span of 4.5 years. We now release the pre-processed data of a 3-month period for research into user navigation behavior. We preprocessed the data so that only successful GET requests of web pages by non-bot users are kept. The resulting 3-month collection comprises 9.6M page visits (190K unique URLs) by 744K unique visitors.
Suzan Verberne, Bram Arends, Wessel Kraaij, Arjen P. de Vries
SIGIR1
2016 Evaluation and analysis of term scoring methods for term extraction
abstract
We evaluate five term scoring methods for automatic term extraction on four different types of text collections: personal document collections, news articles, scientific articles and medical discharge summaries. Each collection has its own use case: author profiling, boolean query term suggestion, personalized query suggestion and patient query expansion. The methods for term scoring that have been proposed in the literature were designed with a specific goal in mind. However, it is as yet unclear how these methods perform on collections with characteristics different than what they were designed for, and which method is the most suitable for a given (new) collection. In a series of experiments, we evaluate, compare and analyse the output of six term scoring methods for the collections at hand. We found that the most important factors in the success of a term scoring method are the size of the collection and the importance of multi-word terms in the domain. Larger collections lead to better terms; all methods are hindered by small collection sizes (below 1000 words). The most flexible method for the extraction of single-word and multi-word terms is pointwise Kullback–Leibler divergence for informativeness and phraseness. Overall, we have shown that extracting relevant terms using unsupervised term scoring methods is possible in diverse use cases, and that the methods are applicable in more contexts than their original design purpose.
Suzan Verberne, Maya Sappelli, Djoerd Hiemstra, Wessel Kraaij
Inf. Retr. J.1
2016 Assessing e-mail intent and tasks in e-mail messages
Maya Sappelli, Gabriella Pasi, Suzan Verberne, Maaike de Boer, Wessel Kraaij
Inf. Sci.3
2015 User Simulations for Interactive Search: Evaluating Personalized Query Suggestion
Suzan Verberne, Maya Sappelli, Kalervo Järvelin, Wessel Kraaij
ECIR1
2014 Query Term Suggestion in Academic Search
Suzan Verberne, Maya Sappelli, Wessel Kraaij
ECIR1
2014 Automatic thematic classification of election manifestos
Suzan Verberne, Eva D'hondt, Antal van den Bosch, Maarten Marx
Inf. Process. Manag.1
2014 Dealing with temporal variation in patent categorization
Eva D'hondt, Suzan Verberne, Nelleke Oostdijk, Jean Beney, Cornelis H. A. Koster, Lou Boves
Inf. Retr.2
2013 Recommending personalized touristic sights using google places
abstract
The purpose of the Contextual Suggestion track, an evaluation task at the TREC 2012 conference, is to suggest personalized tourist activities to an individual, given a certain location and time. In our content-based approach, we collected initial recommendations using the location context as search query in Google Places. We first ranked the recommendations based on their textual similarity to the user profiles. In order to improve the ranking of popular sights, we combined the initial ranking with rankings based on Google Search, popularity and categories. Finally, we performed filtering based on the temporal context. Overall, our system performed well above average and median, and outperformed the baseline - Google Places only -- run.
Maya Sappelli, Suzan Verberne, Wessel Kraaij
SIGIR2
2013 Reliability and validity of query intent assessments
abstract
In most intent recognition studies, annotations of query intent are created post hoc by external assessors who are not the searchers themselves. It is important for the field to get a better understanding of the quality of this process as an approximation for determining the searcher's actual intent. Some studies have investigated the reliability of the query intent annotation process by measuring the interassessor agreement. However, these studies did not measure the validity of the judgments, that is, to what extent the annotations match the searcher's actual intent. In this study, we asked both the searchers themselves and external assessors to classify queries using the same intent classification scheme. We show that of the seven dimensions in our intent classification scheme, four can reliably be used for query annotation. Of these four, only the annotations on the topic and spatial sensitivity dimension are valid when compared with the searcher's annotations. The difference between the interassessor agreement and the assessor‐searcher agreement was significant on all dimensions, showing that the agreement between external assessors is not a good estimator of the validity of the intent classifications. Therefore, we encourage the research community to consider using query intent classifications by the searchers themselves as test data.
Suzan Verberne, Maarten van der Heijden, Max Hinne, Maya Sappelli, Saskia Koldijk, Eduard Hoenkamp, Wessel Kraaij
J. Assoc. Inf. Sci. Technol.1
2011 Bringing Why-QA to Web Search
Suzan Verberne, Lou Boves, Wessel Kraaij
ECIR1
2011 Learning to rank for why-question answering
abstract
In this paper, we evaluate a number of machine learning techniques for the task of ranking answers to why-questions. We use TF-IDF together with a set of 36 linguistically motivated features that characterize questions and answers. We experiment with a number of machine learning techniques (among which several classifiers and regression techniques, Ranking SVM and SVM map ) in various settings. The purpose of the experiments is to assess how the different machine learning approaches can cope with our highly imbalanced binary relevance data, with and without hyperparameter tuning. We find that with all machine learning techniques, we can obtain an MRR score that is significantly above the TF-IDF baseline of 0.25 and not significantly lower than the best score of 0.35. We provide an in-depth analysis of the effect of data imbalance and hyperparameter tuning, and we relate our findings to previous research on learning to rank for Information Retrieval.
Suzan Verberne, Hans van Halteren, Daphne Theijssen, Stephan Raaijmakers, Lou Boves
Inf. Retr.1
2009 Annotation of URLs: more than the sum of parts
abstract
Recently a number of studies have demonstrated that search engine logfiles are an important resource to determine the relevance relation between URLs and query terms. We hypothesized that the queries associated with a URL could also be presented as useful URL metadata in a search engine result list, e.g. for helping to determine the semantic category of a URL. We evaluated this hypothesis by a classification experiment based on the DMOZ dataset. Our method can also annotate URLs that have no associated queries.
Max Hinne, Wessel Kraaij, Stephan Raaijmakers, Suzan Verberne, Theo P. van der Weide, Maarten van der Heijden
SIGIR4
2008 Evaluating Paragraph Retrieval for
Suzan Verberne, Lou Boves, Nelleke Oostdijk, Peter-Arno Coppen
ECIR1
2007 Paragraph retrieval for why-question answering
abstract
No abstract available.
Suzan Verberne
SIGIR1
2007 Evaluating discourse-based answer extraction for why-question answering
abstract
No abstract available.
Suzan Verberne, Lou Boves, Nelleke Oostdijk, Peter-Arno Coppen
SIGIR1