Marijn Koolen

dblp:82/1466 · DBLP profile ↗
← Back
37ranked-venue papers in the field
11as first author
8since 2021 · last 2026
0000-0002-0301-2029ORCID · verified

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 34 (9 first)Data Mining & Knowledge Discovery · 2 (1 first)Knowledge Engineering, Semantic Web & Information Systems · 1 (1 first)
YearPublicationVenuePosition
2026 Tip-of-the-Tongue Search in the Wild: Analyzing Human and LLM Performance and Success Factors on Complex Search Requests
abstract
Users often turn to online forums when searching for known books, movies, or games that they cannot identify through conventional search engines. These “tip-of-the tongue” requests present a unique challenge, appearing highly variable in formulation, context, and specificity. So far, these could mostly only be solved by other humans answering in forums. Generative AI is believed to help solve these specific questions. In this work, we manually annotated 150 requests each for books, games, and movies in the casual leisure domain to study the differences between solved and unsolved requests and identify factors that influence their difficulty. We compare human responses in forum threads with the performance of a Large Language Model (LLM) under similar conditions. Specifically, we investigate how the formulation of requests affects human and LLM success; how item properties impact LLM retrieval; how interaction and feedback within a thread shape human and LLM performance; and whether increasing the information provided to an LLM improves its chances of solving the request. Our findings offer new insights into what makes these known-item search problems easier or harder to solve. This study contributes to a better understanding of complex search behavior and the role of LLMs in helping with difficult casual-leisure information needs.
Toine Bogers, Maria Gäde, Mark M. Hall, Marijn Koolen, Vivien Petras, Mette Skov
CHIIR4
2025 Exploring the Zero-Shot Known-Item Retrieval Capabilities of LLMs for Casual Leisure Information Needs
abstract
The rapidly increasing popularity of LLM-powered chatbots has led to them being used for a increasing number of different tasks by the general public.One of these tasks is searching for information instead of using a search engine.Previous work has shown that complex search tasks can be problematic for traditional search engines to solve, but little is known about the capability of LLMs on the same task.We compared four LLMs on their capability to answer a specific type of complex search task: known-item requests from casual leisure domains.We constructed a test collection by gathering known-item requests for books, games and movies from online forums along with verified answers by the original requester.We prompted four LLMs multiple times with the same prompt and analyzed the results with respect to accuracy and the degree to which answers were fabricated by the LLM.Our results show that LLMs are not particularly effective in fulfilling these complex casual leisure needs, but there are are big differences between LLMs and across domains.
Toine Bogers, Maria Gäde, Mark M. Hall, Marijn Koolen, Vivien Petras, Mette Skov
CHIIR4
2025 A Corpus of Early Modern Decision-Making - the Resolutions of the States General of the Dutch Republic
abstract
This paper presents a corpus of early modern Dutch resolutions made in the daily meetings of the States General, the central governing body of the Dutch Republic, over a period of 220 years, from 1576 to 1796. This corpus has been digitised from over half a million scans of mostly handwritten text, segmented into individual resolutions (decisions) and enriched with named entities and metadata extracted from the text of the resolutions. We developed a pipeline for automatic text recognition for historic Dutch, and a document segmentation approach that combines ML classifiers trained on annotated data with rule-based fuzzy matching of the highly formulaic language of the resolutions. The decisions that the States General made were often based on propositions (requests or proposals) submitted in writing, by other governing bodies and by citizens of the republic. The resolutions contain information about these submitted propositions, including the persons and organisations who submitted them. The second part of this paper includes an analysis of the information about these proposition documents that can be extracted from the resolutions, and the potential to link the resolutions to their corresponding propositions using named entities and extracted metadata. We find that for the overwhelming majority of propositions, we can identify the name of person or organisation who submitted it, making it feasible to (semi-)automatically link the resolutions to their corresponding proposition documents. This will allow historians and genealogists to study not only the decision making of the States General in the early modern period, but also the concerns put forward by both high-ranking officials and regular citizens of the Republic.
Marijn Koolen, Rik Hoekstra
LDK1
2024 The Impact of CHIIR Publications: A Study of Eight Years of CHIIR
abstract
Across all scientific fields, there is an increased focus on the impact of scientific research: what academic and societal benefits does it provide? This question has spurred the development of a variety of different approaches to impact assessment, each appropriate in different circumstances. In this paper, we study the academic impact of the CHIIR community through a comprehensive analysis of the work published in the 2016-2023 CHIIR conference series. We collect citation counts, citing documents, and altmetrics scores for all CHIIR publications to determine their academic impact across a variety of different attributes of the CHIIR publications. In addition, we analyze a subset of citation contexts in the papers that have cited CHIIR publications to analyze how they are being used and what that means for their potential impact. Finally, we attempt to predict which properties of CHIIR publications are most predictive of future impact.
Maria Gäde, Toine Bogers, Mark M. Hall, Marijn Koolen, Vivien Petras, Birger Larsen
CHIIR4
2023 How we Work, Share, and Re-use at CHIIR
abstract
In this paper, we present the results of an initial study of the research, sharing, and re-use practices at the CHIIR conference through a systematic analysis of all CHIIR papers published from 2016 to 2022. We find that CHIIR is a conference predominantly focused on empirical, multi-methods research that over the years has undergone a focusing in terms of the type of research methods that are being used. A modest number of papers re-use existing data and design resources, but infrastructure component re-use is much more rare. Only a fraction of CHIIR papers actually share their own resources, which suggests that there is much to gain in terms of reproducibility of research presented at CHIIR and could potentially be used to support changes in reviewing practices.
Toine Bogers, Maria Gäde, Mark M. Hall, Marijn Koolen, Vivien Petras, Birger Larsen
CHIIR4
2023 Collaboration Patterns and Impact of Sharing at CHIIR
abstract
We studied the collaboration patterns of CHIIR authors, and found that most papers are collaborative. A core of 33% of the CHIIR researchers are directly connected and frequently co-author, and several disconnected clusters also make frequent CHIIR contributions. We also studied citation impact of the CHIIR papers and show that in relation to research design type, theoretical and empirical papers tend to receive more citations than resource papers. With regards to sharing and re-use, papers that share at least one resource tend to have significantly higher citation impact—in particular when sharing data resources and design resources. Re-using resources does not significantly increase citation impact in itself.
Toine Bogers, Birger Larsen, Marijn Koolen, Maria Gäde, Mark M. Hall, Vivien Petras
CHIIR3
2022 Third Workshop on Building towards Information Interaction and Retrieval Resources Re-use (BIIRRR 2022)
abstract
Work in Progress Share on Third Workshop on Building towards Information Interaction and Retrieval Resources Re-use (BIIRRR 2022) Authors: Toine Bogers Aalborg University, Denmark Aalborg University, DenmarkView Profile , Maria Gäde Humboldt-Universität zu Berlin, Germany Humboldt-Universität zu Berlin, GermanyView Profile , Mark Michael Hall The Open University, United Kingdom The Open University, United KingdomView Profile , Marijn Koolen Huygens Institute for the History of the Netherlands, Royal Netherlands Academy of Arts and Sciences, Netherlands Huygens Institute for the History of the Netherlands, Royal Netherlands Academy of Arts and Sciences, NetherlandsView Profile , Vivien Petras Humboldt-Universität zu Berlin, Germany Humboldt-Universität zu Berlin, GermanyView Profile , Paul Thomas Microsoft, Australia Microsoft, AustraliaView Profile Authors Info & Claims CHIIR '22: ACM SIGIR Conference on Human Information Interaction and RetrievalMarch 2022 Pages 374–376https://doi.org/10.1145/3498366.3505838Published:14 March 2022Publication History 0citation24DownloadsMetricsTotal Citations0Total Downloads24Last 12 Months24Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Toine Bogers, Maria Gäde, Mark M. Hall, Marijn Koolen, Vivien Petras, Paul Thomas 0001
CHIIR4
2021 A Manifesto on Resource Re-Use in Interactive Information Retrieval
abstract
This perspective paper on resource re-use intends to draw the attention of the interactive information retrieval (IIR) community to the challenges of research documentation and archiving for future use. Resources are understood as encompassing research designs, research data and research infrastructures. It proposes eight principles for improving the re-use of resources in the IIR community and presents concrete steps on how to achieve them. A five-level system for data archiving and documentation envisions increasingly open and stable documentation and access infrastructures.
Maria Gäde, Marijn Koolen, Mark M. Hall, Toine Bogers, Vivien Petras
CHIIR2
2020 A Workflow Analysis Perspective to Scholarly Research Tasks
abstract
Since the appearance of digital research infrastructures in the humanities in the last decade, important efforts are being made to understand and model scholarly processes. Different methods are used in those investigations, which often result in abstract representations of research phases, taxonomies of scholarly activities, in conceptual frameworks, or in scholarly ontologies. While the aim of these representations is to inform the design of the digital infrastructures, the complexity and diversity of scholarly work pose the question about the applicability of those models for design and evaluation of research infrastructures and tools. In this paper, we explore a methodology to analyze workflows from a micro-perspective, which aims at capturing the transitions between activities. We use two scholarly projects as case studies, describe their research activities in detail by using existing ontologies and describe the connections between activities, and analyse generic transitions. We discuss what kinds of implications this approach has to evaluation and design of information systems and services to facilitate scholars' complex and varied research processes.
Marijn Koolen, Sanna Kumpulainen, Liliana Melgar
CHIIR1
2020 ComplexRec 2020: Workshop on Recommendation in Complex Environments
abstract
During the past decade, recommender systems have rapidly become an indispensable element of websites, apps, and other platforms that are looking to provide personalized interaction to their users. As recommendation technologies are applied to an ever-growing array of non-standard problems and scenarios, researchers and practitioners are also increasingly faced with challenges of dealing with greater variety and complexity in the inputs to those recommender systems. For example, there has been more reliance on fine-grained user signals as inputs rather than simple ratings or likes. Many applications also require more complex domain-specific constraints on inputs to the recommender systems. The outputs of recommender systems are also moving towards more complex composite items, such as package or sequence recommendations. This increasing complexity requires smarter recommender algorithms that can deal with this diversity in inputs and outputs. The ComplexRec workshop series offers an interactive venue for discussing approaches to recommendation in complex scenarios that have no simple one-size-fits-all solution.
Toine Bogers, Marijn Koolen, Casper Petersen, Bamshad Mobasher, Alexander Tuzhilin
RecSys2
2019 Workshop on Barriers to Interactive IR Resources Re-use (BIIRRR 2019)
abstract
Share on Workshop on Barriers to Interactive IR Resources Re-use (BIIRRR 2019) Authors: Toine Bogers Aalborg University Copenhagen, Copenhagen, Denmark Aalborg University Copenhagen, Copenhagen, DenmarkView Profile , Samuel Dodson University of British Columbia, Vancouver, Canada University of British Columbia, Vancouver, CanadaView Profile , Luanne Freund University of British Columbia, Vancouver, Canada University of British Columbia, Vancouver, CanadaView Profile , Maria Gäde Humboldt-Universität zu Berlin, Berlin, Germany Humboldt-Universität zu Berlin, Berlin, GermanyView Profile , Mark Hall Martin-Luther-Universität Halle-Wittenberg, Halle, Germany Martin-Luther-Universität Halle-Wittenberg, Halle, GermanyView Profile , Marijn Koolen Royal Netherlands Academy of Arts and Sciences, Amsterdam, Netherlands Royal Netherlands Academy of Arts and Sciences, Amsterdam, NetherlandsView Profile , Vivien Petras Humboldt-Universität zu Berlin, Berlin, Germany Humboldt-Universität zu Berlin, Berlin, GermanyView Profile , Nils Pharo Oslo Metropolitan University, Oslo, Norway Oslo Metropolitan University, Oslo, NorwayView Profile , Mette Skov Aalborg University, Aalborg, Denmark Aalborg University, Aalborg, DenmarkView Profile Authors Info & Claims CHIIR '19: Proceedings of the 2019 Conference on Human Information Interaction and RetrievalMarch 2019 Pages 389–392https://doi.org/10.1145/3295750.3298965Published:08 March 2019Publication History 1citation75DownloadsMetricsTotal Citations1Total Downloads75Last 12 Months8Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Toine Bogers, Samuel Dodson, Luanne Sinnamon, Maria Gäde, Mark M. Hall, Marijn Koolen, Vivien Petras, Nils Pharo, Mette Skov
CHIIR6
2019 The CLARIAH Media Suite: a Hybrid Approach to System Design in the Humanities
abstract
The practices of digital humanists are evolving, highly diversified and experimental. There is also a lack of agreement about whether or not digital humanists should have data and programming skills. Thus, their underlying needs for higher levels of flexibility and transparency may be contradicted by their explicit requests for user-friendly graphic user interfaces (GUIs), creating challenges for designing information systems in the digital humanities. This paper describes the experience of designing the Media Suite, which provides access to important Dutch audiovisual collections and is part of the Dutch infrastructure for digital humanities. We outline a solution to the conflicting needs of scholars, by combining a semi-traditional GUI with Jupyter Notebooks. This solution tackles the needs of both novice and advanced users in digital research methods in the humanities. This demonstration paper explains how the Media Suite and the Jupyter notebooks work together, and elaborates on the rationale behind the design choices. We also outline the implications this hybrid and extensible approach has for interface design for the information science and scholarly community.
Liliana Melgar, Marijn Koolen, Kaspar Beelen, Hugo C. Huurdeman, Mari Wigham, Carlos Martinez-Ortiz, Jaap Blom, Roeland Ordelman
CHIIR2
2019 Third workshop on recommendation in complex scenarios (ComplexRec 2019)
abstract
Over the past decade, recommendation algorithms for ratings prediction and item ranking have steadily matured. However, these state-of-the-art algorithms are typically applied in relatively straightforward and static scenarios: given information about a user's past item preferences in isolation, can we predict whether they will like a new item or rank all unseen items based on predicted interest? In reality, recommendation is often a more complex problem: the evaluation of a list of recommended items never takes place in a vacuum, and it is often a single step in the user's more complex background task or need. The goal of the ComplexRec 2019 workshop is to offer an interactive venue for discussing approaches to recommendation in complex scenarios that have no simple one-size-fits-all solution.
Marijn Koolen, Toine Bogers, Bamshad Mobasher, Alexander Tuzhilin
RecSys1
2018 Workshop on Barriers to Interactive IR Resources Re-use
abstract
The goal of this workshop is to serve as a starting point for a community-driven effort to design and implement a platform for the collection, organization, maintenance, and sharing of resources for IIR experimentation. As in all scientific endeavors, progress in IIR research is contingent on the ability to build on previous ideas, approaches, and resources. However, we believe there to be a number of barriers to reproducibility and re-use of resources in IIR research: the fragmentary nature of how the community»s resources are organized, the lack of awareness of their existence, documentation and organization of the resources, the nature of the typical research publication cycle, and the effort required to make such resources available. We believe that an online platform dedicated to the collection and organization of IIR resources could be a promising way of overcoming these barriers. The workshop therefore aims to serve both as a brainstorming opportunity about the shape this iRepository should take, as well as a way of building support in the community for its implementation.
Toine Bogers, Maria Gäde, Luanne Sinnamon, Mark M. Hall, Marijn Koolen, Vivien Petras, Mette Skov
CHIIR5
2018 2nd workshop on recommendation in complex scenarios (complexrec 2018)
abstract
Over the past decade, recommendation algorithms for ratings prediction and item ranking have steadily matured. However, these state-of-the-art algorithms are typically applied in relatively straightforward scenarios. In reality, recommendation is often a more complex problem: it is usually just a single step in the user's more complex background need. These background needs can often place a variety of constraints on which recommendations are interesting to the user and when they are appropriate. However, relatively little research has been done on these complex recommendation scenarios. The ComplexRec 2018 workshop addresses this by providing an interactive venue for discussing approaches to recommendation in complex scenarios that have no simple one-size-fits-all solution.
Toine Bogers, Marijn Koolen, Bamshad Mobasher, Alan Said, Casper Petersen
RecSys2
2017 Second Workshop on Supporting Complex Search Tasks
abstract
There is broad consensus in the field of IR that search is complex in many use cases and applications, both on the Web and in domain specific collections, and both professionally and in our daily life. Yet our understanding of complex search tasks, in comparison to simple look up tasks, is fragmented at best. The workshop addresses many open research questions: What are the obvious use cases and applications of complex search? What are essential features of work tasks and search tasks to take into account? And how do these evolve over time--With a multitude of information, varying from introductory to specialized, and from authoritative to speculative or opinionated, when to show what sources of information? How does the information seeking process evolve and what are relevant differences between different stages? With complex task and search process management, blending searching, browsing, and recommendations, and supporting exploratory search to sensemaking and analytics, UI and UX design pose an overconstrained challenge. How do we evaluate and compare approaches? Which measures should be taken into account? Supporting complex search tasks requires new collaborations across the fields of CHI and IR, and the proposed workshop will bring together a diverse group of researchers to work together on one of the greatest challenges of our field.
Nicholas J. Belkin, Toine Bogers, Jaap Kamps, Diane Kelly 0001, Marijn Koolen, Emine Yilmaz
CHIIR5
2017 A Process Model of Scholarly Media Annotation
abstract
Annotation has been identified as one of the "scholarly primitives", and plays a pivotal role in facilitating access to audio-visual (AV) media in a scholarly context. However, there is a lack of understanding of scholars' annotation needs and behavior. This paper is part of a group of studies aiming to understand how to improve annotation support of AV media, in order to facilitate research activities of media scholars and other scholars who make intensive use of AV media.
Liliana Melgar, Marijn Koolen, Hugo C. Huurdeman, Jaap Blom
CHIIR2
2017 Defining and Supporting Narrative-driven Recommendation
abstract
Research into recommendation algorithms has made great strides in recent years. However, these algorithms are typically applied in relatively straightforward scenarios: given information about a user's past preferences, what will they like in the future? Recommendation is often more complex: evaluating recommended items never takes place in a vacuum, and it is often a single step in the user's more complex background task. In this paper, we define a specific type of recommendation scenario called narrative-driven recommendation, where the recommendation process is driven by both a log of the user's past transactions as well as a narrative description of their current interest(s). Through an analysis of a set of real-world recommendation narratives from the LibraryThing forums, we demonstrate the uniqueness and richness of this scenario and highlight common patterns and properties of such narratives.
Toine Bogers, Marijn Koolen
RecSys2
2017 Workshop on Recommendation in Complex Scenarios: (ComplexRec 2017)
abstract
Recommendation algorithms for ratings prediction and item ranking have steadily matured during the past decade. However, these state-of-the-art algorithms are typically applied in relatively straightforward scenarios. In reality, recommendation is often a more complex problem: it is usually just a single step in the user's more complex background need. These background needs can often place a variety of constraints on which recommendations are interesting to the user and when they are appropriate. However, relatively little research has been done on these complex recommendation scenarios. The ComplexRec 2017 workshop addressed this by providing an interactive venue for discussing approaches to recommendation in complex scenarios that have no simple one-size-fits-all-solution.
Toine Bogers, Marijn Koolen, Bamshad Mobasher, Alan Said, Alexander Tuzhilin
RecSys2
2016 Third Workshop on New Trends in Content-based Recommender Systems (CBRecSys 2016)
abstract
While content-based recommendation has been applied successfully in many different domains, it has not seen the same level of attention as collaborative filtering techniques have. However, there are many recommendation domains and applications where content and metadata play a key role, either in addition to or instead of ratings and implicit usage data. For some domains, such as movies, the relationship between content and usage data has seen thorough investigation already, but for many other domains, such as books, news, scientific articles, and Web pages we still do not know if and how these data sources should be combined to provided the best recommendation performance. The CBRecSys 2016 workshop provides a dedicated venue for papers dedicated to all aspects of content-based recommendation.
Toine Bogers, Marijn Koolen, Cataldo Musto, Pasquale Lops, Giovanni Semeraro
RecSys2
2015 Supporting Complex Search Tasks - ECIR 2015 Workshop
Maria Gäde, Mark M. Hall, Hugo C. Huurdeman, Jaap Kamps, Marijn Koolen, Mette Skov, Elaine Toms, David Walsh 0001
ECIR5
2015 Looking for Books in Social Media: An Analysis of Complex Search Requests
Marijn Koolen, Toine Bogers, Antal van den Bosch, Jaap Kamps
ECIR1
2015 Second Workshop on New Trends in Content-based Recommender Systems (CBRecSys 2015)
Toine Bogers, Marijn Koolen
RecSys2
2014 "User Reviews in the Search Index? That'll Never Work!"
Marijn Koolen
ECIR1
2014 Workshop on new trends in content-based recommender systems: (CBRecSys 2014)
abstract
While content-based recommendation has been applied successfully in many different domains, it has not seen the same level of attention as collaborative filtering techniques have. However, there are many recommendation domains and applications where content and metadata play a key role, either in addition to or instead of ratings and implicit usage data. For some domains, such as movies, the relationship between content and usage data has seen thorough investigation already, but for many other domains, such as books, news, scientific articles, and Web pages we still do not know if and how these data sources should be combined to provided the best recommendation performance. The CBRecSys 2014 workshop aims to address this by providing a dedicated venue for papers dedicated to all aspects of content-based recommendation.
Toine Bogers, Marijn Koolen, Iván Cantador
RecSys2
2012 Social book search: comparing topical relevance judgements and book suggestions for evaluation
abstract
The Web and social media give us access to a wealth of information, not only different in quantity but also in character---traditional descriptions from professionals are now supplemented with user generated content. This challenges modern search systems based on the classical model of topical relevance and ad hoc search: How does their effectiveness transfer to the changing nature of information and to the changing types of information needs and search tasks? We use the INEX 2011 Books and Social Search Track's collection of book descriptions from Amazon and social cataloguing site LibraryThing. We compare classical IR with social book search in the context of the LibraryThing discussion forums where members ask for book suggestions. Specifically, we compare book suggestions on the forum with Mechanical Turk judgements on topical relevance and recommendation, both the judgements directly and their resulting evaluation of retrieval systems. First, the book suggestions on the forum are a complete enough set of relevance judgements for system evaluation. Second, topical relevance judgements result in a different system ranking from evaluation based on the forum suggestions. Although it is an important aspect for social book search, topical relevance is not sufficient for evaluation. Third, professional metadata alone is often not enough to determine the topical relevance of a book. User reviews provide a better signal for topical relevance. Fourth, user-generated content is more effective for social book search than professional metadata. Based on our findings, we propose an experimental evaluation that better reflects the complexities of social book search.
Marijn Koolen, Jaap Kamps, Gabriella Kazai
CIKM1
2011 Are Semantically Related Links More Effective for Retrieval?
Marijn Koolen, Jaap Kamps
ECIR1
2011 Crowdsourcing for book search evaluation: impact of hit design on comparative system ranking
abstract
The evaluation of information retrieval (IR) systems over special collections, such as large book repositories, is out of reach of traditional methods that rely upon editorial relevance judgments. Increasingly, the use of crowdsourcing to collect relevance labels has been regarded as a viable alternative that scales with modest costs. However, crowdsourcing suffers from undesirable worker practices and low quality contributions. In this paper we investigate the design and implementation of effective crowdsourcing tasks in the context of book search evaluation. We observe the impact of aspects of the Human Intelligence Task (HIT) design on the quality of relevance labels provided by the crowd. We assess the output in terms of label agreement with a gold standard data set and observe the effect of the crowdsourced relevance judgments on the resulting system rankings. This enables us to observe the effect of crowdsourcing on the entire IR evaluation process. Using the test set and experimental runs from the INEX 2010 Book Track, we find that varying the HIT design, and the pooling and document ordering strategies leads to considerable differences in agreement with the gold set labels. We then observe the impact of the crowdsourced relevance label sets on the relative system rankings using four IR performance metrics. System rankings based on MAP and Bpref remain less affected by different label sets while the [email protected] and [email protected] lead to dramatically different system rankings, especially for labels acquired from HITs with weaker quality controls. Overall, we find that crowdsourcing can be an effective tool for the evaluation of IR systems, provided that care is taken when designing the HITs.
Gabriella Kazai, Jaap Kamps, Marijn Koolen, Natasa Milic-Frayling
SIGIR3
2010 The importance of anchor text for ad hoc search revisited
abstract
It is generally believed that propagated anchor text is very important for effective Web search as offered by the commercial search engines. "Google Bombs" are a notable illustration of this. However, many years of TREC Web retrieval research failed to establish the effectiveness of link evidence for ad hoc retrieval on Web collections. The ultimate resolution to this dilemma was that typical Web search is very different from the traditional ad hoc methodology. So far, however, no one has established why link information, like incoming link degree or anchor text, does not help ad hoc retrieval effectiveness. Several possible explanations were given, including the collections being too small for anchors to be effective, and the density of the link graph being too low.
Marijn Koolen, Jaap Kamps
SIGIR1
2010 The impact of collection size on relevance and diversity
abstract
It has been observed that precision increases with collection size. One explanation could be that the redundancy of information increases, making it easier to find multiple documents conveying the same information. Arguably, a user has no interest in reading the same information over and over, but would prefer a set of diverse search results covering multiple aspects of the search topic. In this paper, we look at the impact of the collection size on the relevance and diversity of retrieval results by down-sampling the collection.
Marijn Koolen, Jaap Kamps
SIGIR1
2009 Using wikipedia categories for ad hoc search
abstract
In this paper we explore the use of category information for ad hoc retrieval in Wikipedia. We show that techniques for entity ranking exploiting this category information can also be applied to ad hoc topics and lead to significant improvements. Automatically assigned target categories are good surrogates for manually assigned categories, which perform only slightly better.
Rianne Kaptein, Marijn Koolen, Jaap Kamps
SIGIR2
2009 Is Wikipedia link structure different?
abstract
In this paper, we investigate the difference between Wikipedia and Web link structure with respect to their value as indicators of the relevance of a page for a given topic of request. Our experimental evidence is from two IR test-collections: the .GOV collection used at the TREC Web tracks and the Wikipedia XML Corpus used at INEX. We first perform a comparative analysis of Wikipedia and .GOV link structure and then investigate the value of link evidence for improving search on Wikipedia and on the .GOV domain. Our main findings are: First, Wikipedia link structure is similar to the Web, but more densely linked. Second, Wikipedia's outlinks behave similar to inlinks and both are good indicators of relevance, whereas on the Web the inlinks are more important. Third, when incorporating link evidence in the retrieval model, for Wikipedia the global link evidence fails and we have to take the local context into account.
Jaap Kamps, Marijn Koolen
WSDM2
2009 Wikipedia pages as entry points for book search
abstract
A lot of the world's knowledge is stored in books, which, as a result of recent mass-digitisation efforts, are increasingly available online. Search engines, such as Google Books, provide mechanisms for searchers to enter this vast knowledge space using queries as entry points. In this paper, we view Wikipedia as a summary of this world knowledge and aim to use this resource to guide users to relevant books. Thus, we investigate possible ways of using Wikipedia as an intermediary between the user's query and a collection of books being searched. We experiment with traditional query expansion techniques, exploiting Wikipedia articles as rich sources of information that can augment the user's query. We then propose a novel approach based on link distance in an extended Wikipedia graph: we associate books with Wikipedia pages that cite these books and use the link distance between these nodes and the pages that match the user query as an estimation of a book's relevance to the query. Our results show that a) classical query expansion using terms extracted from query pages leads to increased precision, and b) link distance between query and book pages in Wikipedia provides a good indicator of relevance that can boost the retrieval score of relevant books in the result ranking of a book search engine.
Marijn Koolen, Gabriella Kazai, Nick Craswell
WSDM1
2008 The Importance of Link Evidence in Wikipedia
Jaap Kamps, Marijn Koolen
ECIR2
2008 Locating relevant text within XML documents
abstract
Traditional document retrieval has shown to be a competitive approach in XML element retrieval, which is counter-intuitive since the element retrieval task requests all and only relevant document parts to be retrieved. This paper conducts a comparative analysis of document and element retrieval, highlights the relative strengths and weaknesses of both approaches, and explains the relative effectiveness of document retrieval approaches at element retrieval tasks.
Jaap Kamps, Marijn Koolen, Mounia Lalmas-Roelleke
SIGIR2
2007 Where to start reading a textual XML document?
abstract
In structured information retrieval, the aim is to exploit document structure to retrieve relevant components, allowing the user to go straight to the relevant material. This paper looks at the so-called best entry points (BEPs), which are intended to give the user the best starting point to access the relevant information in the document. We examine the relationship between BEPs and relevant components in the INEX 2006 ad hoc assessments. Our main findings are the following: First, although documents are short, assessors often choose the best entry point some distance from the start of the document. Second, many of the best entry points coincide with the first relevant character in relevant documents, showing a strong relation between the BEP and relevant text. Third, we find browsing BEPs in articles with a single relevant passages, and container BEPs or context BEPs in articles with more relevant passages.
Jaap Kamps, Marijn Koolen, Mounia Lalmas-Roelleke
SIGIR2
2006 A Cross-Language Approach to Historic Document Retrieval
Marijn Koolen, Frans Adriaans, Jaap Kamps, Maarten de Rijke
ECIR1