Wouter Weerkamp

dblp:52/5250 · DBLP profile ↗
← Back
34ranked-venue papers
11as first author
0since 2021 · last 2016
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 28 · 8 first-authorArtificial intelligence and machine learning · 13 · 4 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
17 papers
Information retrieval · 83% Knowledge graphs · 15% Data mining · 2%
Artificial intelligence
2 papers
Information extraction and text analysis · 100%

Topics — the 30 heaviest of 41, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval
retrieval models
0.332011
Hypergeometric language models for republished article finding · SIGIR 2011
A Generative Blog Post Retrieval Model that Uses Query Expansion based on External Collections · ACL/IJCNLP 2009
Parsimonious relevance models · SIGIR 2008
Information retrieval › search engines › semantic search › entity retrieval
entity ranking
0.212016
Dynamic Collective Entity Representations for Entity Ranking · WSDM 2016
Knowledge graphs
entity representation
0.212016
Dynamic Collective Entity Representations for Entity Ranking · WSDM 2016
Information retrieval › ranking
learning to rank
0.212016
Dynamic Collective Entity Representations for Entity Ranking · WSDM 2016
Information retrieval › evaluation
online evaluation
0.212016
Building a Self-Learning Search Engine: From Research to Business · SIGIR 2016
Information retrieval
retrieval evaluation
0.212016
Building a Self-Learning Search Engine: From Research to Business · SIGIR 2016
Information retrieval
search engines
0.212016
Building a Self-Learning Search Engine: From Research to Business · SIGIR 2016
Information retrieval › search engines › semantic search › entity retrieval
people search
0.222011
People searching for people: analysis of a people search engine log · SIGIR 2011
Finding people and their utterances in social media · SIGIR 2010
Knowledge graphs › knowledge graph explanation
entity relationship explanation
0.212015
Learning to Explain Entity Relationships in Knowledge Graphs · ACL (1) 2015
Knowledge graphs
knowledge graph explanation
0.212015
Learning to Explain Entity Relationships in Knowledge Graphs · ACL (1) 2015
Information retrieval › retrieval models
language model
0.222011
Hypergeometric language models for republished article finding · SIGIR 2011
Parsimonious relevance models · SIGIR 2008
Information retrieval › query understanding
query modeling
0.222011
Linking online news and social media · WSDM 2011
A few examples go a long way: constructing query models from elaborate query formulations · SIGIR 2008
Information retrieval
evaluation
0.222013
Pseudo test collections for training and tuning microblog rankers · SIGIR 2013
A comparison of retrieval-based hierarchical clustering approaches to person name disambiguation · SIGIR 2009
Information retrieval › query reformulation
query expansion
0.222009
A Generative Blog Post Retrieval Model that Uses Query Expansion based on External Collections · ACL/IJCNLP 2009
A few examples go a long way: constructing query models from elaborate query formulations · SIGIR 2008
Information retrieval › web search › web information retrieval › social media retrieval
microblog retrieval
0.212013
Pseudo test collections for training and tuning microblog rankers · SIGIR 2013
Information retrieval › evaluation
test collection
0.212013
Pseudo test collections for training and tuning microblog rankers · SIGIR 2013
Natural language and speech › Information extraction and text analysis
entity linking
0.112012
Adding semantics to microblog posts · WSDM 2012
Knowledge graphs
knowledge base linking
0.112012
Adding semantics to microblog posts · WSDM 2012
Information retrieval › distributed information retrieval
multi-source retrieval
0.112011
Linking online news and social media · WSDM 2011
Information retrieval › similarity search
near-duplicate detection
0.112011
Hypergeometric language models for republished article finding · SIGIR 2011
Information retrieval
ranking
0.122010
Credibility Improves Topical Blog Post Retrieval · ACL 2008
A two-stage model for blog feed search · SIGIR 2010
Natural language and speech › Information extraction and text analysis
sentiment analysis
0.112010
Generating Focused Topic-Specific Sentiment Lexicons · ACL 2010
Information retrieval › web search
blog feed search
0.112010
A two-stage model for blog feed search · SIGIR 2010
Information retrieval › web search › web information retrieval › social media retrieval
blog retrieval
0.112010
A two-stage model for blog feed search · SIGIR 2010
Information retrieval › web search › web information retrieval
social media retrieval
0.112010
Finding people and their utterances in social media · SIGIR 2010
Information retrieval › retrieval models
generative retrieval
0.112009
A Generative Blog Post Retrieval Model that Uses Query Expansion based on External Collections · ACL/IJCNLP 2009
Data mining › clustering
hierarchical clustering
0.112009
A comparison of retrieval-based hierarchical clustering approaches to person name disambiguation · SIGIR 2009
Information retrieval › text analysis
name disambiguation
0.112009
A comparison of retrieval-based hierarchical clustering approaches to person name disambiguation · SIGIR 2009
Information retrieval › search engines
expert finding
0.112008
Bloggers as experts: feed distillation using expert retrieval models · SIGIR 2008
Information retrieval › retrieval models › language model
parsimonious language model
0.112008
Parsimonious relevance models · SIGIR 2008

Methods — techniques the papers use, named apart from their topics

machine learning with engineered features · 0.3online learning to rank · 0.2machine learning · 0.2dynamic representation learning · 0.2unsupervised query generation · 0.2language modeling · 0.2hashtag-based relevance judgments · 0.2hypergeometric distribution · 0.1graph-based term selection · 0.1data fusion · 0.1
YearPublicationVenuePosition
2016 Building a Self-Learning Search Engine: From Research to Business
abstract
904Labs B.V. was founded in 2014 by Wouter Weerkamp, Manos Tsagkias, and Maarten de Rijke to commercialize state-of-the-art search engine technology. 904Labs' strategic product is a self-learning search engine for online retailers, which uses some of the most recent scientific developments in machine learning and search engine evaluation. 904Labs has raised about 200K in funding and has signed pilots with large international and national companies. Since its start, 904Labs has grown with two developers and two experienced business people. In this presentation we tell how to go from research to business and the challenges it brings along?
Manos Tsagkias, Wouter Weerkamp
SIGIR2
2016 Dynamic Collective Entity Representations for Entity Ranking
abstract
Entity ranking, i.e., successfully positioning a relevant entity at the top of the ranking for a given query, is inherently difficult due to the potential mismatch between the entity's description in a knowledge base, and the way people refer to the entity when searching for it. To counter this issue we propose a method for constructing dynamic collective entity representations. We collect entity descriptions from a variety of sources and combine them into a single entity representation by learning to weight the content from different sources that are associated with an entity for optimal retrieval effectiveness. Our method is able to add new descriptions in real time and learn the best representation as time evolves so as to capture the dynamics of how people search entities. Incorporating dynamic description sources into dynamic collective entity representations improves retrieval effectiveness by 7% over a state-of-the-art learning to rank baseline. Periodic retraining of the ranker enables higher ranking effectiveness for dynamic collective entity representations.
David Graus, Manos Tsagkias, Wouter Weerkamp, Edgar Meij, Maarten de Rijke
WSDM3
2015 Learning to Explain Entity Relationships in Knowledge Graphs
abstract
Nikos Voskarides, Edgar Meij, Manos Tsagkias, Maarten de Rijke, Wouter Weerkamp. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Nikos Voskarides, Edgar Meij, Manos Tsagkias, Maarten de Rijke, Wouter Weerkamp
ACL (1)5
2014 Time-Aware Rank Aggregation for Microblog Search
abstract
We tackle the problem of searching microblog posts and frame it as a rank aggregation problem where we merge result lists generated by separate rankers so as to produce a final ranking to be returned to the user. We propose a rank aggregation method, TimeRA, that is able to infer the rank scores of documents via latent factor modeling. It is time-aware and rewards posts that are published in or near a burst of posts that are ranked highly in many of the lists being aggregated. Our experimental results show that it significantly outperforms state-of-the-art rank aggregation and time-sensitive microblog search algorithms.
Shangsong Liang, Zhaochun Ren, Wouter Weerkamp, Edgar Meij, Maarten de Rijke
CIKM3
2014 Query-Dependent Contextualization of Streaming Data
Nikos Voskarides, Daan Odijk, Manos Tsagkias, Wouter Weerkamp, Maarten de Rijke
ECIR4
2014 Recipient recommendation in enterprises using communication graphs and email content
abstract
We address the task of recipient recommendation for emailing in enterprises. We propose an intuitive and elegant way of modeling the task of recipient recommendation, which uses both the communication graph (i.e., who are most closely connected to the sender) and the content of the email. Additionally, the model can incorporate evidence as prior probabilities. Experiments on two enterprise email collections show that our model achieves very high scores, and that it outperforms two variants that use either the communication graph or the content in isolation.
David Graus, David van Dijk, Manos Tsagkias, Wouter Weerkamp, Maarten de Rijke
SIGIR4
2013 Inside the world's playlist
abstract
We describe Streamwatchr, a real-time system for analyzing the music listening behavior of people around the world. Streamwatchr collects music-related tweets, extracts artists and songs, and visualizes the results in three ways: (i) currently trending songs and artists, (ii) newly discovered songs, and (iii) popularity statistics per country and world-wide for both songs and artists.
Wouter Weerkamp, Manos Tsagkias, Maarten de Rijke
CIKM1
2013 Pseudo test collections for training and tuning microblog rankers
abstract
Recent years have witnessed a persistent interest in generating pseudo test collections, both for training and evaluation purposes. We describe a method for generating queries and relevance judgments for microblog search in an unsupervised way. Our starting point is this intuition: tweets with a hashtag are relevant to the topic covered by the hashtag and hence to a suitable query derived from the hashtag. Our baseline method selects all commonly used hashtags, and all associated tweets as relevance judgments; we then generate a query from these tweets. Next, we generate a timestamp for each query, allowing us to use temporal information in the training process. We then enrich the generation process with knowledge derived from an editorial test collection for microblog search.
Richard Berendsen, Manos Tsagkias, Wouter Weerkamp, Maarten de Rijke
SIGIR3
2012 Result Disambiguation in Web People Search
Richard Berendsen, Bogomil Kovachev, Evangelia-Paraskevi Nastou, Maarten de Rijke, Wouter Weerkamp
ECIR5
2012 A Framework for Unsupervised Spam Detection in Social Networking Sites
Maarten Bosma, Edgar Meij, Wouter Weerkamp
ECIR3
2012 Adaptive Temporal Query Modeling
Maria-Hendrike Peetz, Edgar Meij, Maarten de Rijke, Wouter Weerkamp
ECIR4
2012 Adding semantics to microblog posts
abstract
Microblogs have become an important source of information for the purpose of marketing, intelligence, and reputation management. Streams of microblogs are of great value because of their direct and real-time nature. Determining what an individual microblog post is about, however, can be non-trivial because of creative language usage, the highly contextualized and informal nature of microblog posts, and the limited length of this form of communication. We propose a solution to the problem of determining what a microblog post is about through semantic linking: we add semantics to posts by automatically identifying concepts that are semantically related to it and generating links to the corresponding Wikipedia articles. The identified concepts can subsequently be used for, e.g., social media mining, thereby reducing the need for manual inspection and selection. Using a purpose-built test collection of tweets, we show that recently proposed approaches for semantic linking do not perform well, mainly due to the idiosyncratic nature of microblog posts. We propose a novel method based on machine learning with a set of innovative features and show that it is able to achieve significant improvements over all other methods, especially in terms of precision.
Edgar Meij, Wouter Weerkamp, Maarten de Rijke
WSDM2
2012 Credibility-inspired ranking for blog post retrieval
abstract
Credibility of information refers to its believability or the believability of its sources. We explore the impact of credibility-inspired indicators on the task of blog post retrieval, following the intuition that more credible blog posts are preferred by searchers. Based on a previously introduced credibility framework for blogs, we define several credibility indicators, and divide them into post-level (e.g., spelling, timeliness, document length) and blog-level (e.g., regularity, expertise, comments) indicators. The retrieval task at hand is precision-oriented, and we hypothesize that the use of credibility-inspired indicators will positively impact precision. We propose to use ideas from the credibility framework in a reranking approach to the blog post retrieval problem: We introduce two simple ways of reranking the top n of an initial run. The first approach, Credibility-inspired reranking, simply reranks the top n of a baseline based on the credibility-inspired score. The second approach, Combined reranking, multiplies the credibility-inspired score of the top n results by their retrieval score, and reranks based on this score. Results show that Credibility-inspired reranking leads to larger improvements over the baseline than Combined reranking, but both approaches are capable of improving over an already strong baseline. For Credibility-inspired reranking the best performance is achieved using a combination of all post-level indicators. Combined reranking works best using the post-level indicators combined with comments and pronouns. The blog-level indicators expertise, regularity, and coherence do not contribute positively to the performance, although analysis shows that they can be useful for certain topics. Additional analysis shows that a relative small value of n (15–25) leads to the best results, and that posts that move up the ranking due to the integration of reranking based on credibility-inspired indicators do indeed appear to be more credible than the ones that go down.
Wouter Weerkamp, Maarten de Rijke
Inf. Retr.1
2012 Exploiting External Collections for Query Expansion
abstract
A persisting challenge in the field of information retrieval is the vocabulary mismatch between a user’s information need and the relevant documents. One way of addressing this issue is to apply query modeling: to add terms to the original query and reweigh the terms. In social media, where documents usually contain creative and noisy language (e.g., spelling and grammatical errors), query modeling proves difficult. To address this, attempts to use external sources for query modeling have been made and seem to be successful. In this article we propose a general generative query expansion model that uses external document collections for term generation: the External Expansion Model (EEM). The main rationale behind our model is our hypothesis that each query requires its own mixture of external collections for expansion and that an expansion model should account for this. For some queries we expect, for example, a news collection to be most beneficial, while for other queries we could benefit more by selecting terms from a general encyclopedia. EEM allows for query-dependent weighing of the external collections. We put our model to the test on the task of blog post retrieval and we use four external collections in our experiments: (i) a news collection, (ii) a Web collection, (iii) Wikipedia, and (iv) a blog post collection. Experiments show that EEM outperforms query expansion on the individual collections, as well as the Mixture of Relevance Models that was previously proposed by Diaz and Metzler [2006]. Extensive analysis of the results shows that our naive approach to estimating query-dependent collection importance works reasonably well and that, when we use “oracle” settings, we see the full potential of our model. We also find that the query-dependent collection importance has more impact on retrieval performance than the independent collection importance (i.e., a collection prior).
Wouter Weerkamp, Krisztian Balog, Maarten de Rijke
ACM Trans. Web1
2011 Incorporating Query Expansion and Quality Indicators in Searching Microblog Posts
Kamran Massoudi, Manos Tsagkias, Maarten de Rijke, Wouter Weerkamp
ECIR4
2011 Hypergeometric language models for republished article finding
abstract
Republished article finding is the task of identifying instances of articles that have been published in one source and republished more or less verbatim in another source, which is often a social media source. We address this task as an ad hoc retrieval problem, using the source article as a query. Our approach is based on language modeling. We revisit the assumptions underlying the unigram language model taking into account the fact that in our setup queries are as long as complete news articles. We argue that in this case, the underlying generative assumption of sampling words from a document with replacement, i.e., the multinomial modeling of documents, produces less accurate query likelihood estimates.
Manos Tsagkias, Maarten de Rijke, Wouter Weerkamp
SIGIR3
2011 People searching for people: analysis of a people search engine log
abstract
Recent years show an increasing interest in vertical search: searching within a particular type of information. Understanding what people search for in these "verticals" gives direction to research and provides pointers for the search engines themselves. In this paper we analyze the search logs of one particular vertical: people search engines. Based on an extensive analysis of the logs of a search engine geared towards finding people, we propose a classification scheme for people search at three levels: (a) queries, (b) sessions, and (c) users. For queries, we identify three types, (i) event-based high-profile queries (people that become "popular" because of an event happening), (ii) regular high-profile queries (celebrities), and (iii) low-profile queries (other, less-known people). We present experiments on automatic classification of queries. On the session level, we observe five types: (i) family sessions (users looking for relatives), (ii) event sessions (querying the main players of an event), (iii) spotting sessions (trying to "spot" different celebrities online), (iv) polymerous sessions (sessions without a clear relation between queries), and (v) repetitive sessions (query refinement and copying). Finally, for users we identify four types: (i) monitors, (ii) spotters, (iii) followers, and (iv) polymers.
Wouter Weerkamp, Richard Berendsen, Bogomil Kovachev, Edgar Meij, Krisztian Balog, Maarten de Rijke
SIGIR1
2011 Linking online news and social media
abstract
Much of what is discussed in social media is inspired by events in the news and, vice versa, social media provide us with a handle on the impact of news events. We address the following linking task: given a news article, find social media utterances that implicitly reference it. We follow a three-step approach: we derive multiple query models from a given source news article, which are then used to retrieve utterances from a target social media index, resulting in multiple ranked lists that we then merge using data fusion techniques. Query models are created by exploiting the structure of the source article and by using explicitly linked social media utterances that discuss the source article. To combat query drift resulting from the large volume of text, either in the source news article itself or in social media utterances explicitly linked to it, we introduce a graph-based method for selecting discriminative terms.
Manos Tsagkias, Maarten de Rijke, Wouter Weerkamp
WSDM3
2011 Blog feed search with a post index
abstract
User generated content forms an important domain for mining knowledge. In this paper, we address the task of blog feed search: to find blogs that are principally devoted to a given topic, as opposed to blogs that merely happen to mention the topic in passing. The large number of blogs makes the blogosphere a challenging domain, both in terms of effectiveness and of storage and retrieval efficiency. We examine the effectiveness of an approach to blog feed search that is based on individual posts as indexing units (instead of full blogs). Working in the setting of a probabilistic language modeling approach to information retrieval, we model the blog feed search task by aggregating over a blogger’s posts to collect evidence of relevance to the topic and persistence of interest in the topic. This approach achieves state-of-the-art performance in terms of effectiveness. We then introduce a two-stage model where a pre-selection of candidate blogs is followed by a ranking step. The model integrates aggressive pruning techniques as well as very lean representations of the contents of blog posts, resulting in substantial gains in efficiency while maintaining effectiveness at a very competitive level.
Wouter Weerkamp, Krisztian Balog, Maarten de Rijke
Inf. Retr.1
2010 Generating Focused Topic-Specific Sentiment Lexicons
Valentin Jijkoun, Maarten de Rijke, Wouter Weerkamp
ACL3
2010 News Comments: Exploring, Modeling, and Online Prediction
Manos Tsagkias, Wouter Weerkamp, Maarten de Rijke
ECIR2
2010 Finding people and their utterances in social media
abstract
Since its introduction, social media, "a group of internet-based applications that (...) allow the creation and exchange of user generated content" [1], has attracted more and more users. Over the years, many platforms have arisen that allow users to publish information, communicate with others, connect to like-minded, and share anything a users wants to share. Text-centric examples are mailing lists, forums, blogs, community question answering, collaborative knowledge sources, social networks, and microblogs, with new platforms starting all the time. Given the volume of information available in social media, ways of accessing this information intelligently are needed; this is the scope of my research.
Wouter Weerkamp
SIGIR1
2010 A two-stage model for blog feed search
abstract
We consider blog feed search: identifying relevant blogs for a given topic. An individual's search behavior often involves a combination of exploratory behavior triggered by salient features of the information objects being examined plus goal-directed in-depth information seeking behavior. We present a two-stage blog feed search model that directly builds on this insight. We first rank blog posts for a given topic, and use their parent blogs as selection of blogs that we rank using a blog-based model.
Wouter Weerkamp, Krisztian Balog, Maarten de Rijke
SIGIR1
2009 A Generative Blog Post Retrieval Model that Uses Query Expansion based on External Collections
Wouter Weerkamp, Krisztian Balog, Maarten de Rijke
ACL/IJCNLP1
2009 A query model based on normalized log-likelihood
abstract
Leveraging information from relevance assessments has been proposed as an effective means for improving retrieval. We introduce a novel language modeling method which uses information from each assessed document and their aggregate. While most previous approaches focus either on features of the entire set or on features of the individual relevant documents, our model exploits features of both the documents and the set as a whole. When evaluated, we show that our model is able to significantly improve over state-of-art feedback methods.
Edgar Meij, Wouter Weerkamp, Maarten de Rijke
CIKM2
2009 Predicting the volume of comments on online news stories
abstract
On-line news agents provide commenting facilities for readers to express their views with regard to news stories. The number of user supplied comments on a news article may be indicative of its importance or impact. We report on exploratory work that predicts the comment volume of news articles prior to publication using five feature sets. We address the prediction task as a two stage classification task: a binary classification identifies articles with the potential to receive comments, and a second binary classification receives the output from the first step to label articles "low" or "high" comment volume. The results show solid performance for the former task, while performance degrades for the latter.
Manos Tsagkias, Wouter Weerkamp, Maarten de Rijke
CIKM2
2009 Using Contextual Information to Improve Search in Email Archives
Wouter Weerkamp, Krisztian Balog, Maarten de Rijke
ECIR1
2009 A comparison of retrieval-based hierarchical clustering approaches to person name disambiguation
abstract
This paper describes a simple clustering approach to person name disambiguation of retrieved documents. The methods are based on standard IR concepts and do not require any task-specific features. We compare different term-weighting and indexing methods and evaluate their performance against the Web People Search task (WePS). Despite their simplicity these approaches achieve very competitive performance.
Christof Monz, Wouter Weerkamp
SIGIR2
2009 An effective coherence measure to determine topical consistency in user-generated content
abstract
When searching for blogs on a specific topic, information seekers prefer blogs that place a central focus on that topic over blogs whose mention of the topic is diffuse or incidental. In order to present users with better blog feed search results, we developed a measure of topical consistency that is able to capture whether or not a blog is topically focused. The measure, called the coherence score , is inspired by the genetics literature and captures the tightness of the clustering structure of a data set relative to a background collection. In a set of experiments on synthetic data, the coherence score is shown to provide a faithful reflection of topic clustering structure. The properties that make the coherence score more appropriate than lexical cohesion, a common measure of topical structure, are discussed. Retrieval experiments show that integrating the coherence score as a prior in a language modeling-based approach to blog feed search improves retrieval effectiveness. The coherence score must, however, be used judiciously in order to avoid boosting the ranking of irrelevant but topically focused blogs. To this end, we experiment with a series of weighting schemes that adjust the contribution of the coherence score according to the relevance of a blog to the user query. An appropriate weighting scheme is able to improve retrieval performance. Finally, we show that the coherence score can be reliably estimated with a sample exceeding 20 posts in size. Consistent with this finding, experiments show that the best retrieval performance is achieved if coherence scores are used when a blog contains more than 20 posts.
Jiyin He, Wouter Weerkamp, Martha A. Larson, Maarten de Rijke
Int. J. Document Anal. Recognit.2
2008 Credibility Improves Topical Blog Post Retrieval
Wouter Weerkamp, Maarten de Rijke
ACL1
2008 Finding Key Bloggers, One Post At A Time
abstract
User generated content in general, and blogs in particular, form an interesting and relatively little explored domain for mining knowledge. We address the task of blog distillation: to find blogs that are principally devoted to a given topic, as opposed to blogs that merely happen to discuss the topic in passing. Working in the setting of statistical language modeling, we model the task by aggregating a blogger's blog posts to collect evidence of relevance to the topic and persistence of interest in the topic. This approach achieves state-of-the-art performance. On top of this baseline, we extend our model by incorporating a number of blog-specific features, concerning document structure, social structure, and temporal structure. These blog-specific features yield further improvements.
Wouter Weerkamp, Krisztian Balog, Maarten de Rijke
ECAI1
2008 Bloggers as experts: feed distillation using expert retrieval models
abstract
We address the task of (blog) feed distillation: to find blogs that are principally devoted to a given topic. The task may be viewed as an association finding task, between topics and bloggers. Under this view, it resembles the expert finding task, for which a range of models have been proposed. We adopt two language modeling-based approaches to expert finding, and determine their effectiveness as feed distillation strategies. The two models capture the idea that a human will often search for key blogs by spotting highly relevant posts (the Posting model) or by taking global aspects of the blog into account (the Blogger model). Results show the Blogger model outperforms the Posting model and delivers state-of-the art performance, out-of-the-box.
Krisztian Balog, Maarten de Rijke, Wouter Weerkamp
SIGIR3
2008 A few examples go a long way: constructing query models from elaborate query formulations
abstract
We address a specific enterprise document search scenario, where the information need is expressed in an elaborate manner. In our scenario, information needs are expressed using a short query (of a few keywords) together with examples of key reference pages. Given this setup, we investigate how the examples can be utilized to improve the end-to-end performance on the document retrieval task. Our approach is based on a language modeling framework, where the query model is modified to resemble the example pages. We compare several methods for sampling expansion terms from the example pages to support query-dependent and query-independent query expansion; the latter is motivated by the wish to increase "aspect recall", and attempts to uncover aspects of the information need not captured by the query.
Krisztian Balog, Wouter Weerkamp, Maarten de Rijke
SIGIR2
2008 Parsimonious relevance models
abstract
We describe a method for applying parsimonious language models to re-estimate the term probabilities assigned by relevance models. We apply our method to six topic sets from test collections in five different genres. Our parsimonious relevance models (i) improve retrieval effectiveness in terms of MAP on all collections, (ii) significantly outperform their non-parsimonious counterparts on most measures, and (iii) have a precision enhancing effect, unlike other blind relevance feedback methods.
Edgar Meij, Wouter Weerkamp, Krisztian Balog, Maarten de Rijke
SIGIR2