VLDB 2026 Research / reviewers in the wild / expert
Wouter Weerkamp
dblp:52/5250
· DBLP profile ↗
34ranked-venue papers
11as first author
0since 2021 · last 2016
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 28 · 8 first-authorArtificial intelligence and machine learning · 13 · 4 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
17 papers |
Information retrieval · 83% Knowledge graphs · 15% Data mining · 2% | |
| Artificial intelligence
2 papers |
Information extraction and text analysis · 100% |
Topics — the 30 heaviest of 41, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval
retrieval models |
0.3 | 3 | 2011 | Hypergeometric language models for republished article finding · SIGIR 2011 A Generative Blog Post Retrieval Model that Uses Query Expansion based on External Collections · ACL/IJCNLP 2009 Parsimonious relevance models · SIGIR 2008 |
Information retrieval › search engines › semantic search › entity retrieval
entity ranking |
0.2 | 1 | 2016 | Dynamic Collective Entity Representations for Entity Ranking · WSDM 2016 |
Knowledge graphs
entity representation |
0.2 | 1 | 2016 | Dynamic Collective Entity Representations for Entity Ranking · WSDM 2016 |
Information retrieval › ranking
learning to rank |
0.2 | 1 | 2016 | Dynamic Collective Entity Representations for Entity Ranking · WSDM 2016 |
Information retrieval › evaluation
online evaluation |
0.2 | 1 | 2016 | Building a Self-Learning Search Engine: From Research to Business · SIGIR 2016 |
Information retrieval
retrieval evaluation |
0.2 | 1 | 2016 | Building a Self-Learning Search Engine: From Research to Business · SIGIR 2016 |
Information retrieval
search engines |
0.2 | 1 | 2016 | Building a Self-Learning Search Engine: From Research to Business · SIGIR 2016 |
Information retrieval › search engines › semantic search › entity retrieval
people search |
0.2 | 2 | 2011 | People searching for people: analysis of a people search engine log · SIGIR 2011 Finding people and their utterances in social media · SIGIR 2010 |
Knowledge graphs › knowledge graph explanation
entity relationship explanation |
0.2 | 1 | 2015 | Learning to Explain Entity Relationships in Knowledge Graphs · ACL (1) 2015 |
Knowledge graphs
knowledge graph explanation |
0.2 | 1 | 2015 | Learning to Explain Entity Relationships in Knowledge Graphs · ACL (1) 2015 |
Information retrieval › retrieval models
language model |
0.2 | 2 | 2011 | Hypergeometric language models for republished article finding · SIGIR 2011 Parsimonious relevance models · SIGIR 2008 |
Information retrieval › query understanding
query modeling |
0.2 | 2 | 2011 | Linking online news and social media · WSDM 2011 A few examples go a long way: constructing query models from elaborate query formulations · SIGIR 2008 |
Information retrieval
evaluation |
0.2 | 2 | 2013 | Pseudo test collections for training and tuning microblog rankers · SIGIR 2013 A comparison of retrieval-based hierarchical clustering approaches to person name disambiguation · SIGIR 2009 |
Information retrieval › query reformulation
query expansion |
0.2 | 2 | 2009 | A Generative Blog Post Retrieval Model that Uses Query Expansion based on External Collections · ACL/IJCNLP 2009 A few examples go a long way: constructing query models from elaborate query formulations · SIGIR 2008 |
Information retrieval › web search › web information retrieval › social media retrieval
microblog retrieval |
0.2 | 1 | 2013 | Pseudo test collections for training and tuning microblog rankers · SIGIR 2013 |
Information retrieval › evaluation
test collection |
0.2 | 1 | 2013 | Pseudo test collections for training and tuning microblog rankers · SIGIR 2013 |
Natural language and speech › Information extraction and text analysis
entity linking |
0.1 | 1 | 2012 | Adding semantics to microblog posts · WSDM 2012 |
Knowledge graphs
knowledge base linking |
0.1 | 1 | 2012 | Adding semantics to microblog posts · WSDM 2012 |
Information retrieval › distributed information retrieval
multi-source retrieval |
0.1 | 1 | 2011 | Linking online news and social media · WSDM 2011 |
Information retrieval › similarity search
near-duplicate detection |
0.1 | 1 | 2011 | Hypergeometric language models for republished article finding · SIGIR 2011 |
Information retrieval
ranking |
0.1 | 2 | 2010 | Credibility Improves Topical Blog Post Retrieval · ACL 2008 A two-stage model for blog feed search · SIGIR 2010 |
Natural language and speech › Information extraction and text analysis
sentiment analysis |
0.1 | 1 | 2010 | Generating Focused Topic-Specific Sentiment Lexicons · ACL 2010 |
Information retrieval › web search
blog feed search |
0.1 | 1 | 2010 | A two-stage model for blog feed search · SIGIR 2010 |
Information retrieval › web search › web information retrieval › social media retrieval
blog retrieval |
0.1 | 1 | 2010 | A two-stage model for blog feed search · SIGIR 2010 |
Information retrieval › web search › web information retrieval
social media retrieval |
0.1 | 1 | 2010 | Finding people and their utterances in social media · SIGIR 2010 |
Information retrieval › retrieval models
generative retrieval |
0.1 | 1 | 2009 | A Generative Blog Post Retrieval Model that Uses Query Expansion based on External Collections · ACL/IJCNLP 2009 |
Data mining › clustering
hierarchical clustering |
0.1 | 1 | 2009 | A comparison of retrieval-based hierarchical clustering approaches to person name disambiguation · SIGIR 2009 |
Information retrieval › text analysis
name disambiguation |
0.1 | 1 | 2009 | A comparison of retrieval-based hierarchical clustering approaches to person name disambiguation · SIGIR 2009 |
Information retrieval › search engines
expert finding |
0.1 | 1 | 2008 | Bloggers as experts: feed distillation using expert retrieval models · SIGIR 2008 |
Information retrieval › retrieval models › language model
parsimonious language model |
0.1 | 1 | 2008 | Parsimonious relevance models · SIGIR 2008 |
Methods — techniques the papers use, named apart from their topics
machine learning with engineered features · 0.3online learning to rank · 0.2machine learning · 0.2dynamic representation learning · 0.2unsupervised query generation · 0.2language modeling · 0.2hashtag-based relevance judgments · 0.2hypergeometric distribution · 0.1graph-based term selection · 0.1data fusion · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2016 | Building a Self-Learning Search Engine: From Research to Businessabstract904Labs B.V. was founded in 2014 by Wouter Weerkamp, Manos Tsagkias, and Maarten de Rijke to commercialize state-of-the-art search engine technology. 904Labs' strategic product is a self-learning search engine for online retailers, which uses some of the most recent scientific developments in machine learning and search engine evaluation. 904Labs has raised about 200K in funding and has signed pilots with large international and national companies. Since its start, 904Labs has grown with two developers and two experienced business people. In this presentation we tell how to go from research to business and the challenges it brings along? Manos Tsagkias, Wouter Weerkamp |
SIGIR | 2 |
| 2016 | Dynamic Collective Entity Representations for Entity RankingabstractEntity ranking, i.e., successfully positioning a relevant entity at the top of the ranking for a given query, is inherently difficult due to the potential mismatch between the entity's description in a knowledge base, and the way people refer to the entity when searching for it. To counter this issue we propose a method for constructing dynamic collective entity representations. We collect entity descriptions from a variety of sources and combine them into a single entity representation by learning to weight the content from different sources that are associated with an entity for optimal retrieval effectiveness. Our method is able to add new descriptions in real time and learn the best representation as time evolves so as to capture the dynamics of how people search entities. Incorporating dynamic description sources into dynamic collective entity representations improves retrieval effectiveness by 7% over a state-of-the-art learning to rank baseline. Periodic retraining of the ranker enables higher ranking effectiveness for dynamic collective entity representations. David Graus, Manos Tsagkias, Wouter Weerkamp, Edgar Meij, Maarten de Rijke |
WSDM | 3 |
| 2015 | Learning to Explain Entity Relationships in Knowledge GraphsabstractNikos Voskarides, Edgar Meij, Manos Tsagkias, Maarten de Rijke, Wouter Weerkamp. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Nikos Voskarides, Edgar Meij, Manos Tsagkias, Maarten de Rijke, Wouter Weerkamp |
ACL (1) | 5 |
| 2014 | Time-Aware Rank Aggregation for Microblog SearchabstractWe tackle the problem of searching microblog posts and frame it as a rank aggregation problem where we merge result lists generated by separate rankers so as to produce a final ranking to be returned to the user. We propose a rank aggregation method, TimeRA, that is able to infer the rank scores of documents via latent factor modeling. It is time-aware and rewards posts that are published in or near a burst of posts that are ranked highly in many of the lists being aggregated. Our experimental results show that it significantly outperforms state-of-the-art rank aggregation and time-sensitive microblog search algorithms. Shangsong Liang, Zhaochun Ren, Wouter Weerkamp, Edgar Meij, Maarten de Rijke |
CIKM | 3 |
| 2014 | Query-Dependent Contextualization of Streaming Data
Nikos Voskarides, Daan Odijk, Manos Tsagkias, Wouter Weerkamp, Maarten de Rijke |
ECIR | 4 |
| 2014 | Recipient recommendation in enterprises using communication graphs and email contentabstractWe address the task of recipient recommendation for emailing in enterprises. We propose an intuitive and elegant way of modeling the task of recipient recommendation, which uses both the communication graph (i.e., who are most closely connected to the sender) and the content of the email. Additionally, the model can incorporate evidence as prior probabilities. Experiments on two enterprise email collections show that our model achieves very high scores, and that it outperforms two variants that use either the communication graph or the content in isolation. David Graus, David van Dijk, Manos Tsagkias, Wouter Weerkamp, Maarten de Rijke |
SIGIR | 4 |
| 2013 | Inside the world's playlistabstractWe describe Streamwatchr, a real-time system for analyzing the music listening behavior of people around the world. Streamwatchr collects music-related tweets, extracts artists and songs, and visualizes the results in three ways: (i) currently trending songs and artists, (ii) newly discovered songs, and (iii) popularity statistics per country and world-wide for both songs and artists. Wouter Weerkamp, Manos Tsagkias, Maarten de Rijke |
CIKM | 1 |
| 2013 | Pseudo test collections for training and tuning microblog rankersabstractRecent years have witnessed a persistent interest in generating pseudo test collections, both for training and evaluation purposes. We describe a method for generating queries and relevance judgments for microblog search in an unsupervised way. Our starting point is this intuition: tweets with a hashtag are relevant to the topic covered by the hashtag and hence to a suitable query derived from the hashtag. Our baseline method selects all commonly used hashtags, and all associated tweets as relevance judgments; we then generate a query from these tweets. Next, we generate a timestamp for each query, allowing us to use temporal information in the training process. We then enrich the generation process with knowledge derived from an editorial test collection for microblog search. Richard Berendsen, Manos Tsagkias, Wouter Weerkamp, Maarten de Rijke |
SIGIR | 3 |
| 2012 | Result Disambiguation in Web People Search
Richard Berendsen, Bogomil Kovachev, Evangelia-Paraskevi Nastou, Maarten de Rijke, Wouter Weerkamp |
ECIR | 5 |
| 2012 | A Framework for Unsupervised Spam Detection in Social Networking Sites
Maarten Bosma, Edgar Meij, Wouter Weerkamp |
ECIR | 3 |
| 2012 | Adaptive Temporal Query Modeling
Maria-Hendrike Peetz, Edgar Meij, Maarten de Rijke, Wouter Weerkamp |
ECIR | 4 |
| 2012 | Adding semantics to microblog postsabstractMicroblogs have become an important source of information for the purpose of marketing, intelligence, and reputation management. Streams of microblogs are of great value because of their direct and real-time nature. Determining what an individual microblog post is about, however, can be non-trivial because of creative language usage, the highly contextualized and informal nature of microblog posts, and the limited length of this form of communication. We propose a solution to the problem of determining what a microblog post is about through semantic linking: we add semantics to posts by automatically identifying concepts that are semantically related to it and generating links to the corresponding Wikipedia articles. The identified concepts can subsequently be used for, e.g., social media mining, thereby reducing the need for manual inspection and selection. Using a purpose-built test collection of tweets, we show that recently proposed approaches for semantic linking do not perform well, mainly due to the idiosyncratic nature of microblog posts. We propose a novel method based on machine learning with a set of innovative features and show that it is able to achieve significant improvements over all other methods, especially in terms of precision. Edgar Meij, Wouter Weerkamp, Maarten de Rijke |
WSDM | 2 |
| 2012 | Credibility-inspired ranking for blog post retrievalabstractCredibility of information refers to its believability or the believability of its sources. We explore the impact of credibility-inspired indicators on the task of blog post retrieval, following the intuition that more credible blog posts are preferred by searchers. Based on a previously introduced credibility framework for blogs, we define several credibility indicators, and divide them into post-level (e.g., spelling, timeliness, document length) and blog-level (e.g., regularity, expertise, comments) indicators. The retrieval task at hand is precision-oriented, and we hypothesize that the use of credibility-inspired indicators will positively impact precision. We propose to use ideas from the credibility framework in a reranking approach to the blog post retrieval problem: We introduce two simple ways of reranking the top n of an initial run. The first approach, Credibility-inspired reranking, simply reranks the top n of a baseline based on the credibility-inspired score. The second approach, Combined reranking, multiplies the credibility-inspired score of the top n results by their retrieval score, and reranks based on this score. Results show that Credibility-inspired reranking leads to larger improvements over the baseline than Combined reranking, but both approaches are capable of improving over an already strong baseline. For Credibility-inspired reranking the best performance is achieved using a combination of all post-level indicators. Combined reranking works best using the post-level indicators combined with comments and pronouns. The blog-level indicators expertise, regularity, and coherence do not contribute positively to the performance, although analysis shows that they can be useful for certain topics. Additional analysis shows that a relative small value of n (15–25) leads to the best results, and that posts that move up the ranking due to the integration of reranking based on credibility-inspired indicators do indeed appear to be more credible than the ones that go down. Wouter Weerkamp, Maarten de Rijke |
Inf. Retr. | 1 |
| 2012 | Exploiting External Collections for Query ExpansionabstractA persisting challenge in the field of information retrieval is the vocabulary mismatch between a user’s information need and the relevant documents. One way of addressing this issue is to apply query modeling: to add terms to the original query and reweigh the terms. In social media, where documents usually contain creative and noisy language (e.g., spelling and grammatical errors), query modeling proves difficult. To address this, attempts to use external sources for query modeling have been made and seem to be successful. In this article we propose a general generative query expansion model that uses external document collections for term generation: the External Expansion Model (EEM). The main rationale behind our model is our hypothesis that each query requires its own mixture of external collections for expansion and that an expansion model should account for this. For some queries we expect, for example, a news collection to be most beneficial, while for other queries we could benefit more by selecting terms from a general encyclopedia. EEM allows for query-dependent weighing of the external collections. We put our model to the test on the task of blog post retrieval and we use four external collections in our experiments: (i) a news collection, (ii) a Web collection, (iii) Wikipedia, and (iv) a blog post collection. Experiments show that EEM outperforms query expansion on the individual collections, as well as the Mixture of Relevance Models that was previously proposed by Diaz and Metzler [2006]. Extensive analysis of the results shows that our naive approach to estimating query-dependent collection importance works reasonably well and that, when we use “oracle” settings, we see the full potential of our model. We also find that the query-dependent collection importance has more impact on retrieval performance than the independent collection importance (i.e., a collection prior). Wouter Weerkamp, Krisztian Balog, Maarten de Rijke |
ACM Trans. Web | 1 |
| 2011 | Incorporating Query Expansion and Quality Indicators in Searching Microblog Posts
Kamran Massoudi, Manos Tsagkias, Maarten de Rijke, Wouter Weerkamp |
ECIR | 4 |
| 2011 | Hypergeometric language models for republished article findingabstractRepublished article finding is the task of identifying instances of articles that have been published in one source and republished more or less verbatim in another source, which is often a social media source. We address this task as an ad hoc retrieval problem, using the source article as a query. Our approach is based on language modeling. We revisit the assumptions underlying the unigram language model taking into account the fact that in our setup queries are as long as complete news articles. We argue that in this case, the underlying generative assumption of sampling words from a document with replacement, i.e., the multinomial modeling of documents, produces less accurate query likelihood estimates. Manos Tsagkias, Maarten de Rijke, Wouter Weerkamp |
SIGIR | 3 |
| 2011 | People searching for people: analysis of a people search engine logabstractRecent years show an increasing interest in vertical search: searching within a particular type of information. Understanding what people search for in these "verticals" gives direction to research and provides pointers for the search engines themselves. In this paper we analyze the search logs of one particular vertical: people search engines. Based on an extensive analysis of the logs of a search engine geared towards finding people, we propose a classification scheme for people search at three levels: (a) queries, (b) sessions, and (c) users. For queries, we identify three types, (i) event-based high-profile queries (people that become "popular" because of an event happening), (ii) regular high-profile queries (celebrities), and (iii) low-profile queries (other, less-known people). We present experiments on automatic classification of queries. On the session level, we observe five types: (i) family sessions (users looking for relatives), (ii) event sessions (querying the main players of an event), (iii) spotting sessions (trying to "spot" different celebrities online), (iv) polymerous sessions (sessions without a clear relation between queries), and (v) repetitive sessions (query refinement and copying). Finally, for users we identify four types: (i) monitors, (ii) spotters, (iii) followers, and (iv) polymers. Wouter Weerkamp, Richard Berendsen, Bogomil Kovachev, Edgar Meij, Krisztian Balog, Maarten de Rijke |
SIGIR | 1 |
| 2011 | Linking online news and social mediaabstractMuch of what is discussed in social media is inspired by events in the news and, vice versa, social media provide us with a handle on the impact of news events. We address the following linking task: given a news article, find social media utterances that implicitly reference it. We follow a three-step approach: we derive multiple query models from a given source news article, which are then used to retrieve utterances from a target social media index, resulting in multiple ranked lists that we then merge using data fusion techniques. Query models are created by exploiting the structure of the source article and by using explicitly linked social media utterances that discuss the source article. To combat query drift resulting from the large volume of text, either in the source news article itself or in social media utterances explicitly linked to it, we introduce a graph-based method for selecting discriminative terms. Manos Tsagkias, Maarten de Rijke, Wouter Weerkamp |
WSDM | 3 |
| 2011 | Blog feed search with a post indexabstractUser generated content forms an important domain for mining knowledge. In this paper, we address the task of blog feed search: to find blogs that are principally devoted to a given topic, as opposed to blogs that merely happen to mention the topic in passing. The large number of blogs makes the blogosphere a challenging domain, both in terms of effectiveness and of storage and retrieval efficiency. We examine the effectiveness of an approach to blog feed search that is based on individual posts as indexing units (instead of full blogs). Working in the setting of a probabilistic language modeling approach to information retrieval, we model the blog feed search task by aggregating over a blogger’s posts to collect evidence of relevance to the topic and persistence of interest in the topic. This approach achieves state-of-the-art performance in terms of effectiveness. We then introduce a two-stage model where a pre-selection of candidate blogs is followed by a ranking step. The model integrates aggressive pruning techniques as well as very lean representations of the contents of blog posts, resulting in substantial gains in efficiency while maintaining effectiveness at a very competitive level. Wouter Weerkamp, Krisztian Balog, Maarten de Rijke |
Inf. Retr. | 1 |
| 2010 | Generating Focused Topic-Specific Sentiment Lexicons
Valentin Jijkoun, Maarten de Rijke, Wouter Weerkamp |
ACL | 3 |
| 2010 | News Comments: Exploring, Modeling, and Online Prediction
Manos Tsagkias, Wouter Weerkamp, Maarten de Rijke |
ECIR | 2 |
| 2010 | Finding people and their utterances in social mediaabstractSince its introduction, social media, "a group of internet-based applications that (...) allow the creation and exchange of user generated content" [1], has attracted more and more users. Over the years, many platforms have arisen that allow users to publish information, communicate with others, connect to like-minded, and share anything a users wants to share. Text-centric examples are mailing lists, forums, blogs, community question answering, collaborative knowledge sources, social networks, and microblogs, with new platforms starting all the time. Given the volume of information available in social media, ways of accessing this information intelligently are needed; this is the scope of my research. Wouter Weerkamp |
SIGIR | 1 |
| 2010 | A two-stage model for blog feed searchabstractWe consider blog feed search: identifying relevant blogs for a given topic. An individual's search behavior often involves a combination of exploratory behavior triggered by salient features of the information objects being examined plus goal-directed in-depth information seeking behavior. We present a two-stage blog feed search model that directly builds on this insight. We first rank blog posts for a given topic, and use their parent blogs as selection of blogs that we rank using a blog-based model. Wouter Weerkamp, Krisztian Balog, Maarten de Rijke |
SIGIR | 1 |
| 2009 | A Generative Blog Post Retrieval Model that Uses Query Expansion based on External Collections
Wouter Weerkamp, Krisztian Balog, Maarten de Rijke |
ACL/IJCNLP | 1 |
| 2009 | A query model based on normalized log-likelihoodabstractLeveraging information from relevance assessments has been proposed as an effective means for improving retrieval. We introduce a novel language modeling method which uses information from each assessed document and their aggregate. While most previous approaches focus either on features of the entire set or on features of the individual relevant documents, our model exploits features of both the documents and the set as a whole. When evaluated, we show that our model is able to significantly improve over state-of-art feedback methods. Edgar Meij, Wouter Weerkamp, Maarten de Rijke |
CIKM | 2 |
| 2009 | Predicting the volume of comments on online news storiesabstractOn-line news agents provide commenting facilities for readers to express their views with regard to news stories. The number of user supplied comments on a news article may be indicative of its importance or impact. We report on exploratory work that predicts the comment volume of news articles prior to publication using five feature sets. We address the prediction task as a two stage classification task: a binary classification identifies articles with the potential to receive comments, and a second binary classification receives the output from the first step to label articles "low" or "high" comment volume. The results show solid performance for the former task, while performance degrades for the latter. Manos Tsagkias, Wouter Weerkamp, Maarten de Rijke |
CIKM | 2 |
| 2009 | Using Contextual Information to Improve Search in Email Archives
Wouter Weerkamp, Krisztian Balog, Maarten de Rijke |
ECIR | 1 |
| 2009 | A comparison of retrieval-based hierarchical clustering approaches to person name disambiguationabstractThis paper describes a simple clustering approach to person name disambiguation of retrieved documents. The methods are based on standard IR concepts and do not require any task-specific features. We compare different term-weighting and indexing methods and evaluate their performance against the Web People Search task (WePS). Despite their simplicity these approaches achieve very competitive performance. Christof Monz, Wouter Weerkamp |
SIGIR | 2 |
| 2009 | An effective coherence measure to determine topical consistency in user-generated contentabstractWhen searching for blogs on a specific topic, information seekers prefer blogs that place a central focus on that topic over blogs whose mention of the topic is diffuse or incidental. In order to present users with better blog feed search results, we developed a measure of topical consistency that is able to capture whether or not a blog is topically focused. The measure, called the coherence score , is inspired by the genetics literature and captures the tightness of the clustering structure of a data set relative to a background collection. In a set of experiments on synthetic data, the coherence score is shown to provide a faithful reflection of topic clustering structure. The properties that make the coherence score more appropriate than lexical cohesion, a common measure of topical structure, are discussed. Retrieval experiments show that integrating the coherence score as a prior in a language modeling-based approach to blog feed search improves retrieval effectiveness. The coherence score must, however, be used judiciously in order to avoid boosting the ranking of irrelevant but topically focused blogs. To this end, we experiment with a series of weighting schemes that adjust the contribution of the coherence score according to the relevance of a blog to the user query. An appropriate weighting scheme is able to improve retrieval performance. Finally, we show that the coherence score can be reliably estimated with a sample exceeding 20 posts in size. Consistent with this finding, experiments show that the best retrieval performance is achieved if coherence scores are used when a blog contains more than 20 posts. Jiyin He, Wouter Weerkamp, Martha A. Larson, Maarten de Rijke |
Int. J. Document Anal. Recognit. | 2 |
| 2008 | Credibility Improves Topical Blog Post Retrieval
Wouter Weerkamp, Maarten de Rijke |
ACL | 1 |
| 2008 | Finding Key Bloggers, One Post At A TimeabstractUser generated content in general, and blogs in particular, form an interesting and relatively little explored domain for mining knowledge. We address the task of blog distillation: to find blogs that are principally devoted to a given topic, as opposed to blogs that merely happen to discuss the topic in passing. Working in the setting of statistical language modeling, we model the task by aggregating a blogger's blog posts to collect evidence of relevance to the topic and persistence of interest in the topic. This approach achieves state-of-the-art performance. On top of this baseline, we extend our model by incorporating a number of blog-specific features, concerning document structure, social structure, and temporal structure. These blog-specific features yield further improvements. Wouter Weerkamp, Krisztian Balog, Maarten de Rijke |
ECAI | 1 |
| 2008 | Bloggers as experts: feed distillation using expert retrieval modelsabstractWe address the task of (blog) feed distillation: to find blogs that are principally devoted to a given topic. The task may be viewed as an association finding task, between topics and bloggers. Under this view, it resembles the expert finding task, for which a range of models have been proposed. We adopt two language modeling-based approaches to expert finding, and determine their effectiveness as feed distillation strategies. The two models capture the idea that a human will often search for key blogs by spotting highly relevant posts (the Posting model) or by taking global aspects of the blog into account (the Blogger model). Results show the Blogger model outperforms the Posting model and delivers state-of-the art performance, out-of-the-box. Krisztian Balog, Maarten de Rijke, Wouter Weerkamp |
SIGIR | 3 |
| 2008 | A few examples go a long way: constructing query models from elaborate query formulationsabstractWe address a specific enterprise document search scenario, where the information need is expressed in an elaborate manner. In our scenario, information needs are expressed using a short query (of a few keywords) together with examples of key reference pages. Given this setup, we investigate how the examples can be utilized to improve the end-to-end performance on the document retrieval task. Our approach is based on a language modeling framework, where the query model is modified to resemble the example pages. We compare several methods for sampling expansion terms from the example pages to support query-dependent and query-independent query expansion; the latter is motivated by the wish to increase "aspect recall", and attempts to uncover aspects of the information need not captured by the query. Krisztian Balog, Wouter Weerkamp, Maarten de Rijke |
SIGIR | 2 |
| 2008 | Parsimonious relevance modelsabstractWe describe a method for applying parsimonious language models to re-estimate the term probabilities assigned by relevance models. We apply our method to six topic sets from test collections in five different genres. Our parsimonious relevance models (i) improve retrieval effectiveness in terms of MAP on all collections, (ii) significantly outperform their non-parsimonious counterparts on most measures, and (iii) have a precision enhancing effect, unlike other blind relevance feedback methods. Edgar Meij, Wouter Weerkamp, Krisztian Balog, Maarten de Rijke |
SIGIR | 2 |