VLDB 2026 Research / reviewers in the wild / expert
Michael Völske
dblp:59/10289
· DBLP profile ↗
13ranked-venue papers in the field
2as first author
4since 2021 · last 2022
0000-0002-9283-6846ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 11 (2 first)Data Mining & Knowledge Discovery · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Axiomatic Retrieval Experimentation with ir_axiomsabstractAxiomatic approaches to information retrieval have played a key role in determining basic constraints that characterize good retrieval models. Beyond their importance in retrieval theory, axioms have been operationalized to improve an initial ranking, to "guide" retrieval, or to explain some model's rankings. However, recent open-source retrieval frameworks like PyTerrier and Pyserini, which made it easy to experiment with sparse and dense retrieval models, have not included any retrieval axiom support so far. Alexander Bondarenko 0001, Maik Fröbe, Jan Heinrich Merker, Benno Stein 0001, Michael Völske, Matthias Hagen |
SIGIR | 5 |
| 2021 | CopyCat: Near-Duplicates Within and Between the ClueWeb and the Common CrawlabstractThe amount of near-duplicates in web crawls like the ClueWeb or Common Crawl demands from their users either to develop a preprocessing pipeline for deduplication, which is costly both computationally and in person hours, or accepting the undesired effects that near-duplicates have on reliability and validity of experiments. We introduce ChatNoir-CopyCat-21, which simplifies deduplication significantly. It comes in two parts: (1) A compilation of near-duplicate documents within the ClueWeb09, the ClueWeb12, and two Common Crawl snapshots, as well as between selections of these crawls, and (2) a software library that implements the deduplication of arbitrary document sets. Our analysis shows that 14--52, of the documents within a crawl and around~0.7--2.5, between the crawls are near-duplicates. Two showcases demonstrate the application and usefulness of our resource. Maik Fröbe, Janek Bevendorff, Lukas Gienapp, Michael Völske, Benno Stein 0001, Martin Potthast, Matthias Hagen |
SIGIR | 4 |
| 2021 | The Information Retrieval AnthologyabstractWe present the IR Anthology, a corpus of information retrieval publications accessible via a metadata browser and a full-text search engine. Following the example of the well-known ACL Anthology, the IR Anthology serves as a hub for researchers interested in information retrieval. Our search engine ChatNoir indexes the publications' full texts, enabling a focused search and linking users to the respective publisher's site for personal access. Listing more than 40,000 publications at the time of writing, the IR Anthology can be freely accessed at https://IR.webis.de. Martin Potthast, Sebastian Günther 0002, Janek Bevendorff, Jan Philipp Bittner, Alexander Bondarenko 0001, Maik Fröbe, Christian Kahmann, Andreas Niekler, Michael Völske, Benno Stein 0001, Matthias Hagen |
SIGIR | 9 |
| 2021 | Predicting essay quality from search and writing behaviorabstractAbstract Few studies have investigated how search behavior affects complex writing tasks. We analyze a dataset of 150 long essays whose authors searched the ClueWeb09 corpus for source material, while all querying, clicking, and writing activity was meticulously recorded. We model the effect of search and writing behavior on essay quality using path analysis. Since the boil‐down and build‐up writing strategies identified in previous research have been found to affect search behavior, we model each writing strategy separately. Our analysis shows that the search process contributes significantly to essay quality through both direct and mediated effects, while the author's writing strategy moderates this relationship. Our models explain 25–35% of the variation in essay quality through rather simple search and writing process characteristics alone, a fact that has implications on how search engines could personalize result pages for writing tasks. Authors' writing strategies and associated searching patterns differ, producing differences in essay quality. In a nutshell: essay quality improves if search and writing strategies harmonize—build‐up writers benefit from focused, in‐depth querying, while boil‐down writers fare better with a broader and shallower querying strategy. Pertti Vakkari, Michael Völske, Martin Potthast, Matthias Hagen, Benno Stein 0001 |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2020 | Comparative Web Search Questionsabstract\beginabstract We analyze comparative questions, i.e., questions asking to compare different items, that were submitted to Yandex in 2012. Responses to such questions might be quite different from the simple "ten blue links'' and could, for example, aggregate pros and cons of the different options as direct answers. However, changing the result presentation is an intricate decision such that the classification of comparative questions forms a highly precision-oriented task. Alexander Bondarenko 0001, Pavel Braslavski 0001, Michael Völske, Rami Aly, Maik Fröbe, Alexander Panchenko, Chris Biemann, Benno Stein 0001, Matthias Hagen |
WSDM | 3 |
| 2019 | Wikipedia Text Reuse: Within and Without
Milad Alshomary, Michael Völske, Tristan Licht, Henning Wachsmuth, Benno Stein 0001, Matthias Hagen, Martin Potthast |
ECIR (1) | 2 |
| 2019 | Query-Task MappingabstractSeveral recent task-based search studies aim at splitting query logs into sets of queries for the same task or information need. We address the natural next step: mapping a currently submitted query to an appropriate task in an already task-split log. This query-task mapping can, for instance, enhance query suggestions---rendering efficiency of the mapping, besides accuracy, a key objective. Our main contributions are three large benchmark datasets and preliminary experiments with four query-task mapping approaches: (1) a Trie-based approach, (2) MinHash~LSH, (3) word movers distance in a Word2Vec setup, and (4) an inverted index-based approach. The experiments show that the fast and accurate inverted index-based method forms a strong baseline. Michael Völske, Ehsan Fatehifar, Benno Stein 0001, Matthias Hagen |
SIGIR | 1 |
| 2019 | Modeling the usefulness of search results as measured by information use
Pertti Vakkari, Michael Völske, Martin Potthast, Matthias Hagen, Benno Stein 0001 |
Inf. Process. Manag. | 2 |
| 2018 | Predicting Retrieval Success Based on Information Use for Writing Tasks
Pertti Vakkari, Michael Völske, Martin Potthast, Matthias Hagen, Benno Stein 0001 |
TPDL | 2 |
| 2016 | How Writers Search: Analyzing the Search and Writing Logs of Non-fictional EssaysabstractMany writers of non-fictional texts engage intensively in exploratory web search scenarios during their background research on the essay topic. Though understanding such search behavior is necessary for the development of search engines that specifically support writing tasks, it has neither been systematically recorded nor analyzed. This paper contributes part of the missing research: We report on the outcomes of a large-scale corpus construction initiative to acquire detailed interaction logs of writers who were given a writing task on 150 pre-defined TREC topics. The corpus is freely available to foster research on exploratory search. Each essay is at least 5000 words long and comes with a chronological log of search queries, result clicks, web browsing trails, and fine-grained writing revisions that reflect the task completion status. To ensure reproducibility, a fully-fledged, static web search environment has been created on top of the ClueWeb09 corpus as part of our initiative. Matthias Hagen, Martin Potthast, Michael Völske, Jakob Gomoll, Benno Stein 0001 |
CHIIR | 3 |
| 2016 | Axiomatic Result Re-RankingabstractWe consider the problem of re-ranking the top-k documents returned by a retrieval system given some search query. This setting is common to learning-to-rank scenarios, and it is often solved with machine learning and feature weighting based on user preferences such as clicks, dwell times, etc. In this paper, we combine the learning-to-rank paradigm with the recent developments on axioms for information retrieval. In particular, we suggest to re-rank the top-k documents of a retrieval system using carefully chosen axiom combinations. In recent years, research on axioms for information retrieval has focused on identifying reasonable constraints that retrieval systems should fulfill. Researchers have analyzed a wide range of standard retrieval models for conformance to the proposed axioms and, at times, suggested certain adjustments to the models. We take up this axiomatic view---but, instead of adjusting the retrieval models themselves, we suggest the following innovation: to adopt the learning-to-rank idea and to re-rank the top-k results directly using promising axiom combinations. This way, we can turn every reasonable basic retrieval model into an axiom-based retrieval model. In large-scale experiments on the ClueWeb corpora, we identify promising axiom combinations for a variety of retrieval models. Our experiments show that for most of these models our axiom-based re-ranking significantly improves the original retrieval performance. Matthias Hagen, Michael Völske, Steve Goering, Benno Stein 0001 |
CIKM | 2 |
| 2015 | What Users Ask a Search Engine: Analyzing One Billion Russian Question QueriesabstractWe analyze the question queries submitted to a large commercial web search engine to get insights about what people ask, and to better tailor the search results to the users' needs. Based on a dataset of about one billion question queries submitted during the year 2012, we investigate askers' querying behavior with the support of automatic query categorization. While the importance of question queries is likely to increase, at present they only make up 3-4% of the total search traffic. Michael Völske, Pavel Braslavski 0001, Matthias Hagen, Galina Lezina, Benno Stein 0001 |
CIKM | 1 |
| 2013 | Learning Overlap Optimization for Domain Decomposition Methods
Steven Burrows, Jörg Frochte, Michael Völske, Ana Belén Martínez Torres, Benno Stein 0001 |
PAKDD (1) | 3 |