VLDB 2026 Research / reviewers in the wild / expert
Birger Larsen
dblp:34/421
· DBLP profile ↗
22ranked-venue papers in the field
4as first author
4since 2021 · last 2026
0000-0002-3622-2698ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 21 (4 first)Big Data, Cloud & Distributed Data Systems · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Formalized Information Needs Improve Large-Language-Model Relevance JudgmentsabstractCranfield-style retrieval evaluations with too few or too many relevant documents or with low inter-assessor agreement on relevance can reduce the reliability of observations. In evaluations with human assessors, information needs are often formalized as retrieval topics to avoid an excessive number of relevant documents while maintaining good agreement. However, emerging evaluation setups that use Large Language Models (LLMs) as relevance assessors often use only queries, potentially decreasing the reliability. To study whether LLM relevance assessors benefit from formalized information needs, we synthetically formalize information needs with LLMs into topics that follow the established structure from previous human relevance assessments (i.e., descriptions and narratives). We compare assessors using synthetically formalized topics against the LLM-default query-only assessor on the~2019/2020~editions of TREC Deep Learning and Robust04. We find that assessors without formalization judge many more documents relevant and have a lower agreement, leading to reduced reliability in retrieval evaluations. Furthermore, we show that the formalized topics improve agreement between human and LLM relevance judgments, even when the topics are not highly similar to their human counterparts. Our findings indicate that LLM relevance assessors should use formalized information needs, as is standard for human assessment, and synthetically formalize topics when no human formalization exists to improve evaluation reliability. Jüri Keller, Maik Fröbe, Björn Engelmann 0002, Fabian Haak 0001, Timo Breuer 0002, Birger Larsen, Philipp Schaer |
SIGIR | 6 |
| 2024 | The Impact of CHIIR Publications: A Study of Eight Years of CHIIRabstractAcross all scientific fields, there is an increased focus on the impact of scientific research: what academic and societal benefits does it provide? This question has spurred the development of a variety of different approaches to impact assessment, each appropriate in different circumstances. In this paper, we study the academic impact of the CHIIR community through a comprehensive analysis of the work published in the 2016-2023 CHIIR conference series. We collect citation counts, citing documents, and altmetrics scores for all CHIIR publications to determine their academic impact across a variety of different attributes of the CHIIR publications. In addition, we analyze a subset of citation contexts in the papers that have cited CHIIR publications to analyze how they are being used and what that means for their potential impact. Finally, we attempt to predict which properties of CHIIR publications are most predictive of future impact. Maria Gäde, Toine Bogers, Mark M. Hall, Marijn Koolen, Vivien Petras, Birger Larsen |
CHIIR | 6 |
| 2023 | How we Work, Share, and Re-use at CHIIRabstractIn this paper, we present the results of an initial study of the research, sharing, and re-use practices at the CHIIR conference through a systematic analysis of all CHIIR papers published from 2016 to 2022. We find that CHIIR is a conference predominantly focused on empirical, multi-methods research that over the years has undergone a focusing in terms of the type of research methods that are being used. A modest number of papers re-use existing data and design resources, but infrastructure component re-use is much more rare. Only a fraction of CHIIR papers actually share their own resources, which suggests that there is much to gain in terms of reproducibility of research presented at CHIIR and could potentially be used to support changes in reviewing practices. Toine Bogers, Maria Gäde, Mark M. Hall, Marijn Koolen, Vivien Petras, Birger Larsen |
CHIIR | 6 |
| 2023 | Collaboration Patterns and Impact of Sharing at CHIIRabstractWe studied the collaboration patterns of CHIIR authors, and found that most papers are collaborative. A core of 33% of the CHIIR researchers are directly connected and frequently co-author, and several disconnected clusters also make frequent CHIIR contributions. We also studied citation impact of the CHIIR papers and show that in relation to research design type, theoretical and empirical papers tend to receive more citations than resource papers. With regards to sharing and re-use, papers that share at least one resource tend to have significantly higher citation impact—in particular when sharing data resources and design resources. Re-using resources does not significantly increase citation impact in itself. Toine Bogers, Birger Larsen, Marijn Koolen, Maria Gäde, Mark M. Hall, Vivien Petras |
CHIIR | 2 |
| 2020 | Hands-free but not Eyes-free: A Usability Evaluation of Siri while DrivingabstractDistractions while driving are a major cause of traffic accidents and chief among these is the use of mobile phones. Driver distractions typically fall into four categories-visual, cognitive, bio-mechanical, and auditory-and different technological solutions have been proposed to address these. Intelligent Personal Assistants (IPAs), such as Siri, is a recent example of such a technological solution that offers the potential for hands-free phone interaction through a voice-controlled interface. IPAs could potentially reduce visual and bio-mechanical distractions if they are usable enough to not increase a driver's cognitive load. We present the results of a controlled experiment with the aim of understanding how the use of Siri while driving compares to manual interaction in terms of usability and distractions. We also tested these two interaction types in the lab in order to understand how the main driving task influences Siri's (perceived) usability. Our study shows that Siri is not ready for every-day use in the car: interacting with Siri while driving is likely to be unsafe for most participants, especially less experienced drivers. Participants were distracted by Siri due to its over-reliance on visual feedback as well as frequent time-outs by Siri when waiting for a response from a driver occupied with the road environment. Speech recognition quality in a noisy car as well as problematic multi-lingual speech recognition in general are other issues that resulted in low usability and more cognitive distractions. While interacting with Siri may be hands-free, it does not provide an eyes-free and distraction-free experience yet. Helene Høgh Larsen, Alexander Nuka Scheel, Toine Bogers, Birger Larsen |
CHIIR | 4 |
| 2020 | Factuality Checking in News Headlines with Eye TrackingabstractWe study whether it is possible to infer if a news headline is true or false using only the movement of the human eyes when reading news headlines. Our study with 55 participants who are eye-tracked when reading 108 news headlines (72 true, 36 false) shows that false headlines receive statistically significantly less visual attention than true headlines. We further build an ensemble learner that predicts news headline factuality using only eye-tracking measurements. Our model yields a mean AUC of 0.688 and is better at detecting false than true headlines. Through a model analysis, we find that eye-tracking 25 users when reading 3-6 headlines is sufficient for our ensemble learner. Christian Hansen 0004, Casper Hansen, Jakob Grue Simonsen, Birger Larsen, Stephen Alstrup, Christina Lioma |
SIGIR | 4 |
| 2018 | I just scroll through my stuff until I find it or give up: A Contextual Inquiry of PIM on Private Handheld DevicesabstractWhile ownership and usage of handheld devices such as smartphones and tablets continues to grow at a rapid pace, we do not have complete picture of how people manage personal information on these devices. The few existing studies have typically used interview or survey methods to focus on personal information management (PIM) practices on smartphones. We present the results of an exploratory contextual inquiry study of PIM practices aimed at providing a structured, naturalistic overview of PIM on both smartphones and tablets. We find that people use multiple complementary strategies to acquire different types of information on their devices, and that people rely strongly on automatic chronological ordering instead of organization by subject, although this pays off most for smaller information collections. Deletion of information is strongly influenced by usefulness and personal attachment. Finally, we find that people strongly prefer browsing over search when retrieving information from their devices. Amalie Enshelm Jensen, Caroline Møller Jægerfelt, Sanne Francis, Birger Larsen, Toine Bogers |
CHIIR | 4 |
| 2018 | First International Workshop on Professional Search (ProfS2018)abstractProfessional search is a problem area in which many facets of information retrieval are addressed, both system-related (e.g. distributed search) and user-related (e.g. complex information needs), and the interface between user and system (e.g. supporting exploratory search tasks). Professional search tasks have specific requirements, different from the requirements of generic web search. The aim of this workshop is to bring together researchers to work on the requirements and challenges of professional search from different angles. We will have an interactive workshop where researchers not only present their scientific results but also work together on the definition of future challenges and solutions with input from information professionals. The workshop will deliver a roadmap of research directions for the years to come. Suzan Verberne, Jiyin He, Udo Kruschwitz, Birger Larsen, Tony Russell-Rose, Arjen P. de Vries |
SIGIR | 4 |
| 2016 | A study of factuality, objectivity and relevance: three desiderata in large-scale information retrieval?abstractMuch of the information processed by Information Retrieval (IR) systems is unreliable, biased, and generally untrust-worthy [15, 45, 48]. Yet, factuality & objectivity detection is not a standard component of IR systems, even though it has been possible in Natural Language Processing (NLP) in the last decade. Motivated by this, we ask if and how factuality & objectivity detection may benefit IR. We answer this in two parts. First, we use state-of-the-art NLP to compute the probability of document factuality & objectivity in two TREC collections, and analyse its relation to document relevance. We find that factuality is strongly and positively correlated to document relevance, but objectivity is not. Second, we study the impact of factuality & objectivity to retrieval effectiveness by treating them as query independent features that we combine with a competitive language modelling baseline. Experiments with 450 TREC queries show that factuality improves precision by more than 10% over strong baselines, especially for the type of uncurated data typically used in web search; objectivity gives mixed results. An overall clear trend is that document factuality & objectivity is much more beneficial to IR when searching uncurated (e.g. web) documents vs. curated (e.g. state documentation and newswire articles). Christina Lioma, Birger Larsen, Wei Lu 0019, Yong Huang 0008 |
BDCAT | 2 |
| 2015 | Non-Compositional Term Dependence for Information RetrievalabstractModelling term dependence in IR aims to identify co-occurring terms that are too heavily dependent on each other to be treated as a bag of words, and to adapt the indexing and ranking accordingly. Dependent terms are predominantly identified using lexical frequency statistics, assuming that (a) if terms co-occur often enough in some corpus, they are semantically dependent; (b) the more often they co-occur, the more semantically dependent they are. This assumption is not always correct: the frequency of co-occurring terms can be separate from the strength of their semantic dependence. E.g. "red tape" might be overall less frequent than "tape measure" in some corpus, but this does not mean that "red"+"tape" are less dependent than "tape"+"measure". This is especially the case for non-compositional phrases, i.e. phrases whose meaning cannot be composed from the individual meanings of their terms (such as the phrase "red tape" meaning bureaucracy). Motivated by this lack of distinction between the frequency and strength of term dependence in IR, we present a principled approach for handling term dependence in queries, using both lexical frequency and semantic evidence. We focus on non-compositional phrases, extending a recent unsupervised model for their detection (Kiela & Clark 2013) to IR. Our approach, integrated into ranking using Markov Random Fields (Metzler & Croft 2005), yields effectiveness gains over competitive TREC baselines, showing that there is still room for improvement in the very well-studied area of term dependence in IR. Christina Lioma, Jakob Grue Simonsen, Birger Larsen, Niels Dalum Hansen |
SIGIR | 3 |
| 2014 | Bibliometric-Enhanced Information Retrieval
Philipp Mayr 0001, Andrea Scharnhorst, Birger Larsen, Philipp Schaer, Peter Mutschke |
ECIR | 3 |
| 2013 | Integrating IR Technologies for Professional Search - (Full-Day Workshop)
Michail Salampasis, Norbert Fuhr, Allan Hanbury, Mihai Lupu, Birger Larsen, Henrik Strindberg |
ECIR | 5 |
| 2013 | EuroHCIR2013: the 3rd European workshop on human-computer interaction and information retrievalabstractA proposal summary for the EuroHCIR workshop at SIGIR2013. Max L. Wilson 0001, Birger Larsen, Preben Hansen, Kristian Norling, Tony Russell-Rose |
SIGIR | 2 |
| 2012 | Preliminary study of technical terminology for the retrieval of scientific book metadata recordsabstractBooks only represented by brief metadata (book records) are particularly hard to retrieve. One way of improving their retrieval is by extracting retrieval enhancing features from them. This work focusses on scientific (physics) book records. We ask if their technical terminology can be used as a retrieval enhancing feature. A study of 18,443 book records shows a strong correlation between their technical terminology and their likelihood of relevance. Using this finding for retrieval yields >+5% precision and recall gains. Birger Larsen, Christina Lioma, Ingo Frommholz, Hinrich Schütze |
SIGIR | 1 |
| 2012 | Rhetorical relations for information retrievalabstractTypically, every part in most coherent text has some plausible reason for its presence, some function that it performs to the overall semantics of the text. Rhetorical relations, e.g. contrast, cause, explanation, describe how the parts of a text are linked to each other. Knowledge about this so-called discourse structure has been applied successfully to several natural language processing tasks. This work studies the use of rhetorical relations for Information Retrieval (IR): Is there a correlation between certain rhetorical relations and retrieval performance? Can knowledge about a document's rhetorical relations be useful to IR? We present a language model modification that considers rhetorical relations when estimating the relevance of a document to a query. Empirical evaluation of different versions of our model on TREC settings shows that certain rhetorical relations can benefit retrieval effectiveness notably (>10% in mean average precision over a state-of-the-art baseline). Christina Lioma, Birger Larsen, Wei Lu 0019 |
SIGIR | 2 |
| 2012 | The tipping point: F-score as a function of the number of retrieved items
Raf Guns, Christina Lioma, Birger Larsen |
Inf. Process. Manag. | 3 |
| 2010 | Developing a Test Collection for the Evaluation of Integrated Search
Marianne Lykke, Birger Larsen, Haakon Lund, Peter Ingwersen |
ECIR | 2 |
| 2009 | Representing User Navigation in XML Retrieval with Structural Summaries
Mir Sadek Ali, Mariano P. Consens, Birger Larsen |
ECIR | 3 |
| 2009 | Data fusion according to the principle of polyrepresentationabstractAbstract We report data fusion experiments carried out on the four best‐performing retrieval models from TREC 5. Three were conceptually/algorithmically very different from one another; one was algorithmically similar to one of the former. The objective of the test was to observe the performance of the 11 logical data fusion combinations compared to the performance of the four individual models and their intermediate fusions when following the principle of polyrepresentation. This principle is based on cognitive IR perspective (Ingwersen & Järvelin, 2005) and implies that each retrieval model is regarded as a representation of a unique interpretation of information retrieval (IR). It predicts that only fusions of very different, but equally good, IR models may outperform each constituent as well as their intermediate fusions. Two kinds of experiments were carried out. One tested restricted fusions, which entails that only the inner disjoint overlap documents between fused models are ranked. The second set of experiments was based on traditional data fusion methods. The experiments involved the 30 TREC 5 topics that contain more than 44 relevant documents. In all tests, the Borda and CombSUM scoring methods were used. Performance was measured by precision and recall, with document cutoff values (DCVs) at 100 and 15 documents, respectively. Results show that restricted fusions made of two, three, or four cognitively/algorithmically very different retrieval models perform significantly better than do the individual models at DCV100. At DCV15, however, the results of polyrepresentative fusion were less predictable. The traditional fusion method based on polyrepresentation principles demonstrates a clear picture of performance at both DCV levels and verifies the polyrepresentation predictions for data fusion in IR. Data fusion improves retrieval performance over their constituent IR models only if the models all are quite conceptually/algorithmically dissimilar and equally and well performing, in that order of importance. Birger Larsen, Peter Ingwersen, Berit Lund |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2008 | Inter and intra-document contexts applied in polyrepresentation for best match IR
Mette Skov, Birger Larsen, Peter Ingwersen |
Inf. Process. Manag. | 2 |
| 2006 | Is XML retrieval meaningful to users?: searcher preferences for full documents vs. elementsabstractThe aim of this study is to investigate whether element retrieval (as opposed to full-text retrieval) is meaningful and useful for searchers when carrying out information-seeking tasks. Our results suggest that searchers find the structural breakdown of documents useful when browsing within retrieved documents, and provide support for the usefulness of element retrieval in interactive settings. Birger Larsen, Anastasios Tombros, Saadia Malik |
SIGIR | 1 |
| 2002 | The boomerang effect: retrieving scientific documents via the network of references and citationsabstractNo abstract available. Birger Larsen, Peter Ingwersen |
SIGIR | 1 |