VLDB 2026 Research / reviewers in the wild / expert
Maarten Marx
dblp:m/MaartenMarx
· DBLP profile ↗
46ranked-venue papers in the field
8as first author
7since 2021 · last 2025
0000-0003-3255-3729ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 25 (1 first)Database Systems & Data Management · 15 (6 first)Knowledge Engineering, Semantic Web & Information Systems · 5 (1 first)Other / Interdisciplinary · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Lost but Not Only in the Middle - Positional Bias in Retrieval Augmented Generation
Jan Hutter, David Rau, Maarten Marx, Jaap Kamps |
ECIR (1) | 3 |
| 2025 | Spoken Question Answering on Municipal Council Meetings
Pepijn van Wijk, Maarten Marx |
ECIR (5) | 2 |
| 2024 | OpenPSS: An Open Page Stream Segmentation Benchmark
Ruben van Heusden, Jaap Kamps, Maarten Marx |
TPDL (1) | 3 |
| 2024 | Bcubed revisited: elements like meabstractAbstract BCubed is a mathematically clean, elegant and intuitively well behaved external performance metric for clustering tasks. BCubed compares a predicted clustering to a known ground truth clustering through elementwise precision and recall scores. For each element, the predicted and ground truth clusters containing the element are compared, and the mean over all elements is taken. We argue that BCubed overestimates performance, for the intuitive reason that the clustering gets credit for putting an element into its own cluster. This is repaired, and we investigate the repaired version, called “Elements Like Me (ELM)”. We extensively evaluate ELM from both a theoretical and empirical perspective, and conclude that it retains all of its positive properties, and yields a minimum zero score when it should. Synthetic experiments show that ELM can produce different rankings of predicted clusterings when compared to BCubed, and that the ELM scores are distributed with lower mean and a larger variance than BCubed. Ruben van Heusden, Jaap Kamps, Maarten Marx |
Discov. Comput. | 3 |
| 2023 | Enticing Local Governments to Produce FAIR Freedom of Information Act Dossiers
Maarten Marx, Maik Larooij, Filipp Perasedillo, Jaap Kamps |
ECIR (3) | 1 |
| 2023 | Making PDFs Accessible for Visually Impaired Users (and Findable for Everybody Else)
Ruben van Heusden, Hazel Ling, Lars Nelissen, Maarten Marx |
TPDL | 4 |
| 2023 | Detection of Redacted Text in Legal Documents
Ruben van Heusden, Aron de Ruijter, Roderick Majoor, Maarten Marx |
TPDL | 4 |
| 2019 | HiTR: Hierarchical Topic Model Re-Estimation for Measuring Topical Diversity of Documents
Hosein Azarbonyad, Mostafa Dehghani 0001, Tom Kenter, Maarten Marx, Jaap Kamps, Maarten de Rijke |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2017 | Telling How to Narrow it Down: Browsing Path Recommendation for Exploratory SearchabstractSupporting exploratory search tasks with the help of structured data is an effective way to go beyond keyword search, as it provides an overview of the data, enables users to zoom in on their intent, and provides assistance during their navigation trails. However, finding a good starting point for a search episode in the given structure can still pose a considerable challenge, as users tend to be unfamiliar with exact, complex hierarchical structure. Thus, providing lookahead clues can be of great help and allow users to make better decisions on their search trajectory. Mostafa Dehghani 0001, Glorianna Jagfeld, Hosein Azarbonyad, Alex Olieman, Jaap Kamps, Maarten Marx |
CHIIR | 6 |
| 2017 | Words are Malleable: Computing Semantic Shifts in Political and Media DiscourseabstractRecently, researchers started to pay attention to the detection of temporal shifts in the meaning of words. However, most (if not all) of these approaches restricted their efforts to uncovering change over time, thus neglecting other valuable dimensions such as social or political variability. We propose an approach for detecting semantic shifts between different viewpoints---broadly defined as a set of texts that share a specific metadata feature, which can be a time-period, but also a social entity such as a political party. For each viewpoint, we learn a semantic space in which each word is represented as a low dimensional neural embedded vector. The challenge is to compare the meaning of a word in one space to its meaning in another space and measure the size of the semantic shifts. We compare the effectiveness of a measure based on optimal transformations between the two spaces with a measure based on the similarity of the neighbors of the word in the respective spaces. Our experiments demonstrate that the combination of these two performs best. We show that the semantic shifts not only occur over time but also along different viewpoints in a short period of time. For evaluation, we demonstrate how this approach captures meaningful semantic shifts and can help improve other tasks such as the contrastive viewpoint summarization and ideology detection (measured as classification accuracy) in political texts. We also show that the two laws of semantic change which were empirically shown to hold for temporal shifts also hold for shifts across viewpoints. These laws state that frequent words are less likely to shift meaning while words with many senses are more likely to do so. Hosein Azarbonyad, Mostafa Dehghani 0001, Kaspar Beelen, Alexandra Arkut, Maarten Marx, Jaap Kamps |
CIKM | 5 |
| 2017 | Hierarchical Re-estimation of Topic Models for Measuring Topical Diversity
Hosein Azarbonyad, Mostafa Dehghani 0001, Tom Kenter, Maarten Marx, Jaap Kamps, Maarten de Rijke |
ECIR | 4 |
| 2017 | Containment of acyclic conjunctive queries with negated atoms or arithmetic comparisons
Evgeny Sherkhonov, Maarten Marx |
Inf. Process. Lett. | 2 |
| 2016 | Generalized Group Profiling for Content CustomizationabstractThere is an ongoing debate on personalization, adapting results to the unique user exploiting a user's personal history, versus customization, adapting results to a group profile sharing one or more characteristics with the user at hand. Personal profiles are often sparse, due to cold start problems and the fact that users typically search for new items or information, necessitating to back-off to customization, but group profiles often suffer from accidental features brought in by the unique individual contributing to the group. In this paper we propose a generalized group profiling approach that teases apart the exact contribution of the individual user level and the `abstract' group level by extracting a latent model that captures all, and only, the essential features of the whole group. Our main findings are the followings. Mostafa Dehghani 0001, Hosein Azarbonyad, Jaap Kamps, Maarten Marx |
CHIIR | 4 |
| 2016 | Luhn Revisited: Significant Words Language ModelsabstractUsers tend to articulate their complex information needs in only a few keywords, making underspecified statements of request the main bottleneck for retrieval effectiveness. Taking advantage of feedback information is one of the best ways to enrich the query representation, but can also lead to loss of query focus and harm performance in particular when the initial query retrieves only little relevant information when overfitting to accidental features of the particular observed feedback documents. Inspired by the early work of Luhn [23], we propose significant words language models of feedback documents that capture all, and only, the significant shared terms from feedback documents. We adjust the weights of common terms that are already well explained by the document collection as well as the weight of rare terms that are only explained by specific feedback documents, which eventually results in having only the significant terms left in the feedback model. Mostafa Dehghani 0001, Hosein Azarbonyad, Jaap Kamps, Djoerd Hiemstra, Maarten Marx |
CIKM | 5 |
| 2016 | Containment for queries over trees with attribute value comparisons
Maarten Marx, Evgeny Sherkhonov |
Inf. Syst. | 1 |
| 2015 | Sources of Evidence for Automatic Indexing of Political Texts
Mostafa Dehghani 0001, Hosein Azarbonyad, Maarten Marx, Jaap Kamps |
ECIR | 3 |
| 2015 | Time-Aware Authorship Attribution for Short Text StreamsabstractIdentifying authors of short texts on Internet or social media based communication systems is an important tool against fraud and cybercrimes. Besides the challenges raised by the limited length of these short messages, evolving language and writing styles of authors of these texts makes authorship attribution difficult. Most current short text authorship attribution approaches only address the challenge of limited text length. However, neglecting the second challenge may lead to poor performance of authorship attribution for authors who change their writing styles. Hosein Azarbonyad, Mostafa Dehghani 0001, Maarten Marx, Jaap Kamps |
SIGIR | 3 |
| 2014 | Automatic thematic classification of election manifestos
Suzan Verberne, Eva D'hondt, Antal van den Bosch, Maarten Marx |
Inf. Process. Manag. | 4 |
| 2013 | PoliticalMashup Ngramviewer - Tracking Who Said What and When in Parliament
Bart de Goede, Justin van Wees, Maarten Marx, Ridho Reinanda |
TPDL | 3 |
| 2013 | Linking the kingdom: enriched access to a historiographical textabstractDigital history is a branch of digital humanities concerned using ICT to improve study of history. Linked Data provides a way of effective enriched digital access to scientific texts about history (historiographies). In this paper, we present a method for connecting a historiographical text to the Linked Data cloud. We present the method and tools that we use in each of the method's steps. We focus on one extensive case study: the enriched access of an important work of Dutch World War II historiography "Het Koninkrijk der Nederlanden in de Tweede Wereldoorlog". We describe the digitization and present two sources of structured knowledge that link to individual text sources, retrievable on the Web of Data. The first is the manually constructed and highly curated "Back of the Book Index". The second is a list of extracted Named Entities. We compare both structured sources as stepping stones to the Web of Data and present a number of use cases relevant for both historical researchers as well as for the general public. Victor de Boer, Johan van Doornik, Lars Buitinck, Maarten Marx, Tim Veken, Kees Ribbens |
K-CAP | 4 |
| 2013 | Containment for Tree Patterns with Attribute Value Comparisons
Evgeny Sherkhonov, Maarten Marx |
WebDB | 2 |
| 2013 | The quality of the XML Web
Steven Grijzenhout, Maarten Marx |
J. Web Semant. | 2 |
| 2012 | Two-Stage Named-Entity Recognition Using Averaged Perceptrons
Lars Buitinck, Maarten Marx |
NLDB | 2 |
| 2011 | The quality of the XML webabstractWe collect evidence to answer the following question: Is the quality of the XML documents found on the web sufficient to apply XML technology like XQuery, XPath and XSLT? XML collections from the web have been previously studied statistically, but no detailed information about the quality of the XML documents on the web is available to date. We address this shortcoming in this study. We gathered 180K XML documents from the web. Their quality is surprisingly good; 85.4% is well-formed and 99.5% of all specified encodings is correct. Validity needs serious attention. Only 25% of all files contain a reference to a DTD or XSD, of which just one third is actually valid. Errors are studied in detail. Automatic error repair seems promising. Our study is well documented and easily repeatable. This paves the way for a periodic quality assessment of the XML web. Steven Grijzenhout, Maarten Marx |
CIKM | 2 |
| 2010 | Tree patterns with Full Text SearchabstractTree patterns with full text search form the core of both XQuery Full Text and the NEXI query language. On such queries, users expect a relevance-ranked list of XML elements as an answer. But this requirement may lead to undesirable behavior of XML retrieval systems: two queries which are intuitively (e.g., without ranking) equivalent return differently ordered lists of elements. We show that the best performing XML retrieval semantics has this behavior. We also show how minimization of tree patterns can efficiently solve this problem. Maria-Hendrike Peetz, Maarten Marx |
WebDB | 2 |
| 2010 | Focused retrieval and result aggregation with political dataabstractThis paper presents a case-study in which we use a large semi-structured data set consisting of official transcripts of meetings of the Dutch parliament for focused retrieval and result aggregation. Transcripts of meetings are a document genre characterized by a complex narrative structure. The essence is not only what is said, but also by who and to whom. We have notes of more than 40 years of Dutch parliamentary debates where this structure is exploited to automatically make semantic annotations. These annotations yield numerous new ways of searching, browsing, mining and summarizing these documents. Concerning result aggregation, we summarise and visualise the structure of meetings into tables of content and interruption graphs. The contents of meetings or parts of meetings are condensed into word clouds that are created using a parsimonious language model. Furthermore, we have developed a search engine that exploits the structure and annotations of our data making it possible to provide entry points, to group search results, and to use faceted search techniques for data-exploration. Evaluation shows that our content and structure summarization tools provide a good first impression of a debate. Users reported that, compared to a standard document retrieval system, our search engine gives a better overview of the data. Search tasks are performed faster and the users felt more certain of their answers. Rianne Kaptein, Maarten Marx |
Inf. Retr. | 2 |
| 2009 | Helping people to choose for whom to vote. a web information system for the 2009 European electionsabstractWe demonstrate a web information system created for the European elections in June 2009. Based on their speeches in the EU parliament and their written questions, we created language models for each of the 736 members of the EU parliament. These language models were used to search for politicians responsible for a given topic, similar to expert search applications. Users prefer to see some kind of evidence for returning a hit after a search. We created a profile of each EU parlementarian by comparing her personal language model to the language model created from all EU parlementarians. The top 50 words best separating the individual from the avarage were shown as a wordcloud. These top 50 words and their scores were derived from a parsimonious language model. Arjan Nusselder, Maria-Hendrike Peetz, Anne Schuth, Maarten Marx |
CIKM | 4 |
| 2009 | Recursion in XQuery: put your distributivity safety belt onabstractWe introduce a controlled form of recursion in XQuery, an inflationary fixed point operator, familiar from the context of relational databases. This operator imposes restrictions on the expressible types of recursion, but it is sufficiently versatile to capture a wide range of interesting use cases, including Regular XPath and its core transitive closure operator. Loredana Afanasiev, Torsten Grust, Maarten Marx, Jan Rittinger, Jens Teubner |
EDBT | 3 |
| 2009 | Who said what to whom?: capturing the structure of debatesabstractTranscripts of meetings are a document genre characterized by a complex narrative structure. The essence is not only what is said, but also by who and to whom. This paper investigates whether we can use semantic annotations like the speaker in order to capture this debate structure, as well as the related content of the debate. The structure is visualized in a graph, while the content is condensed into word clouds, that are created using a parsimonious language model. Evaluation shows that both tools adequately capture the structure and content of the debate at an aggregated level. Rianne Kaptein, Maarten Marx, Jaap Kamps |
SIGIR | 2 |
| 2008 | An Inflationary Fixed Point Operator in XQueryabstractWe introduce a controlled form of recursion in XQuery, an inflationary fixed point operator, familiar from the context of relational databases. This operator imposes restrictions on the expressible types of recursion, but we show that it is sufficiently versatile to capture a wide range of interesting use cases, including Regular XPath and its core transitive closure operator. While the optimization of general user-defined recursive functions in XQuery appears elusive, we describe how inflationary fixed points can be efficiently evaluated, provided that the recursive XQuery expressions are distributive. We test distributivity syntactically and algebraically, and provide experimental evidence that XQuery processors can benefit substantially from this mode of evaluation. Loredana Afanasiev, Torsten Grust, Maarten Marx, Jan Rittinger, Jens Teubner |
ICDE | 3 |
| 2008 | An analysis of XQuery benchmarks
Loredana Afanasiev, Maarten Marx |
Inf. Syst. | 2 |
| 2007 | Axiomatizing the Logical Core of XPath 2.0
Balder ten Cate, Maarten Marx |
ICDT | 2 |
| 2007 | Queries determined by views: pack your viewsabstractA query Q is determined by a set of views V if, whenever V (I1) = V (I2) for two database instances I1, I2 then also Q(I1) = Q(I2). Does this imply that Q can be rewritten as a query Q0 that only uses the views V?. Maarten Marx |
PODS | 1 |
| 2007 | Electoral search using the VerkiezingsKijker: an experience reportabstractThe Netherlands had parliamentary elections on November 22, 2006. We built a system which helped voters to make an informed choice among the many participating parties. One of the most important pieces of information in the Dutch election and subsequent coalition government formation is the party program, a text document with an average length of 45 pages. Our system provides the voter with focused access to party programs, enabling her to make a topic-wise comparison of parties' viewpoints. We complemented this type of access ("What do the parties promise?") with access to news ("What happens around these topics?") and blogs ("What do people say about them?"). We describe the system, including design technical details, and user statistics. Valentin Jijkoun, Maarten Marx, Maarten de Rijke, Frank van Waveren |
WWW | 2 |
| 2006 | XCheck: A Platform for Benchmarking XQuery Engines
Loredana Afanasiev, Massimo Franceschet, Maarten Marx, Enrico Zimuel |
VLDB | 3 |
| 2006 | Articulating information needs in XML query languagesabstractDocument-centric XML is a mixture of text and structure. With the increased availability of document-centric XML documents comes a need for query facilities in which both structural constraints and constraints on the content of the documents can be expressed. How does the expressiveness of languages for querying XML documents help users to express their information needs? We address this question from both an experimental and a theoretical point of view. Our experimental analysis compares a structure-ignorant with a structure-aware retrieval approach using the test suite of the INEX XML Retrieval Evaluation Initiative. Theoretically, we create two mathematical models of users' knowledge of a set of documents and define query languages which exactly fit these models. One of these languages corresponds to an XML version of fielded search, the other to the INEX query language.Our main experimental findings are: First, while structure is used in varying degrees of complexity, two-thirds of the queries can be expressed in a fielded-search-like format which does not use the hierarchical structure of the documents. Second, three-quarters of the queries use constraints on the context of the elements to be returned; these contextual constraints cannot be captured by ordinary keyword queries. Third, structure is used as a search hint, and not as a strict requirement, when judged against the underlying information need. Fourth, the use of structure in queries functions as a precision enhancing device. Jaap Kamps, Maarten Marx, Maarten de Rijke, Börkur Sigurbjörnsson |
ACM Trans. Inf. Syst. | 2 |
| 2005 | Structured queries in XML retrievalabstractDocument-centric XML is a mixture of text and structure. With the increased availability of document-centric XML content comes a need for query facilities in which both structural constraints and constraints on the content of the documents can be expressed. How does the expressiveness of languages for querying XML documents help users to express their information needs? We address this question from both an experimental and a theoretical point of view. Our experimental analysis compares a structure-ignorant with a structure-aware retrieval approach using the test-suite of the 2004 edition of the INEX XML retrieval evaluation initiative. Theoretically, we create mathematical models of users' knowledge of a set of documents and define query languages which exactly fit these models. One of these languages corresponds to an XML version of fielded search, the other to the INEX query language. Our main findings are: First, while structure is used in varying degrees of complexity, over half of the queries can be expressed in a fielded-search like format which does not use the hierarchical structure of the documents. Second, structure is used as a search hint, and not a strict requirement, when judged against the underlying information need. Third, the use of structure in queries functions as a precision enhancing device. Jaap Kamps, Maarten Marx, Maarten de Rijke, Börkur Sigurbjörnsson |
CIKM | 2 |
| 2005 | First Order Paths in Ordered Trees
Maarten Marx |
ICDT | 1 |
| 2005 | Conditional XPathabstractXPath 1.0 is a variable free language designed to specify paths between nodes in XML documents. Such paths can alternatively be specified in first-order logic. The logical abstraction of XPath 1.0, usually called Navigational or Core XPath, is not powerful enough to express every first-order definable path. In this article, we show that there exists a natural expansion of Core XPath in which every first-order definable path in XML document trees is expressible. This expansion is called Conditional XPath. It contains additional axis relations of the form (child::n[F])+, denoting the transitive closure of the path expressed by child::n[F]. The difference with XPath's descendant::n[F] is that the path (child::n[F])+ is conditional on the fact that all nodes in between the start and end node of the path should also be labeled by n and should make the predicate F true. This result can be viewed as the XPath analogue of the expressive completeness of the relational algebra with respect to first-order logic. Maarten Marx |
ACM Trans. Database Syst. | 1 |
| 2004 | XPath with Conditional Axis Relations
Maarten Marx |
EDBT | 1 |
| 2004 | Conditional XPath, the First Order Complete XPath DialectabstractXPath is the W3C -- standard node addressing language for XML documents. XPath is still under development and its technical aspects are intensively studied. What is missing at present is a clear characterization of the expressive power of XPath, be it either semantical or with reference to some well established existing (logical) formalism. Core XPath (the logical core of XPath 1.0 defined by Gottlob et al.) cannot express queries with conditional paths as exemplified by "do a child step, while test is true at the resulting node." In a first-order complete extension of Core XPath, such queries are expressible, We add conditional axis relations to Core XPath and show that the resulting language, called conditional XPath, is equally expressive as first-order logic when interpreted on ordered trees. Both the result, the extended XPath language, and the proof are closely related to temporal logic. Specifically, while Core XPath may be viewed as a simple temporal logic, conditional XPath extends this with (counterparts of) the since and until operators. Maarten Marx |
PODS | 1 |
| 2004 | Information Retrieval Support for Ontology Construction and Use
Willem Robert van Hage, Maarten de Rijke, Maarten Marx |
ISWC | 3 |
| 2004 | Best-Match Querying from Document-Centric XMLabstractOn the Web, there is a pervasive use of XML to give lightweight semantics to textual collections. Such document-centric XML collections require a query language that can gracefully handle structural constraints as well as constraints on the free text of the documents. Our main contributions are three-fold. First, we outline two fragments of XPath tailored to users that have varying degrees of understanding of the XML structure used, and give both syntactic and semantic characterizations of these fragments. Second, we extend XPath with an about function having a best-match semantics based on the relevance of the document component for the expressed information need. Third, we evaluate the resulting query language using the INEX 2003 test suite, and show that best-match approaches outperform exact-match approaches for evaluating content-and-structure queries. Jaap Kamps, Maarten Marx, Maarten de Rijke, Börkur Sigurbjörnsson |
WebDB | 2 |
| 2003 | XML retrieval: what to retrieve?abstractThe fundamental difference between standard information retrieval and XML retrieval is the unit of retrieval. In traditional IR, the unit of retrieval is fixed: it is the complete document. In XML retrieval, every XML element in a document is a retrievable unit. This makes XML retrieval more difficult: besides being relevant, a retrieved unit should be neither too large nor too small. The research presented here, a comparative analysis of two approaches to XML retrieval, aims to shed light on which XML elements should be retrieved. The experimental evaluation uses data from the Initiative for the Evaluation of XML retrieval (INEX 2002). Jaap Kamps, Maarten Marx, Maarten de Rijke, Börkur Sigurbjörnsson |
SIGIR | 2 |
| 2002 | Notions of Indistinguishability for Semantic Web Languages
Jaap Kamps, Maarten Marx |
ISWC | 2 |
| 1999 | Relation Algebras can Tile
Maarten Marx |
Inf. Sci. | 1 |