EDBT 2026 Demo / reviewers in the wild / expert
Edgar Meij
dblp:58/2923
· DBLP profile ↗
47ranked-venue papers
12as first author
4since 2021 · last 2025
0000-0003-0516-3688ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 43 · 12 first-author · 4 since 2021Artificial intelligence and machine learning · 13 · 3 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Software engineering, systems software and programming languages · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Benchmark Granularity and Model Robustness for Image-Text Retrieval: A Reproducibility StudyabstractImage-Text Retrieval (ITR) systems are central to multimodal information access, with Vision-Language Models (VLMs) showing strong performance on standard benchmarks. However, these benchmarks predominantly rely on coarse-grained annotations, limiting their ability to reveal how models would perform under real-world conditions, where query granularity varies. Motivated by this gap, we examine how dataset granularity and query perturbations affect retrieval performance and robustness across four architecturally diverse VLMs (ALIGN, AltCLIP, CLIP, and GroupViT). Using both standard benchmarks (MS-COCO, Flickr30k) and their fine-grained variants, we show that richer captions consistently enhance retrieval, especially in text-to-image tasks, where we observe an average improvement of 16.23%, compared to 6.44% in image-to-text. To assess robustness, we introduce a taxonomy of perturbations and conduct extensive experiments, revealing that while perturbations typically degrade performance, they can also unexpectedly improve retrieval, exposing nuanced model behaviors. Notably, word order emerges as a critical factor - contradicting prior assumptions of model insensitivity to it. Our results highlight variation in model robustness and a dataset-dependent relationship between caption granularity and perturbation sensitivity and emphasize the necessity of evaluating models on datasets of varying granularity. Mariya Hendriksen, Shuo Zhang 0006, Ridho Reinanda, Mohamed Yahya 0001, Edgar Meij, Maarten de Rijke |
SIGIR | 5 |
| 2023 | ECIR 23 Tutorial: Neuro-Symbolic Approaches for Information Retrieval
Laura Dietz, Hannah Bast, Shubham Chatterjee, Jeff Dalton 0001, Edgar Meij, Arjen P. de Vries |
ECIR (3) | 5 |
| 2022 | Similarity-based Multi-Domain Dialogue State Tracking with Copy Mechanisms for Task-based Virtual Personal AssistantsabstractTask-based Virtual Personal Assistants (VPAs) rely on multi-domain Dialogue State Tracking (DST) models to monitor goals throughout a conversation. Previously proposed models show promising results on established benchmarks, but they have difficulty adapting to unseen domains due to domain-specific parameters in their model architectures. We propose a new Similarity-based Multi-domain Dialogue State Tracking model (SM-DST) that uses retrieval-inspired and fine-grained contextual token-level similarity approaches to efficiently and effectively track dialogue state. The key difference with state-of-the-art DST models is that SM-DST has a single model with shared parameters across domains and slots. Because we base SM-DST on similarity it allows the transfer of tracking information between semantically related domains as well as to unseen domains without retraining. Furthermore, we leverage copy mechanisms that consider the system’s response and the dialogue state from previous turn predictions, allowing it to more effectively track dialogue state for complex conversations. We evaluate SM-DST on three variants of the MultiWOZ DST benchmark datasets. The results demonstrate that SM-DST significantly and consistently outperforms state-of-the-art models across all datasets by absolute 5-18% and 3-25% in the few- and zero-shot settings, respectively. Jarana Manotumruksa, Jeff Dalton 0001, Edgar Meij, Emine Yilmaz |
WWW | 3 |
| 2021 | Contextualizing Trending Entities in News StoriesabstractTrends are those keywords, phrases, or names that are mentioned most often on social media or in news in a particular timeframe.They are an effective way for human news readers to both discover and stay focused on the most relevant information of the day. In this work, we consider trends that correspond to an entity in a knowledge graph and introduce the new and as-yet unexplored task of identifying other entities that may help explain the "why" an entity is trending. We refer to these retrieved entities as contextual entities. Some of them are more important than others in the context of the trending entity and we thus determine a ranking of entities according to how useful they are in contextualizing the trend.We propose two solutions for ranking contextual entities. The first one is fully unsupervised and based on Personalized PageRank, calculated over a trending entity-specific graph of other entities where the edges encode a notion of directional similarity based on embedded background knowledge. Our second method is based on learning to rank and combines the intuitions behind the unsupervised model with signals derived from hand-crafted features in a supervised setting. We compare our models on this novel task by using a new, purpose-built test collection created using crowdsourcing. Our methods improve over the strongest baseline in terms ofPrecision at 1 by 7% (unsupervised) and 13% (supervised). We find that the salience of a contextual entity and how coherent it is with respect to the news story are strong indicators of relevance in both unsupervised and supervised settings. Marco Ponza, Diego Ceccarelli, Paolo Ferragina, Edgar Meij, Sambhav Kothari |
WSDM | 4 |
| 2020 | Identifying Notable News Stories
Antonia Saravanou, Giorgio Stefanoni, Edgar Meij |
ECIR (2) | 3 |
| 2020 | Evaluating the Calibration of Knowledge Graph Embeddings for Trustworthy Link PredictionabstractLittle is known about the trustworthiness of predictions made by knowledge graph embedding (KGE) models.In this paper we take initial steps toward this direction by investigating the calibration of KGE models, or the extent to which they output confidence scores that reflect the expected correctness of predicted knowledge graph triples.We first conduct an evaluation under the standard closed-world assumption (CWA), in which predicted triples not already in the knowledge graph are considered false, and show that existing calibration techniques are effective for KGE under this common but narrow assumption.Next, we introduce the more realistic but challenging open-world assumption (OWA), in which unobserved predictions are not considered true or false until ground-truth labels are obtained.Here, we show that existing calibration techniques are much less effective under the OWA than the CWA, and provide explanations for this discrepancy.Finally, to motivate the utility of calibration for KGE from a practitioner's perspective, we conduct a unique case study of human-AI collaboration, showing that calibrated predictions can improve human performance in a knowledge graph completion task. Tara Safavi, Danai Koutra, Edgar Meij |
EMNLP (1) | 3 |
| 2020 | CrossBERT: A Triplet Neural Architecture for Ranking Entity PropertiesabstractTask-based Virtual Personal Assistants (VPAs) such as the Google Assistant, Alexa, and Siri are increasingly being adopted for a wide variety of tasks. These tasks are grounded in real-world entities and actions (e.g., book a hotel, organise a conference, or requesting funds). In this work we tackle the task of automatically constructing actionable knowledge graphs in response to a user query in order to support a wider variety of increasingly complex assistant tasks. We frame this as an entity property ranking task given a user query with annotated properties. We propose a new method for property ranking, CrossBERT. CrossBERT builds on the Bidirectional Encoder Representations from Transformers (BERT) and creates a new triplet network structure on cross query-property pairs that is used to rank properties. We also study the impact of using external evidence for query entities from textual entity descriptions. We perform experiments on two standard benchmark collections, the NTCIR-13 Actionable Knowledge Graph Generation (AKGG) task and Entity Property Identification (EPI) task. The results demonstrate that CrossBERT significantly outperforms the best performing runs from AKGG and EPI, as well as previous state-of-the-art BERT-based models. In particular, CrossBERT significantly improves Recall and NDCG by approximately 2-12% over the BERT models across the two used datasets. Jarana Manotumruksa, Jeff Dalton 0001, Edgar Meij, Emine Yilmaz |
SIGIR | 3 |
| 2020 | Novel Entity Discovery from Web TablesabstractWhen working with any sort of knowledge base (KB) one has to make sure it is as complete and also as up-to-date as possible. Both tasks are non-trivial as they require recall-oriented efforts to determine which entities and relationships are missing from the KB. As such they require a significant amount of labor. Tables on the Web on the other hand are abundant and have the distinct potential to assist with these tasks. In particular, we can leverage the content in such tables to discover new entities, properties, and relationships. Because web tables typically only contain raw textual content we first need to determine which cells refer to which known entities—a task we dub table-to-KB matching. This first task aims to infer table semantics by linking table cells and heading columns to elements of a KB. We propose a feature-based method and on two public test collections we demonstrate substantial improvements over the state-of-the-art in terms of precision whilst also improving recall. Then second task builds upon these linked entities and properties to not only identify novel ones in the same table but also to bootstrap their type and additional relationships. We refer to this process as novel entity discovery and, to the best of our knowledge, it is the first endeavor on mining the unlinked cells in web tables. Our method identifies not only out-of-KB (“novel”) information but also novel aliases for in-KB (“known”) entities. When evaluated using three purpose-built test collections, we find that our proposed approaches obtain a marked improvement in terms of precision over our baselines whilst keeping recall stable. Shuo Zhang 0006, Edgar Meij, Krisztian Balog, Ridho Reinanda |
WWW | 2 |
| 2019 | Special issue on knowledge graphs and semantics in text analysis and retrieval
Laura Dietz, Chenyan Xiong, Jeff Dalton 0001, Edgar Meij |
Inf. Retr. J. | 4 |
| 2018 | Related Entity Finding on Highly-heterogeneous Knowledge GraphsabstractIn this paper we study the problem of domain-specific related entity finding on highly-heterogeneous knowledge graphs where the task is to find related entities with respect to a query entity. As we are operating in the context of knowledge graphs, our solutions will need to be able to deal with heterogeneous data with multiple objects and a high number of relationship types, and be able to leverage direct and indirect connections between entities. We propose two novel graph-based related entity finding methods: one based on learning to rank and the other based on subgraph propagation in a Bayesian framework. We perform contrastive experiments using a publicly available knowledge graph and show that both our proposed models manage to outperform a strong baseline based on supervised random walks. Ridho Reinanda, Edgar Meij, Joshua Pantony, Jonathan J. Dorando |
ASONAM | 2 |
| 2018 | Utilizing Knowledge Graphs for Text-Centric Information RetrievalabstractThe past decade has witnessed the emergence of several publicly available and proprietary knowledge graphs (KGs). The depth and breadth of content in these KGs made them not only rich sources of structured knowledge by themselves, but also valuable resources for search systems. A surge of recent developments in entity linking and entity retrieval methods gave rise to a new line of research that aims at utilizing KGs for text-centric retrieval applications. This tutorial is the first to summarize and disseminate the progress in this emerging area to industry practitioners and researchers. Laura Dietz, Alexander Kotov 0001, Edgar Meij |
SIGIR | 3 |
| 2018 | The Second Workshop on Knowledge Graphs and Semantics for Text Retrieval, Analysis, and Understanding (KG4IR)abstractSemantic technologies such as controlled vocabularies, thesauri, and knowledge graphs have been used throughout the history of information retrieval for a variety of tasks. Recent advances in knowledge acquisition, alignment, and utilization have given rise to a body of new approaches for utilizing knowledge graphs in text retrieval tasks and it is therefore time to consolidate the community efforts and study how such technologies can be employed in information retrieval systems in the most effective way. It is also time to start and deepen the dialogue between researchers and practitioners in order to ensure that breakthroughs, technologies, and algorithms in this space are widely disseminated. The goal of this workshop is to bring together and grow a community of researchers and practitioners who are interested in using, aligning, and constructing knowledge graphs and similar semantic resources for information retrieval applications. Laura Dietz, Chenyan Xiong, Jeff Dalton 0001, Edgar Meij |
SIGIR | 4 |
| 2018 | Weakly-supervised Contextualization of Knowledge Graph FactsabstractKnowledge graphs (KGs) model facts about the world; they consist of nodes (entities such as companies and people) that are connected by edges (relations such as founderOf ). Facts encoded in KGs are frequently used by search applications to augment result pages. When presenting a KG fact to the user, providing other facts that are pertinent to that main fact can enrich the user experience and support exploratory information needs. \em KG fact contextualization is the task of augmenting a given KG fact with additional and useful KG facts. The task is challenging because of the large size of KGs; discovering other relevant facts even in a small neighborhood of the given fact results in an enormous amount of candidates. We introduce a neural fact contextualization method (\em NFCM ) to address the KG fact contextualization task. NFCM first generates a set of candidate facts in the neighborhood of a given fact and then ranks the candidate facts using a supervised learning to rank model. The ranking model combines features that we automatically learn from data and that represent the query-candidate facts with a set of hand-crafted features we devised or adjusted for this task. In order to obtain the annotations required to train the learning to rank model at scale, we generate training data automatically using distant supervision on a large entity-tagged text corpus. We show that ranking functions learned on this data are effective at contextualizing KG facts. Evaluation using human assessors shows that it significantly outperforms several competitive baselines. Nikos Voskarides, Edgar Meij, Ridho Reinanda, Abhinav Khaitan, Miles Osborne, Giorgio Stefanoni, Prabhanjan Kambadur, Maarten de Rijke |
SIGIR | 2 |
| 2017 | Generating Descriptions of Entity Relationships
Nikos Voskarides, Edgar Meij, Maarten de Rijke |
ECIR | 2 |
| 2017 | The First Workshop on Knowledge Graphs and Semantics for Text Retrieval and Analysis (KG4IR)abstractKnowledge graphs have been used throughout the history of information retrieval for a variety of tasks. Technological advances in knowledge acquisition and alignment technology from the last few years gave rise to a body of new approaches for utilizing knowledge graphs in text retrieval tasks. It is therefore time to consolidate the community efforts in studying how knowledge graph technology can be employed in information retrieval systems in the most effective way. It is also time to start a dialogue with researchers working on knowledge acquisition and alignment to ensure that resulting technologies and algorithms meet the demands posed by information retrieval tasks. The goal of this workshop is to bring together a community of researchers and practitioners who are interested in using, aligning, and constructing knowledge graphs and similar semantic resources for information retrieval applications. Laura Dietz, Chenyan Xiong, Edgar Meij |
SIGIR | 3 |
| 2017 | Utilizing Knowledge Graphs in Text-centricInformation RetrievalabstractThe past decade has witnessed the emergence of several publicly available and proprietary knowledge graphs (KGs). The increasing depth and breadth of content in KGs makes them not only rich sources of structured knowledge by themselves but also valuable resources for search systems. A surge of recent developments in entity linking and retrieval methods gave rise to a new line of research that aims at utilizing KGs for text-centric retrieval applications, making this an ideal time to pause and report current findings to the community, summarizing successful approaches, and soliciting new ideas. This tutorial is the first to disseminate the progress in this emerging field to researchers and practitioners. Laura Dietz, Alexander Kotov 0001, Edgar Meij |
WSDM | 3 |
| 2016 | Document Filtering for Long-tail EntitiesabstractFiltering relevant documents with respect to entities is an essential task in the context of knowledge base construction and maintenance. It entails processing a time-ordered stream of documents that might be relevant to an entity in order to select only those that contain vital information. State-of-the-art approaches to document filtering for popular entities are entity-dependent: they rely on and are also trained on the specifics of differentiating features for each specific entity. Moreover, these approaches tend to use so-called extrinsic information such as Wikipedia page views and related entities which is typically only available only for popular head entities. Entity-dependent approaches based on such signals are therefore ill-suited as filtering methods for long-tail entities. In this paper we propose a document filtering method for long-tail entities that is entity-independent and thus also generalizes to unseen or rarely seen entities. It is based on intrinsic features, i.e., features that are derived from the documents in which the entities are mentioned. We propose a set of features that capture informativeness, entity-saliency, and timeliness. In particular, we introduce features based on entity aspect similarities, relation patterns, and temporal expressions and combine these with standard features for document filtering. Experiments following the TREC KBA 2014 setup on a publicly available dataset show that our model is able to improve the filtering performance for long-tail entities over several baselines. Results of applying the model to unseen entities are promising, indicating that the model is able to learn the general characteristics of a vital document. The overall performance across all entities---i.e., not just long-tail entities---improves upon the state-of-the-art without depending on any entity-specific training data. Ridho Reinanda, Edgar Meij, Maarten de Rijke |
CIKM | 2 |
| 2016 | Dynamic Collective Entity Representations for Entity RankingabstractEntity ranking, i.e., successfully positioning a relevant entity at the top of the ranking for a given query, is inherently difficult due to the potential mismatch between the entity's description in a knowledge base, and the way people refer to the entity when searching for it. To counter this issue we propose a method for constructing dynamic collective entity representations. We collect entity descriptions from a variety of sources and combine them into a single entity representation by learning to weight the content from different sources that are associated with an entity for optimal retrieval effectiveness. Our method is able to add new descriptions in real time and learn the best representation as time evolves so as to capture the dynamics of how people search entities. Incorporating dynamic description sources into dynamic collective entity representations improves retrieval effectiveness by 7% over a state-of-the-art learning to rank baseline. Periodic retraining of the ranker enables higher ranking effectiveness for dynamic collective entity representations. David Graus, Manos Tsagkias, Wouter Weerkamp, Edgar Meij, Maarten de Rijke |
WSDM | 4 |
| 2015 | Learning to Explain Entity Relationships in Knowledge GraphsabstractNikos Voskarides, Edgar Meij, Manos Tsagkias, Maarten de Rijke, Wouter Weerkamp. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Nikos Voskarides, Edgar Meij, Manos Tsagkias, Maarten de Rijke, Wouter Weerkamp |
ACL (1) | 2 |
| 2015 | Dynamic Query Modeling for Related Content FindingabstractWhile watching television, people increasingly consume additional content related to what they are watching. We consider the task of finding video content related to a live television broadcast for which we leverage the textual stream of subtitles associated with the broadcast. We model this task as a Markov decision process and propose a method that uses reinforcement learning to directly optimize the retrieval effectiveness of queries generated from the stream of subtitles. Our dynamic query modeling approach significantly outperforms state-of-the-art baselines for stationary query modeling and for text-based retrieval in a television setting. In particular we find that carefully weighting terms and decaying these weights based on recency significantly improves effectiveness. Moreover, our method is highly efficient and can be used in a live television setting, i.e., in near real time. Daan Odijk, Edgar Meij, Isaac Sijaranamual, Maarten de Rijke |
SIGIR | 2 |
| 2015 | Mining, Ranking and Recommending Entity AspectsabstractEntity queries constitute a large fraction of web search queries and most of these queries are in the form of an entity mention plus some context terms that represent an intent in the context of that entity. We refer to these entity-oriented search intents as entity aspects. Recognizing entity aspects in a query can improve various search applications such as providing direct answers, diversifying search results, and recommending queries. In this paper we focus on the tasks of identifying, ranking, and recommending entity aspects, and propose an approach that mines, clusters, and ranks such aspects from query logs. We perform large-scale experiments based on users' search sessions from actual query logs to evaluate the aspect ranking and recommendation tasks. In the aspect ranking task, we aim to satisfy most users' entity queries, and evaluate this task in a query-independent fashion. We find that entropy-based methods achieve the best performance compared to maximum likelihood and language modeling approaches. In the aspect recommendation task, we recommend other aspects related to the aspect currently being queried. We propose two approaches based on semantic relatedness and aspect transitions within user sessions and find that a combined approach gives the best performance. As an additional experiment, we utilize entity aspects for actual query recommendation and find that our approach improves the effectiveness of query recommendations built on top of the query-flow graph. Ridho Reinanda, Edgar Meij, Maarten de Rijke |
SIGIR | 2 |
| 2015 | Fast and Space-Efficient Entity Linking for QueriesabstractEntity linking deals with identifying entities from a knowledge base in a given piece of text and has become a fundamental building block for web search engines, enabling numerous downstream improvements from better document ranking to enhanced search results pages. A key problem in the context of web search queries is that this process needs to run under severe time constraints as it has to be performed before any actual retrieval takes place, typically within milliseconds. Roi Blanco, Giuseppe Ottaviano, Edgar Meij |
WSDM | 3 |
| 2014 | Time-Aware Rank Aggregation for Microblog SearchabstractWe tackle the problem of searching microblog posts and frame it as a rank aggregation problem where we merge result lists generated by separate rankers so as to produce a final ranking to be returned to the user. We propose a rank aggregation method, TimeRA, that is able to infer the rank scores of documents via latent factor modeling. It is time-aware and rewards posts that are published in or near a burst of posts that are ranked highly in many of the lists being aggregated. Our experimental results show that it significantly outperforms state-of-the-art rank aggregation and time-sensitive microblog search algorithms. Shangsong Liang, Zhaochun Ren, Wouter Weerkamp, Edgar Meij, Maarten de Rijke |
CIKM | 4 |
| 2014 | Entity linking and retrieval for semantic searchabstractMore and more search engine users are expecting direct answers to their information needs, rather than links to documents. Semantic search and its recent applications enabled search engines to organize their wealth of information around entities. Entity linking and retrieval provide the building stones for organizing the web of entities. This tutorial aims to cover all facets of semantic search from a unified point of view and connect real-world applications with results from scientific publications. We provide a comprehensive overview of entity linking and retrieval in the context of semantic search and thoroughly explore techniques for query understanding, entity-based retrieval and ranking on unstructured text, structured knowledge repositories, and a mixture of these. We point out the connections between published approaches and applications, and provide hands-on examples on real-world use cases and datasets. Edgar Meij, Krisztian Balog, Daan Odijk |
WSDM | 1 |
| 2014 | Using temporal bursts for query modeling
Maria-Hendrike Peetz, Edgar Meij, Maarten de Rijke |
Inf. Retr. | 2 |
| 2013 | Entity linking and retrievalabstractThis full-day tutorial presents a comprehensive introduction to entity linking and retrieval. Part I provides a detailed overview of entity linking: identifying and disambiguating entity occurrences in unstructured text. Part II focuses on entity retrieval, by first considering scenarios where explicit representations of entities are available, and then moving to a setting where evidence needs to be collected and aggregated from multiple documents or even collections, thereby combining techniques from both entity linking and entity retrieval. Part III concludes the tutorial with an overview and hands-on comparative analysis of applications and publicly available toolkits and web services. Edgar Meij, Krisztian Balog, Daan Odijk |
SIGIR | 1 |
| 2013 | Personalized time-aware tweets summarizationabstractWe focus on the problem of selecting meaningful tweets given a user's interests; the dynamic nature of user interests, the sheer volume, and the sparseness of individual messages make this an challenging problem. Specifically, we consider the task of time-aware tweets summarization, based on a user's history and collaborative social influences from ``social circles.'' We propose a time-aware user behavior model, the Tweet Propagation Model (TPM), in which we infer dynamic probabilistic distributions over interests and topics. We then explicitly consider novelty, coverage, and diversity to arrive at an iterative optimization algorithm for selecting tweets. Experimental results validate the effectiveness of our personalized time-aware tweets summarization method based on TPM. Zhaochun Ren, Shangsong Liang, Edgar Meij, Maarten de Rijke |
SIGIR | 3 |
| 2012 | A Framework for Unsupervised Spam Detection in Social Networking Sites
Maarten Bosma, Edgar Meij, Wouter Weerkamp |
ECIR | 2 |
| 2012 | Adaptive Temporal Query Modeling
Maria-Hendrike Peetz, Edgar Meij, Maarten de Rijke, Wouter Weerkamp |
ECIR | 2 |
| 2012 | Identifying entity aspects in microblog postsabstractOnline reputation management is about monitoring and handling the public image of entities (such as companies) on the Web. An important task in this area is identifying "aspects" of the entity of interest (such as products, services, competitors, key people, etc.) given a stream of microblog posts referring to the entity. In this paper we compare different IR techniques and opinion target identification methods for automatically identifying aspects and find that (i) simple statistical methods such as TF.IDF are a strong baseline for the task, significantly outperforming opinion-oriented methods, and (ii) only considering terms tagged as nouns improves the results for all the methods analyzed. Damiano Spina, Edgar Meij, Maarten de Rijke, Andrei Oghina, Minh Thuong Bui, Mathias Breuss |
SIGIR | 2 |
| 2012 | Adding semantics to microblog postsabstractMicroblogs have become an important source of information for the purpose of marketing, intelligence, and reputation management. Streams of microblogs are of great value because of their direct and real-time nature. Determining what an individual microblog post is about, however, can be non-trivial because of creative language usage, the highly contextualized and informal nature of microblog posts, and the limited length of this form of communication. We propose a solution to the problem of determining what a microblog post is about through semantic linking: we add semantics to posts by automatically identifying concepts that are semantically related to it and generating links to the corresponding Wikipedia articles. The identified concepts can subsequently be used for, e.g., social media mining, thereby reducing the need for manual inspection and selection. Using a purpose-built test collection of tweets, we show that recently proposed approaches for semantic linking do not perform well, mainly due to the idiosyncratic nature of microblog posts. We propose a novel method based on machine learning with a set of innovative features and show that it is able to achieve significant improvements over all other methods, especially in terms of precision. Edgar Meij, Wouter Weerkamp, Maarten de Rijke |
WSDM | 1 |
| 2011 | People searching for people: analysis of a people search engine logabstractRecent years show an increasing interest in vertical search: searching within a particular type of information. Understanding what people search for in these "verticals" gives direction to research and provides pointers for the search engines themselves. In this paper we analyze the search logs of one particular vertical: people search engines. Based on an extensive analysis of the logs of a search engine geared towards finding people, we propose a classification scheme for people search at three levels: (a) queries, (b) sessions, and (c) users. For queries, we identify three types, (i) event-based high-profile queries (people that become "popular" because of an event happening), (ii) regular high-profile queries (celebrities), and (iii) low-profile queries (other, less-known people). We present experiments on automatic classification of queries. On the session level, we observe five types: (i) family sessions (users looking for relatives), (ii) event sessions (querying the main players of an event), (iii) spotting sessions (trying to "spot" different celebrities online), (iv) polymerous sessions (sessions without a clear relation between queries), and (v) repetitive sessions (query refinement and copying). Finally, for users we identify four types: (i) monitors, (ii) spotters, (iii) followers, and (iv) polymers. Wouter Weerkamp, Richard Berendsen, Bogomil Kovachev, Edgar Meij, Krisztian Balog, Maarten de Rijke |
SIGIR | 4 |
| 2011 | Result diversification based on query-specific cluster rankingabstractAbstract Result diversification is a retrieval strategy for dealing with ambiguous or multi‐faceted queries by providing documents that cover as many facets of the query as possible. We propose a result diversification framework based on query‐specific clustering and cluster ranking, in which diversification is restricted to documents belonging to clusters that potentially contain a high percentage of relevant documents. Empirical results show that the proposed framework improves the performance of several existing diversification methods. The framework also gives rise to a simple yet effective cluster‐based approach to result diversification that selects documents from different clusters to be included in a ranked list in a round robin fashion. We describe a set of experiments aimed at thoroughly analyzing the behavior of the two main components of the proposed diversification framework, ranking and selecting clusters for diversification. Both components have a crucial impact on the overall performance of our framework, but ranking clusters plays a more important role than selecting clusters. We also examine properties that clusters should have in order for our diversification framework to be effective. Most relevant documents should be contained in a small number of high‐quality clusters, while there should be no dominantly large clusters. Also, documents from these high‐quality clusters should have a diverse content. These properties are strongly correlated with the overall performance of the proposed diversification framework. Jiyin He, Edgar Meij, Maarten de Rijke |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2011 | Mapping queries to the Linking Open Data cloud: A case study using DBpedia
Edgar Meij, Marc Bron, Laura Hollink, Bouke Huurnink, Maarten de Rijke |
J. Web Semant. | 1 |
| 2010 | Supervised query modeling using wikipediaabstractWe use Wikipedia articles to semantically inform the generation of query models. To this end, we apply supervised machine learning to automatically link queries to Wikipedia articles and sample terms from the linked articles to re-estimate the query model. On a recent large web corpus, we observe substantial gains in terms of both traditional metrics and diversity measures. Edgar Meij, Maarten de Rijke |
SIGIR | 1 |
| 2010 | Conceptual language models for domain-specific retrieval
Edgar Meij, Dolf Trieschnigg, Maarten de Rijke, Wessel Kraaij |
Inf. Process. Manag. | 1 |
| 2009 | The impact of document structure on keyphrase extractionabstractKeyphrases are short phrases that reflect the main topic of a document. Because manually annotating documents with keyphrases is a time-consuming process, several automatic approaches have been developed. Typically, candidate phrases are extracted using features such as position or frequency in the document text. Document structure may contain useful information about which parts or phrases of a document are important, but has rarely been considered as a source of information for keyphrase extraction. Katja Hofmann, Manos Tsagkias, Edgar Meij, Maarten de Rijke |
CIKM | 3 |
| 2009 | A query model based on normalized log-likelihoodabstractLeveraging information from relevance assessments has been proposed as an effective means for improving retrieval. We introduce a novel language modeling method which uses information from each assessed document and their aggregate. While most previous approaches focus either on features of the entire set or on features of the individual relevant documents, our model exploits features of both the documents and the set as a whole. When evaluated, we show that our model is able to significantly improve over state-of-art feedback methods. Edgar Meij, Wouter Weerkamp, Maarten de Rijke |
CIKM | 1 |
| 2009 | Learning Semantic Query Suggestions
Edgar Meij, Marc Bron, Laura Hollink, Bouke Huurnink, Maarten de Rijke |
ISWC | 1 |
| 2009 | Investigating the Semantic Gap through Query Log Analysis
Peter Mika, Edgar Meij, Hugo Zaragoza |
ISWC | 2 |
| 2009 | An evaluation of entity and frequency based query completion methodsabstractWe present a semantic approach to suggesting query completions which leverages entity and type information. When compared to a frequency-based approach, we show that such information mostly helps rare queries. Edgar Meij, Peter Mika, Hugo Zaragoza |
SIGIR | 1 |
| 2009 | Structuring and extracting knowledge for the support of hypothesis generation in molecular biologyabstractBACKGROUND: Hypothesis generation in molecular and cellular biology is an empirical process in which knowledge derived from prior experiments is distilled into a comprehensible model. The requirement of automated support is exemplified by the difficulty of considering all relevant facts that are contained in the millions of documents available from PubMed. Semantic Web provides tools for sharing prior knowledge, while information retrieval and information extraction techniques enable its extraction from literature. Their combination makes prior knowledge available for computational analysis and inference. While some tools provide complete solutions that limit the control over the modeling and extraction processes, we seek a methodology that supports control by the experimenter over these critical processes. RESULTS: We describe progress towards automated support for the generation of biomolecular hypotheses. Semantic Web technologies are used to structure and store knowledge, while a workflow extracts knowledge from text. We designed minimal proto-ontologies in OWL for capturing different aspects of a text mining experiment: the biological hypothesis, text and documents, text mining, and workflow provenance. The models fit a methodology that allows focus on the requirements of a single experiment while supporting reuse and posterior analysis of extracted knowledge from multiple experiments. Our workflow is composed of services from the 'Adaptive Information Disclosure Application' (AIDA) toolkit as well as a few others. The output is a semantic model with putative biological relations, with each relation linked to the corresponding evidence. CONCLUSION: We demonstrated a 'do-it-yourself' approach for structuring and extracting knowledge in the context of experimental research on biomolecular mechanisms. The methodology can be used to bootstrap the construction of semantically rich biological models using the results of knowledge extraction processes. Models specific to particular experiments can be constructed that, in turn, link with other semantic models, creating a web of knowledge that spans experiments. Mapping mechanisms can link to other knowledge resources such as OBO ontologies or SKOS vocabularies. AIDA Web Services can be used to design personalized knowledge extraction procedures. In our example experiment, we found three proteins (NF-Kappa B, p21, and Bax) potentially playing a role in the interplay between nutrients and epigenetic gene regulation. Marco Roos, M. Scott Marshall, Andrew P. Gibson, Martijn J. Schuemie, Edgar Meij, Sophia Katrenko, Willem Robert van Hage, Konstantinos Krommydas, Pieter W. Adriaans |
BMC Bioinform. | 5 |
| 2008 | Enabling Data Transport between Web Services through alternative protocols and StreamingabstractA significant challenge for Web services is to find reliable and efficient methods to transfer large data between them. This paper describes the problem of scalable data transport between Web services, and proposes a solution: the development of a modular server/client library that uses SOAP as a control channel while the actual data transport is accomplished by various protocol implementation. Apart from file transport, the proposed approach offers the facility of direct data streaming between Web services, an approach that could benefit workflow execution time. Spiros Koulouzis, Edgar Meij, M. Scott Marshall, Adam Belloum |
eScience | 2 |
| 2008 | Towards a combined model for search and navigation of annotated documentsabstractNo abstract available. Edgar Meij |
SIGIR | 1 |
| 2008 | Parsimonious concept modelingabstractNo abstract available. Edgar Meij, Dolf Trieschnigg, Maarten de Rijke, Wessel Kraaij |
SIGIR | 1 |
| 2008 | Parsimonious relevance modelsabstractWe describe a method for applying parsimonious language models to re-estimate the term probabilities assigned by relevance models. We apply our method to six topic sets from test collections in five different genres. Our parsimonious relevance models (i) improve retrieval effectiveness in terms of MAP on all collections, (ii) significantly outperform their non-parsimonious counterparts on most measures, and (iii) have a precision enhancing effect, unlike other blind relevance feedback methods. Edgar Meij, Wouter Weerkamp, Krisztian Balog, Maarten de Rijke |
SIGIR | 1 |
| 2008 | Measuring concept relatedness using language modelsabstractOver the years, the notion of concept relatedness has attracted considerable attention. A variety of approaches, based on ontology structure, information content, association, or context have been proposed to indicate the relatedness of abstract ideas. We propose a method based on the cross entropy reduction between language models of concepts which are estimated based on document-concept assignments. The approach shows improved or competitive results compared to state-of-the-art methods on two test sets in the biomedical domain. Dolf Trieschnigg, Edgar Meij, Maarten de Rijke, Wessel Kraaij |
SIGIR | 2 |