EDBT 2026 Demo / reviewers in the wild / expert
Hugo Zaragoza
dblp:11/4382
· DBLP profile ↗
36ranked-venue papers
6as first author
0since 2021 · last 2016
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 27 · 4 first-authorArtificial intelligence and machine learning · 19 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 2Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
13 papers |
Information retrieval · 84% Indexing and storage engines · 5% Knowledge graphs · 5% | |
| Artificial intelligence
2 papers |
Learning theory · 65% Question answering and dialogue systems · 35% |
Topics — the 30 heaviest of 37, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval › search engines
search engine caching |
0.2 | 2 | 2010 | Caching search engine results over incremental indices · WWW 2010 Caching search engine results over incremental indices · SIGIR 2010 |
Information retrieval
retrieval models |
0.2 | 3 | 2010 | Ad-hoc object retrieval in the web of data · WWW 2010 Parsimonious language models for information retrieval · SIGIR 2004 Learning to Rank Answers on Large Online QA Collections · ACL 2008 |
Information retrieval
ranking |
0.2 | 2 | 2010 | Early exit optimizations for additive machine learned ranking systems · WSDM 2010 Relevance weighting for query independent evidence · SIGIR 2005 |
Indexing and storage engines › index maintenance
incremental indexing |
0.1 | 2 | 2010 | Caching search engine results over incremental indices · WWW 2010 Caching search engine results over incremental indices · SIGIR 2010 |
Information retrieval › retrieval models
ad-hoc retrieval |
0.1 | 2 | 2010 | Ad-hoc object retrieval in the web of data · WWW 2010 Bayesian extension to the language model for ad hoc information retrieval · SIGIR 2003 |
Information retrieval › search engines › semantic search › entity retrieval
ad-hoc object retrieval |
0.1 | 1 | 2010 | Ad-hoc object retrieval in the web of data · WWW 2010 |
Database system architecture and tuning
cache invalidation |
0.1 | 1 | 2010 | Caching search engine results over incremental indices · SIGIR 2010 |
Information retrieval › search engines › semantic search › entity retrieval
entity ranking |
0.1 | 1 | 2010 | Finding support sentences for entities · SIGIR 2010 |
Information retrieval › search engines › semantic search
entity retrieval |
0.1 | 1 | 2010 | Entity summarization of news articles · SIGIR 2010 |
Knowledge graphs › knowledge graph analytics
entity summarization |
0.1 | 1 | 2010 | Entity summarization of news articles · SIGIR 2010 |
Information retrieval › ranking
learning to rank |
0.1 | 1 | 2010 | Early exit optimizations for additive machine learned ranking systems · WSDM 2010 |
Information retrieval
search engines |
0.1 | 1 | 2010 | Early exit optimizations for additive machine learned ranking systems · WSDM 2010 |
Information retrieval › search engines
search engine architecture |
0.1 | 1 | 2010 | Caching search engine results over incremental indices · WWW 2010 |
Information retrieval › search engines
semantic search |
0.1 | 1 | 2010 | Ad-hoc object retrieval in the web of data · WWW 2010 |
Information retrieval › query suggestion
query auto-completion |
0.1 | 1 | 2009 | An evaluation of entity and frequency based query completion methods · SIGIR 2009 |
Information retrieval › retrieval models
language model |
0.1 | 2 | 2004 | Parsimonious language models for information retrieval · SIGIR 2004 Bayesian extension to the language model for ad hoc information retrieval · SIGIR 2003 |
Information retrieval › ranking › graph-based ranking
link-based ranking |
0.1 | 2 | 2007 | Hits on the web: how does it compare? · SIGIR 2007 Relevance weighting for query independent evidence · SIGIR 2005 |
Natural language and speech › Question answering and dialogue systems › community question answering
answer ranking |
0.1 | 1 | 2008 | Learning to Rank Answers on Large Online QA Collections · ACL 2008 |
Machine learning › Learning theory › ranking
learning to rank |
0.1 | 1 | 2008 | Learning to Rank Answers on Large Online QA Collections · ACL 2008 |
Information retrieval › query understanding
query classification |
0.1 | 1 | 2008 | Inferring the most important types of a query: a semantic approach · SIGIR 2008 |
Information retrieval
query understanding |
0.1 | 1 | 2008 | Inferring the most important types of a query: a semantic approach · SIGIR 2008 |
Information retrieval › web search › link analysis
HITS algorithm |
0.1 | 1 | 2007 | Hits on the web: how does it compare? · SIGIR 2007 |
Information retrieval
web search |
0.1 | 1 | 2007 | Hits on the web: how does it compare? · SIGIR 2007 |
Machine learning and data management
feature transformation |
0.1 | 1 | 2005 | Relevance weighting for query independent evidence · SIGIR 2005 |
Information retrieval › retrieval models › term weighting
relevance weighting |
0.1 | 1 | 2005 | Relevance weighting for query independent evidence · SIGIR 2005 |
Information retrieval › retrieval models › language model
parsimonious language model |
0.0 | 1 | 2004 | Parsimonious language models for information retrieval · SIGIR 2004 |
Machine learning › Learning theory › generalization bounds
margin theory |
0.0 | 1 | 2002 | The Perceptron Algorithm with Uneven Margins · ICML 2002 |
Machine learning › Learning theory › online learning
perceptron |
0.0 | 1 | 2002 | The Perceptron Algorithm with Uneven Margins · ICML 2002 |
Knowledge graphs › semantic web
linked data |
0.0 | 1 | 2010 | Ad-hoc object retrieval in the web of data · WWW 2010 |
Information retrieval › document retrieval › domain-specific retrieval
news retrieval |
0.0 | 1 | 2010 | Entity summarization of news articles · SIGIR 2010 |
Methods — techniques the papers use, named apart from their topics
learning to rank · 0.2formal modeling · 0.1feature engineering · 0.1early exit optimization · 0.1bag-of-words models · 0.1semantic approach · 0.1mean reciprocal rank · 0.1mean average precision · 0.1link analysis · 0.1BM25F · 0.1uneven margins · 0.0perceptron algorithm · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2016 | The OnForumS corpus from the Shared Task on Online Forum Summarisation at MultiLing 2015
Mijail A. Kabadjov, Udo Kruschwitz, Massimo Poesio, Josef Steinberger, Marc Poch, Hugo Zaragoza |
LREC | 6 |
| 2012 | Measuring website similarity using an entity-aware click graphabstractQuery logs record the actual usage of search systems and their analysis has proven critical to improving search engine functionality. Yet, despite the deluge of information, query log analysis often suffers from the sparsity of the query space. Based on the observation that most queries pivot around a single entity that represents the main focus of the user's need, we propose a new model for query log data called the entity-aware click graph. In this representation, we decompose queries into entities and modifiers, and measure their association with clicked pages. We demonstrate the benefits of this approach on the crucial task of understanding which websites fulfill similar user needs, showing that using this representation we can achieve a higher precision than other query log-based approaches. Pablo N. Mendes, Peter Mika, Hugo Zaragoza, Roi Blanco |
CIKM | 3 |
| 2011 | Learning to Rank Answers to Non-Factoid Questions from Web CollectionsabstractThis work investigates the use of linguistically motivated features to improve search, in particular for ranking answers to non-factoid questions. We show that it is possible to exploit existing large collections of question–answer pairs (from online social Question Answering sites) to extract such features and train ranking models which combine them effectively. We investigate a wide range of feature types, some exploiting natural language processing such as coarse word sense disambiguation, named-entity identification, syntactic parsing, and semantic role labeling. Our experiments demonstrate that linguistic features, in combination, yield considerable improvements in accuracy. Depending on the system settings we measure relative improvements of 14% to 21% in Mean Reciprocal Rank and Precision@1, providing one of the most compelling evidence to date that complex linguistic features such as word senses and semantic roles can have a significant impact on large-scale information retrieval tasks. Mihai Surdeanu, Massimiliano Ciaramita, Hugo Zaragoza |
Comput. Linguistics | 3 |
| 2010 | TAER: time-aware entity retrieval-exploiting the past to find relevant entities in news articlesabstractRetrieving entities instead of just documents has become an important task for search engines. In this paper we study entity retrieval for news applications, and in particular the importance of the news trail history (i.e., past related articles) in determining the relevant entities in current articles. This is an important problem in applications that display retrieved entities to the user, together with the news article. Gianluca Demartini, Malik Muhammad Saad Missen, Roi Blanco, Hugo Zaragoza |
CIKM | 4 |
| 2010 | Web search solved?: all result rankings the same?abstractThe objective of this work is to derive quantitative statements about what fraction of web search queries issued to the state-of-the-art commercial search engines lead to excellent results or, on the contrary, poor results. To be able to make such statements in an automated way, we propose a new measure that is based on lower and upper bound analysis over the standard relevance measures. Moreover, we extend this measure to carry out comparisons between competing search engines by introducing the concept of disruptive sets, which we use to estimate the degree to which a search engine solves queries that are not solved by its competitors. We report empirical results on a large editorial evaluation of the three largest search engines in the US market. Hugo Zaragoza, Berkant Barla Cambazoglu, Ricardo Baeza-Yates |
CIKM | 1 |
| 2010 | Active Learning for Building a Corpus of Questions for Parsing
Jordi Atserias Batalla, Giuseppe Attardi, Maria Simi, Hugo Zaragoza |
LREC | 4 |
| 2010 | Caching search engine results over incremental indicesabstractA Web search engine must update its index periodically to incorporate changes to the Web. We argue in this paper that index updates fundamentally impact the design of search engine result caches, a performance-critical component of modern search engines. Index updates lead to the problem of cache invalidation: invalidating cached entries of queries whose results have changed. Naive approaches, such as flushing the entire cache upon every index update, lead to poor performance and in fact, render caching futile when the frequency of updates is high. Solving the invalidation problem efficiently corresponds to predicting accurately which queries will produce different results if re-evaluated, given the actual changes to the index. Roi Blanco, Edward Bortnikov, Flavio Paiva Junqueira, Ronny Lempel, Luca Telloli, Hugo Zaragoza |
SIGIR | 6 |
| 2010 | Finding support sentences for entitiesabstractWe study the problem of finding sentences that explain the relationship between a named entity and an ad-hoc query, which we refer to as entity support sentences. Thisisanimportant sub-problem of entity ranking which, to the best of our knowledge, has not been addressed before. In this paper we give the first formalization of the problem, how it can be evaluated, and present a full evaluation dataset. We propose several methods to rank these sentences, namely retrievalbased, entity-ranking based and position-based. We found that traditional bag-of-words models perform relatively well when there is a match between an entity and a query in a given sentence, but they fail to find a support sentence for a substantial portion of entities. This can be improved by incorporating small windows of context sentences and ranking them appropriately. Roi Blanco, Hugo Zaragoza |
SIGIR | 2 |
| 2010 | Entity summarization of news articlesabstractinc.com In this paper we study the problem of entity retrieval for news applications and the importance of the news trail his-tory (i.e. past related articles) to determine the relevant entities in current articles. We construct a novel entity-labeled corpus with temporal information out of the TREC 2004 Novelty collection. We develop and evaluate several features, and show that an article’s history can be exploited to improve its summarization. Gianluca Demartini, Malik Muhammad Saad Missen, Roi Blanco, Hugo Zaragoza |
SIGIR | 4 |
| 2010 | Early exit optimizations for additive machine learned ranking systemsabstractSome commercial web search engines rely on sophisticated machine learning systems for ranking web documents. Due to very large collection sizes and tight constraints on query response times, online efficiency of these learning systems forms a bottleneck. An important problem in such systems is to speedup the ranking process without sacrificing much from the quality of results. In this paper, we propose optimization strategies that allow short-circuiting score computations in additive learning systems. The strategies are evaluated over a state-of-the-art machine learning system and a large, real-life query log, obtained from Yahoo!. By the proposed strategies, we are able to speedup the score computations by more than four times with almost no loss in result quality. Berkant Barla Cambazoglu, Hugo Zaragoza, Olivier Chapelle, Ciya Liao, Zhaohui Zheng 0001, Jon Degenhardt |
WSDM | 2 |
| 2010 | Caching search engine results over incremental indicesabstractA Web search engine must update its index periodically to incorporate changes to the Web, and we argue in this work that index updates fundamentally impact the design of search engine result caches. Index updates lead to the problem of cache invalidation: invalidating cached entries of queries whose results have changed. To enable efficient invalidation of cached results, we propose a framework for developing invalidation predictors and some concrete predictors. Evaluation using Wikipedia documents and a query log from Yahoo! shows that selective invalidation of cached search results can lower the number of query re-evaluations by as much as 30% compared to a baseline time-to-live scheme, while returning results of similar freshness. Roi Blanco, Edward Bortnikov, Flavio Paiva Junqueira, Ronny Lempel, Luca Telloli, Hugo Zaragoza |
WWW | 6 |
| 2010 | Ad-hoc object retrieval in the web of dataabstractSemantic Search refers to a loose set of concepts, challenges and techniques having to do with harnessing the information of the growing Web of Data (WoD) for Web search. Here we propose a formal model of one specific semantic search task: ad-hoc object retrieval. We show that this task provides a solid framework to study some of the semantic search problems currently tackled by commercial Web search engines. We connect this task to the traditional ad-hoc document retrieval and discuss appropriate evaluation metrics. Finally, we carry out a realistic evaluation of this task in the context of a Web search application. Jeffrey Pound, Peter Mika, Hugo Zaragoza |
WWW | 3 |
| 2010 | Structure of morphologically expanded queries: A genetic algorithm approach
Lourdes Araujo, Hugo Zaragoza, José R. Pérez-Agüera, Joaquín Pérez-Iglesias |
Data Knowl. Eng. | 2 |
| 2010 | Introduction
Omar Alonso, Hugo Zaragoza |
Inf. Process. Manag. | 2 |
| 2009 | Company-Oriented Extractive Summarization of Financial News
Katja Filippova, Mihai Surdeanu, Massimiliano Ciaramita, Hugo Zaragoza |
EACL | 4 |
| 2009 | Investigating the Semantic Gap through Query Log Analysis
Peter Mika, Edgar Meij, Hugo Zaragoza |
ISWC | 3 |
| 2009 | An evaluation of entity and frequency based query completion methodsabstractWe present a semantic approach to suggesting query completions which leverages entity and type information. When compared to a frequency-based approach, we show that such information mostly helps rare queries. Edgar Meij, Peter Mika, Hugo Zaragoza |
SIGIR | 3 |
| 2008 | Learning to Rank Answers on Large Online QA Collections
Mihai Surdeanu, Massimiliano Ciaramita, Hugo Zaragoza |
ACL | 3 |
| 2008 | Exploiting Semantic Annotations in Information Retrieval
Omar Alonso, Hugo Zaragoza |
ECIR | 2 |
| 2008 | Semantically Annotated Snapshot of the English Wikipedia
Jordi Atserias Batalla, Hugo Zaragoza, Massimiliano Ciaramita, Giuseppe Attardi |
LREC | 2 |
| 2008 | Towards Semantic Search
Ricardo Baeza-Yates, Massimiliano Ciaramita, Peter Mika, Hugo Zaragoza |
NLDB | 4 |
| 2008 | Exploiting Morphological Query Structure Using Genetic Optimisation
José R. Pérez-Agüera, Hugo Zaragoza, Lourdes Araujo |
NLDB | 2 |
| 2008 | Inferring the most important types of a query: a semantic approachabstractIn this paper we present a technique for ranking the most important types or categories for a given query. Rather than trying to find the category of the query, known as query categorization, our approach seeks to find the most important types related to the query results. Not necessarily the query category falls into this ranking of types and therefore our approach can be complementary. David Vallet, Hugo Zaragoza |
SIGIR | 2 |
| 2007 | Predictive user click models based on click-through historyabstractWeb search engines consistently collect information about users interaction with the system: they record the query they issued, the URL of presented and selected documents along with their ranking. This information is very valuable: It is a poll over millions of users on the most various topics and it has been used in many ways to mine users interests and preferences. Query logs have the potential to partially alleviate the search engines from thousand of searches by providing a way to predict answers for a subset of queries and users without knowing the content of a document. Even if the predicted result is at rank one, this analysis might be of interest: If there is enough confidence on a user's click, we might redirect the user directly to the page whose link would be clicked. In this paper, we present three different models for predicting user clicks, ranging from most specific ones (using only past user history for the query) to very general ones (aggregating data over all users for a given query). The former model has a very high precision at low recall values, while the latter can achieve high recalls. We show that it is possible to combine the different models to predict with high accuracy (over 90%) a high subset of query sessions (24% of all the sessions). Benjamin Piwowarski, Hugo Zaragoza |
CIKM | 2 |
| 2007 | Ranking very many typed entities on wikipediaabstractWe discuss the problem of ranking very many entities of different types. In particular we deal with a heterogeneous set of types, some being very generic and some very specific. We discuss two approaches for this problem: i) exploiting the entity containment graph and ii) using a Web search engine to compute entity relevance. We evaluate these approaches on the real task of ranking Wikipedia entities typed with a state-of-the-art named-entity tagger. Results show that both approaches can greatly increase the performance of methods based only on passage retrieval. Hugo Zaragoza, Henning Rode, Peter Mika, Jordi Atserias Batalla, Massimiliano Ciaramita, Giuseppe Attardi |
CIKM | 1 |
| 2007 | Hits on the web: how does it compare?abstractThis paper describes a large-scale evaluation of the effectiveness of HITS in comparison with other link-based ranking algorithms, when used in combination with a state-of-the-art text retrieval algorithm exploiting anchor text. We quantified their effectiveness using three common performance measures: the mean reciprocal rank, the mean average precision, and the normalized discounted cumulative gain measurements. The evaluation is based on two large data sets: a breadth-first search crawl of 463 million web pages containing 17.6 billion hyperlinks and referencing 2.9 billion distinct URLs; and a set of 28,043 queries sampled from a query log, each query having on average 2,383 results, about 17 of which were labeled by judges. We found that HITS outperforms PageRank, but is about as effective as web-page in-degree. The same holds true when any of the link-based features are combined with the text retrieval algorithm. Finally, we studied the relationship between query specificity and the effectiveness of selected features, and found that link-based features perform better for general queries, whereas BM25F performs better for specific queries. Marc Najork, Hugo Zaragoza, Michael J. Taylor 0001 |
SIGIR | 2 |
| 2007 | On rank-based effectiveness measures and optimization
Stephen E. Robertson, Hugo Zaragoza |
Inf. Retr. | 2 |
| 2006 | Optimisation methods for ranking functions with multiple parametersabstractOptimising the parameters of ranking functions with respect to standard IR rank-dependent cost functions has eluded satisfactory analytical treatment. We build on recent advances in alternative differentiable pairwise cost functions, and show that these techniques can be successfully applied to tuning the parameters of an existing family of IR scoring functions (BM25), in the sense that we cannot do better using sensible search heuristics that directly optimize the rank-based cost function NDCG. We also demonstrate how the size of training set affects the number of parameters we can hope to tune this way. Michael J. Taylor 0001, Hugo Zaragoza, Nick Craswell, Stephen E. Robertson, Christopher J. C. Burges |
CIKM | 2 |
| 2005 | Relevance weighting for query independent evidenceabstractA query independent feature, relating perhaps to document content, linkage or usage, can be transformed into a static, per-document relevance weight for use in ranking. The challenge is to find a good function to transform feature values into relevance scores. This paper presents FLOE, a simple density analysis method for modelling the shape of the transformation required, based on training data and without assuming independence between feature and baseline. For a new query independent feature, it addresses the questions: is it required for ranking, what sort of transformation is appropriate and, after adding it, how successful was the chosen transformation? Based on this we apply sigmoid transformations to PageRank, indegree, URL Length and ClickDistance, tested in combination with a BM25 baseline. Nick Craswell, Stephen E. Robertson, Hugo Zaragoza, Michael J. Taylor 0001 |
SIGIR | 3 |
| 2004 | Simple BM25 extension to multiple weighted fieldsabstractThis paper describes a simple way of adapting the BM25 ranking formula to deal with structured documents. In the past it has been common to compute scores for the individual fields (e.g. title and body) independently and then combine these scores (typically linearly) to arrive at a final score for the document. We highlight how this approach can lead to poor performance by breaking the carefully constructed non-linear saturation of term frequency in the BM25 function. We propose a much more intuitive alternative which weights term frequencies before the nonlinear term frequency saturation function is applied. In this scheme, a structured document with a title weight of two is mapped to an unstructured document with the title content repeated twice. This more verbose unstructured document is then ranked in the usual way. We demonstrate the advantages of this method with experiments on Reuters Vol1 and the TREC dotGov collection. Stephen E. Robertson, Hugo Zaragoza, Michael J. Taylor 0001 |
CIKM | 2 |
| 2004 | Parsimonious language models for information retrievalabstractWe systematically investigate a new approach to estimating the parameters of language models for information retrieval, called parsimonious language models. Parsimonious language models explicitly address the relation between levels of language models that are typically used for smoothing. As such, they need fewer (non-zero) parameters to describe the data. We apply parsimonious models at three stages of the retrieval process: 1) at indexing time; 2) at search time; 3) at feedback time. Experimental results show that we are able to build models that are significantly smaller than standard models, but that still perform at least as well as the standard approaches. Djoerd Hiemstra, Stephen E. Robertson, Hugo Zaragoza |
SIGIR | 3 |
| 2003 | Bayesian extension to the language model for ad hoc information retrievalabstractWe propose a Bayesian extension to the ad-hoc Language Model. Many smoothed estimators used for the multinomial query model in ad-hoc Language Models (including Laplace and Bayes-smoothing) are approximations to the Bayesian predictive distribution. In this paper we derive the full predictive distribution in a form amenable to implementation by classical IR models, and then compare it to other currently used estimators. In our experiments the proposed model outperforms Bayes-smoothing, and its combination with linear interpolation smoothing outperforms all other estimators. Hugo Zaragoza, Djoerd Hiemstra, Michael E. Tipping |
SIGIR | 1 |
| 2002 | The Perceptron Algorithm with Uneven Margins
Yaoyong Li, Hugo Zaragoza, Ralf Herbrich, John Shawe-Taylor, Jaz S. Kandola |
ICML | 2 |
| 2002 | Information Retrieval: Algorithms and Heuristics
Hugo Zaragoza |
Inf. Retr. | 1 |
| 1998 | Multiple multivariate regression and global sequence optimization: : An application to large-scale models of radiation intensity
Hugo Zaragoza, Patrick Gallinari, R. Curtelin, F. Leglaye |
Signal Process. | 1 |
| 1997 | Multiple Multivariate Regression and Global Optimization in a Large Scale Thermodynamical Application
Hugo Zaragoza, Patrick Gallinari |
ICANN | 1 |