Hugo Zaragoza

dblp:11/4382 · DBLP profile ↗
← Back
36ranked-venue papers
6as first author
0since 2021 · last 2016
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 27 · 4 first-authorArtificial intelligence and machine learning · 19 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 2Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
13 papers
Information retrieval · 84% Indexing and storage engines · 5% Knowledge graphs · 5%
Artificial intelligence
2 papers
Learning theory · 65% Question answering and dialogue systems · 35%

Topics — the 30 heaviest of 37, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval › search engines
search engine caching
0.222010
Caching search engine results over incremental indices · WWW 2010
Caching search engine results over incremental indices · SIGIR 2010
Information retrieval
retrieval models
0.232010
Ad-hoc object retrieval in the web of data · WWW 2010
Parsimonious language models for information retrieval · SIGIR 2004
Learning to Rank Answers on Large Online QA Collections · ACL 2008
Information retrieval
ranking
0.222010
Early exit optimizations for additive machine learned ranking systems · WSDM 2010
Relevance weighting for query independent evidence · SIGIR 2005
Indexing and storage engines › index maintenance
incremental indexing
0.122010
Caching search engine results over incremental indices · WWW 2010
Caching search engine results over incremental indices · SIGIR 2010
Information retrieval › retrieval models
ad-hoc retrieval
0.122010
Ad-hoc object retrieval in the web of data · WWW 2010
Bayesian extension to the language model for ad hoc information retrieval · SIGIR 2003
Information retrieval › search engines › semantic search › entity retrieval
ad-hoc object retrieval
0.112010
Ad-hoc object retrieval in the web of data · WWW 2010
Database system architecture and tuning
cache invalidation
0.112010
Caching search engine results over incremental indices · SIGIR 2010
Information retrieval › search engines › semantic search › entity retrieval
entity ranking
0.112010
Finding support sentences for entities · SIGIR 2010
Information retrieval › search engines › semantic search
entity retrieval
0.112010
Entity summarization of news articles · SIGIR 2010
Knowledge graphs › knowledge graph analytics
entity summarization
0.112010
Entity summarization of news articles · SIGIR 2010
Information retrieval › ranking
learning to rank
0.112010
Early exit optimizations for additive machine learned ranking systems · WSDM 2010
Information retrieval
search engines
0.112010
Early exit optimizations for additive machine learned ranking systems · WSDM 2010
Information retrieval › search engines
search engine architecture
0.112010
Caching search engine results over incremental indices · WWW 2010
Information retrieval › search engines
semantic search
0.112010
Ad-hoc object retrieval in the web of data · WWW 2010
Information retrieval › query suggestion
query auto-completion
0.112009
An evaluation of entity and frequency based query completion methods · SIGIR 2009
Information retrieval › retrieval models
language model
0.122004
Parsimonious language models for information retrieval · SIGIR 2004
Bayesian extension to the language model for ad hoc information retrieval · SIGIR 2003
Information retrieval › ranking › graph-based ranking
link-based ranking
0.122007
Hits on the web: how does it compare? · SIGIR 2007
Relevance weighting for query independent evidence · SIGIR 2005
Natural language and speech › Question answering and dialogue systems › community question answering
answer ranking
0.112008
Learning to Rank Answers on Large Online QA Collections · ACL 2008
Machine learning › Learning theory › ranking
learning to rank
0.112008
Learning to Rank Answers on Large Online QA Collections · ACL 2008
Information retrieval › query understanding
query classification
0.112008
Inferring the most important types of a query: a semantic approach · SIGIR 2008
Information retrieval
query understanding
0.112008
Inferring the most important types of a query: a semantic approach · SIGIR 2008
Information retrieval › web search › link analysis
HITS algorithm
0.112007
Hits on the web: how does it compare? · SIGIR 2007
Information retrieval
web search
0.112007
Hits on the web: how does it compare? · SIGIR 2007
Machine learning and data management
feature transformation
0.112005
Relevance weighting for query independent evidence · SIGIR 2005
Information retrieval › retrieval models › term weighting
relevance weighting
0.112005
Relevance weighting for query independent evidence · SIGIR 2005
Information retrieval › retrieval models › language model
parsimonious language model
0.012004
Parsimonious language models for information retrieval · SIGIR 2004
Machine learning › Learning theory › generalization bounds
margin theory
0.012002
The Perceptron Algorithm with Uneven Margins · ICML 2002
Machine learning › Learning theory › online learning
perceptron
0.012002
The Perceptron Algorithm with Uneven Margins · ICML 2002
Knowledge graphs › semantic web
linked data
0.012010
Ad-hoc object retrieval in the web of data · WWW 2010
Information retrieval › document retrieval › domain-specific retrieval
news retrieval
0.012010
Entity summarization of news articles · SIGIR 2010

Methods — techniques the papers use, named apart from their topics

learning to rank · 0.2formal modeling · 0.1feature engineering · 0.1early exit optimization · 0.1bag-of-words models · 0.1semantic approach · 0.1mean reciprocal rank · 0.1mean average precision · 0.1link analysis · 0.1BM25F · 0.1uneven margins · 0.0perceptron algorithm · 0.0
YearPublicationVenuePosition
2016 The OnForumS corpus from the Shared Task on Online Forum Summarisation at MultiLing 2015
Mijail A. Kabadjov, Udo Kruschwitz, Massimo Poesio, Josef Steinberger, Marc Poch, Hugo Zaragoza
LREC6
2012 Measuring website similarity using an entity-aware click graph
abstract
Query logs record the actual usage of search systems and their analysis has proven critical to improving search engine functionality. Yet, despite the deluge of information, query log analysis often suffers from the sparsity of the query space. Based on the observation that most queries pivot around a single entity that represents the main focus of the user's need, we propose a new model for query log data called the entity-aware click graph. In this representation, we decompose queries into entities and modifiers, and measure their association with clicked pages. We demonstrate the benefits of this approach on the crucial task of understanding which websites fulfill similar user needs, showing that using this representation we can achieve a higher precision than other query log-based approaches.
Pablo N. Mendes, Peter Mika, Hugo Zaragoza, Roi Blanco
CIKM3
2011 Learning to Rank Answers to Non-Factoid Questions from Web Collections
abstract
This work investigates the use of linguistically motivated features to improve search, in particular for ranking answers to non-factoid questions. We show that it is possible to exploit existing large collections of question–answer pairs (from online social Question Answering sites) to extract such features and train ranking models which combine them effectively. We investigate a wide range of feature types, some exploiting natural language processing such as coarse word sense disambiguation, named-entity identification, syntactic parsing, and semantic role labeling. Our experiments demonstrate that linguistic features, in combination, yield considerable improvements in accuracy. Depending on the system settings we measure relative improvements of 14% to 21% in Mean Reciprocal Rank and Precision@1, providing one of the most compelling evidence to date that complex linguistic features such as word senses and semantic roles can have a significant impact on large-scale information retrieval tasks.
Mihai Surdeanu, Massimiliano Ciaramita, Hugo Zaragoza
Comput. Linguistics3
2010 TAER: time-aware entity retrieval-exploiting the past to find relevant entities in news articles
abstract
Retrieving entities instead of just documents has become an important task for search engines. In this paper we study entity retrieval for news applications, and in particular the importance of the news trail history (i.e., past related articles) in determining the relevant entities in current articles. This is an important problem in applications that display retrieved entities to the user, together with the news article.
Gianluca Demartini, Malik Muhammad Saad Missen, Roi Blanco, Hugo Zaragoza
CIKM4
2010 Web search solved?: all result rankings the same?
abstract
The objective of this work is to derive quantitative statements about what fraction of web search queries issued to the state-of-the-art commercial search engines lead to excellent results or, on the contrary, poor results. To be able to make such statements in an automated way, we propose a new measure that is based on lower and upper bound analysis over the standard relevance measures. Moreover, we extend this measure to carry out comparisons between competing search engines by introducing the concept of disruptive sets, which we use to estimate the degree to which a search engine solves queries that are not solved by its competitors. We report empirical results on a large editorial evaluation of the three largest search engines in the US market.
Hugo Zaragoza, Berkant Barla Cambazoglu, Ricardo Baeza-Yates
CIKM1
2010 Active Learning for Building a Corpus of Questions for Parsing
Jordi Atserias Batalla, Giuseppe Attardi, Maria Simi, Hugo Zaragoza
LREC4
2010 Caching search engine results over incremental indices
abstract
A Web search engine must update its index periodically to incorporate changes to the Web. We argue in this paper that index updates fundamentally impact the design of search engine result caches, a performance-critical component of modern search engines. Index updates lead to the problem of cache invalidation: invalidating cached entries of queries whose results have changed. Naive approaches, such as flushing the entire cache upon every index update, lead to poor performance and in fact, render caching futile when the frequency of updates is high. Solving the invalidation problem efficiently corresponds to predicting accurately which queries will produce different results if re-evaluated, given the actual changes to the index.
Roi Blanco, Edward Bortnikov, Flavio Paiva Junqueira, Ronny Lempel, Luca Telloli, Hugo Zaragoza
SIGIR6
2010 Finding support sentences for entities
abstract
We study the problem of finding sentences that explain the relationship between a named entity and an ad-hoc query, which we refer to as entity support sentences. Thisisanimportant sub-problem of entity ranking which, to the best of our knowledge, has not been addressed before. In this paper we give the first formalization of the problem, how it can be evaluated, and present a full evaluation dataset. We propose several methods to rank these sentences, namely retrievalbased, entity-ranking based and position-based. We found that traditional bag-of-words models perform relatively well when there is a match between an entity and a query in a given sentence, but they fail to find a support sentence for a substantial portion of entities. This can be improved by incorporating small windows of context sentences and ranking them appropriately.
Roi Blanco, Hugo Zaragoza
SIGIR2
2010 Entity summarization of news articles
abstract
inc.com In this paper we study the problem of entity retrieval for news applications and the importance of the news trail his-tory (i.e. past related articles) to determine the relevant entities in current articles. We construct a novel entity-labeled corpus with temporal information out of the TREC 2004 Novelty collection. We develop and evaluate several features, and show that an article’s history can be exploited to improve its summarization.
Gianluca Demartini, Malik Muhammad Saad Missen, Roi Blanco, Hugo Zaragoza
SIGIR4
2010 Early exit optimizations for additive machine learned ranking systems
abstract
Some commercial web search engines rely on sophisticated machine learning systems for ranking web documents. Due to very large collection sizes and tight constraints on query response times, online efficiency of these learning systems forms a bottleneck. An important problem in such systems is to speedup the ranking process without sacrificing much from the quality of results. In this paper, we propose optimization strategies that allow short-circuiting score computations in additive learning systems. The strategies are evaluated over a state-of-the-art machine learning system and a large, real-life query log, obtained from Yahoo!. By the proposed strategies, we are able to speedup the score computations by more than four times with almost no loss in result quality.
Berkant Barla Cambazoglu, Hugo Zaragoza, Olivier Chapelle, Ciya Liao, Zhaohui Zheng 0001, Jon Degenhardt
WSDM2
2010 Caching search engine results over incremental indices
abstract
A Web search engine must update its index periodically to incorporate changes to the Web, and we argue in this work that index updates fundamentally impact the design of search engine result caches. Index updates lead to the problem of cache invalidation: invalidating cached entries of queries whose results have changed. To enable efficient invalidation of cached results, we propose a framework for developing invalidation predictors and some concrete predictors. Evaluation using Wikipedia documents and a query log from Yahoo! shows that selective invalidation of cached search results can lower the number of query re-evaluations by as much as 30% compared to a baseline time-to-live scheme, while returning results of similar freshness.
Roi Blanco, Edward Bortnikov, Flavio Paiva Junqueira, Ronny Lempel, Luca Telloli, Hugo Zaragoza
WWW6
2010 Ad-hoc object retrieval in the web of data
abstract
Semantic Search refers to a loose set of concepts, challenges and techniques having to do with harnessing the information of the growing Web of Data (WoD) for Web search. Here we propose a formal model of one specific semantic search task: ad-hoc object retrieval. We show that this task provides a solid framework to study some of the semantic search problems currently tackled by commercial Web search engines. We connect this task to the traditional ad-hoc document retrieval and discuss appropriate evaluation metrics. Finally, we carry out a realistic evaluation of this task in the context of a Web search application.
Jeffrey Pound, Peter Mika, Hugo Zaragoza
WWW3
2010 Structure of morphologically expanded queries: A genetic algorithm approach
Lourdes Araujo, Hugo Zaragoza, José R. Pérez-Agüera, Joaquín Pérez-Iglesias
Data Knowl. Eng.2
2010 Introduction
Omar Alonso, Hugo Zaragoza
Inf. Process. Manag.2
2009 Company-Oriented Extractive Summarization of Financial News
Katja Filippova, Mihai Surdeanu, Massimiliano Ciaramita, Hugo Zaragoza
EACL4
2009 Investigating the Semantic Gap through Query Log Analysis
Peter Mika, Edgar Meij, Hugo Zaragoza
ISWC3
2009 An evaluation of entity and frequency based query completion methods
abstract
We present a semantic approach to suggesting query completions which leverages entity and type information. When compared to a frequency-based approach, we show that such information mostly helps rare queries.
Edgar Meij, Peter Mika, Hugo Zaragoza
SIGIR3
2008 Learning to Rank Answers on Large Online QA Collections
Mihai Surdeanu, Massimiliano Ciaramita, Hugo Zaragoza
ACL3
2008 Exploiting Semantic Annotations in Information Retrieval
Omar Alonso, Hugo Zaragoza
ECIR2
2008 Semantically Annotated Snapshot of the English Wikipedia
Jordi Atserias Batalla, Hugo Zaragoza, Massimiliano Ciaramita, Giuseppe Attardi
LREC2
2008 Towards Semantic Search
Ricardo Baeza-Yates, Massimiliano Ciaramita, Peter Mika, Hugo Zaragoza
NLDB4
2008 Exploiting Morphological Query Structure Using Genetic Optimisation
José R. Pérez-Agüera, Hugo Zaragoza, Lourdes Araujo
NLDB2
2008 Inferring the most important types of a query: a semantic approach
abstract
In this paper we present a technique for ranking the most important types or categories for a given query. Rather than trying to find the category of the query, known as query categorization, our approach seeks to find the most important types related to the query results. Not necessarily the query category falls into this ranking of types and therefore our approach can be complementary.
David Vallet, Hugo Zaragoza
SIGIR2
2007 Predictive user click models based on click-through history
abstract
Web search engines consistently collect information about users interaction with the system: they record the query they issued, the URL of presented and selected documents along with their ranking. This information is very valuable: It is a poll over millions of users on the most various topics and it has been used in many ways to mine users interests and preferences. Query logs have the potential to partially alleviate the search engines from thousand of searches by providing a way to predict answers for a subset of queries and users without knowing the content of a document. Even if the predicted result is at rank one, this analysis might be of interest: If there is enough confidence on a user's click, we might redirect the user directly to the page whose link would be clicked. In this paper, we present three different models for predicting user clicks, ranging from most specific ones (using only past user history for the query) to very general ones (aggregating data over all users for a given query). The former model has a very high precision at low recall values, while the latter can achieve high recalls. We show that it is possible to combine the different models to predict with high accuracy (over 90%) a high subset of query sessions (24% of all the sessions).
Benjamin Piwowarski, Hugo Zaragoza
CIKM2
2007 Ranking very many typed entities on wikipedia
abstract
We discuss the problem of ranking very many entities of different types. In particular we deal with a heterogeneous set of types, some being very generic and some very specific. We discuss two approaches for this problem: i) exploiting the entity containment graph and ii) using a Web search engine to compute entity relevance. We evaluate these approaches on the real task of ranking Wikipedia entities typed with a state-of-the-art named-entity tagger. Results show that both approaches can greatly increase the performance of methods based only on passage retrieval.
Hugo Zaragoza, Henning Rode, Peter Mika, Jordi Atserias Batalla, Massimiliano Ciaramita, Giuseppe Attardi
CIKM1
2007 Hits on the web: how does it compare?
abstract
This paper describes a large-scale evaluation of the effectiveness of HITS in comparison with other link-based ranking algorithms, when used in combination with a state-of-the-art text retrieval algorithm exploiting anchor text. We quantified their effectiveness using three common performance measures: the mean reciprocal rank, the mean average precision, and the normalized discounted cumulative gain measurements. The evaluation is based on two large data sets: a breadth-first search crawl of 463 million web pages containing 17.6 billion hyperlinks and referencing 2.9 billion distinct URLs; and a set of 28,043 queries sampled from a query log, each query having on average 2,383 results, about 17 of which were labeled by judges. We found that HITS outperforms PageRank, but is about as effective as web-page in-degree. The same holds true when any of the link-based features are combined with the text retrieval algorithm. Finally, we studied the relationship between query specificity and the effectiveness of selected features, and found that link-based features perform better for general queries, whereas BM25F performs better for specific queries.
Marc Najork, Hugo Zaragoza, Michael J. Taylor 0001
SIGIR2
2007 On rank-based effectiveness measures and optimization
Stephen E. Robertson, Hugo Zaragoza
Inf. Retr.2
2006 Optimisation methods for ranking functions with multiple parameters
abstract
Optimising the parameters of ranking functions with respect to standard IR rank-dependent cost functions has eluded satisfactory analytical treatment. We build on recent advances in alternative differentiable pairwise cost functions, and show that these techniques can be successfully applied to tuning the parameters of an existing family of IR scoring functions (BM25), in the sense that we cannot do better using sensible search heuristics that directly optimize the rank-based cost function NDCG. We also demonstrate how the size of training set affects the number of parameters we can hope to tune this way.
Michael J. Taylor 0001, Hugo Zaragoza, Nick Craswell, Stephen E. Robertson, Christopher J. C. Burges
CIKM2
2005 Relevance weighting for query independent evidence
abstract
A query independent feature, relating perhaps to document content, linkage or usage, can be transformed into a static, per-document relevance weight for use in ranking. The challenge is to find a good function to transform feature values into relevance scores. This paper presents FLOE, a simple density analysis method for modelling the shape of the transformation required, based on training data and without assuming independence between feature and baseline. For a new query independent feature, it addresses the questions: is it required for ranking, what sort of transformation is appropriate and, after adding it, how successful was the chosen transformation? Based on this we apply sigmoid transformations to PageRank, indegree, URL Length and ClickDistance, tested in combination with a BM25 baseline.
Nick Craswell, Stephen E. Robertson, Hugo Zaragoza, Michael J. Taylor 0001
SIGIR3
2004 Simple BM25 extension to multiple weighted fields
abstract
This paper describes a simple way of adapting the BM25 ranking formula to deal with structured documents. In the past it has been common to compute scores for the individual fields (e.g. title and body) independently and then combine these scores (typically linearly) to arrive at a final score for the document. We highlight how this approach can lead to poor performance by breaking the carefully constructed non-linear saturation of term frequency in the BM25 function. We propose a much more intuitive alternative which weights term frequencies before the nonlinear term frequency saturation function is applied. In this scheme, a structured document with a title weight of two is mapped to an unstructured document with the title content repeated twice. This more verbose unstructured document is then ranked in the usual way. We demonstrate the advantages of this method with experiments on Reuters Vol1 and the TREC dotGov collection.
Stephen E. Robertson, Hugo Zaragoza, Michael J. Taylor 0001
CIKM2
2004 Parsimonious language models for information retrieval
abstract
We systematically investigate a new approach to estimating the parameters of language models for information retrieval, called parsimonious language models. Parsimonious language models explicitly address the relation between levels of language models that are typically used for smoothing. As such, they need fewer (non-zero) parameters to describe the data. We apply parsimonious models at three stages of the retrieval process: 1) at indexing time; 2) at search time; 3) at feedback time. Experimental results show that we are able to build models that are significantly smaller than standard models, but that still perform at least as well as the standard approaches.
Djoerd Hiemstra, Stephen E. Robertson, Hugo Zaragoza
SIGIR3
2003 Bayesian extension to the language model for ad hoc information retrieval
abstract
We propose a Bayesian extension to the ad-hoc Language Model. Many smoothed estimators used for the multinomial query model in ad-hoc Language Models (including Laplace and Bayes-smoothing) are approximations to the Bayesian predictive distribution. In this paper we derive the full predictive distribution in a form amenable to implementation by classical IR models, and then compare it to other currently used estimators. In our experiments the proposed model outperforms Bayes-smoothing, and its combination with linear interpolation smoothing outperforms all other estimators.
Hugo Zaragoza, Djoerd Hiemstra, Michael E. Tipping
SIGIR1
2002 The Perceptron Algorithm with Uneven Margins
Yaoyong Li, Hugo Zaragoza, Ralf Herbrich, John Shawe-Taylor, Jaz S. Kandola
ICML2
2002 Information Retrieval: Algorithms and Heuristics
Hugo Zaragoza
Inf. Retr.1
1998 Multiple multivariate regression and global sequence optimization: : An application to large-scale models of radiation intensity
Hugo Zaragoza, Patrick Gallinari, R. Curtelin, F. Leglaye
Signal Process.1
1997 Multiple Multivariate Regression and Global Optimization in a Large Scale Thermodynamical Application
Hugo Zaragoza, Patrick Gallinari
ICANN1