EDBT 2026 Demo / reviewers in the wild / expert
Roi Blanco
dblp:30/3795
· DBLP profile ↗
67ranked-venue papers
26as first author
0since 2021 · last 2019
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 64 · 25 first-authorArtificial intelligence and machine learning · 19 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 2 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
29 papers |
Information retrieval · 84% Web and social media mining · 4% Knowledge graphs · 4% | |
| Artificial intelligence
1 paper |
Information extraction and text analysis · 100% | |
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Energy-efficient computing · 39% Cloud and datacenter computing · 37% Hardware accelerators and domain-specific architectures · 19% |
Topics — the 30 heaviest of 69, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval › ranking
learning to rank |
0.8 | 3 | 2019 | Joint Optimization of Cascade Ranking Models · WSDM 2019 Efficient Cost-Aware Cascade Ranking in Multi-Stage Retrieval · SIGIR 2017 Ranking related news predictions · SIGIR 2011 |
Information retrieval › ranking › ranking algorithms
cascade ranking |
0.6 | 2 | 2018 | Query Driven Algorithm Selection in Early Stage Retrieval · WSDM 2018 Efficient Cost-Aware Cascade Ranking in Multi-Stage Retrieval · SIGIR 2017 |
Information retrieval
query processing |
0.6 | 3 | 2018 | Query Driven Algorithm Selection in Early Stage Retrieval · WSDM 2018 Online News Tracking for Ad-Hoc Queries · SIGIR 2015 Fast and Space-Efficient Entity Linking for Queries · WSDM 2015 |
Information retrieval
ranking |
0.5 | 2 | 2017 | Click Through Rate Prediction for Local Search Results · WSDM 2017 Local Ranking Problem on the BrowseGraph · SIGIR 2015 |
Information retrieval
web search |
0.5 | 3 | 2016 | Exploiting Green Energy to Reduce the Operational Costs of Multi-Center Web Search Engines · WWW 2016 Energy-price-driven query processing in multi-center web search engines · SIGIR 2011 Enhanced results for web search · SIGIR 2011 |
Information retrieval
retrieval models |
0.4 | 2 | 2017 | Click Through Rate Prediction for Local Search Results · WSDM 2017 Extending BM25 with multiple query operators · SIGIR 2012 |
Information retrieval › ranking › text ranking › document ranking
first-stage retrieval |
0.3 | 1 | 2018 | Query Driven Algorithm Selection in Early Stage Retrieval · WSDM 2018 |
Information retrieval › search engines › semantic search
entity retrieval |
0.3 | 2 | 2015 | From "Selena Gomez" to "Marlon Brando": Understanding Explorative Entity Search · WWW 2015 Entity summarization of news articles · SIGIR 2010 |
Natural language and speech › Information extraction and text analysis › named entity recognition
cross-lingual named entity recognition |
0.3 | 1 | 2017 | Lightweight Multilingual Entity Extraction and Linking · WSDM 2017 |
Natural language and speech › Information extraction and text analysis › entity linking
entity disambiguation |
0.3 | 1 | 2017 | Lightweight Multilingual Entity Extraction and Linking · WSDM 2017 |
Natural language and speech › Information extraction and text analysis
entity linking |
0.3 | 1 | 2017 | Lightweight Multilingual Entity Extraction and Linking · WSDM 2017 |
Natural language and speech › Information extraction and text analysis
named entity recognition |
0.3 | 1 | 2017 | Lightweight Multilingual Entity Extraction and Linking · WSDM 2017 |
Recommender systems
click-through rate prediction |
0.3 | 1 | 2017 | Click Through Rate Prediction for Local Search Results · WSDM 2017 |
Information retrieval › interactive information retrieval › search interaction
collaborative search |
0.3 | 1 | 2017 | Modelling Information Needs in Collaborative Search Conversations · SIGIR 2017 |
Information retrieval › interactive information retrieval › conversational information seeking
conversational search |
0.3 | 1 | 2017 | Modelling Information Needs in Collaborative Search Conversations · SIGIR 2017 |
Information retrieval › web search
local search |
0.3 | 1 | 2017 | Click Through Rate Prediction for Local Search Results · WSDM 2017 |
Information retrieval
multi-stage retrieval |
0.3 | 1 | 2017 | Efficient Cost-Aware Cascade Ranking in Multi-Stage Retrieval · SIGIR 2017 |
Information retrieval › interactive information retrieval › search tasks
re-finding |
0.3 | 1 | 2017 | Re-Finding Behaviour in Vertical Domains · ACM Trans. Inf. Syst. 2017 |
Information retrieval
search result diversification |
0.3 | 1 | 2017 | A Concise Integer Linear Programming Formulation for Implicit Search Result Diversification · WSDM 2017 |
Information retrieval
evaluation |
0.2 | 1 | 2016 | Building Test Collections for Evaluating Temporal IR · SIGIR 2016 |
Information retrieval › query suggestion
query auto-completion |
0.2 | 1 | 2016 | Term-by-Term Query Auto-Completion for Mobile Search · WSDM 2016 |
Information retrieval › evaluation › test collection
test collection construction |
0.2 | 1 | 2016 | Building Test Collections for Evaluating Temporal IR · SIGIR 2016 |
Energy-efficient computing
datacenter power management |
0.2 | 1 | 2016 | Exploiting Green Energy to Reduce the Operational Costs of Multi-Center Web Search Engines · WWW 2016 |
Query processing and optimization
ad-hoc query |
0.2 | 1 | 2015 | Online News Tracking for Ad-Hoc Queries · SIGIR 2015 |
Web and social media mining › web usage mining
clickstream analysis |
0.2 | 1 | 2015 | Local Ranking Problem on the BrowseGraph · SIGIR 2015 |
Knowledge graphs
entity linking |
0.2 | 1 | 2015 | Fast and Space-Efficient Entity Linking for Queries · WSDM 2015 |
Information retrieval › ranking
graph-based ranking |
0.2 | 1 | 2015 | Local Ranking Problem on the BrowseGraph · SIGIR 2015 |
Information retrieval
search engines |
0.2 | 1 | 2015 | From "Selena Gomez" to "Marlon Brando": Understanding Explorative Entity Search · WWW 2015 |
Information retrieval › search engines
search engine caching |
0.2 | 2 | 2010 | Caching search engine results over incremental indices · WWW 2010 Caching search engine results over incremental indices · SIGIR 2010 |
Information retrieval › text analysis › topic analysis
topic detection and tracking |
0.2 | 1 | 2015 | Online News Tracking for Ad-Hoc Queries · SIGIR 2015 |
Methods — techniques the papers use, named apart from their topics
LambdaMART · 0.7integer linear programming · 0.6feature engineering · 0.4gradient boosted trees · 0.4backpropagation · 0.4crowdsourcing · 0.4performance prediction · 0.3latency optimization · 0.3machine-learned models · 0.3gradient boosted decision trees · 0.3entity embeddings · 0.3click-log-based candidate retrieval · 0.3workload shifting · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | A Rank-biased Neural Network Model for Click ModelingabstractQuery logs contain rich feedback information from a large number of users interacting with search engines. Various click models have been developed to decode users' search behavior and to extract useful knowledge from query logs. Although the state-of-the-art neural click models have been shown to be very effective in click modeling, the input representations of queries and documents rely on either manually crafted features or on automatic methods suffering from the high-dimensionality issue. Moreover, these neural click models are still rather restrictive when coping with commonly biased user clicks. In this paper, we investigate how to effectively deploy a neural network model for decoding users' click behavior. First, we present two novel rank-biased neural network models ($RBNN$ and $RBNN^* $) for click modeling. The key idea is to deploy different weight matrices across different rank positions. Second, we introduce a new method ($QD\mymathhyphen DCCA$) for automatically learning the vector representations for both queries and documents within the same low-dimensional space, which provides high-quality inputs for $RBNN$ and $RBNN^* $. Finally, a series of experiments are conducted on two different real query logs to validate the effectiveness and efficiency of the proposed neural click models. The experiments demonstrate that: (1) The proposed models can achieve substantially improved performance over the state-of-the-art baseline on two datasets across multiple metrics. By incorporating rank-specific weight matrices, $RBNN$ and $RBNN^* $ are more capable of dealing with the position-bias problem. (2) The input representations of queries, documents and context information significantly affect the performance of neural click models. Thanks to the application of $QD\mymathhyphen DCCA$, not only $RBNN$ and $RBNN^* $ but also the baseline method exhibit enhanced performance. Furthermore, the training cost under the proposed models is greatly reduced. Hai-Tao Yu 0003, Adam Jatowt, Roi Blanco, Joemon M. Jose, Ke Zhou 0003 |
CHIIR | 3 |
| 2019 | Joint Optimization of Cascade Ranking ModelsabstractReducing excessive costs in feature acquisition and model evaluation has been a long-standing challenge in learning-to-rank systems. A cascaded ranking architecture turns ranking into a pipeline of multiple stages, and has been shown to be a powerful approach to balancing efficiency and effectiveness trade-offs in large-scale search systems. However, learning a cascade model is often complex, and usually performed stagewise independently across the entire ranking pipeline. In this work we show that learning a cascade ranking model in this manner is often suboptimal in terms of both effectiveness and efficiency. We present a new general framework for learning an end-to-end cascade of rankers using backpropagation. We show that stagewise objectives can be chained together and optimized jointly to achieve significantly better trade-offs globally. This novel approach is generalizable to not only differentiable models but also state-of-the-art tree-based algorithms such as LambdaMART and cost-efficient gradient boosted trees, and it opens up new opportunities for exploring additional efficiency-effectiveness trade-offs in large-scale search systems. Luke Gallagher, Ruey-Cheng Chen, Roi Blanco, J. Shane Culpepper |
WSDM | 3 |
| 2018 | Query Driven Algorithm Selection in Early Stage RetrievalabstractLarge scale retrieval systems often employ cascaded ranking architectures, in which an initial set of candidate documents are iteratively refined and re-ranked by increasingly sophisticated and expensive ranking models. In this paper, we propose a unified framework for predicting a range of performance-sensitive parameters based on minimizing end-to-end effectiveness loss. The framework does not require relevance judgments for training, is amenable to predicting a wide range of parameters, allows for fine tuned efficiency-effectiveness trade-offs, and can be easily deployed in large scale search systems with minimal overhead. As a proof of concept, we show that the framework can accurately predict a number of performance parameters on a query-by-query basis, allowing efficient and effective retrieval, while simultaneously minimizing the tail latency of an early-stage candidate generation system. On the 50 million document ClueWeb09B collection, and across 25,000 queries, our hybrid system can achieve superior early-stage efficiency to fixed parameter systems without loss of effectiveness, and allows more finely-grained efficiency-effectiveness trade-offs across the multiple stages of the retrieval system. Joel Mackenzie, J. Shane Culpepper, Roi Blanco, Matt Crane, Charles L. A. Clarke, Jimmy Lin |
WSDM | 3 |
| 2018 | Revisiting the cluster-based paradigm for implicit search result diversification
Hai-Tao Yu 0003, Adam Jatowt, Roi Blanco, Hideo Joho, Joemon M. Jose, Long Chen 0008, Fajie Yuan |
Inf. Process. Manag. | 3 |
| 2018 | Characterizing, predicting, and handling web search queries that match very few or no resultsabstractA non‐negligible fraction of user queries end up with very few or even no matching results in leading commercial web search engines. In this work, we provide a detailed characterization of such queries and show that search engines try to improve such queries by showing the results of related queries. Through a user study, we show that these query suggestions are usually perceived as relevant. Also, through a query log analysis, we show that the users are dissatisfied after submitting a query that match no results at least 88.5% of the time. As a first step towards solving these no‐answer queries, we devised a large number of features that can be used to identify such queries and built machine‐learning models. These models can be useful for scenarios such as the mobile‐ or meta‐search, where identifying a query that will retrieve no results at the client device (i.e., even before submitting it to the search engine) may yield gains in terms of the bandwidth usage, power consumption, and/or monetary costs. Experiments over query logs indicate that, despite the heavy skew in class sizes, our models achieve good prediction quality, with accuracy (in terms of area under the curve) up to 0.95. Erdem Sarigil, Ismail Sengör Altingövde, Roi Blanco, Berkant Barla Cambazoglu, Rifat Ozcan, Özgür Ulusoy |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2018 | Characterizing and Predicting Users' Behavior on Local Search QueriesabstractThe use of queries to find products and services that are located nearby is increasing rapidly due mainly to the ubiquity of internet access and location services provided by smartphone devices. Local search engines help users by matching queries with a predefined geographical connotation (“local queries”) against a database of local business listings. Local search differs from traditional Web search because, to correctly capture users’ click behavior, the estimation of relevance between query and candidate results must be integrated with geographical signals, such as distance. The intuition is that users prefer businesses that are physically closer to them or in a convenient area (e.g., close to their home). However, this notion of closeness depends upon other factors, like the business category, the quality of the service provided, the density of businesses in the area of interest, the hour of the day, or even the day of the week. In this work, we perform an extensive analysis of online users’ interactions with a local search engine, investigating their intent, temporal patterns, and highlighting relationships between distance-to-business and other factors, such as business reputation, Furthermore, we investigate the problem of estimating the click-through rate on local search ( LCTR ) by exploiting the combination of standard retrieval methods with a rich collection of geo-, user-, and business-dependent features. We validate our approach on a large log collected from a real-world local search service. Our evaluation shows that the non-linear combination of business and user information, geo-local and textual relevance features leads to a significant improvements over existing alternative approaches based on a combination of relevance, distance, and business reputation [1]. Fidel Cacheda, Roi Blanco, Nicola Barbieri |
ACM Trans. Web | 2 |
| 2017 | Efficient Cost-Aware Cascade Ranking in Multi-Stage RetrievalabstractComplex machine learning models are now an integral part of modern, large-scale retrieval systems. However, collection size growth continues to outpace advances in efficiency improvements in the learning models which achieve the highest effectiveness. In this paper, we re-examine the importance of tightly integrating feature costs into multi-stage learning-to-rank (LTR) IR systems. We present a novel approach to optimizing cascaded ranking models which can directly leverage a variety of different state-of-the-art LTR rankers such as LambdaMART and Gradient Boosted Decision Trees. Using our cascade model, we conclusively show that feature costs and the number of documents being re-ranked in each stage of the cascade can be balanced to maximize both efficiency and effectiveness. Finally, we also demonstrate that our cascade model can easily be deployed on commonly used collections to achieve state-of-the-art effectiveness results while only using a subset of the features required by the full model. Ruey-Cheng Chen, Luke Gallagher, Roi Blanco, J. Shane Culpepper |
SIGIR | 3 |
| 2017 | Modelling Information Needs in Collaborative Search ConversationsabstractThe increase of voice-based interaction has changed the way people seek information, making search more conversational. Development of effective conversational approaches to search requires better understanding of how people express information needs in dialogue. This paper describes the creation and examination of over 32K spoken utterances collected during 34 hours of collaborative search tasks. The contribution of this work is three-fold. First, we propose a model of conversational information needs (CINs) based on a synthesis of relevant theories in Information Seeking and Retrieval. Second, we show several behavioural patterns of CINs based on the proposed model. Third, we identify effective feature groups that may be useful for detecting CINs categories from conversations. This paper concludes with a discussion of how these findings can facilitate advance of conversational search applications. Sosuke Shiga, Hideo Joho, Roi Blanco, Johanne R. Trippas, Mark Sanderson |
SIGIR | 3 |
| 2017 | Click Through Rate Prediction for Local Search ResultsabstractWith the ubiquity of internet access and location services provided by smartphone devices, the volume of queries issued by users to find products and services that are located near them is rapidly increasing. Local search engines help users in this task by matching queries with a predefined geographical connotation ("local queries") against a database of local business listings. Fidel Cacheda, Nicola Barbieri, Roi Blanco |
WSDM | 3 |
| 2017 | Lightweight Multilingual Entity Extraction and LinkingabstractText analytics systems often rely heavily on detecting and linking entity mentions in documents to knowledge bases for downstream applications such as sentiment analysis, question answering and recommender systems. A major challenge for this task is to be able to accurately detect entities in new languages with limited labeled resources. In this paper we present an accurate and lightweight, multilingual named entity recognition (NER) and linking (NEL) system. The contributions of this paper are three-fold: 1) Lightweight named entity recognition with competitive accuracy; 2) Candidate entity retrieval that uses search click-log data and entity embeddings to achieve high precision with a low memory footprint; and 3) efficient entity disambiguation. Our system achieves state-of-the-art performance on TAC KBP 2013 multilingual data and on English AIDA CONLL data. Aasish Pappu, Roi Blanco, Yashar Mehdad, Amanda Stent, Kapil Thadani |
WSDM | 2 |
| 2017 | A Concise Integer Linear Programming Formulation for Implicit Search Result DiversificationabstractTo cope with ambiguous and/or underspecified queries, search result diversification (SRD) is a key technique that has attracted a lot of attention. This paper focuses on implicit SRD, where the possible subtopics underlying a query are unknown beforehand. We formulate implicit SRD as a process of selecting and ranking k exemplar documents that utilizes integer linear programming (ILP). Unlike the common practice of relying on approximate methods, this formulation enables us to obtain the optimal solution of the objective function. Based on four benchmark collections, our extensive empirical experiments reveal that: (1) The factors, such as different initial runs, the number of input documents, query types and the ways of computing document similarity significantly affect the performance of diversification models. Careful examinations of these factors are highly recommended in the development of implicit SRD methods. (2) The proposed method can achieve substantially improved performance over the state-of-the-art unsupervised methods for implicit SRD. Hai-Tao Yu 0003, Adam Jatowt, Roi Blanco, Hideo Joho, Joemon M. Jose, Long Chen 0008, Fajie Yuan |
WSDM | 3 |
| 2017 | An in-depth study on diversity evaluation: The importance of intrinsic diversity
Hai-Tao Yu 0003, Adam Jatowt, Roi Blanco, Hideo Joho, Joemon M. Jose |
Inf. Process. Manag. | 3 |
| 2017 | Decoding multi-click search behavior based on marginal utility
Hai-Tao Yu 0003, Adam Jatowt, Roi Blanco, Hideo Joho, Joemon M. Jose |
Inf. Retr. J. | 3 |
| 2017 | Re-Finding Behaviour in Vertical DomainsabstractRe-finding is the process of searching for information that a user has previously encountered and is a common activity carried out with information retrieval systems. In this work, we investigate re-finding in the context of vertical search, differentiating and modeling user re-finding behavior within different media and topic domains, including images, news, reference material, and movies. We distinguish the re-finding behavior in vertical domains from re-finding in a general search context and engineer features that are effective in differentiating re-finding across the domains. The features are then used to build machine-learned models, achieving an accuracy of re-finding detection in verticals of 85.7% on average. Our results demonstrate that detecting re-finding in specific verticals is more difficult than examining re-finding for general search tasks. We then investigate the effectiveness of differentiating re-finding behavior in two restricted contexts: We consider the case where the history of a searcher’s interactions with the search system is not available. In this scenario, our features and models achieve an average accuracy of 77.5% across the domains. We then examine the detection of re-finding during the early part of a search session. Both of these restrictions represent potential real-world search scenarios, where a system is attempting to learn about a user but may have limited information available. Finally, we investigate in which types of domains re-finding is most difficult. Here, it would appear that re-finding images is particularly challenging for users. This research has implications for search engine design, in terms of adapting search results by predicting the type of user tasks and potentially enabling the presentation of vertical-specific results when re-finding is identified. To the best of our knowledge, this is the first work to investigate the issue of vertical re-finding. Seyedeh Sargol Sadeghi, Roi Blanco, Peter Mika, Mark Sanderson, Falk Scholer, David Vallet |
ACM Trans. Inf. Syst. | 2 |
| 2016 | Memory-based Recommendations of Entities for Web Search UsersabstractModern search engines have evolved from mere document retrieval systems to platforms that assist the users in discovering new information. In this context, entity recommendation systems exploit query log data to proactively provide the users with suggestions of entities (people, movies, places, etc.) from knowledge bases that are relevant for their current information need. Previous works consider the problem of ranking facts and entities related to the user's current query, or focus on specific recommendation domains requiring supervised selection and extraction of features from knowledge bases. In this paper we propose a set of domain-agnostic methods based on nearest neighbors collaborative filtering that exploit query log data to generate entity suggestions, taking into account the user's full search session. Our experimental results on a large dataset from a commercial search engine show that the proposed methods are able to compute relevant entity recommendations outperforming a number of baselines. Finally, we perform an analysis on a cross-domain scenario using different entity types, and conclude that even if knowing the right target domain is important for providing effective recommendations, some inter-domain user interactions are helpful for the task at hand. Ignacio Fernández-Tobías, Roi Blanco |
CIKM | 2 |
| 2016 | Building Test Collections for Evaluating Temporal IRabstractResearch on temporal aspects of information retrieval has recently gained considerable interest within the Information Retrieval (IR) community. This paper describes our efforts for building test collections for the purpose of fostering temporal IR research. In particular, we overview the test collections created at the two recent editions of Temporal Information Access (Temporalia) task organized at NTCIR-11 and NTCIR-12, report on selected results and discuss several observations we made during the task design and implementation. Finally, we outline further directions for constructing test collections suitable for temporal IR. Hideo Joho, Adam Jatowt, Roi Blanco, Hai-Tao Yu 0003, Shuhei Yamamoto |
SIGIR | 3 |
| 2016 | Term-by-Term Query Auto-Completion for Mobile SearchabstractWith the ever increasing usage of mobile search, where text input is typically slow and error-prone, assisting users to formulate their queries contributes to a more satisfactory search experience. Query auto-completion (QAC) techniques, which predict possible completions for user queries, are the archetypal example of query assistance and are present in most search engines. We argue, however, that classic QAC, which operates by suggesting whole-query completions, may be sub-optimal for the case of mobile search as the available screen real estate to show suggestions is limited and editing is typically slower than in desktop search. In this paper we propose the idea of term-by-term QAC, which is a new technique inspired by predictive keyboards that suggests to the user one term at a time, instead of whole-query completions. We describe an efficient mechanism to implement this technique and an adaptation of a prior user model to evaluate the effectiveness of both standard and term-by-term QAC approaches using query log data. Our experiments with a mobile query log from a commercial search engine show the validity of our approach according to this user model with respect to saved characters, saved terms and examination effort. Finally, a user study provides further insights about our term-by-term technique compared with standard QAC with respect to the variables analyzed in the query log-based evaluation and additional variables related to the successfulness, the speed of the interactions and the properties of the submitted queries. Saul Vargas, Roi Blanco, Peter Mika |
WSDM | 2 |
| 2016 | Exploiting Green Energy to Reduce the Operational Costs of Multi-Center Web Search EnginesabstractCarbon dioxide emissions resulting from fossil fuels (brown energy) combustion are the main cause of global warming due to the greenhouse effect. Large IT companies have recently increased their efforts in reducing the carbon dioxide footprint originated from their data center electricity consumption. On one hand, better infrastructure and modern hardware allow for a more efficient usage of electric resources. On the other hand, data-centers can be powered by renewable sources (green energy) that are both environmental friendly and economically convenient. In this paper, we tackle the problem of targeting the usage of green energy to minimize the expenditure of running multi-center Web search engines, i.e., systems composed by multiple, geographically remote, computing facilities. Roi Blanco, Matteo Catena, Nicola Tonellotto |
WWW | 1 |
| 2016 | Using graph distances for named-entity linking
Roi Blanco, Paolo Boldi, Andrea Marino 0001 |
Sci. Comput. Program. | 1 |
| 2015 | Predicting Re-finding Activity and Difficulty
Seyedeh Sargol Sadeghi, Roi Blanco, Peter Mika, Mark Sanderson, Falk Scholer, David Vallet |
ECIR | 2 |
| 2015 | Timely Semantics: A Study of a Stream-Based Ranking System for Entity Relationships
Lorenz Fischer, Roi Blanco, Peter Mika, Abraham Bernstein |
ISWC (2) | 2 |
| 2015 | Local Ranking Problem on the BrowseGraphabstractThe "Local Ranking Problem" (LRP) is related to the computation of a centrality-like rank on a local graph, where the scores of the nodes could significantly differ from the ones computed on the global graph. Previous work has studied LRP on the hyperlink graph but never on the BrowseGraph, namely a graph where nodes are webpages and edges are browsing transitions. Recently, this graph has received more and more attention in many different tasks such as ranking, prediction and recommendation. However, a web-server has only the browsing traffic performed on its pages (local BrowseGraph) and, as a consequence, the local computation can lead to estimation errors, which hinders the increasing number of applications in the state of the art. Also, although the divergence between the local and global ranks has been measured, the possibility of estimating such divergence using only local knowledge has been mainly overlooked. These aspects are of great interest for online service providers who want to: (i) gauge their ability to correctly assess the importance of their resources only based on their local knowledge, and (ii) take into account real user browsing fluxes that better capture the actual user interest than the static hyperlink network. We study the LRP problem on a BrowseGraph from a large news provider, considering as subgraphs the aggregations of browsing traces of users coming from different domains. We show that the distance between rankings can be accurately predicted based only on structural information of the local graph, being able to achieve an average rank correlation as high as 0.8. Michele Trevisiol, Luca Maria Aiello, Paolo Boldi, Roi Blanco |
SIGIR | 4 |
| 2015 | Online News Tracking for Ad-Hoc QueriesabstractFollowing news about a specific event can be a difficult task as new information is often scattered across web pages. An up-to-date summary of the event would help to inform users and allow them to navigate to articles that are likely to contain relevant and novel details. We demonstrate an approach that is feasible for online tracking of news that is relevant to a user's ad-hoc query. Jeroen B. P. Vuurens, Arjen P. de Vries, Roi Blanco, Peter Mika |
SIGIR | 3 |
| 2015 | Fast and Space-Efficient Entity Linking for QueriesabstractEntity linking deals with identifying entities from a knowledge base in a given piece of text and has become a fundamental building block for web search engines, enabling numerous downstream improvements from better document ranking to enhanced search results pages. A key problem in the context of web search queries is that this process needs to run under severe time constraints as it has to be performed before any actual retrieval takes place, typically within milliseconds. Roi Blanco, Giuseppe Ottaviano, Edgar Meij |
WSDM | 1 |
| 2015 | From "Selena Gomez" to "Marlon Brando": Understanding Explorative Entity SearchabstractConsider a user who submits a search query "Shakira" having a specific search goal in mind (such as her age) but at the same time willing to explore information for other entities related to her, such as comparable singers. In previous work, a system called Spark, was developed to provide such search experience. Given a query submitted to the Yahoo search engine, Spark provides related entity suggestions for the query, exploiting, among else, public knowledge bases from the Semantic Web. We refer to this search scenario as explorative entity search. The effectiveness and efficiency of the approach has been demonstrated in previous work. The way users interact with these related entity suggestions and whether this interaction can be predicted have however not been studied. In this paper, we perform a large-scale analysis into how users interact with the entity results returned by Spark. We characterize the users, queries and sessions that appear to promote an explorative behavior. Based on this analysis, we develop a set of query and user-based features that reflect the click behavior of users and explore their effectiveness in the context of a prediction task. Iris Miliaraki, Roi Blanco, Mounia Lalmas-Roelleke |
WWW | 2 |
| 2015 | Predicting primary categories of business listings for local search ranking
Changsung Khan, Jeehaeng Lee, Roi Blanco, Yi Chang 0001 |
Neurocomputing | 3 |
| 2015 | Ranking of daily deals with concept expansion
Roi Blanco, Michael Matthews, Peter Mika |
Inf. Process. Manag. | 1 |
| 2015 | IntoNews: Online news retrieval using closed captions
Roi Blanco, Gianmarco De Francisci Morales, Fabrizio Silvestri |
Inf. Process. Manag. | 1 |
| 2015 | Temporal information searching behaviour and strategies
Hideo Joho, Adam Jatowt, Roi Blanco |
Inf. Process. Manag. | 3 |
| 2014 | Focused Crawling for Structured DataabstractThe Web is rapidly transforming from a pure document collection to the largest connected public data space. Semantic annotations of web pages make it notably easier to extract and reuse data and are increasingly used by both search engines and social media sites to provide better search experiences through rich snippets, faceted search, task completion, etc. In our work, we study the novel problem of crawling structured data embedded inside HTML pages. We describe Anthelion, the first focused crawler addressing this task. We propose new methods of focused crawling specifically designed for collecting data-rich pages with greater efficiency. In particular, we propose a novel combination of online learning and bandit-based explore/exploit approaches to predict data-rich web pages based on the context of the page as well as using feedback from the extraction of metadata from previously seen pages. We show that these techniques significantly outperform state-of-the-art approaches for focused crawling, measured as the ratio of relevant pages and non-relevant pages collected within a given budget. Robert Meusel, Peter Mika, Roi Blanco |
CIKM | 3 |
| 2013 | Influence of Timeline and Named-Entity Components on User Engagement
Yashar Moshfeghi, Michael Matthews, Roi Blanco, Joemon M. Jose |
ECIR | 3 |
| 2013 | Learning Relevance of Web Resources across Domains to Make RecommendationsabstractMost traditional recommender systems focus on the objective of improving the accuracy of recommendations in a single domain. However, preferences of users may extend over multiple domains, especially in the Web where users often have browsing preferences that span across different sites, while being unaware of relevant resources on other sites. This work tackles the problem of recommending resources from various domains by exploiting the semantic content of these resources in combination with patterns of user browsing behavior. We overcome the lack of overlaps between domains by deriving connections based on the explored semantic content of Web resources. We present an approach that applies Support Vector Machines for learning the relevance of resources and predicting which ones are the most relevant to recommend to a user, given that the user is currently viewing a certain page. In real-world datasets of semantically-enriched logs of user browsing behavior at multiple Web sites, we study the impact of structure in generating accurate recommendations and conduct experiments that demonstrate the effectiveness of our approach. Julia Hoxha, Peter Mika, Roi Blanco |
ICMLA (2) | 3 |
| 2013 | Entity Recommendations in Web Search
Roi Blanco, Berkant Barla Cambazoglu, Peter Mika, Nicolas Torzec |
ISWC (2) | 1 |
| 2013 | Federated Entity Search Using On-the-Fly Consolidation
Daniel M. Herzig, Peter Mika, Roi Blanco, Thanh Tran 0001 |
ISWC (1) | 3 |
| 2013 | Web usage mining with semantic analysisabstractWeb usage mining has traditionally focused on the individual queries or query words leading to a web site or web page visit, mining patterns in such data. In our work, we aim to characterize websites in terms of the semantics of the queries that lead to them by linking queries to large knowledge bases on the Web. We demonstrate how to exploit such links for more effective pattern mining on query log data. We also show how such patterns can be used to qualitatively describe the differences between competing websites in the same domain and to quantitatively predict website abandonment. Laura Hollink, Peter Mika, Roi Blanco |
WWW | 3 |
| 2013 | Repeatable and reliable semantic search evaluation
Roi Blanco, Harry Halpin, Daniel M. Herzig, Peter Mika, Jeffrey Pound, Henry S. Thompson, Thanh Tran 0001 |
J. Web Semant. | 1 |
| 2012 | Characterizing web search queries that match very few or no resultsabstractDespite the continuous efforts to improve the web search quality, a non-negligible fraction of user queries end up with very few or even no matching results in leading web search engines. In this work, we provide a detailed characterization of such queries based on an analysis of a real-life query log. Our experimental setup allows us to characterize the queries with few/no results and compare the mechanisms employed by the major search engines in handling them. Ismail Sengör Altingövde, Roi Blanco, Berkant Barla Cambazoglu, Rifat Ozcan, Erdem Sarigil, Özgür Ulusoy |
CIKM | 2 |
| 2012 | You should read this! let me explain you why: explaining news recommendations to usersabstractRecommender systems have become ubiquitous in content-based web applications, from news to shopping sites. Nonetheless, an aspect that has been largely overlooked so far in the recommender system literature is that of automatically building explanations for a particular recommendation. This paper focuses on the news domain, and proposes to enhance effectiveness of news recommender systems by adding, to each recommendation, an explanatory statement to help the user to better understand if, and why, the item can be her interest. We consider the news recommender system as a black-box, and generate different types of explanations employing pieces of information associated with the news. In particular, we engineer text-based, entity-based, and usage-based explanations, and make use of a Markov Logic Networks to rank the explanations on the basis of their effectiveness. The assessment of the model is conducted via a user study on a dataset of news read consecutively by actual users. Experiments show that news recommender systems can greatly benefit from our explanation module as it allows users to discriminate between interesting and not interesting news in the majority of the cases. Roi Blanco, Diego Ceccarelli, Claudio Lucchese, Raffaele Perego 0001, Fabrizio Silvestri |
CIKM | 1 |
| 2012 | Measuring website similarity using an entity-aware click graphabstractQuery logs record the actual usage of search systems and their analysis has proven critical to improving search engine functionality. Yet, despite the deluge of information, query log analysis often suffers from the sparsity of the query space. Based on the observation that most queries pivot around a single entity that represents the main focus of the user's need, we propose a new model for query log data called the entity-aware click graph. In this representation, we decompose queries into entities and modifiers, and measure their association with clicked pages. We demonstrate the benefits of this approach on the crucial task of understanding which websites fulfill similar user needs, showing that using this representation we can achieve a higher precision than other query log-based approaches. Pablo N. Mendes, Peter Mika, Hugo Zaragoza, Roi Blanco |
CIKM | 4 |
| 2012 | Extending BM25 with multiple query operatorsabstractTraditional probabilistic relevance frameworks for informational retrieval refrain from taking positional information into account, due to the hurdles of developing a sound model while avoiding an explosion in the number of parameters. Nonetheless, the well-known BM25F extension of the successful Okapi ranking function can be seen as an embryonic attempt in that direction. In this paper, we proceed along the same line, defining the notion of virtual region: a virtual region is a part of the document that, like a BM25F-field, can provide a (larger or smaller, depending on a tunable weighting parameter) evidence of relevance of the document; differently from BM25F fields, though, virtual regions are generated implicitly by applying suitable (usually, but not necessarily, positional-aware) operators to the query. This technique fits nicely in the eliteness model behind BM25 and provides a principled explanation to BM25F; it specializes to BM25(F) for some trivial operators, but has a much more general appeal. Our experiments (both on standard collections, such as TREC, and on Web-like repertoires) show that the use of virtual regions is beneficial for retrieval effectiveness. Roi Blanco, Paolo Boldi |
SIGIR | 1 |
| 2012 | Language intent models for inferring user browsing behaviorabstractModeling user browsing behavior is an active research area with tangible real-world applications, e.g., organizations can adapt their online presence to their visitors browsing behavior with positive effects in user engagement, and revenue. We concentrate on online news agents, and present a semi-supervised method for predicting news articles that a user will visit after reading an initial article. Our method tackles the problem using language intent models trained on historical data which can cope with unseen articles. We evaluate our method on a large set of articles and in several experimental settings. Our results demonstrate the utility of language intent models for predicting user browsing behavior within online news sites. Manos Tsagkias, Roi Blanco |
SIGIR | 2 |
| 2012 | Graph-based term weighting for information retrieval
Roi Blanco, Christina Lioma |
Inf. Retr. | 1 |
| 2011 | Hybrid models for future event predictionabstractWe present a hybrid method to turn off-the-shelf information retrieval (IR) systems into future event predictors. Given a query, a time series model is trained on the publication dates of the retrieved documents to capture trends and periodicity of the associated events. The periodicity of historic data is used to estimate a probabilistic model to predict future bursts. Finally, a hybrid model is obtained by intertwining the probabilistic and the time-series model. Our empirical results on the New York Times corpus show that autocorrelation functions of time-series suffice to classify queries accurately and that our hybrid models lead to more accurate future event predictions than baseline competitors. Giuseppe Amodeo, Roi Blanco, Ulf Brefeld |
CIKM | 2 |
| 2011 | Assigning documents to master sites in distributed searchabstractAn appealing solution to scale Web search with the growth of the Internet is the use of distributed architectures. Distributed search engines rely on multiple sites deployed in distant regions across the world, where each site is specialized to serve queries issued by the users of its region. This paper investigates the problem of assigning each document to a master site. We show that by leveraging similarities between a document and the activity of the users, we can accurately detect which site is the most relevant to place a document. We conduct various experiments using two document assignment approaches, showing performance improvements of up to 20.8% over a baseline technique which assigns the documents to search sites based on their language. Roi Blanco, Berkant Barla Cambazoglu, Flavio Paiva Junqueira, Ivan Kelly, Vincent Leroy 0001 |
CIKM | 1 |
| 2011 | Coreference aware web object retrievalabstractAs user demands become increasingly sophisticated, search engines today are competing in more than just returning document results from the Web. One area of competition is providing web object results from structured data extracted from a multitude of information sources. We address the problem of performing keyword retrieval over a collection of objects containing a large degree of duplication as different Web-based information sources provide descriptions of the same object. We develop a method for coreference aware retrieval that performs topic-specific coreference resolution on retrieved objects in order to improve object search results. Our results demonstrate that coreference has a significant impact on the effectiveness of retrieval in the domain of local search. Our results show that a coreference aware system outperforms naive object retrieval by more than 20% in P5 and P10. Jeff Dalton 0001, Roi Blanco, Peter Mika |
CIKM | 2 |
| 2011 | Keyword search over RDF graphsabstractLarge knowledge bases consisting of entities and relationships between them have become vital sources of information for many applications. Most of these knowledge bases adopt the Semantic-Web data model RDF as a representation model. Querying these knowledge bases is typically done using structured queries utilizing graph-pattern languages such as SPARQL. However, such structured queries require some expertise from users which limits the accessibility to such data sources. To overcome this, keyword search must be supported. In this paper, we propose a retrieval model for keyword queries over RDF graphs. Our model retrieves a set of subgraphs that match the query keywords, and ranks them based on statistical language models. We show that our retrieval model outperforms the-state-of-the-art IR and DB models for keyword search over structured data using experiments over two real-world datasets. Shady Elbassuoni, Roi Blanco |
CIKM | 2 |
| 2011 | Effective and Efficient Entity Search in RDF Data
Roi Blanco, Peter Mika, Sebastiano Vigna |
ISWC (1) | 1 |
| 2011 | Repeatable and reliable search system evaluation using crowdsourcingabstractThe primary problem confronting any new kind of search task is how to boot-strap a reliable and repeatable evaluation campaign, and a crowd-sourcing approach provides many advantages. However, can these crowd-sourced evaluations be repeated over long periods of time in a reliable manner? To demonstrate, we investigate creating an evaluation campaign for the semantic search task of keyword-based ad-hoc object retrieval. In contrast to traditional search over web-pages, object search aims at the retrieval of information from factual assertions about real-world objects rather than searching over web-pages with textual descriptions. Using the first large-scale evaluation campaign that specifically targets the task of ad-hoc Web object retrieval over a number of deployed systems, we demonstrate that crowd-sourced evaluation campaigns can be repeated over time and still maintain reliable results. Furthermore, we show how these results are comparable to expert judges when ranking systems and that the results hold over different evaluation and relevance metrics. This work provides empirical support for scalable, reliable, and repeatable search system evaluation using crowdsourcing. Roi Blanco, Harry Halpin, Daniel M. Herzig, Peter Mika, Jeffrey Pound, Henry S. Thompson, Thanh Tran 0001 |
SIGIR | 1 |
| 2011 | Enhanced results for web searchabstract"Ten blue links" have defined web search results for the last fifteen years -- snippets of text combined with document titles and URLs. In this paper, we establish the notion of enhanced search results that extend web search results to include multimedia objects such as images and video, intent-specific key value pairs, and elements that allow the user to interact with the contents of a web page directly from the search results page. We show that users express a preference for enhanced results both explicitly, and when observed in their search behavior. We also demonstrate the effectiveness of enhanced results in helping users to assess the relevance of search results. Lastly, we show that we can efficiently generate enhanced results to cover a significant fraction of search result pages. Kevin Haas, Peter Mika, Paul Tarjan, Roi Blanco |
SIGIR | 4 |
| 2011 | Ranking related news predictionsabstractWe estimate that nearly one third of news articles contain references to future events. While this information can prove crucial to understanding news stories and how events will develop for a given topic, there is currently no easy way to access this information. We propose a new task to address the problem of retrieving and ranking sentences that contain mentions to future events, which we call ranking related news predictions. In this paper, we formally define this task and propose a learning to rank approach based on 4 classes of features: term similarity, entity-based similarity, topic similarity, and temporal similarity. Through extensive evaluations using a corpus consisting of 1.8 millions news articles and 6,000 manually judged relevance pairs, we show that our approach is able to retrieve a significant number of relevant predictions related to a given topic. Nattiya Kanhabua, Roi Blanco, Michael Matthews |
SIGIR | 2 |
| 2011 | Energy-price-driven query processing in multi-center web search enginesabstractConcurrently processing thousands of web queries, each with a response time under a fraction of a second, necessitates maintaining and operating massive data centers. For large-scale web search engines, this translates into high energy consumption and a huge electric bill. This work takes the challenge to reduce the electric bill of commercial web search engines operating on data centers that are geographically far apart. Based on the observation that energy prices and query workloads show high spatio-temporal variation, we propose a technique that dynamically shifts the query workload of a search engine between its data centers to reduce the electric bill. Experiments on real-life query workloads obtained from a commercial search engine show that significant financial savings can be achieved by this technique. Enver Kayaaslan, Berkant Barla Cambazoglu, Roi Blanco, Flavio Paiva Junqueira, Cevdet Aykanat |
SIGIR | 3 |
| 2010 | TAER: time-aware entity retrieval-exploiting the past to find relevant entities in news articlesabstractRetrieving entities instead of just documents has become an important task for search engines. In this paper we study entity retrieval for news applications, and in particular the importance of the news trail history (i.e., past related articles) in determining the relevant entities in current articles. This is an important problem in applications that display retrieved entities to the user, together with the news article. Gianluca Demartini, Malik Muhammad Saad Missen, Roi Blanco, Hugo Zaragoza |
CIKM | 3 |
| 2010 | Caching search engine results over incremental indicesabstractA Web search engine must update its index periodically to incorporate changes to the Web. We argue in this paper that index updates fundamentally impact the design of search engine result caches, a performance-critical component of modern search engines. Index updates lead to the problem of cache invalidation: invalidating cached entries of queries whose results have changed. Naive approaches, such as flushing the entire cache upon every index update, lead to poor performance and in fact, render caching futile when the frequency of updates is high. Solving the invalidation problem efficiently corresponds to predicting accurately which queries will produce different results if re-evaluated, given the actual changes to the index. Roi Blanco, Edward Bortnikov, Flavio Paiva Junqueira, Ronny Lempel, Luca Telloli, Hugo Zaragoza |
SIGIR | 1 |
| 2010 | Finding support sentences for entitiesabstractWe study the problem of finding sentences that explain the relationship between a named entity and an ad-hoc query, which we refer to as entity support sentences. Thisisanimportant sub-problem of entity ranking which, to the best of our knowledge, has not been addressed before. In this paper we give the first formalization of the problem, how it can be evaluated, and present a full evaluation dataset. We propose several methods to rank these sentences, namely retrievalbased, entity-ranking based and position-based. We found that traditional bag-of-words models perform relatively well when there is a match between an entity and a query in a given sentence, but they fail to find a support sentence for a substantial portion of entities. This can be improved by incorporating small windows of context sentences and ranking them appropriately. Roi Blanco, Hugo Zaragoza |
SIGIR | 1 |
| 2010 | Entity summarization of news articlesabstractinc.com In this paper we study the problem of entity retrieval for news applications and the importance of the news trail his-tory (i.e. past related articles) to determine the relevant entities in current articles. We construct a novel entity-labeled corpus with temporal information out of the TREC 2004 Novelty collection. We develop and evaluate several features, and show that an article’s history can be exploited to improve its summarization. Gianluca Demartini, Malik Muhammad Saad Missen, Roi Blanco, Hugo Zaragoza |
SIGIR | 3 |
| 2010 | Caching search engine results over incremental indicesabstractA Web search engine must update its index periodically to incorporate changes to the Web, and we argue in this work that index updates fundamentally impact the design of search engine result caches. Index updates lead to the problem of cache invalidation: invalidating cached entries of queries whose results have changed. To enable efficient invalidation of cached results, we propose a framework for developing invalidation predictors and some concrete predictors. Evaluation using Wikipedia documents and a query log from Yahoo! shows that selective invalidation of cached search results can lower the number of query re-evaluations by as much as 30% compared to a baseline time-to-live scheme, while returning results of similar freshness. Roi Blanco, Edward Bortnikov, Flavio Paiva Junqueira, Ronny Lempel, Luca Telloli, Hugo Zaragoza |
WWW | 1 |
| 2010 | Probabilistic static pruning of inverted filesabstractInformation retrieval (IR) systems typically compress their indexes in order to increase their efficiency. Static pruning is a form of lossy data compression: it removes from the index, data that is estimated to be the least important to retrieval performance, according to some criterion. Generally, pruning criteria are derived from term weighting functions, which assign weights to terms according to their contribution to a document's contents. Usually, document-term occurrences that are assigned a low weight are ruled out from the index. The main assumption is that those entries contribute little to the document content. We present a novel pruning technique that is based on a probabilistic model of IR. We employ the Probability Ranking Principle as a decision criterion over which posting list entries are to be pruned. The proposed approach requires the estimation of three probabilities, combining them in such a way that we gather all the necessary information to apply the aforementioned criterion. We evaluate our proposed pruning technique on five TREC collections and various retrieval tasks, and show that in almost every situation it outperforms the state of the art in index pruning. The main contribution of this work is proposing a pruning technique that stems directly from the same source as probabilistic retrieval models, and hence is independent of the final model used for retrieval. Roi Blanco, Álvaro Barreiro |
ACM Trans. Inf. Syst. | 1 |
| 2009 | Part of Speech Based Term Weighting for Information Retrieval
Christina Lioma, Roi Blanco |
ECIR | 2 |
| 2009 | Mixed monolingual homepage finding in 34 languages: the role of language script and search domain
Roi Blanco, Christina Lioma |
Inf. Retr. | 1 |
| 2008 | Probabilistic Document Length Priors for Language Models
Roi Blanco, Álvaro Barreiro |
ECIR | 1 |
| 2008 | Efficiency Issues in Information Retrieval Workshop
Roi Blanco, Fabrizio Silvestri |
ECIR | 1 |
| 2007 | Static Pruning of Terms in Inverted Files
Roi Blanco, Álvaro Barreiro |
ECIR | 1 |
| 2007 | Boosting static pruning of inverted filesabstractThis paper revisits the static term-based pruning technique presented in Carmel et al., SIGIR 2001 for ad-hoc retrieval, addressing different issues concerning its algorithmic design not yet taken into account. Although the original technique is able to retain precision when a considerable part of the inverted file is removed, we show that it is possible to improve precision in some scenarios if some key design features are properly selected. Roi Blanco, Álvaro Barreiro |
SIGIR | 1 |
| 2007 | Random walk term weighting for information retrievalabstractWe present a way of estimating term weights for Information Retrieval (IR), using term co-occurrence as a measure of dependency between terms.We use the random walk graph-based ranking algorithm on a graph that encodes terms and co-occurrence dependencies in text, from which we derive term weights that represent a quantification of how a term contributes to its context. Evaluation on two TREC collections and 350 topics shows that the random walk-based term weights perform at least comparably to the traditional tf-idf term weighting, while they outperform it when the distance between co-occurring terms is between 6 and 30 terms. Roi Blanco, Christina Lioma |
SIGIR | 1 |
| 2006 | TSP and cluster-based solutions to the reassignment of document identifiers
Roi Blanco, Álvaro Barreiro |
Inf. Retr. | 1 |
| 2005 | Document Identifier Reassignment Through Dimensionality Reduction
Roi Blanco, Álvaro Barreiro |
ECIR | 1 |
| 2005 | Characterization of a simple case of the reassignment of document identifiers as a pattern sequencing problemabstractIn this poster, we analyze recent work in the document identifiers reassignment problem. After that, we present a formalization of a simple case of the problem as a PSP (Pattern Sequencing Problem). This may facilitate future work as it opens a new research line to solve the general problem. Roi Blanco, Álvaro Barreiro |
SIGIR | 1 |