Berkant Barla Cambazoglu

dblp:57/2006 · DBLP profile ↗
← Back
72ranked-venue papers
17as first author
2since 2021 · last 2021
0000-0003-2192-3819ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 62 · 16 first-author · 2 since 2021Artificial intelligence and machine learning · 20 · 2 first-authorSystems, architecture and hardware · 5 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 first-authorComputer networks · 2Human-computer interaction and ubiquitous computing · 2 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
26 papers
Information retrieval · 85% Query processing and optimization · 6% Indexing and storage engines · 4%
Computer architecture, parallel and distributed computing, and storage systems
11 papers
Parallel and multicore computing · 40% Distributed systems · 32% Cloud and datacenter computing · 10%

Topics — the 30 heaviest of 57, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval
search engines
1.492016
Scalability and Efficiency Challenges in Large-Scale Web Search Engines · SIGIR 2016
Scalability and Efficiency Challenges in Large-Scale Web Search Engines · WSDM 2015
Unconscious Physiological Effects of Search Latency on Users and Their Click Behaviour · SIGIR 2015
Information retrieval
web search
0.962017
Understanding and Leveraging the Impact of Response Latency on User Behaviour in Web Search · ACM Trans. Inf. Syst. 2017
Improving the efficiency of multi-site web search engines · WSDM 2014
Prefetching query results and its impact on search engines · SIGIR 2012
Information retrieval
user behavior
0.732017
Understanding and Leveraging the Impact of Response Latency on User Behaviour in Web Search · ACM Trans. Inf. Syst. 2017
Unconscious Physiological Effects of Search Latency on Users and Their Click Behaviour · SIGIR 2015
Impact of response latency on user behavior in web search · SIGIR 2014
Information retrieval › search engines
web search engine
0.632016
Scalability and Efficiency Challenges in Large-Scale Web Search Engines · SIGIR 2016
Scalability and efficiency challenges in large-scale web search engines · SIGIR 2014
Scalability and efficiency challenges in commercial web search engines · SIGIR 2013
Information retrieval › search engines
web crawling
0.532016
Optimal Web Page Download Scheduling Policies for Green Web Crawling · IEEE J. Sel. Areas Commun. 2016
A Random Walk Model for Optimization of Search Impact in Web Frontier Ranking · SIGIR 2015
Quantifying performance and quality gains in distributed web search engines · SIGIR 2009
Information retrieval › search engines
search engine architecture
0.532015
Scalability and Efficiency Challenges in Large-Scale Web Search Engines · WSDM 2015
Timestamp-based result cache invalidation for web search engines · SIGIR 2011
A refreshing perspective of search engine caching · WWW 2010
Query processing and optimization
query result caching
0.432013
A financial cost metric for result caching · SIGIR 2013
Prefetching query results and its impact on search engines · SIGIR 2012
Timestamp-based result cache invalidation for web search engines · SIGIR 2011
Information retrieval › user behavior › search behavior
click behavior
0.422017
Understanding and Leveraging the Impact of Response Latency on User Behaviour in Web Search · ACM Trans. Inf. Syst. 2017
Unconscious Physiological Effects of Search Latency on Users and Their Click Behaviour · SIGIR 2015
Information retrieval
query processing
0.342017
Posting list intersection on multicore architectures · SIGIR 2011
On efficient posting list intersection with multicore processors · SIGIR 2009
Understanding and Leveraging the Impact of Response Latency on User Behaviour in Web Search · ACM Trans. Inf. Syst. 2017
Information retrieval
distributed information retrieval
0.322014
Workshop on large-scale and distributed systems for information retrieval (LSDS-IR 2014) · WSDM 2014
Quantifying performance and quality gains in distributed web search engines · SIGIR 2009
Indexing and storage engines
caching
0.212016
Improved Caching Techniques for Large-Scale Image Hosting Services · SIGIR 2016
Information retrieval › distributed information retrieval
distributed search
0.222011
Document assignment in multi-site search engines · WSDM 2011
Query forwarding in geographically distributed search engines · SIGIR 2010
Information retrieval › query processing
list intersection
0.222011
Posting list intersection on multicore architectures · SIGIR 2011
On efficient posting list intersection with multicore processors · SIGIR 2009
Information retrieval › evaluation
search effectiveness
0.212015
A Random Walk Model for Optimization of Search Impact in Web Frontier Ranking · SIGIR 2015
Information retrieval
interactive information retrieval
0.212014
Impact of response latency on user behavior in web search · SIGIR 2014
Information retrieval
retrieval evaluation
0.212014
Impact of response latency on user behavior in web search · SIGIR 2014
Distributed systems › peer-to-peer systems
distributed search
0.212014
Improving the efficiency of multi-site web search engines · WSDM 2014
Distributed systems
query forwarding
0.212014
Improving the efficiency of multi-site web search engines · WSDM 2014
Distributed systems
query result caching
0.212014
Improving the efficiency of multi-site web search engines · WSDM 2014
Parallel and multicore computing
scheduling algorithms
0.212014
Improving the Performance of IndependentTask Assignment Heuristics MinMin, MaxMin and Sufferage · IEEE Trans. Parallel Distributed Syst. 2014
Parallel and multicore computing
task allocation
0.212014
Improving the Performance of IndependentTask Assignment Heuristics MinMin, MaxMin and Sufferage · IEEE Trans. Parallel Distributed Syst. 2014
Information retrieval
evaluation
0.212013
A financial cost metric for result caching · SIGIR 2013
Query processing and optimization › runtime optimization › prefetching
query result prefetching
0.112012
Prefetching query results and its impact on search engines · SIGIR 2012
Data mining › text mining
sentiment analysis
0.112012
A large-scale sentiment analysis for Yahoo! answers · WSDM 2012
Database system architecture and tuning
cache invalidation
0.112011
Timestamp-based result cache invalidation for web search engines · SIGIR 2011
Information retrieval › distributed information retrieval
document allocation
0.112011
Document assignment in multi-site search engines · WSDM 2011
Indexing and storage engines
index maintenance
0.112011
Timestamp-based result cache invalidation for web search engines · SIGIR 2011
Cloud and datacenter computing › resource management
datacenter resource management
0.112011
Energy-price-driven query processing in multi-center web search engines · SIGIR 2011
Hardware accelerators and domain-specific architectures › query processing
energy-efficient query processing
0.112011
Energy-price-driven query processing in multi-center web search engines · SIGIR 2011
Parallel and multicore computing › parallelization strategies
fine-grained parallelization
0.112011
Posting list intersection on multicore architectures · SIGIR 2011

Methods — techniques the papers use, named apart from their topics

simulation · 0.7query log analysis · 0.7machine learning · 0.6greedy policy · 0.5gain-based caching · 0.2correlation analysis · 0.2random walk · 0.2physiological measurement · 0.2link graph enrichment · 0.2controlled experiment · 0.2sufferage · 0.2minmin · 0.2maxmin · 0.2hybrid algorithms · 0.2controlled user study · 0.2LRU caching · 0.2workload shifting · 0.1sparse matrix partitioning · 0.1
YearPublicationVenuePosition
2021 Quantifying Human-Perceived Answer Utility in Non-factoid Question Answering
abstract
Taking a user-centric approach, we study the features that render an answer to a non-factoid question useful in the eyes of the person who asked that question. An editorial study, where participants assess the usefulness of the answers they received in response to their questions, as well as 12 different aspects associated with the answers, indicates considerable correlation between certain aspects such as relevance, correctness, and completeness with the user-perceived usefulness of answers. Moreover, we investigate the effectiveness of some commonly used answer quality measures, such as ROGUE, BLEU, METEOR, and BERTScore, demonstrating that these measures are limited in their ability to capture the aspects of usefulness and have room for improvement. The question answering dataset created in our work was made publicly available.
Berkant Barla Cambazoglu, Valeria Bolotova-Baranova, Falk Scholer, Mark Sanderson, Leila Tavakoli, W. Bruce Croft
CHIIR1
2021 An Intent Taxonomy for Questions Asked in Web Search
abstract
We present a new, multi-faceted taxonomy to classify questions asked in web search engines based on the question intent, types of entities mentioned, types of question words, and granularity of the expected answer. Built based on the inspection of 1,000 real-life questions issued to a web search engine, the taxonomy reflects the recent search behavior of users and enables deep understanding of user intents, goals, and expected answers. This taxonomy is more fine-grained than previous query taxonomies, and is designed with the ultimate goal of reducing the inherent ambiguity in determining the intent of questions. In addition, we describe the formal procedure for conducting an editorial study of the taxonomy including its evaluation. The adopted procedure aims to increase assessor agreement without incurring too much overhead. Our results demonstrate that, despite being more fine-grained, the proposed intent categories result in higher agreement between assessors compared to an existing, commonly used taxonomy.
Berkant Barla Cambazoglu, Leila Tavakoli, Falk Scholer, Mark Sanderson, W. Bruce Croft
CHIIR1
2020 Providing Direct Answers in Search Results: A Study of User Behavior
abstract
To study the impact of providing direct answers in search results on user behavior, we conducted a controlled user study to analyze factors including reading time, eye-tracked attention, and the influence of the quality of answer module content. We also studied a more advanced answer interface, where multiple answers are shown on the search engine results page (SERP). Our results show that users focus more extensively than normal on the top items in the result list when answers are provided. The existence of the answer module helps to improve user engagement on SERPs, reduces user effort, and promotes user satisfaction during the search process. Furthermore, we investigate how the question type -- factoid or non-factoid -- affects user interaction patterns. This work provides insight into the design of SERPs that includes direct answers to queries, including when answers should be shown.
Zhijing Wu 0001, Mark Sanderson, Berkant Barla Cambazoglu, W. Bruce Croft, Falk Scholer
CIKM3
2020 Feature Extraction for Large-Scale Text Collections
abstract
Feature engineering is a fundamental but poorly documented component in Learning-to-Rank (LTR) search engines. Such features are commonly used to construct learning models for web and product search engines, recommender systems, and question-answering tasks. In each of these domains, there is a growing interest in the creation of open-access test collections that promote reproducible research. However, there are still few open-source software packages capable of extracting high-quality machine learning features from large text collections. Instead, most feature-based LTR research relies on "canned" test collections, which often do not expose critical details about the underlying collection or implementation details of the extracted features. Both of these are crucial to collection creation and deployment of a search engine into production. So in this regard, the experiments are rarely reproducible with new features or collections, or helpful for companies wishing to deploy LTR systems.
Luke Gallagher, Antonio Mallia, J. Shane Culpepper, Torsten Suel, Berkant Barla Cambazoglu
CIKM5
2020 Pre-indexing Pruning Strategies
Soner Altin, Ricardo Baeza-Yates, Berkant Barla Cambazoglu
SPIRE3
2019 Impact of response latency on sponsored search
Xiao Bai 0002, Berkant Barla Cambazoglu
Inf. Process. Manag.2
2018 Characterizing, predicting, and handling web search queries that match very few or no results
abstract
A non‐negligible fraction of user queries end up with very few or even no matching results in leading commercial web search engines. In this work, we provide a detailed characterization of such queries and show that search engines try to improve such queries by showing the results of related queries. Through a user study, we show that these query suggestions are usually perceived as relevant. Also, through a query log analysis, we show that the users are dissatisfied after submitting a query that match no results at least 88.5% of the time. As a first step towards solving these no‐answer queries, we devised a large number of features that can be used to identify such queries and built machine‐learning models. These models can be useful for scenarios such as the mobile‐ or meta‐search, where identifying a query that will retrieve no results at the client device (i.e., even before submitting it to the search engine) may yield gains in terms of the bandwidth usage, power consumption, and/or monetary costs. Experiments over query logs indicate that, despite the heavy skew in class sizes, our models achieve good prediction quality, with accuracy (in terms of area under the curve) up to 0.95.
Erdem Sarigil, Ismail Sengör Altingövde, Roi Blanco, Berkant Barla Cambazoglu, Rifat Ozcan, Özgür Ulusoy
J. Assoc. Inf. Sci. Technol.4
2017 A machine learning approach for result caching in web search engines
Tayfun Küçükyilmaz, Berkant Barla Cambazoglu, Cevdet Aykanat, Ricardo Baeza-Yates
Inf. Process. Manag.2
2017 Exploiting search history of users for news personalization
Xiao Bai 0002, Berkant Barla Cambazoglu, Francesco Gullo, Amin Mantrach, Fabrizio Silvestri
Inf. Sci.2
2017 On the feasibility of predicting popular news at cold start
abstract
Prominent news sites on the web provide hundreds of news articles daily. The abundance of news content competing to attract online attention, coupled with the manual effort involved in article selection, necessitates the timely prediction of future popularity of these news articles. The future popularity of a news article can be estimated using signals indicating the article's penetration in social media (e.g., number of tweets) in addition to traditional web analytics (e.g., number of page views). In practice, it is important to make such estimations as early as possible, preferably before the article is made available on the news site (i.e., at cold start). In this paper we perform a study on cold‐start news popularity prediction using a collection of 13,319 news articles obtained from Yahoo News, a major news provider. We characterize the popularity of news articles through a set of online metrics and try to predict their values across time using machine learning techniques on a large collection of features obtained from various sources. Our findings indicate that predicting news popularity at cold start is a difficult task, contrary to the findings of a prior work on the same topic. Most articles' popularity may not be accurately anticipated solely on the basis of content features, without having the early‐stage popularity values.
Ioannis Arapakis, Berkant Barla Cambazoglu, Mounia Lalmas-Roelleke
J. Assoc. Inf. Sci. Technol.2
2017 Understanding and Leveraging the Impact of Response Latency on User Behaviour in Web Search
abstract
The interplay between the response latency of web search systems and users’ search experience has only recently started to attract research attention, despite the important implications of response latency on monetisation of such systems. In this work, we carry out two complementary studies to investigate the impact of response latency on users’ searching behaviour in web search engines. We first conduct a controlled user study to investigate the sensitivity of users to increasing delays in response latency. This study shows that the users of a fast search system are more sensitive to delays than the users of a slow search system. Moreover, the study finds that users are more likely to notice the response latency delays beyond a certain latency threshold, their search experience potentially being affected. We then analyse a large number of search queries obtained from Yahoo Web Search to investigate the impact of response latency on users’ click behaviour. This analysis demonstrates the significant change in click behaviour as the response latency increases. We also find that certain user, context, and query attributes play a role in the way increasing response latency affects the click behaviour. To demonstrate a possible use case for our findings, we devise a machine-learning framework that leverages the latency impact, together with other features, to predict whether a user will issue any clicks on web search results. As a further extension of this use case, we investigate whether this machine-learning framework can be exploited to help search engines reduce their energy consumption during query processing.
Xiao Bai 0002, Ioannis Arapakis, Berkant Barla Cambazoglu, Ana Freire
ACM Trans. Inf. Syst.3
2016 Linguistic Benchmarks of Online News Article Quality
abstract
Online news editors ask themselves the same question many times: what is missing in this news article to go online?This is not an easy question to be answered by computational linguistic methods.In this work, we address this important question and characterise the constituents of news article editorial quality.More specifically, we identify 14 aspects related to the content of news articles.Through a correlation analysis, we quantify their independence and relation to assessing an article's editorial quality.We also demonstrate that the identified aspects, when combined together, can be used effectively in quality control methods for online news.
Ioannis Arapakis, Filipa Peleja, Berkant Barla Cambazoglu, João Magalhães
ACL (1)3
2016 Improved Caching Techniques for Large-Scale Image Hosting Services
abstract
Commercial image serving systems, such as Flickr and Facebook, rely on large image caches to avoid the retrieval of requested images from the costly backend image store, as much as possible. Such systems serve the same image in different resolutions and, thus, in different sizes to different clients, depending on the properties of the clients' devices. The requested resolutions of images can be cached individually, as in the traditional caches, reducing the backend workload. However, a potentially better approach is to store relatively high-resolution images in the cache and resize them during the retrieval to obtain lower-resolution images. Having this kind of on-the-fly image resizing capability enables image serving systems to deploy more sophisticated caching policies and improve their serving performance further. In this paper, we formalize the static caching problem in image serving systems which provide on-the-fly image resizing functionality in their edge caches or regional caches. We propose two gain-based caching policies that construct a static, fixed-capacity cache to reduce the average serving time of images. The basic idea in the proposed policies is to identify the best resolution(s) of images to be cached so that the average serving time for future image retrieval requests is reduced. We conduct extensive experiments using real-life data access logs obtained from Flickr. We show that one of the proposed caching policies reduces the average response time of the service by up to 4.2% with respect to the best-performing baseline that mainly relies on the access frequency information to make the caching decisions. This improvement implies about 25% reduction in cache size under similar serving time constraints.
Xiao Bai 0002, Berkant Barla Cambazoglu, Archie Russell
SIGIR2
2016 Scalability and Efficiency Challenges in Large-Scale Web Search Engines
abstract
Commercial web search engines need to process thousands of queries every second and provide responses to user queries within a few hundred milliseconds. As a consequence of these tight performance constraints, search engines construct and maintain very large computing infrastructures for crawling the Web, indexing discovered pages, and processing user queries. The scalability and efficiency of these infrastructures require careful performance optimizations in every major component of the search engine.
Berkant Barla Cambazoglu, Ricardo Baeza-Yates
SIGIR1
2016 Optimal Web Page Download Scheduling Policies for Green Web Crawling
abstract
A web crawler is responsible for discovering and downloading new pages on the Web as well as refreshing previously downloaded pages. During these operations, the crawler issues a large number of HTTP requests to web servers. These requests increase the energy consumption and carbon footprint of the web servers since computational resources are used while serving the requests. In this work, we introduce the problem of green web crawling, where the objective is to devise a page refresh policy that minimizes the total staleness of pages in the repository of a web crawler, subject to a constraint on the amount of carbon emissions due to the processing on web servers. For the case of one web server and one crawling thread, the optimal policy turns out to be a greedy one. At each iteration, the page to be refreshed is selected based on a metric that considers the page's staleness, its size, and the greenness of the energy consumed at the web server premises. We then extend the optimal policy to the cases of 1) many servers; 2) multiple threads; and 3) pages with variable freshness requirements. We conduct simulations on a real data set that involves a large web server collection hosting around two billion pages. We present experimental results for the optimal page refresh policy as well as for various heuristics, in an effort to study the effect of different factors on performance.
Vassiliki Hatzi, Berkant Barla Cambazoglu, Iordanis Koutsopoulos
IEEE J. Sel. Areas Commun.2
2015 LSDS-IR'15: 2015 Workshop on Large-Scale and Distributed Systems for Information Retrieval
abstract
The growth of the Web and other Big Data sources lead to important performance problems for large-scale and distributed information retrieval systems. The scalability and efficiency of such information retrieval systems have an impact on their effectiveness, eventually affecting the experience of their users and monetization as well. The LSDS-IR'15 workshop will provide space for researchers to discuss the existing performance problems in the context of large-scale and distributed information retrieval systems and define new research directions in the modern Big Data era. The workshop expects to bring together information retrieval practitioners from the industry, as well as academic researchers concerned with any aspect of large-scale and distributed information retrieval systems.
Ismail Sengör Altingövde, Berkant Barla Cambazoglu, Nicola Tonellotto
CIKM2
2015 Know Your Onions: Understanding the User Experience with the Knowledge Module in Web Search
abstract
The increasing availability of large volumes of human-curated content is shifting web search towards a paradigm that introduces seamlessly more semantic information to search engine result pages. This trend has resulted in the design of a new element known as the knowledge module (KM) where certain facts about named entities, obtained from various knowledge bases, are shown to users. So far, little has been done to uncover the role that this module plays on user experience in web search and whether it is perceived by users as a useful aid for their search tasks. Our work is an early attempt to bridge this gap. To this end, we conducted a crowdsourcing study aimed at understanding the effect of the KM on users' search experience and its overall utility. In particular, our study is the first to provide insights about the noticeability and usefulness of the KM in web search, together with comprehensive analyses of usability and workload.
Ioannis Arapakis, Luis A. Leiva, Berkant Barla Cambazoglu
CIKM3
2015 Unconscious Physiological Effects of Search Latency on Users and Their Click Behaviour
abstract
Understanding the impact of a search system's response latency on its users' searching behaviour has been recently an active research topic in the information retrieval and human-computer interaction areas. Along the same line, this paper focuses on the user impact of search latency and makes the following two contributions. First, through a controlled experiment, we reveal the physiological effects of response latency on users and show that these effects are present even at small increases in response latency. We compare these effects with the information gathered from self-reports and show that they capture the nuanced attentional and emotional reactions to latency much better. Second, we carry out a large-scale analysis using a web search query log obtained from Yahoo to understand the change in the way users engage with a web search engine under varying levels of increasing response latency. In particular, we analyse the change in the click behaviour of users when they are subject to increasing response latency and reveal significant behavioural differences.
Miguel Barreda-Ángeles, Ioannis Arapakis, Xiao Bai 0002, Berkant Barla Cambazoglu, Alexandre Pereda-Baños
SIGIR4
2015 A Random Walk Model for Optimization of Search Impact in Web Frontier Ranking
abstract
Large-scale web search engines need to crawl the Web continuously to discover and download newly created web content. The speed at which the new content is discovered and the quality of the discovered content can have a big impact on the coverage and quality of the results provided by the search engine. In this paper, we propose a search-centric solution to the problem of prioritizing the pages in the frontier of a crawler for download. Our approach essentially orders the web pages in the frontier through a random walk model that takes into account the pages' potential impact on user-perceived search quality. In addition, we propose a link graph enrichment technique that extends this solution. Finally, we explore a machine learning approach that combines different frontier prioritization approaches. We conduct experiments using two very large, real-life web datasets to observe various search quality metrics. Comparisons with several baseline techniques indicate that the proposed approaches have the potential to improve the user-perceived quality of web search results considerably.
Giang Tran, Ata Turk, Berkant Barla Cambazoglu, Wolfgang Nejdl
SIGIR3
2015 Scalability and Efficiency Challenges in Large-Scale Web Search Engines
abstract
Commercial web search engines need to process thousands of queries every second and provide responses to user queries within a few hundred milliseconds. As a consequence of these tight performance constraints, search engines construct and maintain very large computing infrastructures for crawling the Web, indexing discovered pages, and processing user queries. The scalability and efficiency of these infrastructures require careful performance optimizations in every major component of the search engine. This tutorial aims to provide a fairly comprehensive overview of the scalability and efficiency challenges in large-scale web search engines. In particular, the tutorial provides an in-depth architectural overview of a web search engine, mainly focusing on the web crawling, indexing, and query processing components. The scalability and efficiency issues encountered in the above-mentioned components are presented at four different granularities: at the level of a single computer, a cluster of computers, a single data center, and a multi-center search engine. The tutorial also points at the open research problems and provides recommendations to researchers who are new to the field.
Berkant Barla Cambazoglu, Ricardo Baeza-Yates
WSDM1
2015 Task allocation in volunteer computing networks under monetary budget constraints
Huseyin Guler, Berkant Barla Cambazoglu, Öznur Özkasap
Peer-to-Peer Netw. Appl.2
2014 Impact of response latency on user behavior in web search
abstract
Traditionally, the efficiency and effectiveness of search systems have both been of great interest to the information retrieval community. However, an in-depth analysis on the interplay between the response latency of web search systems and users' search experience has been missing so far. In order to fill this gap, we conduct two separate studies aiming to reveal how response latency affects the user behavior in web search. First, we conduct a controlled user study trying to understand how users perceive the response latency of a search system and how sensitive they are to increasing delays in response. This study reveals that, when artificial delays are introduced into the response, the users of a fast search system are more likely to notice these delays than the users of a slow search system. The introduced delays become noticeable by the users once they exceed a certain threshold value. Second, we perform an analysis using a large-scale query log obtained from Yahoo web search to observe the potential impact of increasing response latency on the click behavior of users. This analysis demonstrates that latency has an impact on the click behavior of users to some extent. In particular, given two content-wise identical search result pages, we show that the users are more likely to perform clicks on the result page that is served with lower latency.
Ioannis Arapakis, Xiao Bai 0002, Berkant Barla Cambazoglu
SIGIR3
2014 Scalability and efficiency challenges in large-scale web search engines
abstract
Large-scale web search engines rely on massive compute infrastructures to be able to cope with the continuous growth of the Web and their user bases. In such search engines, achieving scalability and efficiency requires making careful architectural design choices while devising algorithmic performance optimizations. Unfortunately, most details about the internal functioning of commercial web search engines remain undisclosed due to their financial value and the high level of competition in the search market. The main objective of this tutorial is to provide an overview of the fundamental scalability and efficiency challenges in commercial web search engines, bridging the existing gap between the industry and academia.
Berkant Barla Cambazoglu, Ricardo Baeza-Yates
SIGIR1
2014 Workshop on large-scale and distributed systems for information retrieval (LSDS-IR 2014)
abstract
The LSDS-IR'14 workshop aims to bring together information retrieval practitioners from industry and academic researchers concerned with efficient and distributed IR systems. The workshop also welcomes contributions that propose different ways of leveraging diversity and multiplicity of resources available in distributed systems. The main goal of the workshop is to attract people from industry and academia to present and discuss ideas, problems, and results related to the efficiency of large scale and distributed information retrieval systems.
Ismail Sengör Altingövde, Berkant Barla Cambazoglu, Craig Macdonald, Nicola Tonellotto
WSDM2
2014 Improving the efficiency of multi-site web search engines
abstract
A multi-site web search engine is composed of a number of search sites geographically distributed around the world. Each search site is typically responsible for crawling and indexing the web pages that are in its geographical neighborhood. A query is selectively processed on a subset of search sites that are predicted to return the best-matching results. The scalability and efficiency of multi-site web search engines have attracted a lot of research attention in recent years. In particular, research has focused on replicating important web pages across sites, forwarding queries to relevant sites, and caching results of previous queries. Yet, these problems have only been studied in isolation, but no prior work has properly investigated the interplay between them.
Guillem Francès, Xiao Bai 0002, Berkant Barla Cambazoglu, Ricardo Baeza-Yates
WSDM3
2014 User engagement in online News: Under the scope of sentiment, interest, affect, and gaze
abstract
Online content providers, such as news portals and social media platforms, constantly seek new ways to attract large shares of online attention by keeping their users engaged. A common challenge is to identify which aspects of online interaction influence user engagement the most. In this article, through an analysis of a news article collection obtained from Yahoo News US, we demonstrate that news articles exhibit considerable variation in terms of the sentimentality and polarity of their content, depending on factors such as news provider and genre. Moreover, through a laboratory study, we observe the effect of sentimentality and polarity of news and comments on a set of subjective and objective measures of engagement. In particular, we show that attention, affect, and gaze differ across news of varying interestingness. As part of our study, we also explore methods that exploit the sentiments expressed in user comments to reorder the lists of comments displayed in news pages. Our results indicate that user engagement can be anticipated predicted if we account for the sentimentality and polarity of the content as well as other factors that drive attention and inspire human curiosity.
Ioannis Arapakis, Mounia Lalmas-Roelleke, Berkant Barla Cambazoglu, Mari-Carmen Marcos, Joemon M. Jose
J. Assoc. Inf. Sci. Technol.3
2014 Improving the Performance of IndependentTask Assignment Heuristics MinMin, MaxMin and Sufferage
abstract
MinMin, MaxMin, and Sufferage are constructive heuristics that are widely and successfully used in assigning independent tasks to processors in heterogeneous computing systems. All three heuristics are known to run in O(KN2) time in assigning N tasks to K processors. In this paper, we propose an algorithmic improvement that asymptotically decreases the running time complexity of MinMin to O(KN log N) without affecting its solution quality. Furthermore, we combine the newly proposed MinMin algorithm with MaxMin as well as Sufferage, obtaining two hybrid algorithms. The motivation behind the former hybrid algorithm is to address the drawback of MaxMin in solving problem instances with highly skewed cost distributions while also improving the running time performance of MaxMin. The latter hybrid algorithm improves the running time performance of Sufferage without degrading its solution quality. The proposed algorithms are easy to implement and we illustrate them through detailed pseudocodes. The experimental results over a large number of real-life data sets show that the proposed fast MinMin algorithm and the proposed hybrid algorithms perform significantly better than their traditional counterparts as well as more recent state-of-the-art assignment heuristics. For the large data sets used in the experiments, MinMin, MaxMin, and Sufferage, as well as recent state-of-the-art heuristics, require days, weeks, or even months to produce a solution, whereas all of the proposed algorithms produce solutions within only two or three minutes.
E. Kartal Tabak, Berkant Barla Cambazoglu, Cevdet Aykanat
IEEE Trans. Parallel Distributed Syst.2
2014 Sentiment-Focused Web Crawling
abstract
Sentiments and opinions expressed in Web pages towards objects, entities, and products constitute an important portion of the textual content available in the Web. In the last decade, the analysis of such content has gained importance due to its high potential for monetization. Despite the vast interest in sentiment analysis, somewhat surprisingly, the discovery of sentimental or opinionated Web content is mostly ignored. This work aims to fill this gap and addresses the problem of quickly discovering and fetching the sentimental content present in the Web. To this end, we design a sentiment-focused Web crawling framework. In particular, we propose different sentiment-focused Web crawling strategies that prioritize discovered URLs based on their predicted sentiment scores. Through simulations, these strategies are shown to achieve considerable performance improvement over general-purpose Web crawling strategies in discovery of sentimental Web content.
A. Gural Vural, Berkant Barla Cambazoglu, Pinar Karagöz
ACM Trans. Web2
2013 Incorporating the surfing behavior of web users into pagerank
abstract
In large-scale commercial web search engines, estimating the importance of a web page is a crucial ingredient in ranking web search results. So far, to assess the importance of web pages, two different types of feedback have been taken into account, independent of each other: the feedback obtained from the hyperlink structure among the web pages (e.g., PageRank) or the web browsing patterns of users (e.g., BrowseRank). Unfortunately, both types of feedback have certain drawbacks. While the former lacks the user preferences and is vulnerable to malicious intent, the latter suffers from sparsity and hence low web coverage. In this work, we combine these two types of feedback under a hybrid page ranking model in order to alleviate the above-mentioned drawbacks. Our empirical results indicate that the proposed model leads to better estimation of page importance according to an evaluation metric that relies on user click feedback obtained from web search query logs. We conduct all of our experiments in a realistic setting, using a very large scale web page collection (around 6.5 billion web pages) and web browsing data (around two billion web page visits).
Shatlyk Ashyralyyev, Berkant Barla Cambazoglu, Cevdet Aykanat
CIKM2
2013 Strategies for setting time-to-live values in result caches
abstract
In web query result caching, staleness of queries are often bounded via a time-to-live (TTL) mechanism, which expires the validity of cached query results at some point in time. In this work, we evaluate the performance of three alternative TTL mechanisms: time-based TTL, frequency-based TTL, and click-based TTL. Moreover, we propose hybrid approaches obtained by pair-wise combination of these mechanisms. Our results indicate that combining time-based TTL with frequency-based TTL yields superior performance (i.e., lower stale query traffic and less redundant computation) than using a particular mechanism in isolation.
Fethi Burak Sazoglu, Berkant Barla Cambazoglu, Rifat Ozcan, Ismail Sengör Altingövde, Özgür Ulusoy
CIKM2
2013 Entity Recommendations in Web Search
Roi Blanco, Berkant Barla Cambazoglu, Peter Mika, Nicolas Torzec
ISWC (2)2
2013 Scalability and efficiency challenges in commercial web search engines
abstract
Commercial web search engines rely on very large compute infrastructures to be able to cope with the continuous growth of the Web and user bases. Achieving scalability and efficiency in such large-scale search engines requires making careful architectural design choices while devising algorithmic performance optimizations. Unfortunately, most details about the internal functioning of commercial web search engines remain undisclosed due to their financial value and the high level of competition in the search market. The main objective of this tutorial is to provide an overview of the fundamental scalability and efficiency challenges in commercial web search engines, bridging the existing gap between the industry and academia.
Berkant Barla Cambazoglu, Ricardo Baeza-Yates
SIGIR1
2013 A financial cost metric for result caching
abstract
Web search engines cache results of frequent and/or recent queries. Result caching strategies can be evaluated using different metrics, hit rate being the most well-known. Recent works take the processing overhead of queries into account when evaluating the performance of result caching strategies and propose cost-aware caching strategies. In this paper, we propose a financial cost metric that goes one step beyond and takes also the hourly electricity prices into account when computing the cost. We evaluate the most well-known static, dynamic, and hybrid result caching strategies under this new metric. Moreover, we propose a financial-cost-aware version of the well-known LRU strategy and show that it outperforms the original LRU strategy in terms of the financial cost metric.
Fethi Burak Sazoglu, Berkant Barla Cambazoglu, Rifat Ozcan, Ismail Sengör Altingövde, Özgür Ulusoy
SIGIR2
2013 Document replication strategies for geographically distributed web search engines
Enver Kayaaslan, Berkant Barla Cambazoglu, Cevdet Aykanat
Inf. Process. Manag.2
2013 A term-based inverted index partitioning model for efficient distributed query processing
abstract
In a shared-nothing, distributed text retrieval system, queries are processed over an inverted index that is partitioned among a number of index servers. In practice, the index is either document-based or term-based partitioned. This choice is made depending on the properties of the underlying hardware infrastructure, query traffic distribution, and some performance and availability constraints. In query processing on retrieval systems that adopt a term-based index partitioning strategy, the high communication overhead due to the transfer of large amounts of data from the index servers forms a major performance bottleneck, deteriorating the scalability of the entire distributed retrieval system. In this work, to alleviate this problem, we propose a novel inverted index partitioning model that relies on hypergraph partitioning. In the proposed model, concurrently accessed index entries are assigned to the same index servers, based on the inverted index access patterns extracted from the past query logs. The model aims to minimize the communication overhead that will be incurred by future queries while maintaining the computational load balance among the index servers. We evaluate the performance of the proposed model through extensive experiments using a real-life text collection and a search query sample. Our results show that considerable performance gains can be achieved relative to the term-based index partitioning strategies previously proposed in literature. In most cases, however, the performance remains inferior to that attained by document-based partitioning.
Berkant Barla Cambazoglu, Enver Kayaaslan, Simon Jonassen, Cevdet Aykanat
ACM Trans. Web1
2013 Second Chance: A Hybrid Approach for Dynamic Result Caching and Prefetching in Search Engines
abstract
Web search engines are known to cache the results of previously issued queries. The stored results typically contain the document summaries and some data that is used to construct the final search result page returned to the user. An alternative strategy is to store in the cache only the result document IDs, which take much less space, allowing results of more queries to be cached. These two strategies lead to an interesting trade-off between the hit rate and the average query response latency. In this work, in order to exploit this trade-off, we propose a hybrid result caching strategy where a dynamic result cache is split into two sections: an HTML cache and a docID cache. Moreover, using a realistic cost model, we evaluate the performance of different result prefetching strategies for the proposed hybrid cache and the baseline HTML-only cache. Finally, we propose a machine learning approach to predict singleton queries, which occur only once in the query stream. We show that when the proposed hybrid result caching strategy is coupled with the singleton query predictor, the hit rate is further improved.
Rifat Ozcan, Ismail Sengör Altingövde, Berkant Barla Cambazoglu, Özgür Ulusoy
ACM Trans. Web3
2012 Characterizing web search queries that match very few or no results
abstract
Despite the continuous efforts to improve the web search quality, a non-negligible fraction of user queries end up with very few or even no matching results in leading web search engines. In this work, we provide a detailed characterization of such queries based on an analysis of a real-life query log. Our experimental setup allows us to characterize the queries with few/no results and compare the mechanisms employed by the major search engines in handling them.
Ismail Sengör Altingövde, Roi Blanco, Berkant Barla Cambazoglu, Rifat Ozcan, Erdem Sarigil, Özgür Ulusoy
CIKM3
2012 Sentiment-focused web crawling
abstract
The sentiments and opinions that are expressed in web pages towards objects, entities, and products constitute an important portion of the textual content available in the Web. Despite the vast interest in sentiment analysis and opinion mining, somewhat surprisingly, the discovery of the sentimental or opinionated web content is mostly ignored. This work aims to fill this gap and address the problem of quickly discovering and fetching the sentimental content present in the Web. To this end, we design a sentiment-focused web crawling framework for faster discovery and retrieval of such content. In particular, we propose different sentiment-focused web crawling strategies that prioritize discovered URLs based on their predicted sentiment scores. Through simulations, these strategies are shown to achieve considerable performance improvement over general-purpose web crawling strategies in discovering sentimental content.
A. Gural Vural, Berkant Barla Cambazoglu, Pinar Karagöz
CIKM2
2012 Adaptive Time-to-Live Strategies for Query Result Caching in Web Search Engines
Sadiye Alici, Ismail Sengör Altingövde, Rifat Ozcan, Berkant Barla Cambazoglu, Özgür Ulusoy
ECIR4
2012 Prefetching query results and its impact on search engines
abstract
We investigate the impact of query result prefetching on the efficiency and effectiveness of web search engines. We propose offline and online strategies for selecting and ordering queries whose results are to be prefetched. The offline strategies rely on query log analysis and the queries are selected from the queries issued on the previous day. The online strategies select the queries from the result cache, relying on a machine learning model that estimates the arrival times of queries. We carefully evaluate the proposed prefetching techniques via simulation on a query log obtained from Yahoo! web search. We demonstrate that our strategies are able to improve various performance metrics, including the hit rate, query response time, result freshness, and query degradation rate, relative to a state-of-the-art baseline.
Simon Jonassen, Berkant Barla Cambazoglu, Fabrizio Silvestri
SIGIR2
2012 Impact of Regionalization on Performance of Web Search Engine Result Caches
Berkant Barla Cambazoglu, Ismail Sengör Altingövde
SPIRE1
2012 A large-scale sentiment analysis for Yahoo! answers
abstract
Sentiment extraction from online web documents has recently been an active research topic due to its potential use in commercial applications. By sentiment analysis, we refer to the problem of assigning a quantitative positive/negative mood to a short bit of text. Most studies in this area are limited to the identification of sentiments and do not investigate the interplay between sentiments and other factors. In this work, we use a sentiment extraction tool to investigate the influence of factors such as gender, age, education level, the topic at hand, or even the time of the day on sentiments in the context of a large online question answering site. We start our analysis by looking at direct correlations, e.g., we observe more positive sentiments on weekends, very neutral ones in the Science & Mathematics topic, a trend for younger people to express stronger sentiments, or people in military bases to ask the most neutral questions. We then extend this basic analysis by investigating how properties of the (asker, answerer) pair affect the sentiment present in the answer. Among other things, we observe a dependence on the pairing of some inferred attributes estimated by a user's ZIP code. We also show that the best answers differ in their sentiments from other answers, e.g., in the Business & Finance topic, best answers tend to have a more neutral sentiment than other answers. Finally, we report results for the task of predicting the attitude that a question will provoke in answers. We believe that understanding factors influencing the mood of users is not only interesting from a sociological point of view, but also has applications in advertising, recommendation, and search.
Onur Küçüktunç, Berkant Barla Cambazoglu, Ingmar Weber, Hakan Ferhatosmanoglu
WSDM2
2012 A five-level static cache architecture for web search engines
Rifat Ozcan, Ismail Sengör Altingövde, Berkant Barla Cambazoglu, Flavio Paiva Junqueira, Özgür Ulusoy
Inf. Process. Manag.3
2012 Cache-Based Query Processing for Search Engines
abstract
In practice, a search engine may fail to serve a query due to various reasons such as hardware/network failures, excessive query load, lack of matching documents, or service contract limitations (e.g., the query rate limits for third-party users of a search service). In this kind of scenarios, where the backend search system is unable to generate answers to queries, approximate answers can be generated by exploiting the previously computed query results available in the result cache of the search engine. In this work, we propose two alternative strategies to implement this cache-based query processing idea. The first strategy aggregates the results of similar queries that are previously cached in order to create synthetic results for new queries. The second strategy forms an inverted index over the textual information (i.e., query terms and result snippets) present in the result cache and uses this index to answer new queries. Both approaches achieve reasonable result qualities compared to processing queries with an inverted index built on the collection.
Berkant Barla Cambazoglu, Ismail Sengör Altingövde, Rifat Ozcan, Özgür Ulusoy
ACM Trans. Web1
2011 Discovering URLs through user feedback
abstract
Search engines rely upon crawling to build their Web page collections. A Web crawler typically discovers new URLs by following the link structure induced by links on Web pages. As the number of documents on the Web is large, discovering newly created URLs may take arbitrarily long, and depending on how a given page is connected to others, such a crawler may miss the pages altogether. In this paper, we evaluate the benefits of integrating a passive URL discovery mechanism into a Web crawler. This mechanism is passive in the sense that it does not require the crawler to actively fetch documents from the Web to discover URLs. We focus here on a mechanism that uses toolbar data as a representative source for new URL discovery. We use the toolbar logs of Yahoo! to characterize the URLs that are accessed by users via their browsers, but not discovered by Yahoo! Web crawler. We show that a high fraction of URLs that appear in toolbar logs are not discovered by the crawler. We also reveal that a certain fraction of URLs are discovered by the crawler later than the time they are first accessed by users. One important conclusion of our work is that web search engines can highly benefit from user feedback in the form of toolbar logs for passive URL discovery.
Xiao Bai 0002, Berkant Barla Cambazoglu, Flavio Paiva Junqueira
CIKM2
2011 Assigning documents to master sites in distributed search
abstract
An appealing solution to scale Web search with the growth of the Internet is the use of distributed architectures. Distributed search engines rely on multiple sites deployed in distant regions across the world, where each site is specialized to serve queries issued by the users of its region. This paper investigates the problem of assigning each document to a master site. We show that by leveraging similarities between a document and the activity of the users, we can accurately detect which site is the most relevant to place a document. We conduct various experiments using two document assignment approaches, showing performance improvements of up to 20.8% over a baseline technique which assigns the documents to search sites based on their language.
Roi Blanco, Berkant Barla Cambazoglu, Flavio Paiva Junqueira, Ivan Kelly, Vincent Leroy 0001
CIKM2
2011 LSDS-IR'11: the 9th workshop on large-scale and distributed systems for information retrieval
abstract
The growth of the Web and user bases lead to important performance problems for large-scale Web search engines. The LSDS- IR '11 workshop focuses on research contributions related to the scalability and efficiency of distributed information retrieval (IR) systems. The workshop also encourages contributions that propose different ways of leveraging diversity and multiplicity of resources available in distributed systems. More specifically, we are interested in novel applications, models, and architectures that deal with efficiency and scalability of distributed IR systems.
Claudio Lucchese, Berkant Barla Cambazoglu
CIKM2
2011 Second Chance: A Hybrid Approach for Dynamic Result Caching in Search Engines
Ismail Sengör Altingövde, Rifat Ozcan, Berkant Barla Cambazoglu, Özgür Ulusoy
ECIR3
2011 Machine learned job recommendation
abstract
We address the problem of recommending suitable jobs to people who are seeking a new job. We formulate this recommendation problem as a supervised machine learning problem. Our technique exploits all past job transitions as well as the data associated with employees and institutions to predict an employee's next job transition. We train a machine learning model using a large number of job transitions extracted from the publicly available employee profiles in the Web. Experiments show that job transitions can be accurately predicted, significantly improving over a baseline that always predicts the most frequent institution in the data.
Ioannis K. Paparrizos, Berkant Barla Cambazoglu, Aristides Gionis
RecSys2
2011 Timestamp-based result cache invalidation for web search engines
abstract
The result cache is a vital component for efficiency of large-scale web search engines, and maintaining the freshness of cached query results is the current research challenge. As a remedy to this problem, our work proposes a new mechanism to identify queries whose cached results are stale. The basic idea behind our mechanism is to maintain and compare generation time of query results with update times of posting lists and documents to decide on staleness of query results. The proposed technique is evaluated using a Wikipedia document collection with real update information and a real-life query log. We show that our technique has good prediction accuracy, relative to a baseline based on the time-to-live mechanism. Moreover, it is easy to implement and incurs less processing overhead on the system relative to a recently proposed, more sophisticated invalidation mechanism.
Sadiye Alici, Ismail Sengör Altingövde, Rifat Ozcan, Berkant Barla Cambazoglu, Özgür Ulusoy
SIGIR4
2011 Energy-price-driven query processing in multi-center web search engines
abstract
Concurrently processing thousands of web queries, each with a response time under a fraction of a second, necessitates maintaining and operating massive data centers. For large-scale web search engines, this translates into high energy consumption and a huge electric bill. This work takes the challenge to reduce the electric bill of commercial web search engines operating on data centers that are geographically far apart. Based on the observation that energy prices and query workloads show high spatio-temporal variation, we propose a technique that dynamically shifts the query workload of a search engine between its data centers to reduce the electric bill. Experiments on real-life query workloads obtained from a commercial search engine show that significant financial savings can be achieved by this technique.
Enver Kayaaslan, Berkant Barla Cambazoglu, Roi Blanco, Flavio Paiva Junqueira, Cevdet Aykanat
SIGIR2
2011 Posting list intersection on multicore architectures
abstract
In current commercial Web search engines, queries are processed in the conjunctive mode, which requires the search engine to compute the intersection of a number of posting lists to determine the documents matching all query terms. In practice, the intersection operation takes a significant fraction of the query processing time, for some queries dominating the total query latency. Hence, efficient posting list intersection is critical for achieving short query latencies. In this work, we focus on improving the performance of posting list intersection by leveraging the compute capabilities of recent multicore systems. To this end, we consider various coarse-grained and fine-grained parallelization models for list intersection. Specifically, we present an algorithm that partitions the work associated with a given query into a number of small and independent tasks that are subsequently processed in parallel. Through a detailed empirical analysis of these alternative models, we demonstrate that exploiting parallelism at the finest-level of granularity is critical to achieve the best performance on multicore systems. On an eight-core system, the fine-grained parallelization method is able to achieve more than five times reduction in average query processing time while still exploiting the parallelism for high query throughput.
Shirish Tatikonda, Berkant Barla Cambazoglu, Flavio Paiva Junqueira
SIGIR2
2011 Document assignment in multi-site search engines
abstract
Assigning documents accurately to sites is critical for the performance of multi-site Web search engines. In such settings, sites crawl only documents they index and forward queries to obtain best-matching documents from other sites. Inaccurate assignments may lead to inefficiencies when crawling Web pages or processing user queries. In this work, we propose a machine-learned document assignment strategy that uses the locality of document views in search results to decide upon assignments. We evaluate the performance of our strategy using various document features extracted from a large Web collection. Our experimental setup uses query logs from a number of search front-ends spread across different geographic locations and uses these logs to learn the document access patterns. We compare our technique against baselines such as region- and language-based document assignment and observe that our technique achieves substantial performance improvements with respect to recall. With our technique, we are able to obtain a small query forwarding rate (0.04) requiring roughly 45% less replication of documents compared to replicating all documents across all sites.
Ulf Brefeld, Berkant Barla Cambazoglu, Flavio Paiva Junqueira
WSDM2
2011 Site-Based Partitioning and Repartitioning Techniques for Parallel PageRank Computation
abstract
The PageRank algorithm is an important component in effective web search. At the core of this algorithm are repeated sparse matrix-vector multiplications where the involved web matrices grow in parallel with the growth of the web and are stored in a distributed manner due to space limitations. Hence, the PageRank computation, which is frequently repeated, must be performed in parallel with high-efficiency and low-preprocessing overhead while considering the initial distributed nature of the web matrices. Our contributions in this work are twofold. We first investigate the application of state-of-the-art sparse matrix partitioning models in order to attain high efficiency in parallel PageRank computations with a particular focus on reducing the preprocessing overhead they introduce. For this purpose, we evaluate two different compression schemes on the web matrix using the site information inherently available in links. Second, we consider the more realistic scenario of starting with an initially distributed data and extend our algorithms to cover the repartitioning of such data for efficient PageRank computation. We report performance results using our parallelization of a state-of-the-art PageRank algorithm on two different PC clusters with 40 and 64 processors. Experiments show that the proposed techniques achieve considerably high speedups while incurring a preprocessing overhead of several iterations (for some instances even less than a single iteration) of the underlying sequential PageRank algorithm.
Ali Cevahir, Cevdet Aykanat, Ata Turk, Berkant Barla Cambazoglu
IEEE Trans. Parallel Distributed Syst.4
2010 Web search solved?: all result rankings the same?
abstract
The objective of this work is to derive quantitative statements about what fraction of web search queries issued to the state-of-the-art commercial search engines lead to excellent results or, on the contrary, poor results. To be able to make such statements in an automated way, we propose a new measure that is based on lower and upper bound analysis over the standard relevance measures. Moreover, we extend this measure to carry out comparisons between competing search engines by introducing the concept of disruptive sets, which we use to estimate the degree to which a search engine solves queries that are not solved by its competitors. We report empirical results on a large editorial evaluation of the three largest search engines in the US market.
Hugo Zaragoza, Berkant Barla Cambazoglu, Ricardo Baeza-Yates
CIKM2
2010 Cold start link prediction
abstract
In the traditional link prediction problem, a snapshot of a social network is used as a starting point to predict, by means of graph-theoretic measures, the links that are likely to appear in the future. In this paper, we introduce cold start link prediction as the problem of predicting the structure of a social network when the network itself is totally missing while some other information regarding the nodes is available. We propose a two-phase method based on the bootstrap probabilistic graph. The first phase generates an implicit social network under the form of a probabilistic graph. The second phase applies probabilistic graph-based measures to produce the final prediction. We assess our method empirically over a large data collection obtained from Flickr, using interest groups as the initial information. The experiments confirm the effectiveness of our approach.
Vincent Leroy 0001, Berkant Barla Cambazoglu, Francesco Bonchi
KDD2
2010 Query forwarding in geographically distributed search engines
abstract
Query forwarding is an important technique for preserving the result quality in distributed search engines where the index is geographically partitioned over multiple search sites. The key component in query forwarding is the thresholding algorithm by which the forwarding decisions are given. In this paper, we propose a linear-programming-based thresholding algorithm that significantly outperforms the current state-of-the-art in terms of achieved search efficiency values. Moreover, we evaluate a greedy heuristic for partial index replication and investigate the impact of result cache freshness on query forwarding performance. Finally, we present some optimizations that improve the performance further, under certain conditions. We evaluate the proposed techniques by simulations over a real-life setting, using a large query log and a document collection obtained from Yahoo!.
Berkant Barla Cambazoglu, Emre Varol, Enver Kayaaslan, Cevdet Aykanat, Ricardo Baeza-Yates
SIGIR1
2010 Early exit optimizations for additive machine learned ranking systems
abstract
Some commercial web search engines rely on sophisticated machine learning systems for ranking web documents. Due to very large collection sizes and tight constraints on query response times, online efficiency of these learning systems forms a bottleneck. An important problem in such systems is to speedup the ranking process without sacrificing much from the quality of results. In this paper, we propose optimization strategies that allow short-circuiting score computations in additive learning systems. The strategies are evaluated over a state-of-the-art machine learning system and a large, real-life query log, obtained from Yahoo!. By the proposed strategies, we are able to speedup the score computations by more than four times with almost no loss in result quality.
Berkant Barla Cambazoglu, Hugo Zaragoza, Olivier Chapelle, Ciya Liao, Zhaohui Zheng 0001, Jon Degenhardt
WSDM1
2010 A refreshing perspective of search engine caching
abstract
Commercial Web search engines have to process user queries over huge Web indexes under tight latency constraints. In practice, to achieve low latency, large result caches are employed and a portion of the query traffic is served using previously computed results. Moreover, search engines need to update their indexes frequently to incorporate changes to the Web. After every index update, however, the content of cache entries may become stale, thus decreasing the freshness of served results. In this work, we first argue that the real problem in today's caching for large-scale search engines is not eviction policies, but the ability to cope with changes to the index, i.e., cache freshness. We then introduce a novel algorithm that uses a time-to-live value to set cache entries to expire and selectively refreshes cached results by issuing refresh queries to back-end search clusters. The algorithm prioritizes the entries to refresh according to a heuristic that combines the frequency of access with the age of an entry in the cache. In addition, for setting the rate at which refresh queries are issued, we present a mechanism that takes into account idle cycles of back-end servers. Evaluation using a real workload shows that our algorithm can achieve hit rate improvements as well as reduction in average hit ages. An implementation of this algorithm is currently in production use at Yahoo!.
Berkant Barla Cambazoglu, Flavio Paiva Junqueira, Vassilis Plachouras, Scott A. Banachowski, Baoqiu Cui, Swee Lim, Bill Bridge
WWW1
2010 Review of "Search Engines: Information Retrieval in Practice" by Croft, Metzler and Strohman
Berkant Barla Cambazoglu
Inf. Process. Manag.1
2010 A link-based storage scheme for efficient aggregate query processing on clustered road networks
Engin Demir, Cevdet Aykanat, Berkant Barla Cambazoglu
Inf. Syst.3
2009 Quantifying performance and quality gains in distributed web search engines
abstract
Distributed search engines based on geographical partitioning of a central Web index emerge as a feasible solution to the immense growth of the Web, user bases, and query traffic. However, there is still lack of research in quantifying the performance and quality gains that can be achieved by such architectures. In this paper, we develop various cost models to evaluate the performance benefits of a geographically distributed search engine architecture based on partial index replication and query forwarding. Specifically, we focus on possible performance gains due to the distributed nature of query processing and Web crawling processes. We show that any response time gain achieved by distributed query processing can be utilized to improve search relevance as the use of complex but more accurate algorithms can now be enabled for document ranking. We also show that distributed Web crawling leads to better Web coverage and try to see if this improves the search quality. We verify the validity of our claims over large, real-life datasets via simulations.
Berkant Barla Cambazoglu, Vassilis Plachouras, Ricardo Baeza-Yates
SIGIR1
2009 On efficient posting list intersection with multicore processors
abstract
No abstract available.
Shirish Tatikonda, Flavio Paiva Junqueira, Berkant Barla Cambazoglu, Vassilis Plachouras
SIGIR3
2008 Chat mining: Predicting user and message attributes in computer-mediated communication
Tayfun Küçükyilmaz, Berkant Barla Cambazoglu, Cevdet Aykanat, Fazli Can
Inf. Process. Manag.2
2008 Clustering spatial networks for aggregate query processing: A hypergraph approach
Engin Demir, Cevdet Aykanat, Berkant Barla Cambazoglu
Inf. Syst.3
2008 Model Formulation: Sharing Data and Analytical Resources Securely in a Biomedical Research Grid Environment
abstract
OBJECTIVES: To develop a security infrastructure to support controlled and secure access to data and analytical resources in a biomedical research Grid environment, while facilitating resource sharing among collaborators. DESIGN: A Grid security infrastructure, called Grid Authentication and Authorization with Reliably Distributed Services (GAARDS), is developed as a key architecture component of the NCI-funded cancer Biomedical Informatics Grid (caBIG). The GAARDS is designed to support in a distributed environment 1) efficient provisioning and federation of user identities and credentials; 2) group-based access control support with which resource providers can enforce policies based on community accepted groups and local groups; and 3) management of a trust fabric so that policies can be enforced based on required levels of assurance. MEASUREMENTS: GAARDS is implemented as a suite of Grid services and administrative tools. It provides three core services: Dorian for management and federation of user identities, Grid Trust Service for maintaining and provisioning a federated trust fabric within the Grid environment, and Grid Grouper for enforcing authorization policies based on both local and Grid-level groups. RESULTS: The GAARDS infrastructure is available as a stand-alone system and as a component of the caGrid infrastructure. More information about GAARDS can be accessed at http://www.cagrid.org. CONCLUSIONS: GAARDS provides a comprehensive system to address the security challenges associated with environments in which resources may be located at different sites, requests to access the resources may cross institutional boundaries, and user credentials are created, managed, revoked dynamically in a de-centralized manner.
Stephen Langella, Shannon Hastings, Scott Oster, Tony Pan, Ashish Sharma 0001, Justin Permar, David Ervin, Berkant Barla Cambazoglu, Tahsin M. Kurç, Joel H. Saltz
J. Am. Medical Informatics Assoc.8
2008 Multi-level direct K-way hypergraph partitioning with multiple constraints and fixed vertices
Cevdet Aykanat, Berkant Barla Cambazoglu, Bora Uçar
J. Parallel Distributed Comput.2
2007 Computerized Pathological Image Analysis For Neuroblastoma Prognosis
Metin Nafi Gürcan, Jun Kong 0002, Olcay Sertel, Berkant Barla Cambazoglu, Joel H. Saltz, Ümit V. Çatalyürek
AMIA4
2007 Architecture of a grid-enabled Web search engine
Berkant Barla Cambazoglu, Evren Karaca, Tayfun Küçükyilmaz, Ata Turk, Cevdet Aykanat
Inf. Process. Manag.1
2007 Adaptive decomposition and remapping algorithms for object-space-parallel direct volume rendering of unstructured grids
Cevdet Aykanat, Berkant Barla Cambazoglu, Ferit Findik, Tahsin M. Kurç
J. Parallel Distributed Comput.2
2007 Hypergraph-Partitioning-Based Remapping Models for Image-Space-Parallel Direct Volume Rendering of Unstructured Grids
Berkant Barla Cambazoglu, Cevdet Aykanat
IEEE Trans. Parallel Distributed Syst.1
2006 Performance of query processing implementations in ranking-based text retrieval systems using inverted indices
Berkant Barla Cambazoglu, Cevdet Aykanat
Inf. Process. Manag.1