EDBT 2026 Demo / reviewers in the wild / expert
Berkant Barla Cambazoglu
dblp:57/2006
· DBLP profile ↗
72ranked-venue papers
17as first author
2since 2021 · last 2021
0000-0003-2192-3819ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 62 · 16 first-author · 2 since 2021Artificial intelligence and machine learning · 20 · 2 first-authorSystems, architecture and hardware · 5 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 first-authorComputer networks · 2Human-computer interaction and ubiquitous computing · 2 · 2 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
26 papers |
Information retrieval · 85% Query processing and optimization · 6% Indexing and storage engines · 4% | |
| Computer architecture, parallel and distributed computing, and storage systems
11 papers |
Parallel and multicore computing · 40% Distributed systems · 32% Cloud and datacenter computing · 10% |
Topics — the 30 heaviest of 57, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval
search engines |
1.4 | 9 | 2016 | Scalability and Efficiency Challenges in Large-Scale Web Search Engines · SIGIR 2016 Scalability and Efficiency Challenges in Large-Scale Web Search Engines · WSDM 2015 Unconscious Physiological Effects of Search Latency on Users and Their Click Behaviour · SIGIR 2015 |
Information retrieval
web search |
0.9 | 6 | 2017 | Understanding and Leveraging the Impact of Response Latency on User Behaviour in Web Search · ACM Trans. Inf. Syst. 2017 Improving the efficiency of multi-site web search engines · WSDM 2014 Prefetching query results and its impact on search engines · SIGIR 2012 |
Information retrieval
user behavior |
0.7 | 3 | 2017 | Understanding and Leveraging the Impact of Response Latency on User Behaviour in Web Search · ACM Trans. Inf. Syst. 2017 Unconscious Physiological Effects of Search Latency on Users and Their Click Behaviour · SIGIR 2015 Impact of response latency on user behavior in web search · SIGIR 2014 |
Information retrieval › search engines
web search engine |
0.6 | 3 | 2016 | Scalability and Efficiency Challenges in Large-Scale Web Search Engines · SIGIR 2016 Scalability and efficiency challenges in large-scale web search engines · SIGIR 2014 Scalability and efficiency challenges in commercial web search engines · SIGIR 2013 |
Information retrieval › search engines
web crawling |
0.5 | 3 | 2016 | Optimal Web Page Download Scheduling Policies for Green Web Crawling · IEEE J. Sel. Areas Commun. 2016 A Random Walk Model for Optimization of Search Impact in Web Frontier Ranking · SIGIR 2015 Quantifying performance and quality gains in distributed web search engines · SIGIR 2009 |
Information retrieval › search engines
search engine architecture |
0.5 | 3 | 2015 | Scalability and Efficiency Challenges in Large-Scale Web Search Engines · WSDM 2015 Timestamp-based result cache invalidation for web search engines · SIGIR 2011 A refreshing perspective of search engine caching · WWW 2010 |
Query processing and optimization
query result caching |
0.4 | 3 | 2013 | A financial cost metric for result caching · SIGIR 2013 Prefetching query results and its impact on search engines · SIGIR 2012 Timestamp-based result cache invalidation for web search engines · SIGIR 2011 |
Information retrieval › user behavior › search behavior
click behavior |
0.4 | 2 | 2017 | Understanding and Leveraging the Impact of Response Latency on User Behaviour in Web Search · ACM Trans. Inf. Syst. 2017 Unconscious Physiological Effects of Search Latency on Users and Their Click Behaviour · SIGIR 2015 |
Information retrieval
query processing |
0.3 | 4 | 2017 | Posting list intersection on multicore architectures · SIGIR 2011 On efficient posting list intersection with multicore processors · SIGIR 2009 Understanding and Leveraging the Impact of Response Latency on User Behaviour in Web Search · ACM Trans. Inf. Syst. 2017 |
Information retrieval
distributed information retrieval |
0.3 | 2 | 2014 | Workshop on large-scale and distributed systems for information retrieval (LSDS-IR 2014) · WSDM 2014 Quantifying performance and quality gains in distributed web search engines · SIGIR 2009 |
Indexing and storage engines
caching |
0.2 | 1 | 2016 | Improved Caching Techniques for Large-Scale Image Hosting Services · SIGIR 2016 |
Information retrieval › distributed information retrieval
distributed search |
0.2 | 2 | 2011 | Document assignment in multi-site search engines · WSDM 2011 Query forwarding in geographically distributed search engines · SIGIR 2010 |
Information retrieval › query processing
list intersection |
0.2 | 2 | 2011 | Posting list intersection on multicore architectures · SIGIR 2011 On efficient posting list intersection with multicore processors · SIGIR 2009 |
Information retrieval › evaluation
search effectiveness |
0.2 | 1 | 2015 | A Random Walk Model for Optimization of Search Impact in Web Frontier Ranking · SIGIR 2015 |
Information retrieval
interactive information retrieval |
0.2 | 1 | 2014 | Impact of response latency on user behavior in web search · SIGIR 2014 |
Information retrieval
retrieval evaluation |
0.2 | 1 | 2014 | Impact of response latency on user behavior in web search · SIGIR 2014 |
Distributed systems › peer-to-peer systems
distributed search |
0.2 | 1 | 2014 | Improving the efficiency of multi-site web search engines · WSDM 2014 |
Distributed systems
query forwarding |
0.2 | 1 | 2014 | Improving the efficiency of multi-site web search engines · WSDM 2014 |
Distributed systems
query result caching |
0.2 | 1 | 2014 | Improving the efficiency of multi-site web search engines · WSDM 2014 |
Parallel and multicore computing
scheduling algorithms |
0.2 | 1 | 2014 | Improving the Performance of IndependentTask Assignment Heuristics MinMin, MaxMin and Sufferage · IEEE Trans. Parallel Distributed Syst. 2014 |
Parallel and multicore computing
task allocation |
0.2 | 1 | 2014 | Improving the Performance of IndependentTask Assignment Heuristics MinMin, MaxMin and Sufferage · IEEE Trans. Parallel Distributed Syst. 2014 |
Information retrieval
evaluation |
0.2 | 1 | 2013 | A financial cost metric for result caching · SIGIR 2013 |
Query processing and optimization › runtime optimization › prefetching
query result prefetching |
0.1 | 1 | 2012 | Prefetching query results and its impact on search engines · SIGIR 2012 |
Data mining › text mining
sentiment analysis |
0.1 | 1 | 2012 | A large-scale sentiment analysis for Yahoo! answers · WSDM 2012 |
Database system architecture and tuning
cache invalidation |
0.1 | 1 | 2011 | Timestamp-based result cache invalidation for web search engines · SIGIR 2011 |
Information retrieval › distributed information retrieval
document allocation |
0.1 | 1 | 2011 | Document assignment in multi-site search engines · WSDM 2011 |
Indexing and storage engines
index maintenance |
0.1 | 1 | 2011 | Timestamp-based result cache invalidation for web search engines · SIGIR 2011 |
Cloud and datacenter computing › resource management
datacenter resource management |
0.1 | 1 | 2011 | Energy-price-driven query processing in multi-center web search engines · SIGIR 2011 |
Hardware accelerators and domain-specific architectures › query processing
energy-efficient query processing |
0.1 | 1 | 2011 | Energy-price-driven query processing in multi-center web search engines · SIGIR 2011 |
Parallel and multicore computing › parallelization strategies
fine-grained parallelization |
0.1 | 1 | 2011 | Posting list intersection on multicore architectures · SIGIR 2011 |
Methods — techniques the papers use, named apart from their topics
simulation · 0.7query log analysis · 0.7machine learning · 0.6greedy policy · 0.5gain-based caching · 0.2correlation analysis · 0.2random walk · 0.2physiological measurement · 0.2link graph enrichment · 0.2controlled experiment · 0.2sufferage · 0.2minmin · 0.2maxmin · 0.2hybrid algorithms · 0.2controlled user study · 0.2LRU caching · 0.2workload shifting · 0.1sparse matrix partitioning · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Quantifying Human-Perceived Answer Utility in Non-factoid Question AnsweringabstractTaking a user-centric approach, we study the features that render an answer to a non-factoid question useful in the eyes of the person who asked that question. An editorial study, where participants assess the usefulness of the answers they received in response to their questions, as well as 12 different aspects associated with the answers, indicates considerable correlation between certain aspects such as relevance, correctness, and completeness with the user-perceived usefulness of answers. Moreover, we investigate the effectiveness of some commonly used answer quality measures, such as ROGUE, BLEU, METEOR, and BERTScore, demonstrating that these measures are limited in their ability to capture the aspects of usefulness and have room for improvement. The question answering dataset created in our work was made publicly available. Berkant Barla Cambazoglu, Valeria Bolotova-Baranova, Falk Scholer, Mark Sanderson, Leila Tavakoli, W. Bruce Croft |
CHIIR | 1 |
| 2021 | An Intent Taxonomy for Questions Asked in Web SearchabstractWe present a new, multi-faceted taxonomy to classify questions asked in web search engines based on the question intent, types of entities mentioned, types of question words, and granularity of the expected answer. Built based on the inspection of 1,000 real-life questions issued to a web search engine, the taxonomy reflects the recent search behavior of users and enables deep understanding of user intents, goals, and expected answers. This taxonomy is more fine-grained than previous query taxonomies, and is designed with the ultimate goal of reducing the inherent ambiguity in determining the intent of questions. In addition, we describe the formal procedure for conducting an editorial study of the taxonomy including its evaluation. The adopted procedure aims to increase assessor agreement without incurring too much overhead. Our results demonstrate that, despite being more fine-grained, the proposed intent categories result in higher agreement between assessors compared to an existing, commonly used taxonomy. Berkant Barla Cambazoglu, Leila Tavakoli, Falk Scholer, Mark Sanderson, W. Bruce Croft |
CHIIR | 1 |
| 2020 | Providing Direct Answers in Search Results: A Study of User BehaviorabstractTo study the impact of providing direct answers in search results on user behavior, we conducted a controlled user study to analyze factors including reading time, eye-tracked attention, and the influence of the quality of answer module content. We also studied a more advanced answer interface, where multiple answers are shown on the search engine results page (SERP). Our results show that users focus more extensively than normal on the top items in the result list when answers are provided. The existence of the answer module helps to improve user engagement on SERPs, reduces user effort, and promotes user satisfaction during the search process. Furthermore, we investigate how the question type -- factoid or non-factoid -- affects user interaction patterns. This work provides insight into the design of SERPs that includes direct answers to queries, including when answers should be shown. Zhijing Wu 0001, Mark Sanderson, Berkant Barla Cambazoglu, W. Bruce Croft, Falk Scholer |
CIKM | 3 |
| 2020 | Feature Extraction for Large-Scale Text CollectionsabstractFeature engineering is a fundamental but poorly documented component in Learning-to-Rank (LTR) search engines. Such features are commonly used to construct learning models for web and product search engines, recommender systems, and question-answering tasks. In each of these domains, there is a growing interest in the creation of open-access test collections that promote reproducible research. However, there are still few open-source software packages capable of extracting high-quality machine learning features from large text collections. Instead, most feature-based LTR research relies on "canned" test collections, which often do not expose critical details about the underlying collection or implementation details of the extracted features. Both of these are crucial to collection creation and deployment of a search engine into production. So in this regard, the experiments are rarely reproducible with new features or collections, or helpful for companies wishing to deploy LTR systems. Luke Gallagher, Antonio Mallia, J. Shane Culpepper, Torsten Suel, Berkant Barla Cambazoglu |
CIKM | 5 |
| 2020 | Pre-indexing Pruning Strategies
Soner Altin, Ricardo Baeza-Yates, Berkant Barla Cambazoglu |
SPIRE | 3 |
| 2019 | Impact of response latency on sponsored search
Xiao Bai 0002, Berkant Barla Cambazoglu |
Inf. Process. Manag. | 2 |
| 2018 | Characterizing, predicting, and handling web search queries that match very few or no resultsabstractA non‐negligible fraction of user queries end up with very few or even no matching results in leading commercial web search engines. In this work, we provide a detailed characterization of such queries and show that search engines try to improve such queries by showing the results of related queries. Through a user study, we show that these query suggestions are usually perceived as relevant. Also, through a query log analysis, we show that the users are dissatisfied after submitting a query that match no results at least 88.5% of the time. As a first step towards solving these no‐answer queries, we devised a large number of features that can be used to identify such queries and built machine‐learning models. These models can be useful for scenarios such as the mobile‐ or meta‐search, where identifying a query that will retrieve no results at the client device (i.e., even before submitting it to the search engine) may yield gains in terms of the bandwidth usage, power consumption, and/or monetary costs. Experiments over query logs indicate that, despite the heavy skew in class sizes, our models achieve good prediction quality, with accuracy (in terms of area under the curve) up to 0.95. Erdem Sarigil, Ismail Sengör Altingövde, Roi Blanco, Berkant Barla Cambazoglu, Rifat Ozcan, Özgür Ulusoy |
J. Assoc. Inf. Sci. Technol. | 4 |
| 2017 | A machine learning approach for result caching in web search engines
Tayfun Küçükyilmaz, Berkant Barla Cambazoglu, Cevdet Aykanat, Ricardo Baeza-Yates |
Inf. Process. Manag. | 2 |
| 2017 | Exploiting search history of users for news personalization
Xiao Bai 0002, Berkant Barla Cambazoglu, Francesco Gullo, Amin Mantrach, Fabrizio Silvestri |
Inf. Sci. | 2 |
| 2017 | On the feasibility of predicting popular news at cold startabstractProminent news sites on the web provide hundreds of news articles daily. The abundance of news content competing to attract online attention, coupled with the manual effort involved in article selection, necessitates the timely prediction of future popularity of these news articles. The future popularity of a news article can be estimated using signals indicating the article's penetration in social media (e.g., number of tweets) in addition to traditional web analytics (e.g., number of page views). In practice, it is important to make such estimations as early as possible, preferably before the article is made available on the news site (i.e., at cold start). In this paper we perform a study on cold‐start news popularity prediction using a collection of 13,319 news articles obtained from Yahoo News, a major news provider. We characterize the popularity of news articles through a set of online metrics and try to predict their values across time using machine learning techniques on a large collection of features obtained from various sources. Our findings indicate that predicting news popularity at cold start is a difficult task, contrary to the findings of a prior work on the same topic. Most articles' popularity may not be accurately anticipated solely on the basis of content features, without having the early‐stage popularity values. Ioannis Arapakis, Berkant Barla Cambazoglu, Mounia Lalmas-Roelleke |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2017 | Understanding and Leveraging the Impact of Response Latency on User Behaviour in Web SearchabstractThe interplay between the response latency of web search systems and users’ search experience has only recently started to attract research attention, despite the important implications of response latency on monetisation of such systems. In this work, we carry out two complementary studies to investigate the impact of response latency on users’ searching behaviour in web search engines. We first conduct a controlled user study to investigate the sensitivity of users to increasing delays in response latency. This study shows that the users of a fast search system are more sensitive to delays than the users of a slow search system. Moreover, the study finds that users are more likely to notice the response latency delays beyond a certain latency threshold, their search experience potentially being affected. We then analyse a large number of search queries obtained from Yahoo Web Search to investigate the impact of response latency on users’ click behaviour. This analysis demonstrates the significant change in click behaviour as the response latency increases. We also find that certain user, context, and query attributes play a role in the way increasing response latency affects the click behaviour. To demonstrate a possible use case for our findings, we devise a machine-learning framework that leverages the latency impact, together with other features, to predict whether a user will issue any clicks on web search results. As a further extension of this use case, we investigate whether this machine-learning framework can be exploited to help search engines reduce their energy consumption during query processing. Xiao Bai 0002, Ioannis Arapakis, Berkant Barla Cambazoglu, Ana Freire |
ACM Trans. Inf. Syst. | 3 |
| 2016 | Linguistic Benchmarks of Online News Article QualityabstractOnline news editors ask themselves the same question many times: what is missing in this news article to go online?This is not an easy question to be answered by computational linguistic methods.In this work, we address this important question and characterise the constituents of news article editorial quality.More specifically, we identify 14 aspects related to the content of news articles.Through a correlation analysis, we quantify their independence and relation to assessing an article's editorial quality.We also demonstrate that the identified aspects, when combined together, can be used effectively in quality control methods for online news. Ioannis Arapakis, Filipa Peleja, Berkant Barla Cambazoglu, João Magalhães |
ACL (1) | 3 |
| 2016 | Improved Caching Techniques for Large-Scale Image Hosting ServicesabstractCommercial image serving systems, such as Flickr and Facebook, rely on large image caches to avoid the retrieval of requested images from the costly backend image store, as much as possible. Such systems serve the same image in different resolutions and, thus, in different sizes to different clients, depending on the properties of the clients' devices. The requested resolutions of images can be cached individually, as in the traditional caches, reducing the backend workload. However, a potentially better approach is to store relatively high-resolution images in the cache and resize them during the retrieval to obtain lower-resolution images. Having this kind of on-the-fly image resizing capability enables image serving systems to deploy more sophisticated caching policies and improve their serving performance further. In this paper, we formalize the static caching problem in image serving systems which provide on-the-fly image resizing functionality in their edge caches or regional caches. We propose two gain-based caching policies that construct a static, fixed-capacity cache to reduce the average serving time of images. The basic idea in the proposed policies is to identify the best resolution(s) of images to be cached so that the average serving time for future image retrieval requests is reduced. We conduct extensive experiments using real-life data access logs obtained from Flickr. We show that one of the proposed caching policies reduces the average response time of the service by up to 4.2% with respect to the best-performing baseline that mainly relies on the access frequency information to make the caching decisions. This improvement implies about 25% reduction in cache size under similar serving time constraints. Xiao Bai 0002, Berkant Barla Cambazoglu, Archie Russell |
SIGIR | 2 |
| 2016 | Scalability and Efficiency Challenges in Large-Scale Web Search EnginesabstractCommercial web search engines need to process thousands of queries every second and provide responses to user queries within a few hundred milliseconds. As a consequence of these tight performance constraints, search engines construct and maintain very large computing infrastructures for crawling the Web, indexing discovered pages, and processing user queries. The scalability and efficiency of these infrastructures require careful performance optimizations in every major component of the search engine. Berkant Barla Cambazoglu, Ricardo Baeza-Yates |
SIGIR | 1 |
| 2016 | Optimal Web Page Download Scheduling Policies for Green Web CrawlingabstractA web crawler is responsible for discovering and downloading new pages on the Web as well as refreshing previously downloaded pages. During these operations, the crawler issues a large number of HTTP requests to web servers. These requests increase the energy consumption and carbon footprint of the web servers since computational resources are used while serving the requests. In this work, we introduce the problem of green web crawling, where the objective is to devise a page refresh policy that minimizes the total staleness of pages in the repository of a web crawler, subject to a constraint on the amount of carbon emissions due to the processing on web servers. For the case of one web server and one crawling thread, the optimal policy turns out to be a greedy one. At each iteration, the page to be refreshed is selected based on a metric that considers the page's staleness, its size, and the greenness of the energy consumed at the web server premises. We then extend the optimal policy to the cases of 1) many servers; 2) multiple threads; and 3) pages with variable freshness requirements. We conduct simulations on a real data set that involves a large web server collection hosting around two billion pages. We present experimental results for the optimal page refresh policy as well as for various heuristics, in an effort to study the effect of different factors on performance. Vassiliki Hatzi, Berkant Barla Cambazoglu, Iordanis Koutsopoulos |
IEEE J. Sel. Areas Commun. | 2 |
| 2015 | LSDS-IR'15: 2015 Workshop on Large-Scale and Distributed Systems for Information RetrievalabstractThe growth of the Web and other Big Data sources lead to important performance problems for large-scale and distributed information retrieval systems. The scalability and efficiency of such information retrieval systems have an impact on their effectiveness, eventually affecting the experience of their users and monetization as well. The LSDS-IR'15 workshop will provide space for researchers to discuss the existing performance problems in the context of large-scale and distributed information retrieval systems and define new research directions in the modern Big Data era. The workshop expects to bring together information retrieval practitioners from the industry, as well as academic researchers concerned with any aspect of large-scale and distributed information retrieval systems. Ismail Sengör Altingövde, Berkant Barla Cambazoglu, Nicola Tonellotto |
CIKM | 2 |
| 2015 | Know Your Onions: Understanding the User Experience with the Knowledge Module in Web SearchabstractThe increasing availability of large volumes of human-curated content is shifting web search towards a paradigm that introduces seamlessly more semantic information to search engine result pages. This trend has resulted in the design of a new element known as the knowledge module (KM) where certain facts about named entities, obtained from various knowledge bases, are shown to users. So far, little has been done to uncover the role that this module plays on user experience in web search and whether it is perceived by users as a useful aid for their search tasks. Our work is an early attempt to bridge this gap. To this end, we conducted a crowdsourcing study aimed at understanding the effect of the KM on users' search experience and its overall utility. In particular, our study is the first to provide insights about the noticeability and usefulness of the KM in web search, together with comprehensive analyses of usability and workload. Ioannis Arapakis, Luis A. Leiva, Berkant Barla Cambazoglu |
CIKM | 3 |
| 2015 | Unconscious Physiological Effects of Search Latency on Users and Their Click BehaviourabstractUnderstanding the impact of a search system's response latency on its users' searching behaviour has been recently an active research topic in the information retrieval and human-computer interaction areas. Along the same line, this paper focuses on the user impact of search latency and makes the following two contributions. First, through a controlled experiment, we reveal the physiological effects of response latency on users and show that these effects are present even at small increases in response latency. We compare these effects with the information gathered from self-reports and show that they capture the nuanced attentional and emotional reactions to latency much better. Second, we carry out a large-scale analysis using a web search query log obtained from Yahoo to understand the change in the way users engage with a web search engine under varying levels of increasing response latency. In particular, we analyse the change in the click behaviour of users when they are subject to increasing response latency and reveal significant behavioural differences. Miguel Barreda-Ángeles, Ioannis Arapakis, Xiao Bai 0002, Berkant Barla Cambazoglu, Alexandre Pereda-Baños |
SIGIR | 4 |
| 2015 | A Random Walk Model for Optimization of Search Impact in Web Frontier RankingabstractLarge-scale web search engines need to crawl the Web continuously to discover and download newly created web content. The speed at which the new content is discovered and the quality of the discovered content can have a big impact on the coverage and quality of the results provided by the search engine. In this paper, we propose a search-centric solution to the problem of prioritizing the pages in the frontier of a crawler for download. Our approach essentially orders the web pages in the frontier through a random walk model that takes into account the pages' potential impact on user-perceived search quality. In addition, we propose a link graph enrichment technique that extends this solution. Finally, we explore a machine learning approach that combines different frontier prioritization approaches. We conduct experiments using two very large, real-life web datasets to observe various search quality metrics. Comparisons with several baseline techniques indicate that the proposed approaches have the potential to improve the user-perceived quality of web search results considerably. Giang Tran, Ata Turk, Berkant Barla Cambazoglu, Wolfgang Nejdl |
SIGIR | 3 |
| 2015 | Scalability and Efficiency Challenges in Large-Scale Web Search EnginesabstractCommercial web search engines need to process thousands of queries every second and provide responses to user queries within a few hundred milliseconds. As a consequence of these tight performance constraints, search engines construct and maintain very large computing infrastructures for crawling the Web, indexing discovered pages, and processing user queries. The scalability and efficiency of these infrastructures require careful performance optimizations in every major component of the search engine. This tutorial aims to provide a fairly comprehensive overview of the scalability and efficiency challenges in large-scale web search engines. In particular, the tutorial provides an in-depth architectural overview of a web search engine, mainly focusing on the web crawling, indexing, and query processing components. The scalability and efficiency issues encountered in the above-mentioned components are presented at four different granularities: at the level of a single computer, a cluster of computers, a single data center, and a multi-center search engine. The tutorial also points at the open research problems and provides recommendations to researchers who are new to the field. Berkant Barla Cambazoglu, Ricardo Baeza-Yates |
WSDM | 1 |
| 2015 | Task allocation in volunteer computing networks under monetary budget constraints
Huseyin Guler, Berkant Barla Cambazoglu, Öznur Özkasap |
Peer-to-Peer Netw. Appl. | 2 |
| 2014 | Impact of response latency on user behavior in web searchabstractTraditionally, the efficiency and effectiveness of search systems have both been of great interest to the information retrieval community. However, an in-depth analysis on the interplay between the response latency of web search systems and users' search experience has been missing so far. In order to fill this gap, we conduct two separate studies aiming to reveal how response latency affects the user behavior in web search. First, we conduct a controlled user study trying to understand how users perceive the response latency of a search system and how sensitive they are to increasing delays in response. This study reveals that, when artificial delays are introduced into the response, the users of a fast search system are more likely to notice these delays than the users of a slow search system. The introduced delays become noticeable by the users once they exceed a certain threshold value. Second, we perform an analysis using a large-scale query log obtained from Yahoo web search to observe the potential impact of increasing response latency on the click behavior of users. This analysis demonstrates that latency has an impact on the click behavior of users to some extent. In particular, given two content-wise identical search result pages, we show that the users are more likely to perform clicks on the result page that is served with lower latency. Ioannis Arapakis, Xiao Bai 0002, Berkant Barla Cambazoglu |
SIGIR | 3 |
| 2014 | Scalability and efficiency challenges in large-scale web search enginesabstractLarge-scale web search engines rely on massive compute infrastructures to be able to cope with the continuous growth of the Web and their user bases. In such search engines, achieving scalability and efficiency requires making careful architectural design choices while devising algorithmic performance optimizations. Unfortunately, most details about the internal functioning of commercial web search engines remain undisclosed due to their financial value and the high level of competition in the search market. The main objective of this tutorial is to provide an overview of the fundamental scalability and efficiency challenges in commercial web search engines, bridging the existing gap between the industry and academia. Berkant Barla Cambazoglu, Ricardo Baeza-Yates |
SIGIR | 1 |
| 2014 | Workshop on large-scale and distributed systems for information retrieval (LSDS-IR 2014)abstractThe LSDS-IR'14 workshop aims to bring together information retrieval practitioners from industry and academic researchers concerned with efficient and distributed IR systems. The workshop also welcomes contributions that propose different ways of leveraging diversity and multiplicity of resources available in distributed systems. The main goal of the workshop is to attract people from industry and academia to present and discuss ideas, problems, and results related to the efficiency of large scale and distributed information retrieval systems. Ismail Sengör Altingövde, Berkant Barla Cambazoglu, Craig Macdonald, Nicola Tonellotto |
WSDM | 2 |
| 2014 | Improving the efficiency of multi-site web search enginesabstractA multi-site web search engine is composed of a number of search sites geographically distributed around the world. Each search site is typically responsible for crawling and indexing the web pages that are in its geographical neighborhood. A query is selectively processed on a subset of search sites that are predicted to return the best-matching results. The scalability and efficiency of multi-site web search engines have attracted a lot of research attention in recent years. In particular, research has focused on replicating important web pages across sites, forwarding queries to relevant sites, and caching results of previous queries. Yet, these problems have only been studied in isolation, but no prior work has properly investigated the interplay between them. Guillem Francès, Xiao Bai 0002, Berkant Barla Cambazoglu, Ricardo Baeza-Yates |
WSDM | 3 |
| 2014 | User engagement in online News: Under the scope of sentiment, interest, affect, and gazeabstractOnline content providers, such as news portals and social media platforms, constantly seek new ways to attract large shares of online attention by keeping their users engaged. A common challenge is to identify which aspects of online interaction influence user engagement the most. In this article, through an analysis of a news article collection obtained from Yahoo News US, we demonstrate that news articles exhibit considerable variation in terms of the sentimentality and polarity of their content, depending on factors such as news provider and genre. Moreover, through a laboratory study, we observe the effect of sentimentality and polarity of news and comments on a set of subjective and objective measures of engagement. In particular, we show that attention, affect, and gaze differ across news of varying interestingness. As part of our study, we also explore methods that exploit the sentiments expressed in user comments to reorder the lists of comments displayed in news pages. Our results indicate that user engagement can be anticipated predicted if we account for the sentimentality and polarity of the content as well as other factors that drive attention and inspire human curiosity. Ioannis Arapakis, Mounia Lalmas-Roelleke, Berkant Barla Cambazoglu, Mari-Carmen Marcos, Joemon M. Jose |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2014 | Improving the Performance of IndependentTask Assignment Heuristics MinMin, MaxMin and SufferageabstractMinMin, MaxMin, and Sufferage are constructive heuristics that are widely and successfully used in assigning independent tasks to processors in heterogeneous computing systems. All three heuristics are known to run in O(KN2) time in assigning N tasks to K processors. In this paper, we propose an algorithmic improvement that asymptotically decreases the running time complexity of MinMin to O(KN log N) without affecting its solution quality. Furthermore, we combine the newly proposed MinMin algorithm with MaxMin as well as Sufferage, obtaining two hybrid algorithms. The motivation behind the former hybrid algorithm is to address the drawback of MaxMin in solving problem instances with highly skewed cost distributions while also improving the running time performance of MaxMin. The latter hybrid algorithm improves the running time performance of Sufferage without degrading its solution quality. The proposed algorithms are easy to implement and we illustrate them through detailed pseudocodes. The experimental results over a large number of real-life data sets show that the proposed fast MinMin algorithm and the proposed hybrid algorithms perform significantly better than their traditional counterparts as well as more recent state-of-the-art assignment heuristics. For the large data sets used in the experiments, MinMin, MaxMin, and Sufferage, as well as recent state-of-the-art heuristics, require days, weeks, or even months to produce a solution, whereas all of the proposed algorithms produce solutions within only two or three minutes. E. Kartal Tabak, Berkant Barla Cambazoglu, Cevdet Aykanat |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2014 | Sentiment-Focused Web CrawlingabstractSentiments and opinions expressed in Web pages towards objects, entities, and products constitute an important portion of the textual content available in the Web. In the last decade, the analysis of such content has gained importance due to its high potential for monetization. Despite the vast interest in sentiment analysis, somewhat surprisingly, the discovery of sentimental or opinionated Web content is mostly ignored. This work aims to fill this gap and addresses the problem of quickly discovering and fetching the sentimental content present in the Web. To this end, we design a sentiment-focused Web crawling framework. In particular, we propose different sentiment-focused Web crawling strategies that prioritize discovered URLs based on their predicted sentiment scores. Through simulations, these strategies are shown to achieve considerable performance improvement over general-purpose Web crawling strategies in discovery of sentimental Web content. A. Gural Vural, Berkant Barla Cambazoglu, Pinar Karagöz |
ACM Trans. Web | 2 |
| 2013 | Incorporating the surfing behavior of web users into pagerankabstractIn large-scale commercial web search engines, estimating the importance of a web page is a crucial ingredient in ranking web search results. So far, to assess the importance of web pages, two different types of feedback have been taken into account, independent of each other: the feedback obtained from the hyperlink structure among the web pages (e.g., PageRank) or the web browsing patterns of users (e.g., BrowseRank). Unfortunately, both types of feedback have certain drawbacks. While the former lacks the user preferences and is vulnerable to malicious intent, the latter suffers from sparsity and hence low web coverage. In this work, we combine these two types of feedback under a hybrid page ranking model in order to alleviate the above-mentioned drawbacks. Our empirical results indicate that the proposed model leads to better estimation of page importance according to an evaluation metric that relies on user click feedback obtained from web search query logs. We conduct all of our experiments in a realistic setting, using a very large scale web page collection (around 6.5 billion web pages) and web browsing data (around two billion web page visits). Shatlyk Ashyralyyev, Berkant Barla Cambazoglu, Cevdet Aykanat |
CIKM | 2 |
| 2013 | Strategies for setting time-to-live values in result cachesabstractIn web query result caching, staleness of queries are often bounded via a time-to-live (TTL) mechanism, which expires the validity of cached query results at some point in time. In this work, we evaluate the performance of three alternative TTL mechanisms: time-based TTL, frequency-based TTL, and click-based TTL. Moreover, we propose hybrid approaches obtained by pair-wise combination of these mechanisms. Our results indicate that combining time-based TTL with frequency-based TTL yields superior performance (i.e., lower stale query traffic and less redundant computation) than using a particular mechanism in isolation. Fethi Burak Sazoglu, Berkant Barla Cambazoglu, Rifat Ozcan, Ismail Sengör Altingövde, Özgür Ulusoy |
CIKM | 2 |
| 2013 | Entity Recommendations in Web Search
Roi Blanco, Berkant Barla Cambazoglu, Peter Mika, Nicolas Torzec |
ISWC (2) | 2 |
| 2013 | Scalability and efficiency challenges in commercial web search enginesabstractCommercial web search engines rely on very large compute infrastructures to be able to cope with the continuous growth of the Web and user bases. Achieving scalability and efficiency in such large-scale search engines requires making careful architectural design choices while devising algorithmic performance optimizations. Unfortunately, most details about the internal functioning of commercial web search engines remain undisclosed due to their financial value and the high level of competition in the search market. The main objective of this tutorial is to provide an overview of the fundamental scalability and efficiency challenges in commercial web search engines, bridging the existing gap between the industry and academia. Berkant Barla Cambazoglu, Ricardo Baeza-Yates |
SIGIR | 1 |
| 2013 | A financial cost metric for result cachingabstractWeb search engines cache results of frequent and/or recent queries. Result caching strategies can be evaluated using different metrics, hit rate being the most well-known. Recent works take the processing overhead of queries into account when evaluating the performance of result caching strategies and propose cost-aware caching strategies. In this paper, we propose a financial cost metric that goes one step beyond and takes also the hourly electricity prices into account when computing the cost. We evaluate the most well-known static, dynamic, and hybrid result caching strategies under this new metric. Moreover, we propose a financial-cost-aware version of the well-known LRU strategy and show that it outperforms the original LRU strategy in terms of the financial cost metric. Fethi Burak Sazoglu, Berkant Barla Cambazoglu, Rifat Ozcan, Ismail Sengör Altingövde, Özgür Ulusoy |
SIGIR | 2 |
| 2013 | Document replication strategies for geographically distributed web search engines
Enver Kayaaslan, Berkant Barla Cambazoglu, Cevdet Aykanat |
Inf. Process. Manag. | 2 |
| 2013 | A term-based inverted index partitioning model for efficient distributed query processingabstractIn a shared-nothing, distributed text retrieval system, queries are processed over an inverted index that is partitioned among a number of index servers. In practice, the index is either document-based or term-based partitioned. This choice is made depending on the properties of the underlying hardware infrastructure, query traffic distribution, and some performance and availability constraints. In query processing on retrieval systems that adopt a term-based index partitioning strategy, the high communication overhead due to the transfer of large amounts of data from the index servers forms a major performance bottleneck, deteriorating the scalability of the entire distributed retrieval system. In this work, to alleviate this problem, we propose a novel inverted index partitioning model that relies on hypergraph partitioning. In the proposed model, concurrently accessed index entries are assigned to the same index servers, based on the inverted index access patterns extracted from the past query logs. The model aims to minimize the communication overhead that will be incurred by future queries while maintaining the computational load balance among the index servers. We evaluate the performance of the proposed model through extensive experiments using a real-life text collection and a search query sample. Our results show that considerable performance gains can be achieved relative to the term-based index partitioning strategies previously proposed in literature. In most cases, however, the performance remains inferior to that attained by document-based partitioning. Berkant Barla Cambazoglu, Enver Kayaaslan, Simon Jonassen, Cevdet Aykanat |
ACM Trans. Web | 1 |
| 2013 | Second Chance: A Hybrid Approach for Dynamic Result Caching and Prefetching in Search EnginesabstractWeb search engines are known to cache the results of previously issued queries. The stored results typically contain the document summaries and some data that is used to construct the final search result page returned to the user. An alternative strategy is to store in the cache only the result document IDs, which take much less space, allowing results of more queries to be cached. These two strategies lead to an interesting trade-off between the hit rate and the average query response latency. In this work, in order to exploit this trade-off, we propose a hybrid result caching strategy where a dynamic result cache is split into two sections: an HTML cache and a docID cache. Moreover, using a realistic cost model, we evaluate the performance of different result prefetching strategies for the proposed hybrid cache and the baseline HTML-only cache. Finally, we propose a machine learning approach to predict singleton queries, which occur only once in the query stream. We show that when the proposed hybrid result caching strategy is coupled with the singleton query predictor, the hit rate is further improved. Rifat Ozcan, Ismail Sengör Altingövde, Berkant Barla Cambazoglu, Özgür Ulusoy |
ACM Trans. Web | 3 |
| 2012 | Characterizing web search queries that match very few or no resultsabstractDespite the continuous efforts to improve the web search quality, a non-negligible fraction of user queries end up with very few or even no matching results in leading web search engines. In this work, we provide a detailed characterization of such queries based on an analysis of a real-life query log. Our experimental setup allows us to characterize the queries with few/no results and compare the mechanisms employed by the major search engines in handling them. Ismail Sengör Altingövde, Roi Blanco, Berkant Barla Cambazoglu, Rifat Ozcan, Erdem Sarigil, Özgür Ulusoy |
CIKM | 3 |
| 2012 | Sentiment-focused web crawlingabstractThe sentiments and opinions that are expressed in web pages towards objects, entities, and products constitute an important portion of the textual content available in the Web. Despite the vast interest in sentiment analysis and opinion mining, somewhat surprisingly, the discovery of the sentimental or opinionated web content is mostly ignored. This work aims to fill this gap and address the problem of quickly discovering and fetching the sentimental content present in the Web. To this end, we design a sentiment-focused web crawling framework for faster discovery and retrieval of such content. In particular, we propose different sentiment-focused web crawling strategies that prioritize discovered URLs based on their predicted sentiment scores. Through simulations, these strategies are shown to achieve considerable performance improvement over general-purpose web crawling strategies in discovering sentimental content. A. Gural Vural, Berkant Barla Cambazoglu, Pinar Karagöz |
CIKM | 2 |
| 2012 | Adaptive Time-to-Live Strategies for Query Result Caching in Web Search Engines
Sadiye Alici, Ismail Sengör Altingövde, Rifat Ozcan, Berkant Barla Cambazoglu, Özgür Ulusoy |
ECIR | 4 |
| 2012 | Prefetching query results and its impact on search enginesabstractWe investigate the impact of query result prefetching on the efficiency and effectiveness of web search engines. We propose offline and online strategies for selecting and ordering queries whose results are to be prefetched. The offline strategies rely on query log analysis and the queries are selected from the queries issued on the previous day. The online strategies select the queries from the result cache, relying on a machine learning model that estimates the arrival times of queries. We carefully evaluate the proposed prefetching techniques via simulation on a query log obtained from Yahoo! web search. We demonstrate that our strategies are able to improve various performance metrics, including the hit rate, query response time, result freshness, and query degradation rate, relative to a state-of-the-art baseline. Simon Jonassen, Berkant Barla Cambazoglu, Fabrizio Silvestri |
SIGIR | 2 |
| 2012 | Impact of Regionalization on Performance of Web Search Engine Result Caches
Berkant Barla Cambazoglu, Ismail Sengör Altingövde |
SPIRE | 1 |
| 2012 | A large-scale sentiment analysis for Yahoo! answersabstractSentiment extraction from online web documents has recently been an active research topic due to its potential use in commercial applications. By sentiment analysis, we refer to the problem of assigning a quantitative positive/negative mood to a short bit of text. Most studies in this area are limited to the identification of sentiments and do not investigate the interplay between sentiments and other factors. In this work, we use a sentiment extraction tool to investigate the influence of factors such as gender, age, education level, the topic at hand, or even the time of the day on sentiments in the context of a large online question answering site. We start our analysis by looking at direct correlations, e.g., we observe more positive sentiments on weekends, very neutral ones in the Science & Mathematics topic, a trend for younger people to express stronger sentiments, or people in military bases to ask the most neutral questions. We then extend this basic analysis by investigating how properties of the (asker, answerer) pair affect the sentiment present in the answer. Among other things, we observe a dependence on the pairing of some inferred attributes estimated by a user's ZIP code. We also show that the best answers differ in their sentiments from other answers, e.g., in the Business & Finance topic, best answers tend to have a more neutral sentiment than other answers. Finally, we report results for the task of predicting the attitude that a question will provoke in answers. We believe that understanding factors influencing the mood of users is not only interesting from a sociological point of view, but also has applications in advertising, recommendation, and search. Onur Küçüktunç, Berkant Barla Cambazoglu, Ingmar Weber, Hakan Ferhatosmanoglu |
WSDM | 2 |
| 2012 | A five-level static cache architecture for web search engines
Rifat Ozcan, Ismail Sengör Altingövde, Berkant Barla Cambazoglu, Flavio Paiva Junqueira, Özgür Ulusoy |
Inf. Process. Manag. | 3 |
| 2012 | Cache-Based Query Processing for Search EnginesabstractIn practice, a search engine may fail to serve a query due to various reasons such as hardware/network failures, excessive query load, lack of matching documents, or service contract limitations (e.g., the query rate limits for third-party users of a search service). In this kind of scenarios, where the backend search system is unable to generate answers to queries, approximate answers can be generated by exploiting the previously computed query results available in the result cache of the search engine. In this work, we propose two alternative strategies to implement this cache-based query processing idea. The first strategy aggregates the results of similar queries that are previously cached in order to create synthetic results for new queries. The second strategy forms an inverted index over the textual information (i.e., query terms and result snippets) present in the result cache and uses this index to answer new queries. Both approaches achieve reasonable result qualities compared to processing queries with an inverted index built on the collection. Berkant Barla Cambazoglu, Ismail Sengör Altingövde, Rifat Ozcan, Özgür Ulusoy |
ACM Trans. Web | 1 |
| 2011 | Discovering URLs through user feedbackabstractSearch engines rely upon crawling to build their Web page collections. A Web crawler typically discovers new URLs by following the link structure induced by links on Web pages. As the number of documents on the Web is large, discovering newly created URLs may take arbitrarily long, and depending on how a given page is connected to others, such a crawler may miss the pages altogether. In this paper, we evaluate the benefits of integrating a passive URL discovery mechanism into a Web crawler. This mechanism is passive in the sense that it does not require the crawler to actively fetch documents from the Web to discover URLs. We focus here on a mechanism that uses toolbar data as a representative source for new URL discovery. We use the toolbar logs of Yahoo! to characterize the URLs that are accessed by users via their browsers, but not discovered by Yahoo! Web crawler. We show that a high fraction of URLs that appear in toolbar logs are not discovered by the crawler. We also reveal that a certain fraction of URLs are discovered by the crawler later than the time they are first accessed by users. One important conclusion of our work is that web search engines can highly benefit from user feedback in the form of toolbar logs for passive URL discovery. Xiao Bai 0002, Berkant Barla Cambazoglu, Flavio Paiva Junqueira |
CIKM | 2 |
| 2011 | Assigning documents to master sites in distributed searchabstractAn appealing solution to scale Web search with the growth of the Internet is the use of distributed architectures. Distributed search engines rely on multiple sites deployed in distant regions across the world, where each site is specialized to serve queries issued by the users of its region. This paper investigates the problem of assigning each document to a master site. We show that by leveraging similarities between a document and the activity of the users, we can accurately detect which site is the most relevant to place a document. We conduct various experiments using two document assignment approaches, showing performance improvements of up to 20.8% over a baseline technique which assigns the documents to search sites based on their language. Roi Blanco, Berkant Barla Cambazoglu, Flavio Paiva Junqueira, Ivan Kelly, Vincent Leroy 0001 |
CIKM | 2 |
| 2011 | LSDS-IR'11: the 9th workshop on large-scale and distributed systems for information retrievalabstractThe growth of the Web and user bases lead to important performance problems for large-scale Web search engines. The LSDS- IR '11 workshop focuses on research contributions related to the scalability and efficiency of distributed information retrieval (IR) systems. The workshop also encourages contributions that propose different ways of leveraging diversity and multiplicity of resources available in distributed systems. More specifically, we are interested in novel applications, models, and architectures that deal with efficiency and scalability of distributed IR systems. Claudio Lucchese, Berkant Barla Cambazoglu |
CIKM | 2 |
| 2011 | Second Chance: A Hybrid Approach for Dynamic Result Caching in Search Engines
Ismail Sengör Altingövde, Rifat Ozcan, Berkant Barla Cambazoglu, Özgür Ulusoy |
ECIR | 3 |
| 2011 | Machine learned job recommendationabstractWe address the problem of recommending suitable jobs to people who are seeking a new job. We formulate this recommendation problem as a supervised machine learning problem. Our technique exploits all past job transitions as well as the data associated with employees and institutions to predict an employee's next job transition. We train a machine learning model using a large number of job transitions extracted from the publicly available employee profiles in the Web. Experiments show that job transitions can be accurately predicted, significantly improving over a baseline that always predicts the most frequent institution in the data. Ioannis K. Paparrizos, Berkant Barla Cambazoglu, Aristides Gionis |
RecSys | 2 |
| 2011 | Timestamp-based result cache invalidation for web search enginesabstractThe result cache is a vital component for efficiency of large-scale web search engines, and maintaining the freshness of cached query results is the current research challenge. As a remedy to this problem, our work proposes a new mechanism to identify queries whose cached results are stale. The basic idea behind our mechanism is to maintain and compare generation time of query results with update times of posting lists and documents to decide on staleness of query results. The proposed technique is evaluated using a Wikipedia document collection with real update information and a real-life query log. We show that our technique has good prediction accuracy, relative to a baseline based on the time-to-live mechanism. Moreover, it is easy to implement and incurs less processing overhead on the system relative to a recently proposed, more sophisticated invalidation mechanism. Sadiye Alici, Ismail Sengör Altingövde, Rifat Ozcan, Berkant Barla Cambazoglu, Özgür Ulusoy |
SIGIR | 4 |
| 2011 | Energy-price-driven query processing in multi-center web search enginesabstractConcurrently processing thousands of web queries, each with a response time under a fraction of a second, necessitates maintaining and operating massive data centers. For large-scale web search engines, this translates into high energy consumption and a huge electric bill. This work takes the challenge to reduce the electric bill of commercial web search engines operating on data centers that are geographically far apart. Based on the observation that energy prices and query workloads show high spatio-temporal variation, we propose a technique that dynamically shifts the query workload of a search engine between its data centers to reduce the electric bill. Experiments on real-life query workloads obtained from a commercial search engine show that significant financial savings can be achieved by this technique. Enver Kayaaslan, Berkant Barla Cambazoglu, Roi Blanco, Flavio Paiva Junqueira, Cevdet Aykanat |
SIGIR | 2 |
| 2011 | Posting list intersection on multicore architecturesabstractIn current commercial Web search engines, queries are processed in the conjunctive mode, which requires the search engine to compute the intersection of a number of posting lists to determine the documents matching all query terms. In practice, the intersection operation takes a significant fraction of the query processing time, for some queries dominating the total query latency. Hence, efficient posting list intersection is critical for achieving short query latencies. In this work, we focus on improving the performance of posting list intersection by leveraging the compute capabilities of recent multicore systems. To this end, we consider various coarse-grained and fine-grained parallelization models for list intersection. Specifically, we present an algorithm that partitions the work associated with a given query into a number of small and independent tasks that are subsequently processed in parallel. Through a detailed empirical analysis of these alternative models, we demonstrate that exploiting parallelism at the finest-level of granularity is critical to achieve the best performance on multicore systems. On an eight-core system, the fine-grained parallelization method is able to achieve more than five times reduction in average query processing time while still exploiting the parallelism for high query throughput. Shirish Tatikonda, Berkant Barla Cambazoglu, Flavio Paiva Junqueira |
SIGIR | 2 |
| 2011 | Document assignment in multi-site search enginesabstractAssigning documents accurately to sites is critical for the performance of multi-site Web search engines. In such settings, sites crawl only documents they index and forward queries to obtain best-matching documents from other sites. Inaccurate assignments may lead to inefficiencies when crawling Web pages or processing user queries. In this work, we propose a machine-learned document assignment strategy that uses the locality of document views in search results to decide upon assignments. We evaluate the performance of our strategy using various document features extracted from a large Web collection. Our experimental setup uses query logs from a number of search front-ends spread across different geographic locations and uses these logs to learn the document access patterns. We compare our technique against baselines such as region- and language-based document assignment and observe that our technique achieves substantial performance improvements with respect to recall. With our technique, we are able to obtain a small query forwarding rate (0.04) requiring roughly 45% less replication of documents compared to replicating all documents across all sites. Ulf Brefeld, Berkant Barla Cambazoglu, Flavio Paiva Junqueira |
WSDM | 2 |
| 2011 | Site-Based Partitioning and Repartitioning Techniques for Parallel PageRank ComputationabstractThe PageRank algorithm is an important component in effective web search. At the core of this algorithm are repeated sparse matrix-vector multiplications where the involved web matrices grow in parallel with the growth of the web and are stored in a distributed manner due to space limitations. Hence, the PageRank computation, which is frequently repeated, must be performed in parallel with high-efficiency and low-preprocessing overhead while considering the initial distributed nature of the web matrices. Our contributions in this work are twofold. We first investigate the application of state-of-the-art sparse matrix partitioning models in order to attain high efficiency in parallel PageRank computations with a particular focus on reducing the preprocessing overhead they introduce. For this purpose, we evaluate two different compression schemes on the web matrix using the site information inherently available in links. Second, we consider the more realistic scenario of starting with an initially distributed data and extend our algorithms to cover the repartitioning of such data for efficient PageRank computation. We report performance results using our parallelization of a state-of-the-art PageRank algorithm on two different PC clusters with 40 and 64 processors. Experiments show that the proposed techniques achieve considerably high speedups while incurring a preprocessing overhead of several iterations (for some instances even less than a single iteration) of the underlying sequential PageRank algorithm. Ali Cevahir, Cevdet Aykanat, Ata Turk, Berkant Barla Cambazoglu |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2010 | Web search solved?: all result rankings the same?abstractThe objective of this work is to derive quantitative statements about what fraction of web search queries issued to the state-of-the-art commercial search engines lead to excellent results or, on the contrary, poor results. To be able to make such statements in an automated way, we propose a new measure that is based on lower and upper bound analysis over the standard relevance measures. Moreover, we extend this measure to carry out comparisons between competing search engines by introducing the concept of disruptive sets, which we use to estimate the degree to which a search engine solves queries that are not solved by its competitors. We report empirical results on a large editorial evaluation of the three largest search engines in the US market. Hugo Zaragoza, Berkant Barla Cambazoglu, Ricardo Baeza-Yates |
CIKM | 2 |
| 2010 | Cold start link predictionabstractIn the traditional link prediction problem, a snapshot of a social network is used as a starting point to predict, by means of graph-theoretic measures, the links that are likely to appear in the future. In this paper, we introduce cold start link prediction as the problem of predicting the structure of a social network when the network itself is totally missing while some other information regarding the nodes is available. We propose a two-phase method based on the bootstrap probabilistic graph. The first phase generates an implicit social network under the form of a probabilistic graph. The second phase applies probabilistic graph-based measures to produce the final prediction. We assess our method empirically over a large data collection obtained from Flickr, using interest groups as the initial information. The experiments confirm the effectiveness of our approach. Vincent Leroy 0001, Berkant Barla Cambazoglu, Francesco Bonchi |
KDD | 2 |
| 2010 | Query forwarding in geographically distributed search enginesabstractQuery forwarding is an important technique for preserving the result quality in distributed search engines where the index is geographically partitioned over multiple search sites. The key component in query forwarding is the thresholding algorithm by which the forwarding decisions are given. In this paper, we propose a linear-programming-based thresholding algorithm that significantly outperforms the current state-of-the-art in terms of achieved search efficiency values. Moreover, we evaluate a greedy heuristic for partial index replication and investigate the impact of result cache freshness on query forwarding performance. Finally, we present some optimizations that improve the performance further, under certain conditions. We evaluate the proposed techniques by simulations over a real-life setting, using a large query log and a document collection obtained from Yahoo!. Berkant Barla Cambazoglu, Emre Varol, Enver Kayaaslan, Cevdet Aykanat, Ricardo Baeza-Yates |
SIGIR | 1 |
| 2010 | Early exit optimizations for additive machine learned ranking systemsabstractSome commercial web search engines rely on sophisticated machine learning systems for ranking web documents. Due to very large collection sizes and tight constraints on query response times, online efficiency of these learning systems forms a bottleneck. An important problem in such systems is to speedup the ranking process without sacrificing much from the quality of results. In this paper, we propose optimization strategies that allow short-circuiting score computations in additive learning systems. The strategies are evaluated over a state-of-the-art machine learning system and a large, real-life query log, obtained from Yahoo!. By the proposed strategies, we are able to speedup the score computations by more than four times with almost no loss in result quality. Berkant Barla Cambazoglu, Hugo Zaragoza, Olivier Chapelle, Ciya Liao, Zhaohui Zheng 0001, Jon Degenhardt |
WSDM | 1 |
| 2010 | A refreshing perspective of search engine cachingabstractCommercial Web search engines have to process user queries over huge Web indexes under tight latency constraints. In practice, to achieve low latency, large result caches are employed and a portion of the query traffic is served using previously computed results. Moreover, search engines need to update their indexes frequently to incorporate changes to the Web. After every index update, however, the content of cache entries may become stale, thus decreasing the freshness of served results. In this work, we first argue that the real problem in today's caching for large-scale search engines is not eviction policies, but the ability to cope with changes to the index, i.e., cache freshness. We then introduce a novel algorithm that uses a time-to-live value to set cache entries to expire and selectively refreshes cached results by issuing refresh queries to back-end search clusters. The algorithm prioritizes the entries to refresh according to a heuristic that combines the frequency of access with the age of an entry in the cache. In addition, for setting the rate at which refresh queries are issued, we present a mechanism that takes into account idle cycles of back-end servers. Evaluation using a real workload shows that our algorithm can achieve hit rate improvements as well as reduction in average hit ages. An implementation of this algorithm is currently in production use at Yahoo!. Berkant Barla Cambazoglu, Flavio Paiva Junqueira, Vassilis Plachouras, Scott A. Banachowski, Baoqiu Cui, Swee Lim, Bill Bridge |
WWW | 1 |
| 2010 | Review of "Search Engines: Information Retrieval in Practice" by Croft, Metzler and Strohman
Berkant Barla Cambazoglu |
Inf. Process. Manag. | 1 |
| 2010 | A link-based storage scheme for efficient aggregate query processing on clustered road networks
Engin Demir, Cevdet Aykanat, Berkant Barla Cambazoglu |
Inf. Syst. | 3 |
| 2009 | Quantifying performance and quality gains in distributed web search enginesabstractDistributed search engines based on geographical partitioning of a central Web index emerge as a feasible solution to the immense growth of the Web, user bases, and query traffic. However, there is still lack of research in quantifying the performance and quality gains that can be achieved by such architectures. In this paper, we develop various cost models to evaluate the performance benefits of a geographically distributed search engine architecture based on partial index replication and query forwarding. Specifically, we focus on possible performance gains due to the distributed nature of query processing and Web crawling processes. We show that any response time gain achieved by distributed query processing can be utilized to improve search relevance as the use of complex but more accurate algorithms can now be enabled for document ranking. We also show that distributed Web crawling leads to better Web coverage and try to see if this improves the search quality. We verify the validity of our claims over large, real-life datasets via simulations. Berkant Barla Cambazoglu, Vassilis Plachouras, Ricardo Baeza-Yates |
SIGIR | 1 |
| 2009 | On efficient posting list intersection with multicore processorsabstractNo abstract available. Shirish Tatikonda, Flavio Paiva Junqueira, Berkant Barla Cambazoglu, Vassilis Plachouras |
SIGIR | 3 |
| 2008 | Chat mining: Predicting user and message attributes in computer-mediated communication
Tayfun Küçükyilmaz, Berkant Barla Cambazoglu, Cevdet Aykanat, Fazli Can |
Inf. Process. Manag. | 2 |
| 2008 | Clustering spatial networks for aggregate query processing: A hypergraph approach
Engin Demir, Cevdet Aykanat, Berkant Barla Cambazoglu |
Inf. Syst. | 3 |
| 2008 | Model Formulation: Sharing Data and Analytical Resources Securely in a Biomedical Research Grid EnvironmentabstractOBJECTIVES: To develop a security infrastructure to support controlled and secure access to data and analytical resources in a biomedical research Grid environment, while facilitating resource sharing among collaborators. DESIGN: A Grid security infrastructure, called Grid Authentication and Authorization with Reliably Distributed Services (GAARDS), is developed as a key architecture component of the NCI-funded cancer Biomedical Informatics Grid (caBIG). The GAARDS is designed to support in a distributed environment 1) efficient provisioning and federation of user identities and credentials; 2) group-based access control support with which resource providers can enforce policies based on community accepted groups and local groups; and 3) management of a trust fabric so that policies can be enforced based on required levels of assurance. MEASUREMENTS: GAARDS is implemented as a suite of Grid services and administrative tools. It provides three core services: Dorian for management and federation of user identities, Grid Trust Service for maintaining and provisioning a federated trust fabric within the Grid environment, and Grid Grouper for enforcing authorization policies based on both local and Grid-level groups. RESULTS: The GAARDS infrastructure is available as a stand-alone system and as a component of the caGrid infrastructure. More information about GAARDS can be accessed at http://www.cagrid.org. CONCLUSIONS: GAARDS provides a comprehensive system to address the security challenges associated with environments in which resources may be located at different sites, requests to access the resources may cross institutional boundaries, and user credentials are created, managed, revoked dynamically in a de-centralized manner. Stephen Langella, Shannon Hastings, Scott Oster, Tony Pan, Ashish Sharma 0001, Justin Permar, David Ervin, Berkant Barla Cambazoglu, Tahsin M. Kurç, Joel H. Saltz |
J. Am. Medical Informatics Assoc. | 8 |
| 2008 | Multi-level direct K-way hypergraph partitioning with multiple constraints and fixed vertices
Cevdet Aykanat, Berkant Barla Cambazoglu, Bora Uçar |
J. Parallel Distributed Comput. | 2 |
| 2007 | Computerized Pathological Image Analysis For Neuroblastoma Prognosis
Metin Nafi Gürcan, Jun Kong 0002, Olcay Sertel, Berkant Barla Cambazoglu, Joel H. Saltz, Ümit V. Çatalyürek |
AMIA | 4 |
| 2007 | Architecture of a grid-enabled Web search engine
Berkant Barla Cambazoglu, Evren Karaca, Tayfun Küçükyilmaz, Ata Turk, Cevdet Aykanat |
Inf. Process. Manag. | 1 |
| 2007 | Adaptive decomposition and remapping algorithms for object-space-parallel direct volume rendering of unstructured grids
Cevdet Aykanat, Berkant Barla Cambazoglu, Ferit Findik, Tahsin M. Kurç |
J. Parallel Distributed Comput. | 2 |
| 2007 | Hypergraph-Partitioning-Based Remapping Models for Image-Space-Parallel Direct Volume Rendering of Unstructured Grids
Berkant Barla Cambazoglu, Cevdet Aykanat |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2006 | Performance of query processing implementations in ranking-based text retrieval systems using inverted indices
Berkant Barla Cambazoglu, Cevdet Aykanat |
Inf. Process. Manag. | 1 |