EDBT 2026 Demo / reviewers in the wild / expert
Ronny Lempel
dblp:60/838
· DBLP profile ↗
42ranked-venue papers
12as first author
1since 2021 · last 2022
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 37 · 9 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 9 · 1 first-authorComputer networks · 2 · 2 first-authorSoftware engineering, systems software and programming languages · 1Theory of computation · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
20 papers |
Information retrieval · 51% Recommender systems · 24% Spatial and temporal data management · 6% | |
| Theoretical computer science
5 papers |
Algorithmic game theory and mechanism design · 80% Mathematical optimization · 15% Approximation and online algorithms · 3% | |
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Cloud and datacenter computing · 46% Parallel and multicore computing · 46% Memory systems · 5% | |
| Artificial intelligence
2 papers |
Reinforcement learning · 50% Face, body and person analysis · 44% Information extraction and text analysis · 6% |
Topics — the 30 heaviest of 58, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Cloud and datacenter computing › resource management
cloud resource management |
0.6 | 1 | 2022 | Truthful Online Scheduling of Cloud Workloads under Uncertainty · WWW 2022 |
Parallel and multicore computing › task scheduling
online scheduling |
0.6 | 1 | 2022 | Truthful Online Scheduling of Cloud Workloads under Uncertainty · WWW 2022 |
Algorithmic game theory and mechanism design
mechanism design |
0.6 | 1 | 2022 | Truthful Online Scheduling of Cloud Workloads under Uncertainty · WWW 2022 |
Algorithmic game theory and mechanism design › resource allocation
truthful scheduling |
0.6 | 1 | 2022 | Truthful Online Scheduling of Cloud Workloads under Uncertainty · WWW 2022 |
Recommender systems
collaborative filtering |
0.4 | 2 | 2015 | Budget-Constrained Item Cold-Start Handling in Collaborative Filtering Recommenders via Optimal Design · WWW 2015 Care to comment?: recommendations for commenting on news stories · WWW 2012 |
Recommender systems
cold-start recommendation |
0.3 | 2 | 2015 | Budget-Constrained Item Cold-Start Handling in Collaborative Filtering Recommenders via Optimal Design · WWW 2015 Adaptive bootstrapping of recommender systems using decision trees · WSDM 2011 |
Information retrieval › search engines
search engine caching |
0.2 | 2 | 2010 | Caching search engine results over incremental indices · WWW 2010 Caching search engine results over incremental indices · SIGIR 2010 |
Mathematical optimization › experimental design
optimal experimental design |
0.2 | 1 | 2015 | Budget-Constrained Item Cold-Start Handling in Collaborative Filtering Recommenders via Optimal Design · WWW 2015 |
Machine learning › Reinforcement learning › exploration › multi-robot exploration
distributed exploration |
0.2 | 1 | 2013 | Distributed Exploration in Multi-Armed Bandits · NIPS 2013 |
Machine learning › Reinforcement learning
multi-armed bandit |
0.2 | 1 | 2013 | Distributed Exploration in Multi-Armed Bandits · NIPS 2013 |
Information retrieval
query log analysis |
0.2 | 1 | 2013 | Expediting search trend detection via prediction of query counts · WSDM 2013 |
Information retrieval › ranking
rank aggregation |
0.2 | 1 | 2013 | Rank quantization · WSDM 2013 |
Computer vision › Face, body and person analysis › face recognition
face annotation |
0.1 | 1 | 2012 | Lightweight automatic face annotation in media pages · WWW 2012 |
Computer vision › Face, body and person analysis
face recognition |
0.1 | 1 | 2012 | Lightweight automatic face annotation in media pages · WWW 2012 |
Indexing and storage engines › index maintenance
incremental indexing |
0.1 | 2 | 2010 | Caching search engine results over incremental indices · WWW 2010 Caching search engine results over incremental indices · SIGIR 2010 |
Information retrieval › distributed information retrieval
distributed search |
0.1 | 1 | 2011 | Inverted index compression via online document routing · WWW 2011 |
Information retrieval › indexing › index compression
inverted index compression |
0.1 | 1 | 2011 | Inverted index compression via online document routing · WWW 2011 |
Information retrieval › indexing
search engine indexing |
0.1 | 1 | 2011 | Inverted index compression via online document routing · WWW 2011 |
Database system architecture and tuning
cache invalidation |
0.1 | 1 | 2010 | Caching search engine results over incremental indices · SIGIR 2010 |
Spatial and temporal data management › location data
points of interest |
0.1 | 1 | 2010 | Constructing travel itineraries from tagged geo-temporal breadcrumbs · WWW 2010 |
Information retrieval › search engines
search engine architecture |
0.1 | 1 | 2010 | Caching search engine results over incremental indices · WWW 2010 |
Information retrieval
search engines |
0.1 | 3 | 2003 | Predictive caching and prefetching of query results in search engines · WWW 2003 Knowledge encapsulation for focused search from pervasive devices · ACM Trans. Inf. Syst. 2002 PicASHOW: pictorial authority search by hyperlinks on the web · ACM Trans. Inf. Syst. 2002 |
Information retrieval › search interfaces
faceted search |
0.1 | 1 | 2008 | Beyond basic faceted search · WSDM 2008 |
Query processing and optimization › runtime optimization › prefetching
query result prefetching |
0.1 | 2 | 2003 | Predictive caching and prefetching of query results in search engines · WWW 2003 Optimizing Result Prefetching in Web Search Engines with Segmented Indices · VLDB 2002 |
Information retrieval › web search
link analysis |
0.1 | 3 | 2002 | PicASHOW: pictorial authority search by hyperlinks on the web · ACM Trans. Inf. Syst. 2002 SALSA: the stochastic approach for link-structure analysis · ACM Trans. Inf. Syst. 2001 PicASHOW: pictorial authority search by hyperlinks on the Web · WWW 2001 |
Information retrieval › document retrieval › structured document retrieval
focused retrieval |
0.1 | 2 | 2002 | Knowledge encapsulation for focused search from pervasive devices · ACM Trans. Inf. Syst. 2002 Knowledge encapsulation for focused search from pervasive devices · WWW 2001 |
Machine learning and data management
data management for machine learning |
0.1 | 1 | 2015 | Budget-Constrained Item Cold-Start Handling in Collaborative Filtering Recommenders via Optimal Design · WWW 2015 |
Information retrieval › ranking › graph-based ranking
authority ranking |
0.1 | 2 | 2001 | SALSA: the stochastic approach for link-structure analysis · ACM Trans. Inf. Syst. 2001 PicASHOW: pictorial authority search by hyperlinks on the Web · WWW 2001 |
Data mining
clustering |
0.1 | 1 | 2006 | Cluster Ranking with an Application to Mining Mailbox Networks · ICDM 2006 |
Data mining › clustering
cluster ranking |
0.1 | 1 | 2006 | Cluster Ranking with an Application to Mining Mailbox Networks · ICDM 2006 |
Methods — techniques the papers use, named apart from their topics
social welfare optimization · 1.1online algorithm design · 1.1optimal design · 0.4approximation algorithm · 0.4quadratic-time algorithm · 0.3greedy approximation · 0.3transfer learning · 0.3ensemble of text analysis and vision components · 0.3latent factor modeling · 0.1content-based filtering · 0.1document reordering · 0.1decision tree learning · 0.1knowledge encapsulation · 0.1integrated cohesion · 0.1probabilistic user model · 0.0SLRU · 0.0LRU · 0.0topic-specific repository · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Truthful Online Scheduling of Cloud Workloads under UncertaintyabstractCloud computing customers often submit repeating jobs and computation pipelines on approximately regular schedules, with arrival and running times that exhibit variance. This pattern, typical of training tasks in machine learning, allows customers to partially predict future job requirements. We develop a model of cloud computing platforms that receive statements of work (SoWs) in an online fashion. The SoWs describe future jobs whose arrival times and durations are probabilistic, and whose utility to the submitting agents declines with completion time. The arrival and duration distributions, as well as the utility functions, are considered private customer information and are reported by strategic agents to a scheduler that is optimizing for social welfare. Moshe Babaioff, Ronny Lempel, Brendan Lucier, Ishai Menache, Aleksandrs Slivkins, Sam Chiu-wai Wong |
WWW | 2 |
| 2017 | Personalization is a Two-Way StreetabstractRecommender systems are first and foremost about matching users with items the systems believe will delight them. The "main street" of personalization is thus about modeling users and items, and matching per user the items predicted to best satisfy the user. This holds for both collaborative filtering and content-based methods. In content discovery engines, difficulties arise from the fact that the content users natively consume on publisher sites does not necessarily match the sponsored content that drives the monetization and sustains those engines. The first part of this talk addresses this gap by sharing lessons learned and by discussing how the gap may be bridged at scale with proper techniques. Ronny Lempel |
RecSys | 1 |
| 2016 | Powering Content Discovery through Scalable, Realtime Profiling of Users' Content PreferencesabstractOutbrain is the Web's leading content discovery service, recommending billions of stories daily to a global audience across many of the world's most prestigious and respected publishers. Outbrain's recommendation technology com- bines contextual cues with personalization, where the per- sonalization aspects are a combination of content-based and collaborative filtering techniques. This paper, and the accompanying demo, offer a behind- the-scenes view of the content-based aspects of Outbrain's personalization technology. We detail the types of features we extract from content, as well as the attributes we keep in each user's content-affinity profile. We then describe and demonstrate how we update each user's profile, in real time, as the user consumes content while browsing the Web. Ido Tamir, Roy Bass, Guy Kobrinsky, Baruch Brutman, Ronny Lempel, Yoram Dayagi |
RecSys | 5 |
| 2015 | Watch-It-Next: A Contextual TV Recommendation System
Michal Aharon, Eshcar Hillel, Amit Kagian, Ronny Lempel, Hayim Makabee, Raz Nissim |
ECML/PKDD (3) | 4 |
| 2015 | Budget-Constrained Item Cold-Start Handling in Collaborative Filtering Recommenders via Optimal DesignabstractIt is well known that collaborative filtering (CF) based recommender systems provide better modeling of users and items associated with considerable rating history. The lack of historical ratings results in the user and the item cold-start problems. The latter is the main focus of this work. Most of the current literature addresses this problem by integrating content-based recommendation techniques to model the new item. However, in many cases such content is not available, and the question arises is whether this problem can be mitigated using CF techniques only. We formalize this problem as an optimization problem: given a new item, a pool of available users, and a budget constraint, select which users to assign with the task of rating the new item in order to minimize the prediction error of our model. We show that the objective function is monotone-supermodular, and propose efficient optimal design based algorithms that attain an approximation to its optimum. Our findings are verified by an empirical study using the Netflix dataset, where the proposed algorithms outperform several baselines for the problem at hand. Oren Anava, Shahar Golan, Nadav Golbandi, Zohar S. Karnin, Ronny Lempel, Oleg Rokhlenko, Oren Somekh |
WWW | 5 |
| 2014 | Approximately optimal facet value selection
Sonya Liberman, Ronny Lempel |
Sci. Comput. Program. | 2 |
| 2013 | Distributed Exploration in Multi-Armed BanditsabstractWe study exploration in Multi-Armed Bandits (MAB) in a setting where~$k$ players collaborate in order to identify an $\epsilon$-optimal arm. Our motivation comes from recent employment of MAB algorithms in computationally intensive, large-scale applications. Our results demonstrate a non-trivial tradeoff between the number of arm pulls required by each of the players, and the amount of communication between them. In particular, our main result shows that by allowing the $k$ players to communicate \emph{only once}, they are able to learn $\sqrt{k}$ times faster than a single player. That is, distributing learning to $k$ players gives rise to a factor~$\sqrt{k}$ parallel speed-up. We complement this result with a lower bound showing this is in general the best possible. On the other extreme, we present an algorithm that achieves the ideal factor $k$ speed-up in learning performance, with communication only logarithmic in~$1/\epsilon$. Eshcar Hillel, Zohar S. Karnin, Tomer Koren, Ronny Lempel, Oren Somekh |
NIPS | 4 |
| 2013 | OFF-set: one-pass factorization of feature sets for online recommendation in persistent cold start settingsabstractOne of the most challenging recommendation tasks is recommending to a new, previously unseen user. This is known as the user cold start problem. Assuming certain features or attributes of users are known, one approach for handling new users is to initially model them based on their features. Michal Aharon, Natalie Aizenberg, Edward Bortnikov, Ronny Lempel, Roi Adadi, Tomer Benyamini, Liron Levin, Ran Roth, Ohad Serfaty |
RecSys | 4 |
| 2013 | Expediting search trend detection via prediction of query countsabstractThe massive volume of queries submitted to major Web search engines reflects human interest at a global scale. While the popularity of many search queries is stable over time or fluctuates with periodic regularity, some queries experience a sudden and ephemeral rise in popularity that is unexplained by their past volumes. Typically the popularity surge is precipitated by some real-life event in the news cycle. Such queries form what are known as search trends. All major search engines, using query log analysis and other signals, invest in detecting such trends. The goal is to surface trends accurately, with low latency relative to the actual event that sparked the trend. Nadav Golbandi, Liran Katzir 0001, Yehuda Koren, Ronny Lempel |
WSDM | 4 |
| 2013 | Rank quantizationabstractWe study the problem of aggregating and summarizing partial orders, on a large scale. Our motivation is two-fold: to discover elements at similar preference levels and to reduce the number of bits needed to store an element's position in a full ranking.We proceed in two steps: first, we find a total order by linearizing the rankings induced by the multiple partial orders and removing potentially inconsistent pairwise preferences. Next, given a total order, we introduce and formalize the rank quantization problem, which intuitively aims to bucketize the total order in a manner that mostly preserves the relations appearing in the partial orders. We show an exact quadratic-time quantization algorithm, as well as a greedy 2/3-approximation algorithm whose running is substantially faster on sparse instances. As an application, we aggregate rankings of top-10 search results over millions of search engine queries, approximately reproducing and then efficiently encoding the underlying static ranks used by the engine. We evaluate the performance of our algorithms on a web dataset of 12 million(2^{23.5}) unique pages and show that we can quantize the pages' static ranks using as few as eight bits, with only a minor degradation in search quality. Ravi Kumar 0001, Ronny Lempel, Roy Schwartz 0002, Sergei Vassilvitskii |
WSDM | 2 |
| 2012 | Modeling Transactional Queries via Templates
Edward Bortnikov, Pinar Donmez, Amit Kagian, Ronny Lempel |
ECIR | 4 |
| 2012 | Dynamic personalized recommendation of comment-eliciting storiesabstractMedia Websites often solicit users' comments on content items such as videos, news stories, blog posts, etc. Commenting activity increases user engagement with the sites, by both comment writers and readers, and so sites are looking for ways to increase the volume of comments. This work develops a recommender system aiming to present users with items -- news stories, in our case -- on which they are likely to comment. We combine items' content with a collaborative-filtering approach (utilizing users' co-commenting patterns) in a latent factor modeling framework. Building upon previous work, we focus on a continuous, real-time approach to address the problem above. After an initial training period during which commenting activity of users is observed, the system is tested at each subsequent comment submission event by predicting which story is being commented on by a given user at a given time. Our results show that we are able to overcome the site's inherent presentation bias and outperform a strong baseline as users' commenting history grows. Michal Aharon, Amit Kagian, Ronny Lempel, Yehuda Koren |
RecSys | 3 |
| 2012 | Recommendation challenges in web media settingsabstractThis paper calls out several research challenges in the art of recommendation technology as applied in Web media sites. One particular characteristic of such recommendation settings is the relative low cost of falsely recommending an irrelevant item, which means that recommendation schemes can be less conservative and more exploratory. This also creates opportunities for better item cold-start handling. Other technical difficulties include analyzing offline data that is heavily biased by the site's appearance, and in a related vein -- once a recommendation module's appearance has been designed -- defining the correct metrics by which to measure it. Also called out are tradeoffs between personalization and contextualization, as are novel schemes that aim at recommending sets and sequences of items. Ronny Lempel |
RecSys | 1 |
| 2012 | Lightweight automatic face annotation in media pagesabstractLabeling human faces in images contained in Web media stories enables enriching the user experience offered by media sites. We propose a lightweight framework for automatic image annotation that exploits named entities mentioned in the article to significantly boost the accuracy of face recognition. While previous works in the area labor to train comprehensive offline visual models for a pre-defined universe of candidates, our approach models the people mentioned in a given story on the y, using a standard Web image search engine as an image sampling mechanism. We overcome multiple sources of noise introduced by this ad-hoc process, to build a fast and robust end-to-end system from off-the-shelf error-prone text analysis and machine vision components. In experiments conducted on approximately 900 faces depicted in 500 stories from a major celebrity news website, we were able to correctly label 81.5% of the faces while mislabeling 14.8% of them. Dmitri Perelman, Edward Bortnikov, Ronny Lempel, Roman Sandler |
WWW | 3 |
| 2012 | Care to comment?: recommendations for commenting on news storiesabstractMany websites provide commenting facilities for users to express their opinions or sentiments with regards to content items, such as, videos, news stories, blog posts, etc. Previous studies have shown that user comments contain valuable information that can provide insight on Web documents and may be utilized for various tasks. This work presents a model that predicts, for a given user, suitable news stories for commenting. The model achieves encouraging results regarding the ability to connect users with stories they are likely to comment on. This provides grounds for personalized recommendations of stories to users who may want to take part in their discussion. We combine a content-based approach with a collaborative-filtering approach (utilizing users' co-commenting patterns) in a latent factor modeling framework. We experiment with several variations of the model's loss function in order to adjust it to the problem domain. We evaluate the results on two datasets and show that employing co-commenting patterns improves upon using content features alone, even with as few as two available comments per story. Finally, we try to incorporate available social network data into the model. Interestingly, the social data does not lead to substantial performance gains, suggesting that the value of social data for this task is quite negligible. Erez Shmueli, Amit Kagian, Yehuda Koren, Ronny Lempel |
WWW | 4 |
| 2011 | Caching for Realtime Search
Edward Bortnikov, Ronny Lempel, Kolman Vornovitsky |
ECIR | 2 |
| 2011 | Adaptive bootstrapping of recommender systems using decision treesabstractRecommender systems perform much better on users for which they have more information. This gives rise to a problem of satisfying users new to a system. The problem is even more acute considering that some of these hard to profile new users judge the unfamiliar system by its ability to immediately provide them with satisfying recommendations, and may quickly abandon the system when disappointed. Rapid profiling of new users by a recommender system is often achieved through a bootstrapping process - a kind of an initial interview - that elicits users to provide their opinions on certain carefully chosen items or categories. The elicitation process becomes particularly effective when adapted to users' responses, making best use of users' time by dynamically modifying the questions to improve the evolving profile. In particular, we advocate a specialized version of decision trees as the most appropriate tool for this task. We detail an efficient tree learning algorithm, specifically tailored to the unique properties of the problem. Several extensions to the tree construction are also introduced, which enhance the efficiency and utility of the method. We implemented our methods within a movie recommendation service. The experimental study delivered encouraging results, with the tree-based bootstrapping process significantly outperforming previous approaches. Nadav Golbandi, Yehuda Koren, Ronny Lempel |
WSDM | 3 |
| 2011 | Inverted index compression via online document routingabstractModern search engines are expected to make documents searchable shortly after they appear on the ever changing Web. To satisfy this requirement, the Web is frequently crawled. Due to the sheer size of their indexes, search engines distribute the crawled documents among thousands of servers in a scheme called local index-partitioning, such that each server indexes only several million pages. To ensure documents from the same host (e.g., www.nytimes.com) are distributed uniformly over the servers, for load balancing purposes, random routing of documents to servers is common. To expedite the time documents become searchable after being crawled, documents may be simply appended to the existing index partitions. However, indexing by merely appending documents, results in larger index sizes since document reordering for index compactness is no longer performed. This, in turn, degrades search query processing performance which depends heavily on index sizes. Gal Lavee, Ronny Lempel, Edo Liberty, Oren Somekh |
WWW | 2 |
| 2010 | On bootstrapping recommender systemsabstractRecommender systems perform much better on users for which they have more information. This gives rise to a problem of satisfying users new to a system. The problem is even more acute considering that some of these hard to profile new users judge the unfamiliar system by its ability to immediately provide them with satisfying recommendations, and may be the quickest to abandon the system when disappointed. Rapid profiling of new users is often achieved through a bootstrapping process - a kind of an initial interview - that elicits users to provide their opinions on certain carefully chosen items or categories. This work offers a new bootstrapping method, which is based on a concrete optimization goal, thereby handily outperforming known approaches in our tests. Nadav Golbandi, Yehuda Koren, Ronny Lempel |
CIKM | 3 |
| 2010 | Caching search engine results over incremental indicesabstractA Web search engine must update its index periodically to incorporate changes to the Web. We argue in this paper that index updates fundamentally impact the design of search engine result caches, a performance-critical component of modern search engines. Index updates lead to the problem of cache invalidation: invalidating cached entries of queries whose results have changed. Naive approaches, such as flushing the entire cache upon every index update, lead to poor performance and in fact, render caching futile when the frequency of updates is high. Solving the invalidation problem efficiently corresponds to predicting accurately which queries will produce different results if re-evaluated, given the actual changes to the index. Roi Blanco, Edward Bortnikov, Flavio Paiva Junqueira, Ronny Lempel, Luca Telloli, Hugo Zaragoza |
SIGIR | 4 |
| 2010 | Caching search engine results over incremental indicesabstractA Web search engine must update its index periodically to incorporate changes to the Web, and we argue in this work that index updates fundamentally impact the design of search engine result caches. Index updates lead to the problem of cache invalidation: invalidating cached entries of queries whose results have changed. To enable efficient invalidation of cached results, we propose a framework for developing invalidation predictors and some concrete predictors. Evaluation using Wikipedia documents and a query log from Yahoo! shows that selective invalidation of cached search results can lower the number of query re-evaluations by as much as 30% compared to a baseline time-to-live scheme, while returning results of similar freshness. Roi Blanco, Edward Bortnikov, Flavio Paiva Junqueira, Ronny Lempel, Luca Telloli, Hugo Zaragoza |
WWW | 4 |
| 2010 | Constructing travel itineraries from tagged geo-temporal breadcrumbsabstractVacation planning is a frequent laborious task which requires skilled interaction with a multitude of resources. This paper develops an end-to-end approach for constructing intra-city travel itineraries automatically by tapping a latent source reflecting geo-temporal breadcrumbs left by millions of tourists. In particular, the popular rich media sharing site, Flickr, allows photos to be stamped by the date and time of when they were taken, and be mapped to Points Of Interest (POIs) by latitude-longitude information as well as semantic metadata (e.g., tags) that describe them. Munmun De Choudhury, Moran Feldman, Sihem Amer-Yahia, Nadav Golbandi, Ronny Lempel, Cong Yu 0001 |
WWW | 5 |
| 2008 | Beyond basic faceted searchabstractThis paper extends traditional faceted search to support richer information discovery tasks over more complex data models. Our first extension adds exible, dynamic business intelligence aggregations to the faceted application, enabling users to gain insight into their data that is far richer than just knowing the quantities of documents belonging to each facet. We see this capability as a step toward bringing OLAP capabilities, traditionally supported by databases over relational data, to the domain of free-text queries over metadata-rich content. Our second extension shows how one can efficiently extend a faceted search engine to support correlated facets - a more complex information model in which the values associated with a document across multiple facets are not independent. We show that by reducing the problem to a recently solved tree-indexing scenario, data with correlated facets can be efficiently indexed and retrieved Ori Ben-Yitzhak, Nadav Golbandi, Nadav Har'El, Ronny Lempel, Andreas Neumann 0001, Shila Ofek-Koifman, Dafna Sheinwald, Eugene J. Shekita, Benjamin Sznajder, Sivan Yogev |
WSDM | 4 |
| 2008 | Cluster ranking with an application to mining mailbox networks
Ziv Bar-Yossef, Ido Guy, Ronny Lempel, Yoelle Maarek, Vladimir Soroka |
Knowl. Inf. Syst. | 3 |
| 2007 | Just in time indexing for up to the second searchabstractE-commerce and intranet search systems require newly arriving content to be indexed and made available for search within minutes or hours of arrival. Applications such as file system and email search demand even faster turnaround from search systems, requiring new content to become available for search almost instantaneously. However, incrementally updating inverted indices, which are the predominant datastructure used in search engines, is an expensive operation that most systems avoid performing at high rates. Ronny Lempel, Yosi Mass, Shila Ofek-Koifman, Dafna Sheinwald, Yael Petruschka, Ron Sivan |
CIKM | 1 |
| 2007 | Efficient Indexing of Versioned Document Sequences
Michael Herscovici, Ronny Lempel, Sivan Yogev |
ECIR | 2 |
| 2006 | Indexing Shared Content in Information Retrieval Systems
Andrei Z. Broder, Nadav Eiron, Marcus Fontoura, Michael Herscovici, Ronny Lempel, John McPherson, Runping Qi, Eugene J. Shekita |
EDBT | 5 |
| 2006 | Cluster Ranking with an Application to Mining Mailbox NetworksabstractWe initiate the study of a new clustering framework, called cluster ranking. Rather than simply partitioning a network into clusters, a cluster ranking algorithm also orders the clusters by their strength. To this end, we introduce a novel strength measure for clusters - the integrated cohesion - which is applicable to arbitrary weighted networks. We then present C-Rank: a new cluster ranking algorithm. Given a network with arbitrary pairwise similarity weights, C-Rank creates a list of overlapping clusters and ranks them by their integrated cohesion. We provide extensive theoretical and empirical analysis of C-Rank and show that it is likely to have high precision and recall. Our experiments focus on mining mailbox networks. A mailbox network is an egocentric social network, consisting of contacts with whom an individual exchanges email. Ties among contacts are represented by the frequency of their co-occurrence on message headers. C-Rank is well suited to mine such networks, since they are abundant with overlapping communities of highly variable strengths. We demonstrate the effectiveness of C-Rank on the Enron data set, consisting of 130 mailbox networks. Ziv Bar-Yossef, Ido Guy, Ronny Lempel, Yoelle Maarek, Vladimir Soroka |
ICDM | 3 |
| 2006 | Efficient PageRank approximation via graph aggregation
Andrei Z. Broder, Ronny Lempel, Farzin Maghoul, Jan O. Pedersen 0001 |
Inf. Retr. | 2 |
| 2005 | Rank-Stability and Rank-Similarity of Link-Based Web Ranking Algorithms in Authority-Connected Graphs
Ronny Lempel, Shlomo Moran |
Inf. Retr. | 1 |
| 2004 | Scaling IR-system evaluation using term relevance setsabstractThis paper describes an evaluation method based on Term Relevance Sets Trels that measures an IR system's quality by examining the content of the retrieved results rather than by looking for pre-specified relevant pages. Trels consist of a list of terms believed to be relevant for a particular query as well as a list of irrelevant terms. The proposed method does not involve any document relevance judgments, and as such is not adversely affected by changes to the underlying collection. Therefore, it can better scale to very large, dynamic collections such as the Web. Moreover, this method can evaluate a system's effectiveness on an updatable "live" collection, or on collections derived from different data sources. Our experiments show that the proposed method is very highly correlated with official TREC measures. Einat Amitay, David Carmel, Ronny Lempel, Aya Soffer |
SIGIR | 3 |
| 2004 | Trend detection through temporal link analysisabstractAbstract Although time has been recognized as an important dimension in the co‐citation literature, to date it has not been incorporated into the analogous process of link analysis on the Web. In this paper, we discuss several aspects and uses of the time dimension in the context of Web information retrieval. We describe the ideal case—where search engines track and store temporal data for each of the pages in their repository, assigning timestamps to the hyperlinks embedded within the pages. We introduce several applications which benefit from the availability of such timestamps. To demonstrate our claims, we use a somewhat simplistic approach, which dates links by approximating the age of the page's content. We show that by using this crude measure alone it is possible to detect and expose significant events and trends. We predict that by using more robust methods for tracking modifications in the content of pages, search engines will be able to provide results that are more timely and better reflect current real‐life trends than those they provide today. Einat Amitay, David Carmel, Michael Herscovici, Ronny Lempel, Aya Soffer |
J. Assoc. Inf. Sci. Technol. | 4 |
| 2004 | Competitive caching of query results in search engines
Ronny Lempel, Shlomo Moran |
Theor. Comput. Sci. | 1 |
| 2004 | Optimizing result prefetching in web search engines with segmented indicesabstractWe study the process in which search engines with segmented indices serve queries. In particular, we investigate the number of result pages that search engines should prepare during the query processing phase.Search engine users have been observed to browse through very few pages of results for queries that they submit. This behavior of users suggests that prefetching many results upon processing an initial query is not efficient, since most of the prefetched results will not be requested by the user who initiated the search. However, a policy that abandons result prefetching in favor of retrieving just the first page of search results might not make optimal use of system resources either.We argue that for a certain behavior of users, engines should prefetch a constant number of result pages per query. We define a concrete query processing model for search engines with segmented indices, and analyze the cost of such prefetching policies. Based on these costs, we show how to determine the constant that optimizes the prefetching policy. Our results are mostly applicable to local index partitions of the inverted files, but are also applicable to processing short queries in global index architectures. Ronny Lempel, Shlomo Moran |
ACM Trans. Internet Techn. | 1 |
| 2003 | Predictive caching and prefetching of query results in search enginesabstractWe study the caching of query result pages in Web search engines. Popular search engines receive millions of queries per day, and efficient policies for caching query results may enable them to lower their response time and reduce their hardware requirements. We present PDC (probability driven cache), a novel scheme tailored for caching search results, that is based on a probabilistic model of search engine users. We then use a trace of over seven million queries submitted to the search engine AltaVista to evaluate PDC, as well as traditional LRU and SLRU based caching schemes. The trace driven simulations show that PDC outperforms the other policies. We also examine the prefetching of search results, and demonstrate that prefetching can increase cache hit ratios by 50% for large caches, and can double the hit ratios of small caches. When integrating prefetching into PDC, we attain hit ratios of over 0.53. Ronny Lempel, Shlomo Moran |
WWW | 1 |
| 2002 | Optimizing Result Prefetching in Web Search Engines with Segmented Indices
Ronny Lempel, Shlomo Moran |
VLDB | 1 |
| 2002 | Knowledge encapsulation for focused search from pervasive devicesabstractMobile knowledge seekers often need access to information on the Web during a meeting or on the road, while away from their desktop. A common practice today is to use pervasive devices such as Personal Digital Assistants or mobile phones. However, these devices have inherent constraints (e.g., slow communication, form factor) which often make information discovery tasks impractical.In this paper, we present a new focused-search approach specifically oriented for the mode of work and the constraints dictated by pervasive devices. It combines focused search within specific topics with encapsulation of topic-specific information in a persistent repository. One key characteristic of these persistent repositories is that their footprint is small enough to fit on local devices, and yet they are rich enough to support many information discovery tasks in disconnected mode. More specifically, we suggest a representation for topic-specific information based on "knowledge-agent bases" that comprise all the information necessary to access information about a topic (under the form of key concepts and key Web pages) and assist in the full search process from query formulation assistance to result scanning on the device itself. The key contribution of our work is the coupling of focused search with encapsulated knowledge representation making information discovery from pervasive devices practical as well as efficient. We describe our model in detail and demonstrate its aspects through sample scenarios. Yariv Aridor, David Carmel, Yoelle Maarek, Aya Soffer, Ronny Lempel |
ACM Trans. Inf. Syst. | 5 |
| 2002 | PicASHOW: pictorial authority search by hyperlinks on the webabstractWe describe PicASHOW, a fully automated WWW image retrieval system that is based on several link-structure analyzing algorithms. Our basic premise is that a pagepdisplays (or links to) an image when the author ofpconsiders the image to be of value to the viewers of the page. We thus extend some well known link-based WWWpage retrievalschemes to the context of image retrieval.PicASHOW's analysis of the link structure enables it to retrieve relevant images even when those are stored in files with meaningless names. The same analysis also allows it to identifyimage containersandimage hubs. We define these as Web pages that are rich in relevant images, or from which many images are readily accessible.PicASHOW requires no image analysis whatsoever and no creation of taxonomies for preclassification of the Web's images. It can be implemented by standard WWW search engines with reasonable overhead, in terms of both computations and storage, and with no change to user query formats. It can thus be used to easily add image retrieving capabilities to standard search engines.Our results demonstrate that PicASHOW, while relying almost exclusively on link analysis, compares well with dedicated WWW image retrieval systems. We conclude that link analysis, a proven effective technique for Web page search, can improve the performance of Web image retrieval, as well as extend its definition to include the retrieval of image hubs and containers. Ronny Lempel, Aya Soffer |
ACM Trans. Inf. Syst. | 1 |
| 2001 | Knowledge encapsulation for focused search from pervasive devicesabstractArticle Knowledge encapsulation for focused search from pervasive devices Share on Authors: Yariv Aridor IBM Research Lab in Haifa, Matam, Haifa 31905, Israel IBM Research Lab in Haifa, Matam, Haifa 31905, IsraelView Profile , David Carmel IBM Research Lab in Haifa, Matam, Haifa 31905, Israel IBM Research Lab in Haifa, Matam, Haifa 31905, IsraelView Profile , Yoëlle S. Maarek IBM Research Lab in Haifa, Matam, Haifa 31905, Israel IBM Research Lab in Haifa, Matam, Haifa 31905, IsraelView Profile , Aya Soffer IBM Research Lab in Haifa, Matam, Haifa 31905, Israel IBM Research Lab in Haifa, Matam, Haifa 31905, IsraelView Profile , Ronny Lempel Department of Computer Science, Technion, Haifa 32000, Israel Department of Computer Science, Technion, Haifa 32000, IsraelView Profile Authors Info & Claims WWW '01: Proceedings of the 10th international conference on World Wide WebMay 2001 Pages 754–764https://doi.org/10.1145/371920.372195Online:01 April 2001Publication History 8citation368DownloadsMetricsTotal Citations8Total Downloads368Last 12 Months2Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Yariv Aridor, David Carmel, Yoelle Maarek, Aya Soffer, Ronny Lempel |
WWW | 5 |
| 2001 | PicASHOW: pictorial authority search by hyperlinks on the WebabstractArticle Share on PicASHOW: pictorial authority search by hyperlinks on the Web Authors: Ronny Lempel Department of Computer Science, The Technion, Haifa 32000, Israel Department of Computer Science, The Technion, Haifa 32000, IsraelView Profile , Aya Soffer IBM Research Lab in Haifa, Matam, Haifa 31905, Israel IBM Research Lab in Haifa, Matam, Haifa 31905, IsraelView Profile Authors Info & Claims WWW '01: Proceedings of the 10th international conference on World Wide WebMay 2001 Pages 438–448https://doi.org/10.1145/371920.372098Online:01 April 2001Publication History 37citation449DownloadsMetricsTotal Citations37Total Downloads449Last 12 Months15Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Ronny Lempel, Aya Soffer |
WWW | 1 |
| 2001 | SALSA: the stochastic approach for link-structure analysisabstractToday, when searching for information on the WWW, one usually performs a query through a term-based search engine. These engines return, as the query's result, a list of Web pages whose contents matches the query. For broad-topic queries, such searches often result in a huge set of retrieved documents, many of which are irrelevant to the user. However, much information is contained in the link-structure of the WWW. Information such as which pages are linked to others can be used to augment search algorithms. In this context, Jon Kleinberg introduced the notion of two distinct types of Web pages: hubs and authorities . Kleinberg argued that hubs and authorities exhibit a mutually reinforcing relationship : a good hub will point to many authorities, and a good authority will be pointed at by many hubs. In light of this, he dervised an algoirthm aimed at finding authoritative pages. We present SALSA, a new stochastic approach for link-structure analysis, which examines random walks on graphs derived from the link-structure. We show that both SALSA and Kleinberg's Mutual Reinforcement approach employ the same metaalgorithm. We then prove that SALSA is quivalent to a weighted in degree analysis of the link-sturcutre of WWW subgraphs, making it computationally more efficient than the Mutual reinforcement approach. We compare that results of applying SALSA to the results derived through Kleinberg's approach. These comparisions reveal a topological Phenomenon called the TKC effect which, in certain cases, prevents the Mutual reinforcement approach from identifying meaningful authorities. Ronny Lempel, Shlomo Moran |
ACM Trans. Inf. Syst. | 1 |
| 2000 | The stochastic approach for link-structure analysis (SALSA) and the TKC effect
Ronny Lempel, Shlomo Moran |
Comput. Networks | 1 |