EDBT 2026 Demo / reviewers in the wild / expert
Elad Yom-Tov
dblp:47/4458
· DBLP profile ↗
45ranked-venue papers
14as first author
4since 2021 · last 2026
0000-0002-2380-4584ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 31 · 9 first-author · 3 since 2021Artificial intelligence and machine learning · 11 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 3 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2Systems, architecture and hardware · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning To Generate Effective Health Advertisements
Yuval Saadaty, Elad Yom-Tov |
WWW | 2 |
| 2025 | Disparate Conditional Prediction in Multiclass ClassifiersabstractWe propose methods for auditing multiclass classifiers for fairness under multiclass equalized odds, by estimating the deviation from equalized odds when the classifier is not completely fair. We generalize to multiclass classifiers the measure of Disparate Conditional Prediction (DCP), originally suggested by Sabato & Yom-Tov (2020) for binary classifiers. DCP is defined as the fraction of the population for which the classifier predicts with conditional prediction probabilities that differ from the closest common baseline. We provide new local-optimization methods for estimating the multiclass DCP under two different regimes, one in which the conditional confusion matrices for each protected sub-population are known, and one in which these cannot be estimated, for instance, because the classifier is inaccessible or
because good-quality individual-level data is not available. These methods can be used to detect classifiers that likely treat a significant fraction of the population unfairly. Experiments demonstrate the accuracy of the methods. The code for the experiments is provided as supplementary material. Sivan Sabato, Eran Treister, Elad Yom-Tov |
ICML | 3 |
| 2025 | The In-Situ Effect of Offensive Ads on Search Engine UsersabstractUnscrupulous advertisers may try to increase attention to search ads by using offensive ads, which can increase attention and recall to the detriment of individuals and society. Here, we investigate whether offensive ads, when shown to search engine users, have such effects. We developed 12 search scenarios and created 4 versions of the search results page (SERP) for each scenario, where some of the ads were changed to be irrelevant and/or offensive. Crowdsourced judges found a strong correlation ( \(\geq 0.63\) ) between the reported number of annoying ads and the actual number of offensive and irrelevant ads, suggesting people conflate these attributes. Furthermore, we found that judges who assessed the SERPs for themselves reported lower positive affect and higher negative affect than judges asked to imagine the results were provided to someone else. In the latter case offensive ads also lead to slightly lower positive ( \(-4\%\) ) and higher negative affect ( \(+61\%\) ). Finally, in a recall test, only 6% of judges reported seeing an offensive ad when using search engines. Our work should further detract advertisers from using offensive ads since, in addition to previously documented adverse effects, such ads have a small but statistically significant negative effect on people’s emotional experience. Elad Yom-Tov, Liat Levontin |
ACM Trans. Inf. Syst. | 1 |
| 2021 | Algorithmic copywriting: automated generation of health-related advertisements to improve their performance
Brit Youngmann, Elad Yom-Tov, Ran Gilad-Bachrach, Danny Karmon |
Inf. Retr. J. | 2 |
| 2020 | Bounding the fairness and accuracy of classifiers from population statisticsabstractWe consider the study of a classification model whose properties are impossible to estimate using a validation set, either due to the absence of such a set or because access to the classifier, even as a black-box, is impossible. Instead, only aggregate statistics on the rate of positive predictions in each of several sub-populations are available, as well as the true rates of positive labels in each of these sub-populations. We show that these aggregate statistics can be used to lower-bound the discrepancy of a classifier, which is a measure that balances inaccuracy and unfairness. To this end, we define a new measure of unfairness, equal to the fraction of the population on which the classifier behaves differently, compared to its global, ideally fair behavior, as defined by the measure of equalized odds. We propose an efficient and practical procedure for finding the best possible lower bound on the discrepancy of the classifier, given the aggregate statistics, and demonstrate in experiments the empirical tightness of this lower bound, as well as its possible uses on various types of problems, ranging from estimating the quality of voting polls to measuring the effectiveness of patient identification from internet search queries. The code and data are available at https://github.com/sivansabato/bfa. Sivan Sabato, Elad Yom-Tov |
ICML | 2 |
| 2020 | The Automated Copywriter: Algorithmic Rephrasing of Health-Related Advertisements to Improve their PerformanceabstractSearch advertising is one of the most commonly-used methods of advertising. Past work has shown that search advertising can be employed to improve health by eliciting positive behavioral change. However, writing effective advertisements requires expertise and (possible expensive) experimentation, both of which may not be available to public health authorities wishing to elicit such behavioral changes, especially when dealing with a public health crises such as epidemic outbreaks. Brit Youngmann, Elad Yom-Tov, Ran Gilad-Bachrach, Danny Karmon |
WWW | 2 |
| 2020 | Screening for Cancer Using a Learning Internet Advertising SystemabstractStudies have shown that search engine queries are indicative of future diagnosis of several types of cancer. These studies were based on self-identification of illness and were limited in that diagnostic information could not be shared with screened individuals. Here I report on two studies that overcome these limitations. Advertisements were displayed on the Bing and Google ads systems to people who sought to self-diagnose one of three types of cancer. People who clicked on these ads were provided with clinically verified questionnaires and the outcomes of these questionnaires. A classifier trained to predict suspected cancer, inferred from questionnaire responses, from past Bing queries reached an area under the curve of 0.64. People who received information that their symptoms were consistent with suspected cancer increased searches for healthcare utilization. In a second study, questionnaire responses provided to the conversion optimization mechanism of the Google advertisement system enabled it to learn to identify people who were likely to have suspected cancer. Following a training period of approximately 10 days, 11% of people selected for showing of targeted campaign ads were found to have suspected cancer. These results demonstrate the utility of using modern advertising systems to identify people who are likely suffering from serious medical conditions. Elad Yom-Tov |
ACM Trans. Comput. Heal. | 1 |
| 2019 | Demographic differences in search engine use with implications for cohort selection
Elad Yom-Tov |
Inf. Retr. J. | 1 |
| 2018 | Discriminative Learning of Prediction IntervalsabstractIn this work we consider the task of constructing prediction intervals in an inductive batch setting. We present a discriminative learning framework which optimizes the expected error rate under a budget constraint on the interval sizes. Most current methods for constructing prediction intervals offer guarantees for a single new test point. Applying these methods to multiple test points can result in a high computational overhead and degraded statistical guarantees. By focusing on expected errors, our method allows for variability in the per-example conditional error rates. As we demonstrate both analytically and empirically, this flexibility can increase the overall accuracy, or alternatively, reduce the average interval size. While the problem we consider is of a regressive flavor, the loss we use is combinatorial. This allows us to provide PAC-style, finite-sample guarantees. Computationally, we show that our original objective is NP-hard, and suggest a tractable convex surrogate. We conclude with a series of experimental evaluations. Nir Rosenfeld, Yishay Mansour, Elad Yom-Tov |
AISTATS | 3 |
| 2018 | Detecting Parkinson's Disease from Interactions with a Search Engine: Is Expert Knowledge Sufficient?abstractParkinson's disease (PD) is a slowly progressing neurodegenerative disease with early manifestation of motor signs. Recently, there has been a growing interest in developing automatic tools that can assess motor function in PD patients. Here we show that mouse tracking data collected during people's interaction with a search engine can be used to distinguish PD patients from similar, non-diseased users and present a methodology developed for the diagnosis of PD from these data. The main challenge we address is the extraction of informative features from raw mouse tracking data. We do so in two complementary ways: First, we manually construct expert-recommended features, aiming to identify abnormalities in motor behaviors. Second, we use an unsupervised representation learning technique to map these raw data to high-level features. Using all the extracted features, a Random Forest classifier is then used to distinguish PD patients from controls, achieving an AUC of 0.92, while results using only expert-generated or auto-generated features are 0.87 and 0.83, respectively. Our results indicate that mouse tracking data can help in detecting users at early stages of the disease and that both expert-generated features and unsupervised techniques for feature generation are required to achieve the best possible performance. Liron I. Allerhand, Brit Youngmann, Elad Yom-Tov, David Arkadir |
CIKM | 3 |
| 2018 | Anxiety and Information Seeking: Evidence From Large-Scale Mouse TrackingabstractPeople seeking information through search engines are assumed to behave similarly, regardless of the topic which they are searching. Here we use mouse tracking, which is correlated with gaze, to show that the information seeking patterns of people differ dramatically depending on their level of anxiety at the time of the search. We investigate the behavior of people during searches for medical symptoms, ranging from benign indications, where users are not usually anxious, to ones which could harbinger life-threatening conditions, where extreme anxiety is expected. We show that for the latter, 90% of people never saw more than the top 67% of the screen, compared to over 95% scanned by people seeking information on benign symptoms, even though relevant documents are similarly distributed in the results pages to these queries. Based on this observation, we develop a model which can predict the level of anxiety experienced by a user, using attributes derived from mouse tracking data and other user interactions. The model achieves Kendall's Tau of 0.48 with the medical severity of the symptoms searched. We show the importance of using information about the users? level of anxiety as predicted by the model, when measuring search engine performance. Our results prove that ignoring this information can lead to significant over-estimation of performance. Additionally, we show the utility of the model in three special instances: where multiple symptoms are searched concurrently; where the searcher has an underlying medical condition; and when users seek information on ways to commit suicide. In the latter, our results demonstrate the importance of help-line notices, and emphasize the need to measure the effective number of results seen by the user. Our results indicate that measures of relevance which use anxiety information can lead to more accurate understanding of the quality of search results, especially when delivering potentially life-saving information to users. Brit Youngmann, Elad Yom-Tov |
WWW | 2 |
| 2017 | Inferring Individual Attributes from Search Engine Queries and Auxiliary InformationabstractInternet data has surfaced as a primary source for investigation of different aspects of human behavior. A crucial step in such studies is finding a suitable cohort (i.e., a set of users) that shares a common trait of interest to researchers. However, direct identification of users sharing this trait is often impossible, as the data available to researchers is usually anonymized to preserve user privacy. To facilitate research on specific topics of interest, especially in medicine, we introduce an algorithm for identifying a trait of interest in anonymous users. We illustrate how a small set of labeled examples, together with statistical information about the entire population, can be aggregated to obtain labels on unseen examples. We validate our approach using labeled data from the political domain. Luca Soldaini, Elad Yom-Tov |
WWW | 2 |
| 2016 | Social Media Research in the Health Domain
Luis Fernández-Luque, Fernando Martín-Sánchez, Ingmar Weber, Elad Yom-Tov, Carolyn Petersen |
AMIA | 4 |
| 2016 | Recommendations meet web browsing: enhancing collaborative filtering using internet browsing logsabstractCollaborative filtering (CF) recommendation systems are one of the most popular and successful methods for recommending products to people. CF systems work by finding similarities between different people according to their past purchases, and using these similarities to suggest possible items of interest. In this work we show that CF systems can be enhanced using Internet browsing data and search engine query logs, both of which represent a rich profile of individuals' interests. Royi Ronen, Elad Yom-Tov, Gal Lavee |
ICDE | 2 |
| 2016 | The Quality of Online Answers to Parents Who Suspect That Their Child Has an Autism Spectrum Disorder
Ayelet Ben-Sasson, Dan Pelleg, Elad Yom-Tov |
ICWSM | 3 |
| 2016 | Enhancing web search in the medical domain via query clarification
Luca Soldaini, Andrew Yates, Elad Yom-Tov, Ophir Frieder, Nazli Goharian |
Inf. Retr. J. | 3 |
| 2015 | Early Detection of Fraud Storms in the Cloud
Hani Neuvirth, Yehuda Finkelstein, Amit Hilbuch, Shai Nahum, Daniel Alon, Elad Yom-Tov |
ECML/PKDD (3) | 6 |
| 2015 | Modularity-Based Query Clustering for Identifying Users Sharing a Common ConditionabstractWe present an algorithm for identifying users who share a common condition from anonymized search engine logs. Input to the algorithm is a set of seed phrases that identify users with the condition of interest with high precision albeit at a very low recall. We expand the set of seed phrases by clustering queries according to the pages users clicked following these queries and the temporal ordering of queries within sessions, emphasizing the subgraph containing seed phrases. To this end, we extend modularity-based clustering such that it uses the information in the initial seed phrases as well as other queries of users in the population of interest. We evaluate the performance of the proposed method on two datasets, one of mood disorders and the other of anorexia, by classifying users according to the clusters in which they appeared and the phrases contained thereof, and show that the area under the receiver operating characteristic curve (AUC) obtained by these methods exceeds 0.87. These results demonstrate the value of our algorithm for both identifying users for future research and to gain better understanding of the language associated with the condition. Maayan Harel, Elad Yom-Tov |
SIGIR | 2 |
| 2015 | Learning About Health and Medicine from Internet DataabstractSurveys show that around 70% of US Internet users consult the Internet when they require medical information. People seek this information using both traditional search engines and via social media. The information created using the search process offers an unprecedented opportunity for applications to monitor and improve the quality of life of people with a variety of medical conditions. In recent years, research in this area has addressed public-health questions such as the effect of media on development of anorexia, developed tools for measuring influenza rates and assessing drug safety, and examined the effects of health information on individual wellbeing. This tutorial will show how Internet data can facilitate medical research, providing an overview of the state-of-the-art in this area. During the tutorial we will discuss the information which can be gleaned from a variety of Internet data sources, including social media, search engines, and specialized medical websites. We will provide an overview of analysis methods used in recent literature, and show how results can be evaluated using publicly-available health information and online experimentation. Finally, we will discuss ethical and privacy issues and possible technological solutions. This tutorial is intended for researchers of user generated content who are interested in applying their knowledge to improve health and medicine. Elad Yom-Tov, Ingemar J. Cox, Vasileios Lampos |
WSDM | 1 |
| 2015 | Assessing the impact of a health intervention via user-generated Internet contentabstractAssessing the effect of a health-oriented intervention by traditional epidemiological methods is commonly based only on population segments that use healthcare services. Here we introduce a complementary framework for evaluating the impact of a targeted intervention, such as a vaccination campaign against an infectious disease, through a statistical analysis of user-generated content submitted on web platforms. Using supervised learning, we derive a nonlinear regression model for estimating the prevalence of a health event in a population from Internet data. This model is applied to identify control location groups that correlate historically with the areas, where a specific intervention campaign has taken place. We then determine the impact of the intervention by inferring a projection of the disease rates that could have emerged in the absence of a campaign. Our case study focuses on the influenza vaccination program that was launched in England during the 2013/14 season, and our observations consist of millions of geo-located search queries to the Bing search engine and posts on Twitter. The impact estimates derived from the application of the proposed statistical framework support conventional assessments of the campaign. Vasileios Lampos, Elad Yom-Tov, Richard Pebody, Ingemar J. Cox |
Data Min. Knowl. Discov. | 2 |
| 2014 | Information is in the eye of the beholder: Seeking information on the MMR vaccine through an Internet search engine
Elad Yom-Tov, Luis Fernández-Luque |
AMIA | 1 |
| 2013 | Measuring inter-site engagementabstractMany large online providers offer a variety of content sites (e.g. news, sport, e-commerce). These providers endeavor to keep users accessing and interacting with their sites, that is to engage users by spending time using their sites and to return regularly to them. They do so by serving users the most relevant content in an attractive and enticing manner. Due to their highly varied content, each site is usually studied and optimized separately. However, these online providers aim not only to engage users with individual sites, but across all sites in their network. In these cases, site engagement should be examined not only within individual sites, but also across the entire content provider network. This paper investigates intersite engagement, that is, site engagement within a network of sites, by defining a global measure of engagement that captures the effect sites have on the engagement on other sites. As an application, we look at the effect of web page layout and structure, which we refer to as web page stylistics, on intersite engagement on Yahoo! properties. Through the analysis of 50 popular Yahoo! sites and a sample of 265,000 users and 19.4M online sessions, we demonstrate that the stylistic components of a web page on a site can be used to predict inter-site engagement across the Yahoo! network of sites. Intersite engagement is a new big data problem as overall it implies analyzing dozen of sites visited by hundreds of millions of people generating billions of sessions. Elad Yom-Tov, Mounia Lalmas-Roelleke, Ricardo Baeza-Yates, Georges Dupret, Janette Lehmann, Pinar Donmez |
IEEE BigData | 1 |
| 2013 | Updating Users about Time Critical Events
Qi Guo 0002, Fernando Diaz 0001, Elad Yom-Tov |
ECIR | 3 |
| 2013 | Workshop on health search and discovery: helping users and advancing medicineabstractThis workshop brings together researchers and practitioners from industry and academia to discuss search and discovery in the medi-cal domain. The event focuses on ways to make medical and health information more accessible to laypeople (including enhancements to ranking algorithms and search interfaces), and how we can dis-cover new medical facts and phenomena from information sought online, as evidenced in query streams and other sources such as social media. This domain also offers many opportunities for appli-cations that monitor and improve quality of life of those affected by medical conditions, by providing tools to support their health-related information behavior. Ryen W. White, Elad Yom-Tov, Eric Horvitz, Eugene Agichtein, William R. Hersh |
SIGIR | 2 |
| 2013 | The Effect of Social and Physical Detachment on Information NeedabstractThe information need of users and the documents which answer this need are frequently contingent on the different characteristics of users. This is especially evident during natural disasters, such as earthquakes and violent weather incidents, which create a strong transient information need. In this article, we investigate how the information need of users, as expressed by their queries, is affected by their physical detachment, as estimated by their physical location in relation to that of the event, and by their social detachment, as quantified by the number of their acquaintances who may be affected by the event. Drawing on large-scale data from ten major events, we show that social and physical detachment levels of users are a major influence on their search engine queries. We demonstrate how knowing social and physical detachment levels can assist in improving retrieval for two applications: identifying search queries related to events and ranking results in response to event-related queries. We find that the average precision in identifying relevant search queries improves by approximately 18%, and that the average precision of ranking that uses detachment information improves by 10%. Using both types of detachment achieved a larger gain in performance than each of them separately. Elad Yom-Tov, Fernando Diaz 0001 |
ACM Trans. Inf. Syst. | 1 |
| 2012 | Models of User Engagement
Janette Lehmann, Mounia Lalmas-Roelleke, Elad Yom-Tov, Georges Dupret |
UMAP | 3 |
| 2012 | On the Relationship between Novelty and Popularity of User-Generated ContentabstractThis work deals with the task of predicting the popularity of user-generated content. We demonstrate how the novelty of newly published content plays an important role in affecting its popularity. More specifically, we study three dimensions of novelty. The first one, termed contemporaneous novelty , models the relative novelty embedded in a new post with respect to contemporary content that was generated by others. The second type of novelty, termed self novelty , models the relative novelty with respect to the user’s own contribution history. The third type of novelty, termed discussion novelty , relates to the novelty of the comments associated by readers with respect to the post content. We demonstrate the contribution of the new novelty measures to estimating blog-post popularity by predicting the number of comments expected for a fresh post. We further demonstrate how novelty based measures can be utilized for predicting the citation volume of academic papers. David Carmel, Haggai Roitman, Elad Yom-Tov |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2011 | Active Online Classification via Information MaximizationabstractWe propose an online classification approach for co-occurrence data which is based on a simple information theoretic principle. We further show how to properly estimate the uncertainty associated with each prediction of our scheme and demonstrate how to exploit these uncertainty estimates. First, in order to abstain highly uncertain predictions. And second, within an active learning framework, in order to preserve classification accuracy while substantially reducing training set size. Our method is highly efficient in terms of run-time and memory footprint requirements. Experimental results in the domain of text classification demonstrate that the classification accuracy of our method is superior or comparable to other state-of-the-art online classification algorithms. Noam Slonim, Elad Yom-Tov, Koby Crammer |
IJCAI | 2 |
| 2011 | Out of sight, not out of mind: on the effect of social and physical detachment on information needabstractThe information needs of users and the documents which answer it are frequently contingent on the different characteristics of users. This is especially evident during natural disasters, such as earthquakes and violent weather incidents, which create a strong transient information need. In this paper we investigate how the information need of users is affected by their physical detachment, as estimated by their physical location in relation to that of the event, and by their social detachment, as quantified by the number of their acquaintances who may be affected by the event. Drawing on large-scale data from three major events, we show that social and physical detachment levels of users are a major influence on their information needs, as manifested by their search engine queries. We demonstrate how knowing social and physical detachment levels can assist in improving retrieval for two applications: identifying search queries related to events and ranking results in response to event-related queries. We find that the average precision in identifying relevant search queries improves by approximately 18%, and that the average precision of ranking that uses detachment information improves by 10%. Elad Yom-Tov, Fernando Diaz 0001 |
SIGIR | 1 |
| 2011 | Location and timeliness of information sources during news eventsabstractPeople nowadays can obtain information on current news events through media outlets, social media, and by actively seeking information using search engines. In this paper we investigate the temporal relationship between news coverage by media outlets, social media, and query logs and show that social media frequently precedes other information sources. Additionally, we demonstrate that there is strong negative correlation between the probability for reporting of an event and the distance of the information source from the event. Elad Yom-Tov, Fernando Diaz 0001 |
SIGIR | 1 |
| 2010 | On the relationship between novelty and popularity of user-generated contentabstractThis work deals with the task of predicting the popularity of user-generated content. We demonstrate how the novelty of newly published content plays an important role in affecting its popularity. We study three dimensions of novelty: contemporaneous novelty, self novelty, and discussion novelty. We demonstrate the contribution of the new novelty measures to estimating blog-post popularity by predicting the number of comments expected for a fresh post. We further demonstrate how novelty based measures can be utilized for predicting the citation volume of academic papers. David Carmel, Haggai Roitman, Elad Yom-Tov |
CIKM | 3 |
| 2010 | Predicting Customer Churn in Mobile Networks through Analysis of Social GroupsabstractChurn prediction aims to identify subscribers who are about to transfer their business to a competitor. Since the cost associated with customer acquisition is much greater than the cost of customer retention, churn prediction has emerged as a crucial Business Intelligence (BI) application for modern telecommunication operators. The dominant approach to churn prediction is to model individual customers and derive their likelihood of churn using a predictive model. Recent work has shown that analyzing customers' interactions by assessing the social vicinity of recent churners can improve the accuracy of churn prediction. We propose a novel framework, termed Group-First Churn Prediction, which eliminates the a priori requirement of knowing who recently churned. Specifically, our approach exploits the structure of customer interactions to predict which groups of subscribers are most prone to churn, before even a single member in the group has churned. Our method works by identifying closely-knit groups of subscribers using second order social metrics derived from information theoretic principles. The interactions within each group are then analyzed to identify social leaders. Based on Key Performance Indicators that are derived from these groups, a novel statistical model is used to predict the churn of the groups and their members. Our experimental results, which are based on data from a telecommunication operator with approximately 16 million subscribers, demonstrate the unique advantages of the proposed method. We further provide empirical evidence that our method captures social phenomena in a highly significant manner. Yossi Richter, Elad Yom-Tov, Noam Slonim |
SDM | 2 |
| 2010 | Estimating the query difficulty for information retrievalabstractMany information retrieval (IR) systems suffer from a radical variance in performance when responding to users' queries. Even for systems that succeed very well on average, the quality of results returned for some of the queries is poor. Thus, it is desirable that IR systems will be able to identify "difficult" queries in order to handle them properly. Understanding why some queries are inherently more difficult than others is essential for IR, and a good answer to this important question will help search engines to reduce the variance in performance, hence better servicing their customer needs. David Carmel, Elad Yom-Tov |
SIGIR | 2 |
| 2010 | Social bookmark weighting for search and recommendation
David Carmel, Haggai Roitman, Elad Yom-Tov |
VLDB J. | 3 |
| 2009 | Who tags the tags?: a framework for bookmark weightingabstractIn this work we propose a novel framework for bookmark weighting which allows us to estimate the effectiveness of each of the bookmarks individually. We show that by weighting bookmarks according to their estimated quality we can significantly improve search effectiveness. Using empirical evaluation on real data gathered from two large bookmarking systems, we demonstrate the effectiveness of the new framework for search enhancement. David Carmel, Haggai Roitman, Elad Yom-Tov |
CIKM | 3 |
| 2009 | Parallel Pairwise ClusteringabstractGiven the pairwise affinity relations associated with a set of data items, the goal of a clustering algorithm is to automatically partition the data into a small number of homogeneous clusters. However, since the input size is quadratic in the number of data points, existing algorithms are non feasible for many practical applications. Here, we propose a simple strategy to cluster massive data by randomly splitting the original affinity matrix into small manageable affinity matrices that are clustered independently. Our proposal is most appealing in a parallel computing environment where at each iteration, each worker node clusters a subset of the input data and the results from all workers are then integrated in a master node to create a new clustering partition over the entire data. We demonstrate that this approach yields high quality clustering partitions for various real world problems, even though at each iteration only small fractions of the original data matrix are examined and at no point is the entire affinity matrix stored in memory or even computed. Furthermore, we demonstrate that the proposed algorithm has intriguing stochastic convergence properties that provide further insight into the clustering problem. Elad Yom-Tov, Noam Slonim |
SDM | 1 |
| 2008 | Automatic Debugging of Concurrent Programs through Active Sampling of Low Dimensional Random ProjectionsabstractConcurrent computer programs are fast becoming prevalent in many critical applications. Unfortunately, these programs are especially difficult to test and debug. Recently, it has been suggested that injecting random timing noise into many points within a program can assist in eliciting bugs within the program. Upon eliciting the bug, it is necessary to identify a minimal set of points that indicate the source of the bug to the programmer. In this paper, we pose this problem as an active feature selection problem. We propose an algorithm called the iterative group sampling algorithm that iteratively samples a lower dimensional projection of the program space and identifies candidate relevant points. We analyze the convergence properties of this algorithm. We test the proposed algorithm on several real-world programs and show its superior performance. Finally, we show the algorithms' performance on a large concurrent program. Elad Yom-Tov, Rachel Tzoref, Shmuel Ur, Shlomo Hoory |
ASE | 1 |
| 2008 | Better multiclass classification via a margin-optimized single binary problem
Ran El-Yaniv, Dmitry Pechyony, Elad Yom-Tov |
Pattern Recognit. Lett. | 3 |
| 2008 | Maintaining dynamic channel profiles on the webabstractThis work addresses a novel problem of maintaining channel proflies on the Web. Such channel maintenance is essential for next generation of Web 2.0 applications that provide sophisticated search and discovery services over Web information channels. Maintaining a fresh channel profile is extremely difficult due to the the dynamic nature of the channel, especially under the constraint of a limited monitoring budget. We propose a novel monitoring scheme that learns the channels' monitoring rates. The monitoring scheme is further extended to consider the content that is published on the channels. We describe a novelty detection filter that refines the monitoring rate according to the expected rate of novel content published on the channels. We further show how inter-channel profile similarities can be utilized to refine the channel monitoring rates. Using real-world data of Web feeds we study the performance of the monitoring scheme. We experiment with several monitoring policies over a large set of Web feeds and show that a policy based on learning the monitoring rate of the channels, combined with novelty detection, outperforms alternative channel monitoring policies. Our results show that the suggested content-based policy is able to maintain high quality channel profiles under limited monitoring resources. Haggai Roitman, David Carmel, Elad Yom-Tov |
Proc. VLDB Endow. | 3 |
| 2008 | A probabilistic alternative to regression suites
Shady Copty, Shai Fine, Shmuel Ur, Elad Yom-Tov, Avi Ziv |
Theor. Comput. Sci. | 4 |
| 2007 | Instrumenting where it hurts: an automatic concurrent debugging techniqueabstractAs concurrent and distributive applications are becoming more common and debugging such applications is very difficult, practical tools for automatic debugging of concurrent applications are in demand. In previous work, we applied automatic debugging to noise-based testing of concurrent programs. The idea of noise-based testing is to increase the probability of observing the bugs by adding, using instrumentation, timing to the execution of the program. The technique of finding a small subset of points that causes the bug to manifest can be used as an automatic debugging technique. Previously, we showed that Delta Debugging can be used to pinpoint the bug location on some small programs.In the work reported in this paper, we create and evaluate two algorithms for automatically pinpointing program locations that are in the vicinity of the bugs on a number of industrial programs. We discovered that the Delta Debugging algorithms do not scale due to the non-monotonic nature of the concurrent debugging problem. Instead we decided to try a machine learning feature selection algorithm. The idea is to consider each instrumentation point as a feature, execute the program many times with different instrumentations, and correlate the features (instrumentation points) with the executions in which the bug was revealed. This idea works very well when the bug is very hard to reveal using instrumentation, correlating to the case when a very specific timing window is needed to reveal the bug. However, in the more common case, when the bugs are easy to find using instrumentation points ranked high by the feature selection algorithm is not high enough. We show that for these cases, the important value is not the absolute value of the evaluation of the feature but the derivative of that value along the program execution path.As a number of groups expressed interest in this research, we built an open infrastructure for automatic debugging algorithms for concurrent applications, based on noise injection based concurrent testing using instrumentation. The infrastructure is described in this paper. Rachel Tzoref, Shmuel Ur, Elad Yom-Tov |
ISSTA | 3 |
| 2007 | A Self-optimized Job Scheduler for Heterogeneous Server Clusters
Elad Yom-Tov, Yariv Aridor |
JSSPP | 1 |
| 2006 | Improving Resource Matching Through Estimation of Actual Job RequirementsabstractHeterogeneous clusters and grid infrastructures are becoming increasingly popular. In these computing infrastructures, machines have different resources (e.g., memory sizes, disk space, and installed software packages). These differences give rise to a problem of over-provisioning, that is, sub-optimal utilization of a cluster due to users requesting resource capacities greater than what their jobs actually need. Our analysis of a real workload file (LANL CM 5) revealed differences of up to two orders of magnitude between requested memory capacity and actual memory usage. The problem of over-provisioning has received very little attention so far. We discuss different approaches for applying machine learning methods to estimate the actual resource capacities used by jobs. These approaches are independent of the scheduling policies and the dynamic resource-matching schemes used. Our simulations show that these methods can yield an improvement of over 50% in utilization (throughput) of heterogeneous clusters Elad Yom-Tov, Yariv Aridor |
HPDC | 1 |
| 2006 | What makes a query difficult?abstractThis work tries to answer the question of what makes a query difficult. It addresses a novel model that captures the main components of a topic and the relationship between those components and topic difficulty. The three components of a topic are the textual expression describing the information need (the query or queries), the set of documents relevant to the topic (the Qrels), and the entire collection of documents. We show experimentally that topic difficulty strongly depends on the distances between these components. In the absence of knowledge about one of the model components, the model is still useful by approximating the missing component based on the other components. We demonstrate the applicability of the difficulty model for several uses such as predicting query difficulty, predicting the number of topic aspects expected to be covered by the search results, and analyzing the findability of a specific domain. David Carmel, Elad Yom-Tov, Adam Darlow, Dan Pelleg |
SIGIR | 2 |
| 2005 | Learning to estimate query difficulty: including applications to missing content detection and distributed information retrievalabstractIn this article we present novel learning methods for estimating the quality of results returned by a search engine in response to a query. Estimation is based on the agreement between the top results of the full query and the top results of its sub-queries. We demonstrate the usefulness of quality estimation for several applications, among them improvement of retrieval, detecting queries for which no relevant content exists in the document collection, and distributed information retrieval. Experiments on TREC data demonstrate the robustness and the effectiveness of our learning algorithms. Elad Yom-Tov, Shai Fine, David Carmel, Adam Darlow |
SIGIR | 1 |