VLDB 2026 Research / reviewers in the wild / expert
Evgeniy Gabrilovich
dblp:84/1449
· DBLP profile ↗
61ranked-venue papers
13as first author
1since 2021 · last 2022
0000-0001-7933-1926ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 52 · 6 first-author · 1 since 2021Artificial intelligence and machine learning · 30 · 8 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
40 papers |
Information retrieval · 52% Knowledge graphs · 28% Data mining · 8% | |
| Artificial intelligence
11 papers |
Information extraction and text analysis · 39% Knowledge representation and reasoning · 23% Question answering and dialogue systems · 18% | |
| Human-computer interaction and pervasive computing
1 paper |
Collaborative and social computing · 100% |
Topics — the 30 heaviest of 83, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Knowledge graphs
knowledge graph construction |
1.3 | 7 | 2016 | A Review of Relational Machine Learning for Knowledge Graphs · Proc. IEEE 2016 Constructing and Mining Web-scale Knowledge Graphs · SIGIR 2016 From Data Fusion to Knowledge Fusion · Proc. VLDB Endow. 2014 |
Information retrieval › online advertising
sponsored search |
0.6 | 7 | 2011 | Bid generation for advanced match in sponsored search · WSDM 2011 Using landing pages for sponsored search ad selection · WWW 2010 The anatomy of an ad: structured indexing and retrieval for sponsored search · WWW 2010 |
Information retrieval
retrieval models |
0.5 | 4 | 2012 | Joint relevance and freshness learning from clickthroughs for news search · WWW 2012 To each his own: personalized content selection based on text comprehensibility · WSDM 2012 Concept-Based Information Retrieval Using Explicit Semantic Analysis · ACM Trans. Inf. Syst. 2011 |
Knowledge graphs
knowledge graph mining |
0.4 | 2 | 2016 | Constructing and Mining Web-scale Knowledge Graphs · SIGIR 2016 Constructing and mining web-scale knowledge graphs: KDD 2014 tutorial · KDD 2014 |
Knowledge graphs
link prediction |
0.4 | 2 | 2016 | A Review of Relational Machine Learning for Knowledge Graphs · Proc. IEEE 2016 Knowledge base completion via search-based question answering · WWW 2014 |
Information retrieval
online advertising |
0.3 | 4 | 2010 | Automatic generation of bid phrases for online advertising · WSDM 2010 Information retrieval challenges in computational advertising · SIGIR 2010 Towards intent-driven bidterm suggestion · WWW 2009 |
Information retrieval
web search |
0.3 | 3 | 2016 | Generalized link suggestions via web site clustering · WWW 2011 Predicting web searcher satisfaction with existing community-based answers · SIGIR 2011 Constructing and Mining Web-scale Knowledge Graphs · SIGIR 2016 |
Data mining › text mining
information extraction |
0.3 | 2 | 2016 | A Review of Relational Machine Learning for Knowledge Graphs · Proc. IEEE 2016 Information organization and retrieval with collaboratively generated content · SIGIR 2011 |
Information retrieval
ranking |
0.3 | 2 | 2012 | Joint relevance and freshness learning from clickthroughs for news search · WWW 2012 To each his own: personalized content selection based on text comprehensibility · WSDM 2012 |
Natural language and speech › Information extraction and text analysis › lexical semantics
semantic relatedness |
0.3 | 2 | 2012 | Large-scale learning of word relatedness with constraints · KDD 2012 A word at a time: computing word relatedness using temporal semantic analysis · WWW 2011 |
Information retrieval › document retrieval
concept-based retrieval |
0.2 | 2 | 2011 | Concept-Based Information Retrieval Using Explicit Semantic Analysis · ACM Trans. Inf. Syst. 2011 Information organization and retrieval with collaboratively generated content · SIGIR 2011 |
Knowledge graphs
knowledge graph embedding |
0.2 | 1 | 2016 | A Review of Relational Machine Learning for Knowledge Graphs · Proc. IEEE 2016 |
Data mining › relational learning
statistical relational learning |
0.2 | 1 | 2016 | A Review of Relational Machine Learning for Knowledge Graphs · Proc. IEEE 2016 |
Information retrieval › search interfaces
search result presentation |
0.2 | 2 | 2011 | Generalized link suggestions via web site clustering · WWW 2011 Competing for users' attention: on the interplay between organic and sponsored search results · WWW 2010 |
Knowledge graphs › knowledge graph construction
knowledge extraction |
0.2 | 2 | 2014 | Knowledge vault: a web-scale approach to probabilistic knowledge fusion · KDD 2014 Mining, searching and exploiting collaboratively generated content on the web · WSDM 2012 |
Information retrieval › online advertising › sponsored search
ad retrieval |
0.2 | 2 | 2010 | The anatomy of an ad: structured indexing and retrieval for sponsored search · WWW 2010 Information retrieval challenges in computational advertising · SIGIR 2010 |
Data integration and cleaning
truth discovery |
0.2 | 1 | 2015 | Knowledge-Based Trust: Estimating the Trustworthiness of Web Sources · Proc. VLDB Endow. 2015 |
Collaborative and social computing
crowdsourcing |
0.2 | 1 | 2015 | Getting More for Less: Optimized Crowdsourcing with Dynamic Tasks and Goals · WWW 2015 |
Collaborative and social computing › crowdsourcing
worker retention |
0.2 | 1 | 2015 | Getting More for Less: Optimized Crowdsourcing with Dynamic Tasks and Goals · WWW 2015 |
Information retrieval › document processing › document analysis › document representation
explicit semantic analysis |
0.2 | 2 | 2011 | Concept-Based Information Retrieval Using Explicit Semantic Analysis · ACM Trans. Inf. Syst. 2011 Computing Semantic Relatedness Using Wikipedia-based Explicit Semantic Analysis · IJCAI 2007 |
Information retrieval › query understanding
query classification |
0.2 | 3 | 2009 | Cross-language query classification using web search for exogenous knowledge · WSDM 2009 Robust classification of rare queries using web knowledge · SIGIR 2007 Context transfer in search advertising · SIGIR 2009 |
Information retrieval › online advertising
ad targeting |
0.2 | 1 | 2014 | Quizz: targeted crowdsourcing with a billion (potential) users · WWW 2014 |
Data integration and cleaning
data fusion |
0.2 | 1 | 2014 | From Data Fusion to Knowledge Fusion · Proc. VLDB Endow. 2014 |
Knowledge graphs › knowledge graph management
knowledge base curation |
0.2 | 1 | 2014 | Trust, but verify: predicting contribution quality for knowledge base construction and curation · WSDM 2014 |
Knowledge graphs
knowledge base integration |
0.2 | 1 | 2014 | From Data Fusion to Knowledge Fusion · Proc. VLDB Endow. 2014 |
Information retrieval
evaluation |
0.2 | 3 | 2014 | Measuring the reusability of test collections · WSDM 2010 ERD'14: entity recognition and disambiguation challenge · SIGIR 2014 Parameterized generation of labeled datasets for text categorization based on a hierarchical directory · SIGIR 2004 |
Natural language and speech › Information extraction and text analysis
text classification |
0.2 | 3 | 2007 | Harnessing the Expertise of 70, 000 Human Editors: Knowledge-Based Feature Generation for Text Categorization · J. Mach. Learn. Res. 2007 Feature Generation for Text Categorization Using World Knowledge · IJCAI 2005 Text categorization with many redundant features: using aggressive feature selection to make SVMs competitive with C4.5 · ICML 2004 |
Data mining › text mining
text classification |
0.2 | 4 | 2010 | Overcoming the Brittleness Bottleneck using Wikipedia: Enhancing Text Categorization with Encyclopedic Knowledge · AAAI 2006 Parameterized generation of labeled datasets for text categorization based on a hierarchical directory · SIGIR 2004 Information retrieval challenges in computational advertising · SIGIR 2010 |
Natural language and speech › Question answering and dialogue systems
community question answering |
0.2 | 2 | 2012 | Predicting web searcher satisfaction with existing community-based answers · SIGIR 2011 To each his own: personalized content selection based on text comprehensibility · WSDM 2012 |
Information retrieval
indexing |
0.1 | 2 | 2011 | The anatomy of an ad: structured indexing and retrieval for sponsored search · WWW 2010 Concept-Based Information Retrieval Using Explicit Semantic Analysis · ACM Trans. Inf. Syst. 2011 |
Methods — techniques the papers use, named apart from their topics
multi-layer probabilistic model · 0.4joint inference · 0.4explicit semantic analysis · 0.4tensor factorization · 0.2rule mining · 0.2multiway neural networks · 0.2latent feature model · 0.2optimization · 0.2supervised machine learning · 0.2query learning · 0.2probabilistic inference · 0.2probabilistic aggregation · 0.2information theory · 0.2feature engineering · 0.2data fusion techniques · 0.2ad targeting · 0.2user profiling · 0.1low-dimensional embedding · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Implicit User-Generated Content in the Service of Public HealthabstractEvery day millions of people use online products and services to satisfy their information needs. While doing so, they produce large volumes of user-generated content (UGC). In this talk, we will distinguish between "explicit" UGC, which is intended to be made public (such as product ratings or reviews), and "implicit" UGC, which can be responsibly anonymized and aggregated in a privacy-preserving way to improve public health. We will analyze implicit UGC as a positive consumption externality, and will discuss its beneficial uses across a range of public health applications. Evgeniy Gabrilovich |
CIKM | 1 |
| 2016 | Constructing and Mining Web-scale Knowledge GraphsabstractRecent years have witnessed a proliferation of large-scale knowledge graphs, from purely academic projects such as YAGO to major commercial projects such as Google's Knowledge Graph and Microsoft's Satori. Whereas there is a large body of research on mining homogeneous graphs, this new generation of information networks are highly heterogeneous, with thousands of entity and relation types and billions of instances of those types (graph vertices and edges). In this tutorial, we present the state of the art in constructing, mining, and growing knowledge graphs. The purpose of the tutorial is to equip newcomers to this exciting field with an understanding of the basic concepts, tools and methodologies, open research challenges, as well as pointers to available datasets and relevant literature. Knowledge graphs have become an enabling resource for a plethora of new knowledge-rich applications. Consequently, the tutorial will also discuss the role of knowledge bases in empowering a range of web applications, from web search to social networks to digital assistants. A publicly available knowledge base (Freebase) will be used throughout the tutorial to exemplify the different techniques. Evgeniy Gabrilovich, Nicolas Usunier |
SIGIR | 1 |
| 2016 | A Review of Relational Machine Learning for Knowledge GraphsabstractRelational machine learning studies methods for the statistical analysis of relational, or graph-structured, data. In this paper, we provide a review of how such statistical models can be “trained” on large knowledge graphs, and then used to predict new facts about the world (which is equivalent to predicting new edges in the graph). In particular, we discuss two fundamentally different kinds of statistical relational models, both of which can scale to massive data sets. The first is based on latent feature models such as tensor factorization and multiway neural networks. The second is based on mining observable patterns in the graph. We also show how to combine these latent and observable models to get improved modeling power at decreased computational cost. Finally, we discuss how such statistical models of graphs can be combined with text-based information extraction methods for automatically constructing knowledge graphs from the Web. To this end, we also discuss Google's knowledge vault project as an example of such combination. Maximilian Nickel, Kevin Murphy 0002, Volker Tresp, Evgeniy Gabrilovich |
Proc. IEEE | 4 |
| 2015 | Getting More for Less: Optimized Crowdsourcing with Dynamic Tasks and GoalsabstractIn crowdsourcing systems, the interests of contributing participants and system stakeholders are often not fully aligned. Participants seek to learn, be entertained, and perform easy tasks, which offer them instant gratification; system stakeholders want users to complete more difficult tasks, which bring higher value to the crowdsourced application. We directly address this problem by presenting techniques that optimize the crowdsourcing process by jointly maximizing the user longevity in the system and the true value that the system derives from user participation. Ari Kobren, Chun How Tan, Panagiotis G. Ipeirotis, Evgeniy Gabrilovich |
WWW | 4 |
| 2015 | Knowledge-Based Trust: Estimating the Trustworthiness of Web SourcesabstractThe quality of web sources has been traditionally evaluated using exogenous signals such as the hyperlink structure of the graph. We propose a new approach that relies on endogenous signals, namely, the correctness of factual information provided by the source. A source that has few false facts is considered to be trustworthy. The facts are automatically extracted from each source by information extraction methods commonly used to construct knowledge bases. We propose a way to distinguish errors made in the extraction process from factual errors in the web source per se, by using joint inference in a novel multi-layer probabilistic model. We call the trustworthiness score we computed Knowledge-Based Trust (KBT) . On synthetic data, we show that our method can reliably compute the true trustworthiness levels of the sources. We then apply it to a database of 2.8B facts extracted from the web, and thereby estimate the trustworthiness of 119M webpages. Manual evaluation of a subset of the results confirms the effectiveness of the method. Xin Dong 0001, Evgeniy Gabrilovich, Kevin Murphy 0002, Van Dang, Wilko Horn, Camillo Lugaresi, Shaohua Sun, Wei Zhang 0152 |
Proc. VLDB Endow. | 2 |
| 2014 | Knowledge vault: a web-scale approach to probabilistic knowledge fusionabstractRecent years have witnessed a proliferation of large-scale knowledge bases, including Wikipedia, Freebase, YAGO, Microsoft's Satori, and Google's Knowledge Graph. To increase the scale even further, we need to explore automatic methods for constructing knowledge bases. Previous approaches have primarily focused on text-based extraction, which can be very noisy. Here we introduce Knowledge Vault, a Web-scale probabilistic knowledge base that combines extractions from Web content (obtained via analysis of text, tabular data, page structure, and human annotations) with prior knowledge derived from existing knowledge repositories. We employ supervised machine learning methods for fusing these distinct information sources. The Knowledge Vault is substantially bigger than any previously published structured knowledge repository, and features a probabilistic inference system that computes calibrated probabilities of fact correctness. We report the results of multiple studies that explore the relative utility of the different information sources and extraction methods. Xin Dong 0001, Evgeniy Gabrilovich, Geremy Heitz, Wilko Horn, Ni Lao, Kevin Murphy 0002, Thomas Strohmann, Shaohua Sun, Wei Zhang 0152 |
KDD | 2 |
| 2014 | Constructing and mining web-scale knowledge graphs: KDD 2014 tutorialabstractRecent years have witnessed a proliferation of large-scale knowledge graphs, such as Freebase, YAGO, Google's Knowledge Graph, and Microsoft's Satori. Whereas there is a large body of research on mining homogeneous graphs, this new generation of information networks are highly heterogeneous, with thousands of entity and relation types and billions of instances of vertices and edges. In this tutorial, we will present the state of the art in constructing, mining, and growing knowledge graphs. The purpose of the tutorial is to equip newcomers to this exciting field with an understanding of the basic concepts, tools and methodologies, available datasets, and open research challenges. A publicly available knowledge base (Freebase) will be used throughout the tutorial to exemplify the different techniques. Antoine Bordes, Evgeniy Gabrilovich |
KDD | 2 |
| 2014 | ERD'14: entity recognition and disambiguation challengeabstractNo abstract available. David Carmel, Ming-Wei Chang, Evgeniy Gabrilovich, Bo-June Paul Hsu, Kuansan Wang |
SIGIR | 3 |
| 2014 | Trust, but verify: predicting contribution quality for knowledge base construction and curationabstractThe largest publicly available knowledge repositories, such as Wikipedia and Freebase, owe their existence and growth to volunteer contributors around the globe. While the majority of contributions are correct, errors can still creep in, due to editors' carelessness, misunderstanding of the schema, malice, or even lack of accepted ground truth. If left undetected, inaccuracies often degrade the experience of users and the performance of applications that rely on these knowledge repositories. We present a new method, CQUAL, for automatically predicting the quality of contributions submitted to a knowledge base. Significantly expanding upon previous work, our method holistically exploits a variety of signals, including the user's domains of expertise as reflected in her prior contribution history, and the historical accuracy rates of different types of facts. In a large-scale human evaluation, our method exhibits precision of 91% at 80% recall. Our model verifies whether a contribution is correct immediately after it is submitted, significantly alleviating the need for post-submission human reviewing. Chun How Tan, Eugene Agichtein, Panagiotis G. Ipeirotis, Evgeniy Gabrilovich |
WSDM | 4 |
| 2014 | Quizz: targeted crowdsourcing with a billion (potential) usersabstractWe describe Quizz, a gamified crowdsourcing system that simultaneously assesses the knowledge of users and acquires new knowledge from them. Quizz operates by asking users to complete short quizzes on specific topics; as a user answers the quiz questions, Quizz estimates the user's competence. To acquire new knowledge, Quizz also incorporates questions for which we do not have a known answer; the answers given by competent users provide useful signals for selecting the correct answers for these questions. Quizz actively tries to identify knowledgeable users on the Internet by running advertising campaigns, effectively leveraging the targeting capabilities of existing, publicly available, ad placement services. Quizz quantifies the contributions of the users using information theory and sends feedback to the advertisingsystem about each user. The feedback allows the ad targeting mechanism to further optimize ad placement. Panagiotis G. Ipeirotis, Evgeniy Gabrilovich |
WWW | 2 |
| 2014 | Knowledge base completion via search-based question answeringabstractOver the past few years, massive amounts of world knowledge have been accumulated in publicly available knowledge bases, such as Freebase, NELL, and YAGO. Yet despite their seemingly huge size, these knowledge bases are greatly incomplete. For example, over 70% of people included in Freebase have no known place of birth, and 99% have no known ethnicity. In this paper, we propose a way to leverage existing Web-search-based question-answering technology to fill in the gaps in knowledge bases in a targeted way. In particular, for each entity attribute, we learn the best set of queries to ask, such that the answer snippets returned by the search engine are most likely to contain the correct value for that attribute. For example, if we want to find Frank Zappa's mother, we could ask the query `who is the mother of Frank Zappa'. However, this is likely to return `The Mothers of Invention', which was the name of his band. Our system learns that it should (in this case) add disambiguating terms, such as Zappa's place of birth, in order to make it more likely that the search results contain snippets mentioning his mother. Our system also learns how many different queries to ask for each attribute, since in some cases, asking too many can hurt accuracy (by introducing false positives). We discuss how to aggregate candidate answers across multiple queries, ultimately returning probabilistic predictions for possible values for each attribute. Finally, we evaluate our system and show that it is able to extract a large number of facts with high confidence. Robert West 0001, Evgeniy Gabrilovich, Kevin Murphy 0002, Shaohua Sun, Dekang Lin |
WWW | 2 |
| 2014 | From Data Fusion to Knowledge FusionabstractThe task of data fusion is to identify the true values of data items ( e.g. , the true date of birth for Tom Cruise ) among multiple observed values drawn from different sources ( e.g. , Web sites) of varying (and unknown) reliability. A recent survey [20] has provided a detailed comparison of various fusion methods on Deep Web data. In this paper, we study the applicability and limitations of different fusion techniques on a more challenging problem: knowledge fusion . Knowledge fusion identifies true subject-predicate-object triples extracted by multiple information extractors from multiple information sources. These extractors perform the tasks of entity linkage and schema alignment, thus introducing an additional source of noise that is quite different from that traditionally considered in the data fusion literature, which only focuses on factual errors in the original sources. We adapt state-of-the-art data fusion techniques and apply them to a knowledge base with 1.6B unique knowledge triples extracted by 12 extractors from over 1B Web pages, which is three orders of magnitude larger than the data sets used in previous data fusion papers. We show great promise of the data fusion approaches in solving the knowledge fusion problem, and suggest interesting research directions through a detailed error analysis of the methods. Xin Dong 0001, Evgeniy Gabrilovich, Geremy Heitz, Wilko Horn, Kevin Murphy 0002, Shaohua Sun, Wei Zhang 0152 |
Proc. VLDB Endow. | 2 |
| 2013 | Sixth workshop on exploiting semantic annotations in information retrieval (ESAIR'13)abstractThere is an increasing amount of structure on the web as a result of modern web languages, user tagging and annotation, emerging robust NLP tools, and an ever growing volume of linked data. These meaningful, semantic, annotations hold the promise to significantly enhance information access, by enhancing the depth of analysis of today's systems. Currently, we have only started exploring the possibilities and only begin to understand how these valuable semantic cues can be put to fruitful use. ESAIR'13 focuses on two of the most challenging aspects to address in the coming years. First, there is a need to include the currently emerging knowledge resources (such as DBpedia, Freebase) as underlying semantic model giving access to an unprecedented scope and detail of factual information. Second, there is a need to include annotations beyond the topical dimension (think of sentiment, reading level, prerequisite level, etc) that contain vital cues for matching the specific needs and profile of the searcher at hand. Paul N. Bennett, Evgeniy Gabrilovich, Jaap Kamps, Jussi Karlgren |
CIKM | 2 |
| 2012 | Large-scale learning of word relatedness with constraintsabstractPrior work on computing semantic relatedness of words focused on representing their meaning in isolation, effectively disregarding inter-word affinities. We propose a large-scale data mining approach to learning word-word relatedness, where known pairs of related words impose constraints on the learning process. We learn for each word a low-dimensional representation, which strives to maximize the likelihood of a word given the contexts in which it appears. Our method, called CLEAR, is shown to significantly outperform previously published approaches. The proposed method is based on first principles, and is generic enough to exploit diverse types of text corpora, while having the flexibility to impose constraints on the derived word similarities. We also make publicly available a new labeled dataset for evaluating word relatedness algorithms, which we believe to be the largest such dataset to date. Guy Halawi, Gideon Dror, Evgeniy Gabrilovich, Yehuda Koren |
KDD | 3 |
| 2012 | Mining, searching and exploiting collaboratively generated content on the webabstractProliferation of ubiquitous access to the Internet enables millions of Web users to collaborate online on a variety of activities. Many of these activities result in the construction of large repositories of knowledge, either as their primary aim (e.g., Wikipedia) or as a by-product (e.g., Yahoo! Answers). In this tutorial, we will discuss organizing and exploiting Collaboratively Generated Content (CGC) for information organization and retrieval. Specifically, we intend to cover two complementary areas of the problem: (1) using such content as a powerful enabling resource for knowledge-enriched, intelligent representations and new information retrieval algorithms, and (2) development of supporting technologies for extracting, filtering, and organizing collaboratively created content. Eugene Agichtein, Evgeniy Gabrilovich |
WSDM | 2 |
| 2012 | To each his own: personalized content selection based on text comprehensibilityabstractImagine a physician and a patient doing a search on antibiotic resistance. Or a chess amateur and a grandmaster conducting a search on Alekhine's Defence. Although the topic is the same, arguably the two users in each case will satisfy their information needs with very different texts. Yet today search engines mostly adopt the one-size-fits-all solution, where personalization is restricted to topical preference. We found that users do not uniformly prefer simple texts, and that the text comprehensibility level should match the user's level of preparedness. Consequently, we propose to model the comprehensibility of texts as well as the users' reading proficiency in order to better explain how different users choose content for further exploration. We also model topic-specific reading proficiency, which allows us to better explain why a physician might choose to read sophisticated medical articles yet simple descriptions of SLR cameras. We explore different ways to build user profiles, and use collaborative filtering techniques to overcome data sparsity. We conducted experiments on large-scale datasets from a major Web search engine and a community question answering forum. Our findings confirm that explicitly modeling text comprehensibility can significantly improve content ranking (search results or answers, respectively). Chenhao Tan, Evgeniy Gabrilovich, Bo Pang 0001 |
WSDM | 2 |
| 2012 | Joint relevance and freshness learning from clickthroughs for news searchabstractIn contrast to traditional Web search, where topical relevance is often the main selection criterion, news search is characterized by the increased importance of freshness. However, the estimation of relevance and freshness, and especially the relative importance of these two aspects, are highly specific to the query and the time when the query was issued. In this work, we propose a unified framework for modeling the topical relevance and freshness, as well as their relative importance, based on click logs. We use click statistics and content analysis techniques to define a set of temporal features, which predict the right mix of freshness and relevance for a given query. Experimental results on both historical click data and editorial judgments demonstrate the effectiveness of the proposed approach. Hongning Wang, Anlei Dong, Lihong Li 0001, Yi Chang 0001, Evgeniy Gabrilovich |
WWW | 5 |
| 2012 | Introduction to the Special Section on Computational Models of Collective Intelligence in the Social WebabstractNo abstract available. Evgeniy Gabrilovich, Zhong Su, Jie Tang 0001 |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2011 | Information retrieval challenges in computational advertisingabstractNo abstract available. Andrei Z. Broder, Evgeniy Gabrilovich, Vanja Josifovski |
CIKM | 2 |
| 2011 | Retrieval models for audience selection in display advertisingabstractWeb applications often rely on user profiles of observed user actions, such as queries issued, page views, etc. In audience selection for display advertising, the audience that is likely to be responsive to a given ad campaign is identified via such profiles. We formalize the audience selection problem as a ranked retrieval task over an index of known users. We focus on the common case of audience selection where a small seed set of users who have previously responded positively to the campaign is used to identify a broader target audience. The actions of the users in the seed set are aggregated to construct a query, the query is then executed against an index of other user profiles to retrieve the highest scoring profiles. We validate our approach on a real-world dataset, demonstrating the trade-offs of different user and query models and that our approach is particularly robust for small campaigns. The proposed user modeling framework is applicable to many other applications requiring user profiles such as content suggestion and personalization. Sarah K. Tyler, Sandeep Pandey, Evgeniy Gabrilovich, Vanja Josifovski |
CIKM | 3 |
| 2011 | Ad Retrieval Systems in vitro and in vivo: Knowledge-Based Approaches to Computational Advertising
Evgeniy Gabrilovich |
ECIR | 1 |
| 2011 | Information organization and retrieval with collaboratively generated contentabstractProliferation of ubiquitous access to the Internet enables millions of Web users to collaborate online on a variety of activities. Many of these activities result in the construction of large repositories of knowledge, either as their primary aim (e.g., Wikipedia) or as a by-product (e.g., Yahoo! Answers). In this tutorial, we will discuss organizing and exploiting Collaboratively Generated Content (CGC) for information organization and retrieval. Specifically, we intend to cover two complementary areas of the problem: (1) using such content as a powerful enabling resource for knowledge-enriched, intelligent representations and new information retrieval algorithms, and (2) development of supporting technologies for extracting, filtering, and organizing collaboratively created content. The unprecedented amounts of information in CGC enable new, knowledge-rich approaches to information access, which are significantly more powerful than the conventional word-based methods. Considerable progress has been made in this direction over the last few years. Examples include explicit manipulation of human-defined concepts and their use to augment the bag of words (cf. Explicit Semantic Analysis), using large-scale taxonomies of topics from Wikipedia or the Open Directory Project to construct additional class-based features, or using Wikipedia for better word sense disambiguation. However, the quality and comprehensiveness of collaboratively created content vary widely, and in order for this resource to be useful, a significant amount of preprocessing, filtering, and organization is necessary. Consequently, new methods for analyzing CGC and corresponding user interactions are required to effectively harness the resulting knowledge. Thus, not only the content repositories can be used to improve IR methods, but the reverse pollination is also possible, as better information extraction methods can be used for automatically collecting more knowledge, or verifying the contributed content. This natural connection between modeling the generation process of CGC and effectively using the accumulated knowledge suggests covering both areas together in a single tutorial. The intended audience of the tutorial includes IR researchers and graduate students, who would like to learn about the recent advances and research opportunities in working with collaboratively generated content. The emphasis of the tutorial is on comparing the existing approaches and presenting practical techniques that IR practitioners can use in their research. We also cover open research challenges, as well as survey available resources (software tools and data) for getting started in this research field. Eugene Agichtein, Evgeniy Gabrilovich |
SIGIR | 2 |
| 2011 | Predicting web searcher satisfaction with existing community-based answersabstractCommunity-based Question Answering (CQA) sites, such as Yahoo! Answers, Baidu Knows, Naver, and Quora, have been rapidly growing in popularity. The resulting archives of posted answers to questions, in Yahoo! Answers alone, already exceed in size 1 billion, and are aggressively indexed by web search engines. In fact, a large number of search engine users benefit from these archives, by finding existing answers that address their own queries. This scenario poses new challenges and opportunities for both search engines and CQA sites. To this end, we formulate a new problem of predicting the satisfaction of web searchers with CQA answers. We analyze a large number of web searches that result in a visit to a popular CQA site, and identify unique characteristics of searcher satisfaction in this setting, namely, the effects of query clarity, query-to-question match, and answer quality. We then propose and evaluate several approaches to predicting searcher satisfaction that exploit these characteristics. To the best of our knowledge, this is the first attempt to predict and validate the usefulness of CQA archives for external searchers, rather than for the original askers. Our results suggest promising directions for improving and exploiting community question answering services in pursuit of satisfying even more Web search queries. Qiaoling Liu, Eugene Agichtein, Gideon Dror, Evgeniy Gabrilovich, Yoelle Maarek, Dan Pelleg, Idan Szpektor |
SIGIR | 4 |
| 2011 | Bid generation for advanced match in sponsored searchabstractSponsored search is a three-way interaction between advertisers, users, and the search engine. The basic ad selection in sponsored search, lets the advertiser choose the exact queries where the ad is to be shown. To increase advertising volume, many advertisers opt into advanced match, where the search engine can select additional queries that are deemed relevant for the advertiser's ad. In advanced match, the search engine is effectively bidding on the behalf of the advertisers. While advanced match has been extensively studied in the literature from the ad relevance perspective there is little work that discusses how to infer the appropriate bid value for a given advanced match. The bid value is crucial as it affects both the ad placement in revenue reordering and the amount advertisers are charged in case of a click. Andrei Z. Broder, Evgeniy Gabrilovich, Vanja Josifovski, George Mavromatis, Alexander J. Smola |
WSDM | 2 |
| 2011 | A word at a time: computing word relatedness using temporal semantic analysisabstractComputing the degree of semantic relatedness of words is a key functionality of many language applications such as search, clustering, and disambiguation. Previous approaches to computing semantic relatedness mostly used static language resources, while essentially ignoring their temporal aspects. We believe that a considerable amount of relatedness information can also be found in studying patterns of word usage over time. Consider, for instance, a newspaper archive spanning many years. Two words such as "war" and "peace" might rarely co-occur in the same articles, yet their patterns of use over time might be similar. In this paper, we propose a new semantic relatedness model, Temporal Semantic Analysis (TSA), which captures this temporal information. The previous state of the art method, Explicit Semantic Analysis (ESA), represented word semantics as a vector of concepts. TSA uses a more refined representation, where each concept is no longer scalar, but is instead represented as time series over a corpus of temporally-ordered documents. To the best of our knowledge, this is the first attempt to incorporate temporal evidence into models of semantic relatedness. Empirical evaluation shows that TSA provides consistent improvements over the state of the art ESA results on multiple benchmarks. Kira Radinsky, Eugene Agichtein, Evgeniy Gabrilovich, Shaul Markovitch |
WWW | 3 |
| 2011 | Generalized link suggestions via web site clusteringabstractProactive link suggestion leads to improved user experience by allowing users to reach relevant information with fewer clicks, fewer pages to read, or simply faster because the right pages are prefetched just in time. In this paper we tackle two new scenarios for link suggestion, which were not covered in prior work owing to scarcity of historical browsing data. In the web search scenario, we propose a method for generating quick links - additional entry points into Web sites, which are shown for top search results for navigational queries - for tail sites, for which little browsing statistics is available. Beyond Web search, we also propose a method for link suggestion in general web browsing, effectively anticipating the next link to be followed by the user. Our approach performs clustering of Web sites in order to aggregate information across multiple sites, and enables relevant link suggestion for virtually any site, including tail sites and brand new sites for which little historical data is available. Empirical evaluation confirms the validity of our method using editorially labeled data as well as real-life search and browsing data from a major US search engine. Jangwon Seo, Fernando Diaz 0001, Evgeniy Gabrilovich, Vanja Josifovski, Bo Pang 0001 |
WWW | 3 |
| 2011 | Web Page Summarization for Just-in-Time Contextual AdvertisingabstractContextual advertising is a type of Web advertising, which, given the URL of a Web page, aims to embed into the page the most relevant textual ads available. For static pages that are displayed repeatedly, the matching of ads can be based on prior analysis of their entire content; however, often ads need to be matched to new or dynamically created pages that cannot be processed ahead of time. Analyzing the entire content of such pages on-the-fly entails prohibitive communication and latency costs. To solve the three-horned dilemma of either low relevance or high latency or high load, we propose to use text summarization techniques paired with external knowledge (exogenous to the page) to craft short page summaries in real time. Empirical evaluation proves that matching ads on the basis of such summaries does not sacrifice relevance, and is competitive with matching based on the entire page content. Specifically, we found that analyzing a carefully selected 6% fraction of the page text can sacrifice only 1%--3% in ad relevance. Furthermore, our summaries are fully compatible with the standard JavaScript mechanisms used for ad placement: they can be produced at ad-display time by simple additions to the usual script, and they only add 500--600 bytes to the usual request. We also compared our summarization approach, which is based on structural properties of the HTML content of the page, with a more principled one based on one of the standard text summarization tools (MEAD), and found their performance to be comparable. Aris Anagnostopoulos, Andrei Z. Broder, Evgeniy Gabrilovich, Vanja Josifovski, Lance Riedel |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2011 | Concept-Based Information Retrieval Using Explicit Semantic AnalysisabstractInformation retrieval systems traditionally rely on textual keywords to index and retrieve documents. Keyword-based retrieval may return inaccurate and incomplete results when different keywords are used to describe the same concept in the documents and in the queries. Furthermore, the relationship between these related keywords may be semantic rather than syntactic, and capturing it thus requires access to comprehensive human world knowledge. Concept-based retrieval methods have attempted to tackle these difficulties by using manually built thesauri, by relying on term cooccurrence data, or by extracting latent word relationships and concepts from a corpus. In this article we introduce a new concept-based retrieval approach based on Explicit Semantic Analysis (ESA), a recently proposed method that augments keyword-based text representation with concept-based features, automatically extracted from massive human knowledge repositories such as Wikipedia. Our approach generates new text features automatically, and we have found that high-quality feature selection becomes crucial in this setting to make the retrieval more focused. However, due to the lack of labeled data, traditional feature selection methods cannot be used, hence we propose new methods that use self-generated labeled training data. The resulting system is evaluated on several TREC datasets, showing superior performance over previous state-of-the-art results. Ofer Egozi, Shaul Markovitch, Evgeniy Gabrilovich |
ACM Trans. Inf. Syst. | 3 |
| 2010 | Using the past to score the present: extending term weighting models through revision history analysisabstractThe generative process underlies many information retrieval models, notably statistical language models. Yet these models only examine one (current) version of the document, effectively ignoring the actual document generation process. We posit that a considerable amount of information is encoded in the document authoring process, and this information is complementary to the word occurrence statistics upon which most modern retrieval models are based. We propose a new term weighting model, Revision History Analysis (RHA), which uses the revision history of a document (e.g., the edit history of a page in Wikipedia) to redefine term frequency - a key indicator of document topic/relevance for many retrieval models and text processing tasks. We then apply RHA to document ranking by extending two state-of-the-art text retrieval models, namely, BM25 and the generative statistical language model (LM). To the best of our knowledge, our paper is the first attempt to directly incorporate document authoring history into retrieval models. Empirical results show that RHA provides consistent improvements for state-of-the-art retrieval models, using standard retrieval tasks and benchmarks. Ablimit Aji, Yu Wang 0022, Eugene Agichtein, Evgeniy Gabrilovich |
CIKM | 4 |
| 2010 | Exploiting site-level information to improve web searchabstractRanking Web search results has long evolved beyond simple bag-of-words retrieval models. Modern search engines routinely employ machine learning ranking that relies on exogenous relevance signals. Yet the majority of current methods still evaluate each Web page out of context. In this work, we introduce a novel source of relevance information for Web search by evaluating each page in the context of its host Web site. For this purpose, we devise two strategies for compactly representing entire Web sites. We formalize our approach by building two indices, a traditional page index and a new site index, where each "document" represents the an entire Web site. At runtime, a query is first executed against both indices, and then the final page score for a given query is produced by combining the scores of the page and its site. Experimental results carried out on a large-scale Web search test collection from a major commercial search engine confirm the proposed approach leads to consistent and significant improvements in retrieval effectiveness. Andrei Z. Broder, Evgeniy Gabrilovich, Vanja Josifovski, George Mavromatis, Donald Metzler |
CIKM | 2 |
| 2010 | Information retrieval challenges in computational advertisingabstractComputational advertising is an emerging scientific sub-discipline, at the intersection of large scale search and text analysis, information retrieval, statistical modeling, machine learning, classification, optimization, and microeconomics. The central challenge of computational advertising is to find the best match between a given user in a given context and a suitable advertisement. The aim of this tutorial is to present the state of the art in Computational Advertising, in particular in its IR-related aspects, and to expose the participants to the current research challenges in this field. The tutorial does not assume any prior knowledge of Web advertising, and will begin with a comprehensive background survey. Going deeper, our focus will be on using a textual representation of the user context to retrieve relevant ads. At first approximation, this process can be reduced to a conventional setup by constructing a query that describes the user context and executing the query against a large inverted index of ads. We show how to augment this approach using query expansion and text classification techniques tuned for the ad-retrieval problem. In particular, we show how to use the Web as a repository of query-specific knowledge and use the Web search results retrieved by the query as a form of a relevance feedback and query expansion. We also present solutions that go beyond the conventional bag of words indexing by constructing additional features using a large external taxonomy and a lexicon of named entities obtained by analyzing the entire Web as a corpus. The last part of the tutorial will be devoted to a potpourri of recent research results and open problems inspired by Computational Advertising challenges in text summarization, natural language generation, named entity recognition, computer-human interaction, and other SIGIR-relevant areas. Andrei Z. Broder, Evgeniy Gabrilovich, Vanja Josifovski |
SIGIR | 2 |
| 2010 | Measuring the reusability of test collectionsabstractWhile test collection construction is a time-consuming and expensive process, the true cost is amortized by reusing the collection over hundreds or thousands of experiments. Some of these experiments may involve systems that retrieve documents not judged during the initial construction phase, and some of these systems may be "hard" to evaluate: depending on which judgments are missing and which judged documents were retrieved, the experimenter's confidence in an evaluation could potentially be very low. We propose two methods for quantifying the reusability of a test collection for evaluating new systems. The proposed methods provide simple yet highly effective tests for determining whether an existing set of judgments is useful for evaluating a new system. Empirical evaluations using TREC datasets confirm the usefulness of our proposed reusability measures. In particular, we show that our methods can reliably estimate confidence intervals that are indicative of collection reusability. Ben Carterette, Evgeniy Gabrilovich, Vanja Josifovski, Donald Metzler |
WSDM | 2 |
| 2010 | Anatomy of the long tail: ordinary people with extraordinary tastesabstractThe success of "infinite-inventory" retailers such as Amazon.com and Netflix has been ascribed to a "long tail" phenomenon. To wit, while the majority of their inventory is not in high demand, in aggregate these "worst sellers," unavailable at limited-inventory competitors, generate a significant fraction of total revenue. The long tail phenomenon, however, is in principle consistent with two fundamentally different theories. The first, and more popular hypothesis, is that a majority of consumers consistently follow the crowds and only a minority have any interest in niche content; the second hypothesis is that everyone is a bit eccentric, consuming both popular and specialty products. Based on examining extensive data on user preferences for movies, music, Web search, and Web browsing, we find overwhelming support for the latter theory. However, the observed eccentricity is much less than what is predicted by a fully random model whereby every consumer makes his product choices independently and proportional to product popularity; so consumers do indeed exhibit at least some a priori propensity toward either the popular or the exotic. Sharad Goel, Andrei Z. Broder, Evgeniy Gabrilovich, Bo Pang 0001 |
WSDM | 3 |
| 2010 | Automatic generation of bid phrases for online advertisingabstractOne of the most prevalent online advertising methods is textual advertising. To produce a textual ad, an advertiser must craft a short creative (the text of the ad) linking to a landing page, which describes the product or service being promoted. Furthermore, the advertiser must associate the creative to a set of manually chosen bid phrases representing those Web search queries that should trigger the ad. For efficiency, given a landing page, the bid phrases are often chosen first, and then for each bid phrase the creative is produced using a template. Nevertheless, an ad campaign (e.g., for a large retailer) might involve thousands of landing pages and tens or hundreds of thousands of bid phrases, hence the entire process is very laborious. Sujith Ravi, Andrei Z. Broder, Evgeniy Gabrilovich, Vanja Josifovski, Sandeep Pandey, Bo Pang 0001 |
WSDM | 3 |
| 2010 | The anatomy of an ad: structured indexing and retrieval for sponsored searchabstractThe core task of sponsored search is to retrieve relevant ads for the user's query. Ads can be retrieved either by exact match, when their bid term is identical to the query, or by advanced match, which indexes ads as documents and is similar to standard information retrieval (IR). Recently, there has been a great deal of research into developing advanced match ranking algorithms. However, no previous research has addressed the ad indexing problem. Unlike most traditional search problems, the ad corpus is defined hierarchically in terms of advertiser accounts, campaigns, and ad groups, which further consist of creatives and bid terms. This hierarchical structure makes indexing highly non-trivial, as naively indexing all possible displayable ads leads to a prohibitively large and ineffective index. We show that ad retrieval using such an index is not only slow, but its precision is suboptimal as well. We investigate various strategies for compact, hierarchy-aware indexing of sponsored search ads through adaptation of standard IR indexing techniques. We also propose a new ad retrieval method that yields more relevant ads by exploiting the structured nature of the ad corpus. Experiments carried out over a large ad test collection from a commercial search engine show that our proposed methods are highly effective and efficient compared to more standard indexing and retrieval approaches. Michael Bendersky, Evgeniy Gabrilovich, Vanja Josifovski, Donald Metzler |
WWW | 2 |
| 2010 | Using landing pages for sponsored search ad selectionabstractWe explore the use of the landing page content in sponsored search ad selection. Specifically, we compare the use of the ad's intrinsic content to augmenting the ad with the whole, or parts, of the landing page. We explore two types of extractive summarization techniques to select useful regions from the landing pages: out-of-context and in-context methods. Out-of-context methods select salient regions from the landing page by analyzing the content alone, without taking into account the ad associated with the landing page. In-context methods use the ad context (including its title, creative, and bid phrases) to help identify regions of the landing page that should be used by the ad selection engine. In addition, we introduce a simple yet effective unsupervised algorithm to enrich the ad context to further improve the ad selection. Experimental evaluation confirms that the use of landing pages can significantly improve the quality of ad selection. We also find that our extractive summarization techniques reduce the size of landing pages substantially, while retaining or even improving the performance of ad retrieval over the method that utilize the entire landing page. Yejin Choi 0001, Marcus Fontoura, Evgeniy Gabrilovich, Vanja Josifovski, Maurício R. Mediano, Bo Pang 0001 |
WWW | 3 |
| 2010 | Competing for users' attention: on the interplay between organic and sponsored search resultsabstractQueries on major Web search engines produce complex result pages, primarily composed of two types of information: organic results, that is, short descriptions and links to relevant Web pages, and sponsored search results, the small textual advertisements often displayed above or to the right of the organic results. Strategies for optimizing each type of result in isolation and the consequent user reaction have been extensively studied; however, the interplay between these two complementary sources of information has been ignored, a situation we aim to change. Our findings indicate that their perceived relative usefulness (as evidenced by user clicks) depends on the nature of the query. Specifically, we found that, when both sources focus on the same intent, for navigational queries there is a clear competition between ads and organic results, while for non-navigational queries this competition turns into synergy. Cristian Danescu-Niculescu-Mizil, Andrei Z. Broder, Evgeniy Gabrilovich, Vanja Josifovski, Bo Pang 0001 |
WWW | 3 |
| 2009 | Translating relevance scores to probabilities for contextual advertisingabstractInformation retrieval systems conventionally assess document relevance using the bag of words model. Consequently, relevance scores of documents retrieved for different queries are often difficult to compare, as they are computed on different (or even disjoint) sets of textual features. Many tasks, such as federation of search results or global thresholding of relevance scores, require that scores be globally comparable. To achieve this, in this paper we propose methods for non-monotonic transformation of relevance scores into probabilities for a contextual advertising selection engine that uses a vector space model. The calibration of the raw scores is based on historical click data. Deepak Agarwal, Evgeniy Gabrilovich, Rob Hall 0001, Vanja Josifovski, Rajiv Khanna |
CIKM | 2 |
| 2009 | What happens after an ad click?: quantifying the impact of landing pages in web advertisingabstractUnbeknownst to most users, when a query is submitted to a search engine two distinct searches are performed: the organic or algorithmic search that returns relevant Web pages and related data (maps, images, etc.), and the sponsored search that returns paid advertisements. While an enormous amount of work has been invested in understanding the user interaction with organic search, surprisingly little research has been dedicated to what happens after an ad is clicked, a situation we aim to correct. Hila Becker, Andrei Z. Broder, Evgeniy Gabrilovich, Vanja Josifovski, Bo Pang 0001 |
CIKM | 3 |
| 2009 | Context transfer in search advertisingabstractWe define and study the process of context transfer in search advertising, which is the transition of a user from the context of Web search to the context of the landing page that follows an ad-click. We conclude that in the vast majority of cases, the user is shown one of three types of pages, which can be accurately distinguished using automatic text classification. Hila Becker, Andrei Z. Broder, Evgeniy Gabrilovich, Vanja Josifovski, Bo Pang 0001 |
SIGIR | 3 |
| 2009 | Cross-language query classification using web search for exogenous knowledgeabstractThe non-English Web is growing at phenomenal speed, but available language processing tools and resources are predominantly English-based. Taxonomies are a case in point: while there are plenty of commercial and non-commercial taxonomies for the English Web, taxonomies for other languages are either not available or of arguable quality. Given that building comprehensive taxonomies for each language is prohibitively expensive, it is natural to ask whether existing English taxonomies can be leveraged, possibly via machine translation, to enable text processing tasks in other languages. Our experimental results confirm that the answer is affirmative with respect to at least one task. In this study we focus on query classification, which is essential for understanding the user intent both in Web search and in online advertising. We propose a robust method for classifying non-English queries into an English taxonomy, using an existing English text classifier and off-the-shelf machine translation systems. In particular, we show that by considering the Web search results in the query's original language as additional sources of information, we can alleviate the effect of erroneous machine translation. Empirical evaluation on query sets in languages as diverse as Chinese and Russian yields very encouraging results; consequently, we believe that our approach is also applicable to many additional languages. Andrei Z. Broder, Evgeniy Gabrilovich, Vanja Josifovski, Bo Pang 0001 |
WSDM | 3 |
| 2009 | Online expansion of rare queries for sponsored searchabstractSponsored search systems are tasked with matching queries Andrei Z. Broder, Peter Ciccolo, Evgeniy Gabrilovich, Vanja Josifovski, Donald Metzler, Lance Riedel, Jeffrey Yuan |
WWW | 3 |
| 2009 | Towards intent-driven bidterm suggestionabstractIn online advertising, pervasive in commercial search engines, advertisers typically bid on few terms, and the scarcity of data makes ad matching difficult. Suggesting additional bidterms can significantly improve ad clickability and conversion rates. In this paper, we present a large-scale bidterm suggestion system that models an advertiser's intent and finds new bidterms consistent with that intent. Preliminary experiments show that our system significantly increases the coverage of a state of the art production system used at Yahoo while maintaining comparable precision. Patrick Pantel, Ana-Maria Popescu, Evgeniy Gabrilovich |
WWW | 4 |
| 2009 | Wikipedia-based Semantic Interpretation for Natural Language ProcessingabstractAdequate representation of natural language semantics requires access to vast amounts of common sense and domain-specific world knowledge. Prior work in the field was based on purely statistical techniques that did not make use of background knowledge, on limited lexicographic knowledge bases such as WordNet, or on huge manual efforts such as the CYC project. Here we propose a novel method, called Explicit Semantic Analysis (ESA), for fine-grained semantic interpretation of unrestricted natural language texts. Our method represents meaning in a high-dimensional space of concepts derived from Wikipedia, the largest encyclopedia in existence. We explicitly represent the meaning of any text in terms of Wikipedia-based concepts. We evaluate the effectiveness of our method on text categorization and on computing the degree of semantic relatedness between fragments of natural language text. Using ESA results in significant improvements over the previous state of the art in both tasks. Importantly, due to the use of natural concepts, the ESA model is easy to explain to human users. Evgeniy Gabrilovich, Shaul Markovitch |
J. Artif. Intell. Res. | 1 |
| 2009 | Classifying search queries using the Web as a source of knowledgeabstractWe propose a methodology for building a robust query classification system that can identify thousands of query classes, while dealing in real time with the query volume of a commercial Web search engine. We use a pseudo relevance feedback technique: given a query, we determine its topic by classifying the Web search results retrieved by the query. Motivated by the needs of search advertising, we primarily focus on rare queries, which are the hardest from the point of view of machine learning, yet in aggregate account for a considerable fraction of search engine traffic. Empirical evaluation confirms that our methodology yields a considerably higher classification accuracy than previously reported. We believe that the proposed methodology will lead to better matching of online ads to rare queries and overall to a better user experience. Evgeniy Gabrilovich, Andrei Z. Broder, Marcus Fontoura, Amruta Joshi, Vanja Josifovski, Lance Riedel, Tong Zhang 0001 |
ACM Trans. Web | 1 |
| 2008 | Concept-Based Feature Generation and Selection for Information Retrieval
Ofer Egozi, Evgeniy Gabrilovich, Shaul Markovitch |
AAAI | 2 |
| 2008 | To swing or not to swing: learning when (not) to advertiseabstractWeb textual advertising can be interpreted as a search problem over the corpus of ads available for display in a particular context. In contrast to conventional information retrieval systems, which always return results if the corpus contains any documents lexically related to the query, in Web advertising it is acceptable, and occasionally even desirable, not to show any results. When no ads are relevant to the user's interests, then showing irrelevant ads should be avoided since they annoy the user and produce no economic benefit. In this paper we pose a decision problem to swing, that is, whether or not to show any of the ads for the incoming request. We propose two methods for addressing this problem, a simple thresholding approach and a machine learning approach, which collectively analyzes the set of candidate ads augmented with external knowledge. Our experimental evaluation, based on over 28,000 editorial judgments, shows that we are able to predict, with high accuracy, when to swing for both content match and sponsored search advertising. Andrei Z. Broder, Massimiliano Ciaramita, Marcus Fontoura, Evgeniy Gabrilovich, Vanja Josifovski, Donald Metzler, Vanessa Murdock 0001, Vassilis Plachouras |
CIKM | 4 |
| 2008 | Search advertising using web relevance feedbackabstractThe business of Web search, a $10 billion industry, relies heavily on sponsored search, whereas a few carefully-selected paid advertisements are displayed alongside algorithmic search results. A key technical challenge in sponsored search is to select ads that are relevant for the user's query. Identifying relevant ads is challenging because queries are usually very short, and because users, consciously or not, choose terms intended to lead to optimal Web search results and not to optimal ads. Furthermore, the ads themselves are short and usually formulated to capture the reader's attention rather than to facilitate query matching. Andrei Z. Broder, Peter Ciccolo, Marcus Fontoura, Evgeniy Gabrilovich, Vanja Josifovski, Lance Riedel |
CIKM | 4 |
| 2008 | Optimizing relevance and revenue in ad search: a query substitution approachabstractThe primary business model behind Web search is based on textual advertising, where contextually relevant ads are displayed alongside search results. We address the problem of selecting these ads so that they are both relevant to the queries and profitable to the search engine, showing that optimizing ad relevance and revenue is not equivalent. Selecting the best ads that satisfy these constraints also naturally incurs high computational costs, and time constraints can lead to reduced relevance and profitability. We propose a novel two-stage approach, which conducts most of the analysis ahead of time. An offine preprocessing phase leverages additional knowledge that is impractical to use in real time, and rewrites frequent queries in a way that subsequently facilitates fast and accurate online matching. Empirical evaluation shows that our method optimized for relevance matches a state-of-the-art method while improving expected revenue. When optimizing for revenue, we see even more substantial improvements in expected revenue. Filip Radlinski, Andrei Z. Broder, Peter Ciccolo, Evgeniy Gabrilovich, Vanja Josifovski, Lance Riedel |
SIGIR | 4 |
| 2007 | Just-in-time contextual advertisingabstractContextual Advertising is a type of Web advertising, which, given the URL of a Web page, aims to embed into the page (typically via JavaScript) the most relevant textual ads available. For static pages that are displayed repeatedly, the matching of ads can be based on prior analysis of their entire content; however, ads need to be matched also to new or dynamically created pages that cannot be processed ahead of time. Analyzing the entire body of such pages on-the-fly entails prohibitive communication and latency costs. To solve the three-horned dilemma of either low-relevance or high-latency or high-load, we propose to use text summarization techniques paired with external knowledge (exogenous to the page) to craft short page summaries in real time. Empirical evaluation proves that matching ads on the basis of such summaries does not sacrifice relevance, and is competitive with matching based on the entire page content. Specifically, we found that analyzing a carefully selected 5% fraction of the page text sacrifices only 1%-3% in ad relevance. Furthermore, our summaries are fully compatible with the standard JavaScript mechanisms used for ad placement: they can be produced at ad-display time by simple additions to the usual script, and they only add 500-600 bytes to the usual request. Aris Anagnostopoulos, Andrei Z. Broder, Evgeniy Gabrilovich, Vanja Josifovski, Lance Riedel |
CIKM | 3 |
| 2007 | Computing Semantic Relatedness Using Wikipedia-based Explicit Semantic Analysis
Evgeniy Gabrilovich, Shaul Markovitch |
IJCAI | 1 |
| 2007 | Robust classification of rare queries using web knowledgeabstractWe propose a methodology for building a practical robust query classification system that can identify thousands of query classes with reasonable accuracy, while dealing in real-time with the query volume of a commercial web search engine. We use a blind feedback technique: given a query, we determine its topic by classifying the web search results retrieved by the query. Motivated by the needs of search advertising, we primarily focus on rare queries, which are the hardest from the point of view of machine learning, yet in aggregation account for a considerable fraction of search engine traffic. Empirical evaluation confirms that our methodology yields a considerably higher classification accuracy than previously reported. We believe that the proposed methodology will lead to better matching of online ads to rare queries and overall to a better user experience. Andrei Z. Broder, Marcus Fontoura, Evgeniy Gabrilovich, Amruta Joshi, Vanja Josifovski, Tong Zhang 0001 |
SIGIR | 3 |
| 2007 | Harnessing the Expertise of 70, 000 Human Editors: Knowledge-Based Feature Generation for Text Categorization
Evgeniy Gabrilovich, Shaul Markovitch |
J. Mach. Learn. Res. | 1 |
| 2006 | Overcoming the Brittleness Bottleneck using Wikipedia: Enhancing Text Categorization with Encyclopedic Knowledge
Evgeniy Gabrilovich, Shaul Markovitch |
AAAI | 1 |
| 2005 | Feature Generation for Text Categorization Using World Knowledge
Evgeniy Gabrilovich, Shaul Markovitch |
IJCAI | 1 |
| 2004 | Text categorization with many redundant features: using aggressive feature selection to make SVMs competitive with C4.5abstractText categorization algorithms usually represent documents as bags of words and consequently have to deal with huge numbers of features. Most previous studies found that the majority of these features are relevant for classification, and that the performance of text categorization with support vector machines peaks when no feature selection is performed. We describe a class of text categorization problems that are characterized with many redundant features. Even though most of these features are relevant, the underlying concepts can be concisely captured using only a few features, while keeping all of them has substantially detrimental effect on categorization accuracy. We develop a novel measure that captures feature redundancy, and use it to analyze a large collection of datasets. We show that for problems plagued with numerous redundant features the performance of C4.5 is significantly superior to that of SVM, while aggressive feature selection allows SVM to beat C4.5 by a narrow margin. Evgeniy Gabrilovich, Shaul Markovitch |
ICML | 1 |
| 2004 | Parameterized generation of labeled datasets for text categorization based on a hierarchical directoryabstractAlthough text categorization is a burgeoning area of IR research, readily available test collections in this field are surprisingly scarce. We describe a methodology and system (named ACCIO) for automatically acquiring labeled datasets for text categorization from the World Wide Web, by capitalizing on the body of knowledge encoded in the structure of existing hierarchical directories such as the Open Directory. We define parameters of categories that make it possible to acquire numerous datasets with desired properties, which in turn allow better control over categorization experiments. In particular, we develop metrics that estimate the difficulty of a dataset by examining the host directory structure. These metrics are shown to be good predictors of categorization accuracy that can be achieved on a dataset, and serve as efficient heuristics for generating datasets subject to user's requirements. A large collection of automatically generated datasets are made available for other researchers to use. Dmitry Davidov, Evgeniy Gabrilovich, Shaul Markovitch |
SIGIR | 2 |
| 2004 | Newsjunkie: providing personalized newsfeeds via analysis of information noveltyabstractWe present a principled methodology for filtering news stories by formal measures of information novelty, and show how the techniques can be usedto custom-tailor news feeds based on information that a user has already reviewed. We review methods for analyzing novelty and then describe Newsjunkie, a system that personalizes news for users by identifying the novelty of stories in the context of stories they have already reviewed. Newsjunkie employs novelty-analysis algorithms that represent articles as words and named entities. The algorithms analyze inter-andintra-document dynamics by considering how information evolves over timefrom article to article, as well as within individual articles. We review the results of a user study undertaken to gauge the value of the approachover legacy time-based review of newsfeeds, and also to compare the performance of alternate distance metrics that are used to estimate the dissimilarity between candidate new articles and sets of previously reviewed articles. Evgeniy Gabrilovich, Susan T. Dumais, Eric Horvitz |
WWW | 1 |
| 2002 | Placing search in context: the concept revisitedabstractKeyword-based search engines are in widespread use today as a popular means for Web-based information retrieval. Although such systems seem deceptively simple, a considerable amount of skill is required in order to satisfy non-trivial information needs. This paper presents a new conceptual paradigm for performing search in context, that largely automates the search process, providing even non-professional users with highly relevant results. This paradigm is implemented in practice in the IntelliZap system, where search is initiated from a text query marked by the user in a document she views, and is guided by the text surrounding the marked query in that document ("the context"). The context-driven information retrieval process involves semantic keyword extraction and clustering to automatically generate new, augmented queries. The latter are submitted to a host of general and domain-specific search engines. Search results are then semantically reranked, using context. Experimental results testify that using context to guide search, effectively offers even inexperienced users an advanced search tool on the Web. Lev Finkelstein, Evgeniy Gabrilovich, Yossi Matias, Ehud Rivlin, Zach Solan, Gadi Wolfman, Eytan Ruppin |
ACM Trans. Inf. Syst. | 2 |
| 2001 | Placing search in context: the concept revisitedabstractArticle Share on Placing search in context: the concept revisited Authors: Lev Finkelstein Zapper Technologies Inc., 3 Azrieli Center, Tel Aviv 67023, Israel Zapper Technologies Inc., 3 Azrieli Center, Tel Aviv 67023, IsraelView Profile , Evgeniy Gabrilovich Zapper Technologies Inc., 3 Azrieli Center, Tel Aviv 67023, Israel Zapper Technologies Inc., 3 Azrieli Center, Tel Aviv 67023, IsraelView Profile , Yossi Matias Zapper Technologies Inc., 3 Azrieli Center, Tel Aviv 67023, Israel Zapper Technologies Inc., 3 Azrieli Center, Tel Aviv 67023, IsraelView Profile , Ehud Rivlin Zapper Technologies Inc., 3 Azrieli Center, Tel Aviv 67023, Israel Zapper Technologies Inc., 3 Azrieli Center, Tel Aviv 67023, IsraelView Profile , Zach Solan Zapper Technologies Inc., 3 Azrieli Center, Tel Aviv 67023, Israel Zapper Technologies Inc., 3 Azrieli Center, Tel Aviv 67023, IsraelView Profile , Gadi Wolfman Zapper Technologies Inc., 3 Azrieli Center, Tel Aviv 67023, Israel Zapper Technologies Inc., 3 Azrieli Center, Tel Aviv 67023, IsraelView Profile , Eytan Ruppin Zapper Technologies Inc., 3 Azrieli Center, Tel Aviv 67023, Israel Zapper Technologies Inc., 3 Azrieli Center, Tel Aviv 67023, IsraelView Profile Authors Info & Claims WWW '01: Proceedings of the 10th international conference on World Wide WebMay 2001 Pages 406–414https://doi.org/10.1145/371920.372094Online:01 April 2001Publication History 319citation2,268DownloadsMetricsTotal Citations319Total Downloads2,268Last 12 Months193Last 6 weeks26 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Lev Finkelstein, Evgeniy Gabrilovich, Yossi Matias, Ehud Rivlin, Zach Solan, Gadi Wolfman, Eytan Ruppin |
WWW | 2 |
| 1998 | System Demonstration Natural Language Generation With Abstract Machine
Evgeniy Gabrilovich, Nissim Francez, Shuly Wintner |
INLG | 1 |