Mihajlo Grbovic

dblp:77/1208 · DBLP profile ↗
← Back
30ranked-venue papers in the field
16as first author
7since 2021 · last 2025
0009-0003-0928-8776ORCID · corroborated

Domains — venue-derived; a paper can count in several

Data Mining & Knowledge Discovery · 16 (10 first)Information Retrieval & Web Search · 11 (4 first)Big Data, Cloud & Distributed Data Systems · 3 (2 first)
YearPublicationVenuePosition
2025 BiListing: Modality Alignment for Listings
abstract
Airbnb is a leader in offering travel accommodations. Airbnb has historically relied on structured data to understand, rank, and recommend listings to guests due to the limited capabilities and associated complexity arising from extracting meaningful information from text and images. With the rise of representation learning, leveraging rich information from text and photos has become easier. A popular approach has been to create embeddings for text documents and images to enable use cases of computing similarities between listings or using embeddings as features in an ML model. However, an Airbnb listing has diverse unstructured data: multiple images, various unstructured text documents such as title, description, and reviews, making this approach challenging. Specifically, it is a non-trivial task to combine multiple embeddings of different pieces of information to reach a single representation. This paper proposes BiListing, for Bimodal Listing, an approach to align text and photos of a listing by leveraging large-language models and pretrained language-image models. The BiListing approach has several favorable characteristics: capturing unstructured data into a single embedding vector per listing and modality, enabling zero-shot capability to search inventory efficiently in user-friendly semantics, overcoming the cold start problem, and enabling listing-to-listing search along a single modality, or both. We conducted offline and online tests to leverage the BiListing embeddings in the Airbnb search ranking model, and successfully deployed it in production, achieved 0.425% of NDCB gain, and drove tens of millions in incremental revenue.
Guillaume Guy, Mihajlo Grbovic, Chun How Tan
CIKM2
2025 TSMO 2025: Two-sided Marketplace Optimization: Search, Discovery, Matching, Pricing & Growth
abstract
In recent years, two-sided marketplaces have emerged as viable business models in many real-world applications. In particular, we have moved from the social network paradigm to a network with two distinct types of participants representing the supply and demand of a specific good. Examples of industries include but are not limited to accommodation (Airbnb, Booking.com), video content (YouTube, Instagram, TikTok), ridesharing (Uber, Lyft), online shops (Etsy, Ebay, Facebook Marketplace), music (Spotify, Amazon), app stores (Apple App Store, Google App Store) or job sites (LinkedIn). The traditional research in most of these industries focused on satisfying the demand. OTAs would sell hotel accommodation, TV networks would broadcast their own content, or taxi companies would own their own vehicle fleet. In modern examples like Airbnb, YouTube, Instagram, or Uber, the platforms operate by outsourcing the service they provide to their users, whether they are hosts, content creators or drivers, and have to develop their models considering their needs and goals.
Mihajlo Grbovic, Vladan Radosavljevic, Rui Song 0006, Minmin Chen, Zhiwei (Tony) Qin, Katerina Iliakopoulou-Zanos, Thanasis Noulas, Hongtu Zhu, Fabrizio Silvestri
KDD (2)1
2024 TSMO 2024: Two-sided Marketplace Optimization
abstract
In recent years, two-sided marketplaces have emerged as viable business models in many real-world applications. In particular, we have moved from the social network paradigm to a network with two distinct types of participants representing the supply and demand of a specific good. Examples of industries include but are not limited to accommodation (Airbnb, Booking.com), video content (YouTube, Instagram, TikTok), ridesharing (Uber, Lyft), online shops (Etsy, Ebay, Facebook Marketplace), music (Spotify, Amazon), app stores (Apple App Store, Google App Store) or job sites (LinkedIn). The traditional research in most of these industries focused on satisfying the demand. OTAs would sell hotel accommodation, TV networks would broadcast their own content, or taxi companies would own their own vehicle fleet. In modern examples like Airbnb, YouTube, Instagram, or Uber, the platforms operate by outsourcing the service they provide to their users, whether they are hosts, content creators or drivers, and have to develop their models considering their needs and goals.
Mihajlo Grbovic, Vladan Radosavljevic, Minmin Chen, Katerina Iliakopoulou-Zanos, Thanasis Noulas, Fabrizio Silvestri
KDD1
2024 SURE 2024: Workshop on Strategic and Utility-aware REcommendation
Himan Abdollahpouri, Tonia Danylenko, Masoud Mansoury, Babak Loni, Daniel Russo 0001, Mihajlo Grbovic
RecSys6
2022 AdKDD 2022
abstract
An average consumer spends 8+ hours a day across all devices interacting with online content almost entirely sponsored by advertisements. At over $450B global market size in 2022 and expected to pass $1T by 2027, online advertising has already surpassed traditional ads in global spend. Moreover, computational advertising in particular is perhaps the most visible and ubiquitous application of machine learning and one that interacts directly with consumers. When done right, ads help us enrich our lives and creep us out when done badly. Looking at the published literature over the last few years, many researchers might consider computational advertising as a mature field. Yet, the opposite is true. The field is evolving, however, from ads controlled by monolithic publishers and randomly rotating banner ads to highly personalized content experiences in news feeds on mobile devices and even on TV-all utilizing data amassed from petabytes of stored user data. Ads are far from done.
Abraham Bagherjeiran, Nemanja Djuric, Mihajlo Grbovic, Kuang-chih Lee, Wei Liu 0007, Linsey Pang, Vladan Radosavljevic, Suju Rajan, Kexin Xie
KDD3
2021 AdKDD 2021
abstract
The digital advertising field has always had challenging ML problems, learning from petabytes of data that is highly imbalanced, reactivity times in the milliseconds and more recently compounded with the complex user's path to purchase across devices, across platforms and even online/real-world behavior. The AdKDD workshop continues to be a forum for researchers in advertising, during and after KDD. Our website which hosts slides and abstracts receives approximately 2,000 monthly visits. In surveys during AdKDD 2019 and 2020, over 60% agreed that AdKDD is the reason they attended KDD and over 90% indicated they would attend next year. The 2021 edition is particularly timely because of ongoing developments in ad tracking. We will aim to discuss notions of privacy and tracking enforced by GDPR and through company policies. In addition, we will seek papers that discuss fairness in the context of advertising, to what extent does hyper-personalization work, and on whether the ad industry as a whole needs to think through more effective business models such as incrementality. Ad tech is in an interesting place of evolution/maturity now and we would like to use the AdKDD forum to get the researchers to think not only about the ML aspects but also spark conversations about the societal ones.
Abraham Bagherjeiran, Nemanja Djuric, Mihajlo Grbovic, Kuang-chih Lee, Vladan Radosavljevic, Suju Rajan
KDD3
2021 Toward User Engagement Optimization in 2D Presentation
abstract
Given a collection of items to display, such as news, videos, or products, how can we optimize their presentation order to maximize user engagements, such as click-through rate, viewing time, and the number of purchases? The problem becomes more complicated when the items are displayed in a grid-based, 2-dimensional presentation on a widescreen. For example, many E-Commerce websites such as Amazon and Etsy are displaying their products in a grid-like format, and so are streaming services like Youtube and Netflix. Unlike 1-dimensional space, where products can be naturally ranked in a vertical order, the presentation in 2-dimensional space poses a novel challenge about how to find the best presentation order - should we put the best listing on the top left corner, or the central position on the second row? We are aware that many traditional methods can be applied to solve the problem, such as conducting an attention heatmap web test, or a randomization experiment by shuffling positions of listings. However, both tests are costly to perform and they may downgrade the quality of users' search experience. By contrast, we focus on utilizing existing search log data to reveal propensity of positions, which is readily available and ubiquitously abundant.
Liang Wu 0006, Mihajlo Grbovic, Jundong Li
WSDM2
2020 How Airbnb Tells You Will Enjoy Sunset Sailing in Barcelona? Recommendation in a Two-Sided Travel Marketplace
abstract
A two-sided travel marketplace is an E-Commerce platform where users can both host tours or activities and book them as a guest. When a new guest visits the platform, given tens of thousands of available listings, a natural question is that what kind of activities or trips are the best fit. In order to answer the question, a recommender system needs to both understand characteristics of its inventories, and to know the preferences of each individual guest. In this work, we present our efforts on building a recommender system for Airbnb Experiences, a two-sided online marketplace for tours and activities. Traditional recommender systems rely on abundant user-listing interactions. Airbnb Experiences is an emerging business where many listings and guests are new to the platform. Instead of passively waiting for data to accumulate, we propose novel approaches to identify key features of a listing and estimate guest preference with limited data availability.
Liang Wu 0006, Mihajlo Grbovic
SIGIR2
2018 Real-time Personalization using Embeddings for Search Ranking at Airbnb
abstract
Search Ranking and Recommendations are fundamental problems of crucial interest to major Internet companies, including web search engines, content publishing websites and marketplaces. However, despite sharing some common characteristics a one-size-fits-all solution does not exist in this space. Given a large difference in content that needs to be ranked, personalized and recommended, each marketplace has a somewhat unique challenge. Correspondingly, at Airbnb, a short-term rental marketplace, search and recommendation problems are quite unique, being a two-sided marketplace in which one needs to optimize for host and guest preferences, in a world where a user rarely consumes the same item twice and one listing can accept only one guest for a certain set of dates. In this paper we describe Listing and User Embedding techniques we developed and deployed for purposes of Real-time Personalization in Search Ranking and Similar Listing Recommendations, two channels that drive 99% of conversions. The embedding models were specifically tailored for Airbnb marketplace, and are able to capture guest's short-term and long-term interests, delivering effective home listing recommendations. We conducted rigorous offline testing of the embedding models, followed by successful online tests before fully deploying them into production.
Mihajlo Grbovic, Haibin Cheng
KDD1
2018 Modeling Mobile User Actions for Purchase Recommendation using Deep Memory Networks
abstract
Rapid expansion of mobile devices has brought an unprecedented opportunity for mobile operators and content publishers to reach many users at any point in time. Understanding usage patterns of mobile applications (apps) is an integral task that precedes advertising efforts of providing relevant recommendations to users. However, this task can be very arduous due to the unstructured nature of app data, with sparseness in available information. This study proposes a novel approach to learn representations of mobile user actions using Deep Memory Networks. We validate the proposed approach on millions of app usage sessions built from large scale feeds of mobile app events and mobile purchase receipts. The empirical study demonstrates that the proposed approach performed better compared to several competitive baselines in terms of recommendation precision quality. To the best of our knowledge this is the first study analyzing app usage patterns for purchase recommendation.
Djordje Gligorijevic, Jelena Gligorijevic, Aravindan Raghuveer, Mihajlo Grbovic, Zoran Obradovic
SIGIR4
2018 Workshop on Two-sided Marketplace Optimization: Search, Pricing, Matching & Growth
abstract
The 1st International Workshop on Two-sided Marketplace Optimization: Search, Pricing, Matching & Growth(TSMO) will be held in Los Angeles, California, USA on February 9th, 2018, co-located with the 11th ACM International Conference on Web Search and Data Mining(WSDM). The main objective of the workshop is to address the challenges of two-sided marketplace optimization in web-scale settings. The workshop brings together interdisciplinary researchers in information retrieval, recommender systems, personalization, and related areas, to share, exchange, learn, and develop preliminary results, new concepts, ideas, principles, and methodologies on applying data mining technologies to marketplace optimization. We have constructed an exciting program papers and invited talks that will help us better understand the future of two-sided marketplaces
Mihajlo Grbovic, Thanasis Noulas
WSDM1
2017 Search Ranking And Personalization at Airbnb
abstract
Search ranking is a fundamental problem of crucial interest to major Internet companies, including web search engines, content publishing websites and marketplaces. However, despite sharing some common characteristics a one-size-fits-all solution does not exist in this space. Given a large difference in content that needs to be ranked and the parties affected by ranking, each search ranking problem is somewhat specific. Correspondingly, search ranking at Airbnb is quite unique, being a two-sided marketplace in which one needs to optimize for host and guest preferences, in a world where a user rarely consumes the same item twice and one listing can accept only one guest for a certain set of dates. In this talk, I will discuss challenges we have encountered and Machine Learning solutions we have developed for listing ranking at Airbnb. Specifically, the listing ranking problem boils down to prioritizing listings that are appealing to the guest but at the same time demoting listings that would likely reject the guest, which is not easily solvable using basic matrix completion or a straightforward linear model. I will shed the light on how we jointly optimize the two objectives by leveraging listing quality, location relevance, reviews, host response time as well as guest and host preferences and past booking history. Finally, we will talk about our recent work on using neural network models to train listing and query embeddings for purposes of enhancing search personalization, broad search and type-ahead suggestions, which are core concepts in any modern search.
Mihajlo Grbovic
RecSys1
2017 iPhone's Digital Marketplace: Characterizing the Big Spenders
abstract
With mobile shopping surging in popularity, people are spending ever more money on digital purchases through their mobile devices and phones. However, few large-scale studies of mobile shopping exist. In this paper we analyze a large data set consisting of more than 776M digital purchases made on Apple mobile devices that include songs, apps, and in-app purchases. We find that 61% of all the spending is on in-app purchases and that the top 1% of users are responsible for 59% of all the spending. These big spenders are more likely to be male and older, and less likely to be from the US. We study how they adopt and abandon individual app, and find that, after an initial phase of increased daily spending, users gradually lose interest: the delay between their purchases increases and the spending decreases with a sharp drop toward the end. Finally, we model the in-app purchasing behavior in multiple steps: 1) we model the time between purchases; 2) we train a classifier to predict whether the user will make a purchase from a new app or continue purchasing from the existing app; and 3) based on the outcome of the previous step, we attempt to predict the exact app, new or existing, from which the next purchase will come. The results yield new insights into spending habits in the mobile digital marketplace.
Farshad Kooti, Mihajlo Grbovic, Luca Maria Aiello, Eric Bax, Kristina Lerman
WSDM2
2016 Network-Efficient Distributed Word2vec Training System for Large Vocabularies
abstract
Word2vec is a popular family of algorithms for unsupervised training of dense vector representations of words on large text corpuses. The resulting vectors have been shown to capture semantic relationships among their corresponding words, and have shown promise in reducing a number of natural language processing (NLP) tasks to mathematical operations on these vectors. While heretofore applications of word2vec have centered around vocabularies with a few million words, wherein the vocabulary is the set of words for which vectors are simultaneously trained, novel applications are emerging in areas outside of NLP with vocabularies comprising several 100 million words. Existing word2vec training systems are impractical for training such large vocabularies as they either require that the vectors of all vocabulary words be stored in the memory of a single server or suffer unacceptable training latency due to massive network data transfer. In this paper, we present a novel distributed, parallel training system that enables unprecedented practical training of vectors for vocabularies with several 100 million words on a shared cluster of commodity servers, using far less network traffic than the existing solutions. We evaluate the proposed system on a benchmark data set, showing that the quality of vectors does not degrade relative to non-distributed training. Finally, for several quarters, the system has been deployed for the purpose of matching queries to ads in Gemini, the sponsored search advertising platform at Yahoo, resulting in significant improvement of business metrics.
Erik Ordentlich, Lee Yang, Andy Feng, Peter Cnudde, Mihajlo Grbovic, Nemanja Djuric, Vladan Radosavljevic, Gavin Owens
CIKM5
2016 Scalable Semantic Matching of Queries to Ads in Sponsored Search Advertising
abstract
Sponsored search represents a major source of revenue for web search engines. The advertising model brings a unique possibility for advertisers to target direct user intent communicated through a search query, usually done by displaying their ads alongside organic search results for queries deemed relevant to their products or services. However, due to a large number of unique queries, it is particularly challenging for advertisers to identify all relevant queries. For this reason search engines often provide a service of advanced matching, which automatically finds additional relevant queries for advertisers to bid on. We present a novel advance match approach based on the idea of semantic embeddings of queries and ads. The embeddings were learned using a large data set of user search sessions, consisting of search queries, clicked ads and search links, while utilizing contextual information such as dwell time and skipped ads. To address the large-scale nature of our problem, both in terms of data and vocabulary size, we propose a novel distributed algorithm for training of the embeddings. Finally, we present an approach for overcoming a cold-start problem associated with new ads and queries. We report results of editorial evaluation and online tests on actual search traffic. The results show that our approach significantly outperforms baselines in terms of relevance, coverage and incremental revenue. Lastly, as part of this study, we open sourced query embeddings that can be used to advance the field.
Mihajlo Grbovic, Nemanja Djuric, Vladan Radosavljevic, Fabrizio Silvestri, Ricardo Baeza-Yates, Andrew Feng, Erik Ordentlich, Lee Yang, Gavin Owens
SIGIR1
2016 TargetAd2016: 2nd International Workshop on Ad Targeting at Scale
abstract
The 2nd International Workshop on Ad Targeting at Scale will be held in San Francisco, California, USA on February 22nd, 2016, co-located with the 9th ACM International Conference on Web Search and Data Mining (WSDM). The main objective of the workshop is to address the challenges of ad targeting in web-scale settings. The workshop brings together interdisciplinary researchers in computational advertising, recommender systems, personalization, and related areas, to share, exchange, learn, and develop preliminary results, new concepts, ideas, principles, and methodologies on applying data mining technologies to ad targeting. We have constructed an exciting program of eight refereed papers and several invited talks that will help us better understand the future of ad targeting.
Mihajlo Grbovic, Nemanja Djuric, Vladan Radosavljevic
WSDM1
2016 Portrait of an Online Shopper: Understanding and Predicting Consumer Behavior
abstract
Consumer spending accounts for a large fraction of economic footprint of modern countries. Increasingly, consumer activity is moving to the web, where digital receipts of online purchases provide valuable data sources detailing consumer behavior. We consider such data extracted from emails and combined with with consumers' demographic information, which we use to characterize, model, and predict purchasing behavior. We analyze such behavior of consumers in different age and gender groups, and find interesting, actionable patterns that can be used to improve ad targeting systems. For example, we found that the amount of money spent on online purchases grows sharply with age, peaking in the late 30s, while shoppers from wealthy areas tend to purchase more expensive items and buy them more frequently. Furthermore, we look at the influence of social connections on purchasing habits, as well as at the temporal dynamics of online shopping where we discovered daily and weekly behavioral patterns. Finally, we build a model to predict when shoppers are most likely to make a purchase and how much will they spend, showing improvement over baseline approaches. The presented results paint a clear picture of a modern online shopper, and allow better understanding of consumer behavior that can help improve marketing efforts and make shopping more pleasant and efficient experience for online customers.
Farshad Kooti, Kristina Lerman, Luca Maria Aiello, Mihajlo Grbovic, Nemanja Djuric, Vladan Radosavljevic
WSDM4
2015 Gender and Interest Targeting for Sponsored Post Advertising at Tumblr
abstract
As one of the leading platforms for creative content, Tumblr offers advertisers a unique way of creating brand identity. Advertisers can tell their story through images, animation, text, music, video, and more, and can promote that content by sponsoring it to appear as an advertisement in the users' live feeds. In this paper, we present a framework that enabled two of the key targeted advertising components for Tumblr, gender and interest targeting. We describe the main challenges encountered during the development of the framework, which include the creation of a ground truth for training gender prediction models, as well as mapping Tumblr content to a predefined interest taxonomy. For purposes of inferring user interests, we propose a novel semi-supervised neural language model for categorization of Tumblr content (i.e., post tags and post keywords). The model was trained on a large-scale data set consisting of $6.8$ billion user posts, with a very limited amount of categorized keywords, and was shown to have superior performance over the baseline approaches. We successfully deployed gender and interest targeting capability in Yahoo production systems, delivering inference for users that covers more than 90% of daily activities on Tumblr. Online performance results indicate advantages of the proposed approach, where we observed 20% increase in user engagement with sponsored posts in comparison to untargeted campaigns.
Mihajlo Grbovic, Vladan Radosavljevic, Nemanja Djuric, Narayan L. Bhamidipati, Ananth Nagarajan
KDD1
2015 E-commerce in Your Inbox: Product Recommendations at Scale
abstract
In recent years online advertising has become increasingly ubiquitous and effective. Advertisements shown to visitors fund sites and apps that publish digital content, manage social networks, and operate e-mail services. Given such large variety of internet resources, determining an appropriate type of advertising for a given platform has become critical to financial success. Native advertisements, namely ads that are similar in look and feel to content, have had great success in news and social feeds. However, to date there has not been a winning formula for ads in e-mail clients. In this paper we describe a system that leverages user purchase history determined from e-mail receipts to deliver highly personalized product ads to Yahoo Mail users. We propose to use a novel neural language-based algorithm specifically tailored for delivering effective product recommendations, which was evaluated against baselines that included showing popular products and products predicted based on co-occurrence. We conducted rigorous offline testing using a large-scale product purchase data set, covering purchases of more than 29 million users from 172 e-commerce websites. Ads in the form of product recommendations were successfully tested on online traffic, where we observed a steady 9% lift in click-through rates over other ad formats in mail, as well as comparable lift in conversion rates. Following successful tests, the system was launched into production during the holiday season of 2014.
Mihajlo Grbovic, Vladan Radosavljevic, Nemanja Djuric, Narayan L. Bhamidipati, Jaikit Savla, Varun Bhagwan, Doug Sharp
KDD1
2015 Context- and Content-aware Embeddings for Query Rewriting in Sponsored Search
abstract
Search engines represent one of the most popular web services, visited by more than 85% of internet users on a daily basis. Advertisers are interested in making use of this vast business potential, as very clear intent signal communicated through the issued query allows effective targeting of users. This idea is embodied in a sponsored search model, where each advertiser maintains a list of keywords they deem indicative of increased user response rate with regards to their business. According to this targeting model, when a query is issued all advertisers with a matching keyword are entered into an auction according to the amount they bid for the query, and the winner gets to show their ad. One of the main challenges is the fact that a query may not match many keywords, resulting in lower auction value, lower ad quality, and lost revenue for advertisers and publishers. Possible solution is to expand a query into a set of related queries and use them to increase the number of matched ads, called query rewriting. To this end, we propose rewriting method based on a novel query embedding algorithm, which jointly models query content as well as its context within a search session. As a result, queries with similar content and context are mapped into vectors close in the embedding space, which allows expansion of a query via simple K-nearest neighbor search in the projected space. The method was trained on more than 12 billion sessions, one of the largest corpuses reported thus far, and evaluated on both public TREC data set and in-house sponsored search data set. The results show the proposed approach significantly outperformed existing state-of-the-art, strongly indicating its benefits and the monetization potential.
Mihajlo Grbovic, Nemanja Djuric, Vladan Radosavljevic, Fabrizio Silvestri, Narayan L. Bhamidipati
SIGIR1
2015 Hierarchical Neural Language Models for Joint Representation of Streaming Documents and their Content
abstract
We consider the problem of learning distributed representations for documents in data streams. The documents are represented as low-dimensional vectors and are jointly learned with distributed vector representations of word tokens using a hierarchical framework with two embedded neural language models. In particular, we exploit the context of documents in streams and use one of the language models to model the document sequences, and the other to model word sequences within them. The models learn continuous vector representations for both word tokens and documents such that semantically similar documents and words are close in a common vector space. We discuss extensions to our model, which can be applied to personalized recommendation and social relationship mining by adding further user layers to the hierarchy, thus learning user-specific vectors to represent individual preferences. We validated the learned representations on a public movie rating data set from MovieLens, as well as on a large-scale Yahoo News data comprising three months of user activity logs collected on Yahoo servers. The results indicate that the proposed model can learn useful representations of both documents and word tokens, outperforming the current state-of-the-art by a large margin.
Nemanja Djuric, Vladan Radosavljevic, Mihajlo Grbovic, Narayan L. Bhamidipati
WWW4
2015 Evolution of Conversations in the Age of Email Overload
abstract
Email is a ubiquitous communications tool in the workplace and plays an important role in social interactions. Previous studies of email were largely based on surveys and limited to relatively small populations of email users within organizations. In this paper, we report results of a large-scale study of more than 2 million users exchanging 16 billion emails over several months. We quantitatively characterize the replying behavior in conversations within pairs of users. In particular, we study the time it takes the user to reply to a received message and the length of the reply sent. We consider a variety of factors that affect the reply time and length, such as the stage of the conversation, user demographics, and use of portable devices. In addition, we study how increasing load affects emailing behavior. We find that as users receive more email messages in a day, they reply to a smaller fraction of them, using shorter replies. However, their responsiveness remains intact, and they may even reply to emails faster. Finally, we predict the time to reply, length of reply, and whether the reply ends a conversation. We demonstrate considerable improvement over the baseline in all three prediction tasks, showing the significant role that the factors that we uncover play, in determining replying behavior. We rank these factors based on their predictive power. Our findings have important implications for understanding human behavior and designing better email management applications for tasks like ranking unread emails.
Farshad Kooti, Luca Maria Aiello, Mihajlo Grbovic, Kristina Lerman, Amin Mantrach
WWW3
2014 How Many Folders Do You Really Need?: Classifying Email into a Handful of Categories
abstract
Email classification is still a mostly manual task. Consequently, most Web mail users never define a single folder. Recently however, automatic classification offering the same categories to all users has started to appear in some Web mail clients, such as AOL or Gmail. We adopt this approach, rather than previous (unsuccessful) personalized approaches because of the change in the nature of consumer email traffic, which is now dominated by (non-spam) machine-generated email. We propose here a novel approach for (1) automatically distinguishing between personal and machine-generated email and (2) classifying messages into latent categories, without requiring users to have defined any folder. We report how we have discovered that a set of 6 "latent" categories (one for human- and the others for machine-generated messages) can explain a significant portion of email traffic. We describe in details the steps involved in building a Web-scale email categorization system, from the collection of ground-truth labels, the selection of features to the training of models. Experimental evaluation was performed on more than 500 billion messages received during a period of six months by users of Yahoo mail service, who elected to be part of such research studies. Our system achieved precision and recall rates close to 90% and the latent categories we discovered were shown to cover 70% of both email traffic and email search queries. We believe that these results pave the way for a change of approach in the Web mail industry, and could support the invention of new large-scale email discovery paradigms that had not been possible before.
Mihajlo Grbovic, Guy Halawi, Zohar S. Karnin, Yoelle Maarek
CIKM1
2014 Hidden Conditional Random Fields with Deep User Embeddings for Ad Targeting
abstract
Estimating a user's propensity to click on a display ad or purchase a particular item is a critical task in targeted advertising, a burgeoning online industry worth billions of dollars. Better and more accurate estimation methods result in improved online user experience, as only relevant and interesting ads are shown, and may also lead to large benefits for advertisers, as targeted users are more likely to click or make a purchase. In this paper we address this important problem, and propose an approach for improved estimation of ad click or conversion probability based on a sequence of user's online actions, modeled using Hidden Conditional Random Fields (HCRF) model. In addition, in order to address the sparsity issue at the input side of the HCRF model, we propose to learn distributed, low-dimensional representations of user actions through a directed skip-gram, a neural architecture suitable for sequential data. Experimental results on a real-world data set comprising thousands of user sessions collected at Yahoo servers clearly indicate the benefits and the potential of the proposed approach, which outperformed competing state-of-the-art algorithms and obtained significant improvements in terms of retrieval measures.
Nemanja Djuric, Vladan Radosavljevic, Mihajlo Grbovic, Narayan L. Bhamidipati
ICDM3
2013 Distributed confidence-weighted classification on MapReduce
abstract
Explosive growth in data size, data complexity, and data rates, triggered by emergence of high-throughput technologies such as remote sensing, crowd-sourcing, social networks, or computational advertising, in recent years has led to an increasing availability of data sets of unprecedented scales, with billions of high-dimensional data examples stored on hundreds of terabytes of memory. In order to make use of this large-scale data and extract useful knowledge, researchers in machine learning and data mining communities are faced with numerous challenges, since the classification algorithms designed for standard desktop computers are not capable of addressing these problems due to memory and time constraints. As a result, there exists an evident need for development of novel, more scalable algorithms that can handle large data sets. In this paper we propose such method, named AROW-MR, a linear SVM solver for efficient training of recently proposed confidence-weighted (CW) classifiers. Linear CW models maintain a Gaussian distribution over parameter vectors, thus allowing a user to estimate, in addition to separating hyperplane between two classes, parameter confidence as well. The proposed method employs MapReduce framework to train CW classifier in a distributed way, obtaining significant improvements in both training time and accuracy. This is achieved through training of local CW classifiers on each mapper, followed by optimally combining local classifiers on the reducer to obtain aggregated, more accurate CW linear model. We validated the proposed algorithm on synthetic data, and further showed that AROW-MR algorithm outperforms the baseline classifiers on an industrial, large-scale task of Ad Latency prediction, with nearly one billion examples.
Nemanja Djuric, Mihajlo Grbovic, Slobodan Vucetic
IEEE BigData2
2013 Large scale ad latency analysis
abstract
Late web display advertisements are problematic for both the user experience and the monetary machinery powering the display advertising industry. If a web page is delivered to a user but the ad fails to load in time, the publisher cannot charge the advertiser for that impression. Detecting whether a specific ad will render in time could give the publisher a choice to show that ad or another one. Further, discovering the root causes of latency, possibly over time as new violators emerge, would allow the publisher to address the actionable issues. We propose a system that predicts, at serve time, which ads are likely to have high latency. Once identified we can either ignore those ads, even if they win the auction, or apply a penalty to those ads. In addition, our system collects the daily impression logs, consisting of different types of observations measured at serve time and the associated latency in milliseconds, and analyzes the data to identify the features associated with late ads and likely to be causing the delay.
Mihajlo Grbovic, Jon Malkin, Hirakendu Das
IEEE BigData1
2012 Supervised Clustering of Label Ranking Data
abstract
In this paper we study supervised clustering in the context of label ranking data. Segmentation of such complex data has many potential real-world applications. For example, in target marketing, the goal is to cluster customers in the feature space by taking into consideration the assigned, potentially incomplete product preferences, such that the preferences of instances within a cluster are more similar than the preferences of customers in the other clusters. We establish several heuristic baselines for this application that make use of well-known algorithms such as K-means, and propose a principled algorithm specifically tailored for this type of clustering. It is based on the Plackett-Luce (PL) probabilistic ranking model. Each cluster is represented as a union of Voronoi cells defined by a set of prototypes and is assigned a set of PL label scores that determine the cluster-specific label ranking. The unknown cluster PL parameters and prototype positions are determined using a supervised learning technique. Cluster membership and ranking for a new instance is determined by membership of its nearest prototype. The proposed algorithms were empirically evaluated on synthetic and reallife label ranking data. The PL-based method was superior to the heuristically-based supervised clustering approaches. The proposed PL-based algorithm was also evaluated on the task of label ranking prediction. The results showed that it is highly competitive to the state of the art label ranking algorithms, and that it is particularly accurate on data with partial rankings.
Mihajlo Grbovic, Nemanja Djuric, Slobodan Vucetic
SDM1
2011 Tracking Concept Change with Incremental Boosting by Minimization of the Evolving Exponential Loss
Mihajlo Grbovic, Slobodan Vucetic
ECML/PKDD (1)1
2009 Decentralized Estimation Using Learning Vector Quantization
abstract
A decentralized estimation system consists of n distributed data sources S1... Snand a fusion center. The data sources produce multivariate random vectors X1... Xnthat are transmitted to the fusion center in the form of messages Z1... Zn, Zi= alpha1(X1). Due to communication constraints, Ziis a discrete variable with cardinality Mirepresented as an integer from a set {1...Mi}. At the fusion center, the goal is to estimate the conditional expectation of unobserved variable Y, E(Y|x1... xn), by fusion function h(z1... zn). The challenge is to find quantization functions alpha1... alphanand fusion function h such that the estimation error is minimized under given communication constraints.
Mihajlo Grbovic, Slobodan Vucetic
DCC1
2009 Regression Learning Vector Quantization
abstract
Learning vector quantization (LVQ) is a popular class of nearest prototype classifiers for multiclass classification. Learning algorithms from this family are widely used because of their intuitively clear learning process and ease of implementation. In this paper we propose an extension of the LVQ algorithm to regression. Just like the LVQ algorithm, the proposed modification uses a supervised learning procedure to learn the best prototype positions, but unlike LVQ algorithm for classification, it also learns the best prototype target values. This results in the effective partition of the feature space, similar to the one the K-means algorithm would make. Experimental results on benchmark datasets showed that the proposed regression LVQ algorithm performs better than the nearest prototype competitors that choose prototypes randomly or through K-means clustering, classification LVQ on quantized target values, and similarly to the memory-based Parzen window and nearest neighbor algorithms.
Mihajlo Grbovic, Slobodan Vucetic
ICDM1