VLDB 2026 Research / reviewers in the wild / expert
Vladan Radosavljevic
dblp:17/1226
· DBLP profile ↗
25ranked-venue papers in the field
2as first author
8since 2021 · last 2026
0009-0006-9128-1101ORCID · reported
Domains — venue-derived; a paper can count in several
Data Mining & Knowledge Discovery · 19 (2 first)Information Retrieval & Web Search · 6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GenAI4SM: Generative AI for Streaming MediaabstractStreaming media has become a popular medium for consumers of all ages, with people spending several hours a day streaming videos, games, music, audiobooks or podcasts across devices. Most global streaming services have introduced Generative Artificial Intelligence (GenAI) into their operations to personalize consumer experience, improve content, and further enhance the value proposition of streaming services. Despite the rapid growth, there is a need to bridge the gap between academic research and industry requirements and build connections between researchers and practitioners in the field. This workshop aims to provide a unique forum for practitioners and researchers interested in GenAI to get together, exchange ideas and get a pulse for the state of the art in research and burning issues in the industry. Vladan Radosavljevic, Sudarshan Lamkhede, Praveen Chandar, Arnab Bhadury, Tao Ye 0001 |
WSDM | 2 |
| 2025 | AdKDD 2025abstractThe digital advertising field has always had challenging ML problems, learning from petabytes of data that is highly imbalanced, reactivity times in the milliseconds, and more recently compounded with the complex user's path to purchase across devices, across platforms, and even online/real-world behavior. The AdKDD workshop continues to be a forum for researchers in advertising, during and after KDD. Our website, which hosts slides and abstracts, continues to receive a large number of monthly visits and active users. In surveys during AdKDD 2019 and 2020, over 60% agreed that AdKDD is the reason they attended KDD, and over 90% indicated they would attend next year. The 2025 edition is particularly timely because of the increasing application of Graph-based NN and Generative AI models in advertising. Coupled with privacy-preserving initiatives, such as those enforced by GDPR or CCPA, the future of computational advertising is at an interesting crossroads. For this edition, we plan to solicit papers that span the spectrum of deep user understanding while remaining privacy-preserving. In addition, we will seek papers that discuss fairness in the context of advertising, to what extent does hyper-personalization work, and whether the ad industry as a whole needs to think through more effective business models such as incrementality. We have hosted several academic and industry luminaries as keynote speakers and have found our invited speaker series hosting expert practitioners to be an audience favorite. We will continue fielding a diverse set of keynote speakers and invited talks for this edition as well. As with past editions, we hope to motivate researchers in this space to think not only about the ML aspects but also to spark conversations about the societal impact of online advertising. Abraham Bagherjeiran, Nemanja Djuric, Kuang-chih Lee, Linsey Pang, Vladan Radosavljevic, Suju Rajan |
KDD (2) | 5 |
| 2025 | TSMO 2025: Two-sided Marketplace Optimization: Search, Discovery, Matching, Pricing & GrowthabstractIn recent years, two-sided marketplaces have emerged as viable business models in many real-world applications. In particular, we have moved from the social network paradigm to a network with two distinct types of participants representing the supply and demand of a specific good. Examples of industries include but are not limited to accommodation (Airbnb, Booking.com), video content (YouTube, Instagram, TikTok), ridesharing (Uber, Lyft), online shops (Etsy, Ebay, Facebook Marketplace), music (Spotify, Amazon), app stores (Apple App Store, Google App Store) or job sites (LinkedIn). The traditional research in most of these industries focused on satisfying the demand. OTAs would sell hotel accommodation, TV networks would broadcast their own content, or taxi companies would own their own vehicle fleet. In modern examples like Airbnb, YouTube, Instagram, or Uber, the platforms operate by outsourcing the service they provide to their users, whether they are hosts, content creators or drivers, and have to develop their models considering their needs and goals. Mihajlo Grbovic, Vladan Radosavljevic, Rui Song 0006, Minmin Chen, Zhiwei (Tony) Qin, Katerina Iliakopoulou-Zanos, Thanasis Noulas, Hongtu Zhu, Fabrizio Silvestri |
KDD (2) | 2 |
| 2024 | AdKDD 2024abstractThe digital advertising field has always had challenging ML problems, learning from petabytes of data that is highly imbalanced, reactivity times in the milliseconds, and more recently compounded with the complex user's path to purchase across devices, across platforms, and even online/real-world behavior. The AdKDD workshop continues to be a forum for researchers in advertising, during and after KDD. Our website which hosts slides and abstracts receives approximately 2,000 monthly visits and 1,800 active users during the KDD 2021. In surveys during AdKDD 2019 and 2020, over 60% agreed that AdKDD is the reason they attended KDD, and over 90% indicated they would attend next year. The 2024 edition is particularly timely because of the increasing application of Graph-based NN and Generative AI models in advertising. Coupled with privacy-preserving initiatives enforced by GDPR, CCPA the future of computational advertising is at an interesting crossroads. For this edition, we plan to solicit papers that span the spectrum of deep user understanding while remaining privacy-preserving. In addition, we will seek papers that discuss fairness in the context of advertising, to what extent does hyper-personalization work, and whether the ad industry as a whole needs to think through more effective business models such as incrementality. We have hosted several academic and industry luminaries as keynote speakers and have found our invited speaker series hosting expert practitioners to be an audience favorite. We will continue fielding a diverse set of keynote speakers and invited talks for this edition as well. As with past editions, we hope to motivate researchers in this space to think not only about the ML aspects but also to spark conversations about the societal impact of online advertising. Abraham Bagherjeiran, Nemanja Djuric, Kuang-chih Lee, Linsey Pang, Vladan Radosavljevic, Suju Rajan |
KDD | 5 |
| 2024 | TSMO 2024: Two-sided Marketplace OptimizationabstractIn recent years, two-sided marketplaces have emerged as viable business models in many real-world applications. In particular, we have moved from the social network paradigm to a network with two distinct types of participants representing the supply and demand of a specific good. Examples of industries include but are not limited to accommodation (Airbnb, Booking.com), video content (YouTube, Instagram, TikTok), ridesharing (Uber, Lyft), online shops (Etsy, Ebay, Facebook Marketplace), music (Spotify, Amazon), app stores (Apple App Store, Google App Store) or job sites (LinkedIn). The traditional research in most of these industries focused on satisfying the demand. OTAs would sell hotel accommodation, TV networks would broadcast their own content, or taxi companies would own their own vehicle fleet. In modern examples like Airbnb, YouTube, Instagram, or Uber, the platforms operate by outsourcing the service they provide to their users, whether they are hosts, content creators or drivers, and have to develop their models considering their needs and goals. Mihajlo Grbovic, Vladan Radosavljevic, Minmin Chen, Katerina Iliakopoulou-Zanos, Thanasis Noulas, Fabrizio Silvestri |
KDD | 2 |
| 2023 | AdKDD 2023abstractThe digital advertising field has always had challenging ML problems, learning from petabytes of data that is highly imbalanced, reactivity times in the milliseconds, and more recently compounded with the complex user's path to purchase across devices, across platforms, and even online/real-world behavior. The AdKDD workshop continues to be a forum for researchers in advertising, during and after KDD. Our website which hosts slides and abstracts receives approximately 2,000 monthly visits and 1,800 active users during the KDD 2021. In surveys during AdKDD 2019 and 2020, over 60% agreed that AdKDD is the reason they attended KDD, and over 90% indicated they would attend next year. The 2023 edition is particularly timely because of the increasing application of Graph-based NN and Generative AI models in advertising. Coupled with privacy-preserving initiatives enforced by GDPR, CCPA the future of computational advertising is at an interesting crossroads. For this edition, we plan to solicit papers that span the spectrum of deep user understanding while remaining privacy-preserving. In addition, we will seek papers that discuss fairness in the context of advertising, to what extent does hyper-personalization work, and whether the ad industry as a whole needs to think through more effective business models such as incrementality. We have hosted several academic and industry luminaries as keynote speakers and have found our invited speaker series hosting expert practitioners to be an audience favorite. We will continue fielding a diverse set of keynote speakers and invited talks for this edition as well. As with past editions, we hope to motivate researchers in this space to think not only about the ML aspects but also to spark conversations about the societal impact of online advertising. Abraham Bagherjeiran, Nemanja Djuric, Kuang-chih Lee, Linsey Pang, Vladan Radosavljevic, Suju Rajan |
KDD | 5 |
| 2022 | AdKDD 2022abstractAn average consumer spends 8+ hours a day across all devices interacting with online content almost entirely sponsored by advertisements. At over $450B global market size in 2022 and expected to pass $1T by 2027, online advertising has already surpassed traditional ads in global spend. Moreover, computational advertising in particular is perhaps the most visible and ubiquitous application of machine learning and one that interacts directly with consumers. When done right, ads help us enrich our lives and creep us out when done badly. Looking at the published literature over the last few years, many researchers might consider computational advertising as a mature field. Yet, the opposite is true. The field is evolving, however, from ads controlled by monolithic publishers and randomly rotating banner ads to highly personalized content experiences in news feeds on mobile devices and even on TV-all utilizing data amassed from petabytes of stored user data. Ads are far from done. Abraham Bagherjeiran, Nemanja Djuric, Mihajlo Grbovic, Kuang-chih Lee, Wei Liu 0007, Linsey Pang, Vladan Radosavljevic, Suju Rajan, Kexin Xie |
KDD | 8 |
| 2021 | AdKDD 2021abstractThe digital advertising field has always had challenging ML problems, learning from petabytes of data that is highly imbalanced, reactivity times in the milliseconds and more recently compounded with the complex user's path to purchase across devices, across platforms and even online/real-world behavior. The AdKDD workshop continues to be a forum for researchers in advertising, during and after KDD. Our website which hosts slides and abstracts receives approximately 2,000 monthly visits. In surveys during AdKDD 2019 and 2020, over 60% agreed that AdKDD is the reason they attended KDD and over 90% indicated they would attend next year. The 2021 edition is particularly timely because of ongoing developments in ad tracking. We will aim to discuss notions of privacy and tracking enforced by GDPR and through company policies. In addition, we will seek papers that discuss fairness in the context of advertising, to what extent does hyper-personalization work, and on whether the ad industry as a whole needs to think through more effective business models such as incrementality. Ad tech is in an interesting place of evolution/maturity now and we would like to use the AdKDD forum to get the researchers to think not only about the ML aspects but also spark conversations about the societal ones. Abraham Bagherjeiran, Nemanja Djuric, Mihajlo Grbovic, Kuang-chih Lee, Vladan Radosavljevic, Suju Rajan |
KDD | 6 |
| 2020 | PodRecs: Workshop on Podcast RecommendationsabstractThe last year has been a breakout year for podcasts. There are now over 1 million podcast shows and over 64 million podcast episodes available through public RSS feeds. In the United States, 32% of all people listened to a podcast every month, and forecasts point to global podcast listenership to reach 2.2 billion monthly listeners by 2024. The workshop on Podcast Recommendations (PodRecs), collocated with RecSys 2020, introduces researchers in other domains of recommender systems to the special characteristics and challenges of podcast recommendations: how podcasts are as created, consumed, and how we might see algorithms being designed specifically for podcast content. We hope this workshop will help grow a community of researchers and foster an active research and innovation in the field. Ching-Wei Chen, Longqi Yang 0001, Hongyi Wen, Rosie Jones, Vladan Radosavljevic, Hugues Bouchard |
RecSys | 5 |
| 2019 | Homepage personalization at spotifyabstractWe aim to surface the best of Spotify for each user on the Home page by providing a personalized space where users can find recommendations of playlists, albums, artists, podcasts tailored to their individual preferences. Hundreds of millions of users listen to music on Spotify each month, with more than 50 million daily active users on the Homepage alone. The quality of the recommendations on Home depends on a multi-armed bandit framework that balances exploration and exploitation and allows us to adapt quickly to changes in user preferences. We employ counterfactual training and reasoning to evaluate new algorithms without having to always rely on A/B testing or randomized data collection experiments [3]. Oguz Semerci, Alois Gruson, Catherine M. Edwards, Benjamin Lacker, Clay Gibson, Vladan Radosavljevic |
RecSys | 6 |
| 2016 | Network-Efficient Distributed Word2vec Training System for Large VocabulariesabstractWord2vec is a popular family of algorithms for unsupervised training of dense vector representations of words on large text corpuses. The resulting vectors have been shown to capture semantic relationships among their corresponding words, and have shown promise in reducing a number of natural language processing (NLP) tasks to mathematical operations on these vectors. While heretofore applications of word2vec have centered around vocabularies with a few million words, wherein the vocabulary is the set of words for which vectors are simultaneously trained, novel applications are emerging in areas outside of NLP with vocabularies comprising several 100 million words. Existing word2vec training systems are impractical for training such large vocabularies as they either require that the vectors of all vocabulary words be stored in the memory of a single server or suffer unacceptable training latency due to massive network data transfer. In this paper, we present a novel distributed, parallel training system that enables unprecedented practical training of vectors for vocabularies with several 100 million words on a shared cluster of commodity servers, using far less network traffic than the existing solutions. We evaluate the proposed system on a benchmark data set, showing that the quality of vectors does not degrade relative to non-distributed training. Finally, for several quarters, the system has been deployed for the purpose of matching queries to ads in Gemini, the sponsored search advertising platform at Yahoo, resulting in significant improvement of business metrics. Erik Ordentlich, Lee Yang, Andy Feng, Peter Cnudde, Mihajlo Grbovic, Nemanja Djuric, Vladan Radosavljevic, Gavin Owens |
CIKM | 7 |
| 2016 | Scalable Semantic Matching of Queries to Ads in Sponsored Search AdvertisingabstractSponsored search represents a major source of revenue for web search engines. The advertising model brings a unique possibility for advertisers to target direct user intent communicated through a search query, usually done by displaying their ads alongside organic search results for queries deemed relevant to their products or services. However, due to a large number of unique queries, it is particularly challenging for advertisers to identify all relevant queries. For this reason search engines often provide a service of advanced matching, which automatically finds additional relevant queries for advertisers to bid on. We present a novel advance match approach based on the idea of semantic embeddings of queries and ads. The embeddings were learned using a large data set of user search sessions, consisting of search queries, clicked ads and search links, while utilizing contextual information such as dwell time and skipped ads. To address the large-scale nature of our problem, both in terms of data and vocabulary size, we propose a novel distributed algorithm for training of the embeddings. Finally, we present an approach for overcoming a cold-start problem associated with new ads and queries. We report results of editorial evaluation and online tests on actual search traffic. The results show that our approach significantly outperforms baselines in terms of relevance, coverage and incremental revenue. Lastly, as part of this study, we open sourced query embeddings that can be used to advance the field. Mihajlo Grbovic, Nemanja Djuric, Vladan Radosavljevic, Fabrizio Silvestri, Ricardo Baeza-Yates, Andrew Feng, Erik Ordentlich, Lee Yang, Gavin Owens |
SIGIR | 3 |
| 2016 | TargetAd2016: 2nd International Workshop on Ad Targeting at ScaleabstractThe 2nd International Workshop on Ad Targeting at Scale will be held in San Francisco, California, USA on February 22nd, 2016, co-located with the 9th ACM International Conference on Web Search and Data Mining (WSDM). The main objective of the workshop is to address the challenges of ad targeting in web-scale settings. The workshop brings together interdisciplinary researchers in computational advertising, recommender systems, personalization, and related areas, to share, exchange, learn, and develop preliminary results, new concepts, ideas, principles, and methodologies on applying data mining technologies to ad targeting. We have constructed an exciting program of eight refereed papers and several invited talks that will help us better understand the future of ad targeting. Mihajlo Grbovic, Nemanja Djuric, Vladan Radosavljevic |
WSDM | 3 |
| 2016 | Portrait of an Online Shopper: Understanding and Predicting Consumer BehaviorabstractConsumer spending accounts for a large fraction of economic footprint of modern countries. Increasingly, consumer activity is moving to the web, where digital receipts of online purchases provide valuable data sources detailing consumer behavior. We consider such data extracted from emails and combined with with consumers' demographic information, which we use to characterize, model, and predict purchasing behavior. We analyze such behavior of consumers in different age and gender groups, and find interesting, actionable patterns that can be used to improve ad targeting systems. For example, we found that the amount of money spent on online purchases grows sharply with age, peaking in the late 30s, while shoppers from wealthy areas tend to purchase more expensive items and buy them more frequently. Furthermore, we look at the influence of social connections on purchasing habits, as well as at the temporal dynamics of online shopping where we discovered daily and weekly behavioral patterns. Finally, we build a model to predict when shoppers are most likely to make a purchase and how much will they spend, showing improvement over baseline approaches. The presented results paint a clear picture of a modern online shopper, and allow better understanding of consumer behavior that can help improve marketing efforts and make shopping more pleasant and efficient experience for online customers. Farshad Kooti, Kristina Lerman, Luca Maria Aiello, Mihajlo Grbovic, Nemanja Djuric, Vladan Radosavljevic |
WSDM | 6 |
| 2015 | Gender and Interest Targeting for Sponsored Post Advertising at TumblrabstractAs one of the leading platforms for creative content, Tumblr offers advertisers a unique way of creating brand identity. Advertisers can tell their story through images, animation, text, music, video, and more, and can promote that content by sponsoring it to appear as an advertisement in the users' live feeds. In this paper, we present a framework that enabled two of the key targeted advertising components for Tumblr, gender and interest targeting. We describe the main challenges encountered during the development of the framework, which include the creation of a ground truth for training gender prediction models, as well as mapping Tumblr content to a predefined interest taxonomy. For purposes of inferring user interests, we propose a novel semi-supervised neural language model for categorization of Tumblr content (i.e., post tags and post keywords). The model was trained on a large-scale data set consisting of $6.8$ billion user posts, with a very limited amount of categorized keywords, and was shown to have superior performance over the baseline approaches. We successfully deployed gender and interest targeting capability in Yahoo production systems, delivering inference for users that covers more than 90% of daily activities on Tumblr. Online performance results indicate advantages of the proposed approach, where we observed 20% increase in user engagement with sponsored posts in comparison to untargeted campaigns. Mihajlo Grbovic, Vladan Radosavljevic, Nemanja Djuric, Narayan L. Bhamidipati, Ananth Nagarajan |
KDD | 2 |
| 2015 | E-commerce in Your Inbox: Product Recommendations at ScaleabstractIn recent years online advertising has become increasingly ubiquitous and effective. Advertisements shown to visitors fund sites and apps that publish digital content, manage social networks, and operate e-mail services. Given such large variety of internet resources, determining an appropriate type of advertising for a given platform has become critical to financial success. Native advertisements, namely ads that are similar in look and feel to content, have had great success in news and social feeds. However, to date there has not been a winning formula for ads in e-mail clients. In this paper we describe a system that leverages user purchase history determined from e-mail receipts to deliver highly personalized product ads to Yahoo Mail users. We propose to use a novel neural language-based algorithm specifically tailored for delivering effective product recommendations, which was evaluated against baselines that included showing popular products and products predicted based on co-occurrence. We conducted rigorous offline testing using a large-scale product purchase data set, covering purchases of more than 29 million users from 172 e-commerce websites. Ads in the form of product recommendations were successfully tested on online traffic, where we observed a steady 9% lift in click-through rates over other ad formats in mail, as well as comparable lift in conversion rates. Following successful tests, the system was launched into production during the holiday season of 2014. Mihajlo Grbovic, Vladan Radosavljevic, Nemanja Djuric, Narayan L. Bhamidipati, Jaikit Savla, Varun Bhagwan, Doug Sharp |
KDD | 2 |
| 2015 | Context- and Content-aware Embeddings for Query Rewriting in Sponsored SearchabstractSearch engines represent one of the most popular web services, visited by more than 85% of internet users on a daily basis. Advertisers are interested in making use of this vast business potential, as very clear intent signal communicated through the issued query allows effective targeting of users. This idea is embodied in a sponsored search model, where each advertiser maintains a list of keywords they deem indicative of increased user response rate with regards to their business. According to this targeting model, when a query is issued all advertisers with a matching keyword are entered into an auction according to the amount they bid for the query, and the winner gets to show their ad. One of the main challenges is the fact that a query may not match many keywords, resulting in lower auction value, lower ad quality, and lost revenue for advertisers and publishers. Possible solution is to expand a query into a set of related queries and use them to increase the number of matched ads, called query rewriting. To this end, we propose rewriting method based on a novel query embedding algorithm, which jointly models query content as well as its context within a search session. As a result, queries with similar content and context are mapped into vectors close in the embedding space, which allows expansion of a query via simple K-nearest neighbor search in the projected space. The method was trained on more than 12 billion sessions, one of the largest corpuses reported thus far, and evaluated on both public TREC data set and in-house sponsored search data set. The results show the proposed approach significantly outperformed existing state-of-the-art, strongly indicating its benefits and the monetization potential. Mihajlo Grbovic, Nemanja Djuric, Vladan Radosavljevic, Fabrizio Silvestri, Narayan L. Bhamidipati |
SIGIR | 3 |
| 2015 | Hierarchical Neural Language Models for Joint Representation of Streaming Documents and their ContentabstractWe consider the problem of learning distributed representations for documents in data streams. The documents are represented as low-dimensional vectors and are jointly learned with distributed vector representations of word tokens using a hierarchical framework with two embedded neural language models. In particular, we exploit the context of documents in streams and use one of the language models to model the document sequences, and the other to model word sequences within them. The models learn continuous vector representations for both word tokens and documents such that semantically similar documents and words are close in a common vector space. We discuss extensions to our model, which can be applied to personalized recommendation and social relationship mining by adding further user layers to the hierarchy, thus learning user-specific vectors to represent individual preferences. We validated the learned representations on a public movie rating data set from MovieLens, as well as on a large-scale Yahoo News data comprising three months of user activity logs collected on Yahoo servers. The results indicate that the proposed model can learn useful representations of both documents and word tokens, outperforming the current state-of-the-art by a large margin. Nemanja Djuric, Vladan Radosavljevic, Mihajlo Grbovic, Narayan L. Bhamidipati |
WWW | 3 |
| 2014 | Hidden Conditional Random Fields with Deep User Embeddings for Ad TargetingabstractEstimating a user's propensity to click on a display ad or purchase a particular item is a critical task in targeted advertising, a burgeoning online industry worth billions of dollars. Better and more accurate estimation methods result in improved online user experience, as only relevant and interesting ads are shown, and may also lead to large benefits for advertisers, as targeted users are more likely to click or make a purchase. In this paper we address this important problem, and propose an approach for improved estimation of ad click or conversion probability based on a sequence of user's online actions, modeled using Hidden Conditional Random Fields (HCRF) model. In addition, in order to address the sparsity issue at the input side of the HCRF model, we propose to learn distributed, low-dimensional representations of user actions through a directed skip-gram, a neural architecture suitable for sequential data. Experimental results on a real-world data set comprising thousands of user sessions collected at Yahoo servers clearly indicate the benefits and the potential of the proposed approach, which outperformed competing state-of-the-art algorithms and obtained significant improvements in terms of retrieval measures. Nemanja Djuric, Vladan Radosavljevic, Mihajlo Grbovic, Narayan L. Bhamidipati |
ICDM | 2 |
| 2014 | Utilizing temporal patterns for estimating uncertainty in interpretable early decision makingabstractEarly classification of time series is prevalent in many time-sensitive applications such as, but not limited to, early warning of disease outcome and early warning of crisis in stock market. \textcolor{black}{ For example,} early diagnosis allows physicians to design appropriate therapeutic strategies at early stages of diseases. However, practical adaptation of early classification of time series requires an easy to understand explanation (interpretability) and a measure of confidence of the prediction results (uncertainty estimates). These two aspects were not jointly addressed in previous time series early classification studies, such that a difficult choice of selecting one of these aspects is required. In this study, we propose a simple and yet effective method to provide uncertainty estimates for an interpretable early classification method. The question we address here is "how to provide estimates of uncertainty in regard to interpretable early prediction." In our extensive evaluation on twenty time series datasets we showed that the proposed method has several advantages over the state-of-the-art method that provides reliability estimates in early classification. Namely, the proposed method is more effective than the state-of-the-art method, is simple to implement, and provides interpretable results. Mohamed F. Ghalwash, Vladan Radosavljevic, Zoran Obradovic |
KDD | 2 |
| 2014 | Neural Gaussian Conditional Random Fields
Vladan Radosavljevic, Slobodan Vucetic, Zoran Obradovic |
ECML/PKDD (2) | 1 |
| 2013 | Which links should I use?: a variogram-based selection of relationship measures for prediction of node attributes in temporal multigraphsabstractWhen faced with the task of forming predictions for nodes in a social network, it can be quite difficult to decide which of the available connections among nodes should be used for the best results. This problem is further exacerbated when temporal information is available, prompting the question of whether this information should be aggregated or not, and if not, which portions of it should be used. With this challenge in mind, we propose a novel utilization of variograms for selecting potentially useful relationship types, whose merits are then evaluated using a Gaussian Conditional Random Field model for node attribute prediction of temporal social networks with a multigraph structure. Our flexible model allows for measuring many kinds of relationships between nodes in the network that evolve over time, as well as using those relationships to augment the outputs of various unstructured predictors to further improve performance. The experimental results exhibit the effectiveness of using particular relationships to boost performance of unstructured predictors, show that using other relationships could actually impede performance, and also indicate that while variograms alone are not necessarily sufficient to identify a useful relationship, they greatly help in removing obviously useless measures, and can be combined with intuition to identify the optimal relationships. Alexey Uversky, Dusan Ramljak, Vladan Radosavljevic, Kosta Ristovski, Zoran Obradovic |
ASONAM | 3 |
| 2013 | Extraction of Interpretable Multivariate Patterns for Early DiagnosticsabstractLeveraging temporal observations to predict a patient's health state at a future period is a very challenging task. Providing such a prediction early and accurately allows for designing a more successful treatment that starts before a disease completely develops. Information for this kind of early diagnosis could be extracted by use of temporal data mining methods for handling complex multivariate time series. However, physicians usually prefer to use interpretable models that can be easily explained, rather than relying on more complex black-box approaches. In this study, a temporal data mining method is proposed for extracting interpretable patterns from multivariate time series data, which can be used to assist in providing interpretable early diagnosis. The problem is formulated as an optimization based binary classification task addressed in three steps. First, the time series data is transformed into a binary matrix representation suitable for application of classification methods. Second, a novel convex-concave optimization problem is defined to extract multivariate patterns from the constructed binary matrix. Then, a mixed integer discrete optimization formulation is provided to reduce the dimensionality and extract interpretable multivariate patterns. Finally, those interpretable multivariate patterns are used for early classification in challenging clinical applications. In the conducted experiments on two human viral infection datasets and a larger myocardial infarction dataset, the proposed method was more accurate and provided classifications earlier than three alternative state-of-the-art methods. Mohamed F. Ghalwash, Vladan Radosavljevic, Zoran Obradovic |
ICDM | 2 |
| 2008 | Spatio-Temporal Partitioning for Improving Aerosol Prediction AccuracyabstractIn supervised learning, on data collected over space and time, different relationships can be found over different spatio-temporal regions. In such situations, an appropriate spatio-temporal data partitioning followed by building specialized predictors could often achieve higher overall prediction accuracy than when learning a single predictor on all the data. In practice, partitions are typically decided based on prior knowledge. As an alternative to domain-based partitioning, we propose a method that automatically discovers a spatio-temporal partitioning through the competition of regression models. The method is evaluated on a challenging problem using satellite observations to predict Aerosol Optical Depth (AOD), which represents the amount of depletion that a beam of radiation undergoes as it passes through the atmosphere. Our experiments used more than 20,000 labeled data points collected during 3 years from more than 100 sites worldwide. Our partitioning-based approach was compared to the recently developed operational AOD prediction algorithm, called C5, which uses domain knowledge for spatio-temporal partitioning of the Earth and implements a region-specific deterministic predictor that utilizes forward simulations from the postulated physical models. Data partitioning used in C5 divides the world into three spatio-temporal regions that differ based on the location and the time of the year as decided by domain experts. The results showed that a neural network predictor trained on all the data has accuracy comparable to C5. When specialized neural network predictors were learned on C5-based partitions, the overall prediction accuracy was not improved. On the other hand, our competition-based spatio-temporal data partitioning approach resulted in large accuracy improvements. The most accurate results were obtained when (1) the data from each of the sites were split into two temporal subsets, one for winter-spring months and another for summer-fall months; and (2) two neural network predictors were competing for each of the identified spatio-temporal subsets. Vladan Radosavljevic, Slobodan Vucetic, Zoran Obradovic |
SDM | 1 |
| 2008 | Aerosol Optical Depth Prediction from Satellite Obsercations by Multiple Instance RegressionabstractAerosols are small airborne particles that both reflect and absorb incoming solar radiation and whose effect on the Earth's radiation budget is one of the biggest challenges of current climate research. To help address this challenge, numerous satellite sensors are employed to achieve global-scale monitoring of aerosols. Given the satellite measurements, the common objective is prediction of Aerosol Optical Depth (AOD). An important property of AOD is its low spatial variability on a scale of tens of kilometers. On the other hand, satellite sensors gather information in the form of multi-spectral images with high spatial resolution where pixels could be as small as a few hundred meters. Given an accurate ground-based AOD measurement over a specific location and time, all the pixels in the vicinity can be assumed to have the same AOD. If we treat satellite measurement at a single pixel as an instance, all pixels from the neighborhood can be considered as a bag of instances labeled with the same AOD. Given a number of bags obtained at numerous locations and at different times we can treat the problem of AOD prediction from satellite attributes as Multiple Instance Regression (MIR). An important challenge is that because of rapidly changing surface properties attribute values of pixels from a bag can vary a lot. This study evaluated several MIR approaches on several synthetic data sets and on a data set consisting of 800 labeled bags, each containing hundreds of pixel instances observed over the Continental U.S. by the MISR satellite instrument. The results indicate that the most successful MIR approach consists of an iterative procedure that detects and discards outlying instances and trains a predictor on the remaining ones. Vladan Radosavljevic, Zoran Obradovic |
SDM | 2 |