Yu Sun 0021

dblp:62/3689-21 · DBLP profile ↗
← Back
24ranked-venue papers
8as first author
4since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 15 · 8 first-author · 2 since 2021Artificial intelligence and machine learning · 14 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-authorHuman-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Question answering and dialogue systems · 32% Efficient and distributed learning · 23% Reinforcement learning · 17%
Databases, data mining, and information retrieval
11 papers
Recommender systems · 67% Information retrieval · 14% Spatial and temporal data management · 10%

Topics — the 30 heaviest of 38, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Recommender systems › collaborative filtering
rating prediction
0.922021
Combating Selection Biases in Recommender Systems with a Few Unbiased Ratings · WSDM 2021
Doubly Robust Joint Learning for Recommendation on Data Missing Not at Random · ICML 2019
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
0.822021
Adversarial Distillation for Learning with Privileged Provisions · IEEE Trans. Pattern Anal. Mach. Intell. 2021
KDGAN: Knowledge Distillation with Generative Adversarial Networks · NeurIPS 2018
Recommender systems
context-aware recommendation
0.832017
Collaborative Intent Prediction with Real-Time Contextual Data · ACM Trans. Inf. Syst. 2017
Collaborative Nowcasting for Contextual Recommendation · WWW 2016
Contextual Intent Tracking for Personal Assistants · KDD 2016
Machine learning › Efficient and distributed learning › distillation
adversarial distillation
0.512021
Adversarial Distillation for Learning with Privileged Provisions · IEEE Trans. Pattern Anal. Mach. Intell. 2021
Recommender systems
collaborative filtering
0.512021
Combating Selection Biases in Recommender Systems with a Few Unbiased Ratings · WSDM 2021
Recommender systems › recommender system evaluation › off-policy evaluation
inverse propensity scoring
0.512021
Combating Selection Biases in Recommender Systems with a Few Unbiased Ratings · WSDM 2021
Recommender systems › interactive recommendation
proactive recommendation
0.522016
Collaborative Nowcasting for Contextual Recommendation · WWW 2016
Contextual Intent Tracking for Personal Assistants · KDD 2016
Recommender systems › debiased recommendation
selection bias
0.512021
Combating Selection Biases in Recommender Systems with a Few Unbiased Ratings · WSDM 2021
Robotics › Autonomous driving › planning for self-driving vehicles
action decision
0.412020
MALA: Cross-Domain Dialogue Generation with Action Learning · AAAI 2020
Natural language and speech › Question answering and dialogue systems
dialogue generation
0.412020
MALA: Cross-Domain Dialogue Generation with Action Learning · AAAI 2020
Natural language and speech › Question answering and dialogue systems › dialogue management
dialogue planning
0.412020
MALA: Cross-Domain Dialogue Generation with Action Learning · AAAI 2020
Natural language and speech › Question answering and dialogue systems › task-oriented dialogue
dialogue policy learning
0.412020
Semi-Supervised Dialogue Policy Learning via Stochastic Reward Estimation · ACL 2020
Natural language and speech › Language models and text generation › preference optimization
direct preference optimization
0.412020
Semi-Supervised Dialogue Policy Learning via Stochastic Reward Estimation · ACL 2020
Machine learning › Reinforcement learning › function approximation › representation learning for reinforcement learning
latent action learning
0.412020
MALA: Cross-Domain Dialogue Generation with Action Learning · AAAI 2020
Machine learning › Reinforcement learning
reward learning
0.412020
Semi-Supervised Dialogue Policy Learning via Stochastic Reward Estimation · ACL 2020
Natural language and speech › Question answering and dialogue systems
task-oriented dialogue
0.412020
Semi-Supervised Dialogue Policy Learning via Stochastic Reward Estimation · ACL 2020
Natural language and speech › Question answering and dialogue systems › dialogue generation
task-oriented dialogue generation
0.412020
MALA: Cross-Domain Dialogue Generation with Action Learning · AAAI 2020
Spatial and temporal data management
spatial keyword query
0.412020
Efficient processing of moving collective spatial keyword queries · VLDB J. 2020
Machine learning › Trustworthy machine learning
debiasing
0.412019
Doubly Robust Joint Learning for Recommendation on Data Missing Not at Random · ICML 2019
Machine learning › Reinforcement learning › off-policy evaluation
doubly robust estimation
0.412019
Doubly Robust Joint Learning for Recommendation on Data Missing Not at Random · ICML 2019
Recommender systems › debiased recommendation
missing not at random
0.412019
Doubly Robust Joint Learning for Recommendation on Data Missing Not at Random · ICML 2019
Machine learning › Generative modeling
generative adversarial network
0.312018
KDGAN: Knowledge Distillation with Generative Adversarial Networks · NeurIPS 2018
Machine learning › Efficient and distributed learning
model compression
0.312018
KDGAN: Knowledge Distillation with Generative Adversarial Networks · NeurIPS 2018
Data mining › predictive modeling › forecasting
sales prediction
0.312017
App Download Forecasting: An Evolutionary Hierarchical Competition Approach · IJCAI 2017
Data mining › time series analysis
time series forecasting
0.312017
App Download Forecasting: An Evolutionary Hierarchical Competition Approach · IJCAI 2017
Recommender systems › user modeling › user behavior prediction
app usage prediction
0.212016
A contextual collaborative approach for app usage forecasting · UbiComp 2016
Spatial and temporal data management
reverse nearest neighbor
0.212016
Reverse nearest neighbor heat maps: A tool for influence exploration · ICDE 2016
Machine learning › Trustworthy machine learning
robustness
0.112021
Combating Selection Biases in Recommender Systems with a Few Unbiased Ratings · WSDM 2021
Machine learning › Trustworthy machine learning › dataset bias
selection bias
0.112021
Combating Selection Biases in Recommender Systems with a Few Unbiased Ratings · WSDM 2021
Natural language and speech › Question answering and dialogue systems › dialogue generation
dialogue response generation
0.112020
MALA: Cross-Domain Dialogue Generation with Action Learning · AAAI 2020

Methods — techniques the papers use, named apart from their topics

propensity weighting · 1.8unbiased performance estimation · 1.0content moderation · 0.9doubly robust estimation · 0.8minimax game · 0.5gumbel-softmax · 0.5variational inference · 0.4reward learning · 0.4dynamics model · 0.4contrastive learning · 0.4action embedding · 0.4error imputation · 0.4sequential modeling · 0.3nowcasting model · 0.3multi-scale time series forecasting · 0.3evolutionary hierarchical competition model · 0.3kalman filter · 0.2
YearPublicationVenuePosition
2025 PhysFFTFormer: A Frequency Domain-based Vision Transformer for Efficient Remote Physiological Measurement
abstract
Remote Photoplethysmography (rPPG) is a non-contact technique for extracting physiological signals from facial videos. Recently, Transformer-based architectures have exhibited remarkable performance in rPPG estimation, owing to their excellent long-range spatial-temporal modeling capacities. However, challenges persist in applying Transformer-based rPPG methods, where quadratic computational costs and inadequate feature modeling diversity remain formidable. To address these challenges, we leverage the power of the Frequency Domain-based Vision Transformer and propose an end-to-end model PhysFFTFormer. Specifically, by integrating customized Frequency-Domain Spatiotemporal Self-Attention and Frequency-Domain Discriminative Feed-Forward modules, PhysFFTFormer efficiently captures spatial-temporal dependencies with reduced computational complexity. To extract rich and diverse feature representations, we design a dual-pathway architecture to utilize both raw and differential video frames. Furthermore, the Frequency-Domain Spatiotemporal Cross-Attention module is introduced to enhance information exchange and enable feature complementation between the two pathways. Extensive experiments on multiple benchmark datasets demonstrate PhysFFTFormer’s state-of-the-art performance, robustness, and potential for real-world non-contact health monitoring.
Sirui Zhao, Tong Xu 0001, Yu Sun 0021, Hao Wang 0076, Suojuan Zhang, Enhong Chen
ICME4
2021 Combating Selection Biases in Recommender Systems with a Few Unbiased Ratings
abstract
Recommendation datasets are prone to selection biases due to self-selection behavior of users and item selection process of systems. This makes explicitly combating selection biases an essential problem in training recommender systems. Most previous studies assume no unbiased data available for training. We relax this assumption and assume that a small subset of training data is unbiased. Then, we propose a novel objective that utilizes the unbiased data to adaptively assign propensity weights to biased training ratings. This objective, combined with unbiased performance estimators, alleviates the effects of selection biases on the training of recommender systems. To optimize the objective, we propose an efficient algorithm that minimizes the variance of propensity estimates for better generalized recommender systems. Extensive experiments on two real-world datasets confirm the advantages of our approach in significantly reducing both the error of rating prediction and the variance of propensity estimation.
Xiaojie Wang 0003, Rui Zhang 0003, Yu Sun 0021, Jianzhong Qi 0001
WSDM3
2021 Integrity 2021: Integrity in Social Networks and Media
abstract
The second Workshop on Integrity in Social Networks and Media is held in conjunction with the 14th ACM Conference on Web Search and Data Mining (WSDM) in Jerusalem, Israel. The goal of the workshop is to bring together researchers and practitioners to discuss content and interaction integrity challenges in social networks and social media platforms.
Lluís Garcia Pueyo, Anand Bhaskar, Roelof van Zwol, Timos K. Sellis, Gireeja Ranade, Prathyusha Senthil Kumar, Yu Sun 0021, Joy Zhang
WSDM7
2021 Adversarial Distillation for Learning with Privileged Provisions
abstract
Knowledge distillation aims to train a student (model) for accurate inference in a resource-constrained environment. Traditionally, the student is trained by a high-capacity teacher (model) whose training is resource-intensive. The student trained this way is suboptimal because it is difficult to learn the real data distribution from the teacher. To address this issue, we propose to train the student against a discriminator in a minimax game. Such a minimax game has an issue that it can take an excessively long time for the training to converge. To address this issue, we propose adversarial distillation consisting of a student, a teacher, and a discriminator. The discriminator is now a multi-class classifier that distinguishes among the real data, the student, and the teacher. The student and the teacher aim to fool the discriminator via adversarial losses, while they learn from each other via distillation losses. By optimizing the adversarial and the distillation losses simultaneously, the student and the teacher can learn the real data distribution. To accelerate the training, we propose to obtain low-variance gradient updates from the discriminator using a Gumbel-Softmax trick. We conduct extensive experiments to demonstrate the superiority of the proposed adversarial distillation under both accuracy and training speed.
Xiaojie Wang 0003, Rui Zhang 0003, Yu Sun 0021, Jianzhong Qi 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2020 MALA: Cross-Domain Dialogue Generation with Action Learning
abstract
Response generation for task-oriented dialogues involves two basic components: dialogue planning and surface realization. These two components, however, have a discrepancy in their objectives, i.e., task completion and language quality. To deal with such discrepancy, conditioned response generation has been introduced where the generation process is factorized into action decision and language generation via explicit action representations. To obtain action representations, recent studies learn latent actions in an unsupervised manner based on the utterance lexical similarity. Such an action learning approach is prone to diversities of language surfaces, which may impinge task completion and language quality. To address this issue, we propose multi-stage adaptive latent action learning (MALA) that learns semantic latent actions by distinguishing the effects of utterances on dialogue progress. We model the utterance effect using the transition of dialogue states caused by the utterance and develop a semantic similarity measurement that estimates whether utterances have similar effects. For learning semantic actions on domains without dialogue states, MALA extends the semantic similarity measurement across domains progressively, i.e., from aligning shared actions to learning domain-specific actions. Experiments using multi-domain datasets, SMD and MultiWOZ, show that our proposed model achieves consistent improvements over the baselines models in terms of both task completion and language quality.
Xinting Huang, Jianzhong Qi 0001, Yu Sun 0021, Rui Zhang 0003
AAAI3
2020 Semi-Supervised Dialogue Policy Learning via Stochastic Reward Estimation
abstract
Dialogue policy optimization often obtains feedback until task completion in taskoriented dialogue systems.This is insufficient for training intermediate dialogue turns since supervision signals (or rewards) are only provided at the end of dialogues.To address this issue, reward learning has been introduced to learn from state-action pairs of an optimal policy to provide turn-by-turn rewards.This approach requires complete state-action annotations of human-to-human dialogues (i.e., expert demonstrations), which is labor intensive.To overcome this limitation, we propose a novel reward learning approach for semisupervised policy learning.The proposed approach learns a dynamics model as the reward function which models dialogue progress (i.e., state-action sequences) based on expert demonstrations, either with or without annotations.The dynamics model computes rewards by predicting whether the dialogue progress is consistent with expert demonstrations.We further propose to learn action embeddings for a better generalization of the reward function.The proposed approach outperforms competitive policy learning baselines on MultiWOZ, a benchmark multi-domain dataset.
Xinting Huang, Jianzhong Qi 0001, Yu Sun 0021, Rui Zhang 0003
ACL3
2020 Relevance Ranking for Real-Time Tweet Search
abstract
Relevance ranking is a key component of many search engines, including the Tweet search engine at Twitter. Users often use Tweet search to discover live discussions and different voices on trending topics or recent events. Tweet search is thus unique due to its focus on real-time content, where both the retrieved content and queries change drastically on an hourly basis. Another important property of Tweet search is that its relevance ranking takes the social endorsements from other users into account, e.g., "likes" and "retweets", which is different from mainly relying on clicks as implicit feedback. The relevance ranking of Tweet search is also subject to strict latency constraints, because every second, a large amount of Tweets are posted and indexed, while tens of thousands of queries are issued to search posted Tweets. Considering the above properties and constraints, we present a relevance ranking system for Tweet search addressing all these challenges at Twitter. We first discuss the formation of the relevance ranking pipeline, which consists of a series of ranking models. We then present the methodology for training the models and the various groups of features we use, including real-time and personalized features. We also investigate approaches of achieving unbiased model training and building up automatic online tuning of system parameters. Experiments using online A/B testing demonstrate the effectiveness of the proposed approaches and we have deployed the proposed relevance ranking system in production for more than three years.
Yu Sun 0021, Juan Caicedo Carvajal, Jinliang Fan, Bhargav Mangipudi, Lisa Huang, Yatharth Saraf
CIKM2
2020 Integrity 2020: Integrity in Social Networks and Media
abstract
The first Workshop on Integrity in Social Networks and Media is held in conjunction with the 13th ACM Conference on Web Search and Data Mining (WSDM) in Houston, Texas, USA. The goal of the workshop is to bring together researchers and practitioners to discuss content and interaction integrity challenges in social networks and social media platforms.
Lluís Garcia Pueyo, Anand Bhaskar, Panayiotis Tsaparas, Aristides Gionis, Tina Eliassi-Rad, Maria Daltayanni, Yu Sun 0021, Panagiotis Papadimitriou 0002
WSDM7
2020 Efficient processing of moving collective spatial keyword queries
Hongfei Xu, Yu Gu 0002, Yu Sun 0021, Jianzhong Qi 0001, Ge Yu 0001, Rui Zhang 0003
VLDB J.3
2019 Doubly Robust Joint Learning for Recommendation on Data Missing Not at Random
abstract
In recommender systems, usually the ratings of a user to most items are missing and a critical problem is that the missing ratings are often missing not at random (MNAR) in reality. It is widely acknowledged that MNAR ratings make it difficult to accurately predict the ratings and unbiasedly estimate the performance of rating prediction. Recent approaches use imputed errors to recover the prediction errors for missing ratings, or weight observed ratings with the propensities of being observed. These approaches can still be severely biased in performance estimation or suffer from the variance of the propensities. To overcome these limitations, we first propose an estimator that integrates the imputed errors and propensities in a doubly robust way to obtain unbiased performance estimation and alleviate the effect of the propensity variance. To achieve good performance guarantees, based on this estimator, we propose joint learning of rating prediction and error imputation, which outperforms the state-of-the-art approaches on four real-world datasets.
Xiaojie Wang 0003, Rui Zhang 0003, Yu Sun 0021, Jianzhong Qi 0001
ICML3
2019 CARL: Aggregated Search with Context-Aware Module Embedding Learning
abstract
Aggregated search aims to construct search result pages (SERPs) from blue-links and heterogeneous modules (such as news, images, and videos). Existing studies have largely ignored the correlations between blue-links and heterogeneous modules when selecting the heterogeneous modules to be presented. We observe that the top ranked blue-links, which we refer to as the context, can provide important information about query intent and helps identify the relevant heterogeneous modules. For example, informative terms like "streamed" and "recorded" in the context imply that a video module may better satisfy the query. To model and utilize the context information for aggregated search, we propose a model with context attention and representation learning (CARL). Our model applies a recurrent neural network with attention mechanism to encode the context, and incorporates the encoded context information into module embeddings. The context-aware module embeddings together with the ranking policy are jointly optimized under the Markov decision process (MDP) formulation. To achieve a more effective joint learning, we further propose an optimization function with self-supervision loss to provide auxiliary supervision signals. Experimental results based on two public datasets demonstrate the superiority of CARL over multiple baseline approaches, and confirm the effectiveness of the proposed optimization function in boosting the joint learning process.
Xinting Huang, Jianzhong Qi 0001, Yu Sun 0021, Rui Zhang 0003, Hai-Tao Zheng 0002
IJCNN3
2018 Learning Effective Embeddings for Machine Generated Emails with Applications to Email Category Prediction
abstract
Machine generated business-to-consumer (B2C) emails such as receipts, newsletters, and promotions constitute a large portion of users' inboxes today. These emails reflect the users' interests and often are sequentially correlated, e.g., users interested in relocating may receive a sequence of messages on housing, moving, job availability, etc. We aim to infer (and eventually serve) the users' future interests by predicting the categories of their future emails. There are many useful methods, such as recurrent neural networks, that can be applied for such predictions, but in all cases the key to better performance is an effective representation of emails and users. To this end, we propose a general framework for learning embeddings for emails and users, using as input only the sequence of B2C templates users receive and open. (A template is a B2C email stripped of all transient information related to specific users.) These learned embeddings allow us to identify both sequentially correlated emails and users with similar sequential interests. We can also use the learned embeddings either as input features or embedding initializers for email category prediction tasks. Extensive experiments with millions of fully anonymized B2C emails demonstrate that the learned embeddings can significantly improve the prediction accuracy for future email categories. We hope that this effective yet simple embedding learning framework will inspire new machine intelligence applications that will improve the users' email experience.
Yu Sun 0021, Lluís Garcia Pueyo, James B. Wendt, Marc Najork, Andrei Z. Broder
IEEE BigData1
2018 KDGAN: Knowledge Distillation with Generative Adversarial Networks
abstract
Knowledge distillation (KD) aims to train a lightweight classifier suitable to provide accurate inference with constrained resources in multi-label learning. Instead of directly consuming feature-label pairs, the classifier is trained by a teacher, i.e., a high-capacity model whose training may be resource-hungry. The accuracy of the classifier trained this way is usually suboptimal because it is difficult to learn the true data distribution from the teacher. An alternative method is to adversarially train the classifier against a discriminator in a two-player game akin to generative adversarial networks (GAN), which can ensure the classifier to learn the true data distribution at the equilibrium of this game. However, it may take excessively long time for such a two-player game to reach equilibrium due to high-variance gradient updates. To address these limitations, we propose a three-player game named KDGAN consisting of a classifier, a teacher, and a discriminator. The classifier and the teacher learn from each other via distillation losses and are adversarially trained against the discriminator via adversarial losses. By simultaneously optimizing the distillation and adversarial losses, the classifier will learn the true data distribution at the equilibrium. We approximate the discrete distribution learned by the classifier (or the teacher) with a concrete distribution. From the concrete distribution, we generate continuous samples to obtain low-variance gradient updates, which speed up the training. Extensive experiments using real datasets confirm the superiority of KDGAN in both accuracy and training speed.
Xiaojie Wang 0003, Rui Zhang 0003, Yu Sun 0021, Jianzhong Qi 0001
NeurIPS3
2018 A Joint Optimization Approach for Personalized Recommendation Diversification
Xiaojie Wang 0003, Jianzhong Qi 0001, Kotagiri Ramamohanarao, Yu Sun 0021, Bo Li 0026, Rui Zhang 0003
PAKDD (3)4
2018 Context-Uncertainty-Aware Chatbot Action Selection via Parameterized Auxiliary Reinforcement Learning
Chuandong Yin, Rui Zhang 0003, Jianzhong Qi 0001, Yu Sun 0021, Tenglun Tan
PAKDD (1)4
2017 App Download Forecasting: An Evolutionary Hierarchical Competition Approach
abstract
Product sales forecasting enables comprehensive understanding of products' future development, making it of particular interest for companies to improve their business, for investors to measure the values of firms, and for users to capture the trends of a market. Recent studies show that the complex competition interactions among products directly influence products' future development. However, most existing approaches fail to model the evolutionary competition among products and lack the capability to organically reflect multi-level competition analysis in sales forecasting. To address these problems, we propose the Evolutionary Hierarchical Competition Model (EHCM), which effectively considers the time-evolving multi-level competition among products. The EHCM model systematically integrates hierarchical competition analysis with multi-scale time series forecasting. Extensive experiments using a real-world app download dataset show that EHCM outperforms state-of-the-art methods in various forecasting granularities.
Yingzi Wang, Nicholas Jing Yuan, Yu Sun 0021, Chuan Qin 0002, Xing Xie 0001
IJCAI3
2017 Collaborative Intent Prediction with Real-Time Contextual Data
abstract
Intelligent personal assistants on mobile devices such as Apple’s Siri and Microsoft Cortana are increasingly important. Instead of passively reacting to queries, they provide users with brand new proactive experiences that aim to offer the right information at the right time. It is, therefore, crucial for personal assistants to understand users’ intent, that is, what information users need now. Intent is closely related to context. Various contextual signals, including spatio-temporal information and users’ activities, can signify users’ intent. It is, however, challenging to model the correlation between intent and context. Intent and context are highly dynamic and often sequentially correlated. Contextual signals are usually sparse, heterogeneous, and not simultaneously available. We propose an innovative collaborative nowcasting model to jointly address all these issues. The model effectively addresses the complex sequential and concurring correlation between context and intent and recognizes users’ real-time intent with continuously arrived contextual signals. We extensively evaluate the proposed model with real-world data sets from a commercial personal assistant. The results validate the effectiveness the proposed model, and demonstrate its capability of handling the real-time flow of contextual signals. The studied problem and model also provide inspiring implications for new paradigms of recommendation on mobile intelligent devices.
Yu Sun 0021, Nicholas Jing Yuan, Xing Xie 0001, Kieran McDonald, Rui Zhang 0003
ACM Trans. Inf. Syst.1
2016 A contextual collaborative approach for app usage forecasting
abstract
Fine-grained long-term forecasting enables many emerging recommendation applications such as forecasting the usage amounts of various apps to guide future investments, and forecasting users' seasonal demands for a certain commodity to find potential repeat buyers. For these applications, there often exists certain homogeneity in terms of similar users and items (e.g., apps), which also correlates with various contexts like users' spatial movements and physical environments. Most existing works only focus on predicting the upcoming situation such as the next used app or next online purchase, without considering the long-term temporal co-evolution of items and contexts and the homogeneity among all dimensions. In this paper, we propose a contextual collaborative forecasting (CCF) model to address the above issues. The model integrates contextual collaborative filtering with time series analysis, and simultaneously captures various components of temporal patterns, including trend, seasonality, and stationarity. The approach models the temporal homogeneity of similar users, items, and contexts. We evaluate the model on a large real-world app usage dataset, which validates that CCF outperforms state-of-the-art methods in terms of both accuracy and efficiency for long-term app usage forecasting.
Yingzi Wang, Nicholas Jing Yuan, Yu Sun 0021, Xing Xie 0001, Qi Liu 0003, Enhong Chen
UbiComp3
2016 Reverse nearest neighbor heat maps: A tool for influence exploration
abstract
We study the problem of constructing a reverse nearest neighbor (RNN) heat map by finding the RNN set of every point in a two-dimensional space. Based on the RNN set of a point, we obtain a quantitative influence (i.e., heat) for the point. The heat map provides a global view on the influence distribution in the space, and hence supports exploratory analyses in many applications such as marketing and resource management. To construct such a heat map, we first reduce it to a problem called Region Coloring (RC), which divides the space into disjoint regions within which all the points have the same RNN set. We then propose a novel algorithm named CREST that efficiently solves the RC problem by labeling each region with the heat value of its containing points. In CREST, we propose innovative techniques to avoid processing expensive RNN queries and greatly reduce the number of region labeling operations. We perform detailed analyses on the complexity of CREST and lower bounds of the RC problem, and prove that CREST is asymptotically optimal in the worst case. Extensive experiments with both real and synthetic data sets demonstrate that CREST outperforms alternative algorithms by several orders of magnitude.
Yu Sun 0021, Rui Zhang 0003, Andy Yuan Xue, Jianzhong Qi 0001, Xiaoyong Du 0001
ICDE1
2016 Contextual Intent Tracking for Personal Assistants
abstract
A new paradigm of recommendation is emerging in intelligent personal assistants such as Apple's Siri, Google Now, and Microsoft Cortana, which recommends "the right information at the right time" and proactively helps you "get things done". This type of recommendation requires precisely tracking users' contemporaneous intent, i.e., what type of information (e.g., weather, stock prices) users currently intend to know, and what tasks (e.g., playing music, getting taxis) they intend to do. Users' intent is closely related to context, which includes both external environments such as time and location, and users' internal activities that can be sensed by personal assistants. The relationship between context and intent exhibits complicated co-occurring and sequential correlation, and contextual signals are also heterogeneous and sparse, which makes modeling the context intent relationship a challenging task. To solve the intent tracking problem, we propose the Kalman filter regularized PARAFAC2 (KP2) nowcasting model, which compactly represents the structure and co-movement of context and intent. The KP2 model utilizes collaborative capabilities among users, and learns for each user a personalized dynamic system that enables efficient nowcasting of users' intent. Extensive experiments using real-world data sets from a commercial personal assistant show that the KP2 model significantly outperforms various methods, and provides inspiring implications for deploying large-scale proactive recommendation systems in personal assistants.
Yu Sun 0021, Nicholas Jing Yuan, Yingzi Wang, Xing Xie 0001, Kieran McDonald, Rui Zhang 0003
KDD1
2016 Collaborative Nowcasting for Contextual Recommendation
abstract
Mobile digital assistants such as Microsoft Cortana and Google Now currently offer appealing proactive experiences to users, which aim to deliver the right information at the right time. To achieve this goal, it is crucial to precisely predict users' real-time intent. Intent is closely related to context, which includes not only the spatial-temporal information but also users' current activities that can be sensed by mobile devices. The relationship between intent and context is highly dynamic and exhibits chaotic sequential correlation. The context itself is often sparse and heterogeneous. The dynamics and co-movement among contextual signals are also elusive and complicated. Traditional recommendation models cannot directly apply to proactive experiences because they fail to tackle the above challenges. Inspired by the nowcasting practice in meteorology and macroeconomics, we propose an innovative collaborative nowcasting model to effectively resolve these challenges. The proposed model successfully addresses sparsity and heterogeneity of contextual signals. It also effectively models the convoluted correlation within contextual signals and between context and intent. Specifically, the model first extracts collaborative latent factors, which summarize shared temporal structural patterns in contextual signals, and then exploits the collaborative Kalman Filter to generate serially correlated personalized latent factors, which are utilized to monitor each user's real-time intent. Extensive experiments with real-world data sets from a commercial digital assistant demonstrate the effectiveness of the collaborative nowcasting model. The studied problem and model provide inspiring implications for new paradigms of recommendations on mobile intelligent devices.
Yu Sun 0021, Nicholas Jing Yuan, Xing Xie 0001, Kieran McDonald, Rui Zhang 0003
WWW1
2015 K-Nearest Neighbor Temporal Aggregate Queries
abstract
We study a new type of queries called the k-nearest neigh-bor temporal aggregate (kNNTA) query. Given a query point and a time interval, it returns the top-k locations that have the smallest weighted sums of (i) the spatial distance to the query point and (ii) a temporal aggregate on a cer-tain attribute over the time interval. For example, find a nearby club that has the largest number of people visiting in the last hour. This type of queries has emerging applica-tions in location-based social networks, location-based mo-bile advertising and social event recommendation. It is a great challenge to efficiently answer the query due to the highly dynamic nature and the large volume of the data and queries. To address this challenge, we propose an index named TAR-tree, which organizes locations by integrating the spatial and temporal aggregate information. We per-form a detailed analysis on the cost of processing kNNTA queries using the TAR-tree. The analysis shows that the TAR-tree results in much fewer node accesses than alterna-tives. Furthermore, we propose two enhancements for the kNNTA query: (i) an algorithm suggesting the least amount of weights to be adjusted to explore different query results and (ii) a collective processing scheme to share index traver-sal among a batch of queries. We conduct extensive exper-iments using real-world data sets. The results validate the accuracy of the cost analysis and show that the TAR-tree outperforms alternatives by up to ten times in node accesses. The results also show that the weight adjustment algorithm and collective processing scheme outperform their baselines by significant margins. 1.
Yu Sun 0021, Jianzhong Qi 0001, Yu Zheng 0004, Rui Zhang 0003
EDBT1
2012 Location selection for utility maximization with capacity constraints
abstract
Given a set of client locations, a set of facility locations where each facility has a service capacity, and the assumptions that: (i) a client seeks service from its nearest facility; (ii) a facility provides service to clients in the order of their proximity, we study the problem of selecting all possible locations such that setting up a new facility with a given capacity at these locations will maximize the number of served clients. This problem has wide applications in practice, such as setting up new distribution centers for online sales business and building additional base stations for mobile subscribers. We formulate the problem as location selection query for utility maximization. After applying three pruning rules to a baseline solution,we obtain an efficient algorithm to answer the query. Extensive experiments confirm the efficiency of our proposed algorithm.
Yu Sun 0021, Jin Huang 0003, Yueguo Chen, Rui Zhang 0003, Xiaoyong Du 0001
CIKM1
2012 Top-k Most Incremental Location Selection with Capacity Constraint
Yu Sun 0021, Jin Huang 0003, Yueguo Chen, Xiaoyong Du 0001, Rui Zhang 0003
WAIM1