VLDB 2026 Research / reviewers in the wild / expert
Xian Wu 0003
dblp:03/5595-3
· DBLP profile ↗
16ranked-venue papers
9as first author
2since 2021 · last 2022
0000-0003-0840-5857ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 16 · 9 first-author · 2 since 2021Artificial intelligence and machine learning · 10 · 5 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 first-authorTheory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Representation Learning on Variable Length and Incomplete Wearable-Sensory Time SeriesabstractThe prevalence of wearable sensors (e.g., smart wristband) is creating unprecedented opportunities to not only inform health and wellness states of individuals, but also assess and infer personal attributes, including demographic and personality attributes. However, the data captured from wearables, such as heart rate or number of steps, present two key challenges: (1) the time series is often of variable length and incomplete due to different data collection periods (e.g., wearing behavior varies by person); and (2) there is inter-individual variability to external factors like stress and environment. This article addresses these challenges and brings us closer to the potential of personalized insights about an individual, taking the leap from quantified self to qualified self. Specifically, HeartSpace proposed in this article learns embedding of the time-series data with variable length and missing values via the integration of a time-series encoding module and a pattern aggregation network. Additionally, HeartSpace implements a Siamese-triplet network to optimize representations by jointly capturing intra- and inter-series correlations during the embedding learning process. The empirical evaluation over two different real-world data presents significant performance gains over state-of-the-art baselines in a variety of applications, including user identification, personality prediction, demographics inference, job performance prediction, and sleep duration estimation. Xian Wu 0003, Chao Huang 0001, Pablo Robles-Granda, Nitesh V. Chawla |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2021 | motif2vec: Semantic-aware Representation Learning for Wearables' Time Series DataabstractThe proliferation of wearable sensors allows for the continuous collection of temporal characterization of an individual's physical activity and physiological data. This is enabling an unprecedented opportunity to delve into a deeper analysis of the underlying patterns of such temporal data and to infer attributes associated with health, behaviors, and well-being. However, there remain several challenges to fully discover both structural and temporal patterns (motifs) in these data streams and to leverage the semantic relationship among these motifs. These include: i) the temporal data of variable length and high resolution leads to the motifs of various sizes; ii) periodic occurrences and hierarchical overlaps of these motifs further challenge the modeling of their complex structural and semantic relations. We propose a semantic-aware unsupervised representation learning model, motif2vec, to learn the latent representation of time series data collected from wearable sensors. The motif2vec consists of three major components: 1) transforming the time series into a set of variable-length motif sequences; 2) formalizing random walks to construct the neighborhood of motifs and thus to extract structural and semantic relationship among motifs; 3) learning time series latent features to capture the motif neighborhood structure with a skip-gram model. Experiments on two real-world datasets, derived from two different wearables and population groups, show motif2vec outperforms six state-of-the-art benchmarks on various tasks. Suwen Lin, Xian Wu 0003, Nitesh V. Chawla |
DSAA | 2 |
| 2020 | Personalized Imputation on Wearable-Sensory Time Series via Knowledge TransferabstractThe analysis of wearable-sensory time series data (e.g., heart rate records) benefits many applications (e.g., activity recognition, disease diagnosis). However, sensor measurements usually contain missing values due to various factors (e.g., user behavior, lack of charging), which may degrade the performance of downstream analytical tasks (e.g., regression, prediction). Thus, time series imputation is desired, which is capable of making sensory time series complete. Existing time series imputation methods generally employ various deep neural network models (e.g., GRU and GAN) to fill missing values by leveraging temporal patterns extracted from the contextual observations. Despite their effectiveness, we argue that most existing models can only achieve sub-optimal imputation performance due to the fact that they are inherently limited in sharing only one single set of model parameters to perform imputation on all individuals. Relying on one set of parameters limits the expressiveness of the imputation model as such models are bound to fail in capturing various complex personal characteristics. Therefore, most existing models tend to achieve inferior imputation performance, especially when a long duration of missing values, i.e., a large gap, is observed in the time series data. To address the limitation, this work develops a new imputation framework--Personalized Wearable-Sensory Time Series Imputation framework (PTSI) to provide a fully personalized treatment for time series imputation via effective knowledge transfer. In particular, PTSI first leverages a meta-learning paradigm to learn a well-generalized initialization to facilitate the adaption process for each user. To make the time series imputation be reflective of an individual's unique characteristics, we further endow PTSI with the capability of learning personalized model parameters, which is achieved by designing a parameter initialization modulating component. Extensive experiments on real-world human heart rate datasets demonstrate that our PTSI framework outperforms various state-of-the-art methods by a large margin consistently. Xian Wu 0003, Stephen M. Mattingly, Shayan Mirjafari, Chao Huang 0001, Nitesh V. Chawla |
CIKM | 1 |
| 2020 | Filling Missing Values on Wearable-Sensory Time Series DataabstractMissing data points is a common problem associated with data collected from wearables. This problem is particularly compounded if different subjects have different aspects of missingness associated with them – that is varying degrees of compliance behavior of individuals (participants) with respect to wearables as well as personal changes in lifestyle and health impacting heart rate. Moreover, despite the varying degree of compliance behavior, the wearable in itself might have glitches that lead to observations being dropped. Thus, any missing value imputation in such data has to not only generalize to the wearable behavior but also to the participant behavior. In this paper, we present a deep learning based approach for imputing missing values in heart rate time series data collected from a participant's wearable. In particular, for each participant, we first leverage his/her historical heart rate records as a reference set to extract the underlying personalized characteristics, and then impute the missing heart rate values by considering both contextual information of the current observations and the user's features learned from previous records. Adversarial training is applied to guide the learning process, which imputed more reasonable heart rate series with the consideration of human health conditions, e.g., heart rate fluctuations. Extensive experiments are conducted on two real-world data to show the superiority of our proposed method over state-of-the-art baselines. Suwen Lin, Xian Wu 0003, Gonzalo J. Martínez, Nitesh V. Chawla |
SDM | 2 |
| 2020 | Learning from Cross-Modal Behavior Dynamics with Graph-Regularized Neural Contextual BanditabstractContextual multi-armed bandit algorithms have received significant attention in modeling users’ preferences for online personalized recommender systems in a timely manner. While significant progress has been made along this direction, a few major challenges have not been well addressed yet: (i) a vast majority of the literature is based on linear models that cannot capture complex non-linear inter-dependencies of user-item interactions; (ii) existing literature mainly ignores the latent relations among users and non-recommended items: hence may not properly reflect users’ preferences in the real-world; (iii) current solutions are mainly based on historical data and are prone to cold-start problems for new users who have no interaction history. Xian Wu 0003, Suleyman Cetintas, Deguang Kong, Miao Lu, Jian Yang 0002, Nitesh V. Chawla |
WWW | 1 |
| 2020 | Hierarchically Structured Transformer Networks for Fine-Grained Spatial Event ForecastingabstractSpatial event forecasting is challenging and crucial for urban sensing scenarios, which is beneficial for a wide spectrum of spatial-temporal mining applications, ranging from traffic management, public safety, to environment policy making. In spite of significant progress has been made to solve spatial-temporal prediction problem, most existing deep learning based methods based on a coarse-grained spatial setting and the success of such methods largely relies on data sufficiency. In many real-world applications, predicting events with a fine-grained spatial resolution do play a critical role to provide high discernibility of spatial-temporal data distributions. However, in such cases, applying existing methods will result in weak performance since they may not well capture the quality spatial-temporal representations when training triple instances are highly imbalanced across locations and time. Xian Wu 0003, Chao Huang 0001, Chuxu Zhang, Nitesh V. Chawla |
WWW | 1 |
| 2019 | Similarity-Aware Network Embedding with Self-Paced LearningabstractNetwork embedding, which aims to learn low-dimensional vector representations for nodes in a network, has shown promising performance for many real-world applications, such as node classification and clustering. While various embedding methods have been developed for network data, they are limited in their assumption that nodes are correlated with their neighboring nodes with the same similarity degree. As such, these methods can be suboptimal for embedding network data. In this paper, we propose a new method named SANE, short for Similarity-Aware Network Embedding, to learn node representations by explicitly considering different similarity degrees between connected nodes in a network. In particular, we develop a new framework based on self-paced learning by accounting for both the explicit relations (i.e., observed links) and implicit relations (i.e., unobserved node similarities) in network representation learning. To justify our proposed model, we perform experiments on two real-world network data. Experiments results show that SNAE outperforms state-of-the-art embedding models on the tasks of node classification and node clustering. Chao Huang 0001, Baoxu Shi, Xuchao Zhang, Xian Wu 0003, Nitesh V. Chawla |
CIKM | 4 |
| 2019 | Deep Prototypical Networks for Imbalanced Time Series Classification under Data ScarcityabstractWith the increase of temporal data availability, time series classification has drawn a lot of attention in the literature because of its wide spectrum of applications in diverse domains (e.g., healthcare, bioinformatics and finance), ranging from human activity recognition to financial pattern identification. While significant progress has been made to solve time series classification problem, the success of such methods relies on data sufficiency, and may not well capture the quality embeddings when training triple instances are scarce and highly imbalance across classes. To address these challenges, we propose a prototype embedding framework-Deep Prototypical Networks (DPN), which leverages a main embedding space to capture the discrepancies of difference time series classes for alleviating data scarcity. In addition, we further augment DPN framework with a relationship-dependent masking module to automatically fuse relevant information with a distance metric learning process, which addresses the data imbalance issue and performs robust time series classification. Experimental results show significant and consistent improvements compared to state-of-the-art techniques. Chao Huang 0001, Xian Wu 0003, Xuchao Zhang, Suwen Lin, Nitesh V. Chawla |
CIKM | 2 |
| 2019 | Online Purchase Prediction via Multi-Scale Modeling of Behavior DynamicsabstractOnline purchase forecasting is of great importance in e-commerce platforms, which is the basis of how to present personalized interesting product lists to individual customers. However, predicting online purchases is not trivial as it is influenced by many factors including: (i) the complex temporal pattern with hierarchical inter-correlations; (ii) arbitrary category dependencies. To address these factors, we develop a Graph Multi-Scale Pyramid Networks (GMP) framework to fully exploit users' latent behavioral patterns with both multi-scale temporal dynamics and arbitrary inter-dependencies among product categories. In GMP, we first design a multi-scale pyramid modulation network architecture which seamlessly preserves the underlying hierarchical temporal factors--governing users' purchase behaviors. Then, we employ convolution recurrent neural network to encode the categorical temporal pattern at each scale. After that, we develop a resolution-wise recalibration gating mechanism to automatically re-weight the importance of each scale-view representations. Finally, a context-graph neural network module is proposed to adaptively uncover complex dependencies among category-specific purchases. Extensive experiments on real-world e-commerce datasets demonstrate the superior performance of our method over state-of-the-art baselines across various settings. Chao Huang 0001, Xian Wu 0003, Xuchao Zhang, Chuxu Zhang, Jiashu Zhao, Dawei Yin 0001, Nitesh V. Chawla |
KDD | 2 |
| 2019 | Neural Tensor Factorization for Temporal Interaction LearningabstractNeural collaborative filtering (NCF) and recurrent recommender systems (RRN) have been successful in modeling relational data (user-item interactions). However, they are also limited in their assumption of static or sequential modeling of relational data as they do not account for evolving users' preference over time as well as changes in the underlying factors that drive the change in user-item relationship over time. We address these limitations by proposing a Neural network based Tensor Factorization (NTF) model for predictive tasks on dynamic relational data. The NTF model generalizes conventional tensor factorization from two perspectives: First, it leverages the long short-term memory architecture to characterize the multi-dimensional temporal interactions on relational data. Second, it incorporates the multi-layer perceptron structure for learning the non-linearities between different latent factors. Our extensive experiments demonstrate the significant improvement in both the rating prediction and link prediction tasks on various dynamic relational data by our NTF model over both neural network based factorization models and other traditional methods. Xian Wu 0003, Baoxu Shi, Yuxiao Dong, Chao Huang 0001, Nitesh V. Chawla |
WSDM | 1 |
| 2019 | MiST: A Multiview and Multimodal Spatial-Temporal Learning Framework for Citywide Abnormal Event ForecastingabstractCitywide abnormal events, such as crimes and accidents, may result in loss of lives or properties if not handled efficiently. It is important for a wide spectrum of applications, ranging from public order maintaining, disaster control and people's activity modeling, if abnormal events can be automatically predicted before they occur. However, forecasting different categories of citywide abnormal events is very challenging as it is affected by many complex factors from different views: (i) dynamic intra-region temporal correlation; (ii) complex inter-region spatial correlations; (iii) latent cross-categorical correlations. In this paper, we develop a Multi-View and Multi-Modal Spatial-Temporal learning (MiST) framework to address the above challenges by promoting the collaboration of different views (spatial, temporal and semantic) and map the multi-modal units into the same latent space. Specifically, MiST can preserve the underlying structural information of multi-view abnormal event data and automatically learn the importance of view-specific representations, with the integration of a multi-modal pattern fusion module and a hierarchical recurrent framework. Extensive experiments on three real-world datasets, i.e., crime data and urban anomaly data, demonstrate the superior performance of our MiST method over the state-of-the-art baselines across various settings. Chao Huang 0001, Chuxu Zhang, Jiashu Zhao, Xian Wu 0003, Nitesh V. Chawla, Dawei Yin 0001 |
WWW | 4 |
| 2018 | RESTFul: Resolution-Aware Forecasting of Behavioral Time Series DataabstractLeveraging historical behavioral data (e.g., sales volume and email communication) for future prediction is of fundamental importance for practical domains ranging from sales to temporal link prediction. Current forecasting approaches often use only a single time resolution (e.g., daily or weekly), which truncates the range of observable temporal patterns. However, real-world behavioral time series typically exhibit patterns across multi-dimensional temporal patterns, yielding dependencies at each level. To fully exploit these underlying dynamics, this paper studies the forecasting problem for behavioral time series data with the consideration of multiple time resolutions and proposes a multi-resolution time series forecasting framework, RESolution-aware Time series Forecasting (RESTFul). In particular, we first develop a recurrent framework to encode the temporal patterns at each resolution. In the fusion process, a convolutional fusion framework is proposed, which is capable of learning conclusive temporal patterns for modeling behavioral time series data to predict future time steps. Our extensive experiments demonstrate that the RESTFul model significantly outperforms the state-of-the-art time series prediction techniques on both numerical and categorical behavioral time series data. Xian Wu 0003, Baoxu Shi, Yuxiao Dong, Chao Huang 0001, Louis Faust, Nitesh V. Chawla |
CIKM | 1 |
| 2018 | Who will Attend This Event Together? Event Attendance Prediction via Deep LSTM NetworksabstractEvent-based social network (EBSN) services have emerged as a new platform on which users can choose events of interest to attend in the physical world. Over years, there are growing research interests in predicting whether certain actors will participate in an event together. In this work, we refer to this task as the event attendance prediction problem and aim to address the predictability of individuals' event attendance. In real-world settings, the factors that influence an individual's attendance may change over time, leading to the dynamic nature of individuals' behavior. However, existing event attendance prediction methods cannot deal with such dynamic scenarios. To address this issue, we propose an end-to-end Deep Event Attendance Prediction (DEAP) framework—a three-level hierarchical LSTM architecture—to explicitly model users' multi-dimensional and evolving preferences. Extensive experiments on three real-world datasets demonstrate that DEAP significantly outperforms the state-of-the-art techniques across various settings. Xian Wu 0003, Yuxiao Dong, Baoxu Shi, Ananthram Swami, Nitesh V. Chawla |
SDM | 1 |
| 2017 | Reliable fake review detection via modeling temporal and behavioral patternsabstractFake reviews have become a pervasive problem in online review systems, wherein fraudulent users manipulate the perception of an object (e.g., a restaurant) by fabricating fake reviews. Extensive work has been devoted to identifying fake reviews via modeling different factors separately, such as user features, object characteristics, and user-object bipartite relations. However, this problem remains challenging due to the fact that more advanced camouflage strategies are utilized by malicious users. In real-world scenarios, spammers may pretend to be normal users by giving fake reviews with the similar score distribution as normal users. To address these issues, we propose to explore the temporal patterns of users' review behavior, because spammers prefer to promote or demote the target businesses in a short period of time. In this work, we present a unified framework Reliable Fake Review Detection (RFRD) that explicitly models temporal patterns of users' review behavior into a probabilistic generative model. Moreover, the RFRD framework models users' underlying review credibility and objects' highly-skewed review distributions. We conduct experiments on two Yelp datasets, demonstrating the effectiveness of the proposed RFRD framework. Xian Wu 0003, Yuxiao Dong, Jun Tao 0002, Chao Huang 0001, Nitesh V. Chawla |
IEEE BigData | 1 |
| 2017 | UAPD: Predicting Urban Anomalies from Spatial-Temporal Data
Xian Wu 0003, Yuxiao Dong, Chao Huang 0001, Jian Xu 0019, Dong Wang 0002, Nitesh V. Chawla |
ECML/PKDD (2) | 1 |
| 2016 | Crowdsourcing-based Urban Anomaly Prediction System for Smart CitiesabstractCrowdsourcing has become an emerging data collection paradigm for smart city applications. A new category of crowdsourcing-based urban anomaly reporting systems have been developed to enable pervasive and real-time reporting of anomalies in cities (e.g., noise, illegal use of public facilities, urban infrastructure malfunctions). An interesting challenge in these applications is how to accurately predict an anomaly in a given region of the city before it happens. Prior works have made significant progress in anomaly detection. However, they can only detect anomalies after they happen, which may lead to significant information delay and lack of preparedness to handle the anomalies in an efficient way. In this paper, we develop a Crowdsourcing-based Urban Anomaly Prediction Scheme (CUAPS) to accurately predict the anomalies of a city by exploring both spatial and temporal information embedded in the crowdsourcing data. We evaluated the performance of our scheme and compared it to the state-of-the-art baselines using four real-world datasets collected from 311 service in the city of New York. The results showed that our scheme can predict different categories of anomalies in a city more accurately than the baselines. Chao Huang 0001, Xian Wu 0003, Dong Wang 0002 |
CIKM | 2 |