Kunqing Xie

dblp:06/1913 · DBLP profile ↗
← Back
52ranked-venue papers
0as first author
5since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 26 · 3 since 2021Databases, data management, data science and information retrieval · 23 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 13Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 since 2021Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2022 GraphAD: A Graph Neural Network for Entity-Wise Multivariate Time-Series Anomaly Detection
abstract
In recent years, the emergence and development of third-party platforms have greatly facilitated the growth of the Online to Offline (O2O) business. However, the large amount of transaction data raises new challenges for retailers, especially anomaly detection in operating conditions. Thus, platforms begin to develop intelligent business assistants with embedded anomaly detection methods to reduce the management burden on retailers. Traditional time-series anomaly detection methods capture underlying patterns from the perspectives of time and attributes, ignoring the difference between retailers in this scenario. Besides, similar transaction patterns extracted by the platforms can also provide guidance to individual retailers and enrich their available information without privacy issues. In this paper, we pose an entity-wise multivariate time-series anomaly detection problem that considers the time-series of each unique entity. To address this challenge, we propose GraphAD, a novel multivariate time-series anomaly detection model based on the graph neural network. GraphAD decomposes the Key Performance Indicator (KPI) into stable and volatility components and extracts their patterns in terms of attributes, entities and temporal perspectives via graph neural networks. We also construct a real-world entity-wise multivariate time-series dataset from the business data of Ele.me. The experimental results on this dataset show that GraphAD significantly outperforms existing anomaly detection methods.
Xu Chen 0022, Qiu Qiu, Changshan Li, Kunqing Xie
SIGIR4
2022 Understanding and Improvement of Adversarial Training for Network Embedding from an Optimization Perspective
abstract
Network Embedding aims to learn a function mapping the nodes to Euclidean space contribute to multiple learning analysis tasks on networks. However, both the noisy information behind the real-world networks and the overfitting problem negatively impact the quality of embedding vectors. To tackle these problems, researchers utilize Adversarial Perturbations on Parameters (APP) and achieve state-of-the-art performance. Unlike the mainstream methods introducing perturbations on the network structure or the data feature, Adversarial Training for Network Embedding (AdvTNE) adopts APP to directly perturb the model parameters, thus providing a new chance to understand the mechanism behind it. In this paper, we explain APP theoretically from an optimization perspective. Considering the Power-law property of networks and the optimization objective, we analyze the reason for its remarkable results on network embedding. Based on the above analysis and the Sigmoid saturation region problem, we propose a new Sine-base activation to enhance the performance of AdvTNE. We conduct extensive experiments on four real networks to validate the effectiveness of our method in node classification and link prediction. The results demonstrate that our method is competitive with state-of-the-art methods.
Lun Du, Xu Chen 0022, Qiang Fu 0015, Kunqing Xie, Shi Han, Dongmei Zhang 0001
WSDM5
2021 Fast Hierarchy Preserving Graph Embedding via Subspace Constraints
abstract
Hierarchy preserving network embedding is a method that project nodes into feature space by preserving the hierarchy property of networks. Recently, researches on network representation have considerably profited from taking hierarchy into consideration. Among these works, SpaceNE1[1] stands out by preserving hierarchy with the help of subspace constraints on the hierarchy subspace system. However, like all other hierarchy preserving network embedding methods, SpaceNE is time-consuming and cannot generalize to new nodes. In this paper, we propose an inductive method, FastHGE, to learn node representations more efficiently and generalize to new nodes more easily. Empirically, the experiment of node classification demonstrates that the convergence speed of FastHGE is increased by 30 times in the case of the same accuracy with SpaceNE.
Xu Chen 0022, Lun Du, Yun Wang 0012, Qingqing Long, Kunqing Xie
ICASSP6
2021 TrafficStream: A Streaming Traffic Flow Forecasting Framework Based on Graph Neural Networks and Continual Learning
abstract
With the rapid growth of traffic sensors deployed, a massive amount of traffic flow data are collected, revealing the long-term evolution of traffic flows and the gradual expansion of traffic networks. How to accurately forecasting these traffic flow attracts the attention of researchers as it is of great significance for improving the efficiency of transportation systems. However, existing methods mainly focus on the spatial-temporal correlation of static networks, leaving the problem of efficiently learning models on networks with expansion and evolving patterns less studied. To tackle this problem, we propose a Streaming Traffic Flow Forecasting Framework, TrafficStream, based on Graph Neural Networks (GNNs) and Continual Learning (CL), achieving accurate predictions and high efficiency. Firstly, we design a traffic pattern fusion method, cleverly integrating the new patterns that emerged during the long-term period into the model. A JS-divergence-based algorithm is proposed to mine new traffic patterns. Secondly, we introduce CL to consolidate the knowledge learned previously and transfer them to the current model. Specifically, we adopt two strategies: historical data replay and parameter smoothing. We construct a streaming traffic data set to verify the efficiency and effectiveness of our model. Extensive experiments demonstrate its excellent potential to extract traffic patterns with high efficiency on long-term streaming network scene. The source code is available at https://github.com/AprLie/TrafficStream.
Junshan Wang, Kunqing Xie
IJCAI3
2021 Spatial-Temporal Graph ODE Networks for Traffic Flow Forecasting
abstract
Spatial-temporal forecasting has attracted tremendous attention in a wide range of applications, and traffic flow prediction is a canonical and typical example. The complex and long-range spatial-temporal correlations of traffic flow bring it to a most intractable challenge. Existing works typically utilize shallow graph convolution networks (GNNs) and temporal extracting modules to model spatial and temporal dependencies respectively. However, the representation ability of such models is limited due to: (1) shallow GNNs are incapable to capture long-range spatial correlations, (2) only spatial connections are considered and a mass of semantic connections are ignored, which are of great importance for a comprehensive understanding of traffic networks. To this end, we propose Spatial-Temporal Graph Ordinary Differential Equation Networks (STGODE).1 Specifically, we capture spatial-temporal dynamics through a tensor-based ordinary differential equation (ODE), as a result, deeper networks can be constructed and spatial-temporal features are utilized synchronously. To understand the network more comprehensively, semantical adjacency matrix is considered in our model, and a well-design temporal dialated convolution structure is used to capture long term temporal dependencies. We evaluate our model on multiple real-world traffic datasets and superior performance is achieved over state-of-the-art baselines.
Zheng Fang 0007, Qingqing Long, Guojie Song, Kunqing Xie
KDD4
2020 TSSRGCN: Temporal Spectral Spatial Retrieval Graph Convolutional Network for Traffic Flow Forecasting
abstract
Traffic flow forecasting is of great significance for improving the efficiency of transportation systems and preventing emergencies. Due to the highly non-linearity and intricate evolutionary patterns of short-term and long-term traffic flow, existing methods often fail to take full advantage of spatial-temporal information, especially the various temporal patterns with different period shifting and the characteristics of road segments. Besides, the globality representing the absolute value of traffic status indicators and the locality representing the relative value have not been considered simultaneously. This paper proposes a neural network model that focuses on the globality and locality of traffic networks as well as the temporal patterns of traffic data. The cycle-based dilated deformable convolution block is designed to capture different time-varying trends on each node accurately. Our model can extract both global and local spatial information since we combine two graph convolutional network methods to learn the representations of nodes and edges. Experiments on two real-world datasets show that the model can scrutinize the spatial-temporal correlation of traffic data, and its performance is better than the compared state-of-the-art methods. Further analysis indicates that the locality and globality of the traffic networks are critical to traffic flow prediction and the proposed TSSRGCN model can adapt to the various temporal traffic patterns.
Xu Chen 0022, Yuanxing Zhang, Lun Du, Zheng Fang 0007, Kaigui Bian, Kunqing Xie
ICDM7
2019 Transfer Knowledge Between Sub-regions for Traffic Prediction Using Deep Learning Method
Kunqing Xie
IDEAL (1)2
2017 Discovering spatio-temporal dependencies based on time-lag in intelligent transportation data
Xiabing Zhou, Haikun Hong, Xingxing Xing, Kaigui Bian, Kunqing Xie
Neurocomputing5
2016 Structure Feature Learning Method for Incomplete Data
abstract
Learning with incomplete data remains challenging in many real-world applications especially when the data is high-dimensional and dynamic. Many imputation-based algorithms have been proposed to handle with incomplete data, where these algorithms use statistics of the historical information to remedy the missing parts. However, these methods merely use the structural information existing in the data, which are very helpful for sharing between the complete entries and the missing ones. For example, in traffic system, some group information and temporal smoothness exist in the data structure. In this paper, we propose to incorporate these structural information and develop structural feature leaning method for learning with incomplete data (SFLIC). The SFLIC model adopt a fused Lasso based regularizer and a group Lasso style regularizer to enlarge the data sharing along both the temporal smoothness level and the feature group level to fill the gap where the data entries are missing. The proposed SFLIC model is a nonsmooth function according to the model parameters, and we adopt the smoothing proximal gradient (SPG) method to seek for an efficient solution. We evaluate our model on both synthetic and real-world highway traffic datasets. Experimental results show that our method outperforms the state-of-the-art methods.
Xiabing Zhou, Xingxing Xing, Lei Han 0001, Haikun Hong, Kaigui Bian, Kunqing Xie
Int. J. Pattern Recognit. Artif. Intell.6
2015 Learning Common Metrics for Homogenous Tasks in Traffic Flow Prediction
abstract
Nearest neighbor based nonparametric regression is a classic data-driven method for traffic flow prediction in intelligent transportation systems (ITS). Performances of those models depend heavily on the similarity or distance metric used to search nearest neighborhood. Metric learning algorithms have been developed to learn the distance metrics from data in recent years. In real-world transportation application, multiple forecasting tasks are set since there are lots of road sections and detector points in the traffic network. Previous works tend to learn only one global metric to be used for all the tasks or learn multiple local metrics for each task which may lead to under-fitting or over-fitting problem. To balance these two kinds of methods and improve the generalization of learned metrics, we propose a common metric learning algorithm under the intuition that homogenous tasks tend to have similar local metrics. Then the learned common metrics are used in common metric KNN (CM-KNN) for traffic flow prediction. Experimental results show that our algorithm to learn common metrics are reasonable and CM-KNN method for traffic flow prediction outperforms other competing methods.
Haikun Hong, Xiabing Zhou, Wenhao Huang 0001, Xingxing Xing, Kaigui Bian, Kunqing Xie
ICMLA8
2015 Improving deep neural network ensembles using reconstruction error
abstract
Ensemble learning of neural network is a learning paradigm where ensembles of several neural networks show improved generalization capabilities that outperform those of single networks. For deep learning of multi-layer neural networks, ensemble learning is still applicable. In addition, characteristics of deep neural networks can provide potential opportunities to improve the performance of traditional neural network ensembles. In this paper, we propose an ensemble criterion of deep neural networks that is based on the reconstruction error and present two strategies to solve the most important issues in ensemble learning of neural networks: component dataset sampling and output averaging. Component training datasets are selected according to the reconstruction error instead of random bootstrap sampling or re-weighting. Moreover, for each testing instance, we can compute the reconstruction error yielded by the sub-model simultaneously with the output. The reconstruction error is used as the weights in output averaging. From the perspectives of prediction interval and confidence interval, we demonstrated that smaller reconstruction error could ensure smaller prediction interval. We also incorporate the famous structure ensemble approach “Dropout” into the proposed approach to achieve the best performance. We conduct experiments on classification and regression datasets to validate the effectiveness of our approach.
Wenhao Huang 0001, Haikun Hong, Kaigui Bian, Xiabing Zhou, Guojie Song, Kunqing Xie
IJCNN6
2015 Probabilistic dynamic causal model for temporal data
abstract
Learning temporal causal structures between time series is one of key tools for analyzing time series data. Most previous works focuse on learning with static temporal causal relationships. However, in many real world applications, such as climate environment and transportation system, the causal structures vary dramatically over time. In this paper, we propose a probabilistic dynamic causal (PDC) model based on Lasso-Granger to uncover the dynamic temporal dependencies. Specifically, the PDC model infers different state varying of temporal data and causal structures of each state in one unified model. We devise the expectation-maximization (EM) algorithm to infer the model parameters. Furthermore, to address the smoothness of state varying in adjacent time, we extend the PDC model with a regularization term encouraging states to be similar in adjacent time. Though it may slightly decrease the precision on training data, it improves the generalization capability of the model. We conduct experiments on synthetic dataset as well as two real-world datasets of climate and traffic to evaluate the effectiveness of the PDC model. Experimental results show that the proposed model is effective in discovering the dynamic causal factors of Particulate Matter 2.5 (PM2.5) and traffic spatial causalities.
Xiabing Zhou, Wenhao Huang 0001, Weisong Hu, Sizhen Du, Guojie Song, Kunqing Xie
IJCNN7
2015 On Influential Nodes Tracking in Dynamic Social Networks
abstract
Real world marketing campaign utilizing the word-of-mouth effect usually lasts a long time, where multiple sets of influential users need to be mined and targeted at different times to fully utilize the power of viral marketing. As both social network structure and strength of influence between individuals evolve constantly, it requires to track the influential nodes under a dynamic setting. To address the above problem, we explore the Influential Node Tracking (INT) problem as an extension to the traditional Influence Maximization problem under dynamic social networks. While Influence Maximization problem aims at identifying a set of k nodes to maximize the joint influence under one static network, INT problem focuses on tracking a set of influential nodes that keeps maximizing the influence as the network evolves. Utilizing the smoothness of the evolution of the network structure, we propose an efficient algorithm, Upper Bound Interchange Greedy (UBI) to solve the INT problem. Instead of constructing the seed set from the ground, we start from the influential seed set we find previously and implement node replacement to improve the influence coverage. Furthermore, by using a fast update method to maintain an upper bound on the node replacing gain, our algorithm can scale to dynamic social networks with millions of nodes. Empirical experiments on three real large-scale dynamic social networks show that our UBI algorithm achieves better performance in terms of both influence coverage and running time.
Guojie Song, Xinran He, Kunqing Xie
SDM4
2015 Mining Dependencies Considering Time Lag in Spatio-Temporal Traffic Data
Xiabing Zhou, Haikun Hong, Xingxing Xing, Wenhao Huang 0001, Kaigui Bian, Kunqing Xie
WAIM6
2015 Overlapping Decomposition for Gaussian Graphical Modeling
abstract
Correlation based graphical models are developed to detect the dependence relationships among random variables and provide intuitive explanations for these relationships in complex systems. Most of the existing works focus on learning a single correlation based graphical model for all the random variables. However, it is difficult to understand and interpret the massive dependencies of the variables learned from a single graphical model at a global level especially when the graph is large. In order to provide a clearer understanding for the dependence relationships among a large number of random variables, in this paper, we propose the problem of estimating an overlapping decomposition for the Gaussian graphical model of a large scale to generate overlapping sub-graphical models, where strong and meaningful correlations remain in each subgraph with a small scale. Specifically, we propose a greedy algorithm to achieve the overlapping decomposition for the Gaussian graphical model. A key technique of the algorithm is that the problem of solving a (k + 1)-node Gaussian graphical model can be approximately reduced to the problem of solving a one-step vector regularization problem based on a solved k-node Gaussian graphical model with theoretical guarantee. Based on this technique, a greedy expansion algorithm is proposed to generate the overlapping subgraphs. Moreover, we extend the proposed method to deal with dynamic graphs where the dependence relationships among random variables vary with the time. We evaluate the proposed methods on synthetic dataset and a real-life traffic dataset, and the experimental results show the superiority of the proposed methods.
Guojie Song, Lei Han 0001, Kunqing Xie
IEEE Trans. Knowl. Data Eng.3
2015 Influence Maximization on Large-Scale Mobile Social Network: A Divide-and-Conquer Method
abstract
With the proliferation of mobile devices and wireless technologies, mobile social network systems are increasingly available. A mobile social network plays an essential role as the spread of information and influence in the form of “word-of-mouth”. It is a fundamental issue to find a subset of influential individuals in a mobile social network such that targeting them initially (e.g., to adopt a new product) will maximize the spread of the influence (further adoptions of the new product). The problem of finding the most influential nodes is unfortunately NP-hard. It has been shown that a Greedy algorithm with provable approximation guarantees can give good approximation; However, it is computationally expensive, if not prohibitive, to run the greedy algorithm on a large mobile social network. In this paper, a divide-and-conquer strategy with parallel computing mechanism has been adopted. We first propose an algorithm called Community-based Greedy algorithm for mining top-K influential nodes. It encompasses two components: dividing the large-scale mobile social network into several communities by taking into account information diffusion and selecting communities to find influential nodes by a dynamic programming. Then, to further improve the performance, we parallelize the influence propagation based on communities and consider the influence propagation crossing communities. Also, we give precision analysis to show approximation guarantees of our models. Experiments on real large-scale mobile social networks show that the proposed methods are much faster than previous algorithms, meanwhile, with high accuracy.
Guojie Song, Xiabing Zhou, Kunqing Xie
IEEE Trans. Parallel Distributed Syst.4
2014 Encoding Tree Sparsity in Multi-Task Learning: A Probabilistic Framework
abstract
Multi-task learning seeks to improve the generalization performance by sharing common information among multiple related tasks. A key assumption in most MTL algorithms is that all tasks are related, which, however, may not hold in many real-world applications. Existing techniques, which attempt to address this issue, aim to identify groups of related tasks using group sparsity. In this paper, we propose a probabilistic tree sparsity (PTS) model to utilize the tree structure to obtain the sparse solution instead of the group structure. Specifically, each model coefficient in the learning model is decomposed into a product of multiple component coefficients each of which corresponds to a node in the tree. Based on the decomposition, Gaussian and Cauchy distributions are placed on the component coefficients as priors to restrict the model complexity. We devise an efficient expectation maximization algorithm to learn the model parameters. Experiments conducted on both synthetic and real-world problems show the effectiveness of our model compared with state-of-the-art baselines.
Lei Han 0001, Yu Zhang 0006, Guojie Song, Kunqing Xie
AAAI4
2014 Deep process neural network for temporal deep learning
abstract
Process neural network is widely used in modeling temporal process inputs in neural networks. Traditional process neural network is usually limited in structure of single hidden layer due to the unfavorable training strategies of neural network with multiple hidden layers and complex temporal weights in process neural network. Deep learning has emerged as an effective pre-training method for neural network with multiple hidden layers. Though deep learning is usually limited in static inputs, it provided us a good solution for training neural network with multiple hidden layers. In this paper, we extended process neural network to deep process neural network. Two basic structures of deep process neural network are discussed. One is the accumulation first deep process neural network and the other is accumulation last deep process neural network. We could build any architecture of deep process neural network based on those two structures. Temporal process inputs are represented as sequences in this work for the purpose of unsupervised feature learning with less prior knowledge. Based on this, we proposed learning algorithms for two basic structures inspired by the numerical learning approach for process neural network and the auto-encoder in deep learning. Finally, extensive experiments demonstrated that deep process neural network is effective in tasks with temporal process inputs. Accuracy of deep process neural network is higher than traditional process neural network while time complexity is near in the task of traffic flow prediction in highway system.
Wenhao Huang 0001, Haikun Hong, Guojie Song, Kunqing Xie
IJCNN4
2014 Dynamic boosting in deep learning using reconstruction error
abstract
Deep learning has attracted a lot of attention in research and industry in recent years. Behind the success of deep learning, there is much space for improvement. It is difficult to identify if a testing sample can be represented by the deep network effectively before we examining the final result. In this paper, we proposed a dynamic boosting strategy according to reconstruction error in deep networks. We use reconstruction error to determine whether the result is reliable or not. From the perspective of prediction interval, we demonstrated that with the increase of reconstruction error, the prediction interval would become bigger. Therefore, the classification result is not reliable when the reconstruction error exceeds the predetermined threshold. Since we can record the reconstruction error as well as the classification error for all training samples in training set. We can learn an extra boosting model besides the deep network in training set to improve the performance of the model. An important factor in learning the boosting model is to determine an appropriate threshold for selecting training samples. In testing, we first examine whether the reconstruction error of a testing sample exceeds the threshold to determine if we should use the boosting model. If the boosting model is used, the final result is the average of the output of the deep network and the boosting model. We conducted experiments on two widely used classification datasets and an air quality dataset. From the experiments, we see that our boosting strategy is effective in improving the performance of classification. We tested several boosting models in this paper. They can all reduce the test error to some extent under appropriate parameter settings.
Wenhao Huang 0001, Weisong Hu, Haikun Hong, Guojie Song, Kunqing Xie
IJCNN6
2014 Metric-Based Multi-Task Grouping Neural Network for Traffic Flow Forecasting
Haikun Hong, Wenhao Huang 0001, Guojie Song, Kunqing Xie
ISNN4
2014 A Spatial-temporal Topic Segmentation Model for Human Mobile Behavior
Xingxing Xing, Weisong Hu, Wenhao Huang 0001, Guojie Song, Kunqing Xie
WAIM6
2014 Deep Architecture for Traffic Flow Prediction: Deep Belief Networks With Multitask Learning
abstract
Traffic flow prediction is a fundamental problem in transportation modeling and management. Many existing approaches fail to provide favorable results due to being: 1) shallow in architecture; 2) hand engineered in features; and 3) separate in learning. In this paper we propose a deep architecture that consists of two parts, i.e., a deep belief network (DBN) at the bottom and a multitask regression layer at the top. A DBN is employed here for unsupervised feature learning. It can learn effective features for traffic flow prediction in an unsupervised fashion, which has been examined and found to be effective for many areas such as image and audio classification. To the best of our knowledge, this is the first paper that applies the deep learning approach to transportation research. To incorporate multitask learning (MTL) in our deep architecture, a multitask regression layer is used above the DBN for supervised prediction. We further investigate homogeneous MTL and heterogeneous MTL for traffic flow prediction. To take full advantage of weight sharing in our deep architecture, we propose a grouping method based on the weights in the top layer to make MTL more effective. Experiments on transportation data sets show good performance of our deep architecture. Abundant experiments show that our approach achieved close to 5% improvements over the state of the art. It is also presented that MTL can improve the generalization performance of shared tasks. These positive results demonstrate that deep learning and MTL are promising in transportation research.
Wenhao Huang 0001, Guojie Song, Haikun Hong, Kunqing Xie
IEEE Trans. Intell. Transp. Syst.4
2013 Deep Architecture for Traffic Flow Prediction
Wenhao Huang 0001, Haikun Hong, Weisong Hu, Guojie Song, Kunqing Xie
ADMA (2)6
2012 An Adaption of Relief for Redundant Feature Elimination
Tianshu Wu, Kunqing Xie, Chengkai Nie, Guojie Song
ISNN (2)2
2012 Overlapping decomposition for causal graphical modeling
abstract
Causal graphical models are developed to detect the dependence relationships between random variables and provide intuitive explanations for the relationships in complex systems. Most of existing work focuses on learning a single graphical model for all the variables. However, a single graphical model cannot accurately characterize the complicated causal relationships for a relatively large graph. In this paper, we propose the problem of estimating an overlapping decomposition for Gaussian graphical models of a large scale to generate overlapping sub-graphical models. Specifically, we formulate an objective function for the overlapping decomposition problem and propose an approximate algorithm for it. A key theory of the algorithm is that the problem of solving a κ+1 node graphical model can be reduced to the problem of solving a one-step regularization based on a solved κ node graphical model. Based on this theory, a greedy expansion algorithm is proposed to generate the overlapping subgraphs. We evaluate the effectiveness of our model on both synthetic datasets and real traffic dataset, and the experimental results show the superiority of our method.
Lei Han 0001, Guojie Song, Gao Cong, Kunqing Xie
KDD4
2011 Simulated Annealing Based Influence Maximization in Social Networks
abstract
The problem of influence maximization, i.e., mining top-k influential nodes from a social network such that the spread of influence in the network is maximized, is NP-hard. Most of the existing algorithms for the prob- lem are based on greedy algorithm. Although greedy algorithm can achieve a good approximation, it is computational expensive. In this paper, we propose a totally different approach based on Simulated Annealing(SA) for the influence maximization problem. This is the first SA based algorithm for the problem. Additionally, we propose two heuristic methods to accelerate the con- vergence process of SA, and a new method of comput- ing influence to speed up the proposed algorithm. Experimental results on four real networks show that the proposed algorithms run faster than the state-of-the-art greedy algorithm by 2-3 orders of magnitude while being able to improve the accuracy of greedy algorithm.
Qingye Jiang, Guojie Song, Gao Cong, Wenjun Si, Kunqing Xie
AAAI6
2011 Transportation Modes Identification from Mobile Phone Data Using Probabilistic Models
Dafeng Xu, Guojie Song, Rongzeng Cao, Xinwei Nie, Kunqing Xie
ADMA (2)6
2011 Discrete Trajectory Prediction on Mobile Data
Wenhao Huang 0001, Guojie Song, Kunqing Xie
APWeb4
2011 Efficient approaches for summarizing subspace clusters into k representatives
Xiuli Ma, Dongqing Yang, Shiwei Tang, Meng Shuai, Kunqing Xie
Soft Comput.6
2010 Anchor Points Seeking of Large Urban Crowd Based on the Mobile Billing Data
Wenhao Huang 0001, Zhengbin Dong, Guojie Song, Kunqing Xie
ADMA (1)8
2010 Community-based greedy algorithm for mining top-K influential nodes in mobile social networks
abstract
With the proliferation of mobile devices and wireless technologies, mobile social network systems are increasingly available. A mobile social network plays an essential role as the spread of information and influence in the form of "word-of-mouth". It is a fundamental issue to find a subset of influential individuals in a mobile social network such that targeting them initially (e.g. to adopt a new product) will maximize the spread of the influence (further adoptions of the new product). The problem of finding the most influential nodes is unfortunately NP-hard. It has been shown that a Greedy algorithm with provable approximation guarantees can give good approximation; However, it is computationally expensive, if not prohibitive, to run the greedy algorithm on a large mobile network.
Gao Cong, Guojie Song, Kunqing Xie
KDD4
2009 Mining the Structure and Evolution of the Airport Network of China over the Past Twenty Years
Zhengbin Dong, Xiujun Ma, Kunqing Xie, Fengjun Jin
ADMA4
2009 Numerical Learning Method for Process Neural Network
Tianshu Wu, Kunqing Xie, Guojie Song, Xingui He
ISNN (1)2
2009 Accelerating sequence searching: dimensionality reduction method
Guojie Song, Bin Cui 0001, Baihua Zheng, Kunqing Xie, Dongqing Yang
Knowl. Inf. Syst.4
2008 Squeezing Long Sequence Data for Efficient Similarity Search
Guojie Song, Bin Cui 0001, Baihua Zheng, Kunqing Xie, Dongqing Yang
APWeb4
2008 An online approach based on locally weighted learning for short-term traffic flow prediction
abstract
Traffic flow prediction is a basic function of Intelligent Transportation System. Due to the complexity of traffic phenomenon, most existing methods build complex models such as neural networks for traffic flow prediction. As a model may lose effect with time lapse, it is important to update the model on line. However, the high computational cost of maintaining a complex model puts great challenge for model updating. The high computation cost lies in two aspects: computation of complex model coefficients and huge amount training data for it. In this paper, we propose to use a nonparametric approach based on locally weighted learning to predict traffic flow. Our approach incrementally incorporates new data to the model and is computationally efficient, which makes it suitable for online model updating and predicting. In addition, we adopt wavelet analysis to extract the periodic characteristic of the traffic data, which is then used for the input of the prediction model instead of the raw traffic flow data. The primary experiments on real data demonstrate the effectiveness and efficiency of our approach.
Meng Shuai, Kunqing Xie, Wen Pu, Guojie Song, Xiujun Ma
GIS2
2008 Integrating Map Services and Location-based Services for Geo-Referenced Individual Data Collection
abstract
With the rapid advance of location-based services (LBS) and online map services, it is now more feasible than before to collect the geo-referenced individual level data. However, privacy is always an issue whenever personal location is traced and recorded. This paper proposes a reactive location-based service to collect and process individual location data by different privacy policies. The reactive LBS provides user with an active pull mode to collect his/her location information. With different privacy policies, the user can enter his/her current address manually to the server or automatically generated by the LBS location provider. In order to gain the accurate reconstruction of individuals' activity-travel patterns with considerable space-time details, the LBS server invokes an online map service to geo-reference the location data and to derive the route between locations. In this paper, a household's daily activity survey scenario is showed in Beijing city.
Xiujun Ma, Zhongya Wei, Yanwei Chai, Kunqing Xie
IGARSS (5)4
2008 Real-Time Short-Term Traffic Flow Forecasting Based on Process Neural Network
Guojie Song, Kunqing Xie, Yizhou Sun
ISNN (2)4
2007 CLAIM: An Efficient Method for Relaxed Frequent Closed Itemsets Mining over Stream Data
Guojie Song, Dongqing Yang, Bin Cui 0001, Baihua Zheng, Kunqing Xie
DASFAA6
2007 An Optimized Process Neural Network Model
Guojie Song, Dongqing Yang, Bin Cui 0001, Kunqing Xie
DASFAA6
2007 Local Word Bag Model for Text Categorization
abstract
Many text processing applications adopted the bag of words (BOW) model representation of documents, in which each document is represented as a vector of weighted terms or n-grams, and then the cosine distance between two vectors is used as the similarity measurement. Although the great success in information retrieval and text categorization, the conventional BOW model ignores the detailed local text information, i.e. the co-occurrence pattern of words at sentence or paragraph level. In this paper, we propose a novel approach to represent a document as a set of local tf-idf vectors, or what we called local word bags (LWB). By encapsulating local information distributed around a document into multiple LWBs, we can measure the similarity of two documents via the partial match of their corresponding local bags. To perform the matching efficiently, we introduce the local word bag kernel (LWB kernel), a variant of VG-Pyramid match kernel. The new kernel enables the discriminative machine learning methods like SVM to compute the partial matching between two sets of LWBs in linear time after an one time hierarchical clustering procedure over all local bags at the initialization stage. Experiments on real world datasets demonstrate the effectiveness of our new approach.
Wen Pu, Ning Liu 0001, Shuicheng Yan, Jun Yan 0001, Kunqing Xie, Zheng Chen 0001
ICDM5
2007 Causal relation of queries from temporal logs
abstract
In this paper, we study a new problem of mining causal relation of queries in search engine query logs. Causal relation between two queries means event on one query is the causation of some event on the other. We first detect events in query logs by efficient statistical frequency threshold. Then the causal relation of queries is mined by the geometric features of the events. Finally the Granger Causality Test (GCT) is utilized to further re-rank the causal relation of queries according to their GCT coefficients. In addition, we develop a 2-dimensional visualization tool to display the detected relationship of events in a more intuitive way. The experimental results on the MSN search engine query logs demonstrate that our approach can accurately detect the events in temporal query logs and the causal relation of queries is detected effectively.
Yizhou Sun, Kunqing Xie, Ning Liu 0001, Shuicheng Yan, Benyu Zhang, Zheng Chen 0001
WWW2
2005 G-WSDL: a data-oriented approach to model GIS Web services
abstract
GIS Web service turns out to be one of the new directions of the paradigm of GIS, due to the popular use of the Internet and the dramatic progress of telecommunications technology. However, focused on the description of the interface of functions by Web service description language (WSDL), the general Web service technologies are not sufficient to model GIS Web service with data-oriented characteristic, especially those data provider services and data transform services. In this paper, we propose geographic-Web service description language (abbr. G-WSDL) which enhances the ability of describing such data-driven services by extending the current WSDL with service metadata. Two application scenarios will be presented at last to show its strength in automatic GIS Web service discovery and dynamic services combination.
Kunqing Xie, Xiujun Ma, Lebin Sun
IGARSS2
2005 A spatio-temporal database prototype for managing moving objects in GIS
abstract
The location technologies, such as GPS and telegraphy, are producing more and more data of moving objects. Spatio-temporal database is needed to manage these data, so as to solve the problems in spatio-temporal applications. However, there are few prototypes of spatio-temporal database systems yet. This is because attention is always paid on some specified aspects not the whole system. In this paper, we present the design and implementation of a spatio-temporal database prototype for managing moving objects. A Trajectory Model is proposed in this paper, where the moving objects' movements in 2D plane over time are treated as trajectories in the 3D spatio-temporal space. Referring to R tree and the trajectory oriented indexing method TB tree, a new indexing method: S-TB tree is proposed in the prototype. Distance based queries and trajectory topology based queries are all supported by the prototype. The distance based queries includes point queries, range queries. This paper also presents different strategies to process these queries. Finally, two application cases are presented: traffic control and historical queries of the storms on Earth.
Kunqing Xie, Xiujun Ma, Huibin Zhang
IGARSS3
2005 A novel method to integrate spatial data mining and geographic information system
abstract
GIS is a good spatial analysis tool, nevertheless, as the accumulation of spatial data, the functions that GIS offers are not enough. Spatial data mining, which can automatically discover implicit knowledge from spatial data, has recently received wide attention. However, the spatial data mining systems are not competent for preprocessing and presenting spatial data. Hence, the idea to integrate spatial data mining and GIS is straightforward. In this paper, we propose a novel method to integrate spatial data mining and GIS. This method is implemented with SPMML, an XML-based language that is extended from PMML. We build a prototype system with this method and it works well.
Xingxing Jin, Yingkun Cai, Kunqing Xie, Xiujun Ma, Cuo Cai
IGARSS3
2005 A spatio-temporal aquarium for visual exploration on geographic phenomena
Xiujun Ma, Kunqing Xie, Cuo Cai, Wen Pu
IGARSS3
2005 A compensation mechanism in GIS Web service composition
abstract
With the evolution of GIS from stand-alone systems with geo-data tightly coupled with systems to an increasingly distributed model based on independently-provided, interoperable GIS Web service, much more research has been focused on GIS Web service composition. However, little research works concern on control mechanisms of improving availability and reliability in GIS Web service composition. Considering that GIS Web services are in essence loosely-coupled and hosted by different providers. As a result, any update of any service, might affect critically the overall composition consistency and execution; Moreover because most of geographic operations are CPU-intensive, it means that GIS Web service compositions consisting of basic geographic operations are more CPU-intensive - it will perhaps take several days to finish the execution of one single GIS Web service composition! Without support of high availability and reliability, any mistake will lead to the failure of the whole composition execution and the waste of much computing resource. So availability and reliability in GIS Web service composition should be paid more attention. This paper proposes a mechanism of service compensation to achieve high availability and reliability in GIS Web service composition. The basic idea is that when one Web service of composition fails, the execution of the whole composition is not aborted immediately; on the contrary, the compensation mechanism tries to find another Web service which can provide the same function to "compensate" it. The abortion of composition only happens when the "compensating" Web services can not be found or all of them fail.
Xiujun Ma, Kunqing Xie
IGARSS3
2005 Detecting spatio-temporal outliers in climate dataset: a method study
abstract
Outlier detecting is one of the most important data analysis technologies in data mining, which can be used to discover anomalous phenomena in huge dataset. Many literatures on spatial outlier detecting and time series outlier detecting have appeared, while the area of spatio-temporal outliers considering both spatial and temporal dimensions has still rarely been touched. Defining outliers in traditional dataset is more explicit because the data structure we need to focus on is very straightforward (e.g., a spatial point or a transaction record). However, it is much more difficult to give outlier a definite characterization in spatio-temporal lattice data, since there are so many data structures we can pay attention to. With the aim of detecting useful and meaningful outliers in climate dataset, we introduce a formalized way to define outliers in spatio-temporal lattice data, in which the importance of clarifying basic data structure (we call it basic element in our paper) is stressed. As a case study, we define two kinds of spatio-temporal outliers based on a global climate dataset, according to the three aspects we propose in defining an outlier. The introduction of basic element and the formulation of outlier definition process make it easier and clearer to define meaningful outliers. Thus outlier detecting in spatio-temporal lattice data will provide us with really interesting and useful knowledge.
Kunqing Xie, Xiujun Ma, Xingxing Jin, Wen Pu
IGARSS2
2005 Adaptive sampling for selectivity estimation in spatial database
abstract
Spatial sampling is a significant part of query processing in spatial database. In this paper, an adaptive data-driven sampling for selectivity estimation is proposed in spatial database. This technique presents an efficient sampling for spatial data, especially for two-dimensional line and polygon data, and it make the sample size fit the limit of time and memory, or a user-defined parameter. The data-driven sampling technique is compared with various techniques on different type of datasets in our experimental study, and it out outperforms the other techniques over a broad range of query workloads and datasets.
Xiujun Ma, Kunqing Xie, Huibin Zhang, Wen Pu
IGARSS3
2005 Spatial data cube: provides better support for spatial data mining
abstract
Spatial data mining is a promising technique that deals with extraction of implicit knowledge or other interesting patterns from large amount of spatial data. Though most data mining systems work with data stored in flat files or operational database, it has been recognized that mining in a data warehouse usually result in better information. Because data are usually cleansed before they are stored into data warehouse. Furthermore, data warehouse provides data with different levels of summarization for the clients, which will lead to fruitful data mining. However, current techniques of data warehouse can not handle spatial data well. Both dimensions and measures in the data model of data warehouse are nonspatial data. In this paper, we propose a new data model called spatial data cube for data warehouse. Spatial data cube supports both spatial and nonspatial data. We also introduce how to construct a spatial data cube that can answer queries efficiently by selective materialization. We believe that the spatial data cube can provide better support for spatial data mining.
Kunqing Xie, Xiujun Ma, Cuo Cai, Shiwei Tang
IGARSS2
2005 Efficient discovery of multilevel spatial association rules using partitions
Kunqing Xie, Xiuli Ma
Inf. Softw. Technol.2
2004 A Comparative Study on Feature Weight in Text Categorization
Zhi-Hong Deng 0001, Shiwei Tang, Dongqing Yang, Ming Zhang 0004, Liyu Li, Kunqing Xie
APWeb6