Nan Han

dblp:84/8629 · DBLP profile ↗
← Back
34ranked-venue papers
2as first author
21since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 5 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ACJoin: A Low-Latency Multi-Table Join Order Selection Model With Minimum Cost Using Asynchronous Advantage Actor-Critic
abstract
Join order selection is one of the most challenging problems in query optimization and plays an essential role in providing high query performance in Big Data management. Currently, researchers have applied deep reinforcement learning methods, for example, Rejoin and DQ, to join order selection in order to obtain high query performance. However, Rejoin and DQ cannot capture the structural characteristic of the join tree, which may lead to similar encoding structure for different execution plans. To tackle these challenges, we propose a new learning optimizer called ACJoin (asynchronous advantage Actor-Critic for multi-table Join order selection). ACJoin employs a new encoding method to capture the structural characteristics of the join tree through integrating GRU (Gated Recurrent Unit). In particular, ACJoin can distinguish different execution plans. It uses A3C (Asynchronous Advantage Actor-Critic) to guide the join order selection and reduce the time taken to find the best query plan with the minimum cost. Compared with existing search strategies, ACJoin can find the globally optimal solution with efficient and stable query performance. Extensive experiments are conducted on the real JOB and the synthetic TPC-H datasets. The results show that ACJoin outperforms the state-of-the-art join order selection methods and DRL Deep Reinforcement Learning)-based methods in cost and latency.
Shaojie Qiao, Nan Han, Kanglei Xu, Ruiwei Gao, Xindong Wu 0001
IEEE Trans. Big Data3
2026 Where to Go: A Spatial Social Force Graph Neural Network for Predicting Pedestrian Trajectories From Videos With Complex Motion Scenarios
abstract
Traditional pedestrian trajectory prediction models focus on spatio–temporal data without proper consideration of individual interactions with the environment, mutual interactions, and contextual information, resulting in low prediction performance in real applications. In this article, we propose a new pedestrian trajectory prediction model called spatial social force graph neural network (SSF-GNN). First, SSF-GNN adopts a gate recurrent unit (GRU) network and a CenterNet network to capture pedestrian trajectory features and environmental features from historical trajectory sequences. Particularly, SSF-GNN can quantify pedestrian interactions and context-awareness information based on social force. Second, SSF-GNN employs a graph neural network to integrate social influence and hidden states of pedestrians. The distance between adjacent trajectory points is approximated by the weighted average summation of pedestrian historical trajectories. Third, SSF-GNN employs a new interaction function between pedestrians by considering the distance between pedestrians, as well as the movement speed of pedestrians in the social force model, to accurately predict trajectories of pedestrians. Extensive experiments are conducted on two famous datasets, and the results demonstrate SSF-GNN’s outperforms the state-of-the-art models, where average displacement error (ADE) is reduced by more than 25.6%, and final displacement error (FDE) is reduced by more than 15.4%. When predicting a pedestrian’s trajectory in the next eight frames of locations, SSF-GNN outperforms other models significantly with an accuracy of 69.71%.
Shaojie Qiao, Rongmin Tang, Leying Pan, Haosong Gou, Nan Han, Chunfang Yang, Guan Yuan, Tao Wu 0003, Xindong Wu 0001
IEEE Trans. Comput. Soc. Syst.5
2026 CMA+DB: How to Automatically Tune Database Parameters Through Collaborative Multi-Agents
abstract
Database parameter automatic tuning is one of the challenging and difficult tasks that database administrators (DBAs) frequently encounter in artificial intelligence (AI) enabled database (DB) systems. Preferentially optimizing key parameters emerges as a critical point in addressing this issue, and it can help identify important parameters by exploring the interactions between parameters. Aiming to overcome the disadvantages of existing methods, we propose a collaborative multi-agents model called CMA+DB to automatically tune DB parameters in an effective and efficient fashion. CMA+DB integrates three components including SAPM (Single-Agent Pre-trained Model), MATM (Multi-Agent Joint Training Model), and PJTM (Probability-based Joint Training Model). SAPM applies the deep deterministic policy gradient to explore the impact of one single agent on DB performance, MATM uses multi-agent deep deterministic policy gradients to find agents that collaboratively work to improve DB performance, and PJTM can enhance parameter tuning by important agents based on a probabilistic selection factor. In the CMA+DB model, each agent is responsible for tuning a portion of the parameters, and multiple agents collaborate to recommend the optimal parameter configuration. This hybrid model can expand the number of tunable parameters in order to perform parameter tuning from the aspects of functions and parameter levels (i.e., global, DB, and session level). Experimental results reveal that CMA+DB obtains the fastest convergence performance (when reaching the largest throughput) of 14.83% faster than the state-of-the-art (SOTA) algorithms in the TPC-C benchmark on average. Essentially, after the phase of SAPM model training, CMA+DB outperforms the performance of the SOTA models in throughput. Furthermore, DB performance of CMA+DB can be improved by 1.758% through the phases of MATM and PJTM model training.
Shaojie Qiao, Rongmin Tang, Jiangmin Li, Yunjun Gao, Quanqing Xu, Nan Han, Bangping Wang, Guan Yuan, Xindong Wu 0001
IEEE Trans. Knowl. Data Eng.6
2025 Learning database optimization techniques: the state-of-the-art and prospects
abstract
Abstract Artificial intelligence-enabled database technology, known as AI4DB (Artificial Intelligence for Databases), is an active research area attracting significant attention and innovation. This survey first introduces the background of learning-based database techniques. It then reviews advanced query optimization methods for learning databases, focusing on four popular directions: cardinality/cost estimation, learning-based join order selection, learning-based end-to-end optimizers, and text-to-SQL models. Cardinality/cost estimation is classified into supervised and unsupervised methods based on learning models, with illustrative examples provided to explain the working mechanisms. Detailed descriptions of various query optimizers are also given to elucidate the working mechanisms of each component in learning query optimizers. Additionally, we discuss the challenges and development opportunities of learning query optimizers. The survey further explores text-to-SQL models, a new research area within AI4DB. Finally, we consider the future development prospects of learning databases.
Shaojie Qiao, Han-Lin Fan, Nan Han, Yu-Han Peng, Rong-Min Tang, Xiao Qin 0005
Frontiers Comput. Sci.3
2025 LBFL: A Lightweight Blockchain-Based Federated Learning Framework With Proof-of-Contribution Committee Consensus
abstract
Blockchain technology makes it possible to design robust decentralized federated learning (FL). Minimizing the communication cost and storage consumption incurred is one of the essential challenges. In addition, maintaining the security and privacy of Big Data raises to be a difficult problem. Aiming to tackle these challenges, this paper presents LBFL (aLightweightBlockchain-basedFLframework) that offers three novel features. First, it employs a new committee consensus mechanism called Proof-of-Contribution, which is used to avoid the selection latency from the competition of miners and alleviate the congestion in cross-validation of parameters in an asynchronous fashion. Second, LBFL employs a role-adaptive incentive mechanism to estimate devices’ workloads and identify malicious nodes effectively. Third, to cope with the excessive storage overheads incurred in full-replication, LBFL applies a new storage partition mechanism that distributes triple redundant chunks in Reed-Solomon coding (RSC) evenly to participating devices with high fault tolerance and recovery efficiency. To evaluate LBFL, empirical studies are performed on the famousMNISTdataset and LBFL is compared with the state-of-the-art FL frameworks. The results demonstrate that LBFL can reduce evaluation latency and storage consumption by 69.2% and 72.1%, respectively, and the learning efficiency of LBFL is higher than the state-of-the-art methods. In particular, important findings are obtained: the proposed role-adaptive incentive mechanism can properly identify malicious devices and switch the roles of legitimate devices to achieve good decentralization.
Shaojie Qiao, Yuhe Jiang, Nan Han, Yufeng Lin, Shengjie Min, Xindong Wu 0001
IEEE Trans. Big Data3
2024 REXIO: Indexing for Low Write Amplification by Reducing Extra I/Os in Key-Value Store Under Mixed Read/Write Workloads
Qiang Qu 0001, Nan Han, Zhelang Deng, Yizhuo Ma, Jintao Meng 0001
WISE (1)3
2024 A three-in-one dynamic shared bicycle demand forecasting model under non-classical conditions
Shaojie Qiao, Nan Han, He Li 0006, Guan Yuan, Tao Wu 0003, Yuzhong Peng, Hongguo Cai, Jiangtao Huang
Appl. Intell.2
2024 GTR: An SQL Generator With Transition Representation in Cross-Domain Database Systems
abstract
Recent studies have focused on using natural language (NL) to automatically retrieve useful data from database (DB) systems. As an important component of autonomous DB systems, the NL-to-SQL technique can assist DB administrators in writing high-quality SQL statements and make persons with no SQL background knowledge learn complex SQL languages. However, existing studies cannot deal with the issue that the expression of NL inevitably mismatches the implementation details of SQLs, and the large number of out-of-domain (OOD) words makes it difficult to predict table columns. In particular, it is difficult to accurately convert NL into SQL in an end-to-end fashion. Intuitively, it facilitates the model to understand the relations if a "bridge" [transition representation (TR)] is employed to make it compatible with both NL and SQL in the phase of conversion. In this article, we propose an automatic SQL generator with TR called GTR in cross-domain DB systems. Specifically, GTR contains three SQL generation steps: 1) GTR learns the relation between questions and DB schemas; 2) GTR uses a grammar-based model to synthesize a TR; and 3) GTR predicts SQL from TR based on the rules. We conduct extensive experiments on two commonly used datasets, that is, WikiSQL and Spider. On the testing set of the Spider and WikiSQL datasets, the results show that GTR achieves 58.32% and 71.29% exact matching accuracy which outperforms the state-of-the-art methods, respectively.
Shaojie Qiao, Nan Han, Yuhan Peng, Lingchun Wu, He Li 0006, Guan Yuan
IEEE Trans. Neural Networks Learn. Syst.4
2023 Imbalanced data classification: Using transfer learning and active sampling
Shaojie Qiao, Meiqi Liu, Lulu Qu, Nan Han, Guan Yuan, Tao Wu 0003, Yuzhong Peng
Eng. Appl. Artif. Intell.6
2022 LMNNB: Two-in-One imbalanced classification approach by combining metric learning and ensemble learning
Shaojie Qiao, Nan Han, Faliang Huang, Kun Yue, Tao Wu 0003, Yugen Yi, Rui Mao 0001, Chang-an Yuan 0001
Appl. Intell.2
2022 Geometric correction method for Tibetan woodcut document images
Lijia Xiawu, Shaojie Qiao, Wenrong Tan, Tao Liu 0027, Nan Han
Multim. Tools Appl.7
2022 Correction to: Geometric correction method for Tibetan woodcut document images
Lijia Xiawu, Shaojie Qiao, Wenrong Tan, Tao Liu 0027, Nan Han
Multim. Tools Appl.7
2022 Affective Impression: Sentiment-Awareness POI Suggestion via Embedding in Heterogeneous LBSNs
abstract
Location-based social networks (LBSNs) add geographical information into traditional social networks and link people’s virtual and physical lives. As an important application of LBSNs, point-of-interest (POI) suggestion has become an important method to help users explore interesting and attractive locations in LBSNs. The main problems of POI suggestion include data sparsity and cold start, which have been paid much attention by existing techniques. There are two major challenges which can greatly influence the performance of suggestion accuracy. One is the fuzzy boundary between sentiments, i.e., the fine distinction between sentiments makes it difficult to classify words and texts after word-sentiment mapping operation. The other challenge is the unreliability of data quality represented by similarity metrics, which relies on data integrity and path reachability of a heterogeneous network. To cope with the above two challenges, we present a novel framework calledCommunity-based SentimentExtraction andNetworkEmbedding for POIRecommendation (CENTER) for suggesting impressive POIs to a specific user in an effective fashion. The CENTER framework contains two essential techniques: (1) a latent probabilistic generative model calledCommunity-basedSentimentExtraction (CSE), which can accurately capture the sentiments from review content in LBSNs by taking into consideration the characteristics of social communities. The parameters of the CSE model can be inferred effectively by the Gibbs sampling method. The primary sentiments are obtained based on the distribution of sentiments; (2) a network embedding model calledSentiment-awareNeworkEmbedding for POIRecommendation (SNER) is employed to learn the representation of the factors including POIs, users and textual sentiments in a low-dimensional embedding space. The joint training is utilized to alternatively sample all sets of edges in a heterogeneous information network. Extensive experiments were conducted on two large-scale real datasets, in order to evaluate the performance of the proposed CENTER framework. The results demonstrate that CENTER is superior to the state-of-the-art baseline methods in the effectiveness and efficiency of POI suggestion.
Shaojie Qiao, Nan Han, Ling He 0003
IEEE Trans. Affect. Comput.3
2022 Algorithms for Trajectory Points Clustering in Location-based Social Networks
abstract
Recent advances in localization techniques have fundamentally enhanced social networking services, allowing users to share their locations and location-related contents. This has further increased the popularity of location-based social networks (LBSNs) and produces a huge amount of trajectories composed of continuous and complex spatio-temporal points from people’s daily lives. How to accurately aggregate large-scale trajectories is an important and challenging task. Conventional clustering algorithms (e.g., k -means or k -mediods) cannot be directly employed to process trajectory data due to their serialization, triviality and redundancy. Aiming to overcome the drawbacks of traditional k -means algorithm and k -mediods, including their sensitivity to the selection of the initial k value, the cluster centers and easy convergence to a locally optimal solution, we first propose an optimized k -means algorithm (namely OKM ) to obtain k optimal initial clustering centers based on the density of trajectory points. Second, because k -means is sensitive to noisy points, we propose an improved k -mediods algorithm called IKMD based on an acceptable radius r by considering users’ geographic location in LBSNs. The value of k can be calculated based on r , and the optimal k points are selected as the initial clustering centers with high densities to reduce the cost of distance calculation. Thirdly, we thoroughly analyze the advantages of IKMD by comparing it with the commonly used clustering approaches through illustrative examples. Last, we conduct extensive experiments to evaluate the performance of IKMD against seven clustering approaches including the proposed optimized k -means algorithm, k -mediods algorithm, traditional density-based k -mediods algorithm and the state-of-the-arts trajectory clustering methods. The results demonstrate that IKMD significantly outperforms existing algorithms in the cost of distance calculation and the convergence speed. The methods proposed is proved to contribute to a larger effort targeted at advancing the study of intelligent trajectory data analytics.
Nan Han, Shaojie Qiao, Kun Yue, Qiang He 0001, Tingting Tang, Faliang Huang, Chang-an Yuan 0001
ACM Trans. Intell. Syst. Technol.1
2021 CAE-CNN: Predicting transcription factor binding site with convolutional autoencoder and convolutional neural network
Yongqing Zhang 0001, Shaojie Qiao, Yuanqi Zeng, Dongrui Gao, Nan Han, Jiliu Zhou
Expert Syst. Appl.5
2021 Cardinality Estimator: Processing SQL with a Vertical Scanning Convolutional Neural Network
Shaojie Qiao, Nan Han, Faliang Huang, Kun Yue, Yugen Yi, Chang-an Yuan 0001
J. Comput. Sci. Technol.3
2021 Algorithm for detecting anomalous hosts based on group activity evolution
Xiaoming Ye, Shaojie Qiao, Nan Han, Kun Yue, Tao Wu 0003, Faliang Huang, Chang-an Yuan 0001
Knowl. Based Syst.3
2021 Which, Who and How: Detecting fraudulent sale accumulation behavior from multi-dimensional sparse data
Jiaoling Zheng, Shaojie Qiao, Nan Han, Kun Yue, Qiang He 0001, Shengjie Min, Guanghua Ying, Xindong Wu 0001
Knowl. Based Syst.3
2021 A Dynamic Convolutional Neural Network Based Shared-Bike Demand Forecasting Model
abstract
Bike-sharing systems are becoming popular and generate a large volume of trajectory data. In a bike-sharing system, users can borrow and return bikes at different stations. In particular, a bike-sharing system will be affected by weather, the time period, and other dynamic factors, which challenges the scheduling of shared bikes. In this article, a new shared-bike demand forecasting model based on dynamic convolutional neural networks, called SDF , is proposed to predict the demand of shared bikes. SDF chooses the most relevant weather features from real weather data by using the Pearson correlation coefficient and transforms them into a two-dimensional dynamic feature matrix, taking into account the states of stations from historical data. The feature information in the matrix is extracted, learned, and trained with a newly proposed dynamic convolutional neural network to predict the demand of shared bikes in a dynamical and intelligent fashion. The phase of parameter update is optimized from three aspects: the loss function, optimization algorithm, and learning rate. Then, an accurate shared-bike demand forecasting model is designed based on the basic idea of minimizing the loss value. By comparing with classical machine learning models, the weight sharing strategy employed by SDF reduces the complexity of the network. It allows a high prediction accuracy to be achieved within a relatively short period of time. Extensive experiments are conducted on real-world bike-sharing datasets to evaluate SDF. The results show that SDF significantly outperforms classical machine learning models in prediction accuracy and efficiency.
Shaojie Qiao, Nan Han, Kun Yue, Rui Mao 0001, Hongping Shu, Qiang He 0001, Xindong Wu 0001
ACM Trans. Intell. Syst. Technol.2
2021 Dynamic Community Evolution Analysis Framework for Large-Scale Complex Networks Based on Strong and Weak Events
abstract
Community evolution remains a heavily researched and challenging area in the analysis of dynamic complex network structures. Currently, the primary limitation of traditional event-based approaches for community evolution analysis is the lack of strict constraint conditions for distinguishing evolutionary events, which entails that as the cardinality of discovered events increases, so does the number of redundant events. Another limitation of existing approaches is the lack of consideration for weak events. Weak events can be generated by small changes in communities, which are empirically prevalent, and are typically not captured by traditional events. To manage these two aforementioned limitations, this research aims to formalize a weak and strong events-based framework, which includes the following newly discovered events: “weak shrink,” “weak expand,” “weak merge,” and “weak splity” predicated on the community overlapping degree and community degree membership, this article refines these traditional strong events, as well as new constraints for weak events. In addition, a community evolution mining framework, which is based on both strong and weak events, is proposed and denoted by a weak-event-based community evolution method (WECEM). The framework can be summarized by the following: 1) communities in complex networks with adjacent time-stamps are compared to determine the community overlapping degree and community membership degree; 2) the values of the community overlapping degree and membership degree meet the definition of events; and 3) weak events are effectively identified. Extensive experimental results, on real and synthetic data sets consisting of dynamic complex networks and online social networks, demonstrate that WECEM is able to identify weak events more effectively than traditional frameworks. Specifically, WECEM outperforms traditional frameworks by 22.9% in the number of discovered strong events. The detection accuracy of evolutionary events is approximately 12.2% higher than that of traditional event-based frameworks. It is also worth noting that, as the cardinality of the data grows, the proposed framework, when compared with traditional frameworks, can more effectively, and efficiently, mine large-scale complex networks.
Shaojie Qiao, Nan Han, Yunjun Gao, Rong-Hua Li 0001, Heli Sun, Xindong Wu 0001
IEEE Trans. Syst. Man Cybern. Syst.2
2021 Robust Graph Structure Learning for Multimedia Data Analysis
abstract
With the rapid development of computer network technology, we can acquire a large amount of multimedia data, and it becomes a very important task to analyze these data. Since graph construction or graph learning is a powerful tool for multimedia data analysis, many graph‐based subspace learning and clustering approaches have been proposed. Among the existing graph learning algorithms, the sample reconstruction‐based approaches have gone the mainstream. Nevertheless, these approaches not only ignore the local and global structure information but also are sensitive to noise. To address these limitations, this paper proposes a graph learning framework, termed Robust Graph Structure Learning (RGSL). Different from the existing graph learning approaches, our approach adopts the self‐expressiveness of samples to capture the global structure, meanwhile utilizing data locality to depict the local structure. Specially, in order to improve the robustness of our approach against noise, we introduce l2,1‐norm regularization criterion and nonnegative constraint into the graph construction process. Furthermore, an iterative updating optimization algorithm is designed to solve the objective function. A large number of subspace learning and clustering experiments are carried out to verify the effectiveness of the proposed approach.
Wei Zhou 0003, Zhaoxuan Gong, Wei Guo 0016, Nan Han, Shaojie Qiao
Wirel. Commun. Mob. Comput.4
2020 A point-of-interest suggestion algorithm in Multi-source geo-social networks
Shaojie Qiao, Nan Han, Guan Yuan, Yongqing Zhang 0001
Eng. Appl. Artif. Intell.4
2020 Where to go: An effective point-of-interest recommendation framework for heterogeneous social networks
Shaojie Qiao, Nan Han, Zhan Bu, Rong-Hua Li 0001, Kun Yue, Guan Yuan
Neurocomputing3
2019 An effective image classification method for shallow densely connected convolution networks through squeezing and splitting techniques
Chang-an Yuan 0001, Yong Wu 0006, Xiao Qin 0005, Shaojie Qiao, Yonghua Pan, Dunhu Liu, Nan Han
Appl. Intell.8
2019 A novel Chinese herbal medicine clustering algorithm via artificial bee colony optimization
Nan Han, Shaojie Qiao, Guan Yuan, Dingxiang Liu, Kun Yue
Artif. Intell. Medicine1
2019 How to balance the bioinformatics data: pseudo-negative sampling
abstract
BACKGROUND: Imbalanced datasets are commonly encountered in bioinformatics classification problems, that is, the number of negative samples is much larger than that of positive samples. Particularly, the data imbalance phenomena will make us underestimate the performance of the minority class of positive samples. Therefore, how to balance the bioinformatic data becomes a very challenging and difficult problem. RESULTS: In this study, we propose a new data sampling approach, called pseudo-negative sampling, which can be effectively applied to handle the case that: negative samples greatly dominate positive samples. Specifically, we design a supervised learning method based on a max-relevance min-redundancy criterion beyond Pearson correlation coefficient (MMPCC), which is used to choose pseudo-negative samples from the negative samples and view them as positive samples. In addition, MMPCC uses an incremental searching technique to select optimal pseudo-negative samples to reduce the computation cost. Consequently, the discovered pseudo-negative samples have strong relevance to positive samples and less redundancy to negative ones. CONCLUSIONS: To validate the performance of our method, we conduct experiments base on four UCI datasets and three real bioinformatics datasets. According to the experimental results, we clearly observe the performance of MMPCC is better than other sampling methods in terms of Sensitivity, Specificity, Accuracy and the Mathew's Correlation Coefficient. This reveals that the pseudo-negative samples are particularly helpful to solve the imbalance dataset problem. Moreover, the gain of Sensitivity from the minority samples with pseudo-negative samples grows with the improvement of prediction accuracy on all dataset.
Yongqing Zhang 0001, Shaojie Qiao, Rongzhao Lu, Nan Han, Dingxiang Liu, Jiliu Zhou
BMC Bioinform.4
2019 Identification of DNA-protein binding sites by bootstrap multiple convolutional neural networks on sequence information
Yongqing Zhang 0001, Shaojie Qiao, Shengjie Ji, Nan Han, Dingxiang Liu, Jiliu Zhou
Eng. Appl. Artif. Intell.4
2019 ADPDF: A Hybrid Attribute Discrimination Method for Psychometric Data With Fuzziness
abstract
The existing approaches for attribute discrimination are applied to clinical data with unambiguous boundaries, and rarely take into careful consideration on how to utilize psychometric data with fuzziness. In addition, it is difficult for conventional attribute reduction methods to reduce attributes of psychometric data which are composed of a lot of attributes and contain a relatively small-scale samples. Importantly, these methods cannot be used to reduce options which are relevant to each other. In this paper, we first introduce new concepts, that is, option entropy and option influence degree, which are employed to describe the relation and distribution of options. Then, we propose a hybrid attribute discrimination method for psychometric data with fuzziness, called a hybrid attribute discrimination for psychometric data with fuzziness (ADPDF). ADPDF contains three essential techniques: 1) a fuzzy option reduction method, which aims to combine a fuzzy option to adjacent options, and is used to reduce the fuzziness of options in a psychometry and 2) k -fold attribute reduction method, which partitions all samples into several subsets and negotiates the reduction results of different subsets, and reduces the noise for the purpose of accurately discovering key attributes. In order to show the advantages of the proposed approach, we conducted experiments on two real datasets collected from clinical diagnoses. The experimental results show that the proposed method can decrease the correlation between options effectively. Interestingly, we find three reserved options and one hundred samples in each subset show the best classification performance. Finally, we compare the proposed method with typical attribute discrimination algorithms. The results reveal that our method can improve the classification accuracy with the guarantee of time performance.
Shaojie Qiao, Haiqing Zhang, Nan Han, Rong-Hua Li 0001
IEEE Trans. Syst. Man Cybern. Syst.6
2018 SocialMix: A familiarity-based and preference-aware location suggestion approach
Shaojie Qiao, Nan Han, Jiliu Zhou, Rong-Hua Li 0001, Cheqing Jin, Louis Alberto Gutierrez
Eng. Appl. Artif. Intell.2
2018 Predicting Long-Term Trajectories of Connected Vehicles via the Prefix-Projection Technique
abstract
The vehicle location prediction based on their spatial and temporal information is an important and difficult task in many applications. In the last few years, devices, such as connected vehicles, smart phones, GPS navigation systems, and smart home appliances, have amassed the large stores of geographic data. The task of leveraging this data by employing moving objects database techniques to predict spatio-temporal locations in an accurate and efficient fashion, comprising a complete trajectory remains an actively researched area. Existing methods for frequent sequential pattern mining tend to be limited to predicting short-term partial trajectories, at extremely high computational costs. In order to address these limitations, we designed a prefix-projection-based trajectory prediction algorithm called PrefixTP, which contains three essential phases. First, data collection, connected vehicles equipped with sensors comprise a vehicle grid and generate copious amounts of spatio-temporal data, in order to communicate and share traffic information. Second, model training, examining only the prefix subsequences, and projecting only their corresponding postfix subsequences into projected sets. Finally, trajectory matching, recursively finding postfix sequences meeting the requirement of minimum support count, and outputting the most frequent sequential pattern as the most probable trajectory. Fundamentally, PrefixTP supports three trajectory matching strategies which encompass all possibilities of prediction. Extensive experiments were conducted using real world GPS data sets, and the results show, when comparing predicted complete trajectories against partial short-term trajectories with a guarantee of real-time forecasting, that PrefixTP outperforms first-order, second-order Markov models, and Apriori-based trajectory prediction algorithm.
Shaojie Qiao, Nan Han, Junfeng Wang 0003, Rong-Hua Li 0001, Louis Alberto Gutierrez, Xindong Wu 0001
IEEE Trans. Intell. Transp. Syst.2
2018 A Fast Parallel Community Discovery Model on Complex Networks Through Approximate Optimization
abstract
Community discovery plays an essential role in the analysis of the structural features of complex networks. Since online networks grow increasingly large and complex over time, the methods traditionally used for community discovery cannot efficiently handle large-scale network data. This introduces the important problem of how to effectively and efficiently discover large communities from complex networks. In this study, we propose a fast parallel community discovery model called picaso (a parallel community discovery algorithm based on approximate optimization), which integrates two new techniques: (1) Mountain model, which works by utilizing graph theory to approximate the selection of nodes needed for merging, and (2) Landslide algorithm, which is used to update the modularity increment based on the approximated optimization. In addition, the GraphX distribution computing framework is employed in order to achieve parallel community detection over complex networks. In the proposed model, clustering on modularity is used to initialize the Mountain model as well as to compute the weight of each edge in the networks. The relationships among the communities are then simplified by applying the Landslide algorithm, which allows us to obtain the community structures of the complex networks. Extensive experiments were conducted on real and synthetic complex network datasets, and the results demonstrate that the proposed algorithm can outperform the state of the art methods, in effectiveness and efficiency, when working to solve the problem of community detection. Moreover, we demonstratively prove that overall time performance approximates to four times faster than similar approaches. Effectively our results suggest a new paradigm for large-scale community discovery of complex networks.
Shaojie Qiao, Nan Han, Yunjun Gao, Rong-Hua Li 0001, Louis Alberto Gutierrez, Xindong Wu 0001
IEEE Trans. Knowl. Data Eng.2
2015 TraPlan: An Effective Three-in-One Trajectory-Prediction Model in Transportation Networks
abstract
The existing approaches for trajectory prediction (TP) are primarily concerned with discovering frequent trajectory patterns (FTPs) from historical movement data. Moreover, most of these approaches work by using a linear TP model to depict the positions of objects, which does not lend itself to the complexities of most real-world applications. In this research, we propose a three-in-one TP model in road-constrained transportation networks called TraPlan. TraPlan contains three essential techniques: 1) constrained network R-tree (CNR-tree), which is a two-tiered dynamic index structure of moving objects based on transportation networks; 2) a region-of-interest (RoI) discovery algorithm is employed to partition a large number of trajectory points into distinct clusters; and 3) a FTP-tree-based TP approach, called FTP-mining, is proposed to discover FTPs to infer future locations of objects moving within RoIs. In order to evaluate the results of the proposed CNR-tree index structure, we conducted experiments on synthetically generated data sets taken from real-world transportation networks. The results show that the CNR-tree can reduce the time cost of index maintenance by an average gap of about 40% when compared with the traditional NDTR-tree, as well as reduce the time cost of trajectory queries. Moreover, compared with fixed network R-Tree (FNR-trees), the accuracy of range queries has shown an on average improvement of about 32%. Furthermore, the experimental results show that the TraPlan demonstrates accurate and efficient prediction of possible motion curves of objects in distinct trajectory data sets by over 80% on average. Finally, we evaluate these results and the performance of the TraPlan model in regard to TP by comparing it with other TP algorithms.
Shaojie Qiao, Nan Han, William Zhu 0001, Louis Alberto Gutierrez
IEEE Trans. Intell. Transp. Syst.2
2015 A Self-Adaptive Parameter Selection Trajectory Prediction Approach via Hidden Markov Models
abstract
Trajectory prediction of objects in moving objects databases (MODs) has garnered wide support in a variety of applications and is gradually becoming an active research area. The existing trajectory prediction algorithms focus on discovering frequent moving patterns or simulating the mobility of objects via mathematical models. While these models are useful in certain applications, they fall short in describing the position and behavior of moving objects in a network-constraint environment. Aiming to solve this problem, a hidden Markov model (HMM)-based trajectory prediction algorithm is proposed, called Hidden Markov model-based Trajectory Prediction (HMTP). By analyzing the disadvantages of HMTP, a self-adaptive parameter selection algorithm called HMTP* is proposed, which captures the parameters necessary for real-world scenarios in terms of objects with dynamically changing speed. In addition, a density-based trajectory partition algorithm is introduced, which helps improve the efficiency of prediction. In order to evaluate the effectiveness and efficiency of the proposed algorithms, extensive experiments were conducted, and the experimental results demonstrate that the effect of critical parameters on the prediction accuracy in the proposed paradigm, with regard to HMTP*, can greatly improve the accuracy when compared with HMTP, when subjected to randomly changing speeds. Moreover, it has higher positioning precision than HMTP due to its capability of self-adjustment.
Shaojie Qiao, Dayong Shen, Xiaoteng Wang, Nan Han, William Zhu 0001
IEEE Trans. Intell. Transp. Syst.4
2010 KISTCM: knowledge discovery system for traditional Chinese medicine
Shaojie Qiao, Changjie Tang, Huidong Jin 0001, Jing Peng 0002, Darren Davis, Nan Han
Appl. Intell.6