EDBT 2026 Demo / reviewers in the wild / expert
Jiajie Xu 0001
dblp:20/1590 · also Jia-Jie Xu 0001
· DBLP profile ↗
100ranked-venue papers in the field
7as first author
46since 2021 · last 2026
0000-0001-8227-8636ORCID · conflict
Domains — venue-derived; a paper can count in several
Database Systems & Data Management · 56 (6 first)Information Retrieval & Web Search · 31Data Mining & Knowledge Discovery · 7 (1 first)Other / Interdisciplinary · 4Big Data, Cloud & Distributed Data Systems · 1Knowledge Engineering, Semantic Web & Information Systems · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Zero-Shot Cellular Trajectory Map MatchingabstractCellular Trajectory Map-Matching (CTMM) aims to align cellular location sequences to road networks, which is a necessary preprocessing in location-based services on web platforms like Google Maps, including navigation and route optimization. Current approaches mainly rely on ID-based features and region-specific data to learn correlations between cell towers and roads, limiting their adaptability to unexplored areas. To enable high-accuracy CTMM without additional training in target regions, Zero-shot CTMM requires to extract not only region-adaptive features, but also sequential and location uncertainty to alleviate positioning errors in cellular data. In this paper, we propose a pixel-based trajectory calibration assistant for zero-shot CTMM, which takes advantage of transferable geospatial knowledge to calibrate pixelated trajectory, and then guide the path-finding process at the road network level. To enhance knowledge sharing across similar regions, a Gaussian mixture model is incorporated into VAE, enabling the identification of scenario-adaptive experts through soft clustering. To mitigate high positioning errors, a spatial-temporal awareness module is designed to capture sequential features and location uncertainty, thereby facilitating the inference of approximate user positions. Finally, a constrained path-finding algorithm is employed to reconstruct the road ID sequence, ensuring topological validity within the road network. This process is guided by the calibrated trajectory while optimizing for the shortest feasible path, thus minimizing unnecessary detours. Extensive experiments demonstrate that our model outperforms existing methods in zero-shot CTMM by 16.8\%. Yue Cui 0001, Mengze Li 0001, Jia Zhu 0003, Jiajie Xu 0001, Xiaofang Zhou 0001 |
IEEE Trans. Knowl. Data Eng. | 7 |
| 2025 | SWIFT: Scene-Aware Dual Cross-Attention Flight Trajectory Prediction
Pingfu Chao, Junhua Fang, Jiajie Xu 0001 |
IEEE Big Data | 5 |
| 2025 | TKHist: Cardinality Estimation for Join Queries via Histograms with Dominant Attribute Correlation FindingabstractCardinality estimation has long been crucial for cost-based database optimizers in identifying optimal query execution plans, attracting significant attention over the past decades. While recent advancements have significantly improved the accuracy of multi-table join query estimations, these methods introduce challenges such as higher space overhead, increased latency, and greater complexity, especially when integrated with the binary join framework. In this paper, we introduce a novel cardinality estimation method named TKHist, which addresses these challenges by relaxing the uniformity assumption in histograms. TKHist captures bin-wise non-uniformity information, enabling accurate cardinality estimation for join queries without filter predicates. Furthermore, we explore the attribute independent assumption, which can lead to significant over-estimation rather than under-estimation in multi-table join queries. To address this issue, we propose the dominating join path correlation discovery algorithm to highlight and manage correlations between join keys and filter predicates. Our extensive experiments on popular benchmarks demonstrate that TKHist reduces error variance by 2-3 orders of magnitude compared to SOTA methods, while maintaining comparable or lower memory usage. Renrui Li, Qingzhi Ma, Jiajie Xu 0001, Lei Zhao 0001, An Liu 0002 |
CIKM | 3 |
| 2025 | CF-TS: A General Coarse-to-Fine Method for Trajectory Simplification
Junhua Fang, Pingfu Chao, Jiajie Xu 0001, Pengpeng Zhao 0001 |
DASFAA (1) | 5 |
| 2025 | LODC: A Lightweight Online Update Method for Density-Based Clustering
Jiajie Xu 0001, Junhua Fang, Pingfu Chao, Pengpeng Zhao 0001, An Liu 0002 |
DASFAA (1) | 1 |
| 2025 | Enhancing Large-Scale Entity Alignment with Critical Structure and High-Quality ContextabstractEntity Alignment (EA) aims to identify equivalent entities across multiple Knowledge Graphs (KGs). However, when applied to larger-scale KGs, most existing EA approaches suffer from the scalability issue due to excessive GPU memory and time consumption. To mitigate this, recent advances have introduced the Large-scale EA (LsEA) task, which divides large-scale KG pairs into smaller sub-graph pairs. Despite their promising results, several notable challenges remain, preventing these advances from achieving optimal performance: 1) How to effectively utilize critical structures when generating sub-tasks? 2) How to supplement high-quality context to enhance LsEA performance? 3) How to address scenarios without alignment seeds? To tackle these challenges, we propose a novel method called ELsEA. It comprises three main components: (1) Source and Target Graph Partition, using a Metis-based weighted partitioner and a counter-part candidate generator to partition source and target graphs respectively, aiming to utilize critical structures effectively; (2) Supplement High-quality Context, which utilizes a value-based informativeness-evaluation module and a neighbor enrichment module to assess each entity's informativeness effectively, then supplement high-quality context based on this informativeness; and (3) Seed-free Setup, introducing a mixed-info pseudo-seed generation strategy to mitigate name bias, generating accurate pseudo-seeds when alignment seeds are unavailable. Extensive experiments demonstrate that ELsEA outperforms state-of-the-art baselines. The code of ELsEA is available online11https://githuh.com/wx-qzhou/ELsEA.git. Wei Chen 0070, Li Zhang 0004, Pengpeng Zhao 0001, Jiajie Xu 0001, Lei Zhao 0001 |
ICDE | 5 |
| 2025 | Disentangled Graph Debiasing for Next POI RecommendationabstractGraph neural networks play a pivotal role in various location-based applications, showcasing their compelling ability to capture collaborative signals across user check-in sequences. Recent advancements in next POI recommendation have further leveraged spatio-temporal graphs to uncover the transitional and geographical regularities. However, these methods are usually vulnerable due to the presence of data biases in real-life scenarios, which may mislead the model to disproportionately favoring certain POIs. To this end, this paper proposes a new graph debiasing paradigm for POI recommendation, which disentangles causal and bias knowledge within spatio-temporal graphs, allowing for not only the mitigation of bias issues, but also the utilization of causal information from spatial and temporal perspectives. Specifically, to facilitate graph debiasing at its topological level, an adaptive edge mask generator is first designed to explicitly decompose an entangled graph into causal and bias subgraphs. We encourage the stable relationships between the causal subgraph and the prediction, while the bias subgraph targets at the skewed bias distribution. We further enhance the independence between such two parts by employing a causal-bias disagreement regularization to encourage their distribution in separate semantic spaces. In addition, an inter-view contrastive learning module is also applied to maintain the relation discriminability of transitional and geographical representations. Extensive experiments on three real-world datasets demonstrate the superiority of our proposed model on recommendation performance, as well as its robustness against data bias. Hailun Zhou, Jiajie Xu 0001, Qiaoming Zhu, Chengfei Liu |
SIGIR | 2 |
| 2025 | An efficient distributed co-movement pattern detection framework for streaming trajectory
Tong Cheng, Pingfu Chao, Kenan Zhang, Junhua Fang, Jiajie Xu 0001 |
Knowl. Inf. Syst. | 5 |
| 2025 | TMLKD: Few-shot Trajectory Metric Learning via Knowledge DistillationabstractTrajectory metric learning, which supports the trajectory similarity search, is one of the most fundamental tasks in spatial-temporal data analysis. However, existing trajectory metric learning methods rely on massive labels of pairwise trajectory distance, and thus cannot be applied to few-shot scenarios frequently occurring in real-world applications. Though performance drops caused by insufficient labels can be alleviated by knowledge distillation, we demonstrate that they cannot be directly applied to few-shot trajectory metric learning due to the domain shift problem. To this end, this paper proposes invariant and relaxed learning enhanced knowledge distillation method TMLKD for few-shot trajectory metric learning, such that domain-invariant representation and rank knowledge can be distilled. Specifically, in the representation learning phase, it first employs an adversarial sub-network to distinguish domain-specific and domain-invariant information, so as to distill transferable representation knowledge from teacher models. To mitigate the few-shot problem in student model training, we further enrich sparse labels of the target domain by utilizing the rank knowledge revealed in teachers' predictions. Particularly, TMLKD employs a list-wise learning-to-rank approach to learn the relaxed trajectory ranking orders instead of focusing on all the samples inefficiently. Finally, to guide accurate distillation, we adaptively assign reliability of teacher prediction by utilizing the ground-truth labels, to avoid misleading the student model with low-quality teacher predictions. Extensive experiments on three real-world datasets demonstrate the superiority of our model. Danling Lai, Jiajie Xu 0001, Jianfeng Qu, Pingfu Chao, Junhua Fang, Chengfei Liu |
Proc. VLDB Endow. | 2 |
| 2025 | Towards DS-NER: Unveiling and Addressing Latent Noise in Distant AnnotationsabstractDistantly supervised named entity recognition (DS-NER) has emerged as a cheap and convenient alternative to traditional human annotation methods, enabling the automatic generation of training data by aligning text with external resources. Despite the many efforts in noise measurement methods, few works focus on the latent noise distribution between different distant annotation methods. In this work, we explore the effectiveness and robustness of DS-NER by two aspects: (1) distant annotation techniques, which encompasses both traditional rule-based methods and the innovative large language model supervision approach, and (2) noise assessment, for which we introduce a novel framework. This framework addresses the challenges by distinctly categorizing them into theunlabeled-entity problem (UEP)and thenoisy-entity problem (NEP), subsequently providing specialized solutions for each. Our proposed method achieves significant improvements on eight real-world distant supervision datasets originating from three different data sources and involving four distinct annotation techniques, confirming its superiority over current state-of-the-art methods. Yuyang Ding, Juntao Li 0005, Jiajie Xu 0001, Pingfu Chao, Xiaofang Zhou 0001, Min Zhang 0005 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2025 | Contrastive Variational Group Recommendation With Data-Agnostic AugmentationabstractGroup recommendation aims to recommend desired items for a group of users. Existing methods mainly adopt deterministic networks to represent groups as fixed-point vectors, assuming their preferences be highly close to these vectors in interest space. However, each group tends to have various interests, which cannot be fully captured by fixed-point vectors and thus calls for probabilistic modeling of interests as density instead. Although this can be supported by Variational AutoEncoder (VAE), interaction data in group recommendation are highly sparse and insufficient for VAE model training, resulting in high risks of posterior collapse and deficiency in personalization. To this end, this paper proposes a contrastive variational learning model boosted by variational model augmentation and an easyto-hard paradigm. Specifically, VAE with tailored attention is first employed to represent group preferences as variational vectors for probabilistic preference modeling. Additionally, we conduct data-agnostic augmentation via learnable variational dropout, which removes redundant or irrelevant neurons in VAE to generate meaningful augmented views adequately for contrastive learning in spite of data sparsity. Difficulty-aware negative sampling is further applied to generate high-quality negative samples adapting to varying requirements of task difficulty according to the training process. Finally, we utilize density-based variational alignment to guide the optimization process of contrastive learning. Experiments on four real-world datasets are conducted to demonstrate the significant performance improvements of our model compared with SOTA methods for group recommendation. Wen Yang 0018, Jiajie Xu 0001, Rui Zhou 0001, Lu Chen 0008, Jianxin Li 0001, Pengpeng Zhao 0001, Chengfei Liu |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2024 | Ocean: Online Clustering and Evolution Analysis for Dynamic Streaming DataabstractWith the popularization of mobile applications and the timely acquisition of fresh data, real-time clustering and its evolution analysis have become the primary operations for data processing and knowledge discovery. Such continuous queries on massive objects are computation-intensive tasks in dynamic scenarios. However, existing clustering techniques are incompetent to achieve decent performance when computation-intensive operations frequently occur in streaming scenarios, which is caused by two challenges: (i) uncertainty of the clustering frequency; (ii) unpredictable distribution evolution. Hence, it is critical to find a lightweight model that can cluster the high-speed dynamic instances while exploiting the evolution amid different clustering results. This paper focuses on the problem of real-time clustering on streaming data in computation-intensive and high-dynamics tasks, through a framework Ocean, consisting of the Online clustering algorithm and evolution analysis. Particularly, the framework conceives a flexible composite window to augment the knowledge mining, achieving a proper real-time response in various scenarios. The evolution analysis supports full life-cycle detection, improving the adaptability to dynamic concept drifts and multiple patterns. Inspired by the grid partition strategy, this framework adopts grid feature vectors to capture the significant changes in streaming data. Furthermore, we propose an optimization that removes sparse grids timely and performs the online clustering adaptively for space and time efficiency. It is proven to be effective both theoretically and experimentally. This strategy enables real-time clustering for dynamic streaming data without degrading the clustering quality or increasing the computation cost. Experiments on real datasets and synthetic datasets verify the accuracy and effectiveness of Ocean compared to the state-of-the-art approaches, as well as the superior ability to perform clustering in a real-time manner. Chunhui Feng, Junhua Fang, Yue Xia, Pingfu Chao, Pengpeng Zhao 0001, Jiajie Xu 0001, Xiaofang Zhou 0001 |
ICDE | 6 |
| 2024 | Multi-view Attentive Variational Learning for Group RecommendationabstractGroup recommendation aims to recommend desired items for a group of users. Due to the sparsity of group-item interactions, existing methods mainly model group preferences by aggregating member-level preference. However, they not only ignore possible user interest drift in specific groups, but also adopt deterministic models to represent group preferences using fixed-points, which are weak in characterizing uncertain group preferences. To this end, following the paradigm of variational learning, this paper proposes a multi-view attentive variational preference aggregation network called GroupAV for group rec-ommendation, so as to conduct user/group preference modeling and aggregation in a density-based manner. Specifically, we first adopt Variational AutoEncoder (VAE) to capture member-level preferences by variational vectors as density. To address user interest drift in groups, a variational preference adapter module is designed to learn group-contextualized preferences via rational transformation in variational space. Next, attentive variational aggregation networks are carefully designed for group-level preference aggregation in two different views (i.e., group-interactions and member-consensus views). Besides, we apply contrastive learning and gating fusion to optimize the multi-view learning process for the final group preference modeling of Mixture-of-Gaussian distribution. Finally, we conduct experiments on real-world datasets and demonstrate GroupAV's significant performance improvements compared to state-of-the-art group recommendation methods. Wen Yang 0018, Jiajie Xu 0001, Rui Zhou 0001, Lu Chen 0008, Jianxin Li 0001, Pengpeng Zhao 0001, Chengfei Liu |
ICDE | 2 |
| 2024 | Reliability-Driven Local Community Search in Dynamic NetworksabstractCommunity search over large dynamic graph has become an important research problem in modern complex networks, such as the online social network, collaboration network and biological networks. Network data in the time-varied environment has motivated several recent studies to identify the evolution of the communities. However, these studies mostly match communities of different snapshot or utilize the aggregation of the disjoint structural information and ignores the cohesion continuity. To fill this research gap, in this work, we propose a novel$(\theta ,k)$-core reliable community (CRC) and define the reliable community search problem which jointly considers member engagement, connection strength and cohesion continuity of the community in the dynamic network. We propose an online search algorithm based on eligible edge filtering and we further construct the Weighted Core Forest-Index (WCF-index) and develop efficient index-based querying algorithm with strong pruning properties. We also propose top-$l$reliable community search problem that couples query based distance to reduce the free rider effect in local community search and support flexible multiple query vertices. Extensive experiments are conducted to show the efficiency and effectiveness of the proposed algorithms. Yifu Tang, Jianxin Li 0001, Nur Al Hasan Haldar, Ziyu Guan, Jiajie Xu 0001, Chengfei Liu |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2024 | Evidence Reasoning and Curriculum Learning for Document-Level Relation ExtractionabstractDocument-level Relation Extraction (RE) is a promising task aiming at identifying relations of multiple entity pairs in a document. Compared with the sentence-level counterpart, it has raised two significant challenges: a) In most cases, a relational fact can be adequately expressed via a small subset of sentences from the document, namely evidence. But the traditional method cannot model such strong semantic correlations between evidence sentences that collaborate to describe a specific relation; b) The data of this task is extremely long-tail in terms of too many NA instances and imbalanced relational types. Such data can mislead the tail prediction bias to the head categories in the RE model. In this paper, we present a novelEvidence reasoning andCurriculum learning method forDocRE(DRE-EC) to address these challenges. Particularly, we first formulate evidence extraction as a sequential decision problem through a crafted reinforcement learning mechanism with an efficient path searching strategy to reduce the action space. Providing the evidence for each entity pair as a customized-filtered document in advance helps infer the relations better. To address the long-tail issue, we further develop a hybrid curriculum learning method at the NA-level (NC) and relation-level (RC) with our customized difficulty measure score. In NC, the NA samples are scheduled in an easy-to-hard scheme and gradually added, resulting in the data distribution from ideal and balanced to real and unbalanced. In RC, the scheme is switched into hard-to-easy to enhance the hard and tail samples. In addition, we propose a new Equalization adaptive Focal Loss(EFLoss) that can adjust to the changing data distribution and focus more on the tail categories. We conduct various experiments on two document-level RE benchmarks and achieve a remarkable improvement over previous competitive baselines. Furthermore, we provide detailed analyses of the advantages and effectiveness of our method. Tianyu Xu 0004, Jianfeng Qu, Wen Hua, Zhixu Li, Jiajie Xu 0001, An Liu 0002, Lei Zhao 0001, Xiaofang Zhou 0001 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2024 | CCML: Curriculum and Contrastive Learning Enhanced Meta-Learner for Personalized Spatial Trajectory PredictionabstractSpatial trajectory prediction is a fundamental problem for diverse location-based applications. However, existing methods fall short in learning and generalization, and cannot sufficiently capture users’ spatiotemporal preferences, especially for cold-start users. Moreover, these methods do not explicitly consider the diversity of moving patterns among users and trajectories, i.e., the learning difficulty of different user and trajectory samples, thus hindering the improvement of prediction accuracy. To solve these problems, we propose a novel Curriculum and Contrastive Learning Enhanced Meta-Learner (CCML) that transfers knowledge from users with rich data to cold-start users. Specifically, a Contrastive-based Trajectory Predictor (CTP) is designed as the base model, which utilizes contrastive learning technique on both user-level and trajectory-level, aiming to facilitate a more profound understanding and differentiation of the varied travel behaviors and preferences exhibited by individuals. Meanwhile, CCML also incorporates the curriculum learning and the hard sample mining strategies. It simultaneously considers the learning difficulty of both user and trajectory samples, and presents the learning tasks by an easy-to-hard curriculum. By learning more challenging combinations of user and trajectory samples in each meta-learning iteration, the meta-learner can converge to a better status. Extensive experiments on two real-world datasets demonstrate the superiority of our models. Jing Zhao 0040, Jiajie Xu 0001, Yuan Xu 0008, Junhua Fang, Pingfu Chao, Xiaofang Zhou 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2023 | LHMM: A Learning Enhanced HMM Model for Cellular Trajectory Map MatchingabstractMap matching is a problem to align recorded location data to a digital map. It has been well studied to map GPS data collected from vehicles to paths in a road network. The problem of Cellular Trajectory Map-Matching (CTMM) is a new problem that deals with trajectories of cellular-based positioning data. It has a wide range of applications, for example, for telecommunication companies to understand and predict traffic information based on telecom tokens obtained from vehicles. CTMM is a significantly more challenging task that faces much lower data precision and higher positioning errors. While Hidden Markov Model (HMM) based methods can achieve satisfactory results for GPS-based map matching, we show that they cannot be directly applied to the CTMM problem. In this paper, we aim at reducing the impact of positioning errors by incorporating knowledge obtained by neural networks into learned probabilities. A multi-relational graph learning method is developed to generate meaningful embedding, with multi-relational useful information fully preserved in a shared space. An attentive neural network is then designed as the learner for observation probability, incorporating the knowledge of the dynamic correlation between roads and cell towers under varying trajectory contexts. A transition probability learner is used to capture implicit deep features for enhanced transition probability modeling. Finally, the learned observation and transition probabilities are seamlessly integrated into HMM to guide more accurate path-finding. Extensive experiments on two large-scale cellular datasets reveal that our approach achieves high accuracy and robustness on CTMM. Jiajie Xu 0001, Junhua Fang, Pingfu Chao, An Liu 0002, Xiaofang Zhou 0001 |
ICDE | 2 |
| 2023 | Garden: a real-time processing framework for continuous top-k trajectory similarity search
Pingfu Chao, Junhua Fang, Wei Chen 0070, Jiajie Xu 0001, Lei Zhao 0001 |
Knowl. Inf. Syst. | 5 |
| 2023 | Densest Multipartite Subgraph Search in Heterogeneous Information NetworksabstractCohesive multipartite subgraphs (CMS) in heterogeneous information networks (HINs) uncover closely connected vertex groups of multiple types, enhancing real applications like community search and anomaly detection. However, existing works for HINs pay less attention to searching CMS. In this paper, we leverage well-established concepts of meta-path and densest subgraph to propose a novel CMS model called the densest P -partite subgraph. Given a multipartite subgraph of an HIN induced by i =| P | types of vertices defined in a query meta-path P (i.e., a P -partite subgraph), we devise a novel density function which is the number of the instances of P over the geometric mean of the sizes of i different types of vertex sets in the subgraph. A P -partite subgraph with the highest density serves as the optimum result. To find the densest P -partite subgraph in an HIN with n vertices, we first design an exact algorithm with a runtime cost equivalent to solving Θ(|M|) instances of the min-cut problem where |M|= O (( n/i ) i ). Then, we attempt a more efficient approximation algorithm that achieves a ratio of 1/ i but still incurs the cost of solving Θ(|M|) instances of our proposed peeling problem. Both approaches struggle with scalability due to Θ(|M|). To overcome this bottleneck, we improve the exact algorithm with novel pruning rules that non-trivially reduce the number of min-cut problem instances to solve to O (|M|). Empirically, 70-90% instances are pruned, making the improved exact algorithm significantly faster than the approximation algorithm. Extensive experiments on real datasets demonstrate the effectiveness of the proposed model and the efficiency of our algorithms. Lu Chen 0008, Chengfei Liu, Rui Zhou 0001, Kewen Liao, Jiajie Xu 0001, Jianxin Li 0001 |
Proc. VLDB Endow. | 5 |
| 2023 | Feature-Level Deeper Self-Attention Network With Contrastive Learning for Sequential RecommendationabstractSequential recommendation, which aims to recommend next item that the user will likely interact in a near future, has become essential in various Internet applications. Existing methods usually consider the transition patterns between items, but ignore the transition patterns between features of items. We argue that only the item-level sequences cannot reveal the full sequential patterns, while explicit and implicit feature-level sequences can help extract the full sequential patterns. Meanwhile, the item-level sequential recommendation also suffers from limited supervised signal issues. In this article, we propose a novel model Feature-level Deeper Self-Attention Network with Contrastive Learning (FDSA-CL) for sequential recommendation. Specifically, FDSA-CL first integrates various heterogeneous features of items into feature-level sequences with different weights through a vanilla attention mechanism. After that, FDSA-CL applies separated self-attention blocks on item-level sequences and feature-level sequences, respectively, to model item transition patterns and feature transition patterns. Moreover, we propose contrastive learning and item feature recommendation tasks to capture the embedding commonality and further utilize the beneficial interaction among the two levels, so as to alleviate the sparsity of the supervised signal and extract the most critical information. Finally, we jointly optimize the above tasks. We evaluate the proposed model using two real-world datasets and experimental results show that our model significantly outperforms the state-of-the-art approaches. Yongjing Hao, Pengpeng Zhao 0001, Yanchi Liu, Victor S. Sheng, Jiajie Xu 0001, Guanfeng Liu 0001, Xiaofang Zhou 0001 |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2023 | Empowering A* Algorithm With Neuralized Variational Heuristics for Fastest Route RecommendationabstractFastest route recommendation (FRR) is crucial for intelligent transportation systems. The existing methods treat it as a pathfinding problem on dynamic graphs, and extend A* algorithm with neuralized travel time estimators as cost functions. However, they fail to provide effective heuristic cost due to the neglect of its admissibility and the utilization of noise path information, resulting in sub-optimal results and inefficiency. Besides, path sequentiality is also ignored, affecting algorithm accuracy as well. In this paper, we propose a variational inference based fastest route recommendation method, which follows the framework of A* algorithm and provides effective costs for routing. Specifically, we first adopt a sequential estimator to accurately estimate the travel time of a specific path. More importantly, we design a variational inference based estimator, which models the distribution of travel time between two nodes and provides an effective heuristic cost with high probability of being admissible. We further take advantage of adversarial learning to enrich the fastest path information. To the best of our knowledge, we are the first to use variational estimator to consider the admissibility of heuristics in FRR. Extensive experiments are conducted on two real-world datasets. The results verify the performance advantage of our proposed method. Minrui Xu, Jiajie Xu 0001, Rui Zhou 0001, Jianxin Li 0001, Kai Zheng 0001, Pengpeng Zhao 0001, Chengfei Liu |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2022 | Aries: Accurate Metric-based Representation Learning for Fast Top-k Trajectory Similarity QueryabstractWith the prevalence of location-based services (LBS), trajectories are being generated rapidly. As is widely used in LBS, top-k trajectory similarity query serves as a key operation, deeply empowering applications such as travel route recommendation and carpooling. Given the rise of deep learning, trajectory representation has been well-proven to speed up this operator. However, existing representation-based computing modes remain two major problems understudied: the low quality of trajectory representation and insufficient support for various trajectory similarity metrics, which make them difficult to apply in practice. Therefore, we propose an Accurate metric-based representation learning approach for fast top-k trajectory similarity query, named Aries. Specifically, Aries has two sophisticated modules: (1) An novel trajectory embedding strategy enhanced by the bidirectional LSTM encoder and spatial attention mechanism, which can extract more precise and comprehensive knowledge. (2) A deep metric learning network aggregating multiple measures for better top-k query. Extensive experiments conducted on real trajectory dataset show that Aries achieves both impressive accuracy and lower training time compared with state-of-the-art solutions. In particular, it achieves 5x-10x speedup and 10%-20% accuracy improvement over Euclidean, Hausdorff, DTW, and EDR measures. Besides, our method can maintain stable performance when handling various scenarios, without repeated training in order to adapt to diverse similarity metrics. Chunhui Feng, Junhua Fang, Jiajie Xu 0001, Pengpeng Zhao 0001, Lei Zhao 0001 |
CIKM | 4 |
| 2022 | Evidence-aware Document-level Relation ExtractionabstractDocument-level Relation Extraction (RE) is a promising task aiming at identifying relations of multiple entity pairs in a document. However, in most cases, a relational fact can be expressed enough via a small subset of sentences from the document, namely evidence sentence. Moreover, there often exist strong semantic correlations between evidence sentences that collaborate together to describe a specific relation. To address these challenges, we propose a novel evidence-aware model for document-level RE. Particularly, we formulate evidence sentence selection as a sequential decision problem through a crafted reinforcement learning mechanism. Considering the explosive search space of our agent, an efficient path searching strategy is executed on the converted document graph to heuristically obtain hopeful sentences and feed them to reinforcement learning. Finally, each entity pair owns a customized-filtered document for further inferring the relation between them. We conduct various experiments on two document-level RE benchmarks and achieve a remarkable improvement over previous competitive baselines, verifying the effectiveness of our method. Tianyu Xu 0004, Wen Hua, Jianfeng Qu, Zhixu Li, Jiajie Xu 0001, An Liu 0002, Lei Zhao 0001 |
CIKM | 5 |
| 2022 | Transportation-Mode Aware Travel Time Estimation via Meta-learning
Jiajie Xu 0001, Rui Zhou 0001, Chengfei Liu |
DASFAA (2) | 2 |
| 2022 | When Multitask Learning Make a Difference: Spatio-Temporal Joint Prediction for Cellular Trajectories
Yuan Xu 0008, Jiajie Xu 0001, Junhua Fang, An Liu 0002, Lei Zhao 0001 |
DASFAA (1) | 2 |
| 2022 | MetaPTP: An Adaptive Meta-optimized Model for Personalized Spatial Trajectory PredictionabstractTrajectory prediction is a fundamental problem for a wide spectrum of location-based applications. Existing methods can achieve inspiring results in predicting personal frequent routes conditioned on massive historical data. However, trajectory estimation may involve cold-start routes or users due to the data sparsity problem, which severely limits the performance of spatial trajectory prediction. Although meta-learning models can alleviate the cold-start problem, they simply utilize the same initialization for all tasks and thus cannot fit each user well due to users' varying travel preferences. To this end, we propose an adaptive meta-optimized model called MetaPTP for personalized spatial trajectory prediction. Specifically, it adopts a soft-clustering based method to guide the network initialization in a finer granularity, so that shared knowledge can be better transferred across users with similar travel preferences. Besides, towards model fine-tuning, an effective trajectory sampling method is introduced to generate meaningful support set, which simultaneously considers user preference and spatial trace similarities to provide task-related information for model adaptation. In addition, we design a weight generator to adaptively assign reasonable weights to trajectories in support set to avoid sub-optimal results which will occur when fine-tuning the initial network with the same weight for trajectories with different user preferences and spatial distributions. Finally, extensive experiments on two real-world datasets demonstrate the superiority of our model. Yuan Xu 0008, Jiajie Xu 0001, Jing Zhao 0040, Kai Zheng 0001, An Liu 0002, Lei Zhao 0001, Xiaofang Zhou 0001 |
KDD | 2 |
| 2022 | Conats: A Novel Framework for Cross-Modal Map Extraction
Junhua Fang, Pingfu Chao, Jianfeng Qu, Pengpeng Zhao 0001, Jiajie Xu 0001 |
WISE | 6 |
| 2022 | Bitcoin Transaction Confirmation Time Prediction: A Classification View
Limeng Zhang, Rui Zhou 0001, Qing Liu 0001, Jiajie Xu 0001, Chengfei Liu |
WISE | 4 |
| 2022 | When Research Topic Trend Prediction Meets Fact-Based AnnotationsabstractAbstract The unprecedented growth of publications in many research domains brings the great convenience for tracing and analyzing the evolution and development of research topics. Despite the significant contributions made by existing studies, they usually extract topics from the titles of papers, instead of obtaining topics from the authoritative sessions provided by venues (e.g., AAAI, NeurIPS, and SIGMOD). To make up for the shortcoming of existing work, we develop a novel framework namely RTTP(Research Topic Trend Prediction). Specifically, the framework contains the following two components: (1) a topic alignment strategy called TAS is designed to obtain the detailed contents of research topics in each year, (2) an enhanced prediction network called EPN is designed to capture the research trend of known years for prediction. In addition, we construct two real-world datasets of specific research domains in computer science, i.e., database and data mining, computer architecture and parallel programming. The experimental results demonstrate that the problem is well solved and our solution outperforms the state-of-the-art methods. Jiajie Xu 0001, Wei Chen 0070, Lei Zhao 0001 |
Data Sci. Eng. | 2 |
| 2022 | MTLM: a multi-task learning model for travel time estimation
Saijun Xu, Ruoqian Zhang, Wanjun Cheng, Jiajie Xu 0001 |
GeoInformatica | 4 |
| 2022 | Efficient Maximal Biclique Enumeration for Large Sparse Bipartite GraphsabstractMaximal bicliques are effective to reveal meaningful information hidden in bipartite graphs. Maximal biclique enumeration (MBE) is challenging since the number of the maximal bicliques grows exponentially w.r.t. the number of vertices in a bipartite graph in the worst case. However, a large bipartite graph is usually very sparse, which is against the worst case and may lead to fast MBE algorithms. The uncharted opportunity is taking advantage of the sparsity to substantially improve the MBE efficiency for large sparse bipartite graphs. We observe that for a large sparse bipartite graph, a vertex u may converge to a few vertices in the same vertex set as u via its neighbours, which reveals that the enumeration scope for a vertex could be very small. Based on this observation, we propose novel concepts: unilateral coreness for individual vertices, unilateral order for each vertex set and unilateral convergence (ζ) for a large sparse bipartite graph, ζ could be a few thousand for a large sparse bipartite graph with hundreds of million edges. Using the unilateral order, every vertex with τ unilateral coreness only needs to check at most 2 τ combinations so that all maximal bicliques can be enumerated and τ is bounded by ζ, which leads to a novel MBE algorithm running in O * (2 ζ ). We then propose a batch-pivots technique to eliminate all enumerations resulting in non-maximal bicliques, which guarantees that every maximal biclique is reported in O (ζ e )-delay, where e is the number of edges. We devise novel data structures that allow storing subgraphs at omissible space for further speeding up MBE. Extensive experiments are conducted on synthetic and real large datasets to justify that our proposed algorithm is faster and more scalable than the existing algorithms. Lu Chen 0008, Chengfei Liu, Rui Zhou 0001, Jiajie Xu 0001, Jianxin Li 0001 |
Proc. VLDB Endow. | 4 |
| 2022 | Reliable Community Search in Dynamic NetworksabstractSearching for local communities is an important research problem that supports advanced data analysis in various complex networks, such as social networks, collaboration networks, cellular networks, etc. The evolution of such networks over time has motivated several recent studies to identify local communities in dynamic networks. However, these studies only utilize the aggregation of disjoint structural information to measure the quality and ignore the reliability of the communities in a continuous time interval. To fill this research gap, we propose a novel (θ, k )- core reliable community (CRC) model in the weighted dynamic networks, and define the problem of most reliable community search that couples the desirable properties of connection strength, cohesive structure continuity, and the maximal member engagement. To solve this problem, we first develop a novel edge filtering based online CRC search algorithm that can effectively filter out the trivial edge information from the networks while searching for a reliable community. Further, we propose an index structure, Weighted Core Forest-Index (WCF-index), and devise an index-based dynamic programming CRC search algorithm, that can prune a large number of insignificant intermediate results and support efficient query processing. Finally, we conduct extensive experiments systematically to demonstrate the efficiency and effectiveness of our proposed algorithms on eight real datasets under various experimental settings. Yifu Tang, Jianxin Li 0001, Nur Al Hasan Haldar, Ziyu Guan, Jiajie Xu 0001, Chengfei Liu |
Proc. VLDB Endow. | 5 |
| 2022 | A Survey and Quantitative Study on Map Inference Algorithms From GPS TrajectoriesabstractMap inference algorithm aims to construct a digital map from other data sources automatically. Due to the labour intensity of traditional map creation and the frequent road change nowadays, map inference is deemed to be a promising solution to automatic map construction and update. However, existing map inference from GPS trajectories suffers from low GPS data quality, which makes the quality of the constructed map unsatisfactory. In this paper, we study the existing map inference algorithms using GPS trajectories. Different from previous surveys, we (1) include the most recent solutions and propose a new categorisation of method; (2) study how different types of GPS errors affect the quality of inference results; (3) evaluate the existing map inference quality measures regarding their ability to identify map quality issues. To achieve these goals, we conduct a comprehensive experimental study on several representative algorithms using both real-world datasets and synthetic datasets, which are generated from our proposed synthetic trajectory generator and artificial map generator. Overall, our study provides insightful observations regarding (1) which inference method performs better in each working scenario, (2) the general data quality requirements for map inference, (3) the direction of future works for quantitative map quality measures. Pingfu Chao, Wen Hua, Rui Mao 0001, Jiajie Xu 0001, Xiaofang Zhou 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2022 | Trajectory-Based Spatiotemporal Entity LinkingabstractTrajectory-based spatiotemporal entity linking is to match the same moving object in different datasets based on their movement traces. It is a fundamental step to support spatiotemporal data integration and analysis. In this paper, we study the problem of spatiotemporal entity linking using effective and concise signatures extracted from their trajectories. This linking problem is formalized as a$k$-nearest neighbor ($k$-NN) query on the signatures. Four representation strategies (sequential, temporal, spatial, and spatiotemporal) and two quantitative criteria (commonality and unicity) are investigated for signature construction. A simple yet effective dimension reduction strategy is developed together with a novel indexing structure called the WR-tree to speed up the search. A number of optimization methods are proposed to improve the accuracy and robustness of the linking. Our extensive experiments on real-world datasets verify the superiority of our approach over the state-of-the-art solutions in terms of both accuracy and efficiency. Fengmei Jin, Wen Hua, Thomas Zhou, Jiajie Xu 0001, Matteo Francia, Maria E. Orlowska, Xiaofang Zhou 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2022 | Index-Based Solutions for Efficient Density Peak ClusteringabstractDensity Peak Clustering (DPC), a popular density-based clustering approach, has received considerable attention from the research community primarily due to its simplicity and fewer-parameter requirement. However, the resultant clusters obtained using DPC are influenced by the sensitive parameter$d_c$, which depends on data distribution and requirements of different users. Besides, the original DPC algorithm requires visiting a large number of objects, making it slow. To this end, this paper investigates index-based solutions for DPC. Specifically, we propose two list-based index methods viz. (i) a simple List Index, and (ii) an advanced Cumulative Histogram Index. Efficient query algorithms are proposed for these indices which significantly avoids irrelevant comparisons at the cost of space. For memory-constrained systems, we further introduce an approximate solution to the above indices which allows substantial reduction in the space cost, provided that slight inaccuracies are admissible. Furthermore, owing to considerably lower memory requirements of existing tree-based index structures, we also present effective pruning techniques and efficient query algorithms to support DPC using the popular Quadtree Index and R-tree Index. Finally, we practically evaluate all the above indices and present the findings and results, obtained from a set of extensive experiments on six synthetic and real datasets. The experimental insights obtained can help to guide in selecting a befitting index. Zafaryab Rasool, Rui Zhou 0001, Lu Chen 0008, Chengfei Liu, Jiajie Xu 0001 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2022 | Where to Go Next: A Spatio-Temporal Gated Network for Next POI RecommendationabstractNext Point-of-Interest (POI) recommendation which is of great value to both users and POI holders is a challenging task since complex sequential patterns and rich contexts are contained in extremely sparse user check-in data. Recently proposed embedding techniques have shown promising results in alleviating the data sparsity issue by modeling context information, and Recurrent Neural Network (RNN) has been proved effective in the sequential prediction. However, existing next POI recommendation approaches train the embedding and network model separately, which cannot fully leverage rich contexts. In this paper, we propose a novel unified neural network framework, named NeuNext, which leverages POI context prediction to assist next POI recommendation by joint learning. Specifically, the Spatio-Temporal Gated Network (STGN) is proposed to model personalized sequential patterns for users’ long and short term preferences in the next POI recommendation. In the POI context prediction, rich contexts on POI sides are used to construct graph, and enforce the smoothness among neighboring POIs. Finally, we jointly train the POI context prediction and the next POI recommendation to fully leverage labeled and unlabeled data. Extensive experiments on real-world datasets show that our method outperforms other approaches for next POI recommendation in terms of Accuracy and MAP. Pengpeng Zhao 0001, Anjing Luo, Yanchi Liu, Jiajie Xu 0001, Zhixu Li, Fuzhen Zhuang, Victor S. Sheng, Xiaofang Zhou 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2021 | CBML: A Cluster-based Meta-learning Model for Session-based RecommendationabstractSession-based recommendation is to predict an anonymous user's next action based on the user's historical actions in the current session. However, the cold-start problem of limited number of actions at the beginning of an anonymous session makes it difficult to model the user's behavior, i.e., hard to capture the user's various and dynamic preferences within the session. This severely affects the accuracy of session-based recommendation. Although some existing meta-learning based approaches have alleviated the cold-start problem by borrowing preferences from other users, they are still weak in modeling the behavior of the current user. To tackle the challenge, we propose a novel cluster-based meta-learning model for session-based recommendation. Specially, we adopt a soft-clustering method and design a parameter gate to better transfer shared knowledge across similar sessions and preserve the characteristics of the session itself. Besides, we apply two self-attention blocks to capture the transition patterns of sessions in both item and feature aspects. Finally, comprehensive experiments are conducted on two real-world datasets and demonstrate the superior performance of CBML over existing approaches. Jiayu Song, Jiajie Xu 0001, Rui Zhou 0001, Lu Chen 0008, Jianxin Li 0001, Chengfei Liu |
CIKM | 2 |
| 2021 | Meta-Learning Based Hyper-Relation Feature Modeling for Out-of-Knowledge-Base EmbeddingabstractKnowledge graph (KG) embedding aims to encode both entities and relations into a continuous vector space. Most existing methods require that all entities should be observed during training while ignoring the evolving nature of KG. Major recent efforts on this issue embed new entities by aggregating neighborhood information from existing entities and relations with Graph Neural Network (GNN). However, these methods rely on the neighbors seen during training and suffer from the embedding of new entities with insufficient triplets or triplets with the unseen-to-unseen form. To relieve this problem, we propose a two-stage learning model referred as Hyper-Relation Feature Learning Network (HRFN) for effective out-of-knowledge-base embedding. For the first stage, HRFN learns pre-representations for emerging entities using hyper-relation features meta-learned from the training set. A novel feature aggregating network that involves an entity-centered Graph Convolutional Network (GCN) and a relation-centered GCN is proposed to aggregate information from both new entities themselves and their neighbors. For stage two, a transductive learning network is employed to learn finer-grained embeddings based on above-mentioned pre-representations of new entities. Experimental results on the link prediction task demonstrate the superiority of our model. Further analysis is also done to validate the effectiveness and efficiency of pre-representing emerging entities with the hyper-relation feature. Weiqing Wang 0001, Wei Chen 0070, Jiajie Xu 0001, An Liu 0002, Lei Zhao 0001 |
CIKM | 4 |
| 2021 | SSRGAN: A Generative Adversarial Network for Streaming Sequential Recommendation
Yao Lv, Jiajie Xu 0001, Rui Zhou 0001, Junhua Fang, Chengfei Liu |
DASFAA (3) | 2 |
| 2021 | Index-based Solutions for Efficient Density Peak Clustering (Extended Abstract)abstractClusters reflect a potential relationship among different entities of data. This data can be sourced from a wide range of domains like market research, spatial data analysis, etc. Many clustering algorithms have been developed in the last few decades in response to the proliferating demands across industries and organizations, which help them make operational and strategic decisions. Among them, density-based clustering algorithms are popular, which find subsets of objects in "dense regions" separated by not-so-dense regions, where each subset represents a cluster. In this paper, our focal point will be Density Peak Clustering (DPC) [1] , a popular approach towards obtaining density-based clusters. Zafaryab Rasool, Rui Zhou 0001, Lu Chen 0008, Chengfei Liu, Jiajie Xu 0001 |
ICDE | 5 |
| 2021 | Efficient Exact Algorithms for Maximum Balanced Biclique Search in Bipartite GraphsabstractGiven a bipartite graph, the maximum balanced biclique (MBB) problem, discovering a mutually connected while disjoint sets of equal size with the maximum cardinality, plays a significant role for mining the bipartite graph and has numerous applications. Despite the NP-hardness of the MBB problem, in this paper, we show that an exact MBB can be discovered extremely fast in bipartite graphs for real applications. We propose two exact algorithms dedicated for small dense and large sparse bipartite graphs respectively. For dense bipartite graphs, an O*(1.3803n) algorithm is proposed. This algorithm in fact can find an MBB very fast for small dense bipartite graphs that are common for applications such as VLSI design. This is because, using our proposed novel techniques, the search can fast converge to sufficiently dense bipartite graphs which we prove to be polynomial-time solvable. For large sparse bipartite graphs typical for applications such as biological data analysis, an O*(1.3803 δ) algorithm is proposed, where δ is only a few hundred for large sparse bipartite graphs with millions of vertices. The indispensible optimization that leads to this time complexity is: we transform a large sparse bipartite graph into a limited number of dense subgraphs such that each of the dense subgraphs has up to δ vertices and then apply our proposed algorithm for dense bipartite graphs on each of the subgraphs. To further speed up this algorithm, tighter upper bounds, faster heuristics and more effective reductions are proposed, allowing an MBB to be discovered within a few seconds for bipartite graphs with millions of vertices. Extensive experiments are conducted on synthetic and real large bipartite graphs to demonstrate the efficiency and effectiveness of our proposed algorithms and techniques. Lu Chen 0008, Chengfei Liu, Rui Zhou 0001, Jiajie Xu 0001, Jianxin Li 0001 |
SIGMOD Conference | 4 |
| 2021 | Disatra: A Real-Time Distributed Abstract Trajectory Clustering
Pingfu Chao, Junhua Fang, Wei Chen 0070, Jiajie Xu 0001, Lei Zhao 0001 |
WISE (1) | 5 |
| 2021 | Transaction Confirmation Time Estimation in the Bitcoin Blockchain
Limeng Zhang, Rui Zhou 0001, Qing Liu 0001, Jiajie Xu 0001, Chengfei Liu |
WISE (1) | 4 |
| 2021 | TAML: A Traffic-aware Multi-task Learning Model for Estimating Travel TimeabstractTravel time estimation has been recognized as an important research topic that can find broad applications. Existing approaches aim to explore mobility patterns via trajectory embedding for travel time estimation. Though state-of-the-art methods utilize estimated traffic condition (by explicit features such as average traffic speed) for auxiliary supervision of travel time estimation, they fail to model their mutual influence and result in inaccuracy accordingly. To this end, in this article, we propose an improved traffic-aware model, called TAML, which adopts a multi-task learning network to integrate a travel time estimator and a traffic estimator in a shared space and improves the accuracy of estimation by enhanced representation of traffic condition, such that more meaningful implicit features are fully captured. In TAML, multi-task learning is further applied for travel time estimation in multi-granularities (including road segment, sub-path, and entire path). The multiple loss functions are combined by considering the homoscedastic uncertainty of each task. Extensive experiments on two real trajectory datasets demonstrate the effectiveness of our proposed methods. Jiajie Xu 0001, Saijun Xu, Rui Zhou 0001, Chengfei Liu, An Liu 0002, Lei Zhao 0001 |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2021 | Improving the Quality of Web-Based Data Imputation With Crowd InterventionabstractData incompleteness is a common data quality problem in databases. Recent work proposes to retrieve missing string values from the World Wide Web for higher imputation recall, but on the other hand, takes the risk of introducing web noises into the imputation results. So far there lacks an effective way to control the quality of web-based data imputation, given the complexity of the quality model and lacking of enough ground truth data. In this article, an EM-based quality model is first built for web-based data imputation which investigates three key factors jointly, i.e., precision of web sources, correlation among web sources, and precision and recall of the employed extractors. However, the accuracy of the EM-based quality model could be harmed when the EM (Expectation Maximization) assumption that “the majority agree on the truth” does not hold in some cases. To solve this problem, we introduce crowd intervention to help improve the quality model. While a straightforward but expensive way is to let the crowd to identify all these undesirable cases and provide the right imputation values for these blanks, a most crowd-economic way is to select a small set of blanks for crowd-based imputation, whose results could help to adjust the EM-based quality model towards a better one. To achieve this, an adaptive blank selection strategy is proposed to select a sequence of blanks for crowd-based imputation. Also, we work on finding a proper time to stop further crowd intervention for the balance of crowd efficiency and quality improvement. Our experiments performed on three real world and one simulated data collections prove that the proposed quality model can effectively help improve the quality of the web-based imputation results by more than 15 percent, while our crowd cost saving strategy saves more than 75 percent crowd cost. Binbin Gu, Zhixu Li, An Liu 0002, Jiajie Xu 0001, Lei Zhao 0001, Xiaofang Zhou 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2021 | Predicting Destinations by a Deep Learning based ApproachabstractDestination prediction is known as an important problem for many location based services (LBSs). Existing solutions generally apply probabilistic models or neural network models to predict destinations over a subtrajectory, and adopt the standard attention mechanism to improve the prediction accuracy. However, the standard attention mechanism uses fixed feature representations, and has a limited ability to represent distinct features of locations. Besides, existing methods rarely take the impact of spatial and temporal characteristics of the trajectory into account. Their accuracies in fine-granularity prediction are always not satisfactory due to the data sparsity problem. Thus, in this paper, a carefully designed deep learning model called LATL model is presented. It not only adopts an adaptive attention network to model the distinct features of locations, but also implements time gates and distance gates into the Long Short-Term Memory (LSTM) network to capture the spatial-temporal relation between consecutive locations. Furthermore, to better understand the mobility patterns in different spatial granularities, and explore the fusion of multi-granularity learning capability, a hierarchical model that utilizes tailored combination of different neural networks under multiple spatial granularities is further proposed. Extensive empirical studies verify that the newly proposed models perform effectively and settle the problem nicely. Jiajie Xu 0001, Jing Zhao 0040, Rui Zhou 0001, Chengfei Liu, Pengpeng Zhao 0001, Lei Zhao 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2020 | MemTimes: Temporal Scoping of Facts with Memory Network
Siyuan Cao, Qiang Yang 0015, Zhixu Li, Guanfeng Liu 0001, Detian Zhang, Jiajie Xu 0001 |
DASFAA (3) | 6 |
| 2020 | Modeling Periodic Pattern with Self-Attention Network for Sequential Recommendation
Pengpeng Zhao 0001, Yanchi Liu, Victor S. Sheng, Jiajie Xu 0001, Lei Zhao 0001 |
DASFAA (3) | 5 |
| 2020 | C2TTE: Cross-city Transfer Based Model for Travel Time Estimation
Jiayu Song, Jiajie Xu 0001, Xinghong Ling, Junhua Fang, Rui Zhou 0001, Chengfei Liu |
DASFAA (1) | 2 |
| 2020 | MTGCN: A Multitask Deep Learning Model for Traffic Flow Prediction
Fucheng Wang, Jiajie Xu 0001, Chengfei Liu, Rui Zhou 0001, Pengpeng Zhao 0001 |
DASFAA (1) | 2 |
| 2020 | TADNM: A Transportation-Mode Aware Deep Neural Model for Travel Time Estimation
Saijun Xu, Jiajie Xu 0001, Rui Zhou 0001, Chengfei Liu, Zhixu Li, An Liu 0002 |
DASFAA (1) | 2 |
| 2020 | Finding Effective Geo-social Group for Impromptu Activities with Diverse DemandsabstractGeo-social group search aims to find a group of people proximate to a location while socially related. One of the driven applications for geo-social group search is organizing an impromptu activity. This is because the social cohesiveness of a found geo-social group ensures a good communication atmosphere for the activity and the spatial closeness of the geo-social group reduces the preparation time for the activity. Most existing works treat geo-social group search as a problem that finds a group satisfying a single social constraint while optimizing the spatial proximity. However, since different impromptu activities have diverse demands on attendees, e.g. an activity could require (or prefer) the attendees to have skills (or favorites) related to the activity, the existing works cannot find this kind of geo-social groups effectively. In this paper, we propose a novel geo-social group model, equipped with elegant keyword constraints, to fill this gap. We propose a novel search framework which first significantly narrows down the search space with theoretical guarantees and then efficiently finds the optimum result. To evaluate the effectiveness, we conduct experiments on real datasets, demonstrating the superiority of our proposed model. We conduct extensive experiments on large semi-synthetic datasets for justifying the efficiency of the proposed search algorithms. Lu Chen 0008, Chengfei Liu, Rui Zhou 0001, Jiajie Xu 0001, Jeffrey Xu Yu, Jianxin Li 0001 |
KDD | 4 |
| 2020 | User Profile Linkage Across Multiple Social Platforms
Manman Wang, Wei Chen 0070, Jiajie Xu 0001, Pengpeng Zhao 0001, Lei Zhao 0001 |
WISE (1) | 3 |
| 2020 | Exploiting Aesthetic Preference in Deep Cross Networks for Cross-domain RecommendationabstractVisual aesthetics of products plays an important role in the decision process when purchasing appearance-first products, e.g., clothes. Indeed, user’s aesthetic preference, which serves as a personality trait and a basic requirement, is domain independent and could be used as a bridge between domains for knowledge transfer. However, existing work has rarely considered the aesthetic information in product images for cross-domain recommendation. To this end, in this paper, we propose a new deep Aesthetic Cross-Domain Networks (ACDN), in which parameters characterizing personal aesthetic preferences are shared across networks to transfer knowledge between domains. Specifically, we first leverage an aesthetic network to extract aesthetic features. Then, we integrate these features into a cross-domain network to transfer users’ domain independent aesthetic preferences. Moreover, network cross-connections are introduced to enable dual knowledge transfer across domains. Finally, the experimental results on real-world datasets show that our proposed model ACDN outperforms benchmark methods in terms of recommendation accuracy. Jian Liu 0001, Pengpeng Zhao 0001, Fuzhen Zhuang, Yanchi Liu, Victor S. Sheng, Jiajie Xu 0001, Xiaofang Zhou 0001, Hui Xiong 0001 |
WWW | 6 |
| 2020 | Multi-objective spatial keyword query with semantics: a distance-owner based approach
Jiajie Xu 0001, Lihua Yin |
Distributed Parallel Databases | 1 |
| 2020 | On accurate POI recommendation via transfer learning
Siyi Wei, Xiaojiao Hu, Jiajie Xu 0001 |
Distributed Parallel Databases | 5 |
| 2020 | S2R-tree: a pivot-based indexing structure for semantic-aware spatial keyword search
Jiajie Xu 0001, Rui Zhou 0001, Pengpeng Zhao 0001, Chengfei Liu, Junhua Fang, Lei Zhao 0001 |
GeoInformatica | 2 |
| 2020 | Collective spatial keyword search on activity trajectories
Xiaozhao Song, Jiajie Xu 0001, Rui Zhou 0001, Chengfei Liu, Kai Zheng 0001, Pengpeng Zhao 0001, Nick Falkner |
GeoInformatica | 2 |
| 2020 | Privacy-preserving shared collaborative web services QoS prediction
An Liu 0002, Xindi Shen, Haoran Xie 0001, Zhixu Li, Guanfeng Liu 0001, Jiajie Xu 0001, Lei Zhao 0001, Fu Lee Wang |
J. Intell. Inf. Syst. | 6 |
| 2019 | Multiple Query Point Based Collective Spatial Keyword Querying
Jiajie Xu 0001 |
ADMA | 5 |
| 2019 | Attention and Convolution Enhanced Memory Network for Sequential Recommendation
Jian Liu 0001, Pengpeng Zhao 0001, Yanchi Liu, Jiajie Xu 0001, Junhua Fang, Lei Zhao 0001, Victor S. Sheng, Zhiming Cui 0002 |
DASFAA (2) | 4 |
| 2019 | Adaptive Attention-Aware Gated Recurrent Unit for Sequential Recommendation
Anjing Luo, Pengpeng Zhao 0001, Yanchi Liu, Jiajie Xu 0001, Zhixu Li, Lei Zhao 0001, Victor S. Sheng, Zhiming Cui 0002 |
DASFAA (2) | 4 |
| 2019 | AdaCML: Adaptive Collaborative Metric Learning for Recommendation
Pengpeng Zhao 0001, Yanchi Liu, Jiajie Xu 0001, Junhua Fang, Lei Zhao 0001, Victor S. Sheng, Zhiming Cui 0002 |
DASFAA (2) | 4 |
| 2019 | Moving Object Linking Based on Historical TraceabstractThe prevalent adoption of GPS-enabled devices has witnessed an explosion of various location-based services which produce a huge amount of trajectories monitoring an individual's movement. This triggers an interesting question: is movement history sufficiently representative and distinctive to identify an individual? In this work, we study the problem of moving object linking based on their historical traces. However, it is non-trivial to extract effective patterns from moving history and meanwhile conduct object linking efficiently. To this end, we propose four representation strategies (sequential, temporal, spatial, and spatiotemporal) and two quantitative criteria (commonality and unicity) to construct the personalised signature from the historical trace. Moreover, we formalise the problem of moving object linking as a k-nearest neighbour (k-NN) search on the collection of signatures, and aim to improve efficiency considering the high dimensionality of signatures and the large cardinality of the candidate object set. A simple but effective dimension reduction strategy is introduced in this work, which empirically outperforms existing algorithms including PCA and LSH. We propose a novel indexing structure, Weighted R-tree (WR-tree), and two pruning methods to further speed up k-NN search by combining weight and spatial information contained in the signature. Our extensive experimental results on a real world dataset verify the superiority of our proposals, in terms of both accuracy and efficiency, over state-of-the-art approaches. Fengmei Jin, Wen Hua, Jiajie Xu 0001, Xiaofang Zhou 0001 |
ICDE | 3 |
| 2019 | Multiple Interaction Attention Model for Open-World Knowledge Graph Completion
Chenpeng Fu, Zhixu Li, Qiang Yang 0015, Zhigang Chen 0003, Junhua Fang, Pengpeng Zhao 0001, Jiajie Xu 0001 |
WISE | 7 |
| 2019 | Locking Mechanism for Concurrency Conflicts on Hyperledger Fabric
Wei Chen 0070, Zhixu Li, Jiajie Xu 0001, An Liu 0002, Lei Zhao 0001 |
WISE | 4 |
| 2019 | Recurrent Convolutional Neural Network for Sequential RecommendationabstractThe sequential recommendation, which models sequential behavioral patterns among users for the recommendation, plays a critical role in recommender systems. However, the state-of-the-art Recurrent Neural Networks (RNN) solutions rarely consider the non-linear feature interactions and non-monotone short-term sequential patterns, which are essential for user behavior modeling in sparse sequence data. In this paper, we propose a novel Recurrent Convolutional Neural Network model (RCNN). It not only utilizes the recurrent architecture of RNN to capture complex long-term dependencies, but also leverages the convolutional operation of Convolutional Neural Network (CNN) model to extract short-term sequential patterns among recurrent hidden states. Specifically, we first generate a hidden state at each time step with the recurrent layer. Then the recent hidden states are regarded as an “image”, and RCNN searches non-linear feature interactions and non-monotone local patterns via intra-step horizontal and inter-step vertical convolutional filters, respectively. Moreover, the output of convolutional filters and the hidden state are concatenated and fed into a fully-connected layer to generate the recommendation. Finally, we evaluate the proposed model using four real-world datasets from various application scenarios. The experimental results show that our model RCNN significantly outperforms the state-of-the-art approaches on sequential recommendation. Chengfeng Xu, Pengpeng Zhao 0001, Yanchi Liu, Jiajie Xu 0001, Victor S. Sheng, Zhiming Cui 0002, Xiaofang Zhou 0001, Hui Xiong 0001 |
WWW | 4 |
| 2018 | On Prediction of User Destination by Sub-Trajectory Understanding: A Deep Learning based ApproachabstractDestination prediction is known as an important problem for many location based services (LBSs). Existing solutions generally apply probabilistic models to predict destinations over a sub-trajectory, but their accuracies in fine-granularity prediction are always not satisfactory due to the data sparsity problem. This paper presents a carefully designed deep learning model called TALL model for destination prediction. It not only takes advantage of the bidirectional Long Short-Term Memory (LSTM) network for sequence modeling, but also gives more attention to meaningful locations that have strong correlations w.r.t. destination by adopting attention mechanism. Furthermore, a hierarchical model that explores the fusion of multi-granularity learning capability is further proposed to improve the accuracy of prediction. Extensive experiments on Beijing and Chengdu real datasets finally demonstrate that our proposed models outperform existing methods without considering external features. Jing Zhao 0040, Jiajie Xu 0001, Rui Zhou 0001, Pengpeng Zhao 0001, Chengfei Liu, Feng Zhu 0011 |
CIKM | 2 |
| 2018 | A Privacy-Preserving Framework for Subgraph Pattern Matching in Cloud
Jiuru Gao, Jiajie Xu 0001, Guanfeng Liu 0001, Wei Chen 0070, Hongzhi Yin, Lei Zhao 0001 |
DASFAA (1) | 2 |
| 2018 | Discovering Expert Drivers from TrajectoriesabstractDiscovering expert drivers is highly important for a broad range of location based services, but this issue is largely untouched in previous trajectory mining and search studies. In this paper, we study the problem of trajectory data driven expert driver discovery. It aims to find out top-k expert drivers about a region of interest, based on the understanding of their historical trajectories. To this end, we first investigate the construction of reference system, which collectively describes exemplar routes among important junctions, so that the driving behaviors embedded in each trajectory can be evaluated. To discover expert drivers accurately, a novel tf-idf concept based measure is proposed afterwards, such that the rationality of their trajectories are not only precisely evaluated by the match to reference system, but also properly aggregated for modelling expert drivers. Extensive experimental evaluation using real trajectory datasets demonstrates the effectiveness and efficiency of our proposed solutions. Jiabao Sun, Jiajie Xu 0001, Rui Zhou 0001, Kai Zheng 0001, Chengfei Liu |
ICDE | 2 |
| 2018 | Eliminating Temporal Conflicts in Uncertain Temporal Knowledge Graphs
Lingjiao Lu, Junhua Fang, Pengpeng Zhao 0001, Jiajie Xu 0001, Hongzhi Yin, Lei Zhao 0001 |
WISE (1) | 4 |
| 2017 | Interactive Spatial Keyword Querying with SemanticsabstractConventional spatial keyword queries confront the difficulty of returning desired objects that are synonyms but morphologically different to query keywords. To overcome this flaw, this paper investigates the interactive spatial keyword querying with semantics. It aims to enhance the conventional queries by not only making sense of the query keywords, but also refining the understanding of query semantics through interactions. On top of the probabilistic topic model, a novel interactive strategy is proposed to precisely infer the latent query semantics by learning from user feedbacks. In each interaction, the returned objects are carefully selected to ensure effective inference of user intended query semantics. Query processing is carried out on a small candidate object set at each round of interaction, and the whole querying process terminates when the latent query semantics learned from user feedback becomes explicit enough. The experimental results on real check-in dataset demonstrates that the quality of results has been significantly improved through limited number of interactions. Jiabao Sun, Jiajie Xu 0001, Kai Zheng 0001, Chengfei Liu |
CIKM | 2 |
| 2017 | Multi-objective Spatial Keyword Query with Semantics
Jiajie Xu 0001, Chengfei Liu, Zhixu Li, An Liu 0002, Zhiming Ding |
DASFAA (2) | 2 |
| 2017 | Outlier Trajectory Detection: A Trajectory Analytics Based Approach
Zhongjian Lv, Jiajie Xu 0001, Pengpeng Zhao 0001, Guanfeng Liu 0001, Lei Zhao 0001, Xiaofang Zhou 0001 |
DASFAA (1) | 2 |
| 2017 | Social Personalized Ranking Embedding for Next POI Recommendation
Pengpeng Zhao 0001, Victor S. Sheng, Guanfeng Liu 0001, Jiajie Xu 0001, Jian Wu 0002, Zhiming Cui 0002 |
WISE (1) | 5 |
| 2017 | Semantic-aware Query Processing for Activity TrajectoriesabstractNowadays, users of social networks like tweets and weibo have generated massive geo-tagged records, and these records reveal their activities in the physical world together with spatio-temporal dynamics. Existing trajectory data management studies mainly focus on analyzing the spatio-temporal properties of trajectories, while leaving the understanding of their activities largely untouched. In this paper, we incorporate the semantic analysis of the activity information embedded in trajectories into query modelling and processing, with the aim of providing end users more accurate and meaningful trip recommendations. To this end, we propose a novel trajectory query that not only considers the spatio-temporal closeness but also, more importantly, leverages probabilistic topic modelling to capture the semantic relevance of the activities between data and query. To support efficient query processing, we design a novel hybrid index structure, namely ST-tree, to organize the trajectory points hierarchically, which enables us to prune the search space in spatial and topic dimensions simultaneously. The experimental results on real datasets demonstrate the efficiency and scalability of the proposed index structure and search algorithms. Jiajie Xu 0001, Kai Zheng 0001, Chengfei Liu, Lan Du 0002 |
WSDM | 2 |
| 2016 | A Hybrid Method for POI Recommendation: Combining Check-In Count, Geographical Information and Reviews
Xiefeng Xu, Pengpeng Zhao 0001, Guanfeng Liu 0001, Caidong Gu, Jiajie Xu 0001, Jian Wu 0002, Zhiming Cui 0002 |
APWeb (2) | 5 |
| 2016 | An Efficient Location-Aware Top-k Subscription Matching for Publish/Subscribe with Boolean Expressions
Hanhan Jiang, Pengpeng Zhao 0001, Victor S. Sheng, Jiajie Xu 0001, An Liu 0002, Jian Wu 0002, Zhiming Cui 0002 |
DASFAA (2) | 4 |
| 2016 | On Efficient Spatial Keyword Querying with Semantics
Zhihu Qian, Jiajie Xu 0001, Kai Zheng 0001, Zhixu Li, Haoming Guo |
DASFAA (2) | 2 |
| 2015 | A Secure and Efficient Framework for Privacy Preserving Social Recommendation
Shushu Liu, An Liu 0002, Guanfeng Liu 0001, Zhixu Li, Jiajie Xu 0001, Pengpeng Zhao 0001, Lei Zhao 0001 |
APWeb | 5 |
| 2015 | PPS-POI-Rec: A Privacy Preserving Social Point-of-Interest Recommender System
Xiao Liu 0043, An Liu 0002, Guanfeng Liu 0001, Zhixu Li, Jiajie Xu 0001, Pengpeng Zhao 0001, Lei Zhao 0001 |
APWeb | 5 |
| 2015 | Making Sense of Spatial TrajectoriesabstractSpatial trajectory data is widely available today. Over a sustained period of time, trajectory data has been collected from numerous GPS devices, smartphones, sensors and social media applications. Daily increases of real-time trajectory data have also been phenomenal in recent years. More and more new applications have emerged to derive business values from both trajectory data warehouses and real-time trajectory data. Due to their very large volumes, their nature of streaming, their highly variable levels of data quality, as well as many possible links with other types of data, making sense of spatial trajectory data becomes one of the crucial areas for big data analytics. In this paper we will present a review of the extensive work in spatiotemporal data management and trajectory mining, and discuss new challenges and new opportunities in the context of new applications, focusing on recent advances in trajectory data management and trajectory mining from their foundations to high performance processing with modern computing infrastructure. Xiaofang Zhou 0001, Kai Zheng 0001, Hoyoung Jeung, Jiajie Xu 0001, Shazia Sadiq |
CIKM | 4 |
| 2015 | An Efficient Method to Find the Optimal Social Trust Path in Contextual Social Graphs
Guanfeng Liu 0001, Lei Zhao 0001, Kai Zheng 0001, An Liu 0002, Jiajie Xu 0001, Zhixu Li, Athman Bouguettaya |
DASFAA (2) | 5 |
| 2015 | On Efficient Passenger Assignment for Group Transportation
Jiajie Xu 0001, Guanfeng Liu 0001, Kai Zheng 0001, Chengfei Liu, Haoming Guo, Zhiming Ding |
DASFAA (1) | 1 |
| 2015 | Scalable Top- k Spatial Image Search on Road Networks
Pengpeng Zhao 0001, Xiaopeng Kuang, Victor S. Sheng, Jiajie Xu 0001, Jian Wu 0002, Zhiming Cui 0002 |
DASFAA (2) | 4 |
| 2015 | Efficient Trip Planning for Maximizing User Satisfaction
Jiajie Xu 0001, Chengfei Liu, Pengpeng Zhao 0001, An Liu 0002, Lei Zhao 0001 |
DASFAA (1) | 2 |
| 2015 | Interactive Top-k Spatial Keyword queriesabstractConventional top-k spatial keyword queries require users to explicitly specify their preferences between spatial proximity and keyword relevance. In this work we investigate how to eliminate this requirement by enhancing the conventional queries with interaction, resulting in Interactive Top-k Spatial Keyword (ITkSK) query. Having confirmed the feasibility by theoretical analysis, we propose a three-phase solution focusing on both effectiveness and efficiency. The first phase substantially narrows down the search space for subsequent phases by efficiently retrieving a set of geo-textual k-skyband objects as the initial candidates. In the second phase three practical strategies for selecting a subset of candidates are developed with the aim of maximizing the expected benefit for learning user preferences at each round of interaction. Finally we discuss how to determine the termination condition automatically and estimate the preference based on the user's feedback. Empirical study based on real PoI datasets verifies our theoretical observation that the quality of top-k results in spatial keyword queries can be greatly improved through only a few rounds of interactions. Kai Zheng 0001, Han Su 0001, Bolong Zheng, Shuo Shang, Jiajie Xu 0001, Jiajun Liu 0004, Xiaofang Zhou 0001 |
ICDE | 5 |
| 2015 | Effective Sampling of Points of Interests on Maps Based on Road Networks
Ziting Zhou, Pengpeng Zhao 0001, Victor S. Sheng, Jiajie Xu 0001, Zhixu Li, Jian Wu 0002, Zhiming Cui 0002 |
WAIM | 4 |
| 2015 | Ranked Reverse Boolean Spatial Keyword Nearest Neighbors Search
Hailin Fang, Pengpeng Zhao 0001, Victor S. Sheng, Zhixu Li, Jiajie Xu 0001, Jian Wu 0002, Zhiming Cui 0002 |
WISE (1) | 5 |
| 2015 | HV: A Feature Based Method for Trajectory Dataset Profiling
Jie Zhu 0009, Jiajie Xu 0001, Zhixu Li, Pengpeng Zhao 0001, Lei Zhao 0001 |
WISE (1) | 3 |
| 2015 | Efficient route search on hierarchical dynamic road networks
Jiajie Xu 0001, Yunjun Gao, Chengfei Liu, Lei Zhao 0001, Zhiming Ding |
Distributed Parallel Databases | 1 |
| 2014 | A Social Trust Path Recommendation System in Contextual Online Social Networks
Guohao Sun 0001, Guanfeng Liu 0001, Lei Zhao 0001, Jiajie Xu 0001, An Liu 0002, Xiaofang Zhou 0001 |
APWeb | 4 |
| 2014 | SharkDB: An In-Memory Column-Oriented Trajectory StorageabstractThe last decade has witnessed the prevalence of sensor and GPS technologies that produce a high volume of trajectory data representing the motion history of moving objects. However some characteristics of trajectories such as variable lengths and asynchronous sampling rates make it difficult to fit into traditional database systems that are disk-based and tuple-oriented. Motivated by the success of column store and recent development of in-memory databases, we try to explore the potential opportunities of boosting the performance of trajectory data processing by designing a novel trajectory storage within main memory. In contrast to most existing trajectory indexing methods that keep consecutive samples of the same trajectory in the same disk page, we partition the database into frames in which the positions of all moving objects at the same time instant are stored together and aligned in main memory. We found this column-wise storage to be surprisingly well suited for in-memory computing since most frames can be stored in highly compressed form, which is pivotal for increasing the memory throughput and reducing CPU-cache miss. The independence between frames also makes them natural working units when parallelizing data processing on a multi-core environment. Lastly we run a variety of common trajectory queries on both real and synthetic datasets in order to demonstrate advantages and study the limitations of our proposed storage. Haozhou Wang, Kai Zheng 0001, Jiajie Xu 0001, Bolong Zheng, Xiaofang Zhou 0001, Shazia Sadiq |
CIKM | 3 |
| 2014 | Ranking Based Activity Trajectory Search
Wei Chen 0070, Lei Zhao 0001, Jiajie Xu 0001, Kai Zheng 0001, Xiaofang Zhou 0001 |
WISE (1) | 3 |
| 2013 | On Efficient Map-Matching According to Intersections You Pass By
Chengfei Liu, Kuien Liu, Jiajie Xu 0001, Fengcheng He, Zhiming Ding |
DEXA (2) | 4 |
| 2013 | Exploiting Structural Similarity for Automatic Information Extraction from Lists
Dat T. Huynh, Jiajie Xu 0001, Shazia Sadiq, Xiaofang Zhou 0001 |
WISE (2) | 2 |
| 2012 | Traffic Aware Route Planning in Dynamic Road Networks
Jiajie Xu 0001, Limin Guo 0002, Zhiming Ding, Xiling Sun, Chengfei Liu |
DASFAA (1) | 1 |
| 2012 | Effective map-matching on the most simplified road networkabstractThe effectiveness of map-matching algorithms highly depends on the accuracy and correctness of underlying road networks. In practice, the storage capacity of certain hardware, e.g. mobile devices and embedded systems, is sometimes insufficient to maintain a large digital map for map-matching. Unfortunately, most existing map-matching approaches consider little about this problem. They only apply to environments with information-rich maps, but turn out to be unacceptable for map-matching on simplified road networks. In this paper, we propose a novel map-matching algorithm called Passby to work on most simplified road networks. The storage size of a digital map in disk or memory can be greatly reduced after the simplification. Even under the most simplified situation, i.e., each road segment only consists of a couple of intersection points and omits any other information of it, the experimental results on real dataset show that our Passby algorithm significantly maintains high matching accuracy. Benefiting from the small size of map, simple index structure and heuristic foresight strategy, Passby improves matching accuracy as well as efficiency. Kuien Liu, Fengcheng He, Jiajie Xu 0001, Zhiming Ding |
SIGSPATIAL/GIS | 4 |
| 2010 | An Artifact-Centric Approach to Generating Web-Based Business Process Driven User Interfaces
Sira Yongchareon, Chengfei Liu, Xiaohui Zhao 0001, Jiajie Xu 0001 |
WISE | 4 |
| 2006 | Calculation of Target Locations for Web Resources
Saeid Asadi, Jiajie Xu 0001, Joachim Diederich, Xiaofang Zhou 0001 |
WISE | 2 |