Xujun Zhao

dblp:97/9831 · DBLP profile ↗
← Back
28ranked-venue papers
4as first author
22since 2021 · last 2026
0000-0003-4246-4383ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 3 first-author · 15 since 2021Databases, data management, data science and information retrieval · 6 · 6 since 2021Systems, architecture and hardware · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Multi-source Anomaly Detection Using Feature Selection and Relevant Subspace
Xujun Zhao, Jifu Zhang, Jianghui Cai, Haifeng Yang 0001
ICIC2
2026 A Novel Framework for temporal knowledge graph reasoning with complex causal relations
Jianghui Cai, Cuicui Xu, Haifeng Yang 0001, Xin Chen 0070, Aiyu Zheng, Yaling Xun, Xujun Zhao
Expert Syst. Appl.8
2026 Low-reliability characteristics augment and recognition based on kinship features space
Yanting He, Haifeng Yang 0001, Jianghui Cai, Chenhui Shi 0002, Meihong Su, Xujun Zhao, Yaling Xun
Expert Syst. Appl.8
2026 Content suppression mechanisms-based recommendation systems
Haifeng Yang 0001, Jianghui Cai, Jie Wang 0046, Yaling Xun, Xujun Zhao
Expert Syst. Appl.8
2026 Multi-view clustering based on heterogeneous representation learning and tensor weighted low-rank constraints
Haifeng Yang 0001, Chenhui Shi 0002, Yongjie Xin, Jianghui Cai, Lichan Zhou, Meihong Su, Yanting He, Xujun Zhao, Yaling Xun
Neurocomputing8
2026 Multi-view clustering based on the association of graph structure and feature distribution
Chenhui Shi 0002, Yongjie Xin, Haifeng Yang 0001, Jianghui Cai, Jie Wang 0046, Lichan Zhou, Yanting He, Fuxing Cui, Xujun Zhao, Yaling Xun
Inf. Process. Manag.9
2026 A confidence-aware active learning framework for cross-modal inconsistency in clustering
Chenhui Shi 0002, Haifeng Yang 0001, Jianghui Cai, Yanting He, Meihong Su, Xujun Zhao, Yaling Xun
Knowl. Based Syst.6
2026 Dual-channel hard negative sample generation for graph contrastive learning
Jianghui Cai, Haifeng Yang 0001, Jie Wang 0046, Guojiao An, Yaling Xun, Xujun Zhao
Neural Networks8
2025 Interpretable deep classification of time series based on class discriminative prototype learning
abstract
Prototypes help to explain the predictions of deep classification models for time series. However, most models learn prototypes by randomly initializing an uncertain number of low-discriminative prototypes, which may lead to unstable models and unreliable results. To address these issues, we propose a new class D iscriminative P rototype L earning Net work (DPL-Net), which learns an appropriate number of class-discriminative prototypes, thus improving classification performance. Specifically, the proposed P rototype I nitialization M echanism (PIM) introduces a new proximity metric based on the silhouette coefficient and statistical metrics. It facilitates the automatic determination of the class-discriminative prototypes for each class. Then, the encoder layer encodes the prototypes derived from PIM and the input series using one-dimensional convolutional neural networks (1D-CNN). Finally, the prototype classification layer optimizes the prototypes according to the regularization terms, while simultaneously classifying the input sequence based on its similarity to the updated prototypes. The comparison experiments are conducted on 26 UCR datasets compared with 10 baselines. The results show that our proposed approach achieves the best accuracy on 11 datasets. Specifically, our method outperforms PIP, CSSL, and LSS by an average of 16.33%, 9.77% and 5.96% on 22, 14 and 16 datasets, respectively. The interpretability experimental results and the application analysis on spectral data indicate that the learned prototypes can provide reasonable explanations for the classification results of the model.
Jianghui Cai, Haifeng Yang 0001, Chenhui Shi 0002, Min Zhang 0047, Jie Wang 0046, Xujun Zhao
Intell. Data Anal.8
2025 Three-way clustering based on the graph of local density trend
Haifeng Yang 0001, Jianghui Cai, Jie Wang 0046, Yaling Xun, Xujun Zhao
Int. J. Approx. Reason.7
2025 Cross-domain pedestrian trajectory prediction via behavioral pattern-aware multi-instance GCN
Haifeng Yang 0001, Jianghui Cai, Lichan Zhou, Jianing Tian, Yan Li 0152, Yaling Xun, Xujun Zhao
Knowl. Based Syst.9
2024 A new community detection method for simplified networks by combining structure and attribute information
Jianghui Cai, Haifeng Yang 0001, Xujun Zhao, Yaling Xun, Dongchao Zhang
Expert Syst. Appl.5
2024 A novel graph-attention based multimodal fusion network for joint classification of hyperspectral image and LiDAR data
Jianghui Cai, Min Zhang 0047, Haifeng Yang 0001, Yanting He, Chenhui Shi 0002, Xujun Zhao, Yaling Xun
Expert Syst. Appl.7
2023 Multi-scale fusion and adaptively attentive generative adversarial network for image de-raining
Haifeng Yang 0001, Yongjie Xin, Jianghui Cai, Min Zhang 0047, Xujun Zhao, Yingyue Zhao, Yanting He
Appl. Intell.6
2023 A review on semi-supervised clustering
Jianghui Cai, Haifeng Yang 0001, Xujun Zhao
Inf. Sci.4
2023 A new interest extraction method based on multi-head attention mechanism for CTR prediction
Haifeng Yang 0001, Linjing Yao, Jianghui Cai, Xujun Zhao
Knowl. Inf. Syst.5
2023 A New MC-LSTM Network Structure Designed for Regression Prediction of Time Series
Haifeng Yang 0001, Juanjuan Hu, Jianghui Cai, Xin Chen 0070, Xujun Zhao
Neural Process. Lett.6
2022 Vehicle anomalous trajectory detection algorithm based on road network partition
Xujun Zhao, Jianhua Su, Jianghui Cai, Haifeng Yang 0001, Ting-ting Xi
Appl. Intell.1
2022 ISBFK-means: A new clustering algorithm based on influence space
Jianghui Cai, Haifeng Yang 0001, Xujun Zhao
Expert Syst. Appl.5
2022 Density clustering with divergence distance and automatic center selection
Jianghui Cai, Haifeng Yang 0001, Xujun Zhao
Inf. Sci.4
2022 ARIS: A Noise Insensitive Data Pre-Processing Scheme for Data Reduction Using Influence Space
abstract
The extensive growth of data quantity has posed many challenges to data analysis and retrieval. Noise and redundancy are typical representatives of the above-mentioned challenges, which may reduce the reliability of analysis and retrieval results and increase storage and computing overhead. To solve the above problems, a two-stage data pre-processing framework for noise identification and data reduction, called ARIS, is proposed in this article. The first stage identifies and removes noises by the following steps: First, the influence space (IS) is introduced to elaborate data distribution. Second, a ranking factor (RF) is defined to describe the possibility that the points are regarded as noises, then, the definition of noise is given based on RF. Third, a clean dataset (CD) is obtained by removing noise from the original dataset. The second stage learns representative data and realizes data reduction. In this process, CD is divided into multiple small regions by IS. Then the reduced dataset is formed by collecting the representations of each region. The performance of ARIS is verified by experiments on artificial and real datasets. Experimental results show that ARIS effectively weakens the impact of noise and reduces the amount of data and significantly improves the accuracy of data analysis within a reasonable time cost range.
Jianghui Cai, Haifeng Yang 0001, Xujun Zhao
ACM Trans. Knowl. Discov. Data4
2021 Outlier detection from multiple data sources
Xujun Zhao, Chaowei Zhang 0001, Jifu Zhang, Xiao Qin 0001
Inf. Sci.2
2020 TAD: A trajectory clustering algorithm based on spatial-temporal density analysis
abstract
In this paper, a novel trajectory clustering algorithm - TAD - is proposed to extract trajectory Stays based on spatial-temporal density analysis of data. Two new metrics - NMAST (Neighbourhood Move Ability and Stay Time) density function and NT (Noise Tolerance) factor - are defined in this algorithm. Firstly, NMAST integrates the characteristics of Neighbourhood Move Ability (NMA, extended from the concept of Move Ability MA), Stay Time (ST), and evaluation factor Eμ to measure the spatial-temporal density of data. Secondly, NT utilizes the features of noise to dynamically evaluate and reduce the influence of noise. The experimental results on Geolife dataset shows that the distributions hidden in data are extracted more realistically, especially for various complex or special trajectories with long-duration gaps. Furthermore, our analytical method of trajectory data is particularly applied in the spectra of LAMOST survey to analyse the variation characteristics of sky-background. The results show a regular distribution on observational date which is relatively concentrated in the month of 1, 10, 11, 12 in each year. The laws discovered in this work would provide a reasonable support for the designation of observational plans, and the new trajectory analysis method would also provide the services for the astronomical data analysis and then for the further studies of formation and evolution of the universe.
Jianghui Cai, Haifeng Yang 0001, Jifu Zhang, Xujun Zhao
Expert Syst. Appl.5
2019 Parallel mining of contextual outlier using sparse subspace
Xujun Zhao, Jifu Zhang, Xiao Qin 0001, Jianghui Cai
Expert Syst. Appl.1
2018 kNN-DP: Handling Data Skewness in kNN Joins Using MapReduce
abstract
In this study, we discover that the data skewness problem imposes adverse impacts on MapReduce-based parallel kNN-join operations running clusters. We propose a data partitioning approach-called kNN-DP-to alleviate load imbalance incurred by data skewness. The overarching goal of kNN-DP is to equally divide data objects into a large number of partitions, which are processed by mappers and reducers in parallel. At the heart of kNN-DP is a data partitioning module, which dynamically and judiciously partitions data to optimize kNN-join performance by suppressing data skewness on Hadoop clusters. Data partitioning decisions largely depends on data properties (e.g., distributions), the analysis of which is highly expensive for a massive amount of data. To speed up the data-property analysis, we incorporate a sampling technique to profile the data distribution of a small sample dataset representing big datasets. After building a data-partitioning cost model for parallel kNN-joins, we derive the time-complexity upper and lower bounds of parallel kNN-join algorithms. The cost model offers us a guidance to systematically investigate kNN-DP's performance. kNN-DP obtains global nearest neighbors using local nearest neighbors. To improve the accuracy of such an approximation solution, we augment each node's local data by a small amount of redundant data. We develop two kNN-DP-based schemes called LSH+ and z-value+, which seamlessly integrate kNN-DP with the existing LSH and z-value algorithms for kNN-join computing. We implement and evaluate LSH+ and z-value+ on a 24-node Hadoop cluster driven by both synthetic and real-world high-dimensional datasets. The experimental results show that kNN-DP significantly improves the performance of LSH and z-value while offering high extensibility and scalability on Hadoop clusters.
Xujun Zhao, Jifu Zhang, Xiao Qin 0001
IEEE Trans. Parallel Distributed Syst.1
2017 LOMA: A local outlier mining algorithm based on attribute relevance analysis
Xujun Zhao, Jifu Zhang, Xiao Qin 0001
Expert Syst. Appl.1
2017 FiDoop-DP: Data Partitioning in Frequent Itemset Mining on Hadoop Clusters
abstract
Traditional parallel algorithms for mining frequent itemsets aim to balance load by equally partitioning data among a group of computing nodes. We start this study by discovering a serious performance problem of the existing parallel Frequent Itemset Mining algorithms. Given a large dataset, data partitioning strategies in the existing solutions suffer high communication and mining overhead induced by redundant transactions transmitted among computing nodes. We address this problem by developing a data partitioning approach called FiDoop-DP using the MapReduce programming model. The overarching goal of FiDoop-DP is to boost the performance of parallel Frequent Itemset Mining on Hadoop clusters. At the heart of FiDoop-DP is the Voronoi diagram-based data partitioning technique, which exploits correlations among transactions. Incorporating the similarity metric and the Locality-Sensitive Hashing technique, FiDoop-DP places highly similar transactions into a data partition to improve locality without creating an excessive number of redundant transactions. We implement FiDoop-DP on a 24-node Hadoop cluster, driven by a wide range of datasets created by IBM Quest Market-Basket Synthetic Data Generator. Experimental results reveal that FiDoop-DP is conducive to reducing network and computing loads by the virtue of eliminating redundant transactions on Hadoop nodes. FiDoop-DP significantly improves the performance of the existing parallel frequent-pattern scheme by up to 31 percent with an average of 18 percent.
Yaling Xun, Jifu Zhang, Xiao Qin 0001, Xujun Zhao
IEEE Trans. Parallel Distributed Syst.4
2013 Interrelation analysis of celestial spectra data using constrained frequent pattern trees
Jifu Zhang, Xujun Zhao, Sulan Zhang, Shu Yin 0001, Xiao Qin 0001
Knowl. Based Syst.2