Jifu Zhang

dblp:34/6923 · DBLP profile ↗
← Back
66ranked-venue papers
4as first author
38since 2021 · last 2026
0000-0002-0396-8901ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 27 · 3 first-author · 17 since 2021Systems, architecture and hardware · 13 · 4 since 2021Databases, data management, data science and information retrieval · 10 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Computer networks · 1Security and privacy · 1 · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2026 Multi-source Anomaly Detection Using Feature Selection and Relevant Subspace
Xujun Zhao, Jifu Zhang, Jianghui Cai, Haifeng Yang 0001
ICIC3
2026 Improved differential and linear cryptanalysis on round-reduced SIMON
Muzhou Li, Jifu Zhang
Des. Codes Cryptogr.3
2026 Dynamic UAV task offloading combining deep reinforcement learning and two-stage stochastic optimization
Yaling Xun, Jifu Zhang, Haifeng Yang 0001, Jianghui Cai
Future Gener. Comput. Syst.3
2026 Multimodal interpretable image recognition network via language-guided global-local collaboratively alignment
Sulan Zhang, Peijun Zhang, Lihua Hu, Jifu Zhang
Knowl. Based Syst.5
2026 Adaptive deep multi-view clustering via quality-aware representation learning
Xueying Niu, Jifu Zhang
Knowl. Based Syst.4
2026 A progressive attention network with transformer for multi-label image recognition
Sulan Zhang, Zhenwen Liao, Jianeng Li, Lihua Hu, Jifu Zhang
Pattern Recognit.5
2025 A Relevant Subspace-Based Contextual Outlier Detection Using Chebyshev Coulomb Resultant Force
Jianying Liu, Jifu Zhang
ICIC (8)3
2025 Similarity Attribute-Based Categorical Attribute Grouping for Outlier Detecting
Yijing Song, Jianying Liu, Min Zhang 0049, Xiao Qin 0001, Jifu Zhang
ICIC (8)5
2025 Weakly supervised semantic segmentation for ancient architecture based on multiscale adaptive fusion and spectral clustering
Ruifei Sun, Sulan Zhang, Meihong Su, Lihua Hu, Jifu Zhang
Comput. Graph.5
2025 Attribute Grouping-Based Outlier Detection Using Dynamic Dimensional Embedding Representations
abstract
ABSTRACT Attribute grouping serves as one of the effective steps in high‐dimensional outlier detection. However, existing outlier detection methods merely capture the local coupling between pairs of attribute values within each attribute group, and overlook the global coupling among combinations of multiple attribute values, thereby limiting the effectiveness of outlier detection. In this paper, a novel attribute grouping‐based outlier detection method for categorical data is proposed by using dynamic dimensional embedding representations to characterize the global coupling among all variables. First, within each attribute group, we dynamically construct the embedding dimensions of all attributes to avoid redundant representations caused by static dimensions. Second, Multi‐Head Self‐Attention is used to construct the dynamic dimensional embedding of any data object, effectively characterizing the global coupling. Then, the reconstruction error derived from the dynamic dimensional embedding is employed to quantify the outlier degree of data objects. Based on this, we propose an attribute grouping‐based outlier detection method, in which dynamic dimensional embedding representations are used for each attribute group. In the end, experimental results on the UCI and synthetic datasets validate that the algorithm has good outlier detection performance. Importantly, compared with the competitive methods, the algorithm bolsters the AUC index and the detection efficiency by an average of 7.76% and 48.49%, respectively.
Yijing Song, Jifu Zhang
Concurr. Comput. Pract. Exp.3
2025 Two-Stage Feature Selection for Fine-Grained Image Recognition Via Partial Order Analysis and Heterogeneity Evaluation
abstract
ABSTRACT The core challenge of fine‐grained image recognition (FGIR) tasks is distinguishing highly similar subclasses within the same base category. Most CNN‐based deep learning methods typically focus on extracting information from local regions while overlook the inherent structure between subclasses and the complex relationships between features. This paper presents a two‐stage feature selection method based on partial order analysis (POA) and heterogeneity evaluation (HE) for FGIR tasks, guiding the model to focus on distinctive features while reducing uncertainty caused by interfering information. Specifically, in the POA stage, clustering first groups similar subcategories into a medium‐granularity category. Formal concept analysis then models their hierarchical partial order, identifying “shared features” among subcategories and “exclusive features” unique to each. This structured representation highlights key contrastive cues. In the HE stage, a novel heterogeneity index is introduced to measure the fluctuation of low‐level features within each fine‐grained category. This index guides the model to suppress pseudo‐discriminative features with high heterogeneity, mitigating the impact of noisy and unstable information on decision‐making. We perform comprehensive experiments on three commonly used benchmark datasets (CUB‐200‐2011, Stanford Cars, and FGVC‐Aircraft). Experimental results show that the proposed method outperforms classic FGIC methods, validating the effectiveness of our approach.
Hongli Gao, Sulan Zhang, Huiyuan Zhou, Lihua Hu, Jifu Zhang
IET Image Process.5
2025 Sensitivity-propagated dual-frequency graph neural network for multivariate time series forecasting
Yaling Xun, Jianghui Cai, Haifeng Yang 0001, Jifu Zhang
Neurocomputing5
2025 ICE: Incremental Subspace Clustering of High-Dimensional Categorical Data
abstract
Subspace clustering is an effective way to analyze high-dimensional data. The main problems of the conventional subspace clustering techniques are as follows: first, conventional clustering methods can not describe categorical attribute space in more detail; second, most subspace-clustering techniques failure to process dynamic data effectively; finally, lack of effective noise recognition leads to the decline of the efficiency of incremental subspace-clustering analysis. We address the above problems by an incremental subspace-clustering algorithm — called ICE. With attribute subspace constructed by a rough set-based weight computing method, ICE obtains clustering results through initial and incremental clustering stage. Utilizing the original cluster results generated from initial clustering stage, we adopt merging and splitting operation to dynamic adjust cluster-structure in incremental clustering stage. Before achieving the final results, a polymerization-based noise recognition technique is employed to automatically identify noise from sparse clusters without human threshold intervention. We implement ICE on synthetic and real-world datasets. The experimental results reveal that incremental subspace-clustering method can achieves satisfactory performance on extensibility, accuracy and robustness.
Ning Pang, Chaowei Zhang 0001, Jifu Zhang, Xiao Qin 0001
Int. J. Uncertain. Fuzziness Knowl. Based Syst.3
2025 ARFIS: An adaptive robust model for regression with heavy-tailed distribution
Meihong Su, Jifu Zhang, Yaqing Guo, Wenjian Wang 0001
Inf. Sci.2
2025 Reinforcement negative sampling recommendation based on collaborative knowledge graph
Yaling Xun, Jifu Zhang
J. Intell. Inf. Syst.3
2025 Joint entropy minimization-based statistical dependency for deep multi-view clustering
Xueying Niu, Jifu Zhang
Knowl. Inf. Syst.4
2025 Context feature fusion and enhanced non-maximum suppression for pedestrian detection in crowded scenes
Lihua Hu, Jifu Zhang, Xinbo Wang
Multim. Tools Appl.4
2025 Similarity Metrics: Chebyshev Coulomb Force and Resultant Force for High-Dimensional Data
abstract
The similarity metric has garnered widespread attention thanks to its potential applications in the fields of data mining, machine learning, and so on. Due to the interference of “distance concentration” caused by “Curse of dimensionality,” however, existing similarity metrics are inadequate in high-dimensional data analysis. In this study, we propose two innovative similarity metrics—Chebyshev Coulomb force and Chebyshev Coulomb resultant force—anchored on Chebyshev p-norms. In the initial phase, we eliminate dependency relationships among attributes by applying a metric matrix—and the theoretical analysis reveals that the Chebyshev p-norms is capable of mitigating the effect of “distance concentration” among high-dimensional data objects. Next, we devise two similarity metrics—Chebyshev Coulomb force and Chebyshev Coulomb resultant force—by adopting the metric matrix and Chebyshev p-norms. Chebyshev Coulomb force and Chebyshev Coulomb resultant force, being effective in characterizing the similarity among data objects, quantify the deviation of data objects from their respective dataset centers. Additionally, the two metrics alleviate the interference of “distance concentration.” Importantly, the discrepancy of data objects in attribute dimensions is captured by Chebyshev Coulomb force vector, rendering the similarity metric interpretable. By utilizing the UCI dataset, the experimental validation demonstrates the superiority of our similarity metrics, confirming their efficacy in mitigating the interference of “distance concentration.” Compared with the existing similarity metric approaches, the AUC index of outlier detection shows an average improvement of 8.18%—and the ARI, NMI, and F_score indices of clustering are revamped by averages 6.56%, 6.87%, and 6.01%, respectively.
Jian Ying Liu, Chaowei Zhang 0001, Min Zhang 0049, Xiao Qin 0001, Jifu Zhang
ACM Trans. Knowl. Discov. Data5
2024 Unsupervised single image-based depth estimation powered by coplanarity-driven disparity derivation
Xiaoling Yao, Lihua Hu, Jifu Zhang
Eng. Appl. Artif. Intell.4
2024 A multi-view ensemble clustering approach using joint entropy
Xueying Niu, Jifu Zhang
Expert Syst. Appl.4
2024 Higher-order embedded learning for heterogeneous information networks and adaptive POI recommendation
Yaling Xun, Jifu Zhang, Haifeng Yang 0001, Jianghui Cai
Inf. Process. Manag.3
2024 A Progressive Stacking Pseudoinverse Learning Framework via Active Learning in Random Subspaces
abstract
Stacking pseudoinverse learner (SP) is an ensemble learning technology, and its generalization performance greatly affects the effect of image classification. Currently, most SPs randomly initialize the input weight matrix in a random subspace without limiting the random initial values, resulting in unstable training results and a decrease in generalization performance; in addition, training all samples at once may cause the classifier redundant and also affect the generalization performance of the model. To efficiently address the above issues, we propose a new framework called progressive stacking pseudoinverse learner (PSP), which aims to enhance the generalization performance of SP via active learning (AL) in random subspaces. Specifically, on the one hand, a random feature SP (RFSP) model is proposed, which constrains the random subspace by initializing the input weight matrix into different random specific distributions to improve the generalization performance of SP. On the other hand, an AL progressive (ALP) model based on RFSP is proposed. By iteratively selecting useful samples to optimize the classification results, the training sample information is effectively used to progressively enhance the generalization performance of the model. Experimental results on three public datasets show that our proposed PSP algorithm achieves better performance in accuracy, precision, recall, and$F1$score, and the results are competitive with state-of-the-art methods.
Zhenjiao Cai, Sulan Zhang, Ping Guo 0002, Jifu Zhang, Lihua Hu
IEEE Trans. Syst. Man Cybern. Syst.4
2024 AM-RP Stacking PILers: Random projection stacking pseudoinverse learning algorithm based on attention mechanism
Zhenjiao Cai, Sulan Zhang, Ping Guo 0002, Jifu Zhang, Lihua Hu
Vis. Comput.4
2023 A multi-view ensemble clustering approach using joint affinity matrix
Xueying Niu, Chaowei Zhang 0001, Lihua Hu, Jifu Zhang
Expert Syst. Appl.5
2023 A density connection weight-based clustering approach for dataset with density-sparse region
Min Zhang 0049, Junli Li 0005, Jifu Zhang
Expert Syst. Appl.4
2023 A user-guided reduction concept lattice and its algebraic structure
Sulan Zhang, Jifu Zhang, Jianeng Li, Ping Guo 0002, Witold Pedrycz
Expert Syst. Appl.2
2023 A multi-view subspace representation learning approach powered by subspace transformation relationship
Xueying Niu, Chaowei Zhang 0001, Lihua Hu, Jifu Zhang
Knowl. Based Syst.5
2023 Multi-level Self-supervised Representation Learning via Triple-way Attention Fusion and Local Similarity Optimization
Sulan Zhang, Jifu Zhang, Aiqin Liu
Neural Process. Lett.3
2023 AdaHOSVD: an adaptive higher-order singular value decomposition method for point cloud denoising
Lihua Hu, Wenhao Liang, Yuting Bai, Jifu Zhang
Pattern Anal. Appl.4
2023 A High-Dimensional Outlier Detection Approach Based on Local Coulomb Force
abstract
Traditional outlier detections are inadequate for high-dimensional data analysis due to the interference of distance tending to be concentrated (“curse of dimensionality”). Inspired by the Coulomb’s law, we propose a new high-dimensional data similarity measure vector, which consists of outlier Coulomb force and outlier Coulomb resultant force. Outlier Coulomb force not only effectively gauges similarity measures among data objects, but also fully reflects differences among dimensions of data objects by vector projection in each dimension. More importantly, Coulomb resultant force can effectively measure deviations of data objects from a data center, making detection results interpretable. We introduce a new neighborhood outlier factor, which drives the development of a high-dimensional outlier detection algorithm. In our approach, attribute values with a high deviation degree is treated as interpretable information of outlier data. Finally, we implement and evaluate our algorithm using the UCI and synthetic datasets. Our experimental results show that the algorithm effectively alleviates the interference of “Curse of Dimensionality”. The findings confirm that high-dimensional outlier data originated by the algorithm are interpretable.
Pengyun Zhu, Chaowei Zhang 0001, Jifu Zhang, Xiao Qin 0001
IEEE Trans. Knowl. Data Eng.4
2022 MiCS-P: Parallel mutual-information computation of big categorical data on spark
Junli Li 0005, Chaowei Zhang 0001, Jifu Zhang, Xiao Qin 0001, Lihua Hu
J. Parallel Distributed Comput.3
2022 Image annotation of ancient chinese architecture based on visual attention mechanism and GCN
Sulan Zhang, Songzan Chen, Jifu Zhang, Zhenjiao Cai, Lihua Hu
Multim. Tools Appl.3
2022 GMC_FM : a grid and multi-density-based method for matching ancient Chinese architectural images
Lihua Hu, Yaoyao Nie, Jifu Zhang, Sulan Zhang
Mach. Vis. Appl.3
2021 A novel discretization algorithm based on multi-scale and information entropy
Yaling Xun, Qingxia Yin, Jifu Zhang, Haifeng Yang 0001, Xiaohui Cui
Appl. Intell.3
2021 KR-DBSCAN: A density-based clustering algorithm based on reverse nearest neighbor and influence space
Lihua Hu, Hongkai Liu, Jifu Zhang, Aiqin Liu
Expert Syst. Appl.3
2021 Incremental frequent itemsets mining based on frequent pattern tree and multi-scale
Yaling Xun, Xiaohui Cui, Jifu Zhang, Qingxia Yin
Expert Syst. Appl.3
2021 Outlier detection from multiple data sources
Xujun Zhao, Chaowei Zhang 0001, Jifu Zhang, Xiao Qin 0001
Inf. Sci.4
2021 HBPFP-DC: A parallel frequent itemset mining using Spark
Yaling Xun, Jifu Zhang, Haifeng Yang 0001, Xiao Qin 0001
Parallel Comput.2
2020 Computing Mutual Information of Big Categorical Data and Its Application to Feature Grouping
abstract
This paper develops a parallel computing system - MiCS - for mutual information of big categorical data on the Spark computing platform. The MiCS algorithm is conductive to processing a large amount and strong repeatability of mutual-information calculation among feature pairs by applying a column-wise transformation scheme. And to improve the efficiency of the MiCS and the utilization rate of Spark cluster resources, we adopt a virtual partitioning scheme to achieve balanced load while mitigating the data skewness problem in the Spark Shuffle process.
Junli Li 0005, Chaowei Zhang 0001, Jifu Zhang, Xiao Qin 0001
ICDE3
2020 TAD: A trajectory clustering algorithm based on spatial-temporal density analysis
abstract
In this paper, a novel trajectory clustering algorithm - TAD - is proposed to extract trajectory Stays based on spatial-temporal density analysis of data. Two new metrics - NMAST (Neighbourhood Move Ability and Stay Time) density function and NT (Noise Tolerance) factor - are defined in this algorithm. Firstly, NMAST integrates the characteristics of Neighbourhood Move Ability (NMA, extended from the concept of Move Ability MA), Stay Time (ST), and evaluation factor Eμ to measure the spatial-temporal density of data. Secondly, NT utilizes the features of noise to dynamically evaluate and reduce the influence of noise. The experimental results on Geolife dataset shows that the distributions hidden in data are extracted more realistically, especially for various complex or special trajectories with long-duration gaps. Furthermore, our analytical method of trajectory data is particularly applied in the spectra of LAMOST survey to analyse the variation characteristics of sky-background. The results show a regular distribution on observational date which is relatively concentrated in the month of 1, 10, 11, 12 in each year. The laws discovered in this work would provide a reasonable support for the designation of observational plans, and the new trajectory analysis method would also provide the services for the astronomical data analysis and then for the further studies of formation and evolution of the universe.
Jianghui Cai, Haifeng Yang 0001, Jifu Zhang, Xujun Zhao
Expert Syst. Appl.4
2020 EDOM: Improving energy efficiency of database operations on multicore servers
Yi Zhou 0009, Shubbhi Taneja, Xiao Qin 0001, Wei-Shinn Ku, Jifu Zhang
Future Gener. Comput. Syst.5
2020 ThermoBench: A thermal efficiency benchmark for clusters in data centers
Yi Zhou 0009, Yuanqi Chen, Shubbhi Taneja, Ajit Chavan, Xiao Qin 0001, Jifu Zhang
Parallel Comput.6
2020 Weighted Outlier Detection of High-Dimensional Categorical Data Using Feature Grouping
abstract
We propose a weighted outlier mining method called WATCH to identify outliers in high-dimensional categorical datasets. WATCH is composed of two distinctive modules: 1) feature grouping by the virtue of correlation measurement among features and 2) outlier mining by assigning scores to objects in each feature groups. At the heart of WATCH is the feature grouping module, which groups an array of features into multiple groups to discover various aspects of feature patterns in each group. The outlier mining module detects outliers from high-dimensional categorical datasets. Except for the number of outliers specified by users, WATCH is conducive to bypassing the optimization of any user-given parameter. We implement and evaluate WATCH using synthetic and real-world datasets. Our experimental results show that WATCH is a promising and practical algorithm to detect outliers in high-dimensional categorical datasets, because WATCH achieves high performance in terms of precision, efficiency, and interpretability.
Junli Li 0005, Jifu Zhang, Ning Pang, Xiao Qin 0001
IEEE Trans. Syst. Man Cybern. Syst.2
2020 Scalable Mining of Contextual Outliers Using Relevant Subspace
abstract
In this paper, we propose a scalable mining algorithm to discover contextual outliers using relevant subspaces. We develop the mining algorithm using the MapReduce programming model running on a Hadoop cluster. Relevant subspaces, which effectively capture the local distribution of various datasets, are quantified using local sparseness of attribute dimensions. We design a novel way of calculating local outlier factors in a relevant subspace with the probability density of local datasets; this new approach can effectively reflect the outlier degree of a data object that does not satisfy the distribution of the local dataset in the relevant subspace. Attribute dimensions of a relevant subspace, and local outlier factors are expressed as vital contextual information, which improves the interpretability of outliers. Importantly, the selection of N data objects with the largest local outlier factor value is categorized as contextual outliers in our solution. To this end, our scalable mining algorithm, which incorporates the locality sensitive hashing distributed strategy, is implemented on a Hadoop cluster. The experimental results validate the effectiveness, interpretability, scalability, and extensibility of the algorithm using both synthetic data and stellar spectral data as experimental datasets.
Jifu Zhang, Yaling Xun, Sulan Zhang, Xiao Qin 0001
IEEE Trans. Syst. Man Cybern. Syst.1
2019 PUMA: Parallel subspace clustering of categorical data using multi-attribute weights
Ning Pang, Jifu Zhang, Chaowei Zhang 0001, Xiao Qin 0001, Jianghui Cai
Expert Syst. Appl.2
2019 Parallel mining of contextual outlier using sparse subspace
Xujun Zhao, Jifu Zhang, Xiao Qin 0001, Jianghui Cai
Expert Syst. Appl.2
2019 Feature grouping-based parallel outlier mining of categorical data using spark
Junli Li 0005, Jifu Zhang, Xiao Qin 0001, Yaling Xun
Inf. Sci.2
2019 Parallel Hierarchical Subspace Clustering of Categorical Data
abstract
Parallel clustering is an important research area of big data analysis. The conventional Hierarchical Agglomerative Clustering (HAC) techniques are inadequate to handle big-scale categorical datasets due to two drawbacks. First, HAC consumes excessive CPU time and memory resources; and second, it is non-trivial to decompose clustering tasks into independent sub-tasks executed in parallel. We solve these two problems by a MapReduce-based hierarchical subspace-clustering algorithm - called PAPU - using LSH-based data partitioning. PAPU is conducive to partitioning a large-scale dataset into multiple independent sub-datasets, into which similar data objects are mapped. Advocating parallel computing, PAPU obtains sub-clusters corresponding to respective attribute subspaces from independent chunks in the local clustering phase. To improve the accuracy of approximated clustering results, PAPU measures various scale clusters by applying the hierarchical clustering scheme to iteratively merge sub-clusters during the global clustering phase. We implement PAPU on a 24-node Hadoop computing platform. The experimental results reveal that hierarchical subspace-clustering coupled with the data-partitioning strategy achieves high clustering efficiency on both synthetic and real-world large-scale datasets. The experiments also demonstrate that PAPU delivers superior performance in terms of extensibility and scalability (e.g., a nearly linear speedup).
Ning Pang, Jifu Zhang, Chaowei Zhang 0001, Xiao Qin 0001
IEEE Trans. Computers2
2018 Towards thermal-aware Hadoop clusters
Yi Zhou 0009, Shubbhi Taneja, Gautam Dudeja, Xiao Qin 0001, Jifu Zhang, Minghua Jiang, Mohammed I. Alghamdi
Future Gener. Comput. Syst.5
2018 kNN-DP: Handling Data Skewness in kNN Joins Using MapReduce
abstract
In this study, we discover that the data skewness problem imposes adverse impacts on MapReduce-based parallel kNN-join operations running clusters. We propose a data partitioning approach-called kNN-DP-to alleviate load imbalance incurred by data skewness. The overarching goal of kNN-DP is to equally divide data objects into a large number of partitions, which are processed by mappers and reducers in parallel. At the heart of kNN-DP is a data partitioning module, which dynamically and judiciously partitions data to optimize kNN-join performance by suppressing data skewness on Hadoop clusters. Data partitioning decisions largely depends on data properties (e.g., distributions), the analysis of which is highly expensive for a massive amount of data. To speed up the data-property analysis, we incorporate a sampling technique to profile the data distribution of a small sample dataset representing big datasets. After building a data-partitioning cost model for parallel kNN-joins, we derive the time-complexity upper and lower bounds of parallel kNN-join algorithms. The cost model offers us a guidance to systematically investigate kNN-DP's performance. kNN-DP obtains global nearest neighbors using local nearest neighbors. To improve the accuracy of such an approximation solution, we augment each node's local data by a small amount of redundant data. We develop two kNN-DP-based schemes called LSH+ and z-value+, which seamlessly integrate kNN-DP with the existing LSH and z-value algorithms for kNN-join computing. We implement and evaluate LSH+ and z-value+ on a 24-node Hadoop cluster driven by both synthetic and real-world high-dimensional datasets. The experimental results show that kNN-DP significantly improves the performance of LSH and z-value while offering high extensibility and scalability on Hadoop clusters.
Xujun Zhao, Jifu Zhang, Xiao Qin 0001
IEEE Trans. Parallel Distributed Syst.2
2017 LOMA: A local outlier mining algorithm based on attribute relevance analysis
Xujun Zhao, Jifu Zhang, Xiao Qin 0001
Expert Syst. Appl.2
2017 Towards two-phase scheduling of real-time applications in distributed systems
Mohammed I. Alghamdi, Xunfei Jiang, Ji Zhang 0002, Jifu Zhang, Minghua Jiang, Xiao Qin 0001
J. Netw. Comput. Appl.4
2017 A parallel algorithm for mining constrained frequent patterns using MapReduce
Xiaowu Yan, Jifu Zhang, Yaling Xun, Xiao Qin 0001
Soft Comput.2
2017 FiDoop-DP: Data Partitioning in Frequent Itemset Mining on Hadoop Clusters
abstract
Traditional parallel algorithms for mining frequent itemsets aim to balance load by equally partitioning data among a group of computing nodes. We start this study by discovering a serious performance problem of the existing parallel Frequent Itemset Mining algorithms. Given a large dataset, data partitioning strategies in the existing solutions suffer high communication and mining overhead induced by redundant transactions transmitted among computing nodes. We address this problem by developing a data partitioning approach called FiDoop-DP using the MapReduce programming model. The overarching goal of FiDoop-DP is to boost the performance of parallel Frequent Itemset Mining on Hadoop clusters. At the heart of FiDoop-DP is the Voronoi diagram-based data partitioning technique, which exploits correlations among transactions. Incorporating the similarity metric and the Locality-Sensitive Hashing technique, FiDoop-DP places highly similar transactions into a data partition to improve locality without creating an excessive number of redundant transactions. We implement FiDoop-DP on a 24-node Hadoop cluster, driven by a wide range of datasets created by IBM Quest Market-Basket Synthetic Data Generator. Experimental results reveal that FiDoop-DP is conducive to reducing network and computing loads by the virtue of eliminating redundant transactions on Hadoop nodes. FiDoop-DP significantly improves the performance of the existing parallel frequent-pattern scheme by up to 31 percent with an average of 18 percent.
Yaling Xun, Jifu Zhang, Xiao Qin 0001, Xujun Zhao
IEEE Trans. Parallel Distributed Syst.2
2016 Miner*: A Weighted Distance Sum based Outlier Mining System of Star Spectrum Data
abstract
Existing distance-based outlier mining methods do not consider the impact of each attribute's importance degree, thereby resulting in poor mining accuracies. To address this problem, we propose a new outlier mining algorithm – Miner* – that makes use of information entropy and Weighted Distance Sum to substantially improve mining accuracies. Miner* employs information entropy to determine weight values indicating the importance degrees of data attributes. An input dataset is reduced by Miner* through the neighbour-radius-based pruning technologies. Thus, Miner* obtains a candidate outlier set by removing any data objects that are unlikely to be outliers. Miner* calculates the weighted distance sum value Wkof each object in the candidate outlier set; Wkvalue ranks the top n to be regarded as outliers. Due to the sum of distance, which takes full advantage of the clustering characteristics of the dataset, edge distribution data objects and local outliers can be effectively mined out. To demonstrate the effectiveness of the Miner* algorithm, we implement Miner* in a prototype system to detect star spectrum data objects with abnormal characteristic lines. Our experimental results show that the algorithm in Miner* achieves high accuracy, high scalability, and low man-made influence by utilizing UCI and star spectrum dataset. Our results also confirm that Miner* is feasible and effective in mining spectrum data with abnormal characteristic lines from massive star spectrum dataset.
Chaowei Zhang 0001, Jifu Zhang, Xiao Qin 0001, Sulan Zhang
Int. J. Uncertain. Fuzziness Knowl. Based Syst.2
2016 A relevant subspace based contextual outlier mining algorithm
Jifu Zhang, Sulan Zhang, Yaling Xun, Xiao Qin 0001
Knowl. Based Syst.1
2016 A FWCL-based method for visual vocabulary formation
Sulan Zhang, Jifu Zhang, Ping Guo 0002, Meng Chu, Kai-Hsiung Chang
Multim. Tools Appl.2
2016 g-Good-neighbor conditional diagnosability measures for 3-ary n-cube networks
Jun Yuan 0001, Aixia Liu, Xiao Qin 0001, Jifu Zhang, Jing Li 0048
Theor. Comput. Sci.4
2016 TIGER: Thermal-Aware File Assignment in Storage Clusters
abstract
In this paper, we present a thermal-aware file assignment technique called TIGER for reducing the cooling cost of storage clusters in data centers. We show that peak inlet temperatures of storage nodes depend on not only CPU utilization but also I/O activities, which rely on file assignments in a cluster. The TIGER scheme aims to lower peak inlet temperatures of storage clusters by dynamic thermal management through file placements. TIGER makes use of cross-interference coefficients to estimate the re-circulation of hot air from the outlets to the inlets of data nodes. TIGER first calculates the thresholds of disks in each data node based on its contribution to heat re-circulation in a data center. TIGER undertakes two steps to achieve high I/O performance while reducing cooling cost. First, TIGER assigns groups of files with similar service times to shorten I/O response times. Second, TIGER ensures that load imbalance does not exceed a specified threshold. We evaluate performance of TIGER in terms of both cooling energy conservation and response time of a storage cluster. Our results confirm that TIGER reduces cooling-power requirements for clusters by offering about 10 to 15 percent cooling-energy savings without significantly degrading I/O performance.
Ajit Chavan, Mohammed I. Alghamdi, Xunfei Jiang, Xiao Qin 0001, Meikang Qiu, Minghua Jiang, Jifu Zhang
IEEE Trans. Parallel Distributed Syst.7
2016 FiDoop: Parallel Mining of Frequent Itemsets Using MapReduce
abstract
Existing parallel mining algorithms for frequent itemsets lack a mechanism that enables automatic parallelization, load balancing, data distribution, and fault tolerance on large clusters. As a solution to this problem, we design a parallel frequent itemsets mining algorithm called FiDoop using the MapReduce programming model. To achieve compressed storage and avoid building conditional pattern bases, FiDoop incorporates the frequent items ultrametric tree, rather than conventional FP trees. In FiDoop, three MapReduce jobs are implemented to complete the mining task. In the crucial third MapReduce job, the mappers independently decompose itemsets, the reducers perform combination operations by constructing small ultrametric trees, and the actual mining of these trees separately. We implement FiDoop on our in-house Hadoop cluster. We show that FiDoop on the cluster is sensitive to data distribution and dimensions, because itemsets with different lengths have different decomposition and construction costs. To improve FiDoop's performance, we develop a workload balance metric to measure load balance across the cluster's computing nodes. We develop FiDoop-HD, an extension of FiDoop, to speed up the mining performance for high-dimensional data analysis. Extensive experiments using real-world celestial spectral data demonstrate that our proposed solution is efficient and scalable.
Yaling Xun, Jifu Zhang, Xiao Qin 0001
IEEE Trans. Syst. Man Cybern. Syst.2
2015 The g-Good-Neighbor Conditional Diagnosability of k-Ary n-Cubes under the PMC Modeland MM* Model
abstract
The diagnosability of a system is defined as the maximum number of faulty processors that the system can guarantee to identify, which plays an important role in measuring of the reliability of multiprocessor systems. In the work of Peng et al. in 2012, they proposed a new measure for fault diagnosis of systems, namely,$g$-good-neighbor conditional diagnosability. It is defined as the diagnosability of a multiprocessor system under the assumption that every fault-free node contains at least$g$fault-free neighbors, which can measure the reliability of interconnection networks in heterogeneous environments more accurately than traditional diagnosability. The$k$-ary$n$-cube is a family of popular networks. In this study, we first investigate and determine the$R_g$-connectivity of$k$-ary$n$-cube for$0\le g\le n.$Based on this, we determine the$g$-good-neighbor conditional diagnosability of$k$-ary$n$-cube under the PMC model and MM* model for$k\ge 4, n\ge 3$and$0\le g\le n.$Our study shows the$g$-good-neighbor conditional diagnosability of$k$-ary$n$-cube is several times larger than the classical diagnosability of$k$-ary$n$-cube.
Jun Yuan 0001, Aixia Liu, Xiao Qin 0001, Jifu Zhang
IEEE Trans. Parallel Distributed Syst.6
2013 TIGER: Thermal-aware file assignment in storage clusters
abstract
In this paper, we present thermal-aware file assignment technique called TIGER for reducing cooling cost of storage clusters in data centers. TIGER first calculates the thresholds of disks in each node based on its contribution to heat recirculation in a data center. Next, TIGER assigns files to data nodes according to calculated thresholds. We evaluated performance of TIGER in terms of both cooling energy conservation and response time of a storage cluster. Our results confirm that TIGER reduces cooling-power requirements for clusters by offering about 10 to 15 percent cooling energy savings without significantly degrading I/O performance.
Ajit Chavan, Xunfei Jiang, Mohammed I. Alghamdi, Xiao Qin 0001, Minghua Jiang, Jifu Zhang
MSST6
2013 Interrelation analysis of celestial spectra data using constrained frequent pattern trees
Jifu Zhang, Xujun Zhao, Sulan Zhang, Shu Yin 0001, Xiao Qin 0001
Knowl. Based Syst.1
2012 A completeness analysis of frequent weighted concept lattices and their algebraic properties
Sulan Zhang, Ping Guo 0002, Jifu Zhang, Witold Pedrycz
Data Knowl. Eng.3
2009 A concept lattice based outlier mining method in low-dimensional subspaces
Jifu Zhang, Yiyong Jiang, Kai-Hsiung Chang, Sulan Zhang, Jianghui Cai, Lihua Hu
Pattern Recognit. Lett.1
2006 A Pruning Based Incremental Construction of Horizontal Partitioned Concept Lattice
Lihua Hu, Jifu Zhang, Sulan Zhang
ICIC (2)2