Jifu Zhang

dblp:34/6923 · DBLP profile ↗
← Back
10ranked-venue papers in the field
0as first author
7since 2021 · last 2025
0000-0002-0396-8901ORCID · verified

Domains — venue-derived; a paper can count in several

Knowledge Engineering, Semantic Web & Information Systems · 4Database Systems & Data Management · 3Data Mining & Knowledge Discovery · 2Information Retrieval & Web Search · 1
YearPublicationVenuePosition
2025 ARFIS: An adaptive robust model for regression with heavy-tailed distribution
Meihong Su, Jifu Zhang, Yaqing Guo, Wenjian Wang 0001
Inf. Sci.2
2025 Reinforcement negative sampling recommendation based on collaborative knowledge graph
Yaling Xun, Jifu Zhang
J. Intell. Inf. Syst.3
2025 Joint entropy minimization-based statistical dependency for deep multi-view clustering
Xueying Niu, Jifu Zhang
Knowl. Inf. Syst.4
2025 Similarity Metrics: Chebyshev Coulomb Force and Resultant Force for High-Dimensional Data
abstract
The similarity metric has garnered widespread attention thanks to its potential applications in the fields of data mining, machine learning, and so on. Due to the interference of “distance concentration” caused by “Curse of dimensionality,” however, existing similarity metrics are inadequate in high-dimensional data analysis. In this study, we propose two innovative similarity metrics—Chebyshev Coulomb force and Chebyshev Coulomb resultant force—anchored on Chebyshev p-norms. In the initial phase, we eliminate dependency relationships among attributes by applying a metric matrix—and the theoretical analysis reveals that the Chebyshev p-norms is capable of mitigating the effect of “distance concentration” among high-dimensional data objects. Next, we devise two similarity metrics—Chebyshev Coulomb force and Chebyshev Coulomb resultant force—by adopting the metric matrix and Chebyshev p-norms. Chebyshev Coulomb force and Chebyshev Coulomb resultant force, being effective in characterizing the similarity among data objects, quantify the deviation of data objects from their respective dataset centers. Additionally, the two metrics alleviate the interference of “distance concentration.” Importantly, the discrepancy of data objects in attribute dimensions is captured by Chebyshev Coulomb force vector, rendering the similarity metric interpretable. By utilizing the UCI dataset, the experimental validation demonstrates the superiority of our similarity metrics, confirming their efficacy in mitigating the interference of “distance concentration.” Compared with the existing similarity metric approaches, the AUC index of outlier detection shows an average improvement of 8.18%—and the ARI, NMI, and F_score indices of clustering are revamped by averages 6.56%, 6.87%, and 6.01%, respectively.
Jian Ying Liu, Chaowei Zhang 0001, Min Zhang 0049, Xiao Qin 0001, Jifu Zhang
ACM Trans. Knowl. Discov. Data5
2024 Higher-order embedded learning for heterogeneous information networks and adaptive POI recommendation
Yaling Xun, Jifu Zhang, Haifeng Yang 0001, Jianghui Cai
Inf. Process. Manag.3
2023 A High-Dimensional Outlier Detection Approach Based on Local Coulomb Force
abstract
Traditional outlier detections are inadequate for high-dimensional data analysis due to the interference of distance tending to be concentrated (“curse of dimensionality”). Inspired by the Coulomb’s law, we propose a new high-dimensional data similarity measure vector, which consists of outlier Coulomb force and outlier Coulomb resultant force. Outlier Coulomb force not only effectively gauges similarity measures among data objects, but also fully reflects differences among dimensions of data objects by vector projection in each dimension. More importantly, Coulomb resultant force can effectively measure deviations of data objects from a data center, making detection results interpretable. We introduce a new neighborhood outlier factor, which drives the development of a high-dimensional outlier detection algorithm. In our approach, attribute values with a high deviation degree is treated as interpretable information of outlier data. Finally, we implement and evaluate our algorithm using the UCI and synthetic datasets. Our experimental results show that the algorithm effectively alleviates the interference of “Curse of Dimensionality”. The findings confirm that high-dimensional outlier data originated by the algorithm are interpretable.
Pengyun Zhu, Chaowei Zhang 0001, Jifu Zhang, Xiao Qin 0001
IEEE Trans. Knowl. Data Eng.4
2021 Outlier detection from multiple data sources
Xujun Zhao, Chaowei Zhang 0001, Jifu Zhang, Xiao Qin 0001
Inf. Sci.4
2020 Computing Mutual Information of Big Categorical Data and Its Application to Feature Grouping
abstract
This paper develops a parallel computing system - MiCS - for mutual information of big categorical data on the Spark computing platform. The MiCS algorithm is conductive to processing a large amount and strong repeatability of mutual-information calculation among feature pairs by applying a column-wise transformation scheme. And to improve the efficiency of the MiCS and the utilization rate of Spark cluster resources, we adopt a virtual partitioning scheme to achieve balanced load while mitigating the data skewness problem in the Spark Shuffle process.
Junli Li 0005, Chaowei Zhang 0001, Jifu Zhang, Xiao Qin 0001
ICDE3
2019 Feature grouping-based parallel outlier mining of categorical data using spark
Junli Li 0005, Jifu Zhang, Xiao Qin 0001, Yaling Xun
Inf. Sci.2
2012 A completeness analysis of frequent weighted concept lattices and their algebraic properties
Sulan Zhang, Ping Guo 0002, Jifu Zhang, Witold Pedrycz
Data Knowl. Eng.3