Jun Zhang 0003

dblp:z/JunZhang3 · DBLP profile ↗
← Back
29ranked-venue papers in the field
2as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — venue-derived; a paper can count in several

Database Systems & Data Management · 11 (1 first)Knowledge Engineering, Semantic Web & Information Systems · 11Information Retrieval & Web Search · 6 (1 first)Data Mining & Knowledge Discovery · 1
YearPublicationVenuePosition
2026 FedMM: Federated Collaborative Signal Quantization for Multi-Market CTR Prediction
abstract
Online platforms such as Amazon and Netflix serve users across multiple countries and regions, underscoring the importance of multi-market recommendation (MMR). Most MMR methods adopt a pre-training and fine-tuning paradigm, in which a unified model is first trained on centralized, global data and subsequently adapted to specific markets. However, this approach ignores the privacy of market data. While traditional federated learning preserves privacy, it typically aims to obtain a global model by aggregating model parameters and does not account for significant market heterogeneity. Additionally, because ID spaces are disjoint across markets, embedding-based aggregation strategies become ineffective. To overcome these challenges, we propose a federated collaborative signal quantization (FedMM) method for multi-market click-through rate (CTR) prediction. Our core idea leverages a discrete codebook mechanism to achieve privacy-preserving transmission and align disjoint ID spaces. We further employ a hierarchical codebook structure to capture cross-market shared patterns and market-specific characteristics. Specifically, we deploy a residual quantized variational autoencoder (RQ-VAE) with a dual-layer codebook mechanism for each market to quantize collaborative embeddings. The first layer utilizes a global federated codebook, updated via aggregation to capture universally shared collaborative patterns, while the second layer maintains a local codebook to learn market-specific semantics. Finally, the learned discrete codes, which integrate both general and specific collaborative signals, are incorporated into downstream CTR models to enhance prediction accuracy across all markets. Extensive experiments on benchmark datasets demonstrate that FedMM significantly improves recommendation performance with privacy guarantees.
Jun Zhang 0003, Dugang Liu, Xing Tang 0007, Xiuqiang He 0001, Zhong Ming 0001
SIGIR1
2026 Large language model as meta-surrogate for offline data-driven many-task optimization: A proof-of-principle study
Xian-Rong Zhang, Yue-Jiao Gong, Yuanting Zhong, Ting Huang 0001, Jun Zhang 0003
Inf. Sci.5
2025 A Comparative Study on Sub-route Merging Ways for Clustering Assisted Ant Colony Optimization to Solve Large-Scale Traveling Salesman Problem
Zhongheng Jiang, Qiang Yang 0008, Danting Duan, Zhenyu Lu 0002, Jun Zhang 0003
WISE (2)5
2025 A Comparative Analysis of Ant Colony Optimization for Mobile Robot Route Optimization
Wen-Jun Zheng, Qiang Yang 0008, Danting Duan, Zhenyu Lu 0002, Jun Zhang 0003
WISE (2)5
2025 A probabilistic tournament learning swarm optimizer for large-scale optimization
Li-Ting Xu, Qiang Yang 0008, Jian-Yu Li, Peilan Xu, Xin Lin 0004, Xu-Dong Gao 0003, Zhenyu Lu 0002, Jun Zhang 0003
Inf. Sci.8
2024 A Benchmark Test Suite for Multiple Traveling Salesmen Problem with Pivot Cities
Zi-Yang Bo, Danting Duan, Qiang Yang 0008, Xu-Dong Gao 0003, Peilan Xu, Xin Lin 0004, Zhenyu Lu 0002, Jun Zhang 0003
WISE (4)8
2024 Niche center identification differential evolution for multimodal optimization problems
Shao-Min Liang, Zijia Wang 0001, Yi-Biao Huang, Zhi-hui Zhan, Sam Kwong, Jun Zhang 0003
Inf. Sci.6
2024 EvoS&R: Evolving Multiple Seeds and Radii for Varying Density Data Clustering
abstract
Density clustering has shown advantages over other types of clustering methods for processing arbitrarily shaped datasets. In recent years, extensive research efforts has been made on the improvements of DBSCAN or the algorithms incorporating the concept of density peaks. However, these previous studies remain the problems of being sensitive to the parameter settings, and some of them will stuck in weak results when encountering the situations of varying-density distributions. To overcome these issues, we propose an evolution framework named EvoS&R that evolves multiple seeds and the corresponding radii for varying-density data clustering. Compared with the traditional methods, EvoS&R handles the parameter tuning and multi-density fitting problems in an integrated and straightforward manner. Note that, however, the underlying task in EvoS&R is a mixed-variable optimization problem that is challenging in nature. We specifically design a hybrid encoding differential evolution algorithm with novel encoding, mutation, etc., to solve the optimization problem efficiently. Extensive experiments on density-based datasets shows that our algorithm outperforms the other state-of-the-arts in most cases, which validates the effectiveness of the proposed method.
Jun-Xian Chen, Yue-Jiao Gong, Weineng Chen, Jun Zhang 0003
IEEE Trans. Knowl. Data Eng.4
2023 A Privacy-Preserving Evolutionary Computation Framework for Feature Selection
Jian-Yu Li, Xiao Fang Liu, Qiang Yang 0008, Zhi-hui Zhan, Jun Zhang 0003
WISE6
2023 Heterogeneous cognitive learning particle swarm optimization for large-scale optimization problems
En Zhang, Zihao Nie, Qiang Yang 0008, Yiqiao Wang 0002, Dong Liu 0008, Sang-Woon Jeon, Jun Zhang 0003
Inf. Sci.7
2023 Enhanced Multi-Task Learning and Knowledge Graph-Based Recommender System
abstract
In recent years, themulti-task learning forknowledge graph-basedrecommender system, termed MKR, has shown its promising performance and has attracted increasing interest, because a recommendation task and a knowledge graph embedding (KGE) task can help each other to improve the recommendation. However, MKR still has two difficult issues. The first is how fully to capture users’ historical behavior pattern in the recommendation task and how fully to utilize deep multi-relation semantic information in the KGE task. The second is how to deal with datasets with different sparsity. Tackling these challenging issues, this paper proposes an enhanced MKR (EMKR) approach with two novelties. First, we propose to utilize the attention mechanism to aggregate users’ historical behavior for more accurately mining preferences in the recommendation task, and utilize the relation-aware graph convolutional neural network to fully capture the deep multi-relation neighborhood features in the KGE task, so as to address the first issue. Second, a two-part modeling strategy is proposed for a better representation of users in the recommendation task to expand the expressive ability of the model for adapting to datasets with different sparsity, so as to address the second issue. Extensive experiments are conducted on widely-used datasets and 11 approaches are used for comparison. The results show that the proposed EMKR can achieve substantial gains over the compared state-of-the-art approaches, especially in the situation where user-item interactions are sparse.
Min Gao 0012, Jian-Yu Li, Chun-Hua Chen 0002, Yun Li 0002, Jun Zhang 0003, Zhi-hui Zhan
IEEE Trans. Knowl. Data Eng.5
2022 DSGA: A Distributed Segment-Based Genetic Algorithm for Multi-Objective Outsourced Database Partitioning
Yong-Feng Ge, Zhi-hui Zhan, Jinli Cao, Hua Wang 0002, Yanchun Zhang, Kuei-Kuei Lai, Jun Zhang 0003
Inf. Sci.7
2022 Random neighbor elite guided differential evolution for global numerical optimization
Qiang Yang 0008, Xu-Dong Gao 0003, Dong-Dong Xu, Zhenyu Lu 0002, Jun Zhang 0003
Inf. Sci.6
2022 Toward Predicting Active Participants in Tweet Streams: A Case Study on Two Civil Rights Events
abstract
Online social media have aroused much research interest in recent years. In contrast to previous work that focused on the detection of emerging topics, this article undertakes the prediction of active users in online social events, which is so far rarely explored. This prediction task is formulated as a binary classification problem that built on real-world tweet streams, taking Ferguson event and New York Chockhold event as examples. Then, a comprehensive user feature system is designed to characterize the events’ online participants, which includes not only basic statistical characteristics and image-pixel-level features, but also some emotional features and personality features. Next, the Weighted Random Forest (Weighted-RF) classifier is adopted to solve the classification problem. Based on the user feature system and the classifier, the experience of a previous event can be archived and applied to the prediction of later similar events. Experimental results show that the Weighted-RF trained by samples of Ferguson event can effectively predict active users in NYC event, with an AUC value around 0.8392. Besides, the image-content based personality model provides a new tool for depicting user portraits, which further contributes to the quantitative analysis of online social events.
Xiaokun Wu 0004, Tianfang Zhao, Weineng Chen, Jun Zhang 0003
IEEE Trans. Knowl. Data Eng.4
2019 Supervised Group Embedding for Rumor Detection in Social Media
Xingming Chen, Yanghui Rao, Haoran Xie 0001, Qing Li 0001, Jun Zhang 0003, Yingchao Zhao 0001, Fu Lee Wang
ICWE6
2018 Multiobjective optimization with ϵ-constrained method for solving real-parameter constrained optimization problems
Jing-Yu Ji, Wei-jie Yu 0001, Yue-Jiao Gong, Jun Zhang 0003
Inf. Sci.4
2018 A hybrid differential evolution algorithm for mixed-variable optimization problems
Ying Lin 0001, Weineng Chen, Jun Zhang 0003
Inf. Sci.4
2018 A tri-objective differential evolution approach for multimodal optimization
Wei-jie Yu 0001, Jing-Yu Ji, Yue-Jiao Gong, Qiang Yang 0008, Jun Zhang 0003
Inf. Sci.5
2018 Semi-Supervised Ensemble Clustering Based on Selected Constraint Projection
abstract
Traditional cluster ensemble approaches have several limitations. (1) Few make use of prior knowledge provided by experts. (2) It is difficult to achieve good performance in high-dimensional datasets. (3) All of the weight values of the ensemble members are equal, which ignores different contributions from different ensemble members. (4) Not all pairwise constraints contribute to the final result. In the face of this situation, we propose double weighting semi-supervised ensemble clustering based on selected constraint projection(DCECP) which applies constraint weighting and ensemble member weighting to address these limitations. Specifically, DCECP first adopts the random subspace technique in combination with the constraint projection procedure to handle high-dimensional datasets. Second, it treats prior knowledge of experts as pairwise constraints, and assigns different subsets of pairwise constraints to different ensemble members. An adaptive ensemble member weighting process is designed to associate different weight values with different ensemble members. Third, the weighted normalized cut algorithm is adopted to summarize clustering solutions and generate the final result. Finally, nonparametric statistical tests are used to compare multiple algorithms on real-world datasets. Our experiments on 15 high-dimensional datasets show that DCECP performs better than most clustering algorithms.
Zhiwen Yu 0002, Peinan Luo, Jiming Liu 0001, Hau-San Wong, Jane You, Guoqiang Han 0002, Jun Zhang 0003
IEEE Trans. Knowl. Data Eng.7
2017 Cooperation coevolution with fast interdependency identification for large scale optimization
Xiaomin Hu, Fei-Long He, Weineng Chen, Jun Zhang 0003
Inf. Sci.4
2017 Adaptive Ensembling of Semi-Supervised Clustering Solutions
abstract
Conventional semi-supervised clustering approaches have several shortcomings, such as (1) not fully utilizing all useful must-link and cannot-link constraints, (2) not considering how to deal with high dimensional data with noise, and (3) not fully addressing the need to use an adaptive process to further improve the performance of the algorithm. In this paper, we first propose the transitive closure based constraint propagation approach, which makes use of the transitive closure operator and the affinity propagation to address the first limitation. Then, the random subspace based semi-supervised clustering ensemble framework with a set of proposed confidence factors is designed to address the second limitation and provide more stable, robust, and accurate results. Next, the adaptive semi-supervised clustering ensemble framework is proposed to address the third limitation, which adopts a newly designed adaptive process to search for the optimal subspace set. Finally, we adopt a set of nonparametric tests to compare different semi-supervised clustering ensemble approaches over multiple datasets. The experimental results on 20 real high dimensional cancer datasets with noisy genes and 10 datasets from UCI datasets and KEEL datasets show that (1) The proposed approaches work well on most of the real-world datasets. (2) It outperforms other state-of-the-art approaches on 12 out of 20 cancer datasets, and 8 out of 10 UCI machine learning datasets.
Zhiwen Yu 0002, Zongqiang Kuang, Jiming Liu 0001, Jun Zhang 0003, Jane You, Hau-San Wong, Guoqiang Han 0002
IEEE Trans. Knowl. Data Eng.5
2016 Adaptive noise immune cluster ensemble using affinity propagation
abstract
Cluster ensemble, as one of the important research directions in the ensemble learning area, is gaining more and more attention, due to its powerful capability to integrate multiple clustering solutions and provide a more accurate, stable and robust result. Cluster ensemble has a lot of useful applications in a large number of areas. Although most of traditional cluster ensemble approaches obtain good results, few of them consider how to achieve good performance for noisy datasets. Some noisy datasets have a number of noisy attributes which may degrade the performance of conventional cluster ensemble approaches. Some noisy datasets which contain noisy samples will affect the final results. Other noisy datasets may be sensitive to distance functions.
Zhiwen Yu 0002, Guoqiang Han 0002, Le Li 0002, Jiming Liu 0001, Jun Zhang 0003
ICDE5
2016 Incremental semi-supervised clustering ensemble for high dimensional data clustering
abstract
Recently, cluster ensemble approaches have gained more and more attention [1]–[2], due to useful applications in the areas of pattern recognition, data mining, bioinformatics, and so on. When compared with traditional single clustering algorithms, cluster ensemble approaches are able to integrate multiple clustering solutions obtained from different data sources into a unified solution, and provide a more robust, stable and accurate final result.
Zhiwen Yu 0002, Peinan Luo, Si Wu 0002, Guoqiang Han 0002, Jane You, Hareton K. N. Leung, Hau-San Wong, Jun Zhang 0003
ICDE8
2016 Incremental Semi-Supervised Clustering Ensemble for High Dimensional Data Clustering
abstract
Traditional cluster ensemble approaches have three limitations: (1) They do not make use of prior knowledge of the datasets given by experts. (2) Most of the conventional cluster ensemble methods cannot obtain satisfactory results when handling high dimensional data. (3) All the ensemble members are considered, even the ones without positive contributions. In order to address the limitations of conventional cluster ensemble approaches, we first propose an incremental semi-supervised clustering ensemble framework (ISSCE) which makes use of the advantage of the random subspace technique, the constraint propagation approach, the proposed incremental ensemble member selection process, and the normalized cut algorithm to perform high dimensional data clustering. The random subspace technique is effective for handling high dimensional data, while the constraint propagation approach is useful for incorporating prior knowledge. The incremental ensemble member selection process is newly designed to judiciously remove redundant ensemble members based on a newly proposed local cost function and a global cost function, and the normalized cut algorithm is adopted to serve as the consensus function for providing more stable, robust, and accurate results. Then, a measure is proposed to quantify the similarity between two sets of attributes, and is used for computing the local cost function in ISSCE. Next, we analyze the time complexity of ISSCE theoretically. Finally, a set of nonparametric tests are adopted to compare multiple semisupervised clustering ensemble approaches over different datasets. The experiments on 18 real-world datasets, which include six UCI datasets and 12 cancer gene expression profiles, confirm that ISSCE works well on datasets with very high dimensionality, and outperforms the state-of-the-art semi-supervised clustering ensemble approaches.
Zhiwen Yu 0002, Peinan Luo, Jane You, Hau-San Wong, Hareton K. N. Leung, Si Wu 0002, Jun Zhang 0003, Guoqiang Han 0002
IEEE Trans. Knowl. Data Eng.7
2015 Competitive and cooperative particle swarm optimization with information sharing mechanism for global optimization problems
Yuhua Li 0002, Zhi-hui Zhan, Shujin Lin, Jun Zhang 0003
Inf. Sci.4
2015 Adaptive Noise Immune Cluster Ensemble Using Affinity Propagation
abstract
Cluster ensemble is one of the main branches in the ensemble learning area which is an important research focus in recent years. The objective of cluster ensemble is to combine multiple clustering solutions in a suitable way to improve the quality of the clustering result. In this paper, we design a new noise immune cluster ensemble framework named as AP2CE to tackle the challenges raised by noisy datasets. AP2CE not only takes advantage of the affinity propagation algorithm (AP) and the normalized cut algorithm (Ncut), but also possesses the characteristics of cluster ensemble. Compared with traditional cluster ensemble approaches, AP2CE is characterized by several properties. (1) It adopts multiple distance functions instead of a single Euclidean distance function to avoid the noise related to the distance function. (2) AP2CE applies AP to prune noisy attributes and generate a set of new datasets in the subspaces consists of representative attributes obtained by AP. (3) It avoids the explicit specification of the number of clusters. (4) AP2CE adopts the normalized cut algorithm as the consensus function to partition the consensus matrix and obtain the final result. In order to improve the performance of AP2CE, the adaptive AP2CE is designed, which makes use of an adaptive process to optimize a newly designed objective function. The experiments on both synthetic and real datasets show that (1) AP2CE works well on most of the datasets, in particular the noisy datasets; (2) AP2CE is a better choice for most of the datasets when compared with other cluster ensemble approaches; (3) AP2CE has the capability to provide more accurate, stable and robust results.
Zhiwen Yu 0002, Le Li 0002, Jiming Liu 0001, Jun Zhang 0003, Guoqiang Han 0002
IEEE Trans. Knowl. Data Eng.4
2014 Special issue on big data research in China
Nanning Zheng 0001, Jun Zhang 0003, Chenghong Wang
Knowl. Inf. Syst.2
2013 Dirichlet Process Mixture Model for Document Clustering with Feature Partition
abstract
Finding the appropriate number of clusters to which documents should be partitioned is crucial in document clustering. In this paper, we propose a novel approach, namely DPMFP, to discover the latent cluster structure based on the DPM model without requiring the number of clusters as input. Document features are automatically partitioned into two groups, in particular, discriminative words and nondiscriminative words, and contribute differently to document clustering. A variational inference algorithm is investigated to infer the document collection structure as well as the partition of document words at the same time. Our experiments indicate that our proposed approach performs well on the synthetic data set as well as real data sets. The comparison between our approach and state-of-the-art document clustering approaches shows that our approach is robust and effective for document clustering.
Rui-zhang Huang, Guan Yu, Zhaojun Wang, Jun Zhang 0003, Liangxing Shi
IEEE Trans. Knowl. Data Eng.4
2008 Chaotic Time Series Prediction Using a Neuro-Fuzzy System with Time-Delay Coordinates
abstract
This paper presents an investigation into the use of the time delay coordinate embedding technique in the multi-input-multi-output-adaptive-network-based fuzzy inference system (MANFIS) for chaotic time series prediction. The inputs of the MANFIS are embedded-phase-space (EPS) vectors preprocessed from the time series under test while the output time series is extracted from the EPS vectors. With such EPS preprocessing, the prediction accuracy of the MANFIS is found to be significantly improved. The proposed system will be tested with a periodic and the Mackey-Glass chaotic time series by comparing the prediction accuracy with and without EPS preprocessing. A moving root-mean-square error is used to monitor the error along the prediction horizon and to tune the membership functions in the MANFIS.
Jun Zhang 0003, Henry S. H. Chung
IEEE Trans. Knowl. Data Eng.1