VLDB 2026 Research / reviewers in the wild / expert
Zhaojun Wang
dblp:85/391
· DBLP profile ↗
17ranked-venue papers
4as first author
13since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 2 first-author · 7 since 2021Systems, architecture and hardware · 4 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2Theory of computation · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Masked Genetic Operators with Causal Grouping for Constrained Multi-Objective OptimizationabstractUncovering the direct causal relationships between decision variables and optimization objectives can significantly simplify the complexity of optimization problems. However, most existing constrained multi-objective evolutionary algorithms (CMOEAs) fail to address constrained multi-objective optimization from this perspective. To bridge this gap, this study introduces a novel algorithm, CI-CMOEA (Causal Intervention-based CMOEA), which leverages causal intervention techniques to enhance optimization performance. CI-CMOEA begins by constructing a causal relationship network that captures the interactions between decision variables and optimization objectives. Using this network, a genetic operator with a causal relationship mask is designed to group decision variables based on their causal impact on the objectives. By focusing genetic operations on key variables with significant causal influence, the algorithm effectively guides the evolutionary optimization process towards better solutions. To further improve performance, CI-CMOEA employs a dual-population collaboration mechanism. One population operates under relaxed epsilon constraints to explore the solution space, while the other disregards constraints to enhance convergence. Preliminary experiments on the LIR-CMOP test suite demonstrate that CI-CMOEA not only accurately identifies the causal relationships between decision variables and objectives but also outperforms eight state-of-the-art CMOEAs in terms of IGD, IGD+ and HV metrics, showcasing its superior optimization performance and reliability. Zhaojun Wang, Jiachun Huang, Wenji Li, Shunge Wang, Yifeng Qiu, Jiafan Zhuang, Zhun Fan |
CEC | 1 |
| 2025 | Dynamic Adaptive Fault-Tolerance in Stream Computing Systems Under Resource Constraints
Zhaojun Wang, Dawei Sun 0001, Xuan Zang, Atul Sajjanhar, Rajkumar Buyya |
ICA3PP (2) | 1 |
| 2025 | Conditional Testing based on Localized Conformal p-valuesabstractIn this paper, we address conditional testing problems through the conformal inference framework. We define the localized conformal $p$-values by inverting prediction intervals and prove their theoretical properties. These defined $p$-values are then applied to several conditional testing problems to illustrate their practicality. Firstly, we propose a conditional outlier detection procedure to test for outliers in the conditional distribution with finite-sample false discovery rate (FDR) control. We also introduce a novel conditional label screening problem with the goal of screening multivariate response variables and propose a screening procedure to control the family-wise error rate (FWER). Finally, we consider the two-sample conditional distribution test and define a weighted U-statistic through the aggregation of localized $p$-values. Numerical simulations and real-data examples validate the superior performance of our proposed strategies. Zhaojun Wang, Changliang Zou |
ICLR | 3 |
| 2025 | Conformal Prediction with Cellwise Outliers: A Detect-then-Impute ApproachabstractConformal prediction is a powerful tool for constructing prediction intervals for black-box models, providing a finite sample coverage guarantee for exchangeable data. However, this exchangeability is compromised when some entries of the test feature are contaminated, such as in the case of cellwise outliers. To address this issue, this paper introduces a novel framework called *detect-then-impute conformal prediction*. This framework first employs an outlier detection procedure on the test feature and then utilizes an imputation method to fill in those cells identified as outliers. To quantify the uncertainty in the processed test feature, we adaptively apply the detection and imputation procedures to the calibration set, thereby constructing exchangeable features for the conformal prediction interval of the test label. We develop two practical algorithms, $\texttt{PDI-CP}$ and $\texttt{JDI-CP}$, and provide a distribution-free coverage analysis under some commonly used detection and imputation procedures. Notably, $\texttt{JDI-CP}$ achieves a finite sample $1-2\alpha$ coverage guarantee. Numerical experiments on both synthetic and real datasets demonstrate that our proposed algorithms exhibit robust coverage properties and comparable efficiency to the oracle baseline. Yajie Bao, Haojie Ren, Zhaojun Wang, Changliang Zou |
ICML | 4 |
| 2025 | CE-DCVSI: Multimodal relational extraction based on collaborative enhancement of dual-channel visual semantic information
Yunchao Gong, Xueqiang Lv, Zangtai Cai, Yuzhong Chen 0003, Zhaojun Wang, Xindong You |
Expert Syst. Appl. | 7 |
| 2025 | Time-Varying Target Predictive Entrapment Based on Gene Regulatory Network and Sliding Mode ControlabstractTo address slow convergence and formation maintenance challenges in swarm robotic entrapment tasks, a predictive entrapment control algorithm that combines gene regulatory network and sliding mode control (GRN-SMC) is proposed. First, to stabilize the target position information generated by the hierarchical GRN, a novel sorting rule is designed. Then, an artificial neural network (ANN) is employed to perform on-line prediction of the swarm robots’ kinematic states. These predicted values are then fed into a specifically designed sliding mode controller, which ultimately outputs the optimal control velocities for the swarm robots. Comparative simulation experiments with three state-of-the-art algorithms demonstrate that our method achieves significant improvements in tracking accuracy(error reduced by 82%), single-iteration execution time(reduced by 29%), and formation maintenance (formation integrity increased by 34%). Furthermore, physical robot experiments demonstrate that even in the presence of unknown external disturbances (such as ground slippage) and robot positioning errors (±0.1 m, ±5°), the proposed algorithm still exhibits excellent robustness. Ziling Wen, Zhaojun Wang, Dawei Huang, Binghao Yang, Wenji Li, Zhun Fan, An-Min Zou |
IEEE Internet Things J. | 3 |
| 2025 | Online Multiple Changepoint Detection With False Discovery Rate ControlabstractTechnological advances have led to the emergence of an increasing number of applications requiring the analysis of datastreams, that are characterized by an indefinitely long and time-evolving sequence, particularly in the healthcare domain. In such applications, the status of a stream can alternate, possibly many times, between a regular status and an irregular status. Consequently, it is necessary to develop statistical methodologies that constantly detect multiple changepoints in an online manner. While we may employ conventional methods of sequential change detection to trigger signals after the change occurs, no online procedure is available to quantify the uncertainty of the detected changes. In this work, we fill this gap by framing online multiple changepoint detection into an online multiple testing problem and proposing a new framework to test the null hypothesis that there is no change between neighboring signalled points. To obtain valid p-values for online multiple testing, we propose a data-fission-based procedure that is a simple yet effective way of dealing with the post-detection uncertainty quantification. It is shown that popular online false discovery rate control methods with those p-values can achieve finite-sample false discovery rate control. We evaluate the proposed method in simulation studies. The method is applied to health monitoring dataset, alleviating the false alarm issue in online data analysis. Haoyu Geng, Haojie Ren, Zhaojun Wang, Changliang Zou |
IEEE Trans. Inf. Theory | 4 |
| 2024 | Well Trajectory Design Based on Constrained Many-Objective Optimization AlgorithmsabstractIn the field of drilling engineering, the design and optimization of well trajectories are crucial. This study focuses on optimizing key aspects such as the length of the well trajectory, drill string torque, the energy of the well-profile, and accuracy in reaching the target. This problem encompasses eleven complex nonlinear constraints and four conflicting objectives, presenting challenges for traditional mathematical programming methods. To tackle this problem, we introduce a novel constrained many-objective optimization algorithm, named PPS-NSGA-III. The proposed algorithm partitions the objective space into subspaces, using NSGA-III to find Pareto optimal solutions in each, enhancing diversity. The push-and-pull search framework is employed to overcome local optima in each subproblem, accelerating overall convergence. Through a comparative analysis with some evolutionary algorithms, PPS-NSGA-III has shown superior performance. It delivers more effective design solutions with lower risk, reduced cost, and a higher drilling encounter rate in the proposed well trajectory optimization model. Zhaojun Wang, Chenwen Ding, Wenji Li, Yifeng Qiu, Jiafan Zhuang, Zhun Fan |
CEC | 1 |
| 2024 | Robust Estimation of High-Dimensional Linear Regression With ChangepointsabstractThe identification of changes in linear models is a fundamental problem encountered in various applications. Traditional methods often encounter difficulties when attempting to identify changepoints in the presence of heavy-tailed distribution. This paper focuses on the study of high-dimensional linear models with multiple structural changes in the presence of heavy-tailed errors, especially for those errors without moment conditions. We first propose a robust method that simultaneously estimates regression coefficient and changepoint by incorporating$\ell _{1}$norm penalized Wilcoxon rank loss minimization for single changepoint estimation. Furthermore, we extend it to multiple changepoints estimation based on a novel two-step moving window mechanism that combines fast coarse grid screening and an efficient refinement. Our method exhibits robustness against heavy-tailed random errors while maintaining high efficiency for normal random errors. Theoretically, we establish non-asymptotic error bounds with a near-oracle rate for the estimates of both the coefficient and the changepoint under weak conditions on the random error distribution. Numerical results provide evidence for the validity and effectiveness of the proposed approach. Haoyu Geng, Zhaojun Wang, Changliang Zou |
IEEE Trans. Inf. Theory | 3 |
| 2024 | Multimodal heterogeneous graph entity-level fusion for named entity recognition with multi-granularity visual guidance
Yunchao Gong, Xueqiang Lv, Zhaojun Wang, Xindong You |
J. Supercomput. | 4 |
| 2024 | A relation enhanced model for temporal knowledge graph alignment
Zhaojun Wang, Xindong You, Xueqiang Lv |
J. Supercomput. | 1 |
| 2023 | A Surrogate-Ensemble Assisted Coevolutionary Algorithm for Expensive Constrained Multi-Objective Optimization ProblemsabstractIn real-world applications, there are some constrained multi-objective problems where the evaluation of objectives is expensive and the evaluation of constraints is cheap. Currently, few studies have focused on solving expensive constrained multi-objective optimization problems (ECMOPs), and they usually assume that the constraints of ECMOPs are also expensive. In this paper, we propose a surrogate-ensemble assisted coevolutionary algorithm (SEACoEA) for ECMOPs with inexpensive constraint evaluation. First, a feasible sampling strategy is designed to initialize the population in the feasible regions. Next, two populations are set to optimize the original ECMOP and the problem without considering constraints, respectively. To improve the search efficiency, we redesigned the objective function of the surrogate-ensemble model. Finally, a new infill strategy is proposed to select candidate individuals from each population for real evaluation. Experimental results show that the proposed algorithm performs significantly better on most MW problems compared to several state-of-the-art algorithms. Wenji Li, Ruitao Mai, Pengxiang Ren, Zhaojun Wang, Qinchang Zhang, Zhun Fan |
CEC | 4 |
| 2021 | An Improved Epsilon Method with M2M for Solving Imbalanced CMOPs with Simultaneous Convergence-Hard and Diversity-Hard Constraints
Zhun Fan, Zhi Yang 0007, Yajuan Tang, Wenji Li, Zhaojun Wang, Fuzan Sun, Zhoubin Long, Guijie Zhu |
EMO | 6 |
| 2018 | LSHADE44 with an Improved $\epsilon$ Constraint-Handling Method for Solving Constrained Single-Objective Optimization ProblemsabstractThis paper proposes an improved$\epsilon$constrained handling method (IEpsilon) for solving constrained single-objective optimization problems (CSOPs). The IEpsilon method adaptively adjusts the value of$\epsilon$according to the proportion of feasible solutions in the current population, which has an ability to balance the search between feasible regions and infeasible regions during the evolutionary process. The proposed constrained handling method is embedded to the differential evolutionary algorithm LSHADE44 to solve CSOPs. Furthermore, a new mutation operator DE/randr1*/1 is proposed in the LSHADE44-IEpsilon. In this paper, twenty-eight CSOPs given by “Problem Definitions and Evaluation Criteria for the CEC 2017 Competition on Constrained Real-Parameter Optimization” are tested by the LSHADE44-IEpsilon and four other differential evolution algorithms CAL-SHADE, LSHADE44+IDE, LSHADE44 and UDE. The experimental results show that the LSHADE44-IEpsilon outperforms these compared algorithms, which indicates that the IEpsilon is an effective constraint-handling method to solve the CEC2017 benchmarks. Zhun Fan, Yi Fang 0007, Wenji Li, Yutong Yuan, Zhaojun Wang, Xinchao Bian |
CEC | 5 |
| 2013 | Dirichlet Process Mixture Model for Document Clustering with Feature PartitionabstractFinding the appropriate number of clusters to which documents should be partitioned is crucial in document clustering. In this paper, we propose a novel approach, namely DPMFP, to discover the latent cluster structure based on the DPM model without requiring the number of clusters as input. Document features are automatically partitioned into two groups, in particular, discriminative words and nondiscriminative words, and contribute differently to document clustering. A variational inference algorithm is investigated to infer the document collection structure as well as the partition of document words at the same time. Our experiments indicate that our proposed approach performs well on the synthetic data set as well as real data sets. The comparison between our approach and state-of-the-art document clustering approaches shows that our approach is robust and effective for document clustering. Rui-zhang Huang, Guan Yu, Zhaojun Wang, Jun Zhang 0003, Liangxing Shi |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2010 | DSpam: Defending Against Spam in Tagging Systems via Users' ReliabilityabstractResisting spam in tagging system is very challenging. This paper presents DSpam, a novel spam-resistant tagging system which can significantly diminish spam in tag search results with users’ reliabilities. DSpam client groups other users into unfamiliar users and interacted users according to the fact whether the client has interacted with such users. For an unfamiliar user, the client computes his reliability by tagging behavior-based mechanism which reflects correlation of annotations between them. For an interacted user, the reliability includes two parts: feedback-based reliability, which indicates direct interactions between that user and the client, and recommendation reliability, which indicates the evaluation about that user from the client’s friends. The client ranks search result with the average reliabilities of himself with respect to annotators of each result. Experimental results show DSpam can effectively resist tag spam and work better than existing tag search schemes. Ennan Zhai, Cui Cao, Yongqiang Xie, Zhaojun Wang, Jian-bin Hu, Zhong Chen 0001 |
ICPADS | 5 |
| 2010 | Document clustering via dirichlet process mixture model with feature selectionabstractOne essential issue of document clustering is to estimate the appropriate number of clusters for a document collection to which documents should be partitioned. In this paper, we propose a novel approach, namely DPMFS, to address this issue. The proposed approach is designed 1) to group documents into a set of clusters while the number of document clusters is determined by the Dirichlet process mixture model automatically; 2) to identify the discriminative words and separate them from irrelevant noise words via stochastic search variable selection technique. We explore the performance of our proposed approach on both a synthetic dataset and several realistic document datasets. The comparison between our proposed approach and stage-of-the-art document clustering approaches indicates that our approach is robust and effective for document clustering. Guan Yu, Rui-zhang Huang, Zhaojun Wang |
KDD | 3 |