VLDB 2026 Research / reviewers in the wild / expert
Chengzhen Zhang
dblp:219/7201
· DBLP profile ↗
4ranked-venue papers
0as first author
4since 2021 · last 2026
0009-0001-3272-0825ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MIReNB: misclassified instances retraining-based Naive Bayes algorithmsabstractNaive Bayes is renowned for its simplicity and efficiency, occupying a significant position within the domains of data mining and machine learning. However, its performance optimisation is constrained by the adequacy of instance training and the effectiveness of feature selection. To address these limitations, this paper first proposes a novel algorithm called Misclassified Instances Re-training Naive Bayes (MIReNB). This methodology employs Leave-One-Out Cross Validation (LOOCV) to meticulously discern and iteratively reuse misclassified instances, dynamically adjusting the frequency table of the conditional probability distribution for each attribute-value and class, followed by retraining procedures to augment predictive performance. Building upon MIReNB, we subsequently propose two extensions. The first extension, designated as MIReNB Select, incorporates an additional feature selection stage before identifying misclassified instances to determine the optimal feature subset for retraining, thereby improving prediction accuracy and robustness. The second extension, termed MIReNB Twice, involves a dual iterative process that deeply integrates misclassified instances, achieving adaptive refinement of the conditional probability distribution while enhancing classification accuracy and generalisation ability. Numerous empirical investigations have shown that MIReNB, along with its two extensions, can enhance classification accuracy notably while preserving the inherent simplicity and efficiency of the Naive Bayes classifier. Huihang Ke, Shenglei Chen, Chengzhen Zhang, Qingwu Gao |
J. Exp. Theor. Artif. Intell. | 3 |
| 2022 | Identifying a Minimum Sequence of High-Level Changes Between WorkflowsabstractAdaptive workflow management systems allow workflows to be changed in both the modeling and runtime stages, resulting in many workflow variants. Identifying a minimum sequence of high-level changes between two workflows represents a fundamental yet critical issue. The state-of-the-art approach utilizes digital logic to seek the optimal solution; however, this approach may face difficulties when advanced workflow patterns (e.g., loops) are involved, and it does not scale well. To address this problem, we first propose a naive approach that applies all valid changes to one workflow until the other workflow is found. Then, the approach is optimized from two aspects. First, we present advanced heuristics that significantly reduce the search space without pruning the optimal solution. Second, we employ the A$^\ast$search algorithm to direct the search procedure. Because the heuristic function used in the A$^\ast$algorithm is problem specific, we devise a consistent heuristic function to approximate the edit distance between two workflows, thereby accelerating the search. We implement our approach in a prototype tool and conduct extensive experiments on two data sets to evaluate its effectiveness and efficiency. The experimental results demonstrate that our approach outperforms the state of the art in terms of both application scope and scalability. Wei Song 0003, Fangfei Chen, Hans-Arno Jacobsen, Chengzhen Zhang |
IEEE Trans. Serv. Comput. | 4 |
| 2021 | An Empirical Study on Data Flow Bugs in Business ProcessesabstractAn increasing number of service-based business processes are being developed with the booming of BPaaS (Business Process as a Service) in cloud computing. The profits and performance of enterprises strongly depend on the soundness of their processes being bereft of control flow and data flow bugs. Although some work has focused on the detection of control flow bugs, few studies have comprehensively and empirically investigated data flow bugs in business processes. To this end, we report an empirical study on data flow bugs in business (BPEL) processes. Our analysis of 178 real-world BPEL processes reveals that data flow bugs are surprisingly common: 94 BPEL processes involve data flow bugs, among which redundant output is predominant. The distribution and common scenarios of data flow bugs provide a reference for BPEL process designers. We also investigate the correlation between process complexity metrics and data flow bugs. Based on the statistics of the process complexity metrics and data flow bugs in our empirical study, we present a method to select appropriate metrics as features of BPEL processes and utilize state-of-the-art supervised learning algorithms to predict data flow bugs in an unseen BPEL process. The prediction accuracies of the different classification algorithms exceed 90 percent on average when using our selected metrics. Wei Song 0003, Chengzhen Zhang, Hans-Arno Jacobsen |
IEEE Trans. Cloud Comput. | 2 |
| 2021 | Dependence-Based Data-Aware Process Conformance CheckingabstractData-aware executable processes are an effective and efficient means to build service-oriented applications. However, since the services involved are loosely-coupled and self-managed, the process is flexible by nature and it executions may deviate from their specifications. In contrast to existing approaches that focus on control flow deviations, we leverage activity dependences for data-aware process conformance checking. To analyze the conformance of a process instance to its process definition, we seek a process reference trace “best-fitting” the instance trace such that the conformance degree of the input trace to the process equals the consistency degree of both traces. We measure the consistency between two traces based on their activity dependences. Since finding the reference trace is NP-hard, we resort to heuristics based on process decomposition and trace replaying to determine the trace. Our approach can identify conformance decrease caused by activity dependence deviations, thus, complementing existing approaches. We implement our approach as a ProM plugin. Experimental results on 102 real-world WS-BPEL processes and 26,880 synthetic input traces confirm the effectiveness and efficiency of our approach. Wei Song 0003, Hans-Arno Jacobsen, Chengzhen Zhang, Xiaoxing Ma |
IEEE Trans. Serv. Comput. | 3 |