Zhengyuan Wei

dblp:41/5172 · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 9 · 4 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 HiCert: Toward Patch Robustness Certification and Detection for Deep Learning Systems Beyond Consistent Samples
abstract
Patch robustness certification is an emerging kind of provable defense technique against adversarial patch attacks for deep learning systems. Certified detection ensures the detection of all patched harmful versions of certified samples, which mitigates the failures of empirical defense techniques that could (easily) be compromised. However, existing certified detection methods are ineffective in certifying samples that are misclassified or whose mutants are inconsistently predicted to different labels. This paper proposes HiCert, a novel masking-based certified detection technique. By focusing on the problem of mutants predicted with a label different from the true label with our formal analysis, HiCert formulates a novel formal relation between harmful samples generated by identified loopholes and their benign counterparts. By checking the bound of the maximum confidence among these potentially harmful (i.e., inconsistent) mutants of each benign sample, HiCert ensures that each harmful sample either has the minimum confidence among mutants that are predicted the same as the harmful sample itself below this bound, or has at least one mutant predicted with a label different from the harmful sample itself, formulated after two novel insights. As such, HiCert systematically certifies those inconsistent samples and consistent samples to a large extent. To our knowledge, HiCert is thefirstwork capable of providing such a comprehensive patch robustness certification for certified detection. Our experiments show the high effectiveness of HiCert with a new state-of-the-art performance: It certifies significantly more benign samples, including those inconsistent and consistent, and achieves significantly higher accuracy on those samples without warnings and a significantly lower false silent ratio. Moreover, on actual patch attacks, its defense success ratio is significantly higher than its peers.
Qilin Zhou, Zhengyuan Wei, Haipeng Wang 0005, Wing Kwong Chan
IEEE Trans. Reliab.2
2025 AugOracle: In-Capability Raw Input Validation for Deep Learning Models in Deployment
abstract
Software engineering community has developed effective techniques to detect the failures of software systems with deep learning model components from out-of-distribution inputs. However, failures from in-distribution inputs are underexplored. In this paper, we formulate the problem of in-capability raw input validation regarding the in-distribution failures and propose AugOracle, an effective yet efficient input validation technique, to address this problem. AugOracle differs from previous works in two key aspects: (1) efficiency - it provides robustness requirement rules for detection rather than heavy training-based methods; (2) scalability - it can scale to larger deep learning models in constant overhead. We evaluate AugOracle on 27 configurations to show its effectiveness and efficiency and open-source the implementation.
Zhengyuan Wei
COMPSAC1
2025 Scalable and Precise Patch Robustness Certification for Deep Learning Models with Top- $k$ Predictions
abstract
Patch robustness certification is an emerging verification approach for defending against adversarial patch attacks with provable guarantees for deep learning systems. Certified recovery techniques guarantee the prediction of the sole true label of a certified sample. However, existing techniques, if applicable to top-$k$predictions, commonly conduct pairwise comparisons on those votes between labels, failing to certify the sole true label within the top$k$prediction labels precisely due to the inflation on the number of votes controlled by the attacker (i.e., attack budget); yet enumerating all combinations of vote allocation suffers from the combinatorial explosion problem. We propose CostCert, a novel, scalable, and precise voting-based certified recovery defender. CostCert verifies the true label of a sample within the top$k$predictions without pairwise comparisons and combinatorial explosion through a novel design: whether the attack budget on the sample is infeasible to cover the smallest total additional votes on top of the votes uncontrollable by the attacker to exclude the true labels from the top$k$prediction labels. Experiments show that CostCert significantly outperforms the current state-of-the-art defender PatchGuard, such as retaining up to 57.3% in certified accuracy when the patch size is 96, whereas PatchGuard has already dropped to zero.
Qilin Zhou, Haipeng Wang 0005, Zhengyuan Wei, Wing Kwong Chan
QRS3
2025 Context-Aware Fuzzing for Robustness Enhancement of Deep Learning Models
abstract
In the testing-retraining pipeline for enhancing the robustness property of deep learning (DL) models, many state-of-the-art robustness-oriented fuzzing techniques are metric-oriented. The pipeline generates adversarial examples as test cases via such a DL testing technique and retrains the DL model under test with test suites that contain these test cases. On the one hand, the strategies of these fuzzing techniques tightly integrate the key characteristics of their testing metrics. On the other hand, they are often unaware of whether their generated test cases are different from the samples surrounding these test cases and whether there are relevant test cases of other seeds when generating the current one. We propose a novel testing metric called Contextual Confidence (CC). CC measures a test case through the surrounding samples of a test case in terms of their mean probability predicted to the prediction label of the test case. Based on this metric, we further propose a novel fuzzing technique Clover as a DL testing technique for the pipeline. In each fuzzing round, Clover first finds a set of seeds whose labels are the same as the label of the seed under fuzzing. At the same time, it locates the corresponding test case that achieves the highest CC values among the existing test cases of each seed in this set of seeds and shares the same prediction label as the existing test case of the seed under fuzzing that achieves the highest CC value. Clover computes the piece of difference between each such pair of a seed and a test case. It incrementally applies these pieces of differences to perturb the current test case of the seed under fuzzing that achieves the highest CC value and to perturb the resulting samples along the gradient to generate new test cases for the seed under fuzzing. Clover finally selects test cases among the generated test cases of all seeds as much as possible and with a preference to select test cases with higher CC values for improving model robustness. The experiments show that Clover outperforms the state-of-the-art coverage-based technique Adapt and loss-based fuzzing technique RobOT by 67%–129% and 48%–100% in terms of robustness improvement ratio, respectively, delivered through the same testing-retraining pipeline. For test case generation, in terms of numbers of unique adversarial labels and unique categories for the constructed test suites, Clover outperforms Adapt by \(2.0\times\) and \(3.5\times\) and RobOT by \(1.6\times\) and \(1.7\times\) on fuzzing clean models, and also outperforms Adapt by \(3.4\times\) and \(4.5\times\) and RobOT by \(9.8\times\) and \(11.0\times\) on fuzzing adversarially trained models, respectively.
Haipeng Wang 0005, Zhengyuan Wei, Qilin Zhou, Wing Kwong Chan
ACM Trans. Softw. Eng. Methodol.2
2024 Toward AI-facilitated Learning Cycle in Integration Course Through Pair Programming with AI Agents
abstract
We propose a new methodology that harnesses recent advancements in AI techniques to formulate an AI-facilitating code learning cycle for students. The approach builds on an existing learning process and innovatively incorporates pair programming into the learning cycle. It first transforms the example code into scaffold code as exercises through an instructor-AI pairing. The scaffold code serves as an exercise for students to complete and debug on a hardware platform iteratively with an expert AI assistant. This design alleviates instructors' burden of crafting new exercises for new scenarios and offers students the advantage of interactive learning with scenario diversity. We evaluate the methodology using a suite of example codes and assess the semantic similarity among different code versions produced by AI assistants. The case study shows promising results of the methodology. We further discuss our findings and outline future work for the proposed methodology.
Zhengyuan Wei, Albert T. L. Lee, Victor C. S. Lee, Wing Kwong Chan
CSEE&T1
2023 A Majority Invariant Approach to Patch Robustness Certification for Deep Learning Models
abstract
Patch robustness certification ensures no patch within a given bound on a sample can manipulate a deep learning model to predict a different label. However, existing techniques cannot certify samples that cannot meet their strict bars at the classifier level or the patch region level. This paper proposes MajorCert. MajorCert firstly finds all possible label sets manipulatable by the same patch region on the same sample across the underlying classifiers, then enumerates their combinations element-wise, and finally checks whether the majority invariant of all these combinations is intact to certify samples.
Qilin Zhou, Zhengyuan Wei, Haipeng Wang 0005, Wing Kwong Chan
ASE2
2023 Aster: Encoding Data Augmentation Relations into Seed Test Suites for Robustness Assessment and Fuzzing of Data-Augmented Deep Learning Models
abstract
Data-augmented deep learning models are widely used in real-world applications. However, many state-of the-art loss-based or coverage-based fuzzing techniques fail to produce fuzzing samples for them from many seeds. This paper proposes Aster, a novel technique to address this problem to enhance their fuzzing effectiveness for deep learning models trained with multi-sample data augmentation methods. Aster formulates a novel reachability-based strategy to encode the insights of every seed’s direct and indirect data augmentation relation instances into the replacement seed of that seed systematically. Our experiment shows that Aster is highly effective. On average, loss-based and coverage-based fuzzing techniques can generate 166% and 110% more fuzzing samples and reduce 31% and 22% unsuccessful seeds, respectively, after adopting the replacement seeds generated by Aster to replace their original seeds. Their improved models also become up to 55% and 40% on average more robust against FGSM and PGD attacks in the experiment.
Haipeng Wang 0005, Zhengyuan Wei, Qilin Zhou, Bo Jiang 0001, Wing Kwong Chan
QRS2
2023 DeepPatch: Maintaining Deep Learning Model Programs to Retain Standard Accuracy with Substantial Robustness Improvement
abstract
Maintaining a deep learning (DL) model by making the model substantially more robust through retraining with plenty of adversarial examples of non-trivial perturbation strength often reduces the model’s standard accuracy. Many existing model repair or maintenance techniques sacrifice standard accuracy to produce a large gain in robustness or vice versa. This article proposes DeepPatch, a novel technique to maintain filter-intensive DL models. To the best of our knowledge, DeepPatch is the first work to address the challenge of standard accuracy retention while substantially improving the robustness of DL models with plenty of adversarial examples of non-trivial and diverse perturbation strengths. Rather than following the conventional wisdom to generalize all the components of a DL model over the union set of clean and adversarial samples, DeepPatch formulates a novel division of labor method to adaptively activate a subset of its inserted processing units to process individual samples. Its produced model can generate the original or replacement feature maps in each forward pass of the patched model, making the patched model carry an intrinsic property of behaving like the model under maintenance on demand. The overall experimental results show that DeepPatch successfully retains the standard accuracy of all pretrained models while improving the robustness accuracy substantially. However, the models produced by the peer techniques suffer from either large standard accuracy loss or small robustness improvement compared with the models under maintenance, rendering them unsuitable in general to replace the latter.
Zhengyuan Wei, Haipeng Wang 0005, Imran Ashraf 0001, Wing Kwong Chan
ACM Trans. Softw. Eng. Methodol.1
2022 Predictive Mutation Analysis of Test Case Prioritization for Deep Neural Networks
abstract
Testing deep neural networks requires high-quality test cases, but using new test cases would incur the labor-intensive test case labeling issue in the test oracle problem. Test case prioritization for failure-revealing test cases alleviates the problem. Existing metric-based techniques analyze vector-based prediction outputs. They cannot handle regression models. Existing mutation-based techniques either remain ineffective or incur high computational costs. In this paper, we propose EffiMAP, an effective and efficient test case prioritization technique with predictive mutation analysis. In the test phase, without performing a comprehensive mutation analysis, EffiMAP predicts whether model mutants are killed by a test case by the information extracted from the execution trace of the test case. Our experiment shows that EffiMAP significantly outperforms the previous state-of-the-art technique in both effectiveness and efficiency in the test phase of handling test cases of both classification and regression models. This paper is the first work to show the feasibility of predictive mutation analysis to rank test cases with a higher probability of exposing model prediction failures in the domain of deep neural network testing.
Zhengyuan Wei, Haipeng Wang 0005, Imran Ashraf 0001, Wing Kwong Chan
QRS1
2021 A Multi-Modal Transformer-based Code Summarization Approach for Smart Contracts
abstract
Code comment has been an important part of computer programs, greatly facilitating the understanding and maintenance of source code. However, high-quality code comments are often unavailable in smart contracts, the increasingly popular programs that run on the blockchain. In this paper, we propose a Multi-Modal Transformer-based (MMTrans) code summarization approach for smart contracts. Specifically, the MMTrans learns the representation of source code from the two heterogeneous modalities of the Abstract Syntax Tree (AST), i.e., Structure-based Traversal (SBT) sequences and graphs. The SBT sequence provides the global semantic information of AST, while the graph convolution focuses on the local details. The MMTrans uses two encoders to extract both global and local semantic information from the two modalities respectively, and then uses a joint decoder to generate code comments. Both the encoders and the decoder employ the multi-head attention structure of the Transformer to enhance the ability to capture the long-range dependencies between code tokens. We build a dataset with over 300Kpairs of smart contracts, and evaluate the MMTrans on it. The experimental results demonstrate that the MMTrans outperforms the state-of-the-art baselines in terms of four evaluation metrics by a substantial margin, and can generate higher quality comments.
Zhen Yang 0022, Jacky W. Keung, Xiao Yu 0008, Xiaodong Gu 0002, Zhengyuan Wei, Miao Zhang 0025
ICPC5
2021 Fuzzing Deep Learning Models against Natural Robustness with Filter Coverage‡
abstract
Coverage-guided fuzzing on deep learning (DL) models can generate natural adversarial variants. However, no existing fuzzing work can show that covering or uncovering a coverage element of a test adequacy criterion can significantly change the accuracy of the DL model under test. This paper proposes a novel testing criterion, Filter Coverage, and a novel fuzzing technique FilterFuzz guided by this criterion. To the best of our knowledge, Filter Coverage is the first test adequacy criterion able to identify such a coverage element-blamed filter-in a DL model using convolutional filters by demonstrating a cause-and-effect chain. Both Filter Coverage and Neuron Coverage measure whether individual neurons are activated, but Filter Coverage is more selective and collective and higher in the abstraction level. FilterFuzz, guided by Filter Coverage, tackles the challenge of achieving a higher rate of generating natural adversarial variants against natural robustness. Our case study shows that FilterFuzz is significantly more effective than fuzzing guided by Neuron Coverage by 33% and produces more diverse kinds of such variants. Moreover, when applying to the problem of labeling cost reduction on generated natural variants, Filter Coverage identifies 4.3x to 4.7x of natural adversarial variants than random reordering. Our work also calls for a re-examination of the previous ineffective conclusions of Neuron Coverage.
Zhengyuan Wei, Wing Kwong Chan
QRS1