VLDB 2026 Research / reviewers in the wild / expert
Miao Zhang 0025
dblp:60/7041-25
· DBLP profile ↗
15ranked-venue papers
4as first author
10since 2021 · last 2025
0000-0001-9659-9393ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 15 · 4 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | LLM-Based Semantic Modeling and Cooperative Evolutionary Fuzzing for Traffic Violation Scenario GenerationabstractEnsuring the safety of autonomous driving systems (ADS) in a cost-effective and efficient manner remains a critical challenge. Existing law-guided scenario generation approaches are typically limited to a narrow subset of legal rules, resulting in insufficient scenario diversity, and search-based methods often struggle with large and sparse search spaces. To address these limitations, we propose SLaFE (Semantic Law Modeling and Fuzzing based on Cooperative Evolution), a novel traffic violation scenario generation framework designed to systematically evaluate the safety of ADS. SLaFE harnesses the reasoning capabilities of large language models (LLMs) to convert traffic laws into structured scenario constraints. These scenarios are then optimized via a cooperative evolutionary fuzzing algorithm that explores the parameter space to identify boundary cases likely to trigger abnormal ADS behaviors. We evaluate SLaFE on the Apollo platform within the LGSVL simulator using ten realworld traffic regulations. Experimental results show that SLaFE successfully triggered all 10 types of traffic law violations (10/10), outperforming the best existing method, VioHawk (9/10), while others detected no more than 3. Moreover, SLaFE achieved an average triggering time of 5.1 minutes per law type, significantly faster than VioHawk (9.0 minutes) and other baselines. These results highlight SLaFE’s effectiveness in discovering diverse and critical law-violating scenarios for ADS testing. Yan Xiao 0002, Miao Zhang 0025, Pengcheng Zhang 0001 |
APSEC | 4 |
| 2025 | An Empirical Study of Reinforcement Learning-based Class Integration Test Order GenerationabstractComplicated class dependencies in object-oriented systems challenge traditional integration testing methods. Given varying efforts required to construct test stubs, determining optimal class test orders to minimize stubbing complexity is critical in integration testing. Reinforcement learning (RL), with its strengths in solving complex optimizations, has attracted substantial research attention. To fill the gap in the existing research that lacks of cross-strategy comparisons of RL, we systematically compare five RL algorithms’ performance using varied stubbing complexity weighting methods and critical class identification. Experimental results indicate that when fixed weights are used to calculate stubbing complexity, D3QN achieved the minimum overall stubbing complexity in seven out of nine tested programs. However, PPO attained the lowest mean overall stubbing complexity in five programs, demonstrating the best overall performance. DQN exhibited improved performance when stub complexity was calculated using the entropy-weighted method. After incorporating importance scores as a factor, D3QN achieved the best performance. The RL reward function with introduced importance values is more suitable for systems that are highly centralized and rely on core classes. This work provides robust empirical support and selection guidelines for the practical application of RL algorithms in CITO generation, as well as targeted insights for optimizing RL applications in this field. Shuxiang Zheng, Miao Zhang 0025, Yan Xiao 0002, Peihong Chen, Xiaoxing Yang |
APSEC | 3 |
| 2022 | The impact of the distance metric and measure on SMOTE-based techniques in software defect prediction
Shuo Feng 0003, Jacky W. Keung, Peichang Zhang, Yan Xiao 0002, Miao Zhang 0025 |
Inf. Softw. Technol. | 5 |
| 2021 | ROCT: Radius-based Class Overlap Cleaning Technique to Alleviate the Class Overlap Problem in Software Defect PredictionabstractThe training data commonly used in software defect prediction (SDP) usually contains some instances that have similar values on features but are in different classes, which significantly degrades the performance of prediction models trained using these instances. This is referred to as the class overlap problem (COP). Previous studies have concluded that COP has a more negative impact on the performance of prediction models than the class imbalance problem (CIP). However, less research has been conducted on COP than CIP. Moreover, the performance of the existing class overlap cleaning techniques heavily relies on the settings of hyperparameters such as the value of K in the K-nearest neighbor algorithm or the K-means algorithm, but how to find those optimal hyperparameters is still a challenge. In this study, we propose a novel technique named the radius-based class overlap cleaning technique (ROCT) to better alleviate COP without tuning hyperparameters in SDP. The basic idea of ROCT is to take each instance as the center of a hypersphere and directly optimize the radius of the hypersphere. Then ROCT identifies those instances with the opposite label of the center instance as the overlapping instance and removes them. To investigate the performance of ROCT, we conduct the empirical experiment across 29 datasets collected from various software repositories on the K-nearest neighbor, random forest, logistic regression, and naive Bayes classifiers measured by AUC, balance, pd, and pf. The experimental results show that ROCT performs the best and significantly improves the performance of prediction models by as much as 15.2% and 29.9% in terms of AUC and balance compared with the existing class overlap cleaning techniques. The superior performance of ROCT indicates that ROCT should be recommended as an efficient alternative to alleviate COP in SDP. Shuo Feng 0003, Jacky W. Keung, Jie Liu 0016, Yan Xiao 0002, Xiao Yu 0008, Miao Zhang 0025 |
COMPSAC | 6 |
| 2021 | A Multi-Modal Transformer-based Code Summarization Approach for Smart ContractsabstractCode comment has been an important part of computer programs, greatly facilitating the understanding and maintenance of source code. However, high-quality code comments are often unavailable in smart contracts, the increasingly popular programs that run on the blockchain. In this paper, we propose a Multi-Modal Transformer-based (MMTrans) code summarization approach for smart contracts. Specifically, the MMTrans learns the representation of source code from the two heterogeneous modalities of the Abstract Syntax Tree (AST), i.e., Structure-based Traversal (SBT) sequences and graphs. The SBT sequence provides the global semantic information of AST, while the graph convolution focuses on the local details. The MMTrans uses two encoders to extract both global and local semantic information from the two modalities respectively, and then uses a joint decoder to generate code comments. Both the encoders and the decoder employ the multi-head attention structure of the Transformer to enhance the ability to capture the long-range dependencies between code tokens. We build a dataset with over 300Kpairs of smart contracts, and evaluate the MMTrans on it. The experimental results demonstrate that the MMTrans outperforms the state-of-the-art baselines in terms of four evaluation metrics by a substantial margin, and can generate higher quality comments. Zhen Yang 0022, Jacky W. Keung, Xiao Yu 0008, Xiaodong Gu 0002, Zhengyuan Wei, Miao Zhang 0025 |
ICPC | 7 |
| 2021 | COSTE: Complexity-based OverSampling TEchnique to alleviate the class imbalance problem in software defect prediction
Shuo Feng 0003, Jacky W. Keung, Xiao Yu 0008, Yan Xiao 0002, Kwabena Ebo Bennin, Md. Alamgir Kabir, Miao Zhang 0025 |
Inf. Softw. Technol. | 7 |
| 2021 | Investigation on the stability of SMOTE-based oversampling techniques in software defect prediction
Shuo Feng 0003, Jacky W. Keung, Xiao Yu 0008, Yan Xiao 0002, Miao Zhang 0025 |
Inf. Softw. Technol. | 5 |
| 2021 | Validating class integration test order generation systems with Metamorphic Testing
Miao Zhang 0025, Jacky W. Keung, Tsong Yueh Chen, Yan Xiao 0002 |
Inf. Softw. Technol. | 1 |
| 2021 | Evaluating the effects of similar-class combination on class integration test order generation
Miao Zhang 0025, Jacky W. Keung, Yan Xiao 0002, Md. Alamgir Kabir |
Inf. Softw. Technol. | 1 |
| 2021 | An Integration Test Order Strategy to Consider Control CouplingabstractIntegration testing is a very important step in software testing. Existing methods evaluate the stubbing cost for class integration test orders by considering only the interclass direct relationships such as inheritance, aggregation, and association, but they omit the interclass indirect relationship caused by control coupling, which can also affect the test orders and the stubbing cost. In this paper, we introduce an integration test order strategy to consider control coupling. We advance the concept of transitive relationship to describe this kind of interclass dependency and propose a new measurement method to estimate the complexity of control coupling, which is the complexity of stubs created for a transitive relationship. We evaluate our integration test order strategy on 10 programs on various scales. The results show that considering the transitive relationship when generating class integration test orders can significantly reduce the stubbing cost for most programs and that our integration test order strategy obtains satisfactory results more quickly than other methods. Shujuan Jiang, Miao Zhang 0025, Rongcun Wang, Qiao Yu 0001, Jacky W. Keung |
IEEE Trans. Software Eng. | 2 |
| 2020 | A Drift Propensity Detection Technique to Improve the Performance for Cross-Version Software Defect PredictionabstractIn cross-version defect prediction (CVDP), historical data is derived from the prior version of the same project to predict defects of the current version. Recent studies in CVDP focus on subset selection to deal with the changes of the data distributions. No prior study has focused on training data arriving in streaming fashion across the versions where the significant differences between versions make the prediction unreliable. We refer to this situation as Drift Propensity (DP). By identifying DP, necessary steps can be taken (e.g., updating or retraining the model) to improve the prediction performance. In this paper, we investigate the chronological defect datasets and identify DP in the datasets. The no-memory data management technique is employed to manage the data distributions and a DP detection technique is proposed. The idea behind the proposed DP detection technique is to monitor the algorithm's error-rate. The DP detector triggers DP, warning, and control flags to take necessary steps. The proposed technique is significantly superior in identifying the distribution differences (p-value <; 0.05). The DP's identified in the data distributions achieve large effect sizes (Hedges' g ≥ 0.80) during the pair-wise comparisons. We observe that if the error-rate exponentially increases, it causes DP, resulting in prediction performance deterioration. We thus recommend researches and practitioners to address DP in the chronological datasets. Due to its potential effects in the datasets, the prediction models could be enhanced to get the best results in CVDP. Md. Alamgir Kabir, Jacky W. Keung, Kwabena Ebo Bennin, Miao Zhang 0025 |
COMPSAC | 4 |
| 2020 | Smart Contracts Vulnerability Auditing with Multi-semanticsabstractSmart contracts vulnerability auditing is vitally critical to ensure transaction execution in normal on blockchain. The current data-driven approaches normally tokenize smart contracts into a series of sequences according to only one tokenization standard for vulnerability detection purpose, resulting some of the semantic contexts could not be reflected within restricted sequence length. To address this limitation, we generate sequences from smart contracts in three tokenization standards for which we utilize n-gram language model to capture semantic contexts respectively, and finally exploiting our effective combination strategy of Intersection or Union to integrate the audited results from multiple semantic contexts. In order to evaluate the proposed approach, we applied it on over 7200 Ethereum smart contract samples. Experimental result shows our proposed method is capable of detecting vulnerabilities and competitive with the baseline in test sets, with improved precision of over 44% when Intersection is applied in their results, as well as improved Recall measure up by over 300% and F-measure up by 220% when Union is applied. Our proposed method for smart contract vulnerability detection, an important tool for developing quality decentralized software applications, is able to analyze multiple semantic contexts and successfully detects more true vulnerabilities with high precision, outperforming that of the baseline approaches. Zhen Yang 0022, Jacky W. Keung, Miao Zhang 0025, Yan Xiao 0002, Yangyang Huang, Tik Hui |
COMPSAC | 3 |
| 2019 | Assessing the Significant Impact of Concept Drift in Software Defect PredictionabstractConcept drift is a known phenomenon in software data analytics. It refers to the changes in the data distribution over time. The performance of analytic and prediction models degrades due to the changes in the data over time. To improve prediction performance, most studies propose that the prediction model be updated when concept drift occurs. In this work, we investigate the existence of concept drift and its associated effects on software defect prediction performance. We adopt the strategy of an empirically proven method DDM (Drift Detection Method) and evaluate its statistical significance using the chi-square test with Yates continuity correction. The objective is to empirically determine the concept drift and to calibrate the base model accordingly. The empirical study indicates that the concept drift occurs in software defect datasets, and its existence subsequently degrades the performance of prediction models. Two types of concept drifts (gradual and sudden drifts) were identified using the chi-square test with Yates continuity correction in the software defect datasets studied. We suggest concept drift should be considered by software quality assurance teams when building prediction models. Md. Alamgir Kabir, Jacky W. Keung, Kwabena Ebo Bennin, Miao Zhang 0025 |
COMPSAC (1) | 4 |
| 2019 | A Heuristic Approach to Break Cycles for the Class Integration Test Order GenerationabstractIt is a general objective to minimize overall stubbing cost when performing class integration test order generation. Existing approaches are unable to obtain a cost-optimal class test order, this is largely due to the lack of a comprehensive analysis on the factors that affect overall stubbing cost, i.e., the number of required test stubs and the corresponding stubbing complexity. To address this issue, we propose an approach called HBCITO (Heuristic approach to Break Cycles for the class Integration Test Order generation). Given a set of removed dependencies, a heuristic algorithm is employed to search for a near ideal set of class dependencies. Such dependencies break the same or greater number of cycles as the initialized dependencies but attract less stubbing cost. The experimental results show that HBCITO is capable of generating class test orders with significantly lower stubbing cost compared with other approaches. Miao Zhang 0025, Jacky W. Keung, Yan Xiao 0002, Md. Alamgir Kabir, Shuo Feng 0003 |
COMPSAC (1) | 1 |
| 2017 | A multi-level feedback approach for the class integration and test order problem
Miao Zhang 0025, Shujuan Jiang, Xingya Wang, Qiao Yu 0001 |
J. Syst. Softw. | 1 |