VLDB 2026 Research / reviewers in the wild / expert
Wei Zheng 0006
dblp:44/4773-6
· DBLP profile ↗
38ranked-venue papers
17as first author
25since 2021 · last 2025
0000-0003-2905-7193ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 28 · 11 first-author · 17 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Code-DiTing: Automatic Evaluation of Code Generation without References or Test CasesabstractTrustworthy evaluation methods for code snippets play a crucial role in neural code generation. Traditional methods, which either rely on reference solutions or require executable test cases, have inherent limitation in flexibility and scalability. The recent LLM-as-Judge methodology offers a promising alternative by directly evaluating functional consistency between the problem description and the generated code. To systematically understand the landscape of these LLM-as-Judge methods, we conduct a comprehensive empirical study across three diverse datasets. Our investigation reveals the pros and cons of two categories of LLM-as-Judge methods: the methods based on general foundation models can achieve good performance but require complex prompts and lack explainability, while the methods based on reasoning foundation models provide better explainability with simpler prompts but demand substantial computational resources due to their large parameter sizes. To address these limitations, we propose Code-DiTing, a novel code evaluation method that balances accuracy, efficiency and explainability. We develop a data distillation framework that effectively transfers reasoning capabilities from DeepSeek-R1-671B to our Code-DiTing 1.5B and 7B models, significantly enhancing evaluation explainability and reducing the computational cost. With the majority vote strategy in the inference process, Code-DiTing 1.5B outperforms all models with the same magnitude of parameters and achieves performance which would normally exhibit in a model with 5 times of parameter scale. Code-DiTing 7B surpasses GPT-4o and DeepSeek-V3 671B, even though it only uses 1% of the parameter volume of these large models. Further experiments show that Code-DiTing is robust to preference leakage and can serve as a promising alternative for code evaluation. Guang Yang 0019, Yu Zhou 0010, Xiang Chen 0005, Wei Zheng 0006, Xing Hu 0008, Xin Zhou 0014, David Lo 0001, Taolue Chen 0001 |
ASE | 4 |
| 2025 | Adversarial generation method for smart contract fuzz testing seeds guided by chain-based LLM
Jiaze Sun, Zhiqiang Yin, Hengshan Zhang, Xiang Chen 0005, Wei Zheng 0006 |
Autom. Softw. Eng. | 5 |
| 2025 | Enhancing concurrency vulnerability detection through AST-based static fuzz mutation
Wei Zheng 0006, Peiran Deng, Xiang Chen 0005, Xiaoxue Wu 0001 |
J. Syst. Softw. | 1 |
| 2025 | NG_MDERANK: A software vulnerability feature knowledge extraction method based on N-gram similarityabstractAbstract As software grows in size and complexity, software vulnerabilities are increasing, leading to a range of serious insecurity issues. Open‐source software vulnerability reports and documentation can provide researchers with great convenience for analysis and detection. However, the quality of different data sources varies, the data are duplicated and lack of correlation, which often requires a lot of manual management and analysis. In order to solve the problems of scattered and heterogeneous data and lack of correlation in traditional vulnerability repositories, this paper proposes a software vulnerability feature knowledge extraction method that combines the N‐gram model and mask similarity. The method generates mask text data based on the extraction of N‐gram candidate keywords and extracts vulnerability feature knowledge by calculating the similarity of mask text. This method analyzes the samples efficiently and stably in the environment of large sample size and complex samples and can obtain high‐value semi‐structured data. Then, the final node, relationship, and attribute information are obtained by secondary knowledge cleaning and extraction of the extracted semi‐structured data results. And based on the extraction results, the corresponding software vulnerability domain knowledge graph is constructed to deeply explore the semantic information features and entity relationships of vulnerabilities, which can help to efficiently study software security problems and solve vulnerability problems. The effectiveness and superiority of the proposed method is verified by comparing it with several traditional keyword extraction algorithms on Common Weakness Enumeration (CWE) and Common Vulnerabilities and Exposures (CVE) vulnerability data. Xiaoxue Wu 0001, Shiyu Weng, Wei Zheng 0006, Xiang Chen 0005, Xiaobing Sun 0001 |
J. Softw. Evol. Process. | 4 |
| 2024 | Product promotion copywriting from multimodal data: New benchmark and modelabstractIn our latest project, we devise a comprehensive corpus for product promotion text generation, named Video-Enabled Product Promotion Corpus (VPPC), which integrates multimodal and multi-structural information of products such as visual spatial details and fine structural specifics. It is crucial to highlight that this is one of the largest datasets available in the field of video captioning. Notably, conventional multimodal text generation often focuses on regular descriptions of entities and events, which doesn not suffice the real-world requirements of product promotion copywriting, as it necessitates a more lively language style and a high degree of authenticity. Regrettably, there is an evident lack of reusable evaluation frameworks and sufficient datasets at the current stage. To address these challenges, we have proposed a unique baseline approach and authenticity evaluation metric, both tailored to meet the realistic demands of our dataset. The results are promising, as our method surpasses previous approaches across all evaluation metrics. Jinjin Ren, Liming Lin, Wei Zheng 0006 |
Neurocomputing | 3 |
| 2024 | Duplicate Bug Report detection using Named Entity Recognition
Wei Zheng 0006, Xiaoxue Wu 0001, Jingyuan Cheng |
Knowl. Based Syst. | 1 |
| 2024 | ISTA+: Test case generation and optimization for intelligent systems based on coverage analysis
Xiaoxue Wu 0001, Yizeng Gu, Lidan Lin, Wei Zheng 0006, Xiang Chen 0005 |
Sci. Comput. Program. | 4 |
| 2024 | A Novel Tensor Learning Model for Joint Relational Triplet ExtractionabstractThe relational triplet is a format to represent relational facts in the real world, which consists of two entities and a semantic relation between these two entities. Since the relational triplet is the essential component in a knowledge graph (KG), extracting relational triplets from unstructured texts is vital for KG construction and has attached increasing research interest in recent years. In this work, we find that relation correlation is common in real life and could be beneficial for the relational triplet extraction task. However, existing relational triplet extraction works neglect to explore the relation correlation that bottlenecks the model performance. Therefore, to better explore and take advantage of the correlation among semantic relations, we innovatively utilize a three-dimension word relation tensor to describe relations between words in a sentence. Then, we treat the relation extraction task as a tensor learning problem and propose an end-to-end tensor learning model based on Tucker decomposition. Compared with directly capturing correlation among relations in a sentence, learning the correlation of elements in a three-dimension word relation tensor is more feasible and could be addressed through tensor learning methods. To verify the effectiveness of the proposed model, extensive experiments are also conducted on two widely used benchmark datasets, that is, NYT and WebNLG. Results show that our model outperforms the state-of-the-art by a large margin of F1 scores, such as the developed model has an improvement of 3.2% on the NYT dataset compared to the state-of-the-art. Source codes and data can be found at https://github.com/Sirius11311/TLRel.git. Zhen Wang 0004, Hongyi Nie, Wei Zheng 0006, Yaqing Wang 0002, Xuelong Li 0001 |
IEEE Trans. Cybern. | 3 |
| 2024 | An Empirical Study on Correlations Between Deep Neural Network Fairness and Neuron Coverage CriteriaabstractRecently, with the widespread use of deep neural networks (DNNs) in high-stakes decision-making systems (such as fraud detection and prison sentencing), concerns have arisen about the fairness of DNNs in terms of the potential negative impact they may have on individuals and society. Therefore, fairness testing has become an important research topic in DNN testing. At the same time, the neural network coverage criteria (such as criteria based on neuronal activation) is considered as an adequacy test for DNN white-box testing. It is implicitly assumed that improving the coverage can enhance the quality of test suites. Nevertheless, the correlation between DNN fairness (a test property) and coverage criteria (a test method) has not been adequately explored. To address this issue, we conducted a systematic empirical study on seven coverage criteria, six fairness metrics, three fairness testing techniques, and five bias mitigation methods on five DNN models and nine fairness datasets to assess the correlation between coverage criteria and DNN fairness. Our study achieved the following findings: 1) with the increase in the size of the test suite, some of the coverage and fairness metrics changed significantly, as the size of the test suite increased; 2) the statistical correlation between coverage criteria and DNN fairness is limited; and 3) after bias mitigation for improving the fairness of DNN, the change pattern in coverage criteria is different; 4) Models debiased by different bias mitigation methods have a lower correlation between coverage and fairness compared to the original models. Our findings cast doubt on the validity of coverage criteria concerning DNN fairness (i.e., increasing the coverage may even have a negative impact on the fairness of DNNs). Therefore, we warn DNN testers against blindly pursuing higher coverage of coverage criteria at the cost of test properties of DNNs (such as fairness). Wei Zheng 0006, Lidan Lin, Xiaoxue Wu 0001, Xiang Chen 0005 |
IEEE Trans. Software Eng. | 1 |
| 2023 | An Intelligent Duplicate Bug Report Detection Method Based on Technical Term ExtractionabstractAs the bug description data generated during the software maintenance cycle, bug reports are usually hastily written by different users, resulting in many redundant and duplicate bug reports (DBRs). Once the DBRs are repeatedly assigned to developers, it will inevitably lead to a serious waste of human resources, especially for large-scale open-source projects. Recently, many experts and scholars have devoted themselves to researching the detection of DBRs and put forward a series of detection methods for DBRs. However, there is still much room for improvement in the performance of DBR prediction. Therefore, this paper proposes a new method for detecting DBR based on technical term extraction, CTEDB (Combination of Term Extraction and DeBERTaV3) for short. This method first extracts technical terms from the text information of bug reports based on Word2Vec and TextRank algorithms. Then it calculates the semantic similarity of technical terms between different bug reports by combining Word2Vec and SBERT models. Finally, it completes the DBR detection task by combining the DeBERTaV3 model. The experimental results show that CTEDB has achieved good results in detecting DBR, and has obviously improved the accuracy, F1-score, recall and precision compared with the baseline approaches. Xiaoxue Wu 0001, Wenjing Shan, Wei Zheng 0006, Xiaobing Sun 0001 |
AST | 3 |
| 2023 | ISTA: Automatic Test Case Generation and Optimization for Intelligent Systems based on Coverage AnalysisabstractWith the applications of intelligent systems in areas (such as self-driving cars, robotics, and smart cities), the impact of these intelligent systems’ defects cannot be ignored. For example, in a recent report, the self-driving car collided with another self-driving car because it incorrectly identified a roadblock. Therefore, it is necessary to conduct adequate testing of intelligent systems to avoid dangerous behaviors as much as possible. However, due to the particularity of its own structure, the low efficiency, and the high cost of manual collection the large-scale test cases, it is important and challenging to design tools to test the adequacy of intelligent systems.To overcome the above problems, we propose an intelligent system test adequacy evaluation tool ISTA. ISTA implements the automatic generation and optimization of test cases based on coverage analysis, which can improve the test adequacy of the intelligent system while expanding the dataset. To evaluate the usefulness of our developed tool, we analyze the application of ISTA on the five-layer fully-connected dnn model and german credit dataset (text data type) for binary classification as well as on the Rambo model and hmb dataset (image data type) for self-driving car. The evaluation results show that the test dataset is expanded and the models are more fully tested after ISTA’s test case generation and optimization for both text and image data types, with a corresponding increase in the average 80% coverage criteria used. Wei Zheng 0006, Lidan Lin, Xiang Chen 0005, Jinjin Shen, Qingqing Xu, Yizeng Gu |
SANER | 1 |
| 2023 | An Abstract Syntax Tree based static fuzzing mutation for vulnerability evolution analysis
Wei Zheng 0006, Peiran Deng, Kui Gui, Xiaoxue Wu 0001 |
Inf. Softw. Technol. | 1 |
| 2023 | RNNtcs: A test case selection method for Recurrent Neural Networks
Xiaoxue Wu 0001, Jinjin Shen, Wei Zheng 0006, Lidan Lin, Yulei Sui, Abubakar Omari Abdallah Semasaba |
Knowl. Based Syst. | 3 |
| 2023 | An empirical evaluation of deep learning-based source code vulnerability detection: Representation versus modelsabstractAbstract Vulnerabilities in the source code of the software are critical issues in the realm of software engineering. Coping with vulnerabilities in software source code is becoming more challenging due to several aspects such as complexity and volume. Deep learning has gained popularity throughout the years as a means of addressing such issues. This paper proposes an evaluation of vulnerability detection performance on source code representations and evaluates how machine learning (ML) strategies can improve them. The structure of our experiment consists of three deep neural networks (DNNs) in conjunction with five different source code representations: abstract syntax trees (ASTs), code gadgets (CGs), semantics‐based vulnerability candidates (SeVCs), lexed code representations (LCRs), and composite code representations (CCRs). Experimental results show that employing different ML strategies in conjunction with the base model structure influences the performance results to a varying degree. However, ML‐based techniques suffer from poor performance on class imbalance handling and dimensionality reduction when used in conjunction with source code representations. Abubakar Omari Abdallah Semasaba, Wei Zheng 0006, Xiaoxue Wu 0001, Samuel Akwasi Agyemang |
J. Softw. Evol. Process. | 2 |
| 2022 | Interpretability application of the Just-in-Time software defect prediction model
Wei Zheng 0006, Tianren Shen, Xiang Chen 0005, Peiran Deng |
J. Syst. Softw. | 1 |
| 2022 | Domain knowledge-based security bug reports prediction
Wei Zheng 0006, Jingyuan Cheng, Xiaoxue Wu 0001, Ruiyang Sun, Xiaobing Sun 0001 |
Knowl. Based Syst. | 1 |
| 2022 | ReCDroid+: Automated End-to-End Crash Reproduction from Bug Reports for Android AppsabstractThe large demand of mobile devices creates significant concerns about the quality of mobile applications (apps). Developers heavily rely on bug reports in issue tracking systems to reproduce failures (e.g., crashes). However, the process of crash reproduction is often manually done by developers, making the resolution of bugs inefficient, especially given that bug reports are often written in natural language. To improve the productivity of developers in resolving bug reports, in this paper, we introduce a novel approach, called ReCDroid+, that can automatically reproduce crashes from bug reports for Android apps. ReCDroid+ uses a combination of natural language processing (NLP) , deep learning, and dynamic GUI exploration to synthesize event sequences with the goal of reproducing the reported crash. We have evaluated ReCDroid+ on 66 original bug reports from 37 Android apps. The results show that ReCDroid+ successfully reproduced 42 crashes (63.6% success rate) directly from the textual description of the manually reproduced bug reports. A user study involving 12 participants demonstrates that ReCDroid+ can improve the productivity of developers when resolving crash bug reports. Yu Zhao 0010, Ting Su 0001, Yang Liu 0003, Wei Zheng 0006, Xiaoxue Wu 0001, Ramakanth Kavuluru, William G. J. Halfond, Tingting Yu 0001 |
ACM Trans. Softw. Eng. Methodol. | 4 |
| 2022 | Data Quality Matters: A Case Study on Data Label Correctness for Security Bug Report PredictionabstractIn the research of mining software repositories, we need to label a large amount of data to construct a predictive model. The correctness of the labels will affect the performance of a model substantially. However, limited studies have been performed to investigate the impact of mislabeled instances on a predictive model. To bridge the gap, in this article, we perform a case study on the security bug report (SBR) prediction. We found five publicly available datasets for SBR prediction contains many mislabeled instances, which lead to the poor performance of SBR prediction models of recent studies (e.g., the work of Peterset al.and Shuet al.). Furthermore, it might mislead the research direction of SBR prediction. In this article, we first improve the label correctness of these five datasets by manually analyzing each bug report, and we find 749 SBRs, which are originally mislabeled as Non-SBRs (NSBRs). We then evaluate the impacts of datasets label correctness by comparing the performance of the classification models on both the noisy (i.e., before our correction) and the clean (i.e., after our correction) datasets. The results show that the cleaned datasets result in improvement in the performance of classification models. The performance of the approaches proposed by Peterset al.and Shuet al.on the clean datasets is much better than on the noisy datasets. Furthermore, with the clean datasets, the simple text classification models could significantly outperform the security keywords-matrix-based approaches applied by Peterset al.and Shuet al. Xiaoxue Wu 0001, Wei Zheng 0006, Xin Xia 0001, David Lo 0001 |
IEEE Trans. Software Eng. | 2 |
| 2021 | Automatically Identifying Bug Reports with Tactical Vulnerabilities by Deep Feature LearningabstractIdentifying and fixing bug reports with tactical vul-nerabilities in a timely and accurate manner is essential to ensure the security of the software architecture. Manually identifying the bug reports with tactical vulnerabilities is labor-intensive and challenging. This paper presents Itactivul, an approach to automatically identify bug reports with tactical vulnerabilities and recommend their tactical categories to guide the fix. Unlike the existing security bug report prediction approach, we are the first attempt to use deep learning to mine discriminative tactical text features only from the vulnerability descriptions of the National Vulnerability Database (NVD) and apply them to identify bug reports with tactical vulnerabilities. We evaluate Itactivul on three bug reports datasets gathered from three large-scale open-source projects, including Chromium, PHP, and Thunderbird. The experimental results show that Itactivul outperforms baselines by an average of 8.88 %, 13.58 %, and 6.61 % in the F1-score of three datasets, respectively. To improve the explainability of the features mined by Itactivul, we manually analyze the high-weight phrases extracted by using attention backtracking. The results show that Itactivul can mine key and potential tactical vulnerabilities text features. Wei Zheng 0006, Manqing Zhang, Yuanfang Cai, Xiang Chen 0005, Xiaoxue Wu 0001, Abubakar Omari Abdallah Semasaba |
ISSRE | 1 |
| 2021 | Research Progress of Flaky TestsabstractA flaky test is a test that both passes and fails periodically without any code changes, and its uncontrolled uncertainty will destroy the value of the test suites and even cause developers to distrust the test results. Recently, researches in the flaky test have received broad attention in the software test community to reduce the manual maintenance cost of flaky tests by developers. In this survey, we conducted comprehensive research progress on the flaky test We identified 31 relevant studies and summarized the following aspects of the flaky test: root causes and factors, analyzing the impact, detecting and classifying techniques, and fixing approaches. This survey also identifies open research challenges to be further explored in future work. Wei Zheng 0006, Manqing Zhang, Xiang Chen 0005, Wenqiao Zhao |
SANER | 1 |
| 2021 | Representation vs. Model: What Matters Most for Source Code Vulnerability DetectionabstractVulnerabilities in the source code of software are critical issues in the realm of software engineering. Coping with vulnerabilities in software source code is becoming more challenging due to several aspects of complexity and volume. Deep learning has gained popularity throughout the years as a means of addressing such issues. In this paper, we propose an evaluation of vulnerability detection performance on source code representations and evaluate how Machine Learning (ML) strategies can improve them. The structure of our experiment consists of 3 Deep Neural Networks (DNNs) in conjunction with five different source code representations; Abstract Syntax Trees (ASTs), Code Gadgets (CGs), Semantics-based Vulnerability Candidates (SeVCs), Lexed Code Representations (LCRs), and Composite Code Representations (CCRs). Experimental results show that employing different ML strategies in conjunction with the base model structure influences the performance results to a varying degree. However, ML-based techniques suffer from poor performance on class imbalance handling when used in conjunction with source code representations for software vulnerability detection. Wei Zheng 0006, Abubakar Omari Abdallah Semasaba, Xiaoxue Wu 0001, Samuel Akwasi Agyemang |
SANER | 1 |
| 2021 | A survey of Intel SGX and its applications
Wei Zheng 0006, Xiaoxue Wu 0001, Chen Feng 0005, Yulei Sui, Xiapu Luo, Yajin Zhou |
Frontiers Comput. Sci. | 1 |
| 2021 | WRTRe: Weighted relative position transformer for joint entity and relation extraction
Wei Zheng 0006, Zhen Wang 0004, Quanming Yao, Xuelong Li 0001 |
Neurocomputing | 1 |
| 2021 | Improving high-impact bug report prediction with combination of interactive machine learning and active learning
Xiaoxue Wu 0001, Wei Zheng 0006, Xiang Chen 0005, Yu Zhao 0010, Tingting Yu 0001 |
Inf. Softw. Technol. | 2 |
| 2021 | A Comparative Study of Class Rebalancing Methods for Security Bug Report ClassificationabstractIdentifying security bug reports (SBRs) accurately from a bug repository can reduce a software product’s security risk. However, the class imbalance problem exists for SBR prediction since the number of SBRs is often limited, and this issue has not been thoroughly investigated in previous studies. In our study, we choose six real-world projects of different sizes with over 120 000 bug reports in total as our empirical subjects. We first analyze the impact of the class imbalance issue on SBR prediction and confirm its negative impact on prediction performance. Then we perform a comparative study of six state-of-the-art class rebalancing methods combined with five popular classification algorithms for SBR prediction. By comparing with the baseline method Farsec, using the class rebalancing methods can improve the performance in 78% of cases in the worst case. Moreover, the combination of the Rose and random forest classification algorithm can construct the model with the best performance, which increases the performance by 267% in the best case and 75% on average in terms ofF1-score. Finally, we summarize eight main findings based on our empirical studies’ results, which can provide guidelines for choosing appropriate class rebalancing methods and classifiers for SBR prediction in practice. Wei Zheng 0006, Yuxing Xun, Xiaoxue Wu 0001, Xiang Chen 0005, Yulei Sui |
IEEE Trans. Reliab. | 1 |
| 2020 | How Well Just-In-Time Defect Prediction Techniques Enhance Software Reliability?abstractMany Just-In-Time defect prediction (JIT) techniques, which anticipate defect-prone software changes, have been proposed in recent years. Researchers have evaluated these techniques from different perspectives and have drawn inconsistent conclusions about which JIT defect prediction techniques are the most effective and efficient. This paper evaluates JIT techniques from a reliability perspective. For short-term early evaluation, we measure JIT predictive performance on early exposed defects. While for long-term evaluation, we quantify the overall reliability improvement resulted from JIT. A case study applying 11 state-of-the-art JIT methods on 18 large open-source projects has shown: 1) Different JIT methods have their own individual strengths for different purposes, 2) in general, RandomForest is the most effective method in short-term software reliability improvement, and CBS+ performs best in long-term reliability improvement; 3) JIT prediction accuracy is highly correlated to overall reliability improvement. Yuli Tian, Ning Li 0022, Jeff Tian, Wei Zheng 0006 |
QRS | 4 |
| 2020 | Literature survey of deep learning-based vulnerability analysis on source codeabstractVulnerabilities in software source code are one of the critical issues in the realm of software code auditing. Due to their high impact, several approaches have been studied in the past few years to mitigate the damages from such vulnerabilities. Among the approaches, deep learning has gained popularity throughout the years to address such issues. In this literature survey, the authors provide an extensive review of the many works in the field software vulnerability analysis that utilise deep learning-based techniques. The reviewed works are systemised according to their objectives (i.e. the type of vulnerability analysis aspect), the area of focus (i.e. the focus area of the analysis), what information about source code is used (i.e. the features), and what deep learning techniques they employ (i.e. what algorithm is used to process the input and produce the output). They also study the limitations of the papers and topical trends concerning vulnerability analysis. Abubakar Omari Abdallah Semasaba, Wei Zheng 0006, Xiaoxue Wu 0001, Samuel Akwasi Agyemang |
IET Softw. | 2 |
| 2020 | CVE-assisted large-scale security bug report dataset construction method
Xiaoxue Wu 0001, Wei Zheng 0006, Xiang Chen 0005 |
J. Syst. Softw. | 2 |
| 2020 | The impact factors on the performance of machine learning-based vulnerability detection: A comparative study
Wei Zheng 0006, Jialiang Gao, Xiaoxue Wu 0001, Yuxing Xun, Xiang Chen 0005 |
J. Syst. Softw. | 1 |
| 2020 | Invalid bug reports complicate the software aging situation
Xiaoxue Wu 0001, Wei Zheng 0006, Minchao Pu |
Softw. Qual. J. | 2 |
| 2019 | ReCDroid: automatically reproducing Android application crashes from bug reportsabstractThe large demand of mobile devices creates significant concerns about the quality of mobile applications (apps). Developers heavily rely on bug reports in issue tracking systems to reproduce failures (e.g., crashes). However, the process of crash reproduction is often manually done by developers, making the resolution of bugs inefficient, especially that bug reports are often written in natural language. To improve the productivity of developers in resolving bug reports, in this paper, we introduce a novel approach, called ReCDroid, that can automatically reproduce crashes from bug reports for Android apps. ReCDroid uses a combination of natural language processing (NLP) and dynamic GUI exploration to synthesize event sequences with the goal of reproducing the reported crash. We have evaluated ReCDroid on 51 original bug reports from 33 Android apps. The results show that ReCDroid successfully reproduced 33 crashes (63.5% success rate) directly from the textual description of bug reports. A user study involving 12 participants demonstrates that ReCDroid can improve the productivity of developers when resolving crash bug reports. Yu Zhao 0010, Tingting Yu 0001, Ting Su 0001, Yang Liu 0003, Wei Zheng 0006, Jingzhi Zhang, William G. J. Halfond |
ICSE | 5 |
| 2019 | Automatically Extracting Bug Reproducing Steps from Android Bug Reports
Yu Zhao 0010, Kye Miller, Tingting Yu 0001, Wei Zheng 0006, Minchao Pu |
ICSR | 4 |
| 2019 | Event trace reduction for effective bug replay of Android apps via differential GUI state analysisabstractExisting Android testing tools, such as Monkey, generate a large quantity and a wide variety of user events to expose latent GUI bugs in Android apps. However, even if a bug is found, a majority of the events thus generated are often redundant and bug-irrelevant. In addition, it is also time-consuming for developers to localize and replay the bug given a long and tedious event sequence (trace). Yulei Sui, Yifei Zhang 0001, Wei Zheng 0006, Manqing Zhang, Jingling Xue |
ESEC/SIGSOFT FSE | 3 |
| 2019 | Towards understanding bugs in an open source cloud management stack: An empirical study of OpenStack software bugs
Wei Zheng 0006, Chen Feng 0005, Tingting Yu 0001, Xibing Yang, Xiaoxue Wu 0001 |
J. Syst. Softw. | 1 |
| 2018 | MS-guided many-objective evolutionary optimisation for test suite minimisationabstractTest suite minimisation is a process that seeks to identify and then eliminate the obsolete orredundant test cases from the test suite. It is a trade‐off between cost andother value criteria and is appropriate to be described as a many‐objectiveoptimisation problem. This study introduces a mutation score (MS)‐guidedmany‐objective optimisation approach, which prioritises the fault detectionability of test cases and takes MS, cost and three standard code coveragecriteria as objectives for the test suite minimisation process. They use sixclassical evolutionary many‐objective optimisation algorithms to identifyefficient test suite, and select three small programs from the Software‐ArtefactInfrastructure Repository (SIR) and two larger program space and gzip forexperimental evaluation as well as statistical analysis. The experiment resultsof the three small programs show non‐dominated sorting genetic algorithm II(NSGA‐II) with tuning was the most effective approach. However, MOEA/D‐PBI andMOEA/D‐WS outperform NSGA‐II in the cases of two large programs. On the otherhand, the test cost of the optimal test suite obtained by their proposedMS‐guided many‐objective optimisation approach is much lower than the onewithout it in most situation for both small programs and large programs. Wei Zheng 0006, Xiaoxue Wu 0001, Shichao Cao |
IET Softw. | 1 |
| 2017 | Rapid Line-Extraction Method for SAR Images Based on Edge-Field FeaturesabstractThis letter proposes a rapid line-extraction (RLE) method for synthetic aperture radar (SAR) images. RLE first transforms an image in the space domain into an image in the frequency domain. Then, using the central-slice theorem, RLE skilfully maps the image in the frequency domain into a parameter space, which effectively accelerates the straight-line extraction process. Unlike the traditional Hough transform, RLE is performed directly on an edge-field image rather than on a binary edge map. Theoretical analysis proves the advantages of using the edge-field map. Notably, the computational complexity can be greatly reduced relative to the complexity of obtaining a binary edge map, and the method can efficiently avoid the negative influence of false edges in the binary edge map. More importantly, because speckle, clutter, and blurred edges in real-world images decrease the sharpness of peaks, edge-field images that include the strength and direction information of SAR images are adopted to reduce the diffusion of peaks and improve the detection accuracy. Experimental studies show that RLE works independently, is robust to noise, has low computational complexity, achieves high true-positive detection rates, and yields satisfactory detection precision. Qian-Ru Wei, Da-Zheng Feng, Wei Zheng 0006, Jiangbin Zheng 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2016 | Multi-objective optimisation for regression testing
Wei Zheng 0006, Robert M. Hierons, Miqing Li, Xiaohui Liu 0001, Veronica Vinciotti |
Inf. Sci. | 1 |
| 2016 | SIP: Optimal Product Selection from Feature Models Using Many-Objective Evolutionary OptimizationabstractA feature model specifies the sets of features that define valid products in a software product line. Recent work has considered the problem of choosing optimal products from a feature model based on a set of user preferences, with this being represented as a many-objective optimization problem. This problem has been found to be difficult for a purely search-based approach, leading to classical many-objective optimization algorithms being enhanced either by adding in a valid product as a seed or by introducing additional mutation and replacement operators that use an SAT solver. In this article, we instead enhance the search in two ways: by providing a novel representation and by optimizing first on the number of constraints that hold and only then on the other objectives. In the evaluation, we also used feature models with realistic attributes, in contrast to previous work that used randomly generated attribute values. The results of experiments were promising, with the proposed (SIP) method returning valid products with six published feature models and a randomly generated feature model with 10,000 features. For the model with 10,000 features, the search took only a few minutes. Robert M. Hierons, Miqing Li, Xiaohui Liu 0001, Sergio Segura, Wei Zheng 0006 |
ACM Trans. Softw. Eng. Methodol. | 5 |