Shihao Weng

dblp:58/10149 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
6since 2021 · last 2026
0009-0004-8817-8723ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 5 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Data preparation and quality for code-centric generative software engineering tasks: a systematic literature review
abstract
Abstract The rapid advancements in Deep Neural Networks (DNNs) have revolutionized generative software engineering tasks, including code summarization, program repair, code generation, and code translation. However, the performance of DNN models in these tasks heavily depends on the quality of their training and evaluation datasets. This systematic literature review examines 70 primary studies to comprehensively analyze dataset construction methodologies, prevalent data quality challenges, and solutions proposed to address these challenges. Our findings reveal that dataset construction processes significantly influence quality, with common issues such as noise, redundancy, imbalance, and insufficient granularity undermining model effectiveness. We identify key strategies to mitigate these problems, including data augmentation, automated cleaning techniques, and standardized validation frameworks. Furthermore, we highlight the critical role of dataset diversity and timeliness in improving model generalization. This study provides actionable insights for researchers and practitioners in the era of generative AI, where high-quality datasets are essential for developing reliable language models as software engineering tools. By emphasizing rigorous dataset curation and innovative quality assurance methods, our work bridges the gap between theoretical advancements and practical applications, enabling the creation of robust, generalizable models for real-world code-related tasks. The synthesized recommendations aim to guide future research in optimizing dataset design, fostering reproducibility, and addressing evolving challenges in data-driven software engineering.
Shihao Weng, Yang Feng 0003, Yining Yin, Zhenlun Zhang, Baowen Xu
Frontiers Comput. Sci.1
2025 Lightweight Probabilistic Coverage Metrics for Efficient Testing of Deep Neural Networks
abstract
Deep neural networks (DNNs) have been deployed in many software systems to assist in various tasks.Accompanying with great performance, however, DNNs could also exhibit erroneous behaviors and cause massive losses.To assist the quality assurance and measure the testing adequacy of DNNs, recent research has proposed many neuron coverage (NC) metrics that measure the proportion of neurons activated in executions.While neuron coverage metrics are an analogy to structural code coverage for conventional software programs and reflect the internal behaviors of DNN models in executions, we still lack a comprehensive understanding about the application effectiveness of neuron coverage for deep learning testing.Besides, technologies like DeepGini and ATS have demonstrated the superiority of output probability vectors over neuron coverage for test selection, these techniques do not serve as coverage metrics and thus cannot be directly compared with neuron coverage in other deep learning testing tasks.This paper systematically evaluates the effectiveness of neuron activation-based coverage in multiple testing application scenarios.In addition, to better understand neuron coverage bottlenecks, we further propose an output-probability vector-based coverage metric (named Pt) inspired by existing test selection technique.We perform a comprehensive experiments across three prevalent application scenarios: assessing dataset diversity, improving model retraining, and guiding test generation.Experimental results show that most neuron coverage techniques are not very effective in deep learning testing.Coverage based on neuron activation state do not improve testing efficiency like code coverage.In contrast, the output-based coverage we introduced demonstrates significantly enhanced effectiveness.Our study improves the comprehension of * Yang Feng is the corresponding author.
Yining Yin, Yang Feng 0003, Shihao Weng, Jia Liu 0015
Internetware3
2024 Datactive: Data Fault Localization for Object Detection Systems
abstract
Object detection (OD) models are seamlessly integrated into numerous intelligent software systems, playing a crucial role in various tasks. These models are typically constructed upon humanannotated datasets, whose quality can greatly affect their performance and reliability. Erroneous and inadequate annotated datasets can induce classification/localization inaccuracies during deployment, precipitating security breaches or traffic accidents that inflict property damage or even loss of life. Therefore, ensuring and improving data quality is a crucial issue for the reliability of the object detection system. This paper introduces Datactive, a data fault localization technique for object detection systems. Datactive is designed to locate various types of data faults including mislocalization and missing objects, without utilizing the prediction of object detection models trained on dirty datasets. To achieve this, we first construct foreground-only and background-included datasets via data disassembling strategies, and then employ a robust learning method to train classifiers using disassembled datasets. Based on the classifier predictions, Datactive produces a unified suspiciousness score for both foreground annotations and image backgrounds. It allows testers to easily identify and correct faulty or missing annotations with minimal effort. To validate the effectiveness, we conducted experiments on three datasets with 6 baselines, and demonstrated the superiority of Datactive from various aspects. We also explored Datactive's ability to find natural data faults and its application in both training and evaluation scenarios.
Yining Yin, Yang Feng 0003, Shihao Weng, Yuan Yao 0001, Jia Liu 0015
ISSTA3
2024 Seeing the invisible: test prioritization for object detection system
Shihao Weng, Yang Feng 0003, Yining Yin, Yuxuan Dai
Empir. Softw. Eng.1
2023 Prioritizing Testing Instances to Enhance the Robustness of Object Detection Systems
abstract
Object detection models have been widely deployed in military and life-related intelligent software systems. However, along with the outstanding success of object detection, it may exhibit abnormal behavior and lead to severe accidents and losses. During the development and evaluation process, training and evaluating an object detection model are computationally intensive, while preparing annotated tests requires extremely heavy manual labor. Therefore, reducing the annotation budget of test data collection becomes a challenging and necessary task. Although many test prioritization approaches for DNN-based systems have been proposed, the large differences between classification and object detection make them difficult to apply to testing object detection models.
Shihao Weng, Yang Feng 0003, Yining Yin, Jia Liu 0015
Internetware1
2023 Dynamic Data Fault Localization for Deep Neural Networks
abstract
Rich datasets have empowered various deep learning (DL) applications, leading to remarkable success in many fields. However, data faults hidden in the datasets could result in DL applications behaving unpredictably and even cause massive monetary and life losses. To alleviate this problem, in this paper, we propose a dynamic data fault localization approach, namely DFauLo, to locate the mislabeled and noisy data in the deep learning datasets. DFauLo is inspired by the conventional mutation-based code fault localization, but utilizes the differences between DNN mutants to amplify and identify the potential data faults. Specifically, it first generates multiple DNN model mutants of the original trained model. Then it extracts features from these mutants and maps them into a suspiciousness score indicating the probability of the given data being a data fault. Moreover, DFauLo is the first dynamic data fault localization technique, prioritizing the suspected data based on user feedback, and providing the generalizability to unseen data faults during training. To validate DFauLo, we extensively evaluate it on 26 cases with various fault types, data types, and model structures. We also evaluate DFauLo on three widely-used benchmark datasets. The results show that DFauLo outperforms the state-of-the-art techniques in almost all cases and locates hundreds of different types of real data faults in benchmark datasets.
Yining Yin, Yang Feng 0003, Shihao Weng, Yuan Yao 0001, Zhenyu Chen 0001
ESEC/SIGSOFT FSE3