VLDB 2026 Research / reviewers in the wild / expert
Chenhui Cui
dblp:240/2312
· DBLP profile ↗
9ranked-venue papers
1as first author
8since 2021 · last 2026
0009-0004-8746-316XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Large Language Models for Automated Web-Form-Test Generation: An Empirical StudyabstractTesting web forms is an essential activity for ensuring the quality of web applications. It typically involves evaluating the interactions between users and forms. Automated test-case generation remains a challenge for web-form testing: Due to the complex, multi-level structure of Web pages, it can be difficult to automatically capture their inherent contextual information for inclusion in the tests. Large Language Models (LLMs) have shown great potential for contextual text generation. This motivated us to explore how they could generate automated tests for web forms, making use of the contextual information within form elements. To the best of our knowledge, no comparative study examining different LLMs has yet been reported for web-form-test generation. To address this gap in the literature, we conducted a comprehensive empirical study investigating the effectiveness of 11 LLMs on 146 web forms from 30 open source Java web applications. In addition, we propose three HTML-structure-pruning methods to extract key contextual information. The experimental results show that different LLMs can achieve different testing effectiveness, with the GPT-4, GLM-4, and Baichuan2 LLMs generating the best web-form tests. Compared with GPT-4, the other LLMs had difficulty generating appropriate tests for the web forms: Their Successfully Submitted Rates (SSRs)—the proportions of the LLMs-generated web-form tests that could be successfully inserted into the web forms and submitted—decreased by 9.10% to 74.15%. Our findings also show that, for all LLMs, when the designed prompts include complete and clear contextual information about the web forms, more effective web-form tests were generated. Specifically, when using Parser-Processed HTML for Task Prompt (PH-P), the SSR averaged 70.63%, higher than the 60.21% for Raw HTML for Task Prompt (RH-P) and 50.27% for LLM-Processed HTML for Task Prompt (LH-P). With RH-P, GPT-4’s SSR was 98.86%, outperforming models like LLaMa2 (7B) with 34.47% and GLM-4V with 0%. Similarly, with PH-P, GPT-4 reached an SSR of 99.54%, the highest among all models and prompt types. Finally, this article also highlights strategies for selecting LLMs based on performance metrics, and for optimizing the prompt design to improve the quality of the web-form tests. Chenhui Cui, Rubing Huang, Dave Towey, Lei Ma 0003 |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2026 | A Novel Vision-Based Approach to Test Sequence Generation for Mobile GUI TestingabstractMobileGraphical User Interface(GUI) testing is a critical component of quality assurance for mobile applications. As GUI designs grow increasingly complex, vision-based testing methods have become essential for improving the quality of test scripts and reports. However, current approaches face significant limitations. For instance, the process of cropping GUI widgets requires significant manual effort and consumes considerable time. Meanwhile, existing automated testing tools still fail to generate satisfactory test sequences. In this paper, we proposeVision-based Test Sequence Generation(VTSG), a novel perception-driven approach for mobile GUI testing. More specifically: (1) VTSG employs a light-weight GUI element detection model to crop widgets from GUI pages automatically; (2) Guided by human perception principles, it sequences widget screenshots through saturation and spatial layout analysis; (3) The system then integrates GUI widget interaction instructions to generate vision-based test scripts that accurately simulate human interaction patterns on mobile devices. We evaluate VTSG against four state-of-the-art (SOTA) GUI testing tools across five applications. Experimental results demonstrate that VTSG significantly outperforms existing methods, achieving 47.44% code coverage and 51.33% activity coverage, respectively, compared to the other approaches. Additionally, we conduct a series of supplementary experiments on two mainstream commercial applications (i.e.,ToutiaoandDouyin). The results further confirm that VTSG maintains higher activity coverage even on these real-world commercial apps. Chenhui Cui, Yinming Huang, Rubing Huang, Ling Zhou 0005, Rongcun Wang |
IEEE Trans. Reliab. | 1 |
| 2025 | Short-Term Electricity-Load Forecasting by deep learning: A comprehensive survey
Rubing Huang, Chenhui Cui, Dave Towey, Ling Zhou 0005, Jinyu Tian 0001 |
Eng. Appl. Artif. Intell. | 3 |
| 2025 | Applying Lexicographical Ordering to Software Product Line TestingabstractTest case prioritization (TCP) has been widely used in software testing, which aims to execute test cases that are more likely to detect faults earlier than others. Among many proposed TCP approaches, lexicographical ordering-based TCP (LO-TCP) can effectively resolve ties encountered in the prioritization process, leading to better performance than original TCP approaches. However, the current LO-TCP needs to use the white-box information such as the code coverage of the program under test, which may be infeasible in some black-box testing applications such as software product lines (SPLs). In this article, we transfer the traditional LO-TCP to SPL testing by leveraging test configuration coverage instead of code coverage, and also empirically conduct some simulations and evaluate the large-scale real-world programs with real faults. The experimental results show that LO-TCP can have better performance for testing SPLs, as compared with traditional TCP approaches. Chenhui Cui, Yinyin Xu, Rubing Huang |
IEEE Trans. Reliab. | 2 |
| 2025 | Adaptive Random Testing of Deep Learning Systems Using Image HashingabstractIn recent years, deep learning (DL) systems have been applied in many areas, including image processing and autonomous driving. Software testing is an important way to ensure the quality of software. Among various testing methods, random testing (RT) has been widely used for DL systems, due to its simplicity and efficiency. However, it has been criticized for its poor fault-detection effectiveness. As an enhancement of RT, adaptive random testing (ART) attempts to evenly spread test cases over the input domain, aiming to improve the distribution diversity. However, current ART methods for DL systems have low testing efficiency, particularly for image-based DL systems. This is because of the current reliance on visual geometry group network-16 (VGGNet-16) to extract image features to represent image inputs—VGGNet-16 is a 16-layer, deep convolutional neural network that has been widely used for image classification and feature extraction. Feature extraction with VGGNet-16 is very time-consuming, with each image being extracted as a high-dimensional vector. The (dis)similarity calculations for images with such high-dimensional vectors incur heavy computational overheads. To overcome these challenges, we propose a new ART approach:image-hashing-based ART(IHART). IHART uses image hashing to quickly extract features from each image, storing them as a low-dimensional binary vector. This can significantly reduce the computational costs for dissimilarity calculations during test-case generation. We report on a series of experiments, using several well-known datasets and DL systems, to evaluate the IHART performance. Our results show that, of the three mainstream image-hashing strategies studied, perceptual hashing delivers the best ART test-case generation performance—perceptual hashing, which is used in image deduplication and content searching, uses content features in the hashing process. Compared with current approaches, IHART performs very well in fault-detection effectiveness across most datasets and models, and significantly better fault-detection efficiency. Linwei Yi, Chenhui Cui, Rubing Huang, Dave Towey, Rongcun Wang |
IEEE Trans. Reliab. | 2 |
| 2024 | Toward Cost-Effective Adaptive Random Testing: An Approximate Nearest Neighbor ApproachabstractAdaptive Random Testing(ART) enhances the testing effectiveness (including fault-detection capability) ofRandom Testing(RT) by increasing the diversity of the random test cases throughout the input domain. Many ART algorithms have been investigated such asFixed-Size-Candidate-Set ART(FSCS) andRestricted Random Testing(RRT), and have been widely used in many practical applications. Despite its popularity, ART suffers from the problem of high computational costs during test-case generation, especially as the number of test cases increases. Although several strategies have been proposed to enhance the ART testing efficiency, such as theforgetting strategyand thek-dimensional tree strategy, these algorithms still face some challenges, including: (1) Although these algorithms can reduce the computation time, their execution costs are still very high, especially when the number of test cases is large; and (2) To achieve low computational costs, they may sacrifice some fault-detection capability. In this paper, we propose an approach based onApproximate Nearest Neighbors(ANNs), calledLocality-Sensitive Hashing ART(LSH-ART). When calculating distances among different test inputs, LSH-ART identifies the approximate (not necessarily exact) nearest neighbors for candidates in an efficient way. LSH-ART attempts to balance ART testing effectiveness and efficiency. Rubing Huang, Chenhui Cui, Junlong Lian, Dave Towey, Weifeng Sun 0004, Haibo Chen 0005 |
IEEE Trans. Software Eng. | 2 |
| 2023 | VPP-ART: An Efficient Implementation of Fixed-Size-Candidate-Set Adaptive Random Testing Using Vantage Point PartitioningabstractAdaptive random testing(ART) is an enhancement ofrandom testing(RT), and aims to improve the RT failure-detection effectiveness by distributing test cases more evenly in the input domain. Many ART algorithms have been proposed, withfixed-size-candidate-setART (FSCS-ART) being one of the most effective and popular. FSCS-ART ensures high failure-detection effectiveness by selecting as the next test case the candidate farthest from previously executed test cases. Although FSCS-ART has good failure-detection effectiveness, it also faces some challenges, including heavy computational overheads. In this article, we propose an enhanced version of FSCS-ART,vantage point partitioning ART(VPP-ART). VPP-ART addresses the FSCS-ART computational overhead problem using VPP, while maintaining the failure-detection effectiveness. VPP-ART partitions the input domain space using amodified vantage point tree(VP-tree) and finds the approximate nearest executed test cases of a candidate test case in the partitioned subdomains—thereby significantly reducing the time overheads compared with the searches required for FSCS-ART. To enable the FSCS-ART dynamic insertion process, we modify the traditional VP-tree to support dynamic data. The simulation results show that VPP-ART has a much lower time overhead compared to FSCS-ART, but also delivers similar (or better) failure-detection effectiveness, especially in the higher dimensional input domains. According to statistical analyses, VPP-ART can improve on the FSCS-ART failure-detection effectiveness by approximately 50–58%. VPP-ART also compares favorably with theKD-tree-enhanced fixed-size-candidate-set ART(KDFC-ART) algorithms (a series of enhanced ART algorithms based on the KD-tree). Our experiments also show that VPP-ART is more cost-effective than FSCS-ART and KDFC-ART. Rubing Huang, Chenhui Cui, Dave Towey, Weifeng Sun 0004, Junlong Lian |
IEEE Trans. Reliab. | 2 |
| 2022 | A nearest-neighbor divide-and-conquer approach for adaptive random testing
Rubing Huang, Weifeng Sun 0004, Haibo Chen 0005, Chenhui Cui |
Sci. Comput. Program. | 4 |
| 2020 | Poster: Is Euclidean Distance the best Distance Measurement for Adaptive Random Testing?abstractAdaptive random testing (ART) aims at enhancing the testing effectiveness of random testing (RT) by more evenly spreading test cases over the input domain. Many ART methods have been proposed, based on various, different notions. For example, distance-based ART (DART) makes use of the concept of distance to implement ART, attempting to generate new test cases that are far away from previously executed ones. The Euclidean distance has been a popular choice of distance metric, used in DART to evaluate the differences between test cases. However, is the Euclidean distance the most suitable choice for DART? To answer this question, we conducted a series of simulations to investigate the impact that the Euclidean distance, and its many variations, has on the testing effectiveness of DART. The results show that when the dimensionality of the input domain is low, the Euclidean distance may indeed be a good choice. However, when the dimensionality is high, it appears to be less suitable. Rubing Huang, Chenhui Cui, Weifeng Sun 0004, Dave Towey |
ICST | 2 |