VLDB 2026 Research / reviewers in the wild / expert
Haonan Zhang 0006
dblp:238/5984-6
· DBLP profile ↗
5ranked-venue papers
3as first author
5since 2021 · last 2025
0000-0002-6874-5581ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 5 · 3 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Who's to Blame? Rethinking the Brittleness of Automated Web GUI Testing from a Pragmatic PerspectiveabstractAutomated web GUI testing is important for software quality, however, its effectiveness is often undermined by test case brittleness, especially in continuously evolving real-world applications. In this experience paper, we pragmatically investigate the root causes of brittleness. We first analyze why legacy test cases, derived from the Mind2Web dataset, fail when executed on current web application versions. Our findings reveal that brittleness stems from multifaceted factors, including test script design, web application complexity, and automation framework limitations. A longitudinal study further shows that 81.7% of repaired tests break again within six months, primarily due to similar recurring issues, highlighting the persistent nature of brittleness. We further demonstrate that Large Language Models, when provided with human-like diagnostic context, can successfully repair a substantial portion of these brittle tests, though human expertise remains important for more complex scenarios. Our findings emphasize that brittleness is a multifaceted problem requiring collaboration between different parts involved in the automation testing. Haonan Zhang 0006, Kundi Yao, Zishuo Ding, Lizhi Liao, Weiyi Shang |
ASE | 1 |
| 2025 | An Empirical Study of Logging Practice in CUDA-Based Deep Learning SystemsabstractAlthough logging practices have been extensively explored in conventional software systems, there remains a lack of understanding of how logging is applied in CUDAbased deep learning (DL) systems, despite their growing adoption in practice. In this paper, we conduct an empirical study to examine the characteristics and rationales of logging practices in these systems. We analyze logging statements from 33 CUDA-based open-source DL projects, covering both general-purpose logging libraries and DL-specific logging frameworks. For each type, we identify the development or execution phases in which the logs are used and investigate the reasoning behind their usage. Our quantitative analysis reveals that the majority of logging statements occur during the model training phase, with significant usage also in the model loading phase and model evaluation/validation phase. Furthermore, we observe that logging is predominantly used for monitoring purposes and tracking model-related information. Our findings not only shed light on current logging practices in CUDAbased DL development but also provide practical guidance on when to use DL-specific versus general-purpose logging, helping practitioners make more informed decisions and guiding the evolution of DL-focused logging tools to better support developer needs. Kundi Yao, Haonan Zhang 0006, Yiming Tang 0002, Weiyi Shang |
QRS | 3 |
| 2025 | Improving Qa System Testing Efficiency Through White-Box Test PrioritizationabstractEffective testing of sequence-to-sequence (seq2seq) models, such as those used in question answering (QA) systems, is essential for ensuring their reliability. While recent efforts have introduced metamorphic testing strategies to detect bugs without requiring ground-truth labels, the efficiency of these methods remains limited by their lack of test case prioritization. Executing all test cases uniformly can lead to wasted resources and slower fault discovery. In this paper, we propose a white-box prioritization framework that ranks test cases based on internal signals extracted from the underlying model. Building upon a prior work that introduced two whitebox techniques (i.e., GRI and WALI) for identifying vulnerable tokens, we adapt these techniques to the task of test prioritization. Instead of generating new test inputs, our methods analyze test cases produced by QAQA and prioritize those most likely to uncover faults. We evaluate our approaches on three widely-used QA datasets: BoolQ, NarrativeQA, and SQuAD2. Experimental results show that GRI significantly improves the rate of bug detection under constrained testing budgets, while WALI achieves comparable performance to baseline methods. Our findings demonstrate the value of incorporating white-box insights into the prioritization process, offering a more efficient and effective way to test QA systems. Hanying Shao, Zishuo Ding, Kundi Yao, Haonan Zhang 0006, Weiyi Shang |
QRS | 4 |
| 2024 | Towards a Robust Waiting Strategy for Web GUI Testing for an Industrial Software SystemabstractAutomated web GUI testing has been widely adopted since manual testing is time-consuming and tedious. Waiting strategy plays a vital role in automated web GUI testing since it significantly impacts the testing performance. Though important, little focus has been set on the waiting strategies in web GUI testing. Existing waiting strategies either wait for a predetermined time, which is not reliable in a dynamic environment, or only wait for a specific condition to be verified, which is often not robust enough to handle the complicated testing scenarios. In this work, we introduce a robust waiting strategy. Instead of waiting for a predetermined time or waiting for the availability of a particular element, our approach waits for a desired state to reach. This is achieved by capturing the Document Object Models (DOM) at the desired point, followed by an offline analysis to identify the differences between the DOMs associated with every two consecutive test actions. Such differences are used to determine the appropriate waiting time when automatically generating tests. Evaluation results with an industrial web application indicate that our approach produces more robust tests than the conventional waiting strategies used in web GUI testing. Furthermore, our generated tests are more representative of the recorded usage scenarios and are efficient with low overhead in test execution time. Haonan Zhang 0006, Lizhi Liao, Zishuo Ding, Weiyi Shang, Nidhi Narula, Catalin Sporea, Andrei Toma, Sarah Sajedi |
ASE | 1 |
| 2022 | Studying logging practice in test code
Haonan Zhang 0006, Yiming Tang 0002, Maxime Lamothe, Heng Li 0007, Weiyi Shang |
Empir. Softw. Eng. | 1 |