VLDB 2026 Research / reviewers in the wild / expert
Zebao Gao
dblp:129/2283
· DBLP profile ↗
7ranked-venue papers
4as first author
0since 2021 · last 2016
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 7 · 4 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
3 papers |
Software testing · 87% Program analysis · 13% |
Topics — the 7 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Software testing
test repair |
0.2 | 1 | 2016 | SITAR: GUI Test Script Repair · IEEE Trans. Software Eng. 2016 |
Program analysis › specification mining
invariant detection |
0.2 | 1 | 2015 | Making System User Interactive Tests Repeatable: When and What Should We Control? · ICSE (1) 2015 |
Software testing
system testing |
0.2 | 1 | 2015 | Making System User Interactive Tests Repeatable: When and What Should We Control? · ICSE (1) 2015 |
Software testing
test oracle |
0.2 | 1 | 2015 | Making System User Interactive Tests Repeatable: When and What Should We Control? · ICSE (1) 2015 |
Software testing › test adequacy
coverage criteria |
0.2 | 1 | 2014 | Virtual DOM coverage for effective testing of dynamic web applications · ISSTA 2014 |
Software testing › test adequacy
test adequacy criteria |
0.2 | 1 | 2014 | Virtual DOM coverage for effective testing of dynamic web applications · ISSTA 2014 |
Software testing
web application testing |
0.2 | 1 | 2014 | Virtual DOM coverage for effective testing of dynamic web applications · ISSTA 2014 |
Methods — techniques the papers use, named apart from their topics
test script synthesis · 0.2reverse engineering · 0.2event-flow graph · 0.2empirical study · 0.2controlled experimentation · 0.2code coverage analysis · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2016 | Making System User Interactive Tests Repeatable: When and What Should we Control?abstractSystem user interactive tests are widely used to evaluate the behavior of an application as a whole. To automate this process, many techniques are proposed whose effectiveness are evaluated by metrics such as code coverage and fault detection. However, most of previous work assumes determinism in the outputs of interactive tests. In this paper, we propose three layers of testing outputs to examine: the code layer (codecoverage), the behavioral layer (invariant detection) and the user interaction layer (fault detection with GUI oracle). We further study the impact of common set of factors such as operating system, Java version, initial starting state and time delay on these metrics. A comprehensive experiment has been conducted on Java Swing applications, and the results show that as many as184 lines can be covered differently and up to 96% false positives with respect to fault detection. We plan to study the repeatability of interactive tests on the Android platform. Zebao Gao |
ICST | 1 |
| 2016 | SITAR: GUI Test Script RepairabstractSystem testing of a GUI-based application requires that test cases, consisting of sequences of user actions/events, be executed and the software's output be verified. To enable automated re-testing, such test cases are increasingly being coded as low-level test scripts, to be replayed automatically using test harnesses. Whenever the GUI changes—widgets get moved around, windows get merged—some scripts become unusable because they no longer encode valid input sequences. Moreover, because the software's output may have changed, theirtest oracles—assertions and checkpoints—encoded in the scripts may no longer correctly check the intended GUI objects. We presentScrIpT repAireR(SITAR), a technique to automaticallyrepairunusable low-level test scripts.SITARuses reverse engineering techniques to create an abstract test for each script, maps it to an annotated event-flow graph (EFG), uses repairing transformations and human input to repair the test, and synthesizes a new “repaired” test script. During this process,SITARalso repairs the reference to the GUI objects used in the checkpoints yielding a final test script that can be executed automatically to validate the revised software.SITARamortizes the cost of human intervention across multiple scripts by accumulating the human knowledge as annotations on the EFG. An experiment using QTP test scripts suggests thatSITARis effective in that 41-89 percent unusable test scripts were repaired. Annotations significantly reduced human cost after 20 percent test scripts had been repaired. Zebao Gao, Zhenyu Chen 0001, Yunxiao Zou, Atif M. Memon |
IEEE Trans. Software Eng. | 1 |
| 2015 | Making System User Interactive Tests Repeatable: When and What Should We Control?abstractSystem testing and invariant detection is usually conducted from the user interface perspective when the goal is to evaluate the behavior of an application as a whole. A large number of tools and techniques have been developed to generate and automate this process, many of which have been evaluated in the literature or internally within companies. Typical metrics for determining effectiveness of these techniques include code coverage and fault detection, however, with the assumption that there is determinism in the resulting outputs. In this paper we examine the extent to which a common set of factors such as the system platform, Java version, application starting state and tool harness configurations impact these metrics. We examine three layers of testing outputs: the code layer, the behavioral (or invariant) layer and the external (or user interaction) layer. In a study using five open source applications across three operating system platforms, manipulating several factors, we observe as many as 184 lines of code coverage difference between runs using the same test cases, and up to 96 percent false positives with respect to fault detection. We also see some a small variation among the invariants inferred. Despite our best efforts, we can reduce, but not completely eliminate all possible variation in the output. We use our findings to provide a set of best practices that should lead to better consistency and smaller differences in test outcomes, allowing more repeatable and reliable testing and experimentation. Zebao Gao, Yalan Liang, Myra B. Cohen, Atif M. Memon |
ICSE (1) | 1 |
| 2015 | Conceptualization and Evaluation of Component-Based Testing Unified with Visual GUI Testing: An Empirical StudyabstractIn this paper we present the results of a two-phase empirical study where we evaluate and compare the applicability of automated component-based Graphical User Interface (GUI) testing and Visual GUI Testing (VGT) in the tools GUITAR and a prototype tool we refer to as VGT GUITAR. First, GUI mutation operators are defined to create 18 faulty versions of an application on which both tools are then applied in an experiment. Results from 456 test case executions in each tool show, with statistical significance, that the component-based approach reports more false negatives than VGT for acceptance tests but that the VGT approach reports more false positives for system tests. Second, a case study is performed with larger open source applications, ranging from 8,803-55,006 lines of code. Results show that GUITAR is applicable in practice but has some challenges related to GUI component states. The results also show that VGT GUITAR is currently not applicable in practice and therefore requires further research and development. Based on the study's results we present areas of future work for both test approaches and conclude that the approaches have different benefits and drawbacks. The component-based approach is robust and executes tests faster than the VGT approach, with a factor of 3. However, the VGT approach can perform visual assertions and is perceived more flexible than the component- based approach. These conclusions let us hypothesize that a combination of the two approaches is the most suitable in practice and therefore warrants future research. Emil Alégroth, Zebao Gao, Rafael Alves Paes de Oliveira, Atif M. Memon |
ICST | 2 |
| 2015 | Pushing the limits on automation in GUI regression testingabstractAlthough there has been much work on automated GUI regression testing of software, full automation continues to etude us. There are two significant impediments to full automation: obtaining (J) test inputs and (2) test oracle. We now push the envelope 011 full automation of GUI regression testing by fully automatically generating test cases as well as the test oracle, completely eliminating manual work. This allows us to study issues of false positives/negatives in test failure; we provide ways to minimize these. The results of our empirical studies suggest that our approach of using workflow-based test eases, derived front the software under test, may help empower the end user to perform regression testing before applying software updates. Zebao Gao, Chunrong Fang, Atif M. Memon |
ISSRE | 1 |
| 2014 | Virtual DOM coverage for effective testing of dynamic web applicationsabstractTest adequacy criteria are fundamental in software testing. Among them, code coverage criterion is widely used due to its simplicity and effectiveness. However, in dynamic web application testing, merely covering server-side script code is inadequate because it neglects client-side execution, which plays an important role in triggering client-server interactions to reach important execution states. Similarly, a criterion aiming at covering the UI elements on client-side pages ignores the server-side execution, leading to insufficiency. Yunxiao Zou, Zhenyu Chen 0001, Yunhui Zheng, Xiangyu Zhang 0001, Zebao Gao |
ISSTA | 5 |
| 2014 | GUI testing assisted by human knowledge: Random vs. functional
Weiran Yang, Zhenyu Chen 0001, Zebao Gao, Yunxiao Zou |
J. Syst. Softw. | 3 |