EDBT 2026 Demo / reviewers in the wild / expert
Wing Lam
dblp:25/1049
· DBLP profile ↗
33ranked-venue papers
15as first author
12since 2021 · last 2025
0000-0003-2243-1218ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 30 · 13 first-author · 12 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Ranking Relevant Tests for Order-Dependent Flaky TestsabstractOne major challenge of regression testing are flaky tests, i.e., tests that may pass in one run but fail in another run for the same version of code. One prominent category of flaky tests is order-dependent (OD) flaky tests, which can pass or fail depending on the order in which the tests are run. To help developers debug and fix OD tests, prior work attempts to automatically find OD-relevant tests, which are tests that determine whether an OD test passes or fails, depending on whether the OD-relevant tests run before or after the OD test. Prior work found OD-relevant tests by running different tests before the OD test, without considering each test's likelihood of being OD-relevant tests. We propose RankF to rank tests in order of likelihood of being OD-relevant tests, finding the first OD-relevant test for a given OD test more quickly. We propose two ranking approaches, each requiring different information. Our first approach,$\boldsymbol{RankF}_L$, relies on training a large-language model to analyze test code. Our second approach,$\boldsymbol{RankF}_O$, relies on analyzing prior test-order execution information. We evaluate our approaches on 155 OD tests across 24 open-source projects. We compare RankF against baselines from prior work, where we find that RankF finds the first OD-relevant test for an OD test faster than the best baseline; depending on the type of OD-relevant test, RankF takes 9.4 to 14.1 seconds on median, compared to the baseline's 34.2 to 118.5 seconds on median. Shanto Rahman, Bala Naren Chanumolu, Suzzana Rafi, August Shi, Wing Lam |
ICSE | 5 |
| 2024 | Test Scheduling Across Heterogeneous Machines While Balancing Running Time, Price, and FlakinessabstractScheduling tests to run in parallel across different machines is an effective way to reduce overall test running time. Prior work has focused on scheduling tests across homogeneous machines, namely, machines all of the same configuration. However, using all the same configuration may not be the most cost-effective way to reduce test running time. We propose scheduling tests across machines with different configurations, namely heterogeneous machines. Doing so allows us to balance various factors, e.g., price, as tests may have similar running times on different machine configurations but result in drastically different monetary prices. Furthermore, there can be flaky tests that fail more often on different machine configurations, so scheduling them across heterogeneous machines gives better control over their flaky-failure rates. Our approach, GASearch, leverages genetic algorithms and a fitness function to balance running time and price to efficiently generate a heterogeneous machine configuration on which to run tests. We also model flaky-failure rate of tests on different machines within the fitness function as a factor of running time, where a failing flaky test would be rerun until it passes (or to a maximum number of runs) to confirm if it is a flaky failure, so we can balance all factors at once. We evaluate our approach on test suites from 24 modules in open-source Maven projects. Compared against baselines that schedule across homogeneous machines, we find that scheduling across heterogeneous ones can achieve a lower running time and price. Hengchen Yuan, Jiefang Lin, Wing Lam, August Shi |
ICSME | 3 |
| 2024 | Aurora: Navigating UI Tarpits via Automated Neural Screen UnderstandingabstractNearly a decade of research in software engineering has focused on automating mobile app testing to help engineers in overcoming the unique challenges associated with the software platform. Much of this work has come in the form of Automated Input Generation tools (AIG tools) that dynamically explore app screens. However, such tools have repeatedly been demonstrated to achieve lower-than-expected code coverage - particularly on sophisticated proprietary apps. Prior work has illustrated that a primary cause of these coverage deficiencies is related to so-called tarpits, or complex screens that are difficult to navigate. In this paper, we take a critical step toward enabling AIG tools to effectively navigate tarpits during app exploration through a new form of automated semantic screen understanding. That is, we introduce Aurora,a technique that learns from the visual and textual patterns that exist in mobile app UIs to automatically detect common screen designs and navigate them accordingly. The key idea of Aurorais that there are a finite number of mobile app screen designs, albeit with subtle variations, such that the general patterns of different categories of UI designs can be learned. As such, Auroraemploys a multi-modal, neural screen classifier that is able to recognize the most common types of UI screen designs. After recognizing a given screen, it then applies a set of flexible and generalizable heuristics to properly navigate the screen. We evaluated Auroraboth on a set of 12 apps with known tarpits from prior work, and on a new set of five of the most popular apps from the Google Play store. Our results indicate that Aurorais able to effectively navigate tarpit screens, outperforming prior approaches that avoid tarpits by 19.6% in terms of method coverage. Our analysis of the results finds that the improvements can be attributed to AURORA's VI design classification and heuristic navigation techniques. Safwat Ali Khan, Yiran Ren, Jiangfan Shi, Alyssa McGowan, Wing Lam, Kevin Moran |
ICST | 7 |
| 2024 | Automatically Reproducing Timing-Dependent Flaky-Test FailuresabstractWhen developers run tests after making code changes, they may encounter test failures from flaky tests, which are tests that can non-deterministically pass or fail on the same version of code. Prior work has found “timing dependence” to be a top cause of this non-determinism, i.e., tests may pass or fail depending on the timing of asynchronous callbacks or different thread interleavings that can occur when thread executions run faster or slower relative to others. Similar to how one debugs and fixes normal test failures, developers need to be able to reliably reproduce flaky-test failures. However, many of these failures can be extremely unlikely to occur (e.g., failing only once out of 10,000 runs in prior work), making it costly for developers to reproduce the failures. We present FlakeRake, an automated approach for reproducing timing-dependent (TD) flaky-test failures by inserting well-placed sleep calls, which temporarily pauses one thread or task and allows another to overtake it. When applied to an existing dataset of known flaky-test failures, FlakeRake is able to reproduce the exact same failure at least once for 136 failures, whereas simply rerunning each test 10,000 times reproduces only 115 failures or rerunning the entire test suites 10,000 times reproduces only 127 failures. For each failure that can be reproduced, we find that FlakeRake can reliably reproduce (>50% of the time) 107 failures, while rerunning just the flaky test or the entire test suite could not reliably reproduce any failure. We also find that if a developer needs to reproduce a failure six or more times, using FlakeRake (including the one-time cost to search for sleep calls) takes less time to reproduce that many failures than continually rerunning just the flaky test. Lastly, we inspect the sleep locations that FlakeRake outputs and provide insights for how one should cope with TD flaky tests. Shanto Rahman, Aaron Massey, Wing Lam, August Shi, Jonathan Bell 0001 |
ICST | 3 |
| 2024 | Hierarchy-Aware Regression Test PrioritizationabstractRegression testing is widely used to check whether software changes lead to test failures. Regression Test Prioriti-zation (RTP) aims to order tests such that tests that are more likely to fail are run earlier. Prior RTP techniques—which we call hierarchy-unaware (HU)—ignored an important aspect: real test suites are organized hierarchically, and individual tests belong to composites that can be hierarchically nested. Prior RTP work overlooked the runtime cost to switch across hierarchical test compositesand used the APFDcmetric, which represents the runtime of tests till test failures, to rank orders generated by RTP techniques. However, APFDccan misleadingly rank orders if their runtimes differ (e.g., two orders may have different numbers of composite switches and, consequently, runtimes). To account for runtime differences, we propose a new metric, HAPFDc. Unlike APFDc, HAPFDcenables proper comparison of test orders with different runtimes by "extending" runtimes as needed. To reduce the cost of composite switching, we introduce hierarchy-aware (HA) RTP by presenting meta-techniques that first prioritize composites and then tests within composites. We evaluate HA RTP on test classes in multi-module Java and Maven projects from two large datasets used in prior work. The results show that our HA RTP improves both HAPFDcvalues and time-based metrics over HU RTP. Hao Wang 0112, Pu Yi 0001, Jeremias Parladorio, Wing Lam, Darko Marinov, Tao Xie 0001 |
ISSRE | 4 |
| 2024 | The Effects of Computational Resources on Flaky TestsabstractFlaky tests are tests that non-deterministically pass and fail in unchanged code. These tests can be detrimental to developers’ productivity. Particularly when tests run in continuous integration environments, the tests may be competing for access to limited computational resources (CPUs, memory etc.), and we hypothesize that resource (un)-availability may be a significant factor in the failure rate of flaky tests. We present the first assessment of the impact that computational resources have on flaky tests, including a total of 52 projects written in Java, JavaScript and Python, and 27 different resource configurations. Using a rigorous statistical methodology, we determine which tests are RAFTs (Resource-Affected Flaky Tests). We find that 46.5% of the flaky tests in our dataset are RAFTs, indicating that a substantial proportion of flaky-test failures happen depending on the resources available when running tests. We report RAFTs and configurations to avoid them to developers, and received interest to either fix the RAFTs or to improve the specifications of the projects so that tests would be run only in configurations that are unlikely to encounter RAFT failures. Although most test suites in our dataset are executed quite quickly (under one minute) in a baseline configuration, our results highlight the possibility of using this methodology to detect RAFT to reduce the cost of cloud infrastructure for reliably running larger test suites. Denini Silva, Martin Gruber, Satyajit Gokhale, Ellen Arteca, Alexi Turcotte, Marcelo d'Amorim, Wing Lam, Stefan Winter 0001, Jonathan Bell 0001 |
IEEE Trans. Software Eng. | 7 |
| 2023 | Systematically Producing Test Orders to Detect Order-Dependent Flaky TestsabstractSoftware testing suffers from the presence of flaky tests, which can pass or fail when run on the same version of code. Order- dependent tests (OD tests) are flaky tests whose outcome depends on the order in which they are run. An OD test can be detected if specific tests are run or not run before it, resulting in a difference in test outcome. While prior work has proposed rerunning tests in different random test orders, this approach does not provide guarantees toward detecting all OD tests. Later work that proposed a more systematic approach to ordering tests still fails to account for the relationships between all tests in the test suite. We propose three new techniques to detect OD tests through a more systematic means of producing test orders. Our techniques build upon prior work in Tuscan squares to cover test pairs in a minimal set of test orders while also obeying the constraints of how tests can be positioned in a test order w.r.t. their test classes. Further, as there are many test pairs that need to be covered, we develop a technique that can take a specified set of test pairs to cover and produce test orders that aim to cover just those test pairs. Our evaluation with 289 known OD tests across 47 test suites from open-source projects shows that our most cost-effective technique can detect 97.2% of the known OD tests with 104.7 test orders, on average, per subject. While all techniques produce a relatively large number of test orders, our analysis of the minimal set of test orders needed to detect OD tests shows a tremendous reduction in the test orders needed to detect OD tests – representing an opportunity for future work to prioritize test orders. Chengpeng Li 0002, Mohammad Mahdi Khosravi, Wing Lam, August Shi |
ISSTA | 3 |
| 2023 | Optimizing Continuous Development by Detecting and Preventing Unnecessary Content GenerationabstractContinuous development (CD) helps developers quickly release and update their software. To enact CD, developers customize their CD builds to perform several tasks, including compiling, testing, static analysis checks, etc. However, as developers add more tasks to their builds, the builds take longer to run, therefore slowing down the entire CD process. Furthermore, developers may unknowingly include tasks into their builds whose results are not used (e.g., generating coverage files that are never read or uploaded anywhere), therefore wasting build runtime doing unnecessary tasks. We propose OptCD, a technique to dynamically detect unnecessary work within CD builds. Our intuition is that unnecessary work can be identified by the generation of files that are not used by any other task within the build. OptCD runs alongside a CD build, tracking the generated files during the build and which files are read/written. Files that are written to but are never read from are unnecessary content from a build. Based on the names of the unnecessary files, OptCD then maps the files to the specific build tasks responsible for generating or writing to those files. Finally, OptCD leverages ChatGPT to suggest changing the build configuration to disable generating these unnecessary files. Our evaluation of OptCD on 22 open-source projects finds that 95.6% of projects generate at least one unused directory, a directory whose contents are all unnecessarily generated. OptCD identifies the correct task that generates 92.0% of the unused directories. Further, OptCD can produce a patch for the CD configuration file to prevent generating 72.0% of the unused directories. Using the patches, we reduce the runtime by 7.0% on average for the projects we studied. We submitted 26 pull requests for the unused directories that we could disable. Developers have accepted 12 of them, with five rejected, and nine still pending. Talank Baral, Shanto Rahman, Bala Naren Chanumolu, Basak Balci, Tuna Tuncer, August Shi, Wing Lam |
ASE | 7 |
| 2022 | Preempting Flaky Tests via Non-Idempotent-Outcome TestsabstractRegression testing can greatly help in software development, but it can be seriously undermined by flaky tests, which can both pass and fail, seemingly nondeterministically, on the same code commit. Flaky tests are an emerging topic in both research and industry. Prior work has identified multiple categories of flaky tests, developed techniques for detecting these flaky tests, and analyzed some detected flaky tests. Anjiang Wei, Pu Yi 0001, Zhengxi Li, Tao Xie 0001, Darko Marinov, Wing Lam |
ICSE | 6 |
| 2022 | A Theoretical Analysis of Random Regression Test PrioritizationabstractAbstract Regression testing is an important activity to check software changes by running the tests in a test suite to inform the developers whether the changes lead to test failures. Regression test prioritization (RTP) aims to inform the developers faster by ordering the test suite so that tests likely to fail are run earlier. Many RTP techniques have been proposed and are often compared with the random RTP baseline by sampling some of the n! different test-suite orders for a test suite with n tests. However, there is no theoretical analysis of random RTP. We present such an analysis, deriving probability mass functions and expected values for metrics and scenarios commonly used in RTP research. Using our analysis, we revisit some of the most highly cited RTP papers and find that some presented results may be due to insufficient sampling. Future RTP research can leverage our analysis and need not use random sampling but can use our simple formulas or algorithms to more precisely compare with random RTP. Pu Yi 0001, Hao Wang 0112, Tao Xie 0001, Darko Marinov, Wing Lam |
TACAS (2) | 5 |
| 2021 | An infrastructure approach to improving effectiveness of Android UI testing toolsabstractDue to the importance of Android app quality assurance, many Android UI testing tools have been developed by researchers over the years. However, recent studies show that these tools typically achieve low code coverage on popular industrial apps. In fact, given a reasonable amount of run time, most state-of-the-art tools cannot even outperform a simple tool, Monkey, on popular industrial apps with large codebases and sophisticated functionalities. Our motivating study finds that these tools perform two types of operations, UI Hierarchy Capturing (capturing information about the contents on the screen) and UI Event Execution (executing UI events, such as clicks), often inefficiently using UIAutomator, a component of the Android framework. In total, these two types of operations use on average 70% of the given test time. Wing Lam, Tao Xie 0001 |
ISSTA | 2 |
| 2021 | Probabilistic and Systematic Coverage of Consecutive Test-Method Pairs for Detecting Order-Dependent Flaky TestsabstractAbstract Software developers frequently check their code changes by running a set of tests against their code. Tests that can nondeterministically pass or fail when run on the same code version are called flaky tests. These tests are a major problem because they can mislead developers to debug their recent code changes when the failures are unrelated to these changes. One prominent category of flaky tests is order-dependent (OD) tests, which can deterministically pass or fail depending on the order in which the set of tests are run. By detecting OD tests in advance, developers can fix these tests before they change their code. Due to the high cost required to explore all possible orders (n! permutations for n tests), prior work has developed tools that randomize orders to detect OD tests. Experiments have shown that randomization can detect many OD tests, and that most OD tests depend on just one other test to fail. However, there was no analysis of the probability that randomized orders detect OD tests. In this paper, we present the first such analysis and also present a simple change for sampling random test orders to increase the probability. We finally present a novel algorithm to systematically explore all consecutive pairs of tests, guaranteeing to detect all OD tests that depend on one other test, while running substantially fewer orders and tests than simply running all test pairs. Anjiang Wei, Pu Yi 0001, Tao Xie 0001, Darko Marinov, Wing Lam |
TACAS (1) | 5 |
| 2020 | A study on the lifecycle of flaky testsabstractDuring regression testing, developers rely on the pass or fail outcomes of tests to check whether changes broke existing functionality. Thus, flaky tests, which nondeterministically pass or fail on the same code, are problematic because they provide misleading signals during regression testing. Although flaky tests are the focus of several existing studies, none of them study (1) the reoccurrence, runtimes, and time-before-fix of flaky tests, and (2) flaky tests in-depth on proprietary projects. Wing Lam, Kivanç Muslu, Hitesh Sajnani, Suresh Thummalapenta |
ICSE | 1 |
| 2020 | Understanding Reproducibility and Characteristics of Flaky Tests Through Test Reruns in Java ProjectsabstractFlaky tests are tests that can non-deterministically pass and fail. They pose a major impediment to regression testing, because they provide an inconclusive assessment on whether recent code changes contain faults or not. Prior studies of flaky tests have proposed tools to detect flaky tests and identified various sources of flakiness in tests, e.g., order-dependent (OD) tests that deterministically fail for some order of tests in a test suite but deterministically pass for some other orders. Several of these studies have focused on OD tests. We focus on an important and under-explored source of flakiness in tests: non-order-dependent tests that can nondeterministically pass and fail even for the same order of tests. Instead of using specialized tools that aim to detect flaky tests, we run tests using the tool configured by the developers. Specifically, we perform our empirical evaluation on Java projects that rely on the Maven Surefire plugin to run tests. We re-execute each test suite 4000 times, potentially in different test-class orders, and we label tests as flaky if our runs have both pass and fail outcomes across these reruns. We obtain a dataset of 107 flaky tests and study various characteristics of these tests. We find that many tests previously called “non-order-dependent” actually do depend on the order and can fail with very different failure rates for different orders. Wing Lam, Stefan Winter 0001, Angello Astorga, Victoria Stodden, Darko Marinov |
ISSRE | 1 |
| 2020 | Dependent-test-aware regression testing techniquesabstractDevelopers typically rely on regression testing techniques to ensure that their changes do not break existing functionality. Unfortunately, these techniques suffer from flaky tests, which can both pass and fail when run multiple times on the same version of code and tests. One prominent type of flaky tests is order-dependent (OD) tests, which are tests that pass when run in one order but fail when run in another order. Although OD tests may cause flaky-test failures, OD tests can help developers run their tests faster by allowing them to share resources. We propose to make regression testing techniques dependent-test-aware to reduce flaky-test failures. Wing Lam, August Shi, Reed Oei, Sai Zhang 0001, Michael D. Ernst, Tao Xie 0001 |
ISSTA | 1 |
| 2020 | A large-scale longitudinal study of flaky testsabstractFlaky tests are tests that can non-deterministically pass or fail for the same code version. These tests undermine regression testing efficiency, because developers cannot easily identify whether a test fails due to their recent changes or due to flakiness. Ideally, one would detect flaky tests right when flakiness is introduced, so that developers can then immediately remove the flakiness. Some software organizations, e.g., Mozilla and Netflix, run some tools—detectors—to detect flaky tests as soon as possible. However, detecting flaky tests is costly due to their inherent non-determinism, so even state-of-the-art detectors are often impractical to be used on all tests for each project change. To combat the high cost of applying detectors, these organizations typically run a detector solely on newly added or directly modified tests, i.e., not on unmodified tests or when other changes occur (including changes to the test suite, the code under test, and library dependencies). However, it is unclear how many flaky tests can be detected or missed by applying detectors in only these limited circumstances. To better understand this problem, we conduct a large-scale longitudinal study of flaky tests to determine when flaky tests become flaky and what changes cause them to become flaky. We apply two state-of-theart detectors to 55 Java projects, identifying a total of 245 flaky tests that can be compiled and run in the code version where each test was added. We find that 75% of flaky tests (184 out of 245) are flaky when added, indicating substantial potential value for developers to run detectors specifically on newly added tests. However, running detectors solely on newly added tests would still miss detecting 25% of flaky tests. The percentage of flaky tests that can be detected does increase to 85% when detectors are run on newly added or directly modified tests. The remaining 15% of flaky tests become flaky due to other changes and can be detected only when detectors are always applied to all tests. Our study is the first to empirically evaluate when tests become flaky and to recommend guidelines for applying detectors in the future. Wing Lam, Stefan Winter 0001, Anjiang Wei, Tao Xie 0001, Darko Marinov, Jonathan Bell 0001 |
Proc. ACM Program. Lang. | 1 |
| 2019 | iDFlakies: A Framework for Detecting and Partially Classifying Flaky TestsabstractRegression testing is increasingly important with the wide use of continuous integration. A desirable requirement for regression testing is that a test failure reliably indicates a problem in the code under test and not a false alarm from the test code or the testing infrastructure. However, some test failures are unreliable, stemming from flaky tests that can nondeterministically pass or fail for the same code under test. There are many types of flaky tests, with order-dependent tests being a prominent type. To help advance research on flaky tests, we present (1) a framework, iDFlakies, to detect and partially classify flaky tests; (2) a dataset of flaky tests in open-source projects; and (3) a study with our dataset. iDFlakies automates experimentation with our tool for Maven-based Java projects. Using iDFlakies, we build a dataset of 422 flaky tests, with 50.5% order-dependent and 49.5% not. Our study of these flaky tests finds the prevalence of two types of flaky tests, probability of a test-suite run to have at least one failure due to flaky tests, and how different test reorderings affect the number of detected flaky tests. We envision that our work can spur research to alleviate the problem of flaky tests. Wing Lam, Reed Oei, August Shi, Darko Marinov, Tao Xie 0001 |
ICST | 1 |
| 2019 | Root causing flaky tests in a large-scale industrial settingabstractIn today’s agile world, developers often rely on continuous integration pipelines to help build and validate their changes by executing tests in an efficient manner. One of the significant factors that hinder developers’ productivity is flaky tests—tests that may pass and fail with the same version of code. Since flaky test failures are not deterministically reproducible, developers often have to spend hours only to discover that the occasional failures have nothing to do with their changes. However, ignoring failures of flaky tests can be dangerous, since those failures may represent real faults in the production code. Furthermore, identifying the root cause of flakiness is tedious and cumbersome, since they are often a consequence of unexpected and non-deterministic behavior due to various factors, such as concurrency and external dependencies. Wing Lam, Patrice Godefroid, Suman Nath, Anirudh Santhiar, Suresh Thummalapenta |
ISSTA | 1 |
| 2019 | Neural detection of semantic code clones via tree-based convolutionabstractCode clones are similar code fragments that share the same semantics but may differ syntactically to various degrees. Detecting code clones helps reduce the cost of software maintenance and prevent faults. Various approaches of detecting code clones have been proposed over the last two decades, but few of them can detect semantic clones, i.e., code clones with dissimilar syntax. Recent research has attempted to adopt deep learning for detecting code clones, such as using tree-based LSTM over Abstract Syntax Tree (AST). However, it does not fully leverage the structural information of code fragments, thereby limiting its clone-detection capability. To fully unleash the power of deep learning for detecting code clones, we propose a new approach that uses tree-based convolution to detect semantic clones, by capturing both the structural information of a code fragment from its AST and lexical information from code tokens. Additionally, our approach addresses the limitation that source code has an unlimited vocabulary of tokens and models, and thus exploiting lexical information from code tokens is often ineffective when dealing with unseen tokens. Particularly, we propose a new embedding technique called position-aware character embedding (PACE), which essentially treats any token as a position-weighted combination of character one-hot embeddings. Our experimental results show that our approach substantially outperforms an existing state-of-the-art approach with an increase of 0.42 and 0.15 in F1-score on two popular code-clone benchmarks (OJClone and BigCloneBench), respectively, while being more computationally efficient. Our experimental results also show that PACE enables our approach to be substantially more effective when code clones contain unseen tokens. Hao Yu 0016, Wing Lam, Ge Li 0001, Tao Xie 0001, Qianxiang Wang |
ICPC | 2 |
| 2019 | iFixFlakies: a framework for automatically fixing order-dependent flaky testsabstractRegression testing provides important pass or fail signals that developers use to make decisions after code changes. However, flaky tests, which pass or fail even when the code has not changed, can mislead developers. A common kind of flaky tests are order-dependent tests, which pass or fail depending on the order in which the tests are run. Fixing order-dependent tests is often tedious and time-consuming. August Shi, Wing Lam, Reed Oei, Tao Xie 0001, Darko Marinov |
ESEC/SIGSOFT FSE | 2 |
| 2018 | A Characteristic Study of Parameterized Unit Tests in .NET Open Source Projects
Wing Lam, Siwakorn Srisakaokul, Blake Bassett, Peyman Mahdian, Tao Xie 0001, Pratap Lakshman, Jonathan de Halleux |
ECOOP | 1 |
| 2018 | Bugs.jar: a large-scale, diverse dataset of real-world Java bugsabstractWe present Bugs.jar, a large-scale dataset for research in automated debugging, patching, and testing of Java programs. Bugs.jar is comprised of 1,158 bugs and patches, drawn from 8 large, popular open-source Java projects, spanning 8 diverse and prominent application categories. It is an order of magnitude larger than Defects4J, the only other dataset in its class. We discuss the methodology used for constructing Bugs.jar, the representation of the dataset, several use-cases, and an illustration of three of the use-cases through the application of 3 specific tools on Bugs.jar, namely our own tool, Elixir, and two third-party tools, Ekstazi and JaCoCo. Ripon K. Saha, Yingjun Lyu, Wing Lam, Hiroaki Yoshida, Mukul R. Prasad |
MSR | 3 |
| 2017 | Record and replay for Android: are we there yet in industrial cases?abstractMobile applications, or apps for short, are gaining popularity. The input sources (e.g., touchscreen, sensors, transmitters) of the smart devices that host these apps enable the apps to offer a rich experience to the users, but these input sources pose testing complications to the developers (e.g., writing tests to accurately utilize multiple input sources together and be able to replay such tests at a later time). To alleviate these complications, researchers and practitioners in recent years have developed a variety of record-and-replay tools to support the testing expressiveness of smart devices. These tools allow developers to easily record and automate the replay of complicated usage scenarios of their app. Due to Android's large share of the smart-device market, numerous record-and-replay tools have been developed using a variety of techniques to test Android apps. To better understand the strengths and weaknesses of these tools, we present a comparison of popular record-and-replay tools from researchers and practitioners, by applying these tools to test three popular industrial apps downloaded from the Google Play store. Our comparison is based on three main metrics: (1) ability to reproduce common usage scenarios, (2) space overhead of traces created by the tools, and (3) robustness of traces created by the tools (when being replayed on devices with different resolutions). The results from our comparison show which record-and-replay tools may be the best for developers and identify future directions for improving these tools to better address testing complications of smart devices. Wing Lam, Zhengkai Wu, Dengfeng Li 0003, Haibing Zheng, Yuetang Deng, Tao Xie 0001 |
ESEC/SIGSOFT FSE | 1 |
| 2016 | Repairing test dependenceabstractIn a test suite, all the tests should be independent: no test should affect another test's result, and running the tests in any order should yield the same test results. The assumption of such test independence is important so that tests behave consistently as designed. However, this critical assumption often does not hold in practice due to test dependence. Wing Lam |
SIGSOFT FSE | 1 |
| 2016 | Automated test input generation for Android: are we really there yet in an industrial case?abstractGiven the ever increasing number of research tools to automatically generate inputs to test Android applications (or simply apps), researchers recently asked the question "Are we there yet?" (in terms of the practicality of the tools). By conducting an empirical study of the various tools, the researchers found that Monkey (the most widely used tool of this category in industrial practices) outperformed all of the research tools that they studied. In this paper, we present two significant extensions of that study. First, we conduct the first industrial case study of applying Monkey against WeChat, a popular messenger app with over 762 million monthly active users, and report the empirical findings on Monkey's limitations in an industrial setting. Second, we develop a new approach to address major limitations of Monkey and accomplish substantial code-coverage improvements over Monkey, along with empirical insights for future enhancements to both Monkey and our approach. Xia Zeng, Dengfeng Li 0003, Wujie Zheng, Yuetang Deng, Wing Lam, Wei Yang 0013, Tao Xie 0001 |
SIGSOFT FSE | 6 |
| 2014 | Empirically revisiting the test independence assumptionabstractIn a test suite, all the test cases should be independent: no test should affect any other test’s result, and running the tests in any order should produce the same test results. Techniques such as test prioritization generally assume that the tests in a suite are independent. Test dependence is a little-studied phenomenon. This paper presents five results related to test dependence. Sai Zhang 0001, Darioush Jalali, Jochen Wuttke, Kivanç Muslu, Wing Lam, Michael D. Ernst, David Notkin |
ISSTA | 5 |
| 2006 | Center for Army Lessons Learned: Knowledge Application Process in the MilitaryabstractThis paper is an instructional case that describes how the Center for Army Lessons Learned (CALL) has developed a unique, institutionalised knowledge application process. The paper highlights several issues related to knowledge application, including the collection, distillation, and dissemination of knowledge, the role of subject experts in the knowledge application process, and how technology facilitates knowledge application. Interested readers can contact the lead author for a list of questions and suggested answers intended for teaching the process of knowledge application to graduate students. Alton Yeow-Kuan Chua, Wing Lam |
Int. J. Knowl. Manag. | 2 |
| 2005 | Investigating success factors in enterprise application integration: a case-driven analysisabstractThis paper investigates Critical Success Factors (CSFs) in Enterprise Application Integration (EAI). An initial set of CSFs for EAI projects was created based on a review and synthesis of the literature in the general area of integration, including Enterprise Resource Planning (ERP) projects. A case analysis, involving a large financial services provider integrating its consumer banking systems, was used to validate the CSFs. The findings resulted in a more structured and holistic CSF model which identifies three broad groups of CSFs, namely (1) top management support, (2) overall integration strategy, and (3) EAI project planning and execution. Although EAI projects share many of the same CSFs as ERP projects and other information systems projects, issues related to the selection of the right EAI tool and emphasis on technology planning and enterprise architecture are distinguishing features of EAI projects. Some of practical implications of the research are that EAI projects require personnel with specific skills and expertise, business integration should precede technology integration, that availability of adapters is an important criteria in EAI tool selection and that some custom adapter development may be unavoidable if custom applications need to be integrated. Wing Lam |
Eur. J. Inf. Syst. | 1 |
| 2005 | An Enterprise Application Integration (EAI) Case-Study: Seamless Mortgage Processing at Harmond Bank
Wing Lam |
J. Comput. Inf. Syst. | 1 |
| 1998 | Change Analysis and Management in a Reuse-Oriented Software Development Setting
Wing Lam |
CAiSE | 1 |
| 1998 | Managing Requirements Change: A Set of Good Practices
Wing Lam, Venky Shankararaman, Sara Jones 0001 |
REFSQ | 1 |
| 1997 | Scenario reuse: a technique for complementing scenario-based requirements engineering approachesabstractIn recent years, the use of 'scenarios' have been recognised as an important tool in requirements engineering (RE), and has led to a number of scenario-based RE approaches being proposed (Potts et al., 1995; Jacobson 1995; Hsai et al., 1995). This paper presents an approach for reusing scenarios, called SCORE (SCenario-Oriented REuse), which is intended to complement existing scenario-based RE approaches. We describe our application of SCORE on two case studies, and discuss the issues that have surfaced from each. We also compare SCORE with other approaches to reuse at the RE level. Finally, we summarise our experience of scenario reuse by outlining the questions that need to be addressed by a more comprehensive model of the scenario reuse process. Wing Lam |
APSEC | 1 |
| 1997 | Mechanising Requirements Engineering: Reuse and the Application of Domain Analysis TechnologyabstractThe paper describes efforts that have made to mechanise the requirements engineering process in an industrial avionics domain. The authors' approach is based on an analysis of both the application domain and the task domain. The paper describes the processes they have used for domain analysis, and the tool they have developed to support mechanisation. They give an initial evaluation of the approach and close with a summary of lessons learnt. Wing Lam, Sara Jones 0001 |
ASE | 1 |