VLDB 2026 Research / reviewers in the wild / expert
Wing Kwong Chan
dblp:c/WingKwonChan · also W. K. Chan 0001, Wing-Kwong Chan
· DBLP profile ↗
147ranked-venue papers
14as first author
34since 2021 · last 2026
0000-0001-7726-6235ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 117 · 11 first-author · 20 since 2021Applied, interdisciplinary, general and emerging computing · 31 · 4 first-author · 6 since 2021Systems, architecture and hardware · 6Databases, data management, data science and information retrieval · 5 · 3 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HiCert: Toward Patch Robustness Certification and Detection for Deep Learning Systems Beyond Consistent SamplesabstractPatch robustness certification is an emerging kind of provable defense technique against adversarial patch attacks for deep learning systems. Certified detection ensures the detection of all patched harmful versions of certified samples, which mitigates the failures of empirical defense techniques that could (easily) be compromised. However, existing certified detection methods are ineffective in certifying samples that are misclassified or whose mutants are inconsistently predicted to different labels. This paper proposes HiCert, a novel masking-based certified detection technique. By focusing on the problem of mutants predicted with a label different from the true label with our formal analysis, HiCert formulates a novel formal relation between harmful samples generated by identified loopholes and their benign counterparts. By checking the bound of the maximum confidence among these potentially harmful (i.e., inconsistent) mutants of each benign sample, HiCert ensures that each harmful sample either has the minimum confidence among mutants that are predicted the same as the harmful sample itself below this bound, or has at least one mutant predicted with a label different from the harmful sample itself, formulated after two novel insights. As such, HiCert systematically certifies those inconsistent samples and consistent samples to a large extent. To our knowledge, HiCert is thefirstwork capable of providing such a comprehensive patch robustness certification for certified detection. Our experiments show the high effectiveness of HiCert with a new state-of-the-art performance: It certifies significantly more benign samples, including those inconsistent and consistent, and achieves significantly higher accuracy on those samples without warnings and a significantly lower false silent ratio. Moreover, on actual patch attacks, its defense success ratio is significantly higher than its peers. Qilin Zhou, Zhengyuan Wei, Haipeng Wang 0005, Wing Kwong Chan |
IEEE Trans. Reliab. | 5 |
| 2025 | Scalable and Precise Patch Robustness Certification for Deep Learning Models with Top- $k$ PredictionsabstractPatch robustness certification is an emerging verification approach for defending against adversarial patch attacks with provable guarantees for deep learning systems. Certified recovery techniques guarantee the prediction of the sole true label of a certified sample. However, existing techniques, if applicable to top-$k$predictions, commonly conduct pairwise comparisons on those votes between labels, failing to certify the sole true label within the top$k$prediction labels precisely due to the inflation on the number of votes controlled by the attacker (i.e., attack budget); yet enumerating all combinations of vote allocation suffers from the combinatorial explosion problem. We propose CostCert, a novel, scalable, and precise voting-based certified recovery defender. CostCert verifies the true label of a sample within the top$k$predictions without pairwise comparisons and combinatorial explosion through a novel design: whether the attack budget on the sample is infeasible to cover the smallest total additional votes on top of the votes uncontrollable by the attacker to exclude the true labels from the top$k$prediction labels. Experiments show that CostCert significantly outperforms the current state-of-the-art defender PatchGuard, such as retaining up to 57.3% in certified accuracy when the patch size is 96, whereas PatchGuard has already dropped to zero. Qilin Zhou, Haipeng Wang 0005, Zhengyuan Wei, Wing Kwong Chan |
QRS | 4 |
| 2025 | Context-Aware Fuzzing for Robustness Enhancement of Deep Learning ModelsabstractIn the testing-retraining pipeline for enhancing the robustness property of deep learning (DL) models, many state-of-the-art robustness-oriented fuzzing techniques are metric-oriented. The pipeline generates adversarial examples as test cases via such a DL testing technique and retrains the DL model under test with test suites that contain these test cases. On the one hand, the strategies of these fuzzing techniques tightly integrate the key characteristics of their testing metrics. On the other hand, they are often unaware of whether their generated test cases are different from the samples surrounding these test cases and whether there are relevant test cases of other seeds when generating the current one. We propose a novel testing metric called Contextual Confidence (CC). CC measures a test case through the surrounding samples of a test case in terms of their mean probability predicted to the prediction label of the test case. Based on this metric, we further propose a novel fuzzing technique Clover as a DL testing technique for the pipeline. In each fuzzing round, Clover first finds a set of seeds whose labels are the same as the label of the seed under fuzzing. At the same time, it locates the corresponding test case that achieves the highest CC values among the existing test cases of each seed in this set of seeds and shares the same prediction label as the existing test case of the seed under fuzzing that achieves the highest CC value. Clover computes the piece of difference between each such pair of a seed and a test case. It incrementally applies these pieces of differences to perturb the current test case of the seed under fuzzing that achieves the highest CC value and to perturb the resulting samples along the gradient to generate new test cases for the seed under fuzzing. Clover finally selects test cases among the generated test cases of all seeds as much as possible and with a preference to select test cases with higher CC values for improving model robustness. The experiments show that Clover outperforms the state-of-the-art coverage-based technique Adapt and loss-based fuzzing technique RobOT by 67%–129% and 48%–100% in terms of robustness improvement ratio, respectively, delivered through the same testing-retraining pipeline. For test case generation, in terms of numbers of unique adversarial labels and unique categories for the constructed test suites, Clover outperforms Adapt by \(2.0\times\) and \(3.5\times\) and RobOT by \(1.6\times\) and \(1.7\times\) on fuzzing clean models, and also outperforms Adapt by \(3.4\times\) and \(4.5\times\) and RobOT by \(9.8\times\) and \(11.0\times\) on fuzzing adversarially trained models, respectively. Haipeng Wang 0005, Zhengyuan Wei, Qilin Zhou, Wing Kwong Chan |
ACM Trans. Softw. Eng. Methodol. | 4 |
| 2024 | Enhancing Valid Test Input Generation with Distribution Awareness for Deep Neural NetworksabstractComprehensive testing is important in improving the reliability of Deep Learning (DL)-based systems. Various Test Input Generators (TIGs) have been proposed to generate misbehavior-inducing test inputs. However, the lack of validity checking in TIGs often results in the generation of invalid inputs (i.e., out of the learned distribution), leading to unreliable testing. To save the effort of manually checking the validity and improve test efficiency, it is important to assess the effectiveness and reliability of automated validators. In this study, we comprehensively assess four automated Input Validators (IV s), Our findings show that the accuracy of IVs ranges from 49% to 77%. Distance-based IVs generally outperform reconstruction-based and density-based IVs for both classification and regression tasks. Based on the findings, we enhance existing testing frameworks by incorporating distribution awareness through joint optimization. The results demonstrate our framework leads to a 2 % to 10% increase in the number of valid inputs, which establishes our method as an effective technique for valid test input generation. Jacky W. Keung, Yan Xiao 0002, Yishu Li, Wing Kwong Chan |
COMPSAC | 7 |
| 2024 | Toward AI-facilitated Learning Cycle in Integration Course Through Pair Programming with AI AgentsabstractWe propose a new methodology that harnesses recent advancements in AI techniques to formulate an AI-facilitating code learning cycle for students. The approach builds on an existing learning process and innovatively incorporates pair programming into the learning cycle. It first transforms the example code into scaffold code as exercises through an instructor-AI pairing. The scaffold code serves as an exercise for students to complete and debug on a hardware platform iteratively with an expert AI assistant. This design alleviates instructors' burden of crafting new exercises for new scenarios and offers students the advantage of interactive learning with scenario diversity. We evaluate the methodology using a suite of example codes and assess the semantic similarity among different code versions produced by AI assistants. The case study shows promising results of the methodology. We further discuss our findings and outline future work for the proposed methodology. Zhengyuan Wei, Albert T. L. Lee, Victor C. S. Lee, Wing Kwong Chan |
CSEE&T | 4 |
| 2024 | Improving Interpretable Embeddings for Ad-hoc Video Search with Generative Captions and Multi-word Concept BankabstractAligning a user query and video clips in cross-modal latent space and that with semantic concepts are two mainstream approaches for ad-hoc video search (AVS). However, the effectiveness of existing approaches is bottlenecked by the small sizes of available video-text datasets and the low quality of concept banks, which results in the failures of unseen queries and the out-of-vocabulary problem. This paper addresses these two problems by constructing a new dataset and developing a multi-word concept bank. Specifically, capitalizing on a generative model, we construct a new dataset consisting of 7 million generated text and video pairs for pre-training. To tackle the out-of-vocabulary problem, we develop a multi-word concept bank based on syntax analysis to enhance the capability of a state-of-the- art interpretable AVS method in modelling relationships between query words. We also study the impact of current advanced features on the method. Experimental results show that the integration of the above-proposed elements doubles the R@1 performance of the AVS method on the MSRVTT dataset and improves the xinfAP on the TRECVid AVS query sets for 2016-2023 (eight years) by a margin from 2% to 77%, with an average about 20%. The code and model are available at https://github.com/nikkiwoo-gh/Improved-ITV. Jiaxin Wu 0001, Chong-Wah Ngo, Wing Kwong Chan |
ICMR | 3 |
| 2024 | FoodMask: Real-time food instance counting, segmentation and recognition
Huu-Thanh Nguyen 0003, Chong-Wah Ngo, Wing Kwong Chan |
Pattern Recognit. | 4 |
| 2024 | Identifying metamorphic relations: A data mutation directed approachabstractSummary Metamorphic testing (MT) is an effective technique to alleviate the test oracle problem. The principle of MT is to detect failures by checking whether some necessary properties, commonly known as metamorphic relations (MRs), of software under test (SUT) hold among multiple executions of source and follow‐up test cases. Since both the generation of follow‐up test cases and test result verification depend on MRs, the identification of MRs plays a key role in MT, which is an important yet difficult task requiring deep domain knowledge of the SUT. Accordingly, techniques that can direct a tester to identify MRs effectively are desirable. In this paper, we propose MT, a data mutation directed approach to identifying MRs. MT guides a tester to identify MRs by providing a set of data mutation operators and template‐style mapping rules, which not only alleviates the difficulties faced in the process of MR identification but also improves the identification effectiveness. We have further developed a tool to implement the proposed approach and conducted an empirical study to evaluate the MR identification effectiveness of MT and the performance of MRs identified by MT with respect to fault detection capability and statement coverage. The empirical results show that MT is able to identify MRs for numeric programs effectively, and the identified MRs have high fault detection capability and statement coverage. The work presented in this paper advances the field of MT by providing a simple yet practical approach to the MR identification problem. Chang-Ai Sun, An Fu, Zuoyi Wang, Wing Kwong Chan |
Softw. Pract. Exp. | 6 |
| 2024 | (Un)likelihood Training for Interpretable EmbeddingabstractCross-modal representation learning has become a new normal for bridging the semantic gap between text and visual data. Learning modality agnostic representations in a continuous latent space, however, is often treated as a black-box data-driven training process. It is well known that the effectiveness of representation learning depends heavily on the quality and scale of training data. For video representation learning, having a complete set of labels that annotate the full spectrum of video content for training is highly difficult, if not impossible. These issues, black-box training and dataset bias, make representation learning practically challenging to be deployed for video understanding due to unexplainable and unpredictable results. In this article, we propose two novel training objectives, likelihood and unlikelihood functions, to unroll the semantics behind embeddings while addressing the label sparsity problem in training. The likelihood training aims to interpret semantics of embeddings beyond training labels, while the unlikelihood training leverages prior knowledge for regularization to ensure semantically coherent interpretation. With both training objectives, a new encoder-decoder network, which learns interpretable cross-modal representation, is proposed for ad-hoc video search. Extensive experiments on TRECVid and MSR-VTT datasets show that the proposed network outperforms several state-of-the-art retrieval models with a statistically significant performance margin. Jiaxin Wu 0001, Chong-Wah Ngo, Wing Kwong Chan, Zhijian Hou |
ACM Trans. Inf. Syst. | 3 |
| 2023 | CONE: An Efficient COarse-to-fiNE Alignment Framework for Long Video Temporal GroundingabstractZhijian Hou, Wanjun Zhong, Lei Ji, Difei Gao, Kun Yan, W.k. Chan, Chong-Wah Ngo, Mike Zheng Shou, Nan Duan. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Zhijian Hou, Wanjun Zhong, Lei Ji 0001, Difei Gao, Kun Yan 0004, Wing Kwong Chan, Chong-Wah Ngo, Zheng Shou 0001, Nan Duan 0001 |
ACL (1) | 6 |
| 2023 | Toward AI-assisted Exercise Creation for First Course in Programming through Adversarial Examples of AI ModelsabstractWe propose a new methodology, the Exercise Creation Methodology (ECM), that leverages recent AI technology advancements to create ChatGPT-assisted programming exercises for beginners. ECM takes an existing exercise as input and mutates it by removing some contents into semantically equivalent but syntactically different versions. The pair of versions are labeled as answered correctly and misleadingly by ChatGPT. The removed contents are re-inserted incrementally with further mutation, ensuring the labels remain unchanged. Using the version with the misleading answer and the ChatGPT elaboration on the other version, we construct a ChatGPT-assisted exercise. The latter version may also serve as a solution. We illustrate ECM using a case study. Wing Kwong Chan, Y. T. Yu, Jacky W. Keung, Victor C. S. Lee |
CSEE&T | 1 |
| 2023 | A Majority Invariant Approach to Patch Robustness Certification for Deep Learning ModelsabstractPatch robustness certification ensures no patch within a given bound on a sample can manipulate a deep learning model to predict a different label. However, existing techniques cannot certify samples that cannot meet their strict bars at the classifier level or the patch region level. This paper proposes MajorCert. MajorCert firstly finds all possible label sets manipulatable by the same patch region on the same sample across the underlying classifiers, then enumerates their combinations element-wise, and finally checks whether the majority invariant of all these combinations is intact to certify samples. Qilin Zhou, Zhengyuan Wei, Haipeng Wang 0005, Wing Kwong Chan |
ASE | 4 |
| 2023 | A study on the impact of pre-trained model on Just-In-Time defect predictionabstractPrevious researchers conducting Just-In-Time (JIT) defect prediction tasks have primarily focused on the performance of individual pre-trained models, without exploring the relationship between different pre-trained models as backbones. In this study, we build six models: RoBERTaJIT, CodeBERTJIT, BARTJIT, PLBARTJIT, GPT2JIT, and CodeGPTJIT, each with a distinct pre-trained model as its backbone. We systematically explore the differences and connections between these models. Specifically, we investigate the performance of the models when using Commit code and Commit message as inputs, as well as the relationship between training efficiency and model distribution among these six models. Additionally, we conduct an ablation experiment to explore the sensitivity of each model to inputs. Furthermore, we investigate how the models perform in zero-shot and few-shot scenarios. Our findings indicate that each model based on different backbones shows improvements, and when the backbone’s pre-training model is similar, the training resources that need to be consumed are closer. We also observe that Commit code plays a significant role in defect detection, and different pre-trained models demonstrate better defect detection ability with a balanced dataset under few-shot scenarios. These results provide new insights for optimizing JIT defect prediction tasks using pre-trained models and highlight the factors that require more attention when constructing such models. Additionally, CodeGPTJIT and GPT2JIT achieved better performance than DeepJIT and CC2Vec on the two datasets respectively under 2000 training samples. These findings emphasize the effectiveness of transformer-based pre-trained models in JIT defect prediction tasks, especially in scenarios with limited training data. Yuxiang Guo 0004, Xiaopeng Gao, Zhenyu Zhang 0004, Wing Kwong Chan, Bo Jiang 0001 |
QRS | 4 |
| 2023 | Aster: Encoding Data Augmentation Relations into Seed Test Suites for Robustness Assessment and Fuzzing of Data-Augmented Deep Learning ModelsabstractData-augmented deep learning models are widely used in real-world applications. However, many state-of the-art loss-based or coverage-based fuzzing techniques fail to produce fuzzing samples for them from many seeds. This paper proposes Aster, a novel technique to address this problem to enhance their fuzzing effectiveness for deep learning models trained with multi-sample data augmentation methods. Aster formulates a novel reachability-based strategy to encode the insights of every seed’s direct and indirect data augmentation relation instances into the replacement seed of that seed systematically. Our experiment shows that Aster is highly effective. On average, loss-based and coverage-based fuzzing techniques can generate 166% and 110% more fuzzing samples and reduce 31% and 22% unsuccessful seeds, respectively, after adopting the replacement seeds generated by Aster to replace their original seeds. Their improved models also become up to 55% and 40% on average more robust against FGSM and PGD attacks in the experiment. Haipeng Wang 0005, Zhengyuan Wei, Qilin Zhou, Bo Jiang 0001, Wing Kwong Chan |
QRS | 5 |
| 2023 | OAT: An Optimized Android Testing Framework Based on Reinforcement Learning
Mengjun Du, Lian Song, Wing Kwong Chan, Bo Jiang 0001 |
TASE | 4 |
| 2023 | Introduction to the Special Issue on Software-Intensive Autonomous Systems: Methods and applicationsabstractIdentifying self-admitted technical debt (SATD) plays an important role in maintaining software stability and improving software quality. Although existing methods can detect SATD and researchers have identified design debt and requirement debt, an approach to realize multiple classification of SATD, including defect, test, and documentation, is still lacking. In this paper, we combine text generation oversampling and the Convolutional Neural Networks-Gated Recurrent Unit (CNNGRU) model, and propose an approach called SCGRU to classify multiple debt, including defect, test, documentation, design, and requirement. First, SeqGAN-based text generation is employed to generate new samples by learning the original SATD data, thereby increasing the number of SATD samples such as defect debt and reducing data imbalance. Then, we apply the CNNGRU model to refine SATD into multiple classes. An experiment with cross-project identification of 10 projects shows that our approach is more effective than existing methods such as CNN and text mining. The proposed SCGRU approach has strong advantages especially in cases of flawed debt with very unbalanced data such as test debt and documention debt. Nesrine Khabou, Ismael Bouassida Rodriguez, Khalil Drira, Paris Avgeriou, David C. Shepherd, Wing Kwong Chan, Raffaela Mirandola |
J. Syst. Softw. | 6 |
| 2023 | DeepPatch: Maintaining Deep Learning Model Programs to Retain Standard Accuracy with Substantial Robustness ImprovementabstractMaintaining a deep learning (DL) model by making the model substantially more robust through retraining with plenty of adversarial examples of non-trivial perturbation strength often reduces the model’s standard accuracy. Many existing model repair or maintenance techniques sacrifice standard accuracy to produce a large gain in robustness or vice versa. This article proposes DeepPatch, a novel technique to maintain filter-intensive DL models. To the best of our knowledge, DeepPatch is the first work to address the challenge of standard accuracy retention while substantially improving the robustness of DL models with plenty of adversarial examples of non-trivial and diverse perturbation strengths. Rather than following the conventional wisdom to generalize all the components of a DL model over the union set of clean and adversarial samples, DeepPatch formulates a novel division of labor method to adaptively activate a subset of its inserted processing units to process individual samples. Its produced model can generate the original or replacement feature maps in each forward pass of the patched model, making the patched model carry an intrinsic property of behaving like the model under maintenance on demand. The overall experimental results show that DeepPatch successfully retains the standard accuracy of all pretrained models while improving the robustness accuracy substantially. However, the models produced by the peer techniques suffer from either large standard accuracy loss or small robustness improvement compared with the models under maintenance, rendering them unsuitable in general to replace the latter. Zhengyuan Wei, Haipeng Wang 0005, Imran Ashraf 0001, Wing Kwong Chan |
ACM Trans. Softw. Eng. Methodol. | 4 |
| 2023 | Davida: A Decentralization Approach to Localizing Transaction Sequences for Debugging Transactional Atomicity ViolationsabstractAtomicity is a desirable property for multithreaded programs. In such programs, a transaction is an execution of an atomic code region that may contain memory accesses on an arbitrary number of shared variables. When transactions are not conflicting with one another in a trace, they greatly simplify the reasoning of the program correctness. If a transaction incurs an atomicity violation in a trace, developers have to debug the code, but this is challenging. To achieve practical runtime performances, existing dynamic techniques for detecting such atomicity violations face a challenge: They are designed for either detecting all such atomicity violations without the capability of localizing the corresponding cross-thread transaction sequences or deliberately missing some atomicity violations in the trade of localizing some of them to support their atomicity violation detection. In this article, we propose Davida, a novel technique to address this problem. Davida efficiently tracks selective transactions and cross-thread dependency sequences over transactions reachable from the currently active transactions of all the threads in a decentralized manner. We prove that Davida precisely accomplishes every atomicity violation in a trace with an actual sequence of transactions triggering the violation. The experimental results on 15 subjects showed that Davida outperformed Velodrome, the previous graph-based state-of-the-art technique, in both performance and completeness. Imran Ashraf 0001, Wing Kwong Chan |
IEEE Trans. Reliab. | 3 |
| 2022 | An Empirical Study on the Effects of Entry Function Pairs in Fuzzing Smart ContractsabstractEthereum smart contracts may incur security vulnerabilities. Fuzzing is an industry-standard practice to detect them in improving the dependability of programs. Existing fuzz testing techniques for Ethereum smart contracts are insensitive to whether consecutive seeds of the same function are used for fuzzing the smart contract under test. Nonetheless, smart contracts are often designed to have collaborations among different functions for business activity to complete. We wonder whether this mismatch will make fuzzing techniques less effective than they should be. In this paper, to the best of our knowledge, we present the first work to show that security vulnerability detection can be significantly more effective in smart contract fuzzing if the entry functions of recent past test cases can be distinct. The empirical results show that the performance boost can be as large as 10.4% by simply enabling any test case not invoking the same entry functions as a few recent past test cases. The empirical result also shows that the cost-effectiveness also increases by up to 21.9%. Imran Ashraf 0001, Wing Kwong Chan |
COMPSAC | 2 |
| 2022 | Cross-lingual Adaptation for Recipe Retrieval with MixupabstractCross-modal recipe retrieval has attracted research attention in recent years, thanks to the availability of large-scale paired data for training. Nevertheless, obtaining adequate recipe-image pairs covering the majority of cuisines for supervised learning is difficult if not impossible. By transferring knowledge learnt from a data-rich cuisine to a data-scarce cuisine, domain adaptation sheds light on this practical problem. Nevertheless, existing works assume recipes in source and target domains are mostly originated from the same cuisine and written in the same language. This paper studies unsupervised domain adaptation for image-to-recipe retrieval, where recipes in source and target domains are in different languages. Moreover, only recipes are available for training in the target domain. A novel recipe mixup method is proposed to learn transferable embedding features between the two domains. Specifically, recipe mixup produces mixed recipes to form an intermediate domain by discretely exchanging the section(s) between source and target recipes. To bridge the domain gap, recipe mixup loss is proposed to enforce the intermediate domain to locate in the shortest geodesic path between source and target domains in the recipe embedding space. By using Recipe 1M dataset as source domain (English) and Vireo-FoodTransfer dataset as target domain (Chinese), empirical experiments verify the effectiveness of recipe mixup for cross-lingual adaptation in the context of image-to-recipe retrieval. Bin Zhu 0006, Chong-Wah Ngo, Jingjing Chen 0001, Wing Kwong Chan |
ICMR | 4 |
| 2022 | Predictive Mutation Analysis of Test Case Prioritization for Deep Neural NetworksabstractTesting deep neural networks requires high-quality test cases, but using new test cases would incur the labor-intensive test case labeling issue in the test oracle problem. Test case prioritization for failure-revealing test cases alleviates the problem. Existing metric-based techniques analyze vector-based prediction outputs. They cannot handle regression models. Existing mutation-based techniques either remain ineffective or incur high computational costs. In this paper, we propose EffiMAP, an effective and efficient test case prioritization technique with predictive mutation analysis. In the test phase, without performing a comprehensive mutation analysis, EffiMAP predicts whether model mutants are killed by a test case by the information extracted from the execution trace of the test case. Our experiment shows that EffiMAP significantly outperforms the previous state-of-the-art technique in both effectiveness and efficiency in the test phase of handling test cases of both classification and regression models. This paper is the first work to show the feasibility of predictive mutation analysis to rank test cases with a higher probability of exposing model prediction failures in the domain of deep neural network testing. Zhengyuan Wei, Haipeng Wang 0005, Imran Ashraf 0001, Wing Kwong Chan |
QRS | 4 |
| 2022 | WasmFuzzer: A Fuzzer for WasAssembly Virtual MachinesabstractWebAssembly is a fast, safe, and portable low-level language suitable for diverse application scenarios.And The WebAssembly virtual machines are widely used by Web browsers or Blockchain platforms as execution engine.When there is a bug in the implementation of the Wasm virtual machine, the execution of WebAssembly may lead to errors or vulnerability in the application.Due to the grammar checks by WASM VMs, fuzzing at the binary level is ineffective to expose the bugs because most inputs cannot reach the deep logic within the WASM VM.In this work, we propose WasmFuzzer, a bytecode level fuzzing tool for WASM VMs.WasmFuzzer proposes to generate initial seeds for Fuzzing at the Wasm bytecode level and it also designs a systematic set of mutation operators for Wasm bytecode.Furthermore, WasmFuzzer proposes an adaptive mutation strategy to search for the best mutation operators for different fuzzing targets.Our evaluation on 3 real-life Wasm VMs shows that WasmFuzzer can significantly outperform AFL in terms of both code coverage and unique crash. Bo Jiang 0001, Zichao Li 0008, Yuhe Huang, Zhenyu Zhang 0004, Wing Kwong Chan |
SEKE | 5 |
| 2022 | SibNet: Food instance counting and segmentation
Huu-Thanh Nguyen 0003, Chong-Wah Ngo, Wing Kwong Chan |
Pattern Recognit. | 3 |
| 2022 | Learning From Web Recipe-Image Pairs for Food Recognition: Problem, Baselines and PerformanceabstractCross-modal recipe retrieval has recently been explored for food recognition and understanding. Text-rich recipe provides not only visual content information (e.g., ingredients, dish presentation) but also procedure of food preparation (cutting and cooking styles). The paired data is leveraged to train deep models to retrieve recipes for food images. Most recipes on the Web include sample pictures as the references. The paired multimedia data is not noise-free, due to errors such as pairing of images containing partially prepared dishes with recipes. The content of recipes and food images are not always consistent due to free-style writing and preparation of food in different environments. As a consequence, the effectiveness of learning cross-modal deep models from such noisy web data is questionable. This paper conducts an empirical study to provide insights whether the features learnt with noisy pair data are resilient and could capture the modality correspondence between visual and text. Bin Zhu 0006, Chong-Wah Ngo, Wing Kwong Chan |
IEEE Trans. Multim. | 3 |
| 2021 | CONQUER: Contextual Query-aware Ranking for Video Corpus Moment RetrievalabstractThis paper tackles a recently proposed Video Corpus Moment Retrieval task. This task is essential because advanced video retrieval applications should enable users to retrieve a precise moment from a large video corpus. We propose a novel CONtextual QUery-awarE Ranking~(CONQUER) model for effective moment localization and ranking. CONQUER explores query context for multi-modal fusion and representation learning in two different steps. The first step derives fusion weights for the adaptive combination of multi-modal video content. The second step performs bi-directional attention to tightly couple video and query as a single joint representation for moment localization. As query context is fully engaged in video representation learning, from feature fusion to transformation, the resulting feature is user-centered and has a larger capacity in capturing multi-modal signals specific to query. We conduct studies on two datasets, TVR for closed-world TV episodes and DiDeMo for open-world user-generated videos, to investigate the potential advantages of fusing video and query online as a joint representation for moment retrieval. Zhijian Hou, Chong-Wah Ngo, Wing Kwong Chan |
ACM Multimedia | 3 |
| 2021 | WANA: Symbolic Execution of Wasm Bytecode for Extensible Smart Contract Vulnerability DetectionabstractMany popular blockchain platforms support smart contracts for building decentralized applications. However, the vulnerabilities within smart contracts have demonstrated to lead to serious financial loss to their end users. In particular, the smart contracts on EOSIO smart contract platform have resulted in the loss of around 380K EOS tokens, which was around 1.9 million worth of USD at the time of attack. The EOSIO smart contract platform is based on the Wasm VM, which is also the underlying system supporting other smart contract platforms as well as Web application. In this work, we present WANA, an extensible smart contract vulnerability detection tool based on the symbolic execution for Wasm bytecode. WANA proposes a set of algorithms to detect the vulnerabilities in EOSIO smart contracts based on Wasm bytecode analysis. Our experimental analysis shows that WANA can effectively and efficiently detect vulnerabilities in EOSIO smart contracts. Furthermore, our case study also demonstrates that WANA can be extended to effectively detect vulnerabilities in Ethereum smart contracts. Bo Jiang 0001, Imran Ashraf 0001, Wing Kwong Chan |
QRS | 5 |
| 2021 | DroidGamer: Android Game Testing with Operable Widget Recognition by Deep LearningabstractAndroid game applications are an important type of application widely used by end users. Bugs in such applications can significantly affect user experience. Due to the use of rendered Graphical User Interface (GUI) widgets, automated testing of Android game application becomes challenging because such GUI widgets cannot be queried with Android system APIs, making existing GUI-based testing techniques blind to the locations of widgets. In this work, we propose DroidGamer, a novel GUI traversal-based Android game testing technique, which relies on deep learning models to recognize operable GUI widgets. DroidGamer adopts a novel GUI model traversal algorithm and a new GUI state equivalence criterion over the widget recognition results of the deep learning models. Our experiment on 10 open-source Android games shows that DroidGamer is significantly more effective than existing techniques including Monkey, Stoat, and PUMA for testing Android games in terms of both code coverage and fault detection ability. Bo Jiang 0001, Wenlin Wei, Wing Kwong Chan |
QRS | 4 |
| 2021 | Sound Predictive Atomicity Violation Detection§abstractMany concurrency bugs are hidden deeply behind thread interleaving and are hard to detect. Existing dynamic predictive checkers can analyze execution traces to expose predictive cases of data races and deadlocks hidden by the thread interleavings of execution traces and inferable from these inter-leavings. To the best of our knowledge, however, no existing work for the detection of atomicity violation at the transaction level (AV) can expose predictive atomicity violations. In this paper, we present the first work to address this problem. Our technique, Meteor, formulates a novel algorithm with a thread-centric pipeline to capture enumerable dependency sequences containing reversible dependencies incrementally, implicitly, and soundly. It detects predictive atomicity violations without producing false positives. We prove its soundness by theorems. We have evaluated Meteor on 19 subjects, which confirms the soundness and effectiveness of Meteor to detect predictive atomicity violations in the programs. Imran Ashraf 0001, Wing Kwong Chan |
QRS | 3 |
| 2021 | Fuzzing Deep Learning Models against Natural Robustness with Filter Coverage‡abstractCoverage-guided fuzzing on deep learning (DL) models can generate natural adversarial variants. However, no existing fuzzing work can show that covering or uncovering a coverage element of a test adequacy criterion can significantly change the accuracy of the DL model under test. This paper proposes a novel testing criterion, Filter Coverage, and a novel fuzzing technique FilterFuzz guided by this criterion. To the best of our knowledge, Filter Coverage is the first test adequacy criterion able to identify such a coverage element-blamed filter-in a DL model using convolutional filters by demonstrating a cause-and-effect chain. Both Filter Coverage and Neuron Coverage measure whether individual neurons are activated, but Filter Coverage is more selective and collective and higher in the abstraction level. FilterFuzz, guided by Filter Coverage, tackles the challenge of achieving a higher rate of generating natural adversarial variants against natural robustness. Our case study shows that FilterFuzz is significantly more effective than fuzzing guided by Neuron Coverage by 33% and produces more diverse kinds of such variants. Moreover, when applying to the problem of labeling cost reduction on generated natural variants, Filter Coverage identifies 4.3x to 4.7x of natural adversarial variants than random reordering. Our work also calls for a re-examination of the previous ineffective conclusions of Neuron Coverage. Zhengyuan Wei, Wing Kwong Chan |
QRS | 2 |
| 2021 | OPE: Transforming Programs with Clean and Precise Separation of Tested Intraprocedural Program Paths with Path ProfilingabstractExecuting program paths outside the ones tested means that the program is executing scenarios not tested before deployment. No existing technique can produce a program that precisely contains an arbitrary set of tested program paths in each procedure of a tested program. This paper presents the first work, a novel technique called OPE, to address this problem. OPE first builds a transformed procedure that contains the target set of tested paths for every procedure in a tested program. It extends the transformed procedure with additional branches and basic blocks of code to include all remaining paths of the given procedure. The resultant transformed program is functionally equivalent to the tested program. OPE achieves an inherent strict separation of the tested paths from the rest ready for deployment or follow-up program testing and analysis tasks. The experiment confirms that OPE generates programs with clean path separations and outperforms the previous state-of-the-art path encoding technique when applied to path profiling. Chunbai Yang, Imran Ashraf 0001, Hao Zhang 0085, Wing Kwong Chan |
QRS | 5 |
| 2021 | RegionTrack: A Trace-Based Sound and Complete Checker to Debug Transactional Atomicity Violations and Non-Serializable TracesabstractAtomicity is a correctness criterion to reason about isolated code regions in a multithreaded program when they are executed concurrently. However, dynamic instances of these code regions, called transactions , may fail to behave atomically, resulting in transactional atomicity violations. Existing dynamic online atomicity checkers incur either false positives or false negatives in detecting transactions experiencing transactional atomicity violations. This article proposes RegionTrack. RegionTrack tracks cross-thread dependences at the event, dynamic subregion, and transaction levels. It maintains both dynamic subregions within selected transactions and transactional happens-before relations through its novel timestamp propagation approach. We prove that RegionTrack is sound and complete in detecting both transactional atomicity violations and non-serializable traces. To the best of our knowledge, it is the first online technique that precisely captures the transitively closed set of happens-before relations over all conflicting events with respect to every running transaction for the above two kinds of issues. We have evaluated RegionTrack on 19 subjects of the DaCapo and the Java Grande Forum benchmarks. The empirical results confirm that RegionTrack precisely detected all those transactions which experienced transactional atomicity violations and identified all non-serializable traces. The overall results also show that RegionTrack incurred 1.10x and 1.08x lower memory and runtime overheads than Velodrome and 2.10x and 1.21x lower than Aerodrome, respectively. Moreover, it incurred 2.89x lower memory overhead than DoubleChecker. On average, Velodrome detected about 55% fewer violations than RegionTrack, which in turn reported about 3%–70% fewer violations than DoubleChecker. Shangru Wu, Ernest Bota Pobee, Xiupei Mei, Hao Zhang 0085, Bo Jiang 0001, Wing Kwong Chan |
ACM Trans. Softw. Eng. Methodol. | 7 |
| 2021 | DeepEutaxy: Diversity in Weight Search Direction for Fixing Deep Learning Model Training Through Batch PrioritizationabstractDeveloping a deep learning (DL) based software system is slow. One of the critical issues is to conduct many trials and errors in developing a DL model that usually serves as the major component of such a system. A major reason for this inefficiency is the progress of gradual reduction of the gap between the DL model under training and the ground truths. Prior techniques commonly focus on optimizing such errors after the errors have formed. They are insensitive to how a training dataset is provided to the DL model under training in batches, making their approaches nonproactive to deal with such errors. In this article, we propose DeepEutaxy, the first work to repair the model convergence problem from the batch prioritization perspective. Our key insight is that increasing the diversity (i.e., dissimilarity) of the corresponding weights of complex DL models before and after each training step can make the models learn faster and optimize the training errors quicker. DeepEutaxy first trains a DL model with several epochs for initialization. It then partitions and continually prioritizes the training batches for subsequent training epochs based on our novel notion of diversity between the pair of models before and after training on each batch, capturing the strength of the search direction to deal with training errors impacted by that batch. The experiment on six DL models over the MNIST and CIFAR-10 datasets shows that DeepEutaxy can accelerate the convergence of DL models on these two datasets with speedups of 1.75-8.45 and 2.67-15.15 times with respect to the training and test accuracies, respectively. DeepEutaxy can also be integrated into the existing techniques and compare favorably with the prior art in the experiment. Hao Zhang 0085, Wing Kwong Chan |
IEEE Trans. Reliab. | 2 |
| 2021 | Guest Editorial: Special Section on IEEE International Conference on Software Quality, Reliability, and Security (QRS) 2020abstractThe IEEE International Conference on Software Quality, Reliability, and Security celebrated its 20th anniversary in 2020.We received 156 submissions to the regular and short articles tracks. After the process of desk rejection, all remaining submissions were reviewed by three or more program committee members. Each submission was discussed online by the respective reviewers. Finally, the program chairs recommended six submissions of top quality out of the accepted articles to the IEEE TRANSACTIONS ON RELIABILITY for potential publication. We invited the authors of the six submissions to submit the corresponding revised versions that had addressed the reviewers’ comments to the IEEE TRANSACTIONS ON RELIABILITY for further evaluation. We invited the original reviewers of these QRS submissions to review these revised articles using the standard of the journal for article acceptance. In addition, the abstracts of these six selected articles were published in the QRS 2020 conference proceedings. The six articles we invited span the topics of both reliability and security in hardware and software with a new focus on AI-assisted techniques. The first three articles were published in theMarch 2021 issue; the fourth and fifth articles appear in this issue; and the last article will be included in the September 2021 issue. Wing Kwong Chan, Meiyappan Nagappan, Christof J. Budnik |
IEEE Trans. Reliab. | 1 |
| 2021 | Sifter: A Service Isolation Strategy for Internet ApplicationsabstractService oriented architecture (SOA) provides a flexible platform to build collaborative Internet applications by composing existing self-contained and autonomous services. However, the implicit interactions among the concurrently provisioned services may introduce interference to Internet applications and cause them behave abnormally. It is thus desirable to isolate services to safeguard their application consistency. Existing approaches mostly address this problem by restricting concurrent execution of services to avoid all the implicit interactions. These approaches, however, compromise the performance and flexibility of Internet applications due to the long running nature of services. This paper presents Sifter, a new service isolation strategy for Internet applications. We devise in this strategy a novel static approach to analyze the potential implicit interactions among the services and their impacts on the consistency of the associated Internet applications. By locating only those afflicted implicit interactions that may violate the application consistency, a novel approach based on exception handling and behavior constraints is customized to involved services to eliminate their impacts. We show that this approach exempts the consistency property of Internet applications from being interfered at runtime. The experimental results show that our approach has a better performance than existing solutions. Chunyang Ye, Shing-Chi Cheung, Wing Kwong Chan |
IEEE Trans. Serv. Comput. | 3 |
| 2020 | CUDAsmith: A Fuzzer for CUDA CompilersabstractCUDA is a parallel computing platform and programming model for the graphics processing unit (GPU) of NVIDIA. With CUDA programming, general purpose computing on GPU (GPGPU) is possible. However, the correctness of CUDA programs relies on the correctness of CUDA compilers, which is difficult to test due to its complexity. In this work, we propose CUDAsmith, a fuzzing framework for CUDA compilers. Our tool can randomly generate deterministic and valid CUDA kernel code with several different strategies. Moreover, it adopts random differential testing and EMI testing techniques to solve the test oracle problems of CUDA compiler testing. In particular, we lift live code injection to CUDA compiler testing to help generate EMI variants. Our fuzzing experiments with both the NVCC compiler and the Clang compiler for CUDA have detected thousands of failures, some of which have been confirmed by compiler developers. Finally, the cost-effectiveness of CUDAsmith is also thoroughly evaluated in our fuzzing experiment. Bo Jiang 0001, Wing Kwong Chan, T. H. Tse, Yongfeng Yin, Zhenyu Zhang 0004 |
COMPSAC | 3 |
| 2020 | EOSFuzzer: Fuzzing EOSIO Smart Contracts for Vulnerability DetectionabstractEOSIO is one typical public blockchain platform. It is scalable in terms of transaction speeds and has a growing ecosystem supporting smart contracts and decentralized applications. However, the vulnerabilities within the EOSIO smart contracts have led to serious attacks, which caused serious financial loss to its end users. In this work, we systematically analyzed three typical EOSIO smart contract vulnerabilities and their related attacks. Then we presented EOSFuzzer, a general black-box fuzzing framework to detect vulnerabilities within EOSIO smart contracts. In particular, EOSFuzzer proposed effective attacking scenarios and test oracles for EOSIO smart contract fuzzing. Our fuzzing experiment on 3963 EOSIO smart contracts shows that EOSFuzzer is both effective and efficient to detect EOSIO smart contract vulnerabilities with high accuracy. Yuhe Huang, Bo Jiang 0001, Wing Kwong Chan |
Internetware | 3 |
| 2020 | Improving Fault-Localization Accuracy by Referencing Debugging History to Alleviate Structure Bias in Code SuspiciousnessabstractSpectrum-based fault localization (SBFL) techniques can automatically localize software faults. They employ the program spectrum, such as code coverage profile with test verdicts, to rank the program entities based on their code suspiciousness. In the past decades, researchers have proposed many approaches to optimize these techniques; however, the program structure, which can influence their performance, is not taken into consideration in developing and improving these techniques. In this article, we identify and analyze the effect of the program structure on the application of SBFL techniques. We observe that some specific program structures may introduce structure bias to code suspiciousness and negatively influence the output of SBFL techniques. To mitigate these effects and improve the performance of fault localization, we propose Delta4Ts, a structure-aware technique. Delta4Ts references debugging history to alleviate the impact of structure bias in the calculation of code suspiciousness. It reasons from the observable suspicious value towards the desired suspicious value and the impact of structure bias. To evaluate Delta4Ts under practical constraints, we conduct a controlled experiment using nine widely-studied SBFL formulae on 12 C programs and 6 Java programs. The experiment results show that Delta4Ts can significantly improve the accuracy of the studied SBFL formulae by an average of 34.8% on 12 C programs and 30.6% on 6 Java programs, and improve more on subject programs associated with more history versions or having larger code sizes. Yang Feng 0003, Zhenyu Zhang 0004, Wing Kwong Chan, Jian Zhang 0001, Yuming Zhou |
IEEE Trans. Reliab. | 5 |
| 2019 | ConRS: A Requests Scheduling Framework for Increasing Concurrency Degree of Server ProgramsabstractServer programs always store a great deal of data and respond to plenty of requests in a short time. Most server programs are highly concurrent. Testing concurrent programs is difficult and costly because of their non-determinism. Researches have shown that increasing the degree of concurrency is more effective for testing process. This paper uses synchronization-pair and data race to quantify the degree of concurrency. The more synchronization-pairs and data races in a trace are, the higher concurrency degree is. A requests scheduling framework ConRS is proposed in this paper to reschedule requests in an existing test case and make the program "more concurrent". Compared with server stress testing tools, ConRS includes more categories of requests. Compared with coverage-guided bug detectors and data race detectors, ConRS are lower overhead and easier to understand. The experiments on MySQL database server show the effectiveness of ConRS. The synchronization-pairs and data races of test cases have been increased by at least 10% and 30% respectively after applying ConRS. Biyun Zhu, Ruijie Meng, Zhenyu Zhang 0004, Wing Kwong Chan |
COMPSAC (1) | 4 |
| 2019 | Efficient Transaction-Based Deterministic Replay for Multi-threaded ProgramsabstractExisting deterministic replay techniques propose strategies which attempt to reduce record log sizes and achieve successful replay. However, these techniques still generate large logs and achieve replay only under certain conditions. We propose a solution based on the division of the sequence of events of each thread into sequential blocks called transactions. Our insight is that there are usually few to no atomicity violations among transactions reported during a program execution. We present TPLAY, a novel deterministic replay technique which records thread access interleavings on shared memory locations at the transactional level. TPLAY also generates an artificial pair of interleavings when an atomicity violation is reported on a transaction. We present an experiment using the Splash2x extension of the PARSEC benchmark suite. Experimental results indicate that TPLAY experiences a 13-fold improvement in record log sizes and achieves a higher replay probability in comparison to existing work. Ernest Bota Pobee, Xiupei Mei, Wing Kwong Chan |
ASE | 3 |
| 2019 | Apricot: A Weight-Adaptation Approach to Fixing Deep Learning ModelsabstractA deep learning (DL) model is inherently imprecise. To address this problem, existing techniques retrain a DL model over a larger training dataset or with the help of fault injected models or using the insight of failing test cases in a DL model. In this paper, we present Apricot, a novel weight-adaptation approach to fixing DL models iteratively. Our key observation is that if the deep learning architecture of a DL model is trained over many different subsets of the original training dataset, the weights in the resultant reduced DL model (rDLM) can provide insights on the adjustment direction and magnitude of the weights in the original DL model to handle the test cases that the original DL model misclassifies. Apricot generates a set of such reduced DL models from the original DL model. In each iteration, for each failing test case experienced by the input DL model (iDLM), Apricot adjusts each weight of this iDLM toward the average weight of these rDLMs correctly classifying the test case and/or away from that of these rDLMs misclassifying the same test case, followed by training the weight-adjusted iDLM over the original training dataset to generate a new iDLM for the next iteration. The experiment using five state-of-the-art DL models shows that Apricot can increase the test accuracy of these models by 0.87%-1.55% with an average of 1.08%. The experiment also reveals the complementary nature of these rDLMs in Apricot. Hao Zhang 0085, Wing Kwong Chan |
ASE | 2 |
| 2019 | AggrePlay: efficient record and replay of multi-threaded programsabstractDeterministic replay presents challenges and often results in high memory and runtime overheads. Previous studies deterministically reproduce program outputs often only after several replay iterations or may produce a non-deterministic sequence of output to external sources. In this paper, we propose AggrePlay, a deterministic replay technique which is based on recording read-write interleavings leveraging thread-local determinism and summarized read values. During the record phase, AggrePlay records a read count vector clock for each thread on each memory location. Each thread checks the logged vector clock against the current read count in the replay phase before a write event. We present an experiment and analyze the results using the Splash2x benchmark suite as well as two real-world applications. The experimental results show that on average, AggrePlay experiences a better reduction in compressed log size, and 56% better runtime slowdown during the record phase, as well as a 41.58% higher probability in the replay phase than existing work. Ernest Bota Pobee, Wing Kwong Chan |
ESEC/SIGSOFT FSE | 2 |
| 2019 | A Systematic Study on Factors Impacting GUI Traversal-Based Test Case Generation Techniques for Android ApplicationsabstractMany test case generation algorithms have been proposed to test Android apps through their graphical user interfaces. However, no systematic study on the impact of the core design elements in these algorithms on effectiveness and efficiency has been reported. This paper presents the first controlled experiment to examine three key design factors, each of which is popularly used in GUI traversal-based test case generation techniques. These three major factors are definition of GUI state equivalence, state search strategy, and waiting time strategy in between input events. The empirical results on 33 Android apps with real faults revealed interesting results. First, different choices of GUI state equivalence led to significant difference on failure detection rate and extent of code coverage. Second, searching the GUI state hierarchy randomly is as effective as searching it systematically. Last but not the least, the choices on when to fire the next input event to the app under test is immaterial so long as the length of the test session is practically long enough such as 1 h. We also found two new GUI state equivalence definitions that are statistically as effective as the existing best strategy for GUI state equivalence. Bo Jiang 0001, Yaoyue Zhang, Wing Kwong Chan, Zhenyu Zhang 0004 |
IEEE Trans. Reliab. | 3 |
| 2018 | Fuse: An Architecture for Smart Contract Fuzz Testing ServiceabstractIn this paper, we report our project Fuse, which is a fuzz testing service. It presents the Fuse architecture, and discusses the progress and technical issues to be addressed to fuzz-test smart contracts and support fuzz-testing of Dapps. Wing Kwong Chan, Bo Jiang 0001 |
APSEC | 1 |
| 2018 | Message from SETA 2018 Symposium ChairsabstractPresents the introductory welcome message from the conference proceedings. May include the conference officers' congratulations to all involved with the conference event and publication of the proceedings record. Wing Kwong Chan, Hong Zhu 0002 |
COMPSAC (1) | 1 |
| 2018 | ReTestDroid: Towards Safer Regression Test Selection for Android ApplicationabstractMobile applications are widely used in our daily life and Android is the most popular open source mobile operating system. Because mobile applications update frequently, it is important developers to perform regression testing to ensure their quality. Modeling the control flow of an android application based on the activity lifecycle model only is imprecise for regression testing. Because many Android applications use asynchronous tasks, fragments, and native code frequently, which must be considered during change impact analysis. Otherwise, regression test selection techniques may miss some failure-revealing test cases, compromising the safety of these techniques. In this work, we propose a novel approach to model asynchronous task invocations, fragment-based activity lifecycle, and native code within the control flow graph of an Android application. Furthermore, we designed a regression test selection tool ReTestDroid based on our graph model. Our experiments on five real-life Android applications showed that our approach could enable much safer regression test selection while significantly saving regression-testing time. Bo Jiang 0001, Yongfei Zhang, Zhenyu Zhang 0004, Wing Kwong Chan |
COMPSAC (1) | 5 |
| 2018 | GBRAD: A General Framework to Evaluate Design Strategies for Hybrid Race DetectionabstractData race detection is a method in testing multithreaded programs to ensure their reliability against concurrency errors. In this paper, we present the GBRAD framework to support the initialization of various hybrid race detection techniques, which also supports the evaluation of these strategies at two decision points based on two major design factors of hybrid race detectors. In the GBRAD frame-work, one decision point consists of six skipping strategies and another decision point consists of eight reduction strategies. By combining these strategies, 48 hybrid detection techniques are initialized. We report a controlled experiment on the PARSEC benchmark suite as well as four real-world applications to evaluate these 48 techniques and their strategies in terms of runtime slowdown, memory overhead, and race detection effectiveness. The experiment identified 9 previously unknown techniques that are comparable to the state-of-the-art hybrid race detection technique. Wing Kwong Chan, Yuen-Tak Yu, Jacky W. Keung |
COMPSAC (1) | 2 |
| 2018 | ContractFuzzer: fuzzing smart contracts for vulnerability detectionabstractDecentralized cryptocurrencies feature the use of blockchain to transfer values among peers on networks without central agency. Smart contracts are programs running on top of the blockchain consensus protocol to enable people make agreements while minimizing trusts. Millions of smart contracts have been deployed in various decentralized applications. The security vulnerabilities within those smart contracts pose significant threats to their applications. Indeed, many critical security vulnerabilities within smart contracts on Ethereum platform have caused huge financial losses to their users. In this work, we present ContractFuzzer, a novel fuzzer to test Ethereum smart contracts for security vulnerabilities. ContractFuzzer generates fuzzing inputs based on the ABI specifications of smart contracts, defines test oracles to detect security vulnerabilities, instruments the EVM to log smart contracts runtime behaviors, and analyzes these logs to report security vulnerabilities. Our fuzzing of 6991 smart contracts has flagged more than 459 vulnerabilities with high precision. In particular, our fuzzing tool successfully detects the vulnerability of the DAO contract that leads to USD 60 million loss and the vulnerabilities of Parity Wallet that have led to the loss of USD 30 million and the freezing of USD 150 million worth of Ether. Bo Jiang 0001, Wing Kwong Chan |
ASE | 3 |
| 2018 | Special issue on software engineering technology and applications
Wing Kwong Chan, Xiaodong Liu 0002, Hridesh Rajan |
J. Syst. Softw. | 1 |
| 2018 | HistLock+: Precise Memory Access Maintenance Without Lockset Comparison for Complete Hybrid Data Race DetectionabstractDynamic hybrid data race detectors alleviate the detection imprecision problem incurred by pure lockset-based race detectors and the thread interleaving sensitive problem incurred by pure happens-before race detectors. Nonetheless, to ensure at least one data race on every memory location to be detected, keeping all historical memory access events in the analysis state of such a detector is impractical. Existing complete hybrid race detectors perform extensive comparisons among the locksets of the memory accesses on each memory location to identify which of them to be retained in its analysis state, which incurs significant runtime overhead. In this paper, we investigate to what extent a complete hybrid data race detector able to perform such identifications without lockset comparison. We present HistLock+, which is built atop thread epoch and lock release events to infer whether two memory accesses on the same memory location from the same thread in between consecutive lock release operations have any lock subset relation without performing expensive lockset comparison. HistLock+ guarantees exactly one racy memory access event to be reported on each thread segment separated by lock releases and hard-order thread synchronizations, and it never reports false positive on lockset violation. We have validated HistLock+ using the PARSEC benchmark suite and four real-world applications. The experimental results showed that HistLock+ was 122% faster and 28% more memory-efficient than the previous state-of-the-art complete hybrid race detector. Moreover, HistLock+ achieved the highest effectiveness in race detection among all evaluated race detectors in our experiment. Bo Jiang 0001, Wing Kwong Chan |
IEEE Trans. Reliab. | 3 |
| 2017 | Theoretical, Weak and Strong Accuracy Graphs of Spectrum-Based Fault Localization FormulasabstractDriven by the need to know which spectrum-based fault localization techniques are more effective in locating faults, many studies have sought to compare the accuracy of different formulas used in these techniques, resulting in findings of both theoretical and empirical accuracy relations of these formulas. Theoretical accuracy relations are independent of the specific programs and other settings involved, but limited by underlying assumptions and manual work in proofs. An accuracy graph can be constructed to holistically represent the proved relations. On the other hand, empirical studies are free of specific theoretical assumptions and can be highly automatable and scalable. A recent study has developed a systematic methodology based on statistical tests to reveal consistent and statistically sound empirical accuracy relations. That work has demonstrated the merits of empirical accuracy graphs in revealing relations that can be hard to prove. In this paper, we propose to use a stronger criterion for comparing formulas, describe an exploratory experiment to construct accuracy graphs based on the criterion, and report interesting relations found from the resulting accuracy graphs. Chung Man Tang, Wing Kwong Chan, Yuen-Tak Yu |
COMPSAC (2) | 2 |
| 2017 | A Case Study on Context Maintenance in Dynamic Hybrid Race DetectorsabstractMany dynamic hybrid race detectors aim at detecting violations of the lockset discipline in execution traces of multithreaded programs. They are designed to abstract memory accesses appearing in traces as contexts. Nonetheless, they keep these contexts in different extents and partition the sets of contexts into equivalent classes of different granularity. In our case study, we compare three detectors using the PARSEC benchmark suite to examine the impact of using unrestricted strategy or restricted strategy for keeping these contexts in sequence on detection effectiveness, and the impact of partitioning context sequences by different granularities on scalability in time cost. The case study results indicate that using restricted context sequences sufficed to detect very high proportions of locking discipline violations detectable by using unrestricted context sequences, and the partitioning of context sets into finer equivalent classes significantly lowers the scalability in time cost with increasing number of threads to handle the same input workload. Ernest Bota Pobee, Wing Kwong Chan |
COMPSAC (2) | 3 |
| 2017 | SimplyDroid: efficient event sequence simplification for Android applicationabstractTo ensure the quality of Android applications, many automatic test case generation techniques have been proposed. Among them, the Monkey fuzz testing tool and its variants are simple, effective and widely applicable. However, one major drawback of those Monkey tools is that they often generate many events in a failure-inducing input trace, which makes the follow-up debugging activities hard to apply. It is desirable to simplify or reduce the input event sequence while triggering the same failure. In this paper, we propose an efficient event trace representation and the SimplyDroid tool with three hierarchical delta-debugging algorithms each operating on this trace representation to simplify crash traces. We have evaluated SimplyDroid on a suite of real-life Android applications with 92 crash traces. The empirical result shows that our new algorithms in SimplyDroid are both efficient and effective in reducing these event traces. Bo Jiang 0001, Wing Kwong Chan |
ASE | 4 |
| 2017 | An Empirical Analysis of Three-Stage Data-Preprocessing for Analogy-Based Software Effort Estimation on the ISBSG DataabstractAnalogy-based software effort estimation is a method to estimate the project cost of an unseen project based on analogies against previous projects sharing selected features. The validity of the selected features depends on many factors, and one of most crucial factors is the effectiveness of the datapreprocessing techniques applied to the datasets of the previous projects. In this paper, we report the first controlled experiment that studies the class of three-stage data-preprocessing techniques with stages of missing data imputation, data normalization, and feature selection for analogy-based effort estimation. We conducted our investigation on the ISBSG data. The experimental results show that three-stage data-preprocessing techniques have significant impacts on the resultant effort estimation accuracy. The results also indicate that the combined use of Z-Score normalization, kNN imputation and mutual information based feature weighting can be an effective choice for analogy-based effort estimation. Jianglin Huang, Yan-Fu Li, Jacky W. Keung, Yuen-Tak Yu, Wing Kwong Chan |
QRS | 5 |
| 2017 | Which Factor Impacts GUI Traversal-Based Test Case Generation Technique Most? A Controlled Experiment on Android ApplicationsabstractThere are many research works on automated GUI traversal-based test case generation techniques for Android application. However, the effect of different factors used in a GUI traversal algorithm has not been systematically explored. In this work, we report a controlled experiment on 33 real-world applications to expose their real failures to systematically study three major factors that are commonly observed in testing tools for this class of applications. They include the notion of GUI state equivalence, the state search (or exploration) strategy, and the amount of time to wait between two input events. Our experimental results clearly show that different notions of GUI state equivalences have significantly different effects on failure detection rate and code coverage, randomized search is comparable to systematic search, and different choices of waiting time strategies do not make significant differences in terms of testing effectiveness. We also report other interesting results in this paper. Bo Jiang 0001, Yaoyue Zhang, Wing Kwong Chan, Zhenyu Zhang 0004 |
QRS | 3 |
| 2017 | Special Issue on Software Engineering Technology and Applications
Doris L. Carver, Wing Kwong Chan, Carl K. Chang |
J. Syst. Softw. | 2 |
| 2017 | Cross-validation based K nearest neighbor imputation for software quality datasets: An empirical study
Jianglin Huang, Jacky W. Keung, Federica Sarro, Yan-Fu Li, Yuen-Tak Yu, Wing Kwong Chan, Hongyi Sun |
J. Syst. Softw. | 6 |
| 2017 | A theoretical analysis on cloning the failed test cases to improve spectrum-based fault localization
Lanfei Yan, Zhenyu Zhang 0004, Jian Zhang 0001, Wing Kwong Chan, Zheng Zheng 0001 |
J. Syst. Softw. | 5 |
| 2017 | Accuracy Graphs of Spectrum-Based Fault Localization FormulasabstractThe effectiveness of spectrum-based fault localization techniques primarily relies on the accuracy of their fault localization formulas. Theoretical studies prove the relative accuracy orders of selected formulas under certain assumptions, forming a graph of their theoretical accuracy relations. However, it is unclear whether in such a graph the relative positions of these formulas may change when some assumptions are relaxed. On the other hand, empirical studies can measure the actual accuracy of any formula in controlled settings that more closely approximate practical scenarios but in less general contexts. In this paper, we propose an empirical framework of accuracy graphs and their construction that reveal the relative accuracy of formulas. Our work not only evaluates the association between certain assumptions and the theoretical relations among formulas, but also expands our knowledge to reveal new potential accuracy relationships of other formulas which have not been discovered by theoretical analysis. Using our proposed framework, we identified a list of formula pairs in which a formula is consistently statistically more accurate than or similar in accuracy to another, enlightening directions for further theoretical analysis. Chung Man Tang, Wing Kwong Chan, Yuen-Tak Yu, Zhenyu Zhang 0004 |
IEEE Trans. Reliab. | 2 |
| 2017 | SDA-CLOUD: A Multi-VM Architecture for Adaptive Dynamic Data Race DetectionabstractA concrete service consists of a number of program components, each of which is integrated to the service at either design time or runtime. In testing a concrete service, testers should validate the correctness of each of its components under diverse service consumption scenarios. Analyzing the program executions of these components under different configurations allows developers to compare and pinpoint issues therein. There is surprisingly little work in bridging this gap. In this paper, to the best of our knowledge, we propose the first work in designing dynamic analysis-as-a-service using a multi-virtual machine (multi-VM) approach to dynamic data race detection. Almost all existing work on dynamic data race detection focuses on improving detection precision, efficiency, or coverage of thread interleaving scenarios on the same but single compiled concurrent program component. Our model continually selects VM instances, each hosting a different compiled version of the same program component and running a state-of-the-art detector to detect data races. As such, our model innovatively takes existing race detectors as building blocks and operates at a higher level of abstraction. We have evaluated our proposal through an experiment. The experiment reveals that the multi-VM approach is feasible in monitoring multiple compiled versions and can detect different races both in amount and in detection probability. Under a limited execution budget constraint, the multi-VM approach is also significantly more effective in detecting races than approaches that use single compiled versions only. Some races hidden deeply in one compiled version have been found to be significantly more detectable in some other compiled versions of the same service component. Changjiang Jia, Chunbai Yang, Wing Kwong Chan, Yuen-Tak Yu |
IEEE Trans. Serv. Comput. | 3 |
| 2016 | Message from the SETA Organizing CommitteeabstractPresents the introductory welcome message from the conference proceedings. May include the conference officers' congratulations to all involved with the conference event and publication of the proceedings record. Doris L. Carver, Wing Kwong Chan, Xiaodong Liu 0002, Carl K. Chang, Hridesh Rajan |
COMPSAC | 2 |
| 2016 | Testing and Debugging in Continuous Integration with Budget Quotas on Test ExecutionsabstractIn Continuous Integration, a software application is developed through a series of development sessions, each with limited time allocated to testing and debugging on each of its modules. Test Case Prioritization can help execute test cases with higher failure estimate earlier in each session. When the testing time is limited, executing such prioritized test cases may only produce partial and prioritized execution coverage data. To identify faulty code, existing Spectrum-Based Fault Localization techniques often use execution coverage data but without the assumption of execution coverage priority. Is it possible to decompose these two steps for optimization within individual steps? In this paper, we study to what extent the selection of test case prioritization techniques may reduce its influence on the effectiveness of spectrum-based fault localization, thereby showing the possibility to decompose the process of continuous integration for optimization in workflow steps. We present a controlled experiment using the Siemens suite as subjects, nine test case prioritization techniques and four spectrum-based fault localization techniques. The findings showed that the studied test cases prioritization and spectrum-based fault localization can be customized separately, and, interestingly, prioritization over a smaller test suite can enable spectrum-based fault localization to achieve higher accuracy by assigning faulty statements with higher ranks. Bo Jiang 0001, Wing Kwong Chan |
QRS | 2 |
| 2016 | Facilitating Monkey Test by Detecting Operable Regions in Rendered GUI of Mobile Game AppsabstractGraphical User Interface (GUI) is a component of many software applications. Many mobile game applications in particular have to provide excellent user experiences using graphical engines to render GUI screens. On a rendered GUI screen such as a treasury map, no GUI widget is embodied in it and the operable GUI regions, each of which is a region that triggers actions when certain events acting on these regions, may only be implicitly determinable. Traditional testing tools like monkey test do not effectively generate effective event sequences over such operable GUI regions. Our insight is that operable regions in a rendered GUI screen of many mobile game applications are given with visible hints to catch user attentions. In this paper, we propose Smart Monkey, which uses the fundamental features of a screen, including color, intensity, and texture, as visual signals to detect operable GUI region candidates, and iteratively identifies and confirms the real operable GUI regions by launching GUI events to the region. We have implemented Smart Monkey as a testing tool for Android apps and conducted case studies on real-world applications to compare it with a peer technique. The empirical results show that it effective in identifying such operable regions and thus able to generate functional event sequences more efficiently. Chenglong Sun, Zhenyu Zhang 0004, Bo Jiang 0001, Wing Kwong Chan |
QRS | 4 |
| 2016 | DFL: Dual-Service Fault LocalizationabstractIn engineering a service, software developers often construct and deploy a newer (forthcoming) version of the service to replace the current version. A forthcoming version is often placed online for users to consume and report feedback. In the case of observed failures, the forthcoming version should be debugged and further evolved. In this paper, we propose the model of dual-service fault localization (DFL) to aid this evolution process. Many prior research studies on spectrum-based fault localization (SBFL) consider each version separately. The DFL model correlates the dynamic execution spectra of the current and the forthcoming versions of the same service placed for live test of the forthcoming version, and dynamically generates an adaptive fault localization formula to estimate the code regions in the forthcoming service responsible for the observed failures. We report an experiment in which we initialized the DFL model into six instances, each using an ensemble technique dynamically composed from 11 existing SBFL formulas, and applied the model to four benchmarks. The results show that DFL is feasible and multiple instances are statistically more effective than, if not as effective as, the best of these individual SBFL formulas on each benchmark. Chung Man Tang, Jacky W. Keung, Yuen-Tak Yu, Wing Kwong Chan |
QRS | 4 |
| 2016 | HistLock: Efficient and Sound Hybrid Detection of Hidden Predictive Data Races with Functional ContextsabstractState-of-the-art hybrid data race detectors combine the happens-before relation and the locking discipline to alleviate the imprecision problem incurred by lockset-based detector and the thread interleaving problem incurred by happens-before detectors. However, they incur high runtime overheads. In this paper, we present HistLock, a novel and sound hybrid dynamic race detector, which attains high precision, low slowdown and memory overheads, and thread insensitivity. It formulates a novel context-based strategy to phrase out non-redundant memory access events and check races to conserve both time and memory computation. It ensures each race warning to be the one violating the locking discipline in the original event history. Our experiment compared HistLock to FastTrack, AccuLock, and MultiLock-HB, which were a precise happens-before race detector, an imprecise hybrid race detector, and a precise hybrid race detector, respectively, on 13 benchmark subjects. HistLock was found to be of higher precision, 156% faster and 33% more memory-efficient than MultiLock-HB. It detected 59 more race warnings than AccuLock, attaining higher effectiveness, but ran slower by 43%. In most cases, in a single run, HistLock reported all the races detected by FastTrack in 100 runs. Chunbai Yang, Wing Kwong Chan |
QRS | 3 |
| 2016 | Editorial of the special issue to celebrate the 35th anniversary of JSS
W. Eric Wong, Wing Kwong Chan |
J. Syst. Softw. | 2 |
| 2016 | Hierarchical Program PathsabstractComplete dynamic control flow is a fundamental kind of execution profile about program executions with a wide range of applications. Tracing the dynamic control flow of program executions for a brief period easily generates a trace consisting of billions of control flow events. The number of events in such a trace is large, making both path tracing and path querying to incur significant slowdowns. A major class of path tracing techniques is to design novel trace representations that can be generated efficiently, and encode the inputted sequences of such events so that the inputted sequences are still derivable from the encoded and smaller representations. The control flow semantics in such representations have, however, become obscure, which makes implementing path queries on such a representation inefficient and the design of such queries complicated. We propose a novel two-phase path tracing framework— Hierarchical Program Path (HPP)—to model the complete dynamic control flow of an arbitrary number of executions of a program. In Phase 1, HPP monitors each execution, and efficiently generates a stream of events, namely HPPTree, representing a novel tree-based representation of control flow for each thread of control in the execution. In Phase 2, given a set of such event streams, HPP identifies all the equivalent instances of the same exercised interprocedural path in all the corresponding HPPTree instances, and represents each such equivalent set of paths with a single subgraph, resulting in our compositional graph-based trace representation, namely, HPPDAG. Thus, an HPPDAG instance has the potential to be significantly smaller in size than the corresponding HPPTree instances, and still completely preserves the control flow semantics of the traced executions. Control flow queries over all the traced executions can also be directly performed on the single HPPDAG instance instead of separately processing the trace representation of each execution followed by aggregating their results. We validate HPP using the SPLASH2 and SPECint 2006 benchmarks. Compared to the existing technique, named BLPT (Ball-Larus-based Path Tracing), HPP generates significantly smaller trace representations and incurs fewer slowdowns to the native executions in online tracing of Phase 1. The HPPDAG instances generated in Phase 2 are significantly smaller than their corresponding BLPT and HPPTree traces. We show that HPPDAG supports efficient backtrace querying, which is a representative path query based on complete control flow trace. Finally, we illustrate the ease of use of HPPDAG by building a novel and highly efficient path profiling technique to demonstrate the applicability of HPPDAG. Chunbai Yang, Shangru Wu, Wing Kwong Chan |
ACM Trans. Softw. Eng. Methodol. | 3 |
| 2016 | ASP: Abstraction Subspace Partitioning for Detection of Atomicity Violations with an Empirical StudyabstractDynamic concurrency bug detectors predict and then examine suspicious instances of atomicity violations from executions of multithreaded programs. Only few predicted instances are real bugs. Prioritizing such instances can make the examinations cost-effective, but is there any design factor exhibiting significant influence? This work presents the first controlled experiment that studies two design factors, abstraction level and subspace, in partitioning such instances through 35 resultant partition-based techniques on 10 benchmarks with known vulnerability-related bugs. The empirical analysis reveals significant findings. First, partition-based prioritization can significantly improve the fault detection rate. Second, coarse-grained techniques are more effective than fine-grained ones, and using some one-dimensional subspaces is more effective than using other dimensional subspaces. Third, eight previously unknown techniques can be more effective than the technique modeled after a state-of-the-art dynamic detector. Shangru Wu, Chunbai Yang, Changjiang Jia, Wing Kwong Chan |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2016 | To What Extent is Stress Testing of Android TV Applications Automated in Industrial Environments?abstractAn Android-based smart television (TV) must reliably run its applications in an embedded program environment under diverse hardware resource conditions. Owing to the diverse hardware components used to build numerous TV models, TV simulators are usually not sufficiently high in fidelity to simulate various TV models and thus are only regarded as unreliable alternatives when stress testing such applications. Therefore, even though stress testing on real TV sets is tedious, it is the de facto approach to ensure the reliability of these applications in the industry. In this paper, we study to what extent stress testing of smart TV applications can be fully automated in the industrial environments. To the best of our knowledge, no previous work has addressed this important question. We summarize the findings collected from ten industrial test engineers who have tested 20 such TV applications in a real production environment. Our study shows that the industry required test automation supports on high-level GUI object controls and status checking, setup of resource conditions, and the interplay between the two. With such supports, 87% of the industrial test specifications of one TV model can be fully automated, and 71.4% of them were found to be fully reusable to test a subsequent TV model with major upgrades of hardware, operating system, and application. It represents a significant improvement with margins of 28% and 38%, respectively, compared with stress testing without such supports. Bo Jiang 0001, Wing Kwong Chan, Xinchao Zhang |
IEEE Trans. Reliab. | 3 |
| 2016 | Editorial: Software Engineering and Applications for Cloud-Based Mobile SystemsabstractPresents an editorial which explores the market fr cloud-based mobile computing systems. The concept of cloud-based mobile applications has emerged recently to address the limitations of pure mobile applications by leveraging the vast resources available at computing clouds. Traditionally, mobile applications benefit by outsourcing a part of the program execution to a cloud-based system to compute the (intermediate) results and send back these results for integration by the relevant mobile application. By doing so, the program code, data as well as device-specific information, such as user locations, can be made available to a cloud. Infrastructures to support computation outsourcing of the above type and its variants have been widely studied; and the associated security and privacy issues are hot topics in research. Doris L. Carver, Wing Kwong Chan, Carl K. Chang |
IEEE Trans. Serv. Comput. | 2 |
| 2015 | Message from SETA Symposium Organizing CommitteeabstractPresents a listing of the Symposium organizing committee. Doris L. Carver, Wing Kwong Chan, Carl K. Chang |
COMPSAC | 3 |
| 2015 | Architecturing Dynamic Data Race Detection as a Cloud-Based ServiceabstractA web-based service consists of layers of programs (components) in the technology stack. Analyzing program executions of these components separately allows service vendors to acquire insights into specific program behaviors or problems in these components, thereby pinpointing areas of improvement in their offering services. Many existing approaches for testing as a service take an orchestration approach that splits components under test and the analysis services into a set of distributed modules communicating through message-based approaches. In this paper, we present the first work in providing dynamic analysis as a service using a virtual machine (VM)-based approach on dynamic data race detection. Such a detection needs to track a huge number of events performed by each thread of a program execution of a service component, making such an analysis unsuitable to use message passing to transit huge numbers of events individually. In our model, we instruct VMs to perform holistic dynamic race detections on service components and only transfer the detection results to our service selection component. With such result data as the guidance, the service selection component accordingly selects VM instances to fulfill subsequent analysis requests. The experimental results show that our model is feasible. Changjiang Jia, Chunbai Yang, Wing Kwong Chan |
ICWS | 3 |
| 2015 | PORA: Proportion-Oriented Randomized Algorithm for Test Case PrioritizationabstractEffective testing is essential for assuring software quality. While regression testing is time-consuming, the fault detection capability may be compromised if some test cases are discarded. Test case prioritization is a viable solution. To the best of our knowledge, the most effective test case prioritization approach is still the additional greedy algorithm, and existing search-based algorithms have been shown to be visually less effective than the former algorithms in previous empirical studies. This paper proposes a novel Proportion-Oriented Randomized Algorithm (PORA) for test case prioritization. PORA guides test case prioritization by optimizing the distance between the prioritized test suite and a hierarchy of distributions of test input data. Our experiment shows that PORA test case prioritization techniques are as effective as, if not more effective than, the total greedy, additional greedy, and ART techniques, which use code coverage information. Moreover, the experiment shows that PORA techniques are more stable in effectiveness than the others. Bo Jiang 0001, Wing Kwong Chan, T. H. Tse |
QRS | 2 |
| 2015 | ASR: Abstraction Subspace Reduction for Exposing Atomicity Violation Bugs in Multithreaded ProgramsabstractMany two-phase based dynamic concurrency bug detectors predict suspicious instances of atomicity violation from one execution trace, and examine each such instance by scheduling a confirmation run. If the amount of suspicious instances predicted is large, confirming all these instances becomes a burden. In this paper, we present the first controlled experiment that evaluates the efficiency, effectiveness, and cost-effectiveness of reduction on suspicious instances in the detection of atomicity violations. A novel form of reduction technique named ASR is proposed. Our empirical analysis reveals many interesting findings: First, the reduced sets of instances produced by ASR significantly improve the efficiency of atomicity violation detection without significantly compromising the effectiveness. Second, ASR is significantly more cost-effective than random reduction and untreated reduction by 8.5 folds and 60.7 folds, respectively, in terms of mean normalized bug detection ratio. Third, six ASR techniques can be significantly more cost-effective than the technique modeled after a state-of-the-art detector. Shangru Wu, Chunbai Yang, Wing Kwong Chan |
QRS | 3 |
| 2015 | Input-based adaptive randomized test case prioritization: A local beam search approachabstractTest case prioritization assigns the execution priorities of the test cases in a given test suite. Many existing test case prioritization techniques assume the full-fledged availability of code coverage data, fault history, or test specification, which are seldom well-maintained in real-world software development projects. This paper proposes a novel family of input-based local-beam-search adaptive-randomized techniques. They make adaptive tree-based randomized explorations with a randomized candidate test set strategy to even out the search space explorations among the branches of the exploration trees constructed by the test inputs in the test suite. We report a validation experiment on a suite of four medium-size benchmarks. The results show that our techniques achieve either higher APFD values than or the same mean APFD values as the existing code-coverage-based greedy or search-based prioritization techniques, including Genetic, Greedy and ART, in both our controlled experiment and case study. Our techniques are also significantly more efficient than the Genetic and Greedy, but are less efficient than ART. Bo Jiang 0001, Wing Kwong Chan |
J. Syst. Softw. | 2 |
| 2015 | ASN: A Dynamic Barrier-Based Approach to Confirmation of Deadlocks from Warnings for Large-Scale Multithreaded ProgramsabstractMany large-scale multithreaded programs incur deadlock bugs. Existing deadlock warning detection techniques only report warning scenarios, which may or may not be real deadlocks. Each warning should be further verified on whether it may manifest into a real deadlock. For this purpose, a number of active randomized testing schedulers have been developed to trigger them, and yet pervious experiments show that their deadlock confirmation probability can be low. This paper presents ASN, a novel barrier-based randomized scheduler that triggers real deadlocks with high probabilities. We exploit the insights that in a confirmation run, the threads involved in a real deadlock should properly acquire one or more sets of locks prior to deadlocking. ASN automatically identifies three interesting sets of such positions. It guides the threads participating in a given warning to stay at these position sets in turn. When all the threads are staying at the last position set, ASN checks whether any deadlock that matches with the given warning has been triggered. We have evaluated ASN on 15 deadlock bugs in a suite of real-world multithreaded programs. The results show that ASN either confirms more deadlocks from the benchmark suite or triggers the same deadlocks with significantly higher probabilities than existing schedulers. Yan Cai 0001, Changjiang Jia, Shangru Wu, Ke Zhai 0002, Wing Kwong Chan |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2015 | A Subsumption Hierarchy of Test Case Prioritization for Composite ServicesabstractMany composite workflow services utilize non-imperative XML technologies such as WSDL, XPath, XML schema, and XML messages. Regression testing should assure the services against regression faults that appear in both the workflows and these artifacts. In this paper, we propose a refinement-oriented level-exploration strategy and a multilevel coverage model that captures progressively the coverage of different types of artifacts by the test cases. We show that by using them, the test case prioritization techniques initialized on top of existing greedy-based test case prioritization strategy form a subsumption hierarchy such that a technique can produce more test suite permutations than a technique that subsumes it. Our experimental study of a model instance shows that a technique generally achieves a higher fault detection rate than a subsumed technique, which validates that the proposed hierarchy and model have the potential to improve the cost-effectiveness of test case prioritization techniques. Lijun Mei, Yan Cai 0001, Changjiang Jia, Bo Jiang 0001, Wing Kwong Chan, Zhenyu Zhang 0004, T. H. Tse |
IEEE Trans. Serv. Comput. | 5 |
| 2015 | Preemptive Regression Testingof Workflow-Based Web ServicesabstractAn external web service may evolve without prior notification. In the course of the regression testing of a workflow-based web service, existing test case prioritization techniques may only verify the latest service composition using the not-yet-executed test cases, overlooking high-priority test cases that have already been applied to the service composition before the evolution. In this paper, we propose Preemptive Regression Testing (PRT), an adaptive testing approach to addressing this challenge. Whenever a change in the coverage of any service artifact is detected, PRT recursively preempts the current session of regression test and creates a sub-session of the current test session to assure such lately identified changes in coverage by adjusting the execution priority of the test cases in the test suite. Then, the sub-session will resume the execution from the suspended position. PRT terminates only when each test case in the test suite has been executed at least once without any preemption activated in between any test case executions. The experimental result confirms that testing workflow-based web service in the face of such changes is very challenging; and one of the PRT-enriched techniques shows its potential to overcome the challenge. Lijun Mei, Wing Kwong Chan, T. H. Tse, Bo Jiang 0001, Ke Zhai 0002 |
IEEE Trans. Serv. Comput. | 2 |
| 2014 | CrowdAdaptor: A Crowd Sourcing Approach toward Adaptive Energy-Efficient Configurations of Virtual Machines Hosting Mobile ApplicationsabstractApplications written by end-user programmers are hardly energy-optimized by these programmers. The end users of such applications thus suffer significant energy issues. In this paper, we propose CrowdAdaptor, a novel approach toward locating energy-efficient configurations to execute the applications hosted in virtual machines on handheld devices. CrowdAdaptor innovatively makes use of the development artifacts (test cases) and the very large installation base of the same application to distribute the test executions and performance data collection of the whole test suites against many different virtual machine configurations among these installation bases. It synthesizes these data, continuously discovers better energy-efficient configurations, and makes them available to all the installations of the same applications. We report a multi-subject case study on the ability of the framework to discover energy-efficient configurations in three power models. The results show that Crowd Adaptor can achieve up to 50% of energy savings based on a conservative linear power model. Edward Y. Y. Kan, Wing Kwong Chan, T. H. Tse |
COMPSAC | 2 |
| 2014 | Extending the Theoretical Fault Localization Effectiveness Hierarchy with Empirical Results at Different Code Abstraction LevelsabstractSpectrum-based fault localization techniques are semi-automated program debugging techniques that address the bottleneck of finding suspicious program locations for diagnosis. They assess the fault suspiciousness of individual program locations based on the code coverage data achieved by executing the program under debugging over a test suite. A program location can be viewed at different abstraction levels, such as a statement in the source code or an instruction compiled from the source code. In general, a program location at one code abstraction level can be transformed into zero to more program locations at another abstraction level. Although programmers usually debug at the source code level, the code is actually executed at a lower level. It is unclear whether the same techniques applied at different code abstraction levels may achieve consistent results. In this paper, we study a suite of spectrum-based fault localization techniques at both the source and instruction code levels in the context of an existing theoretical hierarchy to assess whether their effectiveness is consistent across the two levels. Our study extends the theoretical hierarchy with empirically validated relationships across two code abstraction levels toward an integration of the theory and practice of fault localization. Chung Man Tang, Wing Kwong Chan, Yuen-Tak Yu |
COMPSAC | 2 |
| 2014 | ConLock: a constraint-based approach to dynamic checking on deadlocks in multithreaded programsabstractMany predictive deadlock detection techniques analyze multithreaded programs to suggest potential deadlocks (referred to as cycles or deadlock warnings). Nonetheless, many of such cycles are false positives. On checking these cycles, existing dynamic deadlock confirmation techniques may frequently encounter thrashing or result in a low confirmation probability. This paper presents a novel technique entitled ConLock to address these problems. ConLock firstly analyzes a given cycle and the execution trace that produces the cycle. It identifies a set of thread scheduling constraints based on a novel should-happen-before relation. ConLock then manipulates a confirmation run with the aim to not violate a reduced set of scheduling constraints and to trigger an occurrence of the deadlock if the cycle is a real deadlock. If the cycle is a false positive, ConLock reports scheduling violations. We have validated ConLock using a suite of real-world programs with 11 deadlocks. The result shows that among all 741 cycles reported by Magiclock, ConLock confirms all 11 deadlocks with a probability of 71%−100%. On the remaining 730 cycles, ConLock reports scheduling violations on each. We have systematically sampled 87 out of the 730 cycles and confirmed that all these cycles are false positives. Yan Cai 0001, Shangru Wu, Wing Kwong Chan |
ICSE | 3 |
| 2014 | Is XML-Based Test Case Prioritization for Validating WS-BPEL Evolution Effective in Both Average and Adverse Scenarios?abstractIn real life, a tester can only afford to apply one test case prioritization technique to one test suite against a service-oriented workflow application once in the regression testing of the application, even if it results in an adverse scenario such that the actual performance in the test session is far below the average. It is unclear whether the factors of test case prioritization techniques known to be significant in terms of average performance can be extrapolated to adverse scenarios. In this paper, we examine whether such a factor or technique may consistently affect the rate of fault detection in both the average and adverse scenarios. The factors studied include prioritization strategy, artifacts to provide coverage data, ordering direction of a strategy, and the use of executable and non-executable artifacts. The results show that only a minor portion of the 10 studied techniques, most of which are based on the iterative strategy, are consistently effective in both average and adverse scenarios. To the best of our knowledge, this paper presents the first piece of empirical evidence regarding the consistency in the effectiveness of test case prioritization techniques and factors of service-oriented workflow applications between average and adverse scenarios. Changjiang Jia, Lijun Mei, Wing Kwong Chan, Yuen-Tak Yu, T. H. Tse |
ICWS | 3 |
| 2014 | Improving the Effectiveness of Testing Pervasive Software via Context DiversityabstractContext-aware pervasive software is responsive to various contexts and their changes. A faulty implementation of the context-aware features may lead to unpredictable behavior with adverse effects. In software testing, one of the most important research issues is to determine the sufficiency of a test suite to verify the software under test. Existing adequacy criteria for testing traditional software, however, have not explored the dimension of serial test inputs and have not considered context changes when constructing test suites. In this article, we define the concept of context diversity to capture the extent of context changes in serial inputs and propose three strategies to study how context diversity may improve the effectiveness of the data-flow testing criteria. Our case study shows that the strategy that uses test cases with higher context diversity can significantly improve the effectiveness of existing data-flow testing criteria for context-aware pervasive software. In addition, test suites with higher context diversity are found to execute significantly longer paths, which may provide a clue that reveals why context diversity can contribute to the improvement of effectiveness of test suites. Huai Wang, Wing Kwong Chan, T. H. Tse |
ACM Trans. Auton. Adapt. Syst. | 2 |
| 2014 | Prioritizing Test Cases for Regression Testing of Location-Based Services: Metrics, Techniques, and Case StudyabstractLocation-based services (LBS) are widely deployed. When the implementation of an LBS-enabled service has evolved, regression testing can be employed to assure the previously established behaviors not having been adversely affected. Proper test case prioritization helps reveal service anomalies efficiently so that fixes can be scheduled earlier to minimize the nuisance to service consumers. A key observation is that locations captured in the inputs and the expected outputs of test cases are physically correlated by the LBS-enabled service, and these services heuristically use estimated and imprecise locations for their computations, making these services tend to treat locations in close proximity homogenously. This paper exploits this observation. It proposes a suite of metrics and initializes them to demonstrate input-guided techniques and point-of-interest (POI) aware test case prioritization techniques, differing by whether the location information in the expected outputs of test cases is used. It reports a case study on a stateful LBS-enabled service. The case study shows that the POI-aware techniques can be more effective and more stable than the baseline, which reorders test cases randomly, and the input-guided techniques. We also find that one of the POI-aware techniques, cdist, is either the most effective or the second most effective technique among all the studied techniques in our evaluated aspects, although no technique excels in all studied SOA fault classes. Ke Zhai 0002, Bo Jiang 0001, Wing Kwong Chan |
IEEE Trans. Serv. Comput. | 3 |
| 2014 | Magiclock: Scalable Detection ofPotential Deadlocks in Large-ScaleMultithreaded ProgramsabstractWe present Magiclock, a novel potential deadlock detection technique by analyzing execution traces (containing no deadlock occurrence) of large-scale multithreaded programs. Magiclock iteratively eliminates removable lock dependencies before potential deadlock localization. It divides lock dependencies into thread specific partitions, consolidates equivalent lock dependencies, and searches over the set of lock dependency chains without the need to examine any duplicated permutations of the same lock dependency chains. We validate Magiclock through a suite of real-world, large-scale multithreaded programs. The experimental results show that Magiclock is significantly more scalable and efficient than existing dynamic detectors in analyzing and detecting potential deadlocks in execution traces of large-scale multithreaded programs. Yan Cai 0001, Wing Kwong Chan |
IEEE Trans. Software Eng. | 2 |
| 2013 | Bypassing Code Coverage Approximation Limitations via Effective Input-Based Randomized Test Case PrioritizationabstractTest case prioritization assigns the execution priorities of the test cases in a given test suite with the aim of achieving certain goals. Many existing test case prioritization techniques however assume the full-fledged availability of code coverage data, fault history, or test specification, which are seldom well-maintained in many software development projects. This paper proposes a novel family of LBS techniques. They make adaptive tree-based randomized explorations with an adaptive randomized candidate test set strategy to diversify the explorations among the branches of the exploration trees constructed by the test inputs in the test suite. They get rid of the assumption on the historical correlation of code coverage between program versions. Our techniques can be applied to programs with or without any previous versions, and hence are more general than many existing test case prioritization techniques. The empirical study on four popular UNIX utility benchmarks shows that, in terms of APFD, our LBS techniques can be as effective as some of the best code coverage-based greedy prioritization techniques ever proposed. We also show that they are significantly more efficient and scalable than the latter techniques. Bo Jiang 0001, Wing Kwong Chan |
COMPSAC | 2 |
| 2013 | Prioritizing Structurally Complex Test Pairs for Validating WS-BPEL EvolutionsabstractMany web services represent their artifacts in the semi-structural format. Such artifacts may or may not be structurally complex. Many existing test case prioritization techniques however treat test cases of different complexity generically. In this paper, we exploit the insights on the structural similarity of XML-based artifacts between test cases, and propose a family of test case prioritization techniques that iteratively selects test case pairs without replacement. The validation experiment shows that these techniques can be more cost-effective than the studied existing techniques in exposing faults. Lijun Mei, Yan Cai 0001, Changjiang Jia, Bo Jiang 0001, Wing Kwong Chan |
ICWS | 5 |
| 2013 | TeamWork: synchronizing threads globally to detect real deadlocks for multithreaded programsabstractThis paper presents the aim of TeamWork, our ongoing effort to develop a comprehensive dynamic deadlock confirmation tool for multithreaded programs. It also presents a refined object abstraction algorithm that refines the existing stack hash abstraction. Yan Cai 0001, Ke Zhai 0002, Shangru Wu, Wing Kwong Chan |
PPoPP | 4 |
| 2013 | On the adoption of MC/DC and control-flow adequacy for a tight integration of program testing and statistical fault localization
Bo Jiang 0001, Ke Zhai 0002, Wing Kwong Chan, T. H. Tse, Zhenyu Zhang 0004 |
Inf. Softw. Technol. | 3 |
| 2013 | A general noise-reduction framework for fault localization of Java programs
Zhenyu Zhang 0004, Wing Kwong Chan, T. H. Tse, Shanping Li |
Inf. Softw. Technol. | 3 |
| 2013 | In quest of the science in statistical fault localizationabstractSUMMARY Many researchers employ various statistical methods to locate faults in faulty programs. Like other researchers, we sometimes have made mistakes in the quest of making statistical fault localization both practical and scientific. In this experience report, we reflect on our work conducted on this topic, organize our isolated experiences in the format of models and errors, and cast them in the context of statistics. Copyright © 2011 John Wiley & Sons, Ltd. Wing Kwong Chan, Yan Cai 0001 |
Softw. Pract. Exp. | 1 |
| 2013 | Lock Trace Reduction for Multithreaded ProgramsabstractMany happened-before-based detectors for debugging multithreaded programs implement vector clocks to incrementally track the casual relations among synchronization events produced by concurrent threads and generate trace logs. They update the vector clocks via vector-based comparison and content assignment in every case. We observe that many such tracking comparison and assignment operations are removable in part or in whole, which if identified and used properly, have the potential to reduce the log traces thus produced. This paper presents our analysis to identify such removable tracking operations and shows how they could be used to reduce log traces. We implement our analysis result as a technique entitled LOFT. We evaluate LOFT on the well-studied PARSEC benchmarking suite and five large-scale real-world applications. The main experimental result shows that on average, LOFT identifies 63.9 percent of all synchronization operations incurred by the existing approach as removable and does not compromise the efficiency of the latter. Yan Cai 0001, Wing Kwong Chan |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2012 | Preemptive Regression Test Scheduling Strategies: A New Testing Approach to Thriving on the Volatile Service EnvironmentsabstractA workflow-based web service may use ultra-late binding to invoke external web services to concretize its implementation at run time. Nonetheless, such external services or the availability of recently used external services may evolve without prior notification, dynamically triggering the workflow-based service to bind to new replacement external services to continue the current execution. Any integration mismatch may cause a failure. In this paper, we propose Preemptive Regression Testing (PRT), a novel testing approach that addresses this adaptive issue. Whenever such a late-change on the service under regression test is detected, PRT preempts the currently executed regression test suite, searches for additional test cases as fixes, runs these fixes, and then resumes the execution of the regression test suite from the preemption point. Lijun Mei, Ke Zhai 0002, Bo Jiang 0001, Wing Kwong Chan, T. H. Tse |
COMPSAC | 4 |
| 2012 | MagicFuzzer: Scalable deadlock detection for large-scale applicationsabstractWe present MagicFuzzer, a novel dynamic deadlock detection technique. Unlike existing techniques to locate potential deadlock cycles from an execution, it iteratively prunes lock dependencies that each has no incoming or outgoing edge. Combining with a novel thread-specific strategy, it dramatically shrinks the size of lock dependency set for cycle detection, improving the efficiency and scalability of such a detection significantly. In the real deadlock confirmation phase, it uses a new strategy to actively schedule threads of an execution against the whole set of potential deadlock cycles. We have implemented a prototype and evaluated it on large-scale C/C++ programs. The experimental results confirm that our technique is significantly more effective and efficient than existing techniques. Yan Cai 0001, Wing Kwong Chan |
ICSE | 2 |
| 2012 | CARISMA: a context-sensitive approach to race-condition sample-instance selection for multithreaded applicationsabstractDynamic race detectors can explore multiple thread schedules of a multithreaded program over the same input to detect data races. Although existing sampling-based precise race detectors reduce overheads effectively so that lightweight precise race detection can be performed in testing or post-deployment environments, they are ineffective in detecting races if the sampling rates are low. This paper presents CARISMA to address this problem. CARISMA exploits the insight that along an execution trace, a program may potentially handle many accesses to the memory locations created at the same site for similar purposes. Iterating over multiple execution trials of the same input, CARISMA estimates and distributes the sampling budgets among such location creation sites, and probabilistically collects a fraction of all accesses to the memory locations associated with such sites for subsequent race detection. Our experiment shows that, compared with PACER on the same platform and at the same sampling rate (such as 1%), CARISMA is significantly more effective. Ke Zhai 0002, Boni Xu, Wing Kwong Chan, T. H. Tse |
ISSTA | 3 |
| 2012 | How well does test case prioritization integrate with statistical fault localization?
Bo Jiang 0001, Zhenyu Zhang 0004, Wing Kwong Chan, T. H. Tse, Tsong Yueh Chen |
Inf. Softw. Technol. | 3 |
| 2012 | Human and program factors affecting the maintenance of programs with deployed design patterns
Tsz Hin Ng, Yuen-Tak Yu, Shing-Chi Cheung, Wing Kwong Chan |
Inf. Softw. Technol. | 4 |
| 2012 | EClass: An execution classification approach to improving the energy-efficiency of software via machine learning
Edward Y. Y. Kan, Wing Kwong Chan, T. H. Tse |
J. Syst. Softw. | 2 |
| 2012 | Special Issue on Dynamic Analysis and Testing of Embedded Software
W. Eric Wong, Wing Kwong Chan, T. H. Tse, Fei-Ching Kuo |
J. Syst. Softw. | 2 |
| 2012 | An empirical evaluation of several test-a-few strategies for testing particular conditionsabstractSUMMARY Existing specification‐based testing techniques often generate comprehensive test suites to cover diverse combinations of test‐relevant aspects. Such a test suite can be prohibitively expensive to execute exhaustively because of its large size. A pragmatic strategy often adopted in practice, called test‐once strategy, is to identify certain particular conditions from the specification and to test each such condition once only. This strategy is implicitly based on the uniformity assumption that the implementation will process a particular condition uniformly, regardless of other parameters or inputs. As the decision of adopting the test‐once strategy is often based on the specification, whether the uniformity assumption actually holds in the implementation needs to be critically assessed, or else the risk of inadequate testing could be non‐negligible. As viable alternatives to reduce such a risk, a family of test‐a‐few strategies for the testing of particular conditions is proposed in this paper. Two rounds of experiments that evaluate the effectiveness of the test‐a‐few strategies as compared with the test‐once strategy are further reported. Our experiments do the following: (1) provide clear evidence that the uniformity assumption often, but not always, holds and that the assumption usually fails to hold when the implementation is faulty; (2) demonstrate that all our proposed test‐a‐few strategies are statistically more reliable than the test‐once strategy in revealing faulty programs; (3) show that random sampling is already substantially more effective than the test‐once strategy; and (4) indicate that, compared with other test‐a‐few strategies under study, choice coverage seems to achieve a better trade‐off between test effort and effectiveness. Copyright © 2011 John Wiley & Sons, Ltd. Eric Ying Kwong Chan, Wing Kwong Chan, Pak-Lok Poon, Yuen-Tak Yu |
Softw. Pract. Exp. | 2 |
| 2011 | Precise Propagation of Fault-Failure Correlations in Program Flow GraphsabstractStatistical fault localization techniques find suspicious faulty program entities in programs by comparing passed and failed executions. Existing studies show that such techniques can be promising in locating program faults. However, coincidental correctness and execution crashes may make program entities indistinguishable in the execution spectra under study, or cause inaccurate counting, thus severely affecting the precision of existing fault localization techniques. In this paper, we propose a Block Rank technique, which calculates, contrasts, and propagates the mean edge profiles between passed and failed executions to alleviate the impact of coincidental correctness. To address the issue of execution crashes, Block Rank identifies suspicious basic blocks by modeling how each basic block contributes to failures by apportioning their fault relevance to surrounding basic blocks in terms of the rate of successful transition observed from passed and failed executions. Block Rank is empirically shown to be more effective than nine representative techniques on four real-life medium-sized programs. Zhenyu Zhang 0004, Wing Kwong Chan, T. H. Tse, Bo Jiang 0001 |
COMPSAC | 2 |
| 2011 | LOFT: Redundant Synchronization Event Removal for Data Race DetectionabstractMany happens-before based techniques for multithreaded programs implement vector clocks to track incrementally the causal relations among the synchronization operations acting on threads and locks. In these detectors, every such operation results in a vector-based assignment to a vector clock, even though the assigned value is the same as the value of the vector clock right before the assignment. The cost of such vector-based operations however grows with the number of threads and the amount of such operations. It is unclear to what extent redundant assignments can be removed. Whether two consecutive assignments to the same vector clock of a thread result in the same content critically depends on the operations on the locks occurred in between these assignments. In this paper, we systematically explore the said insight and quantify a sufficient condition that can soundly remove such operations without affecting the precision of such tracking. We applied our approach on Fast Track to formulate LOFT. We evaluate LOFT using the PARSEC benchmarking suite. The result shows that, on average, LOFT removes 58.0% of all such operations incurred by Fast Track, and runs 16.2% faster than the latter in tracking the causal relations among these operations. Yan Cai 0001, Wing Kwong Chan |
ISSRE | 2 |
| 2011 | A web search-centric approach to recommender systems with URLs as minimal user contexts
Wing Kwong Chan, Yuen Yau Chiu, Yuen-Tak Yu |
J. Syst. Softw. | 1 |
| 2011 | XML-manipulating test case prioritization for XML-manipulating services
Lijun Mei, Wing Kwong Chan, T. H. Tse, Robert G. Merkel |
J. Syst. Softw. | 2 |
| 2011 | Non-parametric statistical fault localization
Zhenyu Zhang 0004, Wing Kwong Chan, T. H. Tse, Yuen-Tak Yu, Peifeng Hu |
J. Syst. Softw. | 2 |
| 2011 | Assuring the model evolution of protocol software specifications by regression testing process improvementabstractSUMMARY Model‐based testing helps test engineers automate their testing tasks so that they are more cost‐effective. When the model is changed because of the evolution of the specification, it is important to maintain the test suites up to date for regression testing. A complete regeneration of the whole test suite from the new model, although inefficient, is still frequently used in the industry, including Microsoft. To handle specification evolution effectively, we propose a test case reusability analysis technique to identify reusable test cases of the original test suite based on graph analysis. We also develop a test suite augmentation technique to generate new test cases to cover the change‐related parts of the new model. The experiment on four large protocol document testing projects shows that our technique can successfully identify a high percentage of reusable test cases and generate low‐redundancy new test cases. When compared with a complete regeneration of the whole test suite, our technique significantly reduces regression testing time while maintaining the stability of requirement coverage over the evolution of requirements specifications. Copyright © 2011 John Wiley & Sons, Ltd. Bo Jiang 0001, T. H. Tse, Wolfgang Grieskamp, Nicolas Kicillof, Wing Kwong Chan |
Softw. Pract. Exp. | 7 |
| 2011 | Introduction to the Special Issue for the 10th International Conference on Quality Software (QSIC 2010)abstractEditorial for the special issue of Software: Practice and Experience which consists of six papers that are extended from the best papers of the 10th International Conference on Quality Software (QSIC 2010), Zhangjiajie, China, 14-15 July 2010. Ji Wang 0001, Wing Kwong Chan, Fei-Ching Kuo |
Softw. Pract. Exp. | 2 |
| 2011 | Guest editors' introduction to the special section on exploring the boundaries of software test automation
Christof J. Budnik, Wing Kwong Chan, Gregory M. Kapfhammer, Hong Zhu 0002 |
Softw. Qual. J. | 2 |
| 2010 | Bridging the Gap Between the Theory and Practice of Software Test AutomationabstractIn software development practice, testing often accounts for as much as 50% of the total development effort. It is therefore imperative to reduce the cost and improve the effectiveness of software testing by automating the testing process. In the past decades, a substantial amount of research effort has been invested into the development and study of automatic test case generation, automatic test oracles, and other (semi-)automated testing techniques. As the theory and practice of software testing becomes more mature, a deeper and more meaningful automation of the testing process is possible. Therefore, the automation of various testing activities is now becoming an integral part of industrial practice. In response to and in support of these exciting developments, the 5th Workshop on the Automation of Software Test provides a publication forum that bridges the gap between the theory and practice of automated testing. Christof J. Budnik, Wing Kwong Chan, Gregory M. Kapfhammer |
ICSE (2) | 2 |
| 2010 | Detecting atomic-set serializability violations in multithreaded programs through active randomized testingabstractConcurrency bugs are notoriously difficult to detect because there can be vast combinations of interleavings among concurrent threads, yet only a small fraction can reveal them. Atomic-set serializability characterizes a wide range of concurrency bugs, including data races and atomicity violations. In this paper, we propose a two-phase testing technique that can effectively detect atomic-set serializability violations. In Phase I, our technique infers potential violations that do not appear in a concrete execution and prunes those interleavings that are violation-free. In Phase II, our technique actively controls a thread scheduler to enumerate these potential scenarios identified in Phase I to look for real violations. We have implemented our technique as a prototype system AssetFuzzer and applied it to a number of subject programs for evaluating concurrency defect analysis techniques. The experimental results show that AssetFuzzer can identify more concurrency bugs than two recent testing tools RaceFuzzer and AtomFuzzer. Zhifeng Lai, Shing-Chi Cheung, Wing Kwong Chan |
ICSE (1) | 3 |
| 2010 | Taking Advantage of Service Selection: A Study on the Testing of Location-Based Web Services Through Test Case PrioritizationabstractDynamic service compositions pose new verification and validation challenges such as uncertainty in service membership. Moreover, applying an entire test suite to loosely coupled services one after another in the same composition can be too rigid and restrictive. In this paper, we investigate the impact of service selection on service-centric testing techniques. Specifically, we propose to incorporate service selection in executing a test suite and develop a suite of metrics and test case prioritization techniques for the testing of location-aware services. A case study shows that a test case prioritization technique that incorporates service selection can outperform their traditional counterpart - the impact of service selection is noticeable on software engineering techniques in general and on test case prioritization techniques in particular. Further-more, we find that points-of-interest-aware techniques can be significantly more effective than input-guided techniques in terms of the number of invocations required to expose the first failure of a service composition. Ke Zhai 0002, Bo Jiang 0001, Wing Kwong Chan, T. H. Tse |
ICWS | 3 |
| 2010 | Fault localization through evaluation sequences
Zhenyu Zhang 0004, Bo Jiang 0001, Wing Kwong Chan, T. H. Tse |
J. Syst. Softw. | 3 |
| 2010 | Finding failures from passed test cases: improving the pattern classification approach to the testing of mesh simplification programsabstractAbstract Mesh simplification programs create three‐dimensional polygonal models similar to an original polygonal model, and yet use fewer polygons. They produce different graphics even though they are based on the same original polygonal model. This results in a test oracle problem. To address the problem, our previous work has developed a technique that uses a reference model of the program under test to train a classifier. Using such an approach may mistakenly mark a failure‐causing test case as passed. It lowers the testing effectiveness of revealing failures. This paper suggests piping the test cases marked as passed by a statistical pattern classification module to an analytical metamorphic testing (MT) module. We evaluate our approach empirically using three subject programs with over 2700 program mutants. The result shows that, using a resembling reference model to train a classifier, the integrated approach can significantly improve the failure detection effectiveness of the pattern classification approach. We also explain how MT in our design trades specificity for sensitivity. Copyright © 2009 John Wiley & Sons, Ltd. Wing Kwong Chan, Jeffrey C. F. Ho, T. H. Tse |
Softw. Test. Verification Reliab. | 1 |
| 2010 | Partial constraint checking for context consistency in pervasive computingabstractPervasive computing environments typically change frequently in terms of available resources and their properties. Applications in pervasive computing use contexts to capture these changes and adapt their behaviors accordingly. However, contexts available to these applications may be abnormal or imprecise due to environmental noises. This may result in context inconsistencies, which imply that contexts conflict with each other. The inconsistencies may set such an application into a wrong state or lead the application to misadjust its behavior. It is thus desirable to detect and resolve the context inconsistencies in a timely way. One popular approach is to detect context inconsistencies when contexts breach certain consistency constraints. Existing constraint checking techniques recheck the entire expression of each affected consistency constraint upon context changes. When a changed context affects only a constraint's subexpression, rechecking the entire expression can adversely delay the detection of other context inconsistencies. This article proposes a rigorous approach to identifying the parts of previous checking results that are reusable without entire rechecking. We evaluated our work on the Cabot middleware through both simulation experiments and a case study. The experimental results reported that our approach achieved over a fifteenfold performance improvement on context inconsistency detection than conventional approaches. Chang Xu 0001, Shing-Chi Cheung, Wing Kwong Chan, Chunyang Ye |
ACM Trans. Softw. Eng. Methodol. | 3 |
| 2009 | Modeling and testing of cloud applicationsabstractWhat is a cloud application precisely? In this paper, we formulate a computing cloud as a kind of graph, a computing resource such as services or intellectual property access rights as an attribute of a graph node, and the use of a resource as a predicate on an edge of the graph. We also propose to model cloud computation semantically as a set of paths in a subgraph of the cloud such that every edge contains a predicate that is evaluated to be true. Finally, we present algorithms to compose cloud computations and a family of model-based testing criteria to support the testing of cloud applications. Wing Kwong Chan, Lijun Mei, Zhenyu Zhang 0004 |
APSCC | 1 |
| 2009 | Taming coincidental correctness: Coverage refinement with context patterns to improve fault localizationabstractRecent techniques for fault localization leverage code coverage to address the high cost problem of debugging. These techniques exploit the correlations between program failures and the coverage of program entities as the clue in locating faults. Experimental evidence shows that the effectiveness of these techniques can be affected adversely by coincidental correctness, which occurs when a fault is executed but no failure is detected. In this paper, we propose an approach to address this problem. We refine code coverage of test runs using control- and data-flow patterns prescribed by different fault types. We conjecture that this extra information, which we call context patterns, can strengthen the correlations between program failures and the coverage of faulty program entities, making it easier for fault localization techniques to locate the faults. To evaluate the proposed approach, we have conducted a mutation analysis on three real world programs and cross-validated the results with real faults. The experimental results consistently show that coverage refinement is effective in easing the coincidental correctness problem in fault localization techniques. Shing-Chi Cheung, Wing Kwong Chan, Zhenyu Zhang 0004 |
ICSE | 3 |
| 2009 | Adaptive Random Test Case PrioritizationabstractRegression testing assures changed programs against unintended amendments. Rearranging the execution order of test cases is a key idea to improve their effectiveness. Paradoxically, many test case prioritization techniques resolve tie cases using the random selection approach, and yet random ordering of test cases has been considered as ineffective. Existing unit testing research unveils that adaptive random testing (ART) is a promising candidate that may replace random testing (RT). In this paper, we not only propose a new family of coverage-based ART techniques, but also show empirically that they are statistically superior to the RT-based technique in detecting faults. Furthermore, one of the ART prioritization techniques is consistently comparable to some of the best coverage-based prioritization techniques (namely, the "additional" techniques) and yet involves much less time cost. Bo Jiang 0001, Zhenyu Zhang 0004, Wing Kwong Chan, T. H. Tse |
ASE | 3 |
| 2009 | Weaving Context Sensitivity into Test Suite ConstructionabstractContext-aware applications capture environmental changes as contexts and self-adapt their behaviors dynamically. Existing testing research has not explored context evolutions or their patterns inherent to individual test cases when constructing test suites. We propose the notation of context diversity as a metric to measure how many changes in contextual values of individual test cases. In this paper, we discuss how this notion can be incorporated in a test case generation process by pairing it with coverage-based test data selection criteria. Huai Wang, Wing Kwong Chan |
ASE | 2 |
| 2009 | Data flow testing of service choreographyabstractService computing has increasingly been adopted by the industry, developing business applications by means of orchestration and choreography. Choreography specifies how services collaborate with one another by defining, say, the message exchange, rather than via the process flow as in the case of orchestration. Messages sent from one service to another may require the use of different XPaths to manipulate or extract message contents. Mismatches in XML manipulations through XPaths (such as to relate incoming and outgoing messages in choreography specifications) may result in failures. In this paper, we propose to associate XPath Rewriting Graphs (XRGs), a structure that relates XPath and XML schema, with actions of choreography applications that are skeletally modeled as labeled transition systems. We develop the notion of XRG patterns to capture how different XRGs are related even though they may refer to different XML schemas or their tags. By applying XRG patterns, we successfully identify new data flow associations in choreography applications and develop new data flow testing criteria. Finally, we report an empirical case study that evaluates our techniques. The result shows our techniques are promising in detecting failures in choreography applications. Lijun Mei, Wing Kwong Chan, T. H. Tse |
ESEC/SIGSOFT FSE | 2 |
| 2009 | Capturing propagation of infected program statesabstractCoverage-based fault-localization techniques find the fault-related positions in programs by comparing the execution statistics of passed executions and failed executions. They assess the fault suspiciousness of individual program entities and rank the statements in descending order of their suspiciousness scores to help identify faults in programs. However, many such techniques focus on assessing the suspiciousness of individual program entities but ignore the propagation of infected program states among them. In this paper, we use edge profiles to represent passed executions and failed executions, contrast them to model how each basic block contributes to failures by abstractly propagating infected program states to its adjacent basic blocks through control flow edges. We assess the suspiciousness of the infected program states propagated through each edge, associate basic blocks with edges via such propagation of infected program states, calculate suspiciousness scores for each basic block, and finally synthesize a ranked list of statements to facilitate the identification of program faults. We conduct a controlled experiment to compare the effectiveness of existing representative techniques with ours using standard bench-marks. The results are promising. Zhenyu Zhang 0004, Wing Kwong Chan, T. H. Tse, Bo Jiang 0001 |
ESEC/SIGSOFT FSE | 2 |
| 2009 | Where to adapt dynamic service compositionsabstractPeer services depend on one another to accomplish their tasks, and their structures may evolve. A service composition may be designed to replace its member services whenever the quality of the composite service fails to meet certain quality-of-service (QoS) requirements. Finding services and service invocation endpoints having the greatest impact on the quality are important to guide subsequent service adaptations. This paper proposes a technique that samples the QoS of composite services and continually analyzes them to identify artifacts for service adaptation. The preliminary results show that our technique has the potential to effectively find such artifacts in services. Bo Jiang 0001, Wing Kwong Chan, Zhenyu Zhang 0004, T. H. Tse |
WWW | 2 |
| 2009 | Test case prioritization for regression testing of service-oriented business applicationsabstractRegression testing assures the quality of modified service-oriented business applications against unintended changes. However, a typical regression test suite is large in size. Earlier execution of those test cases that may detect failures is attractive. Many existing prioritization techniques order test cases according to their respective coverage of program statements in a previous version of the application. On the other hand, industrial service-oriented business applications are typically written in orchestration languages such as WS-BPEL and integrated with workflow steps and web services via XPath and WSDL. Faults in these artifacts may cause the application to extract wrong data from messages, leading to failures in service compositions. Surprisingly, current regression testing research hardly considers these artifacts. We propose a multilevel coverage model to capture the business process, XPath, and WSDL from the perspective of regression testing. We develop a family of test case prioritization techniques atop the model. Empirical results show that our techniques can achieve significantly higher rates of fault detection than existing techniques. Lijun Mei, Zhenyu Zhang 0004, Wing Kwong Chan, T. H. Tse |
WWW | 3 |
| 2009 | Is non-parametric hypothesis testing model robust for statistical fault localization?
Zhenyu Zhang 0004, Wing Kwong Chan, T. H. Tse, Peifeng Hu |
Inf. Softw. Technol. | 2 |
| 2009 | PAT: A pattern classification approach to automatic reference oracles for the testing of mesh simplification programs
Wing Kwong Chan, Shing-Chi Cheung, Jeffrey C. F. Ho, T. H. Tse |
J. Syst. Softw. | 1 |
| 2009 | Resource prioritization of code optimization techniques for program synthesis of wireless sensor network applications
Zhenyu Zhang 0004, Wing Kwong Chan, T. H. Tse, Heng Lu 0001, Lijun Mei |
J. Syst. Softw. | 2 |
| 2009 | Atomicity Analysis of Service Composition across OrganizationsabstractAtomicity is a highly desirable property for achieving application consistency in service compositions. To achieve atomicity, a service composition should satisfy the atomicity sphere, a structural criterion for the backend processes of involved services. Existing analysis techniques for atomicity sphere generally assume complete knowledge of all involved backend processes. Such an assumption is invalid when some service providers do not release all details of their backend processes to service consumers outside the organizations. To address this problem, we propose a process algebraic framework to publish atomicity-equivalent public views from the backend processes. These public views extract relevant task properties and reveal only partial process details that service providers need to expose. Our framework enables the analysis of atomicity sphere for service compositions using these public views instead of their backend processes. This allows service consumers to choose suitable services such that their composition satisfies the atomicity sphere without disclosing the details of their backend processes. Based on the theoretical result, we present algorithms to construct atomicity-equivalent public views and to analyze the atomicity sphere for a service composition. Two case studies from supply chain and insurance domains are given to evaluate our proposal and demonstrate the applicability of our approach. Chunyang Ye, Shing-Chi Cheung, Wing Kwong Chan, Chang Xu 0001 |
IEEE Trans. Software Eng. | 3 |
| 2008 | A Tale of Clouds: Paradigm Comparisons and Some Thoughts on Research IssuesabstractCloud computing is an emerging computing paradigm. It aims to share data, calculations, and services transparently among users of a massive grid. Although the industry has started selling cloud-computing products, research challenges in various areas, such as UI design, task decomposition, task distribution, and task coordination, are still unclear. Therefore, we study the methods to reason and model cloud computing as a step toward identifying fundamental research questions in this paradigm. In this paper, we compare cloud computing with service computing and pervasive computing. Both the industry and research community have actively examined these three computing paradigms. We draw a qualitative comparison among them based on the classic model of computer architecture. We finally evaluate the comparison results and draw up a series of research questions in cloud computing for future exploration. Lijun Mei, Wing Kwong Chan, T. H. Tse |
APSCC | 2 |
| 2008 | Debugging through Evaluation Sequences: A Controlled Experimental StudyabstractPredicate-based statistical fault-localization techniques locate fault-relevant predicates in a program by contrasting the statistics of the values of individual predicates between successful and failure-causing runs. While short-circuit evaluations are common in program execution, treating predicates as atomic units ignores this fact, masking out various types of important statistics. On the contrary, are such statistics useful for debugging? In this paper, we investigate experimentally the impact of the use of short-circuit evaluation information on fault localization. The results show that, by doing so, it significantly improves predicate-based statistical fault-localization techniques. Zhenyu Zhang 0004, Bo Jiang 0001, Wing Kwong Chan, T. H. Tse |
COMPSAC | 3 |
| 2008 | Heuristics-Based Strategies for Resolving Context Inconsistencies in Pervasive Computing ApplicationsabstractContext-awareness allows pervasive applications to adapt to changeable computing environments. Contexts, the pieces of information that capture the characteristics of environments, are often error-prone and inconsistent due to noises. Various strategies have been proposed to enable automatic context inconsistency resolution. They are formulated on different assumptions that may not hold in practice. This causes applications to be less context-aware to different extents. In this paper, we investigate such impacts and propose our new resolution strategy. We conducted experiments to compare our work with major existing strategies. The results showed that our strategy is both effective in resolving context inconsistencies and promising in its support of applications using contexts. Chang Xu 0001, Shing-Chi Cheung, Wing Kwong Chan, Chunyang Ye |
ICDCS | 3 |
| 2008 | Testing pervasive software in the presence of context inconsistency resolution servicesabstractPervasive computing software adapts its behavior according to the changing contexts. Nevertheless, contexts are often noisy. Context inconsistency resolution provides a cleaner pervasive computing environment to context-aware applications. A faulty context-aware application may, however, mistakenly mix up inconsistent contexts and resolved ones, causing incorrect results. This paper studies how such faulty context-aware applications may be affected by these services. We model how programs should handle contexts that are continually checked and resolved by context inconsistency resolution, develop novel sets of data flow equations to analyze the potential impacts, and thus formulate a new family of test adequacy criteria for testing these applications. Experimentation shows that our approach is promising. Heng Lu 0001, Wing Kwong Chan, T. H. Tse |
ICSE | 2 |
| 2008 | Data flow testing of service-oriented workflow applicationsabstractWS-BPEL applications are a kind of service-oriented application. They use XPath extensively to integrate loosely-coupled workflow steps. However, XPath may extract wrong data from the XML messages received, resulting in erroneous results in the integrated process. Surprisingly, although XPath plays a key role in workflow integration, inadequate researches have been conducted to address the important issues in software testing. This paper tackles the problem. It also demonstrates a novel transformation strategy to construct artifacts. We use the mathematical definitions of XPath constructs as rewriting rules, and propose a data structure called XPath Rewriting Graph (XRG), which not only models how an XPath is conceptually rewritten but also tracks individual rewritings progressively. We treat the mathematical variables in the applied rewriting rules as if they were program variables, and use them to analyze how information may be rewritten in an XPath conceptually. We thus develop an algorithm to construct XRGs and a novel family of data flow testing criteria to test WS-BPEL applications. Experiment results show that our testing approach is promising. Lijun Mei, Wing Kwong Chan, T. H. Tse |
ICSE | 2 |
| 2008 | An Adaptive Service Selection Approach to Service CompositionabstractIn service computing, the behavior of a service may evolve. When an organization develops a service-oriented application in which certain services are provided by external partners, the organization should address the problem of uninformed behavior evolution of external services. This paper proposes an adaptive framework that bars problematic external services to be used in the service-oriented application of an organization. We use dynamic WSDL information in public service registries to approximate a snapshot of a network of services, and apply link analysis on the snapshot to identify services that are popularly used by different service consumers at the moment. As such, service composition can be strategically formed using the highly referenced services. We evaluate our proposal through a simulation study. The results show that, in terms of the number of failures experienced by service consumers, our proposal significantly outperforms the random approach in selecting reliable services to form service compositions. Lijun Mei, Wing Kwong Chan, T. H. Tse |
ICWS | 2 |
| 2008 | Inter-context control-flow and data-flow test adequacy criteria for nesC applicationsabstractNesC is a programming language for applications that run on top of networked sensor nodes. Such an application mainly uses an interrupt to trigger a sequence of operations, known as contexts, to perform its actions. However, a high degree of inter-context interleaving in an application can cause it to be error-prone. For instance, a context may mistakenly alter another context's data kept at a shared variable. Existing concurrency testing techniques target testing programs written in general-purpose programming languages, where a small scale of inter-context interleaving between program executions may make these techniques inapplicable. We observe that nesC blocks new context interleaving when handling interrupts, and this feature significantly restricts the scale of inter-context interleaving that may occur in a nesC application. This paper models how operations on different contexts may interleave as inter-context flow graphs. Based on these graphs, it proposes two test adequacy criteria, one on inter-context data-flows and another on inter-context control-flows. It evaluates the proposal by a real-life open-source nesC application. The empirical results show that the new criteria detect significantly more failures than their conventional counterparts. Zhifeng Lai, Shing-Chi Cheung, Wing Kwong Chan |
SIGSOFT FSE | 3 |
| 2008 | Enhancing adaptive random testing for programs with high dimensional input domains or failure-unrelated parameters
Fei-Ching Kuo, Tsong Yueh Chen, Huai Liu, Wing Kwong Chan |
Softw. Qual. J. | 4 |
| 2007 | Piping Classification to Metamorphic Testing: An Empirical Study towards Better Effectiveness for the Identification of Failures in Mesh Simplification ProgramsabstractMesh simplification is a mainstream technique to render graphics responsively in modern graphical software. However, the graphical nature of the output poses a test oracle problem in testing. Previous work uses pattern classification to identify failures. Although such an approach may be promising, it may conservatively mark the test result of a failure-causing test case as passed. This paper proposes a methodology that pipes the test cases marked as passed by the pattern classification component to a metamorphic testing component to look for missed failures. The empirical study uses three simple and general metamorphic relations as subjects, and the experimental results show a 10 percent improvement of effectiveness in the identification of failures. Wing Kwong Chan, Jeffrey C. F. Ho, T. H. Tse |
COMPSAC (1) | 1 |
| 2007 | Do Maintainers Utilize Deployed Design Patterns Effectively?abstractOne claimed benefit of deploying design patterns is facilitating maintainers to perform anticipated changes. However, it is not at all obvious that the relevant design patterns deployed in software will invariably be utilized for the changes. Moreover, we observe that many well-known design patterns consist of three types of programming elements (called participants), and that performing an anticipated change typically entails multiple tasks related to different types of participants. This paper studies empirically whether maintainers utilize deployed design patterns, and when they do, which tasks they more commonly perform. Our experiments show that almost all subjects perform the task of adding new concrete participants, fewer perform the tasks involving clients, whereas even fewer perform the tasks involving abstract participants. Furthermore, utilizing deployed design patterns (by performing whichever of the corresponding tasks) is found to be statistically associated with the delivery of less faulty codes. Tsz Hin Ng, Shing-Chi Cheung, Wing Kwong Chan, Yuen-Tak Yu |
ICSE | 3 |
| 2007 | On impact-oriented automatic resolution of pervasive context inconsistencyabstractContext-awareness is a capability that allows applications in pervasive computing to adapt themselves continuously to changing contexts of their environments. However, contexts from physical environments may be inconsistent. It affects the correctness of these applications. Existing resolution strategies for context inconsistency have diverse adverse impacts on the context awareness of applications, such as feeding different amounts of contexts to the applications. In this paper, we examine the impacts of inconsistency resolution and study the extent to which their effects on context-awareness can be reduced. We conduct simulation experiments of two pervasive computing applications. The experimental results show that existing inconsistency resolution strategies adversely affect the context-awareness of applications. This motivates the importance of deploying an impact-oriented approach to respect context-awareness in inconsistency resolution. Chang Xu 0001, Shing-Chi Cheung, Wing Kwong Chan, Chunyang Ye |
ESEC/SIGSOFT FSE | 3 |
| 2007 | Detection and resolution of atomicity violation in service compositionabstractAtomicity is a desirable property that safeguards application consistency for service compositions. A service composition exhibiting this property could either complete or cancel itself without any side effects. It is possible to achieve this property for a service composition by selecting suitable web services to form an atomicity sphere. However, this property might still be breached at runtime due to the interference between various service compositions caused by implicit interactions. Existing approaches to addressing this problem by restricting concurrent execution of services to avoid all implicit interactions however compromise the performance of service compositions due to the long running nature of web services. In this paper, we propose a novel static approach to analyzing the implicit interactions a web service may incur and their impacts on the atomicity property in each of its service compositions. By locating afflicted implicit interactions in a service composition, behavior constraints based on property propagation are formulated as local safety properties, which can then be enforced by the affected web services at runtime to suppress the impacts of the afflicted implicit interactions. We show that the satisfaction of these safety properties exempts the atomicity property of this service composition from being interfered by other services at runtime. The approach is illustrated using two service applications. Chunyang Ye, Shing-Chi Cheung, Wing Kwong Chan, Chang Xu 0001 |
ESEC/SIGSOFT FSE | 3 |
| 2006 | Reference Models and Automatic Oracles for the Testing of Mesh Simplification Software for Graphics RenderingabstractSoftware with graphics rendering is an important class of applications. Many of them use polygonal models to represent the graphics. Mesh simplification is a vital technique to vary the levels of object details and, hence, improve the overall performance of the rendering process. It progressively enhances the effectiveness of rendering from initial reference systems. As such, the quality of its implementation affects that of the associated graphics rendering application. Testing of mesh simplification is essential towards assuring the quality of the applications. Is it feasible to use the reference systems to serve as automated test oracles for mesh simplification programs? If so, how well are they useful for this purpose? We present a novel approach in this paper. We propose to use pattern classification techniques to address the above problem. We generate training samples from the reference system to test samples from the implementation. Our experimentation shows that the approach is promising Wing Kwong Chan, Shing-Chi Cheung, Jeffrey C. F. Ho, T. H. Tse |
COMPSAC (1) | 1 |
| 2006 | Incremental consistency checking for pervasive contextabstractApplications in pervasive computing are typically required to interact seamlessly with their changing environments. To provide users with smart computational services, these applications must be aware of incessant context changes in their environments and adjust their behaviors accordingly. As these environments are highly dynamic and noisy, context changes thus acquired could be obsolete, corrupted or inaccurate. This gives rise to the problem of context inconsistency, which must be timely detected in order to prevent applications from behaving anomalously. In this paper, we propose a formal model of incremental consistency checking for pervasive contexts. Based on this model, we further propose an efficient checking algorithm to detect inconsistent contexts. The performance of the algorithm and its advantages over conventional checking techniques are evaluated experimentally using Cabot middleware. Chang Xu 0001, Shing-Chi Cheung, Wing Kwong Chan |
ICSE | 3 |
| 2006 | Publishing and composition of atomicity-equivalent services for B2B collaborationabstractException handling resolves inconsistency by backward or forward error recovery methods or both in Business-to-Business (B2B) process collaboration. To avoid committing irrevocable tasks followed by exceptions, B2B processes, which guarantee the atomicity sphere property, are attractive. While atomicity sphere ensures its outcomes to be either all or nothing, conflicting local recoveries may lead to global B2B inconsistencies. Existing (global) analysis techniques however mandate every process unveiling all individual tasks. Such an analysis is infeasible when some business parties refuse to disclose their process details for privacy or business reasons. To address this problem, we propose a process algebraic technique to prove, construct, and check atomicity-equivalent public views from B2B processes. By checking atomicity spheres in the composition of these public views, business parties can identify suitable services that respect their individual and overall atomicity requirements. An example based on a real-life multilateral supply chain process is included. Chunyang Ye, Shing-Chi Cheung, Wing Kwong Chan |
ICSE | 3 |
| 2006 | Testing context-aware middleware-centric programs: a data flow approach and an RFID-based experimentationabstractPervasive context-aware software is an emerging kind of application. Many of these systems register parts of their context-aware logic in the middleware. On the other hand, most conventional testing techniques do not consider such kind of application logic. This paper proposes a novel family of testing criteria to measure the comprehensiveness of their test sets. It stems from context-aware data flow information. Firstly, it studies the evolution of contexts, which are environmental information relevant to an application program. It then proposes context-aware data flow associations and testing criteria. Corresponding algorithms are given. It uses a prototype testing tool to conduct experimentation on an RFID-based location sensing software running on top of context-aware middleware. The experimental results show that our approach is applicable, effective, and promising. Heng Lu 0001, Wing Kwong Chan, T. H. Tse |
SIGSOFT FSE | 2 |
| 2006 | Work experience versus refactoring to design patterns: a controlled experimentabstractProgram refactoring using design patterns is an attractive approach for facilitating anticipated changes. Its benefit depends on at least two factors, namely the effort involved in the refactoring and how effective it is. For example, the benefit would be small if too much effort is required to translate a program correctly into a refactorized form, and whether such a form could effectively guide maintainers to complete anticipated changes is unknown. A metric of effectiveness is the maintainers' performance, which can be affected by their work experience, in realizing the changes. Hence, an interesting question arises. Is program refactoring to introduce additional patterns beneficial regardless of the work experience of the maintainers? In this paper, we report a controlled experiment on maintaining JHotDraw, an open source system deployed with multiple patterns. We compared maintainers with and without work experience. Our empirical results show that, to complete a maintenance task of perfective nature, the time spent even by the inexperienced maintainers on a refactorized version is much shorter than that of the experienced subjects on the original version. Moreover, the quality of their delivered programs, in terms of correctness, is found to be comparable. Tsz Hin Ng, Shing-Chi Cheung, Wing Kwong Chan, Yuen-Tak Yu |
SIGSOFT FSE | 3 |
| 2006 | Local analysis of atomicity sphere for B2B collaborationabstractAtomicity is a desirable property for business processes to conduct transactions in Business-to-Business (B2B) collaboration. Although it is possible to reason about atomicity of B2B collaboration using the public views, yet such reasoning requires the presence of a trustworthy party who has complete knowledge of these views. It is inapplicable when some parties may want to keep the confidentiality of their collaborative partners for privacy and other business reasons, or the trustworthy party is not available. To address this problem, we propose a novel approach that allows each party to jointly conduct local atomicity checking with its direct partners. It is based on iterative forwarding and regression of compensability properties between each pair of direct partners. This approach is applied to a case study based on a real-life insurance process in the motor damage claims domain. Chunyang Ye, Shing-Chi Cheung, Wing Kwong Chan, Chang Xu 0001 |
SIGSOFT FSE | 3 |
| 2006 | Integration Testing of Context-sensitive Middleware-based Applications: a Metamorphic ApproachabstractDuring the testing of context-sensitive middleware-based software, the middleware checks the current situation to invoke the appropriate functions of the applications. Since the middleware remains active and the situation may continue to evolve, however, the conclusion of some test cases may not easily be identified. Moreover, failures appearing in one situation may be superseded by subsequent correct outcomes and, therefore, be hidden. We alleviate the above problems by making use of a special kind of situation, which we call checkpoints, such that the middleware will not activate the functions under test. We recommend testers to generate test cases that start at a checkpoint and end at another. Testers may identify relations that associate different execution sequences of a test case. They then check the results of each test case to detect any contravention of such relations. We illustrate our technique with an example that shows how hidden failures can be detected. We also report the experimentation carried out on an RFID-based location-sensing application on top of a context-sensitive middleware. Wing Kwong Chan, Tsong Yueh Chen, Heng Lu 0001, T. H. Tse, Stephen S. Yau |
Int. J. Softw. Eng. Knowl. Eng. | 1 |
| 2006 | Automatic goal-oriented classification of failure behaviors for testing XML-based multimedia software applications: An experimental case study
Wing Kwong Chan, M. Y. Cheng, Shing-Chi Cheung, T. H. Tse |
J. Syst. Softw. | 1 |
| 2004 | Testing Context-Sensitive Middleware-Based Software ApplicationsabstractContext-sensitive middleware-based software is an emerging kind of ubiquitous computing application. The components of such software communicate proactively among themselves according to the situational attributes of their environments, known as the "contexts". The actual process of accessing and updating the contexts lies with the middleware. The latter invokes the relevant local and remote operations whenever any context inscribed in the situation-aware interface is satisfied. Since the applications operate in a highly dynamic environment, the testing of context-sensitive software is challenging. Metamorphic testing is a property-based testing strategy. It recommends that, even if a test case does not reveal any failure, follow-up test cases should be further constructed from the original to check whether the software satisfies some necessary conditions of the problem to be implemented. This work proposes to use isotropic properties of contexts as metamorphic relations for testing context-sensitive software. For instance, distinct points on the same isotropic curve of contexts would entail comparable responses by the components. This notion of testing context relations is novel, robust, and intuitive to users. T. H. Tse, Stephen S. Yau, Wing Kwong Chan, Heng Lu 0001, Tsong Yueh Chen |
COMPSAC | 3 |
| 2004 | The System of Vector Quasi-Equilibrium Problems with Applications
Qamrul Hasan Ansari, Wing Kwong Chan, X. Q. Yang |
J. Glob. Optim. | 2 |