Bo Jiang 0001

dblp:34/2005-1 · DBLP profile ↗
← Back
57ranked-venue papers
20as first author
18since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 38 · 17 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 6 first-author · 1 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 4 since 2021Systems, architecture and hardware · 4 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2Graphics, computer vision, multimedia, augmented reality and games · 1Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Integrating retinex theory with atmospheric scattering model for mixed-degradation image recovery
Yonglong Jiang, Bo Jiang 0001, Jiahe Zhu, Zelong Tan, Zehua Ji, Hongbing Ma
Expert Syst. Appl.2
2025 TODO: Enhancing LLM Alignment with Ternary Preferences
abstract
Aligning large language models (LLMs) with human intent is critical for enhancing their performance across a variety of tasks. Standard alignment techniques, such as Direct Preference Optimization (DPO), often rely on the binary Bradley-Terry (BT) model, which can struggle to capture the complexities of human preferences—particularly in the presence of noisy or inconsistent labels and frequent ties. To address these limitations, we introduce the Tie-rank Oriented Bradley-Terry model (TOBT), an extension of the BT model that explicitly incorporates ties, enabling more nuanced preference representation. Building on this, we propose Tie-rank Oriented Direct Preference Optimization (TODO), a novel alignment algorithm that leverages TOBT's ternary ranking system to improve preference alignment. In evaluations on Mistral-7B and Llama 3-8B models, TODO consistently outperforms DPO in modeling preferences across both in-distribution and out-of-distribution datasets. Additional assessments using MT Bench and benchmarks such as Piqa, ARC-c, and MMLU further demonstrate TODO's superior alignment performance. Notably, TODO also shows strong results in binary preference alignment, highlighting its versatility and potential for broader integration into LLM alignment. The code for TODO is made publicly available.
Yuxiang Guo 0004, Lu Yin 0006, Bo Jiang 0001
ICLR3
2024 Unishyper: A Rust-based unikernel enhancing reliability and efficiency of embedded systems
Keyang Hu, Wang Huang, Lei Wang 0126, Ce Mo, Runxiang Wang, Yu Chen 0004, Ju Ren 0001, Bo Jiang 0001
J. Syst. Archit.8
2023 Work-in-Progress: Unishyper, A Reliable Rust-based Unikernel for Embedded Scenarios
abstract
Unishyper is a reliable Rust-based unikernel with good performance for embedded scenarios. It relies on Rust Features to enable customization at fine granularity. To achieve high reliability, Unishyper makes full use of the Rust language features to reduce memory safety bugs, ensure safe resource management, and achieve fault handling and recovery. Finally, Unishyper achieves high performance through safe multi-threading model as well as the single-privilege-level and single-address-space design.
Keyang Hu, Lei Wang 0126, Ce Mo, Bo Jiang 0001
EMSOFT4
2023 A study on the impact of pre-trained model on Just-In-Time defect prediction
abstract
Previous researchers conducting Just-In-Time (JIT) defect prediction tasks have primarily focused on the performance of individual pre-trained models, without exploring the relationship between different pre-trained models as backbones. In this study, we build six models: RoBERTaJIT, CodeBERTJIT, BARTJIT, PLBARTJIT, GPT2JIT, and CodeGPTJIT, each with a distinct pre-trained model as its backbone. We systematically explore the differences and connections between these models. Specifically, we investigate the performance of the models when using Commit code and Commit message as inputs, as well as the relationship between training efficiency and model distribution among these six models. Additionally, we conduct an ablation experiment to explore the sensitivity of each model to inputs. Furthermore, we investigate how the models perform in zero-shot and few-shot scenarios. Our findings indicate that each model based on different backbones shows improvements, and when the backbone’s pre-training model is similar, the training resources that need to be consumed are closer. We also observe that Commit code plays a significant role in defect detection, and different pre-trained models demonstrate better defect detection ability with a balanced dataset under few-shot scenarios. These results provide new insights for optimizing JIT defect prediction tasks using pre-trained models and highlight the factors that require more attention when constructing such models. Additionally, CodeGPTJIT and GPT2JIT achieved better performance than DeepJIT and CC2Vec on the two datasets respectively under 2000 training samples. These findings emphasize the effectiveness of transformer-based pre-trained models in JIT defect prediction tasks, especially in scenarios with limited training data.
Yuxiang Guo 0004, Xiaopeng Gao, Zhenyu Zhang 0004, Wing Kwong Chan, Bo Jiang 0001
QRS5
2023 Aster: Encoding Data Augmentation Relations into Seed Test Suites for Robustness Assessment and Fuzzing of Data-Augmented Deep Learning Models
abstract
Data-augmented deep learning models are widely used in real-world applications. However, many state-of the-art loss-based or coverage-based fuzzing techniques fail to produce fuzzing samples for them from many seeds. This paper proposes Aster, a novel technique to address this problem to enhance their fuzzing effectiveness for deep learning models trained with multi-sample data augmentation methods. Aster formulates a novel reachability-based strategy to encode the insights of every seed’s direct and indirect data augmentation relation instances into the replacement seed of that seed systematically. Our experiment shows that Aster is highly effective. On average, loss-based and coverage-based fuzzing techniques can generate 166% and 110% more fuzzing samples and reduce 31% and 22% unsuccessful seeds, respectively, after adopting the replacement seeds generated by Aster to replace their original seeds. Their improved models also become up to 55% and 40% on average more robust against FGSM and PGD attacks in the experiment.
Haipeng Wang 0005, Zhengyuan Wei, Qilin Zhou, Bo Jiang 0001, Wing Kwong Chan
QRS4
2023 CAPS: An Efficient Whole-Program Critical Paths Search Framework for Large-Scale Software
abstract
Tracking the flow of external inputs in a program with taint-analysis techniques can help developers better identify potential security vulnerabilities in the software.However, directly using the static taint analysis provided by Clang Static Analyzer is inefficient for large-scale software due to the huge but redundant ExplodedGraph generated.Therefore, we propose an efficient Whole-Program Critical Paths Search (CAPS) framework.It first performs a set of optimizations to reduce the ExplodedGraph of each function.Then, it constructs a global exploded graph by inserting call edges among the reduced ExplodedGraphs for each function within the Neo4j graph database.Finally, it proposes loop removal and graph segmentation to optimize the search process for critical paths on the global exploded graph.Our experiments on 3 large-scale software show that CAPS can significantly improve the efficiency of critical path search for large-scale software.
Yuening Su, Bo Jiang 0001
SEKE5
2023 OAT: An Optimized Android Testing Framework Based on Reinforcement Learning
Mengjun Du, Lian Song, Wing Kwong Chan, Bo Jiang 0001
TASE5
2023 Rust-Shyper: A reliable embedded hypervisor supporting VM migration and hypervisor live-update
Ce Mo, Lei Wang 0126, Keyang Hu, Bo Jiang 0001
J. Syst. Archit.5
2023 Limon: A Scalable and Stable Key-Value Engine for Fast NVMe Devices
abstract
Modern fast NVMe devices with high throughput and ultra-low latency have brought new opportunities for persistent key-value (KV) engines. In this paper, we propose Limon, a persistent KV engine to exploit the performance potentials of fast NVMe devices. Limon targets three practical design aspects that existing KV engines fail to consider simultaneously: functionality, scalability, and stability. Limon carefully redesigns the index structure, on-disk KV record layout, and I/O processing of a persistent KV engine for these aspects. Specifically, Limon (i) proposes a semi-shared global index to improve scalability and range queries. (ii) employs a fast slab-based record layout with light-weight defragmentation to enable stable performance, and (iii) uses efficient asynchronous per-core I/O processing with two optimizations: DMA-backed buffer pool and page deduplication, to further improve scalability. Our evaluations with the YCSB benchmark and one production workload show that Limon outperforms state-of-the-art persistent key-value engines (i.e., SpanDB, KVell, and uDepot) by up to 1.2x to 3.8x and has the best scalability. Moreover, Limon has stable and predictable performance due to its novel record layout strategy.
Baoyue Yan, Jinbin Zhu, Bo Jiang 0001
IEEE Trans. Computers3
2022 Shyper: An embedded hypervisor applying hierarchical resource isolation strategies for mixed-criticality systems
abstract
With the development of the IoT, modern embedded systems are evolving to general-purpose and mixed-criticality systems, where virtualization has become the key to guarantee the isolation between tasks with different criticality. Traditional server-based hypervisors (KVM and Xen) are difficult to use in embedded scenarios due to performance and security reasons. As a result, several new hypervisors (Jailhouse and Bao) have been proposed in recent years, which effectively solve the problems above through static partitioning. However, this inflexible resource isolation strategy assumes no resource sharing across guests, which greatly reduces the resource utilization and VM scalability. This prevents themselves from simultaneously fulfilling the differentiated demands from VMs conducting different tasks. This paper proposes an efficient and real-time embedded hypervisor “Shyper”, aiming at providing differentiated services for VMs with different criticality. To achieve that, Shyper supports fine-grained hierarchical resource isolation strategies and introduces several novel “VM-Exit-less” real-time virtualization techniques, which grants users the flexibility to strike a trade-off between VM's resource utilization and real-time performance. In this paper, we also compare Shyper with other mainstream hypervisors (KVM, Jailhouse, etc.) to evaluate its feasibility and effectiveness.
Yicong Shen, Lei Wang 0126, Yuanzhi Liang, Bo Jiang 0001
DATE5
2022 Towards Group Fairness via Semi-Centralized Adversarial Training in Federated Learning
abstract
As federated learning increasingly performs better on tasks of decision-making scenarios such as medical care or commercial area, there have been concerns about discrimination against certain populations with sensitive attributes (e.g., race, gender). In this work, we propose to improve group fairness with semi-centralized adversarial training. And we adopt Variational AutoEncoder (VAE) for federated learning scenarios to generate adversarial samples. We keep VAE decoder at server side and leave encoder at client side to encode local samples into feature dimensions for transmitting, which ensures the privacy of user data. Our proposal further performs sensitive attribute alignment to improve group fairness. Our experimental evaluation shows that our approach outperforms the state-of-the-art federated learning frameworks in terms of group fairness and communication resource consumption.
Yurui Yang, Bo Jiang 0001
MDM2
2022 WasmFuzzer: A Fuzzer for WasAssembly Virtual Machines
abstract
WebAssembly is a fast, safe, and portable low-level language suitable for diverse application scenarios.And The WebAssembly virtual machines are widely used by Web browsers or Blockchain platforms as execution engine.When there is a bug in the implementation of the Wasm virtual machine, the execution of WebAssembly may lead to errors or vulnerability in the application.Due to the grammar checks by WASM VMs, fuzzing at the binary level is ineffective to expose the bugs because most inputs cannot reach the deep logic within the WASM VM.In this work, we propose WasmFuzzer, a bytecode level fuzzing tool for WASM VMs.WasmFuzzer proposes to generate initial seeds for Fuzzing at the Wasm bytecode level and it also designs a systematic set of mutation operators for Wasm bytecode.Furthermore, WasmFuzzer proposes an adaptive mutation strategy to search for the best mutation operators for different fuzzing targets.Our evaluation on 3 real-life Wasm VMs shows that WasmFuzzer can significantly outperform AFL in terms of both code coverage and unique crash.
Bo Jiang 0001, Zichao Li 0008, Yuhe Huang, Zhenyu Zhang 0004, Wing Kwong Chan
SEKE1
2022 VM Migration and Live-Update for Reliable Embedded Hypervisor
Lei Wang 0126, Keyang Hu, Ce Mo, Bo Jiang 0001
SETTA5
2021 WANA: Symbolic Execution of Wasm Bytecode for Extensible Smart Contract Vulnerability Detection
abstract
Many popular blockchain platforms support smart contracts for building decentralized applications. However, the vulnerabilities within smart contracts have demonstrated to lead to serious financial loss to their end users. In particular, the smart contracts on EOSIO smart contract platform have resulted in the loss of around 380K EOS tokens, which was around 1.9 million worth of USD at the time of attack. The EOSIO smart contract platform is based on the Wasm VM, which is also the underlying system supporting other smart contract platforms as well as Web application. In this work, we present WANA, an extensible smart contract vulnerability detection tool based on the symbolic execution for Wasm bytecode. WANA proposes a set of algorithms to detect the vulnerabilities in EOSIO smart contracts based on Wasm bytecode analysis. Our experimental analysis shows that WANA can effectively and efficiently detect vulnerabilities in EOSIO smart contracts. Furthermore, our case study also demonstrates that WANA can be extended to effectively detect vulnerabilities in Ethereum smart contracts.
Bo Jiang 0001, Imran Ashraf 0001, Wing Kwong Chan
QRS1
2021 DroidGamer: Android Game Testing with Operable Widget Recognition by Deep Learning
abstract
Android game applications are an important type of application widely used by end users. Bugs in such applications can significantly affect user experience. Due to the use of rendered Graphical User Interface (GUI) widgets, automated testing of Android game application becomes challenging because such GUI widgets cannot be queried with Android system APIs, making existing GUI-based testing techniques blind to the locations of widgets. In this work, we propose DroidGamer, a novel GUI traversal-based Android game testing technique, which relies on deep learning models to recognize operable GUI widgets. DroidGamer adopts a novel GUI model traversal algorithm and a new GUI state equivalence criterion over the widget recognition results of the deep learning models. Our experiment on 10 open-source Android games shows that DroidGamer is significantly more effective than existing techniques including Monkey, Stoat, and PUMA for testing Android games in terms of both code coverage and fault detection ability.
Bo Jiang 0001, Wenlin Wei, Wing Kwong Chan
QRS1
2021 Revisiting the Design of LSM-tree Based OLTP Storage Engine with Persistent Memory
abstract
The recent byte-addressable and large-capacity commercialized persistent memory (PM) is promising to drive database as a service (DBaaS) into unchartered territories. This paper investigates how to leverage PMs to revisit the conventional LSM-tree based OLTP storage engines designed for DRAM-SSD hierarchy for DBaaS instances. Specifically we (1) propose a light-weight PM allocator named Hal-loc customized for LSM-tree, (2) build a high-performance Semi-persistent Memtable utilizing the persistent in-memory writes of PM, (3) design a concurrent commit algorithm named Reorder Ring to aschieve log-free transaction processing for OLTP workloads and (4) present a Global Index as the new globally sorted persistent level with non-blocking in-memory compaction. The design of Reorder Ring and Semi-persistent Memtable achieves fast writes without synchronized logging overheads and achieves near instant recovery time. Moreover, the design of Semi-persistent Memtable and Global Index with in-memory compaction enables the byte-addressable persistent levels in PM, which significantly reduces the read and write amplification as well as the background compaction overheads. The overall evaluation shows that the performance of our proposal over PM-SSD hierarchy outperforms the baseline by up to 3.8x in YCSB benchmark and by 2x in TPC-C benchmark.
Baoyue Yan, Xuntao Cheng, Bo Jiang 0001, Shibin Chen, Canfang Shang, Kenry Huang, Xinjun Yang, Wei Cao 0006, Feifei Li 0001
Proc. VLDB Endow.3
2021 RegionTrack: A Trace-Based Sound and Complete Checker to Debug Transactional Atomicity Violations and Non-Serializable Traces
abstract
Atomicity is a correctness criterion to reason about isolated code regions in a multithreaded program when they are executed concurrently. However, dynamic instances of these code regions, called transactions , may fail to behave atomically, resulting in transactional atomicity violations. Existing dynamic online atomicity checkers incur either false positives or false negatives in detecting transactions experiencing transactional atomicity violations. This article proposes RegionTrack. RegionTrack tracks cross-thread dependences at the event, dynamic subregion, and transaction levels. It maintains both dynamic subregions within selected transactions and transactional happens-before relations through its novel timestamp propagation approach. We prove that RegionTrack is sound and complete in detecting both transactional atomicity violations and non-serializable traces. To the best of our knowledge, it is the first online technique that precisely captures the transitively closed set of happens-before relations over all conflicting events with respect to every running transaction for the above two kinds of issues. We have evaluated RegionTrack on 19 subjects of the DaCapo and the Java Grande Forum benchmarks. The empirical results confirm that RegionTrack precisely detected all those transactions which experienced transactional atomicity violations and identified all non-serializable traces. The overall results also show that RegionTrack incurred 1.10x and 1.08x lower memory and runtime overheads than Velodrome and 2.10x and 1.21x lower than Aerodrome, respectively. Moreover, it incurred 2.89x lower memory overhead than DoubleChecker. On average, Velodrome detected about 55% fewer violations than RegionTrack, which in turn reported about 3%–70% fewer violations than DoubleChecker.
Shangru Wu, Ernest Bota Pobee, Xiupei Mei, Hao Zhang 0085, Bo Jiang 0001, Wing Kwong Chan
ACM Trans. Softw. Eng. Methodol.6
2020 Generative Attention Networks for Multi-Agent Behavioral Modeling
abstract
Understanding and modeling behavior of multi-agent systems is a central step for artificial intelligence. Here we present a deep generative model which captures behavior generating process of multi-agent systems, supports accurate predictions and inference, infers how agents interact in a complex system, as well as identifies agent groups and interaction types. Built upon advances in deep generative models and a novel attention mechanism, our model can learn interactions in highly heterogeneous systems with linear complexity in the number of agents. We apply this model to three multi-agent systems in different domains and evaluate performance on a diverse set of tasks including behavior prediction, interaction analysis and system identification. Experimental results demonstrate its ability to model multi-agent systems, yielding improved performance over competitive baselines. We also show the model can successfully identify agent groups and interaction types in these systems. Our model offers new opportunities to predict complex multi-agent behaviors and takes a step forward in understanding interactions in multi-agent systems.
Max Guangyu Li, Bo Jiang 0001, Zhengping Che, Yan Liu 0002
AAAI2
2020 CUDAsmith: A Fuzzer for CUDA Compilers
abstract
CUDA is a parallel computing platform and programming model for the graphics processing unit (GPU) of NVIDIA. With CUDA programming, general purpose computing on GPU (GPGPU) is possible. However, the correctness of CUDA programs relies on the correctness of CUDA compilers, which is difficult to test due to its complexity. In this work, we propose CUDAsmith, a fuzzing framework for CUDA compilers. Our tool can randomly generate deterministic and valid CUDA kernel code with several different strategies. Moreover, it adopts random differential testing and EMI testing techniques to solve the test oracle problems of CUDA compiler testing. In particular, we lift live code injection to CUDA compiler testing to help generate EMI variants. Our fuzzing experiments with both the NVCC compiler and the Clang compiler for CUDA have detected thousands of failures, some of which have been confirmed by compiler developers. Finally, the cost-effectiveness of CUDAsmith is also thoroughly evaluated in our fuzzing experiment.
Bo Jiang 0001, Wing Kwong Chan, T. H. Tse, Yongfeng Yin, Zhenyu Zhang 0004
COMPSAC1
2020 MCFL: Improving Fault Localization by Differentiating Missing Code and Other Faults
abstract
Software testing is a popular practice to evaluate the software quality, and debugging is one of the most time-consuming tasks. In the last decades, spectrum-based fault localization (SBFL) techniques have been extensively studied and empirically shown effective in locating faults in a program. However, recent researches demonstrated that the accuracy of an SBFL technique may decrease when it is applied to a program containing code-omission faults. In this paper, we present a novel approach - MCFL. It models the behavior of code omission, embeds code-omission probes into programs to identify potential locations of missing code, captures spectra of program execution, and evaluates the suspiciousness of program entities being related to faults. Different from existing SBFL techniques, MCFL synthesizes a ranked list consisting of both suspicious statements and suspicious code-omission sites, which reflect the probability of a normal statement being faulty and the probability of missing code at specific positions in the program, respectively. We conducted a controlled experiment to compare the fault-localization accuracy of MCFL with those of four popular SBFL techniques. Six real-world projects from the dataset Defects4J are used as the experiment subjects. The experiment result showed that (i) MCFL outperforms the experimented SBFL techniques on most subjects, and on average has a 17.47% improvement; (ii) For more than 60% of the faults, MCFL successfully tells whether they are due to code omission.
Zhenyu Zhang 0004, Bo Jiang 0001
COMPSAC4
2020 EOSFuzzer: Fuzzing EOSIO Smart Contracts for Vulnerability Detection
abstract
EOSIO is one typical public blockchain platform. It is scalable in terms of transaction speeds and has a growing ecosystem supporting smart contracts and decentralized applications. However, the vulnerabilities within the EOSIO smart contracts have led to serious attacks, which caused serious financial loss to its end users. In this work, we systematically analyzed three typical EOSIO smart contract vulnerabilities and their related attacks. Then we presented EOSFuzzer, a general black-box fuzzing framework to detect vulnerabilities within EOSIO smart contracts. In particular, EOSFuzzer proposed effective attacking scenarios and test oracles for EOSIO smart contract fuzzing. Our fuzzing experiment on 3963 EOSIO smart contracts shows that EOSFuzzer is both effective and efficient to detect EOSIO smart contract vulnerabilities with high accuracy.
Yuhe Huang, Bo Jiang 0001, Wing Kwong Chan
Internetware2
2020 An Empirical Study of Regression Bug Chains in Linux
abstract
Regression bugs are a type of bugs that cause a feature of software that worked correctly but stop working after a certain software commit. This paper presents a systematic study of regression bug chains, an important but unexplored phenomenon of regression bugs. Our paper is based on the observation that a commit c1, which fixes a regression bug b1, may accidentally introduce another regression bug b2. Likewise, commit c2 repairing b2 may cause another regression bug b3, resulting in a bug chain, i.e., b1 → c1 → b2 → c2 → b3. We have conducted a large-scale study by collecting 1579 regression bugs and 2630 commits from 57 Linux versions (from 2.6.12 to 4.9). The relationships between regression bugs and commits are modeled as a directed bipartite network. Our major contributions and findings are fourfold: 1) a novel concept of regression bug chains and their formulation; 2) compared to an isolated regression bug, a bug on a regression bug chain is much more difficult to repair, costing 2.4× more fixing time, involving 1.3× more developers and 2.8× more comments; 3) 85.8% of bugs on the chains in Linux reside in Drivers, ACPI, Platform Specific/Hardware, and Power Management; and 4) 83% of the chains affect only a single Linux subsystem, while 68% of the chains propagate across Linux versions.
Guanping Xiao, Zheng Zheng 0001, Bo Jiang 0001, Yulei Sui
IEEE Trans. Reliab.3
2020 Adaptive Testing Based on Moment Estimation
abstract
Adaptive testing (AT) is a software testing approach that uses a feedback mechanism to enhance test effectiveness. Its testing strategy can be adjusted online by using the testing data collected during the software testing process. However, it requires complex parameter estimation which results in excessive computational overhead that may hinder the applicability of AT. In this paper, we propose an approach called AT based on moment estimation (AT-ME) to address this problem. The proposed approach uses moment estimation to serve as the algorithm of parameter estimation, which reduces the complexity of AT-ME. In addition, a dynamic length for testing action is set to limit the number of decisions without influencing the test effectiveness. The proposed approach has been validated on the Siemens test suite, which includes seven real programs. The experiments show that AT-ME can reduce the computational overhead of AT without compromising overall testing efficiency. Results demonstrate that AT-ME is a feasible and effective AT strategy.
Peng Xiao 0003, Yongfeng Yin, Bin Liu 0032, Bo Jiang 0001, Yashwant K. Malaiya
IEEE Trans. Syst. Man Cybern. Syst.4
2019 A Systematic Study on Factors Impacting GUI Traversal-Based Test Case Generation Techniques for Android Applications
abstract
Many test case generation algorithms have been proposed to test Android apps through their graphical user interfaces. However, no systematic study on the impact of the core design elements in these algorithms on effectiveness and efficiency has been reported. This paper presents the first controlled experiment to examine three key design factors, each of which is popularly used in GUI traversal-based test case generation techniques. These three major factors are definition of GUI state equivalence, state search strategy, and waiting time strategy in between input events. The empirical results on 33 Android apps with real faults revealed interesting results. First, different choices of GUI state equivalence led to significant difference on failure detection rate and extent of code coverage. Second, searching the GUI state hierarchy randomly is as effective as searching it systematically. Last but not the least, the choices on when to fire the next input event to the app under test is immaterial so long as the length of the test session is practically long enough such as 1 h. We also found two new GUI state equivalence definitions that are statistically as effective as the existing best strategy for GUI state equivalence.
Bo Jiang 0001, Yaoyue Zhang, Wing Kwong Chan, Zhenyu Zhang 0004
IEEE Trans. Reliab.1
2018 Fuse: An Architecture for Smart Contract Fuzz Testing Service
abstract
In this paper, we report our project Fuse, which is a fuzz testing service. It presents the Fuse architecture, and discusses the progress and technical issues to be addressed to fuzz-test smart contracts and support fuzz-testing of Dapps.
Wing Kwong Chan, Bo Jiang 0001
APSEC2
2018 ReTestDroid: Towards Safer Regression Test Selection for Android Application
abstract
Mobile applications are widely used in our daily life and Android is the most popular open source mobile operating system. Because mobile applications update frequently, it is important developers to perform regression testing to ensure their quality. Modeling the control flow of an android application based on the activity lifecycle model only is imprecise for regression testing. Because many Android applications use asynchronous tasks, fragments, and native code frequently, which must be considered during change impact analysis. Otherwise, regression test selection techniques may miss some failure-revealing test cases, compromising the safety of these techniques. In this work, we propose a novel approach to model asynchronous task invocations, fragment-based activity lifecycle, and native code within the control flow graph of an Android application. Furthermore, we designed a regression test selection tool ReTestDroid based on our graph model. Our experiments on five real-life Android applications showed that our approach could enable much safer regression test selection while significantly saving regression-testing time.
Bo Jiang 0001, Yongfei Zhang, Zhenyu Zhang 0004, Wing Kwong Chan
COMPSAC (1)1
2018 The Impact of Lightweight Disassembler on Malware Detection: An Empirical Study
abstract
Malicious software poses serious threats to our lives, and the activity to detect malware is becoming more and more important. An effective approach is to train a classifier using known software samples and malware samples, and recognize malware from new software. To do that, a recent popular trend is to use OpCode, which is extracted from executable modules, as an expression of software entities to drive machine learning. However, we found that the effectiveness of such a framework highly suffers from having insufficient samples, which is caused by the low success rate of disassembly due to the intrinsic complexity of the problem. In this paper, we propose to increase the success rate of disassembly by allowing inaccurate disassembling, with the attempt to increase the number of successful disassembled samples to improve OpCode-driven malware detection. We built a lightweight disassembler D-light based on the linear swap disassembly method to avoid known issues with the recursive descent manner of IDA Pro. We carried out experiment to evaluate the performance, effectiveness, and other design factors of adopting D-light and IDA Pro as disassemblers for malware detection. The empirical study shows the D-light is both more efficient and more effective than IDA Pro in supporting malware detection.
Donghong Zhang, Zhenyu Zhang 0004, Bo Jiang 0001, T. H. Tse
COMPSAC (1)3
2018 Hierarchical Deep Generative Models for Multi-Rate Multivariate Time Series
abstract
Multi-Rate Multivariate Time Series (MR-MTS) are the multivariate time series observations which come with various sampling rates and encode multiple temporal dependencies. State-space models such as Kalman filters and deep learning models such as deep Markov models are mainly designed for time series data with the same sampling rate and cannot capture all the dependencies present in the MR-MTS data. To address this challenge, we propose the Multi-Rate Hierarchical Deep Markov Model (MR-HDMM), a novel deep generative model which uses the latent hierarchical structure with a learnable switch mechanism to capture the temporal dependencies of MR-MTS. Experimental results on two real-world datasets demonstrate that our MR-HDMM model outperforms the existing state-of-the-art deep learning and state-space models on forecasting and interpolation tasks. In addition, the latent hierarchies in our model provide a way to show and interpret the multiple temporal dependencies.
Zhengping Che, Sanjay Purushotham, Max Guangyu Li, Bo Jiang 0001, Yan Liu 0002
ICML4
2018 ContractFuzzer: fuzzing smart contracts for vulnerability detection
abstract
Decentralized cryptocurrencies feature the use of blockchain to transfer values among peers on networks without central agency. Smart contracts are programs running on top of the blockchain consensus protocol to enable people make agreements while minimizing trusts. Millions of smart contracts have been deployed in various decentralized applications. The security vulnerabilities within those smart contracts pose significant threats to their applications. Indeed, many critical security vulnerabilities within smart contracts on Ethereum platform have caused huge financial losses to their users. In this work, we present ContractFuzzer, a novel fuzzer to test Ethereum smart contracts for security vulnerabilities. ContractFuzzer generates fuzzing inputs based on the ABI specifications of smart contracts, defines test oracles to detect security vulnerabilities, instruments the EVM to log smart contracts runtime behaviors, and analyzes these logs to report security vulnerabilities. Our fuzzing of 6991 smart contracts has flagged more than 459 vulnerabilities with high precision. In particular, our fuzzing tool successfully detects the vulnerability of the DAO contract that leads to USD 60 million loss and the vulnerabilities of Parity Wallet that have led to the loss of USD 30 million and the freezing of USD 150 million worth of Ether.
Bo Jiang 0001, Wing Kwong Chan
ASE1
2018 HistLock+: Precise Memory Access Maintenance Without Lockset Comparison for Complete Hybrid Data Race Detection
abstract
Dynamic hybrid data race detectors alleviate the detection imprecision problem incurred by pure lockset-based race detectors and the thread interleaving sensitive problem incurred by pure happens-before race detectors. Nonetheless, to ensure at least one data race on every memory location to be detected, keeping all historical memory access events in the analysis state of such a detector is impractical. Existing complete hybrid race detectors perform extensive comparisons among the locksets of the memory accesses on each memory location to identify which of them to be retained in its analysis state, which incurs significant runtime overhead. In this paper, we investigate to what extent a complete hybrid data race detector able to perform such identifications without lockset comparison. We present HistLock+, which is built atop thread epoch and lock release events to infer whether two memory accesses on the same memory location from the same thread in between consecutive lock release operations have any lock subset relation without performing expensive lockset comparison. HistLock+ guarantees exactly one racy memory access event to be reported on each thread segment separated by lock releases and hard-order thread synchronizations, and it never reports false positive on lockset violation. We have validated HistLock+ using the PARSEC benchmark suite and four real-world applications. The experimental results showed that HistLock+ was 122% faster and 28% more memory-efficient than the previous state-of-the-art complete hybrid race detector. Moreover, HistLock+ achieved the highest effectiveness in race detection among all evaluated race detectors in our experiment.
Bo Jiang 0001, Wing Kwong Chan
IEEE Trans. Reliab.2
2017 Introducing parallel computing concepts in computer system related courses
abstract
All semiconductor market domains are converging to concurrent platforms. This trend has certainly led real challenge to develop applications software that effectively uses these concurrent processors to achieve efficiency and performance goals. This paper argues that the Computer System related courses are natural places to introduce the parallelism, and the earlier to parallel computing concepts will have a wide reach. This paper showed how digital logic classes can motivate topics from parallel computing through common logic structures. We provided an alternative view of the digital logic topics, including: binary representation of integers using decision tree with recursive thinking; use a carry look-ahead adder to show how sequential operations can be parallelized. Another part to introduce parallel concepts focused on write high performance code with specific emphasis on graphic processing unit (GPU). In our teaching experience, parallel pattern teaching has been confirmed to be a useful pedagogical method for teaching parallel concepts. Finally, we report course experience in teaching parallel computing injected course, which resulted in positive student feedback.
Han Wan, Xiaopeng Gao, Xiang Long, Bo Jiang 0001
FIE4
2017 SimplyDroid: efficient event sequence simplification for Android application
abstract
To ensure the quality of Android applications, many automatic test case generation techniques have been proposed. Among them, the Monkey fuzz testing tool and its variants are simple, effective and widely applicable. However, one major drawback of those Monkey tools is that they often generate many events in a failure-inducing input trace, which makes the follow-up debugging activities hard to apply. It is desirable to simplify or reduce the input event sequence while triggering the same failure. In this paper, we propose an efficient event trace representation and the SimplyDroid tool with three hierarchical delta-debugging algorithms each operating on this trace representation to simplify crash traces. We have evaluated SimplyDroid on a suite of real-life Android applications with 92 crash traces. The empirical result shows that our new algorithms in SimplyDroid are both efficient and effective in reducing these event traces.
Bo Jiang 0001, Wing Kwong Chan
ASE1
2017 Which Factor Impacts GUI Traversal-Based Test Case Generation Technique Most? A Controlled Experiment on Android Applications
abstract
There are many research works on automated GUI traversal-based test case generation techniques for Android application. However, the effect of different factors used in a GUI traversal algorithm has not been systematically explored. In this work, we report a controlled experiment on 33 real-world applications to expose their real failures to systematically study three major factors that are commonly observed in testing tools for this class of applications. They include the notion of GUI state equivalence, the state search (or exploration) strategy, and the amount of time to wait between two input events. Our experimental results clearly show that different notions of GUI state equivalences have significantly different effects on failure detection rate and code coverage, randomized search is comparable to systematic search, and different choices of waiting time strategies do not make significant differences in terms of testing effectiveness. We also report other interesting results in this paper.
Bo Jiang 0001, Yaoyue Zhang, Wing Kwong Chan, Zhenyu Zhang 0004
QRS1
2016 Testing and Debugging in Continuous Integration with Budget Quotas on Test Executions
abstract
In Continuous Integration, a software application is developed through a series of development sessions, each with limited time allocated to testing and debugging on each of its modules. Test Case Prioritization can help execute test cases with higher failure estimate earlier in each session. When the testing time is limited, executing such prioritized test cases may only produce partial and prioritized execution coverage data. To identify faulty code, existing Spectrum-Based Fault Localization techniques often use execution coverage data but without the assumption of execution coverage priority. Is it possible to decompose these two steps for optimization within individual steps? In this paper, we study to what extent the selection of test case prioritization techniques may reduce its influence on the effectiveness of spectrum-based fault localization, thereby showing the possibility to decompose the process of continuous integration for optimization in workflow steps. We present a controlled experiment using the Siemens suite as subjects, nine test case prioritization techniques and four spectrum-based fault localization techniques. The findings showed that the studied test cases prioritization and spectrum-based fault localization can be customized separately, and, interestingly, prioritization over a smaller test suite can enable spectrum-based fault localization to achieve higher accuracy by assigning faulty statements with higher ranks.
Bo Jiang 0001, Wing Kwong Chan
QRS1
2016 Facilitating Monkey Test by Detecting Operable Regions in Rendered GUI of Mobile Game Apps
abstract
Graphical User Interface (GUI) is a component of many software applications. Many mobile game applications in particular have to provide excellent user experiences using graphical engines to render GUI screens. On a rendered GUI screen such as a treasury map, no GUI widget is embodied in it and the operable GUI regions, each of which is a region that triggers actions when certain events acting on these regions, may only be implicitly determinable. Traditional testing tools like monkey test do not effectively generate effective event sequences over such operable GUI regions. Our insight is that operable regions in a rendered GUI screen of many mobile game applications are given with visible hints to catch user attentions. In this paper, we propose Smart Monkey, which uses the fundamental features of a screen, including color, intensity, and texture, as visual signals to detect operable GUI region candidates, and iteratively identifies and confirms the real operable GUI regions by launching GUI events to the region. We have implemented Smart Monkey as a testing tool for Android apps and conducted case studies on real-world applications to compare it with a peer technique. The empirical results show that it effective in identifying such operable regions and thus able to generate functional event sequences more efficiently.
Chenglong Sun, Zhenyu Zhang 0004, Bo Jiang 0001, Wing Kwong Chan
QRS3
2016 To What Extent is Stress Testing of Android TV Applications Automated in Industrial Environments?
abstract
An Android-based smart television (TV) must reliably run its applications in an embedded program environment under diverse hardware resource conditions. Owing to the diverse hardware components used to build numerous TV models, TV simulators are usually not sufficiently high in fidelity to simulate various TV models and thus are only regarded as unreliable alternatives when stress testing such applications. Therefore, even though stress testing on real TV sets is tedious, it is the de facto approach to ensure the reliability of these applications in the industry. In this paper, we study to what extent stress testing of smart TV applications can be fully automated in the industrial environments. To the best of our knowledge, no previous work has addressed this important question. We summarize the findings collected from ten industrial test engineers who have tested 20 such TV applications in a real production environment. Our study shows that the industry required test automation supports on high-level GUI object controls and status checking, setup of resource conditions, and the interplay between the two. With such supports, 87% of the industrial test specifications of one TV model can be fully automated, and 71.4% of them were found to be fully reusable to test a subsequent TV model with major upgrades of hardware, operating system, and application. It represents a significant improvement with margins of 28% and 38%, respectively, compared with stress testing without such supports.
Bo Jiang 0001, Wing Kwong Chan, Xinchao Zhang
IEEE Trans. Reliab.1
2016 FLANDROID: Energy-Efficient Recommendations of Reliable Context Providers for Android Applications
abstract
Mobile applications are becoming more and more popular with the prevalence of mobile operating systems and mobile Internet. Many of them consume services provided by the underlying infrastructure and platforms as a part of their application environmental contexts. However, application failures or downgrade in performance may be the results due to inadequate provisions of these environmental issues in the implementations of the mobile applications. In this paper, we propose a framework to enable mobile applications to consume services offered by a reliable context provider with high probability in run time. We report a case study on a suite of five real-world mobile applications with 74 real faults on real mobile phones involving 50 users. The results of the case study show that our framework can significantly improve the reliability of mobile applications with respect to the failures due to buggy-context-provider faults with low slowdown and energy overheads.
Bo Jiang 0001
IEEE Trans. Serv. Comput.1
2015 PORA: Proportion-Oriented Randomized Algorithm for Test Case Prioritization
abstract
Effective testing is essential for assuring software quality. While regression testing is time-consuming, the fault detection capability may be compromised if some test cases are discarded. Test case prioritization is a viable solution. To the best of our knowledge, the most effective test case prioritization approach is still the additional greedy algorithm, and existing search-based algorithms have been shown to be visually less effective than the former algorithms in previous empirical studies. This paper proposes a novel Proportion-Oriented Randomized Algorithm (PORA) for test case prioritization. PORA guides test case prioritization by optimizing the distance between the prioritized test suite and a hierarchy of distributions of test input data. Our experiment shows that PORA test case prioritization techniques are as effective as, if not more effective than, the total greedy, additional greedy, and ART techniques, which use code coverage information. Moreover, the experiment shows that PORA techniques are more stable in effectiveness than the others.
Bo Jiang 0001, Wing Kwong Chan, T. H. Tse
QRS1
2015 Input-based adaptive randomized test case prioritization: A local beam search approach
abstract
Test case prioritization assigns the execution priorities of the test cases in a given test suite. Many existing test case prioritization techniques assume the full-fledged availability of code coverage data, fault history, or test specification, which are seldom well-maintained in real-world software development projects. This paper proposes a novel family of input-based local-beam-search adaptive-randomized techniques. They make adaptive tree-based randomized explorations with a randomized candidate test set strategy to even out the search space explorations among the branches of the exploration trees constructed by the test inputs in the test suite. We report a validation experiment on a suite of four medium-size benchmarks. The results show that our techniques achieve either higher APFD values than or the same mean APFD values as the existing code-coverage-based greedy or search-based prioritization techniques, including Genetic, Greedy and ART, in both our controlled experiment and case study. Our techniques are also significantly more efficient than the Genetic and Greedy, but are less efficient than ART.
Bo Jiang 0001, Wing Kwong Chan
J. Syst. Softw.1
2015 A Subsumption Hierarchy of Test Case Prioritization for Composite Services
abstract
Many composite workflow services utilize non-imperative XML technologies such as WSDL, XPath, XML schema, and XML messages. Regression testing should assure the services against regression faults that appear in both the workflows and these artifacts. In this paper, we propose a refinement-oriented level-exploration strategy and a multilevel coverage model that captures progressively the coverage of different types of artifacts by the test cases. We show that by using them, the test case prioritization techniques initialized on top of existing greedy-based test case prioritization strategy form a subsumption hierarchy such that a technique can produce more test suite permutations than a technique that subsumes it. Our experimental study of a model instance shows that a technique generally achieves a higher fault detection rate than a subsumed technique, which validates that the proposed hierarchy and model have the potential to improve the cost-effectiveness of test case prioritization techniques.
Lijun Mei, Yan Cai 0001, Changjiang Jia, Bo Jiang 0001, Wing Kwong Chan, Zhenyu Zhang 0004, T. H. Tse
IEEE Trans. Serv. Comput.4
2015 Preemptive Regression Testingof Workflow-Based Web Services
abstract
An external web service may evolve without prior notification. In the course of the regression testing of a workflow-based web service, existing test case prioritization techniques may only verify the latest service composition using the not-yet-executed test cases, overlooking high-priority test cases that have already been applied to the service composition before the evolution. In this paper, we propose Preemptive Regression Testing (PRT), an adaptive testing approach to addressing this challenge. Whenever a change in the coverage of any service artifact is detected, PRT recursively preempts the current session of regression test and creates a sub-session of the current test session to assure such lately identified changes in coverage by adjusting the execution priority of the test cases in the test suite. Then, the sub-session will resume the execution from the suspended position. PRT terminates only when each test case in the test suite has been executed at least once without any preemption activated in between any test case executions. The experimental result confirms that testing workflow-based web service in the face of such changes is very challenging; and one of the PRT-enriched techniques shows its potential to overcome the challenge.
Lijun Mei, Wing Kwong Chan, T. H. Tse, Bo Jiang 0001, Ke Zhai 0002
IEEE Trans. Serv. Comput.4
2014 Prioritizing Test Cases for Regression Testing of Location-Based Services: Metrics, Techniques, and Case Study
abstract
Location-based services (LBS) are widely deployed. When the implementation of an LBS-enabled service has evolved, regression testing can be employed to assure the previously established behaviors not having been adversely affected. Proper test case prioritization helps reveal service anomalies efficiently so that fixes can be scheduled earlier to minimize the nuisance to service consumers. A key observation is that locations captured in the inputs and the expected outputs of test cases are physically correlated by the LBS-enabled service, and these services heuristically use estimated and imprecise locations for their computations, making these services tend to treat locations in close proximity homogenously. This paper exploits this observation. It proposes a suite of metrics and initializes them to demonstrate input-guided techniques and point-of-interest (POI) aware test case prioritization techniques, differing by whether the location information in the expected outputs of test cases is used. It reports a case study on a stateful LBS-enabled service. The case study shows that the POI-aware techniques can be more effective and more stable than the baseline, which reorders test cases randomly, and the input-guided techniques. We also find that one of the POI-aware techniques, cdist, is either the most effective or the second most effective technique among all the studied techniques in our evaluated aspects, although no technique excels in all studied SOA fault classes.
Ke Zhai 0002, Bo Jiang 0001, Wing Kwong Chan
IEEE Trans. Serv. Comput.2
2013 An Efficient Grouped Virtual Mapreduce Cluster
abstract
Virtualization technology and MapReduce program model are sharp swords for the big data and cloud computing era. The combination of them exhibits powerful ability of easy-management, fast-deployment, feasible-scalability and high-efficiency. However, the downside is that the performance is limited by the I/O bottleneck of Virtual Machine(VM). A huge number of data should be handled in MapReduce cluster which is deployed in VMs. Luckily, data locality, a very crucial issue affecting performance in a shared clusters environment, is used to ease this conflict and improve the execution time of applications. We present a framework of Grouped Virtual MapReduce Cluster(GVMC) which takes fully advantage of VM data locality to exhibit high performance of Virtual MapReduce Cluster(VMC). The introduction of local-master nodes in GVMC not only offloads the pressure of the master node, but also lowers the communication cost. We compare the organization of three different VMC, describe the architecture of our cluster framework and do the performance analysis. Our experiments demonstrate that the framework of GVMC achieves higher locality and reduces the execution time in both CPU-intensive applications and I/O-intensive applications. Compared to Original Virtual MapReduce Cluster(OVMC), the performance of GVMC improvement is up to 16.5% and 36.2% for CPU-intensive applications and I/O-intensive applications respectively.
Xiang Long, Bo Jiang 0001
AINA3
2013 Bypassing Code Coverage Approximation Limitations via Effective Input-Based Randomized Test Case Prioritization
abstract
Test case prioritization assigns the execution priorities of the test cases in a given test suite with the aim of achieving certain goals. Many existing test case prioritization techniques however assume the full-fledged availability of code coverage data, fault history, or test specification, which are seldom well-maintained in many software development projects. This paper proposes a novel family of LBS techniques. They make adaptive tree-based randomized explorations with an adaptive randomized candidate test set strategy to diversify the explorations among the branches of the exploration trees constructed by the test inputs in the test suite. They get rid of the assumption on the historical correlation of code coverage between program versions. Our techniques can be applied to programs with or without any previous versions, and hence are more general than many existing test case prioritization techniques. The empirical study on four popular UNIX utility benchmarks shows that, in terms of APFD, our LBS techniques can be as effective as some of the best code coverage-based greedy prioritization techniques ever proposed. We also show that they are significantly more efficient and scalable than the latter techniques.
Bo Jiang 0001, Wing Kwong Chan
COMPSAC1
2013 Prioritizing Structurally Complex Test Pairs for Validating WS-BPEL Evolutions
abstract
Many web services represent their artifacts in the semi-structural format. Such artifacts may or may not be structurally complex. Many existing test case prioritization techniques however treat test cases of different complexity generically. In this paper, we exploit the insights on the structural similarity of XML-based artifacts between test cases, and propose a family of test case prioritization techniques that iteratively selects test case pairs without replacement. The validation experiment shows that these techniques can be more cost-effective than the studied existing techniques in exposing faults.
Lijun Mei, Yan Cai 0001, Changjiang Jia, Bo Jiang 0001, Wing Kwong Chan
ICWS4
2013 On the adoption of MC/DC and control-flow adequacy for a tight integration of program testing and statistical fault localization
Bo Jiang 0001, Ke Zhai 0002, Wing Kwong Chan, T. H. Tse, Zhenyu Zhang 0004
Inf. Softw. Technol.1
2012 Preemptive Regression Test Scheduling Strategies: A New Testing Approach to Thriving on the Volatile Service Environments
abstract
A workflow-based web service may use ultra-late binding to invoke external web services to concretize its implementation at run time. Nonetheless, such external services or the availability of recently used external services may evolve without prior notification, dynamically triggering the workflow-based service to bind to new replacement external services to continue the current execution. Any integration mismatch may cause a failure. In this paper, we propose Preemptive Regression Testing (PRT), a novel testing approach that addresses this adaptive issue. Whenever such a late-change on the service under regression test is detected, PRT preempts the currently executed regression test suite, searches for additional test cases as fixes, runs these fixes, and then resumes the execution of the regression test suite from the preemption point.
Lijun Mei, Ke Zhai 0002, Bo Jiang 0001, Wing Kwong Chan, T. H. Tse
COMPSAC3
2012 How well does test case prioritization integrate with statistical fault localization?
Bo Jiang 0001, Zhenyu Zhang 0004, Wing Kwong Chan, T. H. Tse, Tsong Yueh Chen
Inf. Softw. Technol.1
2011 Precise Propagation of Fault-Failure Correlations in Program Flow Graphs
abstract
Statistical fault localization techniques find suspicious faulty program entities in programs by comparing passed and failed executions. Existing studies show that such techniques can be promising in locating program faults. However, coincidental correctness and execution crashes may make program entities indistinguishable in the execution spectra under study, or cause inaccurate counting, thus severely affecting the precision of existing fault localization techniques. In this paper, we propose a Block Rank technique, which calculates, contrasts, and propagates the mean edge profiles between passed and failed executions to alleviate the impact of coincidental correctness. To address the issue of execution crashes, Block Rank identifies suspicious basic blocks by modeling how each basic block contributes to failures by apportioning their fault relevance to surrounding basic blocks in terms of the rate of successful transition observed from passed and failed executions. Block Rank is empirically shown to be more effective than nine representative techniques on four real-life medium-sized programs.
Zhenyu Zhang 0004, Wing Kwong Chan, T. H. Tse, Bo Jiang 0001
COMPSAC4
2011 Assuring the model evolution of protocol software specifications by regression testing process improvement
abstract
SUMMARY Model‐based testing helps test engineers automate their testing tasks so that they are more cost‐effective. When the model is changed because of the evolution of the specification, it is important to maintain the test suites up to date for regression testing. A complete regeneration of the whole test suite from the new model, although inefficient, is still frequently used in the industry, including Microsoft. To handle specification evolution effectively, we propose a test case reusability analysis technique to identify reusable test cases of the original test suite based on graph analysis. We also develop a test suite augmentation technique to generate new test cases to cover the change‐related parts of the new model. The experiment on four large protocol document testing projects shows that our technique can successfully identify a high percentage of reusable test cases and generate low‐redundancy new test cases. When compared with a complete regeneration of the whole test suite, our technique significantly reduces regression testing time while maintaining the stability of requirement coverage over the evolution of requirements specifications. Copyright © 2011 John Wiley & Sons, Ltd.
Bo Jiang 0001, T. H. Tse, Wolfgang Grieskamp, Nicolas Kicillof, Wing Kwong Chan
Softw. Pract. Exp.1
2010 Taking Advantage of Service Selection: A Study on the Testing of Location-Based Web Services Through Test Case Prioritization
abstract
Dynamic service compositions pose new verification and validation challenges such as uncertainty in service membership. Moreover, applying an entire test suite to loosely coupled services one after another in the same composition can be too rigid and restrictive. In this paper, we investigate the impact of service selection on service-centric testing techniques. Specifically, we propose to incorporate service selection in executing a test suite and develop a suite of metrics and test case prioritization techniques for the testing of location-aware services. A case study shows that a test case prioritization technique that incorporates service selection can outperform their traditional counterpart - the impact of service selection is noticeable on software engineering techniques in general and on test case prioritization techniques in particular. Further-more, we find that points-of-interest-aware techniques can be significantly more effective than input-guided techniques in terms of the number of invocations required to expose the first failure of a service composition.
Ke Zhai 0002, Bo Jiang 0001, Wing Kwong Chan, T. H. Tse
ICWS2
2010 Fault localization through evaluation sequences
Zhenyu Zhang 0004, Bo Jiang 0001, Wing Kwong Chan, T. H. Tse
J. Syst. Softw.2
2009 Adaptive Random Test Case Prioritization
abstract
Regression testing assures changed programs against unintended amendments. Rearranging the execution order of test cases is a key idea to improve their effectiveness. Paradoxically, many test case prioritization techniques resolve tie cases using the random selection approach, and yet random ordering of test cases has been considered as ineffective. Existing unit testing research unveils that adaptive random testing (ART) is a promising candidate that may replace random testing (RT). In this paper, we not only propose a new family of coverage-based ART techniques, but also show empirically that they are statistically superior to the RT-based technique in detecting faults. Furthermore, one of the ART prioritization techniques is consistently comparable to some of the best coverage-based prioritization techniques (namely, the "additional" techniques) and yet involves much less time cost.
Bo Jiang 0001, Zhenyu Zhang 0004, Wing Kwong Chan, T. H. Tse
ASE1
2009 Capturing propagation of infected program states
abstract
Coverage-based fault-localization techniques find the fault-related positions in programs by comparing the execution statistics of passed executions and failed executions. They assess the fault suspiciousness of individual program entities and rank the statements in descending order of their suspiciousness scores to help identify faults in programs. However, many such techniques focus on assessing the suspiciousness of individual program entities but ignore the propagation of infected program states among them. In this paper, we use edge profiles to represent passed executions and failed executions, contrast them to model how each basic block contributes to failures by abstractly propagating infected program states to its adjacent basic blocks through control flow edges. We assess the suspiciousness of the infected program states propagated through each edge, associate basic blocks with edges via such propagation of infected program states, calculate suspiciousness scores for each basic block, and finally synthesize a ranked list of statements to facilitate the identification of program faults. We conduct a controlled experiment to compare the effectiveness of existing representative techniques with ours using standard bench-marks. The results are promising.
Zhenyu Zhang 0004, Wing Kwong Chan, T. H. Tse, Bo Jiang 0001
ESEC/SIGSOFT FSE4
2009 Where to adapt dynamic service compositions
abstract
Peer services depend on one another to accomplish their tasks, and their structures may evolve. A service composition may be designed to replace its member services whenever the quality of the composite service fails to meet certain quality-of-service (QoS) requirements. Finding services and service invocation endpoints having the greatest impact on the quality are important to guide subsequent service adaptations. This paper proposes a technique that samples the QoS of composite services and continually analyzes them to identify artifacts for service adaptation. The preliminary results show that our technique has the potential to effectively find such artifacts in services.
Bo Jiang 0001, Wing Kwong Chan, Zhenyu Zhang 0004, T. H. Tse
WWW1
2008 Debugging through Evaluation Sequences: A Controlled Experimental Study
abstract
Predicate-based statistical fault-localization techniques locate fault-relevant predicates in a program by contrasting the statistics of the values of individual predicates between successful and failure-causing runs. While short-circuit evaluations are common in program execution, treating predicates as atomic units ignores this fact, masking out various types of important statistics. On the contrary, are such statistics useful for debugging? In this paper, we investigate experimentally the impact of the use of short-circuit evaluation information on fault localization. The results show that, by doing so, it significantly improves predicate-based statistical fault-localization techniques.
Zhenyu Zhang 0004, Bo Jiang 0001, Wing Kwong Chan, T. H. Tse
COMPSAC2