Yangfan Zhou 0002

dblp:27/6390-2 · DBLP profile ↗
← Back
75ranked-venue papers
11as first author
29since 2021 · last 2025
0000-0002-9184-7383ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 36 · 1 first-author · 20 since 2021Computer networks · 13 · 7 first-authorSystems, architecture and hardware · 9 · 1 first-author · 4 since 2021Security and privacy · 7Human-computer interaction and ubiquitous computing · 6 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 COAS2W: A Chinese Older-Adults Spoken-to-Written Transformation Corpus with Context Awareness
abstract
Spoken language from older adults often deviates from written norms due to omission, disordered syntax, constituent errors, and redundancy, limiting the usefulness of automatic transcripts in downstream tasks.We present COAS2W, a Chinese spoken-to-written corpus of 10,004 utterances from older adults, each paired with a written version, finegrained error labels, and four-sentence context.Fine-tuned lightweight open-source models on COAS2W outperform larger closed-source models.Context ablation shows the value of multi-sentence input, and normalization improves performance on downstream translation tasks.COAS2W supports the development of inclusive, context-aware language technologies for older speakers.Our annotation convention, data, and code are publicly available at https://github.com/Springrx/COAS2W.
Chun Kang, Zhigu Qian, Zhen Fu, Jiaojiao Fu, Yangfan Zhou 0002
EMNLP5
2025 Can User Feedback Help Issue Detection? An Empirical Study on a One-Billion-User Online Service System
abstract
Background: It has long been suggested that user feedback, typically written in natural language by end-users, can help issue detection. However, for large-scale online service systems that receive a tremendous amount of feedback, it remains a challenging task to identify severe issues from user feedback. Aims: To develop a better feedback-based issue detection approach, it is crucial first to gain a comprehensive understanding of the characteristics of user feedback in real production systems. Method: In this paper, we conduct an empirical study on 50,378,766 user feedback items from six real-world services in a one-billion-user online service system. We first study what users provide in their feedback. We then examine whether certain features of feedback items can be good indicators of severe issues. Finally, we investigate whether adopting machine learning techniques to analyze user feedback is reasonable. Results: Our results show that a large proportion of user feedback provides irrelevant information about system issues. As a result, it is crucial to filter out issue-irrelevant information when processing user feedback. Moreover, we find severe issues that cannot be easily detected based solely on user feedback characteristics. Finally, we find that the distributions of the feedback topics in different time intervals are similar. This confirms that designing machine learning-based approaches is a viable direction for better analyzing user feedback. Conclusions: We consider that our findings can serve as an empirical foundation for feedback-based issue detection in large-scale service systems, which sheds light on the design and implementation of practical issue detection approaches.
Shuyao Jiang, Jiazhen Gu, Wujie Zheng, Yangfan Zhou 0002, Michael R. Lyu
ESEM4
2025 Minuku: Detecting Diverse Display Issues in Mobile Apps with Small-scale Dataset
abstract
User interface (UI) display issues, such as widgets occlusion, missing elements, and screen overflow, are emerging as a non-negligible source of user complaints in commercial mobile apps. However, existing automated testing tools typically rely on a vast amount of high-quality training data, making them cost-ineffective for industrial practice. Given that display issues are intuitively recognizable by humans, their diverse appearances can be abstracted by the violation of human commonsense expectations of UI appearance. Therefore, this paper proposes to reduce data requirements in display issue detection through commonsense simulation. Although leveraging large vision-language models (VLMs) to replicate human visual ability looks straightforward, off-the-shelf VLMs lack task-specific knowledge of UI designs and display correctness. To address this, we fine-tune a VLM to learn what constitutes an expected display and to reason potential display issues. This approach is termed as Minuku, an industrial data-efficient UI display issue detector. We evaluate the design effectiveness of Minuku via a set of ablation experiments. Moreover, real-world deployments in one of the largest E-commerce app providers further demonstrate that Minuku can effectively detect 40 previously unknown UI display issues and significantly reduce manual effort in industrial settings.
Yongxiang Hu 0003, Hailiang Jin, Juxing Yuan, Xin Wang 0002, Yangfan Zhou 0002
ASE7
2025 From Redundancy to Efficiency: Exploiting Shared UI Interactions towards Efficient LLM-Based Testing
abstract
Redundant test cases, although well-studied in software engineering, are previously underexplored in UI testing of mobile apps. Our study of real-world test suites shows that, equipped with large-scale testing suites, redundancy in UI testing often manifests as redundant UI interactions. Although negligible in traditional script-based workflows, such redundancy severely impacts the efficiency of emerging Large Language Model (LLM)-based UI agents, which incur substantial decision latency and token costs from repeated LLM queries for the same interactions. To this end, based on the idea of reusing LLMs’ former decisions, we present TestWeaver, a cost-effective LLM-based testing framework. Leveraging a semantic annotated UI Transition Graph (UTG), TestWeaver is capable of detecting shared interactions across test cases. It processes each interaction with a single LLM query, and reuses the result whenever the same interaction occurs. We evaluate TestWeaver on real-world test suites from Meituan. It achieves a 92% success rate with an average cost of $0.11 and 89.7 seconds per case, outperforming the state-of-the-art. We have also deployed TestWeaver in a real-world testing workflow at Meituan for over six months. TestWeaver has executed nearly 2,000 test cases and uncovered 10 previously undetected bugs, while reducing manual testing effort by 75%.
Yingchuan Wang, Yongxiang Hu 0003, Yu Zhang 0165, Hailiang Jin, Juxing Yuan, Yangfan Zhou 0002
ASE8
2025 Distinguishability-Guided Test Program Generation for WebAssembly Runtime Performance Testing
abstract
WebAssembly (Wasm) is a binary instruction format designed as a portable compilation target, which has been widely used on both the web and server sides in recent years. As high performance is a critical design goal of Wasm, it is essential to conduct performance testing for Wasm runtimes. However, existing research on Wasm runtime performance testing still suffers from insufficient high-quality test programs. To solve this problem, we propose a novel test program generation approach WarpGen. It first extracts code snippets from historical issue-triggering test programs as initial operators, then inserts an operator into a seed program to synthesize a new test program. To verify the quality of generated programs, we propose an indicator called distinguishability, which refers to the ability of a test program to distinguish abnormal performance of specific Wasm runtimes. We apply WarpGen for performance testing on four Wasm runtimes and verify its effectiveness compared with baseline approaches. In particular, WarpGen has identified seven new performance issues in three Wasm runtimes.
Shuyao Jiang, Ruiying Zeng, Yangfan Zhou 0002, Michael R. Lyu
SANER3
2025 Exploring the Feasibility and Challenges of Treating Follow-up Patients via a Mobile Platform in China
abstract
Mobile technology is being increasingly adopted in teleconsultation for its convenience and mobility. Although widely accepted by physicians for informal online consultations, the effectiveness of mobile platforms in formal medical interventions, such as follow-up treatments, remains largely underexplored. This research presents a case study from a Chinese hospital, examining how physicians use mobile platforms to treat chronic patients, the challenges they encounter, and their overall experience. The study aims to determine whether mobile technology can enable physicians to fulfill their responsibilities in managing chronic disease follow-ups online. Through observations and interviews, we found that using mobile platforms introduces significant challenges. Physicians operate in complex and varying environments, such as workplaces, homes, and even public areas, making them struggle to access comprehensive information, communicate effectively, and maintain detailed medical records. As a result, physicians face poor working conditions, reduced efficiency, and struggle to make accurate treatment decisions. This study highlights the need for policy reforms and technological innovations to ensure sustainable teleconsultation practices on mobile platforms. It contributes to the HCI and CSCW communities by highlighting the feasibility of using mobile platforms in online chronic disease follow-ups.
Jiaojiao Fu, Yangfan Zhou 0002, Xin Wang 0002, Yi Guo 0009
Proc. ACM Hum. Comput. Interact.2
2025 Seeking a Sense of Meaning and Companionship in Life: Informal Learning on Douyin Among Chinese Older Adults
abstract
This study examines how Chinese older adults leverage Douyin, a short video platform, for informal learning purposes, analyzing their usage patterns, motivations, and encountered challenges. Although Douyin was not explicitly designed with educational features, it has emerged as a significant informal learning platform for this demographic. Through a qualitative investigation comprising participant observations and semi-structured interviews with 17 participants, we reveal the distinctive learning experience that Douyin facilitates. The platform's unique combination of short-form videos, live streaming capabilities, and interactive community features creates an engaging learning environment that particularly resonates with older adults. Our findings demonstrate that participants derive substantial social-emotional benefits beyond knowledge acquisition, including strengthened social connections, enhanced companionship, and a reinforced sense of purpose through their Douyin engagement. These social-emotional aspects emerge as crucial factors driving older adults' selection of Douyin as their preferred learning platform. By providing comprehensive insights into older adults' engagement with digital platforms for lifelong learning, this research offers valuable implications for HCI, particularly in understanding how technology can be optimized to support the learning needs of older populations.
Zhigu Qian, Jiaojiao Fu, Yangfan Zhou 0002
Proc. ACM Hum. Comput. Interact.3
2024 Is unsafe an Achilles' Heel? A Comprehensive Study of Safety Requirements in Unsafe Rust Programming
abstract
Rust is an emerging, strongly-typed programming language focusing on efficiency and memory safety. With increasing projects adopting Rust, knowing how to use Unsafe Rust is crucial for Rust security. We observed that the description of safety requirements needs to be unified in Unsafe Rust programming. Current unsafe API documents in the standard library exhibited variations, including inconsistency and insufficiency. To enhance Rust security, we suggest unsafe API documents to list systematic descriptions of safety requirements for users to follow.
Mohan Cui, Shuran Sun, Hui Xu 0009, Yangfan Zhou 0002
ICSE4
2024 Wapplique: Testing WebAssembly Runtime via Execution Context-Aware Bytecode Mutation
abstract
Reliability is the top concern to runtimes. This paper studies how to test Wasm runtime, by presenting Wapplique, the first Wasm bytecode mutation-based fuzzing tool. Wapplique solves the diversity/efficiency dilemma in generating test cases with a specifically-tailored code-fragment substitution approach for Wasm. In particular, Wapplique appliqués code fragments from real-world programs to seed programs to enhance the diversity of the seeds. Via sophisticated code analysis algorithms we design, Wapplique also guarantees the validity of the resulting programs. This allows Wapplique to generate tremendous valid and diverse Wasm programs as test cases to well exercise target runtimes. Our experiences on applying Wapplique in testing four prevalent real-world runtimes indicate that it can generate test cases efficiently, achieve high coverage, and find 20 previously unknown bugs.
Ruiying Zeng, Yangfan Zhou 0002
ISSTA3
2024 CHIME: A Cache-Efficient and High-Performance Hybrid Index on Disaggregated Memory
abstract
Disaggregated memory (DM) is a widely discussed datacenter architecture in academia and industry. It decouples computing and memory resources from monolithic servers into two network-connected resource pools. Range indexes are widely adopted by storage systems on DM to efficiently locate and query remote data. However, existing range indexes on DM suffer from either high computing-side cache consumption or high memory-side read amplifications. In this paper, we propose CHIME, a hybrid index combining B+ trees with hopscotch hashing, to achieve low cache consumption and low read amplifications simultaneously. There are three challenges in constructing CHIME on DM, i.e., the complicated optimistic synchronization, the extra metadata access, and the read amplifications introduced by hopscotch hashing. CHIME leverages 1) a three-level optimistic synchronization scheme to synchronize read and write operations with various granularities, 2) an access-aggregated metadata management technique to eliminate extra metadata accesses by piggybacking and replicating metadata, and 3) an effective hotness-aware speculative read mechanism to mitigate the read amplifications of hopscotch hashing. Experimental results show that CHIME outperforms the state-of-the-art range indexes on DM by up to 5.1× with the same cache size and achieves similar performance with up to 8.7× lower cache consumption.
Xuchuan Luo, Jiacheng Shen, Pengfei Zuo, Xin Wang 0002, Michael R. Lyu, Yangfan Zhou 0002
SOSP6
2024 Gaining Technological Autonomy and Soci-emotional Support: A Case Study of How and Why Chinese Older Adults Engage with a Semi-acquaintance Online Community
abstract
Older adults are often underserved and marginalized in technology engagement due to their reluctance and the barriers they face in adopting and engaging with mainstream technology. However, Pinxiaoquan, a social feature of an e-commerce platform in China, has gained a large number of older users. This work investigates how and why Chinese older adults use Pinxiaoquan, aiming to unveil the underlying logic and inspire technology-inclusive design for older adults. To this end, we conducted a mixed-methods qualitative study over two years, which included online observation, and semi-interview. We found that Pinxiaoquan's success among Chinese older adults is mainly due to its ability to inspire technological autonomy and provide social-emotional support. Rather than simply lowering technical barriers or asking them to seek assistance outside the platform, Pinxiaoquan builds a semi-acquaintance online community based on location and social ties that allows older adults to realize technical support mutually. Pinxiaoquan also fulfills their social-emotional needs, such as free online expression, memory creation and preservation, relationship expansion and maintenance, and a sense of value. Our research contributes to the HCI community by highlighting the importance of improving older adults' technology autonomy following their specific social and cultural background and social-emotional needs. This work also provides unique insights and implications for building inclusive technology for the growing aging population.
Zhigu Qian, Jiaojiao Fu, Yangfan Zhou 0002
Proc. ACM Hum. Comput. Interact.3
2024 A Memory-Disaggregated Radix Tree
abstract
Disaggregated memory (DM) is an increasingly prevalent architecture with high resource utilization. It separates computing and memory resources into two pools and interconnects them with fast networks. Existing range indexes on DM are based on B+ trees, which suffer from large inherent read and write amplifications. The read and write amplifications rapidly saturate the network bandwidth, resulting in low request throughput and high access latency of B+ trees on DM. In this article, we propose that the radix tree is more suitable for DM than the B+ tree due to smaller read and write amplifications. However, constructing a radix tree on DM is challenging due to the costly lock-based concurrency control, the bounded memory-side IOPS, and the complicated computing-side cache validation. To address these challenges, we design SMART , the first radix tree for disaggregated memory with high performance. Specifically, we leverage (1) a hybrid concurrency control scheme including lock-free internal nodes and fine-grained lock-based leaf nodes to reduce lock overhead, (2) a computing-side read-delegation and write-combining technique to break through the IOPS upper bound by reducing redundant I/Os, and (3) a simple yet effective reverse check mechanism for computing-side cache validation. Experimental results show that SMART achieves 6.1× higher throughput under typical write-intensive workloads and 2.8× higher throughput under read-only workloads in YCSB benchmarks, compared with state-of-the-art B+ trees on DM.
Xuchuan Luo, Pengfei Zuo, Jiacheng Shen, Jiazhen Gu, Xin Wang 0002, Michael R. Lyu, Yangfan Zhou 0002
ACM Trans. Storage7
2024 rCanary: Detecting Memory Leaks Across Semi-Automated Memory Management Boundary in Rust
abstract
Rust is an effective system programming language that guarantees memory safety via compile-time verifications. It employs a novel ownership-based resource management model to facilitate automated deallocation. This model is anticipated to eliminate memory leaks. However, we observed that user intervention drives it into semi-automated memory management and makes it error-prone to cause leaks. In contrast to violating memory-safety guarantees restricted by theunsafekeyword, the boundary of leaking memory is implicit, and the compiler would not emit any warnings for developers. In this paper, we presentrCanary, a static, non-intrusive, and fully automated model checker to detect leaks across the semi-automated boundary. We design an encoder to abstract data with heap allocation and formalize a refined leak-free memory model based on boolean satisfiability. It can generate SMT-Lib2 format constraints for Rust MIR and is implemented as a Cargo component. We evaluaterCanaryby using flawed package benchmarks collected from the pull requests of open-source Rust projects. The results indicate that it is possible to recall all these defects with acceptable false positives. We further apply our tool to more than 1,200 real-world crates from crates.io and GitHub, identifying 19 crates having memory leaks. Our analyzer is also efficient, that costs 8.4 seconds per package.
Mohan Cui, Hui Xu 0009, Hongliang Tian, Yangfan Zhou 0002
IEEE Trans. Software Eng.4
2023 FUSEE: A Fully Memory-Disaggregated Key-Value Store
Jiacheng Shen, Pengfei Zuo, Xuchuan Luo, Yuxin Su 0001, Yangfan Zhou 0002, Michael R. Lyu
FAST6
2023 Revealing Performance Issues in Server-Side WebAssembly Runtimes Via Differential Testing
abstract
WebAssembly (Wasm) is a bytecode format originally serving as a compilation target for Web applications. It has recently been used increasingly on the server side, e.g., providing a safer, faster, and more portable alternative to Linux containers. With the popularity of server-side Wasm applications, it is essential to study performance issues (i.e., abnormal latency) in Wasm runtimes, as they may cause a significant impact on server-side applications. However, there is still a lack of attention to performance issues in server-side Wasm runtimes. In this paper, we design a novel differential testing approach WarpDiff to identify performance issues in server-side Wasm runtimes. The key insight is that in normal cases, the execution time of the same test case on different Wasm runtimes should follow an oracle ratio. We identify abnormal cases where the execution time ratio significantly deviates from the oracle ratio and subsequently locate the Wasm runtimes that cause the performance issues. We apply WarpDiff to test five popular server-side Wasm runtimes using 123 test cases from the LLVM test suite and demonstrate the top 10 abnormal cases we identified. We further conduct an in-depth analysis of these abnormal cases and summarize seven performance issues, all of which have been confirmed by the developers. We hope our work can inspire future investigation on improving Wasm runtime implementation and thus promoting the development of server-side Wasm applications.
Shuyao Jiang, Ruiying Zeng, Zihao Rao, Jiazhen Gu, Yangfan Zhou 0002, Michael R. Lyu
ASE5
2023 SMART: A High-Performance Adaptive Radix Tree for Disaggregated Memory
Xuchuan Luo, Pengfei Zuo, Jiacheng Shen, Jiazhen Gu, Xin Wang 0002, Michael R. Lyu, Yangfan Zhou 0002
OSDI7
2023 Appaction: Automatic GUI Interaction for Mobile Apps via Holistic Widget Perception
abstract
In industrial practice, GUI (Graphic User Interface) testing of mobile apps still inevitably relies on huge manual efforts. The major efforts are those on understanding the GUIs, so that testing scripts can be written accordingly. Quality assurance could therefore be very labor-intensive, especially for modern commercial mobile apps, where one may include tremendous, diverse, and complex GUIs, e.g., those for placing orders of different commercial items. To reduce such human efforts, we propose Appaction, a learning-based automatic GUI interaction approach we developed for Meituan, one of the largest E-commerce providers with over 600 million users. Appaction can automatically analyze the target GUI and understand what each input of the GUI is about, so that corresponding valid inputs can be entered accordingly. To this end, Appaction adopts a multi-modal model to learn from human experiences in perceiving a GUI. This allows it to infer corresponding valid input events that can properly interact with the GUI. In this way, the target app can be effectively exercised. We present our experiences in Meituan on applying Appaction to popular commercial apps. We demonstrate the effectiveness of Appaction in GUI analysis, and it can perform correct interactions for numerous form pages.
Yongxiang Hu 0003, Jiazhen Gu, Shuqing Hu, Yu Zhang 0165, Chaoyi Chen, Yangfan Zhou 0002
ESEC/SIGSOFT FSE8
2023 Ditto: An Elastic and Adaptive Memory-Disaggregated Caching System
abstract
In-memory caching systems are fundamental building blocks in cloud services. However, due to the coupled CPU and memory on monolithic servers, existing caching systems cannot elastically adjust resources in a resource-efficient and agile manner. To achieve better elasticity, we propose to port in-memory caching systems to the disaggregated memory (DM) architecture, where compute and memory resources are decoupled and can be allocated flexibly. However, constructing an elastic caching system on DM is challenging since accessing cached objects with CPU-bypass remote memory accesses hinders the execution of caching algorithms. Moreover, the elastic changes of compute and memory resources on DM affect the access patterns of cached data, compromising the hit rates of caching algorithms. We design Ditto, the first caching system on DM, to address these challenges. Ditto first proposes a client-centric caching framework to efficiently execute various caching algorithms in the compute pool of DM, relying only on remote memory accesses. Then, Ditto employs a distributed adaptive caching scheme that adaptively switches to the best-fit caching algorithm in real-time based on the performance of multiple caching algorithms to improve cache hit rates. Our experiments show that Ditto effectively adapts to the changing resources on DM and outperforms the state-of-the-art caching systems by up to 3.6× in real-world workloads and 9× in YCSB benchmarks.
Jiacheng Shen, Pengfei Zuo, Xuchuan Luo, Yuxin Su 0001, Jiazhen Gu, Yangfan Zhou 0002, Michael R. Lyu
SOSP7
2023 SafeDrop: Detecting Memory Deallocation Bugs of Rust Programs via Static Data-flow Analysis
abstract
Rust is an emerging programming language that aims to prevent memory-safety bugs. However, the current design of Rust also brings side effects, which may increase the risk of memory-safety issues. In particular, it employs ownership-based resource management and enforces automatic deallocation of unused resources without using the garbage collector. It may therefore falsely deallocate reclaimed memory and lead to use-after-free or double-free issues. In this article, we study the problem of invalid memory deallocation and propose SafeDrop , a static path-sensitive data-flow analysis approach to detect such bugs. Our approach analyzes each function of a Rust crate iteratively in a flow-sensitive and field-sensitive way. It leverages a modified Tarjan algorithm to achieve scalable path-sensitive analysis and a cache-based strategy for efficient inter-procedural analysis. We have implemented our approach and integrated it into the Rust compiler. Experiment results show that the approach can successfully detect all such bugs in our experiments with a limited number of false positives and incurs a very small overhead compared to the original compilation time.
Mohan Cui, Hui Xu 0009, Yangfan Zhou 0002
ACM Trans. Softw. Eng. Methodol.4
2023 Adaptive Data Placement in Multi-Cloud Storage: A Non-Stationary Combinatorial Bandit Approach
abstract
Multi-cloud storage is recently a viable approach to solve the vendor lock-in, reliability, and security issues in cloud storage systems. As a key concern, data placement influences the cost and performance of storage services. Yet, in practice it remains challenging to address the huge solution space. Previous studies typically focus on constructing efficient data placement schemes based on the predicted pattern of workloads or assuming fully a-priori known network conditions. They cannot be easily applied in multi-cloud storage scenarios, which typically involve dynamic network conditions and time-varying workloads. To this end, we formulate the data placement optimization in a combinatorial multi-arm bandit (CMAB) perspective and solve it by learning placement strategy online. In contrast to a stationary setting where reward distributions are unknown but identical over time, we consider a realistic multi-cloud environment with non-stationary conditions, i.e., reward distributions change over time. To swiftly accommodate this, we propose an adaptive window combinatorial upper confidence bound based data placement (AW-CUCB-DP) scheme to reduce latency and cost. In AW-CUCB-DP, a simple and efficient change detector, i.e.,Page-Hinkley testwith forgetting mechanism (FM-PHT), is employed to enable variable-size sliding windows to handle both gradual and abrupt variations in network conditions or workloads. We establish that AW-CUCB-DP is asymptotically optimal in the non-stationary multi-cloud environment. Trace-driven experiments further verify that our scheme outperforms alternatives, especially in highly dynamic environments.
Li Li 0111, Jiajie Shen, Bochun Wu, Yangfan Zhou 0002, Xin Wang 0134, Keqin Li 0001
IEEE Trans. Parallel Distributed Syst.4
2023 Towards Usable Neural Comment Generation via Code-Comment Linkage Interpretation: Method and Empirical Study
abstract
Code comment is important to facilitate code comprehension for developers. Recent studies suggest to generate comments automatically with deep learning, in particular, based on neural machine translation models. However, such a promising Neural Comment Generation (NCG) technique suffers from unsatisfactory performance, as well as poor usability, i.e., developers cannot easily understand and modify the auto-generated comments. This paper suggests that a proper interpretation of how the comments are generated can significantly improve the usability of NCG approaches. We propose a novel model-independent framework, namely CCLink, to interpret the auto-generated comments. CCLink generates a set of code mutants and obtains their corresponding comments. Based on these data, several contribution mining algorithms are designed to infer the key elements in code that contributes to the generation of the key phrases in the comments. The links between code and its auto-generated comment can thus be constructed. This in turn allows CCLink to visualize the links as the comment interpretations to developers. It greatly facilitates manual verification and correction of the comments. We examine the performance of CCLink with different contribution mining algorithms, NCG approaches, and real-world datasets. We also conduct an empirical study on 32 experienced Java programmers to evaluate the effectiveness of CCLink. The results show that CCLink is promising in making NCG more usable with a proper interpretation of the auto-generated comments.
Shuyao Jiang, Jiacheng Shen, Yue Yu 0001, Yangfan Zhou 0002
IEEE Trans. Software Eng.6
2022 Muffin: Testing Deep Learning Libraries via Neural Architecture Fuzzing
abstract
Deep learning (DL) techniques are proven effective in many challenging tasks, and become widely-adopted in practice. However, previous work has shown that DL libraries, the basis of building and executing DL models, contain bugs and can cause severe consequences. Unfortunately, existing testing approaches still cannot comprehensively exercise DL libraries. They utilize existing trained models and only detect bugs in model inference phase. In this work we propose Muffin to address these issues. To this end, Muffin applies a specifically-designed model fuzzing approach, which allows it to generate diverse DL models to explore the target library, instead of relying only on existing trained models. Muffin makes differential testing feasible in the model training phase by tailoring a set of metrics to measure the inconsistencies between different DL libraries. In this way, Muffin can best exercise the library code to detect more bugs. To evaluate the effectiveness of Muffin, we conduct experiments on three widely-used DL libraries. The results demonstrate that Muffin can detect 39 new bugs in the latest release versions of popular DL libraries, including Tensorflow, CNTK, and Theano.
Jiazhen Gu, Xuchuan Luo, Yangfan Zhou 0002, Xin Wang 0002
ICSE3
2022 How resource utilization influences UI responsiveness of Android software
Jiaojiao Fu, Yaohui Wang 0003, Yangfan Zhou 0002, Xin Wang 0002
Inf. Softw. Technol.3
2022 Unveiling High-speed Follow-up Consultation for Chronic Disease Treatment: A Pediatric Hospital Case in China
abstract
In developing countries suffering from a severe shortage of physicians, the follow-up clinical consultation of long-term chronic care becomes a high-speed collaborative work, which has to be conducted in a few minutes. Although much existing work studies how time factors affect physicians' workflows and requirements for information systems, high-speed chronic care is yet to be well investigated. This work bridges the gap by presenting a case of a pediatric hospital in China. We focus on the processes of follow-up consultations, the factors enabling physicians to complete consultation in several minutes, as well as the challenges faced by physicians and patients. Through observations and interviews, we find that physicians conduct multiple tasks (information acquisition, patient-provider communication, and medical data documentation) simultaneously to reduce the consultation duration. Adopting an information summary alternative is the key to fast information acquisition. Templates and references in EMR contribute to rapid documentation and prescription. However, multitasking brings physicians a heavy cognitive load. It also severely compresses the duration of patient-provider communication. As a result, some of the patients' needs, especially emotional ones, are neglected. Based on these findings, we discuss the characteristics and requirements of high-speed chronic care and accordingly propose design suggestions.
Jiaojiao Fu, Yangfan Zhou 0002, Xin Wang 0002
Proc. ACM Hum. Comput. Interact.2
2022 Memory-Safety Challenge Considered Solved? An In-Depth Study with All Rust CVEs
abstract
Rust is an emerging programming language that aims at preventing memory-safety bugs without sacrificing much efficiency. The claimed property is very attractive to developers, and many projects start using the language. However, can Rust achieve the memory-safety promise? This article studies the question by surveying 186 real-world bug reports collected from several origins, which contain all existing Rust common vulnerability and exposures (CVEs) of memory-safety issues by 2020-12-31. We manually analyze each bug and extract their culprit patterns. Our analysis result shows that Rust can keep its promise that all memory-safety bugs require unsafe code, and many memory-safety bugs in our dataset are mild soundness issues that only leave a possibility to write memory-safety bugs without unsafe code. Furthermore, we summarize three typical categories of memory-safety bugs, including automatic memory reclaim, unsound function, and unsound generic or trait. While automatic memory claim bugs are related to the side effect of Rust newly-adopted ownership-based resource management scheme, unsound function reveals the essential challenge of Rust development for avoiding unsound code, and unsound generic or trait intensifies the risk of introducing unsoundness. Based on these findings, we propose two promising directions toward improving the security of Rust development, including several best practices of using specific APIs and methods to detect particular bugs involving unsafe code. Our work intends to raise more discussions regarding the memory-safety issues of Rust and facilitate the maturity of the language.
Hui Xu 0009, Zhuangbin Chen, Mingshen Sun, Yangfan Zhou 0002, Michael R. Lyu
ACM Trans. Softw. Eng. Methodol.4
2021 Defuse: A Dependency-Guided Function Scheduler to Mitigate Cold Starts on FaaS Platforms
abstract
Function-as-a-Service (FaaS) is becoming a prevalent paradigm in developing cloud applications. With FaaS, clients can develop applications as serverless functions, leaving the burden of resource management to cloud providers. However, FaaS platforms suffer from the performance degradation caused by the cold starts of serverless functions. Cold starts happen when serverless functions are invoked before they have been loaded into the memory. The problem is unavoidable because the memory in datacenters is typically too limited to hold all serverless functions simultaneously. The latency of cold function invocations will greatly degenerate the performance of FaaS platforms. Currently, FaaS platforms employ various scheduling methods to reduce the occurrences of cold starts. However, they do not consider the ubiquitous dependencies between serverless functions. Observing the potential of using dependencies to mitigate cold starts, we propose Defuse, a Dependency-guided Function Scheduler on FaaS platforms. Specifically, Defuse identifies two types of dependencies between serverless functions, i.e., strong dependencies and weak ones. It uses frequent pattern mining and positive point-wise mutual information to mine such dependencies respectively from function invocation histories. In this way, Defuse constructs a function dependency graph. The connected components (i.e., dependent functions) on the graph can be scheduled to diminish the occurrences of cold starts. We evaluate the effectiveness of Defuse by applying it to an industrial serverless dataset. The experimental results show that Defuse can reduce 22% of memory usage while having a 35% decrease in function cold-start rates compared with the state-of-the-art method.
Jiacheng Shen, Yuxin Su 0001, Yangfan Zhou 0002, Michael R. Lyu
ICDCS4
2021 Fast Outage Analysis of Large-scale Production Clouds with Service Correlation Mining
abstract
Cloud-based services are surging into popularity in recent years. However, outages, i.e., severe incidents that always impact multiple services, can dramatically affect user experience and incur severe economic losses. Locating the root-cause service, i.e., the service that contains the root cause of the outage, is a crucial step to mitigate the impact of the outage. In current industrial practice, this is generally performed in a bootstrap manner and largely depends on human efforts: the service that directly causes the outage is identified first, and the suspected root cause is traced back manually from service to service during diagnosis until the actual root cause is found. Unfortunately, production cloud systems typically contain a large number of interdependent services. Such a manual root cause analysis is often time-consuming and labor-intensive. In this work, we propose COT, the first outage triage approach that considers the global view of service correlations. COT mines the correlations among services from outage diagnosis data. After learning from historical outages, COT can infer the root cause of emerging ones accurately. We implement COT and evaluate it on a real-world dataset containing one year of data collected from Microsoft Azure, one of the representative cloud computing platforms in the world. Our experimental results show that COT can reach a triage accuracy of 82.1%-83.5%, which outperforms the state-of-the-art triage approach by 28.0%-29.7%.
Yaohui Wang 0003, Guo-Zheng Li 0001, Yu Kang 0006, Yangfan Zhou 0002, Hongyu Zhang 0002, Feng Gao 0022, Jeffrey Sun, Pochian Lee, Zhangwei Xu, Pu Zhao 0004, Bo Qiao 0001, Liqun Li, Xu Zhang 0024, Qingwei Lin
ICSE5
2021 Boosting symbolic execution via constraint solving time prediction (experience paper)
abstract
Symbolic execution is an essential approach for automated test case generation. However, the approach is generally not scalable to large programs. One critical reason is that the constraint solving problems in symbolic execution are generally hard. Consequently, the symbolic execution process may get stuck in solving such hard problems. To mitigate this issue, symbolic execution tools generally rely on a timeout threshold to terminate the solving. Such a timeout is generally set to a fixed, predefined value, e.g., five minutes in angr. Nevertheless, how to set a proper timeout is critical to the tool’s efficiency. This paper proposes an approach to tackle the problem by predicting the time required for solving a constraint model so that the symbolic execution engine could base on the information to determine whether to continue the current solving process. Due to the cost of the prediction itself, our approach triggers the predictor only when the solving time has exceeded a relatively small value. We have shown that such a predictor can achieve promising performance with several different machine learning models and datasets. By further employing an adaptive design, the predictor can achieve an F1-score ranging from 0.743 to 0.800 on these datasets. We then apply the predictor to eight programs and conduct simulation experiments. Results show that the efficiency of constraint solving for symbolic execution can be improved by 1.25x to 3x, depending on the distribution of the hardness of their constraint models.
Sicheng Luo, Hui Xu 0009, Yanxiang Bi, Xin Wang 0002, Yangfan Zhou 0002
ISSTA5
2021 RULF: Rust Library Fuzzing via API Dependency Graph Traversal
abstract
Robustness is a key concern for Rust library development because Rust promises no risks of undefined behaviors if developers use safe APIs only. Fuzzing is a practical approach for examining the robustness of programs. However, existing fuzzing tools are not directly applicable to library APIs due to the absence of fuzz targets. It mainly relies on human efforts to design fuzz targets case by case which is labor-intensive. To address this problem, this paper proposes a novel automated fuzz target generation approach for fuzzing Rust libraries via API dependency graph traversal. We identify several essential requirements for library fuzzing, including validity and effectiveness of fuzz targets, high API coverage, and efficiency. To meet these requirements, we first employ breadth-first search with pruning to find API sequences under a length threshold, then we backward search longer sequences for uncovered APIs, and finally we optimize the sequence set as a set covering problem. We implement our fuzz target generator and conduct fuzzing experiments with AFL++ on several real-world popular Rust projects. Our tool finally generates 7 to 118 fuzz targets for each library with API coverage up to 0.92. We exercise each target with a threshold of 24 hours and find 30 previously-unknown bugs from seven libraries.
Hui Xu 0009, Yangfan Zhou 0002
ASE3
2020 Towards intelligent incident management: why we need it and how we make it
abstract
The management of cloud service incidents (unplanned interruptions or outages of a service/product) greatly affects customer satisfaction and business revenue. After years of efforts, cloud enterprises are able to solve most incidents automatically and timely. However, in practice, we still observe critical service incidents that occurred in an unexpected manner and orchestrated diagnosis workflow failed to mitigate them. In order to accelerate the understanding of unprecedented incidents and provide actionable recommendations, modern incident management system employs the strategy of AIOps (Artificial Intelligence for IT Operations). In this paper, to provide a broad view of industrial incident management and understand the modern incident management system, we conduct a comprehensive empirical study spanning over two years of incident management practices at Microsoft. Particularly, we identify two critical challenges (namely, incomplete service/resource dependencies and imprecise resource health assessment) and investigate the underlying reasons from the perspective of cloud system design and operations. We also present IcM BRAIN, our AIOps framework towards intelligent incident management, and show its practical benefits conveyed to the cloud services of Microsoft.
Zhuangbin Chen, Yu Kang 0006, Liqun Li, Xu Zhang 0024, Hongyu Zhang 0002, Hui Xu 0009, Yangfan Zhou 0002, Jeffrey Sun, Zhangwei Xu, Yingnong Dang, Feng Gao 0022, Pu Zhao 0004, Bo Qiao 0001, Qingwei Lin, Dongmei Zhang 0001, Michael R. Lyu
ESEC/SIGSOFT FSE7
2020 Efficient incident identification from multi-dimensional issue reports via meta-heuristic search
abstract
In large-scale cloud systems, unplanned service interruptions and outages may cause severe degradation of service availability. Such incidents can occur in a bursty manner, which will deteriorate user satisfaction. Identifying incidents rapidly and accurately is critical to the operation and maintenance of a cloud system. In industrial practice, incidents are typically detected through analyzing the issue reports, which are generated over time by monitoring cloud services. Identifying incidents in a large number of issue reports is quite challenging. An issue report is typically multi-dimensional: it has many categorical attributes. It is difficult to identify a specific attribute combination that indicates an incident. Existing methods generally rely on pruning-based search, which is time-consuming given high-dimensional data, thus not practical to incident detection in large-scale cloud systems. In this paper, we propose MID (Multi-dimensional Incident Detection), a novel framework for identifying incidents from large-amount, multi-dimensional issue reports effectively and efficiently. Key to the MID design is encoding the problem into a combinatorial optimization problem. Then a specific-tailored meta-heuristic search method is designed, which can rapidly identify attribute combinations that indicate incidents. We evaluate MID with extensive experiments using both synthetic data and real-world data collected from a large-scale production cloud system. The experimental results show that MID significantly outperforms the current state-of-the-art methods in terms of effectiveness and efficiency. Additionally, MID has been successfully applied to Microsoft's cloud systems and helped greatly reduce manual maintenance effort.
Jiazhen Gu, Chuan Luo 0002, Si Qin, Bo Qiao 0001, Qingwei Lin, Hongyu Zhang 0002, Ze Li 0005, Yingnong Dang, Shaowei Cai 0001, Wei Wu 0011, Yangfan Zhou 0002, Murali Chintalapati, Dongmei Zhang 0001
ESEC/SIGSOFT FSE11
2020 Efficient customer incident triage via linking with system incidents
abstract
In cloud service systems, customers will report the service issues they have encountered to cloud service providers. Despite many issues can be handled by the support team, sometimes the customer issues can not be easily solved, thus raising customer incidents. Quick troubleshooting of a customer incident is critical. To this end, a customer incident should be assigned to its responsible team accurately in a timely manner.
Jiazhen Gu, Jiaqi Wen, Pu Zhao 0004, Chuan Luo 0002, Yu Kang 0006, Yangfan Zhou 0002, Jeffrey Sun, Zhangwei Xu, Bo Qiao 0001, Liqun Li, Qingwei Lin, Dongmei Zhang 0001
ESEC/SIGSOFT FSE7
2020 Layered obfuscation: a taxonomy of software obfuscation techniques for layered security
abstract
Abstract Software obfuscation has been developed for over 30 years. A problem always confusing the communities is what security strength the technique can achieve. Nowadays, this problem becomes even harder as the software economy becomes more diversified. Inspired by the classic idea of layered security for risk management, we propose layered obfuscation as a promising way to realize reliable software obfuscation. Our concept is based on the fact that real-world software is usually complicated. Merely applying one or several obfuscation approaches in an ad-hoc way cannot achieve good obscurity. Layered obfuscation, on the other hand, aims to mitigate the risks of reverse software engineering by integrating different obfuscation techniques as a whole solution. In the paper, we conduct a systematic review of existing obfuscation techniques based on the idea of layered obfuscation and develop a novel taxonomy of obfuscation techniques. Following our taxonomy hierarchy, the obfuscation strategies under different branches are orthogonal to each other. In this way, it can assist developers in choosing obfuscation techniques and designing layered obfuscation solutions based on their specific requirements.
Hui Xu 0009, Yangfan Zhou 0002, Jiang Ming 0002, Michael R. Lyu
Cybersecur.2
2020 Information Summary for Chronic Disease Treatment: A Pediatric Hospital Case in China
abstract
The electronic medical record (EMR) systems face many challenges in supporting chronic disease treatment, especially in medical information summary. Many solutions have recently been proposed for hospitals in developed countries. However, these solutions maybe not suitable for hospitals in developing countries because their workflow and patterns may be quite different due to their resource limitations, especially in shortage of physicians. Investigating the information summary alternatives in treating chronic diseases in such hospitals can shed light on EMR system design, especially on that for developing countries. We study one of the best pediatric hospitals in China. It suffers from a severe shortage of physicians. We introduce its information summary alternative, \fs. In particular, we study how and why pediatricians treat chronic diseases with the sheet. Our work unveils the intense work patterns and their corresponding information requirements of hospitals in China. We also demonstrate the characteristics of the \fs\ and accordingly discuss how EMR systems can be optimized.
Jiaojiao Fu, Yangfan Zhou 0002, Xin Wang 0002
Proc. ACM Hum. Comput. Interact.2
2020 Benchmarking the Capability of Symbolic Execution Tools with Logic Bombs
abstract
Symbolic execution has become an indispensable technique for software testing and program analysis. However, since several symbolic execution tools are presently available off-the-shelf, there is a need for a practical benchmarking approach. This paper introduces a fresh approach that can help benchmark symbolic execution tools in a fine-grained and efficient manner. The approach evaluates the performance of such tools against known challenges faced by general symbolic execution techniques, e.g., floating-point numbers and symbolic memories. We first survey related papers and systematize the challenges of symbolic execution. We extract 12 distinct challenges from the literature and categorize them into two categories: symbolic-reasoning challenges and path-explosion challenges. Next, we develop a dataset of logic bombs and a framework for benchmarking symbolic execution tools automatically. For each challenge, our dataset contains several logic bombs, each addressing a specific challenging problem. Triggering one or more logic bombs confirms that the symbolic execution tool in question is able to handle the corresponding problem. Real-world experiments with three popular symbolic execution tools, namely, KLEE, angr, and Triton have shown that our approach can reveal the capabilities and limitations of the tools in handling specific issues accurately and efficiently. The benchmarking process generally takes only a few dozens of minutes to evaluate a tool. We have released our dataset on GitHub as open source, with an aim to better facilitate the community to conduct future work on benchmarking symbolic execution tools.
Hui Xu 0009, Zirui Zhao, Yangfan Zhou 0002, Michael R. Lyu
IEEE Trans. Dependable Secur. Comput.3
2019 Textout: Detecting Text-Layout Bugs in Mobile Apps via Visualization-Oriented Learning
abstract
Layout bugs commonly exist in mobile apps. Due to the fragmentation issues of smartphones, a layout bug may occur only on particular versions of smartphones. It is quite challenging to detect such bugs for state-of-the-art commercial automated testing platforms, although they can test an app with thousands of different smartphones in parallel. The main reason is that typical layout bugs neither crash an app nor generate any error messages. In this paper, we present our work for detecting text-layout bugs, which account for a large portion of layout bugs. We model text-layout bug detection as a classification problem. This then allows us to address it with sophisticated image processing and machine learning techniques. To this end, we propose an approach which we call Textout. Textout takes screenshots as its input and adopts a specifically-tailored text detection method and a convolutional neural network (CNN) classifier to perform automatic text-layout bug detection. We collect 33,102 text-region images as our training dataset and verify the effectiveness of our tool with 1,481 text-region images collected from real-world apps. Textout achieves an AUC (area under the curve) of 0.956 on the test dataset and shows an acceptable overhead. The dataset is open-source released for follow-up research.
Yaohui Wang 0003, Hui Xu 0009, Yangfan Zhou 0002, Michael R. Lyu, Xin Wang 0002
ISSRE3
2019 Understanding I/O performance of IPFS storage: a client's perspective
abstract
IPFS has surged into popularity in recent years. It organizes user data as multiple objects where users can obtain the objects according to their Content IDentifiers (CIDs). As a storage system, it is of great importance to understand its data I/O performance. But existing work still lacks such a comprehensive study. In this work, we deploy an IPFS storage system with geographically-distributed storage nodes on Amazon EC2. We then conduct extensive experiments to evaluate the performance of data I/O operations from a client's perspective. We find that the access patterns of I/O operations (e.g., request size) severely affect the I/O performance, since IPFS typically uses multiple I/O strategies to perform different I/O requests. Moreover, for the read operations, IPFS requires to resolve remote nodes and downloading objects via the internet. Our experimental study reveals that both resolving and downloading operations can become bottlenecks. Our results can shed light to optimizing IPFS in avoiding high-latency I/O operations.
Jiajie Shen, Yangfan Zhou 0002, Xin Wang 0002
IWQoS3
2019 Component-based permission management of Android applications
abstract
Summary Most Android applications include third‐party libraries (3PLs) to make revenues, to facilitate their development, and to track user behaviors. 3PLs generally require specific permissions to realize their functionalities. Current Android systems manage permissions in app (process) granularity. As a result, the permission sets of apps with 3PLs (3PL‐apps) may be augmented, introducing overprivilege risks. In this paper, we firstly study how severe the problem is by analyzing the permission sets of 27 718 real‐world Android apps with and without 3PLs downloaded in both 2016 and 2017. We find that the usage of 3PLs and the permissions required by 3PL‐apps have increased over time. As a result, the possibility of overprivilege risks increases. We then propose Perman, a fine‐grained permission management mechanism for Android. Perman isolates the permissions of the host app and those of the 3PLs through dynamic code instrumentation. It allows users to manage permission requests of different modules of 3PL‐apps during app runtime. Unlike existing tools, Perman does not need to redesign Android apps and systems. Therefore, it can be applied to millions of existing apps and various Android devices. We conduct experiments to evaluate the effectiveness and efficiency of Perman. The experimental results verify that Perman is capable of managing permission requests of the host app and those of the 3PLs. We also confirm that the overhead introduced by Perman is comparable to that by existing commercial permission management tools.
Jiaojiao Fu, Yangfan Zhou 0002, Xin Wang 0002
Softw. Pract. Exp.2
2018 Manufacturing Resilient Bi-Opaque Predicates Against Symbolic Execution
abstract
Control-flow obfuscation increases program complexity by semantic-preserving transformation. Opaque predicates are essential gadgets to achieve such transformation. However, we observe that real-world opaque predicates are generally very simple and engage little security consideration. Recently, such insecure opaque predicates have been severely attacked by symbolic execution-based adversaries and jeopardize the security of control-flow obfuscation. This paper, therefore, proposes symbolic opaque predicates which can be resilient to symbolic execution-based adversaries. We design a general framework to compose such opaque predicates, which requires introducing challenging symbolic analysis problems (e.g., symbolic memory) in each opaque predicate. In this way, we may mislead symbolic execution engines into reaching false conclusions. We observe a novel bi-opaque property about symbolic opaque predicates, which can incur not only false negative issues but also false positive issues to attackers. To evaluate the efficacy of our idea, we have implemented a prototype obfuscation tool based on Obfuscator-LLVM and conduct experiments with real-world programs. Our evaluation results show that symbolic opaque predicates demonstrate excellent resilience to prevalent symbolic execution engines, such as BAP, Triton, and Angr. Moreover, although the costs of symbolic opaque predicates may vary for different problem settings, some predicates can be very efficient. Therefore, our framework is both secure and usable. Users can follow the framework to introduce symbolic opaque predicates into their obfuscation tools and made them more powerful.
Hui Xu 0009, Yangfan Zhou 0002, Yu Kang 0006, Fengzhi Tu, Michael R. Lyu
DSN2
2018 Mobile Cloud-of-Clouds Storage Made Efficient: A Network Coding Based Approach
abstract
Cloud-of-clouds storage is a viable means to ensure security and reliability of distributed data storage, where data are encrypted, encoded, and stored in multiple clouds. However, it is a great challenge to adopt such a paradigm in mobile devices (e.g., smartphone). Mobile devices are generally incapable to perform the heavy-weight operations (i.e., data encryption, encoding, and transmission) required in such a paradigm, given the limited resources in such devices. This paper focuses on addressing this challenge, i.e., improving data storage performance in mobile cloud-of-clouds storage systems. The key of our proposal is to allow the low-capability mobile devices to offload the computational and transmission overhead to the clouds. In other words, we propose a Network Coding based Cloud-of-clouds Storage (NCCS) scheme, where the clouds can encode and exchange data collaboratively. We consider two state-of-the-art cloud-of-clouds storage approaches, i.e., AONT-RS and CAONT-RS, as example cases to deploy our scheme. Accordingly, we propose their network coding-based enhancements, namely NAONT-RS and NCAONT-RS. We implement a prototype cloud-of-clouds system to verify the efficiency of our proposal. We deploy the prototype on Microsoft Azure and conduct extensive experiments with real-world traces. The experimental results show that NAONT-RS and NCAONT-RS can reduce the time of data storage process by up to 50% and improve the throughput by up to 110% compared with their original versions, i.e., AONT-RS and CAONT-RS.
Jiajie Shen, Yangfan Zhou 0002, Xin Wang 0002
SRDS3
2018 Efficient Scheduling for Multi-Block Updates in Erasure Coding Based Storage Systems
abstract
This paper considers the problem of how to reduce the I/O overhead of data update operations in erasure coding based storage systems. To this end, we first analyze the I/O overhead of update operations with current update approaches. We find the key to reduce such I/O overhead is designing a scheduling algorithm to construct the sequence of update operations. Such an algorithm needs to execute with a time limit, since update requests work under a stringent latency constraint. To quickly schedule the order of update operations, we propose an efficient algorithm, namely UCODR. Our theoretical analysis verifies that UCODR can effectively reduce the I/O overhead of update operations when multiple blocks are updated. To further confirm its effectiveness, we implement a prototype storage system to deploy UCODR with different erasure codes. Extensive experiments are conducted on the prototype storage system with real-world traces. The experimental results show that UCODR can reduce the time of update operations by up to 35 percent and improve the throughput of the storage system by up to 67 percent, compared with the state-of-the-art update approaches.
Jiajie Shen, Jiazhen Gu, Yangfan Zhou 0002, Xin Wang 0002
IEEE Trans. Computers4
2017 Concolic Execution on Small-Size Binaries: Challenges and Empirical Study
abstract
Concolic execution has achieved great success in many binary analysis tasks. However, it is still not a primary option for industrial usage. A well-known reason is that concolic execution cannot scale up to large-size programs. Many research efforts have focused on improving its scalability. Nonetheless, we find that, even when processing small-size programs, concolic execution suffers a great deal from the accuracy and scalability issues. This paper systematically investigates the challenges that can be introduced even by small-size programs, such as symbolic array and symbolic jump. We further verify that the proposed challenges are non-trivial via real-world experiments with three most popular concolic execution tools: BAP, Triton, and Angr. Among a set of 22 logic bombs we designed, Angr can solve only four cases correctly, while BAP and Triton perform much worse. The results imply that current tools are still primitive for practical industrial usage. We summarize the reasons and release the bombs as open source to facilitate further study.
Hui Xu 0009, Yangfan Zhou 0002, Yu Kang 0006, Michael R. Lyu
DSN2
2017 Perman: Fine-Grained Permission Management for Android Applications
abstract
Third-party libraries (3PLs) are widely introduced into Android apps and they typically request permissions for their own functionalities. Current Android systems manage permissions in process (app) granularity. Hence, the host app and the 3PLs share the same permission set. 3PL-apps may therefore introduce security risks. Separating the permission sets of the 3PLs and those of the host app are critical to alleviate such security risks. In this paper, we provide Perman, a tool that allows users to manage permissions of different modules (i.e., a 3PL or the host app) of an app at runtime. Perman relies on dynamic code instrumentation to intercept permission requests, and accordingly provide a policy-based permission control. Unlike existing tools that generally require to redesign 3PL-apps, it can thus be applied to the existing apps in market. We evaluate Perman on real-world apps. The experiment results verify its effectiveness in fine-grained permission management.
Jiaojiao Fu, Yangfan Zhou 0002, Yu Kang 0006, Xin Wang 0002
ISSRE2
2016 How does regression test prioritization perform in real-world software evolution?
abstract
In recent years, researchers have intensively investigated various topics in test prioritization, which aims to re-order tests to increase the rate of fault detection during regression testing. While the main research focus in test prioritization is on proposing novel prioritization techniques and evaluating on more and larger subject systems, little effort has been put on investigating the threats to validity in existing work on test prioritization. One main threat to validity is that existing work mainly evaluates prioritization techniques based on simple artificial changes on the source code and tests. For example, the changes in the source code usually include only seeded program faults, whereas the test suite is usually not augmented at all. On the contrary, in real-world software development, software systems usually undergo various changes on the source code and test suite augmentation. Therefore, it is not clear whether the conclusions drawn by existing work in test prioritization from the artificial changes are still valid for real-world software evolution. In this paper, we present the first empirical study to investigate this important threat to validity in test prioritization. We reimplemented 24 variant techniques of both the traditional and time-aware test prioritization, and investigated the impacts of software evolution on those techniques based on the version history of 8 real-world Java programs from GitHub. The results show that for both traditional and time-aware test prioritization, test suite augmentation significantly hampers their effectiveness, whereas source code changes alone do not influence their effectiveness much.
Yafeng Lu, Yiling Lou, Shiyang Cheng 0002, Lingming Zhang 0001, Dan Hao 0001, Yangfan Zhou 0002, Lu Zhang 0023
ICSE6
2016 Cloud-of-Clouds Storage Made Efficient: A Pipeline-Based Approach
abstract
Cloud-of-clouds storage is a recent approach to improve the security and reliability of data storage for online applications. It encrypts and encodes the user data, and disperses the results to multiple clouds. Thus, the data can tolerate cloud failures, while cannot be inferred even when some clouds are compromised. However, efficiency is a well-known challenge to such a paradigm, since its data storing process (also known as the dispersal process) is time-consuming involving encryptions, encoding, and transmissions, posing a barrier to its wide application. How to speed up the dispersal process is yet to be well addressed. We observe that the dispersal process consists of two types of operations: calculation and transmission. We find that they can execute simultaneously. Hence, the process can be optimized with a pipelined architecture. To this end, we propose the pipelined versions of two state-of-the-art cloud-of-clouds storage approaches, i.e., AONT-RS and CAONT-RS. We implement both proposals and release them open-source online. To verify their effectiveness, extensive experiments are conducted on a prototype storage system with real-world traces. The results show that the pipelined architecture can improve the performance of the dispersal process.
Jiajie Shen, Jiazhen Gu, Yangfan Zhou 0002, Xin Wang 0003
ICWS3
2016 Experience Report: Detecting Poor-Responsive UI in Android Applications
abstract
Good user interface (UI) design is key to successful mobile apps. UI latency, which can be considered as the time between the commencement of a UI operation and its intended UI update, is a critical consideration for app developers. Current literature still lacks a comprehensive study on how much UI latency a user can tolerate or how to identify UI design defects that cause intolerably long UI latency. As a result, bad UI apps are still common in app markets, leading to extensive user complaints. This paper examines user expectations of UI latency, anddevelops a tool to pinpoint intolerable UI latency in Android apps. To this end, we design an app to conduct a user survey of app UI latency. Through the survey, we find the tendency between user patience and UI latency. Therefore a timely screen update (e.g., loading animations) is critical to heavy-weighted UI operations (i.e. those that incur a long execution time before the final UI update is available). We then design a tool that, by monitoring the UI inputs and updates, can detect apps that do not follow this criterion. The survey and the tool are open-source released on-line. We also apply the tool to many real-world apps. The results demonstrate the effectiveness of the tool in combating app UI design defects.
Yu Kang 0006, Yangfan Zhou 0002, Yixia Sun, Michael R. Lyu
ISSRE2
2016 Bandwidth-aware delayed repair in distributed storage systems
abstract
In data storage systems, data are typically stored in redundant storage nodes to ensure storage reliability. When storage nodes fail, with the help of the redundant nodes, the lost data can be restored in new storage nodes. Such a regeneration process may be aborted, since storage nodes may fail during the process. Therefore, reducing the time of regeneration process is a well-known challenge to improve the reliability of storage systems. Delayed repair is a typical repair scheme in real-world storage systems. It reduces the overhead of the regeneration process by recovering multiple node failures simultaneously. How to reduce the regeneration time of delayed repair is yet to be well addressed. Since available bandwidth is flowing in storage systems and the regeneration time is seriously affected by the available bandwidth, we find the key to solve this problem is determining the start time of the regeneration process. Via modeling this problem with Lyaponuv optimization framework, we propose an OMFR scheme to reduce the regeneration time. The experimental results show that OMFR scheme can reduce cumulative regeneration time by up to 78% compared with traditional delayed repair schemes.
Jiajie Shen, Jiazhen Gu, Yangfan Zhou 0002, Xin Wang 0003
IWQoS3
2016 DiagDroid: Android performance diagnosis via anatomizing asynchronous executions
abstract
Rapid UI responsiveness is a key consideration to Android app developers. However, the complicated concurrency model of Android makes it hard for developers to understand and further diagnose the UI performance. This paper presents DiagDroid, a tool specifically designed for Android UI performance diagnosis. The key notion of DiagDroid is that UI-triggered asynchronous executions contribute to the UI performance, and hence their performance and their runtime dependency should be properly captured to facilitate performance diagnosis. However, there are tremendous ways to start asynchronous executions, posing a great challenge to profiling such executions and their runtime dependency. To this end, we properly abstract five categories of asynchronous executions as the building basis. As a result, they can be tracked and profiled based on the specifics of each category with a dynamic instrumentation approach carefully tailored for Android. DiagDroid can then accordingly profile the asynchronous executions in a task granularity, equipping it with low-overhead and high compatibility merits. The tool is successfully applied in diagnosing 33 real-world open-source apps, and we find 14 of them contain 27 performance issues. It shows the effectiveness of our tool in Android UI performance diagnosis. The tool is open-source released online.
Yu Kang 0006, Yangfan Zhou 0002, Hui Xu 0009, Michael R. Lyu
SIGSOFT FSE2
2015 PAID: Prioritizing app issues for developers by tracking user reviews over versions
abstract
User review analysis is critical to the bug-fixing and version-modification process for app developers. Many research efforts have been put to user review mining in discovering app issues, including laggy user interface, high memory overhead, privacy leakage, etc. Existing exploration of app reviews generally depends on static collections. As a result, they largely ignore the fact that user reviews are tightly related to app versions. Furthermore, the previous approaches require a developer to spend much time on filtering out trivial comments and digesting the informative textual data. This would be labor-intensive especially to popular apps with tremendous reviews. In the paper, we target at designing a framework in Prioritizing App Issues for Developers (PAID) with minimal manual power and good accuracy. The PAID design is based on the fact that the issues presented in the level of phrase, i.e., a couple of consecutive words, can be more easily understood by developers than in long sentences. Hence, we aim at recommending phrase-level issues of an app to its developers by tracking reviews over the release versions of the app. To assist developers in better comprehending the app issues, PAID employs ThemeRiver to visualize the analytical results to developers. Finally, PAID also allows the developers to check the most related reviews, when they want to obtain a deep insight of a certain issue. In contrast to the traditional evaluation methods such as manual labeling or examining the discussion forum, our experimental study exploits the first-hand information from developers, i.e., app changelogs, to measure the performance of PAID. We analyze millions of user reviews from 18 apps with 117 app versions and the results show that the prioritized issues generated by PAID match the official changelogs with high precision.
Cuiyun Gao 0001, Baoxiang Wang 0001, Pinjia He, Jieming Zhu, Yangfan Zhou 0002, Michael R. Lyu
ISSRE5
2015 SpyAware: Investigating the privacy leakage signatures in app execution traces
abstract
A new security problem on smartphones is the wide spread of spyware nested in apps, which occasionally and silently collects user's private data in the background. The state-of-the-art work for privacy leakage detection is dynamic taint analysis, which, however, suffers usability issues because it requires flashing a customized system image to track the taint propagation and consequently incurs great overhead. Through a real-world privacy leakage case study, we observe that the spyware behaviors share some common features during execution, which may further indicate a correlation between the data flow of privacy leakage and some specific features of program execution traces. In this work, we examine such a hypothesis using the newly proposed SpyAware framework, together with a customized TaintDroid as the ground truth. SpyAware includes a profiler to automatically profile app executions in binder calls and system calls, a feature extractor to extract feature vectors from execution traces, and a classifier to train and predict spyware executions based on the feature vectors. We conduct an evaluation experiment with 100 popular apps downloaded from Google Play. Experimental results show that our approach can achieve promising performance with 67.4% accuracy in detecting device id spyware executions and 78.4% in recognizing location spyware executions.
Hui Xu 0009, Yangfan Zhou 0002, Cuiyun Gao 0001, Yu Kang 0006, Michael R. Lyu
ISSRE2
2015 Towards Operational Cost Minimization in Hybrid Clouds for Dynamic Resource Provisioning with Delay-Aware Optimization
abstract
Recently, hybrid cloud computing paradigm has be widely advocated as a promising solution for Software-as-a-Service (SaaS) providers to effectively handle the dynamic user requests. With such a paradigm, the SaaS providers can extend their local services into the public clouds seamlessly so that the dynamic user request workload to a SaaS can be elegantly processed with both the local servers and the rented computing capacity in the public cloud. However, although it is suggested that a hybrid cloud may save cost compared with building a powerful private cloud, considerable renting cost and communication cost are still introduced in such a paradigm. How to optimize such operational cost becomes one major concern for the SaaS providers to adopt the hybrid cloud computing paradigm. However, this critical problem remains unanswered in the current state of the art. In this paper, we focus on optimizing the operational cost for the hybrid cloud paradigm by theoretically analyzing the problem with a Lyapunov optimization framework. This allows us to design an online dynamic provision algorithm. In this way, our approach can address the real-world challenges where no a priori information of public cloud renting prices is available and the future probability distribution of user requests is unknown. We then conduct extensive experimental study based on a set of real-world data, and the results confirm that our algorithm can work effectively in reducing the operational cost.
Yangfan Zhou 0002, Lei Jiao 0002, Xinya Yan, Xin Wang 0003, Michael R. Lyu
IEEE Trans. Serv. Comput.2
2014 Delay-Aware Cost Optimization for Dynamic Resource Provisioning in Hybrid Clouds
abstract
Hybrid cloud computing paradigm has recently be widely advocated, where Software-as-a-Service (SaaS) providers can extend their local services into the public clouds seamlessly. In this way, dynamic user request workload to a SaaS can be elegantly handled with the rented computing capacity in public cloud. However, although a hybrid cloud may save cost compared with the private cloud, it still introduces considerable renting cost and communication cost. How to optimize such an operational cost becomes one major concern for the SaaS providers to adopt such a hybrid cloud computing paradigm. However, this critical problem remains unanswered in the current state of the art. In this paper, we focus on optimizing the operational cost for the hybrid cloud model by theoretically analyzing the problem with a Lyapunov optimization framework, and accordingly providing an online dynamic provision algorithm. In this way, our approach can address the real-world challenges where no a priori information of public cloud renting prices is available and the future probability distribution of user requests is unknown. We then conduct experimental study based on a set of real-world data, and the results confirm that our algorithm can work well in reducing the cost.
Yangfan Zhou 0002, Lei Jiao 0002, Xinya Yan, Xin Wang 0003, Michael R. Lyu
ICWS2
2014 Towards Continuous and Passive Authentication via Touch Biometrics: An Experimental Study on Smartphones
Hui Xu 0009, Yangfan Zhou 0002, Michael R. Lyu
SOUPS2
2014 An Automatic Framework for Detecting and Characterizing Performance Degradation of Software Systems
abstract
Software systems that run continuously over a long time have been frequently reported encountering gradual degradation issues. That is, as time progresses, software tends to exhibit degraded performance, deflated service capacity, or deteriorated QoS. Currently, the state-of-the-art approach of Mann-Kendall Test & Seasonal Kendall Test & Sen's Slope Estimator & Seasonal Sen's Slope Estimator (MKSK) detects and characterizes degradation via a combination of techniques in statistical trend analysis. Nevertheless, we pinpoint some drawbacks of MKSK in this paper: 1) MKSK cannot be automated for large scale software degradation analysis, 2) MKSK estimates the degradation trend of software in an oversimplified linear way, 3) MKSK is sensitive to noise, and 4) MKSK suffers from high computational complexity. To overcome all these limitations, we propose a more advanced approach called Modified Cox-Stuart Test & Iterative Hodrick-Prescott Filter (CSHP). The superiority of our CSHP approach over MKSK is validated through extensive Monte Carlo simulations, as well as a real performance dataset measured from 99 real-world web servers.
Yong Qi 0001, Yangfan Zhou 0002, Pengfei Chen 0002, Jianfeng Zhan, Michael R. Lyu
IEEE Trans. Reliab.3
2013 An online service-oriented performance profiling tool for cloud computing systems
Haibo Mi, Huaimin Wang 0001, Yangfan Zhou 0002, Michael R. Lyu, Hua Cai, Gang Yin
Frontiers Comput. Sci.3
2013 MDiag: Mobility-assisted diagnosis for wireless sensor networks
Yangfan Zhou 0002, Michael R. Lyu, Evangeline F. Y. Young
J. Netw. Comput. Appl.2
2013 Toward Fine-Grained, Unsupervised, Scalable Performance Diagnosis for Production Cloud Computing Systems
abstract
Performance diagnosis is labor intensive in production cloud computing systems. Such systems typically face many real-world challenges, which the existing diagnosis techniques for such distributed systems cannot effectively solve. An efficient, unsupervised diagnosis tool for locating fine-grained performance anomalies is still lacking in production cloud computing systems. This paper proposes CloudDiag to bridge this gap. Combining a statistical technique and a fast matrix recovery algorithm, CloudDiag can efficiently pinpoint fine-grained causes of the performance problems, which does not require any domain-specific knowledge to the target system. CloudDiag has been applied in a practical production cloud computing systems to diagnose performance problems. We demonstrate the effectiveness of CloudDiag in three real-world case studies.
Haibo Mi, Huaimin Wang 0001, Yangfan Zhou 0002, Michael R. Lyu, Hua Cai
IEEE Trans. Parallel Distributed Syst.3
2012 P-Tracer: Path-Based Performance Profiling in Cloud Computing Systems
abstract
In large-scale cloud computing systems, the growing scale and complexity of component interactions pose great challenges for operators to understand the characteristics of system performance. Performance profiling has long been proved to be an effective approach to performance analysis; however, existing approaches do not consider two new requirements that emerge in cloud computing systems. First, the efficiency of the profiling becomes of critical concern; second, visual analytics should be utilized to make profiling results more readable. To address the above two issues, in this paper, we present P-Tracer, an online performance profiling approach specifically tailored for large-scale cloud computing systems. P-Tracer constructs a specific search engine that adopts a proactive way to process performance logs and generates particular indices for fast queries; furthermore, PTracer provides users with a suite of web-based interfaces to query statistical information of all kinds of services, which helps them quickly and intuitively understand system behavior. The approach has been successfully applied in Alibaba Cloud Computing Inc. to conduct online performance profiling both in production clusters and test clusters. Experience with one real-world case demonstrates that P-Tracer can effectively and efficiently help users conduct performance profiling and localize the primary causes of performance anomalies.
Haibo Mi, Huaimin Wang 0001, Hua Cai, Yangfan Zhou 0002, Michael R. Lyu, Zhenbang Chen 0001
COMPSAC4
2012 Online Protocol Verification in Wireless Sensor Networks via Non-intrusive Behavior Profiling
Yangfan Zhou 0002, Michael R. Lyu, Jiangchuan Liu
WASA1
2012 Localizing root causes of performance anomalies in cloud computing systems by analyzing request trace logs
Haibo Mi, Huaimin Wang 0001, Yangfan Zhou 0002, Michael R. Lyu, Hua Cai
Sci. China Inf. Sci.3
2011 A User Experience-Based Cloud Service Redeployment Mechanism
abstract
Cloud computing has attracted much interest recently from both industry and academic. Nowadays, more and more Internet applications are moving to the cloud environment. Making optimal deployment of cloud applications is critical for providing good performance to attract users. Optimizing user experience is usually required for cloud service deployment. However, it is a challenging task to know the user experience of end users, since there is generally no proactive connection between a user to the machine that will host the service instance. To attack this challenge, in this paper, we first propose a framework to model cloud features and capture user experience. Then based on the collected user connection information, we formulate the redeployment of service instances as k-median and max k-cover problems. We proposed several approximation algorithms to efficiently solve these problems. Comprehensive experiments are conducted by employing a real-world QoS dataset of service invocation. The experimental results show the effectiveness of our proposed redeployment approaches.
Yu Kang 0006, Yangfan Zhou 0002, Zibin Zheng, Michael R. Lyu
IEEE CLOUD2
2011 RealProct: Reliable Protocol Conformance Testing with Real Nodes for Wireless Sensor Networks
abstract
Despite the various applications of wireless sensor network (WSN), experiences from real WSN deployments show that protocol implementations in sensor nodes are susceptible to software failures, which may cause network failures or even breakdown. Pre-deployment protocol conformance testing is essential to ensure reliable communications for WSNs. Unfortunately, existing solutions with simulators cannot test the exact hardware and implementation environment as real sensors, whereas testbeds are expensive and limited to small scale networks and topologies. In this paper, we present RealProct, a novel and reliable framework for testing protocol implementations against their specifications in WSNs. RealProct utilizes real sensors for protocol conformance testing to ensure that the results are close to the real deployment. Using different techniques from those in simulations and real deployments, RealProct virtualizes a large network with any topology and generate non-deterministic events using only a small number of sensors to provide flexibility and to reduce the cost. The framework is carefully designed to support efficient testing in resource-limited sensors. Moreover, test execution and verdict are optimized to minimize the number of runs, while guaranteeing satisfactory false posi tive and false negative rates. We implement RealProct and test it with the IIP TCP/IP protocol stack and a routing protocol developed for WSNs in Contiki-2.4. The results demonstrate the effectiveness of RealProct by detecting several new bugs and all previously discovered bugs in various versions of the μIP TCP/IP protocol stack.
Edith C. H. Ngai, Yangfan Zhou 0002, Michael R. Lyu
TrustCom3
2010 Sentomist: Unveiling Transient Sensor Network Bugs via Symptom Mining
abstract
Wireless Sensor Network (WSN) applications are typically event-driven. While the source codes of these applications may look simple, they are executed with a complicated concurrency model, which frequently introduces software bugs, in particular, transient bugs. Such buggy logics may only be triggered by some occasionally interleaved events that bear implicit dependency, but can lead to fatal system failures. Unfortunately, these deeply-hidden bugs or even their symptoms can hardly be identified by state-of-the-art debugging tools, and manual identification from massive running traces can be prohibitively expensive. In this paper, we present Sentomist (Sensor application anatomist), a novel tool for identifying potential transient bugs in WSN applications. The Sentomist design is based on a key observation that transient bugs make the behaviors of a WSN system deviate from the normal, and thus outliers (i.e., abnormal behaviors) are good indicators of potential bugs. Sentomist introduces the notion of event-handling interval to systematically anatomize the long-term execution history of an event-driven WSN system into groups of intervals. It then applies a customized outlier detection algorithm to quickly identify and rank abnormal intervals. This dramatically reduces the human efforts of inspection (otherwise, we have to manually check tremendous data samples, typically with brute force inspection) and thus greatly speeds up debugging. We have implemented Sentomist based on the concurrency model of TinyOS. We apply Sentomist to test a series of representative real-life WSN applications that contain transient bugs. These bugs, though caused by complicated interactions that can hardly be predicted during the programming stage, are successfully confined by Sentomist.
Yangfan Zhou 0002, Michael R. Lyu, Jiangchuan Liu
ICDCS1
2010 A delay-aware reliable event reporting framework for wireless sensor-actuator networks
Edith C. H. Ngai, Yangfan Zhou 0002, Michael R. Lyu, Jiangchuan Liu
Ad Hoc Networks2
2009 Surviving Holes and Barriers in Geographic Data Reporting for Wireless Sensor Networks
abstract
Geographic forwarding is a favorable scheme for data reporting in wireless sensor networks (WSNs) due to its simplicity and low-overhead. However, WSNs are usually subject to complicated environmental factors. Network holes (i.e., the areas where no nodes inside) and barriers (i.e., those blocking the communication between two close nodes) are inevitable in practical deploying environments. These issues pose an obstacle to adopting geographic forwarding in WSNs, while current approaches lack an efficient method to tolerate such negative factors. In this paper we specifically tailor a waypoint-based geographic data reporting protocol (GDRP) for WSNs. Inherited from geographic forwarding, GDRP is light-weighted and hence well-suits WSNs. But unlike current approaches that often find suboptimal paths, GDRP adopts an intelligent strategy to select a best set of waypoints via which packets can efficiently circumvent holes and barriers, and it can thus find better paths. Extensive simulations are conducted to verify the advantages of GDRP in tolerating network holes and obstacles in WSNs.
Yangfan Zhou 0002, Michael R. Lyu, Jiangchuan Liu
MASS1
2009 Energy-efficient On-demand Active Contour Service for Sensor Networks
abstract
Contour mapping is an important technique for Wireless Sensor Networks (WSNs) in environmental monitoring to abstract the information of a monitored field. State-of-the-art approaches for contour mapping, however, are neither energy-optimized, nor capable of handling heterogeneous user requests. In this paper, we develop a novel energy-efficient On-demand Active Contour Service (OACS) for power-constrained WSNs. OACS regresses the field intensity function with kernel Support Vector Regression (SVR), a novel machine learning tool that flexibly handles both contour line and contour map requests. OACS also adaptively accommodates a wide range of contour line/map precision requirements: (1) For applications of low precision, only a minimum set of nodes are scheduled in working mode while others are sleeping for conserving energy. (2) For applications of high precision, through an active and progressive learning algorithm, OACS determines the best set of nodes that should be turned on for improving the contour line/map precision. Evaluation based on diverse realistic models demonstrates that OACS provides quality and seamless contour services for various application requirements yet significantly conserves energy.
Yangfan Zhou 0002, Michael R. Lyu, Jiangchuan Liu, Kam-Wing Ng
MASS1
2009 On Sensor Network Reconfiguration for Downtime-Free System Migration
Yangfan Zhou 0002, Michael R. Lyu, Jiangchuan Liu
Mob. Networks Appl.1
2008 An Index-Based Sensor-Grouping Mechanism for Efficient Field-Coverage Wireless Sensor Networks
abstract
This paper discusses a point-distribution index, l, which measures the normalized minimum distance between sensors. Maximizing l of a set of points causes the Delaunay triangulation graph of these points to be a net of equilateral triangles. Such a structure indicates the lowest redundancy of coverage if each point represents the center of a disc. Thus l can serve as a promising measure for solving a critical problem in field coverage: How to group a set of sensor nodes into disjoint subsets so that each subset can cover the entire field? Based on the l index, we develop an effective algorithm, MAXINE (MAXimizing-l Node-redundancy Exploiting), for the sensor- grouping problem. We evaluate the performance of MAXINE through extensive simulations and compare it with existing algorithms. The results demonstrate the effectiveness of MAXINE and verify the superiority of employing i for the sensor-grouping problem.
Yangfan Zhou 0002, Michael R. Lyu, Jiangchuan Liu
ICC1
2008 On sensor network reconfiguration for downtime-free system migrations
abstract
Many state-of-the-art wireless sensor networks have been equipped with reprogramming modules, e.g., those for software/firmware updates. System migration tasks such as software reprogramming however will interrupt normal sensing and data reporting operations of a sensor node. Although such tasks are
Yangfan Zhou 0002, Michael R. Lyu, Jiangchuan Liu
QSHINE1
2007 LOFT: A Latency-Oriented Fault Tolerant Transport Protocol for Wireless Sensor-Actuator Networks
abstract
Wireless sensor-actuator networks, or WSANs, refer to a group of sensors and actuators which collect data from the environment and perform application-specific actions in response. To act responsively and accurately, an efficient and reliable data transport protocol is crucial for the sensors to inform the actuators about the environmental events. Unfortunately, the low-power multi-hop communications in WSANs are inherently unreliable; the frequent sensor and link failures as well as the excessive delays due to congestion further aggravate the problem. In this paper, we propose a latency-oriented fault tolerant data transport protocol in WSANs. We argue that reliable data transport in such a real-time system should resist to the transmission failures, and should also consider the importance and freshness of the reported data. We articulate this argument and provide a cross-layer two-step data transport protocol for on- time and fault tolerant data delivery from sensors to actuators. Our protocol adopts smart priority scheduling that differentiates the event data of non-uniform importance. It balances the workload of sensors by checking their queue utilization and copes with node and link failures by an adaptive replication algorithm. We evaluate our protocol through extensive simulations, and the results demonstrate that it achieves the desirable reliability for WSANs.
Edith C. H. Ngai, Yangfan Zhou 0002, Michael R. Lyu, Jiangchuan Liu
GLOBECOM2
2007 POWER-SPEED: A Power-Controlled Real-Time Data Transport Protocol for Wireless Sensor-Actuator Networks
abstract
This paper investigates the data transport problem for reporting delay-sensitive events in wireless sensor-actuator networks (WSANs). We specifically tailor the protocol design according to the features of WSANs and propose POWER-SPEED, a real-time data transport protocol for WSANs to achieve energy-efficient data transport for delay-sensitive event reporting. In POWER-SPEED, sensor nodes select the next-hop neighbor to actuators according to the spatio-temporal historic data of the upstream QoS condition, which completely avoids control packets. With an adaptive transmitter power control scheme, POWER-SPEED conveys packets in an energy-efficient manner while maintaining soft real-time packet transport. It thus reduces the energy consumption of data transport while ensuring the QoS requirement in timeliness domain. We demonstrate the effectiveness of POWER-SPEED through simulations with NS2.
Yangfan Zhou 0002, Edith C. H. Ngai, Michael R. Lyu, Jiangchuan Liu
WCNC1
2006 A point-distribution index and its application to sensor-grouping in wireless sensor networks
abstract
Abstract — We propose ι, a novel index for evaluation of pointdistribution. ι is the minimum distance between each pair of points normalized by the average distance between each pair of points. We find that a set of points that achieve a maximum value of ι result in a honeycomb structure. We propose that ι can serve as a good index to evaluate the distribution of the points, which can be employed in coverage-related problems in wireless sensor networks (WSNs). To validate this idea, we formulate a general sensor-grouping problem for WSNs and provide a general sensing model. We show that locally maximizing ι at sensor nodes is a good approach to solve this problem with an algorithm called Maximizing-ι Node-Deduction (MIND). Simulation results verify that MIND outperforms a greedy algorithm that exploits sensorredundancy we design. This demonstrates a good application of employing ι in coverage-related problems for WSNs. I.
Yangfan Zhou 0002, Haixuan Yang, Michael R. Lyu, Edith C. H. Ngai
IWCMC1
2006 Reliable Reporting of Delay-Sensitive Events in Wireless Sensor-Actuator Networks
abstract
Wireless sensor-actuator networks, or WSANs, greatly enhance the existing wireless sensor network architecture by introducing powerful and even mobile actuators. The actuators work with the sensor nodes, but can perform much richer application-specific actions. To act responsively and accurately, an efficient and reliable reporting scheme is crucial for the sensors to inform the actuators about the environmental events. Unfortunately, the low-power multi-hop communications in a WSAN are inherently unreliable; the frequent sensor failures and the excessive delays due to congestion or in-network data aggregation further aggravate the problem. In this paper, we propose a general reliability-centric framework for event reporting in WSANs. We argue that the reliability in such a real-time system depends not only on the accuracy, but also the importance and freshness of the reported data. Our design follows this argument and seamlessly integrates three key modules that process the event data, namely, an efficient and fault-tolerant event data aggregation algorithm, a delay-aware data transmission protocol, and an adaptive actuator allocation algorithm for unevenly distributed events. Our transmission protocol also adopts smart priority scheduling that differentiates the event data of non-uniform importance. We evaluate our framework through extensive simulations, and the results demonstrate that it achieves desirable reliability with minimized delay.
Edith C. H. Ngai, Yangfan Zhou 0002, Michael R. Lyu, Jiangchuan Liu
MASS2
2005 PORT: A Price-Oriented Reliable Transport Protocol for Wireless Sensor Networks
abstract
In wireless sensor networks, to obtain reliability and minimize energy consumption, a dynamic rate-control and congestion-avoidance transport scheme is very important. We notice that reporting packets may contribute to the sink's fidelity of its knowledge on the phenomenon of interest to different extents. Thus, reliability cannot simply be measured by the sink's total incoming packet rate as considered in current schemes. Also, communication costs between sources and the sink may be different and may change dynamically. Based on these considerations, we propose PORT (price-oriented reliable transport protocol) to facilitate the sink to achieve reliability. Under the constraint that the sink must obtain enough fidelity for reliability purpose, PORT minimizes energy consumption with two schemes. One is based on the sink's application-based optimization approach that feeds back the optimal reporting rates. The other is a locally optimal routing scheme according to the feedback of downstream communication conditions. PORT can adapt well to the communication conditions for energy saving while maintaining the necessary level of reliability. Simulation results in an application case study demonstrate the effectiveness of PORT
Yangfan Zhou 0002, Michael R. Lyu, Jiangchuan Liu, Hui Wang 0011
ISSRE1
2005 On setting up energy-efficient paths with transmitter power control in wireless sensor networks
abstract
Energy-efficiency is an important design consideration of communication schemes for wireless sensor networks (WSNs). In this paper, we investigate the problem of energy-minimized sensor-to-sink communications with adaptive transmitter power settings. We devise a novel network-and application-aware model for this problem, and present a broadcast-on-update (BOU) solution. However, BOU suffers from the high overhead due to explosive broadcasting in path setup. We then show a waiting scheme, BOU-WA, that effectively mitigates the broadcast explosion. In BOU-WA, the waiting time before each broadcast is proportional to the probability that a node could find a more energy-efficient path to the sink. We provide an efficient approximation algorithm to calculate this probability. The performance of BOU-WA is evaluated under diverse network configurations, and the results demonstrate its superiority in conserving energy
Yangfan Zhou 0002, Michael R. Lyu, Jiangchuan Liu
MASS1