VLDB 2026 Research / reviewers in the wild / expert
Zhengwei Qi
dblp:17/1970
· DBLP profile ↗
103ranked-venue papers
5as first author
41since 2021 · last 2026
0000-0003-2730-2319ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 53 · 4 first-author · 21 since 2021Software engineering, systems software and programming languages · 28 · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 since 2021Artificial intelligence and machine learning · 6 · 4 since 2021Computer networks · 6 · 1 first-author · 2 since 2021Security and privacy · 4Databases, data management, data science and information retrieval · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FractalGPU: Fair and Elastic GPU Sharing for General-Purpose Computing
Kaicheng Guo, Lingyun Yang, Wenda Tang, Pengwei Du, Qian Da, Zhengwei Qi |
ICDCS | 9 |
| 2026 | SpiderSense: Lightweight Last-Level Cache Management via Time Period Tagging for LLC-Critical WorkloadsabstractMulti-tenant clouds enhance resource sharing among Virtual Machines (VMs) to boost overall utilization and reduce power consumption. However, this also introduces interference among workloads from different tenants and impedes VM performance isolation. In this article, we first demonstrate that the last-level cache (LLC) in CPUs, which is inherently shared by all VMs on the same physical machine, becomes a significant contending resource for LLC-critical workloads, leading to notable performance imbalances under the default hardware caching strategy. Although recent studies on LLC scheduling have progressed, they often require detailed profiling of user workloads or rely on hyperparameter tuning, limiting their applicability to private clusters or specific scenarios. We propose SpiderSense, a software-initiated LLC partitioner for managing Virtual Machine Monitors (VMM), to address these limitations. SpiderSense leverages modern yet off-the-shelf server CPU features to adaptively orchestrate LLC allocation among running black-boxed user VMs. SpiderSense dynamically samples VMs and calculates their fair share of LLC to allocate them while fully improving performance isolation among VMs. We experiment with SpiderSense using typical LLC-critical workloads, representative of the types of applications that stress LLC performance, such as Memcached and Llama. Our results show that SpiderSense improves performance by up to 40% in numerous colocation scenarios compared to current solutions. Zhixiang Wei, Zhibai Huang, James Yen, Tianlei Xiong, Kailiang Xu, Yucheng Zheng, Xingzi Yu, Yun Wang 0039, Zhengwei Qi |
ACM Trans. Archit. Code Optim. | 10 |
| 2026 | gPooling: An Elastic GPU Resource Management Framework for On-Demand Virtualization in Shared Accelerator Clusters
Kaicheng Guo, Chen Chen 0067, Yun Wang 0039, Pengwei Du, Zhengwei Qi, Haibing Guan |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2025 | Design and Operation of Elastic GPU-Pooling on Campus
Kaicheng Guo, Yun Wang 0039, Semakin Anton, Tovmachenko Dmitry, Jiajie Sheng, Jianwen Wei, James Lin 0001, Zhengwei Qi, Haibing Guan |
Euro-Par (1) | 9 |
| 2025 | DevTrace: Lightweight Plug-In Design for PCIe Transaction Tracing in Edge Intelligence WorkloadsabstractThe complexity of host-peripheral interactions during high-load tasks poses significant challenges for system optimization, with existing tracing tools degrading performance by up to 5.39×. We introduce DevTrace, a novel low-overhead tracing framework for peripheral interactions. Its modular architecture separates data collection from kernel-level operations, enabling lightweight tracing with minimal driver modifications across entire classes of devices. By eliminating heavy kernel tracing interrupts, DevTrace reduces overhead to negligible levels while maintaining data accuracy. In edge-based intelligence deployments, DevTrace achieves a 128× reduction in memory usage and approximately 10× lower CPU overhead compared to page-fault-based solutions. It significantly reduces data loss and performance degradation under high-load conditions, establishing it as a reliable tool for analyzing host-peripheral interactions and optimizing performance in resource-constrained environments. We also discuss potential extensions to eBPF to further decouple tracing from driver frameworks. Zhibai Huang, Kailiang Xu, Zhixiang Wei, Yinghao Deng, Chen Chen 0067, Yun Wang 0039, Fangxin Liu, Mingyuan Xia 0001, Zhengwei Qi |
ICCAD | 9 |
| 2025 | gFlow: Distributed Real-Time Reverse Remote Rendering System Model
Yixiao Xu, Wanzhao Xu, Yicheng Gu, Yun Wang 0039, Jiangyuan Ma, Zhengwei Qi |
MMM (2) | 7 |
| 2025 | ARMing x86 Games: Accelerating Binary Translation Using Software-Only Validated Flag Speculation
James Yen, Zhibai Huang, Zhixiang Wei, Chen Chen 0067, Senhao Yu, Yun Wang 0039, Hao Wang 0022, Zhengwei Qi |
MobiSys | 10 |
| 2025 | To PRI or Not To PRI, That's the question
Yun Wang 0039, Xianting Tian, Ben Luo, Zhixiang Wei, Zhibai Huang, Kailiang Xu, Kaihuan Peng, Kaijie Guo, Guangjian Wang, Shengdong Dai, Yibin Shen, Jiesheng Wu, Zhengwei Qi |
OSDI | 16 |
| 2025 | Effectively Virtual Page Prefetching via Spatial-Temporal Patterns for Memory-intensive Cloud ApplicationsabstractIn today's data-driven era, the explosive growth of global data volume has led to an increasing consumption of computing and storage resources. Effective management of virtual machines (VMs) memory usage is critical for cloud vendors to optimize system performance and resource utilization. Existing memory prefetching methods often slow down system performance, creating a difficult balance between maintaining service quality and optimizing resource use. For instance, Leap, which primarily utilizes address information, performs poorly in VM environments. The main issue is the performance drop caused by the reuse of memory resources in virtualized environments, a common situation in public clouds. Yun Wang 0039, Tianmai Deng, Ben Luo, Yibin Shen, Zhixiang Wei, Yixiao Xu, Minglang Huang, Zhengwei Qi |
PPoPP | 9 |
| 2025 | gCom: Fine-grained Compressors in Graphics Memory of Mobile GPUabstractToday, GPUs significantly boost rendering performance. However, the high memory requirements limit their use, especially on low-end mobile platforms. Compression techniques have been widely adopted to reduce memory consumption but face two primary issues when applied to mobile GPUs: (1) low repetition ratio caused by small raw data sizes and concurrency, and (2) low locality caused by unpredictable rendering behaviors. These two limitations result in a low compression ratio when compressors are applied to low-end mobile devices. This article introduces gCom , a fine-grained rendering compressor accelerated by GPUs. To improve the compression ratio, gCom incorporates the following innovations. First, unlike other compression techniques that use frames or tiles as basic processing units, gCom is the first to employ a fine-grained processing unit (i.e., the color channel), enhancing repetition amplification without increasing raw data. Second, gCom introduces two key features— Hierarchical Delta and Channel Decorrelator —which maximize the locality of adjacent channels and reduce raw data size. Third, to maintain the original GPU throughput, gCom revolutionizes the Golomb-Rice algorithm and proposes a new compression approach, the Parallel-Oriented Golomb-Rice algorithm, enabling parallel execution of both decompression and compression processes. The entire design of gCom utilizes only idle resources and existing commands on mobile GPUs, thus keeping purchasing costs low. To date, gCom has improved the channel locality by nearly 50%. The best compression achievement received by gCom has reached around 20%. Dongjie Tang, Yun Wang 0039, Yicheng Gu, Fangxin Liu, Zhengwei Qi |
ACM Trans. Archit. Code Optim. | 6 |
| 2025 | Exploring Efficient Hardware Accelerator for Learning-Based Image CompressionabstractRecently, learning-based image compression (LIC) methods have surpassed manually designed approaches in both compression quality and bitrate. However, increasing computational demands and insufficient optimizations in codec performance have hindered the advancement of LIC acceleration. Most researches focus on optimizing specific components, often neglecting the sources of underutilization during the execution of LIC models. Generally, efficient LIC acceleration encounters three primary challenges: 1) extra overheads introduced by individual optimizations; 2) load and computation imbalances in small kernels; and 3) mismatches between hardware configurations and the LIC models. To address these challenges, we propose a framework named extensive accelerator for LIC (X-LIC) for efficiently exploring the design space under constrained resources. First, we quantitatively characterize a representative LIC model, including its latency, computation size, and temporal utilization across various accelerators. We design a hardware-optimized quantization method to compensate for the lack of LIC-oriented research, particularly regarding data precision, distortion, and resource consumption. Additionally, we propose a parameterized LIC accelerator architecture that integrates seamlessly with existing loop optimization models and supports various LIC operators. Two optimization schemes are proposed for redundant computation in transposed convolution and load and computation imbalance in small kernels. Experimental results show that our framework demonstrates significant flexibility across a broad design space, achieving an average of 78%–95% of the theoretical peak performance and up to 688.2/759.1 GOP/s en/de-coder performance with INT8 precision. As a result, the en/de-coder performance can reach up to 33/36 FPS in 720P resolution. An FPGA demo of X-LIC is available athttps://github.com/sjtu-tcloud/X-LIC. Chen Chen 0067, Kaicheng Guo, Xingzi Yu, Weidong Qiu, Zhengwei Qi, Haibing Guan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2025 | Rethinking Virtual Machines Live Migration for Memory DisaggregationabstractResource underutilization has troubled data centers for several decades. On the CPU front, live migration plays a crucial role in reallocating CPU resources. Nevertheless, contemporary Virtual Machine (VM) live migration methods are burdened by substantial resource consumption. In terms of memory management, disaggregated memory offers an effective solution to enhance memory utilization, but leaves a gap in addressing CPU underutilization. Our findings highlight a considerable opportunity to optimize live migration in the context of disaggregated memory systems. We introduce Anemoi, a resource management system that seamlessly integrates VM live migration with memory disaggregation to address the aforementioned gap. In the context of disaggregated memory, remote memory becomes accessible from destination nodes, effectively eliminating the need for extensive network transmission of memory pages, and thereby significantly reducing migration time. In addition, we propose using memory replicas as an optimization to the live migration system. To mitigate the overhead of potential excessive memory consumption, we develop a dedicated compression algorithm. Our evaluations demonstrate that Anemoi leads to a notable 69% reduction in network bandwidth utilization and an impressive 83% reduction in migration time compared to traditional VM live migration. Additionally, our compression algorithm achieves an outstanding space-saving rate of 83.6%. Xingzi Yu, Xingguo Jia, Yun Wang 0039, Senhao Yu, Zhengwei Qi |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2025 | Sudoku: Scalable High-Density Cloud Rendering Multi-Client Architecture
Yun Wang 0039, Bing Deng, Xia Jiang, Xuyan Hu, Dongjie Tang, Randy Xu, Yijin Sun, Zhengwei Qi |
IEEE Trans. Serv. Comput. | 10 |
| 2024 | Foliage: Nourishing Evolving Software by Characterizing and Clustering Field BugsabstractModern programs, characterized by their complex functionalities, high integration, and rapid iteration cycles, are prone to errors. This complexity poses challenges in program analysis and software testing, making it difficult to achieve comprehensive bug coverage during the development phase. As a result, many bugs are only discovered during the software’s production phase. Tracking and understanding these field bugs is essential but challenging: the uploaded field error reports are extensive, and trivial yet high-frequency bugs can overshadow important low-frequency bugs. Additionally, application codebases evolve rapidly, causing a single bug to produce varied exceptions and stack traces across different code releases. In this paper, we introduce Foliage, a bug tracking and clustering toolchain designed to trace and characterize field bugs in JavaScript applications, aiding developers in locating and fixing these bugs. To address the challenges of efficiently tracking and analyzing the dynamic and complex nature of software bugs, Foliage proposes an error message enhancement technique. Foliage also introduces the verbal-characteristic-based clustering technique, along with three evaluation metrics for bug clustering: V-measure, cardinality bias, and hit rate. The results show that Foliage’s verbal-characteristic-based bug clustering outperforms previous bug clustering approaches by an average of 31.1% across these three metrics. We present an empirical study of Foliage applied to a complex real-world application over a two-year production period, capturing over 250,000 error reports and clustering them into 132 unique bugs. Finally, we open-source a bug dataset consisting of real and labeled error reports, which can be used to benchmark bug clustering techniques. Zhanyao Lei, Yixiong Chen, Mingyuan Xia 0001, Zhengwei Qi |
ISSTA | 4 |
| 2024 | gHermes: Application-Unaware Acceleration for Cloud Rendering and Computing with Efficient GPU UtilizationabstractGPUs, as crucial tools for intensive computing tasks, are primarily used for rendering and computation.However, due to the high cost of GPUs and their frequent updates, owning powerful local GPUs is considered a luxury.Offloading tasks to the cloud is an effective solution.This method allows users to leverage the powerful capabilities of cloud-based high-performance computing resources to perform complex rendering tasks more quickly and cost-effectively.However, there has not yet been a solution that simultaneously considers both rendering and computation without requiring any modifications to local applications.This paper introduces gHermes, an application-unaware acceleration solution that automatically generates code to intercept rendering and computing tasks, overcoming the limitations of existing technologies.Additionally, for computational tasks, gHermes also enables fine-grained control of GPU utilization.Experimental results show that gHermes can handle both computing and rendering tasks efficiently.While ensuring the Quality of Service (QoS) requirements for rendering tasks (30-60 FPS), it effectively utilizes GPU resources for computation services. Xuyan Hu, Yicheng Gu, Zhengwei Qi |
SEKE | 3 |
| 2024 | gVulkan: Scalable GPU Pooling for Pixel-Grained Rendering in Ray Tracing
Yicheng Gu, Yun Wang 0039, Yunfan Sun, Yuxin Xiang, Xuyan Hu, Zhengwei Qi, Haibing Guan |
USENIX ATC | 6 |
| 2024 | Enhancing embedded systems development with TS-
Xingzi Yu, Tianlei Xiong, Wengang Chen, Zhengwei Qi |
Autom. Softw. Eng. | 7 |
| 2024 | CARE: Cloudified Android With Optimized Rendering PlatformabstractDue to the excellent rendering capabilities, GPUs are mainstream accelerators in the Cloud-rendering industry. However, current Cloud-rendering systems suffer from a CPU-GPU workload imbalance that not only degrades application performance but also causes a significant waste of GPU resources. Recent proposals (such as API-forwarding and c-GPU) for improving CPU-GPU balance are promising but fail to solve system-resource redundancy issues (i.e., each instance tends to occupy all resources, exceeding its requirements). Such behavior will increase CPU load and lower effective GPU utilization. To demonstrate the severity of the issue, we evaluated real-world applications and results show that in most cases, nearly 50% of resources are useless. To solve this problem, we present CARE, the first framework intended to reduce the system-level redundancy by cloudifying the system from monolithic to Cloud-native. To allow users to configure required services, CARE puts forward a functional unit calledConfigurable Android (CA). To allow multiple instances to share certain types of resources, CARE innovatesSharing Resource (SR). To reduce the unused services, CARE introducesPruning Resources (PR). To further alleviate the CPU pressure and achieve CPU-GPU balance, we propose rShare, a system aiming at enhancing CPU effective utilization and increasing Android instance density of the Cloud-rendering platform. Based on Kubernetes, rShare divides all the CPUs into non-overlapping shared CPU pools, allocates instances to pools within milliseconds, and dynamically migrates them by tracking their QoS status. So far, CARE primarily focuses on Android systems and can handle 60 heavyweight instances (e.g., KOG (King of Glory)) on Intel SG1. rShare can apply instance allocation within milliseconds and increase the platform density by 39.4%. Yuxin Xiang, Dongjie Tang, Qiming Shi, Randy Xu, Mohammad R. Haghighat, Cathy Bao, Yicheng Gu, Zhengwei Qi, Haibing Guan |
IEEE Trans. Multim. | 11 |
| 2023 | Rethinking Virtual Machines Live Migration for Memory DisaggregationabstractResource underutilization has troubled data centers for several decades. Memory disaggregation provides an efficient way to improve memory utilization while leaving a missing puzzle piece on CPU underutilization. Live migration is an essential method for CPU resource reallocation. However, the state-of-the-art Virtual Machines (VM) live migration suffers from significant resource consumption. We discover the substantial potential for optimizing live migration in disaggregated memory systems. We propose Anemoi, a source management system incorporating VM live migration into memory disaggregation to fill in the missing piece. Disaggregated memory enables the read-only replica to be accessible from destination nodes, eliminating considerable network transmission for memory pages and saving migration time. Regarding the potential overwhelming memory consumption of duplicate read-only replicas, we design a dedicated compression algorithm. The evaluation shows that Anemoi reduces the network bandwidth use and the migration time by 69% and 83%, respectively, compared to VM live migration. The compression can achieve a space-saving of 83.6%. Xingguo Jia, Xingzi Yu, Yun Wang 0039, Senhao Yu, Zhengwei Qi |
CLUSTER | 5 |
| 2023 | Dissecting Scale-Out Applications Performance on Diverse TLB Designs (S)abstractScale-out applications, such as various big data systems and memory computing programs comprise an important software stack in clouds.Such applications usually have large memory data footprint as well as code sizes, thus stressing the CPU's TLB efficiency.In this paper, we experimentally evaluate how various TLB design choices in modern off-the-shelf x86 CPUs impact the performance of scale-out applications.The findings aim to guide the partitioning schemes and capacity planning of TLBs, and software-hardware co-design for emerging applications. Tianmai Deng, Yixiao Xu, Zhengwei Qi |
SEKE | 4 |
| 2023 | HaeNAS: Hardware-Aware Efficient Neural Architecture Search via Zero-Cost ProxyabstractThe practical use of advanced DNN models is hindered by limited hardware resources and high computation demands.Neural Architecture Search (NAS) is becoming a default technique which automatically discovers architectures that are competitive with handcraft ones.However, existing methods only prioritize accuracy and overlook hardware-related factors.To address this, we introduce HaeNAS (Hardwareaware efficient NAS), which considers both accuracy and computation cost on specific hardware platforms.The search space of HaeNAS consists of several stages, each allowing different convolution kernels, layer numbers, layer widths, and operation types.We use a data-driven approach to predict the latency and energy consumption on target hardware, and we improve the zero-cost proxy based on network pruning research to speed up the NAS process.With these techniques, HaeNAS finds a target network within 160 GPU hours, which achieves 80.7% top-1 accuracy on the ImageNet, with a latency of 10.4ms and energy consumption of 931mJ. Yaodanjun Ren, Chen Chen 0067, Zhengwei Qi |
SEKE | 3 |
| 2023 | Bindox: An Efficient and Secure Cross-System IPC Mechanism for Multi-Platform ContainersabstractContainerization is widely used for isolation in various applications because it is lightweight, scalable, and portable.In modern distributed systems, seamless inter-process communication (IPC) between multi-platform containers is essential for a range of applications and services, including microservices, cloud computing, and Internet of Things (IoT) devices.However, secure and efficient communication between containers on the same host is challenging, especially when different operating systems are involved.This paper introduces Bindox, a lightweight, efficient, and secure IPC mechanism that enables seamless communication across multiple platforms, including Android and Linux.Bindox uses shared memory for data transfer and implements a stable client-server architecture, ensuring high performance and ease of maintenance.Additionally, Bindox provides a robust security mechanism that guarantees confidentiality, integrity, and availability of the communication channel.Experimental results demonstrate that Bindox outperforms existing networking and IPC methods in terms of memory use, latency, and CPU usage, making it a promising solution for efficient and secure communication between multi-platform containers. Yuxin Xiang, Bing Deng, Randy Xu, Marc Mao, Yun Wang 0039, Zhengwei Qi |
SEKE | 7 |
| 2023 | Optimum: Runtime optimization for multiple mixed model deployment deep learning inference
Kaicheng Guo, Yixiao Xu, Zhengwei Qi, Haibing Guan |
J. Syst. Archit. | 3 |
| 2023 | rShare: Alleviating long startup on the Cloud-rendering platform through de-systemization
Dongjie Tang, Marc Mao, Cathy Bao, Qiming Shi, Randy Xu, Mohammad R. Haghighat, Yun Wang 0039, Zhengwei Qi, Haibing Guan, Xiaojie Cao |
J. Syst. Archit. | 10 |
| 2023 | DVHN: A Deep Hashing Framework for Large-Scale Vehicle Re-IdentificationabstractVehicle re-identification is a pervasive technology in real-world intelligence transportation systems. Conventional methods generally perform re-identification tasks by representing vehicle images as real-valued feature vectors and then ranking the gallery images by computing the corresponding Euclidean distances. Despite achieving remarkable retrieval accuracy, these high-dimensional real-valued feature vectors are not tailored for fast indexing and matching and require tremendous memory and computation when the gallery set is large, making them inapplicable in a large-scale real-world retrieval setting. In light of this limitation, in this paper, we make the very first attempt to develop an efficient vehicle re-identification system (DVHN) for real-world large-scale retrieval tasks with deep hashing learning. It could substantially reduce memory usage and enhances retrieval efficiency while maintaining retrieval accuracy. Concretely, DVHN directly learns discrete compact binary hashing codes for each image by jointly optimizing the feature learning network and the hash code generating module. Specifically, we directly constrain the output from the convolutional neural network to be discrete binary codes and ensure the learned binary codes are optimal for classification. To optimize the deep discrete hashing framework, we further propose an alternating minimization method for learning binary similarity-preserved hashing codes. Extensive experiments on two widely-studied vehicle re-identification datasets- VehicleID and VeRi- have demonstrated the superiority of our method against the state-of-the-art deep hash methods. DVHN of 2048 bits can achieve 13.94% and 10.21% accuracy improvement in terms of mAP and Rank@1 for VehicleID (800) dataset. For VeRi, we achieve 35.45% and 32.72% performance gains for Rank@1 and mAP, respectively. Yongbiao Chen, Fangxin Liu, Kaicheng Guo, Zhengwei Qi |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2023 | Energy- and Quality of Experience-Aware Dynamic Resource Allocation for Massively Multiplayer Online Games in Heterogeneous Cloud Computing SystemsabstractMassively multiplayer online games (MMOGs) are a new type of large-scale interactive applications providing a seamless virtual world for millions of gamers from all over the world. Traditionally, MMOGs are implemented using a client-server architecture. With the maturity of cloud computing, many MMOG providers such as Blizzard Entertainment and Jagex Limited have begun to utilize virtualized machines to serve their consumers due to the elasticity, adaptability, and cost advantages of clouds. It is a major challenge for MMOG providers how to dynamically allocate resources in an on-demand pattern in order to reduce the energy cost associated with operating a MMOG cloud while providing good-enough quality of experience (QoE) for MMOG gamers. In this paper, we propose a dynamic resource allocation scheme for MMOGs in heterogeneous cloud computing systems, which makes use of both dynamic virtual machine (VM) consolidation among physical machines (PMs) and long short-term memory based VM resizing at physical machine level to achieve the energy efficiency and desired QoE requirements. Our scheme considers multiple types of resources, characteristic of AFK (away from keyboard) gamers, heterogeneous PMs and VMs, strict QoE requirements and overheads incurred due to migrating VMs. Furthermore, a novel hybrid algorithm based on differential evolution and modified first-fit heuristic is presented for dynamic consolidation of VMs in heterogeneous cloud data centers. The experiment results show that, compared to the traditional over-provisioning policy, our resource allocation scheme can achieve up to 44.8% energy savings while ensuring the QoE requirements of MMOG gamers under rapidly changing workloads. Yongqiang Gao, Zhulong Xie, Zhengwei Qi, Jiantao Zhou 0002 |
IEEE Trans. Serv. Comput. | 4 |
| 2023 | Bootstrapping Automated Testing for RESTful Web ServicesabstractModern RESTful services expose RESTful APIs to integrate with diversified applications. Most RESTful API parameters are weakly typed, which greatly increases the possible input value space. Weakly-typed parameters pose difficulties for automated testing tools to generate effective test cases to reveal web service defects related to parameter validation. We call this phenomenon the type collapse problem. To remedy this problem, we introduce FET (Format-encoded Type) techniques, including the FET, the FET lattice, and the FET inference to model fine-grained information for API parameters. Inferred FET can enhance parameter validation, such as generating a parameter validator for a certain RESTful server. Enhanced by FET techniques, automated testing tools can generate targeted test cases. We demonstrate Leif, a trace-driven fuzzing tool, as a proof-of-concept implementation of FET techniques. Experiment results on 27 commercial services show that FET inference precisely captures documented parameter definitions, which helps Leif discover 11 new bugs and reduce$72\% - 86\%$fuzzing time compared to state-of-the-art fuzzers. Leveraged by the inter-parameter dependency inference, Leif saves$15\%$fuzzing time. Zhanyao Lei, Yixiong Chen, Mingyuan Xia 0001, Zhengwei Qi |
IEEE Trans. Software Eng. | 5 |
| 2022 | Dissecting the Workload of Cloud Storage SystemabstractThe innovation and evolution of file and storage systems have been influenced by workload analysis. Though cloud storage systems have been widely deployed and used, real-world and large-scale cloud storage workload studies are rare. Previous large-scale distributed storage systems can meet versatility, stability, and reliability requirements. Furthermore, modern cloud storage systems need to meet additional challenges, such as coping with surges in peak loads and rapid expansion of requests. These changes may lead to different characteristics.In this work, we propose DiTing data tracing system and collect workloads with over 242,000 billion requests from the Alibaba cloud. By comparing the normal days and the Single’s Day (the world’s largest online shopping festival), we analyze characteristics such as I/O scale, latency, locality, and load distribution. Our analysis reveals four key observations as follows. First, the virtual layer is the performance bottleneck of modern cloud storage systems during extreme peak periods. Second, the write operations dominate the data access because the application and operating system buffers absorb reads better than writes. Third, the workload is heavily skewed toward a small percentage of virtual cloud disks, with 20% of cloud disks accounting for 80% of I/O requests. Finally, data access shows poor temporal and spatial locality, and the I/O requests are mostly small-scaled. Based on these observations, we propose several suggestions for cloud storage systems, including separating I/O processing from the virtual layer to the proxy layer, deploying heavy and light workload applications on the same node, and adopting a write-friendly cloud disk design for write-skewed requests, etc. In summary, these workload characteristics and suggestions are useful for designing and implementing next-generation cloud storage systems. Yaodanjun Ren, Xiaoyi Sun, Kaishi Li, Jiale Lin, Shuzhi Feng, Zhenyu Ren, Jian Yin 0023, Zhengwei Qi |
ICDCS | 8 |
| 2022 | Falcon: A Timestamp-based Protocol to Maximize the Cache Efficiency in the Distributed Shared MemoryabstractDistributed shared memory (DSM) systems can handle data-intensive applications and recently receiving more attention. A majority of existing DSM implementations are based on write-invalidation (WI) protocols, which achieve sub-optimal performance when the cache size is small. Specifically, the vast majority of invalidation messages become useless when evictions are frequent. The problem is troublesome regarding scarce memory resources in data centers. To this end, we propose a self-invalidation protocol Falcon to eliminate invalidation messages. It relies on per-operation timestamps to achieve the global memory order required by sequential consistency (SC). Furthermore, we conduct a comprehensive discussion on the two protocols with an emphasis on the cache size impact. We also implement both protocols atop a recent DSM system, Grappa. The evaluation shows that the optimal protocol can improve the performance of a KV database by 27% and a graph processing application by 71.4% against the vanilla cache-free scheme. Xiangyao Yu, Zhengwei Qi, Haibing Guan |
IPDPS | 3 |
| 2022 | Supervised Contrastive Vehicle Quantization for Efficient Vehicle RetrievalabstractThis paper considers large-scale efficient vehicle re-identification (Vehicle ReID). Existing works adopting deep hashing techniques function by projecting vehicle images into compact binary codes in the Hamming space. Since Hamming distance is less distinct, a considerable amount of discriminative information will be lost, leading to degraded retrieval performances. Inspired by the recent advancements in contrastive learning, we put forward the very first product quantization based framework for large-scale efficient vehicle re-identification: Supervised Contrastive Vehicle Quantization (SCVQ). Specifically, we integrate the product quantization process into deep supervised learning by designing a differentiable quantization network. In addition, we propose a novel supervised cross-quantized contrastive quantization (SCQC) loss for similarity-preserving learning, which is tailored for the asymmetric retrieval in the product quantization process. Comprehensive experiments on two public benchmarks have evidenced the superiority of our framework against the state-of-the-arts. Our work is open-sourced at https://github.com/chrisbyd/ContrastiveVehicleQuant Yongbiao Chen, Kaicheng Guo, Fangxin Liu, Zhengwei Qi |
ICMR | 5 |
| 2022 | TransHash: Transformer-based Hamming Hashing for Efficient Image RetrievalabstractDeep hashing has gained growing popularity in approximate nearest neighbor search for large-scale image retrieval. Until now, the deep hashing for the image retrieval community has been dominated by convolutional neural network architectures, e.g. Resnet [22]. In this paper, inspired by the recent advancements of vision transformers, we present Transhash, a pure transformer-based framework for deep hashing learning. Concretely, our framework is composed of two major modules: (1) Based onVision Transformer (ViT), we design a siamese Multi-Granular Vision Tansformer backbone (MGVT) for image feature extraction. To learn fine-grained features, we innovate a dual-stream multi-granular feature learning on top of the transformer to learn discriminative global and local features. (2) Besides, we adopt a Bayesian learning scheme with a dynamically constructed similarity matrix to learn compact binary hash codes. The entire framework is jointly trained in an end-to-end manner. To the best of our knowledge, this is the first work to tackle deep hashing learning problems without convolutional neural networks (CNNs). We perform comprehensive experiments on three widely-studied datasets: CIFAR-10, NUSWIDE and IMAGENET. The experiments have evidenced our superiority against the existing state-of-the-art deep hashing methods. Specifically, we achieve 8.2%, 2.6%, 12.7% performance gains in terms of average mAP for different hash bit lengths on three public datasets, respectively. Yongbiao Chen, Fangxin Liu, Zhigang Chang, Mang Ye, Zhengwei Qi |
ICMR | 6 |
| 2022 | AppSPIN: reconfiguration-based responsiveness testing and diagnosing for Android Apps
Zhanyao Lei, Wenhua Zhao, Zhenkai Ding, Mingyuan Xia 0001, Zhengwei Qi |
Autom. Softw. Eng. | 5 |
| 2022 | Cost-Efficient and Quality-of-Experience-Aware Player Request Scheduling and Rendering Server Allocation for Edge-Computing-Assisted Multiplayer Cloud GamingabstractPrompted by the remarkable progress in both cloud computing and GPU virtualization, cloud gaming has been attracting more and more attention in the gaming industry. With the cloud gaming model, players do not need to download or install the game on local devices, and constantly upgrade their devices. Despite these advantages, cloud gaming faces several challenges for its success, including long response delay, poor game fairness, and high operational cost. To this end, this article proposes an edge computing-assisted multiplayer cloud gaming system named ECACG to improve multiplayer cloud gaming experiences and operating costs by offloading the game rendering task to the nearby edge server. Based on the ECACG, two decision processes are completed. One is player request scheduling and the other is rendering server allocation. The decision problem is formulated into a constrained multiobjective optimization model. A novel hybrid algorithm based on deep reinforcement learning and heuristic strategy is developed to solve the optimization problem. The effectiveness of the proposed ECACG is evaluated by simulation experiments based on the real-world parameters. The simulation results show that compared with the existing schemes, the proposed ECACG can achieve lower rental costs and better fairness, while providing the good-enough response delay for players. Yongqiang Gao, Chaoyu Zhang, Zhulong Xie, Zhengwei Qi, Jiantao Zhou 0002 |
IEEE Internet Things J. | 4 |
| 2022 | DSPR: Secure decentralized storage with proof-of-replication for edge devices
Yongbiao Chen, Zhengwei Qi, Haibing Guan |
J. Syst. Archit. | 3 |
| 2022 | GiantVM: A Novel Distributed Hypervisor for Resource Aggregation with DSM-aware OptimizationsabstractWe present GiantVM, 1 an open-source distributed hypervisor that provides the many-to-one virtualization to aggregate resources from multiple physical machines. We propose techniques to enable distributed CPU and I/O virtualization and distributed shared memory (DSM) to achieve memory aggregation. GiantVM is implemented based on the state-of-the-art type-II hypervisor QEMU-KVM, and it can currently host conventional OSes such as Linux. (1) We identify the performance bottleneck of GiantVM to be DSM, through a top-down performance analysis. Although GiantVM offers great opportunities for CPU-intensive applications to enjoy the aggregated CPU resources, memory-intensive applications could suffer from cross-node page sharing, which requires frequent DSM involvement and leads to performance collapse. We design the guest-level thread scheduler, DaS (DSM-aware Scheduler), to overcome the bottleneck. When benchmarking with NAS Parallel Benchmarks, the DaS could achieve a performance boost of up to 3.5×, compared to the default Linux kernel scheduler. (2) While evaluating DaS, we observe the advantage of GiantVM as a resource reallocation facility. Thanks to the SSI abstraction of GiantVM, migration could be done by guest-level scheduling. DSM allows standby pages in the migration destination, which need not be transferred through the network. The saved network bandwidth is 68% on average, compared to VM live migration. Resource reallocation with GiantVM increases the overall CPU utilization by 14.3% in a co-location experiment. Xingguo Jia, Boshi Yu, Xingyue Qian, Zhengwei Qi, Haibing Guan |
ACM Trans. Archit. Code Optim. | 5 |
| 2021 | MEGATRON: Software-Managed Device TLB for Shared-Memory FPGA VirtualizationabstractFPGAs are being virtualized to improve resource utilization in data centers. Memory access performance is essential to FPGA hypervisors for shared-memory FPGA platform, where accelerators access memory spontaneously. DMA remapping with IOMMU provides a handy solution; however, fixed IOMMU can not benefit from the reconfigurability of FPGAs. In this work, we propose MEGATRON, a hybrid address translation service consisting of a hardware TLB and a software page table walker. By integrating MEGATRON into an existing FPGA hypervisor, we conduct a comprehensive analysis of link performance of a multi-link CPU-FPGA platform, and demonstrate the competitiveness of the customizable translation service. Yanqiang Liu, Jiacheng Ma 0001, Zhengjun Zhang, Linsheng Li, Zhengwei Qi, Haibing Guan |
DAC | 5 |
| 2021 | Bootstrapping Automated Testing for RESTful Web ServicesabstractAbstract Modern RESTful services expose RESTful APIs to integrate with diversified applications. Most RESTful API parameters are weakly typed, which greatly increases the possible input value space. This poses difficulties for automated testing tools to generate effective test cases to reveal web service defects related to parameter validation. We call this phenomenon the type collapse problem. To remedy this problem, we introduce FET (Format-encoded Type) techniques, including the FET, the FET lattice, and the FET inference to model fine-grained information for API parameters. Enhanced by FET techniques, automated testing tools can generate targeted test cases. We demonstrate Leif, a trace-driven fuzzing tool, as a proof-of-concept implementation of FET techniques. Experiment results on 27 commercial services show that FET inference precisely captures documented parameter definitions, which helps Leif to discover 11 new bugs and reduce $$72\% \sim 86\%$$ 72 % ∼ 86 % fuzzing time as compared to state-of-the-art fuzzers. Yixiong Chen, Zhanyao Lei, Mingyuan Xia 0001, Zhengwei Qi |
FASE | 5 |
| 2021 | Performance Analysis of Open-Source Hypervisors for Automotive SystemsabstractNowadays, automotive products are intelligence intensive and thus inevitably handle multiple functionalities under the current high-speed networking environment. The embedded virtualization has high potentials in the automotive industry, thanks to its advantages in function integration, resource utilization, and security. The invention of ARM virtualization extensions has made it possible to run open-source hypervisors, such as Xen and KVM, for embedded applications. Nevertheless, there is little work to investigate the performance of these hypervisors on automotive platforms. This paper presents a detailed analysis of different types of open-source hypervisors that can be applied in the ARM platform. We carry out the virtualization performance experiment from the perspectives of CPU, memory, file I/O, and some OS operation performance on Xen and Jailhouse. A series of microbenchmark programs have been designed, specifically to evaluate the real-time performance of various hypervisors and the relevant overhead. Compared with Xen, Jailhouse has better latency performance, stable latency, and little interference jitter. The performance experiment results help us summarize the advantages and disadvantages of these hypervisors in automotive applications. Zhengjun Zhang, Yanqiang Liu, Jiangtao Chen, Zhengwei Qi, Huai Liu |
ICPADS | 4 |
| 2021 | CARE: Cloudified Android OSes on the Cloud RenderingabstractGPUs have become ubiquitous in the Cloud-rendering areas due to the outstanding rendering performance. However, many existing Cloud-rendering systems suffer from low GPU utilization caused by the CPU bottleneck. Recent proposals (e.g., API-forwarding and c-GPU) for GPU-usage optimization are promising but fail to address the system-resource redundancy issues (i.e., each instance tends to occupy all the system resources exceeding their requirements), leading to unnecessary CPU consumption and lowering GPU utilization. We conducted an experiment by testing real-world applications on the percentage of unused resources to demonstrate the severity of this issue. Nearly 50% of resources are unused. Dongjie Tang, Cathy Bao, Qiming Shi, Marc Mao, Randy Xu, Linsheng Li, Mohammad R. Haghighat, Zhengwei Qi, Haibing Guan |
ACM Multimedia | 10 |
| 2021 | Efficient shuffle management for DAG computing frameworks based on the FRQ model
Chunghsuan Wu, Zhouwang Fu, Tao Song 0003, Yanqiang Liu, Zhengwei Qi, Haibing Guan |
J. Parallel Distributed Comput. | 6 |
| 2021 | gRemote: Cloud rendering on GPU resource pool based on API-forwarding
Dongjie Tang, Linsheng Li, Jiacheng Ma 0001, Xue (Steve) Liu, Zhengwei Qi, Haibing Guan |
J. Syst. Archit. | 5 |
| 2020 | A Hypervisor for Shared-Memory FPGA PlatformsabstractCloud providers widely deploy FPGAs as application-specific accelerators for customer use. These providers seek to multiplex their FPGAs among customers via virtualization, thereby reducing running costs. Unfortunately, most virtualization support is confined to FPGAs that expose a restrictive, host-centric programming model in which accelerators cannot issue direct memory accesses (DMAs). The host-centric model incurs high runtime overhead for workloads that exhibit pointer chasing. Thus, FPGAs are beginning to support a shared-memory programming model in which accelerators can issue DMAs. However, virtualization support for shared-memory FPGAs is limited. This paper presents Optimus, the first hypervisor that supports scalable shared-memory FPGA virtualization. Optimus offers both spatial multiplexing and temporal multiplexing to provide efficient and flexible sharing of each accelerator on an FPGA. To share the FPGA-CPU interconnect at a high clock frequency, Optimus implements a multiplexer tree. To isolate each guest's address space, Optimus introduces the technique of page table slicing as a hardware-software co-design. To support preemptive temporal multiplexing, Optimus provides an accelerator preemption interface. We show that Optimus supports eight physical accelerators on a single FPGA and improves the aggregate throughput of twelve real-world benchmarks by 1.98x-7x. Jiacheng Ma 0001, Gefei Zuo, Kevin Loughlin, Xiaohe Cheng, Yanqiang Liu, Abel Mulugeta Eneyew, Zhengwei Qi, Baris Kasikci |
ASPLOS | 7 |
| 2020 | gRemote: API-Forwarding Powered Cloud RenderingabstractTraditional GPU resource allocation approaches, widely adopted in today's data centers, only focus on the server-side functions while ignoring the client-side. These approaches waste client-side hardware resources. To solve this problem, remote API-forwarding architectures appear. Through running applications on the client-side, remote API-forwarding architectures offload some workloads to the client. However, many remote API-forwarding systems suffer from one big issue: shared-resource interference, stemming from two reasons: (a) GPU resource racing caused by resource overuse for a single client, and (b) CPU resource racing caused by resource shortage among clients. This paper presents gRemote, an open-source GPU-remoting system that can address this issue. To mitigate the CPU resource shortage, gRemote improves CPU configurations by expanding CPU resources from the server-side to both server- and client-side. To maintain the reasonable GPU usage for individual tasks, we innovate a new resource-sharing mechanism called GPU throttle. gRemote supports 1,228 OpenGL commands with around 10% shared-resource interference. Dongjie Tang, Yun Wang 0039, Linsheng Li, Jiacheng Ma 0001, Xue (Steve) Liu, Zhengwei Qi, Haibing Guan |
HPDC | 6 |
| 2020 | OPS: Optimized Shuffle Management System for Apache SparkabstractIn recent years, distributed computing frameworks, such as Hadoop MapReduce and Spark, are widely used for big data processing. With the explosive growth of the amount of data, companies tend to store intermediate data of the shuffle phase on disk instead of memory. Therefore, intensive network and disk I/O are both involved in the shuffle phase. To optimize the overhead of the shuffle phase, we propose OPS, an open-source distributed computing shuffle management system based on Spark, which provides an independent shuffle service for Spark. By using early-merge and early-shuffle strategy, OPS alleviates the I/O overhead in the shuffle phase and efficiently schedules the I/O and computing resources. OPS also proposes a slot-based scheduling algorithm to predict and calculate the optimal scheduling result of the reduce task. Besides, OPS provides a taint-redo strategy to ensure the fault tolerance of computing jobs. We evaluated the performance of OPS on a 100-node Amazon AWS EC2 cluster. Overall, OPS optimizes the overhead of shuffle by nearly 50%. In the test cases of HiBench, OPS improves end-to-end completion time by nearly 30% on average. Yuchen Cheng, Chunghsuan Wu, Yanqiang Liu, Hong Xu 0001, Zhengwei Qi |
ICPP | 7 |
| 2020 | MAENet: Boosting Feature Representation for Cross-Modal Person Re-Identification with Pairwise SupervisionabstractPerson re-identification aims at successfully retrieving the images of a specific person in the gallery dataset given a probe image. Among all the existing research areas related to person re-identification, visible to thermal person re-identification (VT-REID) has gained proliferating momentum. VT-REID is deemed to be a rather challenging task owing to the large cross-modality gap [25], cross-modality variation and intra-modality variation. Existing techniques generally tackle this problem by embedding cross-modality data with convolutional neural networks into shared feature space to bridge the cross-modality discrepancy, and subsequently, devise hinge losses on similarity learning to alleviate the variation. However, feature extraction methods based simply on convolutional neural networks may fail to capture the distinctive and modality-invariant features, resulting in noises for further re-identification techniques. In this work, we present a novel modality and appearance invariant embedding learning framework equipped with maximum likelihood learning to perform cross-modal person re-identification. Extensive and comprehensive experiments are conducted to test the effectiveness of our framework. Results demonstrated that the proposed framework yields state-of-the-art Re-ID accuracy on RegDB and SYSU-MM01 datasets. Yongbiao Chen, Zhengwei Qi |
ICMR | 3 |
| 2020 | DroidCloud: Scalable High Density AndroidTM Cloud RenderingabstractCloud rendering is an emerging technology in which rendering-heavy applications run on the cloud server and then stream the rendered contents to the end-user device. High density and high scalability of the cloud rendering services are crucial to support millions of users concurrently and cost-effectively. However, it is still challenging to run Android OS in cloud smoothly with high density and high scalability without compromising user experience. This paper presents DroidCloud, the first open-source Android\footnoteAndroid is a trademark of Google LLC. cloud rendering solution focusing on the scalable design and density aspect optimization to the best of our knowledge. To cloudify Android OS, DroidCloud utilizes thevHAL technology in order to support remote devices and keep transparent to Android applications. And aFlexible rendering scheduling policy is introduced to break the boundary of GPU physical locations. Thus, both remote GPUs and local GPUs can accommodate render tasks by forwarding rendering tasks and making it possible to support multiple Android OSes with GPU acceleration. Besides, to further improve the density, DroidCloud optimizes the resource cost both in a single instance and across instances. We show that DroidCloud can run hundreds of Android OSes on a single Intel Xeon server with GPU acceleration simultaneously, increasing the density at the scale of one order of magnitude compared to current cloud gaming systems. Further experimental results demonstrate that DroidCloud can transparently run Android applications at native speed with lower CPU, memory, and storage utilization. Linsheng Li, Cathy Bao, Randy Xu, Mohammad R. Haghighat, Jerry W. Hu, Shoumeng Yan, Zhengwei Qi |
ACM Multimedia | 10 |
| 2020 | GiantVM: a type-II hypervisor implementing many-to-one virtualizationabstractIn recent years, since scale-up machines are not economical and may not be affordable for small businesses, scale-out has become the standard answer to data analysis, machine learning, and many other fields. However, these frameworks introduce complex programming models that put a burden on developers. Therefore, Single System Image (SSI), which means a cluster of machines that appears to be one single system, has been proposed to hide the complexity of distributed systems. Unfortunately, due to the mature ecosystem of current mainstream Operating Systems (OSes), it might be non-trivial and even unaffordable to modify the current OS to implement SSI. With the wide use of virtualization, we believe that it is appealing to support SSI at the hypervisor, without modifying guest OSes. Zhuocheng Ding, Yubin Chen, Xingguo Jia, Boshi Yu, Zhengwei Qi, Haibing Guan |
VEE | 6 |
| 2020 | TimelyRep: Timing deterministic replay for Android web applicationsabstractSummary With the constantly growing and changing requirements of app users, web techniques are used in mobile application development for better cross‐platform compatibility and online update. As the embedded web contents gain complexity, debugging web apps become a critical demand. Web replay tools can record program inputs and reproduce the same execution for debugging and performance tuning. However, traditional replay approaches are largely intended for apps with desktop interaction methods (keyboard, mouse) and require modification to the browser, which limits their applicability in mobile platforms. In this paper, we develop TimelyRep, which provides deterministic record‐and‐replay as a software library, running on commodity Android. TimelyRep can be used for app development with unmodified Android devices and for production to collect faulty execution from users. Also, we propose an efficient replay timing control mechanism and achieve higher timing precision as facing higher event rate on touchscreen devices. TimelyRep also supports cross‐device replay and can replay logged event traces on different devices, which is useful for developers to reproduce user inputs on their own devices. We evaluate TimelyRep with real‐world web applications. The results show that TimelyRep is useful for recreating program bugs and maintaining low delays for touch‐intensive web games. Yanqiang Liu, Fangge Yan, Mingyuan Xia 0001, Zhengwei Qi, Xue (Steve) Liu |
Softw. Test. Verification Reliab. | 4 |
| 2020 | gMig: Efficient vGPU Live Migration with Overlapped Software-Based Dirty Page VerificationabstractThis paper introduces gMig, an open-source and practical vGPU live migration solution for full virtualization. Taking the advantage of the dirty pattern of GPU workloads, gMig presents the One-Shot Pre-Copy mechanism combined with the hashing based Software Dirty Page technique to achieve efficient vGPU live migration. Particularly, we propose three core techniques for gMig: 1) Dynamic Graphics Address Remapping, which parses and manipulates GPU commands to adjust the address mapping and adapt to a different environment after migration, 2) Software Dirty Page, which utilizes a hashing based approach with sampling pre-filtering to detect page modification, overcomes the commodity GPU's hardware limitation, and speeds up the migration by only sending the dirtied pages, 3) Overlapped Migration Process, which significantly compresses the hanging overhead by overlapping the dirty page verification and transmission concurrently. Our evaluation shows that gMig achieves GPU live migration with an average downtime of 302 ms on Windows and 119 ms on Linux. With the help of Software Dirty Page, the number of GPU pages transferred during the downtime is effectively reduced by up to 80.0 percent . The design of sampling filter and overlapped processing can bring about further 30.0 and 10.0 percent improvements in page processing. Qiumin Lu, Jiacheng Ma 0001, Yaozu Dong, Zhengwei Qi, Jianguo Yao 0002, Bingsheng He, Haibing Guan |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2019 | Systematically Testing and Diagnosing Responsiveness for Android AppsabstractApp responsiveness is the most intuitive interpretation of app performance from user's perspective. Traditional performance profilers only focus on one kind of program activities (e.g., CPU profiling), while the cause for slow responsiveness is diverse or even due to the joint effect of multiple kinds. Also, various test configurations, such as device hardware and wireless connectivity can have dramatic impact on particular program activities and indirectly affect app responsiveness. Conventional mobile testing lacks mechanisms to reveal configuration-sensitive bugs. In this paper, we propose AppSPIN, a tool to automatically diagnose app responsiveness bugs and systematically explore configuration-sensitive bugs. AppSPIN instruments the app to collect program events and UI responsiveness. The instrumented app is exercised with automated monkey testers and AppSPIN correlates excessive and lengthy program events with bad responsiveness detected at runtime. The diagnosis process also synthesizes the major resource bottleneck for the app. After one test run, AppSPIN automatically alters the test configuration to with most bottlenecked resource to further explore responsiveness bugs happened only with particular test configurations. Our preliminary experiments with 30 real-world apps show that AppSPIN can detect 123 responsiveness bugs and successfully diagnose the cause for 87% cases, within an average of 15-minute test time. Also with altered test configurations, AppSPIN uncovers a notable number of new bugs within four extra test runs. Wenhua Zhao, Zhenkai Ding, Mingyuan Xia 0001, Zhengwei Qi |
ICSME | 4 |
| 2019 | A distributed hypervisor for resource aggregation: posterabstractScale-out has become the standard answer to data analysis, machine learning and many other fields. Contrary to common belief, scale-up machines can outperform scale-out clusters for a considerable portion of tasks. However, those scale-up machines are not economical and may not be affordable for small businesses. This paper presents GiantVM, a distributed hypervisor that aggregates resources from multiple physical machines, providing the guest OS with a uniform hardware abstraction. We propose techniques to deal with the challenges of CPU, Memory, and I/O virtualization in distributed environments. Yubin Chen, Zhuocheng Ding, Yun Wang 0039, Zhengwei Qi, Haibing Guan |
PPoPP | 5 |
| 2019 | A scala based framework for developing acceleration systems with FPGAs
Yanqiang Liu, Yao Li 0004, Zhengwei Qi, Haibing Guan |
J. Syst. Archit. | 3 |
| 2019 | When I/O Interrupt Becomes System Bottleneck: Efficiency and Scalability Enhancement for SR-IOV Network VirtualizationabstractHigh performance networking interface cards (NIC) have become essential networking devices in commercial cloud computing environments. Therefore, efficient and scalable I/O virtualization is one of the primary challenges on virtualized cloud computing platforms. Single Root I/O Virtualization (SR-IOV) is a network interface technology that eliminates the overhead of redundant data copies and the virtual network switches through direct I/O in order to achieve nearly natural I/O performance. However, the SR-IOV still suffers from serious problems due to the high overhead for processing excessive network interrupts as well as the unpredictable and bursty traffic load in high-speed networking connections. In this paper, the defects of SR-IOV with 10 Gigabit Ethernet networking are studied first and two major challenges are identified: excessive interrupt rate and single threaded virtual network driver. Second, two interrupt rate control optimization schemes, called coarse-grained interrupt rate (CGR) control and adaptive interrupt rate (AIR) control are proposed. The proposed control schemes can significantly reduce the overhead and enhance the SR-IOV performance compared with the traditional driver with fixed interrupt throttle rate (FIR). In addition, multi-threaded VF driver (MTVD) is proposed that allows the SR-IOV VFs to leverage multi-core resources in order to achieve high scalability. Finally, these optimizations are implemented and detailed performance evaluations are conducted. The results show that CGR and AIR can improve the throughput by 2.26× and 2.97× while saving the CPU resources by 1.23 core and 1.44 core, respectively. The MTVD can achieve 2.03× performance with additional 1.46 cores consumption for VM using the SR-IOV driver. Jian Li 0021, Ruhui Ma, Zhengwei Qi, Haibing Guan |
IEEE Trans. Cloud Comput. | 5 |
| 2019 | Fairness-Efficiency Allocation of CPU-GPU Heterogeneous ResourcesabstractConsidering the performance improvement the cloud technology provides by processing workloads in parallel, applications and services are now migrating to online clouds. In a cloud platform, workloads can be executed in a virtualized environment to have a great improvement of the resource utilization. However, there is a new challenge in the allocation problem, which is quantifying and optimizing the fairness and efficiency of heterogeneous resources (CPUs and GPUs) required by applications such as cloud gaming. The solving approach needs scalarization methods of the requirement vector, relevant functions for fairness metrics, and an acceptable algorithm to solve that, where the difficulties mainly locate. We design an iterative, dynamic-adaptive heuristic solving algorithm Fairness-Efficiency Allocation (FEA) and optimize the implementation on a virtualized platform, which collects runtime data, allocates resources and reports differences. Data are recorded and analyzed to discover the effect of the allocation in different situations, including the promotion of fairness and the effect on the frame rate of the workloads. The result indicates that there is a considerable fairness improvement after the resource allocation, especially in situations that many virtual machines are executing simultaneously. Compared with the VGASA strategy, the fairness metric value improved 45 percent in three virtual machines' situation. Qiumin Lu, Jianguo Yao 0002, Zhengwei Qi, Bingsheng He, Haibing Guan |
IEEE Trans. Serv. Comput. | 3 |
| 2018 | Reproducible Interference-Aware Mobile TestingabstractMobile apps are born to work in an environment with ever-changing network connectivity, random hardware interruption, unanticipated task switches, etc. However, such interference cases are often oblivious in traditional mobile testing but happen frequently and sophisticatedly in the field, causing various robustness, responsiveness and consistency problems. In this paper, we propose JazzDroid to introduce interference to mobile testing. JazzDroid adopts a gray-box approach to instrument apps at binary level such that interference logic is inlined with app execution and can be triggered to effectively affect normal execution. Then, JazzDroid repeatedly orchestrates the instrumented app through app developers' existing tests and continuously randomizes interference on the fly to reveal possible faulty executions. Upon discovering problems, JazzDroid generates a test script with the user inputs from developers' tests and the interference injected for developers to reproduce the problems. At a high level, JazzDroid can be seamlessly integrated into app developers' testing procedures, detecting more problems from existing tests. We implement JazzDroid to function on unmodified apps directly from app markets and interface with de facto industrial testing toolchain. JazzDroid improves mobile testing by discovering 6x more problems, including crashes, functional bugs, UI consistency issues and common bug patterns that fail numerous apps. Weilun Xiong, Shihao Chen, Mingyuan Xia 0001, Zhengwei Qi |
ICSME | 5 |
| 2018 | HybridPass: Hybrid Scheduling for Mixed Flows in Datacenter NetworksabstractModern cloud applications generate millions of mixed flows transmitted between distributed nodes, and the typical latency-sensitive and throughput-intensive flows coexist in datacenter networks. Scheduling those mixed flows presents new challenges when meeting both low latency and high throughput requirements. This paper introduces HybridPass, a novel hybrid network architecture for datacenter networks, which is the first attempt to support both time-triggered and event-triggered scheduling in respect to latency-sensitive and throughput-intensive flows. To this end, we develop an arbiter, which uses a loosely synchronized time-triggered manner to allocate the network bandwidth for latency-sensitive and throughput-intensive flows from a global perspective. Specifically, the time-triggered scheduling aims to minimize the latency through establishing flow-level and task-level models. Then the event-triggered scheduling is developed to utilize the leftover bandwidth for throughput-intensive flows without any impact on latency-sensitive flows. Our experiments show that HybridPass can achieve up to 40.71% latency reduction for latency-sensitive flows compared with the baseline DCTCP while maintaining the high throughput with inappreciable 0.77% throughput sacrifice for throughput-intensive flows. Bo Peng 0043, Jianguo Yao 0002, Zhengwei Qi, Haibing Guan |
IPDPS | 3 |
| 2018 | Efficient shuffle management with SCache for DAG computing frameworksabstractIn large-scale data-parallel analytics, shuffle, or the cross-network read and aggregation of partitioned data between tasks with data dependencies, usually brings in large overhead. To reduce shuffle overhead, we present SCache, an open source plug-in system that particularly focuses on shuffle optimization. By extracting and analyzing shuffle dependencies prior to the actual task execution, SCache can adopt heuristic pre-scheduling combining with shuffle size prediction to pre-fetch shuffle data and balance load on each node. Meanwhile, SCache takes full advantage of the system memory to accelerate the shuffle process. We have implemented SCache and customized Spark to use it as the external shuffle service and co-scheduler. The performance of SCache is evaluated with both simulations and testbed experiments on a 50-node Amazon EC2 cluster. Those evaluations have demonstrated that, by incorporating SCache, the shuffle overhead of Spark can be reduced by nearly 89%, and the overall completion time of TPC-DS queries improves 40% on average. Zhouwang Fu, Tao Song 0003, Zhengwei Qi, Haibing Guan |
PPoPP | 3 |
| 2018 | gMig: Efficient GPU Live Migration Optimized by Software Dirty Page for Full VirtualizationabstractThis paper introduces gMig, an open-source and practical GPU live migration solution for full virtualization. By taking advantage of the dirty pattern of GPU workloads, gMig presents the One-Shot Pre-Copy combined with the hashing based Software Dirty Page technique to achieve efficient GPU live migration. Particularly, we propose three approaches for gMig: 1) Dynamic Graphics Address Remapping, which parses and manipulates GPU commands to adjust the address mapping to adapt to a different environment after migration, 2) Software Dirty Page, which utilizes a hashing based approach to detect page modification, overcomes the commodity GPU's hardware limitation, and speeds up the migration by only sending the dirtied pages, 3) One-Shot Pre-Copy, which greatly reduces the rounds of pre-copy of graphics memory. Our evaluation shows that gMig achieves GPU live migration with an average downtime of 302 ms on Windows and 119 ms on Linux. With the help of Software Dirty Page, the number of GPU pages transferred during the downtime is effectively reduced by 80.0%. Jiacheng Ma 0001, Yaozu Dong, Wentai Li, Zhengwei Qi, Bingsheng He, Haibing Guan |
VEE | 5 |
| 2018 | FastDesk: A remote desktop virtualization system for multi-tenant
Tao Song 0003, Jiewei Wu, Ruhui Ma, Alei Liang, Tao Gu 0001, Zhengwei Qi |
Future Gener. Comput. Syst. | 7 |
| 2018 | D2FL: Design and Implementation of Distributed Dynamic Fault LocalizationabstractCompromised or misconfigured routers have been a major concern in large-scale networks. Such routers sabotage packet delivery, and thus hurt network performance. Data-plane fault localization (FL) promises to solve this problem. Regrettably, the path-based FL fails to support dynamic routing, and the neighbor-based FL requires a centralized trusted administrative controller (AC) or global clock synchronization in each router and introduces storage overhead for caching packets. To address these problems, we introduce a dynamic distributed and low-cost model, D2FL. Using random two-hop neighborhood authentication, D2FL supports volatile path without the AC or global clock synchronization. Besides, D2FL requires only constant tens of KB for caching which is independent of the packet transmission rate. This is much less than the cache size of DynaFL or DFL which consumes several MB. The simulations show that D2FL achieves low false positive and false negative rate with no more than 3 percent bandwidth overhead. We also implement an open source prototype and evaluate its effect. The result shows that the performance burden in user space is less than 10 percent with the dynamic sampling algorithm. Fanfu Zhou, Zhengwei Qi, Jianguo Yao 0002, Ruhui Ma, Bin Wang 0062, Athanasios V. Vasilakos, Haibing Guan |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2018 | Scalable GPU Virtualization with Dynamic Sharing of Graphics Memory SpaceabstractWith increasing GPU-intensive workloads deployed on cloud, cloud service providers are seeking for practical and efficient GPU virtualization solutions. However, the cutting-edge GPU virtualization techniques such as gVirt still suffer from the restriction of scalability, which constrains the number of guest virtual GPU instances. This paper presents gScale, a scalable and practical open source GPU virtualization solution based on gVirt. gScale presents a sharing mechanism which combines partition and sharing together to break the hardware limitation of global graphics memory space. Particularly, we propose two approaches for gScale: (1) the private shadow graphics translation table (GTT) , which enables global graphics memory space sharing among virtual GPUs, (2) ladder mapping and fence memory space pool, which allows CPU access host physical memory space (serving the graphics memory) to bypass global graphics memory space. Furthermore, to mitigate the performance degradation caused by switching private shadow GTT when the number of vGPUs scales up, four other mechanisms are proposed: (1) slot sharing, which improves the performance of vGPU by dividing the high global graphics memory into multiple slots, (2) fine-grained slotting, which provides a flexible virtual graphics memory configuration, (3) predictive GTT copy mechanism, which reduces the performance loss by switching private shadow GTT before context switch, (4) predictive-copy aware scheduling, which maximizes the improvement of predictive GTT copy mechanism in cloud environment. Evaluation shows that gScale scales up to 15 guest virtual GPU instances in Linux or 12 guest virtual GPU instances in Windows, which is 5x and 4x, respectively, that of gVirt. At the same time, gScale incurs a slight but acceptable runtime overhead when hosting multiple virtual GPU instances. Mochi Xue, Jiacheng Ma 0001, Wentai Li, Yaozu Dong, Zhengwei Qi, Bingsheng He, Haibing Guan |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2017 | Scala Based FPGA Design Flow (Abstract Only)
Yanqiang Liu, Yao Li 0004, Weilun Xiong, Meng Lai, Zhengwei Qi, Haibing Guan |
FPGA | 6 |
| 2017 | Automated Resource Sharing for Virtualized GPU with Self-ConfigurationabstractIn this paper, we propose Auto-vGPU, a framework of automated resource sharing for virtualized GPU with self-configuration, to reduce manual intervention in system management while ensuring Service Level Agreement (SLA) targets. Auto-vGPU automatically collects the measurements of system metrics and learns a linear model for each application with dimension reduction. In order to fulfill the automated configuration of controller parameters, we propose a self-control-configuration method featuring the theory of automatic tuning of proportional-integral (PI) regulators. The experimental results of cloud gaming implementation demonstrate that Auto-vGPU is able to automatically build the low-dimension model and configure the control parameters without any manual interventions and the derived controller can adaptively allocate virtualized GPU resource to ensure the high performance of cloud applications. Jianguo Yao 0002, Qiumin Lu, Zhengwei Qi |
SRDS | 3 |
| 2017 | Ashman: A Bandwidth Fragmentation-Based Dynamic Flow Scheduling for Data Center NetworksabstractCurrent data center (DC) topologies provide redundant paths to satisfy significant communication requirements. Researchers have proposed various approaches including static and dynamic ones to utilize path diversity more effectively. The static methods adopt hash-based way to distribute flows onto multiple paths randomly, while the dynamic ones relying on centralized controller place flows on candidate paths. Unfortunately, the existing flow-scheduling policies have ignored and even resulted in network bandwidth fragmentation which adversely slows down transmission rate of new flows or affects chances of accepting new flow requests, especially when the granularity of flows is larger. This phenomenon is considered as a bottleneck for higher utilization of bandwidth resource in DCs. In this paper, we are the first to identify and define network bandwidth fragmentation within DC. Accordingly, we present Ashman, a flow-based dynamic scheduling approach to reduce bandwidth fragmentation by proactively considering potential large flows. This policy dynamically schedules the existing large flows to resolve congestions and maximize bandwidth utilization and network throughput. We describe our design and implementation in OpenFlow framework with unmodified hosts. Our evaluation on the software-defined networking simulator called Mininet demonstrates that Ashman effectively reduces bandwidth fragmentation, and thereby achieves higher network bandwidth utilization and the overall throughput of DC. Tao Song 0003, Ruhui Ma, Alei Liang, Zhengwei Qi, Haibing Guan |
Comput. J. | 6 |
| 2017 | Cost-Performance Modeling with Automated Benchmarking on Elastic Computing Clouds
Hongyan Mao, Zhengwei Qi, Jiangang Duan, Xinni Ge |
J. Grid Comput. | 2 |
| 2017 | Evolution of Cloud Operating System: From Technology to Ecosystem
Zuoning Chen, Kang Chen 0001, Jinlei Jiang, Lufei Zhang, Song Wu 0001, Zhengwei Qi, Chunming Hu, Yongwei Wu 0001, Yuzhong Sun, Aobing Sun, Zilu Kang |
J. Comput. Sci. Technol. | 6 |
| 2017 | ForenVisor: A Tool for Acquiring and Preserving Reliable Data in Cloud Live ForensicsabstractLive forensics is an important technique in cloud security but is facing the challenge of reliability. Most of the live forensic tools in cloud computing run either in the target Operating System (OS), or as an extra hypervisor. The tools in the target OS are not reliable, since they might be deceived by the compromised OS. Furthermore, traditional general purpose hypervisors are vulnerable due to their huge code size. However, some modules of a general purpose hypervisor, such as device drivers, are indeed unnecessary for forensics. In this paper, we propose a special purpose hypervisor, called ForenVisor, which is dedicated to reliable live forensics. The reliability is improved in three ways: reducing Trusted Computing Base (TCB) size by leveraging a lightweight architecture, collecting evidence directly from the hardware, and protecting the evidence and other sensitive files with Filesafe module. We have implemented a proof-of-concept prototype on the Windows platform, which can acquire the process data, raw memory, and I/O data, such as keystrokes and network traffic. Furthermore, we evaluate ForenVisor in terms of code size, functionality, and performance. The experiment results show that ForenVisor has a relatively small TCB size of about 13 KLOC, and only causes less than 10 percent performance reduction to the target system. In particular, our experiments verify that ForenVisor can guarantee that the protected files remain untampered, even when the guest OS is compromised by viruses, such as `ILOVEYOU' and Worm.WhBoy. Also, our system can be loaded as a hypervisor without needing to pause the target OS. This allows it to not only avoid destructing but also to gather the live evidence of the target OS. We also posted the source code of ForenVisor on Github. Zhengwei Qi, Chengcheng Xiang, Ruhui Ma, Jian Li 0021, Haibing Guan, David S. L. Wei |
IEEE Trans. Cloud Comput. | 1 |
| 2016 | gScale: Scaling up GPU Virtualization with Dynamic Sharing of Graphics Memory Space
Mochi Xue, Yaozu Dong, Jiacheng Ma 0001, Zhengwei Qi, Bingsheng He, Haibing Guan |
USENIX ATC | 6 |
| 2016 | AutoBench: Finding Workloads That You Need Using Pluggable Hybrid AnalysesabstractResearchers often rely on benchmarks to demonstrate feasibility or efficiency of their contributions. However, finding the right benchmark suite can be a daunting task - existing benchmark suites may be outdated, known to be flawed, or simply irrelevant for the proposed approach. Creating a proper benchmark suite is challenging, extremely time consuming, and also - unless it becomes widely popular - a thankless endeavor. In this paper, we introduce a novel approach to help researchers find relevant workloads for their experimental evaluation needs. Our approach relies on the huge number of open-source projects available in public repositories, and on unit testing having become best practice in software development. Using a repository crawler employing pluggable static and dynamic analyses for filtering and workload characterization, we allow users to automatically find projects with relevant workloads. Preliminary results presented here show that unit tests can provide a viable source of workloads, and that the combination of static and dynamic analyses improves the ability to identify relevant workloads that can serve as the basis for custom benchmark suites. Yudi Zheng, Andrea Rosà, Luca Salucci, Yao Li 0004, Haiyang Sun 0003, Omar Javed, Lubomír Bulej, Lydia Y. Chen, Zhengwei Qi, Walter Binder |
SANER | 9 |
| 2016 | A user mode CPU-GPU scheduling framework for hybrid workloads
Bin Wang 0062, Ruhui Ma, Zhengwei Qi, Jianguo Yao 0002, Haibing Guan |
Future Gener. Comput. Syst. | 3 |
| 2015 | Seagull - A Real-Time Coflow Scheduling SystemabstractData-parallel applications often generate hundreds of flows at the same time in data centers. Since these flows are always connected with application context, traditional flow-level optimization policies are hard to perform well in such collections. The coflow abstraction brings hope and opportunity to make the scheduling much more efficient. But exsiting schedule systems based on that related concept are either static (such as Varys) or impracticable (such as Baraat). In this paper, we address these limitations by presenting Seagull -- a dynamic precise coflow scheduling system to optimize the average CCT (Coflow Completion Time) and guaranteeing predictable completions within coflow deadlines. It's a centralized system which can share the bandwidth resources with background flows in the data center. Our experiments show that 80% CCT of the coflows is about 1.7× faster than Varys. As for deadline meeting, Seagull can guarantee about 50% of admitted coflows finishing within their deadline, which is 10% more precise than Varys. Zhouwang Fu, Tao Song 0003, Fuzong Wang, Zhengwei Qi |
CSCloud | 5 |
| 2015 | FLOWPROPHET: Generic and Accurate Traffic Prediction for Data-Parallel Cluster ComputingabstractData-parallel computing frameworks (DCF) such as MapReduce, Spark, and Dryad etc. Have tremendous applications in big data and cloud computing, and throw tons of flows into data center networks. In this paper, we design and implement FLOW PROPHET, a general framework to predict traffic flows for DCFs. To this end, we analyze and summarize the common features of popular DCFs, and gain a key insight: since application logic in DCFs is naturally expressed by directed acyclic graphs (DAG), DAG contains necessary time and data dependencies for accurate flow prediction. Based on the insight, FLOW PROPHET extracts DAGs from user applications, and uses the time and data dependencies to calculate flow information 4-tuple, (source, destination, flow size, establish time), ahead-of-time for all flows. We also provide generic programming interface to FLOW PROPHET, so that current and future DCFs can deploy FLOW PROPHET readily. We implement FLOW PROPHET on both Spark and Hadoop, and perform extensive evaluations on a testbed with 37 physical servers. Our implementation and experiments demonstrate that, with time in advance and minimal cost, FLOW PROPHET can achieve almost 100% accuracy in source, destination, and flow size predictions. With accurate prediction from FLOW PROPHET, the job completion time of a Hadoop TeraSort benchmark is reduced by 12.52% on our cluster with a simple network scheduler. Hao Wang 0022, Li Chen 0008, Kai Chen 0005, Ziyang Li 0003, Yiming Zhang 0003, Haibing Guan, Zhengwei Qi, Dongsheng Li 0001, Yanhui Geng |
ICDCS | 7 |
| 2015 | DefDroid: Securing Android with Fine-Grained Security PolicyabstractAndroid occupies the absolute dominant position in mobile operating system and has the largest market share.Meanwhile, Android faces the risk of malicious insiders leaking sensitive information.In this paper, we present DefDroid, a repackaging tool for enforcing security policies by modifying Android applications without root privilege.The main advantages of DefDroid are that it provides a user-friendly interface to configure fine-grained policies and it supplies multiple deployment methods.We have implemented policies aimed at three types of services of Android system, i.e., content provider, file system, and network.We choose 74 arbitrary applications from Android market and the experimental results show that the successful rate of repackaging applications is about 94.6% which effectively improve the privacy security of Android system while the increased overhead can be tolerated. Shuohong Wang, Haiyang Sun 0003, Zhengwei Qi |
SEKE | 4 |
| 2015 | Flexible and Extensible Runtime Verification for JavaabstractRuntime verification validates the correctness of a program's execution trace.Much work has been done on improving the expressiveness and efficiency of runtime verification.However, current approaches require static deployment of the verification logic and are often restricted to a limited set of events that can be captured and analyzed, hindering the adoption of runtime verification in production systems.A popular system for runtime verification in Java, JavaMOP (Monitor-Oriented Programming in Java), suffers from the aforementioned limitations due to its dependence on AspectJ, which supports neither dynamic weaving nor an extensible join-point model.In this paper, we extend the JavaMOP framework with a dynamic deployment API and a new MOP specification translator, which targets the domain-specific aspect language DiSL instead of AspectJ; DiSL offers an open join-point model that allows for extensions.A case study on lambda expressions in Java8 demonstrates the extensibility of our approach.Moreover, in comparison with JavaMOP using load-time weaving, our implementation reduces runtime overhead by 21%, and heap memory usage by 16%, on average. Chengcheng Xiang, Zhengwei Qi, Walter Binder |
SEKE | 2 |
| 2015 | Effective Real-Time Android Application AuditingabstractMobile applications can access both sensitive personal data and the network, giving rise to threats of data leaks. App auditing is a fundamental program analysis task to reveal such leaks. Currently, static analysis is the de facto technique which exhaustively examines all data flows and pinpoints problematic ones. However, static analysis generates false alarms for being over-estimated and requires minutes or even hours to examine a real app. These shortcomings greatly limit the usability of automatic app auditing. To overcome these limitations, we design AppAudit that relies on the synergy of static and dynamic analysis to provide effective real-time app auditing. AppAudit embodies a novel dynamic analysis that can simulate the execution of part of the program and perform customized checks at each program state. AppAudit utilizes this to prune false positives of an efficient but over-estimating static analysis. Overall, AppAudit makes app auditing useful for app market operators, app developers and mobile end users, to reveal data leaks effectively and efficiently. We apply AppAudit to more than 1,000 known malware and 400 real apps from various markets. Overall, AppAudit reports comparative number of true data leaks and eliminates all false positives, while being 8.3x faster and using 90% less memory compared to existing approaches. AppAudit also uncovers 30 data leaks in real apps. Our further study reveals the common patterns behind these leaks: 1) most leaks are caused by 3rd-party advertising modules; 2) most data are leaked with simple unencrypted HTTP requests. We believe AppAudit serves as an effective tool to identify data-leaking apps and provides implications to design promising runtime techniques against data leaks. Mingyuan Xia 0001, Lu Gong, Yuanhao Lyu, Zhengwei Qi, Xue (Steve) Liu |
IEEE Symposium on Security and Privacy | 4 |
| 2015 | Boosting GPU Virtualization Performance with Hybrid Shadow Page Tables
Yaozu Dong, Mochi Xue, Zhengwei Qi, Haibing Guan |
USENIX ATC | 5 |
| 2015 | A survey on data center networking for cloud computing
Bin Wang 0062, Zhengwei Qi, Ruhui Ma, Haibing Guan, Athanasios V. Vasilakos |
Comput. Networks | 2 |
| 2015 | Flexible and Extensible Runtime Verification for Java (Extended Version)abstractRuntime verification validates the correctness of a program’s execution trace. Much work has been done on improving the expressiveness and efficiency of runtime verification. However, current approaches require static deployment of the verification logic and are often restricted to a limited set of events that can be captured and analyzed, hindering the adoption of runtime verification in production systems. A popular system for runtime verification in Java, JavaMOP (Monitor-Oriented Programming in Java), suffers from the aforementioned limitations due to its dependence on AspectJ, which supports neither dynamic weaving nor an extensible join-point model. In this article, we extend the JavaMOP framework with a dynamic deployment API and a new MOP specification translator, which targets the domain-specific aspect language DiSL instead of AspectJ; DiSL offers an open join-point model that allows for extensions. A case study on lambda expressions in Java8 demonstrates the extensibility of our approach. Moreover, in comparison with JavaMOP using load-time weaving, our implementation reduces runtime overhead by 32%, and heap memory usage by 13%, on average. Chengcheng Xiang, Zhengwei Qi, Walter Binder |
Int. J. Softw. Eng. Knowl. Eng. | 2 |
| 2015 | Energy-Efficient SLA Guarantees for Virtualized GPU in Cloud GamingabstractBoth power consumption and SLA guarantees are important concerns for cloud gaming. Recently, various approaches have been developed to effectively reduce GPU power consumption by making GPU run at low frequencies. However, virtual machines (VMs) running on the same physical GPU with virtualization technology are correlated, because the change of GPU frequencies will affect the SLA performance of all the VMs. In fact, both reducing power consumption and guaranteeing SLA should work together under the considerations of their correlations. This paper proposes a novel two-layer control architecture called energy-efficient SLA guarantees for virtualized GPU (EvGPU) based on well-established feedback control techniques. The first control loop adopts a proportional-integral (PI) controller to ensure SLA guarantees, which in particular are measured at a predefined level of the frames per second (FPS) for each online game. The secondary power control loop then adjusts GPU frequency through dynamic voltage/frequency scaling (DVFS) to reduce power consumption based on the current FPS achieved by the first loop. Empirical results demonstrate that the proposed solution can effectively reduce GPU power consumption, while achieving the required SLA performance in virtualized GPU for cloud gaming. Haibing Guan, Jianguo Yao 0002, Zhengwei Qi |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2014 | ScalaHDL: Express and test hardware designs in a Scala DSLabstractField Programmable Gate Arrays, or FPGAs, allow designers to implement hardware designs using hardware description languages (HDLs). This type of designs have been gaining significant popularity since improvements in clock frequencies, of high-end CPUs, have started to level off and other alternatives have been explored to accelerate computations. However, traditional HDLs lack a number of modern facilities and a rich ecosystem to express and test designs, which severely restricts the productivity of designers. In this paper, we propose ScalaHDL, an open-source domain-specific language (DSL) built on top of Scala, that enables designers to describe algorithms using a multi-paradigm programming language, and generate the required Verilog code to implement such systems. In addition, these designs can be simulated so that values can be tested programmatically using unit-tests. With ScalaHDL, designers can also leverage the rich and mature ecosystems provided by Java and Scala. Yao Li 0004, Antonio Roldao Lopes, Zhouyun Xu, Zhengwei Qi, Haibing Guan |
ICCD | 4 |
| 2014 | Pinso: Precise Isolation of Concurrency Bugs via Delta TriagingabstractConcurrent programs are known to be difficult to test and maintain. These programs often fail because of concurrency bugs caused by non-deterministic interleavings among shared memory accesses. Even though a concurrency bug can be detected, it is still hard to isolate the root cause of the bug, due to the challenge in understanding the complex thread interleavings or schedules. In this paper, we propose a practical and precise isolation technique for concurrent bugs called Pinso that seeks to exploit the non-deterministic nature of concurrency bugs and accurately find the root causes of program error, to further help developers maintain concurrent programs. Pinso profiles runtime inter-thread interleavings based on a set of summarized memory access patterns, and then, isolates suspicious interleaving patterns in the triaging phase. Using a filtration-oriented scheduler, Pinso effectively eliminates false positives that are irrelevant to the bug. We evaluate Pinso with 11 real-world concurrency bugs, including single- and multi-variable violation, from sever/desktop concurrent applications (MySQL, Apache, and several others). Experiments indicate that our tool accurately isolates the root causes of all the bugs. Bo Liu 0001, Zhengwei Qi, Bin Wang 0062, Ruhui Ma |
ICSME | 2 |
| 2014 | Dynamic program analysis - Reconciling developer productivity and tool performance
Aibek Sarimbekov, Yudi Zheng, Danilo Ansaloni, Lubomír Bulej, Lukás Marek, Walter Binder, Petr Tuma 0001, Zhengwei Qi |
Sci. Comput. Program. | 8 |
| 2014 | VGRIS: Virtualized GPU Resource Isolation and Scheduling in Cloud GamingabstractTo achieve efficient resource management on a graphics processing unit (GPU), there is a demand to develop a framework for scheduling virtualized resources in cloud gaming. In this article, we propose VGRIS, a resource management framework for virtualized GPU resource isolation and scheduling in cloud gaming. A set of application programming interfaces (APIs) is provided so that a variety of scheduling algorithms can be implemented within the framework without modifying the framework itself. Three scheduling algorithms are implemented by the APIs within VGRIS. Experimental results show that VGRIS can effectively schedule GPU resources among various workloads. Zhengwei Qi, Jianguo Yao 0002, Chao Zhang 0115, Zhizhou Yang, Haibing Guan |
ACM Trans. Archit. Code Optim. | 1 |
| 2014 | Multi-Granularity Memory Mirroring via Binary Translation in Cloud EnvironmentsabstractAs the size of DRAM memory grows in clusters, memory errors are common. Current memory availability strategies mostly focus on memory backup and error recovery. Hardware solutions like mirror memory needs costly peripheral equipments while existing software approaches reduce the expense but are limited by the high overhead in practical usage. Moreover, in cloud environments, containers such as LXC now can be used as process and application-level virtualization to run multiple isolated systems on a single host. In this paper, we present a novel system called Memvisor to provide high availability memory mirroring. It is a software approach achieving flexible multi-granularity memory mirroring based on virtualization and binary translation. We can flexibly set memory areas to be mirrored or not from process level to the whole user mode applications. Then, all memory write instructions are duplicated. Data written to memory are synchronized to backup space in the instruction level. If memory failures happen, Memvisor will recover the data from the backup space. Compared with traditional software approaches, the instruction level synchronization lowers the probability of data loss and reduces the backup overhead. The results show that Memvisor outperforms the state-of-the-art software approaches even in the worst case. Zhengwei Qi, Haoliang Dong, Yaozu Dong, Haibing Guan |
IEEE Trans. Netw. Serv. Manag. | 1 |
| 2014 | vGASA: Adaptive Scheduling Algorithm of Virtualized GPU Resource in Cloud GamingabstractAs the virtualization technology for GPUs matures, cloud gaming has become an emerging application among cloud services. In addition to the poor default mechanisms of GPU resource sharing, the performance of cloud games is inevitably undermined by various runtime uncertainties such as rendering complex game scenarios. The question of how to handle the runtime uncertainties for GPU resource sharing remains unanswered. To address this challenge, we propose vGASA, a virtualized GPU resource adaptive scheduling algorithm in cloud gaming. vGASA interposes scheduling algorithms in the graphics API of the operating system, and hence the host graphic driver or the guest operating system remains unmodified. To fulfill the service level agreement as well as maximize GPU usage, we propose three adaptive scheduling algorithms featuring feedback control that mitigates the impact of the runtime uncertainties on the system performance. The experimental results demonstrate that vGASA is able to maintain frames per second of various workloads at the desired level with the performance overhead limited to 5-12 percent. Chao Zhang 0115, Jianguo Yao 0002, Zhengwei Qi, Haibing Guan |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2013 | kMemvisor: flexible system wide memory mirroring in virtual environments
Bin Wang 0062, Zhengwei Qi, Haibing Guan, Haoliang Dong, Yaozu Dong |
HPDC | 2 |
| 2013 | VGRIS: virtualized GPU resource isolation and scheduling in cloud gaming
Chao Zhang 0115, Zhengwei Qi, Jianguo Yao 0002, Yin Wang 0001, Haibing Guan |
HPDC | 3 |
| 2013 | An Automation-Assisted Empirical Study on Lock Usage for Concurrent ProgramsabstractNowadays concurrent programs are becoming more and more important with the development of hardware and network technologies. However, it is not easy for programmers to write reliable concurrent programs. Concurrency characteristics such as thread-interleaving make it difficult to debug or maintain concurrent programs. Although there are lots of research work on concurrency such as multi-thread testing tools, concurrent program verification and data race detection, all of them leave open problems. For instance, some are not scalable enough for large real world applications and some may report false warnings. Since locks are widely used to protect shared memory, it is beneficial for both programmers and tool designers in all fields to have a good understanding of common lock usage patterns in real world concurrent programs. This paper reports an empirical study on lock usage in concurrent programs. It is based on our automatic lock analysis tool called LUPA. The study analyzes how lock is used in concurrent programs and how lock usage changes throughout the product environment. In this study, four representative concurrent programs (Apache httpd, Mysql, Aget, Pbzip2) are selected, of which both lock manifestation and lock usage pattern in different versions are studied. This study reveals some interesting findings including but not limited to: (1) about 80.5% of the lock related functions acquire only one lock, (2) simple lock patterns account for 54.5% of all lock usage in real world applications, (3) only 12 out of 527 detected patterns belong to condition lock pattern which may lead to vulnerabilities easily, (4) only 0.65% of the functions are lock related. Additionally, a potential bug caused by problematic locking pattern is found. Zhengwei Qi, Shiqiu Huang, Chengcheng Xiang, Yudi Zheng, Yin Wang 0001, Haibing Guan |
ICSM | 2 |
| 2013 | SRIDesk: A Streaming based Remote Interactivity architecture for desktop virtualization systemabstractIn recent years, desktop virtualization trends to be a new extension of virtualization framework. The existing desktop virtualization systems suffer performance degradation in terms of response time and video quality. However, previous remote access approaches are designed for standalone architectures and require semantic information which is not transparent to OS. So they are not feasible in desktop virtualization systems. In this paper, we propose SRIDesk, a Streaming based Remote Interactivity architecture for Desktop virtualization system. SRIDesk resides in the host through intercepting virtual display device, which is transparent to guest OS and its applications. SRIDesk integrates server-push streaming mechanism with H.264 encoder into virtualization system, which provides high quality display with low bandwidth consumption and low latency of interaction. We have implemented the SRIDesk prototype in a KVM system. Experimental results show that SRIDesk has low CPU-load, low bandwidth and good scalability. We compared SRIDesk with other popular platforms, including X, VNC, RDP and THiNC. SRIDesk outperformed other systems in bandwidth with no more than 2Mbps and 94% video quality. SRIDesk also achieved lowest latency in WAN environment among all systems. Jiewei Wu, Zhengwei Qi, Haibing Guan |
ISCC | 3 |
| 2013 | Introduction to dynamic program analysis with DiSLabstractDiSL is a new domain-specific language for bytecode instrumentation with complete bytecode coverage. It reconciles expressiveness and efficiency of low-level bytecode manipulation libraries with a convenient, high-level programming model inspired by aspect-oriented programming. This paper summarizes the language features of DiSL and gives a brief overview of several dynamic program analysis tools that were ported to DiSL. DiSL is available as open-source under the Apache 2.0 license. Lukás Marek, Yudi Zheng, Danilo Ansaloni, Lubomír Bulej, Aibek Sarimbekov, Walter Binder, Zhengwei Qi |
ICPE | 7 |
| 2013 | A multi-objective ant colony system algorithm for virtual machine placement in cloud computing
Yongqiang Gao, Haibing Guan, Zhengwei Qi, Liang Liu 0010 |
J. Comput. Syst. Sci. | 3 |
| 2013 | Quality of service aware power management for virtualized data centers
Yongqiang Gao, Haibing Guan, Zhengwei Qi, Bin Wang 0062, Liang Liu 0010 |
J. Syst. Archit. | 3 |
| 2013 | A refined decompiler to generate C code with high readabilityabstractSUMMARY As a key part of reverse engineering, decompilation plays a very important role in software security and maintenance. A number of tools, such as Boomerang and IDA Hex_rays, have been developed to translate executable programs into source code in a relatively high‐level language. Unfortunately, most existing decompilation tools suffer from low accuracy in identifying variables, functions, and composite structures, resulting in poor readability. To address these limitations, we present a practical decompiler called C‐Decompiler for Windows C programs that (i) uses a shadow stack to perform refined data flow analysis, (ii) adopts inter‐basic‐block register propagation to reduce redundant variables, and (iii) recognizes library (i.e., Standard Template Library) functions by signatures. We evaluate and compare the decompilation quality of C‐Decompiler with two existing tools, Boomerang and IDA Hex_rays, considering four aspects: function analysis, variable expansion rate, total percentage reduction, and cyclomatic complexity. Our experimental results show that on average, C‐Decompiler has the highest total percentage reduction of 55.91%, lowest variable expansion rate of 55.79%, and the same cyclomatic complexity as the original source code for each considered application. Furthermore, in our experiments, C‐Decompiler is able to recognize functions with a lower false positive and false negative rate than the other decompilers. A case study and our evaluation results confirm that C‐Decompiler is a practical tool to produce highly readable C‐style code. Copyright © 2012 John Wiley & Sons, Ltd. Gengbiao Chen, Zhengwei Qi, Shiqiu Huang, Kangqi Ni, Yudi Zheng, Walter Binder, Haibing Guan |
Softw. Pract. Exp. | 2 |
| 2012 | Java Bytecode Instrumentation Made Easy: The DiSL Framework for Dynamic Program Analysis
Lukás Marek, Yudi Zheng, Danilo Ansaloni, Aibek Sarimbekov, Walter Binder, Petr Tuma 0001, Zhengwei Qi |
APLAS | 7 |
| 2012 | Memvisor: Application Level Memory Mirroring via Binary TranslationabstractMemory failures are common in clusters, and their destructive effects (e.g., increasing downtime and losing data) make users suffer great loss. Current memory availability strategies mostly require extra expensive hardware. Software approaches based on check pointing technologies intend to reduce the expense, but their high overhead limits the practical usage. In this paper, we present a novel system called Memvisor to provide software mirrored memory for applications. Specifically, all memory write instructions are duplicated. Data written to memory are synchronized to backup space. If memory failures happen, Memvisor will recover the data from the backup space. Compared with traditional software approaches, the instruction-level synchronization lowers the probability of data loss and reduces backup overhead. The results show that even in the worst case, Memvisor outperforms the state-of-the-art software approaches. Haoliang Dong, Bin Wang 0062, Haiyang Sun 0003, Zhengwei Qi, Haibing Guan, Yaozu Dong |
CLUSTER | 5 |
| 2012 | Optimizing virtual machines using hybrid virtualization
Qian Lin 0002, Zhengwei Qi, Jiewei Wu, Yaozu Dong, Haibing Guan |
J. Syst. Softw. | 2 |
| 2011 | DsVD: An Effective Low-Overhead Dynamic Software Vulnerability DiscovererabstractDynamic taint analysis based software vulnerability and malware detection is an effective method to detect a wide range of vulnerabilities. Unfortunately, existing systems suffer from requirement of source code, high overhead or shortage of discovery rules, which limit their usage. This paper proposes a low-overhead vulnerability discovery system called DsVD (Dynamic Software Vulnerabilities Discoverer). DsVD works on X86 executables and does not need any hardware change. A new taint state called controlled-taint is introduced to detect more types of vulnerabilities. Our experiments show that DsVD can effectively detect various software vulnerabilities. DsVD incurs very low overhead, only 3.1 times on average for SPECINT2006 benchmarks. With some optimizations such as Irrelevant API Filter and Basic Block Handling, it can reduce runtime overhead by a factor of 4-11 times. Zhushou Tang, Kan Zhou, Zhengwei Qi, Haibing Guan |
ISADS | 5 |
| 2010 | Real-time Enhancement for Xen HypervisorabstractSystem virtualization, which provides good isolation, is now widely used in server consolidation. Meanwhile, one of the hot topics in this field is to extend virtualization for embedded systems. However, current popular virtualization platforms do not support real-time operating systems such as embedded Linux well because the platform is not real-time ware, which will bring low-performance I/O and high scheduling latency. The goal of this paper is to optimize the Xen virtualization platform to be real-time operating system friendly. We improve two aspects of the Xen virtualization platform. First, we improve the xen scheduler to manage the scheduling latency and response time of the real-time operating system. Second, we import multiple real-time operating systems balancing method. Our experiment demonstrates that our enhancement to the Xen virtualization platform support real-time operating system well and the improvement to the real-time performance is about 20%. Peijie Yu, Mingyuan Xia 0001, Qian Lin 0002, Shang Gao 0009, Zhengwei Qi, Kai Chen 0006, Haibing Guan |
EUC | 6 |
| 2010 | Enhanced Privilege Separation for Commodity Software on Virtualized PlatformabstractConventional privilege separation can effectively reduce the TCB size by granting privilege to only the privileged compartments. However, since they this approach relies on process isolation to ensure security assurance, malware exploiting against kernel components can easily compromise. Meanwhile, the frequent inter-process communications between separated processes inevitably incur notable overhead. To ameliorate these problems, we propose to perform privilege separation without partitioning application into two processes. Instead, we leverage virtualization to enforce the isolation of sensitive portions from other untrusted code. The virtual machine monitor intercepts all the code context switches transparently without requiring the application to explicitly use IPC as privilege context transition. We have implemented a prototype of our system, named Coir, based on commodity hypervisor Xen. Evaluation of our prototype includes a real-world remote control application, which is partitioned and protected in Coir-enabled hypervisor on unmodified Windows XP. We discuss the isolation strength as well as the performance penalty of our system based on the practical case. Mingyuan Xia 0001, Qian Lin 0002, Zhengwei Qi, Haibing Guan |
ICPADS | 4 |
| 2010 | CoDBT: A multi-source dynamic binary translator using hardware-software collaborative techniques
Haibing Guan, Bo Liu 0001, Zhengwei Qi, Yindong Yang, Alei Liang |
J. Syst. Archit. | 3 |
| 2009 | A Heuristic Policy-based System Call Interposition in Dynamic Binary TranslationabstractDynamic binary translation (DBT) is a well known software technology that enables seamless cross-ISA execution. Unfortunately, many malicious programs that may lead to unauthorized access can run easily and unrestrictedly under the DBT system. Because these malicious programs must go through the system call interface to take malicious action, system call interposition has become a widely used technique for intrusion detection and prevention. In this paper, we present HPSCIBit, a solution that efficiently confines malicious applications, supports automatic policy generation and interactive policy generation, intrusion detection and prevention in the DBT system. The experimental result on SPEC2000 CINT benchmarks shows that HPSCIBit is an effective and low overhead solution to the cross-ISA security issues. Deen Zheng, Zhengwei Qi, Alei Liang, Haibing Guan, Liang Liu 0010 |
MASS | 2 |
| 2008 | An Online Model Checking Tool for Safety and Liveness BugsabstractModern software model checkers are usually used to find safety violations. However, checking liveness properties can offer a more natural and effective way to detect errors, particularly in complex concurrent and distributed e-business systems. Specifying global liveness properties which should always eventually be true proves to be more desirable, but it is hard for existing software model checkers to verify liveness in real codes because doing so requires finding an infinite execution. For solving such a challenge, this paper proposes an online checking tool to verify the safety and liveness properties of complex systems. We adopt the linear temporal logic to describe the semantics of the finite model checking, use binary instrumentation to obtain the distribute states and apply a checking engine to dynamically verify the finite trace linear temporal logic properties. At last, we demonstrate the method in a distributed system using distributed protocol Paxos and achieve good results by experiments. Zhengwei Qi, Liang Liu 0010, Alei Liang, Hao Wang 0022, Ying Chen 0004 |
ICPADS | 1 |
| 2006 | Membrane Calculus: a formal method for Grid transactionsabstractAbstract The research of transaction processing in Web Services and Grid Services is active in academic and engineering areas. However, the formal method of transaction processing has not been fully investigated in the literature. This paper proposes a preliminary theoretical model called Membrane Calculus based on Membrane Computing and Petri nets to formalize Grid transactions. Five kinds of transition rules in Membrane Calculus (including object rules and membrane rules) are introduced and the operational semantics of transition rules are defined. Then, a typical long‐running transaction example is presented to demonstrate the use of Membrane Calculus. Finally, the rewriting logic tool Maude is adopted to specify and execute the specification of this example. Copyright © 2006 John Wiley & Sons, Ltd. Zhengwei Qi, Minglu Li 0001, Dongyu Shi, Jinyuan You |
Concurr. Comput. Pract. Exp. | 1 |