VLDB 2026 Research / reviewers in the wild / expert
Binyu Zang
dblp:86/680
· DBLP profile ↗
119ranked-venue papers
1as first author
36since 2021 · last 2026
0000-0002-1968-7645ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 76 · 24 since 2021Software engineering, systems software and programming languages · 25 · 10 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 3 since 2021Security and privacy · 9Databases, data management, data science and information retrieval · 3 · 2 since 2021Computer networks · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Artificial intelligence and machine learning · 1Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | miniK8s: A Pedagogical Cloud-Native SystemabstractThe rapid adoption of cloud-native technologies, particularly containerization and orchestration with systems like Kubernetes, necessitates their effective integration into undergraduate computer science curricula. However, the complexity of production-grade cloud-native systems present a steep learning curve. Traditional pedagogical approaches often involve either oversimplified toy projects built from scratch, which lack real-world relevance, or the direct use of complex cloud systems, which can obscure fundamental concepts. To overcome these barriers, we introduce miniK8s, a lightweight, Kubernetes-like platform designed for an undergraduate cloud computing course. miniK8s distinguishes itself by promoting the pedagogical vision of teaching students to build a substantial and realistic system by integrating existing, robust open-source components with a hand-written, simplified cornerstone component. With miniK8s, students can both learn the key concepts inside the cloud architecture (e.g., resource scaling) and build practical systems with open-source building blocks. We have successfully used miniK8s as a project in a cloud computing course for hundreds of undergraduate students, which can significantly enhance students' understanding of how real-world cloud-native systems work and their ability to build a complex system. Dong Du 0003, Mingyu Wu 0001, Haibo Chen 0001, Binyu Zang |
SIGCSE (1) | 4 |
| 2025 | D-VSync: Decoupled Rendering and Displaying for Smartphone GraphicsabstractRendering service, which typically orchestrates screen display and UI through Vertical Synchronization (VSync), is an indispensable system service for user experiences of smartphone OSes (e.g., Android, OpenHarmony, and iOS). The recent trend of large high-frame-rate screens, stunning visual effects, and physics-based animations has placed unprecedented pressure on the VSync-based rendering architecture, leading to higher frame drops and longer rendering latency. Yuanpei Wu, Dong Du 0003, Yubin Xia, Ming Fu, Binyu Zang, Haibo Chen 0001 |
ASPLOS (1) | 6 |
| 2025 | OS Rendering Service Made Parallel with Out-of-Order Execution and In-Order Commit
Yuanpei Wu, Yubin Xia, Yang Yu 0002, Ming Fu, Binyu Zang, Haibo Chen 0001 |
OSDI | 6 |
| 2025 | Harmonizing Security and Performance in Microkernel File Servers
Wentai Li, Jinyu Gu 0001, Yubin Xia, Binyu Zang |
J. Comput. Sci. Technol. | 5 |
| 2024 | Jade: A High-throughput Concurrent Copying Garbage CollectorabstractGarbage collection (GC) pauses are a notorious issue threatening the latency of applications. To mitigate this problem, state-of-the-art concurrent copying collectors allow GC threads to run simultaneously with application threads (mutators) in nearly all GC phases. However, the design of concurrent copying collectors does not always lead to low application latency. To this end, this work studies the behaviors of mainstream concurrent copying collectors in OpenJDK and mainly focuses on long application pauses under heavy workloads. By analyzing the design of those collectors, this work uncovers that lengthy pre-reclamation cycles (including GC phases before actual memory release), high GC frequency, and large metadata maintenance overhead are major factors for long pauses. Therefore, this work proposes Jade, a concurrent copying collector aiming to achieve both short pauses and high GC efficiency. Compared with existing collectors, Jade provides a group-wise collection mechanism to shorten pre-reclamation cycles while controlling GC frequency. It also embraces a generational heap layout and a single-phase algorithm to maximize young GC's throughput. The evaluation results on representative latency-critical applications show that Jade can reach sub-millisecond-level pauses even under heavy workloads and significantly improve applications' peak throughput compared with state-of-the-art concurrent collectors. Mingyu Wu 0001, Yude Lin, Yifeng Jin, Zhe Li 0037, Hongtao Lyu, Denghui Dong, Haibo Chen 0001, Binyu Zang |
EuroSys | 12 |
| 2024 | Characterization and Reclamation of Frozen Garbage in Managed FaaS WorkloadsabstractFaaS (function-as-a-service) is becoming a popular workload in cloud environments due to its virtues such as auto-scaling and pay-as-you-go. High-level languages like JavaScript and Java are commonly used in FaaS for programmability, but their managed runtimes complicate memory management in the cloud. This paper first observes the issue of frozen garbage, which is caused by freezing cached function instances where their threads have been paused but the unused memory (e.g., garbage) is not reclaimed due to the semantic gap between FaaS and the managed runtime. This paper presents the first characterization of the negative effects induced by frozen garbage with various functions, which uncovers that it can occupy more than half of FaaS instances' memory resources on average. To this end, this paper proposes Desiccant, a freeze-aware memory manager for managed workloads in FaaS, which reclaims idle memory resources consumed by frozen garbage from managed runtime instances and thus notably improves memory efficiency. The evaluation on various FaaS workloads shows that Desiccant can reduce FaaS functions' peak memory consumption by up to 6.72×. Such saved memory consumption allows caching more FaaS instances to reduce the frequency of cold boots (creating instances before function execution) and p99 latency by up to 4.49× and 37.5%, respectively. Ziming Zhao 0003, Mingyu Wu 0001, Haibo Chen 0001, Binyu Zang |
EuroSys | 4 |
| 2024 | Toward an SGX-Friendly Java RuntimeabstractHardware enclaves assist in constructing a trusted execution environment (TEE) to store private code and data and thus become an appealing solution to enhance applications’ security. Nevertheless, state-of-the-art enclave implementations like Intel Software Guard Extensions (SGX) have severe performance issues and hinder the deployment of more complicated applications, especially those written in high-level languages like Java. To reduce the performance overhead, prior work has partitioned applications or rebuilt lightweight language runtimes, but they either require manual labor from developers or fail to provide full-fledged support for existing applications. This work instead providesSAJ, a runtime built upon a full-fledged Java virtual machine (JVM) and thus requires no modifications to applications.SAJfirst analyzes the performance of vanilla JVMs running in enclaves and finds that the memory management overhead and boot phase are culprits for performance slowdown. For memory management,SAJintroduces SGX-aware heap layout and garbage collector, which reduces both GC and application execution time. As for the boot phase,SAJintroduces an address-conscious launching mechanism to improve the boot performance. The evaluation under representative Java applications shows thatSAJcan reduce the overall GC pause time, application time, and boot time by 2.93$\boldsymbol{\times}$, 2.58$\boldsymbol{\times}$, and 2.73$\boldsymbol{\times}$on average, respectively. Mingyu Wu 0001, Zhe Li 0037, Haibo Chen 0001, Binyu Zang, Sanhong Li, Haitao Song 0001 |
IEEE Trans. Computers | 4 |
| 2024 | Ad Hoc Transactions through the Looking Glass: An Empirical Study of Application-Level Transactions in Web ApplicationsabstractMany transactions in web applications are constructed ad hoc in the application code. For example, developers might explicitly use locking primitives or validation procedures to coordinate critical code fragments. We refer to database operations coordinated by application code as ad hoc transactions . Until now, little is known about them. This paper presents the first comprehensive study on ad hoc transactions. By studying 91 ad hoc transactions among eight popular open-source web applications, we found that (i) every studied application uses ad hoc transactions (up to 16 per application), 71 of which play critical roles; (ii) compared with database transactions, concurrency control of ad hoc transactions is much more flexible; (iii) ad hoc transactions are error-prone—53 of them have correctness issues, and 33 of them are confirmed by developers; and (iv) ad hoc transactions have the potential for improving performance in contentious workloads by utilizing application semantics such as access patterns. Based on these findings, we discuss the implications of ad hoc transactions to the database research community. Chuzhe Tang, Qianmian Yu, Binyu Zang, Haibing Guan, Haibo Chen 0001 |
ACM Trans. Database Syst. | 5 |
| 2023 | BeeHive: Sub-second Elasticity for Web Services with Semi-FaaS ExecutionabstractFunction-as-a-service (FaaS), an emerging cloud computing paradigm, is expected to provide strong elasticity due to its promise to auto-scale fine-grained functions rapidly. Although appealing for applications with good parallelism and dynamic workload, this paper shows that it is non-trivial to adapt existing monolithic applications (like web services) to FaaS due to their complexity. To bridge the gap between complicated web services and FaaS, this paper proposes a runtime-based Semi-FaaS execution model, which dynamically extracts time-consuming code snippets (closures) from applications and offloads them to FaaS platforms for execution. It further proposes BeeHive, an offloading framework for Semi-FaaS, which relies on the managed runtime to provide a fallback-based execution model and addresses the performance issues in traditional offloading mechanisms for FaaS. Meanwhile, the runtime system of BeeHive selects offloading candidates in a user-transparent way and supports efficient object sharing, memory management, and failure recovery in a distributed environment. The evaluation using various web applications suggests that the Semi-FaaS execution supported by BeeHive can reach sub-second resource provisioning on commercialized FaaS platforms like AWS Lambda, which is up to two orders of magnitude better than other alternative scaling approaches in cloud computing. Ziming Zhao 0003, Mingyu Wu 0001, Binyu Zang, Haibo Chen 0001 |
ASPLOS (2) | 4 |
| 2023 | CPS: A Cooperative Para-virtualized Scheduling Framework for Manycore MachinesabstractToday's cloud platforms offer large virtual machine (VM) instances with multiple virtual CPUs (vCPU) on manycore machines. These machines typically have a deep memory hierarchy to enhance communication between cores. Although previous researches have primarily focused on addressing the performance scalability issues caused by the double scheduling problem in virtualized environments, they mainly concentrated on solving the preemption problem of synchronization primitives and the traditional NUMA architecture. This paper specifically targets a new aspect of scalability issues caused by the absence of runtime hypervisor-internal states (RHS). We demonstrate two typical RHS problems, namely the invisible pCPU (physical CPU) load and dynamic cache group mapping. These RHS problems result in a collapse in VM performance and low CPU utilization because the guest VM lacks visibility into the latest runtime internal states maintained by the hypervisor, such as pCPU load and vCPU-pCPU mappings. Consequently, the guest VM makes inefficient scheduling decisions. Yuxuan Liu 0019, Tianqiang Xu, Zeyu Mi, Zhichao Hua 0001, Binyu Zang, Haibo Chen 0001 |
ASPLOS (4) | 5 |
| 2023 | ISA-Grid: Architecture of Fine-grained Privilege Control for Instructions and RegistersabstractIsolation is a critical mechanism for enhancing the security of computer systems. By controlling the access privileges of software and hardware resources, isolation mechanisms can decouple software into multiple isolated components and enforce the principle of least privilege. While existing isolation systems primarily focus on memory isolation, they overlook the isolation of instruction and register resources, which we refer to as ISA (Instruction Set Architecture) resources. However, previous works have shown that exploiting ISA resources can lead to serious security problems, such as breaking the system's memory isolation property by abusing x86's CR3 register. Furthermore, existing hardware only provides privilege-level-based access control for ISA resources, which is too coarse-grained for software decoupling. For example, ARM Cortex A53 has several hundred system instructions/registers, but only four exception levels (EL0 to EL3) are provided. Additionally, more than 100 instructions/registers for system control are available in only EL1 (the kernel mode). To address this problem, this paper proposes ISA-Grid, an architecture of fine-grained privilege control for instructions and registers. ISA-Grid is a hardware extension that enables the creation of multiple ISA domains, with each domain having different privileges to access instructions and registers. The ISA domain can provide bit-level fine-grained privilege control for registers. We implemented prototypes of ISA-Grid based on two different CPU cores: 1) a RISC-V CPU core on an FPGA board and 2) an x86 CPU core on a simulator. We applied ISA-Grid to different cases, including Linux kernel decomposition and enhancing existing security systems, to demonstrate how ISA-Grid can isolate ISA resources and mitigate attacks based on abusing them. The performance evaluation results on both x86 and RISC-V platforms with real-world applications showed that ISA-Grid has negligible runtime overhead (less than 1%). Shulin Fan, Zhichao Hua 0001, Yubin Xia, Haibo Chen 0001, Binyu Zang |
ISCA | 5 |
| 2023 | Security and Performance in the Delegated User-level Virtualization
Dingji Li, Zeyu Mi, Yuxuan Liu 0019, Binyu Zang, Haibing Guan, Haibo Chen 0001 |
OSDI | 5 |
| 2023 | Bifrost: Analysis and Optimization of Network I/O Tax in Confidential Virtual Machines
Dingji Li, Zeyu Mi, Chenhui Ji, Yifan Tan, Binyu Zang, Haibing Guan, Haibo Chen 0001 |
USENIX ATC | 5 |
| 2023 | Bridging the Gap between Relational OLTP and Graph-based OLAP
Sijie Shen, Zihang Yao, Lei Wang 0004, Longbin Lai, Li Su 0005, Rong Chen 0001, Wenyuan Yu, Haibo Chen 0001, Binyu Zang, Jingren Zhou 0001 |
USENIX ATC | 11 |
| 2023 | Understanding and Mitigating Twin Function Misuses in Operating System KernelabstractMajor operating system kernels expose twin functions, which are groups of internal primitives that have mostly common but slightly diverging semantics, to kernel modules and subsystems. They are created to make the basic primitives work well in various scenarios. Unfortunately, though being expected as solutions, twin functions may turn to problem-makers in practice. As we have observed from over 500 patches applied to upstream Linux and FreeBSD, developers choose an improper one from the twins, leaving the kernel with stability and security bugs as well as error-prone code. In this paper, we aim to understand and mitigate the twin function misuse problem. First, we provide an informative discussion on the misuse-fix patches. We find that violating the constraints from calling context, missing the primitives with better performance, lacking the necessary security enhancements, and breaking the kernel coding style are the four major factors that lead to misuse. We then identify the programming rules from the patches and apply them with a static program analysis tool extended from Coccinelle, including callgraph tainting and type-based function pointer resolving. We have 136 patches accepted by the Linux community and fix 320 new misuses in the upstream Linux kernel. Jinyu Gu 0001, Jiacheng Shi 0002, Haroran Su, Wentai Li, Binyu Zang, Haibing Guan, Haibo Chen 0001 |
IEEE Trans. Computers | 5 |
| 2023 | Flock: Towards Multitasking Virtual Machines for Function-as-a-ServiceabstractFaaS, or function as a service, promises unprecedented cost-efficiency and elasticity thanks to its on-demand and fine-grained execution nature. However, modern FaaS platforms mainly adopt virtual machines (VMs) or containers as a computing abstraction, which incurs costs like high startup latency, large memory footprint, and high communication overhead. Multi-tasking virtual machines (MVMs), which allow co-executing multiple functions in the same managed language runtime, are appealing for FaaS due to their lightweight nature. Unfortunately, existing MVMs are not designed for FaaS. The proposed abstraction of MVMs does not provide specialized support for fine-grained, latency-sensitive functions and their chain-like execution patterns. Meanwhile, the underlying runtime still contains many global modules and lacks essential support for function-level resource accounting and isolation. To this end, this work proposesFlock, a retrofitted MVM for FaaS execution, which provides FaaS-aware abstractions namedfuncletsand enhanced runtime support for isolation.Flockis implemented atop the HotSpot JVM of OpenJDK 8. Performance evaluation shows thatFlockresults in up to three orders of magnitude performance improvement over state-of-the-art FaaS platforms like OpenWhisk while providing sufficient isolation support for FaaS functions. Ziming Zhao 0003, Mingyu Wu 0001, Xujie Cao, Haibo Chen 0001, Binyu Zang |
IEEE Trans. Computers | 5 |
| 2022 | Serverless computing on heterogeneous computersabstractExisting serverless computing platforms are built upon homogeneous computers, limiting the function density and restricting serverless computing to limited scenarios. We introduce Molecule, the first serverless computing system utilizing heterogeneous computers. Molecule enables both general-purpose devices (e.g., Nvidia DPU) and domain-specific accelerators (e.g., FPGA and GPU) for serverless applications that significantly improve function density (50% higher) and application performance (up to 34.6x). To achieve these results, we first propose XPU-Shim, a distributed shim to bridge the gap between underlying multi-OS systems (when using general-purpose devices) and our serverless runtime (i.e., Molecule). We further introduce vectorized sandbox, a sandbox abstraction to abstract hardware heterogeneity (when using domain-specific accelerators). Moreover, we also review state-of-the-art serverless optimizations on startup and communication latency and overcome the challenges to implement them on heterogeneous computers. We have implemented Molecule on real platforms with Nvidia DPUs and Xilinx FPGAs and evaluate it using benchmarks and real-world applications. Dong Du 0003, Xueqiang Jiang, Yubin Xia, Binyu Zang, Haibo Chen 0001 |
ASPLOS | 5 |
| 2022 | Asymmetry-aware scalable lockingabstractThe pursuit of power-efficiency is popularizing asymmetric multicore processors (AMP) such as ARM big.LITTLE, Apple M1 and recent Intel Alder Lake with big and little cores. However, we find that existing scalable locks fail to scale on AMP and cause collapses in either throughput or latency, or both, because their implicit assumption of symmetric cores no longer holds. To address this issue, we propose the first asymmetry-aware scalable lock named LibASL. LibASL provides a new lock ordering guided by applications' latency requirements, which allows big cores to reorder with little cores for higher throughput under the condition of preserving applications' latency requirements. Using LibASL only requires linking the applications with it and, if latency-critical, inserting few lines of code to annotate the coarse-grained latency requirement. We evaluate LibASL in various benchmarks including five popular databases on Apple M1. Evaluation results show that LibASL can improve the throughput by up to 5 times while precisely preserving the tail latency designated by applications. Jinyu Gu 0001, Dahai Tang, Kenli Li 0001, Binyu Zang, Haibo Chen 0001 |
PPoPP | 5 |
| 2022 | Ad Hoc Transactions in Web Applications: The Good, the Bad, and the UglyabstractMany transactions in web applications are constructed ad hoc in the application code. For example, developers might explicitly use locking primitives or validation procedures to coordinate critical code fragments. We refer to database operations coordinated by application code as ad hoc transactions. Until now, little is known about them. This paper presents the first comprehensive study on ad hoc transactions. By studying 91 ad hoc transactions among 8 popular open-source web applications, we find that (i) every studied application uses ad hoc transactions (up to 16 per application), 71 of which play critical roles; (ii) compared with database transactions, concurrency control of ad hoc transactions is much more flexible; (iii) ad hoc transactions are error-prone-53 of them have correctness issues, and 33 of them are confirmed by developers; and (iv) ad hoc transactions have the potential to improve performance in contentious workloads by utilizing application semantics such as access patterns. Based on the findings, we discuss the implications of ad hoc transactions to the database research community. Chuzhe Tang, Qianmian Yu, Binyu Zang, Haibing Guan, Haibo Chen 0001 |
SIGMOD Conference | 5 |
| 2022 | Zero-Change Object Transmission for Distributed Big Data Analytics
Mingyu Wu 0001, Shuaiwei Wang, Haibo Chen 0001, Binyu Zang |
USENIX ATC | 4 |
| 2022 | General and Fast Inter-Process Communication via Bypassing Privileged SoftwareabstractIPC (Inter-Process Communication) is a widely used operating system (OS) technique that allows one process to invoke the services of other processes. The IPC participants may share the same OS (internal IPC) or use a separate OS (external IPC). Even though a long line of researches has optimized the performance of IPC, it is still a major factor of the run-time overhead of IPC-intensive applications. Furthermore, there is no one-size-fits-all solution for both internal and external IPC. This paper presents SkyBridge, a general communication technique designed and optimized for both types of IPC. SkyBridge requires no involvement of the privileged software (the kernel or the hypervisor) and enables a process to directly switch to the virtual address space of the target process, regardless of whether they are running on the same OS or not. We have implemented SkyBridge on two microkernels (seL4 and Google Zircon) as well as an open-source serverless hypervisor (Firecracker). The evaluation results show that SkyBridge improves the latency of internal IPC and external IPC by up to 19.6x and 1265.7x, respectively. Zeyu Mi, Haoqi Zhuang, Binyu Zang, Haibo Chen 0001 |
IEEE Trans. Computers | 3 |
| 2022 | Colony: A Privileged Trusted Execution Environment With ExtensibilityabstractThe code base of system software is growing fast, which results in a large number of vulnerabilities: for example, 296 CVEs have been found in Xen hypervisor and 2195 CVEs in Linux kernel. To reduce the reliance on the trust of system software, many researchers try to provide trusted execution environments (TEEs), which can be categorized into two types: non-privileged TEEs and privileged TEEs. Non-privileged TEEs (e.g., Intel SGX) are extensible, but cannot protect security services like virtual machine introspection (VMI) due to the lack of system-level semantics. On the contrary, privileged TEEs (e.g., the secure world of ARM TrustZone) have system-level semantics, but any additional service implemented in the privileged TEE directly increases the TCB of the entire system. In this article, we propose a new design of TEE to support system-level security services and achieve better extensibility with a small TCB. Each TEE instance of the proposed design is named aColony. Specifically, we introduce asecure monitorfor isolation and capability management. EachColonyis assigned capabilities to access only necessary system-level semantics. We use the new TEE to build four security services, including secure device accessing, VMI tools, a system call tracer, and a much more complex service to virtualize ARM TrustZone with multipleColonies. We have implemented the system on ARMv7 and ARMv8 platforms, in Xen hypervisor and Linux kernel, and perform a detailed evaluation to show its efficiency.11.This paper is an extended version of the conference paper published in USENIX Security’17: vTZ: Virtualizing ARM TrustZone[29]. A brief summary of differences is in Section8. Yubin Xia, Zhichao Hua 0001, Yang Yu 0002, Jinyu Gu 0001, Haibo Chen 0001, Binyu Zang, Haibing Guan |
IEEE Trans. Computers | 6 |
| 2022 | DrTM+B: Replication-Driven Live Reconfiguration for Fast and General Distributed Transaction ProcessingabstractRecent in-memory database systems leverage advanced hardware features like RDMA to provide transaction processing at millions of transactions per second. Distributed transaction processing systems can scale to even higher rates, especially for partitionable workloads. Unfortunately, it is challenging to sustain such high rates during live reconfiguration of partitions. In this article, we observe that state-of-the-art approaches would cause notable performance disruption under fast transaction processing. To this end, this article presents DrTM+B, a live reconfiguration approach that seamlessly repartitions data with little performance disruption to running transactions. DrTM+B uses a pre-copy-based mechanism to avoid excessive data transfer by leveraging common properties in recent transactional systems. DrTM+B's reconfiguration plans reduce data movement by preferring existing data replicas, while copying data from multiple replicas asynchronously and in parallel. It further reuses the log forwarding mechanism in primary-backup replication to seamlessly track and forward dirty database tuples and avoids iterative copying costs. To commit a reconfiguration plan in a transactional-safe way, DrTM+B designs a cooperative commit protocol for synchronization of data and state among replicas. To boost the performance during data migration, DrTM+B combines the pre-copy and post-copy schemes to propose a hybrid copy scheme. The live reconfiguration approach can also coexist with fault-tolerance mechanisms of primary-backup replication to provide high availability. Evaluation on a working system based on DrTM+R with 3-way replication using typical OLTP workloads like TPC-C and SmallBank shows that DrTM+B incurs only very small performance degradation during live reconfiguration and provides high availability. Both the reconfiguration time and the downtime are also minimal. Sijie Shen, Xingda Wei, Rong Chen 0001, Haibo Chen 0001, Binyu Zang |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2022 | Wukong+G: Fast and Concurrent RDF Query Processing Using RDMA-Assisted GPU Graph ExplorationabstractRDF graph has been increasingly used to store and represent information shared over the Web, including social graphs and knowledge bases. With the increasing scale of RDF graphs and the concurrency level of SPARQL queries, current RDF systems are confronted with inefficient concurrent query processing on massive data parallelism. The situation becomes more severe in the face of data-intensive queries (aka heavy query), which usually lead to suboptimal response time (latency) as well as throughput collapse. In this article, we present Wukong+G, the first graph-based distributed RDF query processing system that efficiently exploits the hybrid parallelism of CPU and GPU. Wukong+G is made fast and concurrent with four key designs. First, Wukong+G tames massive random memory accesses in graph exploration by efficiently mapping data between CPU and GPU for latency hiding, including a set of techniques like query-aware prefetching, pattern-aware pipelining and fine-grained swapping. Second, Wukong+G scales up by introducing a GPU-friendly RDF store to support RDF graphs exceeding GPU memory size, by using techniques like predicate-based grouping, pairwise caching and look-ahead replacing to narrow the gap between host and device memory scale. Third, Wukong+G scales out through a communication layer that decouples the transferring process for query metadata and intermediate results, and further leverages both native and GPUDirect RDMA to enable efficient communication on a CPU/GPU cluster. Finally, Wukong+G simultaneously runs multiple queries on a single GPU to improve overall throughput and fully exploits hardware heterogeneity (CPU/GPU) by scheduling a single query on CPU and GPU adaptively. We have implemented Wukong+G by extending a state-of-the-art distributed RDF store (i.e., Wukong) with distributed GPU support. Evaluation on a heterogeneous CPU/GPU cluster with RDMA-capable network shows that Wukong+G outperforms Wukong by up to 9.0× (from 2.3×) and scales well on 10 GPU cards for heavy queries. Wukong+G can also improve both latency and throughput by more than one order of magnitude when facing hybrid workloads. Zihang Yao, Rong Chen 0001, Binyu Zang, Haibo Chen 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2021 | Bridging the performance gap for copy-based garbage collectors atop non-volatile memoryabstractNon-volatile memory (NVM) is expected to revolutionize the memory hierarchy with not only non-volatility but also large capacity and power efficiency. Memory-intensive applications, which are often written in managed languages like Java, would run atop NVM for better cost-efficiency. Unfortunately, such applications may suffer from performance slowdown due to the unmanaged performance gap between DRAM and NVM. This paper studies the performance of a series of Java applications atop NVM and uncovers that the copy-based garbage collection (GC), the mainstream GC algorithm, is an NVM-unfriendly component in JVM. GC becomes a severe performance bottleneck especially when memory resource is scarce. To this end, this paper analyzes the memory behavior of copy-based GC and uncovers that its inappropriate usage on NVM bandwidth is the main reason for its performance slowdown. This paper thus proposes two NVM-aware optimizations: write cache and header map, to effectively manage the limited NVM bandwidth. It further improves the GC performance with hardware instructions like non-temporal memory accesses and prefetching. We have implemented the optimizations on two mainstream copy-based garbage collectors in OpenJDK. Evaluation with various memory-intensive applications shows that our optimizations can improve the GC time, application execution time, application tail latency by up to 2.69×, 11.0%, and 5.09×, respectively. Yanfei Yang, Mingyu Wu 0001, Haibo Chen 0001, Binyu Zang |
EuroSys | 4 |
| 2021 | Efficiently Recovering Stateful System Components of Multi-server MicrokernelsabstractMicrokernel OSes provide OS services through mutually-isolated system servers running in different user processes, which brings stronger fault isolation than monolithic OSes. Nevertheless, considering the fault recovery capability of system servers, most existing microkernel OSes usually do no more than restarting a fault server, which will cause a server to lose all its running states and then may affect all the applications relying on it. In this paper, we present a mechanism named TxIPC that can efficiently recover stateful system servers on microkernel OSes. Since a system server provides the service by inter-process communication (IPC), TxIPC makes it fault resilient by handling each IPC in a transaction-like manner. Specifically, if a fault happens in a server (during one IPC handling procedure), TxIPC aborts all the updates made by the IPC and thus recovers the server from that fault. Evaluations show that TxIPC can enable servers to recover from 99.8% (injected) faults with 3%-45 % performance overhead on application benchmarks, which significantly outperforms existing counterparts. Wentai Li, Jinyu Gu 0001, Binyu Zang |
ICDCS | 4 |
| 2021 | Unifying Timestamp with Transaction Ordering for MVCC with Decentralized Scalar Timestamp
Xingda Wei, Rong Chen 0001, Haibo Chen 0001, Zhenhan Gong, Binyu Zang |
NSDI | 6 |
| 2021 | Scalable Memory Protection in the PENGLAI Enclave
Erhu Feng, Dong Du 0003, Bicheng Yang, Xueqiang Jiang, Yubin Xia, Binyu Zang, Haibo Chen 0001 |
OSDI | 7 |
| 2021 | Retrofitting High Availability Mechanism to Tame Hybrid Transaction/Analytical Processing
Sijie Shen, Rong Chen 0001, Haibo Chen 0001, Binyu Zang |
OSDI | 4 |
| 2021 | TwinVisor: Hardware-isolated Confidential Virtual Machines for ARMabstractConfidential VM, which offers an isolated execution environment for cloud tenants with limited trust in the cloud provider, has recently been deployed in major clouds such as AWS and Azure. However, while ARM has become increasingly popular in cloud data centers, existing confidential VM designs mainly leverage specialized x86 hardware extensions (e.g., AMD SEV and Intel TDX) to isolate VMs upon a shared hypervisor. Dingji Li, Zeyu Mi, Yubin Xia, Binyu Zang, Haibo Chen 0001, Haibing Guan |
SOSP | 4 |
| 2021 | Characterizing and Optimizing Remote Persistent Memory with RDMA and NVM
Xingda Wei, Xiating Xie, Rong Chen 0001, Haibo Chen 0001, Binyu Zang |
USENIX ATC | 5 |
| 2021 | TZ-Container: protecting container from untrusted OS with ARM TrustZone
Zhichao Hua 0001, Yang Yu 0002, Jinyu Gu 0001, Yubin Xia, Haibo Chen 0001, Binyu Zang |
Sci. China Inf. Sci. | 6 |
| 2021 | Revisiting Persistent Indexing Structures on Intel Optane DC Persistent Memory
Heng Bu, Mingkai Dong 0002, Jifei Yi, Binyu Zang, Haibo Chen 0001 |
J. Comput. Sci. Technol. | 4 |
| 2021 | Enclavisor: A Hardware-Software Co-Design for Enclaves on Untrusted CloudabstractThe releases of Intel SGX and AMD SEV mark the transition of hardware-based enclaves from research prototypes to mainstream products. These two paradigms of secure enclaves are attractive to both the cloud providers and tenants, since security is one of the key pillars of cloud computing. However, it is found that current hardware-defined enclaves are not flexible and efficient enough for the cloud. For example, although SGX can provide strong memory protection with both confidentiality and integrity, the size of secure memory is tightly restricted. On the contrary, SEV enables enclaves to use more memory but has critical security flaws due to no memory integrity protection. Meanwhile, both types of enclaves have relatively long booting latency, which makes them not suitable for short-term tasks like serverless workloads. After an in-depth analysis, we find that there are some intrinsic tradeoffs between security and performance due to the limitation of architectural designs. In this article, we investigate a novel hardware-software co-design of enclaves to meet the requirements of cloud by placing a part of the logic of the enclave mechanism into a lightweight software layer, named Enclavisor, to achieve a balance between security, performance, and flexibility. Specifically, our implementation is based on AMD's SEV and, Enclavisor is placed in the guest kernel mode of SEV's secure virtual machines. Enclavisor inherently supports memory encryption with no memory limitation and also achieves efficient booting, multiple enclave granularities, and post-launch remote attestation. Meanwhile, we also propose hardware/software solutions to mitigate the security flaws caused by the lack of memory integrity. We implement a prototype of Enclavisor on an AMD SEV server. The experiments on both micro-benchmarks and application benchmarks show that enclaves on Enclavisor can have close-to-native performance. Jinyu Gu 0001, Xinyue Wu, Bojun Zhu, Yubin Xia, Binyu Zang, Haibing Guan, Haibo Chen 0001 |
IEEE Trans. Computers | 5 |
| 2021 | Boosting Inter-process Communication with Architectural SupportabstractIPC (inter-process communication) is a critical mechanism for modern OSes, including not only microkernels such as seL4, QNX, and Fuchsia where system functionalities are deployed in user-level processes, but also monolithic kernels like Android where apps frequently communicate with plenty of user-level services. However, existing IPC mechanisms still suffer from long latency. Previous software optimizations of IPC usually cannot bypass the kernel that is responsible for domain switching and message copying/remapping across different address spaces; hardware solutions such as tagged memory or capability replace page tables for isolation, but usually require non-trivial modification to existing software stack to adapt to the new hardware primitives. In this article, we propose a hardware-assisted OS primitive, XPC (Cross Process Call), for efficient and secure synchronous IPC. XPC enables direct switch between IPC caller and callee without trapping into the kernel and supports secure message passing across multiple processes without copying. We have implemented a prototype of XPC based on the ARM AArch64 with Gem5 simulator and RISC-V architecture with FPGA boards. The evaluation shows that XPC can reduce IPC call latency from 664 to 21 cycles, 14×–123× improvement on Android Binder (ARM), and improve the performance of real-world applications on microkernels by 1.6× on Sqlite3. Yubin Xia, Dong Du 0003, Zhichao Hua 0001, Binyu Zang, Haibo Chen 0001, Haibing Guan |
ACM Trans. Comput. Syst. | 4 |
| 2021 | XStore: Fast RDMA-Based Ordered Key-Value Store Using Remote Learned CacheabstractRDMA ( Remote Direct Memory Access ) has gained considerable interests in network-attached in-memory key-value stores. However, traversing the remote tree-based index in ordered key-value stores with RDMA becomes a critical obstacle, causing an order-of-magnitude slowdown and limited scalability due to multiple round trips. Using index cache with conventional wisdom—caching partial data and traversing them locally—usually leads to limited effect because of unavoidable capacity misses, massive random accesses, and costly cache invalidations. We argue that the machine learning (ML) model is a perfect cache structure for the tree-based index, termed learned cache . Based on it, we design and implement XStore , an RDMA-based ordered key-value store with a new hybrid architecture that retains a tree-based index at the server to perform dynamic workloads (e.g., inserts) and leverages a learned cache at the client to perform static workloads (e.g., gets and scans). The key idea is to decouple ML model retraining from index updating by maintaining a layer of indirection from logical to actual positions of key-value pairs. It allows a stale learned cache to continue predicting a correct position for a lookup key. XStore ensures correctness using a validation mechanism with a fallback path and further uses speculative execution to minimize the cost of cache misses. Evaluations with YCSB benchmarks and production workloads show that a single XStore server can achieve over 80 million read-only requests per second. This number outperforms state-of-the-art RDMA-based ordered key-value stores (namely, DrTM-Tree, Cell, and eRPC+Masstree) by up to 5.9× (from 3.7×). For workloads with inserts, XStore still provides up to 3.5× (from 2.7×) throughput speedup, achieving 53M reqs/s. The learned cache can also reduce client-side memory usage and further provides an efficient memory-performance tradeoff, e.g., saving 99% memory at the cost of 20% peak throughput. Xingda Wei, Rong Chen 0001, Haibo Chen 0001, Binyu Zang |
ACM Trans. Storage | 4 |
| 2020 | Catalyzer: Sub-millisecond Startup for Serverless Computing with Initialization-less BootingabstractServerless computing promises cost-efficiency and elasticity for high-productive software development. To achieve this, the serverless sandbox system must address two challenges: strong isolation between function instances, and low startup latency to ensure user experience. While strong isolation can be provided by virtualization-based sandboxes, the initialization of sandbox and application causes non-negligible startup overhead. Conventional sandbox systems fall short in low-latency startup due to their application-agnostic nature: they can only reduce the latency of sandbox initialization through hypervisor and guest kernel customization, which is inadequate and does not mitigate the majority of startup overhead. Dong Du 0003, Yubin Xia, Binyu Zang, Guanglu Yan, Chenggang Qin, Qixuan Wu, Haibo Chen 0001 |
ASPLOS | 4 |
| 2020 | Characterizing serverless platforms with serverlessbenchabstractServerless computing promises auto-scalability and cost-efficiency (in "pay-as-you-go" manner) for high-productive software development. Because of its virtue, serverless computing has motivated increasingly new applications and services in the cloud. This, however, also presents new challenges including how to efficiently design high-performance serverless platforms and how to efficiently program on the platforms. Dong Du 0003, Yubin Xia, Binyu Zang, Ziqian Lu, Pingchao Yang, Chenggang Qin, Haibo Chen 0001 |
SoCC | 5 |
| 2020 | TEEp: Supporting Secure Parallel Processing in ARM TrustZoneabstractMachine learning applications are getting prevelent on various computing platforms, including cloud servers, smart phones, IoT devices, etc. For these applications, security is one of the most emergent requirements. While trusted execution environment (TEE) like ARM TrustZone has been widely used to protect critical prodecures including fingerprint authentication and mobile payment, state-of-the-art implementations of TEE OS lack the support for multi-threading and are not suitable for computing-intensive workloads. This is because current TEE OSes are usually designed for hosting security critical tasks, which are typically small and non-computing-intensive. Thus, most of TEE OSes do not support multi-threading in order to minimize the size of the trusted computing base (TCB). In this paper, we propose TEEp, a system that enables multi-threading in TEE without weakening security, and supports existing multi-threaded applications to run directly in TEE. Our design includes a novel multithreading mechanism based on the cooperation between the TEE OS and the host OS, without trusting the host OS. We implement our system based on OP-TEE and port it to two platforms: a HiKey 970 development board as mobile platform, and a Huawei Hi1610 ARM server as server platform. We run TensorFlow Lite on the development board and TensorFlow on the server for performance evaluation in TEE. The result shows that our system can improve the throughput of TensorFlow Lite on 5 models to 3.2x when 4 cores are available, with 13.5% overhead compared with Linux on average. Zinan Li, Wenhao Li 0009, Yubin Xia, Binyu Zang |
ICPADS | 4 |
| 2020 | No barrier in the road: a comprehensive study and optimization of ARM barriersabstractIn this paper, we present the first comprehensive performance characterization and optimization of ARM barriers on both mobile and server platforms. We draw a set of observations through several abstracted models and validate them in scenarios where barriers are intensively used. We find that (1) order-preserving approaches without involving the bus significantly outperform other approaches, and (2) the tremendous overhead mostly comes from barriers strictly following remote memory references. Usually, such barriers are inserted when threads are exchanging data, and they are used to ensure the relative order between storing the data to a shared buffer and setting a flag to inform the receiver. Based on the observations, we propose a new mechanism, Pilot, to remove such barriers by leveraging the single-copy atomicity to piggyback the flag with the data. Applying Pilot only requires minor changes to applications and provides 10%-360% performance improvements in multiple benchmarks, which are close to the ideal performance without barriers. Binyu Zang, Haibo Chen 0001 |
PPoPP | 2 |
| 2020 | Platinum: A CPU-Efficient Concurrent Garbage Collector for Tail-Reduction of Interactive Services
Mingyu Wu 0001, Ziming Zhao 0003, Yanfei Yang, Haibo Chen 0001, Binyu Zang, Haibing Guan, Sanhong Li, Chuansheng Lu, Tongbao Zhang |
USENIX ATC | 6 |
| 2020 | (Mostly) Exitless VM Protection from Untrusted Hypervisor through Disaggregated Nested Virtualization
Zeyu Mi, Dingji Li, Haibo Chen 0001, Binyu Zang, Haibing Guan |
USENIX Security Symposium | 4 |
| 2020 | GCPersist: an efficient GC-assisted lazy persistency framework for resilient Java applications on NVMabstractThe emergence of non-volatile memory (NVM) has stimulated broad interests in building efficient and persistent systems and programming models. However, most prior work is built atop an eager persistency model, which mandates applications to persist their data as soon as possible and thus causes considerable overhead. Besides, prior work mainly focuses on native languages and overlooks the interactions with the managed runtime system in a high-level language. Such issues limit the scope of applications on NVM, especially for resilient applications that already have reliable but inefficient recovery mechanisms. This paper proposes GCPersist, an easy-to-use NVM programming framework atop a lazy persistency model to defer the persistency of user data for better performance, with the assistance of the garbage collection (GC) module in the managed runtime. GCPersist further provides differentiated persistency modes to reduce the runtime overhead. We have implemented GCPersist on the HotSpot JVM of OpenJDK and the evaluation results on Intel Optane DC persistent memory devices show that GCPersist performs well with resilient applications (like Spark) by reducing the recovery time by up to 3.26X while introducing only 1--6% runtime overhead during normal execution. Mingyu Wu 0001, Haibo Chen 0001, Binyu Zang, Haibing Guan |
VEE | 4 |
| 2020 | Optimistic Transaction Processing in Deterministic Database
Zhiyuan Dong, Chuzhe Tang, Haibo Chen 0001, Binyu Zang |
J. Comput. Sci. Technol. | 6 |
| 2019 | XPC: architectural support for secure and efficient cross process callabstractMicrokernel has many intriguing features like security, fault-tolerance, modularity and customizability, which recently stimulate a resurgent interest in both academia and industry (including seL4, QNX and Google's Fuchsia OS). However, IPC (inter-process communication), which is known as the Achilles' Heel of microkernels, is still the major factor for the overall (poor) OS performance. Besides, IPC also plays a vital role in monolithic kernels like Android Linux, as mobile applications frequently communicate with plenty of user-level services through IPC. Previous software optimizations of IPC usually cannot bypass the kernel which is responsible for domain switching and message copying/remapping; hardware solutions like tagged memory or capability replace page tables for isolation, but usually require non-trivial modification to existing software stack to adapt the new hardware primitives. In this paper, we propose a hardware-assisted OS primitive, XPC (Cross Process Call), for fast and secure synchronous IPC. XPC enables direct switch between IPC caller and callee without trapping into the kernel, and supports message passing across multiple processes through the invocation chain without copying. The primitive is compatible with the traditional address space based isolation mechanism and can be easily integrated into existing microkernels and monolithic kernels. We have implemented a prototype of XPC based on a Rocket RISC-V core with FPGA boards and ported two microkernel implementations, seL4 and Zircon, and one monolithic kernel implementation, Android Binder, for evaluation. We also implement XPC on GEM5 simulator to validate the generality. The result shows that XPC can reduce IPC call latency from 664 to 21 cycles, up to 54.2x improvement on Android Binder, and improve the performance of real-world applications on microkernels by 1.6x on Sqlite3 and 10x on an HTTP server with minimal hardware resource cost. Dong Du 0003, Zhichao Hua 0001, Yubin Xia, Binyu Zang, Haibo Chen 0001 |
ISCA | 4 |
| 2019 | Pisces: A Scalable and Efficient Persistent Transactional Memory
Jinyu Gu 0001, Xiayang Wang, Binyu Zang, Haibing Guan, Haibo Chen 0001 |
USENIX ATC | 5 |
| 2019 | ScissorGC: scalable and efficient compaction for Java full garbage collectionabstractJava runtime frees applications from manual memory management through automatic garbage collection (GC). This, however, is usually at the cost of stop-the-world pauses. State-of-the-art collectors leverage multiple generations, which will inevitably suffer from a full GC phase scanning and compacting the whole heap. This induces a pause tens of times longer than normal collections, which largely affects both throughput and latency of applications. Mingyu Wu 0001, Binyu Zang, Haibo Chen 0001 |
VEE | 3 |
| 2019 | TEEv: virtualizing trusted execution environments on mobile platformsabstractTrusted Execution Environments (TEE) are widely deployed, especially on smartphones. A recent trend in TEE development is the transition from vendor-controlled, single-purpose TEEs to open TEEs that host Trusted Applications (TAs) from multiple sources with independent tasks. This transition is expected to create a TA ecosystem needed for providing stronger and customized security to apps and OS running in the Rich Execution Environment (REE). However, the transition also poses two security challenges: enlarged attack surface resulted from the increased complexity of TAs and TEEs; the lack of trust (or isolation) among TAs and the TEE. Wenhao Li 0009, Yubin Xia, Long Lu, Haibo Chen 0001, Binyu Zang |
VEE | 5 |
| 2019 | Scaling out NUMA-Aware Applications with RDMA-Based Distributed Shared Memory
Yang Hong 0007, Fan Yang 0024, Binyu Zang, Haibing Guan, Haibo Chen 0001 |
J. Comput. Sci. Technol. | 4 |
| 2018 | Espresso: Brewing Java For More Non-Volatility with Non-volatile MemoryabstractFast, byte-addressable non-volatile memory (NVM) embraces both near-DRAM latency and disk-like persistence, which has generated considerable interests to revolutionize system software stack and programming models. However, it is less understood how NVM can be combined with managed runtime like Java virtual machine (JVM) to ease persistence management. This paper proposes Espresso, a holistic extension to Java and its runtime, to enable Java programmers to exploit NVM for persistence management with high performance. Espresso first provides a general persistent heap design called Persistent Java Heap (PJH) to manage persistent data as normal Java objects. The heap is then strengthened with a recoverable mechanism to provide crash consistency for heap metadata. Espresso further provides a new abstraction called Persistent Java Object (PJO) to provide an easy-to-use but safe persistence programming model for programmers to persist application data. Evaluation confirms that Espresso significantly outperforms state-of-art NVM support for Java (i.e., JPA and PCJ) while being compatible to data structures in existing Java programs. Mingyu Wu 0001, Ziming Zhao 0003, Heting Li, Haibo Chen 0001, Binyu Zang, Haibing Guan |
ASPLOS | 6 |
| 2018 | Comprehensive VM Protection Against Untrusted Hypervisor Through Retrofitted AMD Memory EncryptionabstractThe confidentiality of tenant's data is confronted with high risk when facing hardware attacks and privileged malicious software. Hardware-based memory encryption is one of the promising means to provide strong guarantees of data security. Recently AMD has proposed its new memory encryption hardware called SME and SEV, which can selectively encrypt memory regions in a fine-grained manner, e.g., by setting the C-bits in the page table entries. More importantly, SEV further supports encrypted virtual machines. This, intuitively, has provided a new opportunity to protect data confidentiality in guest VMs against an untrusted hypervisor in the cloud environment. In this paper, we first provide a security analysis on the (in)security of SEV and uncover a set of security issues of using SEV as a means to defend against an untrusted hypervisor. Based on the study, we then propose a software-based extension to the SEV feature, namely Fidelius, to address those issues while retaining performance efficiency. Fidelius separates the management of critical resources from service provisioning and revokes the permissions of accessing specific resources from the un-trusted hypervisor. By adopting a sibling-based protection mechanism with non-bypassable memory isolation, Fidelius embraces both security and efficiency, as it introduces no new layer of abstraction. Meanwhile, Fidelius reuses the SEV API to provide a full VM life-cycle protection, including two sets of para-virtualized I/O interfaces to encode the I/O data, which is not considered in the SEV hardware design. A detailed and quantitative security analysis shows its effectiveness in protecting tenant's data from a variety of attack surfaces, and the performance evaluation confirms the performance efficiency of Fidelius. Yuming Wu, Ruifeng Liu, Haibo Chen 0001, Binyu Zang, Haibing Guan |
HPCA | 5 |
| 2018 | VButton: Practical Attestation of User-driven Operations in Mobile AppsabstractMore and more malicious apps and mobile rootkits are found to perform sensitive operations on behalf of legitimate users without their awareness. Malware does so by either forging user inputs or tricking users into making unintended requests to online service providers. Such malware is hard to detect and generates large revenues for cybercriminals, which is often used for committing ad/click frauds, faking reviews/ratings, promoting people or business on social networks, etc. Wenhao Li 0009, Shiyu Luo, Zhichuang Sun, Yubin Xia, Long Lu, Haibo Chen 0001, Binyu Zang, Haibing Guan |
MobiSys | 7 |
| 2018 | EPTI: Efficient Defence against Meltdown Attack for Unpatched VMs
Zhichao Hua 0001, Dong Du 0003, Yubin Xia, Haibo Chen 0001, Binyu Zang |
USENIX ATC | 5 |
| 2018 | SplitPass: A Mutually Distrusting Two-Party Password Manager
Dong Du 0003, Yubin Xia, Haibo Chen 0001, Binyu Zang, Zhenkai Liang |
J. Comput. Sci. Technol. | 5 |
| 2018 | ShadowEth: Private Smart Contract on Public Blockchain
Yubin Xia, Haibo Chen 0001, Binyu Zang, Jan Xie |
J. Comput. Sci. Technol. | 4 |
| 2018 | Replication-Based Fault-Tolerance for Large-Scale Graph ProcessingabstractThe increasing algorithmic complexity and dataset sizes necessitate the use of networked machines for many graph-parallel algorithms, which also makes fault tolerance a must due to the increasing scale of machines. Unfortunately, existing large-scale graph-parallel systems usually adopt a distributed checkpoint mechanism for fault tolerance, which incurs not only notable performance overhead but also lengthy recovery time. This paper observes that the vertex replicas created for distributed graph computation can be naturally extended for fast in-memory recovery of graph states. This paper describes Imitator, a new fault tolerance mechanism, which supports cheap maintenance of vertex states by replicating them to their replicas during normal message exchanges, and provides fast in-memory reconstruction of failed vertices from replicas in other machines. Imitator has been implemented on Cyclops with edge-cut and PowerLyra with vertex-cut. Evaluation on a 50-node EC-2 like cluster shows that Imitator incurs an average of 1.37 and 2.32 percent performance overhead (ranging from -0.6 to 3.7 percent) for Cyclops and PowerLyra respectively, and can recover from failures of more than one million of vertices with less than 3.4 seconds. Rong Chen 0001, Youyang Yao, Kaiyuan Zhang 0005, Haibing Guan, Binyu Zang, Haibo Chen 0001 |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2017 | Secure Live Migration of SGX Enclaves on Untrusted CloudabstractThe recent commercial availability of Intel SGX (Software Guard eXtensions) provides a hardware-enabled building block for secure execution of software modules in an untrusted cloud. As an untrusted hypervisor/OS has no access to an enclave's running states, a VM (virtual machine) with enclaves running inside loses the capability of live migration, a key feature of VMs in the cloud. This paper presents the first study on the support for live migration of SGX-capable VMs. We identify the security properties that a secure enclave migration process should meet and propose a software-based solution. We leverage several techniques such as two-phase checkpointing and self-destroy to implement our design on a real SGX machine. Security analysis confirms the security of our proposed design and performance evaluation shows that it incurs negligible performance overhead. Besides, we give suggestions on the future hardware design for supporting transparent enclave migration. Jinyu Gu 0001, Zhichao Hua 0001, Yubin Xia, Haibo Chen 0001, Binyu Zang, Haibing Guan |
DSN | 5 |
| 2017 | Transparent and Efficient CFI Enforcement with Intel Processor TraceabstractCurrent control flow integrity (CFI) enforcement approaches either require instrumenting application executables and even shared libraries, or are unable to defend against sophisticated attacks due to relaxed security policies, or both; many of them also incur high runtime overhead. This paper observes that the main obstacle of providing transparent and strong defense against sophisticated adversaries is the lack of sufficient runtime control flow information. To this end, this paper describes FlowGuard, a lightweight, transparent CFI enforcement approach by a novel reuse of Intel Processor Trace (IPT), a recent hardware feature that efficiently captures the entire runtime control flow. The main challenge is that IPT is designed for offline performance analysis and software debugging such that decoding collected control flow traces is prohibitively slow on the fly. FlowGuard addresses this challenge by reconstructing applications' conservative control flow graphs (CFG) to be compatible with the compressed encoding format of IPT, and labeling the CFG edges with credits in the help of fuzzing-like dynamic training. At runtime, FlowGuard separates fast and slow paths such that the fast path compares the labeled CFGs with the IPT traces for fast filtering, while the slow path decodes necessary IPT traces for strong security. We have implemented and evaluated FlowGuard on a commodity Intel Skylake machine with IPT support. Evaluation results show that FlowGuard is effective in enforcing CFI for several applications, while introducing only small performance overhead. We also show that, with minor hardware extensions, the performance overhead can be further reduced. Peitao Shi, Haibo Chen 0001, Binyu Zang, Haibing Guan |
HPCA | 5 |
| 2017 | Deconstructing Xen
Le Shi, Yuming Wu, Yubin Xia, Nathan Dautenhahn, Haibo Chen 0001, Binyu Zang |
NDSS | 6 |
| 2017 | POSTER: Recovering Performance for Vector-based Machine Learning on Managed RuntimeabstractNo abstract available. Mingyu Wu 0001, Haibing Guan, Binyu Zang, Haibo Chen 0001 |
PPoPP | 3 |
| 2017 | vTZ: Virtualizing ARM TrustZone
Zhichao Hua 0001, Jinyu Gu 0001, Yubin Xia, Haibo Chen 0001, Binyu Zang, Haibing Guan |
USENIX Security Symposium | 5 |
| 2017 | Characterizing and optimizing Java-based HPC applications on Intel many-core architecture
Yang Yu 0002, Tianyang Lei, Haibo Chen 0001, Binyu Zang |
Sci. China Inf. Sci. | 4 |
| 2017 | Secure Outsourcing of Virtual ApplianceabstractComputation outsourcing using virtual appliance is getting prevalent in cloud computing. However, with both hardware and software being controlled by potentially curious or even malicious cloud operators, it is no surprise to see frequent reports of security accidents, like data leakages or abuses. This paper proposes Kite, a hardware-software framework that guards the security of tenant's virtual machine (VM), in which the outsourced computation is encapsulated. Kite only trusts the processor and makes no security assumption on external memory, devices, or hypervisor. Unlike prior hardware-based approaches, Kite retains transparency with existing VM and requires few changes to the (untrusted) hypervisor by introducing VM-Shim mechanism. Each VM-Shim instance runs in between its VM and the hypervisor, which only transfers necessary information designated by the VM to the hypervisor and external environments. Kite also considers the high-level semantic of interaction between VM and hypervisor to defend against attacks through legitimate operations or interfaces. We have implemented a prototype of Kite's secure processor in a QEMU-based full-system emulator and its software components on real machine. Evaluation shows that the performance overhead of Kite ranges from 0.5-14.0 percent on simulated platform and 0.4-7.3 percent on real hardware. Yubin Xia, Haibing Guan, Yunji Chen, Tianshi Chen 0002, Binyu Zang, Haibo Chen 0001 |
IEEE Trans. Cloud Comput. | 6 |
| 2017 | Fast In-Memory Transaction Processing Using RDMA and HTMabstractDrTM is a fast in-memory transaction processing system that exploits advanced hardware features such as remote direct memory access (RDMA) and hardware transactional memory (HTM). To achieve high efficiency, it mostly offloads concurrency control such as tracking read/write accesses and conflict detection into HTM in a local machine and leverages the strong consistency between RDMA and HTM to ensure serializability among concurrent transactions across machines. To mitigate the high probability of HTM aborts for large transactions, we design and implement an optimized transaction chopping algorithm to decompose a set of large transactions into smaller pieces such that HTM is only required to protect each piece. We further build an efficient hash table for DrTM by leveraging HTM and RDMA to simplify the design and notably improve the performance. We describe how DrTM supports common database features like read-only transactions and logging for durability. Evaluation using typical OLTP workloads including TPC-C and SmallBank shows that DrTM has better single-node efficiency and scales well on a six-node cluster; it achieves greater than 1.51, 34 and 5.24, 138 million transactions per second for TPC-C and SmallBank on a single node and the cluster, respectively. Such numbers outperform a state-of-the-art single-node system (i.e., Silo) and a distributed transaction system (i.e., Calvin) by at least 1.9X and 29.6X for TPC-C. Haibo Chen 0001, Rong Chen 0001, Xingda Wei, Jiaxin Shi, Yanzhe Chen, Binyu Zang, Haibing Guan |
ACM Trans. Comput. Syst. | 7 |
| 2017 | Efficient and Available In-Memory KV-Store with Hybrid Erasure Coding and ReplicationabstractIn-memory key/value store (KV-store) is a key building block for many systems like databases and large websites. Two key requirements for such systems are efficiency and availability, which demand a KV-store to continuously handle millions of requests per second. A common approach to availability is using replication, such as primary-backup (PBR), which, however, requires M +1 times memory to tolerate M failures. This renders scarce memory unable to handle useful user jobs. This article makes the first case of building highly available in-memory KV-store by integrating erasure coding to achieve memory efficiency, while not notably degrading performance. A main challenge is that an in-memory KV-store has much scattered metadata. A single KV put may cause excessive coding operations and parity updates due to excessive small updates to metadata. Our approach, namely Cocytus, addresses this challenge by using a hybrid scheme that leverages PBR for small-sized and scattered data (e.g., metadata and key), while only applying erasure coding to relatively large data (e.g., value). To mitigate well-known issues like lengthy recovery of erasure coding, Cocytus uses an online recovery scheme by leveraging the replicated metadata information to continuously serve KV requests. To further demonstrate the usefulness of Cocytus, we have built a transaction layer by using Cocytus as a fast and reliable storage layer to store database records and transaction logs. We have integrated the design of Cocytus to Memcached and extend it to support in-memory transactions. Evaluation using YCSB with different KV configurations shows that Cocytus incurs low overhead for latency and throughput, can tolerate node failures with fast online recovery, while saving 33% to 46% memory compared to PBR when tolerating two failures. A further evaluation using the SmallBank OLTP benchmark shows that in-memory transactions can run atop Cocytus with high throughput, low latency, and low abort rate and recover fast from consecutive failures. Haibo Chen 0001, Mingkai Dong 0002, Yubin Xia, Haibing Guan, Binyu Zang |
ACM Trans. Storage | 7 |
| 2017 | Fence-Free Synchronization with Dynamically Serialized Synchronization VariablesabstractMemory fences are widely used to ensure the correctness for synchronization constructs on machines with relaxed consistency models. However, they are expensive and usually impose over-constrained ordering that causes unnecessary CPU stalls. In this paper, we observe that memory fences in TSO are merely intended to order synchronization variables. Based on this observation, we rethink the hardware-software interface of synchronization constructs on multicore processors and propose a new design called Sync-Order that differentiates synchronization variables (sync-vars) from normal ones. Sync-Order reduces hardware complexity such that the processor only needs to serialize the ordering among sync-vars. Its simplicity makes it easy to be integrated to the directory controller and it supports distributed directory, a missing feature in prior designs. We show that Sync-Order eliminates traditional fences on all sides of synchronization constructs (instead of only one side in prior work) and requires small effort for a programmer or compiler to annotate sync-vars. Our experimental results show that Sync-Order significantly reduces CPU stalls and boosts the performance of a set of synchronization constructs and concurrent data structures by 10 percent; meanwhile, the fence overhead of full applications from SPLASH-2 and PARSEC is reduced from 42 to 3 percent. Yang Hong 0007, Haibing Guan, Binyu Zang, Haibo Chen 0001 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2016 | A Case for Virtualizing Persistent MemoryabstractWith the proliferation of software and hardware support for persistent memory (PM) like PCM and NV-DIMM, we envision that PM will soon become a standard component of commodity cloud, especially for those applications demanding high performance and low latency. Yet, current virtualization software lacks support to efficiently virtualize and manage PM to improve cost-effectiveness, performance, and endurance. Rong Chen 0001, Haibo Chen 0001, Yubin Xia, KwanJong Park, Binyu Zang, Haibing Guan |
SoCC | 6 |
| 2016 | Performance Analysis and Optimization of Full Garbage Collection in Memory-hungry EnvironmentsabstractGarbage collection (GC), especially full GC, would non- trivially impact overall application performance, especially for those memory-hungry ones handling large data sets. This paper presents an in-depth performance analysis on the full GC performance of Parallel Scavenge (PS), a state-of-the-art and the default garbage collector in the HotSpot JVM, using traditional and big-data applications running atop JVM on CPU (e.g., Intel Xeon) and many-integrated cores (e.g., Intel Xeon i). The analysis uncovers that unnecessary memory accesses and calculations during reference updating in the compaction ase are the main causes of lengthy full GC. To this end, this paper describes an incremental query model for reference calculation, which is further embodied with three schemes (namely optimistic, sort-based and region-based) for different query patterns. Performance evaluation shows that the incremental query model leads to averagely 1.9X (up to 2.9X) in full GC and 19.3% (up to 57.2%) improvement in application throughput, as well as 31.2% reduction in pause time over the vanilla PS collector on CPU, and the numbers are 2.1X (up to 3.4X), 11.1% (up to 41.2%) and 34.9% for Xeon i accordingly. Yang Yu 0002, Tianyang Lei, Haibo Chen 0001, Binyu Zang |
VEE | 5 |
| 2016 | Fast Consensus Using Bounded Staleness for Scalable Read-Mostly SynchronizationabstractReader-mostly synchronization schemes, such as rwlocks and RCU, aim to maximize parallelism among readers, but many existing designs either cause readers to contend, or significantly extend writer latency, or both. This paper attributes such a problem to the lack of a fast consensus protocol between readers and writers, by which the two parts cooperate to obey the semantics of a synchronization construct. This paper describes FCP, a fast consensus protocol among readers and writers that provides scalable read-side performance as well as small writer latency for TSO architectures. The heart of FCP is a version-based consensus protocol between multiple non-communicating readers and a pending writer. FCP leverages bounded staleness of memory consistency to avoid atomic instructions and memory barriers in readers' common paths, and uses message-passing (e.g., IPI) for straggling readers so that the writer latency can be bounded. To demonstrate the effectiveness of FCP, this paper applies FCP to construct a scalable reader-writers lock (rwlock) and a scalable RCU implementation. Evaluation on a 64-core machine shows that FCP significantly boosts the performance of the Linux virtual memory subsystem, a concurrent hashtable and an in-memory database. Micro-benchmarks show that FCP achieves smaller reader-side latency and lower writer-side latency when compared to state-of-the-art rwlocks and RCU implementation. Haibo Chen 0001, Ran Liu 0003, Binyu Zang, Haibing Guan |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2015 | TinMan: eliminating confidential mobile data exposure with security oriented offloadingabstractThe wide adoption of smart devices has stimulated a fast shift of security-critical data from desktop to mobile devices. However, recurrent device theft and loss expose mobile devices to various security threats and even physical attacks. This paper presents TinMan, a system that protects confidential data such as web site password and credit card number (we use the term cor to represent these data, which is short for Confidential Record) from being leaked or abused even under device theft. TinMan separates accesses of cor from the rest of the functionalities of an app, by introducing a trusted node to store cor and offloading any code from a mobile device to the trusted node to access cor. This completely eliminates the exposure of cor on the mobile devices. The key challenges to TinMan include deciding when and how to efficiently and transparently offload execution; TinMan addresses these challenges with security-oriented offloading with a low-overhead tainting scheme called asymmetric tainting to track accesses to cor to trigger offloading, as well as transparent SSL session injection and TCP pay-load replacement to offload accesses to cor. We have implemented a prototype of TinMan based on Android and demonstrated how TinMan protects the information of user's bank account and credit card number without modifying the apps. Evaluation results also show that TinMan incurs only a small amount of performance and power overhead. Yubin Xia, Cheng Tan 0005, Haibing Guan, Binyu Zang, Haibo Chen 0001 |
EuroSys | 6 |
| 2015 | Reducing world switches in virtualized environment with flexible cross-world callsabstractModern computers are built with increasingly complex software stack crossing multiple layers (i.e., worlds), where cross-world call has been a necessity for various important purposes like security, reliability, and reduced complexity. Unfortunately, there is currently limited cross-world call support (e.g., syscall, vmcall), and thus other calls need to be emulated by detouring multiple times to the privileged software layer (i.e., OS kernel and hypervisor). This causes not only significant performance degradation, but also unnecessary implementation complexity. Wenhao Li 0009, Yubin Xia, Haibo Chen 0001, Binyu Zang, Haibing Guan |
ISCA | 4 |
| 2015 | SYNC or ASYNC: time to fuse for distributed graph-parallel computationabstractLarge-scale graph-structured computation usually exhibits iterative and convergence-oriented computing nature, where input data is computed iteratively until a convergence condition is reached. Such features have led to the development of two different computation modes for graph-structured programs, namely synchronous (Sync) and asynchronous (Async) modes. Unfortunately, there is currently no in-depth study on their execution properties and thus programmers have to manually choose a mode, either requiring a deep understanding of underlying graph engines, or suffering from suboptimal performance. This paper makes the first comprehensive characterization on the performance of the two modes on a set of typical graph-parallel applications. Our study shows that the performance of the two modes varies significantly with different graph algorithms, partitioning methods, execution stages, input graphs and cluster scales, and no single mode consistently outperforms the other. To this end, this paper proposes Hsync, a hybrid graph computation mode that adaptively switches a graph-parallel program between the two modes for optimal performance. Hsync constantly collects execution statistics on-the-fly and leverages a set of heuristics to predict future performance and determine when a mode switch could be profitable. We have built online sampling and offline profiling approaches combined with a set of heuristics to accurately predicting future performance in the two modes. A prototype called PowerSwitch has been built based on PowerGraph, a state-of-the-art distributed graph-parallel system, to support adaptive execution of graph algorithms. On a 48-node EC2-like cluster, PowerSwitch consistently outperforms the best of both modes, with a speedup ranging from 9% to 73% due to timely switch between two modes. Chenning Xie, Rong Chen 0001, Haibing Guan, Binyu Zang, Haibo Chen 0001 |
PPoPP | 4 |
| 2015 | You Shouldn't Collect My Secrets: Thwarting Sensitive Keystroke Leakage in Mobile IME Apps
Haibo Chen 0001, Erick Bauman, Zhiqiang Lin 0001, Binyu Zang, Haibing Guan |
USENIX Security Symposium | 5 |
| 2015 | Bipartite-Oriented Distributed Graph Partitioning for Big Learning
Rong Chen 0001, Jiaxin Shi, Haibo Chen 0001, Binyu Zang |
J. Comput. Sci. Technol. | 4 |
| 2014 | Greedy map generalization by iterative point removalabstractThis paper describes a map generalization program we submitted to the ACM SIGSPATIAL Cup 2014. In this competition, the goal is to remove as many points in a set of polygonal lines as quickly as possible with respect to two constraints. The topological relationships among the lines must not change, and the relationships between a set of control points and the lines must not change. Inspired by Visvalingam-Whyatt Algorithm, we iteratively examine successive triplets along each line, and remove the middle point if no control point or point of other lines is in the associated triangle. Based on the features of the training datasets, we further introduce many optimization techniques to speed up the computation. Yanzhe Chen, Yin Wang 0001, Rong Chen 0001, Haibo Chen 0001, Binyu Zang |
SIGSPATIAL/GIS | 5 |
| 2014 | Concurrent and consistent virtual machine introspection with hardware transactional memoryabstractVirtual machine introspection, which provides tamperresistant, high-fidelity “out of the box” monitoring of virtual machines, has many prominent security applications including VM-based intrusion detection, malware analysis and memory forensic analysis. However, prior approaches are either intrusive in stopping the world to avoid race conditions between introspection tools and the guest VM, or providing no guarantee of getting a consistent state of the guest VM. Further, there is currently no effective means for timely examining the VM states in question. In this paper, we propose a novel approach, called TxIntro, which retrofits hardware transactional memory (HTM) for concurrent, timely and consistent introspection of guest VMs. Specifically, TxIntro leverages the strong atomicity of HTM to actively monitor updates to critical kernel data structures. Then TxIntro can mount introspection to timely detect malicious tampering. To avoid fetching inconsistent kernel states for introspection, TxIntro uses HTM to add related synchronization states into the read set of the monitoring core and thus can easily detect potential inflight concurrent kernel updates. We have implemented and evaluated TxIntro based on Xen VMM on a commodity Intel Haswell machine that provides restricted transactional memory (RTM) support. To demonstrate the effectiveness of TxIntro, we implemented a set of kernel rootkit detectors using TxIntro. Evaluation results show that TxIntro is effective in detecting these rootkits, and is efficient in adding negligible performance overhead. Yubin Xia, Haibing Guan, Binyu Zang, Haibo Chen 0001 |
HPCA | 4 |
| 2014 | Computation and communication efficient graph processing with distributed immutable viewabstractCyclops is a new vertex-oriented graph-parallel framework for writing distributed graph analytics. Unlike existing distributed graph computation models, Cyclops retains simplicity and computation-efficiency by synchronously computing over a distributed immutable view, which grants a vertex with read-only access to all its neighboring vertices. The view is provided via read- only replication of vertices for edges spanning machines during a graph cut. Cyclops follows a centralized computation model by assigning a master vertex to update and propagate the value to its replicas unidirectionally in each iteration, which can significantly reduce messages and avoid contention on replicas. Being aware of the pervasively available multicore-based clusters, Cyclops is further extended with a hierarchical processing model, which aggregates messages and replicas in a single multicore machine and transparently decomposes each worker into multiple threads on-demand for different stages of computation. We have implemented Cyclops based on an open-source Pregel clone called Hama. Our evaluation using a set of graph algorithms on an in-house multicore cluster shows that Cyclops outperforms Hama from 2.06X to 8.69X and 5.95X to 23.04X using hash-based and Metis partition algorithms accordingly, due to the elimination of contention on messages and hierarchical optimization for the multicore-based clusters. Cyclops (written in Java) also has comparable performance with PowerGraph (written in C++) despite the language difference, due to the significantly lower number of messages and avoided contention. Rong Chen 0001, Haibo Chen 0001, Binyu Zang, Haibing Guan |
HPDC | 5 |
| 2014 | Hydra: Efficient Detection of Multiple Concurrency Bugs on Fused CPU-GPU ArchitectureabstractDetecting concurrency bugs, such as data race, atomicity violation and order violation, is a cumbersome task for programmers. This situation is further being exacerbated due to the increasing number of cores in a single machine and the prevalence of threaded programming models. Unfortunately, many existing software-based approaches usually incur high runtime overhead or accuracy loss, while most hardware-based proposals usually focus on a specific type of bugs and thus are inflexible to detect a variety of concurrency bugs. In this paper, we propose Hydra, an approach that leverages massive parallelism and programmability of fused GPU architecture to simultaneously detect multiple types of concurrency bugs, including data race, atomicity violation and order violation. Hydra instruments and collects program behavior on CPU and transfers the traces to GPU for bug detection through on-chip interconnect. Furthermore, to achieve high speed, Hydra exploits bloom filter to filter out unnecessary detection traces. Hydra incurs small hardware complexity and requires no changes to internal critical-path processor components such as cache and its coherence protocol, and is with about 1.1% hardware overhead under a 32-core configuration. Experimental results show that Hydra only introduces about 0.35% overhead on average for detecting one type of bugs and 0.92% overhead for simultaneously detecting multiple bugs, yet with the similar detectability of a heavyweight software bug detector (e.g., Helgrind). Zhuofang Dai, Haojun Wang, Haibo Chen 0001, Binyu Zang |
ICPP | 5 |
| 2014 | X10-FT: Transparent fault tolerance for APGAS language and runtime
Zhijun Hao, Chenning Xie, Haibo Chen 0001, Binyu Zang |
Parallel Comput. | 4 |
| 2014 | Measuring Microarchitectural Details of Multi- and Many-Core Memory Systems through MicrobenchmarkingabstractAs multicore and many-core architectures evolve, their memory systems are becoming increasingly more complex. To bridge the latency and bandwidth gap between the processor and memory, they often use a mix of multilevel private/shared caches that are either blocking or nonblocking and are connected by high-speed network-on-chip. Moreover, they also incorporate hardware and software prefetching and simultaneous multithreading (SMT) to hide memory latency. On such multi- and many-core systems, to incorporate various memory optimization schemes using compiler optimizations and performance tuning techniques, it is crucial to have microarchitectural details of the target memory system. Unfortunately, such details are often unavailable from vendors, especially for newly released processors. In this article, we propose a novel microbenchmarking methodology based on short elapsed-time events (SETEs) to obtain comprehensive memory microarchitectural details in multi- and many-core processors. This approach requires detailed analysis of potential interfering factors that could affect the intended behavior of such memory systems. We lay out effective guidelines to control and mitigate those interfering factors. Taking the impact of SMT into consideration, our proposed methodology not only can measure traditional cache/memory latency and off-chip bandwidth but also can uncover the details of software and hardware prefetching units not attempted in previous studies. Using the newly released Intel Xeon Phi many-core processor (with in-order cores) as an example, we show how we can use a set of microbenchmarks to determine various microarchitectural features of its memory system (many are undocumented from vendors). To demonstrate the portability and validate the correctness of such a methodology, we use the well-documented Intel Sandy Bridge multicore processor (with out-of-order cores) as another example, where most data are available and can be validated. Moreover, to illustrate the usefulness of the measured data, we do a multistage coordinated data prefetching case study on both Xeon Phi and Sandy Bridge and show that by using the measured data, we can achieve 1.3X and 1.08X performance speedup, respectively, compared to the state-of-the-art Intel ICC compiler. We believe that these measurements also provide useful insights into memory optimization, analysis, and modeling of such multicore and many-core architectures. Zhenman Fang, Sanyam Mehta, Pen-Chung Yew, Antonia Zhai, James B. S. G. Greensky, Gautham Beeraka, Binyu Zang |
ACM Trans. Archit. Code Optim. | 7 |
| 2014 | Permission Use Analysis for Vetting Undesirable Behaviors in Android AppsabstractThe android platform adopts permissions to protect sensitive resources from untrusted apps. However, after permissions are granted by users at install time, apps could use these permissions (sensitive resources) with no further restrictions. Thus, recent years have witnessed the explosion of undesirable behaviors in Android apps. An important part in the defense is the accurate analysis of Android apps. However, traditional syscall-based analysis techniques are not well-suited for Android, because they could not capture critical interactions between the application and the Android system. This paper presents VetDroid, a dynamic analysis platform for generally analyzing sensitive behaviors in Android apps from a novel permission use perspective. VetDroid proposes a systematic permission use analysis technique to effectively construct permission use behaviors, i.e., how applications use permissions to access (sensitive) system resources, and how these acquired permission-sensitive resources are further utilized by the application. With permission use behaviors, security analysts can easily examine the internal sensitive behaviors of an app. Using real-world Android malware, we show that VetDroid can clearly reconstruct fine-grained malicious behaviors to ease malware analysis. We further apply VetDroid to 1249 top free apps in Google Play. VetDroid can assist in finding more information leaks than TaintDroid, a state-of-the-art technique. In addition, we show how we can use VetDroid to analyze fine-grained causes of information leaks that TaintDroid cannot reveal. Finally, we show that VetDroid can help to identify subtle vulnerabilities in some (top free) applications otherwise hard to detect. Yuan Zhang 0009, Min Yang 0002, Zhemin Yang, Guofei Gu, Peng Ning, Binyu Zang |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2013 | Vetting undesirable behaviors in android apps with permission use analysisabstractAndroid platform adopts permissions to protect sensitive resources from untrusted apps. However, after permissions are granted by users at install time, apps could use these permissions (sensitive resources) with no further restrictions. Thus, recent years have witnessed the explosion of undesirable behaviors in Android apps. An important part in the defense is the accurate analysis of Android apps. However, traditional syscall-based analysis techniques are not well-suited for Android, because they could not capture critical interactions between the application and the Android system. Yuan Zhang 0009, Min Yang 0002, Bingquan Xu, Zhemin Yang, Guofei Gu, Peng Ning, Xiaoyang Sean Wang, Binyu Zang |
CCS | 8 |
| 2013 | Multi-level phase analysis for sampling simulationabstractExtremely long simulation time of architectural simulators has been a major impediment to their wide applicability. To accelerate architectural simulation, prior researchers have proposed representative sampling simulation to trade small loss of accuracy for notable speed improvement. Generally, they use fine-grained phase analysis to select only a small representative portion of program execution intervals for detailed cycle-accurate simulation, while functionally simulating the remaining portion. However, though phase granularity is one of the most important factors to simulation speed, it has not been well investigated and most prior researches explore a fine-grained scheme. This limits their effectiveness in further improving simulation speed with the requirement of increasingly complex architectural designs and new lengthy benchmarks. In this paper, by analyzing the impact of phase granularity on simulation speed, we observe that coarse-grained phases can better capture the overall program characteristics with a less number of phases and the last representative phase could be classified in a very early program position, leading to fewer execution internals being functionally simulated. By contrast, fine-grained phases usually have much shorter execution intervals and thus the overall detailed simulation time could be reduced. Based on the above observation, we design a multi-level sampling simulation technique that combines both fine-grained and coarse-grained phase analysis for sampling simulation. Such a scheme uses fine-grained simulation points to represent only the selected coarse-grained simulation points instead of the entire program execution, thus it could further reduce both the functional and detailed simulation time. Experimental results using SPEC2000 show such a framework is effective: using the SimPoint method as baseline, it can reduce about 90% functional simulation time and about 50% detailed simulation time. It finally achieves a geometric average speedup of 14.04X over SimPoint with comparable accuracy. Haibo Chen 0001, Binyu Zang |
DATE | 4 |
| 2013 | Understanding architectural characteristics of multimedia retrieval workloadsabstractNo abstract available. Chen Dai, Binyu Zang |
SIGMETRICS | 5 |
| 2012 | Transformer: a functional-driven cycle-accurate multicore simulatorabstractFull-system simulators are extremely useful in evaluating design alternatives for multicore. However, state-of-the-art multicore simulators either lack good extensibility due to their tightly-coupled design between functional model (FM) and timing model (TM), or cannot guarantee cycle-accuracy. This paper conducts a comprehensive study on factors affecting cycle-accuracy and uncovers several contributing factors ignored before. Based on the study, we propose a loosely-coupled functional-driven full-system simulator for multicore, namely Transformer. To ensure extensibility and cycle-accuracy, Transformer leverages an architecture-independent interface between FM and TM and uses a lightweight scheme to detect and recover from execution divergence between FM and TM. Based on Transformer, a graduate student only needs to write about 180 lines of code and takes about two months to extend an X86 functional model (QEMU) in Transformer. Moreover, the loosely-coupled design also removes the complex interaction between FM and TM and opens the opportunity to parallelize FM and TM to improve performance. Experimental results show that Transformer achieves an average of 8.4% speedup over GEMS while guaranteeing the cycle-accuracy. A further parallelization between FM and TM leads to 35.3% speedup. Zhenman Fang, Qinghao Min, Keyong Zhou, Yibin Hu, Haibo Chen 0001, Jian Li 0059, Binyu Zang |
DAC | 9 |
| 2012 | CFIMon: Detecting violation of control flow integrity using performance countersabstractMany classic and emerging security attacks usually introduce illegal control flow to victim programs. This paper proposes an approach to detecting violation of control flow integrity based on hardware support for performance monitoring in modern processors. The key observation is that the abnormal control flow in security breaches can be precisely captured by performance monitoring units. Based on this observation, we design and implement a system called CFIMon, which is the first non-intrusive system that can detect and reason about a variety of attacks violating control flow integrity without any changes to applications (either source or binary code) or requiring special-purpose hardware. CFIMon combines static analysis and runtime training to collect legal control flow transfers, and leverages the branch tracing store mechanism in commodity processors to collect and analyze runtime traces on-the-fly to detect violation of control flow integrity. Security evaluation shows that CFIMon has low false positives or false negatives when detecting several realistic security attacks. Performance results show that CFIMon incurs only 6.1% performance overhead on average for a set of typical server applications. Yubin Xia, Haibo Chen 0001, Binyu Zang |
DSN | 4 |
| 2012 | Adaptive Pipeline Parallelism for Image Feature Extraction AlgorithmsabstractCurrently, multimedia data has become one of the major data types processed and transferred on the Internet. With the rapid growth of multimedia data, it is vitally important to find an efficient way to extract useful information from a large amount of data. SIFT and SURF, as the most popular multimedia feature extraction algorithms, have been widely used in many applications. However, the limited processing speed~(about 1.8 and 2.6 images or frames per second for SIFT and SURF respectively on an ordinary CPU) makes it impossible to apply them in many real-world applications with real-time requirements. Therefore, it has become one of the major challenges that how to improve the processing speed of these multimedia feature extraction algorithms. The popularity of multi-core architecture and the increase of computation resources on different platforms provide a new opportunity to accelerate the processing speed of these image feature extraction algorithms. In this paper, we first systematically analyze the major parallel constraints in SIFT and SURF, such as imbalanced workload and indeterminate time distribution. Then, based on these analysis, we design and implement an adaptive pipeline parallel scheme (AD-PIPE) for both SIFT and SURF to alleviate these limitations. In our scheme, we dynamically check the workload in different pipeline stages and adjust the thread number in different stages to achieve a balanced partition. Experimental results show that our approach is efficient and scalable. It can achieve a speedup of 16.88X and 20.33X respectively for SIFT and SURF on a 16-core machine and a real time processing speed with about 27 and 52 images or frames per second. Donglei Yang, Binyu Zang, Haibo Chen 0001 |
ICPP | 5 |
| 2012 | Improving dynamic prediction accuracy through multi-level phase analysisabstractPhase analysis, which classifies the set of execution intervals with similar execution behavior and resource requirements, has been widely used in a variety of dynamic systems, including dynamic cache reconfiguration, prefetching and race detection. While phase granularity has been a major factor to the accuracy of phase prediction, it has not been well investigated yet and most dynamic systems usually adopt a fine-grained prediction scheme. However, such a scheme can only take account of recent local phase information and could be frequently interfered by temporary noises due to instant phase changes, which might notably limit the prediction accuracy. Zhenman Fang, Haibo Chen 0001, Binyu Zang |
LCTES | 6 |
| 2012 | Revisiting Software Zero-Copy for Web-caching Applications with Twin Memory Allocation
Jicheng Shi, Haibo Chen 0001, Binyu Zang |
USENIX ATC | 4 |
| 2012 | Swift: a register-based JIT compiler for embedded JVMsabstractCode quality and compilation speed are two challenges to JIT compilers, while selective compilation is commonly used to trade-off these two issues. Meanwhile, with more and more Java applications running in mobile devices, selective compilation meets many problems. Since these applications always have flat execution profile and short live time, a lightweight JIT technique without losing code quality is extremely needed. However, the overhead of compiling stack-based Java bytecode to heterogeneous register-based machine code is significant in embedded devices. This paper presents a fast and effective JIT technique for mobile devices, building on a register-based Java bytecode format which is more similar to the underlying machine architecture. Through a comprehensive study on the characteristics of Java applications, we observe that virtual registers used by more than 90% Java methods can be directly fulfilled by 11 physical registers. Based on this observation, this paper proposes Swift, a novel JIT compiler on register-based bytecode, which generates native code for RISC machines. After mapping virtual registers to physical registers, the code is generated efficiently by looking up a translation table. And the code quality is guaranteed by the static compiler which is used to generate register-based bytecode. Besides, we design two lightweight optimizations and an efficient code unloader to make Swift more suitable for embedded environment. As the prevalence of Android, a prototype of Swift is implemented upon DEX bytecode which is the official distribution format of Android applications. Yuan Zhang 0009, Min Yang 0002, Zhemin Yang, Binyu Zang |
VEE | 6 |
| 2012 | Mercury: Combining Performance with Dependability Using Self-Virtualization
Haibo Chen 0001, Fengzhe Zhang, Rong Chen 0001, Binyu Zang, Pen-Chung Yew |
J. Comput. Sci. Technol. | 4 |
| 2011 | A Hierarchical Approach to Maximizing MapReduce EfficiencyabstractIn this paper, we argued that Hadoop has limitations in exploiting data locality and task parallelism for multi-core platforms. We then extended Hadoop with a hierarchical MapReduce scheme. An in-memory cache scheme is also seamlessly integrated to cache data that is likely to be accessed in memory. Evaluation showed that the hierarchical scheme outperforms Hadoop ranging from 1.4x to 3.5x. Zhiwei Xiao, Haibo Chen 0001, Binyu Zang |
PACT | 3 |
| 2011 | A case for scaling applications to many-core with OS clusteringabstractThis paper proposes an approach to scaling UNIX-like operating systems for many cores in a backward-compatible way, which still enjoys common wisdom in new operating system designs. The proposed system, called Cerberus, mitigates contention on many shared data structures within OS kernels by clustering multiple commodity operating systems atop a VMM, and providing applications with the traditional shared memory interface. Cerberus extends a traditional VMMwith efficient support for resource sharing and communication among the clustered operating systems. It also routes system calls of an application among operating systems, to provide applications with the illusion of running on a single operating system. Haibo Chen 0001, Rong Chen 0001, Yuanxuan Wang, Binyu Zang |
EuroSys | 5 |
| 2011 | A comprehensive analysis and parallelization of an image retrieval algorithmabstractThe prevalence of the Internet and cloud computing has made multimedia data, such as image data and video data, become major data types in our daily life. For example, many data-intensive applications, such as health care and video recommendation, involve collecting, indexing and retrieving tera-scale multimedia data every day. With such a huge amount of multimedia data to process, the processing speed has been one of the major challenges to guarantee real-time requirements. The advent of multi-core hardware has opened new opportunities to improve the effectiveness of multimedia data processing. In this paper, we make a comprehensive analysis on different potential parallelism, including pipeline parallelism, task parallelism at both scale level and block level, data parallelism, and their combinations, in a typical image retrieval algorithm called SURF, which is the core algorithm of many multimedia (i.e., image and video) retrieval applications. Experimental results show the following observations of parallelism in SURF: 1) when only one level parallelism is exploited, block-level parallelism is more efficient and scalable than other alternatives; 2) data parallelism cannot be ignored especially when parallel resources increase and 3) the combination of block-level parallelism and pipeline parallelism is the most efficient parallelizing manner for the studied image retrieval algorithm. Based on these observations, we have implemented a parallel image retrieval algorithm. It can be easily mapped onto different multi-core platforms with good scalability. On a commodity server machine with 16-core, the parallel implementation achieves a speedup of 13X, which is 84% faster than P-SURF, a previous state-of-the-art parallelization of SURF on CPU; while on GPGPU, it achieves a speedup of 46X, which is 53% faster than CUDA SURF, a previous state-of-the-art parallelization of SURF on GPGPU. Zhenman Fang, Donglei Yang, Haibo Chen 0001, Binyu Zang |
ISPASS | 5 |
| 2011 | COREMU: a scalable and portable parallel full-system emulatorabstractThis paper presents the open-source COREMU, a scalable and portable parallel emulation framework that decouples the complexity of parallelizing full-system emulators from building a mature sequential one. The key observation is that CPU cores and devices in current (and likely future) multiprocessors are loosely-coupled and communicate through well-defined interfaces. Based on this observation, COREMU emulates multiple cores by creating multiple instances of existing sequential emulators, and uses a thin library layer to handle the inter-core and device communication and synchronization, to maintain a consistent view of system resources. COREMU also incorporates lightweight memory transactions, feedback-directed scheduling, lazy code invalidation and adaptive signal control to provide scalable performance. To make COREMU useful in practice, we also provide some preliminary tools and APIs that can help programmers to diagnose performance problems and (concurrency) bugs. A working prototype, which reuses the widely-used QEMU as the sequential emulator, is with only 2500 lines of code (LOCs) changes to QEMU. It currently supports x64 and ARM platforms, and can emulates up to 255 cores running commodity OSes with practical performance, while QEMU cannot scale above 32 cores. A set of performance evaluation against QEMU indicates that, COREMU has negligible uniprocessor emulation overhead, performs and scales significantly better than QEMU. We also show how COREMU could be used to diagnose performance problems and concurrency bugs of both OS kernel and parallel applications. Ran Liu 0003, Yufei Chen 0005, Xi Wu 0001, Haibo Chen 0001, Binyu Zang |
PPoPP | 7 |
| 2011 | CloudVisor: retrofitting protection of virtual machines in multi-tenant cloud with nested virtualizationabstractMulti-tenant cloud, which usually leases resources in the form of virtual machines, has been commercially available for years. Unfortunately, with the adoption of commodity virtualized infrastructures, software stacks in typical multi-tenant clouds are non-trivially large and complex, and thus are prone to compromise or abuse from adversaries including the cloud operators, which may lead to leakage of security-sensitive data. Fengzhe Zhang, Haibo Chen 0001, Binyu Zang |
SOSP | 4 |
| 2011 | ORDER: Object centRic DEterministic Replay for Java
Zhemin Yang, Min Yang 0002, Lvcai Xu, Haibo Chen 0001, Binyu Zang |
USENIX ATC | 5 |
| 2011 | Dynamic Software Updating Using a Relaxed Consistency ModelabstractSoftware is inevitably subject to changes. There are patches and upgrades that close vulnerabilities, fix bugs, and evolve software with new features. Unfortunately, most traditional dynamic software updating approaches suffer some level of limitations; few of them can update multithreaded applications when involving data structure changes, while some of them lose binary compatibility or incur nonnegligible performance overhead. This paper presents POLUS, a software maintenance tool capable of iteratively evolving running unmodified multithreaded software into newer versions, yet with very low performance overhead. The main idea in POLUS is a relaxed consistency model that permits the concurrent activity of the old and new code. POLUS borrows the idea of cache-coherence protocol in computer architecture and uses a ”bidirectional write-through” synchronization protocol to ensure system consistency. To demonstrate the applicability of POLUS, we report our experience in using POLUS to dynamically update three prevalent server applications: vsftpd, sshd, and Apache HTTP server. Performance measurements show that POLUS incurs negligible runtime overhead on the three applications—a less than 1 percent performance degradation (but 5 percent for one case). The time to apply an update is also minimal. Haibo Chen 0001, Jie Yu 0016, Chengqun Hang, Binyu Zang, Pen-Chung Yew |
IEEE Trans. Software Eng. | 4 |
| 2010 | Tiled-MapReduce: optimizing resource usages of data-parallel applications on multicore with tilingabstractThe prevalence of chip multiprocessor opens opportunities of running data-parallel applications originally in clusters on a single machine with many cores. MapReduce, a simple and elegant programming model to program large scale clusters, has recently been shown to be a promising alternative to harness the multicore platform. Rong Chen 0001, Haibo Chen 0001, Binyu Zang |
PACT | 3 |
| 2010 | Why software hangs and what can be done with itabstractSoftware hang is an annoying behavior and forms a major threat to the dependability of many software systems. To avoid software hang at the design phase or fix it in production runs, it is desirable to understand its characteristics. Unfortunately, to our knowledge, there is currently no comprehensive study on why software hangs and how to deal with it. In this paper, we study the reported hang-related bugs of four typical open-source software applications, aiming to gain insight into characteristics of software hang and provide some guidelines to fix them at the first place or remedy them in production runs. Haibo Chen 0001, Binyu Zang |
DSN | 3 |
| 2010 | Optimizing crash dump in virtualized environmentsabstractCrash dump, or core dump is the typical way to save memory image on system crash for future offline debugging and analysis. However, for typical server machines with likely abundant memory, the time of core dump can significantly increase the mean time to repair (MTTR) by delaying the reboot-based recovery, while not dumping the failure context for analysis would risk recurring crashes on the same problems. In this paper, we propose several optimization techniques for core dump in virtualized environments, in order to shorten the MTTR of consolidated virtual machines during crashes. First, we parallelize the process of crash dump and the process of rebooting the crashed VM, by dynamically reclaiming and allocating memory between the crashed VM and the newly spawned VM. Second, we use the virtual machine management layer to introspect the critical data structures of the crashed VM to filter out the dump of unused memory. Finally, we implement disk I/O rate control between core dump and the newly spawned VM according to user-tuned rate control policy to balance the time of crash dump and quality of services in the recovery VM. We have implemented a working prototype, Vicover, that optimizes core dump on system crash of a virtual machine in Xen, to minimize the MTTR of core dump and recovery as a whole. In our experiment on a virtualized TPC-W server, Vicover shortens the downtime caused by crash dump by around 5X. Yijian Huang, Haibo Chen 0001, Binyu Zang |
VEE | 3 |
| 2009 | Evaluating SPLASH-2 Applications Using MapReduce
Shengkai Zhu, Zhiwei Xiao, Haibo Chen 0001, Rong Chen 0001, Binyu Zang |
APPT | 6 |
| 2009 | Control flow obfuscation with information flow trackingabstractRecent micro-architectural research has proposed various schemes to enhance processors with additional tags to track various properties of a program. Such a technique, which is usually referred to as information flow tracking, has been widely applied to secure software execution (e.g., taint tracking), protect software privacy and improve performance (e.g., control speculation). Haibo Chen 0001, Liwei Yuan 0003, Xi Wu 0001, Binyu Zang, Bo Huang 0002, Pen-Chung Yew |
MICRO | 4 |
| 2008 | From Speculation to Security: Practical and Efficient Information Flow Tracking Using Speculative HardwareabstractDynamic information flow tracking (also known as taint tracking) is an appealing approach to combat various security attacks. However, the performance of applications can severely degrade without hardware support for tracking taints. This paper observes that information flow tracking can be efficiently emulated using deferred exception tracking in microprocessors supporting speculative execution. Based on this observation, we propose SHIFT, a low-overhead, software-based dynamic information flow tracking system to detect a wide range of attacks. The key idea is to treat tainted state (describing untrusted data) as speculative state (describing deferred exceptions). SHIFT leverages existing architectural support for speculative execution to track tainted state in registers and needs to instrument only load and store instructions to track tainted state in memory using a bitmap, which results in significant performance advantages. Moreover, by decoupling mechanisms for taint tracking from security policies, SHIFT can detect a wide range of exploits, including high-level semantic attacks. We have implemented SHIFT using the Itanium processor, which has support for deferred exceptions, and by modifying GCC to instrument loads and stores. A security assessment shows that SHIFT can detect both low-level memory corruption exploits as well as high-level semantic attacks with no false positives. Performance measurements show that SHIFT incurs about 1% overhead for server applications. The performance slowdown for SPEC-INT2000 is 2.81X and 2.27X for tracking at byte-level and wordlevel respectively. Minor architectural improvements to the Itanium processor (adding three simple instructions) can reduce the performance slowdown down to 2.32X and 1.8X for byte-level and word-level tracking, respectively. Haibo Chen 0001, Xi Wu 0001, Liwei Yuan 0003, Binyu Zang, Pen-Chung Yew, Fred Chong |
ISCA | 4 |
| 2008 | An ontological engineering approach for automating inspection and quarantine at airports
Binyu Zang, Zhuangjian Chen, Chen-Fang Tsai, Christopher Laing |
J. Comput. Syst. Sci. | 1 |
| 2007 | Optimizing Bandwidth Constraint through Register Interconnection for Stream Processors
Binyu Zang, Chuanqi Zhu |
PACT | 3 |
| 2007 | Distance Measurement in PanoramaabstractComputing world distances of scene features from the captured images is a common task in image analysis and scene understanding. Previous projective geometry based methods focus on measuring distance from one single image. Hence, the scope of measurable scene is limited by the field-of-view (FOV) of one single camera. In this paper, we propose one method of measuring distances of line segments in real world scene using panorama representation. With a full view panorama, the scope of measurable scene is increased and can fully cover the sphere of 360 times 360 FOV. With panorama representation, distance of long-range features, which can not be fully captured by a single image, can be measured from the panoramic image. A prototype system called PanoMeasure is developed for enabling user to interactively measure the distances of line segments. Experiments with simulated data and real measurement results verify that the method offers high accuracy. Zhongding Jiang, Binyu Zang |
ICIP (6) | 4 |
| 2007 | A High Resolution Video Display System by Seamlessly Tiling Multiple ProjectorsabstractWith the rapid advances of digital photography technology, high resolution video can be recorded using video camera for home entertainment and digital cinema. High end projector used for high resolution video display is bulky and expensive, only suitable for fixed installation, and prevents itself from being widely adopted. In this paper, we present one high resolution video display system using multiple projectors to replace the high end projector playback system. Our system is driven by single PC with PCI-E16 x interface and supports at least triple projectors, which is suitable for building surround video display system. It fully exploits the high bandwidth between CPU and graphics card, and efficiently plays back high resolution video of any format supported by commercial video player. We design geometric alignment and photometric correction methods to display the video content seamlessly on the planar display screen. Experimental results show that our system displays high resolution video at its full frame rate. Zhongding Jiang, Yandong Mao, Binyu Zang |
ICME | 4 |
| 2007 | Mercury: Combining Performance with Dependability Using Self-virtualizationabstractThere has recently been increasing interests in using system virtualization to improve the dependability of HPC cluster systems. However, it is not cost-free and may come with some performance degradation, uncertain QoS and loss of functionalities. Meanwhile, many virtualization-enabled features such as online maintenance and fault tolerance do not require virtualization being always on. This paper proposes a technique, called self-virtualization, that supports dynamically attaching and detaching a full-fledged virtual machine monitor (VMM) beneath an operating system, without disturbing applications thereon, and rid the system of potential overhead when the virtualization is not needed. This technique enables HPC clusters to reap most benefits from virtualization without sacrificing performance. This paper presents the design and implementation of Mercury, a working prototype based on Linux and Xen VMM. Our performance measurement shows that Mercury incurs very little overhead: about 0.2 ms to complete a mode switch, and negligible performance degradation compared to Linux. Haibo Chen 0001, Rong Chen 0001, Fengzhe Zhang, Binyu Zang, Pen-Chung Yew |
ICPP | 4 |
| 2007 | POLUS: A POwerful Live Updating SystemabstractThis paper presents POLUS, a software maintenance tool capable of iteratively evolving running software into newer versions. POLUS's primary goal is to increase the dependability of contemporary server software, which is frequently disrupted either by external attacks or by scheduled upgrades. To render POLUS both practical and powerful, we design and implement POLUS aiming to retain backward binary compatibility, support for multithreaded software and recover already tainted state of running software, yet with good usability and very low runtime overhead. To demonstrate the applicability of POLUS, we report our experience in using POLUS to dynamically update three prevalent server applications: vsftpd, sshd and apache HTTP server. Performance measurements show that POLUS incurs negligible runtime overhead: a less than 1% performance degradation (but 5% for one case). The time to apply an update is also minimal. Haibo Chen 0001, Jie Yu 0016, Rong Chen 0001, Binyu Zang, Pen-Chung Yew |
ICSE | 4 |
| 2007 | Optimizing software cache performance of packet processing applicationsabstractNetwork processors (NPs) are widely used in many types of networking equipment due to their high performance and flexibility. For most NPs, software cache is used instead of hardware cache due to the chip area, cost and power constraints. Therefore, programmers should take full responsibility for software cache management which is neither intuitive nor easy to most of them. Actually, without an effective use of it, long memory access latency will be a critical limiting factor to overall applications. Prior researches like hardware multi-threading, wide-word accesses and packet access combination for caching have already been applied to help programmers to overcome this bottleneck. However, most of them do not make enough use of the characteristics of packet processing applications and often perform intraprocedural optimizations only. As a result, the binary codes generated by those techniques often get lower performance than that comes from hand-tuned assembly programming for some applications. In this paper, we propose an algorithm including two techniques - Critical Path Based Analysis (CPBA) and Global Adaptive Localization (GAL), to optimize the software cache performance of packet processing applications. Packet processing applications usually have several hot paths and CPBA tries to insert localization instructions according to their execution frequencies. For further optimizations, GAL eliminates some redundant localization instructions by interprocedural analysis and optimizations. Our algorithm is applied on some representative applications. Experiment results show that it leads to an average speedup by a factor of 1.974. Junpu Chen, Min Yang 0002, Binyu Zang |
LCTES | 5 |
| 2006 | Optimizing compiler for shared-memory multiple SIMD architectureabstractWith the rapid growth of multimedia and game, these applications put more and more pressure on the processing ability of modern processors. Multiple SIMD architecture is widely used in multimedia processing field as a multimedia accelerator.With the consideration of power consumption and chip size, shared memory multiple SIMD architecture is mainly used in embedded SOCs. In order to further fit mobile environment, there is the constraint of limited register number as well. Although shared memory multiple SIMD architecture simplify the chip design, these constraints are the major obstacles to map the real multimedia applications to these architectures. Until now, to our best knowledge, there is little research on the optimizing techniques for shared memory multiple SIMD architecture.In this paper, we present a compiler framework, which aims at automatically generating high performance codes for shared memory multiple SIMD architecture. In this framework, we reduce the competition of shared data bus through increasing the register locality, improve the utilization of data bus by read-only data vector replication and solve the problem of limited register number through a resource allocation algorithm. The framework also handlers the issues concerning on data transformation. As the experimental results shown, this framework is successful in mapping real multimedia applications to shared memory multiple SIMD architecture. It leads to an average speedup by a factor of 3.19 and an average utilization of SM-SIMD architecture with 8 SIMD units by a factor of 52.6%. Xinglong Qian, Binyu Zang, Chuanqi Zhu |
LCTES | 4 |
| 2006 | Live updating operating systems using virtualizationabstractMany critical IT infrastructures require non-disruptive operations. However, the operating systems thereon are far from perfect that patches and upgrades are frequently applied, in order to close vulnerabilities, add new features and enhance performance. To mitigate the loss of availability, such operating systems need to provide features such as live update through which patches and upgrades can be applied without having to stop and reboot the operating system. Unfortunately, most current live updating approaches cannot be easily applied to existing operating systems: some are tightly bound to specific design approaches (e.g. object-oriented); others can only be used under particular circumstances (e.g. quiescence states).In this paper, we propose using virtualization to provide the live update capability. The proposed approach allows a broad range of patches and upgrades to be applied at any time without the requirement of a quiescence state. Moreover, such approach shares good portability for its OS-transparency and is suitable for inclusion in general virtualization systems. We present a working prototype, LUCOS, which supports live update capability on Linux running on Xen virtual machine monitor. To demonstrate the applicability of our approach, we use real-life kernel patches from Linux kernel 2.6.10 to Linux kernel 2.6.11, and apply some of those kernel patches on the fly. Performance measurements show that our implementation incurs negligible performance overhead: a less than 1% performance degradation compared to a Xen-Linux. The time to apply a patch is also very minimal. Haibo Chen 0001, Rong Chen 0001, Fengzhe Zhang, Binyu Zang, Pen-Chung Yew |
VEE | 4 |
| 2005 | Boosting the Performance of Multimedia Applications Using SIMD Instructions
Weihua Jiang, Bo Huang 0002, Binyu Zang, Chuanqi Zhu |
CC | 6 |
| 2005 | ADDI: an agent-based extension to UDDI for supply chain managementabstractThis paper proposes a new mediate module ADDI to organize the UDDI nodes published by the different corporations on supply chain and establish a UDDI dynamic ranking model. We suggest a novel model for supply chain conduct classification, and define the supply chain entity conduct types as a standard for Web service classification and UDDI ranking. In this model, Web service-oriented technologies and protocols are deployed for modeling, managing and executing business-oriented functionalities and environments. We focus on the efficient integration of supply chain as key points to harmonize these technologies. Agent-orientation concepts and technologies are applied for SCM construction and interaction patterns. Mi Zhang 0001, Zunping Cheng, Yongzun Zhao, Joshua Zhexue Huang, Binyu Zang |
CSCWD (1) | 6 |
| 2005 | A Web Service-Based Framework for Supply Chain ManagementabstractThis paper proposes a framework based on Web service to organize the corporation nodes on supply chain. We have defined a strategy for aggregating the agents, including the normal agent and the mobile agent, into the Web service architecture and the functionalities for them to control the business conducts. We also devise a UDDI ranking frame based on analysis of supply chain activities. In this frame, Web service-oriented technologies and protocols are deployed for modeling, managing and executing business-oriented functionalities. We focus on the efficient integration of supply chain as key points to harmonize these technologies. Agent-orientation concepts and technologies are applied to SCM construction and interaction patterns. Mi Zhang 0001, Zunping Cheng, Binyu Zang |
ISORC | 5 |
| 2002 | Run-Time Data-Flow Analysis
Binyu Zang, Chuanqi Zhu |
J. Comput. Sci. Technol. | 2 |
| 2001 | A New Approach to Pointer Analysis for Assignments
Bo Huang 0002, Binyu Zang, Chuanqi Zhu |
J. Comput. Sci. Technol. | 2 |
| 1997 | Exploiting loop parallelism with redundant execution
Weiyu Tang, Wu Shi, Binyu Zang, Chuanqi Zhu |
J. Comput. Sci. Technol. | 3 |