VLDB 2026 Research / reviewers in the wild / expert
Chenggang Wu 0002
dblp:51/3529-2
· DBLP profile ↗
42ranked-venue papers
2as first author
21since 2021 · last 2026
0000-0003-1777-8110ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 20 · 1 first-author · 5 since 2021Security and privacy · 14 · 1 first-author · 12 since 2021Software engineering, systems software and programming languages · 13 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fuzzing JavaScript JIT Compilers With Optimization Path Feedback
Jiming Wang, Chenggang Wu 0002, Yan Kang 0002, Yuhao Hu, Jikai Ren, Yuanming Lai, Mengyao Xie, Chao Zhang 0008, Tao Li 0022, Zhe Wang 0017 |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2025 | SyzParam: Incorporating Runtime Parameters into Kernel Driver FuzzingabstractUnder the monolithic architecture of the Linux kernel, all its components operate within the same address space. Notably, device drivers constitute over half of the kernel codebase yet are particularly prone to bugs. Therefore, exploring vulnerabilities in drivers is critical for ensuring kernel security. Extensive research has been done to fuzz kernel drivers through system calls and hardware interrupts. Through a comprehensive study of the Linux Kernel Device Model, we identified that the execution of device drivers is also influenced by runtime parameters, including device attributes and kernel module parameters. Our analysis reveals that large portions of the uncovered code are masked by these parameters, which are exposed to the userspace through a specialized virtual file system known as sysfs. Furthermore, adjacent devices interconnected within the same device tree also impact drivers' behavior. Yan Kang 0002, Chenggang Wu 0002, Kangjie Lu, Jiming Wang, Xingwei Li, Yuhao Hu, Jikai Ren, Yuanming Lai, Mengyao Xie, Zhe Wang 0017 |
CCS | 3 |
| 2025 | Tide: An Efficient Kernel-level Isolation Execution Environment on AArch64 via Dynamically Adjusting Output Address SizeabstractTo enforce the privilege separation in the kernel, kernel-level isolated execution environment (IEE) has become a recent research trend because it can protect critical resources and monitors. Our research found that to isolate the IEE memory, all existing IEEs must act as a reference monitor to isolate page tables and validate their updates, bringing a significant performance overhead. Hence, we propose Tide, a new kernel-level IEE based on the output address size hardware feature on AArch64, which could offload such checks to the hardware. However, it still faces the flexibility and security challenges. To address them, Tide presents using the stage-2 translation to expand the physical address range to flexibly map the IEE memory and perform extra access controls on the physical memory; it designs a novel gate to enter (sneak) into the IEE securely by disabling translation temporarily, and ensures it can only be executed at the fixed locations. The experimental results show that Tide is performant than all existing IEEs on protecting critical kernel structures and security tools. Shiyang Zhang, Chenggang Wu 0002, Chengxuan Hou, Jinglin Lv, Yinqian Zhang, Yuanming Lai, Mengyao Xie, Yan Kang 0002, Zhe Wang 0017 |
CCS | 2 |
| 2025 | BCFuzz: Bytecode-Driven Fuzzing for JavaScript EnginesabstractThe interpreter and the Just-In-Time (JIT) compiler are two core components of modern JavaScript engines, both of which take bytecodes as input. Most bugs in these components are closely related to specific bytecodes. Therefore, effective fuzzing should pay close attention to how bytecode is generated and exercised. However, previous work fails to consider this aspect and instead focuses primarily on the syntactic and semantic validity of test cases. This causes two major issues: 1) certain bytecodes are never exercised during fuzzing; 2) some bytecodes are exercised infrequently. In this paper, we propose BCFuzz, a bytecode-driven fuzzing approach designed to enhance the diversity of generated bytecode and increase testing opportunities for low-frequency bytecodes. Specifically, we introduce a parser-oriented probing technique to identify the necessary conditions for generating specific bytecodes and use this information to enhance the input generation process. To better test low-frequency bytecodes, we propose bytecode-aware seed preservation, scheduling, and mutation strategies. We evaluate BCFuzz on four mainstream JavaScript engines. In 72 hours of testing, BCFuzz discovers 1.73× and 1.67× more bugs than DIE and Fuzzilli, respectively. In total, BCFuzz uncovered 20 previously unknown bugs. Of these, 17 have already been fixed and one has been assigned a CVE. All the discovered bugs are related to bytecodes. Jiming Wang, Chenggang Wu 0002, Jikai Ren, Yuhao Hu, Yan Kang 0002, Yuanming Lai, Mengyao Xie, Zhe Wang 0017 |
ASE | 2 |
| 2025 | Shining Light on the Inter-procedural Code Obfuscation: Keep Pace with Progress in Binary DiffingabstractSoftware obfuscation techniques have lost their effectiveness due to the rapid development of binary diffing techniques, which can achieve accurate function matching and identification. In this paper, we propose a new inter-procedural code obfuscation mechanism KHaos , 1 which moves the code across functions to obfuscate the function by using compilation optimizations. Three obfuscation primitives are proposed to separate, aggregate, and hide the function. They can be combined to enhance the obfuscation effect further. This article also reveals distinguishing factors on obfuscation and compiler optimization and presents novel observations to gain insights into the impact of actively utilizing compiler optimization in obfuscation. A prototype of KHaos is implemented and evaluated on a large number of real-world programs. Experimental results show that KHaos outperforms existing code obfuscations and can significantly reduce the accuracy rates of six state-of-the-art binary diffing techniques with lower runtime overhead. Peihua Zhang, Chenggang Wu 0002, Hanzhi Hu, Lichen Jia, Mingfan Peng, Mengyao Xie, Yuanming Lai, Yan Kang 0002, Zhe Wang 0017 |
ACM Trans. Archit. Code Optim. | 2 |
| 2025 | Yesterday Once MorE: Facilitating Linux Kernel Bug Reproduction via Reverse FuzzingabstractThe Linux kernel remains vulnerable to numerous bugs, with approximately 65% detected by Syzkaller lacking Proof-of-Concept (PoC), hampering risk mitigation efforts. These bugs, termed irreproducible kernel bugs, highlight the challenge of statefulness issue-related irreproducibility in kernel fuzzing, which is an open research without definitive solutions. Our investigation reveals that suboptimal seed quality distribution in fuzzing is the root obstacle preventing effective tracking of the states leading to crashes. Inspired by this insight, we introduce Reverse Fuzzing (RF), an innovative approach that infers hard-to- reach states by continuously reverse-oriented deriving from subsequently encountered bridge states to increase reproduction probability. RF differentiates between the “trigger” seed, which directly causes crashes, and “activator” seeds, which establish the necessary preconditions, prioritizing exploration around trigger while simultaneously regenerating and maintaining activators during fuzzing, which effectively facilitate to restructure such elusive states from “yesterday”. We implement YOME, a prototype leveraging RF to strike a balance between fuzzing efficiency and effectiveness through customized scheduling and mutation strategies, armed with a refinement mechanism to improve seed quality distribution. Our evaluations validate that YOME reproduce 110% more bugs than previous kernel fuzzers and demonstrate its practicality in real-world scenarios. YOME generated 125 PoCs (30.1% of the total) and uncovered 23 unique bugs, with 40 confirmed and 5 assigned CVEs. Xingwei Li, Yan Kang 0002, Chenggang Wu 0002, Danjun Liu, Jiming Wang, Zehui Wu, Yunchao Wang, Rongkuan Ma |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2024 | CodeExtract: Enhancing Binary Code Similarity Detection with Code Extraction TechniquesabstractIn the field of binary code similarity detection (BCSD), when dealing with functions in binary form, the conventional approach is to identify a set of functions that are most similar to the target function. These similar functions often originate from the same source code but may differ due to variations in compilation settings. Such analysis is crucial for applications in the security domain, including vulnerability discovery, malware detection, software plagiarism detection, and patch analysis. Function inlining, an optimization technique employed by compilers, embeds the code of callee functions directly into the caller function. Due to different compilation options (such as O1 and O3) leading to varying levels of function inlining, this results in significant discrepancies between binary functions derived from the same source code under different compilation settings, posing challenges to the accuracy of state-of-the-art (SOTA) learning-based binary code similarity detection (LB-BCSD) methods. In contrast to function inlining, code extraction technology can identify and separate duplicate code within a program, replacing it with corresponding function calls. To overcome the impact of function inlining, this paper introduces a novel approach, CodeExtract. This method initially utilizes code extraction techniques to transform code introduced by function inlining back into function calls. Subsequently, it actively inlines functions that cannot undergo code extraction, effectively eliminating the differences introduced by function inlining. Experimental validation shows that CodeExtract enhances the accuracy of LB-BCSD models by 20% in addressing the challenges posed by function inlining. Lichen Jia, Chenggang Wu 0002, Peihua Zhang, Zhe Wang 0017 |
LCTES | 2 |
| 2024 | A Tale of Two Paths: Toward a Hybrid Data Plane for Efficient Far-Memory Applications
Chenxi Wang 0005, Yifan Qiao 0002, Zhe Wang 0017, Chenggang Wu 0002, Youyou Lu, Xiaobing Feng 0002, Huimin Cui, Shan Lu 0001, Guoqing Harry Xu |
OSDI | 7 |
| 2024 | OptFuzz: Optimization Path Guided Fuzzing for JavaScript JIT Compilers
Jiming Wang, Yan Kang 0002, Chenggang Wu 0002, Yuhao Hu, Jikai Ren, Yuanming Lai, Mengyao Xie, Tao Li 0022, Zhe Wang 0017 |
USENIX Security Symposium | 3 |
| 2024 | HIVE: A Hardware-assisted Isolated Execution Environment for eBPF on AArch64
Peihua Zhang, Chenggang Wu 0002, Yinqian Zhang, Mingfan Peng, Shiyang Zhang, Mengyao Xie, Yuanming Lai, Yan Kang 0002, Zhe Wang 0017 |
USENIX Security Symposium | 2 |
| 2023 | PANIC: PAN-assisted Intra-process Memory Isolation on ARMabstractIntra-process memory isolation is a well-known technique to enforce least privilege within a process. In this paper, we propose a generic and efficient intra-process memory isolation technique named PANIC, by leveraging Privileged Access Never (PAN) and load/store unprivileged (LSU) instructions on AArch64. PANIC executes process code in kernel mode and compartments code into trusted and untrusted components. The untrusted code is restricted from accessing the isolated memory region, which is located on user pages, and the trusted code is allowed to access the isolated memory region by using LSU instructions. To mitigate threats induced by running user code in kernel mode, PANIC provides two novel security mechanisms: shim-based memory isolation and sensitive instruction emulation. PANIC provides a generic and efficient isolation primitive that can be applied in three different isolation scenarios: protecting sensitive data in CFI, creating isolated execution environments, and hardening JIT code cache. We have implemented a prototype of PANIC and experimental evaluation shows that PANIC incurs very low performance overhead, and performs better than existing methods. Mengyao Xie, Chenggang Wu 0002, Yinqian Zhang, Qijing Li, Yuanming Lai, Yan Kang 0002, Wei Wang 0385, Zhe Wang 0017 |
CCS | 3 |
| 2023 | Khaos: The Impact of Inter-procedural Code Obfuscation on Binary Diffing TechniquesabstractSoftware obfuscation techniques can prevent binary diffing techniques from locating vulnerable code by obfuscating the third-party code, to achieve the purpose of protecting embedded device software. With the rapid development of binary diffing techniques, they can achieve more and more accurate function matching and identification by extracting the features within the function. This makes existing software obfuscation techniques, which mainly focus on the intra-procedural code obfuscation, no longer effective. Peihua Zhang, Chenggang Wu 0002, Mingfan Peng, Ding Yu, Yuanming Lai, Yan Kang 0002, Wei Wang 0385, Zhe Wang 0017 |
CGO | 2 |
| 2023 | SpecWands: An Efficient Priority-Based Scheduler Against Speculation Contention AttacksabstractTransient execution attacks (TEAs) have gradually become a major security threat to modern high-performance processors. They exploit the vulnerability of speculative execution to illegally access private data, and transmit them through timing-based covert channels. While new vulnerabilities are discovered continuously, the covert channels can be categorized to two types: 1) Persistent Type, in which covert channels are based on the layout changes of buffering, e.g., through caches or TLBs and 2) Volatile Type, in which covert channels are based on the contention of sharing resources, e.g., through execution units or issuing ports. The defenses against the persistent-type covert channels have been well addressed, while those for the volatile-type are still rather inadequate. Existing mitigation schemes for the volatile type such as Speculative Compression and Time-Division-Multiplexing will introduce significant overhead due to the need to stall the pipeline or to disallow resource sharing. In this article, we look into such attacks and defenses with a new perspective, and propose a scheduling-based mitigation scheme, called SpecWands. It consists of three priority-based scheduling policies to prevent an attacker from transmitting the secret in different contention situations. SpecWands not only can defend against both interthread and intrathread-based attacks but also can keep most of the performance benefit from speculative execution and resource-sharing. We evaluate its runtime overhead on SPEC 2017 benchmarks and realistic programs. The experimental results show that SpecWands has a significant performance advantage over the other two representative schemes. Bowen Tang 0001, Chenggang Wu 0002, Pen-Chung Yew, Yinqian Zhang, Mengyao Xie, Yuanming Lai, Yan Kang 0002, Wei Wang 0385, Zhe Wang 0017 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2023 | SpecBox: A Label-Based Transparent Speculation Scheme Against Transient Execution AttacksabstractSpeculative execution techniques have been a cornerstone of modern processors to improve instruction-level parallelism. However, recent studies showed that this kind of techniques could be exploited by attackers to leak secret data via transient execution attacks, such as Spectre. Many defenses are proposed to address this problem, but they all face various challenges: (1) Tracking data flow in the instruction pipeline could comprehensively address this problem, but it could cause pipeline stalls and incur high performance overhead; (2) Making side effect of speculative execution imperceptible to attackers, but it often needs additional storage components and complicated data movement operations. In this article, we propose alabel-based transparent speculationscheme calledSpecBox. It dynamically partitions the cache system to isolate speculative data and non-speculative data, which can prevent transient execution from being observed by subsequent execution. Moreover, it uses thread ownership semaphores to prevent speculative data from being accessed across cores. In addition,SpecBoxalso enhances the auxiliary components in the cache system against transient execution attacks, such as hardware prefetcher. Our security analysis shows thatSpecBoxis secure and the performance evaluation shows that the performance overhead on SPEC CPU 2006 and PARSEC-3.0 benchmarks is small. Bowen Tang 0001, Chenggang Wu 0002, Zhe Wang 0017, Lichen Jia, Pen-Chung Yew, Yueqiang Cheng, Yinqian Zhang, Chenxi Wang 0005, Guoqing Harry Xu |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2023 | Dancing With Wolves: An Intra-Process Isolation Technique With Privileged HardwareabstractIntra-process memory isolation is a cornerstone technique of protecting the sensitive data in memory-corruption defenses, such as the shadow stack in control flow integrity (CFI) and the safe region in code pointer integrity (CPI). In this article, we proposeSEIMI, a highly efficient intra-process memory isolation technique for memory-corruption defenses. The core is to use the efficientSupervisor-mode Access Prevention (SMAP), a hardware feature that is originally used for preventing the kernel from accessing the user space, to achieve intra-process memory isolation. To leverage SMAP,SEIMIcreatively executes the user code in the privileged mode. In addition to enabling the new design of the SMAP-based memory isolation, we further develop multiple new techniques to ensure secure escalation of user code. Extensive experiments show thatSEIMIoutperforms existing isolation mechanisms, including theMemory Protection Keys(MPK) based scheme and theMemory Protection Extensions(MPX) based scheme. Chenggang Wu 0002, Mengyao Xie, Zhe Wang 0017, Yinqian Zhang, Kangjie Lu, Yuanming Lai, Yan Kang 0002, Min Yang 0002, Tao Li 0022 |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2022 | CETIS: Retrofitting Intel CET for Generic and Efficient Intra-process Memory IsolationabstractIntel control-flow enforcement technology (CET) is a new hardware feature available in recent Intel processors. It supports the coarse-grained control-flow integrity for software to defeat memory corruption attacks. In this paper, we retrofit CET, particularly the write-protected shadow pages of CET used for implementing shadow stacks, to develop a generic and efficient intra-process memory isolation mechanism, dubbed CETIS. Mengyao Xie, Chenggang Wu 0002, Yinqian Zhang, Yuanming Lai, Yan Kang 0002, Wei Wang 0385, Zhe Wang 0017 |
CCS | 2 |
| 2022 | KOP-Fuzzer: A Key-Operation-based Fuzzer for Type Confusion Bugs in JavaScript EnginesabstractJavaScript (JS) engines are a core component of a lot of software, such as web browsers, PDF readers and flash players. There has been much research on finding JS engine vulnerabilities. However, due to the fact that a JS engine's input space is infinite and the vulnerability triggering conditions are extremely strict, it is difficult to generate test cases that are able to trigger deep logic errors in fuzzing. This paper aims to explore an approach which incorporates the human experience into fuzzing. We propose a Key-Operation-based Fuzzer (KOP-Fuzzer), to explore the type confusion vulnerabilities in JS engines. Based on human knowledge, we summarize a trigger model and extract key operations for type confusion vulnerabilities in JS engines. We use clustering to extract the key-operation methods from the engine's source code and develop a fuzzing system for key -operation mutation. Our experimental results demonstrate that the KOP-Fuzzer generates valid test cases with 1.5x fewer runtime errors, while also improving the edge coverage (2.082 %) and key-operation coverage (6.452 %), when compared with the state-of-the-art JS engine fuzzers. The KOP-Fuzzer discovered a total of 21 new bugs in ChakraCore and JavaScriptCore, where 16 of them are caused by the engine's incorrect handling of key operations and 5 of them are caused by type confusions. Lili Sun, Chenggang Wu 0002, Zhe Wang 0017, Yan Kang 0002, Bowen Tang 0001 |
COMPSAC | 2 |
| 2022 | SoftTRR: Protect Page Tables against Rowhammer Attacks using Software-only Target Row Refresh
Zhi Zhang 0001, Yueqiang Cheng, Wenhao Wang 0001, Surya Nepal, Yansong Gao 0001, Zhe Wang 0017, Chenggang Wu 0002 |
USENIX ATC | 10 |
| 2022 | Ferry: State-Aware Symbolic Execution for Exploring State-Dependent Program Paths
Shunfan Zhou, Zhemin Yang, Peng Liu 0005, Min Yang 0002, Zhe Wang 0017, Chenggang Wu 0002 |
USENIX Security Symposium | 7 |
| 2022 | Making Information Hiding Effective AgainabstractInformation hiding (IH) is an important building block for many defenses against code reuse attacks, such as code-pointer integrity (CPI), control-flow integrity (CFI) and fine-grained code (re-)randomization, because of its effectiveness and performance. It employs randomization to probabilistically “hide” sensitive memory areas, called safe areas, from attackers and ensures their addresses are not leaked by any pointers directly. These defenses used safe areas to protect their critical data, such as jump targets and randomization secrets. However, recent works have shown that IH is vulnerable to various attacks. In this article, we propose a new IH technique called SafeHidden. It continuously re-randomizes the locations of safe areas and thus prevents the attackers from probing and inferring the memory layout to find its location. A new thread-private memory mechanism is proposed to isolate the thread-local safe areas and prevent adversaries from reducing the randomization entropy. It also randomizes the safe areas after the TLB misses to prevent attackers from inferring the address of safe areas using cache side-channels. Existing IH-based defenses can utilize SafeHidden directly without any change. Our experiments show that SafeHidden not only prevents existing attacks effectively but also incurs low performance overhead. Zhe Wang 0017, Chenggang Wu 0002, Yinqian Zhang, Bowen Tang 0001, Pen-Chung Yew, Mengyao Xie, Yuanming Lai, Yan Kang 0002, Yueqiang Cheng, Zhi-Ping Shi 0002 |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2021 | Editorial for the special issue on reliability and power efficiency for HPC
Jifeng He 0001, Chenggang Wu 0002, Huawei Li 0001, Yang Guo 0003, Tao Li 0022 |
CCF Trans. High Perform. Comput. | 2 |
| 2020 | SEIMI: Efficient and Secure SMAP-Enabled Intra-process Memory IsolationabstractMemory-corruption attacks such as code-reuse attacks and data-only attacks have been a key threat to systems security. To counter these threats, researchers have proposed a variety of defenses, including control-flow integrity (CFI), code-pointer integrity (CPI), and code (re-)randomization. All of them, to be effective, require a security primitive—intra-process protection of confidentiality and/or integrity for sensitive data (such as CFI’s shadow stack and CPI’s safe region).In this paper, we propose SEIMI, a highly efficient intra-process memory isolation technique for memory-corruption defenses to protect their sensitive data. The core of SEIMI is to use the efficient Supervisor-mode Access Prevention (SMAP), a hardware feature that is originally used for preventing the kernel from accessing the user space, to achieve intra-process memory isolation. To leverage SMAP, SEIMI creatively executes the user code in the privileged mode. In addition to enabling the new design of the SMAP-based memory isolation, we further develop multiple new techniques to ensure secure escalation of user code, e.g., using the descriptor caches to capture the potential segment operations and configuring the Virtual Machine Control Structure (VMCS) to invalidate the execution result of the control registers related operations. Extensive experimental results show that SEIMI outperforms existing isolation mechanisms, including both the Memory Protection Keys (MPK) based scheme and the Memory Protection Extensions (MPX) based scheme, while providing secure memory isolation. Zhe Wang 0017, Chenggang Wu 0002, Mengyao Xie, Yinqian Zhang, Kangjie Lu, Yuanming Lai, Yan Kang 0002, Min Yang 0002 |
SP | 2 |
| 2019 | SafeHidden: An Efficient and Secure Information Hiding Technique Using Re-randomization
Zhe Wang 0017, Chenggang Wu 0002, Yinqian Zhang, Bowen Tang 0001, Pen-Chung Yew, Mengyao Xie, Yuanming Lai, Yan Kang 0002, Yueqiang Cheng, Zhi-Ping Shi 0002 |
USENIX Security Symposium | 2 |
| 2018 | Using Local Clocks to Reproduce Concurrency BugsabstractMulti-threaded programs play an increasingly important role in current multi-core environments. Exposing concurrency bugs and debugging such multi-threaded programs are quite challenging due to their inherent non-determinism. In order to mitigate such non-determinism, many approaches such as record-and-replay have been proposed. However, those approaches often suffer significant performance degradation because they require a large amount of recorded information and/or long analysis and replay time. In this paper, we propose an efficient and effective approach, ReCBuLC (reproducing concurrency bugs using local clocks), to take advantage of the hardware clocks available on modern processors. The key idea is to reduce the recording overhead and the time to analyze events’ global order by recording timestamps in each thread. These timestamps are used to determine the global order of shared accesses. To avoid the large overhead in accessing system-wide global clock, we opt to use local per-core clocks that incur much less access overhead. We then propose techniques to resolve skews among local clocks and obtain an accurate global event order. By using per-core clocks, state-of-the-art bug reproducing systems such as PRES and CLAP can reduce their recording overheads by up to 85 percent, and the analysis time up to 84.66%$\sim$99.99%, respectively. Zhe Wang 0017, Chenggang Wu 0002, Zhenjiang Wang, Pen-Chung Yew, Jeff Huang 0001, Xiaobing Feng 0002, Yanyan Lan, Yunji Chen, Yuanming Lai |
IEEE Trans. Software Eng. | 2 |
| 2017 | SysMon: Monitoring Memory Behaviors via OS Approach
Mengyao Xie, Lei Liu 0030, Chenggang Wu 0002, Hongna Geng |
APPT | 4 |
| 2017 | ReRanz: A Light-Weight Virtual Machine to Mitigate Memory Disclosure AttacksabstractRecent code reuse attacks are able to circumvent various address space layout randomization (ASLR) techniques by exploiting memory disclosure vulnerabilities. To mitigate sophisticated code reuse attacks, we proposed a light-weight virtual machine, ReRanz, which deployed a novel continuous binary code re-randomization to mitigate memory disclosure oriented attacks. In order to meet security and performance goals, costly code randomization operations were outsourced to a separate process, called the "shuffling process". The shuffling process continuously flushed the old code and replaced it with a fine-grained randomized code variant. ReRanz repeated the process each time an adversary might obtain the information and upload a payload. Our performance evaluation shows that ReRanz Virtual Machine incurs a very low performance overhead. The security evaluation shows that ReRanz successfully protect the Nginx web server against the Blind-ROP attack. Zhe Wang 0017, Chenggang Wu 0002, Yuanming Lai, Xiangyu Zhang 0001, Wei-Chung Hsu, Yueqiang Cheng |
VEE | 2 |
| 2016 | Memos: A full hierarchy hybrid memory management frameworkabstractIn this paper, we introduce memos, which integrates suitable memory management policies and schedules resources over the entire memory hierarchy in hybrid memory system. Powered by an OS kernel level monitoring tool, memos captures memory patterns online, and then leverages them to guide the memory page placement and data mapping. Experimental results show, on average, memos can benefit memory utilization, contributing to system throughput and QoS by 19.1% and 23.6%. Moreover, memos can reduce the NVM side memory latency by 3∼83.3%, energy consumption by 25.1∼99%, and benefit the NVM lifetime significantly (40× improvement on average). Lei Liu 0030, Mengyao Xie, Chenggang Wu 0002 |
ICCD | 6 |
| 2015 | ReCBuLC: Reproducing Concurrency Bugs Using Local ClocksabstractMulti-threaded programs play an increasingly important role in current multi-core environments. Exposing concurrency bugs and debugging such multi-threaded programs have become quite challenging due to their inherent non-determinism. In order to eliminate such non-determinism, many approaches such as record-and-replay and other similar bug reproducing systems have been proposed. However, those approaches often suffer significant performance degradation because they require a large amount of recorded information and/or long analysis and replay time. In this paper, we propose an effective approach, ReCBuLC, to take advantage of the hardware clocks available on modern processors. The key idea is to reduce the recording overhead and analyzing events' global order by using time stamps recorded in each thread. Those timestamps are used to determine the global orders of shared accesses. To avoid the large overhead incurred in accessing system-wide global clock, we opt to use local per-core clocks that incur much less access overhead. We then propose techniques to resolve differences among local clocks and obtain an accurate global event order. By using per-core clocks, state-of-the-art bug reproducing systems such as PRES and CLAP can reduce the recording overheads by 1% ~ 85%, and the analysis time by 84.66% ~ 99.99%, respectively. Chenggang Wu 0002, Zhenjiang Wang, Pen-Chung Yew, Jeff Huang 0001, Xiaobing Feng 0002, Yanyan Lan, Yunji Chen |
ICSE (1) | 2 |
| 2015 | HSPT: Practical Implementation and Efficient Management of Embedded Shadow Page Tables for Cross-ISA System Virtual MachinesabstractCross-ISA (Instruction Set Architecture) system-level virtual machine has a significant research and practical value. For example, several recently announced virtual smart phones for iOS which run smart phone applications on x86 based PCs are deployed on cross-ISA system level virtual machines. Also, for mobile device application development, by emulating the Android/ARM environment on the more powerful x86-64 platform, application development and debugging become more convenient and productive. However, the virtualization layer often incurs high performance overhead. The key overhead comes from memory virtualization where a guest virtual address (GVA) must go through multi-level address translation to become a host physical address (HPA). The Embedded Shadow Page Table (ESPT) approach has been proposed to effectively decrease this address translation cost. ESPT directly maps GVA to HPA, thus avoid the lengthy guest virtual to guest physical, guest physical to host virtual, and host virtual to host physical address translation. However, the original ESPT work has a few drawbacks. For example, its implementation relies on a loadable kernel module (LKM) to manage the shadow page table. Using LKMs is less desirable for system virtual machines due to portability, security and maintainability concerns. Our work proposes a different, yet more practical, implementation to address the shortcomings. Instead of relying on using LKMs, our approach adopts a shared memory mapping scheme to maintain the shadow page table (SPT) using only ''mmap'' system call. Furthermore, this work studies the support of SPT for multi-processing in greater details. It devices three different SPT organizations and evaluates their strength and weakness with standard and real Android applications on the system virtual machine which emulates the Android/ARM platform on x86-64 systems. Zhe Wang 0017, Chenggang Wu 0002, Dongyan Yang, Zhenjiang Wang, Wei-Chung Hsu |
VEE | 3 |
| 2015 | FPS: A Fair-Progress Process Scheduling Policy on Shared-Memory MultiprocessorsabstractCompetition for shared memory resources on multiprocessors is the dominant cause for slowing down applications and making their performance varies unpredictably. It exacerbates the need for Quality of Service (QoS) on such systems. In this paper, we propose a fair-progress process scheduling (FPS) policy to improve system fairness. The strategy is to force the equally-weighted applications to bear the same amount of slowdown when they run concurrently. When we find an application suffered more slowdown and accumulated less effective work than others, we allocate more CPU time to give it a better parity. This policy can also be applied to threads with different weights. Evaluation results show that FPS can significantly improve system fairness at the expense of a slight loss in throughput. We can also keep the performance information of an application to guide process scheduling when it runs again later on. When FPS uses such performance information from previous runs, fairness can be maintained without the overhead of the training periods required in FPS. Throughput can thus be enhanced. Chenggang Wu 0002, Pen-Chung Yew, Zhenjiang Wang |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2014 | Dynamic and Adaptive Calling Context Encoding
Zhenjiang Wang, Chenggang Wu 0002, Wei-Chung Hsu |
CGO | 3 |
| 2014 | Localization of concurrency bugs using shared memory access pairsabstractWe propose an effective approach to automatically localize buggy shared memory accesses that trigger concurrency bugs. Compared to existing approaches, our approach has two advantages. First, as long as enough successful runs of a concurrent program are collected, our approach can localize buggy shared memory accesses even with only one single failed run captured, as opposed to the requirement of capturing multiple failed runs in existing approaches. This is a significant advantage because it is more difficult to capture the elusive failed runs than the successful runs in practice. Second, our approach exhibits more precise bug localization results because it also captures buggy shared memory accesses in those failed runs that terminate prematurely, which are often neglected in existing approaches. Based on this proposed approach, we also implement a prototype, named LOCON. Evaluation results on 16 common concurrency bugs show that all buggy shared memory accesses that trigger these bugs can be precisely localized by LOCON with only one failed run captured. Wenwen Wang 0001, Zhenjiang Wang, Chenggang Wu 0002, Pen-Chung Yew, Xipeng Shen, Xiaobing Feng 0002 |
ASE | 3 |
| 2014 | Concurrency bug localization using shared memory access pairsabstractNon-determinism in concurrent programs makes their debugging much more challenging than that in sequential programs. To mitigate such difficulties, we propose a new technique to automatically locate buggy shared memory accesses that triggered concurrency bugs. Compared to existing fault localization techniques that are based on empirical statistical approaches, this technique has two advantages. First, as long as enough successful runs of a concurrent program are collected, the proposed technique can locate buggy memory accesses to the shared data even with only one single failed run captured, as opposed to the need of capturing multiple failed runs in other statistical approaches. Second, the proposed technique is more precise because it considers memory accesses in those failed runs that terminate prematurely. Wenwen Wang 0001, Chenggang Wu 0002, Pen-Chung Yew, Zhenjiang Wang, Xiaobing Feng 0002 |
PPoPP | 2 |
| 2014 | Dynamic I/O-Aware Scheduling for Batch-Mode Applications on Chip Multiprocessor Systems of Cluster Platforms
Huimin Cui, Lei Wang 0004, Lei Liu 0030, Chenggang Wu 0002, Xiaobing Feng 0002, Pen-Chung Yew |
J. Comput. Sci. Technol. | 5 |
| 2013 | Synchronization Identification through On-the-Fly Test
Zhenjiang Wang, Chenggang Wu 0002, Pen-Chung Yew, Wenwen Wang 0001 |
Euro-Par | 3 |
| 2012 | Providing fairness on shared-memory multiprocessors via process schedulingabstractCompetition for shared memory resources on multiprocessors is the most dominant cause for slowing down applications and makes their performance varies unpredictably. It exacerbates the need for Quality of Service (QoS) on such systems. In this paper, we propose a fair-progress process scheduling (FPS) policy to improve system fairness. Its strategy is to force the equally-weighted applications to have the same amount of slowdown when they run concurrently. The basic approach is to monitor the progress of all applications at runtime. When we find an application suffered more slowdown and accumulated less effective work than others, we allocate more CPU time to give it a better parity. Our policy also allows different weights to different threads, and provides an effective and robust tuner that allows the OS to freely make tradeoffs between system fairness and higher throughput. Evaluation results show that FPS can significantly improve system fairness by an average of 53.5% and 65.0% on a 4-core processor with a private cache and a 4-core processor with a shared cache, respectively. The penalty is about 1.1% and 1.6% of the system throughput. For memory-intensive workloads, FPS also improves system fairness by an average of 45.2% and 21.1% on 4-core and 8-core system respectively at the expense of a throughput loss of about 2%. Chenggang Wu 0002, Pen-Chung Yew, Zhenjiang Wang |
SIGMETRICS | 2 |
| 2012 | On-the-fly structure splitting for heap objectsabstractWith the advent of multicore systems, the gap between processor speed and memory latency has grown worse because of their complex interconnect. Sophisticated techniques are needed more than ever to improve an application's spatial and temporal locality. This paper describes an optimization that aims to improve heap data layout by structure-splitting. It also provides runtime address checking by piggybacking on the existing page protection mechanism to guarantee the correctness of such optimization that has eluded many previous attempts due to safety concerns. The technique can be applied to both sequential and parallel programs at either compile time or runtime. However, we focus primarily on sequential programs (i.e., single-threaded programs) at runtime in this paper. Experimental results show that some benchmarks in SPEC 2000 and 2006 can achieve a speedup of up to 142.8%. Zhenjiang Wang, Chenggang Wu 0002, Pen-Chung Yew |
ACM Trans. Archit. Code Optim. | 2 |
| 2011 | Dynamic register promotion of stack variablesabstractDynamic Binary Translation (DBT) has been widely used in various applications. Although new architectures and micro-architectures often create performance opportunities for programmers and compilers, such performance opportunities may not be exploited by legacy executables. For example, the additional general-purpose and XMM registers in the Intel64 architecture do not benefit the IA-32 binaries. In this paper, we designed and developed a DBT system to dynamically promote stack variables in the source binaries to the additional registers of the target architecture. One of the most challenging problems is how to deal with the possible but rare memory aliases between promoted stack variables and other implicit memory references. We devised a runtime alias detection approach based on the page protection mechanism in Linux and a novel stack switching method to catch memory aliases at run-time. This approach is much less expensive than traditional approaches like inserting address checking instructions. On an Intel64 platform, our DBT system with speculative stack variable promotion has sped up several SPEC CPU2006 benchmarks in IA-32 code, with the largest performance gain over 45%. Chenggang Wu 0002, Wei-Chung Hsu |
CGO | 2 |
| 2011 | Efficient and effective misaligned data access handling in a dynamic binary translation systemabstractBinary Translation (BT) has been commonly used to migrate application software across Instruction Set Architectures (ISAs). Some architectures, such as X86, allow Misaligned Data Accesses (MDAs), while most modern architectures require natural data alignments. In a binary translation system, where the source ISA allows MDA and the target ISA does not, memory operations must be carefully translated. Naive translation may cause frequent misaligned data access traps to occur at runtime on the target machine and severely slow down the migrated application. This article evaluates different approaches in handling MDA in a binary translation system including how to identify MDA candidates and how to translate such memory instructions. This article also proposes some new mechanisms to more effectively deal with MDAs. Extensive measurements based on SPEC CPU2000 and CPU2006 benchmarks show that the proposed approaches are more effective than existing methods and getting close to the performance upper bound of MDA handling. Chenggang Wu 0002, Wei-Chung Hsu |
ACM Trans. Archit. Code Optim. | 2 |
| 2010 | On mitigating memory bandwidth contention through bandwidth-aware schedulingabstractShared-memory multiprocessors have dominated all platforms from high-end to desktop computers. On such platforms, it is well known that the interconnect between the processors and the main memory has become a major bottleneck. The bandwidth-aware job scheduling is an effective and relatively easy-to-implement way to relieve the bandwidth contention. Previous policies understood that bandwidth saturation hurt the throughput of parallel jobs so they scheduled the jobs to let the total bandwidth requirement equal to the system peak bandwidth. However, we found that intra-quantum fine-grained bandwidth contention still happened due to a program's irregular fluctuation in memory access intensity, which is mostly ignored in previous policies. Chenggang Wu 0002, Pen-Chung Yew |
PACT | 2 |
| 2010 | On improving heap memory layout by dynamic pool allocationabstractDynamic memory allocation is widely used in modern programs. General-purpose heap allocators often focus more on reducing their run-time overhead and memory space utilization, but less on exploiting the characteristics of their allocated heap objects. This paper presents a lightweight dynamic optimizer, named Dynamic Pool Allocation (DPA), which aims to exploit the affinity of the allocated heap objects and improve their layout at run-time. DPA uses an adaptive partial call chain with heuristics to aggregate affinitive heap objects into dedicated memory regions, called memory pools. We examine the factors that could affect the effectiveness of such layout. We have implemented DPA and measured its performance on several SPEC CPU 2000 and 2006 benchmarks that use extensive heap objects. Evaluations show that it could achieve an average speed up of 12.1% and 10.8% on two x86 commodity machines respectively using GCC -O3, and up to 82.2% for some benchmarks. Zhenjiang Wang, Chenggang Wu 0002, Pen-Chung Yew |
CGO | 2 |
| 2009 | An Evaluation of Misaligned Data Access Handling Mechanisms in Dynamic Binary Translation SystemsabstractBinary translation (BT) has been an important approach to migrate application software across instruction set architectures (ISAs). Some architectures, such as X86, allow misaligned data accesses (MDAs), while most modern architectures have the alignment restriction that requires data to be aligned in memory on natural boundaries. In a binary translation system, where the source ISA allows MDA and the target ISA does not, memory operations must be carefully translated to satisfy the alignment restriction. Naive translation will cause frequent misaligned data access traps to occur at runtime on the target machine, and severely slow down the migrated application.This paper evaluates different approaches in handling MDA in binary translation systems. It also proposes a new mechanism to deal with MDAs. Measurements based on SPEC CPU2000 and CPU2006 benchmark show that the proposed approach can significantly outperform existing methods. Chenggang Wu 0002, Wei-Chung Hsu |
CGO | 2 |