Zhenjiang Wang

dblp:00/8025 · DBLP profile ↗
← Back
15ranked-venue papers
3as first author
1since 2021 · last 2025
0000-0001-7783-6648ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 3 first-author · 1 since 2021Software engineering, systems software and programming languages · 7 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Embedded and real-time systems · 58% Parallel and multicore computing · 25% Performance modeling and evaluation · 7%
Software engineering, system software, and programming languages
6 papers
Debugging and program repair · 49% Concurrent programming · 40% Operating systems · 5%

Topics — the 21 heaviest of 23, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Concurrent programming
concurrency bugs
0.942018
Using Local Clocks to Reproduce Concurrency Bugs · IEEE Trans. Software Eng. 2018
ReCBuLC: Reproducing Concurrency Bugs Using Local Clocks · ICSE (1) 2015
Concurrency bug localization using shared memory access pairs · PPoPP 2014
Embedded and real-time systems › real-time scheduling
multiprocessor scheduling
0.622019
Resource-Aware Scheduling for Dependable Multicore Real-Time Systems: Utilization Bound and Partitioning Algorithm · IEEE Trans. Parallel Distributed Syst. 2019
FPS: A Fair-Progress Process Scheduling Policy on Shared-Memory Multiprocessors · IEEE Trans. Parallel Distributed Syst. 2015
Debugging and program repair
fault localization
0.422014
Concurrency bug localization using shared memory access pairs · PPoPP 2014
Localization of concurrency bugs using shared memory access pairs · ASE 2014
Embedded and real-time systems › real-time scheduling
fault-tolerant real-time scheduling
0.412019
Resource-Aware Scheduling for Dependable Multicore Real-Time Systems: Utilization Bound and Partitioning Algorithm · IEEE Trans. Parallel Distributed Syst. 2019
Embedded and real-time systems › real-time scheduling › multiprocessor scheduling
partitioned-EDF scheduling
0.412019
Resource-Aware Scheduling for Dependable Multicore Real-Time Systems: Utilization Bound and Partitioning Algorithm · IEEE Trans. Parallel Distributed Syst. 2019
Embedded and real-time systems
real-time scheduling
0.412019
Resource-Aware Scheduling for Dependable Multicore Real-Time Systems: Utilization Bound and Partitioning Algorithm · IEEE Trans. Parallel Distributed Syst. 2019
Parallel and multicore computing
task partitioning
0.412019
Resource-Aware Scheduling for Dependable Multicore Real-Time Systems: Utilization Bound and Partitioning Algorithm · IEEE Trans. Parallel Distributed Syst. 2019
Debugging and program repair
bug reproduction
0.312018
Using Local Clocks to Reproduce Concurrency Bugs · IEEE Trans. Software Eng. 2018
Debugging and program repair
record and replay
0.322018
ReCBuLC: Reproducing Concurrency Bugs Using Local Clocks · ICSE (1) 2015
Using Local Clocks to Reproduce Concurrency Bugs · IEEE Trans. Software Eng. 2018
Concurrent programming › concurrency bugs
concurrency bug reproduction
0.212015
ReCBuLC: Reproducing Concurrency Bugs Using Local Clocks · ICSE (1) 2015
Parallel and multicore computing › task scheduling
process scheduling
0.212015
FPS: A Fair-Progress Process Scheduling Policy on Shared-Memory Multiprocessors · IEEE Trans. Parallel Distributed Syst. 2015
Performance modeling and evaluation
workload characterization
0.212015
FPS: A Fair-Progress Process Scheduling Policy on Shared-Memory Multiprocessors · IEEE Trans. Parallel Distributed Syst. 2015
Parallel and multicore computing › multiprocessor system
shared-memory multiprocessor
0.222015
Providing fairness on shared-memory multiprocessors via process scheduling · SIGMETRICS 2012
FPS: A Fair-Progress Process Scheduling Policy on Shared-Memory Multiprocessors · IEEE Trans. Parallel Distributed Syst. 2015
Debugging and program repair › fault localization
bug localization
0.212014
Localization of concurrency bugs using shared memory access pairs · ASE 2014
Debugging and program repair › fault localization
concurrency bug localization
0.212014
Concurrency bug localization using shared memory access pairs · PPoPP 2014
Operating systems › resource management › process management
CPU scheduling
0.112012
Providing fairness on shared-memory multiprocessors via process scheduling · SIGMETRICS 2012
Compilers and program optimization › memory optimization
data locality optimization
0.112012
On-the-fly structure splitting for heap objects · ACM Trans. Archit. Code Optim. 2012
Memory systems › data layout optimization
structure splitting
0.112012
On-the-fly structure splitting for heap objects · ACM Trans. Archit. Code Optim. 2012
Embedded and real-time systems › real-time scheduling
schedulability analysis
0.112019
Resource-Aware Scheduling for Dependable Multicore Real-Time Systems: Utilization Bound and Partitioning Algorithm · IEEE Trans. Parallel Distributed Syst. 2019
Memory systems › memory interference
memory contention
0.112015
FPS: A Fair-Progress Process Scheduling Policy on Shared-Memory Multiprocessors · IEEE Trans. Parallel Distributed Syst. 2015
Cloud and datacenter computing
quality of service
0.012012
Providing fairness on shared-memory multiprocessors via process scheduling · SIGMETRICS 2012

Methods — techniques the papers use, named apart from their topics

local clock · 0.8timestamps · 0.4simulation · 0.4linux kernel implementation · 0.4timestamp recording · 0.3scheduling policy design · 0.3runtime progress monitoring · 0.3runtime address checking · 0.3page protection · 0.3performance measurement · 0.2statistical fault localization · 0.2shared memory access pairs · 0.2failed run analysis · 0.2
YearPublicationVenuePosition
2025 ZS-Puffin: Design, Modeling and Implementation of an Unmanned Aerial-Aquatic Vehicle with Amphibious Wings
abstract
Unmanned aerial-aquatic vehicles (UAAVs) can operate both in the air and underwater, giving them broad application prospects. Inspired by the dual-function wings of puffins, we propose a UAAV with amphibious wings to address the challenge posed by medium differences on the vehicle’s propulsion system. The amphibious wing, redesigned based on a fixed-wing structure, features a single degree of freedom in pitch and requires no additional components. It can generate lift in the air and function as a flapping wing for propulsion underwater, reducing disturbance to marine life and making it environmentally friendly. Additionally, an artificial central pattern generator (CPG) is introduced to enhance the smoothness of the flapping motion. This paper presents the prototype, design details, and practical implementation of this concept.
Zhenjiang Wang, Yunhua Jiang, Zikun Zhen, Yubin Tan, Wubin Wang
IROS1
2020 Blocking-Aware Partitioned Real-Time Scheduling for Uniform Heterogeneous Multicore Platforms
abstract
Heterogeneous multicore processors have recently become de facto computing engines for state-of-the-art embedded applications. Nonetheless, very little research focuses on the scheduling of periodic (implicit-deadline) real-time tasks upon heterogeneous multicores under the requirements of task synchronization, which is stemmed from resource access conflicts and can greatly affect the schedulability of tasks. In view of partitioned Earliest Deadline First and Multiprocessor Stack Resource Policy, we first discuss the blocking-aware utilization bound for uniform heterogeneous multicores and then illustrate its non-monotonicity, where the bound may decrease with more deployed cores. Following the insights obtained from the bound analysis, taking the system heterogeneity into consideration, we propose a Synchronization-Aware Task Partitioning Algorithm for Heterogeneous Multicores (SA-TPA-HM)). Several resource-guided and heterogeneity-oriented mapping heuristics are incorporated to reduce the negative impacts of blocking interferences for better schedulability performance of tasks and balanced workload distribution across cores. The extensive simulation results show that SA-TPA-HM can obtain the schedulability ratios approximate to an Integer Non-Linear Programming--based solution, and much higher (e.g., 60% more) in contrast to the existing partitioning algorithms targeted at homogeneous multicores. The measurement results in Linux kernel further reveal the practical viability of SA-TPA-HM that can experience lower runtime overhead (e.g., 15% less) when compared to other mapping schemes.
Jian-Jun Han, Sunlu Gong, Zhenjiang Wang, Wen Cai, Dakai Zhu 0001, Laurence T. Yang
ACM Trans. Embed. Comput. Syst.3
2019 Resource-Aware Scheduling for Dependable Multicore Real-Time Systems: Utilization Bound and Partitioning Algorithm
abstract
As the computing devices and software executions are susceptible to manifold faults, fault tolerance has been an important research topic in safety-critical real-time systems. Moreover, multicore processors have recently emerged as prevailing computing engines for modern embedded systems. However, there exists rather rare work on the fault-tolerant scheduling of real-time tasks executing on multicores with shared resources, where the task synchronization originated from resource access contention may significantly degrade the schedulability of task system. With the focus on the partitioned-EDF scheduler with the MSRP (Multiprocessor Stack Resource Policy) protocol and primary/backup recovery mechanism, we first investigate a utilization bound and then identify its anomaly where the bound may decrease when more cores are deployed. Next, following the insights gained by the analysis of the bound, we propose a reliability and synchronization aware task partitioning algorithm (RSA-TPA) together with an efficient version to implement the joint management of task synchronization and system reliability, where several resource-oriented heuristics are developed to improve both the schedulability performance and workload balancing. The extensive simulation results show that the RSA-TPA schemes can obtain higher acceptance ratio (e.g., 60 percent more) and generate more balanced partitions, when compared to the existing schemes that consider either reliability management or task synchronization. Finally, with the different fault arrival rates being considered, the actual implementation in Linux kernel further demonstrates the applicability of RSA-TPA that has lower run-time overhead (e.g., 20 percent less) in comparison with other mapping algorithms.
Jian-Jun Han, Zhenjiang Wang, Sunlu Gong, Tianpeng Miao, Laurence T. Yang
IEEE Trans. Parallel Distributed Syst.2
2018 Using Local Clocks to Reproduce Concurrency Bugs
abstract
Multi-threaded programs play an increasingly important role in current multi-core environments. Exposing concurrency bugs and debugging such multi-threaded programs are quite challenging due to their inherent non-determinism. In order to mitigate such non-determinism, many approaches such as record-and-replay have been proposed. However, those approaches often suffer significant performance degradation because they require a large amount of recorded information and/or long analysis and replay time. In this paper, we propose an efficient and effective approach, ReCBuLC (reproducing concurrency bugs using local clocks), to take advantage of the hardware clocks available on modern processors. The key idea is to reduce the recording overhead and the time to analyze events’ global order by recording timestamps in each thread. These timestamps are used to determine the global order of shared accesses. To avoid the large overhead in accessing system-wide global clock, we opt to use local per-core clocks that incur much less access overhead. We then propose techniques to resolve skews among local clocks and obtain an accurate global event order. By using per-core clocks, state-of-the-art bug reproducing systems such as PRES and CLAP can reduce their recording overheads by up to 85 percent, and the analysis time up to 84.66%$\sim$99.99%, respectively.
Zhe Wang 0017, Chenggang Wu 0002, Zhenjiang Wang, Pen-Chung Yew, Jeff Huang 0001, Xiaobing Feng 0002, Yanyan Lan, Yunji Chen, Yuanming Lai
IEEE Trans. Software Eng.4
2015 ReCBuLC: Reproducing Concurrency Bugs Using Local Clocks
abstract
Multi-threaded programs play an increasingly important role in current multi-core environments. Exposing concurrency bugs and debugging such multi-threaded programs have become quite challenging due to their inherent non-determinism. In order to eliminate such non-determinism, many approaches such as record-and-replay and other similar bug reproducing systems have been proposed. However, those approaches often suffer significant performance degradation because they require a large amount of recorded information and/or long analysis and replay time. In this paper, we propose an effective approach, ReCBuLC, to take advantage of the hardware clocks available on modern processors. The key idea is to reduce the recording overhead and analyzing events' global order by using time stamps recorded in each thread. Those timestamps are used to determine the global orders of shared accesses. To avoid the large overhead incurred in accessing system-wide global clock, we opt to use local per-core clocks that incur much less access overhead. We then propose techniques to resolve differences among local clocks and obtain an accurate global event order. By using per-core clocks, state-of-the-art bug reproducing systems such as PRES and CLAP can reduce the recording overheads by 1% ~ 85%, and the analysis time by 84.66% ~ 99.99%, respectively.
Chenggang Wu 0002, Zhenjiang Wang, Pen-Chung Yew, Jeff Huang 0001, Xiaobing Feng 0002, Yanyan Lan, Yunji Chen
ICSE (1)3
2015 HSPT: Practical Implementation and Efficient Management of Embedded Shadow Page Tables for Cross-ISA System Virtual Machines
abstract
Cross-ISA (Instruction Set Architecture) system-level virtual machine has a significant research and practical value. For example, several recently announced virtual smart phones for iOS which run smart phone applications on x86 based PCs are deployed on cross-ISA system level virtual machines. Also, for mobile device application development, by emulating the Android/ARM environment on the more powerful x86-64 platform, application development and debugging become more convenient and productive. However, the virtualization layer often incurs high performance overhead. The key overhead comes from memory virtualization where a guest virtual address (GVA) must go through multi-level address translation to become a host physical address (HPA). The Embedded Shadow Page Table (ESPT) approach has been proposed to effectively decrease this address translation cost. ESPT directly maps GVA to HPA, thus avoid the lengthy guest virtual to guest physical, guest physical to host virtual, and host virtual to host physical address translation. However, the original ESPT work has a few drawbacks. For example, its implementation relies on a loadable kernel module (LKM) to manage the shadow page table. Using LKMs is less desirable for system virtual machines due to portability, security and maintainability concerns. Our work proposes a different, yet more practical, implementation to address the shortcomings. Instead of relying on using LKMs, our approach adopts a shared memory mapping scheme to maintain the shadow page table (SPT) using only ''mmap'' system call. Furthermore, this work studies the support of SPT for multi-processing in greater details. It devices three different SPT organizations and evaluates their strength and weakness with standard and real Android applications on the system virtual machine which emulates the Android/ARM platform on x86-64 systems.
Zhe Wang 0017, Chenggang Wu 0002, Dongyan Yang, Zhenjiang Wang, Wei-Chung Hsu
VEE5
2015 FPS: A Fair-Progress Process Scheduling Policy on Shared-Memory Multiprocessors
abstract
Competition for shared memory resources on multiprocessors is the dominant cause for slowing down applications and making their performance varies unpredictably. It exacerbates the need for Quality of Service (QoS) on such systems. In this paper, we propose a fair-progress process scheduling (FPS) policy to improve system fairness. The strategy is to force the equally-weighted applications to bear the same amount of slowdown when they run concurrently. When we find an application suffered more slowdown and accumulated less effective work than others, we allocate more CPU time to give it a better parity. This policy can also be applied to threads with different weights. Evaluation results show that FPS can significantly improve system fairness at the expense of a slight loss in throughput. We can also keep the performance information of an application to guide process scheduling when it runs again later on. When FPS uses such performance information from previous runs, fairness can be maintained without the overhead of the training periods required in FPS. Throughput can thus be enhanced.
Chenggang Wu 0002, Pen-Chung Yew, Zhenjiang Wang
IEEE Trans. Parallel Distributed Syst.6
2014 Dynamic and Adaptive Calling Context Encoding
Zhenjiang Wang, Chenggang Wu 0002, Wei-Chung Hsu
CGO2
2014 EATBit: Effective automated test for binary translation with high code coverage
abstract
Binary translation makes it convenient to emulate one instruction set by another. Nowadays, it is growing in popularity in various applications, especially the embedded platforms. When it comes to the test of binary translators, traditional methodologies which still mainly rely on manual unit test is costly, labor intensive and often not adequate to test complicated algorithms in the translators. Some standard benchmark suites, like SPEC CPU2006, are compiled with different compilation options for further tests. However, the translation modules still have over 30% of their code unexecuted after such tests, according to our experimental results. Methodologies based on randomization can generate a vast variety of tests, thus improve the code coverage in the translation system. In this paper, we propose such an approach named EATBit. Test binaries are generated with randomly selected instructions and operands. The binaries and a large amount of input data are then refined to exclude invalid ones. Experimental results on a real binary translator demonstrate that EATBit can not only improve code coverage by over 20%, but also find some new bugs in the translator successfully.
Hui Guo 0007, Zhenjiang Wang, Ruining He
DATE2
2014 Localization of concurrency bugs using shared memory access pairs
abstract
We propose an effective approach to automatically localize buggy shared memory accesses that trigger concurrency bugs. Compared to existing approaches, our approach has two advantages. First, as long as enough successful runs of a concurrent program are collected, our approach can localize buggy shared memory accesses even with only one single failed run captured, as opposed to the requirement of capturing multiple failed runs in existing approaches. This is a significant advantage because it is more difficult to capture the elusive failed runs than the successful runs in practice. Second, our approach exhibits more precise bug localization results because it also captures buggy shared memory accesses in those failed runs that terminate prematurely, which are often neglected in existing approaches. Based on this proposed approach, we also implement a prototype, named LOCON. Evaluation results on 16 common concurrency bugs show that all buggy shared memory accesses that trigger these bugs can be precisely localized by LOCON with only one failed run captured.
Wenwen Wang 0001, Zhenjiang Wang, Chenggang Wu 0002, Pen-Chung Yew, Xipeng Shen, Xiaobing Feng 0002
ASE2
2014 Concurrency bug localization using shared memory access pairs
abstract
Non-determinism in concurrent programs makes their debugging much more challenging than that in sequential programs. To mitigate such difficulties, we propose a new technique to automatically locate buggy shared memory accesses that triggered concurrency bugs. Compared to existing fault localization techniques that are based on empirical statistical approaches, this technique has two advantages. First, as long as enough successful runs of a concurrent program are collected, the proposed technique can locate buggy memory accesses to the shared data even with only one single failed run captured, as opposed to the need of capturing multiple failed runs in other statistical approaches. Second, the proposed technique is more precise because it considers memory accesses in those failed runs that terminate prematurely.
Wenwen Wang 0001, Chenggang Wu 0002, Pen-Chung Yew, Zhenjiang Wang, Xiaobing Feng 0002
PPoPP5
2013 Synchronization Identification through On-the-Fly Test
Zhenjiang Wang, Chenggang Wu 0002, Pen-Chung Yew, Wenwen Wang 0001
Euro-Par2
2012 Providing fairness on shared-memory multiprocessors via process scheduling
abstract
Competition for shared memory resources on multiprocessors is the most dominant cause for slowing down applications and makes their performance varies unpredictably. It exacerbates the need for Quality of Service (QoS) on such systems. In this paper, we propose a fair-progress process scheduling (FPS) policy to improve system fairness. Its strategy is to force the equally-weighted applications to have the same amount of slowdown when they run concurrently. The basic approach is to monitor the progress of all applications at runtime. When we find an application suffered more slowdown and accumulated less effective work than others, we allocate more CPU time to give it a better parity. Our policy also allows different weights to different threads, and provides an effective and robust tuner that allows the OS to freely make tradeoffs between system fairness and higher throughput. Evaluation results show that FPS can significantly improve system fairness by an average of 53.5% and 65.0% on a 4-core processor with a private cache and a 4-core processor with a shared cache, respectively. The penalty is about 1.1% and 1.6% of the system throughput. For memory-intensive workloads, FPS also improves system fairness by an average of 45.2% and 21.1% on 4-core and 8-core system respectively at the expense of a throughput loss of about 2%.
Chenggang Wu 0002, Pen-Chung Yew, Zhenjiang Wang
SIGMETRICS5
2012 On-the-fly structure splitting for heap objects
abstract
With the advent of multicore systems, the gap between processor speed and memory latency has grown worse because of their complex interconnect. Sophisticated techniques are needed more than ever to improve an application's spatial and temporal locality. This paper describes an optimization that aims to improve heap data layout by structure-splitting. It also provides runtime address checking by piggybacking on the existing page protection mechanism to guarantee the correctness of such optimization that has eluded many previous attempts due to safety concerns. The technique can be applied to both sequential and parallel programs at either compile time or runtime. However, we focus primarily on sequential programs (i.e., single-threaded programs) at runtime in this paper. Experimental results show that some benchmarks in SPEC 2000 and 2006 can achieve a speedup of up to 142.8%.
Zhenjiang Wang, Chenggang Wu 0002, Pen-Chung Yew
ACM Trans. Archit. Code Optim.1
2010 On improving heap memory layout by dynamic pool allocation
abstract
Dynamic memory allocation is widely used in modern programs. General-purpose heap allocators often focus more on reducing their run-time overhead and memory space utilization, but less on exploiting the characteristics of their allocated heap objects. This paper presents a lightweight dynamic optimizer, named Dynamic Pool Allocation (DPA), which aims to exploit the affinity of the allocated heap objects and improve their layout at run-time. DPA uses an adaptive partial call chain with heuristics to aggregate affinitive heap objects into dedicated memory regions, called memory pools. We examine the factors that could affect the effectiveness of such layout. We have implemented DPA and measured its performance on several SPEC CPU 2000 and 2006 benchmarks that use extensive heap objects. Evaluations show that it could achieve an average speed up of 12.1% and 10.8% on two x86 commodity machines respectively using GCC -O3, and up to 82.2% for some benchmarks.
Zhenjiang Wang, Chenggang Wu 0002, Pen-Chung Yew
CGO1