Zhuofang Dai

dblp:154/3429 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
0since 2021 · last 2016
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
1 paper
Concurrent programming · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
GPUs and heterogeneous computing · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Concurrent programming › concurrency bug detection
atomicity violation detection
0.212016
Hardware Support for Concurrent Detection of Multiple Concurrency Bugs on Fused CPU-GPU Architectures · IEEE Trans. Computers 2016
Concurrent programming
concurrency bug detection
0.212016
Hardware Support for Concurrent Detection of Multiple Concurrency Bugs on Fused CPU-GPU Architectures · IEEE Trans. Computers 2016
Concurrent programming › concurrency bug detection
data race detection
0.212016
Hardware Support for Concurrent Detection of Multiple Concurrency Bugs on Fused CPU-GPU Architectures · IEEE Trans. Computers 2016
GPUs and heterogeneous computing › CPU-GPU heterogeneous computing
fused CPU-GPU architecture
0.112016
Hardware Support for Concurrent Detection of Multiple Concurrency Bugs on Fused CPU-GPU Architectures · IEEE Trans. Computers 2016

Methods — techniques the papers use, named apart from their topics

trace collection · 0.5happens-before relation · 0.5bloom filter · 0.5
YearPublicationVenuePosition
2016 Hardware Support for Concurrent Detection of Multiple Concurrency Bugs on Fused CPU-GPU Architectures
abstract
Detecting concurrency bugs, such as data race, atomicity violation and order violation, is a cumbersome task for programmers. This situation is further being exacerbated due to the increasing number of cores in a single machine and the prevalence of threaded programming models. Unfortunately, many existing software-based approaches usually incur high runtime overhead or accuracy loss, while most hardware-based proposals usually focus on a specific type of bugs and thus are inflexible to detect a variety of concurrency bugs. In this paper, we propose Hydra, an approach that leverages massive parallelism and programmability of fused CPU-GPU architectures to simultaneously detect multiple concurrency bugs in threaded software, including data race, atomicity violation and order violation. Hydra extends contemporary fused CPU and GPU by introducing two modules: 1) a trace collecting module (TCM) that instruments and collects program behavior on CPU; 2) a trace preprocessing module (TPM) that processes and then transfers the traces to GPU for bug detection. Furthermore, Hydra exploits three optimizations to improve speed and accuracy, which includes: 1). using the bloom filter to filter out unnecessary traces; 2). avoiding eviction of shared traces; 3). comparing only last-write traces for shared data with the happens-before relation. Hydra incurs small hardware complexity and requires no changes to internal critical-path processor components such as cache and its coherence protocol, and is with about 1.1 percent hardware overhead under a 32-core configuration. Experimental results show that Hydra only introduces about 0.18 percent overhead on average for detecting one type of bugs and 0.46 percent overhead for simultaneously detecting multiple bugs, yet with the similar detectability of a heavyweight software bug detector (e.g., Helgrind).
Shiqiang Yu, Haojun Wang, Zhuofang Dai, Haibo Chen 0001
IEEE Trans. Computers4
2014 Hydra: Efficient Detection of Multiple Concurrency Bugs on Fused CPU-GPU Architecture
abstract
Detecting concurrency bugs, such as data race, atomicity violation and order violation, is a cumbersome task for programmers. This situation is further being exacerbated due to the increasing number of cores in a single machine and the prevalence of threaded programming models. Unfortunately, many existing software-based approaches usually incur high runtime overhead or accuracy loss, while most hardware-based proposals usually focus on a specific type of bugs and thus are inflexible to detect a variety of concurrency bugs. In this paper, we propose Hydra, an approach that leverages massive parallelism and programmability of fused GPU architecture to simultaneously detect multiple types of concurrency bugs, including data race, atomicity violation and order violation. Hydra instruments and collects program behavior on CPU and transfers the traces to GPU for bug detection through on-chip interconnect. Furthermore, to achieve high speed, Hydra exploits bloom filter to filter out unnecessary detection traces. Hydra incurs small hardware complexity and requires no changes to internal critical-path processor components such as cache and its coherence protocol, and is with about 1.1% hardware overhead under a 32-core configuration. Experimental results show that Hydra only introduces about 0.35% overhead on average for detecting one type of bugs and 0.92% overhead for simultaneously detecting multiple bugs, yet with the similar detectability of a heavyweight software bug detector (e.g., Helgrind).
Zhuofang Dai, Haojun Wang, Haibo Chen 0001, Binyu Zang
ICPP1