Zhufan Wang

dblp:167/9591 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
0since 2021 · last 2020
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 1 first-authorDatabases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Storage systems · 52% Parallel and multicore computing · 16% Hardware accelerators and domain-specific architectures · 16%
Artificial intelligence
1 paper
Deep learning architectures and training · 100%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Parallel and multicore computing
data distribution
0.822020
Determining Data Distribution for Large Disk Enclosures with 3-D Data Templates · ACM Trans. Storage 2020
RAID+: Deterministic and Balanced Data Distribution for Large Disk Enclosures · FAST 2018
Storage systems › storage reliability
RAID
0.822020
Determining Data Distribution for Large Disk Enclosures with 3-D Data Templates · ACM Trans. Storage 2020
RAID+: Deterministic and Balanced Data Distribution for Large Disk Enclosures · FAST 2018
Storage systems › repair
data reconstruction
0.412020
Determining Data Distribution for Large Disk Enclosures with 3-D Data Templates · ACM Trans. Storage 2020
Hardware accelerators and domain-specific architectures › machine learning accelerator › CNN accelerator
CNN training accelerator
0.412019
HyConv: Accelerating Multi-Phase CNN Computation by Fine-Grained Policy Selection · IEEE Trans. Parallel Distributed Syst. 2019
Storage systems › storage reliability
data recovery
0.412019
Dayu: Fast and Low-interference Data Recovery in Very-large Storage Systems · USENIX ATC 2019
Distributed systems
fault tolerance
0.412019
Dayu: Fast and Low-interference Data Recovery in Very-large Storage Systems · USENIX ATC 2019
GPUs and heterogeneous computing
GPU computing
0.412019
HyConv: Accelerating Multi-Phase CNN Computation by Fine-Grained Policy Selection · IEEE Trans. Parallel Distributed Syst. 2019
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.412019
HyConv: Accelerating Multi-Phase CNN Computation by Fine-Grained Policy Selection · IEEE Trans. Parallel Distributed Syst. 2019
Storage systems
storage reliability
0.412019
Dayu: Fast and Low-interference Data Recovery in Very-large Storage Systems · USENIX ATC 2019
Storage systems › magnetic storage
disk storage
0.312018
RAID+: Deterministic and Balanced Data Distribution for Large Disk Enclosures · FAST 2018
Machine learning › Deep learning architectures and training
convolutional neural network
0.112019
HyConv: Accelerating Multi-Phase CNN Computation by Fine-Grained Policy Selection · IEEE Trans. Parallel Distributed Syst. 2019
Machine learning › Deep learning architectures and training › convolutional neural network
convolutional neural network training
0.112019
HyConv: Accelerating Multi-Phase CNN Computation by Fine-Grained Policy Selection · IEEE Trans. Parallel Distributed Syst. 2019

Methods — techniques the papers use, named apart from their topics

winner database · 0.8runtime policy selection · 0.8mutually orthogonal latin squares · 0.4deterministic addressing · 0.4low-interference data recovery · 0.4data distribution · 0.3
YearPublicationVenuePosition
2020 Determining Data Distribution for Large Disk Enclosures with 3-D Data Templates
abstract
Conventional RAID solutions with fixed layouts partition large disk enclosures so that each RAID group uses its own disks exclusively. This achieves good performance isolation across underlying disk groups, at the cost of disk under-utilization and slow RAID reconstruction from disk failures. We propose RAID+, a new RAID construction mechanism that spreads both normal I/O and reconstruction workloads to a larger disk pool in a balanced manner. Unlike systems conducting randomized placement, RAID+ employs deterministic addressing enabled by the mathematical properties of mutually orthogonal Latin squares, based on which it constructs 3-D data templates mapping a logical data volume to uniformly distributed disk blocks across all disks. While the total read/write volume remains unchanged, with or without disk failures, many more disk drives participate in data service and disk reconstruction. Our evaluation with a 60-drive disk enclosure using both synthetic and real-world workloads shows that RAID+ significantly speeds up data recovery while delivering better normal I/O performance and higher multi-tenant system throughput.
Guangyan Zhang, Zhufan Wang, Xiaosong Ma, Zican Huang
ACM Trans. Storage2
2019 Dayu: Fast and Low-interference Data Recovery in Very-large Storage Systems
Zhufan Wang, Guangyan Zhang, Yang Wang 0009, Qinglin Yang, Jiaji Zhu
USENIX ATC1
2019 HyConv: Accelerating Multi-Phase CNN Computation by Fine-Grained Policy Selection
abstract
Existing GPU-based approaches cannot yet meet the performance requirement for training very large convolutional neural networks (CNNs), where convolutional layers (Conv-layers) dominate the training time. In this paper, we find that no single convolution policy can always perform the fastest across all the computing phases. Then, we propose an approach called HyConv to accelerating multi-phase CNN computation by fine-grained policy selection. HyConv encapsulates existing convolution policies into a set of modules, and selects the fastest policy (a.k.a., winner policy) via one-round runtime measurement for computing each phase. Furthermore, HyConv uses a winner database to record the current winner policies, avoiding duplicate measurement later for the same parameter configuration. Our experimental results indicate that over all the used real-world CNN networks, HyConv consistently outperforms existing approaches on either a single GPU or four GPUs, with speedups of up to 3.3× and up to 1.6× over cuDNN-MM respectively. Such improvement can be explained by our result that HyConv delivers obviously better performance for most of single Conv-layers. Furthermore, HyConv has the ability to work with any parameter configuration and thus keeps better usability.
Xiaqing Li, Guangyan Zhang, Zhufan Wang
IEEE Trans. Parallel Distributed Syst.3
2018 RAID+: Deterministic and Balanced Data Distribution for Large Disk Enclosures
Guangyan Zhang, Zican Huang, Xiaosong Ma, Zhufan Wang
FAST5
2016 Performance Analysis of GPU-Based Convolutional Neural Networks
abstract
As one of the most important deep learning models, convolutional neural networks (CNNs) have achieved great successes in a number of applications such as image classification, speech recognition and nature language understanding. Training CNNs on large data sets is computationally expensive, leading to a flurry of research and development of open-source parallel implementations on GPUs. However, few studies have been performed to evaluate the performance characteristics of those implementations. In this paper, we conduct a comprehensive comparison of these implementations over a wide range of parameter configurations, investigate potential performance bottlenecks and point out a number of opportunities for further optimization.
Xiaqing Li, Guangyan Zhang, H. Howie Huang, Zhufan Wang
ICPP4