Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Lipeng Wang 0004

dblp:07/10051-4 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
1since 2021 · last 2022
0000-0001-7918-4786ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 4 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Storage systems · 40% Parallel and multicore computing · 34% Distributed systems · 20%
Databases, data mining, and information retrieval
1 paper
Graph data management · 100%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed systems
distributed caching
0.612022
DIESEL+: Accelerating Distributed Deep Learning Tasks on Image Datasets · IEEE Trans. Parallel Distributed Syst. 2022
Parallel and multicore computing
distributed deep learning training
0.612022
DIESEL+: Accelerating Distributed Deep Learning Tasks on Image Datasets · IEEE Trans. Parallel Distributed Syst. 2022
Storage systems
file systems
0.612022
DIESEL+: Accelerating Distributed Deep Learning Tasks on Image Datasets · IEEE Trans. Parallel Distributed Syst. 2022
Storage systems
i/o optimization
0.612022
DIESEL+: Accelerating Distributed Deep Learning Tasks on Image Datasets · IEEE Trans. Parallel Distributed Syst. 2022
Graph data management › graph pattern matching › subgraph matching
subgraph enumeration
0.412019
Efficient Parallel Subgraph Enumeration on a Single Machine · ICDE 2019
Parallel and multicore computing › parallel programming models
shared-memory parallelization
0.412019
Efficient Parallel Subgraph Enumeration on a Single Machine · ICDE 2019

Methods — techniques the papers use, named apart from their topics

simultaneous multithreading · 0.8minimum set cover · 0.8depth-first search · 0.8SIMD · 0.8region-of-interest decoding · 0.6chunk-wise shuffling · 0.6
YearPublicationVenuePosition
2022 DIESEL+: Accelerating Distributed Deep Learning Tasks on Image Datasets
abstract
We observe that data access and processing takes a significant amount of time in large-scale deep learning training tasks (DLTs) on image datasets. Three factors contribute to this problem: (1) the massive and recurrent accesses to large numbers of small files; (2) the repeated, expensive decoding computation on each image, and (3) the frequent communication between computation nodes and storage nodes. Existing work has addressed some aspects of these problems; however, no end-to-end solutions have been proposed. In this article, we propose DIESEL+, an all-in-one system which accelerates the entire I/O pipeline of deep learning training tasks. DIESEL+ contains several components: (1) local metadata snapshot; (2) per-task distributed caching; (3) chunk-wise shuffling; (4) GPU-assisted image decoding and (5) online region-of-interest (ROI) decoding. The metadata snapshot removes the bottleneck on metadata access in frequent reading of large numbers of files. The per-task distributed cache across the worker nodes of a DLT task to reduce the I/O pressure on the underlying storage. The chunk-based shuffle method converts small file reads into large chunk reads, so that the performance is improved without sacrificing the training accuracy. The GPU-assisted image decoding and the online ROI method minimize the image decoding workloads and reduce the cost of data movement between nodes. These techniques are seamlessly integrated into the system. In our experiments, DIESEL+ outperforms existing systems by a factor of two to three times on the overall training time.
Lipeng Wang 0004, Qiong Luo 0001, Shengen Yan
IEEE Trans. Parallel Distributed Syst.1
2020 Accelerating Deep Learning Tasks with Optimized GPU-assisted Image Decoding
abstract
In computer vision deep learning (DL) tasks, most of the input image datasets are stored in the JPEG format. These JPEG datasets need to be decoded before DL tasks are performed on them. We observe two problems in the current JPEG decoding procedures for DL tasks: (1) the decoding of image entropy data in the decoder is performed sequentially, and this sequential decoding repeats with the DL iterations, which takes significant time; (2) Current parallel decoding methods under-utilize the massive hardware threads on GPUs. To reduce the image decoding time, we introduce a pre-scan mechanism to avoid the repeated image scanning in DL tasks. Our pre-scan generates boundary markers for entropy data so that the decoding can be performed in parallel. To cooperate with the existing dataset storage and caching systems, we propose two modes of the pre-scan mechanism: a compatible mode and a fast mode. The compatible mode does not change the image file structure so pre-scanned files can be stored back to disk for subsequent DL tasks. In comparison, the fast mode crafts a JPEG image into a binary format suitable for parallel decoding, which can be processed directly on the GPU. Since the GPU has thousands of hardware threads, we propose a fine-grained parallel decoding method on the pre-scanned dataset. The fine-grained parallelism utilizes the GPU effectively, and achieves speedups of around 1.5× over existing GPU-assisted image decoding libraries on real-world DL tasks.
Lipeng Wang 0004, Qiong Luo 0001, Shengen Yan
ICPADS1
2020 DIESEL: A Dataset-Based Distributed Storage and Caching System for Large-Scale Deep Learning Training
abstract
We observe three problems in existing storage and caching systems for deep-learning training (DLT) tasks: (1) accessing a dataset containing a large number of small files takes a long time, (2) global in-memory caching systems are vulnerable to node failures and slow to recover, and (3) repeatedly reading a dataset of files in shuffled orders is inefficient when the dataset is too large to be cached in memory. Therefore, we propose DIESEL, a dataset-based distributed storage and caching system for DLT tasks. Our approach is via a storage-caching system co-design. Firstly, since accessing small files is a metadata-intensive operation, DIESEL decouples the metadata processing from metadata storage, and introduces metadata snapshot mechanisms for each dataset. This approach speeds up metadata access significantly. Secondly, DIESEL deploys a task-grained distributed cache across the worker nodes of a DLT task. This way node failures are contained within each DLT task. Furthermore, the files are grouped into large chunks in storage, so the recovery time of the caching system is reduced greatly. Thirdly, DIESEL provides chunk-based shuffle so that the performance of random file access is improved without sacrificing training accuracy. Our experiments show that DIESEL achieves a linear speedup on metadata access, and outperforms an existing distributed caching system in both file caching and file reading. In real DLT tasks, DIESEL halves the data access time of an existing storage system, and reduces the training time by hours without changing any training code.
Lipeng Wang 0004, Songgao Ye, Baichen Yang, Youyou Lu, Hequan Zhang, Shengen Yan, Qiong Luo 0001
ICPP1
2019 Efficient Parallel Subgraph Enumeration on a Single Machine
abstract
Subgraph enumeration finds all subgraphs in an unlabeled graph that are isomorphic to another unlabeled graph. Existing depth-first search (DFS) based algorithms work on a single machine, but they are slow on large graphs due to the large search space. In contrast, distributed algorithms on clusters adopt a parallel breadth-first search (BFS) and improve the performance at the cost of large amounts of hardware resources, since the BFS approach incurs expensive data transfer and space cost due to the exponential number of intermediate results. In this paper, we develop an efficient parallel subgraph enumeration algorithm for a single machine, named LIGHT. Our algorithm reduces redundant computation in DFS by deferring the materialization of pattern vertices until necessary and converting the candidate set computation into finding a minimum set cover. Moreover, we parallelize our algorithm with both SIMD (Single-Instruction-Multiple-Data) instructions and SMT (Simultaneous Multi-Threading) technologies in modern CPUs. Our experimental results show that LIGHT running on a single machine outperforms existing single-machine DFS algorithms by more than three orders of magnitude, and is up to two orders of magnitude faster than the state-of-the-art distributed algorithms running on 12 machines. Additionally, LIGHT completed all test cases, whereas the existing algorithms fail in some cases due to either running out of time or running out of available hardware resources.
Shixuan Sun, Yulin Che, Lipeng Wang 0004, Qiong Luo 0001
ICDE3
2019 Accelerating Long Read Alignment on Three Processors
abstract
Sequence alignment is a fundamental task in bioinformatics, because many downstream applications rely on it. The recent emergence of the third-generation sequencing technology requires new sequence alignment algorithms that handle longer read lengths as well as more sequencing errors. Furthermore, the rapidly increasing volume of sequence data calls for efficient analysis solutions. To address this need, we propose to utilize commodity parallel processors to perform the long read alignment. Specifically, we propose manymap, an acceleration of the leading CPU-based long read aligner minimap2 on the CPU, the GPU, and the Intel Xeon Phi processor. We eliminate intra-loop data dependency in the base-level alignment step of the original minimap2 through redesigning memory layouts of dynamic programming (DP) matrices. This change facilitates the effective vectorization of the most time-consuming procedure in alignment. Additionally, we apply architecture-aware optimizations, such as utilizing high bandwidth memory on Xeon Phi and concurrent kernel execution on GPU. We evaluate our manymap in comparison with the extended minimap2 on a Xeon Gold 5115 CPU, a Tesla V100 GPU, and a Xeon Phi 7210 processor. Our results show that manymap outperforms minimap2 by up to 2.3 times on the overall execution time and 4.5 times on the base-level alignment step.
Zonghao Feng, Lipeng Wang 0004, Qiong Luo 0001
ICPP3
2017 Betweenness Centrality Revisited on Four Processors
abstract
The betweenness centrality measure has been widely adopted in various graph analytics applications, such as community detection and brain network analysis. Due to the high intensity of BC computation and rapid data growth, there have been a number of studies on parallel BC computation, either on CPUs or GPUs. However, there has not been a comprehensive comparative study on the BC algorithm on different processors. In this paper, we revisit shared-memory parallel BC computation on four kinds of processors, including multi-core CPUs, many-core GPUs, and two generations of Intel MIC processors. We find that, with suitable parallelization strategies and data-oriented optimizations, commodity multi-core CPUs are the fastest, followed by the second generation MIC. These two processors are faster than the state-of-the-art GPU implementations across all kinds of graphs. In comparison, the GPU outperforms the first generation MIC only on small-diameter graphs and is the slowest on the other kinds of graphs.
Lipeng Wang 0004, Xiaoying Jia 0001, Qiong Luo 0001
ICPADS1