Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Lifan Du

dblp:294/2395 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
GPUs and heterogeneous computing · 56% Distributed systems · 20% Cloud and datacenter computing · 20%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed systems
distributed scheduling
0.612022
AEML: An Acceleration Engine for Multi-GPU Load-Balancing in Distributed Heterogeneous Environment · IEEE Trans. Computers 2022
GPUs and heterogeneous computing › GPU resource management
GPU load balancing
0.612022
AEML: An Acceleration Engine for Multi-GPU Load-Balancing in Distributed Heterogeneous Environment · IEEE Trans. Computers 2022
Cloud and datacenter computing › cluster resource management and scheduling › cluster scheduling
heterogeneous cluster scheduling
0.612022
AEML: An Acceleration Engine for Multi-GPU Load-Balancing in Distributed Heterogeneous Environment · IEEE Trans. Computers 2022
GPUs and heterogeneous computing
multi-GPU computing
0.612022
AEML: An Acceleration Engine for Multi-GPU Load-Balancing in Distributed Heterogeneous Environment · IEEE Trans. Computers 2022
GPUs and heterogeneous computing › multi-GPU computing
distributed GPU computing
0.512021
An Incremental Iterative Acceleration Architecture in Distributed Heterogeneous Environments With GPUs for Deep Learning · IEEE Trans. Parallel Distributed Syst. 2021
Parallel and multicore computing
distributed deep learning training
0.112021
An Incremental Iterative Acceleration Architecture in Distributed Heterogeneous Environments With GPUs for Deep Learning · IEEE Trans. Parallel Distributed Syst. 2021

Methods — techniques the papers use, named apart from their topics

task mapping · 0.6stream adjustment · 0.6resource-aware scheduling · 0.6sliding window caching · 0.5GPU memory management · 0.5
YearPublicationVenuePosition
2022 AEML: An Acceleration Engine for Multi-GPU Load-Balancing in Distributed Heterogeneous Environment
abstract
For the rapid growth computation requirements in big data and artificial intelligence area, CPU-GPU heterogeneous clusters can provide more powerful computing capacity compared to CPU clusters. The number of GPUs on single computing node is scalable, which greatly improves the computing capacity of the cluster under the condition of limited cluster size. However, there is a lack of the effective load-balancing scheduling model in multi-GPU hardware environment. This paper proposes AEML, an acceleration engine for multi-GPU load-balancing in distributed heterogeneous environment. AEML can effectively integrate GPUs into distributed processing framework and achieve great load-balance among multiple heterogeneous GPUs. We propose a heterogeneous task execution model based on multiple GPUs and multiple streams (MGMS), which can effectively balance the workload of multiple GPUs. MGMS model utilizes four core techniques: a fine-grained task mapping mechanism, a device resource unified management scheme, a novel resource-aware GPU task scheduling strategy and a feedback-based streams adjustment scheme. The implementation of AEML system is based on Spark 2.4.1 and NVIDIA CUDA 10.0. We comprehensively evaluate the performance of AEML with multiple typical benchmarks. Experimental results show that AEML can fully exploit the computing power of GPUs and achieve great load-balance among multiple heterogeneous GPUs.
Zhuo Tang, Lifan Du, Li Yang 0012, Kenli Li 0001
IEEE Trans. Computers2
2021 An Incremental Iterative Acceleration Architecture in Distributed Heterogeneous Environments With GPUs for Deep Learning
abstract
The parallel computing capabilities of GPUs have a significant impact on computationally intensive iterative tasks. Offloading part or all of the deep learning tasks from the CPU to the GPU for execution is mainstream. However, a large number of redundant iterative calculations exist in the iterative process of computing tasks. Therefore, we propose a GPU-based distributed incremental iterative computing architecture that can make full use of distributed parallel computing and GPU memory structure. The architecture supports deep learning and other computationally intensive iterative applications by optimizing data placement and reducing redundant iterative calculations. To support block-based data partitioning and coalesced memory access on GPUs, we propose GDataSet, an abstract data set. The GPU incremental iteration manager called GTracker is designed to be responsible for GDataSet cache management on the GPU. In order to solve the limitation of on-chip memory size, we propose a variable sliding window mechanism. It improves the hit rate of cache access and the speed of data access by realizing the best block arrangement between on-chip memory and off-chip memory. Besides, a communication channel based on an incremental iterative model is designed to support data transmission and task communication in cluster computing. Finally, we implement the proposed architecture based on Spark 2.4.1 and CUDA 10.0. Comparative experiments with widely used computationally intensive iterative applications (K-means, LSTM, etc.) show that the incremental iterative acceleration architecture can significantly improve the efficiency of iterative computing.
Zhuo Tang, Lifan Du, Li Yang 0012
IEEE Trans. Parallel Distributed Syst.3