Linnan Wang

dblp:169/9748 · DBLP profile ↗
← Back
16ranked-venue papers
8as first author
6since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 5 first-author · 6 since 2021Systems, architecture and hardware · 5 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
11 papers
Efficient and distributed learning · 56% Optimization for machine learning · 18% Planning, search and constraint satisfaction · 12%
Computer architecture, parallel and distributed computing, and storage systems
3 papers
Hardware accelerators and domain-specific architectures · 34% GPUs and heterogeneous computing · 26% High-performance computing · 20%
Theoretical computer science
2 papers
Mathematical optimization · 100%

Topics — the 28 heaviest of 30, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning › automated machine learning
neural architecture search
2.342024
Multi-Objective Neural Architecture Search by Learning Search Space Partitions · J. Mach. Learn. Res. 2024
Sample-Efficient Neural Architecture Search by Learning Actions for Monte Carlo Tree Search · IEEE Trans. Pattern Anal. Mach. Intell. 2022
Few-Shot Neural Architecture Search · ICML 2021
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › game tree search
monte carlo tree search
1.432022
Sample-Efficient Neural Architecture Search by Learning Actions for Monte Carlo Tree Search · IEEE Trans. Pattern Anal. Mach. Intell. 2022
Learning Search Space Partition for Black-box Optimization using Monte Carlo Tree Search · NeurIPS 2020
Neural Architecture Search Using Deep Neural Networks and Monte Carlo Tree Search · AAAI 2020
Machine learning › Optimization for machine learning
multi-objective optimization
1.322024
Multi-Objective Neural Architecture Search by Learning Search Space Partitions · J. Mach. Learn. Res. 2024
Multi-objective Optimization by Learning Space Partition · ICLR 2022
Machine learning › Optimization for machine learning
black-box optimization
0.922021
Learning Space Partitions for Path Planning · NeurIPS 2021
Learning Search Space Partition for Black-box Optimization using Monte Carlo Tree Search · NeurIPS 2020
Machine learning › Efficient and distributed learning › automated machine learning › neural architecture search
multi-objective neural architecture search
0.812024
Multi-Objective Neural Architecture Search by Learning Search Space Partitions · J. Mach. Learn. Res. 2024
Machine learning › Reinforcement learning › function approximation › representation learning for reinforcement learning
latent action space
0.612022
Sample-Efficient Neural Architecture Search by Learning Actions for Monte Carlo Tree Search · IEEE Trans. Pattern Anal. Mach. Intell. 2022
Machine learning › Efficient and distributed learning
model deployment
0.612022
Searching the Deployable Convolution Neural Networks for GPUs · CVPR 2022
Hardware accelerators and domain-specific architectures
neural architecture search
0.612022
Searching the Deployable Convolution Neural Networks for GPUs · CVPR 2022
Mathematical optimization
multi-objective optimization
0.612022
Multi-objective Optimization by Learning Space Partition · ICLR 2022
Machine learning › Efficient and distributed learning › automated machine learning › neural architecture search
few-shot NAS
0.512021
Few-Shot Neural Architecture Search · ICML 2021
Machine learning › Efficient and distributed learning › automated machine learning › neural architecture search
one-shot neural architecture search
0.512021
Few-Shot Neural Architecture Search · ICML 2021
Robotics › Motion planning and robot control
path planning
0.512021
Learning Space Partitions for Path Planning · NeurIPS 2021
Machine learning › Efficient and distributed learning
distributed training
0.412020
FFT-based Gradient Sparsification for the Distributed Training of Deep Neural Networks · HPDC 2020
Machine learning › Efficient and distributed learning › distributed training
gradient compression
0.412020
FFT-based Gradient Sparsification for the Distributed Training of Deep Neural Networks · HPDC 2020
Machine learning › Efficient and distributed learning › memory management
GPU memory management
0.312018
Superneurons: dynamic GPU memory management for training deep neural networks · PPoPP 2018
Machine learning › Efficient and distributed learning
memory-efficient training
0.312018
Superneurons: dynamic GPU memory management for training deep neural networks · PPoPP 2018
Machine learning › Efficient and distributed learning
model compression
0.312018
Learning Compact Recurrent Neural Networks With Block-Term Tensor Decomposition · CVPR 2018
Machine learning › Deep learning architectures and training
recurrent neural network
0.312018
Learning Compact Recurrent Neural Networks With Block-Term Tensor Decomposition · CVPR 2018
Machine learning › Representation and self-supervised learning
tensor decomposition
0.312018
Learning Compact Recurrent Neural Networks With Block-Term Tensor Decomposition · CVPR 2018
Parallel and multicore computing › parallel scheduling
communication scheduling
0.312018
ADAPT: an event-based adaptive collective communication framework · HPDC 2018
GPUs and heterogeneous computing
GPU memory
0.312018
Superneurons: dynamic GPU memory management for training deep neural networks · PPoPP 2018
High-performance computing › collective communication
MPI collective communication
0.312018
ADAPT: an event-based adaptive collective communication framework · HPDC 2018
Machine learning › Reinforcement learning
model-based reinforcement learning
0.112021
Learning Space Partitions for Path Planning · NeurIPS 2021
Machine learning › Efficient and distributed learning
parameter sharing
0.112021
Few-Shot Neural Architecture Search · ICML 2021
Machine learning › Efficient and distributed learning › distributed training
gradient aggregation
0.112020
FFT-based Gradient Sparsification for the Distributed Training of Deep Neural Networks · HPDC 2020
Mathematical optimization › optimization for machine learning
neural architecture search
0.112020
Learning Search Space Partition for Black-box Optimization using Monte Carlo Tree Search · NeurIPS 2020
Machine learning › Efficient and distributed learning
memory optimization
0.112018
Superneurons: dynamic GPU memory management for training deep neural networks · PPoPP 2018
GPUs and heterogeneous computing
heterogeneous supercomputing
0.112018
ADAPT: an event-based adaptive collective communication framework · HPDC 2018

Methods — techniques the papers use, named apart from their topics

bayesian optimization · 2.2evolutionary algorithm · 1.3pareto frontier search · 1.1learning space partition · 1.1distributed NAS · 1.1LaMOO · 0.8super-network · 0.5sub-supernet · 0.5monte carlo tree search · 0.5latent representation learning · 0.5LOCAL model · 0.4topology-aware communication tree · 0.3recomputation · 0.3liveness analysis · 0.3
YearPublicationVenuePosition
2024 Multi-Objective Neural Architecture Search by Learning Search Space Partitions
abstract
Deploying deep learning models requires taking into consideration neural network metrics such as model size, inference latency, and #FLOPs, aside from inference accuracy. This results in deep learning model designers leveraging multi-objective optimization to design effective deep neural networks in multiple criteria. However, applying multi-objective optimizations to neural architecture search (NAS) is nontrivial because NAS tasks usually have a huge search space, along with a non-negligible searching cost. This requires effective multi-objective search algorithms to alleviate the GPU costs. In this work, we implement a novel multi-objectives optimizer based on a recently proposed meta-algorithm called LaMOO on NAS tasks. In a nutshell, LaMOO speedups the search process by learning a model from observed samples to partition the search space and then focusing on promising regions likely to contain a subset of the Pareto frontier. Using LaMOO, we observe an improvement of more than 200% sample efficiency compared to Bayesian optimization and evolutionary-based multi-objective optimizers on different NAS datasets. For example, when combined with LaMOO, qEHVI achieves a 225% improvement in sample efficiency compared to using qEHVI alone in NasBench201. For real-world tasks, LaMOO achieves 97.36% accuracy with only 1.62M #Params on CIFAR10 in only 600 search samples. On ImageNet, our large model reaches 80.4% top-1 accuracy with only 522M #FLOPs.
Linnan Wang, Tian Guo 0001
J. Mach. Learn. Res.2
2022 Searching the Deployable Convolution Neural Networks for GPUs
abstract
Customizing Convolution Neural Networks (CNN) for production use has been a challenging task for DL practitioners. This paper intends to expedite the model customization with a model hub that contains the optimized models tiered by their inference latency using Neural Architecture Search (NAS). To achieve this goal, we build a distributed NAS system to search on a novel search space that consists of prominent factors to impact latency and accuracy. Since we target GPU, we name the NAS optimized models as GPUNet, which establishes a new SOTA Pareto frontier in inference latency and accuracy. Within 1ms, GPUNet is 2x faster than EfficientNet-X and FBNetV3 with even better accuracy. We also validate GPUNet on detection tasks, and GPUNet consistently outperforms EfficientNet-X and FB-NetV3 on COCO detection tasks in both latency and accuracy. All of these data validate that our NAS system is effective and generic to handle different design tasks. With this NAS system, we expand GPUNet to cover a wide range of latency targets such that DL practitioners can deploy our models directly in different scenarios.
Linnan Wang, Chenhan Yu, Satish Salian, Slawomir Kierat, Szymon Migacz, Alex Fit-Florea
CVPR1
2022 Multi-objective Optimization by Learning Space Partition
Linnan Wang, Kevin Yang, Tianjun Zhang, Tian Guo 0001, Yuandong Tian
ICLR2
2022 Sample-Efficient Neural Architecture Search by Learning Actions for Monte Carlo Tree Search
abstract
Neural Architecture Search (NAS) has emerged as a promising technique for automatic neural network design. However, existing MCTS based NAS approaches often utilize manually designed action space, which is not directly related to the performance metric to be optimized (e.g., accuracy), leading to sample-inefficient explorations of architectures. To improve the sample efficiency, this paper proposes Latent Action Neural Architecture Search (LaNAS), which learns actions to recursively partition the search space into good or bad regions that contain networks with similar performance metrics. During the search phase, as different action sequences lead to regions with different performance, the search efficiency can be significantly improved by biasing towards the good regions. On three NAS tasks, empirical results demonstrate that LaNAS is at least an order more sample efficient than baseline methods including evolutionary algorithms, Bayesian optimizations, and random search. When applied in practice, both one-shot and regular LaNAS consistently outperform existing results. Particularly, LaNAS achieves 99.0 percent accuracy on CIFAR-10 and 80.8 percent top1 accuracy at 600 MFLOPS on ImageNet in only 800 samples, significantly outperforming AmoebaNet with 33× fewer samples. Our code is publicly available at https://github.com/facebookresearch/LaMCTS.
Linnan Wang, Saining Xie, Teng Li 0009, Rodrigo Fonseca, Yuandong Tian
IEEE Trans. Pattern Anal. Mach. Intell.1
2021 Few-Shot Neural Architecture Search
abstract
Efficient evaluation of a network architecture drawn from a large search space remains a key challenge in Neural Architecture Search (NAS). Vanilla NAS evaluates each architecture by training from scratch, which gives the true performance but is extremely time-consuming. Recently, one-shot NAS substantially reduces the computation cost by training only one supernetwork, a.k.a. supernet, to approximate the performance of every architecture in the search space via weight-sharing. However, the performance estimation can be very inaccurate due to the co-adaption among operations. In this paper, we propose few-shot NAS that uses multiple supernetworks, called sub-supernet, each covering different regions of the search space to alleviate the undesired co-adaption. Compared to one-shot NAS, few-shot NAS improves the accuracy of architecture evaluation with a small increase of evaluation cost. With only up to 7 sub-supernets, few-shot NAS establishes new SoTAs: on ImageNet, it finds models that reach 80.5% top-1 accuracy at 600 MB FLOPS and 77.5% top-1 accuracy at 238 MFLOPS; on CIFAR10, it reaches 98.72% top-1 accuracy without using extra data or transfer learning. In Auto-GAN, few-shot NAS outperforms the previously published results by up to 20%. Extensive experiments show that few-shot NAS significantly improves various one-shot methods, including 4 gradient-based and 6 search-based methods on 3 different tasks in NasBench-201 and NasBench1-shot-1.
Linnan Wang, Yuandong Tian, Rodrigo Fonseca, Tian Guo 0001
ICML2
2021 Learning Space Partitions for Path Planning
abstract
Path planning, the problem of efficiently discovering high-reward trajectories, often requires optimizing a high-dimensional and multimodal reward function. Popular approaches like CEM and CMA-ES greedily focus on promising regions of the search space and may get trapped in local maxima. DOO and VOOT balance exploration and exploitation, but use space partitioning strategies independent of the reward function to be optimized. Recently, LaMCTS empirically learns to partition the search space in a reward-sensitive manner for black-box optimization. In this paper, we develop a novel formal regret analysis for when and why such an adaptive region partitioning scheme works. We also propose a new path planning method LaP3 which improves the function value estimation within each sub-region, and uses a latent representation of the search space. Empirically, LaP3 outperforms existing path planning methods in 2D navigation tasks, especially in the presence of difficult-to-escape local optima, and shows benefits when plugged into the planning components of model-based RL such as PETS. These gains transfer to highly multimodal real-world tasks, where we outperform strong baselines in compiler phase ordering by up to 39% on average across 9 tasks, and in molecular design by up to 0.4 on properties on a 0-1 scale. Code is available at https://github.com/yangkevin2/neurips2021-lap3.
Kevin Yang, Tianjun Zhang, Chris Cummins, Brandon Cui, Benoit Steiner, Linnan Wang, Joseph Gonzalez 0001, Daniel Klein 0001, Yuandong Tian
NeurIPS6
2020 Neural Architecture Search Using Deep Neural Networks and Monte Carlo Tree Search
abstract
Neural Architecture Search (NAS) has shown great success in automating the design of neural networks, but the prohibitive amount of computations behind current NAS methods requires further investigations in improving the sample efficiency and the network evaluation cost to get better results in a shorter time. In this paper, we present a novel scalable Monte Carlo Tree Search (MCTS) based NAS agent, named AlphaX, to tackle these two aspects. AlphaX improves the search efficiency by adaptively balancing the exploration and exploitation at the state level, and by a Meta-Deep Neural Network (DNN) to predict network accuracies for biasing the search toward a promising region. To amortize the network evaluation cost, AlphaX accelerates MCTS rollouts with a distributed design and reduces the number of epochs in evaluating a network by transfer learning, which is guided with the tree structure in MCTS. In 12 GPU days and 1000 samples, AlphaX found an architecture that reaches 97.84% top-1 accuracy on CIFAR-10, and 75.5% top-1 accuracy on ImageNet, exceeding SOTA NAS methods in both the accuracy and sampling efficiency. Particularly, we also evaluate AlphaX on NASBench-101, a large scale NAS dataset; AlphaX is 3x and 2.8x more sample efficient than Random Search and Regularized Evolution in finding the global optimum. Finally, we show the searched architecture improves a variety of vision applications from Neural Style Transfer, to Image Captioning and Object Detection.
Linnan Wang, Yuu Jinnai, Yuandong Tian, Rodrigo Fonseca
AAAI1
2020 FFT-based Gradient Sparsification for the Distributed Training of Deep Neural Networks
abstract
The performance and efficiency of distributed training of Deep Neural Networks (DNN) highly depend on the performance of gradient averaging among participating processes, a step bound by communication costs. There are two major approaches to reduce communication overhead: overlap communications with computations (lossless), or reduce communications (lossy). The lossless solution works well for linear neural architectures, e.g. VGG, AlexNet, but more recent networks such as ResNet and Inception limit the opportunity for such overlapping. Therefore, approaches that reduce the amount of data (lossy) become more suitable. In this paper, we present a novel, explainable lossy method that sparsifies gradients in the frequency domain, in addition to a new range-based float point representation to quantize and further compress gradients. These dynamic techniques strike a balance between compression ratio, accuracy, and computational overhead, and are optimized to maximize performance in heterogeneous environments.
Linnan Wang, Wei Wu 0016, Junyu Zhang 0002, Hang Liu 0001, George Bosilca, Maurice Herlihy, Rodrigo Fonseca
HPDC1
2020 Learning Search Space Partition for Black-box Optimization using Monte Carlo Tree Search
abstract
High dimensional black-box optimization has broad applications but remains a challenging problem to solve. Given a set of samples xi, yi, building a global model (like Bayesian Optimization (BO)) suffers from the curse of dimensionality in the high-dimensional search space, while a greedy search may lead to sub-optimality. By recursively splitting the search space into regions with high/low function values, recent works like LaNAS shows good performance in Neural Architecture Search (NAS), reducing the sample complexity empirically. In this paper, we coin LA-MCTS that extends LaNAS to other domains. Unlike previous approaches, LA-MCTS learns the partition of the search space using a few samples and their function values in an online fashion. While LaNAS uses linear partition and performs uniform sampling in each region, our LA-MCTS adopts a nonlinear decision boundary and learns a local model to pick good candidates. If the nonlinear partition function and the local model fits well with ground-truth black-box function, then good partitions and candidates can be reached with much fewer samples. LA-MCTS serves as a meta-algorithm by using existing black-box optimizers (e.g., BO, TuRBO as its local models, achieving strong performance in general black-box optimization and reinforcement learning benchmarks, in particular for high-dimensional problems.
Linnan Wang, Rodrigo Fonseca, Yuandong Tian
NeurIPS1
2018 Learning Compact Recurrent Neural Networks With Block-Term Tensor Decomposition
abstract
Recurrent Neural Networks (RNNs) are powerful sequence modeling tools. However, when dealing with high dimensional inputs, the training of RNNs becomes computational expensive due to the large number of model parameters. This hinders RNNs from solving many important computer vision tasks, such as Action Recognition in Videos and Image Captioning. To overcome this problem, we propose a compact and flexible structure, namely Block-Term tensor decomposition, which greatly reduces the parameters of RNNs and improves their training efficiency. Compared with alternative low-rank approximations, such as tensortrain RNN (TT-RNN), our method, Block-Term RNN (BT-RNN), is not only more concise (when using the same rank), but also able to attain a better approximation to the original RNNs with much fewer parameters. On three challenging tasks, including Action Recognition in Videos, Image Captioning and Image Generation, BT-RNN outperforms TT-RNN and the standard RNN in terms of both prediction accuracy and convergence rate. Specifically, BT-LSTM utilizes 17,388 times fewer parameters than the standard LSTM to achieve an accuracy improvement over 15.6% in the Action Recognition task on the UCF11 dataset.
Jinmian Ye, Linnan Wang, Guangxi Li, Shandian Zhe, Xinqi Chu, Zenglin Xu
CVPR2
2018 ADAPT: an event-based adaptive collective communication framework
abstract
The increase in scale and heterogeneity of high-performance computing (HPC) systems predispose the performance of Message Passing Interface (MPI) collective communications to be susceptible to noise, and to adapt to a complex mix of hardware capabilities. The designs of state of the art MPI collectives heavily rely on synchronizations; these designs magnify noise across the participating processes, resulting in significant performance slowdown. Therefore, such design philosophy must be reconsidered to efficiently and robustly run on the large-scale heterogeneous platforms. In this paper, we present ADAPT, a new collective communication framework in Open MPI, using event-driven techniques to morph collective algorithms to heterogeneous environments. The core concept of ADAPT is to relax synchronizations, while mamtaining the minimal data dependencies of MPI collectives. To fully exploit the different bandwidths of data movement lanes in heterogeneous systems, we extend the ADAPT collective framework with a topology-aware communication tree. This removes the boundaries of different hardware topologies while maximizing the speed of data movements. We evaluate our framework with two popular collective operations: broadcast and reduce on both CPU and GPU clusters. Our results demonstrate drastic performance improvements and a strong resistance against noise compared to other state of the art MPI libraries. In particular, we demonstrate at least 1.3X and 1.5X speedup for CPU data and 2X and 10X speedup for GPU data using ADAPT event-based broadcast and reduce operations.
Wei Wu 0016, George Bosilca, Thananon Patinyasakdikul, Linnan Wang, Jack J. Dongarra
HPDC5
2018 Warp-Consolidation: A Novel Execution Model for GPUs
abstract
With the unprecedented development of compute capability and extension of memory bandwidth on modern GPUs, parallel communication and synchronization soon becomes a major concern for continuous performance scaling. This is especially the case for emerging big-data applications. Instead of relying on a few heavily-loaded CTAs that may expose opportunities for intra-CTA data reuse, current technology and design trends suggest the performance potential of allocating more lightweighted CTAs for processing individual tasks more independently, as the overheads from synchronization, communication and cooperation may greatly outweigh the benefits from exploiting limited data reuse in heavily-loaded CTAs. This paper proceeds this trend and proposes a novel execution model for modern GPUs that hides the CTA execution hierarchy from the classic GPU execution model; meanwhile exposes the originally hidden warp-level execution. Specifically, it relies on individual warps to undertake the original CTAs' tasks. The major observation is that by replacing traditional inter-warp communication (e.g., via shared memory), cooperation (e.g., via bar primitives) and synchronizations (e.g., via CTA barriers), with more efficient intra-warp communication (e.g., via register shuffling), cooperation (e.g., via warp voting) and synchronizations (naturally lockstep execution) across the SIMD-lanes within a warp, significant performance gain can be achieved. We analyze the pros and cons for this design and propose corresponding solutions to counter potential negative effects. Experimental results on a diverse group of thirty-two representative applications show that our proposed Warp-Consolidation execution model can achieve an average speedup of 1.7x, 2.3x, 1.5x and 1.2x (up to 6.3x, 31x, 6.4x and 3.8x) on NVIDIA Kepler (Tesla-K80), Maxwell (Tesla-M40), Pascal (Tesla-P100) and Volta (Tesla-V100) GPUs, respectively, demonstrating its applicability and portability. Our approach can be directly employed to either transform legacy codes or write new algorithms on modern commodity GPUs.
Ang Li 0006, Weifeng Liu 0002, Linnan Wang, Kevin J. Barker, Shuaiwen Song
ICS3
2018 Superneurons: dynamic GPU memory management for training deep neural networks
abstract
Going deeper and wider in neural architectures improves their accuracy, while the limited GPU DRAM places an undesired restriction on the network design domain. Deep Learning (DL) practitioners either need to change to less desired network architectures, or nontrivially dissect a network across multiGPUs. These distract DL practitioners from concentrating on their original machine learning tasks. We present SuperNeurons: a dynamic GPU memory scheduling runtime to enable the network training far beyond the GPU DRAM capacity. SuperNeurons features 3 memory optimizations, Liveness Analysis, Unified Tensor Pool, and Cost-Aware Recomputation; together they effectively reduce the network-wide peak memory usage down to the maximal memory usage among layers. We also address the performance issues in these memory-saving techniques. Given the limited GPU DRAM, SuperNeurons not only provisions the necessary memory for the training, but also dynamically allocates the memory for convolution workspaces to achieve the high performance. Evaluations against Caffe, Torch, MXNet and TensorFlow have demonstrated that SuperNeurons trains at least 3.2432 deeper network than current ones with the leading performance. Particularly, SuperNeurons can train ResNet2500 that has 104 basic network layers on a 12GB K40c.
Linnan Wang, Jinmian Ye, Wei Wu 0016, Ang Li 0006, Shuaiwen Song, Zenglin Xu, Tim Kraska
PPoPP1
2017 Simple and efficient parallelization for probabilistic temporal tensor factorization
abstract
Probabilistic Temporal Tensor Factorization (PTTF) is an effective algorithm to model the temporal tensor data. It leverages a time constraint to capture the evolving properties of tensor data. Nowadays the exploding dataset demands a large scale PTTF analysis, and a parallel solution is critical to accommodate the trend. Whereas, the parallelization of PTTF still remains unexplored. In this paper, we propose a simple yet efficient Parallel Probabilistic Temporal Tensor Factorization, referred to as P2T2F, to provide a scalable PTTF solution. P2T2F is fundamentally disparate from existing parallel tensor factorizations by considering the probabilistic decomposition and the temporal effects of tensor data. It adopts a new tensor data split strategy to subdivide a large tensor into independent sub-tensors, the computation of which is inherently parallel. We train P2T2F with an efficient algorithm of stochastic Alternating Direction Method of Multipliers, and show that the convergence is guaranteed. Experiments on several real-word tensor datasets demonstrate that P2T2F is a highly effective and efficiently scalable algorithm dedicated for large scale probabilistic temporal tensor analysis.
Guangxi Li, Zenglin Xu, Linnan Wang, Jinmian Ye, Irwin King, Michael R. Lyu
IJCNN3
2017 Accelerating deep neural network training with inconsistent stochastic gradient descent
Linnan Wang, Yi Yang 0018, Martin Renqiang Min, Srimat T. Chakradhar
Neural Networks1
2016 BLASX: A High Performance Level-3 BLAS Library for Heterogeneous Multi-GPU Computing
abstract
Basic Linear Algebra Subprograms (BLAS) are a set of low level linear algebra kernels widely adopted by applications involved with the deep learning and scientific computing. The massive and economic computing power brought forth by the emerging GPU architectures drives interest in implementation of compute-intensive level 3 BLAS on multi-GPU systems. In this paper, we investigate existing multi-GPU level 3 BLAS and present that 1) issues, such as the improper load balancing, inefficient communication, insufficient GPU stream level concurrency and data caching, impede current implementations from fully harnessing heterogeneous computing resources; 2) and the inter-GPU Peer-to-Peer(P2P) communication remains unexplored. We then present BLASX: a highly optimized multi-GPU level-3 BLAS. We adopt the concepts of algorithms-by-tiles treating a matrix tile as the basic data unit and operations on tiles as the basic task. Tasks are guided with a dynamic asynchronous runtime, which is cache and locality aware. The communication cost under BLASX becomes trivial as it perfectly overlaps communication and computation across multiple streams during asynchronous task progression. It also takes the current tile cache scheme one step further by proposing an innovative 2-level hierarchical tile cache, taking advantage of inter-GPU P2P communication. As a result, linear speedup is observable with BLASX under multi-GPU configurations; and the extensive benchmarks demonstrate that BLASX consistently outperforms the related leading industrial and academic implementations such as cuBLAS-XT, SuperMatrix, MAGMA.
Linnan Wang, Wei Wu 0016, Zenglin Xu, Jianxiong Xiao, Yi Yang 0018
ICS1