VLDB 2026 Research / reviewers in the wild / expert
Linnan Wang
dblp:169/9748
· DBLP profile ↗
16ranked-venue papers
8as first author
6since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 5 first-author · 6 since 2021Systems, architecture and hardware · 5 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
11 papers |
Efficient and distributed learning · 56% Optimization for machine learning · 18% Planning, search and constraint satisfaction · 12% | |
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Hardware accelerators and domain-specific architectures · 34% GPUs and heterogeneous computing · 26% High-performance computing · 20% | |
| Theoretical computer science
2 papers |
Mathematical optimization · 100% |
Topics — the 28 heaviest of 30, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning › automated machine learning
neural architecture search |
2.3 | 4 | 2024 | Multi-Objective Neural Architecture Search by Learning Search Space Partitions · J. Mach. Learn. Res. 2024 Sample-Efficient Neural Architecture Search by Learning Actions for Monte Carlo Tree Search · IEEE Trans. Pattern Anal. Mach. Intell. 2022 Few-Shot Neural Architecture Search · ICML 2021 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › game tree search
monte carlo tree search |
1.4 | 3 | 2022 | Sample-Efficient Neural Architecture Search by Learning Actions for Monte Carlo Tree Search · IEEE Trans. Pattern Anal. Mach. Intell. 2022 Learning Search Space Partition for Black-box Optimization using Monte Carlo Tree Search · NeurIPS 2020 Neural Architecture Search Using Deep Neural Networks and Monte Carlo Tree Search · AAAI 2020 |
Machine learning › Optimization for machine learning
multi-objective optimization |
1.3 | 2 | 2024 | Multi-Objective Neural Architecture Search by Learning Search Space Partitions · J. Mach. Learn. Res. 2024 Multi-objective Optimization by Learning Space Partition · ICLR 2022 |
Machine learning › Optimization for machine learning
black-box optimization |
0.9 | 2 | 2021 | Learning Space Partitions for Path Planning · NeurIPS 2021 Learning Search Space Partition for Black-box Optimization using Monte Carlo Tree Search · NeurIPS 2020 |
Machine learning › Efficient and distributed learning › automated machine learning › neural architecture search
multi-objective neural architecture search |
0.8 | 1 | 2024 | Multi-Objective Neural Architecture Search by Learning Search Space Partitions · J. Mach. Learn. Res. 2024 |
Machine learning › Reinforcement learning › function approximation › representation learning for reinforcement learning
latent action space |
0.6 | 1 | 2022 | Sample-Efficient Neural Architecture Search by Learning Actions for Monte Carlo Tree Search · IEEE Trans. Pattern Anal. Mach. Intell. 2022 |
Machine learning › Efficient and distributed learning
model deployment |
0.6 | 1 | 2022 | Searching the Deployable Convolution Neural Networks for GPUs · CVPR 2022 |
Hardware accelerators and domain-specific architectures
neural architecture search |
0.6 | 1 | 2022 | Searching the Deployable Convolution Neural Networks for GPUs · CVPR 2022 |
Mathematical optimization
multi-objective optimization |
0.6 | 1 | 2022 | Multi-objective Optimization by Learning Space Partition · ICLR 2022 |
Machine learning › Efficient and distributed learning › automated machine learning › neural architecture search
few-shot NAS |
0.5 | 1 | 2021 | Few-Shot Neural Architecture Search · ICML 2021 |
Machine learning › Efficient and distributed learning › automated machine learning › neural architecture search
one-shot neural architecture search |
0.5 | 1 | 2021 | Few-Shot Neural Architecture Search · ICML 2021 |
Robotics › Motion planning and robot control
path planning |
0.5 | 1 | 2021 | Learning Space Partitions for Path Planning · NeurIPS 2021 |
Machine learning › Efficient and distributed learning
distributed training |
0.4 | 1 | 2020 | FFT-based Gradient Sparsification for the Distributed Training of Deep Neural Networks · HPDC 2020 |
Machine learning › Efficient and distributed learning › distributed training
gradient compression |
0.4 | 1 | 2020 | FFT-based Gradient Sparsification for the Distributed Training of Deep Neural Networks · HPDC 2020 |
Machine learning › Efficient and distributed learning › memory management
GPU memory management |
0.3 | 1 | 2018 | Superneurons: dynamic GPU memory management for training deep neural networks · PPoPP 2018 |
Machine learning › Efficient and distributed learning
memory-efficient training |
0.3 | 1 | 2018 | Superneurons: dynamic GPU memory management for training deep neural networks · PPoPP 2018 |
Machine learning › Efficient and distributed learning
model compression |
0.3 | 1 | 2018 | Learning Compact Recurrent Neural Networks With Block-Term Tensor Decomposition · CVPR 2018 |
Machine learning › Deep learning architectures and training
recurrent neural network |
0.3 | 1 | 2018 | Learning Compact Recurrent Neural Networks With Block-Term Tensor Decomposition · CVPR 2018 |
Machine learning › Representation and self-supervised learning
tensor decomposition |
0.3 | 1 | 2018 | Learning Compact Recurrent Neural Networks With Block-Term Tensor Decomposition · CVPR 2018 |
Parallel and multicore computing › parallel scheduling
communication scheduling |
0.3 | 1 | 2018 | ADAPT: an event-based adaptive collective communication framework · HPDC 2018 |
GPUs and heterogeneous computing
GPU memory |
0.3 | 1 | 2018 | Superneurons: dynamic GPU memory management for training deep neural networks · PPoPP 2018 |
High-performance computing › collective communication
MPI collective communication |
0.3 | 1 | 2018 | ADAPT: an event-based adaptive collective communication framework · HPDC 2018 |
Machine learning › Reinforcement learning
model-based reinforcement learning |
0.1 | 1 | 2021 | Learning Space Partitions for Path Planning · NeurIPS 2021 |
Machine learning › Efficient and distributed learning
parameter sharing |
0.1 | 1 | 2021 | Few-Shot Neural Architecture Search · ICML 2021 |
Machine learning › Efficient and distributed learning › distributed training
gradient aggregation |
0.1 | 1 | 2020 | FFT-based Gradient Sparsification for the Distributed Training of Deep Neural Networks · HPDC 2020 |
Mathematical optimization › optimization for machine learning
neural architecture search |
0.1 | 1 | 2020 | Learning Search Space Partition for Black-box Optimization using Monte Carlo Tree Search · NeurIPS 2020 |
Machine learning › Efficient and distributed learning
memory optimization |
0.1 | 1 | 2018 | Superneurons: dynamic GPU memory management for training deep neural networks · PPoPP 2018 |
GPUs and heterogeneous computing
heterogeneous supercomputing |
0.1 | 1 | 2018 | ADAPT: an event-based adaptive collective communication framework · HPDC 2018 |
Methods — techniques the papers use, named apart from their topics
bayesian optimization · 2.2evolutionary algorithm · 1.3pareto frontier search · 1.1learning space partition · 1.1distributed NAS · 1.1LaMOO · 0.8super-network · 0.5sub-supernet · 0.5monte carlo tree search · 0.5latent representation learning · 0.5LOCAL model · 0.4topology-aware communication tree · 0.3recomputation · 0.3liveness analysis · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Multi-Objective Neural Architecture Search by Learning Search Space PartitionsabstractDeploying deep learning models requires taking into consideration neural network metrics such as model size, inference latency, and #FLOPs, aside from inference accuracy. This results in deep learning model designers leveraging multi-objective optimization to design effective deep neural networks in multiple criteria. However, applying multi-objective optimizations to neural architecture search (NAS) is nontrivial because NAS tasks usually have a huge search space, along with a non-negligible searching cost. This requires effective multi-objective search algorithms to alleviate the GPU costs. In this work, we implement a novel multi-objectives optimizer based on a recently proposed meta-algorithm called LaMOO on NAS tasks. In a nutshell, LaMOO speedups the search process by learning a model from observed samples to partition the search space and then focusing on promising regions likely to contain a subset of the Pareto frontier. Using LaMOO, we observe an improvement of more than 200% sample efficiency compared to Bayesian optimization and evolutionary-based multi-objective optimizers on different NAS datasets. For example, when combined with LaMOO, qEHVI achieves a 225% improvement in sample efficiency compared to using qEHVI alone in NasBench201. For real-world tasks, LaMOO achieves 97.36% accuracy with only 1.62M #Params on CIFAR10 in only 600 search samples. On ImageNet, our large model reaches 80.4% top-1 accuracy with only 522M #FLOPs. Linnan Wang, Tian Guo 0001 |
J. Mach. Learn. Res. | 2 |
| 2022 | Searching the Deployable Convolution Neural Networks for GPUsabstractCustomizing Convolution Neural Networks (CNN) for production use has been a challenging task for DL practitioners. This paper intends to expedite the model customization with a model hub that contains the optimized models tiered by their inference latency using Neural Architecture Search (NAS). To achieve this goal, we build a distributed NAS system to search on a novel search space that consists of prominent factors to impact latency and accuracy. Since we target GPU, we name the NAS optimized models as GPUNet, which establishes a new SOTA Pareto frontier in inference latency and accuracy. Within 1ms, GPUNet is 2x faster than EfficientNet-X and FBNetV3 with even better accuracy. We also validate GPUNet on detection tasks, and GPUNet consistently outperforms EfficientNet-X and FB-NetV3 on COCO detection tasks in both latency and accuracy. All of these data validate that our NAS system is effective and generic to handle different design tasks. With this NAS system, we expand GPUNet to cover a wide range of latency targets such that DL practitioners can deploy our models directly in different scenarios. Linnan Wang, Chenhan Yu, Satish Salian, Slawomir Kierat, Szymon Migacz, Alex Fit-Florea |
CVPR | 1 |
| 2022 | Multi-objective Optimization by Learning Space Partition
Linnan Wang, Kevin Yang, Tianjun Zhang, Tian Guo 0001, Yuandong Tian |
ICLR | 2 |
| 2022 | Sample-Efficient Neural Architecture Search by Learning Actions for Monte Carlo Tree SearchabstractNeural Architecture Search (NAS) has emerged as a promising technique for automatic neural network design. However, existing MCTS based NAS approaches often utilize manually designed action space, which is not directly related to the performance metric to be optimized (e.g., accuracy), leading to sample-inefficient explorations of architectures. To improve the sample efficiency, this paper proposes Latent Action Neural Architecture Search (LaNAS), which learns actions to recursively partition the search space into good or bad regions that contain networks with similar performance metrics. During the search phase, as different action sequences lead to regions with different performance, the search efficiency can be significantly improved by biasing towards the good regions. On three NAS tasks, empirical results demonstrate that LaNAS is at least an order more sample efficient than baseline methods including evolutionary algorithms, Bayesian optimizations, and random search. When applied in practice, both one-shot and regular LaNAS consistently outperform existing results. Particularly, LaNAS achieves 99.0 percent accuracy on CIFAR-10 and 80.8 percent top1 accuracy at 600 MFLOPS on ImageNet in only 800 samples, significantly outperforming AmoebaNet with 33× fewer samples. Our code is publicly available at https://github.com/facebookresearch/LaMCTS. Linnan Wang, Saining Xie, Teng Li 0009, Rodrigo Fonseca, Yuandong Tian |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | Few-Shot Neural Architecture SearchabstractEfficient evaluation of a network architecture drawn from a large search space remains a key challenge in Neural Architecture Search (NAS). Vanilla NAS evaluates each architecture by training from scratch, which gives the true performance but is extremely time-consuming. Recently, one-shot NAS substantially reduces the computation cost by training only one supernetwork, a.k.a. supernet, to approximate the performance of every architecture in the search space via weight-sharing. However, the performance estimation can be very inaccurate due to the co-adaption among operations. In this paper, we propose few-shot NAS that uses multiple supernetworks, called sub-supernet, each covering different regions of the search space to alleviate the undesired co-adaption. Compared to one-shot NAS, few-shot NAS improves the accuracy of architecture evaluation with a small increase of evaluation cost. With only up to 7 sub-supernets, few-shot NAS establishes new SoTAs: on ImageNet, it finds models that reach 80.5% top-1 accuracy at 600 MB FLOPS and 77.5% top-1 accuracy at 238 MFLOPS; on CIFAR10, it reaches 98.72% top-1 accuracy without using extra data or transfer learning. In Auto-GAN, few-shot NAS outperforms the previously published results by up to 20%. Extensive experiments show that few-shot NAS significantly improves various one-shot methods, including 4 gradient-based and 6 search-based methods on 3 different tasks in NasBench-201 and NasBench1-shot-1. Linnan Wang, Yuandong Tian, Rodrigo Fonseca, Tian Guo 0001 |
ICML | 2 |
| 2021 | Learning Space Partitions for Path PlanningabstractPath planning, the problem of efficiently discovering high-reward trajectories, often requires optimizing a high-dimensional and multimodal reward function. Popular approaches like CEM and CMA-ES greedily focus on promising regions of the search space and may get trapped in local maxima. DOO and VOOT balance exploration and exploitation, but use space partitioning strategies independent of the reward function to be optimized. Recently, LaMCTS empirically learns to partition the search space in a reward-sensitive manner for black-box optimization. In this paper, we develop a novel formal regret analysis for when and why such an adaptive region partitioning scheme works. We also propose a new path planning method LaP3 which improves the function value estimation within each sub-region, and uses a latent representation of the search space. Empirically, LaP3 outperforms existing path planning methods in 2D navigation tasks, especially in the presence of difficult-to-escape local optima, and shows benefits when plugged into the planning components of model-based RL such as PETS. These gains transfer to highly multimodal real-world tasks, where we outperform strong baselines in compiler phase ordering by up to 39% on average across 9 tasks, and in molecular design by up to 0.4 on properties on a 0-1 scale. Code is available at https://github.com/yangkevin2/neurips2021-lap3. Kevin Yang, Tianjun Zhang, Chris Cummins, Brandon Cui, Benoit Steiner, Linnan Wang, Joseph Gonzalez 0001, Daniel Klein 0001, Yuandong Tian |
NeurIPS | 6 |
| 2020 | Neural Architecture Search Using Deep Neural Networks and Monte Carlo Tree SearchabstractNeural Architecture Search (NAS) has shown great success in automating the design of neural networks, but the prohibitive amount of computations behind current NAS methods requires further investigations in improving the sample efficiency and the network evaluation cost to get better results in a shorter time. In this paper, we present a novel scalable Monte Carlo Tree Search (MCTS) based NAS agent, named AlphaX, to tackle these two aspects. AlphaX improves the search efficiency by adaptively balancing the exploration and exploitation at the state level, and by a Meta-Deep Neural Network (DNN) to predict network accuracies for biasing the search toward a promising region. To amortize the network evaluation cost, AlphaX accelerates MCTS rollouts with a distributed design and reduces the number of epochs in evaluating a network by transfer learning, which is guided with the tree structure in MCTS. In 12 GPU days and 1000 samples, AlphaX found an architecture that reaches 97.84% top-1 accuracy on CIFAR-10, and 75.5% top-1 accuracy on ImageNet, exceeding SOTA NAS methods in both the accuracy and sampling efficiency. Particularly, we also evaluate AlphaX on NASBench-101, a large scale NAS dataset; AlphaX is 3x and 2.8x more sample efficient than Random Search and Regularized Evolution in finding the global optimum. Finally, we show the searched architecture improves a variety of vision applications from Neural Style Transfer, to Image Captioning and Object Detection. Linnan Wang, Yuu Jinnai, Yuandong Tian, Rodrigo Fonseca |
AAAI | 1 |
| 2020 | FFT-based Gradient Sparsification for the Distributed Training of Deep Neural NetworksabstractThe performance and efficiency of distributed training of Deep Neural Networks (DNN) highly depend on the performance of gradient averaging among participating processes, a step bound by communication costs. There are two major approaches to reduce communication overhead: overlap communications with computations (lossless), or reduce communications (lossy). The lossless solution works well for linear neural architectures, e.g. VGG, AlexNet, but more recent networks such as ResNet and Inception limit the opportunity for such overlapping. Therefore, approaches that reduce the amount of data (lossy) become more suitable. In this paper, we present a novel, explainable lossy method that sparsifies gradients in the frequency domain, in addition to a new range-based float point representation to quantize and further compress gradients. These dynamic techniques strike a balance between compression ratio, accuracy, and computational overhead, and are optimized to maximize performance in heterogeneous environments. Linnan Wang, Wei Wu 0016, Junyu Zhang 0002, Hang Liu 0001, George Bosilca, Maurice Herlihy, Rodrigo Fonseca |
HPDC | 1 |
| 2020 | Learning Search Space Partition for Black-box Optimization using Monte Carlo Tree SearchabstractHigh dimensional black-box optimization has broad applications but remains a challenging problem to solve. Given a set of samples xi, yi, building a global model (like Bayesian Optimization (BO)) suffers from the curse of dimensionality in the high-dimensional search space, while a greedy search may lead to sub-optimality. By recursively splitting the search space into regions with high/low function values, recent works like LaNAS shows good performance in Neural Architecture Search (NAS), reducing the sample complexity empirically. In this paper, we coin LA-MCTS that extends LaNAS to other domains. Unlike previous approaches, LA-MCTS learns the partition of the search space using a few samples and their function values in an online fashion. While LaNAS uses linear partition and performs uniform sampling in each region, our LA-MCTS adopts a nonlinear decision boundary and learns a local model to pick good candidates. If the nonlinear partition function and the local model fits well with ground-truth black-box function, then good partitions and candidates can be reached with much fewer samples. LA-MCTS serves as a meta-algorithm by using existing black-box optimizers (e.g., BO, TuRBO as its local models, achieving strong performance in general black-box optimization and reinforcement learning benchmarks, in particular for high-dimensional problems. Linnan Wang, Rodrigo Fonseca, Yuandong Tian |
NeurIPS | 1 |
| 2018 | Learning Compact Recurrent Neural Networks With Block-Term Tensor DecompositionabstractRecurrent Neural Networks (RNNs) are powerful sequence modeling tools. However, when dealing with high dimensional inputs, the training of RNNs becomes computational expensive due to the large number of model parameters. This hinders RNNs from solving many important computer vision tasks, such as Action Recognition in Videos and Image Captioning. To overcome this problem, we propose a compact and flexible structure, namely Block-Term tensor decomposition, which greatly reduces the parameters of RNNs and improves their training efficiency. Compared with alternative low-rank approximations, such as tensortrain RNN (TT-RNN), our method, Block-Term RNN (BT-RNN), is not only more concise (when using the same rank), but also able to attain a better approximation to the original RNNs with much fewer parameters. On three challenging tasks, including Action Recognition in Videos, Image Captioning and Image Generation, BT-RNN outperforms TT-RNN and the standard RNN in terms of both prediction accuracy and convergence rate. Specifically, BT-LSTM utilizes 17,388 times fewer parameters than the standard LSTM to achieve an accuracy improvement over 15.6% in the Action Recognition task on the UCF11 dataset. Jinmian Ye, Linnan Wang, Guangxi Li, Shandian Zhe, Xinqi Chu, Zenglin Xu |
CVPR | 2 |
| 2018 | ADAPT: an event-based adaptive collective communication frameworkabstractThe increase in scale and heterogeneity of high-performance computing (HPC) systems predispose the performance of Message Passing Interface (MPI) collective communications to be susceptible to noise, and to adapt to a complex mix of hardware capabilities. The designs of state of the art MPI collectives heavily rely on synchronizations; these designs magnify noise across the participating processes, resulting in significant performance slowdown. Therefore, such design philosophy must be reconsidered to efficiently and robustly run on the large-scale heterogeneous platforms. In this paper, we present ADAPT, a new collective communication framework in Open MPI, using event-driven techniques to morph collective algorithms to heterogeneous environments. The core concept of ADAPT is to relax synchronizations, while mamtaining the minimal data dependencies of MPI collectives. To fully exploit the different bandwidths of data movement lanes in heterogeneous systems, we extend the ADAPT collective framework with a topology-aware communication tree. This removes the boundaries of different hardware topologies while maximizing the speed of data movements. We evaluate our framework with two popular collective operations: broadcast and reduce on both CPU and GPU clusters. Our results demonstrate drastic performance improvements and a strong resistance against noise compared to other state of the art MPI libraries. In particular, we demonstrate at least 1.3X and 1.5X speedup for CPU data and 2X and 10X speedup for GPU data using ADAPT event-based broadcast and reduce operations. Wei Wu 0016, George Bosilca, Thananon Patinyasakdikul, Linnan Wang, Jack J. Dongarra |
HPDC | 5 |
| 2018 | Warp-Consolidation: A Novel Execution Model for GPUsabstractWith the unprecedented development of compute capability and extension of memory bandwidth on modern GPUs, parallel communication and synchronization soon becomes a major concern for continuous performance scaling. This is especially the case for emerging big-data applications. Instead of relying on a few heavily-loaded CTAs that may expose opportunities for intra-CTA data reuse, current technology and design trends suggest the performance potential of allocating more lightweighted CTAs for processing individual tasks more independently, as the overheads from synchronization, communication and cooperation may greatly outweigh the benefits from exploiting limited data reuse in heavily-loaded CTAs. This paper proceeds this trend and proposes a novel execution model for modern GPUs that hides the CTA execution hierarchy from the classic GPU execution model; meanwhile exposes the originally hidden warp-level execution. Specifically, it relies on individual warps to undertake the original CTAs' tasks. The major observation is that by replacing traditional inter-warp communication (e.g., via shared memory), cooperation (e.g., via bar primitives) and synchronizations (e.g., via CTA barriers), with more efficient intra-warp communication (e.g., via register shuffling), cooperation (e.g., via warp voting) and synchronizations (naturally lockstep execution) across the SIMD-lanes within a warp, significant performance gain can be achieved. We analyze the pros and cons for this design and propose corresponding solutions to counter potential negative effects. Experimental results on a diverse group of thirty-two representative applications show that our proposed Warp-Consolidation execution model can achieve an average speedup of 1.7x, 2.3x, 1.5x and 1.2x (up to 6.3x, 31x, 6.4x and 3.8x) on NVIDIA Kepler (Tesla-K80), Maxwell (Tesla-M40), Pascal (Tesla-P100) and Volta (Tesla-V100) GPUs, respectively, demonstrating its applicability and portability. Our approach can be directly employed to either transform legacy codes or write new algorithms on modern commodity GPUs. Ang Li 0006, Weifeng Liu 0002, Linnan Wang, Kevin J. Barker, Shuaiwen Song |
ICS | 3 |
| 2018 | Superneurons: dynamic GPU memory management for training deep neural networksabstractGoing deeper and wider in neural architectures improves their accuracy, while the limited GPU DRAM places an undesired restriction on the network design domain. Deep Learning (DL) practitioners either need to change to less desired network architectures, or nontrivially dissect a network across multiGPUs. These distract DL practitioners from concentrating on their original machine learning tasks. We present SuperNeurons: a dynamic GPU memory scheduling runtime to enable the network training far beyond the GPU DRAM capacity. SuperNeurons features 3 memory optimizations, Liveness Analysis, Unified Tensor Pool, and Cost-Aware Recomputation; together they effectively reduce the network-wide peak memory usage down to the maximal memory usage among layers. We also address the performance issues in these memory-saving techniques. Given the limited GPU DRAM, SuperNeurons not only provisions the necessary memory for the training, but also dynamically allocates the memory for convolution workspaces to achieve the high performance. Evaluations against Caffe, Torch, MXNet and TensorFlow have demonstrated that SuperNeurons trains at least 3.2432 deeper network than current ones with the leading performance. Particularly, SuperNeurons can train ResNet2500 that has 104 basic network layers on a 12GB K40c. Linnan Wang, Jinmian Ye, Wei Wu 0016, Ang Li 0006, Shuaiwen Song, Zenglin Xu, Tim Kraska |
PPoPP | 1 |
| 2017 | Simple and efficient parallelization for probabilistic temporal tensor factorizationabstractProbabilistic Temporal Tensor Factorization (PTTF) is an effective algorithm to model the temporal tensor data. It leverages a time constraint to capture the evolving properties of tensor data. Nowadays the exploding dataset demands a large scale PTTF analysis, and a parallel solution is critical to accommodate the trend. Whereas, the parallelization of PTTF still remains unexplored. In this paper, we propose a simple yet efficient Parallel Probabilistic Temporal Tensor Factorization, referred to as P2T2F, to provide a scalable PTTF solution. P2T2F is fundamentally disparate from existing parallel tensor factorizations by considering the probabilistic decomposition and the temporal effects of tensor data. It adopts a new tensor data split strategy to subdivide a large tensor into independent sub-tensors, the computation of which is inherently parallel. We train P2T2F with an efficient algorithm of stochastic Alternating Direction Method of Multipliers, and show that the convergence is guaranteed. Experiments on several real-word tensor datasets demonstrate that P2T2F is a highly effective and efficiently scalable algorithm dedicated for large scale probabilistic temporal tensor analysis. Guangxi Li, Zenglin Xu, Linnan Wang, Jinmian Ye, Irwin King, Michael R. Lyu |
IJCNN | 3 |
| 2017 | Accelerating deep neural network training with inconsistent stochastic gradient descent
Linnan Wang, Yi Yang 0018, Martin Renqiang Min, Srimat T. Chakradhar |
Neural Networks | 1 |
| 2016 | BLASX: A High Performance Level-3 BLAS Library for Heterogeneous Multi-GPU ComputingabstractBasic Linear Algebra Subprograms (BLAS) are a set of low level linear algebra kernels widely adopted by applications involved with the deep learning and scientific computing. The massive and economic computing power brought forth by the emerging GPU architectures drives interest in implementation of compute-intensive level 3 BLAS on multi-GPU systems. In this paper, we investigate existing multi-GPU level 3 BLAS and present that 1) issues, such as the improper load balancing, inefficient communication, insufficient GPU stream level concurrency and data caching, impede current implementations from fully harnessing heterogeneous computing resources; 2) and the inter-GPU Peer-to-Peer(P2P) communication remains unexplored. We then present BLASX: a highly optimized multi-GPU level-3 BLAS. We adopt the concepts of algorithms-by-tiles treating a matrix tile as the basic data unit and operations on tiles as the basic task. Tasks are guided with a dynamic asynchronous runtime, which is cache and locality aware. The communication cost under BLASX becomes trivial as it perfectly overlaps communication and computation across multiple streams during asynchronous task progression. It also takes the current tile cache scheme one step further by proposing an innovative 2-level hierarchical tile cache, taking advantage of inter-GPU P2P communication. As a result, linear speedup is observable with BLASX under multi-GPU configurations; and the extensive benchmarks demonstrate that BLASX consistently outperforms the related leading industrial and academic implementations such as cuBLAS-XT, SuperMatrix, MAGMA. Linnan Wang, Wei Wu 0016, Zenglin Xu, Jianxiong Xiao, Yi Yang 0018 |
ICS | 1 |