VLDB 2026 Research / reviewers in the wild / expert
Wei Zhang 0022
dblp:10/4661-22
· DBLP profile ↗
30ranked-venue papers
9as first author
3since 2021 · last 2022
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 5 first-author · 1 since 2021Software engineering, systems software and programming languages · 7 · 4 first-authorSystems, architecture and hardware · 4 · 3 first-authorDatabases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
14 papers |
Efficient and distributed learning · 46% Optimization for machine learning · 38% Generative modeling · 9% | |
| Software engineering, system software, and programming languages
8 papers |
Concurrent programming · 55% Debugging and program repair · 32% Software maintenance and evolution · 7% | |
| Computer architecture, parallel and distributed computing, and storage systems
6 papers |
Hardware accelerators and domain-specific architectures · 38% Energy-efficient computing · 21% Distributed systems · 15% |
Topics — the 30 heaviest of 47, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
distributed training |
3.2 | 9 | 2021 | Asynchronous Decentralized Distributed Training of Acoustic Models · IEEE ACM Trans. Audio Speech Lang. Process. 2021 A Decentralized Parallel Algorithm for Training Generative Adversarial Nets · NeurIPS 2020 ScaleCom: Scalable Sparsified Gradient Compression for Communication-Efficient Distributed Training · NeurIPS 2020 |
Concurrent programming
concurrency bugs |
1.2 | 8 | 2015 | Fixing, preventing, and recovering from concurrency bugs · Sci. China Inf. Sci. 2015 ConMem: Detecting Crash-Triggering Concurrency Bugs through an Effect-Oriented Approach · ACM Trans. Softw. Eng. Methodol. 2013 Efficient concurrency-bug detection across inputs · OOPSLA 2013 |
Machine learning › Optimization for machine learning › stochastic optimization
adaptive gradient methods |
1.0 | 2 | 2022 | Robustness to Unbounded Smoothness of Generalized SignSGD · NeurIPS 2022 Towards Better Understanding of Adaptive Gradient Algorithms in Generative Adversarial Nets · ICLR 2020 |
Machine learning › Optimization for machine learning › stochastic gradient descent
asynchronous SGD |
0.9 | 3 | 2020 | Map Generation from Large Scale Incomplete and Inaccurate Data Labels · KDD 2020 Staleness-Aware Async-SGD for Distributed Deep Learning · IJCAI 2016 Model Accuracy and Runtime Tradeoff in Distributed Deep Learning: A Systematic Study · ICDM 2016 |
Machine learning › Optimization for machine learning
stochastic gradient descent |
0.9 | 3 | 2018 | Evolutionary Stochastic Gradient Descent for Optimization of Deep Neural Networks · NeurIPS 2018 Can Decentralized Algorithms Outperform Centralized Algorithms? A Case Study for Decentralized Parallel Stochastic Gradient Descent · NIPS 2017 Model Accuracy and Runtime Tradeoff in Distributed Deep Learning: A Systematic Study · ICDM 2016 |
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.7 | 2 | 2019 | Hybrid 8-bit Floating Point (HFP8) Training and Inference for Deep Neural Networks · NeurIPS 2019 Accelerator Design for Deep Learning Training: Extended Abstract: Invited · DAC 2017 |
Machine learning › Efficient and distributed learning › distributed training
distributed stochastic gradient descent |
0.6 | 2 | 2018 | Asynchronous Decentralized Parallel Stochastic Gradient Descent · ICML 2018 Can Decentralized Algorithms Outperform Centralized Algorithms? A Case Study for Decentralized Parallel Stochastic Gradient Descent · NIPS 2017 |
Machine learning › Optimization for machine learning
non-convex optimization |
0.6 | 1 | 2022 | Robustness to Unbounded Smoothness of Generalized SignSGD · NeurIPS 2022 |
Machine learning › Optimization for machine learning › stochastic gradient descent
signSGD |
0.6 | 1 | 2022 | Robustness to Unbounded Smoothness of Generalized SignSGD · NeurIPS 2022 |
Machine learning › Optimization for machine learning › non-convex optimization
unbounded smoothness |
0.6 | 1 | 2022 | Robustness to Unbounded Smoothness of Generalized SignSGD · NeurIPS 2022 |
Machine learning › Efficient and distributed learning › distributed training
stale gradient |
0.5 | 2 | 2017 | Model Accuracy and Runtime Tradeoff in Distributed Deep Learning: A Systematic Study · IJCAI 2017 Model Accuracy and Runtime Tradeoff in Distributed Deep Learning: A Systematic Study · ICDM 2016 |
Natural language and speech › Speech recognition and synthesis
acoustic modeling |
0.5 | 1 | 2021 | Asynchronous Decentralized Distributed Training of Acoustic Models · IEEE ACM Trans. Audio Speech Lang. Process. 2021 |
Machine learning › Efficient and distributed learning › distributed training › decentralized learning
decentralized asynchronous training |
0.5 | 1 | 2021 | Asynchronous Decentralized Distributed Training of Acoustic Models · IEEE ACM Trans. Audio Speech Lang. Process. 2021 |
Machine learning › Efficient and distributed learning › distributed training
communication-efficient training |
0.4 | 1 | 2020 | ScaleCom: Scalable Sparsified Gradient Compression for Communication-Efficient Distributed Training · NeurIPS 2020 |
Machine learning › Generative modeling › generative adversarial network
GAN training |
0.4 | 1 | 2020 | A Decentralized Parallel Algorithm for Training Generative Adversarial Nets · NeurIPS 2020 |
Machine learning › Generative modeling › generative adversarial network › GAN training
GAN training convergence |
0.4 | 1 | 2020 | Towards Better Understanding of Adaptive Gradient Algorithms in Generative Adversarial Nets · ICLR 2020 |
Machine learning › Generative modeling
generative adversarial network |
0.4 | 1 | 2020 | Towards Better Understanding of Adaptive Gradient Algorithms in Generative Adversarial Nets · ICLR 2020 |
Machine learning › Efficient and distributed learning › distributed training
gradient compression |
0.4 | 1 | 2020 | ScaleCom: Scalable Sparsified Gradient Compression for Communication-Efficient Distributed Training · NeurIPS 2020 |
Machine learning › Deep learning architectures and training
training dynamics |
0.4 | 1 | 2020 | Towards Better Understanding of Adaptive Gradient Algorithms in Generative Adversarial Nets · ICLR 2020 |
Debugging and program repair › automated program repair
concurrency bug fixing |
0.4 | 3 | 2013 | ConAir: featherweight concurrency bug recovery via single-threaded idempotent execution · ASPLOS 2013 Automated Concurrency-Bug Fixing · OSDI 2012 Automated atomicity-violation fixing · PLDI 2011 |
Concurrent programming
concurrency bug detection |
0.4 | 3 | 2013 | Efficient concurrency-bug detection across inputs · OOPSLA 2013 ConSeq: detecting concurrency bugs through sequential errors · ASPLOS 2011 ConMem: detecting severe concurrency bugs through an effect-oriented approach · ASPLOS 2010 |
Machine learning › Efficient and distributed learning
low-precision training |
0.4 | 1 | 2019 | Hybrid 8-bit Floating Point (HFP8) Training and Inference for Deep Neural Networks · NeurIPS 2019 |
Machine learning › Efficient and distributed learning
model compression |
0.4 | 1 | 2019 | Hybrid 8-bit Floating Point (HFP8) Training and Inference for Deep Neural Networks · NeurIPS 2019 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
low-precision arithmetic |
0.4 | 1 | 2019 | Hybrid 8-bit Floating Point (HFP8) Training and Inference for Deep Neural Networks · NeurIPS 2019 |
Machine learning › Optimization for machine learning
evolutionary computation |
0.3 | 1 | 2018 | Evolutionary Stochastic Gradient Descent for Optimization of Deep Neural Networks · NeurIPS 2018 |
Machine learning › Optimization for machine learning › stochastic search
population-based optimization |
0.3 | 1 | 2018 | Evolutionary Stochastic Gradient Descent for Optimization of Deep Neural Networks · NeurIPS 2018 |
Distributed systems › distributed machine learning › distributed training
communication-efficient distributed training |
0.3 | 1 | 2018 | Asynchronous Decentralized Parallel Stochastic Gradient Descent · ICML 2018 |
Parallel and multicore computing
parallel programming models |
0.3 | 1 | 2018 | Asynchronous Decentralized Parallel Stochastic Gradient Descent · ICML 2018 |
Machine learning › Optimization for machine learning › distributed optimization
decentralized SGD |
0.3 | 1 | 2017 | Can Decentralized Algorithms Outperform Centralized Algorithms? A Case Study for Decentralized Parallel Stochastic Gradient Descent · NIPS 2017 |
Energy-efficient computing › energy-efficient architecture
energy-efficient accelerator |
0.3 | 1 | 2017 | Accelerator Design for Deep Learning Training: Extended Abstract: Invited · DAC 2017 |
Methods — techniques the papers use, named apart from their topics
stochastic gradient descent · 2.2allreduce · 1.4decentralized training · 1.2delay-by-one scheme · 1.0asynchronous stochastic gradient descent · 0.8momentum · 0.6gradient clipping · 0.6adam · 0.6learning rate modulation · 0.5u-net · 0.4asynchronous distributed stochastic parallel gradient descent · 0.4adaptive gradient algorithms · 0.4CycleGAN · 0.4quantization · 0.4distributed training · 0.48-bit floating point · 0.4single-threaded idempotent execution · 0.2predictive detection · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Robustness to Unbounded Smoothness of Generalized SignSGDabstractTraditional analyses in non-convex optimization typically rely on the smoothness assumption, namely requiring the gradients to be Lipschitz. However, recent evidence shows that this smoothness condition does not capture the properties of some deep learning objective functions, including the ones involving Recurrent Neural Networks and LSTMs. Instead, they satisfy a much more relaxed condition, with potentially unbounded smoothness. Under this relaxed assumption, it has been theoretically and empirically shown that the gradient-clipped SGD has an advantage over the vanilla one. In this paper, we show that clipping is not indispensable for Adam-type algorithms in tackling such scenarios: we theoretically prove that a generalized SignSGD algorithm can obtain similar convergence rates as SGD with clipping but does not need explicit clipping at all. This family of algorithms on one end recovers SignSGD and on the other end closely resembles the popular Adam algorithm. Our analysis underlines the critical role that momentum plays in analyzing SignSGD-type and Adam-type algorithms: it not only reduces the effects of noise, thus removing the need for large mini-batch in previous analyses of SignSGD-type algorithms, but it also substantially reduces the effects of unbounded smoothness and gradient norms. To the best of our knowledge, this work is the first one showing the benefit of Adam-type algorithms compared with non-adaptive gradient algorithms such as gradient descent in the unbounded smoothness setting. We also compare these algorithms with popular optimizers on a set of deep learning tasks, observing that we can match the performance of Adam while beating others. Michael Crawshaw, Francesco Orabona, Wei Zhang 0022, Zhenxun Zhuang |
NeurIPS | 4 |
| 2021 | 4-Bit Quantization of LSTM-Based Speech Recognition ModelsabstractWe investigate the impact of aggressive low-precision representations of weights and activations in two families of large LSTM-based architectures for Automatic Speech Recognition (ASR): hybrid Deep Bidirectional LSTM -Hidden Markov Models (DBLSTM-HMMs) and Recurrent Neural Network -Transducers (RNN-Ts).Using a 4-bit integer representation, a naïve quantization approach applied to the LSTM portion of these models results in significant Word Error Rate (WER) degradation.On the other hand, we show that minimal accuracy loss is achievable with an appropriate choice of quantizers and initializations.In particular, we customize quantization schemes depending on the local properties of the network, improving recognition performance while limiting computational time.We demonstrate our solution on the Switchboard (SWB) and CallHome (CH) test sets of the NIST Hub5-2000 evaluation.DBLSTM-HMMs trained with 300 or 2000 hours of SWB data achieves <0.5% and <1% average WER degradation, respectively.On the more challenging RNN-T models, our quantization strategy limits degradation in 4-bit inference to 1.3%. Andrea Fasoli, Chia-Yu Chen, Mauricio J. Serrano, Xiao Sun 0013, Naigang Wang, Swagath Venkataramani, George Saon, Brian Kingsbury, Wei Zhang 0022, Zoltán Tüske, Kailash Gopalakrishnan |
Interspeech | 10 |
| 2021 | Asynchronous Decentralized Distributed Training of Acoustic ModelsabstractLarge-scale distributed training of deep acoustic models plays an important role in today's high-performance automatic speech recognition (ASR). In this paper we investigate a variety of asynchronous decentralized distributed training strategies based on data parallel stochastic gradient descent (SGD) to show their superior performance over the commonly-used synchronous distributed training via allreduce, especially when dealing with large batch sizes. Specifically, we study three variants of asynchronous decentralized parallel SGD (ADPSGD), namely, fixed and randomized communication patterns on a ring as well as a delay-by-one scheme. We introduce a mathematical model of ADPSGD, give its theoretical convergence rate, and compare the empirical convergence behavior and straggler resilience properties of the three variants. Experiments are carried out on an IBM supercomputer for training deep long short-term memory (LSTM) acoustic models on the 2000-hour Switchboard dataset. Recognition and speedup performance of the proposed strategies are evaluated under various training configurations. We show that ADPSGD with fixed and randomized communication patterns cope well with slow learners. When learners are equally fast, ADPSGD with the delay-by-one strategy has the fastest convergence with large batches. In particular, using the delay-by-one strategy, we can train the acoustic model in less than 2 hours using 128 V100 GPUs with competitive word error rates. Wei Zhang 0022, Abdullah Kayi, Ulrich Finkler, Brian Kingsbury, George Saon, David S. Kung 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | Improving Efficiency in Large-Scale Decentralized Distributed TrainingabstractDecentralized Parallel SGD (D-PSGD) and its asynchronous variant Asynchronous Parallel SGD (AD-PSGD) is a family of distributed learning algorithms that have been demonstrated to perform well for large-scale deep learning tasks. One drawback of (A)D-PSGD is that the spectral gap of the mixing matrix decreases when the number of learners in the system increases, which hampers convergence. In this paper, we investigate techniques to accelerate (A)D-PSGD based training by improving the spectral gap while minimizing the communication cost. We demonstrate the effectiveness of our proposed techniques by running experiments on the 2000-hour Switchboard speech recognition task and the ImageNet computer vision task. On an IBM P9 supercomputer, our system is able to train an LSTM acoustic model in 2.28 hours with 7.5% WER on the Hub5-2000 Switchboard (SWB) test set and 13.3% WER on the CallHome (CH) test set using 64 V100 GPUs and in 1.98 hours with 7.7% WER on SWB and 13.3% WER on CH using 128 V100 GPUs, the fastest training time reported to date. Wei Zhang 0022, Abdullah Kayi, Ulrich Finkler, Brian Kingsbury, George Saon, Youssef Mroueh, Alper Buyuktosunoglu, David S. Kung 0001, Michael Picheny |
ICASSP | 1 |
| 2020 | Towards Better Understanding of Adaptive Gradient Algorithms in Generative Adversarial Nets
Youssef Mroueh, Jerret Ross, Wei Zhang 0022, Tianbao Yang |
ICLR | 4 |
| 2020 | Map Generation from Large Scale Incomplete and Inaccurate Data LabelsabstractAccurately and globally mapping human infrastructure is an important and challenging task with applications in routing, regulation compliance monitoring, and natural disaster response management etc.. In this paper we present progress in developing an algorithmic pipeline and distributed compute system that automates the process of map creation using high resolution aerial images. Unlike previous studies, most of which use datasets that are available only in a few cities across the world, we utilizes publicly available imagery and map data, both of which cover the contiguous United States (CONUS). We approach the technical challenge of inaccurate and incomplete training data adopting state-of-the-art convolutional neural network architectures such as the U-Net and the CycleGAN to incrementally generate maps with increasingly more accurate and more complete labels of man-made infrastructure such as roads and houses. Since scaling the mapping task to CONUS calls for parallelization, we then adopted an asynchronous distributed stochastic parallel gradient descent training scheme to distribute the computational workload onto a cluster of GPUs with nearly linear speed-up. Conrad M. Albrecht, Wei Zhang 0022, Ulrich Finkler, David S. Kung 0001, Siyuan Lu 0003 |
KDD | 3 |
| 2020 | ScaleCom: Scalable Sparsified Gradient Compression for Communication-Efficient Distributed TrainingabstractLarge-scale distributed training of Deep Neural Networks (DNNs) on state-of-the-art platforms are expected to be severely communication constrained. To overcome this limitation, numerous gradient compression techniques have been proposed and have demonstrated high compression ratios. However, most existing compression methods do not scale well to large scale distributed systems (due to gradient build-up) and / or lack evaluations in large datasets. To mitigate these issues, we propose a new compression technique, Scalable Sparsified Gradient Compression (ScaleComp), that (i) leverages similarity in the gradient distribution amongst learners to provide a commutative compressor and keep communication cost constant to worker number and (ii) includes low-pass filter in local gradient accumulations to mitigate the impacts of large batch size training and significantly improve scalability. Using theoretical analysis, we show that ScaleComp provides favorable convergence guarantees and is compatible with gradient all-reduce techniques. Furthermore, we experimentally demonstrate that ScaleComp has small overheads, directly reduces gradient traffic and provides high compression rates (70-150X) and excellent scalability (up to 64-80 learners and 10X larger batch sizes over normal training) across a wide range of applications (image, language, and speech) without significant accuracy loss. Chia-Yu Chen, Jiamin Ni, Songtao Lu, Xiao Sun 0013, Naigang Wang, Swagath Venkataramani, Vijayalakshmi Srinivasan, Wei Zhang 0022, Kailash Gopalakrishnan |
NeurIPS | 10 |
| 2020 | A Decentralized Parallel Algorithm for Training Generative Adversarial NetsabstractGenerative Adversarial Networks (GANs) are a powerful class of generative models in the deep learning community. Current practice on large-scale GAN training utilizes large models and distributed large-batch training strategies, and is implemented on deep learning frameworks (e.g., TensorFlow, PyTorch, etc.) designed in a centralized manner. In the centralized network topology, every worker needs to either directly communicate with the central node or indirectly communicate with all other workers in every iteration. However, when the network bandwidth is low or network latency is high, the performance would be significantly degraded. Despite recent progress on decentralized algorithms for training deep neural networks, it remains unclear whether it is possible to train GANs in a decentralized manner. The main difficulty lies at handling the nonconvex-nonconcave min-max optimization and the decentralized communication simultaneously. In this paper, we address this difficulty by designing the \textbf{first gradient-based decentralized parallel algorithm} which allows workers to have multiple rounds of communications in one iteration and to update the discriminator and generator simultaneously, and this design makes it amenable for the convergence analysis of the proposed decentralized algorithm. Theoretically, our proposed decentralized algorithm is able to solve a class of non-convex non-concave min-max problems with provable non-asymptotic convergence to first-order stationary point. Experimental results on GANs demonstrate the effectiveness of the proposed algorithm. Wei Zhang 0022, Youssef Mroueh, Jarret Ross, Tianbao Yang |
NeurIPS | 2 |
| 2019 | Distributed Deep Learning Strategies for Automatic Speech RecognitionabstractIn this paper, we propose and investigate a variety of distributed deep learning strategies for automatic speech recognition (ASR) and evaluate them with a state-of-the-art Long short-term memory (LSTM) acoustic model on the 2000-hour Switchboard (SWB2000), which is one of the most widely used datasets for ASR performance benchmark. We first investigate what are the proper hyper-parameters (e.g., learning rate) to enable the training with sufficiently large batch size without impairing the model accuracy. We then implement various distributed strategies, including Synchronous (SYNC) , Asynchronous Decentralized Parallel SGD (ADPSGD) and the hybrid of the two HYBRID, to study their runtime/accuracy trade-off. We show that we can train the LSTM model using ADPSGD in 14 hours with 16 NVIDIA P100 GPUs to reach a 7.6% WER on the Hub5-2000 Switchboard (SWB) test set and a 13.1% WER on the Call-Home (CH) test set. Furthermore, we can train the model using HYBRID in 11.5 hours with 32 NVIDIA V100 GPUs without loss in accuracy. Wei Zhang 0022, Ulrich Finkler, Brian Kingsbury, George Saon, David S. Kung 0001, Michael Picheny |
ICASSP | 1 |
| 2019 | Large-Scale Mixed-Bandwidth Deep Neural Network Acoustic Modeling for Automatic Speech RecognitionabstractIn automatic speech recognition (ASR), wideband (WB) and narrowband (NB) speech signals with different sampling rates typically use separate acoustic models. Therefore mixed-bandwidth (MB) acoustic modeling has important practical values for ASR system deployment. In this paper, we extensively investigate large-scale MB deep neural network acoustic modeling for ASR using 1,150 hours of WB data and 2,300 hours of NB data. We study various MB strategies including downsampling, upsampling and bandwidth extension for MB acoustic modeling and evaluate their performance on 8 diverse WB and NB test sets from various application domains. To deal with the large amounts of training data, distributed training is carried out on multiple GPUs using synchronous data parallelism. Khoi-Nguyen C. Mac, Wei Zhang 0022, Michael Picheny |
INTERSPEECH | 3 |
| 2019 | A Highly Efficient Distributed Deep Learning System for Automatic Speech RecognitionabstractModern Automatic Speech Recognition (ASR) systems rely on distributed deep learning to for quick training completion.To enable efficient distributed training, it is imperative that the training algorithms can converge with a large mini-batch size.In this work, we discovered that Asynchronous Decentralized Parallel Stochastic Gradient Descent (ADPSGD) can work with much larger batch size than commonly used Synchronous SGD (SSGD) algorithm.On commonly used public SWB-300 and SWB-2000 ASR datasets, ADPSGD can converge with a batch size 3X as large as the one used in SSGD, thus enable training at a much larger scale.Further, we proposed a Hierarchical-ADPSGD (H-ADPSGD) system in which learners on the same computing node construct a super learner via a fast allreduce implementation, and super learners deploy ADPSGD algorithm among themselves.On a 64 Nvidia V100 GPU cluster connected via a 100Gb/s Ethernet network, our system is able to train SWB-2000 to reach a 7.6% WER on the Hub5-2000 Switchboard (SWB) test-set and a 13.2% WER on the Callhome (CH) test-set in 5.2 hours.To the best of our knowledge, this is the fastest ASR training system that attains this level of model accuracy for SWB-2000 task to be ever reported in the literature. Wei Zhang 0022, Ulrich Finkler, George Saon, Abdullah Kayi, Alper Buyuktosunoglu, Brian Kingsbury, David S. Kung 0001, Michael Picheny |
INTERSPEECH | 1 |
| 2019 | Hybrid 8-bit Floating Point (HFP8) Training and Inference for Deep Neural NetworksabstractReducing the numerical precision of data and computation is extremely effective in accelerating deep learning training workloads. Towards this end, 8-bit floating point representations (FP8) were recently proposed for DNN training. However, its applicability was demonstrated on a few selected models only and significant degradation is observed when popular networks such as MobileNet and Transformer are trained using FP8. This degradation is due to the inherent precision requirement difference in the forward and backward passes of DNN training. Using theoretical insights, we propose a hybrid FP8 (HFP8) format and DNN end-to-end distributed training procedure. We demonstrate, using HFP8, the successful training of deep learning models across a whole spectrum of applications including Image Classification, Object Detection, Language and Speech without accuracy degradation. Finally, we demonstrate that, by using the new 8 bit format, we can directly quantize a pre-trained model down to 8-bits without losing accuracy by simply fine-tuning batch normalization statistics. These novel techniques enable a new generations of 8-bit hardware that are robust for building and deploying neural network models. Xiao Sun 0013, Jungwook Choi, Chia-Yu Chen, Naigang Wang, Swagath Venkataramani, Vijayalakshmi Srinivasan, Wei Zhang 0022, Kailash Gopalakrishnan |
NeurIPS | 8 |
| 2018 | Asynchronous Decentralized Parallel Stochastic Gradient DescentabstractMost commonly used distributed machine learning systems are either synchronous or centralized asynchronous. Synchronous algorithms like AllReduce-SGD perform poorly in a heterogeneous environment, while asynchronous algorithms using a parameter server suffer from 1) communication bottleneck at parameter servers when workers are many, and 2) significantly worse convergence when the traffic to parameter server is congested. Can we design an algorithm that is robust in a heterogeneous environment, while being communication efficient and maintaining the best-possible convergence rate? In this paper, we propose an asynchronous decentralized stochastic gradient decent algorithm (AD-PSGD) satisfying all above expectations. Our theoretical analysis shows AD-PSGD converges at the optimal $O(1/\sqrt{K})$ rate as SGD and has linear speedup w.r.t. number of workers. Empirically, AD-PSGD outperforms the best of decentralized parallel SGD (D-PSGD), asynchronous parallel SGD (A-PSGD), and standard data parallel SGD (AllReduce-SGD), often by orders of magnitude in a heterogeneous environment. When training ResNet-50 on ImageNet with up to 128 GPUs, AD-PSGD converges (w.r.t epochs) similarly to the AllReduce-SGD, but each epoch can be up to 4-8x faster than its synchronous counterparts in a network-sharing HPC environment. To the best of our knowledge, AD-PSGD is the first asynchronous algorithm that achieves a similar epoch-wise convergence rate as AllReduce-SGD, at an over 100-GPU scale. Xiangru Lian, Wei Zhang 0022, Ce Zhang 0001, Ji Liu 0002 |
ICML | 2 |
| 2018 | Evolutionary Stochastic Gradient Descent for Optimization of Deep Neural NetworksabstractWe propose a population-based Evolutionary Stochastic Gradient Descent (ESGD) framework for optimizing deep neural networks. ESGD combines SGD and gradient-free evolutionary algorithms as complementary algorithms in one framework in which the optimization alternates between the SGD step and evolution step to improve the average fitness of the population. With a back-off strategy in the SGD step and an elitist strategy in the evolution step, it guarantees that the best fitness in the population will never degrade. In addition, individuals in the population optimized with various SGD-based optimizers using distinct hyper-parameters in the SGD step are considered as competing species in a coevolution setting such that the complementarity of the optimizers is also taken into account. The effectiveness of ESGD is demonstrated across multiple applications including speech recognition, image recognition and language modeling, using networks with a variety of deep architectures. Wei Zhang 0022, Zoltán Tüske, Michael Picheny |
NeurIPS | 2 |
| 2017 | Accelerator Design for Deep Learning Training: Extended Abstract: InvitedabstractDeep Neural Networks (DNNs) have emerged as a powerful and versatile set of techniques showing successes on challenging artificial intelligence (AI) problems. Applications in domains such as image/video processing, autonomous cars, natural language processing, speech synthesis and recognition, genomics and many others have embraced deep learning as the foundation. DNNs achieve superior accuracy for these applications with high computational complexity using very large models which require 100s of MBs of data storage, exaops of computation and high bandwidth for data movement. In spite of these impressive advances, it still takes days to weeks to train state of the art Deep Networks on large datasets - which directly limits the pace of innovation and adoption. In this paper, we present a multi-pronged approach to address the challenges in meeting both the throughput and the energy efficiency goals for DNN training. Ankur Agrawal, Chia-Yu Chen, Jungwook Choi, Kailash Gopalakrishnan, Jinwook Oh, Sunil Shukla, Vijayalakshmi Srinivasan, Swagath Venkataramani, Wei Zhang 0022 |
DAC | 9 |
| 2017 | Model Accuracy and Runtime Tradeoff in Distributed Deep Learning: A Systematic StudyabstractDeep learning with a large number of parame-ters requires distributed training, where model accuracy and runtime are two important factors to be considered. However, there has been no systematic study of the tradeoff between these two factors during the model training process. This paper presents Rudra, a parameter server based distributed computing framework tuned for training large-scale deep neural networks. Using variants of the asynchronous stochastic gradient descent algorithm we study the impact of synchronization protocol, stale gradient updates, minibatch size, learning rates, and number of learners on runtime performance and model accuracy. We introduce a new learningrate modulation strategy to counter the effect of stale gradients and propose a new synchronization protocol that can effectively bound the staleness in gradients, improve runtime performance and achieve good model accuracy. Our empirical investigation reveals a principled approach for distributed training of neural networks: the mini-batch size per learner should be reduced as more learners are added to the system to preserve the model accuracy. We validate this approach using commonly-used image classification benchmarks: CIFAR10 and ImageNet. Suyog Gupta, Wei Zhang 0022, Fei Wang 0001 |
IJCAI | 2 |
| 2017 | Can Decentralized Algorithms Outperform Centralized Algorithms? A Case Study for Decentralized Parallel Stochastic Gradient DescentabstractMost distributed machine learning systems nowadays, including TensorFlow and CNTK, are built in a centralized fashion. One bottleneck of centralized algorithms lies on high communication cost on the central node. Motivated by this, we ask, can decentralized algorithms be faster than its centralized counterpart? Although decentralized PSGD (D-PSGD) algorithms have been studied by the control community, existing analysis and theory do not show any advantage over centralized PSGD (C-PSGD) algorithms, simply assuming the application scenario where only the decentralized network is available. In this paper, we study a D-PSGD algorithm and provide the first theoretical analysis that indicates a regime in which decentralized algorithms might outperform centralized algorithms for distributed stochastic gradient descent. This is because D-PSGD has comparable total computational complexities to C-PSGD but requires much less communication cost on the busiest node. We further conduct an empirical study to validate our theoretical analysis across multiple frameworks (CNTK and Torch), different network configurations, and computation platforms up to 112 GPUs. On network configurations with low bandwidth or high latency, D-PSGD can be up to one order of magnitude faster than its well-optimized centralized counterparts. Xiangru Lian, Ce Zhang 0001, Huan Zhang 0001, Cho-Jui Hsieh, Wei Zhang 0022, Ji Liu 0002 |
NIPS | 5 |
| 2016 | Model Accuracy and Runtime Tradeoff in Distributed Deep Learning: A Systematic StudyabstractDeep learning with a large number of parametersrequires distributed training, where model accuracy and runtimeare two important factors to be considered. However, there hasbeen no systematic study of the tradeoff between these two factorsduring the model training process. This paper presents Rudra, aparameter server based distributed computing framework tunedfor training large-scale deep neural networks. Using variants ofthe asynchronous stochastic gradient descent algorithm we studythe impact of synchronization protocol, stale gradient updates, minibatch size, learning rates, and number of learners on runtimeperformance and model accuracy. We introduce a new learningrate modulation strategy to counter the effect of stale gradientsand propose a new synchronization protocol that can effectivelybound the staleness in gradients, improve runtime performanceand achieve good model accuracy. Our empirical investigationreveals a principled approach for distributed training of neuralnetworks: the mini-batch size per learner should be reducedas more learners are added to the system to preserve the modelaccuracy. We validate this approach using commonly-used imageclassification benchmarks: CIFAR10 and ImageNet. Suyog Gupta, Wei Zhang 0022, Fei Wang 0001 |
ICDM | 2 |
| 2016 | Staleness-Aware Async-SGD for Distributed Deep Learning
Wei Zhang 0022, Suyog Gupta, Xiangru Lian, Ji Liu 0002 |
IJCAI | 1 |
| 2015 | Fixing, preventing, and recovering from concurrency bugs
Dongdong Deng, Guoliang Jin, Marc de Kruijf, Ben Liblit, Shan Lu 0001, Shanxiang Qi, Jinglei Ren, Karthikeyan Sankaralingam, Linhai Song, Yongwei Wu 0001, Wei Zhang 0022 |
Sci. China Inf. Sci. | 13 |
| 2013 | ConAir: featherweight concurrency bug recovery via single-threaded idempotent executionabstractMany concurrency bugs are hidden in deployed software and cause severe failures for end-users. When they finally manifest and become known by developers, they are difficult to fix correctly. To support end-users, we need techniques that help software survive hidden concurrency bugs during production runs. To help developers, we need techniques that fix exposed concurrency bugs. Wei Zhang 0022, Marc de Kruijf, Shan Lu 0001, Karthikeyan Sankaralingam |
ASPLOS | 1 |
| 2013 | Efficient concurrency-bug detection across inputsabstractIn the multi-core era, it is critical to efficiently test multi-threaded software and expose concurrency bugs before software release. Previous work has made significant progress in detecting and validating concurrency bugs under a given input. Unfortunately, software testing always faces large sets of test inputs, and existing techniques are still too expensive to be applied to every test input in practice. Dongdong Deng, Wei Zhang 0022, Shan Lu 0001 |
OOPSLA | 2 |
| 2013 | ConMem: Detecting Crash-Triggering Concurrency Bugs through an Effect-Oriented ApproachabstractMulticore technology is making concurrent programs increasingly pervasive. Unfortunately, it is difficult to deliver reliable concurrent programs, because of the huge and nondeterministic interleaving space. In reality, without the resources to thoroughly check the interleaving space, critical concurrency bugs can slip into production versions and cause failures in the field. Approaches to making the best use of the limited resources and exposing severe concurrency bugs before software release would be desirable. Unlike previous work that focuses on bugs caused by specific interleavings (e.g., races and atomicity violations), this article targets concurrency bugs that result in one type of severe effect: program crashes. Our study of the error-propagation process of real-world concurrency bugs reveals a common pattern (50% in our nondeadlock concurrency bug set) that is highly correlated with program crashes. We call this pattern concurrency-memory bugs: buggy interleavings directly cause memory bugs (NULL-pointer-dereferences, dangling-pointers, buffer-overflows, uninitialized-reads) on shared memory objects. Guided by this study, we built ConMem to monitor program execution, analyze memory accesses and synchronizations, and predictively detect these common and severe concurrency-memory bugs. We also built a validator,ConMem-v, to automatically prune false positives by enforcing potential bug-triggering interleavings. We evaluated ConMem using 7 open-source programs with 10 real-world concurrency bugs. ConMem detects more tested bugs (9 out of 10 bugs) than a lock-set-based race detector and an unserializable-interleaving detector, which detect 4 and 6 bugs, respectively, with a false-positive rate about one tenth of the compared tools. ConMem-v further prunes out all the false positives. ConMem has reasonable overhead suitable for development usage. Wei Zhang 0022, Junghee Lim, Shan Lu 0001, Thomas W. Reps |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2012 | Automated Concurrency-Bug Fixing
Guoliang Jin, Wei Zhang 0022, Dongdong Deng |
OSDI | 2 |
| 2011 | ConSeq: detecting concurrency bugs through sequential errorsabstractConcurrency bugs are caused by non-deterministic interleavings between shared memory accesses. Their effects propagate through data and control dependences until they cause software to crash, hang, produce incorrect output, etc. The lifecycle of a bug thus consists of three phases: (1) triggering, (2) propagation, and (3) failure. Wei Zhang 0022, Junghee Lim, Ramya Olichandran, Joel Scherpelz, Guoliang Jin, Shan Lu 0001, Thomas W. Reps |
ASPLOS | 1 |
| 2011 | Automated atomicity-violation fixingabstractFixing software bugs has always been an important and time-consuming process in software development. Fixing concurrency bugs has become especially critical in the multicore era. However, fixing concurrency bugs is challenging, in part due to non-deterministic failures and tricky parallel reasoning. Beyond correctly fixing the original problem in the software, a good patch should also avoid introducing new bugs, degrading performance unnecessarily, or damaging software readability. Existing tools cannot automate the whole fixing process and provide good-quality patches. Guoliang Jin, Linhai Song, Wei Zhang 0022, Shan Lu 0001, Ben Liblit |
PLDI | 3 |
| 2010 | ConMem: detecting severe concurrency bugs through an effect-oriented approach
Wei Zhang 0022, Shan Lu 0001 |
ASPLOS | 1 |
| 2010 | Applying scalable phonetic context similarity in unit selection of concatenative text-to-speech
Wei Zhang 0022 |
INTERSPEECH | 1 |
| 2008 | Developing high performance asr in the IBM multilingual speech-to-speech translation systemabstractThis paper presents our recent development of the real-time speech recognition component in the IBM English/Iraqi Arabic speech-to-speech translation system for the DARPA Transtac project. We describe the details of the acoustic and language modeling that lead to high recognition accuracy and noise robustness and give the performance of the system on the evaluation sets of spontaneous conversational speech. We also introduce the streaming decoding structure and several speedup techniques that achieves best recognition accuracy at about 0.3×RT recognition speed. Liang Gu, Bing Xiang, Wei Zhang 0022 |
ICASSP | 4 |
| 2005 | Toward multiple-language TTS: experiments in English and MandarinabstractText-to-speech systems have dramatically improved in recent years through the use of corpus-based concatenative approaches, and we are beginning to see an interest in endowing them with the ability to handle more than the native language for which they have been developed. In this paper we present ongoing work at IBM in text-to-speech systems that can produce high-quality synthesis in more than one language. We illustrate the discussion with a case study in which two systems, originally developed to support English and Mandarin respectively, have been extended to support each other’s languages. We describe the challenges faced when adapting one system to a different target language, propose adaptation solutions, and present the results of perceptual tests carried out to evaluate how the approaches compare with the performance of the native systems. Raul Fernandez, Wei Zhang 0022, Ellen Eide, Raimo Bakis, Wael Hamza, Michael Picheny, John F. Pitrelli, Yong Qing, Zhiwei Shuang, Li Qin Shen |
INTERSPEECH | 2 |