Itay Hubara

dblp:155/1915 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
4since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 4 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
12 papers
Efficient and distributed learning · 61% Deep learning architectures and training · 17% Optimization for machine learning · 7%
Computer architecture, parallel and distributed computing, and storage systems
3 papers
Hardware accelerators and domain-specific architectures · 81% Performance modeling and evaluation · 16% GPUs and heterogeneous computing · 3%

Topics — the 30 heaviest of 38, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
model compression
3.572024
Towards Cheaper Inference in Deep Networks with Lower Bit-Width Accumulators · ICLR 2024
Minimum Variance Unbiased N: M Sparsity for the Neural Gradients · ICLR 2023
Accelerated Sparse Neural Training: A Provable and Efficient Method to Find N: M Transposable Masks · NeurIPS 2021
Machine learning › Efficient and distributed learning › model compression
quantization
1.432024
Towards Cheaper Inference in Deep Networks with Lower Bit-Width Accumulators · ICLR 2024
Scalable methods for 8-bit training of neural networks · NeurIPS 2018
Quantized Neural Networks: Training Neural Networks with Low Precision Weights and Activations · J. Mach. Learn. Res. 2017
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.922024
Towards Cheaper Inference in Deep Networks with Lower Bit-Width Accumulators · ICLR 2024
MLPerf Inference Benchmark · ISCA 2020
Hardware accelerators and domain-specific architectures › machine learning accelerator
low-precision arithmetic
0.812024
Towards Cheaper Inference in Deep Networks with Lower Bit-Width Accumulators · ICLR 2024
Machine learning › Deep learning architectures and training › training optimization
large-batch training
0.722020
Augment Your Batch: Improving Generalization Through Instance Repetition · CVPR 2020
Train longer, generalize better: closing the generalization gap in large batch training of neural networks · NIPS 2017
Machine learning › Optimization for machine learning
stochastic gradient descent
0.722020
Augment Your Batch: Improving Generalization Through Instance Repetition · CVPR 2020
Train longer, generalize better: closing the generalization gap in large batch training of neural networks · NIPS 2017
Machine learning › Efficient and distributed learning › model compression › sparsity
n:m sparsity
0.712023
Minimum Variance Unbiased N: M Sparsity for the Neural Gradients · ICLR 2023
Machine learning › Efficient and distributed learning › model compression
sparsity
0.712023
Minimum Variance Unbiased N: M Sparsity for the Neural Gradients · ICLR 2023
Machine learning › Efficient and distributed learning › model quantization
bit-width allocation
0.512021
Accurate Post Training Quantization With Small Calibration Sets · ICML 2021
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › constraint optimization
integer programming
0.512021
Accurate Post Training Quantization With Small Calibration Sets · ICML 2021
Machine learning › Efficient and distributed learning › model compression › quantization
post-training quantization
0.512021
Accurate Post Training Quantization With Small Calibration Sets · ICML 2021
Machine learning › Efficient and distributed learning › model compression
sparse neural network
0.512021
Accelerated Sparse Neural Training: A Provable and Efficient Method to Find N: M Transposable Masks · NeurIPS 2021
Machine learning › Efficient and distributed learning › model compression › pruning
structured pruning
0.512021
Accelerated Sparse Neural Training: A Provable and Efficient Method to Find N: M Transposable Masks · NeurIPS 2021
Machine learning › Deep learning architectures and training
data augmentation
0.412020
Augment Your Batch: Improving Generalization Through Instance Repetition · CVPR 2020
Machine learning › Efficient and distributed learning › model compression › quantization › post-training quantization
data-free quantization
0.412020
The Knowledge Within: Methods for Data-Free Model Compression · CVPR 2020
Machine learning › Generative modeling › synthetic data generation
synthetic sample generation
0.412020
The Knowledge Within: Methods for Data-Free Model Compression · CVPR 2020
Machine learning › Deep learning architectures and training › regularization
training regularization
0.412020
Augment Your Batch: Improving Generalization Through Instance Repetition · CVPR 2020
Performance modeling and evaluation
benchmarking
0.412020
MLPerf Inference Benchmark · ISCA 2020
Machine learning › Deep learning architectures and training › normalization
batch normalization
0.312018
Scalable methods for 8-bit training of neural networks · NeurIPS 2018
Machine learning › Trustworthy machine learning › calibration
classifier calibration
0.312018
Fix your classifier: the marginal value of training the last weight layer · ICLR (Poster) 2018
Machine learning › Transfer learning and domain adaptation
fine-tuning
0.312018
Fix your classifier: the marginal value of training the last weight layer · ICLR (Poster) 2018
Machine learning › Transfer learning and domain adaptation › fine-tuning
last-layer retraining
0.312018
Fix your classifier: the marginal value of training the last weight layer · ICLR (Poster) 2018
Machine learning › Learning theory › generalization error
generalization gap
0.312017
Train longer, generalize better: closing the generalization gap in large batch training of neural networks · NIPS 2017
Machine learning › Optimization for machine learning
learning rate schedule
0.312017
Train longer, generalize better: closing the generalization gap in large batch training of neural networks · NIPS 2017
Machine learning › Efficient and distributed learning › model compression › quantization › low-precision computation
low-precision neural network
0.312017
Quantized Neural Networks: Training Neural Networks with Low Precision Weights and Activations · J. Mach. Learn. Res. 2017
Machine learning › Efficient and distributed learning › model compression › quantization
quantized training
0.312017
Quantized Neural Networks: Training Neural Networks with Low Precision Weights and Activations · J. Mach. Learn. Res. 2017
Machine learning › Efficient and distributed learning › model compression › quantization › quantized neural network
binary neural network
0.212016
Binarized Neural Networks · NIPS 2016
Machine learning › Efficient and distributed learning › model compression
lightweight neural network
0.212016
Binarized Neural Networks · NIPS 2016
Machine learning › Efficient and distributed learning › model compression › quantization
quantized neural network
0.212016
Binarized Neural Networks · NIPS 2016
Hardware accelerators and domain-specific architectures › quantization
binarized neural network
0.212016
Binarized Neural Networks · NIPS 2016

Methods — techniques the papers use, named apart from their topics

quantization · 2.3gradient approximation · 1.5minimum variance unbiased estimation · 0.7n:m sparsity · 0.5minimum cost flow · 0.5integer programming · 0.5calibration set optimization · 0.5synthetic data generation · 0.4batch normalization statistics · 0.4batch augmentation · 0.4gradient-based training · 0.2bitwise operations · 0.2
YearPublicationVenuePosition
2024 Towards Cheaper Inference in Deep Networks with Lower Bit-Width Accumulators
abstract
The majority of the research on the quantization of Deep Neural Networks (DNNs) is focused on reducing the precision of tensors visible by high-level frameworks (e.g., weights, activations, and gradients). However, current hardware still relies on high-accuracy core operations. Most significant is the operation of accumulating products. This high-precision accumulation operation is gradually becoming the main computational bottleneck. This is because, so far, the usage of low-precision accumulators led to a significant degradation in performance. In this work, we present a simple method to train and fine-tune DNNs, to allow, for the first time, utilization of cheaper, $12$-bits accumulators, with no significant degradation in accuracy. Lastly, we show that as we decrease the accumulation precision further, using fine-grained gradient approximations can improve the DNN accuracy.
Yaniv Blumenfeld, Itay Hubara, Daniel Soudry
ICLR2
2023 Minimum Variance Unbiased N: M Sparsity for the Neural Gradients
Brian Chmiel, Itay Hubara, Ron Banner, Daniel Soudry
ICLR2
2021 Accurate Post Training Quantization With Small Calibration Sets
abstract
Lately, post-training quantization methods have gained considerable attention, as they are simple to use, and require only a small unlabeled calibration set. This small dataset cannot be used to fine-tune the model without significant over-fitting. Instead, these methods only use the calibration set to set the activations’ dynamic ranges. However, such methods always resulted in significant accuracy degradation, when used below 8-bits (except on small datasets). Here we aim to break the 8-bit barrier. To this end, we minimize the quantization errors of each layer or block separately by optimizing its parameters over the calibration set. We empirically demonstrate that this approach is: (1) much less susceptible to over-fitting than the standard fine-tuning approaches, and can be used even on a very small calibration set; and (2) more powerful than previous methods, which only set the activations’ dynamic ranges. We suggest two flavors for our method, parallel and sequential aim for a fixed and flexible bit-width allocation. For the latter, we demonstrate how to optimally allocate the bit-widths for each layer, while constraining accuracy degradation or model compression by proposing a novel integer programming formulation. Finally, we suggest model global statistics tuning, to correct biases introduced during quantization. Together, these methods yield state-of-the-art results for both vision and text models. For instance, on ResNet50, we obtain less than 1% accuracy degradation — with 4-bit weights and activations in all layers, but first and last. The suggested methods are two orders of magnitude faster than the traditional Quantize Aware Training approach used for lower than 8-bit quantization. We open-sourced our code \textit{https://github.com/papers-submission/CalibTIP}.
Itay Hubara, Yury Nahshan, Yair Hanani, Ron Banner, Daniel Soudry
ICML1
2021 Accelerated Sparse Neural Training: A Provable and Efficient Method to Find N: M Transposable Masks
abstract
Unstructured pruning reduces the memory footprint in deep neural networks (DNNs). Recently, researchers proposed different types of structural pruning intending to reduce also the computation complexity. In this work, we first suggest a new measure called mask-diversity which correlates with the expected accuracy of the different types of structural pruning. We focus on the recently suggested N:M fine-grained block sparsity mask, in which for each block of M weights, we have at least N zeros. While N:M fine-grained block sparsity allows acceleration in actual modern hardware, it can be used only to accelerate the inference phase. In order to allow for similar accelerations in the training phase, we suggest a novel transposable fine-grained sparsity mask, where the same mask can be used for both forward and backward passes. Our transposable mask guarantees that both the weight matrix and its transpose follow the same sparsity pattern; thus, the matrix multiplication required for passing the error backward can also be accelerated. We formulate the problem of finding the optimal transposable-mask as a minimum-cost flow problem. Additionally, to speed up the minimum-cost flow computation, we also introduce a fast linear-time approximation that can be used when the masks dynamically change during training. Our experiments suggest a 2x speed-up in the matrix multiplications with no accuracy degradation over vision and language models. Finally, to solve the problem of switching between different structure constraints, we suggest a method to convert a pre-trained model with unstructured sparsity to an N:M fine-grained block sparsity model with little to no training. A reference implementation can be found at https://github.com/papers-submission/structuredtransposablemasks.
Itay Hubara, Brian Chmiel, Moshe Island, Ron Banner, Joseph Naor, Daniel Soudry
NeurIPS1
2020 The Knowledge Within: Methods for Data-Free Model Compression
abstract
Background: Recently, an extensive amount of research has been focused on compressing and accelerating Deep Neural Networks (DNN). So far, high compression rate algorithms require part of the training dataset for a low precision calibration, or a fine-tuning process. However, this requirement is unacceptable when the data is unavailable or contains sensitive information, as in medical and biometric use-cases. Contributions: We present three methods for generating synthetic samples from trained models. Then, we demonstrate how these samples can be used to calibrate and fine-tune quantized models without using any real data in the process. Our best performing method has a negligible accuracy degradation compared to the original training set. This method, which leverages intrinsic batch normalization layers' statistics of the trained model, can be used to evaluate data similarity. Our approach opens a path towards genuine data-free model compression, alleviating the need for training data during model deployment.
Matan Haroush, Itay Hubara, Elad Hoffer, Daniel Soudry
CVPR2
2020 Augment Your Batch: Improving Generalization Through Instance Repetition
abstract
Large-batch SGD is important for scaling training of deep neural networks. However, without fine-tuning hyperparameter schedules, the generalization of the model may be hampered. We propose to use batch augmentation: replicating instances of samples within the same batch with different data augmentations. Batch augmentation acts as a regularizer and an accelerator, increasing both generalization and performance scaling for a fixed budget of optimization steps. We analyze the effect of batch augmentation on gradient variance and show that it empirically improves convergence for a wide variety of networks and datasets. Our results show that batch augmentation reduces the number of necessary SGD updates to achieve the same accuracy as the state-of-the-art. Overall, this simple yet effective method enables faster training and better generalization by allowing more computational resources to be used concurrently.
Elad Hoffer, Tal Ben-Nun, Itay Hubara, Niv Giladi, Torsten Hoefler, Daniel Soudry
CVPR3
2020 MLPerf Inference Benchmark
abstract
Machine-learning (ML) hardware and software system demand is burgeoning. Driven by ML applications, the number of different ML inference systems has exploded. Over 100 organizations are building ML inference chips, and the systems that incorporate existing models span at least three orders of magnitude in power consumption and five orders of magnitude in performance; they range from embedded devices to data-center solutions. Fueling the hardware are a dozen or more software frameworks and libraries. The myriad combinations of ML hardware and ML software make assessing ML-system performance in an architecture-neutral, representative, and reproducible manner challenging. There is a clear need for industry-wide standard ML benchmarking and evaluation criteria. MLPerf Inference answers that call. In this paper, we present our benchmarking method for evaluating ML inference systems. Driven by more than 30 organizations as well as more than 200 ML engineers and practitioners, MLPerf prescribes a set of rules and best practices to ensure comparability across systems with wildly differing architectures. The first call for submissions garnered more than 600 reproducible inference-performance measurements from 14 organizations, representing over 30 systems that showcase a wide range of capabilities. The submissions attest to the benchmark’s flexibility and adaptability.
Vijay Janapa Reddi, David Kanter, Peter Mattson, Guenther Schmuelling, Carole-Jean Wu, Maximilien Breughe, Mark Charlebois, William Chou, Ramesh Chukka, Cody Coleman, Sam Davis, Gregory Frederick Diamos, Jared Duke, David Fick, J. Scott Gardner, Itay Hubara, Sachin Idgunji, Thomas B. Jablin, Jeff Jiao, Tom St. John, Pankaj Kanwar, Jeffery Liao, Anton Lokhmotov, Francisco Massa, Peng Meng, Paulius Micikevicius, Colin Osborne, Gennady Pekhimenko, Arun Tejusve Raghunath Rajan, Dilip Sequeira, Ashish Sirasao, Fei Sun 0002, Michael Thomson, Frank Wei, Ephrem Wu, Lingjie Xu, Koichi Yamada, George Yuan, Aaron Zhong, Peizhao Zhang
ISCA19
2018 Fix your classifier: the marginal value of training the last weight layer
Elad Hoffer, Itay Hubara, Daniel Soudry
ICLR (Poster)2
2018 Scalable methods for 8-bit training of neural networks
abstract
Quantized Neural Networks (QNNs) are often used to improve network efficiency during the inference phase, i.e. after the network has been trained. Extensive research in the field suggests many different quantization schemes. Still, the number of bits required, as well as the best quantization scheme, are yet unknown. Our theoretical analysis suggests that most of the training process is robust to substantial precision reduction, and points to only a few specific operations that require higher precision. Armed with this knowledge, we quantize the model parameters, activations and layer gradients to 8-bit, leaving at higher precision only the final step in the computation of the weight gradients. Additionally, as QNNs require batch-normalization to be trained at high precision, we introduce Range Batch-Normalization (BN) which has significantly higher tolerance to quantization noise and improved computational complexity. Our simulations show that Range BN is equivalent to the traditional batch norm if a precise scale adjustment, which can be approximated analytically, is applied. To the best of the authors' knowledge, this work is the first to quantize the weights, activations, as well as a substantial volume of the gradients stream, in all layers (including batch normalization) to 8-bit while showing state-of-the-art results over the ImageNet-1K dataset.
Ron Banner, Itay Hubara, Elad Hoffer, Daniel Soudry
NeurIPS2
2017 Train longer, generalize better: closing the generalization gap in large batch training of neural networks
abstract
Background: Deep learning models are typically trained using stochastic gradient descent or one of its variants. These methods update the weights using their gradient, estimated from a small fraction of the training data. It has been observed that when using large batch sizes there is a persistent degradation in generalization performance - known as the "generalization gap" phenomenon. Identifying the origin of this gap and closing it had remained an open problem. Contributions: We examine the initial high learning rate training phase. We find that the weight distance from its initialization grows logarithmically with the number of weight updates. We therefore propose a "random walk on a random landscape" statistical model which is known to exhibit similar "ultra-slow" diffusion behavior. Following this hypothesis we conducted experiments to show empirically that the "generalization gap" stems from the relatively small number of updates rather than the batch size, and can be completely eliminated by adapting the training regime used. We further investigate different techniques to train models in the large-batch regime and present a novel algorithm named "Ghost Batch Normalization" which enables significant decrease in the generalization gap without increasing the number of updates. To validate our findings we conduct several additional experiments on MNIST, CIFAR-10, CIFAR-100 and ImageNet. Finally, we reassess common practices and beliefs concerning training of deep models and suggest they may not be optimal to achieve good generalization.
Elad Hoffer, Itay Hubara, Daniel Soudry
NIPS2
2017 Quantized Neural Networks: Training Neural Networks with Low Precision Weights and Activations
Itay Hubara, Matthieu Courbariaux, Daniel Soudry, Ran El-Yaniv, Yoshua Bengio
J. Mach. Learn. Res.1
2016 Binarized Neural Networks
abstract
We introduce a method to train Binarized Neural Networks (BNNs) - neural networks with binary weights and activations at run-time. At train-time the binary weights and activations are used for computing the parameter gradients. During the forward pass, BNNs drastically reduce memory size and accesses, and replace most arithmetic operations with bit-wise operations, which is expected to substantially improve power-efficiency. To validate the effectiveness of BNNs, we conducted two sets of experiments on the Torch7 and Theano frameworks. On both, BNNs achieved nearly state-of-the-art results over the MNIST, CIFAR-10 and SVHN datasets. We also report our preliminary results on the challenging ImageNet dataset. Last but not least, we wrote a binary matrix multiplication GPU kernel with which it is possible to run our MNIST BNN 7 times faster than with an unoptimized GPU kernel, without suffering any loss in classification accuracy. The code for training and running our BNNs is available on-line.
Itay Hubara, Matthieu Courbariaux, Daniel Soudry, Ran El-Yaniv, Yoshua Bengio
NIPS1
2014 Expectation Backpropagation: Parameter-Free Training of Multilayer Neural Networks with Continuous or Discrete Weights
Daniel Soudry, Itay Hubara, Ron Meir
NIPS2