VLDB 2026 Research / reviewers in the wild / expert
Itay Hubara
dblp:155/1915
· DBLP profile ↗
13ranked-venue papers
4as first author
4since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 4 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
12 papers |
Efficient and distributed learning · 61% Deep learning architectures and training · 17% Optimization for machine learning · 7% | |
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Hardware accelerators and domain-specific architectures · 81% Performance modeling and evaluation · 16% GPUs and heterogeneous computing · 3% |
Topics — the 30 heaviest of 38, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
model compression |
3.5 | 7 | 2024 | Towards Cheaper Inference in Deep Networks with Lower Bit-Width Accumulators · ICLR 2024 Minimum Variance Unbiased N: M Sparsity for the Neural Gradients · ICLR 2023 Accelerated Sparse Neural Training: A Provable and Efficient Method to Find N: M Transposable Masks · NeurIPS 2021 |
Machine learning › Efficient and distributed learning › model compression
quantization |
1.4 | 3 | 2024 | Towards Cheaper Inference in Deep Networks with Lower Bit-Width Accumulators · ICLR 2024 Scalable methods for 8-bit training of neural networks · NeurIPS 2018 Quantized Neural Networks: Training Neural Networks with Low Precision Weights and Activations · J. Mach. Learn. Res. 2017 |
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.9 | 2 | 2024 | Towards Cheaper Inference in Deep Networks with Lower Bit-Width Accumulators · ICLR 2024 MLPerf Inference Benchmark · ISCA 2020 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
low-precision arithmetic |
0.8 | 1 | 2024 | Towards Cheaper Inference in Deep Networks with Lower Bit-Width Accumulators · ICLR 2024 |
Machine learning › Deep learning architectures and training › training optimization
large-batch training |
0.7 | 2 | 2020 | Augment Your Batch: Improving Generalization Through Instance Repetition · CVPR 2020 Train longer, generalize better: closing the generalization gap in large batch training of neural networks · NIPS 2017 |
Machine learning › Optimization for machine learning
stochastic gradient descent |
0.7 | 2 | 2020 | Augment Your Batch: Improving Generalization Through Instance Repetition · CVPR 2020 Train longer, generalize better: closing the generalization gap in large batch training of neural networks · NIPS 2017 |
Machine learning › Efficient and distributed learning › model compression › sparsity
n:m sparsity |
0.7 | 1 | 2023 | Minimum Variance Unbiased N: M Sparsity for the Neural Gradients · ICLR 2023 |
Machine learning › Efficient and distributed learning › model compression
sparsity |
0.7 | 1 | 2023 | Minimum Variance Unbiased N: M Sparsity for the Neural Gradients · ICLR 2023 |
Machine learning › Efficient and distributed learning › model quantization
bit-width allocation |
0.5 | 1 | 2021 | Accurate Post Training Quantization With Small Calibration Sets · ICML 2021 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › constraint optimization
integer programming |
0.5 | 1 | 2021 | Accurate Post Training Quantization With Small Calibration Sets · ICML 2021 |
Machine learning › Efficient and distributed learning › model compression › quantization
post-training quantization |
0.5 | 1 | 2021 | Accurate Post Training Quantization With Small Calibration Sets · ICML 2021 |
Machine learning › Efficient and distributed learning › model compression
sparse neural network |
0.5 | 1 | 2021 | Accelerated Sparse Neural Training: A Provable and Efficient Method to Find N: M Transposable Masks · NeurIPS 2021 |
Machine learning › Efficient and distributed learning › model compression › pruning
structured pruning |
0.5 | 1 | 2021 | Accelerated Sparse Neural Training: A Provable and Efficient Method to Find N: M Transposable Masks · NeurIPS 2021 |
Machine learning › Deep learning architectures and training
data augmentation |
0.4 | 1 | 2020 | Augment Your Batch: Improving Generalization Through Instance Repetition · CVPR 2020 |
Machine learning › Efficient and distributed learning › model compression › quantization › post-training quantization
data-free quantization |
0.4 | 1 | 2020 | The Knowledge Within: Methods for Data-Free Model Compression · CVPR 2020 |
Machine learning › Generative modeling › synthetic data generation
synthetic sample generation |
0.4 | 1 | 2020 | The Knowledge Within: Methods for Data-Free Model Compression · CVPR 2020 |
Machine learning › Deep learning architectures and training › regularization
training regularization |
0.4 | 1 | 2020 | Augment Your Batch: Improving Generalization Through Instance Repetition · CVPR 2020 |
Performance modeling and evaluation
benchmarking |
0.4 | 1 | 2020 | MLPerf Inference Benchmark · ISCA 2020 |
Machine learning › Deep learning architectures and training › normalization
batch normalization |
0.3 | 1 | 2018 | Scalable methods for 8-bit training of neural networks · NeurIPS 2018 |
Machine learning › Trustworthy machine learning › calibration
classifier calibration |
0.3 | 1 | 2018 | Fix your classifier: the marginal value of training the last weight layer · ICLR (Poster) 2018 |
Machine learning › Transfer learning and domain adaptation
fine-tuning |
0.3 | 1 | 2018 | Fix your classifier: the marginal value of training the last weight layer · ICLR (Poster) 2018 |
Machine learning › Transfer learning and domain adaptation › fine-tuning
last-layer retraining |
0.3 | 1 | 2018 | Fix your classifier: the marginal value of training the last weight layer · ICLR (Poster) 2018 |
Machine learning › Learning theory › generalization error
generalization gap |
0.3 | 1 | 2017 | Train longer, generalize better: closing the generalization gap in large batch training of neural networks · NIPS 2017 |
Machine learning › Optimization for machine learning
learning rate schedule |
0.3 | 1 | 2017 | Train longer, generalize better: closing the generalization gap in large batch training of neural networks · NIPS 2017 |
Machine learning › Efficient and distributed learning › model compression › quantization › low-precision computation
low-precision neural network |
0.3 | 1 | 2017 | Quantized Neural Networks: Training Neural Networks with Low Precision Weights and Activations · J. Mach. Learn. Res. 2017 |
Machine learning › Efficient and distributed learning › model compression › quantization
quantized training |
0.3 | 1 | 2017 | Quantized Neural Networks: Training Neural Networks with Low Precision Weights and Activations · J. Mach. Learn. Res. 2017 |
Machine learning › Efficient and distributed learning › model compression › quantization › quantized neural network
binary neural network |
0.2 | 1 | 2016 | Binarized Neural Networks · NIPS 2016 |
Machine learning › Efficient and distributed learning › model compression
lightweight neural network |
0.2 | 1 | 2016 | Binarized Neural Networks · NIPS 2016 |
Machine learning › Efficient and distributed learning › model compression › quantization
quantized neural network |
0.2 | 1 | 2016 | Binarized Neural Networks · NIPS 2016 |
Hardware accelerators and domain-specific architectures › quantization
binarized neural network |
0.2 | 1 | 2016 | Binarized Neural Networks · NIPS 2016 |
Methods — techniques the papers use, named apart from their topics
quantization · 2.3gradient approximation · 1.5minimum variance unbiased estimation · 0.7n:m sparsity · 0.5minimum cost flow · 0.5integer programming · 0.5calibration set optimization · 0.5synthetic data generation · 0.4batch normalization statistics · 0.4batch augmentation · 0.4gradient-based training · 0.2bitwise operations · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Towards Cheaper Inference in Deep Networks with Lower Bit-Width AccumulatorsabstractThe majority of the research on the quantization of Deep Neural Networks (DNNs) is focused on reducing the precision of tensors visible by high-level frameworks (e.g., weights, activations, and gradients). However, current hardware still relies on high-accuracy core operations. Most significant is the operation of accumulating products. This high-precision accumulation operation is gradually becoming the main computational bottleneck. This is because, so far, the usage of low-precision accumulators led to a significant degradation in performance. In this work, we present a simple method to train and fine-tune DNNs, to allow, for the first time, utilization of cheaper, $12$-bits accumulators, with no significant degradation in accuracy. Lastly, we show that as we decrease the accumulation precision further, using fine-grained gradient approximations can improve the DNN accuracy. Yaniv Blumenfeld, Itay Hubara, Daniel Soudry |
ICLR | 2 |
| 2023 | Minimum Variance Unbiased N: M Sparsity for the Neural Gradients
Brian Chmiel, Itay Hubara, Ron Banner, Daniel Soudry |
ICLR | 2 |
| 2021 | Accurate Post Training Quantization With Small Calibration SetsabstractLately, post-training quantization methods have gained considerable attention, as they are simple to use, and require only a small unlabeled calibration set. This small dataset cannot be used to fine-tune the model without significant over-fitting. Instead, these methods only use the calibration set to set the activations’ dynamic ranges. However, such methods always resulted in significant accuracy degradation, when used below 8-bits (except on small datasets). Here we aim to break the 8-bit barrier. To this end, we minimize the quantization errors of each layer or block separately by optimizing its parameters over the calibration set. We empirically demonstrate that this approach is: (1) much less susceptible to over-fitting than the standard fine-tuning approaches, and can be used even on a very small calibration set; and (2) more powerful than previous methods, which only set the activations’ dynamic ranges. We suggest two flavors for our method, parallel and sequential aim for a fixed and flexible bit-width allocation. For the latter, we demonstrate how to optimally allocate the bit-widths for each layer, while constraining accuracy degradation or model compression by proposing a novel integer programming formulation. Finally, we suggest model global statistics tuning, to correct biases introduced during quantization. Together, these methods yield state-of-the-art results for both vision and text models. For instance, on ResNet50, we obtain less than 1% accuracy degradation — with 4-bit weights and activations in all layers, but first and last. The suggested methods are two orders of magnitude faster than the traditional Quantize Aware Training approach used for lower than 8-bit quantization. We open-sourced our code \textit{https://github.com/papers-submission/CalibTIP}. Itay Hubara, Yury Nahshan, Yair Hanani, Ron Banner, Daniel Soudry |
ICML | 1 |
| 2021 | Accelerated Sparse Neural Training: A Provable and Efficient Method to Find N: M Transposable MasksabstractUnstructured pruning reduces the memory footprint in deep neural networks (DNNs). Recently, researchers proposed different types of structural pruning intending to reduce also the computation complexity. In this work, we first suggest a new measure called mask-diversity which correlates with the expected accuracy of the different types of structural pruning. We focus on the recently suggested N:M fine-grained block sparsity mask, in which for each block of M weights, we have at least N zeros. While N:M fine-grained block sparsity allows acceleration in actual modern hardware, it can be used only to accelerate the inference phase. In order to allow for similar accelerations in the training phase, we suggest a novel transposable fine-grained sparsity mask, where the same mask can be used for both forward and backward passes. Our transposable mask guarantees that both the weight matrix and its transpose follow the same sparsity pattern; thus, the matrix multiplication required for passing the error backward can also be accelerated. We formulate the problem of finding the optimal transposable-mask as a minimum-cost flow problem. Additionally, to speed up the minimum-cost flow computation, we also introduce a fast linear-time approximation that can be used when the masks dynamically change during training. Our experiments suggest a 2x speed-up in the matrix multiplications with no accuracy degradation over vision and language models. Finally, to solve the problem of switching between different structure constraints, we suggest a method to convert a pre-trained model with unstructured sparsity to an N:M fine-grained block sparsity model with little to no training. A reference implementation can be found at https://github.com/papers-submission/structuredtransposablemasks. Itay Hubara, Brian Chmiel, Moshe Island, Ron Banner, Joseph Naor, Daniel Soudry |
NeurIPS | 1 |
| 2020 | The Knowledge Within: Methods for Data-Free Model CompressionabstractBackground: Recently, an extensive amount of research has been focused on compressing and accelerating Deep Neural Networks (DNN). So far, high compression rate algorithms require part of the training dataset for a low precision calibration, or a fine-tuning process. However, this requirement is unacceptable when the data is unavailable or contains sensitive information, as in medical and biometric use-cases. Contributions: We present three methods for generating synthetic samples from trained models. Then, we demonstrate how these samples can be used to calibrate and fine-tune quantized models without using any real data in the process. Our best performing method has a negligible accuracy degradation compared to the original training set. This method, which leverages intrinsic batch normalization layers' statistics of the trained model, can be used to evaluate data similarity. Our approach opens a path towards genuine data-free model compression, alleviating the need for training data during model deployment. Matan Haroush, Itay Hubara, Elad Hoffer, Daniel Soudry |
CVPR | 2 |
| 2020 | Augment Your Batch: Improving Generalization Through Instance RepetitionabstractLarge-batch SGD is important for scaling training of deep neural networks. However, without fine-tuning hyperparameter schedules, the generalization of the model may be hampered. We propose to use batch augmentation: replicating instances of samples within the same batch with different data augmentations. Batch augmentation acts as a regularizer and an accelerator, increasing both generalization and performance scaling for a fixed budget of optimization steps. We analyze the effect of batch augmentation on gradient variance and show that it empirically improves convergence for a wide variety of networks and datasets. Our results show that batch augmentation reduces the number of necessary SGD updates to achieve the same accuracy as the state-of-the-art. Overall, this simple yet effective method enables faster training and better generalization by allowing more computational resources to be used concurrently. Elad Hoffer, Tal Ben-Nun, Itay Hubara, Niv Giladi, Torsten Hoefler, Daniel Soudry |
CVPR | 3 |
| 2020 | MLPerf Inference BenchmarkabstractMachine-learning (ML) hardware and software system demand is burgeoning. Driven by ML applications, the number of different ML inference systems has exploded. Over 100 organizations are building ML inference chips, and the systems that incorporate existing models span at least three orders of magnitude in power consumption and five orders of magnitude in performance; they range from embedded devices to data-center solutions. Fueling the hardware are a dozen or more software frameworks and libraries. The myriad combinations of ML hardware and ML software make assessing ML-system performance in an architecture-neutral, representative, and reproducible manner challenging. There is a clear need for industry-wide standard ML benchmarking and evaluation criteria. MLPerf Inference answers that call. In this paper, we present our benchmarking method for evaluating ML inference systems. Driven by more than 30 organizations as well as more than 200 ML engineers and practitioners, MLPerf prescribes a set of rules and best practices to ensure comparability across systems with wildly differing architectures. The first call for submissions garnered more than 600 reproducible inference-performance measurements from 14 organizations, representing over 30 systems that showcase a wide range of capabilities. The submissions attest to the benchmark’s flexibility and adaptability. Vijay Janapa Reddi, David Kanter, Peter Mattson, Guenther Schmuelling, Carole-Jean Wu, Maximilien Breughe, Mark Charlebois, William Chou, Ramesh Chukka, Cody Coleman, Sam Davis, Gregory Frederick Diamos, Jared Duke, David Fick, J. Scott Gardner, Itay Hubara, Sachin Idgunji, Thomas B. Jablin, Jeff Jiao, Tom St. John, Pankaj Kanwar, Jeffery Liao, Anton Lokhmotov, Francisco Massa, Peng Meng, Paulius Micikevicius, Colin Osborne, Gennady Pekhimenko, Arun Tejusve Raghunath Rajan, Dilip Sequeira, Ashish Sirasao, Fei Sun 0002, Michael Thomson, Frank Wei, Ephrem Wu, Lingjie Xu, Koichi Yamada, George Yuan, Aaron Zhong, Peizhao Zhang |
ISCA | 19 |
| 2018 | Fix your classifier: the marginal value of training the last weight layer
Elad Hoffer, Itay Hubara, Daniel Soudry |
ICLR (Poster) | 2 |
| 2018 | Scalable methods for 8-bit training of neural networksabstractQuantized Neural Networks (QNNs) are often used to improve network efficiency during the inference phase, i.e. after the network has been trained. Extensive research in the field suggests many different quantization schemes. Still, the number of bits required, as well as the best quantization scheme, are yet unknown. Our theoretical analysis suggests that most of the training process is robust to substantial precision reduction, and points to only a few specific operations that require higher precision. Armed with this knowledge, we quantize the model parameters, activations and layer gradients to 8-bit, leaving at higher precision only the final step in the computation of the weight gradients. Additionally, as QNNs require batch-normalization to be trained at high precision, we introduce Range Batch-Normalization (BN) which has significantly higher tolerance to quantization noise and improved computational complexity. Our simulations show that Range BN is equivalent to the traditional batch norm if a precise scale adjustment, which can be approximated analytically, is applied. To the best of the authors' knowledge, this work is the first to quantize the weights, activations, as well as a substantial volume of the gradients stream, in all layers (including batch normalization) to 8-bit while showing state-of-the-art results over the ImageNet-1K dataset. Ron Banner, Itay Hubara, Elad Hoffer, Daniel Soudry |
NeurIPS | 2 |
| 2017 | Train longer, generalize better: closing the generalization gap in large batch training of neural networksabstractBackground: Deep learning models are typically trained using stochastic gradient descent or one of its variants. These methods update the weights using their gradient, estimated from a small fraction of the training data. It has been observed that when using large batch sizes there is a persistent degradation in generalization performance - known as the "generalization gap" phenomenon. Identifying the origin of this gap and closing it had remained an open problem. Contributions: We examine the initial high learning rate training phase. We find that the weight distance from its initialization grows logarithmically with the number of weight updates. We therefore propose a "random walk on a random landscape" statistical model which is known to exhibit similar "ultra-slow" diffusion behavior. Following this hypothesis we conducted experiments to show empirically that the "generalization gap" stems from the relatively small number of updates rather than the batch size, and can be completely eliminated by adapting the training regime used. We further investigate different techniques to train models in the large-batch regime and present a novel algorithm named "Ghost Batch Normalization" which enables significant decrease in the generalization gap without increasing the number of updates. To validate our findings we conduct several additional experiments on MNIST, CIFAR-10, CIFAR-100 and ImageNet. Finally, we reassess common practices and beliefs concerning training of deep models and suggest they may not be optimal to achieve good generalization. Elad Hoffer, Itay Hubara, Daniel Soudry |
NIPS | 2 |
| 2017 | Quantized Neural Networks: Training Neural Networks with Low Precision Weights and Activations
Itay Hubara, Matthieu Courbariaux, Daniel Soudry, Ran El-Yaniv, Yoshua Bengio |
J. Mach. Learn. Res. | 1 |
| 2016 | Binarized Neural NetworksabstractWe introduce a method to train Binarized Neural Networks (BNNs) - neural networks with binary weights and activations at run-time. At train-time the binary weights and activations are used for computing the parameter gradients. During the forward pass, BNNs drastically reduce memory size and accesses, and replace most arithmetic operations with bit-wise operations, which is expected to substantially improve power-efficiency. To validate the effectiveness of BNNs, we conducted two sets of experiments on the Torch7 and Theano frameworks. On both, BNNs achieved nearly state-of-the-art results over the MNIST, CIFAR-10 and SVHN datasets. We also report our preliminary results on the challenging ImageNet dataset. Last but not least, we wrote a binary matrix multiplication GPU kernel with which it is possible to run our MNIST BNN 7 times faster than with an unoptimized GPU kernel, without suffering any loss in classification accuracy. The code for training and running our BNNs is available on-line. Itay Hubara, Matthieu Courbariaux, Daniel Soudry, Ran El-Yaniv, Yoshua Bengio |
NIPS | 1 |
| 2014 | Expectation Backpropagation: Parameter-Free Training of Multilayer Neural Networks with Continuous or Discrete Weights
Daniel Soudry, Itay Hubara, Ron Meir |
NIPS | 2 |