VLDB 2026 Research / reviewers in the wild / expert
Ron Banner
dblp:03/5857
· DBLP profile ↗
26ranked-venue papers
10as first author
11since 2021 · last 2025
0000-0002-7333-7830ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 2 first-author · 11 since 2021Computer networks · 8 · 8 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Scaling FP8 training to trillion-token LLMsabstractWe train, for the first time, large language models using FP8 precision on datasets up to 2 trillion tokens --- a 20-fold increase over previous limits. Through these extended training runs, we uncover critical instabilities in FP8 training that were not observable in earlier works with shorter durations. We trace these instabilities to outlier amplification by the SwiGLU activation function. Interestingly, we show, both analytically and empirically, that this amplification happens only over prolonged training periods, and link it to a SwiGLU weight alignment process. To address this newly identified issue, we introduce Smooth-SwiGLU, a novel modification that ensures stable FP8 training without altering function behavior. We also demonstrate, for the first time, FP8 quantization of both Adam optimizer moments. Combining these innovations, we successfully train a 7B parameter model using FP8 precision on 256 Intel Gaudi2 accelerators, achieving on-par results with the BF16 baseline while delivering up to a $\sim$ 34 % throughput improvement. A reference implementation is supplied in https://github.com/Anonymous1252022/Megatron-DeepSpeed Maxim Fishman, Brian Chmiel, Ron Banner, Daniel Soudry |
ICLR | 3 |
| 2025 | FP4 All the Way: Fully Quantized Training of Large Language ModelsabstractWe demonstrate, for the first time, fully quantized training (FQT) of large language models (LLMs) using predominantly 4-bit floating-point (FP4) precision for weights, activations, and gradients on datasets up to 200 billion tokens. We extensively investigate key design choices for FP4, including block sizes, scaling formats, and rounding methods. Our analysis shows that the NVFP4 format, where each block of 16 FP4 values (E2M1) shares a scale represented in E4M3, provides optimal results. We use stochastic rounding for backward and update passes and round-to-nearest for the forward pass to enhance stability. Additionally, we identify a theoretical and empirical threshold for effective quantized training: when the gradient norm falls below approximately $\sqrt{3}$ times the quantization noise, quantized training becomes less effective. Leveraging these insights, we successfully train a 7-billion-parameter model on 256 Intel Gaudi2 accelerators. The resulting FP4-trained model achieves downstream task performance comparable to a standard BF16 baseline, confirming that FP4 training is a practical and highly efficient approach for large-scale LLM training. A reference implementation is supplied in https://github.com/Anonymous1252022/fp4-all-the-way. Brian Chmiel, Maxim Fishman, Ron Banner, Daniel Soudry |
NeurIPS | 3 |
| 2023 | Accurate Neural Training with 4-bit Matrix Multiplications at Standard Formats
Brian Chmiel, Ron Banner, Elad Hoffer, Hilla Ben-Yaacov, Daniel Soudry |
ICLR | 2 |
| 2023 | Minimum Variance Unbiased N: M Sparsity for the Neural Gradients
Brian Chmiel, Itay Hubara, Ron Banner, Daniel Soudry |
ICLR | 3 |
| 2023 | DropCompute: simple and more robust distributed synchronous training via compute variance reductionabstractBackground: Distributed training is essential for large scale training of deep neural networks (DNNs). The dominant methods for large scale DNN training are synchronous (e.g. All-Reduce), but these require waiting for all workers in each step. Thus, these methods are limited by the delays caused by straggling workers.
Results: We study a typical scenario in which workers are straggling due to variability in compute time. We find an analytical relation between compute time properties and scalability limitations, caused by such straggling workers. With these findings, we propose a simple yet effective decentralized method to reduce the variation among workers and thus improve the robustness of synchronous training. This method can be integrated with the widely used All-Reduce. Our findings are validated on large-scale training tasks using 200 Gaudi Accelerators. Niv Giladi, Shahar Gottlieb, Moran Shkolnik, Asaf Karnieli, Ron Banner, Elad Hoffer, Kfir Y. Levy, Daniel Soudry |
NeurIPS | 5 |
| 2021 | Neural gradients are near-lognormal: improved quantized and sparse training
Brian Chmiel, Liad Ben-Uri, Moran Shkolnik, Elad Hoffer, Ron Banner, Daniel Soudry |
ICLR | 5 |
| 2021 | GAN "Steerability" without optimization
Nurit Spingarn-Eliezer, Ron Banner, Tomer Michaeli |
ICLR | 2 |
| 2021 | Accurate Post Training Quantization With Small Calibration SetsabstractLately, post-training quantization methods have gained considerable attention, as they are simple to use, and require only a small unlabeled calibration set. This small dataset cannot be used to fine-tune the model without significant over-fitting. Instead, these methods only use the calibration set to set the activations’ dynamic ranges. However, such methods always resulted in significant accuracy degradation, when used below 8-bits (except on small datasets). Here we aim to break the 8-bit barrier. To this end, we minimize the quantization errors of each layer or block separately by optimizing its parameters over the calibration set. We empirically demonstrate that this approach is: (1) much less susceptible to over-fitting than the standard fine-tuning approaches, and can be used even on a very small calibration set; and (2) more powerful than previous methods, which only set the activations’ dynamic ranges. We suggest two flavors for our method, parallel and sequential aim for a fixed and flexible bit-width allocation. For the latter, we demonstrate how to optimally allocate the bit-widths for each layer, while constraining accuracy degradation or model compression by proposing a novel integer programming formulation. Finally, we suggest model global statistics tuning, to correct biases introduced during quantization. Together, these methods yield state-of-the-art results for both vision and text models. For instance, on ResNet50, we obtain less than 1% accuracy degradation — with 4-bit weights and activations in all layers, but first and last. The suggested methods are two orders of magnitude faster than the traditional Quantize Aware Training approach used for lower than 8-bit quantization. We open-sourced our code \textit{https://github.com/papers-submission/CalibTIP}. Itay Hubara, Yury Nahshan, Yair Hanani, Ron Banner, Daniel Soudry |
ICML | 4 |
| 2021 | Accelerated Sparse Neural Training: A Provable and Efficient Method to Find N: M Transposable MasksabstractUnstructured pruning reduces the memory footprint in deep neural networks (DNNs). Recently, researchers proposed different types of structural pruning intending to reduce also the computation complexity. In this work, we first suggest a new measure called mask-diversity which correlates with the expected accuracy of the different types of structural pruning. We focus on the recently suggested N:M fine-grained block sparsity mask, in which for each block of M weights, we have at least N zeros. While N:M fine-grained block sparsity allows acceleration in actual modern hardware, it can be used only to accelerate the inference phase. In order to allow for similar accelerations in the training phase, we suggest a novel transposable fine-grained sparsity mask, where the same mask can be used for both forward and backward passes. Our transposable mask guarantees that both the weight matrix and its transpose follow the same sparsity pattern; thus, the matrix multiplication required for passing the error backward can also be accelerated. We formulate the problem of finding the optimal transposable-mask as a minimum-cost flow problem. Additionally, to speed up the minimum-cost flow computation, we also introduce a fast linear-time approximation that can be used when the masks dynamically change during training. Our experiments suggest a 2x speed-up in the matrix multiplications with no accuracy degradation over vision and language models. Finally, to solve the problem of switching between different structure constraints, we suggest a method to convert a pre-trained model with unstructured sparsity to an N:M fine-grained block sparsity model with little to no training. A reference implementation can be found at https://github.com/papers-submission/structuredtransposablemasks. Itay Hubara, Brian Chmiel, Moshe Island, Ron Banner, Joseph Naor, Daniel Soudry |
NeurIPS | 4 |
| 2021 | CAT: Compression-Aware Training for bandwidth reductionabstractOne major obstacle hindering the ubiquitous use of CNNs for inference is their relatively high memory bandwidth requirements, which can be the primary energy consumer and throughput bottleneck in hardware accelerators. Inspired by quantization-aware training approaches, we propose a compression-aware training (CAT) method that involves training the model to allow better compression of weights and feature maps during neural network deployment. Our method trains the model to achieve low-entropy feature maps, enabling efficient compression at inference time using classical transform coding methods. CAT significantly improves the state-of-the-art results reported for quantization evaluated on various vision and NLP tasks, such as image classification (ImageNet), image detection (Pascal VOC), sentiment analysis (CoLa), and textual entailment (MNLI). For example, on ResNet-18, we achieve near baseline ImageNet accuracy with an average representation of only 1.5 bits per value with 5-bit quantization. Moreover, we show that entropy reduction of weights and activations can be applied together, further improving bandwidth reduction. Reference implementation is available. Chaim Baskin, Brian Chmiel, Evgenii Zheltonozhskii, Ron Banner, Alexander M. Bronstein, Avi Mendelson |
J. Mach. Learn. Res. | 4 |
| 2021 | Loss aware post-training quantizationabstractNeural network quantization enables the deployment of large models on resource-constrained devices. Current post-training quantization methods fall short in terms of accuracy for INT4 (or lower) but provide reasonable accuracy for INT8 (or above). In this work, we study the effect of quantization on the structure of the loss landscape. We show that the structure is flat and separable for mild quantization, enabling straightforward post-training quantization methods to achieve good results. We show that with more aggressive quantization, the loss landscape becomes highly non-separable with steep curvature, making the selection of quantization parameters more challenging. Armed with this understanding, we design a method that quantizes the layer parameters jointly, enabling significant accuracy improvement over current post-training quantization methods. Reference implementation is available at https://github.com/ynahshan/nn-quantization-pytorch/tree/master/lapq . Yury Nahshan, Brian Chmiel, Chaim Baskin, Evgenii Zheltonozhskii, Ron Banner, Alexander M. Bronstein, Avi Mendelson |
Mach. Learn. | 5 |
| 2020 | Thanks for Nothing: Predicting Zero-Valued Activations with Lightweight Convolutional Neural Networks
Gil Shomron, Ron Banner, Moran Shkolnik, Uri C. Weiser |
ECCV (10) | 2 |
| 2020 | Feature Map Transform Coding for Energy-Efficient CNN InferenceabstractConvolutional neural networks (CNNs) achieve state-of-the-art accuracy in a variety of tasks in computer vision and beyond. One of the major obstacles hindering the ubiquitous use of CNNs for inference on low-power edge devices is their high computational complexity and memory bandwidth requirements. The latter often dominates the energy footprint on modern hardware. In this paper, we introduce a lossy transform coding approach, inspired by image and video compression, designed to reduce the memory bandwidth due to the storage of intermediate activation calculation results. Our method does not require fine-tuning the network weights and halves the data transfer volumes to the main memory by compressing feature maps, which are highly correlated, with variable length coding. Our method outperform previous approach in term of the number of bits per value with minor accuracy degradation on ResNet-34 and MobileNetV2. We analyze the performance of our approach on a variety of CNN architectures and demonstrate that FPGA implementation of ResNet-18 with our approach results in a reduction of around 40% in the memory energy footprint, compared to quantized network, with negligible impact on accuracy. When allowing accuracy degradation of up to 2%, the reduction of 60% is achieved. A reference implementation accompanies the paper. Brian Chmiel, Chaim Baskin, Evgenii Zheltonozhskii, Ron Banner, Yevgeny Yermolin, Alex Karbachevsky, Alexander M. Bronstein, Avi Mendelson |
IJCNN | 4 |
| 2020 | Robust Quantization: One Model to Rule Them AllabstractNeural network quantization methods often involve simulating the quantization process during training, making the trained model highly dependent on the target bit-width and precise way quantization is performed. Robust quantization offers an alternative approach with improved tolerance to different classes of data-types and quantization policies. It opens up new exciting applications where the quantization process is not static and can vary to meet different circumstances and implementations. To address this issue, we propose a method that provides intrinsic robustness to the model against a broad range of quantization processes. Our method is motivated by theoretical arguments and enables us to store a single generic model capable of operating at various bit-widths and quantization policies. We validate our method's effectiveness on different ImageNet Models. A reference implementation accompanies the paper. Moran Shkolnik, Brian Chmiel, Ron Banner, Gil Shomron, Yury Nahshan, Alexander M. Bronstein, Uri C. Weiser |
NeurIPS | 3 |
| 2019 | Post training 4-bit quantization of convolutional networks for rapid-deploymentabstractConvolutional neural networks require significant memory bandwidth and storage for intermediate computations, apart from substantial computing resources. Neural network quantization has significant benefits in reducing the amount of intermediate results, but it often requires the full datasets and time-consuming fine tuning to recover the accuracy lost after quantization. This paper introduces the first practical 4-bit post training quantization approach: it does not involve training the quantized model (fine-tuning), nor it requires the availability of the full dataset. We target the quantization of both activations and weights and suggest three complementary methods for minimizing quantization error at the tensor level, two of whom obtain a closed-form analytical solution. Combining these methods, our approach achieves accuracy that is just a few percents less the state-of-the-art baseline across a wide range of convolutional models. The source code to replicate all experiments is available on GitHub: \url{https://github.com/submission2019/cnn-quantization}. Ron Banner, Yury Nahshan, Daniel Soudry |
NeurIPS | 1 |
| 2018 | Scalable methods for 8-bit training of neural networksabstractQuantized Neural Networks (QNNs) are often used to improve network efficiency during the inference phase, i.e. after the network has been trained. Extensive research in the field suggests many different quantization schemes. Still, the number of bits required, as well as the best quantization scheme, are yet unknown. Our theoretical analysis suggests that most of the training process is robust to substantial precision reduction, and points to only a few specific operations that require higher precision. Armed with this knowledge, we quantize the model parameters, activations and layer gradients to 8-bit, leaving at higher precision only the final step in the computation of the weight gradients. Additionally, as QNNs require batch-normalization to be trained at high precision, we introduce Range Batch-Normalization (BN) which has significantly higher tolerance to quantization noise and improved computational complexity. Our simulations show that Range BN is equivalent to the traditional batch norm if a precise scale adjustment, which can be approximated analytically, is applied. To the best of the authors' knowledge, this work is the first to quantize the weights, activations, as well as a substantial volume of the gradients stream, in all layers (including batch normalization) to 8-bit while showing state-of-the-art results over the ImageNet-1K dataset. Ron Banner, Itay Hubara, Elad Hoffer, Daniel Soudry |
NeurIPS | 1 |
| 2018 | Norm matters: efficient and accurate normalization schemes in deep networksabstractOver the past few years, Batch-Normalization has been commonly used in deep networks, allowing faster training and high performance for a wide variety of applications. However, the reasons behind its merits remained unanswered, with several shortcomings that hindered its use for certain tasks. In this work, we present a novel view on the purpose and function of normalization methods and weight-decay, as tools to decouple weights' norm from the underlying optimized objective. This property highlights the connection between practices such as normalization, weight decay and learning-rate adjustments. We suggest several alternatives to the widely used $L^2$ batch-norm, using normalization in $L^1$ and $L^\infty$ spaces that can substantially improve numerical stability in low-precision implementations as well as provide computational and memory benefits. We demonstrate that such methods enable the first batch-norm alternative to work for half-precision implementations. Finally, we suggest a modification to weight-normalization, which improves its performance on large-scale tasks. Elad Hoffer, Ron Banner, Itay Golan, Daniel Soudry |
NeurIPS | 2 |
| 2013 | Image deblurring using maps of highlightsabstractIn deblurring an image, we seek to recover the original sharp image. However, without knowledge of the blurring process, we cannot expect to recover the image perfectly. We propose a deblurring method of a single-image where the blur kernel is directly estimated from highlight spots or streaks with high intensity value. These highlighted points can be represented by specular reflection of light that may appear in the eye, on a shiny surface, or in small light sources present in the image. In this work, we detect automatically these high-lighted points in a blurred image. Therefore, creating a map of highlight, which is used as a guide to extract automatically a single highlight from the blurred image. Due to its unique nature, it is demosntrated that the highlighted points are a good estimation of the blur kernel and a sharp image is restored using these kernel. The experimental results show the performance of this method in comparison to several other deblurring methods. Fabiane Queiroz, Ing Ren Tsang, Lior Shapira, Ron Banner |
ICASSP | 4 |
| 2010 | Designing Low-Capacity Backup Networks for Fast RestorationabstractThere are two basic approaches to allocate protection resources for fast restoration. The first allocates resources upon the arrival of each connection request; yet, it incurs significant set-up time and is often capacity-inefficient. The second approach allocates protection resources during the network configuration phase; therefore, it needs to accommodate any possible arrival pattern of connection requests, hence potentially calling for a substantial over-provisioning of resources. However, in this study we establish the feasibility of this approach. Specifically, we consider a scheme that, during the network configuration phase, constructs an (additional) low-capacity backup network. Upon a failure, traffic is rerouted through a bypass in the backup network. We establish that, with proper design, backup networks induce feasible capacity overhead. We further impose several design requirements (e.g., hop-count limits) on backup networks and their induced bypasses, and prove that, commonly, they also incur minor overhead. Motivated by these findings, we design efficient algorithms for the construction of backup networks. Ron Banner, Ariel Orda |
INFOCOM | 1 |
| 2008 | Multi-Objective Topology Control in Wireless NetworksabstractTopology control is the task of establishing an efficient underlying graph for ad-hoc networks over which high level routing protocols are implemented. The following design goals are of fundamental importance for wireless topologies: (1) low level of interference; (2) minimum energy consumption; (3) high spatial reuse; (4) connectivity; (5) planarity; (6) sparseness; (7) symmetry; (8) small nodal degree; (9) communication-efficient and localized construction. Previous topology control algorithms have usually been designed to provide good performance guarantees in the worst case. Yet, since design goals often conflict, it is usually impossible to construct a single structure that concurrently addresses a large number of goals efficiently. On the other hand, in this paper we show that a substantially larger number of design goals can concurrently be addressed when "pathological" worst case scenarios are ignored. Accordingly, we focus on average performance and establish a protocol that satisfies all the above design goals. Specifically, we formally prove the efficiency of our protocol with respect to all design goals except for high spatial reuse and small nodal degree, for which this is demonstrated by way of simulations. We note that minimum energy consumption and low level of interference have been the main targets of topology control, and our protocol is proven to offer salient performance guarantees with respect to both. Ron Banner, Ariel Orda |
INFOCOM | 1 |
| 2007 | Bottleneck Routing Games in Communication NetworksabstractWe consider routing games where the performance of each user is dictated by the worst (bottleneck) element it employs. We are given a network, finitely many (selfish) users, each associated with a positive flow demand, and a load-dependent performance function for each network element; the social (i.e., system) objective is to optimize the performance of the worst element in the network (i.e., the network bottleneck). Although we show that such "bottleneck" routing games appear in a variety of practical scenarios, they have not been considered yet. Accordingly, we study their properties, considering two routing scenarios, namely when a user can split its traffic over more than one path (splittable bottleneck game) and when it cannot (unsplittable bottleneck game). First, we prove that, for both splittable and unsplittable bottleneck games, there is a (not necessarily unique) Nash equilibrium. Then, we consider the rate of convergence to a Nash equilibrium in each game. Finally, we investigate the efficiency of the Nash equilibria in both games with respect to the social optimum; specifically, while for both games we show that the price of anarchy is unbounded, we identify for each game conditions under which Nash equilibria are socially optimal. Ron Banner, Ariel Orda |
IEEE J. Sel. Areas Commun. | 1 |
| 2007 | Multipath routing algorithms for congestion minimization
Ron Banner, Ariel Orda |
IEEE/ACM Trans. Netw. | 1 |
| 2007 | The power of tuning: a novel approach for the efficient design of survivable networks
Ron Banner, Ariel Orda |
IEEE/ACM Trans. Netw. | 1 |
| 2006 | Bottleneck Routing Games in Communication NetworksabstractWe consider routing games where the performance of each user is dictated by the worst (bottleneck) element it employs. We are given a network, finitely many (selfish) users, each associated with a positive flow demand, and a load- dependent performance function for each network element; the social (i.e., system) objective is to optimize the performance of the worst element in the network (i.e., the network bottleneck). Although we show that such routing games appear in a variety of practical scenarios, they have not been considered yet. Accordingly, we study their properties, considering two routing scenarios, namely when a user can split its traffic over more than one path (splittable bottleneck game) and when it cannot (unsplittable bottleneck game). First, we prove that, for both splittable and unsplittable bottleneck games, there is a (not necessarily unique) Nash equilibrium. Then, we consider the rate of convergence to a Nash equilibrium in each game. Finally, we investigate the efficiency of the Nash equilibria in both games with respect to the social optimum; specifically, while for both games we show that the price of anarchy is unbounded, we identify for each game conditions under which Nash equilibria are socially optimal. Ron Banner, Ariel Orda |
INFOCOM | 1 |
| 2005 | Multipath Routing Algorithms for Congestion Minimization
Ron Banner, Ariel Orda |
NETWORKING | 1 |
| 2004 | The Power of Tuning: A Novel Approach for the Efficient Design of Survivable NetworksabstractCurrent survivability schemes typically offer two degrees of protection, namely full protection (from a single failure) or no protection at all. Full protection translates into rigid design constraints, i.e. the employment of disjoint paths. We introduce the concept of tunable survivability that bridges the gap between full and no protection. First, we establish several fundamental properties of connections with tunable survivability. With that at hand, we devise efficient polynomial (optimal) connection establishment schemes for both 1:1 and 1+1 protection architectures. Then, we show that the concept of tunable survivability gives rise to a novel hybrid protection architecture, which offers improved performance over the standard 1:1 and 1+1 architectures. Next, we investigate some related QoS extensions. Finally, we demonstrate the advantage of tunable survivability over full survivability. In particular, we show that, by just slightly alleviating the requirement of full survivability, we obtain major improvements in terms of the "feasibility" as well as the "quality" of the solution. Ron Banner, Ariel Orda |
ICNP | 1 |