Dar Gilboa

dblp:203/4469 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
6since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 2 first-author · 5 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Efficient and distributed learning · 25% Representation and self-supervised learning · 20% Learning theory · 18%
Theoretical computer science
4 papers
Quantum computing and quantum information · 67% Information theory · 29% Algorithms and data structures · 4%

Topics — the 26 heaviest of 28, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Quantum computing and quantum information
quantum machine learning
1.422024
Exponential Quantum Communication Advantage in Distributed Inference and Learning · NeurIPS 2024
On quantum backpropagation, information reuse, and cheating measurement collapse · NeurIPS 2023
Machine learning › Efficient and distributed learning › distributed training
communication-efficient training
0.812024
Exponential Quantum Communication Advantage in Distributed Inference and Learning · NeurIPS 2024
Machine learning › Efficient and distributed learning
distributed training
0.812024
Exponential Quantum Communication Advantage in Distributed Inference and Learning · NeurIPS 2024
Quantum computing and quantum information › quantum state tomography
shadow tomography
0.712023
On quantum backpropagation, information reuse, and cheating measurement collapse · NeurIPS 2023
Machine learning › Learning theory
generalization bounds
0.622021
Deep Networks Provably Classify Data on Curves · NeurIPS 2021
Efficient Dictionary Learning with Gradient Descent · ICML 2019
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
manifold learning
0.512021
Deep Networks and the Multiple Manifold Problem · ICLR 2021
Machine learning › Learning theory
neural network theory
0.512021
Deep Networks and the Multiple Manifold Problem · ICLR 2021
Machine learning › Learning theory › neural network theory › neural network kernels
neural tangent kernel
0.512021
Deep Networks Provably Classify Data on Curves · NeurIPS 2021
Machine learning › Representation and self-supervised learning › representation learning
representation geometry
0.512021
Deep Networks and the Multiple Manifold Problem · ICLR 2021
Information theory › information measures › information decomposition
partial information decomposition
0.512021
Estimating the Unique Information of Continuous Variables · NeurIPS 2021
Information theory › information measures › information decomposition
synergy and redundancy
0.512021
Estimating the Unique Information of Continuous Variables · NeurIPS 2021
Machine learning › Representation and self-supervised learning
feature diversity
0.412020
Beyond Signal Propagation: Is Feature Diversity Necessary in Deep Neural Network Initialization? · ICML 2020
Machine learning › Deep learning architectures and training
weight initialization
0.412020
Beyond Signal Propagation: Is Feature Diversity Necessary in Deep Neural Network Initialization? · ICML 2020
Machine learning › Optimization for machine learning
convergence guarantees
0.412019
Efficient Dictionary Learning with Gradient Descent · ICML 2019
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › sparse coding
dictionary learning
0.412019
Efficient Dictionary Learning with Gradient Descent · ICML 2019
Machine learning › Optimization for machine learning › gradient-based optimization
gradient descent
0.412019
Efficient Dictionary Learning with Gradient Descent · ICML 2019
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference
mean-field approximation
0.412019
A Mean Field Theory of Quantized Deep Networks: The Quantization-Depth Trade-Off · NeurIPS 2019
Machine learning › Efficient and distributed learning
model compression
0.412019
A Mean Field Theory of Quantized Deep Networks: The Quantization-Depth Trade-Off · NeurIPS 2019
Machine learning › Optimization for machine learning
non-convex optimization
0.412019
Efficient Dictionary Learning with Gradient Descent · ICML 2019
Machine learning › Efficient and distributed learning › model compression › quantization
quantized neural network
0.412019
A Mean Field Theory of Quantized Deep Networks: The Quantization-Depth Trade-Off · NeurIPS 2019
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference
approximate inference
0.312017
Stochastic Bouncy Particle Sampler · ICML 2017
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
markov chain monte carlo
0.312017
Stochastic Bouncy Particle Sampler · ICML 2017
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods › markov chain monte carlo
non-reversible markov chain
0.312017
Stochastic Bouncy Particle Sampler · ICML 2017
Quantum computing and quantum information
quantum communication
0.212024
Exponential Quantum Communication Advantage in Distributed Inference and Learning · NeurIPS 2024
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference
0.112021
Estimating the Unique Information of Continuous Variables · NeurIPS 2021
Algorithms and data structures › numerical linear algebra › dimensionality reduction › nonlinear dimensionality reduction
manifold learning
0.112021
Deep Networks Provably Classify Data on Curves · NeurIPS 2021

Methods — techniques the papers use, named apart from their topics

gradient descent · 2.9quantum state encoding · 1.5variational autoencoder optimization · 1.0neural tangent kernel · 1.0copula decomposition · 1.0shadow tomography · 0.7parameterized quantum circuit · 0.7symmetry breaking · 0.4random initialization · 0.4negative curvature analysis · 0.4mean-field theory · 0.4rejection-free sampling · 0.3
YearPublicationVenuePosition
2025 Consumable Data via Quantum Communication
abstract
Classical data can be copied and re-used for computation, with adverse consequences economically and in terms of data privacy. Motivated by this, we formulate problems in one-way communication complexity where Alice holds some data x and Bob holds m inputs y_1, …, y_m. They want to compute m instances of a bipartite relation R(⋅,⋅) on every pair (x, y_1), …, (x, y_m). We call this the asymmetric direct sum question for one-way communication. We give examples where the quantum communication complexity of such problems scales polynomially with m, while the classical communication complexity depends at most logarithmically on m. Thus, for such problems, data behaves like a consumable resource that is effectively destroyed upon use when the owner stores and transmits it as quantum states, but not when transmitted classically. We show an application to a strategic data-selling game, and discuss other potential economic implications.
Dar Gilboa, Siddhartha Jain 0002, Jarrod R. McClean
APPROX/RANDOM1
2024 Exponential Quantum Communication Advantage in Distributed Inference and Learning
abstract
Training and inference with large machine learning models that far exceed the memory capacity of individual devices necessitates the design of distributed architectures, forcing one to contend with communication constraints. We present a framework for distributed computation over a quantum network in which data is encoded into specialized quantum states. We prove that for models within this framework, inference and training using gradient descent can be performed with exponentially less communication compared to their classical analogs, and with relatively modest overhead relative to standard gradient-based methods. We show that certain graph neural networks are particularly amenable to implementation within this framework, and moreover present empirical evidence that they perform well on standard benchmarks. To our knowledge, this is the first example of exponential quantum advantage for a generic class of machine learning problems that hold regardless of the data encoding cost. Moreover, we show that models in this class can encode highly nonlinear features of their inputs, and their expressivity increases exponentially with model depth. We also delineate the space of models for which exponential communication advantages hold by showing that they cannot hold for linear classification. Communication of quantum states that potentially limit the amount of information that can be extracted from them about the data and model parameters may also lead to improved privacy guarantees for distributed computation. Taken as a whole, these findings form a promising foundation for distributed machine learning over quantum networks.
Dar Gilboa, Hagay Michaeli, Daniel Soudry, Jarrod R. McClean
NeurIPS1
2023 On quantum backpropagation, information reuse, and cheating measurement collapse
abstract
The success of modern deep learning hinges on the ability to train neural networks at scale. Through clever reuse of intermediate information, backpropagation facilitates training through gradient computation at a total cost roughly proportional to running the function, rather than incurring an additional factor proportional to the number of parameters -- which can now be in the trillions. Naively, one expects that quantum measurement collapse entirely rules out the reuse of quantum information as in backpropagation. But recent developments in shadow tomography, which assumes access to multiple copies of a quantum state, have challenged that notion. Here, we investigate whether parameterized quantum models can train as efficiently as classical neural networks. We show that achieving backpropagation scaling is impossible without access to multiple copies of a state. With this added ability, we introduce an algorithm with foundations in shadow tomography that matches backpropagation scaling in quantum resources while reducing classical auxiliary computational costs to open problems in shadow tomography. These results highlight the nuance of reusing quantum information for practical purposes and clarify the unique difficulties in training large quantum models, which could alter the course of quantum machine learning.
Amira Abbas, Robbie King, Hsin-Yuan Huang, William J. Huggins, Ramis Movassagh, Dar Gilboa, Jarrod R. McClean
NeurIPS6
2021 Deep Networks and the Multiple Manifold Problem
Sam Buchanan, Dar Gilboa, John Wright 0001
ICLR2
2021 Estimating the Unique Information of Continuous Variables
abstract
The integration and transfer of information from multiple sources to multiple targets is a core motive of neural systems. The emerging field of partial information decomposition (PID) provides a novel information-theoretic lens into these mechanisms by identifying synergistic, redundant, and unique contributions to the mutual information between one and several variables. While many works have studied aspects of PID for Gaussian and discrete distributions, the case of general continuous distributions is still uncharted territory. In this work we present a method for estimating the unique information in continuous distributions, for the case of one versus two variables. Our method solves the associated optimization problem over the space of distributions with fixed bivariate marginals by combining copula decompositions and techniques developed to optimize variational autoencoders. We obtain excellent agreement with known analytic results for Gaussians, and illustrate the power of our new approach in several brain-inspired neural models. Our method is capable of recovering the effective connectivity of a chaotic network of rate neurons, and uncovers a complex trade-off between redundancy, synergy and unique information in recurrent networks trained to solve a generalized XOR~task.
Ari Pakman, Amin Nejatbakhsh, Dar Gilboa, Abdullah Makkeh, Luca Mazzucato, Michael Wibral, Elad Schneidman
NeurIPS3
2021 Deep Networks Provably Classify Data on Curves
abstract
Data with low-dimensional nonlinear structure are ubiquitous in engineering and scientific problems. We study a model problem with such structure---a binary classification task that uses a deep fully-connected neural network to classify data drawn from two disjoint smooth curves on the unit sphere. Aside from mild regularity conditions, we place no restrictions on the configuration of the curves. We prove that when (i) the network depth is large relative to certain geometric properties that set the difficulty of the problem and (ii) the network width and number of samples is polynomial in the depth, randomly-initialized gradient descent quickly learns to correctly classify all points on the two curves with high probability. To our knowledge, this is the first generalization guarantee for deep networks with nonlinear data that depends only on intrinsic data properties. Our analysis proceeds by a reduction to dynamics in the neural tangent kernel (NTK) regime, where the network depth plays the role of a fitting resource in solving the classification problem. In particular, via fine-grained control of the decay properties of the NTK, we demonstrate that when the network is sufficiently deep, the NTK can be locally approximated by a translationally invariant operator on the manifolds and stably inverted over smooth functions, which guarantees convergence and generalization.
Tingran Wang, Sam Buchanan, Dar Gilboa, John Wright 0001
NeurIPS3
2020 Beyond Signal Propagation: Is Feature Diversity Necessary in Deep Neural Network Initialization?
abstract
Deep neural networks are typically initialized with random weights, with variances chosen to facilitate signal propagation and stable gradients. It is also believed that diversity of features is an important property of these initializations. We construct a deep convolutional network with identical features by initializing almost all the weights to $0$. The architecture also enables perfect signal propagation and stable gradients, and achieves high accuracy on standard benchmarks. This indicates that random, diverse initializations are \emph{not} necessary for training neural networks. An essential element in training this network is a mechanism of symmetry breaking; we study this phenomenon and find that standard GPU operations, which are non-deterministic, can serve as a sufficient source of symmetry breaking to enable training.
Yaniv Blumenfeld, Dar Gilboa, Daniel Soudry
ICML2
2019 Efficient Dictionary Learning with Gradient Descent
abstract
Randomly initialized first-order optimization algorithms are the method of choice for solving many high-dimensional nonconvex problems in machine learning, yet general theoretical guarantees cannot rule out convergence to critical points of poor objective value. For some highly structured nonconvex problems however, the success of gradient descent can be understood by studying the geometry of the objective. We study one such problem – complete orthogonal dictionary learning, and provide converge guarantees for randomly initialized gradient descent to the neighborhood of a global optimum. The resulting rates scale as low order polynomials in the dimension even though the objective possesses an exponential number of saddle points. This efficient convergence can be viewed as a consequence of negative curvature normal to the stable manifolds associated with saddle points, and we provide evidence that this feature is shared by other nonconvex problems of importance as well.
Dar Gilboa, Sam Buchanan, John Wright 0001
ICML1
2019 A Mean Field Theory of Quantized Deep Networks: The Quantization-Depth Trade-Off
abstract
Reducing the precision of weights and activation functions in neural network training, with minimal impact on performance, is essential for the deployment of these models in resource-constrained environments. We apply mean field techniques to networks with quantized activations in order to evaluate the degree to which quantization degrades signal propagation at initialization. We derive initialization schemes which maximize signal propagation in such networks, and suggest why this is helpful for generalization. Building on these results, we obtain a closed form implicit equation for $L_{\max}$, the maximal trainable depth (and hence model capacity), given $N$, the number of quantization levels in the activation function. Solving this equation numerically, we obtain asymptotically: $L_{\max}\propto N^{1.82}$.
Yaniv Blumenfeld, Dar Gilboa, Daniel Soudry
NeurIPS2
2017 Stochastic Bouncy Particle Sampler
abstract
We introduce a stochastic version of the non-reversible, rejection-free Bouncy Particle Sampler (BPS), a Markov process whose sample trajectories are piecewise linear, to efficiently sample Bayesian posteriors in big datasets. We prove that in the BPS no bias is introduced by noisy evaluations of the log-likelihood gradient. On the other hand, we argue that efficiency considerations favor a small, controllable bias, in exchange for faster mixing. We introduce a simple method that controls this trade-off. We illustrate these ideas in several examples which outperform previous approaches.
Ari Pakman, Dar Gilboa, David E. Carlson, Liam Paninski
ICML2