Bobak T. Kiani

dblp:232/4086 · also Bobak Toussi Kiani · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
9since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 4 first-author · 9 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
10 papers
Learning theory · 41% Deep learning architectures and training · 26% Graph learning · 19%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational science and engineering · 100%

Topics — the 29 heaviest of 31, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Graph learning
graph neural network
2.232024
Unitary Convolutions for Learning on Graphs and Groups · NeurIPS 2024
On the hardness of learning under symmetries · ICLR 2024
Equivariant Polynomials for Graph Neural Networks · ICML 2023
Machine learning › Learning theory
statistical query learning
1.522024
Hardness of Learning Neural Networks under the Manifold Hypothesis · NeurIPS 2024
On the hardness of learning under symmetries · ICLR 2024
Machine learning › Deep learning architectures and training
equivariant neural network
1.322024
On the hardness of learning under symmetries · ICLR 2024
Implicit Bias of Linear Equivariant Networks · ICML 2022
Machine learning › Learning theory
generalization
0.822023
The SSL Interplay: Augmentations, Inductive Bias, and Generalization · ICML 2023
Implicit Bias of Linear Equivariant Networks · ICML 2022
Machine learning › Learning theory
computational learning theory
0.812024
Hardness of Learning Neural Networks under the Manifold Hypothesis · NeurIPS 2024
Machine learning › Learning theory › statistical query learning
correlational statistical query
0.812024
On the hardness of learning under symmetries · ICLR 2024
Machine learning › Deep learning architectures and training › equivariant neural network
group convolutional network
0.812024
Unitary Convolutions for Learning on Graphs and Groups · NeurIPS 2024
Machine learning › Learning theory › inductive bias
manifold hypothesis
0.812024
Hardness of Learning Neural Networks under the Manifold Hypothesis · NeurIPS 2024
Machine learning › Learning theory › neural network theory
neural network learnability
0.812024
Hardness of Learning Neural Networks under the Manifold Hypothesis · NeurIPS 2024
Machine learning › Graph learning › graph neural network › deep graph neural network
over-smoothing
0.812024
Unitary Convolutions for Learning on Graphs and Groups · NeurIPS 2024
Machine learning › Deep learning architectures and training
data augmentation
0.712023
The SSL Interplay: Augmentations, Inductive Bias, and Generalization · ICML 2023
Machine learning › Graph learning › graph neural network
expressive power
0.712023
Equivariant Polynomials for Graph Neural Networks · ICML 2023
Machine learning › Representation and self-supervised learning
self-supervised learning theory
0.712023
The SSL Interplay: Augmentations, Inductive Bias, and Generalization · ICML 2023
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
self-supervised representation learning
0.712023
Self-Supervised Learning with Lie Symmetries for Partial Differential Equations · NeurIPS 2023
Computational science and engineering › scientific machine learning
neural network solver
0.712023
Self-Supervised Learning with Lie Symmetries for Partial Differential Equations · NeurIPS 2023
Computational science and engineering
partial differential equations
0.712023
Self-Supervised Learning with Lie Symmetries for Partial Differential Equations · NeurIPS 2023
Computational science and engineering
scientific machine learning
0.712023
Self-Supervised Learning with Lie Symmetries for Partial Differential Equations · NeurIPS 2023
Machine learning › Learning theory
implicit bias
0.612022
Implicit Bias of Linear Equivariant Networks · ICML 2022
Machine learning › Deep learning architectures and training
recurrent neural network
0.612022
projUNN: efficient method for training deep networks with unitary matrices · NeurIPS 2022
Machine learning › Deep learning architectures and training › recurrent neural network
unitary recurrent neural network
0.612022
projUNN: efficient method for training deep networks with unitary matrices · NeurIPS 2022
Machine learning › Trustworthy machine learning › robustness › certified robustness
certified adversarial robustness
0.512021
Adversarial Robustness Guarantees for Random Deep Neural Networks · ICML 2021
Machine learning › Learning theory
neural network theory
0.512021
Adversarial Robustness Guarantees for Random Deep Neural Networks · ICML 2021
Machine learning › Deep learning architectures and training
random neural network
0.512021
Adversarial Robustness Guarantees for Random Deep Neural Networks · ICML 2021
Machine learning › Trustworthy machine learning
robustness
0.512021
Adversarial Robustness Guarantees for Random Deep Neural Networks · ICML 2021
Machine learning › Learning theory
generalization bounds
0.412019
Random deep neural networks are biased towards simple functions · NeurIPS 2019
Machine learning › Learning theory › generalization bounds
PAC-Bayes bounds
0.412019
Random deep neural networks are biased towards simple functions · NeurIPS 2019
Machine learning › Learning theory › inductive bias
simplicity bias
0.412019
Random deep neural networks are biased towards simple functions · NeurIPS 2019
Machine learning › Learning theory › neural network theory
kernel regime
0.212023
The SSL Interplay: Augmentations, Inductive Bias, and Generalization · ICML 2023
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process
0.112021
Adversarial Robustness Guarantees for Random Deep Neural Networks · ICML 2021

Methods — techniques the papers use, named apart from their topics

gradient descent · 1.3joint embedding · 1.3unitary matrices · 0.8statistical query model · 0.8group convolution · 0.8cryptographic hardness · 0.8weisfeiler-lehman hierarchy · 0.7tensor contraction · 0.7self-supervised learning · 0.7kernel methods · 0.7generalization bounds · 0.7
YearPublicationVenuePosition
2024 On the hardness of learning under symmetries
abstract
We study the problem of learning equivariant neural networks via gradient descent. The incorporation of known symmetries ("equivariance") into neural nets has empirically improved the performance of learning pipelines, in domains ranging from biology to computer vision. However, a rich yet separate line of learning theoretic research has demonstrated that actually learning shallow, fully-connected (i.e. non-symmetric) networks has exponential complexity in the correlational statistical query (CSQ) model, a framework encompassing gradient descent. In this work, we ask: are known problem symmetries sufficient to alleviate the fundamental hardness of learning neural nets with gradient descent? We answer this question in the negative. In particular, we give lower bounds for shallow graph neural networks, convolutional networks, invariant polynomials, and frame-averaged networks for permutation subgroups, which all scale either superpolynomially or exponentially in the relevant input dimension. Therefore, in spite of the significant inductive bias imparted via symmetry, actually learning the complete classes of functions represented by equivariant neural networks via gradient descent remains hard.
Bobak T. Kiani, Thien Le, Hannah Lawrence, Stefanie Jegelka, Melanie Weber 0001
ICLR1
2024 Unitary Convolutions for Learning on Graphs and Groups
abstract
Data with geometric structure is ubiquitous in machine learning often arising from fundamental symmetries in a domain, such as permutation-invariance in graphs and translation-invariance in images. Group-convolutional architectures, which encode symmetries as inductive bias, have shown great success in applications, but can suffer from instabilities as their depth increases and often struggle to learn long range dependencies in data. For instance, graph neural networks experience instability due to the convergence of node representations (over-smoothing), which can occur after only a few iterations of message-passing, reducing their effectiveness in downstream tasks. Here, we propose and study unitary group convolutions, which allow for deeper networks that are more stable during training. The main focus of the paper are graph neural networks, where we show that unitary graph convolutions provably avoid over-smoothing. Our experimental results confirm that unitary graph convolutional networks achieve competitive performance on benchmark datasets compared to state-of-the-art graph neural networks. We complement our analysis of the graph domain with the study of general unitary convolutions and analyze their role in enhancing stability in general group convolutional architectures.
Bobak T. Kiani, Lukas Fesser, Melanie Weber 0001
NeurIPS1
2024 Hardness of Learning Neural Networks under the Manifold Hypothesis
abstract
The manifold hypothesis presumes that high-dimensional data lies on or near a low-dimensional manifold. While the utility of encoding geometric structure has been demonstrated empirically, rigorous analysis of its impact on the learnability of neural networks is largely missing. Several recent results have established hardness results for learning feedforward and equivariant neural networks under i.i.d. Gaussian or uniform Boolean data distributions. In this paper, we investigate the hardness of learning under the manifold hypothesis. We ask, which minimal assumptions on the curvature and regularity of the manifold, if any, render the learning problem efficiently learnable. We prove that learning is hard under input manifolds of bounded curvature by extending proofs of hardness in the SQ and cryptographic settings for boolean data inputs to the geometric setting. On the other hand, we show that additional assumptions on the volume of the data manifold alleviate these fundamental limitations and guarantee learnability via a simple interpolation argument. Notable instances of this regime are manifolds which can be reliably reconstructed via manifold learning. Looking forward, we comment on and empirically explore intermediate regimes of manifolds, which have heterogeneous features commonly found in real world data.
Bobak T. Kiani, Melanie Weber 0001
NeurIPS1
2023 The SSL Interplay: Augmentations, Inductive Bias, and Generalization
abstract
Self-supervised learning (SSL) has emerged as a powerful framework to learn representations from raw data without supervision. Yet in practice, engineers face issues such as instability in tuning optimizers and collapse of representations during training. Such challenges motivate the need for a theory to shed light on the complex interplay between the choice of data augmentation, network architecture, and training algorithm. % on the resulting performance in downstream tasks. We study such an interplay with a precise analysis of generalization performance on both pretraining and downstream tasks in kernel regimes, and highlight several insights for SSL practitioners that arise from our theory.
Vivien Cabannes, Bobak T. Kiani, Randall Balestriero, Yann LeCun, Alberto Bietti
ICML2
2023 Equivariant Polynomials for Graph Neural Networks
abstract
Graph Neural Networks (GNN) are inherently limited in their expressive power. Recent seminal works (Xu et al., 2019; Morris et al., 2019b) introduced the Weisfeiler-Lehman (WL) hierarchy as a measure of expressive power. Although this hierarchy has propelled significant advances in GNN analysis and architecture developments, it suffers from several significant limitations. These include a complex definition that lacks direct guidance for model improvement and a WL hierarchy that is too coarse to study current GNNs. This paper introduces an alternative expressive power hierarchy based on the ability of GNNs to calculate equivariant polynomials of a certain degree. As a first step, we provide a full characterization of all equivariant graph polynomials by introducing a concrete basis, significantly generalizing previous results. Each basis element corresponds to a specific multi-graph, and its computation over some graph data input corresponds to a tensor contraction problem. Second, we propose algorithmic tools for evaluating the expressiveness of GNNs using tensor contraction sequences, and calculate the expressive power of popular GNNs. Finally, we enhance the expressivity of common GNN architectures by adding polynomial features or additional operations / aggregations inspired by our theory. These enhanced GNNs demonstrate state-of-the-art results in experiments across multiple graph learning benchmarks.
Omri Puny, Derek Lim, Bobak T. Kiani, Haggai Maron, Yaron Lipman
ICML3
2023 Self-Supervised Learning with Lie Symmetries for Partial Differential Equations
abstract
Machine learning for differential equations paves the way for computationally efficient alternatives to numerical solvers, with potentially broad impacts in science and engineering. Though current algorithms typically require simulated training data tailored to a given setting, one may instead wish to learn useful information from heterogeneous sources, or from real dynamical systems observations that are messy or incomplete. In this work, we learn general-purpose representations of PDEs from heterogeneous data by implementing joint embedding methods for self-supervised learning (SSL), a framework for unsupervised representation learning that has had notable success in computer vision. Our representation outperforms baseline approaches to invariant tasks, such as regressing the coefficients of a PDE, while also improving the time-stepping performance of neural solvers. We hope that our proposed methodology will prove useful in the eventual development of general-purpose foundation models for PDEs.
Grégoire Mialon, Quentin Garrido, Hannah Lawrence, Danyal Rehman, Yann LeCun, Bobak T. Kiani
NeurIPS6
2022 Implicit Bias of Linear Equivariant Networks
abstract
Group equivariant convolutional neural networks (G-CNNs) are generalizations of convolutional neural networks (CNNs) which excel in a wide range of technical applications by explicitly encoding symmetries, such as rotations and permutations, in their architectures. Although the success of G-CNNs is driven by their explicit symmetry bias, a recent line of work has proposed that the implicit bias of training algorithms on particular architectures is key to understanding generalization for overparameterized neural nets. In this context, we show that L-layer full-width linear G-CNNs trained via gradient descent for binary classification converge to solutions with low-rank Fourier matrix coefficients, regularized by the 2/L-Schatten matrix norm. Our work strictly generalizes previous analysis on the implicit bias of linear CNNs to linear G-CNNs over all finite groups, including the challenging setting of non-commutative groups (such as permutations), as well as band-limited G-CNNs over infinite groups. We validate our theorems via experiments on a variety of groups, and empirically explore more realistic nonlinear networks, which locally capture similar regularization patterns. Finally, we provide intuitive interpretations of our Fourier space implicit regularization results in real space via uncertainty principles.
Hannah Lawrence, Bobak T. Kiani, Kristian Georgiev, Andrew Dienes
ICML2
2022 projUNN: efficient method for training deep networks with unitary matrices
abstract
In learning with recurrent or very deep feed-forward networks, employing unitary matrices in each layer can be very effective at maintaining long-range stability. However, restricting network parameters to be unitary typically comes at the cost of expensive parameterizations or increased training runtime. We propose instead an efficient method based on rank-$k$ updates -- or their rank-$k$ approximation -- that maintains performance at a nearly optimal training runtime. We introduce two variants of this method, named Direct (projUNN-D) and Tangent (projUNN-T) projected Unitary Neural Networks, that can parameterize full $N$-dimensional unitary or orthogonal matrices with a training runtime scaling as $O(kN^2)$. Our method either projects low-rank gradients onto the closest unitary matrix (projUNN-T) or transports unitary matrices in the direction of the low-rank gradient (projUNN-D). Even in the fastest setting ($k=1$), projUNN is able to train a model's unitary parameters to reach comparable performances against baseline implementations. In recurrent neural network settings, projUNN closely matches or exceeds benchmarked results from prior unitary neural networks. Finally, we preliminarily explore projUNN in training orthogonal convolutional neural networks, which are currently unable to outperform state of the art models but can potentially enhance stability and robustness at large depth.
Bobak T. Kiani, Randall Balestriero, Yann LeCun, Seth Lloyd
NeurIPS1
2021 Adversarial Robustness Guarantees for Random Deep Neural Networks
abstract
The reliability of deep learning algorithms is fundamentally challenged by the existence of adversarial examples, which are incorrectly classified inputs that are extremely close to a correctly classified input. We explore the properties of adversarial examples for deep neural networks with random weights and biases, and prove that for any p$\geq$1, the \ell^p distance of any given input from the classification boundary scales as one over the square root of the dimension of the input times the \ell^p norm of the input. The results are based on the recently proved equivalence between Gaussian processes and deep neural networks in the limit of infinite width of the hidden layers, and are validated with experiments on both random deep neural networks and deep neural networks trained on the MNIST and CIFAR10 datasets. The results constitute a fundamental advance in the theoretical understanding of adversarial examples, and open the way to a thorough theoretical characterization of the relation between network architecture and robustness to adversarial perturbations.
Giacomo De Palma, Bobak T. Kiani, Seth Lloyd
ICML2
2019 Random deep neural networks are biased towards simple functions
abstract
We prove that the binary classifiers of bit strings generated by random wide deep neural networks with ReLU activation function are biased towards simple functions. The simplicity is captured by the following two properties. For any given input bit string, the average Hamming distance of the closest input bit string with a different classification is at least sqrt(n / (2π log n)), where n is the length of the string. Moreover, if the bits of the initial string are flipped randomly, the average number of flips required to change the classification grows linearly with n. These results are confirmed by numerical experiments on deep neural networks with two hidden layers, and settle the conjecture stating that random deep neural networks are biased towards simple functions. This conjecture was proposed and numerically explored in [Valle Pérez et al., ICLR 2019] to explain the unreasonably good generalization properties of deep learning algorithms. The probability distribution of the functions generated by random deep neural networks is a good choice for the prior probability distribution in the PAC-Bayesian generalization bounds. Our results constitute a fundamental step forward in the characterization of this distribution, therefore contributing to the understanding of the generalization properties of deep learning algorithms.
Giacomo De Palma, Bobak T. Kiani, Seth Lloyd
NeurIPS2