Akhilan Boopathy

dblp:230/8358 · DBLP profile ↗
← Back
10ranked-venue papers
8as first author
7since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 8 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Trustworthy machine learning · 42% Learning theory · 28% Deep learning architectures and training · 17%
Software engineering, system software, and programming languages
1 paper
Program verification · 100%

Topics — the 19 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
robustness
1.742021
Fast Training of Provably Robust Neural Networks by SingleProp · AAAI 2021
Proper Network Interpretability Helps Adversarial Robustness in Classification · ICML 2020
PROVEN: Verifying Robustness of Neural Networks with a Probabilistic Approach · ICML 2019
Machine learning › Learning theory
inductive bias
1.422024
Towards Exact Computation of Inductive Bias · IJCAI 2024
Model-agnostic Measure of Generalization Difficulty · ICML 2023
Machine learning › Trustworthy machine learning › robustness
certified robustness
1.332021
Fast Training of Provably Robust Neural Networks by SingleProp · AAAI 2021
PROVEN: Verifying Robustness of Neural Networks with a Probabilistic Approach · ICML 2019
CNN-Cert: An Efficient Framework for Certifying Robustness of Convolutional Neural Networks · AAAI 2019
Natural language and speech › Language models and text generation
compositional generalization
0.912025
Breaking Neural Network Scaling Laws with Modularity · ICLR 2025
Machine learning › Deep learning architectures and training
modular neural network
0.912025
Breaking Neural Network Scaling Laws with Modularity · ICLR 2025
Machine learning › Learning theory
sample complexity
0.912025
Breaking Neural Network Scaling Laws with Modularity · ICLR 2025
Machine learning › Learning paradigms
continual learning
0.812024
Rapid Learning without Catastrophic Forgetting in the Morris Water Maze · ICML 2024
Machine learning › Learning theory
generalization
0.712023
Model-agnostic Measure of Generalization Difficulty · ICML 2023
Machine learning › Deep learning architectures and training › neural network training
backpropagation-free training
0.612022
How to Train Your Wide Neural Network Without Backprop: An Input-Weight Alignment Perspective · ICML 2022
Machine learning › Deep learning architectures and training
biologically plausible learning
0.612022
How to Train Your Wide Neural Network Without Backprop: An Input-Weight Alignment Perspective · ICML 2022
Machine learning › Learning theory › neural network theory › neural network kernels
neural tangent kernel
0.612022
How to Train Your Wide Neural Network Without Backprop: An Input-Weight Alignment Perspective · ICML 2022
Machine learning › Trustworthy machine learning › robustness
adversarial robustness
0.412020
Proper Network Interpretability Helps Adversarial Robustness in Classification · ICML 2020
Machine learning › Trustworthy machine learning
interpretability
0.412020
Proper Network Interpretability Helps Adversarial Robustness in Classification · ICML 2020
Machine learning › Trustworthy machine learning › interpretability
neural network interpretation
0.412020
Proper Network Interpretability Helps Adversarial Robustness in Classification · ICML 2020
Machine learning › Trustworthy machine learning › verification
probabilistic verification
0.412019
PROVEN: Verifying Robustness of Neural Networks with a Probabilistic Approach · ICML 2019
Machine learning › Trustworthy machine learning
uncertainty estimation
0.412019
PROVEN: Verifying Robustness of Neural Networks with a Probabilistic Approach · ICML 2019
Machine learning › Trustworthy machine learning › robustness › adversarial robustness
adversarial training
0.112021
Fast Training of Provably Robust Neural Networks by SingleProp · AAAI 2021
Machine learning › Deep learning architectures and training
convolutional neural network
0.112019
CNN-Cert: An Efficient Framework for Certifying Robustness of Convolutional Neural Networks · AAAI 2019
Program verification
neural network verification
0.112019
PROVEN: Verifying Robustness of Neural Networks with a Probabilistic Approach · ICML 2019

Methods — techniques the papers use, named apart from their topics

modularity · 0.9learning rule · 0.9remapping · 0.8grid cells · 0.8convolutional network · 0.8content-addressable heteroassociative memory · 0.8intrinsic dimensionality · 0.7hypothesis volume measurement · 0.7input-weight alignment · 0.6gradient descent · 0.6fast-lin · 0.4CROWN · 0.4CNN-Cert · 0.4
YearPublicationVenuePosition
2025 Breaking Neural Network Scaling Laws with Modularity
abstract
Modular neural networks outperform nonmodular neural networks on tasks ranging from visual question answering to robotics. These performance improvements are thought to be due to modular networks' superior ability to model the compositional and combinatorial structure of real-world problems. However, a theoretical explanation of how modularity improves generalizability, and how to leverage task modularity while training networks remains elusive. Using recent theoretical progress in explaining neural network generalization, we investigate how the amount of training data required to generalize on a task varies with the intrinsic dimensionality of a task's input. We show theoretically that when applied to modularly structured tasks, while nonmodular networks require an exponential number of samples with task dimensionality, modular networks' sample complexity is independent of task dimensionality: modular networks can generalize in high dimensions. We then develop a novel learning rule for modular networks to exploit this advantage and empirically show the improved generalization of the rule, both in- and out-of-distribution, on high-dimensional, modular tasks.
Akhilan Boopathy, Sunshine Jiang, William Yue, Jaedong Hwang, Abhiram Iyer, Ila Fiete
ICLR1
2024 Rapid Learning without Catastrophic Forgetting in the Morris Water Maze
abstract
Animals can swiftly adapt to novel tasks, while maintaining proficiency on previously trained tasks. This contrasts starkly with machine learning models, which struggle on these capabilities. We first propose a new task, the sequential Morris Water Maze (sWM), which extends a widely used task in the psychology and neuroscience fields and requires both rapid and continual learning. It has frequently been hypothesized that inductive biases from brains could help build better ML systems, but the addition of constraints typically hurts rather than helping ML performance. We draw inspiration from biology to show that combining 1) a content-addressable heteroassociative memory based on the entorhinal-hippocampal circuit with grid cells that retain shared across-environment structural representations and hippocampal cells that acquire environment-specific information; 2) a spatially invariant convolutional network architecture for rapid adaptation across unfamiliar environments; and 3) the ability to perform remapping, which orthogonalizes internal representations; leads to good generalization, rapid learning, and continual learning without forgetting, respectively. Our model outperforms ANN baselines from continual learning contexts applied to the task. It retains knowledge of past environments while rapidly acquiring the skills to navigate new ones, thereby addressing the seemingly opposing challenges of quick knowledge transfer and sustaining proficiency in previously learned tasks. These biologically motivated results may point the way toward ML algorithms with similar properties.
Raymond Wang, Jaedong Hwang, Akhilan Boopathy, Ila Fiete
ICML3
2024 Resampling-free Particle Filters in High-dimensions
abstract
State estimation is crucial for the performance and safety of numerous robotic applications. Among the suite of estimation techniques, particle filters have been identified as a powerful solution due to their non-parametric nature. Yet, in high-dimensional state spaces, these filters face challenges such as ’particle deprivation’ which hinders accurate representation of the true posterior distribution. This paper introduces a novel resampling-free particle filter designed to mitigate particle deprivation by forgoing the traditional resampling step. This ensures a broader and more diverse particle set, especially vital in high-dimensional scenarios. Theoretically, our proposed filter is shown to offer a near-accurate representation of the desired posterior distribution in high-dimensional contexts. Empirically, the effectiveness of our approach is underscored through a high-dimensional synthetic state estimation task and a 6D pose estimation derived from videos. We posit that as robotic systems evolve with greater degrees of freedom, particle filters tailored for high-dimensional state spaces will be indispensable.
Akhilan Boopathy, Aneesh Muppidi, Peggy Yang, Abhiram Iyer, William Yue, Ila Fiete
ICRA1
2024 Towards Exact Computation of Inductive Bias
Akhilan Boopathy, William Yue, Jaedong Hwang, Abhiram Iyer, Ila Fiete
IJCAI1
2023 Model-agnostic Measure of Generalization Difficulty
abstract
The measure of a machine learning algorithm is the difficulty of the tasks it can perform, and sufficiently difficult tasks are critical drivers of strong machine learning models. However, quantifying the generalization difficulty of machine learning benchmarks has remained challenging. We propose what is to our knowledge the first model-agnostic measure of the inherent generalization difficulty of tasks. Our inductive bias complexity measure quantifies the total information required to generalize well on a task minus the information provided by the data. It does so by measuring the fractional volume occupied by hypotheses that generalize on a task given that they fit the training data. It scales exponentially with the intrinsic dimensionality of the space over which the model must generalize but only polynomially in resolution per dimension, showing that tasks which require generalizing over many dimensions are drastically more difficult than tasks involving more detail in fewer dimensions. Our measure can be applied to compute and compare supervised learning, reinforcement learning and meta-learning generalization difficulties against each other. We show that applied empirically, it formally quantifies intuitively expected trends, e.g. that in terms of required inductive bias, MNIST $<$ CIFAR10 $<$ Imagenet and fully observable Markov decision processes (MDPs) $<$ partially observable MDPs. Further, we show that classification of complex images $<$ few-shot meta-learning with simple images. Our measure provides a quantitative metric to guide the construction of more complex tasks requiring greater inductive bias, and thereby encourages the development of more sophisticated architectures and learning algorithms with more powerful generalization capabilities.
Akhilan Boopathy, Kevin Liu, Jaedong Hwang, Shu Ge, Asaad Mohammedsaleh, Ila Fiete
ICML1
2022 How to Train Your Wide Neural Network Without Backprop: An Input-Weight Alignment Perspective
abstract
Recent works have examined theoretical and empirical properties of wide neural networks trained in the Neural Tangent Kernel (NTK) regime. Given that biological neural networks are much wider than their artificial counterparts, we consider NTK regime wide neural networks as a possible model of biological neural networks. Leveraging NTK theory, we show theoretically that gradient descent drives layerwise weight updates that are aligned with their input activity correlations weighted by error, and demonstrate empirically that the result also holds in finite-width wide networks. The alignment result allows us to formulate a family of biologically-motivated, backpropagation-free learning rules that are theoretically equivalent to backpropagation in infinite-width networks. We test these learning rules on benchmark problems in feedforward and recurrent neural networks and demonstrate, in wide networks, comparable performance to backpropagation. The proposed rules are particularly effective in low data regimes, which are common in biological learning settings.
Akhilan Boopathy, Ila Fiete
ICML1
2021 Fast Training of Provably Robust Neural Networks by SingleProp
abstract
Recent works have developed several methods of defending neural networks against adversarial attacks with certified guarantees. However, these techniques can be computationally costly due to the use of certification during training. We develop a new regularizer that is both more efficient than existing certified defenses, requiring only one additional forward propagation through a network, and can be used to train networks with similar certified accuracy. Through experiments on MNIST and CIFAR-10 we demonstrate improvements in training speed and comparable certified accuracy compared to state-of-the-art certified defenses.
Akhilan Boopathy, Lily Weng, Sijia Liu 0001, Gaoyuan Zhang, Luca Daniel
AAAI1
2020 Proper Network Interpretability Helps Adversarial Robustness in Classification
abstract
Recent works have empirically shown that there exist adversarial examples that can be hidden from neural network interpretability (namely, making network interpretation maps visually similar), or interpretability is itself susceptible to adversarial attacks. In this paper, we theoretically show that with a proper measurement of interpretation, it is actually difficult to prevent prediction-evasion adversarial attacks from causing interpretation discrepancy, as confirmed by experiments on MNIST, CIFAR-10 and Restricted ImageNet. Spurred by that, we develop an interpretability-aware defensive scheme built only on promoting robust interpretation (without the need for resorting to adversarial loss minimization). We show that our defense achieves both robust classification and robust interpretation, outperforming state-of-the-art adversarial training methods against attacks of large perturbation in particular.
Akhilan Boopathy, Sijia Liu 0001, Gaoyuan Zhang, Cynthia Liu, Shiyu Chang, Luca Daniel
ICML1
2019 CNN-Cert: An Efficient Framework for Certifying Robustness of Convolutional Neural Networks
abstract
Verifying robustness of neural network classifiers has attracted great interests and attention due to the success of deep neural networks and their unexpected vulnerability to adversarial perturbations. Although finding minimum adversarial distortion of neural networks (with ReLU activations) has been shown to be an NP-complete problem, obtaining a non-trivial lower bound of minimum distortion as a provable robustness guarantee is possible. However, most previous works only focused on simple fully-connected layers (multilayer perceptrons) and were limited to ReLU activations. This motivates us to propose a general and efficient framework, CNN-Cert, that is capable of certifying robustness on general convolutional neural networks. Our framework is general – we can handle various architectures including convolutional layers, max-pooling layers, batch normalization layer, residual blocks, as well as general activation functions; our approach is efficient – by exploiting the special structure of convolutional layers, we achieve up to 17 and 11 times of speed-up compared to the state-of-the-art certification algorithms (e.g. Fast-Lin, CROWN) and 366 times of speed-up compared to the dual-LP approach while our algorithm obtains similar or even better verification bounds. In addition, CNN-Cert generalizes state-of-the-art algorithms e.g. Fast-Lin and CROWN. We demonstrate by extensive experiments that our method outperforms state-of-the-art lowerbound-based certification algorithms in terms of both bound quality and speed.
Akhilan Boopathy, Tsui-Wei Weng, Sijia Liu 0001, Luca Daniel
AAAI1
2019 PROVEN: Verifying Robustness of Neural Networks with a Probabilistic Approach
abstract
We propose a novel framework PROVEN to \textbf{PRO}babilistically \textbf{VE}rify \textbf{N}eural network’s robustness with statistical guarantees. PROVEN provides probability certificates of neural network robustness when the input perturbation follow distributional characterization. Notably, PROVEN is derived from current state-of-the-art worst-case neural network robustness verification frameworks, and therefore it can provide probability certificates with little computational overhead on top of existing methods such as Fast-Lin, CROWN and CNN-Cert. Experiments on small and large MNIST and CIFAR neural network models demonstrate our probabilistic approach can tighten up robustness certificate to around $1.8 \times$ and $3.5 \times$ with at least a $99.99%$ confidence compared with the worst-case robustness certificate by CROWN and CNN-Cert.
Lily Weng, Lam M. Nguyen, Mark S. Squillante, Akhilan Boopathy, Ivan V. Oseledets, Luca Daniel
ICML5