Abhiram Iyer

dblp:270/9289 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
7since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Trustworthy machine learning · 27% Language models and text generation · 19% Learning theory · 18%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Bioinformatics and computational biology · 100%

Topics — the 14 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology
computational neuroscience
1.522024
Flexible mapping of abstract domains by grid cells via self-supervised extraction and projection of generalized velocity signals · NeurIPS 2024
Flexible Context-Driven Sensory Processing in Dynamical Vision Models · NeurIPS 2024
Natural language and speech › Language models and text generation
compositional generalization
0.912025
Breaking Neural Network Scaling Laws with Modularity · ICLR 2025
Machine learning › Trustworthy machine learning
data leakage
0.912025
Uncovering Latent Memories in Large Language Models · ICLR 2025
Natural language and speech › Language models and text generation › large language model › knowledge in language models
memorization
0.912025
Uncovering Latent Memories in Large Language Models · ICLR 2025
Machine learning › Deep learning architectures and training
modular neural network
0.912025
Breaking Neural Network Scaling Laws with Modularity · ICLR 2025
Machine learning › Trustworthy machine learning
privacy and data protection
0.912025
Uncovering Latent Memories in Large Language Models · ICLR 2025
Machine learning › Learning theory
sample complexity
0.912025
Breaking Neural Network Scaling Laws with Modularity · ICLR 2025
Machine learning › Deep learning architectures and training
biologically inspired neural network
0.812024
Flexible Context-Driven Sensory Processing in Dynamical Vision Models · NeurIPS 2024
Computer vision › 3D vision › multi-view geometry
geometric consistency
0.812024
Flexible mapping of abstract domains by grid cells via self-supervised extraction and projection of generalized velocity signals · NeurIPS 2024
Machine learning › Learning theory
inductive bias
0.812024
Towards Exact Computation of Inductive Bias · IJCAI 2024
Machine learning › Trustworthy machine learning
interpretability
0.812024
Flexible Context-Driven Sensory Processing in Dynamical Vision Models · NeurIPS 2024
Machine learning › Probabilistic and Bayesian machine learning › dynamical system › neural dynamics
neural network dynamics
0.812024
Flexible Context-Driven Sensory Processing in Dynamical Vision Models · NeurIPS 2024
Bioinformatics and computational biology › computational neuroscience
grid cell
0.812024
Flexible mapping of abstract domains by grid cells via self-supervised extraction and projection of generalized velocity signals · NeurIPS 2024
Machine learning › Transfer learning and domain adaptation › knowledge transfer
representation transfer
0.212024
Flexible mapping of abstract domains by grid cells via self-supervised extraction and projection of generalized velocity signals · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

top-down modulation · 1.5self-supervised learning · 1.5neural network · 1.5low-rank modulation · 1.5excitatory-inhibitory networks · 1.5dimensionality reduction · 1.5weight perturbation · 0.9modularity · 0.9learning rule · 0.9cross-entropy loss diagnostic · 0.9
YearPublicationVenuePosition
2025 Breaking Neural Network Scaling Laws with Modularity
abstract
Modular neural networks outperform nonmodular neural networks on tasks ranging from visual question answering to robotics. These performance improvements are thought to be due to modular networks' superior ability to model the compositional and combinatorial structure of real-world problems. However, a theoretical explanation of how modularity improves generalizability, and how to leverage task modularity while training networks remains elusive. Using recent theoretical progress in explaining neural network generalization, we investigate how the amount of training data required to generalize on a task varies with the intrinsic dimensionality of a task's input. We show theoretically that when applied to modularly structured tasks, while nonmodular networks require an exponential number of samples with task dimensionality, modular networks' sample complexity is independent of task dimensionality: modular networks can generalize in high dimensions. We then develop a novel learning rule for modular networks to exploit this advantage and empirically show the improved generalization of the rule, both in- and out-of-distribution, on high-dimensional, modular tasks.
Akhilan Boopathy, Sunshine Jiang, William Yue, Jaedong Hwang, Abhiram Iyer, Ila Fiete
ICLR5
2025 Uncovering Latent Memories in Large Language Models
abstract
Frontier AI systems are making transformative impacts across society, but such benefits are not without costs: models trained on web-scale datasets containing personal and private data raise profound concerns about data privacy and security. Language models are trained on extensive corpora including potentially sensitive or proprietary information, and the risk of data leakage, where the model response reveals pieces of such information, remains inadequately understood. Prior work has investigated that sequence complexity and the number of repetitions are the primary drivers of memorization. In this work, we examine the most vulnerable class of data: highly complex sequences that are presented only once during training. These sequences often contain the most sensitive information and pose considerable risk if memorized. By analyzing the progression of memorization for these sequences throughout training, we uncover a striking observation: many memorized sequences persist in the model's memory, exhibiting resistance to catastrophic forgetting even after just one encounter. Surprisingly, these sequences may not appear memorized immediately after their first exposure but can later be “uncovered” during training, even in the absence of subsequent exposures - a phenomenon we call "latent memorization." Latent memorization presents a serious challenge for data privacy, as sequences that seem hidden at the final checkpoint of a model may still be easily recoverable. We demonstrate how these hidden sequences can be revealed through random weight perturbations, and we introduce a diagnostic test based on cross-entropy loss to accurately identify latent memorized sequences.
Sunny Duan, Mikail Khona, Abhiram Iyer, Rylan Schaeffer, Ila Fiete
ICLR3
2024 Resampling-free Particle Filters in High-dimensions
abstract
State estimation is crucial for the performance and safety of numerous robotic applications. Among the suite of estimation techniques, particle filters have been identified as a powerful solution due to their non-parametric nature. Yet, in high-dimensional state spaces, these filters face challenges such as ’particle deprivation’ which hinders accurate representation of the true posterior distribution. This paper introduces a novel resampling-free particle filter designed to mitigate particle deprivation by forgoing the traditional resampling step. This ensures a broader and more diverse particle set, especially vital in high-dimensional scenarios. Theoretically, our proposed filter is shown to offer a near-accurate representation of the desired posterior distribution in high-dimensional contexts. Empirically, the effectiveness of our approach is underscored through a high-dimensional synthetic state estimation task and a 6D pose estimation derived from videos. We posit that as robotic systems evolve with greater degrees of freedom, particle filters tailored for high-dimensional state spaces will be indispensable.
Akhilan Boopathy, Aneesh Muppidi, Peggy Yang, Abhiram Iyer, William Yue, Ila Fiete
ICRA4
2024 Towards Exact Computation of Inductive Bias
Akhilan Boopathy, William Yue, Jaedong Hwang, Abhiram Iyer, Ila Fiete
IJCAI4
2024 Flexible Context-Driven Sensory Processing in Dynamical Vision Models
abstract
Visual representations become progressively more abstract along the cortical hierarchy. These abstract representations define notions like objects and shapes, but at the cost of spatial specificity. By contrast, low-level regions represent spatially local but simple input features. How do spatially non-specific representations of abstract concepts in high-level areas flexibly modulate the low-level sensory representations in appropriate ways to guide context-driven and goal-directed behaviors across a range of tasks? We build a biologically motivated and trainable neural network model of dynamics in the visual pathway, incorporating local, lateral, and feedforward synaptic connections, excitatory and inhibitory neurons, and long-range top-down inputs conceptualized as low-rank modulations of the input-driven sensory responses by high-level areas. We study this ${\bf D}$ynamical ${\bf C}$ortical ${\bf net}$work ($DCnet$) in a visual cue-delay-search task and show that the model uses its own cue representations to adaptively modulate its perceptual responses to solve the task, outperforming state-of-the-art DNN vision and LLM models. The model's population states over time shed light on the nature of contextual modulatory dynamics, generating predictions for experiments. We fine-tune the same model on classic psychophysics attention tasks, and find that the model closely replicates known reaction time results. This work represents a promising new foundation for understanding and making predictions about perturbations to visual processing in the brain.
Lakshmi Narasimhan Govindarajan, Abhiram Iyer, Valmiki Kothare, Ila Fiete
NeurIPS2
2024 Flexible mapping of abstract domains by grid cells via self-supervised extraction and projection of generalized velocity signals
abstract
Grid cells in the medial entorhinal cortex create remarkable periodic maps of explored space during navigation. Recent studies show that they form similar maps of abstract cognitive spaces. Examples of such abstract environments include auditory tone sequences in which the pitch is continuously varied or images in which abstract features are continuously deformed (e.g., a cartoon bird whose legs stretch and shrink). Here, we hypothesize that the brain generalizes how it maps spatial domains to mapping abstract spaces. To sidestep the computational cost of learning representations for each high-dimensional sensory input, the brain extracts self-consistent, low-dimensional descriptions of displacements across abstract spaces, leveraging the spatial velocity integration of grid cells to efficiently build maps of different domains. Our neural network model for abstract velocity extraction factorizes the content of these abstract domains from displacements within the domains to generate content-independent and self-consistent, low-dimensional velocity estimates. Crucially, it uses a self-supervised geometric consistency constraint that requires displacements along closed loop trajectories to sum to zero, an integration that is itself performed by the downstream grid cell circuit over learning. This process results in high fidelity estimates of velocities and allowed transitions in abstract domains, a crucial prerequisite for efficient map generation in these high-dimensional environments. We also show how our method outperforms traditional dimensionality reduction and deep-learning based motion extraction networks on the same set of tasks. This is the first neural network model to explain how grid cells can flexibly represent different abstract spaces and makes the novel prediction that they should do so while maintaining their population correlation and manifold structure across domains. Fundamentally, our model sheds light on the mechanistic origins of cognitive flexibility and transfer of representations across vastly different domains in brains, providing a potential self-supervised learning (SSL) framework for leveraging similar ideas in transfer learning and data-efficient generalization in machine learning and robotics.
Abhiram Iyer, Sarthak Chandra, Sugandha Sharma, Ila Fiete
NeurIPS1
2021 Plug-And-Play Image Reconstruction Meets Stochastic Variance-Reduced Gradient Methods
abstract
Plug-and-play (PnP) methods have recently emerged as a powerful framework for image reconstruction that can flexibly combine different physics-based observation models with data-driven image priors in the form of denoisers, and achieve state-of-the-art image reconstruction quality in many applications. In this paper, we aim to further improve the computational efficacy of PnP methods by designing a new algorithm that makes use of stochastic variance-reduced gradients (SVRG), a nascent idea to accelerate runtime in stochastic optimization. Compared with existing PnP methods using batch gradients or stochastic gradients, the new algorithm, called PnP-SVRG, achieves comparable or better accuracy of image reconstruction at a much faster computational speed. Extensive numerical experiments are provided to demonstrate the benefits of the proposed algorithm through the application of compressive imaging using partial Fourier measurements in conjunction with a wide variety of popular image denoisers.
Vincent Monardo, Abhiram Iyer, Sean Donegan, Marc De Graef, Yuejie Chi
ICIP2