VLDB 2026 Research / reviewers in the wild / expert
Abhishek Panigrahi
dblp:208/4926
· DBLP profile ↗
15ranked-venue papers
7as first author
11since 2021 · last 2025
0000-0002-5965-4784ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 7 first-author · 11 since 2021Databases, data management, data science and information retrieval · 2Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Progressive distillation induces an implicit curriculumabstractKnowledge distillation leverages a teacher model to improve the training of a student model. A persistent challenge is that a better teacher does not always yield a better student, to which a common mitigation is to use additional supervision from several “intermediate” teachers. One empirically validated variant of this principle is progressive distillation, where the student learns from successive intermediate checkpoints of the teacher. Using sparse parity as a sandbox, we identify an implicit curriculum as one mechanism through which progressive distillation accelerates the student’s learning. This curriculum is available only through the intermediate checkpoints but not the final converged one, and imparts both empirical acceleration and a provable sample complexity benefit to the student. We then extend our investigation to Transformers trained on probabilistic context-free grammars (PCFGs) and real-world pre-training datasets (Wikipedia and Books). Through probing the teacher model, we identify an analogous implicit curriculum where the model progressively learns features that capture longer context. Our theoretical and empirical findings on sparse parity, complemented by empirical observations on more complex tasks, highlight the benefit of progressive distillation via implicit curriculum across setups. Abhishek Panigrahi, Bingbin Liu, Sadhika Malladi, Andrej Risteski, Surbhi Goel |
ICLR | 1 |
| 2025 | Efficient stagewise pretraining via progressive subnetworksabstractRecent developments in large language models have sparked interest in efficient
pretraining methods. Stagewise training approaches to improve efficiency, like
gradual stacking and layer dropping (Reddi et al., 2023; Zhang & He, 2020), have
recently garnered attention. The prevailing view suggests that stagewise dropping
strategies, such as layer dropping, are ineffective, especially when compared to
stacking-based approaches. This paper challenges this notion by demonstrating
that, with proper design, dropping strategies can be competitive, if not better, than
stacking methods. Specifically, we develop a principled stagewise training framework, progressive subnetwork training, which only trains subnetworks within the
model and progressively increases the size of subnetworks during training, until it
trains the full network. We propose an instantiation of this framework — Random
Part Training (RAPTR) — that selects and trains only a random subnetwork (e.g.
depth-wise, width-wise) of the network at each step, progressively increasing the
size in stages. We show that this approach not only generalizes prior works like
layer dropping but also fixes their key issues. Furthermore, we establish a theoretical basis for such approaches and provide justification for (a) increasing complexity of subnetworks in stages, conceptually diverging from prior works on layer
dropping, and (b) stability in loss across stage transitions in presence of key modern architecture components like residual connections and layer norms. Through
comprehensive experiments, we demonstrate that RAPTR can significantly speed
up training of standard benchmarks like BERT and UL2, up to 33% compared to
standard training and, surprisingly, also shows better downstream performance on
UL2, improving QA tasks and SuperGLUE by 1.5%; thereby, providing evidence
of better inductive bias. Abhishek Panigrahi, Nikunj Saunshi, Kaifeng Lyu, Sobhan Miryoosefi, Sashank J. Reddi, Satyen Kale, Sanjiv Kumar |
ICLR | 1 |
| 2025 | Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs?abstractVision Language Models (VLMs) are impressive at visual question answering and image captioning. But they underperform on multi-step visual reasoning—even compared to LLMs on the same tasks presented in text form—giving rise to perceptions of modality imbalance or brittleness. Towards a systematic study of such issues, we introduce a synthetic framework for assessing the ability of VLMs to perform algorithmic visual reasoning, comprising three tasks: Table Readout, Grid Navigation, and Visual Analogy. Each has two levels of difficulty, SIMPLE and HARD, and even the SIMPLE versions are difficult for frontier VLMs. We propose strategies for training on the SIMPLE version of tasks that improve performance on the corresponding HARD task, i.e., simple-to-hard (S2H) generalization. This controlled setup, where each task also has an equivalent text-only version, allows a quantification of the modality imbalance and how it is impacted by training strategy. We show that 1) explicit image-to-text conversion is important in promoting S2H generalization on images, by transferring reasoning from text; 2) conversion can be internalized at test time. We also report results of mechanistic study of this phenomenon. We identify measures of gradient alignment that can identify training strategies that promote better S2H generalization. Ablations highlight the importance of chain-of-thought. Simon Park 0002, Abhishek Panigrahi, Dingli Yu, Anirudh Goyal, Sanjeev Arora |
ICML | 2 |
| 2025 | On the Power of Context-Enhanced Learning in LLMsabstractWe formalize a new concept for LLMs, **context-enhanced learning**. It involves standard gradient-based learning on text except that the context is enhanced with additional data on which no auto-regressive gradients are computed. This setting is a gradient-based analog of usual in-context learning (ICL) and appears in some recent works.
Using a multi-step reasoning task, we prove in a simplified setting that context-enhanced learning can be **exponentially more sample-efficient** than standard learning when the model is capable of ICL. At a mechanistic level, we find that the benefit of context-enhancement arises from a more accurate gradient learning signal.
We also experimentally demonstrate that it appears hard to detect or recover learning materials that were used in the context during training. This may have implications for data security as well as copyright. Abhishek Panigrahi, Sanjeev Arora |
ICML | 2 |
| 2025 | Representing Rule-based Chatbots with TransformersabstractDan Friedman, Abhishek Panigrahi, Danqi Chen. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Dan Friedman, Abhishek Panigrahi, Danqi Chen 0001 |
NAACL (Long Papers) | 2 |
| 2024 | Trainable Transformer in TransformerabstractRecent works attribute the capability of in-context learning (ICL) in large pre-trained language models to implicitly simulating and fine-tuning an internal model (e.g., linear or 2-layer MLP) during inference. However, such constructions require large memory overhead, which makes simulation of more sophisticated internal models intractable. In this work, we propose a new efficient construction, Transformer in Transformer (in short, TINT), that allows a transformer to simulate and fine-tune more complex models during inference (e.g., pre-trained language models). In particular, we introduce innovative approximation techniques that allow a TINT model with less than 2 billion parameters to simulate and fine-tune a 125 million parameter transformer model within a single forward pass. TINT accommodates many common transformer variants and its design ideas also improve the efficiency of past instantiations of simple models inside transformers. We conduct end-to-end experiments to validate the internal fine-tuning procedure of TINT on various language modeling and downstream tasks. For example, even with a limited one-step budget, we observe TINT for a OPT-125M model improves performance by 4 − 16% absolute on average compared to OPT-125M. These findings suggest that large pre-trained language models are capable of performing intricate subroutines. To facilitate further work, a modular and extensible codebase for TINT is included. Abhishek Panigrahi, Sadhika Malladi, Mengzhou Xia, Sanjeev Arora |
ICML | 1 |
| 2023 | Do Transformers Parse while Predicting the Masked Word?abstractPre-trained language models have been shown to encode linguistic structures like parse trees in their embeddings while being trained unsupervised.Some doubts have been raised whether the models are doing parsing or only some computation weakly correlated with it.Concretely: (a) Is it possible to explicitly describe transformers with realistic embedding dimensions, number of heads, etc. that are capable of doing parsing -or even approximate parsing?(b) Why do pre-trained models capture parsing structure?This paper takes a step toward answering these questions in the context of generative modeling with PCFGs.We show that masked language models like BERT or RoBERTa of moderate sizes can approximately execute the Inside-Outside algorithm for the English PCFG (Marcus et al., 1993).We also show that the Inside-Outside algorithm is optimal for masked language modeling loss on the PCFG-generated data.We conduct probing experiments on models pre-trained on PCFG-generated data to show that this not only allows recovery of approximate parse tree, but also recovers marginal span probabilities computed by the Inside-Outside algorithm, which suggests an implicit bias of masked language modeling towards this algorithm.* When | P| < c| Ĩ|, we can simulate the computations in the final layer using c layers with | Ĩ| heads instead of | P| heads.Additionally, we can decrease the embedding size by only storing probabilities for relevant non-terminals. Abhishek Panigrahi, Rong Ge 0001, Sanjeev Arora |
EMNLP | 2 |
| 2023 | Task-Specific Skill Localization in Fine-tuned Language ModelsabstractPre-trained language models can be fine-tuned to solve diverse NLP tasks, including in few-shot settings. Thus fine-tuning allows the model to quickly pick up task-specific "skills," but there has been limited study of where these newly-learnt skills reside inside the massive model. This paper introduces the term skill localization for this problem and proposes a solution. Given the downstream task and a model fine-tuned on that task, a simple optimization is used to identify a very small subset of parameters ($\sim$0.01% of model parameters) responsible for ($>$95%) of the model’s performance, in the sense that grafting the fine-tuned values for just this tiny subset onto the pre-trained model gives performance almost as well as the fine-tuned model. While reminiscent of recent works on parameter-efficient fine-tuning, the novel aspects here are that: (i) No further retraining is needed on the subset (unlike, say, with lottery tickets). (ii) Notable improvements are seen over vanilla fine-tuning with respect to calibration of predictions in-distribution (40-90% error reduction) as well as quality of predictions out-of-distribution (OOD). In models trained on multiple tasks, a stronger notion of skill localization is observed, where the sparse regions corresponding to different tasks are almost disjoint, and their overlap (when it happens) is a proxy for task similarity. Experiments suggest that localization via grafting can assist certain forms continual learning. Abhishek Panigrahi, Nikunj Saunshi, Sanjeev Arora |
ICML | 1 |
| 2022 | Understanding Gradient Descent on the Edge of Stability in Deep LearningabstractDeep learning experiments by \citet{cohen2021gradient} using deterministic Gradient Descent (GD) revealed an Edge of Stability (EoS) phase when learning rate (LR) and sharpness (i.e., the largest eigenvalue of Hessian) no longer behave as in traditional optimization. Sharpness stabilizes around $2/$LR and loss goes up and down across iterations, yet still with an overall downward trend. The current paper mathematically analyzes a new mechanism of implicit regularization in the EoS phase, whereby GD updates due to non-smooth loss landscape turn out to evolve along some deterministic flow on the manifold of minimum loss. This is in contrast to many previous results about implicit bias either relying on infinitesimal updates or noise in gradient. Formally, for any smooth function $L$ with certain regularity condition, this effect is demonstrated for (1) Normalized GD, i.e., GD with a varying LR $\eta_t =\frac{\eta}{\norm{\nabla L(x(t))}}$ and loss $L$; (2) GD with constant LR and loss $\sqrt{L- \min_x L(x)}$. Both provably enter the Edge of Stability, with the associated flow on the manifold minimizing $\lambda_{1}(\nabla^2 L)$. The above theoretical results have been corroborated by an experimental study. Sanjeev Arora, Zhiyuan Li 0005, Abhishek Panigrahi |
ICML | 3 |
| 2022 | On the SDEs and Scaling Rules for Adaptive Gradient AlgorithmsabstractApproximating Stochastic Gradient Descent (SGD) as a Stochastic Differential Equation (SDE) has allowed researchers to enjoy the benefits of studying a continuous optimization trajectory while carefully preserving the stochasticity of SGD. Analogous study of adaptive gradient methods, such as RMSprop and Adam, has been challenging because there were no rigorously proven SDE approximations for these methods. This paper derives the SDE approximations for RMSprop and Adam, giving theoretical guarantees of their correctness as well as experimental validation of their applicability to common large-scaling vision and language settings. A key practical result is the derivation of a square root scaling rule to adjust the optimization hyperparameters of RMSprop and Adam when changing batch size, and its empirical validation in deep learning settings. Sadhika Malladi, Kaifeng Lyu, Abhishek Panigrahi, Sanjeev Arora |
NeurIPS | 3 |
| 2021 | Learning and Generalization in RNNsabstractSimple recurrent neural networks (RNNs) and their more advanced cousins LSTMs etc. have been very successful in sequence modeling. Their theoretical understanding, however, is lacking and has not kept pace with the progress for feedforward networks, where a reasonably complete understanding in the special case of highly overparametrized one-hidden-layer networks has emerged. In this paper, we make progress towards remedying this situation by proving that RNNs can learn functions of sequences. In contrast to the previous work that could only deal with functions of sequences that are sums of functions of individual tokens in the sequence, we allow general functions. Conceptually and technically, we introduce new ideas which enable us to extract information from the hidden state of the RNN in our proofs---addressing a crucial weakness in previous work. We illustrate our results on some regular language recognition problems. Abhishek Panigrahi, Navin Goyal |
NeurIPS | 1 |
| 2020 | Effect of Activation Functions on the Training of Overparametrized Neural Nets
Abhishek Panigrahi, Abhishek Shetty, Navin Goyal |
ICLR | 1 |
| 2019 | Word2Sense: Sparse Interpretable Word EmbeddingsabstractWe present an unsupervised method to generate Word2Sense word embeddings that are interpretable -each dimension of the embedding space corresponds to a fine-grained sense, and the non-negative value of the embedding along the j-th dimension represents the relevance of the j-th sense to the word.The underlying LDA-based generative model can be extended to refine the representation of a polysemous word in a short context, allowing us to use the embeddings in contextual tasks.On computational NLP tasks, Word2Sense embeddings compare well with other word embeddings generated by unsupervised methods.Across tasks such as word similarity, entailment, sense induction, and contextual interpretation, Word2Sense is competitive with the state-of-the-art method for that task.Word2Sense embeddings are at least as sparse and fast to compute as prior art. Abhishek Panigrahi, Harsha Vardhan Simhadri, Chiranjib Bhattacharyya |
ACL (1) | 1 |
| 2019 | DeepTagRec: A Content-cum-User Based Tag Recommendation Framework for Stack Overflow
Suman Kalyan Maity, Abhishek Panigrahi, Sayan Ghosh 0002, Arundhati Banerjee, Pawan Goyal 0002, Animesh Mukherjee 0001 |
ECIR (2) | 2 |
| 2017 | Book Reading Behavior on Goodreads Can Predict the Amazon Best SellersabstractSuccess of a book depends on various intrinsic and extrinsic parameters. We perform a cross-platform study of book reading behavior on Goodreads and attempt to establish the connection between the Goodreads entities and the Amazon best sellers. We analyze the collective reading behavior on Goodreads platform and quantify various characteristic features of the Goodreads entities to identify differences between the Amazon best sellers (ABS) and the other non-best selling books. Using these features we devise a model to predict if a book shall become a best seller after one month (15 days) since its publication. On a balanced set, we are able to achieve a very high avg. accuracy of 88.72% (85.66%) for the prediction where the other competitive class contains books which are randomly selected from the Goodreads dataset. Our method primarily based on features derived from user posts and genre related characteristic properties achieves an improvement of 16.4% over the traditional popularity factors (ratings, reviews) based baseline methods. Suman Kalyan Maity, Abhishek Panigrahi, Animesh Mukherjee 0001 |
ASONAM | 2 |