VLDB 2026 Research / reviewers in the wild / expert
Kumar Krishna Agrawal
dblp:190/7111
· DBLP profile ↗
11ranked-venue papers
1as first author
6since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Representation and self-supervised learning · 40% Generative modeling · 14% Language models and text generation · 14% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Computing education · 100% | |
| Computer graphics and multimedia
1 paper |
Audio and music processing · 100% |
Topics — the 24 heaviest of 26, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
large language model training |
0.9 | 1 | 2025 | Tracing the Representation Geometry of Language Models from Pretraining to Post-training · NeurIPS 2025 |
Machine learning › Representation and self-supervised learning › representation analysis
learned representation analysis |
0.9 | 1 | 2025 | Tracing the Representation Geometry of Language Models from Pretraining to Post-training · NeurIPS 2025 |
Machine learning › Representation and self-supervised learning › representation learning
representation geometry |
0.9 | 1 | 2025 | Tracing the Representation Geometry of Language Models from Pretraining to Post-training · NeurIPS 2025 |
Computing education
AI education |
0.8 | 2 | 2020 | Model AI Assignments 2020 · AAAI 2020 Model AI Assignments 2019 · AAAI 2019 |
Machine learning › Representation and self-supervised learning
contrastive learning |
0.8 | 1 | 2024 | Harnessing small projectors and multiple views for efficient vision pretraining · NeurIPS 2024 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning
self-supervised visual representation learning |
0.8 | 1 | 2024 | Harnessing small projectors and multiple views for efficient vision pretraining · NeurIPS 2024 |
Computer vision › Vision and language
visual question answering |
0.8 | 1 | 2024 | Attribute Diversity Determines the Systematicity Gap in VQA · EMNLP 2024 |
Robotics › Robot navigation and mapping
dynamic environments |
0.6 | 1 | 2022 | Context-Aware Streaming Perception in Dynamic Environments · ECCV (38) 2022 |
Machine learning › Representation and self-supervised learning › representation analysis
representation evaluation |
0.6 | 1 | 2022 | $\alpha$-ReQ : Assessing Representation Quality in Self-Supervised Learning by measuring eigenspectrum decay · NeurIPS 2022 |
Robotics › Autonomous driving › perception › perception systems
streaming perception |
0.6 | 1 | 2022 | Context-Aware Streaming Perception in Dynamic Environments · ECCV (38) 2022 |
Machine learning › Reinforcement learning › imitation learning › occupancy matching
adversarial imitation learning |
0.4 | 1 | 2019 | Discriminator-Actor-Critic: Addressing Sample Inefficiency and Reward Bias in Adversarial Imitation Learning · ICLR (Poster) 2019 |
Machine learning › Generative modeling
autoregressive model |
0.4 | 1 | 2019 | Discrete Flows: Invertible Generative Models of Discrete Data · NeurIPS 2019 |
Machine learning › Generative modeling › normalizing flow
discrete flow model |
0.4 | 1 | 2019 | Discrete Flows: Invertible Generative Models of Discrete Data · NeurIPS 2019 |
Machine learning › Generative modeling › generative model
discrete generative model |
0.4 | 1 | 2019 | Discrete Flows: Invertible Generative Models of Discrete Data · NeurIPS 2019 |
Machine learning › Generative modeling
generative adversarial network |
0.4 | 1 | 2019 | GANSynth: Adversarial Neural Audio Synthesis · ICLR (Poster) 2019 |
Machine learning › Reinforcement learning
imitation learning |
0.4 | 1 | 2019 | Discriminator-Actor-Critic: Addressing Sample Inefficiency and Reward Bias in Adversarial Imitation Learning · ICLR (Poster) 2019 |
Machine learning › Reinforcement learning
sample efficiency |
0.4 | 1 | 2019 | Discriminator-Actor-Critic: Addressing Sample Inefficiency and Reward Bias in Adversarial Imitation Learning · ICLR (Poster) 2019 |
Audio and music processing
sound synthesis |
0.4 | 1 | 2019 | GANSynth: Adversarial Neural Audio Synthesis · ICLR (Poster) 2019 |
Natural language and speech › Language models and text generation
instruction tuning |
0.3 | 1 | 2025 | Tracing the Representation Geometry of Language Models from Pretraining to Post-training · NeurIPS 2025 |
Natural language and speech › Language models and text generation › large language model training
post-training |
0.3 | 1 | 2025 | Tracing the Representation Geometry of Language Models from Pretraining to Post-training · NeurIPS 2025 |
Machine learning › Efficient and distributed learning › efficient training
efficient pre-training |
0.2 | 1 | 2024 | Harnessing small projectors and multiple views for efficient vision pretraining · NeurIPS 2024 |
Machine learning › Learning theory
generalization |
0.2 | 1 | 2022 | $\alpha$-ReQ : Assessing Representation Quality in Self-Supervised Learning by measuring eigenspectrum decay · NeurIPS 2022 |
Machine learning › Reinforcement learning
actor-critic methods |
0.1 | 1 | 2019 | Discriminator-Actor-Critic: Addressing Sample Inefficiency and Reward Bias in Adversarial Imitation Learning · ICLR (Poster) 2019 |
Natural language and speech › Language models and text generation › language modeling
character-level language modeling |
0.1 | 1 | 2019 | Discrete Flows: Invertible Generative Models of Discrete Data · NeurIPS 2019 |
Methods — techniques the papers use, named apart from their topics
spectral analysis · 0.9eigenspectrum decay · 0.9effective rank · 0.9projector dimensionality reduction · 0.8orthogonalization constraint · 0.8neural audio synthesis · 0.8multiple augmentations · 0.8adversarial training · 0.8power-law approximation · 0.6linear probe · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Tracing the Representation Geometry of Language Models from Pretraining to Post-trainingabstractStandard training metrics like loss fail to explain the emergence of complex capabilities in large language models. We take a spectral approach to investigate the geometry of learned representations across pretraining and post-training, measuring effective rank (RankMe) and eigenspectrum decay (αReQ). With OLMo (1B-7B) and Pythia (160M-12B) models, we uncover a consistent non-monotonic
sequence of three geometric phases during autoregressive pretraining. The initial “warmup” phase exhibits rapid representational collapse. This is followed by an “entropy-seeking” phase, where the manifold’s dimensionality expands substantially, coinciding with peak n-gram memorization. Subsequently, a “compression-seeking” phase imposes anisotropic consolidation, selectively preserving variance along dominant eigendirections while contracting others, a transition marked with significant improvement in downstream task performance. We show these phases can emerge from a fundamental interplay of cross-entropy optimization under skewed token frequencies and representational bottlenecks (d ≪ |V|). Post-training further transforms geometry: SFT and DPO drive “entropy-seeking” dynamics to integrate specific instructional or preferential data, improving in-distribution performance while degrading out-of-distribution robustness. Conversely, RLVR induces “compression-seeking” , enhancing reward alignment but reducing generation diversity. Melody Zixuan Li, Kumar Krishna Agrawal, Arna Ghosh, Komal Kumar Teru, Adam Santoro, Guillaume Lajoie, Blake A. Richards |
NeurIPS | 2 |
| 2024 | Attribute Diversity Determines the Systematicity Gap in VQAabstractAlthough modern neural networks often generalize to new combinations of familiar concepts, the conditions that enable such compositionality have long been an open question.In this work, we study the systematicity gap in visual question answering: the performance difference between reasoning on previously seen and unseen combinations of object attributes.To test, we introduce a novel diagnostic dataset, CLEVR-HOPE.We find that the systematicity gap is not reduced by increasing the quantity of training data, but is reduced by increasing the diversity of training data.In particular, our experiments suggest that the more distinct attribute type combinations are seen during training, the more systematic we can expect the resulting model to be.We release our data and code at https://github.com/ ikb-a/systematicity-gap-in-vqa. Ian Berlot-Attwell, Kumar Krishna Agrawal, Annabelle Michael Carrell, Naomi Saphra |
EMNLP | 2 |
| 2024 | Harnessing small projectors and multiple views for efficient vision pretrainingabstractRecent progress in self-supervised (SSL) visual representation learning has led to the development of several different proposed frameworks that rely on augmentations of images but use different loss functions.
However, there are few theoretically grounded principles to guide practice, so practical implementation of each SSL framework requires several heuristics to achieve competitive performance.
In this work, we build on recent analytical results to design practical recommendations for competitive and efficient SSL that are grounded in theory.
Specifically, recent theory tells us that existing SSL frameworks are actually minimizing the same idealized loss, which is to learn features that best match the data similarity kernel defined by the augmentations used.
We show how this idealized loss can be reformulated to a functionally equivalent loss that is more efficient to compute.
We study the implicit bias of using gradient descent to minimize our reformulated loss function, and find that using a stronger orthogonalization constraint with a reduced projector dimensionality should yield good representations.
Furthermore, the theory tells us that approximating the reformulated loss should be improved by increasing the number of augmentations, and as such using multiple augmentations should lead to improved convergence.
We empirically verify our findings on CIFAR, STL and Imagenet datasets, wherein we demonstrate an improved linear readout performance when training a ResNet-backbone using our theoretically grounded recommendations.
Remarkably, we also demonstrate that by leveraging these insights, we can reduce the pretraining dataset size by up to 2$\times$ while maintaining downstream accuracy simply by using more data augmentations.
Taken together, our work provides theoretically grounded recommendations that can be used to improve SSL convergence and efficiency. Arna Ghosh, Kumar Krishna Agrawal, Shagun Sodhani, Adam M. Oberman, Blake A. Richards |
NeurIPS | 2 |
| 2022 | Learning from an Exploring Demonstrator: Optimal Reward Estimation for BanditsabstractWe introduce the “inverse bandit” problem of estimating the rewards of a multi-armed bandit instance from observing the learning process of a low-regret demonstrator. Existing approaches to the related problem of inverse reinforcement learning assume the execution of an optimal policy, and thereby suffer from an identifiability issue. In contrast, we propose to leverage the demonstrator’s behavior en route to optimality, and in particular, the exploration phase, for reward estimation. We begin by establishing a general information-theoretic lower bound under this paradigm that applies to any demonstrator algorithm, which characterizes a fundamental tradeoff between reward estimation and the amount of exploration of the demonstrator. Then, we develop simple and efficient reward estimators for upper-confidence-based demonstrator algorithms that attain the optimal tradeoff, showing in particular that consistent reward estimation—free of identifiability issues—is possible under our paradigm. Extensive simulations on both synthetic and semi-synthetic data corroborate our theoretical results. Wenshuo Guo, Kumar Krishna Agrawal, Aditya Grover, Vidya Muthukumar, Ashwin Pananjady |
AISTATS | 2 |
| 2022 | Context-Aware Streaming Perception in Dynamic Environments
Gur-Eyal Sela, Ionel Gog, Justin Wong, Kumar Krishna Agrawal, Xiangxi Mo, Sukrit Kalra, Peter Schafhalter, Eric Leong, Xin Wang 0066, Bharathan Balaji, Joseph Gonzalez 0001, Ion Stoica |
ECCV (38) | 4 |
| 2022 | $\alpha$-ReQ : Assessing Representation Quality in Self-Supervised Learning by measuring eigenspectrum decayabstractSelf-Supervised Learning (SSL) with large-scale unlabelled datasets enables learning useful representations for multiple downstream tasks. However, assessing the quality of such representations efficiently poses nontrivial challenges. Existing approaches train linear probes (with frozen features) to evaluate performance on a given task. This is expensive both computationally, since it requires retraining a new prediction head for each downstream task, and statistically, requires task-specific labels for multiple tasks. This poses a natural question, how do we efficiently determine the "goodness" of representations learned with SSL across a wide range of potential downstream tasks? In particular, a task-agnostic statistical measure of representation quality, that predicts generalization without explicit downstream task evaluation, would be highly desirable. In this work, we analyze characteristics of learned representations $\mathbf{f_\theta}$, in well-trained neural networks with canonical architectures \& across SSL objectives. We observe that the eigenspectrum of the empirical feature covariance $\mathrm{Cov}(\mathbf{f_\theta}$) can be well approximated with the family of power-law distribution. We analytically and empirically (using multiple datasets, e.g. CIFAR, STL10, MIT67, ImageNet) demonstrate that the decay coefficient $\alpha$ serves as a measure of representation quality for tasks that are solvable with a linear readout, i.e. there exist well-defined intervals for $\alpha$ where models exhibit excellent downstream generalization. Furthermore, our experiments suggest that key design parameters in SSL algorithms, such as BarlowTwins, implicitly modulate the decay coefficient of the eigenspectrum ($\alpha$). As $\alpha$ depends only on the features themselves, this measure for model selection with hyperparameter tuning for BarlowTwins enables search with less compute. Kumar Krishna Agrawal, Arnab Kumar Mondal, Arna Ghosh, Blake A. Richards |
NeurIPS | 1 |
| 2020 | Model AI Assignments 2020abstractThe Model AI Assignments session seeks to gather and disseminate the best assignment designs of the Artificial Intelligence (AI) Education community. Recognizing that assignments form the core of student learning experience, we here present abstracts of nine AI assignments from the 2020 session that are easily adoptable, playfully engaging, and flexible for a variety of instructor needs. Assignment specifications and supporting resources may be found at http://modelai.gettysburg.edu. Todd W. Neller, Stephen Keeley, Michael Guerzhoy, Wolfgang Hönig, Jiaoyang Li 0001, Sven Koenig, Ameet Soni, Krista Thomason, Lisa Zhang 0003, Bibin Sebastian, Cinjon Resnick, Avital Oliver, Surya Bhupatiraju, Kumar Krishna Agrawal, James Allingham, Sejong Yoon, Jonathan Chen, Tom Larsen, Marion Neumann, Narges Norouzi, Ryan Hausen, Matthew Evett |
AAAI | 14 |
| 2019 | Model AI Assignments 2019abstractThe Model AI Assignments session seeks to gather and disseminate the best assignment designs of the Artificial Intelligence (AI) Education community. Recognizing that assignments form the core of student learning experience, we here present abstracts of ten AI assignments from the 2019 session that are easily adoptable, playfully engaging, and flexible for a variety of instructor needs. Assignment specifications and supporting resources may be found at http: //modelai.gettysburg.edu. Todd W. Neller, Raja Sooriamurthi, Michael Guerzhoy, Lisa Zhang 0003, Paul G. Talaga, Christopher Archibald, Adam Summerville, Joseph C. Osborn, Cinjon Resnick, Avital Oliver, Surya Bhupatiraju, Kumar Krishna Agrawal, Nate Derbinsky, Elena Strange, Marion Neumann, Jonathan Chen, Zac Christensen, Michael Wollowski, Oscar Youngquist |
AAAI | 12 |
| 2019 | GANSynth: Adversarial Neural Audio Synthesis
Jesse H. Engel, Kumar Krishna Agrawal, Ishaan Gulrajani, Chris Donahue, Adam Roberts |
ICLR (Poster) | 2 |
| 2019 | Discriminator-Actor-Critic: Addressing Sample Inefficiency and Reward Bias in Adversarial Imitation Learning
Ilya Kostrikov, Kumar Krishna Agrawal, Debidatta Dwibedi, Sergey Levine, Jonathan Tompson |
ICLR (Poster) | 2 |
| 2019 | Discrete Flows: Invertible Generative Models of Discrete DataabstractWhile normalizing flows have led to significant advances in modeling high-dimensional continuous distributions, their applicability to discrete distributions remains unknown. In this paper, we show that flows can in fact be extended to discrete events---and under a simple change-of-variables formula not requiring log-determinant-Jacobian computations. Discrete flows have numerous applications. We consider two flow architectures: discrete autoregressive flows that enable bidirectionality, allowing, for example, tokens in text to depend on both left-to-right and right-to-left contexts in an exact language model; and discrete bipartite flows that enable efficient non-autoregressive generation as in RealNVP. Empirically, we find that discrete autoregressive flows outperform autoregressive baselines on synthetic discrete distributions, an addition task, and Potts models; and bipartite flows can obtain competitive performance with autoregressive baselines on character-level language modeling for Penn Tree Bank and text8. Dustin Tran, Keyon Vafa, Kumar Krishna Agrawal, Laurent Dinh, Ben Poole |
NeurIPS | 3 |