VLDB 2026 Research / reviewers in the wild / expert
Surya Bhupatiraju
dblp:198/1043
· DBLP profile ↗
4ranked-venue papers
0as first author
0since 2021 · last 2020
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4Graphics, computer vision, multimedia, augmented reality and games · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Computing education · 100% | |
| Artificial intelligence
1 paper |
Reinforcement learning · 50% Optimization for machine learning · 50% | |
| Software engineering, system software, and programming languages
1 paper |
Program synthesis and code generation · 100% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computing education
AI education |
0.8 | 2 | 2020 | Model AI Assignments 2020 · AAAI 2020 Model AI Assignments 2019 · AAAI 2019 |
Machine learning › Reinforcement learning › policy optimization
policy gradient |
0.3 | 1 | 2018 | The Mirage of Action-Dependent Baselines in Reinforcement Learning · ICML 2018 |
Machine learning › Optimization for machine learning
variance reduction |
0.3 | 1 | 2018 | The Mirage of Action-Dependent Baselines in Reinforcement Learning · ICML 2018 |
Program synthesis and code generation › inductive program synthesis
neural program induction |
0.3 | 1 | 2017 | RobustFill: Neural Program Learning under Noisy I/O · ICML 2017 |
Program synthesis and code generation
neural program synthesis |
0.3 | 1 | 2017 | RobustFill: Neural Program Learning under Noisy I/O · ICML 2017 |
Methods — techniques the papers use, named apart from their topics
variance decomposition · 0.3value function parameterization · 0.3attention RNN · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | Model AI Assignments 2020abstractThe Model AI Assignments session seeks to gather and disseminate the best assignment designs of the Artificial Intelligence (AI) Education community. Recognizing that assignments form the core of student learning experience, we here present abstracts of nine AI assignments from the 2020 session that are easily adoptable, playfully engaging, and flexible for a variety of instructor needs. Assignment specifications and supporting resources may be found at http://modelai.gettysburg.edu. Todd W. Neller, Stephen Keeley, Michael Guerzhoy, Wolfgang Hönig, Jiaoyang Li 0001, Sven Koenig, Ameet Soni, Krista Thomason, Lisa Zhang 0003, Bibin Sebastian, Cinjon Resnick, Avital Oliver, Surya Bhupatiraju, Kumar Krishna Agrawal, James Allingham, Sejong Yoon, Jonathan Chen, Tom Larsen, Marion Neumann, Narges Norouzi, Ryan Hausen, Matthew Evett |
AAAI | 13 |
| 2019 | Model AI Assignments 2019abstractThe Model AI Assignments session seeks to gather and disseminate the best assignment designs of the Artificial Intelligence (AI) Education community. Recognizing that assignments form the core of student learning experience, we here present abstracts of ten AI assignments from the 2019 session that are easily adoptable, playfully engaging, and flexible for a variety of instructor needs. Assignment specifications and supporting resources may be found at http: //modelai.gettysburg.edu. Todd W. Neller, Raja Sooriamurthi, Michael Guerzhoy, Lisa Zhang 0003, Paul G. Talaga, Christopher Archibald, Adam Summerville, Joseph C. Osborn, Cinjon Resnick, Avital Oliver, Surya Bhupatiraju, Kumar Krishna Agrawal, Nate Derbinsky, Elena Strange, Marion Neumann, Jonathan Chen, Zac Christensen, Michael Wollowski, Oscar Youngquist |
AAAI | 11 |
| 2018 | The Mirage of Action-Dependent Baselines in Reinforcement LearningabstractPolicy gradient methods are a widely used class of model-free reinforcement learning algorithms where a state-dependent baseline is used to reduce gradient estimator variance. Several recent papers extend the baseline to depend on both the state and action and suggest that this significantly reduces variance and improves sample efficiency without introducing bias into the gradient estimates. To better understand this development, we decompose the variance of the policy gradient estimator and numerically show that learned state-action-dependent baselines do not in fact reduce variance over a state-dependent baseline in commonly tested benchmark domains. We confirm this unexpected result by reviewing the open-source code accompanying these prior papers, and show that subtle implementation decisions cause deviations from the methods presented in the papers and explain the source of the previously observed empirical gains. Furthermore, the variance decomposition highlights areas for improvement, which we demonstrate by illustrating a simple change to the typical value function parameterization that can significantly improve performance. George Tucker, Surya Bhupatiraju, Shixiang Gu, Richard E. Turner, Zoubin Ghahramani, Sergey Levine |
ICML | 2 |
| 2017 | RobustFill: Neural Program Learning under Noisy I/OabstractThe problem of automatically generating a computer program from some specification has been studied since the early days of AI. Recently, two competing approaches for `automatic program learning’ have received significant attention: (1) `neural program synthesis’, where a neural network is conditioned on input/output (I/O) examples and learns to generate a program, and (2) `neural program induction’, where a neural network generates new outputs directly using a latent program representation. Here, for the first time, we directly compare both approaches on a large-scale, real-world learning task and we additionally contrast to rule-based program synthesis, which uses hand-crafted semantics to guide the program generation. Our neural models use a modified attention RNN to allow encoding of variable-sized sets of I/O pairs, which achieve 92\% accuracy on a real-world test set, compared to the 34\% accuracy of the previous best neural synthesis approach. The synthesis model also outperforms a comparable induction model on this task, but we more importantly demonstrate that the strength of each approach is highly dependent on the evaluation metric and end-user application. Finally, we show that we can train our neural models to remain very robust to the type of noise expected in real-world data (e.g., typos), while a highly-engineered rule-based system fails entirely. Jacob Devlin, Jonathan Uesato, Surya Bhupatiraju, Rishabh Singh, Abdel-rahman Mohamed, Pushmeet Kohli |
ICML | 3 |