Surya Bhupatiraju

dblp:198/1043 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
0since 2021 · last 2020
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4Graphics, computer vision, multimedia, augmented reality and games · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
2 papers
Computing education · 100%
Artificial intelligence
1 paper
Reinforcement learning · 50% Optimization for machine learning · 50%
Software engineering, system software, and programming languages
1 paper
Program synthesis and code generation · 100%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computing education
AI education
0.822020
Model AI Assignments 2020 · AAAI 2020
Model AI Assignments 2019 · AAAI 2019
Machine learning › Reinforcement learning › policy optimization
policy gradient
0.312018
The Mirage of Action-Dependent Baselines in Reinforcement Learning · ICML 2018
Machine learning › Optimization for machine learning
variance reduction
0.312018
The Mirage of Action-Dependent Baselines in Reinforcement Learning · ICML 2018
Program synthesis and code generation › inductive program synthesis
neural program induction
0.312017
RobustFill: Neural Program Learning under Noisy I/O · ICML 2017
Program synthesis and code generation
neural program synthesis
0.312017
RobustFill: Neural Program Learning under Noisy I/O · ICML 2017

Methods — techniques the papers use, named apart from their topics

variance decomposition · 0.3value function parameterization · 0.3attention RNN · 0.3
YearPublicationVenuePosition
2020 Model AI Assignments 2020
abstract
The Model AI Assignments session seeks to gather and disseminate the best assignment designs of the Artificial Intelligence (AI) Education community. Recognizing that assignments form the core of student learning experience, we here present abstracts of nine AI assignments from the 2020 session that are easily adoptable, playfully engaging, and flexible for a variety of instructor needs. Assignment specifications and supporting resources may be found at http://modelai.gettysburg.edu.
Todd W. Neller, Stephen Keeley, Michael Guerzhoy, Wolfgang Hönig, Jiaoyang Li 0001, Sven Koenig, Ameet Soni, Krista Thomason, Lisa Zhang 0003, Bibin Sebastian, Cinjon Resnick, Avital Oliver, Surya Bhupatiraju, Kumar Krishna Agrawal, James Allingham, Sejong Yoon, Jonathan Chen, Tom Larsen, Marion Neumann, Narges Norouzi, Ryan Hausen, Matthew Evett
AAAI13
2019 Model AI Assignments 2019
abstract
The Model AI Assignments session seeks to gather and disseminate the best assignment designs of the Artificial Intelligence (AI) Education community. Recognizing that assignments form the core of student learning experience, we here present abstracts of ten AI assignments from the 2019 session that are easily adoptable, playfully engaging, and flexible for a variety of instructor needs. Assignment specifications and supporting resources may be found at http: //modelai.gettysburg.edu.
Todd W. Neller, Raja Sooriamurthi, Michael Guerzhoy, Lisa Zhang 0003, Paul G. Talaga, Christopher Archibald, Adam Summerville, Joseph C. Osborn, Cinjon Resnick, Avital Oliver, Surya Bhupatiraju, Kumar Krishna Agrawal, Nate Derbinsky, Elena Strange, Marion Neumann, Jonathan Chen, Zac Christensen, Michael Wollowski, Oscar Youngquist
AAAI11
2018 The Mirage of Action-Dependent Baselines in Reinforcement Learning
abstract
Policy gradient methods are a widely used class of model-free reinforcement learning algorithms where a state-dependent baseline is used to reduce gradient estimator variance. Several recent papers extend the baseline to depend on both the state and action and suggest that this significantly reduces variance and improves sample efficiency without introducing bias into the gradient estimates. To better understand this development, we decompose the variance of the policy gradient estimator and numerically show that learned state-action-dependent baselines do not in fact reduce variance over a state-dependent baseline in commonly tested benchmark domains. We confirm this unexpected result by reviewing the open-source code accompanying these prior papers, and show that subtle implementation decisions cause deviations from the methods presented in the papers and explain the source of the previously observed empirical gains. Furthermore, the variance decomposition highlights areas for improvement, which we demonstrate by illustrating a simple change to the typical value function parameterization that can significantly improve performance.
George Tucker, Surya Bhupatiraju, Shixiang Gu, Richard E. Turner, Zoubin Ghahramani, Sergey Levine
ICML2
2017 RobustFill: Neural Program Learning under Noisy I/O
abstract
The problem of automatically generating a computer program from some specification has been studied since the early days of AI. Recently, two competing approaches for `automatic program learning’ have received significant attention: (1) `neural program synthesis’, where a neural network is conditioned on input/output (I/O) examples and learns to generate a program, and (2) `neural program induction’, where a neural network generates new outputs directly using a latent program representation. Here, for the first time, we directly compare both approaches on a large-scale, real-world learning task and we additionally contrast to rule-based program synthesis, which uses hand-crafted semantics to guide the program generation. Our neural models use a modified attention RNN to allow encoding of variable-sized sets of I/O pairs, which achieve 92\% accuracy on a real-world test set, compared to the 34\% accuracy of the previous best neural synthesis approach. The synthesis model also outperforms a comparable induction model on this task, but we more importantly demonstrate that the strength of each approach is highly dependent on the evaluation metric and end-user application. Finally, we show that we can train our neural models to remain very robust to the type of noise expected in real-world data (e.g., typos), while a highly-engineered rule-based system fails entirely.
Jacob Devlin, Jonathan Uesato, Surya Bhupatiraju, Rishabh Singh, Abdel-rahman Mohamed, Pushmeet Kohli
ICML3