David Bieber

dblp:200/8035 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
5since 2021 · last 2023
0000-0002-1914-8246ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
6 papers
Program synthesis and code generation · 42% Program analysis · 24% Debugging and program repair · 19%
Artificial intelligence
6 papers
Graph learning · 28% Probabilistic and Bayesian machine learning · 26% Deep learning architectures and training · 18%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Electronic design automation · 100%

Topics — the 18 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Program synthesis and code generation
code generation with language models
0.712023
Can Large Language Models Reason about Program Invariants? · ICML 2023
Program verification
invariant generation
0.712023
Can Large Language Models Reason about Program Invariants? · ICML 2023
Program analysis
static analysis
0.712023
Static Prediction of Runtime Errors by Learning to Execute Programs with External Resource Descriptions · ICLR 2023
Program synthesis and code generation
programming by example
0.612022
TF-Coder: Program Synthesis for Tensor Manipulations · ACM Trans. Program. Lang. Syst. 2022
Machine learning › Graph learning
graph representation learning
0.512021
Learning Semantic Representations to Verify Hardware Designs · NeurIPS 2021
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › machine learning for planning
learning to search
0.512021
BUSTLE: Bottom-Up Program Synthesis Through Learning-Guided Exploration · ICLR 2021
Electronic design automation › hardware verification and test
hardware verification
0.512021
Learning Semantic Representations to Verify Hardware Designs · NeurIPS 2021
Electronic design automation › hardware verification and test
test generation
0.512021
Learning Semantic Representations to Verify Hardware Designs · NeurIPS 2021
Machine learning › Graph learning
graph neural network
0.412020
Learning to Execute Programs with Instruction Pointer Attention Graph Neural Networks · NeurIPS 2020
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
monte carlo estimation
0.412020
Incremental Sampling Without Replacement for Sequence Models · ICML 2020
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models
relational model
0.412020
Global Relational Models of Source Code · ICLR 2020
Machine learning › Optimization for machine learning › mini-batch sampling
sampling without replacement
0.412020
Incremental Sampling Without Replacement for Sequence Models · ICML 2020
Machine learning › Deep learning architectures and training
sequence modeling
0.412020
Incremental Sampling Without Replacement for Sequence Models · ICML 2020
Program analysis › program representation
source code representation
0.412020
Global Relational Models of Source Code · ICLR 2020
Debugging and program repair
automated program repair
0.412019
Neural Program Repair by Jointly Learning to Localize and Repair · ICLR (Poster) 2019
Debugging and program repair › automated program repair
neural program repair
0.412019
Neural Program Repair by Jointly Learning to Localize and Repair · ICLR (Poster) 2019
Machine learning › Deep learning architectures and training › deep learning systems
deep learning framework
0.212022
TF-Coder: Program Synthesis for Tensor Manipulations · ACM Trans. Program. Lang. Syst. 2022
Debugging and program repair
fault localization
0.112019
Neural Program Repair by Jointly Learning to Localize and Repair · ICLR (Poster) 2019

Methods — techniques the papers use, named apart from their topics

graph neural network · 2.3value-based pruning · 1.1bottom-up enumerative search · 1.1learning-guided exploration · 1.0deep architecture · 1.0bottom-up synthesis · 1.0scratchpad prompting · 0.7learning to execute · 0.7large language model · 0.7fine-tuning · 0.7sampling without replacement · 0.4recurrent neural network · 0.4expectation estimation · 0.4attention · 0.4joint learning · 0.4
YearPublicationVenuePosition
2023 Static Prediction of Runtime Errors by Learning to Execute Programs with External Resource Descriptions
David Bieber, Rishab Goel, Daniel Zheng, Hugo Larochelle, Daniel Tarlow
ICLR1
2023 Can Large Language Models Reason about Program Invariants?
abstract
Identifying invariants is an important program analysis task with applications towards program understanding, bug finding, vulnerability analysis, and formal verification. Existing tools for identifying program invariants rely on dynamic analysis, requiring traces collected from multiple executions in order to produce reliable invariants. We study the application of large language models to invariant prediction, finding that models trained on source code and fine-tuned for invariant generation can perform invariant prediction as static rather than dynamic analysis. Using a scratchpad approach where invariants are predicted sequentially through a program gives the best performance, finding invariants statically of quality comparable to those obtained by a dynamic analysis tool with access to five program traces.
Kexin Pei, David Bieber, Kensen Shi, Charles Sutton
ICML2
2022 TF-Coder: Program Synthesis for Tensor Manipulations
abstract
The success and popularity of deep learning is on the rise, partially due to powerful deep learning frameworks such as TensorFlow and PyTorch, which make it easier to develop deep learning models. However, these libraries also come with steep learning curves, since programming in these frameworks is quite different from traditional imperative programming with explicit loops and conditionals. In this work, we present a tool called TF-Coder for programming by example in TensorFlow. TF-Coder uses a bottom-up weighted enumerative search, with value-based pruning of equivalent expressions and flexible type- and value-based filtering to ensure that expressions adhere to various requirements imposed by the TensorFlow library. We train models to predict TensorFlow operations from features of the input and output tensors and natural language descriptions of tasks to prioritize relevant operations during search. TF-Coder solves 63 of 70 real-world tasks within 5 minutes, sometimes finding simpler solutions in less time compared to experienced human programmers.
Kensen Shi, David Bieber, Rishabh Singh
ACM Trans. Program. Lang. Syst.2
2021 BUSTLE: Bottom-Up Program Synthesis Through Learning-Guided Exploration
Augustus Odena, Kensen Shi, David Bieber, Rishabh Singh, Charles Sutton, Hanjun Dai
ICLR3
2021 Learning Semantic Representations to Verify Hardware Designs
abstract
Verification is a serious bottleneck in the industrial hardware design cycle, routinely requiring person-years of effort. Practical verification relies on a "best effort" process that simulates the design on test inputs. This suggests a new research question: Can this simulation data be exploited to learn a continuous representation of a hardware design that allows us to predict its functionality? As a first approach to this new problem, we introduce Design2Vec, a deep architecture that learns semantic abstractions of hardware designs. The key idea is to work at a higher level of abstraction than the gate or the bit level, namely the Register Transfer Level (RTL), which is somewhat analogous to software source code, and can be represented by a graph that incorporates control and data flow. This allows us to learn representations of RTL syntax and semantics using a graph neural network. We apply these representations to several tasks within verification, including predicting what cover points of the design will be exercised by a test, and generating new tests that will exercise desired cover points. We evaluate Design2Vec on three real-world hardware designs, including an industrial chip used in commercial data centers. Our results demonstrate that Design2Vec dramatically outperforms baseline approaches that do not incorporate the RTL semantics, scales to industrial designs, and can generate tests that exercise design points that are currently hard to cover with manually written tests by design verification experts.
Shobha Vasudevan, Wenjie Jiang 0001, David Bieber, Rishabh Singh, Hamid Shojaei, Richard Ho 0001, Charles Sutton
NeurIPS3
2020 Global Relational Models of Source Code
Vincent J. Hellendoorn, Charles Sutton, Rishabh Singh, Petros Maniatis, David Bieber
ICLR5
2020 Incremental Sampling Without Replacement for Sequence Models
abstract
Sampling is a fundamental technique, and sampling without replacement is often desirable when duplicate samples are not beneficial. Within machine learning, sampling is useful for generating diverse outputs from a trained model. We present an elegant procedure for sampling without replacement from a broad class of randomized programs, including generative neural models that construct outputs sequentially. Our procedure is efficient even for exponentially-large output spaces. Unlike prior work, our approach is incremental, i.e., samples can be drawn one at a time, allowing for increased flexibility. We also present a new estimator for computing expectations from samples drawn without replacement. We show that incremental sampling without replacement is applicable to many domains, e.g., program synthesis and combinatorial optimization.
Kensen Shi, David Bieber, Charles Sutton
ICML2
2020 Learning to Execute Programs with Instruction Pointer Attention Graph Neural Networks
abstract
Graph neural networks (GNNs) have emerged as a powerful tool for learning software engineering tasks including code completion, bug finding, and program repair. They benefit from leveraging program structure like control flow graphs, but they are not well-suited to tasks like program execution that require far more sequential reasoning steps than number of GNN propagation steps. Recurrent neural networks (RNNs), on the other hand, are well-suited to long sequential chains of reasoning, but they do not naturally incorporate program structure and generally perform worse on the above tasks. Our aim is to achieve the best of both worlds, and we do so by introducing a novel GNN architecture, the Instruction Pointer Attention Graph Neural Networks (IPA-GNN), which achieves improved systematic generalization on the task of learning to execute programs using control flow graphs. The model arises by considering RNNs operating on program traces with branch decisions as latent variables. The IPA-GNN can be seen either as a continuous relaxation of the RNN model or as a GNN variant more tailored to execution. To test the models, we propose evaluating systematic generalization on learning to execute using control flow graphs, which tests sequential reasoning and use of program structure. More practically, we evaluate these models on the task of learning to execute partial programs, as might arise if using the model as a heuristic function in program synthesis. Results show that the IPA-GNN outperforms a variety of RNN and GNN baselines on both tasks.
David Bieber, Charles Sutton, Hugo Larochelle, Daniel Tarlow
NeurIPS1
2019 Neural Program Repair by Jointly Learning to Localize and Repair
Marko Vasic, Aditya Kanade 0001, Petros Maniatis, David Bieber, Rishabh Singh
ICLR (Poster)4
2017 PixColor: Pixel Recursive Colorization
Sergio Guadarrama, Ryan Dahl, David Bieber, Jonathon Shlens, Mohammad Norouzi 0002, Kevin Murphy 0002
BMVC3