Daniel W. Barowy

dblp:120/2091 · DBLP profile ↗
← Back
7ranked-venue papers
6as first author
1since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 4 · 4 first-authorHuman-computer interaction and ubiquitous computing · 2 · 2 first-authorSystems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
4 papers
Software testing · 30% Program analysis · 24% Program synthesis and code generation · 20%
Human-computer interaction and pervasive computing
1 paper
Collaborative and social computing · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Distributed systems · 50% Parallel and multicore computing · 50%
Databases, data mining, and information retrieval
1 paper
Data integration and cleaning · 100%

Topics — the 11 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Program analysis
static analysis
0.522018
ExceLint: automatically finding spreadsheet formula errors · Proc. ACM Program. Lang. 2018
CheckCell: data debugging for spreadsheets · OOPSLA 2014
Software testing
fault detection
0.312018
ExceLint: automatically finding spreadsheet formula errors · Proc. ACM Program. Lang. 2018
Software testing › fault detection
spreadsheet error detection
0.312018
ExceLint: automatically finding spreadsheet formula errors · Proc. ACM Program. Lang. 2018
Collaborative and social computing
crowdsourcing
0.312017
VoxPL: Programming with the Wisdom of the Crowd · CHI 2017
Programming languages and type systems
domain-specific languages
0.312017
VoxPL: Programming with the Wisdom of the Crowd · CHI 2017
Data integration and cleaning › data extraction
spreadsheet data extraction
0.212015
FlashRelate: extracting relational data from semi-structured spreadsheets using examples · PLDI 2015
Program synthesis and code generation › programming by example
data extraction synthesis
0.212015
FlashRelate: extracting relational data from semi-structured spreadsheets using examples · PLDI 2015
Program synthesis and code generation
programming by example
0.212015
FlashRelate: extracting relational data from semi-structured spreadsheets using examples · PLDI 2015
Empirical software engineering › end-user programming
spreadsheet analysis
0.212014
CheckCell: data debugging for spreadsheets · OOPSLA 2014
Distributed systems
crowdsourcing
0.112012
AutoMan: a platform for integrating human-based and digital computation · OOPSLA 2012
Parallel and multicore computing
task scheduling
0.112012
AutoMan: a platform for integrating human-based and digital computation · OOPSLA 2012

Methods — techniques the papers use, named apart from their topics

sample size computation · 0.6quality control algorithm · 0.6regular expressions · 0.4program synthesis · 0.4rectangular region analysis · 0.3information-theoretic analysis · 0.3statistical analysis · 0.2program analysis · 0.2
YearPublicationVenuePosition
2022 Riker: Always-Correct and Fast Incremental Builds from Simple Specifications
Charlie Curtsinger, Daniel W. Barowy
USENIX ATC2
2020 Infrastructor: Flexible, No-Infrastructure Tools for Scaling CS
abstract
Demand for computer science education has skyrocketed in the last decade. Although challenging everywhere, scaling up CS course capacities is especially painful at small, liberal arts colleges (SLACs). SLACs tend to have few instructors, few large-capacity classrooms, and little or no dedicated IT support staff. As CS enrollment growth continues to outpace the ability to hire instructional staff, maintaining the quality of the close, nurturing learning environment that SLACs advertise-and students expect-is a major challenge.
Daniel W. Barowy, William Jannen
SIGCSE1
2018 ExceLint: automatically finding spreadsheet formula errors
abstract
Spreadsheets are one of the most widely used programming environments, and are widely deployed in domains like finance where errors can have catastrophic consequences. We present a static analysis specifically designed to find spreadsheet formula errors. Our analysis directly leverages the rectangular character of spreadsheets. It uses an information-theoretic approach to identify formulas that are especially surprising disruptions to nearby rectangular regions. We present ExceLint, an implementation of our static analysis for Microsoft Excel. We demonstrate that ExceLint is fast and effective: across a corpus of 70 spreadsheets, ExceLint takes a median of 8 seconds per spreadsheet, and it significantly outperforms the state of the art analysis.
Daniel W. Barowy, Emery D. Berger, Benjamin G. Zorn
Proc. ACM Program. Lang.1
2017 VoxPL: Programming with the Wisdom of the Crowd
abstract
Having a crowd estimate a numeric value is the original inspiration for the notion of "the wisdom of the crowd." Quality control for such estimated values is challenging because prior, consensus-based approaches for quality control in labeling tasks are not applicable in estimation tasks. We present VoxPL, a high-level programming framework that automatically obtains high-quality crowdsourced estimates of values. The VoxPL domain-specific language lets programmers concisely specify complex estimation tasks with a desired level of confidence and budget. VoxPL's runtime system implements a novel quality control algorithm that automatically computes sample sizes and obtains high quality estimates from the crowd at low cost. To evaluate VoxPL, we implement four estimation applications, ranging from facial feature recognition to calorie counting. The resulting programs are concise---under 200 lines of code---and obtain high quality estimates from the crowd quickly and inexpensively.
Daniel W. Barowy, Emery D. Berger, Daniel G. Goldstein, Siddharth Suri
CHI1
2015 FlashRelate: extracting relational data from semi-structured spreadsheets using examples
abstract
With hundreds of millions of users, spreadsheets are one of the most important end-user applications. Spreadsheets are easy to use and allow users great flexibility in storing data. This flexibility comes at a price: users often treat spreadsheets as a poor man's database, leading to creative solutions for storing high-dimensional data. The trouble arises when users need to answer queries with their data. Data manipulation tools make strong assumptions about data layouts and cannot read these ad-hoc databases. Converting data into the appropriate layout requires programming skills or a major investment in manual reformatting. The effect is that a vast amount of real-world data is "locked-in" to a proliferation of one-off formats. We introduce FlashRelate, a synthesis engine that lets ordinary users extract structured relational data from spreadsheets without programming. Instead, users extract data by supplying examples of output relational tuples. FlashRelate uses these examples to synthesize a program in Flare. Flare is a novel extraction language that extends regular expressions with geometric constructs. An interactive user interface on top of FlashRelate lets end users extract data by point-and-click. We demonstrate that correct Flare programs can be synthesized in seconds from a small set of examples for 43 real-world scenarios. Finally, our case study demonstrates FlashRelate's usefulness addressing the widespread problem of data trapped in corporate and government formats.
Daniel W. Barowy, Sumit Gulwani, Ted Hart, Benjamin G. Zorn
PLDI1
2014 CheckCell: data debugging for spreadsheets
abstract
Testing and static analysis can help root out bugs in programs, but not in data. This paper introduces data debugging, an approach that combines program analysis and statistical analysis to automatically find potential data errors. Since it is impossible to know a priori whether data are erroneous, data debugging instead locates data that has a disproportionate impact on the computation. Such data is either very important, or wrong. Data debugging is especially useful in the context of data-intensive programming environments that intertwine data with programs in the form of queries or formulas.
Daniel W. Barowy, Dimitar Gochev, Emery D. Berger
OOPSLA1
2012 AutoMan: a platform for integrating human-based and digital computation
abstract
Humans can perform many tasks with ease that remain difficult or impossible for computers. Crowdsourcing platforms like Amazon's Mechanical Turk make it possible to harness human-based computational power at an unprecedented scale. However, their utility as a general-purpose computational platform remains limited. The lack of complete automation makes it difficult to orchestrate complex or interrelated tasks. Scheduling more human workers to reduce latency costs real money, and jobs must be monitored and rescheduled when workers fail to complete their tasks. Furthermore, it is often difficult to predict the length of time and payment that should be budgeted for a given task. Finally, the results of human-based computations are not necessarily reliable, both because human skills and accuracy vary widely, and because workers have a financial incentive to minimize their effort.
Daniel W. Barowy, Charlie Curtsinger, Emery D. Berger, Andrew McGregor 0001
OOPSLA1