VLDB 2026 Research / reviewers in the wild / expert
Christopher Scaffidi
dblp:13/5036 · also Chris Scaffidi
· DBLP profile ↗
38ranked-venue papers
17as first author
0since 2021 · last 2017
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 28 · 12 first-authorSoftware engineering, systems software and programming languages · 7 · 2 first-authorArtificial intelligence and machine learning · 2 · 2 first-authorSystems, architecture and hardware · 1 · 1 first-authorTheory of computation · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
8 papers |
Empirical software engineering · 49% Software maintenance and evolution · 25% Requirements engineering and software design · 16% | |
| Databases, data mining, and information retrieval
2 papers |
Data integration and cleaning · 36% Information retrieval · 32% Data mining · 32% |
Topics — the 17 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Empirical software engineering › user behavior analysis
information foraging theory |
0.6 | 3 | 2016 | Foraging and navigations, fundamentally: developers' predictions of value and cost · SIGSOFT FSE 2016 The whats and hows of programmers' foraging diets · CHI 2013 Reactive information foraging: an empirical investigation of theory-based recommender systems for programmers · CHI 2012 |
Empirical software engineering
developer studies |
0.4 | 2 | 2016 | Foraging and navigations, fundamentally: developers' predictions of value and cost · SIGSOFT FSE 2016 An Information Foraging Theory Perspective on Tools for Debugging, Refactoring, and Reuse Tasks · ACM Trans. Softw. Eng. Methodol. 2013 |
Software maintenance and evolution › program comprehension
code navigation |
0.2 | 1 | 2016 | Foraging and navigations, fundamentally: developers' predictions of value and cost · SIGSOFT FSE 2016 |
Requirements engineering and software design
software notation design |
0.2 | 1 | 2016 | LondonTube: Overcoming Hidden Dependencies in Cloud-Mobile-Web Programming · CHI 2016 |
Empirical software engineering
end-user programming |
0.2 | 3 | 2016 | Tool support for data validation by end-user programmers · ICSE 2008 LondonTube: Overcoming Hidden Dependencies in Cloud-Mobile-Web Programming · CHI 2016 Using assertions to help end-user programmers create dependable web macros · SIGSOFT FSE 2008 |
Requirements engineering and software design
data validation |
0.2 | 2 | 2008 | Tool support for data validation by end-user programmers · ICSE 2008 Topes: reusable abstractions for validating data · ICSE 2008 |
Software maintenance and evolution
refactoring |
0.2 | 1 | 2013 | An Information Foraging Theory Perspective on Tools for Debugging, Refactoring, and Reuse Tasks · ACM Trans. Softw. Eng. Methodol. 2013 |
Software maintenance and evolution › refactoring
refactoring tool support |
0.2 | 1 | 2013 | An Information Foraging Theory Perspective on Tools for Debugging, Refactoring, and Reuse Tasks · ACM Trans. Softw. Eng. Methodol. 2013 |
Data integration and cleaning › data quality
data validation |
0.1 | 1 | 2008 | Topes: reusable abstractions for validating data · ICSE 2008 |
Program verification
assertions |
0.1 | 1 | 2008 | Using assertions to help end-user programmers create dependable web macros · SIGSOFT FSE 2008 |
Software testing
test adequacy |
0.1 | 1 | 2008 | Using assertions to help end-user programmers create dependable web macros · SIGSOFT FSE 2008 |
Empirical software engineering
mining software repositories |
0.1 | 1 | 2016 | Foraging and navigations, fundamentally: developers' predictions of value and cost · SIGSOFT FSE 2016 |
Information retrieval
e-commerce search |
0.1 | 1 | 2007 | Red Opal: product-feature scoring from reviews · EC 2007 |
Data mining › text mining
sentiment analysis |
0.1 | 1 | 2007 | Red Opal: product-feature scoring from reviews · EC 2007 |
Debugging and program repair › human factors in debugging
debugging strategies |
0.0 | 1 | 2013 | The whats and hows of programmers' foraging diets · CHI 2013 |
Software maintenance and evolution
software reuse |
0.0 | 1 | 2013 | An Information Foraging Theory Perspective on Tools for Debugging, Refactoring, and Reuse Tasks · ACM Trans. Softw. Eng. Methodol. 2013 |
Debugging and program repair
fault localization |
0.0 | 1 | 2012 | Reactive information foraging: an empirical investigation of theory-based recommender systems for programmers · CHI 2012 |
Methods — techniques the papers use, named apart from their topics
information foraging theory · 0.6user study · 0.2literature analysis · 0.2cognitive dimensions framework · 0.2qualitative analysis · 0.2scent-based recommendation · 0.1recommender algorithms · 0.1foraging momentum · 0.1empirical study · 0.1assertion generation · 0.1review mining · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2017 | Support for learning while debugging in a distributed visual programming languageabstractThe LondonTube environment includes a visual programming language to ease creation of apps distributed at runtime over mobile devices, browsers and the cloud. However, a typical programmer still learning the language would struggle with debugging a program of realistic size, in large part due to the difficulty of finding and understanding bugs. We have implemented an IDE plugin aimed at showing where in the code the computation breaks down and helping the programmer to understand why that code might not be working. In a between-subjects experiment, novice LondonTube users with these new features asked fewer questions about the language than those without, and they gave the enhanced environment higher usability ratings. Laxmi Ganesan, Christopher Scaffidi, Andrew P. Dove |
VL/HCC | 2 |
| 2017 | Workers who use spreadsheets and who program earn more than similar workers who do neitherabstractPrior work showed that in 2001 and 2003, workers in America who used spreadsheets or databases, and who did programming, earned 9 to 13% more than similar workers who did not use spreadsheets nor did programming. Such a fact, if still true, could help motivate workers to do programming and/or to create spreadsheets. This paper presents a study replicating these analyses using 2012 data. The results show that workers in America who used spreadsheets and who did programming earned approximately 10% more than similar workers who did neither (controlling for employment and demographic characteristics). These results highlight the potential fruitfulness of further research aimed at quantifying the personal financial benefits of using spreadsheets and doing programming. Christopher Scaffidi |
VL/HCC | 1 |
| 2016 | LondonTube: Overcoming Hidden Dependencies in Cloud-Mobile-Web ProgrammingabstractMany disciplines, including health science, increasingly demand custom applications that synthesize cloud, mobile and web functionality. But creating even simple apps is difficult. Why? In this paper, guided by Cognitive Dimensions, we explore the design space for relevant programming notations and supporting tools, and we pinpoint what we hypothesize to be specific obstacles in the creation of cloud-mobile-web apps. Among these is the prevalence of hidden dependencies within code of apps. Based on this analysis, we propose a new notation called LondonTube aimed at making these hidden dependencies visible, thereby helping health scientists to create apps for themselves. A study showed that LondonTube reduced the time to create a cloud-mobile-web app by a factor of over 20, and it reduced questions about hidden dependencies. Christopher Scaffidi, Andrew P. Dove, Tahmid Nabi |
CHI | 1 |
| 2016 | Foraging and navigations, fundamentally: developers' predictions of value and costabstractEmpirical studies have revealed that software developers spend 35%–50% of their time navigating through source code during development activities, yet fundamental questions remain: Are these percentages too high, or simply inherent in the nature of software development? Are there factors that somehow determine a lower bound on how effectively developers can navigate a given information space? Answering questions like these requires a theory that captures the core of developers' navigation decisions. Therefore, we use the central proposition of Information Foraging Theory to investigate developers' ability to predict the value and cost of their navigation decisions. Our results showed that over 50% of developers' navigation choices produced less value than they had predicted and nearly 40% cost more than they had predicted. We used those results to guide a literature analysis, to investigate the extent to which these challenges are met by current research efforts, revealing a new area of inquiry with a rich and crosscutting set of research challenges and open problems. David Piorkowski, Austin Z. Henley, Tahmid Nabi, Scott D. Fleming, Christopher Scaffidi, Margaret M. Burnett |
SIGSOFT FSE | 5 |
| 2016 | Putting information foraging theory to work: Community-based design patterns for programming toolsabstractThe design of programming tools is slow and costly. To ease this process, we developed a design pattern catalog aimed at providing guidance for tool designers. This catalog is grounded in Information Foraging Theory (IFT), which empirical studies have shown to be useful for understanding how developers look for information during development tasks. New design patterns, authored by members of the research community for the catalog, concretely explain how to apply IFT in tool design. In our evaluation, qualitative analyses revealed the community-written design patterns compared well in quality to patterns that we had ourselves published in a smaller, peer-reviewed catalog. Tahmid Nabi, Kyle M. D. Sweeney, Sam Lichlyter, David Piorkowski, Christopher Scaffidi, Margaret M. Burnett, Scott D. Fleming |
VL/HCC | 5 |
| 2016 | Potential financial motivations for end-user programmingabstractResearch has identified multiple reasons why people do end-user programming but has yet to quantify one of the most basic: making more money. This is an important gap in the literature given the current widespread efforts to promote computational thinking skills, because this education campaign is often linked to the argument that end-user programming skills will contribute to workers' long-term career prospects. Therefore, this paper presents a study investigating how the earnings of workers varied as a function of whether they used spreadsheets/databases and/or did programming at work. Examining survey data from 2003 revealed a positive correlation of earnings against these variables, even after controlling for the occupations of workers. Overall, occupation-adjusted earnings were 14% higher for workers who used spreadsheets/databases and who also did programming, versus those who did neither. These results provoke numerous questions for future research regarding what financial benefits workers can expect to obtain in different settings through end-user programming. Christopher Scaffidi |
VL/HCC | 1 |
| 2015 | Linking the Physical with the Perceptual: Health and Exposure Monitoring with Cyber-physical QuestionnairesabstractHealth often depends as much on human choices as on physical phenomena: how people perceive their status and how they decide to respond affect their health and, more generally, their wellness. Supporting health and wellness with a cyber-physical system requires a holistic integration of components for remote monitoring of both physical and perceptual phenomena. This paper presents a system that meets this requirement through cyber-physical questionnaires, which trigger questions based on physical phenomena to record human perceptions. This prototype is a basis for future efforts aimed at evaluating the system in the field and expanding it not only to track human perceptions but also to affect choices and lifestyles. Christopher Scaffidi, Laurel Kincl, Diana Rohlman, Kim Anderson |
DSD | 1 |
| 2015 | A Code-Centric Cluster-Based Approach for Searching Online Support Forums for ProgrammersabstractOnline forums provide peer-to-peer technical support for many user populations, including programmers struggling to master a new language. Programmers can help one another by uploading code samples to such a forum. Unfortunately, finding relevant code samples can prove difficult using existing search engines for large, diverse forums. Therefore, we have prototyped a new kind of code search engine for online forums that draws upon unsupervised machine learning in two ways. First, it displays code samples in visual groupings based on the mutual similarity of code samples. Second, it uses the assignment of code samples to clusters to achieve a form of query expansion, thereby identifying additional search results as potentially useful. We evaluated the system by running it on the forum for the LabVIEW programming language. A textual analysis of posts showed that the unsupervised machine learning algorithm successfully tended to assign code samples to clusters based on topical similarity. An empirical user evaluation confirmed that the new search engine improved on the forum's existing search engine by providing results for more queries, by generating more results per query, and by providing more relevant search results. Christopher Scaffidi, Christopher Chambers, Sheela Surisetty |
ICMLA | 1 |
| 2015 | To fix or to learn? How production bias affects developers' information foraging during debuggingabstractDevelopers performing maintenance activities must balance their efforts to learn the code vs. their efforts to actually change it. This balancing act is consistent with the “production bias” that, according to Carroll's minimalist learning theory, generally affects software users during everyday tasks. This suggests that developers' focus on efficiency should have marked effects on how they forage for the information they think they need to fix bugs. To investigate how developers balance fixing versus learning during debugging, we conducted the first empirical investigation of the interplay between production bias and information foraging. Our theory-based study involved 11 participants: half tasked with fixing a bug, and half tasked with learning enough to help someone else fix it. Despite the subtlety of difference between their tasks, participants foraged remarkably differently-making foraging decisions from different types of “patches,” with different types of information, and succeeding with different foraging tactics. David Piorkowski, Scott D. Fleming, Christopher Scaffidi, Margaret M. Burnett, Irwin Kwan, Austin Z. Henley, Charles Hill 0001, Amber Horvath |
ICSME | 3 |
| 2015 | Behavior-based clustering of visual codeabstractA perennial problem with online repositories of end-user programmers' code is the low level of reuse, including in situations where existing code might aid in learning. This paper presents a formative study of middle-schoolers learning the Scratch animation environment, which revealed that they struggled to find short pieces of code that they could reuse directly, or from which they could discover language primitives (language instructions) to implement desired behavior. In response, we present a model and supporting prototype tool for clustering behaviorally similar code together, as a basis for helping end-user programmers to locate code. We conducted an empirical study confirming that our tool's model for estimating code similarity does correspond well with programmers' perceptions of code's behavioral similarity. Future work will expand on these results by providing new search engines that help end-user programmers to find and reuse visual code from online repositories. Sheela Surisetty, Catherine Law, Christopher Scaffidi |
VL/HCC | 3 |
| 2015 | Idea Garden: Situated Support for Problem Solving by End-User ProgrammersabstractAlthough there have been many advances in end-user programming environments, recent empirical studies report that programming still remains difficult for end-users. We hypothesize that one reason may be lack of effective support for helping end-user programmers problem-solve their own way around barriers they encounter. Therefore, in this paper, we describe the Idea Garden, a concept designed to help end-user programmers generate new ideas and problem-solve when they run into barriers. The Idea Garden has its roots in Minimalist Learning Theory and problem-solving theories. Our proof-of-concept prototype of the Idea Garden concept in the CoScripter end-user programming environment currently targets three barriers reported in end-user programming literature. It does so using an integrated, just-in-time combination of scaffolding for problem-solving strategies, for design patterns and for programming concepts. Our empirical results showed that this approach helped end-user programmers overcome all three types of barriers that our prototype targeted. Jill Cao, Scott D. Fleming, Margaret M. Burnett, Christopher Scaffidi |
Interact. Comput. | 4 |
| 2014 | Code you can use: Searching for web automation scripts based on reusabilityabstractWeb scripting enables users to automate interactions with websites. Online open source repositories provide scripts available for reuse. Yet just because these scripts are open source does not mean they are all reusable: many are specialized and irrelevant to most peoples' needs, while others are hard to understand or learn from. Repositories offer keyword-based search engines to find scripts relevant to specialized needs, but they lack any means for filtering search results according to reusability. To address this shortcoming, we present an approach for creating a model to automatically estimate the reusability of web automation scripts. To test this approach, we prototyped a search engine that uses these reusability estimates to sort one particular kind of web automation scripts, CoScripter macros, according to reusability. An empirical evaluation confirmed that the system's reusability estimates are significantly correlated with user perceptions of macro reusability, thus implying that our approach presents a viable means for helping end-user programmers to find reusable web automation scripts. James Admire, Abbas Al Zawwad, Abdulwahab Almorebah, Sanchit Karve, Christopher Scaffidi |
VL/HCC | 5 |
| 2014 | Towards aiding within-patch information foraging by end-user programmersabstractMany tools help professional programmers with the difficult problem of finding information during code maintenance. The empirical success of these tools can be explained by Information Foraging Theory (IFT) which predicts how a person seeks information by navigating through an information system based on the visual weight of information features presented to the person. Motivated by the success of these tools, we investigated the reasonable expectation that end-user programmers would likewise benefit from tools that increased the relative visual weight of important information features. We prototyped and evaluated two tools, each of which uses an existing algorithm to identify the most important lines of code. One prototype highlights important lines of code; the other prototype hides unimportant lines of code. An empirical study revealed that increasing the relative weight of important information features by highlighting did positively impact the amount of information foraged and the rate of information gained; on the other hand, decreasing the relative weight of unimportant information features by hiding had a modest negative impact. These results reveal opportunities for enhancing existing IFT-based foraging models and applying them to design more effective end-user programming tools for coding, debugging, and code reuse. Balaji Athreya, Christopher Scaffidi |
VL/HCC | 2 |
| 2013 | The whats and hows of programmers' foraging dietsabstractOne of the least studied areas of Information Foraging Theory is diet: the information foragers choose to seek. For example, do foragers choose solely based on cost, or do they stubbornly pursue certain diets regardless of cost? Do their debugging strategies vary with their diets? To investigate "what" and "how" questions like these for the domain of software debugging, we qualitatively analyzed 9 professional developers' foraging goals, goal patterns, and strategies. Participants spent 50% of their time foraging. Of their foraging, 58% fell into distinct dietary patterns - mostly in patterns not previously discussed in the literature. In general, programmers' foraging strategies leaned more heavily toward enrichment than we expected, but different strategies aligned with different goal types. These and our other findings help fill the gap as to what programmers' dietary goals are and how their strategies relate to those goals. David Piorkowski, Scott D. Fleming, Irwin Kwan, Margaret M. Burnett, Christopher Scaffidi, Rachel K. E. Bellamy, Joshua Jordahl |
CHI | 5 |
| 2013 | Smell-driven performance analysis for end-user programmersabstractEnd-user programmers such as scientists and engineers often adopt a visual domain-specific language due to its easy learnability, but then they later encounter problems when trying to create high-performance programs. In response, they typically have had to resort to learning and using general textual languages such as C or Fortran as a supplement or replacement for the visual language. This paper proposes a technique, called Smell-driven performance analysis, for helping end-user programmers to overcome performance problems without leaving the visual dataflow paradigm. The technique involves statically analyzing programs to heuristically detect areas with potential performance problems (“bad smells”), alerting enduser programmers about problems, and advising on how to fix those problems. We have implemented a prototype for applying this technique and conducted a user study in which end-user programmers diagnosed performance problems. The experiment showed our technique increased participants' success rates at finding problems and decreased the time required for finding solutions. Qualitatively, 92% of participants said our technique was helpful, and they listed numerous specific benefits provided. Christopher Chambers, Christopher Scaffidi |
VL/HCC | 2 |
| 2013 | An Information Foraging Theory Perspective on Tools for Debugging, Refactoring, and Reuse TasksabstractTheories of human behavior are an important but largely untapped resource for software engineering research. They facilitate understanding of human developers’ needs and activities, and thus can serve as a valuable resource to researchers designing software engineering tools. Furthermore, theories abstract beyond specific methods and tools to fundamental principles that can be applied to new situations. Toward filling this gap, we investigate the applicability and utility of Information Foraging Theory (IFT) for understanding information-intensive software engineering tasks, drawing upon literature in three areas: debugging, refactoring, and reuse. In particular, we focus on software engineering tools that aim to support information-intensive activities, that is, activities in which developers spend time seeking information. Regarding applicability, we consider whether and how the mathematical equations within IFT can be used to explain why certain existing tools have proven empirically successful at helping software engineers. Regarding utility, we applied an IFT perspective to identify recurring design patterns in these successful tools, and consider what opportunities for future research are revealed by our IFT perspective. Scott D. Fleming, Christopher Scaffidi, David Piorkowski, Margaret M. Burnett, Rachel K. E. Bellamy, Joseph Lawrance, Irwin Kwan |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2012 | Reactive information foraging: an empirical investigation of theory-based recommender systems for programmersabstractInformation Foraging Theory (IFT) has established itself as an important theory to explain how people seek information, but most work has focused more on the theory itself than on how best to apply it. In this paper, we investigate how to apply a reactive variant of IFT (Reactive IFT) to design IFT-based tools, with a special focus on such tools for ill-structured problems. Toward this end, we designed and implemented a variety of recommender algorithms to empirically investigate how to help people with the ill-structured problem of finding where to look for information while debugging source code. We varied the algorithms based on scent type supported (words alone vs. words + code structure), and based on use of foraging momentum to estimate rapidity of foragers' goal changes. Our empirical results showed that (1) using both words and code structure significantly improved the ability of the algorithms to recommend where software developers should look for information; (2) participants used recommendations to discover new places in the code and also as shortcuts to navigate to known places; and (3) low-momentum recommendations were significantly more useful than high-momentum recommendations, suggesting rapid and numerous goal changes in this type of setting. Overall, our contributions include two new recommendation algorithms, empirical evidence about when and why participants found IFT-based recommendations useful, and implications for the design of tools based on Reactive IFT. David Piorkowski, Scott D. Fleming, Christopher Scaffidi, Christopher Bogart, Margaret M. Burnett, Bonnie E. John, Rachel K. E. Bellamy, Calvin Swart |
CHI | 3 |
| 2012 | How well do online forums facilitate discussion and collaboration among novice animation programmers?abstractAnimation programming is a widely-respected approach for helping students to learn programming skills, and online forums are a widely-used approach for helping students to interact with one another. But in what ways, if any, does combining animation programming with online forums lead to useful discussion and collaboration among learners? To answer this question, we analyzed online forum discussions among people who were learning to create animation programs using the Scratch programming environment. We discovered that specific kinds of online posts were more likely than others to be followed by discussion, and we found that the ensuing collaboration often involved the exchange of design ideas and feedback within small groups of users. These findings reveal opportunities for enhancing online forums and surrounding tools so they more effectively facilitate discussion, collaboration, and ultimately development of programming skills. Christopher Scaffidi, Aniket Dahotre |
SIGCSE | 1 |
| 2012 | End-user programmers on the loose: A study of programming on the phone for the phoneabstractMicrosoft TouchDevelop is a programming environment enabling users use their phones to create scripts that run on the mobile phones. This is achieved via a semi-structured editor and a programming language with several distinctive features, such as support for using smartphone hardware. In order to uncover opportunities for future tool development aimed at facilitating end-user programming of phones on phones, we have investigated the kinds of scripts that people are creating with the current tool set as well as what problems they ask for help with solving. This paper is the first to study how end-user programmers “in the wild” are programming mobile phones. In particular, no previous study has investigated the ways in which end users programmatically use mobile phones' special hardware (e.g., GPS, accelerometer, gyroscope) for practical everyday purposes. We discovered that, in essence, people are using TouchDevelop to create apps: downloadable applications with small, fairly reliable feature sets that take advantage of mobile hardware. In addition, we identified several areas for further innovation aimed at enhancing the programming tool and the online repository where users share scripts with one another. Balaji Athreya, Faezeh Bahmani, Alex Diede, Christopher Scaffidi |
VL/HCC | 4 |
| 2012 | From barriers to learning in the idea garden: An empirical studyabstractHow can end-user programming environments better help their users overcome programming barriers? We have been investigating an approach called Idea Gardening, which addresses this problem by helping end users to help themselves overcome barriers in the context of “doing”. In this paper, we report on a qualitative empirical study of how effectively an Idea Garden prototype helped end users overcome programming barriers in the CoScripter environment, and the extent to which participants learned after interacting with our features. Our results showed that 9 out of 10 participants who encountered barriers and then used the Idea Garden, overcame their barriers. Further, all 9 went on to demonstrate evidence of having learned the programming concepts, patterns, and strategies relevant to overcoming these barriers. Jill Cao, Irwin Kwan, Rachel White, Scott D. Fleming, Margaret M. Burnett, Christopher Scaffidi |
VL/HCC | 6 |
| 2012 | Planted-model evaluation of algorithms for identifying differences between spreadsheetsabstractUsers often need to test, debug or reuse spreadsheets. We present a new algorithm that can identify differences between two spreadsheets, providing a basis for future tools to help users compare two versions of a spreadsheet (thereby seeing what is new and needs testing) or two different spreadsheets (thereby seeing which is more appropriate for reuse in a situation). This algorithm, RowColAlign, is a two-dimensional generalization of the classic dynamic programming algorithm for solving the one-dimensional longest common subsequence problem. In addition, we present a new planted model for generating test cases to evaluate this algorithm and others like it, including the greedy SheetDiff algorithm presented in prior work. In our evaluation, our new RowColAlign algorithm made no errors at all on test cases, including test cases comparable to relatively large spreadsheets. Moreover, further analysis revealed that it is unexpected for our new algorithm to make errors except when spreadsheets contain an unrealistically small number of distinct values. These results are extremely encouraging, revealing our algorithm's potential as the basis for future spreadsheet tools. Anna Harutyunyan, Glencora Borradaile, Chris Chambers, Christopher Scaffidi |
VL/HCC | 4 |
| 2012 | Skill Progression Demonstrated by Users in the Scratch Animation EnvironmentabstractThe Scratch environment exemplifies a tool+community approach to teaching elementary programming skills, as it includes a website where users can publish, discuss, and organize animations that are programs. To explore this environment's effectiveness for helping people to develop programming skills, a quantitative analysis of 250 randomly selected users' data, including more than 1,000 of their animations, was performed. Skill based on 4 models that had proven useful in prior empirical studies was measured. Overall, mixed results about the environment's effectiveness were found. Among users who do not drop out, an increasing progression in social skills was found. However, an extremely high drop-out rate was also observed. Moreover, a flat or decreasing level of demonstrated skill was observed on virtually every measure. These results call into question whether simply combining an animation tool and an online community is sufficient for keeping people engaged long enough to learn elementary programming skills. Christopher Scaffidi, Chris Chambers |
Int. J. Hum. Comput. Interact. | 1 |
| 2011 | Obstacles and opportunities with using visual and domain-specific languages in scientific programmingabstractScientific discovery is the lifeblood of technological progress, and end-user programming in turn is increasingly essential to modern science. In order to uncover opportunities to facilitate scientific programming, we interviewed scientists about their choice of tools and languages, as well as the obstacles resulting from those choices. We focused on domain-specific languages (DSLs), particularly visual DSLs, because prior empirical studies had not explored scientists' DSL use in detail. We found that DSLs were indeed used by most of these scientists, and in fact it was typical for scientific projects to use an increasing number of DSLs over time. Our study extended some findings from related work, and it identified obstacles not previously uncovered. In particular, we found that scientists often struggled with managing data complexity, as well as with using version control systems. Our study revealed several opportunities to improve DSLs and related tools, such as for helping scientists to cope with data complexity and for helping them to foresee problems when choosing a language. Christopher Scaffidi |
VL/HCC | 2 |
| 2011 | Modeling programmer navigation: A head-to-head empirical evaluation of predictive modelsabstractSoftware developers frequently need to perform code maintenance tasks, but doing so requires time-consuming navigation through code. A variety of tools are aimed at easing this navigation by using models to identify places in the code that a developer might want to visit, and then providing shortcuts so that the developer can quickly navigate to those locations. To date, however, only a few of these models have been compared head-to-head to assess their predictive accuracy. In particular, we do not know which models are most accurate overall, which are accurate only in certain circumstances, and whether combining models could enhance accuracy. Therefore, we have conducted an empirical study to evaluate the accuracy of a broad range of models for predicting many different kinds of code navigations in sample maintenance tasks. Overall, we found that models tended to perform best if they took into account how recently a developer has viewed pieces of the code, and if models took into account the spatial proximity of methods within the code. We also found that the accuracy of single-factor models can be improved by combining factors, using a spreading-activation based approach, to produce multi-factor models. Based on these results, we offer concrete guidance about how these models could be used to provide enhanced software development tools that ease the difficulty of navigating through code. David Piorkowski, Scott D. Fleming, Christopher Scaffidi, Liza John, Christopher Bogart, Bonnie E. John, Margaret M. Burnett, Rachel K. E. Bellamy |
VL/HCC | 3 |
| 2010 | A qualitative study of animation programming in the wildabstractScratch is the latest iteration in a series of animation tools aimed at teaching programming skills. Scratch, in particular, aims not only to teach technical skills, but also skills related to collaboration and code reuse. In order to assess the strengths and weaknesses of Scratch relative to these goals, we have performed an empirical field study of Scratch animations and associated user comments from the online animation repository. Overall, we found that Scratch represents substantial progress toward its designers' goals, though we also identified several opportunities for significant improvement. In particular, many Scratch programs revealed significant technical mastery of the programming environment by programmers, and some animations even demonstrated design patterns. On the other hand, while the Scratch repository has successfully served as a supportive environment for generating constructive feedback among users, we did not find any occasions within our sample where this interaction led to online collaboration. In addition, we found low levels of code reuse, in terms of both frequency and success. Based on these results, we identify implications for improving the design of animation tools, for using these tools to teach programming skills, and for fostering successful collaboration and code reuse among end-user programmers. Aniket Dahotre, Christopher Scaffidi |
ESEM | 3 |
| 2010 | Struggling to Excel: A Field Study of Challenges Faced by Spreadsheet UsersabstractSpreadsheets have become one of the most widely-adopted software technologies. They have proven useful for performing numeric computations as well as for organizing, manipulating, exploring, and visualizing data. Yet only one aspect of spreadsheets, formulas, has received extensive attention in field studies to date. In this paper, we describe a three-part field study that widens this focus to uncover a broader range of challenges that people encounter when creating and using spreadsheets. This study has revealed several opportunities to improve spreadsheet editors, including developing different modes for spreadsheet creation, improving support for spreadsheet reuse, and helping users to find and use features. Chris Chambers, Christopher Scaffidi |
VL/HCC | 2 |
| 2009 | Intelligently creating and recommending reusable reformatting rulesabstractWhen users combine data from multiple sources into a spreadsheet or dataset, the result is often a mishmash of different formats, since phone numbers, dates, course numbers and other string-like kinds of data can each be written in many different formats. Although spreadsheets provide features for reformatting numbers and a few specific kinds of string data, they do not provide any support for the wide range of other kinds of string data encountered by users. We describe a user interface where a user can describe the formats of each kind of data. We provide an algorithm that uses these formats to automatically generate reformatting rules that transform strings from one format to another. In effect, our system enables users to create a small expert system called a "tope" that can recognize and reformat instances of one kind of data. Later, as the user is working with a spreadsheet, our system recommends appropriate topes for validating and reformatting the data. With a recall of over 80% for a query time of under 1 second, this algorithm is accurate enough and fast enough to make useful recommendations in an interactive setting. A laboratory experiment shows that compared to manual typing, users can reformat sample spreadsheet data more than twice as fast by creating and using topes. Christopher Scaffidi, Brad A. Myers, Mary Shaw |
IUI | 1 |
| 2009 | Predicting reuse of end-user web macro scriptsabstractRepositories of code written by end-user programmers are beginning to emerge, but when a piece of code is new or nobody has yet reused it, then current repositories provide users with no information about whether that code might be appropriate for reuse. Addressing this problem requires predicting reusability based on information that exists when a script is created. To provide such a model for web macro scripts, we identified script traits that might plausibly predict reuse, then used IBM CoScripter repository logs to statistically test how well each corresponded to reuse. We then built a machine learning model that combines the useful traits and evaluated how well it can predict four different types of reuse that we saw in the repository logs. Our model was able to predict reuse from a surprisingly small set of traits. It is simple enough to be explained in only 6-11 rules, making it potentially viable for integration in repository search engines for end-user programmers. Christopher Scaffidi, Christopher Bogart, Margaret M. Burnett, Allen Cypher, Brad A. Myers, Mary Shaw |
VL/HCC | 1 |
| 2008 | Topes: reusable abstractions for validating dataabstractProgrammers often omit input validation when inputs can appear in many different formats or when validation criteria cannot be precisely specified. To enable validation in these situations, we present a new technique that puts valid inputs into a consistent format and that identifies "questionable" inputs which might be valid or invalid, so that these values can be double-checked by a person or a program. Our technique relies on the concept of a "tope", which is an application-independent abstraction describing how to recognize and transform values in a category of data. We present our definition of topes and describe a development environment that supports the implementation and use of topes. Experiments with web application and spreadsheet data indicate that using our technique improves the accuracy and reusability of validation code and also improves the effectiveness of subsequent data cleaning such as duplicate identification. Christopher Scaffidi, Brad A. Myers, Mary Shaw |
ICSE | 1 |
| 2008 | Tool support for data validation by end-user programmersabstractEnd-user programming tools for creating spreadsheets and webforms offer no data types except "string" for storing many kinds of data, such as person names and street addresses. Consequently, these tools cannot automatically validate these data. Christopher Scaffidi, Brad A. Myers, Mary Shaw |
ICSE | 1 |
| 2008 | Using assertions to help end-user programmers create dependable web macrosabstractWeb macros give web browser users ways to "program" tedious tasks, allowing those tasks to be repeated more quickly and reliably than when performed by hand. Web macros face dependability problems of their own, however: changes in websites or failure on the part of end-user programmers to anticipate possible macro behaviors can cause macros to act incorrectly, often in ways that are difficult to detect. We would like to provide at least some of the benefits of software engineering methodologies to the creators of web macros. To do this we adapt assertions to web-macro programming scenarios. While assertions are well-known to professional software engineers, our web macro assertions are unique in their focus on website evolution, are generated automatically, and encode the expectations and assumptions of a rapidly growing group of users who often have limited formal programming expertise. We have integrated our techniques for assertion generation and evaluation into a web macro tool, and performed an empirical study investigating its use. Our results show that the assertions can help web macro users detect macro failures and correct macro faults. Andhy Koesnandar, Sebastian G. Elbaum, Gregg Rothermel, Lorin Hochstein, Christopher Scaffidi, Kathryn T. Stolee |
SIGSOFT FSE | 5 |
| 2008 | End-user programming in the wild: A field study of CoScripter scriptsabstractAlthough a new class of languages has emerged to enable end users to create their own Web applications, little is known about how end-user programmers actually use such languages in the real world. In this paper, we report a field study on over 1400 scripts collected from the Internet which were created by early adopters of CoScripter, a Web macro programming-by-demonstration language. We contrast these Internet scripts with those written by users inside IBM, and describe script usage and re-usage patterns, features used, and users' clever workarounds for features not present in the language. The results show how users grapple with such programming notions as repetition, generalization, and reuse, sometimes inventing their own devices for these. Finally, we discuss the many scripts we found with social implications, whose purposes were to circumvent intended rules, regulations, and usage norm assumptions of a number of Web sites. Christopher Bogart, Margaret M. Burnett, Allen Cypher, Christopher Scaffidi |
VL/HCC | 4 |
| 2007 | Red Opal: product-feature scoring from reviewsabstractOnline shoppers are generally highly task-driven: they have a certain goal in mind, and they are looking for a product with features that are consistent with that goal. Unfortunately, finding a product with specific features is extremely time-consuming using the search functionality provided by existing web sites.In this paper, we present a new search system called Red Opal that enables users to locate products rapidly based on features. Our fully automatic system examines prior customer reviews, identifies product features, and scores each product on each feature. Red Opal uses these scores to determine which products to show when a user specifies a desired product feature. We evaluate our system on four dimensions: precision of feature extraction, efficiency of feature extraction, precision of product scores, and estimated time savings to customers. On each dimension, Red Opal performs better than a comparison system. Christopher Scaffidi, Kevin Bierhoff, Eric Chang, Mikhael Felker, Herman Ng, Chun Jin |
EC | 1 |
| 2007 | A Lightweight Model for End Users' Data: Progress and Future WorkabstractThis research enable end users to create reusable new abstractions for data categories, thereby enabling them to automate these and other tasks by creating programs. "Tope," the Greek word for "place," is the name for such an abstraction in this research, since each abstraction corresponds to a data category that has a natural place in the problem domain (unlike float and int). For example, US phone number would be a tope.Each tope implementation is a small package of executable software functions for recognizing, transforming, and equivalence-testing instances of a data category. The data model underlying a tope is a directed graph. Each graph node corresponds to a format, and each edge corresponds to a transformation between formats. Christopher Scaffidi |
VL/HCC | 1 |
| 2007 | Scenario-Based Requirements for Web Macro ToolsabstractWeb macros automate interactions with Web sites and related information systems. Though Web macro recorders and players have grown in sophistication over the past decade, these tools cannot yet meet many needs of users in daily life. Based on observations of browser users, we have compiled ten scenarios describing tasks that users would benefit from automating. Our analysis of these scenarios yields specific requirements that Web macro tools should support if those tools are to be applicable to these real-life tasks. Our set of requirements constitutes a benchmark for evaluating tools. Christopher Scaffidi, Allen Cypher, Sebastian G. Elbaum, Andhy Koesnandar, Brad A. Myers |
VL/HCC | 1 |
| 2006 | A Lightweight Model for End Users' Domain-Specific DataabstractMany end user programming tools lack adequate support for domain-specific data. We will design a lightweight representation for categories of data, called "topes," and develop simple methods that end users and system administrators can use to define new topes. To evaluate this approach, we will improve programming tools so end users can write programs that recognize data as instances of topes and manipulate them accordingly. We expect that these enhancements will help end users produce higher quality software Christopher Scaffidi |
VL/HCC | 1 |
| 2006 | Dimensions Characterizing Programming Feature Usage by Information WorkersabstractInformation workers such as administrative staff, consultants, and their managers constitute one of the largest groups of end users, yet little research about their usage of programming features is available to guide development of end user programming tools. In this paper, we describe our survey of over 800 information workers and our analysis of their feature usage in applications such as spreadsheets, browsers, and databases. Our factor analysis reveals three clusters of features - macro features, linked structure features, and imperative features - such that information workers with an inclination to use a feature in each cluster also were inclined to use other features in that cluster, even though each cluster spans several tools. We discuss the implications for research aimed at providing end user programming tools for information workers Christopher Scaffidi, Amy J. Ko, Brad A. Myers, Mary Shaw |
VL/HCC | 1 |
| 2005 | Estimating the Numbers of End Users and End User ProgrammersabstractIn 1995, Boehm predicted that by 2005, there would be "55 million performers" of "end user programming" in the United States. The original context and method which generated this number had two weaknesses, both of which we address. First, it relies on undocumented, judgment-based factors to estimate the number of end user programmers based on the total number of end users; we address this weakness by identifying specific end user sub-populations and then estimating their sizes. Second, Boehm's estimate relies on additional undocumented, judgment-based factors to adjust for rising computer usage rates; we address this weakness by integrating fresh Bureau of Labor Statistics (BLS) data and projections as well as a richer estimation method. With these improvements to Boehm's method, we estimate that in 2012 there will be 90 million end users in American workplaces. Of these, we anticipate that over 55 million will use spreadsheets or databases (and therefore may potentially program), while over 13 million will describe themselves as programmers, compared to BLS projections of fewer than 3 million professional programmers. We have validated our improved method by generating estimates for 2001 and 2003, then verifying that our estimates are consistent with existing estimates from other sources. Christopher Scaffidi, Mary Shaw, Brad A. Myers |
VL/HCC | 1 |