Casey Casalnuovo

dblp:167/0168 · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
0since 2021 · last 2020
0000-0002-7112-8189ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 7 · 4 first-authorArtificial intelligence and machine learning · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
5 papers
Empirical software engineering · 64% Program analysis · 12% Operating systems · 11%
Network and information security
1 paper
Systems and software security · 100%

Topics — the 8 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Empirical software engineering
mining software repositories
0.832017
GitcProc: a tool for processing and classifying GitHub commits · ISSTA 2017
The sky is not the limit: multitasking across GitHub projects · ICSE 2016
Assert Use in GitHub Projects · ICSE (1) 2015
Empirical software engineering › mining software repositories › commit analysis
commit classification
0.312017
GitcProc: a tool for processing and classifying GitHub commits · ISSTA 2017
Program analysis › binary analysis
deobfuscation
0.312017
Recovering clear, natural identifiers from obfuscated JS names · ESEC/SIGSOFT FSE 2017
Operating systems › resource management › process management
context switching
0.212016
The sky is not the limit: multitasking across GitHub projects · ICSE 2016
Empirical software engineering
developer studies
0.212016
The sky is not the limit: multitasking across GitHub projects · ICSE 2016
Empirical software engineering › developer studies › software teams
newcomer onboarding
0.212015
Developer onboarding in GitHub: the role of prior social links and language experience · ESEC/SIGSOFT FSE 2015
Systems and software security › software protection
code obfuscation
0.112017
Recovering clear, natural identifiers from obfuscated JS names · ESEC/SIGSOFT FSE 2017
Collaborative and social computing
team collaboration
0.112015
Developer onboarding in GitHub: the role of prior social links and language experience · ESEC/SIGSOFT FSE 2015

Methods — techniques the papers use, named apart from their topics

statistical machine translation · 0.6static analysis · 0.6statistical analysis · 0.4empirical study · 0.4regular expression matching · 0.3survey · 0.2ecosystem-level data analysis · 0.2defect analysis · 0.2call graph analysis · 0.2
YearPublicationVenuePosition
2020 Does Surprisal Predict Code Comprehension Difficulty?
Casey Casalnuovo, Premkumar T. Devanbu, Emily Morgan
CogSci1
2019 Test coverage in python programs
abstract
We study code coverage in several popular Python projects: flask, matplotlib, pandas, scikit-learn, and scrapy. Coverage data on these projects is gathered and hosted on the Codecov website, from where this data can be mined. Using this data, and a syntactic parse of the code, we examine the effect of control flow structure, statement type (e.g., if, for) and code age on test coverage. We find that coverage depends on control flow structure, with more deeply nested statements being significantly less likely to be covered. This is a clear effect, which holds up in every project, even when controlling for the age of the line (as determined by git blame). We find that the age of a line per se has a small (but statistically significant) positive effect on coverage. Finally, we find that the kind of statement (try, if, except, raise, etc) has varying effects on coverage, with exception-handling statements being covered much less often. These results suggest that developers in Python projects have difficulty writing test sets that cover deeply-nested and error-handling statements, and might need assistance covering such code.
Hongyu Zhai, Casey Casalnuovo, Premkumar T. Devanbu
MSR2
2019 Studying the difference between natural and programming language corpora
Casey Casalnuovo, Kenji Sagae, Premkumar T. Devanbu
Empir. Softw. Eng.1
2017 GitcProc: a tool for processing and classifying GitHub commits
abstract
Sites such as GitHub have created a vast collection of software artifacts that researchers interested in understanding and improving software systems can use. Current tools for processing such GitHub data tend to target project metadata and avoid source code processing, or process source code in a manner that requires significant effort for each language supported. This paper presents GitcProc, a lightweight tool based on regular expressions and source code blocks, which downloads projects and extracts their project history, including fine-grained source code information and development time bug fixes. GitcProc can track changes to both single-line and block source code structures and associate these changes to the surrounding function context with minimal set up required from users. We demonstrate GitcProc's ability to capture changes in multiple languages by evaluating it on C, C++, Java, and Python projects, and show it finds bug fixes and the context of source code changes effectively with few false positives.
Casey Casalnuovo, Yagnik Suchak, Baishakhi Ray, Cindy Rubio-González
ISSTA1
2017 Recovering clear, natural identifiers from obfuscated JS names
abstract
Well-chosen variable names are critical to source code readability, reusability, and maintainability. Unfortunately, in deployed JavaScript code (which is ubiquitous on the web) the identifier names are frequently minified and overloaded. This is done both for efficiency and also to protect potentially proprietary intellectual property. In this paper, we describe an approach based on statistical machine translation (SMT) that recovers some of the original names from the JavaScript programs minified by the very popular UglifyJS. This simple tool, Autonym, performs comparably to the best currently available deobfuscator for JavaScript, JSNice, which uses sophisticated static analysis. In fact, Autonym is quite complementary to JSNice, performing well when it does not, and vice versa. We also introduce a new tool, JSNaughty, which blends Autonym and JSNice, and significantly outperforms both at identifier name recovery, while remaining just as easy to use as JSNice. JSNaughty is available online at http://jsnaughty.org.
Bogdan Vasilescu, Casey Casalnuovo, Premkumar T. Devanbu
ESEC/SIGSOFT FSE2
2016 The sky is not the limit: multitasking across GitHub projects
abstract
Software development has always inherently required multitasking: developers switch between coding, reviewing, testing, designing, and meeting with colleagues. The advent of software ecosystems like GitHub has enabled something new: the ability to easily switch between projects. Developers also have social incentives to contribute to many projects; prolific contributors gain social recognition and (eventually) economic rewards. Multitasking, however, comes at a cognitive cost: frequent context-switches can lead to distraction, sub-standard work, and even greater stress. In this paper, we gather ecosystem-level data on a group of programmers working on a large collection of projects. We develop models and methods for measuring the rate and breadth of a developers' context-switching behavior, and we study how context-switching affects their productivity. We also survey developers to understand the reasons for and perceptions of multitasking. We find that the most common reason for multitasking is interrelationships and dependencies between projects. Notably, we find that the rate of switching and breadth (number of projects) of a developer's work matter. Developers who work on many projects have higher productivity if they focus on few projects per day. Developers that switch projects too much during the course of a day have lower productivity as they work on more projects overall. Despite these findings, developers perceptions of the benefits of multitasking are varied.
Bogdan Vasilescu, Kelly Blincoe, Qi Xuan 0001, Casey Casalnuovo, Daniela E. Damian, Premkumar T. Devanbu, Vladimir Filkov
ICSE4
2015 Assert Use in GitHub Projects
abstract
Asserts have long been a strongly recommended (if non-functional) adjunct to programs. They certainly don't add any user-evident feature value; and it can take quite some skill and effort to devise and add useful asserts. However, they are believed to add considerable value to the developer. Certainly, they can help with automated verification; but even in the absence of that, claimed advantages include improved understandability, maintainability, easier fault localization and diagnosis, all eventually leading to better software quality. We focus on this latter claim, and use a large dataset of asserts in C and C++ programs to explore the connection between asserts and defect occurrence. Our data suggests a connection: functions with asserts do have significantly fewer defects. This indicates that asserts do play an important role in software quality; we therefore explored further the factors that play a role in assertion placement: specifically, process factors (such as developer experience and ownership) and product factors, particularly interprocedural factors, exploring how the placement of assertions in functions are influenced by local and global network properties of the callgraph. Finally, we also conduct a differential analysis of assertion use across different application domains.
Casey Casalnuovo, Premkumar T. Devanbu, Abílio Oliveira, Vladimir Filkov, Baishakhi Ray
ICSE (1)1
2015 Developer onboarding in GitHub: the role of prior social links and language experience
abstract
The team aspects of software engineering have been a subject of great interest since early work by Fred Brooks and others: how well do people work together in teams? why do people join teams? what happens if teams are distributed? Recently, the emergence of project ecosystems such as GitHub have created an entirely new, higher level of organization. GitHub supports numerous teams; they share a common technical platform (for work activities) and a common social platform (via following, commenting, etc). We explore the GitHub evidence for socialization as a precursor to joining a project, and how the technical factors of past experience and social factors of past connections to team members of a project affect productivity both initially and in the long run. We find developers preferentially join projects in GitHub where they have pre-existing relationships; furthermore, we find that the presence of past social connections combined with prior experience in languages dominant in the project leads to higher productivity both initially and cumulatively. Interestingly, we also find that stronger social connections are associated with slightly less productivity initially, but slightly more productivity in the long run.
Casey Casalnuovo, Bogdan Vasilescu, Premkumar T. Devanbu, Vladimir Filkov
ESEC/SIGSOFT FSE1