Eddie A. Santos

dblp:174/8571 · also Eddie Antonio Santos · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
5since 2021 · last 2025
0000-0001-5337-715XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 5 · 3 first-authorHuman-computer interaction and ubiquitous computing · 5 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author
YearPublicationVenuePosition
2025 Error Messages are Here to Help!
Dennis J. Bouvier, Ellie Lovellette, Eddie A. Santos, Brett A. Becker, Venu G. Dasigi, Jack Forden, Olga Glebova, Swaroop Joshi, Stanislav Kurkovsky, Seán Russell 0001
ITiCSE (2)3
2024 "It's Weird That it Knows What I Want": Usability and Interactions with Copilot for Novice Programmers
abstract
Recent developments in deep learning have resulted in code-generation models that produce source code from natural language and code-based prompts with high accuracy. This is likely to have profound effects in the classroom, where novices learning to code can now use free tools to automatically suggest solutions to programming exercises and assignments. However, little is currently known about how novices interact with these tools in practice. We present the first study that observes students at the introductory level using one such code auto-generating tool, Github Copilot, on a typical introductory programming (CS1) assignment. Through observations and interviews we explore student perceptions of the benefits and pitfalls of this technology for learning, present new observed interaction patterns, and discuss cognitive and metacognitive difficulties faced by students. We consider design implications of these findings, specifically in terms of how tools like Copilot can better support and scaffold the novice programming experience.
James Prather, Brent N. Reeves, Paul Denny 0001, Brett A. Becker, Juho Leinonen 0001, Andrew Luxton-Reilly, Garrett B. Powell, James Finnie-Ansley, Eddie A. Santos
ACM Trans. Comput. Hum. Interact.9
2023 Programming Is Hard - Or at Least It Used to Be: Educational Opportunities and Challenges of AI Code Generation
abstract
The introductory programming sequence has been the focus of much research in computing education. The recent advent of several viable and freely-available AI-driven code generation tools present several immediate opportunities and challenges in this domain. In this position paper we argue that the community needs to act quickly in deciding what possible opportunities can and should be leveraged and how, while also working on overcoming otherwise mitigating the possible challenges. Assuming that the effectiveness and proliferation of these tools will continue to progress rapidly, without quick, deliberate, and concerted efforts, educators will lose advantage in helping shape what opportunities come to be, and what challenges will endure. With this paper we aim to seed this discussion within the computing education community.
Brett A. Becker, Paul Denny 0001, James Finnie-Ansley, Andrew Luxton-Reilly, James Prather, Eddie A. Santos
SIGCSE (1)6
2023 Applying Software Engineering Anti-patterns to Programming Error Messages
abstract
Programming error messages (PEMs) have long been a hindrance to novice programmers. This work aims to establish a catalog of PEM anti-patterns--- common, reoccurring features of PEMs that make them unhelpful or actively harmful to programmers. The goal is for educators to be aware of, and actively teach concrete ways that PEMs can be misleading to students; to encourage language implementers to be cognizant of these; and avoid them when designing error feedback. A pilot study is being conducted to validate the presence of anti-patterns in error messages.
Eddie A. Santos, Ioannis Karvelas, Brett A. Becker
SIGCSE (2)1
2021 On the Computational Modelling of Michif Verbal Morphology
abstract
This paper presents a finite-state computational model of the verbal morphology of Michif.Michif, the official language of the Métis peoples, is a uniquely mixed language with Algonquian and French origins.It is spoken across the Métis homelands in what is now called Canada and the United States, but it is highly endangered with less than 100 speakers.The verbal morphology is remarkably complex, as the already polysynthetic Algonquian patterns are combined with French elements and unique morpho-phonological interactions.The model presented in this paper, LI VERB KAA-OOSHITAHK DI MICHIF handles this complexity by using a series of composed finite-state transducers to model the concatenative morphology and phonological rule alternations that are unique to Michif.Such a rulebased approach is necessary as there is insufficient language data for an approach that uses machine learning.A language model such as LI VERB KAA-OOSHITAHK DI MICHIF furthers the goals of Indigenous computational linguistics in Canada while also supporting the creation of tools for documentation, education, and revitalization that are desired by the Métis community.
Fineen Davis, Eddie A. Santos, Heather Souter
EACL2
2020 The Indigenous Languages Technology project at NRC Canada: An empowerment-oriented approach to developing language software
abstract
Roland Kuhn, Fineen Davis, Alain Désilets, Eric Joanis, Anna Kazantseva, Rebecca Knowles, Patrick Littell, Delaney Lothian, Aidan Pine, Caroline Running Wolf, Eddie Santos, Darlene Stewart, Gilles Boulianne, Vishwa Gupta, Brian Maracle Owennatékha, Akwiratékha’ Martin, Christopher Cox, Marie-Odile Junker, Olivia Sammons, Delasie Torkornoo, Nathan Thanyehténhas Brinklow, Sara Child, Benoît Farley, David Huggins-Daines, Daisy Rosenblum, Heather Souter. Proceedings of the 28th International Conference on Computational Linguistics. 2020.
Roland Kuhn 0001, Fineen Davis, Alain Désilets, Eric Joanis, Anna Kazantseva, Rebecca Knowles, Patrick Littell, Delaney Lothian, Aidan Pine, Caroline Running Wolf, Eddie A. Santos, Darlene A. Stewart, Gilles Boulianne, Vishwa Gupta, Brian Maracle Owennatékha, Akwiratékha' Martin, Christopher Cox, Marie-Odile Junker, Olivia Sammons, Delasie Torkornoo, Nathan Thanyehténhas Brinklow, Sara Child, Benoit Farley, David Huggins-Daines, Daisy Rosenblum, Heather Souter
COLING11
2018 Syntax and sensibility: Using language models to detect and correct syntax errors
abstract
Syntax errors are made by novice and experienced programmers alike; however, novice programmers lack the years of experience that help them quickly resolve these frustrating errors. Standard LR parsers are of little help, typically resolving syntax errors and their precise location poorly. We propose a methodology that locates where syntax errors occur, and suggests possible changes to the token stream that can fix the error identified. This methodology finds syntax errors by using language models trained on correct source code to find tokens that seem out of place. Fixes are synthesized by consulting the language models to determine what tokens are more likely at the estimated error location. We compare n-gram and LSTM (long short-term memory) language models for this task, each trained on a large corpus of Java code collected from GitHub. Unlike prior work, our methodology does not rely that the problem source code comes from the same domain as the training data. We evaluated against a repository of real student mistakes. Our tools are able to find a syntactically-valid fix within its top-2 suggestions, often producing the exact fix that the student used to resolve the error. The results show that this tool and methodology can locate and suggest corrections for syntax errors. Our methodology is of practical use to all programmers, but will be especially useful to novices frustrated with incomprehensible syntax errors.
Eddie A. Santos, Joshua Charles Campbell, Dhvani Patel, Abram Hindle, José Nelson Amaral
SANER1
2018 How does docker affect energy consumption? Evaluating workloads in and out of Docker containers
Eddie A. Santos, Carson McLean, Christopher Solinas, Abram Hindle
J. Syst. Softw.1
2016 Training & Quality Assessment of an Optical Character Recognition Model for Northern Haida
Isabell Hubert Lyall, Antti Arppe, Jordan Lachler, Eddie A. Santos
LREC4
2016 The unreasonable effectiveness of traditional information retrieval in crash report deduplication
abstract
Organizations like Mozilla, Microsoft, and Apple are flooded with thousands of automated crash reports per day. Although crash reports contain valuable information for debugging, there are often too many for developers to examine individually. Therefore, in industry, crash reports are often automatically grouped together in buckets. Ubuntu's repository contains crashes from hundreds of software systems available with Ubuntu. A variety of crash report bucketing methods are evaluated using data collected by Ubuntu's Apport automated crash reporting system. The trade-off between precision and recall of numerous scalable crash deduplication techniques is explored. A set of criteria that a crash deduplication method must meet is presented and several methods that meet these criteria are evaluated on a new dataset. The evaluations presented in this paper show that using off-the-shelf information retrieval techniques, that were not designed to be used with crash reports, outperform other techniques which are specifically designed for the task of crash bucketing at realistic industrial scales. This research indicates that automated crash bucketing still has a lot of room for improvement, especially in terms of identifier tokenization.
Joshua Charles Campbell, Eddie A. Santos, Abram Hindle
MSR2
2016 Judging a commit by its cover: correlating commit message entropy with build status on travis-CI
abstract
Developers summarize their changes to code in commit messages. When a message seems "unusual", however, this puts doubt into the quality of the code contained in the commit. We trained n-gram language models and used cross-entropy as an indicator of commit message "unusualness" of over 120,000 commits from open source projects. Build statuses collected from Travis-CI were used as a proxy for code quality. We then compared the distributions of failed and successful commits with regards to the "unusualness" of their commit message. Our analysis yielded significant results when correlating cross-entropy with build status.
Eddie A. Santos, Abram Hindle
MSR1
2016 Visualizing Project Evolution through Abstract Syntax Tree Analysis
abstract
What is a developer's contribution to a repository? By only counting commits and number of lines changed, existing tools that visualize source code repositories (such as GitHub's graphs) fall short on showing the effective contributions made by each developer. When many commits are viewed as a group, the details are lost. Commit information can be misleading since lines of code give no indication of what was actually being worked on without careful examination of the changed code. Providing a semantic view of this information could provide deeper insights into how software projects evolve since changes to design and features are not clearly visible from line changes alone. We present TypeV: a method for visualizing Java source code repositories. Instead of counting line changes in a commit we extract detailed type information over time by using the differences between abstract syntax trees (ASTs). We are then able to track the additions and deletions of declarations and invocations for each type. Furthermore, we can track each author's type usage over time. Using TypeV, we examine specific cases in well-known repositories where our tool reveals interesting and useful information. We then compare type coverage information from the AST compared to file coverage to determine if unique information is provided by type information.
Michael D. Feist, Eddie A. Santos, Ian Watts, Abram Hindle
VISSOFT2