Mauricio Soto

dblp:23/6244 · DBLP profile ↗
← Back
11ranked-venue papers
6as first author
2since 2021 · last 2022
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 6 · 5 first-author · 1 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-authorTheory of computation · 3Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
2 papers
Debugging and program repair · 54% Software testing · 39% Empirical software engineering · 7%
Human-computer interaction and pervasive computing
1 paper
Wearable and physiological sensing · 77% Health and well-being technologies · 23%

Topics — the 5 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Debugging and program repair
automated program repair
1.022022
Quality of Automated Program Repair on Real-World Defects · IEEE Trans. Software Eng. 2022
Improving Patch Quality by Enhancing Key Components of Automatic Program Repair · ASE 2019
Software testing › test quality
test suite quality
0.612022
Quality of Automated Program Repair on Real-World Defects · IEEE Trans. Software Eng. 2022
Debugging and program repair › automated program repair
patch quality
0.412019
Improving Patch Quality by Enhancing Key Components of Automatic Program Repair · ASE 2019
Software testing
test suite evaluation
0.412019
Improving Patch Quality by Enhancing Key Components of Automatic Program Repair · ASE 2019
Health and well-being technologies › health monitoring
wellness monitoring
0.112021
Observing and predicting knowledge worker stress, focus and awakeness in the wild · Int. J. Hum. Comput. Stud. 2021

Methods — techniques the papers use, named apart from their topics

test suite generation · 0.6search-based repair · 0.6mutation operator selection · 0.4
YearPublicationVenuePosition
2022 Quality of Automated Program Repair on Real-World Defects
abstract
Automated program repair is a promising approach to reducing the costs of manual debugging and increasing software quality. However, recent studies have shown that automated program repair techniques can be prone to producing patches of low quality, overfitting to the set of tests provided to the repair technique, and failing to generalize to the intended specification. This paper rigorously explores this phenomenon on real-world Java programs, analyzing the effectiveness of four well-known repair techniques, GenProg, Par, SimFix, and TrpAutoRepair, on defects made by the projects’ developers during their regular development process. We find that: (1) When applied to real-world Java code, automated program repair techniques produce patches for between 10.6 and 19.0 percent of the defects, which is less frequent than when applied to C code. (2) The produced patches often overfit to the provided test suite, with only between 13.8 and 46.1 percent of the patches passing an independent set of tests. (3) Test suite size has an extremely small but significant effect on the quality of the patches, with larger test suites producing higher-quality patches, though, surprisingly, higher-coverage test suites correlate with lower-quality patches. (4) The number of tests that a buggy program fails has a small but statistically significant positive effect on the quality of the produced patches. (5) Test suite provenance, whether the test suite is written by a human or automatically generated, has a significant effect on the quality of the patches, with developer-written tests typically producing higher-quality patches. And (6) the patches exhibit insufficient diversity to improve quality through some method of combining multiple patches. We develop JaRFly, an open-source framework for implementing techniques for automatic search-based improvement of Java programs. Our study uses JaRFly to faithfully reimplement GenProg and TrpAutoRepair to work on Java code, and makes the first public release of an implementation of Par. Unlike prior work, our study carefully controls for confounding factors and produces a methodology, as well as a dataset of automatically-generated test suites, for objectively evaluating the quality of Java repair techniques on real-world defects.
Manish Motwani, Mauricio Soto, Yuriy Brun, René Just, Claire Le Goues
IEEE Trans. Software Eng.2
2021 Observing and predicting knowledge worker stress, focus and awakeness in the wild
Mauricio Soto, Chris Satterfield, Thomas Fritz 0001, Gail C. Murphy, David C. Shepherd, Nicholas A. Kraft
Int. J. Hum. Comput. Stud.1
2019 Improving Patch Quality by Enhancing Key Components of Automatic Program Repair
abstract
The error repair process in software systems is, historically, a resource-consuming task that relies heavily in developer manual effort. Automatic program repair approaches enable the repair of software with minimum human interaction, therefore, mitigating the burden from developers. However, a problem automatically generated patches commonly suffer is generating low-quality patches (which overfit to one program specification, thus not generalizing to an independent oracle evaluation). This work proposes a set of mechanisms to increase the quality of plausible patches including an analysis of test suite behavior and their key characteristics for automatic program repair, analyzing developer behavior to inform the mutation operator selection distribution, and a study of patch diversity as a means to create consolidated higher quality fixes.
Mauricio Soto
ASE1
2018 Common statement kind changes to inform automatic program repair
abstract
The search space for automatic program repair approaches is vast and the search for mechanisms to help restrict this search are increasing. We make a granular analysis based on statement kinds to find which statements are more likely to be modified than others when fixing an error. We construct a corpus for analysis by delimiting debugging regions in the provided dataset and recursively analyze the differences between the Simplified Syntax Trees associated with EditEvent's. We build a distribution of statement kinds with their corresponding likelihood of being modified and we validate the usage of this distribution to guide the statement selection. We then build association rules with different confidence thresholds to describe statement kinds commonly modified together for multi-edit patch creation. Finally we evaluate association rule coverage over a held out test set and find that when using a 95% confidence threshold we can create less and more accurate rules that fully cover 93.8% of the testing instances.
Mauricio Soto, Claire Le Goues
MSR1
2018 Using a probabilistic model to predict bug fixes
abstract
Automatic Software Repair (APR) has significant potential to reduce software maintenance costs by reducing the human effort required to localize and fix bugs. State-of-the-art generate-and-validate APR techniques select between and instantiate various mutation operators to construct candidate patches, informed largely by heuristic probability distributions. This may reduce effectiveness in terms of both efficiency and output quality. In practice, human developers have many options in terms of how to edit code to fix bugs, some of which are far more common than others (e.g., deleting a line of code is more common than adding a new class). We mined the most recent 100 bug-fixing commits from each of the 500 most popular Java projects in GitHub (the largest dataset to date) to create a probabilistic model describing edit distributions. We categorize, compare and evaluate the different mutation operators used in state-of-the-art approaches. We find that a probabilistic modelbased APR approach patches bugs more quickly in the majority of bugs studied, and that the resulting patches are of higher quality than those produced by previous approaches. Finally, we mine association rules for multi-edit source code changes, an understudied but important problem. We validate the association rules by analyzing how much of our corpus can be built from them. Our evaluation indicates that 84.6% of the multi-edit patches from the corpus can be built from the association rules, while maintaining 90% confidence.
Mauricio Soto, Claire Le Goues
SANER1
2018 On distance-preserving elimination orderings in graphs: Complexity and algorithms
David Coudert, Guillaume Ducoffe, Nicolas Nisse, Mauricio Soto
Discret. Appl. Math.4
2017 Analyzing the impact of social attributes on commit integration success
abstract
As the software development community makes it easier to contribute to open source projects, the number of commits and pull requests keep increasing. However, this exciting growth renders it more difficult to only accept quality contributions. Recent research has found that both technical and social factors predict the success of project contributions on GitHub. We take this question a step further, focusing on predicting continuous integration build success based on technical and social factors involved in a commit. Specifically, we investigated if social factors (such as being a core member of the development team, having a large number of followers, or contributing a large number of commits) improve predictions of build success. We found that social factors cause a noticeable increase in predictive power (12%), core team members are more likely to pass the build tests (10%), and users with 1000 or more followers are more likely to pass the build tests (10%).
Mauricio Soto, Zack Coker, Claire Le Goues
MSR1
2016 A deeper look into bug fixes: patterns, replacements, deletions, and additions
abstract
Many implementations of research techniques that automatically repair software bugs target programs written in C. Work that targets Java often begins from or compares to direct translations of such techniques to a Java context. However, Java and C are very different languages, and Java should be studied to inform the construction of repair approaches to target it. We conduct a large-scale study of bug-fixing commits in Java projects, focusing on assumptions underlying common search-based repair approaches. We make observations that can be leveraged to guide high quality automatic software repair to target Java specifically, including common and uncommon statement modifications in human patches and the applicability of previously-proposed patch construction operators in the Java context.
Mauricio Soto, Ferdian Thung, Chu-Pan Wong, Claire Le Goues, David Lo 0001
MSR1
2011 Asymptotic Modularity of Some Graph Classes
Fabien de Montgolfier, Mauricio Soto, Laurent Viennot
ISAAC2
2011 Treewidth and Hyperbolicity of the Internet
abstract
We study the measurement of the Internet according to two graph parameters: tree width and hyper bolicity. Both tell how far from a tree a graph is. They are computed from snapshots of the Internet released by CAIDA, DIMES, AQUALAB, UCLA, Rocket fuel and Strasbourg University, at the AS or at the router level. On the one hand, the tree width of the Internet appears to be quite large and being far from a tree with that respect, reflecting some high degree of connectivity. This proves the existence of a well linked core in the Internet. On the other hand, the hyper bolicity (as a graph parameter) appears to be very low, reflecting a tree-like structure with respect to distances. Additionally, we compute the tree width and hyper bolicity obtained for classical Internet models and compare with the snapshots.
Fabien de Montgolfier, Mauricio Soto, Laurent Viennot
NCA2
2009 Adversarial queuing theory with setups
Marcos A. Kiwi, Mauricio Soto, Christopher Thraves
Theor. Comput. Sci.2