John Anvik

dblp:65/4421 · DBLP profile ↗
← Back
30ranked-venue papers
8as first author
8since 2021 · last 2025
0000-0002-6912-1754ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 18 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 10 · 1 first-author · 5 since 2021Systems, architecture and hardware · 5 · 2 first-authorHuman-computer interaction and ubiquitous computing · 3 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Investigating How Course Features Correlate with Student Perceptions of Two-Stage CS Exams
abstract
A two-stage exam (TSE) is an exam format where students write an exam individually and rewrite the same exam in collaboration with a group of peers. This Working Group investigates TSEs in post-secondary computer science courses. By collecting data from a variety of post-secondary institutions, we aim to identify correlations between course environment indicated through the CALI inventory and students' perceptions and performance on TSEs, with consideration for students in under-represented groups. Students will be given surveys upon completion of TSEs which include open-ended questions regarding their group's dynamics and perceptions of participating in a TSE. Instructors will also provide their experiences running and observing TSEs. This research aims to understand best practices for the implementation of TSEs.
Celine Latulipe, Steph McIntyre, John Anvik, Kevin Lin 0001, Sabrin Nowrin, Brian P. Railing, Scott J. Reckinger, Armita Zarnegar
ITiCSE (2)3
2024 Automatically Identifying Planning Comments in Bug Reports (P)
Shraddhaben Devaiya, John Anvik
SEKE2
2024 Comparing Machine Learning and Feature Selection Approaches for Automated Bug Report Assignment (P)
Farjana Yeasmin Omee, John Anvik
SEKE2
2024 Learning Software Engineering Principles with Program Wars v.3.0
abstract
Program Wars is a web-based card game for teaching fundamental concepts of programming and cybersecurity. We propose Program Wars v.3.0, an extension to Program Wars v.2.0, to explore the effectiveness of teaching fundamental software engineering concepts using GBL.
Md. Hasan Tareque, John Anvik
SIGCSE (2)2
2021 Automatically Annotating Sentences for Task-specific Bug Report Summarization
abstract
There is a need to summarize bug reports as they can become long due to many comments from conversations between developers and various DevOps tools. Although automated approaches to bug report summarization have been developed, we believe they aim at the wrong target - getting as close as possible to a gold-standard summary. Instead, researchers should create automated bug report annotation approaches that allow project members to create summaries based on their task-specific information needs. We present such an approach.
Akalanka Galappaththi, John Anvik, Rafat Bin Islam
ASE2
2021 Evaluating Visual Explanation of Bug Report Assignment Recommendations (S)
abstract
Software development projects typically use an issue tracking system where the project members and users can either report faults or request additional features.Each of these reports needs to be triaged to determine such things as the priority of the report or which developers should be assigned to resolve the report.To assist a triager with report assigning, an assignment recommender has been suggested as a means of improving the process.However, proposed assignment recommenders typically present a list of developer names without an explanation of the rationale.This work presents the results of a small user study to validate our approach to visually explaining bug report assignments.
Shayla Azad Bhuyan, John Anvik
SEKE2
2021 Evaluating a Tool for Creating Bug Report Assignment Recommenders (S)
abstract
Large software development projects that use bug tracking systems can become overwhelmed by the number of reports filed.To assist in reducing the workload of project members, researchers have proposed the use of bug report assignment recommenders.To assist project members with the creation of assignment recommenders, we proposed a web-based tool called the Creation Assistant for Supporting Triage Recommenders (CASTR).This paper presents the results of both a laboratory and field study of CASTR.We found that CASTR can create assignment recommenders with accuracy as high as 95%, 80%, and 70% for Top-1, Top-3 and Top-5, respectively.The field study showed that 60% of the participants found CASTR easy to use, whereas the remaining participants found CASTR moderately or slightly easy to use.
Disha Devaiya, John Anvik, Meher Bheree, Farjana Yeasmin Omee
SEKE2
2021 CASTR: Assisting Bug Report Assignment Recommender Creation
abstract
Issue tracking systems are used to make a software development process more manageable, especially for a geographically dispersed team.However, for each bug report, a decision-making process called bug report triage needs to be performed.A common bug report triage decision is the assigning of a developer to a specific bug report.Bug report triage can take significant time and resources.Bug report assignment recommenders have been proposed for reducing the workload of a project member.However, creating a recommender is complex, as project members have to perform several steps such as data preparation and selection of a machine learning algorithm.Although previous work has sought to find specific answers for aspects of the assignment recommender creation process, to the best of our knowledge, only a few other works have examined assisting with this recommender creation process.CASTR (Creation Assistant for Supporting Triage Recommenders) [1, 2] is a platform-independent multi-tier web application.It allows a project member to analyze the dataset using a graphical representation and also assists in configuring project-specific parameters for a machine learning algorithm.This demonstration shows how to use CASTR to create a bug report assignment recommender, which consists of the following steps:
Disha Thakarshibhai Devaiya, John Anvik, Farjana Yeasmin Omee, Meher Bheree
SEKE2
2019 Feature Evaluation for Automatic Bug Report Summarization (S)
abstract
Bug reports can be lengthy due to long descriptions and long conversation threads.Automatic summarization of the text in a bug report can reduce the time spent by software project members on understanding the content of a bug report.Our work further examines Rastkar et al.'s use of a logistic regression model to determine which sentences from the text of a bug report should be extracted for creating a summary.Using their publicly available bug report corpus, which contains manually annotated bug reports, we examined two aspects regarding the features used by the model.First, we examined how much of a reduction occurs in the precision and recall if some of the more complex features are not used.Second, we examined how the use of different feature combinations affects the precision and recall of the models.We found that the absence of some of the complex features resulted in a modest decrease in precision and recall, and confirmed that some features, such as sentence length, were the most significant features for bug report summarization.
Akalanka Galappaththi, John Anvik
SEKE2
2019 Program Wars: A Card Game for Learning Programming and Cybersecurity Concepts
abstract
Although there are many computer science learning games with the goal of teaching programming, such games typically require the person to either learn an existing programming language or the game's own specialized language. This can be intimidating, confusing or frustrating for an individual when they cannot get their "program" to work correctly (e.g. syntax error, infinite loop). Additionally, such games commonly use a puzzle-solving approach that does not appeal to some demographics. This paper presents a programming-language-independent approach to teaching fundamental programming and cybersecurity concepts using simple vocabulary. This approach also uses the familiar activity of playing cards against opponents to create a more dynamic and engaging learning experience. The approach is demonstrated by a web-based game called Program Wars. Results from a user study show that players are able to effectively connect game concepts to actual programming language structures; however, whether players' comprehension of computer programming is improved is unclear.
John Anvik, Vincent Côté, Jace Riehl
SIGCSE1
2017 Parallel Implementation of a Bug Report Assignment Recommender Using Deep Learning
Adrian-Catalin Florea, John Anvik, Razvan Andonie
ICANN (2)2
2016 A feature location approach supported by time-aware weighting of terms associated with developer expertise profiles
Sima Zamani, Sai Peck Lee, Ramin Shokripour, John Anvik
Knowl. Inf. Syst.4
2015 Using Time Series Models for Defect Prediction in Software Release Planning
abstract
A time series model is presented that uses historical project information to predict the number of future defects, given the number of proposed features and improvements to be completed.This allows for hypothetical release plans to be compared by assessing their predicted impact on testing and defect-fixing time.We selected the VARX time series model as a reasonable approach.The accuracy of the model appeared low for a single dataset, but the error was found to be normally distributed.
James Tunnell, John Anvik
SEKE2
2015 A time-based approach to automatic bug report assignment
Ramin Shokripour, John Anvik, Zarinah Mohd Kasirun, Sima Zamani
J. Syst. Softw.2
2014 Assisting Software Projects with Assignment Recomender Creation
John Anvik, Marshall Brooks, Henry Burton, Justin Canada
SEKE1
2014 A noun-based approach to feature location using time-aware term-weighting
Sima Zamani, Sai Peck Lee, Ramin Shokripour, John Anvik
Inf. Softw. Technol.4
2013 Why so complicated? simple term filtering and weighting for location-based bug report assignment recommendation
abstract
Large software development projects receive many bug reports and each of these reports needs to be triaged. An important step in the triage process is the assignment of the report to a developer. Most previous efforts towards improving bug report assignment have focused on using an activity-based approach. We address some of the limitations of activity-based approaches by proposing a two-phased location-based approach where bug report assignment recommendations are based on the predicted location of the bug. The proposed approach utilizes a noun extraction process on several information sources to determine bug location information and a simple term weighting scheme to provide a bug report assignment recommendation. We found that by using a location-based approach, we achieved an accuracy of 89.41% and 59.76% when recommending five developers for the Eclipse and Mozilla projects, respectively.
Ramin Shokripour, John Anvik, Zarinah Mohd Kasirun, Sima Zamani
MSR2
2011 Reducing the effort of bug report triage: Recommenders for development-oriented decisions
abstract
A key collaborative hub for many software development projects is the bug report repository. Although its use can improve the software development process in a number of ways, reports added to the repository need to be triaged. A triager determines if a report is meaningful. Meaningful reports are then organized for integration into the project's development process. To assist triagers with their work, this article presents a machine learning approach to create recommenders that assist with a variety of decisions aimed at streamlining the development process. The recommenders created with this approach are accurate; for instance, recommenders for which developer to assign a report that we have created using this approach have a precision between 70% and 98% over five open source projects. As the configuration of a recommender for a particular project can require substantial effort and be time consuming, we also present an approach to assist the configuration of such recommenders that significantly lowers the cost of putting a recommender in place for a project. We show that recommenders for which developer should fix a bug can be quickly configured with this approach and that the configured recommenders are within 15% precision of hand-tuned developer recommenders.
John Anvik, Gail C. Murphy
ACM Trans. Softw. Eng. Methodol.1
2008 An approach to detecting duplicate bug reports using natural language and execution information
abstract
An open source project typically maintains an open bug repository so that bug reports from all over the world can be gathered. When a new bug report is submitted to the repository, a person, called a triager, examines whether it is a duplicate of an existing bug report. If it is, the triager marks it as DUPLICATE and the bug report is removed from consideration for further work. In the literature, there are approaches exploiting only natural language information to detect duplicate bug reports. In this paper we present a new approach that further involves execution information. In our approach, when a new bug report arrives, its natural language information and execution information are compared with those of the existing bug reports. Then, a small number of existing bug reports are suggested to the triager as the most similar bug reports to the new bug report. Finally, the triager examines the suggested bug reports to determine whether the new bug report duplicates an existing bug report. We calibrated our approach on a subset of the Eclipse bug repository and evaluated our approach on a subset of the Firefox bug repository. The experimental results show that our approach can detect 67%-93% of duplicate bug reports in the Firefox bug repository, compared to 43%-72% using natural language information alone.
Xiaoyin Wang, Lu Zhang 0023, Tao Xie 0001, John Anvik, Jiasu Sun
ICSE4
2008 Task articulation in software maintenance: Integrating source code annotations with an issue tracking system
abstract
Managing and articulating development tasks is an important aspect of software maintenance. Developers already have a variety of specialty tools to support task management, for example, issue tracking and configuration management software, but they also make use of other tools within their software engineering environments to support software task management. An example of one such mechanism is the appropriation of source code comments to document finer grained details of formally specified tasks. In this research, we propose and present a tool that integrates these source code annotations with an issue tracking management system. We describe how this tool addresses deficiencies that occur in task management and propose future research to improve task management.
John Anvik, Margaret-Anne D. Storey
ICSM1
2006 Visual Explanation of Evidence with Additive Classifiers
Brett Poulin, Roman Eisner, Duane Szafron, Paul Lu, Russell Greiner, David S. Wishart, Alona Fyshe, Brandon Pearcy, John Anvik
AAAI10
2006 Automating bug report assignment
abstract
Open-source development projects typically support an open bug repository to which both developers and users can report bugs. A report that appears in this repository must be triaged to determine if the report is one which requires attention and if it is, which developer will be assigned the responsibility of resolving the report. Large open-source developments are burdened by the rate at which new bug reports appear in the bug repository. The thesis of this work is that the task of triage can be eased by using a semi-automated approach to assign bug reports to developers. The approach consists of constructing a recommender for bug assignments; examined are both a range of algorithms that can be used and the various kinds of information provided to the algorithms. The proposed work seeks to determine through human experimentation a sufficient level of precision for the recommendations, and to analytically determine the trade-offs of the various algorithmic and information choices.
John Anvik
ICSE1
2006 Who should fix this bug?
abstract
Open source development projects typically support an open bug repository to which both developers and users can report bugs. The reports that appear in this repository must be triaged to determine if the report is one which requires attention and if it is, which developer will be assigned the responsibility of resolving the report. Large open source developments are burdened by the rate at which new bug reports appear in the bug repository. In this paper, we present a semi-automated approach intended to ease one part of this process, the assignment of reports to a developer. Our approach applies a machine learning algorithm to the open bug repository to learn the kinds of reports each developer resolves. When a new report arrives, the classifier produced by the machine learning technique suggests a small number of developers suitable to resolve the report. With this approach, we have reached precision levels of 57% and 64% on the Eclipse and Firefox development projects respectively. We have also applied our approach to the gcc open source development with less positive results. We describe the conditions under which the approach is applicable and also report on the lessons we learned about applying machine learning to repositories used in open source development.
John Anvik, Lyndon Hiew, Gail C. Murphy
ICSE1
2005 Asserting the utility of CO2P3S using the Cowichan Problem Set
John Anvik, Jonathan Schaeffer 0001, Duane Szafron, Kai Tan 0005
J. Parallel Distributed Comput.1
2004 Predicting subcellular localization of proteins using machine-learned classifiers
abstract
MOTIVATION: Identifying the destination or localization of proteins is key to understanding their function and facilitating their purification. A number of existing computational prediction methods are based on sequence analysis. However, these methods are limited in scope, accuracy and most particularly breadth of coverage. Rather than using sequence information alone, we have explored the use of database text annotations from homologs and machine learning to substantially improve the prediction of subcellular location. RESULTS: We have constructed five machine-learning classifiers for predicting subcellular localization of proteins from animals, plants, fungi, Gram-negative bacteria and Gram-positive bacteria, which are 81% accurate for fungi and 92-94% accurate for the other four categories. These are the most accurate subcellular predictors across the widest set of organisms ever published. Our predictors are part of the Proteome Analyst web-service.
Zhiyong Lu, Duane Szafron, Russell Greiner, Paul Lu, David S. Wishart, Brett Poulin, John Anvik, Roman Eisner
Bioinform.7
2003 Why Not Use a Pattern-Based Parallel Programming System?
John Anvik, Jonathan Schaeffer 0001, Duane Szafron, Kai Tan 0005
Euro-Par1
2003 Using generative design patterns to generate parallel code for a distributed memory environment
abstract
A design pattern is a mechanism for encapsulating the knowledge of experienced designers into a re-usable artifact. Parallel design patterns reflect commonly occurring parallel communication and synchronization structures. Our tools, CO2P3S (Correct Object-Oriented Pattern-based Parallel Programming System) and MetaCO2P3S, use generative design patterns. A programmer selects the parallel design patterns that are appropriate for an application, and then adapts the patterns for that specific application by selecting from a small set of code-configuration options. CO2P3S then generates a custom framework for the application that includes all of the structural code necessary for the application to run in parallel. The programmer is only required to write simple code that launches the application and to fill in some application-specific sequential hook routines. We use generative design patterns to take an application specification (parallel design patterns + sequential user code) and use it to generate parallel application code that achieves good performance in shared memory and distributed memory environments. Although our implementations are for Java, the approach we describe is tool and language independent. This paper describes generalizing CO2P3S to generate distributed-memory parallel solutions.
Kai Tan 0005, Duane Szafron, Jonathan Schaeffer 0001, John Anvik, Steve MacDonald
PPoPP4
2002 Pattern-Based Parallel Programming
abstract
The advantages of pattern-based programming have been well-documented in the sequential programming literature. However patterns have yet to make their way into mainstream parallel computing, even though several research tools support them. There are two critical shortcomings of pattern (or template) based systems for parallel programming: lack of extensibility and performance. This paper describes our approach for addressing these problems in the CO/sub 2/P/sub 3/S parallel programming system. CO/sub 2/P/sub 3/S supports multiple levels of abstraction, allowing the user to design an application with high-level patterns, but move to lower levels of abstraction for performance tuning. Patterns are implemented as parameterized templates, allowing the user the ability to customize the pattern to meet their needs. CO/sub 2/P/sub 3/S generates code that is specific to the pattern/parameter combination selected by the user. The MetaCO/sub 2/P/sub 3/S tool addresses extensibility by giving users the ability to design and add new pattern templates to CO/sub 2/P/sub 3/S. Since the pattern templates are stored in a system-independent format, they are suitable for storing in a repository to be shared throughout the user community.
Steven Bromling, Steve MacDonald, John Anvik, Jonathan Schaeffer 0001, Duane Szafron, Kai Tan 0005
ICPP3
2002 Generative Design Patterns
abstract
A design pattern encapsulates the knowledge of object-oriented designers into re-usable artifacts. A design pattern is a descriptive device that fosters software design re-use. There are several reasons why design patterns are not used as generative constructs that support code re-use. The first reason is that design patterns describe a set of solutions to a family of related design problems and it is difficult to generate a single body of code that adequately solves each problem in the family. A second reason is that it is difficult to construct and edit generative design patterns. A third major impediment is the lack of a tool-independent representation. A common representation could lead to a shared repository to make more patterns available. We describe a new approach to generative design patterns that solves these three difficult problems. We illustrate this approach using tools called CO/sub 2/P/sub 2/S and Meta-CO/sub 2/P/sub 2/S but our approach is tool-independent.
Steve MacDonald, Duane Szafron, Jonathan Schaeffer 0001, John Anvik, Steven Bromling, Kai Tan 0005
ASE4
2002 From patterns to frameworks to parallel programs
Steve MacDonald, John Anvik, Steven Bromling, Jonathan Schaeffer 0001, Duane Szafron, Kai Tan 0005
Parallel Comput.2