VLDB 2026 Research / reviewers in the wild / expert
John Anvik
dblp:65/4421
· DBLP profile ↗
30ranked-venue papers
8as first author
8since 2021 · last 2025
0000-0002-6912-1754ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 18 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 10 · 1 first-author · 5 since 2021Systems, architecture and hardware · 5 · 2 first-authorHuman-computer interaction and ubiquitous computing · 3 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Investigating How Course Features Correlate with Student Perceptions of Two-Stage CS ExamsabstractA two-stage exam (TSE) is an exam format where students write an exam individually and rewrite the same exam in collaboration with a group of peers. This Working Group investigates TSEs in post-secondary computer science courses. By collecting data from a variety of post-secondary institutions, we aim to identify correlations between course environment indicated through the CALI inventory and students' perceptions and performance on TSEs, with consideration for students in under-represented groups. Students will be given surveys upon completion of TSEs which include open-ended questions regarding their group's dynamics and perceptions of participating in a TSE. Instructors will also provide their experiences running and observing TSEs. This research aims to understand best practices for the implementation of TSEs. Celine Latulipe, Steph McIntyre, John Anvik, Kevin Lin 0001, Sabrin Nowrin, Brian P. Railing, Scott J. Reckinger, Armita Zarnegar |
ITiCSE (2) | 3 |
| 2024 | Automatically Identifying Planning Comments in Bug Reports (P)
Shraddhaben Devaiya, John Anvik |
SEKE | 2 |
| 2024 | Comparing Machine Learning and Feature Selection Approaches for Automated Bug Report Assignment (P)
Farjana Yeasmin Omee, John Anvik |
SEKE | 2 |
| 2024 | Learning Software Engineering Principles with Program Wars v.3.0abstractProgram Wars is a web-based card game for teaching fundamental concepts of programming and cybersecurity. We propose Program Wars v.3.0, an extension to Program Wars v.2.0, to explore the effectiveness of teaching fundamental software engineering concepts using GBL. Md. Hasan Tareque, John Anvik |
SIGCSE (2) | 2 |
| 2021 | Automatically Annotating Sentences for Task-specific Bug Report SummarizationabstractThere is a need to summarize bug reports as they can become long due to many comments from conversations between developers and various DevOps tools. Although automated approaches to bug report summarization have been developed, we believe they aim at the wrong target - getting as close as possible to a gold-standard summary. Instead, researchers should create automated bug report annotation approaches that allow project members to create summaries based on their task-specific information needs. We present such an approach. Akalanka Galappaththi, John Anvik, Rafat Bin Islam |
ASE | 2 |
| 2021 | Evaluating Visual Explanation of Bug Report Assignment Recommendations (S)abstractSoftware development projects typically use an issue tracking system where the project members and users can either report faults or request additional features.Each of these reports needs to be triaged to determine such things as the priority of the report or which developers should be assigned to resolve the report.To assist a triager with report assigning, an assignment recommender has been suggested as a means of improving the process.However, proposed assignment recommenders typically present a list of developer names without an explanation of the rationale.This work presents the results of a small user study to validate our approach to visually explaining bug report assignments. Shayla Azad Bhuyan, John Anvik |
SEKE | 2 |
| 2021 | Evaluating a Tool for Creating Bug Report Assignment Recommenders (S)abstractLarge software development projects that use bug tracking systems can become overwhelmed by the number of reports filed.To assist in reducing the workload of project members, researchers have proposed the use of bug report assignment recommenders.To assist project members with the creation of assignment recommenders, we proposed a web-based tool called the Creation Assistant for Supporting Triage Recommenders (CASTR).This paper presents the results of both a laboratory and field study of CASTR.We found that CASTR can create assignment recommenders with accuracy as high as 95%, 80%, and 70% for Top-1, Top-3 and Top-5, respectively.The field study showed that 60% of the participants found CASTR easy to use, whereas the remaining participants found CASTR moderately or slightly easy to use. Disha Devaiya, John Anvik, Meher Bheree, Farjana Yeasmin Omee |
SEKE | 2 |
| 2021 | CASTR: Assisting Bug Report Assignment Recommender CreationabstractIssue tracking systems are used to make a software development process more manageable, especially for a geographically dispersed team.However, for each bug report, a decision-making process called bug report triage needs to be performed.A common bug report triage decision is the assigning of a developer to a specific bug report.Bug report triage can take significant time and resources.Bug report assignment recommenders have been proposed for reducing the workload of a project member.However, creating a recommender is complex, as project members have to perform several steps such as data preparation and selection of a machine learning algorithm.Although previous work has sought to find specific answers for aspects of the assignment recommender creation process, to the best of our knowledge, only a few other works have examined assisting with this recommender creation process.CASTR (Creation Assistant for Supporting Triage Recommenders) [1, 2] is a platform-independent multi-tier web application.It allows a project member to analyze the dataset using a graphical representation and also assists in configuring project-specific parameters for a machine learning algorithm.This demonstration shows how to use CASTR to create a bug report assignment recommender, which consists of the following steps: Disha Thakarshibhai Devaiya, John Anvik, Farjana Yeasmin Omee, Meher Bheree |
SEKE | 2 |
| 2019 | Feature Evaluation for Automatic Bug Report Summarization (S)abstractBug reports can be lengthy due to long descriptions and long conversation threads.Automatic summarization of the text in a bug report can reduce the time spent by software project members on understanding the content of a bug report.Our work further examines Rastkar et al.'s use of a logistic regression model to determine which sentences from the text of a bug report should be extracted for creating a summary.Using their publicly available bug report corpus, which contains manually annotated bug reports, we examined two aspects regarding the features used by the model.First, we examined how much of a reduction occurs in the precision and recall if some of the more complex features are not used.Second, we examined how the use of different feature combinations affects the precision and recall of the models.We found that the absence of some of the complex features resulted in a modest decrease in precision and recall, and confirmed that some features, such as sentence length, were the most significant features for bug report summarization. Akalanka Galappaththi, John Anvik |
SEKE | 2 |
| 2019 | Program Wars: A Card Game for Learning Programming and Cybersecurity ConceptsabstractAlthough there are many computer science learning games with the goal of teaching programming, such games typically require the person to either learn an existing programming language or the game's own specialized language. This can be intimidating, confusing or frustrating for an individual when they cannot get their "program" to work correctly (e.g. syntax error, infinite loop). Additionally, such games commonly use a puzzle-solving approach that does not appeal to some demographics. This paper presents a programming-language-independent approach to teaching fundamental programming and cybersecurity concepts using simple vocabulary. This approach also uses the familiar activity of playing cards against opponents to create a more dynamic and engaging learning experience. The approach is demonstrated by a web-based game called Program Wars. Results from a user study show that players are able to effectively connect game concepts to actual programming language structures; however, whether players' comprehension of computer programming is improved is unclear. John Anvik, Vincent Côté, Jace Riehl |
SIGCSE | 1 |
| 2017 | Parallel Implementation of a Bug Report Assignment Recommender Using Deep Learning
Adrian-Catalin Florea, John Anvik, Razvan Andonie |
ICANN (2) | 2 |
| 2016 | A feature location approach supported by time-aware weighting of terms associated with developer expertise profiles
Sima Zamani, Sai Peck Lee, Ramin Shokripour, John Anvik |
Knowl. Inf. Syst. | 4 |
| 2015 | Using Time Series Models for Defect Prediction in Software Release PlanningabstractA time series model is presented that uses historical project information to predict the number of future defects, given the number of proposed features and improvements to be completed.This allows for hypothetical release plans to be compared by assessing their predicted impact on testing and defect-fixing time.We selected the VARX time series model as a reasonable approach.The accuracy of the model appeared low for a single dataset, but the error was found to be normally distributed. James Tunnell, John Anvik |
SEKE | 2 |
| 2015 | A time-based approach to automatic bug report assignment
Ramin Shokripour, John Anvik, Zarinah Mohd Kasirun, Sima Zamani |
J. Syst. Softw. | 2 |
| 2014 | Assisting Software Projects with Assignment Recomender Creation
John Anvik, Marshall Brooks, Henry Burton, Justin Canada |
SEKE | 1 |
| 2014 | A noun-based approach to feature location using time-aware term-weighting
Sima Zamani, Sai Peck Lee, Ramin Shokripour, John Anvik |
Inf. Softw. Technol. | 4 |
| 2013 | Why so complicated? simple term filtering and weighting for location-based bug report assignment recommendationabstractLarge software development projects receive many bug reports and each of these reports needs to be triaged. An important step in the triage process is the assignment of the report to a developer. Most previous efforts towards improving bug report assignment have focused on using an activity-based approach. We address some of the limitations of activity-based approaches by proposing a two-phased location-based approach where bug report assignment recommendations are based on the predicted location of the bug. The proposed approach utilizes a noun extraction process on several information sources to determine bug location information and a simple term weighting scheme to provide a bug report assignment recommendation. We found that by using a location-based approach, we achieved an accuracy of 89.41% and 59.76% when recommending five developers for the Eclipse and Mozilla projects, respectively. Ramin Shokripour, John Anvik, Zarinah Mohd Kasirun, Sima Zamani |
MSR | 2 |
| 2011 | Reducing the effort of bug report triage: Recommenders for development-oriented decisionsabstractA key collaborative hub for many software development projects is the bug report repository. Although its use can improve the software development process in a number of ways, reports added to the repository need to be triaged. A triager determines if a report is meaningful. Meaningful reports are then organized for integration into the project's development process. To assist triagers with their work, this article presents a machine learning approach to create recommenders that assist with a variety of decisions aimed at streamlining the development process. The recommenders created with this approach are accurate; for instance, recommenders for which developer to assign a report that we have created using this approach have a precision between 70% and 98% over five open source projects. As the configuration of a recommender for a particular project can require substantial effort and be time consuming, we also present an approach to assist the configuration of such recommenders that significantly lowers the cost of putting a recommender in place for a project. We show that recommenders for which developer should fix a bug can be quickly configured with this approach and that the configured recommenders are within 15% precision of hand-tuned developer recommenders. John Anvik, Gail C. Murphy |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2008 | An approach to detecting duplicate bug reports using natural language and execution informationabstractAn open source project typically maintains an open bug repository so that bug reports from all over the world can be gathered. When a new bug report is submitted to the repository, a person, called a triager, examines whether it is a duplicate of an existing bug report. If it is, the triager marks it as DUPLICATE and the bug report is removed from consideration for further work. In the literature, there are approaches exploiting only natural language information to detect duplicate bug reports. In this paper we present a new approach that further involves execution information. In our approach, when a new bug report arrives, its natural language information and execution information are compared with those of the existing bug reports. Then, a small number of existing bug reports are suggested to the triager as the most similar bug reports to the new bug report. Finally, the triager examines the suggested bug reports to determine whether the new bug report duplicates an existing bug report. We calibrated our approach on a subset of the Eclipse bug repository and evaluated our approach on a subset of the Firefox bug repository. The experimental results show that our approach can detect 67%-93% of duplicate bug reports in the Firefox bug repository, compared to 43%-72% using natural language information alone. Xiaoyin Wang, Lu Zhang 0023, Tao Xie 0001, John Anvik, Jiasu Sun |
ICSE | 4 |
| 2008 | Task articulation in software maintenance: Integrating source code annotations with an issue tracking systemabstractManaging and articulating development tasks is an important aspect of software maintenance. Developers already have a variety of specialty tools to support task management, for example, issue tracking and configuration management software, but they also make use of other tools within their software engineering environments to support software task management. An example of one such mechanism is the appropriation of source code comments to document finer grained details of formally specified tasks. In this research, we propose and present a tool that integrates these source code annotations with an issue tracking management system. We describe how this tool addresses deficiencies that occur in task management and propose future research to improve task management. John Anvik, Margaret-Anne D. Storey |
ICSM | 1 |
| 2006 | Visual Explanation of Evidence with Additive Classifiers
Brett Poulin, Roman Eisner, Duane Szafron, Paul Lu, Russell Greiner, David S. Wishart, Alona Fyshe, Brandon Pearcy, John Anvik |
AAAI | 10 |
| 2006 | Automating bug report assignmentabstractOpen-source development projects typically support an open bug repository to which both developers and users can report bugs. A report that appears in this repository must be triaged to determine if the report is one which requires attention and if it is, which developer will be assigned the responsibility of resolving the report. Large open-source developments are burdened by the rate at which new bug reports appear in the bug repository. The thesis of this work is that the task of triage can be eased by using a semi-automated approach to assign bug reports to developers. The approach consists of constructing a recommender for bug assignments; examined are both a range of algorithms that can be used and the various kinds of information provided to the algorithms. The proposed work seeks to determine through human experimentation a sufficient level of precision for the recommendations, and to analytically determine the trade-offs of the various algorithmic and information choices. John Anvik |
ICSE | 1 |
| 2006 | Who should fix this bug?abstractOpen source development projects typically support an open bug repository to which both developers and users can report bugs. The reports that appear in this repository must be triaged to determine if the report is one which requires attention and if it is, which developer will be assigned the responsibility of resolving the report. Large open source developments are burdened by the rate at which new bug reports appear in the bug repository. In this paper, we present a semi-automated approach intended to ease one part of this process, the assignment of reports to a developer. Our approach applies a machine learning algorithm to the open bug repository to learn the kinds of reports each developer resolves. When a new report arrives, the classifier produced by the machine learning technique suggests a small number of developers suitable to resolve the report. With this approach, we have reached precision levels of 57% and 64% on the Eclipse and Firefox development projects respectively. We have also applied our approach to the gcc open source development with less positive results. We describe the conditions under which the approach is applicable and also report on the lessons we learned about applying machine learning to repositories used in open source development. John Anvik, Lyndon Hiew, Gail C. Murphy |
ICSE | 1 |
| 2005 | Asserting the utility of CO2P3S using the Cowichan Problem Set
John Anvik, Jonathan Schaeffer 0001, Duane Szafron, Kai Tan 0005 |
J. Parallel Distributed Comput. | 1 |
| 2004 | Predicting subcellular localization of proteins using machine-learned classifiersabstractMOTIVATION: Identifying the destination or localization of proteins is key to understanding their function and facilitating their purification. A number of existing computational prediction methods are based on sequence analysis. However, these methods are limited in scope, accuracy and most particularly breadth of coverage. Rather than using sequence information alone, we have explored the use of database text annotations from homologs and machine learning to substantially improve the prediction of subcellular location. RESULTS: We have constructed five machine-learning classifiers for predicting subcellular localization of proteins from animals, plants, fungi, Gram-negative bacteria and Gram-positive bacteria, which are 81% accurate for fungi and 92-94% accurate for the other four categories. These are the most accurate subcellular predictors across the widest set of organisms ever published. Our predictors are part of the Proteome Analyst web-service. Zhiyong Lu, Duane Szafron, Russell Greiner, Paul Lu, David S. Wishart, Brett Poulin, John Anvik, Roman Eisner |
Bioinform. | 7 |
| 2003 | Why Not Use a Pattern-Based Parallel Programming System?
John Anvik, Jonathan Schaeffer 0001, Duane Szafron, Kai Tan 0005 |
Euro-Par | 1 |
| 2003 | Using generative design patterns to generate parallel code for a distributed memory environmentabstractA design pattern is a mechanism for encapsulating the knowledge of experienced designers into a re-usable artifact. Parallel design patterns reflect commonly occurring parallel communication and synchronization structures. Our tools, CO2P3S (Correct Object-Oriented Pattern-based Parallel Programming System) and MetaCO2P3S, use generative design patterns. A programmer selects the parallel design patterns that are appropriate for an application, and then adapts the patterns for that specific application by selecting from a small set of code-configuration options. CO2P3S then generates a custom framework for the application that includes all of the structural code necessary for the application to run in parallel. The programmer is only required to write simple code that launches the application and to fill in some application-specific sequential hook routines. We use generative design patterns to take an application specification (parallel design patterns + sequential user code) and use it to generate parallel application code that achieves good performance in shared memory and distributed memory environments. Although our implementations are for Java, the approach we describe is tool and language independent. This paper describes generalizing CO2P3S to generate distributed-memory parallel solutions. Kai Tan 0005, Duane Szafron, Jonathan Schaeffer 0001, John Anvik, Steve MacDonald |
PPoPP | 4 |
| 2002 | Pattern-Based Parallel ProgrammingabstractThe advantages of pattern-based programming have been well-documented in the sequential programming literature. However patterns have yet to make their way into mainstream parallel computing, even though several research tools support them. There are two critical shortcomings of pattern (or template) based systems for parallel programming: lack of extensibility and performance. This paper describes our approach for addressing these problems in the CO/sub 2/P/sub 3/S parallel programming system. CO/sub 2/P/sub 3/S supports multiple levels of abstraction, allowing the user to design an application with high-level patterns, but move to lower levels of abstraction for performance tuning. Patterns are implemented as parameterized templates, allowing the user the ability to customize the pattern to meet their needs. CO/sub 2/P/sub 3/S generates code that is specific to the pattern/parameter combination selected by the user. The MetaCO/sub 2/P/sub 3/S tool addresses extensibility by giving users the ability to design and add new pattern templates to CO/sub 2/P/sub 3/S. Since the pattern templates are stored in a system-independent format, they are suitable for storing in a repository to be shared throughout the user community. Steven Bromling, Steve MacDonald, John Anvik, Jonathan Schaeffer 0001, Duane Szafron, Kai Tan 0005 |
ICPP | 3 |
| 2002 | Generative Design PatternsabstractA design pattern encapsulates the knowledge of object-oriented designers into re-usable artifacts. A design pattern is a descriptive device that fosters software design re-use. There are several reasons why design patterns are not used as generative constructs that support code re-use. The first reason is that design patterns describe a set of solutions to a family of related design problems and it is difficult to generate a single body of code that adequately solves each problem in the family. A second reason is that it is difficult to construct and edit generative design patterns. A third major impediment is the lack of a tool-independent representation. A common representation could lead to a shared repository to make more patterns available. We describe a new approach to generative design patterns that solves these three difficult problems. We illustrate this approach using tools called CO/sub 2/P/sub 2/S and Meta-CO/sub 2/P/sub 2/S but our approach is tool-independent. Steve MacDonald, Duane Szafron, Jonathan Schaeffer 0001, John Anvik, Steven Bromling, Kai Tan 0005 |
ASE | 4 |
| 2002 | From patterns to frameworks to parallel programs
Steve MacDonald, John Anvik, Steven Bromling, Jonathan Schaeffer 0001, Duane Szafron, Kai Tan 0005 |
Parallel Comput. | 2 |