VLDB 2026 Research / reviewers in the wild / expert
Jeremy S. Bradbury
dblp:42/932
· DBLP profile ↗
17ranked-venue papers
3as first author
8since 2021 · last 2025
0000-0002-5204-908XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 12 · 3 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 6 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Contrastive Learning Approach to Bug Severity Classification with Large Language Model EmbeddingsabstractAutomatically classifying bug severity helps reduce manual effort and improve response times in software maintenance. This study leverages Large Language Models (LLMs), specifically CodeBERT, to generate contextual embeddings of bug reports for automated severity classification. To enhance the quality of embeddings, we integrate Contrastive Learning, which structures the embedding space by bringing similar bug reports closer and pushing dissimilar reports apart. Our approach is evaluated on the NASA PITS and Mozilla datasets and compared against LLMs fine-tuned without contrastive learning as well as traditional embedding models like Doc2Vec. Our results demonstrate that contrastive learning consistently enhances performance, particularly on imbalanced and diverse datasets. Furthermore, the results demonstrate that LLMs excel in handling longer bug descriptions, while traditional embedding models like Doc2Vec are suitable for smaller, structured datasets. Mosarrat Rumman, Emon Roy, Anushka Zaman, Jeremy S. Bradbury |
COMPSAC | 4 |
| 2025 | Addressing Data Leakage in HumanEval Using Combinatorial Test DesignabstractThe use of large language models (LLMs) is widespread across many domains, including Software Engineering, where they have been used to automate tasks such as program generation and test classification. As LLM-based methods continue to evolve, it is important that we define clear and robust methods that fairly evaluate performance. Benchmarks are a common approach to assess LLMs with respect to their ability to solve problem-specific tasks as well as assess different versions of an LLM to solve tasks over time. For example, the HumanEval benchmark is composed of 164 hand-crafted tasks and has become an important tool in assessing LLM-based program generation. However, a major barrier to a fair evaluation of LLMs using benchmarks like HumanEval is data contamination resulting from data leakage of benchmark tasks and solutions into the training data set. This barrier is compounded by the black-box nature of LLM training data which makes it difficult to even know if data leakage has occurred. To address the data leakage problem, we propose a new benchmark construction method where a benchmark is composed of template tasks that can be instantiated into new concrete tasks using combinatorial test design. Concrete tasks for the same template task must be different enough that data leakage has minimal impact and similar enough that the tasks are interchangeable with respect to performance evaluation. To assess our benchmark construction method, we propose HumanEval_T, an alternative benchmark to HumanEval that was constructed using template tasks and combinatorial test design. Jeremy S. Bradbury, Riddhi More |
ICST | 1 |
| 2025 | An Analysis of LLM Fine-Tuning and Few-Shot Learning for Flaky Test Detection and ClassificationabstractFlaky tests exhibit non-deterministic behavior during execution and they may pass or fail without any changes to the program under test. Detecting and classifying these flaky tests is crucial for maintaining the robustness of automated test suites and ensuring the overall reliability and confidence in the testing. However, flaky test detection and classification is challenging due to the variability in test behavior, which can depend on environmental conditions and subtle code interactions. Large Language Models (LLMs) offer promising approaches to address this challenge, with fine-tuning and few-shot learning (FSL) emerging as viable techniques. With enough data fine-tuning a pre-trained LLM can achieve high accuracy, making it suitable for organizations with more resources. Alternatively, we introduce FlakyXbert, an FSL approach that employs a Siamese network architecture to train efficiently with limited data. To understand the performance and cost differences between these two methods, we compare fine-tuning on larger datasets with FSL in scenarios restricted by smaller datasets. Our evaluation involves two existing flaky test datasets, FlakyCat and IDoFT. Our results suggest that while fine-tuning can achieve high accuracy, FSL provides a cost-effective approach with competitive accuracy, which is especially beneficial for organizations or projects with limited historical data available for training. These findings underscore the viability of both fine-tuning and FSL in flaky test detection and classification with each suited to different organizational needs and resource availability. Riddhi More, Jeremy S. Bradbury |
ICST | 2 |
| 2025 | How Effective and Efficient are Student-Written Software Tests?abstractMany computer science students complete their undergraduate degrees with insufficient testing skills and knowledge. To understand the gaps in students' testing skills and knowledge, we analyzed 1014 software tests written by 12 groups in an undergraduate Software Quality Assurance (SQA) course project. In the project the student groups were provided a requirements document and were instructed to follow Test Driven Development (TDD) practices using black-box tests. To understand how the groups applied black-box testing in their project, we created an automatic tool to sort the tests into categories or "test buckets." By analyzing the test bucket data, we were able to assess the effectiveness and efficiency of student-written tests. We observed that the student groups were significantly more likely to test for explicit requirements than implicit requirements and significantly more likely to test happy paths than invalid inputs. Furthermore, students inefficiently tested happy paths, invalid inputs and explicit requirements resulting in a higher proportion of software tests with duplicate intent. Based on these results we provide insights into how black-box test education can be improved. Amanda Showler, Michael A. Miljanovic, Jeremy S. Bradbury |
SIGCSE (1) | 3 |
| 2024 | PIE: A Tool for Visualizing the Life Cycle of Design Patterns in Open Source Software ProjectsabstractDesign patterns are employed in source code to solve commonly occurring programming tasks using understood best practices. Object-oriented design patterns usually span multiple classes and objects and play an integral role in the way object-oriented software is built. One challenge with using object-oriented design patterns is that over the life of a project, these patterns can undergo both planned and unplanned changes. Unplanned changes are often the result of bug fixes or code maintenance tasks that modify a design pattern as a side effect. Furthermore, these unplanned changes can result in increased brittleness of the code and can compromise the overall stability of the software. Over the lifetime of a software project, developers may only become aware of these unplanned changes when the code brittleness results in a software bug. To improve developers' understanding of object-oriented design pattern evolution, we introduce the design Pattern Instance Explorer (PIE) _ an exploratory visualization tool that enable developers to visualize a git repository's object-oriented design patterns and their life cycles. In addition to discussing the PIE tool, we provide examples of how this tool can be used to identify and understand design pattern changes. Tool demonstration video: https://www.youtube.com/watch?v=Gkn_5q8_Awg Christopher Collins 0001, Jeremy S. Bradbury |
VISSOFT | 3 |
| 2023 | Adapting Between Parsons Problems and Coding TasksabstractPrevious research has shown that Parsons problems are an effective scaffolding activity for coding. Recently the development of Adaptive Parsons problems has provided more flexible scaffolding for students learning to code. However, there is still a gap between Parsons problems and coding tasks which can both challenge and frustrate novices. As such, we have developed an adaptive learning tool which looks to bridge this gap and support transitioning directly between Parsons problems and code writing. Nadia L. Goralski, Jeremy S. Bradbury |
SIGCSE (2) | 2 |
| 2023 | Run, Llama, Run: A Computational Thinking Game for K-5 Students Designed to Support Equitable AccessabstractComputational thinking is now included in K-5 classrooms and this has led to a demand for new interactive and collaborative learning tools that engage a younger audience. Block-based programming and educational games have both been shown to be effective at engaging children, however they have limitations with respect to supporting collaborative learning and equitable access. Our goal in designing Run, Llama, Run was to build on the positive aspects of block-based programming and educational games while also addressing these limitations. Furthermore, we are using Run, Llama, Run as a platform to explore the trade-offs between digital and tangible interfaces to understand how best to support equitable access while maintaining learning, engagement, and collaboration. Stacey A. Koornneef, Jeremy S. Bradbury, Michael A. Miljanovic |
SIGCSE (2) | 2 |
| 2022 | Run, Llama, Run: A Collaborative Physical and Online Coding Game for Children
Stacey A. Koornneef, Jeremy S. Bradbury, Michael A. Miljanovic |
SIGCSE (2) | 2 |
| 2017 | RoboBUG: A Serious Game for Learning Debugging TechniquesabstractDebugging is an essential but challenging task that can present a great deal of confusion and frustration to novice programmers. It can be argued that Computer Science education does not sufficiently address the challenges that students face when identifying bugs in their programs. To help students learn effective debugging techniques and to provide students a more enjoyable and motivating experience, we have designed the RoboBUG game. RoboBUG is a serious game that can be customized with respect to different programming languages and game levels. Michael A. Miljanovic, Jeremy S. Bradbury |
ICER | 2 |
| 2013 | Special section on Mutation testing (Mutation 2010)
Lydie du Bousquet, Jeremy S. Bradbury, Gordon Fraser 0001 |
Sci. Comput. Program. | 2 |
| 2011 | Implementing and Evaluating a Runtime Conformance Checker for Mobile Agent SystemsabstractA Mobile Agent System (MAS) is a special kind of distributed system in which the agent software can move from one physical host to another. This paper describes a new approach, together with its implementation and evaluation, for checking the conformance of a MAS with respect to an executable model. In order to check the effectiveness of our conformance check, we have built a mutation-based evaluation framework. Part of the framework is a set of 29 new mutation operators for mobile agent systems. Our conformance checking approach is used to compare the mutated agents with the executable model and determine nonconformance. Our experimental results suggest that our approach holds promise for the generation and detection of non-equivalent mutants. Ahmad A. Saifan, Jürgen Dingel, Jeremy S. Bradbury, Ernesto Posse |
ICST | 3 |
| 2011 | Guest Editorial for Special Section on Mutation Testing
Benoit Baudry, Jeremy S. Bradbury, Gordon Fraser 0001 |
Inf. Softw. Technol. | 2 |
| 2010 | Using clone detection to identify bugs in concurrent softwareabstractIn this paper we propose an active testing approach that uses clone detection and rule evaluation as the foundation for detecting bug patterns in concurrent software. If we can identify a bug pattern as being present then we can localize our testing effort to the exploration of interleavings relevant to the potential bug. Furthermore, if the potential bug is indeed a real bug, then targeting specific thread interleavings instead of examining all possible executions can increase the probability of the bug being detected sooner. Kevin Jalbert, Jeremy S. Bradbury |
ICSM | 2 |
| 2010 | How Good is Static Analysis at Finding Concurrency Bugs?abstractDetecting bugs in concurrent software is challenging due to the many different thread interleavings. Dynamic analysis and testing solutions to bug detection are often costly as they need to provide coverage of the interleaving space in addition to traditional black box or white box coverage. An alternative to dynamic analysis detection of concurrency bugs is the use of static analysis. This paper examines the use of three static analysis tools (Find Bugs, J Lint and Chord) in order to assess each tool's ability to find concurrency bugs and to identify the percentage of spurious results produced. The empirical data presented is based on an experiment involving 12 concurrent Java programs. Devin Kester, Martin Mwebesa, Jeremy S. Bradbury |
SCAM | 3 |
| 2006 | Using source transformation to test and model check implicit-invocation systems
Jeremy S. Bradbury, James R. Cordy, Jürgen Dingel |
Sci. Comput. Program. | 2 |
| 2005 | An empirical framework for comparing effectiveness of testing and property-based formal analysisabstractToday, many formal analysis tools are not only used to provide certainty but are also used to debug software systems - a role that has traditional been reserved for testing tools. We are interested in exploring the complementary relationship as well as tradeoffs between testing and formal analysis with respect to debugging and more specifically bug detection. In this paper we present an approach to the assessment of testing and formal analysis tools using metrics to measure the quantity and efficiency of each technique at finding bugs. We also present an assessment framework that has been constructed to allow for symmetrical comparison and evaluation of tests versus properties. We are currently beginning to conduct experiments and this paper presents a discussion of possible outcomes of our proposed empirical study. Jeremy S. Bradbury, James R. Cordy, Jürgen Dingel |
PASTE | 1 |
| 2003 | Evaluating and improving the automatic analysis of implicit invocation systemsabstractModel checking and other finite-state analysis techniques have been very successful when used with hardware systems and less successful with software systems. It is especially difficult to analyze software systems developed with the implicit invocation architectural style because the loose coupling of their components increases the size of the finite state model. In this paper we provide insight into the larger problem of how to make model checking a better analysis and verification tool for software systems. Specifically, we will extend an existing approach to model checking implicit invocation to allow for the modeling of larger and more realistic systems. Our focus will be on improving the representation of events, event delivery policies and event-method bindings. We also evaluate our technique on two non-trivial examples. In one of our examples, we will show how with iterative analysis a system parameter can be chosen to meet the appropriate system requirements. Jeremy S. Bradbury, Jürgen Dingel |
ESEC / SIGSOFT FSE | 1 |