Scott Barnett

dblp:165/5477 · DBLP profile ↗
← Back
20ranked-venue papers
4as first author
13since 2021 · last 2025
0000-0002-3187-4937ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 18 · 3 first-author · 12 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 RAGProbe: Breaking RAG Pipelines with Evaluation Scenarios
abstract
Retrieval Augmented Generation (RAG) is increasingly employed in building Generative AI applications, yet their evaluation often relies on manual, trial-and-error processes. Automating this evaluation process involves generating test data to trigger failures involving context comprehension, data formatting, specificity, and content completeness. Random question-answer generation is insufficient. However, prior works rely on standard QA datasets, benchmarks and tactics that are not tailored to the specific domain requirements. Hence, current approaches and datasets do not trigger sufficiently broad and context-specific failures. In this paper, we introduce evaluation scenarios that describe the process of generating question-answer pairs from content indexed by RAG pipelines, and they are designed to trigger a wider range of failures and to simplify automation. This enables developers to identify and address weaknesses more effectively. We validate our approach on five open-source RAG pipelines using three datasets. Our approach triggers high failure rates, by generating prompts that combine multiple questions (up to 91% failure rate) highlighting the need for developers to prioritize handling such queries. We generated failure rates of 60% in an academic domain dataset and 53% and 64% in open-domain datasets. Compared to existing state-of-the-art methods, our approach triggers 77% more failures on average per RAG pipeline and 53% more failures on average per dataset, offering a mechanism to support developers to improve the RAG pipeline quality.
Shangeetha Sivasothy, Scott Barnett, Stefanus Kurniawan, Zafaryab Rasool, Rajesh Vasa
CAIN2
2024 ML-On-Rails: Safeguarding Machine Learning Models in Software Systems - A Case Study
abstract
Machine learning (ML), especially with the emergence of large language models (LLMs), has significantly transformed various industries. However, the transition from ML model prototyping to production use within software systems presents several challenges. These challenges primarily revolve around ensuring safety, security, and transparency, subsequently influencing the overall robustness and trustworthiness of ML models. In this paper, we introduce ML-On-Rails, a protocol designed to safeguard ML models, establish a well-defined endpoint interface for different ML tasks, and clear communication between ML providers and ML consumers (software engineers). ML-On-Rails enhances the robustness of ML models via incorporating detection capabilities to identify unique challenges specific to production ML. We evaluated the ML-On-Rails protocol through a real-world case study of the MoveReminder application. Through this evaluation, we emphasize the importance of safeguarding ML models in production.
Hala Abdelkader, Mohamed Almorsy, Scott Barnett, Jean-Guy Schneider, Priya Rani, Rajesh Vasa
CAIN3
2024 Seven Failure Points When Engineering a Retrieval Augmented Generation System
abstract
Software engineers are increasingly adding semantic search capabilities to applications using a strategy known as Retrieval Augmented Generation (RAG). A RAG system involves finding documents that semantically match a query and then passing the documents to a large language model (LLM) such as ChatGPT to extract the right answer using an LLM. RAG systems aim to: a) reduce the problem of hallucinated responses from LLMs, b) link sources/references to generated responses, and c) remove the need for annotating documents with meta-data. However, RAG systems suffer from limitations inherent to information retrieval systems and from reliance on LLMs. In this paper, we present an experience report on the failure points of RAG systems from three case studies from separate domains: research, education, and biomedical. We share the lessons learned and present 7 failure points to consider when designing a RAG system. The two key takeaways arising from our work are: 1) validation of a RAG system is only feasible during operation, and 2) the robustness of a RAG system evolves rather than designed in at the start. We conclude with a list of potential research directions on RAG systems for the software engineering community.
Scott Barnett, Stefanus Kurniawan, Srikanth Thudumu, Zach Brannelly, Mohamed Almorsy
CAIN1
2024 Green Runner: A Tool for Efficient Deep Learning Component Selection
abstract
For software that relies on machine-learned functionality, model selection is key to finding the right model for the task with desired performance characteristics. Evaluating a model requires developers to i) select from many models (e.g. the Hugging face model repository), ii) select evaluation metrics and training strategy, and iii) tailor trade-offs based on the problem domain. However, current evaluation approaches are either ad-hoc resulting in sub-optimal model selection or brute force leading to wasted compute. In this work, we present GreenRunner, a novel tool to automatically select and evaluate models based on the application scenario provided in natural language. We leverage the reasoning capabilities of large language models to propose a training strategy and extract desired trade-offs from a problem description. GreenRunner features a resource-efficient experimentation engine that integrates constraints and trade-offs based on the problem into the model selection process. Our preliminary evaluation demonstrates that GreenRunner is both efficient and accurate compared to ad-hoc evaluations and brute force. This work presents an important step toward energy-efficient tools to help reduce the environmental impact caused by the growing demand for software with machine-learned functionality. Our tool is available at Figshare GreenRunner.
Jai Kannan, Scott Barnett, Anj Simmons, Taylan Selvi, Luis Cruz 0002
CAIN2
2024 LLMs for Test Input Generation for Semantic Applications
abstract
Large language models (LLMs) enable state-of-the-art semantic capabilities to be added to software systems such as semantic search of unstructured documents and text generation. However, these models are computationally expensive. At scale, the cost of serving thousands of users increases massively affecting also user experience. To address this problem, semantic caches are used to check for answers to similar queries (that may have been phrased differently) without hitting the LLM service. Due to the nature of these semantic cache techniques that rely on query embeddings, there is a high chance of errors impacting user confidence in the system. Adopting semantic cache techniques usually requires testing the effectiveness of a semantic cache (accurate cache hits and misses) which requires a labelled test set of similar queries and responses which is often unavailable. In this paper, we present VaryGen, an approach for using LLMs for test input generation that produces similar questions from unstructured text documents. Our novel approach uses the reasoning capabilities of LLMs to 1) adapt queries to the domain, 2) synthesise subtle variations to queries, and 3) evaluate the synthesised test dataset. We evaluated our approach in the domain of a student question and answer system by qualitatively analysing 100 generated queries and result pairs, and conducting an empirical case study with an open source semantic cache. Our results show that query pairs satisfy human expectations of similarity and our generated data demonstrates failure cases of a semantic cache. Additionally, we also evaluate our approach on Qasper dataset. This work is an important first step into test input generation for semantic applications and presents considerations for practitioners when calibrating a semantic cache.
Zafaryab Rasool, Scott Barnett, David Willie, Stefanus Kurniawan, Sherwin Balugo, Srikanth Thudumu, Mohamed Almorsy
CAIN2
2024 Comparative analysis of real issues in open-source machine learning projects
abstract
Abstract Context In the last decade of data-driven decision-making, Machine Learning (ML) systems reign supreme. Because of the different characteristics between ML and traditional Software Engineering systems, we do not know to what extent the issue-reporting needs are different, and to what extent these differences impact the issue resolution process. Objective We aim to compare the differences between ML and non-ML issues in open-source applied AI projects in terms of resolution time and size of fix. This research aims to enhance the predictability of maintenance tasks by providing valuable insights for issue reporting and task scheduling activities. Method We collect issue reports from Github repositories of open-source ML projects using an automatic approach, filter them using ML keywords and libraries, manually categorize them using an adapted deep learning bug taxonomy, and compare resolution time and fix size for ML and non-ML issues in a controlled sample. Result 147 ML issues and 147 non-ML issues are collected for analysis. We found that ML issues take more time to resolve than non-ML issues, the median difference is 14 days. There is no significant difference in terms of size of fix between ML and non-ML issues. No significant differences are found between different ML issue categories in terms of resolution time and size of fix. Conclusion Our study provided evidence that the life cycle for ML issues is stretched, and thus further work is required to identify the reason. The results also highlighted the need for future work to design custom tooling to support faster resolution of ML issues.
Tuan Dung Lai, Anj Simmons, Scott Barnett, Jean-Guy Schneider, Rajesh Vasa
Empir. Softw. Eng.3
2023 Decentralized Federated Learning Strategy with Image Classification using ResNet Architecture
abstract
The rapid growth of both the Industrial Internet of Things (IIoT) and Artificial Intelligence (AI) results in a high demand for AI applications in devices. To achieve high levels of accuracy, AI applications typically require a large amount of annotated data. Accessing such data is challenging in various applications such as healthcare, finance and information security. Federated learning (FL) is one of the strategies that was proposed to overcome this challenge. Specifically, FL enables the AI model in the centralized system to be trained without any prior knowledge of the information on the devices. Recent FLs have the disadvantage that they are dependent upon a centralized system, and thus are susceptible to single points of failure. This paper proposes a strategy that employs FL in a decentralized environment where devices can communicate with each other to increase the accuracy of the AI model in each device. Furthermore, we evaluate the proposed strategy in the image classification task with the ResNet50 architecture and the CIFAR-10 dataset. The evaluation shows that the ResNet50 model trained in the decentralized environment can achieve comparable results to the model trained in the centralized environment.
Hung Du, Srikanth Thudumu, Sankhya Singh, Scott Barnett, Irini Logothetis, Rajesh Vasa, Kon Mouzakis
CCNC4
2023 mHealthSwarm: A Unified Platform for mHealth Applications
abstract
Mobile health (mHealth) applications are ubiquitous and offer several benefits such as easier access to one's health and wellness data through smartphones. However, their growth in popularity has also introduced several challenges for both end-users and developers. Users face challenges around the poor user experience (UX) introduced by the need to install several apps and limited customizability which is exacerbated by the limited control over app functionality. While features common across different apps may not directly affect individual developers, the fact that one app may provide a better implementation of features leads to wasted developer effort. A single platform that can satisfy all user needs does not currently exist, and given the diversity of the mHealth domain, a single app would be far too complex if it existed. In this paper, we present a new approach and a platform for mHealth apps - micro-mHealth apps, and we discuss them as an alternative to the current mHealth app development model, which we believe could improve the current state of mHealth app development and adoption. We are currently evaluating our prototype with several micro-mHealth apps built using common features in the mHealth apps available in commercial app stores.
Ben Joseph Philip, Yasmeen Anjeer Alshehhi, Mohamed Almorsy, Scott Barnett, Alessio Bonti, John C. Grundy
ENASE4
2023 Cell2Doc: ML Pipeline for Generating Documentation in Computational Notebooks
abstract
Computational notebooks have become the go-to way for solving data-science problems. While they are designed to combine code and documentation, prior work shows that documentation is largely ignored by the developers because of the manual effort. Automated documentation generation can help, but existing techniques fail to capture algorithmic details and developers often end up editing the generated text to provide more explanation and sub-steps. This paper proposes a novel machine-learning pipeline, Cell2Doc, for code cell documentation in Python data science notebooks. Our approach works by identifying different logical contexts within a code cell, generating documentation for them separately, and finally combining them to arrive at the documentation for the entire code cell. Cell2Doc takes advantage of the capabilities of existing pre-trained language models and improves their efficiency for code cell documentation. We also provide a new benchmark dataset for this task, along with a data-preprocessing pipeline that can be used to create new datasets. We also investigate an appropriate input representation for this task. Our automated evaluation suggests that our best input representation improves the pre-trained model's performance by 2.5x on average. Further, Cell2Doc achieves 1.33x improvement during human evaluation in terms of correctness, informativeness, and readability against the corresponding standalone pretrained model.
Tamal Mondal, Scott Barnett, Akash Lal, Jyothi Vedurada
ASE2
2022 A Framework for Evaluating MRC Approaches with Unanswerable Questions
abstract
Machine reading comprehension (MRC) is a challenging task in natural language processing that demonstrates the language understanding of the machine. An approach to tackle this challenge requires the machine to answer the question about the given context when needed and abstain from answering when there is no answer. Recent works attempted to solve this challenge with various comprehensive neural network architectures for sequences such as SAN, U-Net, EQuANt, and others that were trained on the SQuAD 2.0 dataset containing unanswerable questions. However, the robustness of these approaches has not been evaluated. In this paper, we propose a data augmentation approach that converts answerable questions to unanswerable questions in the SQuAD 2.0 dataset by altering the entities in the question to its antonym from ConceptNet which is a semantic network. The augmented data is, then, fitted into the U-Net question answering model to evaluate the robustness of the model.
Hung Du, Srikanth Thudumu, Sankhya Singh, Scott Barnett, Irini Logothetis, Rajesh Vasa, Kon Mouzakis
e-Science4
2022 PiMS: A Pre-ML Labelling Tool
abstract
Machine Learning (ML) techniques in clinical decision support systems are scarce due to the limited availability of clinically validated and labelled training data sets. We present a framework to (1) enable quality controls at data submission toward ML appropriate data, (2) provide in-situ algorithm assessments, and (3) prepare dataframes for ML training and robust stochastic analysis. We developed and evaluated PiMS (Pandemic Intervention and Monitoring Systems): a remote monitoring solution for patients that are Covid-positive. The system was trialled at two hospitals in Melbourne, Australia (Alfred Health and Monash Health) involving 109 patients and 15 clinicians.
Irini Logothetis, Scott Barnett, Leonard Hoon, Srikanth Thudumu, Joseph Mathew, Carl Luckhoff, Gerard O'Reilly, David Collard, Rajesh Vasa, Kon Mouzakis, Mark Fitzgerald
e-Science2
2022 Subspace based Anomaly Detection Framework for Point Clouds
abstract
In many real-world applications such as the inspection of powerlines, the automated detection of anomalies can minimise damage and reduce costs that result from the presence of unknown anomalies. Technologies such as LiDAR scans obtained from Unmanned Aerial Vehicles (UAV) are becoming prominent due to the data depth they provide. In the context of powerline transmission, investigators must search for anomalous elements such as line defects or obstructions. Such occurrences are not always apparent and detecting them requires extensive analysis of data within vast areas of wilderness. Automating this process can reduce time and labor costs. We propose a methodology to define what constitutes an anomaly within mapped real-world scenes, and a technique to address different types of anomalies. The notion of unknowns and knowns composed of unknown to both human and machine, known to human and unknown to machine, unknown to human and known to machine, and known to both human and machine is considered to develop a novel framework that detects anomalous patterns. For the purpose of evaluation, we introduce synthetic anomalous data points through our data augmentation methods. Our framework achieved 63.78% accuracy in detecting the points known to the machine and unknown to the machine from the Sensat Urban validation scene. Within the scene, 78.22% of the incorrectly classified data were detected as unknown to the machine. Furthermore, our framework achieved 84.34% accuracy in detecting the synthetic data and 35.5% accuracy in detecting those data as anomalies.
Johnahan Van Zyl, Hung Du, Srikanth Thudumu, Irini Logothetis, Scott Barnett, Rajesh Vasa, Kon Mouzakis
e-Science5
2021 Towards a taxonomy for annotation of data science experiment repositories
abstract
Data scientists, like software engineers, use search engines, code repositories, tutorials, and question and answer sites for finding code snippets. The objective of this study is to understand what information can be extracted from data science experiment repositories for quicker availability of relevant information when data scientists search for information. In this paper, we investigated a set of notebooks to identify recurring data science techniques for efficient information retrieval and easy adaptation from online solutions to support their search during experimentation. From the manual annotation of 57 natural language processing notebooks, a taxonomy on 106 data science techniques was developed, grouped by data science workflow stages. The preliminary evaluation shows that our constructed taxonomy is relevant to retrieve information that data scientists are searching for. Future work will continue to investigate the creation of a context aware code snippet engine designed for data scientists.
Shangeetha Sivasothy, Scott Barnett, Niroshinie Fernando, Rajesh Vasa, Roopak Sinha, Anj Simmons
SCAM2
2020 A large-scale comparative analysis of Coding Standard conformance in Open-Source Data Science projects
abstract
Background: Meeting the growing industry demand for Data Science requires cross-disciplinary teams that can translate machine learning research into production-ready code. Software engineering teams value adherence to coding standards as an indication of code readability, maintainability, and developer expertise. However, there are no large-scale empirical studies of coding standards focused specifically on Data Science projects. Aims: This study investigates the extent to which Data Science projects follow code standards. In particular, which standards are followed, which are ignored, and how does this differ to traditional software projects? Method: We compare a corpus of 1048 Open-Source Data Science projects to a reference group of 1099 non-Data Science projects with a similar level of quality and maturity. Results: Data Science projects suffer from a significantly higher rate of functions that use an excessive numbers of parameters and local variables. Data Science projects also follow different variable naming conventions to non-Data Science projects. Conclusions: The differences indicate that Data Science codebases are distinct from traditional software codebases and do not follow traditional software engineering conventions. Our conjecture is that this may be because traditional software engineering conventions are inappropriate in the context of Data Science projects.
Anj Simmons, Scott Barnett, Jessica Rivera-Villicana, Akshat Bajaj, Rajesh Vasa
ESEM2
2020 Interpreting cloud computer vision pain-points: a mining study of stack overflow
abstract
Intelligent services are becoming increasingly more pervasive; application developers want to leverage the latest advances in areas such as computer vision to provide new services and products to users, and large technology firms enable this via RESTful APIs. While such APIs promise an easy-to-integrate on-demand machine intelligence, their current design, documentation and developer interface hides much of the underlying machine learning techniques that power them. Such APIs look and feel like conventional APIs but abstract away data-driven probabilistic behaviour---the implications of a developer treating these APIs in the same way as other, traditional cloud services, such as cloud storage, is of concern. The objective of this study is to determine the various pain-points developers face when implementing systems that rely on the most mature of these intelligent services, specifically those that provide computer vision. We use Stack Overflow to mine indications of the frustrations that developers appear to face when using computer vision services, classifying their questions against two recent classification taxonomies (documentation-related and general questions). We find that, unlike mature fields like mobile development, there is a contrast in the types of questions asked by developers. These indicate a shallow understanding of the underlying technology that empower such systems. We discuss several implications of these findings via the lens of learning taxonomies to suggest how the software engineering community can improve these services and comment on the nature by which developers use them.
Alex Cummaudo, Rajesh Vasa, Scott Barnett, John C. Grundy, Mohamed Almorsy
ICSE3
2020 Beware the evolving 'intelligent' web service! an integration architecture tactic to guard AI-first components
abstract
Intelligent services provide the power of AI to developers via simple RESTful API endpoints, abstracting away many complexities of machine learning. However, most of these intelligent services---such as computer vision---continually learn with time. When the internals within the abstracted 'black box' become hidden and evolve, pitfalls emerge in the robustness of applications that depend on these evolving services. Without adapting the way developers plan and construct projects reliant on intelligent services, significant gaps and risks result in both project planning and development. Therefore, how can software engineers best mitigate software evolution risk moving forward, thereby ensuring that their own applications maintain quality? Our proposal is an architectural tactic designed to improve intelligent service-dependent software robustness. The tactic involves creating an application-specific benchmark dataset baselined against an intelligent service, enabling evolutionary behaviour changes to be mitigated. A technical evaluation of our implementation of this architecture demonstrates how the tactic can identify 1,054 cases of substantial confidence evolution and 2,461 cases of substantial changes to response label sets using a dataset consisting of 331 images that evolve when sent to a service.
Alex Cummaudo, Scott Barnett, Rajesh Vasa, John C. Grundy, Mohamed Almorsy
ESEC/SIGSOFT FSE2
2020 Threshy: supporting safe usage of intelligent web services
abstract
Increased popularity of ‘intelligent’ web services provides end-users with machine-learnt functionality at little effort to developers. However, these services require a decision threshold to be set which is dependent on problem-specific data. Developers lack a systematic approach for evaluating intelligent services and existing evaluation tools are predominantly targeted at data scientists for pre-development evaluation. This paper presents a workflow and supporting tool, Threshy, to help software developers select a decision threshold suited to their problem domain. Unlike existing tools, Threshy is designed to operate in multiple workflows including pre-development, pre-release, and support. Threshy is designed for tuning the confidence scores returned by intelligent web services and does not deal with hyper-parameter optimisation used in ML models. Additionally, it considers the financial impacts of false positives. Threshold configuration files exported by Threshy can be integrated into client applications and monitoring infrastructure. Demo: https://bit.ly/2YKeYhE.
Alex Cummaudo, Scott Barnett, Rajesh Vasa, John C. Grundy
ESEC/SIGSOFT FSE2
2015 Bootstrapping Mobile App Development
abstract
Modern IDEs provide limited support for developers when starting a new data-driven mobile app. App developers are currently required to write copious amounts of boilerplate code, scripts, organise complex directories, and author actual functionality. Although this scenario is ripe for automation, current tools are yet to address it adequately. In this paper we present RAPPT, a tool that generates the scaffolding of a mobile app based on a high level description specified in a Domain Specific Language (DSL). We demonstrate the feasibility of our approach by an example case study and feedback from a professional development team. Demo at: https://www.youtube.com/watch?v=ffquVgBYpLM.
Scott Barnett, Rajesh Vasa, John C. Grundy
ICSE (2)1
2015 A multi-view framework for generating mobile apps
abstract
This paper demonstrates a multi-view framework for Rapid APPlication Tool (RAPPT). RAPPT enables rapid development of mobile applications. It employs a multilevel approach to mobile application development: a Domain Specific Visual Language to define the high level structure of mobile apps, a Domain Specific Textual Language to define behavioural concepts, and concrete source code for fine grained improvements.
Scott Barnett, Iman Avazpour, Rajesh Vasa, John C. Grundy
VL/HCC1
2015 A Conceptual Model for Architecting Mobile Applications
abstract
Quality attributes are essential in software architecture and they are determined by identifying the concerns of the stakeholders of a system. The concerns of constructing mobile applications (apps) are quite specific due to the characteristics of mobile devices. These concerns have not been adequately addressed in industry standards and practices. In this paper, we present a mobile app development conceptual model comprising six key concepts that impact quality. Using two case studies, we show that these interrelated concepts influence the architectural decisions of mobile apps and their tradeoffs need to be well considered. As such, we suggest that these concepts should be first class entities when designing mobile app architecture to ensure that the quality attributes are satisfied.
Scott Barnett, Rajesh Vasa, Antony Tang
WICSA1