VLDB 2026 Research / reviewers in the wild / expert
Rajesh Vasa
dblp:62/3920
· DBLP profile ↗
40ranked-venue papers
2as first author
15since 2021 · last 2026
0000-0003-4805-1467ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 31 · 2 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 4 since 2021Artificial intelligence and machine learning · 5 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 3Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Goal-Oriented Multi-Agent Reinforcement Learning for Decentralized Agent TeamsabstractConnected and autonomous vehicles across land, water, and air must often operate in dynamic, unpredictable environments with limited communication, no centralized control, and partial observability. These real-world constraints pose significant challenges for coordination, particularly when vehicles pursue individual objectives. To address this, we propose a decentralized Multi-Agent Reinforcement Learning (MARL) framework that enables vehicles, acting as agents, to communicate selectively based on local goals and observations. This goal-aware communication strategy allows agents to share only relevant information, enhancing collaboration while respecting visibility limitations. We validate our approach in complex multi-agent navigation tasks featuring obstacles and dynamic agent populations. Results show that our method significantly improves task success rates and reduces time-to-goal compared to non-cooperative baselines. Moreover, task performance remains stable as the number of agents increases, demonstrating scalability. These findings highlight the potential of decentralized, goal-driven MARL to support effective coordination in realistic multi-vehicle systems operating across diverse domains. Hung Du, Hy Nguyen, Srikanth Thudumu, Rajesh Vasa, Kon Mouzakis |
CCNC | 4 |
| 2026 | Mitigating malware prevalence in networks with arbitrary topologies: a Flip-It cyber game approach integrated with epidemic modelingabstractCyber threats have evolved in complexity, aiming at a wide range of sectors using advanced methods and tools. This evolving threat landscape challenges existing cybersecurity frameworks, many of which lack the adaptability to counteract the complex tactics of sophisticated adversaries. Developing robust cyber defense strategies requires simulating dynamic interactions between attackers and defenders across high, moderate, and low-impact scenarios. The Flip-It cyber game serves as an intelligent framework for simulating these interactions, enabling the analysis of adaptive strategies in cybersecurity. This paper aims to address the problem of mitigating malware prevalence in full consideration of attack/defense capabilities in arbitrary network topologies. This paper proposes a sophisticated discrete-time epidemic model to characterize security state transitions over time for all three scenarios within the Flip-It game framework. On this basis, the original problem is modeled as a closed-loop control problem to seek the optimal containment strategy. Deep Reinforcement Learning (DRL) is then used to tackle the problem, generating efficient defense strategies that are well-adapted to changing cybersecurity environments. Numerical simulations based on small-world networks, scale-free networks, and router networks are then carried out to generate corresponding strategies. Additionally, we have evaluated the performance of the proposed method against the State-Of-The-Art (SOTA) in terms of attack/defense objective function, control actions, number of devices under the control of the attacker and defender, stability, execution time, and scalability. This comprehensive approach integrates epidemiological modeling, game theory, and advanced machine learning to effectively tackle the complexities of contemporary cybersecurity threats. • Mitigates malware across low, medium, and high-impact cyberattacks. • Integrates the Flip-It game for attacker-defender dynamic interactions. • Employs DRL to enable adaptive and optimized defense strategies. • Evaluates defense evolution across diverse network topologies. Mousa Tayseer Jafar, Lu-Xing Yang, Gang Li 0009, Robin Doss, Kon Mouzakis, Rajesh Vasa, Helge Janicke, Ahmed Ibrahim 0002, Ahmed Mohsin, Iqbal H. Sarker, Kristen Moore, Seyit Ahmet Çamtepe, Diksha Goel |
Inf. Sci. | 6 |
| 2025 | Safeguarding LLM-Applications: Specify or Train?abstractLarge Language Models (LLMs) are powerful tools used in several applications such as conversational AI, and code generation. However, significant robustness concerns arise with LLMs in production, such as hallucinations, prompt injection attacks, harmful content generation, and challenges in maintaining accurate domain-specific content moderation. Guardrails aim to mitigate these challenges by aligning LLM outputs with desired behaviors without modifying the underlying models. Nvidia NeMo Guardrails, for instance, rely on specifying acceptable/unacceptable behaviours. However, it is challenging to predict and address potential issues of LLMs in advance to create these guardrails. Also, manual updates from software engineers are often required to maintain and refine these guardrails. We introduce LLM-Guards, specialised machine learning (ML) models trained to function as protective guards. Additionally, we present an automation pipeline for training and continual fine-tuning of these guards using reinforcement learning from human feedback (RLHF). We evaluated several small LLMs, including Llama-3, Mistral, and Gemma, as LLM-Guards for challenges such as moderation and detecting off-topic queries, and compared their performance against NeMo Guardrails. The proposed Llama-3 LLM-Guard outperformed NeMo Guardrails in detecting offtopic queries, achieving an accuracy of 98.7% compared to 81%. Furthermore, the LLM-Guard detected 97.86% of harmful queries” surpassing NeMo Guardrails by 19.86%. Hala Abdelkader, Mohamed Almorsy, Sankhya Singh, Irini Logothetis, Priya Rani, Rajesh Vasa, Jean-Guy Schneider |
CAIN | 6 |
| 2025 | RAGProbe: Breaking RAG Pipelines with Evaluation ScenariosabstractRetrieval Augmented Generation (RAG) is increasingly employed in building Generative AI applications, yet their evaluation often relies on manual, trial-and-error processes. Automating this evaluation process involves generating test data to trigger failures involving context comprehension, data formatting, specificity, and content completeness. Random question-answer generation is insufficient. However, prior works rely on standard QA datasets, benchmarks and tactics that are not tailored to the specific domain requirements. Hence, current approaches and datasets do not trigger sufficiently broad and context-specific failures. In this paper, we introduce evaluation scenarios that describe the process of generating question-answer pairs from content indexed by RAG pipelines, and they are designed to trigger a wider range of failures and to simplify automation. This enables developers to identify and address weaknesses more effectively. We validate our approach on five open-source RAG pipelines using three datasets. Our approach triggers high failure rates, by generating prompts that combine multiple questions (up to 91% failure rate) highlighting the need for developers to prioritize handling such queries. We generated failure rates of 60% in an academic domain dataset and 53% and 64% in open-domain datasets. Compared to existing state-of-the-art methods, our approach triggers 77% more failures on average per RAG pipeline and 53% more failures on average per dataset, offering a mechanism to support developers to improve the RAG pipeline quality. Shangeetha Sivasothy, Scott Barnett, Stefanus Kurniawan, Zafaryab Rasool, Rajesh Vasa |
CAIN | 5 |
| 2025 | Optimizing Deep Reinforcement Learning Configurations for Single Object TrackingabstractDeep Reinforcement Learning (DRL) has become a critical approach for Object Tracking (OT) due to its ability to handle the sequential decision-making processes inherent in tracking tasks. By iteratively refining predictions and adapting to changes in object appearance or motion, DRL-based methods offer robust performance in complex tracking scenarios. However, most existing DRL-based OT methods have primarily focused on algorithmic framework design, often overlooking the optimization of configurations, such as action space, state space, reward function, and DRL algorithm. This oversight is significant, as optimal configuration choices can enhance DRL framework performance by up to 64%. Addressing this gap, our study investigates the impact of various configuration setups on the performance of DRL-based systems, specifically for Single Object Tracking (SOT) with a fixed camera view. Through theoretical analyses and experimentation, we demonstrate that appropriate configurations can improve tracking precision and system adaptability. The insights from this study provide a foundational guide for optimizing DRL applications in SOT with a fixed camera view, paving the way for more robust and efficient implementations in practical scenarios. Hy Nguyen, Srikanth Thudumu, Hung Du, Rajesh Vasa, Kon Mouzakis |
eScience | 4 |
| 2025 | CSAOT: Cooperative Multi-Agent System for Active Object TrackingabstractObject Tracking is essential for many computer vision applications, such as autonomous navigation, surveillance, and robotics. Unlike Passive Object Tracking (POT), which relies on static camera viewpoints to detect and track objects across consecutive frames, Active Object Tracking (AOT) requires a controller agent to actively adjust its viewpoint to maintain visual contact with a moving target in complex environments. Existing AOT solutions are predominantly single-agent-based, which struggle in dynamic and complex scenarios due to limited information gathering and processing capabilities, often resulting in suboptimal decision-making. Alleviating these limitations necessitates the development of a multi-agent system where different agents perform distinct roles and collaborate to enhance learning and robustness in dynamic and complex environments. Although some multi-agent approaches exist for AOT, they typically rely on external auxiliary agents, which require additional devices, making them costly. In contrast, we introduce the Collaborative System for Active Object Tracking (CSAOT), a method that leverages multi-agent deep reinforcement learning (MADRL) and a Mixture of Experts (MoE) framework to enable multiple agents to operate on a single device, thereby improving tracking performance and reducing costs. Our approach enhances robustness against occlusions and rapid motion while optimizing camera movements to extend tracking duration. We validated the effectiveness of CSAOT on various interactive maps with dynamic and stationary obstacles. Hy Nguyen, Bao Pham, Srikanth Thudumu, Hung Du, Rajesh Vasa, Kon Mouzakis |
ECAI | 5 |
| 2024 | ML-On-Rails: Safeguarding Machine Learning Models in Software Systems - A Case StudyabstractMachine learning (ML), especially with the emergence of large language models (LLMs), has significantly transformed various industries. However, the transition from ML model prototyping to production use within software systems presents several challenges. These challenges primarily revolve around ensuring safety, security, and transparency, subsequently influencing the overall robustness and trustworthiness of ML models. In this paper, we introduce ML-On-Rails, a protocol designed to safeguard ML models, establish a well-defined endpoint interface for different ML tasks, and clear communication between ML providers and ML consumers (software engineers). ML-On-Rails enhances the robustness of ML models via incorporating detection capabilities to identify unique challenges specific to production ML. We evaluated the ML-On-Rails protocol through a real-world case study of the MoveReminder application. Through this evaluation, we emphasize the importance of safeguarding ML models in production. Hala Abdelkader, Mohamed Almorsy, Scott Barnett, Jean-Guy Schneider, Priya Rani, Rajesh Vasa |
CAIN | 6 |
| 2024 | Towards Robust ML-enabled Software Systems: Detecting Out-of-Distribution data using Gini CoefficientsabstractMachine learning (ML) models have become essential components in software systems across several domains, such as autonomous driving, healthcare, and finance. The robustness of these ML models is crucial for maintaining the software systems performance and reliability. A significant challenge arises when these systems encounter out-of-distribution (OOD) data, examples that differ from the training data distribution. OOD data can cause a degradation of the software systems performance. Therefore, an effective OOD detection mechanism is essential for maintaining software system performance and robustness. Such a mechanism should identify and reject OOD inputs and alert software engineers. Current OOD detection methods rely on hyperparameters tuned with in-distribution and OOD data. However, defining the OOD data that the system will encounter in production is often infeasible. Further, the performance of these methods degrades with OOD data that has similar characteristics to the in-distribution data. In this paper, we propose a novel OOD detection method using the Gini coefficient. Our method does not require prior knowledge of OOD data or hyperparameter tuning. On common benchmark datasets, we show that our method outperforms the existing maximum softmax probability (MSP) baseline. For a model trained on the MNIST dataset, we improve the OOD detection rate by 4% on the CIFAR10 dataset and by more than 50% for the EMNIST dataset. Hala Abdelkader, Jean-Guy Schneider, Mohamed Almorsy, Priya Rani, Rajesh Vasa |
ASE | 5 |
| 2024 | Comparative analysis of real issues in open-source machine learning projectsabstractAbstract Context In the last decade of data-driven decision-making, Machine Learning (ML) systems reign supreme. Because of the different characteristics between ML and traditional Software Engineering systems, we do not know to what extent the issue-reporting needs are different, and to what extent these differences impact the issue resolution process. Objective We aim to compare the differences between ML and non-ML issues in open-source applied AI projects in terms of resolution time and size of fix. This research aims to enhance the predictability of maintenance tasks by providing valuable insights for issue reporting and task scheduling activities. Method We collect issue reports from Github repositories of open-source ML projects using an automatic approach, filter them using ML keywords and libraries, manually categorize them using an adapted deep learning bug taxonomy, and compare resolution time and fix size for ML and non-ML issues in a controlled sample. Result 147 ML issues and 147 non-ML issues are collected for analysis. We found that ML issues take more time to resolve than non-ML issues, the median difference is 14 days. There is no significant difference in terms of size of fix between ML and non-ML issues. No significant differences are found between different ML issue categories in terms of resolution time and size of fix. Conclusion Our study provided evidence that the life cycle for ML issues is stretched, and thus further work is required to identify the reason. The results also highlighted the need for future work to design custom tooling to support faster resolution of ML issues. Tuan Dung Lai, Anj Simmons, Scott Barnett, Jean-Guy Schneider, Rajesh Vasa |
Empir. Softw. Eng. | 5 |
| 2023 | Decentralized Federated Learning Strategy with Image Classification using ResNet ArchitectureabstractThe rapid growth of both the Industrial Internet of Things (IIoT) and Artificial Intelligence (AI) results in a high demand for AI applications in devices. To achieve high levels of accuracy, AI applications typically require a large amount of annotated data. Accessing such data is challenging in various applications such as healthcare, finance and information security. Federated learning (FL) is one of the strategies that was proposed to overcome this challenge. Specifically, FL enables the AI model in the centralized system to be trained without any prior knowledge of the information on the devices. Recent FLs have the disadvantage that they are dependent upon a centralized system, and thus are susceptible to single points of failure. This paper proposes a strategy that employs FL in a decentralized environment where devices can communicate with each other to increase the accuracy of the AI model in each device. Furthermore, we evaluate the proposed strategy in the image classification task with the ResNet50 architecture and the CIFAR-10 dataset. The evaluation shows that the ResNet50 model trained in the decentralized environment can achieve comparable results to the model trained in the centralized environment. Hung Du, Srikanth Thudumu, Sankhya Singh, Scott Barnett, Irini Logothetis, Rajesh Vasa, Kon Mouzakis |
CCNC | 6 |
| 2022 | A Framework for Evaluating MRC Approaches with Unanswerable QuestionsabstractMachine reading comprehension (MRC) is a challenging task in natural language processing that demonstrates the language understanding of the machine. An approach to tackle this challenge requires the machine to answer the question about the given context when needed and abstain from answering when there is no answer. Recent works attempted to solve this challenge with various comprehensive neural network architectures for sequences such as SAN, U-Net, EQuANt, and others that were trained on the SQuAD 2.0 dataset containing unanswerable questions. However, the robustness of these approaches has not been evaluated. In this paper, we propose a data augmentation approach that converts answerable questions to unanswerable questions in the SQuAD 2.0 dataset by altering the entities in the question to its antonym from ConceptNet which is a semantic network. The augmented data is, then, fitted into the U-Net question answering model to evaluate the robustness of the model. Hung Du, Srikanth Thudumu, Sankhya Singh, Scott Barnett, Irini Logothetis, Rajesh Vasa, Kon Mouzakis |
e-Science | 6 |
| 2022 | PiMS: A Pre-ML Labelling ToolabstractMachine Learning (ML) techniques in clinical decision support systems are scarce due to the limited availability of clinically validated and labelled training data sets. We present a framework to (1) enable quality controls at data submission toward ML appropriate data, (2) provide in-situ algorithm assessments, and (3) prepare dataframes for ML training and robust stochastic analysis. We developed and evaluated PiMS (Pandemic Intervention and Monitoring Systems): a remote monitoring solution for patients that are Covid-positive. The system was trialled at two hospitals in Melbourne, Australia (Alfred Health and Monash Health) involving 109 patients and 15 clinicians. Irini Logothetis, Scott Barnett, Leonard Hoon, Srikanth Thudumu, Joseph Mathew, Carl Luckhoff, Gerard O'Reilly, David Collard, Rajesh Vasa, Kon Mouzakis, Mark Fitzgerald |
e-Science | 9 |
| 2022 | Subspace based Anomaly Detection Framework for Point CloudsabstractIn many real-world applications such as the inspection of powerlines, the automated detection of anomalies can minimise damage and reduce costs that result from the presence of unknown anomalies. Technologies such as LiDAR scans obtained from Unmanned Aerial Vehicles (UAV) are becoming prominent due to the data depth they provide. In the context of powerline transmission, investigators must search for anomalous elements such as line defects or obstructions. Such occurrences are not always apparent and detecting them requires extensive analysis of data within vast areas of wilderness. Automating this process can reduce time and labor costs. We propose a methodology to define what constitutes an anomaly within mapped real-world scenes, and a technique to address different types of anomalies. The notion of unknowns and knowns composed of unknown to both human and machine, known to human and unknown to machine, unknown to human and known to machine, and known to both human and machine is considered to develop a novel framework that detects anomalous patterns. For the purpose of evaluation, we introduce synthetic anomalous data points through our data augmentation methods. Our framework achieved 63.78% accuracy in detecting the points known to the machine and unknown to the machine from the Sensat Urban validation scene. Within the scene, 78.22% of the incorrectly classified data were detected as unknown to the machine. Furthermore, our framework achieved 84.34% accuracy in detecting the synthetic data and 35.5% accuracy in detecting those data as anomalies. Johnahan Van Zyl, Hung Du, Srikanth Thudumu, Irini Logothetis, Scott Barnett, Rajesh Vasa, Kon Mouzakis |
e-Science | 6 |
| 2022 | Requirements of API Documentation: A Case Study into Computer Vision ServicesabstractUsing cloud-based computer vision services is gaining traction, where developers access AI-powered components through familiar RESTful APIs, not needing to orchestrate large training and inference infrastructures or curate/label training datasets. However, while these APIsseemfamiliar to use, their non-deterministic run-time behaviour and evolution is not adequately communicated to developers. Therefore, improving these services’ API documentation is paramount—more extensive documentation facilitates the development process of intelligent software. In a prior study, we extracted 34 API documentation artefacts from 21 seminal works, devising a taxonomy of five key requirements to produce quality API documentation. We extend this study in two ways. First, by surveying 104 developers of varying experience to understand what API documentation artefacts are ofmost valueto practitioners. Second, identifying which of these highly-valued artefacts are or are not well-documented through a case study in the emerging computer vision service domain. We identify: (i) several gaps in the software engineering literature, where aspects of API documentation understanding is/is not extensively investigated; and (ii) where industry vendors (in contrast) document artefacts to better serve their end-developers. We provide a set of recommendations to enhance intelligent software documentation for both vendors and the wider research community. Alex Cummaudo, Rajesh Vasa, John C. Grundy, Mohamed Almorsy |
IEEE Trans. Software Eng. | 2 |
| 2021 | Towards a taxonomy for annotation of data science experiment repositoriesabstractData scientists, like software engineers, use search engines, code repositories, tutorials, and question and answer sites for finding code snippets. The objective of this study is to understand what information can be extracted from data science experiment repositories for quicker availability of relevant information when data scientists search for information. In this paper, we investigated a set of notebooks to identify recurring data science techniques for efficient information retrieval and easy adaptation from online solutions to support their search during experimentation. From the manual annotation of 57 natural language processing notebooks, a taxonomy on 106 data science techniques was developed, grouped by data science workflow stages. The preliminary evaluation shows that our constructed taxonomy is relevant to retrieve information that data scientists are searching for. Future work will continue to investigate the creation of a context aware code snippet engine designed for data scientists. Shangeetha Sivasothy, Scott Barnett, Niroshinie Fernando, Rajesh Vasa, Roopak Sinha, Anj Simmons |
SCAM | 4 |
| 2020 | A large-scale comparative analysis of Coding Standard conformance in Open-Source Data Science projectsabstractBackground: Meeting the growing industry demand for Data Science requires cross-disciplinary teams that can translate machine learning research into production-ready code. Software engineering teams value adherence to coding standards as an indication of code readability, maintainability, and developer expertise. However, there are no large-scale empirical studies of coding standards focused specifically on Data Science projects. Aims: This study investigates the extent to which Data Science projects follow code standards. In particular, which standards are followed, which are ignored, and how does this differ to traditional software projects? Method: We compare a corpus of 1048 Open-Source Data Science projects to a reference group of 1099 non-Data Science projects with a similar level of quality and maturity. Results: Data Science projects suffer from a significantly higher rate of functions that use an excessive numbers of parameters and local variables. Data Science projects also follow different variable naming conventions to non-Data Science projects. Conclusions: The differences indicate that Data Science codebases are distinct from traditional software codebases and do not follow traditional software engineering conventions. Our conjecture is that this may be because traditional software engineering conventions are inappropriate in the context of Data Science projects. Anj Simmons, Scott Barnett, Jessica Rivera-Villicana, Akshat Bajaj, Rajesh Vasa |
ESEM | 5 |
| 2020 | Interpreting cloud computer vision pain-points: a mining study of stack overflowabstractIntelligent services are becoming increasingly more pervasive; application developers want to leverage the latest advances in areas such as computer vision to provide new services and products to users, and large technology firms enable this via RESTful APIs. While such APIs promise an easy-to-integrate on-demand machine intelligence, their current design, documentation and developer interface hides much of the underlying machine learning techniques that power them. Such APIs look and feel like conventional APIs but abstract away data-driven probabilistic behaviour---the implications of a developer treating these APIs in the same way as other, traditional cloud services, such as cloud storage, is of concern. The objective of this study is to determine the various pain-points developers face when implementing systems that rely on the most mature of these intelligent services, specifically those that provide computer vision. We use Stack Overflow to mine indications of the frustrations that developers appear to face when using computer vision services, classifying their questions against two recent classification taxonomies (documentation-related and general questions). We find that, unlike mature fields like mobile development, there is a contrast in the types of questions asked by developers. These indicate a shallow understanding of the underlying technology that empower such systems. We discuss several implications of these findings via the lens of learning taxonomies to suggest how the software engineering community can improve these services and comment on the nature by which developers use them. Alex Cummaudo, Rajesh Vasa, Scott Barnett, John C. Grundy, Mohamed Almorsy |
ICSE | 2 |
| 2020 | Beware the evolving 'intelligent' web service! an integration architecture tactic to guard AI-first componentsabstractIntelligent services provide the power of AI to developers via simple RESTful API endpoints, abstracting away many complexities of machine learning. However, most of these intelligent services---such as computer vision---continually learn with time. When the internals within the abstracted 'black box' become hidden and evolve, pitfalls emerge in the robustness of applications that depend on these evolving services. Without adapting the way developers plan and construct projects reliant on intelligent services, significant gaps and risks result in both project planning and development. Therefore, how can software engineers best mitigate software evolution risk moving forward, thereby ensuring that their own applications maintain quality? Our proposal is an architectural tactic designed to improve intelligent service-dependent software robustness. The tactic involves creating an application-specific benchmark dataset baselined against an intelligent service, enabling evolutionary behaviour changes to be mitigated. A technical evaluation of our implementation of this architecture demonstrates how the tactic can identify 1,054 cases of substantial confidence evolution and 2,461 cases of substantial changes to response label sets using a dataset consisting of 331 images that evolve when sent to a service. Alex Cummaudo, Scott Barnett, Rajesh Vasa, John C. Grundy, Mohamed Almorsy |
ESEC/SIGSOFT FSE | 3 |
| 2020 | Threshy: supporting safe usage of intelligent web servicesabstractIncreased popularity of ‘intelligent’ web services provides end-users with machine-learnt functionality at little effort to developers. However, these services require a decision threshold to be set which is dependent on problem-specific data. Developers lack a systematic approach for evaluating intelligent services and existing evaluation tools are predominantly targeted at data scientists for pre-development evaluation. This paper presents a workflow and supporting tool, Threshy, to help software developers select a decision threshold suited to their problem domain. Unlike existing tools, Threshy is designed to operate in multiple workflows including pre-development, pre-release, and support. Threshy is designed for tuning the confidence scores returned by intelligent web services and does not deal with hyper-parameter optimisation used in ML models. Additionally, it considers the financial impacts of false positives. Threshold configuration files exported by Threshy can be integrated into client applications and monitoring infrastructure. Demo: https://bit.ly/2YKeYhE. Alex Cummaudo, Scott Barnett, Rajesh Vasa, John C. Grundy |
ESEC/SIGSOFT FSE | 3 |
| 2020 | A revised open source usability defect classification taxonomy
Nor Shahida Mohamad Yusop, John C. Grundy, Jean-Guy Schneider, Rajesh Vasa |
Inf. Softw. Technol. | 4 |
| 2019 | What should I document? A preliminary systematic mapping study into API documentation knowledgeabstractBackground: Good API documentation facilitates the development process, improving productivity and quality. While the topic of API documentation quality has been of interest for the last two decades, there have been few studies to map the specific constructs needed to create a good document. In effect, we still need a structured taxonomy that captures such knowledge systematically.Aims: This study reports emerging results of a systematic mapping study. We capture key conclusions from previous studies that assess API documentation quality, and synthesise the results into a single framework.Method: By conducting a systematic review of 21 key works, we have developed a five dimensional taxonomy based on 34 categorised weighted recommendations.Results: All studies utilise field study techniques to arrive at their recommendations, with seven studies employing some form of interview and questionnaire, and four conducting documentation analysis. The taxonomy we synthesise reinforces that usage description details (code snippets, tutorials, and reference documents) are generally highly weighted as helpful in API documentation, in addition to design rationale and presentation.Conclusions: We propose extensions to this study aligned to developer utility for each of the taxonomy's categories. Alex Cummaudo, Rajesh Vasa, John C. Grundy |
ESEM | 2 |
| 2019 | Losing Confidence in Quality: Unspoken Evolution of Computer Vision ServicesabstractThe following topics are dealt with: software maintenance; public domain software; program testing; source code (software); program debugging; software quality; program diagnostics; learning (artificial intelligence); mobile computing; data mining. Alex Cummaudo, Rajesh Vasa, John C. Grundy, Mohamed Almorsy, Andrew Cain |
ICSME | 2 |
| 2019 | Merging Intelligent API Responses Using a Proportional Representation Approach
Tomohiro Ohtake, Alex Cummaudo, Mohamed Almorsy, Rajesh Vasa, John C. Grundy |
ICWE | 4 |
| 2019 | Emotion-oriented requirements engineering: A case study in developing a smart home system for the elderly
Maheswaree Kissoon Curumsing, Niroshinie Fernando, Mohamed Almorsy, Rajesh Vasa, Kon Mouzakis, John C. Grundy |
J. Syst. Softw. | 4 |
| 2017 | Analysis of the Textual Content of Mined Open Source Usability Defect ReportsabstractWriting a good usability defect report can be a tedious task, especially in identifying what important information should be included, and capturing the attention of software developers to fix them. This paper is a continuity of our previous studies investigating software development practitioners' day-to-day practices when dealing with usability defects. In this study, we mined 377 developer-tagged usability defect reports from Mozilla Thunderbird, Firefox for Android and Eclipse Platform to confirm what software development practitioners claimed to provide when reporting usability defects. We looked for the presence of key defect attributes - steps to reproduce, impact, software context, expected output, actual output, assumed causes, solution proposal and supplementary information. In addition, we analyzed the trend of different types of usability defects, correlation between usability defects and defect severity, and failure qualifier. Our findings demonstrate a mismatch between what software development practitioners claimed to provide when reporting usability defects, and the information that actually appears in the defect reports. The results of our research have important implications for software defect reporting, especially in designing more effective mechanisms for reporting usability defects. Nor Shahida Mohamad Yusop, Jean-Guy Schneider, John C. Grundy, Rajesh Vasa |
APSEC | 4 |
| 2017 | Spatio-Temporal Reference Frames as Geographic ObjectsabstractIt is often desirable to analyse trajectory data in local coordinates relative to a reference location. Similarly, temporal data also needs to be transformed to be relative to an event. Together, temporal and spatial contextualisation permits comparative analysis of similar trajectories taken across multiple reference locations. To the GIS professional, the procedures to establish a reference frame at a location and reproject the data into local coordinates are well known, albeit tedious. However, GIS tools are now often used by subject matter experts who may not have the deep knowledge of coordinate frames and projections required to use these techniques effectively. Anj Simmons, Rajesh Vasa |
SIGSPATIAL/GIS | 2 |
| 2017 | Reporting Usability Defects: A Systematic Literature ReviewabstractUsability defects can be found either by formal usability evaluation methods or indirectly during system testing or usage. No matter how they are discovered, these defects must be tracked and reported. However, empirical studies indicate that usability defects are often not clearly and fully described. This study aims to identify the state of the art in reporting of usability defects in the software engineering and usability engineering literature. We conducted a systematic literature review of usability defect reporting drawing from both the usability and software engineering literature from January 2000 until March 2016. As a result, a total of 57 studies were identified, in which we classified the studies into three categories: reporting usability defect information, analysing usability defect data and key challenges. Out of these, 20 were software engineering studies and 37 were usability studies. The results of this systematic literature review show that usability defect reporting processes suffer from a number of limitations, including: mixed data, inconsistency of terms and values of usability defect data, and insufficient attributes to classify usability defects. We make a number of recommendations to improve usability defect reporting and management in software engineering. Nor Shahida Mohamad Yusop, John C. Grundy, Rajesh Vasa |
IEEE Trans. Software Eng. | 3 |
| 2016 | What Influences Usability Defect Reporting? - A Survey of Software Development PractitionersabstractSoftware development organizations invest in test automation tools and methods to optimize defect discovery rates. The true value of these tools is realized when the defects are addressed before release, and hence good quality defect reports are critical. We describe a survey we conducted to better understand usability defect reporting, in particular, influences on the quality of usability defect reports. We analyze feedback from nearly 150 software developers and usability defect reporters and identify key determinants of quality defect reports, aspects of usability defects that are challenging to report and directions for future research into usability defect reporting tools to improve usability defect reports quality. Nor Shahida Mohamad Yusop, Jean-Guy Schneider, John C. Grundy, Rajesh Vasa |
APSEC | 4 |
| 2016 | Reporting usability defects: do reporters report what software developers need?abstractReporting usability defects can be a challenging task, especially in convincing the software developers that the reported defect actually requires attention. Stronger evidence in the form of specific details is often needed. However, research to date in software defect reporting has not investigated the value of capturing different information based on defect type. We surveyed practitioners in both open source communities and industrial software organizations about their usability defect reporting practices to better understand information needs to address usability defect reporting issues. Our analysis of 147 responses show that reporters often provide observed result, expected result and steps to reproduce when describing usability defects, similar to the way other types of defects are reported. However, reporters rarely provide usability-related information. In fact, reporters ranked cause of the problem is the most difficult information to provide followed by usability principle, video recoding, UI event trace and title. Conversely, software developers consider cause of the problem as the most helpful information for them to fix usability defects. Our statistical analysis reveals a substantial gap between what reporters provide and what software developers need when fixing usability defects. We propose some remedies to resolve this gap. Nor Shahida Mohamad Yusop, John C. Grundy, Rajesh Vasa |
EASE | 3 |
| 2016 | An empirical study of user perceived usefulness and preference of open learner model visualisationsabstractMany higher education institutions have reformed their academic programmes to adopt unit learning outcomes. It is essential for effective learning management tools to support this fundamental transformation. We see the need for a tool that provides visualisations of Open Learner Models (OLM) and associated e-portfolio content to guide students in achieving intended learning outcomes, help evidence their learning and keep them engaged in their study. OLMs surface the relationship between learning activities and tasks, formative and summative assessment, and intended learning outcomes. We have developed and validated a set of candidate visualisations for such OLMs. We report key findings from our study in terms of potential users' feedback on our tool's support for tracking learning progress against learning outcomes. Our findings can inform and refine learning management systems. Check Yee Law, John C. Grundy, Rajesh Vasa, Andrew Cain |
VL/HCC | 3 |
| 2015 | QoS-Aware Service Selection for Customisable Multi-tenant Service-Based Systems: Maturity and ApproachesabstractMulti-tenant service-based systems (SBSs) have become a major paradigm in software engineering in the cloud environment. Instead of serving a single end-user, a multitenant SBS provides multiple tenants with similar and yet customised functionalities with potentially different quality-of service (QoS) values. Thus, existing approaches to service selection for single-tenant SBSs are no longer suitable. Furthermore, the target multi-tenancy maturity level also needs to be considered in the service selection approach for an SBS. In this paper, we propose three novel QoS-aware service selection approaches for composing multi-tenant SBSs that achieve three different multi-tenancy maturity levels. Extensive and comprehensive experiments are conducted and the experimental results show that our approaches outperform the existing approach in both effectiveness and efficiency. Qiang He 0001, Jun Han 0004, Feifei Chen 0001, Rajesh Vasa, Yun Yang 0001, Hai Jin 0001 |
CLOUD | 5 |
| 2015 | A Preliminary Study of Open Learner Model Representation Formats to Support Formative AssessmentabstractOpen learner models provide a way of showing teachers and students the learning progress of a student against expectations. Creating an effective interface to present a learner model is an important part in open learner modeling. Usefulness of the learner model relies on using a clear and effective representation format to facilitate users' understanding of the information presented. In this paper, we propose a range of open learner model representation formats to display students' learning task status and achievement in terms of learning outcomes. We investigate different types of representation formats and data useful for users to inspect the learner model. An interface design prototype has been built to provide potential users with an evaluation platform that enables investigation of these aspects. The results obtained will allow us to develop a more effective new open learner model visualisation tool. Check Yee Law, John C. Grundy, Andrew Cain, Rajesh Vasa |
COMPSAC | 4 |
| 2015 | Multi-node Multi-agent Cloud Simulation: Approximating SynchronisationabstractTraffic engineering is a key in effective utilisation of the road network infrastructure. Simulation assists traffic engineers making informed decisions on how to operate and direct traffic within the road networks. These simulations are complex, generate big data and require high-powered computers, which can process information faster than real time, to ensure the results can be used to affect traffic. Cloud computing, a relatively new technology paradigm, can meet the essential requirements, such as scalability, interoperability, availability and high-end performance. In this paper, a novel approach to a synchronisation strategy of large-scale complex simulations is proposed. This approach builds upon advancements achieved in distributed computing. The new synchronisation strategy is designed to allow different granularities of synchronisation accuracy. Through this strategy, synchronisation overhead is reduced, thus allowing the computing bandwidth to be applied to simulation performance increases as a result of the trade off between synchronisation accuracy and performance. Antonio Giardina, Yun Yang 0001, Hai Le Vu 0001, Rajesh Vasa |
e-Science | 4 |
| 2015 | Bootstrapping Mobile App DevelopmentabstractModern IDEs provide limited support for developers when starting a new data-driven mobile app. App developers are currently required to write copious amounts of boilerplate code, scripts, organise complex directories, and author actual functionality. Although this scenario is ripe for automation, current tools are yet to address it adequately. In this paper we present RAPPT, a tool that generates the scaffolding of a mobile app based on a high level description specified in a Domain Specific Language (DSL). We demonstrate the feasibility of our approach by an example case study and feedback from a professional development team. Demo at: https://www.youtube.com/watch?v=ffquVgBYpLM. Scott Barnett, Rajesh Vasa, John C. Grundy |
ICSE (2) | 2 |
| 2015 | A multi-view framework for generating mobile appsabstractThis paper demonstrates a multi-view framework for Rapid APPlication Tool (RAPPT). RAPPT enables rapid development of mobile applications. It employs a multilevel approach to mobile application development: a Domain Specific Visual Language to define the high level structure of mobile apps, a Domain Specific Textual Language to define behavioural concepts, and concrete source code for fine grained improvements. Scott Barnett, Iman Avazpour, Rajesh Vasa, John C. Grundy |
VL/HCC | 3 |
| 2015 | Hub Map: A new approach for visualizing traffic data sets with multi-attribute link dataabstractVisualizing road traffic datasets involves representing junctions, their links, and the attributes of those links. Current traffic visualization techniques are not sufficient for professional traffic engineers, as they are limited in the number of attributes that can be represented. This paper proposes a new approach to visualize multiple attributes on graph edges without compromising their visibility. In particular, we introduce a parameterized connector symbol that increases the number of attributes that can be displayed on graph edges. We demonstrate that our approach can significantly increase the number of traffic parameters that can be displayed compared to existing traffic visualizations. Anj Simmons, Iman Avazpour, Hai Le Vu 0001, Rajesh Vasa |
VL/HCC | 4 |
| 2015 | A Conceptual Model for Architecting Mobile ApplicationsabstractQuality attributes are essential in software architecture and they are determined by identifying the concerns of the stakeholders of a system. The concerns of constructing mobile applications (apps) are quite specific due to the characteristics of mobile devices. These concerns have not been adequately addressed in industry standards and practices. In this paper, we present a mobile app development conceptual model comprising six key concepts that impact quality. Using two case studies, we show that these interrelated concepts influence the architectural decisions of mobile apps and their tradeoffs need to be well considered. As such, we suggest that these concepts should be first class entities when designing mobile app architecture to ensure that the quality attributes are satisfied. Scott Barnett, Rajesh Vasa, Antony Tang |
WICSA | 2 |
| 2012 | Learning Better Inspection Optimization PoliciesabstractRecent research has shown the value of social metrics for defect prediction. Yet many repositories lack the information required for a social analysis. So, what other means exist to infer how developers interact around their code? One option is static code metrics that have already demonstrated their usefulness in analyzing change in evolving software systems. But do they also help in defect prediction? To address this question we selected a set of static code metrics to determine what classes are most "active" (i.e., the classes where the developers spend much time interacting with each other's design and implementation decisions) in 33 open-source Java systems that lack details about individual developers. In particular, we assessed the merit of these activity-centric measures in the context of "inspection optimization" — a technique that allows for reading the fewest lines of code in order to find the most defects. For the task of inspection optimization these activity measures perform as well as (usually, within 4%) a theoretical upper bound on the performance of any set of measures. As a result, we argue that activity-centric static code metrics are an excellent predictor for defects. Markus Lumpe, Rajesh Vasa, Tim Menzies, Rebecca Rush, Burak Turhan |
Int. J. Softw. Eng. Knowl. Eng. | 2 |
| 2009 | Comparative analysis of evolving software systems using the Gini coefficientabstractSoftware metrics offer us the promise of distilling useful information from vast amounts of software in order to track development progress, to gain insights into the nature of the software, and to identify potential problems. Unfortunately, however, many software metrics exhibit highly skewed, non-Gaussian distributions. As a consequence, usual ways of interpreting these metrics - for example, in terms of ldquoaveragerdquo values - can be highly misleading. Many metrics, it turns out, are distributed like wealth - with high concentrations of values in selected locations. We propose to analyze software metrics using the Gini coefficient, a higher-order statistic widely used in economics to study the distribution of wealth. Our approach allows us not only to observe changes in software systems efficiently, but also to assess project risks and monitor the development process itself. We apply the Gini coefficient to numerous metrics over a range of software projects, and we show that many metrics not only display remarkably high Gini values, but that these values are remarkably consistent as a project evolves over time. Rajesh Vasa, Markus Lumpe, Philip Branch, Oscar Nierstrasz |
ICSM | 1 |
| 2007 | The Inevitable Stability of Software ChangeabstractReal software systems change and become more complex over time. But which parts change and which parts remain stable? Common wisdom, for example, states that in a well-designed object-oriented system, the more popular a class is, the less likely it is to change from one version to the next, since changes to this class are likely to impact its clients. We have studied consecutive releases of several public domain, object-oriented software systems and analyzed a number of measures indicative of size, popularity, and complexity of classes and interfaces. As it turns out, the distributions of these measures are remarkably stable as an application evolves. The distribution of class size and complexity retains its shape over time. Relatively little code is modified over time. Classes that tend to be modified, however, are also the more popular ones, that is, those with greater Fan-In. In general, the more "complex" a class or interface becomes, the more likely it is to change from one version to the next. Rajesh Vasa, Jean-Guy Schneider, Oscar Nierstrasz |
ICSM | 1 |