VLDB 2026 Research / reviewers in the wild / expert
Gunel Jahangirova
dblp:183/0177
· DBLP profile ↗
20ranked-venue papers
6as first author
13since 2021 · last 2025
0000-0002-1423-1083ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 19 · 6 first-author · 12 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | $\mu \text{PRL}$: A Mutation Testing Pipeline for Deep Reinforcement Learning Based on Real FaultsabstractReinforcement Learning (RL) is increasingly adopted to train agents that can deal with complex sequential tasks, such as driving an autonomous vehicle or controlling a humanoid robot. Correspondingly, novel approaches are needed to ensure that RL agents have been tested adequately before going to production. Among them, mutation testing is quite promising, especially under the assumption that the injected faults (mutations) mimic the real ones. In this paper, we first describe a taxonomy of real RL faults obtained by repository mining. Then, we present the mutation operators derived from such real faults and implemented in the tool$\mu \text{PRL}$. Finally, we discuss the experimental results, showing that$\mu \text{PRL}$is effective at discriminating strong from weak test generators, hence providing useful feedback to developers about the adequacy of the generated test scenarios. Deepak-George Thomas, Matteo Biagiola, Nargiz Humbatova, Mohammad Wardat, Gunel Jahangirova, Hridesh Rajan, Paolo Tonella |
ICSE | 5 |
| 2025 | LLMLOOP: Improving LLM-Generated Code and Tests Through Automated Iterative Feedback LoopsabstractLarge Language Models (LLMs) are showing remarkable performance in generating source code, yet the generated code often has issues like compilation errors or incorrect code. Researchers and developers often face wasted effort in implementing checks and refining LLM-generated code, frequently duplicating their efforts. This paper presents LLMLOOP, a framework that automates the refinement of both source code and test cases produced by LLMs. LLMLOOP employs five iterative loops: resolving compilation errors, addressing static analysis issues, fixing test case failures, and improving test quality through mutation analysis. These loops ensure the generation of high-quality test cases that serve as both a validation mechanism and a regression test suite for the generated code. We evaluated llmloop on HumanEval-X, a recent benchmark of programming tasks. Results demonstrate the tool effectiveness in refining LLM-generated outputs. A demonstration video of the tool is available at https://youtu.be/2CLG9x1fsNI. Ravin Ravi, Dylan Bradshaw, Stefano Ruberto, Gunel Jahangirova, Valerio Terragni |
ICSME | 4 |
| 2025 | An empirical study of fault localisation techniques for deep neural networksabstractWith the increased popularity of Deep Neural Networks (DNNs), increases also the need for tools to assist developers in the DNN implementation, testing and debugging process. Several approaches have been proposed that automatically analyse and localise potential faults in DNNs under test. In this work, we evaluate and compare existing state-of-the-art fault localisation techniques, which operate based on both dynamic and static analysis of the DNN. The evaluation is performed on a benchmark consisting of both real faults obtained from bug reporting platforms and faulty models produced by a mutation tool. Our findings indicate that the usage of a single, specific ground truth (e.g. the human-defined one) for the evaluation of DNN fault localisation tools results in pretty low performance (maximum average recall of 0.33 and precision of 0.21). However, such figures increase when considering alternative, equivalent patches that exist for a given faulty DNN. The results indicate that DeepFD is the most effective tool, achieving an average recall of 0.55 and a precision of 0.37 on our benchmark. Nargiz Humbatova, Jinhan Kim, Gunel Jahangirova, Shin Yoo, Paolo Tonella |
Empir. Softw. Eng. | 3 |
| 2024 | Experience Report of the AWS+KCL Impact Accelerator for Public Sector EngagementabstractThis industry experience report chronicles the experience of developing an impact-focused group project module within a computer science master's programme at King's College London over two years. The module was set up in collaboration with Amazon Web Services to match student teams with public sector challenges requiring innovative technological solutions. An iterative process of modifications based on partner and student feedback aimed to enhance the learning experience and outcomes. Key benefits included providing authentic professional development for students, enabling innovation and entrepreneurship, building partnerships between academia and the public sector, and embedding responsible innovation into projects. However, challenges emerged around managing expectations, ensuring consistent partner engagement, providing support for spin-outs, and handling sensitive data issues. As more projects involved artificial intelligence applications in the second year, developing mechanisms to ethically provide access while protecting sensitive information was an increasingly crucial need. Moreover, understanding the value and impact of this model of software engineering project module requires additional research support. Overall, this collaborative module offers a promising model to deliver impact-driven solutions through coordinating academia, industry, and public sector partners. Further research can help optimise such partnerships for societal impact. Caitlin M. Bentley, Elena Simperl, Mike Bainbridge, Daisy Ogden, Stefanos Leonardos, Gunel Jahangirova, Joanna Walker, Christopher Hampson |
CSEE&T | 6 |
| 2024 | Spectral Analysis of the Relation between Deep Learning Faults and Neural Activation ValuesabstractThe growing adoption of Deep Learning (DL) systems in all areas of life makes the exposure and repair of errors in such systems a task of paramount importance. Existing approaches for measuring test adequacy and producing automated repair patches for such systems rely heavily on the activation values of neurons. However, there is no empirical evidence that links the behaviour of a DL model in terms of its neuron activation patterns to the type of faults present in that DL model. In this work, we perform a large-scale empirical study in which we inject artificial faults into a DL model and check whether the same types of faults lead to similar activation patterns. To analyse the patterns, we propose the notion of the spectrum of a deep neural network (DNN), which is defined as the probability distribution of the activation values of the neurons of the DNN. We perform our analysis on 6 subject systems considering 24 different fault types, each associated with a specific DL mutation operator. Our results show that we can successfully identify the fault type based on the top 3 spectra similarity matches in 75% of the cases and clusters of similar spectra have low impurity (i.e., they contain mostly one or few fault types). To demonstrate a practical application of this finding, we train different classifiers to predict the error type in a DL model from the spectra of the activation values. Our results show that the prediction accuracy can be as high as 79%, when the classifier is trained on subject-specific spectra. Nargiz Humbatova, Gunel Jahangirova, Paolo Tonella |
ICST | 2 |
| 2024 | GenMorph: Automatically Generating Metamorphic Relations via Genetic ProgrammingabstractMetamorphic testing is a popular approach that aims to alleviate the oracle problem in software testing. At the core of this approach are Metamorphic Relations (MRs), specifying properties that hold among multiple test inputs and corresponding outputs. Deriving MRs is mostly a manual activity, since their automated generation is a challenging and largely unexplored problem. This paper presentsGenMorph, a technique to automatically generate MRs for Java methods that involve inputs and outputs that are boolean, numerical, or ordered sequences.GenMorphuses an evolutionary algorithm to search foreffectivetest oracles, i.e., oracles that trigger no false alarms and expose software faults in the method under test. The proposed search algorithm is guided by two fitness functions that measure the number of false alarms and the number of missed faults for the generated MRs. Our results show thatGenMorphgenerates effective MRs for 18 out of 23 methods (mutation score >20%). Furthermore, it can increaseRandoop’s fault detection capability in 7 out of 23 methods, andEvosuite’s in 14 out of 23 methods. When compared with AUTOMR, a state-of-the-art MR generator,GenMorphalso outperformed its fault detection capability in 9 out of 10 methods. Jon Ayerdi, Valerio Terragni, Gunel Jahangirova, Aitor Arrieta, Paolo Tonella |
IEEE Trans. Software Eng. | 3 |
| 2023 | Repairing DNN Architecture: Are We There Yet?abstractAs Deep Neural Networks (DNNs) are rapidly being adopted within large software systems, software developers are increasingly required to design, train, and deploy such models into the systems they develop. Consequently, testing and improving the robustness of these models have received a lot of attention lately. However, relatively little effort has been made to address the difficulties developers experience when designing and training such models: if the evaluation of a model shows poor performance after the initial training, what should the developer change? We survey and evaluate existing state-of-the-art techniques that can be used to repair model performance, using a benchmark of both real-world mistakes developers made while designing DNN models and artificial faulty models generated by mutating the model code. The empirical evaluation shows that random baseline is comparable with or sometimes outperforms existing state-of-the-art techniques. However, for larger and more complicated models, all repair techniques fail to find fixes. Our findings call for further research to develop more sophisticated techniques for Deep Learning repair. Jinhan Kim, Nargiz Humbatova, Gunel Jahangirova, Paolo Tonella, Shin Yoo |
ICST | 3 |
| 2023 | Software testing in the machine learning era
Andrea Stocco 0001, Onn Shehory, Gunel Jahangirova, Vincenzo Riccio, Guy Barash, Eitan Farchi, Diptikalyan Saha |
Empir. Softw. Eng. | 3 |
| 2021 | Quality Metrics and Oracles for Autonomous Vehicles TestingabstractThe race for deploying AI-enabled autonomous vehicles (AVs) on public roads is based on the promise that such self-driving cars will be as safe as or safer than human drivers. Numerous techniques have been proposed to test AVs, which however lack oracle definitions that account for the quality of driving, due to the lack of a commonly used set of metrics. Towards filling this gap, we first performed a systematic analysis of the literature concerning the assessment of the quality of driving of human drivers and extracted 126 metrics. Then, we measured the correlation between such metrics and the human perception of driving quality when AVs are driving. Lastly, we performed a study based on mutation analysis to assess whether the 26 metrics that best capture the quality of AV driving according to the human study can be used as functional oracles. Our results, targeting the Udacity platform, indicate that our automated oracles can kill a high proportion of mutants at a zero or very low false alarm rate, and therefore can be used as effective functional oracles for the quality of driving of AVs. Gunel Jahangirova, Andrea Stocco 0001, Paolo Tonella |
ICST | 1 |
| 2021 | DeepCrime: mutation testing of deep learning systems based on real faultsabstractDeep Learning (DL) solutions are increasingly adopted, but how to test them remains a major open research problem. Existing and new testing techniques have been proposed for and adapted to DL systems, including mutation testing. However, no approach has investigated the possibility to simulate the effects of real DL faults by means of mutation operators. We have defined 35 DL mutation operators relying on 3 empirical studies about real faults in DL systems. We followed a systematic process to extract the mutation operators from the existing fault taxonomies, with a formal phase of conflict resolution in case of disagreement. We have implemented 24 of these DL mutation operators into DeepCrime, the first source-level pre-training mutation tool based on real DL faults. We have assessed our mutation operators to understand their characteristics: whether they produce interesting, i.e., killable but not trivial, mutations. Then, we have compared the sensitivity of our tool to the changes in the quality of test data with that of DeepMutation++, an existing post-training DL mutation tool. Nargiz Humbatova, Gunel Jahangirova, Paolo Tonella |
ISSTA | 2 |
| 2021 | DeepMetis: Augmenting a Deep Learning Test Set to Increase its Mutation ScoreabstractDeep Learning (DL) components are routinely integrated into software systems that need to perform complex tasks such as image or natural language processing. The adequacy of the test data used to test such systems can be assessed by their ability to expose artificially injected faults (mutations) that simulate real DL faults.In this paper, we describe an approach to automatically generate new test inputs that can be used to augment the existing test set so that its capability to detect DL mutations increases. Our tool DeepMetis implements a search based input generation strategy. To account for the non-determinism of the training and the mutation processes, our fitness function involves multiple instances of the DL model under test. Experimental results show that DeepMetis is effective at augmenting the given test set, increasing its capability to detect mutants by 63% on average. A leave-one-out experiment shows that the augmented test set is capable of exposing unseen mutants, which simulate the occurrence of yet undetected faults. Vincenzo Riccio, Nargiz Humbatova, Gunel Jahangirova, Paolo Tonella |
ASE | 3 |
| 2021 | Diversifying Focused Testing for Unit TestingabstractSoftware changes constantly, because developers add new features or modifications. This directly affects the effectiveness of the test suite associated with that software, especially when these new modifications are in a specific area that no test case covers. This article tackles the problem of generating a high-quality test suite to cover repeatedly a given point in a program, with the ultimate goal of exposing faults possibly affecting the given program point. Both search-based software testing and constraint solving offer ready, but low-quality, solutions to this: Ideally, a maximally diverse covering test set is required, whereas search and constraint solving tend to generate test sets with biased distributions. Our approach, Diversified Focused Testing (DFT), uses a search strategy inspired by GödelTest. We artificially inject parameters into the code branching conditions and use a bi-objective search algorithm to find diverse inputs by perturbing the injected parameters, while keeping the path conditions still satisfiable. Our results demonstrate that our technique, DFT, is able to cover a desired point in the code at least 90% of the time. Moreover, adding diversity improves the bug detection and the mutation killing abilities of the test suites. We show that DFT achieves better results than focused testing, symbolic execution, and random testing by achieving from 3% to 70% improvement in mutation score and up to 100% improvement in fault detection across 105 software subjects. Héctor D. Menéndez 0001, Gunel Jahangirova, Federica Sarro, Paolo Tonella, David Clark 0001 |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2021 | An Empirical Validation of Oracle ImprovementabstractWe propose a human-in-the-loop approach for oracle improvement and analyse whether the proposed oracle improvement process is helping developers to create better oracles. For this, we conducted two human studies with 68 participants overall: an oracle assessment study and an oracle improvement study. Our results show that developers exhibit poor performance (29 percent accuracy) when manually assessing whether an assertion oracle contains a false positive, a false negative or none of the two. This shows that automated detection of these oracle deficiencies is beneficial for the users. Our tool OASIs (Oracle ASsessment and Improvement) helps developers produce assertions with higher quality. Participants who used OASIs in the improvement study were able to achieve 33 percent of full and 67 percent of partial correctness as opposed to participants without the tool who achieved only 21 percent of full and 43 percent of partial correctness. Gunel Jahangirova, David Clark 0001, Mark Harman, Paolo Tonella |
IEEE Trans. Software Eng. | 1 |
| 2020 | Taxonomy of real faults in deep learning systemsabstractThe growing application of deep neural networks in safety-critical domains makes the analysis of faults that occur in such systems of enormous importance. In this paper we introduce a large taxonomy of faults in deep learning (DL) systems. We have manually analysed 1059 artefacts gathered from GitHub commits and issues of projects that use the most popular DL frameworks (TensorFlow, Keras and PyTorch) and from related Stack Overflow posts. Structured interviews with 20 researchers and practitioners describing the problems they have encountered in their experience have enriched our taxonomy with a variety of additional faults that did not emerge from the other two sources. Our final taxonomy was validated with a survey involving an additional set of 21 developers, confirming that almost all fault categories (13/15) were experienced by at least 50% of the survey participants. Nargiz Humbatova, Gunel Jahangirova, Gabriele Bavota, Vincenzo Riccio, Andrea Stocco 0001, Paolo Tonella |
ICSE | 2 |
| 2020 | An Empirical Evaluation of Mutation Operators for Deep Learning SystemsabstractDeep Learning (DL) is increasingly adopted to solve complex tasks such as image recognition or autonomous driving. Companies are considering the inclusion of DL components in production systems, but one of their main concerns is how to assess the quality of such systems. Mutation testing is a technique to inject artificial faults into a system, under the assumption that the capability to expose (kilt) such artificial faults translates into the capability to expose also real faults. Researchers have proposed approaches and tools (e.g., Deep-Mutation and MuNN) that make mutation testing applicable to deep learning systems. However, existing definitions of mutation killing, based on accuracy drop, do not take into account the stochastic nature of the training process (accuracy may drop even when re-training the un-mutated system). Moreover, the same mutation operator might be effective or might be trivial/impossible to kill, depending on its hyper-parameter configuration. We conducted an empirical evaluation of existing operators, showing that mutation killing requires a stochastic definition and identifying the subset of effective mutation operators together with the associated most effective configurations. Gunel Jahangirova, Paolo Tonella |
ICST | 1 |
| 2020 | Evolutionary improvement of assertion oraclesabstractAssertion oracles are executable boolean expressions placed inside the program that should pass (return true) for all correct executions and fail (return false) for all incorrect executions. Because designing perfect assertion oracles is difficult, assertions often fail to distinguish between correct and incorrect executions. In other words, they are prone to false positives and false negatives. In this paper, we propose GAssert (Genetic ASSERTion improvement), the first technique to automatically improve assertion oracles. Given an assertion oracle and evidence of false positives and false negatives, GAssert implements a novel co-evolutionary algorithm that explores the space of possible assertions to identify one with fewer false positives and false negatives. Our empirical evaluation on 34 Java methods from 7 different Java code bases shows that GAssert effectively improves assertion oracles. GAssert outperforms two baselines (random and invariant-based oracle improvement), and is comparable with and in some cases even outperformed human-improved assertions. Valerio Terragni, Gunel Jahangirova, Paolo Tonella, Mauro Pezzè |
ESEC/SIGSOFT FSE | 2 |
| 2020 | Testing machine learning based systems: a systematic mappingabstractAbstract Context: A Machine Learning based System (MLS) is a software system including one or more components that learn how to perform a task from a given data set. The increasing adoption of MLSs in safety critical domains such as autonomous driving, healthcare, and finance has fostered much attention towards the quality assurance of such systems. Despite the advances in software testing, MLSs bring novel and unprecedented challenges, since their behaviour is defined jointly by the code that implements them and the data used for training them. Objective: To identify the existing solutions for functional testing of MLSs, and classify them from three different perspectives: (1) the context of the problem they address, (2) their features, and (3) their empirical evaluation. To report demographic information about the ongoing research. To identify open challenges for future research. Method: We conducted a systematic mapping study about testing techniques for MLSs driven by 33 research questions. We followed existing guidelines when defining our research protocol so as to increase the repeatability and reliability of our results. Results: We identified 70 relevant primary studies, mostly published in the last years. We identified 11 problems addressed in the literature. We investigated multiple aspects of the testing approaches, such as the used/proposed adequacy criteria, the algorithms for test input generation, and the test oracles. Conclusions: The most active research areas in MLS testing address automated scenario/input generation and test oracle creation. MLS testing is a rapidly growing and developing research area, with many open challenges, such as the generation of realistic inputs and the definition of reliable evaluation metrics and benchmarks. Vincenzo Riccio, Gunel Jahangirova, Andrea Stocco 0001, Nargiz Humbatova, Michael Weiss 0004, Paolo Tonella |
Empir. Softw. Eng. | 2 |
| 2018 | OASIs: oracle assessment and improvement toolabstractThe oracle problem remains one of the key challenges in software testing, for which little automated support has been developed so far. We introduce OASIs, a search-based tool for Java that assists testers in oracle assessment and improvement. It does so by combining test case generation to reveal false positives and mutation testing to reveal false negatives. In this work, we describe how OASIs works, provide details of its implementation, and explain how it can be used in an iterative oracle improvement process with a human in the loop. Finally, we present a summary of previous empirical evaluation showing that the fault detection rate of the oracles after improvement using OASIs increases, on average, by 48.6%. Gunel Jahangirova, David Clark 0001, Mark Harman, Paolo Tonella |
ISSTA | 1 |
| 2017 | Oracle problem in software testingabstractThe oracle problem remains one of the key challenges in software testing, for which little automated support has been developed so far. In my thesis work we introduce a technique for assessing and improving test oracles by reducing the incidence of both false positives and false negatives. Our technique combines test case generation to reveal false positives and mutation testing to reveal false negatives. The experimental results on five real-world subjects show that the fault detection rate of the oracles after improvement increases, on average, by 48.6% (86% over the implicit oracle). Three actual, exposed faults in the studied systems were subsequently confirmed and fixed by the developers. However, our technique contains a human in the loop, which was represented only by the author during the initial experiments. Our next goal is to conduct further experiments where the human in the loop will be represented by real developers. Our second future goal is to address the oracle placement problem. When testing software, developers can place oracles externally or internally to a method. Given a faulty execution state, i.e., one that differs from the expected one, an oracle might be unable to expose the fault if it is placed at a program point with no access to the incorrect program state or where the program state is no longer corrupted. In such a case, the oracle is subject to failed error propagation. Internal oracles are in principle less subject to failed error propagation than external oracles. However, they are also more difficult to define manually. Hence, a key research question is whether a more intrusive oracle placement is justified by its higher fault detection capability. Gunel Jahangirova |
ISSTA | 1 |
| 2016 | Test oracle assessment and improvementabstractWe introduce a technique for assessing and improving test oracles by reducing the incidence of both false positives and false negatives. We prove that our approach can always result in an increase in the mutual information between the actual and perfect oracles. Our technique combines test case generation to reveal false positives and mutation testing to reveal false negatives. We applied the decision support tool that implements our oracle improvement technique to five real-world subjects. The experimental results show that the fault detection rate of the oracles after improvement increases, on average, by 48.6% (86% over the implicit oracle). Three actual, exposed faults in the studied systems were subsequently confirmed and fixed by the developers. Gunel Jahangirova, David Clark 0001, Mark Harman, Paolo Tonella |
ISSTA | 1 |