VLDB 2026 Research / reviewers in the wild / expert
Sira Vegas
dblp:78/6590
· DBLP profile ↗
46ranked-venue papers
10as first author
17since 2021 · last 2026
0000-0001-8535-9386ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 46 · 10 first-author · 17 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Stop Comparing Apples and Oranges: Matching for Better Results in Mining Software Repositories StudiesabstractConfounders (or confounding variables) pose significant challenges to detecting reliable causal relationships in observational studies. When data are collected from naturally occurring phenomena—e.g., mining software repositories (MSR)—researchers cannot rely on randomization to control for confounders, leading to biased causal inferences. Alternative approaches are required to mitigate confounding bias when exploring causal inferences. This paper explains and exemplifies the use of matching in MSR. Sabato Nocera, Nyyti Saarimäki, Valentina Lenarduzzi, Davide Taibi 0001, Sira Vegas |
MSR | 5 |
| 2026 | Building and validating deep learning models for forecasting the quality of cloud servicesabstractAbstract Cloud services operate in highly dynamic and heterogeneous environments, requiring continuous and accurate assessment of service quality. While Quality of Service (QoS) models are widely used to monitor performance, deep learning (DL) architectures–such as Long Short-Term Memory (LSTM) and Bidirectional Gated Recurrent Units (BI-GRU)–offer enhanced capabilities for forecasting potential Service Level Agreement (SLA) violations. However, many existing experiments in this domain suffer from methodological shortcomings, including the use of outdated or proprietary datasets, a narrow set of QoS metrics, incomplete documentation of model architectures and training procedures, and a lack of statistical rigor, which undermines reproducibility and applicability in industrial contexts. This study empirically compares the performance of BI-GRU, LSTM, and AutoRegressive Integrated Moving Average (ARIMA) models for QoS forecasting using a rigorously designed experimental protocol that addresses these limitations. We build a multi-metric QoS dataset covering five months of operational data from a cloud service in an IT company, comprising 16 QoS metrics. Forecasting models were trained and evaluated using Root Mean Squared Error (RMSE), Mean Absolute Error (MAE), and Mean Absolute Percentage Error (MAPE), with training time considered as an efficiency indicator. BI-GRU outperformed ARIMA across all QoS metrics and achieved statistically significant improvements over LSTM in 9 out of 16 metrics. In contrast, LSTM only significantly outperformed BI-GRU in one metric. Our findings demonstrate that most BI-GRU models provide superior accuracy and efficiency. Furthermore, the methodological rigor of the experimental design supports their applicability for proactive QoS management and informed decision-making in industrial cloud service environments. Ximena Guerron, Marta Fernández-Diego, Silvia Abrahão, Emilio Insfrán, Sira Vegas |
Autom. Softw. Eng. | 5 |
| 2025 | Software Composition Analysis and Supply Chain Security in Apache Projects: an Empirical StudyabstractA software supply chain consists of anything needed to develop and deliver a software project, including (third-party) components. Software Composition Analysis (SCA) allows for managing the security of software supply chains by identifying such components and their (security) vulnerabilities. The main goal of the empirical study presented in this paper is to investigate the effects of adopting/using over time an SCA tool like OWASP Dependency-Check (OWASP DC) in the context of the security of the software supply chain. To this end, following a cohort design, we analyzed the vulnerabilities affecting the components of the open-source (OS) Java Maven projects owned by the Apache Software Foundation (ASF) and publicly hosted on GitHub. These projects could adopt (or not) OWASP DC. The results indicate that the adoption of OWASP DC appears to be causing a significant reduction in the overall number/score of vulnerabilities, including those with a high Common Vulnerability Scoring System (CVSS) severity level. The use of OWASP DC also increased the vulnerabilities with a low severity level. Our results seem to encourage practitioners to adopt SCA to improve the security of their software supply chains. Sabato Nocera, Sira Vegas, Giuseppe Scanniello, Natalia Juristo Juzgado |
MSR | 2 |
| 2025 | Does microservice adoption impact the velocity? A cohort studyabstractAbstract [Context] Microservices enable the decomposition of applications into small, independent, and connected services. The independence between services could positively affect a project’s velocity, which is considered an important maintenance metric measuring the time taken to implement features and fix bugs. However, no studies have investigated the causal relationship between microservices and velocity. [Objective and Method] The goal of this study is to investigate the effect of microservices on velocity which is a common maintenance metric. The study compares projects on GitHub developed with microservices style from the beginning and similar projects using monolithic architectures. The study was conducted as a retrospective cohort study, which is a study type used to assess causality. [Results] The results did not find statistically significant differences in mean velocities in microservice-based and monolithic projects. Furthermore, the statistical adjustment performed to quantify the statistical impact of the use of microservices on velocity considering additional confounders did not find statistically significant impact from these. [Conclusions] The results did not indicate a difference between microservices-based projects and monolithic projects in terms of velocity. In addition, this study will contribute to the body of knowledge of empirical methods and be among the first works to adopt the methodology of the cohort study. Nyyti Saarimäki, Mikel Robredo, Valentina Lenarduzzi, Sira Vegas, Natalia Juristo Juzgado, Davide Taibi 0001 |
Empir. Softw. Eng. | 4 |
| 2025 | Relevant Information in TDD Experiment ReportingabstractExperiments are a commonly used method of research in software engineering (SE). Researchers report their experiments following detailed guidelines. However, researchers do not, in the field of test-driven development (TDD) at least, specify how they operationalized the response variables and, particularly, the measurement process. This article has three aims: (i) identify the response variable operationalization components in TDD experiments that study external quality; (ii) study their influence on the experimental results; (iii) determine if the experiment reports describe the measurement process components that have an impact on the results. We used two-part sequential mixed methods research. The first part of the research adopts a quantitative approach applying a statistical analysis of the impact of the operationalization components on the experimental results. The second part follows with a qualitative approach applying a systematic mapping study (SMS). The test suites, intervention types and measurers have an influence on the measurements and results of the statistical analysis of TDD experiments in SE. The test suites have a major impact on both the measurements and the results of the experiments. The intervention type has less impact on the results than on the measurements. While the measurers have an impact on the measurements, this is not transferred to the experimental results. On the other hand, the results of our SMS confirm that TDD experiments do not usually report either the test suites, the test case generation method, or the details of how external quality was measured. A measurement protocol should be used to ensure that the measurements made by different measurers are similar. It is necessary to report the test cases, the experimental task and the intervention type in order to be able to reproduce the measurements and statistical analyses, as well as to replicate experiments and build dependable families of experiments. Fernando Uyaguari, Silvia Teresita Acuña, John W. Castro, Davide Fucci, Óscar Dieste Tubío, Sira Vegas |
ACM Trans. Softw. Eng. Methodol. | 6 |
| 2024 | Crossover Designs in Software Engineering Experiments: Review of the State of AnalysisabstractExperimentation is an essential method for causal inference in any empirical discipline. Crossover-design experiments are common in Software Engineering (SE) research. In these, subjects apply more than one treatment in different orders. This design increases the amount of obtained data and deals with subject variability but introduces threats to internal validity like the learning and carryover effect. Vegas et al. reviewed the state of practice for crossover designs in SE research and provided guidelines on how to address its threats during data analysis while still harnessing its benefits. In this paper, we reflect on the impact of these guidelines and review the state of analysis of crossover-design experiments in SE publications between 2015 and March 2024. To this end, by conducting a forward snowballing of the guidelines, we survey 136 publications reporting 67 crossover-design experiments and evaluate their data analysis against the provided guidelines. The results show that the validity of data analyses has improved compared to the original state of analysis. Still, despite the explicit guidelines, only 29.5% of all threats to validity were addressed properly. While the maturation and the optimal sequence threats are properly addressed in 35.8% and 38.8% of all studies in our sample respectively, the carryover threat is only modeled in about 3% of the observed cases. The lack of adherence to the analysis guidelines threatens the validity of the conclusions drawn from crossover-design experiments. Julian Frattini, Davide Fucci, Sira Vegas |
ESEM | 3 |
| 2024 | Evidence-Based Commit Message Generation with Deep Learning Techniques (EvidenCoM)abstractContext. Automatic code commit message generation tools using deep learning techniques have gained significant attention in recent years. Although many experiments with these tools have been reported, there is no current evidence on which of them performs best. Objective. This project aims to provide evidence on the set of proposed tools. Method. We will conduct a secondary study, consisting of a systematic review that includes evidence synthesis. Results. We have identified the set of primary studies and extracted the information from them. The next steps include obtaining evidence from the extracted information. Sira Vegas, Xavier Ferré, Hongming Zhu |
ESEM | 1 |
| 2024 | Cohort Studies for Mining Software RepositoriesabstractMining Software Repositories studies have become increasingly popular over the years. However, a notable limitation is that they report correlational relationships rather than establishing causation. In contrast, certain disciplines (e.g. epidemiology) have developed specific methods to address this limitation. The goal of this tutorial is to introduce participants to one such method: cohort studies. By the end of the tutorial, participants will be familiar with the steps and techniques involved in designing and analyzing cohort studies. Nyyti Saarimäki, Sira Vegas, Valentina Lenarduzzi, Davide Taibi 0001, Mikel Robredo |
MSR | 2 |
| 2023 | Comparing 2D and Augmented Reality Visualizations for Microservice System Understandability: A Controlled ExperimentabstractMicroservice-based systems are often complex to understand, especially when their sizes grow. Abstracted views help practitioners with the system understanding from a certain perspective. Recent advancement in interactive data visualization begs the question of whether established software engineering models to visualize system design remain the most suited approach for the service-oriented design of microservices. Our recent work proposed presenting a 3D visualization for microservices in augmented reality. This paper analyzes whether such an approach brings any benefits to practitioners when dealing with selected architectural questions related to system design quality. For this purpose, we conducted a controlled experiment involving 20 participants investigating their performance in identifying service dependency, service cardinality, and bottlenecks. Results show that the 3D enables novices to perform as well as experts in the detection of service dependencies, especially in large systems, while no differences are reported for the identification of service cardinality and bottlenecks. We recommend industry and researchers to further investigate AR for microservice architectural analysis, especially to ease the onboarding of new developers in microservice projects. Amr S. Abdelfattah, Tomás Cerný, Davide Taibi 0001, Sira Vegas |
ICPC | 4 |
| 2023 | Pitfalls in Experiments with DNN4SE: An Analysis of the State of the PracticeabstractSoftware engineering (SE) techniques are increasingly relying on deep learning approaches to support many SE tasks, from bug triaging to code generation. To assess the efficacy of such techniques researchers typically perform controlled experiments. Conducting these experiments, however, is particularly challenging given the complexity of the space of variables involved, from specialized and intricate architectures and algorithms to a large number of training hyper-parameters and choices of evolving datasets, all compounded by how rapidly the machine learning technology is advancing, and the inherent sources of randomness in the training process. In this work we conduct a mapping study, examining 194 experiments with techniques that rely on deep neural networks (DNNs) appearing in 55 papers published in premier SE venues to provide a characterization of the state of the practice, pinpointing experiments’ common trends and pitfalls. Our study reveals that most of the experiments, including those that have received ACM artifact badges, have fundamental limitations that raise doubts about the reliability of their findings. More specifically, we find: 1) weak analyses to determine that there is a true relationship between independent and dependent variables (87% of the experiments), 2) limited control over the space of DNN relevant variables, which can render a relationship between dependent variables and treatments that may not be causal but rather correlational (100% of the experiments), and 3) lack of specificity in terms of what are the DNN variables and their values utilized in the experiments (86% of the experiments) to define the treatments being applied, which makes it unclear whether the techniques designed are the ones being assessed, or how the sources of extraneous variation are controlled. We provide some practical recommendations to address these limitations. Sira Vegas, Sebastian G. Elbaum |
ESEC/SIGSOFT FSE | 1 |
| 2022 | Tutorial 3: Pitfalls in the Measurement Methods Applied in Experimental Software Engineering - Assessment and Suggestions for ImprovementabstractMeasurement is an essential issue in empirical software engineering. It is subject to different sources of error that must be kept as small as possible. Measuring instruments is one of these sources of error. In this tutorial, we provide awareness of potential pitfalls in the measurement methods—specifically measuring instruments—applied in empirical software engineering, describing statistical techniques that can be used for measures assessment, and making recommendations to improve the measurement practice in software engineering. Test suites are used as measuring instruments in many software engineering experiments. We will use the case of test suites when measuring the external quality of the code developed by participants of TDD-related experiments. Óscar Dieste Tubío, Sira Vegas |
EASE | 2 |
| 2021 | Towards a Methodology for Participant Selection in Software Engineering Experiments: A Vision of the FutureabstractBackground. Software Engineering (SE) researchers extensively perform experiments with human subjects. Well-defined samples are required to ensure external validity. Samples are selected purposely or by convenience, limiting the generalizability of results. Objective. We aim to depict the current status of participants selection in empirical SE, identifying the main threats and how they are mitigated. We draft a robust approach to participants' selection. Method. We reviewed existing participants' selection guidelines in SE, and performed a preliminary literature review to find out how participants' selection is conducted in SE in practice. Results. We outline a new selection methodology, by 1) defining the characteristics of the desired population, 2) locating possible sources of sampling available for researchers, and 3) identifying and reducing the "distance" between the selected sample and its corresponding population. Conclusion. We propose a roadmap to develop and empirically validate the selection methodology. Valentina Lenarduzzi, Óscar Dieste Tubío, Davide Fucci, Sira Vegas |
ESEM | 4 |
| 2021 | A family of experiments on test-driven development
Adrián Santos, Sira Vegas, Óscar Dieste Tubío, Fernando Uyaguari, Ayse Tosun Misirli, Davide Fucci, Burak Turhan, Giuseppe Scanniello, Simone Romano 0001, Itir Karac, Marco Kuhrmann, Vladimir Mandic, Robert Ramac, Dietmar Pfahl, Christian Engblom, Jarno Kyykka, Kerli Rungi, Carolina Palomeque, Jaroslav Spisak, Markku Oivo, Natalia Juristo Juzgado |
Empir. Softw. Eng. | 2 |
| 2021 | Comparing the results of replications in software engineering
Adrián Santos, Sira Vegas, Markku Oivo, Natalia Juristo Juzgado |
Empir. Softw. Eng. | 2 |
| 2021 | Evaluating Model-Driven Development Claims with Respect to Quality: A Family of ExperimentsabstractContext: There is a lack of empirical evidence on the differences between model-driven development (MDD), where code is automatically derived from conceptual models, and traditional software development method, where code is manually written. In our previous work, we compared both methods in a baseline experiment concluding that quality of the software developed following MDD was significantly better only for more complex problems (with more function points). Quality was measured through test cases run on a functional system. Objective: This paper reports six replications of the baseline to study the impact of problem complexity on software quality in the context of MDD. Method: We conducted replications of two types: strict replications and object replications. Strict replications were similar to the baseline, whereas we used more complex experimental objects (problems) in the object replications. Results: MDD yields better quality independently of problem complexity with a moderate effect size. This effect is bigger for problems that are more complex. Conclusions: Thanks to the bigger size of the sample after aggregating replications, we discovered an effect that the baseline had not revealed due to the small sample size. The baseline results hold, which suggests that MDD yields better quality for more complex problems. José Ignacio Panach, Óscar Dieste Tubío, Beatriz Marín, Sergio España 0001, Sira Vegas, Oscar Pastor 0001, Natalia Juristo Juzgado |
IEEE Trans. Software Eng. | 5 |
| 2021 | A Procedure and Guidelines for Analyzing Groups of Software Engineering ReplicationsabstractContext: Researchers from different groups and institutions are collaborating on building groups of experiments by means of replication (i.e., conducting groups of replications). Disparate aggregation techniques are being applied to analyze groups of replications. The application of unsuitable techniques to aggregate replication results may undermine the potential of groups of replications to provide in-depth insights from experiment results. Objectives: Provide an analysis procedure with a set of embedded guidelines to aggregate software engineering (SE) replication results. Method: We compare the characteristics of groups of replications for SE and other mature experimental disciplines such as medicine and pharmacology. In view of their differences, the limitations with regard to the joint data analysis of groups of SE replications and the guidelines provided in mature experimental disciplines to analyze groups of replications, we build an analysis procedure with a set of embedded guidelines specifically tailored to the analysis of groups of SE replications. We apply the proposed analysis procedure to a representative group of SE replications to illustrate its use. Results: All the information contained within the raw data should be leveraged during the aggregation of replication results. The analysis procedure that we propose encourages the use of stratified individual participant data and aggregated data in tandem to analyze groups of SE replications. Conclusion: The aggregation techniques used to analyze groups of replications should be justified in research articles. This will increase the reliability and transparency of joint results. The proposed guidelines should ease this endeavor. Adrián Santos, Sira Vegas, Markku Oivo, Natalia Juristo Juzgado |
IEEE Trans. Software Eng. | 2 |
| 2021 | Investigating the Impact of Development Task on External Quality in Test-Driven Development: An Industry ExperimentabstractReviews on test-driven development (TDD) studies suggest that the conflicting results reported in the literature are due to unobserved factors, such as the tasks used in the experiments, and highlight that there are very few industry experiments conducted with professionals. The goal of this study is to investigate the impact of a new factor, the chosentask, and thedevelopment approachon external quality in an industrial experimental setting with 17 professionals. The participants are junior to senior developers in programming with Java, beginner to novice in unit testing, JUnit, and they have no prior experience in TDD. The experimental design is a$2\times 2$cross-over, i.e., we use two tasks for each of the two approaches, namely TDD and incremental test-last development (ITLD). Our results reveal that bothdevelopment approachandtaskare significant factors with regards to the external quality achieved by the participants. More specifically, the participants produce higher quality code during ITLD in which splitting user stories into subtasks, coding, and testing activities are followed, compared to TDD. The results also indicate that the participants produce higher quality code during the implementation of Bowling Score Keeper, compared to that of Mars Rover API, although they perceived both tasks as of similar complexity. An interaction between thedevelopment approachandtaskcould not be observed in this experiment. We conclude that variables that have not been explored so often, such as the extent to which the task is specified in terms of smaller subtasks, and developers’ unit testing experience might be critical factors in TDD experiments. The real-world appliance of TDD and its implications on external quality still remain to be challenging unless these uncontrolled and unconsidered factors are further investigated by researchers in both academic and industrial settings. Ayse Tosun Misirli, Óscar Dieste Tubío, Sira Vegas, Dietmar Pfahl, Kerli Rungi, Natalia Juristo Juzgado |
IEEE Trans. Software Eng. | 3 |
| 2020 | Cohort Studies in Software Engineering: A Vision of the FutureabstractBackground. Most Mining Software Repositories (MSR) studies cannot obtain causal relations because they are not controlled experiments. The use of cohort studies as defined in epidemiology could help to overcome this shortcoming. Nyyti Saarimäki, Valentina Lenarduzzi, Sira Vegas, Natalia Juristo Juzgado, Davide Taibi 0001 |
ESEM | 3 |
| 2020 | On (Mis)perceptions of testing effectiveness: an empirical study
Sira Vegas, Patricia Riofrío, Esperanza Marcos, Natalia Juristo Juzgado |
Empir. Softw. Eng. | 1 |
| 2020 | Impact of usability mechanisms: An experiment on efficiency, effectiveness and user satisfaction
Juan M. Ferreira, Silvia Teresita Acuña, Óscar Dieste Tubío, Sira Vegas, Adrián Santos, Francy D. Rodríguez, Natalia Juristo Juzgado |
Inf. Softw. Technol. | 4 |
| 2020 | Increasing validity through replication: an illustrative TDD caseabstractAbstract Software engineering (SE) experiments suffer from threats to validity that may impact their results. Replication allows researchers building on top of previous experiments’ weaknesses and increasing the reliability of the findings. Illustrating the benefits of replication to increase the reliability of the findings and uncover moderator variables. We replicate an experiment on test-driven development (TDD) and address some of its threats to validity and those of a previous replication. We compare the replications’ results and hypothesize on plausible moderators impacting results. Differences across TDD replications’ results might be due to the operationalization of the response variables, the allocation of subjects to treatments, the allowance to work outside the laboratory, the provision of stubs, or the task. Replications allow examining the robustness of the findings, hypothesizing on plausible moderators influencing results, and strengthening the evidence obtained. Adrián Santos, Sira Vegas, Fernando Uyaguari, Óscar Dieste Tubío, Burak Turhan, Natalia Juristo Juzgado |
Softw. Qual. J. | 2 |
| 2019 | A controlled experiment on time pressure and confirmation bias in functional software testingabstractConfirmation bias is a person’s tendency to look for evidence that strengthens his/her prior beliefs rather than refutes them. Manifestation of confirmation bias in software testing may have adverse effects on software quality. Psychology research suggests that time pressure could trigger confirmation bias. In the software industry, this phenomenon may deteriorate software quality. In this study, we investigate whether testers manifest confirmation bias and how it is affected by time pressure in functional software testing. We performed a controlled experiment with 42 graduate students to assess manifestation of confirmation bias in terms of the conformity of their designed test cases to the provided requirements specification. We employed a one factor with two treatments between-subjects experimental design. We observed, overall, participants designed significantly more confirmatory test cases as compared to disconfirmatory ones, which is in line with previous research. However, we did not observe time pressure as an antecedent to an increased rate of confirmatory testing behaviour. People tend to design confirmatory test cases regardless of time pressure. For practice, we find it necessary that testers develop self-awareness of confirmation bias and counter its potential adverse effects with a disconfirmatory attitude. We recommend further replications to investigate the effect of time pressure as a potential contributor to the manifestation of confirmation bias. Iflaah Salman, Burak Turhan, Sira Vegas |
Empir. Softw. Eng. | 3 |
| 2019 | Adopting configuration management principles for managing experiment materials in families of experiments
Edison G. Espinosa, Silvia Teresita Acuña, Sira Vegas, Natalia Juristo Juzgado |
Inf. Softw. Technol. | 3 |
| 2018 | Content and structure of laboratory packages for software engineering experiments
Martín Solari, Sira Vegas, Natalia Juristo Juzgado |
Inf. Softw. Technol. | 2 |
| 2017 | An industry experiment on the effects of test-driven development on external quality and productivity
Ayse Tosun Misirli, Óscar Dieste Tubío, Davide Fucci, Sira Vegas, Burak Turhan, Hakan Erdogmus, Adrián Santos, Markku Oivo, Kimmo Toro, Janne Järvinen, Natalia Juristo Juzgado |
Empir. Softw. Eng. | 4 |
| 2016 | A Multi-Site Joint Replication of a Design Patterns Experiment Using Moderator Variables to Generalize across ContextsabstractContext.Several empirical studies have explored the benefits of software design patterns, but their collective results are highly inconsistent. Resolving the inconsistencies requires investigating moderators—i.e., variables that cause an effect to differ across contexts.Objectives.Replicate a design patterns experiment at multiple sites and identify sufficient moderators to generalize the results across prior studies.Methods.We perform a close replication of an experiment investigating the impact (in terms of time and quality) of design patterns (Decorator and Abstract Factory) on software maintenance. The experiment was replicated once previously, with divergent results. We execute our replication at four universities—spanning two continents and three countries—using a new method for performing distributed replications based on closely coordinated, small-scale instances (“joint replication”). We perform two analyses: 1) apost-hocanalysis of moderators, based on frequentist and Bayesian statistics; 2) ana priorianalysis of the original hypotheses, based on frequentist statistics.Results.The main effect differs across the previous instances of the experiment and across the sites in our distributed replication. Our analysis of moderators (including developer experience and pattern knowledge) resolves the differences sufficiently to allow for cross-context (and cross-study) conclusions. The final conclusions represent 126 participants from five universities and 12 software companies, spanning two continents and at least four countries.Conclusions.The Decorator pattern is found to be preferable to a simpler solution during maintenance, as long as the developer has at least some prior knowledge of the pattern. For Abstract Factory, the simpler solution is found to be mostly equivalent to the pattern solution. Abstract Factory is shown to require a higher level of knowledge and/or experience than Decorator for the pattern to be beneficial. Jonathan L. Krein, Lutz Prechelt, Natalia Juristo Juzgado, Aziz Nanthaamornphong, Jeffrey C. Carver, Sira Vegas, Charles D. Knutson, Kevin D. Seppi, Dennis Eggett |
IEEE Trans. Software Eng. | 6 |
| 2016 | Crossover Designs in Software Engineering Experiments: Benefits and PerilsabstractIn experiments with crossover design subjects apply more than one treatment. Crossover designs are widespread in software engineering experimentation: they require fewer subjects and control the variability among subjects. However, some researchers disapprove of crossover designs. The main criticisms are: the carryover threat and its troublesome analysis. Carryover is the persistence of the effect of one treatment when another treatment is applied later. It may invalidate the results of an experiment. Additionally, crossover designs are often not properly designed and/or analysed, limiting the validity of the results. In this paper, we aim to make SE researchers aware of the perils of crossover experiments and provide risk avoidance good practices. We study how another discipline (medicine) runs crossover experiments. We review the SE literature and discuss which good practices tend not to be adhered to, giving advice on how they should be applied in SE experiments. We illustrate the concepts discussed analysing a crossover experiment that we have run. We conclude that crossover experiments can yield valid results, provided they are properly designed and analysed, and that, if correctly addressed, carryover is no worse than other validity threats. Sira Vegas, Cecilia Apa, Natalia Juristo Juzgado |
IEEE Trans. Software Eng. | 1 |
| 2014 | A systematic mapping study on testing technique experiments: has the situation changed since 2000?abstractContext: Juristo et al. [7] published a literature review about testing technique experiments. The goal was to provide a picture of which techniques and aspects of techniques had been studied experimentally, and try to compile a body of knowledge on testing techniques. Goal: In this paper, we extend Juristo et al.'s study to cover the years from 2000 (where it ended) until 2013. Method: We have performed a systematic mapping study. Results: The situation in testing experimentation has not changed since Juristo et al.'s study. Conclusions: The research field has the same shortcomings. Jorge E. González, Natalia Juristo Juzgado, Sira Vegas |
ESEM | 3 |
| 2014 | Replications of software engineering experiments
Jeffrey C. Carver, Natalia Juristo Juzgado, Maria Teresa Baldassarre, Sira Vegas |
Empir. Softw. Eng. | 4 |
| 2014 | Understanding replication of experiments in software engineering: A classification
Omar S. Gómez, Natalia Juristo Juzgado, Sira Vegas |
Inf. Softw. Technol. | 3 |
| 2013 | A process for managing interaction between experimenters to get useful similar replications
Natalia Juristo Juzgado, Sira Vegas, Martín Solari, Silvia Abrahão, Isabel Ramos 0002 |
Inf. Softw. Technol. | 2 |
| 2013 | Determining the effectiveness of three software evaluation techniques through informal aggregation
Babatunde Kazeem Olorisade, Sira Vegas, Natalia Juristo Juzgado |
Inf. Softw. Technol. | 2 |
| 2012 | Comparing the Effectiveness of Equivalence Partitioning, Branch Testing and Code Reading by Stepwise Abstraction Applied by SubjectsabstractSome verification and validation techniques have been evaluated both theoretically and empirically. Most empirical studies have been conducted without subjects, passing over any effect testers have when they apply the techniques. We have run an experiment with students to evaluate the effectiveness of three verification and validation techniques (equivalence partitioning, branch testing and code reading by stepwise abstraction). We have studied how well able the techniques are to reveal defects in three programs. We have replicated the experiment eight times at different sites. Our results show that equivalence partitioning and branch testing are equally effective and better than code reading by stepwise abstraction. The effectiveness of code reading by stepwise abstraction varies significantly from program to program. Finally, we have identified project contextual variables that should be considered when applying any verification and validation technique or to choose one particular technique. Natalia Juristo Juzgado, Sira Vegas, Martín Solari, Silvia Abrahão, Isabel Ramos 0002 |
ICST | 2 |
| 2011 | The role of non-exact replications in software engineering experiments
Natalia Juristo Juzgado, Sira Vegas |
Empir. Softw. Eng. | 2 |
| 2010 | Replications types in experimental disciplinesabstractExperiment replication is a key component of the scientific paradigm. The purpose of replication is to verify previously observed findings. Although some Software Engineering (SE) experiments have been replicated, yet, there is still disagreement about how replications should be run in our field. With the aim of gaining a better understanding of how replications are carried out, this paper examines different replication types in other scientific disciplines. We believe that by analysing the replication types proposed in other disciplines it is possible to clarify some of the question marks still hanging over experimental SE replication. Omar S. Gómez, Natalia Juristo Juzgado, Sira Vegas |
ESEM | 3 |
| 2010 | Using differences among replications of software engineering experiments to gain knowledgeabstractIn no science or engineering discipline does it make sense to speak of isolated experiments. The results of a single experiment cannot be viewed as representative of the underlying reality. The concept of experiment is closely related to replication. Experiment replication is the repetition of an experiment to double-check its results. Multiple replications of an experiment increase the credibility of its results. Software engineering has tried its hand at the identical repetition of experiments in the way of the natural sciences (physics, chemistry, etc.). After numerous attempts over the years, excepting experiments repeated by the same researchers at the same site, no exact replications have yet been achieved. One key reason for this is the complexity of the software development setting. This complexity prevents the many experimental conditions from being reproduced identically. This paper reports research into whether non-exact replications can be of any use. We propose a process that allows researchers to generate new knowledge when running non-exact replications. To illustrate the advantages of the proposed process, two different replications of an experiment are shown. Natalia Juristo Juzgado, Sira Vegas |
MSR | 2 |
| 2009 | Using differences among replications of software engineering experiments to gain knowledgeabstractIn no science or engineering discipline does it make sense to speak of isolated experiments. The results of a single experiment cannot be viewed as representative of the underlying reality. The concept of experiment is closely related to replication. Experiment replication is the repetition of an experiment to double-check its results. Multiple replications of an experiment increase the credibility of its results. Software engineering has tried its hand at the identical repetition of experiments in the way of the natural sciences (physics, chemistry, etc.). After numerous attempts over the years, excepting experiments repeated by the same researchers at the same site, no exact replications have yet been achieved. One key reason for this is the complexity of the software development setting. This complexity prevents the many experimental conditions from being reproduced identically. This paper reports research into whether non-exact replications can be of any use. We propose a process that allows researchers to generate new knowledge when running non-exact replications. To illustrate the advantages of the proposed process, two different replications of an experiment are shown. Natalia Juristo Juzgado, Sira Vegas |
ESEM | 2 |
| 2009 | 11th International Workshop on Learning Software Organizations (LSO 2009) New Media in Transfer and Innovation
Andreas Jedlitschka, Sira Vegas |
PROFES | 2 |
| 2009 | Maturing Software Engineering Knowledge through Classifications: A Case Study on Unit Testing TechniquesabstractClassification makes a significant contribution to advancing knowledge in both science and engineering. It is a way of investigating the relationships between the objects to be classified and identifies gaps in knowledge. Classification in engineering also has a practical application; it supports object selection. They can help mature software engineering knowledge, as classifications constitute an organized structure of knowledge items. Till date, there have been few attempts at classifying in software engineering. In this research, we examine how useful classifications in software engineering are for advancing knowledge by trying to classify testing techniques. The paper presents a preliminary classification of a set of unit testing techniques. To obtain this classification, we enacted a generic process for developing useful software engineering classifications. The proposed classification has been proven useful for maturing knowledge about testing techniques, and therefore, SE, as it helps to: 1) provide a systematic description of the techniques, 2) understand testing techniques by studying the relationships among techniques (measured in terms of differences and similarities), 3) identify potentially useful techniques that do not yet exist by analyzing gaps in the classification, and 4) support practitioners in testing technique selection by matching technique characteristics to project characteristics. Sira Vegas, Natalia Juristo Juzgado, Victor R. Basili |
IEEE Trans. Software Eng. | 1 |
| 2008 | The role of replications in Empirical Software Engineering
Forrest Shull, Jeffrey C. Carver, Sira Vegas, Natalia Juristo Juzgado |
Empir. Softw. Eng. | 3 |
| 2006 | Packaging experiences for improving testing technique selection
Sira Vegas, Natalia Juristo Juzgado, Victor R. Basili |
J. Syst. Softw. | 1 |
| 2005 | A Characterisation Schema for Software Testing Techniques
Sira Vegas, Victor R. Basili |
Empir. Softw. Eng. | 1 |
| 2004 | Reviewing 25 Years of Testing Technique Experiments
Natalia Juristo Juzgado, Ana María Moreno 0001, Sira Vegas |
Empir. Softw. Eng. | 3 |
| 2003 | Best papers on Software Engineering from the SEKE'01 Conference
Sira Vegas, Marcelo Estayno |
J. Syst. Softw. | 1 |
| 2002 | What Information is Relevant When Selecting Software Testing Techniques?abstractOne of the main problems in software testing is the development of a suitable set of test cases so that the effectiveness of the test is maximised with a minimum number of test cases. A lot of testing techniques are now available for developing test cases. However, some of them are misused, others are never used and only a few are applied again and again. When developers have to decide what testing techniques(s) they should use in a project, they have little (if any) experiential information about the available testing techniques, their usefulness and, in general, how suited they are to the project. This paper presents the results of developing a characterization scheme for test technique selection. When instantiated for different techniques, the scheme should provide developers with enough information for choosing the best suited to their project. Thus, their decisions would be based on sound knowledge of the techniques, instead of perceptions, suppositions and assumptions. Sira Vegas, Natalia Juristo Juzgado, Victor R. Basili |
Int. J. Softw. Eng. Knowl. Eng. | 1 |
| 2001 | What Information is Relevant when Selecting Testing Techniques?
Sira Vegas |
SEKE | 1 |