VLDB 2026 Research / reviewers in the wild / expert
Óscar Dieste Tubío
dblp:24/1437 · also Oscar Dieste
· DBLP profile ↗
50ranked-venue papers
14as first author
13since 2021 · last 2026
0000-0002-3060-7853ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 50 · 14 first-author · 13 since 2021Artificial intelligence and machine learning · 3 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Visibility of Domain Elements in the Elicitation Process Interviews: A Family of Empirical Studies
Alejandrina Aranda, Óscar Dieste Tubío, José Ignacio Panach, Natalia Juristo Juzgado |
IEEE Trans. Software Eng. | 2 |
| 2025 | Investigation of the Activities Performed by Experimental Researchers in a Software Engineering Lab Using an Ethnographic MethodologyabstractContext Replication plays a critical role in building cumulative scientific knowledge. However, in the context of Empirical Software Engineering (ESE), replication efforts face persistent difficulties, both technical and methodological, which hinder reproducibility and generalizability. We explored whether the problems were due to formal issues, such as under‐specification or miscommunication, or intrinsic reasons, i.e., the existing replication procedures may not meet the researchers’ needs. Objective To understand how ESE researchers conduct experiments in real‐world settings. Method We conducted an ethnographic study with an experimental software engineering group, using interviews, observations, and analysis of internal documentation. Results We have created conceptual and process models representing experimentation, replication, and synthesis in the target research group. These models align with mainstream procedures at a high level but often break down in practice during specialized tasks such as coordination or documentation. The experimental process observed in the group differs from common assumptions and textbook descriptions in terms of (1) the number and diversity of activities involved, (2) the presence of differentiated roles, (3) the granularity of conceptual elements, and (4) the varying perspectives across subareas or families of experiments. Conclusions The discrepancies between actual laboratory practices and idealized process models may hinder knowledge transfer and complicate replication. We plan to extend this research by involving additional research groups, for instance, through surveys, to examine the generalizability of our findings. Efraín R. Fonseca C., Marta López Fernández, Óscar Dieste Tubío, Natalia Juristo Juzgado |
IET Softw. | 3 |
| 2025 | Reliability of systematic literature reviews on test-driven development
Fernando Uyaguari, Silvia Teresita Acuña, John W. Castro, Óscar Dieste Tubío, Natalia Juristo Juzgado |
Inf. Softw. Technol. | 4 |
| 2025 | Relevant Information in TDD Experiment ReportingabstractExperiments are a commonly used method of research in software engineering (SE). Researchers report their experiments following detailed guidelines. However, researchers do not, in the field of test-driven development (TDD) at least, specify how they operationalized the response variables and, particularly, the measurement process. This article has three aims: (i) identify the response variable operationalization components in TDD experiments that study external quality; (ii) study their influence on the experimental results; (iii) determine if the experiment reports describe the measurement process components that have an impact on the results. We used two-part sequential mixed methods research. The first part of the research adopts a quantitative approach applying a statistical analysis of the impact of the operationalization components on the experimental results. The second part follows with a qualitative approach applying a systematic mapping study (SMS). The test suites, intervention types and measurers have an influence on the measurements and results of the statistical analysis of TDD experiments in SE. The test suites have a major impact on both the measurements and the results of the experiments. The intervention type has less impact on the results than on the measurements. While the measurers have an impact on the measurements, this is not transferred to the experimental results. On the other hand, the results of our SMS confirm that TDD experiments do not usually report either the test suites, the test case generation method, or the details of how external quality was measured. A measurement protocol should be used to ensure that the measurements made by different measurers are similar. It is necessary to report the test cases, the experimental task and the intervention type in order to be able to reproduce the measurements and statistical analyses, as well as to replicate experiments and build dependable families of experiments. Fernando Uyaguari, Silvia Teresita Acuña, John W. Castro, Davide Fucci, Óscar Dieste Tubío, Sira Vegas |
ACM Trans. Softw. Eng. Methodol. | 5 |
| 2023 | Perceived usability of collaborative modeling tools
Ranci Ren, John W. Castro, Santiago R. Acuña, Óscar Dieste Tubío, Silvia Teresita Acuña |
J. Syst. Softw. | 4 |
| 2023 | Effect of Requirements Analyst Experience on Elicitation Effectiveness: A Family of Quasi-ExperimentsabstractContext.In software engineering there is a widespread assumption that experience improves requirements analyst effectiveness, although empirical studies demonstrate the opposite.Aim.Determine whether experience (interviews, eliciting, development, professional) influences requirements elicitation using interviews.Method.We ran 12 quasi-experiments recruiting 124 subjects in which we measured analyst effectiveness as the number of items (i.e., concepts, rules, processes) correctly elicited. The experimental task was to elicit requirements using the open interview technique followed by the consolidation of the elicited information in domains with which the analysts were and were not familiar.Results.In unfamiliar domains, interview experience, requirements experience, development experience, and professional experience does not have any relationship with analyst effectiveness. In familiar domains, effectiveness varies depending on the type of experience. Interview experience has a positive effect, whereas professional experience has a moderate negative effect. Requirements experience appears to have a moderately positive effect; however, the statistical power of the analysis is insufficient to be able to confirm this point. Development experience has no effect.Conclusion.Experience impacts analyst effectiveness differently depending on the problem domain type (familiar, unfamiliar). Generally, experience does not account for all the observed variability in effectiveness, so there are other influential factors. Alejandrina Aranda, Óscar Dieste Tubío, José Ignacio Panach, Natalia Juristo Juzgado |
IEEE Trans. Software Eng. | 2 |
| 2023 | Impact of Usability Mechanisms: A Family of Experiments on Efficiency, Effectiveness and User SatisfactionabstractContext: The usability software quality characteristic aims to improve system user performance. In a previous study, we found evidence of the impact of a set of usability features from the viewpoint of users in terms of efficiency, effectiveness and satisfaction. However, the impact level appears to depend on the usability feature and suggest priorities with respect to their implementation depending on how they promote user performance.Objectives: We use a family of three experiments to increase the precision and generalization of the results in the baseline experiment and provide findings regarding the impact on user performance of the Abort Operation, Progress Feedback and Preferences usability mechanisms.Method: We conduct two replications of the baseline experiment in academic settings. We analyse the data of 366 experimental subjects and apply aggregation (meta-analysis) procedures.Results: We find that the Abort Operation and Preferences usability mechanisms appear to improve system usability a great deal with respect to efficiency, effectiveness and user satisfaction.Conclusions: We find that the family of experiments further corroborates the results of the baseline experiment. Most of the results are statistically significant, and, because of the large number of experimental subjects, the evidence that we gathered in the replications is sufficient to outweigh other experiments. Juan M. Ferreira, Francy D. Rodríguez, Adrián Santos, Óscar Dieste Tubío, Silvia Teresita Acuña, Natalia Juristo Juzgado |
IEEE Trans. Software Eng. | 4 |
| 2023 | Using the SOCIO Chatbot for UML Modelling: A Family of ExperimentsabstractContext:Recent developments in natural language processing have facilitated the adoption of chatbots in typically collaborative software engineering tasks (such as diagram modelling). Families of experiments can assess the performance of tools and processes and, at the same time, alleviate some of the typical shortcomings of individual experiments (e.g., inaccurate and potentially biased results due to a small number of participants).Objective:Compare the usability of a chatbot for collaborative modelling (i.e., SOCIO) and an online web tool (i.e., Creately).Method:We conducted a family of three experiments to evaluate the usability of SOCIO against the Creately online collaborative tool in academic settings.Results:The student participants were faster at building class diagrams using the chatbot than with the online collaborative tool and more satisfied with SOCIO. Besides, the class diagrams built using the chatbot tended to be more concise —albeit slightly less complete.Conclusion:Chatbots appear to be helpful for building class diagrams. In fact, our study has helped us to shed light on the future direction for experimentation in this field and lays the groundwork for researching the applicability of chatbots in diagramming. Ranci Ren, John W. Castro, Adrián Santos, Óscar Dieste Tubío, Silvia Teresita Acuña |
IEEE Trans. Software Eng. | 4 |
| 2022 | Tutorial 3: Pitfalls in the Measurement Methods Applied in Experimental Software Engineering - Assessment and Suggestions for ImprovementabstractMeasurement is an essential issue in empirical software engineering. It is subject to different sources of error that must be kept as small as possible. Measuring instruments is one of these sources of error. In this tutorial, we provide awareness of potential pitfalls in the measurement methods—specifically measuring instruments—applied in empirical software engineering, describing statistical techniques that can be used for measures assessment, and making recommendations to improve the measurement practice in software engineering. Test suites are used as measuring instruments in many software engineering experiments. We will use the case of test suites when measuring the external quality of the code developed by participants of TDD-related experiments. Óscar Dieste Tubío, Sira Vegas |
EASE | 1 |
| 2021 | Towards a Methodology for Participant Selection in Software Engineering Experiments: A Vision of the FutureabstractBackground. Software Engineering (SE) researchers extensively perform experiments with human subjects. Well-defined samples are required to ensure external validity. Samples are selected purposely or by convenience, limiting the generalizability of results. Objective. We aim to depict the current status of participants selection in empirical SE, identifying the main threats and how they are mitigated. We draft a robust approach to participants' selection. Method. We reviewed existing participants' selection guidelines in SE, and performed a preliminary literature review to find out how participants' selection is conducted in SE in practice. Results. We outline a new selection methodology, by 1) defining the characteristics of the desired population, 2) locating possible sources of sampling available for researchers, and 3) identifying and reducing the "distance" between the selected sample and its corresponding population. Conclusion. We propose a roadmap to develop and empirically validate the selection methodology. Valentina Lenarduzzi, Óscar Dieste Tubío, Davide Fucci, Sira Vegas |
ESEM | 2 |
| 2021 | A family of experiments on test-driven development
Adrián Santos, Sira Vegas, Óscar Dieste Tubío, Fernando Uyaguari, Ayse Tosun Misirli, Davide Fucci, Burak Turhan, Giuseppe Scanniello, Simone Romano 0001, Itir Karac, Marco Kuhrmann, Vladimir Mandic, Robert Ramac, Dietmar Pfahl, Christian Engblom, Jarno Kyykka, Kerli Rungi, Carolina Palomeque, Jaroslav Spisak, Markku Oivo, Natalia Juristo Juzgado |
Empir. Softw. Eng. | 3 |
| 2021 | Evaluating Model-Driven Development Claims with Respect to Quality: A Family of ExperimentsabstractContext: There is a lack of empirical evidence on the differences between model-driven development (MDD), where code is automatically derived from conceptual models, and traditional software development method, where code is manually written. In our previous work, we compared both methods in a baseline experiment concluding that quality of the software developed following MDD was significantly better only for more complex problems (with more function points). Quality was measured through test cases run on a functional system. Objective: This paper reports six replications of the baseline to study the impact of problem complexity on software quality in the context of MDD. Method: We conducted replications of two types: strict replications and object replications. Strict replications were similar to the baseline, whereas we used more complex experimental objects (problems) in the object replications. Results: MDD yields better quality independently of problem complexity with a moderate effect size. This effect is bigger for problems that are more complex. Conclusions: Thanks to the bigger size of the sample after aggregating replications, we discovered an effect that the baseline had not revealed due to the small sample size. The baseline results hold, which suggests that MDD yields better quality for more complex problems. José Ignacio Panach, Óscar Dieste Tubío, Beatriz Marín, Sergio España 0001, Sira Vegas, Oscar Pastor 0001, Natalia Juristo Juzgado |
IEEE Trans. Software Eng. | 2 |
| 2021 | Investigating the Impact of Development Task on External Quality in Test-Driven Development: An Industry ExperimentabstractReviews on test-driven development (TDD) studies suggest that the conflicting results reported in the literature are due to unobserved factors, such as the tasks used in the experiments, and highlight that there are very few industry experiments conducted with professionals. The goal of this study is to investigate the impact of a new factor, the chosentask, and thedevelopment approachon external quality in an industrial experimental setting with 17 professionals. The participants are junior to senior developers in programming with Java, beginner to novice in unit testing, JUnit, and they have no prior experience in TDD. The experimental design is a$2\times 2$cross-over, i.e., we use two tasks for each of the two approaches, namely TDD and incremental test-last development (ITLD). Our results reveal that bothdevelopment approachandtaskare significant factors with regards to the external quality achieved by the participants. More specifically, the participants produce higher quality code during ITLD in which splitting user stories into subtasks, coding, and testing activities are followed, compared to TDD. The results also indicate that the participants produce higher quality code during the implementation of Bowling Score Keeper, compared to that of Mars Rover API, although they perceived both tasks as of similar complexity. An interaction between thedevelopment approachandtaskcould not be observed in this experiment. We conclude that variables that have not been explored so often, such as the extent to which the task is specified in terms of smaller subtasks, and developers’ unit testing experience might be critical factors in TDD experiments. The real-world appliance of TDD and its implications on external quality still remain to be challenging unless these uncontrolled and unconsidered factors are further investigated by researchers in both academic and industrial settings. Ayse Tosun Misirli, Óscar Dieste Tubío, Sira Vegas, Dietmar Pfahl, Kerli Rungi, Natalia Juristo Juzgado |
IEEE Trans. Software Eng. | 2 |
| 2020 | Publication Bias: A Detailed Analysis of Experiments Published in ESEMabstractBackground: Publication bias is the failure to publish the results of a study based on the direction or strength of the study findings. The existence of publication bias is firmly established in areas like medical research. Recent research suggests the existence of publication bias in Software Engineering. Aims: Finding out whether experiments published in the International Workshop on Empirical Software Engineering and Measurement (ESEM) are affected by publication bias. Method: We review experiments published in ESEM. We also survey with experimental researchers to triangulate our findings. Results: ESEM experiments do not define hypotheses and frequently perform multiple testing. One-tailed tests have a slightly higher rate of achieving statistically significant results. We could not find other practices associated with publication bias. Conclusions: Our results provide a more encouraging perspective of SE research than previous research: (1) ESEM publications do not seem to be strongly affected by biases and (2) we identify some practices that could be associated with p-hacking, but it is more likely that they are related to the conduction of exploratory research. Rolando P. Reyes Ch., Óscar Dieste Tubío, Efraín R. Fonseca C., Natalia Juristo Juzgado |
EASE | 2 |
| 2020 | Impact of usability mechanisms: An experiment on efficiency, effectiveness and user satisfaction
Juan M. Ferreira, Silvia Teresita Acuña, Óscar Dieste Tubío, Sira Vegas, Adrián Santos, Francy D. Rodríguez, Natalia Juristo Juzgado |
Inf. Softw. Technol. | 3 |
| 2020 | Increasing validity through replication: an illustrative TDD caseabstractAbstract Software engineering (SE) experiments suffer from threats to validity that may impact their results. Replication allows researchers building on top of previous experiments’ weaknesses and increasing the reliability of the findings. Illustrating the benefits of replication to increase the reliability of the findings and uncover moderator variables. We replicate an experiment on test-driven development (TDD) and address some of its threats to validity and those of a previous replication. We compare the replications’ results and hypothesize on plausible moderators impacting results. Differences across TDD replications’ results might be due to the operationalization of the response variables, the allocation of subjects to treatments, the allowance to work outside the laboratory, the provision of stubs, or the task. Replications allow examining the robustness of the findings, hypothesizing on plausible moderators influencing results, and strengthening the evidence obtained. Adrián Santos, Sira Vegas, Fernando Uyaguari, Óscar Dieste Tubío, Burak Turhan, Natalia Juristo Juzgado |
Softw. Qual. J. | 4 |
| 2018 | Statistical errors in software engineering experiments: a preliminary literature reviewabstractBackground: Statistical concepts and techniques are often applied incorrectly, even in mature disciplines such as medicine or psychology. Surprisingly, there are very few works that study statistical problems in software engineering (SE). Aim: Assess the existence of statistical errors in SE experiments. Method: Compile the most common statistical errors in experimental disciplines. Survey experiments published in ICSE to assess whether errors occur in high quality SE publications. Results: The same errors as identified in others disciplines were found in ICSE experiments, where 30% of the reviewed papers included several error types such as: a) missing statistical hypotheses, b) missing sample size calculation, c) failure to assess statistical test assumptions, and d) uncorrected multiple testing. This rather large error rate is greater for research papers where experiments are confined to the validation section. The origin of the errors can be traced back to: a) researchers not having sufficient statistical training, and, b) a profusion of exploratory research. Conclusions: This paper provides preliminary evidence that SE research suffers from the same statistical problems as other experimental disciplines. However, the SE community appears to be unaware of any shortcomings in its experiments, whereas other disciplines work hard to avoid these threats. Further research is necessary to find the underlying causes and set up corrective measures, but there are some potentially effective actions and are a priori easy to implement: a) improve the statistical training of SE researchers, and b) enforce quality assessment and reporting guidelines in SE publications. Rolando P. Reyes Ch., Óscar Dieste Tubío, Efraín R. Fonseca C., Natalia Juristo Juzgado |
ICSE | 2 |
| 2018 | Empirical evaluation of the effects of experience on code quality and programmer productivity: an exploratory studyabstractThis extended abstract summarizes an article, which has been published in the Empirical Software Engineering Journal and was selected for the Journal-First presentations at the International Conference on Software and System Process (ICSSP 2018). Óscar Dieste Tubío, Alejandrina Aranda, Fernando Uyaguari, Burak Turhan, Ayse Tosun Misirli, Davide Fucci, Markku Oivo, Natalia Juristo Juzgado |
ICSSP | 1 |
| 2017 | How do Practitioners Perceive the Relevance of Requirements Engineering Research? An Ongoing StudyabstractThe relevance of Requirements Engineering (RE) research to practitioners is a prerequisite for problem-driven research in the area and key for a long-term dissemination of research results to everyday practice. To understand better how industry practitioners perceive the practical relevance of RE research, we have initiated the RE-Pract project, an international collaboration conducting an empirical study. This project opts for a replication of previous work done in two different domains and relies on survey research. To this end, we have designed a survey to be sent to several hundred industry practitioners at various companies around the world and ask them to rate their perceived practical relevance of the research described in a sample of 418 RE papers published between 2010 and 2015 at the RE, ICSE, FSE, ESEC/FSE, ESEM and REFSQ conferences. In this paper, we summarize our research protocol and present the current status of our study and the planned future steps. Xavier Franch, Daniel Méndez 0001, Marc Oriol, Andreas Vogelsang, Rogardt Heldal, Eric Knauss, Guilherme Horta Travassos, Jeffrey C. Carver, Óscar Dieste Tubío, Thomas Zimmermann 0001 |
RE | 9 |
| 2017 | Empirical evaluation of the effects of experience on code quality and programmer productivity: an exploratory study
Óscar Dieste Tubío, Alejandrina Aranda, Fernando Uyaguari, Burak Turhan, Ayse Tosun Misirli, Davide Fucci, Markku Oivo, Natalia Juristo Juzgado |
Empir. Softw. Eng. | 1 |
| 2017 | An industry experiment on the effects of test-driven development on external quality and productivity
Ayse Tosun Misirli, Óscar Dieste Tubío, Davide Fucci, Sira Vegas, Burak Turhan, Hakan Erdogmus, Adrián Santos, Markku Oivo, Kimmo Toro, Janne Järvinen, Natalia Juristo Juzgado |
Empir. Softw. Eng. | 2 |
| 2017 | Contextual attributes impacting the effectiveness of requirements elicitation Techniques: Mapping theoretical and empirical research
Dante Carrizo Moreno, Óscar Dieste Tubío, Natalia Juristo Juzgado |
Inf. Softw. Technol. | 2 |
| 2016 | How Practitioners Perceive the Relevance of ESEM ResearchabstractBackground: The relevance of ESEM research to industry practitioners is key to the long-term health of the conference. Aims: The goal of this work is to understand how ESEM research is perceived within the practitioner community and provide feedback to the ESEM community ensure our research remains relevant. Method: To understand how practitioners perceive ESEM research, we replicated previous work by sending a survey to several hundred industry practitioners at a number of companies around the world. We asked the survey participants to rate the relevance of the research described in 156 ESEM papers published between 2011 and 2015. Results: We received 9,941 ratings by 437 practitioners who labeled ideas as Essential, Worth-while, Unimportant, or Unwise. The results showed that overall, industrial practitioners find the work published in ESEM to be valuable: 67% of all ratings were essential or worthwhile. We found no correlation between citation count and perceived relevance of the papers. Through a qualitative analysis, we also identified a number of research themes on which practitioners would like to see an increased research focus. Conclusions: The work published in ESEM is generally relevant to industrial practitioners. There are a number of topics for which those practitioners would like to see additional research undertaken. Jeffrey C. Carver, Óscar Dieste Tubío, Nicholas A. Kraft, David Lo 0001, Thomas Zimmermann 0001 |
ESEM | 2 |
| 2016 | Effect of Domain Knowledge on Elicitation Effectiveness: An Internally Replicated Controlled ExperimentabstractContext. Requirements elicitation is a highly communicative activity in which human interactions play a critical role. A number of analyst characteristics or skills may influence elicitation process effectiveness. Aim. Study the influence of analyst problem domain knowledge on elicitation effectiveness. Method. We executed a controlled experiment with post-graduate students. The experimental task was to elicit requirements using open interview and consolidate the elicited information immediately afterwards. We used four different problem domains about which students had different levels of knowledge. Two tasks were used in the experiment, whereas the other two were used in an internal replication of the experiment; that is, we repeated the experiment with the same subjects but with different domains. Results. Analyst problem domain knowledge has a small but statistically significant effect on the effectiveness of the requirements elicitation activity. The interviewee has a big positive and significant influence, as does general training in requirements activities and interview experience. Conclusion. During early contacts with the customer, a key factor is the interviewee; however, training in tasks related to requirements elicitation and knowledge of the problem domain helps requirements analysts to be more effective. Alejandrina Aranda, Óscar Dieste Tubío, Natalia Juristo Juzgado |
IEEE Trans. Software Eng. | 2 |
| 2015 | Towards an operationalization of test-driven development skills: An industrial empirical study
Davide Fucci, Burak Turhan, Natalia Juristo Juzgado, Óscar Dieste Tubío, Ayse Tosun Misirli, Markku Oivo |
Inf. Softw. Technol. | 4 |
| 2015 | In search of evidence for model-driven development claims: An experiment on quality, effort, productivity and satisfaction
José Ignacio Panach, Sergio España 0001, Óscar Dieste Tubío, Oscar Pastor 0001, Natalia Juristo Juzgado |
Inf. Softw. Technol. | 3 |
| 2014 | Evidence of the presence of bias in subjective metrics: analysis within a family of experimentsabstractContext: Measurement is crucial and important to empirical software engineering. Although reliability and validity are two important properties warranting consideration in measurement processes, they may be influenced by random or systematic error (bias) depending on which metric is used. Aim: Check whether, the simple subjective metrics used in empirical software engineering studies are prone to bias. Method: Comparison of the reliability of a family of empirical studies on requirements elicitation that explore the same phenomenon using different design types and objective and subjective metrics. Results: The objectively measured variables (experience and knowledge) tend to achieve more reliable results, whereas subjective metrics using Likert scales (expertise and familiarity) tend to be influenced by systematic error or bias. Conclusions: Studies that predominantly use variables measured subjectively, like opinion polls or expert opinion acquisition, must take every care to prevent bias that can result in incorrect results. Alejandrina Aranda, Óscar Dieste Tubío, Natalia Juristo Juzgado |
EASE | 2 |
| 2014 | Replication types: towards a shared taxonomyabstractContext: The software engineering community is becoming more aware of the need for experimental replications. In spite of the importance of this topic, there is still much inconsistency in the terminology used to describe replications. Maria Teresa Baldassarre, Jeffrey C. Carver, Óscar Dieste Tubío, Natalia Juristo Juzgado |
EASE | 3 |
| 2014 | Reviewing technical approaches for sharing and preservation of experimental dataabstractContext: Empirical Software Engineering (ESE) replication researchers need to store and manipulate experimental data for several purposes, in particular analysis and reporting. Current research needs call for sharing and preservation of experimental data as well. In a previous work, we analyzed Replication Data Management (RDM) needs. A novel concept, called Experimental Ecosystem, was proposed to solve current deficiencies in RDM approaches. The empirical ecosystem provides replication researchers with a common framework that integrates transparently local heterogeneous data sources. A typical situation where the Empirical Ecosystem is applicable, is when several members of a research group, or several research groups collaborating together, need to share and access each other experimental results. However, to be able to apply the Empirical Ecosystem concept and deliver all promised benefits, it is necessary to analyze the software architectures and tools that can properly support it. Efraín R. Fonseca C., Óscar Dieste Tubío, Natalia Juristo Juzgado, Estefanía Serral, Stefan Biffl |
ESEM | 2 |
| 2014 | Effectiveness for detecting faults within and outside the scope of testing techniques: an independent replication
Cecilia Apa, Óscar Dieste Tubío, Edison G. Espinosa, Efraín R. Fonseca C. |
Empir. Softw. Eng. | 2 |
| 2014 | Systematizing requirements elicitation technique selection
Dante Carrizo Moreno, Óscar Dieste Tubío, Natalia Juristo Juzgado |
Inf. Softw. Technol. | 2 |
| 2013 | Replication Data Management: Needs and Solutions - An Initial Evaluation of Conceptual Approaches for Integrating Heterogeneous Replication Study Dataabstract[Context] Replication Data Management (RDM) aims at enabling the use of data collections from several itera-tions of an experiment. However, there are several major chal-lenges to RDM from integrating data models and data from em-pirical study infrastructures that were not designed to cooperate, e.g., data model variation of local data sources. [Objective] In this paper we analyze RDM needs and evaluate conceptual RDM approaches to support replication researchers. [Method] We adapted the ATAM evaluation process to (a) analyze RDM use cases and needs of empirical replication study research groups and (b) compare three conceptual approaches to address these RDM needs: central data repositories with a fixed data model, heterogeneous local repositories, and an empirical ecosystem. [Results] While the central and local approaches have major issues that are hard to resolve in practice, the empirical ecosys-tem allows bridging current gaps in RDM from heterogeneous data sources. [Conclusions] The empirical ecosystem approach should be explored in diverse empirical environments. Stefan Biffl, Estefanía Serral, Dietmar Winkler 0001, Nelly Condori-Fernández, Óscar Dieste Tubío, Natalia Juristo Juzgado |
ESEM | 5 |
| 2012 | A systematic mapping study on the open source software development processabstractAbstract?Background: There is no globally accepted open source software development process to define how open source software is developed in practice. A process description is important for coordinating all the software development activities involving both people and technology. Aim: The research question that this study sets out to answer is: What activities do open source software process models contain? The activity groups on which it focuses are Concept Exploration, Software Requirements, Design, Maintenance and Evaluation. Method: We conduct a systematic mapping study (SMS). A SMS is a form of systematic literature review that aims to identify and classify available research papers concerning a particular issue. Results: We located a total of 29 primary studies, which we categorized by the open source software project that they examine and by activity types (Concept Exploration, Software Requirements, Design, Maintenance and Evaluation). The activities present in most of the open source software development processes were Execute Tests and Conduct Reviews, which belong to the Evaluation activities group. Maintenance is the only group that has primary studies addressing all the activities that it contains. Conclusions: The primary studies located by the SMS are the starting point for analyzing the open source software development process and proposing a process model for this community. The papers in our paper pool that describe a specific open source software project provide more regarding our research question than the papers that talk about open source software development without referring to a specific open source software project. Silvia Teresita Acuña, John W. Castro, Óscar Dieste Tubío, Natalia Juristo Juzgado |
EASE | 3 |
| 2012 | Software as a Service: Undo
Hernán Merlino, Óscar Dieste Tubío, Patricia Pesado, Ramón García-Martínez |
SEKE | 2 |
| 2011 | Comparative analysis of meta-analysis methods: When to use which?abstractBackground: Several meta-analysis methods can be used to quantitatively combine the results of a group of experiments, including the weighted mean difference, statistical vote counting, the parametric response ratio and the non-parametric response ratio. The software engineering community has focused on the weighted mean difference method. However, other meta-analysis methods have distinct strengths, such as being able to be used when variances are not reported. There are as yet no guidelines to indicate which method is best for use in each case Aim: Compile a set of rules that SE researchers can use to ascertain which aggregation method is best for use in the synthesis phase of a systematic review. Method: Monte Carlo simulation varying the number of experiments in the meta analyses, the number of subjects that they include, their variance and effect size. We empirically calculated the reliability and statistical power in each case Results: WMD is generally reliable if the variance is low, whereas its power depends on the effect size and number of subjects per meta-analysis; the reliability of RR is generally unaffected by changes in variance, but it does require more subjects than WMD to be powerful; NPRR is the most reliable method, but it is not very powerful; SVC behaves well when the effect size is moderate, but is less reliable with other effect sizes. Detailed tables of results are annexed. Conclusions: Before undertaking statistical aggregation in software engineering, it is worthwhile checking whether there is any appreciable difference in the reliability and power of the methods. If there is, software engineers should select the method that optimizes both parameters. Óscar Dieste Tubío, Enrique Fernández, Ramón García-Martínez, Natalia Juristo Juzgado |
EASE | 1 |
| 2011 | The Risk of Using the Q Heterogeneity Estimator for Software Engineering ExperimentsabstractAll meta-analyses should include a heterogeneity analysis. Even so, it is not easy to decide whether a set of studies are homogeneous or heterogeneous because of the low statistical power of the statistics used (usually the Q test). Objective: Determine a set of rules enabling SE researchers to find out, based on the characteristics of the experiments to be aggregated, whether or not it is feasible to accurately detect heterogeneity. Method: Evaluate the statistical power of heterogeneity detection methods using a Monte Carlo simulation process. Results: The Q test is not powerful when the meta-analysis contains up to a total of about 200 experimental subjects and the effect size difference is less than 1. Conclusions: The Q test cannot be used as a decision-making criterion for meta-analysis in small sample settings like SE. Random effects models should be used instead of fixed effects models. Caution should be exercised when applying Q test-mediated decomposition into subgroups. Óscar Dieste Tubío, Enrique Fernández, Ramón García-Martínez, Natalia Juristo Juzgado |
ESEM | 1 |
| 2011 | Quantitative Determination of the Relationship between Internal Validity and Bias in Software Engineering Experiments: Consequences for Systematic Literature ReviewsabstractQuality assessment is one of the activities performed as part of systematic literature reviews. It is commonly accepted that a good quality experiment is bias free. Bias is considered to be related to internal validity (e.g., how adequately the experiment is planned, executed and analysed). Quality assessment is usually conducted using checklists and quality scales. It has not yet been proven, however, that quality is related to experimental bias. Aim: Identify whether there is a relationship between internal validity and bias in software engineering experiments. Method: We built a quality scale to determine the quality of the studies, which we applied to 28 experiments included in two systematic literature reviews. We proposed an objective indicator of experimental bias, which we applied to the same 28 experiments. Finally, we analysed the correlations between the quality scores and the proposed measure of bias. Results: We failed to find a relationship between the global quality score (resulting from the quality scale) and bias, however, we did identify interesting correlations between bias and some particular aspects of internal validity measured by the instrument. Conclusions: There is an empirically provable relationship between internal validity and bias. It is feasible to apply quality assessment in systematic literature reviews, subject to limits on the internal validity aspects for consideration. Óscar Dieste Tubío, Anna Grimán, Natalia Juristo Juzgado, Himanshu Saxena |
ESEM | 1 |
| 2011 | Systematic Review and Aggregation of Empirical Studies on Elicitation TechniquesabstractWe have located the results of empirical studies on elicitation techniques and aggregated these results to gather empirically grounded evidence. Our chosen surveying methodology was systematic review, whereas we used an adaptation of comparative analysis for aggregation because meta-analysis techniques could not be applied. The review identified 564 publications from the SCOPUS, IEEEXPLORE, and ACM DL databases, as well as Google. We selected and extracted data from 26 of those publications. The selected publications contain 30 empirical studies. These studies were designed to test 43 elicitation techniques and 50 different response variables. We got 100 separate results from the experiments. The aggregation generated 17 pieces of knowledge about the interviewing, laddering, sorting, and protocol analysis elicitation techniques. We provide a set of guidelines based on the gathered pieces of knowledge. Óscar Dieste Tubío, Natalia Juristo Juzgado |
IEEE Trans. Software Eng. | 1 |
| 2010 | Usability evaluation of multi-device/platform user interfaces generated by model-driven engineeringabstractNowadays several Computer-Aided Software Engineering environments exploit Model-Driven Engineering (MDE) techniques in order to generate a single user interface for a given computing platform or multi-platform user interfaces for several computing platforms simultaneously. Therefore, there is a need to assess the usability of those generated user interfaces, either taken in isolation or compared to each other. This paper describes an MDE approach that generates multi-platform graphical user interfaces (e.g., desktop, web) that will be subject to an exploratory controlled experiment. The usability of user interfaces generated for the two mentioned platforms and used on multiple display devices (i.e., standard size, large, and small screens) has been examined in terms of satisfaction, effectiveness and efficiency. An experiment with a factorial design for repeated measures was conducted for 31 participants, i.e., postgraduate students and professors selected by convenience sampling. The data were collected with the help of questionnaires and forms and were analyzed using parametric and non-parametric tests such as ANOVA with repeated measures and Friedman's test, respectively. Efficiency was significantly better in large screens than in small ones as well as in the desktop platform rather than in the web platform, with a confidence level of 95%. The experiment also suggests that satisfaction tends to be better in standard size screens than in small ones. The results suggest that the tested MDE approach should incorporate enhancements in its multi-device/platform user interface generation process in order to improve its generated usability. Nathalie Aquino, Jean Vanderdonckt, Nelly Condori-Fernández, Óscar Dieste Tubío, Oscar Pastor 0001 |
ESEM | 4 |
| 2009 | A systematic mapping study on empirical evaluation of software requirements specifications techniquesabstractThis paper describes an empirical mapping study, which was designed to identify what aspects of software requirement specifications (SRS) are empirically evaluated, in which context, and by using which research method. On the basis of 46 identified and categorized primary studies, we found that understandability is the most commonly evaluated aspect of SRS, experiments are the most commonly used research method, and the academic environment is where most empirical evaluation takes place. Nelly Condori-Fernández, Maya Daneva, Klaas Sikkel, Roel J. Wieringa, Óscar Dieste Tubío, Oscar Pastor 0001 |
ESEM | 5 |
| 2009 | Developing search strategies for detecting relevant experiments
Óscar Dieste Tubío, Anna Grimán, Natalia Juristo Juzgado |
Empir. Softw. Eng. | 1 |
| 2008 | Obtaining Well-Founded Practices about Elicitation Techniques by Means of an Update of a Previous Systematic Review
Óscar Dieste Tubío, Marta López Fernández, Felicidad Ramos |
SEKE | 1 |
| 2008 | Formalizing a Systematic Review Updating ProcessabstractThe objective of a systematic review is to obtain empirical evidence about the topic under review and to allow moving forward the body of knowledge of a discipline. Therefore, systematic reviewing is a tool we can apply in Software Engineering to develop well founded guidelines with the final goal of improving the quality of the software systems. However, we still do not have as much experience in performing systematic reviews as in other disciplines like medicine, and therefore we need detailed guidance. This paper presents a proposal of a improved process to perform systematic reviews in software engineering. This process is the result of the tasks carried out in a first review and a subsequent update concerning the effectiveness of elicitation techniques. Óscar Dieste Tubío, Marta López Fernández, Felicidad Ramos |
SERA | 1 |
| 2007 | Developing Search Strategies for Detecting Relevant Experiments for Systematic ReviewsabstractInformation retrieval is an important problem in any evidence-based discipline. Although Evidence-based Software Engineering (EBSE) is not immune to this fact, this question has not been examined at length. The goal of this paper is to analyse the optimality of search strategies for use in systematic reviews. We tried out 29 search strategies using different terms and combinations of terms. We evaluated their sensitivity and precision with a view to finding an optimum strategy. From this study of search strategies we were able to analyse trends and weaknesses in terminology use in articles reporting experiments. Óscar Dieste Tubío, Anna Grimán |
ESEM | 1 |
| 2007 | A Quantitative Assessment of Requirements Engineering Publications - 1963-2006
Alan M. Davis, Ann M. Hickey, Óscar Dieste Tubío, Natalia Juristo Juzgado, Ana María Moreno 0001 |
REFSQ | 3 |
| 2007 | It-Outsourcing and IT-Offshoring: Trends and Impacts on SE/KE CurriculaabstractAs a result of IT outsourcing and offshoring, IT professionals and educators are faced with the following question: What SE & KE skill sets will make a software engineer or a knowledge engineer immune to the impact of outsourcing and offshoring? This article summarizes the position papers from a panel held during the 2006 International Conference on Software Engineering and Knowledge Engineering from July 5 to 7 at the Hotel Sofitel, Redwood City in California, USA. Bringing software and knowledge engineers closer to the needs of their prospective customers and providing more value than simply pure software development and maintenance, is an open challenge at least for traditional computer science and software engineering curricula. Ron Hira, Óscar Dieste Tubío, George Spanoudakis, Giuseppe Visaggio, Guido Wirtz |
Int. J. Softw. Eng. Knowl. Eng. | 2 |
| 2006 | Effectiveness of Requirements Elicitation Techniques: Empirical Results Derived from a Systematic ReviewabstractThis paper reports a systematic review of empirical studies concerning the effectiveness of elicitation techniques, and the subsequent aggregation of empirical evidence gathered from those studies. The most significant results of the aggregation process are as follows: (I) interviews, preferentially structured, appear to be one of the most effective elicitation techniques; (2) many techniques often cited in the literature, like card sorting, ranking or thinking aloud, tend to be less effective than interviews; (3) analyst experience does not appear to be a relevant factor; and (4) the studies conducted have not found the use of intermediate representations during elicitation to have significant positive effects. It should be noted that, as a general rule, the studies from which these results were aggregated have not been replicated, and therefore the above claims cannot be said to be absolutely certain. However, they can be used by researchers as pieces of knowledge to be further investigated and by practitioners in development projects, always taking into account that they are preliminary findings Alan M. Davis, Óscar Dieste Tubío, Ann M. Hickey, Natalia Juristo Juzgado, Ana María Moreno 0001 |
RE | 2 |
| 2003 | A conceptual model completely independent of the implementation paradigm
Óscar Dieste Tubío, Marcela Genero, Natalia Juristo Juzgado, José Luis Maté, Ana María Moreno 0001 |
J. Syst. Softw. | 1 |
| 2001 | Development-Paradigm Independent Conceptual Models
Óscar Dieste Tubío |
SEKE | 1 |
| 2000 | Integrated Software Engineering and Knowledge Engineering Teaching ExperiencesabstractThis paper presents the motivations, experiences and results of teaching integrated Software Engineering (SE) and Knowledge Engineering (KE), specifically as part of the master course organized by the Polytechnic University of Madrid (School of Computer Science). The paper outlines a possible approach to this instruction, whose aim is for software practitioners thus educated to have a flexible and moldable view of the software systems development process. This broad and malleable approach allows future practitioners to better address the increasingly more complex, divergent and innovative problems and needs raised by users. This approach is the result of a gradual and continuous process. This paper discusses the current stage of integration, giving a detailed description and justification of the scope of the integrated instruction. For the purpose of quantitatively analyzing this experience, the paper also shows the results of the evaluation conducted throughout this process at three levels (industry, students and projects). Óscar Dieste Tubío, Natalia Juristo Juzgado, Ana María Moreno 0001, Marta López Fernández |
Int. J. Softw. Eng. Knowl. Eng. | 1 |