VLDB 2026 Research / reviewers in the wild / expert
Daniel Rodríguez-García
dblp:00/7633 · also Daniel Rodríguez 0001
· DBLP profile ↗
37ranked-venue papers
9as first author
5since 2021 · last 2026
0000-0002-2887-0185ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 19 · 7 first-author · 2 since 2021Artificial intelligence and machine learning · 12 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Classifying illicit dark web content through zero-shot prompting: An empirical study with GPT modelsabstractThis study evaluates the classification performance of four GPT-based models (GPT-4.1, GPT-4.1-mini, GPT-4.1-nano, and o4-mini) under zero-shot prompting conditions on the complete, multilingual CoDA dataset of Dark Web content, comprising 10 illicit activity categories. The models GPT-4.1, GPT-4.1-mini, and o4-mini achieve a weighted F1 score of 0.885, surpassing prior zero-shot baselines on this dataset. Stability analysis using TARa@10 demonstrates high output consistency for GPT-4.1 (0.964) and GPT-4.1-mini (0.970), indicating their reliability for operational use. Multilingual evaluation reveals only a modest English vs. non-English performance gap for GPT-4.1 (0.031), while other models perform comparably across languages. The strongest results appear in Drugs , Gambling , and Porn (F1 0.9), whereas lower scores are observed in ambiguous or overlapping categories like Violence (F1 0.76) or Crypto (F1 0.84). A qualitative review of misclassifications suggests that some model predictions align with reasonable semantic interpretations, potentially highlighting annotation inconsistencies. This work establishes a performance baseline for GPT-based models in zero-shot classification of multilingual Dark Web content and underscores the importance of clear category definitions for effective deployment. Adrián Domínguez-Díaz, Luis de-Marcos, Víctor Pablo Prado-Sánchez, Daniel Rodríguez-García, José-Javier Martínez 0001 |
Inf. Process. Manag. | 4 |
| 2025 | The Challenge of Generating and Evolving Real-Life Like Synthetic Test Data Without Accessing Real-World Raw Data - A Systematic ReviewabstractABSTRACT Background High‐level system testing of applications that use data from e‐Government services as input requires test data that is real‐life‐like but where the privacy of personal information is guaranteed. Applications with such strong requirement include information exchange between countries, medicine, banking, and so on. This review aims to synthesise the current state‐of‐the‐practice in this domain. Objectives The objective of this Systematic Review is to identify existing approaches for creating and evolving synthetic test data without using real‐life raw data. Methods We followed well‐known methodologies for conducting systematic literature reviews, including the ones from Kitchenham and PRISMA as well as guidelines for analysing the limitations of our review and its threats to validity. Results A variety of methods and tools exist for creating privacy‐preserving test data. Our search found 1013 publications in IEEE Xplore, ACM Digital Library, and SCOPUS. We extracted data from 75 of those publications and identified 37 approaches that answer our research question partly. A common prerequisite for using these methods and tools is direct access to real‐life data for data anonymization or synthetic test data generation. Nine existing synthetic test data generation approaches were identified that were closest to answering our research question. Nevertheless, further work would be needed to add the ability to evolve synthetic test data to the existing approaches. Conclusions None of the publications covered our requirements completely, only partially. Synthetic test data evolution is a field that has not received much attention from researchers but needs to be explored in Digital Government Solutions, especially since new legal regulations are being put in force in many countries. Maj-Annika Tammisto, Faiz Ali Shah, Daniel Rodríguez-García, Dietmar Pfahl |
Expert Syst. J. Knowl. Eng. | 3 |
| 2025 | The effect of data complexity on classifier performanceabstractThe research area of Software Defect Prediction (SDP) is both extensive and popular, and is often treated as a classification problem. Improvements in classification, pre-processing and tuning techniques, (together with many factors which can influence model performance) have encouraged this trend. However, no matter the effort in these areas, it seems that there is a ceiling in the performance of the classification models used in SDP. In this paper, the issue of classifier performance is analysed from the perspective of data complexity. Specifically, data complexity metrics are calculated using the Unified Bug Dataset, a collection of well-known SDP datasets, and then checked for correlation with the defect prediction performance of machine learning classifiers (in particular, the classifiers C5.0, Naive Bayes, Artificial Neural Networks, Random Forests, and Support Vector Machines). In this work, different domains of competence and incompetence are identified for the classifiers. Similarities and differences between the classifiers and the performance metrics are found and the Unified Bug Dataset is analysed from the perspective of data complexity. We found that certain classifiers work best in certain situations and that all data complexity metrics can be problematic, although certain classifiers did excel in some situations. Jonas Eberlein, Daniel Rodríguez-García, Rachel Harrison |
Empir. Softw. Eng. | 2 |
| 2024 | Monitoring tools for DevOps and microservices: A systematic grey literature reviewabstractMicroservice-based systems are usually developed according to agile practices like DevOps, which enables rapid and frequent releases to promptly react and adapt to changes. Monitoring is a key enabler for these systems, as they allow to continuously get feedback from the field and support timely and tailored decisions for a quality-driven evolution. In the realm of monitoring tools available for microservices in the DevOps-driven development practice, each with different features, assumptions, and performance, selecting a suitable tool is an as much difficult as impactful task. This article presents the results of a systematic study of the grey literature we performed to identify, classify and analyze the available monitoring tools for DevOps and microservices. We selected and examined a list of 71 monitoring tools, drawing a map of their characteristics, limitations, assumptions, and open challenges, meant to be useful to both researchers and practitioners working in this area. Results are publicly available and replicable. Editor's note: Open Science material was validated by the Journal of Systems and Software Open Science Board. Luca Giamattei, Antonio Guerriero, Roberto Pietrantuono, Stefano Russo 0001, Ivano Malavolta, Tanjina Islam, Madalina Dinga, Anne Koziolek, Snigdha Singh, Martin Armbruster, Jose-Maria Gutierrez-Martinez, Sergio Caro-Álvaro, Daniel Rodríguez-García, Sebastian Weber 0001, Jörg Henß, Estrella Fernández Vogelin, Fernando Simön Panojo |
J. Syst. Softw. | 13 |
| 2021 | Merge Nondominated Sorting Algorithm for Many-Objective OptimizationabstractMany Pareto-based multiobjective evolutionary algorithms require ranking the solutions of the population in each iteration according to the dominance principle, which can become a costly operation particularly in the case of dealing with many-objective optimization problems. In this article, we present a new efficient algorithm for computing the nondominated sorting procedure, called merge nondominated sorting (MNDS), which has a best computational complexity of$O(N\log N)$and a worst computational complexity of$O(MN^{2})$, with$N$being the population size and$M$being the number of objectives. Our approach is based on the computation of thedominance set, that is, for each solution, the set of solutions that dominate it, by taking advantage of the characteristics of the merge sort algorithm. We compare MNDS against six well-known techniques that can be considered as the state-of-the-art. The results indicate that the MNDS algorithm outperforms the other techniques in terms of the number of comparisons as well as the total running time. Javier Moreno 0004, Daniel Rodríguez-García, Antonio J. Nebro, José Antonio Lozano 0001 |
IEEE Trans. Cybern. | 2 |
| 2019 | Distributed correlation-based feature selection in spark
Raul-Jose Palma-Mendoza, Luis de-Marcos, Daniel Rodríguez-García, Amparo Alonso-Betanzos |
Inf. Sci. | 3 |
| 2018 | Competence-based recommender systems: a systematic literature reviewabstractCompetence-based learning is increasingly widespread in many institutions since it provides flexibility, facilitates the self-learning and brings the academic and professional worlds closer together. Thus, the competence-based recommender systems emerged taking the advantages of competences to offer suggestions (performance of a learning experience, assistance of an expert or recommendation of a learning resource) to the user (learner or instructor). The objective of this work is to conduct a new Systematic Literature Review (SLR) concerning competence-based recommender systems to analyse in relation to their nature and assessment of competences an others key factors that provide more flexible and exhaustive recommendations. To do so, a SLR research methodology was followed in which 25 competence-based recommender systems related to learning or instruction environments were classified according to multiple criteria. We evaluate the role of competences in these proposals and enumerate the emerging challenges. Also a critical analysis of current proposals is carried out to determine their strengths and weakness. Finally, future research paths to be explored are grouped around two main axes closely interlinked; first about the typical challenges related to recommender systems and second, concerning ambitious emerging challenges. Héctor Yago Corral, Julia Clemente Párraga, Daniel Rodríguez-García |
Behav. Inf. Technol. | 3 |
| 2018 | ON-SMMILE: Ontology Network-based Student Model for MultIple Learning Environments
Héctor Yago Corral, Julia Clemente Párraga, Daniel Rodríguez-García, Pedro Fernández de Córdoba |
Data Knowl. Eng. | 3 |
| 2018 | Using simulation-based optimization in the context of IT service management change process
Mercedes Ruiz 0001, Javier Moreno 0004, Bernabé Dorronsoro, Daniel Rodríguez-García |
Decis. Support Syst. | 4 |
| 2018 | Distributed ReliefF-based feature selection in Spark
Raul-Jose Palma-Mendoza, Daniel Rodríguez-García, Luis de-Marcos |
Knowl. Inf. Syst. | 2 |
| 2018 | Multiobjective Testing Resource Allocation Under UncertaintyabstractTesting resource allocation is the problem of planning the assignment of resources to testing activities of software components so as to achieve a target goal under given constraints. Existing methods build on software reliability growth models (SRGMs), aiming at maximizing reliability given time/cost constraints, or at minimizing cost given quality/time constraints. We formulate it as a multiobjective debug-aware and robust optimization problem under uncertainty of data, advancing the state-of-the-art in the following ways. Multiobjective optimization produces a set of solutions, allowing to evaluate alternative tradeoffs among reliability, cost, and release time. Debug awareness relaxes the traditional assumptions of SRGMs-in particular the very unrealistic immediate repair of detected faults-and incorporates the bug assignment activity. Robustness provides solutions valid in spite of a degree of uncertainty on input parameters. We show results with a real-world case study. Roberto Pietrantuono, Pasqualina Potena, Antonio Pecchia, Daniel Rodríguez-García, Stefano Russo 0001, Luis Fernández-Sanz |
IEEE Trans. Evol. Comput. | 4 |
| 2017 | The Consolidated Tree Construction algorithm in imbalanced defect prediction datasetsabstractIn this short paper, we compare well-known rule/tree classifiers in software defect prediction with the CTC decision tree classifier designed to deal with class imbalanced. It is well-known that most software defect prediction datasets are highly imbalance (non-defective instances outnumber defective ones). In this work, we focused only on tree/rule classifiers as these are capable of explaining the decision, i.e., describing the metrics and thresholds that make a module error prone. Furthermore, rules/decision trees provide the advantage that they are easily understood and applied by project managers and quality assurance personnel. The CTC algorithm was designed to cope with class imbalance and noisy datasets instead of using preprocessing techniques (oversampling or undersampling), ensembles or cost weights of misclassification. The experimental work was carried out using the NASA datasets and results showed that induced CTC decision trees performed better or similar to the rest of the rule/tree classifiers. Igor Ibarguren, Jesús M. Pérez, Javier Muguerza, Daniel Rodríguez-García, Rachel Harrison |
CEC | 4 |
| 2017 | Preliminary Study on Applying Semi-Supervised Learning to App Store AnalysisabstractSemi-Supervised Learning (SSL) is a data mining technique which comes between supervised and unsupervised techniques, and is useful when a small number of instances in a dataset are labelled but a lot of unlabelled data is also available. This is the case with user reviews in application stores such as the Apple App Store or Google Play, where a vast amount of reviews are available but classifying them into categories such as bug related review or feature request is expensive or at least labor intensive. SSL techniques are well-suited to this problem as classifying reviews not only takes time and effort, but may also be unnecessary. In this work, we analyse SSL techniques to show their viability and their capabilities in a dataset of reviews collected from the App Store for both transductive (predicting existing instance labels during training) and inductive (predicting labels on unseen future data) performance. Roger Deocadez, Rachel Harrison, Daniel Rodríguez-García |
EASE | 3 |
| 2017 | Machine Learning Approach to Detect Falls on Elderly People Using Sound
Armando Collado Villaverde, María Dolores Rodríguez-Moreno, David F. Barrero, Daniel Rodríguez-García |
IEA/AIE (1) | 4 |
| 2016 | Triaxial Accelerometer Located on the Wrist for Elderly People's Fall Detection
Armando Collado Villaverde, María Dolores Rodríguez-Moreno, David F. Barrero, Daniel Rodríguez-García |
IDEAL | 4 |
| 2014 | Preliminary comparison of techniques for dealing with imbalance in software defect predictionabstractImbalanced data is a common problem in data mining when dealing with classification problems, where samples of a class vastly outnumber other classes. In this situation, many data mining algorithms generate poor models as they try to optimize the overall accuracy and perform badly in classes with very few samples. Software Engineering data in general and defect prediction datasets are not an exception and in this paper, we compare different approaches, namely sampling, cost-sensitive, ensemble and hybrid approaches to the problem of defect prediction with different datasets preprocessed differently. We have used the well-known NASA datasets curated by Shepperd et al. There are differences in the results depending on the characteristics of the dataset and the evaluation metrics, especially if duplicates and inconsistencies are removed as a preprocessing step. Daniel Rodríguez-García, Israel Herraiz, Rachel Harrison, José Javier Dolado, José Cristóbal Riquelme Santos |
EASE | 1 |
| 2013 | 2nd international workshop on realizing artificial intelligence synergies in software engineering (RAISE 2013)abstractThe RAISE'13 workshop brought together researchers from the AI and software engineering disciplines to build on the interdisciplinary synergies which exist and to stimulate research across these disciplines. The first part of the workshop was devoted to current results and consisted of presentations and discussion of the state of the art. This was followed by a second part which looked over the horizon to seek future directions, inspired by a number of selected vision statements concerning the AI-and-SE crossover. The goal of the RAISE workshop was to strengthen the AI-and-SE community and also develop a roadmap of strategic research directions for AI and software engineering. Rachel Harrison, Sol J. Greenspan, Tim Menzies, Marjan Mernik, Pedro Rangel Henriques, Daniela Carneiro da Cruz, Daniel Rodríguez-García |
ICSE | 7 |
| 2013 | A study of subgroup discovery approaches for defect prediction
Daniel Rodríguez-García, Roberto Ruiz Sánchez, José Cristóbal Riquelme Santos, Rachel Harrison |
Inf. Softw. Technol. | 1 |
| 2012 | Empirical findings on ontology metrics
Miguel-Ángel Sicilia, Daniel Rodríguez-García, Elena García-Barriocanal, Salvador Sánchez-Alonso |
Expert Syst. Appl. | 2 |
| 2012 | Searching for rules to detect defective modules: A subgroup discovery approach
Daniel Rodríguez-García, Roberto Ruiz Sánchez, José Cristóbal Riquelme Santos, Jesús S. Aguilar-Ruiz |
Inf. Sci. | 1 |
| 2012 | Empirical findings on team size and productivity in software development
Daniel Rodríguez-García, Miguel-Ángel Sicilia, Elena García-Barriocanal, Rachel Harrison |
J. Syst. Softw. | 1 |
| 2011 | Multiobjective simulation optimisation in software project managementabstractTraditionally, simulation has been used by project managers in optimising decision making. However, current simulation packages only include simulation optimisation which considers a single objective (or multiple objectives combined into a single fitness function). This paper aims to describe an approach that consists of using multiobjective optimisation techniques via simulation in order to help software project managers find the best values for initial team size and schedule estimates for a given project so that cost, time and productivity are optimised. Using a System Dynamics (SD) simulation model of a software project, the sensitivity of the output variables regarding productivity, cost and schedule using different initial team size and schedule estimations is determined. The generated data is combined with a well-known multiobjective optimisation algorithm, NSGA-II, to find optimal solutions for the output variables. The NSGA-II algorithm was able to quickly converge to a set of optimal solutions composed of multiple and conflicting variables from a medium size software project simulation model. Multiobjective optimisation and SD simulation modeling are complementary techniques that can generate the Pareto front needed by project managers for decision making. Furthermore, visual representations of such solutions are intuitive and can help project managers in their decision making process. Daniel Rodríguez-García, Mercedes Ruiz 0001, José Cristóbal Riquelme Santos, Rachel Harrison |
GECCO | 1 |
| 2011 | Subgroup Discovery for Defect Prediction
Daniel Rodríguez-García, Roberto Ruiz Sánchez, José Cristóbal Riquelme Santos, Rachel Harrison |
SSBSE | 1 |
| 2011 | Comparing Bayesian inference and case-based reasoning as support techniques in the diagnosis of Acute Bacterial Meningitis
Ernesto Ocampo Edye, Mariana Maceiras Cabrera, Silvia Herrera Delgado, Cecilia Maurente, Daniel Rodríguez-García, Miguel-Ángel Sicilia |
Expert Syst. Appl. | 5 |
| 2010 | Defining the Semantics of It Service Management Models using OWL and SWRL
María-Cruz Valiente, Daniel Rodríguez-García, Cristina Vicente-Chicote |
KEOD | 2 |
| 2010 | Defining Software Process Model Constraints with Rules Using OWL and SWRLabstractThe Software & Systems Process Engineering meta-model (SPEM) allows the modelling of software processes using OMG (Object Management Group) standards such as the MOF (Meta-Object Facility) and UML (Unified Modelling Language) making it possible to represent software processes using tools compliant with UML. Process definition encompasses both the static and dynamic structure of roles, tasks and work products together with imposed constraints on those elements. However, the latter requires support for constraint enforcement that is not always directly available in SPEM. Such constraint-checking behaviour could be used to detect possible mismatches between process definitions and the actual processes being carried out in the course of a project. This paper approaches the modelling of such constraints using the SWRL (Semantic Web Rule Language), which is a W3C recommendation. To do so, we need to first represent generic processes modelled with SPEM using an underlying ontology based on the OWL (Ontology Web Language) representation together with data derived from actual projects. Daniel Rodríguez-García, Elena García-Barriocanal, Salvador Sánchez-Alonso, Carlos Rodríguez-Solano |
Int. J. Softw. Eng. Knowl. Eng. | 1 |
| 2007 | Defining a Legal Risk Management Strategy: Process, Legal Risk and Lifecycle
Ricardo J. Rejas-Muslera, Juan Jose Cuadrado-Gallego, Daniel Rodríguez-García |
EuroSPI | 3 |
| 2007 | Convertibility Between IFPUG and COSMIC Functional Size Measurements
Juan Jose Cuadrado-Gallego, Daniel Rodríguez-García, Fernando Machado, Alain Abran |
PROFES | 2 |
| 2007 | Software Project Effort Estimation Based on Multiple Parametric Models Generated Through Data Clustering
Juan Jose Cuadrado-Gallego, Daniel Rodríguez-García, Miguel-Ángel Sicilia, Miguel Garre, Ángel García-Crespo |
J. Comput. Sci. Technol. | 2 |
| 2006 | An empirical study of process-related attributes in segmented software cost-estimation relationships
Juan Jose Cuadrado-Gallego, Miguel-Ángel Sicilia, Miguel Garre, Daniel Rodríguez-García |
J. Syst. Softw. | 4 |
| 2005 | Ontologies of Software Artifacts and Activities: Resource Annotation and Application to Learning Technologies
Miguel-Ángel Sicilia, Juan Jose Cuadrado-Gallego, Daniel Rodríguez-García |
SEKE | 3 |
| 2004 | Assertions in Object Oriented Software Maintenance: Analysis and a Case StudyabstractAssertions had their origin in program verification. For the systems developed in industry, construction of assertions and their use in showing program correctness is a near-impossible task. However, they can be used to show that some key properties are satisfied during program execution. We first present a survey of the special roles that assertions can play in object oriented software construction. We then analyse such assertions by relating them to the case study of an automatic surveillance system. In particular, we address the following two issues: What types of assertions can be used most effectively in the context of object oriented software? How can you discover them and where should they be placed? During maintenance, both the design and the software are continuously changed. These changes can mean that the original assertions, if present, are no longer valid for the new software. Can we automatically derive assertions for the changed software?. Manoranjan Satpathy, Nils T. Siebel, Daniel Rodríguez-García |
ICSM | 3 |
| 2004 | Effective Software Project Management Education through Simulation Models: An Externally Replicated Experiment
Daniel Rodríguez-García, Manoranjan Satpathy, Dietmar Pfahl |
PROFES | 1 |
| 2003 | Latitudinal and longitudinal process diversityabstractAbstract Software processes vary across organizations and over time. Managing this process diversity is a delicate balancing act between creative, healthy diversity and chaos. In this paper, we examine a particular aspect of this issue, namely some relationships between diversity in software processes, software evolution and the quality of software products and processes. Our main contribution is to distinguish between two broad kinds of process diversity, which we call latitudinal and longitudinal process diversity. To illustrate the differences between these two, we examine the case of a medium‐sized system (50 000 lines of C++ code) which has undergone major changes during its lifetime of 10 years. The software was originally developed by an individual academic using a research‐oriented process to develop a standalone proof‐of‐concept system. In a current multi‐team project, involving three industrial and three academic partners, the software has been adapted for integration as a subsystem of a near‐market product. We suggest ways in which the observed process diversity seems to be linked to a change in the software's propensity for evolution, and we discuss the impact of this on both product and process quality. Copyright © 2003 John Wiley & Sons, Ltd. Nils T. Siebel, Stephen Cook 0002, Manoranjan Satpathy, Daniel Rodríguez-García |
J. Softw. Maintenance Res. Pract. | 4 |
| 2002 | An Investigation of Prediction Models for Project ManagementabstractIt has been claimed that dynamic prediction models can be used to help project managers make more accurate estimates than static prediction models. However, such a claim needs to be validated so that project managers can use dynamic models with confidence. In this paper we discuss an experiment we conducted in an academic environment that compared a dynamic model using Bayesian belief networks (BBN) with a static model involving the COCOMO and Akiyama models. The results from this experiment in fact validate the above claim. However we suggest replication of this experiment in order to increase confidence to our results. Daniel Rodríguez-García, Rachel Harrison, Manoranjan Satpathy, José Javier Dolado |
COMPSAC | 1 |
| 2002 | Maintenance of Object Oriented Systems through Re-Engineering: A Case StudyabstractUnregulated evolution of software often leads to software ageing which not only makes the product difficult to maintain but also breaks the consistency between design and implementation. In such a case, it may become necessary to re-engineer the software so that it becomes maintainable again. In this paper we present the case study of the reengineering of the People Tracking subsystem of a surveillance system written in C++. We discuss the problems, the challenges and the approaches taken, and we show how the re-engineered product is now better maintainable. We also discuss the generation of the relevant artefacts - from requirement document through to design document. Manoranjan Satpathy, Nils T. Siebel, Daniel Rodríguez-García |
ICSM | 3 |
| 2002 | Generation of Management Rules through System Dynamics and Evolutionary Computation
Jesús S. Aguilar-Ruiz, José Cristóbal Riquelme Santos, Daniel Rodríguez-García, Isabel Ramos 0002 |
PROFES | 3 |