Ismael Caballero 0001

dblp:37/3025 · also Ismael Caballero Muñoz-Reja · DBLP profile ↗
← Back
24ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0002-5189-1427ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 10 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 DMN4DQ+: Optimising data repair to enhance data usability
abstract
Data quality has become crucial in decision-making and data analysis. There is an intrinsic relationship between data quality and usability; however, acceptable levels of data quality depend on the contextual requirements and operational priorities of an organisation. The context of use, business needs, and the organisation’s appetite for risk all influence the usability of data. Achieving adequate levels of usability sometimes requires specific corrections, which can be costly and may incur have far-reaching consequences. This paper introduces the concept of target usability, by representing the minimum level of usability determined by business analysts at which data records can be used without compromising organisational performance. When data records fail to meet this threshold, a combination of corrective actions can be implemented to improve both quality and usability. To achieve near-optimal outcomes, the data quality analyst can effectively combine these sets of actions. This paper proposes DMN4DQ+, an extension of DMN4DQ, where the optimal combination of corrective actions can be derived from the application of constraint optimisation techniques based on the data quality rules described in decision models, the cost model of the actions, and on the usability profile of the record to be improved. The development of a technological stack has conveniently supported DMN4DQ+, and it has been evaluated using a real dataset, thereby demonstrating its applicability and performance.
Álvaro Valencia-Parra, Angel Jesus Varela-Vaca, Luisa Parody, Ismael Caballero 0001, María Teresa Gómez-López
Expert Syst. Appl.4
2024 BIGOWL4DQ: Ontology-driven approach for Big Data quality meta-modelling, selection and reasoning
abstract
Data quality should be at the core of many Artificial Intelligence initiatives from the very first moment in which data is required for a successful analysis. Measurement and evaluation of the level of quality are crucial to determining whether data can be used for the tasks at hand. Conscientious of this importance, industry and academia have proposed several data quality measurements and assessment frameworks over the last two decades. Unfortunately, there is no common and shared vocabulary for data quality terms. Thus, it is difficult and time-consuming to integrate data quality analysis within a (Big) Data workflow for performing Artificial Intelligence tasks. One of the main reasons is that, except for a reduced number of proposals, the presented vocabularies are neither machine-readable nor processable, needing human processing to be incorporated. This paper proposes a unified data quality measurement and assessment information model. This model can be used in different environments and contexts to describe data quality measurement and evaluation concerns. The model has been developed as an ontology to make it interoperable and machine-readable. For better interoperability and applicability, this ontology, BIGOWL4DQ, has been developed as an extension of a previously developed ontology for describing knowledge management in Big Data analytics. This extended ontology provides a data quality measurement and assessment framework required when designing Artificial Intelligence workflows and integrated reasoning capacities. Thus, BIGOWL4DQ can be used to describe Big Data analysis and assess the data quality before the analysis. Our proposal has been validated with two use cases. First, the semantic proposal has been assessed using an academic use case. And second, a real-world case study within an Artificial Intelligence workflow has been conducted to endorse our work.
Cristóbal Barba-González, Ismael Caballero 0001, Angel Jesus Varela-Vaca, José A. Cruz-Lemus, María Teresa Gómez-López, Ismael Navas-Delgado
Inf. Softw. Technol.2
2022 BR4DQ: A methodology for grouping business rules for data quality evaluation
abstract
Data quality evaluation is built upon data quality measurement results. “Data quality evaluation” uses the “data quality rules” representing the risk appetite of the organization to decide on the usability of the data; “data quality measurement” uses the business rules describing the “data requirements” or “data specifications” to determine the validity of the data. Consequently, to conduct meaningful and useful data quality evaluations, business rules must be first completely identified and captured at the beginning of the evaluation to perform sound measurements. We propose that the evaluation leads to better and more interpretable and useful results when the potential contribution of these business rules to the measurement of the data quality characteristics is first evaluated, avoiding the inclusion in the evaluation of those not having potential contribution and the resulting waste of resources. Considering this, we feel that for a better management of business rules for data quality evaluation, it makes sense to group all business rules having an important contribution to the evaluation of data quality characteristics, something that other business rules management methodologies have not covered yet. Through our experiences in conducting industrial projects of data quality evaluations we identified six problems when collecting and grouping the business rules. These problems make data quality evaluation processes less efficient and more costly. The main contribution of this paper is a methodology to systematically collect, group and validate the business rules to avoid or to alleviate these problems. For the sake of generalization, comparability, and reusability, we propose to do the grouping for data quality characteristics and properties defined in ISO/IEC 25012 and ISO/IEC 25024, respectively. Lastly, we validate the methodology in three case studies of real projects. From this validation, it is possible to raise the conclusion that the methodology is useful, applicable in the real world, and valid to capture and group the business rules used as a basis for data quality evaluation.
Ismael Caballero 0001, Fernando Gualo, Moisés Rodríguez 0001, Mario Piattini
Inf. Syst.1
2022 Multisource and temporal variability in Portuguese hospital administrative datasets: Data quality implications
abstract
BACKGROUND: Unexpected variability across healthcare datasets may indicate data quality issues and thereby affect the credibility of these data for reutilization. No gold-standard reference dataset or methods for variability assessment are usually available for these datasets. In this study, we aim to describe the process of discovering data quality implications by applying a set of methods for assessing variability between sources and over time in a large hospital database. METHODS: We described and applied a set of multisource and temporal variability assessment methods in a large Portuguese hospitalization database, in which variation in condition-specific hospitalization ratios derived from clinically coded data were assessed between hospitals (sources) and over time. We identified condition-specific admissions using the Clinical Classification Software (CCS), developed by the Agency of Health Care Research and Quality. A Statistical Process Control (SPC) approach based on funnel plots of condition-specific standardized hospitalization ratios (SHR) was used to assess multisource variability, whereas temporal heat maps and Information-Geometric Temporal (IGT) plots were used to assess temporal variability by displaying temporal abrupt changes in data distributions. Results were presented for the 15 most common inpatient conditions (CCS) in Portugal. MAIN FINDINGS: Funnel plot assessment allowed the detection of several outlying hospitals whose SHRs were much lower or higher than expected. Adjusting SHR for hospital characteristics, beyond age and sex, considerably affected the degree of multisource variability for most diseases. Overall, probability distributions changed over time for most diseases, although heterogeneously. Abrupt temporal changes in data distributions for acute myocardial infarction and congestive heart failure coincided with the periods comprising the transition to the International Classification of Diseases, 10th revision, Clinical Modification, whereas changes in the Diagnosis-Related Groups software seem to have driven changes in data distributions for both acute myocardial infarction and liveborn admissions. The analysis of heat maps also allowed the detection of several discontinuities at hospital level over time, in some cases also coinciding with the aforementioned factors. CONCLUSIONS: This paper described the successful application of a set of reproducible, generalizable and systematic methods for variability assessment, including visualization tools that can be useful for detecting abnormal patterns in healthcare data, also addressing some limitations of common approaches. The presented method for multisource variability assessment is based on SPC, which is an advantage considering the lack of gold standard for such process. Properly controlling for hospital characteristics and differences in case-mix for estimating SHR is critical for isolating data quality-related variability among data sources. The use of IGT plots provides an advantage over common methods for temporal variability assessment due its suitability for multitype and multimodal data, which are common characteristics of healthcare data. The novelty of this work is the use of a set of methods to discover new data quality insights in healthcare data.
Júlio Souza, Ismael Caballero 0001, João Vasco Santos, Mariana Lobo, Andreia Pinto, João Viana, Carlos Sáez 0001, Fernando Lopes 0003, Alberto Freitas
J. Biomed. Informatics2
2021 Measuring Variability in Acute Myocardial Infarction Coding Using a Statistical Process Control and Probabilistic Temporal Data Quality Control Approaches
Júlio Souza, Ismael Caballero 0001, João Vasco Santos, Mariana Lobo, Andreia Pinto, João Viana, Carlos Sáez 0001, Alberto Freitas
WorldCIST (2)2
2021 DMN4DQ: When data quality meets DMN
Álvaro Valencia-Parra, Luisa Parody, Angel Jesus Varela-Vaca, Ismael Caballero 0001, María Teresa Gómez-López
Decis. Support Syst.4
2021 Data quality certification using ISO/IEC 25012: Industrial experiences
Fernando Gualo, Moisés Rodríguez 0001, Javier Verdugo, Ismael Caballero 0001, Mario Piattini
J. Syst. Softw.4
2020 Protocol for Analysis of Root Causes of Problems Affecting the Quality of the Diagnosis Related Group-Based Hospital Data: A Rapid Review and Delphi Process
Mariana Lobo, Ana Raquel Oliveira, João Vasco Santos, Vera Alonso, Fernando Lopes 0003, André Ramalho, Júlio Souza, João Viana, Ismael Caballero 0001, Alberto Freitas
WorldCIST (1)10
2020 Towards a software quality certification of master data-based applications
Fernando Gualo, Ismael Caballero 0001, Moisés Rodríguez 0001
Softw. Qual. J.2
2020 Measuring data credibility and medical coding: a case study using a nationwide Portuguese inpatient database
Júlio Souza, Diana Pimenta, Ismael Caballero 0001, Alberto Freitas
Softw. Qual. J.3
2018 Improving the experience of teaching Scrum
abstract
Scrum dramatically shortens the feedback loop between customer and developers as well as between requirement list and functional implementations. Scrum poses to turn small teams into self-managers of their own work. Despite the increasing importance of Scrum in the software development Industry, Academia is not providing a suitable response to meet the demand of professionals. This is because, in many cases, reference curricula do not consider Scrum with a great relevance in knowledge areas nor in time scheduled for it. Even so, Academia should be able to educate and provide software engineers with enough Scrum knowledge and skills. To collaborate to this aim, we present our efforts for improving the teaching/learning experience when dealing with time limitations: we have verified that when using a comparison strategy (regarding traditional software development methodologies like Unified Process, (UP)), students improve their learning process. To validate our results, students were assessed through a twofold assessment: (i) based on Scrum certification exams plus, and with (ii) a questionnaire based on the previously mentioned comparison (Scrum vs UP). The main conclusion raised from this investigation is that students' learning experience is really more satisfactory when Scrum concepts are introduced after a comparison with UP ones.
Ricardo Pérez-Castillo, Ismael Caballero 0001, Moisés Rodríguez 0001
EDUCON2
2016 PAIS-DQ: Extending process-aware information systems to support data quality in PAIS life-cycle
abstract
The successful execution of a Business Process implies to use data with an adequate level of quality, thereby enabling the output of processes to be obtained in accordance with users requirements. The necessity to be aware of the data quality in the business processes is known, but the problem is how the incorporation of data quality management can affect and increase the complexity of the software development that supports the business process life-cycle. In order to gain advantages that data quality management can provide, organizations need to introduce mechanisms aimed at checking whether data satisfies the established data-quality requirements. Desirably, the implementation, deployment and use of these mechanisms should not interfere into the regular working of the business processes. In order to enable this independence, we propose the PAIS-DQ framework as an extension of the classical Process-Aware Information System (PAIS) proposal. The PAIS-DQ addresses the concerns related to data quality management activities by minimizing the required time for the software developers. In addition, with the aim of guiding developers in the use of PAIS-DQ, a methodology has been also provided to facilitate organizations to deal with complex concerns. The methodology renders our proposal applicable in practice, and has been applied to a case study where a service architecture implementing the standard ISO/IEC 8000-100:2009 parts 100 to 140 is included.
Luisa Parody, María Teresa Gómez-López, Isabel Bermejo, Ismael Caballero 0001, Rafael M. Gasca, Mario Piattini
RCIS4
2016 MAMD: Towards a Data Improvement Model Based on ISO 8000-6X and ISO/IEC 33000
Ana G. Carretero, Ismael Caballero 0001, Mario Piattini
SPICE2
2016 A Data Quality in Use model for Big Data
Jorge Merino, Ismael Caballero 0001, Bibiano Rivas, Manuel A. Serrano, Mario Piattini
Future Gener. Comput. Syst.2
2013 Software modernization by recovering Web services from legacy databases
abstract
ABSTRACT Databases are considered to be a valuable asset for organizations because they contain all those organizations’ persistent pieces of data. Both databases and the information systems that use them undergo erosion as a consequence of uncontrolled maintenance over time. However, when information systems evolve to become modernized versions of them, existing databases must not be discarded because they contain much valuable business knowledge that is not present anywhere else. Some of the software industry's current demands, such as time‐to‐market developments and the provision of software as services entail additional challenges in the reuse of legacy systems during software modernization. This paper addresses this problem and proposes a reengineering process that follows model‐driven development principles to recover Web services from legacy databases. The Web services that are mined manage access to legacy databases without discarding them. Legacy databases can thus be used by modernized information systems in service‐oriented environments. The adoption of this process is facilitated by the implementation of a support tool, which is used to conduct an industrial case study involving a real‐life legacy database. The study demonstrates that the proposal reduces development efforts and improves the return of investment by extending the lifespan of legacy databases. Copyright © 2012 John Wiley & Sons, Ltd.
Ricardo Pérez-Castillo, Ignacio García Rodríguez de Guzmán, Ismael Caballero 0001, Mario Piattini
J. Softw. Evol. Process.3
2012 An approach to web-based Personal Health Records filtering using fuzzy prototypes and data quality criteria
Francisco P. Romero 0001, Ismael Caballero 0001, Jesús Serrano-Guerrero, José Angel Olivas
Inf. Process. Manag.2
2010 A Systematic Literature Review of How to Introduce Data Quality Requirements into a Software Product Development
César Guerra-García, Ismael Caballero 0001, Mario Piattini
ENASE2
2010 Preparing Students and Engineers for Global Software Development: A Systematic Review
abstract
In recent years, the evolution of Global Software Development (GSD) has grown both rapidly and significantly, and although the efficiency of this new type of development has been proven, some challenging issues must still be confronted. Of all these, our research line is focused on designing the specific training that members of virtual teams must receive. Universities and companies therefore need to design training schemas to deal with the specifics of GSD, which are principally related to communication difficulties and time and cultural differences. In this work we present the findings of a Systematic Literature Review in the field of GSD training and teaching. Our intention is twofold: on the one hand we wish to discover the existing strategies and proposals available up to the present day, and on the other hand we wish to identify the open challenges, that will be helpful for practitioners and researchers in the future.
Miguel J. Monasor, Aurora Vizcaíno, Mario Piattini, Ismael Caballero 0001
ICGSE4
2008 DQRDFS - Towards a Semantic Web Enhanced with Data Quality
Ismael Caballero 0001, Eugenio Verbo, Coral Calero, Mario Piattini
WEBIST (1)1
2008 A proposal for a set of attributes relevant for Web portal data quality
Angélica Caro, Coral Calero, Ismael Caballero 0001, Mario Piattini
Softw. Qual. J.3
2006 A First Approach to a Data Quality Model for Web Portals
Angélica Caro, Coral Calero, Ismael Caballero 0001, Mario Piattini
ICCSA (3)3
2006 Defining a quality model for portal data
abstract
Advances in technology and the use of the Internet have favoured the appearance of a great variety of Web applications, among them Web Portals. These applications are important information sources and/or means of accessing information. Many people need to obtain information by means of these applications and they need to ensure that this information is suitable for the use they want to give it. In recent years, several research projects were conducted on topic of Web Data Quality. However, there is still a lack of specific proposals for the data quality of portals. In this paper we introduce a model for the data quality in Web portals.
Angélica Caro, Coral Calero, Ismael Caballero 0001, Mario Piattini
ICWE3
2006 Towards a Data Quality Model for Web Portals - Research in Progress
Angélica Caro, Coral Calero, Ismael Caballero 0001, Mario Piattini
WEBIST (1)3
2006 Defining a Data Quality Model for Web Portals
Angélica Caro, Coral Calero, Ismael Caballero 0001, Mario Piattini
WISE3