Petar Jovanovic 0001

dblp:34/1397-1 · DBLP profile ↗
← Back
22ranked-venue papers
8as first author
6since 2021 · last 2026
0000-0003-4635-6646ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 16 · 8 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorTheory of computation · 1
YearPublicationVenuePosition
2026 Operationalizing and Automating Data Validation in Data Spaces
abstract
Abstract Data spaces have recently emerged as an innovative paradigm for cross-organizational data sharing. These decentralized environments require sophisticated data governance protocols to ensure compliance with data standards, roles and policies. While current policy-based solutions address enforcement of data access control and usage rights, they lack mechanisms for automated data validation -essential for ensuring data quality for collaborative analytics. To address this gap, we present a knowledge graph-based framework to automate data validation inline with data policies. This framework relies on the concept of policy checkers, which represent high-level and technology-agnostic data validation plans that can be dynamically translated into technology-specific user defined functions (UDFs) for compliance checking. Importantly, the usage of knowledge graphs to describe the policy checkers enhances the transparency and traceability of data validation processes, while the two-stage process (technology-agnostic policy checkers and technology-specific UDFs) accommodate data validation on multimodal data. We accompany the description of our approach with a proof of concept that demonstrates the feasibility of this solution in real data spaces.
Achraf Hmimou, Petar Jovanovic 0001, Sergi Nadal, Oscar Romero 0001, Anna Queralt
Data Sci. Eng.2
2025 Evaluating Quality of Disparate Data Sources: A Discord-Driven Approach
Yeasmin Ara Akter, Alberto Abelló, Petar Jovanovic 0001, Tomer Sagi, Katja Hose
ADBIS3
2024 There is no Data Science without Data Governance: a Proposal Based on Knowledge Graphs
Besim Bilalli, Petar Jovanovic 0001, Sergi Nadal, Anna Queralt, Oscar Romero 0001
DOLAP2
2024 Web API Change-Proneness Prediction
abstract
Change-proneness of software artifacts has been mainly related to the design characteristics and their previous history of changes. While these two aspects are essential and contribute significantly to the prediction, they leave out a critical factor: how the artifacts are being used. In the context of web APIs, consumers represent one of the main drivers of the change. Therefore, we propose a methodology for predicting the change-proneness of web API endpoint interfaces, taking into account not only design and change history but also their usage. Since the evolution of web APIs is and should be usage-driven, the way consumers use an API affects the future changes implemented by providers. Consequently, consumers' usage behavior contains essential information that contributes to identifying endpoints that are more prone to change. By considering the reasons behind changes, we introduce a set of metrics comprising design and usage aspects to be used as variables in prediction. To demonstrate the usefulness of the approach we perform an initial evaluation using a real-world web API. We quantify the introduced metrics using web API documentation, code, and usage logs in order to build a classifier able to predict with 82 % accuracy if an endpoint will change based on its design, history of changes, and usage characteristics.
Rediana Koçi, Xavier Franch, Petar Jovanovic 0001, Alberto Abelló
SANER3
2023 Web API evolution patterns: A usage-driven approach
abstract
As the use of Application Programming Interfaces (APIs) is increasingly growing, their evolution becomes more challenging in terms of the service provided according to consumers’ needs. In this paper, we address the role of consumers’ needs in WAPIs evolution and introduce a process mining pattern-based method to support providers in WAPIs evolution by analyzing and understanding consumers’ behavior, imprinted in WAPI usage logs. We take the position that WAPIs’ evolution should be mainly usage-based, i.e., the way consumers use them should be one of the main drivers of their changes. We start by characterizing the structural relationships between endpoints, and next, we summarize these relationships into a set of behavioral patterns (i.e., usage patterns whose occurrences indicate specific consumers’ behavior like repetitive or consecutive calls), that can potentially imply the need for changes (e.g., creating new parameters for endpoints, merging endpoints). We analyze the logs and extract several metrics for the endpoints and their relationships, to then detect the patterns. We apply our method in two real-world WAPIs from different domains, education, and health, respectively the WAPI of Barcelona School of Informatics at the Polytechnic University of Catalonia (Facultat d’Informàtica de Barcelona, FIB, UPC), and District Health Information Software 2 (DHIS2) WAPI. The feedback from consumers and providers of these WAPIs proved the effectiveness of the detected patterns and confirmed the promising potential of our approach.
Rediana Koçi, Xavier Franch, Petar Jovanovic 0001, Alberto Abelló
J. Syst. Softw.3
2021 Improving Web API Usage Logging
Rediana Koçi, Xavier Franch, Petar Jovanovic 0001, Alberto Abelló
RCIS3
2020 A Data-Driven Approach to Measure the Usability of Web APIs
abstract
Application Programming Interfaces (APIs) are means of communication between applications, hence they can be seen as user interfaces, just with different kind of users, i.e., software or computers. However, the very first consumers of the APIs are humans, namely programmers. Based on the available documentation and the "ease of use" perception (sometimes led by corporate decisions and/or restrictions) they decide to use or not a specific API. In this paper, we propose a data-driven approach to measure web API usability, expressed through the predicted error rate. Following the reviewed state of the art in API usability, we identify a set of usability attributes, and for each of them we propose indicators that web API providers should refer to when developing usable web APIs. Our focus in this paper is on those indicators that can be quantified using the API logs, which indeed reflect the actual behaviour of programmers. Next, we define metrics for the aforementioned indicators, and exemplify them in our use case, applying them on the logs from the web API of District Health Information System (DHIS2) used at World Health Organization (WHO). Using these metrics as features, we build a classifier model to predict the error rate of API endpoints. Besides finding usability issues, we also drill down into the usage logs and investigate the potential causes of these errors.
Rediana Koçi, Xavier Franch, Petar Jovanovic 0001, Alberto Abelló
SEAA3
2019 Classification of Changes in API Evolution
abstract
Applications typically communicate with each other, accessing and exposing data and features by using Application Programming Interfaces (APIs). Even though API consumers expect APIs to be steady and well established, APIs are prone to continuous changes, experiencing different evolutive phases through their lifecycle. These changes are of different types, caused by different needs and are affecting consumers in different ways. In this paper, we identify and classify the changes that often happen to APIs, and investigate how all these changes are reflected in the documentation, release notes, issue tracker and API usage logs. The analysis of each step of a change, from its implementation to the impact that it has on API consumers, will help us to have a bigger picture of API evolution. Thus, we review the current state of the art in API evolution and, as a result, we define a classification framework considering both the changes that may occur to APIs and the reasons behind them. In addition, we exemplify the framework using a software platform offering a Web API, called District Health Information System (DHIS2), used collaboratively by several departments of World Health Organization (WHO).
Rediana Koçi, Xavier Franch, Petar Jovanovic 0001, Alberto Abelló
EDOC3
2019 Mapreduce performance model for Hadoop 2.x
Daria Glushkova, Petar Jovanovic 0001, Alberto Abelló
Inf. Syst.2
2018 Intermediate Results Materialization Selection and Format for Data-Intensive Flows
abstract
Data-intensive flows deploy a variety of complex data transformations to build information pipelines from data sources to different end users. As data are processed, these workflows generate large intermediate results, typically pipelined from one operator to the following ones. Materializing intermediate results, shared among multiple flows, brings benefits not only in terms of performance but also in resource usage and consistency. Similar ideas have been proposed in the context of data warehouses, which are studied under the materialized view selection problem. With the rise of Big Data systems, new challenges emerge due to new quality metrics captured by service level agreements which must be taken into account. Moreover, the way such results are stored must be reconsidered, as different data layouts can be used to reduce the I/O cost. In this paper, we propose a novel approach for automatic selection of multi-objective materialization of intermediate results in data-intensive flows, which can tackle multiple and conflicting quality objectives. In addition, our approach chooses the optimal storage data format for selected materialized intermediate results based on subsequent access patterns. The experimental results show that our approach provides 40% better average speedup with respect to the current state-of-the-art, as well as an improvement on disk access time of 18% as compared to fixed format solutions.
Rana Faisal Munir, Sergi Nadal, Oscar Romero 0001, Alberto Abelló, Petar Jovanovic 0001, Maik Thiele, Wolfgang Lehner
Fundam. Informaticae5
2017 Data generator for evaluating ETL process quality
Vasileios Theodorou, Petar Jovanovic 0001, Alberto Abelló, Emona Nakuçi
Inf. Syst.2
2016 H-WorD: Supporting Job Scheduling in Hadoop with Workload-Driven Data Redistribution
Petar Jovanovic 0001, Oscar Romero 0001, Toon Calders, Alberto Abelló
ADBIS1
2016 Incremental Consolidation of Data-Intensive Multi-Flows
abstract
Business intelligence (BI) systems depend on efficient integration of disparate and often heterogeneous data. The integration of data is governed by data-intensive flows and is driven by a set of information requirements. Designing such flows is in general a complex process, which due to the complexity of business environments is hard to be done manually. In this paper, we deal with the challenge of efficient design and maintenance of data-intensive flows and propose an incremental approach, namely CoAl , for semi-automatically consolidating data-intensive flows satisfying a given set of information requirements. CoAl works at the logical level and consolidates data flows from either high-level information requirements or platform-specific programs. As CoAl integrates a new data flow, it opts for maximal reuse of existing flows and applies a customizable cost model tuned for minimizing the overall cost of a unified solution. We demonstrate the efficiency and effectiveness of our approach through an experimental evaluation using our implemented prototype.
Petar Jovanovic 0001, Oscar Romero 0001, Alkis Simitsis, Alberto Abelló
IEEE Trans. Knowl. Data Eng.1
2015 Supporting Data Integration Tasks with Semi-Automatic Ontology Construction
abstract
Data integration aims to facilitate the exploitation of heterogeneous data by providing the user with a unified view of data residing in different sources. Currently, ontologies are commonly used to represent this unified view in terms of a global target schema due to their flexibility and expressiveness. However, most approaches still assume a predefined target schema and focus on generating the mappings between this schema and the sources.
Rizkallah Touma, Oscar Romero 0001, Petar Jovanovic 0001
DOLAP3
2015 Quarry: Digging Up the Gems of Your Data Treasury
abstract
The design lifecycle of a data warehousing (DW) system is primarily led by requirements of its end-users and the complexity of underlying data sources. The process of designing a multidimensional (MD) schema and back-end extracttransform-load (ETL) processes, is a long-term and mostly manual task. As enterprises shift to more real-time and ’on-the-fly’ decision making, business intelligence (BI) systems require automated means for efficiently adapting a physical DW design to frequent changes of business needs. To address this problem, we present Quarry, an end-to-end system for assisting users of various technical skills in managing the incremental design and deployment of MD schemata and ETL processes. Quarry automates the physical design of a DW system from high-level information requirements. Moreover, Quarry provides tools for efficiently accommodating MD schema and ETL process designs to new or changed information needs of its end-users. Finally, Quarry facilitates the deployment of the generated DW design over an extensible list of execution engines. On-site, we will use a variety of examples to show how Quarry facilitates the complexity of the DW design lifecycle.
Petar Jovanovic 0001, Oscar Romero 0001, Alkis Simitsis, Alberto Abelló, Héctor Candón, Sergi Nadal
EDBT1
2014 Bijoux: Data Generator for Evaluating ETL Process Quality
abstract
Obtaining the right set of data for evaluating the fulfillment of different quality standards in the extract-transform-load (ETL) process design is rather challenging. First, the real data might be out of reach due to different privacy constraints, while providing a synthetic set of data is known as a labor-intensive task that needs to take various combinations of process parameters into account. Additionally, having a single dataset usually does not represent the evolution of data throughout the complete process lifespan, hence missing the plethora of possible test cases. To facilitate such demanding task, in this paper we propose an automatic data generator (i.e., Bijoux). Starting from a given ETL process model, Bijoux extracts the semantics of data transformations, analyzes the constraints they imply over data, and automatically generates testing datasets. At the same time, it considers different dataset and transformation characteristics (e.g., size, distribution, selectivity, etc.) in order to cover a variety of test scenarios. We report our experimental findings showing the effectiveness and scalability of our approach.
Emona Nakuçi, Vasileios Theodorou, Petar Jovanovic 0001, Alberto Abelló
DOLAP3
2014 Engine independence for logical analytic flows
abstract
A complex analytic flow in a modern enterprise may perform multiple, logically independent, tasks where each task uses a different processing engine. We term these multi-engine flows hybrid flows. Using multiple processing engines has advantages such as rapid deployment, better performance, lower cost, and so on. However, as the number and variety of these engines grows, developing and maintaining hybrid flows is a significant challenge because they are specified at a physical level and, so are hard to design and may break as the infrastructure evolves. We address this problem by enabling flow design at a logical level and automatic translation to physical flows. There are three main challenges. First, we describe how flows can be represented at a logical level, abstracting away details of any underlying processing engine. Second, we show how a physical flow, expressed in a programming language or some design GUI, can be imported and converted to a logical flow. In particular, we show how a hybrid flow comprising subflows in different languages can be imported and composed as a single, logical flow for subsequent manipulation. Third, we describe how a logical flow is translated into one or more physical flows for execution by the processing engines. The paper concludes with experimental results and example transformations that demonstrate the correctness and utility of our system.
Petar Jovanovic 0001, Alkis Simitsis, Kevin Wilkinson
ICDE1
2014 BabbleFlow: a translator for analytic data flow programs
abstract
A complex analytic data flow may perform multiple, inter-dependent tasks where each task uses a different processing engine. Such a multi-engine flow, termed a hybrid flow, may comprise subflows written in more than one programming language. However, as the number and variety of these engines grow, developing and maintaining hybrid flows at the physical level becomes increasingly challenging. To address this problem, we present BabbleFlow, a system for enabling flow design at a logical level and automatic translation to physical flows. BabbleFlow translates a hybrid flow expressed in a number of languages to a semantically equivalent hybrid flow expressed in the same or a different set of languages. To this end, it composes the multiple physical flows of a hybrid flow into a single logical representation expressed in a unified flow language called xLM. In doing so, it enables a number of graph transformations such as (de-)composition and optimization. Then, it converts the, possibly transformed, xLM data flow graph into an executable form by expressing it in one or more target programming languages.
Petar Jovanovic 0001, Alkis Simitsis, Kevin Wilkinson
SIGMOD Conference1
2014 A requirement-driven approach to the design and evolution of data warehouses
Petar Jovanovic 0001, Oscar Romero 0001, Alkis Simitsis, Alberto Abelló, Daria Mayorova
Inf. Syst.1
2013 xPAD: a platform for analytic data flows
abstract
As enterprises become more automated, real-time, and data-driven, they need to integrate new data sources and specialized processing engines. The traditional business intelligence architecture of Extract-Transform-Load (ETL) flows, followed by querying, reporting, and analytic operations, is being generalized to analytic data flows that utilize a variety of data types and operations. These complicated flows are difficult to design, implement and maintain since they span a variety of systems. Additionally, new design requirements may be imposed such as design for fault-tolerance, freshness, maintainability, sampling, etc. To reduce development time and maintenance costs, automation is needed. We present xPAD, our platform to manage analytic data flows. xPAD enables flow design. We show how these designs can be optimized, not just for performance, but for other objectives as well. xPAD is engine-agnostic. We show how it can generate executable code for a number of execution engines. It can also import existing flows from other engines and optimize those flows. In that way, it can transform a flow written for one engine into an optimized flow for a different engine. In our demonstration, we will also use various example flows to show optimization for different objectives and comparison of flow execution on different engines.
Alkis Simitsis, Kevin Wilkinson, Petar Jovanovic 0001
SIGMOD Conference3
2012 Integrating ETL Processes from Information Requirements
Petar Jovanovic 0001, Oscar Romero 0001, Alkis Simitsis, Alberto Abelló
DaWaK1
2012 ORE: an iterative approach to the design and evolution of multi-dimensional schemas
abstract
Designing a data warehouse (DW) highly depends on the information requirements of its business users. However, tailoring a DW design that satisfies all business requirements is not an easy task. In addition, complex and evolving business environments result in a continuous emergence of new or changed business needs. Furthermore, for building a correct multidimensional (MD) schema for a DW, the designer should deal with the semantics and heterogeneity of the underlying data sources. To cope with such an inevitable complexity, both at the beginning of the design process and when a potential evolution event occurs, in this paper we present a semi-automatic method, named ORE, for constructing the MD schema in an iterative fashion based on the information requirements. In our approach, we consider each requirement separately and incrementally build the unified MD schema satisfying the entire set of requirements.
Petar Jovanovic 0001, Oscar Romero 0001, Alkis Simitsis, Alberto Abelló
DOLAP1