EDBT 2026 Demo / reviewers in the wild / expert
Daniel Rosendo
dblp:203/6322
· DBLP profile ↗
17ranked-venue papers
8as first author
4since 2021 · last 2025
0000-0003-1175-8426ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 5 first-author · 3 since 2021Computer networks · 4Software engineering, systems software and programming languages · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PROV-AGENT: Unified Provenance for Tracking AI Agent Interactions in Agentic WorkflowsabstractLarge Language Models (LLMs) and other foundation models are increasingly used as the core of AI agents. In agentic workflows, these agents plan tasks, interact with humans and peers, and influence scientific outcomes across federated and heterogeneous environments. However, agents can hallucinate or reason incorrectly, propagating errors when one agent’s output becomes another’s input. Thus, assuring that agents’ actions are transparent, traceable, reproducible, and reliable is critical to assess hallucination risks and mitigate their workflow impacts. While provenance techniques have long supported these principles, existing methods fail to capture and relate agent-centric metadata such as prompts, responses, and decisions with the broader workflow context and downstream outcomes. In this paper, we introduce PROV-AGENT, a provenance model that extends W3C PROV and leverages the Model Context Protocol (MCP) and data observability to integrate agent interactions into end-to-end workflow provenance. Our contributions include: (1) a provenance model tailored for agentic workflows, (2) a near real-time, open-source system for capturing agentic provenance, and (3) a cross-facility evaluation spanning edge, cloud, and HPC environments, demonstrating support for critical provenance queries and agent reliability analysis. Renan Souza 0001, Amal Gueroudji, Stephen DeWitt, Daniel Rosendo, Tirthankar Ghosal, Robert B. Ross, Prasanna Balaprakash, Rafael Ferreira da Silva |
eScience | 4 |
| 2023 | ProvLight: Efficient Workflow Provenance Capture on the Edge-to-Cloud ContinuumabstractModern scientific workflows require hybrid infrastructures combining numerous decentralized resources on the IoT/Edge interconnected to Cloud/HPC systems (aka the Computing Continuum) to enable their optimized execution. Understanding and optimizing the performance of such complex Edge-to-Cloud workflows is challenging. Capturing the provenance of key performance indicators, with their related data and processes, may assist in understanding and optimizing workflow executions. However, the capture overhead can be prohibitive, particularly in resource-constrained devices, such as the ones on the IoT/Edge.To address this challenge, based on a performance analysis of existing systems, we propose ProvLight, a tool to enable efficient provenance capture on the IoT/Edge. We leverage simplified data models, data compression and grouping, and lightweight transmission protocols to reduce overheads. We further integrate ProvLight into the E2Clab framework to enable workflow provenance capture across the Edge-to-Cloud Continuum. This integration makes E2Clab a promising platform for the performance optimization of applications through reproducible experiments.We validate ProvLight at a large scale with synthetic workloads on 64 real-life IoT/Edge devices in the FIT IoT LAB testbed. Evaluations show that ProvLight outperforms state-of-the-art systems like ProvLake and DfAnalyzer in resource-constrained devices. ProvLight is 26—37x faster to capture and transmit provenance data; uses 5—7x less CPU; 2x less memory; transmits 2x less data; and consumes 2—2.5x less energy. ProvLight [1] and E2Clab [2] are available as open-source tools. Daniel Rosendo, Marta Mattoso, Alexandru Costan, Renan Souza 0001, Débora B. Pina, Patrick Valduriez, Gabriel Antoniu |
CLUSTER | 1 |
| 2022 | Distributed intelligence on the Edge-to-Cloud Continuum: A systematic literature review
Daniel Rosendo, Alexandru Costan, Patrick Valduriez, Gabriel Antoniu |
J. Parallel Distributed Comput. | 1 |
| 2021 | Reproducible Performance Optimization of Complex Applications on the Edge-to-Cloud ContinuumabstractIn more and more application areas, we are witnessing the emergence of complex workflows that combine computing, analytics and learning. They often require a hybrid execution infrastructure with IoT devices interconnected to cloud/HPC systems (aka Computing Continuum). Such workflows are subject to complex constraints and requirements in terms of performance, resource usage, energy consumption and financial costs. This makes it challenging to optimize their configuration and deployment. We propose a methodology to support the optimization of real-life applications on the Edge-to-Cloud Continuum. We implement it as an extension of E2Clab, a previously proposed framework supporting the complete experimental cycle across the Edge-to-Cloud Continuum. Our approach relies on a rigorous analysis of possible configurations in a controlled testbed environment to understand their behaviour and related performance tradeoffs. We illustrate our methodology by optimizing Pl@ntNet, a world-wide plant identification application. Our methodology can be generalized to other applications in the Edge-to-Cloud Continuum. Daniel Rosendo, Alexandru Costan, Gabriel Antoniu, Matthieu Simonin, Jean-Christophe Lombardo, Alexis Joly, Patrick Valduriez |
CLUSTER | 1 |
| 2020 | E2Clab: Exploring the Computing Continuum through Repeatable, Replicable and Reproducible Edge-to-Cloud ExperimentsabstractDistributed digital infrastructures for computation and analytics are now evolving towards an interconnected ecosystem allowing complex applications to be executed from IoT Edge devices to the HPC Cloud (aka the Computing Continuum, the Digital Continuum, or the Transcontinuum). Understanding end-to-end performance in such a complex continuum is challenging. This breaks down to reconciling many, typically contradicting application requirements and constraints with low-level infrastructure design choices. One important challenge is to accurately reproduce relevant behaviors of a given application workflow and representative settings of the physical infrastructure underlying this complex continuum. In this paper we introduce a rigorous methodology for such a process and validate it through E2Clab. It is the first platform to support the complete analysis cycle of an application on the Computing Continuum: (i) the configuration of the experimental environment, libraries and frameworks; (ii) the mapping between the application parts and machines on the Edge, Fog and Cloud; (iii) the deployment of the application on the infrastructure; (iv) the automated execution; and (v) the gathering of experiment metrics. We illustrate its usage with a real-life application deployed on the Grid'5000 testbed, showing that our framework allows one to understand and improve performance, by correlating it to the parameter settings, the resource usage and the specifics of the underlying infrastructure. Daniel Rosendo, Pedro Silva 0007, Matthieu Simonin, Alexandru Costan, Gabriel Antoniu |
CLUSTER | 1 |
| 2020 | Availability analysis of design configurations to compose virtual performance-optimized data center systems in next-generation cloud data centersabstractSummary Next‐generation cloud data centers are based on software‐defined data center infrastructures that promote flexibility, automation, optimization, and scalability. The Redfish standard and the Intel Rack Scale Design technology enable software‐defined infrastructure and disaggregate bare‐metal compute, storage, and networking resources into virtual pools to dynamically compose resources and create virtual performance‐optimized data centers (vPODs) tailored to workload‐specific demands. This article proposes four chassis design configurations based on Distributed Management Task Force's Redfish industry standard applied to compose vPOD systems, namely, a fully shared design, partially shared homogeneous design, partially shared heterogeneous design, and not shared design; their main difference is based on the used hardware disaggregation level. Furthermore, we propose models that combine reliability block diagram and stochastic Petri net modeling approaches to represent the complexity of the relationship between the pool of disaggregated hardware resources and their power and cooling sources in a vPOD. These four proposed design configurations were analyzed and compared in terms of availability and component's sensitivity indexes by scaling their configurations considering different data center infrastructure. From the obtained results, we can state that, in general, when one increases the hardware disaggregation, availability is improved. However, after a given point, the availability level of the fully shared, partially shared homogeneous, and partially shared heterogeneous configurations remain almost equal, while the not shared configuration is still able to improve its availability. Daniel Rosendo, Demis Gomes, Guto Leoni Santos, Leylane Silva, André L. C. Moreira, Judith Kelner, Djamel Fawzi Hadj Sadok, Glauco Estácio Gonçalves, Amardeep Mehta, Mattias Wildeman, Patricia Takako Endo |
Softw. Pract. Exp. | 1 |
| 2019 | A Methodology for Automating the Cloud Data Center Availability Assessment
Guto Leoni Santos, Daniel Rosendo, Demis Gomes, Leylane Ferreira, André L. C. Moreira, Djamel Fawzi Hadj Sadok, Judith Kelner, Glauco Estácio Gonçalves, Mattias Wildeman, Patricia Takako Endo |
AINA | 2 |
| 2019 | Maximizing the Availability of Composable Systems of Next-Generation Data CentersabstractThe next-generation data center introduces the refactoring of the traditional data center in order to create pools of disaggregated resource units, such as processors, memory, storage, network, power, and cooling sources, named composable system (CSs) with the purpose of offering flexibility, automation, optimization, and scalability. In this paper, we solve an optimization problem to allocate CSs considering next- generation data centers. The main goal is to maximize the CS availability for the application owner, having its minimum requirements (in terms of CPU, memory, network and storage), and available budget as restrictions. This problem is modeled as a bounded multidimensional knapsack problem, and we solve it using Dynamic Programming (DP), and two Soft Computing approaches: Differential Evolution (DE) and Particle Swarm optimization (PSO). We consider two different scenarios in order to analyze heterogeneity and variability aspects when allocating CSs in a data center. Moreover, we also analyze the importance of system components to give directions and priorities of actions to upgrade the system design. Leylane Ferreira, Patricia Takako Endo, Glauco Estácio Gonçalves, Daniel Rosendo, Guto Leoni Santos, André L. C. Moreira, Judith Kelner, Djamel Fawzi Hadj Sadok, Mattias Wildeman, Amardeep Mehta |
SMC | 4 |
| 2019 | Measuring the impact of data center failures on a cloud-based emergency medical call systemabstractSummary Emergency call services are expected to be highly available in order to minimize the loss of urgent calls and, as a consequence, minimize loss of life due to lack of timely medical response. This service availability depends heavily on the cloud data center on which it is hosted. However, availability information alone cannot provide sufficient understanding of how failures impact the service and users' perception. In this paper, we evaluate the impact of failures on an emergency call system, considering service‐level metrics such as the number of affected calls per failure and the time an emergency service takes until it recovers from a failure. We analyze a real data set from an emergency call center for a large Brazilian city. From stochastic models that represent a cloud data center, we evaluate different data center architectures to observe the impact of failures on the emergency call service. Results show that changing data center's architecture in order to improve availability from two to three nines cannot decrease the average number of affected calls per failure. On the other hand, it can decrease the probability to affect a considerable number of calls at the same time. Demis Gomes, Guto Leoni Santos, Daniel Rosendo, Glauco Estácio Gonçalves, André L. C. Moreira, Judith Kelner, Djamel Fawzi Hadj Sadok, Patricia Takako Endo |
Concurr. Comput. Pract. Exp. | 3 |
| 2019 | DCAV: A software system to evaluate next-generation cloud data center availability through a friendly graphical interfaceabstractSummary To assess the availability of different data center configurations, understand the main root causes of data center failures and represent its low‐level details, such as subsystem's behavior and their interconnections, we have proposed, in previous works, a set of stochastic models to represent different data center architectures (considering three subsystems: power, cooling, and IT) based on the TIA‐942 standard. In this paper, we propose the Data Center Availability (DCAV), a web‐based software system to allow data center operators to evaluate the availability of their data center infrastructure through a friendly interface, without need of understanding the technical details of the stochastic models. DCAV offers an easy step‐by‐step interface to create and configure a data center model. The main goal of the DCAV system is to abstract low‐level details and modeling complexities, becoming the data center availability analysis a simple and less time‐consuming task. André L. C. Moreira, Daniel Rosendo, Demis Gomes, Guto Leoni Santos, Leylane Silva, Carolina Cani D. L., Judith Kelner, Djamel Fawzi Hadj Sadok, Glauco Estácio Gonçalves, Amardeep Mehta, Mattias Wildeman, Patricia Takako Endo |
Softw. Pract. Exp. | 2 |
| 2019 | A methodology to assess the availability of next-generation data centers
Daniel Rosendo, Demis Gomes, Guto Leoni Santos, Glauco Estácio Gonçalves, André L. C. Moreira, Leylane Ferreira, Patricia Takako Endo, Judith Kelner, Djamel Fawzi Hadj Sadok, Amardeep Mehta, Mattias Wildeman |
J. Supercomput. | 1 |
| 2017 | Modeling and analyzing power system failures on cloud servicesabstractMany enterprises rely on cloud infrastructure to host their critical applications (such as trading, banking transaction, airline reservation system, and credit card authorization). The unavailability of these applications may lead to severe consequences that go beyond the financial losses, reaching the cloud provider reputation too. However, to maintain high availability in a cloud data center is a difficult task due to its complexity. The power subsystem is crucial for the entire operation of the data center because it supplies power for all other subsystems, including IT components and cooling equipment. Some studies have already proposed models to evaluate the availability of the power subsystem, but none of them are based on standard redundancy models. Standards guide cloud providers regarding availability, points of failure, and watts per square foot based on components' redundancy. This paper proposes RBD and Petri Net models based on the TIA-942 standard to estimate the availability of the data center power subsystem and analyze how failures on power subsystem impact the availability of critical applications. These models are important to resource planning and decision making by the cloud providers, because they may identify which components they ought to invest in order to improve the availability level. Daniel Rosendo, Patricia Takako Endo, Guto Leoni Santos, Demis Gomes, Glauco Estácio Gonçalves, André L. C. Moreira, Judith Kelner, Djamel Fawzi Hadj Sadok, Mozhgan Mahloo |
CNSM | 1 |
| 2017 | An autonomic and policy-based authorization framework for OpenFlow networksabstractThe Network Access Control (NAC) management is a critical task, especially in current networks that are composed of many heterogeneous things (Internet of Things) connected to share data, resources and Internet access. The Software-Defined Networking (SDN) simplifies the network design and operation, and offers new opportunities (programmability, flexibility, dy-namicity, and standardization) to manage the network. Despite this, the access control management remains a challenge, once managing security policies involves dealing with a large set of access control rules, detecting conflicting policies, defining priorities, delegating rights, and reacting against network state changes and events. This work presents the HACFlow, a novel, autonomic, and policy-based framework for access control management in OpenFlow networks. HACFlow aims to simplify and automate the network management allowing network operators to govern rights of network entities by defining dynamic, fine-grained, and high-level access control policies. We analyzed the performance of HACFlow and compared it against related approaches. Daniel Rosendo, Patricia Takako Endo, Djamel Fawzi Hadj Sadok, Judith Kelner |
CNSM | 1 |
| 2017 | A Network Access Control solution combining OrBAC and SDNabstractStandard Port-based Network Access Control (NAS) with tagged Virtual Local Area Networks (VLANs) systems are useful to authenticate users within an isolated network environment. This approach on its own, however, lacks the flexibility and granularity level that new generation networks based on SDN (Software Defined Networking) can provide. The flow-based access control provides a more appropriate granularity to enforce network policies. In this paper, we propose a novel solution named SDN-based Network Access Control (S-NAC) that provides authentication and authorization of clients and servers based on high-level policies enforced at flow level. The solution has been implemented, deployed and tested over emulated and real networks. Rafael Roque Aschoff, Daniel Rosendo, Marcos Machado, Alexandre J. T. Santos, Djamel Fawzi Hadj Sadok |
IM | 2 |
| 2017 | Prototyping a high availability PaaS: Performance analysis and lessons learnedabstractWith cloud computing consolidation, Platform-as-a-Service (PaaS) has been used as a solution for developing applications with low cost and maximum flexibility. However, an open challenge related to PaaS is the proper handling of multi-tier and stateful applications with support for high availability (HA); and scalability can be considered an essential feature for HA. However, dealing with several instances of the same application that access its state in a common area is not a simple task. This paper presents a novel PaaS framework, named NoPaaS, that supports the deployment of multi-tier and stateful applications assuring their availability according to the Service Availability Forum (SAF) redundancy model. The primary goal of this work is to present NoPaaS framework and prototype, and highlight challenges and open issues when providing multi-tier and stateful applications in high availability clouds. Marcos Machado, Daniel Rosendo, Demis Gomes, André L. C. Moreira, Moises Bezerra, Djamel Fawzi Hadj Sadok, Patricia Takako Endo, Calin Curescu |
IM | 2 |
| 2017 | Evaluating the cooling subsystem availability on a Cloud data centerabstractA data center is divided into three basic subsystems: information technology (IT), power, and cooling. Cooling plays an important role related to data center availability, and a failure in this subsystem may cause an interruption of services. Generally, a redundant cooling subsystem is implemented based on replacing the failed component by the standby one. However, it also can be based on a rotation of computer room air conditioners (CRACs). This paper proposes scalable models that represent the cooling subsystem behavior to evaluate the impact of cooling failures on the data center availability. Models are based on the TIA-942 standard and represent Tiers I and II. We validate our model by comparing our results with the literatures. Our results show that the CRACs' rotation has similar results in availability when compared to the replace strategy. Demis Gomes, Patricia Takako Endo, Glauco Estácio Gonçalves, Daniel Rosendo, Guto Leoni Santos, Judith Kelner, Djamel Fawzi Hadj Sadok, Mozhgan Mahloo |
ISCC | 4 |
| 2017 | Analyzing the IT subsystem failure impact on availability of cloud servicesabstractCloud computing has gained popularity in recent years due to its pay-as-you-go business model, high availability of services, and scalability. Service unavailability does not affect just user experience but is also translated into direct costs for cloud providers and companies. Part of this costs is due to SLA breaches, once interruption time greater than those signed in the contract generate financial penalties. Thus, cloud providers have tried to identify failure points and estimate the availability of their services. This paper proposes models to assess the availability of services running in a cloud data center infrastructure. The models follow the TIA-942 standard. We propose Tier I and IV models using the Reliability Block Diagram (RBD) to allow modeling of different types of applications, and Stochastic Petri Net (SPN) to represent the failure behavior of information technology (IT) components in a data center. We perform stationary analysis to measure the service availability, and sensitivity analysis to understand which metrics have major impacts on data center availability. Guto Leoni Santos, Patricia Takako Endo, Glauco Estácio Gonçalves, Daniel Rosendo, Demis Gomes, Judith Kelner, Djamel Fawzi Hadj Sadok, Mozhgan Mahloo |
ISCC | 4 |