VLDB 2026 Research / reviewers in the wild / expert
Jordi Guitart
dblp:81/5135
· DBLP profile ↗
61ranked-venue papers
8as first author
15since 2021 · last 2026
0000-0003-0751-3100ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 28 · 5 first-author · 5 since 2021Software engineering, systems software and programming languages · 7 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 5 since 2021Computer networks · 6 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CloudSkin: AI-Based Learning Plane for Autonomic Management in the Cloud-Edge Continuum
Peini Liu, Joan Oliveras Torra, Marc Palacín, Ramon Nou, Josep Lluís Berral, Jordi Guitart |
COMPSAC | 6 |
| 2026 | KC-Agent: A Dual-Process Cognitive Architecture for Efficient ML Model ImprovementabstractData drift poses significant challenges for machine learning systems in production, requiring continuous model updates to maintain performance. We present KC-Agent, a dual-process cognitive architecture for automated ML model improvement that combines fast pattern recognition (System 1) with deliberate incremental updates (System 2). Our approach implements structured memory systems enabling System 1 to leverage successful solutions previously discovered by System 2, achieving efficient pattern-based responses without costly re-computation. KC-Agent incorporates atomic change principles and rollback capabilities to ensure reliable, verifiable updates in production environments. We evaluate our method on five datasets including real-world NASA turbofan data with authentic temporal degradation and synthetic datasets with controlled drift scenarios. KC-Agent achieves state-of-the-art performance (76.8% accuracy) while maintaining optimal efficiency (13.2s execution time), outperforming established cognitive architectures: CodeAct (+2.4%), Tree of Thoughts (+3.6%), ReAct (+8.0%), and Reflexion (+8.9%). Consensus evaluation by a panel of state-of-the-art LLMs confirms superior strategic efficacy (8.33/10 Smartness score), significantly outperforming baseline agents. The knowledge consolidation mechanism delivers 91% speedup over the slow variant while maintaining higher accuracy. Our approach demonstrates both theoretical foundations and practical viability for cognitive-inspired automated ML improvement systems capable of handling complex real-world data drift scenarios. Gusseppe Bravo Rocca, Jordi Guitart, Ajay Dholakia, David Ellison, Puneet Jain |
COMPSAC | 2 |
| 2026 | Scalable energy-aware VM allocation on cloud data centers through mathematical programming models
Roberto Meroni, Jordi Guitart |
Future Gener. Comput. Syst. | 2 |
| 2025 | Dynamic In-node Group-Aware Scheduling for Multi-Tenant Machine Learning Services on KubernetesabstractMachine Learning (ML) is becoming pervasive and integrated into different kinds of businesses. Hosting multi-tenant ML services on a unified platform requires efficient orchestration. Service orchestration considers two levels: an in-cluster scheduling to decide the node allocation, followed by an in-node scheduling to manage the resource distribution within the node. Previous works introduced multi-container deployments for ML services, demonstrating that partitioning the ML service and enabling CPU/Memory affinity for each container improves performance. These multi-container deployments have been enabled for Kubernetes at in-cluster level to allocate multiple groups of containers across nodes. However, when the containers are launched in the node, the in-node scheduler lacks awareness of the group information which challenges those containers for a fine-grained resource allocation, especially when multiple groups of containers share the same node. This paper presents an in-node group-aware scheduling mechanism for multiple ML services, where each service scheduling contains a group resource selection and container resource assignment. Moreover, we also provide a dynamic resource controller (DRC) to dynamically reallocate the resource allocation for containers using the mechanism, which monitors the group changes and acts with the real containers' system cgroups adaptation. Our results show that for deploying ML services with the same ML model and different ML models, DRC throughput outperforms other deployment scenarios by up to 258% and 319% respectively. In an experiment deploying a dynamic workload of multiple mixed ML services, DRC outperforms baseline NONE-Single by 242 %, 75 %, and 28 % for the average throughput of services with Mobilenet, Resnet50, and VGG 16, respectively, and also DRC results in a makespan 44% faster than NONE-Single and 4% faster than CM-Multi. Peini Liu, Jordi Guitart |
CLOUD | 2 |
| 2025 | Mobility Usecase: Intelligent Service Migration in Cloud-Edge ContinuumabstractAs industries increasingly embrace digital and intelligent transformation, enterprises face significant challenges in the containerization upgrades for their Artificial Intelligence(AI) applications and dynamic service migration and management in Cloud-Edge continuum(CEC). This paper presents CloudSkin, an innovative platform designed to realize streamlined, seamless and intelligent service migration in Cloud-Edge Continuum by integrating advanced containerization techniques and AI-driven orchestration capabilities. Our approach, leveraging intelligent algorithms for service migration, can seamlessly transit services between cloud and edge environments, ensuring optimised resource allocation and reducing service latency to assure quality of service(QoS). CloudSkin has been enabled in a Mobility Usecase, empowering Cellnex businesses undergoing digital transformation to achieve higher operational efficiency. The experimental results show that compared to traditional reactive service migration, using intelligent proactive service migration can provide better migration detection, improving F1-Score up to 23.5%, and reducing 28.9% the service running time where the service latency violates SLA. Peini Liu, Joan Oliveras Torra, Marc Palacín, Michail Dalgitsis, Maria A. Serrano, Eftychia G. Datsika, Angelos Antonopoulos 0001, Javier Santaella Sánchez, Jordi Guitart, Josep Lluís Berral, Ramon Nou |
ICNP | 9 |
| 2025 | Feature Engineering for Agents: An Adaptive Cognitive Architecture for Interpretable ML Monitoring
Gusseppe Bravo Rocca, Peini Liu, Jordi Guitart, Rodrigo M. Carrillo-Larco, Ajay Dholakia, David Ellison |
AAMAS | 3 |
| 2025 | Embedded scaffolding for teaching and assessing inquiry-based hands-on laboratory on distributed systemsabstractInformation Technology education must cultivate proficiency on distributed systems, including strong hands-on laboratory skills, to meet the needs of the society and the industry. Given the complexity of distributed systems, any successful methodology to teach them to novice students must be scaffolded appropriately to ensure that the students acquire the required degree of expertise. We propose a comprehensive scaffolding approach for inquiry-based hands-on laboratory on a distributed systems course, which guides not only the learning process, but also its assessment. The approach is based mainly on embedded scaffolds, namely explicit coding and experimental milestones and open questions with predefined grades, but also features contingent scaffolds provided by the teacher when additional assistance is needed. We apply the methodology in the context of the subject ‘Distributed Network Systems’ offered by our university. We compare the students' performance during three academic courses using the proposed methodology with respect to the three previous courses that were still using the former methodology. We use both visual representations and planned Analysis of Variance (ANOVA) tests to verify our hypothesis defined as a complex contrast. We find that there is a statistically significant improvement in the students' performance when using the new methodology, both in their grades of the assignments ( F (1, 75.364) = 17.770, p = 6.85 × 10 − 5 ) and, more importantly, also in their grades of the exam questions about the practicals ( F (1, 123.186) = 13.285, p = 3.93 × 10 − 4 ). Our results encourage other instructors to incorporate embedded scaffolds for teaching and assessing their hands-on laboratories on distributed systems. • A scaffolding approach for inquiry-based hands-on laboratory on distributed systems. • The scaffolds guide not only the learning process, but also its assessment. • Explicit coding and experimental milestones and open questions with predefined grades. • New methodology applied in the context of a real subject on distributed systems. • There was a statistically significant improvement in the students' performance. Jordi Guitart |
J. Parallel Distributed Comput. | 1 |
| 2024 | Data-Connector: An Agent-Based Framework for Autonomous ML-Based Smart Management in Cloud-Edge ContinuumabstractMachine Learning (ML) is becoming pervasive and integrated into different kinds of intelligent applications, and the collaborative Cloud-Edge continuum has been introduced as an emerging trend to support their adoption into use cases. However, managing these ML applications in the CloudEdge continuum is challenging due to the ML application's dynamic resource usage with different user loads and Cloud and Edge's dynamic resource availability. We envision machine learning methods that can be used for smart management in this dynamic environment, but how to deploy and utilize them for the adaptation scenario in Cloud-Edge continuum is unknown. This paper proposes an agent-based framework to enable autonomous smart management mechanisms that can be broadly enabled in diverse adaptation scenarios. The agent acts as a data-connector11https://github.com/bsc-scanflow/data-connector, connecting different sources of data, utilizing ML models for decision-making and triggering adaptations in Cloud-edge platforms. The case study shows the feasibility of our proposed data-connector for smart migration of an ML workload in the Cloud-edge continuum. The result shows that the smart migration-enabled Cloud-edge scenario has 11.9% ML application prediction time better than the Cloud scenario without migration. Moreover, with minimal customization, the data connector agent can be adapted for more use cases. Peini Liu, Joan Oliveras Torra, Marc Palacín, Jordi Guitart, Josep Lluís Berral, Ramon Nou |
ICNP | 4 |
| 2024 | TADIL: Task-Agnostic Domain-Incremental Learning Through Task-ID Inference Using Transformer Nearest-Centroid Embeddings
Gusseppe Bravo Rocca, Peini Liu, Jordi Guitart, Ajay Dholakia, David Ellison |
ICPR (29) | 3 |
| 2024 | Practicable live container migrations in high performance computing clouds: Diskless, iterative, and connection-persistentabstractCheckpoint/Restore techniques had been thoroughly used by the High Performance Computing (HPC) community in the context of failure recovery. Given the current trend in HPC to use containerization to obtain fast, customized, portable, flexible, and reproducible deployments of their workloads, as well as efficient and reliable sharing and management of HPC Cloud infrastructures, there is a need to integrate Checkpoint/Restore with containerization in such a way that the freeze time of the application is minimal and live migrations are practicable. Whereas current Checkpoint/Restore tools (such as CRIU) support several options to accomplish this, most of them are rarely exploited in HPC Clouds and, consequently, their potential impact on the performance is barely known. Therefore, this paper explores the use of CRIU’s advanced features to implement diskless, iterative (pre-copy and post-copy) migrations of containers with external network namespaces and established TCP connections, so that memory-intensive and connection-persistent HPC applications can live-migrate. Our extensive experiments to characterize the performance impact of those features demonstrate that properly-configured live migrations incur low application downtime and memory/disk usage and are indeed feasible in containerized HPC Clouds. Jordi Guitart |
J. Syst. Archit. | 1 |
| 2023 | Performance Characterization of Multi-Container Deployment Schemes for Online Learning InferenceabstractOnline machine learning (ML) inference services provide users with an interactive way to request for predictions in realtime. To meet the notable computational requirements of such services, they are increasingly being deployed in the Cloud. In this context, the efficient provisioning and optimization of ML inference services in the Cloud is critical to achieve the required performance and meet the dynamic queries by end-users. Existing provisioning solutions focus on framework parameter tuning and infrastructure resources scaling, without considering deployments based on containerization technologies. The latter promises reproducibility and portability features for ML inferences services. There is limited knowledge about the impact of distinct deployment schemes at the container-level on the performance of online ML inference services, particularly on how to exploit multi-container deployments and its relation with processor and memory affinity. In light of this, in this paper we investigate experimentally the containerization of ML inference services and analyze the performance of multi-container deployments that partition the threads belonging to an online learning application into multiple containers in each node. This paper shares the findings and lessons learned from conducting realistic client patterns on an image classification model across numerous deployment configurations, especially including the impact of container granularity and its potential to exploit processor and memory affinity. Our results indicate that fine-grained multi-container deployments and affinity are useful for improving performance (both throughput and latency). In particular, our experiments on single-node and four-node clusters show up to 69% and 87% performance improvement compared to the single-container deployment, respectively. Peini Liu, Jordi Guitart, Amirhosein Taherkordi |
CLOUD | 2 |
| 2022 | Scanflow-K8s: Agent-based Framework for Autonomic Management and Supervision of ML Workflows in Kubernetes ClustersabstractMachine Learning (ML) projects are currently heavily based on workflows composed of some reproducible steps and executed as containerized pipelines to build or deploy ML models efficiently because of the flexibility, portability, and fast delivery they provide to the ML life-cycle. However, deployed models need to be watched and constantly managed, supervised, and debugged to guarantee their availability, validity, and robustness in unexpected situations. Therefore, containerized ML workflows would benefit from leveraging flexible and diverse autonomic capabilities. This work presents an architecture for autonomic ML workflows with abilities for multi-layered control, based on an agent-based approach that enables autonomic management and supervision of ML workflows at the application layer and the infrastructure layer (by collaborating with the orchestrator). We redesign the Scanflow ML framework to support such multi-agent approach by using triggers, primitives, and strategies. We also implement a practical platform, so-called Scanflow-K8s, that enables autonomic ML workflows on Kubernetes clusters based on the Scanflow agents. MNIST image classification and MLPerf ImageNet classification benchmarks are used as case studies to show the capabilities of Scanflow-K8s under different scenarios. The experimental results demonstrate the feasibility and effectiveness of our proposed agent approach and the Scanflow-K8s platform for the autonomic management of ML workflows in Kubernetes clusters at multiple layers. Peini Liu, Gusseppe Bravo Rocca, Jordi Guitart, Ajay Dholakia, David Ellison, Miro Hodak |
CCGRID | 3 |
| 2022 | Human-in-the-loop online multi-agent approach to increase trustworthiness in ML models through trust scores and data augmentationabstractIncreasing a ML model accuracy is not enough, we must also increase its trustworthiness. This is an important step for building resilient AI systems for safety-critical applications such as automotive, finance, and healthcare. For that purpose, we propose a multi-agent system that combines both machine and human agents. In this system, a checker agent calculates a trust score of each instance (which penalizes overconfidence in predictions) using an agreement-based method and ranks it; then an improver agent filters the anomalous instances based on a human rule-based procedure (which is considered safe), gets the human labels, applies geometric data augmentation, and retrains with the augmented data using transfer learning. We evaluate the system on corrupted versions of the MNIST and FashionMNIST datasets. We get an improvement in accuracy and trust score with just few additional labels compared to a baseline approach. Gusseppe Bravo Rocca, Peini Liu, Jordi Guitart, Ajay Dholakia, David Ellison, Miro Hodak |
COMPSAC | 3 |
| 2022 | Scanflow: A multi-graph framework for Machine Learning workflow management, supervision, and debugginabstractMachine Learning (ML) is more than just training models, the whole workflow must be considered. Once deployed, a ML model needs to be watched and constantly supervised and debugged to guarantee its validity and robustness in unexpected situations. Debugging in ML aims to identify (and address) the model weaknesses in not trivial contexts. Several techniques have been proposed to identify different types of model weaknesses, such as bias in classification, model decay, adversarial attacks, etc., yet there is not a generic framework that allows them to work in a collaborative, modular, portable, iterative way and, more importantly, flexible enough to allow both human- and machine-driven techniques. In this paper, we propose a novel containerized directed graph framework to support and accelerate end-to-end ML workflow management, supervision, and debugging. The framework allows defining and deploying ML workflows in containers, tracking their metadata, checking their behavior in production, and improving the models by using both learned and human-provided knowledge. We demonstrate these capabilities by integrating in the framework two hybrid systems to detect data drift distribution which identify the samples that are far from the latent space of the original distribution, ask for human intervention, and whether retrain the model or wrap it with a filter to remove the noise of corrupted data at inference time. We test these systems on MNIST-C, CIFAR-10-C, and FashionMNIST-C datasets, obtaining promising accuracy results with the help of human involvement. Gusseppe Bravo Rocca, Peini Liu, Jordi Guitart, Ajay Dholakia, David Ellison, Jeffrey Falkanger, Miro Hodak |
Expert Syst. Appl. | 3 |
| 2021 | Performance comparison of multi-container deployment schemes for HPC workloads: an empirical study
Peini Liu, Jordi Guitart |
J. Supercomput. | 2 |
| 2020 | Big data deployment in containerized infrastructures through the interconnection of network namespacesabstractSummary Big Data applications tackle the challenge of fast handling of large streams of data. Their performance is not only dependent on the data frameworks implementation and the underlying hardware but also on the deployment scheme and its potential for fast scaling. Consequently, several efforts have focused on the ease of deployment of Big Data applications, notably through the use of containerization. This technology was indeed raised to bring multitenancy and multiprocessing out of clusters, providing high deployment flexibility through lightweight container images. Recent studies have focused mostly on Docker containers. Notwithstanding, this article is actually interested in recent Singularity containers as they provide more security and support high‐performance computing (HPC) environments and, in this way, they can make Big Data applications benefit from the specialized hardware of HPC. Singularity 2.x, however, does not isolate network resources as required by most Big Data components. Singularity 3.x allows allocating each container with isolated network resources, but their interconnection requires a nontrivial amount of configuration effort. In this context, this article makes a functional contribution in the form of a deployment scheme based on the interconnection of network namespaces, through underlay and overlay networking approaches, to make Big Data applications easily deployable inside Singularity containers. We provide detailed account of our deployment scheme when using both interconnection approaches in the form of a “how‐to‐do‐it” report, and we evaluate it by comparing three Big Data applications based on Hadoop when performing on a bare‐metal infrastructure and on scenarios involving Singularity and Docker instances. Carla Sauvanaud, Ajay Dholakia, Jordi Guitart, Chulho Kim, Peter Mayes |
Softw. Pract. Exp. | 3 |
| 2017 | PaaS-IaaS Inter-Layer Adaptation in an Energy-Aware Cloud EnvironmentabstractCloud computing providers resort to a variety of techniques to improve energy consumption at each level of the cloud computing stack. Most of these techniques consider resource-level energy optimization at IaaS layer. This paper argues energy gains can be obtained by creating a cooperation between the PaaS layer (in charge of hosting the application/service) and the IaaS layer (in charge of handling the computing resources). It presents a novel method based on steering information and decision taking to trigger the PaaS and IaaS layers to adapt their energy mode in service operation, therefore enabling the Cloud stack to actively adapt to changing situations. Experimental results demonstrate such adaptation achieves dynamic energy management in each of the PaaS and IaaS cloud layers. Karim Djemame, Raimon Bosch, Richard E. Kavanagh, Pol Álvarez, Jorge Ejarque, Jordi Guitart, Lorenzo Blasi |
IEEE Trans. Sustain. Comput. | 6 |
| 2016 | Estimation and forecasting of ecological efficiency of virtual machines
Gregory Katsaros, Pascal Stichler, Josep Subirats, Jordi Guitart |
Future Gener. Comput. Syst. | 4 |
| 2016 | Analysis of a trust model for SLA negotiation and enforcement in cloud markets
Mario Macías, Jordi Guitart |
Future Gener. Comput. Syst. | 2 |
| 2016 | A Risk Assessment Framework for Cloud ComputingabstractCloud service providers offer access to their resources through formal service level agreements (SLA), and need well-balanced infrastructures so that they can maximise the quality of service (QoS) they offer and minimise the number of SLA violations. This paper focuses on a specific aspect of risk assessment as applied in cloud computing: methods within a framework that can be used by cloud service providers and service consumers to assess risk during service deployment and operation. It describes the various stages in the service lifecycle whereas risk assessment takes place, and the corresponding risk models that have been designed and implemented. The impact of risk on architectural components, with special emphasis on holistic management support at service operation, is also described. The risk assessor is shown to be effective through the experimental evaluation of the implementation, and is already integrated in a cloud computing toolkit. Karim Djemame, Django Armstrong, Jordi Guitart, Mario Macías |
IEEE Trans. Cloud Comput. | 3 |
| 2015 | Matching renewable energy supply and demand in green datacenters
Íñigo Goiri, Md. Enamul Haque, Kien Le, Ryan Beauchea, Thu D. Nguyen, Jordi Guitart, Jordi Torres, Ricardo Bianchini |
Ad Hoc Networks | 6 |
| 2015 | Assessing and forecasting energy efficiency on Cloud computing platforms
Josep Subirats, Jordi Guitart |
Future Gener. Comput. Syst. | 2 |
| 2014 | A Risk-Based Model for Service Level Agreement Differentiation in Cloud Market Providers
Mario Macías, Jordi Guitart |
DAIS | 2 |
| 2014 | Trust-Aware Operation of Providers in Cloud Markets
Mario Macías, Jordi Guitart |
DAIS | 2 |
| 2014 | Business-driven management of infrastructure-level risks in Cloud providers
Josep Oriol Fitó, Jordi Guitart |
Future Gener. Comput. Syst. | 2 |
| 2014 | SLA negotiation and enforcement policies for revenue maximization and client classification in cloud providers
Mario Macías, Jordi Guitart |
Future Gener. Comput. Syst. | 2 |
| 2013 | Risk-Driven Proactive Fault-Tolerant Operation of IaaS ProvidersabstractIn order to improve service execution in Clouds, the management of Cloud Infrastructure has to take measures to adhere to Service Level Agreements and Business Level Objectives, from the application layer through to how services are supported at the lowest hardware levels. In this paper a risk model methodology and holistic management approach is developed specific to the operation of the Cloud Infrastructure Provider and is applied through improvements to SLA fault tolerance in Cloud Infrastructure. Risk assessments are used to analyse execution specific data from the Cloud Infrastructure and linked to a business driven holistic management component that is part of a Cloud Manager. Initial results show improved eco-efficiency, virtual machine availability and reductions in SLA failure across the whole Cloud infrastructure by applying our combined risk-based fault tolerance approach. Jordi Guitart, Mario Macías, Karim Djemame, Tom Kirkham, Django Armstrong |
CloudCom (1) | 1 |
| 2013 | A service framework for energy-aware monitoring and VM management in Clouds
Gregory Katsaros, Josep Subirats, Josep Oriol Fitó, Jordi Guitart, Pierre Gilet, Daniel Espling |
Future Gener. Comput. Syst. | 4 |
| 2012 | Client Classification Policies for SLA Enforcement in Shared Cloud DatacentersabstractIn Utility Computing business model, the owners of the computing resources negotiate with their potential clients to sell computing power. The terms of the Quality of Service (QoS) and the economic conditions are established in a Service-Level Agreement (SLA). There are many scenarios in which the agreed QoS cannot be provided because of errors in the service provisioning or failures in the system. Since providers have usually different types of clients, according to their relationship with the provider or by the fee that they pay, it is important to minimize the impact of the SLA violations in preferential clients. This paper proposes a set of policies to provide better QoS to preferential clients in such situations. The criterion to classify clients is established according to the relationship between client and provider (external user, internal or another privileged relationship) and the QoS that the client purchases (cheap contracts or extra QoS by paying an extra fee). Most of the policies use key features of virtualization: Selective Violation of the SLAs, Dynamic Scaling of the Allocated Resources, and Runtime Migration of Tasks. The validity of the policies is demonstrated through exhaustive experiments. Mario Macías, Jordi Guitart |
CCGRID | 2 |
| 2012 | Business-driven IT Management for Cloud computing providersabstractNowadays, enterprises have high expectations on Cloud systems to achieve their Business-Level Objectives (BLOs). However, those systems are becoming very complex, thus leading to critical management issues. For these reasons, new self-management strategies of virtualized entities driven by business interests have to be explored. In this paper, we present a business-driven self-management optimization loop, which can firmly contribute to fulfill business strategies of Cloud providers. In this sense, a Business-Driven IT Management (BDIM) model is aimed to assess the impact of IT-related events on business-level metrics. Thereupon those high-level impacts are consumed by a policy management frame-work able to autonomously determine the most suitable IT-level management actions in terms of provider's BLOs compliance. Furthermore, we show the benefits achieved by a Platform as a Service (PaaS) provider when adopting such business-driven management. The results obtained through experimen-tation demonstrate that it is able to maximize, individually and simultaneously, two high-level objectives, in this case economical profit and ecological efficiency. Josep Oriol Fitó, Mario Macías, Ferran Julià, Jordi Guitart |
CloudCom | 4 |
| 2012 | GreenHadoop: leveraging green energy in data-processing frameworksabstractInterest has been growing in powering datacenters (at least partially) with renewable or "green" sources of energy, such as solar or wind. However, it is challenging to use these sources because, unlike the "brown" (carbon-intensive) energy drawn from the electrical grid, they are not always available. This means that energy demand and supply must be matched, if we are to take full advantage of the green energy to minimize brown energy consumption. In this paper, we investigate how to manage a datacenter's computational workload to match the green energy supply. In particular, we consider data-processing frameworks, in which many background computations can be delayed by a bounded amount of time. We propose GreenHadoop, a MapReduce framework for a datacenter powered by a photovoltaic solar array and the electrical grid (as a backup). GreenHadoop predicts the amount of solar energy that will be available in the near future, and schedules the MapReduce jobs to maximize the green energy consumption within the jobs' time bounds. If brown energy must be used to avoid time bound violations, GreenHadoop selects times when brown energy is cheap, while also managing the cost of peak brown power consumption. Our experimental results demonstrate that GreenHadoop can significantly increase green energy consumption and decrease electricity cost, compared to Hadoop. Íñigo Goiri, Kien Le, Thu D. Nguyen, Jordi Guitart, Jordi Torres, Ricardo Bianchini |
EuroSys | 4 |
| 2012 | OPTIMIS: A holistic approach to cloud service provisioning
Ana Juan Ferrer, Francisco Hernández-Rodriguez, Johan Tordsson, Erik Elmroth, Ahmed Ali-Eldin, Csilla Zsigri, Raül Sirvent, Jordi Guitart, Rosa M. Badia, Karim Djemame, Wolfgang Ziegler, Theodosis Dimitrakos, Srijith Krishnan Nair, George Kousiouris, Kleopatra Konstanteli, Theodora A. Varvarigou, Benoit Hudzia, Alexander Kipp, Stefan Wesner, Marcelo Corrales, Nikolaus Forgó, Tabassum Sharif, Craig Sheridan |
Future Gener. Comput. Syst. | 8 |
| 2012 | Energy-efficient and multifaceted resource management for profit-driven virtualized data centers
Íñigo Goiri, Josep Lluís Berral, Josep Oriol Fitó, Ferran Julià, Ramon Nou, Jordi Guitart, Ricard Gavaldà, Jordi Torres |
Future Gener. Comput. Syst. | 6 |
| 2012 | Supporting CPU-based guarantees in cloud SLAs via resource-level QoS metrics
Íñigo Goiri, Ferran Julià, Josep Oriol Fitó, Mario Macías, Jordi Guitart |
Future Gener. Comput. Syst. | 5 |
| 2011 | Intelligent Placement of Datacenters for Internet ServicesabstractPopular Internet services are hosted by multiple geographically distributed data centers. The location of the data centers has a direct impact on the services' response times, capital and operational costs, and (indirect) carbon dioxide emissions. Selecting a location involves many important considerations, including its proximity to population centers, power plants, and network backbones, the source of the electricity in the region, the electricity, land, and water prices at the location, and the average temperatures at the location. As there can be many potential locations and many issues to consider for each of them, the selection process can be extremely involved and time-consuming. In this paper, we focus on the selection process and its automation. Specifically, we propose a framework that formalizes the process as a non-linear cost optimization problem, and approaches for solving the problem. Based on the framework, we characterize areas across the United States as potential locations for data centers, and delve deeper into seven interesting locations. Using the framework and our solution approaches, we illustrate the selection trade offs by quantifying the minimum cost of (1) achieving different response times, availability levels, and consistency times, and (2) restricting services to green energy and chiller-less data centers. Among other interesting results, we demonstrate that the intelligent placement of data centers can save millions of dollars under a variety of conditions. We also demonstrate that the selection process is most efficient and accurate when it uses a novel combination of linear programming and simulated annealing. Íñigo Goiri, Kien Le, Jordi Guitart, Jordi Torres, Ricardo Bianchini |
ICDCS | 3 |
| 2011 | Initial thoughts on business-driven IT management challenges in Cloud computing providersabstractNowadays Cloud computing is recognized as the most emerging computing paradigm. Because of its promising benefits, every day more and more enterprises are relying on Cloud systems. Furthermore, new Cloud business models are appearing, most of them within the SaaS marketplace, which fully depend on PaaS and IaaS providers. In any case, the expectation from businesses that IT (Cloud) services and infrastructures should bring them closer to the achievement of their Business-Level Objectives (BLOs) is spreading. Due to this fact, the presence in Cloud providers of a self-management of Cloud services and infrastructures driven by business-level aspects is mandatory. In this direction, the Business-Driven IT Management (BDIM) discipline has been evolving as the most promising way in the sense of aligning IT (low-level) management decisions with business-level objectives coming from providers themselves, as well as from their users. In this paper, we expose several BDIM challenges on the Cloud computing paradigm. Consequently, we outline key issues for the inclusion of BDIM-related features into the core operation of Cloud providers. Josep Oriol Fitó, Jordi Guitart |
Integrated Network Management | 2 |
| 2011 | Optimal Resource Allocation in a Virtualized Software Aging Platform with Software RejuvenationabstractNowadays, virtualized platforms have become the most popular option to deploy complex enough services. The reason is that virtualization allows resource providers to increase resource utilization. Deployed services are expected to be always available, but these long-running services are especially sensitive to suffer from software aging phenomenon. This term refers to an accumulation of errors, which usually causes resource exhaustion, and eventually makes the service hang/crash. To counteract this phenomenon, a preventive approach to fault management, called software rejuvenation has been proposed. In this paper, we propose a framework which provides transparent and predictive software rejuvenation to web services that suffer software aging on virtualized platforms, achieving high levels of availability. To exploit the provider resources, the framework also seeks to maximize the number of services running simultaneously on the platform, while guaranteeing the resources needed by each service. Javier Alonso 0001, Íñigo Goiri, Jordi Guitart, Ricard Gavaldà, Jordi Torres |
ISSRE | 3 |
| 2011 | GreenSlot: scheduling energy consumption in green datacentersabstractIn this paper, we propose GreenSlot, a parallel batch job scheduler for a datacenter powered by a photovoltaic solar array and the electrical grid (as a backup). GreenSlot predicts the amount of solar energy that will be available in the near future, and schedules the workload to maximize the green energy consumption while meeting the jobs' deadlines. If grid energy must be used to avoid deadline violations, the scheduler selects times when it is cheap. Our results for production scientific workloads demonstrate that Green-Slot can increase green energy consumption by up to 117% and decrease energy cost by up to 39%, compared to a conventional scheduler. Based on these positive results, we conclude that green datacenters and green-energy-aware scheduling can have a significant role in building a more sustainable IT ecosystem. Íñigo Goiri, Ryan Beauchea, Kien Le, Thu D. Nguyen, Md. Enamul Haque, Jordi Guitart, Jordi Torres, Ricardo Bianchini |
SC | 6 |
| 2010 | Characterizing Cloud Federation for Enhancing Providers' ProfitabstractCloud federation has been proposed as a new paradigm that allows providers to avoid the limitation of owning only a restricted amount of resources, which forces them to reject new customers when they have not enough local resources to fulfill their customers' requirements. Federation allows a provider to dynamically outsource resources to other providers in response to demand variations. It also allows a provider that has underused resources to rent part of them to other providers. Both things could make the provider to get more profit when used adequately. This requires that the provider has a clear understanding of the potential of each federation decision, in order to choose the most convenient depending on the environment conditions. In this paper, we present a complete characterization of providers' federation in the Cloud, including decision equations to outsource resources to other providers, rent free resources to other providers (i.e. insourcing), or shutdown unused nodes to save power, and we characterize these decisions as a function of several parameters. Then, we demonstrate in the evaluation section how a provider can enhance its profit by using these equations to exploit federation, and how the different parameters influence which is the best decision on each situation. Íñigo Goiri, Jordi Guitart, Jordi Torres |
IEEE CLOUD | 2 |
| 2010 | Energy-Aware Scheduling in Virtualized DatacentersabstractThe reduction of energy consumption in large-scale datacenters is being accomplished through an extensive use of virtualization, which enables the consolidation of multiple workloads in a smaller number of machines. Nevertheless, virtualization also incurs some additional overheads (e.g. virtual machine creation and migration) that can influence what is the best consolidated configuration, and thus, they must be taken into account. In this paper, we present a dynamic job scheduling policy for power-aware resource allocation in a virtualized datacenter. Our policy tries to consolidate workloads from separate machines into a smaller number of nodes, while fulfilling the amount of hardware resources needed to preserve the quality of service of each job. This allows turning off the spare servers, thus reducing the overall datacenter power consumption. As a novelty, this policy incorporates all the virtualization overheads in the decision process. In addition, our policy is prepared to consider other important parameters for a datacenter, such as reliability or dynamic SLA enforcement, in a synergistic way with power consumption. The introduced policy is evaluated comparing it against common policies in a simulated environment that accurately models HPC jobs execution in a virtualized datacenter including power consumption modeling and obtains a power consumption reduction of 15% with respect to typical policies. Íñigo Goiri, Ferran Julià, Ramon Nou, Josep Lluís Berral, Jordi Guitart, Jordi Torres |
CLUSTER | 5 |
| 2010 | Toward business-driven risk management for Cloud computingabstractThe Cloud computing paradigm is offering an innovative and promising vision concerning the Information and Communications Technology (ICT). Notwithstanding, the use of Cloud resources, which usually are external assets to their consumers, implies risk issues that must be taken into account. Josep Oriol Fitó, Mario Macías, Jordi Guitart |
CNSM | 3 |
| 2010 | Rule-based SLA management for revenue maximisation in Cloud Computing MarketsabstractThis paper introduces several Business Rules for maximising the revenue of Providers in Cloud Computing Markets. These rules apply in both negotiation and execution time, and enforce the achievement of Business-Level Objectives by establishing a bidirectional data flow between market and resource layers. The experiments demonstrate that the revenue is maximized by using both resource data when negotiating, and economic information when managing the resources. Mario Macías, Josep Oriol Fitó, Jordi Guitart |
CNSM | 3 |
| 2010 | Prediction of Job Resource Requirements for Deadline Schedulers to Manage High-Level SLAs on the CloudabstractFor a non IT expert to use services in the Cloud is more natural to negotiate the QoS with the provider in terms of service-level metrics-e.g. job deadlines-instead of resource-level metrics-e.g. CPU MHz. However, current infrastructures only support resource-level metrics-e.g. CPU share and memory allocation-and there is not a well-known mechanism to translate from service-level metrics to resource-level metrics. Moreover, the lack of precise information regarding the requirements of the services leads to an inefficient resource allocation-usually, providers allocate whole resources to prevent SLA violations. According to this, we propose a novel mechanism to overcome this translation problem using an online prediction system which includes a fast analytical predictor and an adaptive machine learning based predictor. We also show how a deadline scheduler could use these predictions to help providers to make the most of their resources. Our evaluation shows: (i) that fast algorithms are able to make predictions with an 11% and 17% of relative error for the CPU and memory respectively; (ii) the potential of using accurate predictions in the scheduling compared to simple yet well-known schedulers. Gemma Reig, Jordi Guitart |
NCA | 3 |
| 2010 | Checkpoint-based fault-tolerant infrastructure for virtualized service providersabstractCrash and omission failures are common in service providers: a disk can break down or a link can fail anytime. In addition, the probability of a node failure increases with the number of nodes. Apart from reducing the provider's computation power and jeopardizing the fulfillment of his contracts, this can also lead to computation time wasting when the crash occurs before finishing the task execution. In order to avoid this problem, efficient checkpoint infrastructures are required, especially in virtualized environments where these infrastructures must deal with huge virtual machine images. This paper proposes a smart checkpoint infrastructure for virtualized service providers. It uses Another Union File System to differentiate read-only from read-write parts in the virtual machine image. In this way, read-only parts can be checkpointed only once, while the rest of checkpoints must only save the modifications in read-write parts, thus reducing the time needed to make a checkpoint. The checkpoints are stored in a Hadoop Distributed File System. This allows resuming a task execution faster after a node crash and increasing the fault tolerance of the system, since checkpoints are distributed and replicated in all the nodes of the provider. This paper presents a running implementation of this infrastructure and its evaluation, demonstrating that it is an effective way to make faster checkpoints with low interference on task execution and efficient task recovery after a node failure. Íñigo Goiri, Ferran Julià, Jordi Guitart, Jordi Torres |
NOMS | 3 |
| 2010 | Using resource-level information into nonadditive negotiation models for cloud Market environmentsabstractMarkets arise as an efficient way of organising resources in Cloud Computing scenarios. In Cloud Computing Markets, Brokers that represent both Clients and Service Providers meet in a Market and negotiate for the sales of resources or services. This paper defends the idea that efficient negotiations require of the usage of resource-level information for increasing the accuracy of negotiated Service Level Agreements and facilitating the achievement of both performance and business goals. A negotiation model based on the maximisation of nonadditive utility functions that considers multiple objectives is defined, and its validity is demonstrated in the experiments. Mario Macías, Jordi Guitart |
NOMS | 2 |
| 2010 | SLA-driven Elastic Cloud Hosting ProviderabstractIt is clear that Cloud computing is and will be a sea change for the Information Technology by changing the way in which both software and hardware are designed and purchased. In this work we address the use of this emerging computing paradigm into web hosting providers in order to avoid its resource management limitations. Thanks to the Cloud approach, resources can be provided in a dynamic way according with the needs of providers and end-users. In this paper, we present an elastic web hosting provider, namely Cloud Hosting Provider (CHP), that makes use of the outsourcing technique in order to take advantage of Cloud computing infrastructures for providing scalability and high availability capabilities to the web applications deployed on it. Furthermore, we pursue the main goal of maximizing the revenue earned by the provider through both the analysis of Service Level Agreements (SLA) and the employment of an economic model. The evaluation exposed demonstrates that the system proposed is able to properly react to the dynamic load received by the web applications and it also achieve the aforesaid revenue maximization of the provider by performing an SLA-aware resource (i.e. web servers) management. Josep Oriol Fitó, Íñigo Goiri, Jordi Guitart |
PDP | 3 |
| 2010 | Exploiting semantics and virtualization for SLA-driven resource allocation in service providersabstractAbstract Resource management is a key challenge that service providers must adequately face in order to accomplish their business goals. This paper introduces a framework, the semantically enhanced resource allocator (SERA), aimed to facilitate service provider management, reducing costs and at the same time fulfilling the QoS agreed with the customers. The SERA assigns resources depending on the information given by the service providers according to its business goals and on the resource requirements of the tasks. Tasks and resources are semantically described and these descriptions are used to infer the resource assignments. Virtualization is used to provide an application specific and isolated virtual environment for each task. In addition, the system supports fine‐grain dynamic resource distribution among these virtual environments based on Service‐Level Agreements. The required adaptation is implemented using agents, guarantying enough resources to each task in order to meet the agreed performance goals. Copyright © 2009 John Wiley & Sons, Ltd. Jorge Ejarque, Marc de Palol, Íñigo Goiri, Ferran Julià, Jordi Guitart, Rosa M. Badia, Jordi Torres |
Concurr. Comput. Pract. Exp. | 5 |
| 2010 | A survey on performance management for internet applicationsabstractAbstract Internet applications have become indispensable for many business and personal processes, turning the performance of these applications into a key issue. For this reason, recent research has comprehensively explored mechanisms for managing the performance of these applications, with special focus on dealing with overload situations and providing QoS guarantees to clients. This paper makes a survey on the different proposals in the literature for managing Internet applications' performance. We present a complete taxonomy that characterizes and classifies these proposals into several categories including request scheduling, admission control, service differentiation, dynamic resource management, service degradation, control theoretic approaches, works using queuing models, observation‐based approaches that use runtime measurements, and overall approaches combining several mechanisms. For each work, we provide a brief description in order to provide the reader with a global understanding of the research progress in this area. Copyright © 2009 John Wiley & Sons, Ltd. Jordi Guitart, Jordi Torres, Eduard Ayguadé |
Concurr. Comput. Pract. Exp. | 1 |
| 2010 | Maximizing revenue in Grid markets using an economically enhanced resource managerabstractAbstract Traditional resource management has had as its main objective the optimization of throughput, based on parameters such as CPU, memory, and network bandwidth. With the appearance of Grid markets, new variables that determine economic expenditure, benefit and opportunity must be taken into account. The Self‐organizing ICT Resource Management (SORMA) project aims at allowing resource owners and consumers to exploit market mechanisms to sell and buy resources across the Grid. SORMA's motivation is to achieve efficient resource utilization by maximizing revenue for resource providers and minimizing the cost of resource consumption within a market environment. An overriding factor in Grid markets is the need to ensure that the desired quality of service levels meet the expectations of market participants. This paper explains the proposed use of an economically enhanced resource manager (EERM) for resource provisioning based on economic models. In particular, this paper describes techniques used by the EERM to support revenue maximization across multiple service level agreements and provides an application scenario to demonstrate its usefulness and effectiveness. Copyright © 2008 John Wiley & Sons, Ltd. Mario Macías, Omer F. Rana, Garry Smith, Jordi Guitart, Jordi Torres |
Concurr. Comput. Pract. Exp. | 4 |
| 2009 | Introducing Virtual Execution Environments for Application Lifecycle Management and SLA-Driven Resource Distribution within Service ProvidersabstractResource management is a key challenge that service providers must adequately face in order to ensure their profitability. This paper describes a proof-of-concept framework for facilitating resource management in service providers, which allows reducing costs and at the same time fulfilling the quality of service agreed with the customers. This is accomplished by means of virtualization. Our approach provides application-specific virtual environments and consolidates them in order to achieve a better utilization of the providers resources. In addition, it implements self-adaptive capabilities for dynamically distributing the providers resources among these virtual environments based on Service Level Agreements. The proposed solution has been implemented as a part of the Semantically-Enhanced Resource Allocator prototype developed within the BREIN European project. The evaluation shows that our prototype is able to react in very short time under changing conditions and avoid SLA violations by rescheduling efficiently the resources. Íñigo Goiri, Ferran Julià, Jorge Ejarque, Marc de Palol, Rosa M. Badia, Jordi Guitart, Jordi Torres |
NCA | 6 |
| 2009 | Efficient Data Management Support for Virtualized Service ProvidersabstractVirtualization has been lately introduced for supporting and simplifying service providers management with promising results. Nevertheless, using virtualization introduces also new challenges that must be considered. One of them relates with the data management in the provider. This paper proposes an innovative approach for performing efficiently all the data-related processes in a virtualized service provider, namely VM creation, VM migration and data stage-in/out. Our solution provides a global repository where clients can upload the task input files and retrieve the output files. In addition, the provider implements a distributed file system (using NFS) in which each node can access its own local disk and the disk of the other nodes. As demonstrated in the evaluation, this allows efficient VM creation and task execution, but also task migration with minimum overhead, while keeping it accessible during the whole process. Íñigo Goiri, Ferran Julià, Jordi Guitart |
PDP | 3 |
| 2008 | SLA-Driven Semantically-Enhanced Dynamic Resource Allocator for Virtualized Service ProvidersabstractIn order to be profitable, service providers must be able to undertake complex management tasks such as provisioning, deployment, execution and adaptation in an autonomic way. This paper introduces a framework, the Semantically-Enhanced Resource Allocator (SERA), aimed to facilitate service provider management, reducing costs and at the same time fulfilling the QoS agreed with the customers. The SERA assigns resources depending on the information given by service providers according to its business goals and on the resource requirements of the tasks. Tasks and resources are semantically described and these descriptions are used to infer the resource assignments. Virtualization is used to provide a full-customized and isolated virtual environment for each task. In addition, the system supports fine-grain dynamic resource distribution among these virtual environments based on SLAs. The required adaptation is implemented using agents, guarantying to each task enough resources to meet the agreed performance goals. Jorge Ejarque, Marc de Palol, Íñigo Goiri, Ferran Julià, Jordi Guitart, Rosa M. Badia, Jordi Torres |
eScience | 5 |
| 2008 | Dynamic CPU provisioning for self-managed secure web applications in SMP hosting platforms
Jordi Guitart, David Carrera 0001, Vicenç Beltran 0001, Jordi Torres, Eduard Ayguadé |
Comput. Networks | 1 |
| 2007 | Differentiated Quality of Service for e-Commerce Applications through Connection Scheduling based on System-Level Thread PrioritiesabstractThe e-commerce Web sites receive a great and varied number of visitors every day. These visitors share the application server's limited resources and when there are too many clients connecting to the Web site, it is possible that they hinder between them, even to overload the application server. These visitors can be divided in different categories, depending on their importance from site viewpoint. Considering the importance that in these Web sites some client connections (e.g. buyers' connections) finish successfully before other connections, in this paper we propose a mechanism to provide different quality of service to the different client categories by assigning different priorities to the threads attending the connections. After observing that Java thread priorities are only applied within the JVM, and moreover, these priorities do not reach the O.S. threads, we propose to schedule threads using the Linux Real Time priorities. Our results demonstrate that different quality of service classes can be supported using this mechanism Javier Alonso 0001, Jordi Guitart, Jordi Torres |
PDP | 2 |
| 2007 | Economically Enhanced Resource Management for Internet Service Utilities
Tim Püschel, Nikolay Borissov, Mario Macías, Dirk Neumann 0001, Jordi Guitart, Jordi Torres |
WISE | 5 |
| 2007 | Designing an overload control strategy for secure e-commerce applications
Jordi Guitart, David Carrera 0001, Vicenç Beltran 0001, Jordi Torres, Eduard Ayguadé |
Comput. Networks | 1 |
| 2005 | A Hybrid Web Server Architecture for Secure e-Business Web Applications
Vicenç Beltran 0001, David Carrera 0001, Jordi Guitart, Jordi Torres, Eduard Ayguadé |
HPCC | 3 |
| 2005 | Session-Based Adaptive Overload Control for Secure Dynamic Web ApplicationsabstractAs dynamic Web content and security capabilities are becoming popular in current Web sites, the performance demand on application servers that host the sites is increasing, leading sometimes these servers to overload. As a result, response times may grow to unacceptable levels and the server may saturate or even crash. In this paper we present a session-based adaptive overload control mechanism based on SSL (secure socket layer) connections differentiation and admission control. The SSL connections differentiation is a key factor because the cost of establishing a new SSL connection is much greater than establishing a resumed SSL connection (it reuses an existing SSL session on server). Considering this big difference, we have implemented an admission control algorithm that prioritizes the resumed SSL connections to maximize performance on session-based environments and limits dynamically the number of new SSL connections accepted depending on the available resources and the current number of connections in the system to avoid server overload. In order to allow the differentiation of resumed SSL connections from new SSL connections we propose a possible extension of the Java Secure Sockets Extension (JSSE) API. Our evaluation on Tomcat server demonstrates the benefit of our proposal for preventing server overload. Jordi Guitart, David Carrera 0001, Vicenç Beltran 0001, Jordi Torres, Eduard Ayguadé |
ICPP | 1 |
| 2003 | Complete instrumentation requirements for performance analysis of Web based technologiesabstractIn this paper we present the eDragon environment, a research platform created to perform complete performance analysis of new Web-based technologies. eDragon enables the understanding of how application servers work in both sequential and parallel platforms offering a new insight in the usage of system resources. The environment is composed of a set of instrumentation modules, a performance analysis and visualization tool and a set of experimental methodologies to perform complete performance analysis of Web-based technologies. This paper describes the design and implementation of this research platform and highlights some of its main functionalities. We will also show how a detailed analytical view can be obtained through the application of a bottom-up strategy, starting with a group of system events and advancing to more complex performance metrics using a continuous derivation process. David Carrera 0001, Jordi Guitart, Jordi Torres, Eduard Ayguadé, Jesús Labarta |
ISPASS | 2 |
| 2001 | Performance Analysis Tools for Parallel Java Applications on Shared-memory SystemsabstractIn this paper we describe an instrumentation environment for the performance analysis and visualization of parallel applications written in JOMP, an OpenMP-like interface for Java. The environment includes two complementary approaches. The first one has been designed to provide a detailed analysis of the parallel behavior at the JOMP programming model level. At this level, the user is faced with parallel, work-sharing and synchronization constructs, which are the core of JOMP. The second mechanism has been designed to support an in-depth analysis of the threaded execution inside the Java virtual machine (JVM). At this level of analysis, the user is faced with the supporting threads layer monitors and conditional variables. The paper discusses the implementation of both mechanisms and evaluates the overhead incurred by them. Jordi Guitart, Jordi Torres, Eduard Ayguadé, J. Mark Bull |
ICPP | 1 |
| 2001 | Strategies for the efficient exploitation of loop-level parallelism in JavaabstractAbstract This paper analyzes the overheads incurred in the exploitation of loop‐level parallelism using Java Threads and proposes some code transformations that minimize them. The transformations avoid the intensive use of Java Threads and reduce the number of classes used to specify the parallelism in the application (which reduces the time for class loading). The use of such transformations results in promising performance gains that may encourage the use of Java for exploiting loop‐level parallelism in the framework of OpenMP. On average, the execution time for our synthetic benchmarks is reduced by 50% from the simplest transformation when eight threads are used. The paper explores some possible enhancements to the Java threading API oriented towards improving the application–runtime interaction. Copyright © 2001 John Wiley & Sons, Ltd. José Oliver 0002, Jordi Guitart, Eduard Ayguadé, Nacho Navarro, Jordi Torres |
Concurr. Comput. Pract. Exp. | 2 |