EDBT 2026 Demo / reviewers in the wild / expert
Ramon Nou
dblp:93/3274
· DBLP profile ↗
23ranked-venue papers
5as first author
7since 2021 · last 2026
0000-0003-0464-1473ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 2 first-author · 3 since 2021Computer networks · 3 · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CloudSkin: AI-Based Learning Plane for Autonomic Management in the Cloud-Edge Continuum
Peini Liu, Joan Oliveras Torra, Marc Palacín, Ramon Nou, Josep Lluís Berral, Jordi Guitart |
COMPSAC | 4 |
| 2025 | Integrating Reliability into Intent-Driven Orchestration on Kubernetes-based Edge NodesabstractPlacing workloads across the Edge and Cloud can be performed through containerization. However, conditions in Edge nodes are very different from Cloud, as environmental factors such as temperature, humidity, voltage provisioning or dust, heavily affect execution performance and device health. Reliability is an important factor when deciding to deploy such load onto Edge devices, indicating their capability to achieve a desired Quality of Service (QoS). For this, research on Edge computing must focus on how to monitor, estimate and manage devices, in a distributed, autonomous and reliable manner. Our current efforts on performance analysis for devices under "wild conditions" are moving towards integrating reliability into orchestrator systems. Here we present our vision and roadmap for expanding Cloud orchestration towards the Edge with technologies capable of providing knowledge about node environmental conditions and mitigate its impact. Through Out-Of-Band telemetry, we can retrieve temperature and power consumption variables from node components, indicating its health and estimating its reliability given external stress factors. In particular, using intent-based orchestration for containerized platforms, such factors can be used for enforcing reliability as a key-performance indicator. The current work in progress focuses on industrial and commercial scenarios, e.g., Edge computing for urban mobility, with road-side nodes performing AI-based Video-Analytics, exposed to uncontrolled weather conditions. The principal objective is to achieve an Edge network orchestration that takes into account node health in an automatic manner for reducing operational costs such as energy consumption, device repair and replacement, while maintaining QoS in the Edge. Josep Lluís Berral, David Aguilera-Luzón, Peini Liu, Ramon Nou, Maria A. Serrano, Angelos Antonopoulos 0001, Javier Santaella Sánchez, Mario José Diván |
ICNP | 4 |
| 2025 | Mobility Usecase: Intelligent Service Migration in Cloud-Edge ContinuumabstractAs industries increasingly embrace digital and intelligent transformation, enterprises face significant challenges in the containerization upgrades for their Artificial Intelligence(AI) applications and dynamic service migration and management in Cloud-Edge continuum(CEC). This paper presents CloudSkin, an innovative platform designed to realize streamlined, seamless and intelligent service migration in Cloud-Edge Continuum by integrating advanced containerization techniques and AI-driven orchestration capabilities. Our approach, leveraging intelligent algorithms for service migration, can seamlessly transit services between cloud and edge environments, ensuring optimised resource allocation and reducing service latency to assure quality of service(QoS). CloudSkin has been enabled in a Mobility Usecase, empowering Cellnex businesses undergoing digital transformation to achieve higher operational efficiency. The experimental results show that compared to traditional reactive service migration, using intelligent proactive service migration can provide better migration detection, improving F1-Score up to 23.5%, and reducing 28.9% the service running time where the service latency violates SLA. Peini Liu, Joan Oliveras Torra, Marc Palacín, Michail Dalgitsis, Maria A. Serrano, Eftychia G. Datsika, Angelos Antonopoulos 0001, Javier Santaella Sánchez, Jordi Guitart, Josep Lluís Berral, Ramon Nou |
ICNP | 11 |
| 2024 | Data-Connector: An Agent-Based Framework for Autonomous ML-Based Smart Management in Cloud-Edge ContinuumabstractMachine Learning (ML) is becoming pervasive and integrated into different kinds of intelligent applications, and the collaborative Cloud-Edge continuum has been introduced as an emerging trend to support their adoption into use cases. However, managing these ML applications in the CloudEdge continuum is challenging due to the ML application's dynamic resource usage with different user loads and Cloud and Edge's dynamic resource availability. We envision machine learning methods that can be used for smart management in this dynamic environment, but how to deploy and utilize them for the adaptation scenario in Cloud-Edge continuum is unknown. This paper proposes an agent-based framework to enable autonomous smart management mechanisms that can be broadly enabled in diverse adaptation scenarios. The agent acts as a data-connector11https://github.com/bsc-scanflow/data-connector, connecting different sources of data, utilizing ML models for decision-making and triggering adaptations in Cloud-edge platforms. The case study shows the feasibility of our proposed data-connector for smart migration of an ML workload in the Cloud-edge continuum. The result shows that the smart migration-enabled Cloud-edge scenario has 11.9% ML application prediction time better than the Cloud scenario without migration. Moreover, with minimal customization, the data connector agent can be adapted for more use cases. Peini Liu, Joan Oliveras Torra, Marc Palacín, Jordi Guitart, Josep Lluís Berral, Ramon Nou |
ICNP | 6 |
| 2023 | Adaptive multi-tier intelligent data manager for ExascaleabstractThe main objective of the ADMIRE project1 is the creation of an active I/O stack that dynamically adjusts computation and storage requirements through intelligent global coordination, the elasticity of computation and I/O, and the scheduling of storage resources along all levels of the storage hierarchy, while offering quality-of-service (QoS), energy efficiency, and resilience for accessing extremely large data sets in very heterogeneous computing and storage environments. We have developed a framework prototype that is able to dynamically adjust computation and storage requirements through intelligent global coordination, separated control, and data paths, the malleability of computation and I/O, the scheduling of storage resources along all levels of the storage hierarchy, and scalable monitoring techniques. The leading idea in ADMIRE is to co-design applications with ad-hoc storage systems that can be deployed with the application and adapt their computing and I/O behaviour on runtime, using malleability techniques, to increase the performance of applications and the throughput of the applications. Jesús Carretero 0001, Francisco Javier García Blas, Marco Aldinucci, Jean-Baptiste Besnard, Jean-Thomas Acquaviva, André Brinkmann, Marc-Andre Vef, Emmanuel Jeannot, Alberto Miranda, Ramon Nou, Morris Riedel, Massimo Torquati, Felix Wolf 0001 |
CF | 10 |
| 2021 | Streamlining distributed Deep Learning I/O with ad hoc file systemsabstractWith evolving techniques to parallelize Deep Learning (DL) and the growing amount of training data and model complexity, High-Performance Computing (HPC) has become increasingly important for machine learning engineers. Although many compute clusters already use learning accelerators or GPUs, HPC storage systems are not suitable for the I/O requirements of DL workflows. Therefore, users typically copy the whole training data to the worker nodes or distribute partitions. Because DL depends on randomized input data, prior work stated that partitioning impacts DL accuracy. Their solutions focused mainly on training I/O performance on a high-speed network but did not cover the data stage-in process, for example. We show in this paper that, in practice, (unbiased) partitioning is not harmful for distributed DL accuracy. Nevertheless, manual partitioning can be error prone and inefficient. Typically, data must be unpacked and shuffled before it is distributed to nodes. We propose a solution that features both: efficient stage-in and fast access to a global namespace to prevent biases. Our architecture is based around an ad hoc storage system relying on a high-speed interconnect allowing an efficient stage-in of DL data sets into a single global namespace. Our proposed solution does not limit access to parts of the data set or relies on data duplication, also relieving the HPC storage system. We obtain high I/O performance during training and ensure minimal interference with communication of the learning workers. The optimizations are transparent to DL applications and their accuracy is not affected by our architecture. Frederic Schimmelpfennig, Marc-Andre Vef, Reza Salkhordeh, Alberto Miranda, Ramon Nou, André Brinkmann |
CLUSTER | 5 |
| 2021 | Arbitration Policies for On-Demand User-Level I/O Forwarding on HPC PlatformsabstractI/O forwarding is a well-established and widely-adopted technique in HPC to reduce contention in the access to storage servers and transparently improve I/O performance. Rather than having applications directly accessing the shared parallel file system, the forwarding technique defines a set of I/O nodes responsible for receiving application requests and forwarding them to the file system, thus reshaping the flow of requests. The typical approach is to statically assign I/O nodes to applications depending on the number of compute nodes they use, which is not always necessarily related to their I/O requirements. Thus, this approach leads to inefficient usage of these resources. This paper investigates arbitration policies based on the applications I/O demands, represented by their access patterns. We propose a policy based on the Multiple-Choice Knapsack problem that seeks to maximize global bandwidth by giving more I/O nodes to applications that will benefit the most. Furthermore, we propose a user-level I/O forwarding solution as an on-demand service capable of applying different allocation policies at runtime for machines where this layer is not present. We demonstrate our approach's applicability through extensive experimentation and show it can transparently improve global I/O bandwidth by up to 85% in a live setup compared to the default static policy. Jean Luca Bez, Alberto Miranda, Ramon Nou, Francieli Zanon Boito, Toni Cortes, Philippe Olivier Alexandre Navaux |
IPDPS | 3 |
| 2020 | Freezing time emulating new and faster devices with virtual machines
Luis C. E. Bona, Alessandro Elias, Andre P. Ziviani, Ramon Nou, Toni Cortes, Marco A. Z. Alves |
CCF Trans. High Perform. Comput. | 4 |
| 2020 | Adaptive request scheduling for the I/O forwarding layer using reinforcement learning
Jean Luca Bez, Francieli Zanon Boito, Ramon Nou, Alberto Miranda, Toni Cortes, Philippe Olivier Alexandre Navaux |
Future Gener. Comput. Syst. | 3 |
| 2020 | GekkoFS - A Temporary Burst Buffer File System for HPC Applications
Marc-Andre Vef, Nafiseh Moti, Tim Süß, Markus Tacke, Tommaso Tocci, Ramon Nou, Alberto Miranda, Toni Cortes, André Brinkmann |
J. Comput. Sci. Technol. | 6 |
| 2019 | NORNS: Extending Slurm to Support Data-Driven Workflows through Asynchronous Data StagingabstractAs HPC systems move into the Exascale era, parallel file systems are struggling to keep up with the I/O requirements from data-intensive problems. While the inclusion of burst buffers has helped to alleviate this by improving I/O performance, it has also increased the complexity of the I/O hierarchy by adding additional storage layers each with its own semantics. This forces users to explicitly manage data movement between the different storage layers, which, coupled with the lack of interfaces to communicate data dependencies between jobs in a data-driven workflow, prevents resource schedulers from optimizing these transfers to benefit the cluster's overall performance. This paper proposes several extensions to job schedulers, prototyped using the Slurm scheduling system, to enable users to appropriately express the data dependencies between the different phases in their processing workflows. It also introduces a new service for asynchronous data staging called NORNS that coordinates with the job scheduler to orchestrate data transfers to achieve better resource utilization. Our evaluation shows that a workflow-aware Slurm exploits node-local storage more effectively, reducing the filesystem I/O contention and improving job running times. Alberto Miranda, Adrian Jackson, Tommaso Tocci, Iakovos Panourgias, Ramon Nou |
CLUSTER | 5 |
| 2019 | YOLO: Speeding Up VM and Docker Boot Time by Reducing I/O Operations
Thuy-Linh Nguyen 0001, Ramon Nou, Adrien Lèbre |
Euro-Par | 2 |
| 2019 | Detecting I/O Access Patterns of HPC Workloads at RuntimeabstractIn this paper, we seek to guide optimization and tuning strategies by identifying the application's I/O access pattern. We evaluate three machine learning techniques to automatically detect the I/O access pattern of HPC applications at runtime: decision trees, random forests, and neural networks. We focus on the detection using metrics from file-level accesses as seen by the clients, I/O nodes, and parallel file system servers. We evaluated these detection strategies in a case study in which the accurate detection of the current access pattern is fundamental to adjust a parameter of an I/O scheduling algorithm. We demonstrate that such approaches correctly classify the access pattern, regarding file layout and spatiality of accesses - into the most common ones used by the community and by I/O benchmarking tools to test new I/O optimization - with up to 99% precision. Furthermore, when applied to our study case, it guides a tuning mechanism to achieve 99% of the performance of an Oracle solution. Jean Luca Bez, Francieli Zanon Boito, Ramon Nou, Alberto Miranda, Toni Cortes, Philippe Olivier Alexandre Navaux |
SBAC-PAD | 3 |
| 2018 | GekkoFS - A Temporary Distributed File System for HPC ApplicationsabstractWe present GekkoFS, a temporary, highly-scalable burst buffer file system which has been specifically optimized for new access patterns of data-intensive High-Performance Computing (HPC) applications. The file system provides relaxed POSIX semantics, only offering features which are actually required by most (not all) applications. It is able to provide scalable I/O performance and reaches millions of metadata operations already for a small number of nodes, significantly outperforming the capabilities of general-purpose parallel file systems. Marc-Andre Vef, Nafiseh Moti, Tim Süß, Tommaso Tocci, Ramon Nou, Alberto Miranda, Toni Cortes, André Brinkmann |
CLUSTER | 5 |
| 2018 | Freezing Time: A New Approach for Emulating Fast Storage Devices Using VMabstractRecently we are seeing a considerable effort from both academy and industry in proposing new technologies for storage devices. Often these devices are not readily available for evaluation and methods to allow performing their tests just from their performance parameters are an important tool for system administrators. Simulators are a traditional approach for carrying out such evaluations, however, they are more suitable for evaluating the storage device as an isolate component, mostly due to time constraints. In this paper, we propose an approach based on virtual machine technology that is capable of emulate storage devices transparently for the operating system allowing evaluation of simulating devices within a real system using any synthetic or real workload. To emulate devices in real environments it is necessary to use the currently available devices as a storage medium which creates a difficulty when the device to be emulated is faster than this storage medium. To circumvent this limitation we introduce a new technique called Freezing Time, which takes advantage of virtual machine pausing mechanism to manipulate the virtual machine clock and hide the real I/O completion time. Our approach can be implemented just requiring the hypervisor to be modified, providing a high degree of compatibility and flexibility since it is not necessary to modify neither the operating system nor the application. We evaluate our tool under a real system using old magnetic disks to emulate faster storage devices. Experiments using our technique presented an average latency error of 6.08% for read operations and 6.78% for write operations when comparing a real to device. Luis C. E. Bona, Alessandro Elias, Andre P. Ziviani, Toni Cortes, Ramon Nou, Marco A. Z. Alves |
MASCOTS | 5 |
| 2018 | ECHOFS: A Scheduler-Guided Temporary Filesystem to Leverage Node-Local NVMSabstractThe growth in data-intensive scientific applications poses strong demands on the HPC storage subsystem, as data needs to be copied from compute nodes to I/O nodes and vice versa for jobs to run. The emerging trend of adding denser, NVM-based burst buffers to compute nodes, however, offers the possibility of using these resources to build temporary file systems with specific I/O optimizations for a batch job. In this work, we present echofs, a temporary filesystem that coordinates with the job scheduler to preload a job's input files into node-local burst buffers. We present the results measured with NVM emulation, and different FS backends with DAX/FUSE on a local node, to show the benefits of our proposal and such coordination. Alberto Miranda, Ramon Nou, Toni Cortes |
SBAC-PAD | 2 |
| 2015 | Performance Impacts with Reliable Parallel File Systems at Exascale Level
Ramon Nou, Alberto Miranda, Toni Cortes |
Euro-Par | 1 |
| 2012 | Energy-efficient and multifaceted resource management for profit-driven virtualized data centers
Íñigo Goiri, Josep Lluís Berral, Josep Oriol Fitó, Ferran Julià, Ramon Nou, Jordi Guitart, Ricard Gavaldà, Jordi Torres |
Future Gener. Comput. Syst. | 5 |
| 2011 | A path to achieving a self-managed Grid middleware
Ramon Nou, Ferran Julià, Kevin Hogan, Jordi Torres |
Future Gener. Comput. Syst. | 1 |
| 2010 | Energy-Aware Scheduling in Virtualized DatacentersabstractThe reduction of energy consumption in large-scale datacenters is being accomplished through an extensive use of virtualization, which enables the consolidation of multiple workloads in a smaller number of machines. Nevertheless, virtualization also incurs some additional overheads (e.g. virtual machine creation and migration) that can influence what is the best consolidated configuration, and thus, they must be taken into account. In this paper, we present a dynamic job scheduling policy for power-aware resource allocation in a virtualized datacenter. Our policy tries to consolidate workloads from separate machines into a smaller number of nodes, while fulfilling the amount of hardware resources needed to preserve the quality of service of each job. This allows turning off the spare servers, thus reducing the overall datacenter power consumption. As a novelty, this policy incorporates all the virtualization overheads in the decision process. In addition, our policy is prepared to consider other important parameters for a datacenter, such as reliability or dynamic SLA enforcement, in a synergistic way with power consumption. The introduced policy is evaluated comparing it against common policies in a simulated environment that accurately models HPC jobs execution in a virtualized datacenter including power consumption modeling and obtains a power consumption reduction of 15% with respect to typical policies. Íñigo Goiri, Ferran Julià, Ramon Nou, Josep Lluís Berral, Jordi Guitart, Jordi Torres |
CLUSTER | 3 |
| 2009 | Autonomic QoS control in enterprise Grid environments using online simulation
Ramon Nou, Samuel Kounev, Ferran Julià, Jordi Torres |
J. Syst. Softw. | 1 |
| 2007 | Should the grid middleware look to self-managing capabilities?abstractGrid technologies have enabled the clustering of a wide variety of geographically distributed resources and services. While providing this, the grid middleware layers that make up the supporting platform for grid applications have taken an increased level of importance in the overall performance of such distributed applications. In this paper we discuss the necessity of introducing self-managing capabilities to the core functionalities of the grid middleware. Our opinions are based on a simple example that shows how Globus Toolkit 4 (GT4) can be driven to a state of unavailability under certain overloading conditions, and how a simple but effective self-managing policy applied to its resource management mechanisms could overcome such an unfavourable scenario. We present an approach showing the benefits that a resource management introduced in the middleware can provide to the user. This approach is the base of a self-managing layer that is being developed for use under more generic conditions Ramon Nou, Ferran Julià, Jordi Torres |
ISADS | 1 |
| 2007 | Monitoring and Analysis Framework for Grid MiddlewareabstractAs the use of complex grid middleware becomes widespread and more facilites are offered by these pieces of software, distributed grid applications are becoming more and more popular. But as grid middleware grows in size and offers more advanced features, they become more complex and heavier, as well as harder to tune. Since the performance of a distributed grid application can be strongly influenced by the operation of the underlying grid middleware, it becomes of extreme importance to study and analyse its behaviour and performance. In this paper we present the eDragon monitoring framework (eDMF), a set of tools that can be used for the instrumentation and analysis of grid middleware, and which provides a unique environment to study the performance of grid applications. The eDMF is composed of a set of specialised monitoring tools as well as by a flexible and powerful performance analysis platform. Additionally we also provide a practical application of the eDMF to the Globus toolkit 4 (GT4), one of the most extended and popular grid middleware, showing how it helped us in the detection and resolution of several job management problems observed in the GT4 middleware Ramon Nou, Ferran Julià, David Carrera 0001, Kevin Hogan, Jordi Caubet, Jesús Labarta, Jordi Torres |
PDP | 1 |