EDBT 2026 Demo / reviewers in the wild / expert
Javier Prades
dblp:121/1434
· DBLP profile ↗
20ranked-venue papers
8as first author
7since 2021 · last 2026
0000-0003-3349-2200ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 8 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Edge-Based Auto-Labeling for Multiclass TinyML Application in IoT EnvironmentsabstractIntelligent Environments require perception systems capable of adapting to evolving tasks, heterogeneous sensing conditions, and the resource constraints of large-scale IoT deployments. This work shows an edge-centric pipeline in which a highcapacity model, specifically YOLO12n, operates at the gateway to automatically label visual data and guide the specialization of TinyML models deployed on low-power devices. Focusing on two representative mobility-related classes, cat (present in COCO) and scooter (absent from COCO), we analyze the behavior of YOLO models under zero-shot conditions, during finetuning, and when incrementally integrating new classes. Results show that YOLO12n consistently outperforms YOLO11n in zeroshot evaluation and reaches near-perfect detection accuracy (mAP@50 up to 0.995) after only a few epochs of fine-tuning, with inference times below 5 ms on GPU-equipped edge nodes. When adding the new scooter class, the model rapidly adapts, yet subsequent fine-tuning on cats reveals strong catastrophic forgetting. Rehearsal-based retraining effectively mitigates this degradation, even when using replay buffers as small as$\mathbf{1 0 - 2 0 o r i g i n a l}$dataset. These findings demonstrate that lightweight edge auto-labeling combined with efficient replay mechanisms enables sustainable, privacy-preserving, and continuously adaptive perception pipelines for next-generation Intelligent Environments and IoT monitoring infrastructures. Floreal Acebrón, Javier Prades, Erika Rosas, Juan-Carlos Cano, Pietro Manzoni, José M. Cecilia |
IE | 2 |
| 2026 | Enhancing the performance of GPU acceleration in virtual environments: Thoroughly benchmarking the rigidity of mediated device passthroughabstractVirtualization has been the key element for the growth of cloud computing. Historically, GPUs have been complex devices to virtualize efficiently. The mediated passthrough mechanism is usually leveraged. However, it implies a rigid association between the virtual domain and the virtual GPU, which impairs overall system GPU performance. In this paper we propose the use of Network GPGPU system (NGS) to improve the performance of GPUs in virtual domains. Our proposal is compared to NVIDIA vGPU, the most widely used mechanism for virtualizing CUDA-enabled GPUs today. Results show throughput benefits of approximately 20%, speedup of up to 28% in the execution time of the applications, improved overall GPU utilization (over 85%), and lower energy consumption per job (up to 15.34%). Javier Prades, Carlos Reaño, Federico Silla |
Future Gener. Comput. Syst. | 1 |
| 2025 | Towards efficient stream monitoring: A systematic approach for model selection and continuous improvement in Tiny Machine Learning applicationsabstractMeasuring ephemeral stream flows is essential for ecological and hydrological studies. However, their intermittent nature and remote locations pose challenges for conventional monitoring methods, which often consume excessive energy to capture rare events. We address this with BODOQUE (Bimodal Observational Device for Optimizing Quantification of Ephemeral streams), a dual-mode system that leverages Tiny Machine Learning (TinyML) on low-power microcontrollers. The system remains in an energy-saving sensing state and activates high-precision measurements only when water flow is detected. We present a model selection methodology that balances detection accuracy with inference cost, enabling reliable operation within hardware constraints. To enhance adaptability in diverse environments, we developed a specialized component that facilitates dataset expansion through new field samples. This supports ongoing retraining to maintain model performance under changing conditions. A comprehensive evaluation using real-world data demonstrates that our system can achieve up to 97% annual energy savings compared to traditional continuous monitoring approaches. Benjamín Arratia, Erika Rosas, Javier Prades, Salvador Peña-Haro, José M. Cecilia, Pietro Manzoni |
Eng. Appl. Artif. Intell. | 3 |
| 2024 | NGS: A network GPGPU system for orchestrating remote and virtual acceleratorsabstractIn General-Purpose computing on Graphics Processing Unit (GPGPU), the use of CPUs is combined with that of GPUs. CPUs are used for sequential code, while GPUs are used for parallel code. GPGPU has been enabled by two key factors: (i) the massively parallel architecture of GPUs, which allows thousands of single cores to run parallel code; and (ii) the development of platforms, such as CUDA, that simplify implementing code for GPUs. GPGPU has established itself as the standard computing system in most computing fields due to the great improvements it brings. However, its use is not without problems, such as GPU underutilization, high cost, power consumption, etc. In this paper we present NGS (Network GPGPU System) to address the underutilization of GPUs in computing centers. NGS orchestrates the concurrent access to GPGPU resources from different nodes of the cluster by leveraging the remote GPU virtualization mechanism and the NVML library by NVIDIA. In this way, NGS enables different nodes of the cluster to access remote GPUs as if they were local at the same time that this access is guaranteed to be carried out without collisions. The main novelty is that NGS offers a global and standard solution independent of the computing environment used. Experimental results show up to 4x improvements compared to popular approaches. Javier Prades, Carlos Reaño, Federico Silla |
J. Syst. Archit. | 1 |
| 2023 | BODOQUE: An Energy-Efficient Flow Monitoring System for Ephemeral StreamsabstractEffective environmental monitoring is crucial for managing global environmental challenges and providing the necessary data for Environmental Intelligence (EI). This discipline involves the integration of data from various sources to gain a comprehensive understanding of specific regions or processes. In this paper, we introduce BODOQUE, a hardware-software infrastructure to monitor water flows in ephemeral streams where the water rarely flows with great force. BODOQUE uses a low-power TinyML-based camera to detect the presence of water, activating a more complex system to measure flow only when the water flows, thereby optimizing energy consumption. This device is being deployed in the Segura basin, Murcia, Spain. This region is grappling with severe environmental issues that affect the Mar Menor, a unique saltwater lagoon. This paper focuses on the power-saving capabilities of BODOQUE, comparing the energy consumption of different edge devices running the code that measures water flow in the streams. Our goal is to determine the optimal hardware setup for the system based on our experiments, which involve performance and energy consumption tests. The results provide valuable information for future environmental monitoring systems, considering the best balance among the device's cost, performance, and energy consumption. Benjamín Arratia, Javier Prades, Salvador Peña-Haro, José M. Cecilia, Pietro Manzoni |
MobiHoc | 2 |
| 2023 | Using remote GPU virtualization techniques to enhance edge computing devices
José M. Cecilia, Juan Morales-García, Baldomero Imbernon, Javier Prades, Juan-Carlos Cano, Federico Silla |
Future Gener. Comput. Syst. | 4 |
| 2021 | Improving the management efficiency of GPU workloads in data centers through GPU virtualizationabstractSummary Graphics processing units (GPUs) are currently used in data centers to reduce the execution time of compute‐intensive applications. However, the use of GPUs presents several side effects, such as increased acquisition costs and larger space requirements. Furthermore, GPUs require a nonnegligible amount of energy even while idle. Additionally, GPU utilization is usually low for most applications. In a similar way to the use of virtual machines, using virtual GPUs may address the concerns associated with the use of these devices. In this regard, the remote GPU virtualization mechanism could be leveraged to share the GPUs present in the computing facility among the nodes of the cluster. This would increase overall GPU utilization, thus reducing the negative impact of the increased costs mentioned before. Reducing the amount of GPUs installed in the cluster could also be possible. However, in the same way as job schedulers map GPU resources to applications, virtual GPUs should also be scheduled before job execution. Nevertheless, current job schedulers are not able to deal with virtual GPUs. In this paper, we analyze the performance attained by a cluster using the remote Compute Unified Device Architecture middleware and a modified version of the Slurm scheduler, which is now able to assign remote GPUs to jobs. Results show that cluster throughput, measured as jobs completed per time unit, is doubled at the same time that the total energy consumption is reduced up to 40%. GPU utilization is also increased. Sergio Iserte, Javier Prades, Carlos Reaño, Federico Silla |
Concurr. Comput. Pract. Exp. | 2 |
| 2020 | Improving the performance of physics applications in atom-based clusters with rCUDA
Federico Silla, Javier Prades, Elvira Baydal, Carlos Reaño |
J. Parallel Distributed Comput. | 2 |
| 2019 | Analyzing the performance/power tradeoff of the rCUDA middleware for future exascale systems
Carlos Reaño, Javier Prades, Federico Silla |
J. Parallel Distributed Comput. | 2 |
| 2019 | GPU-Job Migration: The rCUDA CaseabstractVirtualization techniques have been shown to report benefits to data centers and other computing facilities. In this regard, not only virtual machines allow to reduce the size of the computing infrastructure while increasing overall resource utilization, but also virtualizing individual components of computers may provide significant benefits. This is the case, for instance, for the remote GPU virtualization technique, implemented in several frameworks during the recent years. The large degree of flexibility provided by the remote GPU virtualization technique can be further increased by applying the migration mechanism to it, so that the GPU part of applications can be live-migrated to another GPU elsewhere in the cluster during execution time in a transparent way. In this paper we present the implementation of the migration mechanism within the rCUDA remote GPU virtualization middleware. Furthermore, we present a thorough performance analysis of the implementation of the migration mechanism within rCUDA. To that end, we leverage both synthetic and real production applications as well as three different generations of NVIDIA GPUs. Additionally, two different versions of the InfiniBand interconnect are used in this study. Several use cases are provided in order to show the extraordinary benefits that the GPU-job migration mechanism can report to data centers. Javier Prades, Federico Silla |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2018 | Enhancing large-scale docking simulation on heterogeneous systems: An MPI vs rCUDA study
Baldomero Imbernon, Javier Prades, Domingo Giménez, José M. Cecilia, Federico Silla |
Future Gener. Comput. Syst. | 2 |
| 2017 | A Live Demo for Showing the Benefits of Applying the Remote GPU Virtualization Technique to Cloud ComputingabstractCloud computing has become pervasive nowadays. Additionally, cloud computing customers increasingly demand the use of accelerators such as CUDA GPUs. This has motivated that Amazon, for example, provides virtual machine instances comprising up to 16 NVIDIA GPUs. However, the use of GPUs in cloud computing deployments is not exempt from important concerns. In order to overcome many of these concerns, the remote GPU virtualization technique can be used. In this paper we present the design of a live demo to be used in exhibitions in order to show the benefits of using such a virtualization technique in the context of cloud computing systems. The demo was designed to be technically sound at the same time that it draws the attention of attendees. The demo was successfully used in the recent SuperComputing'16 exhibition, attracting more than 100 people to the booth. Javier Prades, Federico Silla |
CCGrid | 1 |
| 2017 | On the benefits of the remote GPU virtualization mechanism: The rCUDA caseabstractSummary Graphics processing units (GPUs) are being adopted in many computing facilities given their extraordinary computing power, which makes it possible to accelerate many general purpose applications from different domains. However, GPUs also present several side effects, such as increased acquisition costs as well as larger space requirements. They also require more powerful energy supplies. Furthermore, GPUs still consume some amount of energy while idle, and their utilization is usually low for most workloads. In a similar way to virtual machines, the use of virtual GPUs may address the aforementioned concerns. In this regard, the remote GPU virtualization mechanism allows an application being executed in a node of the cluster to transparently use the GPUs installed at other nodes. Moreover, this technique allows to share the GPUs present in the computing facility among the applications being executed in the cluster. In this way, several applications being executed in different (or the same) cluster nodes can share 1 or more GPUs located in other nodes of the cluster. Sharing GPUs should increase overall GPU utilization, thus reducing the negative impact of the side effects mentioned before. Reducing the total amount of GPUs installed in the cluster may also be possible. In this paper, we explore some of the benefits that remote GPU virtualization brings to clusters. For instance, this mechanism allows an application to use all the GPUs present in the computing facility. Another benefit of this technique is that cluster throughput, measured as jobs completed per time unit, is noticeably increased when this technique is used. In this regard, cluster throughput can be doubled for some workloads. Furthermore, in addition to increase overall GPU utilization, total energy consumption can be reduced up to 40%. This may be key in the context of exascale computing facilities, which present an important energy constraint. Other benefits are related to the cloud computing domain, where a GPU can be easily shared among several virtual machines. Finally, GPU migration (and therefore server consolidation) is one more benefit of this novel technique. Federico Silla, Sergio Iserte, Carlos Reaño, Javier Prades |
Concurr. Comput. Pract. Exp. | 4 |
| 2017 | Multi-tenant virtual GPUs for optimising performance of a financial risk application
Javier Prades, Blesson Varghese, Carlos Reaño, Federico Silla |
J. Parallel Distributed Comput. | 1 |
| 2016 | Increasing the Performance of Data Centers by Combining Remote GPU Virtualization with SlurmabstractThe use of Graphics Processing Units (GPUs) presents several side effects, such as increased acquisition costs as well as larger space requirements. Furthermore, GPUs require a non-negligible amount of energy even while idle. Additionally, GPU utilization is usually low for most applications. Using the virtual GPUs provided by the remote GPU virtualization mechanism may address the concerns associated with the use of these devices. However, in the same way as workload managers map GPU resources to applications, virtual GPUs should also be scheduled before job execution. Nevertheless, current workload managers are not able to deal with virtual GPUs. In this paper we analyze the performance attained by a cluster using the rCUDA remote GPU virtualization middleware and a modified version of the Slurm workload manager, which is now able to map remote virtual GPUs to jobs. Results show that cluster throughput is doubled at the same time that total energy consumption is reduced up to 40%. GPU utilization is also increased. Sergio Iserte, Javier Prades, Carlos Reaño, Federico Silla |
CCGrid | 2 |
| 2016 | CUDA acceleration for Xen virtual machines in infiniband clusters with rCUDAabstractMany data centers currently use virtual machines (VMs) to achieve a more efficient usage of hardware resources. However, current virtualization solutions, such as Xen, do not easily provide graphics processing unit (GPU) accelerators to applications running in the virtualized domain with the flexibility usually required in data centers (i.e., managing virtual GPU instances and concurrently sharing them among several VMs). Remote GPU virtualization frameworks such as the rCUDA solution may address this problem. Javier Prades, Carlos Reaño, Federico Silla |
PPoPP | 1 |
| 2015 | Acceleration-as-a-Service: Exploiting Virtualised GPUs for a Financial ApplicationabstractHow can GPU acceleration be obtained as a service in a cluster? This question has become increasingly significant due to the inefficiency of installing GPUs on all nodes of a cluster. The research reported in this paper is motivated to address the above question by employing rCUDA (remote CUDA), a framework that facilitates Acceleration-as-a-Service (AaaS), such that the nodes of a cluster can request the acceleration of a set of remote GPUs on demand. The rCUDA framework exploits virtualisation and ensures that multiple nodes can share the same GPU. In this paper we test the feasibility of the rCUDA framework on a real-world application employed in the financial risk industry that can benefit from AaaS in the production setting. The results confirm the feasibility of rCUDA and highlight that rCUDA achieves similar performance compared to CUDA, provides consistent results, and more importantly, allows for a single application to benefit from all the GPUs available in the cluster without loosing efficiency. Blesson Varghese, Javier Prades, Carlos Reaño, Federico Silla |
e-Science | 2 |
| 2015 | On the design of a new dynamic credit-based end-to-end flow control mechanism for HPC clusters
Javier Prades, Federico Silla, Holger Fröning, Mondrian Nüssle, José Duato |
Parallel Comput. | 1 |
| 2014 | SLURM Support for Remote GPU Virtualization: Implementation and Performance StudyabstractSLURM is a resource manager that can be leveraged to share a collection of heterogeneous resources among the jobs in execution in a cluster. However, SLURM is not designed to handle resources such as graphics processing units (GPUs). Concretely, although SLURM can use a generic resource plugin (GRes) to manage GPUs, with this solution the hardware accelerators can only be accessed by the job that is in execution on the node to which the GPU is attached. This is a serious constraint for remote GPU virtualization technologies, which aim at providing a user-transparent access to all GPUs in cluster, independently of the specific location of the node where the application is running with respect to the GPU node. In this work we introduce a new type of device in SLURM, "rgpu", in order to gain access from any application node to any GPU node in the cluster using rCUDA as the remote GPU virtualization solution. With this new scheduling mechanism, a user can access any number of GPUs, as SLURM schedules the tasks taking into account all the graphics accelerators available in the complete cluster. We present experimental results that show the benefits of this new approach in terms of increased flexibility for the job scheduler. Sergio Iserte, Adrián Castelló 0001, Rafael Mayo 0002, Enrique S. Quintana-Ortí, Federico Silla, José Duato, Carlos Reaño, Javier Prades |
SBAC-PAD | 8 |
| 2012 | A New End-to-End Flow-Control Mechanism for High Performance Computing ClustersabstractHigh Performance Computing usually leverages messaging libraries such as MPI or GASNet in order to exchange data among processes in large-scale clusters. Furthermore, these libraries make use of specialized low-level networking layers in order to retrieve as much performance as possible from hardware interconnects such as Infini Band or Myrinet, for example. EXTOLL is another emerging technology targeted for high performance clusters. These specialized low-level networking layers require some kind of flow control in order to prevent buffer overflows at the received side. In this paper we present a new flow control mechanism that is able to adapt the buffering resources used by a process according to the parallel application communication pattern and the varying activity among communicating peers. The tests carried out in a 64-node 1024-core EXTOLL cluster show that our new dynamic flow-control mechanism provides extraordinarily high buffer efficiency along with very low overhead, which is reduced between 8 and 10 times. Javier Prades, Federico Silla, José Duato, Holger Fröning, Mondrian Nüssle |
CLUSTER | 1 |