EDBT 2026 Demo / reviewers in the wild / expert
Roberto R. Expósito
dblp:86/11495
· DBLP profile ↗
30ranked-venue papers
12as first author
8since 2021 · last 2026
0000-0002-2077-1473ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 21 · 8 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2Artificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Weight-based Disk I/O Scaling for Serverless ContainersabstractAbstract Disk bandwidth is a critical resource for I/O-intensive applications that must transfer large volumes of data to and from persistent storage. Most multi-tenant infrastructures efficiently allocate CPU and memory resources to concurrent workloads, but typically lack mechanisms for allocating I/O bandwidth. As a result, users often resort to exclusive node reservations to avoid disk contention, which can lead to underutilisation of other node resources if not fully exploited. Another common issue is that users do not know the exact resource requirements of their applications. Even when this is known, applications rarely maintain peak resource usage throughout their entire execution, resulting in wasted resources that could otherwise benefit other users. Today, many users prefer cloud serverless platforms because of their ease of use and flexible billing. However, these platforms have inherent limitations and may not be suitable for workloads with specific requirements. In this paper, we present a serverless scaling mechanism that dynamically adjusts disk I/O bandwidth for containerised applications by scaling their allocation up or down based on real-time usage and configurable weights. In addition, the system incorporates automatic extension management capabilities for virtual disk devices, such as logical volumes. Our approach can be integrated with other serverless scaling mechanisms, such as CPU and memory management, to provide a comprehensive resource scaling solution. The experimental results have shown significant performance improvements, with overall runtime reductions of up to 53% for concurrent I/O-intensive workloads compared to running them without serverless capabilities. Óscar Castellanos-Rodríguez, Roberto R. Expósito, Jonatan Enes, Juan Touriño |
J. Grid Comput. | 2 |
| 2025 | Reproducibility Report for SC25 Paper MXBLAS: Accelerating 8-bit Deep Learning with a Unified Micro-Scaled GEMM LibraryabstractThis reproducibility report provides details about the artifact evaluation done with regards to the Artifact Description and Evaluation appendix of SC25 paper MXBLAS: Accelerating 8-bit Deep Learning with a Unified Micro-Scaled GEMM Library by Weihu Wang et al. The work was done as part of the Reproducibility Initiative of SC25. The author is a member of the SC25 Reproducibilty Committee. Roberto R. Expósito |
SC | 1 |
| 2024 | Automated Approach for Accurate CPU Power ModellingabstractPower supply is a limiting factor when increasing the computing capacity of supercomputers. As a consequence, power consumption has become one of the biggest challenges in the field of High Performance Computing (HPC). In order to develop energy-efficient tools (e.g., frameworks, applications), it is essential to have an accurate power consumption modelling. Al-though previous works proposed a wide variety of approaches to model CPU power consumption, building models in an automated and adaptable way to changing scenarios and predicting power with high precision remains complex due to multiple factors (e.g., training and test workloads, model variables). In this paper, we present a set of tools to fully automate the process of modelling power consumption using CPU time series data. More specifically, our proposal includes two tools: (1) CPUPowerWatcher, which gathers CPU metrics during the execution of user-configurable workloads; and (2) CPUPowerSeer, which builds models to predict CPU power consumption (e.g., polynomial regression) from different CPU variables (e.g., usage, clock frequency) using time series data. Thus, multiple models can be created and evaluated easily, allowing the selection of an optimal model for a specific workload. The experiments conducted by combining these tools allow analysing the impact of novel factors on CPU power consumption, such as the type of CPU usage generated by different workloads or how the CPU cores are allocated to them. In addition, the accuracy of six regression models is compared when predicting CPU- and I/O-intensive workloads using two different core allocations. Tomé Maseda, Jonatan Enes, Roberto R. Expósito, Juan Touriño |
CLUSTER | 3 |
| 2024 | Serverless-like platform for container-based YARN clustersabstractServerless computing is an emerging paradigm that has gained a lot of relevance in recent years, as it allows users to consume computing resources without worrying about the underlying infrastructure and pay only for what they actually use. Most current services that implement this paradigm typically rely on the Function-as-a-Service (FaaS) model, which works perfectly for simple applications based on stateless functions triggered by specific events. However, these services are not designed to run more complex applications with intricate interactions, usually presenting a significant degree of configuration difficulty and/or low ability to customise the execution environment. They also tend to be designed for short and simple workloads, with some services even limiting their maximum runtime to just a few minutes. In this paper, we present a platform based on Hadoop YARN oriented to the execution of Big Data workloads in a containerised and serverless way, so that the resources allocated to such containers are automatically and dynamically scaled according to their actual usage. An experimental evaluation has been carried out to compare our serverless-like platform with a standard YARN deployment when executing Big Data workloads concurrently. Our results have shown experimental evidence of enhancing both performance and overall resource efficiency, providing runtime reductions and resource usage improvements of up to 41% and 50%, respectively. Óscar Castellanos-Rodríguez, Roberto R. Expósito, Jonatan Enes, Guillermo L. Taboada, Juan Touriño |
Future Gener. Comput. Syst. | 2 |
| 2024 | BigDEC: A multi-algorithm Big Data tool based on the k-mer spectrum method for scalable short-read error correctionabstractDespite the significant improvements in both throughput and cost provided by modern Next-Generation Sequencing (NGS) platforms, sequencing errors in NGS datasets can still degrade the quality of downstream analysis. Although state-of-the-art correction tools can provide high accuracy to improve such analysis, they are limited to apply a single correction algorithm while also requiring long runtimes when processing large NGS datasets. Furthermore, current parallel correctors generally only provide efficient support for shared-memory systems lacking the ability to scale out across a cluster of multicore nodes, or they require the availability of specific hardware devices or features. In this paper we present a Big Data Error Correction (BigDEC) tool that overcomes all those limitations by: 1) implementing three different error correction algorithms based on the widely extended k-mer spectrum method; 2) providing scalable performance for large datasets by efficiently exploiting the capabilities of Big Data technologies on multicore clusters based on commodity hardware; 3) supporting two different Big Data processing frameworks (Spark and Flink) to provide greater flexibility to end users; 4) including an efficient, stream-based merge operation to ease downstream processing of the corrected datasets; and 5) significantly outperforming existing parallel tools, being up to 79% faster on a 16-node multicore cluster when using the same underlying correction algorithm. BigDEC is publicly available to download at https://github.com/UDC-GAC/BigDEC. Roberto R. Expósito, Jorge González-Domínguez |
Future Gener. Comput. Syst. | 1 |
| 2023 | SeQual-Stream: approaching stream processing to quality control of NGS datasetsabstractBACKGROUND: Quality control of DNA sequences is an important data preprocessing step in many genomic analyses. However, all existing parallel tools for this purpose are based on a batch processing model, needing to have the complete genetic dataset before processing can even begin. This limitation clearly hinders quality control performance in those scenarios where the dataset must be downloaded from a remote repository and/or copied to a distributed file system for its parallel processing. RESULTS: In this paper we present SeQual-Stream, a streaming tool that allows performing multiple quality control operations on genomic datasets in a fast, distributed and scalable way. To do so, our approach relies on the Apache Spark framework and the Hadoop Distributed File System (HDFS) to fully exploit the stream paradigm and accelerate the preprocessing of large datasets as they are being downloaded and/or copied to HDFS. The experimental results have shown significant improvements in the execution times of SeQual-Stream when compared to a batch processing tool with similar quality control features, providing a maximum speedup of 2.7[Formula: see text] when processing a dataset with more than 250 million DNA sequences, while also demonstrating good scalability features. CONCLUSION: Our solution provides a more scalable and higher performance way to carry out quality control of large genomic datasets by taking advantage of stream processing features. The tool is distributed as free open-source software released under the GNU AGPLv3 license and is publicly available to download at https://github.com/UDC-GAC/SeQual-Stream . Óscar Castellanos-Rodríguez, Roberto R. Expósito, Juan Touriño |
BMC Bioinform. | 2 |
| 2022 | SparkEC: speeding up alignment-based DNA error correction toolsabstractBACKGROUND: In recent years, huge improvements have been made in the context of sequencing genomic data under what is called Next Generation Sequencing (NGS). However, the DNA reads generated by current NGS platforms are not free of errors, which can affect the quality of downstream analysis. Although error correction can be performed as a preprocessing step to overcome this issue, it usually requires long computational times to analyze those large datasets generated nowadays through NGS. Therefore, new software capable of scaling out on a cluster of nodes with high performance is of great importance. RESULTS: In this paper, we present SparkEC, a parallel tool capable of fixing those errors produced during the sequencing process. For this purpose, the algorithms proposed by the CloudEC tool, which is already proved to perform accurate corrections, have been analyzed and optimized to improve their performance by relying on the Apache Spark framework together with the introduction of other enhancements such as the usage of memory-efficient data structures and the avoidance of any input preprocessing. The experimental results have shown significant improvements in the computational times of SparkEC when compared to CloudEC for all the representative datasets and scenarios under evaluation, providing an average and maximum speedups of 4.9[Formula: see text] and 11.9[Formula: see text], respectively, over its counterpart. CONCLUSION: As error correction can take excessive computational time, SparkEC provides a scalable solution for correcting large datasets. Due to its distributed implementation, SparkEC speed can increase with respect to the number of nodes in a cluster. Furthermore, the software is freely available under GPLv3 license and is compatible with different operating systems (Linux, Windows and macOS). Roberto R. Expósito, Marco Martínez-Sánchez, Juan Touriño |
BMC Bioinform. | 1 |
| 2022 | MPI-dot2dot: A parallel tool to find DNA tandem repeats on multicore clustersabstractAbstract Tandem Repeats (TRs) are segments that occur several times in a DNA sequence, and each copy is adjacent to other. In the last few years, TRs have gained significant attention as they are thought to be related with certain human diseases. Therefore, identifying and classifying TRs have become a highly important task in bioinformatics in order to analyze their disorders and relationships with illnesses. Dot2dot, a tool recently developed to find TRs, provides more accurate results than the previous state-of-the-art, but it requires a long execution time even when using multiple threads. This work presents MPI-dot2dot, a novel version of this tool that combines MPI and OpenMP so that it can be executed in a cluster of multicore nodes and thus reduces its execution time. The performance of this new parallel implementation has been tested using different real datasets. Depending on the characteristics of the input genomes, it is able to obtain the same biological results as Dot2dot but more than 100 times faster on a 16-node multicore cluster (384 cores). MPI-dot2dot is publicly available to download from https://sourceforge.net/projects/mpi-dot2dot . Jorge González-Domínguez, José M. Martín-Martínez, Roberto R. Expósito |
J. Supercomput. | 3 |
| 2020 | Power Budgeting of Big Data Applications in Container-based ClustersabstractEnergy consumption is currently highly regarded on computing systems for many reasons, such as improving the environmental impact and reducing operational costs considering the rising price of energy. Previous works have analysed how to improve energy efficiency from the entire infrastructure down to individual computing instances (e.g., virtual machines). However, the research is more scarce when it comes to controlling energy consumption, specially in real time and at the software level. This paper presents a platform that manages a power budget to cap the energy consumed from users to applications and down to individual instances. Using containers as virtualization technology, the energy limitation is implemented thanks to the platform's ability to monitor container energy consumption and dynamically adjust its CPU resources via vertical scaling as required. Representative Big Data applications have been deployed on the platform to prove the feasibility of this approach for energy control, showing that it is possible to distribute and enforce a power budget among users and applications. Jonatan Enes, Guillaume Fieni, Roberto R. Expósito, Romain Rouvoy, Juan Touriño |
CLUSTER | 3 |
| 2020 | Real-time resource scaling platform for Big Data workloads on serverless environments
Jonatan Enes, Roberto R. Expósito, Juan Touriño |
Future Gener. Comput. Syst. | 2 |
| 2020 | SMusket: Spark-based DNA error correction on distributed-memory systems
Roberto R. Expósito, Jorge González-Domínguez, Juan Touriño |
Future Gener. Comput. Syst. | 1 |
| 2020 | CUDA-JMI: Acceleration of feature selection on heterogeneous systems
Jorge González-Domínguez, Roberto R. Expósito, Verónica Bolón-Canedo |
Future Gener. Comput. Syst. | 2 |
| 2019 | Accelerating binary biclustering on platforms with CUDA-enabled GPUs
Jorge González-Domínguez, Roberto R. Expósito |
Inf. Sci. | 2 |
| 2018 | BDWatchdog: Real-time monitoring and profiling of Big Data applications and frameworks
Jonatan Enes, Roberto R. Expósito, Juan Touriño |
Future Gener. Comput. Syst. | 2 |
| 2018 | BDEv 3.0: Energy efficiency and microarchitectural characterization of Big Data processing frameworks
Jorge Veiga, Jonatan Enes, Roberto R. Expósito, Juan Touriño |
Future Gener. Comput. Syst. | 3 |
| 2018 | Big Data-Oriented PaaS Architecture with Disk-as-a-Resource Capability and Container-Based Virtualization
Jonatan Enes, Javier López Cacheiro, Roberto R. Expósito, Juan Touriño |
J. Grid Comput. | 3 |
| 2018 | Enhancing in-memory efficiency for MapReduce-based data processing
Jorge Veiga, Roberto R. Expósito, Guillermo L. Taboada, Juan Touriño |
J. Parallel Distributed Comput. | 2 |
| 2017 | MarDRe: efficient MapReduce-based removal of duplicate DNA reads in the cloudabstractSUMMARY: This article presents MarDRe, a de novo cloud-ready duplicate and near-duplicate removal tool that can process single- and paired-end reads from FASTQ/FASTA datasets. MarDRe takes advantage of the widely adopted MapReduce programming model to fully exploit Big Data technologies on cloud-based infrastructures. Written in Java to maximize cross-platform compatibility, MarDRe is built upon the open-source Apache Hadoop project, the most popular distributed computing framework for scalable Big Data processing. On a 16-node cluster deployed on the Amazon EC2 cloud platform, MarDRe is up to 8.52 times faster than a representative state-of-the-art tool. AVAILABILITY AND IMPLEMENTATION: Source code in Java and Hadoop as well as a user's guide are freely available under the GNU GPLv3 license at http://mardre.des.udc.es . CONTACT: [email protected]. Roberto R. Expósito, Jorge Veiga, Jorge González-Domínguez, Juan Touriño |
Bioinform. | 1 |
| 2016 | Performance evaluation of big data frameworks for large-scale data analyticsabstractThe increasing adoption of Big Data analytics has led to a high demand for efficient technologies in order to manage and process large datasets. Popular MapReduce frameworks such as Hadoop are being replaced by emerging ones like Spark or Flink, which improve both the programming APIs and performance. However, few works have focused on comparing these frameworks. This paper addresses this issue by performing a comparative evaluation of Hadoop, Spark and Flink using representative Big Data workloads and considering factors like performance and scalability. Moreover, the behavior of these frameworks has been characterized by modifying some of the main parameters of the workloads such as HDFS block size, input data size, interconnect network or thread configuration. The analysis of the results has shown that replacing Hadoop with Spark or Flink can lead to a reduction in execution times by 77% and 70% on average, respectively, for non-sort benchmarks. Jorge Veiga, Roberto R. Expósito, Xoán C. Pardo, Guillermo L. Taboada, Juan Touriño |
IEEE BigData | 2 |
| 2016 | Performance Evaluation of Data-Intensive Computing Applications on a Public IaaS CloudabstractThe advent of cloud computing technologies, which dynamically provide on-demand access to computational resources over the Internet, is offering new possibilities to many scientists and researchers. Nowadays, Infrastructure as a Service (IaaS) cloud providers can offset the increasing processing requirements of data-intensive computing applications, becoming an emerging alternative to traditional servers and clusters. In this paper, a comprehensive study of the leading public IaaS cloud platform, Amazon EC2, has been conducted in order to assess its suitability for data-intensive computing. One of the key contributions of this work is the analysis of the storage-optimized family of EC2 instances. Furthermore, this study presents a detailed analysis of both performance and cost metrics. More specifically, multiple experiments have been carried out to analyze the full I/O software stack, ranging from the low-level storage devices and cluster file systems up to real-world applications using representative data-intensive parallel codes and MapReduce-based workloads. The analysis of the experimental results has shown that data-intensive applications can benefit from tailored EC2-based virtual clusters, enabling users to obtain the highest performance and cost-effectiveness in the cloud. Roberto R. Expósito, Guillermo L. Taboada, Sabela Ramos, Juan Touriño, Ramón Doallo |
Comput. J. | 1 |
| 2016 | Flame-MR: An event-driven architecture for MapReduce applications
Jorge Veiga, Roberto R. Expósito, Guillermo L. Taboada, Juan Touriño |
Future Gener. Comput. Syst. | 2 |
| 2015 | Low-latency Java communication devices on RDMA-enabled networksabstractSummary Providing high‐performance inter‐node communication is a key capability for running high performance computing applications efficiently on parallel architectures. In fact, current systems deployments are aggregating a significant number of cores interconnected via advanced networking hardware with Remote Direct Memory Access (RDMA) mechanisms, that enable zero‐copy and kernel‐bypass features. The use of Java for parallel programming is becoming more promising thanks to some useful characteristics of this language, particularly its built‐in multithreading support, portability, easy‐to‐learn properties, and high productivity, along with the continuous increase in the performance of the Java virtual machine. However, current parallel Java applications generally suffer from inefficient communication middleware, mainly based on protocols with high communication overhead that do not take full advantage of RDMA‐enabled networks. This paper presents efficient low‐level Java communication devices that overcome these constraints by fully exploiting the underlying RDMA hardware, providing low‐latency and high‐bandwidth communications for parallel Java applications. The performance evaluation conducted on representative RDMA networks and parallel systems has shown significant point‐to‐point performance increases compared with previous Java communication middleware, allowing to obtain up to 40% improvement in application‐level performance on 4096 cores of a Cray XE6 supercomputer. Copyright © 2015 John Wiley & Sons, Ltd. Roberto R. Expósito, Guillermo L. Taboada, Sabela Ramos, Juan Touriño, Ramón Doallo |
Concurr. Comput. Pract. Exp. | 1 |
| 2015 | Nonblocking collectives for scalable Java communicationsabstractSummary This paper presents a Java implementation of the recently published MPI 3.0 nonblocking message passing collectives in order to analyze and assess the feasibility of taking advantage of these operations in shared memory systems using Java. Nonblocking collectives aim to exploit the overlapping between computation and communication for collective operations to increase scalability of message passing codes, as it has been carried out for nonblocking point‐to‐point primitives. This scalability has become crucial not only for clusters but also for shared memory systems because of the current trend of increasing the number of cores per chip, which is leading to the generalization of multi‐core and many‐core processors. Message passing libraries based on remote direct memory access, thread‐based progression, or implementing pure multi‐threading shared memory support could potentially benefit from the lack of imposed synchronization by nonblocking collectives. But, although the distributed memory scenario has been well studied, the shared memory one has not been tackled yet. Hence, nonblocking collectives support has been included in FastMPJ, a Message Passing in Java (MPJ) implementation, and evaluated on a representative shared memory system, obtaining significant improvements because of overlapping and lack of implicit synchronization, and with barely any overhead imposed over common blocking operations. Copyright © 2014 John Wiley & Sons, Ltd. Sabela Ramos, Guillermo L. Taboada, Roberto R. Expósito, Juan Touriño |
Concurr. Comput. Pract. Exp. | 3 |
| 2013 | Design of Scalable Java Communication Middleware for Multi-Core SystemsabstractThis paper presents smdev, a shared memory communication middleware for multi-core systems. smdev provides a simple and powerful messaging application program interface that is able to exploit the underlying multi-core architecture replacing inter-process and network-based communications by threads and shared memory transfers. The performance evaluation of smdev on several multi-core systems has shown noticeable improvements compared with other Java shared memory solutions, reaching and even overcoming the performance of natively compiled libraries. Thus, smdev has obtained start-up latencies around 0.76 μs and almost 90 Gbps bandwidth for point-to-point communications, as well as high performance and scalability both for collective operations and representative messaging kernels. This fact has motivated the integration of smdev in F-MPJ, our message-passing implementation in Java. Sabela Ramos, Guillermo L. Taboada, Roberto R. Expósito, Juan Touriño, Ramón Doallo |
Comput. J. | 3 |
| 2013 | General-purpose computation on GPUs for high performance cloud computingabstractSUMMARY Cloud computing is offering new approaches for High Performance Computing (HPC) as it provides dynamically scalable resources as a service over the Internet. In addition, General‐Purpose computation on Graphical Processing Units (GPGPU) has gained much attention from scientific computing in multiple domains, thus becoming an important programming model in HPC. Compute Unified Device Architecture (CUDA) has been established as a popular programming model for GPGPUs, removing the need for using the graphics APIs for computing applications. Open Computing Language (OpenCL) is an emerging alternative not only for GPGPU but also for any parallel architecture. GPU clusters, usually programmed with a hybrid parallel paradigm mixing Message Passing Interface (MPI) with CUDA/OpenCL, are currently gaining high popularity. Therefore, cloud providers are deploying clusters with multiple GPUs per node and high‐speed network interconnects in order to make them a feasible option for HPC as a Service (HPCaaS). This paper evaluates GPGPU for high performance cloud computing on a public cloud computing infrastructure, Amazon EC2 Cluster GPU Instances (CGI), equipped with NVIDIA Tesla GPUs and a 10 Gigabit Ethernet network. The analysis of the results, obtained using up to 64 GPUs and 256‐processor cores, has shown that GPGPU is a viable option for high performance cloud computing despite the significant impact that virtualized environments still have on network overhead, which still hampers the adoption of GPGPU communication‐intensive applications. Copyright © 2012 John Wiley & Sons, Ltd. Roberto R. Expósito, Guillermo L. Taboada, Sabela Ramos, Juan Touriño, Ramón Doallo |
Concurr. Comput. Pract. Exp. | 1 |
| 2013 | Performance analysis of HPC applications in the cloud
Roberto R. Expósito, Guillermo L. Taboada, Sabela Ramos, Juan Touriño, Ramón Doallo |
Future Gener. Comput. Syst. | 1 |
| 2013 | Analysis of I/O Performance on an Amazon EC2 Cluster Compute and High I/O Platform
Roberto R. Expósito, Guillermo L. Taboada, Sabela Ramos, Jorge González-Domínguez, Juan Touriño, Ramón Doallo |
J. Grid Comput. | 1 |
| 2013 | Evaluation of messaging middleware for high-performance cloud computing
Roberto R. Expósito, Guillermo L. Taboada, Sabela Ramos, Juan Touriño, Ramón Doallo |
Pers. Ubiquitous Comput. | 1 |
| 2013 | Java in the High Performance Computing arena: Research, practice and experience
Guillermo L. Taboada, Sabela Ramos, Roberto R. Expósito, Juan Touriño, Ramón Doallo |
Sci. Comput. Program. | 3 |
| 2012 | Design of scalable Java message-passing communications over InfiniBand
Roberto R. Expósito, Guillermo L. Taboada, Juan Touriño, Ramón Doallo |
J. Supercomput. | 1 |