Alejandro Calderón 0001

dblp:34/5911-1 · also Alejandro Calderón Mateos · DBLP profile ↗
← Back
30ranked-venue papers
4as first author
8since 2021 · last 2026
0000-0001-6185-653XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 19 · 3 first-author · 4 since 2021Computer networks · 2Artificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Transparent Checkpointing in Parallel Applications Using AD-HOC File Systems
Dario Muñoz-Muñoz, Félix García Carballeira, Alejandro Calderón 0001, Diego Camarmas-Alonso, Jesús H. Carretero
ISPDC3
2026 Hierarchical and distributed data storage for computing continuum
abstract
The Internet of Things (IoT) has transformed how everyone interacts with the environment. Over the past few years, this field has experienced exponential growth , which has led to difficulties in efficiently managing the data generated by these devices and has posed new challenges for cloud infrastructures. As the number of devices involved increases, latency and bandwidth issues become increasingly critical in these systems. To address these issues, architectures such as fog and edge computing emerged that proposed bringing information processing and storage closer to the data generators. This reduced the distance the data had to travel, thereby improving latency and bandwidth and reducing potential bottlenecks. These architectures have evolved into a new concept known as the computing continuum. This approach, based on dynamic collaboration between the cloud, fog, and edge, creates a continuous infrastructure of computational resources that optimizes data processing at each network level as needed. The computing continuum presents important challenges that need to be addressed at multiple levels: the application/algorithmic level (programming paradigms), middleware level (deployment, execution, scheduling, monitoring, data storage, transfer, processing, and analysis), and resource management level. The work introduced in this article addresses the challenges related to the efficient storage and transfer of data between the different levels across the computing continuum. We propose a distributed and parallel file system for this kind of infrastructure that can be used transparently at all levels. This file system can be deployed at different levels (fog, edge, and cloud) hierarchically, allowing the execution of applications at each level and efficient data transfer between these levels and the cloud. It also facilitates the development of IoT applications capable of efficiently transferring data by using typical file system calls. The work presents and validates this data storage system at different levels: two IoT devices (Raspberry Pi 1 & 4), an emulated environment, a controlled simulation, and a performance analysis with Amazon Cloud Services.
Elías Del-Pozo-Puñal, Félix García Carballeira, Diego Camarmas-Alonso, Alejandro Calderón 0001
Future Gener. Comput. Syst.4
2026 CREATOR-Sail: A RISC-V web simulator based on Sail ISA specification
abstract
This article introduces CREATOR-Sail, a RISC-V web simulator that uses Sail as the instruction specification language. This enables binaries written in assembly language to be compiled, executed and debugged while simulating the complete RISC-V architecture and instruction set. The simulator can now simulate both 32-bit and 64-bit RISC-V variants and supports the simulation of the complete instruction set specified in the official standard, including vector and privileged instructions. Both variants include a cache memory module to increase the capacity for architecture simulation, mimicking a real processor. The aim of this work is to provide an industrial-level solution for companies and research groups, offering them a tool with which to customize the environment for their own purposes. The simulator includes an assembly compiler and a debugging module, making it more user-friendly and intuitive than other simulators. All simulator components have been integrated into the web environment using WebAssembly, which enables the execution of native code in a web environment and provides an execution performance similar to that of native applications compared to applications implemented only in JavaScript. The source code of web simulator is available at https://github.com/creatorsim/creator . The web simulator is available at https://creatorsim.github.io/creator/ .
Juan Carlos Cano-Resa, Félix García Carballeira, Diego Camarmas-Alonso, Alejandro Calderón 0001
J. Syst. Archit.4
2025 Improving I/O performance in HPC environments using the Expand Ad-Hoc file system
abstract
This work introduces an exhaustive evaluation, including both benchmarks and real-world data-intensive applications, performed in MareNostrum 4 and HPC4AI Laboratory supercomputers using the Expand Ad-Hoc file system. Expand Ad-Hoc is an ad-hoc file system that dynamically virtualizes the local storage available on compute nodes (i.e., SSD, SHM, etc.) into a fast storage volume to reduce congestion on parallel file systems used as backends in High-Performance Computing (HPC) environments. The main contributions of this work include the design of a new ad-hoc parallel file system, called Expand Ad-Hoc , compatible with POSIX and MPI-IO, and an exhaustive and comprehensive evaluation of Expand. The evaluation compares the performance obtained using IOR and DLIO benchmarks and Nek5000 and Remote Sensing real-world applications on Expand Ad-Hoc , GekkoFS, GPFS, and BeeGFS, showing satisfactory results proving that Expand Ad-Hoc can be used in HPC environments to improve the I/O performance of data-intensive applications transparently.
Diego Camarmas-Alonso, Félix García Carballeira, Alejandro Calderón 0001, Jesús Carretero 0001
J. Supercomput.3
2024 Fault Tolerant in the Expand Ad-Hoc Parallel File System
abstract
Abstract In the last years, applications related to Artificial Intelligence and big data, among others, have been involved. There is a need to improve I/O operations to avoid bottlenecks in accessing a larger amount of data. For this purpose, the Expand Ad-Hoc parallel file system is being designed and developed. Since these applications have very long execution times, fault tolerance mechanisms in the file system are necessary to allow them to continue running in the presence of failures. This work introduces a fault-tolerant design based on data replication for the Expand Ad-Hoc parallel file system and an initial evaluation conducted on the HPC4AI Laboratory supercomputer in Torino. The evaluation of Expand Ad-Hoc with fault-tolerant found that, despite data replication, its performance and scalability are generally better than those of other parallel file systems without fault-tolerant.
Dario Muñoz-Muñoz, Félix García Carballeira, Diego Camarmas-Alonso, Alejandro Calderón 0001, Jesús Carretero 0001
Euro-Par (2)4
2023 A new Ad-Hoc parallel file system for HPC environments based on the Expand parallel file system
abstract
An Ad-Hoc File System dynamically virtualizes storage on compute nodes into a fast storage volume to reduce congestion on parallel file systems used as backends in HPC environments and improve data locality. This paper presents Expand Ad-Hoc, a version of the Expand parallel file system, for use as an Ad-Hoc storage system for HPC environments. Such an update seeks to take better advantage of new storage technologies (on SSDs local to nodes, for example) and to adapt to new demands of current parallel applications (e.g., optimizing data access by analyzing data locality). The paper describes the new system’s features and presents a first evaluation comparing the performance of Expand with another Ad-Hoc File System (GekkoFS), and GPFS. The first results are quite satisfactory and demonstrate the good performance of Expand as an Ad-Hoc storage system.
Félix García Carballeira, Diego Camarmas-Alonso, Alejandro Calderón 0001, Jesús H. Carretero
ISPDC3
2021 A new generic simulator for the teaching of assembly programming
abstract
This article introduces CREATOR, a new generic simulator for assembly programming, developed by the ARCOS group at the UC3M. CREATOR is a new, highly intuitive, and portable simulator that runs from a web browser (no installation needed). This simulator comes with the MIPS32 and RISC-V (32IMF) instruction set. Nevertheless, CREATOR allows, from the simulator itself, to edit and define other instruction sets (instructions, format, registers, etc.). Even more, CREATOR allows the definition of the parameter passing convention to be used in the instruction set. Once each particular instruction set (MIPS32, ARM, RISCV, etc.) has been defined, students can use CREATOR to edit, compile, execute and debug programs written in the associated assembler. The simulator also allows checking that the developed programs comply with the parameter passing convention defined for the instruction set. CREATOR lets us create subroutine libraries that can be loaded and linked to other assembly programs developed in the simulator. All CREATOR features allows teacher to design and deploy practical laboratories more adapted to the desired teaching goals. That improves the teaching experience of the assembly language frequently used in different subjects such as Computer Architecture or Computer Structure. The experience of its use has been very positive in the past courses for students and teachers in both the Universidad Carlos III de Madrid (UC3M) and the Universidad Castilla la Mancha (UCLM).
Diego Camarmas-Alonso, Félix García Carballeira, Elías Del-Pozo-Puñal, Alejandro Calderón 0001
CLEI4
2021 Analyzing the distributed training of deep-learning models via data locality
abstract
In the last few years, deep-learning models are becoming crucial for numerous scientific and industrial applications. Due to the growth and complexity of deep neural networks, researchers have been investigating techniques to train those networks more efficiently. Many efforts have been made to optimize deep-learning models by parallelizing or distributing their training computation across multiple devices. Current state-of-the-art techniques, such as Horovod, have shown to maximize the performance of both the training computation and the inter-node communication of models for different deep-learning frameworks. However, some applications cannot take advantage of the above techniques due to an I/O bottleneck caused by the input data, thus limiting the scalability of the trainings. In this paper, we study an approach based on data locality - that has not been fully studied yet - for those neural networks that cannot benefit from scaling their computation due to a significant bottleneck in the data I/O.
Saul Alonso Monsalve, Alejandro Calderón 0001, Félix García Carballeira, José Rivadeneira
PDP2
2018 A heterogeneous mobile cloud computing model for hybrid clouds
Saul Alonso Monsalve, Félix García Carballeira, Alejandro Calderón 0001
Future Gener. Comput. Syst.3
2018 Assessing and discovering parallelism in C++ code for heterogeneous platforms
David del Rio Astorga, Rafael Sotomayor, Luis Miguel Sánchez, Francisco Javier García Blas, Alejandro Calderón 0001, Javier Fernández 0001
J. Supercomput.5
2017 A new volunteer computing model for data-intensive applications
abstract
Summary Volunteer computing is a type of distributed computing in which ordinary people donate computing resources to scientific projects. BOINC is the main middleware system for this type of distributed computing. The aim of volunteer computing is that organizations be able to attain large computing power thanks to the participation of volunteer clients instead of a high investment in infrastructure. There are projects, like the ATLAS@Home project, in which the number of running jobs has reached a plateau, due to a high load on data servers caused by file transfer. This is why we have designed an alternative, using the same BOINC infrastructure, in order to improve the performance of BOINC projects that have reached their limit due to the I/O bottleneck in data servers. This alternative involves having a percentage of the volunteer clients running as data servers, called data volunteers, that improve the performance of the system by reducing the load on data servers. In addition, our solution takes advantage of data locality, leveraging the low network latencies of closer machines. This paper describes our alternative in detail and shows the performance of the solution, applied to 3 different BOINC projects, using a simulator of our own, ComBoS.
Saul Alonso Monsalve, Félix García Carballeira, Alejandro Calderón 0001
Concurr. Comput. Pract. Exp.3
2016 Improving the Performance of Volunteer Computing with Data Volunteers: A Case Study with the ATLAS@home Project
Saul Alonso Monsalve, Félix García Carballeira, Alejandro Calderón 0001
ICA3PP3
2013 Improving MPI applications with a new MPI_Info and the use of the memoization
abstract
The MPI forum is actively working for a better MPI standard. The results are the new version 3 of the MPI standard, and the efforts for the incoming MPI 3.1/4.0. The technological changes provide many opportunities for improvements and new ideas. This paper introduces two main contributions in this direction: (1) how to improve the MPI_Info object implementation, and (2) a new way of using the former improved MPI_Info object as a storage solution.
Alejandro Calderón 0001, Jesús Carretero 0001, Félix García Carballeira, Javier Fernández 0001, Daniel Higuero, Borja Bergua
EuroMPI1
2012 A Black Box Model for Storage Devices Based on Probability Distributions
abstract
Traditional approaches for storage devices simulation have been based on detailed analytical models. However, detailed models require detailed computations which may be not affordable for large scale simulations. Moreover, highly detailed models cannot be easily generalized. A different approach is the black-box statistical modeling, where the storage device, its interface, and the interconnection mechanisms are modeled as a single stochastic process, defining the request response time as a random variable with an unknown distribution. A random variate generator can be built and integrated into a simulation model. This approach allows to generate a simulation model for both real and synthetic workloads. This article describes a method suitable for building fast simulation models for storage devices. Our method uses as starting point a workload and produces a random variate generator which can be easily integrated into large scale simulation models. A comparison between our variate generator and the widely known simulation tool DiskSim, shows that our variate generator is faster, and can be as accurate as DiskSim.
Laura Prada, Alejandro Calderón 0001, Francisco Javier García Blas, José Daniel García, Jesús Carretero 0001
ISPA2
2012 Expanding the volunteer computing scenario: A novel approach to use parallel applications on volunteer computing
Alejandro Calderón 0001, Félix García Carballeira, Borja Bergua, Luis Miguel Sánchez, Jesús Carretero 0001
Future Gener. Comput. Syst.1
2012 Dynamic-CoMPI: dynamic optimization techniques for MPI parallel applications
Rosa Filgueira, Jesús Carretero 0001, David E. Singh, Alejandro Calderón 0001, Alberto Nuñez
J. Supercomput.4
2010 Branch replication scheme: A new model for data replication in large scale data grids
José María Pérez, Félix García Carballeira, Jesús Carretero 0001, Alejandro Calderón 0001, Javier Fernández 0001
Future Gener. Comput. Syst.4
2010 SENFIS: a Sensor Node File System for increasing the scalability and reliability of Wireless Sensor Networks applications
Soledad Escolar, Florin Isaila, Alejandro Calderón 0001, Luis Miguel Sánchez, David E. Singh
J. Supercomput.3
2009 Fault tolerant file models for parallel file systems: introducing distribution patterns for every file
Alejandro Calderón 0001, Félix García Carballeira, Luis Miguel Sánchez, José Daniel García, Javier Fernández 0001
J. Supercomput.1
2008 Comparing Grid Data Transfer Technologies in the Expand Parallel File System
abstract
Data management is one of the most important problems in grid environments. One important challenge facing grid computing is the design of a grid file system. The Global Grid Forum defines a grid file system as a human-readable resource namespace for management of heterogeneous distributed data resources, that can span across multiple autonomous administrative domains. This paper evaluates Expand, a new grid file system according to the Global Grid Forum recommendations that integrates heterogeneous data storage resources in grids using standard grid technologies: GridFTP and the OGSA ByteIO interface defined by the Open Grid Forum.
Borja Bergua, Félix García Carballeira, Alejandro Calderón 0001, Luis Miguel Sánchez, Jesús Carretero 0001
PDP3
2007 Multiple-Phase Collective I/O Technique for Improving Data Access Locality
abstract
This paper presents multiple-phase collective I/O, a novel collective I/O technique for distributed memory multiprocessors. Multiple-phase collective I/O is a refinement of two-phase collective I/O technique. The communication phase is structured into several steps, which progressively increase the locality of the data to be written to a file system. Besides the description of multiple-phase collective I/O, our paper addresses two additional objectives. First, the authors target to improve the efficiency of the sulphur transport Eurelian model 2 (STEM-II) application. STEM-II is an air quality model that simulates transport, chemical transformations, emission and deposition processes in a unified framework. Due to the large amount of processed data, I/O becomes a critical factor for the application performance. Multiple-phase collective I/O, considerably enhances the performance of the I/O stage in particular and, consequently, of the whole application in general. Second objective consists of evaluating and comparing the performance of multiple-phase collective I/O with that of other well known parallel I/O techniques
David E. Singh, Florin Isaila, Alejandro Calderón 0001, Félix García Carballeira, Jesús Carretero 0001
PDP3
2007 A global and parallel file system for grids
Félix García Carballeira, Jesús Carretero 0001, Alejandro Calderón 0001, José Daniel García, Luis Miguel Sánchez
Future Gener. Comput. Syst.3
2006 On the Reliability of Web Clusters with Partial Replication of Contents
abstract
Traditionally, distributed Web servers have used two strategies for allocating files on server nodes: full replication and full distribution. While full replication provides a highly reliable solution, it limits storage capacity to the capacity of the smallest node. On the other hand, full distribution provides higher storage capacity at the cost of lower reliability. A hybrid solution is partial replication where every file is allocated to a small number of nodes. The most promising architecture for a partial replication strategy is the Web cluster architecture. However, Web clusters present a big flaw from reliability perspective as they contain a single point of failure. To correct this flaw, in this paper we present a modified architecture: the Web cluster with distributed Web switch. Reliability of Web clusters is evaluated for different replication strategies. System evaluations show that our proposal leads to a highly reliable solution with high scalability.
José Daniel García, Jesús Carretero 0001, Javier Fernández 0001, Félix García Carballeira, David E. Singh, Alejandro Calderón 0001
ARES6
2006 Improving the Performance of Cluster Applications through I/O Proxy Architecture
abstract
Clusters are the most common solutions for high performance computing at the present time. In this kind of systems, an important challenge is the I/O subsystem design. Typically, these environments are not flexible enough and the only way to solve performance bottlenecks is adding new hardware. In this paper, we show how an I/O proxy-based architecture can improve the I/O performance of cluster applications in three ways: adapting to the application requirements, reducing the load on the I/O nodes, and finally, increasing the global performance of the storage system
Luis Miguel Sánchez, Florin Isaila, Alejandro Calderón 0001, David E. Singh, José Daniel García
CLUSTER3
2006 A Quantitative Justification to Partial Replication of Web Contents
José Daniel García, Jesús Carretero 0001, Félix García Carballeira, Javier Fernández 0001, Alejandro Calderón 0001, David E. Singh
ICCSA (4)5
2005 High Performance Java Input/Output for Heterogeneous Distributed Computing
abstract
Currently there is a growing interest in using Java for high performance computing. Java has many advantages for high performance computing: it is based on a high-level and object-oriented programming model with support for multithreading and distributed computing. Furthermore, Java 's virtual machine allows applications to run on multiple heterogeneous platforms. A major problem with the use of Java for high performance computing is the I/O. This problem has been solved traditionally in clusters using parallel file systems and parallel I/O libraries, however there is a lack of parallel file systems for Java applications. In this paper, we present a Java parallel I/O library called jExpand. It provides high performance I/O by using several NFS servers in parallel, as NFS can be found in multiple platforms (Linux, Solaris, Windows 2000, etc), we provide a universal parallel file system that can be used everywhere. jExpand requires no changes in the NFS server as it uses RPC operations to provide parallel access to the same file. The paper describes the design, implementation and evaluation of jExpand.
José María Pérez, Luis Miguel Sánchez, Félix García Carballeira, Alejandro Calderón 0001, Jesús Carretero 0001
ISCC4
2004 An Adaptive Cache Coherence Protocol Specification for Parallel Input/Output Systems
abstract
Caching has been intensively used in memory and traditional file systems to improve system performance. However, the use of caching in parallel file systems and I/O libraries has been limited to I/O nodes to avoid cache coherence problems. We specify an adaptive cache coherence protocol that is very suitable for parallel file systems and parallel I/O libraries. This model exploits the use of caching, both at processing and I/O nodes, providing performance improvement mechanisms such as aggressive prefetching and delayed-write techniques. The cache coherence problem is solved by using a dynamic scheme of cache coherence protocols with different sizes and shapes of granularity. The proposed model is very appropriate for parallel I/O interfaces, such as MPI-IO. Performance results, obtained on an IBM SP2, are presented to demonstrate the advantages offered by the cache management methods proposed.
Félix García Carballeira, Jesús Carretero 0001, Alejandro Calderón 0001, José María Pérez, José Daniel García
IEEE Trans. Parallel Distributed Syst.3
2003 Data Allocation and Load Balancing for Heterogeneous Cluster Storage Systems
abstract
Distributed filesystems are a typical solution in networked environments as clusters and grids. Parallel filesystems are a typical solution in order to reach high performance I/O distributed environment, but those filesystems have some limitations in heterogeneous storage systems. Usually in distributed systems, load balancing is used as a solution to improve the performance, but typically the distribution is made between peer-to-peer computational resources and from the processor point of view. In heterogeneous systems, like heterogeneous clusters of workstations, the existing solutions do not work so well. However, the utilization of those systems is more extended every day, having an extreme example in the grid environment. In this paper we bring attention to those aspects of heterogeneous distributed data systems presenting a parallel file system that take into account heterogeneity of storage nodes, the dynamic addition of new storage nodes, and an algorithm to group requests in heterogeneous systems.
José María Pérez, Félix García Carballeira, Jesús Carretero 0001, Alejandro Calderón 0001, Luis Miguel Sánchez
CCGRID4
2003 Video Forwarding Techniques for Mixed Wired and Wireless Networks
abstract
During the last years, Internet video streaming has experiences a phenomenal growth. This is happening despite the notorious difficulties of transmitting data packets with a deadline over the Internet, due to variability in throughput, delays and losses. These problems arise significantly when using wireless networks where the available bandwidth is low and the losses are important due to its error prone transmission nature. In this paper we propose a fast-forwarding technique that is based on segmenting the movie on different files. Normal movie reproduction requires all the files, but fast-forwarding reproduction only requires one file. Those files can me merged by the client or by the server. The segmentation is frame based, grouping all the frames that can be independently decoded together. The resulting file can be showed with any existing player. This group of frames would be the ones to use in a fast-forward reproduction. Our techniques can also be useful in adaptive environments, like wireless networks, because there is no problem for the fast-forward file to use the same optimizations that exist for full movie files. This method also reduces the storage bandwidth and the storage size needed (there is no extra data for fast-forwarding). We also propose a video server architecture that takes advantage of this technique to achieve full interactive video reproduction. The evaluation results shown in this paper demonstrates that our technique enhances video fast-forwarding operations.
Javier Fernández 0001, Jesús Carretero 0001, Félix García Carballeira, José María Pérez, Alejandro Calderón 0001, José J. Muñoz
ISCC5
2001 New Techniques for Collective Communications in Clusters: A Case Study with MPI
abstract
The paper describes new techniques to increase the performance of collective communication operations in clusters. These techniqnes are based in multithreading operations and on-line data compression. The techniques proposed have been implemented in MiMPI, a thread-safe implementation of MPI. We have evaluated, and compared, the performance of MiMPI with other implementations of MPI available for clusters with Linux and Windows 2000. The benchmark used has been MPBench, a flexible and portable framework to allow benchmarking of MPI implementations.
Alejandro Calderón 0001, Félix García Carballeira, Jesús Carretero 0001, Javier Fernández 0001, Oscar Pérez
ICPP1