EDBT 2026 Demo / reviewers in the wild / expert
Francisco Javier García Blas
dblp:86/246 · also Javier García 0005, Javier García Blas, Javier García-Blas
· DBLP profile ↗
55ranked-venue papers
10as first author
15since 2021 · last 2026
0000-0003-1452-1918ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 41 · 7 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Streaming I/O for scientific workflow engine accelerationabstractScientific workflows are increasingly characterized by complex task dependencies and large-scale data exchanges, which place significant pressure on the input/output (I/O) systems of traditional Workflow Engines (WFEs). These challenges are particularly evident in data-intensive and real-time processing contexts, where conventional disk-based I/O mechanisms often become performance bottlenecks. This paper presents an approach to enhancing the DAGonStar scientific workflow engine by integrating CAPIO, a middleware designed to support memory-based streaming I/O. The integration combines DAGonStar’s orchestration capabilities with CAPIO’s efficient data handling to better support workflows operating on continuous or large-scale datasets. We describe the architectural modifications introduced to enable this collaboration and provide an analysis of the resulting system. The proposed solution aims to improve the responsiveness and flexibility of scientific workflows by streamlining data transfers and simplifying task coordination. This work contributes to the evolution of workflow systems toward more efficient and scalable models for scientific computing. • Integration of memory-based streaming I/O into scientific workflow engines. • Automated generation of synchronization rules through workflow dependency analysis of DAGonStar. • Enhanced pipeline’s tasks execution efficiency via system call interception through the usage of CAPIO. • Benchmark evaluation showing up to 33% reduction in execution time with DAGonCAPIO. • Support for both local batch and SLURM-based distributed executions. Simone Perrotta, Ciro Giuseppe De Vita, Gennaro Mellone, Marco Edoardo Santimaria, Massimo Torquati, Francisco Javier García Blas, Raffaele Montella |
Future Gener. Comput. Syst. | 6 |
| 2026 | HERCULES: A scalable and elastic ad-hoc file system for large-scale computing systemsabstractThe increasing demand for data processing by new, data-intensive applications is placing significant strain on the performance and capacity of HPC storage systems. Advancements in storage technologies, such as NVMe and persistent memory, have been introduced to address these demands. However, relying exclusively on ultra-fast storage devices is not cost-effective, necessitating multi-tier storage hierarchies to manage data based on its usage. In response, ad-hoc file systems have been proposed as a solution. These systems use the storage resources available in compute nodes, including memory and persistent storage, to create temporary file systems that adapt to application behavior in the HPC environment. This work presents the design, implementation, and evaluation of HERCULES, a distributed ad-hoc in-memory storage system, with a focus on its new metadata and elasticity model. HERCULES takes advantage of the Unified Communication X (UCX) framework, leveraging RDMA protocols such as Infiniband, Omnipath, shared-memory, and zero-copy transfers for data transfer. It includes elasticity features at runtime and fault-tolerant facilities. The elasticity features, together with flexible policies for data allocation, allow HERCULES to migrate data so that the available resources can be efficiently used. Our exhaustive evaluation results demonstrate a better performance than Lustre and BeeGFS, two parallel file systems heavily used in High-Performance Computing systems, and GekkoFS, an ad-hoc state-of-the-art solution. Genaro Sanchez-Gallegos, Cosmin Petre, Francisco Javier García Blas, Jesús Carretero 0001 |
Future Gener. Comput. Syst. | 3 |
| 2026 | A coupled Lagrangian-AI hierarchical and heterogeneous model for predicting bacteria contamination in farmed mussels
Ciro Giuseppe De Vita, Gennaro Mellone, Diana Di Luccio, Francisco Javier García Blas, Francesca Barchiesi, Raffaele Montella |
Future Gener. Comput. Syst. | 4 |
| 2025 | Automatic Calibration for CT Bed Stitching Based on Fourier-MellinabstractIn Cone Beam CT (CBCT) systems, each acquisition is carried out with a stationary bed. Therefore, the field of view is limited by the size of the detector. To overcome this limitation, it is common to perform successive acquisitions for different bed positions, which are subsequently combined to increase the field of view in the longitudinal direction. To avoid the appearance of double edges in the resulting volume, it is necessary to calculate the exact bed displacement, which has traditionally been obtained by prior geometric calibrations with calibration phantoms. This implies that the calibration must be repeated periodically to adapt to any changes that the equipment may undergo. As an alternative, in this work we propose the use of an automatic calibration algorithm capable of obtaining the misalignment parameters in real time, allowing to obtain multibed images in small and regular animal CBCT systems without the need of a previous calibration. Daniel Sanderson, Daniel Alejandro Rodriguez, Francisco Javier García Blas, Manuel Desco, Mónica Abella |
CBMS | 3 |
| 2025 | Detection of Active and Latent Tuberculosis with Explainable Deep Learning EnsemblesabstractTuberculosis is one of the deadliest diseases in the world, despite being treatable, the number of new cases increases yearly. The situation is even more concerning due to antimicrobial resistance (AMR). To avoid undetected cases and breaking transmission, the research community is using artificial intelligence (AI) to develop new and fast diagnosis tools. Nevertheless, the creation of good quality and well-balanced datasets for training AI tools is challenging. It is especially difficult to collect data on asymptomatic patients who do not go to the professional as they do not feel the symptoms, such as the one who suffered latent tuberculosis. In this work, we propose an explainable deep learning ensemble of convolutional neural networks (CNNs) to classify tuberculosis chest X-ray (CXRs) images. The ensemble was designed to alleviate being influenced by unbalanced data due to the small number of latent patients included in the dataset. The system was trained to detect tuberculosis against healthy patients and other diseases with CXRs as the only input, so as to inform tuberculosis patients of the stage of the disease (active or latent). The model was designed with two parallel CNNs in an ensemble that used a random forest (RF) to overcome the imbalance. The system reported a performance of 96 % accuracy, showing that the ensemble can improve the performance of the underrepresented class. Further, the RF completed the interpretability of the system provided by grad-CAM heatmaps, supporting the evolution and integration of the system in clinical environments. Lara Visuña, Francisco Javier García Blas, Jesús Carretero 0001 |
CBMS | 2 |
| 2025 | Automated monitoring of Mycobacterium population growth in time-lapse microscopy with Deep Learning pseudo-supervised algorithmabstractThe understanding of bacteria evolution through time and the analysis of their interaction with the medium is a main step in biomedical research. Bacteria cause in humans different infectious diseases, such as tuberculosis. Tuberculosis is one of the leading causes of death worldwide, caused by Mycobacterium tuberculosis ( Mtb ). Observation of Mycobacteria under a microscope provides essential information for disease-fighting. Nevertheless, Mtb presents a very particular cell division pattern characterized by a high entanglement in which the frontier between dividing cells are almost indistinguishable to the human eye, which makes the manual analysis of cultures a very challenging task, resulting in extremely long processing times. In this paper, we propose a novel approach to automatically detect bacterial cells and infer the growth rate and growth speed over a time-lapse microscopy (TLM) sequence. This method leverages a semi-supervised Deep Learning approach with a lightweight U-Net for efficient and fast processing. To reduce efforts of database construction, we deploy and exploit a mix of manually and synthetically annotated databases of two different classes of Mycobacterium , i.e. Mtb and Mycobacterium smegmatis ( Msm ). The experiments demonstrated the efficiency of the detection of both bacterial reproduction rate and bacterial death in Mycobacterium species, with the Phase-contrast (PhC) microscopy channels as a unique input. Reaching a segmentation accuracy of 96.3%, recall of 78.2%, and precision of 85.5% on Mtb detection for the model trained with a mix of manually and synthetically annotated databases. The experiments validate as well the synthetic label generation technique, which provides more conscience and accurate results than the manually segmented techniques. The mask generation workflow opens new opportunities for the microbial AI development field. The system implemented may change the paradigm in preclinical anti-mycobacteria drug development for since enables the adoption of TLM, a technology that provides extremely valuable information on drug efficacy but is seldom leveraged due to the complex image processing it involves. • A semi-supervised Deep Learning U-Net based model for TLM bacterial culture analysis. • The method reports bacterial growth and death with PhC as a single input. • Evaluation in two different Mycobacterium ( M. tuberculosis and M. smegmatis ). • Creating a methodology for synthetic annotation of the bacterial population in TLM. • Comparative analysis of manual and synthetic databases. Lara Visuña, Francisco Javier García Blas, Santiago Ferrer-Bazaga, Patricio Lopez-Exposito, Jesús Carretero 0001 |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | A comparative study of ad-hoc file systems for extreme scale computing
Njoud O. Almaaitah, Francisco Javier García Blas, Genaro Sanchez-Gallegos, Jesús Carretero 0001, Marc-Andre Vef, André Brinkmann |
Future Gener. Comput. Syst. | 2 |
| 2025 | Guest Editorial: New Tools and Techniques for the Distributed Computing Continuum
Jesús Carretero 0001, Francisco Javier García Blas, Sameer Shende |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2024 | An Innovative Control Approach for Cyber-Physical Transportation Systems: The Case of Monte-Carlo Workflow ComputationsabstractContemporarily, in light of the intelligent transportation systems (ITS) sector, the tendency can be observed that the solution of the multi-objective cyber-physical optimization problems with imperfect information takes an increasingly weighted role. In the present scientific work, the authors want to take these developments into account by introducing an innovative cyber-physical architectural design and corresponding the two-stage heuristic computing approach. It is utilized in synergy with the MCSA11Multi-tier Cyber-physical System Architecture and DCEx architectural principles for the workflow scheduling of Monte-Carlo simulation, which is based on the intelligent and sustainable route-order dispatching process model. Factors such as emissions, transport costs, risks, and the individual weighting of orders are reflected in the model. In particular, the authors define a stochastic ILP-based22Integer Linear Programming monte-carlo workflow model. They further propose two-stage scheduling heuristic with 6-HEFT DAG relaxation as first stage and apply state-of-the-art techniques as a part of SCIP framework to solve 2nd 1–0 ILP-based stage; evaluate the performance of the scheduling approach. The authors obtain preliminary results of the second stage behavior using a realistic heterogeneous computing scenario and corresponding constraint structures within MACS simulator engine33Modular Architecture for Complex Computing Systems Analysis. The results from the experiments illustrate moderate complexity of the approach. Scalability of the model looks promising for the applicability in various industry-related scenarios and corresponding computing environments. Vladislav Kashansky, Sara Agha Hossein Kashani, Francisco Javier García Blas, Fabrizio Marozzo, Hai Zhuge, Xiaoping Sun |
PDP | 3 |
| 2023 | Adaptive multi-tier intelligent data manager for ExascaleabstractThe main objective of the ADMIRE project1 is the creation of an active I/O stack that dynamically adjusts computation and storage requirements through intelligent global coordination, the elasticity of computation and I/O, and the scheduling of storage resources along all levels of the storage hierarchy, while offering quality-of-service (QoS), energy efficiency, and resilience for accessing extremely large data sets in very heterogeneous computing and storage environments. We have developed a framework prototype that is able to dynamically adjust computation and storage requirements through intelligent global coordination, separated control, and data paths, the malleability of computation and I/O, the scheduling of storage resources along all levels of the storage hierarchy, and scalable monitoring techniques. The leading idea in ADMIRE is to co-design applications with ad-hoc storage systems that can be deployed with the application and adapt their computing and I/O behaviour on runtime, using malleability techniques, to increase the performance of applications and the throughput of the applications. Jesús Carretero 0001, Francisco Javier García Blas, Marco Aldinucci, Jean-Baptiste Besnard, Jean-Thomas Acquaviva, André Brinkmann, Marc-Andre Vef, Emmanuel Jeannot, Alberto Miranda, Ramon Nou, Morris Riedel, Massimo Torquati, Felix Wolf 0001 |
CF | 2 |
| 2023 | Exascale programing models for extreme data processingabstractExtreme Data is an incarnation of Big Data concept distinguished by the massive amounts of data that must be queried, communicated and analysed in (near) real-time by using a very large number of memory/storage elements and Exascale computing systems. Immediate examples are the scientific data produced at a rate of hundreds of gigabits-per-second that must be stored, filtered and analysed, the millions of images per day that must be mined/analyzed in parallel, the one billion of social data posts queried in real-time on an in-memory components database. Traditional disks or commercial storage cannot handle nowadays the extreme scale of such application data. Following the need of improvement of current concepts and technologies, ASPIDE's activities focus on data-intensive applications running on systems composed of up to millions of computing elements (Exascale systems). Practical results will include the methodology and software prototypes that will be designed and used to implement Exascale applications. Francisco Javier García Blas, Jesús Carretero 0001 |
CF | 1 |
| 2023 | Hercules: Scalable and Network Portable In-Memory Ad-Hoc File System for Data-Centric and High-Performance Applications
Francisco Javier García Blas, Genaro Sanchez-Gallegos, Cosmin Petre, Alberto Riccardo Martinelli, Marco Aldinucci, Jesús Carretero 0001 |
Euro-Par | 1 |
| 2023 | A novel approach for large-scale environmental data partitioning on cloud and on-premises storage for compute continuum applicationsabstractSummary Cloud‐based services have proved useful in several research fields, such as engineering, health science, and astrophysics, to mention a few examples. The computational environmental science community developed a strong need for cloud facilities to store, process, and manage data from observations and numerical models for simulations and forecasts. Weather forecast models and global sensor networks deal with multidimensional geo‐referenced data∖sets. However, environmental data consumer applications usually require a relatively small amount of multidimensional input data slice to analyze a specific area or time interval. Hence, reducing data dimension for information retrieval is mandatory. This paper presents a twofold solution: a technique to load and retrieve the sliced multidimensional data set on different cloud services such as Amazon Web Service (AWS), Google Cloud Platform, and Microsoft Azure. The experimental results performed on these cloud services highlight that the proposed method can significantly speed up the process of loading and retrieving the data slices compared to working with the entire data set in bulk or OPeNDAP server. Gennaro Mellone, Ciro Giuseppe De Vita, Dante D. Sánchez-Gallegos, Genaro Sanchez-Gallegos, Catherine Alessandra Torres Charles, Francisco Javier García Blas, Jesús Carretero 0001, José Luis González 0002, Giuliano Laccetti |
Concurr. Comput. Pract. Exp. | 6 |
| 2022 | Multi-bed stitching tool for 3D computed tomography accelerated by GPU devicesabstractIn computed tomography (CT) systems it is common to acquire several consecutive datasets for different bed positions, which are subsequently combined to enlarge the field of view in the longitudinal direction. For this combination, the geometric calibration of the bed motion is necessary to avoid double edges in the overlaped area. This calibration is performed periodically using calibration phantom with markers to guide parameter esti-mation. This work presents a novel correlation-based automatic bed stitching tool for CT that avoids the need of the calibration step. Our approach exploits the massive parallelism offered by GPUs and features an optimized memory model that allows large volumes to be stitched in near-real time. Evaluation in rodent studies demonstrates not only that the offered implementation is able to paste tomographic studies in reduced time, but also that it reduces the memory footprint. Francisco Javier García Blas, Pablo Brox, Manuel Desco, Mónica Abella |
CCGRID | 1 |
| 2022 | Convergence of HPC and Big Data in extreme-scale data analysis through the DCEx programming modelabstractHigh-level programming models can help application developers to access and use resources without the need to manage low-level architectural entities, as a parallel programming model defines a set of programming abstractions that simplify the way by which a programmer structures and expresses her/his algorithm. Early proposals of Exascale programming tools are based on the adaptation of traditional parallel programming languages and hybrid solutions. This incremental approach is too conservative, often resulting in very complex code. This paper describes the design features, the programming constructs, and the runtime mechanisms of the Data Centric programming model for Exascale systems (DCEx). DCEx is based on structuring applications into data-parallel blocks. Blocks are units of shared-and distributed-memory parallel computation, communication, and migration in the memory/storage hierarchy. Blocks and their message queues are mapped onto processes and placed in memory/storage by the DCEx runtime. Those data-parallel blocks are orchestrated by using distributed parallel patterns that simplify the development cost. DCEx aims to reach the convergence of traditional HPC programming models, mainly based on MPI, with the emerging technologies based on the data intensive paradigms. To demonstrate the potential of DCEx, we carried out an experimental evaluation developing a real-world diffusion-weighted magnetic resonance imaging data processing application in a neuroimaging research context. Francisco Javier García Blas, Javier Fernández 0001, Jesús Carretero 0001, Fabrizio Marozzo, Domenico Talia, Paolo Trunfio, Alberto Fernández-Pena, Daniel Martín de Blas |
SBAC-PAD | 1 |
| 2020 | Exposing data locality in HPC-based systems by using the HDFS backendabstractNowadays, there are two main approaches for dealing with data-intensive applications: parallel file systems in classical High-Performance Computing (HPC) centers and Big Data like parallel file system for ensuring the data centric vision. Furthermore, there is a growing overlap between HPC and Big Data applications, given that Big Data paradigm is a growing consumer of HPC resources. HDFS is one of the most important file systems for data intensive applications while, from the parallel file systems point of view, MPI-IO is the most used interface for parallel I/O. In this paper, we propose a novel solution for taking advantage of HDFS through MPI-based parallel applications. To demonstrate its feasibility, we have included our approach in MIMIR, a MapReduce framework for MPI-based applications. We have optimized MIMIR framework by providing data locality features provided by our approach. The experimental evaluation demonstrates that our solution offers around 25% performance for map phase compared with the MIMIR baseline solution. José Rivadeneira, Félix García Carballeira, Jesús Carretero 0001, Francisco Javier García Blas |
HiPC | 4 |
| 2020 | M3AT: Monitoring Agents Assignment Model for Data-Intensive ApplicationsabstractNowadays, massive amounts of data are acquired, transferred, and analyzed nearly in real-time by utilizing a large number of computing and storage elements interconnected through high-speed communication networks. However, one issue that still requires research effort is to enable efficient monitoring of applications and infrastructures of such complex systems. In this paper, we introduce an Integer Linear Programming (ILP) model called M3AT for optimized assignment of monitoring agents and aggregators on large-scale computing systems. We identified a set of requirements from three representative data-intensive applications and exploited them to define the model's input parameters. We evaluated the scalability of M3AT using the Constraint Integer Programing (SCIP) solver with default configuration based on synthetic data sets. Preliminary results show that the model provides optimal assignments for subsystems composed of up to 200 monitoring agents with complex I/O policies, while keeping the number of aggregators constant and demonstrates variable sensitivity with respect to the scale of monitoring data aggregators and limitation policies imposed. Vladislav Kashansky, Dragi Kimovski, Radu Prodan, Prateek Agrawal, Fabrizio Marozzo, Gabriel Iuhasz, Marek Justyna, Francisco Javier García Blas |
PDP | 8 |
| 2020 | Towards enhanced MRI by using a multiple back end programming framework
Francisco Javier García Blas, David del Rio Astorga, Jesús Carretero 0001, José Daniel García |
Future Gener. Comput. Syst. | 1 |
| 2020 | Accelerated iterative image reconstruction for cone-beam computed tomography through Big Data frameworks
Estefania Serrano, Francisco Javier García Blas, Jesús Carretero 0001, Manuel Desco, Mónica Abella |
Future Gener. Comput. Syst. | 2 |
| 2020 | Kulla, a container-centric construction model for building infrastructure-agnostic distributed and parallel applications
Hugo G. Reyes-Anastacio, José Luis González 0002, Víctor Jesús Sosa Sosa, Jesús Carretero 0001, Francisco Javier García Blas |
J. Syst. Softw. | 5 |
| 2019 | Exploiting Stream Parallelism of MRI Reconstruction Using GrPPI over Multiple Back-EndsabstractIn recent years, on-line processing of data streams has been established as a major computing paradigm. This is due mainly to two reasons: first, more and more data are generated in near real-time that need to be processed; the second reason is given by the need of efficient parallel applications. However, the above-mentioned areas expose a tough challenge over traditional data-analysis techniques, which have been forced to evolve to a stream perspective. In this work we present an comparative study of a stream-aware multi-staged application, which has been implemented using GrPPI, a generic and reusable parallel pattern interface for C++ applications. We demonstrate the benefits of using this interface in terms of programability, performance, and scalability. Francisco Javier García Blas, David del Rio Astorga, Javier Daniel Garcia, Jesús Carretero 0001 |
CCGRID | 1 |
| 2019 | Hybrid static-dynamic selection of implementation alternatives in heterogeneous environments
David del Rio Astorga, Manuel F. Dolz, Javier Fernández 0001, Francisco Javier García Blas |
J. Supercomput. | 4 |
| 2018 | GPU-accelerated iterative reconstruction for limited-data tomography in CBCT systemsabstractBACKGROUND: Standard cone-beam computed tomography (CBCT) involves the acquisition of at least 360 projections rotating through 360 degrees. Nevertheless, there are cases in which only a few projections can be taken in a limited angular span, such as during surgery, where rotation of the source-detector pair is limited to less than 180 degrees. Reconstruction of limited data with the conventional method proposed by Feldkamp, Davis and Kress (FDK) results in severe artifacts. Iterative methods may compensate for the lack of data by including additional prior information, although they imply a high computational burden and memory consumption. RESULTS: pixels) using partitioning strategies in forward- and back-projection operations. We evaluated the algorithm on small-animal data for different scenarios with different numbers of projections, angular span, and projection size. Reconstruction time varied linearly with the number of projections and quadratically with projection size but remained almost unchanged with angular span. Forward- and back-projection operations represent 60% of the total computational burden. CONCLUSION: Efficient implementation using parallel processing and large-memory management strategies together with GPU kernels enables the use of advanced reconstruction approaches which are needed in limited-data scenarios. Our GPU implementation showed a significant time reduction (up to 48 ×) compared to a CPU-only implementation, resulting in a total reconstruction time from several hours to few minutes. Claudia de Molina, Estefania Serrano, Francisco Javier García Blas, Jesús Carretero 0001, Manuel Desco, Mónica Abella |
BMC Bioinform. | 3 |
| 2018 | New directions in mobile, hybrid, and heterogeneous clouds for cyberinfrastructures
Jesús Carretero 0001, Francisco Javier García Blas, Gabriel Antoniu, Dana Petcu |
Future Gener. Comput. Syst. | 2 |
| 2018 | Assessing and discovering parallelism in C++ code for heterogeneous platforms
David del Rio Astorga, Rafael Sotomayor, Luis Miguel Sánchez, Francisco Javier García Blas, Alejandro Calderón 0001, Javier Fernández 0001 |
J. Supercomput. | 4 |
| 2017 | Medical Imaging Processing on a Big Data platform using Python: Experiences with Heterogeneous and Homogeneous ArchitecturesabstractThe apparition of new paradigms, programming models, and languages that offer better programmability and better performance turns the implementation of current scientific applications into a less time-consuming task than years ago. One significant example of this trend is the MapReduce programming model and its implementation using Apache Spark. Nowadays, this programming model is mainly used for data analysis and machine learning applications, although it has been expanded to its usage in the HPC community. On the side of programming languages, Python has positioned itself as an alternative to other scientific programming languages, such as Matlab or Julia. In this work we explore the capabilities of Python and Apache Spark as partners in the implementation of the backprojection operator of a CT reconstruction application. We present two interesting approaches with two different types of architectures: a heterogeneous architecture including NVidia GPUs and a full performance CPU mode with the compatibility with C/C++ native source code. We experimentally demonstrate that current CPU-based implementations scale with the number of computational units. Estefania Serrano, Francisco Javier García Blas, Jesús Carretero 0001, Mónica Abella, Manuel Desco |
CCGrid | 2 |
| 2017 | Algorithms and applications towards the convergence of high-end data-intensive and computing systemsabstractWith the increasing availability of data generated by scientific instruments and simulations, today, solving many of our most important scientific and engineering problems requires high-end computing systems (HECS)1 that may be able to process and storage a huge amount of data.2 With this landscape, many synergies between extreme-scale computing, simulations, and data intensive applications might arise.(3, 4) However, the high-performance computing and data analysis platforms, paradigms, and tools have evolved in many cases in different fields, having their own specific methodologies, tools, and techniques. We need to evolve systems and paradigms to create High-End Data-Intensive Computing Systems (HEDICS) to create high-end resources that must be powerful enough in a broad sense (computation, storage, I/O capacity, communications, etc), but at the same time have to provide utilities from the Big Data computing (BDC) space to satisfy the data management and analytics needs of near future applications. Future HECS platforms will be likely characterized by a three to four orders of magnitude, increasing in concurrency, a substantially larger storage capacity, and a deepening of the storage hierarchy. Moreover, the advent of the Big Data challenges5 has generated new initiatives closely related to ultrascale computing systems in large scale distributed systems. The current uncoordinated development model of independently applying optimizations at each layer of the system software I/O software stack will not scale to the required levels of distribution, concurrency, storage hierarchy, and capacity.6 Thus, we need reusable, modular, and scalable frameworks for designing high-end reconfigurable computers, including novel data processing building block and innovative programming models. In those aspects, many new topics are open to research: parallel and distributed algorithms for HEDICS; algorithms for aggressive management of information and knowledge from massive data sources; resource management and scheduling in high-end data and computing systems; tools and environments for parallel/distributed high-end software development; new programming models, as well as machine and application abstractions; resilience issues in HEDICS; adaptive software; architectures, networks, and systems suited for extreme-scale and Big Data; massive distributed and parallel data analytics and feature extraction; new I/O and storage systems valid for HEDICS; and novel and redesigned high-end scientific and engineering computing. This special issue is intended to provide an overview of some key topics and state-of-the-art of recent advances in subjects relevant to High-End Data-Intensive Computing Systems. The general objectives are to address, explore, and exchange information on the challenges and current state-of-the-art in HEDICS, new programming models, run-times, and data facilities design and performance, and their application in various science and engineering domains. This special issue includes research papers addressing the state-of-the-art in high-end data-intensive computing systems. A set of carefully selected works was invited based on the original presentations at the 16th International Conference on Algorithms and Architectures for Parallel Processing (ICA3PP 2016),7 which was held in Granada, Spain, December 2016 and the Third International Workshop of Sustainable Ultrascale Network (NESUS 2016),8 held in Sofia, Bulgaria, October 2016. The extended works have been thoroughly reviewed by an international technical reviewing committee, and only nine papers covering a wide range of relevant challenges in HEDICS were selected for this special issue. The manuscripts present research works showing the convergence of High-End Data and Computing Systems, including new frameworks and platforms, system software enhancements, algorithm design and optimization, programming paradigms and techniques, data processing support in high-end computing systems, and run-time support for HEDICS and performance simulations, measurement, and evaluations. The set of accepted papers can be organized under the following key subjects and subsections and are briefly described in the remaining parts of this section. Current parallel and distributed programming frameworks aid developers to a great extent in implementing applications that exploit homogeneous resources. Nevertheless, it is generally accepted that the ability to develop large-scale distributed applications has lagged seriously behind other developments in cyber-infrastructure.9 Thus, developers strongly require additional expertise to properly port and tune their applications to operate efficiently on specific parallel and distributed platforms, which is not straightforward and demands considerable efforts and specific knowledge. One important cause is the lack of high-level parallel pattern abstractions in the existing frameworks. Dolz et al,10 in their paper A Generic Parallel Pattern Interface for Stream and Data Processing, propose GRPPI, a generic and reusable parallel pattern interface for both stream processing and data-intensive C++ applications available for high-end nodes. GRPPI accommodates a layer between developers and existing parallel programming back-ends targeting multi-core processors, such as C++ threads, OpenMP and Intel TBB, and accelerators back-end like CUDA Thrust. Furthermore, thanks to its high-level C++ API and pattern composability features, GRPPI enables users to easily expose parallelism via stand-alone patterns or pattern compositions matching in sequential applications. The authors evaluate this interface using an image processing use case and demonstrate its benefits from the usability, flexibility, and performance points of views. Furthermore, they analyse the impact of using stream and data pattern compositions on CPUs, GPUs, and heterogeneous configurations. To scale to the next level, as high-end data intensive computing systems become more widespread for scientific applications, there is a necessity of simplifying the development, deployment, and execution of complex data analysis applications for scientific discovery. The scientific workflow model is the leading approach for designing and executing data-intensive applications in high-performance computing infrastructures. Commonly, scientific workflows are built by a set of connected tasks generally arranged in a directed acyclic graph style, which communicate through storage abstractions. Regarding the paper A Data-aware Scheduling Strategy for Workflow Execution in Clouds, Marozzo et al11 present the integration between DMCF and Hercules solutions by using a data-aware scheduling strategy for exploiting data locality in data-intensive workflows. The Data Mining Cloud Framework (DMCF) is a system allowing users to design and execute data analysis workflows on cloud platforms, relying on cloud storage services for every I/O operation, while Hercules is an in-memory I/O solution that can be used in DMCF as an alternative to cloud storage services, providing additional performance and flexibility features. The experimental results demonstrate the performance improvements achieved using the proposed data-aware scheduling strategy in the Microsoft Azure cloud platform. In particular, with the new proposed scheduling strategy, the I/O overhead has been reduced by 55% with respect to the Azure storage, leading to a 20% reduction of the total execution time. In spite of former solutions, network traffic is always a major problem in HEDICS due to data movements. In their paper A scalable synthetic traffic model of Graph500 for computer networks analysis, Fuentes et al12 provide a simulation tool for network architects that need to evaluate the suitability of their interconnect for Big Data applications. Their development is a low computation- and memory-demanding synthetic traffic model that emulates the behaviour of the Graph500 communications and is publicly available in an open-source network simulator. The characterization of network traffic is inferred from a profile of several executions of the benchmark with different input parameters, and the equations in the model have been validated against an execution of benchmarks with a different set of parameters to measure also the impact of the node computation capabilities and network characteristics in the execution time of the model. To cope with huge jobs, some organizations use volunteer computing to get computing resources to scientific projects, so that organizations can be able to attain large computing power from volunteer clients instead of making a high investment in infrastructure. However, there are projects, like the ATLAS@Home project,13 in which the number of running jobs has reached a plateau, due to a high load on data servers and networks caused by file transfers. Alonso et al,14 in the paper A New Volunteer Computing Model for Data-Intensive Applications, provide an alternative, named ComBoS, to improve the performance of volunteer computing projects that have reached their limit due to the I/O bottleneck in data servers by having a percentage of the volunteer clients running as data servers, called data volunteers, to reduce the load on data servers. This solution also improves data locality, leveraging the network latencies of closer machines, as shown by the performance increase provides by their solution, applied to three different BOINC projects. Two current trends in Big Data processing have made the usage of GPGPUs very popular in HEDICS: information discovery and deep-learning techniques15 and collective video games.16 In both cases, there is an increasing trend to discharge client nodes by sending bulk computing to heterogeneous high-end computing nodes for data processing. Data compression is an important area in many data management applications, like training of deep learning, where data must be decompressed many times. Nakano et al,17 in their paper Adaptive Loss-Less Data Compression Method Optimized for GPU Decompression, present a novel lossless data compression method, called Adaptive LossLess (ALL) data compression, designed with the objective of performing decompression very efficiently on the GPU. Evaluations of the ALL data compression method against published lossless data compression methods implemented in GPU show improvements between 1.22 and 23.5 times running on the same GPU. Due to the massive extension of many mobile applications, such as sensors and smart phones, it is crucial for HEDICS to offload applications to high-end nodes so that low-power devices can be used as clients. One possible approach to deal with this problem is the solution proposed in the paper Accelerating Linux and Android applications on low-power devices through remote GPGPU offloading by Montella et al.18 They describe the architecture and integration of RAPID, a complete framework suite for computation offloading to help low-powered devices overcome these limitations. RAPID supports CPU and GPGPU computation offloading on Linux and Android devices, providing lightweight secure data transmission of the offloading operations. The proposed framework is highly modular and exposes a rich Application Programming Interface (API) to developers, making it highly versatile while hiding the complexity of the underlying networking layer. The evaluation results show that Java/Android GPGPU code offloading is possible, through a BioSurveillance application, a commercial real-time face recognition application. High-end networked scientific and engineering applications requires usually HPC for numerical computing and large storage capabilities at end nodes. As the problem grows, the scientific community, in its never-ending road of larger and more efficient computational resources, is in need of more efficient implementations that can adapt better to the current parallel platforms and in need of more new solutions for memory problems that are now memory bound. The memory problem is addressed by Valero19 in the paper Reducing Memory Requirements for Large Size LBM Simulations on GPUs, where he proposes some initiatives to minimize the memory requirements of the Lattice- Boltzmann Method for its usage on GPGPUs to run large scale simulations. The proposed approach allows the author to execute bigger simulations on the same platform without additional memory transfers, those achieving a high performance. In particular, the paper presents two new implementations, LBM-Ghost and LBM-Swap, which are deeply analysed, presenting the pros and cons of each of them. The need of parallelization at high-end nodes is addressed in the paper Parallel solvers for fractional power diffusion problems by Starikoviius et al. 20 The authors construct and investigate parallel solvers for problems described by fractional powers of elliptic operators, like fractional diffusion. Three state-of-the-art approaches are used to transform the non-local fractional-order differential problem into local partial differential equation problems formulated in a space of higher dimension. Scalability of the developed parallel algorithms is investigated, and their parallel performance is compared in the paper. Finally, the problem of accuracy and efficiency for statistic distributions is addressed by Monni et al21 in the paper Fitting Long-Tailed Distribution to Empirical Data. The authors discuss about the limits of the analysis of empirical fat-tailed distributions, which can describe a variety of evolving systems, both natural and man-made. An algorithm to fit fat-tailed distributions is presented and tested against samplings of the power law, the Yule, the log-normal, and Weibull distributions. The algorithm is general and can be applied to any numerical dataset. Thus, the authors compute the parameters defining the shape of each distribution and test the results against simulations. Their method with another state-of- the-art technique to estimate the parameters of empirical distributions. The accuracy of the estimations is discussed, and they conclude that their method based on a weighted iterated χ2 test performs better than the other. Power laws can fit a variety of distributions coming from real data, so a systematic approach to the measurement of the accuracy of fitting algorithms is essential. Articles presented in this special issue provide recent advances in some fields related to high-end data-intensive computing systems. They were selected by invitation of best ranked from two conferences and peer reviewed by journal selection. Acceptance rate for the special issue was below 50% of the invited papers. We hope that the ideas presented in this special issue can contribute to this strategically important, exciting, and fast growing research area and will be of interest for readers of the journal. As guest editors of this special issue, we would like to express our gratitude to all of the authors who submitted their papers to this special issue, and to the Reviewers that helped us with their hard work and the feedback provided to the authors. We also wish to express our gratitude to the Editor-in-Chief Geoffrey C. Fox for the opportunity to edit this special issue and his assistance during the special issue preparation. We acknowledge the following Reviewing Committee members: Pawe Czarnul (Poland), Guilherme Dinis (Sweden), Ece Guran Schmidt (Turkey), Massimiliano Ferrara (Italy), Shih-Hao Hung (Taiwan), Florin Isaila (Spain), Dingde Jiang (China), Amin Khan (Portugal), Marcin Kostur (Poland), Kenli Li (China), Francesco Longo (italy), Francesc Lordan (Spain), Najme Mansouri (Iran), Panagiotis Michailidis (Greece), Eike Mueller (UK), Tomas Potuzak (Cezch Republic), Philipp Reinecke (Germany), Francisco Rodrigo (Spain), Gopal Shyam (India), Shengen Yan (China), Wenwu Tang (USA), and Peng Zhang (USA). Jesús Carretero 0001, Francisco Javier García Blas, Koji Nakano, Peter Mueller |
Concurr. Comput. Pract. Exp. | 2 |
| 2017 | A data-aware scheduling strategy for workflow execution in cloudsabstractSummary As data intensive scientific computing systems become more widespread, there is a necessity of simplifying the development, deployment, and execution of complex data analysis applications for scientific discovery. The scientific workflow model is the leading approach for designing and executing data‐intensive applications in high‐performance computing infrastructures. Commonly, scientific workflows are built by a set of connected tasks arranged in a directed acyclic graph style, which communicate through storage abstractions. The Data Mining Cloud Framework (DMCF) is a system allowing users to design and execute data analysis workflows on cloud platforms, relying on cloud storage services for every I/O operation. Hercules is an in‐memory I/O solution that can be used in DMCF as an alternative to cloud storage services, providing additional performance and flexibility features. This work improves the integration between DMCF and Hercules by using a data‐aware scheduling strategy for exploiting data locality in data‐intensive workflows. This paper presents experimental results demonstrating the performance improvements achieved using the proposed data‐aware scheduling strategy in the Microsoft Azure cloud platform. In particular, with our scheduling strategy, the I/O overhead has been reduced by 55% with respect to the Azure storage, leading to a 20% reduction of the total execution time. Fabrizio Marozzo, Francisco Rodrigo Duro, Francisco Javier García Blas, Jesús Carretero 0001, Domenico Talia, Paolo Trunfio |
Concurr. Comput. Pract. Exp. | 3 |
| 2017 | Boosting analyses in the life sciences via clusters, grids and clouds
Sandra Gesing, Jesús Carretero 0001, Francisco Javier García Blas, Johan Montagnat |
Future Gener. Comput. Syst. | 3 |
| 2017 | Experimental evaluation of a flexible I/O architecture for accelerating workflow engines in ultrascale environments
Francisco Rodrigo Duro, Francisco Javier García Blas, Florin Isaila, Jesús Carretero 0001, Justin M. Wozniak, Robert B. Ross |
Parallel Comput. | 2 |
| 2017 | Virtual Environments and Advanced Interfaces
Daphne Economou, Markos Mentzelopoulos, Nektarios Georgalas, Jesús Carretero 0001, Francisco Javier García Blas |
Pers. Ubiquitous Comput. | 5 |
| 2017 | Model-based energy-aware data movement optimization in the storage I/O stack
Pablo Llopis, Florin Isaila, Francisco Javier García Blas, Jesús Carretero 0001 |
J. Supercomput. | 3 |
| 2016 | Flexible Data-Aware Scheduling for Workflows over an In-memory Object StoreabstractThis paper explores novel techniques for improving the performance of many-task workflows based on the Swift scripting language. We propose novel programmer options for automated distributed data placement and task scheduling. These options trigger a data placement mechanism used for distributing intermediate workflow data over the servers of Hercules, a distributed key-value store that can be used to cache file system data. We demonstrate that these new mechanisms can significantly improve the aggregated throughput of many-task workflows with up to 86x, reduce the contention on the shared file system, exploit the data locality, and trade off locality and load balance. Francisco Rodrigo Duro, Francisco Javier García Blas, Florin Isaila, Justin M. Wozniak, Jesús Carretero 0001, Robert B. Ross |
CCGrid | 2 |
| 2016 | A C++ Generic Parallel Pattern Interface for Stream Processing
David del Rio Astorga, Manuel F. Dolz, Luis Miguel Sánchez, Francisco Javier García Blas, José Daniel García |
ICA3PP | 4 |
| 2016 | Porting Matlab Applications to High-Performance C++ Codes: CPU/GPU-Accelerated Spherical Deconvolution of Diffusion MRI Data
Francisco Javier García Blas, Manuel F. Dolz, José Daniel García, Jesús Carretero 0001, Alessandro Daducci, Yasser Alemán-Gómez, Erick Jorge Canales-Rodríguez |
ICA3PP | 1 |
| 2016 | Improving the Energy Efficiency of MPI Applications by Means of MalleabilityabstractThis work presents two novel techniques for increasing the energy efficiency of parallel applications by means of malleability. These techniques are implemented as an extension of Flex-MPI, a library implemented on top of MPI, which provides performance-aware dynamic reconfiguration for MPI-based applications. During the application execution, Flex-MPI performs energy and performance monitoring by means of energy and performance counters. It leverages this information in order to adapt the program performance using two energy policies: energy minimization and performance-per-watt maximization. The evaluation results show that these new energy-aware capabilities permit MPI applications to be executed in an optimized way both in terms of performance and energy efficiency. Manuel Rodriguez-Gonzalo, David E. Singh, Francisco Javier García Blas, Jesús Carretero 0001 |
PDP | 3 |
| 2016 | Introduction to sustainable ultrascale computing systems and applications
Jesús Carretero 0001, Francisco Javier García Blas, Raimondas Ciegis |
J. Supercomput. | 2 |
| 2016 | Exploiting in-memory storage for improving workflow executions in cloud platforms
Francisco Rodrigo Duro, Fabrizio Marozzo, Francisco Javier García Blas, Domenico Talia, Paolo Trunfio |
J. Supercomput. | 3 |
| 2016 | Analyzing the energy consumption of the storage data path
Pablo Llopis, Manuel F. Dolz, Francisco Javier García Blas, Florin Isaila, Mohammad Reza Heidari, Michael Kuhn 0003 |
J. Supercomput. | 3 |
| 2015 | Simulation Platform for X-Ray Computed Tomography Based on Low-Power Systems
Estefania Serrano, Francisco Javier García Blas, Alberto Verza, Jesús Carretero 0001 |
ICA3PP (4) | 2 |
| 2015 | A comparative study of an X-ray tomography reconstruction algorithm in accelerated and cloud computing systemsabstractSummary With the increase of resolution in medical image scanners and the need of faster reconstruction methods, new ways of exploiting the inherent parallelism of reconstruction algorithms have arisen. In this paper, we present Mangoose++, an application to perform X‐ray computed tomography that supports multiple grades of parallelism. This parallelism is tackled with two different approaches: the usage of parallel nodes with multicore CPUs in a cloud environment and the usage of high‐performance computing (HPC)‐based parallel architectures such as general‐purpose computing on graphics processing unit (GPGPU) or Intel Xeon Phi. In this paper, we show the design and implementation of the application in three types of platforms related to the previous mentioned approaches, comparing and analyzing the performance, resource utilization, and scalability of each platform. Accelerators offer high performance for data sizes that fit inside the accelerator memory. This is the main advantage of Intel Xeon Phi that, in this work, obtains similar performance results than compute unified device architecture (CUDA)‐based GPGPU versions, comparing with cards with less memory capacity. In our evaluation experiments, we additionally analyze and discuss the costs and efficiency of Mangoose++ over Amazon Compute Cloud platform, demonstrating that lower times can be achieved in a reasonable price compared with owned HPC‐based hardware. A comparison between distinct hardware configurations is provided for emphasizing on the advantages and disadvantages of each one. Copyright © 2015 John Wiley & Sons, Ltd. Estefania Serrano, Francisco Javier García Blas, Jesús Carretero 0001 |
Concurr. Comput. Pract. Exp. | 2 |
| 2014 | Evaluation of the Feasibility of Making Large-Scale X-Ray Tomography Reconstructions on CloudsabstractThis work focuses on the evaluation of the suitability of Mangoose++, a medical image application for reconstruction of 3D volumes, by means of Cloud Computing. Due to the increasing resolution of panel detectors in computed tomography and the need of lower execution times, the use of parallel implementations for clusters and accelerators have been generalized. Anyhow, the renewal and maintenance of hardware is expensive which makes Cloud Computing a valuable alternative. In our evaluation, we analyze and discuss the costs and efficiency of the Mangosee++ application over Amazon EC2 platform, demonstrating that lower times can be achieved in a reasonable price compared with owned HPC-based hardware. We also provide a comparison between distinct hardware configurations so that we can infer the advantages and disadvantages of each one. Estefania Serrano, Guzmán Bermejo, Francisco Javier García Blas, Jesús Carretero 0001 |
CCGRID | 3 |
| 2014 | High-performance X-ray tomography reconstruction algorithm based on heterogeneous accelerated computing systemsabstractMany medical image processing applications need high processing speed to achieve almost real-time image reconstruction features. Due to that, massively parallel architectures based on accelerators have become very popular in the area, specially GPGPUs. In this paper we show Mangoose++, an application to perform X-Ray Computed Tomography (CT) from medical image based on a new implementation of the FDK algorithm. Mangoose++ have been designed and implemented to exploit the parallelism existing on several hardware accelerators platforms, as GPGPUs and Intel Xeon Phi accelerators. In this paper we show the design and implementation of the application in three types of platforms, multi-core CPU, GPGPU, and Intel Xeon Phi, and the evaluation made to test the performance, resource utilization, and scalability of each platform. Moreover, to avoid hardware dependencies, we have also implemented the application using the OpenACC runtime to check portability and the overhead incurred when using runtimes. The evaluation results show that our solution is faster than recent related works and that, in terms of computation, Intel Xeon Phi and the CUDA-based GPU versions obtain similar results as the problem size increases. Moreover, the evaluation shows that using OpenACC, we have enhanced programmability because there is a single version of the source code. But it also shows that using OpenACC heavily affects performance of Mangoose++, which is reduced in a 50% when compared with the many-core versions, even when it is not so drastical when compared to the CPU version. Estefania Serrano, Guzmán Bermejo, Francisco Javier García Blas, Jesús Carretero 0001 |
CLUSTER | 3 |
| 2014 | Survey of Energy-Efficient and Power-Proportional Storage SystemsabstractIncreasingly, large-scale computing systems are consuming more power each passing year. As power consumption is on the rise, concern has been raised over the growing implications on power bills, carbon emissions and power supply limitations for data centers. Computing systems use hardware components that are not power proportional and machines comprising large systems tend to be underutilized, resulting in great wastage of energy. Motivated by this fact, researchers aim to improve the energy efficiency of these systems by increasing resource utilization and exploiting system characteristics in order to achieve power proportionality. However, many challenges burden the task of producing a power-aware, energy-efficient large-scale system that provides the same performance as today's systems. Particularly, storage systems are of great concern since storage consumes a large amount of power in the data center. This paper outlines the main problems and challenges of delivering power-efficient storage solutions, and proposes a taxonomy of power-aware techniques. This work aims to provide a detailed exposition of current trends in power-aware storage systems in a comprehensive and organized manner, outlining and comparing current solutions and challenges. Pablo Llopis, Francisco Javier García Blas, Florin Isaila, Jesús Carretero 0001 |
Comput. J. | 2 |
| 2014 | Surfing the optimization space of a multiple-GPU parallel implementation of a X-ray tomography reconstruction algorithm
Francisco Javier García Blas, Mónica Abella, Florin Isaila, Jesús Carretero 0001, Manuel Desco |
J. Syst. Softw. | 1 |
| 2014 | CONDESA: A Framework for Controlling Data Distribution on Elastic Server ArchitecturesabstractApplications running in today's data centers show high workload variability. While seasonal patterns, trends and expected events may help building proactive resource allocation policies, this approach has to be complemented with adaptive strategies which should address unexpected events such as flash crowds and volume spikes. Additionally, the limitations of current I/O infrastructures in the face of dramatic increase of data generation require, the ability to build novel abstractions and models for robust decision making regarding data layout and data locality. In this work, we present CONDESA (CONtrolling Data distribution on Elastic Server Architectures), a framework for exploring adaptive data distribution strategies for elastic server architectures. To the best of our knowledge CONDESA is the first platform that permits to systematically study the interplay between five data related strategies: workload prediction, adaptive control of data distribution and server provisioning, adaptive data grouping, adaptive data placement, and adaptive system sizing. We demonstrate how CONDESA can be used for browsing the design space of adaptive data distribution policies. We show how prediction models can be compared in terms of overhead and accuracy. We evaluate the impact of change detection on prediction accuracy and how CONDESA can be used for choosing an adequate prediction horizon. We demonstrate how adaptive prediction can be used for sizing a server system. Finally, we show how prediction models, change detection strategies, and data placement policies can be combined and compared based on server utilization, load balance, data locality, over- and underprovisioning. Juan Manuel Tirado, Daniel Higuero, Francisco Javier García Blas, Florin Isaila, Jesús Carretero 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2013 | Parallel implementation of a X-ray tomography reconstruction algorithm based on MPI and CUDAabstractMost small-animal X-ray computed tomography (CT) scanners are based on cone-beam geometry with a flat-panel detector orbiting in a circular trajectory. Image reconstruction in these systems is usually performed by approximate methods based on the algorithm proposed by Feldkamp, Davis and Kress (FDK). Currently there is a strong need to speedup the reconstruction of X-Ray CT data in order to extend its clinical applications. The evolution of the semiconductor detector panels has resulted in an increase of detector elements density, which produces a higher amount of data to process. This work focuses on future high-resolution studies (density up to 4096 pixeles), in which multiple level of parallelism will be needed in the reconstruction. In addition, this paper addresses the future challenges of processing high-resolution images in many-core and distributed architectures. In our evaluation section we demonstrate that our solution is 17% faster than recent related works. Francisco Javier García Blas, Florin Isaila, Mónica Abella, Jesús Carretero 0001, Ernesto Liria, Manuel Desco |
EuroMPI | 1 |
| 2013 | A hierarchical parallel storage system based on distributed memory for large scale systemsabstractThis paper presents the design and implementation of a storage system for high performance systems based on a multiple level I/O caching architecture. The solution relies on Memcached as a parallel storage system, preserving its powerful capacities such as transparency, quick deployment, and scalability. The designed parallel storage system targets to reduce the I/O latency in data-intensive high performance applications. The proposed solution consists of a user-level library and extended Memcached servers. The solution aims to be hierarchical by deploying Memcached-based I/O servers across all the infrastructure data path. Our experiments demonstrate that our solution is up to 40% faster than PVFS2. Francisco Rodrigo Duro, Francisco Javier García Blas, Jesús Carretero 0001 |
EuroMPI | 2 |
| 2012 | A Black Box Model for Storage Devices Based on Probability DistributionsabstractTraditional approaches for storage devices simulation have been based on detailed analytical models. However, detailed models require detailed computations which may be not affordable for large scale simulations. Moreover, highly detailed models cannot be easily generalized. A different approach is the black-box statistical modeling, where the storage device, its interface, and the interconnection mechanisms are modeled as a single stochastic process, defining the request response time as a random variable with an unknown distribution. A random variate generator can be built and integrated into a simulation model. This approach allows to generate a simulation model for both real and synthetic workloads. This article describes a method suitable for building fast simulation models for storage devices. Our method uses as starting point a workload and produces a random variate generator which can be easily integrated into large scale simulation models. A comparison between our variate generator and the widely known simulation tool DiskSim, shows that our variate generator is faster, and can be as accurate as DiskSim. Laura Prada, Alejandro Calderón 0001, Francisco Javier García Blas, José Daniel García, Jesús Carretero 0001 |
ISPA | 3 |
| 2011 | A Power-Aware Based Storage Architecture for High Performance ComputingabstractThe energy crisis of the last years and the ever increasing conscience about the negative effects of energy waste on the climate change have brought the sustainability both into public attention, industry, and scientific scrutiny. Energy demand has been increasing in many subsystems, specially in data centers and supercomputers. This paper considers the problem of saving energy on storage systems taking advantage of SSD devices. SSDs and magnetic disk devices offer different power characteristics, being SSD devices much less power consuming than conventional magnetic disk devices. We propose a novel power saving solution based on SSD devices, namely SSD-PASS. Our storage system obtains benefits of permanent caching on SSDs in storage nodes. Nowadays we can find solutions that do not consider the viability and feasibility of the SSD-based storage systems, in terms of monetary cost. We present a cost analysis and evaluate our proposed architecture, in terms of saved energy and performance. Our cost model takes into account magnetic disk and SSD devices reliability metrics, current energy prices, and replacement costs. The experimental results demonstrate that our solution achieves a significant reduction in energy consumption and subsequent monetary savings by up to 66\%. We have evaluated the proposed approach with realistic workloads of three well-known HPC applications. Laura Prada, Francisco Javier García Blas, José Daniel García, Jesús Carretero 0001, Alberto Nuñez |
HPCC | 2 |
| 2011 | Power saving-aware prefetching for SSD-based systems
Laura Prada, Francisco Javier García Blas, José Daniel García, Jesús Carretero 0001 |
J. Supercomput. | 2 |
| 2011 | Design and Evaluation of Multiple-Level Data Staging for Blue Gene SystemsabstractParallel applications currently suffer from a significant imbalance between computational power and available I/O bandwidth. Additionally, the hierarchical organization of current Petascale systems contributes to an increase of the I/O subsystem latency. In these hierarchies, file access involves pipelining data through several networks with incremental latencies and higher probability of congestion. Future Exascale systems are likely to share this trait. This paper presents a scalable parallel I/O software system designed to transparently hide the latency of file system accesses to applications on these platforms. Our solution takes advantage of the hierarchy of networks involved in file accesses, to maximize the degree of overlap between computation, file I/O-related communication, and file system access. We describe and evaluate a two-level hierarchy for Blue Gene systems consisting of client-side and I/O node-side caching. Our file cache management modules coordinate the data staging between application and storage through the Blue Gene networks. The experimental results demonstrate that our architecture achieves significant performance improvements through a high degree of overlap between computation, communication, and file I/O. Florin Isaila, Francisco Javier García Blas, Jesús Carretero 0001, Robert Latham, Robert B. Ross |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2009 | Latency Hiding File I/O for Blue Gene SystemsabstractThis paper presents the design and implementation of a novel file I/O solution for Blue Gene systems. We propose a hierarchical I/O cache architecture based on open source software. Our solution is based on an asynchronous data staging strategy, which hides the latency of file system access from compute nodes. The performance results demonstrate the high scalability and significant performance improvements of our architecture over existing solutions. Florin Isaila, Francisco Javier García Blas, Jesús Carretero 0001, Robert Latham, Samuel Lang, Robert B. Ross |
CCGRID | 2 |
| 2008 | View-Based Collective I/O for MPI-IOabstractThis paper presents the design and implementation of a new file system independent collective I/O optimization based on file views: view-based collective I/O. View-based collective I/O has been implemented and evaluated inside ROMIO implementation of MPI-IO standard. The evaluation section shows that view-based I/O outperforms the original two-phase collective I/O from ROMIO in most of the cases for three well-known parallel I/O benchmarks. This is especially due to a smaller cost of scatter/gather operations, a reduction of the metadata overhead, and a smaller number of collective communication and synchronization primitives used in the implementation. Francisco Javier García Blas, Florin Isaila, David E. Singh, Jesús Carretero 0001 |
CCGRID | 1 |
| 2008 | AHPIOS: An MPI-Based Ad Hoc Parallel I/O SystemabstractThis paper presents the design and implementation of a portable ad-hoc parallel I/O system (AHPIOS). AHPIOS virtualizes on-demand available distributed storage resources and allows the files to be striped over several storage devices. Additionally, the design unifies the configuration of the MPI-IO library and the AHPIOS data servers. By a strong integration of the application, MPI-IO library and file system, a significant performance improvement can be achieved. The experimental section shows that the full MPI-IO integrated AHPIOS implementation of file access operations outperforms the existing MPI-IO implementation by as much as 495% for file writes and 522% for file reads. Florin Isaila, Francisco Javier García Blas, Jesús Carretero 0001, Wei-keng Liao, Alok N. Choudhary |
ICPADS | 2 |