VLDB 2026 Research / reviewers in the wild / expert
Juan Touriño
dblp:t/JuanTourino
· DBLP profile ↗
88ranked-venue papers
4as first author
14since 2021 · last 2026
0000-0001-9670-1933ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 63 · 3 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 4 since 2021Software engineering, systems software and programming languages · 5 · 2 since 2021Computer networks · 4Databases, data management, data science and information retrieval · 3 · 1 first-authorArtificial intelligence and machine learning · 1Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Weight-based Disk I/O Scaling for Serverless ContainersabstractAbstract Disk bandwidth is a critical resource for I/O-intensive applications that must transfer large volumes of data to and from persistent storage. Most multi-tenant infrastructures efficiently allocate CPU and memory resources to concurrent workloads, but typically lack mechanisms for allocating I/O bandwidth. As a result, users often resort to exclusive node reservations to avoid disk contention, which can lead to underutilisation of other node resources if not fully exploited. Another common issue is that users do not know the exact resource requirements of their applications. Even when this is known, applications rarely maintain peak resource usage throughout their entire execution, resulting in wasted resources that could otherwise benefit other users. Today, many users prefer cloud serverless platforms because of their ease of use and flexible billing. However, these platforms have inherent limitations and may not be suitable for workloads with specific requirements. In this paper, we present a serverless scaling mechanism that dynamically adjusts disk I/O bandwidth for containerised applications by scaling their allocation up or down based on real-time usage and configurable weights. In addition, the system incorporates automatic extension management capabilities for virtual disk devices, such as logical volumes. Our approach can be integrated with other serverless scaling mechanisms, such as CPU and memory management, to provide a comprehensive resource scaling solution. The experimental results have shown significant performance improvements, with overall runtime reductions of up to 53% for concurrent I/O-intensive workloads compared to running them without serverless capabilities. Óscar Castellanos-Rodríguez, Roberto R. Expósito, Jonatan Enes, Juan Touriño |
J. Grid Comput. | 4 |
| 2025 | Modular Construction and Optimization of the UZP Sparse Format for SpMV on CPUsabstractSparse data structures are ubiquitous in modern computing, and numerous formats have been designed to represent them. These formats may exploit specific sparsity patterns, aiming to achieve higher performance for key numerical computations than more general-purpose formats such as CSR and COO. In this work presents UZP, a new sparse format based on polyhedral sets of integer points. UZP is a flexible format that subsumes CSR, COO, DIA, BCSR, etc., by raising them to a common mathematical abstraction: a union of integer polyhedra, each intersected with an affine lattice. We present a modular approach to building and optimizing UZP: it captures equivalence classes for the sparse structure, enabling the tuning of the representation for target-specific and application-specific performance considerations. UZP is built from any input sparse structure using integer coordinates, and is interoperable with existing software using CSR and COO data layouts. We provide detailed performance evaluation of UZP on 200+ matrices from SuiteSparse, demonstrating how simple and mostly unoptimized generic executors for UZP can already achieve solid performance by exploiting 𝒵-polyhedra structures. Alonso Rodríguez-Iglesias, Santoshkumar T. Tongli, Emily Tucker, Louis-Noël Pouchet, Gabriel Rodríguez 0001, Juan Touriño |
Proc. ACM Program. Lang. | 6 |
| 2024 | Automated Approach for Accurate CPU Power ModellingabstractPower supply is a limiting factor when increasing the computing capacity of supercomputers. As a consequence, power consumption has become one of the biggest challenges in the field of High Performance Computing (HPC). In order to develop energy-efficient tools (e.g., frameworks, applications), it is essential to have an accurate power consumption modelling. Al-though previous works proposed a wide variety of approaches to model CPU power consumption, building models in an automated and adaptable way to changing scenarios and predicting power with high precision remains complex due to multiple factors (e.g., training and test workloads, model variables). In this paper, we present a set of tools to fully automate the process of modelling power consumption using CPU time series data. More specifically, our proposal includes two tools: (1) CPUPowerWatcher, which gathers CPU metrics during the execution of user-configurable workloads; and (2) CPUPowerSeer, which builds models to predict CPU power consumption (e.g., polynomial regression) from different CPU variables (e.g., usage, clock frequency) using time series data. Thus, multiple models can be created and evaluated easily, allowing the selection of an optimal model for a specific workload. The experiments conducted by combining these tools allow analysing the impact of novel factors on CPU power consumption, such as the type of CPU usage generated by different workloads or how the CPU cores are allocated to them. In addition, the accuracy of six regression models is compared when predicting CPU- and I/O-intensive workloads using two different core allocations. Tomé Maseda, Jonatan Enes, Roberto R. Expósito, Juan Touriño |
CLUSTER | 4 |
| 2024 | Serverless-like platform for container-based YARN clustersabstractServerless computing is an emerging paradigm that has gained a lot of relevance in recent years, as it allows users to consume computing resources without worrying about the underlying infrastructure and pay only for what they actually use. Most current services that implement this paradigm typically rely on the Function-as-a-Service (FaaS) model, which works perfectly for simple applications based on stateless functions triggered by specific events. However, these services are not designed to run more complex applications with intricate interactions, usually presenting a significant degree of configuration difficulty and/or low ability to customise the execution environment. They also tend to be designed for short and simple workloads, with some services even limiting their maximum runtime to just a few minutes. In this paper, we present a platform based on Hadoop YARN oriented to the execution of Big Data workloads in a containerised and serverless way, so that the resources allocated to such containers are automatically and dynamically scaled according to their actual usage. An experimental evaluation has been carried out to compare our serverless-like platform with a standard YARN deployment when executing Big Data workloads concurrently. Our results have shown experimental evidence of enhancing both performance and overall resource efficiency, providing runtime reductions and resource usage improvements of up to 41% and 50%, respectively. Óscar Castellanos-Rodríguez, Roberto R. Expósito, Jonatan Enes, Guillermo L. Taboada, Juan Touriño |
Future Gener. Comput. Syst. | 5 |
| 2024 | CUDA acceleration of MI-based feature selection methodsabstractFeature selection algorithms are necessary nowadays for machine learning as they are capable of removing irrelevant and redundant information to reduce the dimensionality of the data and improve the quality of subsequent analyses. The problem with current feature selection approaches is that they are computationally expensive when processing large datasets. This work presents parallel implementations for Nvidia GPUs of three highly-used feature selection methods based on the Mutual Information (MI) metric: mRMR, JMI and DISR. Publicly available code includes not only CUDA implementations of the general methods, but also an adaptation of them to work with low-precision fixed point in order to further increase their performance on GPUs. The experimental evaluation was carried out on two modern Nvidia GPUs (Turing T4 and Ampere A100) with highly satisfactory results, achieving speedups of up to 283x when compared to state-of-the-art C implementations. Bieito Beceiro, Jorge González-Domínguez, Laura Moran-Fernandez, Verónica Bolón-Canedo, Juan Touriño |
J. Parallel Distributed Comput. | 5 |
| 2023 | Clupiter: a Raspberry Pi mini-supercomputer for educational purposesabstractThe main objective of this work is to bring supercomputing and parallel processing closer to non-specialized audiences by building a Raspberry Pi cluster, called Clupiter, which emulates the operation of a supercomputer. It consists of eight Raspberry Pi devices interconnected to each other so that they can run jobs in parallel. To make it easier to show how it works, a web application has been developed. It allows launching parallel applications and accessing a monitoring system to see the resource usage when these applications are running. The NAS Parallel Benchmarks (NPB) are used as demonstration applications. From this web application a couple of educational videos can also be accessed. They deal, in a very informative way, with the concepts of supercomputing and parallel programming. Alonso Rodríguez-Iglesias, María J. Martín, Juan Touriño |
TrustCom | 3 |
| 2023 | PATO: genome-wide prediction of lncRNA-DNA triple helicesabstractMOTIVATION: Long non-coding RNA (lncRNA) plays a key role in many biological processes. For instance, lncRNA regulates chromatin using different molecular mechanisms, including direct RNA-DNA hybridization via triplexes, cotranscriptional RNA-RNA interactions, and RNA-DNA binding mediated by protein complexes. While the functional annotation of lncRNA transcripts has been widely studied over the last 20 years, barely a handful of tools have been developed with the specific purpose of detecting and evaluating lncRNA-DNA triple helices. What is worse, some of these tools have nearly grown a decade old, making new triplex-centric pipelines depend on legacy software that cannot thoroughly process all the data made available by next-generation sequencing (NGS) technologies. RESULTS: We present PATO, a modern, fast, and efficient tool for the detection of lncRNA-DNA triplexes that matches NGS processing capabilities. PATO enables the prediction of triple helices at the genome scale and can process in as little as 1 h more than 60 GB of sequence data using a two-socket server. Moreover, PATO's efficiency allows a more exhaustive search of the triplex-forming solution space, and so PATO achieves higher levels of prediction accuracy in far less time than other tools in the state of the art. AVAILABILITY AND IMPLEMENTATION: Source code, user manual, and tests are freely available to download under the MIT License at https://github.com/UDC-GAC/pato. Iñaki Amatria-Barral, Jorge González-Domínguez, Juan Touriño |
Bioinform. | 3 |
| 2023 | SeQual-Stream: approaching stream processing to quality control of NGS datasetsabstractBACKGROUND: Quality control of DNA sequences is an important data preprocessing step in many genomic analyses. However, all existing parallel tools for this purpose are based on a batch processing model, needing to have the complete genetic dataset before processing can even begin. This limitation clearly hinders quality control performance in those scenarios where the dataset must be downloaded from a remote repository and/or copied to a distributed file system for its parallel processing. RESULTS: In this paper we present SeQual-Stream, a streaming tool that allows performing multiple quality control operations on genomic datasets in a fast, distributed and scalable way. To do so, our approach relies on the Apache Spark framework and the Hadoop Distributed File System (HDFS) to fully exploit the stream paradigm and accelerate the preprocessing of large datasets as they are being downloaded and/or copied to HDFS. The experimental results have shown significant improvements in the execution times of SeQual-Stream when compared to a batch processing tool with similar quality control features, providing a maximum speedup of 2.7[Formula: see text] when processing a dataset with more than 250 million DNA sequences, while also demonstrating good scalability features. CONCLUSION: Our solution provides a more scalable and higher performance way to carry out quality control of large genomic datasets by taking advantage of stream processing features. The tool is distributed as free open-source software released under the GNU AGPLv3 license and is publicly available to download at https://github.com/UDC-GAC/SeQual-Stream . Óscar Castellanos-Rodríguez, Roberto R. Expósito, Juan Touriño |
BMC Bioinform. | 3 |
| 2023 | pRIblast: A highly efficient parallel application for comprehensive lncRNA-RNA interaction predictionabstractLong non-coding RNAs (lncRNAs) play a key role in several biological processes and scientists are constantly trying to come up with new strategies to elucidate their functions. One common approach to characterize these sequences consists in predicting their interactions with other RNA fragments. Nevertheless, the high computational cost of the bioinformatics tools developed for this purpose prevents their application to large-scale datasets. This paper presents pRIblast, a highly efficient parallel application for comprehensive lncRNA–RNA interaction prediction based on the state-of-the-art RIblast tool, which has been proved to show superior biological accuracy compared to other counterparts in previous experimental evaluations. Benchmarking on a multicore CPU cluster shows that pRIblast is able to compute in a few hours analyses that would need more than three months to complete with the original RIblast algorithm, always achieving the same level of prediction accuracy. Furthermore, this novel application can process large input datasets that cannot be processed with the former tool. pRIblast is free software publicly available to download at https://github.com/UDC-GAC/pRIblast under the MIT license. Iñaki Amatria-Barral, Jorge González-Domínguez, Juan Touriño |
Future Gener. Comput. Syst. | 3 |
| 2023 | ParRADMeth: Identification of Differentially Methylated Regions on Multicore ClustersabstractThe discovery of Differentially Methylated (DM) regions is an important research field in biology, as it can help to anticipate the risk of suffering from specific diseases. Nevertheless, the high computational cost of the bioinformatic tools developed for this purpose prevents their application to large-scale datasets. Hence, much faster tools are required to further progress in this research field. In this work we present ParRADMeth, a parallel tool that applies beta-binomial regression for the identification of these DM regions. It is based on the state-of-the-art sequential tool RADMeth, which proved superior biological accuracy compared to counterparts in previous experimental evaluations. ParRADMeth provides the same DM regions as RADMeth but at significantly reduced runtime thanks to exploiting the compute capabilities of common multicore CPU clusters. For example, our tool is up to 189 times faster for real data experiments on a cluster with 16 nodes, each one containing two eight-core processors. The source code of ParRADMeth, as well as a reference manual, are available at https://github.com/UDC-GAC/ParRADMeth. Alejandro Fernández-Fraga, Jorge González-Domínguez, Juan Touriño |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2022 | Custom High-Performance Vector Code Generation for Data-Specific Sparse ComputationsabstractSparse computations, such as sparse matrix-dense vector multiplication, are notoriously hard to optimize due to their irregularity and memory-boundedness. Solutions to improve the performance of sparse computations have been proposed, ranging from hardware-based such as gather-scatter instructions, to software ones such as generalized and dedicated sparse formats, used together with specialized executor programs for different hardware targets. These sparse computations are often performed on read-only sparse structures: while the data themselves are variable, the sparsity structure itself does not change. Indeed, sparse formats such as CSR have a typically high cost to insert/remove nonzero elements in the representation. The typical use case is to not modify the sparsity during possibly repeated computations on the same sparse structure. Marcos Horro, Louis-Noël Pouchet, Gabriel Rodríguez 0001, Juan Touriño |
PACT | 4 |
| 2022 | MARTA: Multi-configuration Assembly pRofiler and Toolkit for performance AnalysisabstractBenchmarking to characterize specific software or hardware features is an error-prone, arduous and repetitive task. Designing a specialized experimental setup frequently requires writing new scripts or ad-hoc programs in order to properly exhibit interesting performance effects, using code changes and hardware events measurements. These artifacts may have limited reusability for subsequent experiments, since they are dependent on specific problems and, in some cases, platforms. To improve productivity and reproducibility of such experiments, which are often investigative in nature, we introduce MARTA: a fully customizable toolkit that aims to increase productivity by generating benchmark templates, compiling them, and profiling the regions of interest (RoI) specified using hardware events, and performing static code analysis. MARTA can also be applied on existing code regions of interest, it only requires to write a simple configuration file. In an orthogonal dimension, the system is able to run various statistical analyses on the measurements collected. MARTA uses data mining and machine learning or AI-based techniques for classification and regression, automatically extracting the features of the experimental setup which have the most impact on performance or whichever other metric of interest, given a large set of experiments and dimensions to consider. These post-processing tasks are valuable for deriving knowledge from experiments and are not included in most profiling tools. We also provide a set of cases of study to illustrate the ability of MARTA to conveniently create a reliable and reproducible setup for high-performance computing experiments, investigating three vastly different performance effects on modern processors. Marcos Horro, Louis-Noël Pouchet, Gabriel Rodríguez 0001, Juan Touriño |
ISPASS | 4 |
| 2022 | SparkEC: speeding up alignment-based DNA error correction toolsabstractBACKGROUND: In recent years, huge improvements have been made in the context of sequencing genomic data under what is called Next Generation Sequencing (NGS). However, the DNA reads generated by current NGS platforms are not free of errors, which can affect the quality of downstream analysis. Although error correction can be performed as a preprocessing step to overcome this issue, it usually requires long computational times to analyze those large datasets generated nowadays through NGS. Therefore, new software capable of scaling out on a cluster of nodes with high performance is of great importance. RESULTS: In this paper, we present SparkEC, a parallel tool capable of fixing those errors produced during the sequencing process. For this purpose, the algorithms proposed by the CloudEC tool, which is already proved to perform accurate corrections, have been analyzed and optimized to improve their performance by relying on the Apache Spark framework together with the introduction of other enhancements such as the usage of memory-efficient data structures and the avoidance of any input preprocessing. The experimental results have shown significant improvements in the computational times of SparkEC when compared to CloudEC for all the representative datasets and scenarios under evaluation, providing an average and maximum speedups of 4.9[Formula: see text] and 11.9[Formula: see text], respectively, over its counterpart. CONCLUSION: As error correction can take excessive computational time, SparkEC provides a scalable solution for correcting large datasets. Due to its distributed implementation, SparkEC speed can increase with respect to the number of nodes in a cluster. Furthermore, the software is freely available under GPLv3 license and is compatible with different operating systems (Linux, Windows and macOS). Roberto R. Expósito, Marco Martínez-Sánchez, Juan Touriño |
BMC Bioinform. | 3 |
| 2022 | Parallel-FST: A feature selection library for multicore clustersabstractFeature selection is a subfield of machine learning focused on reducing the dimensionality of datasets by performing a computationally intensive process. This work presents Parallel-FST, a publicly available parallel library for feature selection that includes seven methods which follow a hybrid MPI/multithreaded approach to reduce their runtime when executed on high performance computing systems. Performance tests were carried out on a 256-core cluster, where Parallel-FST obtained speedups of up to 229x for representative datasets and it was able to analyze a 512 GB dataset, which was not previously possible with a sequential counterpart library due to memory constraints. Bieito Beceiro, Jorge González-Domínguez, Juan Touriño |
J. Parallel Distributed Comput. | 3 |
| 2020 | Power Budgeting of Big Data Applications in Container-based ClustersabstractEnergy consumption is currently highly regarded on computing systems for many reasons, such as improving the environmental impact and reducing operational costs considering the rising price of energy. Previous works have analysed how to improve energy efficiency from the entire infrastructure down to individual computing instances (e.g., virtual machines). However, the research is more scarce when it comes to controlling energy consumption, specially in real time and at the software level. This paper presents a platform that manages a power budget to cap the energy consumed from users to applications and down to individual instances. Using containers as virtualization technology, the energy limitation is implemented thanks to the platform's ability to monitor container energy consumption and dynamically adjust its CPU resources via vertical scaling as required. Representative Big Data applications have been deployed on the platform to prove the feasibility of this approach for energy control, showing that it is possible to distribute and enforce a power budget among users and applications. Jonatan Enes, Guillaume Fieni, Roberto R. Expósito, Romain Rouvoy, Juan Touriño |
CLUSTER | 5 |
| 2020 | Real-time resource scaling platform for Big Data workloads on serverless environments
Jonatan Enes, Roberto R. Expósito, Juan Touriño |
Future Gener. Comput. Syst. | 3 |
| 2020 | SMusket: Spark-based DNA error correction on distributed-memory systems
Roberto R. Expósito, Jorge González-Domínguez, Juan Touriño |
Future Gener. Comput. Syst. | 3 |
| 2019 | Effect of Distributed Directories in Mesh InterconnectsabstractRecent manycore processors are kept coherent using scalable distributed directories. A paramount example is the Xeon Phi Knights Landing. It features 38 tiles packed in a single die, organized into a 2D mesh. Before accessing remote data, tiles need to query the distributed directory. The effect of this coherence traffic is poorly understood. We show that the apparent UMA behavior results from the degradation of the peak performance. We develop ways to optimize the coherence traffic, the core-to-core-affinity, and the scheduling of a set of tasks on the mesh, leveraging the unique characteristics of processor units stemming from process variations. Marcos Horro, Mahmut T. Kandemir, Louis-Noël Pouchet, Gabriel Rodríguez 0001, Juan Touriño |
DAC | 5 |
| 2019 | Parallel feature selection for distributed-memory clusters
Jorge González-Domínguez, Verónica Bolón-Canedo, Borja Freire, Juan Touriño |
Inf. Sci. | 4 |
| 2019 | Affine Modeling of Program TracesabstractA formal, high-level representation of programs is typically needed for static and dynamic analyses performed by compilers. However, the source code of target applications is not always available in an analyzable form, e.g., to protect intellectual property. To reason on such applications it becomes necessary to build models from observations of its execution. This paper presents an algebraic approach which, taking as input the trace of memory addresses accessed by a single memory reference, synthesizes an affine loop with a single perfectly nested statement that generates the original trace. This approach is extended to support the synthesis of unions of affine loops, useful for minimally modeling traces generated by automatic transformations of polyhedral programs, such as tiling. The resulting system is capable of processing hundreds of gigabytes of trace data in minutes, minimally reconstructing 100 percent of the static control parts in PolyBench/C applications and 99.9 percent in the Pluto-tiled versions of these benchmarks. Gabriel Rodríguez 0001, Mahmut T. Kandemir, Juan Touriño |
IEEE Trans. Computers | 3 |
| 2018 | BDWatchdog: Real-time monitoring and profiling of Big Data applications and frameworks
Jonatan Enes, Roberto R. Expósito, Juan Touriño |
Future Gener. Comput. Syst. | 3 |
| 2018 | BDEv 3.0: Energy efficiency and microarchitectural characterization of Big Data processing frameworks
Jorge Veiga, Jonatan Enes, Roberto R. Expósito, Juan Touriño |
Future Gener. Comput. Syst. | 4 |
| 2018 | Big Data-Oriented PaaS Architecture with Disk-as-a-Resource Capability and Container-Based Virtualization
Jonatan Enes, Javier López Cacheiro, Roberto R. Expósito, Juan Touriño |
J. Grid Comput. | 4 |
| 2018 | Enhancing in-memory efficiency for MapReduce-based data processing
Jorge Veiga, Roberto R. Expósito, Guillermo L. Taboada, Juan Touriño |
J. Parallel Distributed Comput. | 4 |
| 2017 | MarDRe: efficient MapReduce-based removal of duplicate DNA reads in the cloudabstractSUMMARY: This article presents MarDRe, a de novo cloud-ready duplicate and near-duplicate removal tool that can process single- and paired-end reads from FASTQ/FASTA datasets. MarDRe takes advantage of the widely adopted MapReduce programming model to fully exploit Big Data technologies on cloud-based infrastructures. Written in Java to maximize cross-platform compatibility, MarDRe is built upon the open-source Apache Hadoop project, the most popular distributed computing framework for scalable Big Data processing. On a 16-node cluster deployed on the Amazon EC2 cloud platform, MarDRe is up to 8.52 times faster than a representative state-of-the-art tool. AVAILABILITY AND IMPLEMENTATION: Source code in Java and Hadoop as well as a user's guide are freely available under the GNU GPLv3 license at http://mardre.des.udc.es . CONTACT: [email protected]. Roberto R. Expósito, Jorge Veiga, Jorge González-Domínguez, Juan Touriño |
Bioinform. | 4 |
| 2016 | Performance evaluation of big data frameworks for large-scale data analyticsabstractThe increasing adoption of Big Data analytics has led to a high demand for efficient technologies in order to manage and process large datasets. Popular MapReduce frameworks such as Hadoop are being replaced by emerging ones like Spark or Flink, which improve both the programming APIs and performance. However, few works have focused on comparing these frameworks. This paper addresses this issue by performing a comparative evaluation of Hadoop, Spark and Flink using representative Big Data workloads and considering factors like performance and scalability. Moreover, the behavior of these frameworks has been characterized by modifying some of the main parameters of the workloads such as HDFS block size, input data size, interconnect network or thread configuration. The analysis of the results has shown that replacing Hadoop with Spark or Flink can lead to a reduction in execution times by 77% and 70% on average, respectively, for non-sort benchmarks. Jorge Veiga, Roberto R. Expósito, Xoán C. Pardo, Guillermo L. Taboada, Juan Touriño |
IEEE BigData | 5 |
| 2016 | Trace-based affine reconstruction of codesabstractComplete comprehension of loop codes is desirable for a variety of program optimizations. Compilers perform static code analyses and transformations, such as loop tiling or memory partitioning, by constructing and manipulating formal representations of the source code. Runtime systems observe and characterize application behavior to drive resource management and allocation, including dependence detection and parallelization, or scheduling. However, the source codes of target applications are not always available to the compiler or runtime system in an analyzable form. It becomes necessary to find alternate ways to model application behavior. This paper presents a novel mathematical framework to rebuild loops from their memory access traces. An exploration engine traverses a tree-like solution space, driven by the access strides in the trace. It is guaranteed that the engine will find the minimal affine nest capable of reproducing the observed sequence of accesses by exploring this space in a brute force fashion, but most real traces will not be tractable in this way. Methods for an efficient solution space traversal based on mathematical properties of the equation systems which model the solution space are proposed. The experimental evaluation shows that these strategies achieve efficient loop reconstruction, processing hundreds of gigabytes of trace data in minutes. The proposed approach is capable of correctly and minimally reconstructing 100% of the static control parts in PolyBench/C applications. As a side effect, the trace reconstruction process can be used to efficiently compress trace files. The proposed tool can also be used for dynamic access characterization, predicting over 99% of future memory accesses. Gabriel Rodríguez 0001, José M. Andión, Mahmut T. Kandemir, Juan Touriño |
CGO | 4 |
| 2016 | MSAProbs-MPI: parallel multiple sequence aligner for distributed-memory systemsabstractMSAProbs is a state-of-the-art protein multiple sequence alignment tool based on hidden Markov models. It can achieve high alignment accuracy at the expense of relatively long runtimes for large-scale input datasets. In this work we present MSAProbs-MPI, a distributed-memory parallel version of the multithreaded MSAProbs tool that is able to reduce runtimes by exploiting the compute capabilities of common multicore CPU clusters. Our performance evaluation on a cluster with 32 nodes (each containing two Intel Haswell processors) shows reductions in execution time of over one order of magnitude for typical input datasets. Furthermore, MSAProbs-MPI using eight nodes is faster than the GPU-accelerated QuickProbs running on a Tesla K20. Another strong point is that MSAProbs-MPI can deal with large datasets for which MSAProbs and QuickProbs might fail due to time and memory constraints, respectively. AVAILABILITY AND IMPLEMENTATION: Source code in C ++ and MPI running on Linux systems as well as a reference manual are available at http://msaprobs.sourceforge.net CONTACT: [email protected] information: Supplementary data are available at Bioinformatics online. Jorge González-Domínguez, Yongchao Liu 0004, Juan Touriño, Bertil Schmidt |
Bioinform. | 3 |
| 2016 | Performance Evaluation of Data-Intensive Computing Applications on a Public IaaS CloudabstractThe advent of cloud computing technologies, which dynamically provide on-demand access to computational resources over the Internet, is offering new possibilities to many scientists and researchers. Nowadays, Infrastructure as a Service (IaaS) cloud providers can offset the increasing processing requirements of data-intensive computing applications, becoming an emerging alternative to traditional servers and clusters. In this paper, a comprehensive study of the leading public IaaS cloud platform, Amazon EC2, has been conducted in order to assess its suitability for data-intensive computing. One of the key contributions of this work is the analysis of the storage-optimized family of EC2 instances. Furthermore, this study presents a detailed analysis of both performance and cost metrics. More specifically, multiple experiments have been carried out to analyze the full I/O software stack, ranging from the low-level storage devices and cluster file systems up to real-world applications using representative data-intensive parallel codes and MapReduce-based workloads. The analysis of the experimental results has shown that data-intensive applications can benefit from tailored EC2-based virtual clusters, enabling users to obtain the highest performance and cost-effectiveness in the cloud. Roberto R. Expósito, Guillermo L. Taboada, Sabela Ramos, Juan Touriño, Ramón Doallo |
Comput. J. | 4 |
| 2016 | Flame-MR: An event-driven architecture for MapReduce applications
Jorge Veiga, Roberto R. Expósito, Guillermo L. Taboada, Juan Touriño |
Future Gener. Comput. Syst. | 4 |
| 2016 | Parallel Pairwise Epistasis Detection on Heterogeneous Computing ArchitecturesabstractDevelopment of new methods to detect pairwise epistasis, such as SNP-SNP interactions, in Genome-Wide Association Studies is an important task in bioinformatics as they can help to explain genetic influences on diseases. As these studies are time consuming operations, some tools exploit the characteristics of different hardware accelerators (such as GPUs and Xeon Phi coprocessors) to reduce the runtime. Nevertheless, all these approaches are not able to efficiently exploit the whole computational capacity of modern clusters that contain both GPUs and Xeon Phi coprocessors. In this paper we investigate approaches to map pairwise epistasic detection on heterogeneous clusters using both types of accelerators. The runtimes to analyze the well-known WTCCC dataset consisting of about 500 K SNPs and 5 K samples on one and two NVIDIA K20m are reduced by 27 percent thanks to the use of a hybrid approach with one additional single Xeon Phi coprocessor. Jorge González-Domínguez, Sabela Ramos, Juan Touriño, Bertil Schmidt |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2015 | Low-latency Java communication devices on RDMA-enabled networksabstractSummary Providing high‐performance inter‐node communication is a key capability for running high performance computing applications efficiently on parallel architectures. In fact, current systems deployments are aggregating a significant number of cores interconnected via advanced networking hardware with Remote Direct Memory Access (RDMA) mechanisms, that enable zero‐copy and kernel‐bypass features. The use of Java for parallel programming is becoming more promising thanks to some useful characteristics of this language, particularly its built‐in multithreading support, portability, easy‐to‐learn properties, and high productivity, along with the continuous increase in the performance of the Java virtual machine. However, current parallel Java applications generally suffer from inefficient communication middleware, mainly based on protocols with high communication overhead that do not take full advantage of RDMA‐enabled networks. This paper presents efficient low‐level Java communication devices that overcome these constraints by fully exploiting the underlying RDMA hardware, providing low‐latency and high‐bandwidth communications for parallel Java applications. The performance evaluation conducted on representative RDMA networks and parallel systems has shown significant point‐to‐point performance increases compared with previous Java communication middleware, allowing to obtain up to 40% improvement in application‐level performance on 4096 cores of a Cray XE6 supercomputer. Copyright © 2015 John Wiley & Sons, Ltd. Roberto R. Expósito, Guillermo L. Taboada, Sabela Ramos, Juan Touriño, Ramón Doallo |
Concurr. Comput. Pract. Exp. | 4 |
| 2015 | Nonblocking collectives for scalable Java communicationsabstractSummary This paper presents a Java implementation of the recently published MPI 3.0 nonblocking message passing collectives in order to analyze and assess the feasibility of taking advantage of these operations in shared memory systems using Java. Nonblocking collectives aim to exploit the overlapping between computation and communication for collective operations to increase scalability of message passing codes, as it has been carried out for nonblocking point‐to‐point primitives. This scalability has become crucial not only for clusters but also for shared memory systems because of the current trend of increasing the number of cores per chip, which is leading to the generalization of multi‐core and many‐core processors. Message passing libraries based on remote direct memory access, thread‐based progression, or implementing pure multi‐threading shared memory support could potentially benefit from the lack of imposed synchronization by nonblocking collectives. But, although the distributed memory scenario has been well studied, the shared memory one has not been tackled yet. Hence, nonblocking collectives support has been included in FastMPJ, a Message Passing in Java (MPJ) implementation, and evaluated on a representative shared memory system, obtaining significant improvements because of overlapping and lack of implicit synchronization, and with barely any overhead imposed over common blocking operations. Copyright © 2014 John Wiley & Sons, Ltd. Sabela Ramos, Guillermo L. Taboada, Roberto R. Expósito, Juan Touriño |
Concurr. Comput. Pract. Exp. | 4 |
| 2014 | A parallelizing compiler for multicore systemsabstractThis manuscript summarizes the main ideas introduced in [1]. We propose a compiler that automatically transforms a sequential application into a parallel counterpart for multicore processors. It is based on an intermediate representation, named KIR, which exposes multiple levels of parallelism and hides the complexity of the implementation details thanks to the domain-independent kernels (e.g., assignment, reduction). The effectiveness and performance of our approach, built on top of GCC, has been tested with a large variety of codes. José M. Andión, Manuel Arenaz, Gabriel Rodríguez 0001, Juan Touriño |
SCOPES | 4 |
| 2014 | Volatile STT-RAM Scratchpad Design and Data Allocation for Low EnergyabstractOn-chip power consumption is one of the fundamental challenges of current technology scaling. Cache memories consume a sizable part of this power, particularly due to leakage energy. STT-RAM is one of several new memory technologies that have been proposed in order to improve power while preserving performance. It features high density and low leakage, but at the expense of write energy and performance. This article explores the use of STT-RAM--based scratchpad memories that trade nonvolatility in exchange for faster and less energetically expensive accesses, making them feasible for on-chip implementation in embedded systems. A novel multiretention scratchpad partitioning is proposed, featuring multiple storage spaces with different retention, energy, and performance characteristics. A customized compiler-based allocation algorithm suitable for use with such a scratchpad organization is described. Our experiments indicate that a multiretention STT-RAM scratchpad can provide energy savings of 53% with respect to an iso-area, hardware-managed SRAM cache. Gabriel Rodríguez 0001, Juan Touriño, Mahmut T. Kandemir |
ACM Trans. Archit. Code Optim. | 2 |
| 2014 | A 2D algorithm with asymmetric workload for the UPC conjugate gradient method
Jorge González-Domínguez, Osni Marques, María J. Martín, Juan Touriño |
J. Supercomput. | 4 |
| 2013 | Design of Scalable Java Communication Middleware for Multi-Core SystemsabstractThis paper presents smdev, a shared memory communication middleware for multi-core systems. smdev provides a simple and powerful messaging application program interface that is able to exploit the underlying multi-core architecture replacing inter-process and network-based communications by threads and shared memory transfers. The performance evaluation of smdev on several multi-core systems has shown noticeable improvements compared with other Java shared memory solutions, reaching and even overcoming the performance of natively compiled libraries. Thus, smdev has obtained start-up latencies around 0.76 μs and almost 90 Gbps bandwidth for point-to-point communications, as well as high performance and scalability both for collective operations and representative messaging kernels. This fact has motivated the integration of smdev in F-MPJ, our message-passing implementation in Java. Sabela Ramos, Guillermo L. Taboada, Roberto R. Expósito, Juan Touriño, Ramón Doallo |
Comput. J. | 4 |
| 2013 | General-purpose computation on GPUs for high performance cloud computingabstractSUMMARY Cloud computing is offering new approaches for High Performance Computing (HPC) as it provides dynamically scalable resources as a service over the Internet. In addition, General‐Purpose computation on Graphical Processing Units (GPGPU) has gained much attention from scientific computing in multiple domains, thus becoming an important programming model in HPC. Compute Unified Device Architecture (CUDA) has been established as a popular programming model for GPGPUs, removing the need for using the graphics APIs for computing applications. Open Computing Language (OpenCL) is an emerging alternative not only for GPGPU but also for any parallel architecture. GPU clusters, usually programmed with a hybrid parallel paradigm mixing Message Passing Interface (MPI) with CUDA/OpenCL, are currently gaining high popularity. Therefore, cloud providers are deploying clusters with multiple GPUs per node and high‐speed network interconnects in order to make them a feasible option for HPC as a Service (HPCaaS). This paper evaluates GPGPU for high performance cloud computing on a public cloud computing infrastructure, Amazon EC2 Cluster GPU Instances (CGI), equipped with NVIDIA Tesla GPUs and a 10 Gigabit Ethernet network. The analysis of the results, obtained using up to 64 GPUs and 256‐processor cores, has shown that GPGPU is a viable option for high performance cloud computing despite the significant impact that virtualized environments still have on network overhead, which still hampers the adoption of GPGPU communication‐intensive applications. Copyright © 2012 John Wiley & Sons, Ltd. Roberto R. Expósito, Guillermo L. Taboada, Sabela Ramos, Juan Touriño, Ramón Doallo |
Concurr. Comput. Pract. Exp. | 4 |
| 2013 | Performance analysis of HPC applications in the cloud
Roberto R. Expósito, Guillermo L. Taboada, Sabela Ramos, Juan Touriño, Ramón Doallo |
Future Gener. Comput. Syst. | 4 |
| 2013 | Analysis of I/O Performance on an Amazon EC2 Cluster Compute and High I/O Platform
Roberto R. Expósito, Guillermo L. Taboada, Sabela Ramos, Jorge González-Domínguez, Juan Touriño, Ramón Doallo |
J. Grid Comput. | 5 |
| 2013 | Design and Implementation of an Extended Collectives Library for Unified Parallel C
Carlos Teijeiro, Guillermo L. Taboada, Juan Touriño, Ramón Doallo, José Carlos Mouriño, Damián A. Mallón, Brian Wibecan |
J. Comput. Sci. Technol. | 3 |
| 2013 | A novel compiler support for automatic parallelization on multicore systems
José M. Andión, Manuel Arenaz, Gabriel Rodríguez 0001, Juan Touriño |
Parallel Comput. | 4 |
| 2013 | Evaluation of messaging middleware for high-performance cloud computing
Roberto R. Expósito, Guillermo L. Taboada, Sabela Ramos, Juan Touriño, Ramón Doallo |
Pers. Ubiquitous Comput. | 4 |
| 2013 | Java in the High Performance Computing arena: Research, practice and experience
Guillermo L. Taboada, Sabela Ramos, Roberto R. Expósito, Juan Touriño, Ramón Doallo |
Sci. Comput. Program. | 4 |
| 2013 | Performance evaluation of sparse matrix products in UPC
Jorge González-Domínguez, Óscar García-López, Guillermo L. Taboada, María J. Martín, Juan Touriño |
J. Supercomput. | 5 |
| 2013 | Parallel simulation of Brownian dynamics on shared memory systems with OpenMP and Unified Parallel C
Carlos Teijeiro, Godehard Sutmann, Guillermo L. Taboada, Juan Touriño |
J. Supercomput. | 4 |
| 2012 | Design and Performance Issues of Cholesky and LU Solvers Using UPCBLASabstractPartitioned Global Address Space (PGAS) languages offer programmers a shared memory view that increases their productivity and allow locality exploitation to obtain good performance on current large-scale distributed memory systems. UPCBLAS is a parallel numerical library for dense matrix computations using the PGAS Unified Parallel C (UPC) language. The interface of this library exploits the characteristics of the PGAS memory model and thus it is easier to use than MPI-based libraries. This paper addresses the implementation of solvers of systems of equations through Cholesky and LU factorizations in UPC using UPCBLAS. The developed codes are experimentally evaluated and compared to the MPI versions using ScaLAPACK. Parallel solvers of equations are present in many parallel numerical applications and they have been traditionally developed in MPI. This work shows that UPCBLAS can be considered as a good alternative to the MPI-based libraries for increasing the productivity of numerical application developers. Jorge González-Domínguez, Osni Marques, María J. Martín, Guillermo L. Taboada, Juan Touriño |
ISPA | 5 |
| 2012 | Communication avoiding and overlapping for numerical linear algebraabstractTo efficiently scale dense linear algebra problems to future exascale systems, communication cost must be avoided or overlapped. Communication-avoiding 2.5D algorithms improve scalability by reducing inter-processor data transfer volume at the cost of extra memory usage. Communication overlap attempts to hide messaging latency by pipelining messages and overlapping with computational work. We study the interaction and compatibility of these two techniques for two matrix multiplication algorithms (Cannon and SUMMA), triangular solve, and Cholesky factorization. For each algorithm, we construct a detailed performance model that considers both critical path dependencies and idle time. We give novel implementations of 2.5D algorithms with overlap for each of these problems. Our software employs UPC, a partitioned global address space (PGAS) language that provides fast one-sided communication. We show communication avoidance and overlap provide a cumulative benefit as core counts scale, including results using over 24K cores of a Cray XE6 system. Evangelos Georganas, Jorge González-Domínguez, Edgar Solomonik, Yili Zheng, Juan Touriño, Katherine A. Yelick |
SC | 5 |
| 2012 | UPCBLAS: a library for parallel matrix computations in Unified Parallel CabstractSUMMARY The popularity of Partitioned Global Address Space (PGAS) languages has increased during the last years thanks to their high programmability and performance through an efficient exploitation of data locality, especially on hierarchical architectures such as multicore clusters. This paper describes UPCBLAS, a parallel numerical library for dense matrix computations using the PGAS Unified Parallel C language. The routines developed in UPCBLAS are built on top of sequential basic linear algebra subprograms functions and exploit the particularities of the PGAS paradigm, taking into account data locality in order to achieve a good performance. Furthermore, the routines implement other optimization techniques, several of them by automatically taking into account the hardware characteristics of the underlying systems on which they are executed. The library has been experimentally evaluated on a multicore supercomputer and compared with a message‐passing‐based parallel numerical library, demonstrating good scalability and efficiency. Copyright © 2012 John Wiley & Sons, Ltd. Jorge González-Domínguez, María J. Martín, Guillermo L. Taboada, Juan Touriño, Ramón Doallo, Damián A. Mallón, Brian Wibecan |
Concurr. Comput. Pract. Exp. | 4 |
| 2012 | Design of scalable Java message-passing communications over InfiniBand
Roberto R. Expósito, Guillermo L. Taboada, Juan Touriño, Ramón Doallo |
J. Supercomput. | 3 |
| 2012 | F-MPJ: scalable Java message-passing communications on parallel systems
Guillermo L. Taboada, Juan Touriño, Ramón Doallo |
J. Supercomput. | 2 |
| 2011 | Scalable Java Communication Middleware for Hybrid Shared/Distributed Memory ArchitecturesabstractThe up trend in the number of cores in cluster architectures underscores the need for scalable communication middleware on these systems. One of the strategies to take advantage of this increase in the available computational power is the use of efficient message-passing middleware for inter-node communications and thread-based shared memory transfers within each node. This paper presents a Java communication middleware that exploits hybrid shared/distributed memory architectures through the use of scalable Java NIO sockets for inter-node communications and multi-threading on shared memory. Thus, communication-intensive applications running on clusters of multi-core processors can take advantage of the use of this middleware. The performance of these codes generally relies on collective operations, such as broadcasting, scattering or gathering data, which have been optimized to make the most of these architectures. The evaluation of this middleware when relying on multi-core aware communication patterns has shown significant performance improvements both in collective operations and communication-intensive applications. Sabela Ramos, Guillermo L. Taboada, Juan Touriño, Ramón Doallo |
HPCC | 3 |
| 2011 | Design and Implementation of MapReduce Using the PGAS Programming Model with UPCabstractMapReduce is a powerful tool for processing large data sets used by many applications running in distributed environments. However, despite the increasing number of computationally intensive problems that require low-latency communications, the adoption of MapReduce in High Performance Computing (HPC) is still emerging. Here languages based on the Partitioned Global Address Space (PGAS) programming model have shown to be a good choice for implementing parallel applications, in order to take advantage of the increasing number of cores per node and the programmability benefits achieved by their global memory view, such as the transparent access to remote data. This paper presents the first PGAS-based MapReduce implementation that uses the Unified Parallel C (UPC) language, which (1) obtains programmability benefits in parallel programming, (2) offers advanced configuration options to define a customized load distribution for different codes, and (3) overcomes performance penalties and bottlenecks that have traditionally prevented the deployment of MapReduce applications in HPC. The performance evaluation of representative applications on shared and distributed memory environments assesses the scalability of the presented MapReduce framework, confirming its suitability. Carlos Teijeiro, Guillermo L. Taboada, Juan Touriño, Ramón Doallo |
ICPADS | 3 |
| 2011 | Extending the Globus Information Service with the Common Information ModelabstractThe need of task-adapted and complete information for the management of resources is a well known issue in Grid computing. Globus Toolkit 4 (GT4) includes the Monitoring and Discovery System component (MDS4) to carry out resource management. The Common Information Model (CIM) provides a standard conceptual view of the managed environment. This work improves the MDS4 functionality through the use of CIM, with the aim of providing a unified, standard representation of the Grid resources. Since a practical CIM model may contain a large volume of information, a new Index Service that represents the CIM information through Java instances is presented. In addition, a solution that keeps data in persistent storage has also been implemented. The evaluation of the proposed solutions achieves encouraging results, with an important reduction in memory consumption, a good scalability when the number of instances increases, and with a reasonable response time. Iván Díaz, Gracia Fernández, Patricia González, María J. Martín, Juan Touriño |
ISPA | 5 |
| 2011 | Analysis of Performance-impacting Factors on Checkpointing Frameworks: The CPPC Case StudyabstractThis paper focuses on the performance evaluation of Compiler for Portable Checkpointing (CPPC), a tool for the checkpointing of parallel message-passing applications. Its performance and the factors that impact it are transparently and rigorously identified and assessed. The tests were performed on a public supercomputing infrastructure, using a large number of very different applications and showing excellent results in terms of performance and effort required for integration into user codes. Statistical analysis techniques have been used to better approximate the performance of the tool. Quantitative and qualitative comparisons with other rollback-recovery approaches to fault tolerance are also included. All these data and comparisons are then discussed in an effort to extract meaningful conclusions about the state-of-the-art and future research trends in the rollback-recovery field. Gabriel Rodríguez 0001, María J. Martín, Patricia González, Juan Touriño |
Comput. J. | 4 |
| 2011 | Device level communication libraries for high-performance computing in JavaabstractSUMMARY Since its release, the Java programming language has attracted considerable attention from the high‐performance computing (HPC) community because of its portability, high programming productivity, and built‐in multithreading and networking support. As a consequence, several initiatives have been taken to develop a high‐performance Java message‐passing library to program distributed memory architectures, such as clusters. The performance of Java message‐passing applications relies heavily on the communications performance. Thus, the design and implementation of low‐level communication devices that support message‐passing libraries is an important research issue in Java for HPC. MPJ Express is our Java message‐passing implementation for developing high‐performance parallel Java applications. Its public release currently contains three communication devices: the first one is built using the Java New Input/Output (NIO) package for the TCP/IP; the second one is specifically designed for the Myrinet Express library on Myrinet; and the third one supports thread‐based shared memory communications. Although these devices have been successfully deployed in many production environments, previous performance evaluations of MPJ Express suggest that the buffering layer, tightly coupled with these devices, incurs a certain degree of copying overhead, which represents one of the main performance penalties. This paper presents a more efficient Java message‐passing communications device, based on Java Input/Output sockets, that avoids this buffering overhead. Moreover, this device implements several strategies, both in the communication protocol and in the HPC hardware support, which optimizes Java message‐passing communications. In order to evaluate its benefits, this paper analyzes the performance of this device comparatively with other Java and native message‐passing libraries on various high‐speed networks, such as Gigabit Ethernet, Scalable Coherent Interface, Myrinet, and InfiniBand, as well as on a shared memory multicore scenario. The reported communication overhead reduction encourages the upcoming incorporation of this device in MPJ Express ( http://mpj‐express.org ). Copyright © 2011 John Wiley & Sons, Ltd. Guillermo L. Taboada, Juan Touriño, Ramón Doallo, Aamir Shafi, Mark Baker, Bryan Carpenter |
Concurr. Comput. Pract. Exp. | 2 |
| 2011 | Special issue on "Theory and practice of high-performance computing, communications, and security"
Tai-Hoon Kim, Omer F. Rana, Juan Touriño, Isaac Woungang |
J. Supercomput. | 3 |
| 2011 | Design of efficient Java message-passing collectives on multi-core clusters
Guillermo L. Taboada, Sabela Ramos, Juan Touriño, Ramón Doallo |
J. Supercomput. | 3 |
| 2010 | Servet: A benchmark suite for autotuning on multicore clustersabstractThe growing complexity in computer system hierarchies due to the increase in the number of cores per processor, levels of cache (some of them shared) and the number of processors per node, as well as the high-speed interconnects, demands the use of new optimization techniques and libraries that take advantage of their features. In this paper Servet, a suite of benchmarks focused on detecting a set of parameters with high influence in the overall performance of multicore systems, is presented. These benchmarks are able to detect the cache hierarchy, including their size and which caches are shared by each core, bandwidths and bottlenecks in memory accesses, as well as communication latencies among cores. These parameters can be used by auto-tuned codes to increase their performance in multicore clusters. Experimental results using different representative systems show that Servet provides very accurate estimates of the parameters of the machine architecture. Jorge González-Domínguez, Guillermo L. Taboada, Basilio B. Fraguela, María J. Martín, Juan Touriño |
IPDPS | 5 |
| 2010 | CPPC: a compiler-assisted tool for portable checkpointing of message-passing applicationsabstractAbstract With the evolution of high‐performance computing toward heterogeneous, massively parallel systems, parallel applications have developed new checkpoint and restart necessities. Whether due to a failure in the execution or to a migration of the application processes to different machines, checkpointing tools must be able to operate in heterogeneous environments. However, some of the data manipulated by a parallel application are not truly portable. Examples of these include opaque state (e.g. data structures for communications support) or diversity of interfaces for a single feature (e.g. communications, I/O). Directly manipulating the underlyingad hocrepresentations renders checkpointing tools unable to work on different environments. Portable checkpointers usually work around portability issues at the cost of transparency: the user must provide information such as what data need to be stored, where to store them, or where to checkpoint. CPPC (ComPiler for Portable Checkpointing) is a checkpointing tool designed to feature both portability and transparency. It is made up of a library and a compiler. The CPPC library contains routines for variable level checkpointing, using portable code and protocols. The CPPC compiler helps to achieve transparency by relieving the user from time‐consuming tasks, such as data flow and communications analyses and adding instrumentation code. This paper covers both the operation of the CPPC library and its compiler support. Experimental results using benchmarks and large‐scale real applications are included, demonstrating usability, efficiency, and portability. Copyright © 2009 John Wiley & Sons, Ltd. Gabriel Rodríguez 0001, María J. Martín, Patricia González, Juan Touriño, Ramón Doallo |
Concurr. Comput. Pract. Exp. | 4 |
| 2009 | A Parallel Numerical Library for UPC
Jorge González-Domínguez, María J. Martín, Guillermo L. Taboada, Juan Touriño, Ramón Doallo, Andrés Gómez 0002 |
Euro-Par | 4 |
| 2009 | Efficient Java Communication Libraries over InfiniBandabstractThis paper presents our current research efforts on efficient Java communication libraries over InfiniBand. The use of Java for network communications still delivers insufficient performance and does not exploit the performance and other special capabilities (RDMA and QoS) of high-speed networks, especially for this interconnect. In order to increase its Java communication performance, InfiniBand has been supported in our high performance sockets implementation, Java Fast Sockets (JFS), and it has been greatly improved the efficiency of Java Direct InfiniBand (Jdib), our low-level communication layer, enabling zero-copy RDMA capability in Java. According to our experimental results, Java communication performance has been improved significantly, reducing start-up latencies from 34 mus down to 12 and 7 mus for JFS and Jdib, respectively, whereas peak bandwidth has been increased from 0.78 Gbps sending serialized data up to 6.7 and 11.2 Gbps for JFS and Jdib, respectively. Finally, it has been analyzed the impact of these communication improvements on parallel Java applications, obtaining significant speedup increases of up to one order of magnitude on 128 cores. Guillermo L. Taboada, Juan Touriño, Ramón Doallo, Jizhong Han |
HPCC | 2 |
| 2009 | Performance Evaluation of Unified Parallel C Collective CommunicationsabstractUnified Parallel C (UPC) is an extension of ANSI C designed for parallel programming. UPC collective primitives, which are part of the UPC standard, increase programming productivity while reducing the communication overhead. This paper presents an up-to-date performance evaluation of two publicly available UPC collective implementations on three scenarios: shared, distributed, and hybrid shared/distributed memory architectures. The characterization of the throughput of collective primitives is useful for increasing performance through the runtime selection of the appropriate primitive implementation, which depends on the message size and the memory architecture, as well as to detect inefficient implementations. In fact, based on the analysis of the UPC collectives performance, we proposed some optimizations for the current UPC collective libraries. We have also compared the performance of the UPC collective primitives and their MPI counterparts, showing that there is room for improvement. Finally, this paper concludes with an analysis of the influence of the performance of the UPC collectives on a representative communication-intensive application, showing that their optimization is highly important for UPC scalability. Guillermo L. Taboada, Carlos Teijeiro, Juan Touriño, Basilio B. Fraguela, Ramón Doallo, José Carlos Mouriño, Damián A. Mallón, Andrés Gómez 0002 |
HPCC | 3 |
| 2009 | NPB-MPJ: NAS Parallel Benchmarks Implementation for Message-Passing in JavaabstractJava is a valuable and emerging alternative for the development of parallel applications, thanks to the availability of several Java message-passing libraries and its full multithreading support. The combination of both shared and distributed memory programming is an interesting option for parallel programming multi-core systems. However, the concerns about Java performance are hindering its adoption in this field, although it is difficult to evaluate accurately its performance due to the lack of standard benchmarks in Java. This paper presents NPB-MPJ, the first extensive implementation of the NAS Parallel Benchmarks (NPB), the standard parallel benchmark suite, for Message-Passing in Java (MPJ) libraries. Together with the design and implementation details of NPB-MPJ, this paper gathers several optimization techniques that can serve as a guide for the development of more efficient Java applications for High Performance Computing (HPC). NPB-MPJ has been used in the performance evaluation of Java against C/Fortran parallel libraries on two representative multi-core clusters. Thus, NPB-MPJ provides an up-to-date snapshot of MPJ performance, whose comparative analysis of current Java and native parallel solutions confirms that MPJ is an alternative for parallel programming multi-core systems. Damián A. Mallón, Guillermo L. Taboada, Juan Touriño, Ramón Doallo |
PDP | 3 |
| 2008 | Efficiently Building the Gated Single Assignment Form in Codes with Pointers in Modern Optimizing Compilers
Manuel Arenaz, Pedro Amoedo, Juan Touriño |
Euro-Par | 3 |
| 2008 | Topic 6: Grid and Cluster Computing
Marco Danelutto, Juan Touriño, Mark Baker, Rajkumar Buyya, Paraskevi Fragopoulou, Christian Pérez, Erich Schikuta |
Euro-Par | 2 |
| 2008 | Java Fast Sockets: Enabling high-speed Java communications on high performance clusters
Guillermo L. Taboada, Juan Touriño, Ramón Doallo |
Comput. Commun. | 2 |
| 2008 | XARK: An extensible framework for automatic recognition of computational kernelsabstractThe recognition of program constructs that are frequently used by software developers is a powerful mechanism for optimizing and parallelizing compilers to improve the performance of the object code. The development of techniques for automatic recognition of computational kernels such as inductions, reductions and array recurrences has been an intensive research area in the scope of compiler technology during the 90's. This article presents a new compiler framework that, unlike previous techniques that focus on specific and isolated kernels, recognizes a comprehensive collection of computational kernels that appear frequently in full-scale real applications. The XARK compiler operates on top of the Gated Single Assignment (GSA) form of a high-level intermediate representation (IR) of the source code. Recognition is carried out through a demand-driven analysis of this high-level IR at two different levels. First, the dependences between the statements that compose the strongly connected components (SCCs) of the data-dependence graph of the GSA form are analyzed. As a result of this intra-SCC analysis, the computational kernels corresponding to the execution of the statements of the SCCs are recognized. Second, the dependences between statements of different SCCs are examined in order to recognize more complex kernels that result from combining simpler kernels in the same code. Overall, the XARK compiler builds a hierarchical representation of the source code as kernels and dependence relationships between those kernels. This article describes in detail the collection of computational kernels recognized by the XARK compiler. Besides, the internals of the recognition algorithms are presented. The design of the algorithms enables to extend the recognition capabilities of XARK to cope with new kernels, and provides an advanced symbolic analysis framework to run other compiler techniques on demand. Finally, extensive experiments showing the effectiveness of XARK for a collection of benchmarks from different application domains are presented. In particular, the SparsKit-II library for the manipulation of sparse matrices, the Perfect benchmarks, the SPEC CPU2000 collection and the PLTMG package for solving elliptic partial differential equations are analyzed in detail. Manuel Arenaz, Juan Touriño, Ramón Doallo |
ACM Trans. Program. Lang. Syst. | 2 |
| 2007 | Towards Low-Latency Model-Oriented Distributed Systems Management
Iván Díaz, Juan Touriño, Ramón Doallo |
APNOMS | 2 |
| 2007 | Program Behavior Characterization Through Advanced Kernel Recognition
Manuel Arenaz, Juan Touriño, Ramón Doallo |
Euro-Par | 2 |
| 2007 | High Performance Java Sockets for Parallel Computing on ClustersabstractThe use of Java for parallel programming on clusters relies on the need of efficient communication middleware and high-speed cluster interconnect support. Nevertheless, currently there are no solutions that fully fulfill these issues. In this paper, a Java sockets library has been tailored to increase the efficiency of Java parallel applications on clusters. This library supports high-speed cluster interconnects and its API has been extended to meet the requirements of a high performance Java RMI implementation and Java parallel applications on clusters. Thus, it provides Java with a more efficient communication middleware on clusters. The performance evaluation of this middleware on a Gigabit Ethernet (GbE) and a scalable coherent interface (SCI) cluster has shown experimental evidence of throughput increase. Moreover, qualitative aspects of the solution such as transparency to the user, interoperability with other systems and no need of source code modifications are decisive to boost the performance of existing Java parallel applications and their developments in high performance Java cluster computing. Guillermo L. Taboada, Juan Touriño, Ramón Doallo |
IPDPS | 2 |
| 2007 | High Performance Java Remote Method Invocation for Parallel Computing on ClustersabstractThis paper presents a more efficient Java remote method invocation (RMI) implementation for high-speed clusters. The use of Java for parallel programming on clusters is limited by the lack of efficient communication middleware and high-speed cluster interconnect support. This implementation overcomes these limitations through a more efficient Java RMI protocol based on several basic assumptions on clusters. Moreover, the use of a high performance sockets library provides with direct high-speed interconnect support. The performance evaluation of this middleware on a gigabit Ethernet (GbE) and a scalable coherent interface (SCI) cluster shows experimental evidence of throughput increase. Moreover, qualitative aspects of the solution such as transparency to the user, interoperability with other systems and no need of source code modification can augment the performance of existing parallel Java codes and boost the development of new high performance Java RMI applications. Guillermo L. Taboada, Carlos Teijeiro, Juan Touriño |
ISCC | 3 |
| 2007 | Automated and accurate cache behavior analysis for codes with irregular access patternsabstractAbstract The memory hierarchy plays an essential role in the performance of current computers, so good analysis tools that help in predicting and understanding its behavior are required. Analytical modeling is the ideal base for such tools if its traditional limitations in accuracy and scope of application can be overcome. While there has been extensive research on the modeling of codes with regular access patterns, less attention has been paid to codes with irregular patterns due to the increased difficulty in analyzing them. Nevertheless, many important applications exhibit this kind of pattern, and their lack of locality make them more cache‐demanding, which makes their study more relevant. The focus of this paper is the automation of the Probabilistic Miss Equations (PME) model, an analytical model of the cache behavior that provides fast and accurate predictions for codes with irregular access patterns. The information requirements of the PME model are defined and its integration in the XARK compiler, a research compiler oriented to automatic kernel recognition in scientific codes, is described. We show how to exploit the powerful information‐gathering capabilities provided by this compiler to allow the automated modeling of loop‐oriented scientific codes. Experimental results that validate the correctness of the automated PME model are also presented. Copyright © 2007 John Wiley & Sons, Ltd. Diego Andrade, Manuel Arenaz, Basilio B. Fraguela, Juan Touriño, Ramón Doallo |
Concurr. Comput. Pract. Exp. | 4 |
| 2007 | Special Issue: Current Trends in Compilers for Parallel ComputersabstractThis special issue of Concurrency and Computation: Practice and Experience contains a selection of the papers presented at the 12th International Workshop on Compilers for Parallel Computers (CPC'2006), held in A Coruña, Spain, 9-11 January 2006.The CPC Workshop series is well established as an invitational workshop for leading research groups in the field (mainly from Europe, North America and Asia-Pacific) to provide a forum for exchanging and developing new ideas in compiler design for parallel systems and related topics.The Workshop series began in 1989 in Oxford, U.K., and continued every 18 months in a European city: Paris, Juan Touriño, Basilio B. Fraguela, Ramón Doallo, Manuel Arenaz |
Concurr. Comput. Pract. Exp. | 1 |
| 2006 | Efficient Java Communication Protocols on High-speed Cluster InterconnectsabstractThis paper presents communication strategies for achieving efficient parallel and distributed Java applications on clusters with high-speed interconnects. Communication performance is critical for the overall cluster performance. Previous efforts at obtaining efficient Java communications have a limited applicability on high-speed interconnects as they are focused on high level APIs like RMI, ignoring the particularities of these systems and their native high performance communication protocols. By relying on a custom Java socket implementation higher degrees of performance can be achieved exploiting high-speed interconnect facilities. Several protocol definitions are presented, looking for obtaining high performance Java communications. Moreover, the quality of the protocol implementations and their design decisions has been thoroughly evaluated on a scalable coherent interface (SCI) and gigabit Ethernet (GbE) testbed cluster. The results of this analysis have demonstrated that these Java protocols obtain similar results to native communications Guillermo L. Taboada, Juan Touriño, Ramón Doallo |
LCN | 2 |
| 2004 | Compiler Support for Parallel Code Generation through Kernel RecognitionabstractSummary form only given. The automatic parallelization of loops that contain complex computations is still a challenge for current parallelizing compilers. The main limitations are related to the analysis of expressions that contain subscripted subscripts, and the analysis of conditional statements that introduce complex control flows at run-time. We use the term complex loop to designate loops with such characteristics. We describe the parallelization of sequential complex loop nests using a generic compiler framework (proposed in an earlier paper [Arenaz et al., ICS'2003] ) that accomplishes kernel recognition through the analysis of the gated single assignment program representation. Specifically, we focus on an extension of this framework that enables its use as a powerful tool for gathering source code information that is relevant for the parallelization of each computational kernel. A set of example codes are analyzed in detail to illustrate the potential of our approach. Experimental results using a benchmark suite of complex loop nests are also presented. Manuel Arenaz, Juan Touriño, Ramón Doallo |
IPDPS | 2 |
| 2004 | An Inspector-Executor Algorithm for Irregular Assignment Parallelization
Manuel Arenaz, Juan Touriño, Ramón Doallo |
ISPA | 2 |
| 2004 | A middleware architecture for distributed systems management
Jesús Salceda, Iván Díaz, Juan Touriño, Ramón Doallo |
J. Parallel Distributed Comput. | 3 |
| 2004 | A compiler tool to predict memory hierarchy performance of scientific codes
Basilio B. Fraguela, Ramón Doallo, Juan Touriño, Emilio L. Zapata |
Parallel Comput. | 3 |
| 2003 | Performance Analysis of Java Message-Passing Libraries on Fast Ethernet, Myrinet and SCI ClustersabstractThe use of Java for parallel programming on clusters according to the message-passing paradigm is an attractive choice. In this case, the overall application performance will largely depend on the performance of the underlying Java message-passing library. This paper evaluates, models and compares the performance of MPI-like point-to-point and collective communication primitives from selected Java message-passing implementations on clusters with different interconnection networks. We have developed our own micro-benchmark suite to characterize the message-passing communication overhead and thus derive analytical latency models. Guillermo L. Taboada, Juan Touriño, Ramón Doallo |
CLUSTER | 2 |
| 2003 | A GSA-based compiler infrastructure to extract parallelism from complex loopsabstractThis paper presents a new approach for the detection of coarse-grain parallelism in loop nests that contain complex computations, including subscripted subscripts as well as conditional statements that introduce complex control flows at run-time. The approach is based on the recognition of the computational kernels calculated in a loop without considering the semantics of the code. The detection is carried out on top of the Gated Single Assignment (GSA) program representation at two different levels. First, the use-def chains between the statements that compose the strongly connected components (SCCs) of the GSA use-def chain graph are analyzed (intra-SCC analysis). As a result, the kernel computed in each SCC is recognized. Second, the use-def chains between statements of different SCCs are examined (inter-SCC analysis). This second abstraction level enables the detection of more complex computational kernels by the compiler. A prototype was implemented using the infrastructure provided by the Polaris compiler. Experimental results that show the effectiveness of our approach for the detection of coarse-grain parallelism in a suite of real codes are presented. Manuel Arenaz, Juan Touriño, Ramón Doallo |
ICS | 2 |
| 2003 | Research Article: A GIS-embedded system to support land consolidation plans in GaliciaabstractLand consolidation is a strategic instrument for rural planning and thus economic development in the Spanish region of Galicia. This paper describes an experimental system embedded in a GIS environment to aid rural engineers to develop land consolidation plans. The system supports all the stages of the plan and many functionalities are implemented as heuristic processes based on expert knowledge and advice. The overall aim is to overcome administrative and technical problems of traditional consolidation procedures. The system provides an integrated framework for the management of spatial and administrative consolidation information. It also includes optimization-based algorithms for the automated generation of multiple alternative parcel reallocations, as well as an environment to refine and objectively evaluate the proposed solutions. These key capabilities result in a powerful tool for decision making that dramatically reduces the time and cost of land consolidation plans. Pilot experiences in two consolidation zones of Galicia assess the feasibility and effectiveness of the system. Juan Touriño, Jorge Parapar, Ramón Doallo, Marcos Boullón-Magán, Francisco F. Rivera, Javier D. Bruguera, Xesús P. González, Rafael Crecente-Maseda |
Int. J. Geogr. Inf. Sci. | 1 |
| 2002 | Towards Detection of Coarse-Grain Loop-Level Parallelism in Irregular Computations
Manuel Arenaz, Juan Touriño, Ramón Doallo |
Euro-Par | 2 |
| 2002 | Improving Locality in the Parallelization of Doacross Loops (Research Note)
María J. Martín, David E. Singh, Juan Touriño, Francisco F. Rivera |
Euro-Par | 3 |
| 2002 | Exploiting Locality in the Run-Time Parallelization of Irregular LoopsabstractThe goal of this work is the efficient parallel execution of loops with indirect array accesses, in order to be embedded in a parallelizing compiler framework. In this kind of loop pattern, dependences can not always be determined at compile-time as, in many cases, they involve input data that are only known at run-time and/or the access pattern is too complex to be analyzed In this paper we propose runtime strategies for the parallelization of these loops. Our approaches focus not only on extracting parallelism among iterations of the loop, but also on exploiting data access locality to improve memory hierarchy behavior and, thus, the overall program speedup. Two strategies are proposed one based on graph partitioning techniques and other based on a block-cyclic distribution. Experimental results show that both strategies are complementary and the choice of the best alternative depends on some features of the loop pattern. María J. Martín, David E. Singh, Juan Touriño, Francisco F. Rivera |
ICPP | 3 |
| 2001 | Characterization of Message-Passing Overhead on the AP3000 MulticomputerabstractThe performance of the communication primitives of parallel computers is critical for the overall system performance. The characterization of the communication overhead is very important to estimate the global performance of parallel applications and to detect possible bottlenecks. In this paper, we evaluate, model and compare the performance of the message-passing libraries provided by the Fujitsu AP3000 multicomputer: MPI/AP, PVM/AP and APlib. Our aim is to fairly characterize the communication primitives using general models and performance metrics. Juan Touriño, Ramón Doallo |
ICPP | 1 |
| 2001 | Efficient parallel numerical solver for the elastohydrodynamic Reynolds-Hertz problem
Manuel Arenaz, Ramón Doallo, Juan Touriño, Carlos Vázquez 0002 |
Parallel Comput. | 3 |
| 1999 | Performance Evaluation and Modeling of the Fujitsu AP3000 Message-Passing Libraries
Juan Touriño, Ramón Doallo |
Euro-Par | 1 |