Juan Carlos Pichel

dblp:77/5822 · DBLP profile ↗
← Back
38ranked-venue papers
8as first author
11since 2021 · last 2026
0000-0001-9505-6493ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 26 · 7 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 5Artificial intelligence and machine learning · 4 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Unlocking Python Multithreading Capabilities using OpenMP-Based Programming with OMP4Py
abstract
Python exhibits inferior performance relative to traditional high performance computing (HPC) languages such as C, C++, and Fortran. This performance gap is largely due to Python’s interpreted nature and the Global Interpreter Lock (GIL), which restricts multithreading efficiency. However, the introduction of a GIL-free variant in the Python interpreter opens the door to more effective exploitation of multithreading parallelism in Python. Based on this important new feature, we introduce OMP4Py with the aim of bringing OpenMP’s familiar directive-based parallelization paradigm to Python. Its dual-runtime architecture design combines the benefits of a pure Python implementation with the performance and low-level capabilities required to maximize efficiency in compute-intensive tasks. In this way, OMP4Py offers both full Python support and the high performance required by HPC workloads.
César Piñeiro, Juan Carlos Pichel
CGO2
2026 NetQIR: An extension of QIR for distributed quantum computing
abstract
The rapid advancement of quantum computing has highlighted the need for scalable and efficient software infrastructures to fully exploit its potential. Current quantum processors face significant scalability constraints due to the limited number of qubits per chip. In response, distributed quantum computing (DQC) —achieved by networking multiple quantum processor units (QPUs)— is emerging as a promising solution. To support this paradigm, robust intermediate representations (IRs) are needed to translate high-level quantum algorithms into executable instructions suitable for distributed systems. This paper presents NetQIR, an extension of Microsoft’s Quantum Intermediate Representation (QIR), specifically designed to facilitate DQC by incorporating new instruction specifications. NetQIR was developed in response to the lack of abstraction at the network and hardware layers identified in the existing literature as a significant obstacle to effectively implementing distributed quantum algorithms. Based on this analysis, NetQIR introduces new essential abstraction features to support compilers in DQC contexts. It defines network communication instructions independent of specific hardware, abstracting the complexities of inter-QPU communication. Although the proposed work allows abstraction of the underlying network, it is important to note that it is intended for the development of high-performance code on future modular quantum architectures. Leveraging the QIR framework, NetQIR aims to bridge the gap between high-level quantum algorithm design and low-level hardware execution, thus promoting modular and scalable approaches to quantum software infrastructures for distributed applications. Furthermore, its design may serve as a foundational component for future implementations of distributed quantum standards such as the Quantum Message Passing Interface (QMPI).
Francisco Javier Cardama, Jorge Vázquez-Pérez, César Piñeiro, Tomás F. Pena, Juan Carlos Pichel, Andrés Gómez 0002
Future Gener. Comput. Syst.5
2026 OMP4Py: A pure Python implementation of openMP
abstract
Python demonstrates lower performance in comparison to traditional high performance computing (HPC) languages such as C, C++, and Fortran. This performance gap is largely due to Python’s interpreted nature and the Global Interpreter Lock (GIL), which hampers multithreading efficiency. However, the latest version of Python includes the necessary changes to make the interpreter thread-safe, allowing Python code to run without the GIL. This important update will enable users to fully exploit multithreading parallelism in Python. In order to facilitate that task, this paper introduces OMP4Py, the first pure Python implementation of OpenMP. We demonstrate that it is possible to bring OpenMP’s familiar directive-based parallelization paradigm to Python, allowing developers to write parallel code with the same level of control and flexibility as in C, C++, or Fortran. The experimental evaluation shows that OMP4Py significantly impacts the performance of various types of applications, although the current threading limitations of Python’s interpreter (v3.13) reduce its effectiveness for numerical applications. • OMP4Py: First pure Python implementation of OpenMP for parallel programming. • Brings OpenMP’s directive-based parallelism to Python. • Combines with mpi4py for hybrid parallelism on clusters. • Supports broader applications than PyOMP, which is limited to numerical tasks.
César Piñeiro, Juan Carlos Pichel
Future Gener. Comput. Syst.2
2026 Reconstruction of phylogenetic trees via graph-splitting using quantum computing
abstract
Abstract Quantum computing applies principles of quantum mechanics, such as superposition and entanglement, to process information with exponential parallelism. This paradigm offers significant computational advantages over classical methods, particularly for NP-hard problems like phylogenetic tree reconstruction in evolutionary biology. Phylogenetic trees model the evolutionary relationships among species or genes, and their reconstruction is computationally challenging as the number of possible topologies grows exponentially with the number of taxa. To address this, biologists often rely on heuristic methods; however, recent work has shown that recursive graph-cut techniques can achieve high accuracy in phylogenetic inference, though at high computational cost. In this study, we present a quantum algorithm based on the normalized cut ( $$N_{\text {cut}}$$ N cut ) criterion, enabling efficient recursive graph partitioning. Implemented using Quantum Annealing (QA) and the Quantum Approximate Optimization Algorithm (QAOA), demonstrating promising results on real quantum hardware for complex bioinformatics tasks.
Nicolás Fernández-Otero, Tomás F. Pena, Juan Carlos Pichel
J. Supercomput.3
2025 Review of intermediate representations for quantum computing
abstract
Abstract Intermediate representations (IRs) are fundamental to classical and quantum computing, bridging high-level quantum programming languages and the hardware-specific instructions required for execution. This paper reviews the development of quantum IRs, focusing on their evolution and the need for abstraction layers that facilitate portability and optimization. Monolithic quantum IRs, such as QIR (Lubinski et al. in Front Phys 10:940293, 2022. https://doi.org/10.3389/fphy.2022.940293), QSSA (Peduri et al. in Proceedings of the 31st ACM SIGPLAN international conference on compiler construction. CC 2022. Association for Computing Machinery, New York, 2022), or Q-MLIR (McCaskey and Nguyen in Proceedings-2021 IEEE International Conference on Quantum Computing and Engineering, QCE, 2021), their effectiveness in handling abstractions, and their hybrid support between quantum-classical operations are evaluated. However, a key limitation is their inability to address qubit locality, an essential feature for distributed quantum computing (DQC). To overcome this, InQuIR (Nishio and Wakizaka in InQuIR: Intermediate Representation for Interconnected Quantum Computers, 2023. https://arxiv.org/abs/2302.00267) was introduced as an IR specifically designed for distributed systems, providing explicit control over qubit locality and inter-node communication. While effective in managing qubit distribution, InQuIR’s dependence on manual manipulation of communication protocols increases complexity for developers. NetQIR (Vázquez-Pérez et al. in NetQIR: An Extension of QIR for Distributed Quantum Computing, 2024. https://arxiv.org/abs/2408.03712), an extension of QIR for DQC, emerges as a solution to achieve the abstraction of quantum communications protocols. This review emphasizes the need for further advancements in IRs for distributed quantum systems, which will play a crucial role in the scalability and usability of future quantum networks.
Francisco Javier Cardama, Jorge Vázquez-Pérez, César Piñeiro, Juan Carlos Pichel, Tomás F. Pena, Andrés Gómez 0002
J. Supercomput.4
2025 Inqasm: InQuIR compiler to NetQASM
abstract
Abstract Quantum computing is a rapidly evolving field, with almost every aspect open to change or improvement. This includes moving from using a single quantum processing unit to interconnecting multiple quantum processing units (or several of them), establishing a new paradigm called distributed quantum computing and increasing the overall computing capability. Some research is already underway in this area to prepare the ground for an eventual architecture with these characteristics. This is the case of InQuIR (Nishio and Wakizaka in arXiv:2302.00267 2023) and NetQASM (Dahlberg et al in QST 7:035023 2022), two languages developed for distributed quantum computing. This paper presents the development of the InQASM compiler with the aim of translating code from the InQuIR language to NetQASM, establishing a compilation stack for the new distributed paradigm. An example of this compilation and a simulation of the compiled code are shown to showcase it.
Jorge Vázquez-Pérez, Francisco Javier Cardama, César Piñeiro, Juan Carlos Pichel, Tomás F. Pena, Andrés Gómez 0002
J. Supercomput.4
2024 An unsupervised perplexity-based method for boilerplate removal
abstract
Abstract The availability of large web-based corpora has led to significant advances in a wide range of technologies, including massive retrieval systems or deep neural networks. However, leveraging this data is challenging, since web content is plagued by the so-called boilerplate: ads, incomplete or noisy text and rests of the navigation structure, such as menus or navigation bars. In this work, we present a novel and efficient approach to extract useful and well-formed content from web-scraped data. Our approach takes advantage of Language Models and their implicit knowledge about correctly formed text, and we demonstrate here that perplexity is a valuable artefact that can contribute in terms of effectiveness and efficiency. As a matter of fact, the removal of noisy parts leads to lighter AI or search solutions that are effective and entail important reductions in resources spent. We exemplify here the usefulness of our method with two downstream tasks, search and classification, and a cleaning task. We also provide a Python package with pre-trained models and a web demo demonstrating the capabilities of our approach.
Marcos Fernández-Pichel, Manuel de Prada Corral, David E. Losada, Juan Carlos Pichel, Pablo Gamallo 0001
Nat. Lang. Eng.4
2024 QPU integration in OpenCL for heterogeneous programming
abstract
Abstract The integration of quantum processing units (QPUs) in a heterogeneous high-performance computing environment requires solutions that facilitate hybrid classical–quantum programming. Standards such as OpenCL facilitate the programming of heterogeneous environments, consisting of CPUs and hardware accelerators. This study presents an innovative method that incorporates QPU functionality into OpenCL, standardizing quantum processes within classical environments. By leveraging QPUs within OpenCL, hybrid quantum–classical computations can be sped up, impacting domains like cryptography, optimization problems, and quantum chemistry simulations. Using Portable Computing Language (Jääskeläinen et al. in Int J Parallel Program 43(5):752–785, 2014. https://doi.org/10.1007/s10766-014-0320-y ) and the Qulacs library (Suzuki et al. in Quantum 5:559, 2021. https://doi.org/10.22331/q-2021-10-06-559 ), results demonstrate, for instance, the successful execution of Shor’s algorithm (Nielsen and Chuang in Quantum computation and quantum information, 10th anniversary edn. Cambridge University Press, Cambridge, 2010), serving as a proof of concept for extending the approach to larger qubit systems and other hybrid quantum–classical algorithms. This integration approach bridges the gap between quantum and classical computing paradigms, paving the way for further optimization and application to a wide range of computational problems.
Jorge Vázquez-Pérez, César Piñeiro, Juan Carlos Pichel, Tomás F. Pena, Andrés Gómez 0002
J. Supercomput.3
2022 A multistage retrieval system for health-related misinformation detection
Marcos Fernández-Pichel, David E. Losada, Juan Carlos Pichel
Eng. Appl. Artif. Intell.3
2022 A unified framework to improve the interoperability between HPC and Big Data languages and programming models
abstract
One of the most important issues in the path to the convergence of HPC and Big Data is caused by the differences in their software stacks. Despite some research efforts, the interoperability between their programming models and languages is still limited. To deal with this problem we introduce a new computing framework called IgnisHPC, whose main objective is to unify the execution of Big Data and HPC workloads in the same framework. IgnisHPC has native support for multi-language applications using JVM and non-JVM-based languages. Since MPI was used as its backbone technology, IgnisHPC takes advantage of many communication models and network architectures. Moreover, MPI applications can be directly executed in an efficient way in the framework. The main consequence is that users could combine in the same multi-language code HPC tasks (using MPI) with Big Data tasks (using MapReduce operations). The experimental evaluation demonstrates the benefits of our proposal in terms of performance and productivity with respect to other frameworks. IgnisHPC is publicly available for the Big Data and HPC research community.
César Piñeiro, Juan Carlos Pichel
Future Gener. Comput. Syst.2
2021 Reliability Prediction for Health-Related Content: A Replicability Study
Marcos Fernández-Pichel, David E. Losada, Juan Carlos Pichel, David Elsweiler
ECIR (2)3
2020 Very Fast Tree: speeding up the estimation of phylogenies for large alignments through parallelization and vectorization strategies
abstract
MOTIVATION: FastTree-2 is one of the most successful tools for inferring large phylogenies. With speed at the core of its design, there are still important issues in the FastTree-2 implementation that harm its performance and scalability. To deal with these limitations, we introduce VeryFastTree, a highly tuned implementation of the FastTree-2 tool that takes advantage of parallelization and vectorization strategies to boost performance. RESULTS: VeryFastTree is able to construct a tree on a standard server using double-precision arithmetic from an ultra-large 330k alignment in only 4.5 h, which is 7.8× and 3.5× faster than the sequential and best parallel FastTree-2 times, respectively. AVAILABILITY AND IMPLEMENTATION: VeryFastTree is available at the GitHub repository: https://github.com/citiususc/veryfasttree. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
César Piñeiro, José Manuel Abuín, Juan Carlos Pichel
Bioinform.3
2020 A big data approach to metagenomics for all-food-sequencing
abstract
BACKGROUND: All-Food-Sequencing (AFS) is an untargeted metagenomic sequencing method that allows for the detection and quantification of food ingredients including animals, plants, and microbiota. While this approach avoids some of the shortcomings of targeted PCR-based methods, it requires the comparison of sequence reads to large collections of reference genomes. The steadily increasing amount of available reference genomes establishes the need for efficient big data approaches. RESULTS: We introduce an alignment-free k-mer based method for detection and quantification of species composition in food and other complex biological matters. It is orders-of-magnitude faster than our previous alignment-based AFS pipeline. In comparison to the established tools CLARK, Kraken2, and Kraken2+Bracken it is superior in terms of false-positive rate and quantification accuracy. Furthermore, the usage of an efficient database partitioning scheme allows for the processing of massive collections of reference genomes with reduced memory requirements on a workstation (AFS-MetaCache) or on a Spark-based compute cluster (MetaCacheSpark). CONCLUSIONS: We present a fast yet accurate screening method for whole genome shotgun sequencing-based biosurveillance applications such as food testing. By relying on a big data approach it can scale efficiently towards large-scale collections of complex eukaryotic and bacterial reference genomes. AFS-MetaCache and MetaCacheSpark are suitable tools for broad-scale metagenomic screening applications. They are available at https://muellan.github.io/metacache/afs.html (C++ version for a workstation) and https://github.com/jmabuin/MetaCacheSpark (Spark version for big data clusters).
Robin Kobus, José Manuel Abuín, André Müller, Sören Lukas Hellmann, Juan Carlos Pichel, Tomás F. Pena, Andreas Hildebrandt 0001, Thomas Hankeln, Bertil Schmidt
BMC Bioinform.5
2020 Ignis: An efficient and scalable multi-language Big Data framework
César Piñeiro, Rodrigo Martínez-Castaño, Juan Carlos Pichel
Future Gener. Comput. Syst.3
2019 Dataflow Execution of Hierarchically Tiled Arrays
Chih-Chieh Yang, Juan Carlos Pichel, David A. Padua
Euro-Par2
2018 A New Approach for Sparse Matrix Classification Based on Deep Learning Techniques
abstract
In this paper, a new methodology to select the best storage format for sparse matrices based on deep learning techniques is introduced. We focus on the selection of the proper format for the sparse matrix-vector multiplication (SpMV), which is one of the most important computational kernels in many scientific and engineering applications. Our approach considers the sparsity pattern of the matrices as an image, using the RGB channels to code several of the matrix properties. As a consequence, we generate image datasets that include enough information to successfully train a Convolutional Neural Network (CNN). Considering GPUs as target platforms, the trained CNN selects the best storage format 90.1% of the time, obtaining 99.4% of the highest SpMV performance among the tested formats.
Juan Carlos Pichel, Beatriz Pateiro-López
CLUSTER1
2018 A Micromodule Approach for Building Real-Time Systems with Python-Based Models: Application to Early Risk Detection of Depression on Social Media
Rodrigo Martínez-Castaño, Juan Carlos Pichel, David E. Losada, Fabio Crestani
ECIR2
2017 PASTASpark: multiple sequence alignment meets Big Data
abstract
MOTIVATION: One basic step in many bioinformatics analyses is the multiple sequence alignment. One of the state-of-the-art tools to perform multiple sequence alignment is PASTA (Practical Alignments using SATé and TrAnsitivity). PASTA supports multithreading but it is limited to process datasets on shared memory systems. In this work we introduce PASTASpark, a tool that uses the Big Data engine Apache Spark to boost the performance of the alignment phase of PASTA, which is the most expensive task in terms of time consumption. RESULTS: Speedups up to 10× with respect to single-threaded PASTA were observed, which allows to process an ultra-large dataset of 200 000 sequences within the 24-h limit. AVAILABILITY AND IMPLEMENTATION: PASTASpark is an Open Source tool available at https://github.com/citiususc/pastaspark. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
José Manuel Abuín, Tomás F. Pena, Juan Carlos Pichel
Bioinform.3
2015 BigBWA: approaching the Burrows-Wheeler aligner to Big Data technologies
abstract
Abstract Summary: BigBWA is a new tool that uses the Big Data technology Hadoop to boost the performance of the Burrows–Wheeler aligner (BWA). Important reductions in the execution times were observed when using this tool. In addition, BigBWA is fault tolerant and it does not require any modification of the original BWA source code. Availability and implementation: BigBWA is available at the project GitHub repository: https://github.com/citiususc/BigBWA Contact: [email protected] Supplementary information: Supplementary data are available at Bioinformatics online.
José Manuel Abuín, Juan Carlos Pichel, Tomás F. Pena, Jorge Amigo
Bioinform.2
2014 Perldoop: Efficient execution of Perl scripts on Hadoop clusters
abstract
Hadoop is one of the most important implementations of the MapReduce programming model. It is written in Java and most of the programs that run on Hadoop are also written in this language. Hadoop also provides an utility to execute applications written in other languages, known as Hadoop Streaming. However, the ease of use provided by Hadoop Streaming comes at the expense of a noticeable degradation in the performance. In this work, we introduce Perldoop, a new tool that automatically translates Hadoop-ready Perl scripts into its Java counterparts, which can be directly executed on Hadoop while improving their performance significantly. We have tested our tool using several Natural Language Processing (NLP) modules, which consist of hundreds of regular expressions, but Perldoop could be used with any Perl code ready to be executed with Hadoop Streaming. Performance results show that Java codes generated using Perldoop execute up to 12x faster than the original Perl modules using Hadoop Streaming. In this way, the new NLP modules are able to process the whole Wikipedia in less than 2 hours using a Hadoop cluster with 64 nodes.
José Manuel Abuín, Juan Carlos Pichel, Tomás F. Pena, Pablo Gamallo 0001, Marcos García 0001
IEEE BigData2
2014 Multiobjective optimization technique based on monitoring information to increase the performance of thread migration on multicores
abstract
Multicore systems present on-board memory hierarchies and communication networks that influence their performance when they execute shared memory parallel codes. Characterizing this influence is complex, and understanding the effect of particular hardware configurations on different codes is of paramount importance. In this paper, monitoring information extracted from hardware counters in runtime is used to characterize the behaviour of each thread in the parallel code in terms of three values: the number of floating point operations per second, the operational intensity, and the memory access latency. Note that these values characterize the Roofline Model with the inclusion of additional information about memory access latencies. We propose to use this information to guide thread migration strategies that improve the efficiency of the execution of the code by increasing locality and affinity. The idea behind this proposal is to use these three values as objective functions to be optimized as a multiobjective optimization problem. The proposed technique is an iterative method inspired in evolutive optimization algorithms. To this end, an individual utility function is defined to represent the relative importance of these values. This function is a weighted product that can be considered as representative of the performance of each parallel thread. Different configurations of the SAXPY and SDOT kernels on multicores were used to validate the benefits of the proposed thread migration strategies. The results show that our strategy produces improvements up to 25% in scenarios where locality and affinity are low, and negligible degradation is observed when they are high. The use of hardware counters produces low overheads when extracting monitoring information.
Oscar G. Lorenzo, Tomás F. Pena, José Carlos Cabaleiro, Juan Carlos Pichel, Francisco F. Rivera
CLUSTER4
2014 A hardware counter-based toolkit for the analysis of memory accesses in SMPs
abstract
SUMMARY In this paper, a set of three hardware counter (HC)‐based tools to characterise memory access of parallel codes in Symmetric Multiprocessors (SMPs) is presented. This toolkit simplifies accessing and programming HCs, which are included in modern microprocessors. Hardware counters are used to obtain information about memory accesses in a parallel code at very low cost. This information is presented to the user in a friendly way. The first tool can be used to automatically monitor the memory accesses of a system and to analyse a code even if the source is not available. The second tool allows the user to insert in a source code, in a simple and transparent way, the instructions needed to monitor and manage HCs. This way, specific parts of the code can be analysed. The user can either add appropriate directives to a C code or use a graphical interface to select those parts of the code to be analysed. The tool takes this source file and automatically adds the monitoring code. The third tool takes the information gathered by the aforementioned tools, processes it and displays it graphically. This tool shows the information in a comprehensive and simple way, allowing the user to adjust the level of detail. The aim of these tools was to characterise the memory accesses of parallel codes in multicore systems, in which the cache hierarchy can greatly influence the performance. For illustrative purposes, these tools were used to carry out two case studies, a sparse matrix vector product and a dot product. These studies have been made in two different environments. Anyway, they can be used in almost any system as long as the necessary HCs are available.Copyright © 2013 John Wiley & Sons, Ltd.
Oscar G. Lorenzo, Tomás F. Pena, José Carlos Cabaleiro, Juan Carlos Pichel, Juan Ángel Lorenzo del Castillo, Francisco F. Rivera
Concurr. Comput. Pract. Exp.4
2014 Using sampled information: is it enough for the sparse matrix-vector product locality optimization?
abstract
SUMMARY One of the main factors that affect the performance of the sparse matrix–vector product (SpMV) is the low data reuse caused by the irregular and indirect memory access patterns. Different strategies to deal with this problem such as data reordering techniques have been proposed. The computational cost of these techniques is typically high because they consider all the nonzeros of the sparse matrix in order to find an appropriate permutation of rows and columns that improves the SpMV performance. In this paper, we analyze the possibility of increasing the locality of the SpMV using incomplete information in the reordering process. This partial information comes as a consequence of considering only a subset of the nonzero elements of the matrix. These nonzeros are obtained from the original matrix through a sampling process. In particular, two different sampling methods have been considered: a random sampling and an event‐based sampling using hardware counters. We have detected that a small number of samples is enough to obtain quality reorderings. As a consequence, using sampling‐based reorderings leads to noticeable performance improvements with respect to the non‐reordered matrices, reaching speedup values up to 2.1 × . In addition, an important reduction in the computational time required by the reordering technique has been observed. Copyright © 2012 John Wiley & Sons, Ltd.
Juan Carlos Pichel, Juan Ángel Lorenzo del Castillo, Francisco F. Rivera, Dora Blanco Heras, Tomás F. Pena
Concurr. Comput. Pract. Exp.1
2014 3DyRM: a dynamic roofline model including memory latency information
Oscar G. Lorenzo, Tomás F. Pena, José Carlos Cabaleiro, Juan Carlos Pichel, Francisco F. Rivera
J. Supercomput.4
2013 Sparse matrix-vector multiplication on the Single-Chip Cloud Computer many-core processor
Juan Carlos Pichel, Francisco F. Rivera
J. Parallel Distributed Comput.1
2013 A flexible and dynamic page migration infrastructure based on hardware counters
Juan Ángel Lorenzo del Castillo, Juan Carlos Pichel, Francisco F. Rivera, Tomás F. Pena, José Carlos Cabaleiro
J. Supercomput.2
2012 Hardware Counters Based Analysis of Memory Accesses in SMPs
abstract
Modern microprocessors incorporate Hardware Counters (HC) that provide useful information with low overhead. HC are not commonly used because of the lack of tools to get their information in an easy way. In this paper, a set of tools to simplify the accessing and programming of Intel Itanium 2 ™EARs (Event Address Registers) is presented. The aim of these tools is to characterise the memory accesses of parallel codes, in multicore systems, in which the cache hierarchy can greatly influence the performance. The first tool allows the user to insert in the code, in a simple and transparent way, the instructions needed to monitor and manage hardware counters. Two versions of this tool have been implemented. The first one is a command line tool that takes as input a C source file with appropriate directives and outputs it with the monitoring code added. The other one is a graphical interface that allows the user to select the parts of the code to analise. The second tool takes the information gathered by the monitored parallel code provided by the hardware counters and displays it graphically. This tool shows the information in a comprehensive but simple way, allowing the user to adjust the level of detail. These tools were used to carry out a study of parallel irregular codes. Although this study has been made in a specific environment, the tools here presented can be used in any system as long as it is based on hardware counters present in current processors.
Oscar G. Lorenzo, Tomás F. Pena, José Carlos Cabaleiro, Juan Carlos Pichel, Juan Ángel Lorenzo del Castillo, Francisco F. Rivera
ISPA4
2012 A Graphical Tool for Performance Analysis of Multicore Systems Based on the Roofline Model
abstract
A tool to characterize the performance of parallel codes on multicore systems is presented in this paper. This tool allows the user to define the Roofline Model of the target system, to execute the code under study and to represent the performance results in the roofline plot. The final product is an easy to use tool to provide an insightful model which allows to determine, at a glance, performance issues like load balance, locality and those related to thread and memory allocation. Results show that this model provides practical information of the effects that degrade the performance of a code and gives hints to improve it.
Francisco F. Rivera, Ramón Iglesias, Juan Ángel Lorenzo del Castillo, Juan Carlos Pichel, Tomás F. Pena, José Carlos Cabaleiro
ISPA4
2011 Analyzing the execution of sparse matrix-vector product on the Finisterrae SMP-NUMA system
Juan Carlos Pichel, Juan Ángel Lorenzo del Castillo, Dora Blanco Heras, José Carlos Cabaleiro, Tomás F. Pena
J. Supercomput.1
2010 Lessons Learnt Porting Parallelisation Techniques for Irregular Codes to NUMA Systems
abstract
This work presents a study undertaken to characterise the behaviour of some parallelisation techniques for irregular codes, previously developed for SMP architectures, on a several-node SMP NUMA system. The main objective is to determine the performance effect of bus contention and cache coherency in such a complex architecture. Results show that: (1) cores which share a socket can be considered as independent processors in this context; (2) for big data sizes, the effect of sharing a bus degrades the performance but masks the cache coherency effects and (3) the NUMA-ratio is a critical factor on irregular codes. These results allow us to study the effect in performance of the thread-to-core mappings and memory allocation policies.
Juan Ángel Lorenzo del Castillo, Juan Carlos Pichel, David LaFrance-Linden, Francisco F. Rivera, David E. Singh
PDP2
2009 On the Influence of Thread Allocation for Irregular Codes in NUMA Systems
abstract
This work presents a study undertaken to characterise the FINISTERRAE supercomputer, one of the biggest NUMA systems in Europe. The main objective was to determine the performance effect of bus contention and cache coherency as well as the suitability of porting strategies regarding irregular codes in such a complex architecture. Results show that: (1) cores which share a socket can be considered as independent processors in this context; (2) for big data sizes, the effect of sharing a bus degrades the final performance but masks the cache coherency effects; (3) the NUMA factor (remote to local memory latency ratio) is an important factor on irregular codes and (4) the default kernel allocation policy is not optimal in this system. These results allow us to understand the behaviour of thread-to-core mappings and memory allocation policies.
Juan Ángel Lorenzo del Castillo, Francisco F. Rivera, Petr Tuma 0001, Juan Carlos Pichel
PDCAT4
2009 Increasing data reuse of sparse algebra codes on simultaneous multithreading architectures
abstract
Abstract In this paper the problem of the locality of sparse algebra codes on simultaneous multithreading (SMT) architectures is studied. In these kind of architectures many hardware structures are dynamically shared among the running threads. This puts a lot of stress on the memory hierarchy, and a poor locality, both inter‐thread and intra‐thread, may become a major bottleneck in the performance of a code. This behavior is even more pronounced when the code is irregular, which is the case of sparse matrix ones. Therefore, techniques that increase the locality of irregular codes on SMT architectures are important to achieve high performance. This paper proposes a data reordering technique specially tuned for these kind of architectures and codes. It is based on a locality model developed by the authors in previous works. The technique has been tested, first, using a simulator of a SMT architecture, and subsequently, on a real architecture as Intel's Hyper‐Threading. Important reductions in the number of cache misses have been achieved, even when the number of running threads grows. When applying the locality improvement technique, we also decrease the total execution time and improve the scalability of the code. Copyright © 2009 John Wiley & Sons, Ltd.
Juan Carlos Pichel, Dora Blanco Heras, José Carlos Cabaleiro, Francisco F. Rivera
Concurr. Comput. Pract. Exp.1
2009 A collective I/O implementation based on inspector-executor paradigm
David E. Singh, Florin Isaila, Juan Carlos Pichel, Jesús Carretero 0001
J. Supercomput.3
2009 A collective I/O implementation based on inspector-executor paradigm
David E. Singh, Florin Isaila, Juan Carlos Pichel, Jesús Carretero 0001
J. Supercomput.3
2008 Exploiting data compression in collective I/O techniques
abstract
This paper presents Two-Phase Compressed I/O (TPC I/O,) an optimization of the Two-Phase collective I/O technique from ROMIO, the most popular MPI-IO implementation. In order to reduce network traffic, TPC I/O employs LZO algorithm to compress and decompress exchanged data in the inter-node communication operations. The compression algorithm has been fully implemented in the MPI collective technique, allowing to dynamically use (or not) compression. Compared with Two-Phase I/O, Two-Phase Compressed I/O obtains important improvements in the overall execution time for many of the considered scenarios.
Rosa Filgueira, David E. Singh, Juan Carlos Pichel, Jesús Carretero 0001
CLUSTER3
2008 Reordering Algorithms for Increasing Locality on Multicore Processors
abstract
In order to efficiently exploit available parallelism, multicore processors must address contention for shared resources as cache hierarchy. This fact becomes even more important when irregular codes are executed on them, which is the case for sparse matrix ones. In this paper a technique for increasing locality of sparse matrix codes on multicore platforms is presented. The technique consists on reorganizing the data guided by a locality model which introduces the concept of windows of locality. The evaluation of the reordering technique has been performed on two different leading multicore platforms: Intel Core2Duo and Intel Xeon. Experimental results show important performance improvements when using our reordered matrices with respect to original ones. In particular, an average execution time reduction of about 30% is achieved considering different number of running threads. These results are due to an improved overall cache behavior. Likewise, a comparison of our proposal with some standard reordering techniques is included in the paper. Results point out that the reordering technique always outperforms standard algorithms and is effective for matrices with any structure.
Juan Carlos Pichel, David E. Singh, Jesús Carretero 0001
HPCC1
2006 Image segmentation based on merging of sub-optimal segmentations
Juan Carlos Pichel, David E. Singh, Francisco F. Rivera
Pattern Recognit. Lett.1
2005 Performance optimization of irregular codes based on the combination of reordering and blocking techniques
Juan Carlos Pichel, Dora Blanco Heras, José Carlos Cabaleiro, Francisco F. Rivera
Parallel Comput.1