Tomás F. Pena

dblp:65/2366 · also Tomás Fernandez Pena, Tomás Fernández Pena · DBLP profile ↗
← Back
45ranked-venue papers
1as first author
10since 2021 · last 2026
0000-0002-7622-4698ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 31 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 6Security and privacy · 3 · 2 since 2021Databases, data management, data science and information retrieval · 2Artificial intelligence and machine learning · 1Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 NetQIR: An extension of QIR for distributed quantum computing
abstract
The rapid advancement of quantum computing has highlighted the need for scalable and efficient software infrastructures to fully exploit its potential. Current quantum processors face significant scalability constraints due to the limited number of qubits per chip. In response, distributed quantum computing (DQC) —achieved by networking multiple quantum processor units (QPUs)— is emerging as a promising solution. To support this paradigm, robust intermediate representations (IRs) are needed to translate high-level quantum algorithms into executable instructions suitable for distributed systems. This paper presents NetQIR, an extension of Microsoft’s Quantum Intermediate Representation (QIR), specifically designed to facilitate DQC by incorporating new instruction specifications. NetQIR was developed in response to the lack of abstraction at the network and hardware layers identified in the existing literature as a significant obstacle to effectively implementing distributed quantum algorithms. Based on this analysis, NetQIR introduces new essential abstraction features to support compilers in DQC contexts. It defines network communication instructions independent of specific hardware, abstracting the complexities of inter-QPU communication. Although the proposed work allows abstraction of the underlying network, it is important to note that it is intended for the development of high-performance code on future modular quantum architectures. Leveraging the QIR framework, NetQIR aims to bridge the gap between high-level quantum algorithm design and low-level hardware execution, thus promoting modular and scalable approaches to quantum software infrastructures for distributed applications. Furthermore, its design may serve as a foundational component for future implementations of distributed quantum standards such as the Quantum Message Passing Interface (QMPI).
Francisco Javier Cardama, Jorge Vázquez-Pérez, César Piñeiro, Tomás F. Pena, Juan Carlos Pichel, Andrés Gómez 0002
Future Gener. Comput. Syst.4
2026 Reconstruction of phylogenetic trees via graph-splitting using quantum computing
abstract
Abstract Quantum computing applies principles of quantum mechanics, such as superposition and entanglement, to process information with exponential parallelism. This paradigm offers significant computational advantages over classical methods, particularly for NP-hard problems like phylogenetic tree reconstruction in evolutionary biology. Phylogenetic trees model the evolutionary relationships among species or genes, and their reconstruction is computationally challenging as the number of possible topologies grows exponentially with the number of taxa. To address this, biologists often rely on heuristic methods; however, recent work has shown that recursive graph-cut techniques can achieve high accuracy in phylogenetic inference, though at high computational cost. In this study, we present a quantum algorithm based on the normalized cut ( $$N_{\text {cut}}$$ N cut ) criterion, enabling efficient recursive graph partitioning. Implemented using Quantum Annealing (QA) and the Quantum Approximate Optimization Algorithm (QAOA), demonstrating promising results on real quantum hardware for complex bioinformatics tasks.
Nicolás Fernández-Otero, Tomás F. Pena, Juan Carlos Pichel
J. Supercomput.2
2025 Review of intermediate representations for quantum computing
abstract
Abstract Intermediate representations (IRs) are fundamental to classical and quantum computing, bridging high-level quantum programming languages and the hardware-specific instructions required for execution. This paper reviews the development of quantum IRs, focusing on their evolution and the need for abstraction layers that facilitate portability and optimization. Monolithic quantum IRs, such as QIR (Lubinski et al. in Front Phys 10:940293, 2022. https://doi.org/10.3389/fphy.2022.940293), QSSA (Peduri et al. in Proceedings of the 31st ACM SIGPLAN international conference on compiler construction. CC 2022. Association for Computing Machinery, New York, 2022), or Q-MLIR (McCaskey and Nguyen in Proceedings-2021 IEEE International Conference on Quantum Computing and Engineering, QCE, 2021), their effectiveness in handling abstractions, and their hybrid support between quantum-classical operations are evaluated. However, a key limitation is their inability to address qubit locality, an essential feature for distributed quantum computing (DQC). To overcome this, InQuIR (Nishio and Wakizaka in InQuIR: Intermediate Representation for Interconnected Quantum Computers, 2023. https://arxiv.org/abs/2302.00267) was introduced as an IR specifically designed for distributed systems, providing explicit control over qubit locality and inter-node communication. While effective in managing qubit distribution, InQuIR’s dependence on manual manipulation of communication protocols increases complexity for developers. NetQIR (Vázquez-Pérez et al. in NetQIR: An Extension of QIR for Distributed Quantum Computing, 2024. https://arxiv.org/abs/2408.03712), an extension of QIR for DQC, emerges as a solution to achieve the abstraction of quantum communications protocols. This review emphasizes the need for further advancements in IRs for distributed quantum systems, which will play a crucial role in the scalability and usability of future quantum networks.
Francisco Javier Cardama, Jorge Vázquez-Pérez, César Piñeiro, Juan Carlos Pichel, Tomás F. Pena, Andrés Gómez 0002
J. Supercomput.5
2025 Inqasm: InQuIR compiler to NetQASM
abstract
Abstract Quantum computing is a rapidly evolving field, with almost every aspect open to change or improvement. This includes moving from using a single quantum processing unit to interconnecting multiple quantum processing units (or several of them), establishing a new paradigm called distributed quantum computing and increasing the overall computing capability. Some research is already underway in this area to prepare the ground for an eventual architecture with these characteristics. This is the case of InQuIR (Nishio and Wakizaka in arXiv:2302.00267 2023) and NetQASM (Dahlberg et al in QST 7:035023 2022), two languages developed for distributed quantum computing. This paper presents the development of the InQASM compiler with the aim of translating code from the InQuIR language to NetQASM, establishing a compilation stack for the new distributed paradigm. An example of this compilation and a simulation of the compiled code are shown to showcase it.
Jorge Vázquez-Pérez, Francisco Javier Cardama, César Piñeiro, Juan Carlos Pichel, Tomás F. Pena, Andrés Gómez 0002
J. Supercomput.5
2024 Assessing Intel OneAPI capabilities and cloud-performance for heterogeneous computing
abstract
Abstract This work presents a performance-oriented study of a heterogeneous application developed with Intel OneAPI to solve two well-known diffusion problems: heat diffusion and image denoising. We have explored CPU+iGPU and CPU+FPGA schemes, applying dynamic load balancing and conducting experiments on Intel DevCloud. The results demonstrate that the CPU+iGPU scheme outperforms the execution times achieved by the fastest device when the problem is sufficiently computationally demanding. We also found that the performance of the CPU+FPGA scheme is heavily affected by bandwidth limitations and specific strategies to manage memory efficiently are required. Moreover, it was demonstrated that dynamic workload balancing is crucial due to possible performance fluctuations in any of the implicated devices. In conclusion, Intel OneAPI provides a helpful tool for multi-platform development using a unique high-level language, DPC++. However, developing specific code for each platform is necessary to achieve optimal performance.
Silvia R. Alcaraz, Ruben Laso, Oscar G. Lorenzo, David López Vilariño, Tomás F. Pena, Francisco F. Rivera
J. Supercomput.5
2024 QPU integration in OpenCL for heterogeneous programming
abstract
Abstract The integration of quantum processing units (QPUs) in a heterogeneous high-performance computing environment requires solutions that facilitate hybrid classical–quantum programming. Standards such as OpenCL facilitate the programming of heterogeneous environments, consisting of CPUs and hardware accelerators. This study presents an innovative method that incorporates QPU functionality into OpenCL, standardizing quantum processes within classical environments. By leveraging QPUs within OpenCL, hybrid quantum–classical computations can be sped up, impacting domains like cryptography, optimization problems, and quantum chemistry simulations. Using Portable Computing Language (Jääskeläinen et al. in Int J Parallel Program 43(5):752–785, 2014. https://doi.org/10.1007/s10766-014-0320-y ) and the Qulacs library (Suzuki et al. in Quantum 5:559, 2021. https://doi.org/10.22331/q-2021-10-06-559 ), results demonstrate, for instance, the successful execution of Shor’s algorithm (Nielsen and Chuang in Quantum computation and quantum information, 10th anniversary edn. Cambridge University Press, Cambridge, 2010), serving as a proof of concept for extending the approach to larger qubit systems and other hybrid quantum–classical algorithms. This integration approach bridges the gap between quantum and classical computing paradigms, paving the way for further optimization and application to a wide range of computational problems.
Jorge Vázquez-Pérez, César Piñeiro, Juan Carlos Pichel, Tomás F. Pena, Andrés Gómez 0002
J. Supercomput.4
2023 Digital forensic analysis of the private mode of browsers on Android
abstract
The smartphone has become an essential electronic device in our daily lives. We carry our most precious and important data on it, from family videos of the last few years to credit card information so that we can pay with our phones. In addition, in recent years, mobile devices have become the preferred device for surfing the web, already representing more than 50% of Internet traffic. As one of the devices we spend the most time with throughout the day, it is not surprising that we are increasingly demanding a higher level of privacy. One of the measures introduced to help us protect our data by isolating certain activities on the Internet is the private mode integrated in most modern browsers. Of course, this feature is not new, and has been available on desktop platforms for more than a decade. Reviewing the literature, one can find several studies that test the correct functioning of the private mode on the desktop. However, the number of studies conducted on mobile devices is incredibly small. And not only is it small, but also most of them perform the tests using various emulators or virtual machines running obsolete versions of Android. Therefore, in this paper we apply the methodology we presented in a previous work to Google Chrome, Brave, Mozilla Firefox, and Tor Browser running on a tablet with Android 13 and on two virtual devices created with Android Emulator. The results confirm that these browsers do not store information about the browsing performed in private mode in the file system. However, the analysis of the volatile memory made it possible to recover the username and password used to log in to a website or the keywords typed in a search engine, even after the devices had been rebooted.
Xosé Fernández-Fuentes, Tomás F. Pena, José Carlos Cabaleiro
Comput. Secur.2
2022 Digital forensic analysis methodology for private browsing: Firefox and Chrome on Linux as a case study
abstract
The web browser has become one of the basic tools of everyday life. A tool that is increasingly used to manage personal information. This has led to the introduction of new privacy options by the browsers, including private mode. In this paper, a methodology to explore the effectiveness of the private mode included in most browsers is proposed. A browsing session was designed and conducted in Mozilla Firefox and Google Chrome running on four different Linux environments. After analyzing the information written to disk and the information available in memory, it can be observed that Firefox and Chrome did not store any browsing-related information on the hard disk. However, memory analysis reveals that a large amount of information could be retrieved in some of the environments tested. For example, for the case where the browsers were executed in a VMware virtual machine, it was possible to retrieve most of the actions performed, from the keywords entered in a search field to the username and password entered to log in to a website, even after restarting the computer. In contrast, when Firefox was run on a slightly hardened non-virtualized Linux, it was not possible to retrieve any browsing-related artifacts after the browser was closed.
Xosé Fernández-Fuentes, Tomás F. Pena, José Carlos Cabaleiro
Comput. Secur.2
2022 CIMAR, NIMAR, and LMMA: Novel algorithms for thread and memory migrations in user space on NUMA systems using hardware counters
abstract
This paper introduces two novel algorithms for thread migrations, named CIMAR (Core-aware Interchange and Migration Algorithm with performance Record –IMAR–) and NIMAR (Node-aware IMAR), and a new algorithm for the migration of memory pages, LMMA (Latency-based Memory pages Migration Algorithm), in the context of Non-Uniform Memory Access (NUMA) systems. This kind of system has complex memory hierarchies that present a challenging problem in extracting the best possible performance, where thread and memory mapping play a critical role. The presented algorithms gather and process the information provided by hardware counters to make decisions about the migrations to be performed, trying to find the optimal mapping. They have been implemented as a user space tool that looks for improving the system performance, particularly in, but not restricted to, scenarios where multiple programs with different characteristics are running. This approach has the advantage of not requiring any modification on the target programs or the Linux kernel while keeping a low overhead. Two different benchmark suites have been used to validate our algorithms: The NAS parallel benchmark, mainly devoted to computational routines, and the LevelDB database benchmark focused on read–write operations. These benchmarks allow us to illustrate the influence of our proposal in these two important types of codes. Note that those codes are state-of-the-art implementations of the routines, so few improvements could be initially expected. Experiments have been designed and conducted to emulate three different scenarios: a single program running in the system with full resources, an interactive server where multiple programs run concurrently varying the availability of resources, and a queue of tasks where granted resources are limited. The proposed algorithms have been able to produce significant benefits, especially in systems with higher latency penalties for remote accesses. When more than one benchmark is executed simultaneously, performance improvements have been obtained, reducing execution times up to 60%. In this kind of situation, the behaviour of the system is more critical, and the NUMA topology plays a more relevant role. Even in the worst case, when isolated benchmarks are executed using the whole system, that is, just one task at a time, the performance is not degraded.
Ruben Laso, Oscar G. Lorenzo, José Carlos Cabaleiro, Tomás F. Pena, Juan Ángel Lorenzo del Castillo, Francisco F. Rivera
Future Gener. Comput. Syst.4
2021 LBMA and IMAR2: Weighted lottery based migration strategies for NUMA multiprocessing servers
abstract
Summary Multicore NUMA systems present on‐board memory hierarchies and communication networks that influence performance when executing shared memory parallel codes. Characterizing this influence is complex, and understanding the effect of particular hardware configurations on different codes is of paramount importance. In this article, monitoring information extracted from hardware counters at runtime is used to characterize the behavior of each thread for an arbitrary number of multithreaded processes running in a multiprocessing environment. This characterization is given in terms of number of operations per second, operational intensity, and latency of memory accesses. We propose a runtime tool, executed in user space, that uses this information to guide two different thread migration strategies for improving execution efficiency by increasing locality and affinity without requiring any modification in the running codes. Different configurations of NAS Parallel OpenMP benchmarks running concurrently on multicore NUMA systems were used to validate the benefits of our proposal, in which up to four processes are running simultaneously. In more than the 95% of the executions of our tool, results outperform those of the operating system (OS) and produces up to 38% improvement in execution time over the OS for heterogeneous workloads, under different and realistic locality and affinity scenarios.
Ruben Laso, Oscar G. Lorenzo, Francisco F. Rivera, José Carlos Cabaleiro, Tomás F. Pena, Juan Ángel Lorenzo del Castillo
Concurr. Comput. Pract. Exp.5
2020 A big data approach to metagenomics for all-food-sequencing
abstract
BACKGROUND: All-Food-Sequencing (AFS) is an untargeted metagenomic sequencing method that allows for the detection and quantification of food ingredients including animals, plants, and microbiota. While this approach avoids some of the shortcomings of targeted PCR-based methods, it requires the comparison of sequence reads to large collections of reference genomes. The steadily increasing amount of available reference genomes establishes the need for efficient big data approaches. RESULTS: We introduce an alignment-free k-mer based method for detection and quantification of species composition in food and other complex biological matters. It is orders-of-magnitude faster than our previous alignment-based AFS pipeline. In comparison to the established tools CLARK, Kraken2, and Kraken2+Bracken it is superior in terms of false-positive rate and quantification accuracy. Furthermore, the usage of an efficient database partitioning scheme allows for the processing of massive collections of reference genomes with reduced memory requirements on a workstation (AFS-MetaCache) or on a Spark-based compute cluster (MetaCacheSpark). CONCLUSIONS: We present a fast yet accurate screening method for whole genome shotgun sequencing-based biosurveillance applications such as food testing. By relying on a big data approach it can scale efficiently towards large-scale collections of complex eukaryotic and bacterial reference genomes. AFS-MetaCache and MetaCacheSpark are suitable tools for broad-scale metagenomic screening applications. They are available at https://muellan.github.io/metacache/afs.html (C++ version for a workstation) and https://github.com/jmabuin/MetaCacheSpark (Spark version for big data clusters).
Robin Kobus, José Manuel Abuín, André Müller, Sören Lukas Hellmann, Juan Carlos Pichel, Tomás F. Pena, Andreas Hildebrandt 0001, Thomas Hankeln, Bertil Schmidt
BMC Bioinform.6
2020 Next-generation big data federation access control: A reference model
Feras M. Awaysheh, Mamoun Alazab, Maanak Gupta, Tomás F. Pena, José Carlos Cabaleiro
Future Gener. Comput. Syst.4
2020 TrustE-VC: Trustworthy Evaluation Framework for Industrial Connected Vehicles in the Cloud
abstract
The integration between cloud computing and vehicular ad hoc networks, namely, vehicular clouds (VCs), has become a significant research area. This integration was proposed to accelerate the adoption of intelligent transportation systems. The trustworthiness in VCs is expected to carry more computing capabilities that manage large-scale collected data. This trend requires a security evaluation framework that ensures data privacy protection, integrity of information, and availability of resources. To the best of our knowledge, this is the first study that proposes a robust trustworthiness evaluation of vehicular cloud for security criteria evaluation and selection. This article proposes three-level security features in order to develop effectiveness and trustworthiness in VCs. To assess and evaluate these security features, our evaluation framework consists of three main interconnected components: 1) an aggregation of the security evaluation values of the security criteria for each level; 2) a fuzzy multicriteria decision-making algorithm; and 3) a simple additive weight associated with the importance-performance analysis and performance rate to visualize the framework findings. The evaluation results of the security criteria based on the average performance rate and global weight suggest that data residency, data privacy, and data ownership are the most pressing challenges in assessing data protection in a VC environment. Overall, this article paves the way for a secure VC using an evaluation of effective security features and underscores directions and challenges facing the VC community. This article sheds light on the importance of security by design, emphasizing multiple layers of security when implementing industrial VCs.
Mohammad Aladwan, Feras M. Awaysheh, Sadi Alawadi, Mamoun Alazab, Tomás F. Pena, José Carlos Cabaleiro
IEEE Trans. Ind. Informatics5
2019 Leveraging Bitmap Indexing for Subgraph Searching
abstract
Deciding whether a query graph is a subgraph of some other in a very large database of small graphs is a problem of major interest in many application domains. As an example, it arises in the searching for specific molecular substructures in currently available molecular databases, whose sizes may reach levels close to one hundred million. State of the art methods to solve this problem follow a filter-then-verify (FTV) paradigm, where an indexing technique is first used in a filtering stage to obtain result candidates and a subgraph isomorphism algorithm is next applied to the candidates in a verification stage to obtain the final result. Among all the available techniques of the state of the art, two of them have demonstrated better performance when applied to large datasets, namely, the GraphGrepSX (GGSX) and CT-Index (CTI). In this paper, three new indexing techniques, one based on GGSX and two based on CTI, are proposed. In particular, Bitmap GGSX (BM-GGSX) leverages the use of bitmaps in the trie structure used by GGSX to achieve performance gains of around 90% in the filtering stage. Column-Wise CT-Index (CW-CTI) exploits a column-wise representation of the fingerprints (bitmaps) used by CT-Index to reduce the filtering times around 80% for small queries (8 edges). Finally, K-Means CT-Index (KM-CTI), constructs a binary tree of bitmaps from the CT-Index fingerprints to reach filtering time reductions of around 70% for medium queries (20 edges) and 75% for large queries (40 edges).
David Luaces, José R. R. Viqueira, Tomás F. Pena, José Manuel Cotos
EDBT3
2019 Poster: A Pluggable Authentication Module for Big Data Federation Architecture
abstract
This paper intends to propose a trustworthy model for authenticating users and services over a Big Data Federation deployment architecture. The main goal of this model is to provide a Single-Sign-on (SSO) approach for the latest Hadoop 3.x platform. To achieve this, a conceptual model is proposed combining Hadoop access control primitives and the Apache Knox framework. The paper provides various insights regarding the latest ongoing developments and open challenges in this domain.
Feras M. Awaysheh, José Carlos Cabaleiro, Tomás F. Pena, Mamoun Alazab
SACMAT3
2017 EME: An Automated, Elastic and Efficient Prototype for Provisioning Hadoop Clusters On-demand
Feras M. Awaysheh, Tomás F. Pena, José Carlos Cabaleiro
CLOSER2
2017 PASTASpark: multiple sequence alignment meets Big Data
abstract
MOTIVATION: One basic step in many bioinformatics analyses is the multiple sequence alignment. One of the state-of-the-art tools to perform multiple sequence alignment is PASTA (Practical Alignments using SATé and TrAnsitivity). PASTA supports multithreading but it is limited to process datasets on shared memory systems. In this work we introduce PASTASpark, a tool that uses the Big Data engine Apache Spark to boost the performance of the alignment phase of PASTA, which is the most expensive task in terms of time consumption. RESULTS: Speedups up to 10× with respect to single-threaded PASTA were observed, which allows to process an ultra-large dataset of 200 000 sequences within the 24-h limit. AVAILABILITY AND IMPLEMENTATION: PASTASpark is an Open Source tool available at https://github.com/citiususc/pastaspark. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
José Manuel Abuín, Tomás F. Pena, Juan Carlos Pichel
Bioinform.2
2017 Landing sites detection using LiDAR data on manycore systems
Oscar G. Lorenzo, Jorge Martínez Sánchez, David López Vilariño, Tomás F. Pena, José Carlos Cabaleiro, Francisco F. Rivera
J. Supercomput.4
2016 State of the Journal
abstract
Discusses the current state of the journal, reports on current and future areas of exploration and research, and presents new editors.
Paolo Montuschi, Edward J. McCluskey, Samarjit Chakraborty, Jason Cong, Ramón M. Rodríguez-Dagnino, Fred Douglis, Lieven Eeckhout, Gernot Heiser, Sushil Jajodia, Ruby B. Lee, Dinesh Manocha, Tomás F. Pena, Isabelle Puaut, Hanan Samet, Donatella Sciuto
IEEE Trans. Computers12
2015 BigDataDIRAC: Deploying Distributed Big Data Applications
abstract
The Distributed Infrastructure with Remote Agent Control (DIRAC) software framework allows a user community to manage computing activities in different grid and cloud environments. Many communities from several fields (LHCb, Belle II, Creatis, DIRAC4EGI multiple community portal, etc.) use DIRAC to run jobs in distributed environments. Google created the MapReduce programming model offering an efficient way of performing distributed computation over large data sets. Several enterprises are providing Hadoop cloud based resources to their users, and are trying to simplify the usage of Hadoop in the cloud. Based in these two robust technologies, we have created BigDataDIRAC, a solution which gives users the opportunity to access multiple Big Data resources scattered in different geographical areas, such as access to grid resources. This approach opens the possibility of offering not only grid and cloud to the users, but also Big Data resources from the same DIRAC environment. Proof of concept is shown using three computing centers in two countries, and with four Hadoop clusters. Our results demonstrate the ability of BigDataDIRAC to manage jobs driven by dataset location stored in the Hadoop File System (HDFS) of the Hadoop distributed clusters. DIRAC is used to monitor the execution, collect the necessary statistical data, and upload the results from the remote HDFS to the SandBox Storage machine. The tests produced the equivalent of 5 days continuous processing.
Víctor Fernández 0002, Víctor Méndez Muñoz, Tomás F. Pena
CCGRID3
2015 Study of the KVM CPU Performance of Open-Source Cloud Management Platforms
abstract
Nowadays, there are several open-source solutions for building private, public and even hybrid clouds such as Eucalyptus, Apache Cloud Stack and Open Stack. KVM is one of the supported hypervisors for these cloud platforms. Different KVM configurations are being supplied by these platforms and, in some cases, a subset of CPU features are being presented to guest systems, providing a basic abstraction of the underlying CPU. One of the reasons for limiting the features of the Virtual CPU is to guarantee the guest compatibility with different hardware in heterogeneous environments. However, in a large number of situations, the cloud is deployed on an homogeneous set of hosts. In these cases, this limitation can affect the performance of applications being executed in guest systems. In this paper, we have analyzed the architecture, the KVM setup, and the performance of the Virtual Machines deployed by three popular cloud management platforms: Eucalyptus, Apache Cloud Stack and Open Stack, employing a representative set of applications.
Fernando Gomez-Folgar, Antonio J. García-Loureiro, Tomás F. Pena, J. Isaac Zablah, Natalia Seoane
CCGRID3
2015 BigBWA: approaching the Burrows-Wheeler aligner to Big Data technologies
abstract
Abstract Summary: BigBWA is a new tool that uses the Big Data technology Hadoop to boost the performance of the Burrows–Wheeler aligner (BWA). Important reductions in the execution times were observed when using this tool. In addition, BigBWA is fault tolerant and it does not require any modification of the original BWA source code. Availability and implementation: BigBWA is available at the project GitHub repository: https://github.com/citiususc/BigBWA Contact: [email protected] Supplementary information: Supplementary data are available at Bioinformatics online.
José Manuel Abuín, Juan Carlos Pichel, Tomás F. Pena, Jorge Amigo
Bioinform.3
2014 Perldoop: Efficient execution of Perl scripts on Hadoop clusters
abstract
Hadoop is one of the most important implementations of the MapReduce programming model. It is written in Java and most of the programs that run on Hadoop are also written in this language. Hadoop also provides an utility to execute applications written in other languages, known as Hadoop Streaming. However, the ease of use provided by Hadoop Streaming comes at the expense of a noticeable degradation in the performance. In this work, we introduce Perldoop, a new tool that automatically translates Hadoop-ready Perl scripts into its Java counterparts, which can be directly executed on Hadoop while improving their performance significantly. We have tested our tool using several Natural Language Processing (NLP) modules, which consist of hundreds of regular expressions, but Perldoop could be used with any Perl code ready to be executed with Hadoop Streaming. Performance results show that Java codes generated using Perldoop execute up to 12x faster than the original Perl modules using Hadoop Streaming. In this way, the new NLP modules are able to process the whole Wikipedia in less than 2 hours using a Hadoop cluster with 64 nodes.
José Manuel Abuín, Juan Carlos Pichel, Tomás F. Pena, Pablo Gamallo 0001, Marcos García 0001
IEEE BigData3
2014 Multiobjective optimization technique based on monitoring information to increase the performance of thread migration on multicores
abstract
Multicore systems present on-board memory hierarchies and communication networks that influence their performance when they execute shared memory parallel codes. Characterizing this influence is complex, and understanding the effect of particular hardware configurations on different codes is of paramount importance. In this paper, monitoring information extracted from hardware counters in runtime is used to characterize the behaviour of each thread in the parallel code in terms of three values: the number of floating point operations per second, the operational intensity, and the memory access latency. Note that these values characterize the Roofline Model with the inclusion of additional information about memory access latencies. We propose to use this information to guide thread migration strategies that improve the efficiency of the execution of the code by increasing locality and affinity. The idea behind this proposal is to use these three values as objective functions to be optimized as a multiobjective optimization problem. The proposed technique is an iterative method inspired in evolutive optimization algorithms. To this end, an individual utility function is defined to represent the relative importance of these values. This function is a weighted product that can be considered as representative of the performance of each parallel thread. Different configurations of the SAXPY and SDOT kernels on multicores were used to validate the benefits of the proposed thread migration strategies. The results show that our strategy produces improvements up to 25% in scenarios where locality and affinity are low, and negligible degradation is observed when they are high. The use of hardware counters produces low overheads when extracting monitoring information.
Oscar G. Lorenzo, Tomás F. Pena, José Carlos Cabaleiro, Juan Carlos Pichel, Francisco F. Rivera
CLUSTER2
2014 A hardware counter-based toolkit for the analysis of memory accesses in SMPs
abstract
SUMMARY In this paper, a set of three hardware counter (HC)‐based tools to characterise memory access of parallel codes in Symmetric Multiprocessors (SMPs) is presented. This toolkit simplifies accessing and programming HCs, which are included in modern microprocessors. Hardware counters are used to obtain information about memory accesses in a parallel code at very low cost. This information is presented to the user in a friendly way. The first tool can be used to automatically monitor the memory accesses of a system and to analyse a code even if the source is not available. The second tool allows the user to insert in a source code, in a simple and transparent way, the instructions needed to monitor and manage HCs. This way, specific parts of the code can be analysed. The user can either add appropriate directives to a C code or use a graphical interface to select those parts of the code to be analysed. The tool takes this source file and automatically adds the monitoring code. The third tool takes the information gathered by the aforementioned tools, processes it and displays it graphically. This tool shows the information in a comprehensive and simple way, allowing the user to adjust the level of detail. The aim of these tools was to characterise the memory accesses of parallel codes in multicore systems, in which the cache hierarchy can greatly influence the performance. For illustrative purposes, these tools were used to carry out two case studies, a sparse matrix vector product and a dot product. These studies have been made in two different environments. Anyway, they can be used in almost any system as long as the necessary HCs are available.Copyright © 2013 John Wiley & Sons, Ltd.
Oscar G. Lorenzo, Tomás F. Pena, José Carlos Cabaleiro, Juan Carlos Pichel, Juan Ángel Lorenzo del Castillo, Francisco F. Rivera
Concurr. Comput. Pract. Exp.2
2014 Modeling the performance of parallel applications using model selection techniques
abstract
SUMMARY Nowadays, parallel architectures are changing so fast that there is a need for scalable and efficient tools to analyze and predict the performance of parallel applications. Analytical models are proved to be a useful approximation for characterizing parallel algorithms, but developing accurate analytical models is a hard issue, and, in general, they provide coarse performance predictions due to their intrinsic lack of accuracy. In this paper, we describe in detail the Tools for Instrumentation and Analysis (TIA) framework, an easy‐to‐use tool that automatically obtains accurate performance models by means of analytical expressions. This framework automatizes most of its internal tasks, reducing opportunities for human error, and it only requires the user to focus on the metrics and execution parameters that might influence the performance, those that should be considered in the modeling process. Its main advantage over other tools is that TIA uses model selection techniques that allow the automation of the modeling process. As a case of study, the use of TIA to obtain analytical models of different implementations of the broadcast collective communication in a cluster of multicores is shown. The results obtained by TIA are evaluated and compared with theoretical approaches based on the LogGP model. Copyright © 2013 John Wiley & Sons, Ltd.
Diego Rodríguez Martínez, Vicente Blanco 0001, José Carlos Cabaleiro, Tomás F. Pena, Francisco F. Rivera
Concurr. Comput. Pract. Exp.4
2014 Using sampled information: is it enough for the sparse matrix-vector product locality optimization?
abstract
SUMMARY One of the main factors that affect the performance of the sparse matrix–vector product (SpMV) is the low data reuse caused by the irregular and indirect memory access patterns. Different strategies to deal with this problem such as data reordering techniques have been proposed. The computational cost of these techniques is typically high because they consider all the nonzeros of the sparse matrix in order to find an appropriate permutation of rows and columns that improves the SpMV performance. In this paper, we analyze the possibility of increasing the locality of the SpMV using incomplete information in the reordering process. This partial information comes as a consequence of considering only a subset of the nonzero elements of the matrix. These nonzeros are obtained from the original matrix through a sampling process. In particular, two different sampling methods have been considered: a random sampling and an event‐based sampling using hardware counters. We have detected that a small number of samples is enough to obtain quality reorderings. As a consequence, using sampling‐based reorderings leads to noticeable performance improvements with respect to the non‐reordered matrices, reaching speedup values up to 2.1 × . In addition, an important reduction in the computational time required by the reordering technique has been observed. Copyright © 2012 John Wiley & Sons, Ltd.
Juan Carlos Pichel, Juan Ángel Lorenzo del Castillo, Francisco F. Rivera, Dora Blanco Heras, Tomás F. Pena
Concurr. Comput. Pract. Exp.5
2014 3DyRM: a dynamic roofline model including memory latency information
Oscar G. Lorenzo, Tomás F. Pena, José Carlos Cabaleiro, Juan Carlos Pichel, Francisco F. Rivera
J. Supercomput.2
2013 A flexible and dynamic page migration infrastructure based on hardware counters
Juan Ángel Lorenzo del Castillo, Juan Carlos Pichel, Francisco F. Rivera, Tomás F. Pena, José Carlos Cabaleiro
J. Supercomput.4
2012 Performance of the CloudStack KVM Pod Primary Storage under NFS Version 3
abstract
Currently, there are an increasing number of open-source solutions for building Clouds. The performance of the Virtual Machines running in these Clouds is a key point. The way the Virtual Machines are created in Clouds can have important effects upon their disk I/O operations. We have selected CloudStack platform to study the disk I/O performance for KVM Virtual Machines.
Fernando Gomez-Folgar, Antonio J. García-Loureiro, Tomás F. Pena, Raúl Valín Ferreiro
ISPA3
2012 Hardware Counters Based Analysis of Memory Accesses in SMPs
abstract
Modern microprocessors incorporate Hardware Counters (HC) that provide useful information with low overhead. HC are not commonly used because of the lack of tools to get their information in an easy way. In this paper, a set of tools to simplify the accessing and programming of Intel Itanium 2 ™EARs (Event Address Registers) is presented. The aim of these tools is to characterise the memory accesses of parallel codes, in multicore systems, in which the cache hierarchy can greatly influence the performance. The first tool allows the user to insert in the code, in a simple and transparent way, the instructions needed to monitor and manage hardware counters. Two versions of this tool have been implemented. The first one is a command line tool that takes as input a C source file with appropriate directives and outputs it with the monitoring code added. The other one is a graphical interface that allows the user to select the parts of the code to analise. The second tool takes the information gathered by the monitored parallel code provided by the hardware counters and displays it graphically. This tool shows the information in a comprehensive but simple way, allowing the user to adjust the level of detail. These tools were used to carry out a study of parallel irregular codes. Although this study has been made in a specific environment, the tools here presented can be used in any system as long as it is based on hardware counters present in current processors.
Oscar G. Lorenzo, Tomás F. Pena, José Carlos Cabaleiro, Juan Carlos Pichel, Juan Ángel Lorenzo del Castillo, Francisco F. Rivera
ISPA2
2012 Model Selection to Characterize Performance Using Genetic Algorithms
abstract
The TIA modeling framework provides analytical models of the performance of parallel applications. The resulting models are obtained using model selection techniques and are accurate enough for various purposes. Its main drawback is that the completion time depends on the number of candidate models and, in some situations, it becomes critical. In this work, a genetic algorithm is proposed for reducing the time for searching of the best candidate model. The use of this genetic algorithm to obtain the performance model of the linear implementation of the broadcast collective communication in a cluster of multicores is shown.
Diego Rodríguez Martínez, José Carlos Cabaleiro, Tomás F. Pena, Francisco F. Rivera, Vicente Blanco 0001
ISPA3
2012 A Graphical Tool for Performance Analysis of Multicore Systems Based on the Roofline Model
abstract
A tool to characterize the performance of parallel codes on multicore systems is presented in this paper. This tool allows the user to define the Roofline Model of the target system, to execute the code under study and to represent the performance results in the roofline plot. The final product is an easy to use tool to provide an insightful model which allows to determine, at a glance, performance issues like load balance, locality and those related to thread and memory allocation. Results show that this model provides practical information of the effects that degrade the performance of a code and gives hints to improve it.
Francisco F. Rivera, Ramón Iglesias, Juan Ángel Lorenzo del Castillo, Juan Carlos Pichel, Tomás F. Pena, José Carlos Cabaleiro
ISPA5
2011 Estimating the effect of cache misses on the performance of parallel applications using analytical models
abstract
In this paper a methodology to characterize the influence of cache misses on the performance of parallel applications is presented. This methodology is based on analytical models provided by the TIA framework. This framework obtains analytical models of given observable quantities by instrumenting the source code and applying model selection techniques. In particular, two metrics related with the performance are considered in this work: the number of cache misses and the elapsed time. Based on both models, the influence in terms of execution time due to the cache misses can be inferred. Two different versions of the parallel product of dense matrices are used as case of study.
Diego Rodríguez Martínez, Vicente Blanco 0001, José Carlos Cabaleiro, Tomás F. Pena, Francisco F. Rivera
AICCSA4
2011 Using accurate AIC-based performance models to improve the scheduling of parallel applications
Diego Rodríguez Martínez, Julio L. Albín, Tomás F. Pena, José Carlos Cabaleiro, Francisco F. Rivera, Vicente Blanco 0001
J. Supercomput.3
2011 Analyzing the execution of sparse matrix-vector product on the Finisterrae SMP-NUMA system
Juan Carlos Pichel, Juan Ángel Lorenzo del Castillo, Dora Blanco Heras, José Carlos Cabaleiro, Tomás F. Pena
J. Supercomput.5
2010 Performance Modeling of MPI Applications Using Model Selection Techniques
abstract
A new method for obtaining models of the performance of parallel applications based on statistical analysis is presented in this paper. This method is based on the Akaike's information criterion (AIC) that provides an objective mechanism to rank different models by means of an experimental data fit. The input of the modeling process is a set of variables and parameters that can a priori influence the performance of the application. This set can be provided by the user. Using this information, the method automatically generates a set of candidate models. These models are fit to the experimental data and the AIC score of each model is calculated. The model with the best AIC score is selected as the best model. Also, using the AIC scores of all candidate models, useful statistical information is provided to help the user to evaluate the quality of the selected model, as well as indications of how to interactively improve this modeling process. As a first case of study, statistical models obtained for different implementations of the broadcast collective communication in Open MPI are shown. These models are very accurate, exceeding its adjustment to theoretical approaches based on the LogGP model. Finally, the NAS Parallel Benchmark is also characterized using this new method with good results in terms of accuracy.
Diego Rodríguez Martínez, José Carlos Cabaleiro, Tomás F. Pena, Francisco F. Rivera, Vicente Blanco 0001
PDP3
2009 Accurate analytical performance model of communications in MPI applications
abstract
This paper presents a new LogP-based model, called LoOgGP, which allows an accurate characterization of MPI applications based on microbenchmark measurements. This new model is an extension of LogP for long messages in which both overhead and gap parameters perform a linear dependency with message size. The LoOgGP model has been fully integrated into a modelling framework to obtain statistical models of parallel applications, providing the analyst with an easy and automatic tool for LoOgGP parameter set assessment to characterize communications. The use of LoOgGP model to obtain a statistical performance model of an image deconvolution application is illustrated as a case of study.
Diego Rodríguez Martínez, José Carlos Cabaleiro, Tomás F. Pena, Francisco F. Rivera, Vicente Blanco 0001
IPDPS3
2007 An Inspector/Executor Based Strategy to Efficiently Parallelize N-Body Simulation Programs on Shared Memory Systems
abstract
Reordering of data is becoming more and more significant in order to achieve a higher performance in memory data access and, particularly, in program runtime. This fact becomes specially important in parallel applications that are executed in shared memory systems. This work presents a new parallelizing, run time strategy for irregular structures associated to N-Body problem simulation algorithms. Such strategy, so-called STPCLS (Step Classification), is based on the inspector-executor paradigm. It has been tested in a shared memory system using a significant set of irregular loops. The outcomes show that the efficiency of our solution is high, and the benefits overcome the overheads imposed by our algorithm.
Juan Ángel Lorenzo del Castillo, Julio L. Albín, Tomás F. Pena, Francisco F. Rivera, David E. Singh
ISPDC3
2004 Performance Prediction for Parallel Iterative Solvers
Vicente Blanco 0001, Patricia González, José Carlos Cabaleiro, Dora Blanco Heras, Tomás F. Pena, Juan J. Pombo, Francisco F. Rivera
J. Supercomput.5
2003 AVISPA: visualizing the performance prediction of parallel iterative solvers
Vicente Blanco 0001, Patricia González, José Carlos Cabaleiro, Dora Blanco Heras, Tomás F. Pena, Juan J. Pombo, Francisco F. Rivera
Future Gener. Comput. Syst.5
2001 Parallel Computation of Wavelet Transforms Using the Lifting Scheme
Patricia González, José Carlos Cabaleiro, Tomás F. Pena
J. Supercomput.3
2000 On parallel solvers for sparse triangular systems
Patricia González, José Carlos Cabaleiro, Tomás F. Pena
J. Syst. Archit.3
1994 Finite Element Simulation of Semiconductor Devices on Multiprocessor Computers
Tomás F. Pena, Emilio L. Zapata, David J. Evans 0001
Parallel Comput.1
1992 Image reconstruction on hypercube computers: Application to electron microscopy
Emilio L. Zapata, José Ignacio Benavides Benítez, Francisco F. Rivera, Javier D. Bruguera, Tomás F. Pena, José María Carazo
Signal Process.5