EDBT 2026 Demo / reviewers in the wild / expert
Jorge González-Domínguez
dblp:35/7411
· DBLP profile ↗
36ranked-venue papers
16as first author
13since 2021 · last 2025
0000-0002-2602-4874ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 24 · 10 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 4 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Biclustering in bioinformatics using big data and High Performance Computing applications: challenges and perspectives, a reviewabstractAbstract Biclustering is a powerful machine learning technique that simultaneously groups rows and columns in matrix-based datasets. Applied to gene expression data in bioinformatics, its use has expanded alongside the rapid growth of high-throughput sequencing technologies, leading to massive and complex biological datasets. This review aims to examine how biclustering methods and their validation strategies are evolving to meet the demands of High Performance Computing (HPC) and Big Data environments. We present a structured classification of existing approaches based on the computational paradigms they employ, including MPI/OpenMP, Apache Hadoop/Spark, and GPU/CUDA. By synthesising these developments, we highlight current trends and outline key research challenges. The knowledge gathered in this work may support researchers in adapting and scaling biclustering algorithms to analyse large-scale biomedical data more efficiently. Our contribution is intended to bridge the gap between algorithmic innovation and computational scalability in the context of bioinformatics and data-intensive applications. Aurelio López-Fernández, Francisco Gómez-Vela, Domingo S. Rodríguez-Baena, Fernando M. Delgado-Chaves, Jorge González-Domínguez |
J. Supercomput. | 5 |
| 2024 | BigDEC: A multi-algorithm Big Data tool based on the k-mer spectrum method for scalable short-read error correctionabstractDespite the significant improvements in both throughput and cost provided by modern Next-Generation Sequencing (NGS) platforms, sequencing errors in NGS datasets can still degrade the quality of downstream analysis. Although state-of-the-art correction tools can provide high accuracy to improve such analysis, they are limited to apply a single correction algorithm while also requiring long runtimes when processing large NGS datasets. Furthermore, current parallel correctors generally only provide efficient support for shared-memory systems lacking the ability to scale out across a cluster of multicore nodes, or they require the availability of specific hardware devices or features. In this paper we present a Big Data Error Correction (BigDEC) tool that overcomes all those limitations by: 1) implementing three different error correction algorithms based on the widely extended k-mer spectrum method; 2) providing scalable performance for large datasets by efficiently exploiting the capabilities of Big Data technologies on multicore clusters based on commodity hardware; 3) supporting two different Big Data processing frameworks (Spark and Flink) to provide greater flexibility to end users; 4) including an efficient, stream-based merge operation to ease downstream processing of the corrected datasets; and 5) significantly outperforming existing parallel tools, being up to 79% faster on a 16-node multicore cluster when using the same underlying correction algorithm. BigDEC is publicly available to download at https://github.com/UDC-GAC/BigDEC. Roberto R. Expósito, Jorge González-Domínguez |
Future Gener. Comput. Syst. | 2 |
| 2024 | CUDA acceleration of MI-based feature selection methodsabstractFeature selection algorithms are necessary nowadays for machine learning as they are capable of removing irrelevant and redundant information to reduce the dimensionality of the data and improve the quality of subsequent analyses. The problem with current feature selection approaches is that they are computationally expensive when processing large datasets. This work presents parallel implementations for Nvidia GPUs of three highly-used feature selection methods based on the Mutual Information (MI) metric: mRMR, JMI and DISR. Publicly available code includes not only CUDA implementations of the general methods, but also an adaptation of them to work with low-precision fixed point in order to further increase their performance on GPUs. The experimental evaluation was carried out on two modern Nvidia GPUs (Turing T4 and Ampere A100) with highly satisfactory results, achieving speedups of up to 283x when compared to state-of-the-art C implementations. Bieito Beceiro, Jorge González-Domínguez, Laura Moran-Fernandez, Verónica Bolón-Canedo, Juan Touriño |
J. Parallel Distributed Comput. | 2 |
| 2024 | PARamrfinder: detecting allele-specific DNA methylation on multicore clustersabstractAbstract The discovery of Allele-Specific Methylation (ASM) is an important research field in biology as it regulates genomic imprinting, which has been identified as the cause of some genetic diseases. Nevertheless, the high computational cost of the bioinformatic tools developed for this purpose prevents their application to large-scale datasets. Hence, much faster tools are required to further progress in this research field. In this work we presentPARamrfinder, a parallel tool that applies a statistical model to identify ASM in data from high-throughput short-read bisulfite sequencing. It is based on the state-of-the-art sequential toolamrfinder, which is able to detect ASM at regional level from Bisulfite Sequencing (BS-Seq) experiments in the absence of Single Nucleotide Polymorphism information.PARamrfinderprovides the same Allelically Methylated Regions asamrfinderbut at significantly reduced runtime thanks to exploiting the compute capabilities of common multicore CPU clusters and MPI RMA operations to attain an efficient dynamic workload balance. As an example, our tool is up to 567 times faster for real data experiments on a cluster with 8 nodes, each one containing two 16-core processors. The source code of PARamrfinder, as well as a reference manual, is available at https://github.com/UDC-GAC/PARamrfinder . Alejandro Fernández-Fraga, Jorge González-Domínguez, María J. Martín |
J. Supercomput. | 2 |
| 2023 | PATO: genome-wide prediction of lncRNA-DNA triple helicesabstractMOTIVATION: Long non-coding RNA (lncRNA) plays a key role in many biological processes. For instance, lncRNA regulates chromatin using different molecular mechanisms, including direct RNA-DNA hybridization via triplexes, cotranscriptional RNA-RNA interactions, and RNA-DNA binding mediated by protein complexes. While the functional annotation of lncRNA transcripts has been widely studied over the last 20 years, barely a handful of tools have been developed with the specific purpose of detecting and evaluating lncRNA-DNA triple helices. What is worse, some of these tools have nearly grown a decade old, making new triplex-centric pipelines depend on legacy software that cannot thoroughly process all the data made available by next-generation sequencing (NGS) technologies. RESULTS: We present PATO, a modern, fast, and efficient tool for the detection of lncRNA-DNA triplexes that matches NGS processing capabilities. PATO enables the prediction of triple helices at the genome scale and can process in as little as 1 h more than 60 GB of sequence data using a two-socket server. Moreover, PATO's efficiency allows a more exhaustive search of the triplex-forming solution space, and so PATO achieves higher levels of prediction accuracy in far less time than other tools in the state of the art. AVAILABILITY AND IMPLEMENTATION: Source code, user manual, and tests are freely available to download under the MIT License at https://github.com/UDC-GAC/pato. Iñaki Amatria-Barral, Jorge González-Domínguez, Juan Touriño |
Bioinform. | 2 |
| 2023 | pRIblast: A highly efficient parallel application for comprehensive lncRNA-RNA interaction predictionabstractLong non-coding RNAs (lncRNAs) play a key role in several biological processes and scientists are constantly trying to come up with new strategies to elucidate their functions. One common approach to characterize these sequences consists in predicting their interactions with other RNA fragments. Nevertheless, the high computational cost of the bioinformatics tools developed for this purpose prevents their application to large-scale datasets. This paper presents pRIblast, a highly efficient parallel application for comprehensive lncRNA–RNA interaction prediction based on the state-of-the-art RIblast tool, which has been proved to show superior biological accuracy compared to other counterparts in previous experimental evaluations. Benchmarking on a multicore CPU cluster shows that pRIblast is able to compute in a few hours analyses that would need more than three months to complete with the original RIblast algorithm, always achieving the same level of prediction accuracy. Furthermore, this novel application can process large input datasets that cannot be processed with the former tool. pRIblast is free software publicly available to download at https://github.com/UDC-GAC/pRIblast under the MIT license. Iñaki Amatria-Barral, Jorge González-Domínguez, Juan Touriño |
Future Gener. Comput. Syst. | 2 |
| 2023 | ParRADMeth: Identification of Differentially Methylated Regions on Multicore ClustersabstractThe discovery of Differentially Methylated (DM) regions is an important research field in biology, as it can help to anticipate the risk of suffering from specific diseases. Nevertheless, the high computational cost of the bioinformatic tools developed for this purpose prevents their application to large-scale datasets. Hence, much faster tools are required to further progress in this research field. In this work we present ParRADMeth, a parallel tool that applies beta-binomial regression for the identification of these DM regions. It is based on the state-of-the-art sequential tool RADMeth, which proved superior biological accuracy compared to counterparts in previous experimental evaluations. ParRADMeth provides the same DM regions as RADMeth but at significantly reduced runtime thanks to exploiting the compute capabilities of common multicore CPU clusters. For example, our tool is up to 189 times faster for real data experiments on a cluster with 16 nodes, each one containing two eight-core processors. The source code of ParRADMeth, as well as a reference manual, are available at https://github.com/UDC-GAC/ParRADMeth. Alejandro Fernández-Fraga, Jorge González-Domínguez, Juan Touriño |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2022 | PyToxo: a Python tool for calculating penetrance tables of high-order epistasis modelsabstractBACKGROUND: Epistasis is the interaction between different genes when expressing a certain phenotype. If epistasis involves more than two loci it is called high-order epistasis. High-order epistasis is an area under active research because it could be the cause of many complex traits. The most common way to specify an epistasis interaction is through a penetrance table. RESULTS: This paper presents PyToxo, a Python tool for generating penetrance tables from any-order epistasis models. Unlike other tools available in the bibliography, PyToxo is able to work with high-order models and realistic penetrance and heritability values, achieving high-precision results in a short time. In addition, PyToxo is distributed as open-source software and includes several interfaces to ease its use. CONCLUSIONS: PyToxo provides the scientific community with a useful tool to evaluate algorithms and methods that can detect high-order epistasis to continue advancing in the discovery of the causes behind complex diseases. Borja González-Seoane, Christian Ponte-Fernández, Jorge González-Domínguez, María J. Martín |
BMC Bioinform. | 3 |
| 2022 | A SIMD algorithm for the detection of epistatic interactions of any orderabstractEpistasis is a phenomenon in which a phenotype outcome is determined by the interaction of genetic variation at two or more loci and it cannot be attributed to the additive combination of effects corresponding to the individual loci. Although it has been more than 100 years since William Bateson introduced this concept, it still is a topic under active research. Locating epistatic interactions is a computationally expensive challenge that involves analyzing an exponentially growing number of combinations. Authors in this field have resorted to a multitude of hardware architectures in order to speed up the search, but little to no attention has been paid to the vector instructions that current CPUs include in their instruction sets. This work extends an existing third-order exhaustive algorithm to support the search of epistasis interactions of any order and discusses multiple SIMD implementations of the different functions that compose the search using Intel AVX Intrinsics. Results using the GCC and the Intel compiler show that the 512-bit explicit vector implementation proposed here performs the best out of all of the other implementations evaluated. The proposed 512-bit vectorization accelerates the original implementation of the algorithm by an average factor of 7 and 12, for GCC and the Intel Compiler, respectively, in the scenarios tested. Christian Ponte-Fernández, Jorge González-Domínguez, María J. Martín |
Future Gener. Comput. Syst. | 2 |
| 2022 | Parallel-FST: A feature selection library for multicore clustersabstractFeature selection is a subfield of machine learning focused on reducing the dimensionality of datasets by performing a computationally intensive process. This work presents Parallel-FST, a publicly available parallel library for feature selection that includes seven methods which follow a hybrid MPI/multithreaded approach to reduce their runtime when executed on high performance computing systems. Performance tests were carried out on a 256-core cluster, where Parallel-FST obtained speedups of up to 229x for representative datasets and it was able to analyze a 512 GB dataset, which was not previously possible with a sequential counterpart library due to memory constraints. Bieito Beceiro, Jorge González-Domínguez, Juan Touriño |
J. Parallel Distributed Comput. | 2 |
| 2022 | Evaluation of Existing Methods for High-Order Epistasis DetectionabstractFinding epistatic interactions among loci when expressing a phenotype is a widely employed strategy to understand the genetic architecture of complex traits in GWAS. The abundance of methods dedicated to the same purpose, however, makes it increasingly difficult for scientists to decide which method is more suitable for their studies. This work compares the different epistasis detection methods published during the last decade in terms of runtime, detection power and type I error rate, with a special emphasis on high-order interactions. Results show that in terms of detection power, the only methods that perform well across all experiments are the exhaustive methods, although their computational cost may be prohibitive in large-scale studies. Regarding non-exhaustive methods, not one could consistently find epistasis interactions when marginal effects are absent. If marginal effects are present, there are methods that perform well for high-order interactions, such as BADTrees, FDHE-IW, SingleMI or SNPHarvester. As for false-positive control, only SNPHarvester, FDHE-IW and DCHE show good results. The study concludes that there is no single epistasis detection method to recommend in all scenarios. Authors should prioritize exhaustive methods when sufficient computational resources are available considering the data set size, and resort to non-exhaustive methods when the analysis time is prohibitive. Christian Ponte-Fernández, Jorge González-Domínguez, Antonio Carvajal-Rodríguez, María J. Martín |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2022 | MPI-dot2dot: A parallel tool to find DNA tandem repeats on multicore clustersabstractAbstract Tandem Repeats (TRs) are segments that occur several times in a DNA sequence, and each copy is adjacent to other. In the last few years, TRs have gained significant attention as they are thought to be related with certain human diseases. Therefore, identifying and classifying TRs have become a highly important task in bioinformatics in order to analyze their disorders and relationships with illnesses. Dot2dot, a tool recently developed to find TRs, provides more accurate results than the previous state-of-the-art, but it requires a long execution time even when using multiple threads. This work presents MPI-dot2dot, a novel version of this tool that combines MPI and OpenMP so that it can be executed in a cluster of multicore nodes and thus reduces its execution time. The performance of this new parallel implementation has been tested using different real datasets. Depending on the characteristics of the input genomes, it is able to obtain the same biological results as Dot2dot but more than 100 times faster on a 16-node multicore cluster (384 cores). MPI-dot2dot is publicly available to download from https://sourceforge.net/projects/mpi-dot2dot . Jorge González-Domínguez, José M. Martín-Martínez, Roberto R. Expósito |
J. Supercomput. | 1 |
| 2022 | Fiuncho: a program for any-order epistasis detection in CPU clustersabstractAbstract Epistasis can be defined as the statistical interaction of genes during the expression of a phenotype. It is believed that it plays a fundamental role in gene expression, as individual genetic variants have reported a very small increase in disease risk in previous Genome-Wide Association Studies. The most successful approach to epistasis detection is the exhaustive method, although its exponential time complexity requires a highly parallel implementation in order to be used. This work presents Fiuncho, a program that exploits all levels of parallelism present in x86_64 CPU clusters in order to mitigate the complexity of this approach. It supports epistasis interactions of any order, and when compared with other exhaustive methods, it is on average 358, 7 and 3 times faster than MDR, MPI3SNP and BitEpi, respectively. Christian Ponte-Fernández, Jorge González-Domínguez, María J. Martín |
J. Supercomput. | 2 |
| 2020 | Toxo: a library for calculating penetrance tables of high-order epistasis modelsabstractBACKGROUND: Epistasis is defined as the interaction between different genes when expressing a specific phenotype. The most common way to characterize an epistatic relationship is using a penetrance table, which contains the probability of expressing the phenotype under study given a particular allele combination. Available simulators can only create penetrance tables for well-known epistasis models involving a small number of genes and under a large number of limitations. RESULTS: Toxo is a MATLAB library designed to calculate penetrance tables of epistasis models of any interaction order which resemble real data more closely. The user specifies the desired heritability (or prevalence) and the program maximizes the table's prevalence (or heritability) according to the input epistatic model boundaries. CONCLUSIONS: Toxo extends the capabilities of existing simulators that define epistasis using penetrance tables. These tables can be directly used as input for software simulators such as GAMETES so that they are able to generate data samples with larger interactions and more realistic prevalences/heritabilities. Christian Ponte-Fernández, Jorge González-Domínguez, Antonio Carvajal-Rodríguez, María J. Martín |
BMC Bioinform. | 2 |
| 2020 | SMusket: Spark-based DNA error correction on distributed-memory systems
Roberto R. Expósito, Jorge González-Domínguez, Juan Touriño |
Future Gener. Comput. Syst. | 2 |
| 2020 | CUDA-JMI: Acceleration of feature selection on heterogeneous systems
Jorge González-Domínguez, Roberto R. Expósito, Verónica Bolón-Canedo |
Future Gener. Comput. Syst. | 1 |
| 2019 | Accelerating binary biclustering on platforms with CUDA-enabled GPUs
Jorge González-Domínguez, Roberto R. Expósito |
Inf. Sci. | 1 |
| 2019 | Parallel feature selection for distributed-memory clusters
Jorge González-Domínguez, Verónica Bolón-Canedo, Borja Freire, Juan Touriño |
Inf. Sci. | 1 |
| 2018 | MPIGeneNet: Parallel Calculation of Gene Co-Expression Networks on Multicore ClustersabstractIn this work, we present MPIGeneNet, a parallel tool that applies Pearson's correlation and Random Matrix Theory to construct gene co-expression networks. It is based on the state-of-the-art sequential tool RMTGeneNet, which provides networks with high robustness and sensitivity at the expenses of relatively long runtimes for large scale input datasets. MPIGeneNet returns the same results as RMTGeneNet but improves the memory management, reduces the I/O cost, and accelerates the two most computationally demanding steps of co-expression network construction by exploiting the compute capabilities of common multicore CPU clusters. Our performance evaluation on two different systems using three typical input datasets shows that MPIGeneNet is significantly faster than RMTGeneNet. As an example, our tool is up to 175.41 times faster on a cluster with eight nodes, each one containing two 12-core Intel Haswell processors. The source code of MPIGeneNet, as well as a reference manual, are available at https://sourceforge.net/projects/mpigenenet/. Jorge González-Domínguez, María J. Martín |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2017 | MarDRe: efficient MapReduce-based removal of duplicate DNA reads in the cloudabstractSUMMARY: This article presents MarDRe, a de novo cloud-ready duplicate and near-duplicate removal tool that can process single- and paired-end reads from FASTQ/FASTA datasets. MarDRe takes advantage of the widely adopted MapReduce programming model to fully exploit Big Data technologies on cloud-based infrastructures. Written in Java to maximize cross-platform compatibility, MarDRe is built upon the open-source Apache Hadoop project, the most popular distributed computing framework for scalable Big Data processing. On a 16-node cluster deployed on the Amazon EC2 cloud platform, MarDRe is up to 8.52 times faster than a representative state-of-the-art tool. AVAILABILITY AND IMPLEMENTATION: Source code in Java and Hadoop as well as a user's guide are freely available under the GNU GPLv3 license at http://mardre.des.udc.es . CONTACT: [email protected]. Roberto R. Expósito, Jorge Veiga, Jorge González-Domínguez, Juan Touriño |
Bioinform. | 3 |
| 2016 | Combining GPU and FPGA technology for efficient exhaustive interaction analysis in GWASabstractInteraction between genes has become a major topic in quantitative genetics. It is believed that these interactions play a significant role in genetic variations causing complex diseases. Due to the number of tests required for an exhaustive search in genome-wide association studies (GWAS), a large amount of computational power is required. In this paper, we present a hybrid architecture consisting of tightly interconnected CPUs, GPUs and FPGAs and a fine-tuned software suite to outperform other implementations in pairwise interaction analysis while consuming less than 300Watts and fitting into a standard desktop computer case. Jan Christian Kässens, Lars Wienbrandt, Manfred Schimmler, Jorge González-Domínguez, Bertil Schmidt |
ASAP | 4 |
| 2016 | ParDRe: faster parallel duplicated reads removal tool for sequencing studiesabstractUNLABELLED: Current next generation sequencing technologies often generate duplicated or near-duplicated reads that (depending on the application scenario) do not provide any interesting biological information but can increase memory requirements and computational time of downstream analysis. In this work we present ParDRe, a de novo parallel tool to remove duplicated and near-duplicated reads through the clustering of Single-End or Paired-End sequences from fasta or fastq files. It uses a novel bitwise approach to compare the suffixes of DNA strings and employs hybrid MPI/multithreading to reduce runtime on multicore systems. We show that ParDRe is up to 27.29 times faster than Fulcrum (a representative state-of-the-art tool) on a platform with two 8-core Sandy-Bridge processors. AVAILABILITY AND IMPLEMENTATION: Source code in C ++ and MPI running on Linux systems as well as a reference manual are available at https://sourceforge.net/projects/pardre/ CONTACT: [email protected]. Jorge González-Domínguez, Bertil Schmidt |
Bioinform. | 1 |
| 2016 | MSAProbs-MPI: parallel multiple sequence aligner for distributed-memory systemsabstractMSAProbs is a state-of-the-art protein multiple sequence alignment tool based on hidden Markov models. It can achieve high alignment accuracy at the expense of relatively long runtimes for large-scale input datasets. In this work we present MSAProbs-MPI, a distributed-memory parallel version of the multithreaded MSAProbs tool that is able to reduce runtimes by exploiting the compute capabilities of common multicore CPU clusters. Our performance evaluation on a cluster with 32 nodes (each containing two Intel Haswell processors) shows reductions in execution time of over one order of magnitude for typical input datasets. Furthermore, MSAProbs-MPI using eight nodes is faster than the GPU-accelerated QuickProbs running on a Tesla K20. Another strong point is that MSAProbs-MPI can deal with large datasets for which MSAProbs and QuickProbs might fail due to time and memory constraints, respectively. AVAILABILITY AND IMPLEMENTATION: Source code in C ++ and MPI running on Linux systems as well as a reference manual are available at http://msaprobs.sourceforge.net CONTACT: [email protected] information: Supplementary data are available at Bioinformatics online. Jorge González-Domínguez, Yongchao Liu 0004, Juan Touriño, Bertil Schmidt |
Bioinform. | 1 |
| 2016 | Parallel Pairwise Epistasis Detection on Heterogeneous Computing ArchitecturesabstractDevelopment of new methods to detect pairwise epistasis, such as SNP-SNP interactions, in Genome-Wide Association Studies is an important task in bioinformatics as they can help to explain genetic influences on diseases. As these studies are time consuming operations, some tools exploit the characteristics of different hardware accelerators (such as GPUs and Xeon Phi coprocessors) to reduce the runtime. Nevertheless, all these approaches are not able to efficiently exploit the whole computational capacity of modern clusters that contain both GPUs and Xeon Phi coprocessors. In this paper we investigate approaches to map pairwise epistasic detection on heterogeneous clusters using both types of accelerators. The runtimes to analyze the well-known WTCCC dataset consisting of about 500 K SNPs and 5 K samples on one and two NVIDIA K20m are reduced by 27 percent thanks to the use of a hybrid approach with one additional single Xeon Phi coprocessor. Jorge González-Domínguez, Sabela Ramos, Juan Touriño, Bertil Schmidt |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2015 | Parallelizing Epistasis Detection in GWAS on FPGA and GPU-Accelerated Computing SystemsabstractHigh-throughput genotyping technologies (such as SNP-arrays) allow the rapid collection of up to a few million genetic markers of an individual. Detecting epistasis (based on 2-SNP interactions) in Genome-Wide Association Studies is an important but time consuming operation since statistical computations have to be performed for each pair of measured markers. Computational methods to detect epistasis therefore suffer from prohibitively long runtimes; e.g., processing a moderately-sized dataset consisting of about 500,000 SNPs and 5,000 samples requires several days using state-of-the-art tools on a standard 3 GHz CPU. In this paper, we demonstrate how this task can be accelerated using a combination of fine-grained and coarse-grained parallelism on two different computing systems. The first architecture is based on reconfigurable hardware (FPGAs) while the second architecture uses multiple GPUs connected to the same host. We show that both systems can achieve speedups of around four orders-of-magnitude compared to the sequential implementation. This significantly reduces the runtimes for detecting epistasis to only a few minutes for moderately-sized datasets and to a few hours for large-scale datasets. Jorge González-Domínguez, Lars Wienbrandt, Jan Christian Kässens, David Ellinghaus, Manfred Schimmler, Bertil Schmidt |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2014 | UPC++ for bioinformatics: A case study using genome-wide association studiesabstractModern genotyping technologies are able to obtain up to a few million genetic markers (such as SNPs) of an individual within a few minutes of time. Detecting epistasis, such as SNP-SNP interactions, in Genome-Wide Association Studies is an important but time-consuming operation since statistical computations have to be performed for each pair of measured markers. Therefore, a variety of HPC architectures have been used to accelerate these studies. In this work we present a parallel approach for multi-core clusters, which is implemented with UPC++ and takes advantage of the features available in the Partitioned Global Address Space and Object Oriented Programming models. Our solution is based on a well-known regression model (used by the popular BOOST tool) to test SNP-pairs interactions. Experimental results show that UPC++ is suitable for parallelizing data-intensive bioinformatics applications on clusters. For instance, it reduces the time to analyze a real-world dataset with more than 500,000 SNPs and 5,000 individuals from several days when using a single core to less than one minute using 512 nodes (12,288 cores) of a Cray XC30 supercomputer. Jan Christian Kässens, Jorge González-Domínguez, Lars Wienbrandt, Bertil Schmidt |
CLUSTER | 2 |
| 2014 | Hybrid CPU/GPU Acceleration of Detection of 2-SNP Epistatic Interactions in GWAS
Jorge González-Domínguez, Bertil Schmidt, Jan Christian Kässens, Lars Wienbrandt |
Euro-Par | 1 |
| 2014 | Analyzing the Energy Efficiency of the Memory Subsystem in Multicore ProcessorsabstractIn this paper we analyze the energy overhead incurred when operating with data stored in different levels of the memory subsystem (cache levels and DDR chips) of current multicore architectures. Our approach builds upon servet, a portable framework for the memory characterization of multicore processors, extending this suite with a power-related test that, when applied to a platform equipped with a power measurement mechanism, provides information on the efficiency of memory energy usage. As additional contributions, i) we provide a complete experimental study of the impact that the CPU performance states (also known as P-states) exert on the memory energy efficiency of a collection of recent server-oriented and low-power cores, and ii) we show how this framework carries over to cover also the scalability analysis of the memory energy performance on multicore processors. Sandra Catalán, Jorge González-Domínguez, Rafael Mayo 0002, Enrique S. Quintana-Ortí |
ISPA | 2 |
| 2014 | A 2D algorithm with asymmetric workload for the UPC conjugate gradient method
Jorge González-Domínguez, Osni Marques, María J. Martín, Juan Touriño |
J. Supercomput. | 1 |
| 2013 | Analysis of I/O Performance on an Amazon EC2 Cluster Compute and High I/O Platform
Roberto R. Expósito, Guillermo L. Taboada, Sabela Ramos, Jorge González-Domínguez, Juan Touriño, Ramón Doallo |
J. Grid Comput. | 4 |
| 2013 | Performance evaluation of sparse matrix products in UPC
Jorge González-Domínguez, Óscar García-López, Guillermo L. Taboada, María J. Martín, Juan Touriño |
J. Supercomput. | 1 |
| 2012 | Design and Performance Issues of Cholesky and LU Solvers Using UPCBLASabstractPartitioned Global Address Space (PGAS) languages offer programmers a shared memory view that increases their productivity and allow locality exploitation to obtain good performance on current large-scale distributed memory systems. UPCBLAS is a parallel numerical library for dense matrix computations using the PGAS Unified Parallel C (UPC) language. The interface of this library exploits the characteristics of the PGAS memory model and thus it is easier to use than MPI-based libraries. This paper addresses the implementation of solvers of systems of equations through Cholesky and LU factorizations in UPC using UPCBLAS. The developed codes are experimentally evaluated and compared to the MPI versions using ScaLAPACK. Parallel solvers of equations are present in many parallel numerical applications and they have been traditionally developed in MPI. This work shows that UPCBLAS can be considered as a good alternative to the MPI-based libraries for increasing the productivity of numerical application developers. Jorge González-Domínguez, Osni Marques, María J. Martín, Guillermo L. Taboada, Juan Touriño |
ISPA | 1 |
| 2012 | Communication avoiding and overlapping for numerical linear algebraabstractTo efficiently scale dense linear algebra problems to future exascale systems, communication cost must be avoided or overlapped. Communication-avoiding 2.5D algorithms improve scalability by reducing inter-processor data transfer volume at the cost of extra memory usage. Communication overlap attempts to hide messaging latency by pipelining messages and overlapping with computational work. We study the interaction and compatibility of these two techniques for two matrix multiplication algorithms (Cannon and SUMMA), triangular solve, and Cholesky factorization. For each algorithm, we construct a detailed performance model that considers both critical path dependencies and idle time. We give novel implementations of 2.5D algorithms with overlap for each of these problems. Our software employs UPC, a partitioned global address space (PGAS) language that provides fast one-sided communication. We show communication avoidance and overlap provide a cumulative benefit as core counts scale, including results using over 24K cores of a Cray XE6 system. Evangelos Georganas, Jorge González-Domínguez, Edgar Solomonik, Yili Zheng, Juan Touriño, Katherine A. Yelick |
SC | 2 |
| 2012 | UPCBLAS: a library for parallel matrix computations in Unified Parallel CabstractSUMMARY The popularity of Partitioned Global Address Space (PGAS) languages has increased during the last years thanks to their high programmability and performance through an efficient exploitation of data locality, especially on hierarchical architectures such as multicore clusters. This paper describes UPCBLAS, a parallel numerical library for dense matrix computations using the PGAS Unified Parallel C language. The routines developed in UPCBLAS are built on top of sequential basic linear algebra subprograms functions and exploit the particularities of the PGAS paradigm, taking into account data locality in order to achieve a good performance. Furthermore, the routines implement other optimization techniques, several of them by automatically taking into account the hardware characteristics of the underlying systems on which they are executed. The library has been experimentally evaluated on a multicore supercomputer and compared with a message‐passing‐based parallel numerical library, demonstrating good scalability and efficiency. Copyright © 2012 John Wiley & Sons, Ltd. Jorge González-Domínguez, María J. Martín, Guillermo L. Taboada, Juan Touriño, Ramón Doallo, Damián A. Mallón, Brian Wibecan |
Concurr. Comput. Pract. Exp. | 1 |
| 2010 | Servet: A benchmark suite for autotuning on multicore clustersabstractThe growing complexity in computer system hierarchies due to the increase in the number of cores per processor, levels of cache (some of them shared) and the number of processors per node, as well as the high-speed interconnects, demands the use of new optimization techniques and libraries that take advantage of their features. In this paper Servet, a suite of benchmarks focused on detecting a set of parameters with high influence in the overall performance of multicore systems, is presented. These benchmarks are able to detect the cache hierarchy, including their size and which caches are shared by each core, bandwidths and bottlenecks in memory accesses, as well as communication latencies among cores. These parameters can be used by auto-tuned codes to increase their performance in multicore clusters. Experimental results using different representative systems show that Servet provides very accurate estimates of the parameters of the machine architecture. Jorge González-Domínguez, Guillermo L. Taboada, Basilio B. Fraguela, María J. Martín, Juan Touriño |
IPDPS | 1 |
| 2009 | A Parallel Numerical Library for UPC
Jorge González-Domínguez, María J. Martín, Guillermo L. Taboada, Juan Touriño, Ramón Doallo, Andrés Gómez 0002 |
Euro-Par | 1 |