VLDB 2026 Research / reviewers in the wild / expert
Manuel F. Dolz
dblp:88/7977
· DBLP profile ↗
43ranked-venue papers
6as first author
17since 2021 · last 2026
0000-0001-9466-3398ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 32 · 4 first-author · 11 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Cross-platform characterisation and performance analysis of homomorphic matrix multiplicationabstractAbstract Fully Homomorphic Encryption (FHE) enables computation over encrypted data while preserving strong security and privacy guarantees. However, its high computational cost remains a major challenge. This study therefore evaluates the performance of homomorphic matrix multiplication using three Homomorphic Encryption (HE) libraries, Microsoft SEAL, HElib and OpenFHE, across two platforms with AMD EPYC and Intel Xeon CPUs, with a particular focus on the impact of different compiler flags. The results indicate that compiler configurations and hardware selection significantly influence runtime in libraries such as Microsoft SEAL, whereas HElib and OpenFHE show negligible variation under different compilation settings. Furthermore, the analysis reveals that SEAL makes more efficient use of memory bandwidth while OpenFHE achieves the highest overall performance, the lowest mean absolute error and the shortest execution time. Franklin Espinoza, Justo Molina, Darwin Quezada-Gaibor, Sandra Catalán, Manuel F. Dolz |
J. Supercomput. | 5 |
| 2025 | Sinusoidal Initialization, Time for a New StartabstractInitialization plays a critical role in Deep Neural Network training, directly influencing convergence, stability, and generalization. Common approaches such as Glorot and He initializations rely on randomness, which can produce uneven weight distributions across layer connections. In this paper, we introduce the Sinusoidal initialization, a novel deterministic method that employs sinusoidal functions to construct structured weight matrices expressly to improve the spread and balance of weights throughout the network while simultaneously fostering a more uniform, well‑conditioned distribution of neuron activation states from the very first forward pass. Because Sinusoidal initialization begins with weights and activations that are already evenly and efficiently utilized, it delivers consistently faster convergence, greater training stability, and higher final accuracy across a wide range of models, including convolutional neural networks, vision transformers, and large language models. On average, our experiments show an increase of 4.8 % in final validation accuracy and 20.9 % in convergence speed. By replacing randomness with structure, this initialization provides a stronger and more reliable foundation for Deep Learning systems. Alberto Fernández-Hernández, José I. Mestre, Manuel F. Dolz, José Duato, Enrique S. Quintana-Ortí |
NeurIPS | 3 |
| 2025 | Deep learning inference optimisation for IoT: Conv2D-ReLU-BN layer fusion and quantisationabstractAbstract The deployment of deep learning models on resource-constrained devices requires the development of new optimisation techniques to effectively exploit the computational and storage capacities of these devices. Thus, the primary objective of this research is to introduce an innovative and efficient approach for fusing convolution (or fully connected), ReLU, and batch normalisation neural network layers into a unified, single-layer structure, alongside a quantisation method for this new fused layer. This approach has been evaluated using the Arduino BLE Sense ARM Cortex-M4 and the Arduino Portenta H7 Lite ARM Cortex-M4 and M7 processors, known for their widespread adoption in various Internet of Things devices. Depending on the microcontroller unit and compilation flag used, the fused layers can reduce the overall execution time by up to 1.53 $$\times$$ × , and on individual layers it can reach a speedup of 2.95 $$\times$$ × . José I. Mestre, Sergio Barrachina 0001, Darwin Quezada, Manuel F. Dolz |
J. Supercomput. | 4 |
| 2024 | Optimizing Convolutions for Deep Learning Inference on ARM Cortex-M ProcessorsabstractWe perform a series of optimisations on the convolution operator within the ARM CMSIS-NN library to improve the performance of deep learning tasks on Arduino development boards equipped with ARM Cortex-M4 and M7 microcontrollers. To this end, we develop custom microkernels that efficiently handle the internal computations required by the convolution operator via the lowering approach and the direct method, and we design two techniques to avoid register spilling. We also take advantage of all the RAM on the Arduino boards by reusing it as a scratchpad for the convolution filters. The integration of these techniques into CMSIS-NN, when invoked by TensorFlow Lite for microcontrollers for quantised versions of VGG, SqueezeNet, ResNet, and MobileNet-like convolutional neural networks enhances the overall inference speed by a factor ranging from 1.13× to 1.50×. Antonio Maciá-Lillo, Sergio Barrachina 0001, Germán Fabregat, Manuel F. Dolz |
IEEE Internet Things J. | 4 |
| 2024 | Automatic generation of ARM NEON micro-kernels for matrix multiplicationabstractAbstract General matrix multiplication ( gemm ) is a fundamental kernel in scientific computing and current frameworks for deep learning. Modern realisations of gemm are mostly written in C, on top of a small, highly tuned micro-kernel that is usually encoded in assembly. The high performance realisation of gemm in linear algebra libraries in general include a single micro-kernel per architecture, usually implemented by an expert. In this paper, we explore a couple of paths to automatically generate gemm micro-kernels, either using C++ templates with vector intrinsics or high-level Python scripts that directly produce assembly code. Both solutions can integrate high performance software techniques, such as loop unrolling and software pipelining, accommodate any data type, and easily generate micro-kernels of any requested dimension. The performance of this solution is tested on three ARM-based cores and compared with state-of-the-art libraries for these processors: BLIS, OpenBLAS and ArmPL. The experimental results show that the auto-generation approach is highly competitive, mainly due to the possibility of adapting the micro-kernel to the problem dimensions. Guillermo Alaejos, Héctor Martínez 0002, Adrián Castelló 0001, Manuel F. Dolz, Francisco D. Igual, Pedro Alonso 0002, Enrique S. Quintana-Ortí |
J. Supercomput. | 4 |
| 2024 | Urban sound classification using neural networks on embedded FPGAsabstractAbstract Sound classification using neural networks has recently produced very accurate results. A large number of different applications use this type of sound classifiers such as controlling and monitoring the type of activity in a city or identifying different types of animals in natural environments. While traditional acoustic processing applications have been developed on high-performance computing platforms equipped with expensive multi-channel audio interfaces, the Internet of Things (IoT) paradigm requires the use of more flexible and energy-efficient systems. Although software-based platforms exist for implementing general-purpose neural networks, they are not optimized for sound classification, wasting energy and computational resources. In this work, we have used FPGAs to develop an ad hoc system where only the hardware needed for our application is synthesized, resulting in faster and more energy-efficient circuits. The results show that our developments are accelerated by a factor of 35 compared to a software-based implementation on a Raspberry Pi. Jose A. Belloch, Raul Coronado, Óscar Valls, Rocío del Amor, German Leon, Valery Naranjo, Manuel F. Dolz, Adrian Amor-Martin, Gema Piñero |
J. Supercomput. | 7 |
| 2023 | Reformulating the direct convolution for high-performance deep learning inference on ARM processorsabstractWe present two high-performance implementations of the convolution operator via the direct algorithm that outperform the so-called lowering approach based on the im2col transform plus the gemm kernel on an ARMv8-based processor. One of our methods presents the additional advantage of zero-memory overhead while the other employs an additional yet rather moderate workspace, substantially smaller than that required by the im2col+gemm solution. In contrast with a previous implementation of a similar zero-memory overhead direct convolution, this work exhibits the key advantage of preserving the conventional NHWC data layout for the input/output activations of the convolution layers. Sergio Barrachina 0001, Adrián Castelló 0001, Manuel F. Dolz, Tze Meng Low, Héctor Martínez 0002, Enrique S. Quintana-Ortí, Upasana Sridhar, Andrés Tomás |
J. Syst. Archit. | 3 |
| 2023 | Performance-energy trade-offs of deep learning convolution algorithms on ARM processorsabstractAbstract In this work, we assess the performance and energy efficiency of high-performance codes for the convolution operator, based on the direct, explicit/implicit lowering and Winograd algorithms used for deep learning (DL) inference on a series of ARM-based processor architectures. Specifically, we evaluate the NVIDIA Denver2 and Carmel processors, as well as the ARM Cortex-A57 and Cortex-A78AE CPUs as part of a recent set of NVIDIA Jetson platforms. The performance–energy evaluation is carried out using the ResNet-50 v1.5 convolutional neural network (CNN) on varying configurations of convolution algorithms, number of threads/cores, and operating frequencies on the tested processor cores. The results demonstrate that the best throughput is obtained on all platforms with the Winograd convolution operator running on all the cores at their highest frequency. However, if the goal is to reduce the energy footprint, there is no rule of thumb for the optimal configuration. Manuel F. Dolz, Sergio Barrachina 0001, Héctor Martínez 0002, Adrián Castelló 0001, Antonio M. Vidal, Germán Fabregat, Andrés Tomás |
J. Supercomput. | 1 |
| 2023 | Efficient and portable Winograd convolutions for multi-core processorsabstractAbstract We take a step forward towards developing high-performance codes for the convolution operator, based on the Winograd algorithm, that are easy to customise for general-purpose processor architectures. In our approach, augmenting the portability of the solution is achieved via the introduction of vector instructions from Intel SSE/AVX2/AVX512 and ARM NEON/SVE to exploit the single-instruction multiple-data capabilities of current processors as well as OpenMP pragmas to exploit multi-threaded parallelism. While this comes at the cost of sacrificing a fraction of the computational performance, our experimental results on three distinct processors, with Intel Xeon Skylake, ARM Cortex A57 and Fujitsu A64FX processors, show that the impact is affordable and still renders a Winograd-based solution that is competitive when compared with the lowering gemm-based convolution. Manuel F. Dolz, Héctor Martínez 0002, Adrián Castelló 0001, Pedro Alonso 0002, Enrique S. Quintana-Ortí |
J. Supercomput. | 1 |
| 2022 | Towards Portable Realizations of Winograd-based Convolution with Vector Intrinsics and OpenMPabstractWe take a step forward in the direction of developing high performance codes for the convolution, based on the Winograd transformation, that are easy to customize for different processor architectures. In our approach, augmenting the portability of the solution is achieved via the introduction of vector intrinsics to exploit the SIMD (single-instruction multiple-data) capabilities of current processors as well as OpenMP pragmas to exploit multi-thread parallelism. While this comes at the cost of sacrificing a fraction of the computational performance, our experimental results on two distinct processors, with Intel Xeon Skylake and ARM Cortex A57 architectures, show that the impact is affordable, and still renders a Winograd-based solution that is competitive with the general method for the convolution based on the so-called im2col transform followed by a matrix-matrix multiplication. Manuel F. Dolz, Adrián Castelló 0001, Enrique S. Quintana-Ortí |
PDP | 1 |
| 2022 | Convolution Operators for Deep Learning Inference on the Fujitsu A64FX ProcessorabstractThe convolution operator is a crucial kernel for many computer vision and signal processing applications that rely on deep learning (DL) technologies. As such, the efficient implementation of this operator has received considerable attention in the past few years for a fair range of processor architectures. In this paper, we follow the technology trend toward integrating long SIMD (single instruction, multiple data) arithmetic units into high performance multicore processors to analyse the benefits of this type of hardware acceleration for latency-constrained DL workloads. For this purpose, we implement and optimise for the Fujitsu processor A64FX, three distinct methods for the calculation of the convolution, namely, the lowering approach, a blocked variant of the direct convolution algorithm, and the Winograd minimal filtering algorithm. Our experimental results include an extensive evaluation of the parallel scalability of these three methods and a comparison of their global performance using three popular DL models and a representative dataset. Manuel F. Dolz, Héctor Martínez 0002, Pedro Alonso 0002, Enrique S. Quintana-Ortí |
SBAC-PAD | 1 |
| 2022 | Efficient and portable GEMM-based convolution operators for deep neural network training on multicore processorsabstractConvolutional Neural Networks (CNNs) play a crucial role in many image recognition and classification tasks, recommender systems, brain-computer interfaces, etc. As a consequence, there is a notable interest in developing high performance realizations of the convolution operators, which concentrate a significant portion of the computational cost of this type of neural networks. In a previous work, we introduced a portable, high performance convolution algorithm, based on the BLIS realization of matrix multiplication, which eliminates most of the runtime and memory overheads that impair the performance of the convolution operators appearing in the forward training pass, when performed via explicit im2col transform. In this paper, we extend our ideas to the full training process of CNNs on multicore processors, proposing new high performance strategies to tackle the convolution operators that are present in the more complex backward pass of the training process, while maintaining the portability of the realizations. In addition, we conduct a full integration of these algorithms into a framework for distributed training of CNNs on clusters of computers, providing a complete experimental evaluation of the actual benefits in terms of both performance and memory consumption. Compared with baseline implementation, the use of the new convolution operators using pre-allocated memory can accelerate the training by a factor of about 6%–25%, provided there is sufficient memory available. In comparison, the operator variants that do not rely on persistent memory can save up to 70% of memory. Sergio Barrachina 0001, Manuel F. Dolz, Pablo San Juan, Enrique S. Quintana-Ortí |
J. Parallel Distributed Comput. | 2 |
| 2022 | High performance and energy efficient inference for deep learning on multicore ARM processors using general optimization techniques and BLISabstractWe evolve PyDTNN, a framework for distributed parallel training of Deep Neural Networks (DNNs), into an efficient inference tool for convolutional neural networks. Our optimization process on multicore ARM processors involves several high-level transformations of the original framework, such as the development and integration of Cython routines to exploit thread-level parallelism; the design and development of micro-kernels for the matrix multiplication, vectorized with ARM’s NEON intrinsics, that can accommodate layer fusion; and the appropriate selection of several cache configuration parameters tailored to the memory hierarchy of the target ARM processors. Our experiments evaluate both inference throughput (measured in processed images/s) and inference latency (i.e., time-to-response) as well as energy consumption per image when varying the level of thread parallelism and the processor power modes. The experiments with the new inference engine are reported for the ResNet50 v1.5 model on the ImageNet dataset from the MLPerf suite using the ARM v8.2 cores in the NVIDIA Jetson AGX Xavier board. These results show superior performance compared with the well-spread TFLite from Google and slightly inferior results when compared with ArmNN, the native library from ARM for DNN inference. Adrián Castelló 0001, Sergio Barrachina 0001, Manuel F. Dolz, Enrique S. Quintana-Ortí, Pau San Juan, Andrés Tomás |
J. Syst. Archit. | 3 |
| 2022 | BestOf: an online implementation selector for the training and inference of deep neural networksabstractAbstract Tuning and optimising the operations executed in deep learning frameworks is a fundamental task in accelerating the processing of deep neural networks (DNNs). However, this optimisation usually requires extensive manual efforts in order to obtain the best performance for each combination of tensor input size, layer type, and hardware platform. In this work, we present , a novel online auto-tuner that optimises the training and inference phases of DNNs. automatically selects at run time, and among the provided alternatives, the best performing implementation in each layer according to gathered profiling data. The evaluation of is performed on multi-core architectures for different DNNs using , a lightweight library for distributed training and inference. The experimental results reveal that the auto-tuner delivers the same or higher performance than that achieved using a static selection approach. Sergio Barrachina 0001, Adrián Castelló 0001, Manuel F. Dolz, Andrés Tomás |
J. Supercomput. | 3 |
| 2021 | Performance Modeling for Distributed Training of Convolutional Neural NetworksabstractWe perform a theoretical analysis comparing the scalability of data versus model parallelism, applied to the distributed training of deep convolutional neural networks (CNNs), along five axes: batch size, node (floating-point) arithmetic performance, node memory bandwidth, network link bandwidth, and cluster dimension. Our study relies on analytical performance models that can be configured to reproduce the components and organization of the CNN model as well as the hardware configuration of the target distributed platform. In addition, we provide evidence of the accuracy of the analytical models by performing a validation against a Python library for distributed deep learning training. Adrián Castelló 0001, Mar Catalán, Manuel F. Dolz, José I. Mestre, Enrique S. Quintana-Ortí, José Duato |
PDP | 3 |
| 2021 | Evaluation of MPI Allreduce for Distributed Training of Convolutional Neural NetworksabstractTraining deep neural networks is a costly procedure, often performed via sophisticated deep learning frameworks on clusters of computers. As faster processor technologies are integrated into these cluster facilities (e.g., NVIDIA's graphics accelerators or Google's tensor processing units), the communication component of the training process rapidly becomes a performance bottleneck. In this paper, we offer a complete analysis of the key collective communication primitive for the distributed data-parallel training of convolutional network networks (CNNs) focused on three relevant instances of the Message Passing Interface (MPI): MPICH, OpenMPI, and IntelMPI. In addition, our experimental evaluation is extended to expose the practical impact of this collective primitive when the training is performed using TensorFlow+ Horovod on a 16-node cluster. Finally, the theoretical analysis is further refined to a number of accelerated cluster configurations that are emulated by adjusting the communication-arithmetic ratio of the training process. Adrián Castelló 0001, Mar Catalán, Manuel F. Dolz, José I. Mestre, Enrique S. Quintana-Ortí, José Duato |
PDP | 3 |
| 2021 | PyDTNN: A user-friendly and extensible framework for distributed deep learning
Sergio Barrachina 0001, Adrián Castelló 0001, Mar Catalán, Manuel F. Dolz, José I. Mestre |
J. Supercomput. | 4 |
| 2020 | High Performance and Portable Convolution Operators for Multicore ProcessorsabstractThe considerable impact of Convolutional Neural Networks on many Artificial Intelligence tasks has led to the development of various high performance algorithms for the convolution operator present in this type of networks. One of these approaches leverages the IM2COL transform followed by a general matrix multiplication (GEMM) in order to take advantage of the highly optimized realizations of the GEMM kernel in many linear algebra libraries. The main problems of this approach are 1) the large memory workspace required to host the intermediate matrices generated by the IM2COL transform; and 2) the time to perform the IM2COL transform, which is not negligible for complex neural networks. This paper presents a portable high performance convolution algorithm based on the BLIS realization of the GEMM kernel that avoids the use of the intermediate memory by taking advantage of the BLIS structure. In addition, the proposed algorithm eliminates the cost of the explicit IM2COL transform, while maintaining the portability and performance of the underlying realization of GEMM in BLIS. Pablo San Juan, Adrián Castelló 0001, Manuel F. Dolz, Pedro Alonso 0002, Enrique S. Quintana-Ortí |
SBAC-PAD | 3 |
| 2020 | Performance modeling of the sparse matrix-vector product via convolutional neural networks
Maria Barreda, Manuel F. Dolz, M. Asunción Castaño, Pedro Alonso 0002, Enrique S. Quintana-Ortí |
J. Supercomput. | 2 |
| 2020 | Detecting semantic violations of lock-free data structures through C++ contracts
Javier López-Gómez, David del Rio Astorga, Manuel F. Dolz, Javier Fernández 0001, José Daniel García |
J. Supercomput. | 3 |
| 2019 | Theoretical Scalability Analysis of Distributed Deep Convolutional Neural NetworksabstractWe analyze the asymptotic performance of the training process of deep neural networks (NN) on clusters in order to determine the scalability. For this purpose, i) we assume a data parallel implementation of the training algorithm, which distributes the batches among the cluster nodes and replicates the model; ii) we leverage the roofline model to inspect the performance at the node level, taking into account the floating-point unit throughput and memory bandwidth; and iii) we consider distinct collective communication schemes that are optimal depending on the message size and underlying network interconnection topology. We then apply the resulting performance model to analyze the scalability of several well-known deep convolutional neural networks as a function of the batch size, node floating-point throughput, node memory bandwidth, cluster dimension, and link bandwidth. Adrián Castelló 0001, Manuel F. Dolz, Enrique S. Quintana-Ortí, José Duato |
CCGRID | 2 |
| 2019 | Analysis of model parallelism for distributed neural networksabstractWe analyze the performance of model parallelism applied to the training of deep neural networks on clusters. For this study, we elaborate a parameterized analytical performance model that captures the main computational and communication stages in distributed model parallel training. This model is then leveraged to assess the impact on the performance of four representative convolutional neural networks (CNNs) when varying the node throughput in terms of operations per second and memory bandwidth, the number of nodes of the cluster, the bandwidth of the network links, and algorithmic parameters such as the dimension of the batch. Adrián Castelló 0001, Manuel F. Dolz, Enrique S. Quintana-Ortí, José Duato |
EuroMPI | 2 |
| 2019 | Exploring stream parallel patterns in distributed MPI environments
Javier López-Gómez, Javier Fernández 0001, David del Rio Astorga, Manuel F. Dolz, José Daniel García |
Parallel Comput. | 4 |
| 2019 | Hybrid static-dynamic selection of implementation alternatives in heterogeneous environments
David del Rio Astorga, Manuel F. Dolz, Javier Fernández 0001, Francisco Javier García Blas |
J. Supercomput. | 2 |
| 2019 | A pipeline structure for the block QR update in digital signal processing
Manuel F. Dolz, Fran J. Alventosa, Pedro Alonso 0002, Antonio M. Vidal |
J. Supercomput. | 1 |
| 2019 | A similarity study of I/O traces via string kernels
Raul Torres, Julian M. Kunkel, Manuel F. Dolz, Thomas Ludwig 0002 |
J. Supercomput. | 3 |
| 2018 | Parallelizing and Optimizing LHCb-Kalman for Intel Xeon Phi KNL ProcessorsabstractReal time data processing is an important component of particle physics experiments with large computing resource requirements. As the Large Hadron Collider (LHC) at CERN is preparing for its next upgrade the LHCb experiment is upgrading its detector for a 30x increase in data throughput. In preparation for this upgrade the experiment is considering a number of architectural improvements encompassing both its software and hardware infrastructure. One of the hardware platforms under consideration is the Intel Xeon-Phi Knights Landing processor. Thanks to its on-package high-bandwidth memory and many-core architecture it offers an interesting alternative to more traditional server systems. We present a scalable, multi-threaded and NUMA-aware Kalman filter proto-application for particle track fitting expressed in terms of generic parallel patterns using the GrPPI interface. We show how code maintainability and readability improves, while maintaining comparable levels of performance to the baseline implementation. This is achieved by keeping the parallel algorithms in the underlying framework generic, but topology aware through the use of the Portable Hardware Locality (hwloc) library, which allows us to target different architectures with the same program. We measure the performance of our topology-aware GrPPI Kalman filter implementation on the Intel Xeon-Phi Knights Landing platform and conclude on the feasibility of integrating such high-level parallelization libraries in complex software frameworks such as LHCb's Gaudi framework. Placido Fernández, David del Rio Astorga, Manuel F. Dolz, Javier Fernández 0001, Omar Awile, José Daniel García |
PDP | 3 |
| 2018 | Supporting MPI-distributed stream parallel patterns in GrPPIabstractIn the recent years, the large volumes of stream data and the near real-time requirements of data streaming applications have exacerbated the need for new scalable algorithms and programming interfaces for distributed and shared-memory platforms. To contribute in this direction, this paper presents a new distributed MPI back end for GrPPI, a C++ high-level generic interface of data-intensive and stream processing parallel patterns. This back end, as a new execution policy, supports the distributed and hybrid (distributed and shared-memory) parallel execution of the Pipeline and Farm patterns, where the hybrid mode combines the MPI policy with a GrPPI shared-memory one. A detailed analysis of the GrPPI MPI execution policy reports considerable benefits from the programmability, flexibility and readability points of view. The experimental evaluation on a streaming application with different distributed and shared-memory scenarios reports considerable performance gains with respect to the sequential versions at the expense of negligible GrPPI overheads. Javier Fernández 0001, Manuel F. Dolz, David del Rio Astorga, Javier Prieto Cepeda, José Daniel García |
EuroMPI | 2 |
| 2018 | Paving the way towards high-level parallel pattern interfaces for data stream processing
David del Rio Astorga, Manuel F. Dolz, Javier Fernández 0001, José Daniel García |
Future Gener. Comput. Syst. | 2 |
| 2017 | Probabilistic-Based Selection of Alternate Implementations for Heterogeneous Platforms
Javier Fernández 0001, Andrés Sánchez Cuadrado, David del Rio Astorga, Manuel F. Dolz, José Daniel García |
ICA3PP | 4 |
| 2017 | A generic parallel pattern interface for stream and data processingabstractSummary Current parallel programming frameworks aid developers to a great extent in implementing applications that exploit parallel hardware resources. Nevertheless, developers require additional expertise to properly use and tune them to operate efficiently on specific parallel platforms. On the other hand, porting applications between different parallel programming models and platforms is not straightforward and demands considerable efforts and specific knowledge. Apart from that, the lack of high‐level parallel pattern abstractions, in those frameworks, further increases the complexity in developing parallel applications. To pave the way in this direction, this paper proposesGRPPI, a generic and reusable parallel pattern interface for both stream processing and data‐intensive C++ applications.GRPPIaccommodates a layer between developers and existing parallel programming frameworks targeting multi‐core processors, such as C++ threads, OpenMP and Intel TBB, and accelerators, as CUDA Thrust. Furthermore, thanks to its high‐level C++ application programming interface and pattern composability features,GRPPIallows users to easily expose parallelism via standalone patterns or patterns compositions matching in sequential applications. We evaluate this interface using an image processing use case and demonstrate its benefits from the usability, flexibility, and performance points of view. Furthermore, we analyze the impact of using stream and data pattern compositions on CPUs, GPUs and heterogeneous configurations. David del Rio Astorga, Manuel F. Dolz, Javier Fernández 0001, José Daniel García |
Concurr. Comput. Pract. Exp. | 2 |
| 2017 | Enabling semantics to improve detection of data races and misuses of lock-free data structuresabstractSummary The rapid progress of multi/many‐core architectures has caused data‐intensive parallel applications not yet fully optimized to deliver the best performance. In the advent of concurrent programming, frameworks offering structured patterns have alleviated developers' burden adapting such applications to multithreaded architectures. While some of these patterns are implemented using synchronization primitives, others avoid them by means of lock‐free data mechanisms. However, lock‐free programming is not straightforward, ensuring an appropriate use of their interfaces can be challenging, since different memory models plus instruction reordering at compiler/processor levels can interfere in the occurrence of data races. The benefits of race detectors are formidable in this sense; however, they may emit false positives if are unaware of the underlying lock‐free structure semantics. To mitigate this issue, this paper extends ThreadSanitizer, a race detection tool, with the semantics of 2 lock‐free data structures: the single‐producer/single‐consumer and the multiple‐producer/multiple‐consumer queues. With it, we are able to drop false positives and detect potential semantic violations. The experimental evaluation, using different queue implementations on a set ofμbenchmarks and real applications, demonstrates that it is possible to reduce, on average, 60% the number of data race warnings and detect wrong uses of these structures. Manuel F. Dolz, David del Rio Astorga, Javier Fernández 0001, Massimo Torquati, José Daniel García, Félix García Carballeira, Marco Danelutto |
Concurr. Comput. Pract. Exp. | 1 |
| 2017 | Adapting concurrency throttling and voltage-frequency scaling for dense eigensolvers
José Ignacio Aliaga, Maria Barreda, M. Asunción Castaño, Manuel F. Dolz, Enrique S. Quintana-Ortí |
J. Supercomput. | 4 |
| 2016 | A C++ Generic Parallel Pattern Interface for Stream Processing
David del Rio Astorga, Manuel F. Dolz, Luis Miguel Sánchez, Francisco Javier García Blas, José Daniel García |
ICA3PP | 2 |
| 2016 | Porting Matlab Applications to High-Performance C++ Codes: CPU/GPU-Accelerated Spherical Deconvolution of Diffusion MRI Data
Francisco Javier García Blas, Manuel F. Dolz, José Daniel García, Jesús Carretero 0001, Alessandro Daducci, Yasser Alemán-Gómez, Erick Jorge Canales-Rodríguez |
ICA3PP | 2 |
| 2016 | Analyzing the energy consumption of the storage data path
Pablo Llopis, Manuel F. Dolz, Francisco Javier García Blas, Florin Isaila, Mohammad Reza Heidari, Michael Kuhn 0003 |
J. Supercomput. | 2 |
| 2014 | Enhancing performance and energy consumption of runtime schedulers for dense linear algebraabstractSUMMARY The road towards Exascale Computing requires a holistic effort to address three different challenges simultaneously: high performance, energy efficiency, and programmability. The use of runtime task schedulers to orchestrate parallel executions with minimal developer intervention has been introduced in recent years to tackle the programmability issue while maintaining, or even improving, performance. In this paper, we enhance the SuperMatrix runtime task scheduler integrated in the libflame library in two different directions that address high performance and energy efficiency. First, we extend the runtime by accommodating hybrid parallel executions and managing task priorities for dense linear algebra operations, with remarkable performance improvements. Second, we introduce techniques to reduce energy consumption during idle times inherent to parallel executions, attaining important energy savings. In addition, we propose a power consumption model that can be leveraged by runtime task schedulers to make decisions based not only on performance but also on energy considerations. Copyright © 2014 John Wiley & Sons, Ltd. Pedro Alonso 0002, Manuel F. Dolz, Francisco D. Igual, Rafael Mayo 0002, Enrique S. Quintana-Ortí |
Concurr. Comput. Pract. Exp. | 2 |
| 2014 | Modeling power and energy consumption of dense matrix factorizations on multicore processorsabstractSUMMARY In this paper, we propose a model for the energy consumption of the concurrent execution of three key dense matrix factorizations, with task parallelism leveraged via the Symmetric Multi‐Processing Superscalar (SMPSs) runtime, on a multicore processor. Our model decomposes the power dissipation into the system, static and dynamic components, with the former two being estimated from basic, off‐line experiments. The dynamic power, on the other hand, requires significantly more care, and we introduce a contention‐aware model that accommodates for the variability of power consumption due to memory contention. Experimental results on an Intel Xeon E5504 processor with four cores, using an internal powermeter that samples the power drawn by the mainboard with a frequency of 1 KHz, show the reliability of the energy model for the Cholesky, LU, and QR factorizations on this platform. Copyright © 2013 John Wiley & Sons, Ltd. Pedro Alonso 0002, Manuel F. Dolz, Rafael Mayo 0002, Enrique S. Quintana-Ortí |
Concurr. Comput. Pract. Exp. | 2 |
| 2014 | Block pivoting implementation of a symmetric Toeplitz solver
Pedro Alonso 0002, Manuel F. Dolz, Antonio M. Vidal |
J. Parallel Distributed Comput. | 2 |
| 2012 | Tools for Power-Energy Modelling and Analysis of Parallel Scientific ApplicationsabstractUnderstanding power usage in parallel workloads is crucial to develop the energy-aware software that will run in future Exascale systems. In this paper, we contribute towards this goal by introducing an integrated framework to profile, monitor, model and analyze power dissipation in parallel MPI and multi-threaded scientific applications. The framework includes an own-designed device to measure internal DC power consumption and a package offering a simple interface to interact with this design as well as commercial power meters. Combined with the instrumentation package Extrae and the graphical analysis tool Paraver, the result is a useful environment to identify sources of power inefficiency directly in the source application code. For task-parallel codes, we also offer a statistical software module that inspects the execution trace of the application to calculate the parameters of an accurate model for the global energy consumption, which can be then decomposed into the average power usage per task or the nodal power dissipated per core. Pedro Alonso 0002, Rosa M. Badia, Jesús Labarta, Maria Barreda, Manuel F. Dolz, Rafael Mayo 0002, Enrique S. Quintana-Ortí, Ruymán Reyes |
ICPP | 5 |
| 2012 | Reducing Energy Consumption of Dense Linear Algebra Operations on Hybrid CPU-GPU PlatformsabstractWe investigate the balance between the time-to-solution and the energy consumption of a task-parallel execution of the Cholesky and LU factorizations on a hybrid platform, equipped with a multi-core processor and several GPUs. To improve energy efficiency, we incorporate two energy-saving techniques in the runtime in charge of scheduling the computations, to block idle threads and enable the transition to a more energy-friendly state of the general-purpose cores. Experiments on an Intel Xeon-based platform connected to an NVIDIA Tesla server report an average reduction of the energy consumption close to 9% (38% when only the consumption associated with the application is considered), for a minor increase in the execution time of the algorithm. Pedro Alonso 0002, Manuel F. Dolz, Francisco D. Igual, Rafael Mayo 0002, Enrique S. Quintana-Ortí |
ISPA | 2 |
| 2012 | Binding Performance and Power of Dense Linear Algebra OperationsabstractIn this paper we combine a powerful tracing framework with a power measurement setup to perform a visual analysis of the computational performance and the power consumption of tuned implementations for three key dense linear algebra operations: the LU factorization, the Cholesky factorization, and the reduction to tridiagonal form. Our results using 6 and 12 cores of an AMD Opteron-based platform reveal the serial/concurrent phases of the algorithms, and their connection to periods of low/high power consumption, as well as the linear dependency between execution time and energy for this class of operations. Maria Barreda, Manuel F. Dolz, Rafael Mayo 0002, Enrique S. Quintana-Ortí, Ruymán Reyes |
ISPA | 2 |
| 2012 | Saving Energy in the LU Factorization with Partial Pivoting on Multi-core ProcessorsabstractIn this paper we analyze the trade-off between energy and performance for a data-parallel execution of the LU factorization with partial pivoting on a multi-core processor. To improve energy efficiency, we adapt the runtime in charge of controlling the concurrent execution of the algorithm to leverage DVFS and block idle threads. For a CPU-bounded operation like the LU factorization, experiments on an AMD 8-core processor report a reduction around 5% in energy consumption for the largest problem sizes in exchange for a minor increase in the execution time. Pedro Alonso 0002, Manuel F. Dolz, Francisco D. Igual, Rafael Mayo 0002, Enrique S. Quintana-Ortí |
PDP | 2 |