Sergio Barrachina 0001

dblp:b/SergioBarrachina · also Sergio Barrachina-Mir · DBLP profile ↗
← Back
26ranked-venue papers
11as first author
8since 2021 · last 2025
0000-0003-4642-7467ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 7 first-author · 7 since 2021Artificial intelligence and machine learning · 9 · 4 first-authorGraphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-authorComputer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Deep learning inference optimisation for IoT: Conv2D-ReLU-BN layer fusion and quantisation
abstract
Abstract The deployment of deep learning models on resource-constrained devices requires the development of new optimisation techniques to effectively exploit the computational and storage capacities of these devices. Thus, the primary objective of this research is to introduce an innovative and efficient approach for fusing convolution (or fully connected), ReLU, and batch normalisation neural network layers into a unified, single-layer structure, alongside a quantisation method for this new fused layer. This approach has been evaluated using the Arduino BLE Sense ARM Cortex-M4 and the Arduino Portenta H7 Lite ARM Cortex-M4 and M7 processors, known for their widespread adoption in various Internet of Things devices. Depending on the microcontroller unit and compilation flag used, the fused layers can reduce the overall execution time by up to 1.53 $$\times$$ × , and on individual layers it can reach a speedup of 2.95 $$\times$$ × .
José I. Mestre, Sergio Barrachina 0001, Darwin Quezada, Manuel F. Dolz
J. Supercomput.2
2024 Optimizing Convolutions for Deep Learning Inference on ARM Cortex-M Processors
abstract
We perform a series of optimisations on the convolution operator within the ARM CMSIS-NN library to improve the performance of deep learning tasks on Arduino development boards equipped with ARM Cortex-M4 and M7 microcontrollers. To this end, we develop custom microkernels that efficiently handle the internal computations required by the convolution operator via the lowering approach and the direct method, and we design two techniques to avoid register spilling. We also take advantage of all the RAM on the Arduino boards by reusing it as a scratchpad for the convolution filters. The integration of these techniques into CMSIS-NN, when invoked by TensorFlow Lite for microcontrollers for quantised versions of VGG, SqueezeNet, ResNet, and MobileNet-like convolutional neural networks enhances the overall inference speed by a factor ranging from 1.13× to 1.50×.
Antonio Maciá-Lillo, Sergio Barrachina 0001, Germán Fabregat, Manuel F. Dolz
IEEE Internet Things J.2
2023 Reformulating the direct convolution for high-performance deep learning inference on ARM processors
abstract
We present two high-performance implementations of the convolution operator via the direct algorithm that outperform the so-called lowering approach based on the im2col transform plus the gemm kernel on an ARMv8-based processor. One of our methods presents the additional advantage of zero-memory overhead while the other employs an additional yet rather moderate workspace, substantially smaller than that required by the im2col+gemm solution. In contrast with a previous implementation of a similar zero-memory overhead direct convolution, this work exhibits the key advantage of preserving the conventional NHWC data layout for the input/output activations of the convolution layers.
Sergio Barrachina 0001, Adrián Castelló 0001, Manuel F. Dolz, Tze Meng Low, Héctor Martínez 0002, Enrique S. Quintana-Ortí, Upasana Sridhar, Andrés Tomás
J. Syst. Archit.1
2023 Performance-energy trade-offs of deep learning convolution algorithms on ARM processors
abstract
Abstract In this work, we assess the performance and energy efficiency of high-performance codes for the convolution operator, based on the direct, explicit/implicit lowering and Winograd algorithms used for deep learning (DL) inference on a series of ARM-based processor architectures. Specifically, we evaluate the NVIDIA Denver2 and Carmel processors, as well as the ARM Cortex-A57 and Cortex-A78AE CPUs as part of a recent set of NVIDIA Jetson platforms. The performance–energy evaluation is carried out using the ResNet-50 v1.5 convolutional neural network (CNN) on varying configurations of convolution algorithms, number of threads/cores, and operating frequencies on the tested processor cores. The results demonstrate that the best throughput is obtained on all platforms with the Winograd convolution operator running on all the cores at their highest frequency. However, if the goal is to reduce the energy footprint, there is no rule of thumb for the optimal configuration.
Manuel F. Dolz, Sergio Barrachina 0001, Héctor Martínez 0002, Adrián Castelló 0001, Antonio M. Vidal, Germán Fabregat, Andrés Tomás
J. Supercomput.2
2022 Efficient and portable GEMM-based convolution operators for deep neural network training on multicore processors
abstract
Convolutional Neural Networks (CNNs) play a crucial role in many image recognition and classification tasks, recommender systems, brain-computer interfaces, etc. As a consequence, there is a notable interest in developing high performance realizations of the convolution operators, which concentrate a significant portion of the computational cost of this type of neural networks. In a previous work, we introduced a portable, high performance convolution algorithm, based on the BLIS realization of matrix multiplication, which eliminates most of the runtime and memory overheads that impair the performance of the convolution operators appearing in the forward training pass, when performed via explicit im2col transform. In this paper, we extend our ideas to the full training process of CNNs on multicore processors, proposing new high performance strategies to tackle the convolution operators that are present in the more complex backward pass of the training process, while maintaining the portability of the realizations. In addition, we conduct a full integration of these algorithms into a framework for distributed training of CNNs on clusters of computers, providing a complete experimental evaluation of the actual benefits in terms of both performance and memory consumption. Compared with baseline implementation, the use of the new convolution operators using pre-allocated memory can accelerate the training by a factor of about 6%–25%, provided there is sufficient memory available. In comparison, the operator variants that do not rely on persistent memory can save up to 70% of memory.
Sergio Barrachina 0001, Manuel F. Dolz, Pablo San Juan, Enrique S. Quintana-Ortí
J. Parallel Distributed Comput.1
2022 High performance and energy efficient inference for deep learning on multicore ARM processors using general optimization techniques and BLIS
abstract
We evolve PyDTNN, a framework for distributed parallel training of Deep Neural Networks (DNNs), into an efficient inference tool for convolutional neural networks. Our optimization process on multicore ARM processors involves several high-level transformations of the original framework, such as the development and integration of Cython routines to exploit thread-level parallelism; the design and development of micro-kernels for the matrix multiplication, vectorized with ARM’s NEON intrinsics, that can accommodate layer fusion; and the appropriate selection of several cache configuration parameters tailored to the memory hierarchy of the target ARM processors. Our experiments evaluate both inference throughput (measured in processed images/s) and inference latency (i.e., time-to-response) as well as energy consumption per image when varying the level of thread parallelism and the processor power modes. The experiments with the new inference engine are reported for the ResNet50 v1.5 model on the ImageNet dataset from the MLPerf suite using the ARM v8.2 cores in the NVIDIA Jetson AGX Xavier board. These results show superior performance compared with the well-spread TFLite from Google and slightly inferior results when compared with ArmNN, the native library from ARM for DNN inference.
Adrián Castelló 0001, Sergio Barrachina 0001, Manuel F. Dolz, Enrique S. Quintana-Ortí, Pau San Juan, Andrés Tomás
J. Syst. Archit.2
2022 BestOf: an online implementation selector for the training and inference of deep neural networks
abstract
Abstract Tuning and optimising the operations executed in deep learning frameworks is a fundamental task in accelerating the processing of deep neural networks (DNNs). However, this optimisation usually requires extensive manual efforts in order to obtain the best performance for each combination of tensor input size, layer type, and hardware platform. In this work, we present , a novel online auto-tuner that optimises the training and inference phases of DNNs. automatically selects at run time, and among the provided alternatives, the best performing implementation in each layer according to gathered profiling data. The evaluation of is performed on multi-core architectures for different DNNs using , a lightweight library for distributed training and inference. The experimental results reveal that the auto-tuner delivers the same or higher performance than that achieved using a static selection approach.
Sergio Barrachina 0001, Adrián Castelló 0001, Manuel F. Dolz, Andrés Tomás
J. Supercomput.1
2021 PyDTNN: A user-friendly and extensible framework for distributed deep learning
Sergio Barrachina 0001, Adrián Castelló 0001, Mar Catalán, Manuel F. Dolz, José I. Mestre
J. Supercomput.1
2017 Accelerating FaST-LMM for Epistasis Tests
Héctor Martínez 0002, Sergio Barrachina 0001, María Isabel Castillo, Enrique S. Quintana-Ortí, Jordi Rambla De Argila, Xavier Farré, Arcadi Navarro
ICA3PP2
2015 Concurrent and Accurate Short Read Mapping on Multicore Processors
abstract
We introduce a parallel aligner with a work-flow organization for fast and accurate mapping of RNA sequences on servers equipped with multicore processors. Our software, HPG Aligner SA (HPG Aligner SA is an open-source application. The software is available at http://www.opencb.org, exploits a suffix array to rapidly map a large fraction of the RNA fragments (reads), as well as leverages the accuracy of the Smith-Waterman algorithm to deal with conflictive reads. The aligner is enhanced with a careful strategy to detect splice junctions based on an adaptive division of RNA reads into small segments (or seeds), which are then mapped onto a number of candidate alignment locations, providing crucial information for the successful alignment of the complete reads. The experimental results on a platform with Intel multicore technology report the parallel performance of HPG Aligner SA, on RNA reads of 100-400 nucleotides, which excels in execution time/sensitivity to state-of-the-art aligners such as TopHat 2+Bowtie 2, MapSplice, and STAR.
Héctor Martínez 0002, Joaquín Tárraga, Ignacio Medina, Sergio Barrachina 0001, María Isabel Castillo, Joaquín Dopazo, Enrique S. Quintana-Ortí
IEEE ACM Trans. Comput. Biol. Bioinform.4
2013 A dynamic pipeline for RNA sequencing on multicore processors
abstract
We present a concurrent algorithm for mapping short and long RNA sequences on multicore processors. Our solution processes the data, initially stored on disk, in batches of reads which are passed between the consecutive stages of a pipeline. A major operational reorganization of the original static pipeline, combined with a complete reimplementation based on POSIX threads, renders a dissociated execution between threads and stages/task types, so that threads can compute any type of pending task resulting in a dynamic pipeline. The experiments on a multicore platform reveal that this reorganization yields significantly higher performance, specially for architectures equipped with a small to moderate number of cores.
Héctor Martínez 0002, Joaquín Tárraga, Ignacio Medina, Sergio Barrachina 0001, María Isabel Castillo, Joaquín Dopazo, Enrique S. Quintana-Ortí
EuroMPI4
2009 Statistical Approaches to Computer-Assisted Translation
abstract
Current machine translation (MT) systems are still not perfect. In practice, the output from these systems needs to be edited to correct errors. A way of increasing the productivity of the whole translation process (MT plus human work) is to incorporate the human correction activities within the translation process itself, thereby shifting the MT paradigm to that of computer-assisted translation. This model entails an iterative process in which the human translator activity is included in the loop: In each iteration, a prefix of the translation is validated (accepted or amended) by the human and the system computes its best (or n-best) translation suffix hypothesis to complete this prefix. A successful framework for MT is the so-called statistical (or pattern recognition) framework. Interestingly, within this framework, the adaptation of MT systems to the interactive scenario affects mainly the search process, allowing a great reuse of successful techniques and models. In this article, alignment templates, phrase-based models, and stochastic finite-state transducers are used to develop computer-assisted translation systems. These systems were assessed in a European project (TransType2) in two real tasks: The translation of printer manuals; manuals and the translation of the Bulletin of the European Union. In each task, the following three pairs of languages were involved (in both translation directions): English-Spanish, English-German, and English-French.
Sergio Barrachina 0001, Oliver Bender, Francisco Casacuberta, Jorge Civera, Elsa Cubel, Shahram Khadivi, Antonio L. Lagarda, Hermann Ney, Jesús Tomás, Enrique Vidal 0001, Juan Miguel Vilar
Comput. Linguistics1
2009 Exploiting the capabilities of modern GPUs for dense matrix computations
abstract
Abstract We present several algorithms to compute the solution of a linear system of equations on a graphics processor (GPU), as well as general techniques to improve their performance, such as padding and hybrid GPU‐CPU computation. We compare single and double precision performance of a modern GPU with unified architecture, and show how iterative refinement with mixed precision can be used to regain full accuracy in the solution of linear systems, exploiting the potential of the processor for single precision arithmetic. Experimental results on a GTX280 using CUBLAS 2.0, the implementation of BLAS for NVIDIA® GPUs with unified architecture, illustrate the performance of the different algorithms and techniques proposed. Copyright © 2009 John Wiley & Sons, Ltd.
Sergio Barrachina 0001, María Isabel Castillo, Francisco D. Igual, Rafael Mayo 0002, Enrique S. Quintana-Ortí, Gregorio Quintana-Ortí
Concurr. Comput. Pract. Exp.1
2009 Toward the parallelization of GSL
José Ignacio Aliaga, Francisco Almeida, José M. Badía, Sergio Barrachina 0001, Vicente Blanco 0001, María Isabel Castillo, Rafael Mayo 0002, Enrique S. Quintana-Ortí, Gregorio Quintana-Ortí, Alfredo Remón, Casiano Rodríguez, Francisco de Sande, Adrián Santos
J. Supercomput.4
2008 Solving Dense Linear Systems on Graphics Processors
Sergio Barrachina 0001, María Isabel Castillo, Francisco D. Igual, Rafael Mayo 0002, Enrique S. Quintana-Ortí
Euro-Par1
2008 Evaluation and tuning of the Level 3 CUBLAS for graphics processors
abstract
The increase in performance of the last generations of graphics processors (GPUs) has made this class of platform a coprocessing tool with remarkable success in certain types of operations. In this paper we evaluate the performance of the Level 3 operations in CUBLAS, the implementation of BIAS for NVIDIAreg GPUs with unified architecture. From this study, we gain insights on the quality of the kernels in the library and we propose several alternative implementations that are competitive with those in CUBLAS. Experimental results on a GeForce 8800 Ultra compare the performance of CUBLAS and the new variants.
Sergio Barrachina 0001, María Isabel Castillo, Francisco D. Igual, Rafael Mayo 0002, Enrique S. Quintana-Ortí
IPDPS1
2008 The UJIpenchars Database: a Pen-Based Database of Isolated Handwritten Characters
David Llorens, Federico Prat, Andrés Marzal, Juan Miguel Vilar, María José Castro Bleda, Juan-Carlos Amengual, Sergio Barrachina 0001, Antonio Castellanos, Salvador España Boquera, J. A. Gómez, Jorge Gorbe-Moya, Albert Gordo, Vicente Palazón, Guillermo Peris, Rafael Ramos-Garijo, Francisco Zamora-Martínez
LREC7
2006 A Computer-Assisted Translation Tool based on Finite-State Technology
Jorge Civera, Antonio L. Lagarda, Elsa Cubel, Francisco Casacuberta, Enrique Vidal 0001, Juan Miguel Vilar, Sergio Barrachina 0001
EAMT7
2006 An Open Source Web Service Based Platform for Heterogeneous Clusters
Francisco Almeida, Sergio Barrachina 0001, Vicente Blanco 0001, Enrique S. Quintana-Ortí, Adrián Santos
ISPA2
2006 Parallelization of GSL: The Web Service Interface
abstract
We present our joint effort to develop a Web based interface for the GNU Scientific library and its parallelization. The interface has been developed using standard Web services technology to enable the use of non local resources to execute parallel programs. The final result is a computing service where sequential and parallel routines demanding high performance computing are supplied. The design allows to incorporate new servers and platforms with a small number of software requirements.
José Ignacio Aliaga, José M. Badía, Sergio Barrachina 0001, María Isabel Castillo, Rafael Mayo 0002, Enrique S. Quintana-Ortí, Gregorio Quintana-Ortí, Francisco Almeida, Vicente Blanco 0001, Casiano Rodríguez, Francisco de Sande, Adrián Santos
PDP3
2004 Automatic Discovery of Translation Collocations from Bilingual Corpora
Sergio Barrachina 0001, Juan Miguel Vilar
ECAI1
2004 From Machine Translation to Computer Assisted Translation using Finite-State Models
Jorge Civera, Elsa Cubel, Antonio L. Lagarda, David Picó, Enrique Vidal 0001, Francisco Casacuberta, Juan Miguel Vilar, Sergio Barrachina 0001
EMNLP9
2004 Some approaches to statistical and finite-state speech-to-speech translation
Francisco Casacuberta, Hermann Ney, Franz Josef Och, Enrique Vidal 0001, Juan Miguel Vilar, Sergio Barrachina 0001, Ismael García-Varea, David Llorens, Carlos D. Martínez-Hinarejos, Sirko Molau
Comput. Speech Lang.6
2003 Incremental and iterative monolingual clustering algorithms
Sergio Barrachina 0001, Juan Miguel Vilar
INTERSPEECH1
2000 Speeding Up the Computation of the Edit Distance for Cyclic Strings
abstract
A new algorithm to compute the edit distance between cyclic strings is presented. Experimental results with synthetic cyclic strings and a handwritten digits recognition task show that the new algorithm is faster than Maes' (1990) and Gregor and Thomason's (1993) algorithms.
Andrés Marzal, Sergio Barrachina 0001
ICPR2
1999 Automatically deriving categories for translation
abstract
An accurate fundamental frequency (F0) estimation method for non-stationary, speech-like sounds is proposed based on the dif-ferential properties of the instantaneous frequencies of two sets of filter outputs. A specific type of fixed points of mapping from the filter center frequency to the output instantaneous frequency provides frequencies of the constituent sinusoidal components of the input signal. When the filter is made from an isometric Gabor function convoluted with a cardinal B-spline basis func-tion, the differential properties at the fixed points provide prac-tical estimates of the carrier-to-noise ratio of the corresponding components. These estimates are used to select the fundamental component and to integrate the F0 information distributed among the other harmonic components. 1.
Sergio Barrachina 0001, Juan Miguel Vilar
EUROSPEECH1