Juan C. Moure

dblp:59/5413 · also Juan Carlos Moure · DBLP profile ↗
← Back
35ranked-venue papers
6as first author
5since 2021 · last 2023
0000-0001-6697-0331ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 20 · 5 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 2Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
GPUs and heterogeneous computing · 85% Processor architecture and microarchitecture · 8% Parallel and multicore computing · 8%
Artificial intelligence
2 papers
3D vision · 53% Segmentation and scene understanding · 38% Autonomous driving · 9%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Bioinformatics and computational biology · 100%
Computer graphics and multimedia
3 papers
Image and video coding · 75% Image and video processing · 19% Audio and music processing · 6%

Topics — the 19 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
GPUs and heterogeneous computing
GPU computing
1.742023
WFA-GPU: gap-affine pairwise read-alignment using GPUs · Bioinform. 2023
3D Perception With Slanted Stixels on GPU · IEEE Trans. Parallel Distributed Syst. 2021
GPU Implementation of Bitplane Coding with Parallel Coefficient Processing for High Performance Image Compression · IEEE Trans. Parallel Distributed Syst. 2017
Bioinformatics and computational biology
sequence alignment
0.722023
Fast gap-affine pairwise alignment using the wavefront algorithm · Bioinform. 2021
WFA-GPU: gap-affine pairwise read-alignment using GPUs · Bioinform. 2023
GPUs and heterogeneous computing
GPU-accelerated bioinformatics
0.712023
WFA-GPU: gap-affine pairwise read-alignment using GPUs · Bioinform. 2023
Image and video coding
bit-plane coding
0.522017
GPU Implementation of Bitplane Coding with Parallel Coefficient Processing for High Performance Image Compression · IEEE Trans. Parallel Distributed Syst. 2017
Bitplane Image Coding With Parallel Coefficient Processing · IEEE Trans. Image Process. 2016
Computer vision › 3D vision
depth estimation
0.512021
3D Perception With Slanted Stixels on GPU · IEEE Trans. Parallel Distributed Syst. 2021
Bioinformatics and computational biology › sequence alignment
pairwise sequence alignment
0.512021
Fast gap-affine pairwise alignment using the wavefront algorithm · Bioinform. 2021
GPUs and heterogeneous computing › embedded GPU
embedded GPU acceleration
0.512021
3D Perception With Slanted Stixels on GPU · IEEE Trans. Parallel Distributed Syst. 2021
Computer vision › Segmentation and scene understanding
semantic segmentation
0.412019
Slanted Stixels: A Way to Represent Steep Streets · Int. J. Comput. Vis. 2019
Computer vision › 3D vision
stereo vision
0.412019
Slanted Stixels: A Way to Represent Steep Streets · Int. J. Comput. Vis. 2019
Parallel and multicore computing › parallel algorithms › parallel image processing
parallel image coding
0.212016
Bitplane Image Coding With Parallel Coefficient Processing · IEEE Trans. Image Process. 2016
Processor architecture and microarchitecture
SIMD
0.212016
Bitplane Image Coding With Parallel Coefficient Processing · IEEE Trans. Image Process. 2016
Image and video coding › transform coding
discrete wavelet transform
0.212015
Implementation of the DWT in a GPU through a Register-based Strategy · IEEE Trans. Parallel Distributed Syst. 2015
Image and video processing
wavelet transform
0.212015
Implementation of the DWT in a GPU through a Register-based Strategy · IEEE Trans. Parallel Distributed Syst. 2015
Bioinformatics and computational biology › sequence analysis › read mapping
long-read alignment
0.212023
WFA-GPU: gap-affine pairwise read-alignment using GPUs · Bioinform. 2023
Robotics › Autonomous driving
perception
0.112021
3D Perception With Slanted Stixels on GPU · IEEE Trans. Parallel Distributed Syst. 2021
Computer vision › Segmentation and scene understanding › scene understanding
semantic scene understanding
0.112021
3D Perception With Slanted Stixels on GPU · IEEE Trans. Parallel Distributed Syst. 2021
Computer vision › Segmentation and scene understanding
scene understanding
0.112019
Slanted Stixels: A Way to Represent Steep Streets · Int. J. Comput. Vis. 2019
Image and video coding › image compression
wavelet-based image coding
0.112017
GPU Implementation of Bitplane Coding with Parallel Coefficient Processing for High Performance Image Compression · IEEE Trans. Parallel Distributed Syst. 2017
Audio and music processing
decorrelation
0.112015
Implementation of the DWT in a GPU through a Register-based Strategy · IEEE Trans. Parallel Distributed Syst. 2015

Methods — techniques the papers use, named apart from their topics

dynamic programming · 2.8wavefront alignment algorithm · 1.3GPU parallelization · 1.0thread-to-data mapping · 0.6memory management · 0.6context modeling · 0.5arithmetic coding · 0.5SIMD vectorization · 0.5SIMD processing · 0.5register sharing · 0.4CUDA · 0.4over-segmentation · 0.4global energy minimization · 0.4fully convolutional network · 0.4
YearPublicationVenuePosition
2023 WFA-GPU: gap-affine pairwise read-alignment using GPUs
abstract
MOTIVATION: Advances in genomics and sequencing technologies demand faster and more scalable analysis methods that can process longer sequences with higher accuracy. However, classical pairwise alignment methods, based on dynamic programming (DP), impose impractical computational requirements to align long and noisy sequences like those produced by PacBio and Nanopore technologies. The recently proposed wavefront alignment (WFA) algorithm paves the way for more efficient alignment tools, improving time and memory complexity over previous methods. However, high-performance computing (HPC) platforms require efficient parallel algorithms and tools to exploit the computing resources available on modern accelerator-based architectures. RESULTS: This paper presents WFA-GPU, a GPU (graphics processing unit)-accelerated tool to compute exact gap-affine alignments based on the WFA algorithm. We present the algorithmic adaptations and performance optimizations that allow exploiting the massively parallel capabilities of modern GPU devices to accelerate the alignment computations. In particular, we propose a CPU-GPU co-design capable of performing inter-sequence and intra-sequence parallel sequence alignment, combining a succinct WFA-data representation with an efficient GPU implementation. As a result, we demonstrate that our implementation outperforms the original multi-threaded WFA implementation by up to 4.3× and up to 18.2× when using heuristic methods on long and noisy sequences. Compared to other state-of-the-art tools and libraries, the WFA-GPU is up to 29× faster than other GPU implementations and up to four orders of magnitude faster than other CPU implementations. Furthermore, WFA-GPU is the only GPU solution capable of correctly aligning long reads using a commodity GPU. AVAILABILITY AND IMPLEMENTATION: WFA-GPU code and documentation are publicly available at https://github.com/quim0/WFA-GPU.
Quim Aguado-Puig, Max Doblas, Christos Matzoros, Antonio Espinosa 0001, Juan C. Moure, Santiago Marco-Sola, Miquel Moretó
Bioinform.5
2021 OpenCL-based FPGA Accelerator for Semi-Global Approximate String Matching Using Diagonal Bit-Vectors
abstract
An FPGA accelerator for the computation of the semi-global Levenshtein distance between a pattern and a reference text is presented. The accelerator provides an important benefit to reduce the execution time of read-mappers used in short-read genomic sequencing. Previous attempts to solve the same problem in FPGA use the Myers algorithm following a column approach to compute the dynamic programming table. We use an approach based on diagonals that allows for some resource savings while maintaining a very high throughput of 1 alignment per clock cycle. The design is implemented in OpenCL and tested on two FPGA accelerators. The maximum performance obtained is 91.5 MPairs/s for 100 × 120 sequences and 47 MPairs/s for 300 × 360 sequences, the highest ever reported for this problem.
David Castells-Rufas, Santiago Marco-Sola, Quim Aguado-Puig, Antonio Espinosa 0001, Juan C. Moure, Lluc Alvarez, Miquel Moretó
FPL5
2021 Fast gap-affine pairwise alignment using the wavefront algorithm
abstract
MOTIVATION: Pairwise alignment of sequences is a fundamental method in modern molecular biology, implemented within multiple bioinformatics tools and libraries. Current advances in sequencing technologies press for the development of faster pairwise alignment algorithms that can scale with increasing read lengths and production yields. RESULTS: In this article, we present the wavefront alignment algorithm (WFA), an exact gap-affine algorithm that takes advantage of homologous regions between the sequences to accelerate the alignment process. As opposed to traditional dynamic programming algorithms that run in quadratic time, the WFA runs in time O(ns), proportional to the read length n and the alignment score s, using O(s2) memory. Furthermore, our algorithm exhibits simple data dependencies that can be easily vectorized, even by the automatic features of modern compilers, for different architectures, without the need to adapt the code. We evaluate the performance of our algorithm, together with other state-of-the-art implementations. As a result, we demonstrate that the WFA runs 20-300× faster than other methods aligning short Illumina-like sequences, and 10-100× faster using long noisy reads like those produced by Oxford Nanopore Technologies. AVAILABILITY AND IMPLEMENTATION: The WFA algorithm is implemented within the wavefront-aligner library, and it is publicly available at https://github.com/smarco/WFA.
Santiago Marco-Sola, Juan C. Moure, Miquel Moretó, Antonio Espinosa 0001
Bioinform.2
2021 Real-time 16K video coding on a GPU with complexity scalable BPC-PaCo
abstract
The advent of new technologies such as high dynamic range or 8K screens has enhanced the quality of digital images but it has also increased the codecs’ computational demands to process such data. This paper presents a video codec that, while providing the same coding features and performance as those of JPEG2000, can process 16K video in real time using a consumer-grade GPU. This high throughput is achieved with a technique that introduces complexity scalability to a bitplane coding engine, which is the most computationally complex stage of the coding pipeline. The resulting codec can trade throughput for coding performance depending on the user’s needs. Experimental results suggest that our method can double the throughput achieved by CPU implementations of the recently approved High-Throughput JPEG2000 and by hardwired implementations of HEVC in a GPU.
Carlos de Cea-Dominguez, Juan C. Moure, Joan Bartrina-Rapesta, Francesc Aulí Llinàs
Signal Process. Image Commun.2
2021 3D Perception With Slanted Stixels on GPU
abstract
This article presents a GPU-accelerated software design of the recently proposed model of Slanted Stixels, which represents the geometric and semantic information of a scene in a compact and accurate way. We reformulate the measurement depth model to reduce the computational complexity of the algorithm, relying on the confidence of the depth estimation and the identification of invalid values to handle outliers. The proposed massively parallel scheme and data layout for the irregular computation pattern that corresponds to a Dynamic Programming paradigm is described and carefully analyzed in performance terms. Performance is shown to scale gracefully on current generation embedded GPUs. We assess the proposed methods in terms of semantic and geometric accuracy as well as run-time performance on three publicly available benchmark datasets. Our approach achieves real-time performance with high accuracy for 2048 × 1024 image sizes and 4 × 4 Stixel resolution on the low-power embedded GPU of an NVIDIA Tegra Xavier.
Daniel Hernández Juárez, Antonio Espinosa 0001, David Vázquez 0001, Antonio M. López 0001, Juan C. Moure
IEEE Trans. Parallel Distributed Syst.5
2020 Complexity Scalable Bitplane Image Coding With Parallel Coefficient Processing
abstract
Very fast image and video codecs are a pursued goal both in the academia and the industry. This paper presents a complexity scalable and parallel bitplane coding engine for wavelet-based image codecs. The proposed method processes the coefficients in parallel, suiting hardware architectures based on vector instructions. Our previous work is extended with a mechanism that provides complexity scalability to the system. Such a feature allows the coder to regulate the throughput achieved at the expense of slightly penalizing compression efficiency. Experimental results suggests that, when using the fastest speed, the method almost doubles the throughput of our previous engine while penalizing compression efficiency by about 10%.
Carlos de Cea-Dominguez, Juan C. Moure, Joan Bartrina-Rapesta, Francesc Aulí Llinàs
IEEE Signal Process. Lett.2
2019 Slanted Stixels: A Way to Represent Steep Streets
abstract
Abstract This work presents and evaluates a novel compact scene representation based on Stixels that infers geometric and semantic information. Our approach overcomes the previous rather restrictive geometric assumptions for Stixels by introducing a novel depth model to account for non-flat roads and slanted objects. Both semantic and depth cues are used jointly to infer the scene representation in a sound global energy minimization formulation. Furthermore, a novel approximation scheme is introduced in order to significantly reduce the computational complexity of the Stixel algorithm, and then achieve real-time computation capabilities. The idea is to first perform an over-segmentation of the image, discarding the unlikely Stixel cuts, and apply the algorithm only on the remaining Stixel cuts. This work presents a novel over-segmentation strategy based on a fully convolutional network, which outperforms an approach based on using local extrema of the disparity map. We evaluate the proposed methods in terms of semantic and geometric accuracy as well as run-time on four publicly available benchmark datasets. Our approach maintains accuracy on flat road scene datasets while improving substantially on a novel non-flat road dataset.
Daniel Hernández Juárez, Lukas Schneider, Pau Cebrian, Antonio Espinosa 0001, David Vázquez 0001, Antonio M. López 0001, Uwe Franke, Marc Pollefeys, Juan C. Moure
Int. J. Comput. Vis.9
2018 High Throughput Image Codec for High-Resolution Satellite Images
abstract
The growth in the use of satellite images has generated the need for their fast compression, processing, and distribution. JPEG2000 is a widespread standard for the compression and transmission of such images once they are in the ground. Despite its advanced features and excellent coding performance, JPEG2000 demands significant computational resources. This paper introduces a wavelet-based codec that uses the JPEG2000 framework, but replaces its most computationally demanding coding stage by a highly parallel engine. When executed in Graphics Processing Units to code high-resolution satellite images, the proposed codec achieves speed-ups of up to 8× when compared to the fastest implementation of JPEG2000 executed in a multi-core platform.
Carlos de Cea-Dominguez, Pablo Enfedaque, Juan C. Moure, Joan Bartrina-Rapesta, Francesc Aulí Llinàs
IGARSS3
2017 Slanted Stixels: Representing San Francisco's Steepest Streets
Daniel Hernández Juárez, Lukas Schneider, Antonio Espinosa 0001, Juan C. Moure, David Vázquez 0001, Antonio M. López 0001, Uwe Franke, Marc Pollefeys
BMVC4
2017 GPU-Accelerated Real-Time Stixel Computation
abstract
The Stixel World is a medium-level, compact representation of road scenes that abstracts millions of disparity pixels into hundreds or thousands of stixels. The goal of this work is to implement and evaluate a complete multistixel estimation pipeline on an embedded, energy-efficient, GPU-accelerated device. This work presents a full GPUaccelerated implementation of stixel estimation that produces reliable results at 26 frames per second (real-time) on the Tegra X1 for disparity images of 1024×440 pixels and stixel widths of 5 pixels, and achieves more than 400 frames per second on a high-end Titan X GPU card.
Daniel Hernández Juárez, Antonio Espinosa 0001, Juan C. Moure, David Vázquez 0001, Antonio M. López 0001
WACV3
2017 Coalition structure generation problems: optimization and parallelization of the IDP algorithm in multicore systems
abstract
Summary The coalition structure generation problem is well known in the area of multi‐agent systems. Its goal is to establish coalitions between agents while maximizing the global welfare. Among the existing different algorithms designed to solve the coalition structure generation problem, DP and IDP are the ones with smaller temporal complexity. After analyzing the operation of the dynamic programming and improved dynamic programming algorithms, we have identified which are the most frequent operations and propose an optimized method. In addition, we study and implement a method for dividing the work into different threads. To describe incremental improvements of the algorithm design, we first compare performance of an improved single central processing unit core version where we obtain speedups ranging from 7 × to 11 × . Then, we describe the best resource use in a multi‐thread optimized version where we obtain an additional 7.5 × speedup running in a 12‐core machine.
Francisco Cruz-Mencia, Antonio Espinosa 0001, Juan C. Moure, Jesús Cerquides, Juan A. Rodríguez-Aguilar, Kim Svensson, Sarvapali D. Ramchurn
Concurr. Comput. Pract. Exp.3
2017 Introducing computational thinking, parallel programming and performance engineering in interdisciplinary studies
Eduardo César, Ana Cortés, Antonio Espinosa 0001, Tomàs Margalef, Juan C. Moure, Anna Sikora, Remo Suppi
J. Parallel Distributed Comput.5
2017 GPU Implementation of Bitplane Coding with Parallel Coefficient Processing for High Performance Image Compression
abstract
The fast compression of images is a requisite in many applications like TV production, teleconferencing, or digital cinema. Many of the algorithms employed in current image compression standards are inherently sequential. High performance implementations of such algorithms often require specialized hardware like field integrated gate arrays. Graphics Processing Units (GPUs) do not commonly achieve high performance on these algorithms because they do not exhibit fine-grain parallelism. Our previous work introduced a new core algorithm for wavelet-based image coding systems. It is tailored for massive parallel architectures. It is called bitplane coding with parallel coefficient processing (BPC-PaCo). This paper introduces the first high performance, GPUbased implementation of BPC-PaCo. A detailed analysis of the algorithm aids its implementation in the GPU. The main insights behind the proposed codec are an efficient thread-to-data mapping, a smart memory management, and the use of efficient cooperation mechanisms to enable inter-thread communication. Experimental results indicate that the proposed implementation matches the requirements for high resolution (4 K) digital cinema in real time, yielding speedups of 30× with respect to the fastest implementations of current compression standards. Also, a power consumption evaluation shows that our implementation consumes 40× less energy for equivalent performance than state-of-the-art methods.
Pablo Enfedaque, Francesc Aulí Llinàs, Juan C. Moure
IEEE Trans. Parallel Distributed Syst.3
2016 Bitplane Image Coding With Parallel Coefficient Processing
abstract
Image coding systems have been traditionally tailored for multiple instruction, multiple data (MIMD) computing. In general, they partition the (transformed) image in codeblocks that can be coded in the cores of MIMD-based processors. Each core executes a sequential flow of instructions to process the coefficients in the codeblock, independently and asynchronously from the others cores. Bitplane coding is a common strategy to code such data. Most of its mechanisms require sequential processing of the coefficients. The last years have seen the upraising of processing accelerators with enhanced computational performance and power efficiency whose architecture is mainly based on the single instruction, multiple data (SIMD) principle. SIMD computing refers to the execution of the same instruction to multiple data in a lockstep synchronous way. Unfortunately, current bitplane coding strategies cannot fully profit from such processors due to inherently sequential coding task. This paper presents bitplane image coding with parallel coefficient (BPC-PaCo) processing, a coding method that can process many coefficients within a codeblock in parallel and synchronously. To this end, the scanning order, the context formation, the probability model, and the arithmetic coder of the coding engine have been re-formulated. The experimental results suggest that the penalization in coding performance of BPC-PaCo with respect to the traditional strategies is almost negligible.
Francesc Aulí Llinàs, Pablo Enfedaque, Juan C. Moure, Victor Sanchez
IEEE Trans. Image Process.3
2015 Parallelisation and Application of AD 3 as a Method for Solving Large Scale Combinatorial Auctions
Francisco Cruz-Mencia, Jesús Cerquides, Antonio Espinosa 0001, Juan C. Moure, Juan A. Rodríguez-Aguilar
COORDINATION4
2015 Strategy of Microscopic Parallelism for Bitplane Image Coding
abstract
Recent years have seen the upraising of a new type of processors strongly relying on the Single Instruction, Multiple Data (SIMD) architectural principle. The main idea behind SIMD computing is to apply a flow of instructions to multiple pieces of data in parallel and synchronously. This permits the execution of thousands of operations in parallel, achieving higher computational performance than with traditional Multiple Instruction, Multiple Data (MIMD) architectures. The level of parallelism required in SIMD computing can only be achieved in image coding systems via microscopic parallel strategies that code multiple coefficients in parallel. Until now, the only way to achieve microscopic parallelism in bit plane coding engines was by executing multiple coding passes in parallel. Such a strategy does not suit well SIMD computing because each thread executes different instructions. This paper introduces the first bit plane coding engine devised for the fine grain of parallelism required in SIMD computing. Its main insight is to allow parallel coefficient processing in a coding pass. Experimental tests show coding performance results similar to those of JPEG2000.
Francesc Aulí Llinàs, Pablo Enfedaque, Juan C. Moure, Ian Blanes, Victor Sanchez
DCC3
2015 Strategies of SIMD Computing for Image Coding in GPU
abstract
The main difficulty to implement modern image coding systems in a GPU is that the algorithms employed in the core of the coding scheme are inherently sequential. We recently proposed bitplane image coding with parallel coefficient processing (BPC-PaCo), a coding scheme that, contrarily to most systems, permits the processing of multiple coefficients of the image in parallel. This enables the use of SIMD computing, ideal for its implementation in a GPU. This paper introduces and evaluates the GPU implementation of BPC-PaCo employing two different strategies that tradeoff computational throughput and compression efficiency. The proposed implementation is compared to the best CPU and GPU implementations of JPEG2000, the state-of-the-art image compression standard. Experimental results indicate that BPC-PaCo achieves a computational throughput that is an order of magnitude superior to that achieved with such implementations with a small reduction in coding efficiency.
Pablo Enfedaque, Francesc Aulí Llinàs, Juan C. Moure
HiPC3
2015 Boosting the FM-Index on the GPU: Effective Techniques to Mitigate Random Memory Access
abstract
The recent advent of high-throughput sequencing machines producing big amounts of short reads has boosted the interest in efficient string searching techniques. As of today, many mainstream sequence alignment software tools rely on a special data structure, called the FM-index, which allows for fast exact searches in large genomic references. However, such searches translate into a pseudo-random memory access pattern, thus making memory access the limiting factor of all computation-efficient implementations, both on CPUs and GPUs. Here, we show that several strategies can be put in place to remove the memory bottleneck on the GPU: more compact indexes can be implemented by having more threads work cooperatively on larger memory blocks, and a k-step FM-index can be used to further reduce the number of memory accesses. The combination of those and other optimisations yields an implementation that is able to process about two Gbases of queries per second on our test platform, being about 8 × faster than a comparable multi-core CPU version, and about 3 × to 5 × faster than the FM-index implementation on the GPU provided by the recently announced Nvidia NVBIO bioinformatics library.
Alejandro Chacón, Santiago Marco-Sola, Antonio Espinosa 0001, Paolo Ribeca, Juan C. Moure
IEEE ACM Trans. Comput. Biol. Bioinform.5
2015 Implementation of the DWT in a GPU through a Register-based Strategy
abstract
The release of the CUDA Kepler architecture in March 2012 has provided Nvidia GPUs with a larger register memory space and instructions for the communication of registers among threads. This facilitates a new programming strategy that utilizes registers for data sharing and reusing in detriment of the shared memory. Such a programming strategy can significantly improve the performance of applications that reuse data heavily. This paper presents a register-based implementation of the Discrete Wavelet Transform (DWT), the prevailing data decorrelation technique in the field of image coding. Experimental results indicate that the proposed method is, at least, four times faster than the best GPU implementation of the DWT found in the literature. Furthermore, theoretical analysis coincide with experimental tests in proving that the execution times achieved by the proposed implementation are close to the GPU's performance limits.
Pablo Enfedaque, Francesc Aulí Llinàs, Juan C. Moure
IEEE Trans. Parallel Distributed Syst.3
2014 Job scheduling in Hadoop with Shared Input Policy and RAMDISK
abstract
Hadoop Framework is a successful option for industry and academia to handle Big Data applications. Large input data sets are split into smaller chunks, distributed among the cluster nodes and processed in the same nodes where they are stored. However, some Hadoop data-intensive applications generate a very large volume of intermediate data to the local file system of each node. Many data spilled to disk associated with concurrent accesses from different tasks that are executed on the same node overload the input/output system. We propose to extend Shared Input Policy, a Hadoop job scheduler policy developed by our research group, by adding a RAMDISK for temporary storage of intermediate data. Shared Input Policy schedules batches of data-intensive jobs that share the same input data set. We add RAMDISK to improve performance of Shared Input Policy. RAMDISK has high throughput and low latency and this allows quick access to intermediate data relieving hard disk. Experimental results show that our approach outperforms Hadoop default policy from 40% to 60% for data intensive applications.
Aprígio Bezerra, Porfidio Hernández, Antonio Espinosa 0001, Juan C. Moure
CLUSTER4
2014 Thread-cooperative, bit-parallel computation of levenshtein distance on GPU
abstract
Approximate string matching is a very important problem in computational biology; it requires the fast computation of string distance as one of its essential components. Myers' bit-parallel algorithm improves the classical dynamic programming approach to Levenshtein distance computation, and offers competitive performance on CPUs. The main challenge when designing an efficient GPU implementation is to expose enough SIMD parallelism while at the same time keeping a relatively small working set for each thread.
Alejandro Chacón, Santiago Marco-Sola, Antonio Espinosa 0001, Paolo Ribeca, Juan C. Moure
ICS5
2014 FM-Index on GPU: A Cooperative Scheme to Reduce Memory Footprint
abstract
The FM-index is a data structure which is seeing more and more pervasive use, in particular in the field of high-throughput bioinformatics. Algorithms based on it show a pseudo-random memory access pattern. As a consequence, they are usually bound by memory bandwidth rather than CPU usage. Naive GPU implementations are no exception. Here we show that the combination of a compact design of the FM-index and a thread-cooperative approach can be used to restore a proper balance. The resulting solution is less memory-bandwidth intensive, and allows full exploitation of the computational resources of the GPU across several GPU architectures.
Alejandro Chacón, Santiago Marco-Sola, Antonio Espinosa 0001, Paolo Ribeca, Juan C. Moure
ISPA5
2013 Job scheduling for optimizing data locality in Hadoop clusters
abstract
We describe the use of non-dedicated clusters by a known group of local applications sharing the computational resources with additional bioinformatics MapReduce applications. We have studied how to effectively use the resources shared by both application types during their execution. In order to keep local application execution times unaffected we consider the configuration of a group of parameters of the Hadoop platform. One of the most relevant aspects to consider is the job scheduling policy. Our aim is to allow that tasks from different jobs that handle the same data blocks are grouped to be run on the same node where the blocks are allocated. Experimental results show that our approach outperforms traditional policies.
Aprígio Bezerra, Porfidio Hernández, Antonio Espinosa 0001, Juan C. Moure
EuroMPI4
2012 Analysis and improvement of map-reduce data distribution in read mapping applications
Antonio Espinosa 0001, Porfidio Hernández, Juan C. Moure, J. Protasio, Ana Ripoll
J. Supercomput.3
2011 Performance Behavior Prediction Scheme for Shared-Memory Parallel Applications
abstract
A current challenge in computing centers with different clusters to run applications is which multicore systems must we choose to run a given shared-memory parallel application. Our proposal is to generate a node performance profile database (NPPDB), composed by performance profiles given by distinct micro benchmark-target node combination. Then, applications are executed on a base node to identify different execution phases and their weights, and to collect performance and functional data for each phase. For similarity, the information to compare behavior is always obtained on the same node. When we want to project performance behavior, we look for similarity using the information from the performance profiles database with the phase characterization, in order to select the appropriate node for running the application.
John Corredor, Juan C. Moure, Dolores Rexachs, Daniel Franco 0002, Emilio Luque
CLUSTER2
2010 A reconfigurable cache memory with heterogeneous banks
abstract
The optimal size of a large on-chip cache can be different for different programs: at some point, the reduction of cache misses achieved when increasing cache size hits diminishing returns, while the higher cache latency hurts performance. This paper presents the Amorphous Cache (AC), a reconfigurable L2 on-chip cache aimed at improving performance as well as reducing energy consumption. AC is composed of heterogeneous sub-caches as opposed to common caches using homogenous sub-caches. The sub-caches are turned off depending on the application workload to conserve power and minimize latencies. A novel reconfiguration algorithm based on Basic Block Vectors is proposed to recognize program phases, and a learning mechanism is used to select the appropriate cache configuration for each program phase. We compare our reconfigurable cache with existing proposals of adaptive and non-adaptive caches. Our results show that the combination of AC and the novel reconfiguration algorithm provides the best power consumption and performance. For example, on average, it reduces the cache access latency by 55.8%, the cache dynamic energy by 46.5%, and the cache leakage power by 49.3% with respect to a non-adaptive cache.
Domingo Benitez, Juan C. Moure, Dolores Rexachs, Emilio Luque
DATE2
2006 Wide and efficient trace prediction using the local trace predictor
abstract
High prediction bandwidth enables performance improvements and power reduction techniques. This paper explores a mechanism to increase prediction width (instructions per prediction) by predicting instruction traces. Our analysis shows that predicting traces including multiple branches is not significantly less accurate than predicting single branches. A novel Local Trace Predictor organization is proposed. It increases prediction width without reducing the ratio of prediction accuracy versus memory resources with respect to a Basic Block Predictor.Compared to the previously proposed Next-Trace Predictor, the Local Trace Predictor reduces memory requirements by codifying trace predictions, and by limiting the number of traces starting at the same instruction to 2 or 4. The limit lessens prediction width only slightly, and does not affect prediction accuracy. The overall result is that the Local Trace Predictor outperforms the Next-Trace Predictor for sizes higher than 12 KBytes.
Juan C. Moure, Domingo Benitez, Dolores Rexachs, Emilio Luque
ICS1
2005 Target Encoding for Efficient Indirect Jump Prediction
Juan C. Moure, Domingo Benitez, Dolores Rexachs, Emilio Luque
Euro-Par1
2005 Performance and Power Evaluation of an Intelligently Adaptive Data Cache
Domingo Benitez, Juan C. Moure, Dolores Rexachs, Emilio Luque
HiPC2
2004 Graduate students learning strategies through research collaboration
abstract
It is already known that the learning process can be accelerated with the mixture of theoretical classes and experimental work. This paper describes an interesting experiment with that combination in the teaching of computer architecture for Ph.D. students in collaboration with a researcher in a real design investigation. As the work progressed, a simple cyclical methodology arose as reference for future works.
Eduardo Argollo, Mauricio Hanzich, Diego Mostaccio, Germán Bianchini, Paula Cecilia Fritzsche, Ferran Bonàs, Emilio Luque, Juan C. Moure, Dolores Rexachs
ITiCSE8
2003 Optimizing a Decoupled Front-End Architecture: The Indexed Fetch Target Buffer (iFTB)
Juan C. Moure, Dolores Rexachs, Emilio Luque
Euro-Par1
2002 Speeding Up Target Address Generation Using a Self-indexed FTB (Research Note)
Juan C. Moure, Dolores Rexachs, Emilio Luque
Euro-Par1
2002 The KScalar simulator
abstract
Modern processors increase their performance with complex microarchitectural mechanisms, which makes them more and more difficult to understand and evaluate. KScalar is a graphical simulation tool that facilitates the study of such processors. It allows students to analyze the performance behavior of a wide range of processor microarchitectures: from a very simple in-order, scalar pipeline, to a detailed out-of-order, superscalar pipeline with non-blocking caches, speculative execution, and complex branch prediction. The simulator interprets executables for the Alpha AXP instruction set: from very short program fragments to large applications. The object's program execution may be simulated in varying levels of detail: either cycle-by-cycle, observing all the pipeline events that determine processor performance, or million cycles at once, taking statistics of the main performance issues.Instructors may use KScalar in several ways. First, it may be used to provide demonstrations in lectures or online learning environments. Second, it allows students to investigate the characteristics of specific processor microarchitectures as practical short assignments associated to a lecture course. Third, students may undertake major projects involving the optimization of real programs at the software-hardware interface, or involving the optimization of a processor microarchitecture for a given application workload.A preliminary version of KScalar has been successfully used in several lecture courses during the last two years in the University Autónoma of Barcelona. It runs on a x86/Linux/KDE system. The graphical interface has been developed using the KDE and QT libraries. The simulator engine running behind the graphical interface is a heavily-modified version of SimpleScalar. KScalar code is available under the terms of the GNU and SimpleScalar General Public License
Juan C. Moure, Dolores Rexachs, Emilio Luque
ACM J. Educ. Resour. Comput.1
2001 Improving Single-Thread Fetch Performance on a Multithreaded Processor
abstract
Multithreaded processors, by simultaneously using both the thread-level parallelism and the instruction-level parallelism of applications, achieve larger instruction per cycle rate than single-thread processors. On a multi-thread workload, a clustered organization maximizes performances. On a single-thread workload, however, all but one of the clusters are idle, degrading single-thread performance significantly. Using a clustered multi-thread performance as a baseline, we propose and analyze several mechanisms and policies to improve single-thread execution exploiting the existing hardware without a significant multi-thread performance loss. We focus on the fetch unit, which is maybe the most performance-critical stage. Essentially, we analyze three ways of exploiting the idle fetch clusters: allowing a single thread accessing its neighbor clusters, use the idle fetch clusters to provide multiple-path execution, or use them to widen the effective single-three fetch block.
Juan C. Moure, R. B. García, Dolores Rexachs, Emilio Luque
DSD1
1994 Programming environment for a transputer based computer
Emilio Luque, Miquel A. Senar, Daniel Franco 0002, Porfidio Hernández, Elisa Heymann, Juan C. Moure
Future Gener. Comput. Syst.6