EDBT 2026 Demo / reviewers in the wild / expert
Juan C. Moure
dblp:59/5413 · also Juan Carlos Moure
· DBLP profile ↗
35ranked-venue papers
6as first author
5since 2021 · last 2023
0000-0001-6697-0331ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 20 · 5 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 2Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
GPUs and heterogeneous computing · 85% Processor architecture and microarchitecture · 8% Parallel and multicore computing · 8% | |
| Artificial intelligence
2 papers |
3D vision · 53% Segmentation and scene understanding · 38% Autonomous driving · 9% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Bioinformatics and computational biology · 100% | |
| Computer graphics and multimedia
3 papers |
Image and video coding · 75% Image and video processing · 19% Audio and music processing · 6% |
Topics — the 19 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
GPUs and heterogeneous computing
GPU computing |
1.7 | 4 | 2023 | WFA-GPU: gap-affine pairwise read-alignment using GPUs · Bioinform. 2023 3D Perception With Slanted Stixels on GPU · IEEE Trans. Parallel Distributed Syst. 2021 GPU Implementation of Bitplane Coding with Parallel Coefficient Processing for High Performance Image Compression · IEEE Trans. Parallel Distributed Syst. 2017 |
Bioinformatics and computational biology
sequence alignment |
0.7 | 2 | 2023 | Fast gap-affine pairwise alignment using the wavefront algorithm · Bioinform. 2021 WFA-GPU: gap-affine pairwise read-alignment using GPUs · Bioinform. 2023 |
GPUs and heterogeneous computing
GPU-accelerated bioinformatics |
0.7 | 1 | 2023 | WFA-GPU: gap-affine pairwise read-alignment using GPUs · Bioinform. 2023 |
Image and video coding
bit-plane coding |
0.5 | 2 | 2017 | GPU Implementation of Bitplane Coding with Parallel Coefficient Processing for High Performance Image Compression · IEEE Trans. Parallel Distributed Syst. 2017 Bitplane Image Coding With Parallel Coefficient Processing · IEEE Trans. Image Process. 2016 |
Computer vision › 3D vision
depth estimation |
0.5 | 1 | 2021 | 3D Perception With Slanted Stixels on GPU · IEEE Trans. Parallel Distributed Syst. 2021 |
Bioinformatics and computational biology › sequence alignment
pairwise sequence alignment |
0.5 | 1 | 2021 | Fast gap-affine pairwise alignment using the wavefront algorithm · Bioinform. 2021 |
GPUs and heterogeneous computing › embedded GPU
embedded GPU acceleration |
0.5 | 1 | 2021 | 3D Perception With Slanted Stixels on GPU · IEEE Trans. Parallel Distributed Syst. 2021 |
Computer vision › Segmentation and scene understanding
semantic segmentation |
0.4 | 1 | 2019 | Slanted Stixels: A Way to Represent Steep Streets · Int. J. Comput. Vis. 2019 |
Computer vision › 3D vision
stereo vision |
0.4 | 1 | 2019 | Slanted Stixels: A Way to Represent Steep Streets · Int. J. Comput. Vis. 2019 |
Parallel and multicore computing › parallel algorithms › parallel image processing
parallel image coding |
0.2 | 1 | 2016 | Bitplane Image Coding With Parallel Coefficient Processing · IEEE Trans. Image Process. 2016 |
Processor architecture and microarchitecture
SIMD |
0.2 | 1 | 2016 | Bitplane Image Coding With Parallel Coefficient Processing · IEEE Trans. Image Process. 2016 |
Image and video coding › transform coding
discrete wavelet transform |
0.2 | 1 | 2015 | Implementation of the DWT in a GPU through a Register-based Strategy · IEEE Trans. Parallel Distributed Syst. 2015 |
Image and video processing
wavelet transform |
0.2 | 1 | 2015 | Implementation of the DWT in a GPU through a Register-based Strategy · IEEE Trans. Parallel Distributed Syst. 2015 |
Bioinformatics and computational biology › sequence analysis › read mapping
long-read alignment |
0.2 | 1 | 2023 | WFA-GPU: gap-affine pairwise read-alignment using GPUs · Bioinform. 2023 |
Robotics › Autonomous driving
perception |
0.1 | 1 | 2021 | 3D Perception With Slanted Stixels on GPU · IEEE Trans. Parallel Distributed Syst. 2021 |
Computer vision › Segmentation and scene understanding › scene understanding
semantic scene understanding |
0.1 | 1 | 2021 | 3D Perception With Slanted Stixels on GPU · IEEE Trans. Parallel Distributed Syst. 2021 |
Computer vision › Segmentation and scene understanding
scene understanding |
0.1 | 1 | 2019 | Slanted Stixels: A Way to Represent Steep Streets · Int. J. Comput. Vis. 2019 |
Image and video coding › image compression
wavelet-based image coding |
0.1 | 1 | 2017 | GPU Implementation of Bitplane Coding with Parallel Coefficient Processing for High Performance Image Compression · IEEE Trans. Parallel Distributed Syst. 2017 |
Audio and music processing
decorrelation |
0.1 | 1 | 2015 | Implementation of the DWT in a GPU through a Register-based Strategy · IEEE Trans. Parallel Distributed Syst. 2015 |
Methods — techniques the papers use, named apart from their topics
dynamic programming · 2.8wavefront alignment algorithm · 1.3GPU parallelization · 1.0thread-to-data mapping · 0.6memory management · 0.6context modeling · 0.5arithmetic coding · 0.5SIMD vectorization · 0.5SIMD processing · 0.5register sharing · 0.4CUDA · 0.4over-segmentation · 0.4global energy minimization · 0.4fully convolutional network · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | WFA-GPU: gap-affine pairwise read-alignment using GPUsabstractMOTIVATION: Advances in genomics and sequencing technologies demand faster and more scalable analysis methods that can process longer sequences with higher accuracy. However, classical pairwise alignment methods, based on dynamic programming (DP), impose impractical computational requirements to align long and noisy sequences like those produced by PacBio and Nanopore technologies. The recently proposed wavefront alignment (WFA) algorithm paves the way for more efficient alignment tools, improving time and memory complexity over previous methods. However, high-performance computing (HPC) platforms require efficient parallel algorithms and tools to exploit the computing resources available on modern accelerator-based architectures. RESULTS: This paper presents WFA-GPU, a GPU (graphics processing unit)-accelerated tool to compute exact gap-affine alignments based on the WFA algorithm. We present the algorithmic adaptations and performance optimizations that allow exploiting the massively parallel capabilities of modern GPU devices to accelerate the alignment computations. In particular, we propose a CPU-GPU co-design capable of performing inter-sequence and intra-sequence parallel sequence alignment, combining a succinct WFA-data representation with an efficient GPU implementation. As a result, we demonstrate that our implementation outperforms the original multi-threaded WFA implementation by up to 4.3× and up to 18.2× when using heuristic methods on long and noisy sequences. Compared to other state-of-the-art tools and libraries, the WFA-GPU is up to 29× faster than other GPU implementations and up to four orders of magnitude faster than other CPU implementations. Furthermore, WFA-GPU is the only GPU solution capable of correctly aligning long reads using a commodity GPU. AVAILABILITY AND IMPLEMENTATION: WFA-GPU code and documentation are publicly available at https://github.com/quim0/WFA-GPU. Quim Aguado-Puig, Max Doblas, Christos Matzoros, Antonio Espinosa 0001, Juan C. Moure, Santiago Marco-Sola, Miquel Moretó |
Bioinform. | 5 |
| 2021 | OpenCL-based FPGA Accelerator for Semi-Global Approximate String Matching Using Diagonal Bit-VectorsabstractAn FPGA accelerator for the computation of the semi-global Levenshtein distance between a pattern and a reference text is presented. The accelerator provides an important benefit to reduce the execution time of read-mappers used in short-read genomic sequencing. Previous attempts to solve the same problem in FPGA use the Myers algorithm following a column approach to compute the dynamic programming table. We use an approach based on diagonals that allows for some resource savings while maintaining a very high throughput of 1 alignment per clock cycle. The design is implemented in OpenCL and tested on two FPGA accelerators. The maximum performance obtained is 91.5 MPairs/s for 100 × 120 sequences and 47 MPairs/s for 300 × 360 sequences, the highest ever reported for this problem. David Castells-Rufas, Santiago Marco-Sola, Quim Aguado-Puig, Antonio Espinosa 0001, Juan C. Moure, Lluc Alvarez, Miquel Moretó |
FPL | 5 |
| 2021 | Fast gap-affine pairwise alignment using the wavefront algorithmabstractMOTIVATION: Pairwise alignment of sequences is a fundamental method in modern molecular biology, implemented within multiple bioinformatics tools and libraries. Current advances in sequencing technologies press for the development of faster pairwise alignment algorithms that can scale with increasing read lengths and production yields. RESULTS: In this article, we present the wavefront alignment algorithm (WFA), an exact gap-affine algorithm that takes advantage of homologous regions between the sequences to accelerate the alignment process. As opposed to traditional dynamic programming algorithms that run in quadratic time, the WFA runs in time O(ns), proportional to the read length n and the alignment score s, using O(s2) memory. Furthermore, our algorithm exhibits simple data dependencies that can be easily vectorized, even by the automatic features of modern compilers, for different architectures, without the need to adapt the code. We evaluate the performance of our algorithm, together with other state-of-the-art implementations. As a result, we demonstrate that the WFA runs 20-300× faster than other methods aligning short Illumina-like sequences, and 10-100× faster using long noisy reads like those produced by Oxford Nanopore Technologies. AVAILABILITY AND IMPLEMENTATION: The WFA algorithm is implemented within the wavefront-aligner library, and it is publicly available at https://github.com/smarco/WFA. Santiago Marco-Sola, Juan C. Moure, Miquel Moretó, Antonio Espinosa 0001 |
Bioinform. | 2 |
| 2021 | Real-time 16K video coding on a GPU with complexity scalable BPC-PaCoabstractThe advent of new technologies such as high dynamic range or 8K screens has enhanced the quality of digital images but it has also increased the codecs’ computational demands to process such data. This paper presents a video codec that, while providing the same coding features and performance as those of JPEG2000, can process 16K video in real time using a consumer-grade GPU. This high throughput is achieved with a technique that introduces complexity scalability to a bitplane coding engine, which is the most computationally complex stage of the coding pipeline. The resulting codec can trade throughput for coding performance depending on the user’s needs. Experimental results suggest that our method can double the throughput achieved by CPU implementations of the recently approved High-Throughput JPEG2000 and by hardwired implementations of HEVC in a GPU. Carlos de Cea-Dominguez, Juan C. Moure, Joan Bartrina-Rapesta, Francesc Aulí Llinàs |
Signal Process. Image Commun. | 2 |
| 2021 | 3D Perception With Slanted Stixels on GPUabstractThis article presents a GPU-accelerated software design of the recently proposed model of Slanted Stixels, which represents the geometric and semantic information of a scene in a compact and accurate way. We reformulate the measurement depth model to reduce the computational complexity of the algorithm, relying on the confidence of the depth estimation and the identification of invalid values to handle outliers. The proposed massively parallel scheme and data layout for the irregular computation pattern that corresponds to a Dynamic Programming paradigm is described and carefully analyzed in performance terms. Performance is shown to scale gracefully on current generation embedded GPUs. We assess the proposed methods in terms of semantic and geometric accuracy as well as run-time performance on three publicly available benchmark datasets. Our approach achieves real-time performance with high accuracy for 2048 × 1024 image sizes and 4 × 4 Stixel resolution on the low-power embedded GPU of an NVIDIA Tegra Xavier. Daniel Hernández Juárez, Antonio Espinosa 0001, David Vázquez 0001, Antonio M. López 0001, Juan C. Moure |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2020 | Complexity Scalable Bitplane Image Coding With Parallel Coefficient ProcessingabstractVery fast image and video codecs are a pursued goal both in the academia and the industry. This paper presents a complexity scalable and parallel bitplane coding engine for wavelet-based image codecs. The proposed method processes the coefficients in parallel, suiting hardware architectures based on vector instructions. Our previous work is extended with a mechanism that provides complexity scalability to the system. Such a feature allows the coder to regulate the throughput achieved at the expense of slightly penalizing compression efficiency. Experimental results suggests that, when using the fastest speed, the method almost doubles the throughput of our previous engine while penalizing compression efficiency by about 10%. Carlos de Cea-Dominguez, Juan C. Moure, Joan Bartrina-Rapesta, Francesc Aulí Llinàs |
IEEE Signal Process. Lett. | 2 |
| 2019 | Slanted Stixels: A Way to Represent Steep StreetsabstractAbstract This work presents and evaluates a novel compact scene representation based on Stixels that infers geometric and semantic information. Our approach overcomes the previous rather restrictive geometric assumptions for Stixels by introducing a novel depth model to account for non-flat roads and slanted objects. Both semantic and depth cues are used jointly to infer the scene representation in a sound global energy minimization formulation. Furthermore, a novel approximation scheme is introduced in order to significantly reduce the computational complexity of the Stixel algorithm, and then achieve real-time computation capabilities. The idea is to first perform an over-segmentation of the image, discarding the unlikely Stixel cuts, and apply the algorithm only on the remaining Stixel cuts. This work presents a novel over-segmentation strategy based on a fully convolutional network, which outperforms an approach based on using local extrema of the disparity map. We evaluate the proposed methods in terms of semantic and geometric accuracy as well as run-time on four publicly available benchmark datasets. Our approach maintains accuracy on flat road scene datasets while improving substantially on a novel non-flat road dataset. Daniel Hernández Juárez, Lukas Schneider, Pau Cebrian, Antonio Espinosa 0001, David Vázquez 0001, Antonio M. López 0001, Uwe Franke, Marc Pollefeys, Juan C. Moure |
Int. J. Comput. Vis. | 9 |
| 2018 | High Throughput Image Codec for High-Resolution Satellite ImagesabstractThe growth in the use of satellite images has generated the need for their fast compression, processing, and distribution. JPEG2000 is a widespread standard for the compression and transmission of such images once they are in the ground. Despite its advanced features and excellent coding performance, JPEG2000 demands significant computational resources. This paper introduces a wavelet-based codec that uses the JPEG2000 framework, but replaces its most computationally demanding coding stage by a highly parallel engine. When executed in Graphics Processing Units to code high-resolution satellite images, the proposed codec achieves speed-ups of up to 8× when compared to the fastest implementation of JPEG2000 executed in a multi-core platform. Carlos de Cea-Dominguez, Pablo Enfedaque, Juan C. Moure, Joan Bartrina-Rapesta, Francesc Aulí Llinàs |
IGARSS | 3 |
| 2017 | Slanted Stixels: Representing San Francisco's Steepest Streets
Daniel Hernández Juárez, Lukas Schneider, Antonio Espinosa 0001, Juan C. Moure, David Vázquez 0001, Antonio M. López 0001, Uwe Franke, Marc Pollefeys |
BMVC | 4 |
| 2017 | GPU-Accelerated Real-Time Stixel ComputationabstractThe Stixel World is a medium-level, compact representation of road scenes that abstracts millions of disparity pixels into hundreds or thousands of stixels. The goal of this work is to implement and evaluate a complete multistixel estimation pipeline on an embedded, energy-efficient, GPU-accelerated device. This work presents a full GPUaccelerated implementation of stixel estimation that produces reliable results at 26 frames per second (real-time) on the Tegra X1 for disparity images of 1024×440 pixels and stixel widths of 5 pixels, and achieves more than 400 frames per second on a high-end Titan X GPU card. Daniel Hernández Juárez, Antonio Espinosa 0001, Juan C. Moure, David Vázquez 0001, Antonio M. López 0001 |
WACV | 3 |
| 2017 | Coalition structure generation problems: optimization and parallelization of the IDP algorithm in multicore systemsabstractSummary The coalition structure generation problem is well known in the area of multi‐agent systems. Its goal is to establish coalitions between agents while maximizing the global welfare. Among the existing different algorithms designed to solve the coalition structure generation problem, DP and IDP are the ones with smaller temporal complexity. After analyzing the operation of the dynamic programming and improved dynamic programming algorithms, we have identified which are the most frequent operations and propose an optimized method. In addition, we study and implement a method for dividing the work into different threads. To describe incremental improvements of the algorithm design, we first compare performance of an improved single central processing unit core version where we obtain speedups ranging from 7 × to 11 × . Then, we describe the best resource use in a multi‐thread optimized version where we obtain an additional 7.5 × speedup running in a 12‐core machine. Francisco Cruz-Mencia, Antonio Espinosa 0001, Juan C. Moure, Jesús Cerquides, Juan A. Rodríguez-Aguilar, Kim Svensson, Sarvapali D. Ramchurn |
Concurr. Comput. Pract. Exp. | 3 |
| 2017 | Introducing computational thinking, parallel programming and performance engineering in interdisciplinary studies
Eduardo César, Ana Cortés, Antonio Espinosa 0001, Tomàs Margalef, Juan C. Moure, Anna Sikora, Remo Suppi |
J. Parallel Distributed Comput. | 5 |
| 2017 | GPU Implementation of Bitplane Coding with Parallel Coefficient Processing for High Performance Image CompressionabstractThe fast compression of images is a requisite in many applications like TV production, teleconferencing, or digital cinema. Many of the algorithms employed in current image compression standards are inherently sequential. High performance implementations of such algorithms often require specialized hardware like field integrated gate arrays. Graphics Processing Units (GPUs) do not commonly achieve high performance on these algorithms because they do not exhibit fine-grain parallelism. Our previous work introduced a new core algorithm for wavelet-based image coding systems. It is tailored for massive parallel architectures. It is called bitplane coding with parallel coefficient processing (BPC-PaCo). This paper introduces the first high performance, GPUbased implementation of BPC-PaCo. A detailed analysis of the algorithm aids its implementation in the GPU. The main insights behind the proposed codec are an efficient thread-to-data mapping, a smart memory management, and the use of efficient cooperation mechanisms to enable inter-thread communication. Experimental results indicate that the proposed implementation matches the requirements for high resolution (4 K) digital cinema in real time, yielding speedups of 30× with respect to the fastest implementations of current compression standards. Also, a power consumption evaluation shows that our implementation consumes 40× less energy for equivalent performance than state-of-the-art methods. Pablo Enfedaque, Francesc Aulí Llinàs, Juan C. Moure |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2016 | Bitplane Image Coding With Parallel Coefficient ProcessingabstractImage coding systems have been traditionally tailored for multiple instruction, multiple data (MIMD) computing. In general, they partition the (transformed) image in codeblocks that can be coded in the cores of MIMD-based processors. Each core executes a sequential flow of instructions to process the coefficients in the codeblock, independently and asynchronously from the others cores. Bitplane coding is a common strategy to code such data. Most of its mechanisms require sequential processing of the coefficients. The last years have seen the upraising of processing accelerators with enhanced computational performance and power efficiency whose architecture is mainly based on the single instruction, multiple data (SIMD) principle. SIMD computing refers to the execution of the same instruction to multiple data in a lockstep synchronous way. Unfortunately, current bitplane coding strategies cannot fully profit from such processors due to inherently sequential coding task. This paper presents bitplane image coding with parallel coefficient (BPC-PaCo) processing, a coding method that can process many coefficients within a codeblock in parallel and synchronously. To this end, the scanning order, the context formation, the probability model, and the arithmetic coder of the coding engine have been re-formulated. The experimental results suggest that the penalization in coding performance of BPC-PaCo with respect to the traditional strategies is almost negligible. Francesc Aulí Llinàs, Pablo Enfedaque, Juan C. Moure, Victor Sanchez |
IEEE Trans. Image Process. | 3 |
| 2015 | Parallelisation and Application of AD 3 as a Method for Solving Large Scale Combinatorial Auctions
Francisco Cruz-Mencia, Jesús Cerquides, Antonio Espinosa 0001, Juan C. Moure, Juan A. Rodríguez-Aguilar |
COORDINATION | 4 |
| 2015 | Strategy of Microscopic Parallelism for Bitplane Image CodingabstractRecent years have seen the upraising of a new type of processors strongly relying on the Single Instruction, Multiple Data (SIMD) architectural principle. The main idea behind SIMD computing is to apply a flow of instructions to multiple pieces of data in parallel and synchronously. This permits the execution of thousands of operations in parallel, achieving higher computational performance than with traditional Multiple Instruction, Multiple Data (MIMD) architectures. The level of parallelism required in SIMD computing can only be achieved in image coding systems via microscopic parallel strategies that code multiple coefficients in parallel. Until now, the only way to achieve microscopic parallelism in bit plane coding engines was by executing multiple coding passes in parallel. Such a strategy does not suit well SIMD computing because each thread executes different instructions. This paper introduces the first bit plane coding engine devised for the fine grain of parallelism required in SIMD computing. Its main insight is to allow parallel coefficient processing in a coding pass. Experimental tests show coding performance results similar to those of JPEG2000. Francesc Aulí Llinàs, Pablo Enfedaque, Juan C. Moure, Ian Blanes, Victor Sanchez |
DCC | 3 |
| 2015 | Strategies of SIMD Computing for Image Coding in GPUabstractThe main difficulty to implement modern image coding systems in a GPU is that the algorithms employed in the core of the coding scheme are inherently sequential. We recently proposed bitplane image coding with parallel coefficient processing (BPC-PaCo), a coding scheme that, contrarily to most systems, permits the processing of multiple coefficients of the image in parallel. This enables the use of SIMD computing, ideal for its implementation in a GPU. This paper introduces and evaluates the GPU implementation of BPC-PaCo employing two different strategies that tradeoff computational throughput and compression efficiency. The proposed implementation is compared to the best CPU and GPU implementations of JPEG2000, the state-of-the-art image compression standard. Experimental results indicate that BPC-PaCo achieves a computational throughput that is an order of magnitude superior to that achieved with such implementations with a small reduction in coding efficiency. Pablo Enfedaque, Francesc Aulí Llinàs, Juan C. Moure |
HiPC | 3 |
| 2015 | Boosting the FM-Index on the GPU: Effective Techniques to Mitigate Random Memory AccessabstractThe recent advent of high-throughput sequencing machines producing big amounts of short reads has boosted the interest in efficient string searching techniques. As of today, many mainstream sequence alignment software tools rely on a special data structure, called the FM-index, which allows for fast exact searches in large genomic references. However, such searches translate into a pseudo-random memory access pattern, thus making memory access the limiting factor of all computation-efficient implementations, both on CPUs and GPUs. Here, we show that several strategies can be put in place to remove the memory bottleneck on the GPU: more compact indexes can be implemented by having more threads work cooperatively on larger memory blocks, and a k-step FM-index can be used to further reduce the number of memory accesses. The combination of those and other optimisations yields an implementation that is able to process about two Gbases of queries per second on our test platform, being about 8 × faster than a comparable multi-core CPU version, and about 3 × to 5 × faster than the FM-index implementation on the GPU provided by the recently announced Nvidia NVBIO bioinformatics library. Alejandro Chacón, Santiago Marco-Sola, Antonio Espinosa 0001, Paolo Ribeca, Juan C. Moure |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2015 | Implementation of the DWT in a GPU through a Register-based StrategyabstractThe release of the CUDA Kepler architecture in March 2012 has provided Nvidia GPUs with a larger register memory space and instructions for the communication of registers among threads. This facilitates a new programming strategy that utilizes registers for data sharing and reusing in detriment of the shared memory. Such a programming strategy can significantly improve the performance of applications that reuse data heavily. This paper presents a register-based implementation of the Discrete Wavelet Transform (DWT), the prevailing data decorrelation technique in the field of image coding. Experimental results indicate that the proposed method is, at least, four times faster than the best GPU implementation of the DWT found in the literature. Furthermore, theoretical analysis coincide with experimental tests in proving that the execution times achieved by the proposed implementation are close to the GPU's performance limits. Pablo Enfedaque, Francesc Aulí Llinàs, Juan C. Moure |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2014 | Job scheduling in Hadoop with Shared Input Policy and RAMDISKabstractHadoop Framework is a successful option for industry and academia to handle Big Data applications. Large input data sets are split into smaller chunks, distributed among the cluster nodes and processed in the same nodes where they are stored. However, some Hadoop data-intensive applications generate a very large volume of intermediate data to the local file system of each node. Many data spilled to disk associated with concurrent accesses from different tasks that are executed on the same node overload the input/output system. We propose to extend Shared Input Policy, a Hadoop job scheduler policy developed by our research group, by adding a RAMDISK for temporary storage of intermediate data. Shared Input Policy schedules batches of data-intensive jobs that share the same input data set. We add RAMDISK to improve performance of Shared Input Policy. RAMDISK has high throughput and low latency and this allows quick access to intermediate data relieving hard disk. Experimental results show that our approach outperforms Hadoop default policy from 40% to 60% for data intensive applications. Aprígio Bezerra, Porfidio Hernández, Antonio Espinosa 0001, Juan C. Moure |
CLUSTER | 4 |
| 2014 | Thread-cooperative, bit-parallel computation of levenshtein distance on GPUabstractApproximate string matching is a very important problem in computational biology; it requires the fast computation of string distance as one of its essential components. Myers' bit-parallel algorithm improves the classical dynamic programming approach to Levenshtein distance computation, and offers competitive performance on CPUs. The main challenge when designing an efficient GPU implementation is to expose enough SIMD parallelism while at the same time keeping a relatively small working set for each thread. Alejandro Chacón, Santiago Marco-Sola, Antonio Espinosa 0001, Paolo Ribeca, Juan C. Moure |
ICS | 5 |
| 2014 | FM-Index on GPU: A Cooperative Scheme to Reduce Memory FootprintabstractThe FM-index is a data structure which is seeing more and more pervasive use, in particular in the field of high-throughput bioinformatics. Algorithms based on it show a pseudo-random memory access pattern. As a consequence, they are usually bound by memory bandwidth rather than CPU usage. Naive GPU implementations are no exception. Here we show that the combination of a compact design of the FM-index and a thread-cooperative approach can be used to restore a proper balance. The resulting solution is less memory-bandwidth intensive, and allows full exploitation of the computational resources of the GPU across several GPU architectures. Alejandro Chacón, Santiago Marco-Sola, Antonio Espinosa 0001, Paolo Ribeca, Juan C. Moure |
ISPA | 5 |
| 2013 | Job scheduling for optimizing data locality in Hadoop clustersabstractWe describe the use of non-dedicated clusters by a known group of local applications sharing the computational resources with additional bioinformatics MapReduce applications. We have studied how to effectively use the resources shared by both application types during their execution. In order to keep local application execution times unaffected we consider the configuration of a group of parameters of the Hadoop platform. One of the most relevant aspects to consider is the job scheduling policy. Our aim is to allow that tasks from different jobs that handle the same data blocks are grouped to be run on the same node where the blocks are allocated. Experimental results show that our approach outperforms traditional policies. Aprígio Bezerra, Porfidio Hernández, Antonio Espinosa 0001, Juan C. Moure |
EuroMPI | 4 |
| 2012 | Analysis and improvement of map-reduce data distribution in read mapping applications
Antonio Espinosa 0001, Porfidio Hernández, Juan C. Moure, J. Protasio, Ana Ripoll |
J. Supercomput. | 3 |
| 2011 | Performance Behavior Prediction Scheme for Shared-Memory Parallel ApplicationsabstractA current challenge in computing centers with different clusters to run applications is which multicore systems must we choose to run a given shared-memory parallel application. Our proposal is to generate a node performance profile database (NPPDB), composed by performance profiles given by distinct micro benchmark-target node combination. Then, applications are executed on a base node to identify different execution phases and their weights, and to collect performance and functional data for each phase. For similarity, the information to compare behavior is always obtained on the same node. When we want to project performance behavior, we look for similarity using the information from the performance profiles database with the phase characterization, in order to select the appropriate node for running the application. John Corredor, Juan C. Moure, Dolores Rexachs, Daniel Franco 0002, Emilio Luque |
CLUSTER | 2 |
| 2010 | A reconfigurable cache memory with heterogeneous banksabstractThe optimal size of a large on-chip cache can be different for different programs: at some point, the reduction of cache misses achieved when increasing cache size hits diminishing returns, while the higher cache latency hurts performance. This paper presents the Amorphous Cache (AC), a reconfigurable L2 on-chip cache aimed at improving performance as well as reducing energy consumption. AC is composed of heterogeneous sub-caches as opposed to common caches using homogenous sub-caches. The sub-caches are turned off depending on the application workload to conserve power and minimize latencies. A novel reconfiguration algorithm based on Basic Block Vectors is proposed to recognize program phases, and a learning mechanism is used to select the appropriate cache configuration for each program phase. We compare our reconfigurable cache with existing proposals of adaptive and non-adaptive caches. Our results show that the combination of AC and the novel reconfiguration algorithm provides the best power consumption and performance. For example, on average, it reduces the cache access latency by 55.8%, the cache dynamic energy by 46.5%, and the cache leakage power by 49.3% with respect to a non-adaptive cache. Domingo Benitez, Juan C. Moure, Dolores Rexachs, Emilio Luque |
DATE | 2 |
| 2006 | Wide and efficient trace prediction using the local trace predictorabstractHigh prediction bandwidth enables performance improvements and power reduction techniques. This paper explores a mechanism to increase prediction width (instructions per prediction) by predicting instruction traces. Our analysis shows that predicting traces including multiple branches is not significantly less accurate than predicting single branches. A novel Local Trace Predictor organization is proposed. It increases prediction width without reducing the ratio of prediction accuracy versus memory resources with respect to a Basic Block Predictor.Compared to the previously proposed Next-Trace Predictor, the Local Trace Predictor reduces memory requirements by codifying trace predictions, and by limiting the number of traces starting at the same instruction to 2 or 4. The limit lessens prediction width only slightly, and does not affect prediction accuracy. The overall result is that the Local Trace Predictor outperforms the Next-Trace Predictor for sizes higher than 12 KBytes. Juan C. Moure, Domingo Benitez, Dolores Rexachs, Emilio Luque |
ICS | 1 |
| 2005 | Target Encoding for Efficient Indirect Jump Prediction
Juan C. Moure, Domingo Benitez, Dolores Rexachs, Emilio Luque |
Euro-Par | 1 |
| 2005 | Performance and Power Evaluation of an Intelligently Adaptive Data Cache
Domingo Benitez, Juan C. Moure, Dolores Rexachs, Emilio Luque |
HiPC | 2 |
| 2004 | Graduate students learning strategies through research collaborationabstractIt is already known that the learning process can be accelerated with the mixture of theoretical classes and experimental work. This paper describes an interesting experiment with that combination in the teaching of computer architecture for Ph.D. students in collaboration with a researcher in a real design investigation. As the work progressed, a simple cyclical methodology arose as reference for future works. Eduardo Argollo, Mauricio Hanzich, Diego Mostaccio, Germán Bianchini, Paula Cecilia Fritzsche, Ferran Bonàs, Emilio Luque, Juan C. Moure, Dolores Rexachs |
ITiCSE | 8 |
| 2003 | Optimizing a Decoupled Front-End Architecture: The Indexed Fetch Target Buffer (iFTB)
Juan C. Moure, Dolores Rexachs, Emilio Luque |
Euro-Par | 1 |
| 2002 | Speeding Up Target Address Generation Using a Self-indexed FTB (Research Note)
Juan C. Moure, Dolores Rexachs, Emilio Luque |
Euro-Par | 1 |
| 2002 | The KScalar simulatorabstractModern processors increase their performance with complex microarchitectural mechanisms, which makes them more and more difficult to understand and evaluate. KScalar is a graphical simulation tool that facilitates the study of such processors. It allows students to analyze the performance behavior of a wide range of processor microarchitectures: from a very simple in-order, scalar pipeline, to a detailed out-of-order, superscalar pipeline with non-blocking caches, speculative execution, and complex branch prediction. The simulator interprets executables for the Alpha AXP instruction set: from very short program fragments to large applications. The object's program execution may be simulated in varying levels of detail: either cycle-by-cycle, observing all the pipeline events that determine processor performance, or million cycles at once, taking statistics of the main performance issues.Instructors may use KScalar in several ways. First, it may be used to provide demonstrations in lectures or online learning environments. Second, it allows students to investigate the characteristics of specific processor microarchitectures as practical short assignments associated to a lecture course. Third, students may undertake major projects involving the optimization of real programs at the software-hardware interface, or involving the optimization of a processor microarchitecture for a given application workload.A preliminary version of KScalar has been successfully used in several lecture courses during the last two years in the University Autónoma of Barcelona. It runs on a x86/Linux/KDE system. The graphical interface has been developed using the KDE and QT libraries. The simulator engine running behind the graphical interface is a heavily-modified version of SimpleScalar. KScalar code is available under the terms of the GNU and SimpleScalar General Public License Juan C. Moure, Dolores Rexachs, Emilio Luque |
ACM J. Educ. Resour. Comput. | 1 |
| 2001 | Improving Single-Thread Fetch Performance on a Multithreaded ProcessorabstractMultithreaded processors, by simultaneously using both the thread-level parallelism and the instruction-level parallelism of applications, achieve larger instruction per cycle rate than single-thread processors. On a multi-thread workload, a clustered organization maximizes performances. On a single-thread workload, however, all but one of the clusters are idle, degrading single-thread performance significantly. Using a clustered multi-thread performance as a baseline, we propose and analyze several mechanisms and policies to improve single-thread execution exploiting the existing hardware without a significant multi-thread performance loss. We focus on the fetch unit, which is maybe the most performance-critical stage. Essentially, we analyze three ways of exploiting the idle fetch clusters: allowing a single thread accessing its neighbor clusters, use the idle fetch clusters to provide multiple-path execution, or use them to widen the effective single-three fetch block. Juan C. Moure, R. B. García, Dolores Rexachs, Emilio Luque |
DSD | 1 |
| 1994 | Programming environment for a transputer based computer
Emilio Luque, Miquel A. Senar, Daniel Franco 0002, Porfidio Hernández, Elisa Heymann, Juan C. Moure |
Future Gener. Comput. Syst. | 6 |