VLDB 2026 Research / reviewers in the wild / expert
Carlos García 0001
dblp:21/5403-1 · also Carlos García Sánchez 0001
· DBLP profile ↗
27ranked-venue papers
2as first author
9since 2021 · last 2026
0000-0002-3470-1097ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 7Artificial intelligence and machine learning · 4 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A comparative performance and efficiency analysis of Apple's M architectures: A GEMM case studyabstractThis paper evaluates the performance and energy efficiency of Apple processors across multiple ARM-based M-series generations and models (standard and Pro). The study is motivated by the increasing heterogeneity of Apple’s SoC architectures, which integrate multiple computing engines raising the scientific question of which hardware components are best suited for executing general-purpose and domain-specific computations such as the GEneral Matrix Multiply ( GEMM ). The analysis focuses on four key components: the Central Processing Unit (CPU), the Graphics Processing Unit (GPU), the matrix calculation accelerator (AMX), and the Apple Neural Engine (ANE). The assessments use the GEMM as benchmark to characterize the performance of the CPU and GPU, alongside tests on AMX, which is specialized in handling large-scale mathematical operations, and tests on the ANE, which is specifically designed for Deep Learning purposes. Additionally, energy consumption data has been collected to analyze the energy efficiency of the aforementioned resources. Results highlight notable improvements in computational capacity and energy efficiency over successive generations. On one hand, the AMX stands out as the most efficient component for FP32 and FP64 workloads, significantly boosting overall system performance. In the M4 Pro, which integrates two matrix accelerators, it achieves up to 68% of the GPU’s FP32 performance while consuming only 42% of its power. On the other hand, the ANE, although limited to FP16 precision, excels in energy efficiency for low-precision tasks, surpassing other accelerators with over 700 GFLOPs/Watt under batched workloads. This analysis offers a clear understanding of how Apple’s custom ARM designs optimize both performance and energy use, particularly in the context of multi-core processing and specialized acceleration units. In addition, a significant contribution of this study is the comprehensive comparative analysis of Apple’s accelerators, which have previously been poorly documented and scarcely studied. The analysis spans different generations and compares the accelerators against both CPU and GPU performance. Sandra Catalán, Rafael Rodríguez-Sánchez 0001, Carlos García 0001, Luis Piñuel Moreno |
Future Gener. Comput. Syst. | 3 |
| 2025 | Evaluation of Juliana Tool: A translator for Julia's CUDA.jl code into KernelAbstraction.jlabstractJulia is a high-level language that supports the execution of parallel code through various packages. CUDA.jl is widely used for developing GPU-accelerated code in Julia and is integrated into many libraries and programs. In this paper, we present Juliana, a novel tool that automatically translates Julia code utilizing the CUDA.jl package into an abstract, multi-backend representation powered by the KernelAbstractions.jl package. To evaluate the tool’s viability and performance, we analyzed four Julia projects: Rodinia, miniBUDE, BabelStream, and Oceananigans.jl. The performance overhead of this approach was found to be relatively low (under 7% for the Rodinia suite), with performance portability metrics showing results nearly identical to the native implementations. By running the same code across multiple KernelAbstractions’ backends, we successfully executed these translated projects on GPUs from vendors such as NVIDIA, Intel, AMD, and Apple. This ensured compatibility across these platforms and enabled first-time execution on some devices. • The paper proposes Juliana ( Julia unification la yer for Intel, AMD, Nvidia and Apple) tool. • Juliana translates Julia code based on CUDA.jl to KernelAbstractions to use GPUs. • Demonstration of Juliana’s low performance overhead across multiple hardware platforms. • Juliana tested on Rodinia, miniBUDE, BabelStream, and Oceananigans for applicability. • Introduction of new features makes Juliana versatile for real-world applications. Enrique de la Calle, Carlos García 0001 |
Future Gener. Comput. Syst. | 2 |
| 2025 | Analyzing the performance portability of SYCL across CPUs, GPUs, and hybrid systems with SW sequence alignmentabstractThe high-performance computing (HPC) landscape is undergoing rapid transformation, with an increasing emphasis on energy-efficient and heterogeneous computing environments. This comprehensive study extends our previous research on SYCL’s performance portability by evaluating its effectiveness across a broader spectrum of computing architectures , including CPUs, GPUs , and hybrid CPU–GPU configurations from NVIDIA, Intel, and AMD. Our analysis covers single-GPU, multi-GPU, single-CPU, and CPU–GPU hybrid setups, using two common, bioinformatic applications as a case study . The results demonstrate SYCL’s versatility across different architectures, maintaining comparable performance to CUDA on NVIDIA GPUs while achieving similar architectural efficiency rates on AMD and Intel GPUs in the majority of cases tested. SYCL also demonstrated remarkable versatility and effectiveness across CPUs from various manufacturers, including the latest hybrid architectures from Intel. Although SYCL showed excellent functional portability in hybrid CPU–GPU configurations, performance varied significantly based on specific hardware combinations. Some performance limitations were identified in multi-GPU and CPU–GPU configurations, primarily attributed to workload distribution strategies rather than SYCL-specific constraints. These findings position SYCL as a promising unified programming model for heterogeneous computing environments, particularly for bioinformatic applications. Manuel Costanzo, Enzo Rucci, Carlos García 0001, Marcelo R. Naiouf, Manuel Prieto 0001 |
Future Gener. Comput. Syst. | 3 |
| 2024 | Assessing opportunities of SYCL for biological sequence alignment on GPU-based systemsabstractAbstract Bioinformatics and computational biology are two fields that have been exploiting GPUs for more than two decades, with being CUDA the most used programming language for them. However, as CUDA is an NVIDIA proprietary language, it implies a strong portability restriction to a wide range of heterogeneous architectures, like AMD or Intel GPUs. To face this issue, the Khronos group has recently proposed the SYCL standard, which is an open, royalty-free, cross-platform abstraction layer that enables the programming of a heterogeneous system to be written using standard, single-source C++ code. Over the past few years, several implementations of this SYCL standard have emerged, being oneAPI the one from Intel. This paper presents the migration process of theSW# suite, a biological sequence alignment tool developed in CUDA, to SYCL using Intel’s oneAPI ecosystem. The experimental results show thatSW# was completely migrated with a small programmer intervention in terms of hand-coding. In addition, it was possible to port the migrated code between different architectures (considering multiple vendor GPUs and also CPUs), with no noticeable performance degradation on five different NVIDIA GPUs. Moreover, performance remained stable when switching to another SYCL implementation. As a consequence, SYCL and its implementations can offer attractive opportunities for the bioinformatics community, especially considering the vast existence of CUDA-based legacy codes. Manuel Costanzo, Enzo Rucci, Carlos García 0001, Marcelo R. Naiouf, Manuel Prieto 0001 |
J. Supercomput. | 3 |
| 2024 | SYCL in the edge: performance and energy evaluation for heterogeneous accelerationabstractAbstract Edge computing is essential to handle increasing data volumes and processing capacities. It provides real-time and secure data processing near data sources, like smart devices, alleviating cloud computing energy use, and saving network bandwidth. Specialized accelerators, like GPUs and FPGAs, are vital for low-latency edge computing but the requirements to customized code for different hardware and vendors suppose important compatibility issues. This paper evaluates the potential of SYCL in addressing code portability issues encountered in edge computing. We employed the Polybench suite to compare various SYCL implementations, specifically DPC++ and AdaptiveCpp, with the native solution, CUDA. The disparity between SYCL implementations was negligible, at just 5%. Furthermore, we evaluated SYCL in the context of specific edge computing applications such as video processing using three different optical flow algorithms. The results revealed a slight performance gap of 3% when transitioning from CUDA to SYCL. Upon evaluating energy consumption, the observed difference ranged from $$\pm 10\%$$ ± 10 % , depending on the application utilized. These gaps are the price one may need to pay when achieving the ability to successfully run the same code on two distinct edge boards. These findings underscore SYCL’s capacity to increase productivity in terms of development costs and facilitate IoT deployment without being locked into a particular platform or manufacturer. Youssef Faqir-Rhazoui, Carlos García 0001 |
J. Supercomput. | 2 |
| 2023 | Comparing Performance and Portability Between CUDA and SYCL for Protein Database Search on NVIDIA, AMD, and Intel GPUsabstractThe heterogeneous computing paradigm has led to the need for portable and efficient programming solutions that can leverage the capabilities of various hardware devices, such as NVIDIA, Intel, and AMD GPUs. This study evaluates the portability and performance of the SYCL and CUDA languages for one fundamental bioinformatics application (Smith-Waterman protein database search) across different GPU architectures, considering single and multi-GPU configurations from different vendors. The experimental work showed that, while both CUDA and SYCL versions achieve similar performance on NVIDIA devices, the latter demonstrated remarkable code portability to other GPU architectures, such as AMD and Intel. Furthermore, the architectural efficiency rates achieved on these devices were superior in 3 of the 4 cases tested. This brief study highlights the potential of SYCL as a viable solution for achieving both performance and portability in the heterogeneous computing ecosystem. Manuel Costanzo, Enzo Rucci, Carlos García 0001, Marcelo R. Naiouf, Manuel Prieto 0001 |
SBAC-PAD | 3 |
| 2023 | Exploring the performance and portability of the k-means algorithm on SYCL across CPU and GPU architecturesabstractAbstract The aim of SYCL is to reduce the gap between the performance and code portability of the main accelerators used in HPC, such as multi-vendor CPUs, GPUs, and FPGAs. To evaluate SYCL’s performance portability, this paper uses the k-means algorithm as a case study. The k-means algorithm is simple to code but can be complex to optimize. In this research, we compare our developed SYCL version with the most efficient implementations of CUDA and OpenMP. Our resulting SYCL code can potentially run on multi-vendor CPUs and GPUs. Additionally, we have created a hand-tuned SYCL variation that is optimized for specific device architectures (CPU, NVIDIA GPU, and Intel GPU) to evaluate the performance difference between a standard version and an optimized one. The results show that SYCL outperforms Intel GPUs and CPUs compared to the state-of-the-art He-Vialle version, while on NVIDIA GPUs SYCL offers equivalent performance compared to its native CUDA implementation. Youssef Faqir-Rhazoui, Carlos García 0001 |
J. Supercomput. | 2 |
| 2022 | Portability and Performance Assessment of the Non-Negative Matrix Factorization Algorithm with OpenMP and SYCLabstractThe SYCL standard was released to improve code portability across heterogeneous environments. Intel released the oneAPI toolkit, which includes the Data-Parallel C++ (DPC++) compiler which is the Intel’s SYCL implementation. SYCL is designed to use a single source code to target multiple accelerators such as: multi-core CPUs, GPUs and even FPGAs. Additionally, the C/C++ compiler provided in the oneAPI toolkit supports OpenMP which also allows targeting codes on both CPU and GPU devices. In this paper, the performance of SYCL and OpenMP is evaluated using the well-known non-negative matrix factorization (NMF) algorithm. Three different NMF implementations are developed: baseline, SYCL and OpenMP versions to analyze the acceleration on CPU and GPU. Experimental results show that while the two programming models perform almost identically on CPU, on GPU, SYCL outperforms its OpenMP counterpart slightly. Youssef Faqir-Rhazoui, Carlos García 0001, Francisco Tirado |
CLEI | 2 |
| 2022 | Evaluation of Intel's DPC++ Compatibility Tool in heterogeneous computingabstractThe Intel DPC++ Compatibility Tool is a component of the Intel oneAPI Base Toolkit. This tool automatically transforms CUDA code into Data Parallel C++ (DPC++), thus assisting in the migration process. DPC++ is an implementation of the programming standard for heterogeneous computing known as SYCL, which unifies the development of parallel applications on CPUs, GPUs or even FPGAs. This paper analyzes the DPC++ Compatibility Tool by considering the manual intervention required and the problems encountered while migrating the Rodinia benchmarks. For this suite, this tool achieves an impressive rate of almost 87% for code successfully migrated. Moreover, a comparative study of the performance obtained by the migrated code was carried out, showing a moderate overhead in most of the migrated examples. Finally, a performance comparison on different devices was also performed. Germán Castaño, Youssef Faqir-Rhazoui, Carlos García 0001, Manuel Prieto 0001 |
J. Parallel Distributed Comput. | 3 |
| 2020 | HEVC optimization based on human perception for real-time environments
David Guillermo Fernández, Guillermo Botella Juan, Alberto A. Del Barrio, Carlos García 0001, Manuel Prieto 0001, Christos Grecos |
Multim. Tools Appl. | 4 |
| 2019 | Open Multi-Processing Acceleration for Unsupervised Land Cover Categorization Using Probabilistic Latent Semantic AnalysisabstractThe probabilistic Latent Semantic Analysis (pLSA) model has recently shown a great potential to uncover highly descriptive semantic features from limited amounts of remote sensing data. Nonetheless, the high computational cost of this algorithm often constraints its operational application for land cover categorization tasks. In this scenario, this paper presents an Open Multi-Processing (OpenMP) implementation of the pLSA algorithm for unsupervised Synthetic Aperture Radar (SAR) and Multi-Spectral Imaging (MSI) image categorization. The experimental results suggest that multi-core systems are an important architecture for the efficient processing of both SAR and MSI datasets. Specifically, the proposed approach is able to cover a real scenario exhibiting good results in both accuracy and performance terms. Sergio Bernabé, Carlos García 0001, Rubén Fernández-Beltran, Mercedes Eugenia Paoletti, Juan Mario Haut, Javier Plaza, Antonio Plaza |
IGARSS | 2 |
| 2019 | Portability Study of an OpenCL Algorithm for Automatic Target Detection in Hyperspectral ImagesabstractIn the last decades, the problem of target detection has received considerable attention in remote sensing applications. When this problem is tackled using hyperspectral images with hundreds of bands, the use of high-performance computing (HPC) is essential. One of the most popular algorithms in the hyperspectral image analysis community for this purpose is the automatic target detection and classification algorithm (ATDCA). Previous research has already investigated the mapping of ATDCA on HPC platforms such as multicore processors, graphics processing units (GPUs), and field-programmable gate arrays (FPGAs), showing impressive speedup factors (after careful fine-tuning) that allow for its exploitation in time-critical scenarios. However, the lack of standardization resulted in most implementations being too specific to a given architecture, eliminating (or at least making extremely difficult) code reusability across different platforms. In order to address this issue, we present a portability study of an implementation of ATDCA developed using the open computing language (OpenCL). We focus on cross-platform parameters such as performance, energy consumption, and code design complexity, as compared to previously developed (hand-tuned) implementations. Our portability study analyzes different strategies to expose data parallelism as well as enable the efficient exploitation of complex memory hierarchies in heterogeneous devices. We also conduct an assessment of energy consumption and discuss metrics to analyze the quality of our code. The conducted experiments-using synthetic and real hyperspectral data sets collected by the Hyperspectral Digital Imagery Collection Experiment (HYDICE) and NASA's Airborne Visible Infra-Red Imaging Spectrometer (AVIRIS)-demonstrate, for the first time in the literature, that portability across different HPC platforms can be achieved for real-time target detection in hyperspectral missions. Sergio Bernabé, Carlos García 0001, Francisco D. Igual, Guillermo Botella Juan, Manuel Prieto 0001, Antonio Plaza |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2018 | Multicore Real-Time Implementation of a Full Hyperspectral Unmixing ChainabstractSolving the mixture problem in remotely sensed hyperspectral images remains a challenging task. In particular, solutions are needed in order to obtain a response for applications with real-time constraints. In the last decade, several efforts have been developed, many of them using graphics processing units (GPUs) and focused on the exploitation of spectral information alone. However, a few spectral unmixing chains have been developed using other architectures such as multicore processors, field programmable gate arrays, or Intel Xeon Phi coprocessors. In this letter, we develop a new parallel unmixing chain for multicore processors. Compared with other approaches, the proposed spatial-spectral alternative takes advantage of the complementary information provided by the spatial correlation of the pixels in the image in addition to the spectral information. Our implementation has been optimized using the application program interface OpenMP and the Intel Math Kernel Library on two multicore architectures, and using real analysis scenarios. The results reveal competitive real-time performance compared with another compute unified device architecture implementation previously developed for GPUs. Sergio Bernabé, Luis Ignacio Jiménez Gil, Carlos García 0001, Javier Plaza, Antonio Plaza |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2018 | Fast and effective CU size decision based on spatial and temporal homogeneity detection
David Guillermo Fernández, Alberto A. Del Barrio, Guillermo Botella Juan, Carlos García 0001 |
Multim. Tools Appl. | 4 |
| 2017 | Embedded Grammars for Grammatical Evolution on GPGPU
J. Ignacio Hidalgo, Carlos Cervigón, José Manuel Velasco, José Manuel Colmenar, Carlos García 0001, Guillermo Botella Juan |
EvoApplications (1) | 5 |
| 2017 | First Experiences Accelerating Smith-Waterman on Intel's Knights Landing Processor
Enzo Rucci, Carlos García 0001, Guillermo Botella Juan, Armando De Giusti, Marcelo R. Naiouf, Manuel Prieto 0001 |
ICA3PP | 2 |
| 2016 | GPU Implementation of Spatial-Spectral Preprocessing for Hyperspectral UnmixingabstractSpectral unmixing pursues the identification of spectrally pure constituents, called endmembers, and their corresponding abundances in each pixel of a hyperspectral image. Most unmixing techniques have focused on the exploitation of spectral information alone. Recently, some techniques have been developed to take advantage of the complementary information provided by the spatial correlation of the pixels in the image. Computational complexity represents a major problem in these spatial-spectral techniques, as hyperspectral images contain very rich information in both the spatial and spectral domains. In this letter, we develop a computationally efficient implementation of a spatial-spectral processing algorithm that has been successfully applied prior to the spectral unmixing of the hyperspectral data. Our implementation has been optimized for the commodity graphics processing units (GPUs) and is evaluated (using both synthetic and real data) using different GPU architectures. Significant speedups can be achieved when processing hyperspectral images of different sizes. This allows for the inclusion of the proposed parallel preprocessing module in a full hyperspectral unmixing chain able to operate in real time. Luis Ignacio Jiménez Gil, Gabriel Martín, Sergio Sánchez, Carlos García 0001, Sergio Bernabé, Javier Plaza, Antonio Plaza |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2015 | Early Experiences with OpenCL on FPGAs: Convolution Case StudyabstractMany proprietary standards and tools have been designed in order to cover a closed set of architectures, and OpenCL has become a free standard for parallel programming on heterogeneous systems, which include custom devices, CPUs, GPUs, FPGAs. This work evaluates the use of the well-known convolution operator in signal processing disciplines focused on FPGA evaluation under different optimizations with respect to thread and memory level exploitation. Carlos Rodriguez-Donate, Guillermo Botella Juan, Carlos García 0001, Eduardo Cabal-Yepez, Manuel Prieto 0001 |
FCCM | 3 |
| 2015 | NMF-mGPU: non-negative matrix factorization on multi-GPU systemsabstractBACKGROUND: In the last few years, the Non-negative Matrix Factorization ( NMF ) technique has gained a great interest among the Bioinformatics community, since it is able to extract interpretable parts from high-dimensional datasets. However, the computing time required to process large data matrices may become impractical, even for a parallel application running on a multiprocessors cluster. In this paper, we present NMF-mGPU, an efficient and easy-to-use implementation of the NMF algorithm that takes advantage of the high computing performance delivered by Graphics-Processing Units ( GPUs ). Driven by the ever-growing demands from the video-games industry, graphics cards usually provided in PCs and laptops have evolved from simple graphics-drawing platforms into high-performance programmable systems that can be used as coprocessors for linear-algebra operations. However, these devices may have a limited amount of on-board memory, which is not considered by other NMF implementations on GPU. RESULTS: NMF-mGPU is based on CUDA ( Compute Unified Device Architecture ), the NVIDIA's framework for GPU computing. On devices with low memory available, large input matrices are blockwise transferred from the system's main memory to the GPU's memory, and processed accordingly. In addition, NMF-mGPU has been explicitly optimized for the different CUDA architectures. Finally, platforms with multiple GPUs can be synchronized through MPI ( Message Passing Interface ). In a four-GPU system, this implementation is about 120 times faster than a single conventional processor, and more than four times faster than a single GPU device (i.e., a super-linear speedup). CONCLUSIONS: Applications of GPUs in Bioinformatics are getting more and more attention due to their outstanding performance when compared to traditional processors. In addition, their relatively low price represents a highly cost-effective alternative to conventional clusters. In life sciences, this results in an excellent opportunity to facilitate the daily work of bioinformaticians that are trying to extract biological meaning out of hundreds of gigabytes of experimental information. NMF-mGPU can be used "out of the box" by researchers with little or no expertise in GPU programming in a variety of platforms, such as PCs, laptops, or high-end GPU clusters. NMF-mGPU is freely available at https://github.com/bioinfo-cnb/bionmf-gpu . Edgardo Mejía-Roa, Daniel Tabas-Madrid, Javier Setoain, Carlos García 0001, Francisco Tirado, Alberto D. Pascual-Montano |
BMC Bioinform. | 4 |
| 2015 | An energy-aware performance analysis of SWIMM: Smith-Waterman implementation on Intel's Multicore and Manycore architecturesabstractSummary Alignment is essential in many areas such as biological, chemical and criminal forensics. The well‐known Smith–Waterman (SW) algorithm is able to retrieve the optimal local alignment with quadratic time and space complexity. There are several implementations that take advantage of computing parallelization, such as manycores, FPGAs or GPUs, in order to reduce the alignment effort. In this research, we adapt, develop and tune the SW algorithm named SWIMM on a heterogeneous platform based on Intel's Xeon and Xeon Phi coprocessor. SWIMM is a free tool available in a public git repository https://github.com/enzorucci/SWIMM . We efficiently exploit data and thread‐level parallelism, reaching up to 380 GCUPS on heterogeneous architecture, 350 GCUPS for the isolated Xeon and 50 GCUPS on Xeon Phi. Despite the heterogeneous implementation obtaining the best performance, it is also the most energy‐demanding. In fact, we also present a trade‐off analysis between performance and power consumption. The greenest configuration is based on an isolated multicore system that exploits AVX2 instruction set architecture reaching 1.5 GCUPS/Watts. Copyright © 2015 John Wiley & Sons, Ltd. Enzo Rucci, Carlos García 0001, Guillermo Botella Juan, Armando De Giusti, Marcelo R. Naiouf, Manuel Prieto 0001 |
Concurr. Comput. Pract. Exp. | 2 |
| 2014 | Smith-Waterman algorithm on heterogeneous systems: A case studyabstractThe well-known Smith-Waterman (SW) algorithm is a high-sensitivity method for local alignments. However, SW is expensive in terms of both execution time and memory usage, which makes it impractical in many applications. Some heuristics are possible but at the expense of losing sensitivity. Fortunately, previous research have shown that new computing platforms such as GPUs and FPGAs are able to accelerate SW and achieve impressive speedups. In this paper we have explored SW acceleration on a heterogeneous platform equipped with an Intel Xeon Phi coprocessor. Our evaluation, using the well-known Swiss-Prot database as a benchmark, has shown that a hybrid CPU-Phi heterogeneous system is able to achieve competitive performance (62.6 GCUPS), even with moderate low-level optimisations. Enzo Rucci, Armando De Giusti, Marcelo R. Naiouf, Guillermo Botella Juan, Carlos García 0001, Manuel Prieto 0001 |
CLUSTER | 5 |
| 2013 | Non-negative matrix factorization on low-power architectures: a comparative studyabstractPower consumption is emerging as one of the main concerns in the High Performance Computing (HPC) field. Many bioinformatics applications require HPC techniques and parallel architectures to meet performance requirements, but at the same time they can be severely limited by energy consumption restrictions. In this paper, we perform an empirical study of an optimized implementation of the Nonnegative Matrix Factorization (NMF), that is widely used in many fields of bioinformatics. We target different types of architectures, including general-purpose, low-power embedded processors and specific-purpose architectures like graphics processors and digital signal processors. From our study, we gain insights in both performance and energy consumption for each one of them under given experimental conditions, and conclude that the most appropriate architecture is usually a trade-off between performance and power consumption for a given experiment and dataset. Carlos García 0001, Francisco D. Igual, Guillermo Botella Juan, Manuel Prieto 0001, Francisco Tirado |
EuroMPI | 1 |
| 2013 | GPU-based acceleration of bio-inspired motion estimation modelabstractSUMMARY In this paper, we describe the specific and efficient implementation of a gradient‐based optical flow model. This scheme was particularized using a validated neuromorphic motion estimation system for the robust extraction of image velocity. This model contains many characteristics that enhanced the capability when compared with other optical flow gradient family algorithms. Our implementation was performed using specific graphic processing units designed in an ad hoc framework for this model, which could be reused in several low‐level machine‐vision approaches. Observed performance results indicate that these accelerators be highly recommended. Furthermore, the throughput obtained in comparison with a general CPU was analyzed for the accurateness of a system built with regard to other optical flow systems. Additionally, several visual examples, commonly used for testing motion estimation sequences, were shown to reveal implementation behavior features. Copyright © 2012 John Wiley & Sons, Ltd. Fermin Ayuso, Guillermo Botella Juan, Carlos García 0001, Manuel Prieto 0001, Francisco Tirado |
Concurr. Comput. Pract. Exp. | 3 |
| 2012 | OpenIRS-UCM: an open-source multi-platform for interactive response systemsabstractInteractive Response Systems (IRS) have been gaining acceptance within the educational community in recent years and a clear proof is the growing number of commercial systems available today in the market. However, most solutions are based on systems which are closed, rigid and dependent on proprietary keypad or platform. We have developed OpenIRS-UCM, a free teaching tool for interactive polling that solves these drawbacks. It is an open source software so it allows the development of new functions by anybody. It has a friendly interface that anyone without high computer skills can use. It enables the coexistence of several commercial clickers simultaneously with smart-phones, tablets or other modern electronic devices. It is developed in Java, thus its use is not restricted to systems based on Microsoft Windows and it is independent of any proprietary software. Carlos García 0001, Fernando Castro, José Ignacio Gómez, Christian Tenllado, Daniel Chaver, José Antonio López Orozco |
ITiCSE | 1 |
| 2011 | Biclustering and classification analysis in gene expression using Nonnegative Matrix Factorization on multi-GPU systemsabstractA great interest has been given to the Nonnegative Matrix Factorization (NMF) technique due to its ability of extracting highly-interpretable parts from data sets. Gene expression analysis is one of the most popular applications of NMF in Bioinformatics. Nonetheless, its usage is hindered by the computational complexity when processing large data sets. In this paper, we present two parallel implementations of NMF. The first version uses CUDA on a Graphics Processing Unit (GPU). Large input matrices are iteratively blockwise transferred and processed. The second implementation distributes data among multiple GPUs synchronized through MPI (Message Passing Interface). When analyzing large data sets with two and four GPUs, it performs respectively, 2.3 and 4.13 times faster than the single-GPU version. This represents about 120 times faster than a conventional CPU. These super linear speedups are achieved when data portions assigned to each GPU are small enough to be transferred only once. Edgardo Mejía-Roa, Carlos García 0001, José Ignacio Gómez, Manuel Prieto 0001, Francisco Tirado, Rubén Nogales, Alberto D. Pascual-Montano |
ISDA | 2 |
| 2010 | Building efficient multi-threaded search nodesabstractSearch nodes are single-purpose components of large Web search engines and their efficient implementation is critical to sustain thousands of queries per second and guarantee individual query response times within a fraction of a second. Current technology trends indicate that search nodes ought to be implemented as multi-threaded multi-core systems. The straightforward solution that system designers can apply in this case is simply to follow standard practice by deploying one asynchronous thread per active query in the node and attaching each thread to a different core. Each concurrent thread is responsible for sequentially processing a single query at a time. The only potential source of read/write conflicts among threads are the accesses to the different application caches present in the search node. However, new Web applications pose much more demanding requirements in terms of read/write conflicts than recent past applications since now data updates must take place concurrently with query processing. Insisting on the same paradigm of concurrent threads now augmented with a transaction concurrency control protocol is a feasible solution. In this paper we propose a more efficient and much simpler solution which has the additional advantage of enabling a very efficient administration of application caches. We propose performing relaxed bulk-synchronous parallelism at multi-core level. Carolina Bonacic, Carlos García 0001, Mauricio Marín, Manuel Prieto 0001, Francisco Tirado |
CIKM | 2 |
| 2008 | Exploiting Hybrid Parallelism in Web Search Engines
Carolina Bonacic, Carlos García 0001, Mauricio Marín, Manuel Prieto 0001, Francisco Tirado |
Euro-Par | 2 |