EDBT 2026 Demo / reviewers in the wild / expert
Manuel Prieto 0001
dblp:p/ManuelPrietoMatias · also Manuel Prieto-Matías
· DBLP profile ↗
71ranked-venue papers
6as first author
11since 2021 · last 2025
0000-0003-0687-3737ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 54 · 6 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 13Artificial intelligence and machine learning · 4Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Analyzing the performance portability of SYCL across CPUs, GPUs, and hybrid systems with SW sequence alignmentabstractThe high-performance computing (HPC) landscape is undergoing rapid transformation, with an increasing emphasis on energy-efficient and heterogeneous computing environments. This comprehensive study extends our previous research on SYCL’s performance portability by evaluating its effectiveness across a broader spectrum of computing architectures , including CPUs, GPUs , and hybrid CPU–GPU configurations from NVIDIA, Intel, and AMD. Our analysis covers single-GPU, multi-GPU, single-CPU, and CPU–GPU hybrid setups, using two common, bioinformatic applications as a case study . The results demonstrate SYCL’s versatility across different architectures, maintaining comparable performance to CUDA on NVIDIA GPUs while achieving similar architectural efficiency rates on AMD and Intel GPUs in the majority of cases tested. SYCL also demonstrated remarkable versatility and effectiveness across CPUs from various manufacturers, including the latest hybrid architectures from Intel. Although SYCL showed excellent functional portability in hybrid CPU–GPU configurations, performance varied significantly based on specific hardware combinations. Some performance limitations were identified in multi-GPU and CPU–GPU configurations, primarily attributed to workload distribution strategies rather than SYCL-specific constraints. These findings position SYCL as a promising unified programming model for heterogeneous computing environments, particularly for bioinformatic applications. Manuel Costanzo, Enzo Rucci, Carlos García 0001, Marcelo R. Naiouf, Manuel Prieto 0001 |
Future Gener. Comput. Syst. | 5 |
| 2024 | Exploiting Elasticity via OS-Runtime Cooperation to Improve CPU Utilization in Multicore SystemsabstractThe chip multicore processor (CMP) architecture has become the predominant design choice for contemporary general-purpose systems across multiple sectors of commercial technology. Thanks to technological progress, CMP systems can now feature hundreds of cores. While multithreaded applications may potentially benefit from the increasing core counts, leveraging all available cores is not always feasible due to limited Thread-Level Parallelism (TLP), load imbalance among threads, and other scalability bottlenecks. Colocating multiple applications on the same node is becoming a popular practice to maximize processor utilization. In HPC, malleability -the ability to dynamically alter the number of active threads within the same application-, is also being exploited at the runtime-system level to better deal with scenarios exhibiting time-varying scalability. In the cloud, application colocation is leveraged along with different forms of coarse-grained elasticity to cater to the varying resource demands. This work introduces an operating system (OS) level elastic mechanism designed to efficiently leverage idle CPU periods in workloads consisting of unmodified applications, many of which do not rely on a runtime system to function. This mechanism constitutes a form of fine-grained vertical elasticity that leverages cooperation between the runtime sys-tem and the OS to maximize CPU utilization. To this end, it opportunistically increases the active thread count of mal-leable applications during idle periods. We implemented our proposed OS extensions in the Linux kernel, and augmented the GNU's OpenMP runtime to show a proof of concept of the required OS-runtime interaction. By using diverse multi- threaded programs, we demonstrate the ability of the proposed OS support to substantially improve the system throughput. Javier Rubio, Carlos Bilbao, Juan Carlos Saez, Manuel Prieto 0001 |
PDP | 4 |
| 2024 | Big-PERCIVAL: Exploring the Native Use of 64-Bit Posit Arithmetic in Scientific ComputingabstractThe accuracy requirements in many scientific computing workloads result in the use of double-precision floating-point arithmetic in the execution kernels. Nevertheless, emerging real-number representations, such as posit arithmetic, show promise in delivering even higher accuracy in such computations. In this work, we explore the native use of 64-bit posits in a series of numerical benchmarks and compare their timing performance, accuracy and hardware cost to IEEE 754 doubles. In addition, we also study the conjugate gradient method for numerically solving systems of linear equations in real-world applications. For this, we extend the PERCIVAL RISC-V core and the Xposit custom RISC-V extension with posit64 and quire operations. Results show that posit64 can obtain up to 4 orders of magnitude lower mean square error than doubles. This leads to a reduction in the number of iterations required for convergence in some iterative solvers. However, leveraging the quire accumulator register can limit the order of some operations such as matrix multiplications. Furthermore, detailed FPGA and ASIC synthesis results highlight the significant hardware cost of 64-bit posit arithmetic and quire. Despite this, the large accuracy improvements achieved with the same memory bandwidth suggest that posit arithmetic may provide a potential alternative representation for scientific computing. David Mallasén, Alberto A. Del Barrio, Manuel Prieto 0001 |
IEEE Trans. Computers | 3 |
| 2024 | Assessing opportunities of SYCL for biological sequence alignment on GPU-based systemsabstractAbstract Bioinformatics and computational biology are two fields that have been exploiting GPUs for more than two decades, with being CUDA the most used programming language for them. However, as CUDA is an NVIDIA proprietary language, it implies a strong portability restriction to a wide range of heterogeneous architectures, like AMD or Intel GPUs. To face this issue, the Khronos group has recently proposed the SYCL standard, which is an open, royalty-free, cross-platform abstraction layer that enables the programming of a heterogeneous system to be written using standard, single-source C++ code. Over the past few years, several implementations of this SYCL standard have emerged, being oneAPI the one from Intel. This paper presents the migration process of theSW# suite, a biological sequence alignment tool developed in CUDA, to SYCL using Intel’s oneAPI ecosystem. The experimental results show thatSW# was completely migrated with a small programmer intervention in terms of hand-coding. In addition, it was possible to port the migrated code between different architectures (considering multiple vendor GPUs and also CPUs), with no noticeable performance degradation on five different NVIDIA GPUs. Moreover, performance remained stable when switching to another SYCL implementation. As a consequence, SYCL and its implementations can offer attractive opportunities for the bioinformatics community, especially considering the vast existence of CUDA-based legacy codes. Manuel Costanzo, Enzo Rucci, Carlos García 0001, Marcelo R. Naiouf, Manuel Prieto 0001 |
J. Supercomput. | 5 |
| 2023 | PERCIVAL: Deploying Posits and Quire Arithmetic into the CVA6 RISC-V CoreabstractRepresenting and operating on real numbers in a microprocessor presents unique challenges not encountered with the set of integers. Working with real numbers introduces additional concepts such as precision, that is, the error made between the number with which we want to operate and the approximation that we can represent in a finite number of bits. Currently, the universally extended way of representing the set of real numbers is using floating-point numbers defined by the IEEE 754 standard. This format presents a series of difficulties, such as the different rounding schemes, reproducibility problems depending on the implementation, a multitude of ways to represent Not a Numbers (NaNs) or the existence of plus and minus zero. David Mallasén, Raul Murillo 0001, Alberto A. Del Barrio, Guillermo Botella Juan, Luis Piñuel, Manuel Prieto 0001 |
CF | 6 |
| 2023 | Comparing Performance and Portability Between CUDA and SYCL for Protein Database Search on NVIDIA, AMD, and Intel GPUsabstractThe heterogeneous computing paradigm has led to the need for portable and efficient programming solutions that can leverage the capabilities of various hardware devices, such as NVIDIA, Intel, and AMD GPUs. This study evaluates the portability and performance of the SYCL and CUDA languages for one fundamental bioinformatics application (Smith-Waterman protein database search) across different GPU architectures, considering single and multi-GPU configurations from different vendors. The experimental work showed that, while both CUDA and SYCL versions achieve similar performance on NVIDIA devices, the latter demonstrated remarkable code portability to other GPU architectures, such as AMD and Intel. Furthermore, the architectural efficiency rates achieved on these devices were superior in 3 of the 4 cases tested. This brief study highlights the potential of SYCL as a viable solution for achieving both performance and portability in the heterogeneous computing ecosystem. Manuel Costanzo, Enzo Rucci, Carlos García 0001, Marcelo R. Naiouf, Manuel Prieto 0001 |
SBAC-PAD | 5 |
| 2023 | Flexible system software scheduling for asymmetric multicore systems with PMCSched: A case for Intel Alder LakeabstractSummary Asymmetric multicore processors (AMPs) couple high‐performance big cores and power‐efficient small ones, all exposing a shared instruction set architecture to software, but with different microarchitectural features. The energy efficiency benefits of AMPs, together with the general‐purpose nature of the various cores, have led hardware manufacturers to build commercial AMP‐based products, first for the mobile and embedded domains, and more recently, for the desktop market segment, as with the Intel Alder Lake processor family. This trend indicates that AMPs may become a solid and more energy efficient replacement for symmetric multicores in a wide range of application domains. Previous research has demonstrated that the system software can substantially improve scheduling—critical to get the most out of heterogeneous cores—by leveraging hardware facilities that are directly managed by the OS, such as performance monitoring counters, or the recently introduced Intel Thread Director technology. Unfortunately, the OS‐level support enabling access to these hardware facilities may often take a long time to be adopted in operating systems, or may come in forms that make its utilization challenging from specific levels of the system software stack, especially in production systems. To fill this gap, we propose PMCSched, an open‐source framework enabling rapid development and evaluation of custom scheduling‐related support in the Linux kernel. PMCSched greatly simplifies the design and implementation of a wide range of scheduling policies for multicore systems that operate at different system software layers without requiring to patch the kernel. To demonstrate the potential of our framework, we conduct a set of experimental case studies on asymmetry‐aware scheduling for Intel Alder Lake processors. Carlos Bilbao, Juan Carlos Saez, Manuel Prieto 0001 |
Concurr. Comput. Pract. Exp. | 3 |
| 2023 | Divide&Content: A Fair OS-Level Resource Manager for Contention Balancing on NUMA MulticoresabstractChip multicore processors (CMPs) constitute the cherry-picked architecture for high-performance servers employed in supercomputers and cloud datacenters. In the last few years, Non-Uniform Memory Access (NUMA) multicore systems have become the dominant choice in these domains. Regardless of the technology advances enabling to pack an increasing number of cores and bigger caches on the same chip, contention for shared resources still represents an important challenge for the system software. Cores in CMPs typically share multiple resources, such as the last-level cache (LLC) or a DRAM controller. The competition for the usage of these resources leads to uneven performance degradation across co-running applications. Previous research has demonstrated that contention effects on CMPs can be mitigated via smart partitioning of the LLC or by distributing threads across groups of cores so as to even out the degree of competition on multiple LLCs or memory nodes. However, most existing resource-management strategies fail to effectively combine both contention-mitigating techniques, thus providing suboptimal results on NUMA multicores. In this paper, we analyze how to best combine these techniques to improve system-wide fairness, and, based on the conclusions of our analysis, propose a fair OS-level NUMA-aware resource manager that leverages dynamic contention-aware thread-to-socket mappings and cache-partitioning. We implemented our resource manager in the Linux kernel and assessed its effectiveness on a real dual-socket system featuring Intel Skylake processors. Our results show that it reduces unfairness by more than 17% on average compared to Linux and a state-of-the-art NUMA-aware resource manager. Carlos Bilbao, Juan Carlos Saez, Manuel Prieto 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2022 | PERCIVAL: Open-Source Posit RISC-V Core With Quire CapabilityabstractPresents the front cover, title page, cover page, or splash screen of the proceedings record. David Mallasén, Raul Murillo 0001, Alberto A. Del Barrio, Guillermo Botella Juan, Luis Piñuel, Manuel Prieto 0001 |
ARITH | 6 |
| 2022 | Evaluation of Intel's DPC++ Compatibility Tool in heterogeneous computingabstractThe Intel DPC++ Compatibility Tool is a component of the Intel oneAPI Base Toolkit. This tool automatically transforms CUDA code into Data Parallel C++ (DPC++), thus assisting in the migration process. DPC++ is an implementation of the programming standard for heterogeneous computing known as SYCL, which unifies the development of parallel applications on CPUs, GPUs or even FPGAs. This paper analyzes the DPC++ Compatibility Tool by considering the manual intervention required and the problems encountered while migrating the Rodinia benchmarks. For this suite, this tool achieves an impressive rate of almost 87% for code successfully migrated. Moreover, a comparative study of the performance obtained by the migrated code was carried out, showing a moderate overhead in most of the migrated examples. Finally, a performance comparison on different devices was also performed. Germán Castaño, Youssef Faqir-Rhazoui, Carlos García 0001, Manuel Prieto 0001 |
J. Parallel Distributed Comput. | 4 |
| 2022 | LFOC+: A Fair OS-Level Cache-Clustering Policy for Commodity Multicore SystemsabstractCommodity multicore systems are increasingly adopting hardware support that enables the system software to partition the last-level cache (LLC). This support makes it possible for the operating system (OS) to mitigate shared-resource contention effects on multicores by assigning different co-running applications to various cache partitions. Cache-clustering strategies have emerged as a way to improve throughput and fairness on platforms with cache-partitioning support. Unlike strict cache-partitioning, which allocates separate cache partitions to each application, cache-clustering allows partitions to be shared by several applications. In this article we propose LFOC+, a fair OS-level cache-clustering policy for commodity multicores. LFOC+ tries to mimic the behavior of the optimal cache-clustering solution for fairness, which we could obtain for different workloads by using a simulation tool. Our strategy continuously gathers data from performance counters to classify applications based on the degree of cache sensitivity and contentiouness, and separates cache-sensitive applications from aggressor programs to improve fairness, while providing acceptable throughput. We implemented LFOC+ in the Linux kernel and evaluated it on a system featuring an Intel Skylake processor, where we compare its effectiveness to that of four state-of-the-art policies. Our analysis reveals that LFOC+ brings a higher reduction in unfairness, and constitutes a lightweight cache-clustering policy. Juan Carlos Saez, Fernando Castro, Graziano Fanizzi, Manuel Prieto 0001 |
IEEE Trans. Computers | 4 |
| 2020 | Enabling performance portability of data-parallel OpenMP applications on asymmetric multicore processorsabstractAsymmetric multicore processors (AMPs) couple high-performance big cores and low-power small cores with the same instruction-set architecture but different features, such as clock frequency or microarchitecture. Previous work has shown that asymmetric designs may deliver higher energy efficiency than symmetric multicores for diverse workloads. Despite their benefits, AMPs pose significant challenges to runtime systems of parallel programming models. While previous work has mainly explored how to efficiently execute task-based parallel applications on AMPs, via enhancements in the runtime system, improving the performance of unmodified data-parallel applications on these architectures is still a big challenge. In this work we analyze the particular case of loop-based OpenMP applications, which are widely used today in scientific and engineering domains, and constitute the dominant application type in many parallel benchmark suites used for performance evaluation on multicore systems. We observed that conventional loop-scheduling OpenMP approaches are unable to efficiently cope with the load imbalance that naturally stems from the different performance delivered by big and small cores. Juan Carlos Saez, Fernando Castro, Manuel Prieto 0001 |
ICPP | 3 |
| 2020 | HEVC optimization based on human perception for real-time environments
David Guillermo Fernández, Guillermo Botella Juan, Alberto A. Del Barrio, Carlos García 0001, Manuel Prieto 0001, Christos Grecos |
Multim. Tools Appl. | 5 |
| 2020 | STEEL-RT: combining single task-single executor model and expanded scheduling to ease heterogeneity exploitation
Antón Rey, Francisco D. Igual, Manuel Prieto 0001 |
J. Supercomput. | 3 |
| 2019 | LFOC: A Lightweight Fairness-Oriented Cache Clustering Policy for Commodity MulticoresabstractMulticore processors constitute the main architecture choice for modern computing systems in different market segments. Despite their benefits, the contention that naturally appears when multiple applications compete for the use of shared resources among cores, such as the last-level cache (LLC), may lead to substantial performance degradation. This may have a negative impact on key system aspects such as throughput and fairness. Assigning the various applications in the workload to separate LLC partitions with possibly different sizes, has been proven effective to mitigate shared-resource contention effects. Adrian Garcia-Garcia, Juan Carlos Saez, Fernando Castro, Manuel Prieto 0001 |
ICPP | 4 |
| 2019 | Portability Study of an OpenCL Algorithm for Automatic Target Detection in Hyperspectral ImagesabstractIn the last decades, the problem of target detection has received considerable attention in remote sensing applications. When this problem is tackled using hyperspectral images with hundreds of bands, the use of high-performance computing (HPC) is essential. One of the most popular algorithms in the hyperspectral image analysis community for this purpose is the automatic target detection and classification algorithm (ATDCA). Previous research has already investigated the mapping of ATDCA on HPC platforms such as multicore processors, graphics processing units (GPUs), and field-programmable gate arrays (FPGAs), showing impressive speedup factors (after careful fine-tuning) that allow for its exploitation in time-critical scenarios. However, the lack of standardization resulted in most implementations being too specific to a given architecture, eliminating (or at least making extremely difficult) code reusability across different platforms. In order to address this issue, we present a portability study of an implementation of ATDCA developed using the open computing language (OpenCL). We focus on cross-platform parameters such as performance, energy consumption, and code design complexity, as compared to previously developed (hand-tuned) implementations. Our portability study analyzes different strategies to expose data parallelism as well as enable the efficient exploitation of complex memory hierarchies in heterogeneous devices. We also conduct an assessment of energy consumption and discuss metrics to analyze the quality of our code. The conducted experiments-using synthetic and real hyperspectral data sets collected by the Hyperspectral Digital Imagery Collection Experiment (HYDICE) and NASA's Airborne Visible Infra-Red Imaging Spectrometer (AVIRIS)-demonstrate, for the first time in the literature, that portability across different HPC platforms can be achieved for real-time target detection in hyperspectral missions. Sergio Bernabé, Carlos García 0001, Francisco D. Igual, Guillermo Botella Juan, Manuel Prieto 0001, Antonio Plaza |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2019 | Variable intra-task threading for power-constrained performance and energy optimization in DAG scheduling
Antón Rey, Francisco D. Igual, Manuel Prieto 0001 |
J. Supercomput. | 3 |
| 2018 | A CPU-GPU Parallel Ant Colony Optimization Solver for the Vehicle Routing Problem
Antón Rey, Manuel Prieto 0001, José Ignacio Gómez, Christian Tenllado, J. Ignacio Hidalgo |
EvoApplications | 2 |
| 2018 | Reuse Detector: Improving the Management of STT-RAM SLLCsabstractVarious constraints of Static Random Access Memory (SRAM) are leading to consider new memory technologies as candidates for building on-chip shared last-level caches (SLLCs). Spin-Transfer Torque RAM (STT-RAM) is currently postulated as the prime contender due to its better energy efficiency, smaller die footprint and higher scalability. However, STT-RAM also exhibits some drawbacks, like slow and energy-hungry write operations that need to be mitigated before it can be used in SLLCs for the next generation of computers. In this work, we address these shortcomings by leveraging a new management mechanism for STT-RAM SLLCs. This approach is based on the previous observation that although the stream of references arriving at the SLLC of a Chip MultiProcessor (CMP) exhibits limited temporal locality, it does exhibit reuse locality, i.e. those blocks referenced several times manifest high probability of forthcoming reuse. As such, conventional STT-RAM SLLC management mechanisms, mainly focused on exploiting temporal locality, result in low efficient behavior. In this paper, we employ a cache management mechanism that selects the contents of the SLLC aimed to exploit reuse locality instead of temporal locality. Specifically, our proposal consists in the inclusion of a Reuse Detector (RD) between private cache levels and the STT-RAM SLLC. Its mission is to detect blocks that do not exhibit reuse, in order to avoid their insertion in the SLLC, hence reducing the number of write operations and the energy consumption in the STT-RAM. Our evaluation, using multiprogrammed workloads in quad-core, eight-core and 16-core systems, reveals that our scheme reports on average, energy reductions in the SLLC in the range of 37–30%, additional energy savings in the main memory in the range of 6–8% and performance improvements of 3% (quad-core), 7% (eight-core) and 14% (16-core) compared with an STT-RAM SLLC baseline where no RD is employed. More importantly, our approach outperforms DASCA, the state-of-the-art STT-RAM SLLC management, reporting—depending on the specific scenario and the kind of applications used—SLLC energy savings in the range of 4–11% higher than those of DASCA, delivering higher performance in the range of 1.5–14% and additional improvements in DRAM energy consumption in the range of 2–9% higher than DASCA. Roberto Rodríguez-Rodríguez, Fernando Castro, Pablo Ibáñez 0001, Daniel Chaver, Víctor Viñals, Juan Carlos Saez, Manuel Prieto 0001, Luis Piñuel, Teresa Monreal Arnal, José María Llabería |
Comput. J. | 8 |
| 2018 | On the Interplay Between Throughput, Fairness and Energy Efficiency on Asymmetric Multicore ProcessorsabstractAsymmetric single-ISA multicore processors (AMPs), which integrate high-performance big cores and low-power small cores, were shown to deliver higher performance per watt than symmetric multicores. Previous work has highlighted that this potential of AMP systems can be realizable by scheduling the various applications in a workload on the most appropriate core type. A number of scheduling schemes have been proposed to accomplish different goals, such as system throughput optimization, enforcing fairness or reducing energy consumption. While the interrelationship between throughput and fairness on AMPs has been comprehensively studied, the impact that optimizing energy efficiency has on the other two aspects is still unclear. To fill this gap, we carry out a comprehensive analytical and experimental study that illustrates the interplay between throughput, fairness and energy efficiency on AMPs. Our analytical study allowed us to define the energy-efficiency factor (EEF) metric, which aids the OS scheduler in identifying which applications are more suitable for running on the various cores to ensure a good balance between performance and energy consumption. We propose two energy-aware OS-level schedulers that leverage the EEF metric; the first one strives to optimize the energy-delay product and the second scheduler can be configured to optimize different metrics on the AMP. To demonstrate the effectiveness of these proposals, we performed an extensive evaluation and comparison with state-of-the-art schemes by using real asymmetric hardware and scheduler implementations in the Linux kernel. Juan Carlos Saez, Adrian Pousa, Armando De Giusti, Manuel Prieto 0001 |
Comput. J. | 4 |
| 2018 | Contention-Aware Fair Scheduling for Asymmetric Single-ISA Multicore SystemsabstractAsymmetric single-ISA multicore processors (AMPs), which integrate high-performance big cores and low-power small cores, were shown to deliver higher performance per watt than symmetric multicores. Previous work has demonstrated that the OS scheduler plays an important role in realizing the potential of AMP systems. While throughput optimization on AMPs has been extensively studied, delivering fairness on these platforms still constitutes an important challenge to the OS. To this end, the scheduler must be equipped with a mechanism enabling to accurately track the progress that each application in the workload makes as it runs on the various core types throughout the execution. In turn, this progress largely depends on the benefit (or speedup) that an application derives on a big core relative to a small one, which may differ greatly across applications. While existing fairness-aware schedulers take application relative speedup into consideration when tracking progress, they do not cater to the performance degradation that may occur naturally due to contention on shared resources among cores, such as the last-level cache or the memory bus. In this paper, we propose CAMPS, a contention-aware fair scheduler for AMPs that primarily targets long-running compute-intensive workloads. Unlike other schemes, CAMPS does not require special hardware extensions or platform-specific speedup-prediction models to function. Our experimental evaluation, which leverages real asymmetric hardware and scheduler implementations in the Linux kernel, demonstrates that CAMPS improves fairness by up to 11 percent with respect to a state-of-the-art fairness-aware OS-level scheme, while delivering better system throughput. Adrian Garcia-Garcia, Juan Carlos Saez, Manuel Prieto 0001 |
IEEE Trans. Computers | 3 |
| 2017 | First Experiences Accelerating Smith-Waterman on Intel's Knights Landing Processor
Enzo Rucci, Carlos García 0001, Guillermo Botella Juan, Armando De Giusti, Marcelo R. Naiouf, Manuel Prieto 0001 |
ICA3PP | 6 |
| 2017 | PMCTrack: Delivering Performance Monitoring Counter Support to the OS SchedulerabstractHardware performance monitoring counters (PMCs) have proven effective in characterizing application performance. Because PMCs can only be accessed directly at the OS privilege level, kernel-level tools must be developed to enable the end-user and userspace programs to access PMCs. A large body of work has demonstrated that the OS can perform effective runtime optimizations in multicore systems by leveraging performance-counter data. Special attention has been paid to optimizations in the OS scheduler. While existing performance monitoring tools greatly simplify the collection of PMC application data from userspace, they do not provide an architecture-agnostic kernel-level mechanism that is capable of exposing high-level PMC metrics to OS components, such as the scheduler. As a result, the implementation of PMC-based OS scheduling schemes is typically tied to specific processor models. To address this shortcoming we present PMCTrack, a novel tool for the Linux kernel that provides a simple architecture-independent mechanism that makes it possible for the OS scheduler to access per-thread PMC data. Despite being an OS-oriented tool, PMCTrack still allows the gathering of monitoring data from userspace, enabling kernel developers to carry out the necessary offline analysis and debugging to assist them during the scheduler design process. In addition, the tool provides both the OS and the user-space PMCTrack components with other insightful metrics available in modern processors and which are not directly exposed as PMCs, such as cache occupancy or energy consumption. This information is also of great value when it comes to analyzing the potential benefits of novel scheduling policies on real systems. In this paper, we analyze different case studies that demonstrate the flexibility, simplicity and powerful features of PMCTrack. Juan Carlos Saez, Adrian Pousa, Roberto Rodríguez-Rodríguez, Fernando Castro, Manuel Prieto 0001 |
Comput. J. | 5 |
| 2017 | Towards completely fair scheduling on asymmetric single-ISA multicore processors
Juan Carlos Saez, Adrian Pousa, Fernando Castro, Daniel Chaver, Manuel Prieto 0001 |
J. Parallel Distributed Comput. | 5 |
| 2017 | Performance-Power Evaluation of an OpenCL Implementation of the Simplex Growing Algorithm for Hyperspectral UnmixingabstractOver the last few years, several new strategies for spectral unmixing of remotely sensed hyperspectral data have been proposed. Many of them have been developed to solve the most time-consuming and relevant step: endmember extraction. However, unmixing algorithms can be computationally very expensive in terms of processing time and energy consumption, a fact that compromises their use in applications under real-time and energy/power constraints. In this letter, we present a new parallel simplex growing algorithm (SGA) for hyperspectral data which exploits the memory hierarchy with operations in single-precision floating point. Those optimizations accelerate the most time-consuming parts of this method using the open computing language (OpenCL) standard. We have evaluated the performance versus energy consumption using the same open standard for parallel programming over a diverse set of heterogeneous platforms. Experiments have been conducted using real hyperspectral images collected by NASA's Airborne Visible Infrared Imaging Spectrometer and a collection of 24 synthetic hyperspectral images simulated with different sizes and number of endmembers (10-30). Considering the power consumption and OpenCL across all the proposed devices, the analysis presented indicates that the SGA can now be executed in computationally efficient fashion, which was not possible before introducing the parallel implementation described in this letter. Sergio Bernabé, Guillermo Botella Juan, Jose M. R. Navarro, Carlos Orueta, Francisco D. Igual, Manuel Prieto 0001, Antonio Plaza |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2016 | HeSP: A Simulation Framework for Solving the Task Scheduling-Partitioning Problem on Heterogeneous Architectures
Antón Rey, Francisco D. Igual, Manuel Prieto 0001 |
Euro-Par | 3 |
| 2016 | Parallel implementation of the simplex growing algorithm for hyperspectral unmixing using OpenCLabstractMany algorithms for spectral unmixing have been proposed in the last years applied on hyperspectral imaging. This process is composed by three stages where the extraction of endmembers is the most consuming step. However, endmember extraction algorithms (EEAs) can be computationally very expensive and its acceleration on parallel architectures is still an interesting and open problem. In this paper, we present a parallel implementation of the simplex growing algorithm for hyperspectral unmixing called P-SGA on different platforms using the OpenCL framework. The proposed implementation exploits the memory hierarchy to accelerate the parts of this method which are more time-consuming. The proposed algorithm is evaluated in terms of both accuracy and computational performance through Monte Carlo simulations using the following architectures: multi-core Xeon CPU, NVidia GeForce GTX 980 GPU and Intel Xeon Phi accelerator. Experiments are conducted using real hyperspectral data set revealing considerable acceleration factors, which satisfies the real-time constraints given by the data acquisition rate. Sergio Bernabé, Guillermo Botella Juan, Jose M. R. Navarro, Carlos Orueta, Manuel Prieto 0001, Antonio Plaza |
IGARSS | 5 |
| 2015 | Early Experiences with OpenCL on FPGAs: Convolution Case StudyabstractMany proprietary standards and tools have been designed in order to cover a closed set of architectures, and OpenCL has become a free standard for parallel programming on heterogeneous systems, which include custom devices, CPUs, GPUs, FPGAs. This work evaluates the use of the well-known convolution operator in signal processing disciplines focused on FPGA evaluation under different optimizations with respect to thread and memory level exploitation. Carlos Rodriguez-Donate, Guillermo Botella Juan, Carlos García 0001, Eduardo Cabal-Yepez, Manuel Prieto 0001 |
FCCM | 5 |
| 2015 | An energy-aware performance analysis of SWIMM: Smith-Waterman implementation on Intel's Multicore and Manycore architecturesabstractSummary Alignment is essential in many areas such as biological, chemical and criminal forensics. The well‐known Smith–Waterman (SW) algorithm is able to retrieve the optimal local alignment with quadratic time and space complexity. There are several implementations that take advantage of computing parallelization, such as manycores, FPGAs or GPUs, in order to reduce the alignment effort. In this research, we adapt, develop and tune the SW algorithm named SWIMM on a heterogeneous platform based on Intel's Xeon and Xeon Phi coprocessor. SWIMM is a free tool available in a public git repository https://github.com/enzorucci/SWIMM . We efficiently exploit data and thread‐level parallelism, reaching up to 380 GCUPS on heterogeneous architecture, 350 GCUPS for the isolated Xeon and 50 GCUPS on Xeon Phi. Despite the heterogeneous implementation obtaining the best performance, it is also the most energy‐demanding. In fact, we also present a trade‐off analysis between performance and power consumption. The greenest configuration is based on an isolated multicore system that exploits AVX2 instruction set architecture reaching 1.5 GCUPS/Watts. Copyright © 2015 John Wiley & Sons, Ltd. Enzo Rucci, Carlos García 0001, Guillermo Botella Juan, Armando De Giusti, Marcelo R. Naiouf, Manuel Prieto 0001 |
Concurr. Comput. Pract. Exp. | 6 |
| 2014 | Smith-Waterman algorithm on heterogeneous systems: A case studyabstractThe well-known Smith-Waterman (SW) algorithm is a high-sensitivity method for local alignments. However, SW is expensive in terms of both execution time and memory usage, which makes it impractical in many applications. Some heuristics are possible but at the expense of losing sensitivity. Fortunately, previous research have shown that new computing platforms such as GPUs and FPGAs are able to accelerate SW and achieve impressive speedups. In this paper we have explored SW acceleration on a heterogeneous platform equipped with an Intel Xeon Phi coprocessor. Our evaluation, using the well-known Swiss-Prot database as a benchmark, has shown that a hybrid CPU-Phi heterogeneous system is able to achieve competitive performance (62.6 GCUPS), even with moderate low-level optimisations. Enzo Rucci, Armando De Giusti, Marcelo R. Naiouf, Guillermo Botella Juan, Carlos García 0001, Manuel Prieto 0001 |
CLUSTER | 6 |
| 2013 | Multi-level Clustering on Metric Spaces Using a Multi-GPU Platform
Ricardo J. Barrientos, José Ignacio Gómez, Christian Tenllado, Manuel Prieto 0001, Pavel Zezula |
Euro-Par | 4 |
| 2013 | Non-negative matrix factorization on low-power architectures: a comparative studyabstractPower consumption is emerging as one of the main concerns in the High Performance Computing (HPC) field. Many bioinformatics applications require HPC techniques and parallel architectures to meet performance requirements, but at the same time they can be severely limited by energy consumption restrictions. In this paper, we perform an empirical study of an optimized implementation of the Nonnegative Matrix Factorization (NMF), that is widely used in many fields of bioinformatics. We target different types of architectures, including general-purpose, low-power embedded processors and specific-purpose architectures like graphics processors and digital signal processors. From our study, we gain insights in both performance and energy consumption for each one of them under given experimental conditions, and conclude that the most appropriate architecture is usually a trade-off between performance and power consumption for a given experiment and dataset. Carlos García 0001, Francisco D. Igual, Guillermo Botella Juan, Manuel Prieto 0001, Francisco Tirado |
EuroMPI | 4 |
| 2013 | Delivering fairness and priority enforcement on asymmetric multicore systems via OS schedulingabstractSymmetric-ISA (instruction set architecture) asymmetric-performance multicore processors (AMPs) were shown to deliver higher performance per watt and area than symmetric CMPs for applications with diverse architectural requirements. So, it is likely that future multicore processors will combine big power-hungry fast cores and small low-power slow ones. In this paper, we propose a novel thread scheduling algorithm that aims to improve the throughput-fairness trade-off on AMP systems. Our experimental evaluation on real hardware and using scheduler implementations on a general-purpose operating system, reveals that our proposal delivers a better throughput-fairness trade-off than previous schedulers for a wide variety of multi-application workloads including single-threaded and multithreaded applications. Juan Carlos Saez, Fernando Castro, Daniel Chaver, Manuel Prieto 0001 |
SIGMETRICS | 4 |
| 2013 | GPU-based acceleration of bio-inspired motion estimation modelabstractSUMMARY In this paper, we describe the specific and efficient implementation of a gradient‐based optical flow model. This scheme was particularized using a validated neuromorphic motion estimation system for the robust extraction of image velocity. This model contains many characteristics that enhanced the capability when compared with other optical flow gradient family algorithms. Our implementation was performed using specific graphic processing units designed in an ad hoc framework for this model, which could be reused in several low‐level machine‐vision approaches. Observed performance results indicate that these accelerators be highly recommended. Furthermore, the throughput obtained in comparison with a general CPU was analyzed for the accurateness of a system built with regard to other optical flow systems. Additionally, several visual examples, commonly used for testing motion estimation sequences, were shown to reveal implementation behavior features. Copyright © 2012 John Wiley & Sons, Ltd. Fermin Ayuso, Guillermo Botella Juan, Carlos García 0001, Manuel Prieto 0001, Francisco Tirado |
Concurr. Comput. Pract. Exp. | 4 |
| 2013 | System-level memory management based on statistical variability compensation for frame-based applicationsabstractProcess variability and dynamic domains increase the uncertainty of embedded systems and force designers to apply pessimistic designs, which become unnecessarily conservative and have a tremendous impact on both performance and energy consumption. In this context, developing uncertainty-aware design methodologies that take both variation at platform and at application level into account becomes a must. These methodologies should mitigate the effects derived from uncertainty, avoiding worst-case assumptions. In this article we propose a comprehensive methodology to tackle two forms of uncertainty: (1) process variation on the memory system, (2) application dynamism. A statistical model has been developed to deal with variability derived from fabrication process, whereas system scenarios are selected to cope with dynamic domains. Both sources of uncertainty are firstly tackled in combination at design time, to be refined later, at setup. As a result, at run time the platform can be successfully adapted to the current application behaviour as well as the current variations. Our simulations show that this methodology provides significant energy savings while still meeting strict timing constraints. Concepción Sanz, José Ignacio Gómez, Christian Tenllado, Manuel Prieto 0001, Francky Catthoor |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2013 | Survey of Energy-Cognizant Scheduling TechniquesabstractExecution time is no longer the only metric by which computational systems are judged. In fact, explicitly sacrificing raw performance in exchange for energy savings is becoming a common trend in environments ranging from large server farms attempting to minimize cooling costs to mobile devices trying to prolong battery life. Hardware designers, well aware of these trends, include capabilities like DVFS (to throttle core frequency) into almost all modern systems. However, hardware capabilities on their own are insufficient and must be paired with other logic to decide if, when, and by how much to apply energy-minimizing techniques while still meeting performance goals. One obvious choice is to place this logic into the OS scheduler. This choice is particularly attractive due to the relative simplicity, low cost, and low risk associated with modifying only the scheduler part of the OS. Herein we survey the vast field of research on energy-cognizant schedulers. We discuss scheduling techniques to perform energy-efficient computation. We further explore how the energy-cognizant scheduler's role has been extended beyond simple energy minimization to also include related issues like the avoidance of negative thermal effects as well as addressing asymmetric multicore architectures. Sergey Zhuravlev, Juan Carlos Saez, Sergey Blagodurov, Alexandra Fedorova, Manuel Prieto 0001 |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2012 | Range Query Processing in a Multi-GPU EnvironmentabstractSimilarity search has been widely studied in the last years, as it can be applied to several fields such as searching by content in multimedia objects, text retrieval or computational biology. These applications usually work on very large databases that are often indexed off-line to enable the acceleration of on-line searches. However, to maintain an acceptable throughput, it is essential to exploit the intrinsic parallelism of the algorithms used for the on-line query solving process, even with indexed databases. Therefore, many strategies have been proposed in the literature to parallelize these algorithms, both on shared and distributed memory multiprocessor systems. Lately, GPUs have also been used to implement brute-force approaches instead of using indexing structures, due to the difficulties introduced by the index in the efficient exploitation of the GPU resources. In this work we propose a Multi-GPU metric-space technique that efficiently exploits index data structures for similarity search in large databases, and show how it outperforms previous OpenMP and GPU brute-force strategies. Furthermore, our analysis covers the effects of the database size and its nature. Ricardo J. Barrientos, José Ignacio Gómez, Christian Tenllado, Manuel Prieto 0001, Mauricio Marín |
ISPA | 4 |
| 2012 | Block Tridiagonal Solvers on Heterogeneous ArchitecturesabstractModern multi-core and many-core systems offer a very impressive cost/performance ratio. In this paper a set of new parallel implementations for the solution of linear systems with block-tridiagonal coefficient matrix on current parallel architectures is proposed and evaluated: one of them on multi-core, others on many-core and finally, a new heterogeneous implementation on both architectures. The results show a speedup higher than 6 on certain parts of the problem, being the heterogeneous implementation the fastest. Pedro Valero-Lara, Alfredo Pinelli, Julien Favier, Manuel Prieto 0001 |
ISPA | 4 |
| 2012 | Leveraging Core Specialization via OS Scheduling to Improve Performance on Asymmetric Multicore SystemsabstractAsymmetric multicore processors (AMPs) consist of cores with the same ISA (instruction-set architecture), but different microarchitectural features, speed, and power consumption. Because cores with more complex features and higher speed typically use more area and consume more energy relative to simpler and slower cores, we must use these cores for running applications that experience significant performance improvements from using those features. Having cores of different types in a single system allows optimizing the performance/energy trade-off. To deliver this potential to unmodified applications, the OS scheduler must map threads to cores in consideration of the properties of both. Our work describes a Comprehensive scheduler for Asymmetric Multicore Processors (CAMP) that addresses shortcomings of previous asymmetry-aware schedulers. First, previous schedulers catered to only one kind of workload properties that are crucial for scheduling on AMPs; either efficiency or thread-level parallelism (TLP), but not both. CAMP overcomes this limitation showing how using both efficiency and TLP in synergy in a single scheduling algorithm can improve performance. Second, most existing schedulers relying on models for estimating how much faster a thread executes on a “fast” vs. “slow” core (i.e., the speedup factor ) were specifically designed for AMP systems where cores differ only in clock frequency. However, more realistic AMP systems include cores that differ more significantly in their features. To demonstrate the effectiveness of CAMP on more realistic scenarios, we augmented the CAMP scheduler with a model that predicts the speedup factor on a real AMP prototype that closely matches future asymmetric systems. Juan Carlos Saez, Alexandra Fedorova, David A. Koufaty, Manuel Prieto 0001 |
ACM Trans. Comput. Syst. | 4 |
| 2011 | kNN Query Processing in Metric Spaces Using GPUs
Ricardo J. Barrientos, José Ignacio Gómez, Christian Tenllado, Manuel Prieto 0001, Mauricio Marín |
Euro-Par (1) | 4 |
| 2011 | Introduction
Pierre Manneback, Gudula Rünger, Manuel Prieto 0001 |
Euro-Par (2) | 4 |
| 2011 | Biclustering and classification analysis in gene expression using Nonnegative Matrix Factorization on multi-GPU systemsabstractA great interest has been given to the Nonnegative Matrix Factorization (NMF) technique due to its ability of extracting highly-interpretable parts from data sets. Gene expression analysis is one of the most popular applications of NMF in Bioinformatics. Nonetheless, its usage is hindered by the computational complexity when processing large data sets. In this paper, we present two parallel implementations of NMF. The first version uses CUDA on a Graphics Processing Unit (GPU). Large input matrices are iteratively blockwise transferred and processed. The second implementation distributes data among multiple GPUs synchronized through MPI (Message Passing Interface). When analyzing large data sets with two and four GPUs, it performs respectively, 2.3 and 4.13 times faster than the single-GPU version. This represents about 120 times faster than a conventional CPU. These super linear speedups are achieved when data portions assigned to each GPU are small enough to be transferred only once. Edgardo Mejía-Roa, Carlos García 0001, José Ignacio Gómez, Manuel Prieto 0001, Francisco Tirado, Rubén Nogales, Alberto D. Pascual-Montano |
ISDA | 4 |
| 2011 | Leveraging workload diversity through OS scheduling to maximize performance on single-ISA heterogeneous multicore systems
Juan Carlos Saez, Daniel Shelepov, Alexandra Fedorova, Manuel Prieto 0001 |
J. Parallel Distributed Comput. | 4 |
| 2010 | Building efficient multi-threaded search nodesabstractSearch nodes are single-purpose components of large Web search engines and their efficient implementation is critical to sustain thousands of queries per second and guarantee individual query response times within a fraction of a second. Current technology trends indicate that search nodes ought to be implemented as multi-threaded multi-core systems. The straightforward solution that system designers can apply in this case is simply to follow standard practice by deploying one asynchronous thread per active query in the node and attaching each thread to a different core. Each concurrent thread is responsible for sequentially processing a single query at a time. The only potential source of read/write conflicts among threads are the accesses to the different application caches present in the search node. However, new Web applications pose much more demanding requirements in terms of read/write conflicts than recent past applications since now data updates must take place concurrently with query processing. Insisting on the same paradigm of concurrent threads now augmented with a transaction concurrency control protocol is a feasible solution. In this paper we propose a more efficient and much simpler solution which has the additional advantage of enabling a very efficient administration of application caches. We propose performing relaxed bulk-synchronous parallelism at multi-core level. Carolina Bonacic, Carlos García 0001, Mauricio Marín, Manuel Prieto 0001, Francisco Tirado |
CIKM | 4 |
| 2010 | A comprehensive scheduler for asymmetric multicore systemsabstractSymmetric-ISA (instruction set architecture) asymmetric-performance multicore processors were shown to deliver higher performance per watt and area for applications with diverse architectural requirements, and so it is likely that future multicore processors will combine a few fast cores characterized by complex pipelines, high clock frequency, high area requirements and power consumption, and many slow ones, characterized by simple pipelines, low clock frequency, low area requirements and power consumption. Asymmetric multicore processors (AMP) derive their efficiency from core specialization. Efficiency specialization ensures that fast cores are used for "CPU-intensive" applications, which efficiently utilize these cores' "expensive" features, while slow cores would be used for "memory-intensive" applications, which utilize fast cores inefficiently. TLP (thread-level parallelism) specialization ensures that fast cores are used to accelerate sequential phases of parallel applications, while leaving slow cores for energy-efficient execution of parallel phases. Specialization is effected by an asymmetry-aware thread scheduler, which maps threads to cores in consideration of the properties of both. Previous asymmetry-aware schedulers employed one type of specialization (either efficiency or TLP), but not both. As a result, they were effective only for limited workload scenarios. We propose, implement, and evaluate CAMP, a Comprehensive AMP scheduler, which delivers both efficiency and TLP specialization. Furthermore, we propose a new light-weight technique for discovering which threads utilize fast cores most efficiently. Our evaluation in the OpenSolaris operating system demonstrates that CAMP accomplishes an efficient use of an AMP system for a variety of workloads, while existing asymmetry-aware schedulers were effective only in limited scenarios. Juan Carlos Saez, Manuel Prieto 0001, Alexandra Fedorova, Sergey Blagodurov |
EuroSys | 2 |
| 2010 | Improving face recognition by combination of natural and Gabor faces
Christian Tenllado, José Ignacio Gómez, Javier Setoain, Darío Mora, Manuel Prieto 0001 |
Pattern Recognit. Lett. | 5 |
| 2009 | System-level process variability compensation on memory organizations: on the scalability of multi-mode memoriesabstractProcess variation and the dynamism of modern applications can degrade the expected performance of a system. Execution time can be severely affected by both factors, resulting in deadline violations and energy consumption overheads. Memory organizations, which account for a large part of the system energy and the time budgets, are especially vulnerable to process variation. Configurable - multimode - memories are a promising technology to deal with these problems, but they also introduce new issues that need to be solved. Essentially, adding configuration capabilities to the memories comes with a cost, both in memory area and control complexity; hence, we need to evaluate what is the minimum amount of re-configurability to satisfy system's constraints. In this paper, we analyze the scalability of configurable memories and highlight the relationship among mode allocation, memory mapping and data allocation. Concepción Sanz, Manuel Prieto 0001, José Ignacio Gómez, Antonis Papanikolaou, Francky Catthoor |
ASP-DAC | 2 |
| 2009 | Introduction
Alex Szalay, Djoerd Hiemstra, Alfons Kemper, Manuel Prieto 0001 |
Euro-Par | 4 |
| 2009 | Endmember Extraction from Hyperspectral Imagery using a Parallel Ensemble Approach with Consensus AnalysisabstractWe have explored in this paper a framework to test in a quantitative manner the stability of different endmember extraction and spectral unmixing algorithms based on the concept of Consensus Clustering. The idea is to investigate if the sensibility of those algorithms to the number of endmembers can be used to estimate this parameter itself. Preliminary results on synthetic data reveal that the proposed scheme, which can be implemented efficiently in parallel, can compete with state-of-the-art schemes. Fermin Ayuso, Javier Setoain, Manuel Prieto 0001, Christian Tenllado, Francisco Tirado, Javier Plaza, Antonio Plaza |
IGARSS (5) | 3 |
| 2009 | Using age registers for a simple load-store queue filtering
Fernando Castro, Daniel Chaver, Luis Piñuel, Manuel Prieto 0001, Francisco Tirado |
J. Syst. Archit. | 4 |
| 2009 | Replacing Associative Load Queues: A Timing-Centric ApproachabstractOne of the main challenges of modern processor design is the implementation of a scalable and efficient mechanism to detect memory access order violations as a result of out-of-order execution. Traditional age-ordered associative load queues are complex, inefficient, and power hungry. In this paper, we introduce two new dependence checking schemes with different design tradeoffs, but both explicitly rely on timing information as a primary instrument to rule out dependence violation. Our timing-centric designs operate at a fraction of the energy cost of an associative LQ and achieve the same functionality with an insignificant performance impact on average. Studies with parallel benchmarks also show that they are equally effective and efficient in a chip-multiprocessor environment. Fernando Castro, Regana Noor, Alok Garg, Daniel Chaver, Michael C. Huang 0001, Luis Piñuel, Manuel Prieto 0001, Francisco Tirado |
IEEE Trans. Computers | 7 |
| 2008 | Exploiting Hybrid Parallelism in Web Search Engines
Carolina Bonacic, Carlos García 0001, Mauricio Marín, Manuel Prieto 0001, Francisco Tirado |
Euro-Par | 4 |
| 2008 | Improving Priority Enforcement via Non-Work-Conserving SchedulingabstractCurrent operating system schedulers are not fully aware of multi-core and multi-threaded architectures, and as a result, schedule threads in a way that may cause contention for critical resources such as the last level in the cache memory hierarchy or the memory access bandwidth. This contention has a significant impact on the system productivity and the quality of service that each individual thread gets from the platform, which can widely vary depending on the behavior of its simultaneous co-runners.In this paper we describe the design and implementation of a non-work-conserving framework to schedule threads that tries to improve priority enforcement, based on on-line statistics collected through hardware performance counters. We have implemented our scheme in Linux running on both multicore and SMT processors. For synthetic workloads based on the latest SPEC CPU2006 benchmarks, our framework speeds up high-priority threads by up to 50%, while keeping or even slightly improving the overall system throughput. Juan Carlos Saez, José Ignacio Gómez, Manuel Prieto 0001 |
ICPP | 3 |
| 2008 | Combining system scenarios and configurable memories to tolerate unpredictabilityabstractProcess variability and the dynamism of new applications increase the uncertainty of embedded systems and force designers to use pessimistic assumptions, which have a tremendous impact on both the performance and energy consumption of their memory organizations. In this article we introduce an experimental framework which tries to mitigate the effects of both sources of unpredictability. At compile time, an extensive profiling helps us to detect system scenarios and bounds application dynamism. At the organization level, we incorporate a heterogeneous memory architecture composed by several configurable memories. A calibration process and a runtime control system adapt the platform to the current application needs. Our approach manages to reduce significantly the energy overhead associated to both variability and application dynamism (up to 60%, according to our simulations) without compromising the timing constraints existing in our target domain of dynamic periodic multimedia applications. Concepción Sanz, Manuel Prieto 0001, José Ignacio Gómez, Antonis Papanikolaou, Miguel Corbalan, Francky Catthoor |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2008 | Parallel Implementation of the 2D Discrete Wavelet Transform on Graphics Processing Units: Filter Bank versus LiftingabstractThe widespread usage of the DiscreteWaveletTransform (DWT) has motivated the development of fastDWT algorithms and their tuning on all sorts of computersystems. Several studies have compared the performanceof the most popular schemes, known as Filter Bank(FBS) and Lifting (LS), and have always concluded thatLifting is the most efficient option. However, there isno such study on streaming processors such as modernGraphic Processing Units (GPUs). Current trends havetransformed these devices into powerful stream processorswith enough flexibility to perform intensive and complexfloating-point calculations. The opportunities opened upby these platforms, as well as the growing popularityof the DWT within the computer graphics field, make anew performance comparison of great practical interest.Our study indicates that FBS outperforms LS in currentgeneration GPUs. In our experiments, the actual FBS gainsrange between 10% and 140%, depending on the problemsize and the type and length of the wavelet filter. Moreover,design trends suggest higher gains in future generationGPUs. Christian Tenllado, Javier Setoain, Manuel Prieto 0001, Luis Piñuel, Francisco Tirado |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2007 | Parallel Morphological Endmember Extraction Using Commodity Graphics HardwareabstractSpatial/spectral algorithms have been shown in previous work to be a promising approach to the problem of extracting image end members from remotely sensed hyperspectral data. Such algorithms map nicely on high-performance systems such as massively parallel clusters and networks of computers. Unfortunately, these systems are generally expensive and difficult to adapt to onboard data processing scenarios, in which low-weight and low-power integrated components are highly desirable to reduce mission payload. An exciting new development in this context is the emergence of graphics processing units (GPUs), which can now satisfy extremely high computational requirements at low cost. In this letter, we propose a GPU-based implementation of the automated morphological end member extraction algorithm, which is used in this letter as a representative case study of joint spatial/spectral techniques for hyperspectral image processing. The proposed implementation is quantitatively assessed in terms of both end member extraction accuracy and parallel efficiency, using two generations of commercial GPUs from NVidia. Combined, these parts offer a thoughtful perspective on the potential and emerging challenges of implementing hyperspectral imaging algorithms on commodity graphics hardware. Javier Setoain, Manuel Prieto 0001, Christian Tenllado, Antonio Plaza, Francisco Tirado |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2006 | Substituting associative load queue with simple hash tables in out-of-order microprocessorsabstractBuffering more in-flight instructions in an out-of-order microprocessor is a straightforward and effective method to help tolerate the long latencies generally associated with off-chip memory accesses. One of the main challenges of buffering a large number of instructions, however, is the implementation of a scalable and efficient mechanism to detect memory access order violations as a result of out-of-order scheduling of load and store instructions. Traditional CAM-based associative queues can be very slow and energy consuming. In this paper, instead of using the traditional age-based load queue to record load addresses, we explicitly record age information in address-indexed hash tables to achieve the same functionality of detecting premature loads. This alternative design eliminates associative searches and significantly reduces the energy consumption of the load queue. With simple techniques to reduce the number of false positives, performance degradation is kept at a minimum. Alok Garg, Fernando Castro, Michael C. Huang 0001, Daniel Chaver, Luis Piñuel, Manuel Prieto 0001 |
ISLPED | 6 |
| 2006 | DMDC: Delayed Memory Dependence Checking through Age-Based FilteringabstractOne of the main challenges of modern processor design is the implementation of a scalable and efficient mechanism to detect memory access order violations as a result of out-of-order execution of memory instructions. Traditional CAM-based associative queues can be very slow and energy hungry. In this paper we introduce two new management schemes. The first one is a filtering scheme based on simple age-tracking. This scheme can easily avoid 95-98% of associative load queue (LQ) searches using only a few registers. This translates into significant power savings. More importantly, however, this filtering makes our second scheme, delayed memory dependence checking (DMDC), practical. With a small hash table, DMDC completely avoids the need for an associative LQ and relies on indexing-based checking at the commit phase and hence cuts the energy spent on LQ by an average of 95%. At an average of about 0.3%, the performance impact is negligible. When the energy cost of the increased execution time is factored in, the processor still makes net energy savings of about 3-8%, depending on the configuration and the applications Fernando Castro, Luis Piñuel, Daniel Chaver, Manuel Prieto 0001, Michael C. Huang 0001, Francisco Tirado |
MICRO | 4 |
| 2005 | Load-Store Queue Management: an Energy-Efficient Design Based on a State-Filtering MechanismabstractModern microprocessors incorporate sophisticated techniques to allow early execution of loads without compromising program correctness. To do so, the structures that hold the memory instructions (load and store queues) implement several complex mechanisms to dynamically resolve the memory-based dependences. Our main objective in this paper is to design an efficient LQ-SQ structure, which saves energy without sacrificing much performance. We propose a new design that divides the load queue into two structures, a conventional associative queue and a simpler FIFO queue that does not allow associative searching. A dependence predictor predicts whether a load instruction has a memory dependence on any inflight store instruction. If so, the load is sent to the conventional associative queue. Otherwise, it is sent to the non-associative queue which can only detect dependence in an inexact and conservative way. In addition, the load will not check the store queue at execution time. These measures combined save energy consumption. We explore different predictor designs and runtime policies. Our experiments indicate that such a design can reduce the energy consumption in the load-store queue by 35-50% with an insignificant performance penalty of about 1%. When the energy cost of the increased execution time is factored in, the processor still makes net energy savings of about 3-4%. Fernando Castro, Daniel Chaver, Luis Piñuel, Manuel Prieto 0001, Francisco Tirado, Michael C. Huang 0001 |
ICCD | 4 |
| 2005 | Energy-aware fetch mechanism: trace cache and BTB customizationabstractA highly-efficient fetch unit is essential not only to obtain good performance but also to achieve energy efficiency. However, existing designs are inflexible and depending on program behavior, can be either insufficient or an overkill. We introduce a phase-based adaptive fetch mechanism that can be dynamically adjusted based on feedback information of the program behavior. This design adds very little hardware complexity and relegates complex tasks to the software components. It is also very effective: saving 26.8% and 34.1% fetch energy on average compared with a conventional and a trace cache-based fetch unit, respectively. At the same time, performance is improved by 5.7% and 0.6%, respectively Daniel Chaver, Miguel A. Rojas, Luis Piñuel, Manuel Prieto 0001, Francisco Tirado, Michael C. Huang 0001 |
ISLPED | 4 |
| 2003 | Branch prediction on demand: an energy-efficient solutionabstractHigh-end processors typically incorporate complex branch predictors consisting of many large structures that together consume a notable fraction of total chip power (more than 10% in some cases). Depending on the applications, some of these resources may remain underused for long periods of time. We propose a methodology to reduce the energy consumption of the branch predictor by characterizing prediction demand using profiling and dynamically adjusting predictor resources accordingly. Specifically, we disable components of the hybrid direction predictor and resize the branch target buffer. Detailed simulations show that this approach reduces the energy consumption in the branch predictor by an average of 72% and up to 89% with virtually no impact on prediction accuracy and performance. Daniel Chaver, Luis Piñuel, Manuel Prieto 0001, Francisco Tirado, Michael C. Huang 0001 |
ISLPED | 3 |
| 2003 | A parallel multigrid solver for viscous flows on anisotropic structured grids
Manuel Prieto 0001, Rubén S. Montero, Ignacio Martín Llorente, Francisco Tirado |
Parallel Comput. | 1 |
| 2002 | -D Wavelet Transform Enhancement on General-Purpose Microprocessors: Memory Hierarchy and SIMD Parallelism Exploitation
Daniel Chaver, Christian Tenllado, Luis Piñuel, Manuel Prieto 0001, Francisco Tirado |
HiPC | 4 |
| 2001 | A Multigrid Solver for the Incompressible Navier-Stokes Equations on a Beowulf-Class SystemabstractThis paper presents an efficient parallel multigrid solver for speeding up the computation of a 3-D model that treats the flow of a viscous fluid over a solid obstacle. From a numerical point of view, the main interest of this simulation lies in exhibiting some basic difficulties that prevent optimal multigrid efficiencies from being achieved. As the computing platform we have used Coral, a Beowulf-class system based on a high performance GigaNet cLAN interconnect and Dual Pentium III nodes. The solver which has been devised taking both algorithmic and architectural issues into account, not only attains good convergence rates, but also achieves satisfactory efficiencies on the target cluster. Manuel Prieto 0001, Rubén S. Montero, Ignacio Martín Llorente, Francisco Tirado |
ICPP | 1 |
| 2001 | Parallel Multigrid for Anisotropic Elliptic Equations
Manuel Prieto 0001, R. Santiago, David Espadas, Ignacio Martín Llorente, Francisco Tirado |
J. Parallel Distributed Comput. | 1 |
| 2001 | A parallel multigrid solver for 3D convection and convection-diffusion problems
Ignacio Martín Llorente, Manuel Prieto 0001, Boris Diskin |
Parallel Comput. | 2 |
| 2000 | Impact of PE Mapping on Cray T3E Message-Passing Performance
Eduardo Huedo, Manuel Prieto 0001, Ignacio Martín Llorente, Francisco Tirado |
Euro-Par | 2 |
| 2000 | Data Locality Exploitation in the Decomposition of Regular Domain ProblemsabstractThe aim of this paper is to study the effect of local memory hierarchy and communication network exploitation on message sending and the influence of this effect on the decomposition of regular applications. In particular, we have considered two different parallel computers, a Cray T3E-900 and an SGI Origin 2000. In both systems, the bandwidth reduction due to non-unit-stride memory access is quite significant and could be more important than the reduction due to contention in the network. These conclusions affect the choice of optimal decompositions for regular domains problems. Thus, although traditional 3D decompositions lead to lower inherent communication-to-computation ratios and could exploit more efficiently the interconnection network, lower dimensional decompositions are found to be more efficient due to the data decomposition effects on the spatial locality of the messages to be communicated. This increasing importance of local optimisations has also been shown using a well-known communication-computation overlapping technique which increases execution time, instead of reducing it as we could expect, due to poor cache memory exploitation. Manuel Prieto 0001, Ignacio Martín Llorente, Francisco Tirado |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 1999 | Message Passing Evaluation and Analysis on Cray T3E and SGI Origin 2000 Systems
Manuel Prieto 0001, David Espadas, Ignacio Martín Llorente, Francisco Tirado |
Euro-Par | 1 |
| 1999 | Solution of Alternating-Line Processes on Modern Parallel ComputersabstractThe aim of this paper is the study of different methods for the solution of alternating-line problems, taking into account the evolution of architectural parameters on modern parallel computers, i.e. processors, memory hierarchy, and interconnection network performance. Three different kinds of solvers are studied: The Pipelined Gaussian Elimination scheme, the Matrix Transposition scheme, and the new one which is presented in this paper: the Mapping Transposition scheme, whose performance clearly betters, in many cases, that obtained by all the other methods, due to its better fitting to the characteristics of modern parallel computers. The experimental results have been obtained on a Cray T3E and on an SGI Origin 2000, up to 512 and 32 processors, respectively. David Espadas, Manuel Prieto 0001, Ignacio Martín Llorente, Francisco Tirado |
ICPP | 2 |
| 1999 | An environment to develop parallel code for solving partial differential equations based-problems
Manuel Prieto 0001, Ignacio Martín Llorente, Francisco Tirado |
J. Syst. Archit. | 1 |