EDBT 2026 Demo / reviewers in the wild / expert
Adrián Barredo
dblp:252/8055
· DBLP profile ↗
9ranked-venue papers
5as first author
4since 2021 · last 2025
0000-0001-9435-3234ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 5 first-author · 4 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ParaVerser: Harnessing Heterogeneous Parallelism for Affordable Fault Detection in Data CentersabstractData-Center operators have awoken to the fact that silent data corruption resulting from defective silicon compute units is endemic at scale. Software scanners have been deployed to mitigate the issue, but either have low coverage or take months, leaving long windows of incorrect behaviour. By contrast, the redundancy mechanisms used in automotive double the required power and area, so cannot be practically deployed in server-space.We present ParaVerser, a high-coverage, low-overhead solution to hardware-level error detection in servers. Through minor architectural modifications, we enable conventional cores in heterogeneous server-grade processors to act as checker cores, thus exploiting heterogeneity, frequency scaling and the inherent parallelism in repeat runs to provide energy-efficient error checking. By dynamically coupling big.LITTLE-style out-of-order superscalar cores with in-order ones, we reduce energy overheads relative to a typical lockstep system by 70% with identical guarantees, at only 4.3% performance degradation, and 1064B per-core area overhead. Minli Julie Liao, Sam Ainsworth 0001, Lev Mukhanov, Adrián Barredo, Markos Kynigos, Timothy M. Jones 0001 |
DSN | 4 |
| 2022 | Compiler-Assisted Compaction/Restoration of SIMD InstructionsabstractVector processors (e.g., SIMD or GPUs) are ubiquitous in high performance systems. All the supercomputers in the world exploit data-level parallelism (DLP), for example by using single instructions to operate over several data elements. Improving vector processing is therefore key for exascale computing. However, despite its potential, vector code generation and execution have significant challenges. Among these challenges, control flow divergence is one of the main performance limiting factors. Most modern vector instruction sets, including SIMD, rely on predication to support divergence control. Nevertheless, the performance and energy consumption in predicated codes is usually insensitive to the number of active elements in a predicated mask. Since the trend is that vector register size increases, the energy efficiency of exascale computing systems will become sub-optimal. This article proposes a novel approach to improve execution efficiency in predicated vector codes, the Compiler-Assisted Compaction/Restoration (CACR) technique. Baseline CR delays predicated SIMD instructions with inactive elements, compacting active elements from instances of the same instruction of consecutive loop iterations. Compacted elements form an equivalentdensevector instruction. After executing the dense instructions, their results are restored to the original instructions. However, CR has a significant performance and energy penalty when it fails to find active elements, either due to lack of resources when unrolling or because of inter-loop dependencies. In CACR, the compiler analyzes the code looking for key information required to configure CR. Then, it passes this information to the processor via new instructions inserted in the code. This prevents CR from waiting for active elements on scenarios when it would fail to form dense instructions. Simulated results (gem5) show that CACR improves performance by up to 29 percent and reduces dynamic energy by up to 24.2 percent on average, for a a set of applications with predicated execution. The baseline CR only achieves 18.6 percent performance and 14 percent energy improvements for the same configuration and applications. Juan M. Cebrian, Thibaud Balem, Adrián Barredo, Marc Casas, Miquel Moretó, Alberto Ros 0001, Alexandra Jimborean |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2021 | VIA: A Smart Scratchpad for Vector Units with Application to Sparse Matrix ComputationsabstractSparse matrix operations are critical kernels in multiple application domains such as High Performance Computing, artificial intelligence and big data. Vector processing is widely used to improve performance on mathematical kernels with dense matrices. Unfortunately, existing vector architectures do not cope well with sparse matrix computations, achieving much lower performance in comparison with their dense counterparts.To overcome this limitation, we present the Vector Indexed Architecture (VIA), a novel hardware vector architecture that accelerates applications with irregular memory access patterns such as sparse matrix computations. There are two main bottlenecks when computing with sparse matrices: irregular memory accesses and index matching. VIA addresses these two bottlenecks with a smart scratchpad that is tightly coupled to the Vector Functional Units within the core.Thanks to this structure, VIA improves locality for sparse-dense computations and improves the index matching search process for sparse computations. As a result, VIA achieves significant performance speedup over highly optimized state-of-the-art C++ algebra libraries. On average, VIA outperforms sparse matrix vector multiplication, sparse matrix addition and sparse matrix matrix multiplication kernels by 4.22 ×, 6.14 × and 6.00 ×, respectively, when evaluated over a thousand sparse matrices that arise in real applications. In addition, we prove the generality of VIA by showing that it can accelerate histogram and stencil applications by 4.5 × and 3.5 ×, respectively. Julian Pavon, Iván Vargas Valdivieso, Adrián Barredo, Joan Marimon, Miquel Moretó, Francesc Moll, Osman S. Unsal, Mateo Valero, Adrián Cristal |
HPCA | 3 |
| 2021 | PLANAR: a programmable accelerator for near-memory data rearrangementabstractMany applications employ irregular and sparse memory accesses that cannot take advantage of existing cache hierarchies in high performance processors. To solve this problem, Data Layout Transformation (DLT) techniques rearrange sparse data into a dense representation, improving locality and cache utilization. However, prior proposals in this space fail to provide a design that (i) scales with multi-core systems, (ii) hides rearrangement latency, and (iii) provides the necessary interfaces to ease programmability. Adrián Barredo, Adrià Armejach, Jonathan C. Beard, Miquel Moretó |
ICS | 1 |
| 2020 | Improving Predication Efficiency through Compaction/Restoration of SIMD InstructionsabstractVector processors offer a wide range of unexplored opportunities to improve performance and energy efficiency. However, despite its potential, vector code generation and execution have significant challenges, the most relevant ones being control flow divergence. Most modern processors including SIMD extensions (such as AVX) rely on predication to support divergence control. In predicated codes, performance and energy consumption are usually insensitive to the number of true values in a predicated mask. This implies that the system efficiency becomes sub-optimal as vector length increases. In this paper we focus on SIMD extensions and propose a novel approach to improve execution efficiency in predicated SIMD instructions, the Compaction/Restoration (CR) technique. CR delays predicated SIMD instructions with inactive elements and compacts them with instances of the same instruction from different loop iterations to form an equivalent dense vector instruction, where, in the best case, all the elements are active. After executing such dense instructions, their results are restored to the original instructions. Our evaluation shows that CR improves performance by up to 25% and reduces dynamic energy consumption by up to 43% on real unmodified applications with predicated execution. Moreover, CR allows executing unmodified legacy code with short vector instructions (AVX-2) on newer architectures with wider vectors (AVX-512), achieving up to 56% performance benefits. Adrián Barredo, Juan M. Cebrian, Miquel Moretó, Marc Casas, Mateo Valero |
HPCA | 1 |
| 2020 | Semi-automatic validation of cycle-accurate simulation infrastructures: The case for gem5-x86
Juan M. Cebrian, Adrián Barredo, Helena Caminal, Miquel Moretó, Marc Casas, Mateo Valero |
Future Gener. Comput. Syst. | 2 |
| 2020 | Efficiency analysis of modern vector architectures: vector ALU sizes, core counts and clock frequencies
Adrián Barredo, Juan M. Cebrian, Mateo Valero, Marc Casas, Miquel Moretó |
J. Supercomput. | 1 |
| 2019 | POSTER: SPiDRE: Accelerating Sparse Memory Access PatternsabstractDevelopment in process technology has led to an exponential increase in processor speed and memory capacity. However, memory latencies have not improved as dramatically and represent a well-known problem in computer architecture. Cache memories provide more bandwidth with lower latencies than main memories but they are capacity limited. Locality-friendly applications benefit from a large and deep cache hierarchy. Nevertheless, this is a limited solution for applications suffering from sparse and irregular memory access patterns, such as data analytics. In order to accelerate them, we should maximize usable bandwidth, reduce latency and maximize moved data reuse. In this work we explore the Sparse Data Rearrange Engine (SPiDRE), a novel hardware approach to accelerate these applications through near-memory data reorganization. Adrián Barredo, Jonathan C. Beard, Miquel Moretó |
PACT | 1 |
| 2019 | POSTER: An Optimized Predication Execution for SIMD ExtensionsabstractVector processing is a widely used technique to improve performance and energy efficiency in modern processors. Most of them rely on predication to support divergence control. However, performance and energy consumption in predicated instructions are usually independent on the number of true values in a mask. This means that the efficiency of the system becomes sub-optimal as vector length increases. In this work we propose the Optimized Predication Execution (OPE) technique. OPE delays the execution of sparse masked vector instructions sharing the same PC, extracts their active elements and creates a new dense instruction with a higher mask density. After executing such dense instruction, results are restored to the original sparse instructions. Our approach improves performance by up to 25% and reduces dynamic energy consumption by up to 43% on real applications with predication. Adrián Barredo, Juan M. Cebrian, Miquel Moretó, Marc Casas, Mateo Valero |
PACT | 1 |