EDBT 2026 Demo / reviewers in the wild / expert
Diego Andrade
dblp:38/10
· DBLP profile ↗
25ranked-venue papers
9as first author
4since 2021 · last 2025
0000-0001-5670-7425ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 21 · 7 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Adapt-S: Effective DNN Pruning via Unified Accuracy and Performance TuningabstractModel sparsification has emerged as a promising approach to reducing model size with minimum impact on accuracy. This is achieved through the removal of some model parameters, a process also known as Deep Neural Network (DNN) pruning. The irregular nature of the generated sparse tensors poses a great challenge in the development of efficient GPU kernels optimized for these workloads. This challenge has been recently addressed through the use of hardware-aware semistructured sparsification methods designed to conform to specialized sparse formats and codesigned with template-based kernel implementations. These methods are commonly based on grouping the non-pruned values in blocks of a given size to generate regularity or on generating patterns that fit specialized hardware units. This pruning pattern-format-kernel triplet presents a high degree of tunability, both at the pruning and kernel sides, which can be used to fit certain accuracy-to-performance tradeoffs. On the pruning side, using larger blocks of consecutive non-pruned values favors performance over accuracy, as the weight selection for removal policy becomes less flexible. On the kernel side, recent studies have proven that the tuning of the configuration parameters of template-based multi-level tiling kernel implementations can yield an extra performance boost. This paper presents AdAPT-S, an autotuning system that generates DNN pruning recipes and optimized kernel configurations to fit an accuracy-to-performance specification. This is done through a cost model that integrates both aspects. AdAPT-S gets extra benefits from the exploitation of layer sensitivity by providing per-layer pruning recipes and kernel configurations. The results show that our approach can achieve superior accuracy-to-performance trade-offs and that this can be used to produce models that fit the user requirements. Roberto L. Castro, Diego Andrade, Basilio B. Fraguela |
IPDPS | 2 |
| 2023 | VENOM: A Vectorized N: M Format for Unleashing the Power of Sparse Tensor CoresabstractThe increasing success and scaling of Deep Learning models demands higher computational efficiency and power. Sparsification can lead to both smaller models as well as higher compute efficiency, and accelerated hardware is becoming available. However, exploiting it efficiently requires kernel implementations, pruning algorithms, and storage formats, to utilize hardware support of specialized sparse vector units. An example of those are the NVIDIA's Sparse Tensor Cores (SPTCs), which promise a 2× speedup. However, SPTCs only support the 2:4 format, limiting achievable sparsity ratios to 50%. We present the V:N:M format, which enables the execution of arbitrary N:M ratios on SPTCs. To efficiently exploit the resulting format, we propose Spatha, a high-performance sparse-library for DL routines. We show that Spatha achieves up to 37× speedup over cuBLAS. We also demonstrate a second-order pruning technique that enables sparsification to high sparsity ratios with V:N:M and little to no loss in accuracy in modern transformers. Roberto L. Castro, Andrei Ivanov, Diego Andrade, Tal Ben-Nun, Basilio B. Fraguela, Torsten Hoefler |
SC | 3 |
| 2022 | Probing the Efficacy of Hardware-Aware Weight Pruning to Optimize the SpMM Routine on Ampere GPUsabstractThe Deep Learning (DL) community found in pruning techniques a good way to reduce the models' resource and energy consumption. These techniques lead to smaller sparse models, but sparse computations in GPUs only outperform their dense counterparts for extremely high levels of sparsity. However, pruning up to such sparsity levels can seriously harm the accuracy of the Neural Networks (NNs). To alleviate this, novel performance-aware pruning techniques favor the generation of more regular sparse matrices that can improve the exploitation of the underlying hardware. Nevertheless, an important drawback is that these techniques heavily condition the location of the non-pruned values, which can strongly degrade the accuracy of the models. Roberto L. Castro, Diego Andrade, Basilio B. Fraguela |
PACT | 2 |
| 2021 | High-performance dataflow computing in hybrid memory systems with UPC++ DepSpawn
Basilio B. Fraguela, Diego Andrade |
J. Supercomput. | 2 |
| 2020 | An automatic optimizer for heterogeneous devices
Jorge Fernández-Fabeiro, Diego Andrade, Basilio B. Fraguela, Ramón Doallo |
Future Gener. Comput. Syst. | 2 |
| 2019 | Easy Dataflow Programming in Clusters with UPC++ DepSpawnabstractThe Partitioned Global Address Space (PGAS) programming model is one of the most relevant proposals to improve the ability of developers to exploit distributed memory systems. However, despite its important advantages with respect to the traditional message-passing paradigm, PGAS has not been yet widely adopted. We think that PGAS libraries are more promising than languages because they avoid the requirement to (re)write the applications using them, with the implied uncertainties related to portability and interoperability with the vast amount of APIs and libraries that exist for widespread languages. Nevertheless, the need to embed these libraries within a host language can limit their expressiveness and very useful features can be missing. This paper contributes to the advance of PGAS by enabling the simple development of arbitrarily complex task-parallel codes following a dataflow approach on top of the PGAS UPC++ library, implemented in C++. In addition, our proposal, called UPC++ DepSpawn, relies on an optimized multithreaded runtime that provides very competitive performance, as our experimental evaluation shows. Basilio B. Fraguela, Diego Andrade |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2018 | Heterogeneous distributed computing based on high-level abstractionsabstractSummary The rise of heterogeneous systems has given place to great challenges for users as they involve new concepts, restrictions, and frameworks. Their exploitation is further complicated in the context of distributed memory systems, which require the usage of additional different programming paradigms and tools. In this paper, we propose a novel approach to program heterogeneous clusters that is based on high‐level abstractions such as tiles and hierarchical decomposition combined with the powerful APIs that data types and embedded languages can provide in languages such as C++. Rather than building our proposal from scratch, we have implemented it as a natural integration of the existing Hierarchically Tiled Arrays (HTA) and Heterogeneous Programming Library (HPL) projects, ie, the first one being focused on distributed computing and the second one on heterogeneous processing. The result, called Heterogeneous Hierarchically Tiled Arrays (H2TA), is very intuitive and easy to use thanks to the global view of the data and the single‐threaded view of the execution that it provides at cluster level together with the transparency it provides with respect to the management of the heterogeneous devices. An evaluation comparing our proposal with MPI‐based implementations shows its large programmability advantages and the reasonable overhead incurred. Moisés Viñas, Basilio B. Fraguela, Diego Andrade, Ramón Doallo |
Concurr. Comput. Pract. Exp. | 3 |
| 2017 | Facilitating the development of stencil applications using the Heterogeneous Programming LibraryabstractSummary Stencil computations are very common in scientific codes. Heterogeneous systems achieve good results solving these problems, but their programming is complex because of the ghost regions required in multi‐device implementations and the difficulty to properly exploit their hardware. The Heterogeneous Programming Library (HPL) is a recent framework that improves the programmability of heterogeneous devices. This paper describes two extensions of HPL focused on stencil computations. The first one allows to automatically update the ghost regions they involve. The second one automates the implementation of the computational kernels of these algorithms. In our evaluation, the first mechanism reduces on average the number of lines of code and the Halstead programming effort of the host code of comparable HPL baselines by 34% and 64.2%, respectively, while the second contribution reduces these metrics by 72% and 79% in the computational kernels, respectively. Also, the first technique has negligible performance overheads, while the second one matches the performance of manually developed kernels. As an added benefit, the facilitation of the development of these codes thanks to these techniques helps programmers experiment with optimizations suited for this applications such as the ghost cell expansion technique, which provides speedups of up to 13% in our experiments. Moisés Viñas, Basilio B. Fraguela, Diego Andrade, Ramón Doallo |
Concurr. Comput. Pract. Exp. | 3 |
| 2017 | High productivity multi-device exploitation with the Heterogeneous Programming Library
Moisés Viñas, Basilio B. Fraguela, Diego Andrade, Ramón Doallo |
J. Parallel Distributed Comput. | 3 |
| 2016 | Writing a performance-portable matrix multiplication
Jorge Fernández-Fabeiro, Diego Andrade, Basilio B. Fraguela |
Parallel Comput. | 2 |
| 2015 | Automatic Generation of Optimized OpenCL Codes Using OCLoptimizerabstractThe eruption of multicore processors and several kinds of accelerators has generalized the interest in parallel programming. The OpenCL standard is very appealing because it provides code portability across most of these platforms. It defines a programming model where a host code requests the execution of kernels in computational devices. Unfortunately, the host application programming interface of OpenCL is quite verbose, which makes the development of its host code tedious and error-prone. More importantly, OpenCL does not provide automatic performance portability. As a result, users have to hand-tune OpenCL codes for each specific device, which implies trying different versions of the kernels and task partition granularities. As an answer to this situation, we present OCLoptimizer, a tool that automatically generates host codes and optimizes OpenCL kernels for each specific target device based on a user provided configuration file. This configuration file describes basic kernel characteristics and annotations in the kernels that indicate the code transformations to test. Our tool can explore different granularities for the problem decomposition as well as different alternatives for the kernel. This exploration is performed by means of an iterative optimization process whose parameters and search strategy are defined by the user specifications. Support for OpenCL codes composed of multiple kernels is also provided by the tool. Experiments performed on multicore CPUs and different accelerators show that the tool is very effective, generating codes with an average speedup of 2.54 with respect to baseline hand-tuned implementations, in single kernel codes and 1.79 in a code with multiple kernels. Jorge Fernández-Fabeiro, Diego Andrade, Basilio B. Fraguela, Ramón Doallo |
Comput. J. | 2 |
| 2015 | Developing adaptive multi-device applications with the Heterogeneous Programming Library
Moisés Viñas, Zeki Bozkus, Basilio B. Fraguela, Diego Andrade, Ramón Doallo |
J. Supercomput. | 4 |
| 2014 | Writing Self-adaptive Codes for Heterogeneous Systems
Jorge Fernández-Fabeiro, Diego Andrade, Basilio B. Fraguela, Ramón Doallo |
Euro-Par | 2 |
| 2014 | A fine-grained thread-aware management policy for shared cachesabstractSUMMARY Two of the main sources of inefficiency in current caches are the non‐uniform distribution of the memory accesses across the cache sets, which causes misses due to the mapping restrictions of non fully associative caches and the access patterns with little locality that degrade the performance of caches under the traditional least recently used. replacement policy. This paper proposes a technique to tackle in a coordinated way both kinds of problems in the context of chip multiprocessors, whose last level caches can be shared by threads with different patterns of locality. Our proposal, called thread‐aware mapping and replacement miss reduction (TAMR2) policy, tracks the behavior of each thread in each set in order to decide the appropriate combination of policies to deal with these problems. Despite its small overhead, TAMR2 achieved in our experiments average power consumption and memory latency reductions of 10% and 12%, respectively, resulting in an average throughput improvement of 5.6%, relative to a traditional cache design using four cores. TAMR2 also outperformed many recent related approaches in the field. Copyright © 2013 John Wiley & Sons, Ltd. Dyer Rolán, Diego Andrade, Basilio B. Fraguela, Ramón Doallo |
Concurr. Comput. Pract. Exp. | 2 |
| 2013 | Accurate prediction of the behavior of multithreaded applications in shared caches
Diego Andrade, Basilio B. Fraguela, Ramón Doallo |
Parallel Comput. | 1 |
| 2013 | Numerical simulation of pollutant transport in a shallow-water system on the Cell heterogeneous processor
Carlos H. Gonzalez, Basilio B. Fraguela, Diego Andrade, José A. García, Manuel Jesús Castro Díaz |
J. Supercomput. | 3 |
| 2012 | Using an Analytical Model of Shared Caches for Selecting the Optimal Parallelization SchemeabstractMulticores are now the norm. Their cache hierarchy has often a last level shared cache. The performance of this shared cache during the execution of multithreaded applications depends on the parallelization scheme followed. For example, critical parameters for the performance of parallelized loops are the number of threads and the block size. The selection of the optimal scheme in a compiler can be guided using heuristics or the execution time. Heuristics can be imprecise, while an execution time guided search is very time-consuming. This paper shows the usage of an analytical model to predict the cache behavior of shared caches during the execution of multithreaded applications that have been parallelized at loop level. The model predicts the number of misses generated by a given code when different number of threads or block sizes are used. The execution time of the codes analyzed is highly correlated to the number of misses generated in the shared cache, thus, the prediction of the model is a powerful tool to select the best parallelization scheme for them. Diego Andrade, Basilio B. Fraguela, Ramón Doallo |
ISPA | 1 |
| 2012 | Static analysis of the worst-case memory performance for irregular codes with indirectionsabstractReal-time systems are subject to timing constraints, whose upper bound is given by the Worst-Case Execution Time (WCET). Cache memory behavior is difficult to predict analytically and estimating a safe and precise worst-case value is even more challenging. The worst-case memory performance (WCMP) component of the WCET can only be estimated with the precise knowledge of the stream of data addresses accessed by the code, which is determined by the access patterns and the base addresses of the data structures accessed. The regularity of strided access patterns simplifies their analysis, as they are characterized by relatively few parameters, which are often available at compile time. Unfortunately codes may exhibit irregular access patterns, which are much more difficult to statically analyze. As for the base addresses of the data structures, they are not always available at compile-time for many reasons: stack variables, dynamically allocated memory, modules compiled separately, etc. This article addresses these problems by presenting a model that predicts an %safe and upper bound of the data cache performance for codes both with regular and irregular access patterns, which is valid for any possible base addresses of the data structures. The model analyzes irregular access patterns due to the presence of indirections in the code and it can provide two kinds of predictions: a safe hard boundary that is suitable for hard real-time systems and a soft boundary whose safeness is not guaranteed but which is valid most of the times. In fact, in all our experiments the number of misses was below the soft boundary predicted by the model. This turns this soft boundary prediction into a valuable tool, particularly for non and soft real-time systems, which tolerate a percentage of the runs exceeding their deadlines. Diego Andrade, Basilio B. Fraguela, Ramón Doallo |
ACM Trans. Archit. Code Optim. | 1 |
| 2010 | Address-Independent Estimation of the Worst-case Memory PerformanceabstractReal-time systems are subject to temporal constraints and require a schedulability analysis to ensure that task execution finishes within lower and upper specified bounds. Worst-case memory performance (WCMP) plays a key role in the calculation of the upper bound of the execution time. Data caches complicate the calculation of the WCMP, since their behavior is highly dependent on the sequence of memory addresses accessed, which is often not available. For example, the address of a data structure may not be available at compile-time, and it may change between different executions of the program. We present an analytical model that provides fast, safe and tight estimations of the WCMP component of the worst-case execution time, using no information about the data base addresses. The address-independent absolute WCMP for codes with references that follow the same access pattern can be very high with respect to the average behavior because those references may be aligned with respect to the cache, thus generating systematic interferences among them. Our model can also provide a tighter and safe estimation for the WCMP for these codes when the user avoids these alignments. Basilio B. Fraguela, Diego Andrade, Ramón Doallo |
IEEE Trans. Ind. Informatics | 2 |
| 2009 | Task-Parallel versus Data-Parallel Library-Based Programming in Multicore SystemsabstractMulticore machines are becoming common. There are many languages, language extensions and libraries devoted to improve the programmability and performance of these machines. In this paper we compare two libraries, that face the problem of programming multi-cores from two different perspectives, task parallelism and data parallelism. The Intel threading building blocks (TBB) library separates logical task patterns, which are easy to understand, from physical threads, and delegates the scheduling of the tasks to the system. On the other hand, hierarchically tiled arrays (HTAs) are data structures that facilitate locality and parallelism of array intensive computations with a block-recursive nature following a data-parallel paradigm. Our comparison considers both ease of programming and the performance obtained using both approaches. In our experience, HTA programs tend to be smaller or as long as TBB programs, while performance of both approaches is very similar. Diego Andrade, Basilio B. Fraguela, James C. Brodman, David A. Padua |
PDP | 1 |
| 2009 | Static Prediction of Worst-Case Data Cache Performance in the Absence of Base Address InformationabstractWhile caches are essential to reduce execution time and power consumption, they complicate the estimation of the worst-case execution time (WCET), crucial for many real-time systems (RTS). Most research on static worst-case cache behavior prediction has focused on hard RTS, which need complete information on the access patterns and addresses of the data to guarantee the predicted WCET is a safe upper bound of any execution time. Access patterns are available in those codes that have a steady state of access patterns after the first iteration of a loop (in the following regular codes), however, the addresses of the data are not always known at compile time for many reasons: stack variables, dynamically allocated memory, modules compiled separately, etc. Even when available, their usefulness to predict cache behavior in systems with virtual memory decreases in the presence of physically-indexed caches. In this paper we present a model that predicts a reasonable bound of the worst-case behavior of data caches during the execution of regular codes without information on the base address of the data structures. In 99.7% of our tests the number of misses performed below the boundary predicted by the model. This turns the model into a valuable tool, particularly for non-RTS and soft RTS, which tolerate a percentage of the runs exceeding their deadlines. Diego Andrade, Basilio B. Fraguela, Ramón Doallo |
IEEE Real-Time and Embedded Technology and Applications Symposium | 1 |
| 2007 | Automated and accurate cache behavior analysis for codes with irregular access patternsabstractAbstract The memory hierarchy plays an essential role in the performance of current computers, so good analysis tools that help in predicting and understanding its behavior are required. Analytical modeling is the ideal base for such tools if its traditional limitations in accuracy and scope of application can be overcome. While there has been extensive research on the modeling of codes with regular access patterns, less attention has been paid to codes with irregular patterns due to the increased difficulty in analyzing them. Nevertheless, many important applications exhibit this kind of pattern, and their lack of locality make them more cache‐demanding, which makes their study more relevant. The focus of this paper is the automation of the Probabilistic Miss Equations (PME) model, an analytical model of the cache behavior that provides fast and accurate predictions for codes with irregular access patterns. The information requirements of the PME model are defined and its integration in the XARK compiler, a research compiler oriented to automatic kernel recognition in scientific codes, is described. We show how to exploit the powerful information‐gathering capabilities provided by this compiler to allow the automated modeling of loop‐oriented scientific codes. Experimental results that validate the correctness of the automated PME model are also presented. Copyright © 2007 John Wiley & Sons, Ltd. Diego Andrade, Manuel Arenaz, Basilio B. Fraguela, Juan Touriño, Ramón Doallo |
Concurr. Comput. Pract. Exp. | 1 |
| 2007 | Precise automatable analytical modeling of the cache behavior of codes with indirectionsabstractThe performance of memory hierarchies, in which caches play an essential role, is critical in nowadays general-purpose and embedded computing systems because of the growing memory bottleneck problem. Unfortunately, cache behavior is very unstable and difficult to predict. This is particularly true in the presence of irregular access patterns, which exhibit little locality. Such patterns are very common, for example, in applications in which pointers or compressed sparse matrices give place to indirections. Nevertheless, cache behavior in the presence of irregular access patterns has not been widely studied. In this paper we present an extension of a systematic analytical modeling technique based on PMEs (probabilistic miss equations), previously developed by the authors, that allows the automated analysis of the cache behavior for codes with irregular access patterns resulting from indirections. The model generates very accurate predictions despite the irregularities and has very low computing requirements, being the first model that gathers these desirable characteristics that can automatically analyze this kind of codes. These properties enable this model to help drive compiler optimizations, as we show with an example. Diego Andrade, Basilio B. Fraguela, Ramón Doallo |
ACM Trans. Archit. Code Optim. | 1 |
| 2006 | Analytical modeling of codes with arbitrary data-dependent conditional structures
Diego Andrade, Basilio B. Fraguela, Ramón Doallo |
J. Syst. Archit. | 1 |
| 2003 | Cache Behavior Modeling of Codes with Data-Dependent Conditionals
Diego Andrade, Basilio B. Fraguela, Ramón Doallo |
SCOPES | 1 |