EDBT 2026 Demo / reviewers in the wild / expert
Pablo Abad
dblp:42/5356 · also Pablo Abad Fidalgo
· DBLP profile ↗
21ranked-venue papers
13as first author
5since 2021 · last 2026
0000-0002-1262-1256ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 20 · 12 first-author · 5 since 2021Software engineering, systems software and programming languages · 2 · 2 first-authorArtificial intelligence and machine learning · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Design Tradeoffs in Backend Organization in Out-Of-Order RISC-V ProcessorsabstractThe primary role of the processor backend is to schedule instructions exposed by the frontend onto available execution resources while preserving sequential semantics. Optimizing this process is critical in modern superscalar out-of-order architectures, as it involves navigating complex tradeoffs in issue logic design and functional unit (FU) organization. Esther Alonso, Pablo Prieto, Pablo Abad, Valentin Puente |
SPAA | 3 |
| 2025 | Efficient Server Consolidation through a balanced mix of Transformer-based and Conventional ApplicationsabstractAlthough not optimal for Large Language Model (LLM) inference, general-purpose processors remain a practical choice due to their ubiquity and accessibility. This work explores strategies to maximize server utilization when incorporating such applications into the workload mix. To do that, we conduct an exhaustive profiling process to analyze hardware-software interactions, leading to two key observations. First, we have validated and quantified the common intuitions regarding LLM execution, with a specific focus on the microarchitecture's backend. Second, we observe the relatively low contention of LLMs and conventional applications (e.g., SPEC CPU17) when running together. Inspired by these findings, we explore whether combining applications of each type on a common server, if properly balanced, could lead to a better system utilization. Using both state-of-the-art server configurations and slightly older systems, we demonstrate that executing LLMs on general-purpose processors is feasible with minimal impact on co-located applications and lead to a better overall server performance. Pablo Abad, Pablo Prieto, Valentin Puente, José-Ángel Gregorio |
ICS | 1 |
| 2023 | Performance Characterization of Popular DNN Models on Out-of-Order CPUsabstractDNN popularity, which is driving advances in a growing number of fields, has increased the amount of computing resources running this kind of applications at an unprecedent rate. Specialized hardware, such as GPUs or ASIC-based accelerators, has been the preferred platform to run these applications. However, the ubiquity of DNN models is rapidly extending the presence of this software to general-purpose CPUs. For this reason, there is a pressing need to gain understanding of the main features of state-of-the-art DNN models to adapt CPU microarchitecture accordingly. In this paper we investigated a representative set of DNN models and, based on data collected from real hardware, we evaluated how efficiently they utilize the underlying system. We analyzed overall system performance, as well as the amount of vectorization provided by CPU-optimized frameworks. We quantified the performance loss caused by processor backend, and the contribution of memory hierarchy and functional units to it. We compared the backend utilization of DNN applications to popular benchmarks such as SPEC CPU2017 and found a lower balance in the use of the elements that make up the processor microarchitecture. Although many workloads seem to be constrained by functional unit availability, in a significant group of applications we found a non-negligible impact of memory hierarchy on performance. Pablo Prieto, Pablo Abad, José-Ángel Gregorio, Valentin Puente |
PACT | 2 |
| 2022 | Top-Down Performance Profiling on NVIDIA's GPUsabstractThe rise of data-intensive algorithms, such as Machine Learning ones, has meant a strong diversification of Graphics Processing Units (GPU) in fields with intensive Data-Level Parallelism. This trend, known as general-purpose computing on GPU (GP-GPU), makes the execution process on a GPU (seemingly simple in its architecture) far from trivial when targeting performance for many dissimilar applications. A proof of this is the existence of many profiling tools that help programmers to understand how to maximize hardware utilization. In contrast, this paper proposes a profiling tool focused on microarchitecture analysis under large sets of dissimilar applications. Therefore, the tool has a double objective. On the one hand, to check the suitability of a GPU for diverse sets of application kernels. On the other hand, to identify possible bottlenecks in a given GPU microarchitecture, facilitating the improvement of subsequent designs. For this purpose, using Top-Down methodology proposed by Intel for their CPUs as inspiration, we have defined a hierarchical organization for the execution pipeline of the GPU. The proposal makes use of the available hardware performance counters to identify how each component contributes to performance losses. We demonstrate the feasibility of the proposed methodology, analyzing how different modern NVIDIA architectures behave running relevant benchmarks, assessing in which microarchitecture component performance losses are the most significant. Alvaro Saiz, Pablo Prieto, Pablo Abad, José-Ángel Gregorio, Valentin Puente |
IPDPS | 3 |
| 2021 | Fast, Accurate Processor Evaluation Through Heterogeneous, Sample-Based BenchmarkingabstractPerformance evaluation is a key task in computing and communication systems. Benchmarking is one of the most common techniques for evaluation purposes, where the performance of a set of representative applications is used to infer system responsiveness in a general usage scenario. Unfortunately, most benchmarking suites are limited to a reduced number of applications, and in some cases, rigid execution configurations. This makes it hard to extrapolate performance metrics for a general-purpose architecture, supposed to have a multi-year lifecycle, running dissimilar applications concurrently. The main culprit of this situation is that current benchmark-derived metrics lack generality, statistical soundness and fail to represent general-purpose environments. Previous attempts to overcome these limitations through random app mixes significantly increase computational cost (workload population shoots up), making the evaluation process barely affordable. To circumvent this problem, in this article we present a more elaborate performance evaluation methodology named BenchCast. Our proposal provides more representative performance metrics, but with a drastic reduction of computational cost, limiting app execution to a small and representative fraction marked through code annotation. Thanks to this labeling and making use of synchronization techniques, we generate heterogeneous workloads where every app runs simultaneously inside its Region Of Interest, making a few execution seconds highly representative of full application execution. Pablo Prieto, Pablo Abad, José-Ángel Gregorio, Valentin Puente |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2020 | SPECcast: A Methodology for Fast Performance Evaluation with SPEC CPU 2017 Multiprogrammed WorkloadsabstractPerformance comparison is a key task in computer architecture research. These evaluations might need to consider the scenario where resources are shared concurrently by a wide range of application classes. In many cases, well known benchmarking tools, such as SpecCPU do not provide evaluation metrics under such usage circumstances. Previous attempts to fill the gap with realistic workloads have reiled on random combination of applications, formulating performance comparison as a statistical task to reduce the population size. The computational cost of these approaches is substantial, given the large mix required to achieve statistically meaningful results. In this paper, we present SPECcast, a methodology for the SPEC CPU2017 suite, which can circumvent this issue. The idea relies on exploiting the inner application characteristics to minimize the computational cost without degrading the statistical significance of the results. Using manual source-code annotation, we determine a small portion of each application, denoted Region of Interest (ROI), that accurately resembles the whole program's characteristics. Then, we develop synchronization mechanisms that can concurrently run any combination of applications in the cores of the system. This enables us to run multiprogrammed SPEC workloads ∼95% faster without losing statistical significance. Pablo Prieto, Pablo Abad, Jose Angel Herrero, José-Ángel Gregorio, Valentin Puente |
ICPP | 2 |
| 2019 | Architecting Racetrack Memory Preshift through Pattern-Based Prediction MechanismsabstractRacetrack Memories (RM) are a promising spintronic technology able to provide multi-bit storage in a single cell (tape-like) through a ferromagnetic nanowire with multiple domains. This technology offers superior density, non-volatility and low static power compared to CMOS memories. These features have attracted great interest in the adoption of RM as a replacement of RAM technology, from Main memory (DRAM) to maybe on-chip cache hierarchy (SRAM). One of the main drawbacks of this technology is the serialized access to the bits stored in each domain, resulting in unpredictable access time. An appropriate header management policy can potentially reduce the number of shift operations required to access the correct position. Simple policies such as leaving read/write head on the last domain accessed (or on the next) provide enough improvement in the presence of a certain level of locality on data access. However, in those cases with much lower locality, a more accurate behavior from the header management policy would be desirable. In this paper, we explore the utilization of hardware prefetching policies to implement the header management policy. “Predicting” the length and direction of the next displacement, it is possible to reduce shift operations, improving memory access time. The results of our experiments show that, with an appropriate header, our proposal reduces average shift latency by up to 50% in L2 and LLC, improving average memory access time by up to 10%. Adrian Colaso, Pablo Prieto, Pablo Abad, José-Ángel Gregorio, Valentin Puente |
IPDPS | 3 |
| 2018 | Memory Hierarchy Characterization of NoSQL Applications through Full-System SimulationabstractIn this work, we conduct a detailed memory characterization of a representative set of modern data-management software (Cassandra, MongoDB, OrientDB and Redis) running an illustrative NoSQL benchmark suite (YCSB). These applications are widely popular NoSQL databases with different data models and features such as in-memory storage. We compare how these data-serving applications behave with respect to other well-known benchmarks, such as SPEC CPU2006, PARSEC and NAS Parallel Benchmark. The methodology employed for evaluation relies on state-of-the-art full-system simulation tools, such as gem5. This allows us to explore configurations unattainable using performance monitoring units in actual hardware, being able to characterize memory properties. The results obtained suggest that NoSQL application behavior is not dissimilar to conventional workloads. Therefore, some of the optimizations present in state-of-the-art hardware might have a direct benefit. Nevertheless, there are some common aspects that are distinctive of conventional benchmarks that might be sufficiently relevant to be considered in architectural design. Strikingly, we also found that most database engines, independently of aspects such as workload or database size, exhibit highly uniform behavior. Finally, we show that different data-base engines make highly distinctive demands on the memory hierarchy, some being more stringent than others. Adrian Colaso, Pablo Prieto, Jose Angel Herrero, Pablo Abad, Lucia G. Menezo, Valentin Puente, José-Ángel Gregorio |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2016 | AC-WAR: Architecting the Cache Hierarchy to Improve the Lifetime of a Non-Volatile Endurance-Limited Main MemoryabstractThis work shows how by adapting replacement policies in contemporary cache hierarchies it is possible to extend the lifespan of a write endurance-limited main memory by almost one order of magnitude. The inception of this idea is that during cache residency 1) blocks are modified in a bimodal way: either most of the content of the block is modified or most of the content of the block never changes, and 2) in most applications, the majority of blocks are only slightly modified. When those facts are considered by the cache replacement algorithms, it is possible to significantly reduce the number of bit-flips per write-back to main memory. Our proposal favors the off-chip eviction of slightly modified blocks according to an adaptive replacement algorithm that operates coordinately in L2 and L3. This way it is possible to improve significantly system memory lifetime, with negligible performance degradation. We found that using a few bits per block to track changes in cache blocks with respect to the main memory content is enough. With a slightly modified sectored LRU and a simple cache performance predictor it is possible to achieve a simple implementation with minimal cost in area and no impact on cache access time. On average, our proposal increases the memory lifetime obtained with an LRU policy up to 10 times (10×) and 15 times (15×) when combined with other memory centric techniques. In both cases, the performance degradation could be considered negligible. Pablo Abad, Pablo Prieto, Valentin Puente, José-Ángel Gregorio |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2015 | Improving last level shared cache performance through mobile insertion policies (MIP)
Pablo Abad, Pablo Prieto, Valentin Puente, José-Ángel Gregorio |
Parallel Comput. | 1 |
| 2013 | Interaction of NoC Design and Coherence Protocol in 3D-Stacked CMPsabstractComputer architectures have evolved to structures where communication has become an essential part of the system and most of it currently takes place inside the chip. The number of on-Chip cores and the available off-chip bandwidth is not growing at the same rate. This demands for the inclusion of more sophisticated memory hierarchies inside the chip to deal with off-chip latency and bandwidth problems in order to keep on improving performance. The exhaustion of Moore's law will accelerate the use of 3D-Stacked on-chip memory hierarchies to sustain the required scalability of forthcoming CMPs. For this class of systems' memory hierarchy, coherence protocol and interconnection network are two closely related components, but which are usually designed independently. In this work we will demonstrate that network components can be coupled to coherence protocol in order to extract significant performance benefits. Making use of a well-known snoop coherence protocol, we will present different network optimizations, better able to adapt to the communication requirements of this protocol. Evaluation results show that with minimal hardware changes, for some real applications, full system performance can be improved by up to 48%. Pablo Abad, Pablo Prieto, Lucia G. Menezo, Adrian Colaso, Valentin Puente, José-Ángel Gregorio |
DSD | 1 |
| 2013 | Improving Test Generation under Rich Contracts by Tight Bounds and Incremental SAT SolvingabstractWe present a novel and general technique for automated test generation that combines tight bounds with incremental SAT solving. The proposed technique uses incremental SAT to build test suites targeting a specific testing criterion, amongst various black-box and white-box criteria. As our experimental results show, the combination of tight bounds with incremental SAT, and the testing criterion driven approach implemented in our prototype tool FAJITA, enable us to effectively generate test suites for container classes with rich contracts, more efficiently than other state-of-the-art tools. Pablo Abad, Nazareno Aguirre, Valeria S. Bengolea, Daniel Alfredo Ciolek, Marcelo F. Frias, Juan P. Galeotti, T. S. E. Maibaum, Mariano M. Moscato, Nicolás Rosner, Ignacio Vissani |
ICST | 1 |
| 2013 | LIGERO: A light but efficient router conceived for cache-coherent chip multiprocessorsabstractAlthough abstraction is the best approach to deal with computing system complexity, sometimes implementation details should be considered. Considering on-chip interconnection networks in particular, underestimating the underlying system specificity could have nonnegligible impact on performance, cost, or correctness. This article presents a very efficient router that has been devised to deal with cache-coherent chip multiprocessor particularities in a balanced way. Employing the same principles of packet rotation structures as in the rotary router, we present a router configuration with the following novel features: (1) reduced buffering requirements, (2) optimized pipeline under contentionless conditions, (3) more efficient deadlock avoidance mechanism, and (4) optimized in-order delivery guarantee. Putting it all together, our proposal provides a set of features that no other router, to the best of our knowledge, has achieved previously. These are: (1') low implementation cost, (2') low pass-through latency under low load, (3') improved resource utilization through adaptive routing and a buffering scheme free of head-of-line blocking, (4') guarantee of coherence protocol correctness via end-to-end deadlock avoidance and in-order delivery, and (5') improvement of coherence protocol responsiveness through adaptive in-network multicast support. We conduct a thorough evaluation that includes hardware cost estimation and performance evaluation under a wide spectrum of realistic workloads and coherence protocols. Comparing our proposal with VCTM, an optimized state-of-the-art wormhole router, it requires 50% less area, reduces on-chip cache hierarchy energy delay product on average by 20%, and improves the cache-coherency chip multiprocessor performance under realistic working conditions by up to 20%. Pablo Abad, Valentin Puente, José-Ángel Gregorio |
ACM Trans. Archit. Code Optim. | 1 |
| 2012 | BIXBAR: A low cost solution to support dynamic link reconfiguration in networks on chipabstractImproving link utilization is a key aspect in interconnection network design. Reconfigurable-direction interrouter links optimize network resource utilization, which substantially increases the maximum achievable throughput. In the case of On-chip Networks, the short distance between adjacent routers makes feasible fast link arbitration, which makes dynamic link reconfiguration an attractive solution. In this paper we propose a low-cost router micro-architecture that is able to deal with reconfigurable links with a marginal cost over a conventional router. The key element of the proposal is a bidirectional crossbar, which enables reconfiguration of links, without significantly increasing router area and energy. The results obtained indicate that with this proposal, system performance could be improved, for some selected workloads, by up to 25% while energy-performance tradeoff is reduced by 20%, avoiding the additional costs entailed in other state-of-the-art routers capable of performing dynamic link reconfiguration. Pablo Abad, Pablo Prieto, Valentin Puente, José-Ángel Gregorio |
ICCD | 1 |
| 2012 | TOPAZ: An Open-Source Interconnection Network Simulator for Chip Multiprocessors and SupercomputersabstractAs in other computer architecture areas, interconnection networks research relies most of the times on simulation tools. This paper announces the release of an open-source tool suitable to be used for accurate modeling from small CMP to large supercomputer interconnection networks. The cycle-accurate modeling of TOPAZ can be used standalone through synthetic traffic patterns and application-traces or within full-system evaluation systems such as GEMS or GEM5 effortlessly. In fact, we provide an advanced interface that enables the replacement of the original lightweight but optimistic GEMS and GEM5 network simulator with limited performance impact on the simulation time. Our tests indicate that in this context, underestimating network modeling could induce up to 50% error in the performance estimation of the simulated system. To minimize the impact of detailed network modeling on simulation time, we incorporate mechanisms able to attenuate the higher computational effort, reducing in this way the slowdown of the full system simulation with accurate performance estimations. Additionally, in order to evaluate large-scale networks, we parallelize the simulator to be able to optimize memory resources with the growing number of cores available per chip in the simulation farms. This allows us to simulate node networks exceeding one million of routers with up to 70% efficiency in a multithreaded simulation running on twelve cores. Pablo Abad, Pablo Prieto, Lucia G. Menezo, Adrian Colaso, Valentin Puente, José-Ángel Gregorio |
NOCS | 1 |
| 2012 | Balancing Performance and Cost in CMP Interconnection NetworksabstractThis paper presents an innovative router design, called Rotary Router, which successfully addresses CMP cost/performance constraints. The router structure is based on two independent rings, which force packets to circulate either clockwise or counterclockwise, traveling through every port of the router. These two rings constitute a completely decentralized arbitration scheme that enables a simple, but efficient way to connect every input port to every output port. The proposed router is able to avoid network deadlock, livelock, and starvation without requiring data-path modifications. The organization of the router permits the inclusion of throughput enhancement techniques without significantly penalizing the implementation cost. In particular, the router performs adaptive routing, eliminates HOL blocking, and carries out implicit congestion control using simple arbitration and buffering strategies. Additionally, the proposal is capable of avoiding end-to-end deadlock at coherence protocol level with no physical or virtual resource replication, while guaranteeing in-order packet delivery. This facilitates router management and improves storage utilization. Using a comprehensive evaluation framework that includes full-system simulation and hardware description, the proposal is compared with two representative router counterparts. The results obtained demonstrate the Rotary Router's substantial performance and efficiency advantages. Pablo Abad, Valentin Puente, José-Ángel Gregorio |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2012 | Adaptive-Tree Multicast: Efficient Multidestination Support for CMP Communication SubstrateabstractMultidestination communications are a highly necessary capability for many coherence protocols in order to minimize on-chip hit latency. Although CMPs share this necessity, up to now few suitable proposals have been developed. The combination of resource scarcity and the common idea that multicast support requires a substantial amount of extra resources is responsible for this situation. In this work, we propose a new approach suitable for on-chip networks capable of managing multidestination traffic via hardware in an efficient way with negligible complexity. We introduce a novel multicast routing mechanism, able to circumvent many of the limitations of conventional multicast schemes. Adaptive-tree multicasting is able to maintain correctness for multiflit multicast messages without routing restrictions, while also coupling correctness and performance in a natural way. Replication restrictions not only guarantee the presence of enough resources to avoid deadlock, but also dynamically adapt tree shape to network conditions, routing multicast messages through noncongested paths. The performance results, using a state-of-the-art full system simulation framework, show that it improves the average full system performance of a CMP by 20 percent and network ED2P by 15 percent, when compared to a state-of-the-art router with conventional multicast support and similar implementation cost. Pablo Abad, Valentin Puente, Lucia G. Menezo, José-Ángel Gregorio |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2009 | MRR: Enabling fully adaptive multicast routing for CMP interconnection networksabstractOn-network hardware support for multi-destination traffic is a desirable feature in most multiprocessor machines. Multicast hardware capabilities enable much more effective bandwidth utilization as multi-destination packets do not need to repeatedly use the same resources, as occurs when multicast traffic must be decomposed in unicast packets. Although Chip Multiprocessors are not an exception in this interest, up to date, few fitting proposals exist. The combination of the scarcity of available resources and the common idea that multicast support requires a substantial amount of extra resources is responsible for this situation. In this work, we propose a new approach suitable for on-chip networks capable of managing multi-destination traffic via hardware in an efficient way with negligible complexity. We introduce the Multicast Rotary Router (MRR), a router able to: (1) perform on-network multicast support with almost zero cost over the Rotary Router, (2) use a fully adaptive tree to distribute multicast traffic, (3) perform on-network congestion control extending network utilization range. The performance results, using a state-of-the-art full system simulation framework, show that it improves average full system performance of a CMP using a unicast Rotary Router in its interconnection network by 25%, and an input buffered router with multicast support by 20%. Pablo Abad, Valentin Puente, José-Ángel Gregorio |
HPCA | 1 |
| 2009 | Improving topological maps for safer and robust navigationabstractNowadays we frequently find big amounts of data to work with, what facilitates many robotic tasks and helps to solve perception problems. At the same time, this fact origins an interesting ongoing research problem: how to organize and arrange big sets of information to be useful in later uses. Topological mapping is a very useful tool to arrange and deal with big amounts of reference images for robotic tasks. There are many previous works on topological mapping and many others use this kind of maps for topological localization, planning and navigation. This work is focused on the problem of carefully design topological map building processes that facilitate the posterior robot tasks that use them and make them safer. We propose a new hierarchy of topological maps focused on this aspect. The experiments included in this paper were run outdoors using omnidirectional images and GPS information, and show the good topological maps obtained and how they allow robust and safer localization and navigation tasks. Ana Cristina Murillo, Pablo Abad, Josechu J. Guerrero, Carlos Sagüés |
IROS | 2 |
| 2008 | Reducing the Interconnection Network Cost of Chip Multiprocessors
Pablo Abad, Valentin Puente, José-Ángel Gregorio |
NOCS | 1 |
| 2007 | Rotary router: an efficient architecture for CMP interconnection networksabstractThe trend towards increasing the number of processor cores and cache capacity in future Chip-Multiprocessors (CMPs), will require scalable packet-switched interconnection networks adapted to the restrictions imposed by the CMP environment. This paper presents an innovative router design, which successfully addresses CMP cost/performance constraints. The router structure is based on two independent rings, which force packets to circulate either clockwise or anti-clockwise, traveling through every port of the router. It uses a completely decentralized scheduling scheme, which allows the design to: (1) take advantage of wide links, (2) reduce Head of Line blocking, (3) use adaptive routing, (4) be topology agnostic, (5) scale with network degree, and (6) have reasonable power consumption and implementation cost. A thorough comparative performance analysis against competitive conventional routers shows an advantage for our proposal of up to 50 % in terms of raw performance and nearly 60 % in terms of energy-delay product. Pablo Abad, Valentin Puente, José-Ángel Gregorio, Pablo Prieto |
ISCA | 1 |