Pablo Prieto

dblp:34/4315 · DBLP profile ↗
← Back
20ranked-venue papers
5as first author
6since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 19 · 5 first-author · 5 since 2021Software engineering, systems software and programming languages · 2
YearPublicationVenuePosition
2026 Design Tradeoffs in Backend Organization in Out-Of-Order RISC-V Processors
abstract
The primary role of the processor backend is to schedule instructions exposed by the frontend onto available execution resources while preserving sequential semantics. Optimizing this process is critical in modern superscalar out-of-order architectures, as it involves navigating complex tradeoffs in issue logic design and functional unit (FU) organization.
Esther Alonso, Pablo Prieto, Pablo Abad, Valentin Puente
SPAA2
2025 Efficient Server Consolidation through a balanced mix of Transformer-based and Conventional Applications
abstract
Although not optimal for Large Language Model (LLM) inference, general-purpose processors remain a practical choice due to their ubiquity and accessibility. This work explores strategies to maximize server utilization when incorporating such applications into the workload mix. To do that, we conduct an exhaustive profiling process to analyze hardware-software interactions, leading to two key observations. First, we have validated and quantified the common intuitions regarding LLM execution, with a specific focus on the microarchitecture's backend. Second, we observe the relatively low contention of LLMs and conventional applications (e.g., SPEC CPU17) when running together. Inspired by these findings, we explore whether combining applications of each type on a common server, if properly balanced, could lead to a better system utilization. Using both state-of-the-art server configurations and slightly older systems, we demonstrate that executing LLMs on general-purpose processors is feasible with minimal impact on co-located applications and lead to a better overall server performance.
Pablo Abad, Pablo Prieto, Valentin Puente, José-Ángel Gregorio
ICS2
2023 Performance Characterization of Popular DNN Models on Out-of-Order CPUs
abstract
DNN popularity, which is driving advances in a growing number of fields, has increased the amount of computing resources running this kind of applications at an unprecedent rate. Specialized hardware, such as GPUs or ASIC-based accelerators, has been the preferred platform to run these applications. However, the ubiquity of DNN models is rapidly extending the presence of this software to general-purpose CPUs. For this reason, there is a pressing need to gain understanding of the main features of state-of-the-art DNN models to adapt CPU microarchitecture accordingly. In this paper we investigated a representative set of DNN models and, based on data collected from real hardware, we evaluated how efficiently they utilize the underlying system. We analyzed overall system performance, as well as the amount of vectorization provided by CPU-optimized frameworks. We quantified the performance loss caused by processor backend, and the contribution of memory hierarchy and functional units to it. We compared the backend utilization of DNN applications to popular benchmarks such as SPEC CPU2017 and found a lower balance in the use of the elements that make up the processor microarchitecture. Although many workloads seem to be constrained by functional unit availability, in a significant group of applications we found a non-negligible impact of memory hierarchy on performance.
Pablo Prieto, Pablo Abad, José-Ángel Gregorio, Valentin Puente
PACT1
2023 Digital twin based simulation platform for heavy duty hybrid electric vehicles
abstract
The reduction of Greenhouse Gases is of great interest inside the industry and, especially, the road transport sector. Heavy-duty vehicles have been extensively researched due to their significant contribution. This work presents a digital twin (LONGRUN simulation platform) to analyse different heavy-duty vehicle aerodynamic designs and powertrain topologies. Based on a forward-looking formulation, this work allows the analysis of novel hybridisation and electrification control strategies impact. We validated our platform components against a commercial tool (VECTO) adopted by the European Commission for combustion powertrain analysis. In this article, we provide a modular platform for heavy-duty vehicles, in a widely used software inside the automotive industry. Furthermore, our platform offers the possibility to introduce customised control strategies for hybrid-electric vehicles. This work analyses the impact of parallel and serial hybrid powertrain topologies.
Eneko Otaola, Beñat Arteta, Joshué Pérez, Andres Sierra-Gonzalez, Pablo Prieto
VTC2023-Spring5
2022 Top-Down Performance Profiling on NVIDIA's GPUs
abstract
The rise of data-intensive algorithms, such as Machine Learning ones, has meant a strong diversification of Graphics Processing Units (GPU) in fields with intensive Data-Level Parallelism. This trend, known as general-purpose computing on GPU (GP-GPU), makes the execution process on a GPU (seemingly simple in its architecture) far from trivial when targeting performance for many dissimilar applications. A proof of this is the existence of many profiling tools that help programmers to understand how to maximize hardware utilization. In contrast, this paper proposes a profiling tool focused on microarchitecture analysis under large sets of dissimilar applications. Therefore, the tool has a double objective. On the one hand, to check the suitability of a GPU for diverse sets of application kernels. On the other hand, to identify possible bottlenecks in a given GPU microarchitecture, facilitating the improvement of subsequent designs. For this purpose, using Top-Down methodology proposed by Intel for their CPUs as inspiration, we have defined a hierarchical organization for the execution pipeline of the GPU. The proposal makes use of the available hardware performance counters to identify how each component contributes to performance losses. We demonstrate the feasibility of the proposed methodology, analyzing how different modern NVIDIA architectures behave running relevant benchmarks, assessing in which microarchitecture component performance losses are the most significant.
Alvaro Saiz, Pablo Prieto, Pablo Abad, José-Ángel Gregorio, Valentin Puente
IPDPS2
2021 Fast, Accurate Processor Evaluation Through Heterogeneous, Sample-Based Benchmarking
abstract
Performance evaluation is a key task in computing and communication systems. Benchmarking is one of the most common techniques for evaluation purposes, where the performance of a set of representative applications is used to infer system responsiveness in a general usage scenario. Unfortunately, most benchmarking suites are limited to a reduced number of applications, and in some cases, rigid execution configurations. This makes it hard to extrapolate performance metrics for a general-purpose architecture, supposed to have a multi-year lifecycle, running dissimilar applications concurrently. The main culprit of this situation is that current benchmark-derived metrics lack generality, statistical soundness and fail to represent general-purpose environments. Previous attempts to overcome these limitations through random app mixes significantly increase computational cost (workload population shoots up), making the evaluation process barely affordable. To circumvent this problem, in this article we present a more elaborate performance evaluation methodology named BenchCast. Our proposal provides more representative performance metrics, but with a drastic reduction of computational cost, limiting app execution to a small and representative fraction marked through code annotation. Thanks to this labeling and making use of synchronization techniques, we generate heterogeneous workloads where every app runs simultaneously inside its Region Of Interest, making a few execution seconds highly representative of full application execution.
Pablo Prieto, Pablo Abad, José-Ángel Gregorio, Valentin Puente
IEEE Trans. Parallel Distributed Syst.1
2020 SPECcast: A Methodology for Fast Performance Evaluation with SPEC CPU 2017 Multiprogrammed Workloads
abstract
Performance comparison is a key task in computer architecture research. These evaluations might need to consider the scenario where resources are shared concurrently by a wide range of application classes. In many cases, well known benchmarking tools, such as SpecCPU do not provide evaluation metrics under such usage circumstances. Previous attempts to fill the gap with realistic workloads have reiled on random combination of applications, formulating performance comparison as a statistical task to reduce the population size. The computational cost of these approaches is substantial, given the large mix required to achieve statistically meaningful results. In this paper, we present SPECcast, a methodology for the SPEC CPU2017 suite, which can circumvent this issue. The idea relies on exploiting the inner application characteristics to minimize the computational cost without degrading the statistical significance of the results. Using manual source-code annotation, we determine a small portion of each application, denoted Region of Interest (ROI), that accurately resembles the whole program's characteristics. Then, we develop synchronization mechanisms that can concurrently run any combination of applications in the cores of the system. This enables us to run multiprogrammed SPEC workloads ∼95% faster without losing statistical significance.
Pablo Prieto, Pablo Abad, Jose Angel Herrero, José-Ángel Gregorio, Valentin Puente
ICPP1
2019 Modelling and Validation of Full Vehicle Model based on a Novel Multibody Formulation
abstract
Nowadays, the growing functionalities implemented on vehicles make the simulation phase much more important in the design process. For that purpose, a representative model is required, as it allows to reproduce the exact behaviour of the vehicle, and reduce not only the time required for its setup and testing, but also the cost related to these. Due to this, the development of accurate vehicle models has become one of the main areas of interest for the automotive industry. In this work a 16 DOF (degree of freedom) full vehicle model is presented. This model is based on multibody formulation combined with an appropriate solver for real-time execution. In order to validate this model, data from a real test vehicle has been used, comparing the real dynamic response of the vehicle to the one provided by the developed dynamic model. Results show that the presented approach represents effectively the behaviour of a real vehicle, both in longitudinal and lateral terms.
Alberto Parra, Dionisio Cagigas, Asier Zubizarreta-Pico, Antonio Joaquín Rodríguez, Pablo Prieto
IECON5
2019 Geo-Fence Based Route Tracking Diagnosis Strategy for Energy Prediction Strategies Applied to EV
abstract
Nowadays, the shortage of energy and environmental pollution are considered as relevant problems due to the high amount of traditional automotive vehicles with internal combustion engines (ICEs). Electric vehicle (EV) is one of the solutions to localize the energy source and the best choice for saving energy and provide zero emission vehicles. However, their main drawback when compared to conventional vehicles is their limited energy storage capacity, resulting in poor driving ranges. In order to mitigate this issue, the scientific community is extensively researching on energy optimization and prediction strategies to extend the autonomy of EV. In general, such strategies require the knowledge of the route profile, being of capital importance to identify whether the vehicle is on route or not. Considering this, in this paper, a route tracking diagnosis strategy is proposed and tested. The proposed strategy relies on the information provided by the Google Maps API (Application Programming Interface) to calculate the vehicles reference route. Additionally, a Global Positioning System (GPS) device is used to monitor the real vehicle position. The proposed strategy is validated throughout simulation, Driver in the Loop (DiL) test and experimental tests.
Pablo Prieto, Elena Trancho, Beñat Arteta, Alberto Parra, Alvaro Coupeau, Dionisio Cagigas, Edorta Ibarra
IECON1
2019 Architecting Racetrack Memory Preshift through Pattern-Based Prediction Mechanisms
abstract
Racetrack Memories (RM) are a promising spintronic technology able to provide multi-bit storage in a single cell (tape-like) through a ferromagnetic nanowire with multiple domains. This technology offers superior density, non-volatility and low static power compared to CMOS memories. These features have attracted great interest in the adoption of RM as a replacement of RAM technology, from Main memory (DRAM) to maybe on-chip cache hierarchy (SRAM). One of the main drawbacks of this technology is the serialized access to the bits stored in each domain, resulting in unpredictable access time. An appropriate header management policy can potentially reduce the number of shift operations required to access the correct position. Simple policies such as leaving read/write head on the last domain accessed (or on the next) provide enough improvement in the presence of a certain level of locality on data access. However, in those cases with much lower locality, a more accurate behavior from the header management policy would be desirable. In this paper, we explore the utilization of hardware prefetching policies to implement the header management policy. “Predicting” the length and direction of the next displacement, it is possible to reduce shift operations, improving memory access time. The results of our experiments show that, with an appropriate header, our proposal reduces average shift latency by up to 50% in L2 and LLC, improving average memory access time by up to 10%.
Adrian Colaso, Pablo Prieto, Pablo Abad, José-Ángel Gregorio, Valentin Puente
IPDPS2
2018 Memory Hierarchy Characterization of NoSQL Applications through Full-System Simulation
abstract
In this work, we conduct a detailed memory characterization of a representative set of modern data-management software (Cassandra, MongoDB, OrientDB and Redis) running an illustrative NoSQL benchmark suite (YCSB). These applications are widely popular NoSQL databases with different data models and features such as in-memory storage. We compare how these data-serving applications behave with respect to other well-known benchmarks, such as SPEC CPU2006, PARSEC and NAS Parallel Benchmark. The methodology employed for evaluation relies on state-of-the-art full-system simulation tools, such as gem5. This allows us to explore configurations unattainable using performance monitoring units in actual hardware, being able to characterize memory properties. The results obtained suggest that NoSQL application behavior is not dissimilar to conventional workloads. Therefore, some of the optimizations present in state-of-the-art hardware might have a direct benefit. Nevertheless, there are some common aspects that are distinctive of conventional benchmarks that might be sufficiently relevant to be considered in architectural design. Strikingly, we also found that most database engines, independently of aspects such as workload or database size, exhibit highly uniform behavior. Finally, we show that different data-base engines make highly distinctive demands on the memory hierarchy, some being more stringent than others.
Adrian Colaso, Pablo Prieto, Jose Angel Herrero, Pablo Abad, Lucia G. Menezo, Valentin Puente, José-Ángel Gregorio
IEEE Trans. Parallel Distributed Syst.2
2016 Energy minimization at all layers of the data center: The ParaDIME project
Oscar Palomar, Santhosh Kumar Rethinagiri, Gulay Yalcin, J. Rubén Titos Gil, Pablo Prieto, Emma Torrella, Osman S. Unsal, Adrián Cristal, Pascal Felber, Anita Sobe, Yaroslav Hayduk, Mascha Kurpicz, Christof Fetzer, Thomas Knauth, Malte Schneegaß, Jens Struckmeier, Dragomir Milojevic
DATE5
2016 AC-WAR: Architecting the Cache Hierarchy to Improve the Lifetime of a Non-Volatile Endurance-Limited Main Memory
abstract
This work shows how by adapting replacement policies in contemporary cache hierarchies it is possible to extend the lifespan of a write endurance-limited main memory by almost one order of magnitude. The inception of this idea is that during cache residency 1) blocks are modified in a bimodal way: either most of the content of the block is modified or most of the content of the block never changes, and 2) in most applications, the majority of blocks are only slightly modified. When those facts are considered by the cache replacement algorithms, it is possible to significantly reduce the number of bit-flips per write-back to main memory. Our proposal favors the off-chip eviction of slightly modified blocks according to an adaptive replacement algorithm that operates coordinately in L2 and L3. This way it is possible to improve significantly system memory lifetime, with negligible performance degradation. We found that using a few bits per block to track changes in cache blocks with respect to the main memory content is enough. With a slightly modified sectored LRU and a simple cache performance predictor it is possible to achieve a simple implementation with minimal cost in area and no impact on cache access time. On average, our proposal increases the memory lifetime obtained with an LRU policy up to 10 times (10×) and 15 times (15×) when combined with other memory centric techniques. In both cases, the performance degradation could be considered negligible.
Pablo Abad, Pablo Prieto, Valentin Puente, José-Ángel Gregorio
IEEE Trans. Parallel Distributed Syst.2
2015 Improving last level shared cache performance through mobile insertion policies (MIP)
Pablo Abad, Pablo Prieto, Valentin Puente, José-Ángel Gregorio
Parallel Comput.2
2013 Interaction of NoC Design and Coherence Protocol in 3D-Stacked CMPs
abstract
Computer architectures have evolved to structures where communication has become an essential part of the system and most of it currently takes place inside the chip. The number of on-Chip cores and the available off-chip bandwidth is not growing at the same rate. This demands for the inclusion of more sophisticated memory hierarchies inside the chip to deal with off-chip latency and bandwidth problems in order to keep on improving performance. The exhaustion of Moore's law will accelerate the use of 3D-Stacked on-chip memory hierarchies to sustain the required scalability of forthcoming CMPs. For this class of systems' memory hierarchy, coherence protocol and interconnection network are two closely related components, but which are usually designed independently. In this work we will demonstrate that network components can be coupled to coherence protocol in order to extract significant performance benefits. Making use of a well-known snoop coherence protocol, we will present different network optimizations, better able to adapt to the communication requirements of this protocol. Evaluation results show that with minimal hardware changes, for some real applications, full system performance can be improved by up to 48%.
Pablo Abad, Pablo Prieto, Lucia G. Menezo, Adrian Colaso, Valentin Puente, José-Ángel Gregorio
DSD2
2013 CMP off-chip bandwidth scheduling guided by instruction criticality
abstract
This paper explores the benefits of scheduling off-chip memory operations in a Chip Multiprocessor (CMP) according to their execution relevance. Assuming the scenario of having many out-of-order execution cores in the CMP, from the processor perspective, the importance of the instruction that triggers an access to off-chip memory may vary considerably. Consequently, it makes sense to consider this point of view at the memory controller level to reorder outgoing memory accesses. After exploring different processor-centric sorting criteria, we reach the conclusion that the most simple and useful metric for scheduling a memory operation is the position in the reorder buffer of the instruction that triggers the on-chip miss. We propose a simple memory controller scheduling policy that employs this information as its main parameter. This proposal significantly improves system responsiveness, both in terms of throughput and fairness. The idea is analyzed through full-system simulation, running a broad set of workloads with diverse memory behavior. When it is compared with other scheduling algorithms with similar complexity, throughput can be improved by an average of 10% and fairness enhanced by an average of 15% even in very adverse usage scenarios. Moreover, the idea supports the possibility of dynamically favoring throughput or fairness, according to the end-user requirements.
Pablo Prieto, Valentin Puente, José-Ángel Gregorio
ICS1
2012 BIXBAR: A low cost solution to support dynamic link reconfiguration in networks on chip
abstract
Improving link utilization is a key aspect in interconnection network design. Reconfigurable-direction interrouter links optimize network resource utilization, which substantially increases the maximum achievable throughput. In the case of On-chip Networks, the short distance between adjacent routers makes feasible fast link arbitration, which makes dynamic link reconfiguration an attractive solution. In this paper we propose a low-cost router micro-architecture that is able to deal with reconfigurable links with a marginal cost over a conventional router. The key element of the proposal is a bidirectional crossbar, which enables reconfiguration of links, without significantly increasing router area and energy. The results obtained indicate that with this proposal, system performance could be improved, for some selected workloads, by up to 25% while energy-performance tradeoff is reduced by 20%, avoiding the additional costs entailed in other state-of-the-art routers capable of performing dynamic link reconfiguration.
Pablo Abad, Pablo Prieto, Valentin Puente, José-Ángel Gregorio
ICCD2
2012 TOPAZ: An Open-Source Interconnection Network Simulator for Chip Multiprocessors and Supercomputers
abstract
As in other computer architecture areas, interconnection networks research relies most of the times on simulation tools. This paper announces the release of an open-source tool suitable to be used for accurate modeling from small CMP to large supercomputer interconnection networks. The cycle-accurate modeling of TOPAZ can be used standalone through synthetic traffic patterns and application-traces or within full-system evaluation systems such as GEMS or GEM5 effortlessly. In fact, we provide an advanced interface that enables the replacement of the original lightweight but optimistic GEMS and GEM5 network simulator with limited performance impact on the simulation time. Our tests indicate that in this context, underestimating network modeling could induce up to 50% error in the performance estimation of the simulated system. To minimize the impact of detailed network modeling on simulation time, we incorporate mechanisms able to attenuate the higher computational effort, reducing in this way the slowdown of the full system simulation with accurate performance estimations. Additionally, in order to evaluate large-scale networks, we parallelize the simulator to be able to optimize memory resources with the growing number of cores available per chip in the simulation farms. This allows us to simulate node networks exceeding one million of routers with up to 70% efficiency in a multithreaded simulation running on twelve cores.
Pablo Abad, Pablo Prieto, Lucia G. Menezo, Adrian Colaso, Valentin Puente, José-Ángel Gregorio
NOCS2
2010 Design and implementation of a direct RF-to-digital UHF-TV multichannel transceiver
abstract
This manuscript presents the design and implementation of a direct UHF digital transceiver that provides direct sampling of the UHF-TV input signal spectrum. Our SDR-based approach is based on a pair of high-speed ADC/DAC devices along with a channelizer, a dechannelizer and a channel management unit, which is capable of inserting, deleting and/or conmuting individual RF-TV channels. Simulation results and in-lab measurements assess that the proposed system is able to receive, manipulate and retransmit a real UHF-TV spectrum at a negligible quality penalty cost.
Mikel Sánchez, Javier Del Ser, Pablo Prieto
ISCAS3
2007 Rotary router: an efficient architecture for CMP interconnection networks
abstract
The trend towards increasing the number of processor cores and cache capacity in future Chip-Multiprocessors (CMPs), will require scalable packet-switched interconnection networks adapted to the restrictions imposed by the CMP environment. This paper presents an innovative router design, which successfully addresses CMP cost/performance constraints. The router structure is based on two independent rings, which force packets to circulate either clockwise or anti-clockwise, traveling through every port of the router. It uses a completely decentralized scheduling scheme, which allows the design to: (1) take advantage of wide links, (2) reduce Head of Line blocking, (3) use adaptive routing, (4) be topology agnostic, (5) scale with network degree, and (6) have reasonable power consumption and implementation cost. A thorough comparative performance analysis against competitive conventional routers shows an advantage for our proposal of up to 50 % in terms of raw performance and nearly 60 % in terms of energy-delay product.
Pablo Abad, Valentin Puente, José-Ángel Gregorio, Pablo Prieto
ISCA4