Pedro Marcuello

dblp:85/1247 · DBLP profile ↗
← Back
15ranked-venue papers
6as first author
1since 2021 · last 2021
0000-0001-6104-9105ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 14 · 6 first-author · 1 since 2021Software engineering, systems software and programming languages · 3 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
10 papers
GPUs and heterogeneous computing · 37% Processor architecture and microarchitecture · 23% Energy-efficient computing · 18%
Computer graphics and multimedia
2 papers
Rendering · 100%

Topics — the 24 heaviest of 29, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
GPUs and heterogeneous computing
GPU architecture
0.412019
Visibility Rendering Order: Improving Energy Efficiency on Mobile GPUs through Frame Coherence · IEEE Trans. Parallel Distributed Syst. 2019
GPUs and heterogeneous computing › GPU rendering
GPU graphics pipeline
0.412019
Rendering Elimination: Early Discard of Redundant Tiles in the Graphics Pipeline · HPCA 2019
Memory systems › memory bandwidth management
memory bandwidth reduction
0.412019
Rendering Elimination: Early Discard of Redundant Tiles in the Graphics Pipeline · HPCA 2019
GPUs and heterogeneous computing › GPU rendering
mobile GPU rendering
0.412019
Visibility Rendering Order: Improving Energy Efficiency on Mobile GPUs through Frame Coherence · IEEE Trans. Parallel Distributed Syst. 2019
Embedded and real-time systems
collision detection
0.212015
Ultra-low power render-based collision detection for CPU/GPU systems · MICRO 2015
GPUs and heterogeneous computing › CPU-GPU heterogeneous computing
CPU-GPU heterogeneous systems
0.212015
Ultra-low power render-based collision detection for CPU/GPU systems · MICRO 2015
Energy-efficient computing › power management
mobile device energy
0.212015
Ultra-low power render-based collision detection for CPU/GPU systems · MICRO 2015
Energy-efficient computing
power management
0.212015
Ultra-low power render-based collision detection for CPU/GPU systems · MICRO 2015
Processor architecture and microarchitecture › multithreading
speculative multithreading
0.132008
Mitosis: A Speculative Multithreaded Processor Based on Precomputation Slices · IEEE Trans. Parallel Distributed Syst. 2008
Thread-Spawning Schemes for Speculative Multithreading · HPCA 2002
Value Prediction for Speculative Multithreaded Architectures · MICRO 1999
Processor architecture and microarchitecture
instruction-level parallelism
0.122011
Fg-STP: Fine-Grain Single Thread Partitioning on Multicores · HPCA 2011
Value Prediction for Speculative Multithreaded Architectures · MICRO 1999
Processor architecture and microarchitecture
chip multiprocessor
0.112011
Fg-STP: Fine-Grain Single Thread Partitioning on Multicores · HPCA 2011
Processor architecture and microarchitecture
thread partitioning
0.112011
Fg-STP: Fine-Grain Single Thread Partitioning on Multicores · HPCA 2011
Energy-efficient computing › energy-efficient architecture
GPU energy reduction
0.112019
Rendering Elimination: Early Discard of Redundant Tiles in the Graphics Pipeline · HPCA 2019
Parallel and multicore computing › speculative parallelization
thread-level speculation
0.122004
Thread Partitioning and Value Prediction for Exploiting Speculative Thread-Level Parallelism · IEEE Trans. Computers 2004
A framework for modeling and optimization of prescient instruction prefetch · SIGMETRICS 2003
Processor architecture and microarchitecture
speculative execution
0.122005
Mitosis compiler: an infrastructure for speculative threading based on pre-computation slices · PLDI 2005
Value Prediction for Speculative Multithreaded Architectures · MICRO 1999
Processor architecture and microarchitecture
value prediction
0.122004
Thread Partitioning and Value Prediction for Exploiting Speculative Thread-Level Parallelism · IEEE Trans. Computers 2004
Value Prediction for Speculative Multithreaded Architectures · MICRO 1999
Parallel and multicore computing
thread-level parallelism
0.132004
Thread-Spawning Schemes for Speculative Multithreading · HPCA 2002
Thread Partitioning and Value Prediction for Exploiting Speculative Thread-Level Parallelism · IEEE Trans. Computers 2004
Value Prediction for Speculative Multithreaded Architectures · MICRO 1999
Compilers and program optimization › parallelization
thread-level speculation
0.112005
Mitosis compiler: an infrastructure for speculative threading based on pre-computation slices · PLDI 2005
Processor architecture and microarchitecture › multithreading
helper threads
0.012003
A framework for modeling and optimization of prescient instruction prefetch · SIGMETRICS 2003
Processor architecture and microarchitecture › instruction fetch
instruction prefetching
0.012003
A framework for modeling and optimization of prescient instruction prefetch · SIGMETRICS 2003
Processor architecture and microarchitecture
multithreading
0.011999
Value Prediction for Speculative Multithreaded Architectures · MICRO 1999
Performance modeling and evaluation › workload characterization › program behavior
program behavior modeling
0.012003
A framework for modeling and optimization of prescient instruction prefetch · SIGMETRICS 2003
Performance modeling and evaluation
workload characterization
0.012003
A framework for modeling and optimization of prescient instruction prefetch · SIGMETRICS 2003
Compilers and program optimization › program transformation
program partitioning
0.012002
Thread-Spawning Schemes for Speculative Multithreading · HPCA 2002

Methods — techniques the papers use, named apart from their topics

temporal coherence exploitation · 0.8early-depth test · 0.8render-based collision detection · 0.4signature comparison · 0.4speculation · 0.2hardware-software co-design · 0.2instruction replication · 0.1dependence speculation · 0.1pre-computation slices · 0.1control speculation · 0.0profile-based analysis · 0.0heuristics · 0.0
YearPublicationVenuePosition
2021 Mont-Blanc 2020: Towards Scalable and Power Efficient European HPC Processors
abstract
The Mont-Blanc 2020 (MB2020) project has triggered the development of the next generation industrial processor for Big Data and High Performance Computing (HPC). MB2020 is paving the way to the future low-power European processor for exascale, defining the System-on-Chip (SoC) architecture and implementing new critical building blocks to be integrated in such an SoC. In this paper, we first present an overview of the MB2020 project, then we describe our experimental infrastructure, the requirements of relevant applications, and the IP blocks developed in the project. Finally, we present our emulation-based final demonstrator and explain how it integrates within our first generation of HPC processors.
Adrià Armejach, Bine Brank, Jordi Cortina, François Dolique, Timothy Hayes 0001, Nam Ho, Pierre-Axel Lagadec, Romain Lemaire, Guillem López-Paradís, Laurent Marliac, Miquel Moretó, Pedro Marcuello, Dirk Pleiter, Xubin Tan, Said Derradji
DATE12
2019 Rendering Elimination: Early Discard of Redundant Tiles in the Graphics Pipeline
abstract
GPUs are one of the most energy-consuming components for real-time rendering applications, since a large number of fragment shading computations and memory accesses are involved. Main memory bandwidth is especially taxing battery-operated devices such as smart-phones. TileBased Rendering GPUs divide the screen space into multiple tiles that are independently rendered in on-chip buffers, thus reducing memory bandwidth and energy consumption. We have observed that, in many animated graphics workloads, a large number of screen tiles have the same color across adjacent frames. In this paper, we propose Rendering Elimination (RE), a novel micro-architectural technique that accurately determines if a tile will be identical to the same tile in the preceding frame before rasterization by means of comparing signatures. Since RE identifies redundant tiles early in the graphics pipeline, it completely avoids the computation and memory accesses of the most power consuming stages of the pipeline, which substantially reduces the execution time and the energy consumption of the GPU. For widely used Android applications, we show that RE achieves an average speedup of 1.74x and energy reduction of 43% for the GPU/Memory system, surpassing by far the benefits of Transaction Elimination, a state-of-the-art memory bandwidth reduction technique available in some commercial Tile-Based Rendering GPUs.
Martí Anglada, Enrique de Lucas, Joan-Manuel Parcerisa, Juan L. Aragón, Pedro Marcuello, Antonio González 0001
HPCA5
2019 Visibility Rendering Order: Improving Energy Efficiency on Mobile GPUs through Frame Coherence
abstract
During real-time graphics rendering, objects are processed by the GPU in the order they are submitted by the CPU, and occluded surfaces are often processed even though they will end up not being part of the final image, thus wasting precious time and energy. To help discard occluded surfaces, most current GPUs include an Early-Depth test before the fragment processing stage. However, to be effective it requires that opaque objects are processed in a front-to-back order. Depth sorting and other occlusion culling techniques at the object level incur overheads that are only offset for applications having substantial depth and/or fragment shading complexity, which is often not the case in mobile workloads. We propose a novel architectural technique for GPUs, Visibility Rendering Order (VRO), which reorders objects front-to-back entirely in hardware by exploiting the fact that the objects in graphics animated applications tend to keep its relative depth order across consecutive frames (temporal coherence). Since order relationships are already tested by the Depth Test, VRO incurs minimal energy overheads because it just requires adding a small hardware to capture that information and use it later to guide the rendering of the following frame. Moreover, unlike other approaches, this unit works in parallel with the graphics pipeline without any performance overhead. We illustrate the benefits of VRO using various unmodified commercial 3D applications for which VRO achieves 27 percent speed-up and 15.8 percent energy reduction on average over a state-of-the-art mobile GPU.
Enrique de Lucas, Pedro Marcuello, Joan-Manuel Parcerisa, Antonio González 0001
IEEE Trans. Parallel Distributed Syst.2
2015 Ultra-low power render-based collision detection for CPU/GPU systems
abstract
Smartphones have become powerful computing systems able to carry out complex tasks, such as web browsing, image processing and gaming, among others. Graphics animation applications such as 3D games represent a large percentage of downloaded applications for mobile devices and the trend is towards more complex and realistic scenes with accurate 3D physics simulations, like those in laptops and desktops. Collision detection (CD) is one of the main algorithms used in any physics kernel. However, real-time highly accurate CD is very expensive in terms of energy consumption and this parameter is of paramount importance for mobile devices since it has a direct effect on the autonomy of the system.
Enrique de Lucas, Pedro Marcuello, Joan-Manuel Parcerisa, Antonio González 0001
MICRO2
2011 Fg-STP: Fine-Grain Single Thread Partitioning on Multicores
abstract
Power and complexity issues have led the microprocessor industry to shift to Chip Multiprocessors in order to be able to better utilize the additional transistors ensured by Moore's law. While parallel programs are going to be able to take most of the advantage of these CMPs, single thread applications are not equipped to benefit from them. In this paper we propose Fine-Grain Single-Thread Partitioning (Fg-STP), a hardware-only scheme that takes advantage of CMP designs to speedup single-threaded applications. Our proposal improves single thread performance by reconfiguring two cores with the aim of collaborating on the fetching and execution of the instructions. These cores are basically conventional out-of-order cores in which execution is orchestrated using a dedicated hardware that has minimum and localized impact on the original design of the cores. This approach partitions the code at instruction granularity and differs from previous proposals on the extensive use of dependence speculation, replication and communication. These features are combined with the ability to look for parallelism on large instruction windows without any software intervention (no re-compilation or profiling hints are needed). These characteristics allow Fg-STP to speedup single thread by 18% and 7% on average over similar hardware-only approaches like Core Fusion, on medium sized and small sized 2-core CMP respectively for Spec 2006 benchmarks.
Fernando Latorre, Pedro Marcuello, Antonio González 0001
HPCA3
2009 P-slice based efficient speculative multithreading
abstract
Microprocessor industry has recently shifted towards multi-core to take advantage of the ever increasing number of transistors provided by the new technologies. Unfortunately, the multi-core approach does not allow single threaded applications to benefit from the additional cores to improve their execution time. Speculative multithreading (SpMT) has been proposed in the past to boost performance of irregular applications in multi-core environments. In this work, we study the main bottlenecks of these architectures, such as the memory behavior and the pre-computation slices and propose two novel schemes that allow SpMT to get 25% average speedup over single threaded execution. We propose Selective Replication as a technique to improve the performance of the SpMT memory system. This technique does not introduce additional traffic in the bus and improves the performance of a conventional SpMT memory model by 6% on average and up to 21% for some applications. Also, we propose a scheme called Slice Specialization that reduces the number of instructions in the pre-computation slices by adapting the slice to every single speculative thread spawned. The later proposal outperforms previous schemes with slices by 15% and overall, both techniques combined achieve an improvement of 20% over a conventional SpMT processor.
Pedro Marcuello, Fernando Latorre, Antonio González 0001
HiPC2
2008 Mitosis: A Speculative Multithreaded Processor Based on Precomputation Slices
abstract
This paper presents the Mitosis framework, which is a combined hardware-software approach to speculative multithreading, even in the presence of frequent dependences among threads. Speculative multithreading increases single-threaded application performance by exploiting thread-level parallelism speculatively - that is, executing code in parallel even when the compiler or runtime system cannot guarantee the parallelism exists. The proposed approach is based on predicting/computing thread input values via software, through a piece of code that is added at the beginning of each thread (the pre-computation slice). A pre-computation slice is expected to compute the correct thread input values most of the time, but not necessarily always. This allows aggressive optimization techniques to be applied to the slice to make it very short. This paper focuses on the microarchitecture that supports this execution model. The primary novelty of the microarchitecture is the hardware support for the execution and validation of pre-computation slices. Additionally, this paper presents new architectures for the register file and the cache memory in order to support multiple versions of each variable and allow for efficient roll-back in case of misspeculation. We show that the proposed microarchitecture, together with the compiler support, achieves an average speedup of 2.2 for applications that conventional non-speculative approaches are not able to parallelize at all.
Carlos Madriles, Carlos García Quiñones, F. Jesús Sánchez, Pedro Marcuello, Antonio González 0001, Dean M. Tullsen, Hong Wang 0003, John Paul Shen
IEEE Trans. Parallel Distributed Syst.4
2005 Mitosis compiler: an infrastructure for speculative threading based on pre-computation slices
Carlos García Quiñones, Carlos Madriles, F. Jesús Sánchez, Pedro Marcuello, Antonio González 0001, Dean M. Tullsen
PLDI4
2004 Thread Partitioning and Value Prediction for Exploiting Speculative Thread-Level Parallelism
abstract
Speculative thread-level parallelism has been recently proposed as a source of parallelism to improve the performance in applications where parallel threads are hard to find. However, the efficiency of this execution model strongly depends on the performance of the control and data speculation techniques. Several hardware-based schemes for partitioning the program into speculative threads are analyzed and evaluated. In general, we find that spawning threads associated to loop iterations is the most effective technique. We also show that value prediction is critical for the performance of all of the spawning policies. Thus, a new value predictor, the increment predictor, is proposed. This predictor is specially oriented for this kind of architecture and clearly outperforms the adapted versions of conventional value predictors such as the last value, the stride, and the context-based, especially for small-sized history tables.
Pedro Marcuello, Antonio González 0001, Jordi Tubella
IEEE Trans. Computers1
2003 A framework for modeling and optimization of prescient instruction prefetch
abstract
This paper describes a framework for modeling macroscopic program behavior and applies it to optimizing prescient instruction prefetch -- novel technique that uses helper threads to improve single-threaded application performance by performing judicious and timely instruction prefetch. A helper thread is initiated when the main thread encounters a spawn point, and prefetches instructions starting at a distant target point. The target identifies a code region tending to incur I-cache misses that the main thread is likely to execute soon, even though intervening control flow may be unpredictable. The optimization of spawn-target pair selections is formulated by modeling program behavior as a Markov chain based on profile statistics. Execution paths are considered stochastic outcomes, and aspects of program behavior are summarized via path expression mappings. Mappings for computing reaching, and posteriori probability; path length mean, and variance; and expected path footprint are presented. These are used with Tarjan's fast path algorithm to efficiently estimate the benefit of spawn-target pair selections. Using this framework we propose a spawn-target pair selection algorithm for prescient instruction prefetch. This algorithm has been implemented, and evaluated for the Itanium Processor Family architecture. A limit study finds 4.8%to 17% speedups on an in-order simultaneous multithreading processor with eight contexts, over nextline and streaming I-prefetch for a set of benchmarks with high I-cache miss rates. The framework in this paper is potentially applicable to other thread speculation techniques.
Tor M. Aamodt, Pedro Marcuello, Paul Chow, Antonio González 0001, Per Hammarlund, Hong Wang 0003, John Paul Shen
SIGMETRICS2
2002 Thread-Spawning Schemes for Speculative Multithreading
abstract
Speculative multithreading has been recently proposed to boost performance by means of exploiting thread-level parallelism in applications difficult to parallelize. The performance of these processors heavily depends on the partitioning policy used to split the program into threads. Previous work uses heuristics to spawn speculative threads based on easily-detectable program constructs such as loops or subroutines. In this work we propose a profile-based mechanism to divide programs into threads by searching for those parts of the code that have certain features that could benefit from potential thread-level parallelism. Our profile-based spawning scheme is evaluated on a Clustered Speculative Multithreaded Processor and results show large performance benefits. When the proposed spawning scheme is compared with traditional heuristics, we outperform them by almost 20%. When a realistic value predictor and a 8-cycle thread initialization penalty is considered, the performance difference between them is maintained. The speed-up over a single thread execution is higher than 5x for a 16-thread-unit processor and close to 2x for a 4-thread-unit processor.
Pedro Marcuello, Antonio González 0001
HPCA1
2000 A Quantitative Assessment of Thread-Level Speculation Techniques
abstract
Speculative thread-level parallelism has been recently proposed as an alternative source of parallelism that can boost the performance for applications where independent threads are hard to find. Several schemes to exploit thread level parallelism have been proposed and significant performance gains have been reported. However, the sources of the performance gains are poorly understood as well as the impact of some design choices. In this work, the advantages of different thread speculation techniques are analyzed as are the impact of some critical issues including the value predictor, the branch predictor, the thread initialization overhead and the connectivity among thread units.
Pedro Marcuello, Antonio González 0001
IPDPS1
1999 Clustered speculative multithreaded processors
abstract
In this paper we present a processor microarchitecture that can simultaneously execute multiple threads and has a clustered design for scalability purposes.A main feature of the proposed microarchitecture is its capability to spawn speculative threads from a single-thread application at run-time.These speculative threaak use otherwise idle resources of the machine.Spawning a speculative thread involves predicting its control flow as well as its dependences with other threads and the values that flow through them.In this way, threads fhat are not independent can be executed in parallel.Control-Jlow, data value and data dependence predictors particularly designedfor this type of microarchitecture are presented.Results show the potential of the microarchitecture to exploit speculative parallelism in programs that are hard to parallelize at compile-time, such as the SpecInt9.5.For a 4-thread unit configuration, some programs such as ijpeg and Ii can exploit an average degree of parallelism of more than 2 threads per cycle.The average degree ofparallelism for the whole SpecInt95 suite is 1.6 threads per cycle.This speculative parallelism results in significant speedups for all the Speclnt95 programs when compared with a single-thread execution.
Pedro Marcuello, Antonio González 0001
International Conference on Supercomputing1
1999 Value Prediction for Speculative Multithreaded Architectures
abstract
The speculative multithreading paradigm (speculative thread-level parallelism) is based on the concurrent execution of control-speculative threads. The efficiency of microarchitectures that adopt this paradigm strongly depends on the performance of the control and data speculation techniques. While control speculation is used to predict the most effective points where a thread can be spawned, data speculation is required to eliminate the serialization imposed by inter-thread dependences. This work studies the performance of different value predictors for speculative multithreaded processors. We propose a value predictor, the increment predictor, and evaluate its performance for a particular microarchitecture that implements this execution paradigm (Clustered Speculative Multithreaded architecture). The proposed trace-oriented increment predictor clearly outperforms trace-adapted versions of the last value, stride and context-based predictors, specially for small-sized history tables. A 1-KB increment predictor achieves a 73% prediction accuracy and a performance that is just 13% lower than that of a perfect value predictor.
Pedro Marcuello, Jordi Tubella, Antonio González 0001
MICRO1
1998 Speculative Multithreaded Processors
abstract
In this paper we present a novel processor microarchitecture that relieves four of the most important bottlenecks of superscalar processors to exploit instruction level parallelism: the serialization imposed by true dependences, the instruction window size, the complexity of a wide issue machine and the instruction fetch bandwidth requirements. The new microarchitecture executes simultaneously multiple threads of control obtained from a single program by means of control speculation techniques that do not require any compiler/user support. In this way, it works on a large instruction window composed of multiple nonadjacent small windows. Multiple simultaneous threads execute different iterations of the same loop, which requires the same fetch bandwidth as a single thread since they share the same code. Dependences among different threads as well as the values that flow through them are speculated by means of data prediction techniques. The novel processor organization does not require ...
Pedro Marcuello, Antonio González 0001, Jordi Tubella
International Conference on Supercomputing1