EDBT 2026 Demo / reviewers in the wild / expert
Pedro Marcuello
dblp:85/1247
· DBLP profile ↗
15ranked-venue papers
6as first author
1since 2021 · last 2021
0000-0001-6104-9105ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 14 · 6 first-author · 1 since 2021Software engineering, systems software and programming languages · 3 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
10 papers |
GPUs and heterogeneous computing · 37% Processor architecture and microarchitecture · 23% Energy-efficient computing · 18% | |
| Computer graphics and multimedia
2 papers |
Rendering · 100% |
Topics — the 24 heaviest of 29, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
GPUs and heterogeneous computing
GPU architecture |
0.4 | 1 | 2019 | Visibility Rendering Order: Improving Energy Efficiency on Mobile GPUs through Frame Coherence · IEEE Trans. Parallel Distributed Syst. 2019 |
GPUs and heterogeneous computing › GPU rendering
GPU graphics pipeline |
0.4 | 1 | 2019 | Rendering Elimination: Early Discard of Redundant Tiles in the Graphics Pipeline · HPCA 2019 |
Memory systems › memory bandwidth management
memory bandwidth reduction |
0.4 | 1 | 2019 | Rendering Elimination: Early Discard of Redundant Tiles in the Graphics Pipeline · HPCA 2019 |
GPUs and heterogeneous computing › GPU rendering
mobile GPU rendering |
0.4 | 1 | 2019 | Visibility Rendering Order: Improving Energy Efficiency on Mobile GPUs through Frame Coherence · IEEE Trans. Parallel Distributed Syst. 2019 |
Embedded and real-time systems
collision detection |
0.2 | 1 | 2015 | Ultra-low power render-based collision detection for CPU/GPU systems · MICRO 2015 |
GPUs and heterogeneous computing › CPU-GPU heterogeneous computing
CPU-GPU heterogeneous systems |
0.2 | 1 | 2015 | Ultra-low power render-based collision detection for CPU/GPU systems · MICRO 2015 |
Energy-efficient computing › power management
mobile device energy |
0.2 | 1 | 2015 | Ultra-low power render-based collision detection for CPU/GPU systems · MICRO 2015 |
Energy-efficient computing
power management |
0.2 | 1 | 2015 | Ultra-low power render-based collision detection for CPU/GPU systems · MICRO 2015 |
Processor architecture and microarchitecture › multithreading
speculative multithreading |
0.1 | 3 | 2008 | Mitosis: A Speculative Multithreaded Processor Based on Precomputation Slices · IEEE Trans. Parallel Distributed Syst. 2008 Thread-Spawning Schemes for Speculative Multithreading · HPCA 2002 Value Prediction for Speculative Multithreaded Architectures · MICRO 1999 |
Processor architecture and microarchitecture
instruction-level parallelism |
0.1 | 2 | 2011 | Fg-STP: Fine-Grain Single Thread Partitioning on Multicores · HPCA 2011 Value Prediction for Speculative Multithreaded Architectures · MICRO 1999 |
Processor architecture and microarchitecture
chip multiprocessor |
0.1 | 1 | 2011 | Fg-STP: Fine-Grain Single Thread Partitioning on Multicores · HPCA 2011 |
Processor architecture and microarchitecture
thread partitioning |
0.1 | 1 | 2011 | Fg-STP: Fine-Grain Single Thread Partitioning on Multicores · HPCA 2011 |
Energy-efficient computing › energy-efficient architecture
GPU energy reduction |
0.1 | 1 | 2019 | Rendering Elimination: Early Discard of Redundant Tiles in the Graphics Pipeline · HPCA 2019 |
Parallel and multicore computing › speculative parallelization
thread-level speculation |
0.1 | 2 | 2004 | Thread Partitioning and Value Prediction for Exploiting Speculative Thread-Level Parallelism · IEEE Trans. Computers 2004 A framework for modeling and optimization of prescient instruction prefetch · SIGMETRICS 2003 |
Processor architecture and microarchitecture
speculative execution |
0.1 | 2 | 2005 | Mitosis compiler: an infrastructure for speculative threading based on pre-computation slices · PLDI 2005 Value Prediction for Speculative Multithreaded Architectures · MICRO 1999 |
Processor architecture and microarchitecture
value prediction |
0.1 | 2 | 2004 | Thread Partitioning and Value Prediction for Exploiting Speculative Thread-Level Parallelism · IEEE Trans. Computers 2004 Value Prediction for Speculative Multithreaded Architectures · MICRO 1999 |
Parallel and multicore computing
thread-level parallelism |
0.1 | 3 | 2004 | Thread-Spawning Schemes for Speculative Multithreading · HPCA 2002 Thread Partitioning and Value Prediction for Exploiting Speculative Thread-Level Parallelism · IEEE Trans. Computers 2004 Value Prediction for Speculative Multithreaded Architectures · MICRO 1999 |
Compilers and program optimization › parallelization
thread-level speculation |
0.1 | 1 | 2005 | Mitosis compiler: an infrastructure for speculative threading based on pre-computation slices · PLDI 2005 |
Processor architecture and microarchitecture › multithreading
helper threads |
0.0 | 1 | 2003 | A framework for modeling and optimization of prescient instruction prefetch · SIGMETRICS 2003 |
Processor architecture and microarchitecture › instruction fetch
instruction prefetching |
0.0 | 1 | 2003 | A framework for modeling and optimization of prescient instruction prefetch · SIGMETRICS 2003 |
Processor architecture and microarchitecture
multithreading |
0.0 | 1 | 1999 | Value Prediction for Speculative Multithreaded Architectures · MICRO 1999 |
Performance modeling and evaluation › workload characterization › program behavior
program behavior modeling |
0.0 | 1 | 2003 | A framework for modeling and optimization of prescient instruction prefetch · SIGMETRICS 2003 |
Performance modeling and evaluation
workload characterization |
0.0 | 1 | 2003 | A framework for modeling and optimization of prescient instruction prefetch · SIGMETRICS 2003 |
Compilers and program optimization › program transformation
program partitioning |
0.0 | 1 | 2002 | Thread-Spawning Schemes for Speculative Multithreading · HPCA 2002 |
Methods — techniques the papers use, named apart from their topics
temporal coherence exploitation · 0.8early-depth test · 0.8render-based collision detection · 0.4signature comparison · 0.4speculation · 0.2hardware-software co-design · 0.2instruction replication · 0.1dependence speculation · 0.1pre-computation slices · 0.1control speculation · 0.0profile-based analysis · 0.0heuristics · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Mont-Blanc 2020: Towards Scalable and Power Efficient European HPC ProcessorsabstractThe Mont-Blanc 2020 (MB2020) project has triggered the development of the next generation industrial processor for Big Data and High Performance Computing (HPC). MB2020 is paving the way to the future low-power European processor for exascale, defining the System-on-Chip (SoC) architecture and implementing new critical building blocks to be integrated in such an SoC. In this paper, we first present an overview of the MB2020 project, then we describe our experimental infrastructure, the requirements of relevant applications, and the IP blocks developed in the project. Finally, we present our emulation-based final demonstrator and explain how it integrates within our first generation of HPC processors. Adrià Armejach, Bine Brank, Jordi Cortina, François Dolique, Timothy Hayes 0001, Nam Ho, Pierre-Axel Lagadec, Romain Lemaire, Guillem López-Paradís, Laurent Marliac, Miquel Moretó, Pedro Marcuello, Dirk Pleiter, Xubin Tan, Said Derradji |
DATE | 12 |
| 2019 | Rendering Elimination: Early Discard of Redundant Tiles in the Graphics PipelineabstractGPUs are one of the most energy-consuming components for real-time rendering applications, since a large number of fragment shading computations and memory accesses are involved. Main memory bandwidth is especially taxing battery-operated devices such as smart-phones. TileBased Rendering GPUs divide the screen space into multiple tiles that are independently rendered in on-chip buffers, thus reducing memory bandwidth and energy consumption. We have observed that, in many animated graphics workloads, a large number of screen tiles have the same color across adjacent frames. In this paper, we propose Rendering Elimination (RE), a novel micro-architectural technique that accurately determines if a tile will be identical to the same tile in the preceding frame before rasterization by means of comparing signatures. Since RE identifies redundant tiles early in the graphics pipeline, it completely avoids the computation and memory accesses of the most power consuming stages of the pipeline, which substantially reduces the execution time and the energy consumption of the GPU. For widely used Android applications, we show that RE achieves an average speedup of 1.74x and energy reduction of 43% for the GPU/Memory system, surpassing by far the benefits of Transaction Elimination, a state-of-the-art memory bandwidth reduction technique available in some commercial Tile-Based Rendering GPUs. Martí Anglada, Enrique de Lucas, Joan-Manuel Parcerisa, Juan L. Aragón, Pedro Marcuello, Antonio González 0001 |
HPCA | 5 |
| 2019 | Visibility Rendering Order: Improving Energy Efficiency on Mobile GPUs through Frame CoherenceabstractDuring real-time graphics rendering, objects are processed by the GPU in the order they are submitted by the CPU, and occluded surfaces are often processed even though they will end up not being part of the final image, thus wasting precious time and energy. To help discard occluded surfaces, most current GPUs include an Early-Depth test before the fragment processing stage. However, to be effective it requires that opaque objects are processed in a front-to-back order. Depth sorting and other occlusion culling techniques at the object level incur overheads that are only offset for applications having substantial depth and/or fragment shading complexity, which is often not the case in mobile workloads. We propose a novel architectural technique for GPUs, Visibility Rendering Order (VRO), which reorders objects front-to-back entirely in hardware by exploiting the fact that the objects in graphics animated applications tend to keep its relative depth order across consecutive frames (temporal coherence). Since order relationships are already tested by the Depth Test, VRO incurs minimal energy overheads because it just requires adding a small hardware to capture that information and use it later to guide the rendering of the following frame. Moreover, unlike other approaches, this unit works in parallel with the graphics pipeline without any performance overhead. We illustrate the benefits of VRO using various unmodified commercial 3D applications for which VRO achieves 27 percent speed-up and 15.8 percent energy reduction on average over a state-of-the-art mobile GPU. Enrique de Lucas, Pedro Marcuello, Joan-Manuel Parcerisa, Antonio González 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2015 | Ultra-low power render-based collision detection for CPU/GPU systemsabstractSmartphones have become powerful computing systems able to carry out complex tasks, such as web browsing, image processing and gaming, among others. Graphics animation applications such as 3D games represent a large percentage of downloaded applications for mobile devices and the trend is towards more complex and realistic scenes with accurate 3D physics simulations, like those in laptops and desktops. Collision detection (CD) is one of the main algorithms used in any physics kernel. However, real-time highly accurate CD is very expensive in terms of energy consumption and this parameter is of paramount importance for mobile devices since it has a direct effect on the autonomy of the system. Enrique de Lucas, Pedro Marcuello, Joan-Manuel Parcerisa, Antonio González 0001 |
MICRO | 2 |
| 2011 | Fg-STP: Fine-Grain Single Thread Partitioning on MulticoresabstractPower and complexity issues have led the microprocessor industry to shift to Chip Multiprocessors in order to be able to better utilize the additional transistors ensured by Moore's law. While parallel programs are going to be able to take most of the advantage of these CMPs, single thread applications are not equipped to benefit from them. In this paper we propose Fine-Grain Single-Thread Partitioning (Fg-STP), a hardware-only scheme that takes advantage of CMP designs to speedup single-threaded applications. Our proposal improves single thread performance by reconfiguring two cores with the aim of collaborating on the fetching and execution of the instructions. These cores are basically conventional out-of-order cores in which execution is orchestrated using a dedicated hardware that has minimum and localized impact on the original design of the cores. This approach partitions the code at instruction granularity and differs from previous proposals on the extensive use of dependence speculation, replication and communication. These features are combined with the ability to look for parallelism on large instruction windows without any software intervention (no re-compilation or profiling hints are needed). These characteristics allow Fg-STP to speedup single thread by 18% and 7% on average over similar hardware-only approaches like Core Fusion, on medium sized and small sized 2-core CMP respectively for Spec 2006 benchmarks. Fernando Latorre, Pedro Marcuello, Antonio González 0001 |
HPCA | 3 |
| 2009 | P-slice based efficient speculative multithreadingabstractMicroprocessor industry has recently shifted towards multi-core to take advantage of the ever increasing number of transistors provided by the new technologies. Unfortunately, the multi-core approach does not allow single threaded applications to benefit from the additional cores to improve their execution time. Speculative multithreading (SpMT) has been proposed in the past to boost performance of irregular applications in multi-core environments. In this work, we study the main bottlenecks of these architectures, such as the memory behavior and the pre-computation slices and propose two novel schemes that allow SpMT to get 25% average speedup over single threaded execution. We propose Selective Replication as a technique to improve the performance of the SpMT memory system. This technique does not introduce additional traffic in the bus and improves the performance of a conventional SpMT memory model by 6% on average and up to 21% for some applications. Also, we propose a scheme called Slice Specialization that reduces the number of instructions in the pre-computation slices by adapting the slice to every single speculative thread spawned. The later proposal outperforms previous schemes with slices by 15% and overall, both techniques combined achieve an improvement of 20% over a conventional SpMT processor. Pedro Marcuello, Fernando Latorre, Antonio González 0001 |
HiPC | 2 |
| 2008 | Mitosis: A Speculative Multithreaded Processor Based on Precomputation SlicesabstractThis paper presents the Mitosis framework, which is a combined hardware-software approach to speculative multithreading, even in the presence of frequent dependences among threads. Speculative multithreading increases single-threaded application performance by exploiting thread-level parallelism speculatively - that is, executing code in parallel even when the compiler or runtime system cannot guarantee the parallelism exists. The proposed approach is based on predicting/computing thread input values via software, through a piece of code that is added at the beginning of each thread (the pre-computation slice). A pre-computation slice is expected to compute the correct thread input values most of the time, but not necessarily always. This allows aggressive optimization techniques to be applied to the slice to make it very short. This paper focuses on the microarchitecture that supports this execution model. The primary novelty of the microarchitecture is the hardware support for the execution and validation of pre-computation slices. Additionally, this paper presents new architectures for the register file and the cache memory in order to support multiple versions of each variable and allow for efficient roll-back in case of misspeculation. We show that the proposed microarchitecture, together with the compiler support, achieves an average speedup of 2.2 for applications that conventional non-speculative approaches are not able to parallelize at all. Carlos Madriles, Carlos García Quiñones, F. Jesús Sánchez, Pedro Marcuello, Antonio González 0001, Dean M. Tullsen, Hong Wang 0003, John Paul Shen |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2005 | Mitosis compiler: an infrastructure for speculative threading based on pre-computation slices
Carlos García Quiñones, Carlos Madriles, F. Jesús Sánchez, Pedro Marcuello, Antonio González 0001, Dean M. Tullsen |
PLDI | 4 |
| 2004 | Thread Partitioning and Value Prediction for Exploiting Speculative Thread-Level ParallelismabstractSpeculative thread-level parallelism has been recently proposed as a source of parallelism to improve the performance in applications where parallel threads are hard to find. However, the efficiency of this execution model strongly depends on the performance of the control and data speculation techniques. Several hardware-based schemes for partitioning the program into speculative threads are analyzed and evaluated. In general, we find that spawning threads associated to loop iterations is the most effective technique. We also show that value prediction is critical for the performance of all of the spawning policies. Thus, a new value predictor, the increment predictor, is proposed. This predictor is specially oriented for this kind of architecture and clearly outperforms the adapted versions of conventional value predictors such as the last value, the stride, and the context-based, especially for small-sized history tables. Pedro Marcuello, Antonio González 0001, Jordi Tubella |
IEEE Trans. Computers | 1 |
| 2003 | A framework for modeling and optimization of prescient instruction prefetchabstractThis paper describes a framework for modeling macroscopic program behavior and applies it to optimizing prescient instruction prefetch -- novel technique that uses helper threads to improve single-threaded application performance by performing judicious and timely instruction prefetch. A helper thread is initiated when the main thread encounters a spawn point, and prefetches instructions starting at a distant target point. The target identifies a code region tending to incur I-cache misses that the main thread is likely to execute soon, even though intervening control flow may be unpredictable. The optimization of spawn-target pair selections is formulated by modeling program behavior as a Markov chain based on profile statistics. Execution paths are considered stochastic outcomes, and aspects of program behavior are summarized via path expression mappings. Mappings for computing reaching, and posteriori probability; path length mean, and variance; and expected path footprint are presented. These are used with Tarjan's fast path algorithm to efficiently estimate the benefit of spawn-target pair selections. Using this framework we propose a spawn-target pair selection algorithm for prescient instruction prefetch. This algorithm has been implemented, and evaluated for the Itanium Processor Family architecture. A limit study finds 4.8%to 17% speedups on an in-order simultaneous multithreading processor with eight contexts, over nextline and streaming I-prefetch for a set of benchmarks with high I-cache miss rates. The framework in this paper is potentially applicable to other thread speculation techniques. Tor M. Aamodt, Pedro Marcuello, Paul Chow, Antonio González 0001, Per Hammarlund, Hong Wang 0003, John Paul Shen |
SIGMETRICS | 2 |
| 2002 | Thread-Spawning Schemes for Speculative MultithreadingabstractSpeculative multithreading has been recently proposed to boost performance by means of exploiting thread-level parallelism in applications difficult to parallelize. The performance of these processors heavily depends on the partitioning policy used to split the program into threads. Previous work uses heuristics to spawn speculative threads based on easily-detectable program constructs such as loops or subroutines. In this work we propose a profile-based mechanism to divide programs into threads by searching for those parts of the code that have certain features that could benefit from potential thread-level parallelism. Our profile-based spawning scheme is evaluated on a Clustered Speculative Multithreaded Processor and results show large performance benefits. When the proposed spawning scheme is compared with traditional heuristics, we outperform them by almost 20%. When a realistic value predictor and a 8-cycle thread initialization penalty is considered, the performance difference between them is maintained. The speed-up over a single thread execution is higher than 5x for a 16-thread-unit processor and close to 2x for a 4-thread-unit processor. Pedro Marcuello, Antonio González 0001 |
HPCA | 1 |
| 2000 | A Quantitative Assessment of Thread-Level Speculation TechniquesabstractSpeculative thread-level parallelism has been recently proposed as an alternative source of parallelism that can boost the performance for applications where independent threads are hard to find. Several schemes to exploit thread level parallelism have been proposed and significant performance gains have been reported. However, the sources of the performance gains are poorly understood as well as the impact of some design choices. In this work, the advantages of different thread speculation techniques are analyzed as are the impact of some critical issues including the value predictor, the branch predictor, the thread initialization overhead and the connectivity among thread units. Pedro Marcuello, Antonio González 0001 |
IPDPS | 1 |
| 1999 | Clustered speculative multithreaded processorsabstractIn this paper we present a processor microarchitecture that can simultaneously execute multiple threads and has a clustered design for scalability purposes.A main feature of the proposed microarchitecture is its capability to spawn speculative threads from a single-thread application at run-time.These speculative threaak use otherwise idle resources of the machine.Spawning a speculative thread involves predicting its control flow as well as its dependences with other threads and the values that flow through them.In this way, threads fhat are not independent can be executed in parallel.Control-Jlow, data value and data dependence predictors particularly designedfor this type of microarchitecture are presented.Results show the potential of the microarchitecture to exploit speculative parallelism in programs that are hard to parallelize at compile-time, such as the SpecInt9.5.For a 4-thread unit configuration, some programs such as ijpeg and Ii can exploit an average degree of parallelism of more than 2 threads per cycle.The average degree ofparallelism for the whole SpecInt95 suite is 1.6 threads per cycle.This speculative parallelism results in significant speedups for all the Speclnt95 programs when compared with a single-thread execution. Pedro Marcuello, Antonio González 0001 |
International Conference on Supercomputing | 1 |
| 1999 | Value Prediction for Speculative Multithreaded ArchitecturesabstractThe speculative multithreading paradigm (speculative thread-level parallelism) is based on the concurrent execution of control-speculative threads. The efficiency of microarchitectures that adopt this paradigm strongly depends on the performance of the control and data speculation techniques. While control speculation is used to predict the most effective points where a thread can be spawned, data speculation is required to eliminate the serialization imposed by inter-thread dependences. This work studies the performance of different value predictors for speculative multithreaded processors. We propose a value predictor, the increment predictor, and evaluate its performance for a particular microarchitecture that implements this execution paradigm (Clustered Speculative Multithreaded architecture). The proposed trace-oriented increment predictor clearly outperforms trace-adapted versions of the last value, stride and context-based predictors, specially for small-sized history tables. A 1-KB increment predictor achieves a 73% prediction accuracy and a performance that is just 13% lower than that of a perfect value predictor. Pedro Marcuello, Jordi Tubella, Antonio González 0001 |
MICRO | 1 |
| 1998 | Speculative Multithreaded ProcessorsabstractIn this paper we present a novel processor microarchitecture that relieves four of the most important bottlenecks of superscalar processors to exploit instruction level parallelism: the serialization imposed by true dependences, the instruction window size, the complexity of a wide issue machine and the instruction fetch bandwidth requirements. The new microarchitecture executes simultaneously multiple threads of control obtained from a single program by means of control speculation techniques that do not require any compiler/user support. In this way, it works on a large instruction window composed of multiple nonadjacent small windows. Multiple simultaneous threads execute different iterations of the same loop, which requires the same fetch bandwidth as a single thread since they share the same code. Dependences among different threads as well as the values that flow through them are speculated by means of data prediction techniques. The novel processor organization does not require ... Pedro Marcuello, Antonio González 0001, Jordi Tubella |
International Conference on Supercomputing | 1 |