Enrique de Lucas

dblp:173/9802 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
2since 2021 · last 2022
0000-0002-5717-7312ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 3 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
GPUs and heterogeneous computing · 54% Energy-efficient computing · 32% Memory systems · 9%
Computer graphics and multimedia
3 papers
Rendering · 100%

Topics — the 12 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
GPUs and heterogeneous computing › GPU rendering
GPU graphics pipeline
0.822019
Rendering Elimination: Early Discard of Redundant Tiles in the Graphics Pipeline · HPCA 2019
Early Visibility Resolution for Removing Ineffectual Computations in the Graphics Pipeline · HPCA 2019
Energy-efficient computing › energy-efficient architecture
GPU energy efficiency
0.612022
Omega-Test: A Predictive Early-Z Culling to Improve the Graphics Pipeline Energy-Efficiency · IEEE Trans. Vis. Comput. Graph. 2022
GPUs and heterogeneous computing
GPU microarchitecture
0.612022
Omega-Test: A Predictive Early-Z Culling to Improve the Graphics Pipeline Energy-Efficiency · IEEE Trans. Vis. Comput. Graph. 2022
GPUs and heterogeneous computing
GPU architecture
0.412019
Visibility Rendering Order: Improving Energy Efficiency on Mobile GPUs through Frame Coherence · IEEE Trans. Parallel Distributed Syst. 2019
Memory systems › memory bandwidth management
memory bandwidth reduction
0.412019
Rendering Elimination: Early Discard of Redundant Tiles in the Graphics Pipeline · HPCA 2019
GPUs and heterogeneous computing › GPU rendering
mobile GPU rendering
0.412019
Visibility Rendering Order: Improving Energy Efficiency on Mobile GPUs through Frame Coherence · IEEE Trans. Parallel Distributed Syst. 2019
Energy-efficient computing › energy-efficient architecture
GPU energy reduction
0.222019
Rendering Elimination: Early Discard of Redundant Tiles in the Graphics Pipeline · HPCA 2019
Early Visibility Resolution for Removing Ineffectual Computations in the Graphics Pipeline · HPCA 2019
Embedded and real-time systems
collision detection
0.212015
Ultra-low power render-based collision detection for CPU/GPU systems · MICRO 2015
GPUs and heterogeneous computing › CPU-GPU heterogeneous computing
CPU-GPU heterogeneous systems
0.212015
Ultra-low power render-based collision detection for CPU/GPU systems · MICRO 2015
Energy-efficient computing › power management
mobile device energy
0.212015
Ultra-low power render-based collision detection for CPU/GPU systems · MICRO 2015
Energy-efficient computing
power management
0.212015
Ultra-low power render-based collision detection for CPU/GPU systems · MICRO 2015
Rendering
visibility computation
0.212022
Omega-Test: A Predictive Early-Z Culling to Improve the Graphics Pipeline Energy-Efficiency · IEEE Trans. Vis. Comput. Graph. 2022

Methods — techniques the papers use, named apart from their topics

tile-based rendering · 1.1frame-to-frame coherence · 1.1temporal coherence exploitation · 0.8early-depth test · 0.8render-based collision detection · 0.4speculative visibility prediction · 0.4signature comparison · 0.4
YearPublicationVenuePosition
2022 Dynamic sampling rate: harnessing frame coherence in graphics applications for energy-efficient GPUs
abstract
In real-time rendering, a 3D scene is modelled with meshes of triangles that the GPU projects to the screen. They are discretized by sampling each triangle at regular space intervals to generate fragments which are then added texture and lighting effects by a shader program. Realistic scenes require detailed geometric models, complex shaders, high-resolution displays and high screen refreshing rates, which all come at a great compute time and energy cost. This cost is often dominated by the fragment shader, which runs for each sampled fragment. Conventional GPUs sample the triangles once per pixel; however, there are many screen regions containing low variation that produce identical fragments and could be sampled at lower than pixel-rate with no loss in quality. Additionally, as temporal frame coherence makes consecutive frames very similar, such variations are usually maintained from frame to frame. This work proposes Dynamic Sampling Rate (DSR), a novel hardware mechanism to reduce redundancy and improve the energy efficiency in graphics applications. DSR analyzes the spatial frequencies of the scene once it has been rendered. Then, it leverages the temporal coherence in consecutive frames to decide, for each region of the screen, the lowest sampling rate to employ in the next frame that maintains image quality. We evaluate the performance of a state-of-the-art mobile GPU architecture extended with DSR for a wide variety of applications. Experimental results show that DSR is able to remove most of the redundancy inherent in the color computations at fragment granularity, which brings average speedups of 1.68x and energy savings of 40%.
Martí Anglada, Enrique de Lucas, Joan-Manuel Parcerisa, Juan L. Aragón, Antonio González 0001
J. Supercomput.2
2022 Omega-Test: A Predictive Early-Z Culling to Improve the Graphics Pipeline Energy-Efficiency
abstract
The most common task of GPUs is to render images in real time. When rendering a 3D scene, a key step is to determine which parts of every object are visible in the final image. There are different approaches to solve the visibility problem, the Z-Test being the most common. A main factor that significantly penalizes the energy efficiency of a GPU, especially in the mobile arena, is the so-called overdraw, which happens when a portion of an object is shaded and rendered but finally occluded by another object. This useless work results in a waste of energy; however, a conventional Z-Test only avoids a fraction of it. In this article we present a novel microarchitectural technique, the Omega-Test, to drastically reduce the overdraw on a Tile-Based Rendering (TBR) architecture. Graphics applications have a great degree of inter-frame coherence, which makes the output of a frame very similar to the previous one. The proposed approach leverages the frame-to-frame coherence by using the resulting information of the Z-Test for a tile (a buffer containing all the calculated pixel depths for a tile), which is discarded by nowadays GPUs, to predict the visibility of the same tile in the next frame. As a result, the Omega-Test early identifies occluded parts of the scene and avoids the rendering of non-visible surfaces eliminating costly computations and off-chip memory accesses. Our experimental evaluation shows average EDP savings in the overall GPU/Memory system of 26.4 percent and an average speedup of 16.3 percent for the evaluated benchmarks.
David Corbalán-Navarro, Juan L. Aragón, Martí Anglada, Enrique de Lucas, Joan-Manuel Parcerisa, Antonio González 0001
IEEE Trans. Vis. Comput. Graph.4
2019 Early Visibility Resolution for Removing Ineffectual Computations in the Graphics Pipeline
abstract
GPUs' main workload is real-time image rendering. These applications take a description of a (animated) scene and produce the corresponding image(s). An image is rendered by computing the colors of all its pixels. It is normal that multiple objects overlap at each pixel. Consequently, a significant amount of processing is devoted to objects that will not be visible in the final image, in spite of the widespread use of the Early Depth Test in modern GPUs, which attempts to discard computations related to occluded objects. Since animations are created by a sequence of similar images, visibility usually does not change much across consecutive frames. Based on this observation, we present Early Visibility Resolution (EVR), a mechanism that leverages the visibility information obtained in a frame to predict the visibility in the following one. Our proposal speculatively determines visibility much earlier in the pipeline than the Early Depth Test. We leverage this early visibility estimation to remove ineffectual computations at two different granularities: pixel-level and tile-level. Results show that such optimizations lead to 39% performance improvement and 43% energy savings for a set of commercial Android graphics applications running on stateof-the-art mobile GPUs.
Martí Anglada, Enrique de Lucas, Joan-Manuel Parcerisa, Juan L. Aragón, Antonio González 0001
HPCA2
2019 Rendering Elimination: Early Discard of Redundant Tiles in the Graphics Pipeline
abstract
GPUs are one of the most energy-consuming components for real-time rendering applications, since a large number of fragment shading computations and memory accesses are involved. Main memory bandwidth is especially taxing battery-operated devices such as smart-phones. TileBased Rendering GPUs divide the screen space into multiple tiles that are independently rendered in on-chip buffers, thus reducing memory bandwidth and energy consumption. We have observed that, in many animated graphics workloads, a large number of screen tiles have the same color across adjacent frames. In this paper, we propose Rendering Elimination (RE), a novel micro-architectural technique that accurately determines if a tile will be identical to the same tile in the preceding frame before rasterization by means of comparing signatures. Since RE identifies redundant tiles early in the graphics pipeline, it completely avoids the computation and memory accesses of the most power consuming stages of the pipeline, which substantially reduces the execution time and the energy consumption of the GPU. For widely used Android applications, we show that RE achieves an average speedup of 1.74x and energy reduction of 43% for the GPU/Memory system, surpassing by far the benefits of Transaction Elimination, a state-of-the-art memory bandwidth reduction technique available in some commercial Tile-Based Rendering GPUs.
Martí Anglada, Enrique de Lucas, Joan-Manuel Parcerisa, Juan L. Aragón, Pedro Marcuello, Antonio González 0001
HPCA2
2019 Visibility Rendering Order: Improving Energy Efficiency on Mobile GPUs through Frame Coherence
abstract
During real-time graphics rendering, objects are processed by the GPU in the order they are submitted by the CPU, and occluded surfaces are often processed even though they will end up not being part of the final image, thus wasting precious time and energy. To help discard occluded surfaces, most current GPUs include an Early-Depth test before the fragment processing stage. However, to be effective it requires that opaque objects are processed in a front-to-back order. Depth sorting and other occlusion culling techniques at the object level incur overheads that are only offset for applications having substantial depth and/or fragment shading complexity, which is often not the case in mobile workloads. We propose a novel architectural technique for GPUs, Visibility Rendering Order (VRO), which reorders objects front-to-back entirely in hardware by exploiting the fact that the objects in graphics animated applications tend to keep its relative depth order across consecutive frames (temporal coherence). Since order relationships are already tested by the Depth Test, VRO incurs minimal energy overheads because it just requires adding a small hardware to capture that information and use it later to guide the rendering of the following frame. Moreover, unlike other approaches, this unit works in parallel with the graphics pipeline without any performance overhead. We illustrate the benefits of VRO using various unmodified commercial 3D applications for which VRO achieves 27 percent speed-up and 15.8 percent energy reduction on average over a state-of-the-art mobile GPU.
Enrique de Lucas, Pedro Marcuello, Joan-Manuel Parcerisa, Antonio González 0001
IEEE Trans. Parallel Distributed Syst.1
2017 DSPONE48: A methodology for automatically synthesize HDL focus on the reuse of DSP slices
Enrique de Lucas, Marcos Sánchez-Élez Martín, Inmaculada Pardines
J. Parallel Distributed Comput.1
2015 Ultra-low power render-based collision detection for CPU/GPU systems
abstract
Smartphones have become powerful computing systems able to carry out complex tasks, such as web browsing, image processing and gaming, among others. Graphics animation applications such as 3D games represent a large percentage of downloaded applications for mobile devices and the trend is towards more complex and realistic scenes with accurate 3D physics simulations, like those in laptops and desktops. Collision detection (CD) is one of the main algorithms used in any physics kernel. However, real-time highly accurate CD is very expensive in terms of energy consumption and this parameter is of paramount importance for mobile devices since it has a direct effect on the autonomy of the system.
Enrique de Lucas, Pedro Marcuello, Joan-Manuel Parcerisa, Antonio González 0001
MICRO1