Germán Ceballos

dblp:157/6480 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
0since 2021 · last 2019
0000-0003-2314-7307ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Performance modeling and evaluation · 87% Parallel and multicore computing · 13%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Performance modeling and evaluation › simulation › architectural simulation
sampled simulation
0.412019
Sampled Simulation of Task-Based Programs · IEEE Trans. Computers 2019
Performance modeling and evaluation
simulation
0.412019
Sampled Simulation of Task-Based Programs · IEEE Trans. Computers 2019
Parallel and multicore computing › parallel programming models
task-based programming
0.112019
Sampled Simulation of Task-Based Programs · IEEE Trans. Computers 2019

Methods — techniques the papers use, named apart from their topics

analytical performance modeling · 0.4DBSCAN clustering · 0.4
YearPublicationVenuePosition
2019 Sampled Simulation of Task-Based Programs
abstract
Sampled simulation is a mature technique for reducing simulation time of single-threaded programs. Nevertheless, current sampling techniques do not take advantage of other execution models, like task-based execution, to provide both more accurate and faster simulation. Recent multi-threaded sampling techniques assume that the workload assigned to each thread does not change across multiple executions of a program. This assumption does not hold for dynamically scheduled task-based programming models. Task-based programming models allow the programmer to specify program segments as tasks which are instantiated many times and scheduled dynamically to available threads. Due to variation in scheduling decisions, two consecutive executions on the same machine typically result in different instruction streams processed by each thread. In this paper, we propose TaskPoint, a sampled simulation technique for dynamically scheduled task-based programs. We leverage task instances as sampling units and simulate only a fraction of all task instances in detail. Between detailed simulation intervals, we employ a novel fast-forwarding mechanism for dynamically scheduled programs. We evaluate different automatic techniques for clustering task instances and show that DBSCAN clustering combined with analytical performance modeling provides the best trade-off of simulation speed and accuracy. TaskPoint is the first technique combining sampled simulation and analytical modeling and provides a new way to trade off simulation speed and accuracy. Compared to detailed simulation, TaskPoint accelerates architectural simulation with 8 simulated threads by an average factor of 220x at an average error of 0.5 percent and a maximum error of 7.9 percent.
Thomas Grass, Trevor E. Carlson, Alejandro Rico, Germán Ceballos, Eduard Ayguadé, Marc Casas, Miquel Moretó
IEEE Trans. Computers4
2018 Behind the Scenes: Memory Analysis of Graphical Workloads on Tile-Based GPUs
abstract
Graphics rendering is a complex multi-step process whose data demands typically dominate memory system design in SoCs. GPUs create images by merging many simpler scenes for each frame. For performance, scenes are tiled into parallel tasks which produce different parts of the final output. This execution model results in complex memory behavior with bandwidth demands and data sharing varying over time, and which depends heavily on the structure of the application. To design systems that can efficiently accommodate and schedule these workloads we need to understand their behavior and diversity. In this work, we develop a quantitative characterization of the data demands of modern graphics rendering. Our approach uses an architecturally-independent analysis, identifying different types of data sharing present in the applications, independent of their scheduling. From this analysis, we present a limit study into the potential to improve memory system performance by tackling each type of data sharing. We see that there is the potential to reduce graphics bandwidth by 43% if we can take full advantage of data reuse between tasks and scenes within each frame. For the particularly complex benchmarks, capturing inter-task reuse alone has the potential to reduce bandwidth by 15% (up to 31%), while targeting interscene reuse could provide a savings of 60% (up to 75%). These insights provide us the opportunity to understand where we should focus design efforts on graphics memory systems.
Germán Ceballos, Andreas Sembrant, Trevor E. Carlson, David Black-Schaffer
ISPASS1
2018 Analyzing performance variation of task schedulers with TaskInsight
Germán Ceballos, Thomas Grass, Andra Hugo, David Black-Schaffer
Parallel Comput.1