Silvio Heverton Campelo de Santana

dblp:196/4222 · DBLP profile ↗
← Back
3ranked-venue papers
3as first author
3since 2021 · last 2026
0009-0005-1693-9033ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 3 · 3 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Chips Need DIP: Time-Proportional Per-Instruction Cycle Stacks at Dispatch
Silvio Heverton Campelo de Santana, Joseph Rogers, Lieven Eeckhout, Magnus Jahre
ASPLOS (2)1
2026 FYI: A Foundational Counter Architecture for Building CPI Stacks on Out-of-Order Processors
abstract
The end of Dennard scaling and the imminent end of Moore’s law is making it increasingly critical to write software that fully utilizes the hardware compute resources. Software developers hence need to gain insight into the inefficiencies of their applications, and a typical first step is to obtain applicationlevel Cycles-Per-Instruction (CPI) stacks, for example by using a Top-Down methodology. Unfortunately, existing approaches have only been validated against processors with idealized microarchitectural resources. Although this approach adds confidence that the generated CPI stacks make intuitive sense, it does not conclusively document that the CPI stacks break down total execution time into components in an architecturally sound manner. Our goal in this paper is to bridge this knowledge gap. We first define a complete and mutually exclusive reference that accounts for all dispatch slots in all clock cycles across all microoperations ($\mu$ ops), i.e., a $\mu$ op is either (i) dispatching, (ii) stalled while waiting for a structural back-end stall, (iii) not available due to a front-end issue, or (iv) mistakenly executed due to misspeculation. If the reference is also architecturally sound, i.e., it attributes each dispatch slot to the CPI stack component that accurately represents the impact on effective dispatch bandwidth, we refer to it as foundational; second-order overlap effects typically mean that an architecture has multiple foundational references. We derive foundational references for the BOOM out-of-order core, and, surprisingly perhaps, find that they can be implemented in hardware with less than 1 KB of additional state. We hence propose the Foundational Yet Implementable (FYI) counter architecture for CPI stacks that are, by design, an exact match for the target reference.
Silvio Heverton Campelo de Santana, Lieven Eeckhout, Magnus Jahre
ISPASS1
2026 Pesto: Diagnosing Performance Pathologies in Out-of-Order Processors
abstract
High-performance parallel acceleration necessitates high-performance CPUs to stave off Amdahl’s law. Getting the most out of a CPU requires instruction-level performance profiling to understand microarchitecture performance bottlenecks, guiding hardware design for future generations. This depends on effective methods for using instruction profilers to identify the key ways performance losses occur in out-of-order processors. Identifying the root causes of these losses is challenging, but can be greatly assisted by comprehending how they manifest behaviorally in hardware, which we refer to in this work as ‘performance pathologies’. Historically, characterizing these behaviors has been a labor-intensive process that required a detailed understanding of software and microarchitecture. However, the advent of profilers capable of producing accurate instruction-level cycle stacks at both ends of the execution window, i.e., at dispatch and commit, dramatically simplifies this task. In this work, we present Pesto, an instruction-level methodology for diagnosing microarchitecture performance pathologies, i.e., identifying the key microarchitectural manifestations of performance loss in out-of-order processors. Pesto uses k-means clustering to group per-instruction cycle stack components produced at both ends of the execution window. The resulting centroids for each cluster characterize exactly one kind of pathology (e.g., pipeline flush due to branch misprediction, or commit stall due to data cache misses). We apply Pesto to an FPGA-accelerated BOOM core in FireSim running SPEC2017 and discover that the BOOM core has 20 pathologies, corresponding to 20 fundamental ways its microarchitecture can lose performance. Furthermore, we observe that individual instruction types (e.g., loads) lead to performance loss in at most four ways. Finally, we show that Pesto works using statistical sampling, making it easy to adopt.
Silvio Heverton Campelo de Santana, Joseph Rogers, Lieven Eeckhout, Magnus Jahre
ISPASS1