Joseph Rogers

dblp:206/5245 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
5since 2021 · last 2026
0009-0002-6172-593XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 first-author · 4 since 2021Software engineering, systems software and programming languages · 4 · 2 first-author · 4 since 2021
YearPublicationVenuePosition
2026 Chips Need DIP: Time-Proportional Per-Instruction Cycle Stacks at Dispatch
Silvio Heverton Campelo de Santana, Joseph Rogers, Lieven Eeckhout, Magnus Jahre
ASPLOS (2)2
2026 Pesto: Diagnosing Performance Pathologies in Out-of-Order Processors
abstract
High-performance parallel acceleration necessitates high-performance CPUs to stave off Amdahl’s law. Getting the most out of a CPU requires instruction-level performance profiling to understand microarchitecture performance bottlenecks, guiding hardware design for future generations. This depends on effective methods for using instruction profilers to identify the key ways performance losses occur in out-of-order processors. Identifying the root causes of these losses is challenging, but can be greatly assisted by comprehending how they manifest behaviorally in hardware, which we refer to in this work as ‘performance pathologies’. Historically, characterizing these behaviors has been a labor-intensive process that required a detailed understanding of software and microarchitecture. However, the advent of profilers capable of producing accurate instruction-level cycle stacks at both ends of the execution window, i.e., at dispatch and commit, dramatically simplifies this task. In this work, we present Pesto, an instruction-level methodology for diagnosing microarchitecture performance pathologies, i.e., identifying the key microarchitectural manifestations of performance loss in out-of-order processors. Pesto uses k-means clustering to group per-instruction cycle stack components produced at both ends of the execution window. The resulting centroids for each cluster characterize exactly one kind of pathology (e.g., pipeline flush due to branch misprediction, or commit stall due to data cache misses). We apply Pesto to an FPGA-accelerated BOOM core in FireSim running SPEC2017 and discover that the BOOM core has 20 pathologies, corresponding to 20 fundamental ways its microarchitecture can lose performance. Furthermore, we observe that individual instruction types (e.g., loads) lead to performance loss in at most four ways. Finally, we show that Pesto works using statistical sampling, making it easy to adopt.
Silvio Heverton Campelo de Santana, Joseph Rogers, Lieven Eeckhout, Magnus Jahre
ISPASS2
2025 HILP: Accounting for Workload-Level Parallelism in System-on-Chip Design Space Exploration
abstract
High-performance System-on-Chip (SoC) architectures are becoming increasingly complex and heterogeneous, and the days when a single application could utilize all of an SoC’s hardware resources are all but over. The SoC’s workload, i.e., the set of independent applications that the SoC typically executes, therefore has a significant impact on its efficiency. Accounting for Workload-Level Parallelism (WLP) in early-stage design space exploration is thus critical as later-stage analysis steps must focus on favorable design points to yield optimal results. Unfortunately, state-of-the-art MultiAmdahl and Gables fall short because they only model the extremes of minimal and maximal WLP. We hence propose HILP, the first early-stage design space exploration approach for heterogeneous SoCs that fully accounts for WLP. Our key observation is that scheduling a workload of independent multi-phase applications on a heterogeneous SoC is an instance of the classic job-shop scheduling optimization problem and thus can be solved using integer linear programming. HILP therefore uses a high-performance integer linear programming solver to find a near-optimal schedule that minimizes the overall execution time of the workload, i.e., it schedules the dependent phases of all applications in the workload on the cores and accelerators of the target SoC to maximize performance while respecting power consumption and memory bandwidth constraints. We validate HILP by demonstrating that it captures the performance effects of Amdahl’s law, the memory wall, and dark silicon, and then use it to explore the impact of WLP across a large SoC design space, yielding multiple insights. The key takeaway is that modeling WLP is necessary to ensure that more detailed, later-stage design tasks focus on the most favorable parts of the vast design space of heterogeneous SoCs.
Joseph Rogers, Lieven Eeckhout, Magnus Jahre
HPCA1
2025 Neoscope: How Resilient Is My SoC to Workload Churn?
abstract
The lifetime of hardware is increasing, but the lifetime of software is not.This leads to devices that, while performant when released, have fall-off due to changing workload suitability.To ensure that performance is maintained, computer architects must begin considering the effects of evolving workloads early in the design process.The ways that workloads can change, i.e., churn, over time profoundly affect hardware's ability to maintain performance.To better understand churn, we introduce the concepts Magnitude and Disruption, which enable how workloads change to be quantitatively described.By leveraging these terms, we present churn as a span, where the churn of a given workload can be categorized as either Minimal, Perturbing, Escalating, or Volatile.To account for how churn affects performance, we propose Neoscope, the first multi-objective pre-silicon design space exploration tool for investigating System-on-Chip (SoC) architectures that are resilient to workload churn.Neoscope uses integer linear programming and concepts from job-shop scheduling to construct near-optimal SoCs for a given workload.Unlike previous methods, Neoscope approaches (and often finds) the globally optimal SoC with a single invocation, instead of needing to parameter sweep.Neoscope is also multi-objective, i.e., it can optimize for many kinds of metrics other than absolute performance, including area, energy, cost, and carbon efficiencies.Using Neoscope, we explore near-optimal SoCs for many workload-churn configurations, and investigate how changing the objective function changes the ideal SoC.Neoscope shows that small SoCs need high levels of specialization, but that this is risky, as it increases susceptibility to churn.
Joseph Rogers, Lieven Eeckhout, Taha Soliman, Magnus Jahre
ISCA1
2024 AIO: An Abstraction for Performance Analysis Across Diverse Accelerator Architectures
abstract
Specialization is the key approach for continued performance growth beyond the end of Dennard scaling. Academics and industry are hence continuously proposing new accelerator architectures, including conventional Domain-Specific Accelerators (DSAs) and emerging Processing in Memory (PIM) accelerators. We are thus fast approaching an era in which earlystage accelerator analysis is critical for maintaining the productivity of software developers, system software designers, and computer architects — to ensure that they focus time-consuming implementation and optimization efforts on the most favorable class of accelerators for the problem at hand. Unfortunately, existing approaches fall short because they either adopt a level of abstraction that is too high — and therefore are unable to account for key performance phenomena — or too low — because they focus on details that do not generalize across diverse accelerators. Our Architecture-Independent Operation (AIO) abstraction addresses this issue by leveraging that accelerators typically focus on data-level parallelism, and an AIO is hence a key piece of algorithm-level data-parallel work that remains the same across diverse accelerators. To demonstrate that the AIO abstraction can be accurate and useful, we create the AccMe performance model which predicts kernel performance by estimating the number of clock cycles spent on compute, memory, and invocation overhead while accounting for overlap between compute and memory cycles as well as finite memory bandwidth. We demonstrate that AccMe can be accurate, i.e., it yields an average performance prediction error of $5.6 \%$ across our diverse kernels and accelerators. This is a significant improvement over the $\mathbf{2 0. 6 \%}$ average error of curve-fitted Roofline which provides the best-case accuracy of Roofline’s operational intensity abstraction. We further demonstrate that AccMe is useful through three case studies that illustrate (i) how developers can use AccMe for accelerator selection under uncertainty; (ii) how system software can use AccMe for scheduling — and thereby improve throughput by $2.8 \times$ on average compared to Roofline-driven scheduling; and (iii) how computer architects can use AccMe for architectural exploration.
Joseph Rogers, Taha Soliman, Magnus Jahre
ISCA1