EDBT 2026 Demo / reviewers in the wild / expert
Anna Klein
dblp:133/3700
· DBLP profile ↗
2ranked-venue papers
0as first author
1since 2021 · last 2022
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 2 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Performance modeling and evaluation · 61% GPUs and heterogeneous computing · 39% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Performance modeling and evaluation › workload characterization
AI workload characterization |
0.6 | 1 | 2022 | AI-Enabling Workloads on Large-Scale GPU-Accelerated System: Characterization, Opportunities, and Implications · HPCA 2022 |
GPUs and heterogeneous computing › GPU computing
GPU-accelerated systems |
0.6 | 1 | 2022 | AI-Enabling Workloads on Large-Scale GPU-Accelerated System: Characterization, Opportunities, and Implications · HPCA 2022 |
Performance modeling and evaluation
workload characterization |
0.6 | 1 | 2022 | AI-Enabling Workloads on Large-Scale GPU-Accelerated System: Characterization, Opportunities, and Implications · HPCA 2022 |
GPUs and heterogeneous computing
GPU computing |
0.2 | 1 | 2022 | AI-Enabling Workloads on Large-Scale GPU-Accelerated System: Characterization, Opportunities, and Implications · HPCA 2022 |
Methods — techniques the papers use, named apart from their topics
user behavior analysis · 0.6job trace analysis · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | AI-Enabling Workloads on Large-Scale GPU-Accelerated System: Characterization, Opportunities, and ImplicationsabstractProduction high-performance computing (HPC) systems are adopting and integrating GPUs into their design to accommodate artificial intelligence (AI), machine learning, and data visualization workloads. To aid with the design and operations of new and existing GPU-based large-scale systems, we provide a detailed characterization of system operations, job characteristics, user behavior, and trends on a contemporary GPU-accelerated production HPC system. Our insights indicate that the pre-mature phases in modern AI workflow take up significant GPU hours while underutilizing GPUs, which opens up the opportunity for a multi-tier system. Finally, we provide various potential recommendations and areas for future investment for system architects, operators, and users. Baolin Li 0001, Rohin Arora, Siddharth Samsi, Tirthak Patel, William Arcand, David Bestor, Chansup Byun, Rohan Basu Roy, Bill Bergeron, John T. Holodnak, Michael Houle 0001, Matthew Hubbell, Michael Jones 0001, Jeremy Kepner, Anna Klein, Peter Michaleas, Joseph McDonald, Lauren Milechin, Julie Mullen, Andrew Prout, Benjamin Price, Albert Reuther, Antonio Rosa, Matthew L. Weiss, Charles Yee, Daniel Edelman, Allan Vanterpool, Anson Cheng, Vijay Gadepally, Devesh Tiwari |
HPCA | 15 |
| 2013 | P-sync: A Photonically Enabled Architecture for Efficient Non-local Data AccessabstractCommunication in multi- and many-core processors has long been a bottleneck to performance due to the high cost of long-distance electrical transmission. This difficulty has been partially remedied by architectural constructs such as caches and novel interconnect topologies, albeit at a steep cost in terms of complexity. Unfortunately, even these measures are rendered ineffective by certain kinds of communication, most notably scatter and gather operations that exhibit highly nonlocal data access patterns. Much work has gone into examining how the increased bandwidth density afforded by chip-scale silicon photonic interconnect technologies affects computing, but photonics have additional properties that can be leveraged to greatly accelerate performance and energy efficiency under such difficult loads. This paper describes a novel synchronized global photonic bus and system architecture called P-sync that uses photonics' distance independence to greatly improve performance on many important applications previously limited by electronic interconnect. The architecture is evaluated in the context of a non-local yet common application: the distributed Fast Fourier Transform. We show that it is possible to achieve high efficiency by tightly balancing computation and communication latency in P-sync and achieve upwards of a 6× performance increase on gather patterns, even when bandwidth is equalized. David Whelihan, Jeffrey J. Hughes, Scott M. Sawyer, Eric Robinson, Michael M. Wolf, Sanjeev Mohindra, Julie Mullen, Anna Klein, Michelle S. Beard, Nadya Bliss, Johnnie Chan, Robert Hendry, Keren Bergman, Luca P. Carloni |
IPDPS | 8 |