VLDB 2026 Research / reviewers in the wild / expert
Samuel Naffziger
dblp:68/10253 · also Sam Naffziger, Samuel D. Naffziger
· DBLP profile ↗
7ranked-venue papers
2as first author
3since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Realizing the AMD Exascale Heterogeneous Processor Vision : Industry ProductabstractAMD had previously detailed its exascale research journey from initial targets and requirements to the development and evolution of its vision of a high-performance computing (HPC) accelerated processing unit (APU), dubbed the Exascale Heterogeneous Processor or EHP. At the conclusion of that work, the learnings were integrated into the design of the node architecture that went into the Frontier supercomputer, the world’s first exascale machine. However, while the Frontier node architecture embodied many of the attributes of the EHP concept, advanced heterogeneous integration capabilities at the time were not yet sufficiently mature to realize our vision of a fully-integrated APU for HPC and AI. In this paper, we finish the EHP’s story by digging deeper into why an APU was not the right solution at the time of our first exascale architecture, what the shortcomings were of previous EHP concepts, and how AMD further evolved the concept into the AMD Instinct™ MI300A APU. MI300A is the culmination of years of AMD developments in advanced packaging technologies, its APU hardware and software, and the next step in our highly effective chiplet strategy to not only deliver a groundbreaking design for exascale computing, but to also meet the demands of new large-language model and generative AI applications. Alan Smith 0003, Gabriel H. Loh, Michael J. Schulte, Mike Ignatowski, Samuel Naffziger, Mike Mantor, Nathan Kalyanasundharam, Vamsi Alla, Nicholas Malaya, Joseph L. Greathouse, Eric Chapman, Raja Swaminathan |
ISCA | 5 |
| 2021 | Understanding Chiplets Today to Anticipate Future Integration Opportunities and LimitsabstractChiplet-based architectures have recently started attracting a lot of attention, and we are seeing real-world architectures utilizing chiplet technologies in high-volume commercial production in multiple mainstream markets. In this special session paper, we provide a technical overview of the current state of chiplet technology including its benefits and limitations. This provides background and grounding in the current state-of-the-art and also lays out a range of technical areas to consider for the remaining forward-looking papers in this special session. We discuss the benefits and costs of different approaches to splitting and modularizing a monolithic chip into chiplets. In particular, we cover supporting high bandwidth and low latency communication between the die, mixed integration of multiple process technology nodes, and silicon and IP reuse. We then explore future challenges for chiplet architectures looking into the next decade of innovation. Gabriel H. Loh, Samuel Naffziger, Kevin Lepak |
DATE | 2 |
| 2021 | Pioneering Chiplet Technology and Design for the AMD EPYC™ and Ryzen™ Processor Families : Industrial ProductabstractFor decades, Moore’s Law has delivered the ability to integrate an exponentially increasing number of devices in the same silicon area at a roughly constant cost. This has enabled tremendous levels of integration, where the capabilities of computer systems that previously occupied entire rooms can now fit on a single integrated circuit.In recent times, the steady drum beat of Moore’s Law has started to slow down. Whereas device density historically doubled every 18-24 months, the rate of recent silicon process advancements has declined. While improvements in device scaling continue, albeit at a reduced pace, the industry is simultaneously observing increases in manufacturing costs.In response, the industry is now seeing a trend toward reversing direction on the traditional march toward more integration. Instead, multiple industry and academic groups are advocating that systems on chips (SoCs) be "disintegrated" into multiple smaller "chiplets." This paper details the technology challenges that motivated AMD to use chiplets, the technical solutions we developed for our products, and how we expanded the use of chiplets from individual processors to multiple product families. Samuel Naffziger, Noah Beck, Thomas Burd, Kevin Lepak, Gabriel H. Loh, Mahesh Subramony, Sean White |
ISCA | 1 |
| 2016 | Unified Power Frequency Model FrameworkabstractThis paper describes a unified power-frequency model (UPFM) which combines analytical and empirical approaches to ensure a high degree of modeling flexibility and accuracy to measured silicon (Si) results. On one end, System-on-a-Chip (SoC) design teams focus the bulk of their efforts on using detailed low-level models to verify power consumption. Such models are available late in the design cycle, and often limited in number of workloads that can be evaluated. On the other end, FPGA-based modeling and spreadsheet approaches that operate on higher-level abstraction have been proposed. However these are often limited by poor correlation to measured Si results. In addition, extant models typically focus on power projection or prediction of Si speed but not both. A unified approach is much needed since SOCs today have to meet stringent power and performance constraints simultaneously. The proposed UPFM model overcomes these limitations. First actual measured Si results serve as the empirical baseline foundation for projections so that simulated vs. measured differences can be calibrated. Second, each IP is analytically modeled using a large number of relevant parameters. This high level of abstraction allows for the model to be useful from early design cycle all the way to the mature phase (parameters get refined over time). Also, wide-ranging parameters have been carefully chosen (and improved over multiple product generations) so that accuracy is not sacrificed. We demonstrate UPFM as a comprehensive framework where technology, architecture and infrastructure (test/thermal) choices can be modeled with high accuracy and drive optimal perf-per-watt SoC designs. Sriram Sundaram, Warren He, Sriram Sambamurthy, Aaron Grenat, Steven Liepe, Samuel Naffziger |
ISLPED | 6 |
| 2014 | Welcome program chairsabstractWelcome to HOT CHIPS 26, the 26th anniversary conference. A look back over these 26 years shows the amazing progress made in computer engineering. This progress can be measured in terms of performance, cost, density and versatility of function. These technological advances have brought advances in functionality, creating new companies, new markets and indeed new patterns of socialization. Samuel Naffziger, Gurindar S. Sohi |
Hot Chips Symposium | 1 |
| 2003 | Correction to "statistical clock skew modeling with data delay variations"
Samuel Naffziger |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2001 | Statistical clock skew modeling with data delay variationsabstractAccurate clock skew budgets are important for microprocessor designers to avoid hold-time failures and to properly allocate resources when optimizing global and local paths. Many published clock skew budgets neglect voltage jitter and process variation, which are becoming dominant factors in otherwise balanced H-trees. However, worst-case process variation assumptions are severely pessimistic. This paper describes the major sources of clock skew in a microprocessor using a modified H-tree and applies the model to a second-generation Itanium-M processor family microprocessor currently under design. Monte Carlo simulation is used to develop statistical clock skew budgets for setup and hold time constraints in a four-level skew hierarchy. Voltage jitter through the phase locked loop (PLL) and clock buffers accounts for the majority of skew budgets. We show that taking into account the number of nearly critical paths between clocked elements at each level of the skew hierarchy and variations in the data delays of these paths reduces the difference between global and local skew budgets by more than a factor of two. Another insight is that data path delay variability limits the potential cycle-time benefits of active deskew circuits because the paths with the worst skew are unlikely to also be the paths with the longest data delays. Samuel Naffziger |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |