Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Nathan Kalyanasundharam

dblp:58/879 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
1since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Integrated circuit design · 30% GPUs and heterogeneous computing · 30% High-performance computing · 30%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
GPUs and heterogeneous computing › heterogeneous architecture
accelerated processing unit
0.812024
Realizing the AMD Exascale Heterogeneous Processor Vision : Industry Product · ISCA 2024
High-performance computing › supercomputing
exascale computing
0.812024
Realizing the AMD Exascale Heterogeneous Processor Vision : Industry Product · ISCA 2024
Integrated circuit design
heterogeneous integration
0.812024
Realizing the AMD Exascale Heterogeneous Processor Vision : Industry Product · ISCA 2024
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.212024
Realizing the AMD Exascale Heterogeneous Processor Vision : Industry Product · ISCA 2024

Methods — techniques the papers use, named apart from their topics

chiplet integration · 0.8advanced packaging · 0.8
YearPublicationVenuePosition
2024 Realizing the AMD Exascale Heterogeneous Processor Vision : Industry Product
abstract
AMD had previously detailed its exascale research journey from initial targets and requirements to the development and evolution of its vision of a high-performance computing (HPC) accelerated processing unit (APU), dubbed the Exascale Heterogeneous Processor or EHP. At the conclusion of that work, the learnings were integrated into the design of the node architecture that went into the Frontier supercomputer, the world’s first exascale machine. However, while the Frontier node architecture embodied many of the attributes of the EHP concept, advanced heterogeneous integration capabilities at the time were not yet sufficiently mature to realize our vision of a fully-integrated APU for HPC and AI. In this paper, we finish the EHP’s story by digging deeper into why an APU was not the right solution at the time of our first exascale architecture, what the shortcomings were of previous EHP concepts, and how AMD further evolved the concept into the AMD Instinct™ MI300A APU. MI300A is the culmination of years of AMD developments in advanced packaging technologies, its APU hardware and software, and the next step in our highly effective chiplet strategy to not only deliver a groundbreaking design for exascale computing, but to also meet the demands of new large-language model and generative AI applications.
Alan Smith 0003, Gabriel H. Loh, Michael J. Schulte, Mike Ignatowski, Samuel Naffziger, Mike Mantor, Nathan Kalyanasundharam, Vamsi Alla, Nicholas Malaya, Joseph L. Greathouse, Eric Chapman, Raja Swaminathan
ISCA7
2009 Blade computing with the AMD Opteron™ processor ("magny-cours")
Pat Conway, Nathan Kalyanasundharam, Gregg Donley, Kevin Lepak, Bill Hughes
Hot Chips Symposium2
1999 Simultaneous Switching Noise Considerations in the Design of a High Speed, Multiported TLB of a Server-Class Microprocessor
abstract
Noise introduced on the supply networks by simultaneous switching of the nodes of complex macros is becoming an important issue in very deep sub-micron technologies. Server-class microprocessors demand design of very high speed multiported macros that generate high peak currents and current slew-rates. This paper investigates one such macro: a fully-associative, multiported data TLB. Our simulations show a slowdown of 10-20% due to the supply noise despite robust C4-based supply network. The traditional solution of employing decoupling capacitors to combat the supply noise results in an unacceptable area increase. Macro design techniques that can reduce peak current and current slew rate without reducing the speed of critical path are proposed. Employment of hierarchical match line and delayed split precharge techniques reduce the SSN and the required decoupling capacitance by a factor of 5x.
Nathan Kalyanasundharam, Nital Patwa
ICCD1