Archit Vasan

dblp:360/3337 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2025
0000-0002-8299-1033ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
High-performance computing · 77% GPUs and heterogeneous computing · 23%

Topics — the 2 heaviest of 2, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing
scientific computing systems
0.812024
MProt-DPO: Breaking the ExaFLOPS Barrier for Multimodal Protein Design Workflows with Direct Preference Optimization · SC 2024
GPUs and heterogeneous computing
GPU and heterogeneous computing
0.212024
MProt-DPO: Breaking the ExaFLOPS Barrier for Multimodal Protein Design Workflows with Direct Preference Optimization · SC 2024

Methods — techniques the papers use, named apart from their topics

multimodal generative models · 0.8mixed precision · 0.8direct preference optimization · 0.8
YearPublicationVenuePosition
2025 AI and HPC Applications on Leadership Computing Platforms: Performance and Scalability Studies
abstract
As HPC systems move into the exascale era an increasing diversity of processing hardware is being deployed. The last decade saw the ascendance of NVIDIA GPU-accelerated systems among the largest scale HPC systems and spurred the need for application developers to consider approaches to performance portability that preserved developer productivity. This challenge has been compounded in the last several years by the introduction of the first two exascale systems, Frontier and Aurora (\#2 and \#3 on the November 2024 Top 500 list respectively). These systems utilize new and different GPUs, with the AMD MI-250X GPU on Frontier and the Intel Data Center GPU Max 1550 on Aurora. This study investigates the performance and qualitative performance portability of$\mathbf{1 2}$HPC and ML applications on three large scale HPC systems that utilize GPUs from the three different vendors: Frontier (AMD), Aurora (Intel), and Polaris (NVIDIA A100). The performance of these applications is evaluated on single GPU, single node, and multinode scales on each of the systems. We show that the figures-of-merit (FOMs) of the applications on a single GPU of Aurora and Frontier ranged from$0.9-4 x$and$0.8-2.5 x$, respectively, the performance on a GPU of Polaris. We also show that the FOMs on a single node of Aurora and Frontier ranged from 1.3-6.3x and 0.8-2.6x, respectively, a single node of Polaris. The applications were scaled up to 512 nodes showing good scaling efficiency across the board. Finally, we discuss useful concepts and experiences gained in running diverse applications on diverse HPC systems.
JaeHyuk Kwack, Colleen Bertoni, Umesh Unnikrishnan, Riccardo Balin, Khalid Hossain, Yasaman Ghadar, Timothy J. Williams, Abhishek Bagusetty, Mathialakan Thavappiragasam, Väinö Hatanpää, Archit Vasan, John R. Tramm, Scott Parker
IPDPS11
2024 MProt-DPO: Breaking the ExaFLOPS Barrier for Multimodal Protein Design Workflows with Direct Preference Optimization
abstract
We present a scalable, end-to-end workflow for protein design. By augmenting protein sequences with natural language descriptions of their biochemical properties, we train generative models that can be preferentially aligned with protein fitness landscapes. Through complex experimental-and simulation-based observations, we integrate these measures as preferred parameters for generating new protein variants and demonstrate our workflow on five diverse supercomputers. We achieve >1 ExaFLOPS sustained performance in mixed precision on each supercomputer and a maximum sustained performance of 4.11 Ex-aFLOPS and peak performance of 5.57 ExaFLOPS. We establish the scientific performance of our model on two tasks: (1) across a predetermined benchmark dataset of deep mutational scanning experiments to optimize the fitness-determining mutations in the yeast protein HIS7, and (2) in optimizing the design of the enzyme malate dehydrogenase to achieve lower activation barriers (and therefore increased catalytic rates) using simulation data. Our implementation thus sets high watermarks for multimodal protein design workflows.
Gautham Dharuman, Kyle Hippe, Alex Brace, Sam Foreman, Väinö Hatanpää, Varuni Sastry 0001, Huihuo Zheng, Logan T. Ward, Servesh Muralidharan, Archit Vasan, Bharat Kale, Carla M. Mann, Yun-Hsuan Cheng, Yuliana Zamora, Shengchao Liu, Chaowei Xiao, Murali Emani, Tom Gibbs, Mahidhar Tatineni, Deepak Canchi, Jerome Mitchell, Koichi Yamada, María Jesús Garzarán, Michael E. Papka, Ian T. Foster, Rick L. Stevens, Anima Anandkumar, Venkatram Vishwanath, Arvind Ramanathan
SC10