Servesh Muralidharan

dblp:125/7662 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
2since 2021 · last 2025
0000-0002-3541-5658ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
High-performance computing · 31% Processor architecture and microarchitecture · 24% Storage systems · 12%
Software engineering, system software, and programming languages
1 paper
Operating systems · 44% Runtime systems and virtual machines · 44% Programming languages and type systems · 13%

Topics — the 10 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing
scientific computing systems
0.812024
MProt-DPO: Breaking the ExaFLOPS Barrier for Multimodal Protein Design Workflows with Direct Preference Optimization · SC 2024
Performance modeling and evaluation
benchmarking
0.312017
Efficient Multibyte Floating Point Data Formats Using Vectorization · IEEE Trans. Computers 2017
Storage systems
data representation
0.312017
Efficient Multibyte Floating Point Data Formats Using Vectorization · IEEE Trans. Computers 2017
Processor architecture and microarchitecture
instruction set architecture
0.312017
Efficient Multibyte Floating Point Data Formats Using Vectorization · IEEE Trans. Computers 2017
Hardware accelerators and domain-specific architectures › machine learning accelerator › low-precision arithmetic
low-precision floating point
0.312017
Efficient Multibyte Floating Point Data Formats Using Vectorization · IEEE Trans. Computers 2017
Processor architecture and microarchitecture › SIMD
vector instructions
0.312017
Efficient Multibyte Floating Point Data Formats Using Vectorization · IEEE Trans. Computers 2017
GPUs and heterogeneous computing
GPU and heterogeneous computing
0.212024
MProt-DPO: Breaking the ExaFLOPS Barrier for Multimodal Protein Design Workflows with Direct Preference Optimization · SC 2024
Operating systems › resource management › process management
context switching
0.212013
Compiler support for lightweight context switching · ACM Trans. Archit. Code Optim. 2013
Runtime systems and virtual machines › parallel runtime systems
lightweight threads
0.212013
Compiler support for lightweight context switching · ACM Trans. Archit. Code Optim. 2013
Programming languages and type systems
language implementation
0.012013
Compiler support for lightweight context switching · ACM Trans. Archit. Code Optim. 2013

Methods — techniques the papers use, named apart from their topics

multimodal generative models · 0.8mixed precision · 0.8direct preference optimization · 0.8vectorization · 0.3work-stealing scheduler · 0.2LLVM · 0.2
YearPublicationVenuePosition
2025 THAPI: Tracing Heterogeneous APIs
Solomon Abera, Aurelio Vivas, Thomas Applencourt, Servesh Muralidharan, Bryce Allen, Kazutomo Yoshii, Swann Perarnau, Brice Videau
Euro-Par (1)4
2024 MProt-DPO: Breaking the ExaFLOPS Barrier for Multimodal Protein Design Workflows with Direct Preference Optimization
abstract
We present a scalable, end-to-end workflow for protein design. By augmenting protein sequences with natural language descriptions of their biochemical properties, we train generative models that can be preferentially aligned with protein fitness landscapes. Through complex experimental-and simulation-based observations, we integrate these measures as preferred parameters for generating new protein variants and demonstrate our workflow on five diverse supercomputers. We achieve >1 ExaFLOPS sustained performance in mixed precision on each supercomputer and a maximum sustained performance of 4.11 Ex-aFLOPS and peak performance of 5.57 ExaFLOPS. We establish the scientific performance of our model on two tasks: (1) across a predetermined benchmark dataset of deep mutational scanning experiments to optimize the fitness-determining mutations in the yeast protein HIS7, and (2) in optimizing the design of the enzyme malate dehydrogenase to achieve lower activation barriers (and therefore increased catalytic rates) using simulation data. Our implementation thus sets high watermarks for multimodal protein design workflows.
Gautham Dharuman, Kyle Hippe, Alex Brace, Sam Foreman, Väinö Hatanpää, Varuni Sastry 0001, Huihuo Zheng, Logan T. Ward, Servesh Muralidharan, Archit Vasan, Bharat Kale, Carla M. Mann, Yun-Hsuan Cheng, Yuliana Zamora, Shengchao Liu, Chaowei Xiao, Murali Emani, Tom Gibbs, Mahidhar Tatineni, Deepak Canchi, Jerome Mitchell, Koichi Yamada, María Jesús Garzarán, Michael E. Papka, Ian T. Foster, Rick L. Stevens, Anima Anandkumar, Venkatram Vishwanath, Arvind Ramanathan
SC9
2017 Efficient Multibyte Floating Point Data Formats Using Vectorization
abstract
We propose a scheme for reduced-precision representation of floating point data on a continuum between IEEE-754 floating point types. Our scheme enables the use of lower precision formats for a reduction in storage space requirements and data transfer volume. We describe how our scheme can be accelerated using existing hardware vector units on two general-purpose processor (GPP) microarchitectures (Intel Ivy Bridge and Haswell), as well as on a numerical accelerator (Intel Xeon Phi). Our evaluation demonstrates that supporting reduced precision by exploiting native vector instructions can yield a low overhead custom-precision floating point solution that does not require specialized hardware support. In our experiments we find cases where our scheme is actually faster than native floating point types where the underlying vector instruction set supports efficient byte-level permutations.
Andrew Anderson 0001, Servesh Muralidharan, David Gregg
IEEE Trans. Computers2
2013 Compiler support for lightweight context switching
abstract
We propose a new language-neutral primitive for the LLVM compiler, which provides efficient context switching and message passing between lightweight threads of control. The primitive, called Swapstack, can be used by any language implementation based on LLVM to build higher-level language structures such as continuations, coroutines, and lightweight threads. As part of adding the primitives to LLVM, we have also added compiler support for passing parameters across context switches. Our modified LLVM compiler produces highly efficient code through a combination of exposing the context switching code to existing compiler optimizations, and adding novel compiler optimizations to further reduce the cost of context switches. To demonstrate the generality and efficiency of our primitives, we add one-shot continuations to C++, and provide a simple fiber library that allows millions of fibers to run on multiple cores, with a work-stealing scheduler and fast inter-fiber sychronization. We argue that compiler-supported lightweight context switching can be significantly faster than using a library to switch between contexts, and provide experimental evidence to support the position.
Stephen Dolan, Servesh Muralidharan, David Gregg
ACM Trans. Archit. Code Optim.2