EDBT 2026 Demo / reviewers in the wild / expert
Servesh Muralidharan
dblp:125/7662
· DBLP profile ↗
4ranked-venue papers
0as first author
2since 2021 · last 2025
0000-0002-3541-5658ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
High-performance computing · 31% Processor architecture and microarchitecture · 24% Storage systems · 12% | |
| Software engineering, system software, and programming languages
1 paper |
Operating systems · 44% Runtime systems and virtual machines · 44% Programming languages and type systems · 13% |
Topics — the 10 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
High-performance computing
scientific computing systems |
0.8 | 1 | 2024 | MProt-DPO: Breaking the ExaFLOPS Barrier for Multimodal Protein Design Workflows with Direct Preference Optimization · SC 2024 |
Performance modeling and evaluation
benchmarking |
0.3 | 1 | 2017 | Efficient Multibyte Floating Point Data Formats Using Vectorization · IEEE Trans. Computers 2017 |
Storage systems
data representation |
0.3 | 1 | 2017 | Efficient Multibyte Floating Point Data Formats Using Vectorization · IEEE Trans. Computers 2017 |
Processor architecture and microarchitecture
instruction set architecture |
0.3 | 1 | 2017 | Efficient Multibyte Floating Point Data Formats Using Vectorization · IEEE Trans. Computers 2017 |
Hardware accelerators and domain-specific architectures › machine learning accelerator › low-precision arithmetic
low-precision floating point |
0.3 | 1 | 2017 | Efficient Multibyte Floating Point Data Formats Using Vectorization · IEEE Trans. Computers 2017 |
Processor architecture and microarchitecture › SIMD
vector instructions |
0.3 | 1 | 2017 | Efficient Multibyte Floating Point Data Formats Using Vectorization · IEEE Trans. Computers 2017 |
GPUs and heterogeneous computing
GPU and heterogeneous computing |
0.2 | 1 | 2024 | MProt-DPO: Breaking the ExaFLOPS Barrier for Multimodal Protein Design Workflows with Direct Preference Optimization · SC 2024 |
Operating systems › resource management › process management
context switching |
0.2 | 1 | 2013 | Compiler support for lightweight context switching · ACM Trans. Archit. Code Optim. 2013 |
Runtime systems and virtual machines › parallel runtime systems
lightweight threads |
0.2 | 1 | 2013 | Compiler support for lightweight context switching · ACM Trans. Archit. Code Optim. 2013 |
Programming languages and type systems
language implementation |
0.0 | 1 | 2013 | Compiler support for lightweight context switching · ACM Trans. Archit. Code Optim. 2013 |
Methods — techniques the papers use, named apart from their topics
multimodal generative models · 0.8mixed precision · 0.8direct preference optimization · 0.8vectorization · 0.3work-stealing scheduler · 0.2LLVM · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | THAPI: Tracing Heterogeneous APIs
Solomon Abera, Aurelio Vivas, Thomas Applencourt, Servesh Muralidharan, Bryce Allen, Kazutomo Yoshii, Swann Perarnau, Brice Videau |
Euro-Par (1) | 4 |
| 2024 | MProt-DPO: Breaking the ExaFLOPS Barrier for Multimodal Protein Design Workflows with Direct Preference OptimizationabstractWe present a scalable, end-to-end workflow for protein design. By augmenting protein sequences with natural language descriptions of their biochemical properties, we train generative models that can be preferentially aligned with protein fitness landscapes. Through complex experimental-and simulation-based observations, we integrate these measures as preferred parameters for generating new protein variants and demonstrate our workflow on five diverse supercomputers. We achieve >1 ExaFLOPS sustained performance in mixed precision on each supercomputer and a maximum sustained performance of 4.11 Ex-aFLOPS and peak performance of 5.57 ExaFLOPS. We establish the scientific performance of our model on two tasks: (1) across a predetermined benchmark dataset of deep mutational scanning experiments to optimize the fitness-determining mutations in the yeast protein HIS7, and (2) in optimizing the design of the enzyme malate dehydrogenase to achieve lower activation barriers (and therefore increased catalytic rates) using simulation data. Our implementation thus sets high watermarks for multimodal protein design workflows. Gautham Dharuman, Kyle Hippe, Alex Brace, Sam Foreman, Väinö Hatanpää, Varuni Sastry 0001, Huihuo Zheng, Logan T. Ward, Servesh Muralidharan, Archit Vasan, Bharat Kale, Carla M. Mann, Yun-Hsuan Cheng, Yuliana Zamora, Shengchao Liu, Chaowei Xiao, Murali Emani, Tom Gibbs, Mahidhar Tatineni, Deepak Canchi, Jerome Mitchell, Koichi Yamada, María Jesús Garzarán, Michael E. Papka, Ian T. Foster, Rick L. Stevens, Anima Anandkumar, Venkatram Vishwanath, Arvind Ramanathan |
SC | 9 |
| 2017 | Efficient Multibyte Floating Point Data Formats Using VectorizationabstractWe propose a scheme for reduced-precision representation of floating point data on a continuum between IEEE-754 floating point types. Our scheme enables the use of lower precision formats for a reduction in storage space requirements and data transfer volume. We describe how our scheme can be accelerated using existing hardware vector units on two general-purpose processor (GPP) microarchitectures (Intel Ivy Bridge and Haswell), as well as on a numerical accelerator (Intel Xeon Phi). Our evaluation demonstrates that supporting reduced precision by exploiting native vector instructions can yield a low overhead custom-precision floating point solution that does not require specialized hardware support. In our experiments we find cases where our scheme is actually faster than native floating point types where the underlying vector instruction set supports efficient byte-level permutations. Andrew Anderson 0001, Servesh Muralidharan, David Gregg |
IEEE Trans. Computers | 2 |
| 2013 | Compiler support for lightweight context switchingabstractWe propose a new language-neutral primitive for the LLVM compiler, which provides efficient context switching and message passing between lightweight threads of control. The primitive, called Swapstack, can be used by any language implementation based on LLVM to build higher-level language structures such as continuations, coroutines, and lightweight threads. As part of adding the primitives to LLVM, we have also added compiler support for passing parameters across context switches. Our modified LLVM compiler produces highly efficient code through a combination of exposing the context switching code to existing compiler optimizations, and adding novel compiler optimizations to further reduce the cost of context switches. To demonstrate the generality and efficiency of our primitives, we add one-shot continuations to C++, and provide a simple fiber library that allows millions of fibers to run on multiple cores, with a work-stealing scheduler and fast inter-fiber sychronization. We argue that compiler-supported lightweight context switching can be significantly faster than using a library to switch between contexts, and provide experimental evidence to support the position. Stephen Dolan, Servesh Muralidharan, David Gregg |
ACM Trans. Archit. Code Optim. | 2 |