Ankit Bhardwaj 0002

dblp:205/6833-2 · DBLP profile ↗
← Back
7ranked-venue papers
5as first author
6since 2021 · last 2026
0000-0002-5094-9711ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 2 first-author · 2 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Cloud and datacenter computing · 44% Storage systems · 23% Distributed systems · 17%
Software engineering, system software, and programming languages
1 paper
Operating systems · 100%
Artificial intelligence
1 paper
Efficient and distributed learning · 100%

Topics — the 12 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed systems
fault tolerance
1.012026
Checkmate: Zero Performance Overhead Model Checkpointing via Network Gradient Replication · NSDI 2026
Storage systems › flash and SSD
SSD performance
1.012026
Unleashing The Potential of Datacenter SSDs by Taming Performance Variability · NSDI 2026
Cloud and datacenter computing
cluster resource management and scheduling
0.912025
Auto-reconfiguration for Latency Minimization in CPU-based DNN Serving · ICML 2025
Cloud and datacenter computing › inference serving
DNN serving
0.912025
Auto-reconfiguration for Latency Minimization in CPU-based DNN Serving · ICML 2025
Cloud and datacenter computing
inference serving
0.912025
Auto-reconfiguration for Latency Minimization in CPU-based DNN Serving · ICML 2025
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.912025
Auto-reconfiguration for Latency Minimization in CPU-based DNN Serving · ICML 2025
Storage systems › data placement
adaptive data placement
0.412020
Adaptive Placement for In-memory Storage Functions · USENIX ATC 2020
Cloud and datacenter computing
resource management
0.412020
Adaptive Placement for In-memory Storage Functions · USENIX ATC 2020
Machine learning › Efficient and distributed learning
distributed training
0.312026
Checkmate: Zero Performance Overhead Model Checkpointing via Network Gradient Replication · NSDI 2026
Parallel and multicore computing › task allocation
thread placement
0.312025
Auto-reconfiguration for Latency Minimization in CPU-based DNN Serving · ICML 2025
Distributed systems
replication
0.112021
NrOS: Effective Replication and Sharing in an Operating System · OSDI 2021
Storage systems › storage architecture
in-memory storage
0.112020
Adaptive Placement for In-memory Storage Functions · USENIX ATC 2020

Methods — techniques the papers use, named apart from their topics

network gradient replication · 2.0online reconfiguration · 0.9intra-operator parallelism · 0.9
YearPublicationVenuePosition
2026 Dynamic NUMA-Aware Data Structure Replication
Erika Hunhoff, Zack McKevitt, Ankit Bhardwaj 0002, Reto Achermann, Gerd Zellweger, Marcos K. Aguilera, Eric Keller
IPDPS3
2026 Checkmate: Zero Performance Overhead Model Checkpointing via Network Gradient Replication
Ankit Bhardwaj 0002, Weiyang Wang, Jeremy Carin, Adam Belay, Manya Ghobadi
NSDI1
2026 Unleashing The Potential of Datacenter SSDs by Taming Performance Variability
Gohar Irfan Chaudhry, Ankit Bhardwaj 0002, Zhenyuan Ruan, Adam Belay
NSDI2
2025 Auto-reconfiguration for Latency Minimization in CPU-based DNN Serving
abstract
In this paper, we investigate how to push the performance limits of serving Deep Neural Network (DNN) models on CPU-based servers. Specifically, we observe that while intra-operator parallelism across multiple threads is an effective way to reduce inference latency, it provides diminishing returns. Our primary insight is that instead of running a single instance of a model with all available threads on a server, running multiple instances each with smaller batch sizes and fewer threads for intra-op parallelism can provide lower inference latency. However, the right configuration is hard to determine manually since it is workload- (DNN model and batch size used by the serving system) and deployment-dependent (number of CPU cores on server). We present Packrat, a new serving system for online inference that given a model and batch size (𝐵) algorithmically picks the optimal number of instances (𝑖), the number of threads each should be allocated (𝑡), and the batch sizes each should operate on (𝑏) that minimizes latency. Packrat is built as an extension to TorchServe and supports online reconfigurations to avoid serving downtime. Averaged across a range of batch sizes, Packrat improves inference latency by 1.43× to 1.83× on a range of commonly used DNNs.
Ankit Bhardwaj 0002, Amar Phanishayee, Deepak Narayanan, Ryan Stutsman
ICML1
2022 Cache-coherent accelerators for persistent memory crash consistency
abstract
Building persistent memory (PM) data structures is difficult because crashes interrupt operations, leaving data structures in an inconsistent state. Solving this requires augmenting code that modifies PM state to ensure that interrupted operations can be completed or undone. Today, this is done using careful, hand-crafted code, a compiler pass, or page faults. We propose a new, easy way to transform volatile data structure code to work with PM that uses a cache-coherent accelerator to do this augmentation, and we show that it may outperform existing approaches for building PM structures.
Ankit Bhardwaj 0002, Todd Thornley, Vinita Pawar, Reto Achermann, Gerd Zellweger, Ryan Stutsman
HotStorage1
2021 NrOS: Effective Replication and Sharing in an Operating System
Ankit Bhardwaj 0002, Chinmay Kulkarni 0002, Reto Achermann, Irina Calciu, Sanidhya Kashyap, Ryan Stutsman, Amy Tai, Gerd Zellweger
OSDI1
2020 Adaptive Placement for In-memory Storage Functions
Ankit Bhardwaj 0002, Chinmay Kulkarni 0002, Ryan Stutsman
USENIX ATC1