Saksham Agarwal

dblp:190/3721 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 7 · 4 first-author · 6 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 TiNA: Tiered Network Buffer Architecture for Fast Networking in Chiplet-based CPUs
abstract
To manufacture a large CPU cost-effectively, the industry has begun exploiting emerging packaging technologies that integrate multiple chiplets—each comprising a subset of cores and/or memory and I/O subsystems—into a single package. However, such a CPU experiences longer memory access latency with more pronounced variance, especially when its cores in one chiplet access LLC slices or DRAM controllers in other chiplets. This creates unique challenges in μs-scale networking, which is highly sensitive to memory access latency. In this work, we start by proposing exploiting a little-known mode, known as Sub-NUMA Clustering (SNC), in the latest chiplet-based CPUs. As it restricts receiving and processing packets to a particular chiplet unless explicitly specified otherwise, it offers shorter memory access latency and, consequently, lower networking latency than the default mode (non-SNC). Nonetheless, when receiving long bursts of packets, SNC incurs higher networking latency than non-SNC, as it provides less LLC capacity for CPU cores processing the packets, making Direct Cache Access (DCA)—a commonly used CPU feature to reduce memory access latency for packet processing—ineffective. To address this drawback, we propose, a tiered network buffer architecture consisting of an enhanced NIC and networking stack, which opportunistically uses LLC slices in other chiplets for DCA only when receiving long bursts of packets. On average, reduces the mean (tail) latency by 25% (18%) and 28% (22%), compared to SNC and non-SNC, respectively, across diverse network applications and traces.
Siddharth Agarwal, Jinghan Huang 0001, Saksham Agarwal, Nam Sung Kim
ASPLOS (1)4
2026 Interpretable AI Models for Detecting and Classifying Multiple Retinal Conditions Using Hybrid CNN-Transformer-Ensemble Architectures
abstract
Retinal diseases impact a significant portion of the global population, yet specialized care is predominantly available in urban areas. We have developed AI-based methods for diagnosing various retinal conditions using fundus images to address this disparity. We aim to create a highly accurate automatic detection system to bridge the gap between patients and the limited number of retinal specialists. Critical challenges in multi-label classification tasks include limited sample sizes per label and imbalanced class distributions. To overcome these issues and enhance data diversity, we created three composite datasets (MRID) by aggregating multiple open source datasets. We developed hybrid models integrating deep Convolutional Neural Networks (CNNs), Transformer encoders, and ensemble architectures to classify retinal fundus images into 21 distinct labels. Our models demonstrated superior performance compared to baseline methods and the existing state-of-the-art model, with significant improvements in evaluation metrics. Performance was further enhanced by incorporating domain knowledge into the patch extraction step of Vision Transformers used in specific hybrid models. To address user trust and model interpretability, we implemented SHAP-based explanations and used these alongside evaluation metrics for comparative analysis of model performance across datasets. This research aims to advance retinal disease diagnosis, offering accessible healthcare solutions in underserved regions while ensuring accurate and comprehensive disease prediction.
Deependra Singh, Saksham Agarwal, Subhankar Mishra
ACM Trans. Comput. Heal.2
2025 Your network doesn't end at the NIC: A case for unifying the inter-host and intra-host networks in (AI) datacenters
abstract
Modern ML workloads increasingly rely on direct communication between host devices—such as GPUs, NVMe SSDs, and DRAM—spanning intra-host and inter-host networks. However, today's intra-host network lacks hardware-level primitives for routing across heterogeneous interconnects, hindering efficient use of alternative paths and leading to sub-optimal performance under failures or congestion. Furthermore, the inter-host network treats the NIC as the endpoint, with intra-host interconnects like PCIe running oblivious to inter-host network protocols. This prevents leveraging multiple paths for communication between host devices across different servers. To address these limitations, we propose expanding the datacenter network layer to encompass the intra-host network, making intra-host devices first-class network endpoints. Our scheme envisions hardware-level routing and forwarding across multiple intra-host interconnects and makes intra-host devices visible to the inter-host network. This unified approach provides a principled foundation for robust, efficient peer-to-peer communication between storage and compute hardware devices in AI datacenters.
Raj Joshi, Saksham Agarwal, ChonLam Lao, Minlan Yu
HotNets2
2025 Re-architecting End-host Networking with CXL: Coherence, Memory, and Offloading
Houxiang Ji, Yang Zhou 0050, Ipoom Jeong, Ren Wang 0001, Saksham Agarwal, Nam Sung Kim
MICRO6
2024 Harmony: A Congestion-free Datacenter Architecture
Saksham Agarwal, Qizhe Cai, Rachit Agarwal 0001, David B. Shmoys, Amin Vahdat
NSDI1
2024 High-throughput and Flexible Host Networking for Accelerated Computing
Athinagoras Skiadopoulos, Mark Zhao, Qizhe Cai, Saksham Agarwal, Jacob Adelmann, David Ahern, Carlo Contavalli, Michael D. Goldflam, Vitaly Mayatskikh, Raghu Raja, Daniel Walton, Rachit Agarwal 0001, Shrijeet Mukherjee, Christoforos E. Kozyrakis
OSDI5
2024 Understanding the Host Network
abstract
The host network integrates processor, memory, and peripheral interconnects to enable data transfer within the host. Several recent studies from production datacenters show that contention within the host network can have significant impact on end-to-end application performance. The goal of this paper is to build an in-depth understanding of such contention within the host network.
Midhul Vuppalapati, Saksham Agarwal, Henry Schuh, Baris Kasikci, Arvind Krishnamurthy, Rachit Agarwal 0001
SIGCOMM2
2024 Fast & Safe IO Memory Protection
abstract
IO Memory protection mechanisms prevent malicious and/or buggy IO devices from executing errant transfers into memory. Modern servers achieve this using an IOMMU---IO devices operate on virtual addresses, and IOMMU translates virtual addresses to physical addresses (potentially speeding up translations using a cache called IOTLB) before executing memory transfers. Despite their importance, design of memory protection mechanisms that can provide strong safety properties while achieving high performance has remained elusive. Indeed, recent studies from production datacenters demonstrate that inefficiencies within state-of-the-art memory protection mechanisms result in significant throughput degradation, orders-of-magnitude tail latency inflation, and violation of isolation guarantees.
Benny Rubin, Saksham Agarwal, Qizhe Cai, Rachit Agarwal 0001
SOSP2
2023 Host Congestion Control
abstract
The conventional wisdom in systems and networking communities is that congestion happens primarily within the network fabric. However, adoption of high-bandwidth access links and relatively stagnant technology trends for resources within hosts have led to emergence of host congestion---that is, congestion within the host network that enables data exchange between NIC and CPU/memory. Such host congestion alters the many assumptions entrenched within decades of research and practice of congestion control.
Saksham Agarwal, Arvind Krishnamurthy, Rachit Agarwal 0001
SIGCOMM1
2022 Understanding host interconnect congestion
abstract
We present evidence and characterization of host congestion in production clusters: adoption of high-bandwidth access links leading to emergence of bottlenecks within the host interconnect (NIC-to-CPU data path). We demonstrate that contention on existing IO memory management units and/or the memory subsystem can significantly reduce the available NIC-to-CPU bandwidth, resulting in hundreds of microseconds of queueing delays and eventual packet drops at hosts (even when running a state-of-the-art congestion control protocol that accounts for CPU-induced host congestion). We also discuss implications of host interconnect congestion to design of future host architecture, network stacks and network protocols.
Saksham Agarwal, Rachit Agarwal 0001, Behnam Montazeri, Masoud Moshref, Khaled Elmeleegy, Luigi Rizzo, Marc de Kruijf, Gautam Kumar 0001, Sylvia Ratnasamy, David E. Culler, Amin Vahdat
HotNets1
2021 CodedBulk: Inter-Datacenter Bulk Transfers using Network Coding
Shih-Hao Tseng, Saksham Agarwal, Rachit Agarwal 0001, Hitesh Ballani, Ao Tang
NSDI2
2018 Sincronia: near-optimal network design for coflows
abstract
We present Sincronia, a near-optimal network design for coflows that can be implemented on top on any transport layer (for flows) that supports priority scheduling. Sincronia achieves this using a key technical result --- we show that given a "right" ordering of coflows, any per-flow rate allocation mechanism achieves average coflow completion time within 4X of the optimal as long as (co)flows are prioritized with respect to the ordering.
Saksham Agarwal, Shijin Rajakrishnan, Akshay Narayan 0001, Rachit Agarwal 0001, David B. Shmoys, Amin Vahdat
SIGCOMM1