Arjun Singhvi

dblp:160/4979 · DBLP profile ↗
← Back
11ranked-venue papers
6as first author
5since 2021 · last 2025
0009-0008-5235-3592ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 6 · 4 first-author · 3 since 2021Systems, architecture and hardware · 5 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Cloud and datacenter computing · 60% Distributed systems · 20% Memory systems · 11%
Computer networks
2 papers
Datacenter networks · 53% Transport protocols and congestion control · 27% Internet architecture and protocols · 9%
Databases, data mining, and information retrieval
1 paper
Query processing and optimization · 77% Database system architecture and tuning · 23%

Topics — the 15 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Transport protocols and congestion control
delay-based congestion control
0.912025
Falcon: A Reliable, Low Latency Hardware Transport · SIGCOMM 2025
Datacenter networks › load balancing
multipath load balancing
0.912025
Falcon: A Reliable, Low Latency Hardware Transport · SIGCOMM 2025
Cloud and datacenter computing › serverless computing
cold start mitigation
0.612022
Memory deduplication for serverless computing with Medes · EuroSys 2022
Memory systems › memory management
memory deduplication
0.612022
Memory deduplication for serverless computing with Medes · EuroSys 2022
Cloud and datacenter computing
serverless computing
0.612022
Memory deduplication for serverless computing with Medes · EuroSys 2022
Distributed systems
distributed caching
0.512021
CliqueMap: productionizing an RMA-based distributed caching system · SIGCOMM 2021
Distributed systems › distributed communication
remote memory access
0.512021
CliqueMap: productionizing an RMA-based distributed caching system · SIGCOMM 2021
Cloud and datacenter computing
cluster resource management and scheduling
0.412020
Themis: Fair and Efficient GPU Cluster Scheduling · NSDI 2020
Cloud and datacenter computing › job scheduling
fair scheduling
0.412020
Themis: Fair and Efficient GPU Cluster Scheduling · NSDI 2020
Cloud and datacenter computing › cluster resource management and scheduling › cluster scheduling
GPU cluster scheduling
0.412020
Themis: Fair and Efficient GPU Cluster Scheduling · NSDI 2020
Cloud and datacenter computing › multi-tenancy
multi-tenant datacenter
0.412020
1RMA: Re-envisioning Remote Memory Access for Multi-tenant Datacenters · SIGCOMM 2020
Interconnection networks and networks-on-chip
remote direct memory access
0.412020
1RMA: Re-envisioning Remote Memory Access for Multi-tenant Datacenters · SIGCOMM 2020
Wireless networking
fair scheduling
0.312017
Titan: Fair Packet Scheduling for Commodity Multiqueue NICs · USENIX ATC 2017
Internet architecture and protocols
packet scheduling
0.312017
Titan: Fair Packet Scheduling for Commodity Multiqueue NICs · USENIX ATC 2017
Cloud and datacenter computing
datacenter services
0.112021
CliqueMap: productionizing an RMA-based distributed caching system · SIGCOMM 2021

Methods — techniques the papers use, named apart from their topics

programmable engine · 0.9hardware retransmission · 0.9sandbox state management · 0.6memory deduplication · 0.6remote procedure call · 0.5remote memory access · 0.5
YearPublicationVenuePosition
2025 Falcon: A Reliable, Low Latency Hardware Transport
abstract
Hardware transports such as RoCE deliver high performance with minimal host CPU, but are best suited to special-purpose deployments that limit their use, e.g., backend networks or Ethernet with Priority Flow Control (PFC). We introduce Falcon, the first hardware transport that supports multiple Upper Layer Protocols (ULPs) and heterogeneous application workloads in general-purpose Ethernet datacenter environments (with losses and without special switch support). Key design elements include: delay-based congestion control with multipath load balancing; a layered design with a simple request-response transaction interface for multi-ULP support; hardware-based retransmissions and error-handling for scalability; and a programmable engine for flexibility. The first Falcon hardware implementation delivers a peak performance of 200 Gbps, 120 Mops/sec, with near-optimal operation completion times that are up to 8× lower than CX-7 RoCE under network congestion, and up to 65% higher goodput under lossy conditions.
Arjun Singhvi, Nandita Dukkipati, Prashant Chandra, Hassan M. G. Wassel, Naveen Kr. Sharma, Anthony Rebello, Henry Schuh, Praveen Kumar 0003, Behnam Montazeri, Neelesh Bansod, Sarin Thomas, Inho Cho, Hyojeong Lee Seibert, Baijun Wu, Rui Yang 0034, Qianwen Yin, Srinivas Vaduvatha, Weihuang Wang, Masoud Moshref, David Wetherall, Amin Vahdat
SIGCOMM1
2022 Memory deduplication for serverless computing with Medes
abstract
Serverless platforms today impose rigid trade-offs between resource use and user-perceived performance. Limited controls, provided via toggling sandboxes between warm and cold states and keep-alives, force operators to sacrifice significant resources to achieve good performance. We present a serverless framework, Medes, that breaks the rigid trade-off and allows operators to navigate the trade-off space smoothly. Medes leverages the fact that the warm sandboxes running on serverless platforms have a high fraction of duplication in their memory footprints. We exploit these redundant chunks to develop a new sandbox state, called a dedup state, that is more memory-efficient than the warm state and faster to restore from than the cold state. We develop novel mechanisms to identify memory redundancy at minimal overhead while ensuring that the dedup containers' memory footprint is small. Finally, we develop a simple sandbox management policy that exposes a narrow, intuitive interface for operators to trade-off performance for memory by jointly controlling warm and dedup sandboxes. Detailed experiments with a prototype using real-world serverless workloads demonstrate that Medes can provide up to 1×-2.75× improvements in the end-to-end latencies. The benefits of Medes are enhanced in memory pressure situations, where Medes can provide up to 3.8× improvements in end-to-end latencies. Medes achieves this by reducing the number of cold starts incurred by 10--50% against the state-of-the-art baselines.
Divyanshu Saxena, Arjun Singhvi, Junaid Khalid, Aditya Akella
EuroSys3
2021 Atoll: A Scalable Low-Latency Serverless Platform
abstract
With user-facing apps adopting serverless computing, good latency performance of serverless platforms has become a strong fundamental requirement. However, it is difficult to achieve this on platforms today due to the design of their underlying control and data planes that are particularly ill-suited to short-lived functions with unpredictable arrival patterns. We present Atoll, a serverless platform, that overcomes the challenges via a ground-up redesign of the control and data planes. In Atoll, each app is associated with a latency deadline. Atoll achieves its per-app request latency goals by: (a) partitioning the cluster into (semi-global scheduler, worker pool) pairs, (b) performing deadline-aware scheduling and proactive sandbox allocation, and (c) using a load balancing layer to do sandbox-aware routing, and automatically scale the semi-global schedulers per app. Our results show that Atoll reduces missed deadlines by ~66x and tail latencies by ~3x compared to state-of-the-art alternatives.
Arjun Singhvi, Arjun Balasubramanian, Kevin Houck, Mohammed Danish Shaikh, Shivaram Venkataraman, Aditya Akella
SoCC1
2021 Whiz: Data-Driven Analytics Execution
Robert Grandl, Arjun Singhvi, Raajay Viswanathan, Aditya Akella
NSDI2
2021 CliqueMap: productionizing an RMA-based distributed caching system
abstract
Distributed in-memory caching is a key component of modern Internet services. Such caches are often accessed via remote procedure call (RPC), as RPC frameworks provide rich support for productionization, including protocol versioning, memory efficiency, auto-scaling, and hitless upgrades. However, full-featured RPC limits performance and scalability as it incurs high latencies and CPU overheads. Remote Memory Access (RMA) offers a promising alternative, but meeting productionization requirements can be a significant challenge with RMA-based systems due to limited programmability and narrow RMA primitives.
Arjun Singhvi, Aditya Akella, Maggie Anderson, Rob Cauble, Harshad Deshmukh, Dan Gibson, Milo M. K. Martin, Amanda Strominger, Thomas F. Wenisch, Amin Vahdat
SIGCOMM1
2020 SNF: serverless network functions
abstract
Our work addresses how a cloud provider can offer Network Functions (NF) as a Service, or NFaaS, using the emerging serverless computing paradigm. Serverless computing has the right NFaaS building blocks - usage-based billing, event-driven programming model and elastic scaling. But we identify two core limitations of existing serverless platforms that undermine support for NFaaS - coupling of the billing and work assignment granularities, and state sharing via an external store. Our framework, SNF, overcomes these limitations via two ideas. SNF allocates work at the granularity of flowlets observed in network traffic, whereas billing and programming occur at a finer level. SNF embellishes serverless platforms with ephemeral local state that lasts for the flowlet duration and supports high performance state operations. We demonstrate that our SNF prototype matches utilization closely with demand and reduces tail packet processing latency substantially compared to alternatives.
Arjun Singhvi, Junaid Khalid, Aditya Akella, Sujata Banerjee
SoCC1
2020 Themis: Fair and Efficient GPU Cluster Scheduling
Kshiteej Mahajan, Arjun Balasubramanian, Arjun Singhvi, Shivaram Venkataraman, Aditya Akella, Amar Phanishayee, Shuchi Chawla 0001
NSDI3
2020 1RMA: Re-envisioning Remote Memory Access for Multi-tenant Datacenters
abstract
Remote Direct Memory Access (RDMA) plays a key role in supporting performance-hungry datacenter applications. However, existing RDMA technologies are ill-suited to multi-tenant datacenters, where applications run at massive scales, tenants require isolation and security, and the workload mix changes over time. Our experiences seeking to operationalize RDMA at scale indicate that these ills are rooted in standard RDMA's basic design attributes: connectionorientedness and complex policies baked into hardware.
Arjun Singhvi, Aditya Akella, Dan Gibson, Thomas F. Wenisch, Monica Wong-Chan, Sean Clark 0003, Milo M. K. Martin, Moray McLaren, Prashant Chandra, Rob Cauble, Hassan M. G. Wassel, Behnam Montazeri, Simon L. Sabato, Joel Scherpelz, Amin Vahdat
SIGCOMM1
2018 RoGUE: RDMA over Generic Unconverged Ethernet
abstract
RDMA over Converged Ethernet (RoCE) promises low latency and low CPU utilization over commodity networks, and is attractive for cloud infrastructure services. Current implementations require Priority Flow Control (PFC) that uses backpressure-based congestion control to provide lossless networking to RDMA. Unfortunately, PFC compromises network stability. As a result, RoCE's adoption has been slow and requires complex network management. Recent efforts, such as DCQCN, reduce the risk to the network, but do not completely solve the problem.
Yanfang Le, Brent E. Stephens, Arjun Singhvi, Aditya Akella, Michael M. Swift
SoCC3
2017 Granular Computing and Network Intensive Applications: Friends or Foes?
abstract
Computing/infrastructure as a service continues to evolve with bare metal, virtual machines, containers and now serverless granular computing service offerings. Granular computing enables developers to decompose their applications into smaller logical units or functions, and run them on small, low cost and short lived computation containers without having to worry about setting up servers - hence the term serverless computing. While serverless environments can be used very cost effectively for large scale parallel processing data analytics applications, it is less clear if network intensive packet processing applications can also benefit from these new computing services as they do not share the same characteristics. This paper examines the architectural constraints as well as current serverless implementations to develop a position on this topic and influence the next generation of computing services. We support our position through measurement and experimentation on Amazon's AWS Lambda service with a few popular network functions.
Arjun Singhvi, Sujata Banerjee, Yotam Harchol, Aditya Akella, Mark Peek, Pontus Rydin
HotNets1
2017 Titan: Fair Packet Scheduling for Commodity Multiqueue NICs
Brent E. Stephens, Arjun Singhvi, Aditya Akella, Michael M. Swift
USENIX ATC2