VLDB 2026 Research / reviewers in the wild / expert
Brent E. Stephens
dblp:38/4497
· DBLP profile ↗
28ranked-venue papers
9as first author
11since 2021 · last 2026
0009-0006-9471-1648ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 19 · 8 first-author · 7 since 2021Systems, architecture and hardware · 5 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021Artificial intelligence and machine learning · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Critical Path Guided Decision Making with CALLIGATOR
Meghna Pancholi, Lee Baugh, Olaf Schnapauff, David E. Culler, Kostis Kaffes, Yu Gan 0002, Brent E. Stephens |
SIGCOMM | 7 |
| 2025 | MTP: Transport for In-Network Computing
Rohan Vardekar, Balajee Vamanan, Brent E. Stephens, Aditya Akella |
NSDI | 4 |
| 2024 | CC-NIC: a Cache-Coherent Interface to the NICabstractEmerging interconnects make peripherals, such as the network interface controller (NIC), accessible through the processor's cache hierarchy, allowing these devices to participate in the CPU cache coherence protocol. This is a fundamental change from the separate I/O data paths and read-write transaction primitives of today's PCIe NICs. Our experiments show that the I/O data path characteristics cause NICs to prioritize CPU efficiency at the expense of inflated latency, an issue that can be mitigated by the emerging low-latency coherent interconnects. But, the coherence abstraction is not suited to current host-NIC access patterns. Applying existing signaling mechanisms and data structure layouts in a cache-coherent setting results in extraneous communication and cache retention, limiting performance. Redesigning the interface is necessary to minimize overheads and benefit from the new interactions coherence enables. This work contributes CC-NIC, a host-NIC interface design for coherent interconnects. We model CC-NIC using Intel's Ice Lake and Sapphire Rapids UPI interconnects, demonstrating the potential of optimizing for coherence. Our results show a maximum packet rate of 1.5Gpps and 980Gbps packet throughput. CC-NIC has 77% lower minimum latency, and 88% lower at 80% load, than today's PCIe NICs. We also demonstrate application-level core savings. Finally, we show that CC-NIC's benefits hold across a range of interconnect performance characteristics. Henry Schuh, Arvind Krishnamurthy, David E. Culler, Henry M. Levy, Luigi Rizzo, Samira Manabi Khan, Brent E. Stephens |
ASPLOS (1) | 7 |
| 2024 | Poster: Topocloud: Getting Datacenter Network Experiments Into the Right ShapeabstractThe physical topology of a datacenter network is fundamental as it determines latency, bisection bandwidth, the location and properties of congestion, and the network's ability to tolerate failures. Yet, topology is one of the most difficult factors to control during experimentation: when using a production datacenter or a testbed, the only topology available is the one already deployed. If running on an experimenter's own equipment, rewiring may be possible but is cumbersome and of limited scale. This presents a challenge for controlled experimentation, because such experiments should ideally be run on a range of realistic topologies. To tackle this, we present our work on Topocloud, a system that enables experimenters to construct and modify network topologies in a shared public testbed. Topocloud uses hardware P4 switches to fulfill a dual role: each port can act as either a typical Ethernet MAC-learning switch or as a “virtual wire” that connects ports transparently. We demonstrate that virtual wires in Topocloud can be used to construct low latency network topologies incorporating L2 switches. Pavani Kuppili, Aleksander Maricq, Brent E. Stephens, Ryan Stutsman, Robert Ricci |
ICNP | 3 |
| 2023 | Yama: Providing Performance Isolation for Black-Box OffloadsabstractThe sharing of clusters with various on-NIC offloads by high-level entities (users, containers, etc.) has become increasingly common. Performance isolation across these entities is desired because the offloads can become bottlenecks due to the limited capacity of hardware. However, the existing works that provide scheduling and resource management to NIC offloads all require customization of the NIC or offloads, while commodity off-the-shelf NICs and offloads with proprietary implementation have been widely deployed in datacenters. This paper presents Yama, the first solution to enable per-entity isolation in the sharing of such black-box NIC offloads. Yama provides a generic framework that captures a common abstraction to the operation of most offloads, which allows operators to incorporate existing offloads. The framework proactively probes for the performance of the offloads with auxiliary workload and enforces isolation at the initiator side. Yama also accommodates chained offloads. Our evaluation shows that 1) Yama achieves per-entity max-min fairness for various types of offloads and in complicated offload chaining scenarios; 2) Yama quickly converges to changes in equilibrium and 3) Yama adds negligible overhead to application workload. Divyanshu Saxena, Brent E. Stephens, Aditya Akella |
SoCC | 3 |
| 2023 | RingLeader: Efficiently Offloading Intra-Server Orchestration to NICs
Adney Cardoza, Tarannum Khan, Yeonju Ro, Brent E. Stephens, Hassan M. G. Wassel, Aditya Akella |
NSDI | 5 |
| 2023 | A Cloud-Scale Characterization of Remote Procedure CallsabstractThe global scale and challenging requirements of modern cloud applications have led to the development of complex, widely distributed, service-oriented applications. One enabler of such applications is the remote procedure call (RPC), which provides location-independent communication and hides the myriad of cloud communication complexities and requirements within the RPC stack. Understanding RPCs is thus one key to understanding the behavior of cloud applications. While there have been numerous studies of RPCs in distributed systems, as well as attempts to optimize RPC overheads with both software and hardware, there is still a lack of knowledge about the characteristics of RPCs "in the wild" in the modern cloud environment. Korakit Seemakhupt, Brent E. Stephens, Samira Manabi Khan, Sihang Liu 0001, Hassan M. G. Wassel, Soheil Hassas Yeganeh, Alex C. Snoeren, Arvind Krishnamurthy, David E. Culler, Henry M. Levy |
SOSP | 2 |
| 2022 | ADA: Arithmetic Operations with Adaptive TCAM Population in Programmable SwitchesabstractIn-network applications, such as congestion control, load-balancing, and policy enforcement, require complicated arithmetic operations to track networking parameters. Unfortunately, programmable switches that implement protocol independent switch architecture (PISA) support only a limited set of arithmetic operations, such as addition and subtraction, to guarantee high packet throughput. Existing work addresses this problem by implementing unsupported operations (e.g., multiplication) using TCAM match-action tables; they use wildcards to match over a range of operand values. However, because TCAM is a scarce resource, operators must make a difficult trade-off between accuracy and TCAM occupancy. This problem leads to large and unpredictable errors, and also limits the applicability of in-network computing to many applications.In this paper, we propose ADA, a practical, lightweight approach to reduce TCAM entries without sacrificing accuracy by exploiting the value distribution of operands. ADA tracks the operands’ distribution via a simple binning mechanism to determine the most accessed interval in the domain space of operands and allocates more (or less) entries based on the observed distribution. Our proposed mechanism, (1) saves TCAM space for other applications by aggregating entries that are unused or less popular, and (2) reduces average error by assigning more TCAM entries to intervals with a higher probability of occurrence (and sub-divides these intervals further, if needed). We implement ADA on P4 on a 100 Gbps Barefoot Tofino switch and demonstrate its efficacy by deploying it in existing state-of-the-art in-network applications; ADA imposes a negligible overhead of less than 2% in the switch data plane and about 5% in the control plane. We further evaluate ADA using our C++ and ns-3 simulators over two existing arithmetic-heavy applications (i.e., Nimble and RCP) to demonstrate that ADA can achieve performance close to an ideal implementation with unlimited TCAM space. Mojtaba MalekpourShahraki, Brent E. Stephens, Balajee Vamanan |
ICDCS | 2 |
| 2022 | Backdraft: a Lossless Virtual Switch that Prevents the Slow Receiver Problem
Alireza Sanaee, Farbod Shahinfar, Gianni Antichi, Brent E. Stephens |
NSDI | 4 |
| 2022 | Justitia: Software Multi-Tenancy in Hardware Kernel-Bypass Networks
Yiwen Zhang 0008, Brent E. Stephens, Mosharaf Chowdhury |
NSDI | 3 |
| 2021 | TCP is Harmful to In-Network Computing: Designing a Message Transport Protocol (MTP)abstractThis paper presents the motivation and design of MTP, a new offload-friendly message transport protocol. Existing transport protocols like TCP, MPTCP, and UDP/Quic all have key limitations when used in a network that may potentially offload computation from end-servers into NICs, switches, and other network devices. To enable important new in-network computing use cases and correct congestion control in the face of ever changing network paths and application replicas, MTP introduces a new message transport protocol design and pathlet congestion control, a new approach where end-hosts explicitly communicate messaging information to network devices and network devices explicitly communicate network path and congestion information back to end-hosts. Brent E. Stephens, Darius Grassi, Hamidreza Almasi, Balajee Vamanan, Aditya Akella |
HotNets | 1 |
| 2020 | PANIC: A High-Performance Programmable NIC for Multi-tenant Networks
Kiran Patel, Brent E. Stephens, Anirudh Sivaraman, Aditya Akella |
OSDI | 3 |
| 2020 | Equilibrium optimizer: A novel optimization algorithm
Afshin Faramarzi, Mohammad Heidarinejad, Brent E. Stephens, Seyedali Mirjalili |
Knowl. Based Syst. | 3 |
| 2019 | On the Impact of Cluster Configuration on RoCE Application DesignabstractRDMA over Converged Ethernet (RoCE) allows RDMA-enabled NICs to operate in datacenter networks. This study focuses on identifying how different aspects of datacenter cluster configuration impact the latency, and throughput, and CPU utilization of different ways of transferring data in RoCE (RDMA verbs). We look into the impact of colocated applications competing for both the CPU and access to the NIC as well as the impact of the network MTU. We find that RDMA applications do not fairly share the NIC, large frames should not be used, and that correct verb choice is dependent on many variables, including application access patterns, object size, and the load of both the local and remote CPU. Yanfang Le, Mojtaba MalekpourShahraki, Brent E. Stephens, Aditya Akella, Michael M. Swift |
APNet | 3 |
| 2019 | Ether: Providing both Interactive Service and Fairness in Multi-Tenant DatacentersabstractMulti-tenant datacenters and cloud networks must provide both isolation and interactive service to tenant applications, many of which are sensitive to tail flow completion times. Network operators must also ensure high utilization of network capacity to reduce cost. Existing approaches that statically partition network capacity, in either time or space, provide good isolation but suffer from under-utilization. Existing schemes that dynamically allocate capacity to tenants incur either decreased fairness or high tail flow completion times. To overcome these limitations, we propose Ether. Ether is able to overcome these limitations because it can prioritize bursty flows during short congestion episodes while still ensuring fairness at long timescales. In this paper, we present a preliminary design of Ether and discuss its feasibility in today's programmable switches. Our evaluations show that, at high loads, Ether achieves 23% improvement in tail flow completion times (FCT) when compared with idealized fair queueing (FQ) while still providing similar fairness as FQ. In contrast, pFabric, which optimizes FCT, worsens fairness by a factor of 1.8 when compared with Ether. Mojtaba MalekpourShahraki, Brent E. Stephens, Balajee Vamanan |
APNet | 2 |
| 2019 | Loom: Flexible and Efficient NIC Packet Scheduling
Brent E. Stephens, Aditya Akella, Michael M. Swift |
NSDI | 1 |
| 2018 | RoGUE: RDMA over Generic Unconverged EthernetabstractRDMA over Converged Ethernet (RoCE) promises low latency and low CPU utilization over commodity networks, and is attractive for cloud infrastructure services. Current implementations require Priority Flow Control (PFC) that uses backpressure-based congestion control to provide lossless networking to RDMA. Unfortunately, PFC compromises network stability. As a result, RoCE's adoption has been slow and requires complex network management. Recent efforts, such as DCQCN, reduce the risk to the network, but do not completely solve the problem. Yanfang Le, Brent E. Stephens, Arjun Singhvi, Aditya Akella, Michael M. Swift |
SoCC | 2 |
| 2018 | Your Programmable NIC Should be a Programmable SwitchabstractToday's NICs are becoming programmable ("smart"). To support new network protocols, services, and offloads, there are NICs today that have on-board FPGAs, embedded processors, programmable forwarding pipelines, and specialized engines to support features like RDMA. Unfortunately, existing programmable NICs have a number of key limitations. It is difficult to chain offloads, schedule competing accesses to shared resources, and support functions that require variable processing time and thus may not run at line-rate. Brent E. Stephens, Aditya Akella, Michael M. Swift |
HotNets | 1 |
| 2017 | Low Latency Software Rate Limiters for Cloud NetworksabstractA lot of recent work has focused on reducing in network queueing latency in datacenter networks. In this paper, we focus on a less explored topic --- latency increases caused by queueing in rate limiters on the end-host. First, we show that latency can be increased by an order of magnitude by rate limiters in cloud networks. To solve this problem, we extend ECN marking into rate limiters and use a datacenter congestion control algorithm --- DCTCP. Unfortunately, while this reduces latency, it also leads to throughput oscillation. Thus, this solution is not sufficient. In this paper, we also analyze the specific reasons that ECN marking in software rate limiters leads to the throughput oscillation problem. Finally, we propose two potential solutions to design software rate limiters that can achieve stable high throughput and low latency. Keqiang He, Weite Qin, Wenfei Wu, Tian Pan 0001, Chengchen Hu, Jiao Zhang 0002, Brent E. Stephens, Aditya Akella, Ying Zhang 0022 |
APNet | 9 |
| 2017 | Titan: Fair Packet Scheduling for Commodity Multiqueue NICs
Brent E. Stephens, Arjun Singhvi, Aditya Akella, Michael M. Swift |
USENIX ATC | 1 |
| 2016 | Deadlock-free local fast failover for arbitrary data center networksabstractToday, given data center networks' sizes and bursty workloads, it is likely that at any moment there is packet loss due to some type of failure in the network. This paper focuses on solving the two most common types of data center network failures: congestion and routing failures. Recently, there has been demand for lossless Ethernet (DCB) in data center networks as a solution to congestion failures. However, DCB complicates fault tolerance by introducing a new type of failure, deadlock. If DCB is enabled, then all routing must be deadlock free. To the best of our knowledge, this paper describes the first ever deadlock-free approaches to local fast failover that can be combined with DCB, DF-FI and DF-EDST resilience. Moreover, in the evaluation, this paper shows that DF-EDST resilience, which is the paper's main contribution, can improve fault tolerance without adversely impacting performance when compared to a state-of-the-art approach to deadlock-free routing. If, however, a small reduction in aggregate throughput is acceptable, then it is possible to build routes such that only 0.00001% of the total flows in the network are likely to fail given 16 edge failures on networks with 1K-4K hosts. Brent E. Stephens, Alan L. Cox |
INFOCOM | 1 |
| 2014 | Practical DCB for improved data center networksabstractStorage area networking is driving commodity data center switches to support lossless Ethernet (DCB). Unfortunately, to enable DCB for all traffic on arbitrary network topologies, we must address several problems that can arise in lossless networks, e.g., large buffering delays, unfairness, head of line blocking, and deadlock. We propose TCP-Bolt, a TCP variant that not only addresses the first three problems but reduces flow completion times by as much as 70%. We also introduce a simple, practical deadlock-free routing scheme that eliminates deadlock while achieving aggregate network throughput within 15% of ECMP routing. This small compromise in potential routing capacity is well worth the gains in flow completion time. We note that our results on deadlock-free routing are also of independent interest to the storage area networking community. Further, as our hardware testbed illustrates, these gains are achievable today, without hardware changes to switches or NICs. Brent E. Stephens, Alan L. Cox, Ankit Singla, John B. Carter, Colin Dixon, Wes Felter |
INFOCOM | 1 |
| 2014 | Planck: millisecond-scale monitoring and control for commodity networksabstractSoftware-defined networking introduces the possibility of building self-tuning networks that constantly monitor network conditions and react rapidly to important events such as congestion. Unfortunately, state-of-the-art monitoring mechanisms for conventional networks require hundreds of milliseconds to seconds to extract global network state, like link utilization or the identity of "elephant" flows. Such latencies are adequate for responding to persistent issues, e.g., link failures or long-lasting congestion, but are inadequate for responding to transient problems, e.g., congestion induced by bursty workloads sharing a link. In this paper, we present Planck, a novel network measurement architecture that employs oversubscribed port mirroring to extract network information at 280 µs--7 ms timescales on a 1 Gbps commodity switch and 275 µs--4 ms timescales on a 10 Gbps commodity switch,over 11x and 18x faster than recent approaches, respectively (and up to 291x if switch firmware allowed buffering to be disabled on some ports). To demonstrate the value of Planck's speed and accuracy, we use it to drive a traffic engineering application that can reroute congested flows in milliseconds. On a 10 Gbps commodity switch, Planck-driven traffic engineering achieves aggregate throughput within 1--4% of optimal for most workloads we evaluated, even with flows as small as 50 MiB, an improvement of up to 53% over previous schemes. Jeff Rasley, Brent E. Stephens, Colin Dixon, Eric Rozner, Wes Felter, Kanak Agarwal 0001, John B. Carter, Rodrigo Fonseca |
SIGCOMM | 2 |
| 2013 | Plinko: building provably resilient forwarding tablesabstractThis paper introduces Plinko, a network architecture that uses a novel forwarding model and routing algorithm to build networks with forwarding paths that, assuming arbitrarily large forwarding tables, are provably resilient against t link failures, ∀t ∈ N. However, in practice, there are clearly limits on the size of forwarding tables. Nonetheless, when constrained to hardware comparable to modern top-of-rack (TOR) switches, Plinko scales with high resilience to networks with up to ten thousand hosts. Thus, as long as t or fewer links have failed, the only reason packets of any flow in a Plinko network will be dropped are congestion, packet corruption, and a partitioning of the network topology, and, even after t + 1 failures, most, if not all, flows may be unaffected. In addition, Plinko is topology independent, supports arbitrary paths for routing, provably bounds stretch, and does not require any additional computation during forwarding. To the best of our knowledge, Plinko is the first network to have all of these properties. Brent E. Stephens, Alan L. Cox, Scott Rixner |
HotNets | 1 |
| 2012 | PAST: scalable ethernet for data centersabstractWe present PAST, a novel network architecture for data center Ethernet networks that implements a Per-Address Spanning Tree routing algorithm. PAST preserves Ethernet's self-configuration and mobility support while increasing its scalability and usable bandwidth. PAST is explicitly designed to accommodate unmodified commodity hosts and Ethernet switch chips. Surprisingly, we find that PAST can achieve performance comparable to or greater than Equal-Cost Multipath (ECMP) forwarding, which is currently limited to layer-3 IP networks, without any multipath hardware support. In other words, the hardware and firmware changes proposed by emerging standards like TRILL are not required for high-performance, scalable Ethernet networks. We evaluate PAST on Fat Tree, HyperX, and Jellyfish topologies, and show that it is able to capitalize on the advantages each offers. We also describe an OpenFlow-based implementation of PAST in detail. Brent E. Stephens, Alan L. Cox, Wes Felter, Colin Dixon, John B. Carter |
CoNEXT | 1 |
| 2011 | A Scalability Study of Enterprise Network ArchitecturesabstractThe largest enterprise networks already contain hundreds of thousands of hosts. Enterprise networks are composed of Ethernet subnets interconnected by IP routers. These routers require expensive configuration and maintenance. If the Ethernet subnets are made more scalable, the high cost of the IP routers can be eliminated. Unfortunately, it has been widely acknowledged that Ethernet does not scale well because it relies on broadcast, which wastes bandwidth, and a cycle-free topology, which poorly distributes load and forwarding state. There are many recent proposals to replace Ethernet, each with its own set of architectural mechanisms. These mechanisms include eliminating broadcasts, using source routing, and restricting routing paths. Although there are many different proposed designs, there is little data available that allows for comparisons between designs. This study performs simulations to evaluate all of the factors that affect the scalability of Ethernet together, which has not been done in any of the proposals. The simulations demonstrate that, in a realistic environment, source routing reduces the maximum state requirements of the network by over an order of magnitude. About the same level of traffic engineering achieved by load-balancing all the flows at the TCP/UDP flow granularity is possible by routing only the heavy flows at the TCP/UDP granularity. Additionally, requiring routing restrictions, such as deadlock-freedom or minimum-hop routing, can significantly reduce the network's ability to perform traffic engineering across the links. Brent E. Stephens, Alan L. Cox, Scott Rixner, T. S. Eugene Ng |
ANCS | 1 |
| 2010 | Axon: a flexible substrate for source-routed ethernetabstractThis paper introduces the Axon, an Ethernet-compatible device for creating large-scale datacenter networks. Axons are inexpensive, practical devices that are demonstrated using prototype hardware. Functionally, Axons replace Ethernet switches and maintain full compatibility with existing Ethernet hosts. Between themselves, however, Axons transparently use source-routed Ethernet. This unlocks many benefits, such as improved network scalability, performance, and flexibility.In an Axon network, all state required to route a host's packets is placed in the local Axon---the Axon to which the host is directly connected. Therefore, regardless of the scale of the network, the route computation and storage needs of a single Axon device only need to scale with the demands of its locally-connected hosts. This is in stark contrast to conventional switched Ethernet, which requires routing resources proportional to the traffic that flows through the device. Scalability is also increased by eliminating the use of packet flooding for automatic location and address discovery. Further, source-routed Ethernet increases network flexibility by supporting different route selection strategies. For example, shortest-path routing could be employed, or longer paths selected to minimize congestion by balancing traffic across redundant links. Jeffrey Shafer, Brent E. Stephens, Michael Foss, Scott Rixner, Alan L. Cox |
ANCS | 2 |
| 2003 | Using Genetic Algorithms to Discover Selection Criteria for Contradictory Solutions Retrieved by CBR
Costas Tsatsoulis, Brent E. Stephens |
ICCBR | 2 |