VLDB 2026 Research / reviewers in the wild / expert
Neal Cardwell
dblp:82/1529
· DBLP profile ↗
10ranked-venue papers
2as first author
4since 2021 · last 2026
0009-0003-0900-0859ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 6 · 1 first-author · 4 since 2021Systems, architecture and hardware · 3 · 1 first-authorSoftware engineering, systems software and programming languages · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CSIG: Congestion Signaling for Datacenter TransportsabstractOptimizing burst-heavy datacenter workloads necessitates finegrained network control and visibility. We introduce CSIG, a protocol that delivers precise, multi-bit bottleneck congestion signals via a fixed-length Ethernet header. The architecture captures μsgranularity switch metrics, such as available bandwidth, and signals them to end-hosts using in-band, line-rate operations. We propose Fast Ramp-Up, a congestion control primitive that leverages these bottleneck signals to reduce median RPC latency by 20% and unclaimed bandwidth by 60% in production. Beyond transport-level performance, CSIG enables flow-aware observability by embedding μs-scale metrics into every packet, allowing individual application transfers to pinpoint their bottleneck location, such as the topology tier limiting their performance. CSIG thus transforms network telemetry from post-hoc correlation into a real time, context-aware capability. We demonstrate CSIG's broad deployability by validating it across five generations of commodity switch hardware (up to 102.4 Tbps), four NIC generations, and five transport stacks. Our design proves that a streamlined Layer 2 approach, focusing exclusively on the principal path bottleneck, provides transport-agnostic gains without requiring forklift hardware upgrades. Abhiram Ravi, Nandita Dukkipati, Weiwu Pang, Neal Cardwell, Brad Karp, Mohammad Jafar Akhbarizadeh, Weida Huang, Konstantinos Prasopoulos, Kok-Kiong Yap, Amin Vahdat |
SIGCOMM | 4 |
| 2025 | Preventing Network Bottlenecks: Accelerating Datacenter Services with Hotspot-Aware Placement for Compute and Storage
Hamid Hajabdolali Bazzaz, Yingjie Bi, Weiwu Pang, Minlan Yu, Ramesh Govindan, Neal Cardwell, Nandita Dukkipati, Meng-Jung Tsai, Chris DeForeest, Yuxue Jin, Charles J. Carver, Jan Kopanski, Liqun Cheng, Amin Vahdat |
NSDI | 6 |
| 2023 | Poseidon: Efficient, Robust, and Practical Datacenter CC via Deployable INT
Masoud Moshref, Gautam Kumar 0001, T. S. Eugene Ng, Neal Cardwell, Nandita Dukkipati |
NSDI | 6 |
| 2023 | Fathom: Understanding Datacenter Application Network PerformanceabstractWe describe our experience with Fathom, a system for identifying the network performance bottlenecks of any service running in the Google fleet. Fathom passively samples RPCs, the principal unit of work for services. It segments the overall latency into host and network components with kernel and RPC stack instrumentation. It records these detailed latency metrics, along with detailed transport connection state, for every sampled RPC. This lets us determine if the completion is constrained by the client, network or server. To scale while enabling analysis, we also aggregate samples into distributions that retain multi-dimensional breakdowns. This provides us with a macroscopic view of individual services. Fathom runs globally in our datacenters for all production traffic, where it monitors billions of TCP connections 24x7. For five years Fathom has been our primary tool for troubleshooting service network issues and assessing network infrastructure changes. We present case studies to show how it has helped us improve our production services. Mubashir Adnan Qureshi, Junhua Yan, Yuchung Cheng, Soheil Hassas Yeganeh, Yousuk Seung, Neal Cardwell, Willem de Bruijn, Van Jacobson, Jasleen Kaur 0001, David Wetherall, Amin Vahdat |
SIGCOMM | 6 |
| 2013 | Reducing web latency: the virtue of gentle aggressionabstractTo serve users quickly, Web service providers build infrastructure closer to clients and use multi-stage transport connections. Although these changes reduce client-perceived round-trip times, TCP's current mechanisms fundamentally limit latency improvements. We performed a measurement study of a large Web service provider and found that, while connections with no loss complete close to the ideal latency of one round-trip time, TCP's timeout-driven recovery causes transfers with loss to take five times longer on average. Tobias Flach, Nandita Dukkipati, Andreas Terzis, Barath Raghavan, Neal Cardwell, Yuchung Cheng, Shuai Hao 0002, Ethan Katz-Bassett, Ramesh Govindan |
SIGCOMM | 5 |
| 2013 | packetdrill: Scriptable Network Stack Testing, from Sockets to Packets
Neal Cardwell, Yuchung Cheng, Lawrence Brakmo, Matthew Mathis, Barath Raghavan, Nandita Dukkipati, Hsiao-Keng Jerry Chu, Andreas Terzis, Tom Herbert |
USENIX ATC | 1 |
| 2004 | Monkey See, Monkey Do: A Tool for TCP Tracing and Replaying
Yuchung Cheng, Urs Hölzle, Neal Cardwell, Stefan Savage, Geoffrey M. Voelker |
USENIX ATC, General Track | 3 |
| 2000 | Modeling TCP LatencyabstractSeveral analytic models describe the steady-state throughput of bulk transfer TCP flows as a function of round trip time and packet loss rate. These models describe flows based on the assumption that they are long enough to sustain many packet losses. However, most TCP transfers across today's Internet are short enough to see few, if any, losses and consequently their performance is dominated by startup effects such as connection establishment and slow start. This paper extends the steady-state model proposed in Padhye et al. (1998), in order to capture these startup effects. The extended model characterizes the expected value and distribution of TCP connection establishment and data transfer latency as a function of transfer size, round trip time, and packet loss rate. Using simulations, controlled measurements of TCP transfers, and live Web measurements we show that, unlike earlier steady-state models for TCP performance, our extended model describes connection establishment and data transfer latency under a range of packet loss conditions, including no loss. Neal Cardwell, Stefan Savage, Thomas E. Anderson |
INFOCOM | 1 |
| 1999 | On the scale and performance of cooperative Web proxy cachingabstractWhile algorithms for cooperative proxy caching have been widely studied, little is understood about cooperativecaching performance in the large-scale World Wide Web environment. This paper uses both trace-based analysis and analytic modelling to show the potential advantages and drawbacks of inter-proxy cooperation. With our traces, we evaluate quantitatively the performance-improvement potential of cooperation between 200 small-organization proxies within a university environment, and between two largeorganization proxies handling 23,000 and 60,000 clients, respectively. With our model, we extend beyond these populations to project cooperative caching behavior in regions with millions of clients. Overall, we demonstrate that cooperative caching has performance benefits only within limited population bounds. We also use our model to examine the implications of future trends in Web-access behavior and traffic. 1 Introduction Cooperative caching -- the sharing and coordination of cache... Alec Wolman, Geoffrey M. Voelker, Nitin Sharma 0002, Neal Cardwell, Anna R. Karlin, Henry M. Levy |
SOSP | 4 |
| 1997 | The Energy Efficiency of IRAM ArchitecturesabstractPortable systems demand energy efficiency in order to maximize battery life. IRAM architectures, which combine DRAM and a processor on the same chip in a DRAM process, are more energy efficient than conventional systems. The high density of DRAM permits a much larger amount of memory on-chip than a traditional SRAM cache design in a logic process. This allows most or all IRAM memory accesses to be satisfied on-chip. Thus there is much less need to drive high-capacitance off-chip buses, which contribute significantly to the energy consumption of a system. To quantify this advantage we apply models of energy consumption in DRAM and SRAM memories to results from cache simulations of applications reflective of personal productivity tasks on low power systems. We find that IRAM memory hierarchies consume as little as 22% of the energy consumed by a conventional memory hierarchy for memory-intensive applications, while delivering comparable performance. Furthermore, the energy consumed by a system consisting of an IRAM memory hierarchy combined with an energy efficient CPU core is as little as 40% of that of the same CPU core with a traditional memory hierarchy. Richard Fromm, Stylianos Perissakis, Neal Cardwell, Christoforos E. Kozyrakis, Bruce McGaughy, David A. Patterson 0001, Thomas E. Anderson, Katherine A. Yelick |
ISCA | 3 |