Hassan M. G. Wassel

dblp:01/1085 · DBLP profile ↗
← Back
13ranked-venue papers
1as first author
6since 2021 · last 2025
0009-0004-4090-5987ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 6 · 3 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 2 since 2021Systems, architecture and hardware · 4 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Falcon: A Reliable, Low Latency Hardware Transport
abstract
Hardware transports such as RoCE deliver high performance with minimal host CPU, but are best suited to special-purpose deployments that limit their use, e.g., backend networks or Ethernet with Priority Flow Control (PFC). We introduce Falcon, the first hardware transport that supports multiple Upper Layer Protocols (ULPs) and heterogeneous application workloads in general-purpose Ethernet datacenter environments (with losses and without special switch support). Key design elements include: delay-based congestion control with multipath load balancing; a layered design with a simple request-response transaction interface for multi-ULP support; hardware-based retransmissions and error-handling for scalability; and a programmable engine for flexibility. The first Falcon hardware implementation delivers a peak performance of 200 Gbps, 120 Mops/sec, with near-optimal operation completion times that are up to 8× lower than CX-7 RoCE under network congestion, and up to 65% higher goodput under lossy conditions.
Arjun Singhvi, Nandita Dukkipati, Prashant Chandra, Hassan M. G. Wassel, Naveen Kr. Sharma, Anthony Rebello, Henry Schuh, Praveen Kumar 0003, Behnam Montazeri, Neelesh Bansod, Sarin Thomas, Inho Cho, Hyojeong Lee Seibert, Baijun Wu, Rui Yang 0034, Qianwen Yin, Srinivas Vaduvatha, Weihuang Wang, Masoud Moshref, David Wetherall, Amin Vahdat
SIGCOMM4
2023 RingLeader: Efficiently Offloading Intra-Server Orchestration to NICs
Adney Cardoza, Tarannum Khan, Yeonju Ro, Brent E. Stephens, Hassan M. G. Wassel, Aditya Akella
NSDI6
2023 A Cloud-Scale Characterization of Remote Procedure Calls
abstract
The global scale and challenging requirements of modern cloud applications have led to the development of complex, widely distributed, service-oriented applications. One enabler of such applications is the remote procedure call (RPC), which provides location-independent communication and hides the myriad of cloud communication complexities and requirements within the RPC stack. Understanding RPCs is thus one key to understanding the behavior of cloud applications. While there have been numerous studies of RPCs in distributed systems, as well as attempts to optimize RPC overheads with both software and hardware, there is still a lack of knowledge about the characteristics of RPCs "in the wild" in the modern cloud environment.
Korakit Seemakhupt, Brent E. Stephens, Samira Manabi Khan, Sihang Liu 0001, Hassan M. G. Wassel, Soheil Hassas Yeganeh, Alex C. Snoeren, Arvind Krishnamurthy, David E. Culler, Henry M. Levy
SOSP5
2022 Aquila: A unified, low-latency fabric for datacenter networks
Dan Gibson, Hema Hariharan, Eric Lance, Moray McLaren, Behnam Montazeri, Hassan M. G. Wassel, Zhehua Wu, Sunghwan Yoo, Raghuraman Balasubramanian, Prashant Chandra, Michael Cutforth, Peter Cuy, David Decotigny, Rakesh Gautam, Alex Iriza, Milo M. K. Martin, Rick Roy, Zuowei Shen, Monica Wong-Chan, Joe Zbiciak, Amin Vahdat
NSDI8
2022 Carbink: Fault-Tolerant Far Memory
Yang Zhou 0008, Hassan M. G. Wassel, Sihang Liu 0001, James W. Mickens, Minlan Yu, Chris Kennelly, David E. Culler, Henry M. Levy, Amin Vahdat
OSDI2
2022 Hashing Design in Modern Networks: Challenges and Mitigation Techniques
Yunhong Xu, Keqiang He, Rui Wang 0025, Minlan Yu, Nick G. Duffield, Hassan M. G. Wassel, Shidong Zhang, Leonid B. Poutievski, Junlan Zhou, Amin Vahdat
USENIX ATC6
2020 Sundial: Fault-tolerant Clock Synchronization for Datacenters
Gautam Kumar 0001, Hema Hariharan, Hassan M. G. Wassel, Peter Hochschild, Dave Platt, Simon L. Sabato, Minlan Yu, Nandita Dukkipati, Prashant Chandra, Amin Vahdat
OSDI4
2020 Swift: Delay is Simple and Effective for Congestion Control in the Datacenter
abstract
We report on experiences with Swift congestion control in Google datacenters. Swift targets an end-to-end delay by using AIMD control, with pacing under extreme congestion. With accurate RTT measurement and care in reasoning about delay targets, we find this design is a foundation for excellent performance when network distances are well-known. Importantly, its simplicity helps us to meet operational challenges. Delay is easy to decompose into fabric and host components to separate concerns, and effortless to deploy and maintain as a congestion signal while the datacenter evolves. In large-scale testbed experiments, Swift delivers a tail latency of <50μs for short RPCs, with near-zero packet drops, while sustaining ~100Gbps throughput per server. This is a tail of <3x the minimal latency at a load close to 100%. In production use in many different clusters, Swift achieves consistently low tail completion times for short RPCs, while providing high throughput for long RPCs. It has loss rates that are at least 10x lower than a DCTCP protocol, and handles O(10k) incasts that sharply degrade with DCTCP.
Gautam Kumar 0001, Nandita Dukkipati, Keon Jang, Hassan M. G. Wassel, Xian Wu 0001, Behnam Montazeri, Yaogong Wang, Kevin Springborn, Christopher Alfeld, Michael Ryan, David Wetherall, Amin Vahdat
SIGCOMM4
2020 1RMA: Re-envisioning Remote Memory Access for Multi-tenant Datacenters
abstract
Remote Direct Memory Access (RDMA) plays a key role in supporting performance-hungry datacenter applications. However, existing RDMA technologies are ill-suited to multi-tenant datacenters, where applications run at massive scales, tenants require isolation and security, and the workload mix changes over time. Our experiences seeking to operationalize RDMA at scale indicate that these ills are rooted in standard RDMA's basic design attributes: connectionorientedness and complex policies baked into hardware.
Arjun Singhvi, Aditya Akella, Dan Gibson, Thomas F. Wenisch, Monica Wong-Chan, Sean Clark 0003, Milo M. K. Martin, Moray McLaren, Prashant Chandra, Rob Cauble, Hassan M. G. Wassel, Behnam Montazeri, Simon L. Sabato, Joel Scherpelz, Amin Vahdat
SIGCOMM11
2015 TIMELY: RTT-based Congestion Control for the Datacenter
abstract
Datacenter transports aim to deliver low latency messaging together with high throughput. We show that simple packet delay, measured as round-trip times at hosts, is an effective congestion signal without the need for switch feedback. First, we show that advances in NIC hardware have made RTT measurement possible with microsecond accuracy, and that these RTTs are sufficient to estimate switch queueing. Then we describe how TIMELY can adjust transmission rates using RTT gradients to keep packet latency low while delivering high bandwidth. We implement our design in host software running over NICs with OS-bypass capabilities. We show using experiments with up to hundreds of machines on a Clos network topology that it provides excellent performance: turning on TIMELY for OS-bypass messaging over a fabric with PFC lowers 99 percentile tail latency by 9X while maintaining near line-rate throughput. Our system also outperforms DCTCP running in an optimized kernel, reducing tail latency by $13$X. To the best of our knowledge, TIMELY is the first delay-based congestion control protocol for use in the datacenter, and it achieves its results despite having an order of magnitude fewer RTT signals (due to NIC offload) than earlier delay-based schemes such as Vegas.
Radhika Mittal, Vinh The Lam, Nandita Dukkipati, Emily R. Blem, Hassan M. G. Wassel, Manya Ghobadi, Amin Vahdat, Yaogong Wang, David Wetherall, David Zats
SIGCOMM5
2013 SurfNoC: a low latency and provably non-interfering approach to secure networks-on-chip
abstract
As multicore processors find increasing adoption in domains such as aerospace and medical devices where failures have the potential to be catastrophic, strong performance isolation and security become first-class design constraints. When cores are used to run separate pieces of the system, strong time and space partitioning can help provide such guarantees. However, as the number of partitions or the asymmetry in partition bandwidth allocations grows, the additional latency incurred by time multiplexing the network can significantly impact performance.
Hassan M. G. Wassel, Ying Gao 0001, Jason Oberg, Ted Huffmire, Ryan Kastner, Fred Chong, Timothy Sherwood
ISCA1
2009 Complete information flow tracking from the gates up
abstract
For many mission-critical tasks, tight guarantees on the flow of information are desirable, for example, when handling important cryptographic keys or sensitive financial data. We present a novel architecture capable of tracking all information flow within the machine, including all explicit data transfers and all implicit flows (those subtly devious flows caused by not performing conditional operations). While the problem is impossible to solve in the general case, we have created a machine that avoids the general-purpose programmability that leads to this impossibility result, yet is still programmable enough to handle a variety of critical operations such as public-key encryption and authentication. Through the application of our novel gate-level information flow tracking method, we show how all flows of information can be precisely tracked. From this foundation, we then describe how a class of architectures can be constructed, from the gates up, to completely capture all information flows and we measure the impact of doing so on the hardware implementation, the ISA, and the programmer.
Mohit Tiwari, Hassan M. G. Wassel, Bita Mazloom, Shashidhar Mysore, Fred Chong, Timothy Sherwood
ASPLOS2
2009 Execution leases: a hardware-supported mechanism for enforcing strong non-interference
abstract
High assurance systems such as those found in aircraft controls and the financial industry are often required to handle a mix of tasks where some are niceties (such as the control of media for entertainment, or supporting a remote monitoring interface) while others are absolutely critical (such as the control of safety mechanisms, or maintaining the secrecy of a root key). While special purpose languages, careful code reviews, and automated theorem proving can be used to help mitigate the risk of combining these operations onto a single machine, it is difficult to say if any of these techniques are truly complete because they all assume a simplified model of computation far different from an actual processor implementation both in functionality and timing. In this paper we propose a new method for creating architectures that both a) makes the complete information-flow properties of the machine fully explicit and available to the programmer and b) allows those properties to be verified all the way down to the gate-level implementation the design.
Mohit Tiwari, Xun Li 0001, Hassan M. G. Wassel, Fred Chong, Timothy Sherwood
MICRO3