Nadeen Gebara

dblp:243/5218 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
5since 2021 · last 2026
0009-0001-9071-6621ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 1 first-author · 3 since 2021Computer networks · 3 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 EZCache: Easy Action-Enabled FPGA Caches for Non-Stalling Datapaths in SmartNICs and Beyond
abstract
Caches are widely used in FPGA accelerators such as SmartNICs to hide DRAM latency, but conventional designs treat caches as passive storage. When workloads require read–modify–write (RMW) updates – such as flow tables, counters, or per-connection state – existing cache IPs force designers to either stall the pipeline or duplicate hazard-handling logic outside the cache. Both approaches waste bandwidth, complicate datapath design, and require extensive verification.We propose EZCache, a new FPGA cache IP that integrates programmable action blocks directly into the cache pipeline. These blocks perform user-defined operations on cached data in place, allowing RMW updates to complete without stalling other operations or exposing hazards to surrounding logic. By embedding compute into the cache, EZCache transforms it from a passive buffer into an active architectural primitive suitable for a wide range of FPGA accelerators.EZCache is fully parametric in associativity, latency, throughput, and action complexity, enabling a single design to support diverse workloads and memory systems. This parametric abstraction allows the surrounding datapath to remain stable even as cache configurations, action semantics, and external memories evolve across hardware generations. EZCache has been deployed across five generations of SmartNICs and millions of devices worldwide, proving both its maturity and its impact in production. Results across multiple FPGA platforms show EZCache’s generality and efficiency, establishing it as a new foundation for cache-based acceleration in reconfigurable systems.
Ahmed Abdelsalam, Vishal Gondaliya, Ezz Hamed, Pragati Medleri Hire Math, Marc Gepigon, Joshua Landgraf, Nadeen Gebara, Bob Groza, Anshuman Verma, Andrew Putnam
FCCM7
2026 Rules Offload Engine (ROE): Accelerating Host SDN Policy Evaluation
Anshuman Verma, Tian Tan 0007, Ahmed Abdelsalam, Milan Dasgupta, Jonathan Hunter, Zach Libby, Narayanan Ravichandran, Harish Srinivasan, Matt Reat, Nadeen Gebara, Vishal Gondaliya, Ezz Hamed, Rahul Garlapati, Lok Chand Koppaka, Abdullah Mughrabi, Dev Desai, Alexander Malysh, Shwetha Bhat, Rohan Kandi, Megan Sng, Tushar Garg, Muluken Hailesellasie, Andrew Putnam, Derek Chiou, Osman Ertugay, Alireza Dabagh, Vivek Bhanu, Daniel Firestone
SIGCOMM10
2025 Miniature: Fast AI Supercomputer Networks Simulation on FPGAs
Yicheng Qian, Ran Shu 0001, Rui Ma 0021, Yang Wang 0053, Derek Chiou, Nadeen Gebara, Luca Piccolboni, Miriam Leeser, Yongqiang Xiong
APNet6
2025 Uno: A One-Stop Solution for Inter- and Intra-Data Center Congestion Control and Reliable Connectivity
abstract
Cloud computing and AI workloads are driving unprecedented demand for efficient communication within and across datacenters. However, the coexistence of intra- and inter-datacenter traffic within datacenters plus the disparity between the RTTs of intra- and inter-datacenter networks complicates congestion management and traffic routing. Particularly, faster congestion responses of intra-datacenter traffic causes rate unfairness when competing with slower inter-datacenter flows. Additionally, inter-datacenter messages suffer from slow loss recovery and, thus, require reliability. Existing solutions overlook these challenges and handle inter- and intra-datacenter congestion with separate control loops or at different granularities. We propose Uno, a unified system for both inter- and intra-DC environments that integrates a transport protocol for rapid congestion reaction and fair rate control with a load balancing scheme that combines erasure coding and adaptive routing. Our findings show that Uno significantly improves the completion times of both inter- and intra-DC flows compared to state-of-the-art methods such as Gemini.
Tommaso Bonato, Sepehr Abdous, Abdul Kabbani, Ahmad Ghalayini, Nadeen Gebara, Terry Lam, Anup Agarwal, Tiancheng Chen, Zhuolong Yu, Konstantin Taranov, Mahmoud Elhaddad, Daniele De Sensi, Soudeh Ghorbani, Torsten Hoefler
SC5
2022 Hipernetch: High-Performance FPGA Network Switch
abstract
We present Hipernetch, a novel FPGA-based design for performing high-bandwidth network switching. FPGAs have recently become more popular in data centers due to their promising capabilities for a wide range of applications. With the recent surge in transceiver bandwidth, they could further benefit the implementation and refinement of network switches used in data centers. Hipernetch replaces the crossbar with a “combined parallel round-robin arbiter”. Unlike a crossbar, the combined parallel round-robin arbiter is easy to pipeline, and does not require centralised iterative scheduling algorithms that try to fit too many steps in a single or a few FPGA cycles. The result is a network switch implementation on FPGAs operating at a high frequency and with a low port-to-port latency. Our proposed Hipernetch architecture additionally provides a competitive switching performance approaching output-queued crossbar switches. Our implemented Hipernetch designs exhibit a throughput that exceeds 100 Gbps per port for switches of up to 16 ports, reaching an aggregate throughput of around 1.7 Tbps.
Philippos Papaphilippou, Jiuxi Meng, Nadeen Gebara, Wayne Luk
ACM Trans. Reconfigurable Technol. Syst.3
2020 Fast and Accurate Training of Ensemble Models with FPGA-based Switch
abstract
Random projection is gaining more attention in large scale machine learning. It has been proved to reduce the dimensionality of a set of data whilst approximately preserving the pairwise distance between points by multiplying the original dataset with a chosen matrix. However, projecting data to a lower dimension subspace typically reduces the training accuracy. In this paper, we propose a novel architecture that combines an FPGA-based switch with the ensemble learning method. This architecture enables reducing training time while maintaining high accuracy. Our initial result shows a speedup of 2.12-6.77 times using four different high dimensionality datasets.
Jiuxi Meng, Ce Guo 0002, Nadeen Gebara, Wayne Luk
ASAP3
2020 Challenging the Stateless Quo of Programmable Switches
abstract
Programmable switches based on the Protocol Independent Switch Architecture (PISA) have greatly enhanced the flexibility of today's networks by allowing new packet protocols to be deployed without any hardware changes. They have also been instrumental in enabling a new computing paradigm in which parts of an application's logic run within the network core (in-network computing).
Nadeen Gebara, Alberto Lerner, Mingran Yang, Minlan Yu, Paolo Costa, Manya Ghobadi
HotNets1
2019 Investigating the Feasibility of FPGA-based Network Switches
abstract
FPGAs are being increasingly used on network interface cards (NICs) as offload units to accelerate packet processing tasks. The rationale is that by customizing the NIC logic it is possible to achieve higher performance for the most critical tasks while eliminating unnecessary logic, thus improving overall efficiency. In this paper, we aim to investigate if similar benefits can also be extended to network switches. We compare different switch architectures and analyze their suitability to an FPGA implementation. We discuss several optimization techniques to overcome the challenges of limited FPGA resources and assess the scalability of our designs up to 10, 25, and 50~Gb/s throughput per port.
Jiuxi Meng, Nadeen Gebara, Ho-Cheung Ng, Paolo Costa, Wayne Luk
ASAP2
2018 Scheduling Algorithms for High Performance Network Switching on FPGAs: A Survey
abstract
The scheduling algorithm used in a network switch significantly impacts the switch's performance and thereby the performance of the entire network. To keep up with the ongoing demands for higher network performance, a myriad of scheduling algorithms have been investigated. We propose that FPGAs can be outstanding candidates for benchmarking scheduling algorithms, and that it can be beneficial to have customized scheduling algorithms which are enabled by FPGA based switches due to their reconfigurable architectures. This paper presents the first FPGA targeted survey on high performance scheduling algorithms used in the most popular switch architecture, input-buffered crossbars, with the aim of guiding future research on high performance network switching.
Nadeen Gebara, Jiuxi Meng, Wayne Luk, Paolo Costa
FPT1