Behnaz Arzani

dblp:126/1063 · DBLP profile ↗
← Back
26ranked-venue papers
8as first author
15since 2021 · last 2026
0000-0003-2234-3250ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 25 · 8 first-author · 15 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 Heuristic Analysis from Source Code via Symbolic-Guided Optimization
Pantea Karimi, Siva Kesava Reddy K., Ryan Beckett, Santiago Segarra, Pooria Namyar, Mohammad Alizadeh, Behnaz Arzani
NSDI7
2026 Harvest: Adaptive Photonic Switching Schedules for Collective Communication in Scale-up Domains
abstract
As chip-to-chip silicon photonics gain traction for their bandwidth and energy efficiency, their circuit-switched nature raises a fundamental question for collective communication: when and how should the interconnect be reconfigured to realize these benefits? Establishing direct optical paths can reduce congestion and propagation delay; however, each reconfiguration incurs non-negligible overhead, making naive per-step reconfiguration impractical.
Mahir Rahman, Samuel Joseph, Nihar Kodkani, Behnaz Arzani, Vamsi Addanki
SIGCOMM4
2025 Everything Matters in Programmable Packet Scheduling
Albert Gran Alcoz, Balázs Vass, Pooria Namyar, Behnaz Arzani, Gábor Rétvári, Laurent Vanbever
NSDI4
2025 Enhancing Network Failure Mitigation with Performance-Aware Ranking
Pooria Namyar, Arvin Ghavidel, Daniel Crankshaw, Daniel S. Berger, Kevin Hsieh, Srikanth Kandula, Ramesh Govindan, Behnaz Arzani
NSDI8
2025 Raha: A General Tool to Analyze WAN Degradation
abstract
Raha is the first general tool that can analyze probable degradation of traffic engineered networks under arbitrary failures and traffic shifts to prevent outages. Raha addresses a significant gap in prior work which consider only (1) ≤ k failures; (2) specific traffic engineering schemes; and (3) the maximum impact of failures irrespective of the network design point.
Behnaz Arzani, Sina Taheri, Pooria Namyar, Ryan Beckett, Siva Kesava Reddy K., Elnaz Jalilipour
SIGCOMM1
2024 Towards Safer Heuristics With XPlain
abstract
Many problems that cloud operators solve are computationally expensive, and operators often use heuristic algorithms (that are faster and scale better than optimal) to solve them more efficiently. Heuristic analyzers enable operators to find when and by how much their heuristics underperform. However, these tools do not provide enough detail for operators to mitigate the heuristic's impact in practice: they only discover a single input instance that causes the heuristic to underperform (and not the full set) and they do not explain why.
Pantea Karimi, Solal Pirelli, Siva Kesava Reddy K., Ryan Beckett, Santiago Segarra, Beibin Li, Pooria Namyar, Behnaz Arzani
HotNets8
2024 End-to-End Performance Analysis of Learning-enabled Systems
abstract
We propose a performance analysis tool for learning-enabled systems that allows operators to uncover potential performance issues before deploying DNNs in their systems. The tools that exist for this purpose require operators to faithfully model all components (a white-box approach) or do inefficient black-box local search. We propose a gray-box alternative, which eliminates the need to precisely model all the system's components. Our approach is faster and finds substantially worse scenarios compared to prior work. We show that a state-of-the-art learning-enabled traffic engineering pipeline can underperform the optimal by 6× --- a much higher number compared to what the authors found.
Pooria Namyar, Michael Schapira, Ramesh Govindan, Santiago Segarra, Ryan Beckett, Siva Kesava Reddy K., Behnaz Arzani
HotNets7
2024 Finding Adversarial Inputs for Heuristics using Multi-level Optimization
Pooria Namyar, Behnaz Arzani, Ryan Beckett, Santiago Segarra, Himanshu Raj, Umesh Krishnaswamy, Ramesh Govindan, Srikanth Kandula
NSDI2
2024 Solving Max-Min Fair Resource Allocations Quickly on Large Graphs
Pooria Namyar, Behnaz Arzani, Srikanth Kandula, Santiago Segarra, Daniel Crankshaw, Umesh Krishnaswamy, Ramesh Govindan, Himanshu Raj
NSDI2
2024 Rethinking Machine Learning Collective Communication as a Multi-Commodity Flow Problem
abstract
Cloud operators utilize collective communication optimizers to enhance the efficiency of the single-tenant, centrally managed training clusters they manage. However, current optimizers struggle to scale for such use cases and often compromise solution quality for scalability. Our solution, TE-CCL, adopts a traffic-engineering-based approach to collective communication. Compared to a state-of-the-art optimizer, TACCL, TE-CCL produced schedules with 2× better performance on topologies TACCL supports (and its solver took a similar amount of time as TACCL's heuristic-based approach). TECCL additionally scales to larger topologies than TACCL. On our GPU testbed, TE-CCL outperformed TACCL by 2.14× and RCCL by 3.18× in terms of algorithm bandwidth.
Xuting Liu 0003, Behnaz Arzani, Siva Kesava Reddy K., Liangyu Zhao, Vincent Liu 0001, Srikanth Kandula, Luke Marshall
SIGCOMM2
2023 A Holistic View of AI-driven Network Incident Management
abstract
We discuss the potential improvement large language models (LLM) can provide in incident management and how they can overhaul the ways operators conduct incident management today. We propose a holistic framework for building an AI helper for incident management and discuss the several avenues of future research needed to achieve it.
Pouya Hamadanian, Behnaz Arzani, Sadjad Fouladi, Siva Kesava Reddy K., Rodrigo Fonseca, Denizcan Billor, Ahmad Cheema, Edet Nkposong, Ranveer Chandra
HotNets2
2022 Minding the gap between fast heuristics and their optimal counterparts
abstract
Production systems use heuristics because they are faster or scale better than the corresponding optimal algorithms. Yet, practitioners are often unaware of how worse off a heuristic's solution may be with respect to the optimum in realistic scenarios. Leveraging two-stage games and convex optimization, we present a provable framework that unveils settings where a given heuristic underperforms.
Pooria Namyar, Behnaz Arzani, Ryan Beckett, Santiago Segarra, Himanshu Raj, Srikanth Kandula
HotNets2
2021 Interpretable Feedback for AutoML and a Proposal for Domain-customized AutoML for Networking
abstract
The barrier to entry for network operators to use machine learning (ML) is high for operators who are not ML experts. Automated machine learning (AutoML) promises operators the ability to train ML models without requiring the expertise of data scientists or the need to learn ML. However, AutoML today: (a) is black-box; and (b) does not allow operators to leverage domain expertise. We start this paper by describing our broader vision for a domain-customized AutoML platform for networking and propose a set of potential solutions to realize that vision. As the first step, we introduce our feedback solution for AutoML that allows domain experts (who are not experts in ML) to better understand how to improve the input data to AutoML in order to achieve better accuracy.
Behnaz Arzani, Kevin Hsieh, Haoxian Chen 0001
HotNets1
2021 Towards a Cost vs. Quality Sweet Spot for Monitoring Networks
abstract
Continuously monitoring a wide variety of performance and fault metrics has become a crucial part of operating large-scale datacenter networks. In this work, we ask whether we can reduce the costs to monitor - in terms of collection, storage and analysis - by judiciously controlling how much and which measurements we collect. By positing that we can treat almost all measured signals as sampled time-series, we show that we can use signal processing techniques such as the Nyquist-Shannon theorem to avoid wasteful data collection. We show that large savings appear possible by analyzing tens of popular measurement systems from a production datacenter network. We also discuss some challenges that must be solved when applying these techniques in practice.
Nofel Yaseen, Behnaz Arzani, Krishna Chintalapudi, Vaishnavi Nattar Ranganathan, Felipe Vieira Frujeri, Kevin Hsieh, Daniel S. Berger, Vincent Liu 0001, Srikanth Kandula
HotNets2
2021 Contracting Wide-area Network Topologies to Solve Flow Problems Quickly
Firas Abuzaid, Srikanth Kandula, Behnaz Arzani, Ishai Menache, Matei Zaharia, Peter Bailis
NSDI3
2020 PrivateEye: Scalable and Privacy-Preserving Compromise Detection in the Cloud
Behnaz Arzani, Selim Ciraci, Stefan Saroiu, Alec Wolman, Jack W. Stokes, Geoff Outhred, Lechao Diwu
NSDI1
2020 Aragog: Scalable Runtime Verification of Shardable Networked Systems
Nofel Yaseen, Behnaz Arzani, Ryan Beckett, Selim Ciraci, Vincent Liu 0001
OSDI2
2020 Scouts: Improving the Diagnosis Process Through Domain-customized Incident Routing
abstract
Incident routing is critical for maintaining service level objectives in the cloud: the time-to-diagnosis can increase by 10x due to mis-routings. Properly routing incidents is challenging because of the complexity of today's data center (DC) applications and their dependencies. For instance, an application running on a VM might rely on a functioning host-server, remote-storage service, and virtual and physical network components. It is hard for any one team, rule-based system, or even machine learning solution to fully learn the complexity and solve the incident routing problem. We propose a different approach using per-team Scouts. Each teams' Scout acts as its gate-keeper --- it routes relevant incidents to the team and routes-away unrelated ones. We solve the problem through a collection of these Scouts. Our PhyNet Scout alone --- currently deployed in production --- reduces the time-to-mitigation of 65% of mis-routed incidents in our dataset.
Nofel Yaseen, Robert MacDavid, Felipe Vieira Frujeri, Vincent Liu 0001, Ricardo Bianchini, Ramaswamy Aditya, Xiaohang Wang 0008, Henry Lee, David A. Maltz, Minlan Yu, Behnaz Arzani
SIGCOMM12
2020 GRooT: Proactive Verification of DNS Configurations
abstract
The Domain Name System (DNS) plays a vital role in today's Internet but relies on complex distributed management of records. DNS misconfiguration related outages have rendered popular services like GitHub, HBO, LinkedIn, and Azure inaccessible for extended periods. This paper introduces GRoot, the first verifier that performs static analysis of DNS configuration files, enabling proactive and exhaustive checking for common DNS bugs; by contrast, existing solutions are reactive and incomplete. GRoot uses a new, fast verification algorithm based on generating and enumerating DNS query equivalence classes. GRoot symbolically executes the set of queries in each equivalence class to efficiently find (or prove the absence of) any bugs such as rewrite loops. To prove the correctness of our approach, we develop a formal semantic model of DNS resolution. Applied to the configuration files from a campus network with over a hundred thousand records, GRoot revealed 109 bugs within seconds. When applied to internal zone files consisting of over 3.5 million records from a large infrastructure service provider, GRoot revealed around 160k issues of blackholing, initiating a cleanup. Finally, on a synthetic dataset with over 65 million real records, we find GRoot can scale to networks with tens of millions of records.
Siva Kesava Reddy K., Ryan Beckett, Behnaz Arzani, Todd D. Millstein, George Varghese
SIGCOMM3
2019 dShark: A General, Easy to Program and Scalable Framework for Analyzing In-network Packet Traces
Behnaz Arzani, Rodrigo Fonseca, Tianrong Zhang, Karl Deng
NSDI3
2018 Cloud Datacenter SDN Monitoring: Experiences and Challenges
Deepak Bansal, David Brumley, Harish Kumar Chandrappa, Parag Sharma, Rishabh Tewari, Behnaz Arzani, Alex C. Snoeren
Internet Measurement Conference7
2018 007: Democratically Finding the Cause of Packet Drops
Behnaz Arzani, Selim Ciraci, Luiz F. O. Chamon, Yibo Zhu 0001, Hongqiang Liu, Jitendra Padhye, Boon Thau Loo, Geoff Outhred
NSDI1
2016 Taking the Blame Game out of Data Centers Operations with NetPoirot
abstract
Today, root cause analysis of failures in data centers is mostly done through manual inspection. More often than not, cus- tomers blame the network as the culprit. However, other components of the system might have caused these failures. To troubleshoot, huge volumes of data are collected over the entire data center. Correlating such large volumes of diverse data collected from different vantage points is a daunting task even for the most skilled technicians. In this paper, we revisit the question: how much can you infer about a failure in the data center using TCP statistics collected at one of the endpoints? Using an agent that cap- tures TCP statistics we devised a classification algorithm that identifies the root cause of failure using this information at a single endpoint. Using insights derived from this classi- fication algorithm we identify dominant TCP metrics that indicate where/why problems occur in the network. We val- idate and test these methods using data that we collect over a period of six months in a production data center.
Behnaz Arzani, Selim Ciraci, Boon Thau Loo, Assaf Schuster, Geoff Outhred
SIGCOMM1
2014 Multi-path Solutions to Improve Network Performance
abstract
With the increasing number of internet users, providing performance guarantees to these users is becoming increasingly difficult. In wireless settings, the problem of providing network performance guarantees is even more difficult due to variabilities in link characteristics. Changes in wireless link performance, such as packet loss rate, occur at shorter time scales compared to their wired counterparts. Thus, the finite reaction time of traditional solutions, e.g. Routing, make them unreliable in mitigating the impact of these variations on user performance. Such variability in the delivered rates could be detrimental to particular types of applications such as video and audio which are sensitive to short term variations.
Behnaz Arzani
ICNP1
2014 Deconstructing MPTCP Performance
abstract
The paper seeks to broaden our understanding of MPTCP and focuses on the impact that initial sub-path selection can have on performance. Using empirical data, it demonstrates that which sub-path is chosen to start an MPTCP connection can have unintuitive consequences. Using numerical analysis and a model-driven investigation, the paper elucidates and validates the empirical results, and highlights MPTCP's non-linear coupling between paths as a primary cause for this behavior. The findings are both of operational interest and may help design better MPTCP schedulers, as they are also exposed to complex interactions with MPTCP's congestion control.
Behnaz Arzani, Alexander J. T. Gurney, Sitian Cheng, Roch Guérin, Boon Thau Loo
ICNP1
2012 A distributed routing protocol for predictable rates in wireless mesh networksy
abstract
Wireless mesh networks hold the promise of rapid and flexible deployments of communication facilities. This potential notwithstanding, the often erratic behavior of multihop wireless transmissions is limiting the range of applications that such networks can target. In this paper we investigate the feasibility and benefits of a routing protocol explicitly aimed at making wireless mesh networks more predictable while preserving their efficiency and flexibility. The protocol's basic premise is the classical idea that a multipath solution can offer resiliency to unexpected link variations. The paper's contributions are in demonstrating how this can be effectively realized in a wireless context, and in offering initial evidences of its efficacy. In particular, the paper illustrates how routing decisions that account for link variability can be computed in a distributed fashion, and the benefits they afford in improving the stability of end-to-end transmission rates even in the presence of random network fluctuations.
Behnaz Arzani, Roch Guérin, Alejandro Ribeiro
ICNP1