EDBT 2026 Demo / reviewers in the wild / expert
Zhizhen Zhong
dblp:183/6387
· DBLP profile ↗
16ranked-venue papers
5as first author
13since 2021 · last 2025
0000-0003-3131-374XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 15 · 5 first-author · 12 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Toward Scalable Learning-Based Optical Restoration
Siyong Huang, Qingyu Song 0002, Zhaoning Wang, Zhizhen Zhong, Qiao Xiang, Jiwu Shu |
APNet | 5 |
| 2025 | Enabling Efficient GPU Communication over Multiple NICs with FuseLink
Zhenghang Ren, Zilong Wang 0007, Wenxue Li 0004, Kaiqiang Xu, Xudong Liao, Yijun Sun, Bowen Liu 0002, Han Tian, Junxue Zhang 0001, Mingfei Wang, Zhizhen Zhong, Guyue Liu, Ying Zhang 0022, Kai Chen 0005 |
OSDI | 13 |
| 2025 | MixNet: A Runtime Reconfigurable Optical-Electrical Fabric for Distributed Mixture-of-Experts TrainingabstractMixture-of-Expert (MoE) models outperform conventional models by selectively activating different subnets, named experts, on a per-token basis. This gated computation generates dynamic communications that cannot be determined beforehand, challenging the existing GPU interconnects that remain static during distributed training. In this paper, we advocate for a first-of-its-kind system, called MixNet, that unlocks topology reconfiguration during distributed MoE training. Towards this vision, we first perform a production measurement study and show that the MoE dynamic communication pattern has strong locality, alleviating the need for global reconfiguration. Based on this, we design and implement a regionally reconfigurable high-bandwidth domain that augments existing electrical interconnects using optical circuit switching (OCS), achieving scalability while maintaining rapid adaptability. We build a fully functional MixNet prototype with commodity hardware and a customized collective communication runtime. Our prototype trains state-of-the-art MoE models with in-training topology reconfiguration across 32 A100 GPUs. Large-scale packet-level simulations show that MixNet achieves performance comparable to a non-blocking fat-tree fabric while boosting the networking cost efficiency (e.g., performance per dollar) of four representative MoE models by 1.2×–1.5× and 1.9×–2.3× at 100 Gbps and 400 Gbps link bandwidths, respectively. Xudong Liao, Yijun Sun, Han Tian, Xinchen Wan, Yilun Jin, Zilong Wang 0007, Zhenghang Ren, Wenxue Li 0004, Kin Fai Tse, Zhizhen Zhong, Guyue Liu, Ying Zhang 0022, Xiaofeng Ye, Yiming Zhang 0003, Kai Chen 0005 |
SIGCOMM | 11 |
| 2025 | PreTE: Traffic Engineering with Predictive FailuresabstractFiber links in wide-area networks (WANs) are exposed to complicated environments and hence are vulnerable to failures like fiber cuts. The conventional approach of using static probabilistic failures falls short in fiber-cut scenarios because these fiber cuts are rare but disruptive, making it difficult for network operators to balance network utilization and availability in WAN traffic engineering. Our large-scale measurements of per-second optical-layer data reveal that the fiber's failure probability increases by several orders of magnitude when experiencing a rare and ephemeral degradation state. Therefore, we present a novel traffic engineering (TE) system called PreTE to factor in the dynamic fiber cut probabilities directly into TE systems. At the core of the PreTE system, fiber degradation facilitates failure predictions and traffic tunnels to be proactively updated, followed by traffic allocation optimizations among updated tunnels. We evaluate PreTE using a production-level WAN testbed and large-scale simulations. The testbed evaluation quantifies PreTE's runtime to demonstrate the feasibility to implement in large-scale WANs. Our large-scale simulation results show that PreTE can support up to 2× more demand at the same level of availability as compared to existing TE schemes. Congcong Miao, Zhizhen Zhong, Arpit Gupta, Ying Zhang 0022, Zekun He, Xianneng Zou, Jilong Wang 0001 |
SIGCOMM | 2 |
| 2024 | Understanding Communication Characteristics of Distributed TrainingabstractCommunication is pivotal in distributed training and a thorough understanding of its characteristics is essential for future optimizations. However, prior works are limited, either focusing on customized optimizations or conducting incomplete explorations on communication characteristics. In this work, we systematically analyze the communication characteristics of distributed training, considering two key aspects of communication: pattern and overhead, and assessing a broad spectrum of determinant factors. In particular, we extensively investigate the features of communication patterns, such as predictability, and comprehensively evaluate the impact of various factors on communication overhead. Additionally, we develop and validate an analytical formulation to estimate communication overhead, providing a mathematical understanding of models with predictability. Wenxue Li 0004, Xiangzhou Liu, Yilun Jin, Han Tian, Zhizhen Zhong, Guyue Liu, Ying Zhang 0022, Kai Chen 0005 |
APNet | 6 |
| 2024 | MegaTE: Extending WAN Traffic Engineering to Millions of Endpoints in Virtualized CloudabstractIn today's virtualized cloud, containers and virtual machines (VMs) are prevailing methods to deploy applications with different tenant requirements. However, these requirements are at odds with the resource allocation capabilities of conventional networking stacks in wide-area networks (WANs). In particular, existing WAN traffic engineering (TE) systems at the granularity of aggregated traffic flows are not designed to cater to each individual flow. In this paper, we advocate for a radical new approach to extend TE systems to involve millions of virtual instance endpoints. We propose and implement a first-of-its-kind system, called MegaTE, to satisfy the needs of each fine-grained traffic flow at the virtual instance level. At the core of the MegaTE system is the paradigm shift from the top-down centralized control to the bottom-up asynchronous query in the TE control loop, combined with eBPF-based segment routing on the data plane and TE optimization contraction on the control plane. We evaluate MegaTE using flow-level simulations with production traffic traces. Our results show that MegaTE supports 20× more endpoints with the similar algorithm run time compared to prior work. MegaTE has been adopted by large-scale public cloud providers. Notably, Tencent rolled out MegaTE in its cloud WAN since December 2022. Our production analysis shows that MegaTE reduces the packet latency of real-time applications by up to 51%. Congcong Miao, Zhizhen Zhong, Yunming Xiao, Senkuo Zhang, Yinan Jiang, Zizhuo Bai, Chaodong Lu, Jingyi Geng, Zekun He, Yachen Wang, Xianneng Zou, Chuanchuan Yang |
SIGCOMM | 2 |
| 2023 | On-Fiber Photonic ComputingabstractIn the 1800s, Charles Babbage envisioned computers as analog devices. However, it was not until 150 years later that a Mechanical Analog Computer was constructed for the US Navy to solve differential equations. With the end of Moore's Law, photonic computing is revitalizing the promise of analog computing by leveraging photons' speed, bandwidth, and energy efficiency for faster, more efficient, and scalable analog computing systems. This paper argues that the networking community should augment pluggable transponders with photonic computing capabilities to enable a backward-compatible solution for in-network computing. We propose on-fiber photonic computing to perform computing operations inside network transponders while the data is in the optical domain. We discuss the components required to enable the seamless integration of computation into the very fabric of optical communication links. We then discuss several use cases of on-fiber photonic computing, including machine learning inference, video encoding, load balancing, and intrusion detection. Mingran Yang, Zhizhen Zhong, Manya Ghobadi |
HotNets | 2 |
| 2023 | TopoOpt: Co-optimizing Network Topology and Parallelization Strategy for Distributed Training Jobs
Weiyang Wang, Moein Khazraee, Zhizhen Zhong, Manya Ghobadi, Dheevatsa Mudigere, Ying Zhang 0022, Anthony Kewitsch |
NSDI | 3 |
| 2023 | FlexWAN: Software Hardware Co-design for Cost-Effective and Resilient Optical BackbonesabstractThe rising demand for WAN capacity driven by the rapid growth of inter-data center traffic poses new challenges for costly optical networks. Today cloud providers rely on fixed optical backbones, where all hardware devices operate on a rigid spectrum grid, leading to the waste of expensive optical resources and subpar performance in handling failures. In this paper, we introduce FlexWAN, a novel flexible WAN infrastructure designed to provision cost-effective WAN capacity while ensuring resilience to optical failures. FlexWAN achieves this by incorporating spacing-variable hardware at the optical layer, enabling the generated wavelength to optimize the utilization of limited spectrum resources for the WAN capacity. The configuration of spacing-variable hardware in a multi-vendor optical backbone presents challenges related to spectrum management. To address this, FlexWAN leverages a centralized controller to achieve coordinated control of network-wide optical devices in a vendor-agnostic manner. Moreover, the flexibility at the optical layer introduces new algorithmic problems. FlexWAN formulates the problem of provisioning WAN capacity with the goal of minimizing hardware costs. We evaluate the system performance in production and share insights from years of production experience. Compared to existing optical backbones, FlexWAN can save at least 57% of transponders and reduce 36% of spectrum usage while continuing to meet up to 8× the present-day demands using existing hardware and fiber deployments. FlexWAN further incorporates failure resilience that revives 15% more bandwidth capacity in the overloaded optical backbone. Congcong Miao, Zhizhen Zhong, Ying Zhang 0022, Kunling He, Fangchao Li, Minggang Chen, Xiang Li 0223, Zekun He, Xianneng Zou, Jilong Wang 0001 |
SIGCOMM | 2 |
| 2023 | Demo: First Demonstration of Real-Time Photonic-Electronic DNN Acceleration on SmartNICsabstractWe demonstrate Lightning, a reconfigurable photonic-electronic deep learning smartNIC that serves real-time inference requests at 4.055 GHz compute frequency. To do so, Lightning uses a novel datapath to feed traffic from the NIC into its photonic computing cores without incurring digital data movement bottlenecks. Lightning achieves this by employing a reconfigurable count-action abstraction, which decouples the compute control plane from the data plane. The count-action abstraction counts the number of operations for each computation task in the Directed Acyclic Graph (DAG). It then triggers the execution of the next task(s) as soon as the previous task is finished without interrupting the dataflow. Our prototype shows that Lightning achieves 99.25% photonic MAC accuracy. When serving real-time inference requests, Lightning accelerates the end-to-end inference latency of the LeNet DNN by 9.4× and 6.6× compared to Nvidia P4 and A100 GPUs, respectively. Zhizhen Zhong, Mingran Yang, Jay Lang, Dirk R. Englund, Manya Ghobadi |
SIGCOMM | 1 |
| 2023 | Lightning: A Reconfigurable Photonic-Electronic SmartNIC for Fast and Energy-Efficient InferenceabstractThe massive growth of machine learning-based applications and the end of Moore's law have created a pressing need to redesign computing platforms. We propose Lightning, the first reconfigurable photonic-electronic smartNIC to serve real-time deep neural network inference requests. Lightning uses a fast datapath to feed traffic from the NIC into the photonic domain without creating digital packet processing and data movement bottlenecks. To do so, Lightning leverages a novel reconfigurable count-action abstraction that keeps track of the required computation operations of each inference packet. Our count-action abstraction decouples the compute control plane from the data plane by counting the number of operations in each task and triggers the execution of the next task(s) without interrupting the dataflow. We evaluate Lightning's performance using four platforms: a prototype, chip synthesis, emulations, and simulations. Our prototype demonstrates the feasibility of performing 8-bit photonic multiply-accumulate operations with 99.25% accuracy. To the best of our knowledge, our prototype is the highest-frequency photonic computing system, capable of serving real-time inference queries at 4.055 GHz end-to-end. Our simulations with large DNN models show that compared to Nvidia A100 GPU, A100X DPU, and Brainwave smartNIC, Lightning accelerates the average inference serve time by 337×, 329×, and 42×, while consuming 352×, 419×, and 54× less energy, respectively. Zhizhen Zhong, Mingran Yang, Jay Lang, Christian Williams, Liam Kronman, Alex Sludds, Homa Esfahanizadeh, Dirk R. Englund, Manya Ghobadi |
SIGCOMM | 1 |
| 2021 | A Social Network Under Social Distancing: Risk-Driven Backbone Management During COVID-19 and Beyond
Yiting Xia, Ying Zhang 0022, Zhizhen Zhong, Guanqing Yan, Chiunlin Lim, Satyajeet Ahuja, Soshant Bali, Alexander Nikolaidis, Kimia Ghobadi, Manya Ghobadi |
NSDI | 3 |
| 2021 | ARROW: restoration-aware traffic engineeringabstractFiber cut events reduce the capacity of wide-area networks (WANs) by several Tbps. In this paper, we revive the lost capacity by reconfiguring the wavelengths from cut fibers into healthy fibers. We highlight two challenges that made prior solutions impractical and propose a system called Arrow to address them. First, our measurements show that contrary to common belief, in most cases, the lost capacity is only partially restorable. This poses a cross-layer challenge from the Traffic Engineering (TE) perspective that has not been considered before: “Which IP links should be restored and by how much to best match the TE objective?” To address this challenge, Arrow's restoration-aware TE system takes a set of partial restoration candidates (that we call LotteryTickets) as input and proactively finds the best restoration plan. Second, prior work has not considered the reconfiguration latency of amplifiers. However, in practical settings, amplifiers add tens of minutes of reconfiguration delay. To enable fast and practical restoration, Arrow leverages optical noise loading and bypasses amplifier reconfiguration altogether. We evaluate Arrow using large-scale simulations and a testbed. Our testbed demonstrates Arrow's end-to-end restoration latency is eight seconds. Our large-scale simulations compare Arrow to the state-of-the-art TE schemes and show it can support 2.0x--2.4x more demand without compromising 99.99% availability. Zhizhen Zhong, Manya Ghobadi, Alaa Khaddaj, Jonathan Leach, Yiting Xia, Ying Zhang 0022 |
SIGCOMM | 1 |
| 2019 | Provisioning Short-Term Traffic Fluctuations in Elastic Optical NetworksabstractTransient traffic spikes are becoming a crucial challenge for network operators from both user-experience and network-maintenance perspectives. Different from long-term traffic growth, the bursty nature of short-term traffic fluctuations makes it difficult to be provisioned effectively. Luckily, next-generation elastic optical networks (EONs) provide an economical way to deal with such short-term traffic fluctuations. In this paper, we go beyond conventional network reconfiguration approaches by proposing the novel lightpath-splitting scheme in EONs. In lightpath splitting, we introduce the concept of SplitPoints to describe how lightpath splitting is performed. Lightpaths traversing multiple nodes in the optical layer can be split into shorter ones by SplitPoints to serve more traffic demands by raising signal modulation levels of lightpaths accordingly. We formulate the problem into a mathematical optimization model and linearize it into an integer linear program (ILP). We solve the optimization model on a small network instance and design scalable heuristic algorithms based on greedy and simulated annealing approaches. Numerical results show the tradeoff between throughput gain and negative impacts like traffic interruptions. Especially, by selecting SplitPoints wisely, operators can achieve almost twice as much throughput as conventional schemes without lightpath splitting. Zhizhen Zhong, Nan Hua, Massimo Tornatore, Jialong Li 0006, Yanhe Li, Xiaoping Zheng, Biswanath Mukherjee |
IEEE/ACM Trans. Netw. | 1 |
| 2018 | An Online Strategy for Service Degradation with Proportional QoS in Elastic Optical NetworksabstractElastic Optical Networks (EONs) represent a new approach for dealing with the enormous traffic demand in core networks as they can offer bandwidth granularities closer to those requested by the user and hence improve spectral utilization. In current literature there is a lack of dynamic strategies for service degradation which is a possible measure to address problems related to network congestion and consists in reducing the amount of resources provided. Since services of different classes can be requested, we propose in this paper an online strategy for service degradation using proportional Quality of Service (QoS). Our proposed strategy aims at minimizing the number of blocked requests due to lack of resources while provides throughput and delay guarantees for provisioned lightpaths. Thus, in order to quantify the impact of the degradation on the lightpaths we modeled source-destination pairs in an EON as a queuing system working under the Generalized Processor Sharing (GPS) service discipline with admission control of Leaky Bucket policy. The obtained results show that the proposed algorithm can reduce the blocking probability and give network operators more control between different degraded service classes. Alex S. Santos, Andre Horota, Zhizhen Zhong, Juliana de Santi, Gustavo B. Figueiredo, Massimo Tornatore, Biswanath Mukherjee |
ICC | 3 |
| 2016 | On QoS-Assured Degraded Provisioning in Service-Differentiated Multi-Layer Elastic Optical NetworksabstractDegraded provisioning provides an effective solution to flexibly allocate resources in various dimensions to reduce blocking for differentiated demands when network congestion occurs. In this work, we investigate the novel problem of online degraded provisioning in service-differentiated multi-layer networks with optical elasticity. Quality of Service (QoS) is assured by service-holding-time prolongation and immediate access as soon as the service arrives without set-up delay. We decompose the problem into degraded routing and degraded resource allocation stages, and design polynomial-time algorithms with the enhanced multi-layer architecture to exploit network flexibility in temporal and spectral dimensions. Numerical results verify that we can achieve significant blocking reduction, especially for requests with higher priorities. They also indicate that degradation in optical layer can increase the network capacity, while degradation in electric layer provides flexible time-bandwidth exchange. Zhizhen Zhong, Jipu Li, Nan Hua, Gustavo B. Figueiredo, Yanhe Li, Xiaoping Zheng, Biswanath Mukherjee |
GLOBECOM | 1 |