VLDB 2026 Research / reviewers in the wild / expert
Shizhen Zhao
dblp:90/10156
· DBLP profile ↗
46ranked-venue papers
17as first author
34since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 33 · 11 first-author · 23 since 2021Artificial intelligence and machine learning · 11 · 6 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 7 since 2021Systems, architecture and hardware · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ASSIST-3D: Adapted Scene Synthesis for Class-Agnostic 3D Instance SegmentationabstractClass-agnostic 3D instance segmentation tackles the challenging task of segmenting all object instances, including previously unseen ones, without semantic class reliance. Current methods struggle with generalization due to the scarce annotated 3D scene data or noisy 2D segmentations. While synthetic data generation offers a promising solution, existing 3D scene synthesis methods fail to simultaneously satisfy geometry diversity, context complexity, and layout reasonability, each essential for this task. To address these needs, we propose an Adapted 3D Scene Synthesis pipeline for class-agnostic 3D Instance SegmenTation, termed as ASSIST-3D, to synthesize proper data for model generalization enhancement. Specifically, ASSIST-3D features three key innovations, including 1) Heterogeneous Object Selection from extensive 3D CAD asset collections, incorporating randomness in object sampling to maximize geometric and contextual diversity; 2) Scene Layout Generation through LLM-guided spatial reasoning combined with depth-first search for reasonable object placements; and 3) Realistic Point Cloud Construction via multi-view RGB-D image rendering and fusion from the synthetic scenes, closely mimicking real-world sensor data acquisition. Experiments on ScanNetV2, ScanNet++, and S3DIS benchmarks demonstrate that models trained with ASSIST-3D-generated data significantly outperform existing methods. Further comparisons underscore the superiority of our purpose-built pipeline over existing 3D scene synthesis approaches. Shengchao Zhou, Jiehong Lin, Jiahui Liu 0012, Shizhen Zhao, Chirui Chang, Xiaojuan Qi 0001 |
AAAI | 4 |
| 2026 | ATRO: A Fast Algorithm for Topology Engineering of Reconfigurable Datacenter Networks
Yingming Mao, Qiaozhu Zhai, Ximeng Liu, Xinchi Han, Fanfan Li, Shizhen Zhao, Yuzhou Zhou, Zhen Yao 0003 |
INFOCOM | 6 |
| 2026 | Geminet: Learning the Duality-based Topology-Agnostic Update Operator for Lightweight Traffic Engineering in Changing Topologies
Ximeng Liu, Yingming Mao, Yatao Li, Shizhen Zhao, Xinbing Wang |
NSDI | 5 |
| 2026 | NegotiaToR: Toward a Simple Yet Effective On-Demand Reconfigurable Datacenter NetworkabstractRecent advances in fast optical switching show promise in meeting the high goodput and low latency requirements of datacenter networks. We present NegotiaToR, a simple network architecture for optical reconfigurable DCNs that utilizes on-demand scheduling to handle dynamic traffic. In NegotiaToR, racks exchange scheduling messages through an in-band control plane and distributedly calculate non-conflicting paths from binary traffic demand information. Optimized for incasts, it also provides opportunities to bypass scheduling delays. NegotiaToR is compatible with prevalent flat topologies, and is tailored towards a minimalist design for on-demand reconfigurable DCNs, enhancing practicality. Through large-scale simulations, we show that NegotiaToR achieves both small mice flow completion time and high goodput on two representative flat topologies, especially under heavy loads. Particularly, the flow completion time of mice flows is one to two orders of magnitude better than the state-of-the-art traffic-oblivious reconfigurable DCN design. Cong Liang 0005, Xiangli Song, Mowei Wang, Yashe Liu, Zhenhua Liu 0008, Shizhen Zhao, Yong Cui 0001 |
IEEE Trans. Netw. | 8 |
| 2026 | Loss-Tolerant RDMA Network Over Commodity DevicesabstractThis paper proposes the concept of a “loss-tolerant” RDMA network, instantiating as NüWa. It reveals the fundamental issues under a lossy fabric — packet losses and repetitive retransmission timeouts (RTOs) cause severe performance degradation and even service interruption. The loss-tolerant RDMA must avoid “important” packet losses that trigger RTOs. However, existing loss-protection mechanisms fail to identify these packets precisely. They either generate massive misprotection or ignore selective repeat loss recovery, resulting in buffer overflows and failures of RTO protection. To tackle these issues, NüWa thoroughly analyzes distinct loss-recovery schemes and RTO reasons for commodity NICs. It designsswitch modeandNIC modeto accurately identify and protect all important packets. The switch mode inherently supports the widely deployed non-programmable NICs, while the NIC mode offloads identification complexity to advanced programmable NICs. With effective RTO avoidance, it improves flow completion time (FCT) by 2 ∼ 10× compared to state-of-the-art (SOTA) solutions. In severe incast and large-scale networks, it reduces FCT by 100× compared with a lossless fabric. In storage applications, N¨uWa improves IOPS by ∼ 300% compared to vanilla lossy fabric. For typical AI Workloads, it accelerates AllReduce/AlltoAll communication by 5.5 ∼ 13.7× compared to SOTA lossy network solutions. Likai Wang 0013, Zhe Wang 0015, Yimu Yuan, Shuhan Tian, Linghe Kong, Qiao Xiang, Shizhen Zhao, Di Qu, Hexiang Song, Yashar Ganjali, Guihai Chen |
IEEE Trans. Netw. | 10 |
| 2025 | FauTE: Fault-tolerant Traffic Engineering in Data Center Network
Yang Liu 0437, Jingyi Cheng, Ximeng Liu, Shizhen Zhao |
APNet | 5 |
| 2025 | Learning from Neighbors: Category Extrapolation for Long-Tail LearningabstractBalancing training on long-tail data distributions remains a long-standing challenge in deep learning. While methods such as re-weighting and re-sampling help alleviate the imbalance issue, limited sample diversity continues to hinder models from learning robust and generalizable feature representations, particularly for tail classes. In contrast to existing methods, we offer a novel perspective on long-tail learning, inspired by an observation: datasets with finer granularity tend to be less affected by data imbalance. In this paper, we investigate this phenomenon through both quantitative and qualitative studies, showing that increased granularity enhances the generalization of learned features in tail categories. Motivated by these findings, we propose a method to increase dataset granularity through category extrapolation. Specifically, we introduce open-set fine-grained classes that are related to existing ones, aiming to enhance representation learning for both head and tail classes. To automate the curation of auxiliary data, we leverage large language models (LLMs) as knowledge bases to search for auxiliary categories and retrieve relevant images through web crawling. To prevent the overwhelming presence of auxiliary classes from disrupting training, we introduce a neighbor-silencing loss that encourages the model to focus on class discrimination within the target dataset. During inference, the classifier weights for auxiliary categories are masked out, leaving only the target class weights for use. Extensive experiments on three standard long-tail benchmarks demonstrate the effectiveness of our approach, notably outperforming strong baseline methods that use the same amount of data. The project is available at shizhen-zhao.github.io/LT_Neighbors/. Shizhen Zhao, Xin Wen 0004, Jiahui Liu 0012, Chuofan Ma, Chunfeng Yuan, Xiaojuan Qi 0001 |
CVPR | 1 |
| 2025 | Aligning Effective Tokens with Video Anomaly in Large Language ModelsabstractUnderstanding abnormal events in videos is a vital and challenging task that has garnered significant attention in a wide range of applications. Although current video understanding Multi-modal Large Language Models (MLLMs) are capable of analyzing general videos, they often struggle to handle anomalies due to the spatial and temporal sparsity of abnormal events, where the redundant information always leads to suboptimal outcomes. To address these challenges, exploiting the representation and generalization capabilities of Vison Language Models (VLMs) and Large Language Models (LLMs), we propose VA-GPT, a novel MLLM designed for summarizing and localizing abnormal events in various videos. Our approach efficiently aligns effective tokens between visual encoders and LLMs through two key proposed modules: Spatial Effective Token Selection (SETS) and Temporal Effective Token Generation (TETG). These modules enable our model to effectively capture and analyze both spatial and temporal information associated with abnormal events, resulting in more accurate responses and interactions. Furthermore, we construct an instruction-following dataset specifically for fine-tuning video-anomaly-aware MLLMs, and introduce a cross-domain evaluation benchmark based on XD-Violence dataset. Our proposed method outperforms existing state-of-the-art methods on various benchmarks. Yingxian Chen, Jiahui Liu 0012, Ruidi Fan, Chirui Chang, Shizhen Zhao, Wilton W. T. Fok, Xiaojuan Qi 0001, Yik-Chung Wu |
ICCV | 6 |
| 2025 | Equipping Vision Foundation Model with Mixture of Experts for Out-of-Distribution Detection
Shizhen Zhao, Jiahui Liu 0012, Xin Wen 0004, Haoru Tan, Xiaojuan Qi 0001 |
ICCV | 1 |
| 2025 | Data Pruning by Information MaximizationabstractIn this paper, we present InfoMax, a novel data pruning method, also known as coreset selection, designed to maximize the information content of selected samples while minimizing redundancy. By doing so, InfoMax enhances the overall informativeness of the coreset. The information of individual samples is measured by importance scores, which capture their influence or difficulty in model learning. To quantify redundancy, we use pairwise sample similarities, based on the premise that similar samples contribute similarly to the learning process.
We formalize the coreset selection problem as a discrete quadratic programming (DQP) task, with the objective of maximizing the total information content, represented as the sum of individual sample contributions minus the redundancies introduced by similar samples within the coreset.
To ensure practical scalability, we introduce an efficient gradient-based solver, complemented by sparsification techniques applied to the similarity matrix and dataset partitioning strategies.
This enables InfoMax to seamlessly scale to datasets with millions of samples.
Extensive experiments demonstrate the superior performance of InfoMax in various data pruning tasks, including image classification, vision-language pre-training, and instruction tuning for large language models. Haoru Tan, Sitong Wu, Wei Huang 0042, Shizhen Zhao, Xiaojuan Qi 0001 |
ICLR | 4 |
| 2025 | ONCache: A Cache-Based Low-Overhead Container Overlay Network
Shengkai Lin, Shizhen Zhao, Peirui Cao, Xinchi Han, Quan Tian, Donghai Han, Xinbing Wang |
NSDI | 2 |
| 2025 | VEP: A Two-stage Verification Toolchain for Full eBPF Programmability
Xiwei Wu, Yueyang Feng, Tianyi Huang, Xiaoyang Lu, Shengkai Lin, Lihan Xie, Shizhen Zhao, Qinxiang Cao |
NSDI | 7 |
| 2025 | Orderlock: A New Type of Deadlock and its Implications on High Performance Network Protocol DesignabstractIn the pursuit of designing high-performance network (HPN) protocols, three critical features for effective transmission have been extensively studied: In-order Delivery, Lossless Transmission, and Out-of-order Capability. However, no practical implementation has successfully achieved all three simultaneously. We identify and prove that the simultaneous realization of these features constitutes a necessary and sufficient condition for a new type of deadlock, which we term Orderlock. We demonstrate that operating in an Orderlock-risky network is impractical and conduct a comprehensive exploration and comparison of Orderlock-free protocols, through a case study tuning AI workload performance. From an Orderlock-prevention perspective, our findings provide insights into the requirements for future HPN protocol and hardware designs. Peirui Cao, Shizhen Zhao |
SIGCOMM | 5 |
| 2025 | vClos: Network contention aware scheduling for distributed machine learning tasks in multi-tenant GPU clusters
Xinchi Han, Shizhen Zhao, Yongxi Lv, Peirui Cao, Qinwei Yang, Yunzhuo Liu, Shengkai Lin, Bo Jiang 0003, Ximeng Liu, Yong Cui 0001, Chenghu Zhou, Xinbing Wang |
Comput. Networks | 2 |
| 2025 | Segment EDF: A Scheduling Policy With Tight Deterministic Latency Under Multi-Hop NetworksabstractWe investigate tight end-to-end delay guarantees for real-time flows with stochastic arrivals in multi-hop networks. The existing multi-hop scheduling policy either only serves for real-time flows with deterministic arrivals or cannot offer a tight end-to-end delay guarantee for multi-hop flows. In this paper, we prove a closed-form formula, which offers a sufficient and necessary condition for the EDF (Earliest Deadline First) scheduling policy to meet all the end-to-end deadlines for flows with stochastic arrivals in converge-cast tree networks. To the best of our knowledge, this is the first formula that characterizes the exact schedulability region for EDF in multi-hop networks. Moving beyond converge-cast tree networks to general multi-hop networks, we introduce the Segment EDF approach. This method partitions a general network into multiple converge-cast networks using a novel concept called the critical links. By determining segment deadlines within the schedulability region of each converge-cast tree, Segment EDF offers a tight end-to-end delay guarantee for each flow. We present a theoretical analysis showcasing the superior performance of Segment EDF over the existing scheduling policy of hop-by-hop EDF. Furthermore, we evaluate the performance of Segment EDF based on two key metrics: guaranteed flow completion time and admission ratio, in both real-world and synthetic networks. Our simulation results show that Segment EDF can provide 61.31%-96.4% tighter end-to-end delay guarantee and increase admission ratio by about 3.52%-107.59% for sequential arrival flow set and 1.04%-95.50% for batch arrival flow set than hop-by-hop EDF. Shizhen Zhao, Xinbing Wang, Chenghu Zhou |
IEEE Trans. Netw. | 2 |
| 2025 | Pulse+: DetNet Routing Under Delay-Diff ConstraintabstractDeterministic Networking (DetNet) is a rising technology that offers deterministic delay & jitter and extremely low packet loss in large IP networks. To achieve determinism under failure scenarios, DetNet requires finding at least two paths with close end-to-end delay, i.e., adelay-diffconstraint, for mission-critical flows. However, how to find two routing paths subject to thedelay-diffconstraint remains open. We study the DetNet routing problem in two scenarios. First, given a primary path, we propose Pulse+, which finds a secondary path whose end-to-end delay is within a range determined by the end-to-end delay of the primary path and the delay-diff requirement. Second, we propose CoSE-Pulse+, which integrates Pulse+ with a divide-and-conquer approach to find a pair of paths that meet DetNet’s delay-diff constraint. Both Pulse+ and CoSE-Pulse+ guarantee solution optimality. Notably, although Pulse+ and CoSE-Pulse+ do not have a polynomial worst-case time complexity, their empirical solver running time is better than that of other algorithms. We evaluate Pulse+ and CoSE-Pulse+ against the K-Shortest-Path and Lagrangian-dual based algorithms using synthetic test cases generated over networks with up to 10000 nodes. Both Pulse+ and CoSE-Pulse+ can solve more test cases than other algorithms under a predefined time limit. Compared to the second best algorithm, Pulse+ achieves an average-time speedup of$5\times $and CoSE-Pulse+ achieves an average-time speedup of$22\times $. Our code and test cases are available athttps://gitee.com/zsz2019_shizhenzhao/drcr Shizhen Zhao, Ximeng Liu, Xinbing Wang |
IEEE Trans. Netw. | 1 |
| 2024 | LubeRDMA: A Fail-safe Mechanism of RDMAabstractRecent years have witnessed a wide adoption of Remote Direct Memory Access (RDMA) to accelerate distributed systems. As the scale of distributed applications keeps increasing, network failures become more prominent. Although some link/switch failures can be circumvented by in-network rerouting, failures like NIC failure are still fatal in RDMA networks and may cause the entire system to fail. Shengkai Lin, Qinwei Yang, Zengyin Yang, Shizhen Zhao |
APNet | 5 |
| 2024 | Can OOD Object Detectors Learn from Foundation Models?
Jiahui Liu 0012, Xin Wen 0004, Shizhen Zhao, Yingxian Chen, Xiaojuan Qi 0001 |
ECCV (12) | 3 |
| 2024 | OFC: An Original Congestion-Based Fine-grained Priority Flow ControlabstractWith the proliferation of online data intensive applications and virtualized services, the growing complexity of traffic patterns in data centers increases the likelihood of congestion, especially in incast scenarios and with a combination of short and large flows. To ensure lossless transmission, RDMA over Converged Ethernet networks rely on Priority-based Flow Control (PFC) to prevent packet loss due to buffer overflow. However, it is widely acknowledged that PFC gives rise to several issues, such as Congestion Spreading, Head-of-Line Blocking, and Deadlock, which are increasingly prominent in modern highly congested data centers. In this paper, we analyze the primary causes of congestion spreading and head-of-line blocking issues associated with PFC and propose Original congestion-based fine-grained priority Flow Control (OFC) as a solution. The performance of OFC is assessed using the programmable switch Tofino and simulations carried out with a packet-level simulator across various scenarios, encompassing incast, realistic, deadlock, and in-depth scenarios. The validation of the simulation results through testbed evaluation confirmed that OFC effectively reduces flow completion time, buffer occupancy, and deadlock occurrence by up to 60.28%, 51.47%, and 48.7%, respectively. Peirui Cao, Shizhen Zhao, Xinbing Wang |
SECON | 5 |
| 2024 | NegotiaToR: Towards A Simple Yet Effective On-demand Reconfigurable Datacenter NetworkabstractRecent advances in fast optical switching technology show promise in meeting the high goodput and low latency requirements of datacenter networks (DCN). We present NegotiaToR, a simple network architecture for optical reconfigurable DCNs that utilizes on-demand scheduling to handle dynamic traffic. In NegotiaToR, racks exchange scheduling messages through an in-band control plane and distributedly calculate non-conflicting paths from binary traffic demand information. Optimized for incasts, it also provides opportunities to bypass scheduling delays. NegotiaToR is compatible with prevalent flat topologies, and is tailored towards a minimalist design for on-demand reconfigurable DCNs, enhancing practicality. Through large-scale simulations, we show that NegotiaToR achieves both small mice flow completion time (FCT) and high goodput on two representative flat topologies, especially under heavy loads. Particularly, the FCT of mice flows is one to two orders of magnitude better than the state-of-the-art traffic-oblivious reconfigurable DCN design. Cong Liang 0005, Xiangli Song, Mowei Wang, Yashe Liu, Zhenhua Liu 0008, Shizhen Zhao, Yong Cui 0001 |
SIGCOMM | 7 |
| 2024 | FIGRET: Fine-Grained Robustness-Enhanced Traffic EngineeringabstractTraffic Engineering (TE) is critical for improving network performance and reliability. A key challenge in TE is the management of sudden traffic bursts. Existing TE schemes either do not handle traffic bursts or uniformly guard against traffic bursts, thereby facing difficulties in achieving a balance between normal-case performance and burst-case performance. To address this issue, we introduce FIGRET, a Fine-Grained Robustness-Enhanced TE scheme. FIGRET offers a novel approach to TE by providing varying levels of robustness enhancements, customized according to the distinct traffic characteristics of various source-destination pairs. By leveraging a burst-aware loss function and deep learning techniques, FIGRET is capable of generating high-quality TE solutions efficiently. Our evaluations of real-world production networks, including Wide Area Networks and data centers, demonstrate that FIGRET significantly outperforms existing TE schemes. Compared to the TE scheme currently deployed in Google's Jupiter data center networks, FIGRET achieves a 9%-34% reduction in average Maximum Link Utilization and improves solution speed by 35×-1800×. Against DOTE, a state-of-the-art deep learning-based TE method, FIGRET substantially lowers the occurrence of significant congestion events triggered by traffic bursts by 41%-53.9% in topologies with high traffic dynamics. Ximeng Liu, Shizhen Zhao, Yong Cui 0001, Xinbing Wang |
SIGCOMM | 2 |
| 2024 | LPulse: An efficient algorithm for service function chain placement and routing with delay guarantee
Ximeng Liu, Shizhen Zhao, Xinbing Wang, Chenghu Zhou |
Comput. Networks | 2 |
| 2023 | Reducing Reconfiguration Time in Hybrid Optical-Electrical Datacenter NetworksabstractWe study how to reduce the reconfiguration time in hybrid optical-electrical Datacenter Networks (DCNs). With a layer of Optical Circuit Switches (OCSes), hybrid optical-electrical DCNs could reconfigure their logical topologies to better match the on-going traffic patterns, but the reconfiguration time could directly affect the benefits of reconfigurability. The reconfiguration time consists of the topology solver running time and the network convergence time after triggering reconfiguration. However, existing topology solvers either incur high algorithmic complexity or fail to minimize the reconfiguration overhead. Shu Shan, Shizhen Zhao |
APNet | 3 |
| 2023 | Accelerating QUIC with AF_XDP
Tianyi Huang, Shizhen Zhao |
ICA3PP (3) | 2 |
| 2023 | GRAP: Group-level Resource Allocation Policy for Reconfigurable Dragonfly Network in HPCabstractDragonfly is a highly scalable, low-diameter, and cost-efficient network topology, which has been adopted in new exascale High Performance Computing (HPC) systems. However, Dragonfly topology suffers from the limited direct links between groups. The reconfigurable network can solve this problem by reconfiguring topology to adjust the number of direct links between groups. While the performance improvement of a single job on reconfigurable HPC network has been evaluated in previous works, the performance of HPC workloads has not been studied because of the lack of an appropriate resource allocation policy. Guangnan Feng, Dezun Dong, Shizhen Zhao, Yutong Lu |
ICS | 3 |
| 2023 | Libra: Contention-Aware GPU Thread Allocation for Data Parallel Training in High Speed NetworksabstractOverlapping gradient communication with backward computation is a popular technique to reduce communication cost in the widely adopted data parallel S-SGD training. However, the resource contention between computation and All-Reduce communication in GPU-based training reduces the benefits of overlap. With GPU cluster network evolving from low bandwidth TCP to high speed networks, more GPU resources are required to efficiently utilize the bandwidth, making the contention more noticeable. Existing communication libraries fail to account for such contention when allocating GPU threads and have suboptimal performance. In this paper, we propose to mitigate the contention by balancing the overlapped computation and communication time. We formulate an optimization problem that decides the communication thread allocation to reduce overall backward time. We develop a dynamic programming based near-optimal solution and extend it to co-optimize thread allocation with tensor fusion. We conduct simulated study and real-world experiment using an 8-node GPU cluster with 50Gb RDMA network training four representative DNN models. Results show that our method reduces backward time by 10%-20% compared with Horovod-NCCL, by 6%-13% compared with tensor-fusion-optimization-only methods. Simulation shows that our method achieves the best scalability with a training speedup of 1.2x over the best-performing baseline as we scale up cluster size. Yunzhuo Liu, Bo Jiang 0003, Shizhen Zhao, Tao Lin 0001, Xinbing Wang, Chenghu Zhou |
INFOCOM | 3 |
| 2023 | Flattened Clos: Designing High-performance Deadlock-free Expander Data Center Networks Using Graph Contraction
Shizhen Zhao, Qizhou Zhang, Peirui Cao, Xinbing Wang, Chenghu Zhou |
NSDI | 1 |
| 2023 | Threshold-Based Routing-Topology Co-Design for Optical Data CenterabstractDespite the bandwidth scaling limit of electrical switching and the high cost of building Clos data center networks (DCNs), the adoption of optical DCNs is still limited. There are two reasons. First, existing optical DCN designs usually face high deployment complexity. Second, these designs are not full-optical and the performance benefit over the non-blocking Clos DCN is not clear. After exploring the design tradeoffs of the existing optical DCN designs, we propose TROD (ThresholdRouting basedOpticalDatacenter), a low-complexity optical DCN with superior performance than other optical DCNs. There are two novel designs in TROD that contribute to its success. First, TROD performs robust topology optimization based on the recurring traffic patterns and thus does not need to react to every traffic change, which lowers deployment and management complexity. Second, TROD introduces tVLB (threshold-based Valiant Load Balance), which can avoid network congestion as much as possible even under unexpected traffic bursts. We conduct simulation based on both Facebook’s real DCN traces and our synthesized highly bursty DCN traces. TROD reduces flow completion time (FCT) by about 1.15-2.16$\times$compared to Google’s Jupiter DCN, at least 2$\times$compared to other optical DCN designs, and about 2.4-3.2$\times$compared to expander graph DCN. Compared with the non-blocking Clos, TROD reduces the hop count of the majority packets by one, and could even outperform the non-blocking Clos with proper bandwidth over-provision at the optical layer. Note that TROD can be built with commercially available hardware and does not require host modifications. Peirui Cao, Shizhen Zhao, Zhuotao Liu, Mingwei Xu 0001, Min Yee Teh, Yunzhuo Liu, Xinbing Wang, Chenghu Zhou |
IEEE/ACM Trans. Netw. | 2 |
| 2023 | Enabling Quasi-Static Reconfigurable Networks With Robust Topology EngineeringabstractMany optical circuit switched data center networks (DCN) have been proposed in the last decade to attain higher capacity and topology reconfigurability, though commercial adoption of these architectures have been minimal. One major challenge these architectures face is the difficulty of handling uncertain traffic demands using commercial optical circuit switches (OCS) with high switching latency. Prior works have generally focused on developing fast-switching OCS prototypes to quickly react to traffic variations through frequent reconfigurations. This approach, however, adds tremendous complexity overhead to the control plane, and raises the barrier for commercial adoption of optical circuit switched data center networks. We propose, a robust topology and routing optimization framework for reconfigurable optical circuit switched data centers. co-optimizes topology and routing based on a convex set of traffic matrices, and offers strict throughput guarantees for any future traffic matrices bounded by the convex set. For the bursty traffic demands that are unbounded by the convex set, we employ a desensitization technique to reduce performance hit. This enables to generate topology and routing solutions capable of handling unexpected traffic changes without relying on frequent topology reconfigurations. Our extensive evaluations based on Facebook’s production DCN traces show that, even with daily reconfigurations which could be realized by current commercial MEMS-based OCSs from Calient Technologies, achieves about 20% lower max link utilization, and about 32% lower average hop count compared to cost-equivalent static topologies. Our work shows that adoption of reconfigurable topologies in commercial DCNs is feasible even without fast OCSs. Min Yee Teh, Shizhen Zhao, Peirui Cao, Keren Bergman |
IEEE/ACM Trans. Netw. | 2 |
| 2022 | PDE-based consensus control for leader-follower multi-agent systemsabstractIn this paper, we propose a control for a class of multi-agent systems described by diffusion partial differential equations to solve the leader-follower consensus problem. Each group of follower agents changes according to the leader group's formation, obtaining the desired formation by applying boundary control. We use the Lyapunov direct method to prove the system is uniformly ultimately bounded. Numerical simulation results finally show the validity of the proposed control. Xiaofeng Cui, Yankun He, Zhijie Liu 0001, Shizhen Zhao |
ICARCV | 5 |
| 2022 | Prototypical VoteNet for Few-Shot 3D Point Cloud Object DetectionabstractMost existing 3D point cloud object detection approaches heavily rely on large amounts of labeled training data. However, the labeling process is costly and time-consuming. This paper considers few-shot 3D point cloud object detection, where only a few annotated samples of novel classes are needed with abundant samples of base classes. To this end, we propose Prototypical VoteNet to recognize and localize novel instances, which incorporates two new modules: Prototypical Vote Module (PVM) and Prototypical Head Module (PHM). Specifically, as the 3D basic geometric structures can be shared among categories, PVM is designed to leverage class-agnostic geometric prototypes, which are learned from base classes, to refine local features of novel categories. Then PHM is proposed to utilize class prototypes to enhance the global feature of each object, facilitating subsequent object localization and classification, which is trained by the episodic training strategy. To evaluate the model in this new setting, we contribute two new benchmark datasets, FS-ScanNet and FS-SUNRGBD. We conduct extensive experiments to demonstrate the effectiveness of Prototypical VoteNet, and our proposed method shows significant and consistent improvements compared to baselines on two benchmark datasets. Shizhen Zhao, Xiaojuan Qi 0001 |
NeurIPS | 1 |
| 2021 | Weakly Supervised Text-based Person Re-IdentificationabstractThe conventional text-based person re-identification methods heavily rely on identity annotations. However, this labeling process is costly and time-consuming. In this paper, we consider a more practical setting called weakly supervised text-based person re-identification, where only the text-image pairs are available without the requirement of annotating identities during the training phase. To this end, we propose a Cross-Modal Mutual Training (CMMT) framework. Specifically, to alleviate the intra-class variations, a clustering method is utilized to generate pseudo labels for both visual and textual instances. To further re-fine the clustering results, CMMT provides a Mutual Pseudo Label Refinement module, which leverages the clustering results in one modality to refine that in the other modality constrained by the text-image pairwise relationship. Mean-while, CMMT introduces a Text-IoU Guided Cross-Modal Projection Matching loss to resolve the cross-modal matching ambiguity problem. A Text-IoU Guided Hard Sample Mining method is also proposed for learning discriminative textual-visual joint embeddings. We conduct extensive experiments to demonstrate the effectiveness of the proposed CMMT, and the results show that CMMT performs favorably against existing text-based person re-identification methods. Our code will be available at https://github.com/X-BrainLab/WS_Text-ReID. Shizhen Zhao, Changxin Gao, Yuanjie Shao, Wei-Shi Zheng 0001, Nong Sang |
ICCV | 1 |
| 2021 | TROD: Evolving From Electrical Data Center to Optical Data CenterabstractDespite the bandwidth scaling limit of electrical switching and the high cost of building Clos data center networks (DCNs), the adoption of optical DCNs is still limited. There are two reasons. First, existing optical DCN designs usually face tremendous deployment complexity. Second, these designs are not full-optical and the performance benefit against the non-blocking Clos DCN is not clear.After exploring the design tradeoffs of the existing optical DCN designs, we propose TROD (Threshold Routing based Optical Datacenter), a low-complexity optical DCN with superior performance than other optical DCNs. There are two novel designs in TROD that contribute to its success. First, TROD performs robust topology optimization based on the recurring traffic patterns and thus does not need to react to every traffic change, which lowers deployment and management complexity. Second, TROD introduces tVLB (threshold-based VLB), which can avoid network congestion as much as possible even under unexpected traffic bursts. We conduct simulation based on both Facebook’s real DCN traces and our synthesized highly bursty DCN traces. TROD reduces flow completion time (FCT) by at least 2× compared with the existing optical DCN designs, and by approximately 2.4-3.2× compared with expander graph DCN. Compared with the non-blocking Clos, TROD reduces the hop count of the majority packets by one, and could even outperform the non-blocking Clos with proper bandwidth over-provision at the optical layer. Note that TROD can be built with commercially available hardware and does not require host modifications. Peirui Cao, Shizhen Zhao, Min Yee Teh, Yunzhuo Liu, Xinbing Wang |
ICNP | 2 |
| 2021 | Design of Robust and Efficient Edge Server Placement and Server Scheduling PoliciesabstractWe study how to design edge server placement and server scheduling policies under workload uncertainty for 5G networks. We introduce a new metric called resource pooling factor to handle unexpected workload bursts. Maximizing this metric offers a strong enhancement on top of robust optimization against workload uncertainty. Using both real traces and synthetic traces, we show that the proposed server placement and server scheduling policies not only demonstrate better robustness against workload uncertainty than existing approaches, but also significantly reduce the cost of service providers. Specifically, in order to achieve close-to-zero workload rejection rate, the proposed server placement policy reduces the number of required edge servers by about 25% compared with the state-of-the-art approach; the proposed server scheduling policy reduces the energy consumption of edge servers by about 13% without causing much impact on the service quality. Shizhen Zhao, Peirui Cao, Xinbing Wang |
IWQoS | 1 |
| 2020 | GTNet: Generative Transfer Network for Zero-Shot Object DetectionabstractWe propose a Generative Transfer Network (GTNet) for zero-shot object detection (ZSD). GTNet consists of an Object Detection Module and a Knowledge Transfer Module. The Object Detection Module can learn large-scale seen domain knowledge. The Knowledge Transfer Module leverages a feature synthesizer to generate unseen class features, which are applied to train a new classification layer for the Object Detection Module. In order to synthesize features for each unseen class with both the intra-class variance and the IoU variance, we design an IoU-Aware Generative Adversarial Network (IoUGAN) as the feature synthesizer, which can be easily integrated into GTNet. Specifically, IoUGAN consists of three unit models: Class Feature Generating Unit (CFU), Foreground Feature Generating Unit (FFU), and Background Feature Generating Unit (BFU). CFU generates unseen features with the intra-class variance conditioned on the class semantic embeddings. FFU and BFU add the IoU variance to the results of CFU, yielding class-specific foreground and background features, respectively. We evaluate our method on three public datasets and the results demonstrate that our method performs favorably against the state-of-the-art ZSD approaches. Shizhen Zhao, Changxin Gao, Yuanjie Shao, Lerenhan Li, Changqian Yu, Zhong Ji, Nong Sang |
AAAI | 1 |
| 2020 | Do Not Disturb Me: Person Re-identification Under the Interference of Other Pedestrians
Shizhen Zhao, Changxin Gao, Jun Zhang 0018, Hao Cheng 0012, Chuchu Han, Xinyang Jiang, Wei-Shi Zheng 0001, Nong Sang, Xing Sun 0001 |
ECCV (6) | 1 |
| 2019 | Minimal Rewiring: Efficient Live Expansion for Clos Data Center Networks
Shizhen Zhao, Rui Wang 0025, Junlan Zhou, Joon Ong, Jeffrey C. Mogul, Amin Vahdat |
NSDI | 1 |
| 2017 | Timely Wireless Flows With General Traffic Patterns: Capacity Region and Scheduling AlgorithmsabstractMost existing wireless networking solutions are best-effort and do not provide any delay guarantee required by important applications, such as mobile multimedia conferencing and real-time control of cyber-physical systems. Recently, Hou and Kumar provided a novel framework for analyzing and designing delay-guaranteed wireless networking solutions. While inspiring, their idle-time-based analysis applies only to flows with a special traffic pattern called the frame-synchronized setting. The problem remains largely open for general traffic patterns. This paper addresses this challenge by proposing a general framework that characterizes and achieves the complete delay-constrained capacity region with general traffic patterns in single-hop downlink access-point wireless networks. We first show that the timely wireless flow problem is fundamentally an infinite-horizon Markov decision process (MDP). Then, we judiciously combine different simplification methods to prove that the timely capacity region can be characterized by a finite-size convex polygon. This for the first time allows us to characterize the timely capacity region of wireless flows with general traffic patterns. We then design three scheduling policies to optimize network utility and/or support feasible timely throughput vectors for general traffic patterns. The first policy achieves the optimal network utility and supports any feasible timely throughput vector but suffers from the curse of dimensionality. The second and third policies are inspired by our MDP framework and are of much lower complexity. Simulation results show that both achieve near-optimal performance and outperform other existing alternatives. Lei Deng 0001, Chih-Chun Wang, Minghua Chen 0001, Shizhen Zhao |
IEEE/ACM Trans. Netw. | 4 |
| 2016 | Timely wireless flows with arbitrary traffic patterns: Capacity region and scheduling algorithmsabstractMost existing wireless networking solutions are best-effort and do not provide any delay guarantee required by important applications such as the control traffic of cyber-physical systems. Recently, Hou and Kumar provided the first framework for analyzing and designing delay-guaranteed network solutions. While inspiring, their idle-time-based analysis appears to apply only to flows with a special traffic (arrival and expiration) pattern, and the problem remains largely open for general traffic patterns. This paper addresses this challenge by proposing a new framework that characterizes and achieves the complete delay-constrained capacity region with general traffic patterns in single-hop downlink access-point wireless networks. We first formulate the timely capacity problem as an infinite-horizon Markov Decision Process (MDP) and then judiciously combine different simplification methods to convert it to an equivalent finite-size linear program (LP). This allows us to characterize the timely capacity region of flows with general traffic patterns for the first time in the literature. We then design three timely-flow scheduling algorithms for general traffic patterns. The first algorithm achieves the optimal utility but suffers from the curse of dimensionality. The second and third algorithms are inspired by our MDP framework and are of polynomial-time complexity. Simulation results show that both achieve near-optimal performance and outperform other existing alternatives. Lei Deng 0001, Chih-Chun Wang, Minghua Chen 0001, Shizhen Zhao |
INFOCOM | 4 |
| 2016 | Online multi-stage decisions for robust power-grid operations under high renewable uncertaintyabstractIn this paper, we are interested in online multistage decisions to ensure robust power grid operations under high renewable uncertainty. We jointly consider both the reliability assessment commitment (RAC) and the real-time dispatch problems. We first focus on the real-time dispatch problem and define “maximally robust algorithms,” which can provably ensure grid safety whenever there exists any other algorithm that can ensure grid safety under the same level of future uncertainty. We characterize a class of maximally robust algorithms using the concept of “safe dispatch set,” which also provides conditions for verifying grid safety for RAC. However, in general such safe dispatch sets may be difficult to compute. We then develop efficient computational algorithms for characterizing the safe dispatch sets. Specifically, for a simpler single-bus two-generator case, we show that the safe dispatch sets can be exactly characterized by a polynomial number of convex constraints. Then, based on this two-generator characterization, we develop a new solution for the multi-bus multi-generator case using the idea of virtual demand splitting (VDS), which can effectively compute a suitable subset of the safe-dispatch set. Our numerical results demonstrate that a VDS-based economic dispatch algorithm outperforms the standard economic dispatch algorithm in terms of robustness, without sacrificing economy. Shizhen Zhao, Xiaojun Lin 0001, Dionysios Aliprantis, Hugo N. Villegas, Minghua Chen 0001 |
INFOCOM | 1 |
| 2016 | Design of Scheduling Algorithms for End-to-End Backlog Minimization in Wireless Multi-Hop Networks Under K-Hop Interference ModelsabstractIn this paper, we study the problem of link scheduling for multi-hop wireless networks with per-flow delay constraints under the$K$-hop interference model. Specifically, we are interested in algorithms that maximize the asymptotic decay-rate of the probability with which the maximum end-to-end backlog among all flows exceeds a threshold, as the threshold becomes large. We provide both positive and negative results in this direction. By minimizing the drift of the maximum end-to-end backlog in the converge-cast on a tree, we design an algorithm, Largest-Weight-First (LWF), that achieves the optimal asymptotic decay-rate for the overflow probability of the maximum end-to-end backlog as the threshold becomes large. However, such a drift minimization algorithm may not exist for general networks. We provide an example in which no algorithm can minimize the drift of the maximum end-to-end backlog. Finally, we simulate the LWF algorithm together with a well known algorithm (the back-pressure algorithm) and a large-deviations optimal algorithm in terms of the sum-queue (the P-TREE algorithm) in converge-cast networks. Our simulation shows that our algorithm performs significantly better not only in terms of asymptotic decay-rate, but also in terms of the actual overflow probability. Shizhen Zhao, Xiaojun Lin 0001 |
IEEE/ACM Trans. Netw. | 1 |
| 2015 | Peak-minimizing online EV charging: Price-of-uncertainty and algorithm robustificationabstractWe study competitive online algorithms for EV (electrical vehicle) charging under the scenario of an aggregator serving a large number of EVs together with its background load, using both its own renewable energy (for free) and the energy procured from the external grid. The goal of the aggregator is to minimize its peak procurement from the grid, subject to the constraint that each EV has to be fully charged before its deadline. Further, the aggregator can predict the future demand and the renewable energy supply with some levels of uncertainty. The key challenge here is how to develop a model that captures the prior knowledge from such prediction, and how to best utilize this prior knowledge to reduce the peak under future uncertainty. In this paper, we first propose a 2-level increasing precision model (2-IPM), to capture the system uncertainty. We develop a powerful computation approach that can compute the optimal competitive ratio under 2-IPM over any online algorithm, and also online algorithms that can achieve the optimal competitive ratio. A dilemma for online algorithm design is that an online algorithm with good competitive ratio may exhibit poor average-case performance. We then propose a new Algorithm-Robustification procedure that can convert an online algorithm with reasonable average-case performance to one with both the optimal competitive ratio and good average-case performance. The robustified version of a well-known heuristic algorithm, Receding Horizon Control (RHC), is found to demonstrate superior performance via trace-based simulations. Shizhen Zhao, Xiaojun Lin 0001, Minghua Chen 0001 |
INFOCOM | 1 |
| 2014 | Rate-control and multi-channel scheduling for wireless live streaming with stringent deadlinesabstractSVC-based live video-streaming in multi-channel wireless networks leads to a challenging joint rate-control and scheduling problem with stringent deadline constraints. Traditional utility-based approaches often did not explicitly account for deadlines. In this paper, we explicitly account for deadlines and study the problem of optimizing the total reward from packets meeting their deadlines in a modern 4G OFDM system. Motivated by a heuristic utility-based approach, we propose a class of threshold-based rate-control and wireless scheduling policies that can respect the deadline constraints and approach the optimal system reward asymptotically as the system size increases. We also propose a distributed realization of our threshold-based policies that can be easily implemented in practical scenarios. We substantiate the result via both analysis and simulation. Shizhen Zhao, Xiaojun Lin 0001 |
INFOCOM | 1 |
| 2014 | Node Density and Delay in Large-Scale Wireless Networks With Unreliable LinksabstractWe study the delay performance in large-scale wireless multihop networks with unreliable links from percolation perspective. Previous works have showed that the end-to-end delay scales linearly with the source-to-destination distance, and thus the delay performance can be characterized by the delay-distance ratio γ. However, the range of γ, which may be the most important parameter for delay, remains unknown. We expect that γ may depend heavily on the node density λ of a wireless multihop network. In this paper, we investigate the fundamental relationship between γ and λ. Obtaining the exact value of γ(λ) is extremely hard, mainly because of the dynamically changing network topologies caused by the link unreliability. Instead, we provide both upper bound and lower bound to the delay-distance ratio γ(λ). Simulations are conducted to verify our theoretical analysis. Shizhen Zhao, Xinbing Wang |
IEEE/ACM Trans. Netw. | 1 |
| 2012 | On the design of scheduling algorithms for end-to-end backlog minimization in multi-hop wireless networksabstractIn this paper, we study the problem of link scheduling for multi-hop wireless networks with per-flow delay constraints. Specifically, we are interested in algorithms that maximize the asymptotic decay-rate of the probability with which the maximum end-to-end backlog among all flows exceeds a threshold, as the threshold becomes large. We provide both positive and negative results in this direction. By minimizing the drift of the maximum end-to-end backlog in the converge-cast on a tree, we design an algorithm, Largest-Weight-First(LWF), that achieves the optimal asymptotic decay-rate for the overflow probability of the maximum end-to-end backlog as the threshold becomes large. However, such a drift minimization algorithm may not exist for general networks. We provide an example in which no algorithm can minimize the drift of the maximum end-to-end backlog. Finally, we simulate the LWF algorithm together with a well known algorithm (the back-pressure algorithm) and a large-deviations optimal algorithm in terms of the sum-queue (the P-TREE algorithm) in converge-cast networks. Our simulation shows that our algorithm significantly performs better not only in terms of asymptotic decay-rate, but also in terms of the actual overflow probability. Shizhen Zhao, Xiaojun Lin 0001 |
INFOCOM | 1 |
| 2011 | Fundamental relationship between NodeDensity and delay in wireless ad hoc networks with unreliable linksabstractWe investigate the fundamental relationship between node density and transmission delay in large-scale wireless ad hoc networks with unreliable links from percolation perspective. Previous works[11][2][10] have already showed the relationship between transmission delay and distance from source to destination. However, it still remains as an open question how transmission delay varies in accordance with node density. Answering this question can provide guidance for determining the number of nodes to meet the delay requirement when designing ad hoc networks. In this paper, we study the impact of node density λ on the ratio of delay and distance, denoted by γ(λ). We analytically characterize the properties of γ(λ) as a function of λ. And then we present upper and lower bounds to γ(λ). Next, we take propagation delay into consideration and obtain further results on the upper and lower bounds of γ(λ). Finally, we make simulations to verify our theoretical analysis. Shizhen Zhao, Luoyi Fu, Xinbing Wang, Qian Zhang 0001 |
MobiCom | 1 |