VLDB 2026 Research / reviewers in the wild / expert
Mowei Wang
dblp:206/6742
· DBLP profile ↗
19ranked-venue papers
2as first author
19since 2021 · last 2026
0000-0001-9085-2247ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 18 · 2 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | NegotiaToR: Toward a Simple Yet Effective On-Demand Reconfigurable Datacenter NetworkabstractRecent advances in fast optical switching show promise in meeting the high goodput and low latency requirements of datacenter networks. We present NegotiaToR, a simple network architecture for optical reconfigurable DCNs that utilizes on-demand scheduling to handle dynamic traffic. In NegotiaToR, racks exchange scheduling messages through an in-band control plane and distributedly calculate non-conflicting paths from binary traffic demand information. Optimized for incasts, it also provides opportunities to bypass scheduling delays. NegotiaToR is compatible with prevalent flat topologies, and is tailored towards a minimalist design for on-demand reconfigurable DCNs, enhancing practicality. Through large-scale simulations, we show that NegotiaToR achieves both small mice flow completion time and high goodput on two representative flat topologies, especially under heavy loads. Particularly, the flow completion time of mice flows is one to two orders of magnitude better than the state-of-the-art traffic-oblivious reconfigurable DCN design. Cong Liang 0005, Xiangli Song, Mowei Wang, Yashe Liu, Zhenhua Liu 0008, Shizhen Zhao, Yong Cui 0001 |
IEEE Trans. Netw. | 5 |
| 2025 | Modeling Flow-level Traffic Demand for Network Performance Evaluation and OptimizationabstractModeling traffic demand at the flow level is essential for accurate network performance evaluation and optimization. However, despite its prevalence, the common practice is oversimplified and relies on unsubstantiated assumptions of traffic homogeneity and independent arrivals. In this paper, we analyze real-world traffic data collected from production environments to challenge these assumptions. Our findings reveal notable fidelity issues in the common practice, compromising the reliability of network performance evaluation and optimization. To address these limitations, we introduce Encore, a flow-level traffic demand modeling framework that captures key traffic characteristics and generates high-fidelity synthetic traces. Encore adopts a divide-and-conquer strategy, employing tailored machine learning models for distributional and sequential modeling, along with problem-specific enhancements. Systematic evaluations demonstrate that Encore outperforms existing traffic modeling methods in terms of accuracy and coverage in distribution modeling, and fidelity in sequential modeling. In addition to accurately restoring key characteristics of real traffic, Encore improves simulation performance consistency by a factor of 4 to 17 over the common practice. Moreover, Encore achieves a ~0.88 correlation in parameter ranking compared to the ground truth, showcasing its practical utility for network optimization. Sijiang Huang, Xiaohui Xie, Mowei Wang, Lingfeng Peng, Yong Zhang 0062, Yingjie Qin, Yong Cui 0001 |
ICNP | 3 |
| 2025 | xWitch: Towards Fast and Accurate Performance Evaluation for Hierarchical QoSabstractHierarchical Quality of Service (HQoS) has been designed to meet the diverse needs of different users and applications, and it is widely applied in commercial routers. However, the large number of users and applications results in numerous HQoS queue parameters that need to be configured. Fast and accurate performance evaluation of these configurations is crucial. Traditional discrete event network simulators experience significant slowdowns as the traffic increases. Existing machine learning-based methods for performance evaluation also face challenges, such as low accuracy and poor generalization, due to the long input sequences caused by large network traffic.In this paper, we propose xWitch, a packet-level performance evaluation scheme that supports HQoS. xWitch uses a sequence-to-sequence model to achieve short sequence performance prediction. A long sequence parallel prediction scheme based on dependency prediction is proposed to support fast and accurate prediction of long sequences. Experimental results show that xWitch outperforms all baselines. It achieves an average error of under 3% in latency prediction for short sequences. For long sequence prediction, it can reduce the error by more than 20% and consistently maintains latency prediction errors below 10% across different traffic distributions. Gang Yi, Mowei Wang, Chuxuan Zeng, Feipeng Li, Yong Cui 0001 |
IWQoS | 2 |
| 2025 | PRED: Performance-oriented Random Early Detection for Consistently Stable Performance in Datacenters
Xinle Du, Tong Li 0014, Guangmeng Zhou, Zhuotao Liu, Hanlin Huang, Mowei Wang, Kun Tan 0002, Ke Xu 0002 |
NSDI | 7 |
| 2025 | Revisiting Random Early Detection Tuning for High-Performance Datacenter NetworksabstractRandom Early Detection (RED) has been integrated into datacenter switches as a fundamental Active Queue Management (AQM) for decades. The accurate configuration of RED parameters is crucial to achieving high throughput and low latency. However, due to the highly dynamic nature of workloads in datacenter networks, maintaining consistently high performance with statically configured RED thresholds poses a challenge. Prior work applies reinforcement learning to predict proper thresholds, but their real-world deployment has been hindered by poor tail performance caused by instability. In this paper, we propose$\textsf {PRED}$, a novel system that enables automatic and stable RED parameter adjustment in response to traffic dynamics. Specifically, the system employs a Multiplicative-Increase Multiplicative-Decrease (MIMD) strategy to dynamically adapt to flow concurrency while utilizing an Additive-Increase Additive-Decrease (AIAD) mechanism to adapt to flow distribution. We perform extensive evaluations on our physical testbed and large-scale simulations. The results demonstrate that$\textsf {PRED}$can keep up with the real-time network dynamics generated by realistic workloads. For instance, compared with the static-threshold-based methods,$\textsf {PRED}$keeps 66% shorter switch queue length and obtains up to 80% lower Flow Completion Time (FCT). Compared with the state-of-the-art learning-based method,$\textsf {PRED}$reduces the tail FCT by 34%. Tong Li 0014, Xinle Du, Guangmeng Zhou, Hanlin Huang, Zhuotao Liu, Mowei Wang, Kun Tan 0002, Ke Xu 0002 |
IEEE Trans. Netw. | 7 |
| 2024 | Iphicles: Tuning Parameters of Data Center Networks with Differentiable Performance ModelabstractTuning parameters in Data Center Networks (DCN) has long been a nuisance and one of the reasons service providers are reluctant to deploy new mechanisms in their production environments. Despite the excessive time and resources devoted to finding better configurations, a "one-size-fits-all" solution remains elusive. Neither manual configuration by experts nor black-box optimization can address the challenges of network heterogeneity and dynamics. One essential factor impeding efficient and stable parameter optimization is the need to explore in real environments, which has a long convergence time alongside the risk of performance degradation. To address this problem, we build a twin performance model of the physical DCN that approximates the mapping from parameters to Quality of Service (QoS) metrics for fast and safe performance inference and present a DCN configuration framework called Iphicles. Leveraging gradients provided by differentiable performance models built with Graph Neural Networks (GNN), Iphicles can automatically recommend better parameters efficiently and stably. Experimental results based on extensive simulation demonstrate that in complex scenarios with mixed and dynamic traffic, Iphicles can deliver parameters that lead to evident improvements in flow completion time (FCT) for both mice and elephant flows simultaneously, with minimum convergence time while maintaining performance stability during the optimization process. Sijiang Huang, Mowei Wang, Yashe Liu, Zhenhua Liu 0008, Yong Cui 0001 |
IWQoS | 2 |
| 2024 | NegotiaToR: Towards A Simple Yet Effective On-demand Reconfigurable Datacenter NetworkabstractRecent advances in fast optical switching technology show promise in meeting the high goodput and low latency requirements of datacenter networks (DCN). We present NegotiaToR, a simple network architecture for optical reconfigurable DCNs that utilizes on-demand scheduling to handle dynamic traffic. In NegotiaToR, racks exchange scheduling messages through an in-band control plane and distributedly calculate non-conflicting paths from binary traffic demand information. Optimized for incasts, it also provides opportunities to bypass scheduling delays. NegotiaToR is compatible with prevalent flat topologies, and is tailored towards a minimalist design for on-demand reconfigurable DCNs, enhancing practicality. Through large-scale simulations, we show that NegotiaToR achieves both small mice flow completion time (FCT) and high goodput on two representative flat topologies, especially under heavy loads. Particularly, the FCT of mice flows is one to two orders of magnitude better than the state-of-the-art traffic-oblivious reconfigurable DCN design. Cong Liang 0005, Xiangli Song, Mowei Wang, Yashe Liu, Zhenhua Liu 0008, Shizhen Zhao, Yong Cui 0001 |
SIGCOMM | 4 |
| 2024 | Re-Architecting Buffer Management in Lossless EthernetabstractConverged Ethernet employs Priority-based Flow Control (PFC) to provide a lossless network. However, issues caused by PFC, including victim flow, congestion spreading, and deadlock, impede its large-scale deployment in production systems. The fine-grained experimental observations on switch buffer occupancy find that the root cause of these performance problems is a mismatch of sending rates between end-to-end congestion control and hop-by-hop flow control. Resolving this mismatch requires the switch to provide an additional buffer, which is not supported by the classic dynamic threshold (DT) policy in current shared-buffer commercial switches. In this paper, we propose Selective-PFC (SPFC), a practical buffer management scheme that handles such mismatch. Specifically, SPFC incrementally modifies DT by proactively detecting port traffic and adjusting buffer allocation accordingly to trigger PFC PAUSE frames selectively. Extensive case studies demonstrate that SPFC can reduce the number of PFC PAUSEs on non-bursty ports by up to 69.0%, and reduce the average flow completion time by up to 83.5% for large victim flows. Hanlin Huang, Xinle Du, Tong Li 0014, Ke Xu 0002, Mowei Wang, Huichen Dai |
IEEE/ACM Trans. Netw. | 6 |
| 2024 | xNet: Modeling Network Performance With Graph Neural NetworksabstractToday’s network is notorious for its complexity and uncertainty. Network operators often rely on network models for efficient network planning, operation, and optimization. The network model is responsible for understanding the complex relationships between network performance metrics (e.g., delay and jitter) and network characteristics (e.g., traffic and configuration). However, we still lack a systematic approach to developing accurate and lightweight network models that are aware of the impact of network configurations (i.e., expressiveness) and provide fine-grained flow-level temporal predictions (i.e., granularity). In this paper, we propose xNet, a data-driven network modeling framework based on graph neural networks (GNN). It is worth noting that xNet is not a dedicated network model designed for a specific network scenario with constraint considerations. On the contrary, xNet provides a general approach to modeling the network characteristics of concern with relation graph representations and configurable GNN blocks. xNet learns the state transition functions between time steps and rolls them out to obtain the full fine-grained prediction trajectory. We implement and instantiate xNet with three use cases. The experimental results show that xNet can accurately predict different performance metrics (i.e. temporal and steady-state QoS) in different scenarios, with performance comparable to state-of-the-art domain-specific models. Compared with traditional packet-level simulators, xNet achieves a speed improvement of more than two orders of magnitude, demonstrating its promising application in real-time optimization of network configurations. Sijiang Huang, Yunze Wei, Lingfeng Peng, Mowei Wang, Linbo Hui, Zongpeng Du, Zhenhua Liu 0008, Yong Cui 0001 |
IEEE/ACM Trans. Netw. | 4 |
| 2023 | Datacenter Network Deserves Better Traffic ModelsabstractTraffic modeling of Datacenter Network (DCN) today is over-simplified, deviating from the ground truth. Adopted by numerous researchers, the common practice relies on the assumptions of traffic homogeneity and independence for ease of use. Based on our investigation of a real-world traffic dataset, we disprove these assumptions and point out the severe fidelity issue of the common practice that could invalidate many motivations and conclusions from influential research works. In this paper, we present Encore, a DCN traic modeling framework for ine-grained traic modeling and high-fidelity synthetic traffic generation. Leveraging machine learning techniques, Encore effectively extracts and preserves essential distribution and sequential features from raw traic. Preliminary experiments demonstrate that the traic generated by Encore not only restores the key features of real traffic but also achieves high consistency when used to evaluate network performance. We envision further expanding Encore to full-process traffic modeling and generation, and expect these critical improvements in traffic models can facilitate the DCN performance evaluation and optimization. Sijiang Huang, Lingfeng Peng, Mowei Wang, Yashe Liu, Zhenhua Liu 0008, Xin Wang 0001, Yong Cui 0001 |
HotNets | 3 |
| 2023 | Poster: NegotiaToR: A Simple On-Demand Reconfigurable Data Center NetworkabstractOptical switching technology has developed fast in recent years, which has the potential to provide high good put as well as low latency in reconfigurable data center networks (RDCN). However, existing optical RDCN proposals fail to balance good put, latency, and design complexity well. In this paper, we introduce NegotiaToR, a simple optical RDCN architecture. NegotiaToR utilizes on-demand scheduling to ensure high performance and reduces possible deployment complexity with an in-band control protocol to do the schedule distributedly. Our preliminary evaluation shows that NegotiaToR outperforms the state-of-the-art optical RDCN proposal under similar complexity in both goodput and flow completion time (FCT). Cong Liang 0005, Xiangli Song, Mowei Wang, Yashe Liu, Zhenhua Liu 0008, Yong Cui 0001 |
ICNP | 4 |
| 2022 | Locality Matters! Traffic Demand Modeling in Datacenter NetworksabstractUnderstanding and modeling traffic demand characteristics in datacenter networks is of great importance for datacenter network optimization. However, prior traffic models are over-simplified and insufficient in capturing the complex locality properties of traffic demand. We analyze real-world traffic traces and discover strong dependency between the spatial attributes (source, destination) and non-spatial attributes (interarrival time, flow size) of traffic demand. We propose Lomas to model the joint distribution of multi-dimensional traffic demand attributes and generate synthetic traces. Lomas is a novel extension of hierarchical Bayes model that can represent the relationships among these attributes as a dependency graph. We validate Lomas by showing its ability to recreate the flow-level traffic demand patterns of real-world traffic traces. Our approach can be easily adapted to different datacenters with heterogeneous traffic demand patterns, making it a convenient tool for practitioners to utilize. Mowei Wang, Yong Cui 0001 |
APNet | 2 |
| 2022 | xNet: Improving Expressiveness and Granularity for Network Modeling with Graph Neural NetworksabstractToday’s network is notorious for its complexity and uncertainty. Network operators often rely on network models to achieve efficient network planning, operation, and optimization. The network model is responsible for understanding the complex relationships between the network performance metrics (e.g., latency) and the network characteristics (e.g., traffic). However, we still lack a systematic approach to developing accurate and lightweight network models that are aware of the impact of network configurations (i.e., expressiveness) and provide fine-grained flow-level temporal predictions (i.e., granularity).In this paper, we propose xNet, a data-driven network modeling framework based on graph neural networks (GNN). Unlike the previous proposals, xNet is not a dedicated network model designed for specific network scenarios with constraint considerations. On the contrary, xNet provides a general approach to modeling the network characteristics of concern with relation graph representations and configurable GNN blocks. xNet learns the state transition function between time steps and rolls it out to obtain the full fine-grained prediction trajectory. We implement and instantiate xNet with three use cases. The experiment results show that xNet can accurately predict different performance metrics while achieving over two orders of magnitude of speedup compared with the conventional packet-level simulator. Mowei Wang, Linbo Hui, Yong Cui 0001, Ru Liang, Zhenhua Li 0001 |
INFOCOM | 1 |
| 2022 | Learning Buffer Management Policies for Shared Memory SwitchesabstractToday’s network switches often use on-chip shared memory to improve buffer efficiency and absorb bursty traffic. Current buffer management practices usually rely on simple heuristics and have unrealistic assumptions about the traffic pattern, since developing a buffer management policy suited for every scenario is infeasible. We show that modern machine learning techniques can be of essential help to learn efficient policies automatically.In this paper, we propose Neural Dynamic Threshold (NDT) that uses deep reinforcement learning (RL) to learn buffer management policies without human instructions except for a high-level objective. To tackle the high complexity and scale of the buffer management problem, we develop two domain-specific techniques upon off-the-shelf deep RL solutions. First, we design a scalable RL model by leveraging the permutation symmetry of the switch ports. Second, we use a two-level control mechanism to achieve efficient training and decision-making. The buffer allocation is directly controlled by a low-level heuristic during the decision interval, while the RL agent only decides the high-level control factor according to the traffic density. Testbed and simulation experiments demonstrate that NDT generalizes well and outperforms hand-tuned heuristic policies even on workloads for which it was not explicitly trained. Mowei Wang, Sijiang Huang, Yong Cui 0001, Wendong Wang 0003, Zhenhua Li 0001 |
INFOCOM | 1 |
| 2022 | Adaptive Bitrate with User-level QoE Preference for Video StreamingabstractRecent years have witnessed tremendous growth of video streaming applications. To describe users’ expectations of videos, QoE was proposed, which is critical for content providers. Current video delivery systems optimize QoE with ABR algorithms. However, ABR is usually designed for an abstract "average user" without considering that QoE varies with users. In this paper, to investigate the difference in user preferences, we conduct a user study with 90 subjects and find that the average user can not represent all users. This observation inspires us to propose Ruyi, a video streaming system that incorporates preference awareness into the QoE model and the ABR algorithm. Ruyi profiles QoE preference of users and introduces preference-aware weights over different quality metrics into the QoE model. Based on this QoE model, Ruyi’s ABR is designed to directly predict the influence on metrics after taking different actions. With these predicted metrics, Ruyi chooses the bitrate that maximizes user-specific QoE once the preference is given. Consequently, Ruyi is scalable to different user preferences without re-training the learned models for each user. Simulation results show that Ruyi increases QoE for all users with up to 65.22% improvement. Testbed experimental results show that Ruyi has the highest ratings from subjects. Xutong Zuo, Mowei Wang, Yong Cui 0001 |
INFOCOM | 3 |
| 2022 | Traffic-Aware Buffer Management in Shared Memory SwitchesabstractSwitch buffer serves an important role in the modern internet. To achieve efficiency, today’s switches often use on-chip shared memory. Shared memory switches rely on buffer management policies to allocate buffer among ports. To avoid waste of buffer resources or excessive buffer occupation by a few ports, existing policies tend to maximize overall buffer utilization and pursue queue length fairness. However, blind pursuit of utilization and misleading fairness definition based on queue length lead to buffer occupation with no benefit to throughput but extends queuing delay and undermines burst absorption of other ports. With analysis of current dynamic threshold policies, we demonstrate that meaningless buffer occupation can potentially impair the absorption capability of shared buffer, whereas none of the existing policies have addressed this problem. We contend that a buffer management policy should proactively detect port traffic and adjust buffer allocation accordingly. In this paper, we propose Traffic-aware Dynamic Threshold (TDT) policy. On the basis of the classic dynamic threshold policy, TDT proactively raises or lowers port threshold to absorb burst traffic or evacuate meaningless buffer occupation. We present detailed designs of port control state transition and state decision module that detect real-time traffic and change port thresholds accordingly. Simulation and DPDK-based real testbed demonstrate that TDT simultaneously optimizes for throughput, loss and delay, and reduces up to 50% flow completion time. Sijiang Huang, Mowei Wang, Yong Cui 0001 |
IEEE/ACM Trans. Netw. | 2 |
| 2022 | DeepCC: Bridging the Gap Between Congestion Control and Applications via Multiobjective OptimizationabstractThe increasingly complicated and diverse applications have distinct network performance demands, e.g., some desire high throughput while others require low latency. Traditional congestion controls (CC) have no perception of these demands. Consequently, literatures have explored the objective-specific algorithms, which are based on either offline training or online learning, to adapt to certain application demands. However, once generated, such algorithms are tailored to a specific performance objective function. Newly emerged performance demands in a changeable network environment require either expensive retraining (in the case of offline training), or manually redesigning a new objective function (in the case of online learning). To address this problem, we propose a novel architecture, DeepCC. It generates a CC agent that is generically applicable to a wide range of application requirements and network conditions. The key idea of DeepCC is to leverage both offline deep reinforcement learning and online fine-tuning. In the offline phase, instead of training towards a specific objective function, DeepCC trains its deep neural network model using multi-objective optimization. With the trained model, DeepCC offers near Pareto optimal policies w.r.t different user-specified trade-offs between throughput, delay, and loss rate without any redesigning or retraining. In addition, a quick online fine-tuning phase further helps DeepCC achieve the application-specific demands under dynamic network conditions. The simulation and real-world experiments show that DeepCC outperforms state-of-the-art schemes in a wide range of settings. DeepCC gains a higher target completion ratio of application requirements up to 67.4% than that of other schemes, even in an untrained environment. Lei Zhang 0157, Yong Cui 0001, Mowei Wang, Kewei Zhu, Yibo Zhu 0001, Yong Jiang 0001 |
IEEE/ACM Trans. Netw. | 3 |
| 2021 | Traffic-aware Buffer Management in Shared Memory SwitchesabstractSwitch buffer serves an important role in modern internet. To achieve efficiency, today's switches often use on-chip shared memory. Shared memory switches rely on buffer management policies to allocate buffer among ports. To avoid waste of buffer resources or a few ports occupy too much buffer, existing policies tend to maximize overall buffer utilization and pursue queue length fairness. However, blind pursuit of utilization and misleading fairness definition based on queue length leads to buffer occupation with no benefit to throughput but extends queuing delay and undermines burst absorption of other ports. We contend that a buffer management policy should proactively detect port traffic and adjust buffer allocation accordingly. In this paper, we propose Traffic-aware Dynamic Threshold (TDT) policy. On the basis of classic dynamic threshold policy, TDT proactively raise or lower port threshold to absorb burst traffic or evacuate meaningless buffer occupation. We present detailed designs of port control state transition and state decision module that detect real time traffic and change port thresholds accordingly. Simulation and DPDK-based real testbed demonstrate that TDT simultaneously optimizes for throughput, loss and delay, and reduces up to 50% flow completion time. Sijiang Huang, Mowei Wang, Yong Cui 0001 |
INFOCOM | 2 |
| 2021 | The ACM Multimedia 2021 Meet Deadline Requirements Grand ChallengeabstractDelay-sensitive multimedia streaming applications require their data to be delivered before a deadline to be useful. The data transmitted by these applications can usually be partitioned into blocks with different priorities, assigned based on the impact of a block on the Quality of Experience (QoE) if it misses its delivery deadline. Meet their deadline requirements is challenging due to the dynamics of the network and these applications' high demand on network resources. To encourage the research community to address this challenge, we organize the "Meet Deadline Requirements" Grand Challenge at ACM Multimedia 2021. This grand challenge provides a simulation platform onto which the participants can implement their block scheduler and bandwidth estimator and then benchmark against each other using a common set of application traces and network traces. Junjie Deng, Mowei Wang, Yong Cui 0001, Wei Tsang Ooi, Jiangchuan Liu, Xinyu Zhang 0003, Kai Zheng 0003, Yi Li 0015 |
ACM Multimedia | 3 |