VLDB 2026 Research / reviewers in the wild / expert
Dehui Wei
dblp:286/4108
· DBLP profile ↗
12ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0002-6952-5062ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 10 · 4 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HyperEdge: An Edge CDN Infrastructure for Cost Efficient Video Streaming
Dehui Wei, Jiao Zhang 0002, Zhichen Xue, Yajie Peng, Xiaofei Pang, Jialin Li 0001 |
NSDI | 1 |
| 2025 | Efficient Partitioning Deep Learning Models for Medical Image Analysis on Iot DevicesabstractDeep learning models, due to their strong capabilities in learning from data and representing features, have been widely deployed, particularly in the medical field. However, given the limitations of medical devices and deployment scenarios, more and more researchers are focusing on how to better deploy deep learning models on resource-constrained edge devices to enable real-time medical data analysis. Nevertheless, such resourceconstrained IoT devices present significant challenges, making it difficult to achieve both high accuracy and low inference latency. To address this problem, we propose a novel framework, RC-DLM, designed to split and execute complex deep learning models on IoT devices. Specifically, we partition a deep learning model into multiple sub-models according to the computational capacity of each device, with each sub-model responsible for handling a subset of classes. To further reduce computation overhead and inference latency, we integrate a class-wise pruning method to shrink the size of each sub-model. Through large-scale experiments conducted on four popular datasets with three model architectures, we demonstrate that our approach significantly reduces inference latency and model size by up to 5.72 times and 57.5 times, respectively. We further deploy our method on real-world edge devices and compare it with state-of-theart approaches, evaluating the three most important aspects: accuracy, inference time, and model size. The comprehensive experimental results of our RC-DLM confirm the effectiveness of our proposed method. Xiang Liu 0017, Mengyao Zheng, Junyong Cao, Dehui Wei, Kang Lai, Huiying Lan, Yijun Song, Xia Li 0005 |
BIBM | 4 |
| 2025 | HELDR: Packet Loss Detection and Retransmission for Live Streaming Hyper-Edge NetworkabstractLive streaming platforms like Douyin have developed the Live Streaming Hyper-Edge Delivery Network (LSHEDN) to reduce bandwidth cost. In LS-HEDN, the Content Delivery Network (CDN) splits the live streaming into multiple substreams by randomly assigning each frame to them. Hyperedge devices like set-top boxes with cheap and idle bandwidth resources forward a substream from CDN to multiple users. A protocol based on User Datagram Protocol (UDP) is adopted between devices and users, with users detecting packet loss and requesting retransmissions via Negative Acknowledgment (NACK). Given the demand for lower latency and the inherent fluctuations in public network, existing receiver-side packet loss detection and retransmission methods fall short in achieving both timeliness and accuracy simultaneously. This is manifested as frequent rebuffering and excessive redundancy. Notably, when head-of-line blocking(HOL blocking) occurs in the upstream link of the device, these issues become even more pronounced. To address this, we propose Hyper-Edge Loss Detection and Retransmission (HELDR) algorithm. It features a loss detection algorithm tailored to the transmission characteristics in LSHEDN, which improves detection accuracy. Its immediate retransmission mechanism and the backup devices retransmission mechanism enhance timeliness. Large-scale online A/B tests results show that HELDR reduces the average rebuffering rate by 41.2%, reduces the average redundancy rate by 15.7%. Peisheng Guo, Jiao Zhang 0002, Zhichen Xue, Yajie Peng, Xiaofei Pang, Tao Huang 0005, Ruili Fang, Zhenpeng Zhu, Dehui Wei |
IWQoS | 13 |
| 2025 | QoE-Optimized MultiPath Scheduling for Video Services in Large-Scale Peer-to-Peer CDNsabstractVideo content providers such as Douyin implement Peer-to-Peer Content Delivery Networks (PCDNs) to reduce the costs associated with Content Delivery Networks (CDNs) while still maintaining optimal user-perceived quality of experience (QoE). PCDNs rely on the remaining resources of edge devices, such as edge access devices and hosts, to store and distribute data with a Multiple-Server-to-One-Client (MS2OC) communication pattern. MS2OC parallel transmission pattern suffers from severe data out-of-order issues. PCDNs offer significant cost savings by using multiple low-cost edge devices. However, due to its unique characteristics, including pull-based streaming transmission, many heterogeneous paths, and large receiving buffers, directly applying existing schedulers designed for Multipath TCP (MPTCP) to PCDN fails to meet the two goals of high aggregate bandwidth and low end-to-end delivery latency. To tackle this issue, we provide a detailed overview of Douyin’s self-developed PCDN video transmission system and introduce the first QoE-enhanced packet-level scheduler for PCDN systems, named Pscheduler. Pscheduler evaluates path quality with a congestion-control-decoupled algorithm and employs our proposed path-pick-packet method for data distribution, ensuring a smooth video playback experience. Additionally, we propose a redundant transmission algorithm to enhance task download speeds for segmented video transmission. Our extensive online A/B tests, involving 100,000 Douyin users generating tens of millions of video data points, demonstrate that Pscheduler achieves an average improvement of 60% in goodput, a 20% reduction in data delivery waiting time, and a 30% reduction in rebuffering rates. Furthermore, we conducted simulation experiments that further validate the effectiveness of Pscheduler, confirming its improvements in performance metrics under various network conditions. Dehui Wei, Jiao Zhang 0002, Xiang Liu 0017, Zhichen Xue, Tao Huang 0005, Linshan Jiang, Jialin Li 0001 |
IEEE J. Sel. Areas Commun. | 1 |
| 2024 | High Performance Computing Framework for Variable Selection on Genome-wide Association StudiesabstractVariable selection for genome-wide association studies (GWAS) has been a major research focus for decades. With the exponential growth of biological and biomedical data in the era of big data, scientists are confronted with the challenge of extracting meaningful information from vast datasets while managing the inherent heterogeneity in bioinformatics. To date, there are no highly effective tools that support high-dimensional datasets and achieve robust variable selection performance, all while accounting for the non-i.i.d. features and structured relatedness among explanatory and response variables.To address these challenges, we introduce the first high-performance computing framework for variable selection in GWAS. Our framework integrates various state-of-the-art methods, allowing researchers to easily combine different techniques and fully explore their potential. Additionally, our approach employs novel optimization strategies to solve the problem efficiently, even for high-dimensional data with sparse characteristics. By processing the data holistically, the framework delivers comprehensive analysis and accurate linkage mapping associations. Designed for ease of use, the framework is implemented in Python and offers seamless deployment, making it accessible to a wide range of researchers. Xiang Liu 0017, Jing Diao, Mengyao Zheng, Jihe Li, Dehui Wei, Qipeng Xie, Xia Li 0005, Linshan Jiang |
BIBM | 6 |
| 2024 | Pscheduler: QoE-Enhanced MultiPath Scheduler for Video Services in Large-scale Peer-to-Peer CDNsabstractVideo content providers such as Douyin implement Peer-to-Peer Content Delivery Networks (PCDNs) to reduce the costs associated with Content Delivery Networks (CDNs) while still maintaining optimal user-perceived quality of experience (QoE). PCDNs rely on the remaining resources of edge devices, such as edge access devices and hosts, to store and distribute data with a Multiple-Server-to-One-Client (MS2OC) communication pattern. MS2OC parallel transmission pattern suffers from severe data out-of-order issues. However, direct applying existing schedulers designed for MPTCP to PCDN fails to meet the two goals of high aggregate bandwidth and low end-to-end delivery latency.To address this, we present the comprehensive detail of the Douyin self-developed PCDN video transmission system and propose the first QoE-enhanced packet-level scheduler for PCDN systems, called Pscheduler. Pscheduler estimates path quality using a congestion-control-decoupled algorithm and distributes data by the proposed path-pick-packet method to ensure smooth video playback. Additionally, a redundant transmission algorithm is proposed to improve the task download speed for segmented video transmission. Our large-scale online A/B tests, comprising 100,000 Douyin users that generate tens of millions of videos data, show that Pscheduler achieves an average improvement of 60% in goodput, 20% reduction in data delivery waiting time, and 30% reduction in rebuffering rate. Dehui Wei, Zhichen Xue, Yajie Peng, Xiaofei Pang, Yuanjie Liu |
INFOCOM | 1 |
| 2024 | Breaking the Inertial Thinking: Non-Blocking Multipath Congestion Control Based on the Single-Subflow Reinforcement Learning ModelabstractThe Multipath TCP (MPTCP) protocol has received more attention due to the increasing number of terminals with multiple network interfaces. To meet the higher network performance demand of terminal services, many researches leverage reinforcement learning (RL) for MPTCP congestion control (CC) algorithms to improve the performance of MPTCP. However, we observe two limitations of existing RL-based mechanisms that make them impractical: 1) Fail to break the restriction of the input and output dimensions of RL, making the mechanisms unadaptable to the varying number of subflows. 2) Frequent model decisions block packet transmission, leading to under-utilization of bandwidth. This paper breaks the inertial thinking By “inertial thinking” here, we are referring to the initial reaction of others when dealing with CC in MPTCP. Given the interdependence between MPTCP subflows, scholars have traditionally opted for coupled CC. However, we have challenged this conventional thinking by independently handling the CC of different subflows in a single MPTCP flow and ensuring fairness. to overcome the above limitations and proposes Maggey, a non-blocking CC mechanism that applies the single-subflow model to multipath transmission. To this end, Maggey employs loosely coupled design principles and a unique reward function to ensure the fairness of the algorithm. Additionally, Maggey introduces iterative training to ensure the accuracy of training of the single-subflow model. Furthermore, a mode transition framework is artfully designed to avoid blocking, preserving the flexibility of RL-based CCs. These two features enhance the practicability of Maggey and the paper analyze the stability of Maggey. We implement Maggey in the Linux kernel and evaluate the performance of Maggey through extensive emulation and live experiments. The evaluation results show that Maggey boosts 26% throughput over DRL-CC at high bandwidth and improves 2%-60% throughput over traditional algorithms under different network conditions. Besides, Maggey maintains fairness in different scenarios. Dehui Wei, Jiao Zhang 0002, Yuanjie Liu, Tian Pan 0001, Tao Huang 0005 |
IEEE Trans. Netw. Serv. Manag. | 1 |
| 2023 | Performance Modeling and Analysis of Distributed Deep Neural Network Training with Parameter ServerabstractWith the growth of dataset size and the development of hardware accelerators, the application of deep neural networks (DNN) in various fields has made great breakthroughs. In order to improve the training speed of DNN, distributed training has been widely used. However, the imbalance between computation and communication makes distributed training difficult to achieve maximum efficiency. Therefore there is a need to detect the bottleneck state and verify the effect of some optimization schemes. Testing on a physical cluster incurs additional time and cost overhead. This paper builds a DNN-specific performance model that is used for bottleneck detection and tuning at a low cost. We build this model through detailed analysis and reasonable assumptions. We also focus on fine-grained modeling of scalability and network components, which are key factors affecting performance. Then we verify the performance model with an average error of 5% on testbed and emulator. Finally, we provide use cases of the performance model. Jiao Zhang 0002, Dehui Wei, Tian Pan 0001, Tao Huang 0005 |
GLOBECOM | 3 |
| 2023 | Leopard: A Pragmatic Learning-Based Multipath Congestion Control for Rapid Adaptation to New Network ConditionsabstractMultipath TCP is a multipath transport protocol deployed on end devices, and many learning-based multipath congestion control schemes have been proposed and verified. However, these schemes cannot adapt rapidly to new network conditions because of their convergence problems and generalization issues. To rapidly adapt to new network conditions, we propose Leopard, a learning-based multi-path congestion control frame-work that uses reinforcement learning to combine offline learning with online fine-tuning. The extensive experiments in emulated network conditions and the real world demonstrate that Leopard converges quickly and maintains consistent high performance in new network conditions, which avoids long retraining when the network environment changes. Leopard improves throughput by 13% compared with DRL-CC and reduces the convergence time by 20% compared with MPCC in new network conditions. Yuanjie Liu, Jiao Zhang 0002, Dehui Wei |
ICC | 3 |
| 2023 | Fast, Scalable and Robust Centralized Routing for Data Center NetworksabstractThis paper presents a fast and robust centralized data center network (DCN) routing solution, called . For fast routing calculation, uses centralized controllers to collect/disseminate the network’s link-states (LS), and offload the actual routing calculation onto each switch. Observing that the routing changes can be classified into a few fixed patterns in DCNs which have regular topologies, we simplify each switch’s routing calculation into a table-lookup manner, i.e., comparing LS changes with pre-installed base topology and updating routing paths according to predefined rules. As such, the routing calculation time at each switch only needs 10s of us even in a large network topology containing 10K+ switches. For efficient controller fault-tolerance, purposely uses reporter switch to ensure the LS updates successfully delivered to all affected switches. As such, can use multiple stateless controllers and little redundant traffic to tolerate failures, which incurs little overhead under normal case, and keeps 10s of ms fast routing reaction time even under complex data-/control-plane failures. We design, implement and evaluate with extensive experiments on Linux-machine controllers and white-box switches. provides$\sim$1200x and$\sim$100x shorter convergence time than current distributed protocol BGP and the state-of-the-art centralized routing solution, respectively. Furthermore, Primus maintains good routing controllability/manageability thanks to its centralized architecture, which enables us to build several advanced routing features in our testbed, including routing failure visualization and weighted-cost-multi-path routing. Fusheng Lin, Guo Chen 0001, Guihua Zhou, Dehui Wei, Li Chen 0008, Yuanwei Lu, Andrew Qu, Hongbo Jiang 0001 |
IEEE/ACM Trans. Netw. | 6 |
| 2021 | Primus: Fast and Robust Centralized Routing for Large-scale Data Center NetworksabstractThis paper presents a fast and robust centralized data center network (DCN) routing solution called Primus. For fast routing calculation, Primus uses centralized controller to collect/disseminates the network's link-states (LS), and offload the actual routing calculation onto each switch. Observing that the routing changes can be classified into a few fixed patterns in DCNs which have regular topologies, we simplify each switch's routing calculation into a table-lookup manner, i.e., comparing LS changes with pre-installed base topology and updating routing paths according to predefined rules. As such, the routing calculation time at each switch only needs 10s of us even in a large network topology containing 10K+ switches. For efficient controller fault-tolerance, Primus purposely uses reporter switch to ensure the LS updates successfully delivered to all affected switches. As such, Primus can use multiple stateless controllers and little redundant traffic to tolerate failures, which incurs little overhead under normal case, and keeps 10s of ms fast routing reaction time even under complex data-/control-plane failures. We design, implement and evaluate Primus with extensive experiments on Linux-machine controllers and white-box switches. Primus provides ~1200x and ~100x shorter convergence time than current distributed protocol BGP and the state-of-the-art centralized routing solution, respectively. Guihua Zhou, Guo Chen 0001, Fusheng Lin, Dehui Wei, Jianbing Wu, Li Chen 0008, Yuanwei Lu, Andrew Qu, Hongbo Jiang 0001 |
INFOCOM | 5 |
| 2020 | PLB: Adaptive Partial Congestion-aware Load Balancing for Datacenter NetworksabstractIn order to accommodate ever-increasing new tenants and applications, datacenter networks (DCNs) require an efficient load balancing scheme to fully utilize their bisection bandwidth. Equal-cost MultiPath routing (ECMP) is a widely used load-balancing mechanism in the DCN. However, ECMP blindly hashes traffic to parallel paths and results in imbalance and collisions. Motivated by ECMP's shortcomings, some recent schemes provide more visibility into networks via active probing. They could be broadly classified as probing all the paths or a fixed number of paths (e.g., 3 paths) each probe interval. However, they all suffer from some limitations. Probing all paths introduces high probing overhead while probing a fixed number of paths is suboptimal when the network topology and traffic load change. To our best knowledge, none of the existing schemes adapt the number of paths being probed to the network conditions. Enlightened by the defects of previous work, we introduce PLB, an adaptive partial congestion-aware load-balancing mechanism. At its heart, PLB randomly probes partial paths each probe interval and the number of them changes according to the network topology and the traffic load. Besides, PLB splits flow into flowlets and makes careful routing/rerouting decisions for them. Through analysis, we formulate the correlations between the number of paths being probed and the network conditions. Furthermore, simulations with realistic workloads validate our conclusions and show that PLB reduces overall flow completion times compared to the state-of-the-art load balancing schemes both in symmetric and asymmetric topologies. Kefei Liu 0004, Jiao Zhang 0002, Dehui Wei, Tao Huang 0005 |
GLOBECOM | 3 |