EDBT 2026 Demo / reviewers in the wild / expert
Weichao Li 0001
dblp:94/1464-1
· DBLP profile ↗
51ranked-venue papers
5as first author
29since 2021 · last 2026
0000-0002-7620-1955ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 38 · 5 first-author · 21 since 2021Security and privacy · 4 · 2 since 2021Systems, architecture and hardware · 3 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Unveiling Low‑Altitude 5G Performance: Linking Key Influencing Factors with UAV Flight Parameters
Jianer Zhou, Xiaoyong Ni, Ke Luo 0001, Zhenyu Li 0001, Xiaofeng Tao 0001, Weichao Li 0001 |
SIGCOMM | 7 |
| 2026 | Resource Allocation in RIS-Assisted Integrated Sensing, Communication, and Computation NetworkabstractIntegrated sensing and communication (ISAC) is an emerging paradigm designed to support next-generation wireless services and applications. However, ISAC systems with limited computation capabilities are unable to handle computation-intensive and latency-sensitive tasks. This paper proposes a novel integrated sensing, communication, and computation (ISCC) network empowered by a reconfigurable intelligent surface (RIS) to mitigate the performance degradation caused by interference between radar sensing and uplink offloading. To effectively coordinate the cross-layer resource allocation among communication, sensing, and computation, we propose a resource scheduling problem. Specifically, we maximize the total computation rate while satisfying the sensing signal-to-noise ratio (SNR) requirement by jointly optimizing the energy allocation for local computing and offloading, the transmit and receive beamforming at the base station (BS), and the RIS reflective beamforming. To address this complex non-convex problem, we develop an efficient scheduling algorithm based on the block coordinate descent (BCD) framework. The iterative algorithm employs the fractional programming algorithm based on Lagrangian dual transform and quadratic transform, the generalized eigenvector methods, the convex relaxation techniques, and the successive convex approximation (SCA) algorithms to solve each subproblem separately. Experimental results demonstrate that the proposed scheme outperforms several baseline methods, confirming that RIS technology can effectively enhance system performance. In addition, we reveal the impact of various parameters on system performance. Yingsheng Peng, Jinbei Zhang, Jingpu Duan, Weichao Li 0001, Yong Liu 0005 |
IEEE Trans. Commun. | 4 |
| 2026 | Average Transmission Rate on D2D Coded Caching With Nonuniform File Popularity
Jinbei Zhang, Wenjie Guan, Kai Huang 0012, Jingjing Luo, Weichao Li 0001 |
IEEE Trans. Commun. | 5 |
| 2025 | DeSync: Proactive Congestion Control via Random Delay Offsets for Large-Scale ML TrainingabstractSynchronization-induced congestion is a critical performance bottleneck in modern distributed machine learning (ML) training, where simultaneous gradient exchanges create bursty traffic patterns. Existing solutions, both reactive and proactive, struggle to balance throughput and latency in the presence of synchronized flows. We propose DeSync, a proactive traffic shaping scheme that introduces structured random delay to de-synchronize communication rounds. Evaluations with DCQCN, HPCC, DCTCP, and TIMELY demonstrate that DeSync significantly improves FCT, job completion times, and congestion metrics, enhancing existing CC mechanisms without specialized hardware. Xingbo Feng, Zhuyun Qi, Yi Wang 0004, Ziyao Huang 0001, Yan Liu 0062, Jiashuo Lin, Chenxi Ling, Weichao Li 0001, Jin Zhang 0001, Jianping Wang 0001 |
IWQoS | 8 |
| 2025 | Flux: Fine-Grained Communication Scheduling for Distributed Training in Multi-Tenant AI ClustersabstractCommunication overhead is a major bottleneck in distributed AI training, particularly in multi-tenant environments, limiting GPU utilization. Existing job-level scheduling methods fail to address the varying urgency of individual communication operations. We propose Flux, a novel fine-grained scheduler that prioritizes communication operations based on their Urgency Score and job intensity. Our evaluation shows Flux improves GPU utilization by up to 10 % compared to state-of-the-art job-level algorithms. This demonstrates the significant advantage of fine-grained communication scheduling in multitenant AI clusters. Jiashuo Lin, Xingbo Feng, Hanrui Qi, Yan Liu 0062, Chenxi Ling, Bo Tang 0016, Yi Wang 0004, Xiaofeng Tao 0001, Weichao Li 0001 |
IWQoS | 9 |
| 2025 | ReCQF: Enhancing CQF Redundancy with Delay Alignment Scheduling in TSNabstractIntegrating Frame Replication and Elimination for Reliability (FRER) with Cyclic Queuing and Forwarding (CQF) in Time-Sensitive Networks (TSN) encounters redundancy failures and resource reservation inefficiencies due to length disparities across redundant paths. To address these challenges, we propose ReCQF, a Reliability-Enhanced CQF scheduling framework built on Multi-Instance CQF. ReCQF adaptively assigns redundant flows to multiple CQF queue pairs with specific cycles, effectively aligning transmission delays across redundant paths to ensure low delay and inter-path delay differences while significantly reducing resource reservations. Yan Liu 0062, Zhuyun Qi, Xingbo Feng, Shuangping Zhan, Yao Xin, Jiashuo Lin, Chenxi Ling, Ruide Cao, Weichao Li 0001, Yi Wang 0004 |
IWQoS | 9 |
| 2025 | From Limited Resources to Powerful Insights: Empowering Low-Cost Cameras for Efficient Retrospective QueryingabstractUploading videos from low-cost cameras to the cloud for retrospective analysis presents challenges in privacy, network, and computation. To address these issues and achieve low latency, we propose READY, a novel client-cloud collaborative system. READY aims to enhance the quality of uploaded frames by selectively uploading only the frames relevant to queries. To achieve this, READY establishes an index during video capture, recording object categories and probabilities for each frame. READY adopts an innovative semi-supervised approach for frame indexing, wherein frames are indexed through a continuously updated feature distribution space constructed by k-nearest neighbors (KNN). This enables resource-constrained low-cost cameras to independently establish long-term frame indexes. Additionally, READY utilizes progressively improving operators (lightweight classification models) dispatched by the cloud to optimize the upload order of frames, prioritizing positive frames. By sharing the backbone of low-performance operators, high-performance operators can be efficiently executed, significantly enhancing the camera’s frame processing capability. The established frame index also enables efficient multiple consecutive queries on different classes. Over 110 h of diverse queries across 11 videos, READY outperformed competing alternative designs by achieving an average response time of 67.8% and reducing the proportion of uploaded videos by an average of 79.8% (compared to CloudOnly). Qiaodi Wen, Jianer Zhou, Ziqi Luo, Gareth Tyson, Weichao Li 0001, Jinfan Wang |
IEEE Internet Things J. | 6 |
| 2025 | FlexTAS: Flexible Gating Control for Enhanced Time-Sensitive Networking DeploymentabstractTime-sensitive networking (TSN), essential in industrial networks for its promise of reliable and deterministic data transmission, faces deployment challenges due to the limitations of existing time-aware shaper (TAS)-based scheduling algorithms. Specifically, the size of the generated gate control lists (GCLs) is usually too large to be deployed in actual devices. To bridge the gap between theory and practice, we propose FlexTAS, a flexible and practical solution for TSN. The key insight behind FlexTAS is that relaxing gating does not introduce uncertainty, as long as nonoverlap reserved time slots are guaranteed. FlexTAS is comprised of two main components: first, a novel gating model deviates from the conventional TAS model by incorporating selective relaxation of gating at certain nodes; and second, a deep reinforcement learning-based engine to rapidly generate valid schedules. We build a real testbed and validate the effectiveness of our proposed solution. Our evaluation demonstrates that FlexTAS effectively controls the number of gate entries within the GCL capacity of devices, while simultaneously meeting the Quality of Service(QoS) requirements of time-triggered streams. It significantly reduces the number of GCL entries by 60% to 80%, and facilitates deployment in heterogeneous networks, thus offering a practical solution for TSN. Jiashuo Lin, Weichao Li 0001, Xingbo Feng, Shuangping Zhan, Lewei Ning, Yi Wang 0004, Tao Wang 0014, Hai Wan, Bo Tang 0016, Xiaofeng Tao 0001 |
IEEE Trans. Ind. Informatics | 2 |
| 2025 | Efficient Data Center Network Monitoring and Troubleshooting With LMon: Leveraging ECMP Hashing Linearity and Lightweight ProbingabstractNetwork performance monitoring and troubleshooting are crucial yet challenging tasks in datacenter management. Despite the numerous solutions that have been proposed in recent years, their efforts are often hindered by high costs and unreliable failure localization, making it difficult to deploy them in real-world environments. In this paper, we presentLMon, a highly reliable and efficient system for monitoring and troubleshooting in datacenter networks. LMon utilizes the characteristic of ECMP hashing linearity to control probe packet routing, enabling the monitoring of targeted paths without any modification of underlying protocols and devices. Additionally, LMon leverages a lightweight probing technique to reduce monitoring overhead, as well as integrates the improved LASSO regression and hypothesis testing for higher accuracy and faster processing in link failure localization. We evaluate the performance of LMon in our testing environment. Compared to the monitoring system Pingmesh, LMon generates only one-third probes while maintaining 99% accuracy and 1% false negatives. Qinglin Xun, Weichao Li 0001, Jianer Zhou, Jingpu Duan, Yi Wang 0004, Xiaofeng Tao 0001, Jinbei Zhang |
IEEE Trans. Netw. | 2 |
| 2025 | Safety in DRL-Based Congestion Control: A Framework Empowered by Expert RefinementabstractDeep reinforcement learning (DRL) has been used in congestion control algorithms (CCAs) for its ability to adapt to different network environments. However, its effectiveness is often hindered by the limited availability of training data and constrained training scales. While it has been proved that combining rule-based (expert) CCAs as a guide for DRL (namely hybrid CCAs) can address this limitation, we show through experimental measurements that rule-based CCAs potentially restrict action exploration of DRL models and may cause the DRL models to overly rely on them for higher reward gains. To address this gap, this paper proposes Marten, a framework that improves the effectiveness of rule-based CCAs for DRL. Marten’s key innovations include an entropy-based dynamic exploration scheme that expands the exploration of DRL, and a reward adjustment scheme to prevent the DRL models’ over-reliance on experts in hybrid CCAs. We have implemented Marten in both simulation platform OpenAI Gym and deployment platform QUIC. Experimental results in both emulated and production networks demonstrate Marten can improve throughput by 0.31% and reduce latency by 12.69% on average compared to the state-of-the-art hybrid CCAs. Compared to BBR, Marten achieves a 2.79% increase in throughput and an 11.73% reduction in latency on average. Jianer Zhou, Zhiyuan Pan, Zhenyu Li 0001, Gareth Tyson, Weichao Li 0001, Xinyi Qiu, Xinyi Zhang 0004, Gaogang Xie |
IEEE Trans. Netw. | 5 |
| 2024 | RobustTSN: A Framework for Protecting Time-Sensitive Networking against Unexpected DelaysabstractIndustrial networks require deterministic and reliable communication, which can be achieved by Time-Sensitive Networking (TSN), a set of standards that enable precise timing and synchronization of data transmission. However, TSN is susceptible to unexpected delays caused by device malfunction, interference or cyber attacks, which can have a domino effect and disrupt multiple data flows. To address this challenge, we propose RobustTSN, a framework that protects TSN against the domino effect of delayed frames and tolerates harmless accident frames using Per-Stream Filtering and Policing (PSFP) mechanism. We develop algorithms to calculate ingress filtering schedules based on local-safe delay and global-safe interval concepts, which decide whether to accept or discard out-of-schedule frames. We use a finite state machine to model the interaction between frames and evaluate frame safety. We build a software-defined networking based system to dynamically monitor network states and reconfigure device filtering after out-of-schedule transmission occurs. We conduct experiments on practical scenario topologies and large groups of random flows to demonstrate the effectiveness and efficiency of our framework. Xingbo Feng, Yi Wang 0004, Jiashuo Lin, Weichao Li 0001, Shuangping Zhan, Yan Liu 0062, Jin Zhang 0001, Jianping Wang 0001 |
IWQoS | 4 |
| 2024 | rpkt: A Generic, Safe, and Efficient Userspace Packet Processing Library in Rust
Yupeng Xiao, Yaxuan Chen, Jingpu Duan, Xiaoxi Zhang 0001, Weichao Li 0001, Xiaofeng Tao 0001 |
NPC (2) | 6 |
| 2024 | On the Waveform Design and Performance Enhancement of Multi- Target Detection in Dual-Function Radar-Communication SystemabstractDual-function radar-communication (DFRC) system has been recognized as a potential technology to address the issues of radio frequency spectrum congestion. Despite the advan-tages of the existing orthogonal frequency-division multiplexing (OFDM) chirp waveform, such as its high range resolution and low peak-to-average ratio, it is still plagued by issues related to ghost targets in complex communication environments with multiple targets. In this paper, a novel waveform, leveraging trapezoidal frequency modulation OFDM, is introduced to address the challenge of multi-target detection scenarios. Addition-ally, a power allocation and subcarrier assignment algorithm has been developed to maximize communication performance while adhering to the radar performance threshold, thereby achieving overall system optimization while enhancing multi-target detection capabilities. Finally, the effectiveness of the proposed algorithm is demonstrated through simulation experiments, showcasing its ability to handle multi-target detection scenarios and achieve superior system performance while maintaining a delicate equilibrium between communication and radar considerations. Songning Gao, Jiaxiang Geng, Weichao Li 0001, Yan-Zhao Hou, Qimei Cui, Xiaofeng Tao 0001 |
WCNC | 4 |
| 2024 | Saliency-aware regularized graph neural network
Wenjie Pei, Weina Xu, Weichao Li 0001, Jinfan Wang, Guangming Lu 0002, Xiangrong Wang 0002 |
Artif. Intell. | 4 |
| 2024 | Advancing TSN flow scheduling: An efficient framework without flow isolation constraintabstractIn the domain of Time-Sensitive Networking (TSN), the quest for ultra-reliable low-latency communication is paramount. Current scheduling strategies, which hinge on strict isolation to ensure low latency and jitter, confront the challenges of high overhead in worst-case latency evaluation and consequent limitations in network flow capacity. This paper introduces an innovative framework that transcends traditional isolation constraints, thereby expanding the solution space and augmenting network schedulability. At the heart of this framework lies a novel latency jitter analysis method that assesses the viability of non-isolation scenarios with constant time complexity. This method underpins a heuristic scheduling algorithm that not only boasts the smallest time complexity among existing heuristics but also significantly increases the number of scheduled flows. Complementing this, we integrate a discrete time reference approach to hasten time-intensive scheduling operations, achieving an optimal balance between schedulability and runtime efficiency. The framework further incorporates a workload-shifting technique to enhance online scheduling responsiveness. It adeptly manages the variability in scheduling times caused by disharmonious flow periods, further bolstering the framework’s robustness. Experimental validations demonstrate that our framework can increase the scheduled flows up to 269%. It reduces scheduling runtime by up to 98.44% for medium-scale networks while maintaining a flat runtime growth curve, ensuring predictable performance in online scheduling scenarios. Xingbo Feng, Yi Wang 0004, Jiashuo Lin, Weichao Li 0001, Shuangping Zhan, Yan Liu 0062, Jin Zhang 0001, Jianping Wang 0001 |
Comput. Networks | 4 |
| 2024 | Stochastic Long-Term Energy Optimization in Digital Twin-Assisted Heterogeneous Edge NetworksabstractMobile edge computing (MEC) and digital twin (DT) technologies have been recognized as key enabling factors for the next generation of industrial Internet of Things (IoT) applications. In existing works, DT-assisted edge network resource optimization solutions mostly focus on short-term performance optimization, and long-term resource optimization has not been well studied. Thus, this paper introduces a digital twin-assisted heterogeneous edge network (DTHEN), aiming to minimize long-term energy consumption by jointly optimizing transmit power and computing resource. To solve the stochastic optimization problem, we propose a long-term queue-aware energy minimization (LQEM) scheme for joint communication and computing resource management. The proposed scheme uses Lyapunov optimization to transform the original problem with long-term time constraints into a deterministic upper bound problem for each time slot, decouples it into three independent sub-problems, and solves each sub-problem separately. We then theoretically prove the asymptotic optimality of the LQEM scheme and the tradeoff between system energy consumption and task queue backlog. Finally, experimental results verify the performance analysis of the LQEM scheme, demonstrating its superiority over several benchmark schemes, and reveal the impact of various parameters on the system. Yingsheng Peng, Jingpu Duan, Jinbei Zhang, Weichao Li 0001, Yong Liu 0005, Fuli Jiang |
IEEE J. Sel. Areas Commun. | 4 |
| 2024 | Power Demand Reshaping Using Energy Storage for Distributed Edge CloudsabstractThe booming edge computing market that is supported by the edge cloud (EC) infrastructure has brought huge operating costs, mainly the energy cost, to edge service providers. The energy cost in form of electricity bills usually consists of energy charge and demand charge, and the demand charge based on peak power may account for a large proportion of the energy cost given a significant fluctuating power curve. In this work, we investigate the backup battery characteristics and electricity charge tariffs at ECs and explore the corresponding cost-saving potential. Specifically, we transform the backup battery group into distributed battery energy storage system (BESS) and strategically schedule the BESS to minimize the energy cost of service providers. We then propose a deep reinforcement learning (DRL) based approach to BESS charging/discharging in coping with the dynamic power demand and BESS state at each EC. To enable better decision-making and speed up agent training, we further design the customized invalid action masking (IAM) method and apply the prioritized experience replay (PER) scheme. The experiment results based on real-world EC power traces show that the proposed approach can reduce the demand charge and overall electricity bill by up to 27% and 13%, respectively. Dongyu Zheng, Lei Liu 0003, Guoming Tang, Yi Wang 0004, Weichao Li 0001 |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2023 | Marten: A Built-in Security DRL-Based Congestion Control Framework by Polishing the ExpertabstractDeep reinforcement learning (DRL) has been proved to be an effective method to improve the congestion control algorithms (CCAs). However, the lack of training data and training scale affect the effectiveness of DRL model. Combining rule-based CCAs (such as BBR) as a guide for DRL is an effective way to improve learning-based CCAs. By experiment measurement, we find that the rule-based CCAs limit the action exploration and even cause DRL’s excessive dependence to gain higher DRL’s reward gain. To overcome the constraints, we propose Marten, a framework which improves the effectiveness of rule-based CCAs for DRL. Marten uses entropy as the degree of exploration and uses it to expand the exploration of DRL. Furthermore, Marten introduces the shielding mechanism to avoid wrong DRL actions. We have implemented Marten in both simulation platform OpenAI Gym and deployment platform QUIC. The experimental results in production network demonstrate Marten can improve throughput by 0.36% and reduce latency by 14.89% on average compared with Eagle, and improve throughput by 2.79% and reduce latency by 11.73% on average compared with BBR. Zhiyuan Pan, Jianer Zhou, Xinyi Qiu, Weichao Li 0001 |
INFOCOM | 4 |
| 2023 | Enabling Reliable and Efficient Performance Monitoring and Troubleshooting in Datacenter NetworksabstractNetwork performance monitoring and troubleshooting is a crucial but challenging task in datacenter management. Despite the numerous solutions that have been proposed in recent years, their efforts are often hindered by high costs and unreliable fault localization, making it difficult to deploy them in real-world environments. In this paper, we present LMon, a highly reliable and efficient system for monitoring and troubleshooting in datacenter networks. LMon utilizes the characteristic of ECMP hashing linearity to control the packet routing without any modification of the underlying protocols. Additionally, LMon leverages a lightweight probing technique to reduce monitoring overhead. Furthermore, the system integrates improved LASSO regression and statistical hypothesis testing for higher accuracy and faster processing in link failure localization. The effectiveness of LMon is demonstrated through its implementation and evaluation in ns-3 simulation. The results validate the reliability and efficiency of the system, making it a promising option for ensuring long-term network maintenance in datacenters. Qinglin Xun, Weichao Li 0001, Haorui Guo, Qianyi Huang, Jianer Zhou, Jingpu Duan, Yi Wang 0004, Jinbei Zhang |
IWQoS | 2 |
| 2023 | A Scalable Asynchronous Traffic Shaping Mechanism for TSN with Time Slot and PollingabstractThe IEEE 802.1 Time Sensitive Networking (TSN) task group is devoted to improving deterministic delay during data communication. To schedule traffic in TSN, Asynchronous Traffic Shaping (ATS) has been introduced to guarantee bounded maximum delays without complicated time synchronization, and Urgency Based Scheduler (UBS) becomes the default implementation for ATS. However, due to the traffic fluctuation, UBS cannot achieve best delay bounds if some parameters are not configured appropriately in real time. What is more, UBS needs to consumes excessive butter resources, which prevents TSN from being deployed in real world. To solve these problems, we propose Time Sensitive Queuing (TSQ), a novel ATS mechanism. TSQ lowers the difficulty of TSN deployment by removing the parameter that need to be configured in real time. In addition, to reduce the butter consumption, TSQ applies time-slot-based packet allocation mechanism as the enqueue strategy, and time-slot-based packet polling mechanism as the dequeue strategy, respectively. Our Network Calculus analysis shows that TSQ can provide bounded delays and up to 40% butter resource reduction compared to UBS. The extensive experiments implemented in ns-3 show that TSQ consumes 33% less butter resource compared to UBS. Haorui Guo, Weichao Li 0001, Jiashuo Lin, Jianer Zhou, Qinglin Xun, Shuangping Zhan, Yi Wang 0004, Qingsha Cheng |
NOMS | 2 |
| 2023 | Mew: Enabling Large-Scale and Dynamic Link-Flooding Defenses on Programmable SwitchesabstractLink-flooding attacks (LFAs) can cut off the Internet connection to selected server targets and are hard to mitigate because adversaries use normal-looking and low-rate flows and can dynamically adjust the attack strategy. Traditional centralized defense systems cannot locally and efficiently suppress malicious traffic. Though emerging programmable switches offer an opportunity to bring defense systems closer to targeted links, their limited resource and lack of support for runtime reconfiguration limit their usage for link-flooding defenses.We present Mew1, a resource-efficient and runtime adaptable link-flooding defense system. Mew can counter various LFAs even when a massive number of flows are concentrated on a link, or when the attack strategy changes quickly. We design a distributed storage mechanism and a lossless state migration mechanism to reduce the storage bottleneck of programmable networks. We develop cooperative defense APIs to support multi-grained co-detection and co-mitigation without excessive overhead. Mew's dynamic defense mechanism can constantly analyze network conditions and activate corresponding defenses without rebooting devices or interrupting other running functions. We develop a prototype of Mew by using real-world programmable switches, which are located in five cities. Our experiments show that the real-world prototype can defend against large-scale and dynamic LFAs effectively. Huancheng Zhou, Sungmin Hong, Xiapu Luo, Weichao Li 0001, Guofei Gu |
SP | 5 |
| 2023 | Real-Time Super-Resolution: A New Mechanism for XR over 5G-AdvancedabstractExtended Reality (XR) has attracted great attention from both academic and industry, for providing users with an immersive experience anywhere. Nowadays, XR video streaming service is evolving to high definition (HD), which results in massive data traffic with more stringent latency requirement. Due to the two characteristics, it is challenging to support the commercial use for XR service in the current New Ratio (NR) network. In this paper, we propose a real-time super-resolution (RTSR) framework for XR HD video transmission. The basic idea is to utilize the overfitting feature of Deep Neural Network (DNN) to learn the non-linear mapping between low-definition (LD) video frames and HD video frames. The cloud XR server can transmit the LD frames together with the dedicated super resolution (SR) models instead of sending HD frames directly. The receiver can recover the HD frames locally with the inferencing ability of SR model. In addition, by introducing the online training and layered transmission strategy, the SR model update period can be adaptively adjusted according to the scenario changes, which also reduces the transmission overhead. Simulation results demonstrate the superiority of our proposed RTSR, which can save up to 50% traffic and increase the XR capacity about 40% compared with the conventional SR scheme. In terms of the system capacity, our results show that the average number of UEs can reach about 23 per cell under the common settings of Dense Urban. Weichao Chen 0001, Youlong Cao, Erkai Chen, Guohua Zhou, Weichao Li 0001 |
WCNC | 6 |
| 2023 | Cable: A framework for accelerating 5G UPF based on eBPF
Jianer Zhou, Zengxie Ma, Weijian Tu, Xinyi Qiu, Jingpu Duan, Zhenyu Li 0001, Qing Li 0006, Xinyi Zhang 0004, Weichao Li 0001 |
Comput. Networks | 9 |
| 2022 | HybridTSS: A Recursive Scheme Combining Coarse- and Fine- Grained Tuples for Packet ClassificationabstractThe popular OpenFlow virtual switch Open vSwitch (OVS) uses a variant of Tuple Space Search (TSS) for packet classification. Although it is easy for rule updates, the lookup performance is poor. By introducing partial trees into TSS, the recently proposed CutTSS improves the lookup performance of TSS. However, it is challenging to replace TSS in OVS for two reasons: (1) the hand-tuned partitioning heuristics are rule-set dependent; (2) the complex and irregular data structures make it difficult to be integrated and maintained in real systems. To address these issues, we propose HybridTSS, a recursive TSS scheme for fast packet classification in OVS, which exploits three novel ideas: (1) the recursive partitioning based on reinforcement learning balances global rule partitions with low training complexity; (2) a hybrid TSS scheme combining coarse-grained and fine-grained tuples suppresses tuple explosion in TSS; (3) a heterogeneous search algorithm consisting of TSS and linear search adapts to characteristics of rules at different scales for fast lookups. Using ClassBench, we show that, while immune from the main drawbacks of CutTSS, HybridTSS retains the update performance of TSS, and achieves almost an order of magnitude higher lookup performance than TSS, making it an ideal packet classification algorithm for OVS. Yuxi Liu 0017, Yao Xin, Wenjun Li 0004, Haoyu Song 0001, Ori Rottenstreich, Gaogang Xie, Weichao Li 0001, Yi Wang 0004 |
APNet | 7 |
| 2022 | FABMon: Enabling Fast and Accurate Network Available Bandwidth EstimationabstractCharacterizing the end-to-end network available bandwidth (ABW) is an important but challenging task. Although a number of ABW estimation tools have been introduced over the past two decades, applying them to the real-world networks is still difficult because of the biased results, heavy load, and long measurement time. In this paper, we propose a novel Burst Queue Recovery (BQR) model to infer the ABW. BQR first induces an instant network congestion and then observes the one-way delay (OWD) variation until the tight link recovers from the congestion. By correlating the OWDs with the queue length variation, BQR can calculate the ABW accurately. Compared to the traditional probe gap model (PGM) and probe rate model (PRM), our theoretical analysis and simulations show that BQR is more tolerant to the transient traffic burst and supports the scenarios with multiple congestible links. Based on the model, we build FABMon, a fast and accurate ABW estimation tool. Our experiments show that FABMon can measure ABW within 50 milliseconds, and achieve much more accurate measurement results than the existing tools with a very small volume of probe packets. Weichao Li 0001, Qing Li 0006, Qianyi Huang, Yong Jiang 0001, Shutao Xia |
IWQoS | 2 |
| 2022 | Rethinking the Use of Network Cycle in Time-Sensitive Networking (TSN) Flow SchedulingabstractTime-Sensitive Networking (TSN) is an emerging network architecture that provides bounded latency and reliable network services for time-sensitive applications. Since time-triggered flows in TSN are typically periodic, a concept of network cycle is widely used in both standards and academic researches. However, although network cycle has gained popularity, its rationale has not yet been analyzed systematically.In this paper, we mathematically evaluate the performance of several flow scheduling algorithms in terms of flow schedulability with and without employing network cycle. We observe that only when the network cycle is set to a proper value can the performance of flow scheduling be significantly improved. To better evaluate the scheduling effect, a novel assessment metric and a goal-based optimization algorithm are introduced. Our experiment results show that the network cycle-based algorithm can achieve a considerable improvement (40% - 170% improvement in the number of scheduled flows) compared to the ones with network cycle disabled. Jiashuo Lin, Weichao Li 0001, Xingbo Feng, Shuangping Zhan, Jingbin Feng, Jian Cheng 0004, Tao Wang 0014, Qing Li 0006, Yi Wang 0004, Fuliang Li, Bo Tang 0016 |
IWQoS | 2 |
| 2022 | Design and Implementation of Web-Based Speed Test Analysis Tool Kit
Rui Yang 0036, Ricky K. P. Mok, Shuohan Wu, Xiapu Luo, Hongyu Zou, Weichao Li 0001 |
PAM | 6 |
| 2022 | Rethinking Fine-Grained Measurement From Software-Defined Perspective: A SurveyabstractNetwork measurement provides operators an efficient tool for many network management tasks such as performance diagnosis, traffic engineering and intrusion prevention. However, with the rapid and continuous growth of traffic speed, it needs more computing and memory resources to monitor traffic in per-flow or per-packet granularity. Sample-based measurement systems (e.g., NetFlow, sFlow) have been developed to perform coarse-grained measurement, but they may miss part of records, especially for mice flows, which are important for some network management tasks (e.g., anomaly detection, performance diagnosis). To address these issues, data streaming algorithms such as hash tables and sketches have been introduced to balance the trade-off among accuracy, speed, and memory usage. In this article, we present a systematic survey of various data structures, algorithms and systems which have been proposed in recent years to perform fine-grained measurement for high-speed networks. We organize these methods and systems from a software-defined perspective. In particular, we abstract fine-grained network measurement into three-layer architecture. We introduce the responsibility of each layer and categorize existing state-of-the-art works into this architecture. Finally, we conclude the article and discuss the future directions of fine-grained network measurement. Chen Tian 0001, Long Cheng 0005, Qun Huang 0001, Weichao Li 0001, Yi Wang 0004, Qianyi Huang, Jiaqi Zheng 0001, Yi Wang 0071, Wan-Chun Dou, Guihai Chen |
IEEE Trans. Serv. Comput. | 6 |
| 2021 | A measurement study on device-to-device communication technologies for IIoT
Fuliang Li, Zhenbei Guo, Bocheng Liang, Xiushuang Yi, Xingwei Wang 0001, Weichao Li 0001, Yi Wang 0004 |
Comput. Networks | 6 |
| 2020 | Scalable Traffic Engineering for Higher Throughput in Heavily-loaded Software Defined NetworksabstractExisting traffic engineering (TE) solutions perform well for software defined network (SDN) in average cases. However, during peak hours, bursty traffic spikes are challenging to handle, because it is difficult to react in time and guarantee high performance even after failures with limited flow entries.We propose TED, a scalable TE system that can guarantee high throughput in peak hours. TED can quickly compute a group of maximum number of edge-disjoint paths for each ingress-egress switch pair. Such paths are suitable for well connected networks with unique edge capacity and TED is not limited to use only these paths. We design two methods to select paths under the limit of flow table size. We then input the selected paths to TED to minimize the maximum link utilization. In case of large traffic matrix making the maximum link utilization larger than 1, we input the utilization and the traffic matrix to the optimization of maximizing overall throughput under a new constrain. Thus we obtain a realistic traffic matrix, which has the maximum overall throughput and guarantees no traffic starvation. Experiments show that TED has much better performance for heavily-loaded SDN and has 10% higher probability to satisfy all (> 99.99%) the traffic after a single link failure for G-Scale topology than Smore under the same limit of flow table size. Che Zhang, Yi Wang 0004, Weichao Li 0001, Bo Jin 0002, Ricky K. P. Mok, Qing Li 0006, Hong Xu 0001 |
NOMS | 4 |
| 2020 | Software-Defined Networking-Assisted Content Delivery at Edge of Mobile Social NetworksabstractWith the explosive growth of mobile devices at the edge of mobile social networks (MSNs), the amount of the content that needs to be transmitted is exploded. Traditional content delivery mechanisms leverage only local information to make routing decisions, which results in both high latency and low delivery rate. Software-defined networking (SDN) is a novel network paradigm, the design philosophy of which could be applied to MSN for improving the content delivery performance. In this article, the centralized control thought of SDN is introduced into MSN to efficiently process social information. The classical routing algorithm of BubbleRap is improved from the perspective of network density, which is the basis of designing the sparse and dense routing mechanisms for MSNs. In addition, flexibly switching between these two routing mechanisms is implemented by a discriminating scheme, achieving efficient yet adaptive routing. The experimental results show that the delivery ratio of sparse routing is up to 83%, and the dense routing could reach up to 93%. Fuliang Li, Yaoguang Lu, Xingwei Wang 0001, Yuanguo Bi, Tian Pan 0001, Yuchao Zhang 0004, Weichao Li 0001, Yi Wang 0004 |
IEEE Internet Things J. | 7 |
| 2020 | A Local Communication System Over Wi-Fi Direct: Implementation and Performance EvaluationabstractWireless communication demands increase sharply with the explosive growth of mobile devices. The communications mainly depend on the infrastructure-based networks, e.g., WLANs and cellular networks. However, such wireless connections may be unavailable in crowded areas (e.g., concert and conference hall) or interrupted by infrastructure failures caused by earthquake or tsunami. These promote the evolution of local communication systems over device-to-device communication, such as Bluetooth and Wi-Fi Direct (WFD). However, none of the existing studies construct a full-featured local communication system, and they do not consider how to support the user mobility either. In this article, we implement and evaluate the performance of a WFD-based local communication system. First, we improve the intragroup communication by the native implementation of WFD on the Android platform, and propose an application-layer forwarding solution for the intergroup communication, which can be applied to three or more connected groups. Then, we put forward a self-adaptive handover mechanism taking user mobility and node failures into account. To deal with the uncertainty in the handover decision procedure, a fuzzy-logic-based normalized quantitative decision algorithm (FNQD) with the weights derived from the fuzzy analytic hierarchy process (FAHP) is utilized. Finally, we evaluate the performance of the system through both simulation and experiment analysis. Results show that we can get a maximum throughput of 31.7 Mb/s for the intragroup communication and a maximum goodput of 4.76 Mb/s for the intergroup communication. What is more, mobile devices could perform various types of handover according to their roles and status, which could improve the robustness of the local communication system. Fuliang Li, Xingwei Wang 0001, Jiannong Cao 0001, Xuefeng Liu 0001, Yuanguo Bi, Weichao Li 0001, Yi Wang 0004 |
IEEE Internet Things J. | 7 |
| 2019 | LEAP: learning-based smart edge with caching and prefetching for adaptive video streamingabstractDynamic Adaptive Streaming over HTTP (DASH) has emerged as a popular approach for video transmission, which brings a potential benefit for the Quality of Experience (QoE) because of its segment-based flexibility. However, the Internet can only provide no guaranteed delivery. The high dynamic of the available bandwidth may cause bitrate switching or video rebuffering, thus inevitably damaging the QoE. Besides, the frequently requested popular videos are transmitted for multiple times and contribute to most of the bandwidth consumption, which causes massive transmission redundancy. Therefore, we propose a Learning-based Edge with cAching and Prefetching (LEAP) to improve the online user QoE of adaptive video streaming. LEAP introduces caching into the edge to reduce the redundant video transmission and employs prefetching to fight against network jitters. Taking the state information of users into account, LEAP intelligently makes the most beneficial decisions of caching and prefetching by a QoE-oriented deep neural network model. To demonstrate the performance of our scheme, we deploy the implemented prototype of LEAP in both the simulated scenario and the real Internet. Compared with all selected schemes, LEAP at least raises average bitrate by 34.4% and reduces video rebuffering by 42.7%, which leads to at least 15.9% improvement in the user QoE in the simulated scenario. The results in the real Internet scenario further confirm the superiority of LEAP. Wanxin Shi, Qing Li 0006, Gengbiao Shen, Weichao Li 0001, Yu Wu 0010, Yong Jiang 0001 |
IWQoS | 5 |
| 2019 | An empirical study of mobile network behavior and application performance in the wildabstractMonitoring mobile network performance is critical for optimizing the QoE of mobile apps. Until now, few studies have considered the actual network performance that mobile apps experience in a per-app or per-server granularity. In this paper, we analyze a two-year-long dataset collected by a crowdsourcing per-app measurement tool to gain new insights into mobile network behavior and application performance. We observe that only a small portion of WiFi networks can work in high-speed mode, and more than one-third of the observed ISPs still have not deployed 4G networks. For cellular networks, the DNS settings on smartphones can have a significant impact on mobile app network performance. Moreover, we notice that instant messaging (IM) and voice over IP (VoIP) services nowadays are not as performant as Web services, because the traffic using XMPP experiences longer latencies than HTTPS. We propose an automatic performance degradation detection and localization method for finding possible network problems in our huge, imbalanced and sparse dataset. Our evaluation and case studies show that our method is effective and the running time is acceptable. Weichao Li 0001, Daoyuan Wu, Bo Jin 0002, Rocky K. C. Chang, Debin Gao, Yi Wang 0004, Ricky K. P. Mok |
IWQoS | 2 |
| 2018 | Measuring the Control-Data Plane Consistency in Software Defined NetworkingabstractSoftware Defined Networking (SDN) simplifies network management by separating the control plane from the data plane in performance networks. However, the actual packet behaviors, conforming to the rules in the data plane flow tables, may violate the original policies in the controller due to the inconsistency between the data plane and control plane. To address this problem, we propose 2MVeri, a new framework for measuring the control-data plane consistency, defined as the consistency between the control plane policies and data plane rules. 2MVeri uses a Bloom filter and a two-dimensional vector as a tag which is inserted in the packet header and is updated in each switch that the packet traverses. By exploiting path information compressed in the tag, 2MVeri can verify the control-data plane consistency. Moreover, when verification fails, 2MVeri can localize the faulty switch. Evaluations conducted on a datacenter network with the fat tree topology k=4 demonstrate that 2MVeri achieves 100% accuracy in consistency verification and high fault localization performance. Kai Lei, Weichao Li 0001, Yi Wang 0004 |
ICC | 4 |
| 2018 | Toward Accurate Network Delay Measurement on Android PhonesabstractMeasuring and understanding the performance of mobile networks is becoming very important for end users and operators. Despite the availability of many measurement apps, their measurement accuracy has not received sufficient scrutiny. In this paper, we appraise the accuracy of smartphone-based network performance measurement using the Android platform and the network round-trip time (RTT) as the metric. We show that two of the most popular measurement apps-Ookla Speedtest and MobiPerf-have their RTT measurements inflated. We build three test apps for three common measurement methods and evaluate them in a testbed. We overcome the main challenge of obtaining a complete trace of packets and their timestamps using multiple sniffers and frame-based synchronization. Our multi-layer analysis reveals that the delay inflation can be introduced both in the user space and kernel space. The long path of subfunction invocations accounts for the majority of the delay overhead in the Android runtime (both Dalvik VM and ART), and the sleeping functions in the drivers are the major source of the delay overhead between the kernel and physical layer. We propose and implement a native measurement app to mitigate the delay overhead in the Android runtime, and the resulted delay inflation in the user space can be kept under 1.5 ms for almost all cases. Weichao Li 0001, Daoyuan Wu, Rocky K. C. Chang, Ricky K. P. Mok |
IEEE Trans. Mob. Comput. | 1 |
| 2017 | MopEye: Opportunistic Monitoring of Per-app Mobile Network Performance
Daoyuan Wu, Rocky K. C. Chang, Weichao Li 0001, Eric K. T. Cheng, Debin Gao |
USENIX ATC | 3 |
| 2017 | Detecting Low-Quality Workers in QoE Crowdtesting: A Worker Behavior-Based ApproachabstractQoE crowdtesting is increasingly popular among researchers to conduct subjective assessments of network services. Experimenters can easily access a huge pool of human subjects through crowdsourcing platforms. Without any supervision, low-quality workers, however, can threaten the reliability of the assessments. One of the approaches in classifying the quality of workers is to analyze their behavior during the experiments, such as mouse cursor trajectory. However, existing works analyze the trajectory coarsely, which cannot fully extract the imbedded information. In this paper, we propose a novel method to detect low-quality workers in QoE crowdtesting by analyzing the worker behavior. Our approach is to construct a predictive model by using supervised learning algorithms. A quality score is computed by applying existing anti-cheating techniques and human inspections to label the workers. We define a set of ten worker behavior metrics, which quantifies different types of worker behavior, including finer-grained cursor trajectory analysis. A multiclass Naïve Bayes classifier is applied to train a model to predict the quality of workers from the metrics. We have conducted video QoE assessments on Amazon Mechanical Turk and CrowdFlower to collect the worker behavior. Our results show that the error rates of the model trained from four metrics are equal or less than 30%. We further find that combining the predictions from the four different 5-point Likert scale rating methods can improve the success rate in detecting low-quality workers to around 80%. Finally, our method is 16.5% and 42.9% better in precision and recall than CrowdMOS. Ricky K. P. Mok, Rocky K. C. Chang, Weichao Li 0001 |
IEEE Trans. Multim. | 3 |
| 2016 | Demystifying and Puncturing the Inflated Delay in Smartphone-based WiFi Network MeasurementabstractUsing network measurement apps has become a very effective approach to crowdsourcing WiFi network performance data. However, these apps usually measure the user-level performance metrics instead of the network-level performance which is important for diagnosing performance problems. In this paper we report for the first time that a major source of measurement noises comes from the periodical SDIO (Secure Digital Input Output) bus sleep inside the phone. The additional latency introduced by SDIO and Power Saving Mode can inflate and unstablize network delay measurement significantly. We carefully design and implement a scheme to wake up the phone for delay measurement by sending just enough warm-up and background traffic. Our evaluation results show that the overall median delay overheads can be kept within 3ms, regardless of the actual network delay. Weichao Li 0001, Daoyuan Wu, Rocky K. C. Chang, Ricky K. P. Mok |
CoNEXT | 1 |
| 2016 | IRate: Initial Video Bitrate Selection System for HTTP StreamingabstractMany HTTP streaming video systems have been developed and widely deployed in recent years. Previous efforts were mainly spent on improving the caching of videos or proposing mid-stream measurement methods to update the best bitrate. However, since the video length is often short, the mid-stream measurement may not even converge to the best bitrate due to insufficient bandwidth estimates. On the other hand, because of diversified Web infrastructure, estimating the actual network quality at the pre-stream stage is increasingly challenging for video service providers. In this paper, we propose IRate, which enables video service providers to proactively profile clients' streaming performance by carrying out pre-stream measurement in the Content Delivery Network (CDN). With the measurement results, the video stream can start at the best video quality at the onset of streaming. This is especially beneficial to short video clips, which are very popular in the Internet today. IRate is composed of a probe kit and a quality oracle. The probe kit utilizes the pre-stream time window (e.g., user's think time and pre-roll advertisement) for measuring network quality by running a lightweight measurement script on the Web page to induce probe packets from the IRate middlebox on the server side. With the measurement results, the quality oracle estimates the clients' streaming performance by determining the highest initial bitrate with a pre-trained decision tree. Our testbed results show that IRate is able to achieve 80% accuracy in determining the bitrate within 10s. By having a better estimate of the best initial bitrate, the buffering time and rebuffering events are significantly reduced in HTTP streaming. Furthermore, the stability and the efficiency in dynamic adaptive streaming over HTTP streaming are also improved by about 40% and 36%, respectively. Our user quality of experience (QoE) experiment further validates that IRate can improve the QoE by more than 6% and the perceived quality of initial quality by 24% in the actual Internet environment. Ricky K. P. Mok, Weichao Li 0001, Rocky K. C. Chang |
IEEE J. Sel. Areas Commun. | 2 |
| 2015 | On the accuracy of smartphone-based mobile network measurementabstractAs most of mobile apps rely on network connections for their operations, measuring and understanding the performance of mobile networks is becoming very important for end users and operators. Despite the availability of many measurement apps, their measurement accuracy has not received sufficient scrutiny. In this paper, we appraise the accuracy of smartphone-based network performance measurement using the Android platform and the network round-trip time as the metric. We use a multiple-sniffer testbed to overcome the challenge of obtaining a complete trace for acquiring the required timestamps. Our experiment results show that the RTTs measured by the apps are all inflated, ranging from a few milliseconds (ms) to tens of milliseconds. Moreover, the 95% confidence interval can be as high as 2.4ms. A finer-grained analysis reveals that the delay inflation can be introduced both in the Dalvik VM (DVM) and below the Linux kernel. The in-DVM overhead can be mitigated but the other cannot be. Finally, we propose and implement a native app which uses HTTP messages for network measurement, and the delay inflation can be kept under 5ms for almost all cases. Weichao Li 0001, Ricky K. P. Mok, Daoyuan Wu, Rocky K. C. Chang |
INFOCOM | 1 |
| 2015 | Detecting low-quality crowdtesting workersabstractQoE crowdtesting is increasingly popular among researchers to conduct subjective assessments of different services. Experimenters can easily access to a huge pool of human subjects through crowdsourcing platforms. A fundamental problem threatening the integrity of crowdtesting is to detect cheating from the workers who work without any supervision. One of the approaches in classifying the quality of workers is analyzing their behavior during the experiments. A major challenge is to systematically analyze the mouse cursor trajectory. However, existing works usually analyze the trajectory coarsely, which cannot fully extract the information imbedded in the trajectory. In this paper, we propose to use finer-grained cursor trajectory analysis, including submovement analysis, to identify low quality workers. Our approach is to define a set of ten worker behavior metrics to quantify different types of worker behavior. A jQuery-based library was implemented to collect the worker behavior. Moreover, four different 5-point Likert scale rating methods were employed. A number of methods, including question design, instructions, and human inspections, are used to label workers into three categories. We then apply multiclass Naive Bayes classifier to construct different models using all or some of the metrics and the workers' category. Our results show that the error rates of the model trained from four metrics is equal or less than 30% for four rating methods. By combining the predictions from the four rating methods, the successful rate in detecting low-quality workers is around 80%. Ricky K. P. Mok, Weichao Li 0001, Rocky K. C. Chang |
IWQoS | 2 |
| 2015 | Improving the Packet Send-Time Accuracy in Embedded Devices
Ricky K. P. Mok, Weichao Li 0001, Rocky K. C. Chang |
PAM | 2 |
| 2014 | A user behavior based cheat detection mechanism for crowdtestingabstractCrowdtesting is increasingly popular among researchers to carry out subjective assessments of different services. Experimenters can easily assess to a huge pool of human subjects through crowdsourcing platforms. The workers are usually anonymous, and they participate in the experiments independently. Therefore, a fundamental problem threatening the integrity of these platforms is to detect various types of cheating from the workers. In this poster, we propose cheat-detection mechanism based on an analysis of the workers' mouse cursor trajectories. It provides a jQuery-based library to record browser events. We compute a set of metrics from the cursor traces to identify cheaters. We deploy our mechanism to the survey pages for our video quality assessment tasks published on Amazon Mechanical Turk. Our results show that cheaters' cursor movement is usually more direct and contains less pauses. Ricky K. P. Mok, Weichao Li 0001, Rocky K. C. Chang |
SIGCOMM | 2 |
| 2013 | MonoScope: Automating network faults diagnosis based on active measurements
Waiting W. T. Fok, Xiapu Luo, Ricky K. P. Mok, Weichao Li 0001, Edmond W. W. Chan, Rocky K. C. Chang |
IM | 4 |
| 2013 | Appraising the delay accuracy in browser-based network measurementabstractConducting network measurement in a web browser (e.g., speedtest and Netalyzr) enables end users to understand their network and application performance. However, very little is known about the (in)accuracy of the various methods used in these tools. In this paper, we evaluate the accuracy of ten HTTP-based and TCP socket-based methods for measuring the round-trip time (RTT) with the five most popular browsers on Linux and Windows. Our measurement results show that the delay overheads incurred in most of the HTTP-based methods are too large to ignore. Moreover, the overheads incurred by some methods (such as Flash GET and POST) vary significantly across different browsers and systems, making it very difficult to calibrate. The socket-based methods, on the other hand, incur much smaller overhead.Another interesting and important finding is that Date.getTime(), a typical timing API in Java, does not provide the millisecond resolution assumed by many measurement tools on some OSes (e.g., Windows 7). This results in a serious under-estimation of RTT. On the other hand, some tools over-estimate the RTT by including the TCP handshaking phase. Weichao Li 0001, Ricky K. P. Mok, Rocky K. C. Chang, Waiting W. T. Fok |
Internet Measurement Conference | 1 |
| 2011 | TRIO: measuring asymmetric capacity with three minimum round-trip timesabstractMeasuring network path capacity is an important capability to many Internet applications. But despite over ten years of effort, the capacity measurement problem is far from being completely solved. This paper addresses the problem of measuring network paths of asymmetric capacity without requiring the remote node's control or overwhelming the bottleneck link. We first show through analysis and measurement that the current packet-dispersion methods, due to the packet size limitations, can only measure up to a certain degree of capacity asymmetry. Second, we propose TRIO that removes the limitation by using round-trip times (RTTs). TRIO cleverly exploits two types of probes to obtain three minimum RTTs to compute bothforward and reverse capacities, and another minimum RTT for measurement validation. We validate TRIO's accuracy and versatility on a testbed and the Internet, and develop a system to measure path capacity from the server or user side. Edmond W. W. Chan, Ang Chen 0001, Xiapu Luo, Ricky K. P. Mok, Weichao Li 0001, Rocky K. C. Chang |
CoNEXT | 5 |
| 2011 | Planetopus: A system for facilitating collaborative network monitoringabstractMany new methods and tools have been developed to measure the quality of network paths for the last ten years. However, there are relatively few works that consider deploying these methods for collaborative network measurement: a number of measuring points belonging to different autonomous systems collaborate on monitoring and diagnosing their network performance. In this paper, we present Planetopus, a distributed system for facilitating collaborative network monitoring. Planetopus provides a single platform for configuring and scheduling measurement tasks performed on a set of distributed measuring points. Planetopus currently performs measurement mainly using OneProbe and tcptraceroute. Moreover, we introduce two useful facilities for analyzing the measurement data: a new metric for quantifying route changes and a heatmap-based visualization method for discovering patterns and anomalies from a set of path measurements. We demonstrate the utility of Planetopus through several case studies in which poor routes are identified and corrected, different ISPs' network services are compared, and network performance problems are diagnosed. Weichao Li 0001, Waiting W. T. Fok, Edmond W. W. Chan, Xiapu Luo, Rocky K. C. Chang |
Integrated Network Management | 1 |
| 2011 | Non-cooperative Diagnosis of Submarine Cable Faults
Edmond W. W. Chan, Xiapu Luo, Waiting W. T. Fok, Weichao Li 0001, Rocky K. C. Chang |
PAM | 4 |
| 2010 | Neighbor-Cooperative Measurement of Network Path QualityabstractIn the current Internet landscape, a stub autonomous system (AS) could choose from a number of providers and peers to advertise its routes. However, the route selection may not always result in a best choice in terms of end-to-end path performance. Instead of having an AS to monitor all possible paths, we argue that it is much more effective and beneficial for a number of neighboring ASes to cooperate in the path measurement. In this paper, we present a neighbor-cooperative measurement system in which each participating AS conducts measurement using their current routes for the same set of remote endpoints. A collation of the measurement results can help identify and correct poor routes, compare different providers' network services, and diagnose network performance problems. We report measurement results from an actual deployment involving eight neighboring universities for over a year. Rocky K. C. Chang, Waiting W. T. Fok, Weichao Li 0001, Edmond W. W. Chan, Xiapu Luo |
GLOBECOM | 3 |
| 2010 | Measurement of loss pairs in network pathsabstractLoss-pair measurement was proposed a decade ago for discovering network path properties, such as a router's buffer size. A packet pair is regarded as a loss pair if exactly one packet is lost. Therefore, the residual packet's delay can be used to infer the lost packet's delay. Despite this unique advantage shared by no other methods, no loss-pair measurement in actual networks has ever been reported. In this paper, we further develop the loss-pair measurement and make the following contributions. First, we characterize the residual packet's delay by including other important factors (such as the impact of the first packet in the pair) which were ignored before. Second, we employ a novel TCP-based probing method to measure from a single endpoint all four possible loss pairs for a round-trip network path. Third, we conducted loss-pair measurement for 88 round-trip paths continuously for almost three weeks. Being the first set of loss-pair measurement, we obtained a number of original results, such as prevalence of loss pairs, distribution of different types of loss pairs, and effect of route change on the paths' congestion state. Edmond W. W. Chan, Xiapu Luo, Weichao Li 0001, Waiting W. T. Fok, Rocky K. C. Chang |
Internet Measurement Conference | 3 |