VLDB 2026 Research / reviewers in the wild / expert
Zili Meng
dblp:198/8304
· DBLP profile ↗
50ranked-venue papers
9as first author
39since 2021 · last 2026
0000-0003-2009-7180ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 39 · 8 first-author · 29 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 4 since 2021Systems, architecture and hardware · 4 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DeNC++: Efficient Diffusion-Enhanced Neural Codec for End-to-end Semantic Streaming at the EdgeabstractThe neural-enhanced video streaming (NeVS) has been an emerging technique to integrate neural models into video codecs for higher streaming efficiency. The state-of-the-art methods, e.g., DeNC and Gemino, typically compress videos in RGB space and restore video quality via a neural enhancement model hosted on the external media server. However, these methods are not always accessible in resource-constrained edge environments due to their heavy reliance on the media server's computation, which undermines end-to-end performance and restricts NeVS's usage boundary. This limitation raises an interesting question: is it possible to make NeVS lightweight so that all neural codec operations can be handled directly by clients' edge devices? In this paper, we present the answer yes and develop a new plug-and-play module called DeNC++, which significantly improves the compression-restoration-overhead trade-off over existing methods. Our core design philosophy is to wrap all the codec operations within a latent semantic space, in which the original high-dimensional visual signals are efficiently embedded into low-dimensional semantic representations. With this fundamental transformation, DeNC++'s neural encoder introduces the triple semantic-bitwidth-resolution compression to effectively lower the streaming traffic. Meanwhile, we make DeNC++'s neural decoder aware of the perceptual loss caused by its encoder and design tiny generative models to guarantee high restoration quality. We also strictly restrict the runtime computational overhead and accelerate the neural enhancement process, making DeNC++ compatible with commodity edge devices. Real-world evaluations reveal that DeNC++ consistently provides higher restoration quality while achieving 24-55 times higher compression ratio and 5-7 times end-to-end speedup over the latest NeVS solutions. Qihua Zhou, Wangjiang Gong, Zili Meng, Yaxiong Xie, Yaodong Huang, Junchen Jiang, Laizhong Cui |
AAAI | 3 |
| 2026 | ForeQast: Structuring Forward Looking Delay Guarantees in QUICabstractFor many real-time applications, such as video conferencing, the end-to-end latency is a critical metric to the user experience. While the latency over the Internet has been significantly reduced in recent days, we find that the delay at the sender has been increasingly critical and dominant. We identify that the delay comes from the socket buffer at the sender. Specifically, the gap between when the application writes data and when packets are actually transmitted. The existing socket buffer lacks active management and is not optimized for low latency. To solve this, we propose a predictive thresholding approach, ForeQast, that integrates active queue management directly at the end host. We introduce an adaptive buffer threshold to actively manage the socket send buffer to reduce the latency. We carefully design the threshold to maintain throughput in different network conditions. We implement the ForeQast over QUIC protocol and evaluate it using trace-driven emulation. Evaluations on synthetic and real network traces show that the algorithm meets percentile-based delay guarantees across traces, targets, and RTTs, without throughput loss or noticeable CPU overhead, delivering a controllable delay for the socket send buffer. Shao Fu Wang, Ka Chun Mok, Zili Meng |
ICC | 3 |
| 2026 | Hermit: A Flow-Collaborative Transport Scheme for Multi-Source Video On-Demand StreamingabstractToday's fast-growing Video-on-Demand (VoD) service needs efficient content delivery to guarantee the user experience. To reduce costs, the industry has been exploring the adoption of unstable, heterogeneous, low-performance edge nodes as cost-efficient alternatives to expensive CDN servers. To compensate for the resulting degradation in user experience, Multi-source Parallel Downloading (MPD) is becoming a new VoD transport paradigm. However, existing transport optimization solutions face performance obstacles when applied to the MPD scenarios. They cannot handle the contention between MPD flows of the same download task, which is likely to occur at the shared last-hop, and lack the ability to quickly adapt to the unstable network environments brought by dynamic, heterogeneous, and low-performance edge nodes. To fill this gap, we propose Hermit, a VoD-oriented MPD transport algorithm. Hermit (1) continuously monitors the state of the flows and makes timely scheduling decisions, and (2) efficiently coordinates across the flows to mitigate self-contention at the shared last hop. As a client-driven scheme, Hermit does not require cumbersome coordination among edge nodes, nor does it increase server complexity. Through extensive experiments on real-world large-scale testbed and locally emulated network conditions, we demonstrate that Hermit can improve the consistent downloading rate by 9.2% to 21.3%. Shaorui Ren, Enhuan Dong, Haiping Wang 0002, Jia Zhang 0010, Zili Meng, Mingwei Xu 0001, Shu Shi, Hebin Yu, Zhichen Xue, Yajie Peng, Xiaofei Pang |
ICC | 6 |
| 2026 | Confucius: Adapting Home Routers to Congestion Control's Reactions for Consistent Low LatencyabstractEmerging high-quality real-time applications require consistently low latency, which is often disrupted by latency spikes. We identify the reason as the mismatch between the abrupt bandwidth reallocation on routers and gradual sending rate reaction of congestion control. For example, when a burst of new flows arrives, queue schedulers such as fair queueing immediately reallocate the bandwidth for existing and new flows. However, the flow's sending rate, determined by the congestion control algorithm (CCA), needs several RTTs to converge to the new available bandwidth, during which severe stalls occur. This has been increasingly critical with the demand on consistent low latency. In this paper, we present Confucius, a practical queue management scheme that reallocate the bandwidth for flows following CCA's reaction. Confucius slows down bandwidth adjustment to match the reaction of congestion control, so that the end host can reduce the sending rate without overshooting the network. Confucius is designed for offering real-time flows with consistently low latency regardless of uncertain competition. Experiments show that Confucius reduces the stall duration by more than 50% against existing practical schemes, while competing flows also fairly enjoy on-par performance.Available at: https://github.com/hkust-spark/confucius-qdisc Zili Meng, Nirav Atre, Bochun Zhang, Mingwei Xu 0001, Justine Sherry, Maria Apostolaki |
INFOCOM | 1 |
| 2026 | Real-Time Video Gets a Fast Lane via Smart Queue Flushing at the Wireless EdgeabstractReal-time communication (RTC) applications demand not only low average latency but also tight tail-delay bounds to ensure smooth user experience. However, sudden fluctuations in wireless networks can cause in-flight packets to accumulate in bottleneck queues, delaying or invalidating subsequent frames. Traditional mechanisms focus on rate adaptation but largely overlook managing already enqueued packets that contribute to tail delay. We present Gecko, a lightweight end-to-network coordination mechanism that enables frame-aware queue flushing without requiring any in-network packet modification or protocol negotiation. Gecko-enabled routers monitor queuing delay and implicitly signal the sender, which then makes frame-skipping decisions and conveys flushing intent through minimal in-band RTP markings. This approach preserves end-to-end integrity and is broadly compatible with existing RTC applications. We evaluate Gecko via trace-driven simulations and real-world experiments. Results show that Gecko reduces average frame delay by 21.8% and cuts tail-delay frame ratios by 25% to 91%, demonstrating both its effectiveness and deployability in wireless RTC environments. Zili Meng, Enhuan Dong, Yan Zhang 0002, Jia Zhang 0010, Mingwei Xu 0001 |
INFOCOM | 2 |
| 2026 | Make a Video Call with LLM: A Measurement Campaign over Five Mainstream AppsabstractIn 2025, Large Language Model (LLM) services have launched a new feature-AI video chat-allowing users to interact with AI agents via real-Time video communication (RTC), just like chatting with real people. Despite its significance, no systematic study has characterized the performance of existing AI video chat systems. To address this gap, this paper proposes a comprehensive benchmark with carefully designed metrics across three dimensions: quality, latency, and internal mechanisms. Using custom testbeds, we further evaluate five mainstream AI video chatbots with this benchmark. This work provides the research community a baseline of real-world performance and identifies unique system bottlenecks. In the meantime, our benchmarking results also open up several research questions for future optimizations of AI video chatbots. Xiangjie Huang, Zili Meng |
NOSSDAV | 4 |
| 2026 | MAE: More Adaptive Video Encoder for Consistent Low Latency in High-Quality Real-Time Communication
Yufan Zhuang, Yasna Noushirvani, Xiangjie Huang, Zili Meng |
NSDI | 5 |
| 2026 | A Composable Emulation Framework for Whitebox Switches
Congcong Miao, Xianneng Zou, Chuwen Zhang, Qihang Liu, Zhijie Yan, Yanke Zhang, Yong Jiang 0001, Qiao Xiang, Xin Jin 0008, Zili Meng, Ang Chen 0001 |
NSDI | 11 |
| 2026 | Mortise: Auto-tuning Congestion Control to Optimize QoE via Network-Aware Parameter Optimization
Yixin Shen 0002, Ruihua Chen, Bo Wang 0066, Minhu Wang, Mingwei Xu 0001, Zili Meng |
NSDI | 8 |
| 2026 | Law: Towards Consistent Low Latency in 802.11 Home Networks
Yibin Shen, Zili Meng |
NSDI | 2 |
| 2026 | Compass: Congestion Control for Disobedient TrafficabstractIncreasingly stringent service-level objectives demand fast and accurate congestion control (CC) in data center networks. We observe that the feedback loop used by existing data center sender-driven CC schemes is inherently limited as it suffers from an inevitable delay. Consequently, a portion of traffic, which we call disobedient traffic, could finish before the feedback loop, thereby escaping the control of these schemes and exerting a negative impact on their performance. Furthermore, as link speeds continue to climb, the proportion and impact of disobedient traffic are concurrently escalating, exacerbating these issues. In this paper, we propose Compass, a solution implemented entirely on switch data plane to mitigate the impact of disobedient traffic. Compass utilizes sketching techniques to efficiently estimate the sending rate of disobedient traffic, and seamlessly integrates into HPCC and PowerTCP through incorporating the sending rate of disobedient traffic into their high-precision rate control algorithms. Such integration enhances network performance without incurring additional bandwidth overhead or modifications to host-side logic. Additionally, Compass supports incremental brownfield deployment, and all these properties make it highly practical for production deployment. Extensive simulations show that Compass improves both throughput and latency. For instance, Compass reduces tail flow completion times of medium and large flows by up to 35% for PowerTCP and 17% for HPCC. Kaicheng Yang 0001, Tianbao Zhou, Hengyang Zhou, Kaitai Zhang, Yikai Zhao 0001, Yuanpeng Li 0002, Yuhan Wu 0001, Zili Meng, Fengyuan Ren, Tong Yang 0003 |
IEEE Trans. Netw. | 8 |
| 2025 | INDS: Incremental Named Data Streaming for Real-Time Point Cloud VideoabstractReal-time streaming of point cloud video, characterized by massive data volumes and high sensitivity to packet loss, remains a key challenge for immersive applications under dynamic network conditions. While connection-oriented protocols such as TCP and more modern alternatives like QUIC alleviate some transport-layer inefficiencies, including head-of-line blocking, they still retain a coarse-grained, segment-based delivery model and a centralized control loop that limit fine-grained adaptation and effective caching. We introduce INDS (Incremental Named Data Streaming), an adaptive streaming framework based on Information-Centric Networking (ICN) that rethinks delivery for hierarchical, layered media. INDS leverages the Octree structure of point cloud video and expressive content naming to support progressive, partial retrieval of enhancement layers based on consumer bandwidth and decoding capability. By combining time-windows with Group-of-Frames (GoF), INDS's naming scheme supports fine-grained in-network caching and facilitates efficient multi-user data reuse. INDS can be deployed as an overlay, remaining compatible with QUIC-based transport infrastructure as well as future Media-over-QUIC (MoQ) architectures, without requiring changes to underlying IP networks. Our prototype implementation shows up to 80% lower delay, 15-50% higher throughput, and 20-30% increased cache hit rates compared to state-of-the-art DASH-style systems. Together, these results establish INDS as a scalable, cache-friendly solution for real-time point cloud streaming under variable and lossy conditions, while its compatibility with MoQ overlays further positions it as a practical, forward-compatible architecture for emerging immersive media systems. Ruonan Chai, Yixiang Zhu, Xinjiao Li, Jiawei Li 0009, Zili Meng, Dirk Kutscher |
ACM Multimedia | 5 |
| 2025 | Configuring Dynamic Multi-Stage Serverless Pipelines for Video Processing with Minimal Profiling OverheadabstractServerless computing has become a promising paradigm for video processing workflows, offering simplified deployment and flexible management of business logic. However, the dynamic, multi-stage nature of video processing pipelines poses significant challenges for traditional serverless resource management, particularly in efficiently modeling optimal configurations and adapting to rapidly evolving pipeline structures. To address this challenge, we propose ConfigNavigator, a video pipeline resource tuning framework capable of adapting to dynamic inputs and pipeline structures with minimal overhead. In the offline phase, ConfigNavigator models function execution time distributions at the fundamental operation level and leverages graph theory to decompose complex video processing pipelines, thereby obtaining optimal configurations with minimal overhead. In the online phase, it dynamically adjusts function configurations on critical paths through real-time performance feedback, ensuring pipeline performance stability across varying workloads. We evaluate ConfigNavigator using real video streams on the commercial serverless platform AWS Lambda. Compared to state-of-the-art baselines, ConfigNavigator reduces configuration search time by 94.11% while decreasing end-to-end pipeline processing time by 13.97%. Jiaye Zhang, Hongyi Wang 0009, Peiru Yang, Zili Meng, Mingwei Xu 0001 |
ACM Multimedia | 4 |
| 2025 | Active Management of Jammed Packets in Wireless Real-Time CommunicationsabstractToday's real-time communication (RTC) application requires consistent low latency to ensure the user experience. Many end-to-end rate control, as well as in-network active queue management (AQM) methods, have been designed to improve transport latency. However, most of previous work can only address the network issues in a reactive way - packets during the reaction time are still stuck on the way. No matter how fast the sender reacts to network changes in existing schemes, there will still be packets jammed in the network and increases the latency. To enhance the transport performance in wireless network and improve the user experience of RTC application, we propose Gecko, a practical application-oriented end-to-network collaboration scheme. Gecko can efficiently detect and process the congestion signal, early and proactively draining the jammed packets in the bottleneck queue. We conduct both trace-driven simulations and real-world experiments to evaluate the performance of our scheme. Gecko can reduce the overall frame delay by 21.8 % and reduce tail-delay frame ratio by 25% to 91% in our experiments. Zili Meng, Enhuan Dong, Yan Zhang 0002, Mingwei Xu 0001 |
NOSSDAV | 2 |
| 2025 | Tooth: Toward Optimal Balance of Video QoE and Redundancy Cost by Fine-Grained FEC in Cloud Gaming Streaming
Congkai An, Jingyang Kang, Anfu Zhou, Liang Liu 0001, Huadong Ma, Zili Meng, Delei Ma, Yusheng Dong, Xiaogang Lei |
NSDI | 8 |
| 2025 | ACE: Sending Burstiness Control for High-Quality Real-time CommunicationabstractModern real-time communication (RTC) demands both ultra-low latency and consistently high visual quality. Yet, as content becomes more dynamic and RTTs shrink, we reveal a previously overlooked problem: long-tail queuing latency in the sender's pacing queue between encoder and network. This phenomenon is rooted in a mismatch between the bursty frame stream produced by the encoder and the smooth traffic expected by the network. Existing approaches trying to smoothen the bitrate inevitably force an undesirable trade-off between latency and video quality. To address this, we propose a dual-control approach that manages both the encoding and transmission burstiness. At the sender, we dynamically adjust the bucket size of a token-based pacer to control burstiness at the granularity of frame level. Within the encoder, we introduce an adaptive complexity mechanism that smoothens frame sizes without sacrificing quality. Trace-driven emulation and real-world experiments show our solution ACE reduces end-to-end 95th percentile latency by up to 43% while maintaining superior visual quality versus the state of the art. Xiangjie Huang, Haiping Wang 0002, Hebin Yu, Sandesh Dhawaskar Sathyanarayana, Shu Shi, Zili Meng |
SIGCOMM | 7 |
| 2025 | Learning-Enhanced High-Throughput Pattern Matching Based on Programmable Data Plane
Guanglin Duan, Qing Li 0006, Dan Zhao 0003, Zili Meng, Dirk Kutscher, Ruoyu Li 0003, Yong Jiang 0001, Mingwei Xu 0001 |
USENIX ATC | 6 |
| 2025 | Zhuge: Toward Consistent Low Latency With Minimal Control Loop DelayabstractReal-time communication (RTC) applications demand consistent low latency to ensure a smooth and interactive user experience. However, wireless networks, including WiFi and cellular, although they provide satisfactory median latency, often suffer from significant tail latency due to the highly variable network bandwidth. We observe that the control loop for managing the sending rate of RTC applications becomes inflated when congestion occurs at the wireless access point (AP), leading to untimely rate adaptation in response to wireless dynamics. Existing solutions fail to quickly adapt to bandwidth fluctuations due to the inflated control loop. In this paper, we propose Zhuge, a purely wireless AP-based solution that addresses these issues by separating congestion feedback from congested queues. Our approach involves the design of a Fortune Teller, which accurately estimates the wireless latency for each packet upon its arrival at the wireless AP. To ensure scalability, we also develop a Feedback Updater that translates the estimated latency into understandable feedback messages for various end-to-end protocols, delivering them back to the senders immediately for rate adaptation. Our evaluation, based on both trace-driven simulations and real-world scenarios, demonstrates that Zhuge significantly reduces the occurrence of large tail latency and alleviates RTC performance degradation by 22% to 95%. Bo Wang 0066, Xingxing Yang 0008, Zili Meng, Yaning Guo, Chen Sun 0005, Justine Sherry, Hongqiang Harry Liu, Mingwei Xu 0001 |
IEEE Trans. Netw. | 3 |
| 2024 | Inferring in-Network Queue Management from End Hosts in Real-Time CommunicationsabstractActive queue management (AQM) algorithms, widely deployed in the internet, are designed to signal end hosts with network conditions in the format of packet losses. However, real-time communication (RTC) applications adopt delay-sensitive congestion control algorithms (CCAs), which are no longer responsive to losses or explicit notifications from AQMs. Moreover, packet losses introduced by different AQMs will further degrade the performance of RTC applications due to unexpected and unnecessary loss recovery. We are therefore motivated to understand the behaviors of AQMs and take necessary countermeasures for RTC applications proactively. For example, with the help of AQM inference, RTC applications will benefit by using loss recovery mechanisms that adapt to various kinds of AQMs to deal with packet losses. However, it is challenging to infer the AQM from end hosts since numerous AQMs have different configurations after decades of evolution. We analyze the temporal behaviors of loss series, extract the inherent invariant features of different AQMs, and categorize them into three types. Our simulation shows that AQM inference can classify AQMs with an accuracy of 96%. We also evaluate a use case on using the AQM inference to improve the loss recovery mechanism (forward error correction, FEC). Our FEC method based on AQM inference improves the recovery rate by at least 56%, and finally reduces the end-to-end delay by 13%. Yaning Guo, Zili Meng, Bo Wang 0066, Mingwei Xu 0001 |
ICC | 2 |
| 2024 | Beimin: Serverless-based Adaptive Real-Time Video ProcessingabstractVideo-sharing websites need to process the uploaded videos (e.g., face recognition) before distributing them to users. The timely processing of videos is critical for users to always enjoy the latest content. However, videos uploaded by different users are diverse in content, with the volume fluctuating at different times in one day. The static resource allocations will result in frequent overutilization and underutilization when the demands and contents change, while container and virtual machine(VM)-based solutions will incur significant additional overhead. Moreover, it is also challenging to predict the required resources in the future due to the complicated relationship between resources, contents, demands, etc. This paper introduces Beimin, an adaptive video processing framework designed for heterogeneous video processing with flexible demands in real-time. Beimin adopts a serverless framework to efficiently allocate resources, and a deep reinforcement learning (DRL) model to predict the resources to allocate with multi-dimensional inputs (contents, demands, etc.). We conducted tests with Amazon Lambda using a synthetic dataset from Imagenet VID, and the results demonstrate that Beimin reduces cost by 1.61% and processing time by 33.17% compared to existing solutions without harm to accuracy. Jiaye Zhang, Zili Meng, Mingwei Xu 0001 |
ICME | 2 |
| 2024 | Bidirectional Bandwidth Coordination Under Half-Duplex Bottlenecks for Video StreamingabstractMany video streaming applications will simultaneously transfer data in both directions, from the user to the Internet (uplink) and from the Internet to users (downlink). However, for wireless local area networks (WLANs), the dominant scenarios, the uplink and downlink flows share the same half-duplex physical channel and compete for bandwidth resources. Their bandwidths would be fairly apportioned under the existing link layer access method, but a fair share might be suboptimal for applications. For better application performance, we propose Plum, to coordinate the bitrate of uplink and downlink flows, and allocate the bandwidth in both directions to cater to the application's demands. To make the deployment of Plum practical, we aim at not modifying the link layer but optimizing the transport layers and above. We evaluate our mechanisms with simulations based on real-world traces and testbed experiments, and results show that Plum could improve the video bitrate of streaming applications by up to 48-59%. Bo Wang 0066, Yan Zhang 0002, Minhu Wang, Mingwei Xu 0001, Zili Meng |
ICNP | 7 |
| 2024 | Hairpin: Rethinking Packet Loss Recovery in Edge-based Interactive Video Streaming
Zili Meng, Bo Wang 0066, Mingwei Xu 0001, Venkat Arun, Hongxin Hu |
NSDI | 1 |
| 2024 | Prudentia: Findings of an Internet Fairness WatchdogabstractWith the rise of heterogeneous congestion control algorithms and increasingly complex application control loops (e.g. adaptive bitrate algorithms), the Internet community has expressed growing concern that network bandwidth allocations are unfairly skewed, and that some Internet services are 'winners' at the expense of 'losing' services when competing over shared bottlenecks. In this paper, we provide the first study of fairness between live, end-to-end services with distinct workloads. Rather than focusing on individual components of an application stack (e.g., studying the fairness of an individual congestion control algorithm), we want to provide a direct study over real-world deployed applications. Among our findings, we observe that services typically achieve less-than-fair outcomes: on average, the 'losing' service achieves only 72% of its max-min fair share of link bandwidth. We also find that some services are significantly more contentious than others: for example, one popular file distribution service causes competing applications to obtain as low as 16% of their max-min fair share of bandwidth when competing in a moderately-constrained setting. Adithya Abraham Philip, Rukshani Athapathu, Ranysha Ware, Fabian Francis Mkocheko, Alexis Schlomer, Mengrou Shou, Zili Meng, Srinivasan Seshan, Justine Sherry |
SIGCOMM | 7 |
| 2024 | KEPC-Push: A Knowledge-Enhanced Proactive Content Push Strategy for Edge-Assisted Video Feed Streaming
Ziwen Ye, Qing Li 0006, Chunyu Qiao, Xiaoteng Ma, Yong Jiang 0001, Shengbin Meng, Zhenhui Yuan, Zili Meng |
USENIX ATC | 9 |
| 2024 | Cold Start or Hot Start? Robust Slow Start in Congestion Control with A Priori Knowledge for Mobile Web ServicesabstractMobile web services value a quick loading of contents in the first page, which is quantified by the above-the-fold time of the first page (first AFT) and is likely to fall into the slow start phase in congestion control. However, the widely deployed slow start mechanism is "cold start", which manually hardcodes the parameters and is not suitable for the first AFT of heterogeneous mobile web services. We revisit the slow start mechanism and find that it could be optimized with a priori knowledge. However, blindly relying on a priori knowledge is not robust enough to handle the fluctuating mobile networks and unpredictable application traffic. In this paper, we propose WiseStart, a "hot-start-based" slow start mechanism. WiseStart utilizes the priori knowledge to set the initial parameters, continuously probes the new connection to handle the fluctuating network conditions, and carefully adapts to the application-limit scenarios. We implement WiseStart in a popular mobile web service online in production. Comprehensive experiments demonstrate that WiseStart reduces the First AFT by 25.43% and the average RCT at connection establishment by 16.15% compared to the default slow start mechanism and other state-of-the-art baselines. Jia Zhang 0010, Haixuan Tong, Enhuan Dong, Mingwei Xu 0001, Zili Meng |
WWW | 7 |
| 2024 | RoLL+: Real-Time and Accurate Route Leak Locating With AS Triplet Features at ScaleabstractBorder Gateway Protocol (BGP) is the only inter-domain routing protocol that plays an important role on the Internet. However, BGP suffers from route leaks, which can cause serious security threats. To mitigate the effects of route leaks, accurate and timely route leak locating is of great importance. Prior studies leverage AS business relationships to locate route leaks in real time. However, they fail to achieve high locating accuracy. Recent studies apply machine learning to accurately detect route leaks from statistical features of massive BGP messages. Nevertheless, they have high detection latency and cannot further locate route leaks. In this paper, we propose a real-time and accurate route leak locating system named RoLL+. It leverages distinctive AS triplet features to accurately locate AS triplets with route leaks from each BGP message in real time. Considering that RoLL+ may receive a substantial volume of BGP update messages per second, we integrate a cache-like design and a lazy update mechanism into the system to effectively identify route leaks at scale. Our experimental results on real-world BGP route leak data demonstrate that it can achieve 92% locating accuracy with less than 1 ms locating latency. Furthermore, the results show that RoLL+ can process over 7,000 AS triplets per second, meeting real-world throughput requirements. Jiahao Cao 0001, Zili Meng, Renjie Xie, Qi Li 0002, Yuan Yang 0001, Mingwei Xu 0001 |
IEEE/ACM Trans. Netw. | 3 |
| 2023 | RoLL: Real-Time and Accurate Route Leak Location with AS Triplet FeaturesabstractBGP is the only inter-domain routing protocol that plays an important role on the Internet. However, BGP suffers from route leak, which can cause serious security threats. To mitigate the effects of route leak, accurate and timely route leak location is of great importance. Prior studies leverage AS business relationships to locate route leak in real time. However, they fail to achieve high location accuracy. Recent studies apply machine learning to accurately detect route leak from statistical features of massive BGP messages. Nevertheless, they have high detection latency and cannot further locate route leak. In this paper, we propose a real-time and accurate route leak location system named RoLL. It leverages distinctive AS triplet features to accurately locate AS triplets with route leak from each BGP update message in real time. Our experimental results on real-world BGP route leak data demonstrate that RoLL can achieve 91% location accuracy with less than 10 ms location latency. Jiahao Cao 0001, Zili Meng, Renjie Xie, Mingwei Xu 0001 |
ICC | 3 |
| 2023 | Enabling High Quality Real-Time Communications with Adaptive Frame-Rate
Zili Meng, Tingfeng Wang, Yixin Shen 0002, Bo Wang 0066, Mingwei Xu 0001, Venkat Arun, Hongxin Hu |
NSDI | 1 |
| 2023 | Bridging the Gap between QoE and QoS in Congestion Control: A Large-scale Mobile Web Service Perspective
Jia Zhang 0010, Enhuan Dong, Yan Zhang 0002, Shaorui Ren, Zili Meng, Mingwei Xu 0001, Zongzhi Hou, Xiaoming Fu 0001 |
USENIX ATC | 6 |
| 2023 | Automatic Performance-Optimal Offloading of Network Functions on Programmable SwitchesabstractIn network function virtualization (NFV), network functions (NFs) are chained as a service function chain (SFC) to enhance NF management with low cost and high flexibility. Recent NFV solutions indicate that the packet processing performance of SFCs can be significantly improved by offloading NFs to programmable switches. However, such offloading requires a deep understanding of heterogeneous NF properties (e.g., NF resource consumption and NF performance behaviors) to achieve the maximum SFC performance. Unfortunately, none of existing solutions provide automatic analysis of these NF properties. Thus, network administrators have to manually examine the source codes of NFs and profile various NF properties by hand, which is extremely time-consuming and laborious. In this article, we propose LightNF, a novel system that simplifies NF offloading in programmable networks. LightNF automatically dissects comprehensive NF properties by means of code analysis and performance profiling while eliminating manual efforts. It then leverages its analysis results of NF properties in its SFC placement so as to make the performance-optimal offloading decisions. We have implemented LightNF on Tofino-based hardware programmable switches. We perform extensive experiments to evaluate LightNF with a real-world testbed and large-scale simulation. Our experiments show that LightNF outperforms existing solutions with an orders-of-magnitude reduction in per-packet processing latency and 9.5× improvement in SFC throughput. Xiang Chen 0017, Hongyan Liu 0001, Dong Zhang 0010, Zili Meng, Qun Huang 0001, Haifeng Zhou, Chunming Wu 0001, Xuan Liu 0006, Qiang Yang 0004 |
IEEE Trans. Cloud Comput. | 4 |
| 2023 | Reducing Mobile Web Latency Through Adaptively Selecting Transport ProtocolabstractTo improve the performance of mobile web services, a new transport protocol, QUIC, has been recently proposed as a substitute for TCP. However, with pros and cons of QUIC, it is challenging to decide whether and when to use QUIC in large-scale real-world mobile web services. Complex temporal correlation of network conditions, high user heterogeneity in a nationwide deployment, implementation diversity of QUIC variants limited, and resources on mobile devices all affect the selection of transport protocols. In this paper, we present WiseTrans, an adaptive transport protocol selection mechanism, to switch transport protocols for mobile web services online and improve the completion time of web requests. WiseTrans introduces machine learning techniques to deal with temporal heterogeneity, makes decisions with historical information to handle spatial heterogeneity, adopts an online learning method to keep pace with implementation variation, and switches transport protocols at the request level to reach high performance with acceptable overhead. We implement WiseTrans on two platforms (Android and iOS) in a popular mobile web service application of Baidu. Comprehensive experiments demonstrate that WiseTrans can reduce request completion time by up to 25.8% on average compared to the usage of a single protocol. Jia Zhang 0010, Shaorui Ren, Enhuan Dong, Zili Meng, Yuan Yang 0001, Mingwei Xu 0001 |
IEEE/ACM Trans. Netw. | 4 |
| 2022 | Detecting Ephemeral Optical Events with OpTel
Congcong Miao, Minggang Chen, Arpit Gupta, Zili Meng, Lianjin Ye, Jingyu Xiao, Zekun He, Xulong Luo, Jilong Wang 0001, Heng Yu 0005 |
NSDI | 4 |
| 2022 | Achieving consistent low latency for wireless real-time communications with the shortest control loopabstractReal-time communication (RTC) applications like video conferencing or cloud gaming require consistent low latency to provide a seamless interactive experience. However, wireless networks including WiFi and cellular, albeit providing a satisfactory median latency, drastically degrade at the tail due to frequent and substantial wireless bandwidth fluctuations. We observe that the control loop for the sending rate of RTC applications is inflated when congestion happens at the wireless access point (AP), resulting in untimely rate adaption to wireless dynamics. Existing solutions, however, suffer from the inflated control loop and fail to quickly adapt to bandwidth fluctuations. In this paper, we propose Zhuge, a pure wireless AP based solution that reduces the control loop of RTC applications by separating congestion feedback from congested queues. We design a Fortune Teller to precisely estimate per-packet wireless latency upon its arrival at the wireless AP. To make Zhuge deployable at scale, we also design a Feedback Updater that translates the estimated latency to comprehensible feedback messages for various protocols and immediately delivers them back to senders for rate adaption. Trace-driven and real-world evaluation shows that Zhuge reduces the ratio of large tail latency and RTC performance degradation by 17% to 95%. Zili Meng, Yaning Guo, Chen Sun 0005, Bo Wang 0066, Justine Sherry, Hongqiang Harry Liu, Mingwei Xu 0001 |
SIGCOMM | 1 |
| 2021 | Physical-Layer Informed Multipath Redundancy Optimization for Mobile Real-Time CommunicationabstractNo abstract available. Zili Meng, Mingwei Xu 0001 |
APNet | 2 |
| 2021 | Towards Optimization for Large-scale Earth Observation Missions from a Global PerspectiveabstractNo abstract available. Yaning Guo, Zili Meng, Mingwei Xu 0001 |
APNet | 3 |
| 2021 | LightNF: Simplifying Network Function Offloading in Programmable NetworksabstractIn network function virtualization (NFV), network functions (NFs) are chained as a service function chain (SFC) to enhance NF management with high flexibility. Recent solutions indicate that the processing performance of SFCs can be significantly improved by offloading NFs to programmable switches. However, such offloading requires a deep understanding of NF properties to achieve the maximum SFC performance, which brings non-trivial burdens to network administrators. In this paper, we propose LightNF, a novel system that simplifies NF offloading in programmable networks. LightNF automatically dissects comprehensive NF properties (e.g., NF performance behaviors) via code analysis and performance profiling while eliminating manual efforts. It then leverages the analyzed NF properties in its SFC placement so as to produce the performance-optimal offloading. We have implemented a LightNF prototype. Our experiments show that LightNF outperforms state-of-the-art solutions with an orders-of-magnitude reduction in per-packet processing latency and 9.5× improvement in SFC throughput. Xiang Chen 0017, Qun Huang 0001, Peiqiao Wang, Zili Meng, Hongyan Liu 0001, Dong Zhang 0010, Haifeng Zhou, Chunming Wu 0001 |
IWQoS | 4 |
| 2021 | HierTopo: Towards High-Performance and Efficient Topology Optimization for Dynamic NetworksabstractDynamic networks have enabled dynamically adapting the network topology to meet the need of real-time traffic demands. However, due to the complexity of topology optimization, existing solutions suffer from a trade-off between performance and efficiency, which either have large optimality gaps or excessive optimization overhead. To break through this trade-off, our key observation is that we could offload the optimization procedure to every network node to handle the complexity. Thus, we propose HierTopo, a hierarchical topology optimization method for dynamic networks that achieves both high performance and efficiency. HierTopo firstly runs a local policy on each network node to aggregate network information into low-dimension features, then uses these features to make global topology decisions. Evaluation on real-world network traces shows that HierTopo outperforms the state-of-the-art solutions by 11.52-38.91% with only milliseconds of decision latency, and is also superior in generalization ability. Zili Meng, Yaning Guo, Mingwei Xu 0001, Hongxin Hu |
IWQoS | 2 |
| 2021 | WiseTrans: Adaptive Transport Protocol Selection for Mobile Web ServiceabstractTo improve the performance of mobile web service, a new transport protocol, QUIC, has been recently proposed. However, for large-scale real-world deployments, deciding whether and when to use QUIC in mobile web service is challenging. Complex temporal correlation of network conditions, high spatial heterogeneity of users in a nationwide deployment, and limited resources on mobile devices all affect the selection of transport protocols. In this paper, we present WiseTrans to adaptively switch transport protocols for mobile web service online and improve the completion time of web requests. Jia Zhang 0010, Enhuan Dong, Zili Meng, Yuan Yang 0001, Mingwei Xu 0001 |
WWW | 3 |
| 2021 | Practically Deploying Heavyweight Adaptive Bitrate Algorithms With Teacher-Student LearningabstractMajor commercial client-side video players employ adaptive bitrate (ABR) algorithms to improve the user quality of experience (QoE). With the evolvement of ABR algorithms, increasingly complex methods such as neural networks have been adopted to pursue better performance. However, these complex methods are too heavyweight to be directly deployed in client devices with limited resources, such as mobile phones. Existing solutions suffer from a trade-off between algorithm performance and deployment overhead. To make the deployment of sophisticated ABR algorithms practical, we propose PiTree, a general, high-performance, and scalable framework that can faithfully convert sophisticated ABR algorithms into decision trees with teacher-student learning. In this way, network operators can train complex models offline and deploy converted lightweight decision trees online. We also present theoretical analysis on the conversion and provide two upper bounds of the prediction error during the conversion and the generalization loss after conversion. Evaluation on three representative ABR algorithms with both trace-driven emulation and real-world experiments demonstrates that PiTree could convert ABR algorithms into decision trees with <; 3% average performance degradation. Moreover, compared to original deployment solutions, PiTree could save considerable operating expenses for content providers. Zili Meng, Yaning Guo, Yixin Shen 0002, Chao Zhou 0003, Minhu Wang, Jia Zhang 0010, Mingwei Xu 0001, Chen Sun 0005, Hongxin Hu |
IEEE/ACM Trans. Netw. | 1 |
| 2020 | SmartChain: Enabling High-Performance Service Chain Partition between SmartNIC and CPUabstractSmart Network Interface Cards (SmartNICs) have been widely used to accelerate software-based network functions (NFs). However, from the scope of a service chain, a careless selection of NFs to offload onto SmartNIC could severely degrade the performance due to frequent communications between CPU and SmartNIC. In this paper, we present SmartChain, a high performance and efficient framework that achieves optimal partition of service chains between SmartNIC and CPU. SmartChain consists of two logical steps. First, SmartChain analyzes the suitability of elements in a chain to run on SmartNIC to exploit its high performance. Besides, SmartChain also ensures the dependencies between elements. Second, as our key novelty, SmartChain models the service chain latency and resource constraints, and solves the partition problem with 0-1 integer linear programming. We implement a SmartChain prototype based on Netronome SmartNIC. Evaluation results show that when used in real world cases, SmartChain could reduce the service chain latency by up to 87% with throughput maintained compared with strawman solutions. Shuhe Wang, Zili Meng, Chen Sun 0005, Minhu Wang, Mingwei Xu 0001, Jun Bi, Tong Yang 0003, Qun Huang 0001, Hongxin Hu |
ICC | 2 |
| 2020 | Martini: Bridging the Gap between Network Measurement and Control Using Switching ASICsabstractAdvanced network management systems, including network measurement and traffic control, rely on a remote controller to make control decisions. However, this approach incurs a long control loop of a few seconds to minutes. Even if we switch to switch-local controller, the latency is still tens of milliseconds and is unacceptable for many latency-sensitive tasks. In this paper, we propose Martini, a general framework that supports measurement-based timely control. The key idea is to perform measurement, control decision, and control entirely in the switch data plane. This could shorten the control loop of management tasks that require timely control based on only locally measured statistics in the switch. First, Martini introduces a set of primitives to describe management tasks. Next, Martini provides an innovative network-wide task placement mechanism to exploit resources of all switches to accommodate massive management tasks. Finally, Martini provides a code library and a compiler to support measurement and control on a state-of-the-art switching ASIC. Evaluation results show that Martini can effectively support a wide range of fine-timescale management tasks such as microburst detection and fast load balancing by reducing the control loop from seconds to nanoseconds. Shuhe Wang, Chen Sun 0005, Zili Meng, Minhu Wang, Jiamin Cao, Mingwei Xu 0001, Jun Bi, Qun Huang 0001, Masoud Moshref, Tong Yang 0003, Hongxin Hu, Gong Zhang 0001 |
ICNP | 3 |
| 2020 | Interpreting Deep Learning-Based Networking SystemsabstractWhile many deep learning (DL)-based networking systems have demonstrated superior performance, the underlying Deep Neural Networks (DNNs) remain blackboxes and stay uninterpretable for network operators. The lack of interpretability makes DL-based networking systems prohibitive to deploy in practice. In this paper, we propose Metis, a framework that provides interpretability for two general categories of networking problems spanning local and global control. Accordingly, Metis introduces two different interpretation methods based on decision tree and hypergraph, where it converts DNN policies to interpretable rule-based controllers and highlight critical components based on analysis over hypergraph. We evaluate Metis over two categories of state-of-the-art DL-based networking systems and show that Metis provides human-readable interpretations while preserving nearly no degradation in performance. We further present four concrete use cases of Metis, showcasing how Metis helps network operators to design, debug, deploy, and ad-hoc adjust DL-based networking systems. Zili Meng, Minhu Wang, Jiasong Bai, Mingwei Xu 0001, Hongzi Mao, Hongxin Hu |
SIGCOMM | 1 |
| 2019 | When NFV Meets ANN: Rethinking Elastic Scaling for ANN-based NFsabstractNetwork Function Virtualization (NFV) provides middleboxes with substantial elasticity from a system level, and Artificial Neural Network (ANN) empowers middleboxes with great intelligence from an algorithm-level perspective. However, when ANN-based Network Functions (NFs) want to take advantage of the elasticity of NFV, our study finds that huge gaps exist between the existing approaches and the ideal goals for the elasticity control of ANN-based NFs. By revealing the key differences between ANN-based NFs and traditional NFs, we propose LEGO, an innovative framework that provides systematic mechanisms for traffic splitting, instance partition and runtime management to enable correct and efficient scaling of ANN-based NFs. Preliminary implementation and evaluation demonstrate the feasibility and effectiveness of the LEGO system. The major purpose of this paper is to highlight these challenges and sketch out a new roadmap towards ANN-based NFV paradigm. Menghao Zhang 0001, Jiasong Bai, Zili Meng, Hongda Li 0002, Hongxin Hu, Mingwei Xu 0001 |
ICNP | 4 |
| 2019 | PiTree: Practical Implementation of ABR Algorithms Using Decision TreesabstractMajor commercial client-side video players employ adaptive bitrate (ABR) algorithms to improve user quality of experience (QoE). With the evolvement of ABR algorithms, increasingly complex methods such as neural networks have been adopted to pursue better performance. However, these complex methods are too heavyweight to be directly implemented in client devices, especially mobile phones with very limited resources. Existing solutions suffer from a trade-off between algorithm performance and deployment overhead. To make the implementation of sophisticated ABR algorithms practical, we propose PiTree, a general, high-performance and scalable framework that can faithfully convert sophisticated ABR algorithms into lightweight decision trees to reduce deployment overhead. We also provide a theoretical upper bound on the optimization loss during the conversion. Evaluation results on three representative ABR algorithms demonstrate that PiTree could faithfully convert ABR algorithms into decision trees with <3% average performance degradation. Moreover, comparing to original implementation solutions, PiTree could save operating expenses for large content providers. Zili Meng, Yaning Guo, Chen Sun 0005, Hongxin Hu, Mingwei Xu 0001 |
ACM Multimedia | 1 |
| 2019 | Learning scheduling algorithms for data processing clustersabstractEfficiently scheduling data processing jobs on distributed compute clusters requires complex algorithms. Current systems use simple, generalized heuristics and ignore workload characteristics, since developing and tuning a scheduling policy for each workload is infeasible. In this paper, we show that modern machine learning techniques can generate highly-efficient policies automatically. Hongzi Mao, Malte Schwarzkopf, Shaileshh Bojja Venkatakrishnan, Zili Meng, Mohammad Alizadeh |
SIGCOMM | 4 |
| 2019 | MicroNF: An Efficient Framework for Enabling Modularized Service Chains in NFVabstractThe modularization of service function chains (SFCs) in network function virtualization (NFV) could introduce significant performance overhead and resource efficiency degradation due to introducing frequent packet transfer and consuming much more hardware resources. In response, we exploit the reusability, lightweightness, and individual scalability features of elements in modularized SFCs (MSFCs) and propose MicroNF, an efficient framework for MSFC in NFV. MicroNF addresses the performance overhead and resource efficiency problems in three ways. First, MicroNF graph constructor reuses the processing results of elements from different NFs and reconstructs the MSFC after modularization to shorten the chain latency. Second, optimized placer pays attention to the problem of which elements to consolidate and provides a performance-aware placement algorithm to place MSFCs compactly and optimize the global packet transfer cost. Third, MicroNF individual scaler innovatively introduces a push-aside scaling up strategy to avoid degrading performance and taking up new CPU cores. To support MSFC reusing and consolidation, MicroNF also designs a high-performance infrastructure to efficiently forwarding packets with consistency ensured and to automatically scheduling elements with fairness ensured when the elements are consolidated on the CPU core. Our evaluation results show that MicroNF achieves significant performance improvement and efficient resource utilization on several metrics. Zili Meng, Jun Bi, Haiping Wang 0002, Chen Sun 0005, Hongxin Hu |
IEEE J. Sel. Areas Commun. | 1 |
| 2018 | GEN: A GPU-Accelerated Elastic Framework for NFVabstractNetwork Function Virtualization (NFV) has the potential to enhance service delivery flexibility and reduce overall costs by provisioning software-based service function chains (SFCs) on commodity hardware. However, we observe that existing CPU-based SFC solutions cannot achieve both high performance and high elasticity simultaneously. To address such a critical challenge, we seek beyond CPU and exploit the capability of Graphics Processing Unit (GPU) to support NFV. We propose GEN, a GPU-based high performance and elastic framework for NFV. As opposed to pipeline-based SFCs in existing GPU-based NFV systems, GEN proposes to support RTC-based SFCs to improve processing performance. Meanwhile, GEN offers great elasticity of network function (NF) scaling up and down by allocating a different number of fine-grained GPU threads to an NF during runtime. We have implemented a prototype of GEN. Preliminary evaluation results demonstrate that GEN improves performance with RTC-based SFCs, and supports adaptive, precise, and fast NF scaling for NFV. Zhilong Zheng, Jun Bi, Chen Sun 0005, Heng Yu 0005, Hongxin Hu, Zili Meng, Shuhe Wang, Kai Gao 0001 |
APNet | 6 |
| 2018 | CoCo: Compact and Optimized Consolidation of Modularized Service Function Chains in NFVabstractThe modularization of Service Function Chains (SFCs) in Network Function Virtualization (NFV) could introduce significant performance overhead and resource efficiency degradation due to introducing frequent packet transfer and consuming much more hardware resources. In response, we exploit the lightweight and individually scalable features of elements in Modularized SFCs (MSFCs) and propose CoCo, a compact and optimized consolidation framework for MSFC in NFV. CoCo addresses the above problems in two ways. First, CoCo Optimized Placer pays attention to the problem of which elements to consolidate and provides a performance-aware placement algorithm to place MSFCs compactly and optimize the global packet transfer cost. Second, CoCo Individual Scaler innovatively introduces a push-aside scaling up strategy to avoid degrading performance and taking up new CPU cores. To support MSFC consolidation, CoCo also provides an automatic runtime scheduler to ensure fairness when elements are consolidated on CPU core. Our evaluation results show that CoCo achieves significant performance improvement and efficient resource utilization. Zili Meng, Jun Bi, Haiping Wang 0002, Chen Sun 0005, Hongxin Hu |
ICC | 1 |
| 2018 | OFM: Optimized Flow Migration for NFV Elasticity ControlabstractNetwork Function Virtualization (NFV) together with Software Defined Networking (SDN) offers the potential for enhancing service delivery flexibility and reducing overall costs. Based on the capability of dynamic creation and destruction of network function (NF) instances, NFV provides great elasticity in NF control, such as NF scaling out, scaling in, load balancing, etc. To realize NFV elasticity control, network traffic flows need to be redistributed across NF instances. However, deciding which flows are suitable for migration is a critical problem for efficient NFV elasticity control. In this paper, we propose to build an innovative flow migration controller, OFM Controller, to achieve optimized flow migration for NFV elasticity control. We identify the trigger conditions and control goals for different situations, and carefully design models and algorithms to address three major challenges including buffer overflow avoidance, migration cost calculation, and effective flow selection for migration. We implement the OFM Controller on top of NFV and SDN environments. Our evaluation results show that OFM Controller is efficient to support optimized flow migration in NFV elasticity control. Chen Sun 0005, Jun Bi, Zili Meng, Hongxin Hu |
IWQoS | 3 |
| 2018 | Enabling NFV Elasticity Control With Optimized Flow MigrationabstractNetwork function virtualization (NFV) together with software defined networking (SDN) offers the potential for enhancing service delivery flexibility and reducing overall costs. Based on the capability of dynamic creation and destruction of network function (NF) instances, NFV provides great elasticity in NF control, such as NF scaling out, scaling in, and load balancing. To realize NFV elasticity control, network traffic flows need to be redistributed across NF instances. However, deciding which flows are suitable for migration is a critical problem for efficient NFV elasticity control. In this paper, we propose to build an innovative flow migration controller, OFM controller, to achieve optimized flow migration for NFV elasticity control. We identify the trigger conditions and control goals for different situations, and carefully design models and algorithms to address three major challenges including buffer overflow avoidance, migration cost calculation, and effective flow selection for migration. We implement the OFM controller on top of NFV and SDN environments. Our evaluation results show that OFM controller is efficient to support optimized flow migration in NFV elasticity control. Chen Sun 0005, Jun Bi, Zili Meng, Tong Yang 0003, Hongxin Hu |
IEEE J. Sel. Areas Commun. | 3 |