Libin Liu 0001

dblp:59/8294-1 · DBLP profile ↗
← Back
19ranked-venue papers
8as first author
17since 2021 · last 2026
0000-0002-2718-8483ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 12 · 4 first-author · 12 since 2021Systems, architecture and hardware · 3 · 3 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Security and privacy · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 HeraClass: Towards Open-World Network Flow Classification via Traffic-Language Mapping
Ni Jin, Libin Liu 0001, Yukai Miao, Li Chen 0008, Dan Li 0001, Xizheng Wang, Xiuting Xu, Baojiang Cui
IWQoS2
2025 Resolving Packets from Counters: Enabling Multi-scale Network Traffic Super Resolution via Composable Large Traffic Model
Xizheng Wang, Libin Liu 0001, Li Chen 0008, Dan Li 0001, Yukai Miao, Yu Bai 0021
NSDI2
2025 Accelerating point cloud analytics on resource-constrained edge devices
Jingzong Li, Yik Hong Cai, Libin Liu 0001, Yu Mao 0001, Chun Jason Xue, Hong Xu 0001
Comput. Networks3
2025 Reducing Makespan via Optimizing Service Applications Scheduling Without Runtime Estimation
abstract
Efficient scheduling of service applications is critical for improving cluster resource utilization while minimizing makespan and application completion time. However, existing schedulers often struggle with coordinating task placement on worker machines due to the lack of runtime estimations. This limitation leads to two major performance issues: the non-synchronization problem and the contention-oblivious problem, both of which result in suboptimal application completion times. To address these challenges, Morbius is proposed, a scheduler that explicitly leverages the spatial structure of service applications to enhance scheduling decisions. Morbius adopts an all-or-nothing scheduling policy, ensuring that all tasks of an application are scheduled to run simultaneously, thereby effectively mitigating the non-synchronization problem. Within each priority queue, Morbius follows a shortest total time first policy, which facilitates contention-aware scheduling. Moreover, Morbius incorporates work conservation and starvation avoidance policies to better handle execution uncertainties and further improve application completion times. A prototype of Morbius is implemented on Yarn and evaluated in two environments: a homogeneous cluster with 36 machines and a heterogeneous cluster with 122 machines. Experimental results show that Morbius significantly outperforms existing approaches, improving average application completion time by up to 10.41× and reducing makespan by over 32.80%.
Libin Liu 0001, Zhixiong Niu, Xiuting Xu
IEEE Trans. Serv. Comput.1
2024 Marvel: Towards Efficient Federated Learning on IoT Devices
Libin Liu 0001, Xiuting Xu
Comput. Networks1
2024 Efficient Time-Series Data Delivery in IoT With Xender
abstract
Large amounts of time-series data need to be continually delivered from IoT devices to the cloud for real-time data analytics. The data delivery process is intrinsically slow and costly. Therefore, lots of work proposes various data reduction methods to accelerate it. Yet, they are either designed for the simple linear time-series data or computation-intensive, which is not suitable for the IoT devices with limited resources. In this paper, we propose Xender, a system to accelerate time-series data delivery. Xender consists of two key components: data sampler and data generator. Data sampler works on IoT devices to sample time-series data with low resource footprint, and data generator works on the cloud to efficiently generate data that significantly resembles the original. Besides, Xender can adapt to the dynamic characteristics of the time-series data with the content-aware mechanism, as well as the dynamic computation resources by supporting multiple data generation quality levels and using the anytime generation mechanism. We implement Xender and evaluate it with testbed experiments using six real-world datasets. The results show that it can significantly reduce data delivery time by 45.79% on average compared against existing schemes, and adapt to computation resources with up to 1014.40Mbps data generation throughput.
Libin Liu 0001, Jingzong Li, Zhixiong Niu, Wei Zhang 0049, Chun Jason Xue, Hong Xu 0001
IEEE Trans. Mob. Comput.1
2023 Cross-Camera Inference on the Constrained Edge
abstract
The proliferation of edge devices has pushed computing from the cloud to the data sources, and video analytics is among the most promising applications of edge computing. Running video analytics is compute- and latency-sensitive, as video frames are analyzed by complex deep neural networks (DNNs) which put severe pressure on resource-constrained edge devices. To resolve the tension between inference latency and resource cost, we present Polly, a cross-camera inference system that enables co-located cameras with different but overlapping fields of views (FoVs) to share inference results between one another, thus eliminating the redundant inference work for objects in the same physical area. Polly’s design solves two basic challenges of cross-camera inference: how to identify overlapping FoVs automatically, and how to share inference results accurately across cameras. Evaluation on NVIDIA Jetson Nano with a real-world traffic surveillance dataset shows that Polly reduces the inference latency by up to 71.4% while achieving almost the same detection accuracy with state-of-the-art systems.
Jingzong Li, Libin Liu 0001, Hong Xu 0001, Shudeng Wu, Chun Jason Xue
INFOCOM2
2023 Moby: Empowering 2D Models for Efficient Point Cloud Analytics on the Edge
abstract
3D object detection plays a pivotal role in many applications, most notably autonomous driving and robotics. These applications are commonly deployed on edge devices to promptly interact with the environment, and often require near real-time response. With limited computation power, it is challenging to execute 3D detection on the edge using highly complex neural networks. Common approaches such as offloading to the cloud induce significant latency overheads due to the large amount of point cloud data during transmission. To resolve the tension between wimpy edge devices and compute-intensive inference workloads, we explore the possibility of empowering fast 2D detection to extrapolate 3D bounding boxes. To this end, we present Moby, a novel system that demonstrates the feasibility and potential of our approach. We design a transformation pipeline for Moby that generates 3D bounding boxes efficiently and accurately based on 2D detection results without running 3D detectors. Further, we devise a frame offloading scheduler that decides when to launch the 3D detector judiciously in the cloud to avoid the errors from accumulating. Extensive evaluations on NVIDIA Jetson TX2 with real-world autonomous driving datasets demonstrate that Moby offers up to 91.9% latency improvement with modest accuracy loss over state of the art.
Jingzong Li, Yik Hong Cai, Libin Liu 0001, Yu Mao 0001, Chun Jason Xue, Hong Xu 0001
ACM Multimedia3
2023 Efficient Real-time Video Conferencing with Adaptive Frame Delivery
Libin Liu 0001, Jingzong Li, Hong Xu 0001, Chun Jason Xue
Comput. Networks1
2023 Bottleneck-Aware Non-Clairvoyant Coflow Scheduling With Fai
abstract
Coflow scheduling is critical to data-parallel applications in data centers. While schemes like Varys can achieve optimal performance, they require a priori information about coflows which is hard to obtain in practice. Existing non-clairvoyant solutions like Aalo generalize least attained service (LAS) scheduling discipline to address this issue. However, they fail to identify the bottleneck flows in a coflow and tend to allocate excessive bandwidth to the non-bottleneck flows, leading to bandwidth wastage and inferior overall performance. To this end, we present Fai that strives to improve the overall coflow performance by accelerating the bottleneck flows without priori knowledge. Fai employs bottleneck-aware scheduling. It adopts loose coordination to update coflow priority and flow rates based on total bytes sent. In addition, Fai detects bottleneck flows based on a flow’s rate and bytes sent, and de-allocates bandwidth for other flows to match the bottleneck rate without affecting the coflow completion time (CCT). The saved bandwidth is then distributed among coflows according to their priority to improve overall performance. Testbed evaluation on a 40-node cluster shows that Fai improves average (P95) CCT by 1.73× (3.43×), compared to Aalo. Large-scale trace-driven simulations also show that Fai outperforms Aalo substantially.
Libin Liu 0001, Chengxi Gao, Peng Wang 0037, Hongming Huang, Jiamin Li 0002, Hong Xu 0001, Wei Zhang 0049
IEEE Trans. Cloud Comput.1
2022 Efficient Clustered Network Telemetry based on Failure Awareness
abstract
Nowadays, various network telemetry technologies are proposed to monitor the network and detect failures accurately in real-time, which can be categorized into two types, including the proactive network telemetry (NT) and the passive one. The passive NT can monitor the network with low band-width overhead, yet, cannot guarantee full network coverage. The proactive one can achieve full coverage, yet, lead to high bandwidth cost. To deal with the problem, we propose a failure-aware clustered network telemetry approach, called CNT. CNT leverages the practical objective network operating experience: different network links have various failure probabilities. It is aware of the failure probabilities and assigns the network links into two clusters accordingly. Then, based on the original network topology, CNT designs an active path planning algorithm to connect the two clusters of links into two sub-topologies, respectively. Finally, CNT performs network telemetry with different cycles. We evaluate CNT with various simulation experiments. The results show that compared to existing proactive schemes, CNT can achieve comparable network coverage with less cost.
Libin Liu 0001, Lizhuang Tan, Wei Gao 0030, Wei Zhang 0049
APNOMS2
2022 Efficient OFDM Channel Estimation with RRDBNet
abstract
Channel estimation is important for orthogonal frequency division multiplexing (OFDM) in current wireless communication systems. Prevalent channel estimation algorithms, however, cannot be widely deployed due to some practical reasons, such as poor robustness and high computational complexity. To solve the problems for OFDM systems, we propose a new channel estimation scheme with a fine-designed deep learning model, called RRDBNet. RRDBNet can be trained easily while maintaining the advantages of residual learning and increasing the structure capacity, by combining the multi-level residual network and dense links. Our simulation results show that RRDBNet outperforms the traditional least-square algorithm and existing DL-based super-resolution schemes, which ranges from 0.5 to 1dB at low SNR and from 2 to 3dB at high SNR. Besides, in terms of the number of pilots, RRDBNet is also superior to existing schemes and approaches LMMSE.
Wei Gao 0030, Meihong Yang, Wei Zhang 0049, Libin Liu 0001
ISCC4
2022 Software-defined network assimilation: bridging the last mile towards centralized network configuration management with NAssim
abstract
On-boarding new devices into an existing SDN network is a pain for network operations (NetOps) teams, because much expert effort is required to bridge the gap between the configuration models of the new devices and the unified data model in the SDN controller. In this work, we present an assistant framework NAssim, to help NetOps accelerate the process of assimilating a new device into a SDN network. Our solution features a unified parser framework to parse diverse device user manuals into preliminary configuration models, a rigorous validator that confirm the correctness of the models via formal syntax analysis, model hierarchy validation and empirical data validation, and a deep-learning-based mapping algorithm that uses state-of-the-art neural language processing techniques to produce human-comprehensible recommended mapping between the validated configuration model and the one in the SDN controller. In all, NAssim liberates the NetOps from most tedious tasks by learning directly from devices' manuals to produce data models which are comprehensible by both the SDN controller and human experts. Our evaluation shows, NAssim can accelerate the assimilation process by 9.1x. In this process, we also identify and correct 243 errors in four mainstream vendors' device manuals, and release a validated and expert-curated dataset of parsed manual corpus for future research.
Huangxun Chen, Yukai Miao, Li Chen 0008, Haifeng Sun 0001, Hong Xu 0001, Libin Liu 0001, Gong Zhang 0001, Wei Wang 0011
SIGCOMM6
2022 DeepQueueNet: towards scalable and generalized network performance estimation with packet-level visibility
abstract
Network simulators are an essential tool for network operators, and can assist important tasks such as capacity planning, topology design, and parameter tuning. Popular simulators are all based on discrete event simulation, and their performance does not scale with the size of modern networks. Recently, deep-learning-based techniques are introduced to solve the scalability problem, but, as we show with experiments, they have poor visibility in their simulation results, and cannot generalize to diverse scenarios. In this work, we combine scalable and generalized continuous simulation techniques with discrete event simulation to achieve high scalability, while providing packet-level visibility. We start from a solid queueing-theoretic modeling of modern networks, and carefully identify the mathematically-intractable or computationally-expensive parts, only which are then modeled using deep neural networks (DNN). Dubbed DeepQueueNet, our approach combines prior knowledge of networks, and supports arbitrary topology and device traffic management mechanisms (given sufficient training data). Our extensive experiments show that DeepQueueNet achieves near-linear speedup in the number of GPUs, and its estimation accuracy for average and 99th percentile round-trip time outperforms existing end-to-end DNN-based performance estimators in all scenarios.
Xi Peng 0006, Li Chen 0008, Libin Liu 0001, Jingze Zhang, Hong Xu 0001, Baochun Li, Gong Zhang 0001
SIGCOMM4
2022 ScaleFlux: Efficient Stateful Scaling in NFV
abstract
Network function virtualization (NFV) enables elastic scaling to middlebox deployment and management. Therefore, efficient stateful scaling is an important task because operators often need to shift traffic and the associated flow states across VNF instances to deal with time-varying loads. Existing NFV scaling methods, however, typically focus on one aspect of the scaling pipeline and does not offer an end-to-end scaling framework. This article presents ScaleFlux, a complete stateful scaling system that efficiently reduces flow-level latency and achieves near-optimal resource usage. ScaleFlux (1) monitors traffic load for each VNF instance and adopts a queue-based mechanism to detect load burstiness timely, (2) deploys a flow bandwidth predictor to predict flow bandwidth time-series with the ABCNN-LSTM model, and (3) schedules the necessary flow and state migration using the simulated annealing algorithm to achieve both flow-level latency guarantee and resource usage minimization. Testbed evaluation with a five-machine cluster shows that ScaleFlux reduces flow completion time by at least 8.7× for all the workloads and achieves near-optimal CPU usage during scaling.
Libin Liu 0001, Hong Xu 0001, Zhixiong Niu, Jingzong Li, Wei Zhang 0049, Peng Wang 0037, Jiamin Li 0002, Chun Jason Xue, Cong Wang 0001
IEEE Trans. Parallel Distributed Syst.1
2021 Elasecutor: Elastic Executor Scheduling in Data Analytics Systems
abstract
Modern data analytics systems use long-running executors to run an application's entire DAG. Executors exhibit salient time-varying resource requirements. Yet, existing schedulers simply reserve resources for executors statically, and use the peak resource demand to guide executor placement. This leads to low utilization and poor application performance. We present Elasecutor, a novel executor scheduler for data analytics systems. Elasecutor dynamically allocates and explicitly sizes resources to executors over time according to the predicted time/varying resource demands. Rather than placing executors using their peak demand, Elasecutor strategically assigns them to machines based on a concept called dominant remaining resource to minimize resource fragmentation. Elasecutor further adaptively reprovisions resources in order to tolerate inaccurate demand prediction and reschedules tasks to deal with inadequate reprovisioning resources on one machine. Testbed evaluation on a 35-node cluster with our Spark-based prototype implementation shows that Elasecutor reduces makespan by more than 36% on average, and improves cluster utilization by up to 55% compared to existing work.
Libin Liu 0001, Hong Xu 0001
IEEE/ACM Trans. Netw.1
2021 RepNet: Cutting Latency with Flow Replication in Data Center Networks
abstract
Data center networks need to provide low latency, especially at the tail, as demanded by many interactive applications. To improve tail latency, existing approaches require modifications to switch hardware and/or end-host operating systems, making them difficult to be deployed. We present the design, implementation, and evaluation of RepNet, an application layer transport that can be deployed today. RepNet exploits the fact that only a few paths among many are congested at any moment in the network, and applies simple flow replication to mice flows to opportunistically use the less congested path. RepNet has two designs for flow replication: (1) RepSYN, which only replicates SYN packets and uses the first connection that finishes TCP handshaking for data transmission, and (2) RepFlow which replicates the entire mice flow. We implement RepNet on node.js, one of the most commonly used platforms for networked interactive applications. node's single threaded event-loop and non-blocking I/O make flow replication highly efficient. Performance evaluation on a real network testbed and in Mininet reveals that RepNet is able to reduce the tail latency of mice flows, as well as application completion times, by more than 50 percent.
Shuhao Liu 0001, Hong Xu 0001, Libin Liu 0001, Wei Bai 0001, Kai Chen 0005, Zhiping Cai
IEEE Trans. Serv. Comput.3
2018 Elasecutor: Elastic Executor Scheduling in Data Analytics Systems
abstract
Modern data analytics systems use long-running executors to run an application's entire DAG. Executors exhibit salient time-varying resource requirements. Yet, existing schedulers simply reserve resources for executors statically, and use the peak resource demand to guide executor placement. This leads to low utilization and poor application performance.
Libin Liu 0001, Hong Xu 0001
SoCC1
2018 Kuijia: Traffic Rescaling in Software-Defined Data Center WANs
abstract
Network faults like link or switch failures can cause heavy congestion and packet loss. Traffic engineering systems need a lot of time to detect and react to such faults, which results in significant recovery times. Recent work either preinstalls a lot of backup paths in the switches to ensure fast rerouting or proactively prereserves bandwidth to achieve fault resiliency. Our idea agilely reacts to failures in the data plane while eliminating the preinstallation of backup paths. We propose Kuijia, a robust traffic engineering system for data center WANs, which relies on a novel failover mechanism in the data plane called rate rescaling. The victim flows on failed tunnels are rescaled to the remaining tunnels and enter lower priority queues to avoid performance impairment of aboriginal flows. Real system experiments show that Kuijia is effective in handling network faults and significantly outperforms the conventional rescaling method.
Che Zhang, Hong Xu 0001, Libin Liu 0001, Zhixiong Niu, Peng Wang 0037
Secur. Commun. Networks3