VLDB 2026 Research / reviewers in the wild / expert
Jingzong Li
dblp:327/7948
· DBLP profile ↗
12ranked-venue papers
4as first author
12since 2021 · last 2026
0000-0002-2519-2550ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 6 · 2 first-author · 6 since 2021Systems, architecture and hardware · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ReasonCache: Accelerating Large Reasoning Model Serving through KV Cache Sharing
Xin Tan 0004, Minchen Yu, Jingzong Li, Hong Xu 0001 |
IWQoS | 4 |
| 2025 | Easz: An Agile Transformer-based Image Compression Framework for Resource-constrained IoTsabstractNeural image compression, necessary in various machine-to-machine communication scenarios, suffers from its heavy encode-decode structures and inflexibility in switching between different compression levels. Consequently, it raises significant challenges in applying the neural image compression to edge devices that are developed for powerful servers with high computational and storage capacities. We take a step to solve the challenges by proposing a new transformer-based edge-computefree image coding framework called Easz. Easz shifts the computational overhead to the server, and hence avoids the heavy encoding and model switching overhead on the edge. Easz utilizes a patch-erase algorithm to selectively remove image contents using a conditional uniform-based sampler. The erased pixels are reconstructed on the receiver side through a transformer-based framework. To further reduce the computational overhead on the receiver, we then introduce a lightweight transformer-based reconstruction structure to reduce the reconstruction load on the receiver side. Extensive evaluations conducted on a realworld testbed demonstrate multiple advantages of Easz over existing compression approaches, in terms of adaptability to different compression levels, computational efficiency, and image reconstruction quality. Yu Mao 0001, Jingzong Li, Hong Xu 0001, Tei-Wei Kuo, Nan Guan, Chun Jason Xue |
DAC | 2 |
| 2025 | Accelerating point cloud analytics on resource-constrained edge devices
Jingzong Li, Yik Hong Cai, Libin Liu 0001, Yu Mao 0001, Chun Jason Xue, Hong Xu 0001 |
Comput. Networks | 1 |
| 2024 | Arlo: Serving Transformer-based Language Models with Dynamic Input LengthsabstractA prominent challenge in serving requests for NLP tasks is handling the varying length of input texts. Existing solutions, such as uniform zero-padding and compiler support, suffer from either computational inefficiency or suboptimal latency. To address these practical issues, we propose an approach called polymorphing. Polymorphing involves creating and utilizing multiple runtimes of the model, each statically compiled with a different input length, to serve requests accordingly. This fine-grained use of statically-compiled runtimes reduces the overheads of zero-padding while improving latency performance compared to dynamic compilation. To practically realize polymorphing, we have developed an inference scheduling system, Arlo, which leverages the observed input length distribution to periodically allocate compute resources across multiple runtimes by solving an integer linear program. Upon request arrival, Arlo uses a multi-level queue-based heuristic to dispatch requests to the most suitable runtime instances, efficiently adapting to the dynamics of request length and instance load. Extensive testbed evaluations and large-scale simulations using production traces demonstrate Arlo’s promising potential. It achieves 23.7%–98.1% mean latency reductions compared to existing schemes while significantly reducing tail latency. Xin Tan 0004, Jiamin Li 0002, Jingzong Li, Hong Xu 0001 |
ICPP | 4 |
| 2024 | A Learning-only Method for Multi-Cell Multi-User MIMO Sum Rate MaximizationabstractSolving the sum rate maximization problem for interference reduction in multi-cell multi-user multiple-input multiple-output (MIMO) wireless communication systems has been investigated for a decade. Several machine learning-assisted methods have been proposed under conventional sum rate maximization frameworks, such as the Weighted Minimum Mean Square Error (WMMSE) framework. However, existing learning-assisted methods suffer from a deficiency in parallelization, and their performance is intrinsically bounded by WMMSE. In contrast, we propose a structural learning-only framework from the abstraction of WMMSE. Our proposed framework increases the solvability of the original MIMO sum rate maximization problem by dimension expansion via a unitary learnable parameter matrix to create an equivalent problem in a higher dimension. We then propose a structural solution updating method to solve the higher dimensional problem, utilizing neural networks to generate the learnable matrix-multiplication parameters. We show that the proposed structural learning framework achieves lower complexity than WMMSE thanks to its parallel implementation. Simulation results under practical communication network settings demonstrate that our proposed learning-only framework achieves up to 98% optimality over state-of-the-art algorithms while providing up to 47× acceleration in various scenarios. Qingyu Song 0002, Juncheng Wang 0001, Jingzong Li, Guochen Liu, Hong Xu 0001 |
INFOCOM | 3 |
| 2024 | Efficient Time-Series Data Delivery in IoT With XenderabstractLarge amounts of time-series data need to be continually delivered from IoT devices to the cloud for real-time data analytics. The data delivery process is intrinsically slow and costly. Therefore, lots of work proposes various data reduction methods to accelerate it. Yet, they are either designed for the simple linear time-series data or computation-intensive, which is not suitable for the IoT devices with limited resources. In this paper, we propose Xender, a system to accelerate time-series data delivery. Xender consists of two key components: data sampler and data generator. Data sampler works on IoT devices to sample time-series data with low resource footprint, and data generator works on the cloud to efficiently generate data that significantly resembles the original. Besides, Xender can adapt to the dynamic characteristics of the time-series data with the content-aware mechanism, as well as the dynamic computation resources by supporting multiple data generation quality levels and using the anytime generation mechanism. We implement Xender and evaluate it with testbed experiments using six real-world datasets. The results show that it can significantly reduce data delivery time by 45.79% on average compared against existing schemes, and adapt to computation resources with up to 1014.40Mbps data generation throughput. Libin Liu 0001, Jingzong Li, Zhixiong Niu, Wei Zhang 0049, Chun Jason Xue, Hong Xu 0001 |
IEEE Trans. Mob. Comput. | 2 |
| 2023 | Faster and Stronger Lossless Compression with Optimized Autoregressive FrameworkabstractNeural AutoRegressive (AR) framework has been applied in general-purpose lossless compression recently to improve compression performance. However, this paper found that directly applying the original AR framework causes the duplicated processing problem and the in-batch distribution variation problem, which leads to deteriorated compression performance. The key to address the duplicated processing problem is to disentangle the processing of the history symbol set at the input side. Two new types of neural blocks are first proposed. An individual-block performs separate feature extraction on each history symbol while a mix-block models the correlation between extracted features and estimates the probability. A progressive AR-based compression framework (PAC) is then proposed, which only requires one history symbol from the host at a time rather than the whole history symbol set. In addition, we introduced a trainable matrix multiplication to model the ordered importance, replacing previous hardware-unfriendly Gumble-Softmax sampling. The in-batch distribution variation problem is caused by AR-based compression’s structured batch construction. Based on this observation, a batch-location-aware individual block is proposed to capture the heterogeneous in-batch distributions precisely, improving the performance without efficiency losses. Experimental results show the proposed framework can achieve an average of 130% speed improvement with an average of 3% compression ratio gain across data domains compared to the state-of-the-art. Yu Mao 0001, Jingzong Li, Yufei Cui, Chun Jason Xue |
DAC | 2 |
| 2023 | Cross-Camera Inference on the Constrained EdgeabstractThe proliferation of edge devices has pushed computing from the cloud to the data sources, and video analytics is among the most promising applications of edge computing. Running video analytics is compute- and latency-sensitive, as video frames are analyzed by complex deep neural networks (DNNs) which put severe pressure on resource-constrained edge devices. To resolve the tension between inference latency and resource cost, we present Polly, a cross-camera inference system that enables co-located cameras with different but overlapping fields of views (FoVs) to share inference results between one another, thus eliminating the redundant inference work for objects in the same physical area. Polly’s design solves two basic challenges of cross-camera inference: how to identify overlapping FoVs automatically, and how to share inference results accurately across cameras. Evaluation on NVIDIA Jetson Nano with a real-world traffic surveillance dataset shows that Polly reduces the inference latency by up to 71.4% while achieving almost the same detection accuracy with state-of-the-art systems. Jingzong Li, Libin Liu 0001, Hong Xu 0001, Shudeng Wu, Chun Jason Xue |
INFOCOM | 1 |
| 2023 | Moby: Empowering 2D Models for Efficient Point Cloud Analytics on the Edgeabstract3D object detection plays a pivotal role in many applications, most notably autonomous driving and robotics. These applications are commonly deployed on edge devices to promptly interact with the environment, and often require near real-time response. With limited computation power, it is challenging to execute 3D detection on the edge using highly complex neural networks. Common approaches such as offloading to the cloud induce significant latency overheads due to the large amount of point cloud data during transmission. To resolve the tension between wimpy edge devices and compute-intensive inference workloads, we explore the possibility of empowering fast 2D detection to extrapolate 3D bounding boxes. To this end, we present Moby, a novel system that demonstrates the feasibility and potential of our approach. We design a transformation pipeline for Moby that generates 3D bounding boxes efficiently and accurately based on 2D detection results without running 3D detectors. Further, we devise a frame offloading scheduler that decides when to launch the 3D detector judiciously in the cloud to avoid the errors from accumulating. Extensive evaluations on NVIDIA Jetson TX2 with real-world autonomous driving datasets demonstrate that Moby offers up to 91.9% latency improvement with modest accuracy loss over state of the art. Jingzong Li, Yik Hong Cai, Libin Liu 0001, Yu Mao 0001, Chun Jason Xue, Hong Xu 0001 |
ACM Multimedia | 1 |
| 2023 | Efficient Real-time Video Conferencing with Adaptive Frame Delivery
Libin Liu 0001, Jingzong Li, Hong Xu 0001, Chun Jason Xue |
Comput. Networks | 2 |
| 2022 | Alfie: Neural-Reinforced Adaptive Prefetching for Short VideosabstractShort videos have received extraordinary success in recent years. To provide smooth playback and avoid rebuffering delay, prefetching upcoming videos is commonly used in cellular networks. Current prefetching designs fall short in dealing with bandwidth overhead, especially the exit overhead of downloaded but unconsumed chunks due to user exit. Measurement from a large short video platform shows that exit overhead accounts for up to 43.5% of bandwidth overhead. Thus we build Alfie, a bandwidth-efficient short video prefetching algorithm via reinforcement learning. Essentially Alfie adjusts prefetching based upon user viewing patterns in addition to network conditions. We demonstrate that Alfie outperforms the state of the art by up to 26.8% in overall performance while reducing the exit overhead by up to 84.9%. Jingzong Li, Hong Xu 0001, Haopeng Yan, Chun Jason Xue |
ICME | 1 |
| 2022 | ScaleFlux: Efficient Stateful Scaling in NFVabstractNetwork function virtualization (NFV) enables elastic scaling to middlebox deployment and management. Therefore, efficient stateful scaling is an important task because operators often need to shift traffic and the associated flow states across VNF instances to deal with time-varying loads. Existing NFV scaling methods, however, typically focus on one aspect of the scaling pipeline and does not offer an end-to-end scaling framework. This article presents ScaleFlux, a complete stateful scaling system that efficiently reduces flow-level latency and achieves near-optimal resource usage. ScaleFlux (1) monitors traffic load for each VNF instance and adopts a queue-based mechanism to detect load burstiness timely, (2) deploys a flow bandwidth predictor to predict flow bandwidth time-series with the ABCNN-LSTM model, and (3) schedules the necessary flow and state migration using the simulated annealing algorithm to achieve both flow-level latency guarantee and resource usage minimization. Testbed evaluation with a five-machine cluster shows that ScaleFlux reduces flow completion time by at least 8.7× for all the workloads and achieves near-optimal CPU usage during scaling. Libin Liu 0001, Hong Xu 0001, Zhixiong Niu, Jingzong Li, Wei Zhang 0049, Peng Wang 0037, Jiamin Li 0002, Chun Jason Xue, Cong Wang 0001 |
IEEE Trans. Parallel Distributed Syst. | 4 |