VLDB 2026 Research / reviewers in the wild / expert
Jianxin Shi 0005
dblp:78/2327-5
· DBLP profile ↗
19ranked-venue papers
6as first author
18since 2021 · last 2026
0000-0002-7687-8480ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 10 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 8 since 2021Systems, architecture and hardware · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ACPGS: Towards Bandwidth-Efficient Delivery of 3D Gaussian Splatting
Cong Zhang 0002, Jianxin Shi 0005, Xiaoyi Fan 0001, Laizhong Cui, Jiangchuan Liu |
NOSSDAV | 3 |
| 2026 | Implicit Representation-based Volumetric Video Streaming for Photorealistic Full-scene ExperienceabstractThe widespread integration of the Internet of Things with sensors like depth-of-field cameras, LiDAR scanners, and eye-tracking infrared sensors, in head-mounted devices, has ushered in a new era of immersive digital experiences. Full-scene volumetric video (VV), a key innovation in this integration, provides a deeply immersive experience by capturing the richness and detail of the 3D world. However, its massive data volume presents significant streaming challenges. While 3D tile-based viewport approaches have been proposed, they struggle to full-scene VV given the small video buffer limitation, high tile segmentation overhead, and lack of full-scene consideration. In this work, inspired by the advancements of implicit neural radiance field (NeRF), we present \({\mathsf{V}^{2}\mathsf{NeRF}}\) , a novel full-scene VV streaming system featured by layered representation. It harmonizes the NeRF with explicit point clouds to represent the static background and dynamic foreground, thereby avoiding large data transfers and achieving photorealistic content representation. To tackle the issues of intensive computation requirements and multiscale adaptation scheduling within \({\mathsf{V}^{2}\mathsf{NeRF}}\) system, we propose a lightweight non-visible background removal method and a two-stage decoupled architecture. In addition, an efficient buffer-aware simulated annealing algorithm is developed, alongside the utilization of a perceptually learned metric, to enhance user experience. We further discuss the concerns about practical development and deployment. Extensive prototype evaluations demonstrate \({\mathsf{V}^{2}\mathsf{NeRF}}\) ’s superior streaming and viewing performance on a wide variety of networks, viewing motions, and scenes. For instance, compared to state-of-the-art approaches, it achieves a 24% increment in perceptual quality, an 83% reduction in rebuffering time, and a 54% enhancement in user experience on average. Jianxin Shi 0005, Miao Zhang 0003, Linfeng Shen, Jiangchuan Liu, Yuan Zhang 0013, Lingjun Pu, Jingdong Xu |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2025 | Flick: Frame-Perceptive Packet Scheduling for Low-Latency Video Services in Wi-Fi NetworksabstractEmerging low-latency video (LLV) services, such as cloud gaming, video conferencing, and virtual reality, demand ultra-low latency for smooth interaction. However, existing methods often overlook the misalignment between frame-level perception and packet-level scheduling in ubiquitous Wi-Fi networks, causing high tail latency and degraded user experience. To this end, we propose Flick, a frame-perceptive packet scheduling framework at Wi-Fi access points (APs). It leverages the periodic per-frame transmission behavior and the LLV traffic distribution characteristics to infer the end-to-end frame latency at APs. Flick consists of three components: a Frame Boundary Identifier that detects video frame boundaries using only packet size, an End-to-End Frame Latency Estimator that estimates the end-to-end latency without sender or receiver timestamps, and a Fast-Send, Slow-Recovery Scheduler that dynamically adjusts scheduling priority based on inferred latency. We implement Flick on a commercial Wi-Fi AP. Testbed results show that Flick reduces P99 latency and stall rate by 57% and 81%, respectively, while preserving 95% throughput and maintaining high fairness. Qianyun Gong, Jiapei Xu, Jianxin Shi 0005, Xinjing Yuan, Lingjun Pu, Jingdong Xu |
ICNP | 3 |
| 2025 | Lightweight in-Network Flow Classification with Deep Differentiable Logic Gate NetworksabstractDeploying artificial intelligence (AI) models on the programmable data plane is a key direction toward realizing intelligent data planes. However, this vision faces significant challenges: existing approaches often incur excessive consumption of scarce switch hardware resources or introduce packet recirculation, leading to performance bottlenecks. To address these issues, we propose SwitchLGN, a novel Deep Differentiable Logic Gate Network (DDLGN) architecture deployable on programmable switches. The core of SwitchLGN lies in its hardware-aligned design, which decomposes the model into multiple independent sub-layers and maps each to distinct pipeline processing units. This design enables the entire inference process to be executed solely with the switch's native bit-level logic operations, thereby eliminating the challenges of cross-cluster computation and complex arithmetic in the data plane. In addition, we develop a compilation toolchain to support the automated mapping of SwitchLGN models to P4 code. Experimental evaluations show that SwitchLGN achieves accuracies of 99.42% for network anomaly detection and 95.59% for flow size classification, while reducing SRAM usage to below 3% and making TCAM consumption negligible. Notably, it delivers line-rate inference without packet recirculation, achieving a per-packet latency of only 283 ns. Kaiwei Gao, Xinjing Yuan, Jianxin Shi 0005, Lingjun Pu |
ICPADS | 4 |
| 2025 | BAROC: Concealing Packet Losses in LSNs with Bimodal Behavior Awareness for Livecast Ingestion
Haoyuan Zhao, Jianxin Shi 0005, Guanzhen Wu, Hao Fang 0012, Yi Ching Chou, Long Chen 0025, Feng Wang 0001, Jiangchuan Liu |
INFOCOM | 2 |
| 2025 | Libra: Novel LLM Token Streaming via Region-Based Task Scheduling and Token BundlingabstractThe LLM serving systems are increasingly growing in popularity, as they provide various capabilities ranging from realtime translation to AI-driven chatbots. Recently, significant effort has been made to optimize server-side metrics such as token generation throughput, while the optimization of token streaming is simply overlooked, resulting in excessive network traffic and poor network utilization. In this paper, we introduce user regions to relieve the network issues, since users from the same region (e.g., universities and business zones) are likely to share similar behaviors to access LLM serving systems (e.g., they are active in a period of time). In this context, we propose Libra, a proxy-cloud collaborative serving system, where the cloud generates and bundles the tokens in terms of user regions and region-based proxy extracts and repacks the received token bundle to their corresponding users. At its core, we design an online region-based task scheduling algorithm with a provable performance to optimize user QoE and system overhead over time. Our evaluations show that Libra outperforms the state-of-the-art LLM serving systems (without user regions), such as vLLM, VTC and Andes, by up to$2.1 \times$in the Time-Between-Tokens (TBT) metric and$78.4 \times$in the number of packets. In addition, it achieves a 32.6 % reduction in TBT compared to other alternative algorithms (with user regions). Chengjin Zhou, Xinjing Yuan, Jianxin Shi 0005, Yuan Zhang 0013, Lingjun Pu |
IWQoS | 4 |
| 2025 | TrackerSplat: Exploiting Point Tracking for Fast and Robust Dynamic 3D Gaussians ReconstructionabstractRecent advancements in 3D Gaussian Splatting (3DGS) have demonstrated its potential for efficient and photorealistic 3D reconstructions, which is crucial for diverse applications such as robotics and immersive media. However, current Gaussian-based methods for dynamic scene reconstruction struggle with large inter-frame displacements, leading to artifacts and temporal inconsistencies under fast object motions. To address this, we introduce TrackerSplat, a novel method that integrates advanced point tracking methods to enhance the robustness and scalability of 3DGS for dynamic scene reconstruction. TrackerSplat utilizes off-the-shelf point tracking models to extract pixel trajectories and triangulate per-view pixel trajectories onto 3D Gaussians to guide the relocation, rotation, and scaling of Gaussians before training. This strategy effectively handles large displacements between frames, dramatically reducing the fading and recoloring artifacts prevalent in prior methods. By accurately positioning Gaussians prior to gradient-based optimization, TrackerSplat overcomes the quality degradation associated with large frame gaps when processing multiple adjacent frames in parallel across multiple devices, thereby boosting reconstruction throughput while preserving rendering quality. Experiments on real-world datasets confirm the robustness of TrackerSplat in challenging scenarios with significant displacements, achieving superior throughput under parallel settings and maintaining visual quality compared to baselines. The code is available at https://github.com/yindaheng98/TrackerSplat. Daheng Yin, Isaac Ding, Yili Jin 0001, Jianxin Shi 0005, Jiangchuan Liu |
SIGGRAPH Asia | 4 |
| 2025 | Towards Neural Codec-Empowered 360$^\circ$ Video Streaming: A Saliency-Aided Synergistic ApproachabstractNetworked 360$^\circ$video has become increasingly popular. Despite the immersive experience for users, its sheer data volume, even with the latest H.266 coding and viewport adaptation, remains a significant challenge to today's networks. Recent studies have shown that integrating deep learning into video coding can significantly enhance compression efficiency, providing new opportunities for high-quality video streaming. In this work, we conduct a comprehensive analysis of the potential and issues in applying neural codecs to 360$^\circ$video streaming. We accordingly present$\mathsf {NETA}$, a synergistic streaming scheme that merges neural compression with traditional coding techniques, seamlessly implemented within an edge intelligence framework. To address the non-trivial challenges in the short viewport prediction window and time-varying viewing directions, we propose implicit-explicit buffer-based prefetching grounded in content visual saliency and bitrate adaptation with smart model switching around viewports. A novel Lyapunov-guided deep reinforcement learning algorithm is developed to maximize user experience and ensure long-term system stability. We further discuss the concerns towards practical development and deployment and have built a working prototype that verifies$\mathsf {NETA}$’s excellent performance. For instance, it achieves a 27% increment in viewing quality, a 90% reduction in rebuffering time, and a 64% decrease in quality variation on average, compared to state-of-the-art approaches. Jianxin Shi 0005, Miao Zhang 0003, Linfeng Shen, Jiangchuan Liu, Lingjun Pu, Jingdong Xu |
IEEE Trans. Multim. | 1 |
| 2025 | Streaming Media over LEO Satellite Networking: A Measurement-Based Analysis and OptimizationabstractRecently, Low Earth orbit Satellite Networks (LSNs) have been suggested as a critical and promising component toward high-bandwidth and low-latency global coverage in the upcoming 6G communication infrastructure. SpaceX’s Starlink is arguably the largest and most operational LSN to date. There have been practical uses of Starlink across diverse networked applications, including those with stringent demands, such as multimedia applications. Given the mixed and inconsistent feedback from end users, it remains unclear whether today’s LSNs, in particular Starlink, are ready for realtime multimedia. In this article, we present a systematic measurement study on realtime multimedia services over Starlink, seeking insights into their operations and performance in this new generation of networking. Our findings demonstrate that Starlink can handle most video-on-demand (VoD) and live-streaming services with properly configured buffers but suffers from video pauses or audio cut-offs during interactive videoconferencing. We identify the key factors that impact the performance of LSN, particularly for multimedia services, including satellite switching, routing strategies, and weather conditions. Our findings offer valuable hints into future enhancements for multimedia services over LSNs. Specifically, we further propose a Weather Aware Buffer Based Rate Adaption algorithm based on our observations on weather impacts, which is capable of maximizing the quality of experience for VoD applications with seamless integration of dynamic weather conditions. Hao Fang 0012, Haoyuan Zhao, Feng Wang 0001, Yi Ching Chou, Long Chen 0025, Jianxin Shi 0005, Jiangchuan Liu |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2024 | TailClip: Mitigating Tail Latency in Cloud Gaming via Smart Video Frame GenerationabstractLatency is one of the most significant issues in cloud gaming, among which tail latency, mainly attributed to dynamic network environments (i.e., transmission) and limited device computing capacity (e.g., decoding), has attracted increasing attention. To mitigate the tail latency, different from existing researches considering resource adaptation such as bitrate adaptation, we propose TailClip, whose novel idea is to enable video frame generation at the client if the tail latency is about to happen. TailClip consists of two innovative components: a Deep Reinforcement Learning (DRL) driven tail latency trigger that jointly decides a series of subsequent frame generation regarding multidimensional features of historical tail latency; a lightweight frame generation model derived by adaptive pruning in terms of device computing capacity at runtime. Extensive evaluations indicate the superior performance of TailClip. For example, it can remove the high-latency frames (i.e., over 100 ms) by 77% with an acceptable video quality (i.e., 0.45 dB reduction on average). Qianyun Gong, Kunheng Jiang, Jingjing Wen, Xinjing Yuan, Jianxin Shi 0005, Lingjun Pu |
ICME | 5 |
| 2024 | Robust Live Streaming over LEO Satellite Constellations: Measurement, Analysis, and Handover-Aware AdaptationabstractLive streaming has experienced significant growth recently. Yet this rise in popularity contrasts with the reality that a substantial segment of the global population still lacks Internet access. The emergence of Low Earth orbit Satellite Networks (LSNs), such as SpaceX's Starlink and Amazon's Project Kuiper, presents a promising solution to fill this gap. Nevertheless, our measurement study reveals that existing live streaming platforms may not be able to deliver a smooth viewing experience on LSNs due to frequent satellite handovers, which lead to frequent video rebuffering events. Current state-of-the-art learning-based Adaptive Bitrate (ABR) algorithms, even when trained on LSNs' network traces, fail to manage the abrupt network variations associated with satellite handovers effectively. To address these challenges, for the first time, we introduce Satellite-Aware Rate Adaptation (SARA), a versatile and lightweight middleware that can seamlessly integrate with various ABR algorithms to enhance the performance of live streaming over LSNs. SARA intelligently modulates video playback speed and furnishes ABR algorithms with insights derived from the distinctive network characteristics of LSNs, thereby aiding ABR algorithms in making informed bitrate selections and effectively minimizing rebuffering events that occur during satellite handovers. Our extensive evaluation shows that SARA can effectively reduce the rebuffering time by an average of 39.41% and slightly improve latency by 0.65% while only introducing an overall loss in bitrate by 0.13%. Hao Fang 0012, Haoyuan Zhao, Jianxin Shi 0005, Miao Zhang 0003, Guanzhen Wu, Yi Ching Chou, Feng Wang 0001, Jiangchuan Liu |
ACM Multimedia | 3 |
| 2024 | FSVFG: Towards Immersive Full-Scene Volumetric Video Streaming with Adaptive Feature GridabstractGiven the truly immersive viewing experiences, full-scene volumetric videos have received increasing attention from both academia and industry. Their vast data volumes, however, present significant challenges for real-time streaming over today's bandwidth-limited Internet. Considering the vast amount of full-scene volumetric data to be streamed and the limited bandwidth on the Internet, achieving adaptive full-scene volumetric video streaming over the Internet presents a significant challenge. Inspired by the advantages offered by neural fields, especially the feature grid method, we propose FSVFG, a novel full-scene volumetric video streaming system integrated feature grids as the representation of volumetric content. FSVFG employs an incremental training approach for feature grids and stores the features and residuals between adjacent grids as frames. To support adaptive streaming, we delve into the data structure and rendering processes of feature grids and propose bandwidth adaptation mechanisms. The mechanisms involve a coarse ray-marching for the selection of features and residuals to be sent, and achieve variable bitrate streaming by Level-of-Detail (LoD) and residual filtering. Based on these mechanisms, FSVFG achieves adaptive streaming by adaptively balancing the transmission of feature and residual according to the available bandwidth. Our preliminary results demonstrate the effectiveness of FSVFG, demonstrating its ability to improve visual quality and reduce bandwidth requirements of full-scene volumetric video streaming. Daheng Yin, Jianxin Shi 0005, Miao Zhang 0003, Zhaowu Huang, Jiangchuan Liu, Fang Dong 0001 |
ACM Multimedia | 2 |
| 2024 | Towards Full-scene Volumetric Video Streaming via Spatially Layered Representation and NeRF GenerationabstractImmersive full-scene volumetric video (VV) showcases the richness and detail of the 3D world, yet poses significant streaming challenges given its massive data volume. Existing 3D tile-based viewport approaches struggle to effectively adapt to full-scene VV owing to their small video buffer limitation, high tile segmentation overhead, and lack of full-scene consideration. Jianxin Shi 0005, Miao Zhang 0003, Linfeng Shen, Jiangchuan Liu, Yuan Zhang 0013, Lingjun Pu, Jingdong Xu |
NOSSDAV | 1 |
| 2024 | nHAS: Neural-Compensated Hybrid Adaptive Scheduling for Cloud Gaming
Qianyun Gong, Jiapei Xu, Jianxin Shi 0005, Xinjing Yuan, Jingdong Xu, Guanyu Gao, Lingjun Pu |
NPC (1) | 3 |
| 2024 | : Erasure-Coded Multi-Source Streaming for UHD Videos Within Cloud Native 5G NetworksabstractUltra-High-Definition (UHD) videos have been getting increasing attention. However, existing video streaming solutions fail to deliver them due to the extremely high bandwidth requirement. The emerging cloud native 5G networks have opened up the possibility of enhancing UHD video quality by leveraging in-network video streaming. Unfortunately, the restricted storage and bandwidth of in-network servers could become the main bottleneck. To this end, we present${\sf EMS}$, a novel UHD video streaming framework, by integratingErasure-coded storage withMulti-sourceStreaming. We respectively introduce a deadline-aware and a latency-sensitive metric to indicate the service quality of video servers and advocate a federated learning paradigm for the adaptive service quality update, including a reinforcement learning based multi-server selection (i.e., user local training) and a global service quality aggregation. To facilitate user local training without sacrificing streaming Quality-of-Experience (QoE), we cast the multi-server selection associated with the restriction on the average number of selected servers per video chunk into two kinds of Multi-Armed Bandit (MAB) models in terms of the proposed service quality metrics. We design lightweight Upper Confidence Bound (UCB) based algorithms with a theoretical performance guarantee. We implement a prototype of${\sf EMS}$, and extensive experiments confirm the superiority of the proposed algorithms. Lingjun Pu, Jianxin Shi 0005, Xinjing Yuan, Xu Chen 0004, Lei Jiao 0002, Jingdong Xu |
IEEE Trans. Mob. Comput. | 2 |
| 2023 | Muster: Multi-Source Streaming for Tile-Based 360° Videos Within Cloud Native 5G Networksabstract360° videos generally require a large amount of bandwidth between video servers and users, which puts much burden on the current CDN-based single-source video streaming solutions. The emerging cloud native 5G networks can bridge the distance between video servers and users by leveraging in-network single-source video streaming to enhance 360° video quality. Unfortunately, the restricted bandwidth of in-network servers becomes the main bottleneck. Although tile-based video streaming is promising to reduce video transmission size while keeping user QoE, it highly depends on the accuracy of user FoV prediction, which existing prediction methods cannot guarantee. Recently, some researchers advocate the idea of “super FoV” (i.e., an extended range of predicted FoV) to cope with the inaccurate FoV prediction, which however could lower the effect of tile-based video streaming. Alternatively, we present Muster, a multi-source streaming for tile-based 360° videos within cloud native 5G networks. We detail the system components, provide a comprehensive model, formulate joint server selection and tile requesting problems, and correspondingly propose efficient online algorithms with a performance guarantee. Small-scale testbed and large-scale simulation based evaluation confirm the superiority of the proposed algorithms. Xinjing Yuan, Lingjun Pu, Jianxin Shi 0005, Qianyun Gong, Jingdong Xu |
IEEE Trans. Mob. Comput. | 3 |
| 2022 | Sophon: Super-Resolution Enhanced 360° Video Streaming with Visual Saliency-aware Prefetchabstract360° video streaming requires ultra-high bandwidth to provide an excellent immersive experience. Traditional viewport-aware streaming methods are theoretically effective but unreliable in practice due to the adverse effects of time-varying available bandwidth on the small playback buffer. To this end, we ponder the complementarity between the large buffer-based approach and the viewport-aware strategy for 360°video streaming. In this work, we present Sophon, a buffer-based and neural-enhanced streaming framework, which exploits the double buffer design, super-resolution technique, and viewport-aware strategy to improve user experience. Furthermore, we propose two well-suited ideas: visual saliency-aware prefetch and super-resolution model selection scheme to address the challenges of insufficient computing resources and dynamic user preferences. Correspondingly, we respectively introduce the prefetch and model selection metric, and develop a lightweight buffer occupancy-based prefetch algorithm and a deep reinforcement learning method to trade off bandwidth consumption, computing resource utilization, and content quality enhancement. We implement a prototype of Sophon and extensive evaluations corroborate its superior performance over state-of-the-art works. Jianxin Shi 0005, Lingjun Pu, Xinjing Yuan, Qianyun Gong, Jingdong Xu |
ACM Multimedia | 1 |
| 2021 | EC-360: Speeding Up 360° Video Streaming Using Tile-based Online Erasure Codingabstract360° video services require extremely high bitrate and frame rate videos for a good immersive experience. Traditional solutions for adaptive bitrate streaming are still limited by currently insufficient and fluctuating bandwidth. Besides, viewpoint-aware or tile-based solutions would lead more rebuffering due to the short viewpoint prediction window. In this paper, we present EC-360, a novel video streaming framework to speed up 360° video streaming. Specifically, it creatively integrates tilebased online erasure coding into multi-source content delivery to mitigate “cask effect” or impact of straggler node, which oftentimes leads to failure to improve delivery speed in multisource streaming. We formulate the critical tile-based request scheduling problem in EC-360 based on the Combinatorial Multiarmed Bandit (CMAB) model and develop a low complexity online algorithm. We further theoretically verify the efficiency of the CMAB-based algorithm by deriving its regret upper bound. Extensive experiments based on prototype implementation and real network traces corroborate the efficiency, flexibility, and lightweight of proposed solution. EC-360 achieves superior performance improvement compared to state-of-the-art works in various network scenes and system settings. Jianxin Shi 0005, Lingjun Pu, Jingdong Xu |
GLOBECOM | 1 |
| 2020 | Allies: Tile-Based Joint Transcoding, Delivery and Caching of 360° Videos in Edge Cloud Networksabstract360° or panoramic video applications have seen booming development and absorbed great attention in recent years. However, as they are of significant size and usually watched from a close distance, they require an extremely higher bandwidth and frame rate for a good immersible experience, which poses a great challenge on mobile networks. Realizing the great potentials of tile-based transcoding, viewport adaptive streaming, and edge caching, we propose Allies, a tile-based joint transcoding, delivery, and caching framework for 360° video services in edge cloud networks. Meanwhile, an innovative idea about 360° video caching way is proposed and applied to improve the cache utilization of edge clouds. In this framework, we formulate the joint optimization problem as an integer nonlinear program and propose a greedy suboptimal algorithm with polynomial running time to minimize video service costs. Finally, extensive simulations with real user's head movement traces and corresponding 360° video datasets corroborate the efficiency, flexibility, and lightweight of our proposed algorithm; for instance, it achieves over 25% performance improvement compared to state-of-the-art works in various system settings. Jianxin Shi 0005, Lingjun Pu, Jingdong Xu |
CLOUD | 1 |