VLDB 2026 Research / reviewers in the wild / expert
Jie Li 0015
dblp:17/2703-15
· DBLP profile ↗
21ranked-venue papers
12as first author
13since 2021 · last 2025
0000-0001-8483-6240ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 10 · 8 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 6 first-author · 8 since 2021Systems, architecture and hardware · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ViewGauss: A Head Movement Dataset for 6DoF Gaussian Splatting Video ViewingabstractGaussian splatting video has recently emerged as a promising representation for immersive 6-degree-of-freedom (6DoF) content due to its low-latency rendering, compact data structure, and high visual fidelity. In particular, 4D Gaussian splatting video-which models dynamic scenes as temporally evolving Gaussian splats in 3D space-offers an efficient solution for rendering photorealistic, interactive experiences. However, a systematic understanding of user behavior in such environments, especially head movement, remains largely unexplored due to the absence of dedicated datasets tailored to this format. This lack of data severely limits progress in viewpoint prediction, attention modeling, and video streaming optimization. To address this critical gap, we introduce ViewGauss-the first publicly available dataset that captures full 6DoF head movement during the viewing of 4D Gaussian splatting videos. Our dataset is collected from 35 participants using a high-precision Vive Focus Vision headset in a controlled environment, while they freely watched four reconstructed Gaussian splatting video sequences derived from the HiFi4G dataset. The data are recorded with high temporal resolution using position coordinates and unit quaternions, and organized into structured CSV files with precise timestamps for downstream synchronization and behavioral analysis. To demonstrate the practical value of ViewGauss, we conduct a preliminary viewpoint prediction experiment using the iTransformer model. The results show that head orientation patterns in 4D Gaussian splatting video scenes are not only temporally coherent but also learnable, highlighting the potential of ViewGauss as a benchmark for future behavioral modeling and predictive rendering systems. The dataset is publicly available at: https://github.com/Cedarleigh/ViewGauss-DataSet. Zhixia Zhao, Qiyue Li 0001, Jie Li 0015, Richang Hong, Zhi Liu 0002 |
ACM Multimedia | 3 |
| 2025 | PCVD: A Dataset of Point Cloud Video for Dynamic Human InteractionabstractPoint cloud is widely used in computer vision and augmented reality for representing 3D information of real-world scenes. However, challenges such as noise, incompleteness, and quality variations, particularly in dynamic environments, hinder effective processing and analysis. These issues are further complicated in human activity scenarios due to motion and changing lighting. To address these challenges, this paper introduces a point cloud video dataset PCVD captured with synchronized Azure Kinect cameras, designed to support tasks like denoising, segmentation, and motion recognition in single and multi-person scenes. It provides high-quality depth and color data from diverse real-world scenes with human actions. We compare it with existing datasets, and the results show its superiority in uniformity and completeness, making it ideal for dynamic environments. We also evaluate the state-of-the-art denoising schemes on this dataset to demonstrate the practicality and sophistication of the dataset. The dataset is publicly available at https://github.com/Atlantichan/PCVD-A-Dataset-of-Point-Cloud-Video-for-Dynamic-Human-Interaction. Jie Li 0015, Shujiao Chen, Qiyue Li 0001, Zhi Liu 0002 |
MMSys | 1 |
| 2025 | Viewport Prediction for Volumetric Video Streaming by Exploring Video Saliency and User Trajectory InformationabstractVolumetric video, also referred to as hologram video, is an emerging medium that represents 3D content in extended reality. As a next-generation video technology, it is poised to become a key application in 5G and future wireless communication networks. Because each user generally views only a specific portion of the volumetric video, known as the viewport, accurate prediction of the viewport is crucial for ensuring an optimal streaming performance. Despite its significance, research in this area is still in the early stages. To this end, this paper introduces a novel approach called Saliency and Trajectory-based Viewport Prediction (STVP), which enhances the accuracy of viewport prediction in volumetric video streaming by effectively leveraging both video saliency and viewport trajectory information. In particular, we first introduce a novel sampling method, Uniform Random Sampling (URS), which efficiently preserves video features while minimizing computational complexity. Next, we propose a saliency detection technique that integrates both spatial and temporal information to identify visually static and dynamic geometric and luminance-salient regions. Finally, we fuse saliency and trajectory information to achieve more accurate viewport prediction. Extensive experimental results validate the superiority of our method over existing state-of-the-art schemes. To the best of our knowledge, this is the first comprehensive study of viewport prediction in volumetric video streaming. We also make the source code of this work publicly available. Jie Li 0015, Zhi Liu 0002, Peng Yuan Zhou, Richang Hong, Qiyue Li 0001, Han Hu 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | VPFormer: Leveraging Transformer with Voxel Integration for Viewport Prediction in Volumetric VideoabstractWith the continuous advancement of computer vision, image processing technologies, volumetric video, represented by point cloud videos, holds the potential for extensive applications in areas such as Virtual Reality (VR) and Augmented Reality (AR). Viewport prediction, also referred to as Field of View (FoV) prediction, is a crucial component in emerging VR and AR applications, playing a vital role in the transmission of point cloud videos. Currently, models for viewpoint prediction that integrate feature extraction and FoV information heavily rely on the spatial-temporal features extracted by convolutional neural networks. However, the drawback of 3D convolution lies in its inability to effectively capture long-term spatial-temporal dependencies within videos. Moreover, the temporal contrast layer used for time feature extraction only compares features within each block, leading to matching errors and inaccurate temporal feature extraction, consequently diminishing predictive performance. To address these limitations, we propose a Transformer-based Volumetric Point Cloud Video Viewport Prediction Network (VPFormer) that can efficiently extract spatial-temporal features from point cloud videos. VPFormer constitutes a viewport prediction framework that combines the spatial-temporal features of point cloud videos with user trajectory information. Specifically, we introduce a novel sampling method that effectively preserves spatial-temporal information while reducing computational complexity. Additionally, we incorporate context-aware dynamic positional encoding to capture inter-frame spatial-temporal context information. Subsequently, we introduce a voxel-based temporal contrast layer and partition the point cloud into smaller voxel blocks during feature matching, significantly reducing matching errors and enhancing the analysis and extraction of temporal features. Finally, by combining the spatial-temporal features of point cloud videos with user head trajectory information, we successfully predict future user viewpoints. Experimental results demonstrate that this approach outperforms other solutions in terms of performance. Jie Li 0015, Zhixia Zhao, Qiyue Li 0001, Peng Yuan Zhou, Zhi Liu 0002, Hao Zhou 0001, Zhu Li 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2023 | Demo: Landscape: Saliency and Trajectory based Viewport Prediction in Point Cloud Video StreamingabstractEfficient point cloud video streaming requires accurate viewport prediction, and research on this topic is still in its infancy. This paper demonstrates a high-precision scheme for viewport prediction in the point cloud video, named Landscape, exploring both video saliency information and viewport trajectory. Specifically, we first propose a novel point cloud video sampling method, which reduces computational load while preserving video features. Furthermore, we introduce a new saliency detection technique that integrates temporal and spatial information to detect dynamic, static geometric, and color salient regions. Finally, we intelligently fuse saliency and trajectory information to achieve more accurate viewport prediction. We verify the performance of our proposed viewport prediction methods over state-of-the-art wireless networks. Jie Li 0015, Qiyue Li 0001, Wei Sun 0011, Zhi Liu 0002 |
MobiSys | 1 |
| 2023 | Demo: Horizon: a Real-time Point Cloud Video Streaming System over Wireless NetworksabstractAs a popular way of representing holographic video or volumetric video, point cloud video can provide users with a highly immersive viewing experience of 6 degrees of freedom (6DoF) and is expected to become the mainstream video format of the future. However, the real-time transmission of point cloud video faces many challenges due to the huge amount of data and the large search space of the optimization problem with constraints. To this end, we propose Horizon, a novel Dynamic Adaptive Streaming over HTTP (DASH) based real-time point cloud video streaming system, which aims to maximize the user's viewing experience by predicting the next several steps through a rolling framework and uses a Deep Reinforcement Learning (DRL) based algorithm to achieve a real-time solution to the rolling optimization problem. We have prototyped this system and demonstrated its performance on a state-of-the-art wireless network. Jie Li 0015, Qiyue Li 0001, Xin Liu 0104, Zhi Liu 0002 |
MobiSys | 1 |
| 2023 | Toward Optimal Real-Time Volumetric Video Streaming: A Rolling Optimization and Deep Reinforcement Learning Based ApproachabstractVolumetric video provides users with a good viewing experience of six degrees of freedom (DoF) and has wide applications in many fields such as teleconferencing and online games. However, the huge data volume and strict latency requirements of point cloud video, the most popular representative of volumetric video, pose a challenge to its transmission. Existing point cloud video transmission algorithms usually segment a long video by every one or several group of frames, predict network bandwidth and field of view (FoV) information, then perform adaptive transmission by solving the quality of experience (QoE) optimization problem. However, such segmentation neglects the impact of current optimization decisions on the subsequent video streaming process, as well as the accumulated prediction error across a long interval, severely degrading user’s QoE. Moreover, the complex constrained optimization problem makes the solution time too long to meet the real-time video streaming requirements. To this end, in this paper, we propose a rolling prediction-optimization-transmission (POT) framework, which makes predictions of network bandwidth and FoV in each short rolling window to reduce prediction error. And our framework takes into account the upper bounded QoE contribution of the subsequent point cloud video to improve the system performance. In addition, we design a deep reinforcement learning based real-time solver to make decisions for the fixed structure optimization problem in each roll, allowing our system to run in real-time. We have performed simulations and experiments, and the results show that our solution outperforms existing methods. Jie Li 0015, Zhi Liu 0002, Peng Yuan Zhou, Xianfu Chen, Qiyue Li 0001, Richang Hong |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Optimal Volumetric Video Streaming With Hybrid Saliency Based TilingabstractVolumetric video enables a six-degree-of-freedom (6DoF) immersive viewing experience and has a wide range of applications in entertainment and education, among others. Most existing approaches to volumetric video streaming are extensions of VR video streaming solutions that do not take into account user behavior and the properties of the video during the tiling process, and the complexity of decoding is high. To this end, we study volumetric video streaming in this paper and address the research questions mentioned above. In particular, we first propose a hybrid visual saliency and hierarchical clustering empowered 3D tiling scheme that better matches the user’s field of view (FoV). Then, we build a quality of experience (QoE) model considering the volumetric video features as the optimization objective. In addition to the usual encoded version, we introduce the reconstructed version (i.e., decoded version, which allows the user to skip the decoding process and thus reduces the decoding overhead) and propose a joint computational and communication resource allocation scheme to achieve a trade-off between communication and computational resources to maximize the QoE. We perform exhaustive simulations and build a prototype system to verify the performance of the proposed tiling and transmission scheme. The results show that the proposed tiling and transmission scheme performs significantly better than the comparison schemes. Jie Li 0015, Cong Zhang 0002, Zhi Liu 0002, Richang Hong, Han Hu 0003 |
IEEE Trans. Multim. | 1 |
| 2023 | Spherical Convolution Empowered Viewport Prediction in 360 Video Multicast with Limited FoV FeedbackabstractField of view (FoV) prediction is critical in 360-degree video multicast, which is a key component of the emerging virtual reality and augmented reality applications. Most of the current prediction methods combining saliency detection and FoV information neither take into account that the distortion of projected 360-degree videos can invalidate the weight sharing of traditional convolutional networks nor do they adequately consider the difficulty of obtaining complete multi-user FoV information, which degrades the prediction performance. This article proposes a spherical convolution-empowered FoV prediction method, which is a multi-source prediction framework combining salient features extracted from 360-degree video with limited FoV feedback information. A spherical convolutional neural network is used instead of a traditional two-dimensional convolutional neural network to eliminate the problem of weight sharing failure caused by video projection distortion. Specifically, salient spatial-temporal features are extracted through a spherical convolution-based saliency detection model, after which the limited feedback FoV information is represented as a time-series model based on a spherical convolution-empowered gated recurrent unit network. Finally, the extracted salient video features are combined to predict future user FoVs. The experimental results show that the performance of the proposed method is better than other prediction methods. Jie Li 0015, Qiyue Li 0001, Zhi Liu 0002 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2022 | An Energy Efficient Uplink Scheduling and Resource Allocation for M2M Communications in SC-FDMA Based LTE-A Networks
Qiyue Li 0001, Yuling Ge, Yangzhao Yang, Yadong Zhu, Wei Sun 0011, Jie Li 0015 |
Mob. Networks Appl. | 6 |
| 2022 | Resource Orchestration of Cloud-Edge-based Smart Grid Fault DetectionabstractReal-time smart grid monitoring is critical to enhancing resiliency and operational efficiency of power equipment. Cloud-based and edge-based fault detection systems integrating deep learning have been proposed recently to monitor the grid in real time. However, state-of-the-art cloud-based detection may require uploading a large amount of data and suffer from long network delay, while edge-based schemes do not adequately consider the detection requirement and thus cannot provide flexible and optimal performance. To solve these problems, we study a cloud-edge based hybrid smart grid fault detection system. Embedded devices are placed at the edge of the monitored equipment with several lightweight neural networks for fault detection. Considering limited communication resources, relatively low computation capabilities of edge devices, and different monitoring accuracies supported by these neural networks, we design an optimal communication and computational resource allocation method for this cloud-edge based smart grid fault detection system. Our method can maximize the processing throughput of the system and improve resource utilization while satisfying the data transmission and processing latency requirements. Extensive simulations are conducted and the results show the superiority of the proposed scheme over comparison schemes. We have also prototyped this system and verified its feasibility and performance in real-world scenarios. Jie Li 0015, Yuxing Deng, Wei Sun 0011, Ruidong Li 0001, Qiyue Li 0001, Zhi Liu 0002 |
ACM Trans. Sens. Networks | 1 |
| 2021 | An Optimal Uplink Scheduling in Heterogeneous PLC and LTE Communication for Delay-aware Smart Grid Applications
Qiyue Li 0001, Wei Sun 0011, Jinjin Ding, Guojun Luo, Jie Li 0015 |
Mob. Networks Appl. | 7 |
| 2021 | Evaluation on visualization methods of dynamic collaborative relationships for project management
Qiang Lu 0002, Xiaohui Yuan 0001, Jie Li 0015 |
Vis. Comput. | 5 |
| 2020 | Joint Communication and Computational Resource Allocation for QoE-driven Point Cloud Video StreamingabstractPoint cloud video is the most popular representation of hologram, which is the medium to precedent natural content in VR/AR/MR and is expected to be the next generation video. Point cloud video system provides users immersive viewing experience with six degrees of freedom (6DoF) and has wide applications in many fields such as online education and entertainment. To further enhance these applications, point cloud video streaming is in critical demand. The inherent challenges lie in the large size by the necessity of recording the three-dimensional coordinates besides color information, and the associated high computation complexity of encoding/decoding. To this end, this paper proposes a communication and computational resource allocation scheme for QoE-driven point cloud video streaming. In particular, with the goal to maximize the defined QoE by selecting proper quality levels (uncompressed tiles at different quality levels are also considered) for each partitioned point cloud video tile, we formulate this into an optimization problem under the limited communication and computational resources constraints and propose a scheme to solve it. Extensive simulations are conducted and the simulation results show the superior performance of the proposed scheme over the existing schemes. Jie Li 0015, Cong Zhang 0002, Zhi Liu 0002, Wei Sun 0011, Qiyue Li 0001 |
ICC | 1 |
| 2020 | ElectricVIS: visual analysis system for power supply data of smart city
Qiang Lu 0002, Haibo Zhang 0007, Qingpeng Tang, Jie Li 0015 |
J. Supercomput. | 5 |
| 2019 | Expectation-based 3D edge bundling
Guibing Yang, Kunle Ma, Xiaohui Yuan 0001, Jie Li 0015, Qiang Lu 0002 |
Multim. Tools Appl. | 4 |
| 2018 | Modeling QoE of Virtual Reality Video Transmission over Wireless NetworksabstractVirtual Reality (VR) provides an immersive 360 viewing experience and has been widely used in vast areas such as education, entertainment and training. To further widen its applications, networked 360 VR video becomes essential. Quality of Experience (QoE), which objectively measures user experience, is vital for 360 VR video transmission mechanism design. However, to the best of our knowledge, there are few subjective QoE metric for 360 VR video transmission over wireless networks. In this paper, we aim to fill this gap by proposing a general QoE model based on subjective quality evaluation experiments. First, the state-of-the-art 360 VR video processing and wireless transmission schemes are used to conduct subjective experiments according to the international standard. Then, how user experience is affected by different factors, including users' viewing angle, tiling (how the 360 VR video is partitioned into smaller parts to facilitate transmission), stall and resolution switch, is analyzed mathematically. A general QoE model is finally proposed to facilitate the future 360 VR video streaming mechanism design. Jie Li 0015, Ransheng Feng, Zhi Liu 0002, Wei Sun 0011, Qiyue Li 0001 |
GLOBECOM | 1 |
| 2017 | Cramér-Rao Bound Analysis of Wi-Fi Indoor Localization Using Fingerprint and Assistant NodesabstractLocation estimation in Wi-Fi environment has gained considerable attention over the past years, and the Cramer-Rao Lower Bound (CRLB) can be used to evaluate the performance of the localization system. In this paper, we analyze the CRLB of Wi- Fi indoor localization using fingerprint and assistant nodes. This localization method combines received signal strength (RSS) and Time of Arrival (TOA) into together, and constructs a fixed spatial model with several assistant nodes to improve localization performance. There are two purposes of the CRLB analysis framework proposed in this paper. Firstly, the expression of lower bound on location estimation error can help in designing and refining efficient localization algorithm and parameters. Secondly, the error trends can provide suggestions for a positioning system design and deployment. Furthermore, detailed analysis as well as experimental results are both presented in this paper. Qiyue Li 0001, Wei Li 0092, Wei Sun 0011, Jie Li 0015, Zhi Liu 0002 |
VTC Fall | 4 |
| 2016 | A Correlation-Based Energy Balanced Probabilistic Flooding Algorithm in Wireless Sensor NetworkabstractThe costly explicit and implicit acknowledgements (ACKs) are issues that need to be addressed in the existing reliability aware flooding algorithms. This research focuses on energy efficiency on both data transmission and ACKs, while achieving target reliability and balancing the residual energy of sensor nodes. A correlation-based probabilistic flooding algorithm (CPFA) is proposed. It exploits the link correlation between neighbors and tracks aggregate ACKs to decide whether or not to retransmit a packet. Simulation is carried out to reveal that our proposed scheme saves more than 50% energy on explicit and implicit ACKs in most cases while balancing the residual energy of sensor nodes. Qiyue Li 0001, Huihui Rong, Wei Sun 0011, Jianping Wang 0002, Jie Li 0015 |
VTC Spring | 5 |
| 2016 | Joint MCS and power allocation for SVC video multicast over heterogeneous cellular networks
Jie Li 0015, Zhongming Bao, Chenxiang Zhang, Qiyue Li 0001, Zhi Liu 0002 |
Comput. Commun. | 1 |
| 2015 | A Dynamic State Estimation of Power System Harmonics Using Distributed Related Kalman Filter
Wei Sun 0011, Chanjuan Zhao, Jianping Wang 0002, Chenghui Zhu, Daoming Mu, Liangfeng Chen, Jie Li 0015, Qiyue Li 0001 |
ICA3PP (1) | 7 |