VLDB 2026 Research / reviewers in the wild / expert
Jun Xu 0040
dblp:90/514-40
· DBLP profile ↗
8ranked-venue papers
2as first author
7since 2021 · last 2026
0009-0006-2358-4098ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 6 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Volumetric Video on Demand System Based on Scalable Spacetime Gaussian Splatting
Jun Xu 0040, Bingcong Lu, Rong Xie 0004, Li Song 0001 |
ISCAS | 3 |
| 2024 | Pioneer: Offline Reinforcement Learning based Bandwidth Estimation for Real-Time CommunicationabstractFor Real-time Communication (RTC), Bandwidth Estimation (BWE) is crucial for enhancing user Quality of Experience (QoE) by ensuring efficient bandwidth utilization and low latency. Recent advancements have shifted towards machine learning based algorithms, particularly online reinforcement leanring (RL), to dynamically infer future bandwidth using statistical data. However, challenges such as dependency on training settings, the necessity for extensive trial and error, and instability in complex state spaces hinder their efficacy. To address these limitations, we propose Pioneer, a novel offline RL framework for BWE in RTC systems. Unlike its predecessors, Pioneer eliminates the need for real-time environment interaction during training and achieves good performance through lightweight training. Our framework consists of a Trajectory Sampler for state information preprocessing and a Bandwidth Estimator based on offline RL model. Our test results on offline datasets show that Pioneer can achieve better performance than expert algorithms. We also tested Pioneer on online simulation platforms, and Pioneer can improve QoE by 9% compared to other offline algorithm, demonstrating good robustness. Bingcong Lu, Jun Xu 0040, Rong Xie 0004, Li Song 0001, Wenjun Zhang 0001 |
MMSys | 3 |
| 2022 | CNN-Based Fast CU Partitioning Algorithm for VVC Intra CodingabstractOver a year has passed since the finalization of Versatile Video Coding (H.266/VVC), yet it is still far from practical deployment, a major reason being the excessive complexity. The flexible and sophisticated quad-tree with nested multi-type tree partitioning structure in VVC provides considerable performance gains while bringing about an exponential increase in encoding time. To reduce the coding complexity, this paper proposes a Convolutional Neural Network (CNN) based fast Coding Unit (CU) partitioning algorithm for intra coding, which accelerates CU partition through predicting the partition modes with texture information and terminating redundant modes in advance. Corresponding classifiers are designed for different CU sizes to improve prediction accuracy. Low rate-distortion performance degradation is guaranteed by introducing performance loss due to misclassification into the loss function. Experiments show that the proposed method can save encoding time ranging from 38.39% to 62.33% with 0.92% to 2.36% bit rate increase. Jun Xu 0040, Yan Huang 0033, Li Song 0001 |
ICIP | 1 |
| 2022 | A Multi-User Oriented Live Free-Viewpoint Video Streaming System Based on View InterpolationabstractAs an important application form of immersive multimedia services, free-viewpoint video (FVV) enables users with great immersive experience by strong interaction. However, the computational complexity of virtual view synthesis algorithms poses a significant challenge to the real-time performance of an FVV system. Furthermore, the individuality of user interaction makes it difficult to serve multiple users simultaneously for a system with conventional architecture. In this paper, we novelly introduce a CNN-based view interpolation algorithm to synthesis dense virtual views in real time. Based on this, we also build an end-to-end live free-viewpoint system with a multi-user oriented streaming strategy. Our system can utilize a single edge server to serve multiple users at the same time without having to bring a large view synthesis load on the client side. We analyze the whole system and show that our approaches give the user a pleasant immersive experience, in terms of both visual quality and latency. Jingchuan Hu, Shuai Guo 0002, Kai Zhou 0016, Jun Xu 0040, Li Song 0001 |
ICME | 5 |
| 2022 | Complexity-Oriented Per-Shot Video Coding OptimizationabstractCurrent per-shot encoding schemes aim to improve the compression efficiency by shot-level optimization. It splits a source video sequence into shots and imposes optimal sets of encoding parameters on each shot. Per-shot encoding achieved approximately 20% bitrate savings over baseline fixed QP encoding at the expense of pre-processing complexity. However, the adjustable parameter space of the current per-shot encoding schemes only has spatial resolution and QP/CRF, resulting in a lack of encoding flexibility. In this paper, we extend the per-shot encoding framework in the complexity dimension. We believe that per-shot encoding with flexible complexity will help in deploying user-generated content. We propose a rate-distortion-complexity optimization process for encoders and a methodology to determine the coding parameters under the constraints of complexities and bitrate ladders. Experimental results show that our proposed method achieves complexity constraints ranging from 100% to 3% in a dense form compared to the slowest per-shot anchor. With similar complexities of the per-shot scheme fixed in specific presets, our proposed method achieves BDrate gain up to −19.17%. Hongcheng Zhong, Jun Xu 0040, Donghui Feng 0003, Li Song 0001 |
ICME | 2 |
| 2022 | A new free viewpoint video dataset and DIBR benchmarkabstractFree viewpoint video (FVV) has drawn great attention in recent years, which provides viewers with strong interactive and immersive experience. Despite the developments made, further progress of FVV research is limited by existing datasets that mostly have too few number of camera views, or static scenes. To overcome the limitations, in this paper, we present a new dynamic RGB-D video dataset with up to 12 views. Our dataset consists of 13 groups of dynamic video sequences that are taken at the same scene, and a group of video sequences of the empty scene. Each group has 12 HD video sequences taken by synchronized cameras and 12 correspondingly estimated depth video sequences. Moreover, we also introduce a FVV synthesis benchmark on the basis of depth image based rendering (DIBR) to help researchers validate their data-driven methods. We hope our work will inspire more FVV synthesis methods with enhanced robustness, improved performance and deeper understanding. Shuai Guo 0002, Kai Zhou 0016, Jingchuan Hu, Jionghao Wang, Jun Xu 0040, Li Song 0001 |
MMSys | 5 |
| 2022 | Edge-Based Video Compression Texture Synthesis Using Generative Adversarial NetworkabstractIt has been recognized that texture patterns with abundant high-frequency components, such as grass and water, produce visual masking effects, and the distortion in textures is hard to be perceived by human eyes than structure regions. However, modern video codecs in a rate-distortion optimized manner usually consume a lot of bits to encode textures, leading to the insufficiency in perceptual coding performance. Nowadays, with the rapid development of deep learning, learning based texture synthesis methods have been proposed to replace the coding process of prediction residuals to reduce the rate cost. In this paper, we present a deep texture synthesizer named edge-based texture synthesis framework (ETSF). At encoder side, the framework detects texture regions by semantic and fidelity classification criteria, and the detected regions are quantized coarsely by the hybrid coding framework. In texture characterization, ETSF extracts low-level edge features representing pixel intensity variation. Feature processing tools are developed to remove the spatiotemporal redundancy of edges. The processed edge information is compressed and transmitted. To effectively recover textures, we design an edge-based texture synthesis generative adversarial network (ETSGAN) at the decoder of ETSF, which can incorporate edge information into convolutional layers and generate realistic textures. Experimental results on a collected texture dataset show that the proposed ETSF can achieve an average of -12.8%, -14.2% and -9.6% MOS BD-rate under lowdelay_B, lowdelay_P and random_access configurations of VVC coding, respectively. Jun Xu 0040, Donghui Feng 0003, Rong Xie 0004, Li Song 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2007 | Bit Allocation for Fine-Granular SNR Scalability Coding with Hierarchical B PicturesabstractHierarchical B pictures are devised to achieve temporal scalability in the scalable extension of H.264/AVC (SVC) which is under standardization. The fine-granular SNR scalability (FGS) can be provided by progressive refinement (PR) slices in SVC. In this paper, we firstly investigate error propagation in the case of discarding PR slices and obtain a rate difference distortion optimization criterion to improve coding efficiency of base layer. Then we consider the full rate case and propose a rate distortion slope criterion to enhance FGS coding efficiency at high rate. Finally the criterion to boost coding efficiency in the whole range of FGS rate is derived by combing the criterions derived previously. The proposed method is compared to the approach in SVC test model and up to 0.3dB coding gains are achieved. Jun Xu 0040, Li Song 0001, Shibao Zheng, Xiaokang Yang 0001, Rong Xie 0004 |
ICME | 1 |