EDBT 2026 Demo / reviewers in the wild / expert
Junqi Liao
dblp:352/9880
· DBLP profile ↗
14ranked-venue papers
7as first author
14since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 9 since 2021Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Computer networks · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Task-Oriented Video Compression via XAI-Guided Frame FilteringabstractTask-oriented video compression aims to eliminate redundancy while preserving task-critical information. However, existing spatial domain methods incur high computational overhead, whereas temporal domain approaches often rely on black-box confidence scores that may discard taskrelevant frames. These methods follow an implicit informationpreservation paradigm that entangles redundancy and relevance, yielding opaque decision-making. In this paper, we therefore reformulate it as an explicit and explainable frame filtering method. Instead of relying on black-box confidence scores, task relevance is quantified by an explainable saliency score grounded in explicit model attribution and statistical aggregation, providing a transparent criterion for frame filtering. This offline-defined saliency metric supervises a lightweight online regressor, so that each online decision inherits the same explainable semantic meaning while avoiding heavy computation at the edge. Specifically, frame filtering is decomposed into two sequential and explainable steps: a similarity detector first removes structurally redundant frames, and a saliency-guided filter then discards frames with limited contribution to the downstream task. Experimental results demonstrate the effectiveness of the proposed method. Jingyue Tang, Junqi Liao, Lindong Zhao, Xin Wei 0001 |
IEEE Signal Process. Lett. | 3 |
| 2026 | USTC-TD: A Test Dataset and Benchmark for Image and Video Coding in 2020sabstractImage/video coding has been a remarkable research area for both academia and industry for many years. Testing datasets, especially high-quality image/video datasets, are desirable for the justified evaluation of coding-related research, practical applications, and standardization activities. We put forward a test dataset, namely USTC-TD, which has been successfully adopted in the practical end-to-end image/video coding challenge ofIEEE International Conference on Visual Communications and Image Processing (VCIP)in 2022 and 2023. USTC-TD contains 40 images at 4K spatial resolution and 10 video sequences at 1080p spatial resolution, featuring various content due to the diverse environmental factors (e.g., scene type, texture, motion, view) and the designed imaging factors (e.g., illumination, lens, shadow). We quantitatively evaluate USTC-TD on different image/video features (spatial, temporal, color, lightness), and compare it with the previous image/video test datasets, which verifies its excellent compensation for the shortcomings of existing datasets. We also evaluate both classic standardized and recently learned image/video coding schemes on USTC-TD using objective quality metrics (PSNR, MS-SSIM, VMAF) and subjective quality metric (MOS), providing an extensive benchmark for these evaluated schemes. Based on the characteristics and specific design of the proposed test dataset, we analyze the benchmark performance and shed light on the future research and development of image/video coding. All the data are released online:https://esakak.github.io/USTC-TD. Zhuoyuan Li 0001, Junqi Liao, Chuanbo Tang, Haotian Zhang 0009, Yifan Bian, Xihua Sheng, Xinmin Feng, Yao Li 0016, Changsheng Gao, Li Li 0040, Dong Liu 0002, Feng Wu 0005 |
IEEE Trans. Multim. | 2 |
| 2025 | LargeSceneGaussian: High-Efficiency 3D Gaussian Splatting for Large-Scale Scene ReconstructionabstractRecently, significant progress has been made in the field of novel view synthesis. 3D Gaussian Splatting has shown high training efficiency and real-time rendering capabilities in scene modeling. However, large-scale and complex unbounded scenes pose challenges such as data initialization, view dependency, and high memory overhead, affecting rendering efficiency and visual quality. To address these issues, we propose LargeSceneGaussian (LSG), a novel method for efficient large-scale scene reconstruction. LSG incorporates three key strategies: (1) Denoising sparse point cloud data from drone-captured images to enhance quality, (2) K-Means-based partitioning of the scene into balanced sub-blocks for independent training and seamless merging, and (3) View filtering to remove redundant camera views, improving efficiency. The LSG achieves 15% lower memory usage, 10% faster training, and 20% less redundant view data, significantly enhancing rendering quality and efficiency for large-scale 3D modeling. Yifeng Ge, Junqi Liao, Deke Tang |
ICIP | 2 |
| 2025 | EFVC: Error-Propagation-Free Neural Video Coding with Reversible TransformabstractNeural video codecs (NVCs) have received more and more attention. Although NVCs’ performance shows a trend of surpassing traditional video coding, challenges still hinder their practical application. One significant challenge is the error propagation in long prediction chains. Many existing NVCs alleviate the error propagation but do not solve it. In this paper, we analyze that the causes of error propagation are train-test mismatch and explicit hierarchical quality scheme absence in NVC. Moreover, in response to the Grand Challenge on Neural Network-based Video Coding at ISCAS 2025, we propose the first error-propagation-free NVC, EFVC. The EFVC introduces a reversible transform backbone to eliminate the unstable network information loss caused by train-test mismatch. Then, a hierarchical quality strategy is introduced to constrain hierarchical quality explicitly. Furthermore, to eliminate the error propagation further, we restrict the nonlinearity. The experimental results show that EFVC effectively eliminates error propagation. In addition, our proposed EFVC leads to a 21.8% BD-rate reduction on average compared with the reproduced SOTA NVC, DCVC-DC. Junqi Liao, Li Li 0040, Dong Liu 0002, Houqiang Li |
ISCAS | 1 |
| 2025 | IVCA: Inter-relation-aware Video Complexity AnalyzerabstractTo address the real-time analysis requirements of video streaming applications, we propose an innovative inter-relation-aware video complexity analyzer (IVCA) to enhance the existing video complexity analyzer (VCA). The IVCA overcomes the limitations of the VCA by incorporating inter-frame relations, focusing on inter motion and reference structure. To begin with, we improve the accuracy of temporal features by integrating feature-domain motion estimation into the IVCA framework, which allows for a more nuanced understanding of motion across frames. Furthermore, inspired by the hierarchical reference structures utilized in modern codecs, we introduce layer-aware weights that effectively adjust the contributions of frame complexity across different layers, ensuring a more balanced representation of video characteristics. In addition, we broaden the analysis of temporal features by considering reference frames rather than relying solely on the preceding frame, thereby enriching the contextual understanding of video content. Experimental results demonstrate a significant enhancement in complexity estimation accuracy achieved by the IVCA, coupled with a negligible increase in time complexity, indicating its potential for real-time applications in video streaming scenarios. This advancement not only improves video processing efficiency but also paves the way for more sophisticated analytical tools in video technology. Junqi Liao, Yao Li 0016, Zhuoyuan Li 0001, Li Li 0040, Dong Liu 0002 |
ISCAS | 1 |
| 2025 | EHVC: Efficient Hierarchical Reference and Quality Structure for Neural Video CodingabstractNeural video codecs (NVCs), leveraging the power of end-to-end learning, have demonstrated remarkable coding efficiency improvements over traditional video codecs. Recent research has begun to pay attention to the quality structures in NVCs, optimizing them by introducing explicit hierarchical designs. However, less attention has been paid to the reference structure design, which fundamentally should be aligned with the hierarchical quality structure. In addition, there is still significant room for further optimization of the hierarchical quality structure. To address these challenges in NVCs, we propose EHVC, an efficient hierarchical neural video codec featuring three key innovations: (1) a hierarchical multi-reference scheme that draws on traditional video codec design to align reference and quality structures, thereby addressing the reference-quality mismatch; (2) a lookahead strategy to utilize an encoder-side context from future frames to enhance the quality structure; (3) a layer-wise quality scale with random quality training strategy to stabilize quality structures during inference. With these improvements, EHVC achieves significantly superior performance to the state-of-the-art NVCs. Code will be released in: https://github.com/bytedance/NEVC. Junqi Liao, Yaojun Wu 0001, Chaoyi Lin, Zhipin Deng, Li Li 0040, Dong Liu 0002, Xiaoyan Sun 0001 |
ACM Multimedia | 1 |
| 2025 | Cross-Modal Semantic Transmission Strategy for Mobile ScenariosabstractTo fulfill the demands of emerging multi-modal services, the cross-modal semantic communication paradigm comes into being. It fully utilizes potential semantic correlations among modalities to address polysemy and ambiguity issues, enhancing transmission reliability. However, applying cross-modal semantic communication in resource-constrained mobile scenarios introduces new challenges, including radio spectrum bandwidth limitations and fluctuations for the transmitter, and computing resource constraints for the receiver, which leads to potential transmission failures. To bridge this gap, this paper proposes a cross-modal semantic transmission strategy for mobile scenarios (MobileCMST). We first construct the framework for MobileCMST. Within this framework, a semantic encoder is designed to achieve redundancy elimination for visual and haptic signals. Then, a semantic delivery approach is developed to cope with bandwidth fluctuations and multipath fading channels. Finally, an efficient semantic decoder based on a visual-haptic semantic-integrated diffusion model is proposed. It employs the Mamba backbone to reconstruct high-quality signals with lightweight computational complexity. Extensive experiments demonstrate the excellent performance of the proposed MobileCMST strategy in resource-constrained mobile scenarios. Junqi Liao, Xin Wei 0001, Liang Zhou 0002, Weihua Zhuang |
IEEE Trans. Commun. | 1 |
| 2024 | Practical Learned Image Compression with Online Encoder OptimizationabstractLearned image compression methods have shown significant advances in performance. However, they often suffer from higher decoding complexity compared to traditional codecs. In this paper, we present an approach toward practical learned image compression that focuses on faster decoding by employing online encoder optimization. To reduce network complexity, we employ a hyperprior structure with smaller convolution kernels, which enables efficient compression. We also introduce a content adaptive skipping algorithm to accelerate entropy decoding. By jointly optimizing entropy decoding complexity and the distortion of the reconstructed image at the encoder side, we achieve faster decoding while maintaining the quality of the reconstructed image. Furthermore, we incorporate online iterative optimization at the encoder side to enhance rate-distortion performance without increasing decoding complexity. This approach allows us to fine-tune the latent and improve overall performance. Our method achieves a notable improvement over BPG on the VCIP 2023 challenge test set, with more than 1 dB increase in PSNR and a 50% reduction in decoding time. Haotian Zhang 0009, Feihong Mei, Junqi Liao, Li Li 0040, Houqiang Li, Dong Liu 0002 |
PCS | 3 |
| 2024 | Frame Level Content Adaptive λ for Neural Video CompressionabstractNeural video compression (NVC) methods have made significant advances in recent years. In most NVC methods, all frames share the same Rate-Distortion trade-off parameter λ, which might be sub-optimal. Recently, inspired by traditional video codecs' hierarchical quality structure, DCVC-DC proposed allocating periodic weights to λ to equip NVC with the hierarchical quality structure. However, the inspiration from traditional video codecs is designed to complement their complex reference structure. Compared to traditional video codecs, NVC methods' reference structure is much simpler and may not require large fluctuations in their hierarchical quality structure. Moreover, DCVC-DC's fixed hierarchical quality structure ignored the influence of video content. We conduct an elaborate study on the hierarchical quality structure in DCVC-DC, shedding light on the potential for improving compression performance by proposing a content adaptive λ to achieve a more reasonable hierarchical quality structure based on the fixed hierarchical weights. Experimental results demonstrate that the proposed method achieves a better rate-distortion performance than allocating the fixed weights to the fixed λ. On DCVC-DC and DCVC-SDD, we achieved 4.9% and 8.8% bdrate reduction with our method. Zhirui Zuo, Junqi Liao, Xiaomin Song, Huiming Zheng, Dong Liu 0002 |
VCIP | 2 |
| 2024 | Cross-Attention and Cycle-Consistency-Based Haptic to Image InpaintingabstractWith the rapid advancement of deep learning and multimedia technologies, image inpainting has made significant progress in generating desirable content for damaged images. To fill in the corrupted part of the image, existing methods either utilize intrinsic information within the visual modality for single-modal image inpainting or consider semantic correlations from non-visual modalities for cross-modal image inpainting. However, when the corrupted part is extensive, intrinsic information alone within visual modality is insufficient to inpaint the entire image. Cross-modal image inpainting methods cannot guarantee image quality, as they do not explore the relationship among the corrupted part of the image, the remaining part of the image, and the non-visual modality. In order to handle this issue, a haptic to image inpainting scheme is proposed. Specifically, feature extraction and cross-attention-based feature index are firstly constructed to find the correspondence between the corrupted part of the image and the haptic modality. Then, cycle-consistency-based feature translation is implemented to explore the intrinsic correlations among the above two modalities, facilitating the formation of the corrupted part. Finally, the formed corrupted part is combined with the remaining part to realize image inpainting. Experimental results show the effectiveness of the proposed scheme. Junqi Liao, Xin Wei 0001 |
IEEE Signal Process. Lett. | 1 |
| 2024 | Toward Generic Cross-Modal Transmission StrategyabstractMulti-modal services, integrating various modalities such as audio, visual, and haptic, have emerged as leading multimedia applications in the 5G era and beyond. To fulfill the demands for low latency, high reliability, and large capacity, cross-modal transmission schemes have been proposed. Typically, these schemes emphasize on either audio-visual or haptic modality, and prioritize flawless transmission of one modality to assist the other modality streaming. However, these prerequisite and assumption do not hold for generic multi-modal services and communication environments, where determining the priority of modality and guaranteeing flawless transmission becomes challenging. To address this fundamental problem, in this paper, we introduce a strategy toward generic cross-modal transmission, enabling visual and haptic modalities to assist each other as needed. The strategy includes a visual-haptic mutual stream delivery mechanism at the sender and a visual-haptic mutual signal reconstruction approach at the receiver. The former aims to eliminate redundancy in visual and haptic streams through mutual assistance, while the latter adaptively handles impaired, missing, or delayed visual or haptic signals by leveraging modality-aware knowledge transfer and semantic-aware signal generation techniques. The proposed strategy demonstrates excellent performance through experiments conducted on a standard multi-modal dataset and a practical visual-haptic communication platform. Xin Wei 0001, Junqi Liao, Liang Zhou 0002, Hikmet Sari, Weihua Zhuang |
IEEE Trans. Commun. | 2 |
| 2024 | Content-Adaptive Rate-Distortion Modeling for Frame-Level Rate Control in Versatile Video CodingabstractRate control (RC) plays an essential role in video coding. RC algorithms based on more accurate rate-distortion (R-D) models often achieve higher control precision and better R-D performance. The most widely used R-D model for HEVC and VVC is the hyperbolic R-D model, which is equivalent to a linear relationship between$\ln {R}$and$\ln {D}$. Due to its high accuracy, few studies have attempted to improve the accuracy of the hyperbolic R-D model further. Intuitively, we consider that the accuracy of the hyperbolic R-D model could be further improved by increasing the order of the R-D model. However, this may also increase the number of model parameters to be estimated, which may not benefit the one-pass RC precision. In this paper, we first explicitly note that there is a trade-off between the order of the R-D model and the difficulty in estimating the model parameters in one-pass RC. Then, motivated by the trade-off, we propose high-order R-D models and the corresponding one-pass frame-level RC algorithms for video coding. Finally, we introduce the quadratic R-D model into frame-level RC in the VVC Test Model (VTM-19.2) and provide a content-adaptive model selection between the first-order and second-order R-D models. Experimental results show that the proposed frame-level RC algorithm based on the quadratic R-D model reduces the average frame-level bitrate error by 5.30%, 3.11%, and 13.45% and achieves 0.49%, 0.65%, and 0.35% BD-Rate savings under Low Delay B (LDB), Low Delay P (LDP), and Random Access (RA) configurations, respectively, when compared to the default RC algorithm used in VTM-19.2. Junqi Liao, Li Li 0040, Dong Liu 0002, Houqiang Li |
IEEE Trans. Multim. | 1 |
| 2023 | Padding-Aware Learned Image CompressionabstractFor current learned image compression methods, padding input images is necessary to meet the resolution requirements of down-sampling layers. However, the impact of padding has not been studied thoroughly. Most previous studies ignore padded images in the training process. In this paper, we analyze the impact of padding on compression performance. Then, we propose a padding-aware training (PAT) strategy, handling the padding effect during the training. Specifically, our PAT strategy calculates the loss of pre-padding image through a masking operation. Finally, according to our systematic experimental results, we find that images with different resolutions tend to favor different padding modes. Therefore, we further propose to conduct padding mode decision in the encoding process for rate-distortion optimization. Experiments demonstrate that our proposed PAT strategy and padding mode decision effectively compensate for the performance drop caused by padding. Haotian Zhang 0009, Junqi Liao, Yiheng Jiang, Li Li 0040, Dong Liu 0002 |
ISCAS | 2 |
| 2023 | Reinforcement Learning-based Frame-level Bit Allocation for VVCabstractAs frame-level bit allocation is a dependent sequential decision-making problem, it can be modeled as a Markov Decision Process (MDP) and solved by reinforcement learning (RL). Existing reports using RL for coding mainly have two problems: unrepresentative handcrafted features and limited interaction efficiency. In this paper, to address these two problems, we first propose to use a deep neural network to extract features from the state, which is designed as a combination of raw pixel-level information and handcrafted features. The network is able to extract information necessary for the following decision-making neural network automatically. We then propose an efficient training scheme. By designing a parallel interaction algorithm and using a simplified video encoder, our proposed scheme can train on a sophisticated encoder such as VTM. To the best of our knowledge, this is the first RL algorithm implemented for bit allocation in VTM. The experimental results show that our scheme can learn interpretable frame-level bit allocation and achieves a better rate-distortion (R-D) performance compared with VTM. Junqi Liao, Li Li 0040, Dong Liu 0002, Houqiang Li |
VCIP | 1 |