EDBT 2026 Demo / reviewers in the wild / expert
Zhuoyuan Li 0001
dblp:81/2220-1
· DBLP profile ↗
17ranked-venue papers
3as first author
17since 2021 · last 2026
0009-0003-7370-4068ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 3 first-author · 14 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Systems, architecture and hardware · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Neural Video Compression with Reference HierarchyabstractEfficient reference structures are essential in video compression, enabling the exploitation of temporal dependencies across frames to reduce redundancy. In this paper, we delve into the inter-frame reference management mechanism in neural video codecs (NVCs). Previous schemes have inherited the reference propagation mechanism with the guidance of predefined reference structure, but the reference modeling across diverse reference sources remains underexplored. Moreover, the mismatch between the reference structure used for motion estimation and motion compensation limits the effectiveness of inter-frame prediction. To address the above limitations, we propose the unified reference hierarchy that integrates a learned hierarchical reference structure into the existing inherent reference propagation mechanism. Specifically, we first propose the hierarchical reference structure (HRS) to manage the multiple temporal contexts in the propagated reference feature, where a hierarchy-aware reference modulation module is integrated to select the most relevant reference features across different quality levels under the guidance of the reference balance loss. In addition, we propose the HRS-guided feature-wise inter-frame prediction that learns the low-rank approximation of the selected reference feature for ensuring the consistency and improving the inter-frame prediction performance. We conduct experiments on a state-of-the-art NVC, DCVC-DC. Experimental results show that our codec achieves an average 26% bitrate saving over H.266/VVC, and a 28.2% bitrate reduction compared to DCVC-DC without increasing the decoding complexity. Chuanbo Tang, Zhuoyuan Li 0001, Li Li 0040, Dong Liu 0002, Feng Wu 0001 |
AAAI | 2 |
| 2026 | Standard-Compliant Joint Optimization of Rate-Distortion-Decoding-Complexity for Versatile Video Coding
Jiazhen Wang, Yao Li 0016, Xinmin Feng, Zhuoyuan Li 0001, Li Li 0040, Dong Liu 0002 |
ISCAS | 6 |
| 2026 | Partition Map-Based Fast Block Partitioning for VVC Inter CodingabstractAmong the new techniques of Versatile Video Coding (VVC), the quadtree with nested multi-type tree (MTT) block structure yields significant coding gains by providing more flexible block partitioning patterns. However, the recursive partition search in the VVC encoder increases the encoder complexity substantially. To address this issue, we propose a partition map-based algorithm to pursue fast block partitioning in inter coding. Based on our previous work on partition map-based methods for intra coding, we analyze the characteristics of VVC inter coding and improve the partition map by incorporating an MTT mask for early termination. Next, we develop a neural network that uses both spatial and temporal features to predict the partition map. It consists of several special designs, including stacked top-down and bottom-up processing, quantization parameter modulation layers, and partitioning-adaptive warping. Furthermore, we present a dual-threshold decision scheme to achieve a fine-grained trade-off between complexity reduction and rate-distortion performance loss. The experimental results demonstrate that the proposed method achieves an average 51.30% encoding time saving with a 2.12% Bjøntegaard-delta-bit-rate under the random access configuration. The source code is publicly available athttps://github.com/ustc-ivclab/IPM. Xinmin Feng, Zhuoyuan Li 0001, Li Li 0040, Dong Liu 0002, Feng Wu 0005 |
IEEE Trans. Multim. | 2 |
| 2026 | USTC-TD: A Test Dataset and Benchmark for Image and Video Coding in 2020sabstractImage/video coding has been a remarkable research area for both academia and industry for many years. Testing datasets, especially high-quality image/video datasets, are desirable for the justified evaluation of coding-related research, practical applications, and standardization activities. We put forward a test dataset, namely USTC-TD, which has been successfully adopted in the practical end-to-end image/video coding challenge ofIEEE International Conference on Visual Communications and Image Processing (VCIP)in 2022 and 2023. USTC-TD contains 40 images at 4K spatial resolution and 10 video sequences at 1080p spatial resolution, featuring various content due to the diverse environmental factors (e.g., scene type, texture, motion, view) and the designed imaging factors (e.g., illumination, lens, shadow). We quantitatively evaluate USTC-TD on different image/video features (spatial, temporal, color, lightness), and compare it with the previous image/video test datasets, which verifies its excellent compensation for the shortcomings of existing datasets. We also evaluate both classic standardized and recently learned image/video coding schemes on USTC-TD using objective quality metrics (PSNR, MS-SSIM, VMAF) and subjective quality metric (MOS), providing an extensive benchmark for these evaluated schemes. Based on the characteristics and specific design of the proposed test dataset, we analyze the benchmark performance and shed light on the future research and development of image/video coding. All the data are released online:https://esakak.github.io/USTC-TD. Zhuoyuan Li 0001, Junqi Liao, Chuanbo Tang, Haotian Zhang 0009, Yifan Bian, Xihua Sheng, Xinmin Feng, Yao Li 0016, Changsheng Gao, Li Li 0040, Dong Liu 0002, Feng Wu 0005 |
IEEE Trans. Multim. | 1 |
| 2026 | Learning Dual Modality Interactions for Event-Based Motion DeblurringabstractEvent cameras hold great potential for motion deblurring because they capture motion information with microsecond precision, offering robustness to motion blur. However, the limited interaction between RGB frames and event streams presents a significant challenge, preventing the full utilization of the event cameras' unique advantages. To address this, we proposeDual frame-eventInteraction and introduce a multi-scaleNetwork structure, DuInt-Net. DuInt-Net aims to tackle two key challenges: (1) enhancing the representational and interaction capabilities between RGB frames and event streams, and (2) adaptively selecting richer visual features for improved motion deblurring. We introduce an event-frame joint interaction module that consists of three branches: a base branch, a global awareness attention branch, and a local enhancement attention branch. The base branch processes essential pixel-level features that retain the original structural information. The global branch integrates event data to improve large-scale motion understanding, while the local branch uses large-kernel convolutions to refine fine-grained details in RGB frames. For superior reconstruction performance, we also propose the event-guided multi-scale fusion attention module, which effectively combines local visual information and global frame-event relationships. Extensive experiments demonstrate that DuInt-Net achieves superior performance, both quantitatively and qualitatively, showcasing its superior motion deblurring capabilities. Zeyu Xiao 0002, Zhuoyuan Li 0001, Yang Zhao 0002, Yu Liu 0023, Zhao Zhang 0001, Wei Jia 0001 |
IEEE Trans. Multim. | 2 |
| 2025 | Occlusion-Embedded Hybrid Transformer for Light Field Super-ResolutionabstractTransformer-based networks have set new benchmarks in light field super-resolution (SR), but adapting them to capture both global and local spatial-angular correlations efficiently remains challenging. Moreover, many methods fail to account for geometric details like occlusions, leading to performance drops. To tackle these issues, we introduce OHT. This hybrid network leverages occlusion maps through an occlusion-embedded mix layer. It combines the strengths of convolutional networks and Transformers via spatial-angular separable convolution (SASep-Conv) and angular self-attention (ASA). SASep-Conv offers a lightweight alternative to 3D convolution for capturing spatial-angular correlations, while the ASA mechanism applies 3D self-attention across the angular dimension. These designs allow OHT to capture global angular correlations effectively. Extensive experiments on multiple datasets demonstrate OHT's superior performance. Zeyu Xiao 0002, Zhuoyuan Li 0001, Wei Jia 0001 |
AAAI | 2 |
| 2025 | Neural Video Compression with Context ModulationabstractEfficient video coding is highly dependent on exploiting the temporal redundancy, which is usually achieved by extracting and leveraging the temporal context in the emerging conditional coding-based neural video codec (NVC). Although the latest NVC has achieved remarkable progress in improving the compression performance, the inherent temporal context propagation mechanism lacks the ability to sufficiently leverage the reference information, limiting further improvement. In this paper, we address the limitation by modulating the temporal context with the reference frame in two steps. Specifically, we first propose the flow orientation to mine the inter-correlation between the reference frame and prediction frame for generating the additional oriented temporal context. Moreover, we introduce the context compensation to leverage the oriented context to modulate the propagated temporal context generated from the propagated reference feature. Through the synergy mechanism and decoupling loss supervision, the irrelevant propagated information can be effectively eliminated to ensure better context modeling. Experimental results demonstrate that our codec achieves on average 22.7% bitrate reduction over the advanced traditional video codec H.266/VVC, and offers an average 10.1% bitrate saving over the previous state-of-the-art NVC DCVC-FM. The code is available at https://github.com/Austin4USTC/DCMVC. Chuanbo Tang, Zhuoyuan Li 0001, Yifan Bian, Li Li 0040, Dong Liu 0002 |
CVPR | 2 |
| 2025 | Rethinking Joint Optimization in Feature Compression: Insights from Person Re-IdentificationabstractJoint optimization, which jointly optimizes compression and machine vision algorithms, is widely regarded as an effective strategy for enhancing compression performance in the field of coding for machines. However, existing joint optimization methods usually incorporate a semantics parsing module at the end of the pipeline, raising a critical question: Does the performance improvement stem from the joint optimization itself, or is it primarily driven by the tailed semantics parsing module? To address this, we disentangle the tailed semantics parsing module from the joint optimization pipeline by leveraging the simplicity of the person re-identification task, where semantics parsing involves deterministic feature matching rather than a learned neural network. First, we propose a separate optimization pipeline and two joint optimization pipelines to systematically investigate the effectiveness of joint optimization. Our findings reveal that joint optimization alone does not necessarily guarantee performance improvement. Second, we evaluate the influence of the tailed semantics parsing module by equipping it with varying capabilities, demonstrating that higher parsing capability directly correlates with better machine vision performance. These findings underscore the pivotal role of tailed semantics parsing in enhancing machine vision performance and challenge the assumption that joint optimization alone drives improvement. This work offers new insights for designing effective coding methods, emphasizing the interplay between optimization strategies and tailed semantics parsing. Changsheng Gao, Zhuoyuan Li 0001, Li Li 0040, Dong Liu 0002, Feng Wu 0001, Weisi Lin |
ICME | 2 |
| 2025 | Collaborative Decoder-side Motion Vector Refinement for Video CodingabstractMotion compensation prediction (MCP) is a key technology to reduce the temporal redundancy in video coding. Recently, in order to improve its efficiency, the decoder-side MCP schemes are gradually adopted in advanced video coding standards, especially the decoder-side motion vector refinement (DMVR). In DMVR, the bilateral matching scheme is used to refine the motion vector (MV) obtained from merge-based inter mode at the sub-block level, which assumes that the motion vector difference (MVD) in the two reference directions of a bi-prediction block has the symmetric property. Although the bilateral assumption can effectively reduce complexity without any extra signal, the fixed searching rules limit the motion vector accuracy. To address this limitation, we propose a collaborative decoder-side motion vector refinement (C-DMVR) framework. In C-DMVR, the sub-block-based collaborative mechanism is introduced to optimize the distortion calculation (CDC) and bilateral-based searching strategy (CSS) to avoid inaccurate searching results, respectively. In the CDC, the receptive field of the distortion function is enlarged with the collaboration of additional spatial neighbor information to assist the accurate decision of motion vector candidates. In the CSS, the coarse-to-fine candidate list derivation scheme is introduced to construct the accurate searching path with the collaboration of neighbor sub-blocks. The proposed method is implemented into the AOMedia Video 2 reference software, AV2. Experimental results show that the proposed method achieves on average 0.17%, and up to 0.60% BD-rate reduction compared to the AV2 anchor under the random access (RA) configuration, with a slight increase of time complexity on the encoding/decoding side. Zhuoyuan Li 0001, Yao Li 0016, Li Li 0040, Houqiang Li |
ISCAS | 2 |
| 2025 | IVCA: Inter-relation-aware Video Complexity AnalyzerabstractTo address the real-time analysis requirements of video streaming applications, we propose an innovative inter-relation-aware video complexity analyzer (IVCA) to enhance the existing video complexity analyzer (VCA). The IVCA overcomes the limitations of the VCA by incorporating inter-frame relations, focusing on inter motion and reference structure. To begin with, we improve the accuracy of temporal features by integrating feature-domain motion estimation into the IVCA framework, which allows for a more nuanced understanding of motion across frames. Furthermore, inspired by the hierarchical reference structures utilized in modern codecs, we introduce layer-aware weights that effectively adjust the contributions of frame complexity across different layers, ensuring a more balanced representation of video characteristics. In addition, we broaden the analysis of temporal features by considering reference frames rather than relying solely on the preceding frame, thereby enriching the contextual understanding of video content. Experimental results demonstrate a significant enhancement in complexity estimation accuracy achieved by the IVCA, coupled with a negligible increase in time complexity, indicating its potential for real-time applications in video streaming scenarios. This advancement not only improves video processing efficiency but also paves the way for more sophisticated analytical tools in video technology. Junqi Liao, Yao Li 0016, Zhuoyuan Li 0001, Li Li 0040, Dong Liu 0002 |
ISCAS | 3 |
| 2025 | Frequency Domain Intra Pattern Copy for JPEG XS Screen Content CodingabstractJPEG XS is a wavelet-based lightweight image coding standard that features low-complexity and low-latency. As currently there are no efficient intra-compensation prediction techniques conforming to these features, we propose a frequency domain intra-copy prediction framework named Intra Pattern Copy (IPC), to improve its coding efficiency on screen content. In IPC, prediction methods that leverage the diverse decomposed patterns of two-dimensional wavelets, including the directional and frequency characteristics, are proposed to achieve efficient predictions under low-complexity and low-latency constraints. Specifically, we perform in-band compensation predictions in a multi-band synchronized approach, with coefficients of similar pattern distributions predicted simultaneously. A coefficient grouping scheme is derived from the band characteristics to facilitate this compensation process. Based on the grouping scheme, a multi-band synchronized side information coding method is also proposed to code the pattern offset vector of coefficients. Moreover, pattern search schemes incorporating strict limitations on the search range and prediction block size are further developed. Simulation results on JPEG XS demonstrate that an average improvement of 0.75 dB and 1.99 dB in BD-PSNR can be achieved on screen content for two different wavelet decomposition configurations, respectively, with a moderate increase in complexity. Yao Li 0016, Zhuoyuan Li 0001, Dong Liu 0002, Li Li 0040 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Offline and Online Optical Flow Enhancement for Deep Video CompressionabstractVideo compression relies heavily on exploiting the temporal redundancy between video frames, which is usually achieved by estimating and using the motion information. The motion information is represented as optical flows in most of the existing deep video compression networks. Indeed, these networks often adopt pre-trained optical flow estimation networks for motion estimation. The optical flows, however, may be less suitable for video compression due to the following two factors. First, the optical flow estimation networks were trained to perform inter-frame prediction as accurately as possible, but the optical flows themselves may cost too many bits to encode. Second, the optical flow estimation networks were trained on synthetic data, and may not generalize well enough to real-world videos. We address the twofold limitations by enhancing the optical flows in two stages: offline and online. In the offline stage, we fine-tune a trained optical flow estimation network with the motion information provided by a traditional (non-deep) video compression scheme, e.g. H.266/VVC, as we believe the motion information of H.266/VVC achieves a better rate-distortion trade-off. In the online stage, we further optimize the latent features of the optical flows with a gradient descent-based algorithm for the video to be compressed, so as to enhance the adaptivity of the optical flows. We conduct experiments on two state-of-the-art deep video compression schemes, DCVC and DCVC-DC. Experimental results demonstrate that the proposed offline and online enhancement together achieves on average 13.4% bitrate saving for DCVC and 4.1% bitrate saving for DCVC-DC on the tested videos, without increasing the model or computational complexity of the decoder side. Chuanbo Tang, Xihua Sheng, Zhuoyuan Li 0001, Haotian Zhang 0009, Li Li 0040, Dong Liu 0002 |
AAAI | 3 |
| 2024 | Rethinking the Joint Optimization in Video Coding for Machines: A Case StudyabstractIn this work, we investigate the joint optimization strategy in the scenario of video coding for machines (VCM). We formulated two kinds of joint optimization strategies, Opt_JA and Opt_JH , and compared them with the separate optimization strategy Opt_S. The three optimization strategies are illustrated in Fig. 1 . In Opt_S , we separately train the feature compression network with mean squared error (MSE). In Opt_JA , we optimize all modules jointly toward the person re-identification task. In Opt_JH , only the aggregation module and feature compression module are jointly optimized. The feature compression consists of two fully-connected (FC) layers and two batch normalization (BN) layers. Specifically, we set five compression ratios (CR): 256, 128, 64, 32, and 16. Changsheng Gao, Zhuoyuan Li 0001, Li Li 0040, Dong Liu 0002, Feng Wu 0001 |
DCC | 2 |
| 2024 | In-Loop Filtering via Trained Look-Up TablesabstractIn-loop filtering (ILF) is a key technology in image/video coding for reducing the artifacts. Recently, neural network-based in-loop filtering methods achieve remarkable coding gains beyond the capability of advanced video coding standards, establishing themselves a promising candidate tool for future standards. However, the utilization of deep neural networks (DNN) brings high computational complexity and raises high demand of dedicated hardware, which is challenging to apply into general use. To address this limitation, we study an efficient in-loop filtering scheme by adopting look-up tables (LUTs). After training a DNN with a predefined reference range for in-loop filtering, we cache the output values of the DNN into a LUT via traversing all possible inputs. In the coding process, the filtered pixel is generated by locating the input pixels (to-be-filtered pixel and reference pixels) and interpolating between the cached values. To further enable larger reference range within the limited LUT storage, we introduce an enhanced indexing mechanism in the filtering process, and a clipping/finetuning mechanism in the training. The proposed method is implemented into the Versatile Video Coding (VVC) reference software, VTM-11.0. Experimental results show that the proposed method, with three different configurations, achieves on average 0.13%∼0.51%, and 0.10% ∼0.39% BD-rate reduction under the all-intra (AI) and random-access (RA) configurations respectively. The proposed method incurs only 1% ∼8% time increase, an additional computation of 0.13 ∼0.93 kMAC/pixel, and 164 ∼1148 KB storage cost for a single model. Our method has explored a new and more practical approach for neural network-based ILF. Zhuoyuan Li 0001, Jiacheng Li 0004, Yao Li 0016, Li Li 0040, Dong Liu 0002, Feng Wu 0001 |
VCIP | 1 |
| 2024 | Uniformly Accelerated Motion Model for Inter PredictionabstractInter prediction is a key technology in video coding to reduce the temporal redundancy. In natural videos, there are usually moving objects with changing velocity, resulting in complex motion fields that are difficult to represent compactly. In Versatile Video Coding (VVC), existing inter prediction methods usually assume uniform speed motion between consecutive frames, which may not well handle the complex motion fields in the real world. To address these issues, we introduce a uniformly accelerated motion model (UAMM) to exploit velocity and acceleration of moving objects between the video frames, and further combine them to assist in the inter prediction methods to handle the motion change in the temporal domain. First, we review the theory of UAMM. Second, we propose UAMM-based parameter derivation and extrapolation schemes in the coding process. Third, we integrate the UAMM into existing inter prediction modes (Merge, MMVD, CIIP) to achieve higher prediction accuracy. The proposed method is implemented into the VVC reference software, VTM version 12.0. Experimental results show that the proposed method achieves up to 0.38% BD-rate reduction compared to the VTM anchor, under the Low-delay P configuration, with a slight increase of time complexity on the encoding/decoding side. Zhuoyuan Li 0001, Yao Li 0016, Chuanbo Tang, Li Li 0040, Dong Liu 0002, Feng Wu 0001 |
VCIP | 1 |
| 2024 | Temporal Wavelet Transform-Based Low-Complexity Perceptual Quality Enhancement of Compressed VideoabstractThe past few years have witnessed a great success in applying deep learning to enhance the perceptual quality of compressed video. These methods usually perform frame-by-frame quality enhancement, incurring high computational complexity. Low-complexity perceptual quality enhancement is addressed in this paper, motivated by the observation of temporal correlations among video frames. We propose to decompose video content into temporal low-frequency and high-frequency components, and to focus the enhancement of the temporal low-frequency component, which may significantly reduce the computational complexity. Specifically, we employ the temporal wavelet transform (TWT) for the temporal frequency analysis, and build a TWT-based multiple-input multiple-output perceptual quality enhancement scheme. First, we use a motion estimation method on the input video to acquire the motion information, and then use TWT to obtain the temporal low- and high-frequency components. Second, we design a deep network to enhance the quality of the temporal low-frequency component. Finally, the temporal high-frequency component and the enhanced temporal low-frequency component are combined by the temporal wavelet inverse transform (TWIT) to generate the enhanced video. Experimental results show that our method achieves comparable perceptual quality to that of the state-of-the-art methods, but reduces the computational complexity to 1/13. Cunhui Dong, Haichuan Ma, Zhuoyuan Li 0001, Li Li 0040, Dong Liu 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Global Homography Motion Compensation for Versatile Video CodingabstractIn Versatile Video Coding (VVC), local affine motion compensation (LAMC) is adopted to handle complex motions, such as rotation and zooming. However, it is inefficient to use LAMC to handle the global motion due to the following two reasons. First, the use of LAMC may lead to some extra bit cost on the affine motion model parameters. Second, the precision of LAMC is restricted by the MV precision of the control points. Therefore, in this paper, we propose a global homography motion compensation (GHMC) framework to better characterize the global motion. For each coding block, an extra mode is added to perform motion compensation based on an 8-parameter global homography motion model. In addition, an extrapolation scheme is designed to derive the parameters from reference frames to save the bit cost for signaling them. The proposed framework is implemented into the VVC reference software VTM-6.0. Experimental results show that, on average, 0.69% and 0.66% BD-rate reduction is achieved under Low Delay P and Low Delay B configurations, respectively, for sequences with rich complex global motions. Yao Li 0016, Zhuoyuan Li 0001, Li Li 0040, Dong Liu 0002, Houqiang Li |
VCIP | 2 |