Riyu Lu

dblp:355/8424 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
8since 2021 · last 2026
0009-0008-0052-0483ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Learned Reference Picture Resampling Control: A Data-Centric Approach
abstract
Learned reference picture resampling control (LRPRC) adaptively adjusts the coding scale for each frame using an offline-trained neural network. It demonstrates promising promising rate-distortion (R-D) performance improvements over traditional methods, particularly in high-resolution, low-bit-rate video coding scenarios. However, existing LRPRC methods rely exclusively on locally optimal decision labels derived from greedy strategies for network training, leading to suboptimal control performance. To address this limitation, we introduce a novel data-centric solution that substantially improves training label quality, thereby enhancing overall LRPRC performance. Specifically, our key contribution is a parallelized beam search-based coding scale labeling algorithm, which captures decision dependencies across coding steps and produces higher-quality training labels with enhanced R-D performance. By fully exploiting the intra-trellis and inter-trellis parallelism of beam search and hierarchical coding, our proposed labeling algorithm achieves logarithmic-squared time complexity, making it highly suitable for large-scale cluster computing. We validate this simple yet effective data-centric LRPRC approach in the Versatile Video Encoder (VVenC) using 4K video sequences. Experimental results demonstrate that merely upgrading the beam search labels (without any neural architecture re-designs) consistently outperforms the state-of-the-art LRPRC method, achieving BD-rate reductions of 5.09%, 3.98%, and 3.59% under thefast,medium, andslowpresets, respectively.
Riyu Lu, Yingwen Zhang, Hengyu Man, Meng Wang 0017, Long Xu 0001, Shiqi Wang 0001, Xiaopeng Fan 0001
IEEE Trans. Circuits Syst. Video Technol.1
2025 Content-Aware Dynamic In-Loop Filter With Adjustable Complexity for VVC Intra Coding
abstract
Recently, neural network-based in-loop filters have been rapidly developed, effectively improving the reconstruction quality and compression efficiency in video coding. Existing deep in-loop filters typically employed networks with fixed structures to process all image blocks. However, under various bitrate conditions, compressed image blocks with different textures exhibit varying degradations, which poses a challenge for high-quality and low-complexity filtering. Additionally, different complexity requirements for coding tools in various scenarios limit the versatility of fixed models. To address these problems, a content-aware dynamic in-loop filter (dubbed DILF) with adjustable complexity is proposed in this paper. Specifically, DILF comprises a policy network and a filtering network. For each reconstructed image block, the policy network dynamically generates a filtering network topology based on pixel information and the quantization parameter (QP), guiding the filtering network to skip redundant layers and conduct content-aware image enhancement, thereby improving the filtering performance. In addition, by introducing a user-defined balancing factor into the policy network, the content-aware filtering network topology can be further adjusted according to user’s requirements, facilitating adjustable complexity with a single model. We integrate DILF into Versatile Video Coding (VVC) to replace the built-in deblocking filter. Extensive experiments demonstrate the efficiency of DILF in processing image blocks with varying degrees of degradation and its flexibility in controlling complexity. When the balancing factor is set to 2e-5, DILF achieves bitrate savings of 8.07%, 17.97%, and 20.93% on average for YUV components over VVC reference software VTM-11.0 under all-intra configuration. Compared to static networks with fixed structures, DILF demonstrates superior performance and lower computational complexity.
Hengyu Man, Hao Wang 0212, Riyu Lu, Zhaolin Wan, Xiaopeng Fan 0001, Debin Zhao
IEEE Trans. Circuits Syst. Video Technol.3
2025 Learning the Scale in Reference Picture Resampling for Versatile Video Coding
abstract
Compressing high-resolution videos under low bitrate constraints is a challenging task. Resampling-based compression, which reduces the resolution before encoding and restores it after decoding, has great potential to improve the rate-distortion performance in such scenarios. In this paper, we propose a learning-based frame-level coding scale control scheme that enhances the coding performance by adjusting the coding scale for each frame. The scheme cooperates with the Reference Picture Resampling of the latest video coding standard Versatile Video Coding (VVC), which allows coding scale variations on each frame. More specifically, a dataset with 5200 videos is created by a greedy rate-distortion optimization algorithm employed to select the optimal coding scale for each frame. A neural network-based decision model is further incorporated into VVC, learning to predict the coding scale for each frame in one pass. The scheme is implemented into the Fraunhofer Versatile Video Encoder (VVenC), a fast and efficient VVC encoder, and evaluated on 4 K contents. Experimental results show that the proposed scheme outperforms GOP-based coding scale adaptation methods, achieving average bitrate savings of 3.06% and 4.14% in terms of PSNR and MS-SSIM.
Riyu Lu, Yingwen Zhang, Hengyu Man, Meng Wang 0017, Shiqi Wang 0001, Xiaopeng Fan 0001
IEEE Trans. Multim.1
2024 Learned Image Compression for Both Humans and Machines via Dynamic Adaptation
abstract
Recent advancements in neural image compression have shown great potential in outperforming conventional standard codecs in terms of both rate-distortion and rate-analysis performance. However, there is an issue of divergent preferences in information preservation or reconstruction in the process of compression for humans and machines, respectively. Compression for humans tends to retain the signal fidelity or perceptual quality of visual appearance while compression for machines requires preserving critical semantic information, resulting in the limitation of the bitstream supporting only a single requirement during the compression. To bridge this gap, we propose a dynamic adaptation approach that generates a single bitstream serving both humans and machines. This approach aims to mitigate the domain gap among tasks, which facilitates maintaining the performance of out-of-scope tasks. Specifically, the proposed method concentrates on learning a dynamic adaptation process, i.e., optimizing the latent representation in the compressed domain in an end-to-end manner while adhering to the rate-performance constraint. Extensive results reveal that our paradigm significantly reduces the domain gap, surpassing existing codecs.
Lingyu Zhu 0006, Binzhe Li, Riyu Lu, Peilin Chen 0001, Qi Mao 0002, Zhao Wang 0004, Wenhan Yang, Shiqi Wang 0001
ICIP3
2024 Diffusion-Based Bit-Depth Expansion
abstract
Diffusion-based generative models have achieved remarkable success across a variety of applications. However, the potential application for bit-depth expansion has not been extensively studied. This paper introduces a wavelet-based diffusion model for the bit-depth expansion task. In this method, the image is first decomposed into low and high-frequency components via wavelet transformation. This decomposition allows for targeted processing by specialized modules and reduces computational complexity by lowering the image resolution. The low-frequency component is processed in both the forward diffusion and reverse denoising stages. Meanwhile, the high-frequency components are filtered by the High Frequency Denoising Filter (HFDF) to eliminate noise and artifacts. Finally, the low and high-frequency components are recombined into a predicted high-bit-depth image through inverse wavelet transformation. Experimental results demonstrate the superiority of the proposed method in producing perceptually compelling outputs that outperform previous methods.
Riyu Lu, Lingyu Zhu 0006, Baoliang Chen, Xiaopeng Fan 0001, Shiqi Wang 0001
MMSP1
2024 GeneWorker: An end-to-end robotic reinforcement learning approach with collaborative generator and worker networks
Hao Wang 0212, Hengyu Man, Wenxue Cui, Riyu Lu, Chenxin Cai, Xiaopeng Fan 0001
Neural Networks4
2024 MetaIP: Meta-Network-Based Intra Prediction With Customized Parameters for Video Coding
abstract
Intra prediction is a vital tool in video coding that eliminates the spatial redundancy within a frame to enhance compression efficiency. Conventional intra prediction methods employ multiple directional prediction modes to describe textures in local areas. Recently, research on neural network-based intra prediction has achieved great success. The block-context pairs are divided into multiple clusters according to a predefined relationship, and a corresponding network is trained and applied for each cluster. However, the networks in these methods adopt fixed parameters to predict diverse image blocks, making it hard to cope with various textures in natural images. Inspired by recent works on parameter prediction, in this paper, we propose a meta-network-based intra prediction method, called MetaIP, that dynamically customizes the network parameters for each block sample in a given cluster. MetaIP consists of a meta-subnetwork and a prediction subnetwork. For an image block, the meta-subnetwork takes its neighboring reference pixels and some auxiliary information (e.g., quantization parameter) as inputs to generate customized parameters first. Then, the prediction subnetwork uses the customized parameters to infer the predicted block. MetaIP can generate multiple sets of network parameters corresponding to multiple modes for an image block. The optimal mode is determined by the rate-distortion optimization. MetaIP is integrated into VVC to assist or replace the directional prediction modes to evaluate its performance. The experimental results demonstrate that MetaIP with four prediction modes achieves an average of 3.84% and 1.96% bitrate saving for the luma component over VTM-17.0 when assisting or replacing VVC intra modes, respectively.
Hengyu Man, Xiaopeng Fan 0001, Riyu Lu, Chang Yu 0006, Debin Zhao
IEEE Trans. Circuits Syst. Video Technol.3
2023 Meta-ILF: In-Loop Filter with Customized Weights For VVC Intra Coding
abstract
In-Loop filter (ILF) is an essential module in video coding for suppressing compression artifacts and thus improving the quality of reconstructed images. As the state-of-the-art video coding standard, Versatile Video Coding (H.266/VVC) employs three in-loop filters, including deblocking filter, sample adaptive offset, and adaptive loop filter. Recently, many neural network-based in-loop filters have been proposed and shown great success in image restoration. In this paper, we propose a meta-learning-based method called Meta-ILF, which performs filtering with customized weights to enhance the quality of VVC intra-coded images. Meta-ILF consists of a meta-network and a filter network. For each reconstructed image block, the meta-network generates the customized weights first. Then, the filter network uses the customized weights to infer the enhanced reconstruction. By dynamically customizing the network weights for each reconstructed block, meta-ILF can better cope with the diverse compression artifacts. To test the performance, Meta-ILF is integrated into VVC reference software VTM-11.0. The experimental results demonstrate that Meta-ILF can reach an average of 6.77% Bjøntegaard Delta rate (BD-rate) improvement over VVC with all intra configuration.
Hengyu Man, Riyu Lu, Xiaopeng Fan 0001
ICME3