Congrui Fu

dblp:302/9683 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
6since 2021 · last 2025
0009-0001-3721-3291ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 6 since 2021
YearPublicationVenuePosition
2025 Global Spatial-Temporal Information-Based Residual ConvLSTM for Video Space-Time Super-Resolution
abstract
By converting low-frame-rate, low-resolution videos into high-frame-rate, high-resolution ones, space-time video super-resolution techniques can enhance visual experiences and facilitate more efficient information dissemination. We propose a convolutional neural network (CNN) for space-time video super-resolution, namely GIRNet. Our method combines long-term global information and short-term local information from the video to better extract complete and accurate spatial-temporal information. To generate highly accurate features and thus improve performance, the proposed network integrates a feature-level temporal interpolation module with deformable convolutions and a global spatial-temporal information-based residual convolutional long short-term memory (convLSTM) module. In the feature-level temporal interpolation module, we leverage deformable convolution, which adapts to deformations and scale variations of objects across different scene locations. This provides a more efficient solution than conventional convolution for extracting features from moving objects. Our network effectively uses forward and backward feature information to determine inter-frame offsets, leading to the direct generation of interpolated frame features. In the global spatial-temporal information-based residual convLSTM module, the first convLSTM is used to derive global spatial-temporal information from the input features, and the second convLSTM uses the previously computed global spatial-temporal information feature as its initial cell state. This second convLSTM adopts residual connections to preserve spatial information, thereby enhancing the output features. Experiments on the Vimeo90 K dataset show that the proposed method outperforms open source state-of-the-art techniques in peak signal-to-noise-ratio (by 1.45 dB, 1.14 dB, and 0.2 dB over STARnet, TMNet, and 3DAttGAN, respectively), structural similarity index(by 0.027, 0.023, and 0.006 over STARnet, TMNet, and 3DAttGAN, respectively), and visual quality.
Congrui Fu, Hui Yuan 0001, Shiqi Jiang 0006, Liquan Shen, Raouf Hamzaoui
IEEE Trans. Multim.1
2024 A Transformer-Based Intra Luma Enhancement for H.266/VVC
abstract
Intra prediction is essential in reducing spatial domain correlation in video coding. To improve intra prediction accuracy, we introduce a transformer-based quality enhancement method aiming atimproving the luma quality of reconstructed coding tree units (CTUs). Our transformer-based model, namely Enhanceformer, utilizes multi-head attention for comprehensive feature extraction across multiple stages, levels, and scales. By integrating the model into H.266/VVC codec, it not only improves the luma quality of the current reconstructed CTU, but also provides more accurate references for intra prediction of subsequent CTUs. Experimental results show that this method achieves average BD rate savings of 2.63%, 0.21% and 0.48% for Y, Cb and Cr components respectively in all intra configuration, outperforming H.266/Versatile Video Coding (VVC) anchor.
Wenrui Lv, Hui Yuan 0001, Congrui Fu, Shiqi Jiang 0006, Junyan Huo
PCS3
2024 OMR-NET: A Two-Stage Octave Multi-Scale Residual Network for Screen Content Image Compression
abstract
Screen content (SC) differs from natural scene (NS) with unique characteristics such as noise-free, repetitive patterns, and high contrast. Aiming at addressing the inadequacies of current learned image compression (LIC) methods for SC, we propose an improved two-stage octave convolutional residual blocks (IToRB) for high and low-frequency feature extraction and a cascaded two-stage multi-scale residual blocks (CTMSRB) for improved multi-scale learning and nonlinearity in SC. Additionally, we employ a window-based attention module (WAM) to capture pixel correlations, especially for high contrast regions in the image. We also construct a diverse SC image compression dataset (SDU-SCICD2K) for training, including text, charts, graphics, animation, movie, game and mixture of SC images and NS images. Experimental results show our method, more suited for SC than NS data, outperforms existing LIC methods in rate-distortion performance on SC images.
Shiqi Jiang 0006, Ting Ren, Congrui Fu, Shuai Li 0005, Hui Yuan 0001
IEEE Signal Process. Lett.3
2023 Cuboid-Net: A multi-branch convolutional neural network for joint space-time video super resolution
abstract
Abstract The demand for high‐resolution videos has been consistently rising across various domains, propelled by continuous advancements in societal. Nonetheless, limitations in imaging and economic factors often result in obtaining low‐resolution images. The currently available space‐time video super‐resolution methods often fail to fully exploit the information existing within the spatio‐temporal domain. To address this problem, the issue is tackled by conceptualizing the input low‐resolution video as a cuboid structure. An innovative methodology called “Cuboid‐Net”, which incorporates a multi‐branch convolutional neural network, is introduced. Cuboid‐Net is designed to collectively enhance the spatial and temporal resolutions of videos, enabling the extraction of rich and meaningful information across both spatial and temporal dimensions. Specifically, the input video is taken as a cuboid to generate different directional slices as input for different branches of the network. The proposed network contains four modules, that is, a multi‐branch‐based hybrid feature extraction module, a multi‐branch‐based reconstruction module, a first‐stage quality enhancement module, and a second‐stage cross frame quality enhancement module for interpolated frames only. Experimental results demonstrate that the proposed method is not only effective for spatial and temporal super‐resolution of video but also for spatial and angular super‐resolution of light field.
Congrui Fu, Hui Yuan 0001, Hongji Xu, Hao Zhang 0211, Liquan Shen
IET Image Process.1
2023 TMSO-Net: Texture adaptive multi-scale observation for light field image depth estimation
Congrui Fu, Hui Yuan 0001, Hongji Xu, Hao Zhang 0211, Liquan Shen
J. Vis. Commun. Image Represent.1
2021 DRLFNet: A Dense-Connection Residual Learning Neural Network for Light Field Super Resolution
Congrui Fu, Junhui Hou, Hui Yuan 0001
ICIG (3)1