VLDB 2026 Research / reviewers in the wild / expert
Yuantong Zhang
dblp:304/3350
· DBLP profile ↗
16ranked-venue papers
6as first author
16since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 5 first-author · 13 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Continuous Space-Time Video Resampling with Invertible Motion SteganographyabstractSpace-time video resampling aims to conduct both spatial-temporal downsampling and upsampling processes to achieve high-quality video reconstruction. Although there has been much progress, some major challenges still exist, such as how to preserve motion information during temporal resampling while avoiding blurring artifacts, and how to achieve flexible temporal and spatial resampling factors. In this paper, we introduce an Invertible Motion Steganography Module (IMSM), designed to embed motion information from high-frame-rate videos into downsampled frames with lower frame rates in a visually imperceptible manner. Its reversible nature allows the motion information to be recovered, facilitating the reconstruction of high-frame-rate videos. Furthermore, we propose a 3D implicit feature modulation technique that enables continuous spatiotemporal resampling. With tailored training strategies, our method supports flexible frame rate conversions, including non-integer changes like 30 FPS to 24 FPS and vice versa. Extensive experiments show that our method significantly outperforms existing solutions across multiple datasets in various video resampling tasks with high flexibility. Codes will be made available at the URL https://github.com/hahazh/CSTVR. Yuantong Zhang, Zhenzhong Chen 0001 |
CVPR | 1 |
| 2025 | Low-Decoding-Complexity Learned Image Compression with Masked Convolutional Layer Re-parameterization
Wenzhuo Ma, Nianxiang Fu, Junxi Zhang, Yuantong Zhang, Zhenzhong Chen 0001 |
PCS | 4 |
| 2025 | Joint reference frame synthesis and post filter enhancement for Versatile Video Coding
Weijie Bao, Yuantong Zhang, Jianghao Jia, Zhenzhong Chen 0001, Shan Liu 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2025 | Space-Time Video Super-Resolution With Neural OperatorabstractThis paper addresses the task of space-time video super-resolution (STVSR). Existing methods generally suffer from inaccurate motion estimation and motion compensation (MEMC) problems for large motions. Inspired by recent progress in physics-informed neural networks, we model the challenges of MEMC in STVSR as a mapping between two continuous function spaces. Specifically, our approach transforms independent low-resolution representations in the coarse-grained continuous function space into refined representations with enriched spatiotemporal details in the fine-grained continuous function space. To achieve efficient and accurate MEMC, we design a Galerkin-type attention function to perform frame alignment and temporal interpolation. Due to the linear complexity of the Galerkin-type attention mechanism, our model avoids patch partitioning and offers global receptive fields, enabling precise estimation of large motions. The experimental results show that the proposed method surpasses state-of-the-art techniques in both fixed-size and continuous space-time video super-resolution tasks. Code is publicly available at the URL https://github.com/hahazh/STVSR-NO. Yuantong Zhang, Hanyou Zheng, Daiqin Yang, Zhenzhong Chen 0001, Haichuan Ma, Wenpeng Ding |
IEEE Trans. Image Process. | 1 |
| 2024 | Deep Reference Frame for Versatile Video Coding with Structural Re-parameterizationabstractIn video coding, inter-prediction leverages neigh-boring frames to reduce temporal redundancy. The quality of these reference frames is essential for effective inter-prediction. Although many neural network-based methods have been proposed to improve the quality of reference frames, there is still room for the performance and efficiency trade-off. In this paper, we propose an interpolation diverse branch block (InterDBB) suitable for lightweight frame interpolation networks, which optimizes deep reference frame interpolation networks to improve performance without sacrificing speed and increasing complexity. Specifically, we propose a multi-branch structural reparameterization block without batch normalization. This straightforward yet effective modification ensures training stability and performance improvement. Moreover, we propose a parameterized motion estimation strategy based on different input resolution, to achieve a better trade-off between performance and computational complexity. Experimental results demonstrate that our method achieves -2.01%/-2.87%/-2.44% coding efficiency improvements for Y/U/V components under random access (RA) configuration compared to VTM-11.0_NNVC-5.0. Chengzhuo Gui, Yuantong Zhang, Weijie Bao, Zhenzhong Chen 0001, Huairui Wang, Shan Liu 0001 |
VCIP | 2 |
| 2024 | Lightweight Arbitrary-Scale Super-Resolution of Remote Sensing Images via Super-Scale FeatureabstractRemote sensing image (RSI) super-resolution (SR) demands lightweight and efficient methods due to required rapid response in practical applications. Integrating RSIs with different resolutions for diverse applications also requires arbitrary-scale SR, making fix-scaled SR scale inflexible. Therefore, a lightweight SR algorithm capable of arbitrary-scale is necessary for RSIs. To address the above issue, a super-scale feature-based lightweight arbitrary-scale (SFLA) SR network is proposed in this paper. The network consists of two modules: 1) A super-scale feature extraction (SSFE) module that extracts features at both the initial low-resolution (LR) and an integer super-scale resolution, 2) A self-attention implicit function reconstruction (SIFR) module that utilizes multi-layer perceptron (MLP) network and self-attention mechanism for pixel-wise feature mapping to achieve superior SR results. Comparative experiments and ablation results demonstrate that the proposed SFLA algorithm effectively strikes a good balance between performance and complexity. Yifei Long, Yuantong Zhang, Daiqin Yang, Zhenzhong Chen 0001, Huairui Wang, Shan Liu 0001 |
VCIP | 2 |
| 2024 | Active RIS-Assisted Integrated Sensing and Communication Systems: Joint Receive-Transmit Beamforming and Reflection DesignabstractActive reconfigurable intelligent surface (RIS) acts as an enhancement of signal transmitted by the base station (BS) to overcome the effects of multiplicative fading and provide communication services for users. In this paper, we investigate an active RIS-assisted integrated sensing and communication (ISAC) system in which the BS senses the target in the presence of interference and establishes communication with multiple users assisted by the RIS. The optimization problem is formulated to jointly design the BS’s sensing transceiver beamforming, communication beamforming, and the RIS’s reflection coefficients to maximize the sensing SINR while satisfying the communication SINR requirements of users. An algorithm based on the ideas of quadratic transformation (QT), Karush-Kuhn-Tucker (KKT) conditions, semidefinite relaxation (SDR), and alternating optimization (AO) is designed to transform and solve the non-convex problem optimally. Simulation results verify the effectiveness and superiority of active RIS-assisted ISAC systems in improving sensing performance with communication SINR constraints. Yuantong Zhang, Junzhe Lin, Lianfen Huang, Zhibin Gao, Guozhen Xu |
VTC Fall | 1 |
| 2024 | Prediction of blast-hole utilization rate using structured nonlinear support vector machine combined with optimization algorithms
Bingbing Yu, Yuantong Zhang, Guohao Wang |
Appl. Intell. | 4 |
| 2024 | An illumination-guided dual-domain network for image exposure correction
Jie Yang 0002, Yuantong Zhang, Zhenzhong Chen 0001, Daiqin Yang |
J. Vis. Commun. Image Represent. | 2 |
| 2024 | Deep Reference Frame Generation Method for VVC Inter Prediction EnhancementabstractIn video coding, inter prediction aims to reduce temporal redundancy by using previously encoded frames as references. The quality of reference frames is crucial to the performance of inter prediction. This paper presents a deep reference frame generation method to optimize the inter prediction in Versatile Video Coding (VVC). Specifically, reconstructed frames are sent to a well-designed frame generation network to synthesize a picture similar to the current encoding frame. The synthesized picture serves as an additional reference frame inserted into the reference picture list (RPL) to provide a more reliable reference for subsequent motion estimation (ME) and motion compensation (MC). The frame generation network employs optical flow to predict motion precisely. Moreover, an optical flow reorganization strategy is proposed to enable bi-directional and uni-directional predictions with only a single network architecture. To reasonably apply our method to VVC, we further introduce a normative modification of the temporal motion vector prediction (TMVP). Integrated into the VVC reference software VTM-15.0, the deep reference frame generation method achieves coding efficiency improvements of 5.22%, 3.61%, and 3.83% for the Y component under random access (RA), low delay B (LDB), and low delay P (LDP) configurations, respectively. The proposed method has been discussed in Joint Video Exploration Team (JVET) meeting and is currently part of Exploration Experiments (EE) for further study. Jianghao Jia, Yuantong Zhang, Han Zhu 0003, Zhenzhong Chen 0001, Zizheng Liu, Xiaozhong Xu, Shan Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Learning a Single Convolutional Layer Model for Low Light Image EnhancementabstractLow-light image enhancement (LLIE) aims to improve the illuminance of images due to insufficient light exposure. Recently, various lightweight learning-based LLIE methods have been proposed to handle the challenges of unfavorable prevailing low contrast, low brightness, etc. In this paper, we have streamlined the architecture of the network to the utmost degree. By utilizing the effective structural re-parameterization technique, a single convolutional layer model (SCLM) is proposed that provides global low-light enhancement as the coarsely enhanced results. In addition, we introduce a local adaptation module that learns a set of shared parameters to accomplish local illumination correction to address the issue of varied exposure levels in different image regions. Experimental results demonstrate that the proposed method performs favorably against the state-of-the-art LLIE methods in both objective metrics and subjective visual effects. Additionally, our method has fewer parameters and lower inference complexity compared to other learning-based schemes. Code will be made publicly available at the URL https://gitee.com/zhanghahaxixi/SCLM. Yuantong Zhang, Baoxin Teng, Daiqin Yang, Zhenzhong Chen 0001, Haichuan Ma, Wenpeng Ding |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Towards Lightweight Deep Reference Frame for Versatile Video CodingabstractDeep neural network (DNN)-based methods have demonstrated enormous potential for Versatile Video Coding (VVC) inter prediction enhancement. However, due to their typically high computational complexity, implementing them in practical applications can be challenging. In this paper, we propose a lightweight deep reference frame interpolation network to enhance bi-prediction with low complexity. Specifically, given a pair of bi-directional reconstructed frames, first, we down-sample input frames to reduce the complexity before feeding them into the optical flow estimation network. Then the optical flows are utilized to warp extracted features at three different levels. The warped features are fused to generate the output intermediate frame. The additional reference frame is inserted into the reference picture lists to provide an additional reliable reference candidate. In contrast to previous efforts, the proposed method aims at achieving the trade-off between performance and complexity while maintaining a complexity of about 64 kMACs/pix. Experimental results demonstrate that our method achieves 1.82%/2.43%/2.02% coding efficiency improvements for Y/U/V components under random access (RA) configuration compared to the latest NNVC standard software VTM-11.0_NNVC-5.0. Wenhui Meng, Yuantong Zhang, Jianghao Jia, Songtao Chao, Zhenzhong Chen 0001 |
VCIP | 2 |
| 2023 | An Efficient Method for Real-Time Image Exposure CorrectionabstractExposure errors in images, including both underexposure and overexposure, significantly diminish images’ contrast and visual appeal. Existing deep learning-based exposure correction methods either require large networks or longer processing time for inference and are thus not applicable for embedded devices and real-time applications. To address these issues, a lightweight network is proposed in this paper to correct exposure errors with limited memory occupation and inference steps. It adopts the Laplacian pyramid to incrementally recover the color and details of the image through a layer-by-layer procedure. A structural re-parameterization structure is designed to both reduce model size for inference speed up and improve performance with a multi-branch learning structure. Extensive experiments demonstrate that our method achieves a better performance-efficiency trade-off than other exposure correction methods. Jie Yang 0002, Yuantong Zhang, Daiqin Yang, Zhenzhong Chen 0001 |
VCIP | 2 |
| 2023 | KPDFI: Efficient data flow integrity based on key property against data corruption attack
Xiaofan Nie, Haolai Wei, Yuantong Zhang, Ningning Cui |
Comput. Secur. | 4 |
| 2023 | Optical Flow Reusing for High-Efficiency Space-Time Video Super ResolutionabstractIn this paper, we consider the task of space-time video super-resolution (ST-VSR), which can increase the spatial resolution and frame rate for a given video simultaneously. Despite the remarkable progress of recent methods, most of them still suffer from high computational costs and inefficient long-range information usage. To alleviate these problems, we propose a Bidirectional Recurrence Network (BRN) with the optical-flow-reuse strategy to better use temporal knowledge from long-range neighboring frames for high-efficiency reconstruction. Specifically, an efficient and memory-saving multi-frame motion utilization strategy is proposed by reusing the intermediate flow of adjacent frames, which considerably reduces the computation burden of frame alignment compared with traditional LSTM-based designs. In addition, the proposed hidden state in BRN is updated by the reused optical flow and refined by the Feature Refinement Module (FRM) for further optimization. Moreover, by utilizing intermediate flow estimation, the proposed method can inference non-linear motion and restore details better. Extensive experiments demonstrate that our optical-flow-reuse-based bidirectional recurrent network (OFR-BRN) is superior to state-of-the-art methods in accuracy and efficiency. Codes are available on URL:https://github.com/hahazh/OFR-BRN Yuantong Zhang, Huairui Wang, Han Zhu 0003, Zhenzhong Chen 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Controllable Space-Time Video Super-Resolution via Enhanced Bidirectional Flow WarpingabstractSpace-time video super-resolution targets to increase a given video's frame rate and resolution simultaneously. Al-though existing approaches have made great progress, most of them still suffer from the inaccurate approximation of large motions or fail to generate temporal consistent motion trajectory. To alleviate these problems, we carefully review the characteris-tics of different optical flow warping strategies, integrating and enhancing them to achieve more robust capabilities for handling extreme motions and time-modulated interpolation. Specifically, we utilize enhanced backward warping to perform alignment, mine space-time information across low resolution input frames, and propose an enhanced forward warping strategy to interpolate arbitrary intermediate frames. Furthermore, the proposed model can be trained end-to-end and produce intermediate results at any time by merely supervising the center moment. Experimental results show that the proposed algorithm performs favorably against the state-of-the-art methods in objective metrics and subjective visual effects. Yuantong Zhang, Huairui Wang, Zhenzhong Chen 0001 |
VCIP | 1 |