Zheng Chang 0002

dblp:54/2780-2 · DBLP profile ↗
← Back
8ranked-venue papers
7as first author
8since 2021 · last 2025
0000-0002-8986-6841ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021
YearPublicationVenuePosition
2025 STAU: A SpatioTemporal-Aware Unit for Video Prediction and Beyond
abstract
Video prediction aims to predict future frames by modeling the complex spatiotemporal dynamics in videos. However, most existing methods only model the temporal information and the spatial information for videos in an independent manner but have not fully explored the correlations between both terms. In this paper, we propose a SpatioTemporal-Aware Unit (STAU) for video prediction and beyond by exploring the significant spatiotemporal correlations in videos. On the one hand, the motion-aware attention weights are learned from the spatial states to help aggregate the temporal states in the temporal domain. On the other hand, the appearance-aware attention weights are learned from the temporal states to help aggregate the spatial states in the spatial domain. In this way, the temporal information and the spatial information can be greatly aware of each other in both domains, during which, the spatiotemporal receptive field can also be greatly broadened for more reliable spatiotemporal modeling. Experiments are not only conducted on video prediction tasks (deterministic and stochastic), but also another task beyond video prediction, the early action recognition task. Experimental results show that the proposed STAU can achieve satisfactory performance on all tasks compared with other methods.
Zheng Chang 0002, Xinfeng Zhang 0001, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 Learned Image Compression Using Cross-Component Attention Mechanism
abstract
Learned image compression methods have achieved satisfactory results in recent years. However, existing methods are typically designed for RGB format, which are not suitable for YUV420 format due to the variance of different formats. In this paper, we propose an information-guided compression framework using cross-component attention mechanism, which can achieve efficient image compression in YUV420 format. Specifically, we design a dual-branch advanced information-preserving module (AIPM) based on the information-guided unit (IGU) and attention mechanism. On the one hand, the dual-branch architecture can prevent changes in original data distribution and avoid information disturbance between different components. The feature attention block (FAB) can preserve the important information. On the other hand, IGU can efficiently utilize the correlations between Y and UV components, which can further preserve the information of UV by the guidance of Y. Furthermore, we design an adaptive cross-channel enhancement module (ACEM) to reconstruct the details by utilizing the relations from different components, which makes use of the reconstructed Y as the textural and structural guidance for UV components. Extensive experiments show that the proposed framework can achieve the state-of-the-art performance in image compression for YUV420 format. More importantly, the proposed framework outperforms Versatile Video Coding (VVC) with 8.37% BD-rate reduction on common test conditions (CTC) sequences on average. In addition, we propose a quantization scheme for context model without model retraining, which can overcome the cross-platform decoding error caused by the floating-point operations in context model and provide a reference approach for the application of neural codec on different platforms.
Wenhong Duan, Zheng Chang 0002, Chuanmin Jia, Shanshe Wang, Siwei Ma 0001, Li Song 0001, Wen Gao 0001
IEEE Trans. Image Process.2
2023 STAM: A SpatioTemporal Attention Based Memory for Video Prediction
abstract
Video prediction has always been a very challenging problem in video representation learning due to the complexity in spatial structure and temporal variation. However, existing methods mainly predict videos by employing language-based memory structures from the traditional Long Short-Term Memories (LSTMs) or Gated Recurrent Units (GRUs), which may not be powerful enough to model the long-term dependencies in videos, consisting of much more complex spatiotemporal dynamics than sentences. In this paper, we propose a SpatioTemporal Attention based Memory (STAM), which can efficiently improve the long-term spatiotemporal memorizing capacity by incorporating the global spatiotemporal information in videos. In the temporal domain, the proposed STAM aims to observe temporal states from a wider temporal receptive field to capture accurate global motion information. In the spatial domain, the proposed STAM aims to jointly utilize both the high-level semantic spatial state and the low-level texture spatial states to model a more reliable global spatial representation for videos. In particular, the global spatiotemporal information is extracted with the help of an Efficient SpatioTemporal Attention Gate (ESTAG), which can adaptively apply different levels of attention scores to different spatiotemporal states according to their importance. Moreover, the proposed STAM are built with 3D convolutional layers due to their advantages in modeling spatiotemporal dynamics for videos. Experimental results show that the proposed STAM can achieve state-of-the-art performance on widely used datasets by leveraging the proposed spatiotemporal representations for videos.
Zheng Chang 0002, Xinfeng Zhang 0001, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001
IEEE Trans. Multim.1
2022 STRPM: A Spatiotemporal Residual Predictive Model for High-Resolution Video Prediction
abstract
Although many video prediction methods have obtained good performance in low-resolution (64∼128) videos, predictive models for high-resolution (512∼4K) videos have not been fully explored yet, which are more meaningful due to the increasing demand for high-quality videos. Compared with low-resolution videos, high-resolution videos contain richer appearance (spatial) information and more complex motion (temporal) information. In this paper, we propose a Spatiotemporal Residual Predictive Model (STRPM) for high-resolution video prediction. On the one hand, we propose a Spatiotemporal Encoding-Decoding Scheme to preserve more spatiotemporal information for high-resolution videos. In this way, the appearance details for each frame can be greatly preserved. On the other hand, we design a Residual Predictive Memory (RPM) which focuses on modeling the spatiotemporal residual features (STRF) between previous and future frames instead of the whole frame, which can greatly help capture the complex motion information in high-resolution videos. In addition, the proposed RPM can supervise the spatial encoder and temporal encoder to extract different features in the spatial domain and the temporal domain, respectively. Moreover, the proposed model is trained using generative adversarial networks (GANs) with a learned perceptual loss (LP-loss) to improve the perceptual quality of the predictions. Experimental results show that STRPM can generate more satisfactory results compared with various existing methods.
Zheng Chang 0002, Xinfeng Zhang 0001, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001
CVPR1
2021 Iprnn: An Information-Preserving Model For Video Prediction Using Spatiotemporal Grus
abstract
Videos are typically encoded to low-dimensional features to save computation resources for video prediction models. However, the unacceptable information loss while encoding is restricting the performance of the predictive models. To solve this problem, in this paper, we propose an Information-Preserving Spatiotemporal Predictive Model for video prediction, denoted as IPRNN. In our method, we apply multiple skip-connections between the corresponding layers between the encoders and decoders. In this way, more useful information from the encoders can be recalled by the decoders to achieve a more satisfactory performance. Moreover, to further save the computation resources for predictive models, we design a spatiotemporal gated recurrent unit (STGRU), which can efficiently capture the spatial appearance information and temporal motion information for videos. Experimental results show that the proposed method can obtain better performance compared with other state-of-the-art methods.
Zheng Chang 0002, Xinfeng Zhang 0001, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001
ICIP1
2021 STAE: A Spatiotemporal Auto-Encoder for High-Resolution Video Prediction
abstract
Predicting high-resolution videos (≥ 256) is always a very difficult task in video prediction domain. To predict high-quality frames for high-resolution videos, both the challenging spatiotemporal representations and the computation resources are needed to be carefully considered. In this paper, we propose a SpatioTemporal Auto-Encoder for High-Resolution Video Prediction, which is named STAE. In our method, we first jointly utilize the spatial and temporal encoders to extract low-dimensional spatial and temporal features from the high-resolution video input, which can preserve the spatiotemporal information from the input and significantly reduce the computation load for the following modules. In addition, we design a SpatioTemporal Attention based Memory (STAM) to predict the spatiotemporal features for future frames using the encoded low-dimensional features. Then the predicted spatial and temporal features are decoded back to the high-dimensional data space using the spatial and temporal decoders. Finally, the predicted high-dimensional spatial and temporal representations are jointly utilized to predict the future frames. All modules in STAE are built on the basis of 3D neural networks to improve the local perception to videos. Experimental results show the proposed method outperforms diverse state-of-the-arts on widely used datasets and the computation load is relative low.
Zheng Chang 0002, Xinfeng Zhang 0001, Shanshe Wang, Siwei Ma 0001, Yan Ye 0003, Wen Gao 0001
ICME1
2021 ASTM: An Attention based Spatiotemporal Model for Video Prediction Using 3D Convolutional Neural Networks
abstract
Video prediction has always been a challenging task in video representation learning due to the diversity of spatial-temporal evolution in videos. In this paper, we propose an Attention based SpatioTemporal Model for Video Prediction based on 3D Convolutional Neural Networks and Long Short-Term Memory (LSTM), which is named ASTM. In our method, we leverage both multi-term and short-term inter-frame dependencies in temporal domain to capture reliable motion information for videos. In particular, we design an Efficient Inter-Frame Attention Gate (EIFAG) to efficiently aggregate the multi-term inter-frame dependencies and integrate 3D convolutional operations into the proposed model to further improve the local perception to videos by capturing more accurate short-term temporal dependency. In addition, we make use of the multilayer Spatiotemporal LSTM (PredRNN) structure to preserve more spatial appearance details for videos. To evaluate the adaptability of our model on more complex real scenes, we collect a multi-level spatiotemporal (MLST) dataset. Experimental results show that the proposed model can achieve state-of-the-art performance on both widely used datasets and the proposed MLST dataset.
Zheng Chang 0002, Xinfeng Zhang 0001, Shanshe Wang, Siwei Ma 0001, Yan Ye 0003, Wen Gao 0001
ICME1
2021 MAU: A Motion-Aware Unit for Video Prediction and Beyond
abstract
Accurately predicting inter-frame motion information plays a key role in video prediction tasks. In this paper, we propose a Motion-Aware Unit (MAU) to capture reliable inter-frame motion information by broadening the temporal receptive field of the predictive units. The MAU consists of two modules, the attention module and the fusion module. The attention module aims to learn an attention map based on the correlations between the current spatial state and the historical spatial states. Based on the learned attention map, the historical temporal states are aggregated to an augmented motion information (AMI). In this way, the predictive unit can perceive more temporal dynamics from a wider receptive field. Then, the fusion module is utilized to further aggregate the augmented motion information (AMI) and current appearance information (current spatial state) to the final predicted frame. The computation load of MAU is relatively low and the proposed unit can be easily applied to other predictive models. Moreover, an information recalling scheme is employed into the encoders and decoders to help preserve the visual details of the predictions. We evaluate the MAU on both video prediction and early action recognition tasks. Experimental results show that the MAU outperforms the state-of-the-art methods on both tasks.
Zheng Chang 0002, Xinfeng Zhang 0001, Shanshe Wang, Siwei Ma 0001, Yan Ye 0003, Xiang Xinguang, Wen Gao 0001
NeurIPS1